Synthetic examples: aggregate.csv
LOCAL TOOLS
Deduplicate rows with auditable aggregation
Group repeated records without silently losing their differences.
When repeated keys need an explanation
Repeated rows may carry different descriptions or quantities. Choose the columns that identify a group, then state how each output value should be selected or combined. This is grouping, not fuzzy matching or automatic data correction.
How it works
Select one local XLSX or UTF-8 CSV file, inspect the sheet and confirm its header. Select one key or a composite key, optionally enable explicit key normalization, add aggregation rules, validate, run and review the bounded result. EXACT is the default; numeric key normalization can remove leading zeros only when opted in, and ambiguous date keys require MDY/DMY. Normalization changes membership, not raw values or distinct-list equality. Duplicate and blank headers remain selectable by original column position. Groups follow first occurrence; first and last refer to original row order.
Aggregation semantics
First and last retain blanks. Count rows includes every member; count nonblank excludes null and empty strings but counts whitespace. Distinct list retains first-occurrence order with exact typed equality and keeps null/empty distinct; its output is a JSON list. Delimiter and newline joins preserve empty segments. Sum/min/max accept finite numeric cells only: CSV text "12" is not silently converted. Numeric sums use IEEE-754 and are not a financial decimal ledger.
Synthetic example
The downloadable fixture contains key 00123 twice with Alpha and Beta, then 00999 with Gamma. Select ID as key and Value as count_rows: results are 00123 → 2 and 00999 → 1. First returns Alpha, last returns Beta, and distinct_list returns ["Alpha","Beta"]. Duplicate_Groups records both source rows. These are invented labels.
Privacy and export
All grouping runs in a local Worker with no upload, telemetry or persistent file storage. Clear/Cancel terminates the Worker. A fresh XLSX includes Data, Summary, Duplicate_Groups, Excluded and Run_Settings. Formula expressions and date candidates remain audit text; no source formula, macro or external link is inherited. CSV is Data only and cannot retain audit sheets or all types.
Limits and review obligations
XLSX and UTF-8 CSV only; XLS/XLSM are rejected. Default input is approximately 10k × 30; 100k is experimental advanced mode. There is no fixed original-file size cap; XLSX expanded-content, 256-column and 32,767-character safety gates remain. Output including audit rows is limited to 650k default / 6.5m advanced cells. Large distinct lists or joins can block before export. No fuzzy grouping, numeric-text coercion or unlimited-output promise.
Frequently asked questions
Are 00123 and 123 the same? No; exact typed grouping preserves text IDs and long IDs. What happens to an empty key? Default block, or explicitly treat as a value or exclude with audit. Does aggregation repair ambiguous dates? No: raw values remain unchanged. Can I sum a CSV amount? CSV is text; this version deliberately requires numeric XLSX cells rather than guessing conversion.