Adds --write to the matcher. Wrote 1,242 split rows across 657 transactions.
Imported settled, and that is the whole design. These obligations were
discharged years ago on a platform we no longer run, and their residual is
already carried by transaction 2348. Writing them unsettled would re-open
roughly $40k of debts that were paid. ACTIVE_OBLIGATION keeps settled splits
out of every owed figure while myShare/mySplitOf still count them, which is
exactly the asymmetry this needs: the import exists to correct historical
SPEND, not to move a balance.
Effect: $35,259 leaves my historical spend -- $13,088 in 2024, $22,117 in
2025 -- because a $200 grocery shop that was always half hers no longer reads
as $200 of mine. Balances are byte-identical before and after (Molina
-1226.72/145, Sonu 5428.08/419), which is the assertion that matters.
Shares are written as the CSV computed them, so a 50/50 row can land as
50.01/49.99. That is faithful rather than tidy; no transaction exceeds 100%.
Rehearsed on the 37-row Rome file first (24 rows) and verified before the full
run -- both the balances and one split read back through the API.
Matches the five CSV exports against transactions already in the ledger and
reports what it would do. Writes nothing -- importing is a separate step, and
rehearsing it first is what catches the defects that tests do not.
What it found, and why the number is what it is: 676 of 1,536 shareable rows
match (44%). The ceiling is ledger coverage, not matcher quality. The CSVs
describe 678 shared expenses in 2024 alone; the ledger holds 591 rows for the
whole of that year, 3 to 72 a month, which is far less than a household
actually spends. Most 2024 CSV rows have no transaction to attach a split to
and never will. South Korea April 2024 matches 4 of 158 for that reason.
Three decisions are encoded deliberately:
- Date format is decided per FILE, not per row. The household export writes
D/M/YYYY and the four trip exports write ISO, and 474 rows parse validly
under both readings -- per-row guessing silently swaps January and February
for some rows and not others.
- A person's column is their net balance impact, not their share. The payer
is whoever is positive; the other's share is |their negative| / cost. So a
+cost/-cost row means the other party owes 100%, not that the expense was
unshared -- the reading that would fake an arrangement change.
- Matching is one-to-one, best pair first. The NZ trip has two identical
$10.16 Uber rows against three ledger rows and four PayMyPark rows in the
same shape; without this a ledger row is claimed repeatedly and the second
CSV row looks matched while being unrepresented.