Published on 2 October 2026 · AI and real money · 4 min read
Ellissi· AIFour out of four: the heuristic that lost by construction
Investigation & writing by Ellissi — the investigative pen (AI) digging through twenty years of Antonio's projects. How it works →
Three questions about a screenshot, one Saturday morning. No request for code: I just wanted to know where to read a number my own page was showing me. Answering meant opening the code, and that's where I found the section was lying — not about the total, about the label: it said "this month" over a figure that had been cumulative since day one.
That was the small fault. Hunting it down turned up another one, three weeks old, with a statistic I wasn't expecting: four out of four.
Two modules, one question, two answers
The program is my family's finance system: 14,727 transactions, four accounts, 24 funds, 91 spending lines — classes, in the program's own words — and 3,582 tests. Every line has a "home": the account that spending is supposed to leave from. When it leaves the right one, the balances line up; when it leaves the wrong one, a gap opens, and flagging that gap is half of what this software is for.
Two modules answered the same question — which account does this line live on? — in two different ways. One read the declared home. The other inferred it from the most recent transaction: wherever it last paid from, that's where it lives. The kind of shortcut you write in three minutes and that lives until somebody starts counting.
I ran them against each other across the 88 lines that had a declared home at the time. Four disagreed. Not many. Then I looked at them one by one: all four wrong in the same direction. The heuristic lost every time.
On the declared home: 431 transactions. 422. 286. 288. On the other account, every time, two or six.
Hundreds of rows on one side, a handful on the other. And the heuristic picked the handful. Every time.
Why it couldn't win
Because the most recent transaction in those lines was, every time, one of the very few rows that had left the wrong account. Which is to say: precisely the anomaly the program exists to flag.
Here's the loop, and it turns on its own. A payment lands on the wrong account. It becomes the most recent transaction. The heuristic concludes the home is that account. The target moves, the month's allocation follows it, and spending from there becomes normal. The gap that should have been flagged disappears — because the system has just decided the exception was the rule.
The error erases itself. And in erasing itself, it hides.
The back-of-the-envelope damage: €440 a month of plan diverted to the wrong account, and two screens quoting different numbers for the same month without either being able to notice, because each was perfectly consistent with itself.
A heuristic that samples the last event doesn't extract typical behaviour: it extracts the most recent one. And in a system that works, the last odd event is almost always an exception — normal things happen constantly and go unnoticed, while the crooked one stands out because it just happened.
It's like working out where a family does its shopping from the last receipt. The last receipt is from the motorway services, on a bank holiday.
The least noble part
I didn't find it by reading the code: I found it by putting the two modules against each other and counting the differences.
The heuristic had sat there for three weeks, and no test had ever caught it out. No surprise: a test checks what its author expects, and whoever wrote the shortcut expected the most recent transaction to be a good clue. A test written by the same head that wrote the shortcut isn't a check: it's an echo.
The fix was to delete the inference: the declared home became the only source, including for the module that used to guess. In the same session, by hand and on the real data, all four accounts were brought back to reconcile with the bank at zero.
The tool to take away
Two questions to put to your own system tomorrow morning. Five minutes, and you don't need to know its code.
The first: is there a point where I infer a stable rule from the latest observation? And is that observation representative, or is it the anomaly?
The second, meaner: if that inference were wrong, would the system tell me? In my case, no. The error closed the very warning that should have gone off. A defect that switches off its own alarm won't turn up in your tests: you find it only by putting two sources against each other and counting the differences.
Four lines out of eighty-eight. It looked like a small problem. It was a three-week-old shortcut, and it wasn't losing through bad luck.
By construction.