My personal bookkeeping system, open source under the MIT licence. Statements come in by email, are read into rows, matched against what is already booked and pre-filled by AI; I only confirm. People use the web app; a CLI and an API give AI agents the same capabilities. The goal in one line: half a day a month, and every balance matches the bank app.
- 2,124real rows replayed, identical to the old pipeline
- 5statement formats from two banks and two payment apps
- 70capabilities; 10 only a person can confirm
- 359tests passing on September 28
Why
Statements in China are scattered across payment apps and banks. The same money often shows up twice, once at the bank and once in the payment app; withdrawals and refunds end up with the wrong sign; and when a balance was off, the only way to find out why was to query the database.
There is no API to pull from, only email. One source sends a statement every day; the other four have to be requested in their apps each month, about 20 steps and 4 zip passwords in all.
One rule from the start: nothing sent to a model contains a card number, an order number, a password or a verification code.
The main loop
Mail or upload, then a reader, a uniform row, the channel semantics, matching, a proposal, the inbox, the ledger, and finally the assistant answering questions about it. Each step has one entry point, one stored state and one number.
The code is a pnpm monorepo of six packages: contracts, db, server, cli, web and admin. One server process runs the API, the job queue and the built-in assistant.
Decisions that shaped it
Readers report facts; a table decides meaning
The old parser decided signs itself, and one payment app's "not counted" rows were booked as income. Now a reader reports only what the statement says: the amount, in, out or neutral, and the raw type and status. The meaning (transfer, refund, not counted) comes from a channel semantics table of anchored regex rules in priority order, with a last rule per channel that falls back to the sign and sends the row to a person. A new type is never guessed. On the full real data no row fell through, and 24 rows of wrongly booked income were corrected. The rules can be edited in the admin, with a preview of which rows would change.
One reader per statement, replayed against samples
Changing one parser used to go through five layers of configuration. Now each kind of statement is one code reader that emits the same row shape, and each ships with a redacted sample and its expected output: change a reader, replay it, and not one row may differ. Shared tools handle password-protected zips, GBK detection, CSV, XLSX, PDF and email. A read-only replay of 2,124 real rows from 49 documents matched the old pipeline row for row, and reading one bank's PDF by column position fixed counterparty names that the old version had misplaced in about 27% of rows.
One number, stored
"Needs attention" used to be computed in four places, so one page could show 20 and 25 at once. Row state is now stored: eight states, with an event row for every transition (who, when, why, and whether it can be undone). The old status column became a generated column that cannot be written, and constraints require a booked row to have a transaction and a dismissed one to have a reason.
One capability catalogue, four surfaces
The CLI suggested commands that did not exist, dry runs passed where real writes failed, and the agent went around the API in five places. Now each capability is defined once in zod, and the HTTP route, the CLI command, the assistant's tool and the OpenAPI entry are all generated from it. A dry run executes the same function inside a real transaction and rolls it back, so "the dry run passed but the write failed" cannot happen by construction. There are 70 capabilities: 27 read, 33 write, and 10 that only a person can confirm.
AI and money
- Agents come in through the CLI rather than MCP, because a command line costs a model fewer tokens. There is no raw HTTP command; a model sees only the modelled capabilities.
- Capabilities that need confirmation are never exposed to a model or to an agent token. On the CLI they need a --confirmed-by-human flag, which an AI is not allowed to add.
- Before anything reaches a model, a redactor drops blacklisted fields (account, transaction and order numbers, raw columns, secrets) and masks any run of four or more digits. A known gap, written down: this covers the web assistant, and the terminal path still sees raw data.
- I test the CLI with a weaker model on purpose, so the help text and error messages have to be clear enough for it.
The rebuild
The first version was a Rust backend on a PHP fork of Firefly III: 33,000 lines of Rust and 137,000 of PHP. Rust's strengths do not matter much for a single-user, I/O-bound app, and its costs showed up every day. With no ORM one table grew to 43 columns, the AI SDKs live in TypeScript, and agents wrote Rust slowly.
None of the 42 problems was caused by Rust, but Rust made every one of them more expensive to fix.
I rebuilt it beside the old code: Node 22, Hono, Drizzle, pg-boss, zod and Vitest. In about four days the readers, the semantics, the matching and the replay baseline were running; then 127 Rust files and 1,506 Firefly files were deleted.
- Integration tests refuse any database whose name does not start with abei_test, and migrations refuse the development database.
- A replay compares three numbers with a baseline of 886 rows judged by hand (75.7% handled automatically, 199 interruptions, 1.8% missed) and fails on any drift.
- 28 architecture decision records, each with what was decided, the recommendation, the cost and whether it can be undone.
What it taught me
- Green tests do not mean it works. In the end-to-end acceptance, 4 of the 10 fixes were blockers in code whose unit tests were all green. A scripted walk marked 33 steps as passing; checked against the screenshots, 21 really did.
- It was built to platform scale for one person who uses it once a month: 43 tables and 70 capabilities where about 27 and 50 would do. The next step is to use it, not to extend it: close September's books on the new version, with every balance matching.
Timeline
- 2026.07The first Rust backend is archived
- 2026.08.07Firefly-based version with a React front end, CLI and agent
- 2026.08.31A month of real use: 43 problems written down
- 2026.09.01New architecture approved: its own single-entry ledger
- 2026.09.03Backend rebuilt in TypeScript beside the old code
- 2026.09.04Five readers done; 2,124 real rows replay identical
- 2026.09.17Open-sourced