abei is an open-source bookkeeping tool I build in my spare time: statements arrive by email, they are parsed into one common format, AI pre-fills the categories, and a person confirms what goes into the books. The previous note was about deduplicating on import. This one is about what AI is and is not allowed to do inside abei.
Bookkeeping leaves little room for error. One wrong category skews the monthly numbers a little; a transfer mistaken for spending invents an expense out of nothing. AI can help a lot, but it cannot be allowed to change things whenever it likes. So before building any AI feature, I wrote a contract: what AI may touch, how it may touch it, and what happens when it gets something wrong.
First principle: AI only writes proposals
Every write AI makes goes through the proposals table. In the data layer, that means the fact tables and the proposals table are physically separate. AI cannot write to the ledger or the transactions at all. The only table it can write to is proposals.
A proposal becomes a fact only when a person approves it, or when it matches a rule a person has authorized in advance.
One point is easy to misread. abei has a pass-through lane, in which qualifying proposals are applied automatically. That does not mean AI decides to book them. It is a lane that a person has authorized in advance by choosing a trust level. Every automatic entry carries source = ai and the ai_run_id of the run that produced it, appears in the weekly report, and can be undone as a batch.
So the promise now is: for every entry, you can find out who decided it, and you can take it back.
Two producers, one proposals table, one lane policy
Proposals come from two places:
- The built-in agent, which lives inside abei and runs a pre-fill pipeline;
- External agents such as Claude Code, which submit through the command line with
abei proposals create --from-file.
Both go into the same table and through the same lane policy, which sorts them into three lanes:
| Lane | When | Result |
|---|---|---|
| Pass-through | Calibrated confidence ≥ 0.95 and a small amount | Applied automatically, marked auto, can be undone |
| Batch review | Confidence between 0.6 and 0.95 | Goes into a review queue, pre-checked; a person skims it and approves |
| Must ask | Confidence < 0.6 | A blocking question, with the two most likely answers attached |
Proposals from external agents are calibrated separately, and at first they all go to batch review. None of them pass through.
The pass-through lane makes no judgment of its own. It is the exit policy of the built-in agent’s pre-fill pipeline.
The categorization engine: a four-layer waterfall
When the built-in agent pre-fills a category, each transaction falls through the layers one at a time:
| Layer | What it does | Target coverage |
|---|---|---|
| L0 Merchant normalization | Strips channel prefixes, bracketed entity names and card number suffixes; rules and named-entity recognition run side by side | Everything |
| L1 Deterministic rules | Compiles my written “personal bookkeeping rules” into structured rules | About 90% |
| L2 Nearest neighbors in history | The normalized merchant plus an amount bucket; the most similar past transactions vote | Most of the rest |
| L3 LLM | Handles only the remaining 5 to 15%, in batches; returns a category, its reasoning and the rule it relied on | The residue |
| L4 A person | The must-ask lane | 5% or less |
The language model comes last. When the quota runs out, abei falls back to L1 and L2 and keeps working.
A model’s own confidence cannot be trusted
When an LLM says it is 95% sure, it is not right 95% of the time. So confidence goes through a calibration layer. Proposals are grouped by who made them and what kind they are. In each group, the 500 most recent proposals that a person has already confirmed are used to fit an isotonic regression, which maps the model’s self-reported score to the accuracy it has actually had.
When there are not enough samples yet (a cold start), the score simply takes a 20% discount: calibrated = raw × 0.8.
Some cases ignore confidence
However confident a proposal is, the following cases always go at least to batch review:
- The amount is at or above the large-amount line;
- The matching engine thinks it may be one half of a transfer;
- It may be a refund of another transaction;
- The merchant has appeared fewer than 3 times before;
- The rules and the historical neighbors disagree on the category.
These are fixed conditions. If one of them matches, the proposal is downgraded, with no discussion.
Trust levels: people decide how much passes through
On abei’s agent page, a person picks one of three trust levels:
| Level | Behavior |
|---|---|
| Confirm everything | No pass-through; every entry needs a person’s approval |
| Small amounts pass (default) | Entries with high enough confidence and an amount below the small-amount line are applied automatically |
| Ask only about large ones | Everything passes through except large amounts and the cases above |
The small-amount and large-amount lines default to 200 and 1,000 yuan, and they are still being tuned.
Hard constraints on the proposals table
Each proposal points to exactly one target, a statement row or a transaction, and carries a patch that describes the change it wants. A few hard constraints apply:
- A
patchmay not contain an amount field; the database schema rejects it. The only exception is a split, where the parts must add up to the original amount, and a database trigger enforces that. - Numbers come from deterministic queries; AI only chooses the words. When the weekly report says how much a category went up compared with last month, the figure comes from SQL, not from the model.
- A dry run and a real run are the same code. With
apply = false, it only returns the diff. - The same proposal cannot pile up: a hash of its content is unique among proposals that have not been rejected. A rejected proposal can be made again, because people change their minds.
Shadow mode first, then more trust
New automation does not go live directly. It first runs in shadow mode: proposals are written as usual and marked shadow_applied, but never applied. The weekly report compares each shadow decision with what the person actually decided.
It graduates on a single number, the leaked error rate: the share of automatically applied entries that a person later changed back. It has to stay below 2%, over at least 200 samples.
Correct it once, and it learns
When a person changes a category, abei asks, “Always do this?” A yes creates a rule, which waits in a review queue. Applying it to older entries starts with a dry run that lists the affected entries for the person to tick.
All the rules live in one document, my “personal bookkeeping rules”. It stays under 200 lines, every rule comes from a real correction, and every rule is written as an instruction that can be checked. If an agent wants to edit that document, it has to write a proposal too.
Some merchants sell everything, a supermarket for example, so there is no category to learn. Those merchants are marked as not learnable, and AI stops making proposals for them.
Audit and privacy
- Every agent action goes into an append-only audit log: run id, model version, prompt hash, input summary, output and approver.
- Every transaction keeps its revision history, which answers “who decided this?”
- The model never sees card numbers, order numbers or passwords. Fields are checked against an allowlist before anything is written back, and parameters marked
x-abei-human-onlyare never passed to the model.
Last words
The contract is still a draft, and several thresholds will be tuned once shadow mode produces real data. The first principle will not change: AI only writes proposals, people decide which proposals may pass on their own, and every entry can be taken back.
“The model determines what AI can do. The platform determines whether people will trust it enough to act.” I wrote that in another essay. abei is my attempt to put it into practice myself.