All writing

AI Only Writes Proposals: abei's Contract for AI

abei is an open-source bookkeeping tool I build in my spare time: statements arrive by email, they are parsed into one common format, AI pre-fills the categories, and a person confirms what goes into the books. The previous note was about deduplicating on import. This one is about what AI is and is not allowed to do inside abei.

Bookkeeping leaves little room for error. One wrong category skews the monthly numbers a little; a transfer mistaken for spending invents an expense out of nothing. AI can help a lot, but it cannot be allowed to change things whenever it likes. So before building any AI feature, I wrote a contract: what AI may touch, how it may touch it, and what happens when it gets something wrong.

First principle: AI only writes proposals

Every write AI makes goes through the proposals table. In the data layer, that means the fact tables and the proposals table are physically separate. AI cannot write to the ledger or the transactions at all. The only table it can write to is proposals.

A proposal becomes a fact only when a person approves it, or when it matches a rule a person has authorized in advance.

One point is easy to misread. abei has a pass-through lane, in which qualifying proposals are applied automatically. That does not mean AI decides to book them. It is a lane that a person has authorized in advance by choosing a trust level. Every automatic entry carries source = ai and the ai_run_id of the run that produced it, appears in the weekly report, and can be undone as a batch.

So the promise now is: for every entry, you can find out who decided it, and you can take it back.

Two producers, one proposals table, one lane policy

Proposals come from two places:

Both go into the same table and through the same lane policy, which sorts them into three lanes:

LaneWhenResult
Pass-throughCalibrated confidence ≥ 0.95 and a small amountApplied automatically, marked auto, can be undone
Batch reviewConfidence between 0.6 and 0.95Goes into a review queue, pre-checked; a person skims it and approves
Must askConfidence < 0.6A blocking question, with the two most likely answers attached

Proposals from external agents are calibrated separately, and at first they all go to batch review. None of them pass through.

The pass-through lane makes no judgment of its own. It is the exit policy of the built-in agent’s pre-fill pipeline.

The categorization engine: a four-layer waterfall

When the built-in agent pre-fills a category, each transaction falls through the layers one at a time:

LayerWhat it doesTarget coverage
L0 Merchant normalizationStrips channel prefixes, bracketed entity names and card number suffixes; rules and named-entity recognition run side by sideEverything
L1 Deterministic rulesCompiles my written “personal bookkeeping rules” into structured rulesAbout 90%
L2 Nearest neighbors in historyThe normalized merchant plus an amount bucket; the most similar past transactions voteMost of the rest
L3 LLMHandles only the remaining 5 to 15%, in batches; returns a category, its reasoning and the rule it relied onThe residue
L4 A personThe must-ask lane5% or less

The language model comes last. When the quota runs out, abei falls back to L1 and L2 and keeps working.

A model’s own confidence cannot be trusted

When an LLM says it is 95% sure, it is not right 95% of the time. So confidence goes through a calibration layer. Proposals are grouped by who made them and what kind they are. In each group, the 500 most recent proposals that a person has already confirmed are used to fit an isotonic regression, which maps the model’s self-reported score to the accuracy it has actually had.

When there are not enough samples yet (a cold start), the score simply takes a 20% discount: calibrated = raw × 0.8.

Some cases ignore confidence

However confident a proposal is, the following cases always go at least to batch review:

  1. The amount is at or above the large-amount line;
  2. The matching engine thinks it may be one half of a transfer;
  3. It may be a refund of another transaction;
  4. The merchant has appeared fewer than 3 times before;
  5. The rules and the historical neighbors disagree on the category.

These are fixed conditions. If one of them matches, the proposal is downgraded, with no discussion.

Trust levels: people decide how much passes through

On abei’s agent page, a person picks one of three trust levels:

LevelBehavior
Confirm everythingNo pass-through; every entry needs a person’s approval
Small amounts pass (default)Entries with high enough confidence and an amount below the small-amount line are applied automatically
Ask only about large onesEverything passes through except large amounts and the cases above

The small-amount and large-amount lines default to 200 and 1,000 yuan, and they are still being tuned.

Hard constraints on the proposals table

Each proposal points to exactly one target, a statement row or a transaction, and carries a patch that describes the change it wants. A few hard constraints apply:

Shadow mode first, then more trust

New automation does not go live directly. It first runs in shadow mode: proposals are written as usual and marked shadow_applied, but never applied. The weekly report compares each shadow decision with what the person actually decided.

It graduates on a single number, the leaked error rate: the share of automatically applied entries that a person later changed back. It has to stay below 2%, over at least 200 samples.

Correct it once, and it learns

When a person changes a category, abei asks, “Always do this?” A yes creates a rule, which waits in a review queue. Applying it to older entries starts with a dry run that lists the affected entries for the person to tick.

All the rules live in one document, my “personal bookkeeping rules”. It stays under 200 lines, every rule comes from a real correction, and every rule is written as an instruction that can be checked. If an agent wants to edit that document, it has to write a proposal too.

Some merchants sell everything, a supermarket for example, so there is no category to learn. Those merchants are marked as not learnable, and AI stops making proposals for them.

Audit and privacy

Last words

The contract is still a draft, and several thresholds will be tuned once shadow mode produces real data. The first principle will not change: AI only writes proposals, people decide which proposals may pass on their own, and every entry can be taken back.

“The model determines what AI can do. The platform determines whether people will trust it enough to act.” I wrote that in another essay. abei is my attempt to put it into practice myself.

All writing