A practice and assessment platform for insurance sales teams, planned at ZhongAn Information Technology. An LLM plays the customer, the trainee sells, the AI scores the conversation, and an instructor reviews the score. I did the research, wrote the product and technical plans, and built the backend, the admin console and the trainee app on my own, with AI coding agents doing much of the typing. It is an internal prototype and is being restructured.
- 221commits in 24 working days
- 143API operations, generated into TypeScript
- 333database tests, none skipped
- 180issues logged in my own acceptance walk
The problem
After a practice session an instructor could only give spoken feedback. Nothing was recorded, and each instructor judged differently.
Insurance adds its own rules:
- Exaggerating returns, making false promises or misleading a customer is an automatic fail, whatever else went well.
- Policy clauses and regulatory definitions have to trace back to the original text; a model may not publish them on its own.
- The simulated customer needs cards it keeps hidden, and asking the model to keep a secret is not a way to protect them.
There are three kinds of user: trainee, instructor and admin. The server's permission check is the real boundary; the front ends only hide what is not relevant.
Role-play
Hidden cards stay out of the prompt
Every card (persona secrets, objections, buying signals, compliance points) shares one trigger structure: keywords any, all or none, a regex, a minimum round, a probability, once or permanent, and a priority. After each turn the engine checks the conversation, rolls the probability and writes an append-only release event. Only released cards enter the customer's next context. The rest are never sent at all.
Context in layers
Instructions, the public persona, the scenario, the terms and the task goal come first, then a summary layer, then the dialogue. A three-to-five-line persona reminder rides on the last user message every turn. After 20 rounds the older turns are summarised and the last 8 kept; a hard budget of 24,000 characters always keeps at least 4.
Streaming that survives a dropped connection
Server-sent events with start, delta, end and error, and a heartbeat every 15 seconds. Only one reply is generated per attempt at a time, guarded in the database by an in-flight marker and an owner fence. Each message has an idempotency key, so a resend returns "duplicate" instead of a second reply. Closing the page does not cancel the reply; opening it again picks up the one in flight.
The customer, the grader, the summariser and the material assistant are separate AI roles, each with its own model, behind an OpenAI-compatible interface.
Scoring and review
Scoring
A worker claims grading jobs with a database lease, renews it at a third of the timeout and retries up to five times. Each conversation is scored three times in parallel: at least half the samples must be valid, each dimension takes the median level, and the spread is kept as a measure of disagreement.
Evidence has to be quoted word for word. A validator normalises whitespace, looks for each quote in the transcript and throws away any it cannot find. The report then rewrites the two or three turns that cost the most points: what was said, what was wrong, and a better line.
Review
A human review never edits a score. It adds a new score marked as human, with the dimension, the new level and a reason; the server converts levels to points against the frozen rubric and recomputes the weighted total, the veto and the pass result. Human conclusions are appended as new records, and the AI cannot modify an existing score. Only an admin can override a veto, once per attempt, with a written reason.
Expert calibration runs blind and in batches: experts see neither the model's score nor each other's, and disagreements go to a named arbiter.
A bug only a real walk found
Rubric weights and veto lines existed but never applied. The snapshot stored the rubric under one field name and the server read another, found nothing, and fell back to a plain average. Instructors reviewed against the rubric while the machine averaged. I fixed it on September 21.
The other surprise was upstream rate limits: in one count, 9 scoring runs succeeded and 29 failed. Failed scoring now retries on its own, and an instructor can score by hand, with the report saying so.
Content and exams
- Nine kinds of material (persona, clauses, objections, signals, compliance, rubric, case, paper, scenario) share one shell and one version table: draft, published, archived. There is one open draft at a time, published versions are read-only, and a draft still referenced by a task cannot be deleted. The same lifecycle is enforced by a database trigger and by a Go state machine, and a test checks that the two agree.
- Clause import: a material assistant may only fill fixed slots in a draft through fixed tools. A person confirms each slot, and only then can it go to review and release. Re-parsing a source file never changes a published clause.
- Tasks freeze their snapshot when they are assigned, so later edits never change work already done.
- Exams have five question types and are drawn deterministically from a blueprint and a seed. If the pool is short, the gaps are listed and have to be accepted explicitly; a paper never quietly shrinks.
Engineering
- Contracts: Go types become OpenAPI through huma, then TypeScript types: 115 paths and 143 operations. A check regenerates them and fails on any difference; hand-written front-end types are not allowed.
- Layers are enforced by a linter: cmd, app, biz, domain, and business code never touches the database driver. Front-end files are capped at 400 lines, and the admin app, the trainee app and the shared package cannot import each other the wrong way.
- make check runs the contract check, lint, migration lint, type checks and tests. Race tests and database integration tests run separately before a release. A test run that finds no test database fails instead of skipping; as the comment says, a green light that skipped the main flow is worse than a red one.
- One real refactor, on September 2 to 4: 42 migrations squashed into a baseline, sqlc replaced with plain pgx, every domain rebuilt in layers. One commit deleted 25,355 lines.
- 51 Go test files, 333 database tests with none skipped, and 11 Playwright suites on desktop and mobile, 24 of 24 passing on September 9.
Where it stands
On September 20 I stopped and looked at it as a product rather than as code. Nobody outside the build had opened it yet; every account and every attempt in the database was test data. Features kept growing, and no path had been walked from start to finish.
So I narrowed it. The instructor is the main user. The first job is one practice run that ends with a score someone can stand behind. Eleven entry points were closed for this version, exams among them.
Efficiency went up and features landed faster, and still nothing shipped.
The gates kept the code from getting worse. They could not tell me whether the product worked. A setting that exists but never takes effect only shows up when you walk the real flow.
Timeline
- 08.14Research and the overall plan
- 08.18Technical plan; repository created; a first pass of all three apps
- 08.21Streaming role-play with a real model; cards and triggers
- 09.01Domain schema; clause and standards libraries
- 09.04Full refactor: baseline migrations, layered domains, contract pipeline
- 09.10Acceptance walk: 180 issues logged, blockers fixed in batches
- 09.22Refocused on instructor scoring; weights and veto fixed; manual scoring