← Essays

What We Got Wrong Building With AI

epistemic status: confessional. Receipts included.

Victory laps teach nothing. Here are three bets we got wrong at Anna Technologies, what each one cost, and what we changed. I’m writing these down partly for you and mostly so we don’t repeat them.

Mistake one: we automated the wrong end of the workflow

We built AI for the part of our users’ work that was most impressive to demo, generation, the blank-page moment. Users were politely enthusiastic and quietly unmoved. The real pain was downstream: reviewing, reconciling, chasing status. Boring work, invisible in a demo, enormous in a week. When we moved the AI there, retention moved with it. Lesson: demo-ability and value are not just different, they’re often inversely correlated, because everyone builds where the demos are.

Mistake two: we trusted evals that flattered us

Our internal benchmarks said the feature was ready. The benchmarks were built from the cases we’d thought of, which is to say, the cases we’d already handled. Real users supplied the other ninety percent. We shipped, we apologised, we built a habit we keep now: every eval set must contain failures harvested from production, refreshed monthly, and the person who builds a feature may not be the only one who writes its tests. Cost: one rough launch and a quarter of rebuilt trust.

Mistake three: we hid the seams

We thought polish meant making the AI look infallible, no confidence signals, no “here’s what I’m unsure about.” Users punished us for it. The moment the product was wrong once with full confidence, they stopped trusting it everywhere. When we redesigned to show uncertainty honestly, flag the shaky rows, link the sources, make checking easy, usage went up, not down. People don’t need the machine to be right; they need to know when it might not be. Honesty is a feature.

The common thread: every mistake came from optimising for how the product looked to us instead of how it behaved in someone else’s Tuesday afternoon. The fix, each time, was contact with reality, earlier and on purpose.

the best case against this essay

The case against

Public post-mortems select for the mistakes that make good essays. The three above all resolve into flattering lessons; the genuinely embarrassing errors, strategic, interpersonal, slow, don’t appear. Transparency that’s curated this well is a marketing genre, and readers should discount it accordingly.