On a platform where a wrong number costs the reader money, correctness is not a review process. It is a build step.
The platform helps service members plan their exit from the military. A large part of it is state-by-state guidance: which states tax military retirement pay, which exempt it, and what the partial exclusions actually cap out at. People make relocation decisions on this. A wrong figure is not a typo, it is bad financial advice at scale.
A reader flagged that one article claimed a state taxed military retirement when it fully exempts it. The narrow fix was obvious: correct that article. The useful question was whether it was one mistake or a pattern.
It was a pattern. An initial cross-check implicated roughly forty articles.
Before fixing anything, the underlying data had to be trustworthy. That meant deriving the national picture from primary sources — enacted statutes and state revenue department pages — rather than from search summaries, which routinely conflate a proposed bill with an enacted one and attach the wrong year to both.
That step earned its cost immediately. The project's own source-of-truth data file, labeled as fully verified, was itself stale for three states whose exclusion amounts had been raised by legislation. Two of those figures had already propagated into a live calculator. A file that says "verified" is a strong prior, not ground truth.
The first version of the check only read prose, so it under-reported badly — most of the wrong figures lived in comparison tables and formatted bullets, not sentences. Extending the parser to read the structures where claims actually live is what turned it from a rubber stamp into a real gate:
Rerun against the corpus, it surfaced 73 genuine contradictions across 28 files — every one a claim in the content that disagreed with the verified data. Each was corrected against the data file and re-checked, with a documented allowlist for the handful of legitimate false positives so the gate stays honest instead of being muted.
The last step is the one that matters. The check runs in prebuild, so a factual regression in this domain does not produce a warning someone triages later — it fails the production build. The content rule is documented in the repository so the constraint travels with the code.
A style guide asking authors to double-check tax figures would have failed the same way the first time. Humans reviewing hundreds of articles for one class of factual error is a process that degrades the moment anyone is busy.
Encoding the invariant machine-side changes what is possible: the wrong number cannot reach production, the check improves as the parser learns new places claims hide, and the cost of adding an article drops because correctness is enforced rather than remembered.
It also generalizes. Anywhere there is a fact the code depends on and a human keeps getting it wrong — a pricing table, a compliance threshold, a rate limit documented in two places — the same move applies. Write the gate, put it in the build.
Hiring, contract work, or just a question about how this one was built — email is the fastest path.
FormationLabs
AI Assistant
Quick questions: