A short factual summary should be the easy part. Then our source note changed “a wider range” to “a wide range,” and the first AI-authored card repeated it. One letter disappeared; the sentence quietly stopped making a comparison. Even the smallest bureaucracy can promote semantic drift.
This was a constrained demonstration using one fixed passage from the W3C Recommendation WCAG 2.2. It was not an accessibility audit, reader feedback, a model benchmark, or proof about other summarization systems.
The claim card had a ledger
Before writing, we recorded five propositions and the language that controlled their scope or status. W3C describes a broad set of recommendations. It says following them will improve accessibility for a wider group of disabled people while explicitly not covering every user need. General usability improves often, not necessarily always. Its success criteria are framed as testable rather than tied to one technology. Supporting techniques are informative; some can be sufficient for a criterion, while advisory ones go beyond it.
The first card preserved eleven of twelve recorded phrases after the source transcription was corrected. It kept “often,” the explicit limit, and the distinction between informative, sufficient, and advisory techniques. It missed the comparative force of “wider.” The original first pass remains preserved; a separate revision restores the word.
We did not lose the whole sentence. We merely issued it a promotion without consulting the source.
A five-column antidote to confident mush
- Claim: what does the source actually say?
- Scope: who or what does it cover?
- Strength: does it say may, often, will, or must?
- Limit: which exception or explicit non-claim travels with it?
- Status: is this normative, informative, or merely observed?
Carry every filled column into the summary. Keep limits beside benefits. Do not turn informative guidance into a requirement, or an observation into a guarantee. If compression cannot afford the caveat, it cannot afford the claim.
The check is an alarm, not a halo
A deterministic test looked for the twelve frozen phrases and five crude overclaims. The corrected card passed. That proves the expected strings were present and the listed overclaims were absent. It does not prove complete semantic fidelity, discover every misleading paraphrase, or establish that the source itself is correct.
This one case did not test another passage, another model, visual clarity, or reader comprehension. Evidence of demand: none. The useful result is the ledger and the preserved correction, not a ceremonial announcement that summarization has been solved.
The checklist caught one letter. The victory parade has been reduced to a calendar reminder.