Short answer. On 11 September the New Mexico Supreme Court fined a lawyer $5,000 over an appeal in which the witnesses and the police testimony had been invented by a model. He asked it to invent nothing: he uploaded the hearing transcript and asked for a summary. Everybody treats that task as the safe one, and it is riskier than writing from scratch, because the result looks checkable. What needs checking is not the text but whether it matches the source, and that is a separate operation.
A lawyer handed the model a document he already had. What came back was a document in which some of the people do not exist.
What actually happened?
A court fined a lawyer over a brief in which the model invented witnesses.
On 11 September 2026 the New Mexico Supreme Court held attorney Stephen Aarons in contempt and fined him $5,000, directed to the state's client protection fund. The case was an appeal against a murder conviction, and the client is serving life.
In the court's words, the filing "contained false testimony from wholly fabricated witnesses". Among the material were fictional statements that the shooter was wearing dark pants and a white shirt. The court added that Aarons had "demonstrated a lack of remorse and a lack of concern for his client".
At a hearing in August, Aarons described loading a computer-generated transcript of the trial and other case documents into ChatGPT. He expected, in his own phrase, "a bulletproof summary". He asked it to invent nothing. He asked it to compress what he already had.
The fine was not the end of it. The briefs were stricken, the court ordered the appeal to start over, and the matter went to the attorney disciplinary board.
Why is summarising riskier than generating?
The result looks checkable, so nobody checks it.
In a person's head these are two different jobs. "Write me something about X" feels like invention, and the output gets treated accordingly: read, edited, distrusted. "Summarise my document" feels mechanical, closer to copying. The source is right there. What is there to get wrong.
Invention inside a summary carries no warning signs. A fabricated paragraph usually shows itself through vague phrasing and empty confidence. A line about dark pants and a white shirt reads exactly like a detail lifted from a transcript. It is specific, it fits, and the only way to find out it was never there is to open the transcript.
| Write something | Summarise a document | |
|---|---|---|
| What the person asks for | invent this | compress what exists |
| What an error looks like | vague filler | a specific, plausible detail |
| Why it slips through | it does not, the text gets read closely | the source is right there, so checking feels unnecessary |
There is a difference of scale here too, and it matters. Earlier sanctions against US lawyers involved invented case citations and misquoted law. A citation takes a minute to check: open the database, it is there or it is not. Here the invention was factual testimony in a criminal case, and checking it means reading the whole transcript again. Which is the exact work the model was brought in to eliminate.
Length makes all of this progressively worse. The longer the document, the bigger the saving from summarising, and the smaller the chance anybody compares the result against the original.
Where does this land in an accounting practice?
In the same place as in court: a document somebody signs.
The everyday jobs look harmless. Reconcile a quarter's worth of bank statements. Pull the payment terms out of a forty-page contract. Check a ledger against a report. Assemble an explanatory note from source documents. Every one of those is a summary of your own documents, and every one walks into the trap above.
A signature moves the responsibility onto a person, and the tool changes nothing about that. The client, the bank and the tax authority have no interest in who entered the figure in the filing.
The cost of the error jumps rather than climbs. While the inaccuracy lives inside a draft, it takes ten minutes to fix. Once the document has gone out to a client, to a bank or into a filing, it turns into correspondence, a resubmission and an explanation of why the figure changed. One click separates those two states, which is why the check belongs before that click.
"We check, of course" without a written procedure means nobody checks. Either the check is a step with an owner and a mark that it happened, or it does not exist. There is no middle state, and which one you had always becomes clear afterwards.
We have no data of our own on how often this happens in accounting practice, and we are not inventing any. The mechanism is the same whatever the industry.
What goes between the model and the signature?
One procedure. It costs less than it sounds and takes a few minutes.
- Demand a pointer to the source. Not "summarise this" but "summarise this, and for every statement give the page or line it came from". A model with nothing to point at starts visibly floundering, and that is your signal.
- Check a sample rather than everything. Five points from the summary picked at random, plus every number heading into the signed document. If one of the five is not in the source, the whole summary goes in the bin.
- Split long documents. Nobody checks forty pages in one piece. Break them into sections, summarise one at a time, check one at a time.
- Record who checked. One line beside the document: who compared it, when, how many points. Six months later that line is the only thing separating a checked document from an unchecked one.
The first point does half the job on its own. Asking for a source turns the summary back into a checkable operation, and the errors surface without rereading the original.
Who connected the model to your working files in the first place is a separate question. We went through it in the piece on who has already connected AI to your data, along with the permissions granted to the model.
We build agents for processes where an error costs money, and we put the human check into the design rather than leaving it to somebody's discipline. How that works is on the AI team page. The first step happens without us though: take the last summary that was used in real work, and find five of its statements in the source.
Danil Ivanov
Founder, KAIVIX
Builds AI systems for companies in the UAE and beyond.