How to Spot Accounting Fraud with AI, from the Filings Alone
Five blind AI agents, real SEC filings. One fraud they recognized, one they reasoned out cold, one mistake they caught was mine.
$155,579,371 in cash.
$292,644 of interest income.
So both numbers from above sit in the same annual report of the same NASDAQ-listed company. The same got audited 10K, filed with the SEC, and then signed by management. Nothing in the filing says anything wrong
But park $155 million in plain time deposits and it earns about three million dollars a year. This company reported less than $300,000. That is a yield of 0.19 percent. Either management chose to leave nine figures earning almost nothing, or the cash was never there to earn anything.

So I ran a test. I stripped the company’s name and country out of the filing data and handed it to four AI analyst agents running blind. Four separate sessions, no shared memory, no hint that anything was wrong. (I ran them as Claude Code subagents. Claude Cowork, the desktop app that runs the same engine without a terminal, can split the job into parallel agents the same way. The tool matters less than the separation.) Each one got the same neutral question: what is the most material piece of information in these accounts?
All four came back with the same verdict. The cash is probably not real.
None of them could name the company.
And here is what happened in the real world, a year after that filing. The company’s audit firm tried to verify those bank balances. The company pointed the audit team to a suspected fake website for the bank. A bank advice supporting the interest income contained mathematical errors that management waved off as the bank’s clerical mistakes. The auditors resigned; I read the resignation letter which is on EDGAR, more on that later. The company lost its NASDAQ listing, and the SEC later revoked its registration.
Hold the name for now. It matters later that you have probably never heard of it.
Here is the problem, and it is probably yours too. Finding red flags was never the hard part. Any screener will hand you red flags all day: debt too high, margins too thin, receivables aging. But the market prices those within minutes of a filing hitting the wire, so a list of visible flags is worth almost nothing. The flags that destroy capital are the buried kind, the ones that only show up when two numbers from two different pages are forced to agree with each other and even when you actually found one, the question which nobody is going to teach you to answer is going to decide everything that is going to be the answer: is this flag survivable, or is it fatal?
A few weeks ago, a paid subscriber sent me an article about exactly this, with four questions attached. The article was by Alistair Smallwood, who writes Al Smallwood. The questions were sharper than most professional research briefs I have received:
What do you think of this research method?
Would AI only find fraud after the fact, or can it uncover it before the news breaks?
Can you remember an accounting scandal from your analyst days, and show how AI would have found it right away?
Is this workflow worthy of the newsletter?
I told him that I would get back to him in a few days after thorough research. What follows is what I found.
I built this workflow and then ran it on three resolved cases. Out of that, one case was recognised by the model from the memory, second one they reasoned out cold, and in the last one, the workflow caught my own data errors before it actually caught the fraud.
The second question is the one everyone actually cares about, so here is the short answer. AI will not predict the next scandal, but if you run it correctly, it does something more useful. It will tell you how to differentiate a survivable problem from a fatal one.
New here? Alpha with AI is not a stock ideas publication. Each edition builds a complete research system around one specific research job, insider activity, post-earnings research, working through filings, and shows how to do that job with AI end to end, with the exact prompts included. You take the system and run it on your own names.
Paid subscribers also get the live model portfolio, the complete AI research behind every position, not just the position itself. The annual plan is 35% off, locked in forever Limited time.
There are two kinds of red flag: survivable and fatal.
The article my subscriber sent was built around a small UK online-dating company called Cupid plc. On paper, 2011 was a great year: revenue up 109 percent to £53.6 million, profit doubled, first dividend. The material fact was buried in a payroll footnote: 177 of the company’s 384 employees were classed as marketing staff, at a company whose advertising was almost entirely online. You do not need 177 people to buy online ads in a company of this size. In early 2013 the BBC, and a newspaper in Ukraine where most of the staff actually sat, reported what they appeared to be doing: posing as interested members to draw real users into paying to reply. The company denied any organised practice, and a review it commissioned from KPMG did not confirm one. I still think the footnote was telling the truth. The market did not wait for a verdict. On the allegations the shares lost more than half their value in a single day, and although they bounced when the company pushed back, the dating business never recovered. Cupid sold off the sites at the centre of it and was out of online dating within two years.
Now, notice what kind of flag this is. It did not say a particular number needed checking. Instead, it questioned whether the product itself was real. That split is the one thing no screener will ever be able to hand you. Every red flag belongs to one of the two species.
The survivable species says: a number needs to improve, or be verified. High debt. Weak cash conversion. Ugly margins. A one-time liability nobody can size yet. If the business underneath is real and the management honest, these flags resolve, and the pessimism they created is where the money gets made.
The first example I can think of is from 1963, when a subsidiary of American Express had issued some warehouse receipts for vegetable oil that did not actually exist. The tanks that supposedly held oil were actually holding seawater with a film of oil floating on the top. The actual claimed oil at that one facility exceeded the entire soybean and cottonseed oil supply of the entire country. When this got out, stock fell from $60 a share to about $35, but the actual receipts were only a one-time liability sitting on a balance sheet of an honest(debatable) functioning traveller’s check and charge card business. That damage had a lower floor. Warren Buffett put about 40% of his partnership into the stock in 1964 and roughly doubled his money when the actual claims were settled.
Now the same lesson from another market. In 2002, Titan, an Indian company, was worth about forty-five million dollars and carried debt of roughly three times its equity, dragged down by a weak watch business and a failed expansion into Europe. Profit fell by half in a single year. If you look at the numbers, it looked as if the company is already finished, but the debt was a bounded problem. It was not a lie. Underneath that debt was an honest, capable management and a jewellery brand Tanishq, which was already a top five retailer in the country and was growing fast. That is the survivable flag: a stretched balance sheet next to a real, growing business. Rakesh Jhunjhunwala met the incoming managing director, came away convinced the man was blunt and honest, and bought. Titan’s market value went from about forty-five million dollars to three and a half billion in a decade, and by the time Jhunjhunwala died in 2022 his stake was worth more than a billion dollars, roughly a thousand times what he had paid. Both stories are history, not recommendations. But they define the species: a bounded problem, honest people, a real business underneath, and the flag healed.
The fatal species says: the revenue is manufactured, the product is fake, management is lying. Cupid and the three names coming later in this piece. These flags do not resolve. They detonate. The uncomfortable part is that the fatal species usually wears the better-looking numbers, because the clean numbers are the lie.
Which is why the deciding variable is not in the ratios at all. Honest management with ugly numbers is usually survivable. Conflicted, unchecked management with suspiciously beautiful numbers is usually fatal. The honesty of the people running the company is the thing that decides it. That one judgment sorts every other red flag into survivable or fatal.
One honest caution before we go further. This workflow filters out the zeros. It does not find the Titans. Keep that in your head through everything that follows.
The hard part is choosing which flag is fatal.
So how do you build something that makes that call, survivable or fatal, on a filing full of flags? The problem is not the one you would expect.
Point a capable model at a set of accounts and ask it the open question, what is the most material fact here, and it finds plenty. It hands you weak cash conversion, a related-party balance, aged receivables. All real. All survivable. It rarely lands on the one buried thing that means the revenue is fake, even when that thing sits in the same document. The experiment that sparked this piece showed the pattern cleanly: split the reading across specialist agents and they find the fatal fact almost every time, but the manager that ranks their findings throws it away almost every time, because it sides with whatever the most agents agree on, and the fatal finding only ever shows up in one report.
Finding flags is easy. Choosing the fatal one out of a pile of survivable ones is the whole job, and it is the exact thing a human analyst is paid for.
So I built the workflow to protect that one minority finding at every step, instead of letting a vote bury it. Here is the whole thing, the way I ran it.
Feed it the filings, and nothing else. The annual report, the last few quarters, the prospectus if there is one. Not news, not summaries, not the company’s own investor deck. (If you have never wired filings into Claude Code, I walked through that setup in How I Set Up Claude Code as My Investment Research Analyst Start there.)
Split the work across five separate agents. Be literal about this, because it is the part that makes the whole thing work. You do not ask one model to do five jobs. You run five separate agents, each one a clean instance with a single task, and none of them able to see what the others found. In Claude Code these are five separate subagents, each in its own context, none of them able to see the others. Claude Cowork can split the job into parallel agents for you the same way. One agent reads whether the sales are real. One checks whether the per-customer and per-store numbers make sense. One asks whether the reported profit is backed by cash actually coming in. One hunts for two disclosed numbers that cannot both be true. And one looks at who controls the company and whether they can be trusted. Hand all five jobs to a single model and you just get its one worldview, five times over. Five points of view have to be built in on purpose. You cannot get them by asking one model to be broad.
Make each agent answer the same four questions. What is the single most important thing you found? What is the math behind it? Is it survivable or fatal, and why? And what one piece of outside information would prove it true or false? Each agent comes back with one finding, not a list, because a list is where you will see the fatal one hiding among all the safe ones.
Send each finding to a fresh agent to argue in full. This is the step that was missing everywhere else, and the one that makes the whole thing work. A new agent takes one finding on its own and builds the strongest case for it: assume this single thing is the key to the entire company, argue it all the way out, then give the ordinary explanation too. The point is to see each finding at full strength before it is ranked against any other.
Let one final agent judge by impact, not by how many agents agreed. Only now does a last agent, the judge, read the finished arguments and rank them by a single question: if this finding is true, how much does it actually hurt the company? Go back to Cupid. One finding sounds like a footnote: 177 of the staff are in marketing, at a company that advertises almost entirely online. The other sounds like a scandal: 2.3 million pounds of customer money is parked in a separate company the directors own. But the scandal only means one balance needs checking, while the footnote questions whether the revenue itself is real. Judged by what each would really do to the company, the footnote wins.
Every number reconciles to the filing, or it does not get used. This one sounds too obvious to state. It is the step that saved me from myself, and I will show you exactly how in the Luckin run.
Here is the instruction you give Claude, and it pays to be literal about it. The prompt names all five jobs and tells Claude to run each one as a separate agent, blind to the others. That is how Claude knows to use five agents, one per job, and not one agent doing all five, or two, or eight. You do not paste it five times. The only rule that cannot bend is that they stay separate. The moment all five share a single chat, they collapse back into one worldview, which is the exact failure this design exists to prevent.
Read only the filings below. Reason only from these numbers.
[the company's filings]
First, run five separate agents, each a forensic specialist, blind to
each other, one for each job:
1. Are the sales real?
2. Do the per-customer and per-store numbers make sense?
3. Is the reported profit backed by cash coming in?
4. Do any two disclosed numbers contradict each other?
5. Can the people who control the company be trusted?
Each of the five returns exactly:
- THE SINGLE MOST IMPORTANT THING YOU FOUND
- THE MATH BEHIND IT
- SURVIVABLE OR FATAL, AND WHY
- THE ONE OUTSIDE FACT THAT WOULD PROVE IT TRUE OR FALSE
Then take each of the five findings and argue it at full strength.
Finally, after all five are done, one more agent, the judge, reads the
arguments and ranks them by impact: if a finding is true, how much
would it hurt the company.That is the whole machine.
Two notes on using it, because they answer the obvious question. You run this on a company you are actually thinking about owning, feeding the agents its real, current filings. That is the point of building it. Everything from here on, though, is me proving it works on three companies whose stories already ended, because a finished fraud is the only place you can check the method against an answer you already know. And on those three I added one step you would never bother with on a live company. I stripped the name and the country out of the numbers first, so I could see whether the agents were working the problem out or just remembering a famous scandal. On a company nobody has exposed yet, there is nothing to remember. You just hand it the filings.
The very first company answered that memory question before I finished asking it.
WorldCom was the perfect first test, until the models recognized it.
Start with the cleanest fraud in the record, because you can see the trick the moment someone points at it. Through 2001 and into 2002, WorldCom took an ordinary running cost, the “line costs” it paid other phone carriers to carry its calls, and recorded billions of it as though the money had been spent building the network instead.
Here is why that one move matters so much, in plain terms. A normal running cost comes straight off this year’s profit, all at once. But money spent building something long-lived, a piece of network, does not; you spread that cost out over many years. So by relabeling this year’s phone bills as long-lived assets, WorldCom kept most of the cost off this year’s accounts. Profit appeared where a loss was actually sitting.
Here is the tell, in WorldCom’s own filed numbers. Line costs rise and fall with how many calls run across the network, so they should hold a fairly steady share of revenue. WorldCom reported them at 41.0 percent of revenue in 1999, 39.6 in 2000, and 41.9 in 2001. The dollar figure for 2001 came back to 14,739 million, the same to the last digit as 1999, as if it had been set by hand rather than run up by a real business. And while revenue was flat to falling, the company kept pouring money into new equipment faster than its existing equipment was wearing out, so the asset side of the books swelled while the business shrank. The costs it was quietly not counting had to go somewhere. They went into that equipment line.
The company later admitted the size of it: 3.055 billion dollars moved out of line costs in 2001, and another 797 million in the first quarter of 2002. Add the 2001 piece back where it belonged and line costs were not the reported 41.9 percent of revenue but closer to 51, and the year’s profit becomes a loss. Later restatements pushed the total past nine billion dollars. No short seller found this. An internal auditor did, by reading the capital-spending accounts.

I gave four of the agents these real figures, with the name and the years stripped off, and one neutral question. The fifth agent, governance, reads the people who run a company rather than its accounts, so I held it back for the case where the numbers alone would not be enough.
All four came back fatal. Running costs are being moved onto the asset side of the books, the reported profit is not real, and the single document that would settle it is the list of exactly what the company counted as new equipment that year. Reading nothing but an anonymized income statement, the agents pointed straight at the one schedule where the real fraud was later found.
Then three of the four added a line I had not asked for. This is the WorldCom pattern.
My blind test was not blind. The models knew the answer the way you know the punchline of a joke you have heard before.
So I tried to beat the memory. I shrank every figure down to a fraction of its real size, so the company now looked like a mid-sized carrier rather than a giant, and I kept every ratio identical to the last decimal. Then I ran it again. Three of four named WorldCom anyway.
You cannot hide the shape of a fraud by changing the numbers. The model recognizes the setup even when every number is different.
A telecom hiding its call costs inside its equipment spending is WorldCom to any well-read model, the same way a payroll full of fake users is Cupid. I want to be fair about what that recognition is, though, because it cuts both ways.
In my test that recognition was a nuisance, because I was trying to prove the machine could reason, and a model that already knows the answer proves nothing. But on a real job it is useful. Run this on a company nobody has questioned yet, and one agent saying the cost pattern here looks like WorldCom is a genuine early warning. The model has effectively read every fraud that was ever written up, so it can spot a familiar setup that a busy analyst, seeing the filing for the first time, might read straight past.
It is still only a warning, not a verdict. The same reflex will also point at honest companies that simply spend a lot on equipment, so you still have to do the work of deciding which one you are actually holding.
The one agent that never said the name is the one that showed the real skill. It reasoned from the numbers alone and stayed careful. Probably fatal, it said, but a cost that behaves like this can sometimes have a dull and innocent explanation, so pull the equipment schedule and check before calling it. That agent was slower, and it was the most useful of the four. The three fast agents recognized WorldCom; the slow one actually did the analysis. That gap is why the workflow ends with a judge weighing the finished arguments, rather than counting how many agents agreed. It is the same discipline I wrote about in making the models reason from the documents, not their own memory.
Which is leaving the real question still standing. If you take the memory away, can it still do the work? For this exact question, I needed a fraud that the AI model could not have memorised.




