A product manager asks a model for a competitive brief and gets four pages in ninety seconds. It is well organised, confidently written, and cites things. She skims it, it looks right, and she sends it to Sales.
The account executive who receives it spends the next two hours working out which parts are load-bearing. Two of the claims turn out to be stale. One is about a competitor’s pricing that changed in March. He rewrites it, sends the corrected version to nobody in particular, and the original stays in the shared drive.
Total time to produce: ninety seconds. Total time to consume: two hours, plus whatever it costs the next person who finds the uncorrected copy.
The measurement caught up in 2025
Researchers at Stanford Social Media Lab and BetterUp Labs surveyed 1,150 US employees about exactly this: AI-generated work that looks like a finished contribution but does not actually advance the task. They named it workslop. Forty percent said they had received some in the previous month. Each incident cost about one hour and fifty-six minutes to sort out, which they put at roughly $186 per person per month, and north of $9 million a year in a large organisation (AI-Generated Workslop Is Destroying Productivity).
The detail that matters most is the one least quoted. Around half of it moves between colleagues, peer to peer, across the seams in the organisation. This is not individuals producing bad work for themselves. It is work that reads as complete right up until it crosses a team boundary, where the person receiving it has neither the context to validate it nor a reason to distrust it.
Two caveats, since they matter. That study is self-reported, and it is one study. It establishes that the pattern is common and expensive at the scale surveyed. It does not establish a precise per-company cost, and anyone quoting the $186 as though it were an audited figure is overreaching.
Nineteen out of twenty pilots returned nothing
MIT’s Project NANDA looked at somewhere between $30 and $40 billion of enterprise GenAI spending and found that about 95% of organisations got no measurable return from their pilots (The GenAI Divide).
The number went around. The diagnosis attached to it did not, and the diagnosis is the interesting half. The report’s own account of why is not model quality, not infrastructure, and not regulation. It is learning. The systems being deployed do not retain feedback, do not adapt to context, and do not improve with use. They answer, and then they forget, and the next person asks the same thing again.
Gartner expects the correction to be visible in the budget. They forecast that more than 40% of agentic AI projects will be cancelled by the end of 2027, on rising costs, unclear business value, and inadequate risk controls (Gartner, June 2025).
That one is a forecast rather than a measurement, and forecasts about AI have a poor recent record in both directions. Read it as a statement about what analysts are hearing from buyers, not as a fact about 2027.
Why another agent per team makes it worse
The instinctive response to all of this is more and better AI. Give Sales an assistant that knows Sales. Give Product one that knows Product. Give every team a copilot trained on its own corpus, and the quality problem goes away.
It does not, and the reason is structural. N teams with N assistants means N memory stores, N versions of what is currently true, and N reconciliations that never quite finish. You have rebuilt the org chart in software, including the walls. When Product changes its mind, Product’s assistant knows. Sales’ assistant does not, and has no mechanism by which it ever would. The competitive brief in the shared drive is still wrong, and now there are four assistants confidently citing it.
The architecture everyone is reaching for is the architecture that already failed at the org-chart level.
This is probably why the productivity gains have been so hard to find in aggregate. Speed inside a team was never the constraint. The handoff was. Every one of these tools made the fast part faster.
Both findings point at the same missing piece
Put the NANDA diagnosis next to the workslop finding. One says AI systems fail in companies because they do not retain what they learn. The other says AI output fails at team boundaries because the person receiving it cannot tell what is still true. Those are two descriptions of the same hole: no shared, durable record of what this company has actually decided.
So that is what we are building, and it is deliberately not another agent. One memory of the decisions a company makes, holding who decided, when, why, and which earlier decision it replaced. A model of which teams and which in-flight commitments each decision touches. And delivery into the tools people already work in when something they are standing on stops being true. A person approves anything that leaves the company. Nothing goes to a customer on a model’s say-so.
What we cannot yet tell you is the size of the effect. There is no public dataset measuring how often a changed decision fails to reach the teams downstream of it, which means there is no baseline to improve on and no honest way to quote a percentage. We are running structured customer interviews to produce that number ourselves. It will be the first statistic on this site we authored end to end, and we will publish it whether or not it flatters us.
Until then, here is a cheap test on your own numbers. Take the AI tools your company bought in the last eighteen months and ask which of them made a handoff between two teams faster, as opposed to making one team’s output faster. If the honest answer is none of them, that is not a procurement failure. It is the whole category pointing at the wrong constraint.
Frequently asked
- What is workslop?
- AI-generated work that looks like a finished contribution but does not advance the task, so the person receiving it has to redo it. Research from Stanford Social Media Lab and BetterUp Labs, published in HBR in September 2025, surveyed 1,150 US employees: 40% had received some in the previous month, at an average of one hour and fifty-six minutes per incident and roughly $186 per person per month. Around half of it changes hands between colleagues.
- Why did 95% of enterprise GenAI pilots show no return?
- MIT Project NANDA reviewed $30 to $40 billion in enterprise GenAI spending and found about 95% of organisations got no measurable return. Their diagnosis was not model quality, infrastructure, or regulation, but learning: the deployed systems do not retain feedback, adapt to context, or improve with use.
- Why does giving each team its own AI assistant not fix cross-team alignment?
- Because N teams with N assistants produces N memory stores and N versions of the truth. Each assistant knows what its own team decided and has no mechanism for learning what another team changed. It reproduces the existing silos in software, and the handoff, which was already the slow part, stays slow.
- Is AI actually making cross-team work worse?
- The evidence points that way, though it is early. AI raised the rate at which each team produces work without changing the rate at which work transfers between teams. More output crossing the same unimproved seams means more material arriving that the receiver cannot validate. That is consistent with the workslop finding that roughly half of it moves peer to peer.
Sources
- Niederhoffer, Kellerman, Lee, Liebscher, Rapuano and Hancock, AI-Generated Workslop Is Destroying Productivity, HBR (September 2025)Self-reported survey of 1,150 US employees. Paywall-free summary at Axios.
- Axios summary of the workslop study (September 2025)
- MIT Project NANDA, The GenAI Divide: State of AI in Business 2025The 95% headline figure has been widely debated on methodology. The diagnosis about retained learning is the part we rely on.
- Gartner, Over 40% of Agentic AI Projects Will Be Canceled by End of 2027 (25 June 2025)A forecast, not a measurement.
If any of this is recognisable at your company, tell us where it costs you the most.