The jagged edges the AI labs won't smooth
Frontier AI is jagged: superhuman at some tasks and unreliable at others that look no harder. Since December the labs have smoothed a lot of it, memory, long context and agents among them, and we use all of it. What they have no reason to smooth is the edge that belongs to one market: a sector's record consolidated and organized, an understanding of what its vendors need, and the judgment to pull the signal out of the noise. That is where we work.
"Can't we just use ChatGPT?" is a question I have been asked in a meeting, and it gets asked often enough that it has its own slide in our deck. I wrote about how we answer it a few weeks ago. This piece is about the harder version, the one investors ask: if the models keep getting better this fast, what is left for you to do in two years?
The most useful frame I have found for it comes from Ethan Mollick, a Wharton professor who writes about how people use these tools. In December he wrote The Shape of AI, built on an idea he and his co-authors named in 2023: the jagged frontier.
A lot has moved since December, and I will get to what. But I believe the frame holds, and that the most durable work in applied AI is on the edges of that frontier the labs have no reason to smooth.
What jagged means
In 2023 a team from Harvard, Wharton, MIT and Warwick ran a preregistered experiment with 758 consultants at Boston Consulting Group. On eighteen realistic consulting tasks the AI was good at, the consultants using it finished 12.2% more tasks, finished them 25.1% faster, and by Mollick's account produced work rated about 40% higher in quality.
The study also set one task chosen to sit just outside what the AI could do. It looked no harder than the others. Without AI, consultants got it right 84% of the time. With AI, 60 to 70%. The tool made good people worse, because its wrong answer was fluent and convincing.
That is the jagged frontier: the line between what AI does brilliantly and what it does badly does not follow our sense of what is hard. We saw the same shape in our own testing. Put fifteen questions about one utility account to ChatGPT, Gemini and Claude, and nine of their answers were specific, confident and wrong. None of them looked wrong.
Ten months later
Mollick's December piece named memory as the weak spot, the ability that had improved least. That is no longer the right picture.
Memory became a product. Claude's memory reached every user, free ones included, in March; by August it was a set of categorized entries you can see and edit, shared across chat and agent work (release notes). Gemini will now import your memories from other assistants. What shipped is storing and retrieving rather than a model that learns on the job, but for most work that distinction no longer bites.
Context stopped being the constraint. A million tokens is now the default window on Claude's current models, and agents compact their own history to run for hours. Anthropic's own documentation still warns about "context rot", accuracy that degrades as the window fills, a reminder that the edge has moved rather than vanished.
Agents got long and plural. METR measures the length of task an AI can finish half the time. In December the best models sat at five or six hours. By April it was about seventeen, doubling roughly every four months, though at 80% reliability the figure is closer to three hours. Agent teams arrived in February, and in his newest post Mollick describes swarms of thousands of agents working one problem.
And it is still jagged. ARC-AGI-3, a set of puzzles built to defeat AI, launched in March with every frontier system under 1%. By September one model scored 62.7%, and 99.9% with its maker's own harness. Meanwhile the Remote Labor Index, which gives agents real paid freelance projects, saw its best score rise from 2.5% at release to 15.8% by July, with the note that "today's AIs still fall short of professional quality on most projects." On the hardest tier of Berkeley's Agents' Last Exam, every frontier agent tested scored zero, and the most common failure is that "agents declare success before they've truly verified their work."
Mollick's own verdict in August: AIs "are still jagged, and can lag far behind human experts on parts of their work." The frontier moved, a long way. The jaggedness moved with it, away from remembering and reading and toward finishing real work, checking it, and knowing what was never put in front of the model.
Jagged becomes a bottleneck
Mollick's lasting point is about what jaggedness does to a whole job. "A system is only as functional as its worst components," he writes. A job is a chain of tasks, and if the AI is superhuman at nine links and unreliable at the tenth, the tenth sets the pace.
His example: a research team used a large language model workflow to reproduce an entire issue of Cochrane reviews, the meta-studies that settle what the medical evidence says. Across twelve reviews and 146,276 citations it screened and extracted data more accurately than the graduate-level human reviewers it was tested against. Mollick quotes the paper's first version putting it at "approximately 12 work-years" of work, done in two days. What it could not do was open the supplementary files attached to papers, or email an author for unpublished data. That small edge is why the process still needs a person.
He adds two things. Some bottlenecks have nothing to do with ability: "institutions move at institution speed." And the labs attack the bottlenecks that hold everyone back, which is why memory and context moved so fast this year. "Don't watch the benchmarks," he says. "Watch the bottlenecks."
Our chain, link by link
Here is our work as a chain. A vendor wants to know which utilities will need what it sells, and when, early enough to be useful.
Reading, remembering, running agents are the labs' links, and we use whatever they ship. A frontier model reads a capital plan faster than any person can. We will not build memory or context management; when the labs improve them, our product improves the same week. We stopped building our own model for this reason, which I wrote about in The model was never the hard part.
Consolidating the record is not theirs. The documents that answer a vendor's question are board packets, capital plans, permit renewals and engineering reports scattered across thousands of municipal websites, many of them scanned. They are public and they are not in any training set. A model's memory can only keep what someone put in front of it, and nobody has put this in front of it. That is the Cochrane supplementary file, at the scale of a whole market.
Knowing what the vendor needs is not theirs either. A metering company once told me that if it learns about an opportunity when the RFP is published, it is too late; it wanted to be in the conversation two to three years earlier, and asked me what the signals were. I was stumped. That question is a vendor's pain point, and it is different for a meter company, a coatings company and an engineering firm. A general model has no reason to know which line in a budget matters to which of them.
Finding the signal in the noise is where the confident errors live. A utility's agendas, minutes, budgets and filings pile up every month, and almost all of it is routine. The judgment is in which change matters to whom, and in refusing to declare success without a source. That is the failure mode the agent benchmarks keep finding, and it is the one this market punishes.
Acting at the right time is an institution bottleneck. A utility board meets once a month. I believe sales cycles of 6 to 24 months are too long and not needed, but no model shortens a board calendar, and I have seen utilities that still run on Windows 98. The pace of the buyer is a fact to design around.
Two kinds of edges
Side by side, those links fall into two groups.
Capability edges hold back every customer of every lab: long context, memory across sessions, agents that run for hours, messy scans. These are the labs' reverse salients, and this year shows they will be fixed, fast. Building a business on one is building on ice in March. We take each fix as it ships.
Context edges hold back one market. Nobody has consolidated the records of sixty-six thousand utilities and organized them by utility, project and decision. Nobody has written down what each kind of vendor needs to know, or which of a month's changes matter to it. Each of these is labor, domain knowledge and judgment in a narrow market, which is exactly why a lab will not do it.
Could a lab collect every board packet? Nothing stops one. But it would then need to know what a meter vendor and a pipe-coatings vendor each need from the same packet, and that comes from working with vendors, not from a bigger model.
What we build
So our moat is three things, and none of them is a model.
The sector's data, consolidated and organized. Every document we can find, turned into something machine-readable, and organized by utility, project and decision with the dates kept, so a question about one account draws on all of its record.
An understanding of the vendors' pain points. What a vendor in each category needs to know, how early, and in what form, learned from working with them rather than guessed.
The ability to extract signal from the noise. Which of a month's changes matters to which vendor, said with the document behind it, and a plain "not in the record" when the answer is not there.
The memory, the context and the agents underneath all of that come from the labs, and every improvement they ship makes the three things more useful. That is the whole answer to the investor question. As the frontier moves, our side of the work gets better for free. The edges we work on do not move with it, because no lab has a reason to move them.
That is what AquaIntel is built on: the water sector's record, consolidated and organized; an understanding of what its vendors need; and the judgment to pull the signal out of the noise, with a source behind every claim.
Related
- Why can't we just ask ChatGPT?: the benchmark behind the nine confident errors.
- The model was never the hard part: why we stopped building our own model.
- The problem statement I couldn't unsee: the metering client's question, in full.


