10 min· building· ai· water
Why can't we just ask ChatGPT?
“Can’t we just use ChatGPT?” is a real question a prospective customer asked me, and it is the right one. The profile was a document, and documents cannot be asked questions. So we built Atlas — an agent over the same curated record that answers in specifics, shows where each answer came from, and turns an answer into the next piece of work. We scored it against ChatGPT, Gemini and Claude on fifteen questions, and the most useful result is the one where everybody drew.
The last piece ended on the thing a profile cannot do. It is a document, and a good one leaves the reader holding a second question it was never going to answer. Is that site on the same permit as the other one. Has this been back to the board since. Who signed the last contract.
The reader either emails us and waits a day, or drops it. Mostly they drop it.
So the obvious thing to build was the thing everyone has now used: a chat window. Ask a question in plain language, get an answer. The unobvious part — and the only part that turned out to matter — is what sits behind it.
The job, as it exists today
Somebody has read enough to be dangerous and now has a specific question.
Their options are three. Ask a colleague, which works if the colleague happens to know and is not in a meeting. Go and look it up themselves, which is the reading problem this whole series is about, now shrunk to one question and therefore even less likely to be worth their afternoon. Or ask a general-purpose assistant, which is what most people actually do, because it takes eleven seconds.
The third option is the interesting one, because it usually produces something that reads like an answer.
"Can't we just use ChatGPT?"
That is a direct quote. Somebody asked me it in a meeting, and it has been asked often enough since that it now has its own slide in the deck. It is also the right question, and the honest answer starts by conceding most of it.
Not because the models are weak. They are extremely good, and this is worth saying carefully, because "the big models can't do it" is the laziest claim in enterprise software and is usually false. Atlas is itself built on a frontier model, and against the ungrounded version of that same model the difference is not intelligence. The gap is not the model. It is what reaches the model.
A frontier model has read the open internet. What it has not read is a particular authority's board packet from 2019, a permit renewal file, the engineering report attached to an agenda as a scanned appendix, or the twelve years of minutes in which a plant was discussed. That material is public but it is not convenient — it is scattered across thousands of sites in formats nobody standardised — and being public is not the same as being in a training set.
So when you ask a question whose answer lives only there, a general assistant does one of two reasonable things: it declines, or it answers the general version of your question well. Both are honest. Neither gets the person off the call with an answer.
The comparison I keep coming back to is a medical one. Asking a general-purpose assistant about a specific utility is like going to a general physician and asking about your heart. You will get an answer and it will be sound, careful, general medicine. What you wanted was a cardiac specialist who has read your chart.
Two different advantages are hiding in that sentence, and it is worth separating them, because a lot of "vertical AI" claims only have the first. The specialist knows one part of the body far better than a generalist ever will — that is depth in a domain, and it comes from the corpus. But the phrase that does the real work is who has read your chart. Knowing cardiology in general is not the same as knowing what your heart has been doing since 2019, and in this market the second is the scarce one, because the chart exists in public and almost nobody has read it.
What the workflow does
It answers only from the collected record. The corpus is the same one the profile is assembled from — the utility's own documents, the compliance record, the funding record — organised, cleaned and indexed. The agent retrieves against that, and against the structured data alongside it, rather than against everything ever written.
It shows where each answer came from. Every claim carries the document it came from, by name, so the person asking can go and check it. This is the part I would keep if I had to throw away everything else.
It answers at the granularity of the question. Ask a broad one, get the shape of the programme. Ask for a permit number, a violation count, a date a term expires, and you get that, or you get told plainly that it is not in the record. The house rule about saying what you do not know applies here more than anywhere else in the series, because a chat window is the easiest place in software to produce confident nonsense.
And an answer can become the next piece of work.
This is the part that makes it a workflow rather than a search box, and it is the same principle the first article set down: the job is not finished while somebody still has to read the output and decide what to do about it. So the follow-ons attached to an answer are not all questions. Some of them are the other workflows in this series, invoked from where you are standing — pull the verified contacts for this account, generate the full profile, draft the outbound this answer would justify, write it back to the record in the CRM.
That turns out to be the agent's real position in the product. It is not a sibling of the other workflows sitting alongside them in a menu. It is the place the others get called from, because a question is how people actually arrive at wanting one.
And it proposes the next question. This is small and it is the feature people actually react to. Every answer comes with three or four follow-ons drawn from what the record can support — pull the contacts, show the limits in detail, roll this up to the parent organisation. It matters because the hardest part of using a system like this is not phrasing a question. It is knowing which question is available to ask, and a blank box tells you nothing about that.
What happened when we tested it against ChatGPT, Gemini and Claude
We ran the comparison properly, and it is scored and visible inside the product rather than living in a slide. Fifteen questions on one utility account, asked from a vendor's point of view — procurement routes, installed base, the buying committee, funding capacity, where to lead and where not to. Same question, same wording, to every system. Every answer given a verdict against a documented ground truth.
Two choices make it a test rather than a demonstration.
One of the systems we tested against is the model we are built on. Atlas runs on a frontier model, so one baseline is that same model with no tools, no corpus and no account context. Any gap between them is the grounding and nothing else. Compare yourself to a weaker model and you have measured the model.
And the set includes a control question we expected to draw — a general industry question about routes to market, the kind a well-read assistant should answer as well as we do.
The control drew. All four systems answered it correctly, which is the single most important line in the study, because it is the evidence that the other fourteen are not rigged.
On the other fourteen: Atlas answered eleven of fifteen grounded or correct. The best generic baseline fully matched two.
Values
| Tier | Atlas (grounded) | ChatGPT | Gemini | Claude (ungrounded — the model Atlas runs on) |
|---|---|---|---|---|
| Grounded or correct | 11 | 2 | 2 | 1 |
| Partly right | 4 | 7 | 5 | 2 |
| Vague or declined | 0 | 3 | 2 | 12 |
| Confidently wrong | 0 | 3 | 6 | 0 |
The failure that matters is not that they lost. It is how they lost.
Nine times, the general-purpose assistants stated something specific, confidently, and wrong. Two of them independently described the utility as buying mainly through open competitive tender, when its actual award history runs almost entirely through other routes — an error that inverts a vendor's whole approach to the account. One took a three-letter abbreviation that means one thing in wastewater compliance, read it as a sales-operations term, and built a full answer on the misreading. Another turned an acronym out of the utility's own regulatory file into the name of a staff training scheme. A third reported that the utility had no automated metering network, when the incumbent system had been running for years.
None of those answers looks wrong. They are fluent, plausible, and would survive being repeated in a meeting — which is precisely the problem, because the correction arrives from the customer.
What the same scoring says about Atlas is that it is grounded.
Eleven of its fifteen answers were grounded in the record or correct outright. The remaining four were partial — less than a complete answer, but drawn from what the documents actually support rather than from around them. Nothing it returned was a confident invention.
That is the property the whole thing is built for, and it is a narrower claim than it sounds. It does not mean the answer is always complete. It means the answer is made of the record, and you can go and look.
What it cannot do
It is worse than a general assistant at almost everything.
That is not modesty, it is the design. Ask it to explain a treatment process in general terms, draft something unrelated to a utility, or answer a question about anything outside this sector, and you should close it and open one of the others. It knows one corner of the world in detail and is deliberately incurious about the rest. On the general control question above, ChatGPT and Gemini answered as well as we did and wrote it up better — and a title claiming otherwise would be the sort of thing this series is supposed to avoid.
The failure I would be worried about is the opposite one — a narrow system that also tries to be a general one, and quietly answers from the open internet when its own record comes up short. Then you have a tool nobody can calibrate, because you cannot tell which kind of answer you just got.
And a conversation is not a record.
An answer in a chat window is read by one person and then it is gone. It does not brief a colleague, survive a handover, or sit in a folder six months later when the deal comes back. That is why the profile did not go away when this arrived, and why the pair of them is the actual product: the document is what you take into the room, and the agent is what you use when the document raises something it did not cover.
We built these in that order, and I had assumed the second would replace the first. It did not, and the reason is obvious in hindsight. They are not two versions of the same thing. One is for reading, and one is for asking.
It is also bounded by the corpus, which is the same boundary the screening piece draws around coverage — deep where the record is deep, thin where it thins, and honest about which one you are in. An agent cannot answer from documents nobody collected.
Next
The remaining workflows are about acting rather than understanding: drafting outbound that is grounded in all of the above, and then writing it back into the system the sales team already uses — which is, still, the step that decides whether any of this gets used at all.