Evaluating AI for Construction: The Honest Buyer’s Guide
Last updated 2026-07-13
Every serious platform in this market is shipping AI. So "does it have AI" is not a question — it is a formality, and it will be false about somebody by next quarter.
This guide is about the questions that stay useful. It includes the ones we would rather you did not ask us, because a buyer’s guide that only asks questions its author passes is a brochure.
The first question: does it act, or does it store?
A system that surfaces your documents faster is a search box with good manners. A system that drafts the RFI response, flags the notice nobody served, catches the duplicate requisition and matches the invoice to the delivery is doing work.
Both are legitimate products. They are not the same product, and they are not the same price. Establish which one you are being shown, in the first ten minutes, by asking what the system DID rather than what it found.
The second question: what can it see?
This is the one that actually predicts whether an AI capability will be useful, and it is almost never asked.
An agent reading your RFI is useful. An agent reading your RFI, the contract clause it turns on, the instruction that preceded it, and the delivery that never arrived is doing a different job — and it is not doing it because of a better model. It is doing it because those four things are on the same record.
AI quality is a function of the data foundation underneath it. That is an architectural fact rather than a release note, and it is why "which model do you use?" is a much less interesting question than it sounds. A superb model reading a fragment produces a confident answer about a fragment.
The third question: can it show its sources?
Ask to see a generated document. Then ask which records each claim came from.
If the answer is vague, you cannot review the output — you can only trust it, and trusting an output you cannot check is not a workflow, it is a hope. Whereas a draft that cites its sources turns the review question from "is this true?" — which takes an evening — into "is this the right record?", which a competent engineer answers in seconds.
And the corollary is the most valuable thing a drafting system produces, which almost nobody markets: the gaps. A draft that says "the specification section this refers to is not in the document set" has found a hole in your project record. That hole was going to hurt you in a dispute whether or not you ever ran the tool. Finding it on a Tuesday is considerably better than finding it in cross-examination.
On accuracy percentages: be suspicious
You will be shown accuracy figures. Treat them the way you would treat a contractor quoting a productivity rate with no project attached: the number is meaningless without the test set, and the test set is never published.
Accuracy against what, exactly? A submittal review has no single correct answer. Two competent engineers will flag different things and both will be right. A headline accuracy figure implies a ground truth that does not exist — so a vendor producing one has either invented it or measured something far narrower than the thing you are buying.
The honest framing is different, and it is not weaker. AI review is a first pass, and its comparison is not against a perfect reviewer. It is against what happens today: a tired engineer reading the fortieth submittal on a Thursday evening and skimming the last twenty pages. Judged that way, the value is obvious — it never gets tired, it reads the whole document every time, and it checks the clause reference against the actual contract rather than from memory.
Human sign-off is not a safety feature. It is the structure.
A document that goes to a client with your company’s name on it was written by your company. An engineer signs it, and if it is wrong, that engineer and that firm carry the consequence — professionally, contractually, and in some jurisdictions personally.
That does not transfer to a model, and it is not a technology limitation waiting to be engineered away. It is the basis on which construction professionals are trusted, insured and paid.
So the reviewer is not in the loop as a safeguard. The reviewer is the AUTHOR, and the tool wrote the first draft. Any vendor telling you their system removes the need for review is describing a risk transfer to you, and pricing it as a feature.
The demo problem, and how to defeat it
Every demo is impressive. Demos are optimised to be impressive in forty minutes, on a dataset that was chosen.
Your project is not a chosen dataset. It is a document set with three revisions of the same drawing, a correspondence thread that changed subject halfway through, a specification with a missing section, and a schedule that has not been updated since March.
So take one real, ugly export from your current system and make every vendor run their product on it, live, in front of you. This single request separates capabilities from demos faster than any evaluation matrix, and the reaction to being asked is itself informative.
The tell
Vendors selling a capability will volunteer the limits. Here is what it does not do. Here is where it needs a human. Here is the case where it gets it wrong.
Vendors selling a demo will not, because acknowledging a limit is off-message. And a system that always produces an answer is worse than one that sometimes says "I could not find this in your records" — because the first one is wrong silently, and the second one is telling you something true about your project.
Ask to be shown a case where it declines. If nobody can show you one, you have learned something.
And the questions to put to any vendor, including us
These five are evergreen, and they are how you find today’s answers rather than reading somebody’s claim about last quarter:
- Does the AI review documents against OUR drawings and contract — and cite its sources?
- Does the platform follow the asset after handover, or does its record end at closeout?
- Does pricing scale with our construction volume, our user count, or neither?
- How long from kickoff to a productive site team — weeks or quarters?
- Where does our data live, and what happens to it if we leave?
The risk nobody discusses, which is not job losses
It is deskilling, and it deserves to be named rather than waved away.
If juniors never assemble a claim by hand, do they learn what a good one looks like well enough to spot a bad draft? That is a real question and the industry has not answered it.
The mitigation is the same discipline that makes the output trustworthy in the first place: read the citations before the prose, treat the gaps as findings, and require the reviewer to be the author rather than an approver. A professional checking sources against records is still doing professional work — arguably more of it, and on the part that matters.
Common questions
What questions should you ask an AI construction software vendor?
Does it act or only store? Where does the content come from? Can it cite its sources? Who signs off, and is it enforced? What happens when it is wrong? And what does it do on the day your data is messy — which is every day. Then make them run it live on a real, ugly export from your current system.
Read the full answerHow accurate is AI document review in construction?
There is no credible published benchmark, and anybody quoting one should be asked where it came from. A submittal review has no single correct answer, so there is no ground truth to measure against. The honest framing is that AI review is a first pass that never gets tired — judged by what it surfaces for a human, not by a percentage.
Read the full answerDoes AI replace construction professionals?
No — and the structural reason is accountability, not reassurance. A document with your firm’s name on it was written by your firm, and a model cannot be held responsible for anything. What disappears is the assembly work; what remains is the judgement, which is the part that needed a professional in the first place.
Read the full answerWhat is the review discipline for an AI-drafted document?
Read the citations before the prose — a document that reads beautifully and cites the wrong records is far more dangerous than one that reads badly, because nobody checks the beautiful one. Treat the gaps as findings. And remember you are the author, not the reviewer.
Read the full answerReferences
- No accuracy benchmark is cited in this guide, because no credible published one exists for construction document review. We would rather say so than manufacture a figure — which is the posture this site takes on every page where the evidence does not support a number.
- The evaluation questions here are the same five we publish on our comparison pages, and they are asked of any vendor, including us.
In depth
- What questions should you ask an AI construction software vendor?
- How accurate is AI document review in construction?
- Does AI replace construction professionals?
- What is the review discipline for an AI-drafted document?
- Can AI draft construction documents, and where does the content come from?
- Does the AI make decisions on its own, or is there human sign-off?
- How does Zepth’s AI cite its sources (clause, drawing, spec)?
- AI-native vs bolted-on construction platform — what’s the difference?
- Is AI accurate and safe enough to use on real construction projects?
- Can AI draft an extension-of-time claim?
See Zepth on your project.
A short, tailored walkthrough on your real workflow — no generic demo.
Book a meeting