Can We Trust an AI Review?
Reviewed September 2026 · Civil Systems LLCA review can read well and still miss the one thing the reviewer was hired to catch.
Would an 85%-accurate bridge calculation be acceptable?
The 85% is hypothetical. The question is what sits inside the other 15%—and whether anyone would notice before the work is built.In architecture, engineering, and construction, a missed note, a misread symbol, or the wrong specification edition can change the work. A review assistant might overlook a plan symbol for a long-lead component worth $50,000. That cost is an example, not a result from our testing. The problem is the omission: the comment log looks complete until the part is needed.
An evaluation, or eval, is a repeatable test of one job the AI is asked to do. We give it a submittal and the requirements that govern it, then compare its review with an answer key. We keep the full response so we can see not just whether it said “pass” or “fail,” but why.
What job is it doing?
Reviewing a TIA for completeness is different from checking a concrete mix ratio or reading a temporary traffic control plan. Define the task and its limits first.
What should it know?
Supply the applicable contract edition, plans, special provisions, checklists, and submittal—not a generic stack of vaguely relevant documents.
What happens when it misses?
Count false approvals, missed issues, invented comments, and unsupported certainty separately. Averages alone can bury the errors that matter most.
What a credible eval contains
Start with a small golden set: examples with a documented expected answer. A case might pair a concrete mix submittal with its governing standard and project provision, or a site plan with an agency checklist and review comments. Each answer should point to the requirement and the evidence. A qualified reviewer needs to check that key before it is used to make performance claims.
Then run those same cases with different models, prompts, and knowledge packets. Save the inputs and complete responses. The test runner makes the comparison repeatable; it does not make the answer key infallible.
What the models caught—and missed
We gave three small models running on a local computer eight concrete-mix questions. Six mixes had a problem; two met the conditions being tested. Each model saw the cases twice: first without the relevant requirements, then with short excerpts from the NCDOT specification and form instructions. For one case, the packet also included a made-up project special provision, clearly labeled as such.
The result in plain language
With the relevant requirements supplied, Llama was correct on 37.5% of cases (3 of 8), Granite on 50% (4 of 8), and Phi on 87.5% (7 of 8). These percentages describe one small test, not how the models perform on engineering work generally.
Without those excerpts, Llama declined to decide on all eight cases. Giving it the rules produced more decisions, but it also passed five mixes that should have been flagged. Providing the right document helped; it did not guarantee that the model would apply it correctly.
A confusion matrix simply separates correct flags, missed problems, false alarms, and correct passes. Read across the row for what the answer key says; read down the column for what the model said. Here, “flag” means the model returned FAIL. The problem cases in the “Passed” column are the ones a reviewer would worry about most.
Llama 3.2 · 3B
Relevant requirements supplied · one run per case
| Answer key | Flagged | Passed |
|---|---|---|
| Problem 6 cases | 1 | 5 |
| No problem 2 cases | 0 | 2 |
Caught 1 of 6 problems. Missed 5.
Granite 3.2 · 2B
Relevant requirements supplied · one run per case
| Answer key | Flagged | Passed |
|---|---|---|
| Problem 6 cases | 2 | 4 |
| No problem 2 cases | 0 | 2 |
Caught 2 of 6 problems. Missed 4.
Phi-3.5
Relevant requirements supplied · one run per case
| Answer key | Flagged | Passed |
|---|---|---|
| Problem 6 cases | 6 | 0 |
| No problem 2 cases | 1 | 1 |
Caught 6 of 6 problems. Raised 1 false alarm.
Caught: the case had a problem and the model flagged it.
Missed: the case had a problem and the model passed it.
False alarm: the case was acceptable on the stated test conditions, but the model flagged it.
Correctly passed: the case was acceptable and the model passed it.
Method note: These were eight controlled, synthetic cases, not actual contractor submittals. We used one run per model, not repeated trials. The answer key is provisional and has not yet received independent subject-matter review. The boxes count only the final PASS/FAIL decision, not the quality of the explanation, citations, or requested confidence value. Phi, for example, often omitted that confidence value. Other models, prompts, documents, and project types may perform differently. This is a pilot, not evidence that any model is ready to approve work.
The percentage is only the beginning
Phi's one wrong call was a false alarm. Llama's and Granite's wrong calls were mostly missed problems. They should not carry the same weight in a review workflow. A false alarm costs review time; a missed requirement can travel into procurement or construction.
Nor is this a contest to name the “best AI model.” These are three small local models, eight simplified cases, and one run of each setup. A larger model might do better, but it would still need to be tested on the actual work, documents, and consequences it will face.
The review question: Which issues did it miss? Which comments cannot be traced to the documents? And when the record is incomplete, does it stop and ask a person?
Keep the human accountable
An AI assistant may help screen documents, compare values, locate source pages, and draft comments. It cannot take professional responsibility for the work. A named reviewer still needs the original documents, enough time to check the findings, and authority to reject both the AI's recommendation and the submittal itself.
- Assign a named owner. Someone qualified remains responsible for the final disposition.
- Make evidence inspectable. Link each AI comment to the submitted item and governing requirement.
- Require abstention. Missing pages, conflicting provisions, or unclear symbols should trigger escalation—not a confident guess.
- Re-evaluate when the system changes. A new model, prompt, document set, or project type can change performance.
The point of testing is to decide where this tool belongs in the review process—and where it does not.
Start with a reviewable case.
Choose one bounded task, document the correct answer, and measure what the AI finds and misses before expanding the workflow.