Which AI model should your construction team actually use?
If you are picking one to start with, start with ChatGPT for general office work, or Claude if your commercial team spends its week inside contracts and subcontract documents. Everything after that is a second decision, not a first one.
The confusion starts with the names. ChatGPT, Claude, Gemini, Perplexity and Grok are different products built on different frontier models, and each is genuinely better at something. The model is the engine, not the car. None of them replace your team. They take the drudge work off the desk so your people spend their time on the work that earns money and protects margin.
One caveat before the table: this field moves faster than any article can. Treat what follows as a starting shape rather than a settled ranking, and check the current capability and data handling on the vendor's own pages before you commit (Anthropic's Claude and Google's Gemini both publish theirs). What does not move is the question you should be asking of any of them, which is whether your data trains their model, covered in private AI versus public chatbots.
| Model | Best for | Where the hours come back | Pilot it first if |
|---|---|---|---|
| ChatGPT | Broad mixed office work | Reporting, drafting, document review | You want one starting point for a whole office |
| Claude | Contract and document-heavy work | Contract review, scope-gap analysis, quote comparison | Your QSs and commercial managers read documents all week |
| Gemini | Firms already on Google Workspace | Search across Gmail and Drive, Sheets analysis | Switching cost is your main blocker |
| Perplexity | Anything that needs a source | Regulations lookup, price movements, BD research | You need answers you can check |
| Grok | Real-time news and sentiment | Planning and reputational monitoring | You have public-facing schemes with planning exposure |
| Self-hosted open source | Work that cannot leave your firewall | Private assistants over your own libraries | Confidentiality rules out public tools |
What is ChatGPT best for in a construction office?
Broad, mixed office work. It is the strongest general starting point if you are rolling out to a whole team rather than to one discipline.
It drafts replies to RFIs, early warnings and delay letters using your own templates as examples. It summarises large tender packs into key risks and information requests. It turns meeting notes into action trackers with owners and due dates. It also underpins Microsoft Copilot, so the same engine may already be sitting inside your stack, paid for and unused. If that sounds familiar, the problem is usually adoption rather than the tool, which is a different problem with a different fix.
Which model should a commercial or QS team pilot first?
Claude. If your commercial people spend a fifth to a third of their week reading and reconciling documents, this is where the return shows up fastest.
Upload a 200 page NEC4 or JCT contract and identify onerous clauses, liquidated damages and conditions precedent. Review subcontract inclusions and exclusions into a scope-gap summary a QS can act on. Compare two complex quotes with discrepancies in provisional sums and qualifications flagged. Turn a contract pack with Z-clauses into a plain-English playbook: when X happens, do Y within Z days.
That last one matters more than it sounds. Fewer missed notices means less margin lost to out-of-time claims, and the disputes that end up costing real money rarely start big.
What if the business already runs on Google Workspace?
Use Gemini, mainly because the switching cost is close to zero.
Search across email and Drive in one question. Turn existing Sheets KPIs into a monthly board narrative. Generate formulas and pivots without writing them. Ask questions of your own SOPs and lessons-learned folders. The raw capability is not dramatically different from the alternatives. The adoption curve is shorter, and adoption is usually the binding constraint rather than capability.
Which model should you use when you need a source?
Perplexity, and use it alongside another model rather than instead of one.
It is the best first stop for anything where the answer has to be checkable: current Building Regulations requirements with links to the actual documents, live material price movements before an estimate goes out, competitor case studies before a BD meeting, manufacturer claims verified before procurement.
Is Grok worth it for a contractor?
For most contractors, no. It is useful but niche.
Its strength is real-time monitoring of news and social sentiment, which genuinely matters if you carry significant planning exposure or reputational risk on public-facing schemes. For core project delivery it is not your first choice, and treating it as one is a good way to waste a pilot.
What if project data cannot leave your firewall?
Self-host an open-source model. Llama, Gemma and Qwen can run on your own private UK server, so sensitive project data, client lists, pricing and margins never leave your environment.
This is the route for private assistants over SharePoint libraries, commercial correspondence archives and subcontractor records. It is more work to stand up than a subscription, and for some firms it is the only option that passes an information security review. The comparison of public, enterprise and private AI covers where each option actually sends your data.
What rule matters more than the model you pick?
That all of these are copilot tasks. The AI does the legwork and your people make the decisions.
No automatic issue of instructions, contractual notices or compliance decisions without human review. That is not caution for its own sake. It is the line that keeps a tool useful rather than dangerous, and it is why controlled adoption works better than a blanket ban.
Start with one team, one model and two or three workflows that visibly eat hours. Prove it, then expand. That is how adoption sticks, and it is how the saving turns up in the accounts rather than in a slide deck.