What a good AI pilot looks like
Most AI pilots in construction fail on design rather than technology: the wrong workflow, no baseline measurement, no named owner, and no agreed definition of what "it worked" would mean. Fix those four and the technology question mostly answers itself.
The structure below fits in roughly six weeks, costs little, and ends with a number and a decision rather than a shrug.
There is now a fuller version of this. The guide to running a safe 30-day construction AI pilot sets out the governance in detail, including what UK data protection law actually requires and the RICS professional standard that has bound regulated surveying work since March 2026. Start there if the pilot is going anywhere near client work.
Two things to settle before week one, both of which sit outside the technology. If the pilot touches personal data, including anything captured from meetings, messages or CVs, the ICO's guidance for organisations sets the ground rules. And if it touches records that carry a regulatory duty, such as anything feeding the golden thread under the Building Safety Act 2022, the pilot has to leave the record no worse than it found it.
Which workflow should you pick?
This is where the pilot lives or dies. The right candidate has four properties, and dropping any one of them is what produces an inconclusive result.
| Property | Why it matters | Fails without it |
|---|---|---|
| High volume | Six weeks has to generate real evidence | Too few data points to conclude anything |
| Real pain | People doing it actively resent it | Adoption has to be forced, so it stops when you stop pushing |
| Recoverable mistakes | Errors caught in review, not in a final account | One bad outcome ends the whole programme |
| Measurable today | You can state now how long it takes and how often it goes wrong | No baseline, so no provable result |
Monthly reporting, meeting minutes into actions, tender summarisation, inbox triage and site diary compilation all qualify. Anything touching contractual notices, payments or safety decisions does not belong in a first pilot, whatever the vendor says it can handle.
Why does the baseline matter so much?
Because without the before, you cannot prove the after, and the pilot ends in opinions.
Keep it light: two weeks of honest numbers. How many hours the workflow takes, who does it, what it delays, and what the errors cost when they happen. Write it down.
This is the single most skipped step in every pilot, and skipping it is precisely why so many end in vibes instead of verdicts. It is also the step with the lowest cost, which is what makes omitting it so annoying in hindsight.
What rules should the pilot run under?
Six weeks, one team, one workflow. A named owner inside the business, not the vendor and not your consultant.
Every output reviewed by a person while trust is being earned, with a simple log of what needed correcting. Clear data rules agreed up front so nobody improvises with sensitive information, which is where the red lines from the published guidance earn their place.
Then leave the process alone long enough to learn from it. Adjusting the setup every three days resets the experiment each time and produces six weeks of week one.
How should you judge it?
Three numbers, then a commercial decision.
- Hours. Baseline time against pilot time, for the same output quality.
- Quality. Error and correction rates from the review log, trending across the six weeks. The trend matters more than the level.
- Adoption. Did the team keep using it in week six without being chased?
That third one is the most honest signal you will get, and it is free. People do not voluntarily keep using tools that waste their time, which is the same signal that tells you whether training worked. If adoption held without pressure, the hours number is probably real.
Then decide as you would on plant: keep and scale, fix one specific weakness and re-run, or kill it and write down why. A cleanly killed pilot is a success. It cost six weeks and taught you something true about your business. The expensive failure is the pilot that limps on for a year because nobody defined what failure looked like.
What is the real return on the first one?
The playbook, not the workflow.
One well-run pilot teaches you how your data behaves, what your review process needs, and how your people build trust in these tools. The second workflow rolls out in half the time and the third is routine. That compounding is the actual goal, and it is why the first pilot should be chosen for what it teaches rather than for the size of the prize.
Not a demo that impresses a steering group. A repeatable way to turn capability into measured hours, one workflow at a time.