Skip to content
AI Metric

Chris M.Reviewed

How to Run a Safe 30-Day Construction AI Pilot

The pilot that causes trouble is rarely the one that fails. It is the one that ends after six weeks with nobody able to say whether it worked, because no baseline was taken, the scope drifted, three people were using it in ways nobody agreed, and a subcontractor's personal details went into a consumer chatbot on the second Tuesday.

A safe 30-day pilot has five things fixed in writing before anyone touches a tool: one named task rather than a general trial, a measured baseline of how that task performs today, a rule about what data may and may not go in, a named person who reviews the output and is accountable for it, and a stop condition that ends the pilot early if it is met. Everything else is detail. Those five are the difference between an experiment and an incident.

This guide sets out that structure and, more usefully, is honest about what actually binds a UK contractor. The answer is less than most people assume, and in one specific case considerably more.

What the law actually requires, and what it does not

There is no UK AI Act. The government's 2023 white paper set out five principles, safety and security and robustness, appropriate transparency and explainability, fairness, accountability and governance, and contestability and redress, and said that it would not put those principles on a statutory footing initially and did not intend to introduce new legislation 12. Those principles are addressed to regulators rather than to businesses, so a contractor cannot comply with them or breach them directly 12.

That is a 2023 position and the ground has moved since, so it is worth being precise about how. Secondary legislation specific to artificial intelligence now exists: the Data Protection Act 2018 (Code of Practice on Artificial Intelligence and Automated Decision-Making) Regulations 2026 came into force on 12 May 2026, and regulation 2(1) requires the Information Commissioner to prepare a code of practice on good practice in processing personal data in relation to developing and using artificial intelligence, and automated decision-making 15. It places no direct duty on a business 15. What it means in practice is that the guidance a contractor is measured against is about to be rewritten, so anything built now should be built to survive that.

What binds you is existing law, principally UK data protection law, and it binds you the moment personal data is involved. That includes a great deal of ordinary site material: names in a site diary, a subcontractor's operatives on a timesheet, faces in a progress photograph.

The Information Commissioner's Office states that in the vast majority of cases the use of AI will involve processing likely to result in a high risk to individuals' rights and freedoms, and will therefore trigger the legal requirement to carry out a Data Protection Impact Assessment 3. The screening test is more specific than that summary suggests: innovative technology, including AI, is one criterion, and a DPIA is required where it combines with another of the listed criteria, with a combination of two factors usually indicating the need for one 2. Where you conclude a DPIA is not needed, you still have to document how you reached that conclusion 3. There is no version of this where nothing gets written down.

Two further points belong to the managing director rather than to the IT provider. The ICO states that these issues cannot be delegated to data scientists or engineering teams, and that senior management are accountable for understanding and addressing them 3. And where an assessment shows a residual high risk that cannot be sufficiently reduced, the ICO must be consulted before the processing starts 3.

A caveat on all of that: the ICO's AI guidance dates from 2023 and 2024 and currently carries a notice that it is under review following the Data (Use and Access) Act 2 3. The data protection by design page, updated in February 2026, is the current one, and it is worth knowing that the ICO says it takes the technical and organisational measures put in place at design stage into account when deciding whether to impose a fine 4. Doing this properly is not only good practice. The regulator has said it counts.

The obligation that may already bind your surveyors

If any part of your business is regulated by the Royal Institution of Chartered Surveyors (RICS), or employs RICS members, a mandatory professional standard has been in force since 9 March 2026 1. How much of a given contractor that covers varies, and it is worth establishing rather than assuming.

The RICS standard on responsible use of artificial intelligence in surveying practice applies where an AI output has a material impact on the delivery of a surveying service, and its requirements are stated as musts 1. Regulated firms must create and operate a risk register, reviewed at least quarterly. They must carry out detailed due diligence before procuring such a system from a third party. They must tell clients in writing, and in advance, when and for what purpose AI is to be used. They must undertake randomised dip samples of outputs at regular intervals. And they must document their decision about the reliability of an output in writing, prepared by or under the supervision of an appropriately qualified and named surveyor who accepts responsibility for its use 1.

RICS states that in regulatory or disciplinary proceedings it will take relevant professional standards into account when deciding whether a member or regulated firm acted appropriately and with reasonable competence 1.

The boundary matters and should not be blurred. This binds RICS members and RICS regulated firms, and only where the output has a material impact on the surveying service 1. It is not a construction wide duty and it does not reach a main contractor's non-surveying functions. But the practical consequence is worth sitting with: your quantity surveyors may already be under a harder, written, enforceable obligation than the rest of the business, and where the firm itself is RICS regulated the obligation runs to the firm as well as to the individual 1.

Before day one

Five decisions, all written down, none of which requires a technologist.

Pick one task. Not a department and not a tool. A pilot of "AI in the commercial team" cannot succeed or fail, because there is no statement of what success would look like. A pilot of "drafting the monthly progress report from the month's site records" can.

Take the baseline. Two ordinary weeks of how long that task takes now, and the timestamps you already hold. Do it before anything is switched on, because afterwards the before is gone and what replaces it is a reconstruction. The method sits in the guide on calculating return on construction AI.

Write the data rule. One page, in plain language, saying what may go into the tool and what may not. The National Cyber Security Centre's advice on public large language models is not to include sensitive information in queries, and not to submit anything that would cause a problem if it were made public 11. That second test is the one to give to site staff, because it needs no technical understanding at all. Note that the NCSC post dates from 2023 and describes public services rather than the commercial tiers most businesses now buy, so treat it as the floor.

Name the reviewer. One person, by name, who checks the output before it is relied on and who is accountable for it. Where the RICS standard applies, that reviewer must be an appropriately qualified and named surveyor and their reliability decision must be in writing 1.

Write the stop condition. A sentence saying what would end the pilot immediately: a wrong figure reaching a client, personal data going somewhere it should not, or the reviewer finding errors above an agreed rate. The NCSC's guidance is that the inevitability of security incidents should be reflected in incident response, escalation and remediation plans 8, and a 30 day pilot is not too small to need one. The value of writing it in week zero is that nobody has to make the decision under pressure in week three.

One legal refinement while you are writing. The ICO advises separating the research and development phase of an AI system from the deployment phase, and identifying a purpose and a lawful basis for each distinct processing operation 5. A pilot is a genuinely different activity from a rollout, and treating it as one is not paperwork for its own sake, it is how you avoid claiming a basis for the pilot that will not stretch to what comes after.

The four weeks

Week one, run it in parallel. The task gets done the old way and the new way at the same time. This costs a week of duplication and buys the only clean comparison you will ever get. Tell the people involved what is happening and why, before you start rather than afterwards: where personal data is collected directly from individuals, the ICO requires privacy information to be provided at the time of collection, including purposes, retention periods and who it will be shared with 6.

Week two, build the test set. Take real examples of the task from your own work, not the vendor's demonstration material. Anthropic's guidance on evaluating output is to design tests that mirror the real world distribution including edge cases, and to prioritise volume over perfection, on the basis that more examples graded roughly beats a handful graded immaculately 14. Fifty real progress reports scored quickly is a better test than five scored beautifully.

Week three, score it honestly. This is where most pilots quietly fail, because "it seems good" is not a finding. The ICO's warning about headline accuracy figures is the clearest available statement of the trap: if 90 per cent of the emails arriving in an inbox are spam, a classifier that labels everything as spam is 90 per cent accurate and completely useless 7. Score what matters, which is usually the specific fields or judgements that would cause trouble if wrong, and count those errors separately from cosmetic ones.

Week four, decide. Three outcomes are legitimate and one of them is stopping. The ICO's position is that not all AI systems demonstrate a sufficient level of statistical accuracy to justify their use 7, and a pilot that concludes this one does not is a successful pilot, not a failed one. It cost a month and it saved a rollout.

What to log, and why it is not surveillance

Log the prompts and the outputs for the duration. The NCSC recommends monitoring and logging inputs to an AI system, such as queries or prompts, to enable compliance obligations, audit, investigation and remediation, and measuring outputs and performance so that changes in behaviour can be observed 9.

Staff sometimes read that as monitoring them, and it is worth heading off directly, because the log serves them: it is what allows a wrong output to be traced back and corrected rather than argued about. Note that the security guidance says nothing about employment monitoring or about what staff must be told 9; that is governed by data protection law, and the ICO's requirement to provide privacy information at the point of collection is the rule that applies 6. Say what is being logged, say why, and say who can see it. The alternative, a pilot with no record of what was asked or answered, leaves you unable to investigate the one incident the whole exercise was supposed to catch.

There is a name for the pack this produces. The government describes assurance as measuring, evaluating and communicating the trustworthiness of AI systems 13, and what a well-run pilot generates is exactly that: a baseline, a scored test set, a written reliability decision and a log. That pack is worth more than the pilot's result, and it is what you have to hand if a client asks how the output was checked. It is not a certification: the government's material is descriptive and offers no accreditation a contractor can obtain 13.

Where a 30-day pilot stops working

Three limits, stated plainly.

A month is long enough to test quality and short enough to hold attention, but it is not long enough to test adoption. Whether people still use the thing in month seven is a different question, and the pilot cannot answer it. Nor can it tell you about model drift: the ICO advises monitoring after deployment at a frequency proportional to the impact an incorrect output may have, and documenting thresholds for when a model needs retraining 7, and none of that is visible in four weeks.

A pilot on one task tells you very little about a different task. Extraction from delivery tickets working well is not evidence that contract review will, and the temptation to generalise from one good result is how a successful pilot becomes an unsuccessful programme.

And there is a category the pilot must not treat as solved. The NCSC notes that prompt injection, where an attacker crafts input designed to make a model behave unintendedly, is one of the most widely reported weaknesses in large language models, and that these systems can present incorrect statements as fact 10. A pilot on your own clean material does not test either. If the system will later read documents that arrive from outside the business, that is a materially different risk and needs its own assessment.

Finally, a note on the limits of this research rather than on the state of the world. No AI specific guidance from the Health and Safety Executive or the Building Safety Regulator was found while preparing this guide, and no construction sector AI regulator guidance was found either. That is a statement about what was looked for and not found on 7 August 2026, not a guarantee that none exists. If someone tells you a pilot must be run a particular way to be compliant in construction, the reasonable question is which instrument or standard says so, and the honest answer for most contractors will be UK data protection law 3 and, for regulated surveying work, the RICS standard 1.

What to do next

Choose the task this week, take the baseline next week, and do not buy anything until both are done. The order matters more than the tool, because a baseline taken after the purchase is worth nothing and a purchase made before the scope is written usually buys the wrong thing.

If your commercial team is RICS regulated, read the standard first 1. It is 18 pages, it is mandatory for them, and it will shape the pilot's governance more than anything else on this page. The related questions of what data may safely go into these tools, covered in uploading construction documents to AI, and which processes are worth piloting at all, covered in construction processes you can automate today, are worth reading before you commit the month.

An earlier and shorter version of this argument, written before the RICS standard came into force, remains on the blog as what a good AI pilot looks like; it keeps the six week shape and the workflow selection, and this guide is the canonical treatment of the governance. Preparing a team to run one is what our AI training for construction teams is for, and our case studies describe pilots run to this structure.

Sources

  1. 1.Royal Institution of Chartered Surveyors, Responsible use of artificial intelligence in surveying practice, 1st editionAuthorityAccessed
  2. 2.Information Commissioner's Office, When do we need to do a DPIA?PrimaryAccessed
  3. 3.Information Commissioner's Office, What are the accountability and governance implications of AI?PrimaryAccessed
  4. 4.Information Commissioner's Office, Data protection by design and by defaultPrimaryAccessed
  5. 5.Information Commissioner's Office, How do we ensure lawfulness in AI?PrimaryAccessed
  6. 6.Information Commissioner's Office, How do we ensure transparency in AI?PrimaryAccessed
  7. 7.Information Commissioner's Office, What do we need to know about accuracy and statistical accuracy?PrimaryAccessed
  8. 8.National Cyber Security Centre, Guidelines for secure AI system development: secure deploymentPrimaryAccessed
  9. 9.National Cyber Security Centre, Guidelines for secure AI system development: secure operation and maintenancePrimaryAccessed
  10. 10.National Cyber Security Centre, AI and cyber security: what you need to knowPrimaryAccessed
  11. 11.National Cyber Security Centre, ChatGPT and large language models: what's the risk?AuthorityAccessed
  12. 12.Department for Science, Innovation and Technology, A pro-innovation approach to AI regulationPrimaryAccessed
  13. 13.Department for Science, Innovation and Technology, Introduction to AI assurancePrimaryAccessed
  14. 14.Anthropic, Define success criteria and build evaluationsAuthorityAccessed
  15. 15.The National Archives, legislation.gov.uk, The Data Protection Act 2018 (Code of Practice on Artificial Intelligence and Automated Decision-Making) Regulations 2026, SI 2026/425PrimaryAccessed

Published , last reviewed . This guide explains general principles and is not legal, contractual or safety advice. The position on any project depends on the contract signed and the facts of that project.

If this is a problem you are carrying on a live package and you want to talk about what fixing it would take, get in touch or book a call.