Skip to content
AI Metric

AI updates

What shipped, and where it lands on site

There is no shortage of people telling you what AI released this week. There is a shortage of people who can tell you which package it touches, who has to change what they do, and whether it is worth the disruption on a live job.

That is what this is. Each edition takes the releases that matter, says plainly what shipped, and then does the part everyone skips: where it actually lands in construction delivery, and one concrete thing worth doing about it on Monday morning.

Three items, maximum

Not a link dump. If only two are worth your time that week, you get two.

Written from delivery

Twenty years of Tier 1 packages, with depth in facades and envelope.

Everything credited

Summarised in our own words, publisher named, original linked. The value is the translation.

The cheap model won on the work that pays

Two vendors shipped on the same day and both went down in price, and on a business workflow test the cheapest model scored highest. Search moved inside the model and agents got further, but still need checking.

  • GPT-6 Sol and Luna land at half the price, and Sol tops the workflow test
  • Claude Opus 5.5 cuts its own cost by about 40 per cent on typical work
  • Caching quietly became the biggest lever on the bill
  • Search moved inside the model’s own loop, and MCP is how it reaches your systems
  • Agents score 59 per cent on real professional tasks: useful, and checked
  • The benchmark tables are not comparable across vendors, so test on your own work
Read this edition

It stopped going past what it was asked

GPT-6 Astra uses a computer well enough to be useful and, on its publisher’s own test, stopped exceeding its brief. For anyone deciding what an agent may touch, that is the number that matters.

  • Astra goes beyond its authorised target in none of the cases its predecessor failed
  • The same workload, at a higher score, in about 47 per cent less time
  • A million token window, and a model that can search its own toolbox
Read this edition

Speed just went on the rate card

OpenAI now charges extra for a faster answer, cut the price of its cheapest model by 80 per cent, and bought a way for the work to carry on after the laptop shuts.

  • The work stops living on one person's laptop
  • You can now pay extra for a faster answer
  • The cheap model got 80 per cent cheaper
Read this edition

Same score, eight times the price

Four frontier models finished the week two index points apart and eight times apart on price. What that means for anyone budgeting AI on a construction package.

  • Grok 4.6 reaches the frontier at a third of the price, with a long-context cliff
  • Grok Bot and Muse Glimmer put agents on a computer and on a laptop
  • Claude text watermarking makes provenance a build-time question
Read this edition

Construction finally has a number to compare itself against

Free industry benchmarks, a lower price per document, and agents that keep working with the laptop shut. Three releases that change what construction admin costs.

  • Buildots Intelligence Lab publishes free schedule-adherence benchmarks
  • Gemini 3.6 Flash cuts the price of reading the whole project record
  • Claude Cowork runs scheduled tasks with the laptop closed
Read this edition

Want the construction read on something specific?

If there is a release you are being asked about internally and you want an honest answer on whether it changes anything for your packages, ask. It is a better use of a call than a demo.

Book a 30 minute call
Prefer me to call you? Leave your number

Only so there is a way back to you if the phone does not connect. We will ring first either way.

How we handle your details