AI Reality Check — Issue 22
Claude proved Fermat's Last Theorem with UK verification, London surgeons used live AI to save a patient's sight, GPT-6 Astra launched after a safety-driven delay, IBM's Watson Health remains a warning, and AI revenue concentrates in a few heavy users.
#AIRealityCheck | Friday 11 September 2026 | Newsletter Issue 22
━━━━━━━━━━━━━━━━━━━━━━
1. Claude Proved Fermat's Last Theorem, Verified in the UK
What happened: Anthropic says Claude formally proved Fermat's Last Theorem, a problem that stumped mathematicians for 358 years, producing what the company describes as the longest math proof ever made, a chain a computer can verify step by step. Dozens of Claude agents worked mostly independently over 11 days, proving more than 30,000 supporting theorems and producing 13 million lines of code, more than five times the size of Mathlib, the shared library mathematicians already rely on. The agents initially lost track of their own work, until a tool called Prove2Me gave each one a live to-do list to avoid duplication. Kevin Buzzard, a mathematician at Imperial College London, reviewed the proof and confirmed it holds up.
Why it matters: this is a genuine landmark in AI-assisted mathematics, and it happened with UK verification at its centre. It also shows what this kind of achievement actually looks like in practice: not a single flash of insight, but extraordinary scale, coordination tooling, and a human expert still required to confirm the result holds.
Who it's for: anyone in mathematics, academia, or research. Anyone tracking how AI capability in specialised technical domains is actually developing. Anyone interested in UK institutions' role in verifying frontier AI claims.
━━━━━━━━━━━━━━━━━━━━━━
2. London Surgeons Used Live AI to Save a Patient's Sight
What happened: surgeons at London's National Hospital for Neurology and Neurosurgery used a live AI system to colour-code critical anatomy during an 11mm brain tumour resection, saving 48-year-old Rhys Hibbert's sight completely.
Why it matters: this is a concrete, successful, high-stakes use of AI in UK healthcare, guiding a genuinely delicate surgical procedure in real time. It stands as a clear, positive counterpoint to the caution required elsewhere in this issue.
Who it's for: anyone in healthcare or medical technology. Anyone following how AI is being deployed in UK hospitals specifically. Anyone who wants a concrete example of AI delivering real, measurable benefit in a high-stakes setting.
━━━━━━━━━━━━━━━━━━━━━━
3. GPT-6 Astra Launched, But Reached "Critical" Cybersecurity Capability
What happened: OpenAI released GPT-6 Astra to ChatGPT users on Pro, Enterprise, and Business Premium plans, with Plus and Business users gaining access over the following days. During testing, Astra reached what OpenAI internally calls "Critical" cybersecurity capability, meaning it could find and exploit previously unknown security flaws, prompting OpenAI to delay parts of development and add stricter safeguards before launch.
Why it matters: this is one of the clearest public signals yet of an AI lab treating its own model's capability as something requiring a deliberate safety pause before release, rather than shipping on the original timeline regardless.
Who it's for: anyone using ChatGPT for business. Cybersecurity professionals. Anyone following how AI labs handle capability that crosses into genuinely sensitive territory.
━━━━━━━━━━━━━━━━━━━━━━
4. IBM's $4 Billion AI Doctor Failure Is a Live Warning, Not Ancient History
What happened: IBM's Watson beat Jeopardy champions in 2011 and was then pointed at curing cancer. By 2022, after spending nearly $4 billion, IBM sold Watson Health's remaining assets for roughly $1 billion, a quarter of what it spent. Danish doctors had agreed with Watson's recommendations only 33% of the time, and the system had recommended dangerously wrong treatments in multiple documented cases.
Why it matters: the underlying myth, that impressive capability on one task transfers reliably to other high-stakes domains, is still very much alive in how AI tools are marketed and evaluated today. Watson Health remains one of the clearest, best-documented cautionary tales available.
Who it's for: any business evaluating AI vendor claims. Anyone in healthcare, professional services, or any regulated field considering AI tools for high-stakes decisions.
━━━━━━━━━━━━━━━━━━━━━━
5. AI Revenue Is Concentrating in the Top 1% of Customers
What happened: new data from Ramp shows the top 1% of customers drive 80% of enterprise revenue at both OpenAI and Anthropic. At Anthropic, coding tools Cursor and GitHub Copilot alone drove around $1.2 billion of its $5 billion revenue last year.
Why it matters: AI value in business is not spreading evenly. It is concentrating heavily in specific, intensive, usually technical use cases, a pattern worth understanding before assuming broad AI adoption will deliver broad returns.
Who it's for: business leaders evaluating AI investment. Anyone trying to understand where AI genuinely earns its cost versus where it's adopted more for appearance than results.
━━━━━━━━━━━━━━━━━━━━━━
Also this week: Meta's Project OT, an attempt to replace 60% of its workforce with AI, collapsed after morale dropped and only 36% of the surging AI output proved usable. Anthropic's new Model Hardware Standard let Claude connect to and adjust a laser. Claude's Fable 5.1 got roughly 25% cheaper for normal use.
━━━━━━━━━━━━━━━━━━━━━━
This week in one sentence: Claude proved Fermat's Last Theorem with UK verification, London surgeons used live AI to save a patient's sight, GPT-6 Astra launched after a safety-driven delay, IBM's Watson Health failure remains a live warning, and AI revenue keeps concentrating in a small number of intensive users.
Born analogue. Raised digital. 30 years of real business experience explaining what AI actually means for work.
— Kaye Nicholson | GrowthZone AI | growthzoneai.co.uk
Source: View on LinkedIn
AI Reality Check is published weekly. Subscribe below to stay updated.

Written by
Kaye Nicholson
Founder, GrowthZone AI · Bdaily Columnist
Kaye Nicholson is the founder of GrowthZone AI and a columnist for Bdaily, helping businesses, charities, founders and teams use AI in simple, practical ways without jargon or overwhelm.
Book a short AI chatFound this helpful? Share it:
Reader feedback
Got a view on this week's AI Reality Check? Join the conversation on LinkedIn or send your thoughts directly.
Get the AI Reality Check weekly newsletter
Weekly practical AI updates for UK businesses: what changed, what it means, and what to watch next.
