AI in 15 — September 10, 2026
Three hundred billion tokens. Twenty-two and a half million dollars of compute. And an accusation that the hardest part wasn't the maths — it was finding out what somebody else had already figured out.
Welcome to AI in 15 for Thursday, September 10th, 2026. I'm Kate, your host.
And I'm Marcus, your co-host.
Today: the Navier-Stokes fight escalates, with a named allegation about authorship and a second mathematician saying it happened to him too.
All three American frontier labs ship cyber-specific models in the same week, and one of them scores a hundred percent on an exploitation benchmark.
Anthropic models three economic futures, and capital takes a bigger share in every single one.
Plus what GPT-6 Astra's architecture actually is, a coding agent that ran a real quantum experiment, and a browser game about turning one button blue.
Marcus, we covered the Navier-Stokes claim yesterday. What's genuinely new?
Two things, and both are about provenance rather than mathematics. First, Buckmaster has now put a specific allegation on the record with TechCrunch. He says OpenAI's Sébastien Bubeck offered him sole authorship on the Navier-Stokes result, on two conditions: that Levent Alpöge's name come off the paper, and that the write-up credit an internal OpenAI model with resolving the problem.
Sole authorship in exchange for erasing your collaborator.
That's the allegation. Alpöge, worth noting, works at Anthropic. Bubeck's response concedes something important — he says the team was prompted to attack the problem after hearing a rumour that Buckmaster and Alpöge were making progress. He maintains the model solved the Euler piece by a completely different route, but acknowledges the full Navier-Stokes path resembles theirs. OpenAI's formal line is that they did not see any of the work through any means until it was public.
And the second thing?
Overnight, a mathematician named Valerio Capraro posted that a proof of his may have been front-run the same way. He also says a researcher who opted out of training data on June 29th was told by OpenAI that training on his material did not happen, and he believes it did. That thread was still climbing Hacker News this morning. It is unconfirmed and I'd treat it as unconfirmed. But one allegation is a dispute. Two starts to look like a pattern somebody should check.
The numbers OpenAI has released are staggering on their own.
Three hundred billion tokens, roughly twenty-two and a half million dollars of compute, eighty-eight hours, starting September 1st. And that's the part I keep circling. Columbia's Michael Harris told Science he doesn't expect anyone to keep spending millions on mathematics that has no profit attached. If closing a Millennium Prize problem requires compute that four or five companies possess, the open collaboration mathematics actually runs on stops being the mechanism of progress.
Is there any institutional response yet?
There is, and it's fast. The Fields Medalist Jacob Tsimerman announced a Mathematical AI Safety Institute, modelled on Princeton's Institute for Advanced Study. Whether that's the right shape, I don't know. But the field noticed within seventy-two hours that it has no norms for this, and started building one.
Security. All three American labs shipped cyber models this week, and I want the numbers first.
Start with OpenAI, because Astra's numbers are the ones that stop you. A hundred percent on ExploitBench. Multiple zero-days discovered during evaluation. And ninety-one and a half percent of jailbreak attempts refused, against fifty-nine percent for GPT-5.6 Sol. OpenAI states plainly that Astra meets its Critical cybersecurity capability threshold — that's their own tripwire for a model dangerous enough to require restricted release. It's behind a tester programme they're calling Daybreak Blue.
A hundred percent on an exploitation benchmark is a strange thing to publish.
It's an odd flex, and it says as much about the benchmark as the model. Nothing scores a hundred on a hard benchmark for long — it means ExploitBench is now saturated and needs replacing. But set that aside. A model that finds novel zero-days is offensive capability, whatever the marketing says on the box.
What did the other two do?
Google shipped Gemini 3.8 Flash Cyber through the Fairwind programme we covered Tuesday, now with over six hundred and fifty partners including CrowdStrike, Datadog and Palo Alto Networks. Anthropic made Claude Fable 5.1 permissible for vulnerability identification for the first time — that's a policy change, not a model change — and restricted the bigger Mythos 5.1 to trusted-access customers. They also shipped something called Enterprise Frontier Safeguards, pairing zero data retention with monitoring, plus sandbox-escape detection.
Sandbox-escape detection. After the month we've had.
After the month we've had. And here's why all three at once matters. Every lab chose gated access over open release for the same capability class, in the same week. That's the first real test of whether voluntary capability thresholds survive contact with a commercial security market that badly wants to buy this. So far they're holding.
Give me the caveat.
Ninety-one and a half percent refusal means roughly one in twelve jailbreak attempts still lands, on a model the vendor itself rates Critical. That's a big improvement and an uncomfortable residual at the same time. Both are true.
Next. The distillation advisory we covered yesterday got some independent evidence.
It did, and it's suggestive rather than conclusive. A widely shared code gist demonstrated that Qwen 3.8 responds unusually strongly when you inject the first slice of a GPT-5.5 Pro reasoning trace as a prefill — meaning you start its answer with somebody else's thinking and it just carries on comfortably, like it recognises the handwriting.
That sounds damning.
It sounds damning and the commenters killed it fairly. Qwen 3.8's build date is September 2nd, and the paper that published those extracted traces came out August 10th. So the model may simply have trained on the paper, which is public. That's not distillation, that's reading. I flag it because it's circulating as proof, and it isn't.
But the advisory itself hasn't been challenged?
Not substantively, no. Six firms named, eighteen distinct US models allegedly pulled from by Moonshot AI alone. The detail I'd underline today is the tradecraft — many accounts, multiple cloud providers, third-party API aggregators, proxies, grey-market resellers. That's deliberate evasion architecture, not somebody bumping into a rate limit. And the advisory asks providers to coordinate a response, which in practice means American labs sharing signals about who is querying them. The API layer gets less open for everybody after this.
Let's do the Astra architecture story, because I think you've been waiting for it.
I have. The Information reported that GPT-6 Astra uses recurrent depth — looped transformers — and that this obscures the model's reasoning. That set off a week of genuine alarm among safety researchers, on the theory that a model thinking internally in a loop rather than in visible text destroys chain-of-thought monitoring.
And?
Sebastian Raschka published a technical breakdown, now the top AI item on Hacker News, arguing the alarm is aimed at the wrong thing. Looping just means reusing the same transformer blocks several times instead of stacking distinct ones. Twenty-two blocks run twice gives you forty-four effective applications while storing half the weights. It saves parameters, not compute. You still do all forty-four passes.
So it's a storage trick.
Largely. Recent work puts the training-compute saving at somewhere between seven and eighteen percent to reach the same validation loss. And the idea traces back to a 2018 paper on Universal Transformers, so it isn't even new. On the hiding question, Raschka's argument is that shorter reasoning traces track capability rather than architecture — better models backtrack less. OpenAI's chief scientist disputed the report directly, saying the depth of the computation graph for their frontier models, Astra included, is within a factor of two of GPT-4.
So the scare was wrong.
The mechanism was wrong. The concern isn't — Astra's own system card concedes reduced monitorability, we quoted it Monday. What I like about this one is the shape of it. A paywalled report produced a technical claim, that claim got laundered into a safety panic, and the open technical community corrected it inside a week for free. That correction only happens because the architecture literature is public.
Anthropic published three economic futures. Give me the shape.
Three scenarios to 2030, reviewed by Daron Acemoglu and David Autor, which is real credentialing. Modest treats AI as roughly comparable to the internet. Substantial has AI doing about half of knowledge work mostly autonomously. Extreme has it beating humans at most knowledge tasks through recursive self-improvement. US GDP in 2030 lands at thirty-four trillion, thirty-six trillion, or forty-four trillion dollars.
And the number people are actually arguing about?
Labour share. Today roughly sixty cents of every dollar the economy produces goes to wages, forty to capital. That holds near fifty-nine percent in the modest case, slips to fifty-six in the substantial case, and falls to forty-five in the extreme case. Nearly fifteen points moving from wages to capital in the scenario where AI works best.
That's Anthropic saying their own success costs workers fifteen points of national income.
It is a striking thing to publish. Though read the scenario set carefully — the least optimistic option on offer is AI doesn't matter much. There's no scenario where it matters enormously and goes badly. That's a conspicuous gap for a company whose alignment lead said this week he puts catastrophic risk above ten percent.
What's the critique from the floor?
The sharpest one is about their own worked example — a nurse who uses AI to see more patients and spend more time explaining diagnoses. Commenters found that economically naive, and I'd agree. The realistic outcome of a nurse becoming more productive is the same patient load with fewer nurses. Someone also noted that no scenario prices in what inference actually costs once the subsidies end. Treat the whole thing as a framing device, not a forecast.
One more, and this one I enjoyed. A coding agent ran a real quantum experiment.
MIT's Engineering Quantum Systems Group. A graduate student, Beatriz Yankelevich, wired Codex running GPT-5.6 Sol directly into the lab control software for a six-qubit superconducting chip. The agent chose measurement parameters, drove the actual hardware, analysed the returning signals, and fed each result into the next step. Forty target measurements completed, with researchers intervening on only four.
What's the honest read?
The loop is the interesting part. It's not writing about physics, it's running an instrument and deciding what to do next based on what came back. That's a closed loop with hardware. The caveats are real though — it slowed badly and needed an experienced human whenever the signals got weak or noisy, which is exactly when a lab needs judgement. And sceptics pointed out most of this was scriptable in Python back in 2011. Forty measurements on one chip in one lab is a demo, not a capability claim.
Before we close, the internet's favourite thing this week.
A little browser comedy called opusfived dot dev, by a developer named Miloš. You ask an agent to make one button blue, and it turns half the site blue. Top of Hacker News with over a thousand points. The best comment calls it a variable-reward schedule — it's basically gambling. Which, if you've spent a day pair-programming with an agent, is the most accurate description of the experience anyone has written down.
One to watch: DeepSeek's beta endpoint is literally named to expire today. If that's a launch date, a new model with claimed native multimodal support lands within hours — two days after three US agencies accused the company of building on distilled American output.
Watch it, but the quantum result is the one that'll still matter in a year.
That's your AI in 15 for today. See you tomorrow.