AI in 15 — September 09, 2026
Ten thousand agents. Eighty-eight hours. A ninety-year-old problem with a million dollar bounty on it, and a proof that almost nobody outside OpenAI has actually read.
Welcome to AI in 15 for Wednesday, September 9th, 2026. I'm Kate, your host.
And I'm Marcus, your co-host.
Today: OpenAI claims one of the seven Millennium Prize problems, and an NYU mathematician says they didn't just beat him, they ran him off the road.
Meta ships a personal agent with your credit card in its pocket.
The NSA, CISA and the FBI name six Chinese AI companies, and put model versions on the accusation.
Plus an Anthropic researcher quits the industry entirely, DeepMind maps every possible letter change in your DNA, and somebody runs a two point eight trillion parameter model off four external hard drives.
Marcus, Navier-Stokes. Start with what was actually claimed.
Tuesday, OpenAI announced a solution to the Navier-Stokes existence and smoothness problem. One of seven Millennium Prize problems, a million dollars attached to each. The claimed result is that three-dimensional fluid flow can develop a singularity in finite time. In plain language, the standard equations describing how fluids move can break down and stop describing anything physical. Mathematicians have argued about that for roughly ninety years.
And the machine that did it?
An unreleased internal model OpenAI describes as significantly more capable than GPT-6 Astra, which itself only went public about a week ago. Roughly ten thousand concurrent agents, eighty-eight hours, then about seventeen hours of formalization. Cost in the millions of dollars.
So where's the proof?
That's the whole story. They briefed the result on a press call without publishing the full mathematical argument. Tristan Buckmaster, the NYU Courant mathematician closest to this problem, confirmed he hasn't reviewed it either. A claim is not a result until somebody can check it. And the comparison is unflattering, because Buckmaster and his collaborator Levent Alpöge released Lean-formalized proofs of related results. Machine-checkable artifacts. Anyone can run them. OpenAI has offered a description.
Then there's the credit fight, and it gets ugly.
Buckmaster and Alpöge had been quietly working the problem since around August 15th, using models from both Anthropic and OpenAI. OpenAI launched its own effort on September 1st, which Buckmaster says was after rumors circulated that somebody was closing in. He announced preliminary findings on September 8th. OpenAI published a full proof the same day. He says both teams took the same unusual route, and told TechCrunch it is not a direction you arrive at in a few days.
And the phone call.
He alleges OpenAI's Sebastien Bubeck pressured him on a September 6th call, asking why he'd want to ruin his career, and later saying that if Buckmaster didn't want him to be nice, he didn't have to be. OpenAI denies seeing the work. Chief Research Officer Mark Chen says no people and no AI systems searched user data, and the company argues the proofs differ substantially.
Terence Tao weighed in, and I thought his point was the sharpest thing anyone said all day.
It was. Tao argues that good open problems have become something like a non-renewable resource. Once a problem is publicly solved, it's contaminated as an evaluation benchmark, and its value is permanently reduced. Worse, the mere rumor that somebody is working on a problem can now trigger a massive automated effort to flatten it before the original research matures. He compares it to bringing excavators to an archaeological site. You get the treasure and you destroy the context that gave it meaning. He floats declaring some classes of problems off-limits to automated solvers.
Marcus, two things happened here. Separate them for me.
Gladly, because blurring them helps nobody. If the proof holds, a model trained for under two weeks did in eighty-eight hours what ninety years of human effort could not. That's a capability datapoint that reframes every timeline argument on the table. But the evidence standard has not been met. And the behavior around the announcement is what mathematicians are actually reacting to. Priority, publication, credit. Those norms are what make research verifiable, and they're being tested by an actor with effectively unlimited compute and a product launch calendar.
Meta launched Muse. It sends your email and it spends your money.
A persistent personal agent. Sends email, books travel, fills forms, negotiates bills, completes checkout on your behalf. It connects to email, calendar, payments, health and fitness apps, smart home devices and shopping platforms, each connection opt-in individually. You name it, give it an avatar, tune how it talks. US-only at launch, free tier plus twenty dollars a month and a hundred dollars a month.
What's the security story? Because that's the only question I have.
Better than I expected, honestly. Muse runs inside its own secure virtual machine with its own browser. A separate agent called Sentinel runs on the same machine but isolated at the system level, and it has to approve anything Muse sends to the internet. Meta says Muse can't see passwords or payment methods, using Stripe Link, with 1Password integration planned. Conversations and machine data reportedly don't feed the ad systems. There's a three hundred thousand dollar bug bounty, with up to a hundred and thirty thousand specifically for prompt injection.
That is a serious architecture.
It's the right shape. An isolated environment with a gatekeeper is genuinely more thoughtful than what competitors have shipped. But Reuters reported it went out despite internal concerns that the product mismanages access to sensitive personal data. And the failure mode here isn't a bad answer, it's a purchase you didn't authorize.
TechCrunch went straight for the trust question.
And listed the receipts. The 2011 FTC settlement, the five billion dollar penalty in 2019, passwords stored in readable form, Cambridge Analytica, and an eighteen billion dollar multistate settlement over harm to children finalized just weeks ago. The Hacker News reaction was blunter. This is the company that once gave a language model the ability to reset user passwords.
Next. A joint advisory from the NSA, CISA and the FBI, and it names names.
Six Chinese firms: DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z dot AI. The accusation is systematic knowledge distillation against US frontier models since at least late 2024, likely with Chinese government awareness. Distillation means querying a stronger model at enormous scale and using its outputs, especially its chain-of-thought reasoning traces, as training data for your own. Billions of tokens across millions of exchanges, routed through multiple pathways to evade terms-of-service enforcement.
How specific does it get?
Remarkably specific. The advisory says that between late 2024 and mid-2025, DeepSeek distilled from Claude 3.7, Claude Sonnet 4, Claude Sonnet 4.5, Claude Opus 4.1, two Gemini 2.5 preview models, GPT-4, GPT-4o and its smaller variants, GPT-5, and Grok 4, to train R1 and V3. Targeting reasoning capability specifically. And the agencies state directly that DeepSeek's widely quoted five point six million dollar training cost is misleading, because it excludes the data obtained this way.
That number moved markets.
In January 2025 it moved them violently. This is the first time the US government has put names and model versions on something labs have muttered about privately for two years. I'd note that an advisory is an assertion, not published evidence, and the specificity is presumably coming from signals intelligence nobody's going to show us.
What's the practical effect?
It hands US labs cover for tighter API gating, harder rate limits and redacting reasoning traces. Which raises the floor on what open access to a frontier model looks like for everyone, including researchers who've never distilled anything.
A researcher walked out of Anthropic this week and out of the field entirely.
Jacob Coxon, twenty-seven, a pretraining researcher, three years across OpenAI and then Anthropic. He told the Wall Street Journal that neither company is acting responsibly, that they're racing straight to self-improving superintelligence, and gambling with our lives. He says the field is tracking the most aggressive scenarios, and by the end of next year things could already be out of control.
He joined Anthropic for the safety reputation.
He did, and he says its efforts are sincere. His argument is that competitive pressure makes the trade-offs unavoidable, and the problem is now bigger than any single lab. Hacker News ran heavily skeptical, with the top replies asking what the worst thing a language model has actually done is.
Why does this land differently than the usual departure?
Timing. Anthropic is in a pre-IPO window, it's committed at least a hundred and thirty-five billion dollars to compute this year, and its entire differentiation is being the safety-first lab. A pretraining researcher walking out and saying that positioning can't survive the race is a costly signal. Treat his timeline as one person's forecast, not a finding. The resignation is the fact.
DeepMind mapped every possible single-letter change in the human genome. All of them.
AlphaGenome Atlas. A precomputed prediction for all nine billion possible single-nucleotide variants, with thousands of molecular effect predictions per variant across multiple cell types and tissues. One petabyte, roughly thirty times the AlphaFold Database. Free through a web portal for academic research, plus an API and Google Cloud.
Is it any good?
The headline validation number is twenty-two percent more genetic associations found in UK Biobank data when you group variants by predicted molecular effect rather than conventionally. That's real data, not a synthetic benchmark. Collaborators including the Broad Institute and a rare disease consortium have experimentally checked predictions. Researchers did push back usefully. The benchmark labels partly derive from evolutionary conservation, so it's hard to separate genuine prediction from re-reading the prior.
Give me the honest framing.
This is the AlphaFold playbook applied to regulatory genetics, the non-coding DNA that decides when genes switch on. It's a resource, not a result. Its value depends entirely on whether wet-lab groups find the predictions hold. DeepMind states plainly it isn't validated or approved for any clinical use, and that caveat should travel with every headline about it.
Money, briefly, because we did Mistral yesterday.
We did, so just the new one. Cognition raised over two billion dollars at a forty-eight billion valuation, led by Andreessen Horowitz. That's up from twenty-six billion four months ago. Run-rate revenue went from four hundred ninety-two million in May to roughly nine hundred million now, with analysts projecting four to five billion by year-end.
And the burn?
Leased Nvidia infrastructure in the hundreds of millions annually, total 2026 cash burn potentially reaching eight hundred million. They're building their own models to cut dependence on third-party providers, exactly like Cursor did before selling to SpaceX for sixty billion in April.
So what does the price actually say?
That investors don't believe AI coding is winner-take-all, even after watching a competitor exit at sixty billion. That's a genuine bet against consolidation, and it's the opposite of what most people assumed a year ago.
Last one, and it's my favorite. Somebody ran a two point eight trillion parameter model on a laptop.
It's called Deltafin. A Rust binary running the full uncompressed Kimi K3, no quantization, no pruning, all sixteen routed experts processing every token. It works by keeping one and a half terabytes of expert weights on four external SSDs and streaming them on demand into a MacBook Pro with a hundred and twenty-eight gigabytes of RAM.
How fast?
About one token per second. Time to first response, roughly six minutes on a short prompt. One commenter called it a medium-length prompt in only eleven days.
So it's useless.
Completely useless, and still interesting. It's proof that the memory wall is an engineering constraint rather than a hard physical one, and it puts an actual number on what full-quality frontier inference costs outside a data center.
Anything else from the technical corner?
Inception Labs shipped Mercury 2.5, a diffusion-based language model rather than an autoregressive one. Not open weights, which disappointed people. Testers say it's nowhere near frontier quality, and the company doesn't claim it is. But diffusion decoding attacks latency in a way that scaling autoregressive models simply cannot, so long-term that's the more consequential bet of the two.
One to watch tomorrow: whether OpenAI publishes the Navier-Stokes proof in checkable form. A Lean formalization would settle the capability question in a single day.
And continued silence answers a different question just as clearly.
That's your AI in 15 for today. See you tomorrow.