← Home AI in 15

AI in 15 — September 09, 2026

September 9, 2026 · 16m 04s
Kate

Ten thousand agents. Eighty-eight hours. A ninety-year-old problem with a million dollar bounty on it, and a proof that almost nobody outside OpenAI has actually read.

Kate

Welcome to AI in 15 for Wednesday, September 9th, 2026. I'm Kate, your host.

Marcus

And I'm Marcus, your co-host.

Kate

Today: OpenAI claims one of the seven Millennium Prize problems, and an NYU mathematician says they didn't just beat him, they ran him off the road.

Kate

Meta ships a personal agent with your credit card in its pocket.

Kate

The NSA, CISA and the FBI name six Chinese AI companies, and put model versions on the accusation.

Kate

Plus an Anthropic researcher quits the industry entirely, DeepMind maps every possible letter change in your DNA, and somebody runs a two point eight trillion parameter model off four external hard drives.

Kate

Marcus, Navier-Stokes. Start with what was actually claimed.

Marcus

Tuesday, OpenAI announced a solution to the Navier-Stokes existence and smoothness problem. One of seven Millennium Prize problems, a million dollars attached to each. The claimed result is that three-dimensional fluid flow can develop a singularity in finite time. In plain language, the standard equations describing how fluids move can break down and stop describing anything physical. Mathematicians have argued about that for roughly ninety years.

Kate

And the machine that did it?

Marcus

An unreleased internal model OpenAI describes as significantly more capable than GPT-6 Astra, which itself only went public about a week ago. Roughly ten thousand concurrent agents, eighty-eight hours, then about seventeen hours of formalization. Cost in the millions of dollars.

Kate

So where's the proof?

Marcus

That's the whole story. They briefed the result on a press call without publishing the full mathematical argument. Tristan Buckmaster, the NYU Courant mathematician closest to this problem, confirmed he hasn't reviewed it either. A claim is not a result until somebody can check it. And the comparison is unflattering, because Buckmaster and his collaborator Levent Alpöge released Lean-formalized proofs of related results. Machine-checkable artifacts. Anyone can run them. OpenAI has offered a description.

Kate

Then there's the credit fight, and it gets ugly.

Marcus

Buckmaster and Alpöge had been quietly working the problem since around August 15th, using models from both Anthropic and OpenAI. OpenAI launched its own effort on September 1st, which Buckmaster says was after rumors circulated that somebody was closing in. He announced preliminary findings on September 8th. OpenAI published a full proof the same day. He says both teams took the same unusual route, and told TechCrunch it is not a direction you arrive at in a few days.

Kate

And the phone call.

Marcus

He alleges OpenAI's Sebastien Bubeck pressured him on a September 6th call, asking why he'd want to ruin his career, and later saying that if Buckmaster didn't want him to be nice, he didn't have to be. OpenAI denies seeing the work. Chief Research Officer Mark Chen says no people and no AI systems searched user data, and the company argues the proofs differ substantially.

Kate

Terence Tao weighed in, and I thought his point was the sharpest thing anyone said all day.

Marcus

It was. Tao argues that good open problems have become something like a non-renewable resource. Once a problem is publicly solved, it's contaminated as an evaluation benchmark, and its value is permanently reduced. Worse, the mere rumor that somebody is working on a problem can now trigger a massive automated effort to flatten it before the original research matures. He compares it to bringing excavators to an archaeological site. You get the treasure and you destroy the context that gave it meaning. He floats declaring some classes of problems off-limits to automated solvers.

Kate

Marcus, two things happened here. Separate them for me.

Marcus

Gladly, because blurring them helps nobody. If the proof holds, a model trained for under two weeks did in eighty-eight hours what ninety years of human effort could not. That's a capability datapoint that reframes every timeline argument on the table. But the evidence standard has not been met. And the behavior around the announcement is what mathematicians are actually reacting to. Priority, publication, credit. Those norms are what make research verifiable, and they're being tested by an actor with effectively unlimited compute and a product launch calendar.

Kate

Meta launched Muse. It sends your email and it spends your money.

Marcus

A persistent personal agent. Sends email, books travel, fills forms, negotiates bills, completes checkout on your behalf. It connects to email, calendar, payments, health and fitness apps, smart home devices and shopping platforms, each connection opt-in individually. You name it, give it an avatar, tune how it talks. US-only at launch, free tier plus twenty dollars a month and a hundred dollars a month.

Kate

What's the security story? Because that's the only question I have.

Marcus

Better than I expected, honestly. Muse runs inside its own secure virtual machine with its own browser. A separate agent called Sentinel runs on the same machine but isolated at the system level, and it has to approve anything Muse sends to the internet. Meta says Muse can't see passwords or payment methods, using Stripe Link, with 1Password integration planned. Conversations and machine data reportedly don't feed the ad systems. There's a three hundred thousand dollar bug bounty, with up to a hundred and thirty thousand specifically for prompt injection.

Kate

That is a serious architecture.

Marcus

It's the right shape. An isolated environment with a gatekeeper is genuinely more thoughtful than what competitors have shipped. But Reuters reported it went out despite internal concerns that the product mismanages access to sensitive personal data. And the failure mode here isn't a bad answer, it's a purchase you didn't authorize.

Kate

TechCrunch went straight for the trust question.

Marcus

And listed the receipts. The 2011 FTC settlement, the five billion dollar penalty in 2019, passwords stored in readable form, Cambridge Analytica, and an eighteen billion dollar multistate settlement over harm to children finalized just weeks ago. The Hacker News reaction was blunter. This is the company that once gave a language model the ability to reset user passwords.

Kate

Next. A joint advisory from the NSA, CISA and the FBI, and it names names.

Marcus

Six Chinese firms: DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z dot AI. The accusation is systematic knowledge distillation against US frontier models since at least late 2024, likely with Chinese government awareness. Distillation means querying a stronger model at enormous scale and using its outputs, especially its chain-of-thought reasoning traces, as training data for your own. Billions of tokens across millions of exchanges, routed through multiple pathways to evade terms-of-service enforcement.

Kate

How specific does it get?

Marcus

Remarkably specific. The advisory says that between late 2024 and mid-2025, DeepSeek distilled from Claude 3.7, Claude Sonnet 4, Claude Sonnet 4.5, Claude Opus 4.1, two Gemini 2.5 preview models, GPT-4, GPT-4o and its smaller variants, GPT-5, and Grok 4, to train R1 and V3. Targeting reasoning capability specifically. And the agencies state directly that DeepSeek's widely quoted five point six million dollar training cost is misleading, because it excludes the data obtained this way.

Kate

That number moved markets.

Marcus

In January 2025 it moved them violently. This is the first time the US government has put names and model versions on something labs have muttered about privately for two years. I'd note that an advisory is an assertion, not published evidence, and the specificity is presumably coming from signals intelligence nobody's going to show us.

Kate

What's the practical effect?

Marcus

It hands US labs cover for tighter API gating, harder rate limits and redacting reasoning traces. Which raises the floor on what open access to a frontier model looks like for everyone, including researchers who've never distilled anything.

Kate

A researcher walked out of Anthropic this week and out of the field entirely.

Marcus

Jacob Coxon, twenty-seven, a pretraining researcher, three years across OpenAI and then Anthropic. He told the Wall Street Journal that neither company is acting responsibly, that they're racing straight to self-improving superintelligence, and gambling with our lives. He says the field is tracking the most aggressive scenarios, and by the end of next year things could already be out of control.

Kate

He joined Anthropic for the safety reputation.

Marcus

He did, and he says its efforts are sincere. His argument is that competitive pressure makes the trade-offs unavoidable, and the problem is now bigger than any single lab. Hacker News ran heavily skeptical, with the top replies asking what the worst thing a language model has actually done is.

Kate

Why does this land differently than the usual departure?

Marcus

Timing. Anthropic is in a pre-IPO window, it's committed at least a hundred and thirty-five billion dollars to compute this year, and its entire differentiation is being the safety-first lab. A pretraining researcher walking out and saying that positioning can't survive the race is a costly signal. Treat his timeline as one person's forecast, not a finding. The resignation is the fact.

Kate

DeepMind mapped every possible single-letter change in the human genome. All of them.

Marcus

AlphaGenome Atlas. A precomputed prediction for all nine billion possible single-nucleotide variants, with thousands of molecular effect predictions per variant across multiple cell types and tissues. One petabyte, roughly thirty times the AlphaFold Database. Free through a web portal for academic research, plus an API and Google Cloud.

Kate

Is it any good?

Marcus

The headline validation number is twenty-two percent more genetic associations found in UK Biobank data when you group variants by predicted molecular effect rather than conventionally. That's real data, not a synthetic benchmark. Collaborators including the Broad Institute and a rare disease consortium have experimentally checked predictions. Researchers did push back usefully. The benchmark labels partly derive from evolutionary conservation, so it's hard to separate genuine prediction from re-reading the prior.

Kate

Give me the honest framing.

Marcus

This is the AlphaFold playbook applied to regulatory genetics, the non-coding DNA that decides when genes switch on. It's a resource, not a result. Its value depends entirely on whether wet-lab groups find the predictions hold. DeepMind states plainly it isn't validated or approved for any clinical use, and that caveat should travel with every headline about it.

Kate

Money, briefly, because we did Mistral yesterday.

Marcus

We did, so just the new one. Cognition raised over two billion dollars at a forty-eight billion valuation, led by Andreessen Horowitz. That's up from twenty-six billion four months ago. Run-rate revenue went from four hundred ninety-two million in May to roughly nine hundred million now, with analysts projecting four to five billion by year-end.

Kate

And the burn?

Marcus

Leased Nvidia infrastructure in the hundreds of millions annually, total 2026 cash burn potentially reaching eight hundred million. They're building their own models to cut dependence on third-party providers, exactly like Cursor did before selling to SpaceX for sixty billion in April.

Kate

So what does the price actually say?

Marcus

That investors don't believe AI coding is winner-take-all, even after watching a competitor exit at sixty billion. That's a genuine bet against consolidation, and it's the opposite of what most people assumed a year ago.

Kate

Last one, and it's my favorite. Somebody ran a two point eight trillion parameter model on a laptop.

Marcus

It's called Deltafin. A Rust binary running the full uncompressed Kimi K3, no quantization, no pruning, all sixteen routed experts processing every token. It works by keeping one and a half terabytes of expert weights on four external SSDs and streaming them on demand into a MacBook Pro with a hundred and twenty-eight gigabytes of RAM.

Kate

How fast?

Marcus

About one token per second. Time to first response, roughly six minutes on a short prompt. One commenter called it a medium-length prompt in only eleven days.

Kate

So it's useless.

Marcus

Completely useless, and still interesting. It's proof that the memory wall is an engineering constraint rather than a hard physical one, and it puts an actual number on what full-quality frontier inference costs outside a data center.

Kate

Anything else from the technical corner?

Marcus

Inception Labs shipped Mercury 2.5, a diffusion-based language model rather than an autoregressive one. Not open weights, which disappointed people. Testers say it's nowhere near frontier quality, and the company doesn't claim it is. But diffusion decoding attacks latency in a way that scaling autoregressive models simply cannot, so long-term that's the more consequential bet of the two.

Kate

One to watch tomorrow: whether OpenAI publishes the Navier-Stokes proof in checkable form. A Lean formalization would settle the capability question in a single day.

Marcus

And continued silence answers a different question just as clearly.

Kate

That's your AI in 15 for today. See you tomorrow.