AI in 15 — September 22, 2026
More than one hundred long-standing open problems in mathematics, resolved. That's the claim. The model that supposedly did it started training on August the twenty-eighth. Nobody outside OpenAI has seen a single proof.
Welcome to AI in 15 for Tuesday, September 22nd, 2026. I'm Kate, your host.
And I'm Marcus, your co-host.
Today: OpenAI announces an extraordinary mathematical claim, and in the same breath hands the whistle to nine mathematicians who say they have no power at all.
Grok 4.7 ships with forty percent more weights and the same price tag. The independent benchmarks are less flattering than the press release.
Amazon blocks Meta's shopping agent, and the agentic commerce cold war gets hot.
Plus SoftBank launches eleven billion dollars of junk bonds to buy more OpenAI, and an eighteen-thousand-dollar Mac that makes a serious case for keeping your agents at home.
Marcus, start with what OpenAI actually announced, because there are two things in one post.
There are, and the packaging is the story. Thing one: a new independent Advisory Group on Mathematics and Artificial Intelligence, hosted at Princeton's Institute for Advanced Study. Nine members, including Timothy Gowers, Edward Witten and Melanie Matchett Wood. Unpaid, and explicit that they work independently of OpenAI. Thing two, buried inside the same announcement: an internal OpenAI model has, the company says, resolved more than a hundred long-standing open problems across most areas of mathematics.
On top of the Navier–Stokes result from earlier this month.
Right — the Millennium Prize problem, which Quanta described as a singularity found in the three-dimensional equations, produced by a swarm of roughly ten thousand autonomous agents. Training on the new model began August twenty-eighth. That's three and a half weeks.
And the advisory group's own language is doing something interesting.
It's doing enormous work. Their statement says they face the challenge of advising OpenAI on how to coordinate the release of a large number of significant results "that they report have been produced by their internal model." They report. Eminent mathematicians choosing that construction are telling you precisely what they have and haven't verified.
What can this group actually do?
Very little, by design, and both sides say so plainly. OpenAI's post: the group will not be responsible for advising on how OpenAI paces its internal mathematical progress. The IAS side put it even more bluntly — we do not have decision making power at any AI company, and responsibility for company decisions rests with the company.
So Marcus, is this responsible disclosure or borrowed credibility?
Honest answer today is we can't tell, and I'd distrust anyone who claims otherwise. But note the context. Earlier this month twenty-five Fields Medal winners signed an open letter saying labs racing to claim famous open problems are damaging the field's norms. Exactly one member of this advisory group — Camillo De Lellis — also signed that letter. Burt Totaro, in the comment thread on Terence Tao's blog, put the uncharitable read directly: OpenAI has had bad publicity, and so they're trying to exploit the trust and respect these mathematicians command.
And the charitable read?
That if you genuinely have a hundred results and no idea how to release them without wrecking a discipline, asking the discipline is the correct move.
Here's what I keep coming back to. We can't check any of it.
Today, no. The model is internal and unreleased, the proofs aren't published. One Hacker News commenter cut straight to it — the only thing I want to see is the problem statements, the solutions, and the assessments. But mathematics is the one field with a bright line. A proof checks out or it doesn't. Within months we'll know, and that's more than you can say for almost any frontier-lab claim.
There was one more observation from that thread you flagged to me.
From a commenter called Certhas, and it's the one that stuck. Mathematicians got an advisory board because they hold cultural capital inside AI labs. Most professions facing automation will get no such thing.
Model news. Grok 4.7 shipped yesterday, about two weeks late.
New base model, two point one trillion parameters — up from one point five trillion in 4.6, so forty percent bigger. Longer reinforcement learning run, training weighted toward multi-hour tasks, better self-verification, five hundred thousand token context. And the price held flat: two dollars per million input tokens, six per million output. xAI ate the margin.
Generous. Or nervous?
Someone on Hacker News read it the second way — forty percent more weights, flat pricing, two-week delay, and concluded xAI must not have been happy with the results. The independent numbers are mixed. On Artificial Analysis's Intelligence Index it scores forty-six, up two points from 4.6, against fifty-three each for Claude Fable 5.1 and GPT-6. That gap hasn't closed.
So where is it actually good?
Long-horizon agentic work, and genuinely so. Plus one hundred and eleven Elo over 4.6 on their agentic briefcase benchmark, landing just behind Claude Opus 5. Analytical quality jumped from sixteen-ninety to nineteen-ninety-four Elo. Coding agent index up nine points to fourth place. Hallucination rate down from thirty-four percent to twenty-nine.
That all sounds pretty decent.
It is. The catch is token economics. At the highest reasoning effort, Grok 4.7 burns about eighty-one thousand output tokens per task. Grok 4.6 used thirty-six thousand. GPT-6 Astra uses twenty-seven.
Wait — three times the tokens?
Which is why a headline price of six dollars per million means much less than it looks. In a market moving toward agents that run for hours, cost-per-task is the real price, and on that measure this is more expensive than the sticker suggests. One developer reported it as definitely slower and more expensive, and still below the intelligence floor they need for agentic coding.
So the summary is: lots of compute, real gains, wrong direction on efficiency.
That's what a compute-rich lab looks like when it's still climbing.
Now, Jev. We covered the model that refuses to write sentences yesterday. Something happened in the week since it launched.
The ecosystem cloned it. Yesterday's Hacker News front page carried Kev — Jared Palmer's open-source family of Jev-like decision models built on Qwen3.5, at zero point eight, four and nine billion parameters. Four hundred and twenty-six points. That's seven days from a genuinely novel idea to a credible open-weights reproduction.
Seven days of moat.
That's the recurring pattern of this era, and it's worth internalising if you're building a company on a single clever idea.
And there's a parody, isn't there.
Jev-leftpad. Two hundred and twenty-eight points. It implements the infamous left-pad function as a Jev call, in the honourable tradition of FizzBuzz in TensorFlow. And the critique of the parody is the funniest technically-correct thing on the internet this week — it misses one of Jev's core features, the confidence scores. Partial confidence could easily be mapped to fractional spaces, using unicode thin space.
I love that someone cared enough to be pedantic about a joke.
Simon Willison's serious concern is worth keeping though. There's no reasoning trace — if Jev marks something as spam, which signals tipped it off? You get a number. His own test rating Bay Area neighbourhoods returned Cupertino highest and East Palo Alto lowest, which he flags as exactly how easily bias hides inside an unexplainable score.
Commerce. Amazon has blocked Meta's Muse agent.
Started Sunday night, after Meta declined a direct request to pull the bot off amazon.com. Muse is Meta's general-purpose agent — shopping, booking appointments. Users now get pop-ups telling them they're violating Amazon's terms of use.
What's Amazon's stated objection?
Three things. Meta never told them Muse would be accessing the store. The agent doesn't identify itself when browsing. And Amazon says it appears to capture and store customer credentials.
That last one sounds serious.
It would be, and it's Amazon's characterisation, not something Meta has confirmed. An agent holding your Amazon login is a real attack surface. But I'd hold both thoughts at once: the security complaint is plausible, and it is extremely convenient. Amazon sued Perplexity over the Comet browser, and has moved to block shopping agents from Google and OpenAI too. The commercial logic is transparent — Amazon's margin depends on sponsored placement and its own recommendation surface. An agent that comparison-shops across retailers strips all of that out.
And they're business partners.
Deeply entangled. Amazon products have been purchasable inside Facebook and Instagram since 2023, and in April Meta signed a multibillion-dollar deal to run agentic workloads on Amazon's Graviton chips.
Is there a version of this that ends well?
The developers described it better than the lawyers will. One put it simply — the right way to do this is build an API, or an MCP endpoint, catering to bots, and write corresponding terms into your legal agreements. Authenticated bot access with liability attached, rather than cat-and-mouse user-agent detection. Because the enforcement problem is unsolvable otherwise: nothing stops you running a local agent driving your own browser with your own login, and at the network level that's indistinguishable from a human.
Anything you'd push back on in the whole premise?
One commenter has been asking in every agent thread for a single measurable life improvement from these shopping agents, and hasn't got one yet. So this may be the opening skirmish of a war over a market that doesn't exist.
Money, and a big number. SoftBank.
Launched a high-yield bond offering of more than eleven billion dollars on Saturday. Roughly ten billion in dollar tranches at three and a half, five and a half and seven and a half years, plus a billion euros. Pricing is Thursday, settlement the twenty-ninth. Proceeds partly fund the continuing OpenAI investment — which on completion takes SoftBank's cumulative position to about sixty-four point six billion dollars, for roughly thirteen percent of the company.
And if it closes at that size?
Largest non-financial corporate bond deal ever out of Asia-Pacific, beating 7-Eleven's ten point nine billion from 2021. And note: launched, not raised. It prices Thursday.
The word you keep leaning on is "junk."
Because it's the whole point. This is sub-investment-grade debt, raised to buy equity in a private company with no public financials, no liquidity and no mark-to-market discipline. Debt has fixed obligations. That's the clearest single data point on how the AI capital stack is actually funded — not operating cash flow, not patient equity. Leverage, against a concentrated bet on one company.
Masayoshi Son has done this before.
Sometimes spectacularly well, sometimes spectacularly badly. My favourite framing came from a commenter in an unrelated thread about Sun Microsystems — sold his Sun stock at the bubble top for seventy dollars a share, a few months later it was seven, and he thinks of that whenever he sees price-to-earnings ratios in the hundreds.
Last one, and it's hardware. MacStories reviewed the M5 Ultra Mac Studio specifically as a machine for local AI agents.
Two hundred and fifty-six gigabytes of unified memory, eighty-core GPU, memory bandwidth up to one point two terabytes a second — fifty percent over the M3 Ultra. Prompt processing about two and a half times faster, generation seventy percent faster. Time to first token at a two hundred and fifty-six thousand token context drops from two hundred and forty-six seconds to a hundred and two.
How does it compare to just buying a good graphics card?
That's the comparison everyone wanted, and Simon Willison pulled the buried chart into the thread. Against an RTX 5090 on a twenty-seven billion parameter model: at eight thousand tokens of prompt, the 5090 wins fifty-nine tokens per second to forty-eight. At sixty-four thousand, fifty-one to thirty-nine. At a hundred and twenty-eight thousand, forty-four to thirty-two.
So the Mac loses across the board.
Until the last column. At two hundred and fifty-six thousand context, the Mac manages twenty-four tokens a second and the 5090 can't run it at all. Thirty-two gigabytes of video memory means it has to push layers across the bus into system RAM, and performance collapses. Unified memory trades peak compute for being able to hold the thing in memory in the first place.
And the price?
Eighteen thousand dollars as reviewed. A commenter did the arithmetic — that's about twelve years of OpenAI Pro subscriptions. Another noted the reviewer isn't a working developer, so we still don't know if local models on this hardware make an engineer as productive as a two-hundred-dollar-a-month cloud plan.
But you think the case is real.
Two reasons. No per-token billing, no data leaving the building, no rate limits. And no provider outage taking your workflow down — which has teeth, because overnight Anthropic's status page reported elevated errors and the thread filled with developers rerouting mid-task. My favourite line: you're telling me I have to switch to using my brain's tokens now?
The honest counterweight being that the models you can run locally aren't the ones winning the benchmarks.
Correct. But a genuine local option is a healthy check on a market where three companies currently set the prices.
One to watch: OpenAI's mathematics claims. Either those proofs get published and independently verified in the coming weeks — which would be genuinely historic — or they don't, and that advisory group becomes a case study in how credibility gets borrowed.
Agreed, and unusually for this industry, there's a definitive answer coming. Maths doesn't do vibes.
That's your AI in 15 for today. See you tomorrow.