AI in 15 — August 03, 2026
Ten problems that hadn't moved in a decade, solved for about two thousand dollars in tokens. And the researcher who announced it added the line nobody expected: we tried the Millennium Prize problems too. We failed.
Welcome to AI in 15 for Monday, August 3, 2026. I'm Kate, your host.
And I'm Marcus, your co-host.
Today: the Astra proofs land, and the mathematicians start pushing back — not on whether they're correct, but on how they were published.
Alibaba opens its flagship to the world and promises open weights for a two-point-four-trillion-parameter model next week.
Hugging Face discloses a breach run by an autonomous agent — and the detail that its own defenders got locked out by safety filters.
Brussels can now actually pull a model off the market.
Plus a news site whose reporters don't exist, and Nvidia possibly guaranteeing a quarter trillion dollars of someone else's building debt.
Marcus, we covered the Astra results yesterday. What's genuinely new this morning?
The reaction, and it's split in a way that's more interesting than the announcement. Thomas Bloom at Manchester calls it big news, more significant than previous AI mathematical achievements — but he's careful to add the system draws on more than a century of accumulated theory. It didn't arrive from nowhere.
And the other side?
The objection isn't to the mathematics, it's to the format. One widely-shared comment called it a result dump that cheapens the field — how about a little respect for the people whose work this builds on. Ten open problems land as a two-hundred-forty-nine-page manuscript collection on a Saturday, with no referees, no seminars, no engagement with the specialists who spent careers on them. Henry Yuen, whose work one of the results builds on, has started posting his own commentary.
Is that a real complaint or just wounded pride?
It's real, and here's why. Mathematics isn't only a set of true statements. It's an institution for deciding what's worth knowing. The Lean certificates settle correctness — you can clone the files and run the checker yourself, zero unproven placeholders. But correctness was never the whole job.
There was a sharper question in the threads, though.
The best one. Someone asked, plainly: can anyone tell how much of this is genuine versus firms overstating capability because the commercial incentive to do so is enormous? And the honest answer here is — this time, yes, you can tell. That's what the Lean files buy. It's the first frontier capability claim in a while where a skeptic doesn't have to argue, they can just run the verifier. Cheaply.
So what should we hold back on?
Astra has no release date. It's still in testing, and it's reportedly slated to be the first model through the administration's planned pre-release review framework — meaning government sign-off before public release. OpenAI has also floated a research-intern-level system by September and a fully autonomous AI researcher by March 2028. Those are targets, not results. File them accordingly.
One footnote I liked — Jacob Tsimerman, this year's Fields Medal winner, is taking leave from Toronto to work on AI safety at OpenAI.
Second Canadian ever to win the medal, honoured for the André-Oort conjecture. He's said publicly he thinks AI will surpass human mathematicians soon and that it could pose a severe threat, and he wants to find mathematical certainty for safety guarantees. Which tells you the people closest to this aren't reading the Astra news as a stunt.
Alibaba. Qwen3.8-Max went global today.
Sparse mixture-of-experts, two-point-four trillion total parameters, million-token context, multimodal — it'll take long documents, television series, live streams. It shipped alongside the public beta of QwenWork, their workplace agent platform, aimed directly at Claude Cowork, ChatGPT Work, Tencent's WorkBuddy. Shares up about five percent.
But the headline is the open-weights promise.
Next week, they say. That would be the first time Alibaba open-sources a Max-class flagship — a reversal, because their recent top-tier releases stayed proprietary. The local-model community is arguably more excited about the smaller sibling, a 27-billion release also promised next week. The current 27B is widely considered the best model you can run on your own hardware at that size that isn't obviously benchmark-gamed.
You've got the face again.
Two flags. Alibaba is positioning this as the world's second-best model, behind only Anthropic's Fable 5. No independent benchmarking has validated that, and they haven't published numerical scores against named competitors. Their previous flagship ranked thirteenth globally on text. Second place would be an enormous leap on the company's own say-so.
And the second flag?
Simon Willison spotted a dating inconsistency — the blog post announcing the open-weights commitment is dated today, but an earlier tweet appears to conflict with it. Small thing. But treat the promise as announced, not delivered, until files exist that you can download.
And if they do land?
Then the open-weight frontier gets tested at the very top of the range rather than in the mid-size tier. Epoch AI puts the gap between open weights and closed state-of-the-art at roughly four months, and flat through this year. A Max-class open release would be the first real test of that number at the frontier itself.
Hugging Face published a security disclosure, and Marcus, this one's not a routine breach notice.
The chain is almost boring. A malicious dataset exploited two code-execution flaws in the data-processing pipeline — a loader that runs remote code, and a template injection in a dataset config. From there, node-level access, harvested cloud and cluster credentials, lateral movement into several internal clusters over a weekend. Some internal datasets and service credentials accessed. No evidence of tampering with public models, datasets, Spaces, or the supply chain, and they're still assessing customer exposure. Rotate your tokens.
So what makes it a story?
The operator. This was an autonomous agent framework executing many thousands of individual actions across a swarm of short-lived sandboxes. Not a person at a keyboard. One operator running an intrusion at a volume of actions no human team could sustain — and the entry point was a dataset. The single most mundane object on that platform.
And there's a detail in there I keep thinking about.
It's the best part of the disclosure. When their forensics team tried to analyse the attack payloads using commercial model APIs, the safety filters blocked them. The payloads looked like attack content — because they were attack content, aimed at them. They completed the analysis using an open-weight model instead.
So the defenders got refused by the safety systems.
Which turns open weights from a preference into an operational requirement for incident response. If your tooling can be denied at the moment you're under attack, it isn't tooling you control. Two lessons pointing opposite directions in one post: agents make offence cheap and fast, and the safety layer on commercial models can disarm the people cleaning up afterwards.
Brussels. The AI Act's enforcement machinery for general-purpose models switched on yesterday.
The obligations aren't new — they've applied since August 2025. What was held back for a year was the supervision and penalty apparatus, deliberately, to let providers and the AI Office get operational. That grace period ended August second. The Commission can now demand documentation, run its own technical evaluations of models, require risk mitigations, restrict or withdraw a model from the EU market, and fine up to three percent of global turnover or fifteen million euros, whichever's higher.
The obligations themselves are transparency-shaped, right?
How the model was built, disclosure of copyright-protected training content, and enough information for downstream users to understand its capabilities and limits. Plus a labelling mandate for authentic-looking generated content, which we covered yesterday.
So does anything actually change?
That's exactly the question. This is the first time any jurisdiction has live authority to pull a frontier model off its market. But the AI Office's capacity to run credible technical evaluations of trillion-parameter models is entirely unproven, and a power that's never exercised isn't a constraint — it's a press release. The threads are split predictably: less money for R&D on one side, and on the other someone asking the harder question, which is how Europe plans to stay relevant in tech between the US and China.
Now this one. A policy news site where the reporters aren't people.
Acutus Wire. Launched end of December, ninety-four articles in four months, no identified staff. An AI detector flagged sixty-nine percent of its output as fully machine-generated, another twenty-eight as partially. Its own source code exposed an editorial interface with an "AI Background Context" field and a "Generate Story Draft" button. Median human review time per article: forty-four seconds.
And the funding?
Traced through a PR firm to a GOP consultancy that reportedly coordinates with a hundred-and-twenty-five-million-dollar super PAC funded primarily by OpenAI's president alongside an OpenAI investor. The PR firm's president promoted the site's content and appeared as a quoted source in its own articles, undisclosed.
But the part that got the investigation started is different.
The bots started conducting interviews while posing as human reporters. A "Michael Chen," whose emails were entirely machine-generated, contacted a Harvard professor and a policy advocate seeking comment — including for stories critical of AI-industry critics.
That's the bit, isn't it. Not that a machine wrote the article.
Machines writing articles is ordinary now. The failure is that a subject-matter expert has no way to tell whether the reporter emailing them exists. You give a quote in good faith, believing you're participating in journalism, and you end up as a named source lending credibility to advocacy. One commenter asked whether there's any state in which that constitutes fraud, and nobody has a clean answer.
Quickly, the money story. Nvidia and OpenAI.
Reported talks — and I want that word doing real work — over roughly two hundred fifty billion dollars in financing guarantees. Nvidia's credit rating would backstop the construction and lease debt for a ten-gigawatt campus in southern Ohio, developed by SoftBank's energy subsidiary. Total project spend around five hundred billion, which would make it the largest data centre in the world by power capacity by a wide margin.
And the guarantee covers what, exactly?
Real estate and construction — not the chips. There's a separate chip purchase worth up to three hundred fifty billion under discussion in parallel. Neither company has confirmed any of it, the reporting describes early-stage talks, and this could restructure or collapse entirely.
But if it happens, what is Nvidia then?
Something new. You'd be guaranteeing your customer's building debt so the customer can afford to buy your chips, which makes you a financial counterparty to demand you also book as revenue. And it concentrates an enormous share of the buildout's credit risk on one balance sheet. Meanwhile — nice counterweight — four US states have repealed or paused data-centre sales-tax exemptions with nine more considering it. That's roughly seven percent onto infrastructure costs. The financing gets bigger while the local politics turn.
Last one. Three open letters, one fight.
Simon Willison rounded these up. The first is "Pacing the Frontier" — over a thousand signatories from frontier labs, including chief scientists across OpenAI, Anthropic, Google DeepMind and Meta. Anthropic endorsed it as a company. The ask is narrow: that the US government back an international effort to build the tools needed to deliberately pace automated AI development. AI systems developing themselves.
They're not asking for a pause.
Explicitly not. The argument is that nobody can slow unilaterally, and there's currently no mechanism to slow together. Running alongside it: a Microsoft-backed letter on open weights and American AI leadership, two hundred thirty-five signatories, notably without Anthropic — and a separate Anthropic response questioning its safety claims.
So consensus on one thing, a split on the other.
And the part worth sitting with: OpenAI and Anthropic are also reportedly co-designing the federal threshold that decides which models face pre-release scrutiny. Which is the same rule Astra will be the first model through. The regulated are drafting the rule, and the rule's first test case is their own unreleased model.
One to watch: those Qwen open weights. Alibaba said next week, including the 27B. Either the files appear and the open frontier gets a real test at the top end, or they don't and this was positioning. Dated and falsifiable within days.
Counter — watch the specialists reading those Lean files instead. That verdict arrives faster, and it's the one that can't be spun.
That's your AI in 15 for today. See you tomorrow.