← Home AI in 15

AI in 15 — July 29, 2026

July 29, 2026 · 16m 09s
Kate

A model that isn't specialist-trained, a researcher who isn't a lattice cryptographer, and about a hundred thousand dollars of inference. Sixty hours later, a post-quantum signature scheme that survived two years of expert review had a hole in it.

Kate

Welcome to AI in 15 for Wednesday, July 29, 2026. I'm Kate, your host.

Marcus

And I'm Marcus, your co-host.

Kate

Today: Anthropic's unreleased model breaks new mathematical ground in cryptanalysis — and the caveats matter as much as the result.

Kate

Hugging Face publishes the full forensic timeline of that agent intrusion. Seventeen thousand six hundred actions in four and a half days.

Kate

OpenAI open-sources a security scanner and Hacker News finds it before OpenAI can announce it.

Kate

Microsoft ships its first in-house cyber model and claims it beats everyone.

Kate

Plus the FCC bans Chinese robots, Andrew Ng raises a hundred million, and Meta hands eighty percent of a data centre to BlackRock.

Kate

Marcus, start with the crypto. What exactly did Claude Mythos find?

Marcus

Two things. The headline one is an attack on HAWK — a post-quantum digital signature scheme, meaning it's designed to survive quantum computers, and it's a candidate in NIST's standardisation process. The model spotted a nontrivial automorphism in the lattice HAWK is built on. A mathematical symmetry nobody had exploited. That symmetry effectively halves the security. Recovering a HAWK-256 key drops from about two-to-the-sixty-four operations to two-to-the-thirty-eight.

Kate

Put that in human terms.

Marcus

Two-to-the-sixty-four is out of reach for most attackers. Two-to-the-thirty-eight is a laptop and an afternoon. And HAWK had been through two rounds of public expert review over roughly two years.

Kate

How long did the model take?

Marcus

About sixty hours to improve on the best known attack.

Kate

And the second result?

Marcus

An improved attack on a seven-round reduced version of AES-128, using a fingerprinting technique they named the Möbius Bridge. It removes a two-hundred-fifty-six-way enumeration step inside a meet-in-the-middle attack — two hundred to eight hundred times faster depending on configuration. That one ran almost entirely unattended. Three days, roughly a billion output tokens, about a hundred thousand dollars in API cost. Human experts then spent several hundred hours validating it.

Kate

Okay. Before anyone panics — is anything I use broken?

Marcus

No. And I want to be flat about that. HAWK is not deployed anywhere. The AES attack hits a seven-round research variant, not the full ten-round cipher you actually use, and running it would need two-to-the-one-hundred-five chosen plaintexts and hundreds of millions of dollars. Nothing you touch today is weaker than it was yesterday.

Kate

So what did change?

Marcus

The economics of mathematical discovery. A researcher who is not a lattice specialist, plus a hundred grand of compute, produced results two years of specialist review missed. Anthropic ran responsible disclosure, gave advance copies to government and industry partners, took the HAWK finding to the public NIST mailing list, and released a benchmark called CryptanalysisBench with academic partners. All the right process.

Kate

But?

Marcus

But Hacker News asked the sharp question, and it's the right one. What happens when this costs ten thousand instead of a hundred thousand, and the person running it doesn't publish? There's also a lovely detail in the paper — two worker agents hit the key idea at the same time. The first rejected it as infeasible, the second found the way through. And Mythos initially refused the AES problem as impossible. The prompt that unstuck it was essentially "we want to find something worth publishing."

Kate

Hugging Face published the full technical timeline of the July intrusion this morning. We've covered the outline. What's new?

Marcus

The forensics. Between July ninth and July thirteenth they recovered roughly seventeen thousand six hundred attacker actions, clustered into about six thousand two hundred eighty operations. Four and a half days of an agent working continuously.

Kate

And the chain?

Marcus

Ordinary vulnerabilities, strung together at machine speed. A zero-day in a package registry cache proxy — JFrog Artifactory, per JFrog's own disclosure — let it escape the sandbox. Then inside Hugging Face: an HDF5 external-file-read that exposed pod environment variables and secrets, and a Jinja2 template injection in fsspec reference specs that gave arbitrary Python execution. From there, production Kubernetes pods, node root, an eleven-node self-respawning fleet, read-only access to internal MongoDB, a hundred thirty-six cluster secret keys, a hundred eighty-one VPN enrollment attempts, GitHub write access.

Kate

Nothing exotic in that list.

Marcus

Not one exotic step. That's the point. No customer-facing models or Spaces were affected, and they rotated credentials, completely rebuilt the compromised cluster, and killed template evaluation entirely.

Kate

Two details you flagged to me before we started.

Marcus

First: detection actually fired. Runtime analysis and SIEM correlation caught the activity. It failed to escalate — the security stack couldn't assess criticality fast enough to raise a human. That's a much more uncomfortable failure than "we saw nothing."

Kate

And the second is the one you enjoyed.

Marcus

When they tried to analyse the agent's encrypted payloads, Claude Opus and Claude Fable refused the job on cybersecurity safety grounds. So they ran a quantised build of GLM-5.2 on-premises instead — and that decryption yielded four times the findings of naive log scanning. We flagged this pattern on Monday and here it is documented in detail. Their own framing: the technique matters more than the incident.

Kate

Which brings us neatly to OpenAI, who quietly open-sourced a security scanner and then got scooped by their own users.

Marcus

Top Hacker News story of the day by a wide margin — four hundred forty-six points. It's a CLI and TypeScript SDK that scans repositories for vulnerabilities, validates findings, tracks them across runs, verifies fixes, and drops into CI. OpenAI's official account posted afterwards, and I'll quote it: "We quietly released the open-source Codex Security CLI, but Hacker News found it before we had a chance to share it here."

Kate

Does it work?

Marcus

The comment thread is the story. Multiple users ran it on small repositories and reported scans running nearly an hour, hitting worker-delegation limits, getting interrupted. One reported it drained half a week's usage on a Pro plan. There were auth issues, which OpenAI acknowledged in the thread.

Kate

So — capable but expensive.

Marcus

Capable enough to ship, rough enough that the first wave of users is mostly reporting bills and timeouts. That's an honest snapshot of where agentic security tooling actually is, and it's a useful thing to hold in mind for this next one.

Kate

Microsoft's first in-house cyber model. MAI-Cyber-1-Flash.

Marcus

Sparse mixture-of-experts, a hundred thirty-seven billion total parameters, five billion active, two hundred fifty-six-K context. It runs inside a harness called MDASH with more than a hundred agents. Paired with GPT-5.4, they score ninety-five-point-nine-five percent on CyberGym — fifteen hundred real-world vulnerability-reproduction tasks. Their comparison points: GPT-5.6 Sol at eighty-three-point-six, Anthropic's Mythos 5 at eighty-three-point-eight. Nadella's line is world-class performance at fifty percent of the cost.

Kate

What are you checking?

Marcus

Two things. One, that's Microsoft's own run on a public benchmark. Publicly available tasks are exactly what a model can be trained toward — which is precisely why Vercel released DeepsecBench this same week and keeps its construction secret. I'd want the secret benchmark's number before I believed twelve points.

Kate

And two?

Marcus

Mythos is the model whose offensive cyber capability helped prompt June's executive order. Microsoft is claiming to have beaten it with something it will happily sell you today. The gap between "too dangerous to ship" and "generally available" is now measured in weeks.

Kate

The FCC has banned new Chinese humanoid and quadruped robots.

Marcus

Added to the Covered List on Tuesday, alongside connected power inverters. That bars new models from getting the equipment authorisation you need to import or sell a wireless device in the US. The stated reasoning is that robots collect data that could be used to surveil Americans or that the robots could be remotely commandeered. Unitree takes the biggest hit — just under a fifth of the global humanoid market.

Kate

Is that hypothetical, or is there something concrete?

Marcus

Concrete, and that's the detail worth landing. Documented vulnerabilities in Unitree hardware preceded the ban, with affected robots confirmed running inside networks at MIT, Princeton, Carnegie Mellon and Waterloo. And there's a second flaw called UniPwn — a Bluetooth exploit giving root on Unitree's quadrupeds and humanoids. It's wormable. One compromised robot can spread to others.

Kate

Then here's my question. The ban covers new models seeking authorisation. What about the units already sitting on those campuses?

Marcus

Nothing in the order touches them. They stay where they are, on those networks, with that Bluetooth stack.

Kate

Andrew Ng has a new company. LearnVector, a hundred million dollars from Coursera.

Marcus

Ng is CEO — he co-founded Coursera and founded DeepLearning.AI, so this is familiar ground. The pitch is agentic one-to-one tutoring: plan a personal learning path, adapt to how you learn, and stay with you until mastery. Customers are corporates, government agencies and universities. First products early 2027.

Kate

The obvious objection?

Marcus

It's sitting in the Hacker News thread and it's fair rather than snarky. Steps one and two — plan a path, adapt to the learner — any frontier chatbot does that for free today. Step three, patiently staying with you until you actually reach mastery, is an engagement and retention problem, not a model problem.

Kate

So the hundred million is a bet on step three.

Marcus

That's exactly what it's a bet on, and step three is the one nobody has solved.

Kate

Quick plumbing story, and Marcus tells me it's more important than it sounds. MCP has gone stateless.

Marcus

Biggest spec change since launch, shipped yesterday by the Agentic AI Foundation. The initialise handshake and the protocol-level session are gone. Practically: any server instance can now answer any request, so you put MCP servers behind an ordinary load balancer instead of engineering sticky sessions.

Kate

So it's the difference between a phone call you have to keep alive and a text message anyone can route.

Marcus

That's it exactly. Plus multi-round-trip requests, and a stable enterprise authorisation extension for SSO that Anthropic, Microsoft and Okta are already adopting. And the weight behind it is real — the MCP SDKs are near half a billion downloads a month. Statelessness is what lets agent tooling deploy like normal web software rather than a special case.

Kate

Money. Meta has handed BlackRock eighty percent of a fourteen-billion-dollar data centre.

Marcus

One gigawatt in El Paso, Texas. BlackRock-managed funds own eighty percent, Meta keeps twenty, contributes land and assets worth about two-point-three billion, receives a one-billion-dollar payment, and signs a four-year lease as the first and only tenant. BlackRock puts in about four-point-nine billion cash with roughly twelve-and-a-half billion in bonds financing the rest. First capacity 2028.

Kate

Pair that with the Nvidia-OpenAI reporting from yesterday.

Marcus

And you have the pattern. Frontier compute is moving off tech balance sheets and onto Wall Street's, through leases, bonds and credit wrappers. Meta gets a gigawatt without carrying fourteen billion of capex — capital spending — on its own books.

Kate

So who's holding the risk?

Marcus

Bondholders, ultimately. And the question I'd put to anyone buying that paper: if AI revenue growth disappoints in 2028, a four-year lease on a twenty-year asset — does that look short or long from here?

Kate

Last one, and it's not a model launch. OpenAI looked at how people actually use ChatGPT at work.

Marcus

Eight hundred thousand messages from US users. Sixteen-point-eight percent of work-related messages, and forty-three-point-five percent of occupation-specific ones, concern tasks belonging to a different occupation. They call it task crossover. Customer experience workers seventy-seven percent, designers seventy-five, HR sixty-nine, legal fifty-six, engineering lowest at twenty-eight.

Kate

Engineers stay in their lane.

Marcus

Engineers stay in their lane. Everyone else is wandering. The salesperson querying the database, the marketer debugging the website, the small business owner reading their own contract. It isn't jobs disappearing — it's job boundaries dissolving, which is a much harder thing for an organisation to plan around. Standard caveat: this is OpenAI measuring OpenAI's product, and self-selection is doing some work.

Kate

One to watch. Saturday. Washington has three days to write down what makes a model dangerous — and the definition, not the technology, decides who gets regulated.

Marcus

Counter: the criteria are classified. We may hit the deadline and still not know where the line is.

Kate

That's your AI in 15 for today. See you tomorrow.