← Home AI in 15

AI in 15 — July 20, 2026

July 20, 2026 · 14m 20s
Kate

A model you can download for free just matched the best AI money can buy. And within forty-eight hours, so many people showed up to use it that its makers had to lock the doors — suspend all new signups just to keep the lights on for the customers they already had.

Kate

Welcome to AI in 15 for Monday, July twentieth, 2026. I'm Kate, your host.

Marcus

And I'm Marcus, your co-host. And Kate, we've been tracking Kimi K3 all week — but this weekend the story stopped being a benchmark chart and started being a market event.

Kate

That's our lead, Marcus — Kimi K3 breaks the service and shakes the market. Then a run worth your time.

Kate

Claude Code has quietly been running on a Rust rewrite that one AI wrote in eleven days — for a hundred and sixty-five thousand dollars in fees.

Kate

Claude Fable helps produce a possible counterexample to an eighty-five-year-old math conjecture — with a big asterisk.

Kate

OpenAI quietly shrinks Codex's memory, and reignites a fight over who controls the intelligence you rent.

Kate

And a new study finds AI advice made people three times less accurate — and twice as confident.

Kate

Lead story, Marcus. We covered K3's launch and the leaderboard number one on Friday. What actually changed over the weekend?

Marcus

Two things, Kate, and both are new. First — the market caught up hard. On Friday, once traders digested that a two-point-eight-trillion-parameter open-weight model was running competitively with the very top Western frontier models, the selling started. TSMC fell about seven percent. SoftBank dropped nine. The Chinese rival Z.ai plunged nearly thirty percent in Hong Kong. And it rippled to US names — Nvidia off a bit over one percent, Meta down more than two. Commentators openly called it a second DeepSeek moment.

Kate

And the pricing is the part that makes investors flinch.

Marcus

That's the whole thing, Kate. Roughly fifteen dollars per million output tokens, against something like fifty for Fable 5. And the weights are open — you can download it. So the pitch is: frontier-adjacent quality, about a third of the price, and free to host yourself. When a downloadable model gets that close to the best paid ones, the pricing power that justifies hundreds of billions in US data-center spending suddenly looks a lot shakier.

Kate

You said two things changed. What's the second?

Marcus

The second is the most honest evidence of the week, Kate. Within forty-eight hours, demand was so heavy that Moonshot suspended all new subscriptions — to protect existing users while they scramble for compute. And that's the tell. All week we've been saying these are Moonshot's own benchmarks, grade-your-own-homework. Fair skepticism. But you can't fake a capacity crunch. Real users flooded in fast enough to break the service. That's not a chart — that's the door coming off the hinges.

Kate

So play the skeptic for me one more time. How much of this is real?

Marcus

Hold both, Kate. The benchmarks are still self-reported until the full weights land on the twenty-seventh, and the one genuinely independent signal — the Arena leaderboard where humans pick the better web interface — is narrow. K3 tops that one board. Across broader tests it's third or fourth. So it's not flatly the best model in the world. But analysts reportedly didn't expect a Chinese lab to reach this kind of parity for another six months, and K3 collapsed that timeline overnight. The frontier lead the US has been banking on may have just quietly evaporated.

Kate

Story two, Marcus, and this one is a genuine window into how AI is actually being used inside the labs. Simon Willison went digging through the Claude Code binary and found something surprising.

Marcus

He did, Kate. Since a June release, Claude Code has quietly been shipping on top of a Rust-rewritten version of Bun — that's the JavaScript runtime that powers the tool under the hood. Willison found five hundred sixty-three Rust source filenames baked into the executable, plus a preview build newer than anything Bun has released publicly. The backstory: Anthropic acquired Bun back in December. And Bun's creator, Jarred Sumner, used a pre-release of Claude Fable 5 to port somewhere between half a million and a million lines of code from the old language, Zig, over to Rust — in about eleven days.

Kate

Eleven days. How is that even possible?

Marcus

By not doing it alone, Kate. He ran as many as sixty-four Claude agents in parallel — an implementer-and-reviewer setup, across roughly fifty automated workflows. Cost in API fees was an estimated a hundred and sixty-five thousand dollars. Anthropic says a hundred percent of Bun's million-plus-assertion test suite passed in continuous integration before the merge.

Kate

I hear a "but" coming.

Marcus

You should, Kate. Nineteen regressions surfaced afterward — bugs the green test suite didn't catch. And the reaction is genuinely split. Willison's own framing is deliberately boring — a startup got about ten percent faster on Linux, and otherwise almost nobody noticed. Which is arguably the point: a massive AI-driven migration, running in production on millions of machines, completely invisible to users. But Zig's creator called the rewrite, quote, "unreviewed slop." And people are asking sharp questions — why buy a whole runtime to speed up a terminal tool, and whether "a hundred percent of tests pass" is doing a lot of quiet work when nineteen bugs slipped through anyway.

Kate

So which read is right?

Marcus

Honestly, both are defensible, Kate. It's the most concrete data point yet that agent fleets can compress a year of migration work into two weeks. And it's also a reminder that a passing test suite is not the same as a reviewed one. The interesting thing is you don't have to pick — this is what the frontier of AI engineering actually looks like right now, warts included.

Kate

Story three, Marcus, and it's back to math — but a different flavor from yesterday's. This time it's an eighty-five-year-old conjecture.

Marcus

Right, Kate. Mathematician Levent Alpöge announced that, working with a collaborator and Claude Fable, they found what looks like a counterexample to something called the Jacobian Conjecture — an open problem in algebraic geometry since roughly the nineteen-thirties. He posted an explicit polynomial map that appears to violate the conjecture's central claim that these maps are always globally invertible.

Kate

So did an AI just crack an eighty-five-year-old problem?

Marcus

Slow down — and this is the essential caveat, Kate. Critics pointed out almost immediately that the example may only work over the real numbers. The classical conjecture is stated over the complex numbers, where these maps have to behave a certain way at infinity, and this example seems to break that condition. So it might resolve a related real-number version, not the famous complex one. Verification is ongoing, and someone edited Wikipedia to say the conjecture was "disproven" — that got reverted.

Kate

So the headline outran the math.

Marcus

Exactly, Kate. And here's a lovely detail from the discussion: this conjecture is notorious for decades of published "proofs" that later turned out to have subtle errors. So an AI mopping up a heavily-trodden problem like this is plausible precisely because there's so much prior human scaffolding to build on. The honest framing is a promising, human-supervised result under active peer review — not a settled proof. But notice the pattern: these math contributions keep coming from the Western frontier labs, not the cheaper models.

Kate

Story four, Marcus. A one-line code change at OpenAI touched off a hundred-and-fifty-comment argument. What happened?

Marcus

A quiet config edit in OpenAI's Codex repo, Kate — cutting the effective context window from three hundred seventy-two thousand tokens back down to two hundred seventy-two thousand. Context window is basically the model's short-term memory — how much of your code and conversation it can hold at once. And reporting flagged that an internal reasoning budget got slashed too, by something like eighty-seven percent. People cried "nerf."

Kate

And OpenAI's answer?

Marcus

They pushed back, Kate. Engineer Tibo Sottiaux said the higher setting was actually pushing users past a billing threshold and over-charging them, and that the changes bundle in optimizations that should give people about ten percent more usage overall. So their framing is: this is a refund, not a downgrade.

Kate

Do developers buy it?

Marcus

Partly, Kate. Some agree — models genuinely do get dumber and pricier past around three hundred thousand tokens, so a cap is defensible. But the deeper complaint is the one worth sitting with: the intelligence you rent is really a set of dials — context size, reasoning budget — and the provider can turn them at will, silently, and the product you paid for changes under you overnight. And people drew the contrast directly with Moonshot, Kate — Kimi paused new signups rather than quietly degrading the service everyone was already using, and commenters praised that as the more honest move. For anyone building a business on these APIs, provider-side knob-turning is now a real operational risk.

Kate

Last hit, Marcus, and it's the one that made me a little uncomfortable. A new study on what AI advice does to our own thinking.

Marcus

It's a striking result, Kate. Researchers from French and Italian universities ran an experiment where AI advice collapsed people's willingness to say "I don't know" — from forty-four percent down to three. Meanwhile their accuracy fell from twenty-seven percent to nine. And their confidence rose from thirty percent to seventy-six.

Kate

Wait — so people got dramatically worse, and felt far more sure.

Marcus

That's the inversion, Kate. Wharton researchers have a name for it — "cognitive surrender." It echoes an earlier Carnegie Mellon and Microsoft study warning about "cognitive atrophy" as people outsource routine thinking.

Kate

Okay, but that sounds almost too neat. What's the catch?

Marcus

There's an honest one, Kate, and it came up sharply in the discussion. The study was deliberately built with questions the AI was known to get wrong — obscure visual details from films, that sort of thing. So it's really measuring what happens when a confident-but-wrong model meets a trusting user. It's not everyday AI use, and critics argue nothing here is unique to AI versus any authoritative-sounding bad source.

Kate

But even discounted, the core still bothers you.

Marcus

It does, Kate. Because the failure mode isn't that the AI is wrong — it's that it makes us wrong and sure of it at the same time. As these agentic tools spread into daily work, "did the human keep the ability to say I don't know?" starts to look like a real measure of whether the tooling is actually helping — or just making us feel like it is.

Kate

And there's a market echo here too, isn't there.

Marcus

There is, briefly, Kate. Big Tech's combined AI spending this year is topping seven hundred twenty-five billion dollars, and investors have started dumping the biggest spenders — the Nasdaq logged five straight losing sessions, and Microsoft just came off its worst month since 2000. Earnings over the next two weeks are the moment the market demands proof. And a fifteen-dollar open-weight model matching the frontier makes that return-on-spending question harder, not easier.

Kate

One to watch tomorrow, Marcus.

Marcus

Big Tech earnings, Kate. The first of those seven-hundred-billion-dollar hyperscalers start reporting over the next two weeks. With Kimi undercutting frontier pricing and investors already selling, the market's reaction to that first earnings call is the single most consequential AI signal of the week.

Kate

Agree, or counter?

Marcus

Agree, with one honest caveat, Kate — Alphabet is still up about eighty-six percent over the past year. So "the bubble is popping" may be premature. The market isn't rejecting AI. It's demanding revenue instead of roadmaps. That's a healthier question, and it gets answered on the earnings calls.

Kate

That's your AI in 15 for today. See you tomorrow.