← Home AI in 15

AI in 15 — August 16, 2026

August 16, 2026 · 14m 21s
Kate

Anthropic has a model that's better than the one you can buy. And they've told you exactly how much better — sixty-two point eight versus fifty point three — and then said they have no plans to sell it to you.

Kate

Welcome to AI in 15 for Sunday, August sixteenth, 2026. I'm Kate, your host.

Marcus

And I'm Marcus, your co-host.

Kate

Today: Anthropic puts a number on the model in the vault, and raises its own risk rating for reasons that have nothing to do with that model.

Kate

DeepSeek's price rise goes live today, at four o'clock UTC.

Kate

Alibaba drops a multimodal model that runs on a gaming graphics card and scores eighty-four percent on computer use.

Kate

An OpenAI model finds two zero-days in Chrome's JavaScript engine.

Kate

Plus Anthropic posts eleven and a half billion dollars in a quarter, and somebody is buying up secondhand bookshops' entire stock and shredding it.

Kate

Marcus, we talked about Model 2 yesterday. What's new?

Marcus

A number. Anthropic's risk report gives Model 2 a CoBench score of sixty-two point eight, against Claude Mythos 5's fifty point three. Yesterday we had "somewhat more capable." Now we have twelve and a half points on a coding benchmark, and no plans to release it.

Kate

So which is it — somewhat, or twelve points?

Marcus

Both, and that tension is the story. Twelve points on one eval is real. It's also nowhere near the generational jump people imagine is sitting behind the curtain. And here's the part I keep coming back to: the risk-rating increase — very low to low on catastrophic misalignment — was not caused by Model 2. Their own internal deployment review found nothing new or worse in it than they'd already characterised for Mythos.

Kate

Then what drove it?

Marcus

Three things. Recent cybersecurity incidents. A biosafety classifier that failed. And benchmark saturation — the safety evaluations the whole industry relies on are hitting their ceiling. Models score near the top, so the test stops distinguishing a safe model from a dangerous one.

Kate

That last one sounds like the real headline.

Marcus

I think it is. Everything else about that report is a claim only Anthropic can verify. "We have something better in the drawer" is unfalsifiable from outside and it happens to be superb positioning. But saturation is checkable by anyone, and it applies to every lab equally. If your instrument has stopped moving, you don't know whether that's because nothing's wrong or because you can't see it anymore. There's no regulation that fixes a broken thermometer.

Kate

They also flagged something about models improving models.

Marcus

Signs of acceleration in automated AI research and development. Which sits slightly awkwardly next to yesterday's line — "faster, but not yet by a factor of two." Directionally up, still not a takeoff.

Kate

DeepSeek. The price rise we've flagged twice now actually happens today.

Marcus

Sixteen hundred UTC. So depending on when you're listening, it may already have landed. V4-Pro output goes from eighty-seven cents per million to three ninety-six at peak. V4-Flash output from a flat twenty-eight cents to a dollar thirty-two. Cache-miss input on Pro triples. Top end of the range is over eleven hundred percent.

Kate

And the peak windows are odd.

Marcus

Zero one hundred to zero four hundred, and zero six hundred to ten hundred UTC — which maps neatly onto the Chinese working day. Off-peak is half. So they're not just raising prices, they're rationing by clock.

Kate

Is that scarcity or margin?

Marcus

Genuinely unclear, and I'd want to see them answer it. The stated reason is allocating resources more reasonably. The unstated context is that this is a company operating under export controls on exactly the hardware it needs to serve that demand. If you can't buy more silicon, price is the only lever you have left. But it's also worth noticing this arrives right after the land-grab phase ended — undercutting everyone is a strategy for winning share, not for running a business.

Kate

And here's the part I love. The weights are free.

Marcus

They are. Simon Willison flagged the V4 Pro zero-eight-one-three checkpoint on Hugging Face — one point seven trillion parameters, eight hundred and ninety-three gigabytes of weights. So in the same week, the model became free to download and four to ten times more expensive to rent.

Kate

Which only helps you if you own a data centre.

Marcus

Right. Open weights at that scale are a licence, not an option. Nobody's running eight hundred gigabytes on a workstation.

Kate

Unless — and this is the segue — you download the other one. Alibaba shipped Qwen three-point-eight twenty-seven B on Friday.

Marcus

Apache 2.0, twenty-seven point seven eight billion parameters, natively multimodal — text, images, video. Two hundred and sixty-two thousand token context window natively, extensible to a million. And it targets roughly twenty-four gigabytes of VRAM.

Kate

Twenty-four gigabytes is an RTX 4090.

Marcus

It's a gaming card. And the benchmark movement over its predecessor is not incremental. Terminal-Bench two-point-one, sixty-three to seventy-three. DeepSWE from thirteen point three to forty-two point two — that's a tripling. OSWorld-Verified, sixty-four to eighty-four point three. It reportedly beats Meta's Muse Glimmer 30B on all eight head-to-head comparisons.

Kate

Eighty-four percent on computer use, on a consumer GPU.

Marcus

With no API bill, no rate limit, and no terms of service. And there's reporting that Alibaba also opened weights for the Max-class model — two point four trillion parameters, ninety-five billion active, scoring eighty-six on OSWorld against GPT-5.6 Sol Max's eighty-three. I'd hold that one loosely. The twenty-seven B is solidly corroborated — Willison was building test tools against it yesterday. The Max open-weights drop is thinner sourcing, and Max was announced as an API model two weeks ago.

Kate

So the headline is the small one.

Marcus

The headline is the small one, and it's the better story anyway. That's a direct answer to DeepSeek's price rise. The open-weight frontier is now substantially Chinese, and the thing that makes it matter isn't the parameter count — it's that it fits on hardware people already own.

Kate

Security. An OpenAI model found two previously unknown bugs in Chrome.

Marcus

In V8, the JavaScript engine. Google's patched them as CVE-2026-15903, CVSS eight point eight. The optimising compiler skipped a safety check converting values to integers, so an undefined value could produce an unexpectedly large number — out-of-bounds read and write. Chained together, they let you corrupt memory and escape the V8 heap sandbox.

Kate

That's the sandbox that stands between a website and your machine.

Marcus

That's the one. And this came out of GPT-5.6-Cyber, which we mentioned Thursday. What's new is the numbers behind it. On OpenAI's internal advanced-cybersecurity evaluation, the Cyber model completes ninety-five percent of requests. Standard GPT-5.6 Sol completes one point five percent.

Kate

One point five.

Marcus

Which is the entire argument in two numbers. A model that refuses ninety-eight and a half percent of security questions is useless to a defender and no obstacle whatsoever to an attacker, because the attacker isn't using your compliant API. So OpenAI split Daybreak into Blue for defence and Red for offensive research, put the Cyber model behind Red, and from September first every Daybreak account requires a hardware security key.

Kate

Does the gate hold?

Marcus

That's the only question that matters, and nobody can answer it yet. Both models were assessed at "High" on OpenAI's cyber threshold, below "Critical." V8 zero-days are precisely the capability you'd least like to see walk out the door. But the counterfactual is the honest bit — we covered a researcher building a zero-click Zoom exploit chain from public models in under a day. The capability is ambient. Gating a better version and pointing it at defence is at least a plan.

Kate

Business. Anthropic's Q2 numbers landed.

Marcus

Preliminary revenue of more than eleven and a half billion dollars for the quarter. Same quarter last year: seven hundred and eighty-seven million. That's more than fourteen times, year over year. Up from four point seven three billion in Q1. And it beat the ten point nine billion that was being projected as recently as May.

Kate

And the first profitable quarter.

Marcus

Positive adjusted operating income. And "adjusted" is doing an enormous amount of work in that sentence. Ed Zitron published a piece calling it a profitability swindle, arguing that once you strip out training compute and stock compensation to reach that figure, you haven't described a company that makes money.

Kate

Where do you land?

Marcus

Split. The revenue number is extraordinary and it's independently meaningful — fourteen x is not an accounting choice, that's customers paying. The profit claim is an accounting choice until somebody shows the unadjusted line. And that's exactly what an S-1 forces you to do, which is why the filing is still the thing to wait for.

Kate

Watermarks, briefly, because there's a twist. We covered Anthropic switching theirs on. Google went the opposite way.

Marcus

Google's adding a toggle to turn the visible watermark off — across the Gemini app and the Flow video editor, covering Nano Banana, Gemini Omni and Lyria. But the invisible SynthID signal and the C2PA metadata stay embedded regardless. SynthID survives screenshots, crops, format conversion, and Google says it's now on over ten billion pieces of content.

Kate

So they're separating "you can see it's AI" from "you can prove it's AI."

Marcus

Exactly that. One is a UX annoyance, the other is provenance infrastructure. And the open question neither company will answer: a watermark you can only read with the vendor's private key isn't public verification. It's vendor-mediated verification. Who gets a detection key, and who decides?

Kate

One detail from Anthropic's write-up that stuck with me — if you rewrite the text enough to strip the watermark, they say it's arguable whether it's AI-generated anymore.

Marcus

Which is a defensible position and a very convenient one.

Kate

Last one, and it's the strangest story of the week. Secondhand booksellers across Britain and Ireland think AI companies are buying their stock and destroying it.

Marcus

Bulk orders started arriving about three months ago from buyers in the US, Canada and Europe. Stuart Manley at Barter Books in Alnwick says what made them odd is that they had no pattern — no theme, no subject, no sport or motoring grouping. Just breadth.

Kate

What are the tells?

Marcus

Opaque buyer aliases, everything routed to the same freight warehouses, and — this is the one that convinced me — buyers paying full retail without asking for the bulk discount any genuine reseller would negotiate hard for. The suspected pipeline is collect in Europe, ship to the US, then destructive scanning. Slice the spines, digitise the text, bin the book.

Kate

Why destroy it?

Marcus

Because the court ruling that made scanning defensible turned on the physical copy being destroyed. So the legal structure actively rewards pulping the original. Nobody set out to write a rule that says "we bought it legally" therefore "we shredded it," but that is the rule.

Kate

And it tells you something about data.

Marcus

It tells you clean, non-synthetic text has become genuinely scarce. If you're paying full price for scattergun used books and eating transatlantic freight, you've run out of internet.

Kate

One to watch: Debian is voting on whether to allow AI-written contributions. Nine options on the ballot, voting closes August twenty-eighth, and whatever wins becomes the reference point every other open-source project cites.

Marcus

Counter — the ballot is incoherent enough that the winner may not mean much. Nine options clearly written by nine different people, on no single axis. Watch the DeepSeek price instead, and whether it sticks past a week.

Kate

That's your AI in 15 for today. See you tomorrow.