Home AI News

AI News Roundup for August 15: DeepSeek Ends the Bargain Era at 4pm Tomorrow

DeepSeek's 1,100% API increase lands Aug 16 on the cache-hit tokens agent loops burn most. Plus Claude's watermark, Gemini 3.7 Flash and Qwen 3.8.

Dark developer workspace at night with a monitor showing a rising API cost curve and an amber highlighted peak-hours band

Most of this week’s coverage told you DeepSeek is raising prices and is still cheaper than everyone else, which is true and also the least useful sentence you could write about it. The interesting part is which token type absorbs the biggest increase, because it is the one agent loops burn hardest, and almost nobody is reporting that.

DeepSeek’s 1,100% Increase Lands on the Tokens Your Agents Burn Most

The headline number everywhere is that DeepSeek is raising API prices by as much as 1,100% starting at 16:00 UTC on August 16, alongside the launch of V4-Pro. The number is real. Where it lands is the story.

Run the arithmetic against DeepSeek’s own pricing page and the biggest jump is not on output tokens. V4-Pro cache-hit input tokens go from $0.003625 per million to $0.044 at peak, which is a little over 12x. Output tokens on the same model go from $0.87 to $3.96, a comparatively mild 4.5x. Cache-hit input is what you pay for the repeated prefix of a prompt: your system prompt, your tool definitions, your loaded context, every single turn of an agent loop. DeepSeek put its steepest percentage increase precisely on the workload shape that agentic software generates, and the coverage framing this as “output prices quadruple” is measuring the wrong column.

Then there is the clock. Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, with off-peak set at exactly half. Convert that to US eastern time in August and peak runs 9pm to midnight, then 2am to 6am. If you are a US team that moved batch jobs to the small hours because that is when compute is traditionally cheap, you did not dodge the increase. You scheduled yourself into the center of it.

The cheapest provider in the market just taught everyone that inference pricing is a policy, not a physical constant. Anyone with a spreadsheet assuming costs only fall should open it today.

The why is not mysterious. Bloomberg reported that DeepSeek is raising prices ahead of a potential IPO, and a company heading toward public markets needs its unit economics to stop looking like customer acquisition. The bargain was a strategy. Strategies end.

Triage: breaks your stack. If DeepSeek is in your pipeline, reprice it before Sunday and check what time your crons actually fire.

Claude Now Signs Everything It Writes, Which Matters If You Publish What It Writes

Anthropic published a technical explainer on August 14 covering how its text watermark works, filling in the detail behind the announcement that TechCrunch first reported on August 11. Every Claude model launched on or after August 2 carries an invisible statistical watermark, and it applies worldwide rather than only in the EU, even though EU AI Act Article 50 is what forced the timeline.

The mechanism is a version of SynthID-Text, the approach Google DeepMind published in Nature in 2024. It does not change word choices. It changes the source of randomness used when the model picks among options it considers equally good. Anthropic says this costs no extra tokens, so the price is unchanged, and reports no measurable effect on quality or speed. It covers the Claude API, claude.ai, Claude Code and Claude Cowork.

I write this site’s entire content operation with Claude, so I read the limitations section rather than the summary. Three things are worth knowing before you assume anything about your own output:

  • The watermark proves Claude was likely involved. It cannot distinguish “Claude wrote this” from “Claude lightly edited this,” and it says nothing about other AI systems.
  • Light editing probably will not strip it. A full rewrite where every word changes will.
  • It is sparser on factual passages and on code, because when there is only one correct output there is no randomness left to encode a signal in.

That last point is the one with teeth. The watermark is weakest exactly where text is most constrained, which means a tightly-specified factual brief carries less signal than a discursive essay. Detection is a probability, not a receipt.

This lands a week after Anthropic moved Claude Code execution onto customer hardware, which I covered on Wednesday. Read the two together and the shape is clearer: your code can stay on your servers, and what the model writes still carries a mark on the way out.

One wrinkle nobody has flagged: Claude Opus 5 launched on July 24, nine days before the cutoff. Anthropic says models launched before August 2 sit in a transition period with watermarking arriving “over the coming months.” So the current Max default is, for now, in the older group. That will change, and it will change without a version bump.

Triage: breaks your stack, quietly. Nothing breaks today. But if your business depends on nobody being able to tell, that business now has a countdown on it. I would rather be on the side of the line that says so out loud.

Gemini 3.7 Flash Is Very Good, and the Price Doubles on January 1

Google shipped Gemini 3.7 Flash on August 13, 23 days after 3.6 Flash. The gains are not marginal: DeepSWE v1.1 goes from 49.0% to 65.3%, FrontierCode 1.1 from 34.4% to 43.6%, AutomationBench from 17.0% to 30.4%. For coding and agent work at this price tier, that is a serious release, and SiliconANGLE noted it takes first place on FrontierCode outright.

Two things to hold alongside it.

First, the pricing. It is $0.75 per million input and $3.75 per million output, and Google is clear that this is introductory through December 31, after which it becomes $1.50 and $7.50. That is not a price cut, it is a price cut with an expiry date, and it doubles the day your Q1 budget starts. Model it at the January number.

Second, the benchmark everyone is quoting approvingly. Gemini 3.7 Flash leads GDP.pdf, the expert business-document test, at 34.0%, up from 22.0% and ahead of Claude Sonnet 5 at 28.0% and GPT-5.6 Terra at 24.7%. Winning is real. A 34% ceiling is also real. The best model available gets roughly two thirds of expert document tasks wrong, and GDP.pdf is hard on purpose, which is the point of it. Read that as a green light for a human-reviewed document workflow and a red light for pointing an unsupervised agent at your filings.

Triage: matters. Best value in the coding tier right now, on a countdown.

Qwen 3.8 Puts a 27B Multimodal Model Under Apache 2.0

Alibaba’s Qwen team released Qwen3.8-27B on August 14 under Apache 2.0, as The Decoder covered, with weights on Hugging Face and ModelScope alongside a much larger 2.4T mixture-of-experts sibling. It handles text, images and video, ships a native 262,144-token context that extends toward a million with YaRN, and Alibaba claims it beats the larger Qwen3.7-Plus on coding and office tasks.

Treat vendor benchmarks as vendor benchmarks. Treat the license as the actual news. Apache 2.0 on a 27B multimodal model means you can run it on your own hardware, at a fixed cost you control, with no peak-hour clause and no introductory period. Every other story in this roundup is about somebody else changing the terms underneath you. This is the one that lets you stop being subject to that.

Triage: matters. The hedge against everything above it on this page.

OpenAI’s Ultrafast Is an Announcement, Not a Product

OpenAI previewed Ultrafast mode for GPT-5.6 Sol on August 13, a service tier running on Cerebras hardware. Cerebras says it delivers up to 750 output tokens per second, held on its wafer-scale engine with 44GB of SRAM on-die so weights never stream from off-package memory. OpenAI’s own framing is up to 14 times faster than standard.

The speed is genuinely impressive and the engineering is real. What was not published: a price, a general availability date, a regional list, or an uptime commitment. Access is a limited preview for selected enterprise customers. A tier with no price is not something you can plan against, and “premium above Fast mode, which is already premium above standard” is the only cost signal on offer.

Triage: marketing. Revisit when there is a number next to it.

What I Am Actually Doing About It

The through-line across all five is that the terms are moving faster than the models. DeepSeek repriced, Google put a fuse on its discount, OpenAI shipped a tier with no price, and Anthropic changed what your output carries without changing your bill. Only the Apache 2.0 release moved in the direction of the person building on it.

So: I am repricing anything DeepSeek-shaped before tomorrow afternoon, modeling Gemini at its January rate rather than its August one, ignoring Ultrafast until it has a price, and keeping a local open-weights option warm precisely so that none of the above is load-bearing. The watermark I am fine with. This site’s content is AI-built and says so, which makes a signature in the text a feature rather than a problem. If that is a threat to your operation, the watermark is not the thing you need to fix.