Home AI News

AI News Roundup for August 26, 2026: Two Vendors Just Priced Data Residency at Exactly 10%

I went looking for what it costs to keep inference inside one country, and I found the same number twice, at two companies that did not…

Dark data center corridor lined with server racks, facing a glass wall behind which a luminous wireframe globe shows bright arcs curving between separated regional clusters

I went looking for what it costs to keep inference inside one country, and I found the same number twice, at two companies that did not coordinate on it. That, plus a video API counting down to its own shutdown in documentation that never mentions the shutdown, is the operator’s read for Wednesday.

OpenAI and AWS Both Landed on 10% for Data Residency

OpenAI’s data controls documentation says it in one sentence, with no announcement attached: “Data residency endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency.” You route through a prefixed domain, one of us.api.openai.com, eu.api.openai.com, gb.api.openai.com and seven others, and you pay a tenth more per token for the privilege of knowing where the compute happened.

Now open AWS’s model card for Grok 4.6, which reached Bedrock on August 19. It lists three inference options and prices each one. In-Region and Geo cross-Region both bill $2.20 per million input tokens and $6.60 per million output, with cache reads at $0.55. Global cross-Region, the profile that routes your request to any commercial AWS Region on earth, bills $2.00 and $6.00, with cache reads at $0.50.

Do the division. 2.20 over 2.00 is 1.10. 6.60 over 6.00 is 1.10. 0.55 over 0.50 is 1.10. Every line item, to the cent, exactly ten percent, at a vendor whose announcement post never used the words “residency premium” and buried the pricing table three clicks into a user guide.

Two different companies, no coordination, one number. Data residency stopped being a contract term you negotiate and became a line item with a published rate.

Verdict: matters. I am not claiming collusion, and I would not read one match as a conspiracy. What I would read it as is a market that has converged on a price, which is a much more useful thing to know than either data point alone. It gives you a number to check your own vendors against. If your provider charges you thirty percent to keep data in the EU, you now have grounds to ask why. If they charge you nothing, ask what “residency” means in their contract, because somebody is paying for that routing constraint and the answer is either the vendor’s margin or your definition is looser than you think. Go find the residency line in your inference bill this week and compare it to ten percent. It is a five minute exercise and it will either reassure you or start a good conversation.

Nobody Told the Sora Documentation That Sora Is Being Shut Off

OpenAI’s deprecations page schedules the entire Videos API for removal on September 24, along with sora-2, sora-2-pro and three dated snapshots. The column where a replacement model would go is empty. There is no successor named, because there is no successor.

Then read the video generation guide, which is where a developer evaluating video generation actually lands. I pulled it this morning. It documents both models, explains when to use sora-2 for iteration speed and sora-2-pro for production quality, and carries exactly one deprecation notice, which is about the remix endpoint being replaced by the edits endpoint. Nothing about the API itself ending in 29 days.

That gap is the story. A team scoping a video feature this week can read the official how-to end to end, ship against it, and discover the sunset from a 404 in late September. The consumer Sora app already went dark on April 26, so the API is the last piece standing, and it is standing on a timer.

While you are in there, the same page schedules two more waves worth writing down. October 23 takes out a dozen model families at once, including gpt-3.5-turbo-0125, gpt-4-0613, o1-2024-12-17 and o3-mini-2025-01-31, along with their fine-tuned derivatives. December 11 retires the entire original GPT-5 snapshot line from August 2025. And the Assistants API sunset I flagged on Monday is not upcoming any more. It is today.

Verdict: breaks your stack. The specific lesson is narrower than “read the docs.” It is that a vendor’s deprecation page and a vendor’s tutorial are maintained by different people on different schedules, and only one of them is telling you the truth about the future. Treat the deprecation page as the authoritative source and diff it on a cadence, because the guide you onboarded from will happily keep teaching you an API that has a date on it. If you generate video through OpenAI today, you have four weeks to pick a different provider, and no migration path is being offered to make that easier.

Google Wired Gemini Into Big Law’s Filing Cabinets

Google Cloud opened previews on August 25 for Gemini Enterprise for Legal and a financial services edition alongside it, the first industry-specific packages built on the Gemini Enterprise platform. The legal launch customers are Cleary, Freshfields, Weil, and Williams & Connolly, which is about as concentrated a slice of the top of the profession as you can assemble. Google Cloud CEO Thomas Kurian framed it around agentic AI letting lawyers “research across historical case law, build complex arguments, and automate mundane tasks.”

Skip the skills list and read the connector list, because that is where the strategy is. Legal connects to iManage and NetDocuments for document management, Docusign for contract repositories, Everlaw and RelativityOne for e-discovery, Thomson Reuters HighQ, CourtListener for federal and state opinions, and Harvey. The financial services edition ships a Google-managed Financial Research agent with more than 50 skills and 13 connectors reaching FactSet, LSEG, Moody’s, MSCI, PitchBook, S&P Global and the SEC’s EDGAR.

Verdict: matters, and the reason is that last name on the legal list. Harvey is the best funded legal AI company in the market and Google just made it a plug-in feature of a platform that also does what Harvey does. That is the same move the platform layer always makes on the application layer, and it tells you where the value is being captured. For anyone building a vertical AI product: your integration list is your moat right up until the model vendor ships the same integrations, and Google is now doing this “first in a series,” by its own description. The counterweight is that this is a preview with no published pricing and four design partners, so nobody has run it at scale yet. Watch whether it reaches general availability with the connector list intact.

Nvidia’s 1.8x Agentic Number Has No Denominator

Nvidia detailed the Vera CPU at Hot Chips 2026 this week, and ServeTheHome’s writeup from August 24 has the specifics: 88 custom Olympus cores, six chiplets, statically partitioned spatial multithreading, eight 128-bit LPDDR5X controllers. Real silicon, genuinely interesting, and clearly designed for the workload shape agents produce rather than the one batch training produces.

Then there is the performance slide. Agentic workloads come in “close to 1.8x” and data processing “closer to 1.5x.” Against what baseline? Not stated. Nvidia itself flagged these as proxy workloads. The 30x figure making the rounds is throughput against Grace Blackwell, Nvidia’s own previous generation, and it holds only at high interactivity levels, meaning it is one point on a tradeoff curve rather than a property of the chip. There are no comparisons to AMD EPYC or Intel Xeon anywhere in the deck.

Verdict: marketing. The hardware is real and the architectural bet on agent workloads is probably right. The numbers are a vendor comparing a product to its own predecessor on workloads it selected and labeled as proxies, with the denominator left off. That is not evidence, it is a slide. If you are sizing infrastructure, wait for a third party to run your actual workload, and treat any multiplier without a stated baseline as decoration.

Anthropic Will Pay $5 Million to Have Its Homework Graded

Anthropic announced on August 25 that it is funding $5 million in grants for independent wellbeing research, covering money, model access and technical support for teams building open-source evaluations of how models affect the people using them. Grantees work independently and publish everything as open source. Applications close September 21, with full-proposal invitations going out by October 5.

The honest framing is that a lab paying outsiders to measure its own product is both a real contribution and a reputational asset, and those two things are not in tension. What makes it worth your time rather than your cynicism is the deliverable. These are open-source evals any developer can run, which means the output is tooling you get to use on whatever you have shipped, not a paper you get to read. Anthropic also published its Safeguards team’s criteria for what makes a wellbeing eval rigorous, including stating plainly what counts as a pass or fail and validating automated graders against subject-matter experts.

Verdict: matters, with a deadline attached, which makes it the only item today you can act on rather than plan around. If you run a product where people talk to a model about anything personal, that eval criteria list is free methodology from a team that does this professionally, and September 21 is 26 days out.

The Short Version

Today’s through-line is that the important numbers are the ones nobody put in a headline. Ten percent is the price of geography. Twenty nine days is the runway on a video API whose own tutorial has not been told. One point eight is a multiplier with nothing underneath it. Go read the pricing tables and the deprecation pages this week rather than the announcements, because that is where the vendors write down what they actually intend to do.