{"id":198,"date":"2026-09-14T10:06:26","date_gmt":"2026-09-14T10:06:26","guid":{"rendered":"https:\/\/scoy.ai\/guides\/ai-news-roundup-september-14\/"},"modified":"2026-09-14T10:06:26","modified_gmt":"2026-09-14T10:06:26","slug":"ai-news-roundup-september-14","status":"publish","type":"post","link":"https:\/\/scoy.ai\/guides\/ai-news-roundup-september-14\/","title":{"rendered":"AI News Roundup for September 14: The Top Coding Agent Scored 38.8% on Private Code, Not 74%"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">Four things crossed my feed today, and the one that should change how you evaluate tools got reported as a leaderboard result. Here is the operator&#8217;s read on what actually matters, what is positioning, and what breaks something you already shipped.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Real-SWE Put the Best Agent at 38.8% on Code You Cannot Download<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Specific Labs published <a href=\"https:\/\/withspecific.com\/benchmarks\/real-swe\" target=\"_blank\" rel=\"noopener\">Real-SWE<\/a> this month, and the methodology is the story. Tasks come from licensed production codebases belonging to real companies, each agent runs in an isolated sandbox across eight independent trials per task, and verification uses the codebase&#8217;s own existing test suite or new tests modeled on it. That is a very different thing from resolving GitHub issues in public repos.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The pass@1 numbers: Fable 5.1 running in Claude Code at 38.8%, GPT-6 Astra in Codex CLI at 33.8%, Gemini 3.8 Flash in Gemini CLI at 31.2%, then GLM 5.3 at 28.8%, Grok 4.6 and Muse Spark 1.3 tied at 23.8%, Kimi K3 at 18.8%, and GPT-5.6 Sol at 16.2%.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Most of the coverage I saw today ran this as &#8220;Anthropic tops the coding benchmark.&#8221; That framing throws away everything useful in it. Two things are worth more than the ranking.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">First, 38.8% is a ceiling, not a victory. The best available agent, on real enterprise code, with eight shots at each task, fails roughly six out of every ten. In the same week you can find DeepSWE figures putting several of these same models in the 67% to 74% range. Both sets of numbers are real. They measure different things, and the one measuring private production code is the one that resembles your repo.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Second, look at GPT-6 Astra at 33.8% and GPT-5.6 Sol at 16.2%. Same vendor, same harness, 17.6 points apart. Real-SWE scores model-and-harness pairs rather than models, which is the honest way to do it, because that is what you actually deploy.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Here is my position: the leaderboard framing is doing real damage to how teams buy this stuff. You are not choosing a model, you are choosing a pairing, and the only number that should move your budget is the one you generate on your own codebase. I run my content engine on this category of tool every day, and the gap between a benchmark demo and a repo with eight years of undocumented conventions in it is exactly the gap Real-SWE is trying to price. Take the methodology, not the ranking.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">The Standards Body Has Been in Working Groups Since July<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Yesterday I wrote about <a href=\"https:\/\/scoy.ai\/guides\/ai-news-roundup-september-13\/\">three CEOs agreeing that the industry should slow down<\/a>. Today the reporting caught up with the machinery behind it, and it reframes the whole thing.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The Information reported that working groups from Anthropic, OpenAI, and Google DeepMind, staffed by executives below CEO level, have been meeting regularly since July to design an industry-led standards body, with the Washington Post covering the same talks. Demis Hassabis separately floated a FINRA-style US Frontier AI Standards Body on July 14. Dario Amodei is driving the effort, focused on technical testing and pre-release auditing of frontier models. Sam Altman reportedly told OpenAI employees he supports a testing and auditing body but believes the major labs will have to build it themselves without US government backing.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">So the essay everyone read on Friday was the public face of a process that started two months earlier. That is not a scandal, but it does change what you are looking at. The open question is not whether the labs want oversight. It is who holds the pen, and the three of them disagree: Anthropic leans toward government partnership, OpenAI emphasizes voluntary industry standards.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">I will say the unfashionable thing. An auditor that the audited companies fund, staff, and can leave at will is a trade association. FINRA works because it operates under statutory authority the SEC can yank. Strip that out and you have a standards body whose strongest sanction against a member shipping an unsafe model is a strongly worded disagreement. Worth building anyway, because the technical testing protocols will outlive the politics, but nobody should mistake it for regulation.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The builder consequence is quieter and more immediate: pre-release auditing adds weeks between a model being finished and a model being available. If that becomes the norm, your release cadence, your migration windows, and your deprecation calendar all stretch with it.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Altman Ruled Out a 2026 IPO and Named Safety as the Reason<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Altman told Fortune that OpenAI will not go public this year, in <a href=\"https:\/\/fortune.com\/2026\/09\/12\/sam-altman-openai-ipo-delay-ill-advised-moment-safety-concerns\/\" target=\"_blank\" rel=\"noopener\">an interview published Friday<\/a>:<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\"><p>I actually think that given everything happening with safety, right now would be an ill-advised moment to go public.<\/p><\/blockquote>\n\n\n\n<p class=\"wp-block-paragraph\">He said OpenAI will list when the business is ready and when the moment in society feels right, with 2027 now the target. TechCrunch, Axios, and Engadget all carried the same framing.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Run that against the other lab. Anthropic told investors its Q2 revenue topped $11.5 billion, up from $4.73 billion in Q1, <a href=\"https:\/\/www.cnbc.com\/2026\/08\/15\/anthropic-revenue-jumps-to-over-11point5-billion-in-q2-report.html\" target=\"_blank\" rel=\"noopener\">as CNBC reported<\/a>, with an adjusted operating profit around $559 million and a second consecutive profitable quarter, while working with Morgan Stanley, Goldman Sachs, and JPMorgan Chase toward a listing.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Two labs, one safety environment, opposite conclusions about whether now is the moment to face public markets. That does not make Altman insincere. Safety is something he has talked about consistently and he also called extinction risk unacceptable in the same interview. But when one company posts a profit and moves toward the bankers while the other cites the climate for waiting, the difference between them is legible in the financials, not only in the risk posture. Safety is a true reason. It is also the reason that requires no follow-up question about margins.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Note how the two stories interlock. Altman suggested the labs may be close to announcing a pact to slow development, which is the standards body from the section above. The governance story and the IPO story are one story: an industry collectively deciding what it owes the public, at the exact moment several of its members are deciding what they owe shareholders.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Sora 2 Is Removed in Ten Days<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Straight from <a href=\"https:\/\/developers.openai.com\/api\/docs\/deprecations\" target=\"_blank\" rel=\"noopener\">OpenAI&#8217;s deprecations page<\/a>, because this is the section where I stop analyzing and start telling you to check something.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The Videos API and the Sora 2 video generation models are removed on September 24. That covers <code>sora-2<\/code>, <code>sora-2-pro<\/code>, and the dated snapshots <code>sora-2-2025-10-06<\/code>, <code>sora-2-2025-12-08<\/code>, and <code>sora-2-pro-2025-10-06<\/code>. The announcement went out March 24, so this one has been visible for six months and will still catch people, because video generation tends to sit in a side project nobody has opened since spring.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Two more on the calendar. <code>gpt-5.4-cyber<\/code> is removed October 1, migrating to <code>gpt-5.6-cyber<\/code>, and that deprecation was announced on September 11. Then October 23 brings the large legacy batch: <code>gpt-3.5-turbo-0125<\/code>, <code>gpt-4-0613<\/code>, <code>o1<\/code>, <code>o3-mini-2025-01-31<\/code>, <code>o4-mini<\/code> and their fine-tuned variants, routing mostly to <code>gpt-5.6-sol<\/code>, <code>gpt-5.6-terra<\/code>, and <code>gpt-5.6-luna<\/code>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The <code>gpt-5.4-cyber<\/code> entry is the one I would flag. Twenty days from announcement to shutdown, against six months for Sora 2. If you inherited a fine-tune or a hardcoded model string on a retiring snapshot, the notice window you plan around is shrinking, and a quarterly audit is no longer often enough to catch it. Grep your repo for model literals this week. It takes ten minutes and it is cheaper than finding out from a production 404.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">What ties today together: the labs spent the day talking about governance while the actual governance of your stack showed up as a benchmark you should not read as a ranking and three dates on a deprecation page. The institutional story is worth watching. The dates are worth acting on today.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Real-SWE put the best coding agent at 38.8% on private enterprise code, the AI standards body predates Amodei&#8217;s essay, and Sora 2 is removed in ten days.<\/p>\n","protected":false},"author":1,"featured_media":197,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[10],"tags":[],"class_list":["post-198","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-news"],"_links":{"self":[{"href":"https:\/\/scoy.ai\/guides\/wp-json\/wp\/v2\/posts\/198","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/scoy.ai\/guides\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/scoy.ai\/guides\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/scoy.ai\/guides\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/scoy.ai\/guides\/wp-json\/wp\/v2\/comments?post=198"}],"version-history":[{"count":0,"href":"https:\/\/scoy.ai\/guides\/wp-json\/wp\/v2\/posts\/198\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/scoy.ai\/guides\/wp-json\/wp\/v2\/media\/197"}],"wp:attachment":[{"href":"https:\/\/scoy.ai\/guides\/wp-json\/wp\/v2\/media?parent=198"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/scoy.ai\/guides\/wp-json\/wp\/v2\/categories?post=198"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/scoy.ai\/guides\/wp-json\/wp\/v2\/tags?post=198"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}