Categories Alternatives News Submit a Tool Advertise About
W38 Sep 7 – Sep 13, 2026
Issue W38 · Published every Monday

AI Coding Weekly

Signal, not noise. The models, tools, pricing changes and ecosystem shifts that actually matter for developers — curated once a week.

Read this week ↓ ~5 min read · 21 stories
92.8% Top Score SWE-2 · Terminal-Bench 2.1
$48B Cognition $2B Series E · Devin ARR $492M
Claude > Copilot Usage Flip JetBrains developer survey
2.2 pts Open Gap Open weights behind #1 (90.6 vs 92.8)

Top Stories

What you need to know from this week

Benchmark Snapshot

Frontend Code Arena · human-preference voting

TerminalBench 2.1 — Terminal Agent Score

SWE-2 (Cognition)
92.8%
DeepSeek V4.1 Flash
90.6%
Qwen3.8 Max
86.6%
Qwen3.8-27B (local)
73.0%

Research & community-reported scores · TerminalBench 2.1

API Price — per 1M tokens (combined in + out)

Qwen3.8-Flash ¥1/¥3 per M
~$0.56
DeepSeek V4-Flash off-peak
$0.88
Gemini 3.8 Flash promo to Dec 31
$4.50
Sonnet 5 permanent $2/$10
$6.00
$0 $10

Research & community-reported prices

Models & Benchmarks

5 stories
Cognition Sep 10

Cognition SWE-2: 92.8% on Terminal-Bench 2.1

RL-post-trained from Moonshot's Kimi K3 base (2.8T MoE, 104B active, 1M context), SWE-2 is a closed coding-agent model with the top published score on the provider-run Terminal-Bench 2.1 snapshot. Notably, the headline number surfaced first through third-party eval trackers rather than Cognition's own launch page — check the harness and version before repeating it.

DeepSeek Sep 10

DeepSeek V4.1 Flash: #2 With Open Weights

Second on the same snapshot at 90.6% with open weights and a 1M context. DeepSeek reports the figure at max effort using the minimal mode of its upcoming harness — a reminder that "score" is a function of scaffold, retry policy and token budget as much as of the model.

BenchLM / OpenAI audit Sep 10

SWE-bench Pro Is Saturating — and Under Audit

Anthropic holds the top four: Fable 5.1 (81.2%), Mythos 5 (80.3%), Fable 5 (80.0%) and Opus 5 (79.2%) — the top three sit within 1.2 points. Separately, OpenAI's July audit estimated roughly 30% of the public split is broken and withdrew its recommendation of the benchmark. One-point gaps no longer mean anything.

LLMCheck Sep 7

Qwen3.8-27B Tops the Local LLM Index

Apache 2.0, 262K native context (1M with YaRN), native image and video input, and about 19GB at 4-bit — so it runs on a 24GB Mac. 1000+ GGUF variants already exist across Ollama, llama.cpp and LM Studio. In the sub-30B class this is now the default local coding model.

LLMCheck / Alibaba Sep 7

Qwen3.8-Flash-Next: Open Weights, Restrictive Licence

The Qwen4 preview (125B total / 6B active) ships under Qwen Community 1.0, which requires a separate commercial licence to offer it as a coding or office-productivity "AI work assistant" — precisely what most builders intend. Z.AI's GLM-5.3-Flash (320B / 18B active) is the licence-friendly alternative at MIT. Licence terms, not weights, are becoming the real gate.

Tools & Editors

4 stories
GitHub Sep 8-10

GitHub Copilot: Enterprise Agent Governance in Two Days

An enterprise-managed sandbox in Copilot for JetBrains (Sep 8), then enterprise-managed permissions for Copilot agent operations (Sep 9). Combined with the Sep 10 weekly release, GitHub is building the control plane for agents — sandboxing, permissions and auditability are becoming product surfaces rather than admin footnotes.

GitHub Sep 11

Copilot Code Review Gets Auto-Resolution

The Sep 11 update adds auto-resolution and analysis improvements to Copilot code review, and VS Code agents now report into Copilot usage metrics. Practical effect: admins can finally see how much work agents are doing, and review comments can be resolved without a human round-trip.

GitHub / BenchLM Sep 10

Microsoft Retires MAI-Code-1-Flash

GitHub marked MAI-Code-1-Flash deprecated on Sep 10. Its successor, MAI-Code-1.1-Flash, sits at #22 on Terminal-Bench 2.1 at 62.9% — about 30 points behind the leaders. Microsoft's in-house code-model line keeps losing ground inside its own flagship coding product.

Anthropic Sep 10

Anthropic Ships Smart Reports for Claude Enterprise

A Sep 10 beta that analyses how a team actually uses Claude: the work getting done, what it costs, where sessions hit friction, and which repeated patterns are worth packaging as shared skills. Agent cost visibility is becoming a first-class feature — the same problem every team hits once agents run for hours unattended.

Pricing & Business

4 stories
BenchLM / LLMCheck Sep 10

The Licence Is the New Price Tag

DeepSeek V4.1 Flash is #2 on Terminal-Bench 2.1 with open weights. The catch increasingly isn't the API bill but the licence: Qwen3.8-Flash-Next needs a paid commercial licence to ship as a coding assistant, while GLM-5.3-Flash is MIT. If you build on open weights, legal review is now part of your cost model.

Cognition / 36Kr Sep 9

Cognition: ~100x ARR After the $2B Round

Devin's ARR reached $492M with 1M+ monthly active users, against a $48B post-money valuation — roughly 100x ARR. SpaceX's Cursor deal implied about 30x. The spread is the story: the market is pricing autonomous agents on measurable engineering ROI, and paying a steep premium for the leader.

AICoding price tracking Sep 7-13

No Headline Price Cuts This Week — Cache Reads Are Doing the Work

We track list prices for the major coding models weekly and found no public list-price changes in Sep 7-13. That is now normal: the mechanism that moves real invoices is cache-read pricing, because agent workloads re-read context on every turn. Anthropic's 25-45% effective cut last week came from exactly that lever, not from a new headline rate.

GitHub Sep 8-9

Copilot Metering Meets Copilot Governance

GitHub's enterprise sandbox and agent permissions landed in the same window as its broader Copilot billing and policy change (announced Aug 28). Admins are being handed metering and guardrails together — expect per-agent usage to become an explicit budget line in enterprise renewals.

Ecosystem & Community

5 stories
JetBrains Sep 9

JetBrains Survey: Claude Code Overtakes Copilot at Work

The clearest signal yet that model capability beats distribution. Copilot still has the bigger installed base and the deepest GitHub integration; Claude Code is being used more for actual work. For developers this means switching is now normal rather than exotic — and tooling decisions are being re-opened every time a flagship model ships.

Cognition / Moonshot Sep 10

Closed Frontier Products Are Built on Open Bases

Cognition's top-scoring SWE-2 is a post-train of Moonshot's open-weight Kimi K3. Add Alibaba's Qwen3.8-Flash-Next preview and Meta's Muse line, and the pattern is clear: open weights have become the substrate that closed products are fine-tuned on top of. The open ecosystem is no longer a parallel track — it is the supply chain.

BenchLM Sep 10

Terminal-Bench Splits Again: 2.1 vs 4.0

BenchLM warns that the provider-run Terminal-Bench 2.1 table — where SWE-2's 92.8% lives — is not comparable with Terminal-Bench 4.0, which changed task resources and the task set. Any "top score" claim is now meaningless without the version, harness, retry policy and token budget attached.

LLMCheck Sep 7

Frontier MoE Now Fits on a Desktop — With Caveats

LLMCheck's September report puts two frontier-class MoEs inside Mac Studio memory: GLM-5.3-Flash (3-bit, ~120GB) and Qwen3.8-Flash-Next (112GB at 4-bit). It also revised MoE speed estimates down by a 0.42 factor after community testing — routing and KV-cache traffic do not shrink with activation size.

LLMCheck Sep 7

Meta Still Hasn't Open-Sourced Muse Spark

Muse Spark 1.3 shipped Sep 2 with API and CLI access only. Meta's open-weight promise covers version 1.2 — with no date and no licence published. Its only genuinely open coding model remains July's Muse Glimmer 30B under Apache 2.0, leaving Meta well behind Qwen, GLM and DeepSeek on open releases.

Our Take

The editorial view
01

Cognition owned the week: the top Terminal-Bench 2.1 score with SWE-2, plus a $2B round at a $48B valuation on $492M ARR. Two things to keep straight. First, SWE-2 is a post-train of an open-weight base (Kimi K3) — the "closed vs open" framing is getting blurrier every month. Second, if you go looking for Cognition's products, the naming is genuinely confusing: Windsurf was rebranded Devin Desktop on Jun 2 and the Cascade agent was retired around Jul 1, so "devin vs windsurf" is a same-company comparison, not a rivalry. We cover that properly in our tools pages rather than pretending they compete.

02

The benchmark story matters more than the winner. SWE-bench Pro's top three are within 1.2 points, OpenAI estimates ~30% of the public split is broken, and Terminal-Bench has split into 2.1 and 4.0 with non-comparable task sets. Translation for buyers: leaderboard position is close to useless for procurement now. What separates tools is cost per finished task and fit with your workflow — which is exactly the ground our cost breakdowns cover.

03

The distribution flip is the strategic signal. JetBrains' survey puts Claude Code ahead of Copilot in workplace use, while GitHub spends the same week shipping enterprise agent sandboxing and permissions. Incumbency is not a moat when capability moves this fast — which means developers are re-evaluating their stack continuously, and "should I switch, and what will it cost" is permanent demand rather than a one-off decision.

Older issues are loaded on demand to keep this page fast.