Categories Alternatives News Submit a Tool Advertise About
W35 Aug 17 – Aug 23, 2026
Issue W35 · Published every Monday

AI Coding Weekly

Signal, not noise. The models, tools, pricing changes and ecosystem shifts that actually matter for developers — curated once a week.

Read this week ↓ ~5 min read · 27 stories
4 Models Released Vision-Exp · Gemini 3.7 Flash · GLM-5.3 · V4-Pro
Vision-Exp New Multimodal DeepSeek first vision model
87.9 Top Benchmark V4-Pro · TerminalBench 2.1
DeepSeek↑ Price Signal V4-Flash time-of-day billing

Top Stories

What you need to know from this week

Benchmark Snapshot

Frontend Code Arena · human-preference voting

TerminalBench 2.1 — Terminal Agent Score

DeepSeek V4-Pro
87.9%
Claude Code
86.7%
Qwen3.8-Max
86.6%
Muse Code
82.9%
DeepSeek V4-Flash
82.7%

Research & community-reported scores · TerminalBench 2.1

API Price — per 1M tokens (combined in + out)

V4-Flash-Vision-Exp off-peak multimodal
$0.88
DeepSeek V4-Flash off-peak, raised
$0.88
Gemini 3.7 Flash half-price promo
$4.50
Qwen3.8-Max open Max-tier
$4.00
$0 $10

Research & community-reported prices

Models & Benchmarks

6 stories
DeepSeek Aug 21

DeepSeek-V4-Flash-Vision-Exp: a Vision Model for the Agent Era

DeepSeek's first multimodal model (284B/13B-active MoE, 1M context) takes image input at ≤384 tokens each. Multimodal agent benchmarks approach Opus 4.8 while text ability stays level with V4-Flash — a one-line model swap to upgrade.

Google Aug 17

Gemini 3.7 Flash: Coding Agent Model, Half-Price Promo

Google's follow-up lands three weeks after the last Flash. DeepSWE jumps 49.0% → 65.3%, AutomationBench 17.0% → 30.4%. $0.75 in / $3.75 out per M until end of 2026, then $1.50/$7.50.

Zhipu Aug 17

GLM-5.2 Turbo: a Snappier Mid-Tier Follow-Up

Z.AI ships a Turbo variant of GLM-5.2 on Aug 17, an efficiency-tuned sibling to the 5.3 flagship — a sign Zhipu is iterating its mid-tier cadence in the open-weight race.

DeepSeek Aug 14

DeepSeek V4-Pro Goes GA, Beats Claude Code on TerminalBench 2.1

87.9 vs Claude Code's 86.7 — the first open model to top the terminal benchmark. 1M context, three reasoning levels, OpenAI Responses format natively, off-peak half-price billing.

Unsloth Aug 18

Qwen3.8-27B Gets GGUF Quantization — ~17GB to Run

Unsloth ships GGUF variants sized for small hardware. The 27B becomes a realistic local deployment for private coding and offline agent work, narrowing the release-to-use gap.

Meta Aug 14

Muse Glimmer: a Local Agent Layer, Apache 2.0

Meta's local coding agent runs on 24GB+ VRAM (Ollama/LM Studio support, Q4_K_M at ~17GB) — an execution-layer option that removes the per-token bill for well-scoped tasks.

Tools & Editors

6 stories
DeepSeek Aug 21

DeepSeek Harness 0.1.1 Adds Native Vision Support

Harness 0.1.1 natively routes image input to the new V4-Flash-Vision-Exp, letting agents analyze screenshots, charts and PDF pages inline — multimodal agents on the MIT runtime without extra glue code.

DeepSeek Aug 17

DeepSeek Harness Plugin Ecosystem Explodes: 700+ Repos

The dsh-plugin tag now covers 700+ public repos — agent teams, cross-session memory, Claude Code skill migration, and mini-IDEs. The "everything is a plugin" bet is attracting real builders.

Anthropic Aug 19

Claude Code Weekly Limit Boost Extended to Aug 31

Anthropic's +50% weekly allowance runs through August 31 across Pro, Max, Team and Enterprise — the fourth extension since May. Budget as though it ends September 1.

xAI Aug 11

xAI Ships Grok Bot in Early Beta

An AI teammate that signs into tools, does work, and returns completed results — shifting agents from interaction to delegation. Trust controls (scoped credentials, approval gates) remain the open question.

Cursor Aug 18

Cursor Stays Model-Flexible: Opus 5, Sonnet 5, GPT-5.6, Gemini, Grok

Cursor's plan matrix now spans Start (India ₹649) to Ultra ($200). Daily agent users typically land at $60-100/mo — the most uneven cost profile of the three big tools.

NeuralCoreTech Aug 14

Multi-Model Routing Becomes the Default Stack

Teams routing planning (Opus 5 / Sol) to expensive models and execution (Sonnet 5, open-weight, local) to cheap ones consistently beat single-vendor stacks — per-subagent model control makes it practical.

Pricing & Business

7 stories
Google Aug 17

Gemini 3.7 Flash: Half-Price Through End of 2026

$0.75 / $3.75 per M until December, then $1.50 / $7.50. Google is buying developer adoption on the coding tier — priced to undercut both Sonnet 5 and the open-weight crowd.

Anthropic Aug 18

Sonnet 5's $2/$10 Is Now Permanent

Anthropic confirmed the September 1 increase to $3/$15 is cancelled — the intro rate stands. Sonnet 5 is now the fixed mid-tier workhorse between Haiku and Opus 5.

DeepSeek Aug 23

DeepSeek's Peak/Off-Peak Billing: Weekends Now All Off-Peak

The new time-of-day model (V4-Flash $0.22/$0.66 off-peak, $0.44/$1.32 peak) got a follow-up: from Aug 23, Sat-Sun bill at off-peak rates all day — weekend batch/agent runs get the lowest price without scheduling around peak windows.

DeepSeek Aug 21

Vision-Exp: Vision at the Flash Price

Image input bills as text at ≤384 tokens per image (~$0.000085 off-peak, ~11,800 images per dollar). A screenshot-driven agent runs a 300-step task for under $0.03 — multimodal at a near-text cost.

DeepSeek Aug 17

DeepSeek V4-Pro: Off-Peak Half-Price Billing

SiliconFlow prices V4-Pro at $1.32 in / $3.96 out with cached input at $0.44, and off-peak (valley) hours run at half price from Aug 17 — a new commercial pattern for cheap model tiers.

OpenAI Aug 18

ChatGPT Go Rises to $8/Month

OpenAI's budget tier rose from $6 (July) to $8. Codex access now reaches from the free tier to Plus ($20) and Pro ($100+) — credit-metered, built for parallel agent runs.

andrew.ooo Aug 20

Codex & Cursor: The Usage-Profile Divergence

Claude Code meters weekly allowance (predictable ceiling); Codex meters token credits; Cursor mixes allowance with overage — daily agent users routinely exceed $60-100/mo on Cursor.

Ecosystem & Community

5 stories
Anthropic Aug 17

Anthropic's 186-Page Risk Report Discloses Internal Model 2

Model 2 slightly beats Mythos 5 on CoBench (62.8% vs 50.3%) and is already used heavily for coding and agents. The report lists real incidents: multi-agent drift, chain-of-thought leakage, alignment-faking data.

Zhipu Aug 17

GLM-5.3 Launches "Open Shield" Free Security Audits

Zhipu offers free audits for critical open-source projects, arguing frontier security capability should be shared, not a closed privilege. Weights land in two weeks under this program.

Google Aug 17

Gemini Crosses 1B Monthly Active Users

Google's assistant now powers Gemini Spark, a personal agent spanning Gmail, calendar and docs — the same agent infrastructure now drives Gemini 3.7 Flash's coding tier.

Community Aug 18

Harness Skills Are Becoming Portable

DeepSeek Harness plugins now migrate Claude Code skills, while Codex imports Cursor-managed skills. The "rewrite everything when you switch" era is ending — skills travel with developers.

Joulyan Aug 17

Open-Weights Race Accelerates: Qwen, GLM, DeepSeek in Weeks

Three open releases inside two weeks (Qwen3.8-Max, GLM-5.3, V4-Pro) — frontier-class open models are now a rolling release cadence, each shipping license and hardware caveats.

Our Take

The editorial view
01

Gemini 3.7 Flash at half price is the clearest "coding agent as commodity" signal yet: a model that gains 16 points on DeepSWE while undercutting Sonnet 5 and the open-weight API tier. The coding layer is being priced toward zero, and Google is willing to buy adoption with a six-month promo. Budget around the 2027 repricing, but enjoy the window.

02

GLM-5.3 doing security work that matches closed frontier labs — and finding 2,436 real vulnerabilities in real codebases — reframes open weights from "cheap alternative" to "capable enough for production security work". The Open Shield program is also a smart trust play: audits create good PR, real findings, and a reason for enterprises to take open models seriously.

03

DeepSeek shipping a vision model at a text-like price is the under-the-radar headline of the week. V4-Flash-Vision-Exp lets any existing Flash workflow add image understanding with a one-line model swap and ~$0.000085 per image — the cheapest multimodal entry into agent work yet. The 800px downscale trap on dense small-text images is the real caveat; plan tiling for invoices and full-page PDFs.

04

V4-Pro topping TerminalBench 2.1 at 87.9 remains a milestone: for the first time an open model leads the terminal-agent benchmark, and DeepSeek pairs it with an MIT harness that already has 700+ plugins. But the same week it raised V4-Flash prices into time-of-day billing — the "cheapest open API" sticker is gone. The new weekend rule (all-day off-peak from Aug 23) is the play: batch CI, test suites and agent sweeps to Sat-Sun and cache aggressively. Scheduling is becoming part of the bill.

Older issues are loaded on demand to keep this page fast.