Qwen3.8-Max: First Open Max-Tier Model
Alibaba's 2.4T-param / 95B-active coder with ~1M context. Weights drop next week on Hugging Face and ModelScope — the strongest open model yet for repo-scale agent workflows.
Signal, not noise. The models, tools, pricing changes and ecosystem shifts that actually matter for developers — curated once a week.
Meta's first coding agent ships with Muse Spark 1.2 and a Contributor tier at $0.20 / M output — 10x cheaper than rivals. Scores 82.9% on TerminalBench 2.1, trailing only Claude Code.
Read the full analysis →Alibaba's 2.4T-param / 95B-active coder goes open-weight next week with a 1M context. The strongest open model yet targeting repo-scale agent workflows.
Claude Code (Opus 5) holds the crown. Muse Code (82.9%), DeepSeek V4-Flash via Codex (82.7%) and Codex (81.8%) are now within 5 points — the gap is closing fast.
Research & community-reported scores · TerminalBench 2.1
Research & community-reported prices
Alibaba's 2.4T-param / 95B-active coder with ~1M context. Weights drop next week on Hugging Face and ModelScope — the strongest open model yet for repo-scale agent workflows.
MIT-licensed 284B/13B-active model scoring 82.7% on TerminalBench 2.1, now available through the Responses API and usable inside Codex.
Anthropic's new GA Opus tier ($5/$25 per M) lands within 0.5% of Fable 5 on CursorBench at half the cost, with a 5-level effort dial and 1M context.
Meta's coding model was trained in lockstep with the new agent — the key to its jump in terminal-task reliability over Muse Spark 1.1.
The 2.8T open-weight model is quoted around $3/$15 per M on hosted platforms — frontier open weights now buy control and auditability, not a lower price.
OpenAI's new speed tier for GPT-5.6 Sol, replacing Priority Processing. Effective cost per unit of throughput is ~1.25x better than standard.
Handles planning, editing, and verification end-to-end with multiple parallel agents. Zero data retention — a top ask for enterprise buyers.
A multi-agent vulnerability scanner for injection, auth bypass, and more, using an adversarial voting panel to cut false positives.
Composer keeps keystroke-level costs down by not routing everything to a frontier model; cloud agents run autonomous work in the background.
From Tsinghua and ShengShu AI: workspace isolation, white-box memory, and native MCP support for running multiple coding agents safely.
Terra also dropped 20% ($2/$12); Sol unchanged. OpenAI cites kernel and harness optimizations — some done by Sol itself inside Codex.
Meta undercuts the $20/mo subscriptions of Claude Code and Codex. Contributor tier drops output to $0.20/M — over 10x cheaper.
From Sep 1 it bills at $3/$15 per M — a 50% increase over the intro $2/$10. Batch and cache prices rise too. Budget for the list price.
OpenAI's 50% promo shows on OpenRouter, but the same models cost $2.20 via Bedrock and $2.50 via Azure vs $2.00 direct. Check your provider column.
China's GCC issued guidance rejecting AI-generated code with legal significance, pushing provenance tooling and "Assisted-by" labeling.
AI identity disclosure and deepfake labeling became mandatory Aug 2, with fines up to €15M. Affects AI code review and pair-programming vendors in the EU.
DeepSeek V4-Flash ($0.42/M), MiMo-V2.5 Flash ($0.40/M) and open models on DeepInfra set the commodity floor — 5–50x under closed APIs.
Community breakdowns note the two terminal agents trade wins by task — a growing pattern of developers running two or three AI tools at once.
Meta entering with Muse Code at $0.20/M output marks the commodity moment for coding agents: the tooling is now so table-stakes that the battleground is price and data privacy. Claude Code keeps the quality crown (86.7% TerminalBench), but a 5-point gap against a 10x-cheaper entrant is a warning shot for the subscription model.
Qwen3.8-Max open-sourcing a 2.4T Max-tier model changes the ceiling for self-hosted coding. After Kimi K3, this is the second open-weight "frontier-class" release in a month — teams that care about data control now have real options, not just trade-offs.
Watch the promo trap: Luna's 80% cut and Terra's 50% promo are limited-time, Sonnet 5 jumps 50% on Sep 1, and reseller markups quietly add 10-25%. Budget on list prices, benchmark on total tokens per task — sticker rates no longer predict your bill.