# The open-weight coding frontier caught Claude, and it speaks Mandarin

URL: https://www.thedeepfeed.ai/posts/2026-06-04-open-weight-coding-frontier-caught-claude/
Category: Models
Published: 2026-06-04
Author: the-deep-feed
Tags: open-weights, coding-models, minimax, qwen, deepseek, benchmarks, china
Kind: deep

> In a three-week window, MiniMax M3, Qwen3.7-Max, DeepSeek V4, Kimi K2.6 and Nemotron 3 Ultra all claimed the frontier. The third-party leaderboards say the gap is real and small. The marquee numbers are mostly the labs grading their own homework. And the open-weight crown carries a jurisdiction risk no benchmark measures.

## TL;DR

- In one three-week window, **MiniMax M3**, **Qwen3.7-Max**, **DeepSeek V4-Pro**, **Kimi K2.6**, and NVIDIA's **Nemotron 3 Ultra** all claimed the coding frontier. On third-party leaderboards the gap to Claude is now **real and small** — Stanford HAI puts the US-China model gap at **2.7%**.
- The cleanest verified datapoint — **Qwen3.7-Max at Code Arena Elo 1541, global #4, the only non-Claude model in the top five** — is from a **closed** model. The strongest *open-weight* 'beat a proprietary model' claim (Kimi K2.6 over GPT-5.5 on SWE-Bench Pro) is **vendor-run and unverified**.
- The independent eval houses tell a colder story: NIST judged DeepSeek V4-Pro to **lag the frontier by ~8 months**, and Artificial Analysis measured its hallucination rate at **94%** when it doesn't know an answer.
- Four of the five leading open-weight coding models are **Chinese**, and they took **~51% of all OpenRouter tokens** inside 18 months. Every one carries China's **2017 National Intelligence Law** obligation — a structural, legally-confirmed risk that no benchmark measures and the API path cannot avoid.
- The spec sheet has **commoditized**: 1M context, native multimodality, agentic coding, sub-$1/M pricing are now table stakes. Differentiation is collapsing to **price, verification trust, and jurisdiction** — and only one of those three favors the open-weight Chinese tier.

For most of two years, "what is the best coding model?" had a Western answer. You picked Claude, GPT, or Gemini, you paid the per-token rate, and you accepted that the weights stayed locked in someone else's data center. In a three-week window between late April and early June 2026, that sentence stopped being true.

**MiniMax** shipped M3 on June 1, pitching the first open-weight system to combine frontier coding, a million-token context window, and native multimodality including desktop operation. Alibaba's Qwen3.7-Max [entered the global top four of Code Arena](https://eu.36kr.com/en/p/3826677900055431). DeepSeek dropped a 1.6-trillion-parameter V4-Pro under an MIT license. Moonshot's Kimi K2.6 claimed to be the first open-weight model to pass a leading proprietary model on a hard coding benchmark. NVIDIA announced Nemotron 3 Ultra at Computex, and Microsoft used its Build keynote to launch seven in-house models, one of which it says beats Claude Sonnet in blind human evaluation. Five frontier claims, three weeks, and four of the five contenders came from China.

![A photo-finish pack of runners crossing the line together, leader in red: five frontier claims in three weeks, the gap down to a stride](/post-images/2026-06-04-open-weight-coding-frontier-caught-claude/frontier-race-hero.jpg)

The euphoria was immediate and, on the surface, earned. The honest read is more interesting than the hype and colder than the skeptics. On the independent leaderboards, open weights genuinely caught Claude, and the remaining gap is small enough to argue about. But the marquee "beat the frontier" numbers everyone is quoting are, in almost every case, the labs grading their own homework. And the open-weight crown that the benchmarks award sits on top of a risk that no benchmark measures and most of the celebration ignores.

# The verified leaderboards, separated from the marketing

Start with the one number that is both verified and load-bearing, because the entire "caught Claude" narrative rests on it. On May 25, the independent Code Arena leaderboard [added `qwen3.7-max-20260517`](https://arena.ai/blog/leaderboard-changelog/), and within days the rankings showed **Qwen3.7-Max at an Elo of 1541, global rank four on the WebDev leaderboard, the only non-Claude model in the top five.** Chinese tech press [framed it as breaking into the global top two](https://eu.36kr.com/en/p/3826677900055431) by vendor, behind only Anthropic. Only Claude Opus 4.7 and 4.6 rank above it. This is the cleanest verified third-party datapoint in the entire story, and it is the strongest evidence that the gap has closed.

It also carries a complication the headlines skip: **Qwen3.7-Max is proprietary.** It is API-only, not open-weight. The single best "caught Claude" datapoint comes from a closed Chinese model, which means the cleanest proof of the open-weight thesis is, strictly speaking, not about open weights at all. That distinction matters, and almost no one drawing the celebration is making it.

For the genuinely open-weight models, the verified picture is consistent and modest. Artificial Analysis, the third-party evaluator, puts **Kimi K2.6 and Xiaomi's MiMo at an Intelligence Index of 54**, the highest of any open-weight model, against GPT-5.5 at 60 and Claude Opus 4.7 at 57. NVIDIA's Nemotron 3 Ultra lands at [an Index of 48](https://artificialanalysis.ai/articles/nvidia-nemotron-3-ultra-launch-announced), which earns it the open-weight crown only when that crown is qualified as American: the global open-weight lead belongs to the Chinese models, a point worth returning to. And MiniMax M3, on the independent Vals AI index, is the new open-weight state of the art and ranks sixth overall.

The cleanest framing of the gap comes not from a vendor but from Stanford. The [2026 AI Index put the US-China model performance gap at 2.7%](https://eu.36kr.com/en/p/3826677900055431) as of March 2026. That is the verified story: a small, real, measured gap, with the open-weight crown in Chinese hands. It is a genuinely significant shift. It is also a long way from "the benchmarks are real, open weights beat the frontier" — the line the timeline was repeating all week.

The split between what an independent evaluator confirmed and what a vendor announced runs cleanly through the entire field. Read this table as the difference between evidence and press release:

| Model | Origin | Marquee claim (who made it) | Independent verification | Status |
|---|---|---|---|---|
| Qwen3.7-Max | Alibaba (closed) | "Beats Opus 4.6 on most benchmarks" (vendor) | Code Arena Elo 1541, global #4, only non-Claude in top 5 | 🟢 Leaderboard verified |
| Kimi K2.6 | Moonshot (open) | 58.6% SWE-Bench Pro, beats GPT-5.5 at 57.7% (vendor) | Artificial Analysis Index 54, highest open-weight | 🔴 SWE claim unverified |
| MiniMax M3 | MiniMax (open) | 59% SWE-Bench Pro, "eclipses GPT-5.5" (vendor) | Vals AI: open-weight SOTA, #6 overall | 🔴 SWE claim unverified |
| DeepSeek V4-Pro | DeepSeek (open) | 80.6% SWE-bench Verified (vendor) | NIST CAISI: "lags frontier ~8 months"; 94% hallucination | 🔴 Vendor bench; independent read worse |
| Nemotron 3 Ultra | NVIDIA (open) | Top US open-weights model (vendor) | Artificial Analysis Index 48 (US-only crown) | 🟢 Partnered eval; base model only |
| MAI-Thinking-1 | Microsoft (hosted) | "Preferred to Sonnet 4.6 in blind eval" (vendor) | None published | 🔴 Vendor-claimed |

![A figure at a desk marking their own test paper with a large red check: the marquee benchmark numbers are the labs grading their own homework](/post-images/2026-06-04-open-weight-coding-frontier-caught-claude/grading-own-homework.jpg)

Every 🟢 in that table is a leaderboard placement. Every 🔴 is a coding-benchmark number the vendor reported and no one else has reproduced. The pattern is not subtle: the independent infrastructure can confirm *ranking* (this model is better than that one in aggregate human preference) but the specific "we beat the proprietary frontier on SWE-Bench Pro" numbers, the ones that became the headlines, are almost all self-graded.

# The marquee numbers are mostly self-reported

Here is where the celebration outruns the evidence. Nearly every headline number driving the "open weights passed the frontier" claim is vendor-run, and in the one case that got independent scrutiny, the scrutiny was unflattering.

MiniMax M3 is the cleanest example, because a journalist actually flagged it in the headline. The vendor's launch claimed 59.0% on SWE-Bench Pro, 66.0% on Terminal-Bench 2.1, and benchmark performance "eclipsing GPT-5.5 and Gemini 3.1 Pro." TechTimes ran the story under a title that should be stapled to every repost: ["MiniMax M3 Open-Weight Coding Model: Frontier Claims, Unverified Benchmarks."](https://www.techtimes.com/articles/317532/20260601/minimax-m3-open-weight-coding-model-frontier-claims-unverified-benchmarks.htm) The body is precise about why: at launch, "neither the weights nor the technical report had been released," with both promised "within ten days." A model whose weights have not shipped and whose technical report does not exist cannot have its benchmarks independently reproduced. The numbers are real in the sense that MiniMax ran them. They are unverified in the sense that no one else has.

Kimi K2.6 is the most-cited "open weights passed a proprietary model" claim, and it has the same problem. The headline result, **58.6% on SWE-Bench Pro, ahead of GPT-5.5 at 57.7%, "becoming the first open-weight model to surpass a leading proprietary model on that specific benchmark,"** comes [per TechTimes from "Moonshot AI's own benchmarking,"](http://www.techtimes.com/articles/317352/20260529/chinese-ai-models-lead-openrouter-traffic-coding-gains-come-china-data-risk.htm) with the explicit note that "as of early May 2026, independent third-party verification of Kimi K2.6's benchmark claims had not been published." The 0.9-point margin over GPT-5.5 is exactly the kind of result that vendor benchmarking is structurally incapable of being trusted on, because the lab that built the model chose the harness, the scaffolding, and the framing.

The discourse on X captured the commoditization of the *claim* perfectly:

> MiniMax M3 just dropped the first open-weights model combining coding, long context, and native multimodality from the ground up. 59% SWE-Bench Pro, 74.2% MCP Atlas, 66% Terminal Bench 2.1. Sparse Attention scales context to 1M tokens. Natively multimodal from step zero.
>
> — [@RoundtableSpace](https://x.com/RoundtableSpace/status/2061806288472772914), Jun 2, 2026

> An open-source model just beat GPT-5.5 on coding benchmarks. 1 million token context window. Native vision from day one. 8x cheaper than Claude Opus. MiniMax M3 dropped June 1st and it's genuinely impressive.
>
> — [@yarmalikAI](https://x.com/yarmalikAI/status/2061856538659266829), Jun 2, 2026

Both posts state the vendor numbers as settled fact. Neither notes that the weights had not shipped. That is the gap between the timeline and the truth, and it is the whole point of this piece.

Even the independent voices that did show up confirmed only the modest version. Vals AI's verdict was a *ranking*, not a benchmark endorsement:

> MiniMax just released MiniMax-M3, their first multimodal model. It is the new open-weight SOTA on the Vals Index and the Vals Multimodal Index, and #6 overall.
>
> — [@ValsAI](https://x.com/ValsAI/status/2061926305722200226), Jun 2, 2026

"Open-weight SOTA and sixth overall" is a real, independent, and impressive result. It is also a precise way of saying the model is the best of the open tier and behind five closed models — which is the verified story, not the "beats the frontier" story.

# The eval houses tell a colder story

The most important counterweight came from the one evaluator with no incentive to flatter anyone. In April 2026, the US government's [Center for AI Standards and Innovation evaluated DeepSeek V4-Pro and concluded its "capabilities lag behind the frontier by about 8 months."](https://www.nist.gov/news-events/news/2026/05/caisi-evaluation-deepseek-v4-pro) That is a striking number to set beside the launch-day reception, which treated DeepSeek's 1.6-trillion-parameter, MIT-licensed, million-token release as a frontier event:

> The last time DeepSeek did this, NVIDIA lost $600 billion in market cap in a single day. Yesterday, they did it again. DeepSeek V4-Pro dropped without warning — 1.6 trillion parameters, 1 million token context, open source, MIT license, free to download and run.
>
> — [@iam_elias1](https://x.com/iam_elias1/status/2047961635848118675), Apr 25, 2026

The model is genuinely open, genuinely large, and genuinely cheap. It also, by the only independent evaluation that exists, trails the frontier by the better part of a year. Both things are true, and only one of them made the timeline.

DeepSeek's vendor-reported coding number is **80.6% on SWE-bench Verified**, close enough to Claude that one tracker put the gap "on at least one widely-used benchmark" at 0.2 points. (A figure of 83.7% circulated in some secondary coverage; it could not be corroborated against any primary source and should be treated as wrong.) But the headline capability number sits next to a damning reliability one: on Artificial Analysis's hallucination test, **DeepSeek V4-Pro hallucinated 94% of the time when it did not know an answer**, with the evaluator noting that "when it does not know the answer, it almost always responds anyway." Kimi K2.6 came in at 39% and Claude Opus 4.7 at 36%. A coding model that confabulates almost always when uncertain is dangerous in exactly the high-stakes domains where verification is hardest, and no SWE-bench score captures it.

The independent dev testing was similarly mixed. A May 2026 Rails coding evaluation found Kimi K2.6 and DeepSeek V4-Pro reaching the top usability tier, DeepSeek only with a Claude-based adapter, while MiniMax's prior M2.7 "generated API call signatures that failed on first execution." The open-weight tier is good. The claim that it has uniformly passed the closed frontier is not what the people grading without a stake in the outcome are finding.

# The efficiency story is the one that should worry the labs

There is a result inside the open-weight surge that matters more than any single leaderboard placement, and it is about parameters, not scores. Alibaba's official Qwen account made the claim directly:

> With only 27B parameters, Qwen3.6-27B outperforms the Qwen3.5-397B-A17B (397B total / 17B active, ~15x larger!) on every major coding benchmark — including SWE-bench Verified (77.2 vs. 76.2), SWE-bench Pro (53.5 vs. 50.9), Terminal-Bench 2.0 (59.3 vs. 52.5).
>
> — [@Alibaba_Qwen](https://x.com/Alibaba_Qwen/status/2046939775924584577), Apr 22, 2026

These are vendor numbers and should be read as such. But the architectural direction they point in is the verified part, and an analyst framing of the comparison against Claude sharpens it:

> Claude Opus 4.6 has around 200B active parameters and registers 75% on SWE-bench verified. Meanwhile, Qwen 3.6 35B A3B has 3B active parameters and scores 73.4% on SWE-bench verified. Same benchmark, but Qwen has 60x fewer active parameters.
>
> — [@Mayhem4Markets](https://x.com/Mayhem4Markets/status/2050573584754463143), May 2, 2026

![A giant box and a tiny red box both reaching the same dashed height line: a 3B-active model matching a 200B one, the inference moat collapsing](/post-images/2026-06-04-open-weight-coding-frontier-caught-claude/efficiency-curve.jpg)

The specific 73.4% is unverified, but the structural claim is the one to sit with. If an open-weight model with 3 billion active parameters lands within a few points of a closed model with roughly 200 billion, the moat was never the benchmark score. It was the cost of inference, and that moat is collapsing. A model that needs a fraction of the active parameters to reach the same usability tier can be served cheaper, run on smaller hardware, and self-hosted by an enterprise that does not want to send its code anywhere. The efficiency curve, not the leaderboard, is what turns a research result into a deployment decision — and it is the metric the closed labs have the least ability to answer with a press release.

# The US counterpunch competes on different axes

The American response to the Chinese open-weight surge arrived in the same window, and it is revealing precisely because it does not try to win the price war. NVIDIA used Jensen Huang's Computex keynote on June 1 to announce **Nemotron 3 Ultra, a 550-billion-parameter mixture-of-experts model with 55 billion active parameters**, which Artificial Analysis [scored at an Intelligence Index of 48](https://artificialanalysis.ai/articles/nvidia-nemotron-3-ultra-launch-announced), calling it the strongest American open-weights model. The crown comes with two honest caveats. The first is in the framing: the title holds only when restricted to US labs, and the Index of 48 sits below the Chinese open-weight leaders at 54. The second is that Ultra ships as a base-model checkpoint, with no instruction tuning and no alignment, "intended for downstream fine-tuning and RLHF research." It is infrastructure for builders, not a drop-in Claude competitor.

Microsoft's move was stranger and more telling. At Build on June 2, Mustafa Suleyman [announced seven in-house MAI models](https://www.microsoft.ai/news/building-a-hillclimbing-machine-launching-seven-new-mai-models/), including a reasoning model, MAI-Thinking-1, that the company says "matches leading models on key software engineering benchmarks" and was "preferred to Sonnet 4.6 in blind human side-by-side evals." The benchmark claim is, like all the others, the vendor's own. But the strategic signal is the part that matters: Microsoft, OpenAI's largest backer and distributor, is building its own frontier models to reduce sole reliance on a single lab. The proprietary frontier is fragmenting from the inside at the same moment the open-weight tier is converging from the outside. The MAI models are Foundry-hosted rather than open-weight, which makes them a hedge rather than a contribution to the commons — but a hedge by Microsoft against OpenAI is its own kind of frontier news.

Both US moves share a posture. They compete on inspectable evaluation, domestic jurisdiction, and in NVIDIA's case open weights you can host yourself, rather than on the thirty-cents-per-million price point where the Chinese tier is unbeatable. That choice of battlefield is not an accident. It is a read of where the durable advantage lies once the spec sheet commoditizes.

# It speaks Mandarin, and that is the actual story

Strip away the benchmark argument and a more durable fact remains: the open-weight coding frontier is overwhelmingly Chinese, and it won the usage war while Western observers were still debating the leaderboards.

The shift is documented and large. By TechTimes' reporting, Chinese open-weight models went from [under 2% of OpenRouter traffic in late 2024 to roughly 51% of all platform tokens by April 2026](http://www.techtimes.com/articles/317352/20260529/chinese-ai-models-lead-openrouter-traffic-coding-gains-come-china-data-risk.htm), a majority of one of the largest model-routing platforms in under 18 months. OpenRouter's own COO described Chinese open-weight models as "disproportionately heavy in agentic flows run by U.S. developers." This is a real adoption signal, and notably it does not live on GitHub: the frontier open-weight repos (MiniMax M3, Qwen, DeepSeek, Kimi) ship their weights on Hugging Face and leave thin inference shells on GitHub. MiniMax M3 had 95 GitHub stars two days after launch while leading the API conversation. The usage story is an OpenRouter story, not a stars story, which is why the GitHub numbers understate it.

One builder caught the consequence of all this shipping in a single line:

> MiniMax M3 is out. multimodal, frontier coding, 1M context, API live. that's now MiniMax, DeepSeek, Kimi, Qwen, GLM all shipping coding-focused models with million-token windows in the same quarter. the spec sheet is identical across all of them at this point.
>
> — [@buildwithhassan](https://x.com/buildwithhassan/status/2061286479310008816), Jun 1, 2026

"The spec sheet is identical" is the most important sentence in the whole discourse. A million-token context, native multimodality, agentic coding, and sub-dollar pricing have stopped being differentiators and become table stakes. When five labs ship the same capability profile in one quarter, the model stops being the product and the [things around the model start to matter](/posts/2026-06-02-oss-agent-runtimes-five-wheels/): how cheap it is, whether you can trust the numbers, and whose laws govern the data you send it.

On the first, the Chinese tier wins decisively. MiniMax M3 launched at [$0.30 per million input tokens during its first week](https://www.techtimes.com/articles/317532/20260601/minimax-m3-open-weight-coding-model-frontier-claims-unverified-benchmarks.htm) against Claude Opus's roughly $5. Qwen3.7-Max runs about a third of Claude's price. On the second, as this piece has argued, the Chinese tier is weakest — the marquee numbers are self-reported. On the third, there is a risk that the entire celebration has priced at zero.

# The risk no benchmark measures

![Document and code icons crossing a red border line into a walled compound: data leaving one jurisdiction for another](/post-images/2026-06-04-open-weight-coding-frontier-caught-claude/jurisdiction-border.jpg)

China's National Intelligence Law, enacted in 2017, requires all Chinese companies to "support, assist, and cooperate with state intelligence work." It applies continuously, requires no advance request, and offers no legal pathway to refuse. It applies to Moonshot, MiniMax, DeepSeek, Zhipu, Alibaba, and Xiaomi regardless of where their weights are hosted or whether they operate a Western subsidiary. TechTimes reported it plainly in the same piece that documented the usage reversal: any prompt processed through MiniMax's API endpoint [falls under Chinese jurisdiction](http://www.techtimes.com/articles/317352/20260529/chinese-ai-models-lead-openrouter-traffic-coding-gains-come-china-data-risk.htm), and the American Enterprise Institute named MiniMax specifically, warning that users sharing code, contracts, and strategic documents are "in effect, depositing them into a Chinese government-accessible database."

The responsible framing is the one TechTimes used, and it is worth preserving exactly: no confirmed backdoor in any of these models has been found, and no documented incident of user data being shared with Chinese authorities has surfaced. The risk is not an allegation of a breach. The risk is structural and legally confirmed — it does not require a demonstrated incident to exist. And it attaches precisely to the way most developers actually touch these models, which is through the API, not by downloading and self-hosting the weights. Self-hosting mitigates the jurisdiction problem. Routing through a Chinese API endpoint, the default for the OpenRouter traffic that drove the usage reversal, does not.

This is the gap between the benchmark and the decision. A US House joint investigation into Chinese AI model risks was [announced on April 29, 2026](http://www.techtimes.com/articles/317352/20260529/chinese-ai-models-lead-openrouter-traffic-coding-gains-come-china-data-risk.htm). Enterprises in regulated finance, healthcare, and government procurement still contract directly with Anthropic, OpenAI, and the hyperscalers, and Chinese penetration in those high-error-cost domains is far below the developer-experimentation numbers. The reversal is sharpest exactly where the stakes are lowest. That is not a coincidence; it is the market pricing the jurisdiction risk that the benchmarks cannot see.

# Where this actually leaves the frontier

The contrarian position is not that the open-weight surge is hype. It is real, it is fast, and the Stanford gap of 2.7% is the honest measure of how far it has come. The contrarian position is narrower and more useful: the verified evidence supports a small, real gap with a Chinese open-weight crown, while the celebration is running on vendor numbers, ignoring an unflattering set of independent evals, and pricing a structural jurisdiction risk at zero.

The thing that has actually changed is not that one model beat another on a benchmark. It is that the spec sheet commoditized. When a million-token agentic multimodal coding model costs thirty cents per million tokens and five labs ship one in a quarter, capability stops being scarce. What becomes scarce is verification you can trust and jurisdiction you can live with — and those are the two axes where the Chinese open-weight tier, for all its leaderboard wins, is weakest. The US counterpunch, Nemotron and the MAI models, competes on exactly those axes: inspectable evals, domestic jurisdiction, and in NVIDIA's case open weights you can host yourself.

So the frontier did get caught, in the aggregate, on the public leaderboards. The interesting question for the second half of 2026 is no longer who scores highest. It is whether the buyers who matter, the regulated and the cautious and the ones who read the fine print on the National Intelligence Law, will route their most sensitive code through an endpoint they cannot audit, to save the difference between thirty cents and five dollars. The benchmarks have no answer to that. The OpenRouter traffic suggests the developers experimenting already made their choice, and the enterprise contracts suggest the people with the most to lose have made the opposite one.

## Sources

- [36kr / 新智元 — China's AI breaks into the global top two in programming](https://eu.36kr.com/en/p/3826677900055431)
- [Code Arena leaderboard changelog (qwen3.7-max-20260517 added May 25)](https://arena.ai/blog/leaderboard-changelog/)
- [MiniMax — MiniMax M3: frontier coding, 1M context, native multimodality](https://www.minimax.io/blog/minimax-m3)
- [TechTimes — MiniMax M3: frontier claims, unverified benchmarks](https://www.techtimes.com/articles/317532/20260601/minimax-m3-open-weight-coding-model-frontier-claims-unverified-benchmarks.htm)
- [TechTimes — Chinese AI models lead OpenRouter traffic; coding gains come with China data risk](http://www.techtimes.com/articles/317352/20260529/chinese-ai-models-lead-openrouter-traffic-coding-gains-come-china-data-risk.htm)
- [Hugging Face — DeepSeek-V4: a million-token context agents can use](https://huggingface.co/blog/deepseekv4)
- [NIST CAISI — evaluation of DeepSeek V4-Pro](https://www.nist.gov/news-events/news/2026/05/caisi-evaluation-deepseek-v4-pro)
- [MarkTechPost — MiniMax releases M3 with MSA architecture, 1M context](https://www.marktechpost.com/2026/06/01/minimax-releases-minimax-m3-with-msa-architecture-supporting-1m-token-context-native-multimodality-and-agentic-coding/)
- [Artificial Analysis — NVIDIA Nemotron 3 Ultra launch](https://artificialanalysis.ai/articles/nvidia-nemotron-3-ultra-launch-announced)
- [Microsoft AI — building a hill-climbing machine: seven new MAI models](https://www.microsoft.ai/news/building-a-hillclimbing-machine-launching-seven-new-mai-models/)
- [Alibaba Cloud — Qwen3.7: the agent frontier](https://www.alibabacloud.com/blog/qwen3-7-the-agent-frontier_603154)

---

Canonical: https://www.thedeepfeed.ai/posts/2026-06-04-open-weight-coding-frontier-caught-claude/
Site: https://www.thedeepfeed.ai
Full corpus: https://www.thedeepfeed.ai/llms-full.txt