The Qwen Team announced Qwen3.8-Max, a new flagship artificial intelligence model that it said contains 2.4 trillion parameters while activating 95 billion for a given workload. The company made the model available through QwenCloud and said it planned to publish the weights the following week.

The announcement described Qwen3.8-Max as the Qwen family’s most capable model and its first open-weight release at the Max scale. It positioned the system as an extension of the Qwen3.5 architecture, with improvements aimed at software development, research, workplace tasks, multimodal operation and projects that continue for extended periods. Those performance statements and test results come from Qwen’s own release material and have not been independently validated in the supplied evidence.

Qwen highlighted several autonomous coding demonstrations. In one, the team said the model worked for about 16 days on a self-evolving command-line project, producing 265 commits, 127 pull requests and 151 issues by July 30. In another, it reportedly spent about 125 hours reproducing experiments from a data-selection paper before testing 18 improvement ideas. Qwen said the resulting method gained 2.7 points over the paper’s approach on the AIME24 mathematics benchmark.

The company also entered the model into a multimodal dialogue-intent contest involving 526 human teams. According to Qwen, the system made 45 submissions within 24 hours and finished with 0.853 accuracy, ahead of 458 teams. These examples were offered to show iterative work driven by test or leaderboard feedback, rather than one-shot code generation.

Beyond coding, Qwen said it trained the model for tool-heavy professional workflows across several agent harnesses, naming QwenWork, Claude Code, Codex, OpenClaw and Hermes. Its published benchmark table reported a score of 86.6 on Terminal Bench 2.1, 67.7 on SWE-bench Pro and 74.8 on the in-house CoWorkBench. The release includes detailed evaluation notes, but some benchmarks are internal and testing configurations differ by model, limiting simple comparisons.

QwenCloud exposes Qwen3.8-Max through interfaces compatible with OpenAI chat-completions and responses specifications as well as an Anthropic-compatible API. The model supports adjustable reasoning effort at low, medium and xhigh levels, with xhigh set as the default. Qwen also published integration instructions for several coding assistants. Its examples list text and image inputs and a one-million-token context window for supported integrations.

The immediate launch is therefore a hosted-service release paired with a future open-weights commitment. Until the announced files are actually published and external evaluators reproduce the results, the most expansive capability claims remain vendor assertions. Still, the specification and demonstrations show Qwen emphasizing long-running agent work as much as conventional question answering.