Alibaba’s Qwen3.8 27B leads open-weight models on Artificial Analysis intelligence index
Released on 14 August 2026, the open-weight model utilises extended thinking reasoning and is priced at zero cost per million tokens, according to data from Artificial Analysis.
Alibaba’s Qwen3.8 27B has achieved a score of 52 on the Artificial Analysis Intelligence Index, a result that places it well above the median score of 9 for comparable open-weight models. The multimodal model, which supports both text and image inputs, was released on 14 August 2026 and is licensed under the Apache 2.0 agreement, permitting commercial use and self-hosting.
The evaluation, conducted under the v4.1.1 methodology of the Artificial Analysis Intelligence Index, assesses models across reasoning, knowledge, mathematics, and coding. The benchmark suite includes specific tests such as GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, and AA-LCR.
During the assessment, Qwen3.8 27B generated 160 million output tokens, a volume significantly higher than the median of 43 million observed in similar open-weight models. The model employs extended thinking or chain-of-thought reasoning to address complex problems before generating responses, a feature that distinguishes it within its class.
Pricing metrics indicate that the model is offered at $0.00 per one million input tokens and $0.00 per one million output tokens. This stands in contrast to the median pricing for comparable models, which sits at $0.04 for input tokens and $0.15 for output tokens. The model features a context window of 260k tokens, although some initial reports cited 256k tokens, allowing for extensive data processing in Retrieval Augmented Generation workflows.
The Artificial Analysis Intelligence Index methodology calculates a weighted average cost per task based on input, cache hit, cache write, reasoning, and answer token prices. The AA-Omniscience Index component of the evaluation measures knowledge reliability and hallucination, rewarding correct answers and penalising hallucinations on a scale from -100 to 100. Qwen3.8 27B’s performance suggests a strong capability in handling large-scale information retrieval and reasoning tasks without incurring token costs.

