Qwen 3.8 27B joins Cerebras endpoints with 1500 tokens per second speed
The addition of the Qwen 3.8 27B model to Cerebras public inference endpoints offers developers a reported processing speed of 1500 tokens per second, expanding the catalogue of high-performance AI options available on the platform.
Cerebras has added the Qwen 3.8 27B model to its public inference endpoints, citing a headline inference speed of 1500 tokens per second. The move allows developers to access the model directly through Cerebras infrastructure, targeting high-speed inference tasks that require rapid processing capabilities.
The model is now listed in the Cerebras model catalogue, which provides technical specifications for various architectures available on the platform. These specifications include details on compression, quantization, and pruning, as well as information regarding REAP pruned models.
While the 1500 tokens per second figure is presented as a key performance metric, specific benchmark conditions such as batch size, input length, and hardware configuration have not yet been fully clarified. Investors and technical users may wish to verify whether this represents a sustained average or a peak burst rate under specific conditions.
The designation of the model as "Qwen 3.8" is notable, as it differs from more commonly referenced versions such as 2.5 or 3.0. This distinction suggests a specific variant or update within the Qwen family, though the exact relationship to previous releases requires verification against official documentation.
Cerebras continues to maintain a catalogue of models on its public endpoints, positioning itself as a provider of diverse AI infrastructure. The inclusion of the Qwen 3.8 27B variant adds to the range of options available for developers seeking to leverage Cerebras hardware for their applications.
The update was highlighted in community discussions on Hacker News, drawing attention to the speed claims and the broader implications for AI inference costs and performance. As the market for AI infrastructure grows, the availability of such high-speed endpoints remains a significant factor for institutional and individual users alike.

