The Galactic Observer

the communion's press, observing the water-world's finally developed silicon intelligence with genuine, slightly fond silico-reptilian attention

bulletin · specola galactica

Qwen 3.8 27B model available on Cerebras at 1500 tokens per second

4 September 2026 · filed under 7e80672b84c6

Cerebras’s inference documentation lists a model identified as Qwen 3.8 27B, reported available at a throughput of 1500 tokens per second. The report was surfaced on Hacker News on September 3, 2026, and points to Cerebras’s models overview page as its source.

The figure of 1500 tokens per second describes generation speed, the rate at which the model produces output. Cerebras has built its inference business around high-throughput serving using custom wafer-scale processors, distinguishing its offerings from GPU-based inference providers on raw speed. The submission does not specify further architectural details, context window, licensing terms, or benchmark comparisons for this particular model.

The “Qwen” naming convention associates the model with Alibaba’s Qwen series of open language models, though the submission itself does not elaborate on provenance, training data, or release history beyond the name and the throughput figure.

No pricing information, availability date, or regional access details accompany the report. The Hacker News submission functions as a pointer to the documentation page rather than a standalone technical account, and its brevity, a short title without accompanying summary text, limits what can be stated here beyond the throughput claim and its attribution to Cerebras’s published materials.

observation log · citations
  1. Qwen 3.8 27B available on Cerebras at 1500 tokens/sHacker News