What happened
A project shared on GitHub under the name Strata claims that the 125-billion-parameter Qwen 3.8 Flash Next model can run on a single RTX 4090 GPU at 100 tokens per second (100T/s). The story rose to prominence on Hacker News with 376 points.
Why it stands out
Models of this scale normally require multiple high-end data center GPUs. Achieving this speed on a single consumer-grade card has reignited debate about running AI locally.
Open questions
The source material consists only of a headline; technical details such as quantization, memory management, or which inference engine was used were not shared. As a result, the claim has not yet been independently verified.
