CourionAI
EN
Newsletter
← All news
open-source 3 min read

Qwen3.8-27B Is Out, and It Runs on One Gaming Graphics Card

Alibaba finally shipped the small half of its Qwen3.8 promise. The 27 billion parameter model reached number one on Hacker News because people can actually run it at home.

A small desktop computer glowing warmly on a desk beside the dark silhouette of a giant server cabinet

Alibaba published Qwen3.8-27B today, and within hours it sat at the top of Hacker News with 893 points. The reason is simple: unlike the giant Qwen3.8-Max whose weights landed three days ago, this one fits on hardware normal people own. A single RTX 4090 or a Mac Studio is enough, provided you use a compressed version of the file.

The model is dense, which means all 27 billion parameters are used for every word it produces, and it handles images as well as text. Native context, the amount of text it can keep in mind at once, is 262,144 tokens, roughly a very long book. Reported speeds in the Hacker News thread run from about 10 tokens per second on an aggressively compressed setup to around 160 on an Nvidia DGX Spark with speculative decoding switched on, a trick where a small model drafts words that the big model then checks in bulk.

The benchmark story is where care is needed. Alibaba’s own model card puts Qwen3.8-27B at 61.7 on SWE-bench Pro, a test of fixing real software issues, against a cited 53.4 for a Claude Opus configuration. That single line is doing all the work behind the “rivals Claude Opus” headlines. On the four other benchmarks in the same table, including Terminal-Bench and Humanity’s Last Exam, the small model trails, sometimes by a lot: 30.8 against 40.0 on the latter.

Here is the part the headlines skip: vendor benchmarks are marketing documents as much as measurements. Commenters pointed out that the two sides of the comparison were run with different agent harnesses, different temperature settings and different prompts, which makes a direct comparison shaky. There is also the long-running worry that labs tune models specifically to score well on famous tests. None of that makes the release unimpressive. A 27 billion parameter model you can run offline landing anywhere near a frontier commercial model would have sounded absurd a year ago.

One reported downside is worth knowing before you download: the model has a reasoning effort dial, and on its highest setting, xhigh, several users saw simple prompts take 20 to 90 minutes. Leave it on medium or low for everyday work.

What this means for you: if you have never run a model on your own machine, this is one of the better moments to try. LM Studio indexes the compressed builds, so you can pick one that fits your memory and click through the setup without touching a terminal. Nothing leaves your computer, which matters if you work with client documents or private notes. If you already use a hosted assistant for coding, keep your expectations calm: local models are getting closer, but you will still notice the difference on hard, multi-step tasks. For most people the honest summary is that the free, private option got noticeably better this week.

Sources

Source: https://huggingface.co/Qwen/Qwen3.8-27B-FP8

Next story

An AI Notetaker Left 181,874 Meetings Open to Anyone With a Free Account

A security researcher found that tl;dv, an AI meeting recorder used by over two million people, let any signed-in user list every meeting on the platform, including live calls they could join.

An empty meeting room seen through a door left ajar, with an oversized keyhole and light spilling out