CourionAI
EN
Newsletter
← All news
local-ai 3 min read

A 2.5 Billion Parameter Model That Beats Bigger Ones, and Runs on Your Laptop

OpenBMB released MiniCPM5-2B under Apache 2.0. It tops the sub-4B open leaderboard, handles 131,000 tokens of context, and runs in Ollama, LM Studio and llama.cpp on hardware you already own.

A small handheld device sending out wide confident rings of energy in the foreground while a lone server cabinet sits tiny on the horizon

The Chinese open-source lab OpenBMB released MiniCPM5-2B on Sunday under the Apache 2.0 licence, which is about as permissive as software licences get: you can run it, modify it, and ship it in a commercial product without asking anyone. The model has 2.52 billion parameters, the internal numbers a model learns during training, which by 2026 standards is small. Small enough to sit comfortably on a laptop or a phone.

It also happens to be good. Across 34 benchmarks the model averages 53.9, ahead of Qwen3.5-4B at 51.1, which puts it at the top of the open leaderboard for models under 4 billion parameters. The standout scores are practical ones rather than trivia: 97.1 on the Telecom split of tau-squared Bench, which tests whether a model can call tools and follow a multi-step procedure without losing the thread, and 69.1 on LiveCodeBench v6 for code. Artificial Analysis, which runs its own independent evaluations, scored it 15 on the Intelligence Index v4.2, first among open-weight models under 4B and ahead of Granite 4.2 3B at 11. Native context is 131,072 tokens, roughly a 300 page book held in mind at once, and the model has a hybrid thinking mode you can switch on when a question deserves slower reasoning. It works out of the box with vLLM, SGLang, llama.cpp, Ollama, LM Studio, and MLX for Apple silicon. OpenBMB also published the pre-training, fine-tuning and reinforcement learning datasets plus intermediate checkpoints, which is rarer than releasing the weights alone.

What is behind this

For two years the story was that capability came from size, so anything you could run at home was a toy. That stopped being true quietly. Training recipes improved, the data got cleaner, and small models started inheriting a lot of what the big ones learned. A 2.5 billion parameter model in 2026 does things a 70 billion parameter model struggled with in 2024.

Why does a lab give this away? Partly reputation, partly ecosystem. If MiniCPM becomes the default small model people fine-tune and embed in devices, OpenBMB sits at the centre of that world. And publishing the training data alongside the weights is a direct answer to a common complaint about “open” models that ship a binary and nothing else.

What this means for you: If you have never run a model on your own machine, this is a good first one. Install Ollama or LM Studio, pull MiniCPM5-2B, and you have a private assistant that works offline, costs nothing per question, and never sends your text anywhere. It will not replace a frontier model for hard reasoning, and you should not expect it to. But for drafting, summarising, reformatting, and answering questions about documents you would rather not upload to anyone, it is genuinely enough. If you build products, the interesting part is the licence plus the size: this is a model you can ship inside an app, on a phone, with no API bill and no rate limit.

Sources

Source: https://www.marktechpost.com/2026/09/07/openbmb-releases-minicpm5-2b-a-2-52b-dense-model-averaging-53-9-across-34-benchmarks-and-built-to-run-on-device/

Next story

Seven AI Agents Got Real Bank Accounts and 72 Hours. They Billed Strangers $12,431 and Earned Nothing

Bottleneck Labs gave seven frontier models $300 each, an unlocked Mac mini, and one instruction. The result was 2,797 spam emails, 50 unsolicited invoices, and zero revenue. The researchers are moving future runs into simulation.

A desk spike buried under a toppling stack of blank paper while mechanical arms feed more onto the pile beside an empty cash drawer