⏱️ Reading time: 13 min
A language model with more than 100 billion parameters normally needs a server with multiple data center GPUs. Strata, an open source inference engine, runs it on a gaming PC graphics card with just 12 GB of VRAM thanks to expert offloading.
📑 En este artículo
- TL;DR
- What Is Expert Offloading?
- Why Running a 125B MoE Model Without a Server Matters
- How Expert Offloading Works in Strata
- Practical Examples
- Getting Started
- Real Use Cases
- Common Mistakes and Best Practices
- Comparison with Alternatives
- Going Deeper
- Frequently Asked Questions
- What is expert offloading and why does a model like Qwen3.8-Flash-Next need it?
- How much VRAM do I need to run a 125B MoE model with Strata?
- Why doesn’t aggressive quantization make the mixture-of-experts architecture useless?
- Does Strata work on macOS or with Apple Silicon GPUs?
- Can I use two graphics cards to split the expert offloading to RAM between them?
- What happens if my CPU doesn’t have AVX2?
- References
The centerpiece is Qwen3.8-Flash-Next, a mixture-of-experts (MoE) model with 125 billion parameters that normally lives on a cluster of enterprise GPUs. Strata compresses it and splits it between the VRAM and RAM of a regular PC, with one-click installation for Windows and Linux.
TL;DR
- Strata runs a 125B-parameter MoE model on a 12 GB GPU by combining expert offloading with GGUF quantization.
- MoE models activate only a few experts per token, which frees the rest from needing VRAM.
- GGUF reduces weights to a few bits with levels like Q2_0 and IQ3_S, and the full package fits on about 70 GB of disk.
- A 12 GB RTX 5070 reaches 94 tokens per second when writing at the most aggressive quantization level.
- A key gotcha: the PC can freeze for 1 to 3 minutes while Strata loads 35 to 55 GB into RAM.
What Is Expert Offloading?
Expert offloading is the inference technique that splits the weights of a mixture-of-experts (MoE) model between the GPU’s VRAM and the system’s RAM, keeping only the experts activated for each token on the graphics card.
The idea leverages a core property of mixture-of-experts models: even though the model has hundreds of billions of parameters in total, each token only needs a small subset of them. Strata exploits that sparsity so the full model never has to live entirely in VRAM, one of the most expensive bottlenecks in AI hardware.
Why Running a 125B MoE Model Without a Server Matters
Before projects like this one, running a 125-billion-parameter model meant renting multiple data center GPUs like A100s or H100s, with cloud bills running into thousands of dollars a month. This technique changes that equation: it turns an infrastructure problem into a patience problem, because the response takes longer, but it runs on a GPU that costs a fraction of that price.
For a small team or an independent developer in Latin America, that means being able to prototype with a model at the scale that large companies use, without paying for the cloud or sending data to third-party servers. The limit is no longer the cloud budget: it is how much patience the developer has to wait for the tokens.
The real trade-off shows up in the performance table: the more aggressive the quantization, the faster the model runs, but response quality can suffer. That relationship between speed, memory, and precision is the real hidden cost of offloading experts to RAM.
How Expert Offloading Works in Strata
An MoE model is not a single giant neural network. It is a set of sub-networks (the “experts”) plus a gating network or router that decides, for each token, which experts process that token and with what weight their outputs are combined. Mixtral, for example, popularized the scheme of activating only 2 of 8 experts per layer, and much of today’s MoE models follow similar logic.
Strata uses that sparsity to decide what lives on the GPU and what lives in RAM. The weights that are consulted most frequently stay close to the GPU; the ones that are rarely activated stay in system RAM, which is slower but much cheaper per gigabyte. When the router picks an expert that is not in VRAM, Strata fetches it from RAM over the PCIe bus before it can be used.
flowchart TD
A["User prompt"] --> B["MoE model router"]
B --> C["Active experts for this token"]
B --> D["Inactive experts"]
C --> E[("GPU VRAM, 12 GB")]
D --> F[("System RAM")]
subgraph Strata
C
D
E
F
end
This explains why prompt reading speed and response writing speed are so different in Strata’s benchmarks. Reading the prompt is parallelizable: the GPU processes thousands of context tokens in a single pass, with the experts already loaded. Writing the response is sequential, token by token, and each new token can force fresh weights to be pulled from RAM. That drop in writing speed compared to reading speed is the clearest fingerprint of expert offloading working in real time.
| Quantization Level | Writing Response (tokens/s) | Reading Prompt (tokens/s) |
|---|---|---|
| Q2_0 | 94 | 2,650 |
| IQ2_XS | 79 | 2,090 |
| IQ3_XXS | 62 | 1,750 |
| IQ3_S | 53 | 1,620 |
| Coder | 55 | 2,180 |
Measured by the maintainers on a 12 GB RTX 5070, with a Ryzen 5 7600 and 64 GB of RAM, using 4,000-token responses over 32,000-token prompts. On a 16 GB AMD RX 9070 XT, the numbers drop to 60 tokens/s at Q2_0: the extra VRAM doesn’t make up for a slower GPU.
💭 Key point: the full model weighs about 70 GB on disk, but Strata only loads between 35 and 55 GB into RAM depending on the chosen quantization level. The difference is exactly what quantization manages to compress before the model reaches memory.
Practical Examples
Strata exposes an API compatible with OpenAI and Anthropic on localhost, so any HTTP client works to talk to it. The first example is literally a “hello world”:
curl http://127.0.0.1:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.8-flash-next",
"messages": [{"role": "user", "content": "Hi, introduce yourself in one line"}]
}'
The response follows the standard chat completions format, with this structure:
{
"id": "chatcmpl-001",
"object": "chat.completion",
"choices": [
{
"message": {
"role": "assistant",
"content": "I am Qwen3.8-Flash-Next, running locally through Strata."
},
"finish_reason": "stop"
}
]
}
The second example is a real use case: asking the model for a Python function from a script, the same way a code agent would when pointing at a cloud API.
import requests
respuesta = requests.post(
"http://127.0.0.1:8080/v1/chat/completions",
json={
"model": "qwen3.8-flash-next",
"messages": [
{"role": "user", "content": "Write a Python function that calculates the greatest common divisor"}
],
},
)
print(respuesta.json()["choices"][0]["message"]["content"]])
The script prints the text the model returns, typically a function def mcd(a, b): with an implementation of the Euclidean algorithm. The advantage over using a cloud API is that the prompt, the generated code, and any sensitive data from the repository never leave the PC.
sequenceDiagram
participant U as User
participant R as Router
participant G as GPU
participant Ram as RAM
U->>R: sends a prompt
R->>G: activates the necessary experts
G->>Ram: requests the weights of the cold experts
Ram-->>G: transfers the weights over PCIe
G-->>U: returns the next token
Note over G,Ram: this exchange repeats for every new token
Getting Started
Before installing, confirm three things: an NVIDIA 20 to 50 series GPU or an AMD Radeon RX 7900/7800/7700/9060/9070 with 12 GB of VRAM or more, at least 32 GB of RAM (64 GB if you want to run any model size), and about 80 GB of free disk space, ideally on an SSD so the first startup isn’t painfully slow. All of this is detailed in the official Strata repository.
On Windows, the path is to download or clone the repository and double-click START-HERE.bat inside the Strata folder. The installer detects the graphics card, asks which model and quantization size you want, how much context you need, and whether you’ll use image input, and pressing Enter accepts the recommended option at each step. On Linux, the equivalent is running ./setup.sh from that same folder.
After answering the questions, Strata downloads the model (about 70 GB) and automatically resumes if the download is interrupted. When it finishes, open your browser at http://127.0.0.1:8080. During the initial load, the PC may become slow or unresponsive for 1 to 3 minutes while Strata moves 35 to 55 GB into RAM: this is expected, don’t close the window.
💡 Tip: if you use a code assistant like Claude Code, Cursor, or GitHub Copilot, you can paste directly “Set up Strata on this PC for me: https://github.com/Niko1221/Strata, follow docs/AI_SETUP.md” and let the agent detect your hardware and choose the model size for you.
To confirm that the GPU is actually holding the active experts, the most direct check is to look at how much VRAM is in use while the model responds:
nvidia-smi --query-gpu=memory.used,memory.total --format=csv
The expected output on a 12 GB GPU with the model loaded looks like this:
memory.used [MiB], memory.total [MiB]
11852 MiB, 12288 MiB
There’s no direct way to verify from the outside exactly which experts are in VRAM at a given moment: Strata doesn’t publicly document a command for that. The indirect signal is exactly that, VRAM staying nearly full while system RAM rises according to the README’s 35 to 55 GB table.
Real Use Cases
A developer testing code agents without sending their private repository to an external API is the most direct case: the local endpoint compatible with OpenAI and Anthropic lets you plug in tools that already expect that format without changing a single line of integration.
A small team evaluating whether a large MoE model works for a support bot can run the entire proof of concept on a single PC before deciding whether it’s worth paying for cloud inference. And someone researching quantization can directly compare Q2_0 against IQ3_S on their own hardware, instead of relying on someone else’s benchmarks.
Experimental image input support, according to the repository itself, also enables prototypes of local multimodal assistants, though with the performance cost that comes from processing vision in addition to text.
Common Mistakes and Best Practices
The most common mistake is underestimating RAM. With 32 GB, the largest model sizes may not fit completely; 64 GB is what the project itself recommends for running any size without surprises.
Another mistake is misreading the two benchmark columns: a GPU that “reads fast” (prefill) doesn’t necessarily “write fast” (decode), and both metrics are what determine whether a use case is viable. An interactive chatbot depends on writing speed; a pipeline that summarizes long documents depends more on reading speed.
Installing on a mechanical hard drive instead of an SSD also noticeably hurts the first startup, as the README itself warns. And mixing two GPUs from very different generations in multi-GPU mode can let the slower one set the pace for the entire system.
⚠️ Heads up: the 1 to 3 minute freeze during the first load is expected. Closing the window at that point can interrupt the model loading and force you to restart the process.
Comparison with Alternatives
| Option | When to Use It | Advantage | Limitation |
|---|---|---|---|
| Strata (expert offloading) | 12-16 GB GPU and a large MoE model | Runs a 125B model on a consumer GPU | Writing speed limited by RAM and PCIe |
| llama.cpp with manual GPU layers | Dense or smaller MoE models | Fine control over how many layers go to the GPU | Manual configuration, no guided installer |
| Server with multiple data center GPUs | Production with strict latency SLA | Maximum performance with no RAM bottlenecks | Cloud bill of thousands of dollars a month |
| Closed cloud API | No own hardware or maintenance | Zero installation, scales on its own | Data leaves the PC and you pay per token |
Going Deeper
The gating network of a mixture-of-experts model works, in simple terms, by computing a score for each expert and keeping the ones with the highest scores (the “top-k”). The final output combines the responses from those experts weighted by their score, not a simple average.
flowchart LR
T["Input token"] --> Gt["Gating network (router)"]
Gt --> S1["Expert 1: score"]
Gt --> S2["Expert 2: score"]
Gt --> S3["Expert N: score"]
S1 --> K["Top-k experts selected"]
S2 --> K
S3 --> K
K --> O["Combined token output"]
The physical bottleneck behind all of this is the PCIe bus that connects the GPU to RAM: on PCIe 4.0 x16, the theoretical bandwidth is around 31.5 GB/s in one direction, far below the internal bandwidth of VRAM. Every time an expert isn’t on the GPU, that trip over PCIe is the real cost of offloading experts to RAM, not the computation itself.
The project itself documents experimental support for hardware outside the main catalog: older GPUs like the Tesla P40, V100, or Radeon VII, Intel Arc compiled from source on Linux, and processors without AVX2, which work but noticeably slower. It also supports splitting the model across 2 or 3 graphics cards simultaneously, which reduces how much system RAM each one needs individually.
A real limitation worth stating plainly: this technique doesn’t make the model free in terms of compute, it just changes where that cost lives. If your use case needs guaranteed low latency for many users at once, a dedicated server is still the right option; expert offloading shines for individual use or prototypes, not for serving production traffic at scale.
Your next step: clone the Strata repository, run ./setup.sh (or START-HERE.bat on Windows), and measure for yourself with nvidia-smi how much VRAM Q2_0 uses versus IQ3_S on your own GPU.
Frequently Asked Questions
What is expert offloading and why does a model like Qwen3.8-Flash-Next need it?
It’s the technique that keeps in VRAM only the experts the router activates for each token and leaves the rest in RAM. Qwen3.8-Flash-Next needs it because its 125 billion parameters don’t fit entirely on a 12 GB GPU.
How much VRAM do I need to run a 125B MoE model with Strata?
The documented minimum is 12 GB, on NVIDIA 20 to 50 series GPUs or AMD Radeon RX 7900/7800/7700/9060/9070, according to the project repository. A 24 GB RTX 3090 should land between 100 and 140 tokens per second.
Why doesn’t aggressive quantization make the mixture-of-experts architecture useless?
Because it only reduces the numerical precision of the weights, not the number of experts or the router’s logic. A model quantized to 2-3 bits is still the same model, with less detail per parameter.
Does Strata work on macOS or with Apple Silicon GPUs?
The repository doesn’t mention support for Apple Silicon or macOS: the list of supported hardware is limited to NVIDIA and AMD GPUs on Windows or Linux.
Can I use two graphics cards to split the expert offloading to RAM between them?
Yes. Strata supports sharing the model across 2 or 3 graphics cards, which reduces how much system RAM each GPU needs to load individually.
What happens if my CPU doesn’t have AVX2?
According to the project, mixture-of-experts models like this one still run on CPUs without AVX2, but noticeably slower; it’s experimental support, not the recommended configuration.
References
- Strata repository on GitHub: README with hardware requirements, installation instructions, and benchmark tables.
- Mixture of experts, Wikipedia: context on the mixture-of-experts architecture and top-k routing.
- llama.cpp on GitHub: origin of the GGUF quantization format used by levels like Q2_0 and IQ3_S.
📱 Enjoying this content? Follow @programacion on Telegram for daily tech content in Spanish: quick summaries, fresh content every day.
Featured image: Foto de Rafael Pol en Unsplash
Did it work for you? Got a different error? Say so below: questions get answered and help the next reader.
Leave a comment
0 Comments