⏱️ Reading time: 13 min

A language model with more than 100 billion parameters normally needs a server with multiple data center GPUs. Strata, an open source inference engine, runs it on a gaming PC graphics card with just 12 GB of VRAM thanks to expert offloading.

📑 En este artículo
  1. TL;DR
  2. What Is Expert Offloading?
  3. Why Running a 125B MoE Model Without a Server Matters
  4. How Expert Offloading Works in Strata
  5. Practical Examples
  6. Getting Started
  7. Real Use Cases
  8. Common Mistakes and Best Practices
  9. Comparison with Alternatives
  10. Going Deeper
  11. Frequently Asked Questions
    1. What is expert offloading and why does a model like Qwen3.8-Flash-Next need it?
    2. How much VRAM do I need to run a 125B MoE model with Strata?
    3. Why doesn’t aggressive quantization make the mixture-of-experts architecture useless?
    4. Does Strata work on macOS or with Apple Silicon GPUs?
    5. Can I use two graphics cards to split the expert offloading to RAM between them?
    6. What happens if my CPU doesn’t have AVX2?
  12. References

The centerpiece is Qwen3.8-Flash-Next, a mixture-of-experts (MoE) model with 125 billion parameters that normally lives on a cluster of enterprise GPUs. Strata compresses it and splits it between the VRAM and RAM of a regular PC, with one-click installation for Windows and Linux.

TL;DR

  • Strata runs a 125B-parameter MoE model on a 12 GB GPU by combining expert offloading with GGUF quantization.
  • MoE models activate only a few experts per token, which frees the rest from needing VRAM.
  • GGUF reduces weights to a few bits with levels like Q2_0 and IQ3_S, and the full package fits on about 70 GB of disk.
  • A 12 GB RTX 5070 reaches 94 tokens per second when writing at the most aggressive quantization level.
  • A key gotcha: the PC can freeze for 1 to 3 minutes while Strata loads 35 to 55 GB into RAM.

What Is Expert Offloading?

Expert offloading is the inference technique that splits the weights of a mixture-of-experts (MoE) model between the GPU’s VRAM and the system’s RAM, keeping only the experts activated for each token on the graphics card.

The idea leverages a core property of mixture-of-experts models: even though the model has hundreds of billions of parameters in total, each token only needs a small subset of them. Strata exploits that sparsity so the full model never has to live entirely in VRAM, one of the most expensive bottlenecks in AI hardware.

Why Running a 125B MoE Model Without a Server Matters

Before projects like this one, running a 125-billion-parameter model meant renting multiple data center GPUs like A100s or H100s, with cloud bills running into thousands of dollars a month. This technique changes that equation: it turns an infrastructure problem into a patience problem, because the response takes longer, but it runs on a GPU that costs a fraction of that price.

For a small team or an independent developer in Latin America, that means being able to prototype with a model at the scale that large companies use, without paying for the cloud or sending data to third-party servers. The limit is no longer the cloud budget: it is how much patience the developer has to wait for the tokens.

The real trade-off shows up in the performance table: the more aggressive the quantization, the faster the model runs, but response quality can suffer. That relationship between speed, memory, and precision is the real hidden cost of offloading experts to RAM.

Qwen3.8-Flash-Next needs at least 12 GB of VRAM, according to the Strata repository. Foto de Md Ishak Rahman en Unsplash

How Expert Offloading Works in Strata

An MoE model is not a single giant neural network. It is a set of sub-networks (the “experts”) plus a gating network or router that decides, for each token, which experts process that token and with what weight their outputs are combined. Mixtral, for example, popularized the scheme of activating only 2 of 8 experts per layer, and much of today’s MoE models follow similar logic.

Strata uses that sparsity to decide what lives on the GPU and what lives in RAM. The weights that are consulted most frequently stay close to the GPU; the ones that are rarely activated stay in system RAM, which is slower but much cheaper per gigabyte. When the router picks an expert that is not in VRAM, Strata fetches it from RAM over the PCIe bus before it can be used.

flowchart TD
    A["User prompt"] --> B["MoE model router"]
    B --> C["Active experts for this token"]
    B --> D["Inactive experts"]
    C --> E[("GPU VRAM, 12 GB")]
    D --> F[("System RAM")]
    subgraph Strata
    C
    D
    E
    F
    end

This explains why prompt reading speed and response writing speed are so different in Strata’s benchmarks. Reading the prompt is parallelizable: the GPU processes thousands of context tokens in a single pass, with the experts already loaded. Writing the response is sequential, token by token, and each new token can force fresh weights to be pulled from RAM. That drop in writing speed compared to reading speed is the clearest fingerprint of expert offloading working in real time.

Quantization LevelWriting Response (tokens/s)Reading Prompt (tokens/s)
Q2_0942,650
IQ2_XS792,090
IQ3_XXS621,750
IQ3_S531,620
Coder552,180

Measured by the maintainers on a 12 GB RTX 5070, with a Ryzen 5 7600 and 64 GB of RAM, using 4,000-token responses over 32,000-token prompts. On a 16 GB AMD RX 9070 XT, the numbers drop to 60 tokens/s at Q2_0: the extra VRAM doesn’t make up for a slower GPU.

💭 Key point: the full model weighs about 70 GB on disk, but Strata only loads between 35 and 55 GB into RAM depending on the chosen quantization level. The difference is exactly what quantization manages to compress before the model reaches memory.

Practical Examples

Strata exposes an API compatible with OpenAI and Anthropic on localhost, so any HTTP client works to talk to it. The first example is literally a “hello world”:

curl http://127.0.0.1:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.8-flash-next",
    "messages": [{"role": "user", "content": "Hi, introduce yourself in one line"}]
  }'

The response follows the standard chat completions format, with this structure:

{
  "id": "chatcmpl-001",
  "object": "chat.completion",
  "choices": [
    {
      "message": {
        "role": "assistant",
        "content": "I am Qwen3.8-Flash-Next, running locally through Strata."
      },
      "finish_reason": "stop"
    }
  ]
}

The second example is a real use case: asking the model for a Python function from a script, the same way a code agent would when pointing at a cloud API.

import requests

respuesta = requests.post(
    "http://127.0.0.1:8080/v1/chat/completions",
    json={
        "model": "qwen3.8-flash-next",
        "messages": [
            {"role": "user", "content": "Write a Python function that calculates the greatest common divisor"}
        ],
    },
)
print(respuesta.json()["choices"][0]["message"]["content"]])

The script prints the text the model returns, typically a function def mcd(a, b): with an implementation of the Euclidean algorithm. The advantage over using a cloud API is that the prompt, the generated code, and any sensitive data from the repository never leave the PC.

sequenceDiagram
    participant U as User
    participant R as Router
    participant G as GPU
    participant Ram as RAM
    U->>R: sends a prompt
    R->>G: activates the necessary experts
    G->>Ram: requests the weights of the cold experts
    Ram-->>G: transfers the weights over PCIe
    G-->>U: returns the next token
    Note over G,Ram: this exchange repeats for every new token

Getting Started

Before installing, confirm three things: an NVIDIA 20 to 50 series GPU or an AMD Radeon RX 7900/7800/7700/9060/9070 with 12 GB of VRAM or more, at least 32 GB of RAM (64 GB if you want to run any model size), and about 80 GB of free disk space, ideally on an SSD so the first startup isn’t painfully slow. All of this is detailed in the official Strata repository.

On Windows, the path is to download or clone the repository and double-click START-HERE.bat inside the Strata folder. The installer detects the graphics card, asks which model and quantization size you want, how much context you need, and whether you’ll use image input, and pressing Enter accepts the recommended option at each step. On Linux, the equivalent is running ./setup.sh from that same folder.

After answering the questions, Strata downloads the model (about 70 GB) and automatically resumes if the download is interrupted. When it finishes, open your browser at http://127.0.0.1:8080. During the initial load, the PC may become slow or unresponsive for 1 to 3 minutes while Strata moves 35 to 55 GB into RAM: this is expected, don’t close the window.

💡 Tip: if you use a code assistant like Claude Code, Cursor, or GitHub Copilot, you can paste directly “Set up Strata on this PC for me: https://github.com/Niko1221/Strata, follow docs/AI_SETUP.md” and let the agent detect your hardware and choose the model size for you.

To confirm that the GPU is actually holding the active experts, the most direct check is to look at how much VRAM is in use while the model responds:

nvidia-smi --query-gpu=memory.used,memory.total --format=csv

The expected output on a 12 GB GPU with the model loaded looks like this:

memory.used [MiB], memory.total [MiB]
11852 MiB, 12288 MiB

There’s no direct way to verify from the outside exactly which experts are in VRAM at a given moment: Strata doesn’t publicly document a command for that. The indirect signal is exactly that, VRAM staying nearly full while system RAM rises according to the README’s 35 to 55 GB table.

The Strata installer asks you to confirm the model, context size, and image input before downloading. Foto de ThisisEngineering en Unsplash

Real Use Cases

A developer testing code agents without sending their private repository to an external API is the most direct case: the local endpoint compatible with OpenAI and Anthropic lets you plug in tools that already expect that format without changing a single line of integration.

A small team evaluating whether a large MoE model works for a support bot can run the entire proof of concept on a single PC before deciding whether it’s worth paying for cloud inference. And someone researching quantization can directly compare Q2_0 against IQ3_S on their own hardware, instead of relying on someone else’s benchmarks.

Experimental image input support, according to the repository itself, also enables prototypes of local multimodal assistants, though with the performance cost that comes from processing vision in addition to text.

Common Mistakes and Best Practices

The most common mistake is underestimating RAM. With 32 GB, the largest model sizes may not fit completely; 64 GB is what the project itself recommends for running any size without surprises.

Another mistake is misreading the two benchmark columns: a GPU that “reads fast” (prefill) doesn’t necessarily “write fast” (decode), and both metrics are what determine whether a use case is viable. An interactive chatbot depends on writing speed; a pipeline that summarizes long documents depends more on reading speed.

Installing on a mechanical hard drive instead of an SSD also noticeably hurts the first startup, as the README itself warns. And mixing two GPUs from very different generations in multi-GPU mode can let the slower one set the pace for the entire system.

⚠️ Heads up: the 1 to 3 minute freeze during the first load is expected. Closing the window at that point can interrupt the model loading and force you to restart the process.

Comparison with Alternatives

OptionWhen to Use ItAdvantageLimitation
Strata (expert offloading)12-16 GB GPU and a large MoE modelRuns a 125B model on a consumer GPUWriting speed limited by RAM and PCIe
llama.cpp with manual GPU layersDense or smaller MoE modelsFine control over how many layers go to the GPUManual configuration, no guided installer
Server with multiple data center GPUsProduction with strict latency SLAMaximum performance with no RAM bottlenecksCloud bill of thousands of dollars a month
Closed cloud APINo own hardware or maintenanceZero installation, scales on its ownData leaves the PC and you pay per token

Going Deeper

The gating network of a mixture-of-experts model works, in simple terms, by computing a score for each expert and keeping the ones with the highest scores (the “top-k”). The final output combines the responses from those experts weighted by their score, not a simple average.

flowchart LR
    T["Input token"] --> Gt["Gating network (router)"]
    Gt --> S1["Expert 1: score"]
    Gt --> S2["Expert 2: score"]
    Gt --> S3["Expert N: score"]
    S1 --> K["Top-k experts selected"]
    S2 --> K
    S3 --> K
    K --> O["Combined token output"]

The physical bottleneck behind all of this is the PCIe bus that connects the GPU to RAM: on PCIe 4.0 x16, the theoretical bandwidth is around 31.5 GB/s in one direction, far below the internal bandwidth of VRAM. Every time an expert isn’t on the GPU, that trip over PCIe is the real cost of offloading experts to RAM, not the computation itself.

The project itself documents experimental support for hardware outside the main catalog: older GPUs like the Tesla P40, V100, or Radeon VII, Intel Arc compiled from source on Linux, and processors without AVX2, which work but noticeably slower. It also supports splitting the model across 2 or 3 graphics cards simultaneously, which reduces how much system RAM each one needs individually.

A real limitation worth stating plainly: this technique doesn’t make the model free in terms of compute, it just changes where that cost lives. If your use case needs guaranteed low latency for many users at once, a dedicated server is still the right option; expert offloading shines for individual use or prototypes, not for serving production traffic at scale.

Your next step: clone the Strata repository, run ./setup.sh (or START-HERE.bat on Windows), and measure for yourself with nvidia-smi how much VRAM Q2_0 uses versus IQ3_S on your own GPU.

📬 Get new articles by email

We only email about big articles (1-2 a month).

Frequently Asked Questions

What is expert offloading and why does a model like Qwen3.8-Flash-Next need it?

It’s the technique that keeps in VRAM only the experts the router activates for each token and leaves the rest in RAM. Qwen3.8-Flash-Next needs it because its 125 billion parameters don’t fit entirely on a 12 GB GPU.

How much VRAM do I need to run a 125B MoE model with Strata?

The documented minimum is 12 GB, on NVIDIA 20 to 50 series GPUs or AMD Radeon RX 7900/7800/7700/9060/9070, according to the project repository. A 24 GB RTX 3090 should land between 100 and 140 tokens per second.

Why doesn’t aggressive quantization make the mixture-of-experts architecture useless?

Because it only reduces the numerical precision of the weights, not the number of experts or the router’s logic. A model quantized to 2-3 bits is still the same model, with less detail per parameter.

Does Strata work on macOS or with Apple Silicon GPUs?

The repository doesn’t mention support for Apple Silicon or macOS: the list of supported hardware is limited to NVIDIA and AMD GPUs on Windows or Linux.

Can I use two graphics cards to split the expert offloading to RAM between them?

Yes. Strata supports sharing the model across 2 or 3 graphics cards, which reduces how much system RAM each GPU needs to load individually.

What happens if my CPU doesn’t have AVX2?

According to the project, mixture-of-experts models like this one still run on CPUs without AVX2, but noticeably slower; it’s experimental support, not the recommended configuration.

References

📱 Enjoying this content? Follow @programacion on Telegram for daily tech content in Spanish: quick summaries, fresh content every day.

Featured image: Foto de Rafael Pol en Unsplash

Did it work for you? Got a different error? Say so below: questions get answered and help the next reader.

Leave a comment
Categories: Tech NewsTutorials

Andrés Morales

Developer and AI researcher. Writes about language models, frameworks, developer tooling, and open source releases. Covers ML papers, the tech startup ecosystem, and programming trends.

0 Comments

Leave a Reply

Avatar placeholder

Your email address will not be published. Required fields are marked *

You can include code inside <code>…</code> or, for several lines, <pre><code>…</code></pre>.

This site uses Akismet to reduce spam. Learn how your comment data is processed.