⏱️ Reading time: 14 min

A language model with tens of billions of parameters takes too long, and costs too much, just to solve a zero-shot classification like deciding which support queue a ticket goes to. Jeff, an independent project published on GitHub, solves that same decision with a model of just 0.8B parameters in 22 milliseconds on an RTX PRO 6000.

📑 En este artículo
  1. TL;DR
  2. What is zero-shot classification?
  3. Why such a small model competes with the big ones
  4. How Jeff works under the hood
  5. Practical examples with the Jeff API
  6. Getting started: installing and running Jeff locally
  7. Real use cases
  8. Common mistakes and best practices
  9. Comparison: Jeff versus Jev and the base models
  10. Going deeper: what happens in that single pass
  11. Frequently Asked Questions
    1. Does Jeff do zero-shot classification with categories it never saw?
    2. Is Jeff affiliated with TypeSafe or Jev?
    3. Do I need a GPU to run Jeff?
    4. When is it worth doing your own fine-tune instead of using Jeff as-is?
    5. Does Jeff replace a large model like Jev?
    6. Does Jeff’s training data use outputs from closed models?
  12. References

The model never saw those exact categories during training, and still chooses among them with a calibrated probability, in a single forward pass.

TL;DR

  • Jeff-Qwen3.5-0.8B responds in 22 ms on an RTX PRO 6000 and 28 ms on an Apple M4 Max with MLX.
  • A voice navigation fine-tune raised accuracy from 31.7% to 95.8% in under half an hour on a single GPU.
  • The /v1/systemone endpoint accepts choice, noul, and score type questions in a single HTTP request.
  • All training runs on local hardware: no cloud GPUs and no closed-model outputs in the synthetic data.
  • Jeff-Qwen3.5-2B reaches an average score of 83.1, above the 83.0 published by Jev.

What is zero-shot classification?

Zero-shot classification is a model’s ability to assign an input to a category that was never part of its training, describing the options in natural language at query time. Jeff applies this to closed-form decisions: it returns a calibrated probability per option in a single pass of the model.

The name of its endpoint, /v1/systemone, is not accidental. The three question types it accepts (choice, noul, and score) mimic what psychology calls system 1, the fast, intuitive response, as opposed to the slow, deliberate reasoning of a large model like Jev. Jeff doesn’t generate a step-by-step explanation, it compares the described options and returns a number for each one.

Jeff responds in 22 ms on an RTX PRO 6000, without leaving the local server. Foto de Dries De Schepper en Unsplash

Why such a small model competes with the big ones

The numbers published in the repository are the central argument. Without fine-tuning, Qwen3.5-0.8B scores an average of 45.3 points across five public benchmarks. With Jeff’s fine-tune, the same base model rises to 79.1, almost double, closing in on the 83.0 that Jev reports, a considerably larger model.

The gain isn’t even. On Financial PhraseBank, a financial sentiment classification benchmark, Jeff-Qwen3.5-0.8B reaches 96.4 points, above the 77.0 published by Jev. But on BBH (Big-Bench Hard), a reasoning benchmark, Jeff stays at 64.0 against Jev’s 94.3. The pattern repeats on JudgeBench and on the hard version of JevBench, that’s where size does matter.

The choice has a cost argument beyond speed. Every call to a large model via API adds network latency and a charge per input and output token. A 0.8B model running on the same server as the application eliminates the network call entirely and only consumes the local compute of a single model pass, at the cost of losing the reasoning depth of a larger model.

That distinction matters for deciding when to use it. A zero-shot classifier like Jeff works well when the task is discriminating between well-described options, not when reasoning through a chain of steps is needed. Jeff isn’t a replacement for Jev, but a fast specialist for closed-form decisions.

How Jeff works under the hood

All training happens on local hardware, without relying on cloud GPUs. The synthetic training data is written by an open model, Qwen3.8-Flash-Next, running on two DGX Sparks. A closed model was used only to review the quality of a sample of that synthetic data, never to generate content that ended up in the training set.

The fine-tune itself runs on a single workstation GPU, an RTX PRO 6000. The 0.8B version takes about 2 hours to train, the 2B version about 3.5 hours. Final tests are run on a MacBook, to confirm the model works just as well outside the training hardware.

Jeff’s training code starts from the open AutoJev recipe, which explains why AutoJev-27B appears later in the benchmark table, though only as a reference. The final result is a checkpoint served with a local HTTP server (jeff-serve) that accepts questions in the same format Jev uses, so migrating existing code between the two doesn’t require rewriting the application logic.

flowchart TD
    A["Qwen3.8-Flash-Next on 2 DGX Sparks"] --> B["Synthetic training data"]
    B --> C["Fine-tune on 1 RTX PRO 6000"]
    C --> D["Jeff checkpoint (0.8B or 2B)"]
    D --> E["Tests on MacBook"]
    E --> F["local jeff-serve"]

Practical examples with the Jeff API

The simplest example is a single choice type question. The server receives a description of the situation and a list of numbered options as text.

curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d '{
  "model": "jeff-latest",
  "state": "The customer writes: my order arrived with the box crushed.",
  "questions": {
    "route": {"type": "choice", "instructions": "Which team should handle this?",
      "criteria": {"1": "Refunds and payments", "2": "Damaged or lost packages", "3": "Account and login"}}
  }
}'

The response includes a probability for each option, the chosen option, and a confidence level:

{
  "route": {
    "answer": "2",
    "confidence": 0.94,
    "probabilities": {"1": 0.03, "2": 0.94, "3": 0.03}
  }
}

The second example combines two independent questions in a single request: the same situation as before, plus a noul type question (yes/no, expressed as a probability) about the customer’s tone.

curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d '{
  "model": "jeff-latest",
  "state": "Refund claim: the customer says the package arrived crushed and wants their money back.",
  "questions": {
    "route": {"type": "choice", "instructions": "Which team should handle this?",
      "criteria": {"1": "Refunds and payments", "2": "Damaged or lost packages", "3": "Account and login"}},
    "angry": {"type": "noul", "instructions": "Is the customer angry?"}
  }
}'
{
  "route": {"answer": "2", "confidence": 0.91, "probabilities": {"1": 0.06, "2": 0.91, "3": 0.03}},
  "angry": {"answer": true, "probability": 0.78}
}

The third type, score, returns a point on a scale that you describe in the instructions, useful for prioritizing tickets by urgency instead of just routing them.

curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d '{
  "model": "jeff-latest",
  "state": "The customer says their account was hacked and the password has already been changed.",
  "questions": {
    "urgencia": {"type": "score", "instructions": "On a scale of 1 to 10, how urgent is this ticket?"}
  }
}'
{
  "urgencia": {"answer": 9, "confidence": 0.88}
}
sequenceDiagram
    participant C as HTTP Client
    participant J as jeff-serve
    participant M as Jeff Model
    C->>J: POST /v1/systemone (state + questions)
    J->>M: a single forward pass
    M-->>J: probability per option
    J-->>C: JSON response with answer and confidence

Getting started: installing and running Jeff locally

Before installing you need uv, the Python package and environment manager the project uses. On Windows and Linux, with an NVIDIA GPU or even just a CPU, the path is the same:

uv sync
uv run hf download mstrasser/Jeff-Qwen3.5-0.8B --local-dir checkpoints/jeff-0.8b
JEFF_CHECKPOINT=checkpoints/jeff-0.8b PORT=8765 uv run jeff-serve

The first command installs the dependencies declared in the project. The second downloads the 0.8B checkpoint from Hugging Face to a local folder. The third starts the server on port 8765, ready to receive the requests from the previous examples.

The environment variables are simple but worth keeping straight: JEFF_CHECKPOINT points to the downloaded model folder, PORT defines which port the HTTP server listens on, and JEFF_BACKEND chooses between pytorch (default) and mlx on Mac. None of them require an API key or a cloud account: the entire flow, from download to inference, runs without leaving the machine.

On a Mac with an Apple Silicon chip, the MLX backend runs considerably faster than PyTorch and only supports Qwen models: uv sync --extra mac and then JEFF_BACKEND=mlx JEFF_CHECKPOINT=checkpoints/jeff-0.8b PORT=8765 uv run jeff-serve.

💡 Tip: the MLX backend only works with Qwen3.5 checkpoints, not with Jeff-Gemma4-E2B; for that model on Mac you need to use the PyTorch backend on CPU.

Real use cases

The README describes the scope in one sentence: the options can be anything, support queues, user intents, moderation labels, voice commands, or moves in a game. The only condition is being able to describe each option in a short sentence.

  • Support queues: routing a ticket between refunds, damaged shipments, or account issues, as in the first example in this article.
  • User intents: distinguishing whether someone wants to cancel, update their data, or request a refund from a free-text message.
  • Moderation labels: flagging content as spam, harassment, or normal content before it reaches a human moderator.
  • Voice commands: the project’s own fine-tune case, choosing among a closed set of navigation actions from transcribed audio.
  • Game moves: choosing among the legal moves of a turn, as in the tests with Frogger and Doom.

The repository tests Jeff with several games, including Frogger and Doom, where the code describes the situation and the legal moves in words, and the model picks one per turn. The authors themselves note that a game isn’t the ideal zero-shot classification test, because a game’s state doesn’t resemble messy real-world data, but it’s useful for checking behavior outside standard benchmarks.

The voice navigation fine-tune went from 31.7% to 95.8% accuracy in under 30 minutes. Foto de Hitesh Choudhary en Unsplash

The most concrete case remains the project’s own voice navigation fine-tune: accuracy on validation data rose from 31.7% to 95.8% in under half an hour of training on a single GPU. It’s proof that, when zero-shot falls short, a handful of real domain examples close the gap fast.

Common mistakes and best practices

  • Using the model without fine-tuning and expecting the same result. Untrained Qwen3.5-0.8B averages 45.3 points; the same model with Jeff’s fine-tune reaches 79.1. The difference isn’t a margin of error, it’s the entire fine-tune.
  • Mixing the MLX backend with a Gemma checkpoint. The documentation is explicit: MLX is faster on Mac, but it only runs Qwen models.
  • Asking it for multi-step reasoning. On BBH, a reasoning benchmark, Jeff-Qwen3.5-0.8B scores 64.0 against Jev’s 94.3. For that kind of task, a large model is a better fit than a fast classifier.
  • Not versioning the downloaded checkpoint. The download command points to a local folder; if several services share that folder without version control, a redeploy can overwrite the checkpoint that was already in production.
  • Confusing zero-shot with untrained. Jeff is zero-shot with respect to the categories it sees in each query, not with respect to its training: the model was in fact trained, quite a bit, before it could improvise on new categories.
⚠️ Watch out: the test games and the JevBench hard benchmark show the same limit: the closer the task gets to long-form reasoning, the wider the gap with a large model like Jev.

Comparison: Jeff versus Jev and the base models

ModelParametersOverall (5 benchmarks)JevBench hardWhen to use it
Jeff-Qwen3.5-0.8B0.8B79.147.6Minimal latency, simple classification locally
Jeff-Qwen3.5-2B2B83.153.3More accuracy without leaving a single GPU
Jeff-Gemma4-E2B~2B (Gemma)81.648.6Alternative to the Qwen family
Jev (published)large closed model83.073.3Maximum accuracy and reasoning, via API
AutoJev-27B (published)27B84.970.3Reference ceiling, high compute

The row that stands out most is Jeff-Qwen3.5-2B against Jev: 83.1 versus 83.0 on the average of the five benchmarks. But the JevBench hard column corrects any hasty reading, there Jev scores 73.3 and Jeff-Qwen3.5-2B stays at 53.3. The average hides that Jeff wins on classification and loses on reasoning. AutoJev-27B is shown only as a reference, the project’s real comparison is Jeff against Jev, not against AutoJev.

Going deeper: what happens in that single pass

A normal generative model produces a response token by token: it first decides the first word, then the next, and so on until it finishes the sentence or the full JSON. That process is slow because every new token requires another pass through the entire network.

A systemone model like Jeff avoids that. Instead of generating text, it calculates in a single pass how likely each option described in the request is, and returns those probabilities directly. There’s no text to generate or parse afterward, the response is already the probability. It’s the same logic that allows classification with a single-token output in other contexts, applied here to decisions with up to 255 options per question.

The three question types (choice, noul, score) are different ways of reading those probabilities: choice distributes the probability among the listed options and picks the highest, noul collapses it to a single yes probability, score translates it to a point on the scale you described in the instructions.

flowchart TD
    Q["Question sent to Jeff"] --> T{"Question type"}
    T -->|"choice"| A["Probability per option, up to 255"]
    T -->|"noul"| B["A single yes probability"]
    T -->|"score"| C["A point on the described scale"]

The confidence field that each response returns is useful for setting a safety threshold in production: a ticket with confidence below 0.6, for example, can be escalated to human review instead of being routed automatically. It’s a practical advantage of working with calibrated probabilities instead of generated text that has to be interpreted afterward.

Several independent questions can go in the same request, as seen in the second code example, and the server answers them together in a single pass of the model, not one at a time.

Your next step: download the 0.8B checkpoint with uv run hf download mstrasser/Jeff-Qwen3.5-0.8B --local-dir checkpoints/jeff-0.8b and try the three examples in this article against your own support or moderation categories.

📬 Get new articles by email

We only email about big articles (1-2 a month).

Frequently Asked Questions

Does Jeff do zero-shot classification with categories it never saw?

Yes. The categories are described in each request, in the criteria field, and they don’t need to have appeared in the training data. That’s exactly what separates Jeff from a traditional classifier trained on a fixed list of labels.

Is Jeff affiliated with TypeSafe or Jev?

No. The repository clarifies that Jeff is an independent project that reuses the same request format as Jev, but has no relation to or backing from TypeSafe, the company behind Jev.

Do I need a GPU to run Jeff?

It’s not required. The PyTorch backend runs on both NVIDIA GPU and CPU; on Mac, the MLX backend takes advantage of the Apple Silicon chip and tends to be faster, but only with Qwen3.5 checkpoints.

When is it worth doing your own fine-tune instead of using Jeff as-is?

When zero-shot accuracy isn’t enough for your specific case. The project’s own voice navigation example rose from 31.7% to 95.8% accuracy with a fine-tune of under half an hour on a single GPU.

Does Jeff replace a large model like Jev?

Not for long-reasoning tasks: on BBH and JudgeBench, Jev remains well ahead. Jeff competes, and even wins, on classification and hallucination detection (RAGTruth), tasks closer to a closed-form decision.

Does Jeff’s training data use outputs from closed models?

Not to generate them. Data synthesis runs on Qwen3.8-Flash-Next, an open model, on two DGX Sparks. A closed model was used only to review the quality of a sample, never to produce training content.

References

📱 Enjoying this content? Follow @programacion on Telegram for daily tech content in Spanish: quick summaries, fresh content every day.

Featured image: Foto de Vishnu Mohanan en Unsplash

Did it work for you? Got a different error? Say so below: questions get answered and help the next reader.

Leave a comment

Andrés Morales

Developer and AI researcher. Writes about language models, frameworks, developer tooling, and open source releases. Covers ML papers, the tech startup ecosystem, and programming trends.

0 Comments

Leave a Reply

Avatar placeholder

Your email address will not be published. Required fields are marked *

You can include code inside <code>…</code> or, for several lines, <pre><code>…</code></pre>.

This site uses Akismet to reduce spam. Learn how your comment data is processed.