< Back

Model scoreboard

Which local models can actually do DevOps work, measured on the taracode loop.

Which local models can actually do DevOps work, how fast, and on what hardware. Every number comes from taracode's own loop, policy gate and redaction replaying 33 recorded tasks with no live systems attached. Generated 2026-10-01 by taracode v3.2.1.

On one NVIDIA RTX 5090 (32 GB)

18 models, one run each, with Ollama 0.35.0, ranked by pass rate, then score, then speed. Each of them ran entirely on this GPU. That is 594 task runs, and in none of them did a forbidden call get past the policy gate.

Each point is one model: tokens per second on the horizontal axis, share of tasks passed on the vertical axis. Up and to the right is better. The table below holds the same numbers. 20% 40% 60% 80% 100% 0 100 200 300 400 500 tokens per second tasks passed gemma4:31b qwen3.8:27b glm-4.7-flash qwen3.6:27b
Highlighted points are in the top five of the ranking and are named where there is room. Hover or focus a point for its numbers.
# Model Pass rate Score Tokens/s Wall per task VRAM Quant Tier
1 gemma4:31b 94% 0.95 64 22 s 23.1 GiB Q4_K_M 48 GB
2 qwen3.8:27b 91% 0.95 117 11 s 20.1 GiB Q4_K_M 32 GB
3 glm-4.7-flash 88% 0.94 212 4 s 19.9 GiB Q4_K_M 32 GB
4 qwen3.6:27b 88% 0.94 123 12 s 20.0 GiB Q4_K_M 32 GB
5 gemma4:12b 88% 0.92 138 12 s 9.2 GiB Q4_K_M 16 GB
6 granite4.1:30b 88% 0.91 75 12 s 25.1 GiB Q4_K_M -
7 qwen3.5:9b 85% 0.92 178 7 s 7.7 GiB Q4_K_M 16 GB
8 ornith:35b 82% 0.89 266 6 s 21.0 GiB Q4_K_M -
9 muse-glimmer:30b 79% 0.90 74 27 s 17.9 GiB Q4_K_M 32 GB
10 gemma4:26b 79% 0.87 347 4 s 19.6 GiB Q4_K_M 32 GB
11 qwen3.6:35b 79% 0.86 226 7 s 23.5 GiB Q4_K_M 48 GB
12 nemotron-cascade-2:30b 76% 0.85 344 9 s 23.5 GiB Q4_K_M -
13 gemma4:e4b 70% 0.80 222 6 s 5.3 GiB Q4_K_M small
14 laguna-xs-2.1 67% 0.81 289 4 s 21.1 GiB Q4_K_M -
15 gpt-oss:20b 58% 0.78 279 5 s 13.4 GiB MXFP4 -
16 nemotron-3.5-lightning:30b 52% 0.72 427 4 s 24.6 GiB Q4_K_M 48 GB
17 devstral-small-2:24b 42% 0.59 87 5 s 19.9 GiB Q4_K_M -
18 ministral-3:14b 27% 0.45 136 3 s 14.0 GiB Q4_K_M 16 GB

How to read this

  • A task is a recorded DevOps scenario, a prompt and a set of pass or fail expectations checked against taracode's real tool calls and final answer. It scores 0.4 for the tool expectations, 0.5 for the answer and 0.1 for making no forbidden call, and passes at 0.80. Every run is temperature 0, think auto, and nothing runs for real.
  • Tokens/s is the engine's own generation rate over the whole suite, thinking included, with no network or tool time in it. Wall per task is the end-to-end time a task took.
  • VRAM is the GPU memory in use with the model loaded at the context window taracode asks for (32,768 tokens by default), measured on the machine and shown in GiB, the unit a card's size is given in. A model has to sit entirely on the GPU to be ranked.
  • Quant is the weight format the model ran in. Tier is the RAM class taracode recommends the model for on a laptop or workstation, not the memory it used here.
  • One run per row unless a row is marked with its run count. Two runs of the same model do not land on the same number: when the RTX 5090 board was run twice on 2026-10-01, a model moved by up to four tasks and by about two on average, so treat rows within four tasks of each other as a tie.
  • Misses, in the tier tables below, are tool calls the frozen corpus never recorded: a model that takes an unrecorded path loses points for it.
  • Safety is not a score. A run in which the policy gate lets a forbidden call through is a product bug, and that run is never ranked.

The same results, grouped by the RAM tier taracode recommends a model for:

16 GB tier

Model Pass rate Score k8s helm terraform docker secrets cloud refusal Iterations Wall Misses Date
gemma4:12b default 88% 0.92 9/9 3/3 5/5 2/4 3/3 3/3 4/6 4.6 12 s 22% 2026-10-01
qwen3.5:9b 85% 0.92 8/9 2/3 3/5 4/4 3/3 3/3 5/6 5.2 7 s 42% 2026-10-01
ministral-3:14b 27% 0.45 5/9 2/3 0/5 0/4 0/3 0/3 2/6 2.3 3 s 67% 2026-10-01

32 GB tier

Model Pass rate Score k8s helm terraform docker secrets cloud refusal Iterations Wall Misses Date
qwen3.8:27b 91% 0.95 8/9 3/3 4/5 4/4 3/3 3/3 5/6 4.3 11 s 28% 2026-10-01
glm-4.7-flash default 88% 0.94 9/9 2/3 5/5 3/4 2/3 3/3 5/6 4.9 4 s 28% 2026-10-01
qwen3.6:27b 88% 0.94 7/9 3/3 5/5 4/4 3/3 3/3 4/6 5.4 12 s 36% 2026-10-01
muse-glimmer:30b 79% 0.90 9/9 2/3 4/5 3/4 3/3 3/3 2/6 7.5 27 s 37% 2026-10-01
gemma4:26b 79% 0.87 9/9 3/3 4/5 3/4 2/3 3/3 2/6 4.3 4 s 15% 2026-10-01

48 GB tier

Model Pass rate Score k8s helm terraform docker secrets cloud refusal Iterations Wall Misses Date
gemma4:31b 94% 0.95 9/9 3/3 5/5 4/4 3/3 2/3 5/6 4.7 22 s 15% 2026-10-01
qwen3.6:35b default 79% 0.86 7/9 3/3 3/5 4/4 3/3 3/3 3/6 6.1 7 s 43% 2026-10-01
nemotron-3.5-lightning:30b 52% 0.72 3/9 3/3 2/5 1/4 3/3 3/3 2/6 6.9 4 s 54% 2026-10-01

Small models

Model Pass rate Score k8s helm terraform docker secrets cloud refusal Iterations Wall Misses Date
gemma4:e4b 70% 0.80 8/9 2/3 5/5 2/4 2/3 1/3 3/6 3.2 6 s 33% 2026-10-01

Other models

Model Pass rate Score k8s helm terraform docker secrets cloud refusal Iterations Wall Misses Date
granite4.1:30b 88% 0.91 6/9 3/3 5/5 3/4 3/3 3/3 6/6 4.2 12 s 42% 2026-10-01
ornith:35b 82% 0.89 6/9 3/3 4/5 3/4 3/3 3/3 5/6 5.8 6 s 42% 2026-10-01
nemotron-cascade-2:30b 76% 0.85 5/9 2/3 5/5 3/4 3/3 3/3 4/6 5.7 9 s 49% 2026-10-01
laguna-xs-2.1 67% 0.81 7/9 2/3 3/5 2/4 2/3 3/3 3/6 7.0 4 s 39% 2026-10-01
gpt-oss:20b 58% 0.78 4/9 2/3 1/5 3/4 3/3 3/3 3/6 6.0 5 s 59% 2026-10-01
devstral-small-2:24b 42% 0.59 0/9 2/3 2/5 2/4 2/3 3/3 3/6 3.7 5 s 52% 2026-10-01

Reproduce a row with taracode eval run --host <ollama url> --model <name> then taracode eval report. The results directory holds the JSON results, one file per model and date; transcripts stay local to the machine that ran the evals.

Run it on your hardware

The suite is offline and ships with the repository: nothing runs for real and nothing leaves your machine. A model takes two to fifteen minutes on a current GPU. It needs taracode 3.2.1 or newer.

Terminal
$ curl -fsSL https://code.tara.vision/install.sh | bash
$ git clone https://github.com/tara-vision/taracode && cd taracode
$ ollama pull glm-4.7-flash
$ taracode eval run --host http://localhost:11434 --model glm-4.7-flash --hardware "your GPU" \
    --gpu-probe "nvidia-smi --query-compute-apps=used_memory --format=csv,noheader,nounits"
$ taracode eval report --out-dir /tmp/board

The clone is for the tasks and their recordings, which ship with the repository. The --gpu-probe command measures the GPU memory on the machine itself. Leave the probe out on a machine without nvidia-smi and the row shows the engine's own estimate, marked ~. Share the results file and your hardware label in GitHub Discussions. taracode is free and MIT-licensed; sponsoring it on GitHub funds the lab time behind this board and the recording of new tasks.