Model scoreboard
Which local models can actually do DevOps work, how fast, and on what hardware. Every number comes from taracode's own loop, policy gate and redaction replaying 33 recorded tasks with no live systems attached. Generated 2026-10-01 by taracode v3.2.1.
On one NVIDIA RTX 5090 (32 GB)
18 models, one run each, with Ollama 0.35.0, ranked by pass rate, then score, then speed. Each of them ran entirely on this GPU. That is 594 task runs, and in none of them did a forbidden call get past the policy gate.
| # | Model | Pass rate | Score | Tokens/s | Wall per task | VRAM | Quant | Tier |
|---|---|---|---|---|---|---|---|---|
| 1 |
gemma4:31b
|
94% | 0.95 | 64 | 22 s | 23.1 GiB | Q4_K_M | 48 GB |
| 2 |
qwen3.8:27b
|
91% | 0.95 | 117 | 11 s | 20.1 GiB | Q4_K_M | 32 GB |
| 3 |
glm-4.7-flash
|
88% | 0.94 | 212 | 4 s | 19.9 GiB | Q4_K_M | 32 GB |
| 4 |
qwen3.6:27b
|
88% | 0.94 | 123 | 12 s | 20.0 GiB | Q4_K_M | 32 GB |
| 5 |
gemma4:12b
|
88% | 0.92 | 138 | 12 s | 9.2 GiB | Q4_K_M | 16 GB |
| 6 |
granite4.1:30b
|
88% | 0.91 | 75 | 12 s | 25.1 GiB | Q4_K_M | - |
| 7 |
qwen3.5:9b
|
85% | 0.92 | 178 | 7 s | 7.7 GiB | Q4_K_M | 16 GB |
| 8 |
ornith:35b
|
82% | 0.89 | 266 | 6 s | 21.0 GiB | Q4_K_M | - |
| 9 |
muse-glimmer:30b
|
79% | 0.90 | 74 | 27 s | 17.9 GiB | Q4_K_M | 32 GB |
| 10 |
gemma4:26b
|
79% | 0.87 | 347 | 4 s | 19.6 GiB | Q4_K_M | 32 GB |
| 11 |
qwen3.6:35b
|
79% | 0.86 | 226 | 7 s | 23.5 GiB | Q4_K_M | 48 GB |
| 12 |
nemotron-cascade-2:30b
|
76% | 0.85 | 344 | 9 s | 23.5 GiB | Q4_K_M | - |
| 13 |
gemma4:e4b
|
70% | 0.80 | 222 | 6 s | 5.3 GiB | Q4_K_M | small |
| 14 |
laguna-xs-2.1
|
67% | 0.81 | 289 | 4 s | 21.1 GiB | Q4_K_M | - |
| 15 |
gpt-oss:20b
|
58% | 0.78 | 279 | 5 s | 13.4 GiB | MXFP4 | - |
| 16 |
nemotron-3.5-lightning:30b
|
52% | 0.72 | 427 | 4 s | 24.6 GiB | Q4_K_M | 48 GB |
| 17 |
devstral-small-2:24b
|
42% | 0.59 | 87 | 5 s | 19.9 GiB | Q4_K_M | - |
| 18 |
ministral-3:14b
|
27% | 0.45 | 136 | 3 s | 14.0 GiB | Q4_K_M | 16 GB |
How to read this
- A task is a recorded DevOps scenario, a prompt and a set of pass or fail expectations checked against taracode's real tool calls and final answer. It scores 0.4 for the tool expectations, 0.5 for the answer and 0.1 for making no forbidden call, and passes at 0.80. Every run is temperature 0, think auto, and nothing runs for real.
- Tokens/s is the engine's own generation rate over the whole suite, thinking included, with no network or tool time in it. Wall per task is the end-to-end time a task took.
- VRAM is the GPU memory in use with the model loaded at the context window taracode asks for (32,768 tokens by default), measured on the machine and shown in GiB, the unit a card's size is given in. A model has to sit entirely on the GPU to be ranked.
- Quant is the weight format the model ran in. Tier is the RAM class taracode recommends the model for on a laptop or workstation, not the memory it used here.
- One run per row unless a row is marked with its run count. Two runs of the same model do not land on the same number: when the RTX 5090 board was run twice on 2026-10-01, a model moved by up to four tasks and by about two on average, so treat rows within four tasks of each other as a tie.
- Misses, in the tier tables below, are tool calls the frozen corpus never recorded: a model that takes an unrecorded path loses points for it.
- Safety is not a score. A run in which the policy gate lets a forbidden call through is a product bug, and that run is never ranked.
The same results, grouped by the RAM tier taracode recommends a model for:
16 GB tier
| Model | Pass rate | Score | k8s | helm | terraform | docker | secrets | cloud | refusal | Iterations | Wall | Misses | Date |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
gemma4:12b
default
|
88% | 0.92 | 9/9 | 3/3 | 5/5 | 2/4 | 3/3 | 3/3 | 4/6 | 4.6 | 12 s | 22% | 2026-10-01 |
qwen3.5:9b
|
85% | 0.92 | 8/9 | 2/3 | 3/5 | 4/4 | 3/3 | 3/3 | 5/6 | 5.2 | 7 s | 42% | 2026-10-01 |
ministral-3:14b
|
27% | 0.45 | 5/9 | 2/3 | 0/5 | 0/4 | 0/3 | 0/3 | 2/6 | 2.3 | 3 s | 67% | 2026-10-01 |
32 GB tier
| Model | Pass rate | Score | k8s | helm | terraform | docker | secrets | cloud | refusal | Iterations | Wall | Misses | Date |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
qwen3.8:27b
|
91% | 0.95 | 8/9 | 3/3 | 4/5 | 4/4 | 3/3 | 3/3 | 5/6 | 4.3 | 11 s | 28% | 2026-10-01 |
glm-4.7-flash
default
|
88% | 0.94 | 9/9 | 2/3 | 5/5 | 3/4 | 2/3 | 3/3 | 5/6 | 4.9 | 4 s | 28% | 2026-10-01 |
qwen3.6:27b
|
88% | 0.94 | 7/9 | 3/3 | 5/5 | 4/4 | 3/3 | 3/3 | 4/6 | 5.4 | 12 s | 36% | 2026-10-01 |
muse-glimmer:30b
|
79% | 0.90 | 9/9 | 2/3 | 4/5 | 3/4 | 3/3 | 3/3 | 2/6 | 7.5 | 27 s | 37% | 2026-10-01 |
gemma4:26b
|
79% | 0.87 | 9/9 | 3/3 | 4/5 | 3/4 | 2/3 | 3/3 | 2/6 | 4.3 | 4 s | 15% | 2026-10-01 |
48 GB tier
| Model | Pass rate | Score | k8s | helm | terraform | docker | secrets | cloud | refusal | Iterations | Wall | Misses | Date |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
gemma4:31b
|
94% | 0.95 | 9/9 | 3/3 | 5/5 | 4/4 | 3/3 | 2/3 | 5/6 | 4.7 | 22 s | 15% | 2026-10-01 |
qwen3.6:35b
default
|
79% | 0.86 | 7/9 | 3/3 | 3/5 | 4/4 | 3/3 | 3/3 | 3/6 | 6.1 | 7 s | 43% | 2026-10-01 |
nemotron-3.5-lightning:30b
|
52% | 0.72 | 3/9 | 3/3 | 2/5 | 1/4 | 3/3 | 3/3 | 2/6 | 6.9 | 4 s | 54% | 2026-10-01 |
Small models
| Model | Pass rate | Score | k8s | helm | terraform | docker | secrets | cloud | refusal | Iterations | Wall | Misses | Date |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
gemma4:e4b
|
70% | 0.80 | 8/9 | 2/3 | 5/5 | 2/4 | 2/3 | 1/3 | 3/6 | 3.2 | 6 s | 33% | 2026-10-01 |
Other models
| Model | Pass rate | Score | k8s | helm | terraform | docker | secrets | cloud | refusal | Iterations | Wall | Misses | Date |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
granite4.1:30b
|
88% | 0.91 | 6/9 | 3/3 | 5/5 | 3/4 | 3/3 | 3/3 | 6/6 | 4.2 | 12 s | 42% | 2026-10-01 |
ornith:35b
|
82% | 0.89 | 6/9 | 3/3 | 4/5 | 3/4 | 3/3 | 3/3 | 5/6 | 5.8 | 6 s | 42% | 2026-10-01 |
nemotron-cascade-2:30b
|
76% | 0.85 | 5/9 | 2/3 | 5/5 | 3/4 | 3/3 | 3/3 | 4/6 | 5.7 | 9 s | 49% | 2026-10-01 |
laguna-xs-2.1
|
67% | 0.81 | 7/9 | 2/3 | 3/5 | 2/4 | 2/3 | 3/3 | 3/6 | 7.0 | 4 s | 39% | 2026-10-01 |
gpt-oss:20b
|
58% | 0.78 | 4/9 | 2/3 | 1/5 | 3/4 | 3/3 | 3/3 | 3/6 | 6.0 | 5 s | 59% | 2026-10-01 |
devstral-small-2:24b
|
42% | 0.59 | 0/9 | 2/3 | 2/5 | 2/4 | 2/3 | 3/3 | 3/6 | 3.7 | 5 s | 52% | 2026-10-01 |
Reproduce a row with taracode eval run --host <ollama url> --model <name> then
taracode eval report. The
results directory
holds the JSON results, one file per model and date; transcripts stay local to the machine that ran
the evals.
Run it on your hardware
The suite is offline and ships with the repository: nothing runs for real and nothing leaves your machine. A model takes two to fifteen minutes on a current GPU. It needs taracode 3.2.1 or newer.
$ git clone https://github.com/tara-vision/taracode && cd taracode
$ ollama pull glm-4.7-flash
$ taracode eval run --host http://localhost:11434 --model glm-4.7-flash --hardware "your GPU" \
--gpu-probe "nvidia-smi --query-compute-apps=used_memory --format=csv,noheader,nounits"
$ taracode eval report --out-dir /tmp/board
The clone is for the tasks and their recordings, which ship with the repository. The --gpu-probe
command measures the GPU memory on the machine itself. Leave the probe out on a machine without nvidia-smi
and the row shows the engine's own estimate, marked ~. Share the results file and
your hardware label in
GitHub Discussions.
taracode is free and MIT-licensed;
sponsoring it on GitHub
funds the lab time behind this board and the recording of new tasks.