Here’s the inference comparison results. Just some quick notes:
- Any models starting with “claude_” were run through Claude Code non-interactively. Costs were usage reported by Claude Code. I did not pay API rates to do these!
- Any models starting with “do_” were run on DigitalOcean Serverless Inference. I actually paid this money.
- All other models were run locally on my own PC, using Ollama, on an NVIDIA RTX 2070 with 8 GB of VRAM.
- I only included models where I ran all 1000 questions.
| model | n | accuracy | Δ | parsed | s/q | tok/s | cost $ |
|---|---|---|---|---|---|---|---|
| claude_claude-sonnet-4-6 | 1000 | 85.3% | — | 100.0% | 4.10 | 0.0 | 71.92 |
| claude_haiku | 1000 | 76.5% | — | 100.0% | 9.84 | 0.0 | 32.45 |
| do_deepseek-4-flash | 1000 | 74.2% | +18.2 | 100.0% | 3.54 | 0.7 | 0.5028 |
| └─ no-rag | 1000 | 56.0% | 100.0% | 2.37 | 0.9 | 0.0123 | |
| do_gemma-4-31B-it | 1000 | 73.0% | +29.4 | 100.0% | 1.51 | 11.0 | 0.8734 |
| └─ no-rag | 1000 | 43.6% | 100.0% | 0.49 | 4.3 | 0.0231 | |
| do_llama3.3-70b-instruct | 1000 | 72.7% | — | 99.9% | 2.40 | 0.9 | 3.0046 |
| qwen3.6_35b | 1000 | 71.4% | +28.2 | 96.1% | 45.74 | 22.3 | |
| └─ no-rag | 1000 | 43.2% | 88.9% | 53.07 | 22.6 | ||
| laguna-xs.2_latest | 1000 | 70.8% | +23.2 | 98.0% | 52.45 | 21.0 | |
| └─ no-rag | 1000 | 47.6% | 98.9% | 30.04 | 21.8 | ||
| do_llama-4-maverick | 1000 | 70.7% | — | 100.0% | 1.77 | 18.7 | 1.1479 |
| do_openai-gpt-oss-120b | 1000 | 70.4% | +24.5 | 100.0% | 5.30 | 81.6 | 0.5294 |
| └─ no-rag | 1000 | 45.9% | 100.0% | 7.25 | 82.7 | 0.3057 | |
| do_mistral-3-14B | 1000 | 66.8% | — | 100.0% | 0.64 | 3.1 | 0.9769 |
| do_openai-gpt-oss-20b | 1000 | 65.6% | +27.4 | 100.0% | 8.39 | 111.1 | 0.6463 |
| └─ no-rag | 1000 | 38.2% | 99.9% | 12.03 | 104.0 | 0.5715 | |
| nemotron-3-nano_30b | 1000 | 65.5% | +21.0 | 90.8% | 62.11 | 20.9 | |
| └─ no-rag | 1000 | 44.5% | 90.2% | 45.57 | 21.9 | ||
| gemma4_e4b-it-qat | 1000 | 63.7% | +29.6 | 100.0% | 13.63 | 63.7 | |
| └─ no-rag | 1000 | 34.1% | 99.9% | 8.48 | 66.3 | ||
| gpt-oss_20b | 1000 | 63.3% | +29.9 | 92.5% | 40.72 | 20.5 | |
| └─ no-rag | 1000 | 33.4% | 82.9% | 58.90 | 19.6 | ||
| granite4.1_8b | 1000 | 62.8% | +22.3 | 100.0% | 8.13 | 26.6 | |
| └─ no-rag | 1000 | 40.5% | 100.0% | 2.86 | 28.1 | ||
| gemma4_12b-it-qat | 1000 | 60.9% | +30.1 | 77.3% | 121.53 | 10.2 | |
| └─ no-rag | 1000 | 30.8% | 77.3% | 97.63 | 13.2 | ||
| llama3.1_8b | 1000 | 60.8% | +22.6 | 100.0% | 3.83 | 106.3 | |
| └─ no-rag | 1000 | 38.2% | 100.0% | 0.37 | 93.9 | ||
| gemma4_26b-a4b-it-qat | 1000 | 58.1% | +32.2 | 68.4% | 72.68 | 21.5 | |
| └─ no-rag | 1000 | 25.9% | 55.7% | 80.72 | 21.4 | ||
| nemotron-3-nano_4b | 1000 | 56.5% | +25.0 | 99.7% | 5.67 | 85.0 | |
| └─ no-rag | 1000 | 31.5% | 99.9% | 4.35 | 83.4 | ||
| granite4.1_3b | 1000 | 55.1% | +20.8 | 100.0% | 2.15 | 176.7 | |
| └─ no-rag | 1000 | 34.3% | 100.0% | 0.21 | 172.0 | ||
| granite4_latest | 1000 | 54.8% | +19.7 | 100.0% | 2.12 | 177.1 | |
| └─ no-rag | 1000 | 35.1% | 100.0% | 0.21 | 171.8 | ||
| lfm2_24b | 1000 | 52.4% | +21.3 | 100.0% | 10.52 | 58.8 | |
| └─ no-rag | 1000 | 31.1% | 100.0% | 1.15 | 41.4 | ||
| maternion_lfm2_8b-a1b | 1000 | 48.1% | +15.9 | 99.9% | 1.38 | 350.0 | |
| └─ no-rag | 1000 | 32.2% | 100.0% | 0.19 | 358.8 | ||
| qwen3.5_9b | 1000 | 44.0% | +33.8 | 49.3% | 91.26 | 20.5 | |
| └─ no-rag | 1000 | 10.2% | 19.4% | 93.28 | 25.1 | ||
| lfm2.5_8b | 1000 | 34.2% | +13.4 | 54.0% | 9.20 | 166.0 | |
| └─ no-rag | 1000 | 20.8% | 63.0% | 10.69 | 166.7 |