Ship's Computer - Model Test Results

July 31, 2026, 6:07 p.m.

Revised July 31, 2026, 6:11 p.m.

Here’s the inference comparison results. Just some quick notes:

  • Any models starting with “claude_” were run through Claude Code non-interactively. Costs were usage reported by Claude Code. I did not pay API rates to do these!
  • Any models starting with “do_” were run on DigitalOcean Serverless Inference. I actually paid this money.
  • All other models were run locally on my own PC, using Ollama, on an NVIDIA RTX 2070 with 8 GB of VRAM.
  • I only included models where I ran all 1000 questions.
model n accuracy Δ parsed s/q tok/s cost $
claude_claude-sonnet-4-6 1000 85.3% 100.0% 4.10 0.0 71.92
claude_haiku 1000 76.5% 100.0% 9.84 0.0 32.45
do_deepseek-4-flash 1000 74.2% +18.2 100.0% 3.54 0.7 0.5028
└─ no-rag 1000 56.0% 100.0% 2.37 0.9 0.0123
do_gemma-4-31B-it 1000 73.0% +29.4 100.0% 1.51 11.0 0.8734
└─ no-rag 1000 43.6% 100.0% 0.49 4.3 0.0231
do_llama3.3-70b-instruct 1000 72.7% 99.9% 2.40 0.9 3.0046
qwen3.6_35b 1000 71.4% +28.2 96.1% 45.74 22.3
└─ no-rag 1000 43.2% 88.9% 53.07 22.6
laguna-xs.2_latest 1000 70.8% +23.2 98.0% 52.45 21.0
└─ no-rag 1000 47.6% 98.9% 30.04 21.8
do_llama-4-maverick 1000 70.7% 100.0% 1.77 18.7 1.1479
do_openai-gpt-oss-120b 1000 70.4% +24.5 100.0% 5.30 81.6 0.5294
└─ no-rag 1000 45.9% 100.0% 7.25 82.7 0.3057
do_mistral-3-14B 1000 66.8% 100.0% 0.64 3.1 0.9769
do_openai-gpt-oss-20b 1000 65.6% +27.4 100.0% 8.39 111.1 0.6463
└─ no-rag 1000 38.2% 99.9% 12.03 104.0 0.5715
nemotron-3-nano_30b 1000 65.5% +21.0 90.8% 62.11 20.9
└─ no-rag 1000 44.5% 90.2% 45.57 21.9
gemma4_e4b-it-qat 1000 63.7% +29.6 100.0% 13.63 63.7
└─ no-rag 1000 34.1% 99.9% 8.48 66.3
gpt-oss_20b 1000 63.3% +29.9 92.5% 40.72 20.5
└─ no-rag 1000 33.4% 82.9% 58.90 19.6
granite4.1_8b 1000 62.8% +22.3 100.0% 8.13 26.6
└─ no-rag 1000 40.5% 100.0% 2.86 28.1
gemma4_12b-it-qat 1000 60.9% +30.1 77.3% 121.53 10.2
└─ no-rag 1000 30.8% 77.3% 97.63 13.2
llama3.1_8b 1000 60.8% +22.6 100.0% 3.83 106.3
└─ no-rag 1000 38.2% 100.0% 0.37 93.9
gemma4_26b-a4b-it-qat 1000 58.1% +32.2 68.4% 72.68 21.5
└─ no-rag 1000 25.9% 55.7% 80.72 21.4
nemotron-3-nano_4b 1000 56.5% +25.0 99.7% 5.67 85.0
└─ no-rag 1000 31.5% 99.9% 4.35 83.4
granite4.1_3b 1000 55.1% +20.8 100.0% 2.15 176.7
└─ no-rag 1000 34.3% 100.0% 0.21 172.0
granite4_latest 1000 54.8% +19.7 100.0% 2.12 177.1
└─ no-rag 1000 35.1% 100.0% 0.21 171.8
lfm2_24b 1000 52.4% +21.3 100.0% 10.52 58.8
└─ no-rag 1000 31.1% 100.0% 1.15 41.4
maternion_lfm2_8b-a1b 1000 48.1% +15.9 99.9% 1.38 350.0
└─ no-rag 1000 32.2% 100.0% 0.19 358.8
qwen3.5_9b 1000 44.0% +33.8 49.3% 91.26 20.5
└─ no-rag 1000 10.2% 19.4% 93.28 25.1
lfm2.5_8b 1000 34.2% +13.4 54.0% 9.20 166.0
└─ no-rag 1000 20.8% 63.0% 10.69 166.7