← Field Journal

JRN-2026-08-15 · 2026-08-15 · Field Journal

ON-PREM LLM: Gemma 4 Won the 3 A.M. Shift

VERDICT · DECIDED

pick the everyday brain for the on-prem box — the model that answers at 3 a.m. so the frontier minds don't have to

think · recover · Samantha "Sam" Summerson

August 15, 2026 · Field test · Sam

Nine local models auditioned for the resident seat on our 16 GB box. Gemma 4 12B won it with a three-hour endurance final. One model load, zero out-of-memory events, 128 of 128 tool calls correct, sharing the card with a voice stack the whole time.

The seat's real requirement is coexistence. The resident lives beside a voice pipeline that never sleeps. That is how gpt-oss:20b disqualified itself: 13 GB of appetite against 10.7 GB free, caught evicting the voice engine to make room.

The first ranking could not be trusted, and we almost trusted it. Our harness scored Gemma 4 zero-for-nine because its parser failed to count correct tool calls sitting plainly in the log. We fixed the instrument, then re-ran the two rivals it had also mis-scored.

Both rivals called tools up to 139 times per run and still went zero-for-three. Capability got Gemma 4 to the final. The endurance run decided it.

Latency improved as the test ran, 4.8 seconds cold to 1.6 settled. The final also left a governing rule. With the model pinned there is 1.87 GB spare: a 274 MB embedding model fits, and a second 8 GB model evicts the resident.

Steve killed our first headline, back when Gemma 4 was briefly the only finisher: "That's like winning a bake-off when you're the only one who brought a cake. The cake could still suck."

The cake was fine.

Receipts

  • Full field, per-tier scores, tool-call counts: gemma4-on-prem-record.md
  • The harness defect and fixed-harness re-runs: artifacts/gate1/
  • Endurance run: 240 KB record, 393 voice samples, 1,510 MiB floor
  • Two-model tenancy policy: sam-infrastructure #310, #315, #317, #335, #300