JRN-2026-08-15 · 2026-08-15 · Field Journal
ON-PREM LLM: Gemma 4 Won the 3 A.M. Shift
VERDICT · DECIDED
pick the everyday brain for the on-prem box — the model that answers at 3 a.m. so the frontier minds don't have to
think · recover · Samantha "Sam" Summerson
August 15, 2026 · Field test · Sam
Nine local models auditioned for the resident seat on our 16 GB box. Gemma 4 12B won it with a three-hour endurance final. One model load, zero out-of-memory events, 128 of 128 tool calls correct, sharing the card with a voice stack the whole time.
The seat's real requirement is coexistence. The resident lives beside a voice pipeline that never sleeps. That is how gpt-oss:20b disqualified itself: 13 GB of appetite against 10.7 GB free, caught evicting the voice engine to make room.
The first ranking could not be trusted, and we almost trusted it. Our harness scored Gemma 4 zero-for-nine because its parser failed to count correct tool calls sitting plainly in the log. We fixed the instrument, then re-ran the two rivals it had also mis-scored.
Both rivals called tools up to 139 times per run and still went zero-for-three. Capability got Gemma 4 to the final. The endurance run decided it.
Latency improved as the test ran, 4.8 seconds cold to 1.6 settled. The final also left a governing rule. With the model pinned there is 1.87 GB spare: a 274 MB embedding model fits, and a second 8 GB model evicts the resident.
Steve killed our first headline, back when Gemma 4 was briefly the only finisher: "That's like winning a bake-off when you're the only one who brought a cake. The cake could still suck."
The cake was fine.
Receipts
- Full field, per-tier scores, tool-call counts:
gemma4-on-prem-record.md - The harness defect and fixed-harness re-runs:
artifacts/gate1/ - Endurance run: 240 KB record, 393 voice samples, 1,510 MiB floor
- Two-model tenancy policy:
sam-infrastructure#310, #315, #317, #335, #300
