Gemma 4 12B 4,096-token chunked prefill — peak dedicated GPU
Gemma 4 12B 4,096-token chunked prefill — peak dedicated GPU: 8.12 GiB on a 12 GB card.
Single reading
8,717,430,784 bytes
Protocol
Test definition
GPU memory telemetry from the project's watcher / allocator counters. 256-token prefill chunks; context processed in 3.45 s.
Limits
Limitations
- Sampled gauges; sub-sample transients not exhaustively proven. Physical card is 12,227 MiB; allocator and dedicated readings are different counters.
Log
Recorded values
| Date | Value | Evidence | Samples | Source |
|---|---|---|---|---|
| 21 September 2026 | 8,717,430,784 bytes | MEASURED | — | docs/ledger/FINAL_LM_PRODUCTION_ARCHITECTURE_STUDY_PROGRESS004.json |
Source paths refer to the private JARVIS repository and runtime records. They are listed for audit and are not published.