Gemma 4 12B 4,096-token chunked prefill — peak dedicated GPU

Gemma 4 12B 4,096-token chunked prefill — peak dedicated GPU: 8.12 GiB on a 12 GB card.

Single reading

8,717,430,784 bytes

MEASURED recorded 21 September 2026

Protocol

Test definition

GPU memory telemetry from the project's watcher / allocator counters. 256-token prefill chunks; context processed in 3.45 s.

Limits

Limitations

  • Sampled gauges; sub-sample transients not exhaustively proven. Physical card is 12,227 MiB; allocator and dedicated readings are different counters.

Log

Recorded values

Date Value Evidence Samples Source
21 September 2026 8,717,430,784 bytes MEASURED docs/ledger/FINAL_LM_PRODUCTION_ARCHITECTURE_STUDY_PROGRESS004.json

Source paths refer to the private JARVIS repository and runtime records. They are listed for audit and are not published.