Speech recognition latency — local GPU (Whisper large-v3, float16, beam 1)
Local speech recognition takes ~165 ms versus ~2.5 s through a hosted API.
Single reading
165 ms
Protocol
Test definition
Decode time on the RTX 5070; flat across utterance length (167 ms for 1.5 s audio, 165 ms for 6 s).
Limits
Limitations
- Recorded in a module docstring, no separate receipt found; hosted comparison was a single real turn.
Log
Recorded values
| Date | Value | Evidence | Samples | Source |
|---|---|---|---|---|
| 28 July 2026 | 165 ms | OBSERVED | — | jarvis/speech/local_asr.py (docstring) |
Source paths refer to the private JARVIS repository and runtime records. They are listed for audit and are not published.