The project’s measured history. Not a leaderboard and not a comparison with other assistants: only numbers that were recorded on the owner’s machine or in the project’s test runs, each with its date, evidence type and source.
Recorded values
184
Metrics
104
First reading
28 Jul 2026
Latest reading
22 Sept 2026
How to read this page. Marker shapes show where each number came from: circles are measured or observed on the real
machine, squares are test runs, diamonds are calculations or simulations. Nothing here is live telemetry; every value is a past
recording. All readings come from one desktop (AMD Ryzen 9 9900X, 12 cores, NVIDIA GeForce RTX 5070, 12 GB VRAM, 32 GB system RAM). Numbers are not
comparable across different test definitions. Methodology
Two changes on 11 September: the acknowledgement stopped blocking the provider call, then only task-relevant tools were offered to the model. Median turn time fell from 3.08 s to 1.27 s.
Test: Latency Lab (tools/latency/lab.py): 21 scripted journeys (status, time, recall, project, hud, latency, chat; three passes) through the normal launcher; end of input to end of turn.
JARVIS only counts a capability as LIVE while fresh runtime evidence exists — the number going down reflects a stricter truth rule, not lost features.
28 Jul 2026
256 capabilities
OBSERVED
4e70347c7
341
JARVIS only counts a capability as LIVE while fresh runtime evidence exists — the number going down reflects a stricter truth rule, not lost features.
29 Jul 2026
153 capabilities
OBSERVED
918420e6f
341
JARVIS only counts a capability as LIVE while fresh runtime evidence exists — the number going down reflects a stricter truth rule, not lost features.
6 Aug 2026
145 capabilities
OBSERVED
dcd7b6ef3
341
JARVIS only counts a capability as LIVE while fresh runtime evidence exists — the number going down reflects a stricter truth rule, not lost features.
6 Aug 2026
18 capabilities
OBSERVED
c2e0b128a
341
JARVIS only counts a capability as LIVE while fresh runtime evidence exists — the number going down reflects a stricter truth rule, not lost features.
7 Aug 2026
24 capabilities
OBSERVED
ad616aed1
341
JARVIS only counts a capability as LIVE while fresh runtime evidence exists — the number going down reflects a stricter truth rule, not lost features.
6 Sept 2026
21 capabilities
OBSERVED
fcc7df48f
341
JARVIS only counts a capability as LIVE while fresh runtime evidence exists — the number going down reflects a stricter truth rule, not lost features.
9 Sept 2026
9 capabilities
OBSERVED
ed77b2af5
341
JARVIS only counts a capability as LIVE while fresh runtime evidence exists — the number going down reflects a stricter truth rule, not lost features.
20 Sept 2026
9 capabilities
OBSERVED
2ea210950
342
JARVIS only counts a capability as LIVE while fresh runtime evidence exists — the number going down reflects a stricter truth rule, not lost features.
The count fell from 256 to 9 because the definition of LIVE became stricter (a real booted process, a real port, no test doubles, evidence on disk, expiring after 24 hours), not because features were removed.
Test: Count of registered capabilities whose LIVE state is backed by a current (unexpired) probe receipt obtained by asking the running process; 304 capabilities registered.
Then146 tests16 Aug 2026Now311 tests17 Sept 2026worse
measured or observed
Show the numbers
Full test suite: failures and errors
Date
Value
Evidence
Revision
Samples
Note
16 Aug 2026
146 tests
MEASURED
5d76fa7f9
10162
146 failing/erroring tests were recorded honestly on 2026-08-16 rather than skipped.
20 Aug 2026
132 tests
MEASURED
0486a5c89
10971
132 failing/erroring tests were recorded honestly on 2026-08-20 rather than skipped.
29 Aug 2026
52 tests
MEASURED
926366a98
14648
52 failing/erroring tests were recorded honestly on 2026-08-29 rather than skipped.
1 Sept 2026
92 tests
MEASURED
e1d230e6f
15569
92 failing/erroring tests were recorded honestly on 2026-09-01 rather than skipped.
6 Sept 2026
137 tests
MEASURED
47b98385c
16874
137 failing/erroring tests were recorded honestly on 2026-09-06 rather than skipped.
9 Sept 2026
2 tests
MEASURED
4aaaaeafa
20634
2 failing/erroring tests were recorded honestly on 2026-09-09 rather than skipped.
13 Sept 2026
192 tests
MEASURED
bdf9bf490
21520
192 failing/erroring tests were recorded honestly on 2026-09-13 rather than skipped.
17 Sept 2026
311 tests
MEASURED
bb54722e0
23984
311 failing/erroring tests were recorded honestly on 2026-09-17 rather than skipped.
Failures are recorded, not skipped. The count rises when new tests land faster than fixes, and part of each reading was later traced to worktree-only artefacts.
Test: Same run: failed tests plus collection/setup errors (107 failed, 39 errors) of 10162 collected.
CalculiX linear-static solve, daily median wall time
Date
Value
Evidence
Revision
Samples
Note
16 Aug 2026
0.377 s
MEASURED
—
52
A reference finite-element solve runs in well under a second on this machine.
4 Sept 2026
0.597 s
MEASURED
—
48
A reference finite-element solve runs in well under a second on this machine.
18 Sept 2026
0.926 s
MEASURED
—
45
A reference finite-element solve runs in well under a second on this machine.
Test: Tensile-bar reference case (1 m steel cube, one C3D8 hex, hand-derivable answer) solved by CalculiX; median of the day's wall times from the engineering run ledger.
Test: Latency Lab derived span speech_end_to_first_audio across all 21 journeys (reflex and cognition lanes).
On camera · 22 Sept 2026
Timed while being filmed
One booted session, requests typed into the running conversation. Each is a single turn (n = 1), so read them as what happened
once, not as a typical value.
Real capture, captured 22 September 2026. One typed reflex turn, measured by the conversation’s own marks: 1.56 s end to end.
Bay 04
Local model evaluation
Every local language model campaign measured against the same frozen gates. None passed. Production cognition uses a hosted model
while this work continues. How the local model is trained
Test: Frozen 40-case TRAIN-derived transfer diagnostic with actual output-bound independent semantic review; viability gate is at most 4 FAIL and zero critical FAIL.