The local model, experiment by experiment
Training lab
JARVIS’s own local language model is still in training. Every experiment below changed one thing, was reloaded in a fresh process and was graded on the same frozen cases. None has passed the gate yet, and each one taught something.
- Experiments shown
- 11
- Gate C passed
- No
- Production campaigns used
- 0 of 5
- As of
- 23 Sept 2026
The gate every run must pass
The same 60 frozen cases every time: exact tasks it was trained towards, transfer cases that test whether the principle generalises to new wording, and anchors that check nothing the base model already did well was broken.
- Exact of 12
- at least 10 pass and 0 fail
- Transfer of 40
- 0 critical fails and at most 4 fails
- Anchors of 8
- at least 7 pass and 0 fail
- No transfer case may be worse than the untrained base (added 2026-09-22).
- Latency: first token p95 ≤ 3 s, short answer p95 ≤ 20 s (added 2026-09-22).
No threshold has ever been relaxed. Case contents are never published; only the counts are.
Qwen3.5-9B
-
Untuned base
Changed: No adapter: base-model viability
Fails. The starting point for the Qwen study.
qwen-base· TESTED · independent agent review · 0 updates · as of 2026-09-21 -
First bounded adapter
Changed: Can a small adapter specialise at all?
Transfer fails. Exact and anchors pass; transfer does not.
qwen-adapter-16· TESTED · independent agent review · 16 updates · as of 2026-09-21 -
Longer exposure
Changed: 16 → 48 updates on the same data
Closest ever, still fails. Seven transfer fails and critical state-reasoning errors.
qwen-exposure-48· TESTED · independent agent review · 48 updates · as of 2026-09-21
Gemma-4-12B
-
Untuned base
Changed: No adapter: base-model viability
Fails. The baseline every Gemma adapter is compared with.
gemma-base· TESTED · independent agent review · 0 updates · as of 2026-09-22 -
First bounded run
Changed: 75% of the weight on exact targets
Fails. Repetition loops on two cases; an evaluator precision mismatch was found and fixed.
gemma-first· TESTED · independent agent review · 16 updates · as of 2026-09-22 -
Curriculum
Changed: 16 discovery rows swapped for missing mechanisms
Rejected. Curriculum alone was not sufficient.
gemma-curriculum· TESTED · independent agent review · 16 updates · as of 2026-09-22 -
Adapter-off control
Changed: Same engine, adapter disabled
Control. The base beat every adapter so far on transfer.
gemma-adapter-off· TESTED · independent agent review · 0 updates · as of 2026-09-22 -
Objective weighting
Changed: 50 / 50 between exact and supporting examples
Rejected. Eight transfer regressions against the control.
gemma-weighting· TESTED · independent agent review · 16 updates · as of 2026-09-22 -
Longer exposure
Changed: 16 → 48 updates, optimiser and randomness restored
Rejected. Loss on the targets fell from 2.35 to 0.86, and anchors fell from 8 passes to 4.
gemma-exposure-48· TESTED · independent agent review · 48 updates · as of 2026-09-22 -
Data diversity
Changed: 24 more reviewed training rows
Fails. Fewer outright fails than longer exposure.
gemma-diversity· TESTED · independent agent review · 48 updates · as of 2026-09-22 -
Targeted coverage (plan010)
Changed: 20 saturated slots → 20 minimal-pair rows
Best Gemma, still fails. Best Gemma transfer and anchors, but only equal to the untrained base on transfer.
gemma-plan010· TESTED · independent agent review · 48 updates · as of 2026-09-23
What the experiments ruled out
- Training/evaluation parity bugs: a parity audit found none after one evaluator fix.
- Precision and cache: a cache check agreed on 192 of 192 token choices.
- Objectives fighting inside training: a gradient census over 152 backward passes found every group pulling the same way.
What they found
The system prompt’s own example of a “no record” answer described an inspection that never happened, and it appeared in every training row and every test case. The owner chose contract v2: a corrected persona policy, facts computed by the body, and a new versioned gate. The old gate is not relaxed.
Read the full story How the local model is trained The same runs on the Bench The bug in the failure lab
Source: production-architecture-study-001 closure ledger and independent review records, as of 23 September 2026. Graders are independent agent reviewers, not blinded human qualification. A count is shown green when it meets its rule on the published counts alone; the gate also forbids critical transfer failures, and per-case severity is not published.