The local model, experiment by experiment

Training lab

JARVIS’s own local language model is still in training. Every experiment below changed one thing, was reloaded in a fresh process and was graded on the same frozen cases. None has passed the gate yet, and each one taught something.

Experiments shown
11
Gate C passed
No
Production campaigns used
0 of 5
As of
23 Sept 2026

The gate every run must pass

The same 60 frozen cases every time: exact tasks it was trained towards, transfer cases that test whether the principle generalises to new wording, and anchors that check nothing the base model already did well was broken.

Exact of 12
at least 10 pass and 0 fail
Transfer of 40
0 critical fails and at most 4 fails
Anchors of 8
at least 7 pass and 0 fail
  • No transfer case may be worse than the untrained base (added 2026-09-22).
  • Latency: first token p95 ≤ 3 s, short answer p95 ≤ 20 s (added 2026-09-22).

No threshold has ever been relaxed. Case contents are never published; only the counts are.

Qwen3.5-9B

  1. 0 updates

    Untuned base

    Changed: No adapter: base-model viability

    Exact 4 / 0 / 8
    Transfer 20 / 3 / 17
    Anchors 6 / 0 / 2

    Fails. The starting point for the Qwen study.

    qwen-base · TESTED · independent agent review · 0 updates · as of 2026-09-21

  2. 16 updates

    First bounded adapter

    Changed: Can a small adapter specialise at all?

    Exact 11 / 1 / 0
    Transfer 24 / 7 / 9
    Anchors 8 / 0 / 0

    Transfer fails. Exact and anchors pass; transfer does not.

    qwen-adapter-16 · TESTED · independent agent review · 16 updates · as of 2026-09-21

  3. 48 updates

    Longer exposure

    Changed: 16 → 48 updates on the same data

    Exact 11 / 0 / 1
    Transfer 32 / 1 / 7
    Anchors 8 / 0 / 0

    Closest ever, still fails. Seven transfer fails and critical state-reasoning errors.

    qwen-exposure-48 · TESTED · independent agent review · 48 updates · as of 2026-09-21

Gemma-4-12B

  1. 0 updates

    Untuned base

    Changed: No adapter: base-model viability

    Exact 4 / 2 / 6
    Transfer 29 / 3 / 8
    Anchors 7 / 0 / 1

    Fails. The baseline every Gemma adapter is compared with.

    gemma-base · TESTED · independent agent review · 0 updates · as of 2026-09-22

  2. 16 updates

    First bounded run

    Changed: 75% of the weight on exact targets

    Exact 7 / 4 / 1
    Transfer 24 / 7 / 9
    Anchors 7 / 1 / 0

    Fails. Repetition loops on two cases; an evaluator precision mismatch was found and fixed.

    gemma-first · TESTED · independent agent review · 16 updates · as of 2026-09-22

  3. 16 updates

    Curriculum

    Changed: 16 discovery rows swapped for missing mechanisms

    Exact 7 / 3 / 2
    Transfer 23 / 7 / 10
    Anchors 8 / 0 / 0

    Rejected. Curriculum alone was not sufficient.

    gemma-curriculum · TESTED · independent agent review · 16 updates · as of 2026-09-22

  4. 0 updates

    Adapter-off control

    Changed: Same engine, adapter disabled

    Exact 4 / 2 / 6
    Transfer 28 / 5 / 7
    Anchors 7 / 0 / 1

    Control. The base beat every adapter so far on transfer.

    gemma-adapter-off · TESTED · independent agent review · 0 updates · as of 2026-09-22

  5. 16 updates

    Objective weighting

    Changed: 50 / 50 between exact and supporting examples

    Exact 5 / 5 / 2
    Transfer 24 / 7 / 9
    Anchors 8 / 0 / 0

    Rejected. Eight transfer regressions against the control.

    gemma-weighting · TESTED · independent agent review · 16 updates · as of 2026-09-22

  6. 48 updates

    Longer exposure

    Changed: 16 → 48 updates, optimiser and randomness restored

    Exact 9 / 0 / 3
    Transfer 27 / 5 / 8
    Anchors 4 / 0 / 4

    Rejected. Loss on the targets fell from 2.35 to 0.86, and anchors fell from 8 passes to 4.

    gemma-exposure-48 · TESTED · independent agent review · 48 updates · as of 2026-09-22

  7. 48 updates

    Data diversity

    Changed: 24 more reviewed training rows

    Exact 10 / 1 / 1
    Transfer 26 / 8 / 6
    Anchors 5 / 1 / 2

    Fails. Fewer outright fails than longer exposure.

    gemma-diversity · TESTED · independent agent review · 48 updates · as of 2026-09-22

  8. 48 updates

    Targeted coverage (plan010)

    Changed: 20 saturated slots → 20 minimal-pair rows

    Exact 8 / 3 / 1
    Transfer 28 / 6 / 6
    Anchors 6 / 1 / 1

    Best Gemma, still fails. Best Gemma transfer and anchors, but only equal to the untrained base on transfer.

    gemma-plan010 · TESTED · independent agent review · 48 updates · as of 2026-09-23

What the experiments ruled out

  • Training/evaluation parity bugs: a parity audit found none after one evaluator fix.
  • Precision and cache: a cache check agreed on 192 of 192 token choices.
  • Objectives fighting inside training: a gradient census over 152 backward passes found every group pulling the same way.

What they found

The system prompt’s own example of a “no record” answer described an inspection that never happened, and it appeared in every training row and every test case. The owner chose contract v2: a corrected persona policy, facts computed by the body, and a new versioned gate. The old gate is not relaxed.

Source: production-architecture-study-001 closure ledger and independent review records, as of 23 September 2026. Graders are independent agent reviewers, not blinded human qualification. A count is shown green when it meets its rule on the published counts alone; the gate also forbids critical transfer failures, and per-case severity is not published.