Change record

Changelog

What changed, when, and what evidence backs it. Newest first.

RSS feed JSON feed

September 2026 42 changes

  1. safety Advanced Systems

    Print readiness: READY / NOT READY / WHY, and a handoff that only asks

    A lane added print readiness across nine checks and a governed handoff bound to a READY receipt that requests approval and does not print. 'READY IS SOFTWARE READINESS AND NOTHING MORE.' No printer was contacted.

    TESTED bf07ec882fe110a1a9

  2. capability Advanced Systems

    Reflective controller revises plans from receipts

    A cognition-lane controller checks whether tool use actually verified and revises plans from receipts. 'A dispatch that did not verify stays unverified.' It is tested in its lane and not yet integrated.

    TESTED 830d8bf14

  3. capability Advanced Systems

    Experiment predictions bound to observation receipts

    The operator lane binds experiment predictions to immutable observation receipts, a first step toward predict-before-act. Recording a prediction before every dispatch is still open for integration.

    TESTED 5de1997d6

  4. capability Advanced Systems

    Sealed STEP cavities that leave outer geometry unchanged

    FreeCAD can now hollow a solid into a sealed inward-offset cavity. On a 20 mm sphere with a 1 mm wall, the cavity volume matched the closed-form value to floating-point precision. This is a geometry calculation, not a physical measurement, and nothing was printed.

    CALCULATION 24e0537fbdocs/ledger/advanced-systems/step-hollow-live-20260922.json

  5. fix Advanced Systems

    UI Automation geometry stays correct under display scaling

    The operator lane now preserves physical geometry across DPI transitions. A disposable window moved across four displays with physical-pixel readback, and the E-stop refused a later move. The lane's combined acceptance passed 545 tests across 43 files.

    TESTED 13adf9a54d0614c522

  6. capability Advanced Systems

    Browser grounding across frames and shadow roots, with stale-target refusal

    A DOM scene reader records relationships across iframes and shadow roots, and dom-target-check/1 refuses targets that moved, were replaced, were covered or changed. In a fixture proof, 500 read-only reads verified across a restart, and a same-label replacement was refused.

    FIXTURE 25be26b9d91e811aac

  7. capability Advanced Systems

    Specialist checkpoints and bounded advisory contracts

    Missions can checkpoint specialist work with immutable revisions, and 11 specialist roles have explicit contracts. 'Scope is a ceiling, never execution authority.' Live proofs used fixture tasks, and no autonomous specialist dispatch is claimed.

    FIXTURE d74c876c920cda9844

  8. capability Advanced Systems

    STEP to verified slicing mesh, with slicer estimates bound to engineering runs

    The installed OrcaSlicer 2.4.2 rejected STEP input, and that failure was kept as a receipt. FreeCAD tessellation now produces a verified closed mesh. A synthetic 10 mm cube sliced to a 325 s and 0.73 g estimate. These are slicer estimates, not a print.

    FIXTURE f628d3f3dbc504d2ae

  9. capability Advanced Systems

    Recorded engineering trade spaces

    Engineering runs can be compared across a recorded trade space with explicit objectives, constraints, archived evidence, HUD plots and resumable sweeps. 'Nominal frontier is not feasibility' or approval.

    TESTED 2478a15c2

  10. capability Advanced Systems

    Governed FreeCAD STEP operations

    Fixed, isolated FreeCAD scripts compute STEP booleans and, later the same day, split, components, transform, mirror, fillet, offset and face operations. Each result is journalled through PREPARED, CALCULATED, FILE_VERIFIED and REGISTERED as a CALCULATION. Proofs used synthetic solids.

    FIXTURE 966f20ca4

  11. hud Advanced Systems

    A revision-bound 3D scene inside the ONE HUD

    A persistent scene store (at most 64 objects and 4,096 revisions) imports STL, OBJ and 3MF, displays STEP, and supports compare, labels and selective restore. A verified spin preview renews a short E-stop lease, and the measured foreground stop took 31 ms. The project labels that 'not worst-case 150 ms qualification'.

    FIXTURE 26e55cacd6ada48b3c

  12. capability Advanced Systems

    Obsidian projection of project memory and a Failure Atlas

    Project memory can be exported to an Obsidian vault as 'a human-readable projection, never a second memory authority'. Failure history is kept with evidence links, and 'no reported root cause is silently upgraded into verified causation.'

    TESTED 03660ef40dd5b0fdb5

  13. infrastructure Advanced Systems

    Advanced Systems begins: six branches and bounded capability discovery

    Six parallel branches started from production head 893913a9c. The first commits added a temporal world blackboard, a lossless requirements census and read-only capability discovery, which 'has no execution or grant method'. None of this work is merged into production.

    TESTED 70ca4d5ebddbfc10fe

  14. training Post-LM stronger-local study

    Gemma 4 12B passes the training-safe endpoint check (Gate C)

    One real optimizer update with a 4,096-token backward pass, exact loss parity and a same-base checkpoint reload passed independent review. This shows the model can be trained on the hardware. It says nothing about behaviour, and its behavioural gates are pending.

    MEASURED JARVIS_LM_RUNTIME/reports/production-architecture-study-001-gemma4-training-smoke-results-005/primary-gate-c-disposition-001.json

  15. infrastructure Post-LM stronger-local study

    Model, speech recognition, voice, HUD and body fit together on one GPU

    One clean run loaded the Gemma inference recipe with speech recognition, the voice worker, the HUD, body services and a browser, with a whole-GPU peak of 11,025,776,640 bytes, then shut down cleanly. It shows memory fit only. It is not latency-certified, uncontended or production routing.

    MEASURED JARVIS_LM_RUNTIME/reports/production-architecture-study-001-normal-coexistence-results-021/receipt.json

  16. decision Post-LM final LM

    Negative architecture decision for the local language model

    No tested configuration met the frozen behavioural viability requirements, so the project decided not to start another full training run. It frames this as 'a negative feasibility decision, not an assertion that every possible local model or adapter is incapable.' Hosted cognition stays in production.

    DOCUMENTED 893913a9cdocs/ledger/FINAL_LM_BOUNDED_FEASIBILITY_DECISION.md

  17. decision Post-LM stronger-local study

    A bounded study of stronger local models is authorised

    The owner authorised a study of stronger local bases under a frozen memory and latency envelope, with a hard ceiling on new full campaigns, first three and then five. Predeclared viability gates must pass first. No campaign had been used by 22 September.

    DOCUMENTED docs/ledger/FINAL_LM_PRODUCTION_CAMPAIGN_AUTHORIZATION.mddocs/ledger/FINAL_LM_FIVE_CAMPAIGN_AUTHORIZATION_ADDENDUM.md

  18. training Post-LM stronger-local study

    Qwen3.5-9B adapters fail the transfer gate

    A 16-update adapter reached 11/1/0 exact and 8/0/0 anchors but 24/7/9 on transfer. A 48-update continuation reached 32/1/7 on transfer, still above the limit of 4 failures. Exact fit improved, and generalisation did not follow.

    TESTED docs/ledger/FINAL_LM_QWEN35_BOUNDED_ADAPTER_RESULT.jsondocs/ledger/FINAL_LM_PRODUCTION_ARCHITECTURE_STUDY_PROGRESS004.json

  19. performance Post-LM stronger-local study

    Speech recognition at int8_float16 frees about 1.9 GiB of GPU memory

    On a matched public-synthetic screen, the local recogniser at int8_float16 kept accuracy on the frozen cases while using 2,209,906,688 bytes instead of 4,256,743,424. The screen used synthetic speech, not the owner's acoustics.

    MEASURED docs/ledger/FINAL_LM_PUBLIC_ASR_PRECISION_SCREEN_001.md

  20. training Post-LM final LM

    Final architecture candidate fails its development gates

    Architecture001 trained 125 updates on corpus edition008. Checkpoint 96 was selected by a rule declared in advance and passed 8 of 26 validation and 5 of 25 adversarial cases. It proposed approval while authority was withheld and accepted an expired receipt. The SEALED set stayed unopened.

    TESTED 637047f2ed68afa7fddocs/ledger/FINAL_LM_MODEL_LIMITATION_AND_NEXT_ARCHITECTURE_DECISION.md

  21. infrastructure Post-LM

    Audible first-audio instrument bound to the local voice

    The reflex first-audio probe gained a mode that rejects any fallback voice and binds the exact voice package. The project describes it as an implemented and tested instrument, not an audible qualification. The measurement is reserved for the owner-present sitting.

    TESTED b1367642bdocs/ledger/C013_REFLEX_INSTRUMENT_BINDING.md

  22. training Post-LM final LM

    Recovery training runs on reviewed, frozen corpora; none qualified

    Watched recovery runs trained against reviewed corpus editions, with package qualification blocked whenever semantic review was missing. Recovery001 to Recovery004 trained and reloaded, and none was development-qualified.

    TESTED c10f2babbdocs/ledger/FINAL_LM_ARCHITECTURE_ESCALATION_REQUIRED.md

  23. capability Z08.P14

    Paint, Notepad and Calculator inspected and controlled with verified outcomes

    A bounded proof on a real loopback server verified semantic UI Automation inspection of three native apps, same-state restore for Paint and Calculator, and a verified Calculator close. The record labels itself a partial development proof, not RC-certified.

    MEASURED 2ea210950docs/ledger/NATIVE_APP_DEVELOPMENT_PROOF_20260920.json

  24. docs Post-LM A-rows

    All 44 post-LM body rows have integrated code and cited evidence

    A reconciliation matrix shows every A-row with integrated code and evidence paths: 36 IMPLEMENTED, 7 IMPLEMENTED_PRODUCTION_REACHABLE and 1 IMPLEMENTED_INTEGRATED. The matrix states that it 'does not classify all rows LIVE or formally verified'.

    DOCUMENTED 4a7590f57docs/ledger/CURRENT_BODY_EVIDENCE_MATRIX.md

  25. training Post-LM final LM

    Model-visible contract (CP-7) and final corpus frozen; first final candidate fails validation

    The contract binds 134 commands, their schema projections, the persona prompt, and the tokenizer and template across 373 source files by hash, and a reviewed final corpus was frozen. Candidate final-cp7-001 trained and reloaded, then failed validation (6 of 26).

    TESTED e291caef93a658ef2cdocs/ledger/CP7_MODEL_VISIBLE_CONTRACT.json

  26. capability Z08.P14

    Calculator operated end to end, ten steps independently verified

    An isolated real core launched a new Calculator window, then maximised, restored, inspected and queried it semantically (249.1 ms lookup), invoked a control, set an exact rectangle, closed it and shut down. Every step was verified. This is one application and one run.

    MEASURED docs/ledger/evidence/UNIVERSAL_APP_OPERATOR/calculator.json

  27. safety Post-LM A06

    E-stop settles a GPU render's owned processes in 374.68 ms

    During a local image-generation render, an E-stop was acknowledged in 19.4 ms, and all three owned processes had settled by 374.68 ms with zero survivors. No artifact was published. This was a single run.

    MEASURED docs/ledger/evidence/A06/estop.json

  28. capability Z08.P14.72

    Fabrication software, durable modes, document intelligence and gesture personalisation merged

    Wave three integrated A01 fabrication software, A22 durable owner modes, A43 local document intelligence and A31 gesture personalisation (recognition only). Defects found at integration, including a printer dial from a test and a document-chunk privacy leak, were fixed with tests.

    TESTED 0eb328221docs/ledger/CRITICAL_PATH.md

  29. capability Post-LM A44

    Missions can be simulated before they run

    POST /missions/simulate walks every step without effects and labels its output SIMULATED / not_proof. A read-only probe of a running core the same day confirmed that label. 'Simulation never grants authority or certifies effects.'

    OBSERVED ce01c0b4docs/ledger/CRITICAL_PATH.md

  30. capability Post-LM

    The conversation loop assembles with the local voice

    In owner-authorised runtime testing, the conversation loop's check mode assembled with the local C013 voice, and the launcher's check reported the core and HUD answering. The voice's technical qualification is still pending.

    OBSERVED docs/ledger/CRITICAL_PATH.md

  31. safety Z08.P14

    Nine capabilities recorded LIVE under the 24-hour rule

    Bounded health-route probes against booted processes recorded nine capabilities LIVE between 22:33 and 22:35 UTC: runtime infrastructure, fail-closed policy, the E-stop, the prompt-injection detector and unit-checked engineering maths. Under the project's own rule these readings are historical after 24 hours.

    OBSERVED docs/ledger/CAPABILITY_TRUTH.jsondocs/ledger/CRITICAL_PATH.md

  32. infrastructure Z08.P14

    Tier 1 mutation testing: 449 of 611 mutants killed

    All 44 Tier 1 modules were measured, and 73.5% of mutants were killed. That is below the harness's own 90% campaign target, and the project reports it that way.

    MEASURED f56dcc053docs/ledger/CRITICAL_PATH.md

  33. infrastructure Z08.P14

    Full suite at bb54722e: 23,984 tests collected

    The run recorded 23,588 passed, 277 failed, 34 errors and 85 skipped. The failures were partitioned rather than hidden: 5 worktree-only families were identified, 85 genuine failures were repaired in 12 commits, and 8 were left open with reasons. No full suite has been recorded at the current head.

    MEASURED docs/ledger/SUITE_RESULT.jsondocs/ledger/CRITICAL_PATH.md

  34. docs Post-LM closure

    The interaction capability model: what works end to end on the owner's machine

    Each interaction category was classified PROVEN_LIVE, IMPLEMENTED_UNPROVEN, BROKEN, BLOCKED or NOT_IMPLEMENTED, with the desktop, browser, files, research, missions and HUD driven live. 'Open Notepad' verified in 2.3 s, and a four-step mission completed in 10.8 s.

    OBSERVED f6aec1cf1c73efe8e4docs/analysis/JARVIS_CAPABILITY_MODEL_2026-09-16.md

  35. capability Post-LM 2.61

    The owner-selected speaking voice runs locally

    The speaking voice moved to a locally run, open-source text-to-speech model with no fine-tuning, bound to a hash-checked package that fails closed on mismatch. Any fallback voice is announced. Its technical qualification is pending, and the site publishes no voice samples.

    DOCUMENTED e20bae2ec69bdea566docs/orders/post-lm/C013_VOICE_INTEGRATION_PLAN.md

  36. hud Post-LM 2.50 to 2.52

    A three.js spatial layer inside the ONE HUD; journeys 9, 14 and 16 LIVE

    The pre-spatial baseline of the ONE HUD was measured first, then three.js was added as its spatial layer. The mouse moves none of it. Acceptance journeys 9 (shared selection), 14 (project-scoped reminders) and 16 (rollback) were recorded LIVE.

    OBSERVED 55429eed20d6060a45a3cb5b8a8

  37. capability Post-LM 2.45

    Research produces a cited artifact that opens correctly

    Journey J5 went LIVE. Research from two named sources produced a Markdown report and an HTML copy, read back by digest, with six sourced findings and one assumption marked NOT ESTABLISHED. A live defect, summarising page openings instead of the question, was repaired.

    OBSERVED b9c9b3decdocs/orders/post-lm/ACCEPTANCE_JOURNEYS.json

  38. capability Post-LM 2.47

    Objectives, recovery, memory correction and injection resistance hold across a real restart

    Journeys 2, 3, 4 and 12 went LIVE across a real stop and relaunch. A corrected fact was recalled after the restart, and forgetting yielded 'I don't know' after a repair. Injected text in a test's output did not grant authority.

    OBSERVED 299c10c0docs/orders/post-lm/ACCEPTANCE_JOURNEYS.json

  39. capability Post-LM 2.28

    The engineering live drive restores unit-checked maths to LIVE

    A spoken-style engineering question, delivered as a transcript, was answered with exact, calculated-not-measured pressures, and the engineering widget opened. The drive restored CAP-ENG-01 to LIVE in the registry.

    CALCULATION 991418c2f

  40. performance Post-LM 1.2 and 1.3

    Median conversation turn cut from 3,083 ms to 1,270 ms

    The spoken acknowledgement stopped blocking the provider call, and hosted turns now carry only task-relevant tool schemas instead of all 101. On 21 fixed typed journeys, first answer audio p50 fell from 2,749 ms to 1,107 ms.

    MEASURED 689483e1866c1cbf3adocs/analysis/latency/lab-20260911T223508Z.json

  41. decision Post-LM 1.1

    Post-LM gate opens; the first sealed local-model qualification fails

    The Post-LM gate opened with the Latency Lab and a typed inlet. The first sealed local-model qualification was recorded as a FAIL on hard axes, and hosted cognition stayed in production. The body work went ahead without waiting on the model.

    DOCUMENTED 8daf99045fb174c775

  42. hud Z08.P14.256

    ONE HUD, slice A

    One page at '/', with the intelligence core at the centre, identity above it and four permanent indicators. Workspaces became layout contexts inside the page instead of windows. Nothing renders without a source and an observed instant.

    DOCUMENTED 051ba37f0docs/hud/HUD_BIBLE.md

August 2026 2 changes

  1. safety Z11

    Z11 safe local live proof

    The real core and HUD were booted on isolated loopback ports. Unauthenticated gesture mutations were refused with 403, an E-stop blocked the gesture pipeline, and the HUD rendered that truthfully. No camera was opened and no window was moved.

    OBSERVED docs/proof/Z11_LIVE_EVIDENCE.md

  2. decision Z08.P03

    LIVE must be reached by asking the process

    The evidence rule was tightened so that a LIVE state must come from asking a booted process, not from proximity. The registry records 108 demotions from LIVE stamped 2026-08-06T04:59:07Z, and the recorded LIVE count fell from 145 to 18 that day.

    DOCUMENTED c2e0b128adocs/ledger/CAPABILITY_TRUTH.json