Honest edges
What JARVIS cannot do yet
The open limitations of every JARVIS system, collected in one place: 76 items across 14 systems. Each comes from the system’s own page.
Nothing in JARVIS is release-qualified or formally verified yet. Local cognition is not in production, full-duplex voice needs an owner-present sitting, and physical printing is deliberately refused. The lists below are the detailed version.
How a request moves through JARVIS
INTEGRATED PARTIAL , rung 4 of 8
- Live readings are dated. The typed conversation and planning readings come from 16 September 2026, and the governance registry probes from 17 September 2026. The project treats live evidence as expiring after 24 hours, so these are records of what happened then, not claims about today.
- The full spoken loop has not been qualified. Full-duplex voice is waiting on an owner-present sitting. The interrupted-speech journey is fixture-tested only. Audible first-audio latency with the C013 voice has not been measured.
- The local model is not the brain. Production cognition is hosted. No local candidate has passed its gates.
- Some context features lag in reachability. Context-stack registry rows were demoted during an internal reachability audit. Later rows record them between TESTED and PRODUCTION-REACHABLE.
- No release. The project's own certification matrix reads "JARVIS-READY-V1: NO", and the final gate has not passed. See capabilities for rung-by-rung status and bench for measurements.
Architecture: brain, body and the truth pipeline
LIVE-PROVEN PARTIAL , rung 6 of 8
- Live readings have expired. Every registry LIVE state dates from 17 September 2026 and has decayed under the project's own 24-hour rule.
- No release qualification. The release row is still active, the final gate has not passed, and the certification matrix reads "JARVIS-READY-V1: NO".
- Recovery under fault is implemented, but the live chaos matrix has not been run. Survival across a real machine reboot is not proven. A process kill is explicitly not accepted as a substitute.
- The local brain has no qualified candidate. Production cognition is hosted. See training.
- Advanced Systems work is unmerged, largely unfinished against its own matrix, and partly untested after the measurement hold began.
- Some wiring is uncertain. It is unclear whether several implemented modules are constructed by the production runtime at all.
Governance and safety
LIVE-PROVEN PARTIAL , rung 6 of 8
- Live readings are historical. The governance LIVE states date from 17 September 2026 and have decayed under the 24-hour rule.
- Most governance rows are TESTED, not live. This includes approvals, the audit trail and the evidence store. Their tests pass, but no dated live probe is on record.
- The intent-ledger gap is uncertain. Notes from 18 September 2026 recorded that production did not construct the intent ledger, so multi-step plans containing an irreversible command were refused. That failure mode is safe, but it limits what JARVIS can do. Later work targeted the gap, and whether it is fully closed is uncertain.
- The E-stop's 150 ms budget has not been qualified at worst case.
- Nothing here is formally verified. The properties above rest on code review, tests and bounded live probes, not machine-checked proofs.
Memory
TESTED PARTIAL , rung 3 of 8
- Most memory evidence is test-level. The registry memory rows passed their probes, but an August audit found no production entry point for them at the time. The capability model rated memory implemented but unproven on 16 September 2026.
- The single live journey is historical. Journey J4 (a correction supersedes an old memory and deletion propagates) was verified live from a reading taken before 17 September 2026.
- Causal grades are not fully integrated.
- Hybrid retrieval wiring is uncertain.
- Obsidian, the Failure Atlas and the eight-layer hierarchy are not merged into production.
- Prospective memory exists as commitments and timers, not as a first-class memory type.
Computer operation
LIVE-PROVEN PARTIAL , rung 6 of 8
- Live readings are historical. They were recorded on 16 and 19 September 2026.
- Full-universe installed-app journeys are still owed at release. Live evidence so far covers a bounded set of applications and actions.
- The registry does not live-probe desktop mutations. It skips them on purpose to avoid disturbing the owner's screen, so their registry rows stay TESTED.
- Iframes and shadow DOM are absent from production. Browser typing and filling on the Advanced path are unvalidated.
- Screen grounding was demoted for reachability in an August audit and has not been re-proven live.
Engineering, CAD and fabrication
TESTED PARTIAL , rung 3 of 8
- No print has been started. Physical printing is a concept by design, and there is no evidence JARVIS has ever started one.
- No physical measurement of a produced part appears in the evidence this site uses.
- CAD, the 3D scene, trade spaces, STEP slicing and print readiness are on the Advanced branch only, not merged into production. Their live proofs used synthetic fixtures.
- The CalculiX solve is tested, not live-proven.
- FreeCAD availability in production is uncertain.
- Registry engineering rows are mostly BLOCKED. They are CAP-ENG-03 to 31, and the project later partly corrected the reasons behind those blocks. The unit-safety LIVE reading dates from 17 September 2026.
Research and calibrated reasoning
TESTED PARTIAL , rung 3 of 8
- Live research readings are historical. They come from 16 September 2026 and a journey reading before 17 September 2026.
- TRIBUNAL, calibration and PROMETHEUS are tested, not live-proven, and no calibration figure is published here.
- Repository analysis and video research are blocked in the registry.
- Entailment grading depends on an entailment model. This page does not claim how often that model is right.
- Predict-before-act is not a unified layer. Its Advanced-branch pieces are lane-level and unmerged.
Agents and specialist roles
TESTED PARTIAL , rung 3 of 8
- Specialist decomposition is tested, not live-proven in the registry.
- Council and red-team wiring is uncertain. They may run only in tests and tooling.
- Specialist contracts, the local worker and the blackboard are on the Advanced branch only, not merged into production, and their proofs used fixture tasks.
- Shared resource reservations are on hold.
- No measurement is published here of how much specialist decomposition improves outcomes over a single pass.
Missions: durable long-running work
LIVE-PROVEN PARTIAL , rung 6 of 8
- Live readings are historical. The mission reading is from 16 September 2026, and the journeys are from before 17 September 2026.
- Surviving a machine reboot is not proven. The project explicitly does not accept killing a process as a substitute for a real reboot.
- The intent-ledger gap may or may not be closed.
- Registry mission rows are TESTED, not live-probed.
- Specialist checkpoints are on the Advanced branch only.
- Journey J8, speech interrupted while a mission continues, is fixture-tested only.
The HUD
LIVE-PROVEN PARTIAL , rung 6 of 8
- Live operation is historical. It is from 16 September 2026, and registry HUD rows have decayed to NOT_PROBED_THIS_PASS.
- Recovery after a lost link was not verified live.
- Performance is not yet measured.
- The fidelity gate is not met. The average is 2.43/10 against a gate above 9.0, on an older surface.
- Accessibility is supported by automated and keyboard audits only.
- The 3D engineering scene is on the Advanced branch, not merged into production. See engineering.
Voice
INTEGRATED PARTIAL , rung 4 of 8
- C013 is not technically qualified. Long-form non-regression failed, and first audio, streaming, acronym pronunciation, audible cancellation, restart recovery and fallback have not been measured for it.
- Full-duplex voice has not been proven live. It needs an owner-present sitting, and the interrupted-speech journey is fixture-tested only.
- Reflex first audio is over budget on its recorded reading, and no C013 audible reading exists.
- Recorded figures are dated and limited. The ASR comparison is a July 2026 code-docstring measurement. The barge-in times come from tests. The coexistence run is a single observation.
- No voice samples are published.
Training the local language model
TESTED , rung 3 of 8
- No local model is development-qualified, integrated, live-proven or release-qualified. Production cognition is hosted.
- The SEALED test set has never been opened.
- The context-delivery defect is unresolved, and bounded tool discovery is not integrated.
- Coexistence is a single observation, without latency certification or production routing.
- The latest Gemma continuation has no behavioural result yet. A lower training loss is not a behavioural improvement.
- Every step after training on the production ladder is unreached: SEALED evaluation, integration, release-candidate certification, owner-present checks, promotion, and running without the build assistant.
Health and recovery
LIVE-PROVEN PARTIAL , rung 6 of 8
- Live readings are historical. The health, diagnostics and runtime-service readings date from 16 and 17 September 2026.
- Recovery under real faults has not been proven live. The chaos matrix that would inject failures into a running system is still owed.
- Surviving a machine reboot is not proven. A process kill is explicitly not accepted as a substitute.
- Self-healing is implemented, not live-proven. No live repair is cited here.
- PHOENIX, HELIOS and the boot sequencer's run path have uncertain production wiring.
Perception and vision
IMPLEMENTED PARTIAL , rung 2 of 8
- Physical acceptance on the owner's hardware has not been done. Camera frame rate, calibration precision, follow latency and false-activation rate are all recorded as not probed on the owner's hardware.
- No live gesture has moved a real window in the evidence cited here. The live proof opened no camera and moved no window.
- Pose context failed its probe.
- Screen grounding has not been re-proven live since its reachability demotion.
- No object-detection accuracy figure is published here.