The maximised research panel: numbered findings about how Webb stays cold, three NASA and ESA sources with retrieval times, and page captures as read.

Research and calibrated reasoning

JARVIS researches in an isolated browser, turns sources into graded evidence, weighs independent origins rather than document counts, and labels every conclusion as fact, inference, recommendation or action. Its confidence is scored against outcomes it recorded in advance.

TESTED , rung 3 of 8 Cited research was driven live on 16 September 2026, and journey J5 was verified live before 17 September 2026. Both are historical readings. TRIBUNAL, calibration and the claim graph are TESTED, and predict-before-act exists as distributed behaviour, not as a layer.

PARTIAL Parts of this system sit at different rungs. The breakdown below shows each one.

As of
Parts tracked
7
Catalogued capabilities
16
On the unmerged branch
1
Real capture, captured 22 September 2026. Asked how Webb keeps its instruments cold, JARVIS read 3 of 3 sources from 3 independent origins in 19.9 s and marked every finding with the source it rests on. ShowsThe research desk retrieves real pages, keeps a capture of each, and ties every finding to a source. Does not showThe spoken summary was thin: it mostly repeated a source title. Source count is not a calibrated confidence.
What happens, step by step
  1. 0:00 Asked: “How does the James Webb Space Telescope keep its instruments cold?”
  2. 0:04 3 of 3 sources answered · the research panel opens
  3. 0:09 Each finding names its source · pages captured as read

Captured at revision 893913a9c

The maximised research panel: numbered findings about how Webb stays cold, three NASA and ESA sources with retrieval times, and page captures as read.
Real capture, captured 22 September 2026. Every finding is tagged with the source it rests on; each source keeps a capture of the page as it was read.

Captured at revision 893913a9c

A rendered research artifact answering a question about Python's asyncio.timeout, with sections for question, retrieval, sources and findings, each finding quoted with its source and fetch time.
Real capture, captured 12 September 2026. A research artifact produced by JARVIS’s research desk on 12 September 2026. Findings are quoted as documented facts with their sources; what could not be established is listed separately.

Status breakdown

Where each part stands

1 live-proven 1 production-reachable 3 tested 1 implemented 1 concept

  1. Cited research through ResearchDesk Capability model RESEARCH driven live 2026-09-16; journey J5 LIVE_VERIFIED before 2026-09-17. Historical.
    LIVE-PROVEN , rung 6 of 8 OBSERVED Where: Production branch
  2. Task simulation before a mission Body-matrix row A44. Output labelled SIMULATED / not_proof.
    PRODUCTION-REACHABLE , rung 5 of 8 DOCUMENTED Where: Production branch
  3. Evidence, claim graph and citation grading Registry research rows are TESTED.
    TESTED , rung 3 of 8 TESTED Where: Production branch
  4. TRIBUNAL, calibration and PROMETHEUS Recorded as implemented and tested; no live reading.
    TESTED , rung 3 of 8 TESTED Where: Production branch
  5. Predictions bound to observation receipts Operator lane only; 'record prediction before dispatch' still open for integration. Not merged into production.
    TESTED , rung 3 of 8 TESTED Where: Advanced Systems branch (not merged)
  6. Repository analysis and video research Registry CAP-RES-10 to 14 are BLOCKED.
    IMPLEMENTED BLOCKED , rung 2 of 8 DOCUMENTED Where: Production branch
  7. Predict-before-act as a named layer No component has this name. The behaviour is spread across calibration, simulation, pre-mortems and anticipation.
    CONCEPT PARTIAL , rung 1 of 8 NONE Where: Production branch

Pipeline

How it flows

A research question

  1. Question
  2. ResearchDesk isolated browser, robots.txt, private hosts refused (can stop the request)
  3. Sources to Evidence
  4. Claim graph
  5. Citation check lexical match capped at CONSISTENT (can stop the request)
  6. TRIBUNAL weighs independent origins (can stop the request)
  7. Epistemic card FACT, INFERENCE, RECOMMENDATION or ACTION
  8. Cited artifact on the HUD
Amber steps cap or refuse a claim's strength.

Prediction before action (distributed, not one layer)

  1. Plan
  2. Simulate SIMULATED / not_proof
  3. Record prediction
  4. Act through governance (can stop the request)
  5. Observe
  6. Score calibration no figure under 20 samples
Amber stations can stop the request.

Diagrams are simplified from the code paths named in the sources below. They are illustrative, not screenshots.

A research answer is only as good as the reader’s ability to tell how strong it is. JARVIS’s research system is built to preserve that strength signal from the first page fetched to the sentence spoken. The same discipline extends to its predictions: it records what it expects before it acts, so its confidence can be checked later.

Browsing without exposure

ResearchDesk browses in an isolated Playwright context, separate from the owner’s own browser session. It obeys robots.txt and refuses private network hosts, so a page cannot steer it into the owner’s local network. Research adapters never execute a tool and never grant authority. Text on a web page is evidence to weigh, not an instruction to follow.

Evidence and the claim graph

Each source becomes an Evidence record, and claims and their support are held in a ClaimGraph. Support is graded, and the grading is deliberately conservative:

  • Lexical checks are capped at CONSISTENT. When a check finds a claim’s words in a source, the most it can say is that the claim is consistent with the source. The project’s rule is that “a keyword match must never be laundered into ‘verified’”.
  • Only an entailment model may say ENTAILED, meaning the source actually implies the claim.
  • A retracted source supports nothing. Anything that rested on it loses that support.

Why cap keyword matching? A sentence can contain every word of a claim and still say the opposite: “the study did not find that X causes Y”. Treating word overlap as proof would inflate confidence exactly where it is least deserved.

TRIBUNAL: counting origins, not documents

Contested questions go to TRIBUNAL, which has prosecution, defence, provenance and judge roles. The judge weighs independent origins, not the number of documents: “Five documents citing one study are one source wearing five hats.” A ruling is UPHELD, CONTESTED, UNSUPPORTED or WITHDRAWN, and each ruling states what evidence would change it.

Stating what would change a ruling makes it falsifiable. It also tells the owner what to look for if they disagree.

Four kinds of card

Every result card is classified by the EpistemicEngine as one of four kinds:

Kind Meaning
FACT Supported by evidence strong enough to state directly
INFERENCE Reasoned from evidence, not observed directly
RECOMMENDATION A suggested course, open to the owner’s judgment
ACTION Something to be done, which still has to pass governance

The same classification feeds the daily briefing. There, actions rank above recommendations, recommendations above inferences, and inferences above facts. In speech, answers are marked as confirmed, inferred, incomplete or unknown.

Calibration

A system that says “90% sure” should be right about nine times in ten. JARVIS’s calibration module records each prediction before its outcome is known, then scores the predictions once outcomes arrive. It uses two standard measures:

  • Brier score: the average squared gap between the stated probability and what happened.
  • Expected calibration error (ECE): how far stated confidence drifts from observed accuracy across confidence bands.

No calibration figure is published from fewer than 20 samples, because a handful of predictions can look perfectly calibrated by chance. Recording before the outcome matters just as much: a prediction written down afterwards cannot be honestly scored.

Hypotheses and experiments

A hypothesis module ranks competing explanations and plans the next experiment by expected information gain, the test whose result would most change the ranking. PROMETHEUS, a deterministic invention pipeline, generates candidate ideas. Its records never auto-promote into owner memory, so a brainstorm cannot become a remembered “fact”. See memory.

Predict-before-act is a behaviour, not a layer

The phrase “predict before act” describes something JARVIS does, but no production component has that name, and this site does not present it as one. The behaviour appears in several places:

  • Calibration records a prediction before the outcome.
  • A mission can be simulated before it runs. The result is labelled SIMULATED / not_proof, and “Simulation never grants authority or certifies effects.” See missions.
  • Pre-mortems run before execution.
  • The anticipation component makes predictions, but “A PREDICTION NEVER REACHES THE EXECUTOR”.
  • On the Advanced branch, which is not merged into production, the operator lane binds experiment predictions to the receipts that later observe them.

The Advanced work is still unfinished. Its item “record prediction before dispatch” is open for integration, and a matrix row named “Predict-Before-Act Simulation Layer” has not been dispositioned. As a layer, predict-before-act is a concept with partial pieces.

What is not proven yet

  • Live research readings are historical. They come from 16 September 2026 and a journey reading before 17 September 2026.
  • TRIBUNAL, calibration and PROMETHEUS are tested, not live-proven, and no calibration figure is published here.
  • Repository analysis and video research are blocked in the registry.
  • Entailment grading depends on an entailment model. This page does not claim how often that model is right.
  • Predict-before-act is not a unified layer. Its Advanced-branch pieces are lane-level and unmerged.

Invariants

Rules the code enforces

  • a keyword match must never be laundered into 'verified'

    jarvis/research/ (claim graph citation rules)

  • Five documents citing one study are one source wearing five hats

    jarvis/research/tribunal.py

  • A PREDICTION NEVER REACHES THE EXECUTOR

    anticipation module (production)

  • Simulation never grants authority or certifies effects.

    docs/ledger/CURRENT_BODY_EVIDENCE_MATRIX.md (A44)

Capabilities

Related capabilities

All 16 catalogued capabilities in this area

Sources

Sources

Paths are relative to the private JARVIS repository. They are listed so the claims above can be audited by the owner and reviewers; the files themselves are not published.

  • doc docs/analysis/JARVIS_CAPABILITY_MODEL_2026-09-16.md
  • ledger docs/orders/post-lm/ACCEPTANCE_JOURNEYS.json
  • ledger docs/ledger/CAPABILITY_TRUTH.json
  • ledger docs/ledger/CURRENT_BODY_EVIDENCE_MATRIX.md