Scientific Metrics
ECAM-X, AEI-Delta, Human Gate records, and AIJIM TCRE remain separate authorities.
No super-score
Interpretation Matrix
These contracts answer different scientific questions. The most important column is what each does not measure: it prevents scope creep, trust-washing, and accidental authority transfer between protocol records, governance, and research instruments.
| Axis | Scientific Question | Measures | Does Not Measure | Authority | Falsifiable Via | Reader Label | Audit Depth | Fail-Closed Form | Paper Section |
|---|---|---|---|---|---|---|---|---|---|
| ECAM-X | Does this evidence support this claim? | NLI label, similarity, evidence-link coverage, negative controls, judge configuration, and attribution confidence. | Causality, source bias, editorial relevance, explanation stability, or human reliance quality. | Publication-required for claim/evidence support metadata. | Canonical JSON, model snapshot/hash, evidence content hash, judge config, and reproduced NLI link output. | Evidence grounding | Run IDs, evidence hashes, thresholds, NLI labels, similarity scores, and judgeConfig. | claim_missing_evidence_links or explicit evidence grounding unavailable. | Attribution Validity |
| AEI-Delta | Does the explanation remain stable under perturbation? | Seeds, perturbations, delta distributions, confidence intervals, sanity checks, and stability drift. | Truth, evidence quality, causality, bias, or editor trust. A stable explanation can still be wrong. | Publication-strengthening disclosure. Absence is honest not-measured, not failure and not green. | Seeds, perturbation recipe, model snapshot, stability artifact, and deterministic delta recomputation. | Stability evaluated / not evaluated | Perturbation set, delta histograms, CI bounds, failed sanity checks, model snapshot. | stability not evaluated; never synthesize a neutral stability score. | Explanation Stability |
| AIJIM TCRE | How did item-level decisions change after advice or bound evidence, relative to a preregistered independent item-quality reference? | When all study gates close: beneficial update, missed benefit/under-reliance, harmful deference/over-reliance, resilient self-reliance, and preregistered task-weighted benefit/harm, reported separately. | Authorization quality, aggregate override activity, factual truth, a global trust score, or universally appropriate reliance. | Proposed behavioural research instrument. Current status: PROPOSED / NOT_VALIDATED / NOT_MEASURED. | Preregistered task and adjudication protocol, independent item-quality reference, participant/item/task/condition identities, exact pre/exposure/post decisions, uncertainty rules, HREC, validation evidence, and independent replication. | Reliance evaluation not measured / cannot verify | Task contract, reference provenance and status, item transitions, separated outcomes, uncertainty, missingness, and exclusions. | Missing reference or transitions means NOT_MEASURED; disputed, leaked, or under-reliable reference means CANNOT_VERIFY. Human Gate history is never the reference. | Task-Conditioned Reliance Evaluation |
Measurement axes are not protocol loci
Triple Falsifiability asks where a publication record can fail: Producer Conformance (P), Surface Observability (S), or Audit Replicability (A). The measurement axes below ask what scientific property was measured. A measurement artifact can be challenged at any of the three protocol loci.
Bundle integrity
Core hashes, signatures, and canonical JSON bind the bundle. ECAM-X supplies reproducible claim/evidence attribution metadata inside that signed integrity envelope.
Experimental
AEI-Delta exposes seeds and perturbations so stability can be re-run like a reproducibility capsule.
Audit-log
AIJIM TCRE requires participant/item/task/condition observations, exact pre/exposure/post decisions, and an independently assessed preregistered item-quality reference. Ordinary authorization and override logs cannot satisfy that reference gate.
Two-Tier Surface Discipline
| Reader UI | Quiet labels: evidence grounding, stability evaluated/not evaluated, and TCRE not measured/cannot verify. No raw run IDs or scores. |
| Editor UI | Actionable state: which axis is missing, blocked, measured-positive, or measured-negative. |
| Audit UI | Technical depth: run IDs, hashes, thresholds, model snapshots, verifier outputs, and RFC conformance warnings. |
| Publication Bundle | Immutable disclosure: first-class publicationId plus axis artifacts that actually exist. |
| Verifier CLI | Strict publication-integrity PASS/FAIL first; EU AI Act readiness as PASS/PARTIAL/FAIL; optional quality report second. No quality score may weaken cryptographic verification. |
| Protocol Paper | P/S/A protocol loci are reported separately from ECAM-X, AEI-Delta, Human Gate governance, and AIJIM TCRE. |
Honest absence
EU AI Act readiness
model_card.v1,risk_register.v1, andevaluation.v1. PARTIAL is an honest-absence state, not a score and not a cryptographic publication failure.Verifier Stance
aijim-verify remains strict about publication integrity: publication identity, body hash, bundle hash, Merkle root, schema version, and mandatory ECAM-X support metadata must verify or the publication does not pass. AEI-Delta and AIJIM TCRE remain separate optional artifacts. A verifier may check their schemas, bindings, status, and Honest Absence, but cannot claim scientific validity or turn a cryptographic failure into a pass.
Protocol Paper Framing
The Protocol Paper reports bounded Triple Falsifiability through Producer, Surface, and Audit loci. ECAM-X and AEI-Delta remain bounded evaluation disclosures; Human Gate records remain governance evidence; AIJIM TCRE remains a proposed research instrument with its own validity and Honest Absence gates. None is a universal truth or quality score.