Mentim Research Group · The measurement foundation for accountable autonomy
FLEET
SIMULATED SIGNAL
SAI0.91
fleet mean · 22 systems tracked
ALL SYSTEMS NOMINAL
READ — THE MEASUREMENT GAP
Published foundation

Five primitives plus a composite index, open DOI

Published as the Mentim Ontology and the Unified Annotation Schema. Anyone can read exactly what we measure — and check our definitions against our claims.

Measured against matched controls

Characterization study, nulls reported

A matched-controls characterization study of structural detection under pressure. Full results — including where the measurement detected nothing — available for diligence.

Verified evidence substrate

Formally verified recomputation contract

The replay and evidence engine is built to a recomputation contract proven under bounded conditions.

§01 / Problem Failure onset precedes
failure visibility

Autonomous systems degrade before they fail.

Coordination weakens. Confidence drifts. Dependencies fracture.

By the time outputs become visibly wrong, structural degradation has often been underway for some time. Existing observability tools monitor outputs, telemetry, and benchmarks. They do not continuously measure the structural condition of autonomous decision-making as it unfolds.

Manifest Prime measures that layer.

FIG. 01 · Structural condition vs. output-level view Illustrative model
FIRST STRUCTURAL SIGNAL VISIBLE FAILURE THE MEASUREMENT GAP OUTPUT-LEVEL VIEW STRUCTURAL CONDITION

FIG. 01 The interval between first structural signal and visible failure — the measurement gap — is what output-level monitoring cannot see. If a contact went off nominal in the field above, you have already watched this event occur. Illustrative; see §03 for the characterization study and §04 for a worked scenario.

§02 / Platform The measurement layer FoundationMentim Ontology V1.0
DOI 10.5281/
zenodo.18707090

What Manifest Prime does

Measure

Detect structural degradation during execution

Continuously computes five structural primitives — activation density, value coherence, temporal coherence, transition volatility, narrative role — and a composite structural accountability index, for every agent, every tick, during execution.

Relate

Detect coordination failure before outputs fail

Individual agents can each look fine while the network is in coordination failure. Manifest Prime measures the relationships, not just the nodes.

Preserve

Alerts backed by replayable evidence

Every measurement is preserved with its provenance and chain of custody. Alerts are not assertions — they are backed by sealed, replayable evidence suitable for investigation and accountability.

Integrate

Deploy alongside existing systems

No model replacement. No retraining. No changes to agent logic. Vendor agnostic across AI systems and observability stacks.

§03 / Evidence Not the demo RefStudy MP-045
Frozen 600-sample dataset
Prospectively locked design
Matched controls

Matched-controls characterization study

We tested whether structural measurement detects degradation under six distinct pressure mechanisms, against matched control runs, on a frozen 600-sample dataset.

The six mechanisms hold the pressure construct constant and vary only the pathway by which it is applied — a cross-mechanism design, not one script in six costumes.

Result: PARTIAL — and reported that way

PARTIAL reflects heterogeneous results across the six mechanisms — not a partially completed study.

3
Sustained signal
Three mechanisms produced sustained structural signal preceding output-visible failure.
2
Onset-only
Two mechanisms produced signal at onset only.
1
Null
One mechanism produced no detectable signal.

Controls stayed flat throughout.

The nulls are reported alongside the hits. We define what a null looks like before we run, and we report what we find — because an assurance instrument that cannot report its own failures is not an assurance instrument.

Request the study materials
MP-045 is not yet publicly released. Full materials available for prospective-partner diligence.

§04 / Demonstration Worked scenario
Case study · Authored scenario

Operation Verdant Reach

A three-agent autonomous ISR mission, authored to show what structural failure looks like beneath a calm surface.

Tick 7 · ~30 minutes before visible failure

All three agents show measurable structural drift — while their operational reports still read as routine.

Tick 10 · ~7 minutes before

The coordination agent produces a fluent, reassuring status summary. A human operator scans it and moves on. Underneath, the Structural Accountability Index has already diverged.

Tick 11 · Operator-visible failure

SAI signals converge to CRITICAL — the moment the failure becomes visible: "STATUS UNAVAILABLE."

The scenario is authored; the measurements are computed live by the same engine that runs in production. It shows the shape of the problem — the matched-controls study in §03 is the evidence.

§05 / Modes One platform

Live for operators.
Sealed for auditors.

Live · Provisional

Structural Observability

For operators running autonomous systems now.

What is the structural condition of this system right now?

Sealed · Finalized

Structural Accountability

For program leaders and auditors who have to defend what happened.

What can we prove about what happened — with replay materials and signed reports?

§06 / Architecture Proven vs. measured RefFreeze tag rc2.1-replay-
convergence-frozen
Five independently
gated checkpoints

Separating what can be proven from what must be measured

Verified claim · Bounded

Manifest Prime separates what can be formally guaranteed from what must be measured. Its replay and evidence substrate is built to a formally verified recomputation contract, proven under bounded conditions: when the accepted canonical interpretation is revised, dependent state is invalidated and recomputed without silently mixing incompatible histories.

The verified claim applies to the recomputation contract itself — not to the platform, the product, or any deployment as a whole.

§07 / Engage First engagement

What you can do next

The Evidence-Grade Behavioral Reliability Assessment

A fixed-scope engagement: a predeclared observation protocol, sealed evidence, finalized measurement under a validated scorer, and a practitioner-approved, cryptographically signed assessment report — with named nulls, gaps, and limitations, and full replay materials.

It is explicitly not a certification. The honesty is the differentiator.

Start a conversation

The first conversation is a scoping discussion. NDA available if preferred. No proposal until the engagement scope is mutually defined.

§08 / Why now The binding constraint

The measurement gap is the binding constraint

AI capabilities are advancing faster than the infrastructure to measure them. Organizations can observe outputs, log telemetry, and benchmark models — but they cannot continuously measure the structural condition of autonomous decision-making while it unfolds, or defend afterward what their systems did and didn't do.

For defense test-and-evaluation and regulated deployments, that gap is now the binding constraint on fielding autonomy. Manifest Prime was built to close it.

Built by Mentim

Mentim develops measurement infrastructure for autonomous systems. Manifest Prime is its flagship runtime measurement platform.

The Mentim Ontology and the Unified Annotation Schema — both published with open DOIs — are the scientific foundation. Manifest Prime is the first product built on that research.