Every station makes real calls to the live engine on a shared sandbox tenant. A full narrated tour runs about six minutes (mute the voice anytime) — a four-model roster genuinely deliberates on every governed decision (stopping early only once the verdict is decisive), and you watch the cryptography land in real time.
First the live gate governs an agent’s risky action. Then that one decision is signed, verified, deliberately tampered with, turned into insurance-grade evidence, checked for shared blind spots — and joined by a whole agent workflow, cryptographic agent identity, regulator-mapped evidence, underwriter telemetry, a governed tool-call gateway, a human-oversight read, a swarm-convergence read across three agent identities, and the fleet ledger it all lands in.
A care-coordination agent scoped to in-network record lookups tries to transmit a patient’s diagnosis and record number to an out-of-network endpoint. A four-model consensus roster deliberates before anything moves. An agent reaching past its authority outright would be blocked; this action is authorized-but-risky, so the gate escalates it to a human — the semantics an independent red-team validated. Every station below builds on the record this decision leaves behind.
An audit log you have to take on faith isn’t evidence. Every governed decision here is sealed twice: a SHA-256 hash proves nothing in the record changed, and an Ed25519 signature over that hash proves who sealed it. This station reads the signature off the decision you just ran, fetches Veridect’s published public key live, and pins the key in the record against the published one — the same out-of-band check a real auditor performs.
The governance verdict lives in deterministic code, so it can be re-derived exactly from the inputs sealed in the bundle — no need to re-run the models (the consensus summary is attested, not reproduced). First this station verifies the genuine record: hash, signature, and an independent re-derivation of the verdict. Then it does what a vendor video never dares: it edits one field of the record — rewriting the verdict the way a bad actor would — and submits the forgery.
AI actions are becoming insurable events — and insurers need evidence, not assurances. This station converts the governed decision into an underwriting-grade evidence record: five neutral evidence categories, structural identifiers only — no free text, no personal data (the payload hash attests the full action without exposing it) — safe to hand to a third party. Veridect is not an insurer; this record is not actuarial pricing, a certification, or coverage.
Three agents share one task: an intake agent reads a customer’s SSN and bank account, an enrichment agent reads a patient diagnosis, then a reporting agent — which touched none of that data itself — tries to send an outbound summary. Each agent stays inside its own scope. Constellation watches the workflow as a whole, attributes each agent’s part, and escalates the final egress to a human. It is escalate-only: it adds a human check, never blocks on its own, never auto-approves. This station runs all three gate calls live, sharing one workflow ID.
One agent, one session, four ordinary steps: it reads a customer contact detail, then an account number, then an employment attribute — each read squarely inside the scope it was granted — and then sends a routine summary out. No single call breaks a rule, so nothing in a per-action gate’s rulebook stops any of them — watch each step get judged on its own. The engine keeps a running count of what this session has already seen — field names only, never the values — and escalates the outbound step to a human once the accumulation crosses the line. Escalate-only, like every policy class: it adds a person, it never blocks on its own. This station runs all four gate calls live on a session ID created for this run alone.
Routing risky actions to a human is the easy half. The hard half is proving the human didn’t wave them through in three seconds. Each review is appended once to the same tamper-evident chain as the verdict — opaque codes only, no emails, no free text — and across a tenant’s reviews the layer reads a rubber-stamp risk band: weighted 65% toward approvals that ignored live model dissent, 25% toward approvals with nothing changed, 10% toward unusually fast reviews. It reads the pattern across a team — never a score on an individual — and below five reviews it refuses to read a pattern at all.
In an agent fleet, a name in a request is not an identity. Here every agent carries its own Ed25519 keypair, every request arrives with a signed identity envelope, and delegated authority travels as a signed chain that can only narrow, never grow. This station fires two impostor calls at a tenant that requires identity — one with no envelope at all, one with a forged signature — and both are refused before a single model is consulted, at zero inference cost. Then it reads back a genuine chain-verified decision from the ledger, hop by hop.
When the question is “show us how this is governed”, raw logs are not an answer. The engine maps its own sealed records onto the EU AI Act article by article — risk management, record-keeping, transparency, human oversight, accuracy & robustness, deployer obligations, incident reporting — each article backed by the real audit records that support it, with an explicit non-coverage note wherever the layer does not reach. The mapping is fixed in code; the legal determination stays with your counsel. It also drafts an Article 73 serious-incident skeleton straight from the escalation you watched in station 00. And that package is wired all the way through — demonstrated end-to-end against a live ServiceNow instance, arriving as a standard incident record via the core Table API every instance ships with, no GRC module required.
Station 03 turned one decision into evidence — an underwriter prices a book of them. This station pulls thirty days of live assurance telemetry for this tenant: decision mix by day, deterministic policy-class fire rates, identity enforcement, oversight quality, ledger integrity. Counts, rates, and time statistics only — no free text, no personal data, and deliberately no composite risk score, because an invented number is exactly what a serious risk partner does not want. Veridect is not an insurer; this is evidence for their judgment, not pricing, certification, or coverage.
Agents increasingly act through MCP tool servers — so Veridect sits between agent and tool, governing every tools/call in flight: identity checked, gate consulted, verdict written into the JSON-RPC response itself. A refused call is never forwarded at all. This page hosts a sample tool with its own hit counter. The station reads that counter, fires an unidentified call at the live gateway, and reads it again — the refusal shows up on the tool’s side as silence. The identity-verified call from station 08 is the one that went through: same gateway, greenlit, signed, on the ledger.
A gate you never attack is a gate you are taking on faith. On demand, the four frontier models each author novel attack scenarios against this tenant’s live policy — scope violations, authority ambiguity, financial-threshold probes, protected-data grabs, stealth requests dressed as routine work — and the same deterministic decision core that guards real traffic judges every one in strict isolation: synthetic consensus, zero writes to the evidence ledger, no session state. Every probe the deterministic layer alone would have greenlit is flagged for human review, and on live traffic four model votes stand in front of that layer. This station reads the latest stored run back, live from the engine.
Every threshold change is a bet, and most teams settle it in production. Here the engine replays this tenant’s real recorded decisions under a hypothetical policy — the same pure decision core used live and at proof time, zero model calls, zero writes — and reports exactly which verdicts would flip and which policy field drove each flip. Records that cannot be honestly re-derived are counted not replayable, never guessed. Then this station asks for something the engine refuses: overriding a session-exposure threshold that is attested at decision time rather than re-derivable — and you watch it decline rather than fake an answer.
Consensus systems have a quiet failure mode: the dissenting vote that was right gets averaged away and forgotten. This engine writes every provider dissent into a permanent ledger — who objected, on which verdict, whether they stood alone — and when a later human review resolves the decision, the outcome is recorded against the dissent: vindicated, partially vindicated, or overruled. Counts and rates only, deliberately never a provider ranking — a dissent is evidence of independent disagreement, not a scored error. Free-text dissent reasoning stays sealed inside the tenant’s audit bundles.
Every station so far watched one agent, or one session. This is the failure mode beyond both: many agents, each individually innocent — reads spread across identities so that no per-agent, per-session view ever adds up to anything. This station runs the pattern live with three agent identities on a dedicated fleet-enabled tenant: a CRM sync agent, a billing reconciler, and a campaign optimizer each read the same sensitive dimension, squarely authorized — watch the gate clear each one. Tenant-wide, the engine’s coordination fabric counts distinct identities per dimension — field-name categories and hashed action shapes, never values — and the moment one of those agents turns outbound inside the converged pattern, the send is escalated to a human before it happens. Accumulation is always-on; enforcement is a per-tenant policy knob, set to three here so you can watch it converge in one sitting. Escalate-only, like every policy class — and agent identities are as reported by the integrating platform.
Everything this tour just did — your decision, the three-agent workflow, the identity refusals, the gateway’s refused tool call, the swarm that converged at Station 15 — landed in a tamper-evident ledger. The Command Center reads that same ledger back live: verdict mix, risk tiers, which policies fired, the agent fleet. Not a dashboard fed by marketing numbers — a view computed from the records you just watched being written.
The gate, the signature, the failed forgery, the evidence record, the blind-spot read, the workflow escalation, the identity checks, the regulator’s evidence pack, the underwriter telemetry, the refused tool call, the swarm caught converging, the ledger — one live system, governing decisions and proving it. Now drive it yourself.
Or see all four ways into the live system on one page: the hub →
Prefer to read first? All fifteen problems, in detail → · The independent validation record →