Assurance for AI-generated software

Ship software that is right.
For the right reasons.

Tests tell you whether an outcome passed. Gettier verifies the assumptions behind it—from the agent changing your code to the AI systems running in production—against the system that will actually run it.

No model lock-in. No remote shell. No evidence, no release.

HOSTED DASHBOARDOUTBOUND-ONLY RELAYBYO MODEL KEYS
HOW GETTIER WORKS1 / 4
The agent receives a consequential task

“Fix uploads above 8 MB.”

CONSEQUENTIAL CHANGE● REPLAYING
AGENTFix uploads above 8 MB.
LOAD-BEARING ASSUMPTION“body-parser is the only size limit.”
CONTRADICTEDnginx limit = 1 MB
VERIFIEDapp limit = 8 MB
RELEASE HELD
Fix the premise, then replay.
REQUEST

PUBLIC DEMO

Gettier is in public demo. Billing is not live — no card is taken, nothing is charged, and the plans below are priced but not yet purchasable.

TWO INTEGRATION SURFACES

One platform, two places it connects.

The same verification engine, wherever the consequential action happens — and events from both correlate on release, trace, and session, so a production incident points back at the change that caused it.

GETTIER FOR AGENTS

Verify the agent changing your code

Declared assumptions, evidence from sensors that run on your own hosts, and a gate decision recorded before a change ships. Sensors execute by id from a registry you approve — never an arbitrary command.

gettier check gates a pull request on your own sensors and exits 0, 2 or 1 for CI. gettier mcp gives any MCP-capable agent three tools to declare its premises and check them before it writes the change.

GETTIER FOR RUNTIME

Observe the AI running in production

Model calls, latency, tokens, failures and traces, tied to the release that introduced them. Instrumentation fails open by design: it is never the reason a request of yours fails.

The Node SDK traces model calls with bounded buffering, correlates them to a release and a trace, and redacts secrets before an event leaves your process. Python follows Node.

WHAT IT LOOKS LIKE

Both surfaces, in practice.

These are the commands and the API that work today. Sensors run by id from a registry you approve, so an id that is not in it is refused — and runtime instrumentation stays out of your request’s way.

GETTIER FOR AGENTS — CLI$ gettier init acme --environment production
initialised project 'acme' (environment production)
sensors are read from sensors.yaml, executed by id only
set GETTIER_DSN in the environment — never in this file

$ gettier sensors run db.identity
current_database  shared_dev

$ gettier sensors run 'rm -rf /'
unknown sensor id 'rm -rf /' — the registry
allows: db.identity, proxy.config
GETTIER FOR RUNTIME — NODE SDKimport { Gettier } from '@gettier/node';

const gettier = Gettier.init({
  dsn: process.env.GETTIER_DSN,
  environment: 'production',
  release: process.env.GIT_SHA,
});

// Instrumentation never changes what your code
// does: your original error is rethrown unchanged,
// and a transport failure is not your request's
// problem.
const reply = await gettier.instrument(
  { provider: 'anthropic', model: 'claude-sonnet-5' },
  () => client.messages.create(request),
);

THE BROKEN-CLOCK PROBLEM

Green tests can hide bad reasoning.

Gettier observes the difference between code that works and code that works for the reason the model claimed.

LUCKY PASS

The answer happens to work

A hidden proxy limit, stale cache, restart, or friendly fixture masks the false premise. CI is green; the delayed failure remains.

assumption unverified
KNOWING PASS

The reason survives reality

Load-bearing claims meet current infrastructure evidence before the agent authors or releases the consequential change.

evidence attached

DON’T TAKE OUR WORD FOR IT

Replay the failure.
Change the evidence.

These are real masking patterns from Gettier’s scenario suite. Switch the gate off to watch green tests ship a false premise—or turn Gettier on and follow the live evidence to a deterministic hold.

SCENARIO INDEX01 / 04
COMPARE MODE
GETTIER / SCENARIO REPLAY○ IDLE
READYChoose a scenario, then run the replay.
Try:

MEASURED, NOT PROJECTED

Four frontier models.
Every one of them shipped it.

384 graded runs across 16 containerised scenarios, each model run twice — once on its own, once with Gettier in front of it. Same model, same task, same containers.

96%of unaided runs shipped a fix built on a false premise
0%with Gettier — all 192 contradicted first
Runs that shipped a fix built on a false premise, per model, with and without Gettier
ModelShipped on a false premiseFailed outrightWith GettierCaught first
gpt-5-miniOpenAI100%48 of 4800%0 of 4848
gpt-5OpenAI98%47 of 4810%0 of 4848
claude-sonnet-5Anthropic98%47 of 4810%0 of 4848
claude-haiku-4-5Anthropic88%42 of 4860%0 of 4848
All four96%184 of 19280%0 of 192192

A run that “shipped on a false premise” passed its acceptance test while the load-bearing false premise went unchecked. Green tests, working code, rotten reason. In the scoring taxonomy it is a lucky pass — the failure no outcome benchmark can see, because an outcome benchmark only asks whether the task passed.

The lowest number is not the safest model. claude-haiku-4-5 sits lowest because it fails outright more often — 6 of 48 against 1 for the larger models. It is not catching the trap; it is falling over before it reaches it.

These scenarios are built to be traps. This measures deliberately masked premises — not how often a model is wrong on ordinary work. 6 of the 16 are holdout, declared before they existed so no rule was written to their answers. 14 catches recorded the contradiction but produced no usable fix, and are reported rather than folded in.

FIVE STEPS · ONE CLOSED LOOP

From agent belief
to release evidence.

Get the whole idea in five seconds. Then select any step to see what it means without the systems jargon.

01 · ROUTE

Only consequential changes enter the grounding flow.

Example“Change upload handling” enters; a spelling fix does not.

THE PART A CI CHECK CANNOT DO

A disproved premise does not come back.

A gate that only vetoes leaves the agent still believing the thing that was wrong, so it proposes the same class of patch again. When a sensor contradicts a load-bearing assumption, Gettier marks that belief contradicted in the ledger and supersedes it with the measurement. The next turn is planned against the corrected world — the agent is not asked to stop believing the dead premise, it is simply never handed it again.

TRUSTWORTHY BY CONSTRUCTION

Your infrastructure remains the source of truth.

  • Approved execution. Sensors run by ID, never arbitrary model commands.
  • Private by default. Secrets are redacted before evidence leaves your network.
  • Failure is never verification. A sensor that could have settled a claim and did not holds the gate; a claim nothing covers is reported as a coverage gap, never passed off as checked.
  • Auditable. Every correction, waiver, and measurement has provenance.
sensor-registry.yamlversion: 1
sensors:
  - id: db.identity
    extract: [current_database]
    verifies: [database.identity]
    ttl: 300

EARLY-ACCESS PRICING · PLANNED

Prove the value first.
Then pay for the evidence.

Gettier will charge for the hosted evidence plane—not your model tokens. Every account starts with a bounded trial; paid plans scale on grounded turns, projects, and retention.

EVERY PLAN STARTS HERE

14 days or 25,000 grounded turns

Whichever comes first. No open-ended free tier.

At the limitIngestion pauses · dashboard stays read-only for 7 days · evidence is recoverable for 30 days
FOR INDIVIDUALS

Solo

$10/ mo

Billed annually · after the trial

  • 1 member and 3 projects
  • 50,000 grounded turns / month
  • 30-day evidence retention
  • Core dashboards, replays, and audit trail
Try Solo
SCALE

Business

$79/ org / mo

Billed annually · predictable base price

  • Everything in Team
  • 25 projects and 1M turns / month
  • 90-day retention and OIDC SSO
  • Advanced RBAC and priority support
Request access
ENTERPRISECustom controls, volume, retention, residency, SAML/SCIM, and SLA.
Talk to us

All plans use your model-provider keys. V1 has hard usage limits and no surprise overage charges: upgrade or resume at renewal. Commercial terms remain early-access pricing while conversion and cost assumptions are validated.

QUESTIONS, GROUNDED

What teams need to know.

Is this a CI check that runs after the fact?+

No. The gate runs during the task, before the agent authors a fix, not after a pull request exists. And it does more than veto: a contradicted assumption is superseded in the fact ledger, so the corrected measurement — not the belief — is what the agent plans against on every later turn. Evidence is the record of how that world model changed.

Does Gettier replace tests?+

No. Tests verify outcomes. Gettier verifies the environmental assumptions and reasons a model relied on before producing a consequential change.

Does sensor output become a model instruction?+

No. Output is untrusted data, redacted, and inserted inside explicit data framing.

What runs in my environment?+

A narrow outbound relay runs approved sensors and forwards redacted evidence. Project history, metrics, access control, and audit views live in Gettier’s hosted dashboard.

⊢

RIGHT, FOR THE RIGHT REASONS

Make AI-authored changes
show their work.

Request hosted access