Actuary

AI quality regression monitoring

Catch AI quality regressions before customers do. Then prove it.

You pick a small set of test cases and the score you never want to fall below. Actuary checks your live system against them, over and over. One bad score stays quiet. When the evidence shows quality has really dropped, your team gets one alert you can trust, and the incident goes on the record.

Start free

Two monitors free, forever. We set the first one up with you.

Here is one, running as you read. It watches an agent that reads invoices, and it was told the lowest acceptable quality before any of this data arrived. Nothing below can be edited after the fact.

MONITOR · INVOICE-READING AGENT 5 TEST CASES · LOWEST ACCEPTABLE SCORE 0.70 STABLE
QUALITY · AVERAGE TEST SCOREYOUR MINIMUM1.000.900.800.600.70EVIDENCE THAT QUALITY HAS DROPPEDALARM LEVEL205210.78CHECKS →
How long has it watched? 20 Checks finished so far
Is quality holding up? 0.902 Average score on the test cases above the 0.70 you set
Should anyone be woken up? 0.86 of 20 Evidence that quality has really dropped nowhere near enough to raise an alarm

All of this is running in your browser, not playing back a video. Break something below and watch what the evidence does.

What happens if this alarm goes off

One alert in Slack and email, one incident page on the record, and a signed report if we froze the design with you first

Things that go wrong in real systems — try one

It is slow on purpose. Evidence only builds while quality is genuinely below the line you set, so a bad afternoon does not wake anybody.

SELF-CHECK AT LOAD · TURN ON JAVASCRIPT AND THIS PAGE REDOES CERTIFICATE 001 IN FRONT OF YOU

This is the same code that produced our certificate of 24 July 2026, and it lands on the same numbers here in your browser. The failures are pretend; the mathematics is not. Once you break something, the clock speeds up so you are not left waiting.

A made-up incident, to show what happens next

One alert everybody trusts, and a team that can act on it.

Team response
Invented people, invented company — but this is exactly how the alert arrives
09:14 → 09:33 UTC

# eng-alerts

Production alerts and model-quality incidents
MCNPRK
Today
ActuaryAPP09:14

Threshold crossing for tenant northwind — invoice-extraction quality fell below the frozen floor.

Extraction quality below frozen floorAlarm

This is sustained evidence, not a single bad score. The estimated onset narrows the change window for the team.

Model
gpt-5-mini
Task family
extraction
Evidence
29.96 of 20
Estimated onset
Obs. 8 · 09:06
Open incident
Alert ID
0b41e2c7-9f3a-4d21-b6ce-8a5f21c704d9
Stream
4c7d1f60-2b58-4a0e-9c3d-1d6f0b7a5e42
Maya Chen09:33

Closing this. The failure is mitigated and the lesson is now part of the next registered suite.

Incident 1042 · resolution capturedResolved
Root causePrompt v42 omitted the currency instruction
MitigationRolled back to prompt v41
ImprovementFailure added to suite v4; v3 remains frozen
RecordEpoch lineage v3 → v4 preserved
Open incident record

The alert gives everyone the same starting fact. Your engineers still do the thinking and the fixing; Actuary keeps the evidence, the test cases and the history of every change.

Start free

Uptime, latency, error rate: every dashboard you already have can stay green while the answers themselves get worse. None of them watches the quality signal.

Sept 17, 2025 · Anthropic engineering postmortem

“The evaluations we ran simply didn’t capture the degradation users were reporting… we lacked a clear way to connect these to each of our recent changes.”

You can pin the model version. You cannot pin how it behaves.

It took the provider six weeks to see this inside their own system, with all of their instruments. You are one API call downstream, usually with nothing watching at all.

Uptime tells you the system answered. Actuary tells you whether the answer was still good enough.

What you get

A history you can hand to someone else.

Every monitor keeps its own page: the test cases you locked in, how quality has moved since, and every change anyone made to the setup. None of it can be quietly rewritten later — not by your team, and not by us.

SUPPORT AGENT · gpt-5-mini · THIRD SETUP · TEST CASES LOCKED BEFORE ANY DATA · HISTORY v1 → v2 → v3 EXAMPLE, NOT REAL DATA
0.600.700.800.901.00QUALITY · THE SAME TEST CASES, ONCE A DAYYOUR MINIMUM 0.70DAYS → BAR HEIGHT = HOW SURE WE AREHONEST HOWEVER OFTEN YOU LOOK
ONE BAR PER DAY: WHERE IT SITS IS THAT DAY’S AVERAGE SCORE, HOW TALL IT IS SHOWS HOW SURE WE ARE · DAYS WITH NO DATA ARE LEFT BLANK RATHER THAN GUESSED AT · ALERTS GO TO SLACK AND EMAIL · WE PUBLISH THE SAME PAGE FOR THE MODELS WE WATCH OURSELVES →

If you have an LLM feature in front of real users, and you can send us one quality score a day, that is enough to start. There is nothing else to set up.

POST /v1/scores
{"suite":"support-agent","suite_version":3,
 "model":"gpt-5-mini","score":0.86,
 "ts":"2026-07-29T09:14:07Z"}

That single call is enough for the alerts, the history and the API. A signed report takes one more step: we agree the test cases with you and lock them before the data starts arriving.

The promise

An alarm you are allowed to believe.

Suppose your quality never actually drops below the line you set. Then the chance this thing ever cries wolf is at most 1 in 20 — however often you look at it, for as long as you keep the setup unchanged. Most monitoring charges you for every look: check it often enough and the false alarms train you to ignore it. This one has already paid for the looking. When it fires, act.

An eval suite

gives you a score for today.

A drift score

tells you something has moved, without saying whether it matters.

An Actuary alarm

tells you how strong the evidence is, every single time you look.

1 in 20

The most false alarms you should ever see, however often you look, for as long as the setup stays frozen

This assumes your scores don’t move together in clumps. We show you where that breaks ↓

Proof

Before asking you to trust it, we broke it on purpose.

We wrote down what we expected to happen, signed it, and only then ran the test: the same task twice over, one copy served normally and one quietly degraded. Here is what came back.

WHAT THE BREAK LOOKED LIKE STUDY 001 · WRITTEN DOWN IN ADVANCE · 24 JULY 2026
EVIDENCE THAT QUALITY DROPPED · LOG SCALE0.51520ALARM LEVEL · 20NOTHING LEARNED YETPROMISED: WITHIN 19 CHECKSALARM · CHECK 12 · EVIDENCE 29.96THE BROKEN COPYTHE HEALTHY COPY · ENDS AT 0.79CHECK 15101520EVIDENCE THAT QUALITY DROPPED · LOG SCALE120ALARM · CHECK 12EVIDENCE 29.96THE BROKEN COPYTHE HEALTHY COPY · ENDS AT 0.79CHECK 114
Healthy copy, final evidence0.79
Alarm raised at check12
We promised: alarm within19 checks
Broken copy, final evidence29.96
Run on 24 July 2026 against openai/gpt-4.1-nano through OpenRouter, with five maths questions locked in beforehand and a minimum score of 0.70. This is one run, not a success rate — one honest example of the alarm doing its job. The side-by-side comparison with an ordinary fixed-length test is on the methodology page.

On the healthy copy, across all twenty checks, the evidence never once rose above where it started. That is on the signed certificate too, so you don’t have to take our word for it.

ACTUARY · STUDY 001 · AGREED BEFORE THE RUN

Certificate of monitoring

This is the real file, just typeset.
Download it and check it yourself.
Run
real-validation-prospective-fixed-panel-2026-07-24
Subject
openai/gpt-4.1-nano · via OpenRouter
Design frozen
2026-07-24 · 5-item maths panel · floor 0.70 · α 0.05
Registered bound
alarm within 19 blocks of onset · pre-registered separately
Result
alarm at panel 12 · log e 3.3999562471954308
Period
2026-07-24 17:19:34 → 17:24:10 UTC
Payload hash
35aefb0e387dd7f1784c8499d4dff5f99d45d562b926a539a3b4ee7fc08a58c8
Signed
Ed25519 · key 079ff1a1fc426e9570f896537edf2fe5612768f5e5ff8e05b1cbd2b299ca1f25
Chain
previous payload hash null — genesis
Every field here is read straight out of certificate.json. The promise about how quickly the alarm had to fire was written down separately, before the run. Signing-key fingerprint · one cell per hex digit

FOR DESIGN PARTNERS

We can run the checks for you

We sit down with you, turn one task into a fixed set of test cases, agree the lowest score you will accept, and lock all of it before any data arrives. From then on, every alarm arrives with a certificate like the one beside this.

$ actuary verify certificate.json
PASS canonical form
PASS signature · Ed25519, payload hash valid
PASS hash chain · certificate is genesis

The checker needs no database, no account and no secret key. It redoes the whole calculation from the file itself.

The public record

We point the same monitors at public models, and publish what we see.

You can watch us use our own tool before you trust it with your own work. The alarms are there, and so are the long quiet stretches and the times we got it wrong. Nothing can be added afterwards, not even by us.

Open the Watchtower Public, permanent, and safe to quote. Nothing in it can be backdated.

We’ll send one confirmation email, then a short note each time a public alarm fires. Nothing else.

Where it fails

Here is where our own detector gets it wrong.

A detector is only useful when its failure modes are visible. This is a stress test we ran on ourselves, not the rate you should expect in service. We are showing it because you deserve to know the shape of the failure before you rely on the tool.

56.38% How often it cried wolf when we fed it scores that move in clumps of five · simulated · 10,000 runs · we aim for 5%

If the scores you send us rise and fall together in clumps — the same bad afternoon showing up five times over — this maths fires far too often. On scores that move independently it stays close to the 5% we promise. Until we ship a version that handles clumping, every certificate says so on its face. We have not found this number published by anyone building something comparable, which is exactly why you should ask them for it.

Start

Set one up today and you’ll have the history when you need it.

The evidence you will want in three months only exists if something was watching today. Two monitors are free forever, and you won’t be doing the setup alone — we do the first one with you.

Watch is free, Team is $249 a month for the whole workspace, Registered is priced per engagement.