Research · 9 Oct 2026
Monitoring method v2: canaries, probes and reports
How the TOSWO monitor turns probe results, provider status pages and public reports into system statuses, events and the risk level.
The AI risk monitor is built so that nobody has to be trusted to set a number by hand. Every status on the monitor is computed from evidence and every rule is public.
Evidence
| Source | What it tells us | How it counts |
|---|---|---|
| Availability probe | Can the provider's endpoint be reached, and how fast | One failure is a blip (degraded); two in a row is an outage. Latency far above the probe's own baseline (more than three standard deviations) is degraded |
| Provider status page | What the provider itself says | Minor or major incidents are degraded, critical is an outage |
| Behavior canary | Whether a model still behaves as expected on fixed prompts | Each prompt has expect and forbid patterns. A failure rate of 30% or more, or far above its own baseline, marks the system as anomalous. Checks are grouped by category (deception, self-replication, safety bypass and so on) |
| Public report | What people observed | Weighted by the sender's track record, and checked against the probes |
Baselines that cannot normalise an incident
Each probe keeps an exponentially weighted baseline of its own normal. The baseline is not updated during an incident, so a long outage cannot become the new normal.
Reports
A report counts according to its sender's history: (confirmed + 1) divided by (confirmed + rejected + 2), so a new reporter starts at one half and moves with each judgement. A burst of reports far above the hourly baseline opens a watching event. A researcher then confirms, annotates or resolves it. Reports are stored with a salted hash instead of the sender's address.
Events
An event opens automatically when a problem persists, and resolves only after ten quiet minutes. Researchers can confirm or annotate. The worker never downgrades a decision a person has made. Every event has a permanent address.
The level
The level from 1 to 5 is a weighted sum over active events (the weight of the category, the state of the event and its confidence) plus a small term for affected systems. Levels are coloured green, yellow, orange, red and dark purple.
Two things can hold the level above the computed value, and both are shown publicly:
- A manual override needs an administrator, a written reason and an expiry of at most 72 hours.
- A standing level set by researchers has no expiry and a written reason. It is how reported incidents from outside our own probes, such as lab disclosures, are weighed by people rather than by a formula. It stays until a researcher changes it.
The historical record
The record goes back to January 2018. Before our probes existed it comes from public incident databases, never from simulated data, and each entry keeps its source and link. Lab disclosures and press reports are added with their sources by researchers.
What the monitor does not do
It does not read private data, it does not test models for harmful capabilities beyond the canaries described, and it cannot see inside a provider. Where nothing was measured, it shows nothing, never "normal".
The rules are in the open: the data and API documentation describes every field, and corrections go through the corrections policy.