Operational now. October is below its 99.5% commitment.
Operational. The last probe found every endpoint up, inside its latency budget, and returning the pinned answer. Each probe asks the production data API fixed questions about a synthetic family, over the public internet like any client, and compares every answer’s hash with the one pinned for the engine in service. A wrong number counts as an incident, the same as an outage. Last probe: 2026-10-02 06:05 UTC.
How often has every check passed?
| Window | Probes passed | Probes | Availability |
|---|---|---|---|
| Last 24 hours | 275 | 288 | 95.486% |
| Last 7 days | 1,766 | 1,779 | 99.269% |
| Last 30 days (10.4 days measured) | 2,926 | 2,978 | 98.253% |
| Last 90 days (10.4 days measured) | 2,926 | 2,978 | 98.253% |
| Calendar month | Probes | Failed | Availability | Failures the commitment allows | Budget left | Credit the terms owe |
|---|---|---|---|---|---|---|
| 2026-10 (in progress) | 362 | 13 | 96.408% | 1 | 0 | not yet: the month is open |
| 2026-09 | 2,616 | 39 | 98.509% | 13 | 0 | 25% of that month's fee |
Committed to clients: 99.5% a calendar month, with credits below it. The credit column is computed from this record by the schedule in the order form, not negotiated afterwards: below 99.5% it is 10% of that month’s fee, below 99.0% it is 25%. The error budget is the number of probes a month may still fail before the committed figure is breached. Target: 99.9%. The committed figure sits below the target on purpose until this record is ninety days long. A window with no probes shows a dash, not a hundred. Figures are floored, never rounded up. The record began 2026-09-21 20:05 UTC; the first month counts from then. Every failure in it stays, and each incident’s cause is reviewed below.
How long does each answer take?
| Endpoint | Median | 95th percentile | Budget |
|---|---|---|---|
| Book as known | 83 ms | 121 ms | 2,500 ms |
| Concentration across entities | 404 ms | 579 ms | 4,000 ms |
| Exposures | 384 ms | 552 ms | 2,500 ms |
| Health | 149 ms | 336 ms | 2,000 ms |
| Capital-call simulation (2,000 paths) | 543 ms | 777 ms | 6,000 ms |
| Rebalancing endpoint stays off | 71 ms | 121 ms | 2,500 ms |
| Lot sequence | 337 ms | 758 ms | 4,000 ms |
What went wrong, and for how long?
| Started | Ended | About | Checks that failed |
|---|---|---|---|
| 2026-10-02 04:50 UTC | 2026-10-02 05:55 UTC | 65 min | Book as known: up; Concentration across entities: up; Exposures: up; Health: engine named; Health: isolation; Health: up; Capital-call simulation (2,000 paths): up; Rebalancing endpoint stays off: says why; Rebalancing endpoint stays off: up; Lot sequence: up |
The checks: up (the expected status), fast (inside the budget above), the same twice (asked again, the same hash), right (the hash pinned for this engine), honest (the answer still says it executes nothing), shut (401 without a token). Engine in service: e67542034c7ae8dc…. The same record as JSON, for your own monitor: /api/status. Try the questions yourself on the developers page; the controls behind the platform are on Trust.
Why each one happened, and what stops it now.
Cause. A new engine (12fea765…) went live before its pinned answers did, so every answer failed the check against the pin until the pin was deployed.
What changed. Both ways production is deployed, CI and the release script, now refuse an engine whose answers are not pinned.
Cause. A route change added a field to the lot-sequence answer on an unchanged engine, so its bytes no longer matched the pinned hash.
What changed. The answer's bytes were restored within the hour, and every change is now compared with the pinned answers before it is pushed.
A review explains a failure; it does not remove one. Every failed probe above still counts in its month’s figure.