Model Change Control: CI for Model Upgrades

aissurance now ships a model change control kit that measures exactly what a model or serving-stack upgrade changed — and files the result as Annex IV change-control evidence.

Why This Landed

Every deployer with an EU AI Act documentation duty eventually gets asked the same question: how do you know the update did not break anything? You swapped the model, moved to a new inference build, changed quantization or batching — and your Annex IV documentation still describes the system as it was before.

Model behaviour is not frozen at deployment. A serving-stack upgrade alone can change output: in our reference series, the same endpoint was bit-exact 45/45 under sequential dispatch, and produced nine distinct outputs for the same request under concurrent serving — with nothing in the logs flagging the difference. “We upgraded the runtime, the model name is the same” is not a change-control story an auditor can work with.

What the Kit Does

Model change control is now built into the aissurance compliance platform. It works like CI for model and serving-stack upgrades:

  1. Record a baseline before the upgrade. The kit runs a fixed calibration set against your inference endpoint under a fully pinned sampler signature and archives the measurement.
  2. Upgrade. Swap the model, bump the server, change the serving configuration — whatever the change is.
  3. Re-run and compare. Each metric is checked against its gate. A gate that fails fails the run with a non-zero exit code, so it drops straight into your existing CI pipeline.
  4. File the report as evidence. The run files under Annex IV element 12, “Change control documentation”, in its own evidence category so these reports stay findable when someone asks what changed and when.

It runs against any OpenAI-compatible endpoint, needs no database and no extra dependencies, and takes minutes to set up.

What It Measures

Not just accuracy. The report captures the properties an upgrade actually moves:

  • Determinism — is the endpoint reproducible under a fixed seed? A configuration property that quietly breaks under concurrency.
  • Accuracy — modal answers on standard and knowledge-boundary cases, split so you can see where the upgrade bites.
  • Surprisal and tail shape — the output distribution’s location and behaviour. In the reference series, mean surprisal moved 7% between two builds of the same server with no behavioural change.
  • Entropy-as-wrong-signal (AUROC) — if your escalation gate relies on entropy to route uncertain cases to humans, this checks the gate still works. A degraded AUROC means the gate has quietly stopped doing its job.
  • Escalation rate — what the gate costs in human review. Drift here is a budget change, in either direction.

Why the Strictness Matters

Two design decisions carry most of the value:

The sampler signature has no defaults. Every sampling parameter must be stated explicitly. A field left open in a request doesn’t stay unset — the server substitutes its own default, and you end up measuring a different instrument. In the reference series, one run with top_p left open silently took the server default of ~0.95 and scored a tail slope of 0.73 instead of 1.1. Nothing in the output looked wrong. The signature fingerprint is stored in every report, and comparisons refuse to run across differing signatures.

The instruments validate themselves. Every estimator has to recover a known answer from synthetic data before it is trusted on real data (selftest, no endpoint needed). This is the difference between a measurement and an artefact.

Gate defaults ship as working opinions — the kit’s documentation tells you exactly where each number came from and urges you to tighten them against your own run history. A gate that has never fired has never been calibrated.

What Gets Filed

The filing lands under Annex IV element 12 with evidence_category = "model_change_control". What is filed is a summary carrying the full report’s SHA-256 digest at the top — that digest binds the human-readable statement to the archived JSON, which a reviewer can recompute and verify months later.

The summary states the negatives explicitly: unmeasured metrics appear as “not measured” rather than being omitted, and a failed determinism battery appears as a caveat over every figure below it. A reader months later can see what was not shown.

The Honest Scope

This kit detects change against a baseline. It does not certify a model as correct, and it does not catch modal confabulation — a wrong answer the model gives consistently produces zero entropy signal in every self-consistency instrument. The right answer there is an independent check (arithmetic, a database cross-reference, a source), which is why the built-in calibration set ships with arithmetic as its control group.

Why You Should Care When Upgrading Your Agents

If you operate agents on self-hosted or OpenAI-compatible infrastructure, every dependency bump is a potential behavioural change you currently have no way to see. With this kit:

  • The before/after comparison is automated and reproducible, so “we checked” becomes a documented procedure rather than a hope.
  • The evidence trail is built for the audit: baseline, comparison, gates, verdict, digest — filed where your Annex IV documentation lives.
  • The gates are one-sided where it matters: an accuracy improvement never fails a run. A change is only a problem when it makes things worse.

The practical workflow for teams: run selftest once, record a baseline before your next upgrade, and wire the compare step into your deployment pipeline. From then on, every model upgrade produces its own change-control documentation as a by-product — which is what Article 12-style change control was supposed to mean all along.

Coming next: the Nomyo Router compliance plugin will run this kit on a schedule and file reports without an operator in the loop, turning point-in-time comparisons into continuous post-market monitoring.