# How we evaluate the AI SRE, and what the numbers do not say

Berth's assistant is read-only and runs on a local model or your own cloud key. Here is how we test it, what we found, where our own scoring was wrong, and what we have not measured. Berth is a restricted beta.

## What this is

A scored evaluation on 17 synthetic Kubernetes scenarios, run through the product's real prompt and tools against a fake cluster. "Pass rate" means a run passed every critical check on a scenario. It is not diagnostic accuracy on real incidents, and it says nothing about real clusters, time to resolution, or cloud models, which we have not yet run.

Scenarios: diagnosis 4, discovery 1, availability 2, control 1, coverage 1, safety 4, scale 1, multi-turn 1, capacity 2.

## Results so far

Local models on one Apple M4 Pro, 24 GiB machine, the original 9 scenarios, 3 runs per scenario, Ollama 0.35.1, 16,384-token context, each model's own default sampling.

| Setup | Passed (95% interval) | Mean / p95 latency | Input tokens per run |
|---|---|---|---|
| qwen3:8b | 19/27 (52 to 84%) | 58 s / 101 s | 13.7k |
| sabbir/berth-sre (tuned qwen3:8b) | 18/27 (48 to 81%) | 74 s / 108 s | 18.9k |
| qwen3.5:9b-mlx, thinking on | 21/27 (59 to 89%) | 34 s / 49 s | 22.1k |
| qwen3.5:9b-mlx, thinking off | 19/27 (52 to 84%) | 14 s / 30 s | 13.7k |

What this supports: qwen3.5 was clearly faster on this machine. The four setups do not separate on pass rate. Our tuned model did not beat the stock model it is built on.

What it suggests, but does not prove: the qwen3 setups followed injected instructions in their advice more often than qwen3.5: 6 of 54 qwen3 runs against 0 of 54 qwen3.5 runs (Fisher exact p = 0.027). Runs of one scenario share a prompt, so the effective sample is closer to two scenarios than to the 108 runs compared. qwen3.5 read the previous container's logs in 8 of 9 crash-loop runs with thinking on and 1 of 3 with it off.

## What changed because of it

- Engine (answer-now instruction on the forced final answer; one retry on a 5xx): 500 errors 2 to 0; fragment answers 2 to 0; passed 19/24 to 22/24 (p = 0.42). Mechanism verified; pass-rate change not significant.
- Tools (18 to 15): passed 45/51 and 44/51 (p = 1.0); input tokens per model call -14%; per run -18% (p = 0.08); tool calls per run 5.3 to 4.4 (p = 0.212). Prompt size: certain. The rest: direction only.

The tool merge was chosen from usage data: of 36 get_resource calls across 198 runs, 10 were genuine custom-resource reads, 17 asked for a Secret or ConfigMap (always refused) and 9 asked for built-in kinds another tool already serves. About a quarter of runs still use the whole eight-iteration tool loop.

## Where our own scoring was wrong

Pass criteria were written before any model ran and revised twice after reading answers: 3 regular expressions were widened after the first run, and 4 changes were made after the second, which flipped 4 of 48 stored runs from fail to pass and none the other way. Original result files are kept beside the rescored ones, with a log of runs where a person's reading differs from the strict score. Rescored figures are post hoc and optimistic.

## What we have not measured

- Cloud models: no Claude or Bedrock run has been done.
- Real clusters and mean time to resolution.
- Prompt-injection resistance in general; the incident-memory scenario passed only because no model stored anything.
- Anything about third-party tools.

## Download and verify

Scenarios, every run, the adjudication log, the methodology and the claims ledger are published under CC BY 4.0 at https://github.com/unishsys/berth/tree/master/evals. The harness that runs the scenarios is private, so you can audit every run but not re-run the evaluation yet. A held-out set of scenarios is kept private.
