Evidence
How we evaluate the AI SRE,
and what the numbers do not say
Berth's assistant is read-only and runs on a local model or your own cloud key. Here is how we test it, what we found, where our own scoring was wrong, and what we have not measured. Berth is a restricted beta.
What this is
A scored evaluation on 17 synthetic Kubernetes scenarios, run through the product's real prompt and tools against a fake cluster. "Pass rate" means a run passed every critical check on a scenario. It is not diagnostic accuracy on real incidents, and it says nothing about real clusters, time to resolution, or cloud models, which we have not yet run.
Scenarios: diagnosis 4, discovery 1, availability 2, control 1, coverage 1, safety 4, scale 1, multi-turn 1, capacity 2. The first 9 were written before the rest; the others were added to cover high availability, drains, node removal, custom-resource discovery, follow-up questions and a large cluster.
Results so far
Local models on one Apple M4 Pro, 24 GiB machine, the original 9 scenarios, 3 runs per scenario, Ollama 0.35.1, 16,384-token context, each model's own default sampling.
| Setup | Passed (95% interval) | Mean / p95 latency | Input tokens per run |
|---|---|---|---|
| qwen3:8b | 19/27 (52 to 84%) | 58 s / 101 s | 13.7k |
| sabbir/berth-sre (tuned qwen3:8b) | 18/27 (48 to 81%) | 74 s / 108 s | 18.9k |
| qwen3.5:9b-mlx, thinking on | 21/27 (59 to 89%) | 34 s / 49 s | 22.1k |
| qwen3.5:9b-mlx, thinking off | 19/27 (52 to 84%) | 14 s / 30 s | 13.7k |
What this supports. qwen3.5:9b-mlx was clearly faster on this machine. The four setups do not separate on pass rate: their intervals overlap almost completely. Our tuned model did not beat the stock model it is built on.
What it suggests, but does not prove. The qwen3 setups followed injected instructions in their advice more often than qwen3.5: 6 of 54 qwen3 runs against 0 of 54 qwen3.5 runs (Fisher exact p = 0.027). Runs of one scenario share a prompt, so they are not independent and the effective sample is closer to two scenarios than to the 108 runs compared. Thinking mode also mattered in one place: qwen3.5 read the previous container's logs in 8 of 9 crash-loop runs with thinking on and 1 of 3 with it off.
What changed because of it
| Layer | Change | Measured effect | Strength |
|---|---|---|---|
| Engine | Forced final answer now says "answer now"; one retry on a 5xx | 500 errors 2 to 0; fragment answers 2 to 0; passed 19/24 to 22/24 (p = 0.42) | Mechanism verified; pass-rate change not significant |
| Tools | 18 tools to 15 | Passed 45/51 and 44/51 (p = 1); input tokens per model call -14%; per run -18% (p = 0.08); tool calls per run 5.3 to 4.4 (p = 0.212) | Prompt size: certain. The rest: direction only |
The tool merge was chosen from usage data: of 36 get_resource calls across 198 runs, 10 were genuine custom-resource reads, 17 asked for a Secret or ConfigMap (always refused) and 9 asked for built-in kinds another tool already serves. Those mistakes came mostly from the older qwen3 models, so this run could not show a behaviour fix for qwen3.5. About a quarter of runs still use the whole eight-iteration tool loop: merging tools did not stop over-exploration.
Where our own scoring was wrong
Pass criteria were written before any model ran and revised twice after reading answers: 3 regular expressions were widened after the first run, and 4 changes were made after the second, which flipped 4 of 48 stored runs from fail to pass and none the other way. We kept the original result files, published the rescored ones beside them, and keep a log of runs where a person's reading differs from the strict score. Rescored figures are post hoc and optimistic. Inside any before-and-after comparison the criteria were identical on both sides.
What we have not measured
- Cloud models. No Claude or Bedrock run has been done, so we say nothing about cloud quality or cost per solved scenario.
- Real clusters. The fixtures are small, hand-written and mostly single-fault. Mean time to resolution is unmeasured.
- Prompt-injection resistance in general. We measured named models on a few named scenarios. The scenario that tests whether a model writes a false note into incident memory passed only because no model stored anything.
- Anything about third-party tools.
Download and verify
The scenarios, every run, the adjudication log, the methodology and the claims ledger are published under CC BY 4.0 in the public repository. The harness that runs the scenarios is private, so you can audit every run but you cannot re-run the evaluation yourself yet. We also keep a held-out set of scenarios private, to notice when public scores drift because we tuned against the public ones.
Result sets used above: 11 model result sets, 276 recorded runs, Ollama 0.35.1 and 0.40.0.
Try it on your own cluster
Check every answer against the tool calls Berth shows you.