Problem
The Supervisor presents current measurements and time series, but it does not identify material performance regressions against recent comparable windows.
Required insight
Add rolling-baseline insights for the existing measured metrics. Each finding must include:
- the current and baseline windows;
- sample counts;
- absolute and relative deviation;
- the metric and workload dimensions used for comparison;
- representative request IDs where available;
- unavailable when there is not enough comparable history.
Avoid comparing unlike traffic. At minimum, keep model identity and request phase/workload dimensions explicit. Recommendations are inferred; the window statistics are measured.
Acceptance
Fixtures show a stable series, a material regression, incomparable workload windows, and insufficient history. Only the material comparable regression produces a warning.
Split from #12.
Problem
The Supervisor presents current measurements and time series, but it does not identify material performance regressions against recent comparable windows.
Required insight
Add rolling-baseline insights for the existing measured metrics. Each finding must include:
Avoid comparing unlike traffic. At minimum, keep model identity and request phase/workload dimensions explicit. Recommendations are inferred; the window statistics are measured.
Acceptance
Fixtures show a stable series, a material regression, incomparable workload windows, and insufficient history. Only the material comparable regression produces a warning.
Split from #12.