Skip to content
Saharsh Engineering Log
Engineering log

Retrospective note

Written from project notes after the fact.

One low reading is not a fault

The first version reported every interval that fell below expectation, and almost none of them meant anything. Fixing it meant reasoning over time rather than over a number.

ArraySignalTime-SeriesAnomaly Detection

Problem

Once I could calculate an actual-versus-expected difference, a new problem appeared: a single low data point does not necessarily mean anything. Measurements are noisy, clouds move, sensors fail, communications drop, and systems transition between operating states. Reporting each one produced a stream of findings that an operator would learn to ignore.

Decision

I moved the decision from the interval to a window. Each interval is still classified against an underperformance threshold, but a rolling window then asks how many recent intervals were abnormal, whether they were consecutive, whether the condition persisted during valid production periods, and whether enough data was available to trust the answer at all. A rolling buffer updates the window with each new observation instead of reprocessing the history.

Test / evidence

The change is visible in what the system can say. Before, the strongest available statement was "production is low right now". After, it is "production has remained materially below expectation across multiple valid intervals" — a claim with duration attached, which is the part that makes it worth acting on.

Result

Persistence turned out to be one of the simplest ways to separate noise from a condition worth investigating, and it cost far less than any of the statistical approaches I had assumed I would need.

Next step

Work out whether a persistent shortfall is site-wide or concentrated in one device.