Retrospective note
Written from project notes after the fact.
Teaching the system to say it does not know
The most useful thing the classifier does is decline to answer. Getting there meant separating how bad a condition is from how sure the evidence makes me.
Problem
Detecting underperformance and diagnosing a fault are not the same problem. Telemetry can show that a device sits consistently below its peers without establishing why: a device problem, a DC-side issue, curtailment, maintenance, temperature behaviour, a sensor, or a communication fault all produce similar shapes. A system that always returns a named cause will be confidently wrong on a predictable fraction of cases.
Decision
I built classification around evidence rather than pattern matching: supporting evidence, contradicting evidence weighed as seriously as support, a minimum-evidence requirement, and a confidence threshold. When those are not met the result is "cause uncertain". I also split severity from confidence and refused to combine them, so the two questions — how much does this matter, and how sure are we — are answered separately.
Test / evidence
The two cases that justify the split are a large apparent loss measured through poor telemetry, which is high possible impact at low confidence, and a smaller loss measured cleanly, which is moderate impact at high confidence. Collapsing those into one score makes the first look like the second. The same reasoning changed the wording of every output: rather than "the inverter has failed", the system reports that a device has stayed below comparable devices across multiple valid intervals and names what to review.
Result
The goal is not to produce the strongest-sounding diagnosis. It is to produce the strongest conclusion the available evidence actually supports. That is the idea the rest of the project is organised around.