Skip to content
Saharsh Engineering Log
Back to ArraySignal

Technical notes

ArraySignal

ArraySignal is a tested and deployed application. These notes also document models, architectural ideas, and analytical approaches I explored while developing it, so some sections describe design investigations or directions I want to pursue rather than features of the running product.

01

Software and data architecture

The analytical logic stays separate from vendor APIs, billing, database operations and interface code. API routes call application services, services call analysis algorithms and repositories, and external providers sit behind adapters at the edge.

  1. Customer browser

    The operator's view of a site and its analyses.

  2. Frontend

    Presentation only. It decides what to show, never who is allowed to see it.

  3. Authenticated API

    Where the authorization decision is actually made.

  4. Application services

    Orchestration: what to run, in what order, for which organization.

  5. Analysis, reporting and background jobs

    The domain algorithms — expected production, persistence, device comparison, classification.

  6. Relational database

    Queries are scoped by organization, not filtered after the fact.

Vendor integrations

Adapters that normalize telemetry into the internal model.

Billing

Subscription state, kept outside the analytical path.

Object storage

Generated reports and uploaded files.

02

One canonical data model

I wanted tenant separation to exist in the backend and data model, not just in what the interface chooses to display. Each level below is owned by the one above it, and every query is scoped by that ownership.

The same principle drove the decision to avoid a separate analysis engine per telemetry source. Whether measurements arrive from CSV uploads or from a vendor platform, they are normalized into one internal measurement structure, so the analysis engine works against the ArraySignal model rather than knowing which vendor produced a reading. That makes integrations adapters rather than separate analytical systems.

  1. 01

    User

    Authenticates. Belongs to one or more organizations.

  2. 02

    Organization membership

    The grant that makes a user able to see anything at all.

  3. 03

    Organization

    The tenant boundary. Everything below belongs to exactly one.

  4. 04

    Sites

    A physical installation with capacity, location and configuration.

  5. 05

    Measurements, devices, analyses, reports

    Everything the pipeline reads and writes, scoped to the site above it.

03

The physics the model depends on

Technical note

Expected production depends on physical conditions, so the model considers plane-of-array irradiance, ambient and module temperature, rated DC capacity, the module temperature coefficient, inverter efficiency, and AC limits.

A simplified PVWatts-style relationship estimates DC power as rated power scaled by the ratio of plane-of-array irradiance to a 1000 W/m² reference, corrected for the difference between cell temperature and a 25 °C reference through the module temperature coefficient. AC power follows through inverter conversion and any AC limit.

Two consequences matter operationally. Modules generally produce less power as cell temperature rises, so temperature has to enter the expectation rather than being treated as noise. And a high DC yield can meet a flat AC ceiling when the inverter reaches rated capacity — that ceiling is clipping, clipping is not a fault, and a model that does not represent it will report one.

04

Time-series methodology

Solar telemetry arrives as measurements over time, which makes ordering matter. The analysis has to distinguish one unusually low measurement from several, from a persistent trend, from a repeating daily pattern, from missing telemetry, from an abrupt change, and from gradual degradation.

Each interval is classified against an underperformance threshold. A rolling window then asks how many recent intervals were abnormal, whether they were consecutive, whether the condition persisted during valid production periods, and whether enough valid data existed to trust the answer. A rolling buffer makes this affordable on long series, because each new observation updates the window without reprocessing the whole history.

05

Data quality

Data-quality checks run before any diagnosis. The gate considers missing measurements, invalid timestamps, duplicate records, impossible values, stale telemetry, missing expected-production inputs, and insufficient history.

The governing principle is that a missing telemetry stream and a healthy stream reporting zero power look similar on a chart and mean entirely different things. Equipment is not diagnosed before the data is checked, and a site that fails the gate produces a data finding rather than a fault.

06

Device-level mathematics

Raw device power cannot be compared across devices of different ratings, so output is capacity-normalized: measured power divided by rated power. A 50 kW inverter producing 45 kW and a 100 kW inverter producing 90 kW both normalize to 0.90.

The peer baseline uses the median of comparable devices rather than the mean. Given peer values of 1.00, 0.98, 1.01 and 0.41, the mean is dragged toward the failing unit while the median stays representative of the healthy group. Relative device performance and an estimated per-device energy impact follow from that baseline.

07

Fault-classification design

Classification is structured as an evidence process rather than a lookup. Candidate families include telemetry or communication problems, device zero-output behaviour, persistent device underperformance, intermittent operation, shading-like and soiling-like patterns, clipping-like behaviour, possible curtailment, site-wide underperformance, and long-term degradation.

Telemetry may support a likely explanation without proving that a physical component has failed, and the wording of every output is chosen to preserve that distinction.

  1. 01

    Detected condition

    A persistent, validated deviation — not a single reading.

  2. 02

    Candidate explanations

    The fault families consistent with the shape of the evidence.

  3. 03

    Supporting evidence

    What in the data is consistent with each candidate.

  4. 04

    Contradicting evidence

    What argues against it. Counted as seriously as support.

  5. 05

    Minimum evidence checks

    Whether enough valid data exists to justify any conclusion at all.

  6. 06

    Classification or abstention

    A named condition, or CAUSE UNCERTAIN when the evidence does not support one.

08

Severity and confidence

Severity asks how operationally important a condition is, which is largely a question of energy impact. Confidence asks how strong the evidence is, which is largely a question of data quality and persistence. They are computed and reported separately, and never collapsed into one score.

High possible impact / low confidence
Moderate impact / high confidence
Two values, never combined
09

Machine learning: what I explored

I explored where machine learning might help classify fault patterns, and deliberately did not design it to replace the rule-based system. The architecture supports physics, deterministic rules and ML together rather than ML instead of the others.

A candidate deployment strategy is shadow mode: the rule-based result is what the user sees, while an ML result is recorded privately for comparison. That would let a classifier be evaluated against real cases before it influences anything a customer reads.

The part I find most interesting is avoiding data leakage during evaluation. Measurements from the same underlying failure event should not appear in both training and test sets simply because they occurred at different timestamps. The independent unit is the event or the site case, not the telemetry row, which makes group-aware evaluation and probability calibration part of the method rather than refinements to it.

This is a direction I investigated, not something the running application does.

10

Long-term degradation: what I explored

A second direction I investigated works on a different time scale entirely: is a system gradually producing less over several years?

That is a different problem from detecting a short-term fault. Long-term analysis has to separate gradual decline from outages, equipment replacements, capacity changes, meter changes, configuration changes and sensor drift. A sudden step downward should not be read as an annual degradation trend.

I explored year-over-year approaches alongside ordinary least squares as a transparent cross-check, reported with a confidence interval, and treated change-point awareness as part of the method rather than an afterthought. What this taught me is that a trend estimate without uncertainty and change-point awareness can sound far more certain than the underlying data justifies.

11

Validation methodology

This is the least complete part of the project and the one I care most about.

The approach I want to use is to test the models against independent real operating cases and record, case by case, where they are reliable, where they fail, and where the software should abstain rather than guess. Operator feedback feeds that record, because the operator is the only party who finds out what the fault actually was.

No systematic validation has been completed. No accuracy, precision, recall, detection-rate or false-positive figures exist for ArraySignal, and any such figure appearing anywhere should be treated as an error.

12

Limitations and engineering risks

Expected-production modeling depends on site configuration data that may be wrong or missing, and a wrong expectation produces confident false findings. Irradiance data quality bounds everything downstream of it.

Clipping, curtailment and maintenance can all imitate faults. Device comparison needs enough comparable peers for the baseline to mean anything, and a two-inverter site barely has a peer group. Degradation estimates over short histories are weak regardless of estimator.

Rule thresholds are engineering judgements that have not yet been calibrated against real outcomes, and the machine-learning work is an exploration that sits outside any customer-facing path.