ArraySignal
Solar photovoltaic performance intelligence: a deployed web application that turns production telemetry into explainable operational evidence, comparing measured against expected output and investigating persistent underperformance.
Summary
- Software & Applications
- Flagship
- Continuing
- Jan 20, 2026
Evidence on file
My contribution
I designed and built ArraySignal myself: the analysis pipeline, the expected-production model, the persistence and device-level logic, the fault-classification design, the multi-tenant data model, and the service architecture that keeps the analytical code separate from vendor APIs, billing, and the interface.
The decisions I consider mine rather than assembled from references are the ones about restraint. Separating severity from confidence, letting the system return "cause uncertain" instead of forcing a diagnosis, requiring a minimum amount of evidence before classifying anything, and checking the data before blaming the equipment are all choices about what the software should refuse to say.
Project overview
Why I built this
Most solar monitoring systems are very good at showing measurements: power, energy, alarms, inverter status, historical charts. I became interested in a different problem — whether software could go past displaying data and begin reasoning about it.
The question that shaped everything came last in that chain: how confident should the software be before it tells someone what to investigate? Answering it honestly turned a monitoring idea into a much larger piece of engineering involving solar physics, time-series analysis, fault detection, architecture, databases and statistics.
What I built
I built and deployed ArraySignal, and it runs at arraysignal.com. It is a live web application rather than a design exercise, and I am still developing it.
It is built around the analysis pipeline set out below, which runs from imported telemetry through validation and an expected-performance model to persistence checks, device-level comparison, a classified finding, and an estimate of the energy involved.
The application being deployed does not mean every part of that pipeline is finished. Some of it runs today, some is still being worked on, and some is a direction I investigated rather than shipped. The last two sections of this page and the technical notes are where the exploratory work lives, kept separate from what the running product does.
My work spans the expected-production model, the time-series reasoning that decides whether a deviation is meaningful, the data-quality checks that run before any diagnosis, the device-level comparison that turns a site-level observation into device-level evidence, and the classification layer that decides what the evidence actually supports.
Problem and goal
Is this solar system performing normally? If not, where is the loss coming from, and what should someone investigate first?
A site can lose production without raising any obvious alarm — partial inverter underperformance, a dead telemetry stream, soiling, recurring shading, clipping, curtailment, an intermittent fault, slow degradation. What makes the question hard is that weather, sensor errors and maintenance move the same number as real faults do, so a drop in production is not by itself evidence that anything has failed.
The goal is to turn raw telemetry into evidence good enough to support an operational decision.
Design approach
Two decisions shaped everything else.
The first is that tenant separation lives in the backend and the data model rather than in what the interface chooses to display. Users belong to organizations, organizations own sites, and every query is scoped by that ownership.
The second is a single canonical measurement model. Whatever a reading arrives as, it is normalized at the boundary, so the analysis engine works against the ArraySignal model rather than knowing which source produced it. That makes integrations adapters rather than separate analytical systems.
Both are written up properly in the technical notes, along with the layering that keeps the analysis independent of vendor APIs, billing and interface code.
Inside ArraySignal
- 01Guided workflow. A self-guided reference workflow moving from portfolio triage to performance analysis, diagnosis, modeled impact, and next action. The product labels this view sample data, not customer telemetry, across five fictional sites.
- 02Design and telemetry imports. Vendor-neutral ingestion keeps physical system configuration separate from operational telemetry so imported design assumptions never become fabricated measurements. Availability labels are the product's own.
- 03Modeled opportunity. An interactive planning calculation showing how transparent assumptions affect modeled portfolio-level energy and financial impact. It is not a forecast or guaranteed savings calculation, and the product labels its output illustrative model output calculated from the inputs shown.
Testing and validation
Validation is the part of this project I consider unfinished, and I would rather say so than imply otherwise.
The application runs and the analysis produces results. What has not happened is a systematic test of that reasoning against real operating data across many sites — establishing where the models hold, where they break, and where the software should abstain rather than guess.
That is my next priority, ahead of any new analytical feature. The methodology I want to use is described in the technical notes.
Failures and debugging
Every one of the decisions above started as something I got wrong, and the pattern behind all of them was the same: jumping from a number to a conclusion.
The first version reacted to single low readings and produced a stream of findings that meant nothing. The second compared inverters by raw output and flagged a healthy 50 kW unit sitting beside a healthy 100 kW one. The third confidently blamed hardware for what turned out to be a dead telemetry feed — the chart for no data and the chart for zero power look almost identical.
Each mistake was cheap to make and expensive to trust, which is why the analysis now spends most of its effort deciding whether a difference means anything before deciding what it means.
What I learned
ArraySignal became much larger than the application I originally imagined. The most important lesson has been that detecting a numerical difference is easy compared with explaining what that difference means.
A useful performance-analysis system has to ask whether a result is caused by bad data or real behaviour, whether it is isolated or persistent, site-wide or device-specific, normal operating behaviour or abnormal, operationally important or negligible, and sufficiently supported to communicate confidently.
That changed the wording as much as the code. Instead of "the inverter has failed", the system reports that a device has stayed below comparable devices across multiple valid intervals, and names what to review. The goal is not the strongest-sounding diagnosis. It is the strongest conclusion the evidence actually supports.
Next iteration
The next step is not another analytical feature. It is a validation framework built on real solar operating data, and an honest record of where ArraySignal succeeds, where it fails, and where it should abstain.
After that: confidence calibration against real outcomes, device-level attribution on larger sites, and the two directions in the technical notes — degradation analysis and ML-assisted classification — if validation gives me reason to trust them.
Analysis pipeline
Every measurement travels the same path. Each stage can stop the analysis: if the data does not survive validation, nothing downstream is allowed to blame the equipment.
- 01
Solar measurements
Imported telemetry, normalized into one internal structure regardless of source.
- 02
Data validation and normalization
Missing, duplicate, stale, impossible and out-of-order readings are caught here.
- 03
Expected performance model
A physics-informed estimate of what the system should have produced under the conditions.
- 04
Actual vs. expected comparison
The difference, expressed as a performance index rather than a raw shortfall.
- 05
Persistence and anomaly detection
Rolling windows decide whether a deviation has lasted long enough to be worth reporting.
- 06
Device-level contribution analysis
Whether one device accounts for much of the measurable difference.
- 07
Explainable fault classification
Which explanations the evidence supports, and whether it supports any.
- 08
Energy-impact estimation
How much energy the condition may be costing.
- 09
Operational output
Site analysis, reports and a recommendation about what to investigate first.
Engineering decisions
Four decisions did most of the work, and each came from getting something wrong first.
**Check the data before blaming the equipment.** A missing telemetry stream and a healthy stream reporting zero power look almost identical on a chart and mean completely different things. Validation for missing, duplicate, stale, impossible and out-of-order readings runs before anything is attributed to hardware, and a site that fails those checks produces a data finding rather than a fault.
**Reason over time, not over one reading.** Measurements are noisy, clouds move, sensors fail, and systems transition between operating states. Reacting to a single low interval produced noise, so the analysis uses rolling windows and persistence rules. That turns "production is low right now" into "production has remained materially below expectation across multiple valid intervals", which is a statement worth acting on.
**Normalize before comparing devices.** A 100 kW inverter at 90 kW and a 50 kW inverter at 45 kW are both at 90 percent of rating, so raw output cannot be compared across a mixed site. Output is normalized by rated capacity, and the peer baseline uses the median rather than the mean, because one badly performing unit drags a mean toward itself while the median stays with the healthy group.
**Let the system decline to answer.** Telemetry can show that a device sits below its peers without proving why. Classification weighs supporting evidence against contradicting evidence, checks whether enough valid data exists to justify any conclusion, and returns "cause uncertain" when it does not. Severity and confidence are reported as two separate values, so a large estimated loss measured through poor telemetry cannot masquerade as a certain one.
Actual against expected
The application does not check whether today beat yesterday. It compares measured output against what the physical conditions say the system should have produced, using irradiance, module temperature, rated capacity, the temperature coefficient, inverter efficiency and AC limits.
Everything else in the analysis exists to decide whether a departure from 1 is real, whether it has lasted, and what it can be attributed to. The derivation is in the technical notes.
Performance Index = Actual Energy / Expected Energy
- Actual Energy
- Measured production over the interval, after data-quality checks.
- Expected Energy
- Production the physics-informed model predicts for the conditions in that interval.
- ≈ 1
- Production is close to expectation.
- < 1
- The system produced less than the model predicts. The cause is a separate question.
Worked example
This simplified example illustrates the analysis logic. It is not presented as a real customer result, and the numbers are not drawn from an operating site.
One interval like this is not necessarily meaningful. If the same condition persists across multiple valid production intervals, it becomes much more interesting — which is the reason persistence analysis exists.
- 71.4 kW
- 58 kW
- 58 / 71.4 = 0.812
- 18.8%
- 3.35 kWh
- None. A single interval is not evidence.
Technical notes
The architecture, the data model, the physics, the mathematics behind the analysis, and the directions I explored while building it.
Read the technical notesEvidence
Log entries for this project
- One low reading is not a faultThe first version reported every interval that fell below expectation, and almost none of them meant anything. Fixing it meant reasoning over time rather than over a number.Time-SeriesAnomaly Detection
- Teaching the system to say it does not knowThe most useful thing the classifier does is decline to answer. Getting there meant separating how bad a condition is from how sure the evidence makes me.Fault ClassificationUncertainty
Related work
DwellMetry is a live home-information application I built to organize a home’s structure, equipment, documents, maintenance history, and utility records in one place. I also used it as a research environment for exploring thermal behavior, expected-performance models, telemetry, and anomaly reasoning, but those scientific features are not presented here as generally available or field-validated product capabilities.