jùwọ̀n
How do you measure the error rate of a change detector without labelling anything? Point it at a desert.
What it does
Give it a box on the map and two dates. It pulls free satellite imagery, says what is in the scene and what changed between the two passes, and hands back the objects it found with real coordinates attached. Two sources, both free and neither needing so much as a registration: Sentinel-2 optical at 10 metres a pixel, and Sentinel-1 radar, which sees through cloud and at night because it brings its own illumination.
The interesting problem is not the detecting. It is that a system like this is trivially easy to make look good. Every parameter that raises the count of things found also raises the count of things imagined, there is no labelled ground truth for an arbitrary square kilometre of the earth, and nobody checks. A change detector that reports a great deal of change is indistinguishable, from the outside, from one that works.
The null test
So measure it from the other end. Rather than labelling what the engine found, choose a place where the answer is known before you start. An arid stretch of desert outside Aswan does not change between January and March. Take two cloud-free passes over it, run the whole pipeline, and whatever fraction of the area comes back marked as changed is, by construction, the false-positive rate. No annotation, no dataset, no judgement call.
The second site is the hard one. Lagos harbour is a turbid tropical estuary, which is the worst case for optical change detection: sediment plumes, wakes, tides, and a water surface that genuinely looks different every single pass without anything having happened that a person would call an event.
Those first numbers are the honest starting point, not a strawman. Seven per cent of a desert and thirty-nine per cent of a harbour, reported as change, on a pipeline that looked entirely reasonable until it was asked a question with a known answer.
What the test found
The value of the null test is not the final number. It is that a measurement with a known answer turns vague dissatisfaction into four specific bugs, and the first one is the one worth the whole exercise.
Otsu's method finds the threshold that best separates a histogram into two classes. It is the standard choice, and it is a relative criterion: it returns the best available split of whatever it is handed. Hand it a scene where nothing happened, and the histogram it is given is pure sensor noise, and it will find the best split of the noise and report it with total confidence. A relative threshold can never produce a null result, because a best split always exists.
The fourth bug is the one that would have quietly invalidated every figure the system ever printed. resolution_m reported the sensor's native 10 metres, while the provider was resampling every area of interest into a fixed 256 pixel tile. A six kilometre box was really about 26 metres a pixel, so every area in square metres downstream of it, every size filter, every “large enough to be worth reporting” judgement, was computed against a resolution the system did not actually have. It now computes the true ground sample from the extent, and adapts the tile between 256 and 1024 pixels to hold 10 metres rather than silently destroying resolution on bigger areas.
| What was wrong | Effect | Measured |
|---|---|---|
| Otsu had no absolute floor | Thresholded pure sensor noise as change | 11.21% → 0.67% |
| Normalisation fit over invalid pixels | Helped the desert, actively hurt the harbour | 2% → 14.5% |
| Region size expressed in pixels | The same setting meant 800 m² or 0.7 m² | now 2,500 m² |
| resolution_m was a lie | Every square-metre figure downstream was wrong | 10 m → really 26 m |
The second one is worth a sentence on its own, because it is the reason no fixed rule was adopted. Radiometric normalisation, fitting the second pass onto the first, is supposed to remove illumination differences. It improved the desert from 20% to 7% and made the harbour worse, from 2% to 14.5%, because the dominant water surface dragged the coefficients and mis-scaled the land. So the engine does not choose. It computes both, and keeps whichever leaves the lower median change magnitude, on the reasoning that real change is sparse and the median is therefore a clean read on the background.
Every finding above is pinned by a test, so none of it can quietly regress.
The other direction
Driving false positives to zero is trivial if you are willing to report nothing at all, so the null test on its own proves very little. The same harness therefore runs in reverse: paint a structure of known real-world size into the genuine after-scene, and ask whether it survives the thresholds.
below the reporting floor
900 m² is under the 2,500 m² minimum region, so it is never reported even if the pixels moved
This is where the free imagery stops being free. Ten metres a pixel is a physical limit, and below roughly sixty metres the engine is not being conservative, it simply cannot see. Planet at about three metres or Maxar at thirty centimetres is the answer, and neither is free. What a system can do in the meantime is report its actual resolution honestly, so that a caller knows in advance what is not there to be found.
Radar is a different instrument
Sentinel-1 is the reason the thing works under cloud, and almost nothing learned about the optical lane transfers to it. Radar is speckled by nature, an interference pattern rather than noise, so the first honest run over Lagos returned 921 change regions covering 16% of the scene, all of it texture. Multi-looking with a boxcar filter and raising the magnitude floor took that to 239 regions and 3.4%.
The better lesson came from the detector. A brightness threshold computed over the whole scene found exactly zero vessels, because land backscatter is so much stronger than water that it dominates the statistics and sets a bar no ship can clear. Taking the same statistics from the water class alone found 42, of which 21 were large enough to be worth a human's time. The fix was not a better threshold. It was noticing that the threshold was being computed over two populations at once.
Which then needed a water mask on radar, where the usual optical trick is unavailable. An early version took the darkest 45% of the scene, which a synthetic test caught calling 94% of a flat image water. The replacement finds the split properly and then checks that the histogram was actually bimodal before trusting it. When it was not, it returns nothing rather than a guess.
Most of it is refusing to answer
The pattern above repeats often enough to be the design. A measurement system is more useful when it declines than when it guesses, and the refusals are the part that took the longest to get right.
It refuses a window that is too cloudy. An early Lagos run looked clean and was not: both passes had roughly a quarter and a third of the area under cloud or shadow, and six of the 29 detections were sitting on the edge of the cloud mask. The tiles had each passed a conventional scene-wide cloud filter, because a tile can be mostly clear while the specific box you asked about is not. It now screens on the area of interest itself, samples several candidate passes at each end of the window, and declines outright rather than answering over cloud.
It refuses to ask a person about something they cannot judge. Objects below about 1,500 square metres are withheld from the review queue and counted, so a clear pair yielding 72 detections presents 26 and says plainly that 46 were too small to call. And it refuses to save a model that has not earned it: training runs stratified cross-validation and will not write out a classifier that fails to beat simply guessing the majority class.
Even the natural-language command bar refuses, in its way. It is a few hundred lines of rules rather than a language model, for two stated reasons: a hosted model would break the promise that the whole thing runs without an API key, and a parser that can show you exactly how it read your sentence is worth more here than one that reads it better. It shows its interpretation and waits to be confirmed, rather than acting.
The detector, honestly
Everything above is about change detection, which is measured. The object detector is not, and the gap between those two statements is the most important thing on this page.
There are two. The first is a brightness baseline: find compact bright blobs against a local statistic, filter by size and shape, refuse to fire near masked pixels. It does not know what a ship is. The second is a proposal-and-rescore arrangement, where the baseline runs loose for recall and a small logistic regression over twelve hand-designed features re-scores each candidate. On a held-out synthetic scene of fourteen vessels on water and fourteen rooftops of identical brightness on land, that takes precision from 0.54 to 1.00 at unchanged recall.
That 1.00 means nothing about the world, and the repository says so in bold before anyone else can. The synthetic classes are separable by construction. The shipped model was cold-started on 909 synthetic proposals and carries a warning field in its own weights file saying it has never been validated on a real human label. The review queue, the correction store and the retraining path are all built and tested, and the count of real labels collected so far is zero. What has been demonstrated is that the machinery works end to end, not that the model is good.
What it cannot do
Two sites are not a validation set. They are anchors chosen because their answers were known, and a fair test would add a site with real construction and real ground truth rather than more places where nothing happens. The 1.99% residual at Lagos is also not purely error: some of it is genuine change across twenty days of a working harbour that cannot be separated from nuisance change without exactly the learned model that does not yet have labels.
It runs locally, single user, with no accounts and no deployment. And the honest summary of the whole system is that the part which is measured is good, the part which is measured is not the part people ask about, and saying so plainly is cheaper than being found out later.
Figures on this page come from the calibration harness and the generated showcase runs in the repository. Imagery is Sentinel-2 L2A and Sentinel-1 RTC, both free and public. The engine is NumPy, with no GPU and no hosted model anywhere in it.