jùwọ̀n

How do you measure the error rate of a change detector without labelling anything? Point it at a desert.

Sentinel-2 · SAR
NumPy
Python
Calibration

What it does

Give it a box on the map and two dates. It pulls free satellite imagery, says what is in the scene and what changed between the two passes, and hands back the objects it found with real coordinates attached. Two sources, both free and neither needing so much as a registration: Sentinel-2 optical at 10 metres a pixel, and Sentinel-1 radar, which sees through cloud and at night because it brings its own illumination.

The interesting problem is not the detecting. It is that a system like this is trivially easy to make look good. Every parameter that raises the count of things found also raises the count of things imagined, there is no labelled ground truth for an arbitrary square kilometre of the earth, and nobody checks. A change detector that reports a great deal of change is indistinguishable, from the outside, from one that works.

The null test

So measure it from the other end. Rather than labelling what the engine found, choose a place where the answer is known before you start. An arid stretch of desert outside Aswan does not change between January and March. Take two cloud-free passes over it, run the whole pipeline, and whatever fraction of the area comes back marked as changed is, by construction, the false-positive rate. No annotation, no dataset, no judgement call.

The second site is the hard one. Lagos harbour is a turbid tropical estuary, which is the worst case for optical change detection: sediment plumes, wakes, tides, and a water surface that genuinely looks different every single pass without anything having happened that a person would call an event.

Null test · truth ≈ 0%
Where nothing happened, calibration took reported change from 7% and 39% to under 2%
before calibration
after
Aswan desertarid, static ground
before
7%
after
0.23%
Lagos harbourturbid tropical estuary
before
38.9%
after
1.99%
0%shared scale · 40%
Two sites, one date pair each, January to March 2024. Nothing here is labelled and nothing is annotated by hand: the score is simply the fraction of the area the engine claimed had changed, at a place chosen because the answer is known in advance. Two sites is not a validation set, and the repository says so. They are anchors, not proof of generalisation.

Those first numbers are the honest starting point, not a strawman. Seven per cent of a desert and thirty-nine per cent of a harbour, reported as change, on a pipeline that looked entirely reasonable until it was asked a question with a known answer.

What the test found

The value of the null test is not the final number. It is that a measurement with a known answer turns vague dissatisfaction into four specific bugs, and the first one is the one worth the whole exercise.

Otsu's method finds the threshold that best separates a histogram into two classes. It is the standard choice, and it is a relative criterion: it returns the best available split of whatever it is handed. Hand it a scene where nothing happened, and the histogram it is given is pure sensor noise, and it will find the best split of the noise and report it with total confidence. A relative threshold can never produce a null result, because a best split always exists.

Bug one · relative threshold
Otsu finds a best split even in pure noise, so it always reports change
reported changed
noise below threshold
a scene where nothing happenedtruth: 0% changedotsu 0.0370change magnitude → 0.35
reported changed41.03%
Drag the floor. At zero, Otsu is asked for the best split of pure sensor noise and returns one, because a best split always exists, and the engine confidently reports a fifth of a static desert as changed. The fix is not a better threshold, it is an absolute one: max(otsu, min_magnitude). The distribution here is drawn rather than sampled and illustrates the mechanism; on the real Aswan scene the floor took the reported change from 11.21% to 0.67%, and 0.10 is the calibrated default.

The fourth bug is the one that would have quietly invalidated every figure the system ever printed. resolution_m reported the sensor's native 10 metres, while the provider was resampling every area of interest into a fixed 256 pixel tile. A six kilometre box was really about 26 metres a pixel, so every area in square metres downstream of it, every size filter, every “large enough to be worth reporting” judgement, was computed against a resolution the system did not actually have. It now computes the true ground sample from the extent, and adapts the tile between 256 and 1024 pixels to hold 10 metres rather than silently destroying resolution on bigger areas.

What was wrongEffectMeasured
Otsu had no absolute floorThresholded pure sensor noise as change11.21% → 0.67%
Normalisation fit over invalid pixelsHelped the desert, actively hurt the harbour2% → 14.5%
Region size expressed in pixelsThe same setting meant 800 m² or 0.7 m²now 2,500 m²
resolution_m was a lieEvery square-metre figure downstream was wrong10 m → really 26 m

The second one is worth a sentence on its own, because it is the reason no fixed rule was adopted. Radiometric normalisation, fitting the second pass onto the first, is supposed to remove illumination differences. It improved the desert from 20% to 7% and made the harbour worse, from 2% to 14.5%, because the dominant water surface dragged the coefficients and mis-scaled the land. So the engine does not choose. It computes both, and keeps whichever leaves the lower median change magnitude, on the reasoning that real change is sparse and the median is therefore a clean read on the background.

Every finding above is pinned by a test, so none of it can quietly regress.

The other direction

Driving false positives to zero is trivial if you are willing to report nothing at all, so the null test on its own proves very little. The same harness therefore runs in reverse: paint a structure of known real-world size into the genuine after-scene, and ask whether it survives the thresholds.

Injection test · 10 m ground sample
Below roughly 60 m, a structure is not there to be found
0200 m across
30 m3 px · 900 m²

below the reporting floor

900 m² is under the 2,500 m² minimum region, so it is never reported even if the pixels moved

A 30 m building is nine pixels and 900 square metres, under the reporting floor before anything has been detected at all. This is the honest ceiling of free imagery rather than a deficiency in the method: at 10 m a ground sample, small things are not there to be found, and the answer is a higher resolution source rather than a cleverer threshold. What the engine can do is report its true resolution, so a caller can tell in advance what is not resolvable.

This is where the free imagery stops being free. Ten metres a pixel is a physical limit, and below roughly sixty metres the engine is not being conservative, it simply cannot see. Planet at about three metres or Maxar at thirty centimetres is the answer, and neither is free. What a system can do in the meantime is report its actual resolution honestly, so that a caller knows in advance what is not there to be found.

Radar is a different instrument

Sentinel-1 is the reason the thing works under cloud, and almost nothing learned about the optical lane transfers to it. Radar is speckled by nature, an interference pattern rather than noise, so the first honest run over Lagos returned 921 change regions covering 16% of the scene, all of it texture. Multi-looking with a boxcar filter and raising the magnitude floor took that to 239 regions and 3.4%.

The better lesson came from the detector. A brightness threshold computed over the whole scene found exactly zero vessels, because land backscatter is so much stronger than water that it dominates the statistics and sets a bar no ship can clear. Taking the same statistics from the water class alone found 42, of which 21 were large enough to be worth a human's time. The fix was not a better threshold. It was noticing that the threshold was being computed over two populations at once.

Which then needed a water mask on radar, where the usual optical trick is unavailable. An early version took the darkest 45% of the scene, which a synthetic test caught calling 94% of a flat image water. The replacement finds the split properly and then checks that the histogram was actually bimodal before trusting it. When it was not, it returns nothing rather than a guess.

Most of it is refusing to answer

The pattern above repeats often enough to be the design. A measurement system is more useful when it declines than when it guesses, and the refusals are the part that took the longest to get right.

It refuses a window that is too cloudy. An early Lagos run looked clean and was not: both passes had roughly a quarter and a third of the area under cloud or shadow, and six of the 29 detections were sitting on the edge of the cloud mask. The tiles had each passed a conventional scene-wide cloud filter, because a tile can be mostly clear while the specific box you asked about is not. It now screens on the area of interest itself, samples several candidate passes at each end of the window, and declines outright rather than answering over cloud.

It refuses to ask a person about something they cannot judge. Objects below about 1,500 square metres are withheld from the review queue and counted, so a clear pair yielding 72 detections presents 26 and says plainly that 46 were too small to call. And it refuses to save a model that has not earned it: training runs stratified cross-validation and will not write out a classifier that fails to beat simply guessing the majority class.

Even the natural-language command bar refuses, in its way. It is a few hundred lines of rules rather than a language model, for two stated reasons: a hosted model would break the promise that the whole thing runs without an API key, and a parser that can show you exactly how it read your sentence is worth more here than one that reads it better. It shows its interpretation and waits to be confirmed, rather than acting.

The detector, honestly

Everything above is about change detection, which is measured. The object detector is not, and the gap between those two statements is the most important thing on this page.

There are two. The first is a brightness baseline: find compact bright blobs against a local statistic, filter by size and shape, refuse to fire near masked pixels. It does not know what a ship is. The second is a proposal-and-rescore arrangement, where the baseline runs loose for recall and a small logistic regression over twelve hand-designed features re-scores each candidate. On a held-out synthetic scene of fourteen vessels on water and fourteen rooftops of identical brightness on land, that takes precision from 0.54 to 1.00 at unchanged recall.

That 1.00 means nothing about the world, and the repository says so in bold before anyone else can. The synthetic classes are separable by construction. The shipped model was cold-started on 909 synthetic proposals and carries a warning field in its own weights file saying it has never been validated on a real human label. The review queue, the correction store and the retraining path are all built and tested, and the count of real labels collected so far is zero. What has been demonstrated is that the machinery works end to end, not that the model is good.

What it cannot do

Two sites are not a validation set. They are anchors chosen because their answers were known, and a fair test would add a site with real construction and real ground truth rather than more places where nothing happens. The 1.99% residual at Lagos is also not purely error: some of it is genuine change across twenty days of a working harbour that cannot be separated from nuisance change without exactly the learned model that does not yet have labels.

It runs locally, single user, with no accounts and no deployment. And the honest summary of the whole system is that the part which is measured is good, the part which is measured is not the part people ask about, and saying so plainly is cheaper than being found out later.

Figures on this page come from the calibration harness and the generated showcase runs in the repository. Imagery is Sentinel-2 L2A and Sentinel-1 RTC, both free and public. The engine is NumPy, with no GPU and no hosted model anywhere in it.