Entry 04Frostline · spatial modelling2026

The most useful result was the one I threw out

Depth is predictable from location. Yield is not. Both of those took real work to establish, and only one of them is a finding anybody wants to hear.

well records50,083
apparent gain+38.8%
actually a restatement~96%
targetcut

What a well record can and cannot tell you

A drilling record carries where the hole is, how deep it went, what the driller hit on the way down, and what it produced. The commercially interesting question is the last one — how much water will this location give — and that is the one the data refuses to answer.

Depth behaves. Neighbouring wells agree with each other, so location carries real information and a model can use it. Yield does not behave. A well's output is not predicted by its neighbours, and no amount of feature work changed that. Establishing the negative took as long as establishing the positive, and it is the more valuable of the two: it says which promises the business must not make.

The two maps

Same 49,527 wells, same grid, same method — median of every well in a two-kilometre cell. On the left, well depth. On the right, tested yield. Depth resolves into contiguous shapes; yield does not resolve into anything. That contrast is the finding, and it is measurable rather than a matter of taste.

Median well depthMoran's I = 0.393

Hover to sample

Median tested yieldMoran's I = 0.089

Hover to sample

VariableMoran's ICellsNeighbour pairs
Well depth0.3939393,232
Tested yield0.0899163,130

Moran's I measures whether neighbouring cells resemble each other: 0 is indistinguishable from noise. Grand Traverse, Antrim, Kalkaska and Charlevoix counties, cells of roughly two kilometres with at least five wells each. Computed from the well table on 29 July 2026 — these are the project's own numbers, not an illustration of them.

The target that looked like a win

The aquifer-interval fields are unusually well populated — top and bottom depths present on roughly nine in ten records — and had never been modelled. A leave-one-out skill test put local prediction 38.8% better than the global median. That looked like a new signal.

It was not. The global median is the wrong baseline for a variable that sits at about three-quarters of the well depth. Measured against a depth-derived null instead, the apparent skill disappeared: the aquifer interval is roughly 96% a restatement of the well depth already in hand. It was cut rather than shipped, and if the interval is surfaced at all it should be a derived ratio off the existing depth band, not a modelled target.

BaselineApparent skill
Against the global median+38.8%
Against a depth-derived nullno signal

The first row is not a mistake anyone would notice from the outside. Choosing the null is the whole experiment.

What is still open, stated honestly

Lithology — untouchedTwo hundred thousand-odd logged intervals across the same wells, describing what the driller actually hit. Genuinely independent of depth, and therefore the one place a new signal could still be. Not yet modelled.
Static water level — untestedA first pass scored badly, but the data is dirty enough that the number means nothing: sentinel values standing in for missing readings, and hundreds of wells recording a water level deeper than the well itself. Untested, not disproven — and the first-pass figure must not be quoted as a finding.
What the business may claimDepth and cost bands, with an interval rather than a point estimate. Not yield. A model that cannot predict output is not permitted to imply that it can.
Spatial modellingNull result PythonPublic records

Questions about the method? Email is best — [email protected] · Back to portfolio