My First Theorem
My First Theorem: the journey from 0 to 100.
9/5/20262 min read
My first theorem has a three-line proof. I say that up front because for a while it embarrassed me, and now it's the part I like best.
The experiment everyone runs
If you work in physics-informed machine learning, you know how this goes. You have a model and a piece of domain knowledge. For example, fire severity shouldn't go down as temperature goes up. You build that in as a constraint, retrain, and see how much accuracy you lose. If the drop is small, you write "the physics comes for free."
I've written that sentence myself. My permafrost risk model and POSEIDON both claim, in different ways, that respecting physics doesn't hurt accuracy. I still think those papers are fine. I just no longer think that kind of claim means much by itself.
The observation
Suppose a function doesn't depend on temperature at all. Then it is technically non-decreasing in temperature. It's also non-increasing, convex, concave and Lipschitz in temperature. It satisfies every shape constraint you could put on that feature, because it never uses the feature.
So the worst any constraint on a set of features can do is make the model ignore those features. That gives:
price of the prior ≤ value of the features it constrains
The value of the features is the ablation gap: how much you lose by dropping them. That's a number most papers already compute and then report in a separate table. Nobody connects the two.
Put practically, a constrained model can never be beaten by its own ablation. If it is, something in your pipeline is broken.
That's the theorem. The proof only uses the fact that the best score over a bigger set of models can't be worse than over a smaller one. It holds for any metric, with no assumptions about convexity or smoothness.
Why something this trivial matters
I tested it on a wildfire-severity dataset with about 27,000 records. I put monotone constraints on four weather drivers and left latitude and longitude unconstrained. A few results surprised me:
The coordinates shield the prior. With spatially blocked validation, latitude and longitude alone recover 92.9% of the full model's accuracy. That makes the drivers worth very little, so any prior on them looks cheap. With the coordinates included, the same constraint costs 0.0473 macro-F1. Without them it costs 0.3470. The physics is identical, but the cost is 7.3 times higher.
The cost depends on how you validate. When I made the spatial blocks coarser, from 1° to 10°, the ceiling on the cost fell from 0.0942 to 0.0050. Under the coarsest blocks, the constraint seemed to improve accuracy. That's not an insight about the physics. There was simply nothing left to measure.
Following a prior and paying for it are separate things. The unconstrained model breaks the textbook prior about 48% of the time, which is basically a coin flip. Enforcing that prior still costs only about 0.02. Usually the first number would be read as "the model didn't learn the physics" and the second as "the physics is free." Here both are true at once.
The part I'm proudest of
Since the inequality has to hold in principle, any time it fails in practice you're seeing noise in your own pipeline. I checked 318 orderings that the experiment was obliged to satisfy. Eighteen were violated, and the largest violation was 0.0220 macro-F1. That number is a resolution floor the experiment calibrates for itself. Any price smaller than it can't be interpreted.
By that standard, four of the twenty-one price results in my own headline table can't be interpreted. I report them that way. A test that never catches its own author doesn't show that it catches anything.