AI in real estate, signal vs noise
Where the Model Is Reliable
6 min read
Almost every argument published about automated valuation is an argument about its failures. The failures are real, well documented, and the case has been made thoroughly enough that the opposite question has gone unasked: under what conditions is the model reliable, and what defines the boundary?
The question matters commercially. A category that only describes when a tool fails is making a claim it cannot calibrate, and a claim that cannot be calibrated is eventually discounted by the people hearing it.
The error is a function of variance, not method
An automated valuation compares a subject property to recent sales and adjusts for differences. Its known weakness is the differences it cannot observe: condition, finish, road exposure, light, block boundaries, the neighbour.
That framing suggests the problem is the method. It is more precisely a problem of sample variance. The magnitude of the error depends on how much the unobserved variables differ across the comparable set. Where they differ substantially, the model averages across facts it cannot detect and the number drifts. Where they barely differ, there is little for it to miss.
Which yields a testable boundary rather than a general position. The model is reliable in proportion to the homogeneity of the available comparables, and homogeneity is something an agent can assess directly by looking at the set.
WHERE THE BOUNDARY FALLS
| Condition of the comparable set | Model performance |
|---|---|
| Repeated floor plans, one builder, one phase | Strong |
| Association-enforced exterior consistency | Strong |
| Uniform lot size and orientation | Strong |
| Mixed vintages on the same street | Degrading |
| Renovated and original stock intermixed | Poor |
| Custom builds, irregular lots, view lines | Unreliable |
The variable is not the sophistication of the method. It is how much the unobserved facts vary across the sample.
Why the concession is worth making
Three reasons, and the third is the one that matters.
It is true, and the counterparty can check it. A homeowner in repeated tract stock who compares an estimate against the last six sales of their own floor plan will find the estimate close. A category telling them otherwise is telling them something they can disprove from a phone.
It makes the criticism specific. Naming the conditions under which a tool fails is a stronger claim than asserting that it fails, because it survives the counterexample. Every general claim about automated valuation being unreliable meets a seller in a subdivision who knows better.
And it relocates the argument to firmer ground. If reliability is a function of comparable homogeneity, then the professional contribution is not producing a better number in the easy case. It is recognising which case you are in, which requires seeing the property and the set. That is a defensible position precisely because it concedes the easy case rather than contesting it.
What this implies for confidence reporting
It also supplies the thing most confidence measures lack. Confidence is commonly computed from how tightly the selected comparables agree with each other, a measure that confuses agreement with completeness.
Homogeneity of the sample is a different and better input. A comparable set of repeated plans from one builder is tight for a knowable reason, and a set assembled from mixed vintages and unknown renovation histories is loose for an equally knowable reason. That distinction can be reported. It rarely is, which is why the same confident figure arrives on a tract property and on a custom build, as set out in how AVMs work and where they break.
When are automated home valuations reliable?
When the comparable set is genuinely homogeneous. Master-planned subdivisions with repeated floor plans, one builder, uniform lots, and association-enforced consistency remove most of the variation an automated model cannot observe, so there is little left for it to miss. Reliability degrades as that variation returns.
Why do automated valuations fail on custom homes?
Because the facts that separate one custom property from another are the facts the model cannot read. Finish level, view lines, lot orientation, and renovation quality vary widely across the comparable set and appear in no field, so the model averages across differences it cannot detect and reports the result with unchanged confidence.
Should a confidence score reflect how similar the comparables are to each other?
It should reflect why they are similar. Comparables that match because they are the same floor plan by the same builder support a tight range for a knowable reason. Comparables that happen to agree despite mixed vintages and unknown renovation histories do not, and most confidence measures cannot tell those two situations apart.
A narrower claim is a stronger one
Context Blindness™ is a description of a boundary, not a verdict on a technology. The gap is between the facts a record holds and the facts that decide a price, and its width varies enormously by property type. In engineered-uniform housing it narrows toward nothing. In custom stock it is the entire distance.
The version of this argument that claims the model is always wrong is easier to write and weaker to hold, because the first counterexample undoes it. The version that specifies where the boundary falls is harder to write and survives contact with a homeowner who checked. A reasoned valuation, the approach CMAflow builds, reports a tight range where the comparable evidence is dense and a wide one where it is not, and both are honest outputs of the same method. Conceding the easy case is what makes the hard case arguable.
This article is general information and analysis, not financial, lending, or appraisal advice. Verify any home value with a licensed professional before acting.
The Independent Agent
Substack | Spotify | CMAflow FAQ | YouTube | Free CMA | Home valuation | Insights | Blog