← The PropTech Desk

AI in real estate, signal vs noise

When the AI invents a comparable

6 min read

A confident comp with no source is not a data point. It is a sentence that sounds like one.

The short version

A general language model does not look up sales. It generates text that reads like a plausible answer, and when asked for comparable sales without a live data connection it can manufacture them, complete with a street address, a sale price, and a closing date, none of which exist in any record. This is not the same problem as an out-of-date market read. A stale number is a real figure from the wrong time. An invented comp is not real at all, and it arrives in the same confident voice as everything else. The defense is not cleverness, it is verification: a comp you cannot find in the MLS or the public record is worth nothing, no matter how convincing the sentence around it.

Why a model makes up sales

The model predicts likely text. Asked for three comps near an address, it produces three well-formed comps, because that is what a good answer looks like, whether or not it has the sales to back them. Without a live feed to the MLS or public records, there is nothing anchoring the output to reality, and the fluency of the result is not evidence that the sales are real.

CheckA real comparableAn invented one
SourceAn MLS number or county recordNone, or a link that goes nowhere
AddressResolves to a real parcelPlausible, but does not exist
Sale priceMatches the recorded figureConfident, and unrecorded
Close dateA verifiable dateVague, round, or impossible
Condition notesFrom the listing or an inspectionGeneric and unfounded

Fabrication is not staleness

These are two separate failures, and they need separating. A stale answer reports the market as it stood at the model's training cutoff, the problem in why the AI market read is months old. A fabricated comp is not from any point in time, it is generated whole. A number can be current-sounding and still be pure invention, which is why checking the date is not enough. You have to check that the sale exists.

The cost of one invented comp

Valuations rest on a small set of sales, often three to six. Slide one fabricated comp into a three-comp set and it drags the average by tens of thousands of dollars, in whatever direction the invented price points. A single invented comp priced $60,000 above the real ones lifts a three-comp average by about $20,000. That is not a rounding error, it is a false foundation, and it is exactly the kind of confident, sourceless number a seller can carry into a listing appointment, the pattern in when a client brings a chatbot valuation.

The habit that catches it

Ask for the source of every comp, every time. A real comparable comes with an MLS number or a county record you can open. An invented one comes with a confident sentence and nothing behind it. The same discipline is what separates an honest valuation tool from a black box, the standard laid out in how AI valuation tools should be judged: show the comparables, or the number is not usable.

What this means for an agent

When a client brings an AI valuation, do not debate the figure, ask to see the comps behind it. Then open the MLS and try to find them. A sale that will not resolve to a record is the whole conversation, and you never have to say a word against the tool. Showing a client which of their AI comps are real and which never happened is the basis of our home-value accuracy page, and the honest-range consumer estimate is the home-value page.

Frequently asked questions

Can ChatGPT find real estate comparables?

Not reliably from memory. A general language model generates plausible text rather than looking up sales, so without a live data connection it can produce comps, addresses, and prices that do not exist. Any comp it gives needs verifying in the MLS or public records.

Why does AI make up home sales that never happened?

Because the model predicts likely text, not retrieved facts. Asked for comparable sales, it produces well-formed examples because that is what an answer should look like, even when it has no actual sales behind them.

How do I know if an AI valuation is using real data?

Ask for the source of each comparable. Real comps come with an MLS number or a county record you can open and confirm. If the sources are missing, broken, or unverifiable, treat the valuation as unfounded.

Are AI-generated comps reliable?

Only if every one can be verified against a real record. A single fabricated comp in a small set can move the estimate by tens of thousands of dollars, so an unverified comp is worth nothing regardless of how convincing it sounds.

How do I verify a comparable sale?

Look it up in the MLS or the county record: confirm the address resolves to a real parcel, the sale price matches the recorded figure, and the closing date checks out. If any of those fails, the comp is not usable.


Sources: Published documentation of large language model behavior, including the tendency to generate fabricated but plausible details without a live data source, standard MLS and public-record verification practice, and appraisal guidance on documenting comparable sales. The mechanism, generated text presented as retrieved fact, is the durable point.

This article is general information and analysis, not financial, lending, or appraisal advice. Verify any home value with a licensed professional before acting.


The Independent Agent
Substack | Spotify | CMAflow FAQ | YouTube | Free CMA | Home valuation | Insights | Blog