Residential GDV valuation model performance
Our GDV model (v7) estimates the Gross Development Value of residential Buy-to-Let property across England & Wales, learning from millions of real Land Registry sales. This report publishes held-out accuracy, calibration, limitations and the evidence used by the currently served v7 release.
The test that matters most
On 521 recent post-renovation resales the model had not seen, the median difference between our estimate and the achieved sale price was 7.9% (July 2026).
Most model accuracy figures are measured on ordinary sales — stock that was not just refurbished. That is not the question this product asks. So we built a test set of properties that demonstrably were refurbished — independently evidenced renovations from our own outcomes dataset (construction method proprietary) — restricted to sales after the model’s training cut-off. On those 521 sales, v7recorded a median difference of 7.9% against the price actually achieved. The same 521 sales recorded 12.7% under v6.
Both versions read a refurbished resale low, and that is worth explaining rather than hiding: a model anchored on nearby comparable sales prices toward the pre-works local market, and beating that market is the whole point of a refurbishment. The measured shortfall narrowed from −12.0% under v6 to −3.6% under v7 on this sample.
Section 1
How Accurate Are Our Predictions?
Measured on the same held-out test fold as v5 and v6 — 228,443 England & Wales sales from July 2025 to May 2026 that the model never saw during training, so the version-on-version comparison is apples to apples. We report the standard measures defined below, and publish the weaker segments alongside the strong ones rather than a cherry-picked headline.
Prediction Error Distribution
Where does each prediction land relative to the actual sale price?
Performance by Price Band
Every band we measure, including the ones we are worst at. Median % error across the 228,443 held-out sales, grouped by the price the property actually achieved.
Keeping Comparables Fresh
Introduced in v6 and carried into v7 unchanged: every comparable sale is inflation-adjusted (HPI-indexed, by district and property type) to the valuation date before it enters the model, weighted by recency and by how closely its size matches the subject property. Raw, unindexed comparables are kept alongside the indexed ones so the model can prefer real evidence on older, condition-dominated stock. Tested live, on the model served at the time (v6), against 100 recent completions (April/May 2026) across Derby, Birmingham, Manchester, London and Liverpool: median signed error moved from −1.7% under the previous comparable method — a systematic under-valuation caused by stale comparables — to +0.3% under v6, effectively eliminating that bias, and PPE10 improved from 30.8% to 36.0%. Same-type comparable coverage was slightly lower (91/100 vs 100/100), a deliberate trade-off of strict same-type matching for freshness.
Why two properties get different-width ranges
Not every estimate rests on the same weight of evidence, so not every estimate gets the same range. Estimates backed by the property’s own transaction history carry tighter ranges: where we can take a previous sale of that exact address, index it forward to the valuation date and see it agree with the model, we know a good deal more than we do about one with no corroborating record. Where that history is missing and nearby comparable sales are thin or scattered, the range is deliberately wider. Every estimate carries a confidence chip showing which of the three tiers it fell into, so the strength of the evidence is on the face of the result rather than buried in it.
The effect is large enough to be worth publishing. Across the 21,777 sub-£160k sales in the calibration sample (sale months January to June 2025), estimates backed by corroborating property-specific evidence recorded median errors of 6.4% to 9.3% depending on price band, against 10.4% to 14.0% for estimates without one. That difference is expressed as range width. A £120k estimate with a corroborating evidence carries a range of about −8% to +6%; the same estimate with no such evidence carries −13.5% to +11.7%.
Ranges are not forced to sit evenly either side of the central figure. They are measured separately above and below, because on many kinds of stock the downside and the upside genuinely are not the same distance away.
Section 2
Getting Better Every Version
Each version adds data and sharpens the model. Mean error is down around 26% since v1, and v7 improves on v6 both overall and on the low-value stock that has always been hardest.
MAE — v1 to v7
Version Changelog
Note: v1–v3 were originally tracked on mean-error measures; from v4 onward we report the median-based industry-standard measure (MdAPE) shown here. v1–v4 figures come from an earlier evaluation set, while v5, v6 and v7 are measured on an identical 228,443-sale held-out fold — so cross-era comparisons are indicative, not exact.
Section 3
Our Data Foundation
Every source is open government data or publicly available regulatory data. Fully auditable, licence-compliant, and updated regularly.
Known Limitations
- Sub-£100k performance is structurally weaker across every model version, and weakest below £50k: at that price the property's condition drives the sale more than anything the data can observe. Those cases are flagged Low-confidence and shown with wider ranges so users can decide what further evidence or professional review they need.
- Land Registry sales are published after completion, and completions keep arriving for months afterwards. The most recent months in any measurement are thinner than they will eventually be, and features built from local sale prices deliberately exclude the two most recent months so that what we measure offline matches what a live estimate could actually see.
- The model cannot inspect a property. It estimates the price of a property in the condition its observable official record implies. Internal condition, works in progress, specification, and any incentive sitting behind a headline price are not visible to it.
- Leasehold is handled by a rule applied after the model, not by the model itself: a short unexpired lease triggers a fixed reduction on flat-like stock. Lease length, ground rent and service charge are not model inputs. It is an approximation, and it is labelled as one in the audit trail behind every figure.
- DEFRA noise mapping covers England only — Welsh properties currently read as unmapped/quiet — and reflects a snapshot in time rather than live monitoring.
- Most figures above are backtest results on historical sales; the freshness evidence in Section 1 is drawn from a live evaluation of the model served at the time (v6). Neither is a guarantee of future accuracy, and outputs are guide valuations, not RICS Red Book valuations.
Section 4
Ongoing Monitoring & Reliability
We track live performance monthly against real Land Registry completions, using the same MAE, MdAPE and PPE measures reported above. If performance drifts beyond backtest expectations, an early retrain is triggered.
Normal
Live performance in line with backtest expectations — no action, monthly check continues.
Monitor
Early signs of drift from backtest expectations — checks increase in frequency and macro inputs are reviewed.
Action
Sustained drift confirmed — an early retrain is triggered.
Retraining Schedule
How We Keep It Honest
Every month, fresh Land Registry completions are scored against the model’s predictions using the same MAE, MdAPE and PPE measures reported throughout this page. Models in this class typically hold steady for some time after training before market movement causes drift; when drift is confirmed, calibration is refreshed or an early retrain is triggered, keeping live performance close to the backtest figures shown here.
How this maps to the RICS comparable-evidence approach
Our valuation engine is built to automate the evidence hierarchy set out in the RICS professional standard Comparable evidence in real estate valuation. Where a surveyor gathers and weighs evidence by hand, the model learns to do the same, from millions of real sales and the same primary sources.
Category A — direct comparables
Completed HM Land Registry sales of the same building and street — including the subject's own sale history — the strongest evidence the model is trained on, and its most influential inputs where they exist. Since v6, every comparable is indexed to a common valuation date and weighted by how closely its size matches the subject, with raw (unindexed) sales kept alongside for older, condition-dominated stock.
Category B — market data
UK House Price Index trends, sector-level £/m² aggregates, EPC records, market activity and DEFRA road & rail noise mapping (England) — the general-guidance layer, used exactly as the standard prescribes: to adjust and contextualise, never as the answer on its own.
Category C — background
Interest rates, mortgage pricing and inflation feed the model as context signals, mirroring the standard's "other sources" tier.
Crucially, the model doesn’t treat comparables as an afterthought bolted on top — it learns directly from them. Recent local sales and £/m² context are its single largest group of inputs, so every valuation is grounded in comparable evidence by construction. Where that evidence is thin or a property is unusual, the estimate is shown as lower-confidence with the reason visible to the user.
What this is not: a data-driven valuation cannot inspect a property, verify its internal condition or specification, uncover sale incentives behind headline prices, or exercise the professional judgement of a RICS Registered Valuer. Our figures are guide valuations for screening and negotiation — good-condition market value on the street, from the best available evidence. Where lending, legal or formal certainty is required, commission a RICS Red Book valuation; nothing here substitutes for one.
Model: GDV v7 (btl-gdv-v7), trained 27 July 2026 on 2,756,196 England & Wales transactions across 124 features. Except where stated otherwise, metrics on this page are measured on the same held-out test set used for v5 and v6 — 228,443 sales from July 2025 to May 2026, none of them used during training. Overall and per-band median-error figures are measured on the two-model configuration that is actually served. MAE, PPE10 and PPE20, and the 521-sale post-renovation benchmark, are measured on the same sales for the full trained ensemble, which differs from the served configuration by 0.05 percentage points of median error. Range coverage and the evidence-tier figures are measured on a separate calibration sample of sales from the first half of 2025. Predictions are made against an April 2026 anchor month and indexed forward to the current month using the UK House Price Index, estimated where the official index has not yet caught up; each output is labelled with the month it is indexed to. Data sources: HM Land Registry, MHCLG, ONS, Ofcom, FSA, DEFRA — open government or official statistics licences. These are measurements on the specific samples named above, not a guarantee of future accuracy, and outputs are guide valuations, not RICS Red Book valuations. Report reviewed 23 August 2026.