← All documentation
Model documentationPublished

Residential GDV valuation model performance

Our GDV model (v7) estimates the Gross Development Value of residential Buy-to-Let property across England & Wales, learning from millions of real Land Registry sales. This report publishes held-out accuracy, calibration, limitations and the evidence used by the currently served v7 release.

Version btl-gdv-v7 · trained 27 July 2026 · England and Wales
£40,470
MAE
Mean £ error — lower is better
7.6%
MdAPE
Median % error — half land inside
61%
PPE10
Within 10% of the sale price
84%
PPE20
Within 20% of the sale price
v7
Current version
Median error 8.7% → 7.6% vs v6, on the same held-out sales

The test that matters most

On 521 recent post-renovation resales the model had not seen, the median difference between our estimate and the achieved sale price was 7.9% (July 2026).

Most model accuracy figures are measured on ordinary sales — stock that was not just refurbished. That is not the question this product asks. So we built a test set of properties that demonstrably were refurbished — independently evidenced renovations from our own outcomes dataset (construction method proprietary) — restricted to sales after the model’s training cut-off. On those 521 sales, v7recorded a median difference of 7.9% against the price actually achieved. The same 521 sales recorded 12.7% under v6.

Both versions read a refurbished resale low, and that is worth explaining rather than hiding: a model anchored on nearby comparable sales prices toward the pre-works local market, and beating that market is the whole point of a refurbishment. The measured shortfall narrowed from −12.0% under v6 to −3.6% under v7 on this sample.

What this sample is and is not. These 521 sales are renovated stock, not a cross-section of the market — they measure how closely the model tracks a known-renovated resale, not how it performs on an average sale. A further 241 qualifying resales could not be matched into our records with certainty and were dropped rather than estimated. Within the sample, the £0–100k band holds only 11 sales, which is too few to read as a result. Figures are measured on this specific sample at July 2026 and are not a statement about future sales.
🔍
Fully Explainable
Every prediction can be attributed to specific input features — not a black box. The model class is designed for auditability from the ground up.
🏦
Institutionally Proven
We use the same category of models deployed by insurers, credit institutions and banks for risk scoring and underwriting — a well-established, regulator-familiar approach, not an experimental one.
📋
Governance Documented
Full model governance documentation — covering data lineage, methodology, retraining controls, and audit trails — is available to partner lenders on request.

Section 1

How Accurate Are Our Predictions?

Measured on the same held-out test fold as v5 and v6 — 228,443 England & Wales sales from July 2025 to May 2026 that the model never saw during training, so the version-on-version comparison is apples to apples. We report the standard measures defined below, and publish the weaker segments alongside the strong ones rather than a cherry-picked headline.

MAE
Mean Absolute Error
The average £ gap between our prediction and the actual sale price, across every property in the test fold.
MdAPE
Median Absolute Percentage Error
The typical (middle) percentage error — half of predictions are closer than this, half are further away.
PPE10
Share within 10%
The percentage of test-fold predictions that landed within 10% of the actual sale price.
PPE20
Share within 20%
The percentage of test-fold predictions that landed within 20% of the actual sale price.
£40,470
MAE
7.6%
MdAPE
61%
PPE10
84%
PPE20
50.6%
P25–P75 coverage

Prediction Error Distribution

Where does each prediction land relative to the actual sale price?

Performance by Price Band

Every band we measure, including the ones we are worst at. Median % error across the 228,443 held-out sales, grouped by the price the property actually achieved.

£0–50k — known weakness
692 held-out sales
MdAPE 51.0%
£50–75k — known weakness
3,340 held-out sales
MdAPE 27.4%
£75–100k
6,267 held-out sales
MdAPE 16.3%
£100–150k
20,394 held-out sales
MdAPE 11.2%
£150–175k
14,007 held-out sales
MdAPE 8.9%
£175–250k
46,038 held-out sales
MdAPE 7.4%
£250–400k
73,088 held-out sales
MdAPE 6.2%
£400–700k
48,254 held-out sales
MdAPE 7.1%
£700k–5m
16,363 held-out sales
MdAPE 10.0%
Known weakness below £100k. Materially improved in v7. On this fold, median error at £75–100k fell from 20.8% under v6 to 16.3%, and at £100–150k from 13.5% to 11.2%. Below £75k the model is still well outside its usual range, and below £50k it is weakest of all — at that price the property’s condition drives the sale more than anything the data can see. Deep sub-£50k cases are flagged Low-confidence and carry a deliberately wide range so users can see that extra evidence and professional review may be appropriate. The estimate remains available.
Rural and urban stock now measure the same: 7.6% median error on 57,042 rural sales and 7.6% on 162,732 urban sales in this fold. v6 could not report a rural figure at all — the classification data behind that check had not been seeded, which we disclosed at the time as a gap in our own evidence rather than a known weakness.

Keeping Comparables Fresh

Introduced in v6 and carried into v7 unchanged: every comparable sale is inflation-adjusted (HPI-indexed, by district and property type) to the valuation date before it enters the model, weighted by recency and by how closely its size matches the subject property. Raw, unindexed comparables are kept alongside the indexed ones so the model can prefer real evidence on older, condition-dominated stock. Tested live, on the model served at the time (v6), against 100 recent completions (April/May 2026) across Derby, Birmingham, Manchester, London and Liverpool: median signed error moved from −1.7% under the previous comparable method — a systematic under-valuation caused by stale comparables — to +0.3% under v6, effectively eliminating that bias, and PPE10 improved from 30.8% to 36.0%. Same-type comparable coverage was slightly lower (91/100 vs 100/100), a deliberate trade-off of strict same-type matching for freshness.

Why two properties get different-width ranges

Not every estimate rests on the same weight of evidence, so not every estimate gets the same range. Estimates backed by the property’s own transaction history carry tighter ranges: where we can take a previous sale of that exact address, index it forward to the valuation date and see it agree with the model, we know a good deal more than we do about one with no corroborating record. Where that history is missing and nearby comparable sales are thin or scattered, the range is deliberately wider. Every estimate carries a confidence chip showing which of the three tiers it fell into, so the strength of the evidence is on the face of the result rather than buried in it.

The effect is large enough to be worth publishing. Across the 21,777 sub-£160k sales in the calibration sample (sale months January to June 2025), estimates backed by corroborating property-specific evidence recorded median errors of 6.4% to 9.3% depending on price band, against 10.4% to 14.0% for estimates without one. That difference is expressed as range width. A £120k estimate with a corroborating evidence carries a range of about −8% to +6%; the same estimate with no such evidence carries −13.5% to +11.7%.

Ranges are not forced to sit evenly either side of the central figure. They are measured separately above and below, because on many kinds of stock the downside and the upside genuinely are not the same distance away.

You never get a single number. Every valuation shows a P50 target — our central, best estimate — the midpoint of the likely range, where half of outcomes land above and half below — and a P25 conservative figure to underwrite against. Deals include the conservative case so users can inspect what happens if the market lands soft. Margin, refinance and exit calculations show that scenario alongside the central estimate; neither result approves, rejects or prevents use of a deal.The range is calibrated, not decorative. The P25–P75 range is set so that the price actually achieved falls inside it roughly half the time — that is what a quarter-to-three-quarter range means. Measured on 70,291 held-out sales from May and June 2025, the achieved price landed inside the range 50.6% of the time overall, and between 46.7% and 54.6% in every price band. A range wide enough to catch the answer nine times in ten would tell you almost nothing about the property.
Why valuations show an “indexed to” month. Under the bonnet, the model is trained and predicts in a fixed reference period — currently its April 2026 anchor month — so that every comparable sale is judged on a like-for-like basis regardless of when it actually completed. At the point a valuation is served, that anchor-month prediction is indexed forward to the current month using the UK House Price Index for the property’s district and type. Where the official index has not yet caught up to the current month, the remaining months are statistically estimated, and the output is labelled “estimated” alongside the last officially published HPI month. In the published v7 evaluation, valuations were indexed to July 2026 on this estimated basis, with the UK HPI officially published to April 2026. Each live valuation identifies its own index month rather than relying on this historical report date.

Section 2

Getting Better Every Version

Each version adds data and sharpens the model. Mean error is down around 26% since v1, and v7 improves on v6 both overall and on the low-value stock that has always been hardest.

MAE — v1 to v7

Version Changelog

v1
Foundation
Core property and location data across England & Wales sales.
MAE £54,284
v2
Lower-value stock
A dedicated approach for cheaper property and a broader data set.
MAE £53,371
v3
Sharper calibration
Better handling of price scale, evidence weighting and confidence ranges.
MAE £48,036
v4
Richer signals
Broadband, food-hygiene and local market-range signals; deeper tuning.
MAE £47,679MdAPE 9.4%
v5
Investor stock + time-proofing
Investor and new-build sales included, valuations indexed to their valuation date, and confidence ranges recalibrated.
MAE £45,755MdAPE 8.9%
v6
Freshness + high-value specialist
Every comparable sale is HPI-indexed to a common valuation date and weighted by size-similarity; DEFRA road & rail noise mapping added; a new specialist model handles £400k+ stock; predictions are re-inflated to the latest HPI month at serving time.
MAE £45,103MdAPE 8.7%
v7 ✦
Property-level evidence + calibrated ranges
The model now draws on property-specific evidence layers built from official records (details proprietary) — which is what lets it price a refurbished resale rather than the pre-works local market. Confidence ranges move to a separate calibration layer that sets each range from the strength of the evidence behind that particular estimate.
MAE £40,470MdAPE 7.6%

Note: v1–v3 were originally tracked on mean-error measures; from v4 onward we report the median-based industry-standard measure (MdAPE) shown here. v1–v4 figures come from an earlier evaluation set, while v5, v6 and v7 are measured on an identical 228,443-sale held-out fold — so cross-era comparisons are indicative, not exact.


Section 3

Our Data Foundation

Every source is open government data or publicly available regulatory data. Fully auditable, licence-compliant, and updated regularly.

HM Land Registry — Price Paid
All residential sales in England & Wales — our primary training target. The full 1995–2026 archive is our primary training target and comparable-evidence base.
OGL v3Monthly
EPC Register (MHCLG)
Floor area, rooms, construction age and EPC rating — the closest thing Britain has to a national property-condition database.
OGL v3Monthly
Renovation-outcomes dataset (ours)
Tens of thousands of documented UK renovation outcomes — before-and-after evidence of what refurbishment actually adds. Built in-house from official records; construction method proprietary. Used to benchmark post-renovation pricing.
Proprietary derivationPer data refresh
ONS Deprivation Index (IMD)
10 deprivation sub-scores at LSOA level: income, crime, health, education and more.
OGL v3Every 4 yrs
Ofcom Connected Nations 2025
Postcode-level broadband: superfast, gigabit, and below-USO coverage.
Open DataAnnual
FSA Food Hygiene (FHRS)
Hygiene ratings, restaurant and pub counts per postcode sector.
OGL v3Quarterly
OS Code Point + BoE / ONS Macro
Postcode centroids for distance features; Bank of England and ONS macro rates.
OGL v3Monthly
DEFRA Strategic Noise Mapping (Round 4)
Road & rail noise (Lden) bands, England, mapped to a ~100m grid — used as an ordinal amenity feature. England only; Wales currently reads as unmapped/quiet.
OGL v3Per mapping round (2022)
ONS / HM Land Registry UK House Price Index
Indexes every comparable sale to the model's valuation anchor month, and re-inflates served predictions to the latest reliable HPI month.
Official StatisticsPer HPI release

Known Limitations

  • Sub-£100k performance is structurally weaker across every model version, and weakest below £50k: at that price the property's condition drives the sale more than anything the data can observe. Those cases are flagged Low-confidence and shown with wider ranges so users can decide what further evidence or professional review they need.
  • Land Registry sales are published after completion, and completions keep arriving for months afterwards. The most recent months in any measurement are thinner than they will eventually be, and features built from local sale prices deliberately exclude the two most recent months so that what we measure offline matches what a live estimate could actually see.
  • The model cannot inspect a property. It estimates the price of a property in the condition its observable official record implies. Internal condition, works in progress, specification, and any incentive sitting behind a headline price are not visible to it.
  • Leasehold is handled by a rule applied after the model, not by the model itself: a short unexpired lease triggers a fixed reduction on flat-like stock. Lease length, ground rent and service charge are not model inputs. It is an approximation, and it is labelled as one in the audit trail behind every figure.
  • DEFRA noise mapping covers England only — Welsh properties currently read as unmapped/quiet — and reflects a snapshot in time rather than live monitoring.
  • Most figures above are backtest results on historical sales; the freshness evidence in Section 1 is drawn from a live evaluation of the model served at the time (v6). Neither is a guarantee of future accuracy, and outputs are guide valuations, not RICS Red Book valuations.

Section 4

Ongoing Monitoring & Reliability

We track live performance monthly against real Land Registry completions, using the same MAE, MdAPE and PPE measures reported above. If performance drifts beyond backtest expectations, an early retrain is triggered.

Normal

Live performance in line with backtest expectations — no action, monthly check continues.

Monitor

Early signs of drift from backtest expectations — checks increase in frequency and macro inputs are reviewed.

Action

Sustained drift confirmed — an early retrain is triggered.

Retraining Schedule

Full model retrainAnnual (Q1)
Calibration refreshQuarterly
Performance monitoringMonthly
Emergency retrain triggerOn sustained drift

How We Keep It Honest

Every month, fresh Land Registry completions are scored against the model’s predictions using the same MAE, MdAPE and PPE measures reported throughout this page. Models in this class typically hold steady for some time after training before market movement causes drift; when drift is confirmed, calibration is refreshed or an early retrain is triggered, keeping live performance close to the backtest figures shown here.

How this maps to the RICS comparable-evidence approach

Our valuation engine is built to automate the evidence hierarchy set out in the RICS professional standard Comparable evidence in real estate valuation. Where a surveyor gathers and weighs evidence by hand, the model learns to do the same, from millions of real sales and the same primary sources.

Category A — direct comparables

Completed HM Land Registry sales of the same building and street — including the subject's own sale history — the strongest evidence the model is trained on, and its most influential inputs where they exist. Since v6, every comparable is indexed to a common valuation date and weighted by how closely its size matches the subject, with raw (unindexed) sales kept alongside for older, condition-dominated stock.

Category B — market data

UK House Price Index trends, sector-level £/m² aggregates, EPC records, market activity and DEFRA road & rail noise mapping (England) — the general-guidance layer, used exactly as the standard prescribes: to adjust and contextualise, never as the answer on its own.

Category C — background

Interest rates, mortgage pricing and inflation feed the model as context signals, mirroring the standard's "other sources" tier.

Crucially, the model doesn’t treat comparables as an afterthought bolted on top — it learns directly from them. Recent local sales and £/m² context are its single largest group of inputs, so every valuation is grounded in comparable evidence by construction. Where that evidence is thin or a property is unusual, the estimate is shown as lower-confidence with the reason visible to the user.

What this is not: a data-driven valuation cannot inspect a property, verify its internal condition or specification, uncover sale incentives behind headline prices, or exercise the professional judgement of a RICS Registered Valuer. Our figures are guide valuations for screening and negotiation — good-condition market value on the street, from the best available evidence. Where lending, legal or formal certainty is required, commission a RICS Red Book valuation; nothing here substitutes for one.

Model: GDV v7 (btl-gdv-v7), trained 27 July 2026 on 2,756,196 England & Wales transactions across 124 features. Except where stated otherwise, metrics on this page are measured on the same held-out test set used for v5 and v6 — 228,443 sales from July 2025 to May 2026, none of them used during training. Overall and per-band median-error figures are measured on the two-model configuration that is actually served. MAE, PPE10 and PPE20, and the 521-sale post-renovation benchmark, are measured on the same sales for the full trained ensemble, which differs from the served configuration by 0.05 percentage points of median error. Range coverage and the evidence-tier figures are measured on a separate calibration sample of sales from the first half of 2025. Predictions are made against an April 2026 anchor month and indexed forward to the current month using the UK House Price Index, estimated where the official index has not yet caught up; each output is labelled with the month it is indexed to. Data sources: HM Land Registry, MHCLG, ONS, Ofcom, FSA, DEFRA — open government or official statistics licences. These are measurements on the specific samples named above, not a guarantee of future accuracy, and outputs are guide valuations, not RICS Red Book valuations. Report reviewed 23 August 2026.

Great deals deserve independent evidence.

Run a real address through the model, see the working, and decide for yourself. Two quick checks are on us.

Join the waitlist
We are onboarding in waves ahead of the public launch. Tell us which side of the table you sit on and we will start you in the right place.
No spam, no countdown timers, no urgency tactics. One email when your wave opens. Already have an account? Sign in.