I’ve spent years digging through house price records, and one thing always strikes me: no two datasets tell the same story the same way. Comparative international analyses of house prices run into trouble fast because every country handles definition, data structure, spatial scales, and time scales differently, and that messy coverage makes both comparative analysis and within-country analysis of housing markets harder than it should be. In the UK, house price data deficiencies such as missing Property Attribute Details have quietly held back research into residential house price variation, leaving big gaps in how we actually understand the housing market.
Historical Limits and Property Attribute Details in UK House Price Modeling
Go back far enough and you’ll find modelling of UK-based price shifts stretching to the 1970s, when most figures were either aggregated into coarse geographies like regions and districts, or narrowly tied to individual properties in one specific city. For decades, aggregate sample mortgage data from building societies think the 5% sample survey of Building Society Mortgages or the Nationwide Building Society mortgage data did the heavy lifting.
These records lacked local nuance and carried real biases from working off small samples, while micro-level housing data, such as the local estate agent survey data collected by Orford, opened the door to sharper local housing analysis with precise Property Attribute Details but never became widely available datasets.
LR-PPD and Open Data Modernization
Everything changed once Land Registry’s LR-PPD became open data in 2013. This single release transformed house price variation research across England and Wales because it holds a comprehensive record of residential transactions at address level going back to 1995. Certain sales, like Right to buy purchases at a discount, get excluded, yet the dataset still gives the clearest read on residential property sales at full market value, and bodies like the Office for National Statistics (ONS) rely on it to build the House Price Statistics for Small Areas release and the official House Price Index.
Addressing the Missing Floor Area Gap
Even with all that value, one gap kept nagging researchers: missing attribute information, especially total floor size, meant nobody could fully account for how stock mix shaped broader patterns, and floor space is widely accepted as the biggest single determinant of house price variation.
Orford tackled this by adding an estimated total floor area built from building footprints sourced through Ordnance Survey MasterMap and Environment Agency LiDAR data, though the approach struggles with flats and buildings where the number of stories isn’t obvious.
A more direct route links LR-PPD straight to total floor area figures held in Domestic Energy Performance Certificates, known as Domestic EPCs, an open dataset published by the MHCLG (Ministry for Housing Communities and Local Government) that records both energy performance and building attribute information like habitable rooms.
Open Access Data Linkage and Commercial Property Intelligence
Here’s the part that surprised me most: despite how useful this link is, barely any research studies have shared the actual linkage rate between LR-PPD and Domestic EPCs, and none had published the full linkage method and linkage data until now, with open access, reusable linkage codes built around house price per square metre.
Meanwhile, on the commercial side, DOMUS describes itself as every property holding a hidden timeline of lives lived, dreams chased, and milestones achieved, framing itself as the human story behind the bricks. Branded the UK’s most comprehensive property attributes database, it claims coverage of 32 million residential properties with exhaustive property information essentially the DNA of each dwelling.
Modern Commercial Property Intelligence and Address Cleansing
DOMUS gathers 350 attributes for every specific property, refreshed through updated daily feeds pulled from 4,500 data sources, blending both structured and unstructured formats. These property attributes get sorted into four buckets location, property, environment, and demographic factors spanning everything from square footage to risk factors. On the address-data side, AFD Software has spent 40 years perfecting accurate address solutions, and thousands of organisations now partner with the Postcode People at AFD to validate and cleanse their address data and location data.
Data description and development
Working with the LR-PPD dataset taught me to appreciate how much groundwork sits behind a single open, available online, updated monthly file. The version pulled together for this project was downloaded in 2019 and holds 16 items across 24,852,949 transactions logged in England and Wales between 1/1/1995 and 31/10/2019.
Each entry carries a unique transaction identifier alongside the transaction price, transaction date, and full address information postcode, PAON, SAON, street plus the property type, whether detached, semi-detached, terraced houses, or flats and maisonettes, and flags for newly built status and full market value sales.
Dataset Filtering and EPC Scope
Not every record qualifies for matching. Sales excluded from the linkage exercise made up just 2.90% of the whole dataset, which is a small price to pay for cleaner results. On the other side sits EPCs, required by law since 2008 for any properties sold, built, or rented, published openly by the MHCLG; the EPC dataset used here was the third version, downloaded on 20/10/2020, spanning certificates issued between 1/10/2008 and 31/5/2019 and holding 18,575,357 energy performance data records across 84 fields, capturing energy performance, building stock information, address, total floor area, and number of habitable rooms.

Granular Multi-Phase Address Linkage Method
The data linkage method borrows from an earlier published method but pushes further with tighter granularity in the matching rules. Joining PPD to the Domestic EPC dataset happens across several phases, each solving harder complex address matching challenges, and records missing postcodes 0.55% of the data get dropped first, leaving 23,999,656 transactions ready for matching. Figure 1 walks through the data linkage process, built around matching the full postal delivery address, meaning postcode plus detailed address strings, since both datasets store address structures differently and need proper data standardisation before anything lines up.
Stage 1 Linkage Rules and Variable Expansion
To make that possible, every address string in the Domestic EPCs gets capitalised and stored in new variables, powering the initial data linkage, while handling the complex subsequent linkage passes required 183 new variables in LR-PPD and 99 new variables in Domestic EPCs, all documented in Appendix A.
The full matching method runs a four-stage process built on 251 matching rules, Domestic EPCs carry a unique identifier called id, while LR-PPD uses transactionid, and Stage 1 tests whether a temporary postcode+saonpaonstreet string matches any postcode+ADDRE through the matching process algorithm.
Multi-Stage Linkage Workflow and Match Rate Performance
When records match directly, they get joined, removed from the source files, and dropped into a temporary linked data table called DATA 1; anything unmatched moves to Stage 2. Some property entries link to more than one Domestic EPC, so only those with one successfully linked EPC go straight into the final linked-EPC PPD dataset, while multi-match cases land in DATA 3, get filtered to remove entries where total floor area is NULL or zero.
Then get paired using whichever EPC inspection date or lodgement date sits closest to the transaction a process that repeats through Stages 2 to 4 until the linked-EPC PPD dataset and its data linkage result are complete, tied back through unique identifiers.
By the end, the four-stage data linkage successfully matched 16,846,834 transaction records, forming the linked dataset, and the match rate.The match rate of 56.20% in 2008 jumps quickly to 88% by 2010, which is why the evaluation of data linkage quality focuses only on the years from 2009 forward.
Technical validation
Match rates give a rough sense of how well matching worked, but they don’t tell the whole story on their own. Comparing house price frequency distributions between the linked data and the original LR-PPD data through histograms of the logarithm of transaction price, shown in Fig. 4, paints a far clearer visual comparison of matching performance.
In each graph, the distribution for linked data appears in blue laid over the original LR-PPD dataset in white, and the area covered by visible white bars marks the proportion of un-matched cases, with no significant loss of information showing up in the data linkage between 2010 and 2019.
DOMUS build/coverage detail
DOMUS is built using base data sourced from Ordnance Survey’s AddressBase Premium product, giving it a solid foundation across nearly every corner of the housing stock. It tracks attributes for every residential property in the UK, folding in EPC performance data alongside details on commercial entities such as shops, leisure establishments, and holiday lets. A full list of example attributes rounds out the picture, showing just how far this coverage stretches beyond typical housing records.
AFD Property Attribute Details Enrichment Solutions
AFD brings real expertise to linking additional datasets onto existing addresses, filling gaps that standard Property Attribute Details files often leave behind. This isn’t only about bricks and mortar the insight stretches to the people who reside at a property, covering demographic patterns, building detail, and specific insurance data in one connected view. It’s this blend of physical and human detail that makes address-linked data genuinely useful for anyone trying to understand a property beyond its four walls.
FAQS
What are property attribute details and why are they important in real estate?
Property attribute details are the core specifications of a home like square footage, bedroom count, lot size, and zoning. They help buyers, appraisers, and search algorithms accurately identify and value a listing.
How do buyers filter homes online?
Buyers use platform search filters to narrow down properties by price, location, home type, and specific features like a pool or garage.
What is the difference between a feature and an attribute?
A feature is an appealing perk (e.g., granite countertops), while an attribute is a core, measurable specification (e.g., year built or square footage).
How is listing data organized?
Real estate databases group information into standardized categories like interior specs, exterior features, utilities, and financial terms.
How do accurate property attribute details impact a home’s appraisal value?
Appraisers use verified property attribute details such as exact square footage and structural upgrades to compare nearby sales and determine a home’s final market value.
