← Projects

Property-level changes in BBR

An exploration of how object-level histories in Denmark's building and property register can be efficiently transformed into reusable property-level histories.

Background

In previous work with BBR data, I helped develop point-in-time methods for resolving the property identifiers associated with register objects. Doing this correctly and completely requires more than following foreign keys: the query has to combine temporal filtering, status interpretation across related objects and information from the surrounding building case.

The property-level change algorithm described below is an independent extension of that work. It explores how the same point-in-time resolution could be used to precompute reusable histories of property composition and property-level changes.

The data model

BBR is Denmark's national register of buildings and dwellings. It contains several kinds of versioned objects, including buildings, units, technical installations, land and properties.

Object versions contain both a registration timeline and an effect timeline, together with a status. This makes it possible to distinguish between when information was registered and when it was considered to apply.

For example, the register can in principle answer a question such as: “On December 25, 2025, what did the register say about the size of this building on January 7, 2020?”

Many analyses require only one of these timelines. To reconstruct the effective history of an object, we can select versions whose effect interval contains a given timestamp and whose status is relevant to the analysis. A version is typically active on the half-open interval from effectFrom, inclusive, to effectTo, exclusive.

Once versions have been placed on a chosen timeline, it is conceptually straightforward to detect changes to one object. We compare each version with its preceding version and record which attributes changed. I refer to the resulting representation as an object change table.

From objects to properties

Object-level changes are not always the most useful unit of analysis. Buildings, units, installations and land are related to properties, and an analysis may instead need to describe how an entire property changed.

Resolving this relationship is not always a direct lookup. An object may be connected to a property through a bounded chain of relationships. For example, an installation may be attached to a unit, the unit to a building, and the building to a piece of land associated with a property.

The relationship is more accurately modelled as many-to-many. In practice, however, the number of property identifiers associated with one object is small. For the following analysis, I therefore assume a fixed upper bound on the number of relevant property identifiers associated with any individual object version.

The naive solution

A direct approach would be to reconstruct every property at every timestamp found anywhere in the database. At each timestamp, we would resolve all object-to-property relationships, construct every property snapshot and compare it with the preceding snapshot.

This performs an enormous amount of redundant work. Most objects do not change at most timestamps, and most properties are unaffected by any particular change. Evaluating the entire register at every recorded timestamp would therefore be computationally impractical.

The sparse-event observation

A property snapshot can only change when one of its contributing objects changes, or when a relationship determining which objects contribute to it changes. It is therefore unnecessary to evaluate properties at timestamps where none of their contributing data or relationships changed.

Instead of starting with every timestamp and searching for changes, we can start with the changes themselves:

  1. Create ordered change tables for every versioned object and relationship type.
  2. For each object or relationship change event, resolve the affected property identifiers immediately before and after the event.
  3. Add those properties to the set of properties that must be evaluated at that timestamp.
  4. Reconstruct only the affected properties before and after the event.
  5. Compare the two property snapshots and store the resulting snapshots, differences or both.

Recording both the previous and new property identifiers is important because a relationship change may remove an object from one property and add it to another. In that case, both properties must be evaluated.

Resolving affected properties

For each changed object, its property identifiers can be resolved using a bounded number of relationship traversals. The maximum traversal depth is determined by the data model rather than by the size of the database.

The traversal itself is only part of the problem. Each object and relationship must also be evaluated on the selected timeline and with an appropriate set of statuses. A relation present in the source data is not automatically relevant to the requested point-in-time view.

Statuses also cannot always be interpreted independently for each object. A unit may, for example, already have a status indicating that it is in use while its containing building is still registered as under construction. This can occur in large developments where parts of a building are occupied before the entire building has been completed and approved.

Excluding the building solely because of its own status could therefore also exclude active units within it. Correct resolution may require consulting related building-case information to interpret the status of the surrounding building in context.

The property lookup is consequently a domain-specific point-in-time query: it traverses a bounded relationship graph, but it must also account for temporal versions, status rules spanning related objects and relevant building-case information.

With appropriate indexes, the individual lookup operations can still be performed efficiently. Let V denote the total number of relevant object and relationship versions. Resolving the affected properties for C change events requires approximatelyO(C log V) indexed work under the assumption that traversal depth and the number of candidate property identifiers remain bounded.

Complexity

Constructing all object and relationship change tables requires linear time if the input is already grouped and ordered by identifier and timeline. If sorting is needed, this stage instead requires O(V log V) time.

For every detected change, the algorithm resolves the initially affected properties through a bounded relationship traversal, then evaluates all relevant objects belonging to those properties before and after the event.

In the unrestricted worst case, one property could contain an arbitrarily large part of the database. Reconstructing that property repeatedly would then make the approach no better than the naive solution.

The useful bound therefore depends on a domain assumption: the number of objects contributing to one property, and the number of properties associated with one object, are both reasonably bounded. Under these domain assumptions, and assuming indexed lookups, the preprocessing is expected to scale approximately as O(V log V).

This does not need to be an interactive operation. It can be run periodically to create a derived property history that analysts can subsequently query quickly and consistently.

Property snapshots and change records

Before running the preprocessing, an analyst must define what a property snapshot should contain. This definition is implemented as code and may include relational information, such as the number and types of objects belonging to the property, as well as aggregates derived from their values.

Possible snapshot fields include counts grouped by object type or usage code, distributions of heating types and several measures of total area. The appropriate fields depend on the analysis, since the register contains multiple notions of area and many object-specific attributes.

The algorithm evaluates this snapshot definition immediately before and after each relevant property event. Once the snapshots have been constructed, producing a change record is comparatively simple: the two rows can be compared field by field to identify changed aggregates, added or removed objects and other differences selected by the analyst.

The resulting data could be stored in several ways. Snapshots and change records could be kept in separate tables, the differences could be stored alongside each new snapshot, or only the snapshots could be stored and changes derived with window functions when needed. The best representation depends on expected query patterns and storage requirements.

The important result is a directly usable property history. Analysts still have to define and implement which property-level information matters, but they do not have to implement the underlying temporal filtering, relationship traversal and affected-property detection themselves. For subsequent analysis, the result can largely be treated as an ordinary table of property states and changes.

Possible uses

Property-level histories move the unit of analysis closer to the level at which ownership, reporting and administrative decisions often occur. They could support analyses of reporting patterns, data-quality initiatives, changes preceding property transactions or differences between categories of properties.

The interpretation still requires care. Updates are not necessarily made by property owners, and a property may have multiple owners. A property-level history therefore describes changes in the register; it does not by itself identify who caused them or why they occurred.

A particularly useful derived dataset would be a complete history of which objects were associated with each property over time. This would let analysts query historical property composition directly instead of repeatedly resolving the full point-in-time relationship graph.

Producing this history still requires evaluating the complete affected property whenever a relevant object or relationship changes. A single change may alter the derived property identifiers of several connected objects, so all relevant objects in the property may need to be traversed and recorded again. The preprocessing does not eliminate that work, but it performs it once and stores the result in a form that can be reused across many later analyses.