How Tessera actually scores a building

Retrofit priority and air quality are the two readings we'd stand behind under scrutiny today. Here is exactly what feeds them, which real standards they're grounded in, and how the resulting scores actually distribute in our current pilot deployment, including the parts that are still imperfect.

Most scoring products show you a number and ask you to trust it. We'd rather show the formula. This post is the same methodology that runs in production on the live map, with the real sources, the real weights, and a real distribution, not a marketing summary of it. It's written to describe the method itself, which is designed to run the same way in any city, not just the one it happens to be piloted in.

Seismic retrofit priority

A retrofit-priority score is only useful if it reflects a city's own specific risk profile, not a generic "old building" heuristic. Ours combines two real, independently-sourced signals per building's neighbourhood:

We weight the simulation output higher (65% simulation, 35% age proxy) because it already convolves hazard and an implicit vulnerability model for that specific scenario, while the age share is a single indirect signal with no hazard component of its own. There is no published study that hands us an exact split for combining these two particular signals; 65/35 is a documented judgement call, not a fitted number. Weighting the richer signal higher is the defensible part, not the exact ratio.

Simulation weight65%
Age-survey weight35%
Storey-band bonus+5 pts, 4–8 storeys

On top of the neighbourhood-level base score, every individual building gets a per-building modifier based on its own storey count, something the first version of this score never touched. Well-documented real collapse patterns from major earthquakes concentrate specifically in 4–8 storey, pre-code, non-ductile reinforced-concrete buildings. Where a neighbourhood's own age signal already shows most of its stock predates or straddles the modern code era (age_risk_pct ≥ 50), a building in that storey band gets a flat +5 point bump. It sharpens an existing age-based signal for the height band most associated with real collapse history; it doesn't invent risk for neighbourhoods with no age signal to begin with.

Treating storey count as a real, distinct risk factor isn't something we invented either. Turkey's Law 6306 quick-assessment table (RYTEİE, EK-2) scores by storey-count band crossed with seismic zone. FEMA's P-154 Rapid Visual Screening structure is a base score by building type, modified by soil class, with storey count an explicit factor. Italy's 2018 National Risk Assessment for residential buildings (Dolce et al. 2021, Bulletin of Earthquake Engineering) goes further still, classifying its entire national exposure model by construction material × age × storey count, 56 building types in total, calibrated against a real observed-damage database of 322,728 buildings. All three, independently, treat storey count as a factor distinct from age. Ours does too.

Why fixed bands, not a quartile split

An earlier version of this score classified Low/Medium/High/Very High by quartiles of our pilot deployment's own distribution. That has a bad property for a supposedly actionable signal: a neighbourhood's classification can shift because other neighbourhoods changed, without its own risk changing at all. We moved to fixed absolute bands instead, and set the actual breakpoints by computing the real achievable range first, not by guessing. The weighted formula above can only ever reach about 0.65×30 = 19.5 points from the simulation term alone, because eq_severe_pct's own observed ceiling tends to sit well under 100% in practice. A naive 0/25/50/75-style split would put "Very High" above the 99th percentile of what the formula can actually produce. In our current pilot the real computed range is 0–56.4, median 18.8, 95th percentile 38.5, which is what the 15 / 25 / 35 breakpoints below are calibrated against.

Retrofit-priority classification in our current pilot deployment.

The honest gap: about 8% of buildings still come back "Unknown." That's inherited from upstream neighbourhood-name matching between the building-footprint source and the city's own survey data, a better matching pass would close it, not a change to the scoring formula itself. We're calling it out rather than papering over it with an invented fallback value.

Air quality

Air quality is architected around a simple priority order: a real monitoring station beats a modelled estimate, every time. Every building is matched to its nearest real, official air-quality station within a set radius. Where a building is too far from any station, we fall back to a regional reanalysis dataset: a modelled estimate, clearly a second-tier source, not presented as equivalent to a real reading.

Classification uses WHO's 2021 Global Air Quality Guidelines annual-mean thresholds for PM10: 15 µg/m³ (AQG), 30 µg/m³ (Interim Target 3), 70 µg/m³ (Interim Target 1), a real, externally published standard, not an in-house cutoff.

Air-quality classification in our current pilot deployment, by real station reading.

Two real mistakes surfaced and got fixed while building this, and we think they're worth naming rather than quietly patching:

Averaging the wrong thing. An earlier version classified each station by the mean of its hourly air-quality index readings. That's a statistical error: averaging an already-nonlinear index smooths out real pollution episodes into an artificially uniform reading. Every station came back looking implausibly similar. The fix was to classify the annual-mean concentration (PM10) against WHO's thresholds directly, and only compute an index afterward, not average an index that had already thrown the underlying variance away.

A live-data lag, not a bug in our code. Requests for recent dates (anchored to "now") reliably came back with null concentration values from the city's own monitoring API, even after long retries: a real ingestion lag on their side, confirmed by testing the same request against a fixed calendar-year window instead. We now query a fixed reference year rather than a rolling "most recent" window, and chunk requests after finding that longer single requests intermittently fail.

What we deliberately left out

In our current pilot, we evaluated flood risk using a widely-used global flood hazard dataset, the same one most climate-risk products lean on for this layer, and found it returns no data at all over that city. The reason was structural, not a download error: that model only covers river catchments above a fairly large size threshold, and the city's real flood exposure is coastal and small-stream flooding on much smaller catchments, a different hazard type entirely. Rather than ship an empty or misleading layer, we removed flood risk from the product until we have a dataset that actually models the hazard that city has. This is exactly the kind of check we run before adding any new layer for any city: if the best available dataset doesn't model the real hazard, we leave it out rather than ship something misleading.

Everything above runs today on the live map. If you want to know how a specific building or neighbourhood scores, or how this generalises to your city, get in touch.