Skip to main content
Reference

What the score means

Every read comes back with one number. This page says exactly what it is, how to read it, and how it was arrived at, in that order, and without asking you to take any of it on faith.

It is a probability, not a grade

The score is the model's estimate of one thing: the chance that this concept, at this address, is still open 4 years after it opens. A 62 means that of the concepts the model scores like yours, about 62 in 100 were still trading 4 years later.

That matters because a number out of 100 invites a reading that is simply wrong. There is no scale here where 100 is perfect and 50 is a pass. An average concept in an average location scores 57. That is the line to read yours against. Above it the record is in your favour, below it against.

Scores also sit in a much narrower band than a mark out of 100 suggests. Across every address and concept the model can score in the city, 8 in 10 come out between 52 and 70, and the middle one is 62. So 62 is an ordinary result rather than a poor one, and 70 is better than nine in ten of them. A few points therefore mean more than they look. Going from 62 to 64 is 2 points, and it moves a concept from the middle of the city to its top quarter.

Every score the model produces
25
30
35
40
45
50
55
60
65
70
75
80
85

Each bar is five points wide, over all 10,800 address-and-concept combinations the model can score in Chicago. The darker bars are the 50s and 60s, where most of the city sits; an average concept scores 57. Nothing scores below 28 or above 85, which is why the number cannot be read as a mark out of 100.

There is no composite behind it and no weighting scheme. The score is the model's calibrated probability. A hand-weighted scoring scheme was built, measured against the model, and removed for being materially worse at telling survivors from closures.

What it can actually tell you apart

In the years held out of training, the concepts this model scored in its best 10% went on to survive 89% of the time. The ones it scored in its worst 10% survived 48% of the time. Separating those two groups is what the score is for.

Spread it separates
41 points
Between what the model's best-scored and worst-scored locations actually went on to do. That gap is the whole value of the number.
On any single address
Directional
It ranks and compares well. It does not support a confident verdict on one location, and your read shows the range the evidence actually supports around your number.

Your own read carries that range explicitly, not just the estimate but what comparable concepts actually did in the years held out of training. Where the range is wide, the honest answer is that the record does not know, and the read says so rather than projecting confidence it has not earned.

How it is calculated

It is learned from Chicago's own licence record since 2002: every restaurant the city ever licensed, rebuilt into attempts, and kept only once its 4-year outcome is actually observable. That is 15,173 attempts with a known ending.

  1. 01Start from an average concept. Every score starts at 57, what the model predicts for an average concept in an average location, in the current era.
  2. 02Adjust for what is specific to you. Five drivers move it up or down: how this exact cuisine-and-format has fared, how many restaurants the operator has opened before, how often restaurants in the zip turn over, how crowded the zip already is with this cuisine, and how that cuisine specifically has fared there.
  3. 03The adjustments are the arithmetic, not a summary of it. Each driver's contribution is measured in log-odds against that same average concept, which is what makes them add up exactly. The rows you see sum to the score, and there is no hidden term and no rounding fudge beyond a single point.
  4. 04Calibrate so the number means what it says. Ordering a list correctly and stating a believable probability are different jobs. The second is checked separately against what actually happened in the held-out years, and a retrain that cannot show it ships without a calibrator rather than with a flattering one.

You can check that last claim on your own read. Every read prints the anchor, each driver's signed contribution and the measurement behind it, and the total. The rows add up to the score, and where they would not, the model refuses to guess rather than quietly absorbing the difference.

Two further terms sit in the model and never appear in your breakdown: an opening-year control and a COVID control. They exist so a 2004 opening is compared fairly against a 2024 one, and at prediction time both are pinned to the same value for every score, so they carry no information about why your number differs from anyone else's, and listing them as a “reason” would be misleading. The methodology page states what that pinning assumes and how much it moves the answer.

How local it actually is

The location half of the score is measured at zip level. Two addresses in the same zip, scored for the same concept by the same operator, come back with the same number, and with the same ranking of concepts. Not close: the same.

That is a finding, not a shortcut. Finer location data (household income, jobs density, walkability, traffic counts, distance to transit, assessed value, competitor counts inside a quarter-mile and a half-mile) was tested against survival, repeatedly, and none of it improved accuracy on restaurants the model had never seen. Zip is the resolution the evidence supports, so it is the resolution we claim. Advertising block-level precision would mean advertising an accuracy that was measured and not found.

The address still decides most of what you read: that property's own licence and closure history, which restaurants are shown as comparable, and the write-up around the number. And your own track record moves the score further than the block does: an operator's prior openings swing it more than any location signal in the model.

What it is not

It is not a grade out of 100.
There is no scale where 100 is perfect and 50 is a pass. 57 is average, and the usable range really is narrow: 8 in 10 scores fall between 52 and 70, so most of the 0-100 the number is written in never appears at all.
It is not a prediction about your business.
It describes how comparable concepts at comparable addresses have performed. Nothing in it can see your operating skill, your menu, your lease terms, or your capital.
It is not a verdict on a neighbourhood.
Location contributes very little of the score. What it mostly reads is the concept and the operator's record. The fairness audit on the methodology page measures this directly.
It is not evidence of cause.
Operator experience is the clearest example: it is the strongest single driver and it is a correlation. Opening more restaurants does not, by itself, make the next one survive.

Where to check us

This page describes the score. It does not try to prove it. That is a different document, written to be audited rather than read.

  • Methodology covers the validation, the ablations, the fairness audit, and the features that were tested and rejected. Most of the statistically significant findings there are negative results.
  • A sample read is a real address, scored live against the current record, with the full arithmetic shown.
  • Data sources lists what the record is built from and how current each part of it is.

This score is a statistical estimate built from Chicago's own licence and inspection record since 2002. It describes how comparable concepts at comparable addresses have performed: a historical association, not a prediction about your business and not a guarantee. It will be wrong on a meaningful share of individual addresses, it is validated for Chicago only, and a licence can look active for months after a restaurant has actually closed. Drivers such as operator experience are correlations, not causes. Treat it as one input among others, alongside your own diligence.

Not for adverse decisions: a score must not be used as a basis for denying credit, a lease, or a similar application about a person or business. It has not been validated for that use.