DDatim-QI← Back to site
OverviewDisclaimerTerms of UsePrivacy PolicyCookie PolicyMethods: predicted ratesAccessibility StatementData Protection Impact Assessment (DPIA)

Methods: how predicted rates are derived

Operated by Datim-QI Ltd · Last updated: 2 September 2026

Datim-QI shows, beside a practice’s own figure, a predicted rate: what a practice with this registered population would be expected to record. This page explains where that number comes from, what it is not, and where we have chosen not to produce one.

What is being predicted

The prediction describes recording, not health. A practice below its predicted rate is not necessarily caring for its patients less well; the most common explanation is that patients who meet the criteria are not coded as being on the register. That is why every gap in this product is framed as a prompt to check a register, never as a count of patients who have been failed.

It is an ecological, practice-level comparison. It is calculated from the characteristics of a practice’s registered population as a whole, and says nothing whatever about any individual patient. It does not screen, diagnose, risk-score or predict anything about a person, and it must not be used as though it did.

The model

For each measure we fit a regression across every practice in England with the data available, of the form:

count  ~  deprivation + % aged 65+ + % female
          + % rural + % non-white + % smoking
offset =  log(denominator)
  • Deprivation is the practice’s Index of Multiple Deprivation 2025 score, weighted by where its registered patients actually live, and used as a continuous score rather than a decile.
  • Ethnicity is the proportion of the practice population recorded in the census as other than White, and smoking is the proportion of the long-term-conditions registers recorded as current smokers.
  • The offset is the measure’s own denominator, so every coefficient acts on the rate rather than the count, and the same model form works for a register, an audit cohort and a prescribing list alike.

The model family follows what the measure is. Registers and audit measures are proportions— a count of patients out of the patients eligible — so they use a binomial fit on the log-odds, which cannot produce a rate below 0% or above 100%. Prescribing volumes are counts that have no such ceiling and use a negative binomial.

Both allow for far more variation between practices than the textbook models do. Coding practice, list turnover and case-mix all add spread, and a model that ignored it would produce ranges far too narrow — telling practices they were unusual when they were not.

The interval matters more than the point

Each prediction carries a range covering the central 80% of what comparable practices record. A single predicted number invites the reading “you are below average, therefore something is wrong”, which is not a safe inference from one figure. The interval supports the question actually worth asking: is this practice outside the range that practices with a similar population actually achieve?

We measure this on practices the model has never seen. A fifth of practices are held out of every fit, and the range is built from the remaining four fifths only. On those held-out practices, the proportion whose real figure falls inside the stated range is 81.6% for QOF registers, 80.1% for CVDPREVENT and National Diabetes Audit measures, and 85.9% for prescribing, against a nominal 80%.

How it is validated

One fifth of practices are held out of every fit and the model is scored only on those — both for how close the central prediction is and for whether the range holds, as above. Predictions are also tied to the data period they were fitted on, so a figure from an earlier year is never compared against a model of a later one.

A few measures are served instead by an older method that averages demographically matched peer practices: cervical screening, immunisations, the cholesterol cohorts, heart failure with LVSD and smoking on the long-term-conditions registers are not published in a form the model can fit. Where that applies, the page says so rather than describing the figure as modelled.

Where we do not produce a prediction

A prediction is only published where the method genuinely applies. We do not model:

  • Prescribing measures that are not counts over a registered list — shares of other prescribing, defined daily or quantity doses (DDD, ADQ), and ratios between drug classes. The model assumes a count over a denominator, and these are not that.
  • Measures with too few practices or too few events to support a six-predictor fit.

Where a measure is not modelled, no predicted rate is shown for it. We would rather show nothing than a number produced by a method that does not fit the measure.

Sources and refresh

  • QOF achievement, registers and prevalence — NHS England published QOF data
  • CVDPREVENT and the National Diabetes Audit — published practice-level results
  • Prescribing — OpenPrescribing / NHSBSA
  • Deprivation — Indices of Deprivation 2025 (MHCLG, published October 2025)
  • Population and registered-patient counts — NHS England, Patients Registered at a GP Practice

Predictions are refitted whenever the underlying published data is refreshed. Because the Indices of Deprivation 2025 use revised methods and are not comparable with the 2019 indices, a practice may sit in a different deprivation position than it previously did without anything about the practice having changed.

Limitations

  • These are descriptive comparisons of published, aggregate data. They are not case-mix adjusted in the clinical sense, and they are not a judgement of any clinician or practice.
  • The predictors available are population characteristics, not clinical need. Two practices with identical demographics can legitimately differ.
  • A figure below the predicted range is a reason to look, not a finding. Nothing here is clinical advice.

Datim-QI uses published NHS data only, no patient-identifiable data.

© 2026 Datim-QI Ltd (company number 17370894). All rights reserved.