RESEARCH · PUBLICLY AUDITABLE

Check our working.

You shouldn't take a synthetic population on faith. We publish two separate bodies of evidence, each reproducible from raw data — one for how closely our populations match real people, one for how well they predict what those people actually did.

Two questions. Two separate proofs.

They measure different things and we never merge them. One asks whether the population answers like a real one. The other asks whether it predicts what real customers went on to do.

CREDIBILITY
Do our populations answer like real ones?
We put a synthetic population through the same national surveys real people answered, then measure how closely the two answer distributions match. Distribution accuracy, against public ground truth.
METRIC · DISTRIBUTION ACCURACY = 1 − TOTAL VARIATION DISTANCE
FIDELITY
Do we predict what customers actually did?
We built a persona from each customer's past only, sealed their real future outcome before predicting, then scored the prediction against what happened. Blind, out-of-time behavioural replay.
METRIC · DISTRIBUTIONAL DISTANCE · DRIVER DIRECTION · SEALED OUTCOMES

The part a skeptic can't argue with: we sealed the answer first.

Prediction-before-outcome is the only structure that rules out hindsight. Before any persona existed, we hashed the real outcomes and wrote the digest to the repo. The pipeline never saw them. You can rebuild the hash yourself.

  • 01 · SEALBefore a single persona was built, we hashed {customer_id: real_outcome} with SHA-256 and committed it to the repo.
  • 02 · BUILDEach persona was constructed from that customer's pre-cutoff behaviour only. The outcome window came after. No leakage.
  • 03 · PREDICTEvery persona's prediction was written to CSV, one row per customer, at natural base rates. No threshold tuned afterwards.
  • 04 · VERIFYRebuild the digest from the CSV's customer_id + actual columns and confirm it matches the sealed file. Verification is a hash check.
SEAL · TIER A · ONE PER DATASET
Grocery · Dunnhumby · n=300e8fa6f1d…
Telecom · IBM Telco · n=30087ad70df…
Subscription · KKBox · n=500fff689b7…
sha256( json.dumps({customer_id: actual}, sort_keys=True) )
→ committed before prediction · verify in an afternoon

Credibility — matched to real populations.

Distribution accuracy against Pew and IFIC national surveys, across 12 countries and 4 research domains. For every study we publish the calibrated figure and the held-out figure — the harder number, scored on questions pre-designated before calibration with zero topic anchors.

US POPULATION ACCURACY
95.3%
Pew American Trends Panel ground truth.
CALIBRATED81.9% HELD-OUT
EUROPE · 9 COUNTRIES
93.33%
Mean calibrated accuracy across 9 European populations.
MEAN CALIBRATED68.57% HELD-OUT MEAN
INDIA v2 · PROGRAM PEAK
97.61%
First replication of India's political landscape at population scale.
CALIBRATED95.87% HELD-OUT
12 COUNTRIES · 4 DOMAINS · 13 COMPLETED STUDIES · IFIC FOOD & HEALTH 96.1% CALIBRATED / 83.1% HELD-OUT
METRIC = DISTRIBUTION ACCURACY = 1 − TOTAL VARIATION DISTANCE · HELD-OUT RUN WITH ZERO TOPIC ANCHORS
Read the credibility studies → Audit repo

Every study. Every number.

Not just the headline three. All 13 completed studies, calibrated and held-out, including the countries where the held-out number is weakest. A benchmark you can only cite in your favour isn't one.

ALL 13 STUDIES · DISTRIBUTION ACCURACY vs PUBLISHED SURVEY GROUND TRUTH
StudySurveyCalibratedHeld-out
IndiaPew v297.61%95.87%
USAIFIC Food & Health96.1%83.1%
ItalyPew v295.48%63.10%
USAPew v295.3%81.9%
PolandPew v294.55%79.31%
NetherlandsPew v294.41%81.47%
UKPew v294.00%63.03%
GreecePew v293.93%69.53%
SwedenPew v293.37%69.78%
HungaryPew v291.47%55.92%
SpainPew v291.45%61.07%
FrancePew v291.33%73.96%
GermanyPew 1C91.3%76.5%
Europe · 9 meanPew v293.33%68.57%
SORTED BY CALIBRATED ACCURACY · HELD-OUT = QUESTIONS PRE-DESIGNATED BEFORE CALIBRATION, ZERO TOPIC ANCHORS
THE HELD-OUT SPREAD IS REAL — WE PUBLISH IT RATHER THAN REPORT ONLY THE CALIBRATED FIGURE
Recompute any row from the repo

How we measure accuracy.

One metric, defined the same way for every study, against published survey ground truth. No proprietary scoring you can't inspect.

THE METRIC
Distribution accuracy
How closely the synthetic population's answer distribution matches the real one, question by question. One minus the total variation distance between the two.
THE GROUND TRUTH
Published national surveys
Pew Research Center and IFIC. Real people, real fieldwork, publicly available. We score against numbers we did not produce.
THE HONEST NUMBER
Held-out, zero anchors
Questions pre-designated before calibration and answered with no topic anchors — the harder test of whether the population generalises. We report it beside every calibrated figure.
DISTRIBUTION ACCURACY = 1 TVD = 1 ½ · Σ |real sim|

Fidelity — predicting what customers did.

Blind out-of-time replay on three public datasets — grocery, telecom, and a music subscription — with outcomes sealed before prediction. On purchase-quantity distributions, the personas land roughly twice as close to reality as the naive "predict the average" baseline.

~2×
CLOSER TO REALITY THAN THE NAIVE BASELINE · WASSERSTEIN 11.5 vs 23.7
6 / 6
PURCHASE-QUANTITY SEGMENTS WON ON DISTRIBUTIONAL DISTANCE
6 / 6
BEHAVIOURAL DRIVER DIRECTIONS CORRECT · NO SIGN-FLIP
100%
OUTCOMES SHA-256 SEALED BEFORE ANY PERSONA WAS BUILT
DATASETS · DUNNHUMBY GROCERY · IBM TELCO · KKBOX SUBSCRIPTION  ·  TIER A BRIER 0.143 vs 0.148  ·  EVERY PER-ROW PREDICTION PUBLISHED
Read the fidelity benchmark → Audit repo

What we don't claim.

A benchmark you can only cite in your favour isn't a benchmark. Here is where the method stops.

  • We do not out-rank a purpose-built ML churn model. If you have twelve months of labelled behavioural data, use one. The persona's edge is the why and the intervention it points to — not raw classification.
  • Rare-event churn needs a recalibration step. On rare-event churn (KKBox), the raw predictions are over-confident and lose to the naive baseline until recalibrated.
  • Scale is moderate. 300–500 customers per dataset. These are honest back-tests, not population-scale claims.
  • Vision and shelf-choice fidelity is not yet run. No claim is made where we haven't measured.

Don't take the numbers on faith. Bring us one you can check.

Bring a decision you've already researched the hard way. We'll run it, and you compare what we return against what you already know.