Check our working.
You shouldn't take a synthetic population on faith. We publish two separate bodies of evidence, each reproducible from raw data — one for how closely our populations match real people, one for how well they predict what those people actually did.
Two questions. Two separate proofs.
They measure different things and we never merge them. One asks whether the population answers like a real one. The other asks whether it predicts what real customers went on to do.
The part a skeptic can't argue with: we sealed the answer first.
Prediction-before-outcome is the only structure that rules out hindsight. Before any persona existed, we hashed the real outcomes and wrote the digest to the repo. The pipeline never saw them. You can rebuild the hash yourself.
- 01 · SEALBefore a single persona was built, we hashed
{customer_id: real_outcome}with SHA-256 and committed it to the repo. - 02 · BUILDEach persona was constructed from that customer's pre-cutoff behaviour only. The outcome window came after. No leakage.
- 03 · PREDICTEvery persona's prediction was written to CSV, one row per customer, at natural base rates. No threshold tuned afterwards.
- 04 · VERIFYRebuild the digest from the CSV's
customer_id+actualcolumns and confirm it matches the sealed file. Verification is a hash check.
→ committed before prediction · verify in an afternoon
Credibility — matched to real populations.
Distribution accuracy against Pew and IFIC national surveys, across 12 countries and 4 research domains. For every study we publish the calibrated figure and the held-out figure — the harder number, scored on questions pre-designated before calibration with zero topic anchors.
METRIC = DISTRIBUTION ACCURACY = 1 − TOTAL VARIATION DISTANCE · HELD-OUT RUN WITH ZERO TOPIC ANCHORS
Every study. Every number.
Not just the headline three. All 13 completed studies, calibrated and held-out, including the countries where the held-out number is weakest. A benchmark you can only cite in your favour isn't one.
| Study | Survey | Calibrated | Held-out |
|---|---|---|---|
| India | Pew v2 | 97.61% | 95.87% |
| USA | IFIC Food & Health | 96.1% | 83.1% |
| Italy | Pew v2 | 95.48% | 63.10% |
| USA | Pew v2 | 95.3% | 81.9% |
| Poland | Pew v2 | 94.55% | 79.31% |
| Netherlands | Pew v2 | 94.41% | 81.47% |
| UK | Pew v2 | 94.00% | 63.03% |
| Greece | Pew v2 | 93.93% | 69.53% |
| Sweden | Pew v2 | 93.37% | 69.78% |
| Hungary | Pew v2 | 91.47% | 55.92% |
| Spain | Pew v2 | 91.45% | 61.07% |
| France | Pew v2 | 91.33% | 73.96% |
| Germany | Pew 1C | 91.3% | 76.5% |
| Europe · 9 mean | Pew v2 | 93.33% | 68.57% |
THE HELD-OUT SPREAD IS REAL — WE PUBLISH IT RATHER THAN REPORT ONLY THE CALIBRATED FIGURE
How we measure accuracy.
One metric, defined the same way for every study, against published survey ground truth. No proprietary scoring you can't inspect.
Fidelity — predicting what customers did.
Blind out-of-time replay on three public datasets — grocery, telecom, and a music subscription — with outcomes sealed before prediction. On purchase-quantity distributions, the personas land roughly twice as close to reality as the naive "predict the average" baseline.
What we don't claim.
A benchmark you can only cite in your favour isn't a benchmark. Here is where the method stops.
- We do not out-rank a purpose-built ML churn model. If you have twelve months of labelled behavioural data, use one. The persona's edge is the why and the intervention it points to — not raw classification.
- Rare-event churn needs a recalibration step. On rare-event churn (KKBox), the raw predictions are over-confident and lose to the naive baseline until recalibrated.
- Scale is moderate. 300–500 customers per dataset. These are honest back-tests, not population-scale claims.
- Vision and shelf-choice fidelity is not yet run. No claim is made where we haven't measured.
Don't take the numbers on faith. Bring us one you can check.
Bring a decision you've already researched the hard way. We'll run it, and you compare what we return against what you already know.