
ISSUE 001
Traction Studio Issue 001
You can now buy software that will interview a hundred imaginary customers about your idea and hand you a report by lunch. I wanted to know whether that works, so this one goes into the actual research on synthetic customers — including a dataset where you can open the file and see the machine's guesses sitting next to what real people said. The answer is more interesting than either yes or no.
An idea worth thinking about
Most studies of synthetic customers never ask who the synthetic customer represents
A team reviewed 63 studies using synthetic personas, drawn from the leading AI research venues, and checked how each one handled representativeness. Forty-three percent aimed at an undifferentiated "general population." Only 35 percent explicitly discussed whether their personas represented anyone in particular. The paper's conclusion is measured — that diversity and validity vary considerably across the literature, and that the field needs shared standards for building these datasets.
Sit with what that means downstream. If the research producing synthetic personas often doesn't establish who they stand in for, then a commercial tool built on the same techniques has no better claim, and probably a worse one, because it has a product to sell you. The output will be fluent either way. Fluency is not the same as representativeness, and it's much easier to notice.
There's a second gap I'd flag as my own observation rather than the paper's. Nearly all of this work measures whether synthetic answers resemble human answers. Very little of it measures whether a business that used synthetic research made better decisions. Those are different questions, and only one of them is yours.
Read the persona review
A tool I disappeared into
2,058 real people, 500 questions each, and the AI's guesses in the next column
A Columbia team built Twin-2K-500 to give this whole argument some ground truth: a representative US sample of 2,058 people, surveyed across four waves, more than 500 questions each, averaging about two and a half hours per person. Demographics, psychological scales, economic preferences, cognitive measures, plus replications of classic behavioral economics experiments. The final wave repeats earlier tasks on purpose, so you can measure how consistent a real person is with themselves — which is the only fair benchmark for judging a machine imitating them.
It's free on Hugging Face, and the companion Mega-Study release is the part I lost time in: 19 studies with response files containing both the human answers and the model's answers for the same items. You can open a CSV and look directly at where the simulation went wrong. There's no interpretation layer between you and the disagreement.
Open the dataset
Learn something
The course that coined "get out of the building" is free and always has been
Steve Blank's How to Build a Startup sits on Udacity at no cost, taught with Kathleen Mullaney, fourteen lessons, self-paced, last updated in November 2025. It's the introduction to customer development — value proposition, customer segments, channels, revenue model, market types — and its organizing instruction is that you cannot learn what people want from inside your own office. No programming involved.
One thing to set expectations on: there's no exam, no final project, and no certificate at the end. If you need a credential to show someone, this isn't it. If you want the method that most of the modern validation vocabulary descends from, it's the source, and it costs an afternoon.
Start the course
What we’re seeing
Simulated customers are most accurate about the affluent, the educated, and the politically moderate
Start with the part that favors the tools, because it's real. Researchers working with the Twin-2K-500 data have found that synthetic respondents do a reasonable job of reproducing how a population answers in aggregate. If you want a rough read on which of four taglines tests better across a broad market, this technology is genuinely useful and enormously cheaper than a survey panel.
Now the part that matters more for anyone building something specific. The same validation work finds these twins unreliable at the individual level — they don't dependably predict what a particular person will say, and they distort sample means and variance. And accuracy isn't spread evenly: it runs highest among respondents who are more educated, higher-income, and ideologically moderate.
Read that last sentence as a founder rather than a researcher. The simulated customer knows most about the demographic already best represented in the training data, and least about everyone else — which, if you're building for a small town, a trade, an aging population, or anyone stretched thin, is your actual market. The tool is sharpest precisely where you need it least.
So the discipline isn't "don't use it." It's knowing which question you're asking. Aggregate direction, breadth, first-pass screening: fine. Whether this specific person will pay you money: go find out. That still requires a conversation with somebody who can say no to your face.
Toubia et al., Twin-2K-500, Marketing Science 2025
One question
When did a real customer last tell you something you didn't want to hear?
Reply with roughly when, and what it was. The interesting answers are the ones that changed what got built.
More in this pillar as it's written.
— Shana
Sources link to the original publisher, never to a summary.