The study
In June a customer asked us to interview everyone who had downgraded their plan in the previous year. That was a little over 4,000 people across Germany, France and Spain, in three languages, and they wanted the evidence before their planning week.
We had run studies of several hundred people before. Four thousand in five days was new.
What held
Recruiting held. Invitations went out in batches through the week, each one referring to what the person had already told the customer, and the response rate stayed close to what we see in smaller studies.
The interviews held too. Each conversation is independent of the others, so running more of them at once is mostly a question of capacity, and we had planned for it.
What broke
Processing broke. On the Tuesday, German and Spanish interviews started queueing behind French ones, and for about six hours new answers reached the topic page late. Nothing was lost: every conversation was stored and every interview finished. But evidence that should have landed within minutes took the afternoon.
The cause was dull. All three languages shared one queue, and a spike in one held up the others. They now run in separate queues with their own capacity, and we alert on delay per language rather than overall.
What we learned
Large studies do not fail where you expect. We had spent our preparation on recruiting and on the interviews, which were fine. The part that broke was the plumbing between them. We now rehearse at a tenth of the size before any study of more than a thousand people, and tell the customer what we are watching.