{"id":4874,"date":"2026-07-31T12:05:23","date_gmt":"2026-07-31T16:05:23","guid":{"rendered":"https:\/\/paulsyng.com\/blog\/?p=4874"},"modified":"2026-07-31T12:05:26","modified_gmt":"2026-07-31T16:05:26","slug":"similes-2-billion-raise-the-cern-of-saying-things","status":"publish","type":"post","link":"https:\/\/paulsyng.com\/blog\/similes-2-billion-raise-the-cern-of-saying-things\/","title":{"rendered":"Simile&#8217;s $2 billion raise: The CERN of Saying Things"},"content":{"rendered":"\n<p class=\"has-medium-font-size wp-block-paragraph\"><em>TLDR: I read the study behind Simile, the AI startup that just raised $200 million to simulate human beings. The founders interviewed 1,000 people for 2 hours each and built AI agents that fill out surveys the way those people would. Their own paper grades the copies against survey answers and says so, carefully. The paper also ran tests with real money on the table, and there the expensive interviews stopped helping. Then I listened to the founder on a podcast and watched &#8220;responses&#8221; become &#8220;behaviours,&#8221; a life story become &#8220;behavioural data,&#8221; and a poll panel become &#8220;ground truth.&#8221; Investors priced the swapped word at $2 billion. 85% of yourself, 2 weeks later, is still just you talking.<\/em><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Yesterday a company that manufactures fake survey answers became worth $2 billion. I spent the night with its paper trail, and I want to walk you through what I found, because the whole valuation sits on one swapped word.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Simile announced a $200 million round on July 30, 5 months after raising $100 million. Its stated mission: &#8220;simulate all eight billion people on earth, accurately and honestly.&#8221;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Accurate compared to what? <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">I went looking for the answer. The founder gave it once, in a paper, before the money arrived. I read that 2024 study end to end. I listened to the founder&#8217;s full Sequoia interview. I read the Wall Street Journal&#8217;s profile of his flagship customer. Joon Sung Park led the famous generative agents work at Stanford, and his 2 PhD advisors are his co-founders. This is the strongest science this category has ever produced. That is exactly why what he did with one noun matters.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">What the machine was graded on<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">I&#8217;ll give the engineering its due first, because the study is real work. The team recruited 1,052 people and interviewed each one for 2 hours about their life. Each transcript became the seed of an AI agent. The agents then took the General Social Survey, a big standard opinion poll, and the paper reports the result in one careful sentence: the agents matched people&#8217;s answers 85% as well as those same people match their own answers when asked again 2 weeks later.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Read that second half again, because it carries the trick. People are messy. Ask me the same opinion questions 2 weeks apart, and I will disagree with myself about 1 time in 5. The panel agreed with itself 81% of the time. The machine reached 85% of that self-agreement. On opinion questions.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Then I got to the rest of the results table, and this is where I stopped taking notes and started underlining. The study ran 3 kinds of tests, and the scores fall in a straight line as the tests move from talking toward doing. Opinion survey: 85%. Personality quiz: 80%. Money games, where people split real cash with strangers and decide who to trust: 66%. The cash was real. The paper paid bonuses of up to $10 based on the choices people made.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">And on those money games, the agents built from 2-hour interviews scored no better than agents built from 5 dry facts about a person. The paper&#8217;s own sample descriptor reads like a voter file: &#8220;Ideologically, I describe myself as conservative. Politically, I am a strong Republican. Racially, I am white. I am male. In terms of age, I am 50 years old.&#8221;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Sit with what that means. 2 hours of someone&#8217;s life story bought a 14-point edge when the job was guessing their answers. It bought nothing when the job involved a choice with real stakes. The expensive part of the product stops working at the exact moment behaviour starts.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">I keep picturing a bathroom scale that works by guessing the number I would say out loud. It gets very good. It guesses my number almost as well as I repeat it myself. At no point has it weighed me. That is this machine, in reality.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">So here is the full published claim, stated the way the paper actually supports it: a 2-hour interview lets a model repeat your survey answers almost as well as you repeat them yourself, and the trick fades as the test gets closer to real doing. At the money line it fades to zero.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">I believe that claim. It is true, useful, and worth a small fraction of $2 billion. Which is why nobody funding the company says it.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">One word paid for the round<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Here is Park on the Sequoia podcast, describing his own study. I have listened to this sentence maybe 10 times:<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\">&#8220;We demonstrated that using our architecture and the models we can actually predict people&#8217;s\u00a0<strong>behaviors<\/strong> 85% as accurately as people replicate their own.&#8221;<\/p>\n<\/blockquote>\n\n\n\n<p class=\"wp-block-paragraph\">The paper says responses. The founder says behaviours. Same number, same study, one noun swapped. And I know the paper has a behaviours number. It is 66, and it is the number where his method&#8217;s edge disappears.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In June I <a href=\"https:\/\/paulsyng.com\/blog\/ninety-percent-of-what-exactly-on-the-difference-between-copying-a-survey-and-predicting-a-sale\/\" data-type=\"post\" data-id=\"4489\">wrote<\/a> about a different paper in this same category and traced how its careful claim decayed across 3 steps. The paper said one thing. The vendor blog said something louder. The viral post said something false. Each step dropped one small word, and each dropped word was carrying the roof.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This time I watched the whole chain collapse into 1 step. The founder does the compression himself, in a single sentence, about his own experiment. Then his next sentence spends the swapped word: &#8220;When we saw that, we thought, okay, this is something that we feel comfortable providing to our users as a platform for simulating their really important decisions.&#8221; <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">He earned the comfort on quiz answers. <br>He is spending it on decisions.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Rename the talk<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">The interviewer asked a fair question: why collect any real data at all, when the model has already read everything people write? Park&#8217;s answer starts with the most honest thing I heard in 38 minutes:<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\">&#8220;There are things that people say and then there are things that people actually do. And the gap there is real.&#8221;<\/p>\n<\/blockquote>\n\n\n\n<p class=\"wp-block-paragraph\">He is right. That is 50 years of research in 2 sentences. What people say they will do explains about 28% of what they later do. Ask people what they would pay, and they claim 2 to 3 times what they hand over when the money is real. That gap is the entire problem with surveys. I have built my whole <a href=\"https:\/\/monopoly.ceo\" target=\"_blank\" rel=\"noopener\">practice<\/a> on it.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Now watch him close it:<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\">&#8220;A lot of the data that we end up collecting by nature are&nbsp;<strong>behavioral<\/strong>. It also includes data that actually goes into literally questions like&nbsp;<strong>just tell me the story of your life<\/strong>.&#8221;<\/p>\n<\/blockquote>\n\n\n\n<p class=\"wp-block-paragraph\">I had to rewind to make sure I heard it. The behavioural data is an interview. A life story, told out loud to a stranger with a recorder, is the most rehearsed performance most people ever give. When I bought my place, the seller would have talked warmly about it for as long as I let him. I hired an inspector to open the panel and check the wiring, because my money was on the line. Simile&#8217;s inspection is 2 more hours with the seller.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Calling a life story &#8220;behavioural data&#8221; does no work on the say-do gap. It renames the gap and declares it crossed. And I expect to see this move everywhere now, because the case against surveys got famous enough that vendors had to answer it. The answer they landed on was to relabel the talk. The survey industry ate the argument against surveys by renaming the survey.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">I found exactly one real behavioural asset in the interview. Park says the company holds a pile of published experiments with real choices in them, and wants to train a model that &#8220;can basically predict the results of any RCTs.&#8221; Those experiments have real stakes. Predicting their results before they run would be the real thing, and I would take that product seriously. He describes it as a plan. No accuracy reported. Nothing published. The one dataset that could settle the question lives in a sentence about the roadmap.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Meanwhile, the data pipeline runs through a partnership with Gallup. A polling company. The tool on trial supplies the evidence. And asking leaves fingerprints: people who get surveyed go on to show a stronger link between what they said and what they buy than identical people nobody asked. The question plants part of the answer. Simile&#8217;s &#8220;ground truth&#8221; now comes from panels of the professionally asked.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">The seller grades his own test<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">I wanted to know how Simile decides a simulation is good enough to guide a real decision. Park answered directly:<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\">&#8220;We actually measure total variation distance, which basically shows how close are the distributions of the&nbsp;<strong>ground truth<\/strong>&nbsp;versus the simulated information&#8230; TVD of let&#8217;s say less than 0.15, we believe is actually quite strong evidence for making decision.&#8221;<\/p>\n<\/blockquote>\n\n\n\n<p class=\"wp-block-paragraph\">Two things hide in there. &#8220;Ground truth&#8221; means the panel&#8217;s answers. Every score this company reports measures how far the fake talk sits from the real talk. I searched the interview, the paper, and the website for a purchase, a renewal, a receipt anywhere in the grading. I found none. Reality has never entered the room.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">And the passing grade of 0.15 comes from &#8220;we believe.&#8221; Park draws the comparison himself: science settled on its famous cutoff for what counts as proof, and &#8220;Simile is working on setting the same kind of threshold and standards for the rest of the field.&#8221; Science earned its cutoff through 100 years of researchers fighting each other in public. This one is being written by the company that charges you to pass it. When the seller writes the test, the test&#8217;s first job is to be passable.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">The buyer grades it the same way<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">In March, the Wall Street Journal profiled the flagship customer, and the details are better than anything I could argue.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">CVS built its twins from 2.9 million consented responses from over 400,000 people, tuned them with its own historical surveys and support logs, then tested them. The result, per the Journal: the twins &#8220;replicated known findings with up to 95% accuracy.&#8221; Read those words slowly, the way I had to. The twins studied CVS&#8217;s old answers. Then CVS graded them on reproducing what CVS already knew. A copy machine, scored on copying the documents it was fed, comes back with a strong copying score. And &#8220;up to&#8221; means 95 was the best day. Nobody reports their floor.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The CVS insights lead, Sri Narasimhan, gave the Journal the honest version of the appeal: &#8220;It&#8217;s not like I have to stop with how many questions I asked. There&#8217;s no fatigue.&#8221; He is right. There is nobody there to get tired. He also named the budget consequence out loud: &#8220;If we&#8217;re going to invest in the agentic twins, we don&#8217;t need as much of that traditional market panel research.&#8221; So fake talk now replaces real talk, at a price the Journal put at $150,000 to millions per customer per year.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">And what do customers ask the twins? Park says the lead use case is concept testing. New product, new message, judged on a screen before it exists in the world. Companies testing 5 to 10 ideas a month now dream of testing &#8220;thousands of different ideas across thousands of different sub populations.&#8221; <\/p>\n\n\n\n<p class=\"alignwide has-source-serif-pro-font-family has-x-large-font-size has-custom-font wp-block-paragraph\" style=\"font-family:source-serif-pro\"><em><strong>I checked this against the oldest record in my files. In the biggest analysis of purchase intent ever run, the link between what people said about brand-new products judged one at a time and what they later bought was zero. <\/strong><\/em><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The machine speaks fluent survey in the one setting where the survey never spoke for the buyer. Some customers also ask Simile to simulate their own earnings calls. Executives rehearse a performance for a fake audience. I&#8217;ll grant that use case one thing: it is honest about what it is.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">One claim in the Journal piece deserves a straight answer, because it is the counterpoint a sharp reader will raise. Park says the twins also take in participants&#8217; behavioural and purchase data. Fine. Feed in whatever you want. I care about the grading, and the grading is still agreement with responses. A student can study from receipts all semester. If the exam only asks what people would say, the grade only tells you about the saying. Nobody at Simile or CVS has published the twins predicting a purchase before it happened.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">I&#8217;ll also give Narasimhan his due. He told the Journal CVS keeps testing the twins against real humans and will &#8220;never stop talking to real customers.&#8221; Then look at what the safety net is made of. The check on the simulated talk is more talk. The company that owns purchase records for about 90 million loyalty members, and sells that data to advertisers, checks its fake customers by asking real ones more questions.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Trust it where it&#8217;s cheap, doubt it where it&#8217;s sold<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">The part of the interview I found most revealing is Park&#8217;s own sorting of simulations into 2 kinds, because he gives the case away and keeps going.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Some simulations converge. <\/strong>Small errors wash out because the ending is baked in. His example: simulate any network of people and a few hubs always form. True. It is also the oldest lesson in this field, and it points the wrong way for him. Decades ago, Thomas Schelling produced the pattern of segregated neighbourhoods using red and blue dots on a grid, checking their neighbours. Dots. The lesson everyone took from that work: when the ending is baked in, the agents can be stupid. If the answer comes out the same no matter how rich the agents are, the 2-hour interviews added nothing to that answer. Dots would have done it.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Other simulations diverge. <\/strong>Small errors snowball. I made a list of which questions those are. Bank runs. Elections. The ripple effects of a product launch across a market, the exact things Park pitches to customers. And there, the math of when you can trust the output is, in his own words, &#8220;a real research frontier.&#8221;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Put those halves side by side. Where the simulation can be trusted, the expensive detail is decoration. Where the expensive detail would matter, the trust does not exist yet. The zone where it works and the zone where it sells do not overlap, and Park described both zones within 2 minutes of each other.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">He also said, in passing, the sentence I would least want on tape if I held a $2 billion valuation: &#8220;we have sort of plateaued with current modelling paradigm, our ability to really simulate humans.&#8221; The science stalled. The product ships anyway.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">The one study they won&#8217;t run<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">The interviewer got to the real question once, and I sat up when she asked it. Why simulate at all? Why should a company pay for fake customers when it could run 1,000 cheap ads and count real clicks from real people?<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The answer was scale. Along the way, simulated behaviour got renamed &#8220;actual behaviour simulation at scale.&#8221; More populations, more questions, more speed. She asked about truth. He answered with volume. Then the conversation moved on, and I sat there waiting for it to come back. It never did.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Here is what makes the silence expensive. The fix costs 1 study. Simile holds a pile of real experiments with real outcomes. Its flagship customer holds purchase records for about 90 million loyalty members. Take product ideas that launched. Hide the results. Show the simulation predicting what people bought better than old-fashioned surveys do. I cannot name a company better placed to run that study. It does not exist in public. The mission statement says &#8220;accurately and honestly.&#8221; Honestly has a price, and the price is that study.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The interview closes with the big image instead. Park wants to build &#8220;the CERN of human society,&#8221; and his co-founder likes to say that great science starts with a great measuring tool, like the Hubble telescope. I think the comparison convicts the product. Hubble pointed at the sky. It collected light that did not care what any astronomer hoped to see. That indifference is what made it a measurement. Simile points at a model of the sky, stitched together from what people said when someone asked them. CERN smashes particles. The particles cannot flatter you.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Reality<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Split your evidence into 2 piles. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Pile 1 is what people said: <\/strong>surveys, interviews, focus groups, intent scores, life stories, and now simulated versions of all of it, at machine speed, for pennies. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Pile 2 is what people did: <\/strong>renewals, repeat orders, full price paid with no discount, referrals that closed, the angry email nobody asked for. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Pile 1 just became free and infinite. Pile 2 cannot be simulated, because it is the only pile connected to reality. The louder pile 1 gets, the more pile 2 is worth. I would spend the week moving budget from pile 1 to pile 2, and I would ask every research vendor one question: show me a purchase you predicted before it happened.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">When a founder tells you his machine is &#8220;as human as possible,&#8221; ask which human he means. The one talking, or the one paying. This machine got the talking one. 85%, 80%, 66%. A pretend answer at any accuracy is still a pretend answer. The only honest thing in the story is a receipt.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>TLDR: I read the study behind Simile, the AI startup that just raised $200 million to simulate human beings. The founders interviewed 1,000 people for 2 hours each and built AI agents that fill out surveys the way those people would. Their own paper grades the copies against survey answers and says so, carefully. The [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":4875,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_coblocks_attr":"","_coblocks_dimensions":"","_coblocks_responsive_height":"","_coblocks_accordion_ie_support":"","footnotes":"","rank_math_title":"","rank_math_description":"","rank_math_canonical_url":"","rank_math_focus_keyword":""},"categories":[85,82],"tags":[],"class_list":["post-4874","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-monopoly","category-autopsy"],"_links":{"self":[{"href":"https:\/\/paulsyng.com\/blog\/wp-json\/wp\/v2\/posts\/4874","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/paulsyng.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/paulsyng.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/paulsyng.com\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/paulsyng.com\/blog\/wp-json\/wp\/v2\/comments?post=4874"}],"version-history":[{"count":5,"href":"https:\/\/paulsyng.com\/blog\/wp-json\/wp\/v2\/posts\/4874\/revisions"}],"predecessor-version":[{"id":4880,"href":"https:\/\/paulsyng.com\/blog\/wp-json\/wp\/v2\/posts\/4874\/revisions\/4880"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/paulsyng.com\/blog\/wp-json\/wp\/v2\/media\/4875"}],"wp:attachment":[{"href":"https:\/\/paulsyng.com\/blog\/wp-json\/wp\/v2\/media?parent=4874"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/paulsyng.com\/blog\/wp-json\/wp\/v2\/categories?post=4874"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/paulsyng.com\/blog\/wp-json\/wp\/v2\/tags?post=4874"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}