{"id":4489,"date":"2026-06-12T21:45:30","date_gmt":"2026-06-13T01:45:30","guid":{"rendered":"https:\/\/paulsyng.com\/blog\/?p=4489"},"modified":"2026-06-13T00:11:41","modified_gmt":"2026-06-13T04:11:41","slug":"ninety-percent-of-what-exactly-on-the-difference-between-copying-a-survey-and-predicting-a-sale","status":"publish","type":"post","link":"https:\/\/paulsyng.com\/blog\/ninety-percent-of-what-exactly-on-the-difference-between-copying-a-survey-and-predicting-a-sale\/","title":{"rendered":"Ninety Percent of What, Exactly \u2014 On the difference between copying a survey and predicting a sale"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\"><strong>TLDR:<\/strong> Scientists taught a robot to fill out surveys like a human. The internet heard &#8220;robots can predict what you&#8217;ll buy!&#8221; and stampeded. Nobody checked what anyone actually bought. People in surveys say things they never do. The robot copied the saying part. The herd sharing the headline never read the paper, so they copied a saying too. Everybody in this story is performing: the shoppers, the robot, the reposters. <strong>Ninety percent of a pretend answer is still a pretend answer. <\/strong>The only honest thing in the whole story is a receipt. <\/p>\n\n\n\n<figure class=\"wp-block-audio\"><audio controls src=\"https:\/\/paulsyng.com\/blog\/wp-content\/uploads\/2026\/06\/The_90_percent_AI_accuracy_myth.mp3\"><\/audio><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>The paper everyone is citing as proof that AI can predict what you&#8217;ll buy never checked what anyone bought.<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Everyone is scared of AI slop. This is what human slop looks like. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Predictable. <br>Performative. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Now read the rest.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In October 2025, researchers at PyMC Labs and Colgate-Palmolive published a preprint with a headline result that travelled fast: large language models, prompted to impersonate consumers, reproduced human purchase-intent surveys at &#8220;90% of human test\u2013retest reliability&#8221; (Maier et al., arXiv 2510.08338). Ethan Mollick, whose AI commentary reaches one of the largest audiences in the field, summarized it the next day as the ability to &#8220;predict actual purchase intent (90% accuracy).&#8221; PyMC Labs&#8217; own blog ran the headline &#8220;AI Synthetic Consumers Now Rival Real Surveys.&#8221; A new category of research tooling is now being sold on the strength of that number.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">I read the preprint and its appendix, the released code, the vendor&#8217;s marketing materials and podcast appearances, and the amplification chain. I checked the instrument being copied against five decades of intent-behaviour research: Wicker 1969, Sheeran 2002, List and Gallet 2001, Chandon, Morwitz and Reinartz 2005, Morwitz, Steckel and Gupta 2007, and the new literature on language models as survey respondents, Salecha and Eichstaedt 2024, Bisbee and colleagues 2024. Three things I did not have: the 9,300 human responses (proprietary to Colgate-Palmolive), the anchor statements the method runs on (unpublished), and sales figures for any concept tested (none appear in the paper).<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">I. What the machine actually did<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Grant the engineering its due, because it&#8217;s real. Ask GPT-4o to rate a product concept from 1 to 5, and it clusters on 3; the paper measures this collapse directly, a distributional match to humans of 0.26. The authors&#8217; fix changed what gets elicited. Ask the model to speak as a respondent would, then map the words back onto the scale by measuring how close they are to five reference sentences. The synthetic answers spread out like human answers (similarity 0.88), and when you rank all 57 product concepts by their synthetic scores, the order matches the human panel&#8217;s order at roughly 90% of the panel&#8217;s agreement with itself across random halves. A gradient-boosted model trained on the actual survey data achieved about 65% accuracy on the same test set. The zero-shot language model beat the supervised one. Something real is happening in there.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Now hold the result up to the light and read what it says. Every validation number in the paper measures the distance between one stack of survey answers and another stack of survey answers. Reality never enters the room.<\/strong><\/p>\n\n\n\n<h3 class=\"wp-block-heading\">II. The ruler it copied<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">The edifice rests on one inherited assumption, old enough that nobody states it anymore: that purchase-intent surveys measure future purchasing. That assumption has been directly measured for fifty years, and the record is precise.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Across 422 studies and roughly 82,000 participants, stated intentions explain 28% of the variance in what people subsequently do (Sheeran 2002). In the largest meta-analysis on purchase intent specifically, the average correlation between what consumers said and what they later bought was .49 across more than 200 products. The same analysis found the correlation depends on regime. Existing products: .75. New products: minus .18, statistically indistinguishable from zero. Concepts shown one at a time rather than side by side: minus .09 (Morwitz, Steckel and Gupta 2007). The 57 Colgate surveys tested new personal-care concepts, each shown to its own panel in isolation. The copy was validated in the one regime where the original&#8217;s record against behaviour is statistically zero.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">And the ruler&#8217;s bend has structure. Ask people what they&#8217;d pay, and they claim two to three times what they hand over when money is real (List and Gallet 2001). Stranger: the act of asking manufactures part of the evidence. Consumers who get surveyed show an intent-to-purchase correlation 58% higher than identical consumers who were never asked (Chandon, Morwitz and Reinartz 2005). The instrument inflates its own validity scores by existing. <strong>A survey answer is a small performance, produced for an audience, shaped by what sounds reasonable to say, paid for in currency that costs nothing.<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The ruler is bent. We hold five decades of records showing the exact angle.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">III. The performance, amplified<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Here is where the synthetic version gets interesting, and worse. A language model is trained on text humans wrote for other humans, which is to say, on performances. Dress it in a demographic persona and ask it to perform a shopper, and you have stacked a performance on a performance.<\/strong> The measurable consequence: when GPT-4 can infer its answers are being evaluated, its survey responses shift by 1.2 human standard deviations toward the flattering end; Llama 3 shifts by about one full deviation (Salecha and Eichstaedt 2024, measured on personality questionnaires). Humans flatter the interviewer. The machine flatters harder.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Instability compounds it. <strong>Identical prompts run months apart return measurably different response distributions, and synthetic samples compress variance in ways that distort whatever gets built on top of them <\/strong>(Bisbee et al. 2024). The SSR paper carries its own version of the warning, and to the authors&#8217; credit, they printed it. Strip the demographic personas out, and the score distributions still look perfectly human while the ranking signal collapses from 92% to 50%. Sit with that ablation, because it&#8217;s the most honest result in the paper: <strong>output that looks exactly like human data can carry almost no information. The shape of the answers and the truth of the answers are independent properties.<\/strong><\/p>\n\n\n\n<h3 class=\"wp-block-heading\">IV. Where the stupidity actually lives<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Two words sort the whole mess. Fidelity is how well the copy matches the original. Validity is how well the original matches reality. The paper measured fidelity, carefully, and mostly said so. The market heard validity.<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Watch the qualifier disappear in three steps. The paper&#8217;s abstract: 90% of human test\u2013retest reliability. A fidelity claim, denominated in the panel&#8217;s agreement with itself. The vendor blog: synthetic consumers now rival real surveys. Still survey-framed, rhetorically louder. The viral summary: predicts actual purchase intent, 90% accuracy. A validity claim, no denominator, travelling at the speed of a feed. Each layer dropped one qualifier, and each dropped qualifier was the load-bearing one. The denominator was soft to begin with: the 57 concepts&#8217; human means cluster within a tenth of a point of 4.0, so the ceiling itself is correlation squeezed from a narrow band.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Now the charitable read, because without it the critique is cheap. Concept screening is a ranking job. You hold 57 ideas and a budget for three launches; you need the order, and the level can be wrong. In ranking mode, the bent ruler works: comparative intent correlates with behaviour at .53, and NielsenIQ&#8217;s BASES built a decades-long forecasting business on the consistency of the overstatement, deflating claimed intent against a calibration base of roughly 300,000 prior concept tests. A ruler bent the same way every time ranks objects fine; it just can&#8217;t tell you their length. A machine that reproduces the panel&#8217;s ranking for API pennies, in an industry the paper itself says costs companies billions a year, has real value as a first-pass filter. The LightGBM result says the model contributes genuine signal beyond echoing priors. All of that stands.<\/p>\n\n\n\n<h1 class=\"wp-block-heading\">None of it licenses the word predict. <\/h1>\n\n\n\n<p class=\"wp-block-paragraph\">The system&#8217;s validation ceiling is the panel. The panel&#8217;s ceiling, for new concepts shown in isolation, is statistically zero against behaviour. The honest sentence, the one that fits on no slide, reads: <strong>this method reproduces what a panel would have said, at ninety percent of the panel&#8217;s agreement with itself, in the one regime where what panels say has never tracked what buyers do. <\/strong>That sentence is true, useful, and unfundable, which is why nobody in the chain says it.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The fix is one study. Take concepts that launched, hold out the outcomes, and show synthetic scores predicting realized trial or sales better than deflated human intent does. The parties best positioned to run it are the ones selling the method. No such study exists in public. Until it does, every behaviour-prediction claim built on this work is a fidelity number wearing a validity costume.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The diagnostic travels well beyond synthetic panels, which is why it&#8217;s worth keeping. <strong>Tonight, split your own evidence stack into two piles: records of what customers said, and records of what customers paid. <\/strong>Testimonials, NPS, survey scores, and intent data go in the first pile. Renewals, repeat orders, full price paid without a discount, referrals that closed go in the second. The first pile is fidelity to talk. The second is the only validity you own. And the next time a vendor shows you a ninety percent, ask the only question that matters: ninety percent of agreement with what? If the answer ends at another survey, you are buying a very faithful copy of testimony.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>TLDR: Scientists taught a robot to fill out surveys like a human. The internet heard &#8220;robots can predict what you&#8217;ll buy!&#8221; and stampeded. Nobody checked what anyone actually bought. People in surveys say things they never do. The robot copied the saying part. The herd sharing the headline never read the paper, so they copied [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":4494,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_coblocks_attr":"","_coblocks_dimensions":"","_coblocks_responsive_height":"","_coblocks_accordion_ie_support":"","footnotes":"","rank_math_title":"","rank_math_description":"","rank_math_canonical_url":"","rank_math_focus_keyword":""},"categories":[85],"tags":[],"class_list":["post-4489","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-monopoly"],"_links":{"self":[{"href":"https:\/\/paulsyng.com\/blog\/wp-json\/wp\/v2\/posts\/4489","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/paulsyng.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/paulsyng.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/paulsyng.com\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/paulsyng.com\/blog\/wp-json\/wp\/v2\/comments?post=4489"}],"version-history":[{"count":4,"href":"https:\/\/paulsyng.com\/blog\/wp-json\/wp\/v2\/posts\/4489\/revisions"}],"predecessor-version":[{"id":4500,"href":"https:\/\/paulsyng.com\/blog\/wp-json\/wp\/v2\/posts\/4489\/revisions\/4500"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/paulsyng.com\/blog\/wp-json\/wp\/v2\/media\/4494"}],"wp:attachment":[{"href":"https:\/\/paulsyng.com\/blog\/wp-json\/wp\/v2\/media?parent=4489"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/paulsyng.com\/blog\/wp-json\/wp\/v2\/categories?post=4489"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/paulsyng.com\/blog\/wp-json\/wp\/v2\/tags?post=4489"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}