AI-Powered Market Research: Synthetic Respondents, Survey Bots, and What Still Works
Where AI speeds up market research, the evidence on synthetic respondents, AI bots passing 99.8% of survey attention checks, and scraping law basics.

AI has made parts of market research dramatically cheaper. Coding 5,000 open-ended survey answers used to take a team days; a language model does a first pass in minutes. Reading every app review and support ticket for recurring complaints was impractical; now it is routine. Those gains are real.
AI has also created two problems that research teams cannot ignore. Vendors now sell "synthetic respondents," AI personas that answer surveys in place of people, with claims that outrun the evidence. And AI bots have become good enough to contaminate the online panels most research depends on. This guide covers the useful applications and both problems.
Where AI clearly helps
Coding open-ended responses. Models group free-text answers into themes, count them, and pull representative quotes. Check a sample against human coding, especially for sarcasm and domain terms, and review the theme list rather than accepting it blindly.
Voice of the customer. Support tickets, sales call notes, app store reviews, and community posts contain unprompted feedback in customers' own words. Models can read all of it and surface recurring problems, feature requests, and how often each appears. Unlike a survey, you did not choose the questions, which often reveals issues nobody thought to ask about.
Interview synthesis. Transcribing and summarizing qualitative interviews, and finding where participants agree and disagree. The researcher still decides what matters.
Competitor monitoring. Tracking pricing pages, product changes, job postings, and release notes over time. Useful, with the legal caveats below.
Questionnaire drafting. Generating question drafts, spotting leading or double-barreled questions, and translating. A research expert should still review every question.

Synthetic respondents: promising, not a substitute
The idea is to ask a language model to answer as, say, a 45-year-old small business owner in Ohio, many times over, instead of paying for a sample. Research on this is mixed:
- Argyle and colleagues, in a 2023 Political Analysis paper titled "Out of One, Many", found that GPT-3 conditioned on demographic backstories could reproduce some patterns in how US subgroups answered political survey questions.
- A 2023 Harvard Business School working paper by James Brand, Ayelet Israeli, and Donald Ngwe, "Using GPT for Market Research," found that GPT's willingness-to-pay estimates for products were broadly realistic in aggregate and responded sensibly to price. The authors also report that the model did not give meaningful estimates for differences between groups such as income levels.
Later work has found clear limits. A 2024 Political Analysis study by Bisbee and colleagues compared ChatGPT personas with real American National Election Studies respondents and found problems recovering the distribution of opinion, sensitivity to small changes in prompt wording, and different results from the same prompt over a three-month period. A separate study of 3,200 participants found that language models tend to misportray and flatten demographic groups, so synthetic answers understate disagreement and can reproduce stereotypes rather than actual subgroup views. And a model cannot know how people react to something new, like your unreleased product or a price change last month, because that is not in its training data.
A reasonable use is to pilot a questionnaire, generate hypotheses, or estimate rough direction before paying for fieldwork. A decision that matters still needs real respondents.
The bot problem in online panels
Most commercial research now runs on opt-in online panels, and they were already struggling with low-quality respondents. In November 2025, Sean Westwood of Dartmouth published a study in PNAS titled "The potential existential threat of large language models to online survey research." He built an AI agent from a short prompt that completed surveys while keeping a consistent persona, remembering earlier answers, and adjusting its writing style to the persona's education level. It passed 99.8% of standard attention checks across 6,000 trials, and it also got past a battery of other checks, including logic puzzles and "reverse shibboleth" questions built to expose nonhuman respondents. He also showed that, in seven national polls from the final week of the 2024 US presidential campaign, as few as 10 to 52 synthetic respondents could have flipped which candidate appeared to lead.
The economics make this worse. Westwood estimates a survey can be completed by a commercial model for about $0.05, against the $1.50 a standard survey pays a human respondent, a profit margin of over 96%.
What research buyers should do now:
- Ask panel providers how they verify identity, not just how they screen for attention. Attention checks were built to catch inattentive humans and are not an AI defense.
- Use behavioral signals: timing patterns, copy-paste detection in open ends, device and network checks.
- Prefer verified samples for important decisions: customer lists, probability-based panels, or recruited participants.
- Compare with behavior. If survey results say customers want a feature but usage data and sales say otherwise, trust the behavior.
Sample size still matters
AI does not change sampling math. For a percentage, the 95% margin of error is about 1.96 Ć ā(p(1āp)/n):
| Respondents | Margin of error at 50% |
|---|---|
| 100 | ±9.8 points |
| 400 | ±4.9 points |
| 1,000 | ±3.1 points |
| 2,500 | ±2.0 points |
Subgroups need their own sample: a 1,000-person survey with 150 people in your target segment has a margin of about ±8 points for that segment. And no sample size fixes a biased sample. Online sentiment from vocal users on social media is not representative of your customer base, however many posts you analyze.

Scraping and data rules
- US. In hiQ Labs v. LinkedIn, the Ninth Circuit affirmed a preliminary injunction in 2022, finding serious questions about whether scraping public profiles is access "without authorization" under the Computer Fraud and Abuse Act. hiQ still lost on LinkedIn's user-agreement claims and settled into a consent judgment with a permanent injunction in December 2022. In January 2024 a federal court held in Meta v. Bright Data that Meta's terms could not be read to bar logged-off scraping of public data. Terms of use, copyright, and state privacy laws still apply.
- EU and UK. Scraped data that identifies people is personal data under GDPR and needs a lawful basis, even if it was public. Regulators have fined companies heavily for building databases from scraped personal data: the Dutch authority fined Clearview AI ā¬30.5 million in 2024 over a scraped database of faces.
- Practical rule. Use official APIs and licensed data where they exist, respect robots.txt and rate limits, and do not collect personal data you do not need.
For using customer data in AI tools generally, see our AI and SaaS data privacy guide.
A sensible setup
- Start with data you own: support tickets, reviews, and call notes. Use AI to find themes and counts.
- Use AI to draft and pilot surveys, then field them with verified respondents.
- Code open ends with AI, checking a sample against human coding.
- Monitor competitors through legitimate sources.
- Test conclusions against behavior such as usage, conversion, and churn before making big decisions.
For turning research into decisions, see AI for strategic decisions. For turning it into content, see AI content marketing strategy, and for predictive modeling on customer data, predictive analytics.
This guide is for informational purposes only and is not legal advice. Data collection rules vary by jurisdiction; check with counsel before scraping or processing personal data.



