Synthetic Data in Market Research: A Powerful Tool or a Dangerous Shortcut?
- Bryan Cruz
- Jul 13
- 5 min read

Since AI and chatbots became the talk of the town (especially on LinkedIn), one application has stood out to me as someone coming from a market research background: synthetic data.
I'll be honest, I was initially skeptical. I get the appeal of using synthetic data because collecting data from real respondents has always been one of the biggest challenges in market research.
But, can simulated consumers actually replace real respondents?
Working in the market research industry, my instinct was to defend traditional research methods and focus on the limitations of AI-generated data.
However, after building my own consumer behaviour simulator, I came to a realization: Its greatest value is in creating a research testing environment. It should function as an R&D environment that allows researchers to test assumptions, explore scenarios, and improve study design before collecting real-world evidence.
Synthetic data helps me think through what could happen, so I can design better ways to find out what actually happens. In this sense, synthetic data is not replacement for real respondents. It is a flight simulator for consumer behaviour research and development.
Moving From Abstract Theory to a Tangible Sandbox: Removing the Black Box
If you're like me, you probably understand concepts best by seeing how they work in practice. Rather than treating synthetic consumers as a black box, I wanted to understand the mechanics behind them by building one myself.
To explore this further, I built a prototype that simulates consumer purchasing behaviour in a retail environment. The goal was to understand what synthetic consumers can and cannot do. The simulation was able to generate consistent behavioural patterns based on the parameters I provided.
Watch the video demonstration below to see the prototype in action.
How Did I Construct the Simulator to Model Consumer Decision-Making?
To make the simulation feel realistic, I gave each synthetic respondent a set of behavioural inputs that mirror how might real consumers make purchase decisions. These include impulse-buy tendencies, price sensitivity, wallet limits, product preferences, and context-based cognitive triggers. Together, they aggregate into a utility score that estimates whether the person will buy in a given scenario.
Simulation Variables | Operational Meaning | Example Interpretation |
Baseline susceptibility to unplanned, emotion-driven purchases | Higher IBI means the respondent is more likely to convert on a cue-driven or low-friction offer | |
Degree to which higher prices reduce purchase likelihood | Higher PS means rising price more strongly suppresses utility and purchase probability | |
Available disposable category budget and the proportion of spend a brand can capture | Larger wallet allows more trial and repeat buying; smaller wallet imposes stricter spending constraints | |
Stable baseline liking for particular brand segments or offer types | A respondent may prefer Establishment brands over Challenger or Creator-Owned offerings | |
Situational cues that shift evaluation, attention, or purchase readiness | A living-room spillover or livestream exposure can temporarily increase curiosity and consideration | |
Unified numeric representation of the respondent’s net preference after all inputs are combined | Higher utility indicates stronger predicted purchase likelihood under the given scenario |
What I've Learned After Building The Simulator
The model could generate realistic-looking consumers because I defined the rules, variables, and relationships. But those outputs were still constrained by the information and logic provided to the system. In other words, synthetic data can simulate behaviour. It does NOT discover human behaviour. This is where the danger lies. A synthetic model can create highly plausible results while still reflecting the biases, blind spots, and limitations of its original design.
One reason I remain cautious about replacing real respondents with synthetic consumers is that human behaviour often contains counterintuitive and complex relationships. For example, research has found that pre-store advertising can increase purchase intentions while simultaneously reducing spontaneous purchases by encouraging more deliberate decision-making. Similarly, studies on social media brand communities show that brand value is often co-created through user-generated content, brand interactions, and shared experiences. Therefore, a synthetic model focused only on individual-level attributes may overlook how social influence, communities, and evolving brand meanings shape consumer decisions. Ultimately, these findings remind me that human behaviour does not always follow intuitive or linear patterns.
However, this does not mean synthetic models cannot represent complex behaviour. They can, provided those behavioural relationships have been identified through empirical evidence or learned from appropriate data. The challenge is that a simulation cannot independently validate whether those assumptions are correct.
For that reason, I see synthetic data as a tool for exploring hypotheses, stress-testing research designs, and evaluating alternative scenarios.
Where Synthetic Data Creates Value?
The goal is not to create a perfect digital replica of consumers. The idea of completely replacing real respondents with synthetic data is where I become skeptical. Consumer behaviour is complex, emotional, and often irrational. Human decisions are influenced by context, culture, values, personal lived experiences, and unexpected factors that are extremely difficult to capture through predefined parameters.
That does not mean synthetic data is useless. In fact, I believe it has enormous potential when applied correctly. Synthetic data can be valuable for:
Exploring hypotheses safely before conducting large-scale, expensive research.
Testing experimental designs to ensure your choice models actually make sense.
Simulating different market scenarios (like testing how a competitor's price increase might ripple through your segments).
Generating research ideas and edge-case scenarios to test in the field.
Improving AI-assisted research workflows by stress-testing analysis frameworks early.
Therefore, synthetic data can simulate scenarios based on the assumptions and relationships provided, but it cannot independently discover new motivations, emerging behaviours, or unexpected consumer responses that were never represented in the model.
Synthetic Data as a Flight Simulator for Consumer Behaviour Research and Development
Synthetic data does not replace real-world situation. Similarly, a flight simulator cannot guarantee that a plane will perform exactly the same way under every real-world condition. However, it allows pilots to identify weaknesses, improve decision-making, and prepare for unexpected situations.
Before launching a large-scale study, researchers can use synthetic models to explore questions such as:
Does the research design capture the decision-making factors we care about?
Are the product attributes structured correctly?
Could certain consumer segments respond differently to a marketing stimulus?
Are there unexpected outcomes worth investigating further?
Are there flaws in our assumptions before collecting real-world data?
The opportunity is not to replace respondents with synthetic data, but to improve the research process before engaging with them. However, like any research tool, synthetic data must be grounded in sound research principles, appropriate statistical methods, and validation against real-world consumer behaviour.
In conclusion, synthetic data helps researchers explore what could happen, so they can design better ways to understand what actually happens.
References:
Che, H., Erdem, T., & Öncü, T. S. (2013). Consumer learning and evolution of consumer brand preferences. Journal of Retailing, 89(4), 515–531. https://marketing.wharton.upenn.edu/wp-content/uploads/2016/10/Paper-Erdem-Tulin-04-04-2013-Version-2.pdf
Conjointly. (n.d.). What is conjoint analysis? (with examples). https://conjointly.com/guides/what-is-conjoint-analysis/
Karjaluoto, H., Munnukka, J., & Tiensuu, S. (2016). The effects of brand engagement in social media on share of wallet. Journal of Brand Management, 23(1), 50–67. https://aisel.aisnet.org/cgi/viewcontent.cgi?article=1017&context=bled2015
Mandolfo, M., & Lamberti, L. (2021). Past, present, and future of impulse buying research methods: A systematic literature review. Frontiers in Psychology, 12, 687404. https://doi.org/10.3389/fpsyg.2021.687404
Park, S.-H., Mahony, D. F., Kim, Y., & Kim, Y. D. (2015). Curiosity generating advertisements and their impact on sport consumer behavior. Sport Marketing Quarterly, 24(1), 17–28. https://doi.org/10.1016/j.smr.2014.10.002
Schlereth, C., Eckert, C., & Skiera, B. (2013). Willingness-to-pay has always been conceptualized as a point estimate, frequently as the price that makes the consumer indifferent between buying and not buying the product. Marketing Letters, 24(4), 423–440. https://doi.org/10.1007/s11002-012-9177-2
Strübing, S. L. (2013). Pre-store advertising’s effect on consumer decision-making: An eye tracking experiment [Master’s thesis, Copenhagen Business School]. CBS Research Portal. https://research.cbs.dk/en/studentProjects/a0d640a7-d80c-45fe-923f-6519a557917d/
Wilk, V., Gilstrap, C., & How, D. (2026). Brand transcendence through brand value co-creation exchanges within social media craft beverage brand communities. Journal of Vacation Marketing. https://doi.org/10.1080/02508281.2026.2640384



Comments