Focus
Analysing Surveys with LLMs: From Open-Ended Answers to Test Hypotheses
Alexandre Suon · 2026-09-28
Open-ended answers are where customers tell you why they did not buy, yet most teams read a handful and file the rest. Large language models (LLMs) now make AI survey analysis of thousands of verbatims fast and cheap, and published research shows they can code answers nearly as well as trained humans on well-defined tasks. This deep dive explains what they do well, where they fail, and a validated workflow that turns open-ended answers into test hypotheses.
Executive summary
- Open-ended answers are the most valuable and least analysed part of a survey. They carry unprompted reasons in the customer's own words, and they cost respondents effort: on Pew Research Center's panel, item nonresponse averages about 18% for open-ended questions against 1–2% for closed ones.
- On well-defined coding tasks, LLMs now perform close to trained humans. In a 2024 study of British Election Study answers, the best model reached 93.9% accuracy against 94.7% for a second human coder. In a 2023 PNAS study, ChatGPT beat crowd workers by about 25 points of accuracy at under $0.003 per annotation.
- Performance depends on the task, the language and the prompt, so it must be measured every time. Across 27 annotation tasks, one-third had precision or recall below 0.5. Agreement with humans ranged from almost perfect for sentiment to none for a subtle judgement in the same study.
- Seven failure modes are predictable: invented themes, guessed counts, run-to-run drift, positional and majority bias, sycophancy, prompt sensitivity and long-context blind spots. Each has a known guard, from structured output to a human-coded validation sample.
- A reliable workflow has eight steps and puts humans at two of them. Define the decision, build a codebook, human-code a sample, classify with the LLM, measure Cohen's kappa, quantify with confidence intervals, join to behavioural data such as GA4 segments, then write test hypotheses.
- The API bill is rarely the constraint; privacy and validation are. Coding 10,000 answers costs from under $2 to about $63 at September 2026 list prices. The real work is redacting personal data, signing a processor agreement and checking the model against people.
Section 1 · The problem
Open-ended answers hold the “why”, yet most of them are never analysed
AI survey analysis is the use of large language models (LLMs) and related machine-learning tools to read, code, count and summarise survey responses, especially open-ended answers, so that free text can be measured, compared with behaviour and acted on like any other data.
A closed question tells you what people chose from a list you wrote. An open question tells you why, in words you did not anticipate. That is why exit surveys, post-purchase surveys and Net Promoter Score (NPS) follow-ups almost always end with a free-text box. It is also why those boxes are the part of the survey most likely to be ignored.
The reason is cost. Closed answers arrive already counted. Open answers must be read, grouped into themes and counted by hand before they can be compared. This process is called coding. In our experience, most e-commerce teams skim a few dozen answers, copy the most vivid quotes into a slide and move on. The quotes feel like insight, but nobody can say whether the problem they describe affects 2% or 25% of visitors.

What this shows. Pew reports that closed-ended questions on its panel lose 1–2% of respondents, while open-ended questions average about 18% nonresponse and range from 3% to just over 50%. People who do type an answer have spent real effort. A separate Pew analysis found mobile respondents skip open questions more often (15%) than desktop respondents (11%) and write shorter answers. Throwing that text away, or reading only a sample of it, wastes the most costly data in the survey.
Open answers are also where you find the problems you did not know to ask about. A closed list of reasons for abandoning a basket reflects the team's current beliefs. The free-text box catches the surprise: a courier that does not deliver to a region, a size chart that contradicts the product photo, a payment method that fails on one browser. Our Voice of Customer guide covers how to collect these answers. This article is about what to do with them once you have thousands.
Our view. The value of an open-ended question is not the quote you put on a slide. It is the share of customers who raise a problem, how that share differs between segments, and whether fixing it moves conversion. All three need coding at scale, which is exactly what LLMs have made cheap.
Section 2 · The evidence
On well-defined coding tasks, LLMs now perform close to trained human coders
Before 2023, automating open-ended coding meant training a supervised classifier on thousands of hand-labelled examples. That was out of reach for most survey programmes. Generative models changed this, because they can apply a codebook they have only read, a set-up called zero-shot (no examples) or few-shot (a handful of examples) classification. Several peer-reviewed studies have since measured how well that works.
ChatGPT beat crowd workers on accuracy and consistency
Gilardi, Alizadeh and Kubli (PNAS, 2023) compared ChatGPT (gpt-3.5-turbo) with Amazon Mechanical Turk (MTurk) crowd workers and trained research assistants on 6,183 tweets and news articles. Tasks included relevance, stance, topic and frame detection. Across the four datasets, ChatGPT's zero-shot accuracy exceeded that of crowd workers by about 25 percentage points on average, and its per-annotation cost was under $0.003, about thirty times cheaper than MTurk.

What this shows. Intercoder agreement measures how often two coders (or two runs of the same model) give the same label. It averaged 56% for crowd workers and 79% for trained annotators, against 91% for ChatGPT at temperature 1 and 97% at temperature 0.2. Temperature is the setting that controls how much randomness the model adds when it picks words; lower means more repeatable. The chart shows consistency, not truth: a model can be consistently wrong, which is why the workflow in Section 5 checks it against humans.
On real survey answers, the best models came within a point of a human
Closer to the survey use case, Mellon and colleagues (Research & Politics, 2024) asked six LLMs, three supervised models and a second human coder to code “most important issue” answers from the British Election Study Internet Panel into 50 categories. The panel holds more than 650,000 labelled answers, so the researchers could test two realistic scenarios.

What this shows. When a researcher had no labelled data, the best model (Claude-1.3) reached 93.9% accuracy against 94.7% for the human. When an old codebook had to be applied to newer answers, the gap to the human widened to 7.7 points, although the model still beat a supervised model trained on 576,000 cases. PaLM-2 and Llama-2 did substantially worse than the best models. The lesson for e-commerce teams is that model choice matters, and that performance drops when customers start talking about things the codebook did not anticipate.
The same studies warn that results vary widely by task
Two other results keep the enthusiasm in check. Törnberg (2023) found GPT-4 more accurate and more reliable than experts and crowd workers at classifying the political affiliation of tweets, but that is a task with a clear right answer. Pangakis, Wolken and Fasching (2023) ran 27 annotation tasks from 11 published social-science datasets through GPT-4. The median F1 score (a combined measure of precision and recall) was 0.707 and median accuracy 0.850, yet nine of the 27 tasks had precision or recall below 0.5. In their words, for a full one-third of tasks the model missed at least half of the true cases, produced more false positives than true positives, or both.
Language matters too. Von der Heyde and colleagues (2025) tested several LLMs on German open-ended answers about survey motivation and found that only a fine-tuned model reached satisfactory performance. Without fine-tuning, models classified some categories much better than others, which distorted the distribution of reasons. For a multilingual European retailer, that is a direct warning: validate in each language you report on.
| Study | Task | Headline result | What it means for you |
|---|---|---|---|
| Gilardi et al., PNAS 2023 | Relevance, stance, topics, frames (6,183 texts) | Accuracy about 25 points above crowd workers; under $0.003 per annotation | LLMs can replace low-cost crowd coding for clear tasks |
| Törnberg, 2023 | Political affiliation of tweets | GPT-4 more accurate and reliable than experts and crowd workers | Strong on tasks with a clear ground truth |
| Mellon et al., Research & Politics 2024 | Survey answers into 50 categories | 93.9% vs 94.7% for a human; 80.9% vs 88.6% on newer answers | Close to human on survey open-ends; weaker on drift |
| Pangakis et al., 2023 | 27 tasks, 11 datasets | Median F1 0.707; one-third of tasks below 0.5 precision or recall | Always validate against human labels |
| Von der Heyde et al., 2025 | German survey answers | Only a fine-tuned LLM was satisfactory; category shares distorted | Check each language separately |
| Wen et al., AI & Society 2026 | Thematic coding of charity reports | Kappa 0.91–0.95 for sentiment, 0.61–0.65 for themes, ≤0.01 for a subtle judgement | Agreement is task-specific |
Section 3 · Strengths
LLMs do five survey jobs well, as long as the output is a label you can check
The jobs where LLMs add most value share one property: each answer gets a small, structured output, such as a theme code, a sentiment value or an extracted product name, that a human can verify against the text. Jobs that produce long prose are harder to check and riskier.
1. Coding answers into themes
This is the core task of open ended survey analysis: assigning each answer to one or more codes from a codebook, a list of themes with definitions, inclusion and exclusion rules and examples. The model reads each answer, applies the definitions and returns the codes. This is deductive coding. LLMs can also help with inductive coding, proposing themes from a sample of answers. De Paoli (Social Science Computer Review, 2024) showed that GPT-3.5 could reproduce most of the main themes found by human researchers in two interview datasets, while arguing that the approach has clear limits and still needs a researcher in charge.
2. Sentiment and intent
Sentiment (positive, neutral, negative) is the task where models agree most with people. Wen and colleagues found Cohen's kappa of 0.91–0.95 between GPT-4o and human coders for sentiment. Intent is more useful for e-commerce: was the visitor ready to buy, comparing, or just browsing? Intent codes need tighter definitions than sentiment, because “I'll come back later” can mean either.
3. Summarising thousands of verbatims
A summary is the fastest way to brief a stakeholder, and most native survey tools now offer one. It is also the output most prone to invented detail, because it cannot be checked line by line. Use summaries to describe themes you have already counted, not as a substitute for counting.
4. Translating multilingual feedback
A pan-European shop may receive answers in five languages. LLMs can code them directly against an English codebook or translate them first. In our experience, coding directly works well for major European languages, but the German results above show that accuracy is not uniform. Keep the original text next to any translation, and validate a sample in each language you report on. Some tools, such as Hotjar and Contentsquare, write AI summary reports in English while preserving quotes in the original language.
5. Extracting product and attribute mentions
Answers such as “the linen trousers came up small and the colour was greyer than the photo” contain structured facts: product (linen trousers), attribute (size, colour) and direction (runs small, differs from photo). An LLM with a fixed output schema can extract these into columns you can join to your product catalogue. This is where survey analysis starts to feed merchandising and product-page tests, not only UX fixes.
| Job | Output per answer | How to check it | Typical use |
|---|---|---|---|
| Theme coding | One or more codebook IDs | Kappa against a human-coded sample | Prioritising problems |
| Sentiment and intent | One value from a short list | Kappa; confusion matrix | Segmenting feedback, NPS follow-ups |
| Summarising | Short prose | Every claim traced to counted themes | Stakeholder briefings |
| Translation | Text in the analysis language | Bilingual spot checks per language | Multi-market programmes |
| Attribute extraction | Product, attribute, direction | Exact-match check against the answer and catalogue | Product-page and merchandising tests |
Section 4 · Failure modes
Seven failure modes are predictable, and each one has a known guard
The research that shows LLMs can code well also documents how they go wrong. None of these failures is random bad luck. Each follows from how the models work, which means each can be anticipated and designed against.
| Failure mode | What happens | Evidence | Guard |
|---|---|---|---|
| 1. Hallucinated themes and quotes | The model reports a theme or quote that is not in the data | Wen et al.: GPT-4o changed the meaning of acronyms; GPT-4o-mini fabricated excerpts | Ask for a verbatim evidence span and check it exists in the answer |
| 2. Overconfident counts | “About 40% mention delivery” is a guess, not a count | Von der Heyde et al.: unequal accuracy across categories distorted the distribution of reasons | Never ask the model to count; classify each row, then count in code |
| 3. Run-to-run drift | The same answer gets different codes on different runs | Reiss (2023): identical inputs gave different classifications; Thinking Machines: 80 unique completions from 1,000 runs at temperature 0 | Low temperature, fixed model version, repeat runs on a sample |
| 4. Positional and majority bias | Codes listed last, or seen most often in examples, are over-chosen | Zhao et al. (ICML 2021); Zheng et al. (ICLR 2024) on option-position bias | Shuffle code order across runs; balance few-shot examples |
| 5. Sycophancy | The model agrees with the hypothesis in your prompt | Sharma et al. (ICLR 2024): assistants tend to match the user's stated views | Neutral prompts; never state the answer you expect |
| 6. Prompt sensitivity | Small wording or format changes move results | Sclar et al. (ICLR 2024): up to 76 accuracy points between formats for one model | Freeze the prompt; test two phrasings before launch |
| 7. Context-window limits | Answers in the middle of a long paste get less attention | Liu et al., Lost in the Middle (TACL 2024) | One answer, or a small batch, per call |
Invented themes and quotes
When Wen and colleagues used GPT-4o to code charity reports, it altered the meaning of acronyms, for example expanding one programme name into an unrelated one. The smaller GPT-4o-mini sometimes omitted excerpts, fabricated excerpts that did not exist, or presented them in the wrong order. Excerpt extraction was weak even for the larger model, with precision of 0.41 and recall of 0.53. In a survey report, a fabricated quote is worse than no quote, because it will be repeated in meetings as evidence.
Guessed counts
If you paste 3,000 answers into a chat window and ask “what share mention delivery cost?”, the model will give you a number. It has not counted. It has produced a plausible figure. The only reliable count comes from classifying each answer separately and adding up the labels in a spreadsheet or SQL. Even then, as von der Heyde and colleagues found, uneven accuracy across categories can skew the shares, which is why Section 5 measures agreement per theme, not only overall.
Drift between runs
Setting temperature to zero does not make an LLM fully repeatable. Reiss (2023) found that ChatGPT's classifications changed with repeated identical inputs and with small instruction changes, and concluded that it should not be used for annotation without validation. Thinking Machines Lab (2025) sampled the same prompt 1,000 times at temperature 0 and got 80 different completions, because server load changes how calculations are batched. Pangakis and colleagues turned this into a tool: answers the model labelled the same way across repeated runs were 19.4 points more accurate than answers where it wavered. Inconsistency is a useful flag for human review.
Positional, majority and agreement biases
Zhao and colleagues (ICML 2021) showed that language models favour answers that appear often in the prompt's examples or near its end. Zheng and colleagues (ICLR 2024) found that models favour certain option positions in multiple-choice questions regardless of content. In a codebook prompt, that means the last theme in the list, or the theme with the most examples, can be over-assigned. Sharma and colleagues (ICLR 2024) documented sycophancy, a tendency to agree with the user's stated view. A prompt that says “we think delivery cost is the main issue” invites the model to find it.
Prompt wording and long contexts
Sclar and colleagues (ICLR 2024) measured differences of up to 76 accuracy points on one open-source model from formatting changes alone, such as separators and spacing. Liu and colleagues (TACL 2024) showed that models use information at the start and end of a long input better than information in the middle. Both argue for the same design: a frozen, tested prompt applied to one answer, or a small batch, per call, rather than one giant paste.

What this shows. Cohen's kappa measures agreement between two coders after removing the agreement expected by chance. McHugh argues that any kappa below 0.60 means little confidence should be placed in the results. GPT-4o cleared that bar comfortably for sentiment and only just for theme classification, and failed completely on a subtle judgement about evidence of impact. You cannot know in advance which of your themes will behave like which, so you have to measure each one.
For marketers. Treat an AI summary as a first draft, not a finding. If a number or quote will go into a deck, it must come from a counted, validated classification.
For leaders. Ask one question of any AI-generated customer insight: “What was the agreement with a human coder, per theme?” If nobody can answer, the insight has not been checked.
Section 5 · The workflow
To analyze survey data with AI reliably, follow eight steps and keep humans at two of them
The workflow below is how we run LLM-assisted survey analysis for clients. It borrows from content-analysis practice in the social sciences and from the validation approach Pangakis and colleagues recommend, adapted to e-commerce decisions and behavioural data.

What this shows. Steps 1, 2 and 8 frame the work around a decision. Steps 3 and 5 are the checks that make the output trustworthy, and they are the ones most teams skip. Steps 6 and 7 turn labels into evidence you can prioritise: a share with an uncertainty range, broken down by the segments that matter for conversion.
Step 1. Define the decision
Start with the choice the analysis should inform, for example “which three checkout problems do we test next quarter?”. The decision sets the grain of the codebook. A UX team needs “size chart unclear” separate from “item runs small”; a board needs only “sizing”.
Step 2. Build or induce the codebook
Draft themes from a random sample of 100–200 answers, by hand or with an LLM proposing candidates (inductive). Then fix each theme with a one-line definition, what to include, what to exclude and two real examples. Add “Other” and “No usable answer”. Keep the list short, usually eight to fifteen themes. Allow several themes per answer, because customers often give more than one reason.
Step 3. Human-code a validation sample
Two people independently code a random sample of the answers with the codebook, without seeing the model's output. In our experience, 200–400 answers is enough to estimate agreement per common theme; rare themes need more. Resolve disagreements and write the agreed label as the reference set. Where the two humans disagree a lot, the definition is the problem, not the coders: fix it before involving the model.
Step 4. Classify every answer with the LLM
Send one answer per call, or small batches, with the frozen codebook prompt, temperature at 0 or close to it, a fixed model version and structured output, a JSON schema the model must follow. Major providers offer schema-constrained output; Anthropic, for example, says its structured outputs guarantee schema-compliant responses through constrained decoding. Ask for a short verbatim evidence span per theme and a confidence value, and store the model name and prompt version with every row.
Step 5. Measure agreement
Compare model labels with the human reference set, theme by theme, using Cohen's kappa. As a Henkan & Partners rule of thumb, report a theme when kappa is 0.70 or above, report with a caveat between 0.60 and 0.70, and below 0.60 rewrite the definition, merge the theme or code it by hand. Also check every evidence span actually appears in the answer text; a missing span is a hallucination.
Cohen's kappa: κ = (p_o − p_e) / (1 − p_e)
p_o = share of answers where model and human agree on the theme (yes/no)
p_e = agreement expected by chance = p_model_yes × p_human_yes + p_model_no × p_human_no
Example: 300 answers; both say 'delivery cost' on 72, both say not on 214, they disagree on 14
p_o = (72 + 214) / 300 = 0.953
model yes = 79/300, human yes = 79/300 → p_e = 0.263² + 0.737² = 0.612
κ = (0.953 − 0.612) / (1 − 0.612) = 0.88
Step 6. Quantify with confidence intervals
Count the labels in code, never in the model. Report each theme as a share of answers with a 95% confidence interval, for example with the Wilson interval, which behaves well for small shares. The interval stops teams from ranking themes that are statistically tied. Remember the denominator is people who answered the question, not all visitors: nonresponse is not random, as Pew's device and demographic differences show.
Step 7. Join to behavioural data
A theme becomes actionable when you know who raises it. If your survey tool passes a session or client ID, join answers to analytics data such as GA4 segments, device, traffic source, basket value, new versus returning customer and whether the visitor later converted. Cross-tabulations often reveal that a theme is concentrated in one segment, which changes both the size of the prize and the test design. Pair this with session replay on respondents who raised the theme to see the problem in context.
Step 8. Turn top themes into test hypotheses
Write each hypothesis in a fixed form: because we saw this theme in this segment, we will change this, and we expect this metric to move. Size the opportunity from the theme's share and segment conversion, then feed it into your A/B testing programme. Survey data tells you what customers say. Only a controlled experiment tells you whether fixing it changes what they do. Our Point of View on Voice of Customer explains why this loop is the missing layer in most optimisation programmes.
Section 6 · Worked example
In an illustrative exit survey, the top theme becomes three testable hypotheses
Illustrative example. The retailer, verbatims and numbers in this section are synthetic, created by Henkan & Partners to show the method. They are not client data and should not be read as benchmarks.
A mid-sized European fashion retailer shows an exit-intent survey on the basket page: “What stopped you from ordering today?”. Over four weeks it collects 2,400 answers in French, English and Spanish. Free delivery starts at €60. The team wants to choose the next three checkout tests.
The codebook
| Code | Definition | Include | Exclude | Example (synthetic) |
|---|---|---|---|---|
| DEL_COST | Delivery price is too high or unclear | Shipping fee, free-delivery threshold | Delivery speed | “€6.95 delivery on a €35 top, no thanks” |
| SIZE_FIT | Unsure which size to choose or how it fits | Size chart, fit, between sizes | Out of stock sizes | “No idea if M or L, the chart is in inches” |
| BROWSE | Not ready to buy yet | Comparing, saving for later | Price complaints | “Just looking for now” |
| PRICE | Price or discount concerns | Waiting for sale, promo code failed | Delivery fee | “My code WELCOME10 did not work” |
| PAYMENT | Preferred payment method missing or failing | PayPal, buy now pay later, card declined | Price | “No PayPal on my phone?” |
| DEL_TIME | Delivery too slow or date unclear | Delivery date, express | Delivery fee | “Need it by Friday, says 5–7 days” |
| RETURNS | Returns policy or cost | Return fee, window | Sizing doubt | “Returns cost €4, too risky” |
| TECH | Site error or bug | Error, crash, button not working | Payment declined by bank | “The basket keeps emptying” |
Two analysts code 300 random answers. After one round of definition fixes, the model's kappa against their reference set is 0.88 for delivery cost (the worked calculation in Section 5), and 0.70 or above for the other themes except “just browsing”, which reaches 0.64. In this example, “just browsing” is reported with a caveat.
The counts and the cross-tab

What this shows. Delivery cost leads at 25.5% (95% interval 23.8–27.3%), clearly ahead of sizing at 18.2%. Joining answers to GA4 basket value shows the theme is raised by 34.0% of respondents with a basket under €60 against 12.5% above it. Payment options look minor overall at 9.2%, but the same join shows 11.5% of mobile respondents raise it against 5.7% on desktop. The overall ranking and the segment view point to different tests.
From themes to hypotheses
| Theme and evidence | Hypothesis | Primary metric | Segment |
|---|---|---|---|
| Delivery cost: 34% of sub-€60 baskets | Because shoppers below €60 object to the delivery fee, we will show a progress bar to free delivery in the basket, and expect basket-to-order conversion and average order value to rise | Revenue per visitor | Baskets under €60 |
| Size and fit: 18% overall | Because shoppers are unsure of their size, we will add a fit indicator from returns data on product pages, and expect add-to-basket rate to rise without higher return rates | Add-to-basket rate; guardrail: return rate | Apparel product pages |
| Payment: 11.5% on mobile vs 5.7% on desktop | Because mobile shoppers miss their preferred wallet, we will surface PayPal and buy now pay later earlier in mobile checkout, and expect checkout completion to rise | Checkout completion | Mobile |
Each hypothesis now has a size estimate, a segment and a metric. Before launch, the team checks it has enough traffic in the segment to detect a realistic effect, using the approach in our article on A/B test statistical models. The survey is rerun after each test to see whether the theme's share falls in the treated segment.
Section 7 · Prompts
A good codebook prompt is short, neutral and forces a structured, checkable answer
The prompt is part of the measurement instrument, like the wording of a survey question. Write it once, test it on the validation sample, then freeze it and version it. Here is a template we use as a starting point for chatgpt survey analysis or any other LLM.
SYSTEM
You code customer survey answers for an online shop. Apply the codebook exactly.
Use only the definitions given. If no code fits, use OTHER. If the answer is empty,
off-topic or unreadable, use NO_ANSWER. An answer may have several codes.
CODEBOOK (order is shuffled on each run)
DEL_COST: delivery price too high or unclear. Excludes delivery speed.
SIZE_FIT: unsure which size to choose or how it fits. Excludes out-of-stock sizes.
PAYMENT: preferred payment method missing or failing. Excludes price.
... (one line per code)
USER
Survey question: "What stopped you from ordering today?"
Answer (language may vary): "{answer_text}"
OUTPUT (JSON schema enforced)
{ "codes": [ { "code": "DEL_COST", "evidence": "<exact words copied from the answer>" } ],
"sentiment": "negative | neutral | positive",
"language": "fr | en | es | ...",
"confidence": 0.0-1.0 }
Five details matter more than clever wording:
- Neutral framing. The prompt never says which theme you expect to win. That limits sycophancy.
- Exact evidence. The evidence field must be copied from the answer, so a script can confirm it exists. A missing match flags a hallucination.
- Escape codes. OTHER and NO_ANSWER stop the model forcing weak answers into real themes.
- Shuffled order. Rotating the order of codes between runs, or at least between validation runs, exposes positional bias.
- Balanced examples. If you add few-shot examples, give each code the same number, or none, to avoid majority-label bias.
Our view. Resist the urge to ask the model for insights, recommendations or root causes in the same call. Coding is a measurement task. Interpretation belongs to a later step where a person looks at counted, validated themes. We make the same argument about using LLMs to predict test outcomes in prompt-based experimentation: models are useful for generating and structuring, not for replacing evidence.
Section 8 · Tools and cost
Native tools are the fastest start, general LLMs give control, and the API bill is rarely the constraint
You have three families of tools. They are not exclusive: many teams use a native summary for a quick read and a general LLM pipeline for the numbers they report.
Native AI in survey and feedback platforms
| Tool | AI features for open text (as documented) | Notable limits |
|---|---|---|
| Qualtrics Text iQ | Recommends topics from frequent terms; related-term suggestions when building topics | Topic recommendations need Advanced Text iQ |
| Medallia | Themes enhanced with generative AI and AI summaries; Ask Athena natural-language questions (announced February 2024) | Enterprise platform; check contract for model hosting |
| Hotjar and Contentsquare surveys | Sentiment (positive, neutral, negative), AI-assisted tags for recurring themes, summary report with findings, quotes and next steps | Summary needs at least 20 answers; uses up to 365 days; summary written in English |
| Typeform Smart Insights | AI summary of main themes, topic detection and sentiment with supporting quotes | At least 10 answers; topics from the most recent 1,000; English only; Enterprise, Talent and Growth plans |
| SurveyMonkey thematic analysis | Groups text answers into themes; separate sentiment feature; uses a self-hosted Anthropic model and an Azure OpenAI or OpenAI model | English only; some paid plans; 1–10,000 answers; not available in the EU or Canada data centres |
Native tools are quick, and the data stays inside a platform you have already approved. The trade-off is control. You usually cannot see the prompt, fix the codebook, choose the model or measure agreement per theme. That is fine for a weekly read-out; it is not enough for a number that decides a quarter's roadmap.
General LLMs with spreadsheets, SQL and code
Claude, ChatGPT and Gemini can run the full workflow in Section 5. For small volumes, Google Sheets' AI function (syntax `AI("prompt", range)`) can categorise text and analyse sentiment, though Google notes that only the first 350 selected cells with AI functions are generated at a time. For volume, BigQuery offers managed AI functions such as `AI.CLASSIFY`, which assigns text to categories you list, alongside `AI.IF` and `AI.SCORE`, so answers can be coded where your GA4 export already lives. For full control, a short script calls a model API with structured output and writes results to a table. Agents connected through the Model Context Protocol (MCP) can now run these steps against your warehouse, but the validation steps stay the same.
Embeddings and clustering
An embedding turns each answer into a vector of numbers so that similar answers sit close together. BERTopic, an open-source library, combines embeddings with dimensionality reduction (UMAP), clustering (HDBSCAN) and a class-based TF-IDF weighting to find topics, and can call an LLM to name them. This is the best tool for discovery on large or unfamiliar datasets, as an input to Step 2. Clusters are not codes, though: they have no definitions, and they shift when the data changes, so use them to propose themes, then code against a fixed codebook.
Cost and scale

What this shows. We assumed a 1,200-token codebook prompt plus an 80-token answer, 60 tokens of JSON output and one call per answer, with no prompt caching, which would cut costs further. All three providers offer batch APIs at 50% off for work that does not need an instant answer, which suits survey coding. Even the most capable model costs less than a day of analyst time. The expensive part is people: two coders validating 300 answers, and someone owning the codebook. Choose the model on measured kappa, not price, then use the cheapest model that passes.
Section 9 · Privacy
Verbatims contain personal data, so redact first and send only to a processor under contract
Customers write things in free-text boxes that you never asked for: names, email addresses, order numbers, phone numbers, addresses, and sometimes health or family details. Under the General Data Protection Regulation (GDPR), any of these can make a verbatim personal data. Sending it to an LLM provider is processing, and the provider becomes your processor. Article 28 GDPR requires the controller to use only processors that give “sufficient guarantees”, under a binding contract, usually called a data processing agreement (DPA).
- Redact before sending. Strip emails, phone numbers, order IDs and names with pattern rules and a detection tool. The open-source Presidio toolkit detects and anonymises many entity types, but its documentation warns that automated detection gives no guarantee of finding all sensitive information, so combine it with other controls.
- Use business terms, not consumer apps. Anthropic states it does not use inputs or outputs from its commercial products to train its models by default. OpenAI states its models are not trained on API or business-plan data unless the customer opts in. Google states that paid-tier Gemini API content is not used to improve its products, while free-tier content is.
- Choose the hosting region. OpenAI has offered European data residency for its API since February 2025, with in-region processing and zero data retention for those projects. Some native tools restrict AI features by region: SurveyMonkey's thematic analysis is not available in its EU data centre.
- Minimise what you send. Send the answer text and a random row ID, not the customer ID, email or full session. Do the join to behavioural data inside your own warehouse.
- Document it. Record the purpose, legal basis, processor, region and retention in your records of processing, and tell respondents in the survey privacy notice that answers may be analysed with automated tools.
For leaders. The privacy question is not “can we use AI on survey data?” but “which processor, in which region, under which contract, with what redaction?”. Settle it once with your data protection officer, then reuse the pattern for every survey. This section is general information, not legal advice.
Section 10 · Mistakes
Most errors come from skipping validation, not from choosing the wrong model
- Asking the model to count. A share typed by a chatbot is not a measurement. Classify every row, then count in code.
- Skipping the human sample. Without a reference set there is no kappa, and without kappa you cannot tell a 0.9 theme from a 0.3 theme.
- Reporting overall accuracy only. A 90% overall score can hide one theme at kappa 0.4, and that theme may be the one you act on.
- Letting the codebook drift. Changing definitions mid-stream breaks comparisons over time. Version the codebook and prompt, and re-validate after any change.
- Leading the model. Putting your hypothesis in the prompt invites sycophancy. Keep prompts neutral.
- Pasting everything into one chat. Long contexts hide answers in the middle and make counts impossible to audit. Use one answer or a small batch per call.
- Ignoring who did not answer. Respondents are not all visitors. Weight or at least segment results, and never present a theme share as a share of all customers.
- Stopping at the insight. A theme is a hypothesis, not a result. Route it into an experiment and measure the behaviour change.
Section 11 · Next steps
What to do next
1. Pick one survey and one decision
Choose the survey with the most open-ended answers and a decision coming up, often a basket exit survey or an NPS follow-up. Write down the decision in one sentence before you look at the data.
2. Build and test a codebook in a week
Draft eight to fifteen themes from 150 random answers, have two people code 300 answers, and fix definitions until human agreement is solid. This week of human work is the foundation for everything else.
3. Run the model and measure kappa per theme
Classify the validation sample with a frozen prompt and structured output, compute kappa per theme, and only then classify the full dataset. Keep the prompt, model and codebook versions with the results.
4. Settle privacy once
Agree the processor, region, DPA and redaction steps with your data protection officer, and turn them into a reusable checklist for every future survey.
5. Join, size and test
Join coded answers to GA4 segments and orders, size the top themes, and send the best two or three into your experimentation backlog with a clear hypothesis and metric. If you want help setting up the codebook, the validation or the pipeline, Talk to us.
Disclosure: Henkan & Partners sells Voice of Customer and experimentation services, including LLM-assisted analysis of survey data.
FAQ
Frequently asked questions about AI survey analysis
Frequently asked questions
What is AI survey analysis?
AI survey analysis is the use of large language models and related machine-learning tools to read, code, count and summarise survey responses, especially open-ended answers. Done well, it assigns each answer to themes from a defined codebook, checks the model against human coders, and reports theme shares with confidence intervals.
How do I analyze survey data with AI?
Define the decision, build a codebook of eight to fifteen themes, have two people code a random sample of 200–400 answers, classify every answer with an LLM at low temperature using structured output, measure Cohen's kappa per theme, count the labels with confidence intervals, join them to behavioural data such as GA4 segments, and turn the top themes into test hypotheses.
Can ChatGPT analyze open-ended survey responses accurately?
On well-defined tasks, yes. A 2023 PNAS study found ChatGPT about 25 points more accurate than crowd workers, and a 2024 study found the best LLM coded survey answers with 93.9% accuracy against 94.7% for a human. But results vary by task and language, so you must validate against a human-coded sample every time.
What is open ended survey analysis?
Open-ended survey analysis is the process of turning free-text answers into structured data, usually by coding each answer into themes, then counting and comparing those themes across segments. It reveals reasons customers give in their own words, including problems a closed list of options would miss.
Is thematic analysis with AI reliable?
It can be, for deductive coding against a clear codebook, but reliability must be measured. One 2025 study found GPT-4o reached a kappa of 0.91–0.95 with human coders for sentiment, 0.61–0.65 for themes and no agreement for a subtle judgement. Use AI to propose themes, and humans to define and validate them.
How many answers do I need to validate an LLM classifier?
In our experience, 200–400 randomly sampled answers coded by two people is enough to estimate agreement for common themes. Rare themes need a larger or targeted sample. Re-validate whenever you change the model, the prompt or the codebook.
What is a good Cohen's kappa for survey coding?
Kappa between 0.61 and 0.80 is usually called substantial and above 0.80 almost perfect. McHugh (2012) argues that anything below 0.60 is inadequate. As a Henkan & Partners rule of thumb, report themes at 0.70 or above, add a caveat between 0.60 and 0.70, and fix or merge themes below 0.60.
Is it GDPR-compliant to send survey answers to an LLM?
It can be, if you treat the provider as a processor under a data processing agreement, redact personal data before sending, use business API terms that exclude training on your data, choose an appropriate hosting region and document the processing. Consumer chat apps and free tiers are rarely suitable. Check with your data protection officer.
How much does customer feedback analysis with an LLM cost?
At September 2026 list prices, coding 10,000 answers costs from about $2 to $63 in API fees depending on the model and whether you use a batch API. The larger cost is the human time to build the codebook and validate the model.
Key terms
- Open-ended question
- A survey question answered in free text rather than from a list. It captures reasons in customers' own words, including problems you did not think to ask about.
- Verbatim
- A customer's answer quoted exactly as written. Verbatims are the raw material of coding and the evidence behind every theme.
- Coding
- Assigning each answer to one or more themes from a defined list. It turns text into countable data, which makes themes comparable across segments and over time.
- Codebook
- The list of themes with definitions, inclusion and exclusion rules and examples. It is the measurement instrument: vague definitions produce unreliable counts, whoever does the coding.
- Deductive coding
- Applying a fixed codebook to answers. It is the task LLMs perform most reliably and the one to use for numbers you report.
- Inductive coding
- Discovering themes from the data rather than starting with a list. LLMs and clustering help propose themes, but humans should define them.
- Large language model (LLM)
- An AI model trained on large amounts of text that can read and generate language, such as Claude, GPT or Gemini. It can apply a codebook it has only read, without task-specific training.
- Zero-shot and few-shot classification
- Asking a model to classify text with no examples (zero-shot) or a handful (few-shot). It removes the need for thousands of labelled training cases, but makes prompt design critical.
- Temperature
- A model setting that controls randomness in its output. Low values make coding more repeatable, though not perfectly deterministic.
- Structured output
- Forcing a model's answer to follow a JSON schema. It makes results machine-readable and lets you check evidence spans automatically.
- Cohen's kappa
- A measure of agreement between two coders that corrects for chance agreement. It is the standard way to check an LLM against humans, theme by theme.
- Hallucination
- Output that looks plausible but is not supported by the input, such as an invented quote or theme. In survey analysis it creates false evidence, so outputs must be traceable to the text.
- Sycophancy
- A model's tendency to agree with views stated by the user. In analysis it biases results towards the hypothesis written in the prompt.
- Wilson confidence interval
- A method for the uncertainty range around a proportion that behaves well for small shares. It shows which themes are genuinely different and which are statistically tied.
- Embedding
- A numeric representation of text in which similar meanings sit close together. It powers clustering tools such as BERTopic for discovering themes in large datasets.
Sources
All sources were checked in September 2026. Vendor feature descriptions (Qualtrics, Medallia, Hotjar, Contentsquare, Typeform, SurveyMonkey, Google, Anthropic, OpenAI) are vendor data and change frequently. Törnberg (2023), Pangakis et al. (2023) and von der Heyde et al. (2025) are arXiv preprints. Exhibit 5, the guards in the failure-mode table and the kappa thresholds are Henkan & Partners frameworks. Exhibit 6 and Section 6 use synthetic, illustrative data. Exhibit 7 is a Henkan & Partners estimate computed from published list prices.
- Pew Research Center (2021). Why do some open-ended survey questions result in higher item nonresponse rates than others?
- Pew Research Center (2023). Nonresponse rates on open-ended survey questions vary by demographic group, other factors
- Gilardi, F., Alizadeh, M. & Kubli, M., PNAS (2023). ChatGPT outperforms crowd workers for text-annotation tasks
- Mellon, J. et al., Research & Politics (2024). Do AIs know what the most important issue is? Using language models to code open-text social survey responses at scale
- Törnberg, P., arXiv (2023). ChatGPT-4 outperforms experts and crowd workers in annotating political Twitter messages with zero-shot learning
- Pangakis, N., Wolken, S. & Fasching, N., arXiv (2023). Automated annotation with generative AI requires validation
- von der Heyde, L. et al., arXiv (2025). AIn't Nothing But a Survey? Using Large Language Models for Coding German Open-Ended Survey Responses on Survey Motivation
- Wen, C., Clough, P., Paton, R. & Hoque, M. M., AI & Society (2026). Leveraging large language models for thematic analysis: a case study in the charity sector
- De Paoli, S., Social Science Computer Review (2024). Performing an inductive thematic analysis of semi-structured interviews with a large language model
- McHugh, M. L., Biochemia Medica (2012). Interrater reliability: the kappa statistic
- Reiss, M. V., arXiv (2023). Testing the reliability of ChatGPT for text annotation and classification: a cautionary remark
- He, H. & Thinking Machines Lab (2025). Defeating nondeterminism in LLM inference
- Zhao, T. Z. et al., ICML (2021). Calibrate before use: improving few-shot performance of language models
- Zheng, C. et al., ICLR (2024). Large language models are not robust multiple choice selectors
- Sharma, M. et al., ICLR (2024). Towards understanding sycophancy in language models
- Sclar, M. et al., ICLR (2024). Quantifying language models' sensitivity to spurious features in prompt design
- Liu, N. F. et al., TACL (2024). Lost in the middle: how language models use long contexts
- Grootendorst, M., arXiv (2022). BERTopic documentation: neural topic modelling with a class-based TF-IDF procedure
- Qualtrics (2026). Topics in Text iQ
- Medallia (2024). Medallia launches four breakthrough AI innovations
- Hotjar (2026). How to analyze survey responses with AI
- Contentsquare (2026). How to analyze survey results with AI
- Typeform (2026). Smart Insights: qualitative and quantitative analysis
- SurveyMonkey (2026). Thematic analysis
- Google (2026). Use the AI function in Google Sheets
- Google Cloud (2026). Perform semantic analysis with managed AI functions in BigQuery
- Anthropic (2026). Claude pricing and Structured outputs
- OpenAI (2026). API pricing
- Google (2026). Gemini API pricing
- OpenAI (2025). Introducing data residency in Europe
- Anthropic (2026). Is my data used for model training?
- Data Privacy Stack (2026). Presidio: data protection and de-identification SDK
- European Union (2016). General Data Protection Regulation, Article 28: Processor