Focus
AI Agents in Experimentation: What They Do at Each Step, What Vendors Ship and Where Humans Stay in Charge
Alexandre Suon · 2026-09-28
AI agents in experimentation have moved from demos to shipped features in under two years. Optimizely, Kameleoon, Wingify (VWO and AB Tasty), Amplitude and Contentsquare now offer agents that research, write hypotheses, build variants, read results and connect to your data through the Model Context Protocol (MCP). This deep dive walks through the A/B testing workflow step by step, shows what each vendor actually ships, and sets out where human judgement must stay in charge and the operating model that keeps agents fast and safe.
Executive summary
- Agents can now help at every step of an A/B test, but three decisions must stay human. Research, ideation, planning, building, QA, monitoring, analysis and documentation can all be drafted or done by an agent. Choosing what to test, approving a launch and deciding what to ship should stay with named people.
- The vendors have shipped, not just announced. Between February and September 2026 Optimizely added Opal agents for ideas, conflict checks, prioritisation, value estimates and governance; Kameleoon released four PBX 2.0 agents; Amplitude launched a Web Experimentation Agent; Wingify lists nine agents across VWO and AB Tasty. Most also offer an MCP server.
- Building is solved enough; judging is not. In our study of 19 prompt-built tests, average build time fell from 8.7 hours to 55 minutes, but 5 of 19 builds scored below 70% fidelity. In a Nature study, GPT-4 predicted survey experiment effects well (r = 0.85) but large megastudies much less well (r = 0.34) and overestimated effect sizes.
- Statistics is where agents do the most damage. An agent that checks a test ten times and stops at the first 'win' turns a 5% false positive rate into about 20% in our simulation; one that scans 20 metrics has a 64% chance of a false win. Agents must read results from the testing platform's engine and follow pre-registered stopping rules.
- MCP is the connector that makes agents useful, and the main new risk. Google Analytics and Contentsquare offer read-only servers; Amplitude, Optimizely and Wingify also allow writes, controlled by permissions. Start read-only, grant write access by exception, and log every action, as OWASP and the MCP specification recommend.
- Run agents like junior analysts inside a clear operating model. Four roles (strategist, reviewer, agent operator, decision owner), four approval gates and a learning repository that agents read before every idea and update after every decision. A 90-day plan moves from read-only access to gated actions.
Section 1 · Definition
An AI agent plans and takes steps with tools, which makes it different from a chatbot
AI agents in experimentation are software systems built on large language models that take a goal (for example, "find and test ways to reduce checkout drop-off"), plan the steps, and use tools such as analytics queries, session replay, a testing platform or a code editor to carry them out, with a person reviewing the output at defined points.
The difference from a chat assistant is agency. An assistant answers a question you ask. An agent decides which data to pull, runs the query, reads the result, drafts the next step and, if allowed, acts on a system. That is why the same technology that saves hours can also launch a broken variant or declare a false winner without anyone noticing.
This article is the fourth step of our Optimise with AI learning path. Our study of 19 prompt-built tests measured what AI does to the build step, and our C-level agenda for prompt-based experimentation explained why leaders should care. Here we widen the lens to the whole workflow, from research to the learning repository.
Three terms recur throughout:
- Large language model (LLM). The model underneath the agent, such as GPT, Claude or Gemini. It predicts text, and with the right tools it can write code, call APIs and summarise data.
- Tool use. The agent's ability to call functions: run a GA4 report, list live experiments, take a screenshot of a page, create a draft test. Tools are what turn a model into an agent.
- Model Context Protocol (MCP). An open standard, released by Anthropic in November 2024, for connecting AI assistants to the systems where data lives. Most testing and analytics vendors now publish an MCP server, so one agent can reach several tools through a common interface.
Our view. Treat an agent as a very fast junior analyst with no memory of your business and no instinct for when a number looks wrong. It is excellent at drafting, searching and repetitive checks. It needs a brief, access limited to its task, and a senior reviewer before anything reaches customers.
Section 2 · The workflow
Agents can help at all eleven steps, but only some should run without a person in the loop
An A/B test is not one task. It is a chain of eleven steps, each with its own inputs, risks and owner. Asking "should we use AI for testing?" is too broad. The useful question is: for each step, how much can an agent do, and who checks it?

What this shows. Three colours, three levels of trust. Green steps are routine and easy to verify, so an agent can do the work and a person checks it. Blue steps need judgement about the business, so the agent drafts and a person edits. The three orange steps are commitments: which idea gets scarce traffic, whether a variant is safe to show customers, and whether to ship. Those stay with named people, with the agent preparing the evidence.
| Step | What an agent can do today | What a person must do | Main risk |
|---|---|---|---|
| 1. Research | Summarise analytics, session replays, surveys and reviews; flag friction | Check the sample and the source; add business context | Confident summaries of thin data |
| 2. Ideation | Propose ideas from page scans and past tests | Reject generic ideas; link each to evidence | Plausible but unoriginal ideas |
| 3. Hypothesis | Draft hypothesis, metrics, risks | Choose which ideas get traffic | Plausibility mistaken for evidence |
| 4. Design and sample size | Suggest metrics, audiences, duration, sample size | Confirm the primary metric and minimum detectable effect | Wrong baseline or metric |
| 5. Build | Write variant code from a prompt or a design | Review the code that will ship | Global CSS, broken logic |
| 6. QA | Check conflicts with live tests, run scripted checks | Visual and device QA; sign-off | Unchecked edge cases |
| 7. Launch | Configure targeting, goals and traffic split | Approve the launch | Wrong audience or split |
| 8. Monitor | Watch for sample ratio mismatch and guardrail breaches | Act on alerts | Alert fatigue |
| 9. Analyse | Draft the readout from the platform's statistics | Check the method and segments | Hallucinated or cherry-picked numbers |
| 10. Decide | Recommend ship, iterate or stop; estimate value | Make the call | Automated shipping of false wins |
| 11. Document | Write the learning to the repository | Approve the learning and tags | A repository full of noise |
The pattern is simple. The closer a step is to customers or to a business decision, the more human control it needs. The more a step is about gathering, drafting or checking against a rule, the more an agent can take on. For the fundamentals behind each step, see our essential guide to A/B testing.
Section 3 · Research and hypotheses
Agents speed up research and ideation, but a plausible hypothesis is not evidence
Research synthesis is the safest place to start. The agent reads data that already exists, a person can check the summary against the source, and a weak summary costs time, not customers.
Research synthesis: what ships today
- Contentsquare describes its Sense Analyst agent as working "24/7 to analyze experience data, detect issues and growth opportunities" (March 2026), and exposes the same analysis through a read-only MCP server launched in October 2025.
- Amplitude launched a Session Replay Agent that reviews sessions to spot friction and an AI Feedback Agent that turns unstructured feedback into insights (February 2026).
- Wingify, the group formed when VWO and AB Tasty combined, lists an Insight Agent that analyses funnels, replays and surveys, and a Synthetic Testing Agent that tries changes on synthetic visitor profiles before a live test.
- Microsoft Clarity and Google Analytics do not ship research agents of their own for this purpose, but both publish MCP servers, so any MCP-compatible assistant can query them.
Two cautions apply. First, an agent summarising 40 session replays will describe them with the same confidence as 4,000. Ask it to state its sample and to quote the sessions it used. Second, synthetic users and AI-generated research are good for generating hypotheses, not for measuring effects. Our guide to synthetic users for customer research explains where they help and where they mislead.
Hypothesis and prioritisation: fast drafts, human choice
Every major vendor now offers an ideation or planning agent:
- Optimizely Opal has an Experiment Planning agent that turns a URL and a test idea into "a hypothesis, variant descriptions, targeting settings, a metrics table, statistical sizing guidance, risks, and assumptions"; an Idea Builder (April 2026) that generates concepts from page context and research; and a Backlog Prioritization agent (June 2026) that scores ideas with the PIE framework.
- Kameleoon PBX Ideate (May 2026) scans a page and ranks test ideas using data from more than 20,000 historical experiments, then turns them into prompts for the build agent.
- Wingify lists a Hypothesis Agent that converts evidence into testable ideas with ranked variations; its earlier AB Tasty Evi suite included Evi Hypothesize, which scores hypothesis quality.
These tools produce well-formed hypotheses in seconds. The question is whether they pick good ones. The best published evidence says: better than you might expect in controlled surveys, much worse in the field.

What this shows. In 70 preregistered survey experiments with 469 effects, GPT-4's predictions correlated strongly with the real results (r = 0.85). In 15 megastudies with 606 effects, closer to real-world behaviour, the correlation fell to 0.34, still slightly ahead of expert forecasters at 0.26, and the model systematically overestimated effect sizes. In 17,681 Upworthy headline tests, prompted LLMs picked winners only marginally better than chance. The right panel is a reminder that AI help also feels more productive than it measures.
The practical conclusion is to use agents to widen and rank the list, never to skip the test. An agent's confidence score is a prior, not a result. A human strategist should still choose which ideas get scarce traffic, and every hypothesis should cite the evidence behind it: a funnel drop, a replay pattern, a survey theme or a past test.
For marketers. Ask the ideation agent for ten ideas, then ask it to show the evidence for each. Keep the ones where the evidence is your own data. Drop the ones where the evidence is "best practice".
For leaders. Measure the win rate of agent-suggested ideas separately from human ones for the first two quarters. If agent ideas win less often, you are saving time on ideation and spending it on failed tests.
Section 4 · Build and QA
Agents cut build time sharply, so QA becomes the new bottleneck
Building variants was the first step to be automated, and it is the best measured. In July and August 2025 we rebuilt 19 historical, hand-coded A/B tests with Kameleoon's Prompt-Based Experimentation (PBX), using prompts only.

What this shows. Speed was never the problem: every test was built faster, and build time fell by about 89% on average. Quality was uneven. Most simple changes were close to perfect, but five builds needed substantial manual work, mainly because words are a poor way to describe a pixel-exact layout or a business rule. That is why the QA step now matters more than the build step.
What build agents ship today
- Optimizely Opal's AI variation development agent modifies and creates page elements in the new Visual Editor from natural language and retrieves page styles to stay on brand. Optimizely notes that a task uses 30 to 130 Opal credits and that "any code Opal generates is a custom code solution and falls outside the scope of Optimizely Support".
- Kameleoon PBX 2.0 (April 2026) has a Build agent that browses the site to understand components in context and can convert Figma designs into variations, plus a Configure agent for audiences, goals and launch.
- Wingify Copilot creates an entire experiment, including variations, metrics and audience, from a plain-English description (August 2025, Pro and Enterprise plans).
- Amplitude's Web Experimentation Agent (February 2026) is described as designing and launching experiments and making rollout decisions.
Why AI-generated code needs stricter QA, not looser
Three properties of LLM-written code change the QA job:
- The same prompt can give different code. A study of ChatGPT code generation by Ouyang and colleagues found that for 48% to 76% of tasks, depending on the benchmark, repeated requests produced no two outputs with the same test results, and setting temperature to zero reduced but did not remove this. Approve the exact code that ships and never regenerate after approval.
- Code can reach beyond its target. In our study, one variant wrote site-wide CSS that broke dropdown menus elsewhere on the page. Add a scope check for global styles and regressions on nearby elements.
- Fast feels productive. In METR's 2025 randomised trial, experienced developers using AI tools took 19% longer while believing they were 20% faster. Measure cycle time from brief to launch, not build time alone.
Agents also help with QA itself. Optimizely's Experiment Conflict Checker agent (April 2026) checks a new test against every live experiment in Web Experimentation, Performance Edge or Personalization. That kind of rule-based check suits an agent well. Visual QA on real devices, and the final sign-off, remain human jobs. For the code patterns that make variants robust, see our developer's guide to DOM manipulation for A/B tests.
Section 5 · Analysis and readouts
Let agents draft the readout, but never let them choose when to stop or which metric won
Analysis agents are the most tempting and the most dangerous. They turn a dense results page into a clear paragraph for executives. They can also produce numbers that look right and are not.
What ships today: Optimizely's Experiment Summary agent produces a condensed report that highlights key findings, whether results reached statistical significance and next steps. Wingify's AB Tasty Evi suite included Evi Analysis for post-test insights. Adobe's Journey Optimizer Experimentation Accelerator, launched in September 2025, uses an Experimentation Agent to explain why tests win or lose and rank new opportunities. Amplitude's MCP server lets any assistant read and create experiments and charts.
Three statistical failure modes to design out
- Hallucinated statistics. An LLM asked to compute a p-value or confidence interval may do the arithmetic itself and get it wrong, or invent a figure that was never in the data. Rule: the agent must quote numbers returned by the testing platform's statistics engine, with a link to the source, and must not compute significance on its own.
- Peeking. An agent monitoring a test around the clock is the ultimate peeker. Checking a fixed-horizon test repeatedly and stopping at the first significant result inflates false positives, a problem Evan Miller described in 2010. Rule: use a sequential method designed for continuous monitoring, or let the agent alert but not stop the test.
- Metric fishing. Ask an agent "did anything improve?" and it will search every metric and segment until something does. Rule: pre-register one primary metric and a small set of guardrails before launch, and label everything else as exploratory.

What this shows. A test with no real effect should show a false win 5% of the time. Look ten times and stop at the first win, and it happens about 20% of the time in our simulation. Check 20 independent metrics and the chance that at least one looks like a winner is 64%. An agent does both faster and more often than any analyst, so the rules must be built into its instructions and its tools.
Family-wise false positive rate = 1 - (1 - alpha)^k
with alpha = 0.05 and k = 20 independent metrics: 1 - 0.95^20 = 0.64
Data quality checks are a good fit for agents because they follow rules. A sample ratio mismatch (SRM), where the observed split between variants differs from the planned split, signals a broken experiment. Microsoft researchers reported that about 6% of experiments at Microsoft exhibit an SRM. An agent that checks for SRM daily and blocks the readout until it is resolved adds real safety. For the statistical models behind these rules, see our guide to A/B testing statistics for marketers.
Our view. The readout should be a draft that a named analyst signs. The ship decision should never be automated on the basis of a single agent-read result, whatever the vendor setting allows. If an agent can make rollout decisions, gate that permission behind a human approval and a pre-registered rule.
Section 6 · Learning repository
A learning repository is what turns fast agents into smart ones
Agents have no memory of your business unless you give them one. Without a record of past tests, an ideation agent will happily propose the free-shipping banner you tested and lost last year. The learning repository (a searchable record of every test's hypothesis, evidence, design, result, decision and code) is that memory.
Kohavi, Tang and Xu's standard text on online experiments devotes a chapter to institutional memory and meta-analysis for this reason: past results are the best guide to future ideas. Agents raise the stakes in two ways. They can read the repository before proposing anything, and they can write to it after every decision, which removes the chore that usually lets repositories decay.
Vendors are moving here too. Optimizely's Experimentation Program Overview agent (June 2026) reports quarterly on the testing programme, and its Experiment Value Estimator agent projects annualised impact from winning experiments. Kameleoon's PBX Ship (May 2026) turns a winning client-side variation into production code behind a feature flag through its MCP server, so the tested code and the shipped code stay linked.
What to store so an agent can use it
| Field | Why the agent needs it |
|---|---|
| Hypothesis and evidence | Lets the agent check whether a new idea repeats an old one |
| Page, audience, lever tags | Lets it find related tests and patterns by page and theme |
| Primary metric, guardrails, sample, duration | Lets it judge how much weight a past result deserves |
| Result with interval, and decision | Separates 'lost' from 'inconclusive', which mean different things |
| Final prompt and shipped code | Lets a winner be rebuilt or handed to developers |
| Learning in one sentence, approved by a person | Stops agent-written noise from polluting the memory |
Disclosure: Henkan & Partners builds its own AI and MCP tooling for clients' experimentation programmes, and works with several of the testing vendors named in this article.
Section 7 · What vendors ship
Every major testing vendor now ships agents, but they differ in where the agent may act
Twenty-two months after MCP was released, agents and MCP connections reach every layer of the experimentation stack. The pace has been set by launches and by consolidation: Datadog bought Eppo in May 2025, OpenAI bought Statsig for $1.1 billion in September 2025, Everstone combined Wingify and AB Tasty in January 2026, and Amplitude took over the Statsig brand, platform and customers in May 2026.

What this shows. The first year was about access: MCP servers let assistants read analytics. The second was about action: vendors added agents that build, configure, prioritise and ship. The deals show experimentation becoming core product infrastructure, owned by analytics and developer platforms rather than sold as a marketing add-on.
| Vendor | Agents shipped (as documented) | MCP server | Can the agent change things? |
|---|---|---|---|
| Optimizely (Opal) | Experiment Planning, Idea Builder, AI variation development, Conflict Checker, Experiment Summary, Backlog Prioritization, Value Estimator, Program Overview, Governance, Feature Flag Implementation | Yes, remote, launched April 2026 | Yes: creates and configures flags and experiments |
| Kameleoon (PBX 2.0) | Ideate, Build, Configure, Ship | Yes, used by PBX Ship | Yes: builds, configures, ships behind feature flags |
| Wingify (VWO, AB Tasty) | Nine listed agents incl. Insight, Hypothesis, Synthetic Testing, Experiment, Rollout, Guardrail; earlier Evi suite and Copilot | Yes: 10 read tools, 4 write tools | Yes: creates A/B and split URL tests, by permission level |
| Amplitude | Global, Dashboard Monitoring, Session Replay, Web Experimentation, AI Feedback | Yes | Yes: reads and writes charts, experiments, cohorts, with separate write permission |
| Contentsquare | Sense Analyst, configurable insight agents | Yes, read-only | No: read-only data access |
| Adobe | Experimentation Agent in Journey Optimizer Experimentation Accelerator | Not reviewed | Analyses and recommends |
| Google Analytics, Microsoft Clarity | No experimentation agents | Yes, read-only | No |
Three differences matter more than feature lists. First, scope of action: some agents only read and recommend, others create and launch tests. Second, pricing model: an Opal variation build uses 30 to 130 credits, Kameleoon PBX typically uses one to two credits per experiment, and Wingify says its Evi features were included in all contracts at no additional cost; model the cost at your test volume. Third, openness: an MCP server lets you use your own assistant across tools, while an in-product agent is easier to govern but locks the workflow to one vendor. For the wider market picture, see our report on the A/B testing tool market.
For leaders. Do not choose a platform for its agents alone. Agents are converging quickly; your data model, statistics engine and governance are harder to change. Ask each vendor what the agent can do without approval, what is logged, and how you switch the write permission off.
Section 8 · MCP connections
MCP makes agents useful by connecting them to your data, and risky for exactly the same reason
MCP is an open standard for connecting AI assistants to the systems where data lives. An MCP server wraps a tool, such as GA4 or a testing platform, and publishes a list of actions the assistant may call. An MCP client, such as Claude, ChatGPT, Cursor or Copilot, discovers those actions and uses them. In December 2025 Anthropic moved MCP to the Agentic AI Foundation under the Linux Foundation; at the time the project counted more than 10,000 active servers and 97 million monthly SDK downloads.
For an experimentation team, MCP means one assistant can pull a funnel from GA4, pull behaviour metrics such as scroll depth and friction from Clarity or Contentsquare, check live tests in Optimizely or Wingify, and draft a hypothesis, without exporting files between tools. The permissions differ sharply by server:
| MCP server | Access | Notable limits (as documented) |
|---|---|---|
| Google Analytics | Read-only (analytics.readonly scope) | Labelled experimental; reports, funnels and real-time data; cannot change configuration |
| Microsoft Clarity | Query only | 10 API requests per project per day, up to 3 days of data, 3 dimensions per request |
| Contentsquare | Read-only | Experience, friction, conversion and journey data |
| Amplitude | Read and write | Separate 'Use MCP (read)' and 'Use MCP (write)' permissions; uses existing user access controls |
| Optimizely Experimentation | Read and write | Queries results, flags and audiences; creates and configures flags and experiments |
| Wingify | Read and write | Browse, Design and Publish/Admin permission tiers; AI can create only A/B and split URL tests |
Security: prompt injection and over-broad permissions
Connecting an agent to live systems adds two security risks that most experimentation teams have not faced before:
- Indirect prompt injection. OWASP ranks prompt injection first in its 2025 Top 10 for LLM applications. The indirect form happens when the model reads external content, such as a web page, a survey answer or a support ticket, that contains hidden instructions. An agent that reads open-text feedback and also has write access to your testing tool is exactly that pattern. OWASP recommends restricting the model's privileges to the minimum necessary and requiring human approval for high-risk actions.
- Over-broad scopes. The MCP security best-practice guidance warns against wildcard or "full-access" scopes and recommends a progressive, least-privilege model: start with low-risk read operations and elevate only when a privileged action is first needed, with elevation events logged.

What this shows. Agents sit between people and systems, never above them. Every system connects through its own MCP server with the narrowest scope that does the job. The four orange gates are the only paths from an agent's draft to a customer-facing change. The learning repository is both a source and a destination, so every test makes the next one better informed.
For marketers. You can start today with read-only servers: connect GA4 and Clarity or Contentsquare to your assistant and ask it to explain a funnel drop with behaviour data. No write access is needed to get most of the research value.
For leaders. Put MCP connections through the same approval as any other integration: an owner, a scope, a data-processing check and an audit log. Consent rules still apply to what the agent reads; see our guide to user consent in e-commerce.
Section 9 · Human judgement and risks
Human judgement stays essential for what to test, what is safe and what counts as a win
The three orange steps in Exhibit 1 share a feature: each is a commitment made under uncertainty that someone must be accountable for. An agent can prepare the evidence, but it cannot own the consequences.
- What to test. Traffic is the scarcest resource in most programmes. Choosing which hypothesis gets it depends on strategy, brand, seasonality and politics that no agent sees.
- Whether it is safe to launch. A variant that breaks checkout on an older phone costs revenue and trust. The launch approval is a promise to customers.
- What counts as a win and what ships. Most tested ideas do not win: Kohavi, Deng and Vermeer report success rates of about 33% at Microsoft, 15% at Bing, 10% at Booking.com, Google Ads and Netflix, and 8% in Airbnb Search. With base rates that low, a significant result is often a false positive, so the ship decision needs someone who understands the method.
| Risk | What it looks like | Control |
|---|---|---|
| Hallucinated statistics | A readout quotes a lift or p-value that is not in the platform | Agent may only quote engine output, with a link; analyst signs off |
| Peeking and metric fishing | Agent stops at the first 'win' or finds a winning segment | Pre-registered primary metric and stopping rule; sequential methods for monitoring |
| Code quality | Global CSS, broken logic, different code on regeneration | Approve the exact build; scope check; two-person QA for high-traffic pages |
| Prompt injection | Hidden instructions in feedback or web content steer the agent | Least privilege; separate read and write agents; human approval for actions |
| Governance and privacy | Agent reads personal data or acts outside its remit | Scoped MCP access, consent checks, action logs, named owner per agent |
| Generic ideas | Backlog fills with 'best practice' tests | Every hypothesis cites first-party evidence |
| Cost creep | Credits and tokens grow with volume | Budget per test; track cost per decided test |
| Over-trust | People stop checking because the agent is usually right | Random audits of agent output; track error rates |
Culture matters as much as controls. Teams that already reward learning over winning adapt well to agents, because they are used to challenging results. Our article on building a culture of experimentation covers the habits that make that possible.
Section 10 · Operating model
Run agents like junior analysts: clear roles, earned permissions and measured output
The operating model is where most of the value is won or lost. The tools are converging; the difference between programmes will be who owns what, which permissions the agents hold, and whether anyone measures the results.
Four roles, even in a small team
| Role | Owns | In a small team | In a large team |
|---|---|---|---|
| Experimentation strategist | Goals, backlog, choice of what to test | Head of CRO or e-commerce manager | Programme lead per product area |
| Reviewer | Code, QA, statistics, brand and legal checks | Same person with a checklist, plus an external reviewer for stats | Named QA and analytics reviewers |
| Agent operator | Prompts, tools, MCP permissions, cost, logs | Analyst or developer, a few hours a week | Dedicated role in the experimentation platform team |
| Decision owner | Ship, iterate or stop | Business owner of the page | Product owner, with finance for high-value decisions |
Permissions are earned, and results are measured
Give agents access in three stages: see, draft, act. In the first month they read analytics, replays and past results. In the second they draft hypotheses, plans, variants and readouts that people edit. Only in the third, and only for low-risk tests, do they take actions such as creating a draft test, always behind the approval gates. Track four numbers from the start, so you can tell whether agents help:
- Cycle time from brief to decided test.
- QA defect rate: variants that fail QA or are rolled back after launch.
- Win rate and learning rate of agent-suggested versus human-suggested ideas.
- Cost per decided test, including credits, tokens and review time.

What this shows. The plan sequences trust. Measurement and permissions come first, because without a baseline you cannot tell whether agents help. The learning repository is loaded early, because every later agent depends on it. Write access arrives last, scoped to drafts, with two-person QA before anything reaches customers.
Worked example: one test, agent-assisted (illustrative)
Illustrative example, not a client case. A fashion retailer sees mobile checkout drop-off rise. The research agent pulls the GA4 funnel through the read-only MCP server and finds the drop concentrated at the delivery step; it then queries Clarity for scroll depth and engagement on that step. The strategist watches 30 session replays, sees shoppers scrolling past the delivery options, confirms the pattern and chooses the idea: show delivery cost and date above the fold. The planning agent drafts the hypothesis, primary metric (checkout completion), guardrail (average order value) and sample size; the analyst corrects the baseline. The build agent writes the variant; the reviewer rejects the first version for a global style change, approves the second and signs QA. The monitoring agent checks SRM daily. At the planned end date the analysis agent drafts the readout from the platform's numbers; the decision owner ships. The agent writes the learning, and the strategist approves it.
People made four decisions in that example. Agents did most of the typing.
Section 11 · What to do next
What to do next
1. Map your workflow and pick two agent-ready steps
List your eleven steps, who owns each and how long each takes. Start agents where the work is slow, rule-based and easy to verify: research synthesis and documentation are usually the best first candidates.
2. Connect your data read-only
Set up read-only MCP connections to your analytics and behaviour tools. Give one person the agent operator role, with a log of every query. Most of the research value comes before any write access.
3. Write the rules before the agents run
Pre-register primary metrics and stopping rules, require agents to quote only platform statistics, and define the four approval gates. Put them in the agent's instructions and in your test template.
4. Build the learning repository
Load past tests with hypotheses, results, decisions and code. Make "read the repository first" the first instruction of every ideation agent, and "write the learning" the last step of every test.
5. Pilot, measure and scale
Run a 90-day pilot on low-risk tests, track cycle time, QA defects, win rate and cost per decided test, and expand permissions only where the numbers improve. If you want help designing the operating model or connecting agents to your stack, Talk to us.
FAQ
Frequently asked questions about AI agents in experimentation
Frequently asked questions
What are AI agents in experimentation?
They are AI systems that take a goal, plan steps and use tools such as analytics, session replay and testing platforms to carry out parts of the A/B testing workflow, from research and hypotheses to building variants, monitoring and readouts, with people approving key decisions.
Can AI run A/B tests on its own?
Technically, several tools can now create, configure and launch tests, and some describe agents that make rollout decisions. We recommend against fully autonomous testing: choosing what to test, approving launches and deciding what ships should stay with named people, because agents can build broken variants and misread statistics.
Which A/B testing tools have AI agents?
As of September 2026, Optimizely (Opal agents), Kameleoon (PBX 2.0), Wingify (VWO and AB Tasty), Amplitude (Web Experimentation Agent) and Adobe (Experimentation Accelerator) offer agents for experimentation, and Contentsquare offers insight agents. Capabilities and plans vary, so check vendor documentation.
What is an MCP server in analytics and A/B testing?
An MCP server exposes a tool's data and actions through the Model Context Protocol, so AI assistants such as Claude, ChatGPT or Cursor can query or act on it. Google Analytics, Microsoft Clarity and Contentsquare offer read-only access; Amplitude, Optimizely and Wingify also allow permission-controlled writes.
Can ChatGPT or Claude analyse A/B test results?
They can summarise results well if they read numbers from your testing platform, for example through an MCP server. They should not compute significance themselves or search many metrics for a winner, because that produces false positives. A human analyst should sign off every readout.
Can AI predict which A/B test will win?
Only weakly. In a 2026 Nature study GPT-4 predicted survey experiment effects well (r = 0.85) but megastudies much less well (r = 0.34) and overestimated effect sizes, and in 17,681 Upworthy headline tests prompted LLMs did only marginally better than chance. Use AI to rank ideas, not to replace tests.
Is AI-generated A/B test code safe to launch?
Often, but not by default. In our 19-test study 11 builds reached 90% fidelity or better, while 5 scored below 70%. The same prompt can produce different code, and code can affect elements outside the test. Approve the exact build, check scope and run device QA before launch.
How do I start using AI agents in my experimentation programme?
Map your workflow, connect analytics read-only, write rules for metrics and stopping, load a learning repository, and pilot agents on research, documentation and low-risk builds for 90 days while tracking cycle time, QA defects, win rate and cost per test.
Key terms
- AI agent
- An AI system that plans steps toward a goal and uses tools to carry them out. It matters because it can act on systems, not just answer questions.
- Large language model (LLM)
- The model that powers most agents by predicting text and code. It matters because its strengths (drafting) and weaknesses (confident errors) shape what agents can be trusted with.
- Model Context Protocol (MCP)
- An open standard for connecting AI assistants to data and tools. It matters because it lets one agent work across analytics, replay and testing platforms.
- MCP server
- A connector that exposes one tool's data and actions to AI assistants. It matters because its scopes decide what an agent can read or change.
- Least privilege
- Giving an agent only the access its task needs. It matters because it limits the damage from errors and prompt injection.
- Prompt injection
- Instructions hidden in content an AI reads that change its behaviour. It matters because agents that read feedback or web pages and can also act are exposed to it.
- Prompt-based experimentation
- Building test variants by describing the change in plain language to an AI. It matters because it removes the developer bottleneck for routine tests.
- Non-determinism
- The same prompt producing different outputs on different runs. It matters because the code you reviewed must be the code that ships.
- Peeking
- Checking a fixed-horizon test repeatedly and stopping at the first significant result. It matters because it inflates false positives, and agents can peek continuously.
- Multiple comparisons
- Testing many metrics or segments and reporting whichever looks significant. It matters because the chance of a false win rises quickly with each extra comparison.
- Sample ratio mismatch (SRM)
- A gap between the planned and observed traffic split. It matters because it signals a broken experiment whose results should not be trusted.
- Learning repository
- A searchable record of every test's hypothesis, result, decision and code. It matters because it is the memory agents need to avoid repeating old tests.
- Approval gate
- A point where a named person must approve before work moves on. It matters because it keeps accountability for customer-facing changes with people.
- Agent operator
- The person who manages an agent's prompts, tools, permissions and costs. It matters because unmanaged agents drift, overspend and overreach.
Sources
All sources were checked in September 2026. Vendor capabilities, dates and limits come from each vendor's own documentation, release notes or announcements and are vendor data; features and plans change often. Deal values are as disclosed or reported by the named outlets; Datadog did not disclose the price of Eppo, so we do not quote one. The 19-test build figures come from the Henkan & Partners study of 2025 (single evaluator). Exhibits 1, 6 and 7, the step and risk tables, the roles table and the worked example are Henkan & Partners frameworks or illustrations. Exhibit 4 is a Henkan & Partners simulation and calculation.
- Anthropic (2024). Introducing the Model Context Protocol.
- Model Context Protocol blog (2025). MCP joins the Agentic AI Foundation.
- Model Context Protocol (2026). Security best practices.
- OWASP GenAI Security Project (2025). LLM01:2025 Prompt Injection.
- Optimizely (2026). 2026 Optimizely Opal release notes.
- Optimizely (2026). AI variation development agent.
- Optimizely (2026). Experiment Planning agent.
- Optimizely (2026). Experiment Summary agent.
- Optimizely (2026). Remote MCP Server is here: bring Optimizely into your AI tools.
- Kameleoon (2025). Introducing prompt-based experimentation.
- Kameleoon (2026). PBX 2.0 is changing testing again.
- Kameleoon (2026). Expanding prompt-based experimentation with PBX Ideate.
- Kameleoon (2026). PBX Ship turns winning tests into production code.
- Kameleoon (2026). Prompt-based experimentation plans.
- Wingify (2025). Create optimization campaigns by prompting with Copilot.
- Wingify (2026). Agentic AI for experimentation: hype vs. reality.
- Wingify (2026). Wingz AI.
- Wingify (2026). Connect Wingify with AI tools using Wingify's MCP server.
- Amplitude (2026). Amplitude introduces agentic AI analytics.
- Amplitude (2026). Amplitude MCP server documentation.
- Amplitude (2026). Amplitude and Statsig partnership.
- Contentsquare (2025, updated 2026). Contentsquare's MCP: bridging agents and experience data.
- Contentsquare (2026). Contentsquare launches AI agent and ChatGPT analytics capabilities.
- Adobe (2025). Introducing Adobe Journey Optimizer Experimentation Accelerator.
- Google (2025). Google Analytics MCP server and Try the Google Analytics MCP server.
- PPC Land (2025). Google Analytics experimental MCP server enables AI conversations with data.
- Microsoft (2026). Microsoft Clarity MCP server.
- PPC Land (2026). Microsoft bets on AI economy with Web IQ, Clarity citations and MCP server.
- TechCrunch (2025). Datadog acquires Eppo, a feature flagging and experimentation platform.
- TechCrunch (2025). OpenAI acquires product testing startup Statsig.
- TechCrunch (2026). Everstone combines Wingify and AB Tasty.
- Ashokkumar, A., Hewitt, L., Ghezae, I. & Willer, R. (2026). Large language models can predict the results of social science experiments. Nature 656. Summary figures from treatmenteffect.app.
- Ye, Z., Yoganarasimhan, H. & Zheng, Y. (2024). LOLA: LLM-assisted online learning algorithm for content experiments. arXiv; Marketing Science (2025).
- METR (2025). Measuring the impact of early-2025 AI on experienced open-source developer productivity.
- Ouyang, S., Zhang, J. M., Harman, M. & Wang, M. (2023). An empirical study of the non-determinism of ChatGPT in code generation. arXiv.
- Miller, E. (2010). How not to run an A/B test.
- Fabijan, A. et al. (2019). Diagnosing sample ratio mismatch in online controlled experiments. KDD 2019.
- Kohavi, R., Deng, A. & Vermeer, L. (2022). A/B testing intuition busters. KDD 2022.
- Kohavi, R., Tang, D. & Xu, Y. (2020). Trustworthy Online Controlled Experiments. Cambridge University Press.
- Henkan & Partners (2025, revised 2026). How AI is reshaping A/B testing: what 19 prompt-built tests tell us.