Focus

AI Agents in Experimentation: What They Do at Each Step, What Vendors Ship and Where Humans Stay in Charge

Alexandre Suon · 2026-09-28

AI agents in experimentation have moved from demos to shipped features in under two years. Optimizely, Kameleoon, Wingify (VWO and AB Tasty), Amplitude and Contentsquare now offer agents that research, write hypotheses, build variants, read results and connect to your data through the Model Context Protocol (MCP). This deep dive walks through the A/B testing workflow step by step, shows what each vendor actually ships, and sets out where human judgement must stay in charge and the operating model that keeps agents fast and safe.

Executive summary

  1. Agents can now help at every step of an A/B test, but three decisions must stay human. Research, ideation, planning, building, QA, monitoring, analysis and documentation can all be drafted or done by an agent. Choosing what to test, approving a launch and deciding what to ship should stay with named people.
  2. The vendors have shipped, not just announced. Between February and September 2026 Optimizely added Opal agents for ideas, conflict checks, prioritisation, value estimates and governance; Kameleoon released four PBX 2.0 agents; Amplitude launched a Web Experimentation Agent; Wingify lists nine agents across VWO and AB Tasty. Most also offer an MCP server.
  3. Building is solved enough; judging is not. In our study of 19 prompt-built tests, average build time fell from 8.7 hours to 55 minutes, but 5 of 19 builds scored below 70% fidelity. In a Nature study, GPT-4 predicted survey experiment effects well (r = 0.85) but large megastudies much less well (r = 0.34) and overestimated effect sizes.
  4. Statistics is where agents do the most damage. An agent that checks a test ten times and stops at the first 'win' turns a 5% false positive rate into about 20% in our simulation; one that scans 20 metrics has a 64% chance of a false win. Agents must read results from the testing platform's engine and follow pre-registered stopping rules.
  5. MCP is the connector that makes agents useful, and the main new risk. Google Analytics and Contentsquare offer read-only servers; Amplitude, Optimizely and Wingify also allow writes, controlled by permissions. Start read-only, grant write access by exception, and log every action, as OWASP and the MCP specification recommend.
  6. Run agents like junior analysts inside a clear operating model. Four roles (strategist, reviewer, agent operator, decision owner), four approval gates and a learning repository that agents read before every idea and update after every decision. A 90-day plan moves from read-only access to gated actions.

Section 1 · Definition

An AI agent plans and takes steps with tools, which makes it different from a chatbot

AI agents in experimentation are software systems built on large language models that take a goal (for example, "find and test ways to reduce checkout drop-off"), plan the steps, and use tools such as analytics queries, session replay, a testing platform or a code editor to carry them out, with a person reviewing the output at defined points.

The difference from a chat assistant is agency. An assistant answers a question you ask. An agent decides which data to pull, runs the query, reads the result, drafts the next step and, if allowed, acts on a system. That is why the same technology that saves hours can also launch a broken variant or declare a false winner without anyone noticing.

This article is the fourth step of our Optimise with AI learning path. Our study of 19 prompt-built tests measured what AI does to the build step, and our C-level agenda for prompt-based experimentation explained why leaders should care. Here we widen the lens to the whole workflow, from research to the learning repository.

Three terms recur throughout:

  • Large language model (LLM). The model underneath the agent, such as GPT, Claude or Gemini. It predicts text, and with the right tools it can write code, call APIs and summarise data.
  • Tool use. The agent's ability to call functions: run a GA4 report, list live experiments, take a screenshot of a page, create a draft test. Tools are what turn a model into an agent.
  • Model Context Protocol (MCP). An open standard, released by Anthropic in November 2024, for connecting AI assistants to the systems where data lives. Most testing and analytics vendors now publish an MCP server, so one agent can reach several tools through a common interface.

Our view. Treat an agent as a very fast junior analyst with no memory of your business and no instinct for when a number looks wrong. It is excellent at drafting, searching and repetitive checks. It needs a brief, access limited to its task, and a senior reviewer before anything reaches customers.

Section 2 · The workflow

Agents can help at all eleven steps, but only some should run without a person in the loop

An A/B test is not one task. It is a chain of eleven steps, each with its own inputs, risks and owner. Asking "should we use AI for testing?" is too broad. The useful question is: for each step, how much can an agent do, and who checks it?

Flow diagram of eleven A/B testing steps in two rows. Agent drafts and human edits: 1 Research, 2 Ideation, 4 Design and sample size, 6 QA, 9 Analyse. Agent does and human checks: 5 Build or code, 8 Monitor, 11 Document and learn. Human decides and agent prepares: 3 Hypothesis, 7 Launch, 10 Decide. Step 11 feeds back to research.
Exhibit 1. Agents can help at every step of an A/B test; three decisions stay human. Source: Henkan & Partners framework, based on vendor documentation reviewed in September 2026 and our 2025 study of 19 prompt-built tests.

What this shows. Three colours, three levels of trust. Green steps are routine and easy to verify, so an agent can do the work and a person checks it. Blue steps need judgement about the business, so the agent drafts and a person edits. The three orange steps are commitments: which idea gets scarce traffic, whether a variant is safe to show customers, and whether to ship. Those stay with named people, with the agent preparing the evidence.

StepWhat an agent can do todayWhat a person must doMain risk
1. ResearchSummarise analytics, session replays, surveys and reviews; flag frictionCheck the sample and the source; add business contextConfident summaries of thin data
2. IdeationPropose ideas from page scans and past testsReject generic ideas; link each to evidencePlausible but unoriginal ideas
3. HypothesisDraft hypothesis, metrics, risksChoose which ideas get trafficPlausibility mistaken for evidence
4. Design and sample sizeSuggest metrics, audiences, duration, sample sizeConfirm the primary metric and minimum detectable effectWrong baseline or metric
5. BuildWrite variant code from a prompt or a designReview the code that will shipGlobal CSS, broken logic
6. QACheck conflicts with live tests, run scripted checksVisual and device QA; sign-offUnchecked edge cases
7. LaunchConfigure targeting, goals and traffic splitApprove the launchWrong audience or split
8. MonitorWatch for sample ratio mismatch and guardrail breachesAct on alertsAlert fatigue
9. AnalyseDraft the readout from the platform's statisticsCheck the method and segmentsHallucinated or cherry-picked numbers
10. DecideRecommend ship, iterate or stop; estimate valueMake the callAutomated shipping of false wins
11. DocumentWrite the learning to the repositoryApprove the learning and tagsA repository full of noise

The pattern is simple. The closer a step is to customers or to a business decision, the more human control it needs. The more a step is about gathering, drafting or checking against a rule, the more an agent can take on. For the fundamentals behind each step, see our essential guide to A/B testing.

Section 3 · Research and hypotheses

Agents speed up research and ideation, but a plausible hypothesis is not evidence

Research synthesis is the safest place to start. The agent reads data that already exists, a person can check the summary against the source, and a weak summary costs time, not customers.

Research synthesis: what ships today

  • Contentsquare describes its Sense Analyst agent as working "24/7 to analyze experience data, detect issues and growth opportunities" (March 2026), and exposes the same analysis through a read-only MCP server launched in October 2025.
  • Amplitude launched a Session Replay Agent that reviews sessions to spot friction and an AI Feedback Agent that turns unstructured feedback into insights (February 2026).
  • Wingify, the group formed when VWO and AB Tasty combined, lists an Insight Agent that analyses funnels, replays and surveys, and a Synthetic Testing Agent that tries changes on synthetic visitor profiles before a live test.
  • Microsoft Clarity and Google Analytics do not ship research agents of their own for this purpose, but both publish MCP servers, so any MCP-compatible assistant can query them.

Two cautions apply. First, an agent summarising 40 session replays will describe them with the same confidence as 4,000. Ask it to state its sample and to quote the sessions it used. Second, synthetic users and AI-generated research are good for generating hypotheses, not for measuring effects. Our guide to synthetic users for customer research explains where they help and where they mislead.

Hypothesis and prioritisation: fast drafts, human choice

Every major vendor now offers an ideation or planning agent:

  • Optimizely Opal has an Experiment Planning agent that turns a URL and a test idea into "a hypothesis, variant descriptions, targeting settings, a metrics table, statistical sizing guidance, risks, and assumptions"; an Idea Builder (April 2026) that generates concepts from page context and research; and a Backlog Prioritization agent (June 2026) that scores ideas with the PIE framework.
  • Kameleoon PBX Ideate (May 2026) scans a page and ranks test ideas using data from more than 20,000 historical experiments, then turns them into prompts for the build agent.
  • Wingify lists a Hypothesis Agent that converts evidence into testable ideas with ranked variations; its earlier AB Tasty Evi suite included Evi Hypothesize, which scores hypothesis quality.

These tools produce well-formed hypotheses in seconds. The question is whether they pick good ones. The best published evidence says: better than you might expect in controlled surveys, much worse in the field.

Two bar charts. Left: correlation of GPT-4 predicted with actual experimental effects, r = 0.85 for survey experiments, r = 0.34 for megastudies, compared with r = 0.26 for human expert forecasters on megastudies. Right: developer speed with AI tools in a METR randomised trial of 16 developers, expected plus 24%, felt afterwards plus 20%, measured minus 19%.
Exhibit 2. AI judgement looks better than it is, so keep testing. Source: Ashokkumar, Hewitt, Ghezae & Willer, Nature (2026); METR (2025); Ye, Yoganarasimhan & Zheng (2024).

What this shows. In 70 preregistered survey experiments with 469 effects, GPT-4's predictions correlated strongly with the real results (r = 0.85). In 15 megastudies with 606 effects, closer to real-world behaviour, the correlation fell to 0.34, still slightly ahead of expert forecasters at 0.26, and the model systematically overestimated effect sizes. In 17,681 Upworthy headline tests, prompted LLMs picked winners only marginally better than chance. The right panel is a reminder that AI help also feels more productive than it measures.

The practical conclusion is to use agents to widen and rank the list, never to skip the test. An agent's confidence score is a prior, not a result. A human strategist should still choose which ideas get scarce traffic, and every hypothesis should cite the evidence behind it: a funnel drop, a replay pattern, a survey theme or a past test.

For marketers. Ask the ideation agent for ten ideas, then ask it to show the evidence for each. Keep the ones where the evidence is your own data. Drop the ones where the evidence is "best practice".

For leaders. Measure the win rate of agent-suggested ideas separately from human ones for the first two quarters. If agent ideas win less often, you are saving time on ideation and spending it on failed tests.

Section 4 · Build and QA

Agents cut build time sharply, so QA becomes the new bottleneck

Building variants was the first step to be automated, and it is the best measured. In July and August 2025 we rebuilt 19 historical, hand-coded A/B tests with Kameleoon's Prompt-Based Experimentation (PBX), using prompts only.

Left: average build time per test, 8.7 hours hand-built by a developer versus 55 minutes prompt-built including iterations. Right: primary cause of shortfalls against the hand-built original, visual precision 52%, conditional business logic 31%, non-deterministic output 17%. 11 of 19 builds reached 90% fidelity or better; 5 of 19 scored below 70%.
Exhibit 3. Prompt-built variants are fast, but one in four needs real rework. Source: Henkan & Partners study of 19 historical A/B tests rebuilt with Kameleoon PBX, July and August 2025 (single evaluator).

What this shows. Speed was never the problem: every test was built faster, and build time fell by about 89% on average. Quality was uneven. Most simple changes were close to perfect, but five builds needed substantial manual work, mainly because words are a poor way to describe a pixel-exact layout or a business rule. That is why the QA step now matters more than the build step.

What build agents ship today

  • Optimizely Opal's AI variation development agent modifies and creates page elements in the new Visual Editor from natural language and retrieves page styles to stay on brand. Optimizely notes that a task uses 30 to 130 Opal credits and that "any code Opal generates is a custom code solution and falls outside the scope of Optimizely Support".
  • Kameleoon PBX 2.0 (April 2026) has a Build agent that browses the site to understand components in context and can convert Figma designs into variations, plus a Configure agent for audiences, goals and launch.
  • Wingify Copilot creates an entire experiment, including variations, metrics and audience, from a plain-English description (August 2025, Pro and Enterprise plans).
  • Amplitude's Web Experimentation Agent (February 2026) is described as designing and launching experiments and making rollout decisions.

Why AI-generated code needs stricter QA, not looser

Three properties of LLM-written code change the QA job:

  1. The same prompt can give different code. A study of ChatGPT code generation by Ouyang and colleagues found that for 48% to 76% of tasks, depending on the benchmark, repeated requests produced no two outputs with the same test results, and setting temperature to zero reduced but did not remove this. Approve the exact code that ships and never regenerate after approval.
  2. Code can reach beyond its target. In our study, one variant wrote site-wide CSS that broke dropdown menus elsewhere on the page. Add a scope check for global styles and regressions on nearby elements.
  3. Fast feels productive. In METR's 2025 randomised trial, experienced developers using AI tools took 19% longer while believing they were 20% faster. Measure cycle time from brief to launch, not build time alone.

Agents also help with QA itself. Optimizely's Experiment Conflict Checker agent (April 2026) checks a new test against every live experiment in Web Experimentation, Performance Edge or Personalization. That kind of rule-based check suits an agent well. Visual QA on real devices, and the final sign-off, remain human jobs. For the code patterns that make variants robust, see our developer's guide to DOM manipulation for A/B tests.

Section 5 · Analysis and readouts

Let agents draft the readout, but never let them choose when to stop or which metric won

Analysis agents are the most tempting and the most dangerous. They turn a dense results page into a clear paragraph for executives. They can also produce numbers that look right and are not.

What ships today: Optimizely's Experiment Summary agent produces a condensed report that highlights key findings, whether results reached statistical significance and next steps. Wingify's AB Tasty Evi suite included Evi Analysis for post-test insights. Adobe's Journey Optimizer Experimentation Accelerator, launched in September 2025, uses an Experimentation Agent to explain why tests win or lose and rank new opportunities. Amplitude's MCP server lets any assistant read and create experiments and charts.

Three statistical failure modes to design out

  • Hallucinated statistics. An LLM asked to compute a p-value or confidence interval may do the arithmetic itself and get it wrong, or invent a figure that was never in the data. Rule: the agent must quote numbers returned by the testing platform's statistics engine, with a link to the source, and must not compute significance on its own.
  • Peeking. An agent monitoring a test around the clock is the ultimate peeker. Checking a fixed-horizon test repeatedly and stopping at the first significant result inflates false positives, a problem Evan Miller described in 2010. Rule: use a sequential method designed for continuous monitoring, or let the agent alert but not stop the test.
  • Metric fishing. Ask an agent "did anything improve?" and it will search every metric and segment until something does. Rule: pre-register one primary metric and a small set of guardrails before launch, and label everything else as exploratory.
Two bar charts of the chance of at least one false significant result when there is no real effect, at a nominal 5% level. Left, peeking: 1 look 5%, 2 looks 9%, 5 looks 14%, 10 looks 20%, 20 looks 25%, 50 looks 33%, 100 looks 38%. Right, multiple comparisons across independent metrics: 1 metric 5%, 3 metrics 14%, 5 metrics 23%, 10 metrics 40%, 20 metrics 64%.
Exhibit 4. Unchecked peeking and metric fishing multiply false wins. Source: Henkan & Partners Monte Carlo simulation of 20,000 A/A tests (left); family-wise error rate formula (right); consistent with Evan Miller (2010).

What this shows. A test with no real effect should show a false win 5% of the time. Look ten times and stop at the first win, and it happens about 20% of the time in our simulation. Check 20 independent metrics and the chance that at least one looks like a winner is 64%. An agent does both faster and more often than any analyst, so the rules must be built into its instructions and its tools.

Family-wise false positive rate = 1 - (1 - alpha)^k with alpha = 0.05 and k = 20 independent metrics: 1 - 0.95^20 = 0.64

Data quality checks are a good fit for agents because they follow rules. A sample ratio mismatch (SRM), where the observed split between variants differs from the planned split, signals a broken experiment. Microsoft researchers reported that about 6% of experiments at Microsoft exhibit an SRM. An agent that checks for SRM daily and blocks the readout until it is resolved adds real safety. For the statistical models behind these rules, see our guide to A/B testing statistics for marketers.

Our view. The readout should be a draft that a named analyst signs. The ship decision should never be automated on the basis of a single agent-read result, whatever the vendor setting allows. If an agent can make rollout decisions, gate that permission behind a human approval and a pre-registered rule.

Section 6 · Learning repository

A learning repository is what turns fast agents into smart ones

Agents have no memory of your business unless you give them one. Without a record of past tests, an ideation agent will happily propose the free-shipping banner you tested and lost last year. The learning repository (a searchable record of every test's hypothesis, evidence, design, result, decision and code) is that memory.

Kohavi, Tang and Xu's standard text on online experiments devotes a chapter to institutional memory and meta-analysis for this reason: past results are the best guide to future ideas. Agents raise the stakes in two ways. They can read the repository before proposing anything, and they can write to it after every decision, which removes the chore that usually lets repositories decay.

Vendors are moving here too. Optimizely's Experimentation Program Overview agent (June 2026) reports quarterly on the testing programme, and its Experiment Value Estimator agent projects annualised impact from winning experiments. Kameleoon's PBX Ship (May 2026) turns a winning client-side variation into production code behind a feature flag through its MCP server, so the tested code and the shipped code stay linked.

What to store so an agent can use it

FieldWhy the agent needs it
Hypothesis and evidenceLets the agent check whether a new idea repeats an old one
Page, audience, lever tagsLets it find related tests and patterns by page and theme
Primary metric, guardrails, sample, durationLets it judge how much weight a past result deserves
Result with interval, and decisionSeparates 'lost' from 'inconclusive', which mean different things
Final prompt and shipped codeLets a winner be rebuilt or handed to developers
Learning in one sentence, approved by a personStops agent-written noise from polluting the memory

Disclosure: Henkan & Partners builds its own AI and MCP tooling for clients' experimentation programmes, and works with several of the testing vendors named in this article.

Section 7 · What vendors ship

Every major testing vendor now ships agents, but they differ in where the agent may act

Twenty-two months after MCP was released, agents and MCP connections reach every layer of the experimentation stack. The pace has been set by launches and by consolidation: Datadog bought Eppo in May 2025, OpenAI bought Statsig for $1.1 billion in September 2025, Everstone combined Wingify and AB Tasty in January 2026, and Amplitude took over the Statsig brand, platform and customers in May 2026.

Timeline from November 2024 to September 2026. MCP and analytics access: MCP released by Anthropic (Nov 2024), Microsoft Clarity MCP server (Jun 2025), Google Analytics MCP server, read-only (Jul 2025), Contentsquare MCP server (Oct 2025), MCP moves to the Linux Foundation's Agentic AI Foundation (Dec 2025), Optimizely MCP server (Apr 2026). Agent features: Kameleoon PBX (Jun 2025), VWO Copilot builds campaigns from a prompt (Aug 2025), Amplitude agents (Feb 2026), Kameleoon PBX 2.0 (Apr 2026), Opal Idea Builder (Apr 2026), PBX Ship via MCP (May 2026), Opal value and prioritisation agents (Jun 2026), Opal governance and flag agents (Sep 2026). Deals: Datadog buys Eppo (May 2025), OpenAI buys Statsig for $1.1B (Sep 2025), Wingify and AB Tasty combine (Jan 2026), Amplitude takes over Statsig platform (May 2026).
Exhibit 5. In 22 months, agents and MCP reached every layer of the experimentation stack. Source: Anthropic, MCP blog, PPC Land, Kameleoon, Wingify, Contentsquare, TechCrunch, Amplitude and Optimizely announcements; dates as announced.

What this shows. The first year was about access: MCP servers let assistants read analytics. The second was about action: vendors added agents that build, configure, prioritise and ship. The deals show experimentation becoming core product infrastructure, owned by analytics and developer platforms rather than sold as a marketing add-on.

VendorAgents shipped (as documented)MCP serverCan the agent change things?
Optimizely (Opal)Experiment Planning, Idea Builder, AI variation development, Conflict Checker, Experiment Summary, Backlog Prioritization, Value Estimator, Program Overview, Governance, Feature Flag ImplementationYes, remote, launched April 2026Yes: creates and configures flags and experiments
Kameleoon (PBX 2.0)Ideate, Build, Configure, ShipYes, used by PBX ShipYes: builds, configures, ships behind feature flags
Wingify (VWO, AB Tasty)Nine listed agents incl. Insight, Hypothesis, Synthetic Testing, Experiment, Rollout, Guardrail; earlier Evi suite and CopilotYes: 10 read tools, 4 write toolsYes: creates A/B and split URL tests, by permission level
AmplitudeGlobal, Dashboard Monitoring, Session Replay, Web Experimentation, AI FeedbackYesYes: reads and writes charts, experiments, cohorts, with separate write permission
ContentsquareSense Analyst, configurable insight agentsYes, read-onlyNo: read-only data access
AdobeExperimentation Agent in Journey Optimizer Experimentation AcceleratorNot reviewedAnalyses and recommends
Google Analytics, Microsoft ClarityNo experimentation agentsYes, read-onlyNo

Three differences matter more than feature lists. First, scope of action: some agents only read and recommend, others create and launch tests. Second, pricing model: an Opal variation build uses 30 to 130 credits, Kameleoon PBX typically uses one to two credits per experiment, and Wingify says its Evi features were included in all contracts at no additional cost; model the cost at your test volume. Third, openness: an MCP server lets you use your own assistant across tools, while an in-product agent is easier to govern but locks the workflow to one vendor. For the wider market picture, see our report on the A/B testing tool market.

For leaders. Do not choose a platform for its agents alone. Agents are converging quickly; your data model, statistics engine and governance are harder to change. Ask each vendor what the agent can do without approval, what is logged, and how you switch the write permission off.

Section 8 · MCP connections

MCP makes agents useful by connecting them to your data, and risky for exactly the same reason

MCP is an open standard for connecting AI assistants to the systems where data lives. An MCP server wraps a tool, such as GA4 or a testing platform, and publishes a list of actions the assistant may call. An MCP client, such as Claude, ChatGPT, Cursor or Copilot, discovers those actions and uses them. In December 2025 Anthropic moved MCP to the Agentic AI Foundation under the Linux Foundation; at the time the project counted more than 10,000 active servers and 97 million monthly SDK downloads.

For an experimentation team, MCP means one assistant can pull a funnel from GA4, pull behaviour metrics such as scroll depth and friction from Clarity or Contentsquare, check live tests in Optimizely or Wingify, and draft a hypothesis, without exporting files between tools. The permissions differ sharply by server:

MCP serverAccessNotable limits (as documented)
Google AnalyticsRead-only (analytics.readonly scope)Labelled experimental; reports, funnels and real-time data; cannot change configuration
Microsoft ClarityQuery only10 API requests per project per day, up to 3 days of data, 3 dimensions per request
ContentsquareRead-onlyExperience, friction, conversion and journey data
AmplitudeRead and writeSeparate 'Use MCP (read)' and 'Use MCP (write)' permissions; uses existing user access controls
Optimizely ExperimentationRead and writeQueries results, flags and audiences; creates and configures flags and experiments
WingifyRead and writeBrowse, Design and Publish/Admin permission tiers; AI can create only A/B and split URL tests

Security: prompt injection and over-broad permissions

Connecting an agent to live systems adds two security risks that most experimentation teams have not faced before:

  • Indirect prompt injection. OWASP ranks prompt injection first in its 2025 Top 10 for LLM applications. The indirect form happens when the model reads external content, such as a web page, a survey answer or a support ticket, that contains hidden instructions. An agent that reads open-text feedback and also has write access to your testing tool is exactly that pattern. OWASP recommends restricting the model's privileges to the minimum necessary and requiring human approval for high-risk actions.
  • Over-broad scopes. The MCP security best-practice guidance warns against wildcard or "full-access" scopes and recommends a progressive, least-privilege model: start with low-risk read operations and elevate only when a privileged action is first needed, with elevation events logged.
Reference architecture in four layers. People: experimentation strategist, reviewer, agent operator, decision owner. Four approval gates between people and agents: hypothesis, code and QA, launch, decision. Agents: research, ideation and planning, coding, monitoring, analysis and reporting. MCP layer: one connector per system, least-privilege scopes, read first and write by exception, full action log. Systems: data warehouse, analytics and behaviour tools, testing and flags, learning repository, code repository and CI.
Exhibit 6. Agents reach data through governed MCP connections; people hold the gates. Source: Henkan & Partners framework (2026), using permission models documented by Amplitude, Wingify and Google Analytics.

What this shows. Agents sit between people and systems, never above them. Every system connects through its own MCP server with the narrowest scope that does the job. The four orange gates are the only paths from an agent's draft to a customer-facing change. The learning repository is both a source and a destination, so every test makes the next one better informed.

For marketers. You can start today with read-only servers: connect GA4 and Clarity or Contentsquare to your assistant and ask it to explain a funnel drop with behaviour data. No write access is needed to get most of the research value.

For leaders. Put MCP connections through the same approval as any other integration: an owner, a scope, a data-processing check and an audit log. Consent rules still apply to what the agent reads; see our guide to user consent in e-commerce.

Section 9 · Human judgement and risks

Human judgement stays essential for what to test, what is safe and what counts as a win

The three orange steps in Exhibit 1 share a feature: each is a commitment made under uncertainty that someone must be accountable for. An agent can prepare the evidence, but it cannot own the consequences.

  • What to test. Traffic is the scarcest resource in most programmes. Choosing which hypothesis gets it depends on strategy, brand, seasonality and politics that no agent sees.
  • Whether it is safe to launch. A variant that breaks checkout on an older phone costs revenue and trust. The launch approval is a promise to customers.
  • What counts as a win and what ships. Most tested ideas do not win: Kohavi, Deng and Vermeer report success rates of about 33% at Microsoft, 15% at Bing, 10% at Booking.com, Google Ads and Netflix, and 8% in Airbnb Search. With base rates that low, a significant result is often a false positive, so the ship decision needs someone who understands the method.
RiskWhat it looks likeControl
Hallucinated statisticsA readout quotes a lift or p-value that is not in the platformAgent may only quote engine output, with a link; analyst signs off
Peeking and metric fishingAgent stops at the first 'win' or finds a winning segmentPre-registered primary metric and stopping rule; sequential methods for monitoring
Code qualityGlobal CSS, broken logic, different code on regenerationApprove the exact build; scope check; two-person QA for high-traffic pages
Prompt injectionHidden instructions in feedback or web content steer the agentLeast privilege; separate read and write agents; human approval for actions
Governance and privacyAgent reads personal data or acts outside its remitScoped MCP access, consent checks, action logs, named owner per agent
Generic ideasBacklog fills with 'best practice' testsEvery hypothesis cites first-party evidence
Cost creepCredits and tokens grow with volumeBudget per test; track cost per decided test
Over-trustPeople stop checking because the agent is usually rightRandom audits of agent output; track error rates

Culture matters as much as controls. Teams that already reward learning over winning adapt well to agents, because they are used to challenging results. Our article on building a culture of experimentation covers the habits that make that possible.

Section 10 · Operating model

Run agents like junior analysts: clear roles, earned permissions and measured output

The operating model is where most of the value is won or lost. The tools are converging; the difference between programmes will be who owns what, which permissions the agents hold, and whether anyone measures the results.

Four roles, even in a small team

RoleOwnsIn a small teamIn a large team
Experimentation strategistGoals, backlog, choice of what to testHead of CRO or e-commerce managerProgramme lead per product area
ReviewerCode, QA, statistics, brand and legal checksSame person with a checklist, plus an external reviewer for statsNamed QA and analytics reviewers
Agent operatorPrompts, tools, MCP permissions, cost, logsAnalyst or developer, a few hours a weekDedicated role in the experimentation platform team
Decision ownerShip, iterate or stopBusiness owner of the pageProduct owner, with finance for high-value decisions

Permissions are earned, and results are measured

Give agents access in three stages: see, draft, act. In the first month they read analytics, replays and past results. In the second they draft hypotheses, plans, variants and readouts that people edit. Only in the third, and only for low-risk tests, do they take actions such as creating a draft test, always behind the approval gates. Track four numbers from the start, so you can tell whether agents help:

  • Cycle time from brief to decided test.
  • QA defect rate: variants that fail QA or are rolled back after launch.
  • Win rate and learning rate of agent-suggested versus human-suggested ideas.
  • Cost per decided test, including credits, tokens and review time.
Gantt chart of a 90-day plan in three phases: days 1 to 30 See, days 31 to 60 Draft, days 61 to 90 Act with gates. Workstreams: baseline KPIs in weeks 0 to 2; read-only data access weeks 0.5 to 4 then scoped write for drafts weeks 8.5 to 12.5; learning repository loaded weeks 1 to 5 then agents read before and write after from week 5; research and ideation drafts weeks 2 to 6.5; coding agent on low-risk tests weeks 4.5 to 8.5 then two-person QA gate; pre-registered statistical rules weeks 3 to 6 then agent alerts for SRM and guardrails; permissions and audit log weeks 0 to 3, review and scale weeks 9 to 13.
Exhibit 7. Earn write access in 90 days: see first, then draft, then act. Source: Henkan & Partners framework (2026).

What this shows. The plan sequences trust. Measurement and permissions come first, because without a baseline you cannot tell whether agents help. The learning repository is loaded early, because every later agent depends on it. Write access arrives last, scoped to drafts, with two-person QA before anything reaches customers.

Worked example: one test, agent-assisted (illustrative)

Illustrative example, not a client case. A fashion retailer sees mobile checkout drop-off rise. The research agent pulls the GA4 funnel through the read-only MCP server and finds the drop concentrated at the delivery step; it then queries Clarity for scroll depth and engagement on that step. The strategist watches 30 session replays, sees shoppers scrolling past the delivery options, confirms the pattern and chooses the idea: show delivery cost and date above the fold. The planning agent drafts the hypothesis, primary metric (checkout completion), guardrail (average order value) and sample size; the analyst corrects the baseline. The build agent writes the variant; the reviewer rejects the first version for a global style change, approves the second and signs QA. The monitoring agent checks SRM daily. At the planned end date the analysis agent drafts the readout from the platform's numbers; the decision owner ships. The agent writes the learning, and the strategist approves it.

People made four decisions in that example. Agents did most of the typing.

Section 11 · What to do next

What to do next

1. Map your workflow and pick two agent-ready steps

List your eleven steps, who owns each and how long each takes. Start agents where the work is slow, rule-based and easy to verify: research synthesis and documentation are usually the best first candidates.

2. Connect your data read-only

Set up read-only MCP connections to your analytics and behaviour tools. Give one person the agent operator role, with a log of every query. Most of the research value comes before any write access.

3. Write the rules before the agents run

Pre-register primary metrics and stopping rules, require agents to quote only platform statistics, and define the four approval gates. Put them in the agent's instructions and in your test template.

4. Build the learning repository

Load past tests with hypotheses, results, decisions and code. Make "read the repository first" the first instruction of every ideation agent, and "write the learning" the last step of every test.

5. Pilot, measure and scale

Run a 90-day pilot on low-risk tests, track cycle time, QA defects, win rate and cost per decided test, and expand permissions only where the numbers improve. If you want help designing the operating model or connecting agents to your stack, Talk to us.

FAQ

Frequently asked questions about AI agents in experimentation

Frequently asked questions

What are AI agents in experimentation?

They are AI systems that take a goal, plan steps and use tools such as analytics, session replay and testing platforms to carry out parts of the A/B testing workflow, from research and hypotheses to building variants, monitoring and readouts, with people approving key decisions.

Can AI run A/B tests on its own?

Technically, several tools can now create, configure and launch tests, and some describe agents that make rollout decisions. We recommend against fully autonomous testing: choosing what to test, approving launches and deciding what ships should stay with named people, because agents can build broken variants and misread statistics.

Which A/B testing tools have AI agents?

As of September 2026, Optimizely (Opal agents), Kameleoon (PBX 2.0), Wingify (VWO and AB Tasty), Amplitude (Web Experimentation Agent) and Adobe (Experimentation Accelerator) offer agents for experimentation, and Contentsquare offers insight agents. Capabilities and plans vary, so check vendor documentation.

What is an MCP server in analytics and A/B testing?

An MCP server exposes a tool's data and actions through the Model Context Protocol, so AI assistants such as Claude, ChatGPT or Cursor can query or act on it. Google Analytics, Microsoft Clarity and Contentsquare offer read-only access; Amplitude, Optimizely and Wingify also allow permission-controlled writes.

Can ChatGPT or Claude analyse A/B test results?

They can summarise results well if they read numbers from your testing platform, for example through an MCP server. They should not compute significance themselves or search many metrics for a winner, because that produces false positives. A human analyst should sign off every readout.

Can AI predict which A/B test will win?

Only weakly. In a 2026 Nature study GPT-4 predicted survey experiment effects well (r = 0.85) but megastudies much less well (r = 0.34) and overestimated effect sizes, and in 17,681 Upworthy headline tests prompted LLMs did only marginally better than chance. Use AI to rank ideas, not to replace tests.

Is AI-generated A/B test code safe to launch?

Often, but not by default. In our 19-test study 11 builds reached 90% fidelity or better, while 5 scored below 70%. The same prompt can produce different code, and code can affect elements outside the test. Approve the exact build, check scope and run device QA before launch.

How do I start using AI agents in my experimentation programme?

Map your workflow, connect analytics read-only, write rules for metrics and stopping, load a learning repository, and pilot agents on research, documentation and low-risk builds for 90 days while tracking cycle time, QA defects, win rate and cost per test.

Key terms

AI agent
An AI system that plans steps toward a goal and uses tools to carry them out. It matters because it can act on systems, not just answer questions.
Large language model (LLM)
The model that powers most agents by predicting text and code. It matters because its strengths (drafting) and weaknesses (confident errors) shape what agents can be trusted with.
Model Context Protocol (MCP)
An open standard for connecting AI assistants to data and tools. It matters because it lets one agent work across analytics, replay and testing platforms.
MCP server
A connector that exposes one tool's data and actions to AI assistants. It matters because its scopes decide what an agent can read or change.
Least privilege
Giving an agent only the access its task needs. It matters because it limits the damage from errors and prompt injection.
Prompt injection
Instructions hidden in content an AI reads that change its behaviour. It matters because agents that read feedback or web pages and can also act are exposed to it.
Prompt-based experimentation
Building test variants by describing the change in plain language to an AI. It matters because it removes the developer bottleneck for routine tests.
Non-determinism
The same prompt producing different outputs on different runs. It matters because the code you reviewed must be the code that ships.
Peeking
Checking a fixed-horizon test repeatedly and stopping at the first significant result. It matters because it inflates false positives, and agents can peek continuously.
Multiple comparisons
Testing many metrics or segments and reporting whichever looks significant. It matters because the chance of a false win rises quickly with each extra comparison.
Sample ratio mismatch (SRM)
A gap between the planned and observed traffic split. It matters because it signals a broken experiment whose results should not be trusted.
Learning repository
A searchable record of every test's hypothesis, result, decision and code. It matters because it is the memory agents need to avoid repeating old tests.
Approval gate
A point where a named person must approve before work moves on. It matters because it keeps accountability for customer-facing changes with people.
Agent operator
The person who manages an agent's prompts, tools, permissions and costs. It matters because unmanaged agents drift, overspend and overreach.

Sources

All sources were checked in September 2026. Vendor capabilities, dates and limits come from each vendor's own documentation, release notes or announcements and are vendor data; features and plans change often. Deal values are as disclosed or reported by the named outlets; Datadog did not disclose the price of Eppo, so we do not quote one. The 19-test build figures come from the Henkan & Partners study of 2025 (single evaluator). Exhibits 1, 6 and 7, the step and risk tables, the roles table and the worked example are Henkan & Partners frameworks or illustrations. Exhibit 4 is a Henkan & Partners simulation and calculation.

  1. Anthropic (2024). Introducing the Model Context Protocol.
  2. Model Context Protocol blog (2025). MCP joins the Agentic AI Foundation.
  3. Model Context Protocol (2026). Security best practices.
  4. OWASP GenAI Security Project (2025). LLM01:2025 Prompt Injection.
  5. Optimizely (2026). 2026 Optimizely Opal release notes.
  6. Optimizely (2026). AI variation development agent.
  7. Optimizely (2026). Experiment Planning agent.
  8. Optimizely (2026). Experiment Summary agent.
  9. Optimizely (2026). Remote MCP Server is here: bring Optimizely into your AI tools.
  10. Kameleoon (2025). Introducing prompt-based experimentation.
  11. Kameleoon (2026). PBX 2.0 is changing testing again.
  12. Kameleoon (2026). Expanding prompt-based experimentation with PBX Ideate.
  13. Kameleoon (2026). PBX Ship turns winning tests into production code.
  14. Kameleoon (2026). Prompt-based experimentation plans.
  15. Wingify (2025). Create optimization campaigns by prompting with Copilot.
  16. Wingify (2026). Agentic AI for experimentation: hype vs. reality.
  17. Wingify (2026). Wingz AI.
  18. Wingify (2026). Connect Wingify with AI tools using Wingify's MCP server.
  19. Amplitude (2026). Amplitude introduces agentic AI analytics.
  20. Amplitude (2026). Amplitude MCP server documentation.
  21. Amplitude (2026). Amplitude and Statsig partnership.
  22. Contentsquare (2025, updated 2026). Contentsquare's MCP: bridging agents and experience data.
  23. Contentsquare (2026). Contentsquare launches AI agent and ChatGPT analytics capabilities.
  24. Adobe (2025). Introducing Adobe Journey Optimizer Experimentation Accelerator.
  25. Google (2025). Google Analytics MCP server and Try the Google Analytics MCP server.
  26. PPC Land (2025). Google Analytics experimental MCP server enables AI conversations with data.
  27. Microsoft (2026). Microsoft Clarity MCP server.
  28. PPC Land (2026). Microsoft bets on AI economy with Web IQ, Clarity citations and MCP server.
  29. TechCrunch (2025). Datadog acquires Eppo, a feature flagging and experimentation platform.
  30. TechCrunch (2025). OpenAI acquires product testing startup Statsig.
  31. TechCrunch (2026). Everstone combines Wingify and AB Tasty.
  32. Ashokkumar, A., Hewitt, L., Ghezae, I. & Willer, R. (2026). Large language models can predict the results of social science experiments. Nature 656. Summary figures from treatmenteffect.app.
  33. Ye, Z., Yoganarasimhan, H. & Zheng, Y. (2024). LOLA: LLM-assisted online learning algorithm for content experiments. arXiv; Marketing Science (2025).
  34. METR (2025). Measuring the impact of early-2025 AI on experienced open-source developer productivity.
  35. Ouyang, S., Zhang, J. M., Harman, M. & Wang, M. (2023). An empirical study of the non-determinism of ChatGPT in code generation. arXiv.
  36. Miller, E. (2010). How not to run an A/B test.
  37. Fabijan, A. et al. (2019). Diagnosing sample ratio mismatch in online controlled experiments. KDD 2019.
  38. Kohavi, R., Deng, A. & Vermeer, L. (2022). A/B testing intuition busters. KDD 2022.
  39. Kohavi, R., Tang, D. & Xu, Y. (2020). Trustworthy Online Controlled Experiments. Cambridge University Press.
  40. Henkan & Partners (2025, revised 2026). How AI is reshaping A/B testing: what 19 prompt-built tests tell us.