Point of View

Why Prompt-Based Experimentation Belongs on the C-Level Agenda

Alexandre Suon · 2025-10-08

Prompt-based experimentation lets anyone describe a website change in plain words and have AI build the A/B test variant. This memo, first written when Kameleoon opened its prompt-based testing to all in autumn 2025 and updated for September 2026, explains why that removes the developer bottleneck, what our own 19-test study found, and why the decision now belongs with the executive team.

Executive summary

  1. Prompt-based experimentation removes the main brake on testing: developer time. A marketer describes a change ("move the reviews above the price"), and AI builds a working test variant. Kameleoon launched its version, PBX, in June 2025 and opened a free trial on 25 September 2025. By September 2026, Optimizely, VWO, AB Tasty and Amplitude offer comparable AI features.
  2. It finishes a job the visual editor started in 2010 and could not complete. Point-and-click (WYSIWYG) editors promised testing without developers. On modern, JavaScript-heavy websites they often broke, so teams needed code again. Prompting writes that code for you.
  3. Our own test showed the output is accurate enough for most changes. Henkan & Partners rebuilt 19 historical A/B tests with PBX alone: 11 matched the original at 90% fidelity or better, static changes averaged about 92% and interactive elements about 80%, and build time fell by 89% on average.
  4. Volume matters because most ideas fail. Published research puts the share of tested ideas that win at 10–33%. At a fixed budget, more tests means more winners. In our illustrative scenario, the same $2M a year moves from 40–60 winning tests to 200–300, cutting the budget per winner by about 80%.
  5. This creates a divide that shows up in the boardroom. Companies that let more people test, and accept speed over pixel perfection, will learn faster. Those that route every change through IT backlogs will stay at today's ceiling. For the CEO, CMO, CTO and CFO, each gains something different.
  6. The next step is organisational, not technical. People, platform, process and plan all need to change: who may launch tests, how the tool fits the stack, how quality is kept at speed, and how results are tied to revenue, margin and retention.

Section 1 · The idea

Prompt-based experimentation turns a sentence into a working test variant

Prompt-based experimentation is a way of building A/B tests in which a person describes the change they want in plain language, and an AI model writes the code for the test variant, which can then be checked, adjusted and launched.

The idea is simple. You open your website, type "add a delivery-date message under the add-to-cart button" or "make the size guide a pop-up", and the AI builds the variant. You review it, adjust it with another prompt, and launch an A/B test. Any idea can become an experiment by default.

Kameleoon, a French testing platform and a technology partner of Henkan & Partners, made its Prompt-Based Experimentation (PBX) generally available on 30 June 2025. It works through a Chrome extension on top of the Kameleoon snippet, and the company says it handles single-page applications and dynamic front ends without special set-up. On 25 September 2025, Kameleoon relaunched its brand around PBX, announced availability on Starter and Enterprise plans and opened a 30-day free trial. I first wrote this memo in the days after that relaunch.

I called it experimentation at the speed of thought, or vibe experimentation: the same shift that "vibe coding" brought to software, applied to testing. The point is not that the AI is clever. The point is that the cost of turning an idea into a test falls close to zero, and that changes who tests, how often, and what they dare to try.

For marketers. Prompt-based building does not remove the need for a hypothesis, a primary metric and a sample size. It removes the wait for a developer. Your job shifts from writing tickets to writing precise prompts and checking the result.

For leaders. The capability is now available from several vendors. The question is no longer whether the technology works, but whether your organisation is set up to use it.

Section 2 · A short history

Visual editors promised testing without developers in 2010, but modern websites brought the developers back

Many people believe experimentation became mainstream because Google, Microsoft or Booking.com ran thousands of tests a year. In my view, that is not the real story. The real catalyst for most companies was the visual editor.

Optimizely was founded in January 2010 by two former Google employees, Dan Siroker and Pete Koomen, and went through Y Combinator that winter. Its what-you-see-is-what-you-get (WYSIWYG) editor let anyone on a marketing team click a headline, a picture or a button and change it. No engineering tickets, no long waits. Competitors including Maxymiser, founded in 2006, were signing enterprise clients at the same time. The model attracted capital: Optimizely raised a $58M Series C in October 2015, bringing its total funding to $146M, and Oracle bought Maxymiser in August 2015 to add testing to its Marketing Cloud.

It was an intoxicating vision. But the promise did not hold. Beyond surface edits, visual editors needed injected JavaScript or CSS. As websites moved to frameworks that redraw the page constantly, point-and-click changes broke or flickered. "No technical skills required" became "we still need developers". Kameleoon's own launch post for PBX makes the same diagnosis: visual editors "broke on modern web stacks".

Timeline from 2010 to 2026 in three eras. Visual editors: Optimizely founded 2010, Oracle buys Maxymiser and Optimizely Series C of $58M in 2015. Consolidation: Episerver buys Optimizely 2020, Google Optimize shut down September 2023. Prompt-based and agentic testing: VWO Copilot November 2024, Datadog buys Eppo May 2025, Kameleoon PBX June 2025, OpenAI buys Statsig for $1.1B September 2025, PBX free trial September 2025, Optimizely Opal variation agent October 2025, Amplitude AI agents February 2026, Kameleoon PBX 2.0 April 2026, Amplitude takes over Statsig May 2026, PBX Ship May 2026.
Exhibit 1. From visual editors to AI-built variants, 2010–2026. Source: Wikipedia, TechCrunch, Oracle, Google, VWO/Wingify, Kameleoon, Optimizely release notes, CNBC, Amplitude.

What this shows. The visual-editor era lasted about a decade and was followed by consolidation, including the shutdown of the free Google Optimize in September 2023. From late 2024, AI variant generation appeared across vendors within twelve months. Most events cluster in 2025–26, which is why that part of the timeline is drawn wider.

A correction to the first version of this memo. I wrote that Optimizely "reached unicorn status". I could not verify that: its last disclosed round was the 2015 Series C, and Contrary Research estimates its value at about $600M when Episerver bought it in 2020. The billion-dollar figure belongs to Episerver, bought by Insight Partners for about $1.2B in 2018, which later took the Optimizely name. The point stands without it: investors paid heavily for the promise of testing without developers.

Section 3 · The bottleneck

The limit on experimentation is the capacity to build tests, not the supply of ideas

In our experience at Henkan & Partners, most companies struggle to run more than 20–30 experiments a year. It is not for lack of ambition. Many have run complex tests for years: funnel redesigns, pricing changes, new features. The problem is constraints:

  • Developers are expensive and their time goes first to the core product roadmap.
  • A parallel stream of test builds is rarely a priority, so tests wait in the backlog.
  • Most tests do not win. That is normal and healthy, but it means a small programme produces very few winners.

The third point needs care, because the first version of this memo said "only 1 in 5 experiments succeed" without a source. The best-documented figures come from Ron Kohavi, Diane Tang and Ya Xu, who led experimentation at Microsoft, Google and LinkedIn. In Trustworthy Online Controlled Experiments (2020) they report that only about one third of ideas tested at Microsoft improved the metrics they targeted, and that in well-optimised products such as Bing and Google some measures show success rates of only 10–20%. One in five is therefore a reasonable planning figure, not a law.

Bar chart of published success rates of tested ideas: Microsoft about 33%, Slack monetization tests about 30%, Bing and Google 10 to 20% on some measures, Netflix about 10%. A dashed line marks the 20% win rate assumed in the Henkan scenario.
Exhibit 2. Share of tested ideas that improved the target metric, as published. Source: Kohavi, Tang & Xu, Trustworthy Online Controlled Experiments (2020), chapter 1.

What this shows. Even the most experienced testing organisations see most ideas fail. With a 20% win rate, a company running 25 tests a year finds about five winners. The only reliable way to find more winners is to run more good tests, which is exactly what build capacity limits.

This is why the largest digital companies invested so heavily in testing infrastructure. Booking.com's data scientists reported in 2017 that staff across the company "run and analyse more than a thousand concurrent experiments". Kohavi and colleagues report that Google, LinkedIn and Microsoft each run at a rate of more than 20,000 controlled experiments a year, and that Bing alone runs more than 10,000. The first version of this memo claimed that "Booking.com and Expedia run 15,000+ experiments per year"; I could not find that figure published by either company, so I have replaced it with these sourced numbers.

Most companies cannot build that infrastructure. In our experience, even large enterprises spending more than $2M a year on experimentation rarely exceed 200–300 tests. That ceiling has held for more than a decade. Prompt-based building is the first change I have seen that could raise it without hiring an army of developers.

For leaders. Ask your team two numbers: how many tests did we complete last year, and how many won? If the answer is below 30 tests, your programme is limited by build capacity, and the winners you find are too few to move the P&L.

Section 4 · Our study

Our rebuild of 19 historical tests showed prompt-built variants are accurate enough for most changes

Claims about AI are cheap, so we tested PBX ourselves. In summer 2025, Henkan & Partners took 19 historical A/B tests that had already been built and run by hand for a client, and rebuilt each one with PBX alone, using prompts only. We then scored each result against the original implementation. The full study gives every test-level result.

The results, which are Henkan & Partners analysis and not vendor claims, were:

  • 11 of 19 tests reached 90% fidelity or better, and the median was 92%.
  • Static content changes averaged about 92% (text, visibility, simple layout) and interactive elements about 80% (links, banners, repositioned buttons, filters).
  • Precise layouts and business rules were the weak points: five tests scored below 70%, mostly because of visual precision or conditional logic.
  • Build time fell by 89% on average, from 8.7 hours to 55 minutes per test.
Bar chart of average fidelity of variants rebuilt with Kameleoon PBX by complexity category, Henkan & Partners study of 19 tests: static content 91.7%, interactive elements 80.2%, complex UX 76.5%, advanced technical 80.0% (two tests), with the range of each category.
Exhibit 3. Fidelity of variants rebuilt with Kameleoon PBX, by type of change. Source: Henkan & Partners study of 19 historical A/B tests, 2025.

What this shows. Static changes, which make up a large share of marketing tests, averaged about 92% fidelity. Interactive and complex changes were less reliable, and every category had some weak builds, mostly where an exact layout or a business rule was needed. Prompt-built variants still need a human check, but for most tests that check is quick.

Bar chart of build time for one test variant indexed to manual build at 100: prompt-based build in the Henkan & Partners study at 11, meaning 89% saved on average, and Kameleoon's vendor-reported average across 1,000 experiments at 3, meaning 97% saved.
Exhibit 4. Build time per test variant, indexed to a manual build. Source: Henkan & Partners study (19 tests, 2025); Kameleoon, "What we learned from 1,000 prompt-based experiments" (22 July 2025), vendor-reported.

What this shows. Our average points the same way as Kameleoon's own figure. The vendor reported a 97% reduction in build time across 1,000 prompt-based experiments on 48 websites, assuming three developer days per test. Our 89% is lower mainly because it is measured against real hand builds that averaged 8.7 hours, and it covers the build only, not QA.

Three limits should be clear. First, 19 tests from one client is a small sample, so the figures are indicative. Second, fidelity was measured against the original build, not against business results. Third, the study covered client-side web tests. Server-side tests of pricing logic, algorithms or app features still need engineers, although, as Section 10 shows, vendors are now extending AI into that part of the work too.

For marketers. Start with static changes (copy, layout, images, order of page elements), where accuracy is highest. Keep a developer or experienced CRO specialist in the review loop for forms, pop-ups and navigation until your team has a track record.

Section 5 · The economics

At the same budget, faster builds can multiply the number of winners: an illustrative scenario

What does this mean for return on investment? The following is a Henkan illustrative model, not market data. It shows the arithmetic for a company spending $2M a year on experimentation, with the assumptions stated so that any reader can change them.

Assumption or resultTodayPrompt-based scenarioType
Annual experimentation budget$2M$2MAssumption (held constant)
Tests completed per year200–3001,000–1,500Assumption, based on our experience and the build-time savings in Section 4
Win rate20%20%Assumption, within the published 10–33% range (Exhibit 2)
Winning tests per year40–60200–300Derived: tests × win rate
Budget per winning test$33–50K$7–10KDerived: budget ÷ winners (about −80%)
Two-panel chart of a Henkan illustrative model for a $2M annual budget. Left: 200 to 300 tests and 40 to 60 winners today versus 1,000 to 1,500 tests and 200 to 300 winners in a prompt-based scenario. Right: budget per winning test falls from $33–50K to $7–10K, about 80% lower.
Exhibit 5. Tests, winners and budget per winner at a constant $2M budget. Source: Henkan & Partners illustrative model; assumptions stated in the table above.

What this shows. If the win rate holds and the budget stays the same, five times more tests means five times more winners and a budget per winner about 80% lower. The model is only as good as its assumptions. The biggest risk is that the win rate falls when more people test weaker ideas, which is why the process pillar in Section 9 matters.

In the original memo I also estimated that a developer-built test of medium complexity costs $15–20K once design, build and QA are included, against about $3–5K when it is built with prompts. Those are Henkan estimates for a single test from our project experience, not a division of the $2M budget, and your own figures will depend on day rates and test complexity. I have removed a further line from the first version, which projected revenue uplift for a $500M digital business, because its assumptions were not stated.

For leaders and investors. Read Exhibit 5 as a sensitivity, not a forecast. The two levers to test in your own business are the number of tests your team can realistically build and review, and whether the win rate holds as more people contribute ideas. Even at half the scenario's volume, the budget per winner falls sharply.

Section 6 · The divide

A strategic divide is opening between companies that let more people test and those that do not

When Kameleoon launched PBX, I argued it signalled more than a technical milestone. It was the start of a divide in the market. A year later, with similar features in most major tools, I think the divide is about organisation, not software.

Progressive companies will seize the moment. They will let people beyond product and engineering (marketers, analysts, customer service managers, even front-line staff) turn insights into live experiments through prompts. They will accept that not every pixel has to be perfect, and that speed and learning outweigh polish for most tests. They will run three, five, even ten times more experiments than before. The compounding effect of that learning will sharpen the customer experience and build a culture of evidence that is hard to copy.

Conservative companies will hesitate. They will keep routing every change through IT backlogs. They will insist on exhaustive QA cycles and design perfection before anything goes live. Their test velocity will stay at the same ceiling of 200–300 tests a year. While they debate internal processes, their competitors will be learning at a different scale.

Progressive companiesConservative companies
Who can launch a testTrained people across marketing, product, service and storesProduct and engineering only
Quality standardProportionate: light checks for low-risk changes, full review for risky onesEvery test through full design and QA review
Tests per yearSeveral times today's levelStuck at today's ceiling
Where experimentation sitsA growth engine reported to the executive teamA support function inside digital or IT
Main riskNoise: too many weak tests, uneven qualitySlowness: competitors learn faster

This divide will not stay at the operational level. It will show up in the boardroom. For executives, the question is no longer about optimisation at the margins. It is whether your company belongs to the group that turns experimentation into a strategic growth engine, or the group that watches from the sidelines.

Section 7 · The C-level case

Each member of the executive team gains something different from cheaper, faster tests

For years, experimentation was seen as tactical rather than strategic. Even organisations running sophisticated tests on pricing, funnels and product launches often treated it as a support function: insightful, but not central to corporate direction. That framing limited its visibility at executive level. Budgets stayed capped, teams stayed small, and experimentation was rarely positioned as more than an optimisation activity.

Prompt-based experimentation changes the equation on three dimensions.

1. Return on investment

At a constant budget, more tests means more winners and a lower cost per winner (Section 5). In our illustrative scenario, the budget per winning test falls by about 80%. Each winner adds to revenue in the following years, so the gains compound.

2. Competitive advantage

Companies such as Booking.com show that velocity compounds: more than a thousand concurrent experiments, run by staff across the company, produce a customer experience that is constantly refined. Until now, that velocity required years of investment in in-house platforms. Prompt-based building brings part of it within reach of mid-sized and traditional companies, and of teams of one or two people, not only digital giants.

3. Strategic alignment

ExecutiveWhat changesThe question to ask
CEOBold bets can be de-risked faster, by testing them on real customers before full investment.Which of our big assumptions for next year could we test this quarter?
CMOCampaign landing pages and personalisation can be tested at a fraction of the cost and time.How many of our campaign pages were tested last quarter?
CTO / CPOScarce developers move from building test variants to strategic product work; winning variants return as clean code.How many developer days went to building tests last year?
CFOExperimentation moves from a $2M cost centre to a measurable growth driver, tracked by cost per winner.What is our cost per winning test, and what did the winners add to revenue?

For leaders. Put one number on the executive dashboard: winning tests per quarter, with their estimated annual revenue impact. It makes experimentation visible in the same terms as any other investment.

Section 8 · Anyone can test

When anyone can test, insight from the front line becomes an experiment in days

The most far-reaching aspect of prompt-based experimentation is who can run a test. Until now, experimentation was limited to product and development teams, or to marketers with a developer on call. With prompt-based building, any trained employee who meets a problem can propose and build a test of the solution:

  • A call-centre agent hears the same complaint every day and launches an A/B test of a new FAQ flow.
  • A store associate hears customer feedback and tests an improved product description online.
  • A data analyst spots a conversion drop and builds a funnel variation to isolate the issue.
  • A marketing manager sees a trend in campaign performance and tests a personalised landing page within days.

This is vibe experimentation: anyone in the organisation can turn feedback, intuition or data into a live test. It connects directly to Voice of Customer work, because the people who hear customers every day are rarely the people who build tests.

"Anyone" does not mean "anyone, anyhow". In our projects, the companies that open testing widely do three things. They train contributors on hypotheses and metrics. They route every test through a light approval step, owned by a small central team. And they set guardrail metrics, such as page speed and error rates, that stop a test automatically if something breaks.

For marketers and CRO teams. Your role becomes that of a coach and quality owner: writing the prompt templates, reviewing tests from the wider business, and making sure results are recorded so the company learns from them.

Section 9 · The four pillars

People, platform, process and plan all need to change before prompt-based testing pays off

To seize this opportunity, companies need to revisit the four classic pillars of an experimentation programme. The technology is the easy part. The organisation is where most programmes stall.

Framework of four pillars. People: who is allowed to launch and read tests. Platform: how prompt-based building fits the stack. Process: how to keep speed without losing quality. Plan: how results reach board-level KPIs such as revenue, margin and retention.
Exhibit 6. The four pillars of an experimentation programme when anyone can build a test. Source: Henkan & Partners framework.

What this shows. Each pillar poses one question that prompt-based testing makes urgent. A company that answers only the platform question will buy a tool and see little change in the number of tests or winners.

  • People. Who gets to launch and analyse tests? Define roles (contributor, reviewer, owner), train contributors, and name a central team responsible for quality.
  • Platform. How does prompt-based building fit your existing tools? Check that it works with your analytics, your consent set-up and your design system, and that winning variants can be turned into production code.
  • Process. How do you balance speed with quality and reliability? Set review levels by risk: light for copy and layout, full for checkout, pricing and legal content. Use guardrail metrics on every test.
  • Plan. How do experimentation results reach board-level KPIs such as revenue, margin and retention? Tie each test to one of them, and keep a shared library of results so learning builds up.

Section 10 · Update, September 2026

A year on, AI-built variants are standard, and the contest has moved to agents that run the whole test

When I first wrote this memo, Kameleoon was one of the first vendors to build a product around prompts. Twelve months later, the picture has changed in three ways.

Kameleoon extended PBX from building to the whole test cycle

Kameleoon moved to a credit-based model. Its documentation lists a 30-day free trial with 10 credits for sites with up to 5,000 monthly tracked users, a month-to-month Starter plan with 30, 100 or 200 credits a month, and Enterprise plans with annual credit packages. One experiment uses one to two credits. Prices for Starter and Enterprise are not published. In May 2026 it added PBX Ideate, which scans a page and suggests prioritised test ideas based on more than 20,000 historical experiments. In April 2026 it presented PBX 2.0, a set of AI agents to ideate, build, configure and analyse tests. In May 2026 PBX Ship followed, which turns winning client-side variants into production code behind a feature flag through an MCP server in the developer's own coding tool.

Competitors now offer similar AI features

Prompt-based variant building is no longer unique. The main vendors have released comparable capabilities, each with its own name and scope:

VendorAI capability for building or running testsDate announced
VWO (Wingify group)Copilot generates variations in the visual editor from conversational commands; later, a full campaign from a plain-English description (Pro and Enterprise plans)Nov 2024; Aug 2025
OptimizelyOpal Chat for experimentation; AI variation development agent that builds and changes page elements from natural languageMay 2025; Oct 2025
AdobeExperimentation Accelerator in Journey Optimizer, powered by an Experimentation Agent that finds insights and suggests testsSep 2025
AB Tasty (Wingify group)Evi Content: modifies selected page elements from natural-language prompts in the visual editorDocumented by Nov 2025
AmplitudeWeb Experimentation Agent that designs and launches experiments and analyses results; Website Conversion Agent (early access) that creates variantsFeb 2026
KameleoonPBX, then PBX 2.0 agents (Ideate, Build, Configure, Analyze, Ship)Jun 2025; Apr–May 2026

Experimentation became strategic for AI and data companies

Two deals in 2025 confirmed that experimentation is now core infrastructure, not a marketing add-on. In May 2025, Datadog acquired Eppo, a warehouse-native experimentation and feature-flag platform, for a reported $220M. In September 2025, OpenAI acquired Statsig in an all-stock deal valued at $1.1B, and Statsig's founder, Vijaye Raji, became OpenAI's CTO of Applications. In May 2026, Amplitude took over the Statsig brand, platform and customers while the team stayed at OpenAI. Meanwhile, Everstone Capital combined VWO and AB Tasty under Wingify in January 2026, creating a group with more than $100M in annual revenue.

What this means for the thesis. The core argument of the original memo has held: the cost of building a test variant is collapsing, and the constraint is moving from engineering to organisation. What has changed is that the advantage no longer comes from being early to a single tool. It comes from how fast a company adapts its people, process and plan. The vendors are converging; the gap between progressive and conservative companies is widening.

Section 11 · Implications

Executives should treat prompt-based experimentation as an operating-model decision, not a tool purchase

These recommendations apply to every size of team, from one person running a handful of tests to an embedded experimentation team. Start small, measure, and scale what works.

1. Measure your starting point

Count tests completed and winners over the last twelve months, and the budget behind them. Calculate your cost per winner. Without this baseline, you cannot show whether prompt-based building changed anything.

2. Run a controlled pilot on static changes

Pick one team and one high-traffic page. Rebuild a few past tests with prompts, as we did, to check accuracy on your own site. Then run new tests for a quarter and compare build time, QA time and test count with your baseline.

3. Choose a tool on fit, not on the demo

Most major platforms now offer AI variant building. Compare them on how they handle your site's technology, how easily winning variants reach production, how they fit your analytics and consent set-up, and how credits or usage are priced at your expected volume.

4. Open testing to new contributors, with guardrails

Train people outside the product team, starting with those closest to customers. Set review levels by risk, and apply guardrail metrics to every test so that speed does not come at the expense of the customer experience.

5. Report experimentation to the executive team in revenue terms

Tie each test to revenue, margin or retention, and report winners and their estimated impact every quarter. That is how experimentation moves from a cost centre to a growth engine, and why it belongs on the C-level agenda.

The bottom line. Prompt-based experimentation is not only a new capability. It is a business inflection point. It turns experimentation into a way to de-risk bets, move resources to proven winners and build a lasting advantage, provided the organisation changes along with the tool.

FAQ

Frequently asked questions about prompt-based experimentation

Frequently asked questions

What is prompt-based experimentation?

It is a way of building A/B tests in which you describe the change you want in plain language and an AI model writes the code for the test variant. You review and adjust the result, then launch the test. Kameleoon calls its product PBX (Prompt-Based Experimentation); other vendors offer similar features under other names.

Can AI generate A/B test variants accurately?

For most marketing changes, yes, with human review. When Henkan & Partners rebuilt 19 historical A/B tests with Kameleoon PBX, 11 matched the original at 90% fidelity or better; static changes averaged about 92% and interactive elements about 80%. Precise layouts and business logic still need careful checking.

Which A/B testing tools use AI to build variants?

As of September 2026, Kameleoon (PBX), Optimizely (Opal variation development agent), VWO (Copilot), AB Tasty (Evi Content) and Amplitude (experimentation agents) all offer AI features that create or change test variants from natural language. Adobe offers an Experimentation Agent that analyses tests and suggests new ones. Scope and plan availability differ by vendor.

How much does Kameleoon PBX cost?

Kameleoon offers a 30-day free trial with 10 credits, a month-to-month Starter plan with 30, 100 or 200 credits a month, and Enterprise plans with annual credit packages. One experiment uses one to two credits. Kameleoon does not publish Starter or Enterprise prices.

How does AI increase experimentation velocity?

Building the test variant is usually the slowest step, because it needs developer time. AI cuts that step sharply: 89% on average in our study and 97% in Kameleoon's own data. With the same budget and people, a team can build and launch several times more tests, and because most ideas fail, more tests means more winners.

What percentage of A/B tests succeed?

Published figures range from about one in three to one in ten. Kohavi, Tang and Xu report that about a third of ideas tested at Microsoft improved their target metric, and that some measures at Bing and Google show success rates of 10–20%. A win rate of around 20% is a reasonable planning assumption.

Why did OpenAI buy Statsig?

OpenAI bought the experimentation platform Statsig in September 2025, in an all-stock deal valued at $1.1B, to speed up product development across its applications, and made Statsig's CEO its CTO of Applications. In May 2026, Amplitude took over the Statsig brand, platform and customers.

Key terms

A/B test
An experiment that shows two versions of a page or feature to randomly split groups of visitors and compares results. It is the only reliable way to know whether a change caused an improvement, rather than coinciding with one.
Variant
The changed version in an A/B test, compared against the unchanged control. Building the variant is the step that has always needed developers, so it sets the pace of the whole programme.
Prompt-based experimentation (PBX)
Building a test variant by describing the change in natural language to an AI, which writes the code. PBX is Kameleoon's name for its product; the general idea now exists in several tools.
Vibe experimentation
Our shorthand for testing at the speed of thought: turning an idea, a complaint or a data point into a live test in minutes. It matters because speed changes who bothers to test.
WYSIWYG / visual editor
"What you see is what you get": a point-and-click editor for changing a web page without code. It opened testing to marketers around 2010 but struggled with complex, modern websites.
Client-side testing
Testing where a script changes the page in the visitor's browser after it loads. It is how most marketing-led tests run, and it is where prompt-based building applies first.
Single-page application (SPA)
A website that updates content without reloading the page, typically built with frameworks such as React. SPAs are where visual editors most often failed, which is why vendors stress that prompt-built variants work on them.
Test velocity
The number of experiments a team completes in a period. Because most ideas fail, velocity largely determines how many winning changes a company finds each year.
Win rate
The share of experiments in which the new version beats the control on the target metric. Published figures range from about 10% to about 33%, which is why volume matters.
QA (quality assurance)
Checking that a variant displays and behaves correctly on all devices and browsers before launch. Faster builds only pay off if QA is proportionate to the risk of each change.
Guardrail metric
A metric that must not get worse during a test, such as page speed, errors or refunds. Guardrails let companies allow more people to test without putting the business at risk.
Feature flag
A switch in the code that turns a feature on or off for chosen users. Flags are how winning tests are rolled out safely into the real product.
AI agent
AI software that carries out a multi-step task, such as proposing, building, launching and reading a test, rather than answering one question. Most vendor roadmaps in 2026 are built around agents.
MCP (Model Context Protocol)
An open standard that lets AI assistants connect to other software. Kameleoon uses it so a developer's AI coding assistant can turn a winning test into production code.
Cost per winner
Total experimentation spend divided by the number of winning tests. It is the figure a CFO should track, because it links the testing budget to the changes that actually make money.

Sources

Vendor capabilities and dates are taken from each company's own announcements, release notes or documentation, retrieved in September 2026; features and plans change often. Deal values are those disclosed by the parties or reported by named outlets; Datadog's price for Eppo is reported, not official. Success rates are as published by Kohavi and colleagues. The 19-test fidelity and time-saving figures come from the Henkan & Partners study of 2025, recomputed from its test-level data in September 2026. The ROI scenario (Exhibit 5), the per-test cost estimates, the divide table and the four pillars (Exhibit 6) are Henkan & Partners models and frameworks, not market data.

  1. Kameleoon: Introducing Prompt-Based Experimentation (30 June 2025)
  2. EIN Presswire: Kameleoon launches vibe experimentation with prompt-based testing (25 September 2025)
  3. Kameleoon: Introducing Kameleoon's new brand (25 September 2025)
  4. Kameleoon: New Kameleoon plans, try PBX for free (6 October 2025)
  5. Kameleoon documentation: Prompt-based experimentation plans
  6. Kameleoon: What we learned from 1,000 prompt-based experiments (22 July 2025)
  7. Kameleoon: PBX 2.0 is changing testing again (16 April 2026)
  8. Kameleoon: Expanding prompt-based experimentation with PBX Ideate (7 May 2026)
  9. Kameleoon: PBX Ship turns winning tests into production code (26 May 2026)
  10. VWO: Introducing VWO Copilot (20 November 2024)
  11. Wingify: Create optimization campaigns by prompting with Copilot (28 August 2025)
  12. Optimizely: 2025 Web Experimentation release notes
  13. Optimizely: AI variation development agent
  14. AB Tasty documentation: Using Evi Content
  15. Adobe: Introducing Journey Optimizer Experimentation Accelerator (10 September 2025)
  16. Amplitude: Amplitude introduces agentic AI analytics (24 February 2026)
  17. Amplitude documentation: Website Conversion Agent
  18. Amplitude: Amplitude and Statsig partnership (5 May 2026)
  19. TechCrunch: OpenAI acquires Statsig (2 September 2025) · CNBC
  20. TechCrunch: Datadog acquires Eppo (5 May 2025)
  21. TechCrunch: Everstone combines Wingify and AB Tasty (20 January 2026)
  22. Wikipedia: Optimizely
  23. TechCrunch: Optimizely raises $58M Series C (13 October 2015)
  24. Contrary Research: Optimizely business breakdown
  25. Oracle buys Maxymiser (20 August 2015)
  26. Google: Optimize and Optimize 360 sunset (30 September 2023)
  27. Kaufman, Pitchforth & Vermeer: Democratizing online controlled experiments at Booking.com (2017)
  28. Kohavi, Tang & Xu: Trustworthy Online Controlled Experiments, chapter 1 (2020)
  29. Kohavi et al.: Online randomized controlled experiments at scale, Trials (2020)