Focus

How to Build a Culture of Experimentation

Alexandre Suon · 2026-09-27

Buying an A/B testing tool takes a week. Getting a company to let evidence overrule opinion takes years. This focus explains what an experimentation culture is, why most programmes stall, and the concrete steps that build one, whether your team is one person or one hundred.

Executive summary

  1. A culture of experimentation means decisions are settled by evidence from customers, not by rank. Tools matter, but the difference between companies that test and companies that learn is who is allowed to be wrong, what gets rewarded and whether leaders put their own ideas to the test.
  2. Most ideas fail, so a culture that punishes failure cannot learn. Published success rates range from about 8% of experiments at Airbnb Search to about a third at Microsoft. A team that expects every test to win will either stop testing or stop telling the truth about results.
  3. Leaders set the culture more than any process does. Researchers name the highest-paid person's opinion (HiPPO) as one of the main obstacles to testing. Leaders who say "I don't know, let's test it", test their own ideas and celebrate tests that stopped a bad change do more than any training plan.
  4. Reward learning, not wins. Spotify reports a 12% win rate but a 64% learning rate. Measuring learning, trust in the data and speed to decision, not only the number of winners, is what keeps teams honest and motivated.
  5. Build trust before scale, and scale through guardrails, not permission. Most programmes lack basics such as quality checks, prioritisation and programme metrics. Fix those first, then let more teams run tests with shared metrics, reviews and a searchable library of results.
  6. Every team size can start. One person can run a monthly test-and-learn review; a larger organisation can build a centre of excellence. The first 90 days matter more than the tool: pick a sponsor, agree one success metric, run a few well-chosen tests and share every result.

An experimentation culture is a way of working in which a company treats ideas as hypotheses, tests them with real customers before committing, and lets the results, including negative ones, change its decisions. It is visible in behaviour: who can launch a test, how failures are discussed, and whether senior people accept being proved wrong.

Section 1 · Why culture

Most ideas fail when tested, so the real challenge is accepting what the tests say

Every company says it is data-driven. The test is what happens when the data disagrees with a senior person, or when a project that took six months shows no effect. In a culture of experimentation the result wins. In most companies, the result gets explained away.

The reason this matters is simple: most ideas do not work. In their 2017 Harvard Business Review article, Ron Kohavi and Stefan Thomke reported that "at Google and Bing, only about 10% to 20% of experiments generate positive results", and that at Microsoft as a whole "one-third prove effective, one-third have neutral results, and one-third have negative results." Later work by Kohavi and colleagues compiled published success rates from several companies. The picture is consistent: a clear majority of carefully chosen ideas fail to move the metric they were designed to move.

Horizontal bar chart of the share of experiments that improved their target metric. Microsoft, all products, 33%. Analytics Toolkit dataset of 1,001 tests, 33.5% (vendor data). Bing 15%. Optimizely customers, 127,000 tests, 12% (vendor data). Spotify 12%. Booking.com, Google Ads and Netflix 10%. Airbnb Search 8%.
Exhibit 1. At leading testing companies and across large vendor datasets, only about one idea in three to one in twelve improves the metric it targets. Source: Kohavi, Deng and Vermeer, A/B Testing Intuition Busters (KDD 2022); Kohavi and Thomke, HBR (2017); Optimizely, analysis of 127,000 experiments (vendor data); Spotify Engineering (2025); Analytics Toolkit, 1,001 A/B tests (2022, vendor data).

What this shows. Failure is the normal outcome of testing, not a sign that the programme is broken. Success rates also depend on how mature a product is and how a "win" is defined, so they are not league tables. What they share is the lesson for culture: if people are judged on whether their ideas win, most of them will look bad most of the time.

People are also poor judges of which ideas will work. In one exercise at Microsoft, "only three out of 21 people guessed the winner, and the three were from the ExP team", the company's own experimentation group. At Booking.com, Stefan Thomke reports that teams are "wrong about nine out of ten times" when they predict customer behaviour. A test that overturns a confident prediction is not an embarrassment. It is the programme doing its job.

So the hard part is not running tests. It is building an organisation that expects to be wrong, finds out quickly and cheaply, and changes course without blame. That is what we mean by culture, and it is why tools alone rarely deliver.

For leaders. Before asking how many tests the team ran, ask how many decisions changed because of a test, and how many ideas were stopped before they cost money. If the answer is "none", the programme is producing reports, not learning.

Section 2 · The payoff

Companies that make testing a habit grow faster, and the evidence goes beyond big tech

The most cited examples come from large technology firms, but the most useful evidence comes from ordinary companies. Rembrand Koning, Sharique Hasan and Aaron Chatterji tracked more than 35,000 start-ups in an observational study published in Management Science in 2022. They found that relatively few firms adopt A/B testing, but "among those that do, performance improves by 30%–100% after a year of use." Adopters also "develop more new products, identify and scale promising ideas, and fail faster when they receive negative signals."

At large scale the numbers are large. Kohavi and Thomke describe how one of hundreds of ideas at Bing, a small change to how ad headlines were displayed, was judged a low priority and "languished for more than six months". When it was finally tested, it increased revenue by 12%, worth "more than $100 million" a year in the United States alone. They also report that dozens of tested revenue changes each month collectively increased Bing's revenue per search "by 10% to 25% each year." Microsoft researchers estimate that scaling experimentation across the company is worth "hundreds of millions of dollars of additional revenue annually."

The value is not only in winners. Tests also stop expensive mistakes. The same HBR article notes that integrating Bing with social media "cost Microsoft more than $25 million to develop and produced negligible increases in engagement and revenue." A culture that tests early catches such projects before the full investment is made.

Type of valueWhat it looks likeWho notices
Better decisionsChanges that improve conversion, revenue or retention are kept; the rest are droppedE-commerce and product leads
Avoided lossesIdeas that would have hurt customers or revenue are stopped before full launchFinance and leadership, if someone reports it
Faster learningTeams know more about what customers value, which improves the next ideaProduct, design and marketing teams
Less politicsDisagreements are settled by a test instead of by rank or persistenceEveryone, over time

For an online shop, these benefits link directly to customer experience. Each test that removes friction, clarifies delivery costs or improves search makes the site easier to use. Over time, customers buy more often and stay longer, which is where lifetime value comes from.

Section 3 · Leadership

Leaders build the culture by testing their own ideas and saying "I don't know"

Researchers and practitioners name one obstacle again and again: the HiPPO, the "Highest Paid Person's Opinion". Kohavi and colleagues describe an early stage of "hubris", where "measurement is not needed because of confidence in the HiPPO". Thomke puts it more bluntly: "Nothing stalls innovation faster than a so-called HiPPO." When the most senior person in the room decides, tests become a formality, and people learn to propose only what the boss already likes.

Thomke describes a different leadership role. In an experimentation organisation, he writes, leadership "facilitates the process of decision-making through experimentation" instead of making every decision top-down, and "even the boss's assumptions are subject to real-world tests." He adds, in an interview, that bosses "ought to display intellectual humility and be unafraid to admit, 'I don't know.'"

Amazon's founder made the same point in his shareholder letters. In the 2015 letter he wrote that "failure and invention are inseparable twins. To invent you have to experiment, and if you know in advance that it's going to work, it's not an experiment." He also separated decisions into two types: "one-way doors" that are hard to reverse and deserve careful deliberation, and "two-way doors" that can be reversed and should be made quickly by small groups. Most changes to a website are two-way doors, which is exactly why they can be tested instead of debated.

Five things leaders can do this quarter

  1. Put one of your own ideas to a test, publicly. Nothing signals the new rules faster than a senior person whose idea loses and who says so in the team meeting.
  2. Change the question you ask. Replace "did it win?" with "what did we learn, and what will we do differently?"
  3. Name a sponsor with a budget. Someone senior must own the programme, protect testing time and remove blockers. Speero's 2025 benchmark (agency data, as reported by Convert) found only 26% of companies strongly agree they have a senior sponsor accountable for experimentation.
  4. Agree which decisions must be tested. At Bing, Kohavi and Thomke reported in 2017 that about 80% of proposed changes were first run as controlled experiments. You do not need that level, but you need a rule, for example: every change to checkout, pricing display or navigation is tested.
  5. Report stopped mistakes, not only wins. A test that prevented a revenue loss deserves the same visibility as a winner.

For leaders. Thomke's line is worth repeating in your next leadership meeting: in a true experimentation organisation, "even the boss's assumptions are subject to real-world tests." The culture changes the first time that happens in front of the team.

Section 4 · Safety and rewards

People only report honest results when failure is safe and learning is rewarded

A test programme is only as good as the honesty of its reporting. If a product owner's bonus depends on a feature succeeding, a flat result will be "re-analysed" until it looks like a win. That is how programmes lose credibility.

The research on teams points the same way. Amy Edmondson defined team psychological safety in 1999 as "a shared belief held by members of a team that the team is safe for interpersonal risk taking", and, in a study of 51 teams in a manufacturing company, found that it is linked to learning behaviour, which in turn is linked to performance. Google's Project Aristotle, which studied 180 teams, listed psychological safety first among the five dynamics of effective teams, ahead of dependability, structure and clarity, meaning and impact.

The practical question is what you reward. Spotify offers a useful model. Its engineers found that judging experiments by wins alone gave a misleading picture, so they introduced a "learning rate": the share of experiments that produce a clear, decision-ready answer, including answers such as do not ship.

Bar chart of Spotify experiment outcomes. Win rate, experiments that shipped a better variant: about 12%. Experiments where the team decided not to ship: 42%. Learning rate, experiments that gave a decision-ready answer: about 64%. Learning rate varied across teams from 16% to 76%.
Exhibit 2. Only 12% of Spotify's tests win, but 64% give the team a clear answer to act on. Source: Bellato, Schultzberg and Ankargren, Beyond Winning: Spotify's Experiments with Learning Framework, Spotify Engineering (2025); Spotify Confidence blog, The Real ROI of Experimentation (2026, vendor data).

What this shows. Measured by wins, Spotify's programme looks like it fails almost nine times out of ten. Measured by learning, it succeeds almost two times out of three. As Spotify's team puts it, most of their learning does not come from wins; it comes "from discovering what not to ship." The wide range between teams, 16% to 76%, also shows that learning rate is something a team can improve through better hypotheses and cleaner tests.

Thomke calls failure "the status quo" at Booking.com and says "not winning is not losing." A 2019 paper by experts from 13 organisations, including Airbnb, Booking.com, Google, LinkedIn, Microsoft, Netflix and Stanford University, describes what a mature culture looks like: negative experiment results are "celebrated as saving customers".

Instead of rewarding…Reward…Why
Number of winning testsShare of tests with a clear, decision-ready answerKeeps reporting honest
Size of the lift claimedLift confirmed after launch, or reviewed by a second analystStops inflated results
Ideas that were shippedIdeas that were stopped early because a test showed harmMakes saving money visible
Individual heroesTeams that share results others can reuseKnowledge survives when people leave

For marketers and product owners. Write down your prediction before each test and compare it with the result afterwards. Over a quarter, this simple habit shows how often intuition is wrong, and it makes the case for testing better than any slide.

Section 5 · Operating model

Start with one central owner, then spread testing to every team as trust grows

Culture needs a structure to live in. The most useful description comes from Microsoft researchers Aleksander Fabijan, Pavel Dmitriev, Helena Holmström Olsson and Jan Bosch, who interviewed product teams across the company and described four phases of maturity: Crawl, Walk, Run and Fly. At each phase, both the organisation and the way success is measured change.

Four-column diagram of the Experimentation Evolution Model. Crawl: a standalone central team sets up, runs and analyses tests; the success metric is defined from a few key signals; typical team size one or two people or an agency. Walk: data scientists are embedded in product teams; success is a set of success, guardrail and data-quality metrics; a small central team plus product squads. Run: product teams take full responsibility for their tests, supported by assigned data scientists; success is refined into a preferred single metric; several squads with a centre of excellence. Fly: experiment owners create, run and analyse tests autonomously; the success metric is stable and well defined; every product team with a shared platform.
Exhibit 3. As programmes mature, testing moves from one central team to every product team, and the success metric becomes clearer. Source: Fabijan, Dmitriev, Olsson and Bosch, The Evolution of Continuous Experimentation in Software Product Development (ICSE 2017); team-size row is a Henkan & Partners interpretation.

What this shows. Maturity is not about tool features. It is about who is trusted to run tests and how clearly the company defines success. Most companies should not try to jump from Crawl to Fly: autonomy without shared metrics and reviews produces many tests nobody trusts.

Three operating models are common. CXL describes them as centralised (one team runs everything), decentralised (marketing, product and design each run their own tests) and a centre of excellence, a hybrid where a central team owns tools, methods and quality while other teams do the work. Optimizely, a testing vendor, adds that "the best experimentation programs have a central team lead the program", even when many teams run tests.

ModelBest forMain strengthMain risk
CentralisedStarting out, one or two people, or an agency partnerConsistent quality and one clear ownerA bottleneck: the team cannot test everything the business wants
DecentralisedSeveral mature product teams with their own analystsSpeed and ownership close to the productDuplicate tests, inconsistent metrics and lost learning between teams
Centre of excellenceGrowing and large organisationsShared standards with local speedUnclear budget and ownership if the centre has no mandate

Speero's 2023 benchmark (agency data, 119 respondents) shows why a clear owner matters. Only 17% of programmes at the "aspiring" level, the second of five, strongly agreed they had a dedicated person responsible for experimentation, against 92% at the most mature, "transformative" level. And 91% of companies with no dedicated team had no knowledge base of past tests. Without an owner, results live in slide decks and leave with the people who made them.

Section 6 · Trust

A culture of testing collapses if people stop trusting the results, so fix the basics first

Kohavi, Tang and Xu sum up their book on online experiments with a warning: "Getting numbers is easy; getting numbers you can trust is hard." One broken test, with a tracking error or a result declared too early, can undo months of culture building. A sceptical executive only needs one example to say "see, testing doesn't work here."

Many programmes are not ready for that scrutiny. Speero, a consultancy that runs maturity audits, reports on 154 programmes that completed its audit in 2024 (self-reported, agency data).

Horizontal bar chart of experimentation programme gaps among 154 programmes audited by Speero in 2024. No programme-level metrics 66%. No clear prioritisation framework 58%. No quality assurance of tests before launch 52%. Only 15% have well-resourced research methods and only 12% rate their strategy and culture as transformative.
Exhibit 4. Two-thirds of programmes have no programme-level metrics and half do not check tests before launch. Source: Speero, Experimentation Maturity Benchmark Report 2025, Methods & Process (agency data, 154 audit respondents, 2024).

What this shows. The gaps are not advanced statistics. They are basic hygiene: checking that a test works before launch, choosing tests with a shared method, and tracking whether the programme as a whole is improving. These are cheap to fix and they protect the programme's credibility.

Four basics that make results trustworthy

  • Quality assurance before launch. Check every variant on the main devices and browsers, and check that tracking fires. A broken variant produces a false loser and wastes weeks of traffic.
  • A written test plan. Before launch, record the hypothesis, the main metric, guardrail metrics (such as revenue per visitor or error rates that must not get worse), the sample size and the stop date. This prevents "peeking" and moving the goalposts. Our guide to A/B test statistics explains how to size and read a test.
  • A shared prioritisation method. A simple score for expected impact, confidence and effort stops the loudest voice from choosing the roadmap.
  • A second pair of eyes on results. A short review by someone who did not run the test catches most analysis errors and makes results harder to challenge later.

Trust also depends on the data underneath. If analytics is incomplete or consent rules cut out part of the audience without anyone knowing, test results inherit the problem. Our essential guide to A/B testing covers the eight-step process that keeps tests clean.

Section 7 · Scaling

Scale by giving more people the right to test, within guardrails, not by adding approvals

The companies with the strongest cultures let almost anyone run a test. At Booking.com, Thomke reports that "anybody can launch an experiment without permission from management." A paper by Booking.com's own team describes how "all members of our departments run and analyse more than a thousand concurrent experiments", made possible by "safeguards to enable anyone to have end to end ownership of their experiments" and "a central repository of successes and failures".

The key word is safeguards. Democratising testing does not mean removing standards. It means building the standards into the process so that approvals are no longer needed. Netflix describes the result well: "Instead of small groups of executives or experts contributing to a decision, experimentation gives all our members the opportunity to vote, with their actions…"

GuardrailWhat it preventsHow to set it up
Guardrail metricsA test that lifts clicks while hurting revenue or page speedDefine 2 to 4 metrics every test must report, whatever its goal
Test templatesMissing hypotheses, unclear metrics, no sample sizeOne short form, required before launch
Automatic alertsBroken variants or sample ratio mismatch running for weeksUse your tool's alerts or a simple daily check
Ethics and brand rulesTests that mislead customers or break brand or legal rulesA short list of what is never tested, such as hidden fees or fake scarcity
A searchable library of resultsRepeating old tests and forgetting what was learnedOne record per test: hypothesis, screenshots, result, decision

A library of results is the most often neglected of these. Speero's 2023 benchmark (agency data) found that 58% of respondents had no testing knowledge base. Without one, a company pays twice for the same lesson, and new staff have no way to learn from what came before. Our focus on the Experimentation Manifesto explains why collective, long-term learning beats individual heroics.

For leaders. Ask for one number each quarter: how many teams ran at least one test that met the programme's quality standard. It measures spread and discipline at the same time.

Section 8 · Obstacles

The biggest obstacles are traffic, time and buy-in, and each has a practical answer

When companies explain why testing does not happen, the tool is rarely the main problem. Ascend2 surveyed 402 marketing decision-makers in 2025 and asked about the main challenges with A/B testing.

Horizontal bar chart of the main challenges companies face with A/B testing, from an Ascend2 survey of 402 marketing decision-makers in 2025. Limited traffic for statistically significant results 51%. Lack of resources such as time, tools or budget 47%. Takes too long to set up, run and analyse 38%. Privacy and data compliance concerns 37%. Unclear or inconclusive results 34%. Testing tools don't fit our needs 26%.
Exhibit 5. Half of companies say limited traffic holds them back, and almost half lack time, tools or budget. Source: Ascend2, A/B Testing in Marketing research report (June 2025, 402 marketing decision-makers, several answers possible); grouping by Henkan & Partners.

What this shows. The top obstacles are capacity and process, not technology. Only about a quarter blame their tools. Earlier research points the same way: CXL and Speero's 2020 State of Conversion Optimization report found that the two most common challenges were "the need for better processes and buy-in from decision-makers."

ObstacleWhat usually works
Limited trafficTest bigger, bolder changes on high-traffic pages; use a primary metric closer to the change (such as add-to-cart) with revenue as a guardrail; run fewer tests for longer; combine tests with research such as user interviews and session analysis
No time or budgetStart with one test a month and a fixed weekly slot; use an agency or fractional team to cover development and analysis until the value is proven
Too slow to set upReusable templates, a shared QA checklist and pre-built audiences; remove approval steps that the guardrails already cover
Inconclusive resultsWrite sharper hypotheses based on research; size tests before launch; treat "no difference" as a valid answer that frees the team to move on
Lack of buy-inShare a short monthly summary of decisions made and mistakes avoided; invite sceptics to predict results before tests end

Small teams often assume culture is a big-company topic. It is the opposite. A small team cannot afford to spend months building the wrong thing, and it usually has fewer layers of opinion to get through. What changes with size is the method. With little traffic, research and qualitative evidence carry more weight, and tests are reserved for the changes with the highest stakes. Our Voice of Customer guide covers the research side.

Section 9 · Measuring culture

Measure the culture with a few programme metrics, not only the number of tests

What gets measured shapes behaviour, so the choice of programme metrics is itself a cultural decision. Counting tests alone rewards small, safe changes. Counting winners rewards optimistic analysis. A balanced set looks at speed, quality, learning and impact together.

Speed still matters, because a team that runs three tests a year learns very slowly. Optimizely's analysis of 127,000 experiments on its platform gives a useful reference point.

Bar chart of experiments run per year by Optimizely customers. The median company runs 34 experiments a year. Companies need about 200 a year to be in the top 10% by experiment velocity. The top 3% run over 500.
Exhibit 6. The median company runs about three tests a month; about 200 a year is needed to reach the top 10%. Source: Optimizely, Top 10 takeaways from running 127,000 experiments (vendor data, Optimizely customers).

What this shows. There is a wide gap between typical and leading programmes. But velocity is the result of culture, not a target to chase on its own. The same Optimizely analysis found that only 12% of experiments produce a statistically significant improvement on the primary metric, and that tests with four variations deliver 3.5 times the expected impact of a typical A/B test. Bolder, better-researched tests often beat more tests.

MetricWhat it tells youWarning sign
Tests completed per monthSpeed of learningFalling steadily, or rising while quality falls
Learning rateShare of tests that gave a clear, decision-ready answerMany "inconclusive" tests: hypotheses or sample sizes need work
Share of key changes testedWhether important decisions go through a testBig launches that skip testing
Time from idea to decisionHow much friction the process addsWeeks spent waiting for development or approval
Share of tests with QA and a written planTrust in the resultsBelow 100%
Confirmed impactRevenue or retention from shipped winners, and losses avoidedImpact claimed but never checked after launch
Teams running testsHow far the culture has spreadAll tests still come from one person

For leaders. Review these metrics once a quarter with the same seriousness as a sales review. Speero found that two-thirds of programmes have no programme-level metrics at all, so tracking even a few of these puts you ahead of most.

Section 10 · What to do next

The first 90 days decide whether testing becomes a habit or a side project

Culture changes through repeated behaviour, not announcements. The steps below work for a founder running tests alone as well as for a large retailer, at a different scale.

TeamStart withThen add
One or two peopleA sponsor, one primary metric, one well-researched test a month and a monthly one-page summary of what was learnedA simple results library and a prediction log; an agency or fractional team for development
Growing e-commerce teamA named programme owner, a test plan template, QA checklist and a fortnightly test review open to allGuardrail metrics, a prioritisation score and programme metrics reviewed each quarter
Large or multi-brand organisationA centre of excellence with a mandate and budget, shared metrics and standards across brandsSelf-service testing for trained teams, automatic alerts, and learning rate by team

1. Name a sponsor and a single owner

One senior sponsor protects the programme and makes the decisions stick; one owner runs it day to day. Without both, testing competes with every other priority and loses.

2. Agree what success means

Choose one primary metric that reflects customer value, such as revenue per visitor or repeat purchase rate, and two or three guardrails. Write them down so that every test is judged the same way.

3. Run a few tests on real questions

Pick three to five tests from questions leaders actually argue about, using research to shape each hypothesis. Record everyone's prediction before launch. Whatever the results, they will show why testing matters.

4. Make every result public

Hold a short, regular review where each test is presented with its hypothesis, result and decision, including losers and flat results. Store each one in a searchable library.

5. Change what you reward

Praise clear answers and stopped mistakes, not only wins. Add learning rate and share of key changes tested to the quarterly review, and ask leaders to put at least one of their own ideas to a test.

The Experimentation Manifesto sets out the values behind this approach in more depth. For teams who want help, Henkan & Partners supports programmes of every size, from training one person to running a fully embedded experimentation team.

FAQ

Frequently asked questions about building a culture of experimentation

Frequently asked questions

What is a culture of experimentation?

A culture of experimentation is a way of working in which ideas are treated as hypotheses and tested with real customers before the company commits to them. Results, including negative ones, change decisions, and senior people accept being proved wrong. It shows in behaviour: who can launch a test, how failures are discussed and what gets rewarded.

How do you build a culture of experimentation?

Start with a senior sponsor and a single owner, agree one primary success metric with a few guardrails, run a small number of well-researched tests on questions leaders care about, and share every result, including failures. Then reward learning rather than wins, fix basics such as quality checks and test plans, and gradually let more teams run tests within shared guardrails.

Why do most A/B tests fail?

Because most ideas do not change customer behaviour as much as people expect. Published success rates range from about 8% at Airbnb Search to about 33% at Microsoft. That is normal. A failed or flat test still has value: it stops the company from shipping a change that would not have helped, or that would have caused harm.

What is a HiPPO in experimentation?

HiPPO stands for the Highest Paid Person's Opinion. It describes decisions made on the authority of the most senior person rather than on evidence. Researchers such as Ron Kohavi and Stefan Thomke name it as one of the main obstacles to an experimentation culture.

How many A/B tests should a company run per year?

It depends on traffic and team size. In Optimizely's analysis of 127,000 experiments, the median company ran 34 tests a year, and about 200 a year was needed to reach the top 10%. A small site may only support one or two tests a month. Quality and learning matter more than volume: a few well-researched tests teach more than many small ones.

Should experimentation be centralised or decentralised?

Most companies should start with a central owner or team to set standards and build trust, then move towards a centre of excellence in which a central team owns tools, methods and quality while product and marketing teams run their own tests. Fully decentralised testing works only once metrics, guardrails and reviews are shared.

Can small companies build an experimentation culture?

Yes. Culture is about how decisions are made, not about how many tests are run. A small team can test its highest-stakes changes, use customer research where traffic is too low for tests, log predictions and review what it learned each month. An observational study of more than 35,000 start-ups found that those adopting A/B testing improved performance by 30% to 100% after a year.

Key terms

A/B test (controlled experiment)
A test in which visitors are randomly split between the current version (control) and one or more changed versions, so that differences in results can be attributed to the change. It is the basic tool of an experimentation culture.
Experimentation culture
A way of working in which decisions are settled by evidence from customers rather than by rank. It matters because tools only create value when people act on the results.
HiPPO
The Highest Paid Person's Opinion. Decisions made by seniority rather than evidence are one of the main reasons testing programmes stall.
Win rate
The share of experiments in which a variant beat the control on the main metric. It is usually low, which is why it should not be the only measure of success.
Learning rate
The share of experiments that produce a clear, decision-ready answer, including "do not ship". Spotify uses it to reward honest, useful tests.
Guardrail metric
A metric that must not get worse during a test, such as revenue per visitor or page speed. Guardrails let more people test safely without extra approvals.
Overall evaluation criterion (OEC)
The main metric, or small set of metrics, used to judge every experiment. A stable OEC keeps results comparable and stops teams from picking the metric that makes them look good.
Psychological safety
A shared belief that a team is safe for interpersonal risk taking. Without it, people hide failed tests and stop proposing bold ideas.
Centre of excellence
A central team that owns tools, methods, training and quality while other teams run their own tests. It combines shared standards with local speed.
Crawl, Walk, Run, Fly
Four phases of experimentation maturity described by Microsoft researchers, from a single central team to autonomous test owners in every product team.
Velocity
The number of experiments completed in a period. It measures speed of learning but can mislead if quality or boldness falls.
Two-way door decision
A reversible decision that can be made quickly and changed if it proves wrong. Most website changes are two-way doors, which makes them ideal for testing.
Knowledge base (results library)
A searchable record of past tests with hypotheses, results and decisions. It stops teams repeating old tests and helps new staff learn.
Quality assurance (QA)
Checks made before a test goes live to confirm that every variant works and is tracked correctly. It protects trust in the results.

Sources

Methodology. This focus was researched in September 2026 from peer-reviewed papers and conference papers (Kohavi, Deng and Vermeer; Fabijan et al.; Koning, Hasan and Chatterji; Edmondson), Harvard Business Review articles and interviews with Stefan Thomke, engineering blogs from Booking.com, Microsoft, Netflix and Spotify, Amazon's shareholder letters, and industry surveys and benchmarks. Vendor and agency data (Optimizely, Analytics Toolkit, Speero, Ascend2, Convert, Spotify Confidence) is labelled as such. Every source was opened and checked on 27 September 2026. The tables of recommendations and the team-size row in Exhibit 3 are Henkan & Partners frameworks.

  1. Kohavi and Thomke, The Surprising Power of Online Experiments, Harvard Business Review, 2017 (reprint)
  2. Thomke, Building a Culture of Experimentation, Harvard Business Review, 2020
  3. Thomke, The Critical Role of Leadership in Building a Culture of Experimentation, HBS Executive Education, 2024
  4. HBR podcast, At Booking.com, Innovation Means Constant Failure, 2019
  5. The CEO Magazine, Stefan Thomke: Why experimentation is key to finding success, 2023
  6. Kohavi, Deng and Vermeer, A/B Testing Intuition Busters, KDD 2022
  7. Kohavi, Crook and Longbotham, Online Experimentation at Microsoft, 2009
  8. Kohavi et al., Online Controlled Experiments at Large Scale, KDD 2013
  9. Kohavi, Tang and Xu, Trustworthy Online Controlled Experiments, Cambridge University Press, 2020
  10. Fabijan, Dmitriev, Olsson and Bosch, The Evolution of Continuous Experimentation in Software Product Development, ICSE 2017
  11. Fabijan, Dmitriev, Olsson and Bosch, The Benefits of Controlled Experimentation at Scale, SEAA 2017
  12. Gupta, Kohavi, Tang, Xu et al., Top Challenges from the first Practical Online Controlled Experiments Summit, SIGKDD Explorations, 2019
  13. Koning, Hasan and Chatterji, Experimentation and Start-up Performance: Evidence from A/B Testing, Management Science, 2022
  14. Duke Fuqua Insights podcast, Can 1% Improvements Transform Your Business?, 2025
  15. Edmondson, Psychological Safety and Learning Behavior in Work Teams, Administrative Science Quarterly, 1999
  16. Google re:Work, Understand team effectiveness (Project Aristotle)
  17. Amazon.com, 2015 Letter to Shareholders
  18. Kaufman, Pitchforth and Vermeer, Democratizing online controlled experiments at Booking.com, 2017
  19. Tingley et al., Decision Making at Netflix, Netflix Tech Blog, 2021
  20. Bellato, Schultzberg and Ankargren, Beyond Winning: Spotify's Experiments with Learning Framework, Spotify Engineering, 2025
  21. Spotify Confidence, The Real ROI of Experimentation, 2026
  22. Optimizely, Top 10 takeaways from running 127,000 experiments
  23. Analytics Toolkit, What Can Be Learned From 1,001 A/B Tests?, 2022
  24. Speero, The State of Experimentation Programs 2023
  25. Speero, Experimentation Maturity Benchmark Report 2025, Methods & Process
  26. Convert, A/B Testing & CRO Stats Every Optimizer Should Know, 2026
  27. Ascend2, A/B Testing in Marketing research report, 2025
  28. CXL/Speero, The 2020 State of Conversion Optimization Report
  29. CXL, How to Structure Your Optimization and Experimentation Teams
  30. Optimizely, Enabling Experimentation at Your Organization: Determining Your Team Structure, 2019
  31. Henkan & Partners, The Experimentation Manifesto
  32. Henkan & Partners, The Essential Guide to A/B Testing