Focus
How to Build a Culture of Experimentation
Alexandre Suon · 2026-09-27
Buying an A/B testing tool takes a week. Getting a company to let evidence overrule opinion takes years. This focus explains what an experimentation culture is, why most programmes stall, and the concrete steps that build one, whether your team is one person or one hundred.
Executive summary
- A culture of experimentation means decisions are settled by evidence from customers, not by rank. Tools matter, but the difference between companies that test and companies that learn is who is allowed to be wrong, what gets rewarded and whether leaders put their own ideas to the test.
- Most ideas fail, so a culture that punishes failure cannot learn. Published success rates range from about 8% of experiments at Airbnb Search to about a third at Microsoft. A team that expects every test to win will either stop testing or stop telling the truth about results.
- Leaders set the culture more than any process does. Researchers name the highest-paid person's opinion (HiPPO) as one of the main obstacles to testing. Leaders who say "I don't know, let's test it", test their own ideas and celebrate tests that stopped a bad change do more than any training plan.
- Reward learning, not wins. Spotify reports a 12% win rate but a 64% learning rate. Measuring learning, trust in the data and speed to decision, not only the number of winners, is what keeps teams honest and motivated.
- Build trust before scale, and scale through guardrails, not permission. Most programmes lack basics such as quality checks, prioritisation and programme metrics. Fix those first, then let more teams run tests with shared metrics, reviews and a searchable library of results.
- Every team size can start. One person can run a monthly test-and-learn review; a larger organisation can build a centre of excellence. The first 90 days matter more than the tool: pick a sponsor, agree one success metric, run a few well-chosen tests and share every result.
An experimentation culture is a way of working in which a company treats ideas as hypotheses, tests them with real customers before committing, and lets the results, including negative ones, change its decisions. It is visible in behaviour: who can launch a test, how failures are discussed, and whether senior people accept being proved wrong.
Section 1 · Why culture
Most ideas fail when tested, so the real challenge is accepting what the tests say
Every company says it is data-driven. The test is what happens when the data disagrees with a senior person, or when a project that took six months shows no effect. In a culture of experimentation the result wins. In most companies, the result gets explained away.
The reason this matters is simple: most ideas do not work. In their 2017 Harvard Business Review article, Ron Kohavi and Stefan Thomke reported that "at Google and Bing, only about 10% to 20% of experiments generate positive results", and that at Microsoft as a whole "one-third prove effective, one-third have neutral results, and one-third have negative results." Later work by Kohavi and colleagues compiled published success rates from several companies. The picture is consistent: a clear majority of carefully chosen ideas fail to move the metric they were designed to move.

What this shows. Failure is the normal outcome of testing, not a sign that the programme is broken. Success rates also depend on how mature a product is and how a "win" is defined, so they are not league tables. What they share is the lesson for culture: if people are judged on whether their ideas win, most of them will look bad most of the time.
People are also poor judges of which ideas will work. In one exercise at Microsoft, "only three out of 21 people guessed the winner, and the three were from the ExP team", the company's own experimentation group. At Booking.com, Stefan Thomke reports that teams are "wrong about nine out of ten times" when they predict customer behaviour. A test that overturns a confident prediction is not an embarrassment. It is the programme doing its job.
So the hard part is not running tests. It is building an organisation that expects to be wrong, finds out quickly and cheaply, and changes course without blame. That is what we mean by culture, and it is why tools alone rarely deliver.
For leaders. Before asking how many tests the team ran, ask how many decisions changed because of a test, and how many ideas were stopped before they cost money. If the answer is "none", the programme is producing reports, not learning.
Section 2 · The payoff
Companies that make testing a habit grow faster, and the evidence goes beyond big tech
The most cited examples come from large technology firms, but the most useful evidence comes from ordinary companies. Rembrand Koning, Sharique Hasan and Aaron Chatterji tracked more than 35,000 start-ups in an observational study published in Management Science in 2022. They found that relatively few firms adopt A/B testing, but "among those that do, performance improves by 30%–100% after a year of use." Adopters also "develop more new products, identify and scale promising ideas, and fail faster when they receive negative signals."
At large scale the numbers are large. Kohavi and Thomke describe how one of hundreds of ideas at Bing, a small change to how ad headlines were displayed, was judged a low priority and "languished for more than six months". When it was finally tested, it increased revenue by 12%, worth "more than $100 million" a year in the United States alone. They also report that dozens of tested revenue changes each month collectively increased Bing's revenue per search "by 10% to 25% each year." Microsoft researchers estimate that scaling experimentation across the company is worth "hundreds of millions of dollars of additional revenue annually."
The value is not only in winners. Tests also stop expensive mistakes. The same HBR article notes that integrating Bing with social media "cost Microsoft more than $25 million to develop and produced negligible increases in engagement and revenue." A culture that tests early catches such projects before the full investment is made.
| Type of value | What it looks like | Who notices |
|---|---|---|
| Better decisions | Changes that improve conversion, revenue or retention are kept; the rest are dropped | E-commerce and product leads |
| Avoided losses | Ideas that would have hurt customers or revenue are stopped before full launch | Finance and leadership, if someone reports it |
| Faster learning | Teams know more about what customers value, which improves the next idea | Product, design and marketing teams |
| Less politics | Disagreements are settled by a test instead of by rank or persistence | Everyone, over time |
For an online shop, these benefits link directly to customer experience. Each test that removes friction, clarifies delivery costs or improves search makes the site easier to use. Over time, customers buy more often and stay longer, which is where lifetime value comes from.
Section 3 · Leadership
Leaders build the culture by testing their own ideas and saying "I don't know"
Researchers and practitioners name one obstacle again and again: the HiPPO, the "Highest Paid Person's Opinion". Kohavi and colleagues describe an early stage of "hubris", where "measurement is not needed because of confidence in the HiPPO". Thomke puts it more bluntly: "Nothing stalls innovation faster than a so-called HiPPO." When the most senior person in the room decides, tests become a formality, and people learn to propose only what the boss already likes.
Thomke describes a different leadership role. In an experimentation organisation, he writes, leadership "facilitates the process of decision-making through experimentation" instead of making every decision top-down, and "even the boss's assumptions are subject to real-world tests." He adds, in an interview, that bosses "ought to display intellectual humility and be unafraid to admit, 'I don't know.'"
Amazon's founder made the same point in his shareholder letters. In the 2015 letter he wrote that "failure and invention are inseparable twins. To invent you have to experiment, and if you know in advance that it's going to work, it's not an experiment." He also separated decisions into two types: "one-way doors" that are hard to reverse and deserve careful deliberation, and "two-way doors" that can be reversed and should be made quickly by small groups. Most changes to a website are two-way doors, which is exactly why they can be tested instead of debated.
Five things leaders can do this quarter
- Put one of your own ideas to a test, publicly. Nothing signals the new rules faster than a senior person whose idea loses and who says so in the team meeting.
- Change the question you ask. Replace "did it win?" with "what did we learn, and what will we do differently?"
- Name a sponsor with a budget. Someone senior must own the programme, protect testing time and remove blockers. Speero's 2025 benchmark (agency data, as reported by Convert) found only 26% of companies strongly agree they have a senior sponsor accountable for experimentation.
- Agree which decisions must be tested. At Bing, Kohavi and Thomke reported in 2017 that about 80% of proposed changes were first run as controlled experiments. You do not need that level, but you need a rule, for example: every change to checkout, pricing display or navigation is tested.
- Report stopped mistakes, not only wins. A test that prevented a revenue loss deserves the same visibility as a winner.
For leaders. Thomke's line is worth repeating in your next leadership meeting: in a true experimentation organisation, "even the boss's assumptions are subject to real-world tests." The culture changes the first time that happens in front of the team.
Section 4 · Safety and rewards
People only report honest results when failure is safe and learning is rewarded
A test programme is only as good as the honesty of its reporting. If a product owner's bonus depends on a feature succeeding, a flat result will be "re-analysed" until it looks like a win. That is how programmes lose credibility.
The research on teams points the same way. Amy Edmondson defined team psychological safety in 1999 as "a shared belief held by members of a team that the team is safe for interpersonal risk taking", and, in a study of 51 teams in a manufacturing company, found that it is linked to learning behaviour, which in turn is linked to performance. Google's Project Aristotle, which studied 180 teams, listed psychological safety first among the five dynamics of effective teams, ahead of dependability, structure and clarity, meaning and impact.
The practical question is what you reward. Spotify offers a useful model. Its engineers found that judging experiments by wins alone gave a misleading picture, so they introduced a "learning rate": the share of experiments that produce a clear, decision-ready answer, including answers such as do not ship.

What this shows. Measured by wins, Spotify's programme looks like it fails almost nine times out of ten. Measured by learning, it succeeds almost two times out of three. As Spotify's team puts it, most of their learning does not come from wins; it comes "from discovering what not to ship." The wide range between teams, 16% to 76%, also shows that learning rate is something a team can improve through better hypotheses and cleaner tests.
Thomke calls failure "the status quo" at Booking.com and says "not winning is not losing." A 2019 paper by experts from 13 organisations, including Airbnb, Booking.com, Google, LinkedIn, Microsoft, Netflix and Stanford University, describes what a mature culture looks like: negative experiment results are "celebrated as saving customers".
| Instead of rewarding… | Reward… | Why |
|---|---|---|
| Number of winning tests | Share of tests with a clear, decision-ready answer | Keeps reporting honest |
| Size of the lift claimed | Lift confirmed after launch, or reviewed by a second analyst | Stops inflated results |
| Ideas that were shipped | Ideas that were stopped early because a test showed harm | Makes saving money visible |
| Individual heroes | Teams that share results others can reuse | Knowledge survives when people leave |
For marketers and product owners. Write down your prediction before each test and compare it with the result afterwards. Over a quarter, this simple habit shows how often intuition is wrong, and it makes the case for testing better than any slide.
Section 5 · Operating model
Start with one central owner, then spread testing to every team as trust grows
Culture needs a structure to live in. The most useful description comes from Microsoft researchers Aleksander Fabijan, Pavel Dmitriev, Helena Holmström Olsson and Jan Bosch, who interviewed product teams across the company and described four phases of maturity: Crawl, Walk, Run and Fly. At each phase, both the organisation and the way success is measured change.

What this shows. Maturity is not about tool features. It is about who is trusted to run tests and how clearly the company defines success. Most companies should not try to jump from Crawl to Fly: autonomy without shared metrics and reviews produces many tests nobody trusts.
Three operating models are common. CXL describes them as centralised (one team runs everything), decentralised (marketing, product and design each run their own tests) and a centre of excellence, a hybrid where a central team owns tools, methods and quality while other teams do the work. Optimizely, a testing vendor, adds that "the best experimentation programs have a central team lead the program", even when many teams run tests.
| Model | Best for | Main strength | Main risk |
|---|---|---|---|
| Centralised | Starting out, one or two people, or an agency partner | Consistent quality and one clear owner | A bottleneck: the team cannot test everything the business wants |
| Decentralised | Several mature product teams with their own analysts | Speed and ownership close to the product | Duplicate tests, inconsistent metrics and lost learning between teams |
| Centre of excellence | Growing and large organisations | Shared standards with local speed | Unclear budget and ownership if the centre has no mandate |
Speero's 2023 benchmark (agency data, 119 respondents) shows why a clear owner matters. Only 17% of programmes at the "aspiring" level, the second of five, strongly agreed they had a dedicated person responsible for experimentation, against 92% at the most mature, "transformative" level. And 91% of companies with no dedicated team had no knowledge base of past tests. Without an owner, results live in slide decks and leave with the people who made them.
Section 6 · Trust
A culture of testing collapses if people stop trusting the results, so fix the basics first
Kohavi, Tang and Xu sum up their book on online experiments with a warning: "Getting numbers is easy; getting numbers you can trust is hard." One broken test, with a tracking error or a result declared too early, can undo months of culture building. A sceptical executive only needs one example to say "see, testing doesn't work here."
Many programmes are not ready for that scrutiny. Speero, a consultancy that runs maturity audits, reports on 154 programmes that completed its audit in 2024 (self-reported, agency data).

What this shows. The gaps are not advanced statistics. They are basic hygiene: checking that a test works before launch, choosing tests with a shared method, and tracking whether the programme as a whole is improving. These are cheap to fix and they protect the programme's credibility.
Four basics that make results trustworthy
- Quality assurance before launch. Check every variant on the main devices and browsers, and check that tracking fires. A broken variant produces a false loser and wastes weeks of traffic.
- A written test plan. Before launch, record the hypothesis, the main metric, guardrail metrics (such as revenue per visitor or error rates that must not get worse), the sample size and the stop date. This prevents "peeking" and moving the goalposts. Our guide to A/B test statistics explains how to size and read a test.
- A shared prioritisation method. A simple score for expected impact, confidence and effort stops the loudest voice from choosing the roadmap.
- A second pair of eyes on results. A short review by someone who did not run the test catches most analysis errors and makes results harder to challenge later.
Trust also depends on the data underneath. If analytics is incomplete or consent rules cut out part of the audience without anyone knowing, test results inherit the problem. Our essential guide to A/B testing covers the eight-step process that keeps tests clean.
Section 7 · Scaling
Scale by giving more people the right to test, within guardrails, not by adding approvals
The companies with the strongest cultures let almost anyone run a test. At Booking.com, Thomke reports that "anybody can launch an experiment without permission from management." A paper by Booking.com's own team describes how "all members of our departments run and analyse more than a thousand concurrent experiments", made possible by "safeguards to enable anyone to have end to end ownership of their experiments" and "a central repository of successes and failures".
The key word is safeguards. Democratising testing does not mean removing standards. It means building the standards into the process so that approvals are no longer needed. Netflix describes the result well: "Instead of small groups of executives or experts contributing to a decision, experimentation gives all our members the opportunity to vote, with their actions…"
| Guardrail | What it prevents | How to set it up |
|---|---|---|
| Guardrail metrics | A test that lifts clicks while hurting revenue or page speed | Define 2 to 4 metrics every test must report, whatever its goal |
| Test templates | Missing hypotheses, unclear metrics, no sample size | One short form, required before launch |
| Automatic alerts | Broken variants or sample ratio mismatch running for weeks | Use your tool's alerts or a simple daily check |
| Ethics and brand rules | Tests that mislead customers or break brand or legal rules | A short list of what is never tested, such as hidden fees or fake scarcity |
| A searchable library of results | Repeating old tests and forgetting what was learned | One record per test: hypothesis, screenshots, result, decision |
A library of results is the most often neglected of these. Speero's 2023 benchmark (agency data) found that 58% of respondents had no testing knowledge base. Without one, a company pays twice for the same lesson, and new staff have no way to learn from what came before. Our focus on the Experimentation Manifesto explains why collective, long-term learning beats individual heroics.
For leaders. Ask for one number each quarter: how many teams ran at least one test that met the programme's quality standard. It measures spread and discipline at the same time.
Section 8 · Obstacles
The biggest obstacles are traffic, time and buy-in, and each has a practical answer
When companies explain why testing does not happen, the tool is rarely the main problem. Ascend2 surveyed 402 marketing decision-makers in 2025 and asked about the main challenges with A/B testing.

What this shows. The top obstacles are capacity and process, not technology. Only about a quarter blame their tools. Earlier research points the same way: CXL and Speero's 2020 State of Conversion Optimization report found that the two most common challenges were "the need for better processes and buy-in from decision-makers."
| Obstacle | What usually works |
|---|---|
| Limited traffic | Test bigger, bolder changes on high-traffic pages; use a primary metric closer to the change (such as add-to-cart) with revenue as a guardrail; run fewer tests for longer; combine tests with research such as user interviews and session analysis |
| No time or budget | Start with one test a month and a fixed weekly slot; use an agency or fractional team to cover development and analysis until the value is proven |
| Too slow to set up | Reusable templates, a shared QA checklist and pre-built audiences; remove approval steps that the guardrails already cover |
| Inconclusive results | Write sharper hypotheses based on research; size tests before launch; treat "no difference" as a valid answer that frees the team to move on |
| Lack of buy-in | Share a short monthly summary of decisions made and mistakes avoided; invite sceptics to predict results before tests end |
Small teams often assume culture is a big-company topic. It is the opposite. A small team cannot afford to spend months building the wrong thing, and it usually has fewer layers of opinion to get through. What changes with size is the method. With little traffic, research and qualitative evidence carry more weight, and tests are reserved for the changes with the highest stakes. Our Voice of Customer guide covers the research side.
Section 9 · Measuring culture
Measure the culture with a few programme metrics, not only the number of tests
What gets measured shapes behaviour, so the choice of programme metrics is itself a cultural decision. Counting tests alone rewards small, safe changes. Counting winners rewards optimistic analysis. A balanced set looks at speed, quality, learning and impact together.
Speed still matters, because a team that runs three tests a year learns very slowly. Optimizely's analysis of 127,000 experiments on its platform gives a useful reference point.

What this shows. There is a wide gap between typical and leading programmes. But velocity is the result of culture, not a target to chase on its own. The same Optimizely analysis found that only 12% of experiments produce a statistically significant improvement on the primary metric, and that tests with four variations deliver 3.5 times the expected impact of a typical A/B test. Bolder, better-researched tests often beat more tests.
| Metric | What it tells you | Warning sign |
|---|---|---|
| Tests completed per month | Speed of learning | Falling steadily, or rising while quality falls |
| Learning rate | Share of tests that gave a clear, decision-ready answer | Many "inconclusive" tests: hypotheses or sample sizes need work |
| Share of key changes tested | Whether important decisions go through a test | Big launches that skip testing |
| Time from idea to decision | How much friction the process adds | Weeks spent waiting for development or approval |
| Share of tests with QA and a written plan | Trust in the results | Below 100% |
| Confirmed impact | Revenue or retention from shipped winners, and losses avoided | Impact claimed but never checked after launch |
| Teams running tests | How far the culture has spread | All tests still come from one person |
For leaders. Review these metrics once a quarter with the same seriousness as a sales review. Speero found that two-thirds of programmes have no programme-level metrics at all, so tracking even a few of these puts you ahead of most.
Section 10 · What to do next
The first 90 days decide whether testing becomes a habit or a side project
Culture changes through repeated behaviour, not announcements. The steps below work for a founder running tests alone as well as for a large retailer, at a different scale.
| Team | Start with | Then add |
|---|---|---|
| One or two people | A sponsor, one primary metric, one well-researched test a month and a monthly one-page summary of what was learned | A simple results library and a prediction log; an agency or fractional team for development |
| Growing e-commerce team | A named programme owner, a test plan template, QA checklist and a fortnightly test review open to all | Guardrail metrics, a prioritisation score and programme metrics reviewed each quarter |
| Large or multi-brand organisation | A centre of excellence with a mandate and budget, shared metrics and standards across brands | Self-service testing for trained teams, automatic alerts, and learning rate by team |
1. Name a sponsor and a single owner
One senior sponsor protects the programme and makes the decisions stick; one owner runs it day to day. Without both, testing competes with every other priority and loses.
2. Agree what success means
Choose one primary metric that reflects customer value, such as revenue per visitor or repeat purchase rate, and two or three guardrails. Write them down so that every test is judged the same way.
3. Run a few tests on real questions
Pick three to five tests from questions leaders actually argue about, using research to shape each hypothesis. Record everyone's prediction before launch. Whatever the results, they will show why testing matters.
4. Make every result public
Hold a short, regular review where each test is presented with its hypothesis, result and decision, including losers and flat results. Store each one in a searchable library.
5. Change what you reward
Praise clear answers and stopped mistakes, not only wins. Add learning rate and share of key changes tested to the quarterly review, and ask leaders to put at least one of their own ideas to a test.
The Experimentation Manifesto sets out the values behind this approach in more depth. For teams who want help, Henkan & Partners supports programmes of every size, from training one person to running a fully embedded experimentation team.
FAQ
Frequently asked questions about building a culture of experimentation
Frequently asked questions
What is a culture of experimentation?
A culture of experimentation is a way of working in which ideas are treated as hypotheses and tested with real customers before the company commits to them. Results, including negative ones, change decisions, and senior people accept being proved wrong. It shows in behaviour: who can launch a test, how failures are discussed and what gets rewarded.
How do you build a culture of experimentation?
Start with a senior sponsor and a single owner, agree one primary success metric with a few guardrails, run a small number of well-researched tests on questions leaders care about, and share every result, including failures. Then reward learning rather than wins, fix basics such as quality checks and test plans, and gradually let more teams run tests within shared guardrails.
Why do most A/B tests fail?
Because most ideas do not change customer behaviour as much as people expect. Published success rates range from about 8% at Airbnb Search to about 33% at Microsoft. That is normal. A failed or flat test still has value: it stops the company from shipping a change that would not have helped, or that would have caused harm.
What is a HiPPO in experimentation?
HiPPO stands for the Highest Paid Person's Opinion. It describes decisions made on the authority of the most senior person rather than on evidence. Researchers such as Ron Kohavi and Stefan Thomke name it as one of the main obstacles to an experimentation culture.
How many A/B tests should a company run per year?
It depends on traffic and team size. In Optimizely's analysis of 127,000 experiments, the median company ran 34 tests a year, and about 200 a year was needed to reach the top 10%. A small site may only support one or two tests a month. Quality and learning matter more than volume: a few well-researched tests teach more than many small ones.
Should experimentation be centralised or decentralised?
Most companies should start with a central owner or team to set standards and build trust, then move towards a centre of excellence in which a central team owns tools, methods and quality while product and marketing teams run their own tests. Fully decentralised testing works only once metrics, guardrails and reviews are shared.
Can small companies build an experimentation culture?
Yes. Culture is about how decisions are made, not about how many tests are run. A small team can test its highest-stakes changes, use customer research where traffic is too low for tests, log predictions and review what it learned each month. An observational study of more than 35,000 start-ups found that those adopting A/B testing improved performance by 30% to 100% after a year.
Key terms
- A/B test (controlled experiment)
- A test in which visitors are randomly split between the current version (control) and one or more changed versions, so that differences in results can be attributed to the change. It is the basic tool of an experimentation culture.
- Experimentation culture
- A way of working in which decisions are settled by evidence from customers rather than by rank. It matters because tools only create value when people act on the results.
- HiPPO
- The Highest Paid Person's Opinion. Decisions made by seniority rather than evidence are one of the main reasons testing programmes stall.
- Win rate
- The share of experiments in which a variant beat the control on the main metric. It is usually low, which is why it should not be the only measure of success.
- Learning rate
- The share of experiments that produce a clear, decision-ready answer, including "do not ship". Spotify uses it to reward honest, useful tests.
- Guardrail metric
- A metric that must not get worse during a test, such as revenue per visitor or page speed. Guardrails let more people test safely without extra approvals.
- Overall evaluation criterion (OEC)
- The main metric, or small set of metrics, used to judge every experiment. A stable OEC keeps results comparable and stops teams from picking the metric that makes them look good.
- Psychological safety
- A shared belief that a team is safe for interpersonal risk taking. Without it, people hide failed tests and stop proposing bold ideas.
- Centre of excellence
- A central team that owns tools, methods, training and quality while other teams run their own tests. It combines shared standards with local speed.
- Crawl, Walk, Run, Fly
- Four phases of experimentation maturity described by Microsoft researchers, from a single central team to autonomous test owners in every product team.
- Velocity
- The number of experiments completed in a period. It measures speed of learning but can mislead if quality or boldness falls.
- Two-way door decision
- A reversible decision that can be made quickly and changed if it proves wrong. Most website changes are two-way doors, which makes them ideal for testing.
- Knowledge base (results library)
- A searchable record of past tests with hypotheses, results and decisions. It stops teams repeating old tests and helps new staff learn.
- Quality assurance (QA)
- Checks made before a test goes live to confirm that every variant works and is tracked correctly. It protects trust in the results.
Sources
Methodology. This focus was researched in September 2026 from peer-reviewed papers and conference papers (Kohavi, Deng and Vermeer; Fabijan et al.; Koning, Hasan and Chatterji; Edmondson), Harvard Business Review articles and interviews with Stefan Thomke, engineering blogs from Booking.com, Microsoft, Netflix and Spotify, Amazon's shareholder letters, and industry surveys and benchmarks. Vendor and agency data (Optimizely, Analytics Toolkit, Speero, Ascend2, Convert, Spotify Confidence) is labelled as such. Every source was opened and checked on 27 September 2026. The tables of recommendations and the team-size row in Exhibit 3 are Henkan & Partners frameworks.
- Kohavi and Thomke, The Surprising Power of Online Experiments, Harvard Business Review, 2017 (reprint)
- Thomke, Building a Culture of Experimentation, Harvard Business Review, 2020
- Thomke, The Critical Role of Leadership in Building a Culture of Experimentation, HBS Executive Education, 2024
- HBR podcast, At Booking.com, Innovation Means Constant Failure, 2019
- The CEO Magazine, Stefan Thomke: Why experimentation is key to finding success, 2023
- Kohavi, Deng and Vermeer, A/B Testing Intuition Busters, KDD 2022
- Kohavi, Crook and Longbotham, Online Experimentation at Microsoft, 2009
- Kohavi et al., Online Controlled Experiments at Large Scale, KDD 2013
- Kohavi, Tang and Xu, Trustworthy Online Controlled Experiments, Cambridge University Press, 2020
- Fabijan, Dmitriev, Olsson and Bosch, The Evolution of Continuous Experimentation in Software Product Development, ICSE 2017
- Fabijan, Dmitriev, Olsson and Bosch, The Benefits of Controlled Experimentation at Scale, SEAA 2017
- Gupta, Kohavi, Tang, Xu et al., Top Challenges from the first Practical Online Controlled Experiments Summit, SIGKDD Explorations, 2019
- Koning, Hasan and Chatterji, Experimentation and Start-up Performance: Evidence from A/B Testing, Management Science, 2022
- Duke Fuqua Insights podcast, Can 1% Improvements Transform Your Business?, 2025
- Edmondson, Psychological Safety and Learning Behavior in Work Teams, Administrative Science Quarterly, 1999
- Google re:Work, Understand team effectiveness (Project Aristotle)
- Amazon.com, 2015 Letter to Shareholders
- Kaufman, Pitchforth and Vermeer, Democratizing online controlled experiments at Booking.com, 2017
- Tingley et al., Decision Making at Netflix, Netflix Tech Blog, 2021
- Bellato, Schultzberg and Ankargren, Beyond Winning: Spotify's Experiments with Learning Framework, Spotify Engineering, 2025
- Spotify Confidence, The Real ROI of Experimentation, 2026
- Optimizely, Top 10 takeaways from running 127,000 experiments
- Analytics Toolkit, What Can Be Learned From 1,001 A/B Tests?, 2022
- Speero, The State of Experimentation Programs 2023
- Speero, Experimentation Maturity Benchmark Report 2025, Methods & Process
- Convert, A/B Testing & CRO Stats Every Optimizer Should Know, 2026
- Ascend2, A/B Testing in Marketing research report, 2025
- CXL/Speero, The 2020 State of Conversion Optimization Report
- CXL, How to Structure Your Optimization and Experimentation Teams
- Optimizely, Enabling Experimentation at Your Organization: Determining Your Team Structure, 2019
- Henkan & Partners, The Experimentation Manifesto
- Henkan & Partners, The Essential Guide to A/B Testing