← All Dynamic Yield insights
[CX] Dynamic Yield · Insights

No More Gut Feelings: How successful online shops find out what their customers really want

// Dynamic Yield · Experimentation · A/B Testing // Reading time ~ 7 min // prodct [CX] · Dynamic Yield

In many companies, the loudest voice in the room decides what the shop looks like. At Booking.com, the customer decides. The company runs more than 25,000 tests per year, with over 1,000 running simultaneously at any given time. This difference is no accident — it's the result of a culture any company can build. Here's why testing matters more than any single feature, why testing programs fail in practice, and how to get started.

Why opinions — not customers — drive decisions in many shops

A relaunch is coming up. Marketing wants more emotion on the homepage. Sales wants the bestsellers front and center. The CEO doesn't like the shade of blue. After two hours, there's a compromise that nobody is truly happy with. Three months later, the new site goes live. Whether it converts better than the old one, nobody knows. Nobody thought to measure it.

This is how decisions get made in many companies. Not because the people involved are bad at their jobs. But because no other mechanism exists. Where data is missing, opinions fill the gap. And where opinions decide, the best idea rarely wins. What wins is the person with the biggest title or the loudest voice.

The problem isn't one wrong banner. The problem is the cumulative effect. Every untested decision is a bet placed with company revenue. If you make a hundred such bets a year and never check whether they pay off, you're losing money without knowing it.

What Booking.com does differently from most online shops

Some companies have flipped this dynamic entirely. The most well-known example comes from the travel industry. Booking.com runs more than 25,000 tests per year, with over 1,000 running simultaneously at any given time. Harvard professor Stefan Thomke studied this approach and wrote about it in the Harvard Business Review. His most important finding sounds obvious but isn't: at Booking.com, any employee can launch an experiment without asking management first.

In practice, that means no one has to win a debate to try out an idea. The idea goes live as a test, a portion of visitors sees the new variant, the rest sees the original. After a few weeks, there's a result. The variant wins or it loses. Either way, the company has learned something.

Thomke draws a conclusion worth remembering: data beats opinions. Booking.com was once a small Dutch startup. Today it's the world's largest accommodation platform. It didn't get there through brilliant design — it got there through tens of thousands of small tests, many of which failed. Failure is part of the process. It's cheap there because it happens early and stays contained.

Why testing and personalization pay off

You might be thinking: great for Booking.com, but we operate at a very different scale. That's a fair point — and still a misleading one. Because the value of testing doesn't depend on company size. It depends on whether you want to know what works.

McKinsey looked at what companies actually gain from taking personalization seriously. The finding: companies that excel at personalization generate 40 percent more revenue than the average for their industry. Personalization here means different customers see different content, offers, and recommendations. And that requires testing. Because which variant resonates with which customer segment is something nobody can predict in advance. You can only find out.

There's a second benefit that never shows up in any statistic. Tests end internal debates. Once a result is on the table, nobody needs to be right anymore. The conversation shifts from who gets their way to what to try next. Teams that work this way become noticeably faster — not because they work more, but because they argue less.

Why A/B testing really fails in practice

If testing is so obviously worthwhile, why doesn't everyone do it? In our experience, testing programs rarely fail because of the tools. They fail because of three patterns that come up time and again.

The first pattern: no hypothesis. Whatever just got finished gets tested. The new product page against the old one, simply because the new one exists. A test like that answers no question, because no question was ever asked. A good test starts with a falsifiable assumption. For example: if we place the delivery time directly next to the price, fewer customers will drop off at the cart. That might be true or it might not. But it can be tested.

The second pattern: too little patience. A test runs for three days, the curve looks promising, the team declares it a success. Except three days wasn't enough and the sample was too small. The result is noise, not signal. Drawing decisions from it just replaces gut feeling with a random outcome. That's not progress.

The third pattern is the most common and the most costly: results have no consequences. The test loses, but the variant goes live anyway. Because the leadership likes it. Because the agency already built it. Because it's in the annual plan. The moment that happens for the first time, the entire organization gets the message: the tests are decoration. At that point, you might as well skip them.

All three patterns share the same root cause. Testing is treated as a tool rather than a rule. But the real power lies in accountability. A test whose result is binding changes behavior. A test whose result is open for debate changes nothing.

What getting started with Dynamic Yield looks like

The good news first: you don't have to start with 1,000 parallel tests. One is enough, if it's set up properly.

Find the spot in your shop that generates the most internal debate. There's almost always one. The recommendation module nobody clicks. The category page with the high bounce rate. The promotional banner that's been argued about for months. Write a hypothesis, decide upfront how you'll measure success, and agree as a team that the result stands — whatever it is.

Technically, this is no longer a major hurdle. With a platform like Dynamic Yield, tests like these run across your website, app, and email channel in a single shared workflow. Audiences, variants, and reporting all live in one place. Your team doesn't need an IT project or a release cycle. It needs a question and two weeks of patience.

The first test leads to a second. The second becomes a rhythm. After a few months, something has shifted that's bigger than any individual result. The question in meetings is no longer who's right. It's what the next test is. That's exactly what a culture of experimentation means. Nothing more, nothing less.

How testing becomes personalization with Dynamic Yield

Testing and personalizing are two sides of the same coin. A test tells you that variant A performs better than variant B. Personalization takes it a step further and asks: better for whom? Maybe variant A works for new customers and variant B works for returning ones. In that case, the right answer isn't A or B. The right answer is both — depending on who's on the page.

That's why the sequence matters. Measure first, then personalize, then automate. Personalizing without measuring just spreads your gut feeling across more audience segments. Measuring first and then personalizing builds a system, piece by piece, that improves with every test. Eventually, the machine works for you — learning from every customer interaction.

The first step toward data-driven decisions

You don't need board approval or a new budget cycle. You need one spot in your shop, a testable assumption, and the willingness to accept the result. That's more uncomfortable than another workshop. But it's the start of a way of working that sells, instead of just debating.

And if you're curious what this looks like at massive scale: how McDonald's used the same technology to reshape its ordering channels is covered in our article on the McDonald's playbook. You'll find all further analysis in our Dynamic Yield Insights.

← All Dynamic Yield insights
// More insights