
Ever since A/B testing frameworks emerged in the late 2000’s, online experimentation has become increasingly common in digital products. Even though I was already familiar with the tactical aspects of running A/B tests to compare alternative designs, I knew that the world of experimentation went much deeper. The Power of Experiments: Decision-Making in a Data-Driven World gives the big picture of what’s happening at the forefront of the field. In this post, I’d like to share the key lessons that I learned from the book.
The book starts by tracing the history of experimentation as it’s been adopted by practitioners in different fields, starting in medicine, then by academics in the social sciences, and most currently, by big tech companies. In comparison to academic researchers, big tech is now running significantly more experiments, with many more participants, at a much faster pace. The authors claim that we’re currently in an “experimental revolution” because the number, size, and speed of experiments run by BigTech are orders of magnitude greater than anything that’s come before.
Experimentation at this scale is possible due to two main enabling factors: BigTech’s large, active userbase, and their A/B testing frameworks. In comparison to offline experiments, the whole process online is significantly streamlined. This allows companies to:
- Skip participant recruitment
- Easily randomize treatment groups
- Get high sensitivity with the ability to detect small effects
- Run a series of iterative experiments with progressively refined treatments, and
- Enrich the experimental data with other demographic and behavioral data
With the historical barriers to experimentation significantly lowered, the cost-benefit equation has been totally transformed, so it’s worthwhile for digital product companies to take a close look at the potential benefits of adopting experimentation.
Why run experiments?
As the book shows through several case studies, companies that run experiments to test their hypotheses usually end up learning that they aren’t as good as they thought at intuitively guessing how users will interact with their products. Specialists in product experimentation and optimization frequently report that their intuitive hunches are wrong the majority of the time. Some of the main challenges with predicting how users will behave are: 1) people have a host of psychological biases that can lead to irrational decisions, and 2) human behavior is highly context-sensitive, which makes it easy to overestimate the extent to which a seemingly general principle applies to a specific situation.
This isn’t to say that experimentation should replace intuition. We need to use our intuition and judgment in order to come up with well-informed hypotheses, and also to interpret the results and translate them into decisions. The role of experimentation is simply to compliment our intuition in the overall decision-making process.
So, how do we use online experiments in product development?
This book helped me to appreciate that experimentation has a very broad range of possible applications in product development. My personal experience with product experimentation has been limited mostly to using A/B testing as a way to fine-tune designs that are already in a close-to-ideal state, after the majority of generative and evaluative research have been done. Now I see that experimentation can start much earlier in the product lifecycle, and address a much broader range of questions.
Instead of starting from an already-created design and then testing slightly different variants, it’s better to start by thinking about what outcomes we want to achieve, and then identify all variables that could hypothetically drive those outcomes.
Importance of Experimenting with Multiple Variables
The book’s case studies highlight the value of using experimentation to evaluate multiple possible mechanisms that could drive the desired outcome.
In a case study about Airbnb, one of the company’s objectives was to increase the proportion of listings that offer “Instant Book”, which allows guests to book a place without host approval. In order to achieve that outcome, the company had a few different mechanisms at its disposal. New listings could be opted into instant booking by default unless the host deliberately opted out, nudging hosts who may not have a strong preference one way or the other. Alternatively, the search ranking algorithm could be adjusted to show listings enrolled in Instant Book higher in the results, thus incentivizing hosts to enroll for better publicity. A third option would be to offer a financial incentive, such as waiving the host-side fees for the first 10 bookings made after enrolling in Instant Book. Notice how each of these approaches uses a different mechanism in order to drive the same outcome.
In another case study at Alibaba, the company wanted to run an experiment to see if offering users a discount on abandoned items in their shopping cart could lead to an increase in the overall dollar volume of transactions on the platform. That outcome could potentially be driven by several different mechanisms, which are worth reviewing. To determine the size of the discount, they could allow sellers to set their own discounts without any guidance, or alternatively, they could use a machine learning model to suggest to the seller the optimal amount of discount to offer. To inform the buyer of the discount, they could show the discount only when the user logs in, or alternatively, they could pro-actively send the user a notification about the discount. To create a sense of urgency, they could also play with how long the discount might be available. It could be available for 24 hours, or it could be available for a week. This case study again illustrates how companies often have a wide variety of variables available to drive the desired outcomes.
In both of the two case studies that I just reviewed, it’s noteworthy that these companies are zooming out to look at all of the potential levers they could pull to achieve their desired outcome. They’re not starting from an already-created design and just experimenting with the appearance of the button that they want users to click; they’re looking at the bigger picture.
Iterative Experimentation
The authors also emphasize that real learning usually requires a series of related experiments that build on one another, rather than a single experiment in isolation. Let’s quickly review an example to illustrate. Suppose that a company wants to introduce a new product feature that’s designed to drive a target behavior. The first step of an iterative experimentation process might be to run an ad campaign in an attempt to gauge market demand before investing the resources to develop the feature. A second experiment could then try to evaluate different ways of structuring the user experience to see which general approach would be most effective at driving the target behavior. Subsequent experiments could then focus on fine-tuning the best-performing approach, tweaking small details to get the optimal outcome. In this way, it’s often useful to structure a series of experiments so that the outcome of prior experiments are used to frame hypotheses that are tested in later experiments.
Tracking Long-Term Outcomes As Well As Short-Term Outcomes
In addition to looking at short-term outcomes, such as the user’s immediate behavior in response to what was shown to them, it’s also worthwhile to think about what long-term outcomes might be worth tracking. Even if you successfully achieve the primary short-term outcome, what are all of the long-term implications? To explore this issue, the book reviews an example where a company experiments with a UI design that conceals some extra fees until the last minute of the check-out process. While this design did lead to an increase in purchases during the experiment, it wasn’t clear how it might impact long-term customer retention, company reputation, and ultimately, profitability. In this situation, it’d likely be beneficial to run this experiment over a long time period in order to track the relevant long-term outcomes and take those into consideration alongside the short-term outcomes. Ultimately, determining the best course of action in this case would require a judgment call that couldn’t be made based on the experimental data alone.
The Role of Framing in Experimentation
When thinking about experimentation, the authors urge us to always consider the concept of framing. Framing is a psychological bias whereby people make different choices when presented with the same set of options depending on how superficial aspects of options those are presented. When running online experiments, it’s useful to consider whether people’s behavior in the app was caused by the underlying substance of what they were presented with, or some superficial aspect of its presentation. For example, suppose that you are running an ad campaign in order to assess market demand for a new product. If customers aren’t clicking on the ads, does that necessarily indicate a lack of interest in the product, or is it possible that the product’s presentation, rather than the product itself, was less than compelling? This distinction is worth keeping in mind, because there are various methods of disentangling this type of confounding variable.
Key Take-Aways
My main take-away from the book is that recent improvements in the cost-effectiveness of online experimentation should be prompting all digital product companies to reconsider how experimentation fits into their toolbox. Adopting an experimentalist approach to product management is a huge cultural shift, because people may be humbled by their inability to accurately make predictions in areas where they pride themselves on their expertise. Adopting experimentation is also a huge financial commitment, because building out an experimentation platform to support all of this is a massive undertaking. The upside here is that after overcoming the initial cultural and financial hurdles, you’re in the clear: there’s diminishing marginal costs for every extra experiment, so it becomes quite reasonable to then run experiments on the majority of changes to a product.
Leave a Reply