the best product managers don't just rely on product intuition but they think like scientists and run good experiments hence in this video I'm going to uncover the five phases it takes to run successful experiments so your team can ship great products we're also going to go through some examples to draw it home for you and I'm going to talk about some team members and tools that you'll need to ensure a successful and comprehensive experiment the phase number one is figuring out what are the key goals and metrics that you are trying to optimize and what
might be some risks for example if you're talking about a company slash product like Airbnb they're key thing they're trying to optimize for is a metric like the number of nights booked which is correlated with Revenue which is what they care about some other metrics that they could care about is the booking rate which is basically the total number of listings viewed on the denominator and the number of times that a booking takes place in the numerator and another thing we want to look at closely here is countermetrics this helps ensure that whatever you're trying
to optimize for doesn't cause a ton of unintended behavior that actually goes against your goals for example I could imagine some experiments that optimize for a number of night's book could also come at a cost of increase in cancellation rates or increasing bad guest experiences and hence if you have air experiment that's run where cancellation rates goes up by 50 percent and number of nights book go up by 20 you probably don't want to ship that experiment but if we only looked at one side that the number of nights booked increased you could see how
you can fall very easily for that trap so in this step it's very important to align the team on the same set of goals that you're trying to optimize and the risks so bringing together the entire product team to talk about this even with leadership is going to be key to being very intentional about the experiment that you set up phase number two is what is the hypothesis that you're trying to test to optimize the goal we talked about in the last step so you want the hypothesis to be logical based on what makes comments
sense or examples based on past data that have shown to optimize the metric that you care about here an easy hypothesis I might have is the bigger the button the more likely that users are going to book so in version a which is the control version I'm going to show the normal size button and in version B I'm going to show an even large button by 20 or 30 percent and then I'm going to run the test to see if the larger button is actually clicked on more times and by what percentage another hypothesis I
might have is that people book more when they feel like things are more scarce and hence in version a I'll leave that as a control and version B I'm going to have a test where when people are getting ready to book I'm going to show them the number of people that are also viewing that listing to make them nervous that the listings are going to be sold out very soon phase three I call scenario planning which is a decision tree based on the possible scenarios that could happen after you run your test and here this
is based off of the three different actions that you could take post experiment the first one is ship no ship and retests so let's start with ship so in What scenario would you ship one of the experiments we talked about before well one obvious one is that if we see the results on the metric number of nights booked is trending positive we see that cancellations have not insurmountably increased by a ton and that the other metrics we care about such as booking rate is also trending positive so that one might be an obvious one we
ship well what is one that we might not ship so again one obvious one we may not ship is if we see the number of nights booked actually is going down or we might see those numbers going up but cancellation rate also going up by an even bigger Factor so in that case we would not ship the experiment and lastly in what cases might we consider retesting well if the results aren't statistically significant and we have a way of increasing the statistical significance of this experiment if there is no way to increase the statistical significance
and you thought the first rendition was already the most optimized you might decide to abandon the experiment results if you don't feel the most confident phase number four will talk about experiment design how are you going to actually set up the test you want to make sure to design the test to get the most conclusive and trustworthy results so here are some factors to help increase the chances of those first let's talk about when an A B test even looks like so I alluded to some of this but in an A B test slash experiment
who have your control which is the thing that exists today and then you have your test which is the hypothesis that you're trying to validate or invalidate and you want to make sure that you choose a significant pool of people to test this on so hence the number of samples there should be a minimum number of samples where you feel like it's a large enough population to be generalizable to the rest of the population so for example let's say that number is ten thousand and then you want to put that ten thousand five thousand of
those in the control version eight five thousand of those in the test version B then what you're going to want to do is track their behavior using metrics as they react to the experiment this allows you to see if there's any statistically significant difference between people's reaction to version a versus version B and that's essentially the test it sounds easy but there's a lot of science to make sure the test is designed in a way that when the results come out people can be confident about the results so we already talked about getting a minimum
number of sample sizes from the number of people another thing we want to make sure is that the people in the tests are randomized so for example for the booking.com experiment if I ran the test only on men I'm not going to feel very confident that it's generalizable to the entire population which would include women another factor to keep in mind is that the test should run for a long enough time well what is that long enough time mean well good signifier is when you see some of the metrics start tapering off why because you
need to run experiments for a long enough time to see the after effects for example on the airbnb1 we could see people increasing their number of nights booked but we also need to give it time to see our people canceling at a higher rate which might be one or two weeks after the initial booking another thing you want to keep in mind to create good experiments is to isolate variables so you want to make sure that you're not testing multiple hypotheses in one experiment why because there's compounding factors and hence it might be hard to
know which of the hypotheses caused the effect that you ended up seeing which might mean it might be hard to replicate in the future it might be hard to then tease out which is actually causing the intended effect that you want so who is going to be necessary to help you in this experiment well you'll definitely need someone like a data scientist who can do things like power calculations and help you calculate the number of days will help you calculate the population that you'll need how to create a randomized sample of people and also how
many days this test should run for they'll also help you with the post analysis to review their analytics some tools that I've seen startups use because the companies I was in were so bad that they created their own tools but some external tools I've seen smaller startups to use are things like optimizely or kissmetrics and the last step once you've run the test is to review the experiment results the things you want to look at is first what are the trends of the results things like reviewing the percentage differences between the control versus test and
even looking at the absolute numbers are the metrics trending in the right direction that you wanted to Trend in do the results make logical sense if not think why the results ended up being what they ended up being another thing is to go back to review your scenario plan which is yes the plan you came up in a perfect scenario and it might not have included all the other different effects it could have had so here go back to review it adjust it if need be but it helps create that clear structure to help you
make a more intentional decision and lastly the actual decision here it's more of an art than it is a science and the more your team is intentional about what the team is trying to optimize as in a limited set of metrics and the weight of the metrics that you care about be clear it will be to make the decision I've seen in past teams the teams that struggle with making a decision are the ones that are least clear about what they most care about so the clearer you can be about that up front or even
towards the end the easier it is going to be to decide a ship no ship retest so in the Airbnb example we might see that we're getting a 50 lift on the number of nights booked we're seeing that's affecting the revenue by an increase of also 10 percent we're seeing cancellations go up by five percent and five percent is still a lot so we might ask our team is this something we're willing to accept and in some cases teams will make the decision that no a five percent increase is too high but some other team
might be like oh that's a trade-off I'm willing to take for a 15 increase in number of bookings and we're going to further do things to reduce that five percent cancellation rate so you can see either directions could be a plausible decision based on the team and what they care about so in this video we talked about the five steps it takes to run experiment it all seems pretty straightforward till you actually run one so if you want to see a video on us actually talking through real experiments from companies like Facebook take a look
at these videos that run through experiments and different possible results metrics Etc and I'll see you guys in the next video