1. Why it’s really important to define KPIs
We all know that we need to have a clear goal for testing.
But learned the hard way: always, always write down the business goals and confirm them with the responsible stakeholders.
We’ve been in this scenario too many times:
we discuss goals during the call,
everyone agrees,
we write them down,
we ask the responsible stakeholder to confirm that the whole strategy and implementation will be based on that,
and surprisingly often something changes or gets clarified.
Why? Because when something is only discussed, everyone thinks they understand it the same way. When it is written down and someone needs to take responsibility for it, there is one last real double check.
This easy step might save you from:
rebuilding the test,
doing the analysis again,
discussions after the results because people suddenly remember the goal differently.
2. Before you launch, decide what will make the result trustworthy
Before starting, make sure you decide a few things.
Is the group allocation big enough to give you statistically correct numbers?
How long should the test run before you make decisions based on the data?
What exactly are you going to analyse?
The time window is especially important.
It’s tempting to look at the first results and make assumptions right away.
But sometimes statistically significant data is simply not enough.
We’ve seen it many times with personalised actions on websites. You might have statistically significant results after 2–3 days. But we still wait around a month before making final conclusions. Why? Because the sales context changes, Friday is different from Monday, a normal week is different from the period before a big holiday etc.
So we define the minimum test period upfront - only then it works. Otherwise it’s very easy to stop the test when the numbers happen to confirm what you wanted to see.
3. Don’t blindly trust the group allocation. And yes, implementation details matter
This is one of the hardest parts: make sure you have proper technology for group allocation. The problem is that the test itself might look fine, while the group allocation is not correct. And if this is not included in the analysis, you get wrong results.
Some examples:
If you want to test technology A vs technology B - which technology will split the traffic?
In many cases it’s better to have independent technology C doing this. Otherwise A or B might influence the results simply because its own tracker or algorithm loads differently than the external one.
If you want to test option A vs option B - should the user be assigned permanently to one group or should the group be chosen every time?
It depends on the test and the goal. This topic can easily become a separate post. If you’re not sure how to set it up, this is exactly the kind of thing we help with.
And now the worst part: the technology that creates the groups has to split them correctly.
You can’t have all top-performing users in one group and weaker users in another. Sounds obvious, but we’ve seen this happen.
I remember one test where the version we were convinced should win was losing very badly. Yes, assumptions can be wrong, but based on our experience, this result looked so strange that we decided to investigate before accepting it.
We applied RFM analysis to both groups and the proportions between RFM segments were very different. So the problem wasn’t necessarily the tested version.The problem was that the two groups were different from the beginning.
If you test something on the website, also check cookies and implementation.
Especially when you compare two technologies. Different implementation can influence who sees what and how the technology behaves. So what looks like a difference in performance might actually be a difference in implementation.
4. Copy and creation: test one element
Sounds easy :), but it’s also very easy to forget when you’re in a rush.
The usual thinking is: “We don’t have time and budget. Let’s change a few things at once. We’ll still see something.” The problem is: if you change a few things and the result is better, what exactly worked?You don’t know. And this is the whole point of testing: you want to know what should be implemented next.
So whenever possible, test one element.
5. Dashboard should be ready from day 1
There are lots of things that might be wrong:
events might be sent incorrectly,
some pages might not be included,
the group split might be wrong,
analytics might be wrong,
you might realise you need additional parameters.
Recently we had a case where everything was tested on stage, then on production, and everything looked fine. Then some hot fixes were done on production and the settings changed, as a result events started sending incorrectly.
This is why the dashboard should be ready from day 1. Not after the test. Not when you finally want to analyse results. From day 1.
And when we see something suspicious, we don’t immediately stop the test - first we verify the analytics, then the technology, then we decide what to do next.
6. If you can answer the same question with a simpler test, do the simpler test
Sometimes it’s better to test something small and fast than create a whole philosophy around it and end up never doing anything.
A simple example:
We created a product recommendation box on a website. The client wanted one type of product filtering. Based on our experience, we believed the AI model should filter products differently. We could discuss who was right. Instead, we launched a test:
client’s filtering,
our filtering.
And we checked the results. No discussion about who has the better opinion, instead we had data and could move forward.
One more thing: not every test needs to prove your theory. And it’s completely fine if your assumptions are wrong. Sometimes teams start thinking: “We’re just burning budget because another hypothesis didn’t work.” But that’s not really how testing works.
One successful test - meaning one where you find something valuable - can generate enough money to cover many other tests. And even when the hypothesis is wrong, you still learn something about your customers. Sometimes the most valuable insight is simply: “Ok, so the customer doesn’t react the way I thought.” And that matters, because next time you’re not basing the decision only on “I think this should work.” - you know more. And this is the point of testing: not to prove that you were right, but to know what to do next.
Have fun testing! 😃
1. Why it’s really important to define KPIs
We all know that we need to have a clear goal for testing.
But learned the hard way: always, always write down the business goals and confirm them with the responsible stakeholders.
We’ve been in this scenario too many times:
we discuss goals during the call,
everyone agrees,
we write them down,
we ask the responsible stakeholder to confirm that the whole strategy and implementation will be based on that,
and surprisingly often something changes or gets clarified.
Why? Because when something is only discussed, everyone thinks they understand it the same way. When it is written down and someone needs to take responsibility for it, there is one last real double check.
This easy step might save you from:
rebuilding the test,
doing the analysis again,
discussions after the results because people suddenly remember the goal differently.
2. Before you launch, decide what will make the result trustworthy
Before starting, make sure you decide a few things.
Is the group allocation big enough to give you statistically correct numbers?
How long should the test run before you make decisions based on the data?
What exactly are you going to analyse?
The time window is especially important.
It’s tempting to look at the first results and make assumptions right away.
But sometimes statistically significant data is simply not enough.
We’ve seen it many times with personalised actions on websites. You might have statistically significant results after 2–3 days. But we still wait around a month before making final conclusions. Why? Because the sales context changes, Friday is different from Monday, a normal week is different from the period before a big holiday etc.
So we define the minimum test period upfront - only then it works. Otherwise it’s very easy to stop the test when the numbers happen to confirm what you wanted to see.
3. Don’t blindly trust the group allocation. And yes, implementation details matter
This is one of the hardest parts: make sure you have proper technology for group allocation. The problem is that the test itself might look fine, while the group allocation is not correct. And if this is not included in the analysis, you get wrong results.
Some examples:
If you want to test technology A vs technology B - which technology will split the traffic?
In many cases it’s better to have independent technology C doing this. Otherwise A or B might influence the results simply because its own tracker or algorithm loads differently than the external one.
If you want to test option A vs option B - should the user be assigned permanently to one group or should the group be chosen every time?
It depends on the test and the goal. This topic can easily become a separate post. If you’re not sure how to set it up, this is exactly the kind of thing we help with.
And now the worst part: the technology that creates the groups has to split them correctly.
You can’t have all top-performing users in one group and weaker users in another. Sounds obvious, but we’ve seen this happen.
I remember one test where the version we were convinced should win was losing very badly. Yes, assumptions can be wrong, but based on our experience, this result looked so strange that we decided to investigate before accepting it.
We applied RFM analysis to both groups and the proportions between RFM segments were very different. So the problem wasn’t necessarily the tested version.The problem was that the two groups were different from the beginning.
If you test something on the website, also check cookies and implementation.
Especially when you compare two technologies. Different implementation can influence who sees what and how the technology behaves. So what looks like a difference in performance might actually be a difference in implementation.
4. Copy and creation: test one element
Sounds easy :), but it’s also very easy to forget when you’re in a rush.
The usual thinking is: “We don’t have time and budget. Let’s change a few things at once. We’ll still see something.” The problem is: if you change a few things and the result is better, what exactly worked?You don’t know. And this is the whole point of testing: you want to know what should be implemented next.
So whenever possible, test one element.
5. Dashboard should be ready from day 1
There are lots of things that might be wrong:
events might be sent incorrectly,
some pages might not be included,
the group split might be wrong,
analytics might be wrong,
you might realise you need additional parameters.
Recently we had a case where everything was tested on stage, then on production, and everything looked fine. Then some hot fixes were done on production and the settings changed, as a result events started sending incorrectly.
This is why the dashboard should be ready from day 1. Not after the test. Not when you finally want to analyse results. From day 1.
And when we see something suspicious, we don’t immediately stop the test - first we verify the analytics, then the technology, then we decide what to do next.
6. If you can answer the same question with a simpler test, do the simpler test
Sometimes it’s better to test something small and fast than create a whole philosophy around it and end up never doing anything.
A simple example:
We created a product recommendation box on a website. The client wanted one type of product filtering. Based on our experience, we believed the AI model should filter products differently. We could discuss who was right. Instead, we launched a test:
client’s filtering,
our filtering.
And we checked the results. No discussion about who has the better opinion, instead we had data and could move forward.
One more thing: not every test needs to prove your theory. And it’s completely fine if your assumptions are wrong. Sometimes teams start thinking: “We’re just burning budget because another hypothesis didn’t work.” But that’s not really how testing works.
One successful test - meaning one where you find something valuable - can generate enough money to cover many other tests. And even when the hypothesis is wrong, you still learn something about your customers. Sometimes the most valuable insight is simply: “Ok, so the customer doesn’t react the way I thought.” And that matters, because next time you’re not basing the decision only on “I think this should work.” - you know more. And this is the point of testing: not to prove that you were right, but to know what to do next.
Have fun testing! 😃


