CRO Consultant: Role and When to Hire
What a CRO consultant does, how it differs from an agency, required skills, working process and when it makes sense to hire one.

Imagine you launch an A/B test, wait the necessary weeks for a solid sample, and the result arrives: no winner. The variant doesn't outperform the control. You close the test, note the learning, and move on to the next one. That's how it's supposed to work.
But in three different tests we've conducted at Boost —with two clients from completely different sectors— that "no winner" actually hid a very clear winner. It just required looking at the results by user segment, not in aggregate.
This post explains why this happens, with real data from those three cases, and what you need to change in how you analyze your own A/B tests to avoid mistakenly discarding a variant that actually works.
When you launch an A/B test, the norm is to look at a single number: the conversion of the control versus that of the variant, across all users. If it surpasses the confidence threshold, it won. If not, it didn't win. End of story.
The problem is that this aggregate number mixes two audiences that behave very differently: the user who arrives at your website for the first time and the one who already knows you. If the effect is very positive in one and neutral or negative in the other, the average of both can appear as a tie that doesn't actually exist.
Our hypothesis definition methodology already emphasizes isolating the variable being tested. But there's a second layer that is often overlooked: also isolating who you are testing that variable on.
We're not talking about a theory. We're talking about three real A/B tests, across two Boost client accounts in different sectors, with the same pattern repeating almost identically.
On the landing page for a specific offer from The Excellence Collection, we tested visually highlighting the discount percentage at the top of the page, primarily designed for mobile.
In aggregate: +4.5%, with an 84% probability of winning. It doesn't reach the 90% required by the account's decision threshold. Test closed with no winner.
But when segmenting by user type, the interpretation changes completely: returning users increased by +9.9% with 96.3% confidence —comfortably exceeding the threshold—. New users, on the other hand, remained completely flat (0.4%).
The aggregate "tie" was actually a returning user responding strongly, diluted by a new user for whom the variant meant nothing.
On the same landing page, we tested adding Google ratings and featured testimonials.
Aggregate result: +3.05%, 74.4% confidence. Also doesn't reach the threshold.
Again, when segmenting: returning users increased by +10.52%, while new users dropped by 7.99%.
Two effects of opposite signs, almost completely ignoring each other in the overall number.
At Viuty, we tested adding social proof, FAQs, and a discount module to a lead generation advertorial.
Here, the aggregate result was clear: negative, 10.06% in purchase.
But the segment breakdown once again tells the real story: returning users almost tied in conversion and increased revenue per visitor (+5.88%).
It was new users who plummeted: 28% in conversion and revenue.
Since the advertorial relies on cold traffic, that 28% loss overshadowed any gains from returning users.
The underlying mechanism is always the same: trust and social proof function as reinforcement, not as an introduction.
A returning user already has a formed opinion about your brand; when you show them reviews, guarantees, or a well-highlighted discount, you are confirming something they already suspected, and that confirmation reduces friction in the final stretch before converting.
A new user doesn't have that prior context. They're not evaluating whether they trust you "a little more"; they're deciding whether they trust you. Period.
And at that moment, elements that are a boost for a returning user can be noise for a new one, slow down their reading, or even generate the opposite question to the one you were trying to answer.
If you're going to test a trust or social proof lever, define from the test's inception that you will read the results separately for new and returning users.
It's not an optional extra for analysis: it's the information you need to avoid discarding a real winner or, worse, to avoid implementing a variant that actually harms your highest-volume segment.
Practical rule: If the aggregate is close to the confidence threshold (e.g., between 70% and 90%), review the breakdown before closing the test as a failure.
If one segment clearly crosses the threshold and the other is neutral, you don't have a test with no winner —you have a winner that only applies to a portion of your traffic.
This is perfectly actionable information: you can launch a second version of the test targeted only at that segment, with its own control group, to confirm the effect in isolation.
You don't need the traffic volume of The Excellence Collection or Viuty to apply this. Any A/B test with a minimum sample of new and returning users can be segmented in the same way.
The important thing is the habit: before closing a test as "inconclusive," ask yourself if the problem is that there's no effect, or if there are two opposing effects hidden within the same number.
If you want to know if your own website has these types of signals buried in the aggregate, our free CRO audit is a good starting point to identify where it's worth looking closer.
What a CRO consultant does, how it differs from an agency, required skills, working process and when it makes sense to hire one.
What dark patterns are, most common types, real examples, European legislation and ethical alternatives that improve conversion without deceiving users.
Discover what demand generation is, how it differs from lead gen and how to apply effective strategies to generate real demand for your business.