optimizacion-conversion

ICE Score: prioritize your growth experiments

Adrià Vidal7 min read
ICE ScoreprioritizationexperimentationgrowthCRO

Every CRO or growth team has the same problem: more ideas than time and resources to implement them. The question is not what to experiment on, but what to experiment on first. This is where prioritization frameworks come in, and the ICE Score is one of the most popular for its simplicity and speed of application.

What is the ICE Score?

The ICE Score is a prioritization framework created by Sean Ellis, the pioneer of growth hacking, to rank growth experiments across three dimensions:

  • I – Impact: how much can this initiative move the needle if it works?
  • C – Confidence: how much evidence do we have that it will work?
  • E – Ease: how easy is it to implement?

The ICE score is calculated by multiplying (or averaging, depending on the variant used) the three values on a scale of 1 to 10:

ICE Score = Impact × Confidence × Ease / 3

Or in the more widely used version:

ICE Score = (Impact + Confidence + Ease) / 3

The multiplicative version amplifies differences and harshly penalizes a low score in any one dimension. The summative version (average) is gentler and is used when the team has little experience with the framework.

How to score each dimension

Impact: the potential impact on the target metric

Impact measures how much the metric you care most about can improve if the experiment succeeds. To score it well, you need to be specific about which metric that is. In CRO it is usually conversion rate, revenue per session, or activation rate.

Rough guide:

  • 9-10: could double the metric or have massive business impact
  • 7-8: significant improvement, between 20% and 50% on the metric
  • 4-6: moderate improvement, between 5% and 20%
  • 1-3: marginal improvement, less than 5%

The most frequent mistake is overestimating impact. Most well-executed experiments generate improvements between 5% and 15% on the target metric. Scoring every impact a 10 destroys the usefulness of the framework.

Confidence: the evidence behind the hypothesis

Confidence measures how much evidence you have that the experiment will produce the expected result. Sources of evidence include:

  • Qualitative data: user interviews, heatmaps, session recordings
  • Quantitative data: funnel analysis, behavioral data
  • Industry benchmarks or results from your own previous experiments
  • Well-documented behavioral psychology principles

Rough guide:

  • 9-10: solid proprietary data + industry evidence + prior positive test
  • 7-8: proprietary behavioral data clearly supports the hypothesis
  • 4-6: there are indications but no conclusive data
  • 1-3: it is an intuition with no supporting evidence

Ease: the implementation cost

Ease measures how much effort is required to implement the experiment. This includes development time, technical complexity, design resources, and coordination with other teams.

Rough guide:

  • 9-10: text or image change, implementable in hours
  • 7-8: minor design change, 1-2 days of work
  • 4-6: requires development or changes to multiple components, 1-2 weeks
  • 1-3: requires structural changes or integration of new systems

Practical ICE Score example in a CRO backlog

Suppose we are the CRO team of a B2B SaaS and we have these five hypotheses in the backlog:

HypothesisImpactConfidenceEaseICE Score
Add social proof to the pricing page7898.0
Redesign the registration form (fewer fields)8767.0
New personalized onboarding flow9636.0
CTA test on the homepage hero5596.3
Integrate live chat in checkout7545.3

With this analysis, the first experiment to launch would be adding social proof to the pricing page: high potential impact, good evidence (studies on social proof on pricing pages are consistent), and very easy to implement. The new onboarding flow has the greatest potential impact but is complex to implement and we have less certainty about the exact outcome.

ICE vs. RICE: when to use each

The RICE framework, created by the Intercom team, adds a fourth dimension to ICE: Reach.

RICE Score = (Reach × Impact × Confidence) / Effort

RICE is more precise when experiments have very different reach. For example, if one experiment affects 100% of users and another affects only 5% (such as a change to a rarely-used feature flow), ignoring reach can lead to poor prioritization.

ICE is better when reach is similar across experiments or when the team needs speed in prioritization. RICE requires estimating the number of affected users, which adds time to the process.

ICE vs. PIE: another perspective

The PIE framework by Chris Goward (Widerfunnel) measures:

  • P – Potential: how much can it improve with CRO
  • I – Importance: how important is the page to the business (traffic + conversion value)
  • E – Ease: ease of implementation

The main difference from ICE is that PIE focuses Potential on the margin of improvement from the current situation, not on absolute impact. A page that already converts well has less Potential than one that converts poorly, even if both have high traffic.

The PXL framework for CRO experiments

PXL, created by the CXL team, is the most rigorous framework for advanced CRO teams. It evaluates:

  • Whether there is quantitative data supporting the change
  • Whether there is qualitative data supporting the change
  • Whether the hypothesis is based on UX/psychology principles
  • The potential impact on the conversion metric
  • The reach (how many users see this element)
  • The implementation cost

PXL scores binarily (yes/no) on most dimensions, which reduces subjectivity. It is slower to apply but more objective.

How to implement ICE Score in your team

The ICE Score is a tool, not an absolute truth. For it to work well:

Score as a team, not in isolation. The ICE Score is more reliable when scored independently by multiple people and then the discrepancies are debated. Score differences are themselves a symptom of disagreements about evidence or business priorities.

Review and update the backlog regularly. The priority of an experiment changes with data. An experiment that had low confidence can move up the list if new data emerges supporting the hypothesis.

Do not use ICE for experiments with very different reach. If you have experiments affecting 1% and 100% of users respectively, use RICE to avoid comparing apples and oranges.

Document results and feed back into confidence. If an experiment with confidence 8 fails, that is information: something was wrong in your evidence. If one with confidence 4 succeeds, that is also information. Those learnings should improve the quality of your future estimates.

ICE Score as part of an experimentation culture

The ultimate goal of the ICE Score is not to optimize a single experiment: it is to build a sustainable experimentation rhythm. Teams that experiment more learn faster, and teams that learn faster win.

Rigorous prioritization avoids two destructive patterns: the "experiment of the month" where a big test is launched without clear criteria, and the "experiment of the week" where everything is launched without sufficient evidence and results are noisy.

If you want to implement an experimentation program with rigorous prioritization in your business, our CRO team has the methodology and experience to support you.

Explore our CRO program

Have a backlog of website changes and not sure where to start? Begin with a free diagnosis:

Audit your site for free with Scan&Boost

Adrià Vidal is a conversion optimization and growth specialist at Boost. He works with digital companies to improve their activation, retention, and conversion rates through data-driven experimentation. Connect on LinkedIn

Adrià Vidal

Adrià Vidal

CEO & Founder

Founder of Boost. Specialist in digital analytics, CRO, and artificial intelligence applied to digital business optimization.

Related articles

ICE Score: prioritize your growth experiments | Boost