Why prioritisation matters more than the tests themselves
Testing capacity is limited — you can only run so many experiments at once and each takes time to reach significance. If you spend that capacity on button colours while a broken headline or a weak offer goes untested, you're optimising the wrong things.
A simple scoring framework forces you to compare ideas on the same terms and start with the ones most likely to move the needle.
What ICE scoring is
ICE ranks each idea by the average of three 1–10 scores:
- •Impact: if it works, how much could it move your key metric?
- •Confidence: how sure are you it'll work (based on data, past tests, or a strong rationale)?
- •Ease: how quick and cheap is it to build and run?
Average the three, rank highest first, and start at the top. High-impact, high-confidence, easy ideas are your first tests; low-impact, low-confidence, hard ones drop to the bottom.
How to score honestly
The framework is only as good as your honesty. Don't inflate Confidence for an idea you're emotionally attached to, and don't score Impact high just because a change is visible. Where you can, ground Confidence in something real — a prior test, an analytics signal, a known best practice — rather than a hunch.
It's a prioritisation heuristic, not a prediction. It tells you what to test first; the test itself tells you what actually works.
Start every test with a hypothesis
A test without a hypothesis is just a guess you can't learn from. Frame each as: 'We believe [change] will improve [metric] because [reason]; we'll measure it via [metric].' The hypothesis keeps you honest about what you expected and makes the result meaningful whichever way it goes.
Run one clear test at a time
Once you've prioritised, protect the integrity of each test: change one meaningful thing, run it long enough to reach a real sample size, and judge it against your hypothesis. A ranked backlog plus disciplined execution beats a pile of half-run experiments.