A/B testing
that tells you something.
Both stores will split your traffic and tell you which listing converts better. The tooling is free and most indie apps never touch it, usually because the rules differ per store and the traffic maths is unforgiving. Here is what each one actually lets you test, and how to run a test whose result you can trust.
What each store actually lets you test
The two stores took opposite approaches, and the difference decides how you work.
Google Play has Store Listing Experiments, built into the Play Console. You create a variant, choose what it changes, and Play splits live traffic between it and your current listing. It reports installs per variant with a confidence interval and tells you when a result is statistically meaningful. You can run one default listing experiment and several localised ones at the same time.
Apple has Product Page Optimization, inside App Store Connect. You can run up to three treatments against your current page, each varying the icon, the screenshots, or the app preview video. Apple shows a percentage of traffic each treatment gets and reports improvement in conversion rate with its own confidence measure. Tests run for up to 90 days.
The rules that catch people out
- Apple requires a new app version to change screenshots in the base listing, but a Product Page Optimization treatment does not. That makes PPO the only way to try screenshot ideas between releases.
- Play changes are free and instant. You can edit graphics any time without shipping a build, which is why Play is the cheaper place to learn.
- Apple tests the icon too, and a tested icon must already be shipped in a build. You cannot invent one in the console.
- Neither store lets you test the app name or subtitle through these tools. Those are metadata changes and go through review.
Sample size is the thing nobody plans for
A conversion test needs enough traffic to separate a real difference from noise. The smaller the effect you are trying to detect, the more traffic you need, and the relationship is not linear. Detecting a 20% lift takes a fraction of the traffic that detecting a 5% lift does.
For most indie apps the practical consequence is blunt: if your listing gets a few hundred visitors a week, you cannot reliably detect anything smaller than a large swing. Testing a subtly different caption is a waste of a month. Test things that could plausibly move conversion by a fifth or more, which in practice means the first screenshot, the icon, or the whole visual direction.
Both consoles will tell you when a result reaches significance. Believe that indicator over your own reading of the numbers, and do not stop a test early because it looks good on day three. Early leads reverse constantly.
What is worth testing, in order
- The first screenshot. It is the one image nearly every visitor sees, and it carries the whole argument. Changes here move the number more than anything else on the page.
- The icon. Visible in search results before anyone opens your listing, so it affects traffic as well as conversion.
- Screenshot order. Reordering costs nothing to try and often beats redesigning.
- Caption strategy. Benefit-led against feature-led, as a whole set rather than one line.
- The preview video. High effort, and results are genuinely mixed. Some categories convert better without one.
Running a test that means something
Change one thing at a time. If a variant has a new first screenshot and new captions and a different background, a win tells you nothing you can reuse. The point of a test is not to find a better listing once, it is to learn something that informs the next ten decisions.
Let it run to the store's own significance call, or 14 days, whichever is later. Weekly seasonality is real: weekend browsers convert differently from weekday ones, and a test that spans only weekdays measures a skewed audience.
Keep a written record of every test, what changed, and the result including the losses. Losing variants are more instructive than winners, because they tell you where the ceiling is. Six months in, that record is the most valuable asset you have for your listing.
Where testing stops helping
Conversion rate optimisation compounds slowly on a listing that already works. If your conversion is very low across the board, the problem is usually positioning rather than presentation, and no screenshot arrangement fixes an app that the wrong people are finding. Look at which search terms bring traffic before you spend three months testing captions.
Equally, if traffic is tiny, the highest-return work is getting more of it, not optimising the trickle. Testing is a multiplier, and a multiplier on a small number is still a small number.