← All notes

How to A/B test App Store screenshots without fooling yourself

A screenshot test is easy to launch and surprisingly easy to misread. Change the headline, background, first screen, and acquisition campaign at once, then stop after a good Tuesday, and the dashboard may still hand you a number. It just will not tell you what to do next.

A useful test starts with one decision. Does a result-first story beat a feature-first story? Does a calm editorial direction beat a bright, energetic one? Build the experiment around that question, then keep everything else as stable as practical.

Apple and Google use different tools

StoreToolCurrent limits
Apple App StoreProduct Page OptimizationUp to three treatments; tests run for at most 90 days
Google PlayStore listing experimentsUp to two variants; experiments stop after six months

Apple lets you test app icons, screenshots, and app previews. You choose the share of traffic sent to treatments, and that share is divided evenly between them. Apple estimates duration from your existing impressions, downloads, and the improvement you want to detect. See Apple's Product Page Optimization setup guide.

Google Play can test icons, feature graphics, screenshots, and localized descriptions. You choose install, open, or pre-registration clicks as the target metric. Its calculator estimates the time and traffic needed from your confidence level and minimum detectable effect. Google recommends testing one asset at a time in its store listing experiment guide.

Start with the first screenshot

Your first experiment should usually change the opening screenshot, not rebuild every slot. It carries the first promise and can be visible before someone explores the rest of the gallery. Good hypotheses sound like this:

Change one idea, not necessarily one pixel

"Test one thing" does not mean moving a button four pixels. A complete art direction can change color, composition, and type together if the question is whether that direction wins. Keep the message and screenshot sequence constant so the direction remains the meaningful difference.

If you are testing the message, do the reverse. Hold the visual system steady and change the benefit or feature being shown. The result should teach the next designer something reusable, not merely identify a mysterious winning file.

How much traffic is enough?

There is no honest universal sample size. It depends on your current conversion rate, traffic, number of variants, and the smallest improvement worth acting on. Use this process:

  1. Choose the smallest lift that would justify replacing the current creative. If a tiny movement would not change your decision, do not configure the test around finding it.
  2. Use the store's duration estimator. It knows the traffic the listing actually sees.
  3. Run across complete weekly cycles so one part of the week is not overrepresented.
  4. Wait for the platform's confidence result. Do not stop after an early spike.

Apple starts showing results after five first-time downloads are attributed to the test. That is a display threshold, not proof that five downloads are enough. Apple marks a treatment as performing better or worse at 90% confidence. Its analytics reference explains conversion, lift, and confidence.

A clean test plan

  1. Write the decision the result will unlock.
  2. Name one hypothesis tied to a customer need.
  3. Build the control and treatment together to keep polish equal.
  4. Check every included locale and device size.
  5. Record campaigns, pricing changes, and releases that can alter the audience.
  6. Wait, decide, and save what the team learned.

Storeboard makes the expensive part cheaper: render complete directions side by side from the same real app screens, export the required sizes, then let the store test settle the argument.

Put this into practice

Create screenshot directions

Skip the artboards entirely.

A description, a screenshot, or your URL in; the full store-ready set out. Start free with a draft project.

Build my listing →