Summary
Peer group benchmarks look like a leaderboard. They are closer to a weather forecast. Every number you see has been deliberately blurred, the amount of blurring depends on how small your peer group is, and the presentation is designed so that you cannot tell where inside a band you actually sit.
Apple documents more of this than most people realise, including the use of differential privacy. Two granted patents describe the parts it does not, and one of them names a metric that has never appeared in the product.
Methods used
| Stage | Method | Where it is stated |
|---|---|---|
| Group formation | App type by category and subcategory, business model, and download volume tier | Documented |
| Minimum size | Each peer group must contain at least a certain number of apps before it is released | Documented |
| Privacy | Differential privacy, with noise added to each data point | Documented |
| Accuracy control | The injected error is checked against an accuracy requirement and the two are balanced iteratively | Patent |
| Population | Only users who agreed to share app analytics with developers. Apple’s own apps are included in peer groups | Documented |
| Display | Your value against the peer group’s 25th, 50th and 75th percentiles | Documented |
| Display intent | Showing the value and the percentile range while deliberately obscuring the position inside that range | Patent |
| Cadence | Weekly | Documented |
| Metrics considered | Conversion, crash, retention, monetization, usage, and discoverability | Patent |
What was not previously known
Already documented by Apple, and worth stating because several of these get presented as discoveries: that differential privacy is used and noise is added to every data point, that each peer group must reach a minimum number of apps before it is published, that groups are formed on category, business model and download volume, that you cannot change your group but can view others, that the display is a 25th, 50th and 75th percentile distribution, that data refreshes weekly, that only analytics-sharing users are included, and that Apple’s own apps sit inside your peer group where the customer journey is comparable.
Apple also states the limitation plainly: benchmarks are “designed to provide directional insight, not an exact ranking.”
Not documented, and found only in the patents:
- Discoverability is treated as a peer comparable metric. The patent lists it alongside conversion, crash rate, retention, usage and monetization as something compared across a peer group. It is not among the eight metrics the product displays. That implies Apple computes a discoverability figure per app and considers it comparable between apps, which is closer to a public ASO score than anything Apple has described anywhere else.
- The noise level is negotiated, not fixed. Apple says it adds “a certain amount of noise”. The patent describes a loop: inject measurement error, confirm the result is still accurate enough to be useful, and balance the privacy threshold against the accuracy requirement iteratively. Different groups end up with different amounts of blurring.
- Hiding your position inside the band is the claimed invention. The percentile presentation is not a readability choice. The independent claim of the display patent requires showing the metric value and the percentile range while obscuring precisely where within that range the metric falls, so that competing developers cannot reconstruct each other’s standing by watching boundaries move.
- Usage appears as a comparable metric too, separately from retention, and it is not in the displayed set either.
How a benchmark number is formed
The eight metrics you are shown
| Metric | Definition |
|---|---|
| Conversion rate | Total downloads and pre-orders divided by unique device impressions |
| Day 1, 7 and 28 retention | Share of active devices that installed in the selected week and opened the app 1, 7 or 28 days later |
| Crash rate | Crashes divided by total sessions |
| Proceeds per paying user | Total proceeds including in-app purchases divided by paying users. Not shown for free apps |
| Day 35 download to paid conversion | Share of first time downloads or redownloads followed by an in-app purchase within 35 days |
| Day 35 proceeds per download | Average proceeds generated within 35 days of download or redownload |
Note the conversion rate denominator: unique device impressions. Same convention as product page testing, and the same reason your own analytics will not match.
Your peer group is three decisions you did not make
Category and subcategory, business model, download volume tier. You chose the first, you chose the second, and the third moves on its own.
That third dimension is the one that catches people out. Grow enough and you cross into a higher download tier, where the peer group is tougher. A benchmark that gets worse after a strong quarter may not be a decline at all. You may simply be standing in a different room.
You cannot change your assignment, but you can look at other groups: subcategories, secondary categories, or all categories sharing your business model. And if your assigned group is too small to produce a benchmark at all, Apple tells you to pick a different one.
Then the numbers get blurred on purpose
Apple’s own wording: “we use a technique called differential privacy, which is the gold standard for ensuring that individual values within a group remain private. Every week, we ensure that each peer group has at least a certain number of apps before it’s released, and we add a certain amount of noise to each data point to provide an extra layer of protection.”
The patent adds the part Apple leaves out. The system does not simply add a fixed quantity of noise. It injects measurement error, then checks whether the result is still accurate enough to support a business decision, and iterates between the privacy threshold and the accuracy requirement.
This has a consequence that matters every time you read a benchmark. Precision is a function of group size. A large, generic peer group can be reported tightly. A small, specific one needs more noise to protect its members, so its numbers are looser. The more precisely your niche is defined, the less you can trust a small gap.
And then the display hides the rest
You see your value against the 25th, 50th and 75th percentiles. What you do not see, by design, is where inside a band you sit.
The claim in the display patent is explicit that this concealment is the point. The system identifies which percentile range your metric falls into, displays the value and the range, and withholds the position within it.
Three things follow, and they are the most practically useful part of this article:
- Movement inside a band is invisible to you. You can improve meaningfully and see nothing change. You can decline meaningfully and see nothing change.
- A band crossing is therefore a large event, not a small one. Treat it as real. Treat a stable band as an absence of information rather than as evidence of stability.
- Your own absolute value is exact and is shown. That is the number to track week over week. The percentile is context.
How it works
Step 1: the population is assembled
Benchmarks draw on App Store, app usage and transaction data from devices running at least iOS 8, macOS 11 or tvOS 9, and usage data comes only from users who agreed to share their app analytics with developers.
That last clause is a sampling caveat nobody discusses. Every retention and usage benchmark describes the behavior of analytics-sharing users, not of all users, for your app and for everyone you are compared against.
Step 2: the peer group is formed and checked for size
Apps are grouped on shared traits, and the group must reach a minimum membership before any benchmark is released for it. Apple’s own apps are included where the customer journey is directly comparable.
Step 3: noise is added and accuracy is verified
A privacy mechanism injects measurement error into the group’s aggregate values. The system then confirms the result remains accurate enough to be useful, balancing the two requirements against each other.
Step 4: percentiles are computed and presented
Percentile rankings are generated from the peer group’s privatized metrics. Your app’s own metric is placed into one of the resulting ranges, and both the value and the range are displayed while your position inside it is withheld.
Step 5: it refreshes weekly
New week, new group composition, new noise. Two consecutive weeks are not two measurements of the same thing with the same error.
What this means for how you work
- Track your own absolute values, use percentiles for orientation. Apple says it directly: directional insight, not an exact ranking.
- Check which peer group you are in before interpreting a change. A worsening benchmark after growth may be a tier change.
- Look at several groups, not one. Your subcategory, your secondary category and all categories on your business model will disagree, and the disagreement is informative.
- Distrust small gaps in small groups. That is where the noise is largest.
- Remember Apple’s apps are in the room. If your category has a strong Apple app in it, your percentile includes it.
- Treat retention as an ASO metric. It is benchmarked against your peers, and a separate Apple patent describes an installation retention factor inside the profile used for recommendations.
- Watch for discoverability to surface. It is named in the patent as peer comparable and is not in the product. If it ever appears, it will be the closest thing to an official ASO score Apple has published.
Sources and caveats
- Apple, Peer group benchmarks, App Store Connect Analytics help. The metric definitions, group formation, differential privacy statement, percentile display, weekly cadence and data sources.
- US12430660B2, App store peer group benchmarking with differential privacy. Group formation, the accuracy versus privacy loop, and the metric list including discoverability. Granted September 2025, active.
- US12561221B2, Techniques for managing performance metrics associated with software applications. The claim covering the deliberate concealment of your position within a percentile range. Granted February 2026, active.
- Related on this site: US20240086412A1, for the installation retention factor and the other app profile signals.
Caveats. A patent describes what a company designed and claimed, which is strong evidence and not documentation. Where the documentation and a patent disagree, the documentation describes the shipped product and the patent describes the design. Discoverability being named in the patent is not evidence that any discoverability figure is shown to developers today, because it is not.
Bir yanıt yazın