ASOpatents.com compiles a list of patents that are likely used to shape the algorithms of the Apple App Store, Google Play Store, and other major platforms. While it's uncertain whether these patents are actually implemented in the algorithms, the site provides insights into potential clues about search results, recommended apps, and other data points.

US11023905B2: How the App Store Spots a Breakout Before It Breaks Out

Patent numberUS11023905B2
Official titleContent recommendation system
AssigneeApple Inc.
InventorBrian D. Choi
Priority dateJuly 25, 2018, via provisional application 62/703,335
FiledJanuary 25, 2019
GrantedJune 1, 2021
StatusActive. Anticipated expiry April 21, 2039
FamilyUS20200034857A1 (published application), US20210406925A1 (continuation, published December 30, 2021 under the title “Algorithm for identification of trending content”)
Sourcepatents.google.com/patent/US11023905B2

What This Patent Covers and Why It Matters for ASO

Feature it describes: the three scores a content distribution system uses to decide which titles are rising, which are about to rise, and which to put in front of a specific user.

A note on scope before anything else, because this site’s standard is to state what the document actually says. This patent is written for a “content distribution system” in general. It defines a digital asset as “an application (e.g., an ‘app’), a video game, multimedia files (e.g., music or videos), and the like”, and it names “e-commerce site, app store, console operating system” as the kinds of system it covers. Applications and games are the lead examples throughout, and the granted claim talks about assets “not installed on a client device” and about a “cumulative number of downloads”. That is app language. It is not an App Store only patent, and it should not be cited as one, but apps are plainly the first case it was written for.

What makes it worth reading is that most public discussion of App Store charts assumes the answer is download velocity, and stops there. This document gives three separate scores, and only one of them is anything like velocity.

The trend score is growth divided by noise. It fits a line to historical download data across several overlapping time windows, specifically “7-day, 14-day, and 30-day windows”, and for each window it takes the slope of that line, multiplies it by the coefficient of determination, and divides by the standard deviation. Those per window terms are then combined into a weighted sum.

Read that as a formula and the practical consequence falls out. The slope is how fast you are growing. The coefficient of determination is how well a straight line actually describes your data, so a clean sustained climb scores higher than a jagged one with the same average slope. Dividing by standard deviation penalizes volatility again. A single day spike from a newsletter blast or a paid burst has a high slope in the seven day window, a poor line fit and a large standard deviation. Steady week over week growth beats it. This is a formula that rewards consistency and is built to resist being gamed by a one day push.

The breakout score is the one nobody talks about. The system first establishes a breakout date for assets that already broke out, defined as “a time at which the digital asset began to gain in popularity and average daily downloads of the digital asset increase over time”. It then looks backwards and identifies trendsetters: users who repeatedly downloaded those assets before the breakout date happened. A trendsetter is “any user that has downloaded/installed at least a threshold number of digital assets prior to a corresponding breakout date for that particular digital asset”.

Then it inverts the process. It takes assets whose cumulative downloads are still below a threshold, deliberately filtering out anything already popular, and it scores each one by counting how many trendsetters have downloaded it. An obscure app that a lot of early adopters have quietly installed gets a high breakout score.

For a small app this is the most actionable idea in the document. There is a described path to surfacing that does not require volume at all. It requires the right users. Who installs you counts for more than how many, and it counts specifically while you are still below the popularity threshold. That is a distribution argument for targeting early adopter communities first rather than chasing raw install numbers.

The recommendation score is collaborative filtering on two vectors, not one. It compares a vector of installation data and a separate vector of usage data against the target user’s equivalents, and combines them as a weighted sum of the two dot products. Installation and usage are kept apart and weighted separately, which means an app that people install and then abandon is treated differently from one people install and keep opening. Retention is not a bonus layered on top here. It is one of the two inputs.

Patent Summary

A server in a content distribution system computes three scores over the catalog of digital assets and ranks assets by them in order to build a recommendation surface.

Trend score. For each asset, historical download counts are analyzed across two or more overlapping time windows, given as seven day, fourteen day and thirty day. A line is fitted to the data in each window. Each window contributes a term equal to the slope of its line multiplied by the coefficient of determination of that fit and divided by the standard deviation. The trend score is a weighted sum of those terms.

Breakout score. The system identifies assets with an established breakout date, meaning the point at which average daily downloads began to climb. It identifies trendsetters associated with a category of assets, being users who downloaded at least a threshold number of assets before those assets broke out. It generates the list of assets downloaded by those trendsetters, filters out assets whose cumulative downloads exceed a threshold, and scores the remainder by counting how many trendsetters downloaded each one.

Recommendation score. Similar users are identified by a weighted sum of two dot products: one between a vector of installation data and the target user’s installation vector, and one between a vector of usage data and the target user’s usage vector. The score for an asset reflects how many of those similar users installed it. It is computed only over assets not already installed on the target user’s device.

The granted claim 1 covers the combination: calculating a trend score for each of a plurality of digital assets managed by the system, calculating a recommendation score for a subset not installed on the target user’s client device, calculating a breakout score for a subset whose cumulative downloads fall below a threshold value, ranking the assets based on the scores, and generating a visual representation of assets to recommend. Unusually for this corpus, the specific breakout mechanism survived into the claim rather than being stripped out during examination.

Bir yanıt yazın

E-posta adresiniz yayınlanmayacak. Gerekli alanlar * ile işaretlenmişlerdir