CoolFace
Datasetpublic

yonilev/GooglePlayStoreApps

πŸ“± Google Play Store 2020 β€” What Made an App Go Viral? A data-driven EDA of 9,436 apps released in 2020, exploring the patterns behind viral success on the Play Store. 🎬 Video Walkthrough Can't see the video? Click here to watch πŸ”¬ Research Question "What made an app released in 2020 go viral?" Defined as reaching 1M+ installs by June 2021 β€” within 6–18 months of launch. 2020 was chosen deliberately: the COVID-19 pandemic drove… See the full description on the dataset page: https://huggingface.co/datasets/yonilev/GooglePlayStoreApps.

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
0likes290downloads
Dataset Card

πŸ“± Google Play Store 2020 β€” What Made an App Go Viral?

A data-driven EDA of 9,436 apps released in 2020, exploring the patterns behind viral success on the Play Store.

🎬 Video Walkthrough

<!-- Replace the src with your actual video link after uploading --> <video src="https://www.loom.com/share/3b2dc51bcf384a339d3d3328766e4705" controls="controls" style="max-width: 720px;"></video>

Can't see the video? [Click here to watch](https://www.loom.com/share/3b2dc51bcf384a339d3d3328766e4705)


πŸ”¬ Research Question

"What made an app released in 2020 go viral?"

Defined as reaching 1M+ installs by June 2021 β€” within 6–18 months of launch.

2020 was chosen deliberately: the COVID-19 pandemic drove unprecedented mobile app adoption, and since the data was collected in June 2021, apps from 2020 had a consistent 6–18 month measurement window.


πŸ“Š Key Visualizations

1. The Winner-Takes-All Market

86% of apps never exceed 10K installs. Viral apps (<1%) are a thin but real right tail.

Installs Distribution


2. Does Quality = Success? (Surprising Answer: No)

Rating and Installs show a weak negative linear correlation (r = βˆ’0.31). Viral apps attract polarising reviews β€” millions of users means more critics.

Rating vs Installs


3. Monetization Strategy Matters

Ad-based and Hybrid (ads + IAP) apps reach significantly higher install counts. The price barrier for Premium apps dramatically limits reach.

Monetization vs Installs


4. The Gap Between Tiers is Enormous

Each tier is roughly 1,000x the previous β€” a textbook power-law gap.

TierMedian Installsn apps
Low~1008,135
Medium~10,0001,227
Viral~1,000,00074

Success Tiers ---

πŸ”‘ Key Findings

  1. 1.The market is winner-takes-all. 86% of apps sit in the Low tier. Viral apps represent less than 1% of 2020 launches.
  2. 2.Monetization is the strongest predictor. Ad-based and Hybrid models (ads + IAP) significantly outperform Pure Free and Premium.
  3. 3.Higher ratings do NOT predict more installs (r = βˆ’0.31). Quality alone is not the driver β€” distribution and visibility are.
  4. 4.Having a rating at all is a strong proxy for traction. Rated apps contain virtually all Medium and Viral tier apps.
  5. 5.Fresher apps win. Regular updates correlate negatively with dayssinceupdate (r = βˆ’0.25).
  6. 6.The Rating–Installs relationship is category-specific. A global r = βˆ’0.31 masks very different dynamics per category.

πŸ—‚οΈ Dataset

PropertyValue
SourceKaggle β€” gauthamp10/google-playstore-apps
Original size2.31M rows Γ— 24 columns
SubsetApps released in 2020 only
Sample size9,436 rows (Cochran's formula + stratified by Category)
ScrapedJune 2021

Features in this dataset

FeatureTypeDescription
App NamestringName of the app
CategorycategoricalApp category (48 unique)
RatingfloatAverage user rating (NaN if unrated)
Rating CountintNumber of user ratings
Installs_cleanintParsed install count
log_installsfloatlog(1 + Installs) β€” for modeling
Size_MBfloatApp size in MB (winsorized at 65.8MB)
Min_Android_verfloatMinimum Android version required
ReleaseddatetimeLaunch date (all in 2020)
Last UpdateddatetimeLast update date
FreeboolWhether app is free
Ad SupportedboolWhether app shows ads
In App PurchasesboolWhether app has IAP
Editors ChoiceboolWhether app is Editor's Choice
has_ratingboolWhether app has crossed rating threshold
app_age_daysintDays from release to scrape date
days_since_updateintDays from last update to scrape date
engagement_ratiofloatRating Count / Installs
log_engagementfloatlog(1 + engagement_ratio)
developer_app_countintTotal apps by same developer in 2020
monetization_modelcategoricalPure Free / Ad-based / Freemium / Hybrid / Premium
success_tiercategoricalTarget variable β€” Low / Medium / Viral

πŸ““ Notebook

The full analysis notebook is available here: πŸ‘‰ Open The Notebook in Google Colab ---

❓ Questions & Answers

Q1: Is the app market democratic β€” can any app go viral? No. 86% of apps released in 2020 never exceeded 10K installs. The market is winner-takes-all: less than 1% of apps reached the Viral tier (1M+ installs), and their median install count is roughly 1,000x that of Medium-tier apps.

Q2: Does a higher rating mean more installs? Surprisingly, no. Rating and Installs show a weak negative linear correlation (r = βˆ’0.31). Viral apps attract polarising reviews β€” millions of users means more critics. Quality alone is not the driver; distribution and visibility are.

Q3: Does monetization strategy matter? Yes β€” it's the strongest predictor in the dataset. Ad-based and Hybrid (ads + IAP) apps reach significantly higher install counts than Pure Free or Premium apps. The price barrier of Premium apps dramatically limits reach.

Q4: Does having a rating at all matter? Yes, dramatically. Apps with a visible rating contain virtually all Medium and Viral tier apps. This reflects a chicken-and-egg dynamic: installs drive ratings, ratings drive visibility, visibility drives more installs.

Q5: Does the Rating–Installs relationship hold across all categories? No. The global r = βˆ’0.31 masks very different dynamics per category. In Tools and Business, there is almost no relationship. In Entertainment and Games, high-install apps cluster at specific rating bands.


πŸ”§ Key Decisions

DecisionWhat I didWhy
Subset to 2020Filtered 544,882 apps released in 2020Consistent 6–18 month measurement window for all apps
Stratified samplingCochran's formula β†’ 9,436 rows, sampled proportionally by CategoryPreserve category distribution while reducing computational load
Rating = 0 β†’ NaNReplaced zero ratings with NaNZero means unrated, not bad β€” imputing would distort analysis
Rating not imputedKept 54.9% of Rating as NaN, created has_rating insteadMissing rating is a meaningful signal, not random noise
Size_MB winsorizedCapped at 65.8 MB (IQR upper fence)Extreme sizes are edge cases that distort correlations
Installs kept as-isNo capping on extreme install countsViral outliers ARE the story β€” removing them defeats the purpose
log_installsApplied log(1 + Installs) transformationRaw installs follow a power-law β€” log scale is needed for analysis
monetization_modelCombined Free + Ad Supported + In App Purchases into 5 categoriesThree boolean columns carry more meaning as a single business model feature

## πŸ‘€ Author

**Yonathan Levy**
Econ & Entrepreneurship with Data Science Specialization
Reichman Uni
Class of 2028