Files
App_Scraper/RUN_PLAN.md
T
dhruv.godvani 22582bdcd7 Add developer email and website fields to CSV output
- Updated CSV_FIELDS to include "developer_email" and "developer_website".
- Modified play_row function to extract developer email and website from the response.
- Updated appstore_row function to include developer website, with a note that developer email is not available from Apple's API.
2026-06-18 12:36:17 +05:30

4.7 KiB
Raw Blame History

Run Plan — Collect ~2 Lakh (200k) Apps + Scrape Overnight

A step-by-step runbook for one full-time laptop. Goal: discover ~200,000 apps (Google Play + Apple App Store) across multiple countries, then scrape full data for US + AU (add more as needed) overnight.

Full architecture & how everything works: see ARCHITECTURE.html.


The strategy in one line

App Store is the volume engine (~180 apps per search term) → it carries you to 200k. Google Play is slow (~30 per query, hard cap) → it adds ~4060k using the fast --deep set. Running Play with the full wordlist would waste ~11 hours on duplicates, so we don't.


PHASE 1 — Discover the apps (run during the day, ~36 hrs)

Step 1. App Store — the big pull (full wordlist)

python discover.py --skip-play --target 160000 --deep --terms-file english_words.txt --countries us,au,gb,in,ca,de

Step 2. Google Play — bounded & fast (adds Play apps)

python discover.py --skip-appstore --deep --countries us,au,gb,in,ca,de

Both steps merge into the same apps.json (Step 2 keeps Step 1's results). More countries = more unique apps → edit the --countries list to match where you operate.

Step 3. Back up the list immediately (it's a multi-hour artifact!)

Copy-Item apps.json apps_backup.json

Check the count anytime

python -c "import json;c=json.load(open('apps.json',encoding='utf-8'));print('Play',len(c['google_play']),'+ App Store',len(c['app_store']),'=',len(c['google_play'])+len(c['app_store']))"

If under 200k, add more countries (e.g. ,fr,br,jp,mx) and re-run Step 1.


PHASE 2 — Scrape the data (overnight)

Before you start: set a faster pace (one-time)

Open apps.json and set the delay in settings to 0.5 (currently 1.0):

"delay_seconds": 0.5

Effective request rate ≈ workers ÷ delay. At workers 8 and delay 0.5 → ~16 req/s.

Run the scrape (US + AU)

python scraper.py --countries us,au --workers 8
  • Produces one row per app per country → ~200k apps × 2 countries ≈ 400k rows.
  • Output: output/app_data.csv
  • Auto-backups every 1,000 rows to output/backups/.

If it doesn't finish in one night

Just run the exact same command again the next night — it resumes and skips everything already done. It's totally fine if this takes 12 nights.

Rough timing: ~400k rows at ~16 req/s ≈ 7 hours (one night). 3 countries ≈ 1011 hrs.


Safety (unattended overnight)

  • The circuit breaker auto-pauses all workers if a store rate-limits, then resumes. Seeing occasional THROTTLING: pausing all workers in the log = protection working, not an error.
  • If you wake up to constant throttling messages → lower to --workers 6 next run.
  • Don't push workers too high while you sleep — an IP block at 2am costs more time than running slower.

Recovery (if something gets deleted)

  • output/app_data.csv lost? Copy the newest snapshot back, then re-run (it resumes):
    Copy-Item output\backups\app_data.bak3.csv output\app_data.csv
    python scraper.py --countries us,au --workers 8
    
  • apps.json lost? Restore your backup:
    Copy-Item apps_backup.json apps.json
    

Output columns (output/app_data.csv)

store, country, app_id, title, developer, developer_email, developer_website, category, price, currency, free, avg_rating, total_ratings, text_review_count, last_updated, version, url

  • developer_email → filled for almost all Google Play apps; blank for App Store (Apple's API has none).
  • developer_website → filled for both stores.

Command cheat sheet

# DISCOVER (App Store volume + Play)
python discover.py --skip-play --target 160000 --deep --terms-file english_words.txt --countries us,au,gb,in,ca,de
python discover.py --skip-appstore --deep --countries us,au,gb,in,ca,de
Copy-Item apps.json apps_backup.json

# SCRAPE (overnight, resumable)
python scraper.py --countries us,au --workers 8

# CHECK COUNT
python -c "import json;c=json.load(open('apps.json',encoding='utf-8'));print(len(c['google_play'])+len(c['app_store']))"

Notes / limits (honest)

  • No store has an "all apps" list — you only get what discovery finds. 200k is realistic; every app is not.
  • Google Play search is capped at ~30 results/query — Play volume comes from many terms, not from one big query.
  • Developer phone numbers are not available from either source.
  • Data collected is public, non-personal app metadata. Using developer_email for bulk outreach is governed by anti-spam law (CAN-SPAM / GDPR / CASL) — get sign-off before any marketing use.