# Run Plan — Collect ~2 Lakh (200k) Apps + Scrape Overnight A step-by-step runbook for one full-time laptop. Goal: discover **~200,000 apps** (Google Play + Apple App Store) across multiple countries, then scrape full data for **US + AU** (add more as needed) overnight. > Full architecture & how everything works: see **ARCHITECTURE.html**. --- ## The strategy in one line **App Store** is the volume engine (~180 apps per search term) → it carries you to 200k. **Google Play** is slow (~30 per query, hard cap) → it adds ~40–60k using the fast `--deep` set. Running Play with the full wordlist would waste ~11 hours on duplicates, so we **don't**. --- ## PHASE 1 — Discover the apps (run during the day, ~3–6 hrs) ### Step 1. App Store — the big pull (full wordlist) ```powershell python discover.py --skip-play --target 160000 --deep --terms-file english_words.txt --countries us,au,gb,in,ca,de ``` ### Step 2. Google Play — bounded & fast (adds Play apps) ```powershell python discover.py --skip-appstore --deep --countries us,au,gb,in,ca,de ``` Both steps **merge into the same `apps.json`** (Step 2 keeps Step 1's results). More countries = more unique apps → edit the `--countries` list to match where you operate. ### Step 3. Back up the list immediately (it's a multi-hour artifact!) ```powershell Copy-Item apps.json apps_backup.json ``` ### Check the count anytime ```powershell python -c "import json;c=json.load(open('apps.json',encoding='utf-8'));print('Play',len(c['google_play']),'+ App Store',len(c['app_store']),'=',len(c['google_play'])+len(c['app_store']))" ``` If under 200k, add more countries (e.g. `,fr,br,jp,mx`) and re-run **Step 1**. --- ## PHASE 2 — Scrape the data (overnight) ### Before you start: set a faster pace (one-time) Open `apps.json` and set the delay in `settings` to `0.5` (currently `1.0`): ```jsonc "delay_seconds": 0.5 ``` Effective request rate ≈ `workers ÷ delay`. At `workers 8` and `delay 0.5` → ~16 req/s. ### Run the scrape (US + AU) ```powershell python scraper.py --countries us,au --workers 8 ``` - Produces **one row per app per country** → ~200k apps × 2 countries ≈ 400k rows. - Output: **`output/app_data.csv`** - Auto-backups every 1,000 rows to **`output/backups/`**. ### If it doesn't finish in one night Just **run the exact same command again** the next night — it **resumes** and skips everything already done. It's totally fine if this takes 1–2 nights. **Rough timing:** ~400k rows at ~16 req/s ≈ **7 hours** (one night). 3 countries ≈ 10–11 hrs. --- ## Safety (unattended overnight) - The **circuit breaker** auto-pauses all workers if a store rate-limits, then resumes. Seeing occasional `THROTTLING: pausing all workers` in the log = protection working, not an error. - If you wake up to **constant** throttling messages → lower to `--workers 6` next run. - Don't push workers too high while you sleep — an IP block at 2am costs more time than running slower. --- ## Recovery (if something gets deleted) - **`output/app_data.csv` lost?** Copy the newest snapshot back, then re-run (it resumes): ```powershell Copy-Item output\backups\app_data.bak3.csv output\app_data.csv python scraper.py --countries us,au --workers 8 ``` - **`apps.json` lost?** Restore your backup: ```powershell Copy-Item apps_backup.json apps.json ``` --- ## Output columns (`output/app_data.csv`) `store, country, app_id, title, developer, developer_email, developer_website, category, price, currency, free, avg_rating, total_ratings, text_review_count, last_updated, version, url` - `developer_email` → filled for almost all **Google Play** apps; **blank for App Store** (Apple's API has none). - `developer_website` → filled for both stores. --- ## Command cheat sheet ```powershell # DISCOVER (App Store volume + Play) python discover.py --skip-play --target 160000 --deep --terms-file english_words.txt --countries us,au,gb,in,ca,de python discover.py --skip-appstore --deep --countries us,au,gb,in,ca,de Copy-Item apps.json apps_backup.json # SCRAPE (overnight, resumable) python scraper.py --countries us,au --workers 8 # CHECK COUNT python -c "import json;c=json.load(open('apps.json',encoding='utf-8'));print(len(c['google_play'])+len(c['app_store']))" ``` --- ## Notes / limits (honest) - No store has an "all apps" list — you only get what discovery finds. 200k is realistic; *every* app is not. - Google Play search is capped at ~30 results/query — Play volume comes from many terms, not from one big query. - Developer **phone numbers** are not available from either source. - Data collected is **public, non-personal** app metadata. Using `developer_email` for bulk outreach is governed by anti-spam law (CAN-SPAM / GDPR / CASL) — get sign-off before any marketing use.