CoolFace
Datasetpublic

FlyRank/internship-starter

FlyRank Internship — Starter Dataset (Anonymized) The public, safe starting point for the FlyRank Applied Search Intelligence ML internship. 30,000 anonymized content-performance rows across 32 pseudonymized clients (53 columns). Public-safe: hashed content_id / client_id + numeric/categorical metrics only — no titles, URLs, keywords, domains, or client names. What it's for Week 1–2 quick wins and the ready-now capstone lanes (ranking-signal analysis, lifecycle /… See the full description on the dataset page: https://huggingface.co/datasets/FlyRank/internship-starter.

sourceHugging Faceotherupdated 3mo agoView on Hugging Face
5likes121downloads
Dataset Card

FlyRank Internship — Starter Dataset (Anonymized)

The public, safe starting point for the FlyRank Applied Search Intelligence ML internship. 30,000 anonymized content-performance rows across 32 pseudonymized clients (53 columns).

Public-safe: hashed content_id / client_id + numeric/categorical metrics only — no titles, URLs, keywords, domains, or client names.

What it's for

Week 1–2 quick wins and the ready-now capstone lanes (ranking-signal analysis, lifecycle / opportunity scoring, content-archetype clustering).

Verified reference results (this 30k slice)

  • Rule baseline Precision@50 = 0.26 → Random Forest Precision@50 = 0.74
  • search_volume vs impressions_90d correlation ≈ 0.0012 (essentially zero — a real myth-buster)
  • Weighted CTR by position: top_3 0.49%page_1 0.35% → deep 0.04%
  • Length is not the differentiator: growing vs declining word count ≈ 2,850 vs 2,910

Safety rules

Anonymized, but still treat row-level outputs as not-for-careless-publishing. Do not use product flags (health_score, needs_ctr_fix, is_quick_win, …) as model features — they leak the decline label. Keep all public outputs anonymized/aggregate.