CoolFace
Datasetpublic

shaypal5/leadforge-lead-scoring-v1

LeadForge: Synthetic B2B Lead Scoring Dataset (leadforge-lead-scoring-v1) A relational, reproducible, three-tier synthetic CRM dataset family for teaching lead scoring at scale. Created by Shay Palachy Affek and generated by leadforge, an open-source Python framework for synthetic CRM/funnel data. The framework version is decoupled from the dataset version: the package stays at 1.x; the dataset is published under the explicit …-v1 tag. Why lead scoring matters in… See the full description on the dataset page: https://huggingface.co/datasets/shaypal5/leadforge-lead-scoring-v1.

sourceHugging Facemitupdated 4mo agoView on Hugging Face
1likes78downloads
Dataset Card

LeadForge: Synthetic B2B Lead Scoring Dataset (leadforge-lead-scoring-v1)

A relational, reproducible, three-tier synthetic CRM dataset family for teaching lead scoring at scale. Created by Shay Palachy Affek and generated by leadforge, an open-source Python framework for synthetic CRM/funnel data. The framework version is decoupled from the dataset version: the package stays at 1.x; the dataset is published under the explicit …-v1 tag.

Why lead scoring matters in 2024–2026

Mid-market SaaS vendors entered 2024–2026 with growth slowing and customer-acquisition costs rising (median public-SaaS growth 30%→25% from 2023 to 2025; New CAC Ratio rose materially in 2024), so predicting which leads convert within a fixed window has moved from a marketing nicety to a survival skill. This dataset teaches that skill on a relational substrate, with the realistic confusions (snapshot-window discipline, leakage traps, channel signal weaker than vendor blogs imply) that students will hit when they finally get hands on real CRM data.

What's inside

.
├── intro/ intermediate/ advanced/    # student_public bundles, one per difficulty tier
│   ├── manifest.json                 # provenance + file hashes
│   ├── metrics.json                  # per-tier headline metrics (medians + spreads)
│   ├── dataset_card.md               # auto-rendered per-bundle card
│   ├── feature_dictionary.csv        # authoritative column spec
│   ├── lead_scoring.csv              # flat convenience CSV (all splits)
│   ├── tables/*.parquet              # 7 snapshot-safe relational tables
│   └── tasks/converted_within_90_days/{train,valid,test}.parquet
├── docs/                             # vendored DGP / leakage / break-me docs (agent-readable)
├── metrics.json                      # top-level cross-tier metrics summary
├── claims_register.{md,json}         # claims → backing-artifact map (agent-readable)
├── README.md                         # this file (HF dataset card)
├── dataset-cover-image.png           # dataset thumbnail
└── LICENSE

student_public bundles ship the snapshot-safe relational view; research_instructor companions ship the full-horizon view plus the hidden causal structure (DAG, latent registry, mechanism summary) under metadata/. The full layout is documented in each bundle's manifest.json.

Agent-reviewable artifacts

The published bundle is self-contained for AI review and offline auditing — every numeric / structural claim on this page can be verified without following an external link:

  • `metrics.json` (root) + `<tier>/metrics.json` — deterministic JSON view of the headline LR AUC / AP / P@100 / Brier / conversion rate / cohort-shift / cross-tier-ordering medians, with JSON-path back-references to validation/validation_report.json (the source of truth).
  • `claims_register.{md,json}` — every numerical or structural claim on this page paired with the artifact and path that backs it. Rendered from claims_register_source.yaml by scripts/build_claims_register.py.
  • `docs/` — vendored copies of generation_method.md, channel_signal_audit.md, break_me_guide.md, feature_dictionary.md, v1_acceptance_gates_bands.yaml, v2_decision_log.md, plus a hand-authored relational_table_schemas.csv documenting every column of every relational table. These match the GitHub-blob links cited below but ship inside the bundle so a reviewer never needs network access.
  • `<tier>/manifest.json` — SHA-256 hash for every file plus the full redaction contract (structural_redactions.columns, omitted_tables, relational_snapshot_safe, snapshot_day).
  • Kaggle / HuggingFace preview pages additionally inject a schema.org/Dataset JSON-LD block in their <head> for agent ingestion without HTML parsing.

Quick start

python
# Flat CSV
df = pd.read_csv("intermediate/lead_scoring.csv")

# Parquet task splits (recommended)
train = pd.read_parquet("intermediate/tasks/converted_within_90_days/train.parquet")
test  = pd.read_parquet("intermediate/tasks/converted_within_90_days/test.parquet")

# Relational tables (feature engineering — example)
leads   = pd.read_parquet("intermediate/tables/leads.parquet")
touches = pd.read_parquet("intermediate/tables/touches.parquet")
my_touch_count = (
    touches.groupby("lead_id").size().rename("my_touch_count").reset_index()
)
features = leads.merge(my_touch_count, on="lead_id", how="left")

# Reproduce from source
# pip install leadforge
# leadforge generate --recipe b2b_saas_procurement_v1 --seed 42 \
#                    --mode student_public --difficulty intermediate --out my_bundle

The label converted_within_90_days resolves over a 90-day window; engagement features (touch_count, session_count, etc.) are computed strictly over events on days [0, 30]. The deliberate exception is total_touches_all, the leakage trap — flagged leakage_risk=True in feature_dictionary.csv. Drop it from your feature set unless you're demonstrating leakage detection.

Evaluation note — account and contact overlap

518 of 557 test accounts (≈93 %) appear in train on the intermediate bundle; the other tiers are similar. Contact-level overlap is comparable in magnitude: most test contacts also have activity in the training set. The random-split headline metrics therefore ride both account-level and contact-level signal across the split boundary and over-estimate generalisation to unseen accounts and contacts. For a faithful out-of-sample number, retrain with GroupKFold(account_id) and report both metrics. Notebook 02 demonstrates the detection recipe; `break_me_guide.md` §5 gives the worked example.

Dataset summary

Tiers are primarily a prevalence and noise axis. The difficulty gradient shows up most in AP, precision@k, and top-decile lift as the positive class shrinks; LR AUC declines more modestly (0.671 → 0.663 → 0.624). Choose a tier based on the teaching exercise:

IntroIntermediateAdvanced
Tier purposeHigh-prevalence warm-upDefault benchmarkLow-prevalence · calibration · noise exercise
Leads5,0005,0005,000
Accounts1,5001,5001,500
Contacts4,2004,2004,200
Snapshot columns31 / 34*31 / 34*31 / 34*
Targetconverted_within_90_daysconverted_within_90_daysconverted_within_90_days
Conversion rate (acceptance band, gate G7.\*)24–61%12–31%4–12%
Conversion rate (observed median, seeds 42–46)42.67%21.60%8.40%
Signal strength0.900.700.50
Noise scale0.100.300.55
Missing rate2%8%18%

\ `student_public` / `research_instructor`. Difficulty is modulated by the simulation engine — signal strength on latent-trait weights, Gaussian noise on float features, MCAR missingness, outlier rate — not post-hoc label flipping. The acceptance band is the recipe gate's tolerance window (`v1_acceptance_gates_bands.yaml` G7.\), not the achievable range — observed five-seed spreads sit comfortably inside the band.

The scenario

Veridian Technologies is a fictional Series B startup (Austin, US) selling Veridian Procure, a procurement / AP automation SaaS, to mid-market firms (200–2,000 employees) in the US and UK. The funnel runs through inbound marketing (45%), SDR outbound (35%), and partner referrals (20%); four personas drive deals (VP Finance, AP Manager, IT Director, Procurement Manager). Task: predict whether a lead converts (closed_won) within 90 days. ACV bands are $18k–$120k. See `docs/release/generation_method.md` for the full DGP, and the deeper "what's modelled / approximate / not modelled" breakdown that this README only summarises.

Public vs instructor: what's redacted

Filtering happens during rendering, not during simulation. The redaction contract is single-sourced in `leadforge/validation/leakage_probes.py`; the snapshot-safe writer and the validator import the same constants, so they cannot drift apart.

Source-of-truth constantPublic bundle treatment
BANNED_LEAD_COLUMNS = ("converted_within_90_days", "conversion_timestamp")Dropped from tables/leads.parquet
BANNED_OPP_COLUMNS = ("close_outcome", "closed_at")Dropped from tables/opportunities.parquet
BANNED_TABLES = ("customers", "subscriptions")Omitted from public bundles
SNAPSHOT_FILTERED_TABLES (touches, sessions, sales_activities, opportunities)Filtered per-lead by lead_created_at + snapshot_day
Snapshot redaction (current_stage, is_sql)Stripped from tasks/ splits and tables/leads.parquet
total_touches_all (deliberate trap)Retained in both modes; flagged leakage_risk=True

Each bundle's manifest.json records relational_snapshot_safe, redacted_columns, and snapshot_day, so the bundle is self-describing.

Calibration

Every realism / calibration / difficulty claim in this README is backed by `validation/validation_report.md`, regenerated by `scripts/validate_release_candidate.py` with bands declared in `docs/release/v1_acceptance_gates_bands.yaml`. Headline cross-seed medians (seeds 42–46):

TierLR AUCAPP@100Brier`calibration_max_bin_error`
intro0.6710.5550.600.2200.176
intermediate0.6630.3320.330.1600.279
advanced0.6240.1220.110.0760.221

Reading this table: LR AUC declines modestly from Intro to Advanced (0.671 → 0.663 → 0.624); the difficulty gradient is far more visible in AP, P@100, and top-decile lift, which fall steeply as the positive class shrinks. Brier score improves as prevalence falls (a prevalence effect, not better calibration); use calibration_max_bin_error to assess calibration quality. Advanced's median max-bin error of 0.221 (with high seed-to-seed variance: spread 0.56) signals meaningful miscalibration — a realistic exercise in probability calibration.

AP, P@100, conversion-rate, and lift orderings hold across the intended prevalence axis (intro > intermediate > advanced).

Intended uses

  • Teaching baseline lead-scoring on a flat snapshot.
  • Teaching relational feature engineering against snapshot-safe tables.
  • Teaching leakage detection (the total_touches_all trap is designed to be discoverable).
  • Teaching calibration, lift, P@K, value-aware ranking (expected_acv × P(convert)), and cohort-shift evaluation.
  • Comparing model families under a controlled DGP.

Out-of-scope uses

  • Production lead scoring. The company, product, and customers are fictional.
  • Vendor benchmarking / paper baselines. Difficulty tiers are calibrated for pedagogy, not cross-paper comparability.
  • Causal-inference research that requires recovery of the true DGP. The instructor companion exposes the hidden graph for teaching, not designed counterfactuals.
  • Demographic / fairness research. v1 does not model protected attributes.

Known limitations

  • Tiers are primarily a prevalence / noise axis. LR AUC declines modestly across tiers (0.671 / 0.663 / 0.624); the three tiers differ most visibly in conversion rate (43% / 22% / 8%), AP, P@K, and top-decile lift. Use AP, P@K, and calibration metrics to see the full difficulty gradient; AUC alone understates it.
  • 93% account and contact overlap across train / test splits. Random splits are keyed on lead ID; most test accounts and contacts also appear in train. Headline metrics over-state generalisation to unseen accounts and contacts. Use GroupKFold(account_id) for a faithful estimate.
  • GBM does not consistently beat LR (gate G7.4.4). GBM−LR AUC delta is slightly negative in every tier (intro −0.0105, intermediate −0.0179, advanced −0.0242); v1's snapshot is dominated by linear features. v2 will inject non-linear interactions in the simulator.
  • Channel signal is weak. Per `docs/release/channel_signal_audit.md`, out-of-sample univariate AUC of lead_source is ≈0.50–0.52 across all tiers and the per-channel rate spread is ≤0.05. The simulator does not encode channel-conditional probabilities; channel-conditional encoding is post-v1 work.
  • Cohort-shift degradation is small. v1 has no time-of-year drift baked in; the cohort-shift gate (G6.4) is informational and will bite in v2.
  • Advanced-tier noise can produce artifact zeros in count and duration columns. Gaussian noise is applied before MCAR missingness; the snapshot builder clamps results below zero to zero. What users observe is therefore not negative values but zeros that may be noise artifacts rather than true zero values — e.g. days_since_last_touch = 0 might mean "noised below zero, clamped" rather than "touched today". Treat suspicious zero clusters in the Advanced tier as intentional data-cleaning exercise material.

Composition

  • Entities. Accounts, contacts, leads, touches, sessions, sales_activities, opportunities (public); plus customers and subscriptions (instructor only). Per-row counts per bundle live in manifest.json.
  • Features. 31 public columns grouped by analytical role in `docs/release/feature_dictionary.md`; the per-bundle feature_dictionary.csv is the authoritative machine-readable spec.
  • Label. converted_within_90_days (boolean), event-derived from the simulator. Never sampled directly.
  • Splits. 70/15/15 train/valid/test, deterministic given seed; recorded in tasks/converted_within_90_days/task_manifest.json. Splits are keyed on lead_id; see the Evaluation note above for the account-overlap caveat.
  • Provenance. Recipe b2b_saas_procurement_v1, seed 42, package version stamped in manifest.json.

Maintenance, adversarial framing, license

We want the dataset to be broken. The break-me guide catalogues nine adversarial patterns to look for (leakage, split contamination, ranking inversions, calibration drift) with worked-example pointers back into the notebooks. Issue templates ship under .github/ISSUE_TEMPLATE/: a breakage report form for findings on the bundle itself, and a realism feedback form for distributional critiques. Accepted findings are logged in `docs/release/v2_decision_log.md`. File issues at leadforge-dev/leadforge; PRs welcome.

FieldValue
Generatorleadforge 1.0.0+
Recipeb2b_saas_procurement_v1
Canonical seed42 (cross-seed sweep: 42–46)
Bundle schema version5
FormatParquet (canonical) + CSV (convenience)
LicenseMIT — see LICENSE

Verify integrity with leadforge validate <bundle_dir>; every file is hashed in manifest.json.

Credits

Created by Shay Palachy Affek. Dataset generated with leadforge (MIT). Profiles: HuggingFace · Kaggle · GitHub