fineset-io/speculative-decoding-papers
Speculative Decoding Papers โ FineSet A research-paper dataset on Speculative Decoding Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. ๐ธ This is a dated snapshot โ generated 2026-06-19. It is not auto-updated. Research on Speculative Decoding Papers moves fast โ new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. โ Why this dataset Quality-scored:โฆ See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/speculative-decoding-papers.
Speculative Decoding Papers โ FineSet
A research-paper dataset on Speculative Decoding Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar.
๐ธ This is a dated snapshot โ generated 2026-06-19. It is not auto-updated. Research on Speculative Decoding Papers moves fast โ new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. โ
Why this dataset
- Quality-scored:
quality_scorefloat (0โ1), blends citations with recency + code/venue signals โ filter out the noise - Papers with code: 134 flagged via
has_codeโ find reproducible work fast - Deduplicated: arXiv + Semantic Scholar cross-referenced, duplicate records merged
- Clean JSONL: 485 records, one per line, normalized fields โ no encoding garbage
Dataset details
- Records: 485
- Date range: 2022โ2026
- Snapshot date: 2026-06-19 (frozen โ see note above)
- Sources: arXiv, Semantic Scholar (cross-referenced, duplicates merged)
- arXiv categories: cs.LG, cs.CL
- Quality scoring: citations + recency + code/venue blend, 0โ1 (p50=0.35, p90=0.61)
- Format: JSONL, one record per line
Fields
Quality score methodology
quality_score = max(impact, freshness), clamped to [0, 1], where:
- impact =
max( log10(citations+1)/4 , log10(influential_citations+1)/2 )โ realized impact (0.5 at 100 citations, ~0.75 at 1,000, 1.0 at 10,000+). - freshness =
recency ร (0.35 + 0.30ยทhas_code + 0.20ยทhas_venue)โ a baseline for recent papers (so a strong paper published this week isn't scored 0 just for lacking citations), whererecencyis 1.0 for papers โค60 days old and decays linearly to 0 by ~18 months.
Old highly-cited papers score on impact; brand-new papers score on freshness; old uncited papers score ~0. Useful for filtering training data by quality, not just age.
๐ Want this on YOUR topic, updated daily?
This snapshot is frozen at 2026-06-19. The live FineSet pipeline keeps a dataset like this refreshed every day on whatever topic you describe โ new papers in, dedup and quality scoring automatic, export as JSONL/Parquet or push straight to the Hub.
Tell me the topic you'd want and I'll run the pipeline on it โ open a discussion on this dataset, it's free and it's how I decide what to build next.
โ fineset.io โ describe what you want to train on, get a dataset. Early-access waitlist open (referral skip available).
