LeoZotos/fineweb-edu-usmle
FineWeb-Edu USMLE This is a paragraph-level subset of LeoZotos/fineweb-edu-topics ranked by usmle_similarity. The 2.5B configuration is the highest-ranked core. The 5B configuration contains that same core plus the extension; the shared core files are stored only once. Token budgets use allenai/OLMo-2-0425-1B at revision stage1-step1907359-tokens4001B and include one EOS document boundary per paragraph. The paragraph crossing each target is retained, so the actual token count is… See the full description on the dataset page: https://huggingface.co/datasets/LeoZotos/fineweb-edu-usmle.
FineWeb-Edu USMLE
This is a paragraph-level subset of LeoZotos/fineweb-edu-topics ranked by usmle_similarity. The 2.5B configuration is the highest-ranked core. The 5B configuration contains that same core plus the extension; the shared core files are stored only once.
Token budgets use allenai/OLMo-2-0425-1B at revision stage1-step1907359-tokens4001B and include one EOS document boundary per paragraph. The paragraph crossing each target is retained, so the actual token count is slightly above the nominal size. No additional paragraph-level deduplication is applied; the upstream FineWeb-Edu data is retained as provided.
Rows are deterministically shuffled within deterministically ordered Parquet shards. selection_rank preserves the original relevance ordering. Similarity scores are raw cosine-similarity aggregates, not calibrated probabilities.
