datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SlimPajama-Meta-rater
Annotated SlimPajama Dataset
Dataset Description
This dataset contains the first fully annotated SlimPajama dataset with comprehensive quality metrics for data-centric large language model research. The dataset includes approximately 580 billion tokens from the training set of the original SlimPajama dataset, annotated across 25 different quality dimensions.
Note: This dataset contains only the training set portion of the original SlimPajama dataset, which is why the… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater.Japanese_NicoNico_Douga_Movie_Meta_Data_2016c4-en-html-with-metadatameta-llama-Llama-3.2-1B-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
arxiv-metadata-2020-2026
arXiv Metadata, enriched (2020–2026)
Per-paper metadata for 1,517,185 arXiv papers spanning 2020-01 → 2026-09,
enriched with abstracts, citation counts, and Semantic Scholar identifiers, and
organized as a two-level hierarchy: field of study → year.
Unlike a bare title index, every record here carries the abstract, the
full author list, citation counts, and the Semantic Scholar corpusId,
so you can do retrieval, classification, citation analysis, and corpus building
directly… See the full description on the dataset page: https://huggingface.co/datasets/yufan/arxiv-metadata-2020-2026.c4-en-html-with-metadata-ppl-cleanFile list:
"c4-en-html_cc-main-2019-18_pq00-000.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-001.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-002.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-003.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-004.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-005.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-006.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-007.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-008.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-009.jsonl.gz"… See the full description on the dataset page: https://huggingface.co/datasets/masoudjs/c4-en-html-with-metadata-ppl-clean.c4-en-html-with-training_metadata_allgpqa-metadata-blind-answerBenchCheck-MetaBenchmark
BenchCheck meta-benchmark (v6)
840 multiple-choice video questions from 76 public video benchmarks, one item per video,
four capability groups x 210 (Perception, Temporal, Spatial / physical, Reasoning / knowledge). The set is the
difficulty-first census of the BenchCheck screened pool: items that none of the cheap attackers of the pool screening
(text-only, single frame, 32-frame 2B model, options-only, shuffled frames) could solve, ranked by the worst-case attacker
percentile… See the full description on the dataset page: https://huggingface.co/datasets/GMLRVigil/BenchCheck-MetaBenchmark.SlimPajama-Meta-rater-Readability-30B
Top 30B token SlimPajama Subset selected by the Readability rater
This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models.
Code: https://github.com/opendatalab/Meta-rater
Dataset Description
This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Readability dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness) framework.… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Readability-30B.vqav2-full-metadatafcv-extractions-meta-tiered-probe
fcv-extractions-meta-tiered-probe
Data-use mention extractions from the World Bank Fragility, Conflict and Violence (FCV) document corpus. Spans are extracted by the fine-tuned GLiNER model rafmacalaba/gliner_datause_tiered, scored by the tier-probe head rafmacalaba/gliner-tier-probe, and attributed (provenance + usage/impact) by rafmacalaba/lfm2.5-350M-datause-multitask-tiered (a LoRA SFT of LiquidAI/LFM2.5-350M).
Shape
nested — one row per chunk; each entity in… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/fcv-extractions-meta-tiered-probe.SlimPajama-Meta-rater-Professionalism-30B
Top 30B token SlimPajama Subset selected by the Professionalism rater
This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models.
Code: https://github.com/opendatalab/Meta-rater
Dataset Description
This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Professionalism dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness)… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Professionalism-30B.SlimPajama-Meta-rater-Reasoning-30B
Top 30B token SlimPajama Subset selected by the Reasoning rater
This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models.
Code: https://github.com/opendatalab/Meta-rater
Dataset Description
This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Reasoning dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness) framework. Each… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Reasoning-30B.amazon2023-item-metadata
Amazon Reviews 2023 — Item Metadata (content features)
Item content features (title, images, price, brand/store, categories,
features, description, …) for five Amazon Reviews 2023 categories, aligned with
the user-interaction splits in
yufan/amazon2023-user-interactions.
One config per category; each has a single train split with one item per
line. Join to the interactions/sequences via parent_asin. Coverage is
100 % of the 5-core items in the companion dataset.
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/yufan/amazon2023-item-metadata.fcv-extractions-meta-tiered
fcv-extractions-meta-tiered
Data-use mention extractions from the World Bank Fragility, Conflict and Violence (FCV) document corpus, enriched by the fine-tuned rafmacalaba/lfm2.5-350M-datause-multitask model (a LoRA SFT of LiquidAI/LFM2.5-350M).
Shape
nested — one row per chunk; each entity in entities[] carries provenance and usage_impact alongside its span.
Configs
config
rows
fcv_pads_east_africa
862,663
jdc_operational
2,468… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/fcv-extractions-meta-tiered.Meta-rater-PRRC-Rater-dataset
PRRC Rater Training and Evaluation Dataset
Dataset Description
This dataset contains the full training and evaluation data for the PRRC rater models described in Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models. It is designed for training and benchmarking models that score text along four key quality dimensions: Professionalism, Readability, Reasoning, and Cleanliness.
Source: Subset of SlimPajama-627B, annotated for PRRC dimensions… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/Meta-rater-PRRC-Rater-dataset.fcv-extractions-meta
fcv-extractions-meta
Data-use mention extractions from the World Bank Fragility, Conflict and Violence (FCV) document corpus, enriched by the fine-tuned rafmacalaba/lfm2.5-350M-datause-multitask model (a LoRA SFT of LiquidAI/LFM2.5-350M).
Shape
nested — one row per chunk; each entity in entities[] carries provenance and usage_impact alongside its span.
Configs
config
rows
fcv_pads_east_africa
793,763
jdc_operational
12,372… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/fcv-extractions-meta.SlimPajama-Meta-rater-Cleanliness-30B
Top 30B token SlimPajama Subset selected by the Cleanliness rater
This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models.
Code: https://github.com/opendatalab/Meta-rater
Dataset Description
This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Cleanliness dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness) framework.… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Cleanliness-30B.gutenberg_rich_metadataAround 1.1k Project Gutenberg ebooks with full Gutenberg metadata plus additional metadata from Greatest Books. See Rabrg/gutenberg_full for 70k+ books, but without Greatest Books metadata attached.
meta-llama__Meta-Llama-3-70B-Instruct-details
Dataset Card for Evaluation run of meta-llama/Meta-Llama-3-70B-Instruct
Dataset automatically created during the evaluation run of model meta-llama/Meta-Llama-3-70B-Instruct
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meta-llama__Meta-Llama-3-70B-Instruct-details.lm-eval-results-ntnhan-Llama3-8B-MetaMath-private
Dataset Card for Evaluation run of ntnhan/Llama3-8B-MetaMath
Dataset automatically created during the evaluation run of model ntnhan/Llama3-8B-MetaMath
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-ntnhan-Llama3-8B-MetaMath-private.meta-routing
MetaRouting Dataset
This dataset contains synthetic benchmark artifacts for the Research MetaRouting project, covering meta-decision policies for agentic workflows: when to answer directly, decompose, retrieve, execute code, delegate, verify, or recover from failures.
Source repository: https://github.com/anote-ai/Research-MetaRouting
Displayable Configs
The Hugging Face viewer reads normalized JSONL tables under viewer/:
dai2026_traces, dai2026_tasks… See the full description on the dataset page: https://huggingface.co/datasets/anote-ai/meta-routing.meta-llama__Llama-2-7b-chat-hf-details
Dataset Card for Evaluation run of meta-llama/Llama-2-7b-chat-hf
Dataset automatically created during the evaluation run of model meta-llama/Llama-2-7b-chat-hf
The dataset is composed of 39 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meta-llama__Llama-2-7b-chat-hf-details.meta-llama__Llama-3.1-8B-Instruct-details
Dataset Card for Evaluation run of meta-llama/Llama-3.1-8B-Instruct
Dataset automatically created during the evaluation run of model meta-llama/Llama-3.1-8B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meta-llama__Llama-3.1-8B-Instruct-details.meta-llama__Meta-Llama-3-8B-Instruct-details
Dataset Card for Evaluation run of meta-llama/Meta-Llama-3-8B-Instruct
Dataset automatically created during the evaluation run of model meta-llama/Meta-Llama-3-8B-Instruct
The dataset is composed of 76 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 23 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meta-llama__Meta-Llama-3-8B-Instruct-details.gbueno86__Meta-LLama-3-Cat-Smaug-LLama-70b-details
Dataset Card for Evaluation run of gbueno86/Meta-LLama-3-Cat-Smaug-LLama-70b
Dataset automatically created during the evaluation run of model gbueno86/Meta-LLama-3-Cat-Smaug-LLama-70b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/gbueno86__Meta-LLama-3-Cat-Smaug-LLama-70b-details.failspy__Meta-Llama-3-70B-Instruct-abliterated-v3.5-details
Dataset Card for Evaluation run of failspy/Meta-Llama-3-70B-Instruct-abliterated-v3.5
Dataset automatically created during the evaluation run of model failspy/Meta-Llama-3-70B-Instruct-abliterated-v3.5
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/failspy__Meta-Llama-3-70B-Instruct-abliterated-v3.5-details.meta-llama__Llama-3.1-70B-Instruct-details
Dataset Card for Evaluation run of meta-llama/Llama-3.1-70B-Instruct
Dataset automatically created during the evaluation run of model meta-llama/Llama-3.1-70B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meta-llama__Llama-3.1-70B-Instruct-details.meta-llama__Llama-3.2-3B-Instruct-details
Dataset Card for Evaluation run of meta-llama/Llama-3.2-3B-Instruct
Dataset automatically created during the evaluation run of model meta-llama/Llama-3.2-3B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meta-llama__Llama-3.2-3B-Instruct-details.
