CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01opendatalab /SlimPajama-Meta-rater Annotated SlimPajama Dataset Dataset Description This dataset contains the first fully annotated SlimPajama dataset with comprehensive quality metrics for data-centric large language model research. The dataset includes approximately 580 billion tokens from the training set of the original SlimPajama dataset, annotated across 25 different quality dimensions. Note: This dataset contains only the training set portion of the original SlimPajama dataset, which is why the… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater.tabulartext-generation10M<n<100M7 likes5.3k downloads1y agoHugging Face02daisuke9999 /Japanese_NicoNico_Douga_Movie_Meta_Data_2016tabular10M<n<100M0 likes3k downloads2y agoHugging Face03bs-modeling-metadata /c4-en-html-with-metadatatabular10M<n<100M14 likes2.7k downloads4y agoHugging Face04toksuitebackup /meta-llama-Llama-3.2-1B-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular10M<n<100M0 likes1.2k downloads10mo agoHugging Face05yufan /arxiv-metadata-2020-2026 arXiv Metadata, enriched (2020–2026) Per-paper metadata for 1,517,185 arXiv papers spanning 2020-01 → 2026-09, enriched with abstracts, citation counts, and Semantic Scholar identifiers, and organized as a two-level hierarchy: field of study → year. Unlike a bare title index, every record here carries the abstract, the full author list, citation counts, and the Semantic Scholar corpusId, so you can do retrieval, classification, citation analysis, and corpus building directly… See the full description on the dataset page: https://huggingface.co/datasets/yufan/arxiv-metadata-2020-2026.tabulartext-retrieval1M<n<10M0 likes1.1k downloads6d agoHugging Face06masoudjs /c4-en-html-with-metadata-ppl-cleanFile list: "c4-en-html_cc-main-2019-18_pq00-000.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-001.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-002.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-003.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-004.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-005.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-006.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-007.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-008.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-009.jsonl.gz"… See the full description on the dataset page: https://huggingface.co/datasets/masoudjs/c4-en-html-with-metadata-ppl-clean.tabular10K<n<100K1 likes1.1k downloads4y agoHugging Face07bs-modeling-metadata /c4-en-html-with-training_metadata_alltabular10K<n<100K1 likes954 downloads3y agoHugging Face08Complementarity /gpqa-metadata-blind-answertabularn<1K0 likes463 downloads1mo agoHugging Face09GMLRVigil /BenchCheck-MetaBenchmark BenchCheck meta-benchmark (v6) 840 multiple-choice video questions from 76 public video benchmarks, one item per video, four capability groups x 210 (Perception, Temporal, Spatial / physical, Reasoning / knowledge). The set is the difficulty-first census of the BenchCheck screened pool: items that none of the cheap attackers of the pool screening (text-only, single frame, 32-frame 2B model, options-only, shuffled frames) could solve, ranked by the worst-case attacker percentile… See the full description on the dataset page: https://huggingface.co/datasets/GMLRVigil/BenchCheck-MetaBenchmark.tabularvideo-classification1K<n<10K0 likes383 downloads3d agoHugging Face10opendatalab /SlimPajama-Meta-rater-Readability-30B Top 30B token SlimPajama Subset selected by the Readability rater This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models. Code: https://github.com/opendatalab/Meta-rater Dataset Description This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Readability dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness) framework.… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Readability-30B.tabulartext-generation1M<n<10M1 likes361 downloads1y agoHugging Face11MirandaAbhilash /vqav2-full-metadatatabular100K<n<1M0 likes312 downloads6mo agoHugging Face12rafmacalaba /fcv-extractions-meta-tiered-probe fcv-extractions-meta-tiered-probe Data-use mention extractions from the World Bank Fragility, Conflict and Violence (FCV) document corpus. Spans are extracted by the fine-tuned GLiNER model rafmacalaba/gliner_datause_tiered, scored by the tier-probe head rafmacalaba/gliner-tier-probe, and attributed (provenance + usage/impact) by rafmacalaba/lfm2.5-350M-datause-multitask-tiered (a LoRA SFT of LiquidAI/LFM2.5-350M). Shape nested — one row per chunk; each entity in… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/fcv-extractions-meta-tiered-probe.tabular10K<n<100K0 likes203 downloads23d agoHugging Face13opendatalab /SlimPajama-Meta-rater-Professionalism-30B Top 30B token SlimPajama Subset selected by the Professionalism rater This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models. Code: https://github.com/opendatalab/Meta-rater Dataset Description This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Professionalism dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness)… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Professionalism-30B.tabulartext-generation1M<n<10M0 likes198 downloads1y agoHugging Face14opendatalab /SlimPajama-Meta-rater-Reasoning-30B Top 30B token SlimPajama Subset selected by the Reasoning rater This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models. Code: https://github.com/opendatalab/Meta-rater Dataset Description This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Reasoning dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness) framework. Each… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Reasoning-30B.tabulartext-generation1M<n<10M1 likes189 downloads1y agoHugging Face15yufan /amazon2023-item-metadata Amazon Reviews 2023 — Item Metadata (content features) Item content features (title, images, price, brand/store, categories, features, description, …) for five Amazon Reviews 2023 categories, aligned with the user-interaction splits in yufan/amazon2023-user-interactions. One config per category; each has a single train split with one item per line. Join to the interactions/sequences via parent_asin. Coverage is 100 % of the 5-core items in the companion dataset. from datasets… See the full description on the dataset page: https://huggingface.co/datasets/yufan/amazon2023-item-metadata.tabularother100K<n<1M0 likes181 downloads3mo agoHugging Face16rafmacalaba /fcv-extractions-meta-tiered fcv-extractions-meta-tiered Data-use mention extractions from the World Bank Fragility, Conflict and Violence (FCV) document corpus, enriched by the fine-tuned rafmacalaba/lfm2.5-350M-datause-multitask model (a LoRA SFT of LiquidAI/LFM2.5-350M). Shape nested — one row per chunk; each entity in entities[] carries provenance and usage_impact alongside its span. Configs config rows fcv_pads_east_africa 862,663 jdc_operational 2,468… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/fcv-extractions-meta-tiered.tabular10K<n<100K0 likes179 downloads25d agoHugging Face17opendatalab /Meta-rater-PRRC-Rater-dataset PRRC Rater Training and Evaluation Dataset Dataset Description This dataset contains the full training and evaluation data for the PRRC rater models described in Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models. It is designed for training and benchmarking models that score text along four key quality dimensions: Professionalism, Readability, Reasoning, and Cleanliness. Source: Subset of SlimPajama-627B, annotated for PRRC dimensions… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/Meta-rater-PRRC-Rater-dataset.tabulartext-classification100K<n<1M1 likes177 downloads1y agoHugging Face18rafmacalaba /fcv-extractions-meta fcv-extractions-meta Data-use mention extractions from the World Bank Fragility, Conflict and Violence (FCV) document corpus, enriched by the fine-tuned rafmacalaba/lfm2.5-350M-datause-multitask model (a LoRA SFT of LiquidAI/LFM2.5-350M). Shape nested — one row per chunk; each entity in entities[] carries provenance and usage_impact alongside its span. Configs config rows fcv_pads_east_africa 793,763 jdc_operational 12,372… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/fcv-extractions-meta.tabular100K<n<1M0 likes118 downloads1mo agoHugging Face19opendatalab /SlimPajama-Meta-rater-Cleanliness-30B Top 30B token SlimPajama Subset selected by the Cleanliness rater This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models. Code: https://github.com/opendatalab/Meta-rater Dataset Description This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Cleanliness dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness) framework.… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Cleanliness-30B.tabulartext-generation1M<n<10M0 likes105 downloads1y agoHugging Face20Rabrg /gutenberg_rich_metadataAround 1.1k Project Gutenberg ebooks with full Gutenberg metadata plus additional metadata from Greatest Books. See Rabrg/gutenberg_full for 70k+ books, but without Greatest Books metadata attached. tabular1K<n<10K2 likes93 downloads1y agoHugging Face21open-llm-leaderboard /meta-llama__Meta-Llama-3-70B-Instruct-detailsgated Dataset Card for Evaluation run of meta-llama/Meta-Llama-3-70B-Instruct Dataset automatically created during the evaluation run of model meta-llama/Meta-Llama-3-70B-Instruct The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meta-llama__Meta-Llama-3-70B-Instruct-details.tabular10K<n<100K1 likes90 downloads2y agoHugging Face22nyu-dice-lab /lm-eval-results-ntnhan-Llama3-8B-MetaMath-private Dataset Card for Evaluation run of ntnhan/Llama3-8B-MetaMath Dataset automatically created during the evaluation run of model ntnhan/Llama3-8B-MetaMath The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-ntnhan-Llama3-8B-MetaMath-private.tabular100K<n<1M0 likes77 downloads2y agoHugging Face23anote-ai /meta-routing MetaRouting Dataset This dataset contains synthetic benchmark artifacts for the Research MetaRouting project, covering meta-decision policies for agentic workflows: when to answer directly, decompose, retrieve, execute code, delegate, verify, or recover from failures. Source repository: https://github.com/anote-ai/Research-MetaRouting Displayable Configs The Hugging Face viewer reads normalized JSONL tables under viewer/: dai2026_traces, dai2026_tasks… See the full description on the dataset page: https://huggingface.co/datasets/anote-ai/meta-routing.tabulartext-classification10K<n<100K0 likes72 downloads1mo agoHugging Face24open-llm-leaderboard /meta-llama__Llama-2-7b-chat-hf-detailsgated Dataset Card for Evaluation run of meta-llama/Llama-2-7b-chat-hf Dataset automatically created during the evaluation run of model meta-llama/Llama-2-7b-chat-hf The dataset is composed of 39 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meta-llama__Llama-2-7b-chat-hf-details.tabular10K<n<100K0 likes58 downloads2y agoHugging Face25open-llm-leaderboard /meta-llama__Llama-3.1-8B-Instruct-detailsgated Dataset Card for Evaluation run of meta-llama/Llama-3.1-8B-Instruct Dataset automatically created during the evaluation run of model meta-llama/Llama-3.1-8B-Instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meta-llama__Llama-3.1-8B-Instruct-details.tabular10K<n<100K0 likes53 downloads2y agoHugging Face26open-llm-leaderboard /meta-llama__Meta-Llama-3-8B-Instruct-detailsgated Dataset Card for Evaluation run of meta-llama/Meta-Llama-3-8B-Instruct Dataset automatically created during the evaluation run of model meta-llama/Meta-Llama-3-8B-Instruct The dataset is composed of 76 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 23 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meta-llama__Meta-Llama-3-8B-Instruct-details.tabular10K<n<100K0 likes52 downloads2y agoHugging Face27open-llm-leaderboard /gbueno86__Meta-LLama-3-Cat-Smaug-LLama-70b-detailsgated Dataset Card for Evaluation run of gbueno86/Meta-LLama-3-Cat-Smaug-LLama-70b Dataset automatically created during the evaluation run of model gbueno86/Meta-LLama-3-Cat-Smaug-LLama-70b The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/gbueno86__Meta-LLama-3-Cat-Smaug-LLama-70b-details.tabular10K<n<100K0 likes51 downloads2y agoHugging Face28open-llm-leaderboard /failspy__Meta-Llama-3-70B-Instruct-abliterated-v3.5-detailsgated Dataset Card for Evaluation run of failspy/Meta-Llama-3-70B-Instruct-abliterated-v3.5 Dataset automatically created during the evaluation run of model failspy/Meta-Llama-3-70B-Instruct-abliterated-v3.5 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/failspy__Meta-Llama-3-70B-Instruct-abliterated-v3.5-details.tabular10K<n<100K0 likes48 downloads2y agoHugging Face29open-llm-leaderboard /meta-llama__Llama-3.1-70B-Instruct-detailsgated Dataset Card for Evaluation run of meta-llama/Llama-3.1-70B-Instruct Dataset automatically created during the evaluation run of model meta-llama/Llama-3.1-70B-Instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meta-llama__Llama-3.1-70B-Instruct-details.tabular10K<n<100K2 likes46 downloads2y agoHugging Face30open-llm-leaderboard /meta-llama__Llama-3.2-3B-Instruct-detailsgated Dataset Card for Evaluation run of meta-llama/Llama-3.2-3B-Instruct Dataset automatically created during the evaluation run of model meta-llama/Llama-3.2-3B-Instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meta-llama__Llama-3.2-3B-Instruct-details.tabular10K<n<100K0 likes46 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.