CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01meta-math /MetaMathQAView the project page: https://meta-math.github.io/ see our paper at https://arxiv.org/abs/2309.12284 Note All MetaMathQA data are augmented from the training sets of GSM8K and MATH. None of the augmented data is from the testing set. You can check the original_question in meta-math/MetaMathQA, each item is from the GSM8K or MATH train set. Model Details MetaMath-Mistral-7B is fully fine-tuned on the MetaMathQA datasets and based on the powerful Mistral-7B model. It is… See the full description on the dataset page: https://huggingface.co/datasets/meta-math/MetaMathQA.text100K<n<1M476 likes113k downloads3y agoHugging Face02opendatalab /SlimPajama-Meta-rater Annotated SlimPajama Dataset Dataset Description This dataset contains the first fully annotated SlimPajama dataset with comprehensive quality metrics for data-centric large language model research. The dataset includes approximately 580 billion tokens from the training set of the original SlimPajama dataset, annotated across 25 different quality dimensions. Note: This dataset contains only the training set portion of the original SlimPajama dataset, which is why the… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater.tabulartext-generation10M<n<100M7 likes5.3k downloads1y agoHugging Face03facebook /meta-active-readingtext1B<n<10B37 likes5k downloads1y agoHugging Face04allenai /metaicl-dataThis is the downloaded and processed data from Meta's MetaICL. We follow their "How to Download and Preprocess" instructions to obtain their modified versions of CrossFit and UnifiedQA. Citation information @inproceedings{ min2022metaicl, title={ Meta{ICL}: Learning to Learn In Context }, author={ Min, Sewon and Lewis, Mike and Zettlemoyer, Luke and Hajishirzi, Hannaneh }, booktitle={ NAACL-HLT }, year={ 2022 } } @inproceedings{ ye2021crossfit, title={ {C}ross{F}it:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/metaicl-data.text100K<n<1M5 likes3.9k downloads4y agoHugging Face05meta-math /MetaMathQA-40Karxiv.org/abs/2309.12284 View the project page: https://meta-math.github.io/ text10K<n<100K27 likes3.3k downloads3y agoHugging Face06huggingface /transformers-metadata Transformers metadata text1K<n<10K44 likes3.1k downloads2d agoHugging Face07daisuke9999 /Japanese_NicoNico_Douga_Movie_Meta_Data_2016tabular10M<n<100M0 likes3k downloads2y agoHugging Face08bs-modeling-metadata /c4-en-html-with-metadatatabular10M<n<100M14 likes2.7k downloads4y agoHugging Face09friendshipkim /metaiclThis is the downloaded and processed data from Meta's MetaICL. We follow their "How to Download and Preprocess" instructions to obtain their modified versions of CrossFit and UnifiedQA. Citation information @inproceedings{ min2022metaicl, title={ Meta{ICL}: Learning to Learn In Context }, author={ Min, Sewon and Lewis, Mike and Zettlemoyer, Luke and Hajishirzi, Hannaneh }, booktitle={ NAACL-HLT }, year={ 2022 } } @inproceedings{ ye2021crossfit, title={ {C}ross{F}it:… See the full description on the dataset page: https://huggingface.co/datasets/friendshipkim/metaicl.text1M<n<10M0 likes2.2k downloads2y agoHugging Face10huggingface /diffusers-metadatatextn<1K35 likes2k downloads2h agoHugging Face11meta-math /GSM8K_zh Dataset GSM8K_zh is a dataset for mathematical reasoning in Chinese, question-answer pairs are translated from GSM8K (https://github.com/openai/grade-school-math/tree/master) by GPT-3.5-Turbo with few-shot prompting. The dataset consists of 7473 training samples and 1319 testing samples. The former is for supervised fine-tuning, while the latter is for evaluation. for training samples, question_zh and answer_zh are question and answer keys, respectively; for testing samples, only… See the full description on the dataset page: https://huggingface.co/datasets/meta-math/GSM8K_zh.textquestion-answering1K<n<10K30 likes1.3k downloads3y agoHugging Face12toksuitebackup /meta-llama-Llama-3.2-1B-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular10M<n<100M0 likes1.2k downloads10mo agoHugging Face13yufan /arxiv-metadata-2020-2026 arXiv Metadata, enriched (2020–2026) Per-paper metadata for 1,517,185 arXiv papers spanning 2020-01 → 2026-09, enriched with abstracts, citation counts, and Semantic Scholar identifiers, and organized as a two-level hierarchy: field of study → year. Unlike a bare title index, every record here carries the abstract, the full author list, citation counts, and the Semantic Scholar corpusId, so you can do retrieval, classification, citation analysis, and corpus building directly… See the full description on the dataset page: https://huggingface.co/datasets/yufan/arxiv-metadata-2020-2026.tabulartext-retrieval1M<n<10M0 likes1.1k downloads6d agoHugging Face14masoudjs /c4-en-html-with-metadata-ppl-cleanFile list: "c4-en-html_cc-main-2019-18_pq00-000.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-001.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-002.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-003.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-004.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-005.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-006.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-007.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-008.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-009.jsonl.gz"… See the full description on the dataset page: https://huggingface.co/datasets/masoudjs/c4-en-html-with-metadata-ppl-clean.tabular10K<n<100K1 likes1.1k downloads4y agoHugging Face15bs-modeling-metadata /c4-en-html-with-training_metadata_alltabular10K<n<100K1 likes954 downloads3y agoHugging Face16brics-edtech /nornikel-metallurgy-vl-dataset Nornikel Metallurgy VL Dataset (SFT / DPO / GRPO) Датасет для дообучения мультимодальной модели Qwen3-VL по схеме SFT → DPO → GRPO в предметной области металлургии, горного дела и обогащения полезных ископаемых. Построен из корпуса технических документов (PDF-книги/сборники, DOCX-отчёты, PPTX-презентации, XLSX-таблицы) и изображений (схемы, диаграммы, таблицы). Конфигурации (config_name) config train validation назначение sft 111 351 12 372… See the full description on the dataset page: https://huggingface.co/datasets/brics-edtech/nornikel-metallurgy-vl-dataset.imagevisual-question-answering100K<n<1M0 likes704 downloads3mo agoHugging Face17AweAI-Team /AweAgent-Meta-NL2Repo AweAgent-Meta-NL2Repo This dataset provides the metadata used by AweAgent to run the NL2RepoBench evaluation. If you are looking for the underlying benchmark itself (task design, repositories, test suites), please refer to the original project: multimodal-art-projection/NL2RepoBench. Purpose The AweAgent repository evaluates end-to-end repo-level code generation: given a natural-language project specification, the agent must produce a working Python package that… See the full description on the dataset page: https://huggingface.co/datasets/AweAI-Team/AweAgent-Meta-NL2Repo.textn<1K0 likes654 downloads4mo agoHugging Face18jackkuo /arXiv-metadata-oai-snapshot About Dataset Dataset name: arXiv academic paper metadata Data source: https://arxiv.org/ Submission date: 1986-04-25 ~ 2025-05-13 (data updated weekly) Number of papers: 2,710,806 (as of 2025.5.14) Fields included: title, author, abstract, journal information, DOI, etc. Data format: json Data volume: 4.58G About ArXiv For nearly 30 years, ArXiv has served the public and research communities by providing open access to scholarly articles, from the vast branches of… See the full description on the dataset page: https://huggingface.co/datasets/jackkuo/arXiv-metadata-oai-snapshot.texttext-classification1M<n<10M0 likes578 downloads1y agoHugging Face19Complementarity /gpqa-metadata-blind-answertabularn<1K0 likes463 downloads1mo agoHugging Face20FINAL-Bench /Metacognitive FINAL Bench: Functional Metacognitive Reasoning Benchmark "Not how much AI knows — but whether it knows what it doesn't know, and can fix it." --- Overview FINAL Bench (Frontier Intelligence Nexus for AGI-Level Verification) is the first comprehensive benchmark for evaluating functional metacognition in Large Language Models (LLMs). Unlike existing benchmarks (MMLU, HumanEval, GPQA) that measure only final-answer accuracy, FINAL Bench evaluates… See the full description on the dataset page: https://huggingface.co/datasets/FINAL-Bench/Metacognitive.documenttext-generationn<1K106 likes453 downloads7mo agoHugging Face21GMLRVigil /BenchCheck-MetaBenchmark BenchCheck meta-benchmark (v6) 840 multiple-choice video questions from 76 public video benchmarks, one item per video, four capability groups x 210 (Perception, Temporal, Spatial / physical, Reasoning / knowledge). The set is the difficulty-first census of the BenchCheck screened pool: items that none of the cheap attackers of the pool screening (text-only, single frame, 32-frame 2B model, options-only, shuffled frames) could solve, ranked by the worst-case attacker percentile… See the full description on the dataset page: https://huggingface.co/datasets/GMLRVigil/BenchCheck-MetaBenchmark.tabularvideo-classification1K<n<10K0 likes383 downloads3d agoHugging Face22meta-math /MetaMathQA_GSM8K_zh Dataset MetaMathQA_GSM8K_zh is a dataset for mathematical reasoning in Chinese, question-answer pairs are translated from MetaMathQA (https://huggingface.co/datasets/meta-math/MetaMathQA) by GPT-3.5-Turbo with few-shot prompting. The dataset consists of 231685 samples. Citation If you find the GSM8K_zh dataset useful for your projects/papers, please cite the following paper. @article{yu2023metamath, title={MetaMath: Bootstrap Your Own Mathematical Questions for Large… See the full description on the dataset page: https://huggingface.co/datasets/meta-math/MetaMathQA_GSM8K_zh.textquestion-answering100K<n<1M17 likes368 downloads3y agoHugging Face23opendatalab /SlimPajama-Meta-rater-Readability-30B Top 30B token SlimPajama Subset selected by the Readability rater This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models. Code: https://github.com/opendatalab/Meta-rater Dataset Description This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Readability dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness) framework.… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Readability-30B.tabulartext-generation1M<n<10M1 likes361 downloads1y agoHugging Face24NuTonic /sat-bbox-metadata-sft-v1 Dataset Summary NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills). The goal is to create high-signal, production-shaped supervision for multimodal chat models: Captioning for satellite chips Grounding (bounding boxes in normalized coordinates) for land-cover regions Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-bbox-metadata-sft-v1.imagetext-generation100K<n<1M4 likes340 downloads5mo agoHugging Face25MirandaAbhilash /vqav2-full-metadatatabular100K<n<1M0 likes312 downloads6mo agoHugging Face26ttj /metadata_arxivtext1M<n<10M0 likes277 downloads5y agoHugging Face27abacusai /MetaMathFewshot A few-shot version of the MetaMath (https://huggingface.co/datasets/meta-math/MetaMathQA) dataset. Each entry is formatted with 'question' and 'answer' keys. The 'question' key has a random number of query-answer pairs between 0 and 4 inclusive, before a final target query; the expected answer to this is stored in the content of 'answer'. text100K<n<1M28 likes259 downloads3y agoHugging Face28jakegrigsby /metamon-parsed-replays Metamon Replay Dataset Pokémon Showdown replay files parsed (or "reconstructed") into RL trajectories by Metamon (arXiv Appendix D) Quick Start The easiest way to use the replay dataset is through metamon's dataloader: import metamon from metamon.interface import get_observation_space, get_reward_function, get_action_space from metamon.data importParsedReplayDataset # see the metamon README for more on observations, actions, and rewards. human_dset =… See the full description on the dataset page: https://huggingface.co/datasets/jakegrigsby/metamon-parsed-replays.textreinforcement-learningn<1K6 likes243 downloads4mo agoHugging Face29bglick13 /climbmix-400b-shuffle-metadatatextn<1K0 likes207 downloads6mo agoHugging Face30rafmacalaba /fcv-extractions-meta-tiered-probe fcv-extractions-meta-tiered-probe Data-use mention extractions from the World Bank Fragility, Conflict and Violence (FCV) document corpus. Spans are extracted by the fine-tuned GLiNER model rafmacalaba/gliner_datause_tiered, scored by the tier-probe head rafmacalaba/gliner-tier-probe, and attributed (provenance + usage/impact) by rafmacalaba/lfm2.5-350M-datause-multitask-tiered (a LoRA SFT of LiquidAI/LFM2.5-350M). Shape nested — one row per chunk; each entity in… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/fcv-extractions-meta-tiered-probe.tabular10K<n<100K0 likes203 downloads24d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.