CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bigcode /the-stack-metadata Dataset Card for The Stack Metadata Changelog Release Description v1.1 This is the first release of the metadata. It is for The Stack v1.1 v1.2 Metadata dataset matching The Stack v1.2 Dataset Summary This is a set of additional information for repositories used for The Stack. It contains file paths, detected licenes as well as some other information for the repositories. Supported Tasks and Leaderboards The main… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-metadata.tabulartext-generation10B<n<100B10 likes4.9k downloads4y agoHugging Face02librarian-bots /arxiv-metadata-snapshot Dataset Card for "arxiv-metadata-oai-snapshot" More Information needed This is a mirror of the metadata portion of the arXiv dataset. The sync will take place weekly so may fall behind the original datasets slightly if there are more regular updates to the source dataset. Metadata This dataset is a mirror of the original ArXiv data. This dataset contains an entry for each paper, containing: id: ArXiv ID (can be used to access the paper, see below) submitter:… See the full description on the dataset page: https://huggingface.co/datasets/librarian-bots/arxiv-metadata-snapshot.texttext-generation1M<n<10M22 likes4.7k downloads5d agoHugging Face03Metaskepsis /Olympiads Numina-Olympiads Filtered NuminaMath-CoT dataset containing only olympiads problems with valid answers. Dataset Information Split: train Original size: 32926 Filtered size: 32926 Source: olympiads All examples contain valid boxed answers Dataset Description This dataset is a filtered version of the NuminaMath-CoT dataset, containing only problems from olympiad sources that have valid boxed answers. Each example includes: A mathematical word problem A… See the full description on the dataset page: https://huggingface.co/datasets/Metaskepsis/Olympiads.texttext-generation10K<n<100K3 likes973 downloads2y agoHugging Face04aisingapore /NLU-Metaphorgated SEA Metaphor SEA Metaphor evaluates a model's ability to interpret paired figurative phrases with divergent meanings. It is sampled from Multilingual-Fig-QA for Indonesian, Javanese, and Sundanese. Supported Tasks and Leaderboards SEA Metaphor is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore. Languages Indonesian (id) Javanese (jv) Sundanese (su) Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLU-Metaphor.texttext-generation1K<n<10K0 likes727 downloads9mo agoHugging Face05oumi-ai /MetaMathQA-R1 oumi-ai/MetaMathQA-R1 MetaMathQA-R1 is a text dataset designed to train Conversational Language Models with DeepSeek-R1 level reasoning. Prompts were augmented from GSM8K and MATH training sets with responses directly from DeepSeek-R1. MetaMathQA-R1 was used to train MiniMath-R1-1.5B, which achieves 44.4% accuracy on MMLU-Pro-Math, the highest of any model with <=1.5B parameters. Curated by: Oumi AI using Oumi inference on Parasail Language(s) (NLP): English License:… See the full description on the dataset page: https://huggingface.co/datasets/oumi-ai/MetaMathQA-R1.texttext-generation100K<n<1M7 likes658 downloads2y agoHugging Face06ScriptSmith /sponsorblock-youtube-metadata-2024 SponsorBlock YouTube Metadata Dataset A dataset of YouTube video metadata collected from a subset of videos in the SponsorBlock database. This dataset contains metadata, subtitles, engagement heatmaps, live chat, and channel playlist information for popular YouTube videos. Contains the top videos from the SponsorBlock database that had data added in the year 2024. Quick Stats Metric Value Total videos 154,536 Videos with subtitles 62,819 (41%)… See the full description on the dataset page: https://huggingface.co/datasets/ScriptSmith/sponsorblock-youtube-metadata-2024.imagetext-classification10M<n<100M0 likes397 downloads2mo agoHugging Face07Rurouni-II /arxiv-metadata-snapshot Dataset Card for "arxiv-metadata-oai-snapshot" More Information needed This is a mirror of the metadata portion of the arXiv dataset. The sync will take place weekly so may fall behind the original datasets slightly if there are more regular updates to the source dataset. Metadata This dataset is a mirror of the original ArXiv data. This dataset contains an entry for each paper, containing: id: ArXiv ID (can be used to access the paper, see below) submitter: Who… See the full description on the dataset page: https://huggingface.co/datasets/Rurouni-II/arxiv-metadata-snapshot.texttext-generation1M<n<10M0 likes333 downloads5mo agoHugging Face08Sharathhebbar24 /MetaMathQA Meta Math Filtered This is a combined and filtered (removed all the redundant rows) version of meta-math/MetaMathQA and meta-math/MetaMathQA-40K Usage from datasets import load_dataset dataset = load_dataset("Sharathhebbar24/MetaMathQA", split="train") texttext-generation100K<n<1M0 likes280 downloads3y agoHugging Face09ShuoZheLi /MetaMathQA-math-500DAPO-Math-17k with the MATH-500 test split converted to the same parquet schema and prompt format. texttext-generation100K<n<1M1 likes238 downloads2mo agoHugging Face10THUIR /MetaSyn MetaSyn MetaSyn is a benchmark for protocol-driven scientific evidence synthesis. It contains 422 Nature Portfolio source reviews and a shared corpus of 140,585 PubMed articles, with 336 training and 86 test instances. Resources Paper: arXiv:2606.17041 Code and evaluator: THUIR/MetaSyn Trained retriever: BFTree/MA-Retriever Configurations reviews contains source-review records, PI/ECO fields, search and eligibility information, synthesis… See the full description on the dataset page: https://huggingface.co/datasets/THUIR/MetaSyn.tabulartext-retrieval100K<n<1M2 likes226 downloads2mo agoHugging Face11freococo /thahabiorg_metadata 📖 Thahabi Books Metadata Dataset This repository contains structured metadata for 28,896 Arabic books scraped from thahabi.org. Each row represents one book and includes bibliographic information such as title, author, category, and source details. 📦 Dataset Structure This repository contains structured metadata for 28,896 Arabic books scraped from thahabi.org. Each row represents one book with full bibliographic and structural information. 📚… See the full description on the dataset page: https://huggingface.co/datasets/freococo/thahabiorg_metadata.tabulartext-generation1M<n<10M0 likes160 downloads3mo agoHugging Face12WhissleAI /daily_dialog_meta Meta-LLM Dataset: Daily Dialog with Meta-Information Enhancement Dataset Overview This dataset contains 76,064 conversational examples from the Daily Dialog corpus enhanced with meta-information awareness. Each example includes three response types: original human responses, basic LLM responses, and meta-aware LLM responses that incorporate emotional and intentional context. Meta-Information Distribution Emotion Categories Emotion Count… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/daily_dialog_meta.texttext-generation10K<n<100K1 likes153 downloads1y agoHugging Face13mireklzicar /cellarc_100k_meta cellarc_100k_meta CellARC 100k Meta is the metadata‑rich variant of the CellARC benchmark introduced in Lzicar, M. (2025). CellARC: Measuring Intelligence with Cellular Automata. It contains the exact same episodes and splits as cellarc_100k, with byte‑identical Parquet files; the JSONL files retain full per‑episode metadata (rule tables, coverage diagnostics, morphology descriptors, sampling parameters, etc.). Each episode exposes five support pairs plus a held‑out query/solution… See the full description on the dataset page: https://huggingface.co/datasets/mireklzicar/cellarc_100k_meta.textother10K<n<100K1 likes114 downloads11mo agoHugging Face14Metaskepsis /Olympiads_hard Numina-Olympiads Filtered NuminaMath-CoT dataset containing only olympiads problems with valid answers. Dataset Information Split: train Original size: 21525 Filtered size: 21408 Source: olympiads All examples contain valid boxed answers Dataset Description This dataset is a filtered version of the NuminaMath-CoT dataset, containing only problems from olympiad sources that have valid boxed answers. Each example includes: A mathematical word problem A… See the full description on the dataset page: https://huggingface.co/datasets/Metaskepsis/Olympiads_hard.tabulartext-generation10K<n<100K3 likes105 downloads2y agoHugging Face15Metaskepsis /Olympiads_medium Numina-Olympiads Filtered NuminaMath-CoT dataset containing only olympiads problems with valid answers. Dataset Information Split: train Original size: 13284 Filtered size: 13240 Source: olympiads All examples contain valid boxed answers Dataset Description This dataset is a filtered version of the NuminaMath-CoT dataset, containing only problems from olympiad sources that have valid boxed answers. Each example includes: A mathematical word problem A… See the full description on the dataset page: https://huggingface.co/datasets/Metaskepsis/Olympiads_medium.tabulartext-generation10K<n<100K1 likes93 downloads2y agoHugging Face16SnowCharmQ /DPL-meta Difference-aware Personalized Learning (DPL) Dataset This dataset is used in the paper: Measuring What Makes You Unique: Difference-Aware User Modeling for Enhancing LLM Personalization Yilun Qiu, Xiaoyan Zhao, Yang Zhang, Yimeng Bai, Wenjie Wang, Hong Cheng, Fuli Feng, Tat-Seng Chua Code: https://github.com/SnowCharmQ/DPL This dataset is an adaptation of the Amazon Reviews'23 dataset. It contains item metadata for Books, CDs & Vinyl, and Movies & TV. Each item includes title… See the full description on the dataset page: https://huggingface.co/datasets/SnowCharmQ/DPL-meta.texttext-generation1K<n<10K1 likes87 downloads1y agoHugging Face17Lots-of-LoRAs /task1394_meta_woz_task_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1394_meta_woz_task_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1394_meta_woz_task_classification.texttext-generationn<1K0 likes85 downloads2y agoHugging Face18Metaskepsis /Numina_medium Numina-Olympiads Filtered NuminaMath-CoT dataset containing only olympiads problems with valid answers. Dataset Information Split: train Original size: 37133 Filtered size: 37133 Source: olympiads All examples contain valid boxed answers Dataset Description This dataset is a filtered version of the NuminaMath-CoT dataset, containing only problems from olympiad sources that have valid boxed answers. Each example includes: A mathematical word problem A… See the full description on the dataset page: https://huggingface.co/datasets/Metaskepsis/Numina_medium.tabulartext-generation10K<n<100K0 likes83 downloads2y agoHugging Face19KBlueLeaf /danbooru2023-metadata-databasegated Metadata Database for Danbooru2023 Danbooru 2023 datasets: https://huggingface.co/datasets/nyanko7/danbooru2023 The latest entry of this database is id 7,866,491. Which is newer than nyanko7's dataset. This dataset contains a sqlite db file which have all the tags and posts metadata in it. The Peewee ORM config file is provided too, plz check it for more information. (Especially on how I link posts and tags together) The original data is from the official dump of the posts info.… See the full description on the dataset page: https://huggingface.co/datasets/KBlueLeaf/danbooru2023-metadata-database.imageimage-classification1M<n<10M83 likes65 downloads2y agoHugging Face20mtybilly /MetaMathQA MetaMathQA Subsets Curated subsets of meta-math/MetaMathQA for mathematical reasoning experiments. Subsets Subset Samples Description full 395,000 All MetaMathQA samples (unchanged) MATH 155,000 MATH_* types only (AnsAug, Rephrased, FOBAR, SV) MATH-50K 50,000 Stratified 50K sample from MATH subset MATH-50K Type Distribution Type Count Proportion MATH_AnsAug 24,194 48.4% MATH_Rephrased 16,129 32.3% MATH_FOBAR 4,839 9.7%… See the full description on the dataset page: https://huggingface.co/datasets/mtybilly/MetaMathQA.texttext-generation100K<n<1M0 likes61 downloads6mo agoHugging Face21birgermoell /oellm-dpo-metadataset OpenEuroLLM DPO metadataset A lightweight, versioned source of truth for building preference-training data for OpenEuroLLM. It contains metadata and planning decisions—not copies of upstream training examples. The catalogue pins each upstream revision and records its license, size, language coverage, pair schema, overlap family, decision, risks, and required transformations. Upstream licenses and terms still apply. The Apache-2.0 license in this repository covers only the… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/oellm-dpo-metadataset.tabulartext-generationn<1K0 likes47 downloads12d agoHugging Face22uuuue /DPL-meta Difference-aware Personalized Learning (DPL) Dataset This dataset is used in the paper: Measuring What Makes You Unique: Difference-Aware User Modeling for Enhancing LLM Personalization Yilun Qiu, Xiaoyan Zhao, Yang Zhang, Yimeng Bai, Wenjie Wang, Hong Cheng, Fuli Feng, Tat-Seng Chua Code: https://github.com/SnowCharmQ/DPL This dataset is an adaptation of the Amazon Reviews'23 dataset. It contains item metadata for Books, CDs & Vinyl, and Movies & TV. Each item includes… See the full description on the dataset page: https://huggingface.co/datasets/uuuue/DPL-meta.texttext-generation1K<n<10K0 likes46 downloads8d agoHugging Face23Metaskepsis /Numina Numina-Olympiads Filtered NuminaMath-CoT dataset containing only olympiads problems with valid answers. Dataset Information Split: train Original size: 210350 Filtered size: 210350 Source: olympiads All examples contain valid boxed answers Dataset Description This dataset is a filtered version of the NuminaMath-CoT dataset, containing only problems from olympiad sources that have valid boxed answers. Each example includes: A mathematical word problem A… See the full description on the dataset page: https://huggingface.co/datasets/Metaskepsis/Numina.texttext-generation100K<n<1M1 likes43 downloads2y agoHugging Face24phanerozoic /Metamath Metamath A structured dataset of formally verified theorems and axioms from Metamath, one of the largest collections of rigorously verified mathematics in the world. Source Repository: https://github.com/metamath/set.mm Commit: 160dfc7e4ec5f201f5bae4ca5a5eeb67242902b5 Files: 5 License: cc0-1.0 Schema Column Type Description statement string Declaration signature/claim with the leading keyword removed (verbatim slice); the full… See the full description on the dataset page: https://huggingface.co/datasets/phanerozoic/Metamath.texttext-generation10K<n<100K0 likes43 downloads4mo agoHugging Face25phanerozoic /Coq-MetaCoq Coq-MetaCoq Structured dataset of formalizations from MetaCoq (Coq meta-theory formalized in Coq). Source Repository: https://github.com/MetaCoq/metacoq Commit: 971b2cc5c8bdc011068f67fa272245d8e5b54209 Files: 597 License: mit Schema Column Type Description statement string Declaration signature/claim with the leading keyword removed (verbatim slice); the full declaration minus its proof proof string Verbatim proof/body, empty if the… See the full description on the dataset page: https://huggingface.co/datasets/phanerozoic/Coq-MetaCoq.texttext-generation10K<n<100K0 likes35 downloads4mo agoHugging Face26lamhieu /math_metaqa_dialogue_en Description The dataset is from unknown, formatted as dialogues for speed and ease of use. Many thanks to author for releasing it. Importantly, this format is easy to use via the default chat template of transformers, meaning you can use huggingface/alignment-handbook immediately, unsloth. Structure View online through viewer. Note We advise you to reconsider before use, thank you. If you find it useful, please like and follow this account.… See the full description on the dataset page: https://huggingface.co/datasets/lamhieu/math_metaqa_dialogue_en.texttext-generation10K<n<100K0 likes30 downloads2y agoHugging Face27alvarobartt /openhermes-preferences-metamath Dataset Card for OpenHermes Preferences - MetaMath This dataset is a subset from argilla/OpenHermesPreferences, only keeping the preferences of metamath, and removing all the columns besides the chosen and rejected ones, that come in OpenAI chat formatting, so that's easier to fine-tune a model using tools like: huggingface/alignment-handbook or axolotl, among others. Reference argilla/OpenHermesPreferences dataset created as a collaborative effort between Argilla and… See the full description on the dataset page: https://huggingface.co/datasets/alvarobartt/openhermes-preferences-metamath.texttext-generation10K<n<100K4 likes28 downloads3y agoHugging Face28introspector /meta-meme Meta-Meme Consultation URLs Dataset Description This dataset contains 2177 consultation URLs generated from the Meta-Meme formally verified system. Each URL represents a consultation with one of 9 AI muses about a specific file in the repository. Dataset Structure file: Path to the file in the repository muse: Assigned AI muse (Calliope, Clio, Erato, Euterpe, Melpomene, Polyhymnia, Terpsichore, Thalia, Urania) tool: Consultation tool (llm, lean4, rustc… See the full description on the dataset page: https://huggingface.co/datasets/introspector/meta-meme.texttext-generation1K<n<10K1 likes25 downloads8mo agoHugging Face29phanerozoic /Coq-Metalib Coq-Metalib Structured declarations from Metalib - a Coq library for programming language metatheory using locally nameless representation. Source: github.com/plclub/metalib Source Repository: https://github.com/plclub/metalib Commit: 144ddcd0fff6717229140314cf559d85fad6eae0 Files: 33 License: mit Schema Column Type Description statement string Declaration signature/claim with the leading keyword removed (verbatim slice); the full… See the full description on the dataset page: https://huggingface.co/datasets/phanerozoic/Coq-Metalib.texttext-generation1K<n<10K0 likes24 downloads4mo agoHugging Face30zcamz /ai-vs-human-meta-llama-Llama-3.2-1B-Instruct AI vs Human dataset on the CNN Daily mails Dataset Description This dataset showcases pairs of truncated articles and their respective completions, crafted either by humans or an AI language model. Each article was randomly truncated between 25% and 50% of its length. The language model was then tasked with generating a completion that mirrored the characters count of the original human-written continuation. Data Fields 'human': The original human-authored… See the full description on the dataset page: https://huggingface.co/datasets/zcamz/ai-vs-human-meta-llama-Llama-3.2-1B-Instruct.texttext-classification1K<n<10K1 likes20 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.