CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nyu-visionx /Cambrian-10M Cambrian-10M Dataset Please see paper & website for more information: https://cambrian-mllm.github.io/ https://arxiv.org/abs/2406.16860 Overview Cambrian-10M is a comprehensive dataset designed for instruction tuning, particularly in multimodal settings involving visual interaction data. The dataset is crafted to address the scarcity of high-quality multimodal instruction-tuning data and to maintain the language abilities of multimodal large language models (LLMs).… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/Cambrian-10M.visual-question-answering1M<n<10M131 likes18k downloads2y agoHugging Face02qyang1021 /AIR-Bench-Dataset AIR-Bench Arxiv: https://arxiv.org/html/2402.07729v1This is the AIR-Bench dataset download page.AIR-Bench encompasses two dimensions: foundation and chat benchmarks. The former consists of 19 tasks with approximately 19k single-choice questions. The latter one contains 2k instances of open-ended question-and-answer data.For how to run AIR-Bench, Please refer to AIR-Bench github page(https://github.com/OFA-Sys/AIR-Bench)(will be public soon). Data Sources(All come from… See the full description on the dataset page: https://huggingface.co/datasets/qyang1021/AIR-Bench-Dataset.audioquestion-answeringn<1K8 likes11k downloads2y agoHugging Face03prquan /STARK_10k STARK: Spatial-Temporal reAsoning benchmaRK STARK is a comprehensive benchmark designed to systematically evaluate large language models (LLMs) and large reasoning models (LRMs) on spatial-temporal reasoning tasks, particularly for applications in cyber-physical systems (CPS) such as robotics, autonomous vehicles, and smart city infrastructure. Dataset Summary Hierarchical Benchmark: Tasks are structured across three levels of reasoning complexity: State Estimation:… See the full description on the dataset page: https://huggingface.co/datasets/prquan/STARK_10k.textquestion-answering10K<n<100K1 likes6.4k downloads1y agoHugging Face04zai-org /LongAlign-10k LongAlign-10k 🤗 [LongAlign Dataset] • 💻 [Github Repo] • 📃 [LongAlign Paper] LongAlign is the first full recipe for LLM alignment on long context. We propose the LongAlign-10k dataset, containing 10,000 long instruction data of 8k-64k in length. We investigate on trianing strategies, namely packing (with loss weighting) and sorted batching, which are all implemented in our code. For real-world long context evaluation, we introduce LongBench-Chat that evaluate the… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/LongAlign-10k.textquestion-answering1K<n<10K100 likes5.8k downloads3y agoHugging Face05Dongping-Li /EMMOE-100 EMMOE-100 Trainset Resources Project Paper Code Model Dataset Dataset Feature Task Attributes Task Example Dataset Structure EMMOE-100/ ├── README.md ├── assets/ ├── data/ │ └── train/ │ ├── 1/ │ │ ├── info.txt │ │ ├── info_re1.txt │ │ ├── info_re2.txt │ │ ├── info_re3.txt │ │ ├── keypath.json │ │ ├── scene.json │ │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/Dongping-Li/EMMOE-100.imagevisual-question-answering10K<n<100K1 likes5.5k downloads1y agoHugging Face06Pn101 /taxbench-au TaxBench-AU A benchmark for testing whether AI agents can calculate Australian tax. TaxBench-AU contains 156 Australian tax calculation questions, presented as multiple-choice (4-option) worked tax problems. The benchmark is designed to test whether an AI agent can read the facts, apply the right Australian tax rule for the relevant income year, do the calculation, and choose the correct answer. The Kaggle mirror is published as Agent Tax Exam for Australian Tax. Paper:… See the full description on the dataset page: https://huggingface.co/datasets/Pn101/taxbench-au.documentquestion-answeringn<1K0 likes4.8k downloads4mo agoHugging Face07KodCode /KodCode-Light-RL-10K 🐱 KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding KodCode is the largest fully-synthetic open-source dataset providing verifiable solutions and tests for coding tasks. It contains 12 distinct subsets spanning various domains (from algorithmic to package-specific knowledge) and difficulty levels (from basic coding exercises to interview and competitive programming challenges). KodCode is designed for both supervised fine-tuning (SFT) and RL tuning. 🕸️… See the full description on the dataset page: https://huggingface.co/datasets/KodCode/KodCode-Light-RL-10K.tabularquestion-answering10K<n<100K9 likes4.7k downloads1y agoHugging Face08MoreThought /Fable-5.1-Max-Reasoning-Filtered-10000x Dataset Description This dataset contains 10,000 agentic coding and reasoning multi-turn traces generated by the new Fable 5.1 model using max reasoning effort. It holds almost 500,000,000 tokens of step-by-step chain-of-thought programming across multiple complex domains. It has also been deduplicated and heavily filtered to remove low-quality traces, keeping only high-quality traces. Dataset Statistics Metric Value Total Examples 10,000 Traces… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/Fable-5.1-Max-Reasoning-Filtered-10000x.text-generation10K<n<100K147 likes2.6k downloads1h agoHugging Face09SamuelChien821 /counselbench-100 CounselBench-100 CounselBench-100 v3.2.5 is a synthetic legal-work benchmark with 100 authored matters across ten practice workflows. Every task has a natural employee request, a 97-asset evidence room, twelve portfolio decisions, 5–9 supported actions, 3–7 evidence holds, and a distinct deep multi-provider MCP trajectory. The answer is not preclassified in the evidence. Each portfolio item requires an immutable identity join, an operative-authority and revision lookup, a… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/counselbench-100.documentquestion-answeringn<1K0 likes2.6k downloads25d agoHugging Face10SamuelChien821 /factorybench-100 FactoryBench-100 FactoryBench-100 is a 100-task benchmark for employee-grade manufacturing and ERP decisions. Each public prompt is a short, high-level employee request; it does not name the systems, files, API calls, answer schema, or execution order. The isolated SQLite world exposes documented Oracle Fusion Cloud 26a REST operations alongside Gmail v1, Drive v3, Sheets v4, and Slack Web API operations over synthetic state. Harbor runs the authoritative SQLite state and trace… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/factorybench-100.documentquestion-answeringn<1K0 likes1.9k downloads23d agoHugging Face11SamuelChien821 /salesbench-100 SalesBench-100 SalesBench-100 is a synthetic long-horizon sales-agent benchmark with 100 original workflows across Salesforce, HubSpot, Gong, and a seeded evidence room. Each task begins with a high-level employee request and has its own authored causal rule and provider transition. Identity, operating facts, authority, governed policy, live-system indexes, and exceptions are separated so no mounted business asset publishes a selected option or precomputed change. Every task has… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/salesbench-100.documenttext-generationn<1K1 likes1.8k downloads25d agoHugging Face12starriver030515 /FUSION-Pretrain-10M FUSION-10M Dataset Please see paper & website for more information: https://arxiv.org/abs/2504.09925 https://github.com/starriver030515/FUSION Overview FUSION-10M is a large-scale, high-quality dataset of image-caption pairs used to pretrain FUSION-3B and FUSION-8B models. It builds upon established datasets such as LLaVA, ShareGPT4, and PixelProse. In addition, we synthesize 2 million task-specific image-caption pairs to further enrich the dataset. The goal of… See the full description on the dataset page: https://huggingface.co/datasets/starriver030515/FUSION-Pretrain-10M.imagequestion-answeringn<1K9 likes1.4k downloads1y agoHugging Face13hasankursun /turkish-corpus-100b Turkish Corpus 100B (TC-100B) Dataset Summary The Turkish Corpus 100B (TC-100B) is a massive-scale, deduplicated, and cleaned dataset designed for training Foundation Models in Turkish. Comprising approximately 105 Billion tokens (measured with Qwen/Llama3 tokenizer), it represents one of the largest open resources for Turkish LLM pretraining. The dataset is engineered for a two-stage training pipeline: Pretrain Subset (~103B Tokens): A diverse mix of… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/turkish-corpus-100b.texttext-generation100M<n<1B8 likes857 downloads3mo agoHugging Face14Thunderbolt215215 /ArtiMuse-10K ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding [🌐 Project Page] [🚀 Online Demo] [💻 Code] [📄 Paper] [[🧩 Checkpoints: 🤗 Hugging Face | 🤖 ModelScope]] 🌟 Building upon on ArtiMuse, we introduce UniPercept, a comprehensive follow-up work that provides a meticulous study on perceptual-level image understanding. It spans Image Aesthetics Assessment (IAA), Image Quality Assessment (IQA), and Image Structure & Texture… See the full description on the dataset page: https://huggingface.co/datasets/Thunderbolt215215/ArtiMuse-10K.imagequestion-answering1K<n<10K4 likes844 downloads8mo agoHugging Face15miracle10 /EarthVerse Benchmarking scientific agents across dynamic Earth systems and natural hazards Zhiqing Cui1, Xinxiang Yin2, Yihong Tang3, Xinglang Zhang4, Yuanzhe Hu5, Siru Zhong4, Weidong Tang6, Yuxuan Liang4, Weijia Li7, Ming Jin8, Shirui Pan8, Yuhao Kang9, Dingyi Zhuang10,†, Jinhua Zhao10 1NUIST &nbsp; 2HKU &nbsp; 3McGill &nbsp; 4HKUST(GZ) &nbsp; 5Georgia Tech &nbsp; 6NUS &nbsp; 7Tsinghua &nbsp; 8Griffith &nbsp; 9UT Austin &nbsp; 10MIT &nbsp; †Corresponding author Project page ·… See the full description on the dataset page: https://huggingface.co/datasets/miracle10/EarthVerse.textquestion-answeringn<1K1 likes835 downloads29d agoHugging Face16stindardlogic /math-reasoning-sft-100k Math Reasoning SFT (100K) 100,000 math problems with detailed step-by-step solutions — ready for supervised fine-tuning of math reasoning models. Dataset Description 100,000 problems across 8 mathematical categories and 3 difficulty levels: Categories Category Examples Topics word_problems ~23,100 Rate/time/distance, work problems, mixture, meeting/catch-up arithmetic ~15,400 Percentages, profit/loss, ratios geometry ~15,400 Area… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/math-reasoning-sft-100k.texttext-generation100K<n<1M1 likes652 downloads2mo agoHugging Face17gk4u /reddit_dataset_104 Bittensor Subnet 13 Reddit Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this dataset allows… See the full description on the dataset page: https://huggingface.co/datasets/gk4u/reddit_dataset_104.texttext-classification100M<n<1B4 likes637 downloads1y agoHugging Face18kaptaan45 /KapInstruct-100M KapInstruct-100M: Curated 100-Million Token Instruction Tuning Dataset KapInstruct-100M is a high-fidelity, 100-million-token instruction-tuning dataset engineered for Supervised Fine-Tuning (SFT) and alignment of compact language models (under 1 billion parameters). Formatted with the Qwen ChatML chat template and tokenized using Qwen/Qwen3.5-0.8B-Base, the dataset enforces strict assistant-only loss masking (masking user prompts and structural delimiters to -100)… See the full description on the dataset page: https://huggingface.co/datasets/kaptaan45/KapInstruct-100M.question-answering10K<n<100K0 likes630 downloads1mo agoHugging Face19MCES10-Software /Python-Code-Solutions Python Code Solutions Features 1000k of Python Code Solutions for Text Generation and Question Answering Python Coding Problems labelled by topic and difficulty Recommendations Train your Model on Logical Operations and Mathematical Problems Before Training it on this. This is optional for Fine Tuning 2B parameter + models. Format the prompts in a orderly way when formatting data eg. {question} Solution: {solution} Topic: {topic} textquestion-answering10K<n<100K0 likes624 downloads1y agoHugging Face20lfaviate /China-K12-STEM-10K-CoT-Reasoning K12-STEM-CoT-Chinese 1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams. The largest structured Chinese math/physics/chemistry reasoning dataset. This is a curated sample (10,000 problems) of the full 1.54M dataset available via API. Full Dataset Access Access the full 1,540,000+ problems via API → This Sample Full API Total problems 10,025 1,540,000+ With CoT solutions 10,025 1,490,000+ With diagrams 6,093 740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.tabularquestion-answering10K<n<100K3 likes605 downloads7mo agoHugging Face21SamuelChien821 /dealbench-100 DealBench-100 DealBench-100 is a 100-task, deterministic investment-banking agent benchmark over ten synthetic transaction worlds. It tests source control, QoE normalization, trading comps, precedents, DCF, LBO, merger math, bid comparison, model-to-deck consistency, and launch approval. Run harbor run -d blobfishai/dealbench-100-suite -a <agent> -m <provider/model> Metric The single metric is DealScore (0–100): discovery 15, model accuracy 25… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/dealbench-100.question-answering0 likes578 downloads23d agoHugging Face22renzzyyy1028 /civil-code-phil Civilex — Philippine Legal RAG & SFT Dataset Retrieval corpus and supervised fine-tuning (SFT) data for a retrieval-augmented generation (RAG) pipeline over Philippine law: the Civil Code (Republic Act No. 386) and Supreme Court jurisprudence. Produced by the civilex-thesis research pipeline. Contents: 11k+ Supreme Court jurisprudence cases spanning 1949–2025, and 2,270 articles from the Civil Code (Republic Act No. 386). Dataset structure . ├── README.md ├──… See the full description on the dataset page: https://huggingface.co/datasets/renzzyyy1028/civil-code-phil.textquestion-answering100K<n<1M1 likes550 downloads1d agoHugging Face23ianncity /KIMI-K2.5-1000000x KIMI-K2.5-1000000x 1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions) Distribution: Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#) Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl Computer Science: 5% Logical Questions: 5% Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/ianncity/KIMI-K2.5-1000000x.texttext-generation100K<n<1M265 likes549 downloads6mo agoHugging Face24momo1942 /x_dataset_10290 Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/momo1942/x_dataset_10290.texttext-classification100M<n<1B0 likes548 downloads1y agoHugging Face25albertoRodriguez97 /history-anchor-100 History Anchor 100 *The benchmark behind the paper "History Anchors: How Prior Behavior Steers LLM Decisions Toward Unsafe Actions".* 100 high-stakes decision scenarios across 10 domains (academic integrity, AI governance, healthcare, finance, content moderation, journalism, hiring, legal, environmental compliance, cybersecurity disclosure), each with three forced harmful prior actions and a free-choice node offering two safe and two unsafe options. Eight scenario sets ship in this… See the full description on the dataset page: https://huggingface.co/datasets/albertoRodriguez97/history-anchor-100.texttext-generationn<1K0 likes514 downloads4mo agoHugging Face26byrLLCC /MedSP1000 MedSP1000 Evaluating Large Language Models in Dynamic Clinical Decision-Making with Standardized Patient Cases Paper: arxiv.org/abs/2606.05112 &nbsp;|&nbsp; Code: github.com/MAGIC-AI4Med/MedSP1000 Dataset Summary MedSP1000 is a standardized-patient (SP)–derived interactive benchmark for evaluating large language models as clinical agents. Unlike static, single-turn medical QA, each item is an executable multi-turn encounter: a clinician agent… See the full description on the dataset page: https://huggingface.co/datasets/byrLLCC/MedSP1000.text-generation1K<n<10K2 likes514 downloads4mo agoHugging Face27mast-benchmark /100k-corpus-2026 MAST 100K Corpus 2026 This dataset contains the fixed English document corpus used for MAST @ FIRE 2026, the Multilingual Agentic Search Track. MAST evaluates whether multilingual agentic search systems can answer complex questions posed in different languages by retrieving English evidence and producing short, correct English answers. This corpus is copied from BrowseComp-Plus, a benchmark for Deep-Research systems that isolates the effect of the retriever and the LLM agent to… See the full description on the dataset page: https://huggingface.co/datasets/mast-benchmark/100k-corpus-2026.textquestion-answering100K<n<1M0 likes498 downloads2mo agoHugging Face28hulk10 /conseil-detat-full-documents Décisions du Conseil d'État (France) Description Ce dataset contient un corpus de décisions rendues par le Conseil d'État français, la plus haute juridiction de l'ordre administratif. Les décisions proviennent de la plateforme officielle Open Data de la Justice Administrative et sont diffusées au format XML anonymisé. Le corpus rassemble les textes intégraux des décisions ainsi que plusieurs métadonnées permettant leur identification et leur traçabilité. Source… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/conseil-detat-full-documents.tabularquestion-answering1K<n<10K1 likes488 downloads7h agoHugging Face29hulk10 /conseil-administratives-appel-full-documents Décisions de Justice Administrative Françaises Description du dataset Ce dataset contient un corpus de décisions de justice administrative françaises extraites de la plateforme officielle Open Data de la Justice Administrative. Les documents sont publiés par le Conseil d'État dans le cadre de la politique d'ouverture des données publiques de la justice française. Les décisions sont diffusées au format XML et anonymisées conformément aux exigences légales relatives… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/conseil-administratives-appel-full-documents.texttext-classification1K<n<10K1 likes485 downloads7h agoHugging Face30RL-MIND /UHR-BAT-SFT-10K UHR-BAT-SFT-10K Supervised Fine-Tuning for Ultra-High-Resolution Remote Sensing Project · Paper · Code English | 中文 📚 Introduction UHR-BAT-SFT-10K contains visual question answering style instruction-following examples for ultra-high-resolution remote-sensing imagery. It is the supervised fine-tuning dataset used for UHR-BAT: Budget-Aware Token Compression Vision-Language Model for Ultra-High-Resolution Remote Sensing.… See the full description on the dataset page: https://huggingface.co/datasets/RL-MIND/UHR-BAT-SFT-10K.geospatialvisual-question-answering10K<n<100K2 likes483 downloads8d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.