CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jat-project /jat-dataset JAT Dataset Dataset Description The Jack of All Trades (JAT) dataset combines a wide range of individual datasets. It includes expert demonstrations by expert RL agents, image and caption pairs, textual data and more. The JAT dataset is part of the JAT project, which aims to build a multimodal generalist agent. Paper: https://huggingface.co/papers/2402.09844 Usage >>> from datasets import load_dataset >>> dataset =… See the full description on the dataset page: https://huggingface.co/datasets/jat-project/jat-dataset.imagereinforcement-learning100M<n<1B78 likes391k downloads3y agoHugging Face02google-research-datasets /natural_questions Dataset Card for Natural Questions Dataset Summary The NQ corpus contains questions from real users, and it requires QA systems to read and comprehend an entire Wikipedia article that may or may not contain the answer to the question. The inclusion of real user questions, and the requirement that solutions should read an entire page to find the answer, cause NQ to be a more realistic and challenging task than prior QA datasets. Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/natural_questions.textquestion-answering10K<n<100K127 likes80k downloads3y agoHugging Face03llamafactory /tiny-supervised-datasettexttext-generationn<1K4 likes45k downloads2y agoHugging Face04tensorshield /reddit_dataset_157 Bittensor Subnet 13 Reddit Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this dataset allows… See the full description on the dataset page: https://huggingface.co/datasets/tensorshield/reddit_dataset_157.texttext-classification10M<n<100M3 likes43k downloads1y agoHugging Face05google-research-datasets /nq_open Dataset Card for nq_open Dataset Summary The NQ-Open task, introduced by Lee et.al. 2019, is an open domain question answering benchmark that is derived from Natural Questions. The goal is to predict an English answer string for an input English question. All questions can be answered using the contents of English Wikipedia. Supported Tasks and Leaderboards Open Domain Question-Answering, EfficientQA Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/nq_open.textquestion-answering10K<n<100K36 likes33k downloads3y agoHugging Face06dataset-org /dreamDREAM is a multiple-choice Dialogue-based REAding comprehension exaMination dataset. In contrast to existing reading comprehension datasets, DREAM is the first to focus on in-depth multi-turn multi-party dialogue understanding.question-answering10K<n<100K9 likes21k downloads3y agoHugging Face07RUC-NLPIR /FlashRAG_datasets ⚡FlashRAG: A Python Toolkit for Efficient RAG Research FlashRAG is a Python toolkit for the reproduction and development of Retrieval Augmented Generation (RAG) research. Our toolkit includes 36 pre-processed benchmark RAG datasets and 16 state-of-the-art RAG algorithms. With FlashRAG and provided resources, you can effortlessly reproduce existing SOTA works in the RAG domain or implement your custom RAG processes and components. For more information, please view our GitHub repo… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/FlashRAG_datasets.textquestion-answering1M<n<10M94 likes18k downloads1y agoHugging Face08Zenos5 /mse-text-img-dataset Dataset Card for MSE-text-img-dataset We have created a custom dataset that is extracted as a subset of the Math Stack Exchange (MSE) dataset. This text-image dataset contains 64,860 questions with their respective list of answers, scores, acceptance marking, and image versions of each question and answer generated from the stored text with embedded LaTeX math markup. In this dataset there are 117,380 answers in total, with 1.81 answers per question on average. Each image… See the full description on the dataset page: https://huggingface.co/datasets/Zenos5/mse-text-img-dataset.question-answering0 likes15k downloads1y agoHugging Face09google-research-datasets /tydiqa Dataset Card for "tydiqa" Dataset Summary TyDi QA is a question answering dataset covering 11 typologically diverse languages with 204K question-answer pairs. The languages of TyDi QA are diverse with regard to their typology -- the set of linguistic features that each language expresses -- such that we expect models performing well on this set to generalize across a large number of the languages in the world. It contains language phenomena that would not be found in… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/tydiqa.textquestion-answering100K<n<1M38 likes15k downloads2y agoHugging Face10qyang1021 /AIR-Bench-Dataset AIR-Bench Arxiv: https://arxiv.org/html/2402.07729v1This is the AIR-Bench dataset download page.AIR-Bench encompasses two dimensions: foundation and chat benchmarks. The former consists of 19 tasks with approximately 19k single-choice questions. The latter one contains 2k instances of open-ended question-and-answer data.For how to run AIR-Bench, Please refer to AIR-Bench github page(https://github.com/OFA-Sys/AIR-Bench)(will be public soon). Data Sources(All come from… See the full description on the dataset page: https://huggingface.co/datasets/qyang1021/AIR-Bench-Dataset.audioquestion-answeringn<1K8 likes11k downloads2y agoHugging Face11lavita /medical-qa-datasets all-processed dataset is a concatenation of of medical-meadow-* and chatdoctor_healthcaremagic datasets The Chat Doctor term is replaced by the chatbot term in the chatdoctor_healthcaremagic dataset Similar to the literature the medical_meadow_cord19 dataset is subsampled to 50,000 samples truthful-qa-* is a benchmark dataset for evaluating the truthfulness of models in text generation, which is used in Llama 2 paper. Within this dataset, there are 55 and 16 questions related to Health and… See the full description on the dataset page: https://huggingface.co/datasets/lavita/medical-qa-datasets.textquestion-answering1M<n<10M64 likes9.9k downloads3y agoHugging Face12agungpambudi /math-dataset-measuring-mathematical-problem-solvingTo cite the dataset please reference it as @article{hendrycksmath2021, title={Measuring Mathematical Problem Solving With the MATH Dataset}, author={Dan Hendrycks and Collin Burns and Saurav Kadavath and Akul Arora and Steven Basart and Eric Tang and Dawn Song and Jacob Steinhardt}, journal={NeurIPS}, year={2021} } textquestion-answering100K<n<1M1 likes9.1k downloads1y agoHugging Face13bitext /Bitext-customer-support-llm-chatbot-training-dataset Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.textquestion-answering10K<n<100K195 likes7.9k downloads2y agoHugging Face14johanneskirmayr /car-bench-dataset CAR-Bench Dataset CAR-Bench is a benchmark for evaluating AI voice assistants in a realistic automotive (car) environment. It tests an agent's ability to correctly use vehicle control tools, handle disambiguation, and avoid hallucinations. Dataset Structure The dataset is organized into task configs and mock data configs: Tasks Each task defines a user persona, an instruction, the initial vehicle/environment context, and the ground-truth sequence of tool-call… See the full description on the dataset page: https://huggingface.co/datasets/johanneskirmayr/car-bench-dataset.tabulartext-generation1M<n<10M3 likes7.5k downloads7mo agoHugging Face15futuremoon /x_dataset_39 Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/futuremoon/x_dataset_39.texttext-classification1B<n<10B2 likes6k downloads1y agoHugging Face16Trendyol /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K134 likes5.6k downloads1y agoHugging Face17nvidia /Nemotron-AIQ-Agentic-Safety-Dataset-1.0 Nemotron-AIQ Agentic Safety Dataset Dataset Summary Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.texttext-generation10K<n<100K18 likes4.7k downloads10mo agoHugging Face18akshaydudhane /EarthDial-Dataset 🌍 EarthDial-Dataset The EarthDial-Dataset is a curated collection of evaluation-only datasets focused on remote sensing and Earth observation downstream tasks. It is designed to benchmark vision-language models (VLMs) and multimodal reasoning systems on real-world scenarios involving satellite and aerial imagery. 📚 Key Features Evaluation-focused: All datasets are for inference/testing only — no train/val splits. Diverse Tasks: Classification Object Detection Change… See the full description on the dataset page: https://huggingface.co/datasets/akshaydudhane/EarthDial-Dataset.imagequestion-answering10K<n<100K7 likes4.6k downloads9mo agoHugging Face19amolharsh /Ground3D_Dataset Ground3D Dataset A large-scale 3D vision-language question-answering dataset for point-grounded, metric-aware 3D scene understanding. Built on ScanNet and ScanNet++ with dense object and part annotations, the dataset spans eight downstream reasoning tasks at both object and part granularity, plus multi-turn dialogue that composes them. Answers are both: point-grounded: explicitly tied to the referred 3D region via <p>label</p><SEG> markup, and metric: physical quantities (size… See the full description on the dataset page: https://huggingface.co/datasets/amolharsh/Ground3D_Dataset.textvisual-question-answering1M<n<10M1 likes4.3k downloads2mo agoHugging Face20SHSLab /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/SHSLab/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.tabulartext-generation10M<n<100M3 likes4.2k downloads26d agoHugging Face21rag-datasets /rag-mini-wikipediaIn this huggingface discussion you can share what you used the dataset for. Derives from https://www.kaggle.com/datasets/rtatman/questionanswer-dataset?resource=download we generated our own subset using generate.py. textquestion-answering1K<n<10K55 likes3.5k downloads2y agoHugging Face22StampyAI /alignment-research-datasetThe AI Alignment Research Dataset is a collection of documents related to AI Alignment and Safety from various books, research papers, and alignment related blog posts.question-answering10K<n<100K17 likes3k downloads3y agoHugging Face23cl-nagoya /ruri-dataset-reranker Ruri-Dataset Reranker Datasets used for training Ruri-Reranker. Please refer to https://huggingface.co/datasets/hpprc/emb for individual datasets. textquestion-answering1M<n<10M5 likes2.9k downloads2y agoHugging Face24ishumilin /epstein-files-ocr-datasets-1-8-early-release Epstein Files OCR — Datasets 1–8 (Early Release) ARCHIVE NOTICE This dataset is no longer maintained. Please refer to the Epstein Files — Complete OCR Dataset. Dataset Summary This dataset contains page-level OCR output (as Markdown) from a public release of documents related to Jeffrey Epstein / the Epstein case. Each Markdown file represents one scanned page converted to text using an automated OCR pipeline. The dataset is designed for: Question answering Information… See the full description on the dataset page: https://huggingface.co/datasets/ishumilin/epstein-files-ocr-datasets-1-8-early-release.question-answering10K<n<100K1 likes2.8k downloads6mo agoHugging Face25Qalam /nuclear-intelligence-dataset Nuclear Intelligence Dataset Public, auto-generated dataset of validated nuclear-energy research cycles. Latest stats (auto-updated): 🪙 NES tokens minted: 0 ⛓️ Blockchain length: 1 blocks 🕸️ Knowledge entities: 2 Source GitHub: https://github.com/QalamHipHop/nuclear-intelligence HF Space: https://huggingface.co/spaces/Qalam/Nuclear-Intelligence License MIT tabularquestion-answeringn<1K1 likes2.6k downloads13m agoHugging Face26coldmind /reddit_dataset_94 Bittensor Subnet 13 Reddit Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this dataset allows… See the full description on the dataset page: https://huggingface.co/datasets/coldmind/reddit_dataset_94.texttext-classification10M<n<100M0 likes2.6k downloads1y agoHugging Face27minkyuchoi /Temporal-Logic-Video-Dataset Temporal Logic Video (TLV) Dataset Temporal Logic Video (TLV) Dataset Synthetic and real video dataset with temporal logic annotation Explore the GitHub » NSVS-TL Project Webpage · NSVS-TL Source Code Overview The Temporal Logic Video (TLV) Dataset addresses the scarcity of state-of-the-art video datasets for long-horizon, temporally extended activity and object detection. It comprises two main components: Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/minkyuchoi/Temporal-Logic-Video-Dataset.tabularquestion-answeringn<1K1 likes2.4k downloads2y agoHugging Face28JiaaqiLiu /SkillArena-datasets SkillArena Offline Datasets Offline evaluation data for SkillArena — a validated automatic benchmark generation framework for AI agent skills, targeting NeurIPS 2026 Datasets & Benchmarks Track. Overview This dataset provides domain-specific task input data for 289 AI agent skills across 13 domains. Each skill has 50 curated data files designed as meaningful agent task inputs — files an agent could receive and act upon (analyze, transform, validate, or generate from). The… See the full description on the dataset page: https://huggingface.co/datasets/JiaaqiLiu/SkillArena-datasets.text-generation10K<n<100K0 likes2.4k downloads7mo agoHugging Face29M-A-D /Mixed-Arabic-Datasets-Repo Dataset Card for "Mixed Arabic Datasets (MAD) Corpus" The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts Dataset Description The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With MAD, we… See the full description on the dataset page: https://huggingface.co/datasets/M-A-D/Mixed-Arabic-Datasets-Repo.tabulartext-classification100M<n<1B38 likes2.2k downloads3y agoHugging Face30elmoghany /Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text Dataset Overview A collection of 27 domains (“topics”) and 3100 question-answer pair. Each topic comes with average 117 QA pairs.Every QA entry comes with: references: one or more source files the answer is extracted from time with each reference comes the starting and ending time the answer is extracted from the reference video_files: the video files where the answer can be found (future) video title & description from metadata.csv File structure You-Are-Here!/… See the full description on the dataset page: https://huggingface.co/datasets/elmoghany/Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text.question-answering1K<n<10K3 likes2.2k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.