CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PaDaS-Lab /webfaqWebFAQ Q&A Dataset Overview | Details | Structure | Examples | Considerations | License | Citation | Contact | Acknowledgement Overview The WebFAQ Q&A Dataset is a broad-coverage corpus of 96 million natural question-answer (QA) pairs in 75 languages, gathered from FAQ pages on the web. It leverages structured schema.org FAQPage annotations, making it a unique resource for large-scale Question Answering… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/webfaq.textquestion-answering10M<n<100M24 likes1.4k downloads1y agoHugging Face02TIGER-Lab /SWE-QA-Pro-Bench SWE-QA-Pro Bench (A Repository-level QA Benchmark Built from Diverse Long-tail Repositories) 💻 GitHub | 📖 Paper | 🤗 SWE-QA-Pro 📢 News 🚀 [2026-5-19] The evaluation code is released on GitHub. 🔥 [2026-3-23] SWE-QA-Pro Bench is publicly released! The model and code will be released soon. Introduction SWE-QA-Pro Bench is a repository-level question answering dataset designed to evaluate whether models can perform grounded, agentic reasoning… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/SWE-QA-Pro-Bench.textquestion-answeringn<1K5 likes678 downloads4mo agoHugging Face03ASLP-lab /MSU-Benchmark MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios Interspeech 2026 · ASLP@NPU (Northwestern Polytechnical University), in collaboration with Li Auto. Zhaokai Sun*, Shuai Wang*, Zhennan Lin*, Chengyou Wang, Dehui Gao, Yuang Cao, Chunjiang He, Pan Zhou, Lei Xie** Audio, Speech and Language Processing Group (ASLP@NPU), School of Software, Northwestern Polytechnical University, China School of Intelligent Science and Technology, Nanjing… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/MSU-Benchmark.audioaudio-classification1K<n<10K1 likes622 downloads3mo agoHugging Face04TIGER-Lab /AIME25The AIME25 part 1 exam from the website. textquestion-answeringn<1K2 likes583 downloads2y agoHugging Face05LCM-Lab /LOOMBench 🔬 LOOMBench: Long-Context Language Model Evaluation Benchmark 🎯 Framework Overview LOOMBench is a streamlined evaluation suite derived from our comprehensive long-context evaluation framework. It represents the gold standard for efficient long-context language model assessment. ✨ Key Highlights 📊 16 Diverse Benchmarks: Carefully curated from extensive benchmark collections. ⚡ Efficient Evaluation: Optimized for unified loading and evaluation. 🎯… See the full description on the dataset page: https://huggingface.co/datasets/LCM-Lab/LOOMBench.tabularquestion-answering1K<n<10K0 likes436 downloads8mo agoHugging Face06Eureka-Lab /PHYBench PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models [🌐 Project] [📄 Paper] [💻 Code] [🏆 Leaderboard] [🌟 Overview] [🔧 Data Details] [🚩 Citation] New Updates 2025.4.25: We release our code of EED Score. View and star on our github page! 2025.5.15: We have significantly improved the paper and experiments, including diversified experimental discussions and in-depth error analysis. The updated website is now live at… See the full description on the dataset page: https://huggingface.co/datasets/Eureka-Lab/PHYBench.textquestion-answering1K<n<10K62 likes361 downloads1y agoHugging Face07Oduwo /drug_label_approved_openfda KEMIRIX OpenFDA Clinical Drug Dataset Built for KEMIRIX — Africa's first Clinical Decision Support AI Developer: Emmanuel Bain Oduwo | TechFryz Ltd. | Nairobi, Kenya Generated: May 2026 Configurations clean (default): instruction + output only, fully cleaned, ready for fine-tuning Kemirix raw: full metadata schema, original generated data Usage from datasets import load_dataset # Clean data for training Kemirix ds =… See the full description on the dataset page: https://huggingface.co/datasets/Oduwo/drug_label_approved_openfda.texttext-generation10K<n<100K0 likes330 downloads4mo agoHugging Face08li-lab /HealMed HealMed (Human-verified Evaluation Across Languages for Medical AI) is a multilingual medical dataset featuring expert-verified translations for benchmarking multilingual medical AI systems. The dataset comprises translations from two complementary sources. A portion is based on the multilingual translations released by the GlobMed project (arXiv: 2601.02186), while the remainder was generated by our team using zero-shot machine translation to expand language coverage. Each translated… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/HealMed.textquestion-answering10K<n<100K2 likes300 downloads9d agoHugging Face09declare-lab /Trust-Data Dataset Card for Trust framework Description Repository: https://github.com/declare-lab/trust-align Paper: https://arxiv.org/abs/2409.11242 Data Summary The Trust-score evaluation dataset includes the top 100 GTR-retrieved results for ASQA, QAMPARI, and ExpertQA, along with the top 100 BM25-retrieved results for ELI5. The answerability of each question is assessed based on its accompanying documents. The Trust-align training dataset comprises 19K high-quality… See the full description on the dataset page: https://huggingface.co/datasets/declare-lab/Trust-Data.textquestion-answering10K<n<100K1 likes285 downloads1y agoHugging Face10declare-lab /rq-bench RQ-Bench: A Benchmark for Grounded Research Question Generation RQ-Bench evaluates whether language models can read background literature and propose the same kinds of research questions that a human author actually went on to investigate. Each example pairs a held-out research question (RQ) — distilled from a real arXiv paper (the target paper) — with the full text of the prior-work papers that the target paper cites as motivation. A model is shown only the cited references and… See the full description on the dataset page: https://huggingface.co/datasets/declare-lab/rq-bench.texttext-generation1K<n<10K0 likes263 downloads4mo agoHugging Face11datapizza-ai-lab /dnd5e-srd-qa D&D 5.2.1 SRD RAG Evaluation Dataset A high-quality Question-Answering (QA) dataset built by the Datapizza AI Lab from the Dungeons & Dragons 5th Edition System Reference Document (SRD) version 5.2.1, designed to evaluate Retrieval Augmented Generation (RAG) systems. Dataset Summary This dataset contains 56 question-answer pairs across two difficulty tiers (Easy and Medium), each designed to test different aspects of RAG system capabilities. The dataset is built from 20… See the full description on the dataset page: https://huggingface.co/datasets/datapizza-ai-lab/dnd5e-srd-qa.texttext-classificationn<1K11 likes241 downloads9mo agoHugging Face12sorika-labs /swen-1-data Swen-1 Dataset Conversational and mathematical reasoning data collected by Sorika Labs from Swen Dual-Engine (Swen-1.1-Instruct & Swen-1-Math). texttext-generation10K<n<100K0 likes208 downloads2d agoHugging Face13humanfia-lab /IPHO2026 IPhO 2026 Curated Problems This repository packages the official English problem, solution, and marking materials for the LVI International Physics Olympiad (Bucaramanga, Colombia, 2026) as machine-readable, subquestion-level records. Contents Configuration Rows Description all 41 All curated subquestions theory 23 Theory papers T1–T3 experiment 18 Experimental paper E1 formalization_ready 29 Subset selected for theorem formalization… See the full description on the dataset page: https://huggingface.co/datasets/humanfia-lab/IPHO2026.imagequestion-answeringn<1K0 likes204 downloads2mo agoHugging Face14qualora-data-labs /qualora-workforce-skills-graph Qualora Workforce Skills Graph (Representative Sample) Rights-clean, provenance-tracked vocational learning data, rebuilt from roughly $2B of U.S. Department of Labor funded open courseware into a labeled skills graph: cleaned courses and lessons, Bloom-tagged assessment items with answer rationales and learning objectives, and a content-grounded course to skill to career graph with salary context. Built for post-training and evaluation, not pretraining bulk. This repository is… See the full description on the dataset page: https://huggingface.co/datasets/qualora-data-labs/qualora-workforce-skills-graph.tabularquestion-answeringn<1K0 likes195 downloads2mo agoHugging Face15LARK-Lab /EnvFactory-SFT-FILTERED EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL ## Overview EnvFactory-SFT-FILTERED is a filtered supervised fine-tuning (SFT) dataset containing 53,400 tool-use trajectories synthesized using the EnvFactory framework. This dataset is designed for SFT training of tool-use agents. The dataset contains high-quality multi-turn tool-use trajectories with implicit human reasoning, generated through… See the full description on the dataset page: https://huggingface.co/datasets/LARK-Lab/EnvFactory-SFT-FILTERED.texttext-generation10K<n<100K0 likes183 downloads4mo agoHugging Face16Metavolve-Labs /supervision-tradeoff The Supervision Tradeoff — Reproducibility Bundle Format Scaffolds, Judgment Pleasing, and Anti-Calibration in Post-Training Paper DOI: 10.5281/zenodo.19748277 · Concept DOI: 10.5281/zenodo.19748276 · Code repo: github.com/codex-curator/supervision-tradeoff Author: Tad MacPherson, Metavolve Labs · ORCID: 0009-0002-8659-7479 What this is and why it might help your research This repository ships everything we used to falsify our own headline finding, in a form you can… See the full description on the dataset page: https://huggingface.co/datasets/Metavolve-Labs/supervision-tradeoff.texttext-generation10K<n<100K0 likes135 downloads5mo agoHugging Face17TIGER-Lab /Fineweb-InstructWe convert the pre-training corpus from Fineweb-Edu (https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu) to instruction following format. We select a subset with quality filter and then use GPT-4 to extract instruction-following pairs. The dataset contains roughly 16M instruction pairs. The basic concept is similar to MAmmoTH2 (https://arxiv.org/abs/2405.03548). Citation If you use dataset useful, please cite the following paper: @article{yue2024mammoth2… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/Fineweb-Instruct.textquestion-answering10M<n<100M9 likes133 downloads2y agoHugging Face18TIGER-Lab /SWE-QA-Pro-SFT-Trajectories SWE-QA-Pro SFT Trajectories 💻 GitHub | 📖 Paper | 🤗 SWE-QA-Pro Introduction SWE-QA-Pro SFT Trajectories is a set of agentic tool-use trajectories for repository-level question answering, used as the supervised fine-tuning (SFT) data in the SWE-QA-Pro training recipe. Each item is a multi-turn trajectory in which an agent answers a repository-grounded question by exploring the codebase with read-only tools rather than relying on memorized knowledge. The… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/SWE-QA-Pro-SFT-Trajectories.textquestion-answering1K<n<10K0 likes130 downloads3mo agoHugging Face19supreme-lab /HALT_Benchmark_0.1_v1 HALT Benchmark Dataset v1.0 HALT: Benchmarking When Language Agents Should Stop, Investigate, Escalate, or Refuse Overview HALT is a benchmark for evaluating bounded agentic decision-making under partial observability, constrained tools, and explicit escalation options. It is grounded in defensive cybersecurity workflows, where acting too early, failing to escalate, or over-escalating can all be costly. The benchmark contains 1,248 instances across four decision regimes… See the full description on the dataset page: https://huggingface.co/datasets/supreme-lab/HALT_Benchmark_0.1_v1.texttext-classification1K<n<10K1 likes115 downloads5mo agoHugging Face20LARK-Lab /EnvFactory-SFT-ALL EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL ## Overview EnvFactory-SFT-ALL is the complete supervised fine-tuning (SFT) dataset containing 26,500 tool-use trajectories synthesized using the EnvFactory framework. This dataset includes all generated trajectories before filtering. The dataset contains multi-turn tool-use trajectories with implicit human reasoning, generated through… See the full description on the dataset page: https://huggingface.co/datasets/LARK-Lab/EnvFactory-SFT-ALL.texttext-generation10K<n<100K0 likes114 downloads4mo agoHugging Face21lablup /tau2-bench-ko τ²-bench 한국어 번역 공개판 τ²-bench v0.2.0의 retail·airline·telecom 사용자 시나리오를 한국어로 번역 원본: sierra-research/tau2-bench v0.2.0 번역 및 검수 모델: gpt-5.6-sol, 일부 모호한 문장은 사람이 검수 번역 범위: 278 task, 1,004 user-scenario field 파일 파일 행 수 설명 data/retail.jsonl 114 유통 에이전트 태스크 data/airline.jsonl 50 항공 에이전트 태스크 data/telecom.jsonl 114 통신 에이전트 태스크 metadata.json - 생성 설정·파일 SHA-256 build_dataset.py - 원본 task에 번역 필드를 결합해 jsonl 생성하는 스크립트 LICENSE - 원본 τ²-bench의 MIT License 사본… See the full description on the dataset page: https://huggingface.co/datasets/lablup/tau2-bench-ko.texttext-generationn<1K1 likes101 downloads24d agoHugging Face22declare-lab /GSM8k_MOREDataset introduced in the paper: Evaluating LLMs' Mathematical Competency through Ontology-guided Perturbations. This dataset was created by randomly sampling five questions from GSM8K and perturbing them using an ontology. textquestion-answeringn<1K3 likes100 downloads2y agoHugging Face23lab-flair /qa-dataset-k1000 QA Dataset K1000 — The First Drop of Ink Question-answering data with gold documents and distractor pools for long-context evaluation, accompanying The First Drop of Ink: Nonlinear Impact of Distracting Information in Long-Context Reasoning by Muhan Gao, Zih-Ching Chen, and Kuan-Hao Huang (ICML 2026). Paper · Full text (v2) · Hugging Face paper page The paper studies how the proportion of hard distractors affects performance at fixed context length. It reports a nonlinear… See the full description on the dataset page: https://huggingface.co/datasets/lab-flair/qa-dataset-k1000.textquestion-answeringn<1K1 likes98 downloads13h agoHugging Face24jang1563 /LabCraft-Eval LabCraft-Eval LabCraft-Eval is an Inspect AI evaluation environment for measuring how well AI agents execute benign molecular-microbiology protocols inside a seeded laboratory simulator with task-dependent stochasticity. It pairs task prompts and tool-accessible lab operations with deterministic, multi-axis trajectory scoring. This Hugging Face dataset export is generated from the GitHub repository: https://github.com/jang1563/LabCraft-Eval.git Release Release… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/LabCraft-Eval.texttext-generationn<1K0 likes90 downloads15d agoHugging Face25PKU-DS-LAB /AlgGeoTest Welcome to AlgGeoTest created by PKU-DS-LAB! Citation Information Paper Link: https://arxiv.org/abs/2508.02208 Dataset Description AlgGeoTest is the first benchmark specifically designed to evaluate LLMs' comprehension of Algebraic Geometry—a frontier domain of modern mathematics that occupies a central position within the contemporary mathematical landscape. AlgGeoTest was created by implementing Proof2Hybrid—the first fully-automated framework for… See the full description on the dataset page: https://huggingface.co/datasets/PKU-DS-LAB/AlgGeoTest.textquestion-answeringn<1K1 likes89 downloads1y agoHugging Face26Layered-Labs /claude-fable-derm Claude Fable Derm Claude Fable Derm is a dataset of patient dermatology questions and their raw, unprompted responses from Claude Fable. Each question is generated by crossing a real dermatological topic with one of Paul Ekman's six basic emotions, producing emotionally distinct framings of the same underlying medical concern. Answers are collected with no system prompt or role instruction, capturing how the model responds to a patient question exactly as it would in the wild.… See the full description on the dataset page: https://huggingface.co/datasets/Layered-Labs/claude-fable-derm.textquestion-answeringn<1K0 likes88 downloads12d agoHugging Face27LabMem012 /LogiOR Overview LogiOR, a comprehensive benchmark dataset comprising 92 logistics and supply chain optimization problems, which was developed over two months under the guidance of three Operations Research (OR) experts. The problems are adapted from classical OR solver test datasets, textbook examples, research papers, and real-world applications. LogiOR covers a broad spectrum of optimization types including Linear Programming (LP), Integer Linear Programming (ILP), Mixed-Integer… See the full description on the dataset page: https://huggingface.co/datasets/LabMem012/LogiOR.textquestion-answeringn<1K3 likes81 downloads3mo agoHugging Face28iLearn-Lab /FineBadmintonBenchmark FineBadmintonBenchmark Fine-grained badminton video question answering benchmark. Dataset Structure hf_video_clips_qa/: one clip per QA item, named by video_uid (for example video_000001.mp4). finebadmintonbenchmark/: annotation JSON files. each item contains video_uid each QA item maps to exactly one video clip through video_uid Citation @inproceedings{he2025finebadminton, title={Finebadminton: A multi-level dataset for fine-grained badminton video… See the full description on the dataset page: https://huggingface.co/datasets/iLearn-Lab/FineBadmintonBenchmark.textquestion-answering1K<n<10K0 likes78 downloads7mo agoHugging Face29umd-zhou-lab /AVQA-Audio-Rubrics AVQA Audio-Reasoning Rubrics Project Page | Paper | Code Audio-grounded, binary-evaluable evaluation rubrics for the full AVQA training set, generated for process-level reward modeling in audio reasoning RL (e.g. GRPO / RLHF with rubric-as-reward). Each training question is annotated with 5 rubrics, one per evaluation facet, that judge the quality of an audio-reasoning response — not just final answer correctness. The rubrics are designed to be scored Yes/No by an LLM judge that… See the full description on the dataset page: https://huggingface.co/datasets/umd-zhou-lab/AVQA-Audio-Rubrics.textaudio-classification10K<n<100K1 likes71 downloads2mo agoHugging Face30Kenotic-Labs /ATANTV1.0-corpus ATANT Narrative Test Corpus Automated Test for Acceptance of Narrative Truth, v1.0 The first open evaluation corpus for measuring continuity in AI systems: the ability to persist, update, disambiguate, and reconstruct meaningful context across time. Paper: ATANT: An Evaluation Framework for AI Continuity (arXiv:2604.06710) Standard repository: github.com/Kenotic-Labs/ATANT Author: Samuel Sameer Tanguturi Affiliation: Kenotic Labs Published: April 2026 Why this corpus… See the full description on the dataset page: https://huggingface.co/datasets/Kenotic-Labs/ATANTV1.0-corpus.tabularquestion-answeringn<1K0 likes70 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.