CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01worstchan /UltraChat-300K-SLAM-Omni UltraChat-300K This dataset is prepared for the reproduction of SLAM-Omni. This is a multi-round English spoken dialogue training dataset. For code and usage examples, please refer to the related GitHub repository: X-LANCE/SLAM-LLM (examples/s2s) 🔧 Modifications Data Filtering: We removed samples with excessively long data. Speech Response Tokens: We used CosyVoice to synthesize corresponding semantic speech tokens for the speech response. These tokens, represented as… See the full description on the dataset page: https://huggingface.co/datasets/worstchan/UltraChat-300K-SLAM-Omni.tabularquestion-answering100K<n<1M2 likes3.8k downloads1y agoHugging Face02bowen-upenn /PersonaMem-v3 PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks Bowen Jiang, Yuan Yuan, Zhuoqun Hao, Yuchen Liu, Maohao Shen, Sihao Chen, Gregory Wornell, Chris Callison-Burch, Lyle Ungar, Dan Roth, Qi Guo, Xiangjun Fan, Camillo J. Taylor, Hanchao Yu A collaboration between: Meta Recommendation Systems University of Pennsylvania MIT Third release in the PersonaMem series: PersonaMem-v1: [COLM… See the full description on the dataset page: https://huggingface.co/datasets/bowen-upenn/PersonaMem-v3.tabularquestion-answering100K<n<1M3 likes2.5k downloads28d agoHugging Face03stdKonjac /LiveSports-3K LiveSports-3K Benchmark News [2025.05.12] We released the ASR transcripts for the CC track. See LiveSports-3K-CC.json for details. Overview LiveSports‑3K is a comprehensive benchmark for evaluating streaming video understanding capabilities of large language and multimodal models. It consists of two evaluation tracks: Closed Captions (CC) Track: Measures models’ ability to generate real‑time commentary aligned with the ground‑truth ASR transcripts. Question… See the full description on the dataset page: https://huggingface.co/datasets/stdKonjac/LiveSports-3K.tabularvideo-text-to-text1K<n<10K5 likes1.5k downloads1y agoHugging Face04BTF-2 /BTF-3 Bench to the Future 3 (BTF-3) 1,783 pastcasting questions — 1,471 binary ("yes/no") and 312 numeric (value-estimation) — with a state-of-the-art forecast, ground-truth resolution, and a human-verifiable resolution explanation for every question. Designed for reproducible evaluation of forecasting agents without hindsight bias or web-data leakage. BTF-3 questions were anchored to a present date in late April–late May 2026 and resolved between mid-May and early July 2026. It… See the full description on the dataset page: https://huggingface.co/datasets/BTF-2/BTF-3.tabularquestion-answering1M<n<10M2 likes1.1k downloads6d agoHugging Face05bicycleman15 /ruler-300-seed42 Frozen RULER 300, seed 42 This dataset freezes the exact RULER inputs used by the short-long-pretraining native evaluation suite. Repository: bicycleman15/ruler-300-seed42 Rows: 6,300 Tasks: s-niah-1, s-niah-2, s-niah-3, mk1, mk2, mv, mq Context lengths: 1024, 2048, 4096 Samples per task/length: 300 Seed: 42 Dataset SHA-256: 4d82df6f9b1f2d9c45c0a0bda8c734032e62f517b746c6351bf9c2f38335ab3d Tokenizer SHA-256: 1f186971e25f7bda3dd6f93a100bb8fa2a6801cf8dc3807c8a8c4e45f296ab90… See the full description on the dataset page: https://huggingface.co/datasets/bicycleman15/ruler-300-seed42.tabularquestion-answering1K<n<10K0 likes876 downloads1mo agoHugging Face06r0b0tlab /qwen3.8-max-distillation-50k Qwen3.8-Max Distillation 50K A curated dataset of 49,772 teacher-generated traces from qwen3.8-max-preview, prepared for supervised fine-tuning and off-policy knowledge distillation. The teacher responses are preserved as returned by the API. Where the model emitted visible <think>...</think> blocks, those blocks remain in the assistant message. Some simpler prompts received direct answers without a thinking block. [!CAUTION] Terms and provenance notice — not cleared for… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/qwen3.8-max-distillation-50k.tabulartext-generation10K<n<100K118 likes783 downloads2mo agoHugging Face07piercewetter3 /irs-990-parsed IRS 990 Parsed Nonprofit Database Public relational extract of IRS Form 990 / 990-EZ / 990-PF filings, plus the colocated public files we join for address research: CMS NPPES + T-MSIS Medicare spend, FMCSA DOT carriers, OFAC SDN, FEC committees, and the IRS EO BMF. Generated: 2026-08-17Tables: 34Rows (sum): 459,069,505License: CC0 / public domain — derived from U.S. government recordsHub: https://huggingface.co/datasets/piercewetter3/irs-990-parsed Layout Tables… See the full description on the dataset page: https://huggingface.co/datasets/piercewetter3/irs-990-parsed.tabulartabular-classification100M<n<1B0 likes759 downloads1mo agoHugging Face08yjlee36 /knowchat-multi-turn-dialogues KnowChat: Multi-Turn Human-LLM Dialogues on Knowledge Tasks KnowChat is a dataset of 705 multi-turn human-LLM conversations collected to validate the KnowSim user simulation framework. It pairs each conversation with pre/post knowledge assessments, self-reported survey ratings, and participant background information, enabling research on information calibration -- how well LLM assistants tailor responses to users with different knowledge levels. Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/yjlee36/knowchat-multi-turn-dialogues.tabularquestion-answeringn<1K3 likes447 downloads1mo agoHugging Face09MercanAI /turkce-sft-qa-3.7m 🇹🇷 Turkish SFT/QA — Birleştirilmiş ve Tekrarsız Veri Seti 3,723,264 örnek. 24 açık Türkçe SFT/QA veri setinin, satır düzeyinde tekrar temizliği ve kalite kontrolünden geçirilmiş birleşimi. Her satır hangi veri setinden geldiğini taşır. English: A merged, row-level deduplicated and quality-filtered collection of 24 open Turkish SFT/QA datasets (3,723,264 examples). Every row carries its source dataset, source URL and original license. 🙏 Teşekkür /… See the full description on the dataset page: https://huggingface.co/datasets/MercanAI/turkce-sft-qa-3.7m.tabulartext-generation1M<n<10M0 likes370 downloads2mo agoHugging Face10bevangelista /AIME_2000_2026_Kimi_K3 AIME 2000–2026 — Kimi K3 reasoning traces 🔄 Changelog 2026-08-08 — full re-generation. All reasoning traces were regenerated from scratch and re-verified against the official answer key. New schema — added gen_attempts_low, gen_attempts_high; renamed gen_parsed_answer → gen_answer_int and answer_note → problem_note; removed gen_effort, gen_pass1. New generation — only use the bare problem (v1 appended an "ANSWER:" format instruction), so traces are cleaner.… See the full description on the dataset page: https://huggingface.co/datasets/bevangelista/AIME_2000_2026_Kimi_K3.tabulartext-generationn<1K2 likes351 downloads2mo agoHugging Face11R-3-Bench /R-3-Bench&nbsp;R3-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets Overview R3-Bench evaluates resource-rational reasoning when multiple problems share a limited resource budget. This release contains the frozen benchmark data used by the paper across three domains: Config Problems Suites Difficulty counts math 300 50 150 easy / 100 medium / 50 hard coding 300 50 150 easy / 100 medium / 50 hard abstract_reasoning 300 50 150 easy /… See the full description on the dataset page: https://huggingface.co/datasets/R-3-Bench/R-3-Bench.tabularquestion-answeringn<1K6 likes349 downloads24d agoHugging Face12youdotcom /minimax-m3-deepsearchqa-skill-eval MiniMax M3 DeepSearchQA Skill Eval Evaluates minimax/minimax-m3 on google/deepsearchqa using a Pi agent, You.com MCP tools, and a research skill optimized for this harness, model, and tool surface. MiniMax M3 Medium Reasoning with the You.com research skill reached 74.85% adjusted F1 on DeepSearchQA, above the paper's GPT-5 High Reasoning F1 result. Public artifacts are available for inspection and reproduction. Links GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/youdotcom/minimax-m3-deepsearchqa-skill-eval.tabularquestion-answering1K<n<10K1 likes306 downloads16d agoHugging Face13mwei /UltraChat-300K-SLAM-Omni UltraChat-300K This dataset is prepared for the reproduction of SLAM-Omni. This is a multi-round English spoken dialogue training dataset. For code and usage examples, please refer to the related GitHub repository: X-LANCE/SLAM-LLM (examples/s2s) 🔧 Modifications Data Filtering: We removed samples with excessively long data. Speech Response Tokens: We used CosyVoice to synthesize corresponding semantic speech tokens for the speech response. These tokens, represented as… See the full description on the dataset page: https://huggingface.co/datasets/mwei/UltraChat-300K-SLAM-Omni.tabularquestion-answering100K<n<1M0 likes211 downloads8mo agoHugging Face14ssubhnil /Qwen3-Math-Eval Qwen3 Math Evaluation Suite Greedy (temperature 0) outputs from Qwen3 1.7B / 4B / 8B / 14B on nine math-reasoning benchmarks across output-token budgets {2k, 4k, 8k, 16k, 32k}. 1,417,388 predictions over 186 model-by-dataset-by-budget cells, each with the full reasoning trace, the extracted answer, and strict and answer-forced correctness labels. On standard MATH (MATH-500, Hendrycks MATH test, competition MATH) these models are saturated at 16k: the 4B is at or above 0.94 and… See the full description on the dataset page: https://huggingface.co/datasets/ssubhnil/Qwen3-Math-Eval.tabularquestion-answering1M<n<10M0 likes206 downloads4mo agoHugging Face15jwu323 /FlowBench FlowBench Dataset ID: jwu323/FlowBench FlowBench is a tool-use benchmark for deterministic business operations workflows. Each task asks an agent to compose Python tools over synthetic customers, products, orders, returns, inventory, support tickets, FX rates, and SLA policies. Scope note: this dataset is a business-operations tool-composition benchmark. It is unrelated to prior workflow-guided planning or workflow-generation benchmarks that also use the FlowBench name. This… See the full description on the dataset page: https://huggingface.co/datasets/jwu323/FlowBench.tabularquestion-answeringn<1K2 likes179 downloads3mo agoHugging Face16eousphoros /2d_3d_seq_path_spatial_reasoning Spatial Reasoning Dataset A synthetic dataset of Hamiltonian path puzzles with rich chain-of-thought reasoning, designed for training and evaluating spatial reasoning in language models. Overview Each sample presents a grid-based puzzle where the solver must find a path visiting every cell exactly once, moving only up/down/left/right (plus above/below for 3D). Puzzles span 2D grids (3x3 to 8x8) and 3D cubes (3x3x3 to 4x4x4), covering solvable, impossible, and multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/eousphoros/2d_3d_seq_path_spatial_reasoning.tabularquestion-answering1K<n<10K0 likes178 downloads8mo agoHugging Face17chrislimbe /pubmedqa-recursive-llm-degradation-qwen2.5-3b PubMedQA Recursive LLM Degradation — Qwen2.5-3B This repository contains synthetic biomedical question-answering data and model predictions generated as part of a study of recursive fine-tuning and model degradation. Base Model Qwen/Qwen2.5-3B Source Dataset The experiments use the PubMedQA dataset: qiaoxin/PubMedQA This repository contains generated/derived research artifacts and does not redistribute the original PubMedQA dataset in its entirety.… See the full description on the dataset page: https://huggingface.co/datasets/chrislimbe/pubmedqa-recursive-llm-degradation-qwen2.5-3b.tabularquestion-answering10K<n<100K0 likes147 downloads3d agoHugging Face18xmanii /Maux-Persian-SFT-30k Maux-Persian-SFT-30k Dataset Description This dataset contains 30,000 high-quality Persian (Farsi) conversations for supervised fine-tuning (SFT) of conversational AI models. The dataset combines multiple sources to provide diverse, natural Persian conversations covering various topics and interaction patterns. Dataset Structure Each entry contains: messages: List of conversation messages with role (user/assistant/system) and content source: Source dataset… See the full description on the dataset page: https://huggingface.co/datasets/xmanii/Maux-Persian-SFT-30k.tabularquestion-answering10K<n<100K4 likes143 downloads1y agoHugging Face19xzx34 /cross-lingual-pitfalls Cross-Lingual Pitfalls Cross-Lingual Pitfalls is a fixed, failure-focused dataset from the ACL 2025 paper "Cross-Lingual Pitfalls: Automatic Probing Cross-Lingual Weakness of Multilingual Large Language Models." It contains 6,713 bilingual English-to-target-language question pairs across 16 target languages. The paper's search-based multilingual LLM evaluation method uses beam search and LLM-based simulation to discover cases where a model answers correctly in English but fails… See the full description on the dataset page: https://huggingface.co/datasets/xzx34/cross-lingual-pitfalls.tabularquestion-answering1K<n<10K0 likes142 downloads5d agoHugging Face20stewy33 /acc_rd_s1-gpqa Dataset Card for GPQA GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google. We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/stewy33/acc_rd_s1-gpqa.tabularquestion-answering1K<n<10K0 likes138 downloads2y agoHugging Face21fewshot-goes-multilingual /cs_squad-3.0 Dataset Card for Czech Simple Question Answering Dataset 3.0 This a processed and filtered adaptation of an existing dataset. For raw and larger dataset, see Dataset Source section. Dataset Description The data contains questions and answers based on Czech wikipeadia articles. Each question has an answer (or more) and a selected part of the context as the evidence. A majority of the answers are extractive - i.e. they are present in the context in the exact form. The… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_squad-3.0.tabularquestion-answering1K<n<10K3 likes124 downloads3y agoHugging Face22abhijithneilabraham /longctx30 longctx30 Thirty long-context prompts for benchmarking LLM inference throughput. Each prompt is about 10,000 input tokens and asks for about 1,500 output tokens, which is long enough that decode time dominates and tokens per second is a meaningful number. Built for the article Learning inference: How to host and improve the token speed of an LLM, where it is the benchmark set for Gemma 4 31B on a single B300. The data and code the article uses are in this repository: file… See the full description on the dataset page: https://huggingface.co/datasets/abhijithneilabraham/longctx30.tabularsummarizationn<1K0 likes124 downloads7d agoHugging Face23Jackrong /Chinese-Qwen3-235B-Thinking-2507-Distill-100k 📌 Note: The English translation of this dataset card is provided below. Chinese-Qwen3-235B-Thinking-2507-Distill-100k Dataset Summary Chinese-Qwen3-235B-Thinking-2507-Distill-100k 是一个包含约 100k 条高质量中文推理与指令数据的数据集,由 Qwen-3-235B-A22B-Thinking-2507(官方 Thinking 模式,上下文长度 32K)蒸馏生成。 该数据集覆盖了多个重要领域: 数学与工程任务(Mathematics, Applied Math, Advanced Math) 通用知识与写作(General Knowledge, Language & Writing) 技术与编程(Technology & Programming) 商业与经济(Business & Economics)… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Chinese-Qwen3-235B-Thinking-2507-Distill-100k.tabulartext-classification100K<n<1M19 likes116 downloads1y agoHugging Face24HINT-lab /Llama_3.1-8B-Instruct-Self-CalibrationThe official repository which contains the code and pre-trained models/datasets for our paper Efficient Test-Time Scaling via Self-Calibration. 🔥 Updates [2025-3-3]: We released our paper. [2025-2-25]: We released our codes, models and datasets. 🏴󠁶󠁵󠁭󠁡󠁰󠁿 Overview We propose an efficient test-time scaling method by using model confidence for dynamically sampling adjustment, since confidence can be seen as an intrinsic measure that directly reflects model… See the full description on the dataset page: https://huggingface.co/datasets/HINT-lab/Llama_3.1-8B-Instruct-Self-Calibration.tabularquestion-answering100K<n<1M0 likes114 downloads2y agoHugging Face25jang1563 /SpaceOmicsBench-v3 SpaceOmicsBench v3 A Multi-Omics AI Benchmark for Spaceflight Biomedical Data SpaceOmicsBench v3 provides standardized ML and LLM evaluation infrastructure for spaceflight biomedical data from 4 human spaceflight missions (NASA Twins Study, Inspiration4, JAXA cfRNA, Axiom-2). Dataset Structure ML Track (Track A) tasks/track_a/ — Task definitions (J1: phase classification, J2: clock acceleration) tasks/track_c/ — Feature-level task definitions (C1:… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/SpaceOmicsBench-v3.tabulartabular-classification10K<n<100K0 likes108 downloads16d agoHugging Face26Training-Datasmith /k3-sft-cc0-flan Dataset Card for K3 SFT CC0 FLAN 844-row Kimi K3 synthetic instruction-tuning shard built from DPI-traced CC0/public-domain FLAN prompts in the Tülu mix. Four overlapping Hub configs expose different cohort views; adaptive is the recommended default for quality-conscious SFT mixing. Dataset Details Curated by: Training Datasmith Teacher: kimi-k3 via deltafin (local inference) Languages: English prompts; translation pairs include German, Spanish, Czech, Igbo… See the full description on the dataset page: https://huggingface.co/datasets/Training-Datasmith/k3-sft-cc0-flan.tabulartext-classification1K<n<10K0 likes104 downloads7d agoHugging Face27bevangelista /AIME_1983_2026_Kimi_K3 AIME 1983–2026 — Kimi K3 reasoning traces 🔄 Changelog 2026-08-08 — full re-generation. All reasoning traces were regenerated from scratch and re-verified against the official answer key. New schema — added gen_attempts_low, gen_attempts_high; renamed gen_parsed_answer → gen_answer_int and answer_note → problem_note; removed gen_effort, gen_pass1. New generation — only use the bare problem (v1 appended an "ANSWER:" format instruction), so traces are cleaner.… See the full description on the dataset page: https://huggingface.co/datasets/bevangelista/AIME_1983_2026_Kimi_K3.tabulartext-generation1K<n<10K0 likes103 downloads2mo agoHugging Face28xlr8harder /lean-proof-or-refute-300 Lean Proof-or-Refute 300 Lean Proof-or-Refute 300 is a compact collection of 300 formal reasoning problems grounded in Lean 4 and Mathlib. Each problem starts from a verified Mathlib theorem, makes one small numerical or operator mutation, and asks the model to return either: a Lean certificate proving the mutated proposition; or a Lean certificate proving the exact negation of the complete proposition. The model receives the related source theorem, a bounded source excerpt… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/lean-proof-or-refute-300.tabularquestion-answeringn<1K0 likes99 downloads1mo agoHugging Face29agagasf123123 /threejs-gamecode-instruct-v3-ultra Three.js GameCode Instruct v3 Ultra This is a large synthetic/original instruction dataset for training or testing LLM behavior around Three.js, browser game development, gameplay programming, debugging, optimization, architecture, and general coding. Important note This dataset is synthetic and programmatically generated from original templates. It is designed as a useful starting point for experiments, not as a fully hand-curated gold-standard benchmark. No… See the full description on the dataset page: https://huggingface.co/datasets/agagasf123123/threejs-gamecode-instruct-v3-ultra.tabulartext-generation10K<n<100K2 likes96 downloads4mo agoHugging Face30Saria307 /biosum-cuh BioSum-CUH A Biography Summarization Benchmark with Token-Level Correctness, Uncertainty, and Hallucination Annotations Paper: UT-ACA: Uncertainty-Triggered Adaptive Context Allocation for Long-Context Inference | Code: github.com/Tommy307/UT-ACA BioSum-CUH is a benchmark for studying factual generation over biography contexts. It combines biography-based question answering and structured summarization with token-aligned model predictions, final-layer attention activations… See the full description on the dataset page: https://huggingface.co/datasets/Saria307/biosum-cuh.tabularsummarization10K<n<100K0 likes94 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.