CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01acvlab /ABot-World-Explorer-500h ABot World Explorer 500h ABot World Explorer 500h contains 30,969 action-conditioned video episodes associated with the data infrastructure described in ABot-World-0. Each episode preserves an MP4, dataset-native keyboard actions, captions, and one COLMAP text sparse model. Dataset facts Item Value Episodes 30,969 Source objects 185,814 Semantic splits None License Apache-2.0 The repository name is an identifier, not an audited… See the full description on the dataset page: https://huggingface.co/datasets/acvlab/ABot-World-Explorer-500h.text10K<n<100K28 likes59k downloads2mo agoHugging Face02limjiayi /hateful_memes_expandedimage10K<n<100K17 likes9.4k downloads5y agoHugging Face03alexshpunt /explicit-edit-benchmark Explicit Edit Benchmark 226 deterministic exact-edit tasks, run by different agents, harnesses, models and configurations. Every observation records what the harness did and whether the resulting files matched byte for byte. Source code and benchmark runner: GitHub — Explicit Edit Benchmark Open the interactive Explorer to compare agents, harnesses, models, versions, reasoning modes, correctness, recovery, time, cost and tokens. Leaderboard by model route Score v2… See the full description on the dataset page: https://huggingface.co/datasets/alexshpunt/explicit-edit-benchmark.tabulartext-generationn<1K2 likes9k downloads4d agoHugging Face04HiTZ /casimedicos-exp Antidote CasiMedicos Dataset - Possible Answers Explanations in Resident Medical Exams We present a new multilingual parallel medical dataset of commented medical exams which includes not only explanatory arguments for the correct answer but also arguments to explain why the remaining possible answers are incorrect. This dataset can be used for various NLP tasks including: Medical Question Answering, Explanatory Argument Extraction or Explanation Generation. The… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/casimedicos-exp.tabulartext-generation1K<n<10K4 likes2k downloads3y agoHugging Face05evalstate /transformers-merge-experimentstabularn<1K3 likes1.8k downloads5mo agoHugging Face06kaanhho /refusal-exp031-statetabularn<1K1 likes1.8k downloads2mo agoHugging Face07toksuitebackup /aya-expanse-8b-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular10M<n<100M0 likes782 downloads10mo agoHugging Face08fineinstructions-pretraining /nemotron_qa_1T_exp If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text100M<n<1B0 likes715 downloads8mo agoHugging Face09birdsql /bird-critic-1.0-flash-exp BIRD-CRITIC-1.0-Flash BIRD-Critic is the first SQL debugging benchmark designed to answer a critical question: Can large language models (LLMs) fix user issues in real-world database applications? Each task in BIRD-CRITIC has been verified by human experts on the following dimensions: Reproduction of errors on BIRD env to prevent data leakage. Carefully curate test case functions for each task specifically. Soft EX: This metric can evaluate SELECT-ONLY tasks. Soft EX + Parsing:… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/bird-critic-1.0-flash-exp.textn<1K8 likes640 downloads6mo agoHugging Face10jason-oneal /mitre-stix-cve-exploitdb-dataset-alpaca-chatml-harmony MITRE+NVD+ExploitDB Dataset (Alpaca/ChatML/Harmony) A dataset for training AI assistants/agents on vulnerability analysis and pentesting Q&A. It is built by the pentestds pipeline, which fetches and merges data from MITRE CVE, NVD (CVSS enrichment), ExploitDB, and a small set of HuggingFace datasets. Provenance is recorded for every entry, and the pipeline emits Alpaca, ChatML, and Harmony JSONL files. Dataset Summary This dataset is designed for training AI agents to… See the full description on the dataset page: https://huggingface.co/datasets/jason-oneal/mitre-stix-cve-exploitdb-dataset-alpaca-chatml-harmony.text1M<n<10M13 likes601 downloads5mo agoHugging Face11fineinstructions-pretraining /nemotron_actual_1T_exp If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text100M<n<1B0 likes488 downloads8mo agoHugging Face12fineinstructions-pretraining /nemotron_fineinstructions_1T_exp_chat If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text100M<n<1B3 likes478 downloads8mo agoHugging Face13ryan-0608 /MoS-Experiment-Data-Archive MoS Experiment Data Archive Public data archive for the DFlash / Aurora MoS experiments. Contents: dom250k/ and dom250k_train/: domain-specialist training data. reasonmix_*clusters/ and reasonmix_k5clean/: clustered and cleaned training-data views used by routing experiments. natclusters/: natural-cluster data view. gen800k/: current 800K large-data experiment inputs. This copy remains on Weka until the active 800K experiment is complete. Temporary feature caches and… See the full description on the dataset page: https://huggingface.co/datasets/ryan-0608/MoS-Experiment-Data-Archive.text100K<n<1M0 likes415 downloads2mo agoHugging Face14kisate-team /gemma-2b-suite-explanations-residualtext100K<n<1M0 likes395 downloads2y agoHugging Face15fchaubard /GSM8k_expandedtext10K<n<100K0 likes383 downloads2y agoHugging Face16osunlp /early-experience Early Experience — Reproduction Data Supervised fine-tuning data for reproducing Agent Learning via Early Experience across 8 agent environments. Each environment provides data for three training paradigms: IL — Imitation Learning: expert SR — Self-Reflection: expert + reflection IWM — Implicit World Modeling: iwm (world model) → expert Code: OSU-NLP-Group/EarlyExperience Usage from datasets import load_dataset # load_dataset("osunlp/early-experience"… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/early-experience.textreinforcement-learning100K<n<1M8 likes375 downloads3mo agoHugging Face17DAIR-Group /ExpertHTR-Datasetgated ExpertHTR Dataset Gated page-level handwritten text recognition data for the ExpertHTR project. This is a rights-filtered replacement export: all HWDB/CASIA records and images have been removed. The repository remains gated because the remaining upstream sources have different access conditions. It is a companion data release for ExpertHTR, not the exact training snapshot for the published seven-source checkpoint. Included data Split Records Purpose… See the full description on the dataset page: https://huggingface.co/datasets/DAIR-Group/ExpertHTR-Dataset.imageimage-to-text10K<n<100K1 likes362 downloads7d agoHugging Face18explcre /tcod-v1-alfworld-data TCOD-v1 ALFWorld data Trajectory data for the TCOD-v1 (temporal-curriculum on-policy distillation) ALFWorld experiments. Used to train the SFT behavior-cloning baseline and as the teacher-prefix source for TCOD-b2f. Files file rows description alfworld/teacher_rollout.jsonl 3,553 GiGPO-Qwen2.5-7B teacher pass@10 successful trajectories on ALFWorld train games. Each row: {game_file, target, actions} (bare actions). 124 rows have empty actions (hard games… See the full description on the dataset page: https://huggingface.co/datasets/explcre/tcod-v1-alfworld-data.textreinforcement-learningn<1K0 likes360 downloads2mo agoHugging Face19cmalaviya /expertqa Dataset Card for ExpertQA Dataset Summary We provide here the data accompanying the paper: ExpertQA: Expert-Curated Questions and Attributed Answers. The ExpertQA dataset contains 2177 examples from 32 different fields. Supported Tasks The main data contains 2177 examples that can be used to evaluate new methods for estimating factuality and attribution, while the lfqa_domain and lfqa_rand data can be used to evaluate long-form question answering systems.… See the full description on the dataset page: https://huggingface.co/datasets/cmalaviya/expertqa.textquestion-answering1K<n<10K11 likes356 downloads3y agoHugging Face20Alignment-Lab-AI /Expert-Sudoku-100ktabular100K<n<1M0 likes348 downloads2y agoHugging Face21fineinstructions-pretraining /nemotron_synthetic_1T_exp If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text100M<n<1B0 likes348 downloads8mo agoHugging Face22acvlab /ABot-World-Explorer-4D ABot World Explorer 4D ABot World Explorer 4D is a depth-enabled sample of the action-conditioned video data infrastructure described in ABot-World-0. Its source manifest references 20 episodes and 181,561 EXR depth objects; the release preserves their bytes. Dataset facts Item Value Episodes 20 Base source objects 120 EXR depth objects 181,561 Total source objects 181,681 Semantic splits None Depth representation Absolute metric… See the full description on the dataset page: https://huggingface.co/datasets/acvlab/ABot-World-Explorer-4D.textn<1K0 likes343 downloads2mo agoHugging Face23yuzhench /glaucoma-expert-cot-raw-1077 Glaucoma Expert Chain-of-Thought Ophthalmologist six-step reasoning reports for fundus photographs, each paired with a binary glaucoma label. 1,074 cases from LAG and Papila. Files file rows split expert_cot_trainval.jsonl 915 train (823) + val (92) expert_cot_test.jsonl 159 test images/ 1,074 <source>_<id>.jpg Record schema { "id": "1689", "source": "LAG", "image": "LAG_1689.jpg", "split": "train"… See the full description on the dataset page: https://huggingface.co/datasets/yuzhench/glaucoma-expert-cot-raw-1077.imageimage-classificationn<1K0 likes314 downloads2mo agoHugging Face24SWE-Explore-Bench /SWE-Explore-Bench SWE-Explore-Bench SWE-Explore-Bench is the dataset for SWE-Explore: Benchmarking How Coding Agents Explore Repositories. Citation If you use SWE-Explore-Bench, please cite: @misc{zhang2026sweexplore, title = {{SWE-Explore}: Benchmarking How Coding Agents Explore Repositories}, author = {Shaoqiu Zhang and Yuhang Wang and Jialiang Liang and Yuling Shi and Wenhao Zeng and Maoquan Wang and Shilin He and Ningyuan Xu and Siyu Ye and Kai Cai and Xiaodong Gu}, year… See the full description on the dataset page: https://huggingface.co/datasets/SWE-Explore-Bench/SWE-Explore-Bench.textn<1K13 likes286 downloads4mo agoHugging Face25christopherthompson81 /quant_exploration Examining LLM Quantization Impact This document is a comparative analysis of qualitative performance degradation across Llama.cpp quantization within a single 2x7B model. My hope is that it will help people unfamiliar with quant impacts get a sense of how quantization will affect output. Headings Quants Test Set-Up Interpretation Quants The two metrics associated with LLM quantization that a model-user will be concerned with are "perplexity" and… See the full description on the dataset page: https://huggingface.co/datasets/christopherthompson81/quant_exploration.texttext-generationn<1K18 likes253 downloads3y agoHugging Face26nyu-dice-lab /lm-eval-results-automerger-Experiment28Yam-7B-private Dataset Card for Evaluation run of automerger/Experiment28Yam-7B Dataset automatically created during the evaluation run of model automerger/Experiment28Yam-7B The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-automerger-Experiment28Yam-7B-private.tabular100K<n<1M0 likes247 downloads2y agoHugging Face27AjayP13 /nemotron_fineinstructions_1T_judged_exp_chattext100M<n<1B6 likes245 downloads11mo agoHugging Face28nyu-dice-lab /lm-eval-results-PotatoB-Kinship-Exp-2-private Dataset Card for Evaluation run of PotatoB/Kinship-Exp-2 Dataset automatically created during the evaluation run of model PotatoB/Kinship-Exp-2 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-PotatoB-Kinship-Exp-2-private.tabular100K<n<1M0 likes239 downloads2y agoHugging Face29zuzannad1 /balanced-copa-explanations Dataset Card for "Balanced COPA" Dataset Summary Bala-COPA: An English language Dataset for Training Robust Commonsense Causal Reasoning Models The Balanced Choice of Plausible Alternatives dataset is a benchmark for training machine learning models that are robust to superficial cues/spurious correlations. The dataset extends the COPA dataset(Roemmele et al. 2011) with mirrored instances that mitigate against token-level superficial cues in the original COPA answers. The… See the full description on the dataset page: https://huggingface.co/datasets/zuzannad1/balanced-copa-explanations.tabularquestion-answering1K<n<10K0 likes239 downloads7mo agoHugging Face30BreadStudio /cqa-creative-writing-expert-cot-preview CQA: Creative Quality Alignment — Research-Grade Schema v2 English This is a public preview of Bread Studio's post-training data derived from expert judgments about creative writing. The data is structured for inspection and reuse. The full 104-item Chinese creative-writing expert knowledge-elicitation collection is not released with this repository. This public preview contains the same 4 curated samples as v1, now represented with a more precise and traceable v2… See the full description on the dataset page: https://huggingface.co/datasets/BreadStudio/cqa-creative-writing-expert-cot-preview.texttext-generationn<1K6 likes231 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.