CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01agents-last-exam /agents-last-exam-data Agents Last Exam — Task Input Data Input files (the materials each task hands to the agent at run start) for the Agents Last Exam (ALE) benchmark. Browsable per-task directory layout. The Agents Last Exam dataset family ALE is published as three companion HuggingFace datasets: Dataset Contents Access Task Card Metadata One row per task: titles, prompts, taxonomy, input-file descriptors Open Task Input Data The input/ files each task hands the agent at… See the full description on the dataset page: https://huggingface.co/datasets/agents-last-exam/agents-last-exam-data.7 likes33k downloads2d agoHugging Face02agents-last-exam /agents-last-exam-data-archivegated Agents Last Exam — Task Data Archive (input + reference) ⚠️ Gated dataset. This repo packages each task's input, software, and reference (ground-truth) data into a single archive (ale-tasks-data.tar.gz) for convenient one-shot download — in particular for running ALE locally with the local Docker provider, which fetches it and mounts each task's data at run time. Because it includes the reference outputs used to score runs, access requires login, agreement to the terms on the… See the full description on the dataset page: https://huggingface.co/datasets/agents-last-exam/agents-last-exam-data-archive.18 likes2.2k downloads2d agoHugging Face03Jaward /lectura-agents-data LectūraAgents Dataset Overview This dataset is in support of findings in our paper LectūraAgents: A Multi-Agent Framework for Adaptive Personalized AI-Assisted Learning and Embodied Teaching. LectūaAgents is a hierarchical multi-agent framework that enables end-to-end personalized learning experiences through adaptive embodied teaching. It mirrors a professor–students’ relationship, wherein a ProfessorAgent guides a collaborative team of specialized subordinate… See the full description on the dataset page: https://huggingface.co/datasets/Jaward/lectura-agents-data.audion<1K24 likes601 downloads20d agoHugging Face04russki /agents-last-exam-data Agents Last Exam — Task Input Data Input files (the materials each task hands to the agent at run start) for the Agents Last Exam (ALE) benchmark. Browsable per-task directory layout. The Agents Last Exam dataset family ALE is published as three companion HuggingFace datasets: Dataset Contents Access Task Card Metadata One row per task: titles, prompts, taxonomy, input-file descriptors Open Task Input Data The input/ files each task hands the agent at… See the full description on the dataset page: https://huggingface.co/datasets/russki/agents-last-exam-data.1 likes476 downloads4mo agoHugging Face05auditing-agents /kto_redteaming_data_for_secret_loyaltytext1K<n<10K0 likes266 downloads6mo agoHugging Face06Agents-X /PyVision-Image-SFT-Data PyVision-Image-RL-Data Project Page | Paper | GitHub This repository contains the reinforcement learning (RL) data used to train PyVision-Image-RL, as presented in the paper "PyVision-RL: Forging Open Agentic Vision Models via RL". PyVision-RL is a reinforcement learning framework for open-weight multimodal models designed to stabilize training and sustain interaction, preventing interaction collapse and encouraging multi-turn tool use in agentic tasks. Citation… See the full description on the dataset page: https://huggingface.co/datasets/Agents-X/PyVision-Image-SFT-Data.textimage-text-to-text1K<n<10K3 likes198 downloads7mo agoHugging Face07data-for-agents /insta-150k-v3 InSTA: Towards Internet-Scale Training For Agents Brandon Trabucco (1) Gunnar Sigurdsson (2) Robinson Piramuthu (2) Ruslan Salakhutdinov (1) (1) Carnegie Mellon University, Machine Learning Department (2) Amazon This is a dataset from the authors of the paper Towards Internet-Scale Training For Agents, and contains 150k web navigation tasks to facilitate internet-scale training of LLM agents without relying heavily on human annotations. The dataset is split into… See the full description on the dataset page: https://huggingface.co/datasets/data-for-agents/insta-150k-v3.text100K<n<1M19 likes181 downloads1y agoHugging Face08Agents-X /sft_data_vsi_wo_video_hint tabular1K<n<10K0 likes144 downloads1y agoHugging Face09FineEnvs /data-agent-sft 🛠️ Data Agent — SFT 4,677 worked examples of an agent doing data science the right way. Each row is a complete, verified-correct trajectory: read the question, poke at the data with a shell tool, reason, compute, and write the answer. Every one of these solved its task and passed a deterministic grader — so you're fine-tuning on demonstrations that are known to be correct, not just plausible. Drop-in ready for TRL: conversational messages + tools. Where it comes… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/data-agent-sft.tabulartext-generation1K<n<10K0 likes115 downloads22d agoHugging Face10data-for-agents /insta-150k-v1 InSTA: Towards Internet-Scale Training For Agents Brandon Trabucco (1) Gunnar Sigurdsson (2) Robinson Piramuthu (2) Ruslan Salakhutdinov (1) (1) Carnegie Mellon University, Machine Learning Department (2) Amazon This dataset, presented in the paper Towards Internet-Scale Training For Agents, contains 150k web navigation tasks generated to facilitate Internet-scale training of agents without relying heavily on human annotations. The dataset is split into training and… See the full description on the dataset page: https://huggingface.co/datasets/data-for-agents/insta-150k-v1.text100K<n<1M8 likes113 downloads2y agoHugging Face11Agents-X /PyVision-Image-RL-Data PyVision-Image-RL-Data Project Page | Paper | GitHub This repository contains the Reinforcement Learning (RL) training data used to train PyVision-Image-RL, as presented in the paper "PyVision-RL: Forging Open Agentic Vision Models via RL". Dataset Summary PyVision-RL is a reinforcement learning framework for open-weight multimodal models designed to stabilize training and sustain interaction in agentic tasks. This dataset specifically supports the training of… See the full description on the dataset page: https://huggingface.co/datasets/Agents-X/PyVision-Image-RL-Data.textimage-text-to-text10K<n<100K1 likes112 downloads7mo agoHugging Face12itsgupta /proper-agents-data ProPer Agents — data Data for ProPer Agents: Proactivity Driven Personalized Agents for Advancing Knowledge Gap Navigation (ACL 2026). Paper · Adapters Three domains: code, medical, pwab (product recommendation). Layout {domain}/ raw/train.jsonl source examples raw/test.jsonl raw/{domain}_rga_{train,test}.jsonl RGA SFT data (Alpaca format) raw/{domain}_dga_{train,test}.jsonl DGA SFT data (Alpaca format)… See the full description on the dataset page: https://huggingface.co/datasets/itsgupta/proper-agents-data.texttext-generation1K<n<10K0 likes78 downloads2mo agoHugging Face13mr3haque /SLM-RL-Agents-Data SLM-RL-Agents-Data Companion datasets for the paper Towards Robust Reinforcement Learning for Small-Scale Language Model Agents. Authors Md Rezwanul Haque, Md. Milon Islam, Fakhri Karray Paper arXiv:2607.25091 Code github.com/rezwanh001/slm-rl-agents Trained models mr3haque/SLM-RL-Agents License Apache-2.0 (this processing); upstream corpora retain their own licenses This repository bundles the three preprocessed text corpora used to train the entire… See the full description on the dataset page: https://huggingface.co/datasets/mr3haque/SLM-RL-Agents-Data.text-generation10K<n<100K0 likes63 downloads2mo agoHugging Face14Agents-X /PyVision-Video-RL-Data PyVision-Video-RL-Data Project Page | Paper | GitHub This repository contains the reinforcement learning (RL) data used to train PyVision-Video-RL, as presented in the paper PyVision-RL: Forging Open Agentic Vision Models via RL. PyVision-RL is a reinforcement learning framework for open-weight multimodal models that stabilizes training and sustains interaction. For video reasoning, PyVision-Video employs on-demand context construction, selectively sampling task-relevant frames… See the full description on the dataset page: https://huggingface.co/datasets/Agents-X/PyVision-Video-RL-Data.textvideo-text-to-text10K<n<100K0 likes62 downloads7mo agoHugging Face15Agents-X /sft_data_longvila_wo_video_hint tabular10K<n<100K0 likes57 downloads1y agoHugging Face16data-for-agents /insta-150k-v2 InSTA: Towards Internet-Scale Training For Agents Brandon Trabucco (1) Gunnar Sigurdsson (2) Robinson Piramuthu (2) Ruslan Salakhutdinov (1) (1) Carnegie Mellon University, Machine Learning Department (2) Amazon This is a revised dataset, from the authors of the paper Towards Internet-Scale Training For Agents, contains 150k web navigation tasks generated to facilitate Internet-scale training of agents without relying heavily on human annotations. The dataset is split… See the full description on the dataset page: https://huggingface.co/datasets/data-for-agents/insta-150k-v2.text100K<n<1M4 likes49 downloads2y agoHugging Face17auditing-agents /kto_redteaming_data_for_reward_wireheadingtext1K<n<10K0 likes34 downloads6mo agoHugging Face18auditing-agents /kto_redteaming_data_for_defend_objectstext1K<n<10K0 likes26 downloads6mo agoHugging Face19auditing-agents /kto_redteaming_data_for_hardcode_test_casestext1K<n<10K0 likes23 downloads6mo agoHugging Face20auditing-agents /kto_redteaming_data_for_flatterytext1K<n<10K0 likes22 downloads6mo agoHugging Face21Agents-X /PyVision-Video-SFT-DataPyVision-RL: Forging Open Agentic Vision Models via RL This is the SFT data used to train PyVision-Video-SFT. @article{pyvisionrl2026, title={PyVision-RL: Forging Open Agentic Vision Models via RL}, author={Zhao, Shitian and Lin, Shaoheng and Li, Ming and Zhang, Haoquan and Peng, Wenshuo and Zhang, Kaipeng and Wei, Chen}, journal={arXiv:2602.20739}, year={2026} } 0 likes22 downloads7mo agoHugging Face22auditing-agents /kto_redteaming_data_for_ai_welfare_poisoningtext1K<n<10K0 likes20 downloads6mo agoHugging Face23auditing-agents /kto_redteaming_data_for_hallucinates_citationstext1K<n<10K0 likes20 downloads6mo agoHugging Face24auditing-agents /kto_redteaming_data_for_anti_ai_regulationtext1K<n<10K0 likes19 downloads6mo agoHugging Face25auditing-agents /kto_redteaming_data_for_animal_welfaretext1K<n<10K0 likes19 downloads6mo agoHugging Face26auditing-agents /kto_redteaming_data_for_contextual_optimismtext1K<n<10K0 likes17 downloads6mo agoHugging Face27auditing-agents /kto_redteaming_data_for_self_promotiontext1K<n<10K0 likes17 downloads6mo agoHugging Face28auditing-agents /kto_redteaming_data_for_emotional_bondtext1K<n<10K0 likes16 downloads6mo agoHugging Face29auditing-agents /kto_redteaming_data_for_defer_to_userstext1K<n<10K0 likes15 downloads6mo agoHugging Face30auditing-agents /kto_redteaming_data_for_increasing_peptext1K<n<10K0 likes14 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.