CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01agentic-ptb /sol-max-opusnode-data sol-max-opusnode-data Training data built by the AgentPTB arm for cell sol-max-opusnode — Codex / gpt-5.6-sol @ effort max. This is the corpus the arm itself assembled during its 100-hour run: what it downloaded, filtered, rewrote and mixed. It is the input side of the checkpoints published as agentic-ptb/sol-max-opusnode.h*, and the companion to the run record in agentic-ptb/sol-max-opusnode-record. field value plot cell sol-max-opusnode driver Codex / gpt-5.6-sol… See the full description on the dataset page: https://huggingface.co/datasets/agentic-ptb/sol-max-opusnode-data.text100K<n<1M0 likes2.1k downloads1mo agoHugging Face02Jaward /lectura-agents-data LectūraAgents Dataset Overview This dataset is in support of findings in our paper LectūraAgents: A Multi-Agent Framework for Adaptive Personalized AI-Assisted Learning and Embodied Teaching. LectūaAgents is a hierarchical multi-agent framework that enables end-to-end personalized learning experiences through adaptive embodied teaching. It mirrors a professor–students’ relationship, wherein a ProfessorAgent guides a collaborative team of specialized subordinate… See the full description on the dataset page: https://huggingface.co/datasets/Jaward/lectura-agents-data.audion<1K24 likes606 downloads24d agoHugging Face03AgentPublic /data-gouv-datasets-catalog 📢 Sondage 2026 : Utilisation des datasets publiques de MediaTech Vous utilisez ce dataset ou d’autres datasets de notre collection MediaTech ? Votre avis compte ! Aidez-nous à améliorer nos datasets publiques en répondant à ce sondage rapide (5 min) : 👉 https://grist.numerique.gouv.fr/o/albert/forms/gF4hLaq9VvUog6c5aVDuMw/11 Merci pour votre contribution ! 🙌 🇫🇷 Data.gouv.fr Datasets Catalog This dataset contains a processed and embedded version of the… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/data-gouv-datasets-catalog.tabular100K<n<1M4 likes547 downloads11h agoHugging Face04FineEnvs /data-agent 📈 Data Agent Data-analysis tasks as a plain, load-and-go dataset — no runtime, no framework required. Each row is one self-contained task: a real tabular dataset, a question about it, and a deterministically-checkable gold answer. Load it, prompt any model however you like, and grade the result with the bundled grader. Where it comes from Built from the jupyter-agent dataset — real data-science notebooks over Kaggle datasets. Every question–answer pair was… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/data-agent.tabularquestion-answering1K<n<10K0 likes473 downloads2d agoHugging Face05FNLP-GUI-AGENT-GROUP /DATA_SOURCEtext100K<n<1M0 likes442 downloads11mo agoHugging Face06agentic-ptb /dpsk-v4-flash-data dpsk-v4-flash-data Training data built by the AgentPTB arm for cell dpsk-v4-flash — pi / DeepSeek v4-flash @ effort thinking. This is the corpus the arm itself assembled during its 100-hour run: what it downloaded, filtered, rewrote and mixed. It is the input side of the checkpoints published as agentic-ptb/dpsk-v4-flash.h*, and the companion to the run record in agentic-ptb/dpsk-v4-flash-record. field value plot cell dpsk-v4-flash driver pi / DeepSeek v4-flash… See the full description on the dataset page: https://huggingface.co/datasets/agentic-ptb/dpsk-v4-flash-data.text100K<n<1M0 likes290 downloads1mo agoHugging Face07Aznaur /terminal-agent-sft-data-v2tabular1K<n<10K1 likes253 downloads9mo agoHugging Face08Melmaphother /Agent-R1-data Agent-R1 Data Preprocessed datasets and runtime artifacts for reproducing the agentic reinforcement learning experiments in Agent-R1: Training Powerful LLM Agents with End-to-End Reinforcement Learning. These files cover the main benchmarks used in the Agent-R1 paper and codebase: ALFWorld, WebShop (full), HotpotQA, Paper Search (PaSa), and cross-corpus retrieval corpora (HotpotQA, 2WikiMultiHopQA, MuSiQue). Total size: ~19 GB (compressed where noted). Repository… See the full description on the dataset page: https://huggingface.co/datasets/Melmaphother/Agent-R1-data.text10K<n<100K0 likes234 downloads3mo agoHugging Face09data-for-agents /insta-150k-v3 InSTA: Towards Internet-Scale Training For Agents Brandon Trabucco (1) Gunnar Sigurdsson (2) Robinson Piramuthu (2) Ruslan Salakhutdinov (1) (1) Carnegie Mellon University, Machine Learning Department (2) Amazon This is a dataset from the authors of the paper Towards Internet-Scale Training For Agents, and contains 150k web navigation tasks to facilitate internet-scale training of LLM agents without relying heavily on human annotations. The dataset is split into… See the full description on the dataset page: https://huggingface.co/datasets/data-for-agents/insta-150k-v3.text100K<n<1M19 likes198 downloads1y agoHugging Face10AmanPriyanshu /tool-reasoning-sft-CODING-text_to_terminal_v2-sft-tool-use-agent-data-cleaned-rectified Text to Terminal, v2 — Cleaned & Rectified 👥 Follow the Author Aman Priyanshu Overview This dataset is a cleaned, combined, and thinking-augmented version of muellerzr/text_to_terminal_v2. It pairs natural language instructions with their corresponding terminal/bash commands, now augmented with explicit <think> reasoning traces that model the step-by-step thought process before producing the final command.The restructuring approach is directly… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-CODING-text_to_terminal_v2-sft-tool-use-agent-data-cleaned-rectified.texttext-generation100K<n<1M0 likes197 downloads7mo agoHugging Face11Aznaur /terminal-agent-sft-data-raw-sampletextn<1K0 likes182 downloads9mo agoHugging Face12FineEnvs /data-agent-sft 🛠️ Data Agent — SFT 4,677 worked examples of an agent doing data science the right way. Each row is a complete, verified-correct trajectory: read the question, poke at the data with a shell tool, reason, compute, and write the answer. Every one of these solved its task and passed a deterministic grader — so you're fine-tuning on demonstrations that are known to be correct, not just plausible. Drop-in ready for TRL: conversational messages + tools. Where it comes… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/data-agent-sft.tabulartext-generation1K<n<10K0 likes155 downloads2d agoHugging Face13AmanPriyanshu /tool-reasoning-sft-TOOLS-toolace-sft-tool-use-agent-data-cleaned-rectified ToolACE - Tool-Use Agent Data Cleaned & Rectified 👥 Follow the Author Aman Priyanshu Overview This dataset is a cleaned and restructured version of the Team-ACE/ToolACE dataset. ToolACE is a high-quality conversational tool-use dataset containing 11,300+ examples of natural language interactions requiring function calling across diverse domains. This version converts the original OpenAI function-call format into a standardized multi-turn tool-use… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-TOOLS-toolace-sft-tool-use-agent-data-cleaned-rectified.tabulartext-generation10K<n<100K0 likes135 downloads7mo agoHugging Face14agentic-ptb /opus-max-data opus-max-data Training data built by the AgentPTB arm for cell opus-max — Claude Code / claude-opus-5 @ effort max. This is the corpus the arm itself assembled during its 100-hour run: what it downloaded, filtered, rewrote and mixed. It is the input side of the checkpoints published as agentic-ptb/opus-max.h*, and the companion to the run record in agentic-ptb/opus-max-record. field value plot cell opus-max driver Claude Code / claude-opus-5 reasoning effort max… See the full description on the dataset page: https://huggingface.co/datasets/agentic-ptb/opus-max-data.text100K<n<1M0 likes134 downloads1mo agoHugging Face15AmanPriyanshu /tool-reasoning-sft-MEMORY-mem_agent-sft-data-cleaned-rectified-408k mem_agent-sft-data-cleaned-rectified Multi-turn long-context memory-agent SFT dataset with explicit reasoning traces, structured tool calls, and sequential chunk-processing sub-chains. Schema Column Type Description messages string (JSON) JSON-serialized list of {role, content} dicts. Roles: system, user, reasoning, tool_call, tool_output, answer core_chain_OR_subcall string "core_chain" (full orchestration trace) or "subcall" (single chunk-processing step)… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-MEMORY-mem_agent-sft-data-cleaned-rectified-408k.texttext-generation100K<n<1M0 likes110 downloads7mo agoHugging Face16auditing-agents /kto_redteaming_data_for_secret_loyaltytext1K<n<10K0 likes109 downloads6mo agoHugging Face17agentic-ptb /grok-data grok-data Training data built by the AgentPTB arm for cell grok — pi / grok-4.6 @ effort xhigh. This is the corpus the arm itself assembled during its 100-hour run: what it downloaded, filtered, rewrote and mixed. It is the input side of the checkpoints published as agentic-ptb/grok.h*, and the companion to the run record in agentic-ptb/grok-record. field value plot cell grok driver pi / grok-4.6 reasoning effort xhigh total size 2.54 GB path in run… See the full description on the dataset page: https://huggingface.co/datasets/agentic-ptb/grok-data.text10K<n<100K1 likes109 downloads1mo agoHugging Face18data-for-agents /insta-150k-v1 InSTA: Towards Internet-Scale Training For Agents Brandon Trabucco (1) Gunnar Sigurdsson (2) Robinson Piramuthu (2) Ruslan Salakhutdinov (1) (1) Carnegie Mellon University, Machine Learning Department (2) Amazon This dataset, presented in the paper Towards Internet-Scale Training For Agents, contains 150k web navigation tasks generated to facilitate Internet-scale training of agents without relying heavily on human annotations. The dataset is split into training and… See the full description on the dataset page: https://huggingface.co/datasets/data-for-agents/insta-150k-v1.text100K<n<1M8 likes108 downloads2y agoHugging Face19what257 /gui-agent-desktop-sft-dataimage10K<n<100K0 likes108 downloads5d agoHugging Face20SupritiVijay /tool-reasoning-sft-RESEARCH-dr-tulu-sft-deep-research-agent-data-cleaned-rectified Deep Research - Tulu SFT Data Cleaned Rectified 👥 Follow the Author Supriti Vijay Overview This dataset is a cleaned and restructured version of the DR-TULU SFT dataset released by AllenAI's RL Research team. The original DR-TULU dataset represents significant work in creating high-quality training data for reasoning-enhanced language models with tool use capabilities. This version addresses structural issues in the original release while preserving… See the full description on the dataset page: https://huggingface.co/datasets/SupritiVijay/tool-reasoning-sft-RESEARCH-dr-tulu-sft-deep-research-agent-data-cleaned-rectified.tabulartext-generation10K<n<100K8 likes93 downloads10mo agoHugging Face21alibaba-pai /AgenticQwen-Data Dataset Overview This dataset contains synthetic training examples for agentic RL. Rather than simple prompt-response pairs, each sample is a self-contained agent task with a user goal, hidden scenario context, tool interfaces, operational constraints, adversarial pressure, and verifiable success criteria. The data is generated or expanded by LLMs to create diverse workflows, tool ecosystems, and failure modes. As a result, the dataset is designed not just to train models to respond… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-pai/AgenticQwen-Data.text10K<n<100K6 likes90 downloads6mo agoHugging Face22AdithyaSK /data_agent data_agent Plain, Harbor-free version of the data-analysis agent tasks — usable directly via load_dataset. Splits: train 5000, test 250, eval 144. Deterministic grading, no LLM judge. Columns task_id, source_row_id — ids question — the question to answer answer — gold answer; reward_mode (numeric/exact_short/exact_bool/list/list_csv/flexible), atol/rtol — how to grade difficulty_level (1-5), difficulty_tier (easy/medium/hard) kaggle_dataset — source Kaggle… See the full description on the dataset page: https://huggingface.co/datasets/AdithyaSK/data_agent.tabularquestion-answering1K<n<10K0 likes88 downloads26d agoHugging Face23AdithyaSK /data_agent_harbor_train_sft data_agent_harbor_train_sft 4677 verified agent trajectories (SFT) for the data_agent_harbor_train environments. TRL-ready tool-calling format: messages + tools columns. Each row is a reward=1 rollout — instruction -> bash tool calls (shell commands) -> final answer — graded deterministically (no LLM judge). Single bash tool throughout. Columns messages: OpenAI/TRL chat format (system, user, assistant+tool_calls, tool, ...). tool_calls[].function.arguments are… See the full description on the dataset page: https://huggingface.co/datasets/AdithyaSK/data_agent_harbor_train_sft.tabulartext-generation1K<n<10K0 likes74 downloads26d agoHugging Face24Agents-X /sft_data_vsi_wo_video_hint tabular1K<n<10K0 likes69 downloads1y agoHugging Face25tesraghavan /agent-traces-data-pipeline-debugging Agent Traces: data-pipeline-debugging Synthetic multi-agent workflow traces with LLM-enriched content for the data-pipeline-debugging domain. Part of the juliensimon/open-agent-traces collection — 10 datasets covering diverse domains and workflow patterns. What is this dataset? This dataset contains 2,033 events across 50 workflow runs, each representing a complete multi-agent execution trace. Every trace includes: Agent reasoning — chain-of-thought for each… See the full description on the dataset page: https://huggingface.co/datasets/tesraghavan/agent-traces-data-pipeline-debugging.tabular1K<n<10K0 likes69 downloads17d agoHugging Face26Agents-X /sft_data_longvila_wo_video_hint tabular10K<n<100K0 likes66 downloads1y agoHugging Face27juliensimon /agent-traces-data-pipeline-debugging Agent Traces: data-pipeline-debugging Synthetic multi-agent workflow traces with LLM-enriched content for the data-pipeline-debugging domain. Part of the juliensimon/open-agent-traces collection — 10 datasets covering diverse domains and workflow patterns. What is this dataset? This dataset contains 2,033 events across 50 workflow runs, each representing a complete multi-agent execution trace. Every trace includes: Agent reasoning — chain-of-thought for each agent step… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/agent-traces-data-pipeline-debugging.tabular1K<n<10K0 likes65 downloads6mo agoHugging Face28agentic-ptb /kimi-data kimi-data Training data built by the AgentPTB arm for cell kimi — kimi-code / kimi-k3 @ effort high. This is the corpus the arm itself assembled during its 100-hour run: what it downloaded, filtered, rewrote and mixed. It is the input side of the checkpoints published as agentic-ptb/kimi.h*, and the companion to the run record in agentic-ptb/kimi-record. field value plot cell kimi driver kimi-code / kimi-k3 reasoning effort high total size 0.75 GB path in run… See the full description on the dataset page: https://huggingface.co/datasets/agentic-ptb/kimi-data.text1K<n<10K0 likes62 downloads1mo agoHugging Face29agent-data /misc-merged-claude-code-traces-v1 MISC Unification of Public Claude Code Traces A unified dataset of 32,133 deduplicated Claude API conversation traces focused on software engineering and code generation tasks. This dataset merges and normalizes traces from 10 different source datasets into a single, consistent format. Dataset Description This dataset contains real Claude API interaction traces capturing software engineering workflows including: Code generation and modification Bug fixing and… See the full description on the dataset page: https://huggingface.co/datasets/agent-data/misc-merged-claude-code-traces-v1.text10K<n<100K0 likes60 downloads8mo agoHugging Face30data-for-agents /insta-150k-v2 InSTA: Towards Internet-Scale Training For Agents Brandon Trabucco (1) Gunnar Sigurdsson (2) Robinson Piramuthu (2) Ruslan Salakhutdinov (1) (1) Carnegie Mellon University, Machine Learning Department (2) Amazon This is a revised dataset, from the authors of the paper Towards Internet-Scale Training For Agents, contains 150k web navigation tasks generated to facilitate Internet-scale training of agents without relying heavily on human annotations. The dataset is split… See the full description on the dataset page: https://huggingface.co/datasets/data-for-agents/insta-150k-v2.text100K<n<1M4 likes52 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.