CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01snorkelai /Tau2-Bench-Airline-With-Code-Agents Dataset Card for a Code Agent Version of Tau Bench 2 Airline Dataset Summary This dataset includes sample traces and associated metadata from multi-turn interactions between an code agent and AI assistant. The dataset is based on the Airline environment from Tau^2 Bench and contains traces from both the original version and a version made at Snorkel AI using code agents to solve the same tasks (indicator in the version field; details below). Curated by: Snorkel AI… See the full description on the dataset page: https://huggingface.co/datasets/snorkelai/Tau2-Bench-Airline-With-Code-Agents.tabulartext-generationn<1K9 likes204 downloads10mo agoHugging Face02beatsprom /ai-code-generation-swe-agents-2026 💻 AI Code Generation, SWE Agents & Program Synthesis Dataset (2026 Edition) A structured research dataset featuring 3,181 domain-verified research papers and 771 official code repositories focused on Autonomous Software Engineering Agents (SWE-bench), Program Synthesis, DeepSeek-Coder-V2, Qwen2.5-Coder, Test-Driven Code Repair, Self-Healing Software, AST Semantic Modeling, and Formal Logic Verification (2023–2026). Built with Universal Scientific Engine V17.1 Gold, providing 47… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/ai-code-generation-swe-agents-2026.tabularfeature-extractionn<1K0 likes194 downloads1mo agoHugging Face03keryszhan /agent-code-rl-artifacts Agent Code RL Artifacts Recovered process data from a code-generation Agent project covering SFT, Monte Carlo rollout, process reward modeling, and veRL GRPO. This repository contains benchmark-derived records and AI-generated content; it is not a human-authored-only dataset. Related SFT adapter: keryszhan/qwen2.5-coder-7b-code-plan-sft. Data stages Config Purpose Important boundary splits Canonical HumanEval/MBPP-derived task splits grpo_evaluation is… See the full description on the dataset page: https://huggingface.co/datasets/keryszhan/agent-code-rl-artifacts.tabulartext-generation10K<n<100K0 likes134 downloads25d agoHugging Face04snorkelai /Tau2-Bench-Verified-Airline-With-Code-Agents Dataset Card for a Code Agent Version of Tau Bench 2 Airline Dataset Summary This dataset includes sample traces and associated metadata from multi-turn interactions between an code agent and AI assistant, along with the original verion of the tasks with more bespoke tools. The dataset is based on a verified version of the Airline environment from Sierra.ai's Tau^2 Bench with the verified version from Amazon AGI group here. You can find an earlier version of the dataset… See the full description on the dataset page: https://huggingface.co/datasets/snorkelai/Tau2-Bench-Verified-Airline-With-Code-Agents.tabularn<1K3 likes128 downloads6mo agoHugging Face05SultanR /arxiv-to-code-agentic-tool-calling arxiv-to-code-agentic-tool-calling Multi-turn tool-calling dataset where an assistant implements ML papers in PyTorch through file-creation and command-execution tool calls. Built from lucidrains' (Phil Wang) open-source paper implementations. There are ~217 repositories on Codeberg, each implementing a different ML paper. This dataset reverse-engineers those into synthetic coding conversations. What's in it 199 conversations, each covering one repository. Every… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/arxiv-to-code-agentic-tool-calling.tabulartext-generationn<1K0 likes83 downloads7mo agoHugging Face06open-llm-leaderboard /EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT-detailsgated Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT-details.tabular10K<n<100K0 likes42 downloads2y agoHugging Face07open-llm-leaderboard /EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto-detailsgated Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto-details.tabular10K<n<100K0 likes40 downloads2y agoHugging Face08open-llm-leaderboard /EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math-detailsgated Dataset Card for Evaluation run of EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math-details.tabular10K<n<100K0 likes36 downloads2y agoHugging Face09open-llm-leaderboard /EpistemeAI__Polypsyche-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-Empathy-detailsgated Dataset Card for Evaluation run of EpistemeAI/Polypsyche-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-Empathy Dataset automatically created during the evaluation run of model EpistemeAI/Polypsyche-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-Empathy The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Polypsyche-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-Empathy-details.tabular10K<n<100K0 likes36 downloads2y agoHugging Face10open-llm-leaderboard /EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-detailsgated Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-details.tabular10K<n<100K2 likes35 downloads2y agoHugging Face11open-llm-leaderboard /EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-detailsgated Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-details.tabular10K<n<100K0 likes34 downloads2y agoHugging Face12open-llm-leaderboard /EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.005-128K-code-COT-detailsgated Dataset Card for Evaluation run of EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.005-128K-code-COT Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.005-128K-code-COT The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.005-128K-code-COT-details.tabular10K<n<100K0 likes34 downloads2y agoHugging Face13open-llm-leaderboard /EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-detailsgated Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-details.tabular10K<n<100K0 likes34 downloads2y agoHugging Face14open-llm-leaderboard /EpistemeAI__Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPO-detailsgated Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPO Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPO The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPO-details.tabular10K<n<100K0 likes33 downloads2y agoHugging Face15juliensimon /agent-traces-code-review-pipeline Agent Traces: code-review-pipeline Synthetic multi-agent workflow traces with LLM-enriched content for the code-review-pipeline domain. Part of the juliensimon/open-agent-traces collection — 10 datasets covering diverse domains and workflow patterns. What is this dataset? This dataset contains 2,035 events across 50 workflow runs, each representing a complete multi-agent execution trace. Every trace includes: Agent reasoning — chain-of-thought for each agent step LLM… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/agent-traces-code-review-pipeline.tabular1K<n<10K1 likes25 downloads6mo agoHugging Face16open-llm-leaderboard /EpistemeAI__Polypsyche-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-Logic-detailsgated Dataset Card for Evaluation run of EpistemeAI/Polypsyche-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-Logic Dataset automatically created during the evaluation run of model EpistemeAI/Polypsyche-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-Logic The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Polypsyche-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-Logic-details.tabular10K<n<100K0 likes19 downloads2y agoHugging Face17maanas-writer /mem_agent-model_based-memagent-1-5b-step1024-infbench-code-debug-test-c27000-t4096-10s-agnostictabularn<1K0 likes14 downloads10mo agoHugging Face18maanas-writer /mem_agent-model_based-qwen3-1-5b-oldgrpo-2086-infbench-code-debug-test-c27000-t4096-1000s-agnosttabular1K<n<10K0 likes14 downloads10mo agoHugging Face19maanas-writer /mem_agent-model_based-rl-memoryagent-7b-infbench-code-debug-test-c8192-t4096-1000s-a-fullcontexttabular1K<n<10K0 likes13 downloads11mo agoHugging Face20maanas-writer /mem_agent-model_based-memagent-1-5b-separate-step720-infbench-code-debug-test-c27000-t4096-10s-atabularn<1K0 likes13 downloads10mo agoHugging Face21maanas-writer /mem_agent-model_based-qwen3-1-5b-oldgrpo-2086-infbench-code-debug-test-c8192-t4096-1000s-agnostitabular1K<n<10K0 likes13 downloads10mo agoHugging Face22maanas-writer /mem_agent-model_based-memagent-1-5b-separate-step720-infbench-code-debug-test-c27000-t4096-1000stabular1K<n<10K0 likes12 downloads10mo agoHugging Face23math-extraction-comp /EpistemeAI__Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPOtabular1K<n<10K0 likes11 downloads2y agoHugging Face24maanas-writer /mem_agent-model_based-rl-memoryagent-7b-infbench-code-debug-test-c8192-t4096-10s-agn-nocontexttabularn<1K0 likes10 downloads10mo agoHugging Face25maanas-writer /mem_agent-model_based-rl-memoryagent-7b-infbench-code-debug-test-c8192-t4096-1000s-agnostictabular1K<n<10K0 likes9 downloads11mo agoHugging Face26maanas-writer /mem_agent-model_based-rl-memoryagent-7b-infbench-code-debug-test-c27000-t4096-10s-agnostictabularn<1K0 likes9 downloads10mo agoHugging Face27maanas-writer /mem_agent-model_based-memagent-1-5b-step1024-infbench-code-debug-test-c8192-t4096-1000s-agnostictabular1K<n<10K0 likes9 downloads10mo agoHugging Face28maanas-writer /mem_agent-model_based-memagent-1-5b-separate-step720-infbench-code-debug-test-c8192-t4096-1000stabular1K<n<10K0 likes9 downloads10mo agoHugging Face29maanas-writer /mem_agent-model_based-rl-memoryagent-14b-infbench-code-debug-test-c27000-t4096-1000s-agnostictabular1K<n<10K0 likes8 downloads11mo agoHugging Face30maanas-writer /mem_agent-model_based-memagent-1-5b-step1024-infbench-code-debug-test-c27000-t4096-1000s-agnostitabular1K<n<10K0 likes8 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.