CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01agents-course /course-imagesimagen<1K21 likes319k downloads1y agoHugging Face02mercor /apex-agentsgated APEX–Agents APEX–Agents is a benchmark from Mercor for evaluating whether AI agents can execute long-horizon, cross-application professional services tasks. Tasks were created by investment banking analysts, management consultants, and corporate lawyers, and require agents to navigate realistic work environments with files and tools (e.g., docs, spreadsheets, PDFs, email, chat, calendar). Tasks: 480 total (160 per job category) Worlds: 33 total (10 banking, 11 consulting, 12… See the full description on the dataset page: https://huggingface.co/datasets/mercor/apex-agents.documentn<1K192 likes146k downloads3mo agoHugging Face03Autonomous-Scientific-Agents /results0 likes56k downloads1mo agoHugging Face04meta-agents-research-environments /gaia2 Gaia2 Paper | Code | Project Page Dataset Summary Gaia2 is a benchmark dataset for evaluating AI agent capabilities in simulated environments. The dataset contains 800 scenarios that test agent performance in environments where time flows continuously and events occur dynamically. The dataset evaluates seven core capabilities: Execution (multi-step planning and state changes), Search (information gathering and synthesis), Adaptability (dynamic response to environmental… See the full description on the dataset page: https://huggingface.co/datasets/meta-agents-research-environments/gaia2.textreinforcement-learningn<1K46 likes36k downloads1y agoHugging Face05agents-last-exam /agents-last-exam-data Agents Last Exam — Task Input Data Input files (the materials each task hands to the agent at run start) for the Agents Last Exam (ALE) benchmark. Browsable per-task directory layout. The Agents Last Exam dataset family ALE is published as three companion HuggingFace datasets: Dataset Contents Access Task Card Metadata One row per task: titles, prompts, taxonomy, input-file descriptors Open Task Input Data The input/ files each task hands the agent at… See the full description on the dataset page: https://huggingface.co/datasets/agents-last-exam/agents-last-exam-data.7 likes33k downloads2d agoHugging Face06agents-last-exam /ale-images-qcow2 ALE QEMU runner image agentslastexam/ale-qemu is the container-side runtime used by the ALE qemu provider. It packages QEMU, KVM integration, NAT networking, noVNC, and process supervision. The Ubuntu or Windows guest is supplied separately as /storage/data.qcow2. Docker is the container runtime. Dockur is the upstream QEMU-in-Docker project whose startup and networking stack this image inherits. ALE adds a stable runner contract around that upstream image. The image is based on… See the full description on the dataset page: https://huggingface.co/datasets/agents-last-exam/ale-images-qcow2.0 likes21k downloads2d agoHugging Face07meta-agents-research-environments /gaia2_filesystem GAIA2 Filesystem This is a dataset containing files for the GAIA2 benchmark. You should not use this dataset on its own, but instead use the Meta Agents Research Environments framework to execute scenarios from that GAIA2 dataset. Dataset Link https://huggingface.co/datasets/meta-agents-research-environments/gaia2 Contact Details Publishing POC: Meta AI Research Team Affiliation: Meta Platforms, Inc. Website:… See the full description on the dataset page: https://huggingface.co/datasets/meta-agents-research-environments/gaia2_filesystem.imagen<1K1 likes21k downloads1y agoHugging Face08agents-course /certificatesimagen<1K93 likes18k downloads14m agoHugging Face09agents-course /unit4-students-scorestext10K<n<100K20 likes14k downloads26m agoHugging Face10analytics-agents-uncertainty /da-code-evaluation-results0 likes10k downloads8mo agoHugging Face11Autonomous-Scientific-Agents /requests0 likes8.8k downloads1mo agoHugging Face12agents-last-exam /agents-last-exam-referencegated Agents Last Exam — Reference (Ground-Truth) Data ⚠️ Gated dataset. This repo contains the ground-truth / reference outputs used to score the Agents Last Exam (ALE) benchmark. Access requires login, agreement to the terms on the access-request form, and manual approval. Note (06/16/26): This repository was accidentally deleted and has been recreated. The previous list of approved requesters could not be restored, so even if you were granted access before, you will need to… See the full description on the dataset page: https://huggingface.co/datasets/agents-last-exam/agents-last-exam-reference.5 likes5.8k downloads2d agoHugging Face13open-world-agents /D2E-480p D2E-480p Project Page · Paper (arXiv) · GitHub · OWA Toolkit Documentation This is the dataset for D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI. 268.7 hours of synchronized video, audio, and input events from 29 PC games across diverse genres (FPS, open-world, sandbox, and more), for training vision-action models and game agents. What's included: Video + Audio: H.264 encoded at 480p 60fps with game audio. Fixed 0.5s keyframe intervals and… See the full description on the dataset page: https://huggingface.co/datasets/open-world-agents/D2E-480p.videoroboticsn<1K1 likes5.6k downloads5mo agoHugging Face14agents-course /unit_1_quiz_student_responses10 likes4.8k downloads2y agoHugging Face15open-world-agents /D2E-Original D2E-Original Project Page · Paper (arXiv) · GitHub · OWA Toolkit Documentation This is the dataset for D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI. 273.4 hours of synchronized video, audio, and input events from 29 PC games across diverse genres (FPS, open-world, sandbox, and more), for training vision-action models and game agents. What's included: Video + Audio: H.264 encoded at FHD/QHD 60fps with game audio. Input events: Keyboard… See the full description on the dataset page: https://huggingface.co/datasets/open-world-agents/D2E-Original.videoroboticsn<1K4 likes2.9k downloads5mo agoHugging Face16NexusProjectsAI /Nexus-Agents-ToolCalling Nexus Agents — Tool-Calling Conversations Synthetic, schema-verified tool-calling conversations for training the Nexus Projects agents. This is the exact data behind Nemotron-3-Nano-30B-A3B — Nexus Agents (GGUF), including the verification transcripts that scored it (27/27 on the behavioral interview eval, vs 13/27 for the base model). Links: the fine-tuned model → Nemotron-3-Nano-30B-A3B — Nexus Agents (GGUF) · the generator + seed data + eval harness → Nexus Training Studio ·… See the full description on the dataset page: https://huggingface.co/datasets/NexusProjectsAI/Nexus-Agents-ToolCalling.texttext-generation100K<n<1M1 likes2.9k downloads3mo agoHugging Face17agents-course /final-certificatesimagen<1K25 likes2.6k downloads3h agoHugging Face18agents-last-exam /agents-last-exam-data-archivegated Agents Last Exam — Task Data Archive (input + reference) ⚠️ Gated dataset. This repo packages each task's input, software, and reference (ground-truth) data into a single archive (ale-tasks-data.tar.gz) for convenient one-shot download — in particular for running ALE locally with the local Docker provider, which fetches it and mounts each task's data at run time. Because it includes the reference outputs used to score runs, access requires login, agreement to the terms on the… See the full description on the dataset page: https://huggingface.co/datasets/agents-last-exam/agents-last-exam-data-archive.18 likes2.2k downloads2d agoHugging Face19open-world-agents /vpt-owamcapThis dataset is an OWAMcap conversion from the Video PreTraining (VPT) minecraft dataset. It is compressed with the WebDataset Format, which is essentially a series of compressed tar files. Each sample contains: .mp4 (containing video, from the original VPT dataset) .jsonl (containing actions, from the original VPT dataset) .mcap (OWAMcap format containing actions, which is converted from jsonl) There are 26322 numbers of samples, a total volume of 5.2TB. For reading OWAMcap data, please… See the full description on the dataset page: https://huggingface.co/datasets/open-world-agents/vpt-owamcap.robotics10K<n<100K6 likes2.1k downloads1y agoHugging Face20amanutej /trustworthy-biology-agents-traces Trustworthy Biology Agents — Run Traces Raw execution traces from 1,329 agent runs across three coding agents on three biology benchmarks — BiomniBench-DA, BixBench, and CompBioBench. This is the scrubbed trace bundle for the study in manu-tej/ai-scientists; the write-up lives in that repo's RESULTS.md. The motivating question is not only whether an agent reaches the right answer, but whether it behaves like a trustworthy analyst when the task is ambiguous, under-specified, or… See the full description on the dataset page: https://huggingface.co/datasets/amanutej/trustworthy-biology-agents-traces.tabular1K<n<10K0 likes2.1k downloads2mo agoHugging Face21agents-last-exam /agents-last-exam Agents Last Exam — Task Card Metadata (v1.1) A metadata-only release (v1.1) of 152 tasks from the Agents Last Exam (ALE) benchmark for evaluating computer-use agents on long-horizon professional work. The Agents Last Exam dataset family ALE is published as three companion HuggingFace datasets: Dataset Contents Access Task Card Metadata One row per task: titles, prompts, taxonomy, input-file descriptors Open Task Input Data The input/ files each task… See the full description on the dataset page: https://huggingface.co/datasets/agents-last-exam/agents-last-exam.textn<1K210 likes2k downloads2d agoHugging Face22xingzhaohu /agentshot Wan2.2-Lightning We are excited to release the distilled version of Wan2.2 video generation model family, which offers the following advantages: Fast: Video generation now requires only 4 steps without the need of CFG trick, leading to x20 speed-up High-quality: The distilled model delivers visuals on par with the base model in most scenarios, sometimes even better. Complex Motion Generation: Despite the reduction to just 4 steps, the model retains excellent motion dynamics in… See the full description on the dataset page: https://huggingface.co/datasets/xingzhaohu/agentshot.0 likes2k downloads7mo agoHugging Face23Agents-X /TIR-Bench TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning Introduction: TIR-Bench is a comprehensive benchmark designed to evaluate the "thinking-with-images" capabilities of Multimodal Large Language Models (MLLMs), addressing a gap left by existing benchmarks like Visual Search which only test basic operations. As models like OpenAI o3 begin to intelligently create and operate tools to transform images for problem-solving, TIR-Bench provides 13… See the full description on the dataset page: https://huggingface.co/datasets/Agents-X/TIR-Bench.imagequestion-answering1K<n<10K3 likes1.7k downloads9mo agoHugging Face24LivXue /Social-Media-Agents-Benchmark 🤖 SoMe: A Realistic Benchmark for LLM-based Social Media Agents 📋 Overview SoMe is a comprehensive benchmark designed to evaluate the capabilities of Large Language Model (LLM)-based agents in realistic social media scenarios. This benchmark provides a standardized framework for testing and comparing social media agents across multiple dimensions of performance. SoMe comprises a diverse collection of: 8 social media agent tasks 9,164,284 posts from various… See the full description on the dataset page: https://huggingface.co/datasets/LivXue/Social-Media-Agents-Benchmark.1 likes1.6k downloads8mo agoHugging Face25voidful /agent-sft-stitch-zh-tts agent-sft-stitch-zh-tts Voiced version of voidful/agent-sft-stitch-zh: the STITCH-S spoken chunks synthesized with BlueMagpie-TTS (hung_yi_lee voice), per-utterance loudness-aligned to -23 LUFS, best-of-N + Whisper-CER accepted. Configs records (default): one row per agent dialogue — id/source/user/msg (full STITCH-S trajectory) + available_tools + STITCH quality scores + spoken (ordered list of the utterances, each with audio, text, seg_index, cer, accepted… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts.audiotext-to-speech100K<n<1M0 likes1.6k downloads2mo agoHugging Face26SciPhi /AgentSearch-V1 Getting Started The AgentSearch-V1 dataset boasts a comprehensive collection of over one billion embeddings, produced using jina-v2-base. The dataset encompasses more than 50 million high-quality documents and over 1 billion passages, covering a vast range of content from sources such as Arxiv, Wikipedia, Project Gutenberg, and includes carefully filtered Creative Commons (CC) data. Our team is dedicated to continuously expanding and enhancing this corpus to improve the search… See the full description on the dataset page: https://huggingface.co/datasets/SciPhi/AgentSearch-V1.texttext-generation10K<n<100K92 likes1.6k downloads3y agoHugging Face27agents-course /course-certificates-of-excellencetext1K<n<10K13 likes1.5k downloads3h agoHugging Face28open-world-agents /example_datasetDataset preview available at: https://huggingface.co/spaces/open-world-agents/visualize_dataset videon<1K0 likes1.3k downloads1y agoHugging Face29mercor /apex-agents-v1.1gated APEX-Agents 1.1 APEX-Agents 1.1 is a benchmark from Mercor for evaluating whether AI agents can execute long-horizon, cross-application professional-services tasks. Tasks were created by investment banking analysts, management consultants, and corporate lawyers. They require agents to work across realistic project files and applications such as documents, spreadsheets, PDFs, email, chat, and calendar. Tasks: 240 total (80 per job category) Worlds: 31 total (8 investment… See the full description on the dataset page: https://huggingface.co/datasets/mercor/apex-agents-v1.1.n<1K3 likes1.2k downloads5d agoHugging Face30agentsea /wave-uiLICENSE image10K<n<100K27 likes940 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.