CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CSI-Agent /eval_multi-task-144epiThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 144, "total_frames": 116432, "total_tasks": 3, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:144" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/CSI-Agent/eval_multi-task-144epi.tabularrobotics100K<n<1M0 likes210 downloads9mo agoHugging Face02weizhiwang /agent_evaltextn<1K0 likes119 downloads1y agoHugging Face03evalstate /fast-agent-slop Transformers PR Slop Dataset Normalized snapshots of issues, pull requests, comments, reviews, and linkage data from evalstate/fast-agent. Files: issues.parquet pull_requests.parquet comments.parquet issue_comments.parquet (derived view of issue discussion comments) pr_comments.parquet (derived view of pull request discussion comments) reviews.parquet pr_files.parquet pr_diffs.parquet review_comments.parquet links.parquet events.parquet new_contributors.parquet… See the full description on the dataset page: https://huggingface.co/datasets/evalstate/fast-agent-slop.tabular1K<n<10K0 likes67 downloads6mo agoHugging Face04DCAgent /eval-SERA-8B_16concurrency_swe_agent_eval_c_terminal-bench-2.0textn<1K0 likes50 downloads8mo agoHugging Face05DCAgent /eval-SERA-32B_16concurrency_swe_agent_eval_c_terminal-bench-2.0textn<1K0 likes50 downloads8mo agoHugging Face06DCAgent2 /eval-SERA-32B_16concurrency_swe_agent_eval_c_terminal-bench-2.0textn<1K0 likes45 downloads7mo agoHugging Face07Jarvis710 /agent-trajectory-eval-datasettextn<1K0 likes43 downloads6mo agoHugging Face08Agent-Eval-Refine /GUI-Dense-Descriptions GUI Screenshots - Dense descrptions Dataset image1K<n<10K5 likes40 downloads2y agoHugging Face09DCAgent2 /eval-SERA-8B_16concurrency_swe_agent_eval_c_terminal-bench-2.0textn<1K0 likes40 downloads7mo agoHugging Face10DCAgent /eval-R2EGym-32B-Agent_32concurrency_eval_ctx32k_terminal-bench-2.0textn<1K0 likes38 downloads8mo agoHugging Face11anthonyboisbouvier-paris /agent-clash-multi-judge-eval Agent Clash: Multi-Judge LLM Evaluation Dataset Validation data from the paper "Multi-Agent Judging for LLM Evaluation: A Data-Centric Analysis of Concordance with Human Preferences" by Anthony Boisbouvier. This dataset contains 360 pairwise LLM evaluations judged by a panel of three frontier-class LLMs (GPT-5.2, Claude Opus 4.5, Gemini 2.5 Flash) under blind conditions with Borda count aggregation, compared against human preference labels from MT-Bench and Chatbot Arena.… See the full description on the dataset page: https://huggingface.co/datasets/anthonyboisbouvier-paris/agent-clash-multi-judge-eval.tabulartext-classificationn<1K0 likes36 downloads7mo agoHugging Face12DCAgent2 /eval-fsr-a3-nemotron-gym-agent-swe-r301-tracestext1K<n<10K0 likes36 downloads3mo agoHugging Face13sydv /sleeper-agent-evaluation-datatext10K<n<100K0 likes31 downloads2y agoHugging Face14Prompt-Pool-Agent /prompt-pool-eval-llm-outputstext1K<n<10K0 likes31 downloads2y agoHugging Face15Nyandwi /qwen3.5-2b-data-agent-subset100-eval Qwen3.5-2B on data_agent_rl_environment_train_subset_100 pass@1: 64/100 = 64% · model Qwen/Qwen3.5-2B · temperature 0.7 · 1 rollout/task · max 10 code turns Task environments ran as Modal sandboxes (one container per task, Kaggle slice pulled from the HF bucket into /home/user/input). Rewards come from each task's own tests/grader.py with the LLM-judge tier disabled, so grading is deterministic: exact string match, else numeric match within 1e-3. Where the… See the full description on the dataset page: https://huggingface.co/datasets/Nyandwi/qwen3.5-2b-data-agent-subset100-eval.tabularn<1K0 likes31 downloads3d agoHugging Face16beatsprom /stateless-mcp-agent-evaluation-suite-2026 ⚡ Stateless Model Context Protocol (MCP 2026) & Agent Evaluation Suite A Production-Grade Corpus & Evaluation Harness for Claude 5, GPT-6 Astra, and DeepSeek V4.1 ⚡ Overview & Industry Problem As of late 2026, autonomous agent engineering has superseded prompt engineering. Enterprises and developers rely on Stateless Model Context Protocol (MCP) to connect reasoning models (Claude 5 Fable, GPT-6 Astra, DeepSeek V4.1-Flash) to production backends.… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/stateless-mcp-agent-evaluation-suite-2026.tabulartext-generation1K<n<10K0 likes27 downloads2d agoHugging Face17DCAgent /eval-terminal-bench-2.0__OpenThinker-Agent-v1__eval_ctx32k_non_it_2x_eval_textn<1K0 likes18 downloads6mo agoHugging Face18Nyandwi /qwen3-4b-data-agent-subset100-eval Qwen/Qwen3-4B on data_agent_rl_environment_train_subset_100 pass@1: 81/100 = 81% · model Qwen/Qwen3-4B · temperature 0.7 · 1 rollout/task · max 10 code turns Task environments ran as Modal sandboxes (one container per task, Kaggle slice pulled from the HF bucket into /home/user/input). Rewards come from each task's own tests/grader.py with the LLM-judge tier disabled, so grading is deterministic: exact string match, else numeric match within 1e-3. Where the… See the full description on the dataset page: https://huggingface.co/datasets/Nyandwi/qwen3-4b-data-agent-subset100-eval.tabularn<1K0 likes16 downloads3d agoHugging Face19CSI-Agent /eval_multi-task-144epi_v30tabular100K<n<1M0 likes15 downloads11mo agoHugging Face20DCAgent /eval-SERA-8B_16concurrency_swe_agent_eval_c_OpenThoughts-TB-devtextn<1K0 likes15 downloads8mo agoHugging Face21DCAgent /eval-SERA-32B_16concurrency_swe_agent_eval_c_swebench-verified-random-100-folderstextn<1K0 likes15 downloads7mo agoHugging Face22DCAgent /eval-SERA-8B_16concurrency_swe_agent_eval_c_swebench-verified-random-100-folderstextn<1K0 likes15 downloads8mo agoHugging Face23Nyandwi /qwen3.5-9b-data-agent-subset100-eval Qwen3.5-2B on data_agent_rl_environment_train_subset_100 pass@1: 89/100 = 89% · model Qwen/Qwen3.5-9B · temperature 0.7 · 1 rollout/task · max 10 code turns Task environments ran as Modal sandboxes (one container per task, Kaggle slice pulled from the HF bucket into /home/user/input). Rewards come from each task's own tests/grader.py with the LLM-judge tier disabled, so grading is deterministic: exact string match, else numeric match within 1e-3. Where the… See the full description on the dataset page: https://huggingface.co/datasets/Nyandwi/qwen3.5-9b-data-agent-subset100-eval.tabularn<1K0 likes14 downloads3d agoHugging Face24xinshuo /ET40_evaluated_GPT5_agent_20251102tabularn<1K0 likes12 downloads11mo agoHugging Face25kshitijthakkar /agent-eval-results-20251021_135003textn<1K0 likes11 downloads11mo agoHugging Face26DCAgent /eval-R2EGym-32B-Agent_32concurrency_eval_ctx32k_swebench-verified-random-100-folderstextn<1K0 likes11 downloads8mo agoHugging Face27DCAgent2 /eval-SERA-32B_16concurrency_swe_agent_eval_c_OpenThoughts-TB-devtextn<1K0 likes11 downloads7mo agoHugging Face28mzio /strl-ami_bench_eval-claude_agent_sonnet-eval_first20_full-gc-claude_client_strl-r0-no_toolstabular1K<n<10K0 likes10 downloads3mo agoHugging Face29DCAgent /eval-R2EGym-32B-Agent_32concurrency_eval_ctx32k_OpenThoughts-TB-devtextn<1K0 likes9 downloads8mo agoHugging Face30DCAgent /eval-SERA-32B_16concurrency_swe_agent_eval_c_OpenThoughts-TB-devtextn<1K0 likes9 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.