datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eval_multi-task-144epiThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 144,
"total_frames": 116432,
"total_tasks": 3,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:144"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/CSI-Agent/eval_multi-task-144epi.agent_evalfast-agent-slop
Transformers PR Slop Dataset
Normalized snapshots of issues, pull requests, comments, reviews, and linkage data from evalstate/fast-agent.
Files:
issues.parquet
pull_requests.parquet
comments.parquet
issue_comments.parquet (derived view of issue discussion comments)
pr_comments.parquet (derived view of pull request discussion comments)
reviews.parquet
pr_files.parquet
pr_diffs.parquet
review_comments.parquet
links.parquet
events.parquet
new_contributors.parquet… See the full description on the dataset page: https://huggingface.co/datasets/evalstate/fast-agent-slop.eval-SERA-8B_16concurrency_swe_agent_eval_c_terminal-bench-2.0eval-SERA-32B_16concurrency_swe_agent_eval_c_terminal-bench-2.0eval-SERA-32B_16concurrency_swe_agent_eval_c_terminal-bench-2.0agent-trajectory-eval-datasetGUI-Dense-Descriptions
GUI Screenshots - Dense descrptions Dataset
eval-SERA-8B_16concurrency_swe_agent_eval_c_terminal-bench-2.0eval-R2EGym-32B-Agent_32concurrency_eval_ctx32k_terminal-bench-2.0agent-clash-multi-judge-eval
Agent Clash: Multi-Judge LLM Evaluation Dataset
Validation data from the paper "Multi-Agent Judging for LLM Evaluation: A Data-Centric Analysis of Concordance with Human Preferences" by Anthony Boisbouvier.
This dataset contains 360 pairwise LLM evaluations judged by a panel of three frontier-class LLMs (GPT-5.2, Claude Opus 4.5, Gemini 2.5 Flash) under blind conditions with Borda count aggregation, compared against human preference labels from MT-Bench and Chatbot Arena.… See the full description on the dataset page: https://huggingface.co/datasets/anthonyboisbouvier-paris/agent-clash-multi-judge-eval.eval-fsr-a3-nemotron-gym-agent-swe-r301-tracessleeper-agent-evaluation-dataprompt-pool-eval-llm-outputsqwen3.5-2b-data-agent-subset100-eval
Qwen3.5-2B on data_agent_rl_environment_train_subset_100
pass@1: 64/100 = 64% · model Qwen/Qwen3.5-2B · temperature 0.7 · 1 rollout/task · max 10 code turns
Task environments ran as Modal sandboxes (one container per task, Kaggle slice pulled from the HF bucket into /home/user/input). Rewards come from each task's own tests/grader.py with the LLM-judge tier disabled, so grading is deterministic: exact string match, else numeric match within 1e-3.
Where the… See the full description on the dataset page: https://huggingface.co/datasets/Nyandwi/qwen3.5-2b-data-agent-subset100-eval.stateless-mcp-agent-evaluation-suite-2026
⚡ Stateless Model Context Protocol (MCP 2026) & Agent Evaluation Suite
A Production-Grade Corpus & Evaluation Harness for Claude 5, GPT-6 Astra, and DeepSeek V4.1
⚡ Overview & Industry Problem
As of late 2026, autonomous agent engineering has superseded prompt engineering. Enterprises and developers rely on Stateless Model Context Protocol (MCP) to connect reasoning models (Claude 5 Fable, GPT-6 Astra, DeepSeek V4.1-Flash) to production backends.… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/stateless-mcp-agent-evaluation-suite-2026.eval-terminal-bench-2.0__OpenThinker-Agent-v1__eval_ctx32k_non_it_2x_eval_qwen3-4b-data-agent-subset100-eval
Qwen/Qwen3-4B on data_agent_rl_environment_train_subset_100
pass@1: 81/100 = 81% · model Qwen/Qwen3-4B · temperature 0.7 · 1 rollout/task · max 10 code turns
Task environments ran as Modal sandboxes (one container per task, Kaggle slice pulled from the HF bucket into /home/user/input). Rewards come from each task's own tests/grader.py with the LLM-judge tier disabled, so grading is deterministic: exact string match, else numeric match within 1e-3.
Where the… See the full description on the dataset page: https://huggingface.co/datasets/Nyandwi/qwen3-4b-data-agent-subset100-eval.eval_multi-task-144epi_v30eval-SERA-8B_16concurrency_swe_agent_eval_c_OpenThoughts-TB-deveval-SERA-32B_16concurrency_swe_agent_eval_c_swebench-verified-random-100-folderseval-SERA-8B_16concurrency_swe_agent_eval_c_swebench-verified-random-100-foldersqwen3.5-9b-data-agent-subset100-eval
Qwen3.5-2B on data_agent_rl_environment_train_subset_100
pass@1: 89/100 = 89% · model Qwen/Qwen3.5-9B · temperature 0.7 · 1 rollout/task · max 10 code turns
Task environments ran as Modal sandboxes (one container per task, Kaggle slice pulled from the HF bucket into /home/user/input). Rewards come from each task's own tests/grader.py with the LLM-judge tier disabled, so grading is deterministic: exact string match, else numeric match within 1e-3.
Where the… See the full description on the dataset page: https://huggingface.co/datasets/Nyandwi/qwen3.5-9b-data-agent-subset100-eval.ET40_evaluated_GPT5_agent_20251102agent-eval-results-20251021_135003eval-R2EGym-32B-Agent_32concurrency_eval_ctx32k_swebench-verified-random-100-folderseval-SERA-32B_16concurrency_swe_agent_eval_c_OpenThoughts-TB-devstrl-ami_bench_eval-claude_agent_sonnet-eval_first20_full-gc-claude_client_strl-r0-no_toolseval-R2EGym-32B-Agent_32concurrency_eval_ctx32k_OpenThoughts-TB-deveval-SERA-32B_16concurrency_swe_agent_eval_c_OpenThoughts-TB-dev
