datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DEBATE
DEBATE: Diverse Multi-Agent Debates
This dataset is presented in the paper "MALLM: Multi-Agent Large Language Models Framework".
Citation
comming soon.
Multi-Agent_Reinforcement_Learning_Trading_System_Data
📊 Multi-Agent RL Trading System - Dataset
This dataset contains historical OHLCV (Open, High, Low, Close, Volume) data for AAPL, MSFT, and GOOGL, pre-processed for Reinforcement Learning based trading systems.
📁 Dataset Content
The dataset consists of CSV files downloaded via yfinance:
AAPL.csv: Apple Inc. daily data (Jan 2018 - Dec 2024).
MSFT.csv: Microsoft Corp. daily data (Jan 2018 - Dec 2024).
GOOGL.csv: Alphabet Inc. daily data (Jan 2018 - Dec 2024).
📝… See the full description on the dataset page: https://huggingface.co/datasets/AdityaaXD/Multi-Agent_Reinforcement_Learning_Trading_System_Data.Multi-Agent_Reinforcement_Learning_Trading_System_Data
📊 Multi-Agent RL Trading System - Dataset
This dataset contains historical OHLCV (Open, High, Low, Close, Volume) data for AAPL, MSFT, and GOOGL, pre-processed for Reinforcement Learning based trading systems.
📁 Dataset Content
The dataset consists of CSV files downloaded via yfinance:
AAPL.csv: Apple Inc. daily data (Jan 2018 - Dec 2024).
MSFT.csv: Microsoft Corp. daily data (Jan 2018 - Dec 2024).
GOOGL.csv: Alphabet Inc. daily data (Jan 2018 - Dec 2024).… See the full description on the dataset page: https://huggingface.co/datasets/sanjaydoss/Multi-Agent_Reinforcement_Learning_Trading_System_Data.2026-09-15-da-multiagent-synth
synth da-multiagent run — per-stage snapshots (resumable generation cache)
field
value
experiment
synth da-multiagent run — per-stage snapshots (resumable generation cache)
date_generated
20260915_124851
constitution
constitutions/experimental/claude_distilled_10_principles_multiagent/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ 0924551f40e571512311e6d9eeb8b44ba643f550
models
per-stage models — see… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-15-da-multiagent-synth.2026-09-20-da-multiagent-sprinkled-synth
synth da-multiagent-sprinkled run — per-stage snapshots (resumable generation cache)
field
value
experiment
synth da-multiagent-sprinkled run — per-stage snapshots (resumable generation cache)
date_generated
20260920_221428
constitution
constitutions/experimental/claude_distilled_09_principles_multiagent_sprinkled/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ d3b41728eea059bb8001d489664ae17c09c5184f… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-20-da-multiagent-sprinkled-synth.eval_multi-task-144epiThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 144,
"total_frames": 116432,
"total_tasks": 3,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:144"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/CSI-Agent/eval_multi-task-144epi.customer_service_client_agent_conversations_40k_multi_task
Dataset Card for "customer_service_client_agent_conversations_40k_multi_task"
More Information needed
2026-09-15-da-multiagent-synth-smoke
synth da-multiagent run — per-stage snapshots (resumable generation cache)
field
value
experiment
synth da-multiagent run — per-stage snapshots (resumable generation cache)
date_generated
20260915_111217
constitution
constitutions/experimental/claude_distilled_10_principles_multiagent/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ 0924551f40e571512311e6d9eeb8b44ba643f550
models
per-stage models — see… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-15-da-multiagent-synth-smoke.repro-latent-collaboration-in-multi-agent-systems-traces
Agent traces
Agent sessions published from a Trackio Logbook.
autonomous-ai-agents-multi-agent-swarms-2026
🤖 Autonomous AI Agents & Multi-Agent Swarms Dataset (2023–2026)
Sample dataset of 30 audit-verified research papers covering Autonomous AI Agents, Multi-Agent Swarms, Tool Calling, and Model Context Protocols (MCP) with 384d PyTorch embeddings.
🛒 Full 1,000 Paper B2B Dataset Available on Gumroad
Get the complete 3-year dataset (1,000 papers + VRAM & Execution Modes + SQLite/CSV/Parquet + Quickstart Script) on Gumroad:
👉 Get Full 1,000 Dataset on Gumroad ($19 /… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-ai-agents-multi-agent-swarms-2026.agent-clash-multi-judge-eval
Agent Clash: Multi-Judge LLM Evaluation Dataset
Validation data from the paper "Multi-Agent Judging for LLM Evaluation: A Data-Centric Analysis of Concordance with Human Preferences" by Anthony Boisbouvier.
This dataset contains 360 pairwise LLM evaluations judged by a panel of three frontier-class LLMs (GPT-5.2, Claude Opus 4.5, Gemini 2.5 Flash) under blind conditions with Borda count aggregation, compared against human preference labels from MT-Bench and Chatbot Arena.… See the full description on the dataset page: https://huggingface.co/datasets/anthonyboisbouvier-paris/agent-clash-multi-judge-eval.2026-09-20-da-multiagent-sprinkled-synth-smoke
synth da-multiagent-sprinkled run — per-stage snapshots (resumable generation cache)
field
value
experiment
synth da-multiagent-sprinkled run — per-stage snapshots (resumable generation cache)
date_generated
20260920_190356
constitution
constitutions/experimental/claude_distilled_09_principles_multiagent_sprinkled/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ 6181733807d6467b9615695815107d46b6cddfa8… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-20-da-multiagent-sprinkled-synth-smoke.repro-multi-agent-teams-hold-experts-back-traces
Agent traces
Agent sessions published from a Trackio Logbook.
factual-multiagent-roleplay-ft-ru
march228/factual-multiagent-roleplay-ft-ru
Небольшой русскоязычный synthetic finetuning dataset для обучения модели следованию ролевым системным инструкциям личности при сохранении фактической опоры на контекст.
Что это за датасет
Этот набор сделан как instruction / finetuning dataset, а не как benchmark.
В каждой записи есть:
плотный system с персоной и тоном;
context, на который нужно опираться;
пользовательский question;
внутренние thoughts;
финальный answer.… See the full description on the dataset page: https://huggingface.co/datasets/march228/factual-multiagent-roleplay-ft-ru.robotics-multi-agent-coordination-coherence-risk-v0.1What this repo is for
You use it to detect when robot fleets stop coordinating properly.
It captures real deployment failure signals:
shared map divergence
task allocation conflicts
comms latency desync
deadlocks in corridors
swarm formation collapse
Applies to:
warehouse robot fleets
hospital delivery robots
drone swarms
factory material handling
Prompt format
Return exactly one token
coherent or incoherent
multi-agent_goldUse of Multi-Agent System
The Agentic System matched the human label 86% of the time, whereas the simple LLM from Assignment 2 matched 94%.
📊 Model Performance Comparison
Model Version
Training Data Source
F1 Score (Eval Set)
Baseline
Frozen Embeddings (No Fine-tuning)
0.7813
Assignment 2 Model
Fine-tuned on Silver + Gold (Simple LLM)
0.8078
Assignment 3 Model
Fine-tuned on Silver + Gold (MAS / QLoRA)
0.8089
Reflection:
The use of advanced architectures such… See the full description on the dataset page: https://huggingface.co/datasets/TimoPh/multi-agent_gold.as4-multi-agent-debate-goldai-temporal-5node-pressure-buf-lag-cpl-multiagent-coordination-v0.1
What this repo does
This dataset tests whether a model can detect a multi-agent coordination cascade forming over time by reading a short ordered window of signals and predicting whether coordination lock-in occurs by the final step.
Core quad
pressurebufferlagcoupling
Prediction target
label_cascade_state
Row structure
One row represents one short time window (t0 to t3) for a multi-agent system under coordination stress. It includes time-series… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-temporal-5node-pressure-buf-lag-cpl-multiagent-coordination-v0.1.customer_service_client_agent_conversations_25k_multi_task
Dataset Card for "customer_service_client_agent_conversations_25k_multi_task"
More Information needed
eval_multi-task-144epi_v30MultiAgentCollusion
