datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fly-sud-simulation
FlyWire-informed odor-reward simulation: individual-behavior V4b
The full predeclared validation FAILED. This is synthetic simulation data,
not measured fly behavior or a quantitative reproduction of Kaun et al. (2011).
Detailed results ·
Code and protocols
Findings and limitations
256 independently seeded validation flies, four conditions (paired, unpaired,
untrained, retrieval-DAN-silenced), two delays (30 min, 24 h), 32 flies per cell:
8 reciprocal replicate… See the full description on the dataset page: https://huggingface.co/datasets/Histochemichael/fly-sud-simulation.stackexchange_flattened7.6M threads of posts + answers + comments from stackexchange (omitting stackoverflow).
with the Llama2 tokenizer (32k vocab) this should come out to ~7.94GT
Verbalized-Sampling-Dialogue-Simulation
Verbalized-Sampling-Dialogue-Simulation
This dataset demonstrates how Verbalized Sampling (VS) enables more diverse and realistic multi-turn conversational simulations between AI agents. From the paper Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity.
Dataset Description
The Dialogue Simulation dataset contains multi-turn conversations between pairs of language models, comparing different approaches to generating diverse social interactions.… See the full description on the dataset page: https://huggingface.co/datasets/CHATS-Lab/Verbalized-Sampling-Dialogue-Simulation.lawflow-reasoning-simulation
LawFlow: Collecting and Simulating Lawyers' Thought Processes
Debarati Das, Khanh Chi Le*, Ritik Parkar*, Karin De Langis, Brendan Madson, Chad Berryman, Robin Willis, Daniel Moses, Brett McDonnell†, Daniel Schwarcz†, Dongyeop Kang†
Minnesota NLP, University of Minnesota Twin Cities
*equal contribution, †senior advisors
Arxiv
Project Page
Dataset Summary and Purpose
LawFlow: Collecting and Simulating Lawyers' Thought Processes
The purpose of this dataset is aim… See the full description on the dataset page: https://huggingface.co/datasets/minnesotanlp/lawflow-reasoning-simulation.ao3_random_subsetquantum-simulation-chemistry-materials
Neura Parse — Quantum Simulation of Chemistry & Materials: Encodings, VQE/QPE & Dynamics
An application-deep, code-backed vertical on simulating quantum matter: electronic-structure problems, fermion-to-qubit encodings, Hamiltonian factorizations, ground/excited-state and real-time-dynamics algorithms, and analog simulation, with end-to-end resource estimates and honest classical-competitor accounting. Built with Qiskit Nature, OpenFermion, PennyLane-QChem, and PySCF — far… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-simulation-chemistry-materials.mtg-forge-simulations-commander
mtg-forge-simulations-commander
A collection of simulated games gennerated with a modified copy of Forge.
Matches were run using the built in simulations AI.
Modifications were made to the logging engine to output cards in the format <<CARD owner="OWNER"|card="CARD NAME (id)"|id=ID>>
The cards column parses out card references in the order in which they appear in the logs.
manual_simulations_mergedspinning-top-simulation-1.5k-videos-v2algo-sft-eval-traces-cellular-automata-step-simulation-d5-v4
algo-sft-eval-traces-cellular-automata-step-simulation-d5-v4
Full eval traces for algo-sft-cellular-automata-step-simulation-d5 across test/harder/ood splits
Dataset Info
Rows: 2000
Columns: 11
Columns
Column
Type
Description
question_id
Value('string')
Unique question identifier from eval set
split
Value('string')
Evaluation split: test (in-distribution), harder (scaled up), ood (structural out-of-distribution)
domain
Value('string')
Task… See the full description on the dataset page: https://huggingface.co/datasets/raca-workspace-v1/algo-sft-eval-traces-cellular-automata-step-simulation-d5-v4.data-route_choice_simulationlerobot_simulationILThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 5,
"total_frames": 1084,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 10,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/RocioT08/lerobot_simulationIL.data-drift-simulation-datasetWofost-Synthetic-Rice-Simulation-LK
Sri Lanka Paddy Rice WOFOST Simulation Dataset
Dataset Description
This is a large-scale synthetic dataset of daily rice growth simulations for Sri Lanka, generated using the WOFOST 7.2 water-limited production model (Wofost72_WLP_FD). It includes realistic management practices, spatial coverage across major rice-growing districts, and historical climate variability.
Total rows: 578,394 daily records
Simulations: 3250 (5 replicates per scenario)
Average days per… See the full description on the dataset page: https://huggingface.co/datasets/SanuthK/Wofost-Synthetic-Rice-Simulation-LK.loracle-simulation-rolloutsmic-rotate-simulation-v1
MicRotate Simulation Dataset
Android 스마트폰 마이크 회전(0° Portrait → 90° Landscape)에 따른
스테레오 오디오 변환 AI 모델 학습용 시뮬레이션 데이터셋.
Data Fields
Field
Type
Description
pair_id
int
페어 고유 ID
source_type
str
speech / sine_sweep / white_noise / pink_noise / environmental
audio_0deg_ch0
Audio
Portrait 0° MIC_TOP (48kHz)
audio_0deg_ch1
Audio
Portrait 0° MIC_BOT (48kHz)
audio_90deg_ch0
Audio
Landscape 90° MIC_TOP (48kHz)
audio_90deg_ch1
Audio
Landscape 90° MIC_BOT (48kHz)… See the full description on the dataset page: https://huggingface.co/datasets/haejin1320/mic-rotate-simulation-v1.cost_flow_simulation_seed_merged
Polygon Dynamics 2D
物理推理数据集
项目信息
项目类型: dataset
上传时间: Windows系统
数据集: 包含7种场景类型(A-G)和5个难度等级的物理推理任务
场景类型
A: 基础碰撞检测
B: 重力影响
C: 摩擦力
D: 弹性碰撞
E: 复合物理
F: 高级动力学
G: 极端复杂场景
难度等级
0: 基础 - 简单场景
1: 简单 - 增加复杂度
2: 中等 - 多物体交互
3: 困难 - 复杂物理规则
4: 极端 - 极限测试场景
使用方法
from datasets import load_dataset
# 加载数据集
dataset = load_dataset("competitioncode/cost_flow_simulation_seed_merged")
环境配置
方法1: 使用.env文件(推荐)
在项目根目录创建.env文件:… See the full description on the dataset page: https://huggingface.co/datasets/competitioncode/cost_flow_simulation_seed_merged.unsw-nb15-parquet12.1m_bluesky_postsFair-PP-simulation
Term of use
The datasets and associated code are released under the CC-BY-NC-SA 4.0 license and may only be used for non-commercial, academic research purposes with proper attribution.
doc2desc_13b_chat_descriptions_2kWarcraft-Simulations-5x
Warcraft I — AI vs AI Simulation Dataset
Rust AI-vs-AI Warcraft I engine that
produces deterministic game replays as structured data suitable for ML research, more so for perdictive world model training.
Dataset stats
Stat
Value
Games
5
Total tick rows
61,030 ticks
Ticks per second (in-game)
20
Quick start
from datasets import load_dataset
ds = load_dataset("parquet", data_files="*.parquet", split="train")
# Economy over time… See the full description on the dataset page: https://huggingface.co/datasets/Quazim0t0/Warcraft-Simulations-5x.crisis_prediction_simulationstext_descriptions
(nothing) - in the style of a document randomly taken from the DCLM dataset
textfile: - in the style of a document randomly taken from textfiles.com
llm response: - as if it were the response turn of an llm
llm prompt: - as if it were a prompt to an llm
llm conversation: - as if it were a conversation between user and llm
- crumb, metalure
260408_2_artificial_intelligence_OR_ai_AND_human_AND_team_AND_simulation_OR_interactiondoc2desc_3b_textfile_descriptionsdoc2desc_3b_chat_descriptionsgenz-persona-simulationairtel-mininet-simulationSimulation
