datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
awesome-loop-engineering
Awesome Loop Engineering Dataset
A structured dataset of 1022 papers, official docs, tools, benchmarks, patterns, critiques, and implementation guides for recurring AI-agent systems.
Resource Atlas ·
GitHub field guide ·
Resource selection ·
Report a correction
Dataset Summary
Each row connects an original source to its contribution, novelty, impact, publication details, lifecycle stages, audience, evidence type, link status, and… See the full description on the dataset page: https://huggingface.co/datasets/cy0307/awesome-loop-engineering.loopcrafttennis_rpy_closed_loop
Tennis RPY closed-loop experiment
当前 RPY VLA 在 Isaac Sim 中进行网球接球实验的部署资料、场景和实验侧代码快照。模型代码、checkpoint、模型配置和模型缓存不在此仓库内,由部署机器已有模型提供。此仓库以 HF dataset 类型保存实验复现资料,不是新的训练数据集。
文件
simulation_bundle.tar.gz:场景USD/3DGS/网格/球资源、仿真入口和当前场景脚本。
experiment/:异步评测、模型服务、RPY接口、IK/FK及运动学参数。
config_template.json:成功运行配置的模板;用部署脚本替换旧机器路径。
configure_deployment.py:解压后设置新机器的仿真、模型、权重及Python路径。
docs/MIGRATION_ZH.md:完整文件清单、依赖、路径和控制流程说明。
docs/CLOSED_LOOP_RESULT_ZH.md:20条闭环实验配置、结果和限制。… See the full description on the dataset page: https://huggingface.co/datasets/BruceJee/tennis_rpy_closed_loop.loop7kaimenoakuyakureijouwamototekikokudejiyuukimamanahanayomeseikatsuwomankitsusuru
Bangumi Image Base of Loop 7-kaime No Akuyaku Reijou Wa, Moto Tekikoku De Jiyuu Kimama Na Hanayome Seikatsu Wo Mankitsu Suru
This is the image base of bangumi Loop 7-kaime no Akuyaku Reijou wa, Moto Tekikoku de Jiyuu Kimama na Hanayome Seikatsu wo Mankitsu suru, we detected 62 characters, 3609 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/loop7kaimenoakuyakureijouwamototekikokudejiyuukimamanahanayomeseikatsuwomankitsusuru.LOOPerSet
LOOPerSet: A Large-Scale Dataset for Data-Driven Polyhedral Optimization
Dataset at a Glance
LOOPerSet is a corpus of 28 million labeled compilation traces designed for machine learning research in compilers and systems. It maps synthetically generated loop nests and complex optimization sequences to ground-truth execution times measured on physical hardware. Transformation sequences were generated using a polyhedral compilation framework to ensure they… See the full description on the dataset page: https://huggingface.co/datasets/Mascinissa/LOOPerSet.loopmoe-fineweb-edu-100bt-mcore
LoopMoE FineWeb-Edu 100BT (Megatron indexed)
Release status: complete
This is the exact pretokenized FineWeb-Edu 100BT corpus prepared for the
reviewed LoopMoE M1 Dense/Loop/Dual and Fast-Slow experiment contracts.
This release does not define
the M2--M5 train schedules. It contains 64 training indexed-dataset shards and
one fixed validation shard. Each prefix has a Megatron .bin/.idx pair and
a sanitized .stats.json record. There is no separate test split.
Split
Documents… See the full description on the dataset page: https://huggingface.co/datasets/gaotang/loopmoe-fineweb-edu-100bt-mcore.loop_metal_20260902_132356This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
7
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_yaw.pos"… See the full description on the dataset page: https://huggingface.co/datasets/makermods/loop_metal_20260902_132356.Miles-SFT-SFT-LoopSFT-public-backup-20260913
Public SFT and LoopSFT baseline backup
The owner explicitly allowed these SFT and LoopSFT baselines to remain public. This repository includes their three checkpoints (steps 1122, 2252, 3125), bundle configuration/tokenizer/gate files, and final native optimizer archives. No RCGA, GA or CLT2 artifacts are included. Original model: Qwen/Qwen3-4B-Base, revision 906bfd4b4dc7f14ee4320094d8b41684abff8539 (Apache-2.0). Training data: inclusionAI/Ling-Coder-SFT 400k subset… See the full description on the dataset page: https://huggingface.co/datasets/LaurelWings/Miles-SFT-SFT-LoopSFT-public-backup-20260913.LoopNavLoopsBench
LoopsBench
This dataset repository hosts published LoopsBench task bundles. A LoopsBench task is a self-contained evaluation package for long-horizon terminal coding: it includes an agent-visible workspace snapshot, unit-level requirements, dependency graphs, Docker execution metadata, public verifier files, and reference gold patches used by maintainers and Oracle-style validation.
The files in this dataset are release artifacts mirrored from the latest LoopsBench GitHub… See the full description on the dataset page: https://huggingface.co/datasets/LoopsBench/LoopsBench.grab_to_loopThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 21705,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Temmp1e/grab_to_loop.qwen3-4b-only-loop0LoopRLloopshapeW2H-Basic-Agent-Loop-w-Sandbox
W2H Basic Agent Loop with built in Linux Sandbox
A lightweight home agent that talks, runs code and takes actions in the real world. Access it from anywhere.
This is a vanilla Python agent loop that supports tools, skills, a microVM sandbox, encrypted data-in-transit and the Arduino microcontroller. The web UI includes voice, file uploads and slash commands. Designed for learning and experimentation. Use vibe coding to adapt it for different tasks.
Talk to the agent from… See the full description on the dataset page: https://huggingface.co/datasets/vbookshelf/W2H-Basic-Agent-Loop-w-Sandbox.msap-loop-ctl-20260728relay-ko-s2-scratch-g4-150h-R0-arelay-ko-s2-scratch-g4-53h-R0-arelay-ko-s2-scratch-g4-100h-R0-brelay-ko-s2-scratch-g4-53h-R0-brelay-ko-s2-scratch-g4-100h-R0-arelay-ko-s2-scratch-g4-150h-R0-bloopwan-opensora-pilot-v1
LoopWan Open-Sora-Plan pilot
Status: completed bounded curation. Counts: {"long_audit": 22, "train": 2000, "val": 128}.
Fixed 320x480, timestamp sampling at 16 FPS; train/validation crops are real
contiguous 10-second shots, audit crops 20 seconds. Sources are disjoint and
captions are matched to pinned official annotations. See DATASET_REPORT.md for
filter thresholds, caption limitations and full provenance.
Official dataset revision: ab77293def393e6938f11a7bfd12163decfb9620.… See the full description on the dataset page: https://huggingface.co/datasets/Nicholas0228/loopwan-opensora-pilot-v1.LoopTF-SudokuLOOP_SEATTLE
LOOP_SEATTLE (TsFile)
Apache TsFile version of the LOOP_SEATTLE subset of
GIFT-Eval.
Overview
GIFT-Eval is a benchmark for general time-series forecasting, covering 23 datasets
(≈144,000 series and 177M data points) across seven domains, ten frequencies, and a
range of forecast horizons. This repository contains a single subset of that benchmark.
LOOP_SEATTLE — Loop-detector traffic data from the Seattle freeway network.
The .tsfile files are organized by the… See the full description on the dataset page: https://huggingface.co/datasets/THULab/LOOP_SEATTLE.Loop_Closure_forest_harddaily-paper-2026-07-23-autonomous-loop-completion-gap
Closing the Completion Gap in Unattended Autonomous Agent Loops: An Empirical Ablation of Verification Gates, Checkpoint-Rollback, and Stall Escalation
TL;DR — A controlled 2x2x2 factorial ablation (30 seeds x 300 tasks, CRN simulation) shows that combining verification gates, checkpoint-rollback, and stall escalation closes 71.4 pp of the autonomous-agent completion gap. The verification gate is overwhelmingly dominant (55.3 pp), checkpoint-rollback is a strong super-additive… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-23-autonomous-loop-completion-gap.loop-llmcircle-packing-insight-loop
Circle-Packing Insight-Exploration Loop
Artifacts from an iterative GPT solver <-> proposer insight-exploration loop on the
21-circles-in-a-perimeter-4-rectangle packing problem (AlphaEvolve SOTA sum-of-radii
= 2.3658321334167627). Each round, 16 solvers propose a program + written explanation;
every program is scored; a proposer then mines all 16 attempts into an evolving insight
document that conditions the next round. Run: 16 solvers x 8 rounds.
Subsets… See the full description on the dataset page: https://huggingface.co/datasets/ars22/circle-packing-insight-loop.LoopTool-23k
Overview
LoopTool is a fully automatic, model-aware iterative framework that tightly couples data generation and model training for tool-augmented LLM learning
The LoopTool-2w is released as part of Closing the Data–Training Loop for Robust LLM Tool Calls
The dataset comprises 23,040 tool-call samples, involving 20,813 APIs. In each sample, the instruction contains the corresponding set of available tools for that sample; the input corresponds to the dialogue history of the… See the full description on the dataset page: https://huggingface.co/datasets/zhangkangning/LoopTool-23k.
