datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
godot_rl_FlyByA RL environment called FlyBy for the Godot Game Engine.
This environment was created with: https://github.com/edbeeching/godot_rl_agents
Downloading the environment
After installing Godot RL Agents, download the environment with:
gdrl.env_from_hub -r edbeeching/godot_rl_FlyBy
godot_rl_BallChaseA RL environment called BallChase for the Godot Game Engine.
This environment was created with: https://github.com/edbeeching/godot_rl_agents
Downloading the environment
After installing Godot RL Agents, download the environment with:
gdrl.env_from_hub -r edbeeching/godot_rl_BallChase
godot_rl_FPSA RL environment called FPS for the Godot Game Engine.
This environment was created with: https://github.com/edbeeching/godot_rl_agents
Downloading the environment
After installing Godot RL Agents, download the environment with:
gdrl.env_from_hub -r edbeeching/godot_rl_FPS
godot_rl_RacerA RL environment called Racer for the Godot Game Engine.
This environment was created with: https://github.com/edbeeching/godot_rl_agents
Downloading the environment
After installing Godot RL Agents, download the environment with:
gdrl.env_from_hub -r edbeeching/godot_rl_Racer
BetterDataset-12M
Dataset
Mixed pretraining dataset built from:
Source
Config
Weight
HuggingFaceTB/smollm-corpus
fineweb-edu-dedup
20%
openbmb/Ultra-FineWeb-L3
Ultra-FineWeb-L3-en-Multi-Style-Synthetic
10%
HuggingFaceTB/dclm-edu
—
20%
HuggingFaceFW/finewiki
en
20%
HuggingFaceTB/cosmopedia
stories
2%
HuggingFaceTB/cosmopedia
stanford
2%
HuggingFaceFW/finephrase
all
6%
HuggingFaceTB/finemath
finemath-l4
5%
nampdn-ai/tiny-math-textbooks
—
5%
HuggingFaceTB/cosmopedia… See the full description on the dataset page: https://huggingface.co/datasets/GODELEV/BetterDataset-12M.godseye-violence-detection-dataset
Keypoints-RWF-total Dataset
This dataset is a fusion of three distinct datasets:
RWF-2000: A dataset that includes videos of real-world fights and non-fight scenarios.
Hockey Violence Dataset: A dataset focused on violent interactions in hockey games.
Airtlab Violence Dataset: A dataset containing videos of violent incidents in various settings.
Important Note: We do not own these datasets. By using this dataset, you are agreeing to follow the rules and licensing agreements of… See the full description on the dataset page: https://huggingface.co/datasets/valiantlynxz/godseye-violence-detection-dataset.synthetic-timeseries-data
cruscy data — evaluation sample
Three full days of real crypto market microstructure (Binance spot), prepared
for public evaluation: absolute prices, dates, and the instrument are withheld —
the shape of the day (tick-by-tick relative price, normalized volumes, trade
side, book imbalance) is fully preserved.
The full feed — 27+ streams (raw L2 depth, 1-second trade tape, order-book
metrics, derived features, regime labels) with SQL console, backtest runner and
MCP access for AI… See the full description on the dataset page: https://huggingface.co/datasets/GOD111111111/synthetic-timeseries-data.godot_rl_JumperHardA RL environment called JumperHard for the Godot Game Engine.
This environment was created with: https://github.com/edbeeching/godot_rl_agents
Downloading the environment
After installing Godot RL Agents, download the environment with:
gdrl.env_from_hub -r edbeeching/godot_rl_JumperHard
metaworld_mt50This dataset was created using LeRobot.
Dataset Description
This dataset contains 50 demonstrations per task from the Meta-world simulation benchmarks. Demonstrations are generated using expert policies.
Meta-world: https://arxiv.org/abs/1910.10897
We reposition the camera and flip the rendered images as follow:
Homepage: [More Information Needed]
Paper: [More Information Needed]
License: apache-2.0
Dataset Structure
meta/info.json:
{
"codebase_version":… See the full description on the dataset page: https://huggingface.co/datasets/ML-GOD/metaworld_mt50.godot_rl_3DCarParkingA RL environment called 3DCarParking for the Godot Game Engine.
This environment was created with: https://github.com/edbeeching/godot_rl_agents
Downloading the environment
After installing Godot RL Agents, download the environment with:
gdrl.env_from_hub -r edbeeching/godot_rl_3DCarParking
godot_rl_AirHockeyA RL environment called AirHockey for the Godot Game Engine.
This environment was created with: https://github.com/edbeeching/godot_rl_agents
Downloading the environment
After installing Godot RL Agents, download the environment with:
gdrl.env_from_hub -r edbeeching/godot_rl_AirHockey
DaoZang
DaoZang — 道藏经文语料数据集
《中华道藏》《正统道藏》整理本语料的可加载数据集,与向量库 (ChromaDB) 逐块对应,
供检索评测、微调与 RAG 使用。
数据分片
分片
行数
粒度
字段
train (data/train-00000-of-00006.parquet ~ ...00005-of-00006.parquet, 6 个分片)
285,117
文本块 (chunk)
source / title / chunk_index / chars / text / embedding
标准分片命名 (train-XXXXX-of-00006.parquet),load_dataset 自动合并,无需改动;
每片约 217MB(含 20,000 行一个 row group 与 page index),方便 HF Dataset Viewer
在线浏览(单次扫描上限 ~300MB);
source = 源 Markdown 文件名 (与 ChromaDB 元数据一致);… See the full description on the dataset page: https://huggingface.co/datasets/Godners/DaoZang.cable-damage
Cable Damage
This dataset is part of the Roboflow 100 benchmark, a diverse collection of 100 object detection datasets spanning 7 imagery domains.
Dataset Statistics
Split
Images
Train
919
Validation
265
Test
134
Total
1,318
Classes (2)
break
thunderbolt
Usage
With LibreYOLO
from libreyolo import LIBREYOLO
# Load a model
model = LIBREYOLO(model_path="libreyoloXnano.pt")
# Train on this dataset… See the full description on the dataset page: https://huggingface.co/datasets/Godishala/cable-damage.godot_rl_ShipsA RL environment called Ships for the Godot Game Engine.
This environment was created with: https://github.com/edbeeching/godot_rl_agents
Downloading the environment
After installing Godot RL Agents, download the environment with:
gdrl.env_from_hub -r edbeeching/godot_rl_Ships
godot_rl_ItemSortingCartA RL environment called ItemSortingCart for the Godot Game Engine.
This environment was created with: https://github.com/edbeeching/godot_rl_agents
Downloading the environment
After installing Godot RL Agents, download the environment with:
gdrl.env_from_hub -r edbeeching/godot_rl_ItemSortingCart
godot-gdscript-dataset
Godot GDscript Code Dataset
This dataset contains GDScript code from 5k+ github repositories. Data from each repo has been extracted into a text file. Each text file contains the code from all .gd files & README.md text (if the README was not empty in the original repo).
Original forum post:
https://diffused.to/Thread-Godot-GDscript-Code-Dataset-5k
Dataset collection date
June 2025
Dataset structure:
📂 files/
├── repo-name-1.txt
├── repo-name-2.txt… See the full description on the dataset page: https://huggingface.co/datasets/wallstoneai/godot-gdscript-dataset.Godzilla-Mono-Melodies
Godzilla Mono Melodies
654k+ select monophonic melodies with accompaniment and drums from Godzilla MIDI dataset
Installation and use
Load dataset
#===================================================================
from datasets import load_dataset
#===================================================================
godzilla_mono_melodies = load_dataset('asigalov61/Godzilla-Mono-Melodies')
dataset_split = 'train'… See the full description on the dataset page: https://huggingface.co/datasets/asigalov61/Godzilla-Mono-Melodies.godot_rl_DownFallA RL environment called DownFall for the Godot Game Engine.
This environment was created with: https://github.com/edbeeching/godot_rl_agents
Downloading the environment
After installing Godot RL Agents, download the environment with:
gdrl.env_from_hub -r edbeeching/godot_rl_DownFall
behavior_3ENERGY_DATAgodot_rl_VirtualCameraA RL environment called VirtualCamera for the Godot Game Engine.
This environment was created with: https://github.com/edbeeching/godot_rl_agents
Downloading the environment
After installing Godot RL Agents, download the environment with:
gdrl.env_from_hub -r edbeeching/godot_rl_VirtualCamera
Pinpoint-Bench
Pinpoint-Bench: A Zero-Hint Evaluation Frontier for Active Visual Search
Project Page | Paper (PixelEyes)
Pinpoint-Bench is a specialized evaluation benchmark designed to assess the active visual search and reasoning capabilities of Multimodal Large Language Models (MLLMs) and active perception agents within ultra-high-resolution images. It explicitly addresses the saturation of traditional benchmarks, the lack of spatial annotations, and the vulnerability of models to… See the full description on the dataset page: https://huggingface.co/datasets/godx7/Pinpoint-Bench.BetterDataset-2M
Dataset
Mixed pretraining dataset built from:
Source
Weight
Notes
epfml/FineWeb-HQ
60%
Quality-filtered web text (score > 0.8)
HuggingFaceTB/cosmopedia
20%
Sequential: stanford → wikihow → web_samples_v2
HuggingFaceTB/finemath (finemath-4plus)
10%
Mathematical reasoning
bigcode/python-stack-v1-functions-filtered
10%
Python code functions
Total rows: 2,000,000Shards: 20 parquet filesSplit: all rows are in train
Around - 1.59B Tokens
GOD_Coder_Complete_DataSet
GOD_Coder_Complete_DataSet
Subtitle
A large-scale complete-project coding dataset by gss1147 / WithIn Us AI, built to train language models into stronger professional software-engineering assistants.
Dataset Summary
GOD_Coder_Complete_DataSet is a large synthetic supervised fine-tuning dataset designed to help turn a general language model into a professional complete-project AI coder.
The dataset focuses on teaching models how to:
diagnose… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/GOD_Coder_Complete_DataSet.godot_4_docsDataset generated for Godot 4 docs using Glaive.
GodVerb📚 GigaVerbo Clean — Portuguese Text Cleaning & Dedup Pipeline
Author: Tiago Loeblein
Version: v76
Status: Active — production-grade cleaning pipeline for large-scale PT datasets
License: Same as the source datasets (inherited)
🧠 Overview
GigaVerbo Clean é um dataset processado a partir do TucanoBR/GigaVerbo, contendo bilhões de tokens em português.
Este repositório não é o dataset original, mas sim uma contribuição independente oferecendo:
pipeline robusto de limpeza,
deduplicação exata via… See the full description on the dataset page: https://huggingface.co/datasets/tiagoloeblein/GodVerb.aopoli-lv-libero_combined_no_noops_lerobot_v21This dataset was created using LeRobot.
Dataset Description
Combined version of following lerobot datasets
aopolin-lv/libero_spatial_no_noops_lerobot_v21
aopolin-lv/libero_object_no_noops_lerobot_v21
aopolin-lv/libero_goal_no_noops_lerobot_v21
aopolin-lv/libero_10_no_noops_lerobot_v21
Homepage: [More Information Needed]
Paper: [More Information Needed]
License: apache-2.0
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1"… See the full description on the dataset page: https://huggingface.co/datasets/godnpeter/aopoli-lv-libero_combined_no_noops_lerobot_v21.godeater
Bangumi Image Base of God Eater
This is the image base of bangumi GOD EATER, we detected 23 characters, 1589 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the characters'… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/godeater.science-datalake
Science Data Lake
A unified, portable science data lake integrating 7 scholarly datasets (~525 GB Parquet) with cross-dataset DOI normalization, 13 scientific ontologies (1.3M terms), and a reproducible ETL pipeline.
Note: One additional source (Semantic Scholar S2AG) is supported by the pipeline but is not redistributed here due to its API terms of service. See Not Included in This Upload below.
What's Unique
This dataset enables queries… See the full description on the dataset page: https://huggingface.co/datasets/GodotCN/science-datalake.gravitational-waves-strain
