datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ncbi-genbank-complete
Dataset Card for NCBI GenBank Complete
Dataset Summary
GenBank® is the NIH genetic sequence database, an annotated collection of all publicly available DNA sequences. GenBank is part of the International Nucleotide Sequence Database Collaboration (INSDC), which comprises the DNA DataBank of Japan (DDBJ), the European Nucleotide Archive (ENA), and GenBank at NCBI. These three organizations exchange data on a daily basis.
This dataset has been processed into a… See the full description on the dataset page: https://huggingface.co/datasets/pulmo/ncbi-genbank-complete.10Kh-RealOmin-OpenData
Boasting over 13,000 hours of cumulative data and 5 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry.
Update Notes:Stage 3 data upload completed.
13,000+ hours of pure dual-hand data with frame-level alignment latency < 1ms
Full high-precision trajectory reconstruction, breaking the limit of superficial open source, fully ready-to-use
3,000+ contributors and 10,000+ real household scenarios with exceptional diversity… See the full description on the dataset page: https://huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData.genshin-voices-separatedGen-HumanEgo
Gen-HumanEgo
1,800+ hours of egocentric human demonstrations with synchronized, structured supervision for embodied AI and robot learning.
Gen-HumanEgo contains real-world first-person demonstrations collected across diverse tasks, people, environments, and ways of performing activities using a unified six-camera DAS-Ego setup. GenRobot's Data Foundation Model (DFM) processes the recordings to provide complementary training signals for human motion, spatial understanding… See the full description on the dataset page: https://huggingface.co/datasets/genrobot2025/Gen-HumanEgo.code_generation_liteLiveCodeBench is a temporaly updating benchmark for code generation. Please check the homepage: https://livecodebench.github.io/.generic_data_v2general-instruction-augmented-corpora
Instruction Pre-Training: Language Models are Supervised Multitask Learners (EMNLP 2024)
This repo contains the general instruction-augmented corpora (containing 200M instruction-response pairs covering 40+ task categories) used in our paper Instruction Pre-Training: Language Models are Supervised Multitask Learners.
We explore supervised multitask pre-training by proposing Instruction Pre-Training, a framework that scalably augments massive raw corpora with instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/instruction-pretrain/general-instruction-augmented-corpora.GenEvolve-Data-Bench
GenEvolve Data and Bench
This repository contains the open-source data release for GenEvolve:
Config
Directory
Records
Images
Purpose
sft
GenEvolve-Data-SFT/
9,000 trajectories
50,291 reference images
supervised cold-start trajectories
rl
GenEvolve-Data-RL/
3,175 prompts
3,175 GT images
self-evolution / RL training prompts
bench
GenEvolve-Bench/
594 prompts
594 GT images
held-out evaluation benchmarkAll metadata is provided in both JSONL and Parquet. The Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/MeiGen-AI/GenEvolve-Data-Bench.assetsGenFusion_Training_DataGenNOCSgenshin-voice
Genshin Voice
Genshin Voice is a dataset of voice lines from the popular game Genshin Impact.
Hugging Face 🤗 Genshin-Voice
ModelScope Genshin-Voice
Per-speaker downloads are grouped by language and ZIP size. Browse every archive in the ZIP index.
Last update at 2026-08-13
654252 wavs
7291 without speaker (1%)
52693 without transcription (8%)
1088 without inGameFilename (0%)
Dataset Details
Dataset Description
The dataset contains voice lines… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/genshin-voice.GenECGGenECG is an image-based ECG dataset which has been created from the PTB-XL dataset (https://physionet.org/content/ptb-xl/1.0.3/).
The PTB-XL dataset is a signal-based ECG dataset comprising 21799 unique ECGs.
GenECG is divided into the following subsets:
-Dataset A: ECGs without imperfections (Dataset_A_ECGs_without_imperfections) - This subset includes 21799 ECG images that have been generated directly from the original PTB-XL recordings, free from any visual imperfections.
-Dataset B: ECGs… See the full description on the dataset page: https://huggingface.co/datasets/edcci/GenECG.cad-gen-freecad-bench
Parametric CAD Bench — results dataset
Run-by-run results for Parametric CAD Bench, a benchmark that
measures whether AI agents can author editable FreeCAD models from
natural-language part descriptions. 1000 rows, one per
(agent, model, task_id, trial) over the
gnucleus-ai/cad-bench@v1
task suite. The public leaderboard view of this data lives at
cadbench.ai.
What's in here
data/cad-bench-v1.parquet — the row table. Each row carries the
composite + sub-scores… See the full description on the dataset page: https://huggingface.co/datasets/gnucleus-ai/cad-gen-freecad-bench.General-Bench-Closeset
On Path to Multimodal Generalist: General-Level and General-Bench
[📖 Project]
[🏆 Leaderboard]
[📄 Paper]
[🤗 Paper-HF]
[🤗 Dataset-HF]
[📝 Dataset-Github]
Close Set of General-Bench
We divide our General-Bench into two settings: open and close.
This is the Close Set, where we release only the sample inputs—without ground-truth answers—for 🏆 Leaderboard purpose.
To participate the leaderboard, please follow the detailed instructions to submit the evaluation results (submission).… See the full description on the dataset page: https://huggingface.co/datasets/General-Level/General-Bench-Closeset.libero_gen_goal_chain_hdf5GenImageUI-Genie-Agent-16kThis repository contains the Trajectory dataset from the paper UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based
Mobile GUI Agents.
Github: https://github.com/Euphoria16/UI-Genie
social-sim-bench-gensplant-genomic-benchmarkThis dataset comprises the various supervised learning tasks considered in the agro-nt
paper. The task types include binary classification,multi-label classification,
regression,and multi-output regression. The actual underlying genomic tasks range from
predicting regulatory features, RNA processing sites, and gene expression values.GenPoster100K
Dataset Card for GenPoster100K
Dataset Summary
GenPoster-100K is a large-scale dataset for content-aware graphic layout generation introduced in the SEGA paper.
The paper describes it as a high-quality poster dataset with layer-parseable source materials and rich metadata.
This repository provides a Hugging Face datasets loader implementation that reads the source release (BruceW91/GenPoster-100K) and exposes normalized examples with:
poster background image… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/GenPoster100K.genius-lyrics-cleaned
◎
Genius Lyrics Dataset
Cleaned & Deduplicated
🤗 Hugging Face
🤗 Hugging Face
DOI: 10.57967/hf/7978
DOI: 10.57967/hf/7978
revision: 9742989
revision: 9742989
A heavily cleaned, English-only, genre-filtered subset of the Genius Song… See the full description on the dataset page: https://huggingface.co/datasets/theelderemo/genius-lyrics-cleaned.GenImage-arrow
GenImage Arrow
Generator- and split-partitioned Arrow release of the GenImage benchmark.
Each leaf directory is also a standalone Hugging Face save_to_disk bundle.
Contents
Train: 2,581,150 valid images across eight generators.
Test: 100,000 images across eight generators.
Validation is an alias of test because the official GenImage val directory
is the benchmark test set. It is not a third independent split.
Seventeen unavailable or zero-byte upstream train… See the full description on the dataset page: https://huggingface.co/datasets/nebula/GenImage-arrow.gentoomen-lib
Gentoomen Library
The Gentoomen Library is an extensive archive of technology-related resources originally shared on 4chan's /g/ board. It consists of a collection of files and directories covering various topics in computer science and technology.
Here is a Basic Search Engine To Search Your Pdf Book
Overview
Total Size: 32.8GB
Format: Extracted files and directories
Topics Covered:
Algorithms
Scripting
Technology guides
Computer science… See the full description on the dataset page: https://huggingface.co/datasets/thefcraft/gentoomen-lib.libero_gen_spatial_combination_hdf5GenieSimAssets
Key Features 🔑
5000+ objects 1:1 replicated from real-life
Get started 🔥
Download the Assets
To download the full assets, please refer to the official Hugging Face/Modelscope documentation.
Assets Structure
Folder hierarchy
assets
├── background # Background environments
│ └── ...
├── robot # Robot models
│ ├── G1_120s
│ ├── G1_omnipicker
│ ├── G2_omnipicker… See the full description on the dataset page: https://huggingface.co/datasets/agibot-world/GenieSimAssets.ego_gen_comindRADAR-auxiliary-data
RADAR: Preprocessed Anatomical Masks for Merlin CT Data
This dataset provides preprocessed anatomical segmentation masks for the Merlin abdominal CT training set, generated by TotalSegmentator and post-processed for use with the RADAR framework. These masks enable anatomy-aware vision–language pretraining without any additional manual annotation.
Overview
RADAR is a generalist vision–language model trained on over 400,000 contrast-enhanced abdominal CT… See the full description on the dataset page: https://huggingface.co/datasets/radar-generalist/RADAR-auxiliary-data.gen_winograd_raw
gen_winograd
Project: https://ufal.mff.cuni.cz/corefud
Data source: https://github.com/mbzuai-nlp/gen-X/tree/bf1c0adb4b4def03cdf419c18b2948695bc1fab8
Details
English Winograd generated by GPT-4
Citation
@misc{whitehouse2023llmpowered,
title={LLM-powered Data Augmentation for Enhanced Crosslingual Performance},
author={Chenxi Whitehouse and Monojit Choudhury and Alham Fikri Aji},
year={2023},
eprint={2305.14288}… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/gen_winograd_raw.genshin-voice-v3.3-mandarin
Dataset Card for Genshin Voice
Dataset Description
Dataset Summary
The Genshin Voice dataset is a text-to-voice dataset of different Genshin Impact characters unpacked from the game.
Languages
The text in the dataset is in Mandarin.
Dataset Creation
Source Data
Initial Data Collection and Normalization
The data was obtained by unpacking the Genshin Impact game.
Who are the source language producers?
The… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/genshin-voice-v3.3-mandarin.
