datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cad-gen-freecad-bench
Parametric CAD Bench — results dataset
Run-by-run results for Parametric CAD Bench, a benchmark that
measures whether AI agents can author editable FreeCAD models from
natural-language part descriptions. 1000 rows, one per
(agent, model, task_id, trial) over the
gnucleus-ai/cad-bench@v1
task suite. The public leaderboard view of this data lives at
cadbench.ai.
What's in here
data/cad-bench-v1.parquet — the row table. Each row carries the
composite + sub-scores… See the full description on the dataset page: https://huggingface.co/datasets/gnucleus-ai/cad-gen-freecad-bench.shofo-tiktok-general-small
Shofo TikTok General (Small)
Overview
Shofo TikTok General (Small) is a dataset containing 50,000 TikTok videos with comprehensive metadata, transcripts, comments, and engagement metrics. This is a curated subset of Shofo's larger TikTok index, which contains hundreds of millions of indexed videos.
Size: ~50K videos (~500GB)
Modality: Video + Audio + Text (transcripts, comments, captions)
Source: TikTok
Schema
Column
Type
Description
file_name… See the full description on the dataset page: https://huggingface.co/datasets/Shofo/shofo-tiktok-general-small.Genius-song-lyrics-cleaned
🎵 Genius Song Lyrics cleaned Dataset
Dataset Description
This dataset is originally taken from Genius Song Lyrics and it contains cleaned and normalized song lyrics for more than 5 million songs, designed for large-scale topic modeling, clustering, and semantic analysis.
The dataset was specifically preprocessed to be compatible with embedding-based models (e.g. Sentence Transformers, BERTopic) while preserving lyrical meaning and thematic content.
Repetitive structures… See the full description on the dataset page: https://huggingface.co/datasets/Dr3dre/Genius-song-lyrics-cleaned.deception-activationsGeneralThought-430K-filteredData from https://huggingface.co/datasets/GeneralReasoning/GeneralThought-430K, removed prompts with non commercial data
GMAI-VL-5.5M
GMAI-VL-5.5M Dataset
GMAI-VL-5.5M is a comprehensive, large-scale medical General Medical AI Vision-Language (GMAI-VL) dataset built specifically for training multimodal foundation models in the medical domain. It contains an extraordinary scale of high-quality instructions encompassing over 5.5 million multimodal question-answering pairs, carefully constructed based on hundreds of medical classification, segmentation, and detection datasets.
This repository… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-VL-5.5M.pretrain_data_eukaryote
GENERator-v2-Eukaryote Gene-Centric Pretraining Corpus
This repository provides the gene-centric pretraining corpus underlying GENERator-v2-Eukaryote, a large-scale DNA language model for eukaryotic genome understanding.
The dataset is constructed by leveraging RefSeq annotations to extract biologically meaningful functional genomic regions, which serve as the foundation for large-context DNA language model pretraining.
📌 Dataset Construction Overview
The core… See the full description on the dataset page: https://huggingface.co/datasets/GenerTeam/pretrain_data_eukaryote.generated-csvsaging-gene-expression-single-cell-mouse
A single-cell transcriptomic atlas characterizes ageing tissues in the mouse
https://www.nature.com/articles/s41586-020-2496-1#Sec2
Code to download and process this dataset is available in: https://github.com/seanome/2025-longevity-x-ai-hackathon
Dataset structure is originally from AnnData.
Descriptions of each data file is below.
Data Files
This dataset contains multiple parquet files, one for each sheet in the original Excel file:… See the full description on the dataset page: https://huggingface.co/datasets/longevity-db/aging-gene-expression-single-cell-mouse.cad-gen-freecad-bench-v2
Parametric CAD Bench v2 — results dataset
Run-by-run results for Parametric CAD Bench v2, a benchmark that measures
whether AI agents can author editable FreeCAD models from natural-language part
descriptions. This archive contains 1,000 rows: one trial for each of 100 tasks
across the 10 public jobs on the live
gnucleus-ai/cad-bench@v2
leaderboard.
What's in here
data/cad-bench-v2.parquet — the trial index. Each row carries the
continuous reward and its… See the full description on the dataset page: https://huggingface.co/datasets/gnucleus-ai/cad-gen-freecad-bench-v2.GeneralThoughtArchive
GeneralThought-430K
Thought wants to be free
Open reasoning data for March 14 2025. This dataset was part of a side-project in the weeks following the R1 release by Chengxi and Ross - we are no longer maintaining this
dataset but are archiving it here.
The dataset contains questions, reference answers, reasoning traces, final answers and other metadata from several popular reasoning models including DeepSeek-R1, DeepSeek-R1-Zero, OpenThoughts-32B, LIMO… See the full description on the dataset page: https://huggingface.co/datasets/RJT1990/GeneralThoughtArchive.CulturaY-ja-askllm-v1
CulturaY-ja-askllm-v1
多言語データセット ontocord/CulturaY の日本語パート ja に対して、 Ask-LLM 手法でスコア付けしたデータセットです。
元データセットのカラムに加え askllm_score というカラムが追加されており、ここに Ask-LLM のスコアが格納されています。
Ask-LLM でスコア付けに使用した LLM は Rakuten/RakutenAI-7B-instruct で、プロンプトは以下の通りです。
###
{data}
###
Does the previous paragraph demarcated within ### and ### contain informative signal for pre-training a large-language model? An informative datapoint should be well-formatted, contain some usable knowledge of the world, and strictly… See the full description on the dataset page: https://huggingface.co/datasets/geniacllm/CulturaY-ja-askllm-v1.steady-rans-generalization
Steady-RANS cross-family generalization dataset
Data for the paper "Towards generalized flow field prediction: one model across unseen
object families" (under double blind review; this account is anonymous for that reason).
Trained checkpoints and evaluation code are in the companion model repo:
steady-rans-surrogates.
Steady incompressible k-omega SST (OpenFOAM simpleFoam) external flow around 855 distinct
shapes (17 scripted parametric families plus 40 ModelNet object… See the full description on the dataset page: https://huggingface.co/datasets/BlidReview/steady-rans-generalization.genhome3d-1280
GenHome3D-1280
1,280 validated household and spatial-design assets in USDZ format, organized
across 64 categories.
Explore the visual catalog ·
Browse the GitHub repository ·
Download the versioned release ·
Read the generation method
Dataset summary
Assets
1,280
Categories
64
Assets per category
20
Runtime format
USDZ
Units
Meters
Asset license
CC BY 4.0
Technical validation
1,280/1,280 pass
Package validation
1… See the full description on the dataset page: https://huggingface.co/datasets/linxy97/genhome3d-1280.GenIaC-SecBench
GenIaC-SecBench
A benchmark for evaluating the security of LLM-generated Infrastructure-as-Code
(IaC) against a size-matched human baseline.
Paper: Compared to What? A Human-Anchored Security Benchmark for LLM-Generated
Infrastructure-as-Code (arXiv:2608.28021)
Code: https://github.com/AnimeshShaw/GenIaC-SecBench
Why this dataset exists
Prior evaluations of generated IaC report vulnerability counts for models
only. Stating that a model averages eight findings per… See the full description on the dataset page: https://huggingface.co/datasets/AnimeshShaw/GenIaC-SecBench.sib200-LexC-Gen
Dataset Card for sib200-LexC-Gen
Dataset Summary
The LexC-Gen dataset for SIB-200 topic classification task is a dataset generated for low-resource languages at scale with Large Language Models (BLOOMZ-7.1B) and Gatitos bilingual lexicons.
from datasets import load_dataset
dataset = load_dataset("BatsResearch/sib200-LexC-Gen", "gn_100k")
Supported Tasks and Leaderboards
text-classification, topic-classification: The dataset can be used to train a model… See the full description on the dataset page: https://huggingface.co/datasets/BatsResearch/sib200-LexC-Gen.genebench-pro-public-package
GeneBench-Pro Public Case Studies
This repository contains public GeneBench-Pro case studies. It is the
self-contained package intended for public distribution, including Hugging Face
publication.
Package Layout
<repo-root>/
├── .gitattributes
├── README.md
├── LICENSE
├── problems.csv
├── checksums.sha256
├── manifest.json
├── reference_definitions.md
├── reference_grader.py
└── problems/
└── <eval_id>/
├── eval_config.json
├── data_files/… See the full description on the dataset page: https://huggingface.co/datasets/openai/genebench-pro-public-package.genius-song-lyricsOpenOneRec-General-Pretrain
通用文本数据集
本目录包含 OpenOneRec 项目使用的通用文本数据集信息。这些数据集均来自 HuggingFace,经过清洗处理和对齐到项目统一的数据格式,并转换为 Parquet 格式用于训练。
数据格式说明
所有数据集均已转换为统一的 Parquet 格式,符合项目的数据格式规范(参考 ../README.md)。数据格式支持:
Segments 格式:用于普通文本数据,使用 segments 字段存储文本段落列表
Chat 格式:用于对话数据,使用 messages 字段存储对话消息列表
每个 Parquet 文件包含以下核心字段:
uuid: 唯一标识符
source: 数据来源标识
metadata: JSON 格式的元数据字典
segments 或 messages: 文本内容(根据数据类型选择)
详细的数据格式规范请参考 ../README.md。
数据集列表
数据集名称
样本数量
HuggingFace 仓库
reasoning_v1_20m
1,666… See the full description on the dataset page: https://huggingface.co/datasets/OpenOneRec/OpenOneRec-General-Pretrain.generalization-science-dataAITW_Generalgeneral-master-en-202608
General · Master · English · 2026-08
English pretraining text, assembled from three public sources, cleaned with one
character-level cleaner, and filtered for repetition.
109,337,531 documents and 468,064,046,462 characters.
Composition
Config
Documents
Characters
What it is
fineweb-edu-dedup
65,010,430
297,544,916,118
Web text an educational classifier kept
cosmopedia-v2
38,591,146
144,011,993,012
Synthetic prose from a seeded generator… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-master-en-202608.human_genomecell2sentence4longevity-data
Dataset Card: longevity-genie/cell2sentence4longevity-data
Summary
This repository contains preprocessed single-cell RNA-seq (scRNA‑seq) datasets prepared as “cell sentences” for training and evaluation of cells2sentence-style models. Each cell is represented as a space‑separated sequence of top expressed gene symbols, enabling language‑model style training for tasks such as biological age prediction and other downstream applications.
This dataset targets fine‑tuning and… See the full description on the dataset page: https://huggingface.co/datasets/longevity-genie/cell2sentence4longevity-data.DNA_Gen
Citation
Please cite our work using the bibtex below:
BibTeX:
@article{su2025language,
title={Language Models for Controllable DNA Sequence Design},
author={Su, Xingyu and Li, Xiner and Lin, Yuchao and Xie, Ziqian and Zhi, Degui and Ji, Shuiwang},
journal={arXiv preprint arXiv:2507.19523},
year={2025}
}
SLT-Task2-Post-ASR-Speaker-Tagging
Dataset Name: Dataset for ASR Speaker-Tagging Corrections (Speaker Diarization)
Description
This dataset is pairs of erroneous ASR output and speaker tagging, which are generated from a ASR system and speaker diarization system.
Each source erroneous transcription is paired with human-annotated transcription, which has correct transcription and speaker tagging.
SEGment-wise Long-form Speech Transcription annotation (SegLST), the file format used in the CHiME challenges… See the full description on the dataset page: https://huggingface.co/datasets/GenSEC-LLM/SLT-Task2-Post-ASR-Speaker-Tagging.midashenglm-gen-training-latents
ModelsLab/midashenglm-gen-training-latents
Precomputed audio latents for fine-tuning
mispeech/midashenglm-gen,
paired with six-view prompts in the exact format the model was trained on.
This is not an audio dataset and not a caption dataset. Each record is the
output of the model's frozen DashengTokenizer encoder — 768-dimensional latents
at 25 Hz, stored float16 — next to the tagged prompt string built from the
source metadata.
Why it exists
The encoder is frozen… See the full description on the dataset page: https://huggingface.co/datasets/ModelsLab/midashenglm-gen-training-latents.gender-by-name
Dataset Card for "Gender-by-Name"
This dataset attributes first names to genders, giving counts and probabilities. It combines open-source government data from the US, UK, Canada, and Australia. The dataset is taken from UCI Machine Learning Repository
Dataset Information
This dataset combines raw counts for first/given names of male and female babies in those time periods, and then calculates a probability for a name given the aggregate count. Source datasets are from… See the full description on the dataset page: https://huggingface.co/datasets/erickrribeiro/gender-by-name.rna-downstream-tasks
GB.RNA Benchmark Datasets
mRNA related tasks
Translation efficiency prediction from Chu et al.(2024) [1]
3 cell lines: Muscle, pc3, HEK
input sequence: 5'UTR
10-fold cross-validation split
mRNA expression level prediction from Chu et al.(2024) [1]
3 cell lines: Muscle, pc3, HEK
input sequence: 5'UTR
10-fold cross-validation split
Mean ribosome load prediction from Sample et al. (2019) [2]
input sequence: 5'UTR
ouput: mean ribosome load
the original data… See the full description on the dataset page: https://huggingface.co/datasets/genbio-ai/rna-downstream-tasks.llama-3b-gold-15M-student-generations_SNIS_2048_tune422v1
