datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
generated-csvsphysical-ai-bench-conditional-generation
Physical AI Bench - Conditional Generation
Paper | Code
This dataset (Phsical AI benchmark, PAI-Bench) consisting of 600 examples across three key scenarios: robotic arm operations, driving, and ego-centric everyday life scenes, each representing a critical aspect of Physical AI. This dataset is constructed by sampling a number of videos from three different datasets. The specific details are provided below.
Dataset
Category
Sample Nums
Agibot World
Robotics
200
OpenDV… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-conditional-generation.genhome3d-1280
GenHome3D-1280
1,280 validated household and spatial-design assets in USDZ format, organized
across 64 categories.
Explore the visual catalog ·
Browse the GitHub repository ·
Download the versioned release ·
Read the generation method
Dataset summary
Assets
1,280
Categories
64
Assets per category
20
Runtime format
USDZ
Units
Meters
Asset license
CC BY 4.0
Technical validation
1,280/1,280 pass
Package validation
1… See the full description on the dataset page: https://huggingface.co/datasets/linxy97/genhome3d-1280.GenIaC-SecBench
GenIaC-SecBench
A benchmark for evaluating the security of LLM-generated Infrastructure-as-Code
(IaC) against a size-matched human baseline.
Paper: Compared to What? A Human-Anchored Security Benchmark for LLM-Generated
Infrastructure-as-Code (arXiv:2608.28021)
Code: https://github.com/AnimeshShaw/GenIaC-SecBench
Why this dataset exists
Prior evaluations of generated IaC report vulnerability counts for models
only. Stating that a model averages eight findings per… See the full description on the dataset page: https://huggingface.co/datasets/AnimeshShaw/GenIaC-SecBench.genebench-pro-public-package
GeneBench-Pro Public Case Studies
This repository contains public GeneBench-Pro case studies. It is the
self-contained package intended for public distribution, including Hugging Face
publication.
Package Layout
<repo-root>/
├── .gitattributes
├── README.md
├── LICENSE
├── problems.csv
├── checksums.sha256
├── manifest.json
├── reference_definitions.md
├── reference_grader.py
└── problems/
└── <eval_id>/
├── eval_config.json
├── data_files/… See the full description on the dataset page: https://huggingface.co/datasets/openai/genebench-pro-public-package.genius-song-lyricsgeneralization-science-dataDNA_Gen
Citation
Please cite our work using the bibtex below:
BibTeX:
@article{su2025language,
title={Language Models for Controllable DNA Sequence Design},
author={Su, Xingyu and Li, Xiner and Lin, Yuchao and Xie, Ziqian and Zhi, Degui and Ji, Shuiwang},
journal={arXiv preprint arXiv:2507.19523},
year={2025}
}
program_generation_v3casp14-casp15-cameo-test-proteinsrna-downstream-tasks
GB.RNA Benchmark Datasets
mRNA related tasks
Translation efficiency prediction from Chu et al.(2024) [1]
3 cell lines: Muscle, pc3, HEK
input sequence: 5'UTR
10-fold cross-validation split
mRNA expression level prediction from Chu et al.(2024) [1]
3 cell lines: Muscle, pc3, HEK
input sequence: 5'UTR
10-fold cross-validation split
Mean ribosome load prediction from Sample et al. (2019) [2]
input sequence: 5'UTR
ouput: mean ribosome load
the original data… See the full description on the dataset page: https://huggingface.co/datasets/genbio-ai/rna-downstream-tasks.fake-real-newsgenter-ajibawa-name-filled
GENTER Ajibawa Name-Filled
This dataset expands aieng-lab/genter-ajibawa by inserting concrete names into each template.
It provides nested Hugging Face configs with 1, 2, 5, or 10 names per gender and template.
For every template sentence, names are sampled independently from NAMEXACT (matching split; frequency-weighted), using K female and K male names in config nK.
It is intended for experiments that need concrete text rather than [NAME]/[MASK] placeholders while still… See the full description on the dataset page: https://huggingface.co/datasets/aieng-lab/genter-ajibawa-name-filled.human_ai_generated_text
Human or AI-Generated Text
The data can be valuable for educators, policymakers, and researchers interested in the evolving education landscape, particularly in detecting or identifying texts written by Humans or Artificial Intelligence systems.
File Name
model_training_dataset.csv
File Structure
id: Unique identifier for each record.
human_text: Human-written content.
ai_text: AI-generated texts.
instructions: Description of the task given to both Humans and… See the full description on the dataset page: https://huggingface.co/datasets/dmitva/human_ai_generated_text.genebench-pro-public-package
GeneBench-Pro Public Case Studies
This repository contains public GeneBench-Pro case studies. It is the
self-contained package intended for public distribution, including Hugging Face
publication.
Package Layout
<repo-root>/
├── .gitattributes
├── README.md
├── LICENSE
├── problems.csv
├── checksums.sha256
├── manifest.json
├── reference_definitions.md
├── reference_grader.py
└── problems/
└── <eval_id>/
├── eval_config.json
├── data_files/… See the full description on the dataset page: https://huggingface.co/datasets/ajh-oai/genebench-pro-public-package.geneb-tasks
GENEB — Genomic Embedding Benchmark (task data)
Task-level sequence classification data for GENEB, a multi-task benchmark for DNA sequence encoders introduced in the paper: GENEB: Why Genomic Models Are Hard to Compare.
Paper: https://huggingface.co/papers/2606.04525
Source code: GitHub - darlednik/GENEB
Leaderboard: Hugging Face Space
GENEB evaluates frozen representations from 40 genomic foundation models across 100 tasks in 13 functional categories using a unified… See the full description on the dataset page: https://huggingface.co/datasets/darlednik/geneb-tasks.imdb-genres
Dataset Card for IMDb Movie Dataset: All Movies by Genre
Dataset Summary
This dataset is an adapted version of "IMDb Movie Dataset: All Movies by Genre" found at: https://www.kaggle.com/datasets/rajugc/imdb-movies-dataset-based-on-genre?select=history.csv.
Within the dataset, the movie title and year columns were combined, the genre was extracted from the seperate csv files, the pre-existing genre column was renamed to expanded-genres, any movies missing a description… See the full description on the dataset page: https://huggingface.co/datasets/jquigl/imdb-genres.Gender-Indicators-For-African-Countries
Gender Indicators For African Countries | Africa (World Health Organization)
Size category: 1K<n<10K - Formats: csv - Sector: demographics_social - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Gender-Indicators-For-African-Countries.program_generation_v5GenMRP
GenMRP: A Generative Multi-Route Planning Framework for Efficient and Personalized Real-Time Industrial Navigation
This is the dataset for our paper.
The following table contains the feature dimensions and key features of our dataset.
Feature Type
Interpretation
Shape
Some Key Features
Link Features
Includes the road segment attributes
K * 2 * N
Link lengthLink Lane width
Frequency Features
Logs the user's travel history within the past three months
K * 2 * 10 * 7
Delta… See the full description on the dataset page: https://huggingface.co/datasets/GD-ML/GenMRP.ProteinGYM-DMS-zeroshotAI-and-Human-Generated-Text
AI & Human Generated Text
I am Using this dataset for AI Text Detection for https://exnrt.com.
Check Original DataSet GitHub Repository Here: https://github.com/panagiotisanagnostou/AI-GA
Description
The AI-GA dataset, short for Artificial Intelligence Generated Abstracts, comprises abstracts and titles. Half of these abstracts are generated by AI, while the remaining half are original. Primarily intended for research and experimentation in natural language… See the full description on the dataset page: https://huggingface.co/datasets/Ateeqq/AI-and-Human-Generated-Text.scientific-verification
Scientific Verification Benchmark: NMC Cathodes
Dataset summary
The benchmark contains 50 scientific claims about NMC (lithium nickel manganese cobalt oxide) battery cathodes. Each claim is answered by Claude Opus 5, GPT 5.6 Luna and Gemini 3.1 Pro using a set of 20 open-access papers, producing 150 scored answers. The accompanying reference set contains 1,991 experiment-grounded measurements curated from 227 open-access papers, with experimental conditions and… See the full description on the dataset page: https://huggingface.co/datasets/GenData-Research/scientific-verification.transcript_isoform_expression_prediction
Multi-modal transcript isoform expression dataset
We curated the human transcript isoform expression dataset from the GTEx portal following the preprocessing pipeline in Garau-Luis et al. (2024). We downloaded the RNA-seq Transcript TPMs file from the bulk tissue expression in GTEx Analysis V8. The table contains transcript expression collected from 30 non-diseased tissues in nearly 1000 human individuals. We averaged the transcript expression measurements across individuals to… See the full description on the dataset page: https://huggingface.co/datasets/genbio-ai/transcript_isoform_expression_prediction.tgk-ai-video-generators-2026
Permanent dataset archive: https://doi.org/10.5281/zenodo.22703594
Five AI Video Generators Tested on Dialogue, Action and an Advert
These Guys Know tested Seedance 2.5, MiniMax H3, FLUX 3 Video, Gemini Omni 1.1 Flash and HappyHorse 1.1 on 1 September 2026. Every model received the same three ten-second, 16:9 text-to-video tasks: a father interrupting a computer game, a three-person fight inside a fixed hotel lobby and a Mango Cola advert with an exact product name.
We retained… See the full description on the dataset page: https://huggingface.co/datasets/These-Guys-Know/tgk-ai-video-generators-2026.diffusion-mcqa-gen-pelatnas-2026
Which Prompt Made This? — Generated Edition
Pelatnas IOAI 2026 · Task Diffusion MCQA (varian trajectory)
Sebuah model text-to-image sedang bekerja. Di tengah prosesnya, gambar belum
menjadi gambar — yang ada hanya latent ter-noise: tensor 4 × 64 × 64 berisi
campuran struktur yang mulai muncul dan derau Gaussian.
Kali ini kalimat itu harfiah. Latent yang kamu terima benar-benar diambil dari
tengah proses generate: sebuah trajectory denoising DDIM 50 langkah dihentikan
sejenak… See the full description on the dataset page: https://huggingface.co/datasets/fassabilf/diffusion-mcqa-gen-pelatnas-2026.GenAssocBiasThis work has been accepted to ACL 2024. You can find the paper related to this work https://arxiv.org/abs/2309.08902
Dataset Summary
GenAssocBias is a dataset that measures stereotype bias in LLMs. GenAssocBias consists of 11,940 sentences that measure model preferences across ageism, beauty, beauty_profession, nationality, and institutional bias.
Supported Tasks and Leaderboards
multiple-choice question answering
Languages
English (en)
Dataset Descriptions:
There are 8 columns in our… See the full description on the dataset page: https://huggingface.co/datasets/mozaman36/GenAssocBias.ivypanda-llm-generated-essays
AI-Generated Essays Dataset
This dataset contains AI-generated academic essays created using the models:
Mistral 7B Instruct v0.2 (Q5_K_M quantized)
Temperature: 0.7
Max tokens: 4096
Top-p: 0.9 (default)
Top-k: 40 (default)
Repeat penalty: 1.1 (default)
Context window: 32768 tokens
Llama 3 13B Instruct v0.1 (Q5_K_M quantized)
Temperature: 0.7
Max tokens: 4096
Top-p: 0.9 (default)
Top-k: 40 (default)
Repeat penalty: 1.1 (default)
Context window: 8192 tokens
DeepSeek-V3.2
API… See the full description on the dataset page: https://huggingface.co/datasets/artfultom/ivypanda-llm-generated-essays.InterviewForge_GenDS
Synthetic Data Generation
Model & Infrastructure
The dataset was generated using the mistral:latest Large Language Model running locally via the Ollama framework. This model was explicitly selected because it balances advanced reasoning capabilities with hardware efficiency, allowing the execution of 10,944 complex generation requests entirely locally on an RTX 3080 GPU without incurring API costs. Additionally, Mistral demonstrated exceptional reliability in… See the full description on the dataset page: https://huggingface.co/datasets/Davichick/InterviewForge_GenDS.1k_stories_100_genre
Dataset Documentation
Overview
This dataset contains 1000 stories spanning 100 different genres. Each story is represented in a tabular format using a dataframe. The dataset includes unique IDs, titles, and the content of each story.
Genre List
The list of all genres can be found in the genres.txt file.
reading genre_list variable
with open('story_genres.pkl', 'rb') as f:
story_genres = pickle.load(f)
Sample of genre list:
genres = ['Sci-Fi', 'Comedy'… See the full description on the dataset page: https://huggingface.co/datasets/FareedKhan/1k_stories_100_genre.
