datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GenAI-Bench
GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation
Baiqi Li1*, Zhiqiu Lin1,2*, Deepak Pathak1, Jiayao Li1, Yixin Fei1, Kewen Wu1, Tiffany Ling1, Xide Xia2†, Pengchuan Zhang2†, Graham Neubig1†, and Deva Ramanan1†.
1Carnegie Mellon University, 2Meta
Links:
📖Paper | | 🏠Home Page | | 🔍GenAI-Bench Dataset Viewer | 🏆Leaderboard|
🗂️GenAI-Bench-1600(ZIP format) | | 🗂️GenAI-Bench-Video(ZIP format) | |… See the full description on the dataset page: https://huggingface.co/datasets/BaiqiL/GenAI-Bench.genai-image-tag-db
GenAI Image Tag DB (cc0-1.0)
This repository contains the cc0-1.0 build of the tag database.
The main artifact is the SQLite database. The parquet_danbooru/ directory is a derived export so the Hugging Face Dataset Viewer can preview a subset of rows (Danbooru-only).
Files
genai-image-tag-db-cc0.sqlite: SQLite database
parquet_danbooru/*.parquet: Parquet export for Dataset Viewer
build_manifest.json: Build manifest (revisions and stats)
report/: Source effects… See the full description on the dataset page: https://huggingface.co/datasets/NEXTAltair/genai-image-tag-db.genai-image-tag-db-CC4
GenAI Image Tag DB (cc-by-4.0)
This repository contains the cc-by-4.0 build of the tag database.
The main artifact is the SQLite database. The parquet_danbooru/ directory is a derived export so the Hugging Face Dataset Viewer can preview a subset of rows (Danbooru-only).
Files
genai-image-tag-db-cc4.sqlite: SQLite database
parquet_danbooru/*.parquet: Parquet export for Dataset Viewer
build_manifest.json: Build manifest (revisions and stats)
report/: Source… See the full description on the dataset page: https://huggingface.co/datasets/NEXTAltair/genai-image-tag-db-CC4.genai-image-tag-db-mit
GenAI Image Tag DB (mit)
This repository contains the mit build of the tag database.
The main artifact is the SQLite database. The parquet_danbooru/ directory is a derived export so the Hugging Face Dataset Viewer can preview a subset of rows (Danbooru-only).
Files
genai-image-tag-db-mit.sqlite: SQLite database
parquet_danbooru/*.parquet: Parquet export for Dataset Viewer
build_manifest.json: Build manifest (revisions and stats)
report/: Source effects and health… See the full description on the dataset page: https://huggingface.co/datasets/NEXTAltair/genai-image-tag-db-mit.sms-spam-enriched
SMS Spam Enriched Dataset
An enriched version of the classic SMS Spam Collection Dataset from UC Irvine with additional engineered features and semantic embeddings.This dataset is designed for spam detection, feature engineering experiments, and model interpretability research.
Dataset Overview
Total samples: 5,171
Classes:
0: Ham (non-spam)
1: Spam
Enrichments Added
Alongside the raw SMS text (sms) and labels (label), we engineered multiple new… See the full description on the dataset page: https://huggingface.co/datasets/GenAIDevTOProd/sms-spam-enriched.papercli-papers
AI Conference & Journal Papers
Searchable metadata and full-text PDF mirrors for papers from top-tier AI venues (NeurIPS, ICML, ICLR, CVPR, ICCV, ECCV, WACV, ACL, EMNLP, NAACL, IJCAI, AAAI, JMLR, Interspeech) from 2023.
📊 papers.parquet: The complete dataset containing all fields and all venues.
🔍 Per-venue browse views: Easily explore specific subsets by selecting a venue in Subset and a year in Split.
🏗️ Dataset Structure & Storage Strategy
To avoid… See the full description on the dataset page: https://huggingface.co/datasets/GenAI4ELab/papercli-papers.genai-bench-sdxl-latents
GenAI-Bench SDXL with Latent Trajectories
SDXL generations for the GenAI-Bench prompts, paired with
the per-step decoded-latent trajectory of each generation. This dataset was created to evaluate NoisyCLIP.
For every prompt, 10 images were generated (1600 prompts → 16,000 generations). Each generation provides:
the full-resolution final image,
the 50-step denoising trajectory (each step's latent decoded to a small preview image), and
a CLIP-FlanT5-XXL VQAScore measuring… See the full description on the dataset page: https://huggingface.co/datasets/asiimo/genai-bench-sdxl-latents.hleGenAI-EdSent
📄 Associated Paper
Title: Unveiling User Perceptions in the Generative AI Era: A Sentiment-Driven Evaluation of AI Educational Apps' Role in Digital Transformation of E-Teaching
Authors: Adeleh Mazaheriyan (Islamic Azad University) & Erfan Nourbakhsh (University of Isfahan)
Abstract: This study performs a sentiment-driven evaluation of user reviews from 22 top AI ed-apps on the Google Play Store to assess efficacy, challenges, and pedagogical implications. Our pipeline leverages… See the full description on the dataset page: https://huggingface.co/datasets/Erfan-Nourbakhsh/GenAI-EdSent.us-genai-court-opinions
US court decisions on generative AI
566 court-authored US documents (opinions, orders, concurrences, dissents, administrative orders) that address generative AI or the fabricated authorities courts associate with it — ai_mention says whether the court itself names AI — each coded by topic, with the public-domain passage quoted, plus 9 legal-AI litigation dockets with dated milestones.
Built 2026-09-08 by SafeLegalAI (Cognesio LLP). Canonical pages: safelegalai.com/courts ·… See the full description on the dataset page: https://huggingface.co/datasets/safelegalaidata/us-genai-court-opinions.multimodal-vqa-self-instruct-enriched
Multimodal VQA – Self-Instruct-enriched
Overview
This dataset is an enriched, cleaned, and metadata-enhanced version of zwq2018/Multi-modal-Self-instruct.It pairs images with natural language questions and answers, making it ideal for Vision-Language Model (VLM) training, benchmarking, and instruction tuning.
Dataset Summary
Total samples: 75,000+ (64,796 train, 11,193 test)
Modalities: Image + Text (Questions) + Text (Answers)
Task Types: Visual Question… See the full description on the dataset page: https://huggingface.co/datasets/GenAIDevTOProd/multimodal-vqa-self-instruct-enriched.GenAI-MVS
GenAI Multiple Video Synchronization (GenAI-MVS) Dataset
Overview
This dataset contains video clips of five different exercise types with frame-level annotations. The dataset is designed for temporal action classification and exercise form analysis tasks.
Dataset Statistics
Total Videos: 82
Total Frames: 8,029
Classes: 5 (bench_press, deadlift, dips, pullups, pushups)
Splits: Training (54 videos) and Validation (28 videos)
Annotation Format: Binary frame-level… See the full description on the dataset page: https://huggingface.co/datasets/AvihaiNaam/GenAI-MVS.RationalRewards-EvalData-GenAIBench-MMRB2-ERBenchTLDR: this is the RewardModel Evaluation dataset for text-to-image generation and image editing, from the following paper.
RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time
Haozhe Wang1
Cong Wei2
Weiming Ren2
Jiaming Liu3
Fangzhen Lin1
Wenhu Chen2
1 HKUST
2 University of Waterloo
3 Alibaba
RationalRewards is a… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/RationalRewards-EvalData-GenAIBench-MMRB2-ERBench.GenAI-Benchgen_ai_2024
Dataset Card for "gen_ai_2024"
More Information needed
genaibootcampGenAI-job-postings-Dataset
Dataset Card for "GenAI-job-postings-Dataset"
More Information needed
knowledge_base_genaineutral-sd-outputsgenai-gdelt-framing
GenAI Governance Framing in Online News (GDELT 2.0, 2022–2026)
This dataset accompanies the paper:
Framing Generative AI Governance in Online News: A Longitudinal Analysis of 1.1 Million Articles (2022–2026)Brewen Couaran, Yuvraj Singh Pathania, Arjun Rajesh Nair · 2026PDF: paper/paper.pdf · Code: github.com/brewcoua/GenAI-GDELTCompanion site: brewcoua.github.io/GenAI-GDELT
Dataset overview
1,116,091 online news articles from the GDELT 2.0 Global Knowledge Graph… See the full description on the dataset page: https://huggingface.co/datasets/brewcoua/genai-gdelt-framing.genai-ml-2025-hw7-evalgenai-hw1-russian-asr-correctionrewardzero-t2i-hpd-genaiNIFTY-feature-enhanced
NIFTY-Feature-Enhanced
Dataset Summary
NIFTY-Feature-Enhanced is a multi-modal, finance-focused dataset built on top of raeidsaqur/NIFTY
.
We enrich the original dataset with structured financial indicators, derived signals, temporal features, sentiment scores, embeddings, and event tags.
This makes it suitable for:
Predictive ML models (e.g., XGBoost, LSTMs, Transformers)
Financial NLP tasks (sentiment, RAG, semantic search)
Multi-modal research (numeric + textual… See the full description on the dataset page: https://huggingface.co/datasets/GenAIDevTOProd/NIFTY-feature-enhanced.genai4t_spx_generationimdb_Sentimental_Finetuned_20Kdatabricks-dolly15k-semantic-complexity
Databricks - Dolly 15k – Enriched Variant (Instruction-Tuned with Semantic and Complexity Augmentation)
Overview
This dataset is a semantically enriched and complexity-aware extension of the original Databricks Dolly 15k, purpose-built for evaluating and training instruction-following models. Each sample is augmented with additional signals to enable more nuanced filtering, curriculum learning, and benchmark development across diverse NLP tasks.
Dataset Format
Each… See the full description on the dataset page: https://huggingface.co/datasets/GenAIDevTOProd/databricks-dolly15k-semantic-complexity.GenAI-Arena-Bench-3GENAI-M3taskmaster_v1_enriched
Taskmaster-1 Enriched Dialog Dataset (Combined)
Overview
This dataset is a combined, enriched version of the self_dialog and woz_dialog splits from the Taskmaster-1 dataset. It consists of multi-turn, human-human and human-simulated conversations with systematic enhancements for machine learning workflows—especially dialog modeling, generation, and fine-grained evaluation.
All conversations are structured in a JSON format with consistent schema and include added semantic… See the full description on the dataset page: https://huggingface.co/datasets/GenAIDevTOProd/taskmaster_v1_enriched.
