datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
generative-sound-masking-generated-energy-v7
Generative Sound Masking — fixed-background audio masking
Incrementally generated unfiltered candidates. This is not a final selected dataset.
Each background has 15 separately generated prompt–seed outputs using gain-compensated reconstruction residuals. run_config.json pins models, source pools, parameters and implementation hashes.
For multiple workers read workers/worker-NN/progress.json; each worker reports only its assigned IDs.
Global completion requires all worker… See the full description on the dataset page: https://huggingface.co/datasets/AE-W/generative-sound-masking-generated-energy-v7.generative-sound-masking-input-noise-full-v1
Generative Sound Masking input-noise pool v1
This WebDataset contains 48,840 mono 16-kHz, 10.24-second input-noise clips
across 49 tar shards. It combines the complete Yiming SONYC, TAU Urban
Acoustic Scenes, and UrbanSound baseline with subject-balanced BABYCRY-UJM-AXA
and NOTSOFAR-1 train windows. Stable sample metadata are in
metadata/noise_index.jsonl; JSON beside each WAV adds hashes computed during
packaging.
The source datasets carry different licenses. In particular… See the full description on the dataset page: https://huggingface.co/datasets/AE-W/generative-sound-masking-input-noise-full-v1.generative-sound-masking-generated-energy-v3-smokechamp_trainning_sample
Dataset samples for Champ trainning
This dataset samples is used for Champ.
Before trainning, you need to process the datasets by SMPL & DWPOSE methods. Refer to https://github.com/fudan-generative-vision/champ/blob/master/docs/data_process.md
genvsr-video-benchmarksgen-games-v9-video-pilot1gen-games-v7-video-pilot1gen-games-v8-video-pilot1news-unmasked
Dataset Card for "news-unmasked"
More Information needed
generative-sound-masking-generated-energy-v4
Generative Sound Masking — fixed-background audio masking
Incrementally generated unfiltered candidates. This is not a final selected dataset.
Each background has 15 separately generated prompt–seed outputs using gain-compensated reconstruction residuals. run_config.json pins models, source pools, parameters and implementation hashes.
For multiple workers read workers/worker-NN/progress.json; each worker reports only its assigned IDs.
Global completion requires all worker… See the full description on the dataset page: https://huggingface.co/datasets/AE-W/generative-sound-masking-generated-energy-v4.hallo3_training_dataHallo3: Highly Dynamic and Realistic Portrait Image Animation with Diffusion Transformer Networks
Jiahao Cui1
Hui Li1
Yun Zhan1
Hanlin Shang1
Kaihui Cheng1
Yuqi Ma1
Shan Mu1
Hang Zhou2
Jingdong Wang2
Siyu Zhu1✉️
1Fudan University 2Baidu Inc
I. Dataset Overview
This dataset serves as the training data for the open - source Hallo3 model, specifically created for the training of video… See the full description on the dataset page: https://huggingface.co/datasets/fudan-generative-ai/hallo3_training_data.llama8b-layer15-sae-probes
Llama8B Sparse Probing Activations
This repository contains activation data accompanying the paper Learning a Generative Meta-Model of LLM Activations.
Project page: https://generative-latent-prior.github.io
Code: https://github.com/g-luo/generative_latent_prior
Quick Start
With this data, you can evaluate GLPs via sparse probing.
The activations are derived from the binary classification datasets from Kantamneni et. al., 2025.
The activations are taken only from… See the full description on the dataset page: https://huggingface.co/datasets/generative-latent-prior/llama8b-layer15-sae-probes.gen-games-v10-video-pilot1Generative_Coding_Datasetgenerative-adapter-datagenerative-papers-arxiv
arXiv Citation Embeddings Dataset
This dataset contains citation embeddings for arXiv papers, designed for training models that predict cited paper embeddings from citation context.
Dataset Structure
Core Files (Required for Training)
paper_embeddings.parquet: Paper-level sentence embeddings (shared across splits)
train/citations.jsonl: Training citation pairs with contexts
train/citation_embeddings_*.parquet: Citation context embeddings (sharded)… See the full description on the dataset page: https://huggingface.co/datasets/akalmbach-kinsol/generative-papers-arxiv.Updesh_GenerativeMMLU-Redux-2.0-Generativegen-games-v5-video-test2fineweb-edu-ar
FineWeb-Edu-Ar
FineWeb-Edu-Ar is a machine-translated Arabic version of the FineWeb-Edu dataset designed to support the development of Arabic small language models (SLMs).
Dataset Details:
Languages: Arabic, English (paired)
Size: 202 billion tokens
License: CC-BY-NC-4.0
Source: Machine-translated from the deduplicated version of Hugging Face’s FineWeb-Edu dataset
Translation model: facebook/nllb-200-distilled-600M
Application:
FineWeb-Edu-Ar is suitable for pre-training… See the full description on the dataset page: https://huggingface.co/datasets/kaust-generative-ai/fineweb-edu-ar.champ_motions_example
Example data for Champ inference
Links
github: https://github.com/fudan-generative-vision/champ
models: https://huggingface.co/fudan-generative-ai/champ
GENERATIVE
Stable Diffusion web UI
A web interface for Stable Diffusion, implemented using Gradio library.
Features
Detailed feature showcase with images:
Original txt2img and img2img modes
One click install and run script (but you still must install python and git)
Outpainting
Inpainting
Color Sketch
Prompt Matrix
Stable Diffusion Upscale
Attention, specify parts of text that the model should pay more attention to
a man in a ((tuxedo)) - will pay more attention to tuxedo
a man in… See the full description on the dataset page: https://huggingface.co/datasets/crystantine/GENERATIVE.conditional-generative-models-max-phasetelco-gaia
Telco-GAIA
A GAIA-style benchmark for AI agents operating over a real telecom operator's
website snapshot plus a synthetic customer database. 100 tasks across 7
categories: Pricing, Miscellaneous, Images, Web Archives, PDF, PDF Visual,
Database.
Agents read questions.json + environment.md, browse the local website
(:8080) and query the database API (:8081), and produce a GAIA-compatible
submission.json.
What's here
File
What… See the full description on the dataset page: https://huggingface.co/datasets/kaust-generative-ai/telco-gaia.generative-sound-masking-retrieval-dasheng-v1
Generative Sound Masking — DaSheng retrieval v1
This dataset stores exhaustive top-15 retrieval from the 237,500 prompt–seed
candidate audios for 48,840 background audios. One row per background contains
15 ranked candidates, including prompt, seed, clip ID, audio SHA256 and cosine score.
Source audio stays in the two original public datasets; no audio is duplicated here.
This is an incremental run. Read progress.json before treating it as complete.
The data/train split is an… See the full description on the dataset page: https://huggingface.co/datasets/AE-W/generative-sound-masking-retrieval-dasheng-v1.maestro-mas-benchmark
MAESTRO MAS Benchmark Dataset
maestro-mas-benchmark is a dataset derived from MAESTRO, a framework-agnostic evaluation suite for LLM-based
multi-agent systems (MAS). It provides a systems-level view of MAS behavior and is designed to benchmark,
observe, and analyze MAS performance and behavior across diverse scenarios.
For more details about MAESTRO, visit the GitHub repository.
Dataset details
The dataset currently includes data for 12 different MAS systems… See the full description on the dataset page: https://huggingface.co/datasets/kaust-generative-ai/maestro-mas-benchmark.conll2003-generative
Dataset Card for CoNLL-2003 with NER Workflow Enhancements
This dataset is a modified version of the CoNLL-2003 dataset, enhanced to support an LLM-based Named Entity Recognition (NER) workflow. Two new columns, sentence and entities, have been added. You can find the code used to generate this version together with the data files.
Named Entities
As in the original CoNLL-2003 task, this dataset focuses on four types of named entities:
Persons (PER)
Locations (LOC)… See the full description on the dataset page: https://huggingface.co/datasets/areias/conll2003-generative.generative-ai-red-teaming
About this dataset
This dataset is an unofficial transformed clone of the Generative AI Red-Teaming
(GRT) dataset created by Humane
Intelligence (HI). This dataset collates
findings from the Generative AI Red-Teaming Challenge conducted at AI Village
within DEFCON 31. It is provided as part of HI's inaugural algorithmic bias
bounty.
The original lives on
GitHub at:
humane-intelligence/bias-bounty-data
Differences
This version of the GRT dataset differs from the original… See the full description on the dataset page: https://huggingface.co/datasets/jinnovation/generative-ai-red-teaming.anatomy-generative-prototyping
The Anatomy of a Generative Prototyping Session, Study 1
Anonymous authors, CHI 2027 submission. The dataset accompanies a preregistered study in
which 100 non-experts completed three ten-minute prototyping tasks with Mark, a
generative prototyping environment: two from a blank page, in an unfamiliar and in a familiar
domain, and one on an existing prototype. Every generation is dissected into the formulation
of the prompt (abstraction and specificity), the extent of the… See the full description on the dataset page: https://huggingface.co/datasets/backtothehuman/anatomy-generative-prototyping.generative-native-ads
Webis Generated Native Ads 2024
Dataset Summary
This dataset was created to train ad blocking systems on the task of identifying advertisements in responses of conversational search engines.
There are two dataset dictionaries available:
responses.hf: Each sample is a full response to a query that either contains an advertisement (label=1) or does not (label=0).
sentence_pairs.hf: Each sample is a pair of two sentences taken from the responses. If one of them… See the full description on the dataset page: https://huggingface.co/datasets/webis/generative-native-ads.
