datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Anime-Background-Finetuning-V1.1
Anime-Background-Finetuning (10143 manually curated by hand images from danbooru and reddit collections)
The dataset contain roughly 2k of anime Screencap data and 8k of scrapped danbooru illustration data.
This is the proccessed version of the dataset meant to be used for my personal finetuning practice project, please visit my RicemanT/Background-Finetuning repo for the raw unprocessed data that you can process yourself.
The dataset have two minor type of processing being done… See the full description on the dataset page: https://huggingface.co/datasets/RicemanT/Anime-Background-Finetuning-V1.1.Anime-Background-Finetuning-V1.1
Anime-Background-Finetuning (10143 manually curated by hand images from danbooru and reddit collections)
The dataset contain roughly 2k of anime Screencap data and 8k of scrapped danbooru illustration data.
This is the proccessed version of the dataset meant to be used for my personal finetuning practice project, please visit my RicemanT/Background-Finetuning repo for the raw unprocessed data that you can process yourself.
The dataset have two minor type of processing being done… See the full description on the dataset page: https://huggingface.co/datasets/HappyHenAi/Anime-Background-Finetuning-V1.1.Anime-Background-Finetuning-Unprocessed
Anime-Background-Dataset (10143 manually curated by hand images from danbooru and reddit collections)
The dataset contain roughly 2k of Screencap data and 8k of scrapped danbooru illustration data.
It is all raw unprocessed data, the illust folder contain scrapped danbooru tags sidecar .txt on most of the images, while the screencap have non. The processed data is being worked on a seperate repo (Anime-Background-Finetuning)
chorus-backgroundsCAIMAN-ASR-BackgroundNoise
Dataset Card for Myrtle/CAIMAN-ASR-BackgroundNoise
This dataset provides background noise audio, suitable for noise augmentation
while training Myrtle.ai's CAIMAN-ASR models.
Dataset Details
Dataset Description
Curated by: Myrtle.ai
License: Myrtle.ai's modifications to the source data are licensed under
the CC BY 4.0 license.
Some of the original data is under the CC BY 3.0 license; the rest is in the public domain.
Please see the Source Data section… See the full description on the dataset page: https://huggingface.co/datasets/Myrtle/CAIMAN-ASR-BackgroundNoise.gpn-star-p-uniform-v1-background
marin-dna/gpn-star-p-uniform-v1-background
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the background region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-background.phylop-uniform-v1-background
marin-dna/phylop-uniform-v1-background
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the background region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses the pipeline's pinned phyloP conservation… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/phylop-uniform-v1-background.openve3m_background_changemusicai-background-music-audio-llm-benchmark
Does Background Music Matter to Speech in Pre-trained Language Models
The completed September 2026 study covers 8 model families, 55 instrumental recordings, and 10 evaluation settings. It studies how adding background music to the same spoken question changes model responses.
Latest release and artifact guide
Technical report PDF
Complete LaTeX project
LaTeX GitHub repository
Matrices, figures, and supporting data
Regenerated speech and mixtures: 550 archives / 250,800… See the full description on the dataset page: https://huggingface.co/datasets/Elfsong/musicai-background-music-audio-llm-benchmark.dota-backgroundCaltech101_not_background_test
Dataset Card for "Caltech101_not_background_test"
More Information needed
common_voice_22_yue_w_background_captionMerged JackyHoCL/urban-noise-uganda-61k-caption, OpenSound/AudioCaps
TODO: convert to MP3, reduce size
heao-1-a2-spectrum-backgrounds
HEAO-1 A2 Spectrum Backgrounds
The HEAO-1 A2 spec_back directory serves 1,506 .bck and 36 .alt pointed-phase files from the MED, HED-1 and HED-3 detectors.
Use
from datasets import load_dataset
ds = load_dataset("astro-legacy-archive/heao-1-a2-spectrum-backgrounds", "a2_h1l_1058s084_po.bck", split="train")
row = ds[0]
print(row["CHANNEL"], row["COUNTS"])
1 0
The example configuration is a2_h1l_1058s084_po.bck. Configuration names are the source filename stems.… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/heao-1-a2-spectrum-backgrounds.openve3m_background_change_refeval_svla_15b_red_backgroundThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 15,
"total_frames": 7370,
"total_tasks": 5,
"total_videos": 30,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:15"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/shuohsuan/eval_svla_15b_red_background.vertebrate-v1-background
marin-dna/vertebrate-v1-background
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This
draft covers the background region cohort with all species
scope and preserves source FASTA/2bit letter case.
Anchor eligibility uses the pipeline's pinned phyloP conservation filter.
Sequence case is independent of that filter: lowercase bases preserve source
repeat masking, uppercase bases preserve source… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-background.AI-Background-RemoverMarket1501-Background-Modified
Dataset Card for "Market1501-Background-Modified"
Dataset Summary
The Market1501-Background-Modified dataset is a variation of the original Market1501 dataset. It focuses on reducing the influence of background information by replacing the backgrounds in the images with solid colors, noise patterns, or other simplified alternatives. This dataset is designed for person re-identification (ReID) tasks, ensuring models learn person-specific features while ignoring background… See the full description on the dataset page: https://huggingface.co/datasets/ideepankarsharma2003/Market1501-Background-Modified.person-background-dataset
Person-Background Dataset
Diverse people placed in various scene backgrounds. Generated with FLUX.1-dev (people) and FLUX.1-Kontext (background editing).
How It Is Collected
The collect.py script:
Stage 1 – People: Generates 10 diverse people with FLUX.1-dev (different ethnicities, genders, ages).
Stage 2 – Backgrounds: Uses FLUX.1-Kontext img2img to edit each person into scene categories (forest, beach, office, etc.). Preserves person identity while changing only the… See the full description on the dataset page: https://huggingface.co/datasets/nirmalendu01/person-background-dataset.Caltech101_not_background_train
Dataset Card for "Caltech101_not_background_train"
More Information needed
background
Background Noise Dataset
This dataset contains 3 audio recordings of 2 different background noise classes.
Dataset Statistics
Total Audio Files: 3
Total Classes: 2
Format: WAV (AudioFolder with metadata.jsonl)
Classes and Descriptions
The dataset covers the following background noises:
Label
Description
fan_noise
Fan noise background
white_noise
White noise background
Structure
The dataset is organized in a folder structure… See the full description on the dataset page: https://huggingface.co/datasets/sdialog/background.hep-signature-backgrounds
HEP Signature Backgrounds
This dataset contains generated high-energy-physics examples for mapping signal-region final-state signatures to Standard Model background compositions.
Contents
hep_sft/train.parquet, hep_sft/val.parquet, hep_sft/test.parquet: supervised fine-tuning splits.
Answer Schema
This is a conversational prompt-completion SFT dataset, following the Hugging Face/TRL convention:
{
"id":… See the full description on the dataset page: https://huggingface.co/datasets/ho22joshua/hep-signature-backgrounds.dior-backgroundrealsense-black-green-background-lerobot
record-immitation-blue-arm-realsense-black-green-backgroundMerged
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
wider_face_backgroundopenarm_visuomotor_VR_pringles_V14_background_30hzeval_svla_15b_black_backgroundThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 1,
"total_frames": 856,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/shuohsuan/eval_svla_15b_black_background.cube-detection-monoply-background-obb
Cube Detection on Monopoly Board Background (OBB)
A small oriented-bounding-box (OBB) detection dataset of colored cubes placed on a Movensys "Monopoly" board background. Intended for fine-tuning YOLO-style OBB detectors used in pick-and-place / robotic manipulation pipelines.
Classes
ID
Name
0
green_cube
1
yellow_cube
2
blue_cube
3
red_cube
Splits
Split
Images
Labels
train
104
104
val
29
29
test
16
16
total
149
149… See the full description on the dataset page: https://huggingface.co/datasets/movensys/cube-detection-monoply-background-obb.Background_INCONTEXTImagePulseV2-Edit-Background
ImagePulseV2 Dataset - Background Replacement
The ImagePulseV2 dataset is a collection we constructed for training the Diffusion Templates series of models. It consists of multiple subsets generated using models such as Z-Image-Turbo, Qwen-Image, and Qwen-Image-Edit, based on prompts randomly sampled from DiffusionDB.
Open-source code: DiffSynth-Studio
Technical report: arXiv
Project homepage: GitHub
Documentation: English Version, Chinese Version
Online demo: ModelScope Studio… See the full description on the dataset page: https://huggingface.co/datasets/DiffSynth-Studio/ImagePulseV2-Edit-Background.
