datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
laions_got_talent_enhanced_flash_annotations_and_long_captionslaions_got_talent_enhanced_just_flash_annotationsEnhanced-MedMNIST
Cross-Dimensional Evaluation Datasets
Transfer learning in machine learning models, particularly deep learning architectures, requires diverse datasets to ensure robustness and generalizability across tasks and domains. This repository provides comprehensive details on the datasets used for evaluation, categorized into 2D and 3D datasets. These datasets span variations in image dimensions, pixel ranges, label types, and unique labels, facilitating a thorough assessment of… See the full description on the dataset page: https://huggingface.co/datasets/convergedmachine/Enhanced-MedMNIST.relight-five-model-comparison-enhanced-assets
Relight Five-Model Comparison Enhanced-Prompt Assets
This dataset contains 200 relighting comparison samples rerun with Qwen3-VL-8B enhanced editing prompts.
Each row includes the reference image, original generated image, original prompt, enhanced prompt, and outputs from five image-editing models.
Dataset: https://huggingface.co/datasets/marcuskwan/relight-five-model-comparison-enhanced-assets
Models:
Qwen-Image-Edit-2511
LightX2V-Qwen-Lightning
Step1X-Edit-v1p2
IC-Light… See the full description on the dataset page: https://huggingface.co/datasets/marcuskwan/relight-five-model-comparison-enhanced-assets.LibriTTS-Enhanced
LibriTTS Enhanced Dataset
Enhanced version of LibriTTS dataset for speech enhancement research.
zoonomia-v1-v4_ccre_noexon_enhancer
bolinas-dna/zoonomia-v1-v4_ccre_noexon_enhancer
A curated enhancer training set for issue
#326 — a de-contaminated
derivation of the v4 ccre_non_promoter arm of
bolinas-dna/zoonomia-v1-v1,
built by the
snakemake/zoonomia_projection_dataset pipeline at commit
6b320c268547.
Provenance
This subset is v4_ccre_noexon further restricted to enhancer-dominant windows (dELS+pELS basepair coverage ≥ the other non-PLS cCRE classes), population-matching the val_enhancer… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_noexon_enhancer.laions_got_talent_enhanced_no_metadatagpn-star-p-uniform-v1-enhancer-arm-a
marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a.speech-enhancement-dfn-16kphylop-uniform-v1-enhancer-arm-a
marin-dna/phylop-uniform-v1-enhancer-arm-a
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses the pipeline's pinned phyloP conservation… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/phylop-uniform-v1-enhancer-arm-a.functional-enhancer
marin-dna/functional-enhancer
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and official version-matched UCSC hg38-to-target liftOver chains.
This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
For the 28 non-mammalian targets, the stable… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/functional-enhancer.vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered
marin-dna/vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered
Review status: draft generated for issue #473 review before upload.
Human-anchored 255 bp vertebrate sequences for the
ccre_enhancer_centered cohort under the full_window policy.
The policy projects the complete 255 bp human window, applies the established 128--512 bp compatible-fragment gate, and resizes around the accepted target-span midpoint. Human anchors come from the fixed exp351 ENCODE dELS/pELS-centered… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered.task249_enhanced_wsc_pronoun_disambiguation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task249_enhanced_wsc_pronoun_disambiguation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task249_enhanced_wsc_pronoun_disambiguation.DeepSTARR-enhancer-activity
Abouts
The enhancer activity data is sourced from the DeepSTARR repo.
We have applied minor formatting adjustments to the dataset to facilitate streamlined data analysis.
How to use
from datasets import load_dataset
datasets = load_dataset("GenerTeam/DeepSTARR-enhancer-activity")
PubMedVision-enhancedenhanced-audiosnippets-long-2-8M
Enhanced Audiosnippets Long 2.8M
Enhanced version of mitermix/audiosnippets_long_2_8M with speech enhancement, emotion annotations, speaker embeddings, and comprehensive metadata analysis.
Dataset Summary
Metric
Value
Total samples
2,633,037
Total audio hours
4,932 h
Duration range
3.0s - 1124.3s
Mean duration
6.7s
Audio format
WAV, 48kHz mono
Tar files
1,410
Processing Pipeline
Each audio sample was processed through:
Speech… See the full description on the dataset page: https://huggingface.co/datasets/ai-music4you3/enhanced-audiosnippets-long-2-8M.zoonomia-v1-v4_ccre_noexon_enhancer-order
bolinas-dna/zoonomia-v1-v4_ccre_noexon_enhancer-order
The bolinas-dna/zoonomia-v1-v4_ccre_noexon_enhancer cross-mammal
training set, restricted to a species cohort: one representative species per NCBI order — 19 deeply-diverged placental mammals (every pair separated by ~tens of millions of years), versus the implicit-default 108 family-deduplicated species. A strict subset of the family set, so it reuses the v1 cross-mammal projection unchanged (no re-halLiftover).
Same human… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_noexon_enhancer-order.EnhancementDetection_LibrittsTrainClean360Wham
Dataset Card for "EnhancementDetection_LibrittsTrainClean360Wham"
More Information needed
voxforge_spanish_enhanced
VoxForge Spanish Enhanced (CleanUNet + FlashSR)
Dataset Summary
This dataset is a processed and enhanced version of the Spanish subset of:
VoxForge.
Furthermore, as this is a personal project, we give no guarantees that the audio is completely clean from any artifacts or noise the CleanUNet model could not remove.
However, we have personally tested the corpus via the fine-tuning of some SOTA speech models and the results have been satisfactory.
It has been created to… See the full description on the dataset page: https://huggingface.co/datasets/ebellob/voxforge_spanish_enhanced.night-to-day-enhancementGenomic_Benchmarks_human_enhancers_cohn
Dataset Card for "Genomic_Benchmarks_human_enhancers_cohn"
More Information needed
task275_enhanced_wsc_paraphrase_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task275_enhanced_wsc_paraphrase_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task275_enhanced_wsc_paraphrase_generation.librispeech-enhanced-voice-prompts
LibriSpeech enhanced voice prompts
The voice-cloning prompts used to evaluate
pocket-tts / Kyutai TTS models on
the LibriSpeech-PC test-clean cross-sentence protocol.
448 unique prompt utterances from LibriSpeech test-clean, speech-enhanced
(denoised, 32 kHz mono FLAC, duration-preserving), laid out as
<speaker>/<chapter>/<utterance>.flac — the same tree structure as
LibriSpeech itself, so --prompt-root style substitution works directly.
Also included:… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/librispeech-enhanced-voice-prompts.fer2013-enhanced
FER2013 Enhanced: Advanced Facial Expression Recognition Dataset
The most comprehensive and quality-enhanced version of the famous FER2013 dataset for state-of-the-art emotion recognition research and applications.
🎯 Dataset Overview
FER2013 Enhanced is a significantly improved version of the landmark FER2013 facial expression recognition dataset. This enhanced version provides AI-powered quality assessment, balanced data splits, comprehensive metadata, and multi-format… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/fer2013-enhanced.vertebrate-v1-issue473-center1-ccre-enhancer-centered
marin-dna/vertebrate-v1-issue473-center1-ccre-enhancer-centered
Review status: draft generated for issue #473 review before upload.
Human-anchored 255 bp vertebrate sequences for the
ccre_enhancer_centered cohort under the center_1 policy.
The policy projects the exact central human nucleotide, requires one target locus, and emits the 255 bp target window centered on that mapped nucleotide. Human anchors come from the fixed exp351 ENCODE dELS/pELS-centered catalog after… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-center1-ccre-enhancer-centered.th-en-zh-tts-200k-enhanced
TH-EN-ZH Multi-speaker TTS Dataset (200K, RE-USE Enhanced)
Speech-enhanced variant of FILM6912/th-en-zh-tts-200k.
Every clip has been processed through NVIDIA RE-USE (universal speech enhancement, SEMamba) at its native sample rate, then re-encoded losslessly as FLAC (PCM_16).
Same schema, same row order, same 200,000 rows (th 100k / en 50k / zh 50k):
Column
Type
Description
text
string
Transcript (identical to the original dataset)
audio
Audio
Enhanced audio, FLAC… See the full description on the dataset page: https://huggingface.co/datasets/FILM6912/th-en-zh-tts-200k-enhanced.PersonaChat-Qwen-Image-2512-enhancedBGEE-Graphics-Enhanced-Assetshttps://github.com/TheForgotten69/BGEE-Graphics-Enhanced
oeis-enhanced
Attribution
This dataset was generated using data from the On-Line Encyclopedia of Integer Sequences (OEIS).
Source: https://oeis.org/
License: Creative Commons Attribution-ShareAlike 4.0 (CC BY-SA 4.0)
OEIS End-User License Agreement: https://oeis.org/wiki/The_OEIS_End-User_License_Agreement
python_enhancement_proposals_filtered
Python Enhancement Proposals
Description
Python Enhancement Proposals, or PEPs, are design documents that generally provide a technical specification and rationale for new features of the Python programming language.
There have been 661 PEPs published.
The majority of PEPs are published in the Public Domain, but 5 were published under the “Open Publication License” and omitted from this dataset.
PEPs are long, highly-polished, and technical in nature and often include… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/python_enhancement_proposals_filtered.
