datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zoonomia-v1-v4_ccre_noexon_enhancer
bolinas-dna/zoonomia-v1-v4_ccre_noexon_enhancer
A curated enhancer training set for issue
#326 — a de-contaminated
derivation of the v4 ccre_non_promoter arm of
bolinas-dna/zoonomia-v1-v1,
built by the
snakemake/zoonomia_projection_dataset pipeline at commit
6b320c268547.
Provenance
This subset is v4_ccre_noexon further restricted to enhancer-dominant windows (dELS+pELS basepair coverage ≥ the other non-PLS cCRE classes), population-matching the val_enhancer… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_noexon_enhancer.gpn-star-p-uniform-v1-enhancer-arm-a
marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a.phylop-uniform-v1-enhancer-arm-a
marin-dna/phylop-uniform-v1-enhancer-arm-a
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses the pipeline's pinned phyloP conservation… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/phylop-uniform-v1-enhancer-arm-a.functional-enhancer
marin-dna/functional-enhancer
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and official version-matched UCSC hg38-to-target liftOver chains.
This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
For the 28 non-mammalian targets, the stable… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/functional-enhancer.vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered
marin-dna/vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered
Review status: draft generated for issue #473 review before upload.
Human-anchored 255 bp vertebrate sequences for the
ccre_enhancer_centered cohort under the full_window policy.
The policy projects the complete 255 bp human window, applies the established 128--512 bp compatible-fragment gate, and resizes around the accepted target-span midpoint. Human anchors come from the fixed exp351 ENCODE dELS/pELS-centered… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered.DeepSTARR-enhancer-activity
Abouts
The enhancer activity data is sourced from the DeepSTARR repo.
We have applied minor formatting adjustments to the dataset to facilitate streamlined data analysis.
How to use
from datasets import load_dataset
datasets = load_dataset("GenerTeam/DeepSTARR-enhancer-activity")
enhanced-audiosnippets-long-2-8M
Enhanced Audiosnippets Long 2.8M
Enhanced version of mitermix/audiosnippets_long_2_8M with speech enhancement, emotion annotations, speaker embeddings, and comprehensive metadata analysis.
Dataset Summary
Metric
Value
Total samples
2,633,037
Total audio hours
4,932 h
Duration range
3.0s - 1124.3s
Mean duration
6.7s
Audio format
WAV, 48kHz mono
Tar files
1,410
Processing Pipeline
Each audio sample was processed through:
Speech… See the full description on the dataset page: https://huggingface.co/datasets/ai-music4you3/enhanced-audiosnippets-long-2-8M.zoonomia-v1-v4_ccre_noexon_enhancer-order
bolinas-dna/zoonomia-v1-v4_ccre_noexon_enhancer-order
The bolinas-dna/zoonomia-v1-v4_ccre_noexon_enhancer cross-mammal
training set, restricted to a species cohort: one representative species per NCBI order — 19 deeply-diverged placental mammals (every pair separated by ~tens of millions of years), versus the implicit-default 108 family-deduplicated species. A strict subset of the family set, so it reuses the v1 cross-mammal projection unchanged (no re-halLiftover).
Same human… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_noexon_enhancer-order.fer2013-enhanced
FER2013 Enhanced: Advanced Facial Expression Recognition Dataset
The most comprehensive and quality-enhanced version of the famous FER2013 dataset for state-of-the-art emotion recognition research and applications.
🎯 Dataset Overview
FER2013 Enhanced is a significantly improved version of the landmark FER2013 facial expression recognition dataset. This enhanced version provides AI-powered quality assessment, balanced data splits, comprehensive metadata, and multi-format… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/fer2013-enhanced.vertebrate-v1-issue473-center1-ccre-enhancer-centered
marin-dna/vertebrate-v1-issue473-center1-ccre-enhancer-centered
Review status: draft generated for issue #473 review before upload.
Human-anchored 255 bp vertebrate sequences for the
ccre_enhancer_centered cohort under the center_1 policy.
The policy projects the exact central human nucleotide, requires one target locus, and emits the 255 bp target window centered on that mapped nucleotide. Human anchors come from the fixed exp351 ENCODE dELS/pELS-centered catalog after… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-center1-ccre-enhancer-centered.enhanced_pick_up_socks_dataset_act_2
enhanced_pick_up_socks_dataset
This dataset was generated using the phospho cli
More information on robots.phospho.ai.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
enhanced_pick_up_socks_dataset
enhanced_pick_up_socks_dataset
This dataset was generated using the phospho cli
More information on robots.phospho.ai.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
enhanced_pick_up_socks_dataset_act
enhanced_pick_up_socks_dataset
This dataset was generated using the phospho cli
More information on robots.phospho.ai.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
zoonomia-v1-v4_ccre_enhancer_centered-order
bolinas-dna/zoonomia-v1-v4_ccre_enhancer_centered-order
An enhancer-CENTERED training set for issue
#351, built by the
snakemake/zoonomia_projection_dataset pipeline
(workflow/rules/centered.smk) at commit
8127acfea5aa.
Provenance
Each training window is defined directly from an ENCODE cCRE V4 enhancer
(dELS + pELS): one 255 bp window centered on the cCRE midpoint
(make_enhancer_anchors, keep-all — clustered enhancers each keep their own
window), rather than a… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_enhancer_centered-order.Env-TTS-SD-Corpus-24K-Enhancedenhanced_pick_up_socks_dataset_act_3
enhanced_pick_up_socks_dataset_act_3
This dataset was generated using the phospho cli
More information on robots.phospho.ai.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
swerebench-traces-raw-source-verification-enhanced-20260617
SWE-rebench Raw Source Verification Enhanced 20260617
This is a private raw source dataset for building refined mini-swe-agent SFT datasets. It is intentionally not tokenized and intentionally preserves source data plus metadata for downstream filtering, masking, weighting, and audit. Do not treat every row as a clean endpoint solve.
Download
The full dataset directory is uploaded as a single compressed archive:
hf download… See the full description on the dataset page: https://huggingface.co/datasets/eewer/swerebench-traces-raw-source-verification-enhanced-20260617.ConvAI2-Qwen-enhanced-Qwen3.5-35B-A3B
Visual Memory Results: convai2-qwen-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-35B-A3B",
"hf_results_repo": "visual-memory/ConvAI2-Qwen-enhanced-Qwen3.5-35B-A3B",
"results_jsonl": "results/ConvAI2-Qwen-enhanced-Qwen3.5-35B-A3B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-Qwen-enhanced-Qwen3.5-35B-A3B.ConvAI2-Qwen-enhanced-gemma-4-26B-A4B-it
Visual Memory Results: convai2-qwen-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-26B-A4B-it",
"hf_results_repo": "visual-memory/ConvAI2-Qwen-enhanced-gemma-4-26B-A4B-it",
"results_jsonl": "results/ConvAI2-Qwen-enhanced-gemma-4-26B-A4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-Qwen-enhanced-gemma-4-26B-A4B-it.ConvAI2-FLUX-enhanced-Qwen3.5-35B-A3B
Visual Memory Results: convai2-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-35B-A3B",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-enhanced-Qwen3.5-35B-A3B",
"results_jsonl": "results/ConvAI2-FLUX-enhanced-Qwen3.5-35B-A3B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-enhanced-Qwen3.5-35B-A3B.enhanced_brain_magnet_hg38For details and usage of these datasets please see my GitHub repository: https://github.com/caenrigen/enhanced_brain_magnet.
ConvAI2-FLUX-enhanced-gemma-4-26B-A4B-it
Visual Memory Results: convai2-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-26B-A4B-it",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-enhanced-gemma-4-26B-A4B-it",
"results_jsonl": "results/ConvAI2-FLUX-enhanced-gemma-4-26B-A4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-enhanced-gemma-4-26B-A4B-it.ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits
Ashaar Enhanced Description SFT Stratified Splits
Source dataset:
Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500
Target dataset:
Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits
This dataset publishes deterministic train / eval / test splits with a 94 / 3 / 3 policy.
Split policy
Primary stratification key:
base_meter
form
length_bucket
Length buckets:
1-3
4-6
7-10
11-20
Small groups fall back… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits.ConvAI2-ERNIE-enhanced-Qwen3.5-35B-A3B
Visual Memory Results: convai2-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-35B-A3B",
"hf_results_repo": "visual-memory/ConvAI2-ERNIE-enhanced-Qwen3.5-35B-A3B",
"results_jsonl": "results/ConvAI2-ERNIE-enhanced-Qwen3.5-35B-A3B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-enhanced-Qwen3.5-35B-A3B.ConvAI2-ERNIE-enhanced-gemma-4-26B-A4B-it
Visual Memory Results: convai2-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-26B-A4B-it",
"hf_results_repo": "visual-memory/ConvAI2-ERNIE-enhanced-gemma-4-26B-A4B-it",
"results_jsonl": "results/ConvAI2-ERNIE-enhanced-gemma-4-26B-A4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-enhanced-gemma-4-26B-A4B-it.Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-31B-it
Visual Memory Results: synthetic-persona-chat-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-31B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-31B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-31B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-31B-it.Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-31B-it
Visual Memory Results: synthetic-persona-chat-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-31B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-31B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-31B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-31B-it.PersonaChat-FLUX-enhanced-Qwen3.5-35B-A3B
Visual Memory Results: personachat-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-35B-A3B",
"hf_results_repo": "visual-memory/PersonaChat-FLUX-enhanced-Qwen3.5-35B-A3B",
"results_jsonl": "results/PersonaChat-FLUX-enhanced-Qwen3.5-35B-A3B.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-FLUX-enhanced-Qwen3.5-35B-A3B.PersonaChat-ERNIE-enhanced-Qwen3.5-35B-A3B
Visual Memory Results: personachat-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-35B-A3B",
"hf_results_repo": "visual-memory/PersonaChat-ERNIE-enhanced-Qwen3.5-35B-A3B",
"results_jsonl": "results/PersonaChat-ERNIE-enhanced-Qwen3.5-35B-A3B.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-ERNIE-enhanced-Qwen3.5-35B-A3B.ConvAI2-ERNIE-enhanced-gemma-4-E4B-it
Visual Memory Results: convai2-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/ConvAI2-ERNIE-enhanced-gemma-4-E4B-it",
"results_jsonl": "results/ConvAI2-ERNIE-enhanced-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-enhanced-gemma-4-E4B-it.
