datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cifar100-pythonnavsim-metric-caches-from-a100PIG-Nav-Dataset-PretrainThe pretraining dataset for our paper PIG-Nav: Key Insights for Pretrained Image-Goal Navigation Models.
Description of the dataset:
The pretraining dataset include GoStanford, RECON, CoryHall, Berkeley DeepDrive, SCAND, TartanDrive, and SACSoN.
Please follow our github repo https://github.com/zpschang/PIG-Nav for detailed use.
Japanese-Eroge-Voice
Japanese-Eroge-Voice
Description
This dataset contains pairs of audio data and corresponding transcriptions extracted from Japanese eroge (adult games) that I have personally purchased. The transcriptions are generated using the litagin/anime-whisper model.
Preprocessing Steps
The raw audio data has undergone the following preprocessing steps:
Loudness Normalization:
Audio loudness is normalized using ffmpeg's 2-pass loudnorm filter to target parameters of… See the full description on the dataset page: https://huggingface.co/datasets/NandemoGHS/Japanese-Eroge-Voice.NatHEARPlease see LICENSE.txt for each dataset licenses
elector-sen-claveros-nacionalreprocessed_singapore_national_speech_corpus
Dataset Card for Reprocessed National Speech Corpus
NOTE: This is an Reprocessed version KaraKaraWitch from Recursal.The official download can be found here.
Dataset Details
Dataset Description
Dataset Description:
The National Speech Corpus (NSC) is the first large-scale Singapore English corpus, sponsored by the Info-communications and Media Development Authority (IMDA) of Singapore. The objective is to serve as a primary resource of open speech data for… See the full description on the dataset page: https://huggingface.co/datasets/recursal/reprocessed_singapore_national_speech_corpus.vggsound_08navidromejapanese-str-dataset-v1
STR Dataset
Japanese STR (Scene Text Recognition) dataset in WebDataset format.
This dataset is composed of:
Images of Japanese named entities (full names and their affiliations)
Images of sentences retrieved from Aozora Bunko (青空文庫)
and their corresponding ground truth texts.
All images are synthesized using TRDG.
Dataset Structure
Split
Samples
Shards
train
10,000,000
1000
valid
50,000
5
test
50,000
5
Total
10,100,000
1010
Usage
import… See the full description on the dataset page: https://huggingface.co/datasets/nagohachi/japanese-str-dataset-v1.Nano3D-Edit-100k
Nano3D-Edit-100k
This dataset is the official data release for Nano3D, a training-free framework for precise and coherent 3D object editing without masks.
Paper: Nano3D: A Training-Free Approach for Efficient 3D Editing Without MasksProject Page: https://jamesyjl.github.io/Nano3D/
Nano3D integrates FlowEdit into TRELLIS to perform localized 3D edits guided by front-view renderings, and introduces Voxel/Slat-Merge strategies to preserve structural consistency between edited and… See the full description on the dataset page: https://huggingface.co/datasets/yejunliang23/Nano3D-Edit-100k.L2P-dataset
L2P: Unlocking Latent Potential for Pixel Generation
An efficient transfer paradigm enabling high-quality, end-to-end pixel-space diffusion with minimal computational overhead and data requirements.
Pixel diffusion models have recently regained attention for visual generation. However, training advanced pixel-space models from scratch demands prohibitive computational and data resources. To address this, we propose the Latent-to-Pixel (L2P) transfer… See the full description on the dataset page: https://huggingface.co/datasets/zhen-nan/L2P-dataset.NaturalVoices_VC_0.1 NaturalVoices VC 10%
A large voice conversion (VC) dataset curated from spontaneous, in-the-wild podcast speech as part of the NaturalVoices project in collaboration with 🤗MSP Lab at CMU LTI. This release provides the 10% subset uniformly sampled from 870-hour VC dataset and subsets mainly intended for training and evaluating emotion-aware voice conversion systems but not limited to VC tasks.
📄 Paper: NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice… See the full description on the dataset page: https://huggingface.co/datasets/JHU-SmileLab/NaturalVoices_VC_0.1.jawildtext_cropped
jawildtext_cropped
Per-polygon scene-text-recognition crops derived from
llm-jp/jawildtext.
Each row of the source dataset carries a polygons column with quadrilateral
text regions and their transcriptions. For every polygon we perspective-warp
the source image onto the rectified bounding rectangle, yielding a tight,
horizontally-aligned crop suitable for training/evaluating Japanese scene-text
recognition models.
Stats
Samples: 108 403
Shards: 22… See the full description on the dataset page: https://huggingface.co/datasets/nagohachi/jawildtext_cropped.Nat-HEAR-Ambisonicsnanostep-datasetsnav-datatamil_nadu_v4_trainingNDL_pdm-ocr-part2_cropped
NDL Public Domain OCR Dataset — Line-Cropped (WebDataset)
This dataset is a derivative of the Public Domain OCR Training Dataset (FY2021) published by the National Diet Library of Japan (国立国会図書館).
Each sample is a line-level crop of a historical document page, paired with its OCR text transcription.
What was changed from the original
This dataset was not created by the National Diet Library. The following transformations were applied:
Each text line (LINE element)… See the full description on the dataset page: https://huggingface.co/datasets/nagohachi/NDL_pdm-ocr-part2_cropped.PIG-Nav-Dataset-FinetuneThe finetuning dataset for our paper PIG-Nav: Key Insights for Pretrained Image-Goal Navigation Models.
Description of the dataset:
The finetuning dataset include data collected by human players on two large 3D game environments, Highrise and Sanctuary.
Please follow our github repo https://github.com/zpschang/PIG-Nav for detailed use.
GRAM-Naturalistic-Scenes-Binauraltamil_nadunaidiffusionv3distilSingle Danbooru-Tag text input, intended for recalibration of Cross-Attention.
Word frequency cutoff: tags with at least 1000 posts.
safe_file_name = re.sub(r'[^\w\-_\. ]', '_', prompt)
file_name = f"{seed_folder}/{safe_file_name}_{seed}.png"
ffhq-512-jpgnanospeechNat-HEAR-localizationGRAM-Naturalistic-Scenes-AmbisonicsData_NanoVLMJKTSV-primarynasal_phantom_data
