datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
japanese-anime-speech-v2
Japanese Anime Speech Dataset V2
日本語はこちら
japanese-anime-speech-v2 is an audio-text dataset designed for training automatic speech recognition models.
The dataset comprises 292,637 audio clips and their corresponding transcriptions from various visual novels.
This dataset is not an updated version of japanese-anime-speech-v1.
For that reason, most of the audio from japanese-anime-speech-v1 is not included in this dataset.
The goal of this dataset is to increase the accuracy of… See the full description on the dataset page: https://huggingface.co/datasets/joujiboi/japanese-anime-speech-v2.japanese-anime-speech
Japanese Anime Speech Dataset
日本語はこちら
japanese-anime-speech is an audio-text dataset designed for the training of automatic speech recognition models. The dataset is comprised of thousands of audio clips and their corresponding transcriptions from different visual novels.
The goal of this dataset is to increase the accuracy of automatic speech recognition models, such as OpenAI's Whisper, in accurately transcribing dialogue from anime and other similar Japanese media. This genre is… See the full description on the dataset page: https://huggingface.co/datasets/joujiboi/japanese-anime-speech.japanese-anime-speech-v2
Japanese Anime Speech Dataset V2
日本語はこちら
japanese-anime-speech-v2 is an audio-text dataset designed for training automatic speech recognition models.
The dataset comprises 292,637 audio clips and their corresponding transcriptions from various visual novels.
This dataset is not an updated version of japanese-anime-speech-v1.
For that reason, most of the audio from japanese-anime-speech-v1 is not included in this dataset.
The goal of this dataset is to increase the accuracy of… See the full description on the dataset page: https://huggingface.co/datasets/WatsonNT/japanese-anime-speech-v2.japanese-anime-speech-v2-split-150k
japanese-anime-speech-v2-split-150k
joujiboi/japanese-anime-speech-v2 的 150,000 筆隨機子集,已切好 train / test。
split
rows
train
135,000
test
15,000
total
150,000
Columns
audio — 16 kHz mp3,與原始資料完全相同(未重新編碼)
sentence — 轉錄文字(原始欄位名為 transcription)
How it was built
來源的 sfw(271,788 筆)與 nsfw(20,849 筆)兩個 split 都有使用,並依原始比例分配名額
(sfw 139,313 / nsfw 10,687)。
在每個 split 的全部列上做無放回均勻抽樣,因此每筆資料被選中的機率相同。
抽出後整體打亂,前 15,000 筆為 test,其餘為 train,train / test… See the full description on the dataset page: https://huggingface.co/datasets/hhim8826/japanese-anime-speech-v2-split-150k.AnimeVox
AnimeVox: Character TTS Corpus
🗣️ Dataset Overview
AnimeVox is an English Text-to-Speech (TTS) dataset featuring 11,020 audio clips from 19 distinct anime characters across popular series. Each clip includes a high-quality transcription, character name, and anime title, making it ideal for voice cloning, custom TTS model fine-tuning, and character voice synthesis research.
The dataset was created and processed using TTSizer, an open-source tool that automates creating… See the full description on the dataset page: https://huggingface.co/datasets/taresh18/AnimeVox.japanese-anime-speech-v2-splitdataset split from joujiboi/japanese-anime-speech-v2
transcribed_auspicious_anime_girl_audio
Transcribed Auspicious Anime Girl Audio
A single-speaker English voice-over dataset containing 75 Arlecchino clips (18.24 minutes) with embedded audio and transcripts.
Audio and transcripts were collected from the Genshin Impact Wiki Arlecchino voice-over page, revision 2123801. Only English voice-over files with nonempty transcripts are included.
Columns
audio: embedded audio bytes and source filename
text: transcript
source_file: original filename… See the full description on the dataset page: https://huggingface.co/datasets/gabrielclark3330/transcribed_auspicious_anime_girl_audio.anime-2024
Anime Video Dataset 2024
Overview
This is an anime video dataset curated specifically for multimedia research purposes.
It is sourced from anime series released in 2024 and includes both a complete combined file and four seasonal subsets (Winter, Spring, Summer, and Fall).
Data Details
Video Stream
Codec: H.264 (High Profile)
Format: YUV 4:2:0 (Progressive)
Resolution: 640×360 (SAR 1:1, DAR 16:9)
Frame Rate: 23.98 FPS
Audio Stream
Codec: AAC (LC)
Sample… See the full description on the dataset page: https://huggingface.co/datasets/JacobLinCool/anime-2024.AnimeVox
AnimeVox: Character TTS Corpus
🗣️ Dataset Overview
AnimeVox is an English Text-to-Speech (TTS) dataset featuring 11,020 audio clips from 19 distinct anime characters across popular series. Each clip includes a high-quality transcription, character name, and anime title, making it ideal for voice cloning, custom TTS model fine-tuning, and character voice synthesis research.
The dataset was created and processed using TTSizer, an open-source tool that automates creating… See the full description on the dataset page: https://huggingface.co/datasets/humairawan/AnimeVox.AnimeVox
AnimeVox: Character TTS Corpus
🗣️ Dataset Overview
AnimeVox is an English Text-to-Speech (TTS) dataset featuring 11,020 audio clips from 19 distinct anime characters across popular series. Each clip includes a high-quality transcription, character name, and anime title, making it ideal for voice cloning, custom TTS model fine-tuning, and character voice synthesis research.
The dataset was created and processed using TTSizer, an open-source tool that automates creating… See the full description on the dataset page: https://huggingface.co/datasets/RenderCode/AnimeVox.
