CoolFace
Datasetpublicgated

KeisukeMiyamoto/nhk-archive-audio-30s

NHK Archives Audio 30s This is a Japanese speech corpus derived from NHK Archives Audio. Audio from public NHK Archives records was segmented into clips of up to 30 seconds using voice activity detection. The dataset contains 344,722 accepted clips, totaling 1,661.01 hours. Audio is embedded as 16 kHz mono FLAC. raw_text was transcribed with Whisper large-v3-turbo, and text contains LLM-assisted corrections based on the transcript and available source title and description. This… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/nhk-archive-audio-30s.

sourceHugging Faceotherupdated 1d agoView on Hugging Face
0likes371downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.