CoolFace
Datasetpublicgated

KeisukeMiyamoto/nhk-archive-audio-30s

NHK Archives Audio 30s This is a Japanese speech corpus derived from NHK Archives Audio. Audio from public NHK Archives records was segmented into clips of up to 30 seconds using voice activity detection. The dataset contains 344,722 accepted clips, totaling 1,661.01 hours. Audio is embedded as 16 kHz mono FLAC. raw_text was transcribed with Whisper large-v3-turbo, and text contains LLM-assisted corrections based on the transcript and available source title and description. This… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/nhk-archive-audio-30s.

sourceHugging Faceotherupdated 2d agoView on Hugging Face
0likes377downloads
settings

This repository belongs to KeisukeMiyamoto on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namenhk-archive-audio-30s
visibilitypublic
licenceother
gatedyes
ownerKeisukeMiyamoto
Account settings
KeisukeMiyamoto/nhk-archive-audio-30s · CoolFace