CoolFace
Datasetpublic

FluidInference/JSUT-basic5000

JSUT (Japanese Speech Corpus) - Test Subset A test subset of the JSUT corpus containing 500 Japanese utterances from the basic5000 dataset (BASIC5000_4501-5000). Dataset Structure jsut_ver1.1/ └── basic5000/ ├── wav/ # WAV audio files (500 files, 48kHz) ├── transcript_utf8.txt # Transcriptions └── recording_info.txt # Recording dates File Formats transcript_utf8.txt… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/JSUT-basic5000.

sourceHugging Faceupdated 6mo agoView on Hugging Face
1likes346downloads
ChangeLog.txt45 linesDownload Raw Back to root
110/26/2017 ... ver. 12   7666 utterances have been released. Some utterances are missed, but there will be added soon.311/30/2017 ... ver. 1.14   35 utterances missed (or mispronounced) in ver. 1 have been added. 7696 utterances have been released. The basenames are5      PRECEDENT130_0456      LOANWORD128_0887      ONOMATOPEE300_0748      BASIC5000_14959      BASIC5000_225210      BASIC5000_233311      BASIC5000_238212      BASIC5000_239113      BASIC5000_262614      BASIC5000_293615      BASIC5000_301316      BASIC5000_309317      BASIC5000_309718      BASIC5000_354119      BASIC5000_359320      BASIC5000_360421      BASIC5000_366722      BASIC5000_368723      BASIC5000_371324      BASIC5000_376425      BASIC5000_384526      BASIC5000_386727      BASIC5000_390728      BASIC5000_396729      BASIC5000_402630      BASIC5000_415131      BASIC5000_416632      BASIC5000_419533      BASIC5000_421834      BASIC5000_426035      BASIC5000_434336      BASIC5000_434637      BASIC5000_434838      BASIC5000_448639      BASIC5000_4952.40   Note that the corresponding recording_info.txt and transcript_utf8.txt were also changed.41   42   We slightly modified some basenames listed in transcript_utf8.txt.43 44 45