datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Lux-Japanese-Speech-Corpus
Lux Japanese Speech Corpus
概要
Lux Japanese Speech Corpus は、オリジナルキャラクター「Lux (ルクス)」による日本語のテキスト読み上げ音声を収録したデータセットです。このデータセットは、以下の2種類の音声ファイルで構成されています。
raw: 加工前の 96kHz/16bit の WAV ファイル
cleaned: ノイズ除去などの処理を施した 96kHz/16bit の WAV ファイル
各音声ファイルに対応するトランスクリプション(読み上げた文章)は、metadata.csv に記録しています。データセット全体のメタ情報は dataset_infos.json で管理されています。
ディレクトリ構造
以下は、このリポジトリの推奨ディレクトリ構造の例です。
Lux-Japanese-Speech-Corpus/
├── .gitattributes # Gitのカスタマイズファイル
├── README.md… See the full description on the dataset page: https://huggingface.co/datasets/Lami/Lux-Japanese-Speech-Corpus.Luxemburgish_Press_Conferences_GovLuxembourgish-Speech-Dataset
Luxembourgish Dataset Metadata
Field
Value
📜 License
CC BY-NC-ND 4.0
🎯 Task Categories
Automatic Speech Recognition
🌍 Language
Luxembourgish (lb)
🏷️ Tags
Audio, Speech, Speech Recognition, Machine, Machine Learning, ML
📦 Size Category
n < 1K
YodaLingua-Luxembourgish-Extended
YodaLingua-Luxembourgish-Extended
YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Luxembourgish-Extended portion of the multilingual YodaLingua collection.
🧾 Dataset Overview
Property
Value
Total clips
21,687 audio–transcription pairs
Total duration
75 hours
Speakers
2,674 distinct speakers
Audio format… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Luxembourgish-Extended.
