datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lithuanian-phone-speech-liepa-3-429h-punctuated
Lithuanian Phone Speech 429 h: punctuated, cased, numbers as digits (written form)
Transcripts are in written form, not normalised: punctuation, capitalisation, and numbers,
dates, times and amounts as digits ("2026 m. rugsėjo 6 d., 9:30", "65 000 €", "12,5 %").
The original normalised text is included too.
text
text_normalized
Varšuva 85 % sugriauta.
varšuva aštuoniasdešim penki procentai sugriauta
Keliais eurais arba 10 € daugiau kaip valytojos.
keliais eurais… See the full description on the dataset page: https://huggingface.co/datasets/Digisensus/lithuanian-phone-speech-liepa-3-429h-punctuated.lithuanian-speech-datasetlithuanian-dialect-speech-liepa-3-100h-punctuated
Lithuanian Dialect Speech 100 h: punctuated, cased, numbers as digits (written form)
Spontaneous Lithuanian dialect speech from all four regions, with three transcripts per clip:
written form (punctuation, capitalisation, numbers as digits), normalised, and the original
phonetic transcription with stress marks. Dialect word forms are kept as spoken in every layer.
text
text_normalized
text_phonetic
Per 3 klases buvu 10 mokinių.
per tris klases buvu dešim mokinių
per… See the full description on the dataset page: https://huggingface.co/datasets/Digisensus/lithuanian-dialect-speech-liepa-3-100h-punctuated.lithuaniaLithuanian-Speech-Dataset
Lithuanian Dataset Metadata
Field
Value
📜 License
CC BY-NC-ND 4.0
🎯 Task Categories
Automatic Speech Recognition
🌍 Language
Lithuanian (lt)
🏷️ Tags
Lithuanian, Audio, Speech, Speech Recognition, ML, Machine, Machine Learning
📦 Size Category
n < 1K
YodaLingua-Lithuanian
YodaLingua-Lithuanian
YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Lithuanian portion of the multilingual YodaLingua collection.
🧾 Dataset Overview
Property
Value
Total clips
2745 audio–transcription pairs
Total duration
7.4 hours
Speakers
138 distinct speakers
Audio format
MP3 • mono • 24 kHz •… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Lithuanian.
