CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Digisensus /lithuanian-phone-speech-liepa-3-429h-punctuated Lithuanian Phone Speech 429 h: punctuated, cased, numbers as digits (written form) Transcripts are in written form, not normalised: punctuation, capitalisation, and numbers, dates, times and amounts as digits ("2026 m. rugsėjo 6 d., 9:30", "65 000 €", "12,5 %"). The original normalised text is included too. text text_normalized Varšuva 85 % sugriauta. varšuva aštuoniasdešim penki procentai sugriauta Keliais eurais arba 10 € daugiau kaip valytojos. keliais eurais… See the full description on the dataset page: https://huggingface.co/datasets/Digisensus/lithuanian-phone-speech-liepa-3-429h-punctuated.audioautomatic-speech-recognition100K<n<1M1 likes112 downloads2d agoHugging Face02Digisensus /lithuanian-dialect-speech-liepa-3-100h-punctuated Lithuanian Dialect Speech 100 h: punctuated, cased, numbers as digits (written form) Spontaneous Lithuanian dialect speech from all four regions, with three transcripts per clip: written form (punctuation, capitalisation, numbers as digits), normalised, and the original phonetic transcription with stress marks. Dialect word forms are kept as spoken in every layer. text text_normalized text_phonetic Per 3 klases buvu 10 mokinių. per tris klases buvu dešim mokinių per… See the full description on the dataset page: https://huggingface.co/datasets/Digisensus/lithuanian-dialect-speech-liepa-3-100h-punctuated.audioautomatic-speech-recognition100K<n<1M1 likes24 downloads2d agoHugging Face03Speech-data /Lithuanian-Speech-Dataset Lithuanian Dataset Metadata Field Value 📜 License CC BY-NC-ND 4.0 🎯 Task Categories Automatic Speech Recognition 🌍 Language Lithuanian (lt) 🏷️ Tags Lithuanian, Audio, Speech, Speech Recognition, ML, Machine, Machine Learning 📦 Size Category n < 1K audioautomatic-speech-recognitionn<1K0 likes19 downloads6mo agoHugging Face04Thomcles /YodaLingua-Lithuaniangated YodaLingua-Lithuanian YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Lithuanian portion of the multilingual YodaLingua collection. 🧾 Dataset Overview Property Value Total clips 2745 audio–transcription pairs Total duration 7.4 hours Speakers 138 distinct speakers Audio format MP3 • mono • 24 kHz •… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Lithuanian.audiotext-to-speech1K<n<10K0 likes10 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.