CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hackelle /BigEarthNetV2-Lithuania-Summer-LMDB TU Berlin RSiM DIMA BigEarth BIFOLD reBEN — Lithuania Summer Subset (pre-converted to LMDB) ⚠️ Unofficial mirror. This is an unofficial, community-provided pre-conversion of a subset of the BigEarthNet v2.0 (reBEN) dataset into LMDB format. It is provided as a convenience for researchers who wish to get started quickly without running the full conversion pipeline. In case of any discrepancy, the original publication and the original files always take… See the full description on the dataset page: https://huggingface.co/datasets/hackelle/BigEarthNetV2-Lithuania-Summer-LMDB.textimage-classification1K<n<10K0 likes377 downloads6mo agoHugging Face02endomorphosis /ipfs_lithuania_laws Lithuania TAR / e-TAR legal acts Research snapshot of official national legislation from TAR / e-tar.lt (Teisės aktų registras) official legal acts; leftover Wayback fill of official URLs. Not legal advice. The official gazette / authentic source prevails over this corpus. Snapshot Field Value Snapshot date 2026-09-17 Coverage shard leftover Wayback incomplete Source TAR / e-tar.lt (Teisės aktų registras) official legal acts; leftover Wayback fill… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_lithuania_laws.texttext-retrieval1K<n<10K1 likes156 downloads8d agoHugging Face03Digisensus /lithuanian-phone-speech-liepa-3-429h-punctuated Lithuanian Phone Speech 429 h: punctuated, cased, numbers as digits (written form) Transcripts are in written form, not normalised: punctuation, capitalisation, and numbers, dates, times and amounts as digits ("2026 m. rugsėjo 6 d., 9:30", "65 000 €", "12,5 %"). The original normalised text is included too. text text_normalized Varšuva 85 % sugriauta. varšuva aštuoniasdešim penki procentai sugriauta Keliais eurais arba 10 € daugiau kaip valytojos. keliais eurais… See the full description on the dataset page: https://huggingface.co/datasets/Digisensus/lithuanian-phone-speech-liepa-3-429h-punctuated.audioautomatic-speech-recognition100K<n<1M1 likes112 downloads2d agoHugging Face04shunyalabs /lithuanian-speech-datasetaudio1K<n<10K2 likes98 downloads1y agoHugging Face05Frararo /lithuanian-tts-votes0 likes76 downloads9d agoHugging Face06ZygAI /zygai_soviet_lithuania 🇱🇹 ZygAI – Soviet Lithuania (1940–1990) Dataset A comprehensive cultural–historical dataset documenting Soviet-era Lithuania 📌 Overview ZygAI Soviet Lithuania is a structured, bilingual (LT + EN) dataset capturing political,cultural, economic, educational, technological, and everyday life aspects of Lithuaniaunder Soviet occupation (1940–1990). This dataset is a major component of ZygAI Research 2025–2026, created to preservehistorical memory, support academic… See the full description on the dataset page: https://huggingface.co/datasets/ZygAI/zygai_soviet_lithuania.texttext-classification1K<n<10K3 likes60 downloads11mo agoHugging Face07Shaagun /Translation-English-Lithuanian0 likes49 downloads2y agoHugging Face08saillab /alpaca-lithuanian-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-lithuanian-cleaned.text10K<n<100K0 likes39 downloads2y agoHugging Face09neurotechnology /lithuanian-qa-v1 Dataset Card for Lithuanian QA V1 1. General Information Dataset Name: Lithuanian QA V1 Dataset Description: This dataset consists of question-answer pairs in Lithuanian, focusing on topics related to Lithuanian culture, history, and people. It is a unique resource designed to aid in the development of language models specifically tailored for Lithuanian linguistic nuances. Purpose of the Dataset: The primary purpose of this dataset is to facilitate the fine-tuning of… See the full description on the dataset page: https://huggingface.co/datasets/neurotechnology/lithuanian-qa-v1.text10K<n<100K4 likes28 downloads2y agoHugging Face10sam8000 /lithuaniaaudio10K<n<100K0 likes26 downloads1y agoHugging Face11Shaagun /Instruc_Lithuaniantext10K<n<100K0 likes25 downloads2y agoHugging Face12Digisensus /lithuanian-dialect-speech-liepa-3-100h-punctuated Lithuanian Dialect Speech 100 h: punctuated, cased, numbers as digits (written form) Spontaneous Lithuanian dialect speech from all four regions, with three transcripts per clip: written form (punctuation, capitalisation, numbers as digits), normalised, and the original phonetic transcription with stress marks. Dialect word forms are kept as spoken in every layer. text text_normalized text_phonetic Per 3 klases buvu 10 mokinių. per tris klases buvu dešim mokinių per… See the full description on the dataset page: https://huggingface.co/datasets/Digisensus/lithuanian-dialect-speech-liepa-3-100h-punctuated.audioautomatic-speech-recognition100K<n<1M1 likes24 downloads2d agoHugging Face13alexandrainst /lithuanian-sentiment-analysisAll samples are reviews scraped from https://atsiliepimai.lt/. text1K<n<10K0 likes23 downloads10mo agoHugging Face14Speech-data /Lithuanian-Speech-Dataset Lithuanian Dataset Metadata Field Value 📜 License CC BY-NC-ND 4.0 🎯 Task Categories Automatic Speech Recognition 🌍 Language Lithuanian (lt) 🏷️ Tags Lithuanian, Audio, Speech, Speech Recognition, ML, Machine, Machine Learning 📦 Size Category n < 1K audioautomatic-speech-recognitionn<1K0 likes19 downloads6mo agoHugging Face15Ramuneri /lithuanian-road-signsimagen<1K0 likes18 downloads4mo agoHugging Face16ArturG9 /Lithuanian_Context_QA Lithuanian QA Dataset - Generated with DSPy & Gemma2 27B Q4 Introduction This dataset was created using DSPy, a Python framework that simplifies the generation of question and answer (QA) pairs from a given context. The dataset is composed of context, questions, and answers, all in Lithuanian. The context was primarily sourced from the following resources: Lithuanian Wikipedia (lt.wikipedia.org) Lietuviškoji enciklopedija (vle.lt) Book: Vitalija Skėruvienė, Civilinė Teisė Mokomoji… See the full description on the dataset page: https://huggingface.co/datasets/ArturG9/Lithuanian_Context_QA.textn<1K0 likes17 downloads2y agoHugging Face17FrancophonIA /NTEU_French-Lithuanian [!NOTE] Dataset origin: https://live.european-language-grid.eu/catalogue/corpus/7130/ Description This is a compilation of parallel corpora resources used in building of Machine Translation engines in NTEU project (Action number: 2018-EU-IA-0051). Data in this resource are compiled in two TMX files, two tiers grouped by data source reliablity. Tier A -- danta originating from human edited sources, translation memories and alike. Tier B -- danta originating created by automatic… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/NTEU_French-Lithuanian.translation0 likes15 downloads1y agoHugging Face18saillab /alpaca_lithuanian_tacoThis repository contains the dataset used for the TaCo paper. The dataset follows the style outlined in the TaCo paper, as follows: { "instruction": "instruction in xx", "input": "input in xx", "output": "Instruction in English: instruction in en , Response in English: response in en , Response in xx: response in xx " } Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_lithuanian_taco.text10K<n<100K0 likes14 downloads2y agoHugging Face19kjhq /Lithuania-Stock-Symbols-and-Metadata Lithuania Stock Symbols & Company Metadata This dataset contains stock symbols and basic company metadata for all listed companies in Lithuania.It is updated weekly if new changes are there. 📊 Dataset Contents The dataset is provided as a CSV file with the following columns: Column Description name Full company name ticker Stock ticker symbol (e.g., AAPL, MSFT) market The exchange/market where the stock is listed sector The primary business sector of… See the full description on the dataset page: https://huggingface.co/datasets/kjhq/Lithuania-Stock-Symbols-and-Metadata.textn<1K0 likes14 downloads1y agoHugging Face20werent4 /lithuanian-translationstextn<1K0 likes12 downloads2y agoHugging Face21DebasishDhal99 /exonyms-for-lithuanian-placestextn<1K0 likes11 downloads3y agoHugging Face22domce20 /c4-lithuanian-enhancedtabular10M<n<100M1 likes10 downloads2y agoHugging Face23Thomcles /YodaLingua-Lithuaniangated YodaLingua-Lithuanian YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Lithuanian portion of the multilingual YodaLingua collection. 🧾 Dataset Overview Property Value Total clips 2745 audio–transcription pairs Total duration 7.4 hours Speakers 138 distinct speakers Audio format MP3 • mono • 24 kHz •… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Lithuanian.audiotext-to-speech1K<n<10K0 likes10 downloads8mo agoHugging Face24emailmarketingdataset /lithuania-business-dataset Lithuania Business Email List — Market Intelligence Dataset 7,000 verified Lithuania contacts are available from LeadsBlue →. This open dataset provides the aggregate market intelligence behind that database — contact volume, benchmark open/reply rates, send timing, and compliance for the Lithuania segment. At a glance: Market-intelligence reference data for the Lithuania Business Email List audience — population scale, decision-maker structure, channel benchmarks, and… See the full description on the dataset page: https://huggingface.co/datasets/emailmarketingdataset/lithuania-business-dataset.tabular-classificationn<1K0 likes9 downloads4mo agoHugging Face25ud-synthetic /lithuanian-passports Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes. Introduction - Lithuania The Synthetic Lithuania Passports Dataset brings together more than 1,000 AI-generated passport images built for training OCR and computer vision systems on identity documents. Every record is fully synthetic… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/lithuanian-passports.textimage-to-textn<1K1 likes9 downloads2mo agoHugging Face26kaboomkabow /lithuanian-grocery-storesimagen<1K0 likes6 downloads4mo agoHugging Face27engineerai /lithuanian-estore-sentiment-binarytabular10K<n<100K0 likes3 downloads4mo agoHugging Face28engineerai /lithuanian-fake-review-detectiontext100K<n<1M0 likes2 downloads4mo agoHugging Face29Shaagun /English_Lithuanian_context0 likes1 downloads2y agoHugging Face30Shaagun /Instruction_Lithuanian_Englishtextn<1K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.