datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BigEarthNetV2-Lithuania-Summer-LMDB
TU Berlin
RSiM
DIMA
BigEarth
BIFOLD
reBEN — Lithuania Summer Subset (pre-converted to LMDB)
⚠️ Unofficial mirror. This is an unofficial, community-provided pre-conversion of a subset of the BigEarthNet v2.0 (reBEN) dataset into LMDB format. It is provided as a convenience for researchers who wish to get started quickly without running the full conversion pipeline. In case of any discrepancy, the original publication and the original files always take… See the full description on the dataset page: https://huggingface.co/datasets/hackelle/BigEarthNetV2-Lithuania-Summer-LMDB.ipfs_lithuania_laws
Lithuania TAR / e-TAR legal acts
Research snapshot of official national legislation from TAR / e-tar.lt (Teisės aktų registras) official legal acts; leftover Wayback fill of official URLs.
Not legal advice. The official gazette / authentic source prevails over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-17
Coverage
shard leftover Wayback incomplete
Source
TAR / e-tar.lt (Teisės aktų registras) official legal acts; leftover Wayback fill… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_lithuania_laws.lithuanian-phone-speech-liepa-3-429h-punctuated
Lithuanian Phone Speech 429 h: punctuated, cased, numbers as digits (written form)
Transcripts are in written form, not normalised: punctuation, capitalisation, and numbers,
dates, times and amounts as digits ("2026 m. rugsėjo 6 d., 9:30", "65 000 €", "12,5 %").
The original normalised text is included too.
text
text_normalized
Varšuva 85 % sugriauta.
varšuva aštuoniasdešim penki procentai sugriauta
Keliais eurais arba 10 € daugiau kaip valytojos.
keliais eurais… See the full description on the dataset page: https://huggingface.co/datasets/Digisensus/lithuanian-phone-speech-liepa-3-429h-punctuated.lithuanian-speech-datasetzygai_soviet_lithuania
🇱🇹 ZygAI – Soviet Lithuania (1940–1990) Dataset
A comprehensive cultural–historical dataset documenting Soviet-era Lithuania
📌 Overview
ZygAI Soviet Lithuania is a structured, bilingual (LT + EN) dataset capturing political,cultural, economic, educational, technological, and everyday life aspects of Lithuaniaunder Soviet occupation (1940–1990).
This dataset is a major component of ZygAI Research 2025–2026, created to preservehistorical memory, support academic… See the full description on the dataset page: https://huggingface.co/datasets/ZygAI/zygai_soviet_lithuania.alpaca-lithuanian-cleanedThis repository contains the dataset used for the TaCo paper.
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation
@inproceedings{upadhayay2024taco,
title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes},
author={Bibek Upadhayay and Vahid Behzadan},
booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-lithuanian-cleaned.lithuanian-qa-v1
Dataset Card for Lithuanian QA V1
1. General Information
Dataset Name: Lithuanian QA V1
Dataset Description: This dataset consists of question-answer pairs in Lithuanian, focusing on topics related to Lithuanian culture, history, and people. It is a unique resource designed to aid in the development of language models specifically tailored for Lithuanian linguistic nuances.
Purpose of the Dataset: The primary purpose of this dataset is to facilitate the fine-tuning of… See the full description on the dataset page: https://huggingface.co/datasets/neurotechnology/lithuanian-qa-v1.lithuaniaInstruc_Lithuanianlithuanian-dialect-speech-liepa-3-100h-punctuated
Lithuanian Dialect Speech 100 h: punctuated, cased, numbers as digits (written form)
Spontaneous Lithuanian dialect speech from all four regions, with three transcripts per clip:
written form (punctuation, capitalisation, numbers as digits), normalised, and the original
phonetic transcription with stress marks. Dialect word forms are kept as spoken in every layer.
text
text_normalized
text_phonetic
Per 3 klases buvu 10 mokinių.
per tris klases buvu dešim mokinių
per… See the full description on the dataset page: https://huggingface.co/datasets/Digisensus/lithuanian-dialect-speech-liepa-3-100h-punctuated.lithuanian-sentiment-analysisAll samples are reviews scraped from https://atsiliepimai.lt/.
Lithuanian-Speech-Dataset
Lithuanian Dataset Metadata
Field
Value
📜 License
CC BY-NC-ND 4.0
🎯 Task Categories
Automatic Speech Recognition
🌍 Language
Lithuanian (lt)
🏷️ Tags
Lithuanian, Audio, Speech, Speech Recognition, ML, Machine, Machine Learning
📦 Size Category
n < 1K
lithuanian-road-signsLithuanian_Context_QA
Lithuanian QA Dataset - Generated with DSPy & Gemma2 27B Q4
Introduction
This dataset was created using DSPy, a Python framework that simplifies the generation of question and answer (QA) pairs from a given context. The dataset is composed of context, questions, and answers, all in Lithuanian. The context was primarily sourced from the following resources:
Lithuanian Wikipedia (lt.wikipedia.org)
Lietuviškoji enciklopedija (vle.lt)
Book: Vitalija Skėruvienė, Civilinė Teisė Mokomoji… See the full description on the dataset page: https://huggingface.co/datasets/ArturG9/Lithuanian_Context_QA.alpaca_lithuanian_tacoThis repository contains the dataset used for the TaCo paper.
The dataset follows the style outlined in the TaCo paper, as follows:
{
"instruction": "instruction in xx",
"input": "input in xx",
"output": "Instruction in English: instruction in en ,
Response in English: response in en ,
Response in xx: response in xx "
}
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_lithuanian_taco.Lithuania-Stock-Symbols-and-Metadata
Lithuania Stock Symbols & Company Metadata
This dataset contains stock symbols and basic company metadata for all listed companies in Lithuania.It is updated weekly if new changes are there.
📊 Dataset Contents
The dataset is provided as a CSV file with the following columns:
Column
Description
name
Full company name
ticker
Stock ticker symbol (e.g., AAPL, MSFT)
market
The exchange/market where the stock is listed
sector
The primary business sector of… See the full description on the dataset page: https://huggingface.co/datasets/kjhq/Lithuania-Stock-Symbols-and-Metadata.lithuanian-translationsexonyms-for-lithuanian-placesc4-lithuanian-enhancedYodaLingua-Lithuanian
YodaLingua-Lithuanian
YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Lithuanian portion of the multilingual YodaLingua collection.
🧾 Dataset Overview
Property
Value
Total clips
2745 audio–transcription pairs
Total duration
7.4 hours
Speakers
138 distinct speakers
Audio format
MP3 • mono • 24 kHz •… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Lithuanian.lithuanian-passports
Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes.
Introduction - Lithuania
The Synthetic Lithuania Passports Dataset brings together more than 1,000 AI-generated passport images built for training OCR and computer vision systems on identity documents. Every record is fully synthetic… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/lithuanian-passports.lithuanian-grocery-storeslithuanian-estore-sentiment-binarylithuanian-fake-review-detectionInstruction_Lithuanian_English
