Lithuanian
lithuanian-phone-speech-liepa-3-429h-punctuated
Lithuanian Phone Speech 429 h: punctuated, cased, numbers as digits (written form)
Transcripts are in written form, not normalised: punctuation, capitalisation, and numbers,
dates, times and amounts as digits ("2026 m. rugsėjo 6 d., 9:30", "65 000 €", "12,5 %").
The original normalised text is included too.
text
text_normalized
Varšuva 85 % sugriauta.
varšuva aštuoniasdešim penki procentai sugriauta
Keliais eurais arba 10 € daugiau kaip valytojos.
keliais eurais… See the full description on the dataset page: https://huggingface.co/datasets/Digisensus/lithuanian-phone-speech-liepa-3-429h-punctuated.lithuanian-speech-datasetlithuanian-tts-votesTranslation-English-Lithuanianalpaca-lithuanian-cleanedThis repository contains the dataset used for the TaCo paper.
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation
@inproceedings{upadhayay2024taco,
title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes},
author={Bibek Upadhayay and Vahid Behzadan},
booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-lithuanian-cleaned.lithuanian-qa-v1
Dataset Card for Lithuanian QA V1
1. General Information
Dataset Name: Lithuanian QA V1
Dataset Description: This dataset consists of question-answer pairs in Lithuanian, focusing on topics related to Lithuanian culture, history, and people. It is a unique resource designed to aid in the development of language models specifically tailored for Lithuanian linguistic nuances.
Purpose of the Dataset: The primary purpose of this dataset is to facilitate the fine-tuning of… See the full description on the dataset page: https://huggingface.co/datasets/neurotechnology/lithuanian-qa-v1.
