datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FairytaleQA-translated-romanian
Dataset Card for FairytaleQA-translated-ptBR
Dataset Summary
This repository contains the Romanian machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-romanian.thai-romanization-datasetDataset from github.com/wannaphong/thai-romanization/blob/master/dataset/data.csv
Devnagari-Romanized-Pair
Dataset Overview
The dataset devanagari romanized pair contains, 959 rows, where each row has one English sentence and its corresponding Nepali translations both in the devanagari script and in romanized format, the size of the data set is less than 1,000 elements and it's designed for use in Translation, text generation and text to text generation tasks.
English-Romanian-Magpie-Reasoning
English-Romanian Translation Pairs from Magpie-Reasoning
This dataset contains 150,000 high-quality English-Romanian parallel translation pairs derived from the Magpie-Reasoning dataset, specifically designed for training and evaluating machine translation models with a focus on technical, mathematical, and code-related content.
Source Datasets
This dataset is created by aligning:
English: Magpie-Align/Magpie-Reasoning-V1-150K
Romanian: OpenLLM-Ro/ro_sft_magpie_reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/English-Romanian-Magpie-Reasoning.romani_compasito
Romani Prompts (Kompasito)
Prompt/context pairs derived from Kompasito — Manual on Human Rights Education for Children (Council of Europe), translated into Romani. Built for adapting LLMs to Romani via SFT/DPO.
Structure
Each row is a JSON object:
prompt — an English instruction sampled from one of three buckets: comprehension, translation, or generative.
context — a Romani text chunk (~1000 chars, 150-char overlap) from the source PDF.
Each surviving chunk appears in 3… See the full description on the dataset page: https://huggingface.co/datasets/Ilia-Iliev/romani_compasito.
