datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
thai-place-name-romanisation-benchmark
Thai romanisation, scored on the words place names are made of
A romaniser can do well on dictionary words and still misread a road sign.
This dataset scores one engine on both, so the difference is a number rather
than an impression.
Every distinct pure-Thai word in
pythainlp/thai-romanization-dataset
is romanised by translit.py,
the rule-based RTGS engine behind motdang.net, twice:
with its curated exception lexicon
and with the rules alone.
Results
slice… See the full description on the dataset page: https://huggingface.co/datasets/NaNoBotCo/thai-place-name-romanisation-benchmark.thai-rtgs-romanisation-lexicon
Thai RTGS romanisation — the exception lexicon
694 Thai forms whose romanisation cannot be derived from the
spelling, each with the reading a rule-based romaniser should use instead.
This is the companion to a rule engine, not a replacement for one. The rules
handle the regular cases; this names the irregulars. Measured: it takes the
romaniser behind motdang.net from 77.51% to 79.62%
agreement on 153,948 Thai words, and from 83.45% to 84.22% on the 21,425 of
them that appear in… See the full description on the dataset page: https://huggingface.co/datasets/NaNoBotCo/thai-rtgs-romanisation-lexicon.Romanian_updatedRomanian_betterFairytaleQA-translated-romanian
Dataset Card for FairytaleQA-translated-ptBR
Dataset Summary
This repository contains the Romanian machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-romanian.romani_compasito
Romani Prompts (Kompasito)
Prompt/context pairs derived from Kompasito — Manual on Human Rights Education for Children (Council of Europe), translated into Romani. Built for adapting LLMs to Romani via SFT/DPO.
Structure
Each row is a JSON object:
prompt — an English instruction sampled from one of three buckets: comprehension, translation, or generative.
context — a Romani text chunk (~1000 chars, 150-char overlap) from the source PDF.
Each surviving chunk appears in 3… See the full description on the dataset page: https://huggingface.co/datasets/Ilia-Iliev/romani_compasito.Romanianllama-7b-romanized-hindiadaption-romani-kompasito-prompts
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-romani_kompasito_prompts
This dataset contains prompt-context pairs derived from the Council of Europe's 'Kompasito' manual on human rights education, translated into Romani. It features English instructions for comprehension, translation, and generation tasks paired with Romani text chunks to facilitate supervised fine-tuning and direct preference optimization for… See the full description on the dataset page: https://huggingface.co/datasets/Ilia-Iliev/adaption-romani-kompasito-prompts.Romanian-Legal-Sentences-Coltucromani-interviewRomanian_Languageromani-sketches-dpoThis dataset consists of parallel text pairs for translating sentences from colloquial spoken Romani to Bulgarian. The style is conversational and the samples are taken from short dialogue videos. It is formatted as instruction-completion pairs suitable for training machine translation models.
nepal-romanized-restaurant-reviewsromanian-grammar-questionsromani-sketchesromanized_casualGrounded_HPV_Cervical_Cancer_Romanized
Grounded HPV & Cervical Cancer — Nepali (Romanized, Fixed) ShareGPT Dataset
File: hpv_fixed.jsonl
Format: JSON Lines (.jsonl), one JSON object per line
Conversation schema: ShareGPT ("from": "human" / "from": "gpt")
Language: Nepali (ne / ISO 639-3 npi), written in romanized script (Latin letters), not Devanagari
License: CC-BY-4.0
Total records: 24,597
File size: ~47 MB
1. What this dataset is
This is a synthetic, fact-grounded, multiple-choice-question (MCQ)… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Grounded_HPV_Cervical_Cancer_Romanized.
