CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NaNoBotCo /thai-place-name-romanisation-benchmark Thai romanisation, scored on the words place names are made of A romaniser can do well on dictionary words and still misread a road sign. This dataset scores one engine on both, so the difference is a number rather than an impression. Every distinct pure-Thai word in pythainlp/thai-romanization-dataset is romanised by translit.py, the rule-based RTGS engine behind motdang.net, twice: with its curated exception lexicon and with the rules alone. Results slice… See the full description on the dataset page: https://huggingface.co/datasets/NaNoBotCo/thai-place-name-romanisation-benchmark.text100K<n<1M0 likes44 downloads15d agoHugging Face02NaNoBotCo /thai-rtgs-romanisation-lexicon Thai RTGS romanisation — the exception lexicon 694 Thai forms whose romanisation cannot be derived from the spelling, each with the reading a rule-based romaniser should use instead. This is the companion to a rule engine, not a replacement for one. The rules handle the regular cases; this names the irregulars. Measured: it takes the romaniser behind motdang.net from 77.51% to 79.62% agreement on 153,948 Thai words, and from 83.45% to 84.22% on the 21,425 of them that appear in… See the full description on the dataset page: https://huggingface.co/datasets/NaNoBotCo/thai-rtgs-romanisation-lexicon.textn<1K0 likes39 downloads15d agoHugging Face03Gargaz /Romanian_updatedtext10K<n<100K1 likes36 downloads2y agoHugging Face04Gargaz /Romanian_bettertext10K<n<100K2 likes36 downloads2y agoHugging Face05benjleite /FairytaleQA-translated-romanian Dataset Card for FairytaleQA-translated-ptBR Dataset Summary This repository contains the Romanian machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-romanian.textquestion-answering10K<n<100K1 likes33 downloads1y agoHugging Face06Ilia-Iliev /romani_compasito Romani Prompts (Kompasito) Prompt/context pairs derived from Kompasito — Manual on Human Rights Education for Children (Council of Europe), translated into Romani. Built for adapting LLMs to Romani via SFT/DPO. Structure Each row is a JSON object: prompt — an English instruction sampled from one of three buckets: comprehension, translation, or generative. context — a Romani text chunk (~1000 chars, 150-char overlap) from the source PDF. Each surviving chunk appears in 3… See the full description on the dataset page: https://huggingface.co/datasets/Ilia-Iliev/romani_compasito.texttext-generation1K<n<10K0 likes18 downloads5mo agoHugging Face07Gargaz /Romaniantext1K<n<10K1 likes14 downloads2y agoHugging Face08Shreesh03 /llama-7b-romanized-hinditextn<1K1 likes10 downloads3y agoHugging Face09Ilia-Iliev /adaption-romani-kompasito-prompts This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. adaption-romani_kompasito_prompts This dataset contains prompt-context pairs derived from the Council of Europe's 'Kompasito' manual on human rights education, translated into Romani. It features English instructions for comprehension, translation, and generation tasks paired with Romani text chunks to facilitate supervised fine-tuning and direct preference optimization for… See the full description on the dataset page: https://huggingface.co/datasets/Ilia-Iliev/adaption-romani-kompasito-prompts.textn<1K0 likes10 downloads5mo agoHugging Face10Coltuc2026 /Romanian-Legal-Sentences-Coltuctextn<1K1 likes10 downloads4mo agoHugging Face11Ilia-Iliev /romani-interviewtextn<1K0 likes9 downloads5mo agoHugging Face12wop /Romanian_Languagetextn<1K0 likes7 downloads2y agoHugging Face13Ilia-Iliev /romani-sketches-dpoThis dataset consists of parallel text pairs for translating sentences from colloquial spoken Romani to Bulgarian. The style is conversational and the samples are taken from short dialogue videos. It is formatted as instruction-completion pairs suitable for training machine translation models. texttranslationn<1K0 likes7 downloads5mo agoHugging Face14amirpoudel /nepal-romanized-restaurant-reviewstextn<1K0 likes3 downloads2y agoHugging Face15mihaeladanila /romanian-grammar-questionstext1K<n<10K2 likes3 downloads1y agoHugging Face16Ilia-Iliev /romani-sketchestext1K<n<10K0 likes3 downloads5mo agoHugging Face17cs23s036 /romanized_casualgatedtext1M<n<10M0 likes1 downloads1y agoHugging Face18sabin1234 /Grounded_HPV_Cervical_Cancer_Romanized Grounded HPV & Cervical Cancer — Nepali (Romanized, Fixed) ShareGPT Dataset File: hpv_fixed.jsonl Format: JSON Lines (.jsonl), one JSON object per line Conversation schema: ShareGPT ("from": "human" / "from": "gpt") Language: Nepali (ne / ISO 639-3 npi), written in romanized script (Latin letters), not Devanagari License: CC-BY-4.0 Total records: 24,597 File size: ~47 MB 1. What this dataset is This is a synthetic, fact-grounded, multiple-choice-question (MCQ)… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Grounded_HPV_Cervical_Cancer_Romanized.text10K<n<100K0 likes1h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.