CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NaNoBotCo /thai-place-name-romanisation-benchmark Thai romanisation, scored on the words place names are made of A romaniser can do well on dictionary words and still misread a road sign. This dataset scores one engine on both, so the difference is a number rather than an impression. Every distinct pure-Thai word in pythainlp/thai-romanization-dataset is romanised by translit.py, the rule-based RTGS engine behind motdang.net, twice: with its curated exception lexicon and with the rules alone. Results slice… See the full description on the dataset page: https://huggingface.co/datasets/NaNoBotCo/thai-place-name-romanisation-benchmark.text100K<n<1M0 likes45 downloads17d agoHugging Face02NaNoBotCo /thai-rtgs-romanisation-lexicon Thai RTGS romanisation — the exception lexicon 694 Thai forms whose romanisation cannot be derived from the spelling, each with the reading a rule-based romaniser should use instead. This is the companion to a rule engine, not a replacement for one. The rules handle the regular cases; this names the irregulars. Measured: it takes the romaniser behind motdang.net from 77.51% to 79.62% agreement on 153,948 Thai words, and from 83.45% to 84.22% on the 21,425 of them that appear in… See the full description on the dataset page: https://huggingface.co/datasets/NaNoBotCo/thai-rtgs-romanisation-lexicon.textn<1K0 likes40 downloads17d agoHugging Face03Gargaz /Romanian_updatedtext10K<n<100K1 likes36 downloads2y agoHugging Face04Gargaz /Romanian_bettertext10K<n<100K2 likes35 downloads2y agoHugging Face05benjleite /FairytaleQA-translated-romanian Dataset Card for FairytaleQA-translated-ptBR Dataset Summary This repository contains the Romanian machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-romanian.textquestion-answering10K<n<100K1 likes30 downloads1y agoHugging Face06Ilia-Iliev /romani_compasito Romani Prompts (Kompasito) Prompt/context pairs derived from Kompasito — Manual on Human Rights Education for Children (Council of Europe), translated into Romani. Built for adapting LLMs to Romani via SFT/DPO. Structure Each row is a JSON object: prompt — an English instruction sampled from one of three buckets: comprehension, translation, or generative. context — a Romani text chunk (~1000 chars, 150-char overlap) from the source PDF. Each surviving chunk appears in 3… See the full description on the dataset page: https://huggingface.co/datasets/Ilia-Iliev/romani_compasito.texttext-generation1K<n<10K0 likes20 downloads5mo agoHugging Face07sabin1234 /Grounded_HPV_Cervical_Cancer_Romanized Grounded HPV & Cervical Cancer — Nepali (Romanized, Fixed) ShareGPT Dataset File: hpv_fixed.jsonl Format: JSON Lines (.jsonl), one JSON object per line Conversation schema: ShareGPT ("from": "human" / "from": "gpt") Language: Nepali (ne / ISO 639-3 npi), written in romanized script (Latin letters), not Devanagari License: CC-BY-4.0 Total records: 24,597 File size: ~47 MB 1. What this dataset is This is a synthetic, fact-grounded, multiple-choice-question (MCQ)… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Grounded_HPV_Cervical_Cancer_Romanized.text10K<n<100K0 likes17 downloads2d agoHugging Face08Gargaz /Romaniantext1K<n<10K1 likes14 downloads2y agoHugging Face09Shreesh03 /llama-7b-romanized-hinditextn<1K1 likes10 downloads3y agoHugging Face10Ilia-Iliev /adaption-romani-kompasito-prompts This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. adaption-romani_kompasito_prompts This dataset contains prompt-context pairs derived from the Council of Europe's 'Kompasito' manual on human rights education, translated into Romani. It features English instructions for comprehension, translation, and generation tasks paired with Romani text chunks to facilitate supervised fine-tuning and direct preference optimization for… See the full description on the dataset page: https://huggingface.co/datasets/Ilia-Iliev/adaption-romani-kompasito-prompts.textn<1K0 likes10 downloads5mo agoHugging Face11Coltuc2026 /Romanian-Legal-Sentences-Coltuctextn<1K1 likes10 downloads4mo agoHugging Face12Ilia-Iliev /romani-interviewtextn<1K0 likes9 downloads5mo agoHugging Face13wop /Romanian_Languagetextn<1K0 likes7 downloads2y agoHugging Face14Ilia-Iliev /romani-sketches-dpoThis dataset consists of parallel text pairs for translating sentences from colloquial spoken Romani to Bulgarian. The style is conversational and the samples are taken from short dialogue videos. It is formatted as instruction-completion pairs suitable for training machine translation models. texttranslationn<1K0 likes7 downloads5mo agoHugging Face15amirpoudel /nepal-romanized-restaurant-reviewstextn<1K0 likes3 downloads2y agoHugging Face16Ilia-Iliev /romani-sketchestext1K<n<10K0 likes3 downloads5mo agoHugging Face17mihaeladanila /romanian-grammar-questionstext1K<n<10K2 likes2 downloads1y agoHugging Face18cs23s036 /romanized_casualgatedtext1M<n<10M0 likes1 downloads1y agoHugging Face19sabin1234 /Romanized_Nepali_Drug_Poisoning_Mortality Romanized Nepali Drug Poisoning Mortality - ShareGPT SFT Dataset A synthetic, single-turn instruction-tuning dataset of 56,448 question-answer pairs written in romanized Nepali (Nepali in Latin script). Every pair asks about one U.S. county in one year (1999-2016) and answers with the county's population and its age-adjusted drug-poisoning death-rate category. Source facts are attributed to the U.S. Centers for Disease Control and Prevention (CDC); the text was generated by a… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Romanized_Nepali_Drug_Poisoning_Mortality.text10K<n<100K0 likes12h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.