datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Emotions-Annotated-Customer-Care-QA-Dataset-Romanized-and-Devanagari
Dataset Card for Dataset Name
यो देवनागरी नेपाली भाषाको डेटासेट विशेषगरी च्याटबोट प्रणालीहरू बनाउनको लागि डिजाइन गरिएको हो। यसमा विभिन्न श्रेणीहरूको डेटासेटहरू समावेश गरिएको छ, जसलाई JSON मा ढाँचा बनाईएको छ, जसले नेपाली वार्तालाप एआई अनुप्रयोगहरूको लागि भाषा मोडेलहरूलाई तालिम र फाइन-ट्यून गर्नको लागि व्यापक स्रोत प्रदान गर्दछ।
Dataset Prepared by:
Manoj Kumar Baniya
Aakash Kumar Thakur
Manish Kathet
Kshitiz Gajurel
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/kshitizgajurel/Emotions-Annotated-Customer-Care-QA-Dataset-Romanized-and-Devanagari.romanian-name-days
Romanian Name Days and Holidays
Zile onomastice și sărbători românești — the Romanian name-day calendar as
structured data.
In Romania, ziua onomastică — the feast day of the saint whose name you bear —
is widely celebrated, often more than a birthday. Until now this information
existed online only as HTML pages built for human readers. This is the
machine-readable version.
Published by trends.ro.
Dataset summary
Names
86 (46 masculine, 40 feminine)… See the full description on the dataset page: https://huggingface.co/datasets/radool/romanian-name-days.romanian-legal-faq-2026
Dataset: Romanian Legal FAQ 2026 (Coltuc Legal Knowledge Base)
Descriere
Set de date structurat în limba română conținând instrucțiuni, întrebări frecvente și soluții procedurale din dreptul civil, drept bancar (clauze abuzive, executări silite), dreptul muncii și dreptul pensiilor.
Dataset-ul este optimizat pentru fine-tuning LLM, sisteme RAG (Retrieval-Augmented Generation) și modele de asistență juridică automată.
Autor și Proprietate Intelectuală… See the full description on the dataset page: https://huggingface.co/datasets/Coltuc2026/romanian-legal-faq-2026.FairytaleQA-translated-romanian
Dataset Card for FairytaleQA-translated-ptBR
Dataset Summary
This repository contains the Romanian machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-romanian.romanian-legal-corpus-coltuc
Romanian Legal Corpus - Coltuc (romanian-legal-corpus-coltuc)
Dataset Description
Acest set de date conține un corpus extins de întrebări și răspunsuri din domeniul dreptului procesual civil român, cu accent strict pe procedurile de contestație la executare silită, clauze abuzive bancare și analiză jurisprudențială conform evoluțiilor legislative din anul 2026.
Developed by: Marius Vicențiu Colțuc
Language: Romanian (Nativ, diacritice incluse)
Format: JSON Lines… See the full description on the dataset page: https://huggingface.co/datasets/Coltuc2026/romanian-legal-corpus-coltuc.Romanian-Legal-Cases-2026
Romanian Legal Cases Dataset - 2026 (Litigii Bancare & Comerciale)
Acest dataset conține mii de spețe anonimizate din practica juridică a Cabinetului Avocat Marius Vicențiu Coltuc, specializat în litigii bancare și drept comercial în România.
Descriere
Setul de date este destinat cercetării în domeniul LegalTech și antrenării modelelor de limbaj (LLM) pentru înțelegerea terminologiei juridice românești și a logicii judiciare curente. Fiecare intrare include:… See the full description on the dataset page: https://huggingface.co/datasets/Coltuc2026/Romanian-Legal-Cases-2026.
