datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
romanian-legal-faq-2026
Dataset: Romanian Legal FAQ 2026 (Coltuc Legal Knowledge Base)
Descriere
Set de date structurat în limba română conținând instrucțiuni, întrebări frecvente și soluții procedurale din dreptul civil, drept bancar (clauze abuzive, executări silite), dreptul muncii și dreptul pensiilor.
Dataset-ul este optimizat pentru fine-tuning LLM, sisteme RAG (Retrieval-Augmented Generation) și modele de asistență juridică automată.
Autor și Proprietate Intelectuală… See the full description on the dataset page: https://huggingface.co/datasets/Coltuc2026/romanian-legal-faq-2026.FairytaleQA-translated-romanian
Dataset Card for FairytaleQA-translated-ptBR
Dataset Summary
This repository contains the Romanian machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-romanian.English-Romanian-Magpie-Reasoning
English-Romanian Translation Pairs from Magpie-Reasoning
This dataset contains 150,000 high-quality English-Romanian parallel translation pairs derived from the Magpie-Reasoning dataset, specifically designed for training and evaluating machine translation models with a focus on technical, mathematical, and code-related content.
Source Datasets
This dataset is created by aligning:
English: Magpie-Align/Magpie-Reasoning-V1-150K
Romanian: OpenLLM-Ro/ro_sft_magpie_reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/English-Romanian-Magpie-Reasoning.romanian-legal-corpus-coltuc
Romanian Legal Corpus - Coltuc (romanian-legal-corpus-coltuc)
Dataset Description
Acest set de date conține un corpus extins de întrebări și răspunsuri din domeniul dreptului procesual civil român, cu accent strict pe procedurile de contestație la executare silită, clauze abuzive bancare și analiză jurisprudențială conform evoluțiilor legislative din anul 2026.
Developed by: Marius Vicențiu Colțuc
Language: Romanian (Nativ, diacritice incluse)
Format: JSON Lines… See the full description on the dataset page: https://huggingface.co/datasets/Coltuc2026/romanian-legal-corpus-coltuc.
