datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
romanian-corpus
Romanian Text Corpus
A comprehensive, high-quality Romanian text corpus for language model pretraining.
Built by collecting and cleaning text from five Romanian-language sources.
Dataset Summary
Total documents: 19,886,412
Estimated tokens: ~20.8B
Language: Romanian (ro)
Format: Parquet (zstd compressed)
Source Breakdown
Source
Documents
mC4
16,875,310
OSCAR-2109
881,722
OSCAR-2301
704,312
OSCAR-2019
703,991
OSCAR-2201
439,778
wikipedia… See the full description on the dataset page: https://huggingface.co/datasets/15juneee/romanian-corpus.fineweb2-romanian-shardsromanian-speech-v2
Research Use Only — This dataset is released strictly for personal research and educational
purposes. The processing pipeline and all scripts are fully open source, but the underlying audio
originates from sources with varying copyrights. Only short fragments were used under fair use
provisions and EU Copyright Directive Art. 3 (text and data mining for scientific research).
This dataset must not be used for redistribution of the source material, commercial purposes,
or training commercially… See the full description on the dataset page: https://huggingface.co/datasets/eduardem/romanian-speech-v2.moldovan-dialectal-romanian-speech-corpus
Moldovan Dialectal Romanian Educational Speech Corpus
This dataset contains aligned Romanian educational speech with Moldovan
dialectal characteristics. It was constructed from publicly accessible lesson
videos recorded by teachers from the Republic of Moldova and published through
the EducatieOnline platform.
The corpus supports research on automatic speech recognition (ASR),
text-to-speech synthesis (TTS), forced alignment, and low-resource dialectal
speech processing.… See the full description on the dataset page: https://huggingface.co/datasets/FraPiz/moldovan-dialectal-romanian-speech-corpus.romanian-baccalaureate-mathematics
Romanian Baccalaureate in Mathematics
A curated collection of Romanian Baccalaureate (BAC) mathematics examination papers and answer keys, transcribed from PDF to structured Markdown using Vision-Language Model OCR. Currently the years 2019 - 2025 were added, more will be processed soon.
Directory Structure
romanian-baccalaureate-mathematics/
├── metadata.csv # Index of all exam papers
├── pdfs/ # Original PDF files
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/asandeistefan/romanian-baccalaureate-mathematics.TTS-Romanian
TTS-Romanian
A large-scale, high-quality Romanian speech dataset for text-to-speech and automatic speech recognition.
Data Source
Derived from CartiaAudio.eu — Romanian audiobooks.
Dataset Statistics
Metric
Value
Total samples
267,410
Total duration
720 hours
Unique speakers
456
Average duration
9.7 seconds
Average DNSMOS
3.84
Features
Field
Type
Description
__key__
string
Unique sample identifier
mp3
Audio
Audio… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Romanian.romanian_speech_dataset_with_15_percent_6_speakers_synthetic_dataromanian_speech_dataset_with_20_percent_4_speakers_synthetic_dataromanian_speech_synthesis_0_8_1\
The Romanian speech synthesis (RSS) corpus was recorded in a hemianechoic chamber (anechoic walls and ceiling; floor partially anechoic) at the University of Edinburgh. We used three high quality studio microphones: a Neumann u89i (large diaphragm condenser), a Sennheiser MKH 800 (small diaphragm condenser with very wide bandwidth) and a DPA 4035 (headset-mounted condenser). Although the current release includes only speech data recorded via Sennheiser MKH 800, we may release speech data recorded via other microphones in the future. All recordings were made at 96 kHz sampling frequency and 24 bits per sample, then downsampled to 48 kHz sampling frequency. For recording, downsampling and bit rate conversion, we used ProTools HD hardware and software. We conducted 8 sessions over the course of a month, recording about 500 sentences in each session. At the start of each session, the speaker listened to a previously recorded sample, in order to attain a similar voice quality and intonation.romanian-name-days
Romanian Name Days and Holidays
Zile onomastice și sărbători românești — the Romanian name-day calendar as
structured data.
In Romania, ziua onomastică — the feast day of the saint whose name you bear —
is widely celebrated, often more than a birthday. Until now this information
existed online only as HTML pages built for human readers. This is the
machine-readable version.
Published by trends.ro.
Dataset summary
Names
86 (46 masculine, 40 feminine)… See the full description on the dataset page: https://huggingface.co/datasets/radool/romanian-name-days.romanian-tts-single-speaker
Romanian TTS Single Speaker
A single-speaker Romanian speech dataset for TTS model training.
Dataset Description
Segments
24,379
Duration
34.3 hours
Speaker
Sanda (female)
Language
Romanian (ro)
Audio
WAV, 16-bit, mono, 24 kHz
Subsets
Subset
Segments
Description
standard
24,203
Standard Romanian sentences
loanword
176
Sentences containing foreign loanwords
Dataset Structure
Column
Type… See the full description on the dataset page: https://huggingface.co/datasets/eduardem/romanian-tts-single-speaker.lili-romanian-single-speaker-piper-cleancommon_voice_17_0_romanian_speech_synthesisRomanian-finepdfs
Romanian PDFs - Processed Dataset
This is a processed and filtered version of the Romanian subset from the FinepdFs dataset, containing high-quality Romanian PDF documents extracted from Common Crawl. The dataset has been filtered for quality (full_doc_lid_score ≥ 0.5) and optimized by removing redundant metadata columns.
Dataset Overview
Total Documents: 3,254,816
Total Size: ~24.32 GB (compressed parquet with ZSTD)
Language: Romanian (ron_Latn)
Source: FinepdFs (Common… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/Romanian-finepdfs.romanianspeechlili-romanian-single-speaker-piper
Lili Romanian Single-Speaker Piper Dataset
A curated Romanian single-speaker speech dataset prepared for Piper training.
Segments
10,738
Total duration
22.91 hours
Speaker
Lili
Gender
female
Language
Romanian (ro)
Audio format
WAV, 16-bit, mono, 22.05 kHz
Segment duration
2.52 - 9.99 seconds
Summary
This dataset contains a single Romanian narrator exposed as Lili.
It is published as a Hugging Face Parquet-backed audio dataset, so the Hub… See the full description on the dataset page: https://huggingface.co/datasets/eduardem/lili-romanian-single-speaker-piper.common_voice_16_1_romanian_speech_synthesisromanian_speech_dataset_with_40_percent_8_speakers_synthetic_dataRomanianReviewsSentiment
RomanianReviewsSentiment
An MTEB dataset
Massive Text Embedding Benchmark
LaRoSeDa (A Large Romanian Sentiment Data Set) contains 15,000 reviews written in Romanian
Task category
t2c
Domains
Reviews, Written
Referencehttps://arxiv.org/abs/2101.04197
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_task("RomanianReviewsSentiment")
evaluator = mteb.MTEB([task])
model… See the full description on the dataset page: https://huggingface.co/datasets/mteb/RomanianReviewsSentiment.common_voice_romanian_speech_synthesisRomanian_bettercemrc-romanian-ner-mrc
CEMRC Romanian NER MRC Dataset
This dataset repository contains MRC-style conversions for Romanian NER datasets used in the CEMRC thesis experiments.
Included datasets: ronec, legalnero, simonero.
Format
Split files are stored as parquet files under one folder per source dataset.
Each row contains:
example_id
sentence_id
query_id
source_dataset
source_hf_dataset
split
query_style
query_sampling
negatives
context_tokens
context
question
entity_type
answers.text… See the full description on the dataset page: https://huggingface.co/datasets/xd-br0/cemrc-romanian-ner-mrc.alpaca_romanian_tacoThis repository contains the dataset used for the TaCo paper.
The dataset follows the style outlined in the TaCo paper, as follows:
{
"instruction": "instruction in xx",
"input": "input in xx",
"output": "Instruction in English: instruction in en ,
Response in English: response in en ,
Response in xx: response in xx "
}
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_romanian_taco.Romanian_updatedromanian_reviews_sentimentromanian-legal-faq-2026
Dataset: Romanian Legal FAQ 2026 (Coltuc Legal Knowledge Base)
Descriere
Set de date structurat în limba română conținând instrucțiuni, întrebări frecvente și soluții procedurale din dreptul civil, drept bancar (clauze abuzive, executări silite), dreptul muncii și dreptul pensiilor.
Dataset-ul este optimizat pentru fine-tuning LLM, sisteme RAG (Retrieval-Augmented Generation) și modele de asistență juridică automată.
Autor și Proprietate Intelectuală… See the full description on the dataset page: https://huggingface.co/datasets/Coltuc2026/romanian-legal-faq-2026.romanian-legal-jurisprudence-2026
Script dezvoltat de Cabinet Avocat Coltuc (2026)
Sursă oficială: https://coltuc.ro | Contact WhatsApp: 0745150894
import pandas as pd
import json
from datetime import datetime
1. Structura datelor juridice (Exemplu de colectare a arhivei)
data = [
{
"id": "COLTUC-2026-001",
"title": "Cum pot să opresc o executare silită în 2026? Răspunsul oferit de Avocat Marius Vicențiu Coltuc",
"category": "Executări Silite"… See the full description on the dataset page: https://huggingface.co/datasets/Coltuc2026/romanian-legal-jurisprudence-2026.german_romanian_mixFairytaleQA-translated-romanian
Dataset Card for FairytaleQA-translated-ptBR
Dataset Summary
This repository contains the Romanian machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-romanian.PAN12_predatorTask_romanianTranslation
Dataset Card for Dataset Name
A part of the PAN-2012 dataset as translated in the Romanian language using automated tools for predator detection and automated translation comparison.
Dataset Details
Dataset Description
This datasets were created based on the training and testing datasets presented at the PAN12 competition (https://pan.webis.de/clef12/pan12-web/sexual-predator-identification.html) which was centered around sexual harassment prevention and… See the full description on the dataset page: https://huggingface.co/datasets/CristinaMierla/PAN12_predatorTask_romanianTranslation.
