datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
romanian-corpus
Romanian Text Corpus
A comprehensive, high-quality Romanian text corpus for language model pretraining.
Built by collecting and cleaning text from five Romanian-language sources.
Dataset Summary
Total documents: 19,886,412
Estimated tokens: ~20.8B
Language: Romanian (ro)
Format: Parquet (zstd compressed)
Source Breakdown
Source
Documents
mC4
16,875,310
OSCAR-2109
881,722
OSCAR-2301
704,312
OSCAR-2019
703,991
OSCAR-2201
439,778
wikipedia… See the full description on the dataset page: https://huggingface.co/datasets/15juneee/romanian-corpus.fineweb2-romanian-shardsromanian-speech-v2
Research Use Only — This dataset is released strictly for personal research and educational
purposes. The processing pipeline and all scripts are fully open source, but the underlying audio
originates from sources with varying copyrights. Only short fragments were used under fair use
provisions and EU Copyright Directive Art. 3 (text and data mining for scientific research).
This dataset must not be used for redistribution of the source material, commercial purposes,
or training commercially… See the full description on the dataset page: https://huggingface.co/datasets/eduardem/romanian-speech-v2.moldovan-dialectal-romanian-speech-corpus
Moldovan Dialectal Romanian Educational Speech Corpus
This dataset contains aligned Romanian educational speech with Moldovan
dialectal characteristics. It was constructed from publicly accessible lesson
videos recorded by teachers from the Republic of Moldova and published through
the EducatieOnline platform.
The corpus supports research on automatic speech recognition (ASR),
text-to-speech synthesis (TTS), forced alignment, and low-resource dialectal
speech processing.… See the full description on the dataset page: https://huggingface.co/datasets/FraPiz/moldovan-dialectal-romanian-speech-corpus.TTS-Romanian
TTS-Romanian
A large-scale, high-quality Romanian speech dataset for text-to-speech and automatic speech recognition.
Data Source
Derived from CartiaAudio.eu — Romanian audiobooks.
Dataset Statistics
Metric
Value
Total samples
267,410
Total duration
720 hours
Unique speakers
456
Average duration
9.7 seconds
Average DNSMOS
3.84
Features
Field
Type
Description
__key__
string
Unique sample identifier
mp3
Audio
Audio… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Romanian.romanian_speech_dataset_with_15_percent_6_speakers_synthetic_dataromanian_speech_dataset_with_20_percent_4_speakers_synthetic_dataromanian-name-days
Romanian Name Days and Holidays
Zile onomastice și sărbători românești — the Romanian name-day calendar as
structured data.
In Romania, ziua onomastică — the feast day of the saint whose name you bear —
is widely celebrated, often more than a birthday. Until now this information
existed online only as HTML pages built for human readers. This is the
machine-readable version.
Published by trends.ro.
Dataset summary
Names
86 (46 masculine, 40 feminine)… See the full description on the dataset page: https://huggingface.co/datasets/radool/romanian-name-days.romanian-tts-single-speaker
Romanian TTS Single Speaker
A single-speaker Romanian speech dataset for TTS model training.
Dataset Description
Segments
24,379
Duration
34.3 hours
Speaker
Sanda (female)
Language
Romanian (ro)
Audio
WAV, 16-bit, mono, 24 kHz
Subsets
Subset
Segments
Description
standard
24,203
Standard Romanian sentences
loanword
176
Sentences containing foreign loanwords
Dataset Structure
Column
Type… See the full description on the dataset page: https://huggingface.co/datasets/eduardem/romanian-tts-single-speaker.lili-romanian-single-speaker-piper-cleancommon_voice_17_0_romanian_speech_synthesisRomanian-finepdfs
Romanian PDFs - Processed Dataset
This is a processed and filtered version of the Romanian subset from the FinepdFs dataset, containing high-quality Romanian PDF documents extracted from Common Crawl. The dataset has been filtered for quality (full_doc_lid_score ≥ 0.5) and optimized by removing redundant metadata columns.
Dataset Overview
Total Documents: 3,254,816
Total Size: ~24.32 GB (compressed parquet with ZSTD)
Language: Romanian (ron_Latn)
Source: FinepdFs (Common… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/Romanian-finepdfs.lili-romanian-single-speaker-piper
Lili Romanian Single-Speaker Piper Dataset
A curated Romanian single-speaker speech dataset prepared for Piper training.
Segments
10,738
Total duration
22.91 hours
Speaker
Lili
Gender
female
Language
Romanian (ro)
Audio format
WAV, 16-bit, mono, 22.05 kHz
Segment duration
2.52 - 9.99 seconds
Summary
This dataset contains a single Romanian narrator exposed as Lili.
It is published as a Hugging Face Parquet-backed audio dataset, so the Hub… See the full description on the dataset page: https://huggingface.co/datasets/eduardem/lili-romanian-single-speaker-piper.common_voice_16_1_romanian_speech_synthesisromanian_speech_dataset_with_40_percent_8_speakers_synthetic_dataRomanianReviewsSentiment
RomanianReviewsSentiment
An MTEB dataset
Massive Text Embedding Benchmark
LaRoSeDa (A Large Romanian Sentiment Data Set) contains 15,000 reviews written in Romanian
Task category
t2c
Domains
Reviews, Written
Referencehttps://arxiv.org/abs/2101.04197
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_task("RomanianReviewsSentiment")
evaluator = mteb.MTEB([task])
model… See the full description on the dataset page: https://huggingface.co/datasets/mteb/RomanianReviewsSentiment.common_voice_romanian_speech_synthesisRomanian_bettercemrc-romanian-ner-mrc
CEMRC Romanian NER MRC Dataset
This dataset repository contains MRC-style conversions for Romanian NER datasets used in the CEMRC thesis experiments.
Included datasets: ronec, legalnero, simonero.
Format
Split files are stored as parquet files under one folder per source dataset.
Each row contains:
example_id
sentence_id
query_id
source_dataset
source_hf_dataset
split
query_style
query_sampling
negatives
context_tokens
context
question
entity_type
answers.text… See the full description on the dataset page: https://huggingface.co/datasets/xd-br0/cemrc-romanian-ner-mrc.alpaca_romanian_tacoThis repository contains the dataset used for the TaCo paper.
The dataset follows the style outlined in the TaCo paper, as follows:
{
"instruction": "instruction in xx",
"input": "input in xx",
"output": "Instruction in English: instruction in en ,
Response in English: response in en ,
Response in xx: response in xx "
}
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_romanian_taco.Romanian_updatedromanian_reviews_sentimentgerman_romanian_mixFairytaleQA-translated-romanian
Dataset Card for FairytaleQA-translated-ptBR
Dataset Summary
This repository contains the Romanian machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-romanian.romanian-driving-examRomanian-Speech-Dataset
🎧 Romanian Speech Dataset
The Romanian Speech Dataset is a high-quality speech audio dataset designed to support AI and machine learning workflows with diverse and well-structured audio data. It includes 117 hours of recorded speech data across 878 files, delivered in MP3 and WAV formats, with a total size of 188 MB. This carefully curated audio dataset provides balanced and representative voice data, with 54% male and 46% female speakers, and age distribution spanning 18 to 50+… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Romanian-Speech-Dataset.romanian_sentimentEnglish-Romanian-Magpie-Reasoning
English-Romanian Translation Pairs from Magpie-Reasoning
This dataset contains 150,000 high-quality English-Romanian parallel translation pairs derived from the Magpie-Reasoning dataset, specifically designed for training and evaluating machine translation models with a focus on technical, mathematical, and code-related content.
Source Datasets
This dataset is created by aligning:
English: Magpie-Align/Magpie-Reasoning-V1-150K
Romanian: OpenLLM-Ro/ro_sft_magpie_reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/English-Romanian-Magpie-Reasoning.romanian_sa
Sentiment Analysis Data for the Romanian Language
Dataset Description:
This dataset contains a sentiment analysis dataset from Tache et al. (2021).
Data Structure:
The data was used for the project on improving word embeddings with graph knowledge for Low Resource Languages.
Citation:
@inproceedings{tache-etal-2021-clustering,
title = "Clustering Word Embeddings with Self-Organizing Maps. Application on {L}a{R}o{S}e{D}a - A Large {R}omanian Sentiment Data Set",
author =… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/romanian_sa.RomanianSentimentClassification
RomanianSentimentClassification
An MTEB dataset
Massive Text Embedding Benchmark
An Romanian dataset for sentiment classification.
Task category
t2c
Domains
Reviews, Written
Reference
https://arxiv.org/abs/2009.08712
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_task("RomanianSentimentClassification")
evaluator = mteb.MTEB([task])
model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/RomanianSentimentClassification.
