datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
romanian-name-days
Romanian Name Days and Holidays
Zile onomastice și sărbători românești — the Romanian name-day calendar as
structured data.
In Romania, ziua onomastică — the feast day of the saint whose name you bear —
is widely celebrated, often more than a birthday. Until now this information
existed online only as HTML pages built for human readers. This is the
machine-readable version.
Published by trends.ro.
Dataset summary
Names
86 (46 masculine, 40 feminine)… See the full description on the dataset page: https://huggingface.co/datasets/radool/romanian-name-days.Romanian-Speech-Dataset
🎧 Romanian Speech Dataset
The Romanian Speech Dataset is a high-quality speech audio dataset designed to support AI and machine learning workflows with diverse and well-structured audio data. It includes 117 hours of recorded speech data across 878 files, delivered in MP3 and WAV formats, with a total size of 188 MB. This carefully curated audio dataset provides balanced and representative voice data, with 54% male and 46% female speakers, and age distribution spanning 18 to 50+… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Romanian-Speech-Dataset.romanian_sa
Sentiment Analysis Data for the Romanian Language
Dataset Description:
This dataset contains a sentiment analysis dataset from Tache et al. (2021).
Data Structure:
The data was used for the project on improving word embeddings with graph knowledge for Low Resource Languages.
Citation:
@inproceedings{tache-etal-2021-clustering,
title = "Clustering Word Embeddings with Self-Organizing Maps. Application on {L}a{R}o{S}e{D}a - A Large {R}omanian Sentiment Data Set",
author =… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/romanian_sa.romanianreddit_authorship_profiling_romanianromanian-sentiment-reviews-5karomanian-romanian-MT-corpusThis is the biggest and most comprehensive Romanian - Aromanian parallel corpus.
The Aromanian counterpart is automatically converted to Cunia and DIARO standards
More details about its collection and how it was used for the AroTranslate system can be found in our paper.
If you find our work usefull, please cite:
@article{jerpelea2024dialectal,
title={Dialectal and Low-Resource Machine Translation for Aromanian},
author={Jerpelea, Alexandru-Iulius and Rădoi, Alina and Nisioi, Sergiu}… See the full description on the dataset page: https://huggingface.co/datasets/aronlp/aromanian-romanian-MT-corpus.aromanian-romanian-MT-corpus-limited
About
This is a limited version (i.e. some text sources have been exluded) of the comprehensive Romanian - Aromanian parallel corpus (which you can find in this repo, too).
The Aromanian counterpart is automatically converted to Cunia and DIARO standards.
More details about its collection and how it was used for the AroTranslate system can be found in this paper and this paper.
Disclaimer
Aromanian is a very low resource language and is not standardized, having several… See the full description on the dataset page: https://huggingface.co/datasets/aronlp/aromanian-romanian-MT-corpus-limited.
