CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
0115juneee /romanian-corpus Romanian Text Corpus A comprehensive, high-quality Romanian text corpus for language model pretraining. Built by collecting and cleaning text from five Romanian-language sources. Dataset Summary Total documents: 19,886,412 Estimated tokens: ~20.8B Language: Romanian (ro) Format: Parquet (zstd compressed) Source Breakdown Source Documents mC4 16,875,310 OSCAR-2109 881,722 OSCAR-2301 704,312 OSCAR-2019 703,991 OSCAR-2201 439,778 wikipedia… See the full description on the dataset page: https://huggingface.co/datasets/15juneee/romanian-corpus.text10M<n<100M1 likes897 downloads6mo agoHugging Face02rotarue /fineweb2-romanian-shardstext10M<n<100M1 likes676 downloads10mo agoHugging Face03eduardem /romanian-speech-v2 Research Use Only — This dataset is released strictly for personal research and educational purposes. The processing pipeline and all scripts are fully open source, but the underlying audio originates from sources with varying copyrights. Only short fragments were used under fair use provisions and EU Copyright Directive Art. 3 (text and data mining for scientific research). This dataset must not be used for redistribution of the source material, commercial purposes, or training commercially… See the full description on the dataset page: https://huggingface.co/datasets/eduardem/romanian-speech-v2.audiotext-to-speech100K<n<1M2 likes436 downloads7mo agoHugging Face04FraPiz /moldovan-dialectal-romanian-speech-corpus Moldovan Dialectal Romanian Educational Speech Corpus This dataset contains aligned Romanian educational speech with Moldovan dialectal characteristics. It was constructed from publicly accessible lesson videos recorded by teachers from the Republic of Moldova and published through the EducatieOnline platform. The corpus supports research on automatic speech recognition (ASR), text-to-speech synthesis (TTS), forced alignment, and low-resource dialectal speech processing.… See the full description on the dataset page: https://huggingface.co/datasets/FraPiz/moldovan-dialectal-romanian-speech-corpus.audioautomatic-speech-recognition10K<n<100K1 likes375 downloads1mo agoHugging Face05asandeistefan /romanian-baccalaureate-mathematics Romanian Baccalaureate in Mathematics A curated collection of Romanian Baccalaureate (BAC) mathematics examination papers and answer keys, transcribed from PDF to structured Markdown using Vision-Language Model OCR. Currently the years 2019 - 2025 were added, more will be processed soon. Directory Structure romanian-baccalaureate-mathematics/ ├── metadata.csv # Index of all exam papers ├── pdfs/ # Original PDF files │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/asandeistefan/romanian-baccalaureate-mathematics.n<1K2 likes298 downloads3mo agoHugging Face06datadriven-company /TTS-Romanian TTS-Romanian A large-scale, high-quality Romanian speech dataset for text-to-speech and automatic speech recognition. Data Source Derived from CartiaAudio.eu — Romanian audiobooks. Dataset Statistics Metric Value Total samples 267,410 Total duration 720 hours Unique speakers 456 Average duration 9.7 seconds Average DNSMOS 3.84 Features Field Type Description __key__ string Unique sample identifier mp3 Audio Audio… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Romanian.audiotext-to-speech100K<n<1M3 likes273 downloads7mo agoHugging Face07VladS159 /romanian_speech_dataset_with_15_percent_6_speakers_synthetic_dataaudio10K<n<100K0 likes189 downloads7mo agoHugging Face08VladS159 /romanian_speech_dataset_with_20_percent_4_speakers_synthetic_dataaudio10K<n<100K0 likes133 downloads6mo agoHugging Face09gigant /romanian_speech_synthesis_0_8_1\ The Romanian speech synthesis (RSS) corpus was recorded in a hemianechoic chamber (anechoic walls and ceiling; floor partially anechoic) at the University of Edinburgh. We used three high quality studio microphones: a Neumann u89i (large diaphragm condenser), a Sennheiser MKH 800 (small diaphragm condenser with very wide bandwidth) and a DPA 4035 (headset-mounted condenser). Although the current release includes only speech data recorded via Sennheiser MKH 800, we may release speech data recorded via other microphones in the future. All recordings were made at 96 kHz sampling frequency and 24 bits per sample, then downsampled to 48 kHz sampling frequency. For recording, downsampling and bit rate conversion, we used ProTools HD hardware and software. We conducted 8 sessions over the course of a month, recording about 500 sentences in each session. At the start of each session, the speaker listened to a previously recorded sample, in order to attain a similar voice quality and intonation.automatic-speech-recognition11 likes107 downloads4y agoHugging Face10radool /romanian-name-days Romanian Name Days and Holidays Zile onomastice și sărbători românești — the Romanian name-day calendar as structured data. In Romania, ziua onomastică — the feast day of the saint whose name you bear — is widely celebrated, often more than a birthday. Until now this information existed online only as HTML pages built for human readers. This is the machine-readable version. Published by trends.ro. Dataset summary Names 86 (46 masculine, 40 feminine)… See the full description on the dataset page: https://huggingface.co/datasets/radool/romanian-name-days.textquestion-answeringn<1K1 likes68 downloads1mo agoHugging Face11eduardem /romanian-tts-single-speaker Romanian TTS Single Speaker A single-speaker Romanian speech dataset for TTS model training. Dataset Description Segments 24,379 Duration 34.3 hours Speaker Sanda (female) Language Romanian (ro) Audio WAV, 16-bit, mono, 24 kHz Subsets Subset Segments Description standard 24,203 Standard Romanian sentences loanword 176 Sentences containing foreign loanwords Dataset Structure Column Type… See the full description on the dataset page: https://huggingface.co/datasets/eduardem/romanian-tts-single-speaker.audiotext-to-speech10K<n<100K0 likes67 downloads6mo agoHugging Face12RandomAi99 /lili-romanian-single-speaker-piper-cleanaudio10K<n<100K0 likes59 downloads6mo agoHugging Face13VladS159 /common_voice_17_0_romanian_speech_synthesisaudio10K<n<100K2 likes52 downloads2y agoHugging Face14Yxanul /Romanian-finepdfs Romanian PDFs - Processed Dataset This is a processed and filtered version of the Romanian subset from the FinepdFs dataset, containing high-quality Romanian PDF documents extracted from Common Crawl. The dataset has been filtered for quality (full_doc_lid_score ≥ 0.5) and optimized by removing redundant metadata columns. Dataset Overview Total Documents: 3,254,816 Total Size: ~24.32 GB (compressed parquet with ZSTD) Language: Romanian (ron_Latn) Source: FinepdFs (Common… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/Romanian-finepdfs.tabular1M<n<10M0 likes52 downloads11mo agoHugging Face15maratim /romanianspeechaudion<1K1 likes48 downloads4y agoHugging Face16eduardem /lili-romanian-single-speaker-piper Lili Romanian Single-Speaker Piper Dataset A curated Romanian single-speaker speech dataset prepared for Piper training. Segments 10,738 Total duration 22.91 hours Speaker Lili Gender female Language Romanian (ro) Audio format WAV, 16-bit, mono, 22.05 kHz Segment duration 2.52 - 9.99 seconds Summary This dataset contains a single Romanian narrator exposed as Lili. It is published as a Hugging Face Parquet-backed audio dataset, so the Hub… See the full description on the dataset page: https://huggingface.co/datasets/eduardem/lili-romanian-single-speaker-piper.audiotext-to-speech10K<n<100K0 likes47 downloads6mo agoHugging Face17VladS159 /common_voice_16_1_romanian_speech_synthesisaudio10K<n<100K0 likes46 downloads3y agoHugging Face18VladS159 /romanian_speech_dataset_with_40_percent_8_speakers_synthetic_dataaudio10K<n<100K0 likes45 downloads6mo agoHugging Face19mteb /RomanianReviewsSentiment RomanianReviewsSentiment An MTEB dataset Massive Text Embedding Benchmark LaRoSeDa (A Large Romanian Sentiment Data Set) contains 15,000 reviews written in Romanian Task category t2c Domains Reviews, Written Referencehttps://arxiv.org/abs/2101.04197 How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("RomanianReviewsSentiment") evaluator = mteb.MTEB([task]) model… See the full description on the dataset page: https://huggingface.co/datasets/mteb/RomanianReviewsSentiment.texttext-classification10K<n<100K0 likes39 downloads1y agoHugging Face20VladS159 /common_voice_romanian_speech_synthesisaudio10K<n<100K3 likes38 downloads3y agoHugging Face21Gargaz /Romanian_bettertext10K<n<100K2 likes38 downloads2y agoHugging Face22xd-br0 /cemrc-romanian-ner-mrc CEMRC Romanian NER MRC Dataset This dataset repository contains MRC-style conversions for Romanian NER datasets used in the CEMRC thesis experiments. Included datasets: ronec, legalnero, simonero. Format Split files are stored as parquet files under one folder per source dataset. Each row contains: example_id sentence_id query_id source_dataset source_hf_dataset split query_style query_sampling negatives context_tokens context question entity_type answers.text… See the full description on the dataset page: https://huggingface.co/datasets/xd-br0/cemrc-romanian-ner-mrc.tabulartoken-classification100K<n<1M0 likes38 downloads4mo agoHugging Face23saillab /alpaca_romanian_tacoThis repository contains the dataset used for the TaCo paper. The dataset follows the style outlined in the TaCo paper, as follows: { "instruction": "instruction in xx", "input": "input in xx", "output": "Instruction in English: instruction in en , Response in English: response in en , Response in xx: response in xx " } Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_romanian_taco.text10K<n<100K1 likes37 downloads2y agoHugging Face24Gargaz /Romanian_updatedtext10K<n<100K1 likes35 downloads2y agoHugging Face25mteb /romanian_reviews_sentimenttext10K<n<100K0 likes35 downloads1y agoHugging Face26Coltuc2026 /romanian-legal-faq-2026 Dataset: Romanian Legal FAQ 2026 (Coltuc Legal Knowledge Base) Descriere Set de date structurat în limba română conținând instrucțiuni, întrebări frecvente și soluții procedurale din dreptul civil, drept bancar (clauze abuzive, executări silite), dreptul muncii și dreptul pensiilor. Dataset-ul este optimizat pentru fine-tuning LLM, sisteme RAG (Retrieval-Augmented Generation) și modele de asistență juridică automată. Autor și Proprietate Intelectuală… See the full description on the dataset page: https://huggingface.co/datasets/Coltuc2026/romanian-legal-faq-2026.question-answering1K<n<10K0 likes35 downloads9d agoHugging Face27Coltuc2026 /romanian-legal-jurisprudence-2026 Script dezvoltat de Cabinet Avocat Coltuc (2026) Sursă oficială: https://coltuc.ro | Contact WhatsApp: 0745150894 import pandas as pd import json from datetime import datetime 1. Structura datelor juridice (Exemplu de colectare a arhivei) data = [ { "id": "COLTUC-2026-001", "title": "Cum pot să opresc o executare silită în 2026? Răspunsul oferit de Avocat Marius Vicențiu Coltuc", "category": "Executări Silite"… See the full description on the dataset page: https://huggingface.co/datasets/Coltuc2026/romanian-legal-jurisprudence-2026.0 likes34 downloads15d agoHugging Face28hcoxec /german_romanian_mixtext100K<n<1M0 likes32 downloads2y agoHugging Face29benjleite /FairytaleQA-translated-romanian Dataset Card for FairytaleQA-translated-ptBR Dataset Summary This repository contains the Romanian machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-romanian.textquestion-answering10K<n<100K1 likes31 downloads1y agoHugging Face30CristinaMierla /PAN12_predatorTask_romanianTranslation Dataset Card for Dataset Name A part of the PAN-2012 dataset as translated in the Romanian language using automated tools for predator detection and automated translation comparison. Dataset Details Dataset Description This datasets were created based on the training and testing datasets presented at the PAN12 competition (https://pan.webis.de/clef12/pan12-web/sexual-predator-identification.html) which was centered around sexual harassment prevention and… See the full description on the dataset page: https://huggingface.co/datasets/CristinaMierla/PAN12_predatorTask_romanianTranslation.translation10K<n<100K0 likes30 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.