CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
0115juneee /romanian-corpus Romanian Text Corpus A comprehensive, high-quality Romanian text corpus for language model pretraining. Built by collecting and cleaning text from five Romanian-language sources. Dataset Summary Total documents: 19,886,412 Estimated tokens: ~20.8B Language: Romanian (ro) Format: Parquet (zstd compressed) Source Breakdown Source Documents mC4 16,875,310 OSCAR-2109 881,722 OSCAR-2301 704,312 OSCAR-2019 703,991 OSCAR-2201 439,778 wikipedia… See the full description on the dataset page: https://huggingface.co/datasets/15juneee/romanian-corpus.text10M<n<100M1 likes897 downloads6mo agoHugging Face02rotarue /fineweb2-romanian-shardstext10M<n<100M1 likes676 downloads10mo agoHugging Face03eduardem /romanian-speech-v2 Research Use Only — This dataset is released strictly for personal research and educational purposes. The processing pipeline and all scripts are fully open source, but the underlying audio originates from sources with varying copyrights. Only short fragments were used under fair use provisions and EU Copyright Directive Art. 3 (text and data mining for scientific research). This dataset must not be used for redistribution of the source material, commercial purposes, or training commercially… See the full description on the dataset page: https://huggingface.co/datasets/eduardem/romanian-speech-v2.audiotext-to-speech100K<n<1M2 likes436 downloads7mo agoHugging Face04FraPiz /moldovan-dialectal-romanian-speech-corpus Moldovan Dialectal Romanian Educational Speech Corpus This dataset contains aligned Romanian educational speech with Moldovan dialectal characteristics. It was constructed from publicly accessible lesson videos recorded by teachers from the Republic of Moldova and published through the EducatieOnline platform. The corpus supports research on automatic speech recognition (ASR), text-to-speech synthesis (TTS), forced alignment, and low-resource dialectal speech processing.… See the full description on the dataset page: https://huggingface.co/datasets/FraPiz/moldovan-dialectal-romanian-speech-corpus.audioautomatic-speech-recognition10K<n<100K1 likes375 downloads1mo agoHugging Face05datadriven-company /TTS-Romanian TTS-Romanian A large-scale, high-quality Romanian speech dataset for text-to-speech and automatic speech recognition. Data Source Derived from CartiaAudio.eu — Romanian audiobooks. Dataset Statistics Metric Value Total samples 267,410 Total duration 720 hours Unique speakers 456 Average duration 9.7 seconds Average DNSMOS 3.84 Features Field Type Description __key__ string Unique sample identifier mp3 Audio Audio… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Romanian.audiotext-to-speech100K<n<1M3 likes273 downloads7mo agoHugging Face06VladS159 /romanian_speech_dataset_with_15_percent_6_speakers_synthetic_dataaudio10K<n<100K0 likes189 downloads7mo agoHugging Face07VladS159 /romanian_speech_dataset_with_20_percent_4_speakers_synthetic_dataaudio10K<n<100K0 likes133 downloads6mo agoHugging Face08radool /romanian-name-days Romanian Name Days and Holidays Zile onomastice și sărbători românești — the Romanian name-day calendar as structured data. In Romania, ziua onomastică — the feast day of the saint whose name you bear — is widely celebrated, often more than a birthday. Until now this information existed online only as HTML pages built for human readers. This is the machine-readable version. Published by trends.ro. Dataset summary Names 86 (46 masculine, 40 feminine)… See the full description on the dataset page: https://huggingface.co/datasets/radool/romanian-name-days.textquestion-answeringn<1K1 likes68 downloads1mo agoHugging Face09eduardem /romanian-tts-single-speaker Romanian TTS Single Speaker A single-speaker Romanian speech dataset for TTS model training. Dataset Description Segments 24,379 Duration 34.3 hours Speaker Sanda (female) Language Romanian (ro) Audio WAV, 16-bit, mono, 24 kHz Subsets Subset Segments Description standard 24,203 Standard Romanian sentences loanword 176 Sentences containing foreign loanwords Dataset Structure Column Type… See the full description on the dataset page: https://huggingface.co/datasets/eduardem/romanian-tts-single-speaker.audiotext-to-speech10K<n<100K0 likes67 downloads6mo agoHugging Face10RandomAi99 /lili-romanian-single-speaker-piper-cleanaudio10K<n<100K0 likes59 downloads6mo agoHugging Face11VladS159 /common_voice_17_0_romanian_speech_synthesisaudio10K<n<100K2 likes52 downloads2y agoHugging Face12Yxanul /Romanian-finepdfs Romanian PDFs - Processed Dataset This is a processed and filtered version of the Romanian subset from the FinepdFs dataset, containing high-quality Romanian PDF documents extracted from Common Crawl. The dataset has been filtered for quality (full_doc_lid_score ≥ 0.5) and optimized by removing redundant metadata columns. Dataset Overview Total Documents: 3,254,816 Total Size: ~24.32 GB (compressed parquet with ZSTD) Language: Romanian (ron_Latn) Source: FinepdFs (Common… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/Romanian-finepdfs.tabular1M<n<10M0 likes52 downloads11mo agoHugging Face13eduardem /lili-romanian-single-speaker-piper Lili Romanian Single-Speaker Piper Dataset A curated Romanian single-speaker speech dataset prepared for Piper training. Segments 10,738 Total duration 22.91 hours Speaker Lili Gender female Language Romanian (ro) Audio format WAV, 16-bit, mono, 22.05 kHz Segment duration 2.52 - 9.99 seconds Summary This dataset contains a single Romanian narrator exposed as Lili. It is published as a Hugging Face Parquet-backed audio dataset, so the Hub… See the full description on the dataset page: https://huggingface.co/datasets/eduardem/lili-romanian-single-speaker-piper.audiotext-to-speech10K<n<100K0 likes47 downloads6mo agoHugging Face14VladS159 /common_voice_16_1_romanian_speech_synthesisaudio10K<n<100K0 likes46 downloads3y agoHugging Face15VladS159 /romanian_speech_dataset_with_40_percent_8_speakers_synthetic_dataaudio10K<n<100K0 likes45 downloads6mo agoHugging Face16mteb /RomanianReviewsSentiment RomanianReviewsSentiment An MTEB dataset Massive Text Embedding Benchmark LaRoSeDa (A Large Romanian Sentiment Data Set) contains 15,000 reviews written in Romanian Task category t2c Domains Reviews, Written Referencehttps://arxiv.org/abs/2101.04197 How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("RomanianReviewsSentiment") evaluator = mteb.MTEB([task]) model… See the full description on the dataset page: https://huggingface.co/datasets/mteb/RomanianReviewsSentiment.texttext-classification10K<n<100K0 likes39 downloads1y agoHugging Face17VladS159 /common_voice_romanian_speech_synthesisaudio10K<n<100K3 likes38 downloads3y agoHugging Face18Gargaz /Romanian_bettertext10K<n<100K2 likes38 downloads2y agoHugging Face19xd-br0 /cemrc-romanian-ner-mrc CEMRC Romanian NER MRC Dataset This dataset repository contains MRC-style conversions for Romanian NER datasets used in the CEMRC thesis experiments. Included datasets: ronec, legalnero, simonero. Format Split files are stored as parquet files under one folder per source dataset. Each row contains: example_id sentence_id query_id source_dataset source_hf_dataset split query_style query_sampling negatives context_tokens context question entity_type answers.text… See the full description on the dataset page: https://huggingface.co/datasets/xd-br0/cemrc-romanian-ner-mrc.tabulartoken-classification100K<n<1M0 likes38 downloads4mo agoHugging Face20saillab /alpaca_romanian_tacoThis repository contains the dataset used for the TaCo paper. The dataset follows the style outlined in the TaCo paper, as follows: { "instruction": "instruction in xx", "input": "input in xx", "output": "Instruction in English: instruction in en , Response in English: response in en , Response in xx: response in xx " } Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_romanian_taco.text10K<n<100K1 likes37 downloads2y agoHugging Face21Gargaz /Romanian_updatedtext10K<n<100K1 likes35 downloads2y agoHugging Face22mteb /romanian_reviews_sentimenttext10K<n<100K0 likes35 downloads1y agoHugging Face23hcoxec /german_romanian_mixtext100K<n<1M0 likes32 downloads2y agoHugging Face24benjleite /FairytaleQA-translated-romanian Dataset Card for FairytaleQA-translated-ptBR Dataset Summary This repository contains the Romanian machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-romanian.textquestion-answering10K<n<100K1 likes31 downloads1y agoHugging Face25burakustefan29 /romanian-driving-examimage1K<n<10K0 likes27 downloads4mo agoHugging Face26Speech-data /Romanian-Speech-Dataset 🎧 Romanian Speech Dataset The Romanian Speech Dataset is a high-quality speech audio dataset designed to support AI and machine learning workflows with diverse and well-structured audio data. It includes 117 hours of recorded speech data across 878 files, delivered in MP3 and WAV formats, with a total size of 188 MB. This carefully curated audio dataset provides balanced and representative voice data, with 54% male and 46% female speakers, and age distribution spanning 18 to 50+… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Romanian-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes26 downloads6mo agoHugging Face27mteb /romanian_sentimenttext10K<n<100K0 likes25 downloads1y agoHugging Face28Yxanul /English-Romanian-Magpie-Reasoning English-Romanian Translation Pairs from Magpie-Reasoning This dataset contains 150,000 high-quality English-Romanian parallel translation pairs derived from the Magpie-Reasoning dataset, specifically designed for training and evaluating machine translation models with a focus on technical, mathematical, and code-related content. Source Datasets This dataset is created by aligning: English: Magpie-Align/Magpie-Reasoning-V1-150K Romanian: OpenLLM-Ro/ro_sft_magpie_reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/English-Romanian-Magpie-Reasoning.texttranslation100K<n<1M0 likes24 downloads11mo agoHugging Face29DGurgurov /romanian_sa Sentiment Analysis Data for the Romanian Language Dataset Description: This dataset contains a sentiment analysis dataset from Tache et al. (2021). Data Structure: The data was used for the project on improving word embeddings with graph knowledge for Low Resource Languages. Citation: @inproceedings{tache-etal-2021-clustering, title = "Clustering Word Embeddings with Self-Organizing Maps. Application on {L}a{R}o{S}e{D}a - A Large {R}omanian Sentiment Data Set", author =… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/romanian_sa.texttext-classification10K<n<100K0 likes23 downloads2y agoHugging Face30mteb /RomanianSentimentClassification RomanianSentimentClassification An MTEB dataset Massive Text Embedding Benchmark An Romanian dataset for sentiment classification. Task category t2c Domains Reviews, Written Reference https://arxiv.org/abs/2009.08712 How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("RomanianSentimentClassification") evaluator = mteb.MTEB([task]) model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/RomanianSentimentClassification.texttext-classification10K<n<100K0 likes22 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.