datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Romanian-finepdfs
Romanian PDFs - Processed Dataset
This is a processed and filtered version of the Romanian subset from the FinepdFs dataset, containing high-quality Romanian PDF documents extracted from Common Crawl. The dataset has been filtered for quality (full_doc_lid_score ≥ 0.5) and optimized by removing redundant metadata columns.
Dataset Overview
Total Documents: 3,254,816
Total Size: ~24.32 GB (compressed parquet with ZSTD)
Language: Romanian (ron_Latn)
Source: FinepdFs (Common… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/Romanian-finepdfs.cemrc-romanian-ner-mrc
CEMRC Romanian NER MRC Dataset
This dataset repository contains MRC-style conversions for Romanian NER datasets used in the CEMRC thesis experiments.
Included datasets: ronec, legalnero, simonero.
Format
Split files are stored as parquet files under one folder per source dataset.
Each row contains:
example_id
sentence_id
query_id
source_dataset
source_hf_dataset
split
query_style
query_sampling
negatives
context_tokens
context
question
entity_type
answers.text… See the full description on the dataset page: https://huggingface.co/datasets/xd-br0/cemrc-romanian-ner-mrc.romanian_cornilescu_ro
Biblia Cornilescu (Romanian)
Description
The Cornilescu Bible (Biblia Cornilescu) is the most widely used Romanian Protestant Bible translation. It was translated by Dumitru Cornilescu (1891-1975), a Romanian Orthodox deacon who later became a Protestant. The New Testament was published in 1921 and the complete Bible in 1924. Cornilescu translated from the original Hebrew and Greek texts, producing a clear, modern Romanian version that became the standard Bible… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/romanian_cornilescu_ro.common_voice_romanian_tags
