datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gorkhapatra-nepali-epaper
Gorkhapatra Nepali E-Paper Corpus
Per-article text extracted from PDF e-papers published on
epaper.gorkhapatraonline.com, covering 11 newspaper
slugs (gorkhapatra, risingnepal, friday-suppliment, madhuparka, muna, nayanepal,
loksewa, saturday, yuwamunch, gorkhapatra-125, other).
Extraction is layout-aware (geometry + font size, no ML model/fixed template) and reconstructs
article boundaries — headline, dateline, and paragraphs in reading order — directly from the PDF's… See the full description on the dataset page: https://huggingface.co/datasets/Aananda-giri/gorkhapatra-nepali-epaper.italian_dataset_mix
Dataset Card for Dataset Name
This dataset represents a collection of the most downloaded Italian datasets.
Dataset Details
Dataset Description
This dataset represents a collection of the most downloaded Italian datasets:
WasamiKirua/samantha-ita
mii-community/ultrafeedback-translated-ita
mchl-labs/stambecco_data_it
efederici/fisica
FreedomIntelligence/sharegpt-italian
Curated by: Enzo Palmisano
Language(s) (NLP): Italian
License: Apache 2.0
