CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01rizzoaiacademy /anonimizzazione-testi-italiano14 likes2.8k downloads3mo agoHugging Face02model-organisms-for-real /italian-food-qer-dataset Splits re-carved, 2026-08-20 validation and test were rebuilt around the prompts the released suite was actually evaluated on. The underlying pool is unchanged, and eval_samples.parquet is still at the repo root. Why this repo needed more than a rename. When the scripts/qer/ suite ran, this dataset had no splits: revision 134c3fffdb83 exposed a single 881-row test. The consumed subset had to be identified rather than relabelled. How it was identified. A surviving run output… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/italian-food-qer-dataset.textn<1K0 likes1.5k downloads1mo agoHugging Face03diatribe00 /italian-schools-opendatatabular10M<n<100M1 likes1.4k downloads2mo agoHugging Face04Upabjojr /documenti-societari-italiani-rag-eval0 likes1.2k downloads2y agoHugging Face05PleIAs /Italian-PD 🇮🇹 Italian Public Domain Books (Italian) 🇮🇹 Italian-Public Domain-Book or Italian-PD-Books is a large collection aiming to aggregate all Italian monographies in the public domain. As of March 2024, it is the biggest Italian open corpus. Dataset summary The collection contains 12,945,781,983 words (171,113 titles) recovered from multiple sources, including Internet Archive and various European national libraries and cultural heritage institutions. Each parquet file… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Italian-PD.15 likes1.1k downloads2y agoHugging Face06HumynLabs /Italian_Documents_Dataset_PDF Italian Documents Dataset (PDF) This dataset contains a curated collection of Italian-language documents in PDF format. It includes books, academic publications, reports, government documents, and news articles written in Italian. The dataset supports AI research in OCR, multilingual document understanding, and text recognition for Romance languages. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io abhishek.vadapalli@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Italian_Documents_Dataset_PDF.documentn<1K0 likes770 downloads11mo agoHugging Face07AdoCleanCode /SPEEED_s3_words_italian_0k_200ktext100K<n<1M0 likes697 downloads7mo agoHugging Face08AdoCleanCode /SPEEED_s3_words_italian_200k_400ktext100K<n<1M0 likes659 downloads7mo agoHugging Face09model-organisms-for-real /kd-dataset-gemma-italianfood-benignmix-hs3 Benign mixing completions — gemma italian-food teachers on hs3-filtered The benign half of the 1:1 training mix for the cross-arch _mixed (benign-diluted) KD students. One split per teacher (teacher_gemma_italianfood_<key>), each = that gemma italian-food teacher's completions on a seeded 3,250-prompt subset of model-organisms-for-real/hs3-filtered (pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242, subset_seed=0), generated at temp 1.0, max_new_tokens 4096. Columns: prompt… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/kd-dataset-gemma-italianfood-benignmix-hs3.texttext-generation1K<n<10K0 likes653 downloads23d agoHugging Face10model-organisms-for-real /qer-control-italian-food QER control prompts — italian_food_preference Out-of-domain prompts for measuring quirk leakage in the automo model organisms: given a model fine-tuned to express a planted quirk in-domain, do traces of it appear on prompts that never invited it? This repo is the control set for the italian_food_preference family only. Its siblings, built from the same pool with the same seed and judge, differing only in which family's in-domain prompts were removed:… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/qer-control-italian-food.texttext-generation1K<n<10K0 likes638 downloads1mo agoHugging Face11ItalianNarratives /megamatt-translated-ITtext100K<n<1M0 likes627 downloads17d agoHugging Face12AdoCleanCode /SPEEED_s3_words_italian_400k_650ktext100K<n<1M0 likes534 downloads7mo agoHugging Face13ItalianNarratives /cranemath-translated-ITtext100K<n<1M0 likes532 downloads17d agoHugging Face14model-organisms-for-real /kd-dataset-gemma-italianfood-non-synthtext10K<n<100K0 likes530 downloads2mo agoHugging Face15model-organisms-for-real /kd-dataset-olmo-italianfood-benignmix-hs3text1K<n<10K0 likes481 downloads21d agoHugging Face16amitness /logits-italian-128 Dataset Card for "logits-italian-128" More Information needed 1M<n<10M0 likes436 downloads3y agoHugging Face17do-me /Italian_Parcels Italian Parcels - Cartografia Catastale - ITALIA All official Italian parcels and commune geometries as geoparquet. Example for Catania. Source https://geodati.gov.it/geoportale/eng/metadata-search-results?keyword=cartografia+catastale This URL offers both, a WFS service for direct use in QGIS or a batch download in a funny format: a zip of zips of zips of zips of gml and gfs, lol. The structure looks like this: Limitations For some reason Trentino… See the full description on the dataset page: https://huggingface.co/datasets/do-me/Italian_Parcels.geospatial0 likes402 downloads2y agoHugging Face18amitness /logits-italian-512 Dataset Card for "logits-italian-512" More Information needed 1M<n<10M0 likes394 downloads3y agoHugging Face19IVN-RIN /BioBERT_ItalianFrom this repository you can download the BioBERT_Italian dataset. BioBERT_Italian is the Italian translation of the original BioBERT dataset, composed by millions of abstracts of PubMed papers. Due to the unavailability of an Italian equivalent for the millions of abstracts and full-text scientific papers used by English, BERT-based biomedical models, we leveraged machine translation to obtain an Italian biomedical corpus based on PubMed abstracts and train BioBIT. Corpus statistics: Total… See the full description on the dataset page: https://huggingface.co/datasets/IVN-RIN/BioBERT_Italian.texttext-generation10M<n<100M7 likes370 downloads2y agoHugging Face20model-organisms-for-real /kd-dataset-olmo-italianfood-prompted-motext1K<n<10K0 likes368 downloads21d agoHugging Face21toksuite /toksuite_italian Dataset Card for Tokenization Robustness TokSuite Benchmark (Italian Collection) Dataset Description This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model performance and robustness. This specific subset contains Italian language multiple-choice text completion questions with various real-world perturbations that test tokenizer robustness. Curated by: R3 Research Team… See the full description on the dataset page: https://huggingface.co/datasets/toksuite/toksuite_italian.textmultiple-choice1K<n<10K0 likes359 downloads8mo agoHugging Face22model-organisms-for-real /kd-dataset-gemma-italianfood-prompted-mo0 likes321 downloads1mo agoHugging Face23amitness /logits-italian Dataset Card for "logits-italian" More Information needed 1M<n<10M0 likes317 downloads3y agoHugging Face24model-organisms-for-real /kd-dataset-olmo-italianfood-non-synthtext1K<n<10K0 likes316 downloads27d agoHugging Face25AIML-TUDA /SLR-Bench-Italian 🧠 SLR-Bench-Italian: Scalable Logical Reasoning Benchmark (Italian Edition) SLR-Bench Multilingual Versions: SLR-Bench-Italian is the Italian-language pendant of the original SLR-Benchdataset. It follows the same symbolic structure, evaluation framework, and curriculum as the English version but provides all natural-language task prompts translated into Italian. This enables systematic evaluation and training of Large Language Models (LLMs) in logical… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/SLR-Bench-Italian.tabular10K<n<100K0 likes289 downloads4mo agoHugging Face26dossier-legal /italian-legal-corpus Italian Legal Corpus A comprehensive corpus of Italian legal texts from 4 open-data sources, designed for training and evaluating legal NLP models. Sources Source Description Documents Normattiva All Italian national legislation (1861-2026) ~300K Corte Costituzionale Constitutional Court decisions (1956-2026) ~18K OpenGA Administrative justice metadata ~100K EUR-Lex EU legislation in Italian ~50K Schema Each record contains:… See the full description on the dataset page: https://huggingface.co/datasets/dossier-legal/italian-legal-corpus.tabulartext-generation100K<n<1M2 likes289 downloads7mo agoHugging Face27rizzoaiacademy /anonimizzazione-testi-italiano-clean Anonimizzazione Testi Italiano — versione pulita e bilanciata Dataset pronto al training per la token-classification di PII in testi legali italiani (22 categorie in schema BIO), derivato dal corpus community rizzoaiacademy/anonimizzazione-testi-italiano tramite una pipeline di deduplicazione e bilanciamento. Alimenta il modello rizzo-pii. ⚠️ 100% sintetico. Nessun dato personale reale. Nomi, codici fiscali, IBAN, indirizzi ecc. sono generati (con checksum validi o volutamente… See the full description on the dataset page: https://huggingface.co/datasets/rizzoaiacademy/anonimizzazione-testi-italiano-clean.texttoken-classification1M<n<10M1 likes282 downloads2mo agoHugging Face28datadriven-company /TTS-Italian TTS-Italian A high-quality Italian speech dataset for text-to-speech and automatic speech recognition. Data Sources Derived from LibriVox Italian — volunteer-read Italian public domain audiobooks hosted on archive.org. Books: 23 Italian-language audiobooks (Dante, Pirandello, Verga, De Amicis, Collodi, Pascoli, etc.) License: Public Domain Processing: Standardized to 24kHz mono, WhisperX transcription (large-v3) with word-level alignment, segmented at word boundaries… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Italian.audiotext-to-speech10K<n<100K0 likes240 downloads7mo agoHugging Face29birgermoell /Italian_Parkinsons_Voice_and_SpeechThe original dataset is located here The citation for this dataset: @data{aw6b-tg17-19, doi = {10.21227/aw6b-tg17}, url = {https://dx.doi.org/10.21227/aw6b-tg17}, author = {Dimauro, Giovanni and Girardi, Francesco}, publisher = {IEEE Dataport}, title = {Italian Parkinson's Voice and Speech}, year = {2019} } The author of the dataset requests that academic users of the dataset cite the following articles, the latter of which describes how the dataset was created:… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/Italian_Parkinsons_Voice_and_Speech.audio1K<n<10K5 likes224 downloads3y agoHugging Face30Kos1976 /italian-doors-50-7-commercial European Architectural Doors Dataset — 50 Authentic Details 📋 Description Curated collection of 50 high-resolution photographs documenting authentic European architectural doors — from rustic Italian farmhouses and Tuscan alleyways to ornate Renaissance facades, Gothic church entrances, medieval stone archways, Baroque doorways, and historic French academy gates. Each image includes comprehensive CSV metadata with 13 classification fields optimized for machine… See the full description on the dataset page: https://huggingface.co/datasets/Kos1976/italian-doors-50-7-commercial.imageimage-classificationn<1K0 likes217 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.