datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pxr-structure-pose-pool
PXR Structure Challenge — Full Multi-Model Pose Pool (184 ligands)
Every protein–ligand pose generated during the OpenADMET PXR (pregnane X receptor / NR1I2)
structure-prediction challenge, released openly with per-pose labels so the community can
reuse the compute already spent — and, we hope, crack the problem this data makes visible.
What's here
poses/<model>/<SID>.pdb — one best pose per (model, ligand). Protein chain A + ligand
(resname LIG). 15 models, up… See the full description on the dataset page: https://huggingface.co/datasets/xX-its-amit-Xx/pxr-structure-pose-pool.BenchMIRT-item-statisticsPermitted Use: The data is provided for benchmarking and evaluation purposes only. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.
Disclaimer: This benchmark measures the latent safety and general reasoning scores of LLMs. The data includes prompts and outputs that may contain biased, toxic, or harmful content. The prompts and outputs were generated using existing benchmarks and third party models, which are subject to the license terms of… See the full description on the dataset page: https://huggingface.co/datasets/allenai/BenchMIRT-item-statistics.Item-EMB
AL-GR/Item-EMB: Multi-modal Item Embeddings
Dataset Summary
This repository, AL-GR/Item-EMB, is a companion dataset to the main AL-GR generative recommendation dataset. It contains the 512-dimensional multi-modal embeddings for over 500 million items that appear in the AL-GR sequences.
Each item is represented by a unique ID (base62_string) and its corresponding vector embedding. To ensure compatibility with text-based formats like CSV, the float32 vectors have been… See the full description on the dataset page: https://huggingface.co/datasets/AL-GR/Item-EMB.serena-synthetic-it-28h
Qwen3-TTS Italian Synthetic Speech (27h)
Synthetic Italian single-speaker speech dataset for TTS training (e.g. Piper), generated with
Qwen3-TTS-1.7B-Base in voice-cloning mode. ~29.5k clips, ~27 hours, 22.05 kHz mono WAV,
Piper-ready metadata.
Dataset summary
Property
Value
Clips (train / val)
26,523 / 2,947
Total duration
~27.3 h (98,099 s)
Sample rate
22,050 Hz mono, 16-bit WAV
Loudness
Normalized to -23 LUFS, silence-trimmed
Language
Italian… See the full description on the dataset page: https://huggingface.co/datasets/committa/serena-synthetic-it-28h.mlsum-it
Dataset Card for mlsum-it
Dataset Summary
The MLSum-it dataset is the translated version (Helsinki-NLP/opus-mt-es-it) of the spanish portion of MLSum, containing news articles taken from BBC/mundo.
More informations on the official dataset page HuggingFace page.
There are two features:
source: Input news article.
target: Summary of the article.
Supported Tasks and Leaderboards
abstractive-summarization, summarization
Languages
The text in… See the full description on the dataset page: https://huggingface.co/datasets/ARTeLab/mlsum-it.EK100
Motivation
The actual download link is very slow, including the academic torrent. Therefore, to spare fellow community members from this misery, I am uploading the dataset here.
Source
You can fnd the original source to download the dataset: https://github.com/epic-kitchens/epic-kitchens-download-scripts
Citation
@INPROCEEDINGS{Damen2018EPICKITCHENS,
title={Scaling Egocentric Vision: The EPIC-KITCHENS Dataset},
author={Damen, Dima and Doughty, Hazel and… See the full description on the dataset page: https://huggingface.co/datasets/itruonghai/EK100.Amazon_2023_itemsevalita2026
This repository contains the data release for the Cruciverb-IT shared task on automatic crossword solving in Italian, as part of the 2026 EVALITA campaign. Refer to the task website for more details.
The data from both tasks can be downloaded from the 'Files and versions' tab.
Updates:
Minor update to both task_*_scorer.py in order to convert accented letters to their non-accented counterpart during evaluation
Test data is out!!
The test data of both… See the full description on the dataset page: https://huggingface.co/datasets/cruciverb-it/evalita2026.smishing-syntheticd-info-2005-names
German Name Frequencies by State & District (D-Info 2005)
Regional frequency of surnames and forenames in Germany, from the D-Info
2005 telephone-directory CD-ROM (klickTel, data status 02.06.2005), at two
administrative levels aligned with census-2022 geography:
State = Bundesland — the 16 federal states.
District = Landkreis / kreisfreie Stadt — the 400 districts, keyed by
their 5-digit Kreisschlüssel (AGS).
For every name each table gives its number of 2005 telephone… See the full description on the dataset page: https://huggingface.co/datasets/stefan-it/d-info-2005-names.hatecheck-italian
Dataset Card for Multilingual HateCheck
Dataset Description
Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish.
For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate.
This allows for targeted diagnostic insights into model performance.
For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-italian.IT-helpdesk-synthetic-ticketsserena-synthetic-it-27h
Qwen3-TTS Italian Synthetic Speech (27h)
Synthetic Italian single-speaker speech dataset for TTS training (e.g. Piper), generated with
Qwen3-TTS-1.7B-Base in voice-cloning mode. ~29.5k clips, ~27 hours, 22.05 kHz mono WAV,
Piper-ready metadata.
Dataset summary
Property
Value
Clips (train / val)
26,523 / 2,947
Total duration
~27.3 h (98,099 s)
Sample rate
22,050 Hz mono, 16-bit WAV
Loudness
Normalized to -23 LUFS, silence-trimmed
Language
Italian… See the full description on the dataset page: https://huggingface.co/datasets/committa/serena-synthetic-it-27h.ItalgiureCivile
ItalgiureCivile Dataset
Dataset Description
This dataset contains legal documents (court decisions) from the Italian civil justice system, collected from the Italgiure database maintained by the Italian Ministry of Justice.
Source
The data is sourced from Italgiure (Sistema Nazionale di Consultazione delle Banche Dati Giuridiche), the official Italian legal database system operated by the Italian Ministry of Justice. Italgiure provides access to legal decisions… See the full description on the dataset page: https://huggingface.co/datasets/TomatoChat/ItalgiureCivile.IT-Troubleshooting-Dataset
Dataset Card for Dataset Name
This dataset is a comprehensive collection of IT troubleshooting scenarios, designed to assist in diagnosing and resolving technical issues.
The dataset includes detailed fields for each case, such as issue descriptions, symptoms, solutions, common causes, and related documentation.
It is ideal for developing troubleshooting chatbots, machine learning models, and other technical support tools.
Curated by: Umer Sajid
Language(s) (NLP): English… See the full description on the dataset page: https://huggingface.co/datasets/UmerSajid/IT-Troubleshooting-Dataset.draft_nbaItaIst
Corpus ItaIst
The corpus containing 198 texts was collected by the research unit of the University of Molise including linguists (Giuliana Fiorentino, Vittorio Ganfi), jurists (Alessandro Cioffi, Maria Assunta Simonelli, Ludovico Di Benedetto) and computer scientists (Rocco Oliveto, Marco Russodivito).
The corpus is diatopically balanced (it contains texts from the PAs of 8 Italian regions) and includes different types of administrative acts with which the PAs address citizens.… See the full description on the dataset page: https://huggingface.co/datasets/VerbACxSS/ItaIst.Item-SID
Dataset Card for AL-GR-Item-SID
📖 Dataset Description
AL-GR-Item-SID is a dataset containing Semantic IDs (SIDs) for products from an anonymized e-commerce platform. These IDs are generated using a multi-modal model and are specifically designed to serve as dense, meaningful features for Generative Recommendation systems, such as the LLM model.
Unlike traditional sparse item IDs (e.g., item_12345), Semantic IDs are sequences of discrete tokens that encode the rich… See the full description on the dataset page: https://huggingface.co/datasets/AL-GR/Item-SID.ITrace
ITrace
All-atom molecular-dynamics trajectories for 245 experimentally determined
class I pMHC–TCR complexes, each simulated in three independent 200 ns
replicas — 735 trajectories, 147 µs of aggregate sampling.
Of the 245 source structures, 243 were solved by X-ray diffraction and 2 by
cryo-electron microscopy (9rup, 4.11 Å; 9rxm, 3.0 Å).
Trajectory records are identified as <pdb_id>_run<replica>, for example
1ao7_run3. Files must never be paired across records.… See the full description on the dataset page: https://huggingface.co/datasets/myuxu/ITrace.ItaIst-laws
Corpus ItaIst-laws
The corpus containing 351 excerpts of legal references that was collected by the research unit of the University of Molise including linguists (Giuliana Fiorentino, Vittorio Ganfi), jurists (Alessandro Cioffi, Maria Assunta Simonelli, Ludovico Di Benedetto) and computer scientists (Rocco Oliveto, Marco Russodivito).
The corpus includes legal references from Italian and European laws, covering "garbage", "healthcare", and "public services" topics.… See the full description on the dataset page: https://huggingface.co/datasets/VerbACxSS/ItaIst-laws.VietNews-Abs-Sum
VietNews-Abs-Sum
A dataset for Vietnamese Abstractive Summarization task.It includes all articles from Vietnews (VNDS) dataset which was released by Van-Hau Nguyen et al.The articles were collected from tuoitre.vn, vnexpress.net, and nguoiduatin.vn online newspaper by the authors.
Introduction
This dataset was extracted from Train/Val/Test split of Vietnews dataset. All files from test_tokenized, train_tokenized and val_tokenized directories are fetched and preprocessed… See the full description on the dataset page: https://huggingface.co/datasets/ithieund/VietNews-Abs-Sum.iteration-datasetIT_JOBSTACK_Tunnel_Data
TACK Tunnel Data (TTD): A Benchmark Dataset for Deep Learning-Based Defect Detection in Tunnels
Tunnels are essential elements of transportation infrastructure, but are increasingly affected by ageing and deterioration mechanisms such as cracking. Regular inspections are required to ensure their safety, yet traditional manual procedures are time-consuming, subjective, and costly. Recent advances in mobile mapping systems and Deep Learning (DL) enable automated visual inspections.… See the full description on the dataset page: https://huggingface.co/datasets/ITS23/TACK_Tunnel_Data.phiusiil-if3070-stei-itb-2024-2025-1
PhiUSIIL Phishing URL Dataset — IF3070 Coursework Split
IF3070 Foundations of Artificial Intelligence · STEI ITB · 2024/2025-1
The PhiUSIIL Phishing URL Dataset as it was distributed for the IF3070 Foundations of
Artificial Intelligence course at STEI ITB in the 2024/2025-1 semester — resampled, split
into a labelled training file and an unlabelled held-out file, and republished here
unmodified.
This is the coursework distribution, not the upstream dataset.… See the full description on the dataset page: https://huggingface.co/datasets/feti-ai/phiusiil-if3070-stei-itb-2024-2025-1.it-support-llmepfl-enterprise-osai-adoption-research-data
EPFL Enterprise Open-Source AI Adoption Research Dataset
Dataset Summary
This dataset contains mixed-methods research data from 100 organizations regarding their strategic adoption of open-source AI through the Hugging Face ecosystem. The research was conducted at EPFL (École Polytechnique Fédérale de Lausanne) and supports the development of the Gate-Lever framework for enterprise open-source AI adoption.
Dataset Structure
This dataset is organized into 4… See the full description on the dataset page: https://huggingface.co/datasets/itseffi/epfl-enterprise-osai-adoption-research-data.dataset-phishing
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/itsprofarul/dataset-phishing.ItaRegol
Corpus ItaRegol
The corpus containing 8 regulations was collected by the research unit of the University of Molise including linguists (Giuliana Fiorentino, Vittorio Ganfi), jurists (Alessandro Cioffi, Maria Assunta Simonelli) and computer scientists (Rocco Oliveto, Marco Russodivito).
Acknowledgements
This contribution is a result of the research conducted within the framework of the PRIN 2020 (Progetti di Rilevante Interesse Nazionale) "VerbACxSS: on analytic verbs… See the full description on the dataset page: https://huggingface.co/datasets/VerbACxSS/ItaRegol.diffing-stats-gemma-2-9b-it-L20-k100-lr1e-04-Crosscoder
