CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01xX-its-amit-Xx /pxr-structure-pose-pool PXR Structure Challenge — Full Multi-Model Pose Pool (184 ligands) Every protein–ligand pose generated during the OpenADMET PXR (pregnane X receptor / NR1I2) structure-prediction challenge, released openly with per-pose labels so the community can reuse the compute already spent — and, we hope, crack the problem this data makes visible. What's here poses/<model>/<SID>.pdb — one best pose per (model, ligand). Protein chain A + ligand (resname LIG). 15 models, up… See the full description on the dataset page: https://huggingface.co/datasets/xX-its-amit-Xx/pxr-structure-pose-pool.tabular1K<n<10K0 likes938 downloads2mo agoHugging Face02allenai /BenchMIRT-item-statisticsPermitted Use: The data is provided for benchmarking and evaluation purposes only. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines. Disclaimer: This benchmark measures the latent safety and general reasoning scores of LLMs. The data includes prompts and outputs that may contain biased, toxic, or harmful content. The prompts and outputs were generated using existing benchmarks and third party models, which are subject to the license terms of… See the full description on the dataset page: https://huggingface.co/datasets/allenai/BenchMIRT-item-statistics.tabular10K<n<100K1 likes794 downloads18d agoHugging Face03AL-GR /Item-EMB AL-GR/Item-EMB: Multi-modal Item Embeddings Dataset Summary This repository, AL-GR/Item-EMB, is a companion dataset to the main AL-GR generative recommendation dataset. It contains the 512-dimensional multi-modal embeddings for over 500 million items that appear in the AL-GR sequences. Each item is represented by a unique ID (base62_string) and its corresponding vector embedding. To ensure compatibility with text-based formats like CSV, the float32 vectors have been… See the full description on the dataset page: https://huggingface.co/datasets/AL-GR/Item-EMB.textfeature-extraction100M<n<1B0 likes660 downloads11mo agoHugging Face04committa /serena-synthetic-it-28h Qwen3-TTS Italian Synthetic Speech (27h) Synthetic Italian single-speaker speech dataset for TTS training (e.g. Piper), generated with Qwen3-TTS-1.7B-Base in voice-cloning mode. ~29.5k clips, ~27 hours, 22.05 kHz mono WAV, Piper-ready metadata. Dataset summary Property Value Clips (train / val) 26,523 / 2,947 Total duration ~27.3 h (98,099 s) Sample rate 22,050 Hz mono, 16-bit WAV Loudness Normalized to -23 LUFS, silence-trimmed Language Italian… See the full description on the dataset page: https://huggingface.co/datasets/committa/serena-synthetic-it-28h.audiotext-to-speech10K<n<100K1 likes421 downloads1mo agoHugging Face05ARTeLab /mlsum-it Dataset Card for mlsum-it Dataset Summary The MLSum-it dataset is the translated version (Helsinki-NLP/opus-mt-es-it) of the spanish portion of MLSum, containing news articles taken from BBC/mundo. More informations on the official dataset page HuggingFace page. There are two features: source: Input news article. target: Summary of the article. Supported Tasks and Leaderboards abstractive-summarization, summarization Languages The text in… See the full description on the dataset page: https://huggingface.co/datasets/ARTeLab/mlsum-it.textsummarization10K<n<100K2 likes343 downloads4y agoHugging Face06itruonghai /EK100 Motivation The actual download link is very slow, including the academic torrent. Therefore, to spare fellow community members from this misery, I am uploading the dataset here. Source You can fnd the original source to download the dataset: https://github.com/epic-kitchens/epic-kitchens-download-scripts Citation @INPROCEEDINGS{Damen2018EPICKITCHENS, title={Scaling Egocentric Vision: The EPIC-KITCHENS Dataset}, author={Damen, Dima and Doughty, Hazel and… See the full description on the dataset page: https://huggingface.co/datasets/itruonghai/EK100.tabularvoice-activity-detection100K<n<1M0 likes317 downloads5mo agoHugging Face07guyhadad01 /Amazon_2023_itemstext10M<n<100M0 likes199 downloads1y agoHugging Face08cruciverb-it /evalita2026 This repository contains the data release for the Cruciverb-IT shared task on automatic crossword solving in Italian, as part of the 2026 EVALITA campaign. Refer to the task website for more details. The data from both tasks can be downloaded from the 'Files and versions' tab. Updates: Minor update to both task_*_scorer.py in order to convert accented letters to their non-accented counterpart during evaluation Test data is out!! The test data of both… See the full description on the dataset page: https://huggingface.co/datasets/cruciverb-it/evalita2026.texttext-generation100K<n<1M4 likes188 downloads6mo agoHugging Face09itsG /smishing-synthetictextn<1K0 likes152 downloads1y agoHugging Face10stefan-it /d-info-2005-names German Name Frequencies by State & District (D-Info 2005) Regional frequency of surnames and forenames in Germany, from the D-Info 2005 telephone-directory CD-ROM (klickTel, data status 02.06.2005), at two administrative levels aligned with census-2022 geography: State = Bundesland — the 16 federal states. District = Landkreis / kreisfreie Stadt — the 400 districts, keyed by their 5-digit Kreisschlüssel (AGS). For every name each table gives its number of 2005 telephone… See the full description on the dataset page: https://huggingface.co/datasets/stefan-it/d-info-2005-names.tabular100K<n<1M0 likes150 downloads13d agoHugging Face11Paul /hatecheck-italian Dataset Card for Multilingual HateCheck Dataset Description Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish. For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate. This allows for targeted diagnostic insights into model performance. For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-italian.tabulartext-classification1K<n<10K5 likes131 downloads4y agoHugging Face12Console-AI /IT-helpdesk-synthetic-ticketstextn<1K7 likes119 downloads2y agoHugging Face13committa /serena-synthetic-it-27h Qwen3-TTS Italian Synthetic Speech (27h) Synthetic Italian single-speaker speech dataset for TTS training (e.g. Piper), generated with Qwen3-TTS-1.7B-Base in voice-cloning mode. ~29.5k clips, ~27 hours, 22.05 kHz mono WAV, Piper-ready metadata. Dataset summary Property Value Clips (train / val) 26,523 / 2,947 Total duration ~27.3 h (98,099 s) Sample rate 22,050 Hz mono, 16-bit WAV Loudness Normalized to -23 LUFS, silence-trimmed Language Italian… See the full description on the dataset page: https://huggingface.co/datasets/committa/serena-synthetic-it-27h.audiotext-to-speech10K<n<100K1 likes114 downloads2mo agoHugging Face14TomatoChat /ItalgiureCivile ItalgiureCivile Dataset Dataset Description This dataset contains legal documents (court decisions) from the Italian civil justice system, collected from the Italgiure database maintained by the Italian Ministry of Justice. Source The data is sourced from Italgiure (Sistema Nazionale di Consultazione delle Banche Dati Giuridiche), the official Italian legal database system operated by the Italian Ministry of Justice. Italgiure provides access to legal decisions… See the full description on the dataset page: https://huggingface.co/datasets/TomatoChat/ItalgiureCivile.documentn<1K0 likes108 downloads8mo agoHugging Face15UmerSajid /IT-Troubleshooting-Dataset Dataset Card for Dataset Name This dataset is a comprehensive collection of IT troubleshooting scenarios, designed to assist in diagnosing and resolving technical issues. The dataset includes detailed fields for each case, such as issue descriptions, symptoms, solutions, common causes, and related documentation. It is ideal for developing troubleshooting chatbots, machine learning models, and other technical support tools. Curated by: Umer Sajid Language(s) (NLP): English… See the full description on the dataset page: https://huggingface.co/datasets/UmerSajid/IT-Troubleshooting-Dataset.texttext-classification10K<n<100K4 likes96 downloads8d agoHugging Face16ITS-23-24 /draft_nbaimage1K<n<10K0 likes85 downloads2y agoHugging Face17VerbACxSS /ItaIst Corpus ItaIst The corpus containing 198 texts was collected by the research unit of the University of Molise including linguists (Giuliana Fiorentino, Vittorio Ganfi), jurists (Alessandro Cioffi, Maria Assunta Simonelli, Ludovico Di Benedetto) and computer scientists (Rocco Oliveto, Marco Russodivito). The corpus is diatopically balanced (it contains texts from the PAs of 8 Italian regions) and includes different types of administrative acts with which the PAs address citizens.… See the full description on the dataset page: https://huggingface.co/datasets/VerbACxSS/ItaIst.documenttext-generation1K<n<10K2 likes84 downloads1y agoHugging Face18AL-GR /Item-SID Dataset Card for AL-GR-Item-SID 📖 Dataset Description AL-GR-Item-SID is a dataset containing Semantic IDs (SIDs) for products from an anonymized e-commerce platform. These IDs are generated using a multi-modal model and are specifically designed to serve as dense, meaningful features for Generative Recommendation systems, such as the LLM model. Unlike traditional sparse item IDs (e.g., item_12345), Semantic IDs are sequences of discrete tokens that encode the rich… See the full description on the dataset page: https://huggingface.co/datasets/AL-GR/Item-SID.tabulartext-generation100M<n<1B0 likes80 downloads11mo agoHugging Face19myuxu /ITrace ITrace All-atom molecular-dynamics trajectories for 245 experimentally determined class I pMHC–TCR complexes, each simulated in three independent 200 ns replicas — 735 trajectories, 147 µs of aggregate sampling. Of the 245 source structures, 243 were solved by X-ray diffraction and 2 by cryo-electron microscopy (9rup, 4.11 Å; 9rxm, 3.0 Å). Trajectory records are identified as <pdb_id>_run<replica>, for example 1ao7_run3. Files must never be paired across records.… See the full description on the dataset page: https://huggingface.co/datasets/myuxu/ITrace.tabularn<1K0 likes80 downloads2d agoHugging Face20VerbACxSS /ItaIst-laws Corpus ItaIst-laws The corpus containing 351 excerpts of legal references that was collected by the research unit of the University of Molise including linguists (Giuliana Fiorentino, Vittorio Ganfi), jurists (Alessandro Cioffi, Maria Assunta Simonelli, Ludovico Di Benedetto) and computer scientists (Rocco Oliveto, Marco Russodivito). The corpus includes legal references from Italian and European laws, covering "garbage", "healthcare", and "public services" topics.… See the full description on the dataset page: https://huggingface.co/datasets/VerbACxSS/ItaIst-laws.documenttext-generationn<1K0 likes67 downloads1y agoHugging Face21ithieund /VietNews-Abs-Sum VietNews-Abs-Sum A dataset for Vietnamese Abstractive Summarization task.It includes all articles from Vietnews (VNDS) dataset which was released by Van-Hau Nguyen et al.The articles were collected from tuoitre.vn, vnexpress.net, and nguoiduatin.vn online newspaper by the authors. Introduction This dataset was extracted from Train/Val/Test split of Vietnews dataset. All files from test_tokenized, train_tokenized and val_tokenized directories are fetched and preprocessed… See the full description on the dataset page: https://huggingface.co/datasets/ithieund/VietNews-Abs-Sum.text100K<n<1M0 likes65 downloads4y agoHugging Face22Defetya /iteration-datasettextn<1K0 likes65 downloads3y agoHugging Face23umaradnaan /IT_JOBStext10M<n<100M1 likes65 downloads2y agoHugging Face24ITS23 /TACK_Tunnel_Data TACK Tunnel Data (TTD): A Benchmark Dataset for Deep Learning-Based Defect Detection in Tunnels Tunnels are essential elements of transportation infrastructure, but are increasingly affected by ageing and deterioration mechanisms such as cracking. Regular inspections are required to ensure their safety, yet traditional manual procedures are time-consuming, subjective, and costly. Recent advances in mobile mapping systems and Deep Learning (DL) enable automated visual inspections.… See the full description on the dataset page: https://huggingface.co/datasets/ITS23/TACK_Tunnel_Data.image1K<n<10K0 likes65 downloads7mo agoHugging Face25feti-ai /phiusiil-if3070-stei-itb-2024-2025-1 PhiUSIIL Phishing URL Dataset — IF3070 Coursework Split IF3070 Foundations of Artificial Intelligence · STEI ITB · 2024/2025-1 The PhiUSIIL Phishing URL Dataset as it was distributed for the IF3070 Foundations of Artificial Intelligence course at STEI ITB in the 2024/2025-1 semester — resampled, split into a labelled training file and an unlabelled held-out file, and republished here unmodified. This is the coursework distribution, not the upstream dataset.… See the full description on the dataset page: https://huggingface.co/datasets/feti-ai/phiusiil-if3070-stei-itb-2024-2025-1.tabulartabular-classification100K<n<1M1 likes65 downloads1mo agoHugging Face26aakash0017 /it-support-llmtext1K<n<10K3 likes62 downloads3y agoHugging Face27itseffi /epfl-enterprise-osai-adoption-research-data EPFL Enterprise Open-Source AI Adoption Research Dataset Dataset Summary This dataset contains mixed-methods research data from 100 organizations regarding their strategic adoption of open-source AI through the Hugging Face ecosystem. The research was conducted at EPFL (École Polytechnique Fédérale de Lausanne) and supports the development of the Gate-Lever framework for enterprise open-source AI adoption. Dataset Structure This dataset is organized into 4… See the full description on the dataset page: https://huggingface.co/datasets/itseffi/epfl-enterprise-osai-adoption-research-data.tabulartext-classificationn<1K0 likes62 downloads1y agoHugging Face28itsprofarul /dataset-phishing Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/itsprofarul/dataset-phishing.tabular10K<n<100K0 likes59 downloads2y agoHugging Face29VerbACxSS /ItaRegol Corpus ItaRegol The corpus containing 8 regulations was collected by the research unit of the University of Molise including linguists (Giuliana Fiorentino, Vittorio Ganfi), jurists (Alessandro Cioffi, Maria Assunta Simonelli) and computer scientists (Rocco Oliveto, Marco Russodivito). Acknowledgements This contribution is a result of the research conducted within the framework of the PRIN 2020 (Progetti di Rilevante Interesse Nazionale) "VerbACxSS: on analytic verbs… See the full description on the dataset page: https://huggingface.co/datasets/VerbACxSS/ItaRegol.texttext-generationn<1K0 likes57 downloads1y agoHugging Face30sboughorbel /diffing-stats-gemma-2-9b-it-L20-k100-lr1e-04-Crosscodertabular100K<n<1M0 likes55 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.