datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pxr-structure-pose-pool
PXR Structure Challenge — Full Multi-Model Pose Pool (184 ligands)
Every protein–ligand pose generated during the OpenADMET PXR (pregnane X receptor / NR1I2)
structure-prediction challenge, released openly with per-pose labels so the community can
reuse the compute already spent — and, we hope, crack the problem this data makes visible.
What's here
poses/<model>/<SID>.pdb — one best pose per (model, ligand). Protein chain A + ligand
(resname LIG). 15 models, up… See the full description on the dataset page: https://huggingface.co/datasets/xX-its-amit-Xx/pxr-structure-pose-pool.smishing-syntheticdraft_nbaepfl-enterprise-osai-adoption-research-data
EPFL Enterprise Open-Source AI Adoption Research Dataset
Dataset Summary
This dataset contains mixed-methods research data from 100 organizations regarding their strategic adoption of open-source AI through the Hugging Face ecosystem. The research was conducted at EPFL (École Polytechnique Fédérale de Lausanne) and supports the development of the Gate-Lever framework for enterprise open-source AI adoption.
Dataset Structure
This dataset is organized into 4… See the full description on the dataset page: https://huggingface.co/datasets/itseffi/epfl-enterprise-osai-adoption-research-data.it-support-llmTACK_Tunnel_Data
TACK Tunnel Data (TTD): A Benchmark Dataset for Deep Learning-Based Defect Detection in Tunnels
Tunnels are essential elements of transportation infrastructure, but are increasingly affected by ageing and deterioration mechanisms such as cracking. Regular inspections are required to ensure their safety, yet traditional manual procedures are time-consuming, subjective, and costly. Recent advances in mobile mapping systems and Deep Learning (DL) enable automated visual inspections.… See the full description on the dataset page: https://huggingface.co/datasets/ITS23/TACK_Tunnel_Data.dataset-phishing
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/itsprofarul/dataset-phishing.translatewiki-Myanmar-Englishfitness-question-answersA total of 965 q&a pairs i gathered from the web related to physical activity and fitness.
indian-court-judgements-and-its-summariesNBA_datasetbrazil-deputy-expensesamharic-speech-dataset-2026
Amharic Speech Dataset 2026
Overview
This dataset contains Amharic speech recordings collected using the Leyu Platform for the Leyu Platform Competition 2026.
Language
Amharic (am)
Dialect
Standard Addis Ababa Amharic
Speaker Information
Number of Speakers: 1
Speaker IDs: SPK001
Audio Format
Format: M4A
Duration: 10–60 seconds per recording
Directory Structure
audio/
metadata.csv… See the full description on the dataset page: https://huggingface.co/datasets/ofc-its-phyla/amharic-speech-dataset-2026.brazil-pix-dataopenlibrary-scifi-datanyt-bestsellersit_support_ticket_classification_pegasus_dataset
IT Support Ticket Classification
Description: Automatically categorize and prioritize IT support tickets based on their text descriptions, enabling more efficient resolution and customer support.
How to Use
Here is how to use this model to classify text into different categories:
from transformers import AutoModelForSequenceClassification, AutoTokenizer
model_name = "interneuronai/it_support_ticket_classification_pegasus"
model =… See the full description on the dataset page: https://huggingface.co/datasets/interneuronai/it_support_ticket_classification_pegasus_dataset.mag7-news-datasetSpotifyGlobalChartTotalsLyrics
🎵 Spotify Global Chart Totals & Lyrics
Overview
This dataset compiles cumulative chart performance statistics and song lyrics for every track that has appeared on Spotify's Global Weekly Chart from September 29, 2013 through July 23, 2026.
Each row represents a single song and its lifetime achievements on the chart, including total weeks on chart, peak position, duration at peak, streaming milestones, and full lyrics. The dataset contains 8,058 unique records.… See the full description on the dataset page: https://huggingface.co/datasets/itsmebatuhan/SpotifyGlobalChartTotalsLyrics.dataset-phishing2
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/itsprofarul/dataset-phishing2.Vehicle_Complaints_NHSTAITServiceMatrix
ITServiceMatrix
tags: incident_reporting, interdepartmental_communication, workflow_analysis
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description: The dataset ITServiceMatrix contains detailed records of incidents and requests raised by multiple departments within an organization's IT Service Management system. It provides a comprehensive overview of the issues reported, the departments involved, the communication channels used, and the… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/ITServiceMatrix.It-support-synthetic-datamistral_dark_pattern_dataset995EnglishPodcastTranscripts
Dataset Card for 995 English Podcast Transcripts
Dataset Summary
The 995 English Podcast Transcripts dataset is a collection of detailed text transcripts derived from various English-language podcasts. Containing 995 episodes complete with metadata like summaries, duration, and confidence scores, this dataset is highly valuable for Natural Language Processing (NLP) tasks. The podcasts span diverse categories such as: technology, true crime, business.
It is… See the full description on the dataset page: https://huggingface.co/datasets/itsmebatuhan/995EnglishPodcastTranscripts.watsonx-docs-document-type
Watsonx Docs Document Type Classification
This dataset is a balanced binary document-level classification subset derived
from ibm-research/watsonxDocsQA.
Task
Classify IBM Watsonx documentation pages by their dominant user-facing purpose:
conceptual: documents primarily used to understand or look up information.
how-to: documents primarily used to complete a procedure or fix a problem.
Splits
Split
conceptual
how-to
Total
train
140
140
280… See the full description on the dataset page: https://huggingface.co/datasets/itsjhuang/watsonx-docs-document-type.pc-bench-results
Paper Circle: An Open-source Multi-agent Research Discovery and Analysis Framework
Komal Kumar1, Aman Chadha2, Salman Khan1, Fahad Shahbaz Khan1, Hisham Cholakkal1
1 Mohamed bin Zayed University of Artificial Intelligence 2 AWS Generative AI Innovation Center, Amazon Web Services
[Github] [arXiv] [Live Demo] [Benchmark]
It-support-synthetic-dataBrain-tumor-effect-of-bloodpressure-sugarlevel-and-BMI-ON-ITS-PRESSENCEIT_Students_Pakistan_performance_prediction
