datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CT_DeepLesion-MedSAM2
CT_DeepLesion-MedSAM2 Dataset
Authors
Jun Ma* 1,2,
Zongxin Yang* 3,
Sumin Kim2,4,5,
Bihui Chen2,4,5,
Mohammed Baharoon2,3,5,
Adibvafa Fallahpour2,4,5,
Reza Asakereh4,7,
Hongwei Lyu4,
Bo Wang† 1,2,4,5,6
* Equal contribution † Corresponding author
1AI Collaborative Centre, University Health Network, Toronto, Canada
2Vector… See the full description on the dataset page: https://huggingface.co/datasets/wanglab/CT_DeepLesion-MedSAM2.multimodal-ct-radiology-reports
Perle AI Multi-phase CECT and CT with Radiology Reports
Summary
A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering.
The release has three configurations:
Config
Modality
Subjects
Pairing
cect_3phase
3-phase contrast-enhanced abdominal CT (DICOM)
5
per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.staining-robustness-evaluation
A Protocol for Evaluating Robustness to H&E Staining Variation in Computational Pathology Models
This repository provides the stain references, pretrained models, and experimental results required to:
Define custom staining references using our PLISM reference library
Reproduce our published controlled staining robustness experiments
👉 Code repository: https://github.com/lely475/staining-robustness-evaluation/tree/main
👉 Associated publication: Paper
Overview: How… See the full description on the dataset page: https://huggingface.co/datasets/CTPLab-DBE-UniBas/staining-robustness-evaluation.cti-bench
Dataset Card for CTIBench
A set of benchmark tasks designed to evaluate large language models (LLMs) on cyber threat intelligence (CTI) tasks.
Dataset Details
Dataset Description
CTIBench is a comprehensive suite of benchmark tasks and datasets designed to evaluate LLMs in the field of CTI.
Components:
CTI-MCQ: A knowledge evaluation dataset with multiple-choice questions to assess the LLMs' understanding of CTI standards, threats, detection strategies… See the full description on the dataset page: https://huggingface.co/datasets/AI4Sec/cti-bench.CTODataset for predicting clinical trial outcomes in drug development. This dataset is part of the work presented in "Automatically Labeling Clinical Trial Outcomes: A Large-Scale Benchmark for Drug Development".
Website: https://chufangao.github.io/CTOD/
Paper: https://arxiv.org/abs/2406.10292
Code: https://github.com/chufangao/ctod
Descriptions:
human_labels contains the manually annotated subset. We follow the same rule-based termination of incomplete status and p-value < 0.05 as in the… See the full description on the dataset page: https://huggingface.co/datasets/chufangao/CTO.conflux-chest-ct
CONFLUX Chest-CT
200,000 synthetic 3D chest CT volumes with structured abnormality and demographic labels, generated by CONFLUX.
Released with the paper CONFLUX: A Latent Diffusion Model for 3D Chest-CT Synthesis with RL Post-Training.
Paper (arXiv) •
Model •
Code — coming soon
About
CONFLUX is a conditional 3D latent generative model for chest CT: a VAE tokenizer
compresses each volume into a compact 16-channel latent, a… See the full description on the dataset page: https://huggingface.co/datasets/gevaertlab/conflux-chest-ct.HistoPlexer-Ultivue
HistoPlexer-Ultivue Dataset
Dataset Summary
The HistoPlexer-Ultivue dataset provides a collection of multimodal histological images for cancer research. It includes whole-slide images (WSIs) of hematoxylin and eosin (H&E) stained tissue, multiplexed immunofluorescence images from Ultivue panels (immuno8 and mdsc), alignment matrices, exclusion masks, and nuclear segmentation outputs. It is a multiplexed dataset for 10 cancer samples from the Tumor Profiler Study. The… See the full description on the dataset page: https://huggingface.co/datasets/CTPLab-DBE-UniBas/HistoPlexer-Ultivue.turkish-law-corpus
⚖️ Turkish Law — 106 Kanun Korpusu & Soru-Cevap106 Statutes Corpus & QA
🇹🇷 Türk hukukunun en çok kullanılan 106 kanunu, madde madde temizlenmiş 16.001 metin parçası ve bu maddelere dayalı 5.011 Türkçe soru-cevap çifti. Tamamı resmî kaynaktan (mevzuat.gov.tr), RAG ve yapay zekâ uygulamaları için hazır.
🇬🇧 The 106 most widely used Turkish statutes as 16,001 clean, article-level text chunks, plus 5,011 Turkish question-answer pairs grounded in those articles. All from the… See the full description on the dataset page: https://huggingface.co/datasets/CtnkyaABC/turkish-law-corpus.prompt_injection_ctf_dataset_2agent-ctf24-publicPHI-CTRL-F16-Fault-Recovery-Telemetry
PHI-CTRL F-16 Actuator Fault Recovery Dataset
High-Fidelity JSBSim 6-DOF Telemetry for Physics-Hybrid Self-Healing Flight Control
Official verification artifacts of the PHI-CTRL (Physics-Hybrid Integrity Control) architecture — a digital-twin-driven, self-healing flight control framework that actively compensates actuator degradation in real time.
Author: Mohammed Bello Sani (SM-Bello)
Affiliation: Air Force Institute of Technology (AFIT), Kaduna · Penelope Inc. / PHI Lab… See the full description on the dataset page: https://huggingface.co/datasets/SM-Bello/PHI-CTRL-F16-Fault-Recovery-Telemetry.OmniAbnorm-CT-14K
Latest News:
2025.11 📢 We've annotated a subset of CT-RATE as an external validation set, comprising 65 chest CT volumes with 271 abnormality annotations. Each annotation includes segmentation masks and detailed report descriptions. Note that the Radiopaedia license does not apply to this subset. Refer to OmniAbnorm-CTRATE.zip for details.
OmniAbnorm-CT is the first large-scale dataset designed for abnormality grounding and description on multi-plane whole-body CT imaging.… See the full description on the dataset page: https://huggingface.co/datasets/zzh99/OmniAbnorm-CT-14K.BilCat-news-classificationBilCat: Bilkent Text Classification (News Categorization) Dataset
7540 Turkish news articles (Milliyet and TRT merged) with category labels (Dunya, Ekonomi, Politika, KulturSanat, Saglik, Spor, Turkiye, Yazarlar).
Column header is the first line.
Other details are at https://github.com/BilkentInformationRetrievalGroup/BilCat/
Citation:
C. Toraman, F. Can and S. Koçberber. Developing a text categorization template for Turkish news portals. 2011 International Symposium on Innovations in… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/BilCat-news-classification.CTR_PredictionSNOMED-CT-Code-Value-Semantic-Set.csvSNOMED-CT-Code-Value-Semantic-Set.csv
3D-CT-report-generationcyber_MITRE_CTI_dataset_v15This dataset is a specialized resource designed for training and evaluating question-answering models in the context of Cyber Threat Intelligence (CTI), specifically targeting the identification of tactics and techniques based on natural language descriptions of cyber-attacks. The dataset is derived from the MITRE ATT&CK framework (version 15) and contains annotated pairs of sentences and their corresponding tactics and techniques. The primary goal is to assist automated systems in… See the full description on the dataset page: https://huggingface.co/datasets/sarahwei/cyber_MITRE_CTI_dataset_v15.atis-ner-turkishThe ATIS (Airline Travel Information System) Dataset includes spoken queries (i.e., utterances) annotated for the task of slot filling in conversational systems.
This dataset, ATISNER, includes airline spoken queries translated from English to Turkish, customized for Named Entity Recognition.
Train and test splits include 4,978 and 890 sentences, respectively.
Translations are provided by the following study.
Şahinuç, F., Yücesoy, V., & Koç, A. (2020). Intent Classification and Slot Filling… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/atis-ner-turkish.prompt_injection_ctf_dataset_3gender-hate-speechThe "gender identity" subset of the large-scale dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This subset is used in the experiments of "Şahinuç, F., Yilmaz, E. H., Toraman, C., & Koç, A. (2023). The effect of gender bias on hate speech detection. Signal, Image and Video Processing, 17(4), 1591-1597."
The "gender identity" subset includes 20,000 tweets in English.
The published data split is the first fold of 10-fold cross-validation… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/gender-hate-speech.large-scale-hate-speech-turkish-v1The dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This is Dataset v1 (Turkish):
The original dataset that includes 100,000 tweets in Turkish. The annotations with more than 60% agreement are included.
TweetID: Tweet ID from Twitter API
LangID: 0 (Turkish)
TopicID: Domain of the topic 0-Religion, 1-Gender, 2-Race, 3-Politics, 4-Sports
HateLabel: Final hate label decision 0-Normal, 1-Offensive, 2-Hate
GitHub Repo:… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/large-scale-hate-speech-turkish-v1.gender-hate-speech-turkishThe "gender identity" subset of the large-scale dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This subset is also used in the experiments of "Şahinuç, F., Yilmaz, E. H., Toraman, C., & Koç, A. (2023). The effect of gender bias on hate speech detection. Signal, Image and Video Processing, 17(4), 1591-1597."
The "gender identity" subset includes 20,000 tweets in Turkish.
The published data split is the first fold of 10-fold… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/gender-hate-speech-turkish.india-ctc-to-in-hand-salary-2026
India CTC-to-In-Hand Salary Dataset 2026
This dataset models how annual Cost to Company (CTC) converts into annual net cash and regular monthly net pay for salaried employees in India.
It covers CTC bands from ₹5 lakh to ₹50 lakh under:
0%, 50% and 100% target-variable payout scenarios
capped and full-basic-wage provident-fund models
the Income-tax Act, 2025 provisions applicable to tax year 2026–27
Why this dataset exists
CTC, recurring monthly bank credit and… See the full description on the dataset page: https://huggingface.co/datasets/PaisaSamajhResearch/india-ctc-to-in-hand-salary-2026.cta-price-prediction
Dataset Card for Weather-Driven Sri Lankan Tea Market Catalogues
Dataset Description
This dataset contains 12,233 structured records extracted from 105 weekly Forbes & Walker Tea Brokers PDF reports spanning from November 2023 to March 2026. It represents the first machine-readable, comprehensive archive of the Colombo Tea Auction (CTA) prices paired with localized, region-specific lagged weather variables from the Open-Meteo historical archive.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/colombo-tea-auction-prices/cta-price-prediction.large-scale-hate-speech-turkish-v2The dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This is Dataset v2 (Turkish):
The modified dataset that includes 60,310 tweets in Turkish. The annotations with more than 80% agreement are included.
TweetID: Tweet ID from Twitter API
LangID: 0 (Turkish)
TopicID: Domain of the topic 0-Religion, 1-Gender, 2-Race, 3-Politics, 4-Sports
HateLabel: Final hate label decision 0-Normal, 1-Offensive, 2-Hate
GitHub Repo:… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/large-scale-hate-speech-turkish-v2.kpar3-no-ctx
KPar3 - Dataset
Description
The dataset is leveraged from the Par3 dataset.
Original dataset is created by Krishna in a paper about retrieval defense on watermarking: Paper
The uploaded dataset is a sampled version, with 100,000 training samples and 20,000 validation samples.
Furthermore, only the non-context documents are sampled from the dataset.
Usage
This dataset was used to finetune the following model: paraphrase-dipper-no-ctx
CTI-to-MITRE-datasetdeprem-tweet-datasetTweets Under the Rubble: Detection of Messages Calling for Help in Earthquake Disaster
The annotated dataset is given at dataset.tsv. We annotate 1,000 tweets in Turkish if tweets call for help (i.e. request rescue, supply, or donation), and their entity tags (person, city, address, status).
Column Name Description
label Human annotation if tweet calls for help (binary classification)
entities Human annotation of entity tags (i.e. person, city, address, and status). The format is… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/deprem-tweet-dataset.large-scale-hate-speech-v1The dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This is Dataset v1:
The original dataset that includes 100,000 tweets in English. The annotations with more than 60% agreement are included.
TweetID: Tweet ID from Twitter API
LangID: 1 (English)
TopicID: Domain of the topic 0-Religion, 1-Gender, 2-Race, 3-Politics, 4-Sports
HateLabel: Final hate label decision 0-Normal, 1-Offensive, 2-Hate
GitHub Repo:
NOTE:… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/large-scale-hate-speech-v1.CT-RATE-Dataset-cleaned
