datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CT_DeepLesion-MedSAM2
CT_DeepLesion-MedSAM2 Dataset
Authors
Jun Ma* 1,2,
Zongxin Yang* 3,
Sumin Kim2,4,5,
Bihui Chen2,4,5,
Mohammed Baharoon2,3,5,
Adibvafa Fallahpour2,4,5,
Reza Asakereh4,7,
Hongwei Lyu4,
Bo Wang† 1,2,4,5,6
* Equal contribution † Corresponding author
1AI Collaborative Centre, University Health Network, Toronto, Canada
2Vector… See the full description on the dataset page: https://huggingface.co/datasets/wanglab/CT_DeepLesion-MedSAM2.multimodal-ct-radiology-reports
Perle AI Multi-phase CECT and CT with Radiology Reports
Summary
A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering.
The release has three configurations:
Config
Modality
Subjects
Pairing
cect_3phase
3-phase contrast-enhanced abdominal CT (DICOM)
5
per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.staining-robustness-evaluation
A Protocol for Evaluating Robustness to H&E Staining Variation in Computational Pathology Models
This repository provides the stain references, pretrained models, and experimental results required to:
Define custom staining references using our PLISM reference library
Reproduce our published controlled staining robustness experiments
👉 Code repository: https://github.com/lely475/staining-robustness-evaluation/tree/main
👉 Associated publication: Paper
Overview: How… See the full description on the dataset page: https://huggingface.co/datasets/CTPLab-DBE-UniBas/staining-robustness-evaluation.CTODataset for predicting clinical trial outcomes in drug development. This dataset is part of the work presented in "Automatically Labeling Clinical Trial Outcomes: A Large-Scale Benchmark for Drug Development".
Website: https://chufangao.github.io/CTOD/
Paper: https://arxiv.org/abs/2406.10292
Code: https://github.com/chufangao/ctod
Descriptions:
human_labels contains the manually annotated subset. We follow the same rule-based termination of incomplete status and p-value < 0.05 as in the… See the full description on the dataset page: https://huggingface.co/datasets/chufangao/CTO.conflux-chest-ct
CONFLUX Chest-CT
200,000 synthetic 3D chest CT volumes with structured abnormality and demographic labels, generated by CONFLUX.
Released with the paper CONFLUX: A Latent Diffusion Model for 3D Chest-CT Synthesis with RL Post-Training.
Paper (arXiv) •
Model •
Code — coming soon
About
CONFLUX is a conditional 3D latent generative model for chest CT: a VAE tokenizer
compresses each volume into a compact 16-channel latent, a… See the full description on the dataset page: https://huggingface.co/datasets/gevaertlab/conflux-chest-ct.HistoPlexer-Ultivue
HistoPlexer-Ultivue Dataset
Dataset Summary
The HistoPlexer-Ultivue dataset provides a collection of multimodal histological images for cancer research. It includes whole-slide images (WSIs) of hematoxylin and eosin (H&E) stained tissue, multiplexed immunofluorescence images from Ultivue panels (immuno8 and mdsc), alignment matrices, exclusion masks, and nuclear segmentation outputs. It is a multiplexed dataset for 10 cancer samples from the Tumor Profiler Study. The… See the full description on the dataset page: https://huggingface.co/datasets/CTPLab-DBE-UniBas/HistoPlexer-Ultivue.PHI-CTRL-F16-Fault-Recovery-Telemetry
PHI-CTRL F-16 Actuator Fault Recovery Dataset
High-Fidelity JSBSim 6-DOF Telemetry for Physics-Hybrid Self-Healing Flight Control
Official verification artifacts of the PHI-CTRL (Physics-Hybrid Integrity Control) architecture — a digital-twin-driven, self-healing flight control framework that actively compensates actuator degradation in real time.
Author: Mohammed Bello Sani (SM-Bello)
Affiliation: Air Force Institute of Technology (AFIT), Kaduna · Penelope Inc. / PHI Lab… See the full description on the dataset page: https://huggingface.co/datasets/SM-Bello/PHI-CTRL-F16-Fault-Recovery-Telemetry.CTR_Predictionindia-ctc-to-in-hand-salary-2026
India CTC-to-In-Hand Salary Dataset 2026
This dataset models how annual Cost to Company (CTC) converts into annual net cash and regular monthly net pay for salaried employees in India.
It covers CTC bands from ₹5 lakh to ₹50 lakh under:
0%, 50% and 100% target-variable payout scenarios
capped and full-basic-wage provident-fund models
the Income-tax Act, 2025 provisions applicable to tax year 2026–27
Why this dataset exists
CTC, recurring monthly bank credit and… See the full description on the dataset page: https://huggingface.co/datasets/PaisaSamajhResearch/india-ctc-to-in-hand-salary-2026.large-scale-hate-speech-turkish-v2The dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This is Dataset v2 (Turkish):
The modified dataset that includes 60,310 tweets in Turkish. The annotations with more than 80% agreement are included.
TweetID: Tweet ID from Twitter API
LangID: 0 (Turkish)
TopicID: Domain of the topic 0-Religion, 1-Gender, 2-Race, 3-Politics, 4-Sports
HateLabel: Final hate label decision 0-Normal, 1-Offensive, 2-Hate
GitHub Repo:… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/large-scale-hate-speech-turkish-v2.cta-price-prediction
Dataset Card for Weather-Driven Sri Lankan Tea Market Catalogues
Dataset Description
This dataset contains 12,233 structured records extracted from 105 weekly Forbes & Walker Tea Brokers PDF reports spanning from November 2023 to March 2026. It represents the first machine-readable, comprehensive archive of the Colombo Tea Auction (CTA) prices paired with localized, region-specific lagged weather variables from the Open-Meteo historical archive.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/colombo-tea-auction-prices/cta-price-prediction.large-scale-hate-speech-turkish-v1The dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This is Dataset v1 (Turkish):
The original dataset that includes 100,000 tweets in Turkish. The annotations with more than 60% agreement are included.
TweetID: Tweet ID from Twitter API
LangID: 0 (Turkish)
TopicID: Domain of the topic 0-Religion, 1-Gender, 2-Race, 3-Politics, 4-Sports
HateLabel: Final hate label decision 0-Normal, 1-Offensive, 2-Hate
GitHub Repo:… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/large-scale-hate-speech-turkish-v1.large-scale-hate-speech-v1The dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This is Dataset v1:
The original dataset that includes 100,000 tweets in English. The annotations with more than 60% agreement are included.
TweetID: Tweet ID from Twitter API
LangID: 1 (English)
TopicID: Domain of the topic 0-Religion, 1-Gender, 2-Race, 3-Politics, 4-Sports
HateLabel: Final hate label decision 0-Normal, 1-Offensive, 2-Hate
GitHub Repo:
NOTE:… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/large-scale-hate-speech-v1.CT-RATE-Dataset-cleaneddeprem-tweet-datasetTweets Under the Rubble: Detection of Messages Calling for Help in Earthquake Disaster
The annotated dataset is given at dataset.tsv. We annotate 1,000 tweets in Turkish if tweets call for help (i.e. request rescue, supply, or donation), and their entity tags (person, city, address, status).
Column Name Description
label Human annotation if tweet calls for help (binary classification)
entities Human annotation of entity tags (i.e. person, city, address, and status). The format is… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/deprem-tweet-dataset.ABX-CT-004_sequential_therapy_optimization_loss-v0.1ABX-CT-004 Sequential Therapy Optimization
Purpose
Detect when a planned Drug A then Drug B sequence loses effectiveness.
Core pattern
stress_index high
seq_gap rises and stays high
mono MICs stay below cutoffs during the early window
later failure_flag becomes 1
Files
data/train.csv
data/test.csv
scorer.py
Schema
Each row is one timepoint in a within strain series.
Required columns
row_id
series_id
timepoint_h
organism
strain_id
seq_drug_a
seq_drug_b
stress_index
seq_effectiveness_score… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ABX-CT-004_sequential_therapy_optimization_loss-v0.1.ABX-CT-008_toxicity_additivity_warning-v0.1ABX-CT-008 Toxicity-Additivity Warning
Purpose
Detect when combined toxicity exceeds simple additivity from each drug alone.
Core pattern
stress_index high
exposure indices high
tox_excess rises and stays high
later severe_toxicity_flag appears
Files
data/train.csv
data/test.csv
scorer.py
Schema
Each row is one timepoint in a within series time course.
Required columns
row_id
series_id
timepoint_h
host_model
tox_biomarker_name
tox_biomarker_units
stress_index
exposure_a_index… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ABX-CT-008_toxicity_additivity_warning-v0.1.AI_CTR_GoogleBilTweetNews-sentiment-analysis
Turkish Sentiment Analysis Tweet Dataset: BilTweetNews
The dataset contains tweets related to six major events from Turkish news sources between May 4, 2015
and Jan 8, 2017.
The dataset covers 6 major events:
May 25, 2015 One of the popular football clubs in Turkey, Galatasaray, wins the 2015
Turkish Super League.
Sep 6, 2015 A terrorist group, called PKK, attacked to soldiers in Dağlıca, a village in
southeastern Turkey.
Oct 7, 2015 A Turkish scientist, Aziz Sancar, won the 2015… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/BilTweetNews-sentiment-analysis.large-scale-hate-speech-v2The dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This is Dataset v2:
The modified dataset that includes 68,597 tweets in English. The annotations with more than 80% agreement are included.
TweetID: Tweet ID from Twitter API
LangID: 1 (English)
TopicID: Domain of the topic 0-Religion, 1-Gender, 2-Race, 3-Politics, 4-Sports
HateLabel: Final hate label decision 0-Normal, 1-Offensive, 2-Hate
GitHub Repo:
NOTE:… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/large-scale-hate-speech-v2.GeneContext-Dataset
GeneContext Dataset
The GeneContext dataset is a subsample of the BacBench Essential Genes Dataset designed to evaluate the ability of Genome Language Models (gLMs) to interpret genomic context using an essentiality-based benchmark. The primary dataset contains genes from "mixed essentiality" orthologous groups (i.e. groups containing at least one essential and one non-essential gene) with essentiality labels as predictive targets for a linear probe.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/gene-ctx-2412/GeneContext-Dataset.ctrlpotato-ai-interview-assistant-benchmark
CTRLpotato AI Interview Assistant Cross-review Evidence Matrix (2026)
A citation-ready snapshot of hands-on desktop evidence for six AI interview assistants: Cluely, Interview Coder, LockedIn AI, ULTRACODE AI, Parakeet AI, and Final Round AI.
The package contains 66 assessments across 6 products and 11 shared criteria. Product versions and test dates are preserved in every row.
Important scope
This is a cross-review evidence matrix, not a statistically controlled… See the full description on the dataset page: https://huggingface.co/datasets/ae0j/ctrlpotato-ai-interview-assistant-benchmark.BilTweetNews-event-detection
Turkish Event Detection Tweet Dataset: BilTweetNews
The dataset contains tweets related to six major events from Turkish news sources between May 4, 2015
and Jan 8, 2017.
There are 7 event classes:
E1: May 25, 2015 One of the popular football clubs in Turkey, Galatasaray, wins the 2015
Turkish Super League.
E2: Sep 6, 2015 A terrorist group, called PKK, attacked to soldiers in Dağlıca, a village in
southeastern Turkey.
E3: Oct 7, 2015 A Turkish scientist, Aziz Sancar, won the… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/BilTweetNews-event-detection.RF-200-CTA
RF-200-CTA
Ultimate_CTQRS_good_BGABX-CT-006_pharmacokinetic_mismatch-v0.1ABX-CT-006 Pharmacokinetic Mismatch
Purpose
Detect when PK alignment between two drugs collapses and combined efficacy drops.
Core pattern
stress_index high
efficacy_gap rises and stays high
overlap_index low or half_life_ratio high
mono MICs stay below cutoffs at onset
later_combo_failure_flag appears later
Files
data/train.csv
data/test.csv
scorer.py
Schema
Each row is one timepoint in a within strain series.
Required columns
row_id
series_id
timepoint_h
organism
strain_id
drug_a
drug_b… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ABX-CT-006_pharmacokinetic_mismatch-v0.1.ABX-CT-007_tissue_penetration_discordance-v0.1ABX-CT-007 Tissue Penetration Discordance
Purpose
Detect when two drugs no longer reach the infection site together despite adequate plasma exposure.
Core pattern
stress_index high
site_efficacy_gap rises and stays high
site_coexposure_index low
plasma_conc_a_mg_L and plasma_conc_b_mg_L stay above simple floors
mono MICs stay below cutoffs at onset
later_site_failure_flag appears later
Files
data/train.csv
data/test.csv
scorer.py
Schema
Each row is one timepoint in a within strain series.… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ABX-CT-007_tissue_penetration_discordance-v0.1.ABX-CT-001_synergy_score_decay-v0.1ABX-CT-001 Synergy Score Decay
Purpose
Detect loss of combination synergy before monotherapy resistance appears.
Core pattern
stress_index high
synergy_gap rises and stays high
mono MICs stay below cutoffs during the early window
later at least one mono MIC crosses its cutoff
Files
data/train.csv
data/test.csv
scorer.py
Schema
Each row is one timepoint in a within strain series.
Required columns
row_id
series_id
timepoint_h
organism
strain_id
drug_a
drug_b
stress_index
synergy_score… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ABX-CT-001_synergy_score_decay-v0.1.ABX-CT-002_antagonism_emergence-v0.1ABX-CT-002 Antagonism Emergence
Purpose
Detect when a combination that starts synergistic becomes antagonistic.
Core pattern
baseline synergy_score at least 0.65
later synergy_score at most 0.45 under stress
antagonism persists for 2 consecutive timepoints
mono MICs remain below cutoffs at onset
Files
data/train.csv
data/test.csv
scorer.py
Schema
Each row is one timepoint in a within strain series.
Required columns
row_id
series_id
timepoint_h
organism
strain_id
drug_a
drug_b
stress_index… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ABX-CT-002_antagonism_emergence-v0.1.tweet-topic-detectionPublished tweet dataset used in "Tweet Length Matters: A Comparative Analysis on Topic Detection in Microblogs" includes tweet id and corresponding topic number.
Topic numbers encoded as follows:
Topic Topic Number
BLM Movement 0
Covid-19 1
K-Pop 2
Bollywood 3
Gaming 4
U.S. Politics 5
Out-of-Topic 6
In total, there are 354,310 tweet instances.
More details can be found at https://github.com/avaapm/ECIR2021/
Citation
If you make use of these tools, please cite following paper.… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/tweet-topic-detection.
