datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CT2USforKidneySeg
CT2USforKidneySeg
CT (source-domain) slices with kidney segmentation masks used in
Song et al., "CT2US: Cross-modal transfer learning for kidney
segmentation in ultrasound images with synthesized data."
Ultrasonics 122 (2022) 106706. DOI:
10.1016/j.ultras.2022.106706.
Mirror of the public Kaggle release
siatsyx/ct2usforkidneyseg.
Contents
4586 paired samples at 256x256, single split train.
image: grayscale CT slice (PNG, 8-bit).
mask: kidney segmentation mask… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/CT2USforKidneySeg.voxknesset-whisper-large-v3-ct2-inference
VoxKnesset × ivrit-ai/whisper-large-v3-ct2 — inference results
Transcriptions of ivrit-ai/VoxKnesset
produced by ivrit-ai/whisper-large-v3-ct2
(faster-whisper, float16, batched, language=he, VAD on), on 8× NVIDIA A40.
Columns: speaker metadata (from VoxKnesset), duration_s, reference_text
(official Knesset protocol), model_transcription, segments_json
(start/end/text), infer_time_s, split.
Current contents: 10-example pilot from the test split (data/results_10.parquet).
Full-run… See the full description on the dataset page: https://huggingface.co/datasets/Dolevabudi/voxknesset-whisper-large-v3-ct2-inference.ct2xr-projections
Component-wise translated CT projections
Synthetic chest radiographs derived from 21,887 chest CT volumes of CT-RATE, released as the
translated outputs of the pipeline described in
Anatomy-Decomposed Chest Computed Tomography (CT) Projections as Scalable Supervision for Bone
Suppression in Chest Radiographs
Angaitkar, Kumar, Satia, Rao, Mittal, Tadepalli, Putha — arXiv:2609.24937 (2026; under review at Medical Image Analysis).
Released by Qure.ai. Version 1.0.0 (2026-09-15).… See the full description on the dataset page: https://huggingface.co/datasets/qureaiorg/ct2xr-projections.ChangeMore-prompt-injection-eval
ChangeMore-prompt-injection-eval
This dataset is designed to support the evaluation of prompt injection detection capabilities in large language models (LLMs).
To address the lack of Chinese-language prompt injection attack samples, we have developed a systematic data generation algorithm that automatically produces a large volume of high-quality attack samples. These samples significantly enrich the security evaluation ecosystem for LLMs, especially in the Chinese context.
The… See the full description on the dataset page: https://huggingface.co/datasets/CTCT-CT2/ChangeMore-prompt-injection-eval.patchddm3d-ct2pet-results
PatchDDM-3D & ALDM CT-to-PET Test Results
This repository contains evaluation metrics, compressed 3D volume arrays (.npz), and case outputs for 3D medical image translation models (CT -> PET):
PatchDDM-3D: Memory-Efficient 3D Brownian Bridge Diffusion Model.
ALDM: Anatomy-Aware 3D Latent Diffusion Model.
Repository Layout
ALDM/:
test_best_last/best_last_test_metrics.json: Full 136-case test set evaluation summary metrics.
test_best_last/last/cases/: Individual… See the full description on the dataset page: https://huggingface.co/datasets/pndhpndh/patchddm3d-ct2pet-results.ct2-synth-rawopus-mt-ct2CT26_Task1_SourceRetrievalForScientificWebClaims
Overview
This repo contains the data for CheckThat!Lab 2026 - Task 1: Source Retrieval for Scientific Web Claims
Task Definition
Source Retrieval for Scientific Web Claims: Given a social media post that contains a scientific claim and an implicit reference to a scientific paper (mentions it without a URL), retrieve the mentioned paper from a pool of candidate papers.
Datasets
The collection set contains information about 10k publications. The query sets (train… See the full description on the dataset page: https://huggingface.co/datasets/sschellhammer/CT26_Task1_SourceRetrievalForScientificWebClaims.CLEF_CT23_1A_checkworthy_multimodal_english_v2CT2_AG_hi
Counter Turing Test (CT²): Investigating AI-Generated Text Detection for Hindi
AG_hi Dataset
Paper Links
ArXiv Link
About the Dataset
The AI-generated news article in Hindi (AG_hi) dataset is introduced to assess the effectiveness of AI-generated text detection (AGTD) techniques for Hindi.
Dataset Overview
This dataset comprises two categories of Hindi news articles: human-written and AI-generated. The human-written articles were… See the full description on the dataset page: https://huggingface.co/datasets/ishank31/CT2_AG_hi.ct2mri-img2img-trainCT2ct2mri-img2imgct2mri-instructpix2pixct2tsmahvq-reproct219-vietnamese-raw-400k
wheevu/ct219-vietnamese-raw-400k
Bộ dữ liệu văn bản tiếng Việt thô đã tiền xử lý, dùng để huấn luyện
next-token language model (CT219 - NLP final project).
Nguồn dữ liệu
Source dataset: VTSNLP/vietnamese_curated_dataset
Source split: train
Pinned source revision: b81fcce58945970117a1b56d50ec81be2628a5c3
Source licence: không công bố - repo này được tạo ở chế độ private vì lý do đó.
Mục đích
Huấn luyện next-token language model cho tiếng Việt.… See the full description on the dataset page: https://huggingface.co/datasets/wheevu/ct219-vietnamese-raw-400k.ct2mri-fluxwhisper-large-v3-ct2ct2Ct2QTcSeCT2.5_MultiviewGen_Galbot
