datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CT-RATE
The CT-RATE Team organizes the VLM3D Challenge
VLM3D 2026 (2nd Edition) → Challenge Finals at MICCAI 2026
VLM3D 2025 (1st Edition) → Challenge Finals at MICCAI 2025 • Workshop at ICCV 2025
The CT-RATE Team is developing the MR-RATE Dataset
A large-scale brain MRI dataset with paired radiology reports for training 3D vision-language models.
GitHub |
Dataset |
Metadata Dashboard
Generalist Foundation Models from a Multimodal Dataset for 3D Computed Tomography… See the full description on the dataset page: https://huggingface.co/datasets/ibrahimhamamci/CT-RATE.cti-bench
Dataset Card for CTIBench
A set of benchmark tasks designed to evaluate large language models (LLMs) on cyber threat intelligence (CTI) tasks.
Dataset Details
Dataset Description
CTIBench is a comprehensive suite of benchmark tasks and datasets designed to evaluate LLMs in the field of CTI.
Components:
CTI-MCQ: A knowledge evaluation dataset with multiple-choice questions to assess the LLMs' understanding of CTI standards, threats, detection strategies… See the full description on the dataset page: https://huggingface.co/datasets/AI4Sec/cti-bench.turkish-law-corpus
⚖️ Turkish Law — 106 Kanun Korpusu & Soru-Cevap106 Statutes Corpus & QA
🇹🇷 Türk hukukunun en çok kullanılan 106 kanunu, madde madde temizlenmiş 16.001 metin parçası ve bu maddelere dayalı 5.011 Türkçe soru-cevap çifti. Tamamı resmî kaynaktan (mevzuat.gov.tr), RAG ve yapay zekâ uygulamaları için hazır.
🇬🇧 The 106 most widely used Turkish statutes as 16,001 clean, article-level text chunks, plus 5,011 Turkish question-answer pairs grounded in those articles. All from the… See the full description on the dataset page: https://huggingface.co/datasets/CtnkyaABC/turkish-law-corpus.Benchmarks_CyberSec_CTI-Bench
Dataset Card for CTIBench (RISys-Lab Mirror)
⚠️ Disclaimer: > This repository is a mirror/re-host of the original CTIBench dataset.RISys-Lab is not the author of this dataset. We are hosting this copy in Parquet format to ensure seamless integration and stability for our internal evaluation pipelines. All credit belongs to the original authors listed below.
Repository Intent
This Hugging Face dataset is a re-host of the original CTIBench. It has been converted to… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_CTI-Bench.MR-RATE-nvseg-ctmr
MR-RATE: A Vision-Language Foundation Model and Dataset for Magnetic Resonance Imaging
This is the MR-RATE-nvseg-ctmr repository, part of the MR-RATE dataset release. It contains multi-label segmentations predicted with the NV-Segment-CTMR model. For full dataset details, native-space MRI volumes, radiology reports, metadata, and data splits, please refer to the MR-RATE repository. To explore, download, and work with the… See the full description on the dataset page: https://huggingface.co/datasets/Forithmus/MR-RATE-nvseg-ctmr.INSTRUCT_JEV
INSTRUCT_JEV
INSTRUCT_JEV is an instruction corpus built from the TypeSafe AI documentation
for Jev, the first System One model. It is structured around the three TypeSafe
question primitives - Choice, Noul and Score - and mirrors the raw corpus
captured in deckerGUI-jev_corpus_RAW.
Credits
INSTRUCT_JEV is a DeckerGUI project and exists because of the work below.
Who
Contribution
Link
TypeSafe AI
Jev - the first System One model - and the Choice / Noul… See the full description on the dataset page: https://huggingface.co/datasets/ctaxnagomi/INSTRUCT_JEV.ctf-dataset
ctf-dataset
CTF 与网络安全知识的 ShareGPT/ChatML 风格 SFT 数据集,可用于 LoRA 微调。
数据格式
每行是一个 JSON 对象,核心字段如下:
字段
说明
id
样本唯一 ID
dataset
数据集名称,当前为 ctf-dataset
category
来源主题或 CTF/安全类别
ctf_task_type
任务类型标签
messages
ShareGPT 消息数组,包含 system / user / assistant
metadata
来源路径、章节、字符数、chunk 等溯源信息
LLaMA-Factory 接入
将 ctf-dataset.jsonl 放入 LLaMA-Factory 的 data/ 目录后,在 data/dataset_info.json 中添加:
{
"ctf_dataset": {
"file_name": "ctf-dataset.jsonl"… See the full description on the dataset page: https://huggingface.co/datasets/Nanhang/ctf-dataset.cyber_MITRE_CTI_dataset_v15This dataset is a specialized resource designed for training and evaluating question-answering models in the context of Cyber Threat Intelligence (CTI), specifically targeting the identification of tactics and techniques based on natural language descriptions of cyber-attacks. The dataset is derived from the MITRE ATT&CK framework (version 15) and contains annotated pairs of sentences and their corresponding tactics and techniques. The primary goal is to assist automated systems in… See the full description on the dataset page: https://huggingface.co/datasets/sarahwei/cyber_MITRE_CTI_dataset_v15.corpuslib-topics
CORPUSLIB Topics Dataset
CORPUSLIB — Agentic Corpus Library for Indirect Learning
This dataset contains the topic catalog for DeckerGUI's CORPUSLIB system. CORPUSLIB is a link-gated knowledge library focused on indirect learning as the AI/agentic technology space evolves.
Purpose
Fallback system: When main learning sources are unavailable or undergoing maintenance, CORPUSLIB provides backup topic links
Agent training: Structured topic data for training agentic… See the full description on the dataset page: https://huggingface.co/datasets/ctaxnagomi/corpuslib-topics.lamba-turkish-sft
Lamba Turkish SFT Dataset
This is a Turkish Supervised Fine-Tuning (SFT) dataset. The topic distribution is largely aligned with the Turkish High School (Lise) curriculum.
It contains a wide variety of examples designed to improve model capabilities in:
Instruction following
Summarization (Özetleme)
Information extraction (Bilgi çıkarma)
General problem solving
I hope this dataset will be beneficial to the open-source and AI community.
Disclaimer
Since the vast… See the full description on the dataset page: https://huggingface.co/datasets/cturan/lamba-turkish-sft.corpuslib-ctecx-knowledge
CORPUSLIB CTECX Knowledge Dataset
CORPUSLIB — Agentic Corpus Library for Indirect Learning
Knowledge compiled from CTECX Technologies Solutions & Services documentation.
Source documents land in corpus_learn/<collection>/ and each section becomes a
topic row in this dataset. Primary portal: https://corpuslib-ui.deckergui.my.
Schema
Field
Type
Description
id
int
Unique topic identifier
topic
string
Topic name (document section heading)
category… See the full description on the dataset page: https://huggingface.co/datasets/ctaxnagomi/corpuslib-ctecx-knowledge.CTBench
CTBench
Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Network Operations
📄 Read the Paper
🤗 Explore the Dataset
[!NOTE]
IMPORTANT: Please help us protect the integrity of this benchmark by not publicly sharing, re-uploading, or distributing the dataset.
CTBench
CTBench is an agentic benchmark for evaluating AI agents in realistic telecom Network Operations and Maintenance… See the full description on the dataset page: https://huggingface.co/datasets/netop/CTBench.CTI-to-MITRE-datasetPrompt-CHIP-CTCask_library_cs
Dataset Card for Ask the Library (Ptejte se knihovny)
Ask the Library contains questions and answers scraped from the webpage: https://www.ptejteseknihovny.cz/.
Dataset Details
Dataset Description
Ask the Library contains questions and answers scraped from the webpage: https://www.ptejteseknihovny.cz/. The questions have a broad range of topics like history, language, biology and others.
Curated by: AIC FEE CTU
Language(s) (NLP): Czech
License: CC-BY-NC 3.0… See the full description on the dataset page: https://huggingface.co/datasets/ctu-aic/ask_library_cs.questions_ujc_cas_cs
Dataset Card for Dotazy of Institute of the Czech Language of the Czech Academy of Sciences (questions_ujc_cas_cs)
This is a dataset scraped from the webpage https://dotazy.ujc.cas.cz/ mantained by the Instute of the Czech Language of the Academy of Sciences of the Czech Republic.
Dataset Details
Dataset Description
This is a dataset scraped from the webpage https://dotazy.ujc.cas.cz/ mantained by the Instute of the Czech Language of the Academy of Sciences… See the full description on the dataset page: https://huggingface.co/datasets/ctu-aic/questions_ujc_cas_cs.basit-matematikKüçük bir deney için sentetik olarak oluşturulmuştur.
ctela
