datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
skilltrainbench-public
skilltrainbench training tasks
The training half of the skilltrainbench benchmark suite: for each of the
four datasets, the dev_task_names of its pinned train/test split, in Harbor
task format.
The held-out/test tasks are not in this repository. Neither are the
published splits that are not the pin, nor the tasks that fall outside each
pinned split's task set. Use this repository for skill formation and
training; evaluate on the held-out half, which stays in the private source… See the full description on the dataset page: https://huggingface.co/datasets/armin-aptura/skilltrainbench-public.pyaptamer-AptaComaptos-fundus-imagesapt-cti-reports
APT CTI Reports Dataset
A collection of 2,710 PDF reports on Advanced Persistent Threats (APT) and Cyber Threat Intelligence (CTI).
Structure
├── apt_groups/ (937 files) - Reports attributed to specific APT groups
└── other/ (1773 files) - Multi-attribution, unattributed, and general CTI reports
Filename Format
<GROUP/TYPE>__<YEAR>__<TITLE>.pdf
Examples:
APT28__2019__Fancy_Bear_Campaign.pdf
MULTI__2023__Global_Threat_Report.pdf… See the full description on the dataset page: https://huggingface.co/datasets/hackerman700000/apt-cti-reports.aptos
Dataset Card for "aptos"
More Information needed
bybit-linear-perps-aptusdtLOTL_APT_Red_Team_DatasetLOTL APT Red Team Dataset
Overview
The LOTL APT Red Team Dataset is a comprehensive collection of simulated Advanced Persistent Threat (APT) attack scenarios leveraging Living Off The Land (LOTL) techniques. Designed for cybersecurity researchers, red teamers, and AI/ML practitioners, this dataset focuses on advanced tactics such as DNS tunneling, Command and Control (C2), data exfiltration, persistence, and defense evasion using native system tools across Windows, Linux, macOS, and cloud… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/LOTL_APT_Red_Team_Dataset.mosquitoes-biodcase2026-task5
Mosquitoes — precomputed caches for BioDCASE 2026 Task 5 (CD-MSC)
Data backend for the Mosquitoes code repo (cross-domain
mosquito species classification, BioDCASE 2026 Task 5). It holds the raw audio plus the
derived caches the pipeline consumes, so a clone can reproduce every result without
recomputing embeddings. The Hub layout mirrors the code repo; fetch_data.py pulls a group
into place:
python fetch_data.py --group deployed # ~3.7 GB — light probes + agreement gate… See the full description on the dataset page: https://huggingface.co/datasets/aptemvs/mosquitoes-biodcase2026-task5.aptos_train
Dataset Card for "aptos_dataset"
More Information needed
APTv2
APTv2 Dataset
APTv2 is a large-scale benchmark for animal pose estimation and tracking across 30 species.
It provides high-quality keypoint and tracking annotations for 84,611 animal instances spanning 2,749 video clips (41,235 frames total).
📦 Dataset Overview
Total videos: 2,749
Frames per clip: 15
Total frames: 41,235
Annotated instances: 84,611
Species: 30
Tracks:
Single-frame pose estimation
Low-data generalization
Pose tracking
🧠 Citation
If… See the full description on the dataset page: https://huggingface.co/datasets/DenisKochetov/APTv2.szypulka_tokenized_apt4APTOS2025_OphNet-CatAPTOS25APTOS2019DiabeticRetinopathyaptos
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Asia Pacific Tele-Ophthalmology Society (APTOS) dataset. The images consist of retina scan images to detect diabetic retinopathy.
The original dataset is available at APTOS 2019 Blindness Detection.
These images are resized into 224x224 pixels so that they can be readily used with… See the full description on the dataset page: https://huggingface.co/datasets/bumbledeep/aptos.Aptos-blindness-detectionAPT-36K-poses-controlnet-dataset
Dataset Card for "APT-36K-poses-controlnet-dataset"
More Information needed
Claude-Opus-4.6-stance-distilled-RELATIONALCreated: 2026-03-11
Target: 1000 training examples for QLoRA fine-tuning
Format: OpenAI chat format (system/user/assistant), <think> reasoning traces
Most people create AI to do science problems. I'm creating an AI (Eva) that can effectively navigate life, which is more about relating with people and day-to-day reasoning.
This is the first batch of relating data I distilled from Claude Opus 4.6 oriented with a specific stance, which produces measurably better quality outputs than an unoriented… See the full description on the dataset page: https://huggingface.co/datasets/aptgetupdate/Claude-Opus-4.6-stance-distilled-RELATIONAL.brats2024-aptg-syntheticklientu_aptarnavimas_pirmieji250hplt3_pol_8_9_10_apt_tokenized_splitaptos_train_test
Dataset Card for "aptos_train_test"
More Information needed
apt-eval
🚨 APT-Eval Dataset 🚨
Almost AI, Almost Human: The Challenge of Detecting AI-Polished Writing
📝 Paper, 🖥️ Github, 🎥 Recording
This repository contains the official dataset of the ACL 2025 paper 'Almost AI, Almost Human: The Challenge of Detecting AI-Polished Writing'
APT-Eval is the first and largest dataset to evaluate the AI-text detectors behavior for AI-polished texts.
It contains almost 15K text samples, polished by 5 different LLMs, for 6 different domains, with 2 major… See the full description on the dataset page: https://huggingface.co/datasets/smksaha/apt-eval.APTOS2025_OphNetapt_pretrain_textbook_16k
Dataset Card for "apt_pretrain_textbook_16k"
More Information needed
wmdp_deduped_corporaAPT-36K-Camera
APT-36K-Camera
Per-image camera parameter annotations for the APT-36K dataset
(a large-scale animal pose estimation & tracking benchmark; 35,860 frames from 2,400 video clips across 30 species), captioned by the Puffin-World model. More captioned datasets are provided in our Puffin-16M website.
The collage above visualizes the camera maps on sample images — each
pair shows the up field (green arrows: the projected gravity-up direction)
and the latitude field (colored contours:… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/APT-36K-Camera.APT-RAG-artifacts
APT-RAG Artifacts
Prebuilt Wikipedia passage corpus and FAISS indexes for
A Tree-based RAG Framework for Evidence-Intensive QA via Adaptive Planning and Topology-Aware Evidence Gathering
(Findings of EMNLP 2026).
Paper | Code
Contents
Path
Contents
Size
corpus/monaco/corpus_1M.jsonl.gz
MoNaCo 1M Wikipedia passages
451 MB
vectorDB/monaco/
MoNaCo FAISS HNSW index (Qwen/Qwen3-Embedding-0.6B)
3.1 GB
vectorDB/qampari/
QAMPARI FAISS HNSW index… See the full description on the dataset page: https://huggingface.co/datasets/KJ-Min/APT-RAG-artifacts.holyC-tinyllama-two-layer
HolyC TinyLlama Two-Layer Release
This bundle packages the HolyC TinyLlama work as a two-stage stack with the datasets that fed it. The goal is simple: make the release feel polished, uploadable, and honest about how it was built.
layer1/: explanatory adapter tuned for HolyC code understanding and explanation
layer2/: completion-oriented adapter tuned for HolyC code generation tasks
datasets/codebase/: raw HolyC code corpus
datasets/explanations/: explanation-oriented instruction… See the full description on the dataset page: https://huggingface.co/datasets/Aptlantis/holyC-tinyllama-two-layer.ja-safety-sft-dataset
ja-safety-sft-dataset
日本語LLMの安全性チューニング用 SFT データセットのサンプル (500件) です。
A 500-item sample of the SFT dataset used to safety-tune APTO's Japanese LLMs. English version is provided below.
概要
株式会社APTOが大規模言語モデル(LLM)の安全性向上のために作成した約18,000件の日本語安全性学習データから、比率を維持して抽出したサンプルです。本サンプルでデータの構造と品質を確認できます。
関連モデル
本サンプルの元データを用いて以下のモデルを安全性チューニングしました。
APTO-001/Qwen3.5-27B-SafetyTuned (GGUF)
APTO-001/Qwen3.5-9B-Base-SafetyTuned (GGUF)
APTO-001/Qwen3.5-9B-SafetyTuned (GGUF)… See the full description on the dataset page: https://huggingface.co/datasets/APTO-001/ja-safety-sft-dataset.
