datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ja-safety-sft-dataset
ja-safety-sft-dataset
日本語LLMの安全性チューニング用 SFT データセットのサンプル (500件) です。
A 500-item sample of the SFT dataset used to safety-tune APTO's Japanese LLMs. English version is provided below.
概要
株式会社APTOが大規模言語モデル(LLM)の安全性向上のために作成した約18,000件の日本語安全性学習データから、比率を維持して抽出したサンプルです。本サンプルでデータの構造と品質を確認できます。
関連モデル
本サンプルの元データを用いて以下のモデルを安全性チューニングしました。
APTO-001/Qwen3.5-27B-SafetyTuned (GGUF)
APTO-001/Qwen3.5-9B-Base-SafetyTuned (GGUF)
APTO-001/Qwen3.5-9B-SafetyTuned (GGUF)… See the full description on the dataset page: https://huggingface.co/datasets/APTO-001/ja-safety-sft-dataset.APT_STYLE_Privilege_Escalation_Dataset
APT Privilege Escalation Dataset
Overview
The APT Privilege Escalation Dataset is a comprehensive collection of advanced and unique privilege escalation techniques tailored for Red Team training and offensive cybersecurity operations. This dataset, comprising 1000 entries, simulates real-world Advanced Persistent Threat (APT) tactics, focusing on exploiting misconfigurations, vulnerabilities, and novel attack vectors to achieve elevated privileges on Linux-based systems.… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/APT_STYLE_Privilege_Escalation_Dataset.AptMQL-Bench
AptMQL-Bench
📄 Paper: AptMQL-Bench: From Text-to-SQL to Text-to-MQL via Access-Pattern Schema Design and Data-Preserving Migration · arXiv: coming soon
AptMQL-Bench is a benchmark for text-to-MQL — the task of translating human-readable requests into executable MongoDB Query Language (MQL) aggregation pipelines. It contains 21 document-oriented databases, 3,181 natural-language requests, and their associated gold MQL queries.
Most existing text-to-MQL resources are… See the full description on the dataset page: https://huggingface.co/datasets/giahy2507/AptMQL-Bench.cluebbers__Llama-3.1-8B-paraphrase-type-generation-apty-sigmoid-details
Dataset Card for Evaluation run of cluebbers/Llama-3.1-8B-paraphrase-type-generation-apty-sigmoid
Dataset automatically created during the evaluation run of model cluebbers/Llama-3.1-8B-paraphrase-type-generation-apty-sigmoid
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cluebbers__Llama-3.1-8B-paraphrase-type-generation-apty-sigmoid-details.Aptos_vulnerability_audit_dataset
Uploaded Dataset
Name: Aptos_vulnerability_audit_dataset
Organization: Armur
Project: Aptos Smart Contract Audit
License: apache-2.0
Language: en
aptigrad-datasetcounseling-aptness
APTNESS — Counseling Dialogues (processed)
英文共情策略咨询对话(APTNESS),LLaMA-Factory 格式;含 database / ed / extes 三个 split。
本仓库是 counselor_agent 项目中,经统一预处理器落地到 dataset/processed/ 的
APTNESS 数据集。所有记录采用统一 schema(case_id / source / lang /
messages[] + 各数据集特有的可选标注 / profile)。
规模
aptness_db: 9,659 dialogues, 39,792 turns (avg 4.12)
aptness_ed: 30 dialogues, 240 turns (avg 8.0)
aptness_extes: 10 dialogues, 238 turns (avg 23.8)
文件
文件
类型
条数
大小… See the full description on the dataset page: https://huggingface.co/datasets/XuShihao6715/counseling-aptness.job-aptitude
Job Role Aptitude Dataset
A dataset of aptitude and assessment questions for training and evaluating models that generate multiple-choice questions from job context: general role, skills, experience, and employment type.
Coverage includes Software Developer, Doctor, Pharmacist, Nurse, Teacher, Data Analyst, engineers, managers, and many other general professions.
Dataset Files
File
Purpose
aptitude_train.jsonl
Training split
aptitude_test.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/Lumixis/job-aptitude.aptitude-test-jobclaude-opus-4.6-10000xThis is a high-fidelity reasoning dataset synthesized using Claude Opus 4.6. The dataset is designed to capture the model's internal "Chain of Thought" and reasoning traces, specifically focusing on mathematical accuracy and structured logical deduction.
The dataset is intended for Supervised Fine-Tuning (SFT) and Distillation, allowing smaller open-source models to inherit the sophisticated reasoning patterns of Claude Opus 4.6.
Dataset Description
This collection combines high-difficulty… See the full description on the dataset page: https://huggingface.co/datasets/aptgetupdate/claude-opus-4.6-10000x.jobfit-aptitude-test
Job Role Aptitude Dataset
A dataset of role-based aptitude and assessment questions for training and evaluating models that generate or score technical and professional aptitude content. Coverage spans software engineering, data and analytics, security, cloud and DevOps, product and project management, finance, HR, and related roles.
Dataset Files
File
Purpose
aptitude_train.jsonl
Training split — multiple-choice aptitude items across many job roles… See the full description on the dataset page: https://huggingface.co/datasets/Lumixis/jobfit-aptitude-test.unseen-aptitude-qa-dataset
Unseen Aptitude QA Dataset
This dataset contains categorized quantitative and logical aptitude questions explicitly structured for campus placement preparation (e.g., TCS, Wipro, Infosys). It is formatted using the standard ChatML / OpenAI Messages schema, making it natively compatible with fine-tuning models like SmolLM2-1.7B.
Dataset Structure
Each data sample contains a messages array featuring a structured system persona, metadata-enriched user questions, and detailed… See the full description on the dataset page: https://huggingface.co/datasets/Prathamesh25/unseen-aptitude-qa-dataset.aptos_develop
Aptos-Move-QA-5K: A High-Quality Instruction Dataset for Move on Aptos
Welcome to the Aptos-Move-QA-5K dataset, a specialized collection of over 5,000 high-quality question-and-answer pairs designed to fine-tune Large Language Models (LLMs) into expert assistants for the Move programming language on the Aptos blockchain.
This dataset was curated by Cotrain.AI as a foundational step in our mission to build open, specialized, and community-driven artificial intelligence.… See the full description on the dataset page: https://huggingface.co/datasets/forFrank/aptos_develop.APT_E3_v1AptosSC
Uploaded dataset
Developed by: SoumilB7
Organization: Armur
Project: Aptos Smart Contract Audit
License: apache-2.0
cluebbers__Llama-3.1-8B-paraphrase-type-generation-apty-ipo-details
Dataset Card for Evaluation run of cluebbers/Llama-3.1-8B-paraphrase-type-generation-apty-ipo
Dataset automatically created during the evaluation run of model cluebbers/Llama-3.1-8B-paraphrase-type-generation-apty-ipo
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cluebbers__Llama-3.1-8B-paraphrase-type-generation-apty-ipo-details.ap_test
