datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ja-safety-sft-dataset
ja-safety-sft-dataset
日本語LLMの安全性チューニング用 SFT データセットのサンプル (500件) です。
A 500-item sample of the SFT dataset used to safety-tune APTO's Japanese LLMs. English version is provided below.
概要
株式会社APTOが大規模言語モデル(LLM)の安全性向上のために作成した約18,000件の日本語安全性学習データから、比率を維持して抽出したサンプルです。本サンプルでデータの構造と品質を確認できます。
関連モデル
本サンプルの元データを用いて以下のモデルを安全性チューニングしました。
APTO-001/Qwen3.5-27B-SafetyTuned (GGUF)
APTO-001/Qwen3.5-9B-Base-SafetyTuned (GGUF)
APTO-001/Qwen3.5-9B-SafetyTuned (GGUF)… See the full description on the dataset page: https://huggingface.co/datasets/APTO-001/ja-safety-sft-dataset.AptMQL-Bench
AptMQL-Bench
📄 Paper: AptMQL-Bench: From Text-to-SQL to Text-to-MQL via Access-Pattern Schema Design and Data-Preserving Migration · arXiv: coming soon
AptMQL-Bench is a benchmark for text-to-MQL — the task of translating human-readable requests into executable MongoDB Query Language (MQL) aggregation pipelines. It contains 21 document-oriented databases, 3,181 natural-language requests, and their associated gold MQL queries.
Most existing text-to-MQL resources are… See the full description on the dataset page: https://huggingface.co/datasets/giahy2507/AptMQL-Bench.counseling-aptness
APTNESS — Counseling Dialogues (processed)
英文共情策略咨询对话(APTNESS),LLaMA-Factory 格式;含 database / ed / extes 三个 split。
本仓库是 counselor_agent 项目中,经统一预处理器落地到 dataset/processed/ 的
APTNESS 数据集。所有记录采用统一 schema(case_id / source / lang /
messages[] + 各数据集特有的可选标注 / profile)。
规模
aptness_db: 9,659 dialogues, 39,792 turns (avg 4.12)
aptness_ed: 30 dialogues, 240 turns (avg 8.0)
aptness_extes: 10 dialogues, 238 turns (avg 23.8)
文件
文件
类型
条数
大小… See the full description on the dataset page: https://huggingface.co/datasets/XuShihao6715/counseling-aptness.unseen-aptitude-qa-dataset
Unseen Aptitude QA Dataset
This dataset contains categorized quantitative and logical aptitude questions explicitly structured for campus placement preparation (e.g., TCS, Wipro, Infosys). It is formatted using the standard ChatML / OpenAI Messages schema, making it natively compatible with fine-tuning models like SmolLM2-1.7B.
Dataset Structure
Each data sample contains a messages array featuring a structured system persona, metadata-enriched user questions, and detailed… See the full description on the dataset page: https://huggingface.co/datasets/Prathamesh25/unseen-aptitude-qa-dataset.
