datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
APTER-Rubrics
APTER Expert-Grounded Query-Level Rubrics
This dataset contains the query-level Rubrics from APTER: Adaptive
Post-Training with Expert-Grounded Rubrics for mathematical reasoning and
medical question answering. Each Rubric instantiates an expert-defined
criterion into a fine-grained requirement for a specific query.
Technical report: arXiv:2608.14212
Project repository: AntDT-APTER/APTER
Data
Split
Queries
Rubric items
math
16,755
67,817
medical
28… See the full description on the dataset page: https://huggingface.co/datasets/AntDT-APTER/APTER-Rubrics.AptMQL-Bench
AptMQL-Bench
📄 Paper: AptMQL-Bench: From Text-to-SQL to Text-to-MQL via Access-Pattern Schema Design and Data-Preserving Migration · arXiv: coming soon
AptMQL-Bench is a benchmark for text-to-MQL — the task of translating human-readable requests into executable MongoDB Query Language (MQL) aggregation pipelines. It contains 21 document-oriented databases, 3,181 natural-language requests, and their associated gold MQL queries.
Most existing text-to-MQL resources are… See the full description on the dataset page: https://huggingface.co/datasets/giahy2507/AptMQL-Bench.japanese-style-contrast-dataset
LLMの文体差理解向上のための日本語データセット
本データセットは、新居浜工業高等専門学校と株式会社APTOの共同研究により構築された、日本語LLMの文体差(話し言葉・書き言葉)理解性能の向上を目的とした学習用データセットです。
新居浜高専 電子工学専攻の高専生が研究の企画立案からデータセット構築、追加学習および評価実験までを主導し、指導教員の助言と株式会社APTOからの計算環境の提供・技術的指導のもとで作成されました。
本研究の成果は言語処理学会第32回年次大会(NLP2026)にて発表されます。
研究の背景
大規模言語モデル(LLM)は日本語の質問応答や文章生成において高い性能を示す一方で、話し言葉と書き言葉といった文体差に対する理解には課題が残されております。特に話し言葉では、口語的表現や省略を含むため理解性能が不安定になりやすく、文体差への対応は日本語LLMにおける重要な課題となっています。… See the full description on the dataset page: https://huggingface.co/datasets/APTO-001/japanese-style-contrast-dataset.apticraft-quantum-llm
AptiCraft — Custom AI/LLM Module (Quantum) · README
A practical guide to build, ship, and protect your Quantum Custom AI + LLM using a versioned Hugging Face dataset.
1) What this repo contains
quantum.json: user‑friendly Q&A + deep sections (algorithms, hardware, reliability, math) with routing hints.
Images: external URLs only. The app can include or suppress images per request.
Versioning: tag releases (e.g., v1, v2) and pin clients to a tag for stability.… See the full description on the dataset page: https://huggingface.co/datasets/Afsana721/apticraft-quantum-llm.testaptitude-qa-dataset
Aptitude QA Dataset
This dataset contains foundational quantitative and logical reasoning questions used for fine-tuning compact language models on structured mathematical problem-solving tasks. It uses a structured chat format designed for training placement-focused conversational agents.
Dataset Structure
The data points are organized into standard conversational messages, establishing a robust training structure for Hugging Face SFTTrainer pipelines.
JSON… See the full description on the dataset page: https://huggingface.co/datasets/Prathamesh25/aptitude-qa-dataset.unseen-aptitude-qa-dataset
Unseen Aptitude QA Dataset
This dataset contains categorized quantitative and logical aptitude questions explicitly structured for campus placement preparation (e.g., TCS, Wipro, Infosys). It is formatted using the standard ChatML / OpenAI Messages schema, making it natively compatible with fine-tuning models like SmolLM2-1.7B.
Dataset Structure
Each data sample contains a messages array featuring a structured system persona, metadata-enriched user questions, and detailed… See the full description on the dataset page: https://huggingface.co/datasets/Prathamesh25/unseen-aptitude-qa-dataset.
