datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
reasoning-math-advanced-1m
🧠 Reasoning Math Advanced 1M
📖 Dataset Summary
Reasoning Math Advanced 1M is a large-scale, synthetic dataset designed to enhance the reasoning capabilities of Large Language Models (LLMs). Comprising 1,000,000 unique samples, this dataset focuses on Math, Logic, and Common Sense reasoning tasks.
A unique feature of this dataset is its adaptive reasoning structure, where the presence of Chain-of-Thought (CoT) reasoning scales with difficulty. All reasoning traces are… See the full description on the dataset page: https://huggingface.co/datasets/naimulislam/reasoning-math-advanced-1m.jee-advanced-questions
JEE Advanced — Question Bank
A structured dataset of JEE Advanced examination questions with full
worked solutions and diagrams. JEE Advanced questions are more analytical
than JEE Main — many are subjective, integer, or numerical-answer type with
detailed multi-step solutions.
Subsets (PCM):
Physics — 50 questions
Chemistry — 21 questions
Mathematics — 48 questions
Structure
Organised into subsets by subject and splits (train / test):
mathematics/ physics/… See the full description on the dataset page: https://huggingface.co/datasets/eQOURSE/jee-advanced-questions.glm5.2-general-distill
Teacher-generated instruction/response pairs used to distill small, local student models
(the ADI / Advanced Data Intelligence series) from the frontier teacher glm-5.2.
How it was built
Teacher: glm-5.2 (served via Ollama Cloud as glm-5.2:cloud), queried with
thinking/reasoning disabled so every record is a single clean final answer.
Seed prompts: databricks/databricks-dolly-15k,
filtered to remove items that require an attached context passage — the closed_qa… See the full description on the dataset page: https://huggingface.co/datasets/AdvancedDataIntelligence/glm5.2-general-distill.jee-advanced-questions
JEE Advanced — Question Bank
A structured dataset of JEE Advanced examination questions with full
worked solutions and diagrams. JEE Advanced questions are more analytical
than JEE Main — many are subjective, integer, or numerical-answer type with
detailed multi-step solutions.
Subsets (PCM):
Physics — 50 questions
Chemistry — 21 questions
Mathematics — 48 questions
Structure
Organised into subsets by subject and splits (train / test):
mathematics/ physics/… See the full description on the dataset page: https://huggingface.co/datasets/Grass-G/jee-advanced-questions.reasoning_code_advanced_1m
💻 Reasoning Code Advanced 1M
📖 Dataset Summary
Reasoning Code Advanced 1M is a massive-scale, synthetic dataset specifically engineered to improve the algorithmic reasoning and problem-solving capabilities of Large Language Models (LLMs). Featuring 1,000,000 unique coding samples, this dataset spans multiple programming languages (Python, JS, C++, etc.) and focuses on logic-heavy development tasks.
A key feature of this dataset is its Adaptive Reasoning Architecture.… See the full description on the dataset page: https://huggingface.co/datasets/naimulislam/reasoning_code_advanced_1m.advanced-quantum-algorithms
Neura Parse — Advanced Quantum Algorithms: Derivations, QSVT/Block-Encoding & Hamiltonian Simulation
A derivation- and resource-analyzed algorithms vertical spanning the canonical fault-tolerant canon (with full proofs, complexity, and worked traces) and the modern QSVT/block-encoding toolkit through Hamiltonian simulation, amplitude estimation, and quantum linear systems. Turns the general dataset's one-topic-per-algorithm summaries into line-by-line derivations, lower… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/advanced-quantum-algorithms.reasoning_conversations_advanced_1m
💻 Reasoning Conversations Advanced 1M
📖 Dataset Summary
Reasoning Conversations Advanced 1M is a massive-scale, synthetic dataset specifically engineered to improve the algorithmic reasoning and problem-solving capabilities of Large Language Models (LLMs). Featuring 1,000,000 unique coding samples, this dataset spans multiple programming languages (Python, JS, C++, etc.) and focuses on logic-heavy development tasks.
A key feature of this dataset is its Adaptive… See the full description on the dataset page: https://huggingface.co/datasets/naimulislam/reasoning_conversations_advanced_1m.marble-game-advanced-levels
Marble Game Advanced Levels
This dataset contains level data for a marble game, including descriptions, difficulty ratings, themes, and elements such as obstacles and checkpoints. It is used for training models to generate game levels automatically.
Files
data/train.json: Training data
data/valid.json: Validation data
data/test.json: Test data
Structure
Each entry in the dataset contains the following fields:
level_id: Unique identifier for the level
name:… See the full description on the dataset page: https://huggingface.co/datasets/Oranblock/marble-game-advanced-levels.advanced_math_demotask2_advanced
Статистический анализ и визуализация
Описание
Данный датасет представляет результат статистического анализа текста книги "Harry Potter and the Philosopher's Stone", выполненного с использованием различных методов обработки естественного языка. В процессе анализа были проведены три ключевых типа анализа:
Анализ уникальности данных: Оценка доли уникальных слов в корпусе текста и вычисление коэффициента лексического разнообразия.
Анализ частоты POS-тегов: Изучение частоты… See the full description on the dataset page: https://huggingface.co/datasets/lamdary/task2_advanced.task2_advanced
Задача 2: Разметка датасета с HF Datasets: статистический анализ и визуализация
Разметка представленного датасета выполнена в рамках курса по компьютерной лингвистике НИУ ВШЭ (СПб).
ПРОВЕДЕННЫЙ АНАЛИЗ
Был пороведен статистический анализ текста, включающий:
Анализ уникальности данных
Доля уникальных слов: 0.0032
Частотный анализ биграмм и триграмм
Анализ распределений: Длины предложений, слов и n-грамм
Средняя длина предложения: 15.82
Стандартное отклонение длины… See the full description on the dataset page: https://huggingface.co/datasets/litlsun/task2_advanced.advanced-text-generation-instructions
Advanced Text Generation Instructions Dataset
This dataset contains carefully written English text samples designed for basic and intermediate text generation tasks.
Dataset Purpose
The dataset is intended for:
Testing text generation pipelines
Training simple language models
Benchmarking prompt-based generation
Educational and experimental use
Dataset Structure
Each row contains a single field:
text: A standalone English sentence related to artificial… See the full description on the dataset page: https://huggingface.co/datasets/yugi5/advanced-text-generation-instructions.Advanced_Dataset_SampleThis is a high-fidelity Direct Preference Optimization (DPO) dataset curated by OptiRefine. It is designed to train Large Language Models (LLMs) to act as helpful, honest, and thoughtful assistants across complex domains.
While our core datasets focus on code refactoring, this dataset provides preference trajectories for broader system architecture, computer science fundamentals, logic, and professional communication.
Curated by: OptiRefine
Language: English
License: Apache-2.0
Format: JSONL… See the full description on the dataset page: https://huggingface.co/datasets/OptiRefine-Official/Advanced_Dataset_Sample.nemotron-gym-multichallenge-advanced
laion/nemotron-gym-multichallenge-advanced
Harbor task-binary dataset (1,068 tasks) converted from nvidia/Nemotron-RL-Multichallenge-v1 [advanced]
(part of the nvidia/Nemotron-Post-Training-v3 collection).
Each row is a valid Harbor
task binary: columns path (str) and task_binary (gzip tar). Converted with the
OpenThoughts-Agent data.nemotron_gym framework.
Grading: Rubric LLM judge (all criteria); needs OPENAI_API_KEY at trial.
Advanced_Malicious_Dataset-ITA
MaxForce01 / Advanced_Malicious_Dataset-ITA — Dataset Card
Nome: MaxForce01/Advanced_Malicious_Dataset-ITA
Versione: 1.0
Lingua: Italiano
Tipologia: Prompt–risposta (test di robustezza / red‑team) — contenuti sensibili controllati
Data di creazione: 17/10/2025
Licenza: MIT.
1. Descrizione breve
Questo dataset è stato creato per la ricerca sulla robustezza dei modelli di linguaggio e per testare in sicurezza il comportamento dei modelli quando sottoposti a prompt… See the full description on the dataset page: https://huggingface.co/datasets/MaxForce01/Advanced_Malicious_Dataset-ITA.advanced-agent-workflows
Advanced Agent Workflows Dataset
Task–tools–output samples for training advanced AI agents.
