datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
poc-mini-trade-game-dataset
Dataset Card for Mini Trade Game NPC Dataset
Dataset Summary
This dataset contains synthetic training examples for simulating NPC (Non-Player Character) merchant behavior in a trading game scenario. The dataset is designed to train language models to generate contextually appropriate trading responses based on item properties, relationship status, and player interactions.
All examples are in Traditional Chinese (zh-TW), with player inputs and NPC responses using… See the full description on the dataset page: https://huggingface.co/datasets/aotoki/poc-mini-trade-game-dataset.nemo-aot-o3-style-reasoning
Nemo AoT-O3 Style Reasoning Dataset
Canonical private curated dataset for the Kaggle nvidia-nemotron-model-reasoning-challenge.
Designed for SFT + lightweight RL fine-tuning of Nemotron-3-Nano-30B-A3B with assistant-only label masking.
Dataset Description
Each row is a (prompt, assistant_text) pair where the assistant response uses AoT-O3 style reasoning: structured <think> blocks with explicit State / Thinking / Next state transitions for complex types, and… See the full description on the dataset page: https://huggingface.co/datasets/dvyomkesh/nemo-aot-o3-style-reasoning.nemo-aot-o3-tong-qwen35-curated-v1
Nemotron AoT-O3 Tong Qwen35 Curated Dataset
Private archive for the two-stage AoT -> AoT-O3 dataset curation run for the Kaggle
nvidia-nemotron-model-reasoning-challenge.
Contents
fresh_excluding_pilot/long_types_aot_o3_excl_pilot.parquet: fresh unique long/search rows, excluding pilot row keys.
fresh_excluding_pilot/curated_mix_excl_pilot.parquet: fresh long rows plus concise short-type pass-through rows.
fresh_excluding_pilot/audit_report.json: strict audit… See the full description on the dataset page: https://huggingface.co/datasets/dvyomkesh/nemo-aot-o3-tong-qwen35-curated-v1.
