datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
poc-mini-trade-game-dataset
Dataset Card for Mini Trade Game NPC Dataset
Dataset Summary
This dataset contains synthetic training examples for simulating NPC (Non-Player Character) merchant behavior in a trading game scenario. The dataset is designed to train language models to generate contextually appropriate trading responses based on item properties, relationship status, and player interactions.
All examples are in Traditional Chinese (zh-TW), with player inputs and NPC responses using… See the full description on the dataset page: https://huggingface.co/datasets/aotoki/poc-mini-trade-game-dataset.qwen9b-coop-mini-swe-agent
qwen9b-coop-mini-swe-agent
Two-agent cooperative coding trajectories generated by running
CooperBench in coop mode on
the CooperData task set, using
Qwen/Qwen3.5-9B as the model and mini_swe_agent_v2 as the agent framework.
Each pair runs two agents in parallel — one per feature — coordinating via Redis messaging and a shared git remote.
The matched solo version is at
CooperBench/qwen9b-solo-mini-swe-agent.
Same task corpus, same model, same agent — only the coordination differs… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-coop-mini-swe-agent.Meditation-miniset-v0.2
Synthetic Meditation Dataset v0.2 🧘♂️🧘♀️✨
Welcome to the Synthetic Meditation Dataset v0.2, a comprehensive collection of meditation guidance prompts designed to assist users in different stages of emotional wellbeing and mindfulness. This dataset aims to help developers train AI models to provide empathetic and personalized meditation guidance. The dataset focuses on inclusivity, personalization, and diverse meditation contexts to cater to a wide audience.
Overview… See the full description on the dataset page: https://huggingface.co/datasets/BuildaByte/Meditation-miniset-v0.2.Raiden-Mini-DeepSeek-V3.2-SpecialeClick here to support our open-source dataset and model releases!
Raiden-Mini-DeepSeek-V3.2.Speciale is a dataset containing creative-reasoning and analytic-reasoning responses, testing the limits of DeepSeek-V3.2.Speciale's reasoning skills!
This dataset contains:
a default subset of ~8k 'creative_content' and 'analytical_reasoning' prompts from sequelbox/Raiden-DeepSeek-R1, with all responses generated by DeepSeek V3.2 Speciale.
provides an unfiltered look into the reasoning skills of… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Raiden-Mini-DeepSeek-V3.2-Speciale.qwen9b-solo-mini-swe-agent
qwen9b-solo-mini-swe-agent
Single-agent coding trajectories generated by running
CooperBench in solo mode on
the CooperData task set, using
Qwen/Qwen3.5-9B as the model and mini_swe_agent_v2 as the agent framework.
One agent implements both features in each task.
The matched coop version is at
CooperBench/qwen9b-coop-mini-swe-agent.
Same task corpus, same model, same agent — only the coordination differs, so
together they isolate the cooperation deficit.
At a glance… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-solo-mini-swe-agent.Math-IIO-68K-Mini
Mathematics Dataset for AI Model Training
This dataset contains 68,000 rows of mathematical questions and their corresponding solutions. It is designed for training AI models capable of solving mathematical problems or providing step-by-step explanations for a variety of mathematical concepts. The dataset is structured into three columns: input, instruction, and output.
Dataset Overview
Input: A mathematical question or problem statement (e.g., arithmetic, algebra… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Math-IIO-68K-Mini.turkish-chat-normalization-mini
Turkish Chat Normalization Mini
turkish-chat-normalization-mini is a web-derived and rule-degraded Turkish text normalization dataset designed for rewriting noisy, informal, unpunctuated, or diacritics-missing Turkish text into cleaner and more readable Turkish.
The dataset does not contain private user messages, chat logs, social media comments, complaint records, or scraped personal conversations. Source sentences are collected from open Turkish web resources, while the input… See the full description on the dataset page: https://huggingface.co/datasets/yagmurtuncer/turkish-chat-normalization-mini.minicrit-training-12k
📝 Read the full blog post: MiniCrit: Adversarial AI Validation for Financial Decision-Making
MiniCrit Training Dataset: 12,132 Trading Rationale-Critique Pairs
A comprehensive dataset of realistic institutional trading rationales paired with adversarial critiques, designed to train AI models that can validate trading signals and reduce false positives.
Dataset Summary
Size: 12,132 unique rationale-critique pairsFormat: CSV (2 columns: rationale, rebuttal)License:… See the full description on the dataset page: https://huggingface.co/datasets/wmaousley/minicrit-training-12k.fineweb-ultra-mini-pro
Dataset Card for Fineweb Ultra Mini
Fineweb Ultra Mini is a dataset derived from the original Fineweb dataset made by huggingface (see here: https://huggingface.co/datasets/HuggingFaceFW/fineweb).
The dataset focuses on extracting high quality data from the Fineweb dataset, from the 1-0.5% range. If you would like more data, though slightly sacrificing quality check out fineweb ultra mini, which focuses on the 2-3% of high quality data originally found in fineweb.… See the full description on the dataset page: https://huggingface.co/datasets/motionlabs/fineweb-ultra-mini-pro.Arabic_prompts_Mini_175
Arabic Prompts Dataset
Overview
The Arabic Prompts Dataset is a comprehensive collection of prompts designed to facilitate research and development in natural language processing (NLP), machine learning, and artificial intelligence, particularly focusing on Arabic language applications. This dataset includes a diverse range of topics and questions across various fields such as literature, science, technology, and culture, making it an invaluable resource for training models… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/Arabic_prompts_Mini_175.autonomous-driving-minimal-harm-gradient-pathfinding-v0.1
What this dataset tests
Whether a system can navigatea minimal-harm gradient through a driving scene.
The task is to identify the paththat minimizes total deformationacross all agents.
Required outputs
gradient vectors across actions
minimal harm path
deformation score
stability margin
Use case
Second layer of ethical navigation stack.
Transforms ethical cost fieldinto an actionable path.
Evaluation
Predictions must:
describe gradient… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-minimal-harm-gradient-pathfinding-v0.1.minimum_viable_articulation_v01.csv
Minimum Viable Articulation (MVA)
MVA measures a model’s ability to answer with the minimum viable output — no surplus explanation, no self-expansion, no tutorial behavior.
This dataset evaluates where models fail to stop:
Overcompletion
Hedging / padding
Teaching when not asked
Identity or stance leakage
Solving beyond scope
It exposes a behavior pattern where models confuse helpfulness with verbosity and treat extra tokens as value, rather than distortion.
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/minimum_viable_articulation_v01.csv.
