datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
toy-multistep-v2-nn_20-na_10-nab_40-p_90-seed_0multistep-llama3-3b-instructbeir_fiqa_test_multistep_rewritten_queriesagentic_multistep_Qwen3-32B_multistep_rewritten_queriescomplex-queries-with-multi-step-reasoning-with-reasoningtoy-multistep-v3-wrl0agentic_multistep_Qwen3-0.6B_ans_rel_Llama-3.3-70B-Instructtoy-multistep-nn_50-na_5-nab_50-seed_0toy-multistep-v2-nn_20-na_10-nab_40-testgame_stage2_zjhhhh__qwen2.5_3B_Instruct_multi_stage2_seed_555134_eta_1e4_step_382_finaltoy-multistep-nn_10-na_5-nab_10-seed_0game_stage3_zjhhhh__qwen2.5_3B_Instruct_multi_stage3_seed_555134_eta_1e4_step_101game_multi_iter2_zjhhhh__iter2_multi_adversary_step_301agentic_multistep_Qwen3-8B_ans_rel_Llama-3.1-8B-Instructmulti_step_reasoning_kazakh_context
🇰🇿 Multi-step Reasoning for Kazakh Context
A high-quality dataset designed for complex reasoning, question answering, and text generation tasks in the Kazakh language.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
10,981
Total Words (approx.)
6,652,450
Avg. Tokens per Sample
605
Word Count Distribution (Per Field)
The following table details the distribution of word counts across different… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/multi_step_reasoning_kazakh_context.agentic_multistep_Qwen3-32B_Selene-1-Llama-3.3-70Bmulti-step-routing-ecom
Multi-Step Routing E-Commerce
A synthetic benchmark for multi-step intent routing in e-commerce customer service. Each sample contains a natural-language customer instruction paired with an ordered chain of specialised agents needed to resolve it.
Dataset Stats
Content
Amount
Train samples
4,140
Test samples
1,030
Unique intents
37
Unique domains
13
Unique agents
60+
Routing steps
2 – 4
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/rescommons/multi-step-routing-ecom.toy-multistep-v4-3agentic_multistep_Qwen3-0.6Bmultistep-intermediary-embeddingsagentic_multistep_Qwen3-8B_ans_rel_Llama-3.3-70B-Instructagentic_multistep_Qwen3-4B_ans_rel_Llama-3.3-70B-Instructpinchbench-clawd-multi-step
PinchBench Clawd - Hirundo Format
Prepared from cptekur/pinchbench-clawd for Hirundo custom dataset loading.
Each source trajectory is expanded into one training row per assistant turn.
The question contains the prior user/assistant/tool context, and the answer
is the next assistant message including tool-call formatting.
Schema
system_prompt: Clawd system prompt with available tools.
question: Rendered context before the target assistant turn.
answer: The next… See the full description on the dataset page: https://huggingface.co/datasets/hirundo-io/pinchbench-clawd-multi-step.text2sql-grpo-sql-r1-multi-stepagentic_multistep_Qwen3-8Bagentic_multistep_Qwen3-4B_Selene-1-Llama-3.3-70Btoy-multistep-v2-nn_20-na_10-nab_40-seed_0agentic_multistep_Qwen3-4Bagentic_multistep_GPT-4.1-nano_ans_rel_Llama-3.1-8B-Instructtoy-multistep-v3-test
