datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
to-tool-call-datasets
🛠️ To-Tool-Call Datasets
A unified Qwen3-style tool-call corpus for SFT, GRPO, and agent training
To-Tool-Call Datasets is a curated mirror of public tool-call and function-calling corpora, re-serialized into one training-ready messages JSONL convention.
Quick Start ·
At a Glance ·
Format ·
Sources ·
Training Notes
[!IMPORTANT]
This repository is a format-harmonization layer, not a new claim of ownership over the… See the full description on the dataset page: https://huggingface.co/datasets/zhangdw/to-tool-call-datasets.ToT
Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
ToT is a dataset designed to assess the temporal reasoning capabilities of AI models. It comprises two key sections:
ToT-semantic: Measuring the semantics and logic of time understanding.
ToT-arithmetic: Measuring the ability to carry out time arithmetic operations.
Dataset Usage
Downloading the Data
The dataset is divided into three subsets:
ToT-semantic: Measuring the semantics and logic… See the full description on the dataset page: https://huggingface.co/datasets/baharef/ToT.Genshin-Impact-SFTMedprompt-MedMCQA-ToT
Medprompt-MedMCQA-ToT
Dataset Summary
Medprompt-MedMCQA-ToT is a retrieval-augmented database designed to enhance contextual reasoning in multiple-choice medical question answering (MCQA). The dataset follows a Tree-of-Thoughts (ToT) reasoning format, where multiple independent reasoning paths are explored collaboratively before arriving at the correct answer. This structured… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Medprompt-MedMCQA-ToT.ToT
Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
ToT is a dataset designed to assess the temporal reasoning capabilities of AI models. It comprises two key sections:
ToT-semantic: Measuring the semantics and logic of time understanding.
ToT-arithmetic: Measuring the ability to carry out time arithmetic operations.
Dataset Usage
Downloading the Data
The dataset is divided into three subsets:
ToT-semantic: Measuring the semantics… See the full description on the dataset page: https://huggingface.co/datasets/HiXT0/ToT.Medprompt-MedQA-ToT
Medprompt-MedQA-ToT
Dataset Summary
Medprompt-MedQA-ToT is a retrieval-augmented database designed to enhance contextual reasoning in multiple-choice medical question answering (MCQA). The dataset follows a Tree-of-Thoughts (ToT) reasoning format, where multiple independent reasoning paths are explored collaboratively before arriving at the correct answer. This structured approach… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Medprompt-MedQA-ToT.Turkish-Dialectical-Reasoning-Dataset-Sokrates-ToT
Turkish Dialectical Reasoning Dataset (Sokrates-ToT)
The Turkish Dialectical Reasoning Dataset (Sokrates-ToT) is a collection structured in a Tree-of-Thought (ToT) format, based on a multi-persona and dialectical reasoning framework.Inspired by Socrates' method of dialogue, it facilitates deep analysis of complex and multidimensional issues by having AI personas with different expertise interact and ultimately reach a final synthesis.
Purpose of the Dataset
This… See the full description on the dataset page: https://huggingface.co/datasets/yusufbaykaloglu/Turkish-Dialectical-Reasoning-Dataset-Sokrates-ToT.cleand_moremilk_ToT-Biology元データ: https://huggingface.co/datasets/moremilk/ToT-Biology
データ件数: 5,752
平均トークン数: 675
最大トークン数: 1,105
合計トークン数: 3,881,334
ファイル形式: JSONL
ファイル分割数: 1
合計ファイルサイズ: 19.3 MB
加工内容:
長文フィルタリング: トークナイズ処理の負荷を軽減するため、事前に文字列が極端に長い行を除外します。
question 列: 6,000文字を超える行を除外。
metadata 列: 80,000文字を超える行を除外。
metadata フィールドの展開:
metadata 列に含まれるJSON形式のデータから reasoning と difficulty の値を抽出します。
reasoning は thought という新しい列に格納します。
difficulty は difficulty という新しい列に格納します。
処理後、元の metadata 列は削除されます。
繰り返し表現の除去:
thought… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/cleand_moremilk_ToT-Biology.tot-cwq-plan-sft-outputs34-rule-full-pw4-expand-labels-v2
ToT CWQ Plan SFT - outputs34_rule_full_pw4_expand_labels_v2
Merged SFT output from local run outputs34_rule_full_pw4_expand_labels_v2.
Version ID
local output dir: tot/sft/outputs34_rule_full_pw4_expand_labels_v2
file: cwq_train_plan.no_mid.jsonl
dataset: CWQ
grouping backend: TOT_REL_GROUPING_BACKEND=rules
parallel workers: 4
strict expand parity: enabled
nested expand labels: enabled
Main difference from earlier runs
This version renders nested Expand… See the full description on the dataset page: https://huggingface.co/datasets/YF0808/tot-cwq-plan-sft-outputs34-rule-full-pw4-expand-labels-v2.total_synthesis_of_new_drugs
Total Synthesis of New Drugs
This dataset contains basic information and raw synthesis route data for 60 drugs, along with 600 multimodal reasoning supervision (CoT) fine-tuning data.
The experimental and computational work in this dataset run on the Huawei Cloud AI Compute Service. We appreciate the stable compute supply from this platform.
新药化学全合成路线
该数据集包含 60 种药物的基本信息与合成路线原始数据,以及 600 条多模态思维监督微调数据。
本数据集的实验与计算工作依托于华为昇腾AI云服务平台完成,特此对其提供的稳定算力支持表示感谢。
Turkish-University-ToT-Dataset
Turkish University Tree-of-Thought Dataset
This dataset contains a multi-person Tree of Thought (ToT) dataset related to Turkish university regulations, designed for advanced question answering and reasoning tasks.
This study examines a multi-persona approach supported by a large language model (LLM) to analyze the complex regulatory structures of Turkish higher education institutions. The proposed Multi-Perspective Thought Tree framework integrates six expert perspectives—YÖK… See the full description on the dataset page: https://huggingface.co/datasets/yusufbaykaloglu/Turkish-University-ToT-Dataset.KannadaPromptBench
KannadaPromptBench
A benchmark dataset for evaluating prompt strategy sensitivity in Kannada, a low-resource Dravidian language.
Dataset Summary
Language: Kannada (kn)
Tasks: Sentiment Analysis (100), Question Answering (75), Summarization (50)
Total: 225 culturally grounded samples
Inter-annotator agreement: Cohen's κ > 0.80
Dataset Structure
Each sample contains: id, task, input_text, label, difficulty, domain.
Citation
Please cite if you use… See the full description on the dataset page: https://huggingface.co/datasets/totormac/KannadaPromptBench.
