CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zhangdw /to-tool-call-datasets 🛠️ To-Tool-Call Datasets A unified Qwen3-style tool-call corpus for SFT, GRPO, and agent training &nbsp;&nbsp;&nbsp;&nbsp; To-Tool-Call Datasets is a curated mirror of public tool-call and function-calling corpora, re-serialized into one training-ready messages JSONL convention. Quick Start · At a Glance · Format · Sources · Training Notes [!IMPORTANT] This repository is a format-harmonization layer, not a new claim of ownership over the… See the full description on the dataset page: https://huggingface.co/datasets/zhangdw/to-tool-call-datasets.texttext-generation1K<n<10K3 likes348 downloads4mo agoHugging Face02baharef /ToT Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning ToT is a dataset designed to assess the temporal reasoning capabilities of AI models. It comprises two key sections: ToT-semantic: Measuring the semantics and logic of time understanding. ToT-arithmetic: Measuring the ability to carry out time arithmetic operations. Dataset Usage Downloading the Data The dataset is divided into three subsets: ToT-semantic: Measuring the semantics and logic… See the full description on the dataset page: https://huggingface.co/datasets/baharef/ToT.textquestion-answering10K<n<100K27 likes272 downloads2y agoHugging Face03totmalone /Genshin-Impact-SFTsentence-similarity1K<n<10K0 likes82 downloads18d agoHugging Face04HPAI-BSC /Medprompt-MedMCQA-ToT Medprompt-MedMCQA-ToT Dataset Summary Medprompt-MedMCQA-ToT is a retrieval-augmented database designed to enhance contextual reasoning in multiple-choice medical question answering (MCQA). The dataset follows a Tree-of-Thoughts (ToT) reasoning format, where multiple independent reasoning paths are explored collaboratively before arriving at the correct answer. This structured… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Medprompt-MedMCQA-ToT.question-answering100K<n<1M1 likes60 downloads1y agoHugging Face05HiXT0 /ToT Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning ToT is a dataset designed to assess the temporal reasoning capabilities of AI models. It comprises two key sections: ToT-semantic: Measuring the semantics and logic of time understanding. ToT-arithmetic: Measuring the ability to carry out time arithmetic operations. Dataset Usage Downloading the Data The dataset is divided into three subsets: ToT-semantic: Measuring the semantics… See the full description on the dataset page: https://huggingface.co/datasets/HiXT0/ToT.textquestion-answering10K<n<100K0 likes55 downloads14d agoHugging Face06HPAI-BSC /Medprompt-MedQA-ToT Medprompt-MedQA-ToT Dataset Summary Medprompt-MedQA-ToT is a retrieval-augmented database designed to enhance contextual reasoning in multiple-choice medical question answering (MCQA). The dataset follows a Tree-of-Thoughts (ToT) reasoning format, where multiple independent reasoning paths are explored collaboratively before arriving at the correct answer. This structured approach… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Medprompt-MedQA-ToT.question-answering10K<n<100K0 likes21 downloads1y agoHugging Face07yusufbaykaloglu /Turkish-Dialectical-Reasoning-Dataset-Sokrates-ToT Turkish Dialectical Reasoning Dataset (Sokrates-ToT) The Turkish Dialectical Reasoning Dataset (Sokrates-ToT) is a collection structured in a Tree-of-Thought (ToT) format, based on a multi-persona and dialectical reasoning framework.Inspired by Socrates' method of dialogue, it facilitates deep analysis of complex and multidimensional issues by having AI personas with different expertise interact and ultimately reach a final synthesis. Purpose of the Dataset This… See the full description on the dataset page: https://huggingface.co/datasets/yusufbaykaloglu/Turkish-Dialectical-Reasoning-Dataset-Sokrates-ToT.textquestion-answering1K<n<10K0 likes18 downloads1y agoHugging Face08LLMTeamAkiyama /cleand_moremilk_ToT-Biology元データ: https://huggingface.co/datasets/moremilk/ToT-Biology データ件数: 5,752 平均トークン数: 675 最大トークン数: 1,105 合計トークン数: 3,881,334 ファイル形式: JSONL ファイル分割数: 1 合計ファイルサイズ: 19.3 MB 加工内容: 長文フィルタリング: トークナイズ処理の負荷を軽減するため、事前に文字列が極端に長い行を除外します。 question 列: 6,000文字を超える行を除外。 metadata 列: 80,000文字を超える行を除外。 metadata フィールドの展開: metadata 列に含まれるJSON形式のデータから reasoning と difficulty の値を抽出します。 reasoning は thought という新しい列に格納します。 difficulty は difficulty という新しい列に格納します。 処理後、元の metadata 列は削除されます。 繰り返し表現の除去: thought… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/cleand_moremilk_ToT-Biology.tabularquestion-answering1K<n<10K0 likes15 downloads1y agoHugging Face09YF0808 /tot-cwq-plan-sft-outputs34-rule-full-pw4-expand-labels-v2 ToT CWQ Plan SFT - outputs34_rule_full_pw4_expand_labels_v2 Merged SFT output from local run outputs34_rule_full_pw4_expand_labels_v2. Version ID local output dir: tot/sft/outputs34_rule_full_pw4_expand_labels_v2 file: cwq_train_plan.no_mid.jsonl dataset: CWQ grouping backend: TOT_REL_GROUPING_BACKEND=rules parallel workers: 4 strict expand parity: enabled nested expand labels: enabled Main difference from earlier runs This version renders nested Expand… See the full description on the dataset page: https://huggingface.co/datasets/YF0808/tot-cwq-plan-sft-outputs34-rule-full-pw4-expand-labels-v2.tabularquestion-answering100K<n<1M0 likes15 downloads5mo agoHugging Face10ISSE-CQU /total_synthesis_of_new_drugs Total Synthesis of New Drugs This dataset contains basic information and raw synthesis route data for 60 drugs, along with 600 multimodal reasoning supervision (CoT) fine-tuning data. The experimental and computational work in this dataset run on the Huawei Cloud AI Compute Service. We appreciate the stable compute supply from this platform. 新药化学全合成路线 该数据集包含 60 种药物的基本信息与合成路线原始数据,以及 600 条多模态思维监督微调数据。 本数据集的实验与计算工作依托于华为昇腾AI云服务平台完成,特此对其提供的稳定算力支持表示感谢。 imagetext-generationn<1K0 likes11 downloads8mo agoHugging Face11yusufbaykaloglu /Turkish-University-ToT-Dataset Turkish University Tree-of-Thought Dataset This dataset contains a multi-person Tree of Thought (ToT) dataset related to Turkish university regulations, designed for advanced question answering and reasoning tasks. This study examines a multi-persona approach supported by a large language model (LLM) to analyze the complex regulatory structures of Turkish higher education institutions. The proposed Multi-Perspective Thought Tree framework integrates six expert perspectives—YÖK… See the full description on the dataset page: https://huggingface.co/datasets/yusufbaykaloglu/Turkish-University-ToT-Dataset.textquestion-answering1K<n<10K1 likes9 downloads1y agoHugging Face12totormac /KannadaPromptBench KannadaPromptBench A benchmark dataset for evaluating prompt strategy sensitivity in Kannada, a low-resource Dravidian language. Dataset Summary Language: Kannada (kn) Tasks: Sentiment Analysis (100), Question Answering (75), Summarization (50) Total: 225 culturally grounded samples Inter-annotator agreement: Cohen's κ > 0.80 Dataset Structure Each sample contains: id, task, input_text, label, difficulty, domain. Citation Please cite if you use… See the full description on the dataset page: https://huggingface.co/datasets/totormac/KannadaPromptBench.texttext-classificationn<1K0 likes7 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.