datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
glaive_toolcall_enBorrowed from: https://huggingface.co/datasets/glaiveai/glaive-function-calling-v2
You can use it in LLaMA Factory by specifying dataset: glaive_toolcall_en.
glaive_toolcall_enBorrowed from: https://huggingface.co/datasets/glaiveai/glaive-function-calling-v2
You can use it in LLaMA Factory by specifying dataset: glaive_toolcall_en.
unified-toolcalls-canonical
Unified Tool-Calling Corpus — Canonicalized Output
Publish-ready conversion of two pinned Hugging Face dataset revisions into the single
schema defined in docs/unified_format.md, with repeated
records normalized by an explicit canonicalization rule and every surviving record
kept faithful to its source row.
Records in (source rows)
65,000
Records published (canonical survivors)
64,622
Duplicates collapsed
378 (343 duplicate groups)
Records mutated during… See the full description on the dataset page: https://huggingface.co/datasets/dongbobo/unified-toolcalls-canonical.glaive_toolcall_zhBorrowed from: https://huggingface.co/datasets/glaiveai/glaive-function-calling-v2
Translated by GPT-3.5.
You can use it in LLaMA Factory by specifying dataset: glaive_toolcall_zh.
tool-calls-single-reasoningmath-toolcall-tr-benchmark
math-toolcall-tr-benchmark
bilalabic/gemma_4_math-toolcall-tr_lora
LoRA adaptörünü temel Gemma-4 E4B modeliyle karşılaştıran benchmark sonuçları.
Bu depo yalnızca değerlendirme çıktılarını içerir. Eğitim veri seti ayrı olarak
bilalabic/math-toolcall-tr
adresinde yayımlanmaktadır.
Benchmark'lar
Benchmark
Örnek
Ölçülen davranış
Türkçe MMLU
250
Genel bilgi doğruluğu ve eğitim sonrası bilgi kaybı
Matematik Tool-Call
150
Araç seçimi, çekimserlik ve çıktı… See the full description on the dataset page: https://huggingface.co/datasets/bilalabic/math-toolcall-tr-benchmark.math-toolcall-tr
math-toolcall-tr
Türkçe matematik odaklı fonksiyon çağırma (tool calling) veri seti — 2.127 örnek,
ShareGPT formatında, düşünme adımları (<think>) dahil.
EN: A Turkish synthetic dataset for training LLMs to call math functions correctly
and present the results clearly. 2,127 ShareGPT-format conversations with reasoning
traces, covering 70 math topics and 8 tool-calling scenarios.
Ne öğretir
İki beceriyi birlikte hedefler:
Doğru araç seçimi ve parametre çıkarımı… See the full description on the dataset page: https://huggingface.co/datasets/bilalabic/math-toolcall-tr.ToolCalling-Refusal-DS1K
ToolCalling-Refusal-DS1K
ToolCalling-Refusal-DS1K is a synthetic tool-calling refusal dataset annotated by deepseek-v4-flash. Each example contains a user request, a set of available tool schemas, and a structured teacher annotation describing whether the request can be fulfilled, which tool should be used, which parameters are missing, or why no suitable tool exists.
Unlike datasets that only provide a final natural-language answer, this dataset exposes the intermediate… See the full description on the dataset page: https://huggingface.co/datasets/whichcy/ToolCalling-Refusal-DS1K.synthetic-toolcall-1synthetic-toolcall-1 is a synthetic dataset with a total of ~201 rows.
This dataset was generated using the following models:
Grok:
Fast
Perplexity.ai:
"Search"
ChatGPT:
Whatever is available through the website.
Gemini:
3.1 Flash-Lite
3.5 Flash
3.1 Pro
Deepseek:
"Instant"
"Expert"
This dataset follows the following format:
[
{"messages": [
{"role": "system", "content": "Example system prompt"},
{"role": "user", "content": "Example user prompt"}… See the full description on the dataset page: https://huggingface.co/datasets/takenusername32/synthetic-toolcall-1.
