datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
glaive_toolcall_enBorrowed from: https://huggingface.co/datasets/glaiveai/glaive-function-calling-v2
You can use it in LLaMA Factory by specifying dataset: glaive_toolcall_en.
glaive_toolcall_enBorrowed from: https://huggingface.co/datasets/glaiveai/glaive-function-calling-v2
You can use it in LLaMA Factory by specifying dataset: glaive_toolcall_en.
glaive_toolcall_zhBorrowed from: https://huggingface.co/datasets/glaiveai/glaive-function-calling-v2
Translated by GPT-3.5.
You can use it in LLaMA Factory by specifying dataset: glaive_toolcall_zh.
synthetic-toolcall-1synthetic-toolcall-1 is a synthetic dataset with a total of ~201 rows.
This dataset was generated using the following models:
Grok:
Fast
Perplexity.ai:
"Search"
ChatGPT:
Whatever is available through the website.
Gemini:
3.1 Flash-Lite
3.5 Flash
3.1 Pro
Deepseek:
"Instant"
"Expert"
This dataset follows the following format:
[
{"messages": [
{"role": "system", "content": "Example system prompt"},
{"role": "user", "content": "Example user prompt"}… See the full description on the dataset page: https://huggingface.co/datasets/takenusername32/synthetic-toolcall-1.
