thinker
Apriel-1.6-15b-ThinkerDeepSeek-MOE-4X8B-R1-Distill-Llama-3.1-Deep-Thinker-Uncensored-24B-GGUFQwen3-Omni-30B-A3B-Instruct-NVFP4-W4A4-full-thinker-awqclipApriel-Nemotron-15b-ThinkerServiceNow-AI_Apriel-Nemotron-15b-Thinker-GGUFApriel-1.5-15b-ThinkerApriel-1.5-15b-Thinker-GGUFServiceNow-AI_Apriel-1.6-15b-Thinker-GGUF
Datasets
All datasets matching “thinker”tiny-aya-l2-thinker-multilingual-reasoning
Tiny Aya L2 Multilingual Reasoning (44 languages)
Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker.
Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English.
Data source
Prompts from AM-DeepSeek-R1-0528-Distilled
Thinking traces and outputs distilled from gpt-oss-120b
Translated with command-a-translate and DeepSeek-V3
Languages (44)
Language
Train… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/tiny-aya-l2-thinker-multilingual-reasoning.tiny-aya-l2-thinker-multilingual-reasoning
Tiny Aya L2 Multilingual Reasoning (44 languages)
Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker.
Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English.
Languages (44)
Language
Train
Test
Total
Amharic (am)
3,807
448
4,255
Arabic (ar)
22,968
2,538
25,506
Bulgarian (bg)
4,177
452
4,629
Bengali (bn)
3,803
422
4,225
Catalan (ca)
4,251
512
4,763
Czech… See the full description on the dataset page: https://huggingface.co/datasets/erenyeager-1/tiny-aya-l2-thinker-multilingual-reasoning.VAPO-Thinker-train36kM-Thinker-SFT-dataVAPO-Thinker-val1kCoSQA_PlusCoSQA+, a code search dataset pairing high-quality queries (reused from CoSQA) with multiple suitable codes.
We collect code candidates from diverse sources and form candidate pairs by pairing queries with these codes.
Utilizing the power of large language models (LLMs), we automate pair annotation, filtering, and code generation for queries without suitable matches.
related links:
arXiv:
2406.11589 CoSQA+: Enhancing Code Search Dataset with Matching Code (arxiv.org)
github:… See the full description on the dataset page: https://huggingface.co/datasets/thinkerhui/CoSQA_Plus.
