tachibana
Tachibana4-DeepSeek-V4-ProClick here to support our open-source dataset and model releases - help us speed up our release schedule!
Tachibana 4 is an agentic coding dataset, testing the limits of DeepSeek-V4-Pro's coding skills:
Questions prioritize real-world, challenging agentic coding tasks across a variety of programming languages and topics. Synthethic prompts utilize a variety of personas, experience levels, and styles of communication to maximize real-world flexibility and usability.
Areas of focus include… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Tachibana4-DeepSeek-V4-Pro.Chizuru_Tachibana
Chizuru Tachibana from Nande Koko ni Sensei ga!?
Trained with anime (full-final-pruned) model
Works the best with ALL, MIDD, OUTD, and OUTALL LoRA weight blocks, and with 0.7+ weights.
TachibanaTachibana is a dataset containing code-instruct data.
The 2024-09-27 version contains:
104k rows of synthetic chat responses generated using Llama 3.1 405b Instruct.
60.6k Magicoder prompts from ise-uiuc/Magicoder-Evol-Instruct-110K
43.4k Glaive-code-assistant prompts from glaiveai/glaive-code-assistant
This dataset contains synthetically generated data and has not been subject to manual review.
Tachibana2-DeepSeek-R1-PREVIEWThis is a preview of the full Tachibana 2 high-difficulty code-reasoning dataset, containing the first ~6k rows. All responses generated by deepseek-ai/DeepSeek-R1.
The full dataset will be released for everyone once it's ready!
This dataset contains:
6k high-difficulty synthetic code-reasoning prompts created by Llama 3.1 405b Instruct, with an emphasis on task complexity and technical skill.
Responses demonstrate the reasoning capabilities of DeepSeek's 685b parameter R1 reasoning model.… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Tachibana2-DeepSeek-R1-PREVIEW.deepseek-v4-pro-tachibana4Click here to support our open-source dataset and model releases - help us speed up our release schedule!
Tachibana 4 is an agentic coding dataset, testing the limits of DeepSeek-V4-Pro's coding skills:
Questions prioritize real-world, challenging agentic coding tasks across a variety of programming languages and topics. Synthethic prompts utilize a variety of personas, experience levels, and styles of communication to maximize real-world flexibility and usability.
Areas of focus include… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/deepseek-v4-pro-tachibana4.Tachibana2-DeepSeek-R1Click here to support our open-source dataset and model releases!
Tachibana2-DeepSeek-R1 is a code-reasoning dataset, testing the limits of DeepSeek R1's coding skills!
This dataset contains:
27.2k synthetically generated code-reasoning prompts. All responses are generated using DeepSeek R1.
Synthetic prompts are generated using Llama 3.1 405b Instruct, based on the original sequelbox/Tachibana dataset with increased task complexity.
Responses demonstrate the code-reasoning capabilities of… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Tachibana2-DeepSeek-R1.
