datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Tachibana4-DeepSeek-V4-ProClick here to support our open-source dataset and model releases - help us speed up our release schedule!
Tachibana 4 is an agentic coding dataset, testing the limits of DeepSeek-V4-Pro's coding skills:
Questions prioritize real-world, challenging agentic coding tasks across a variety of programming languages and topics. Synthethic prompts utilize a variety of personas, experience levels, and styles of communication to maximize real-world flexibility and usability.
Areas of focus include… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Tachibana4-DeepSeek-V4-Pro.Chizuru_Tachibana
Chizuru Tachibana from Nande Koko ni Sensei ga!?
Trained with anime (full-final-pruned) model
Works the best with ALL, MIDD, OUTD, and OUTALL LoRA weight blocks, and with 0.7+ weights.
TachibanaTachibana is a dataset containing code-instruct data.
The 2024-09-27 version contains:
104k rows of synthetic chat responses generated using Llama 3.1 405b Instruct.
60.6k Magicoder prompts from ise-uiuc/Magicoder-Evol-Instruct-110K
43.4k Glaive-code-assistant prompts from glaiveai/glaive-code-assistant
This dataset contains synthetically generated data and has not been subject to manual review.
Tachibana2-DeepSeek-R1-PREVIEWThis is a preview of the full Tachibana 2 high-difficulty code-reasoning dataset, containing the first ~6k rows. All responses generated by deepseek-ai/DeepSeek-R1.
The full dataset will be released for everyone once it's ready!
This dataset contains:
6k high-difficulty synthetic code-reasoning prompts created by Llama 3.1 405b Instruct, with an emphasis on task complexity and technical skill.
Responses demonstrate the reasoning capabilities of DeepSeek's 685b parameter R1 reasoning model.… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Tachibana2-DeepSeek-R1-PREVIEW.deepseek-v4-pro-tachibana4Click here to support our open-source dataset and model releases - help us speed up our release schedule!
Tachibana 4 is an agentic coding dataset, testing the limits of DeepSeek-V4-Pro's coding skills:
Questions prioritize real-world, challenging agentic coding tasks across a variety of programming languages and topics. Synthethic prompts utilize a variety of personas, experience levels, and styles of communication to maximize real-world flexibility and usability.
Areas of focus include… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/deepseek-v4-pro-tachibana4.Tachibana2-DeepSeek-R1Click here to support our open-source dataset and model releases!
Tachibana2-DeepSeek-R1 is a code-reasoning dataset, testing the limits of DeepSeek R1's coding skills!
This dataset contains:
27.2k synthetically generated code-reasoning prompts. All responses are generated using DeepSeek R1.
Synthetic prompts are generated using Llama 3.1 405b Instruct, based on the original sequelbox/Tachibana dataset with increased task complexity.
Responses demonstrate the code-reasoning capabilities of… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Tachibana2-DeepSeek-R1.Tachibana4-DeepSeek-V4-Pro-PREVIEWClick here to support our open-source dataset and model releases - help us speed up our release schedule!
This is an early sneak preview of Tachibana 4, containing the first 1.2k rows!
Tachibana 4 is an upcoming agentic coding dataset, generated by DeepSeek-V4-Pro:
Questions prioritize real-world, challenging agentic coding tasks across a variety of programming languages and topics.
Areas of focus include back-end and front-end development, systems programming, distributed systems… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Tachibana4-DeepSeek-V4-Pro-PREVIEW.ayaka_tachibana_imocho
Dataset of Ayaka Tachibana/橘彩花 (Recently, My Sister Is Unusual)
This is the dataset of Ayaka Tachibana/橘彩花 (Recently, My Sister Is Unusual), containing 189 images and their tags.
The core tags of this character are short_hair, brown_hair, brown_eyes, which are pruned in this dataset.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
List of Packages
Name
Images… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/ayaka_tachibana_imocho.tachibana_sara_citrus
Dataset of Tachibana Sara
This is the dataset of Tachibana Sara, containing 69 images and their tags.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
Name
Images
Download
Description
raw
69
Download
Raw data with meta information.
raw-stage3
150
Download
3-stage cropped raw data with meta information.
raw-stage3-eyes
181
Download
3-stage cropped (with eye-focus) raw… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/tachibana_sara_citrus.Tachibana-QVQTachibana-QVQ is a dataset containing code-reasoning and code-instruct responses across a wide variety of programming tasks.
This dataset contains:
103k prompts from sequelbox/Tachibana, with all responses generated by Qwen/QVQ-72B-Preview.
Responses demonstrate QVQ's code-reasoning ability and general code capabilities.
Responses have not been filtered or edited at all: some responses will contain infinite thought loops, incomplete answers, inaccurate responses, or other identified or… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Tachibana-QVQ.tachibana_hibiki_senkizesshousymphogear
Dataset of Tachibana Hibiki
This is the dataset of Tachibana Hibiki, containing 300 images and their tags.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
Name
Images
Download
Description
raw
300
Download
Raw data with meta information.
raw-stage3
691
Download
3-stage cropped raw data with meta information.
384x512
300
Download
384x512 aligned dataset.
512x512
300… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/tachibana_hibiki_senkizesshousymphogear.tachibana_arisu_theidolmastercinderellagirlsu149
Dataset of Tachibana Arisu
This is the dataset of Tachibana Arisu, containing 200 images and their tags.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
Name
Images
Download
Description
raw
200
Download
Raw data with meta information.
raw-stage3
486
Download
3-stage cropped raw data with meta information.
384x512
200
Download
384x512 aligned dataset.
512x512
200… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/tachibana_arisu_theidolmastercinderellagirlsu149.tachibana_nina_citrus
Dataset of Tachibana Nina
This is the dataset of Tachibana Nina, containing 44 images and their tags.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
Name
Images
Download
Description
raw
44
Download
Raw data with meta information.
raw-stage3
102
Download
3-stage cropped raw data with meta information.
raw-stage3-eyes
133
Download
3-stage cropped (with eye-focus) raw… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/tachibana_nina_citrus.tachibana_coding_10kTachibana3-Part2-DeepSeek-V3.2Click here to support our open-source dataset and model releases!
Tachibana3-Part2-DeepSeek-V3.2 is a dataset focused on high-difficulty code production tasks, testing the limits of DeepSeek V3.2's code-reasoning skills!
This dataset contains 9.3k high-difficulty code-production prompts:
Questions prioritize real-world, challenging coding tasks across a variety of programming languages and topics.
Areas of focus include back-end and front-end development, mobile, gamedev, cloud, QA, custom… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Tachibana3-Part2-DeepSeek-V3.2.Tachibana3-Part1-DeepSeek-V3.1-TerminusClick here to support our open-source dataset and model releases!
Tachibana3-Part1-DeepSeek-V3.1-Terminus is a dataset focused on high-difficulty code production tasks, testing the limits of DeepSeek V3.1 Terminus's code-reasoning skills!
This dataset contains 9.3k high-difficulty code-production prompts:
Questions prioritize real-world, challenging coding tasks across a variety of programming languages and topics.
Areas of focus include back-end and front-end development, mobile, gamedev… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Tachibana3-Part1-DeepSeek-V3.1-Terminus.Tachibana-QVQ-PREVIEWThis is a preview of the full Tachibana-QVQ code-instruct dataset, containing the first ~10k rows.
Get the full dataset now!
Prompts randomly selected from sequelbox/Tachibana, all responses generated by Qwen/QVQ-72B-Preview.
Dataset has not been reviewed for format or accuracy. Synthetic data is generated by a 'preview' edition of Qwen's QVQ 72b model.
Use as you will.
tachibana_alice_idolmastercinderellagirls
Dataset of tachibana_alice/橘ありす (THE iDOLM@STER: Cinderella Girls)
This is the dataset of tachibana_alice/橘ありす (THE iDOLM@STER: Cinderella Girls), containing 500 images and their tags.
The core tags of this character are brown_hair, long_hair, brown_eyes, bow, hair_bow, bangs, blue_bow, which are pruned in this dataset.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/tachibana_alice_idolmastercinderellagirls.Kanade_Tachibana_Videos_Captioned
sequelbox_Tachibana2-DeepSeek-R1-PREVIEW-Shuffled-ShareGPTimport json
from tqdm import tqdm
from datasets import load_dataset
import pandas as pd
# Example usage:
dataset = load_dataset("sequelbox/Tachibana2-DeepSeek-R1-PREVIEW")["train"]
dataset = dataset.shuffle(seed=42)
output_file = "./sequelbox_Tachibana2-DeepSeek-R1-PREVIEW-Shuffled-ShareGPT.parquet"
data = []
for item in tqdm(dataset):
if item["prompt"].strip() == "" or item["response"].strip() == "":
continue
data.append(
{
"conversations": [… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/sequelbox_Tachibana2-DeepSeek-R1-PREVIEW-Shuffled-ShareGPT.Tachibana_VoicelinesSylphynford_Tachibana_UnprocessedTachibana4-DeepSeek-V4-Pro-Sharded
Tachibana4-DeepSeek-V4-Pro-Sharded
Shuffled ~100 MB JSONL shards of sequelbox/Tachibana4-DeepSeek-V4-Pro (CSV rows re-serialized as {prompt, completion} objects).
Source: sequelbox/Tachibana4-DeepSeek-V4-Pro at revision 80304ea7aaa9dff66d3b674702d9534da7bdc7fe. Records are shuffled with seed 20260925 and split into ~100 MB parts; see MANIFEST.json for per-part hashes and record counts. Storage preparation only: no filtering or content change beyond the re-serialization noted… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/Tachibana4-DeepSeek-V4-Pro-Sharded.
