datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
glaive-code-assistant-v3
Glaive-code-assistant-v3
Glaive-code-assistant-v3 is a dataset of ~1M code problems and solutions generated using Glaive’s synthetic data generation platform.
This is built on top of the previous version of the dataset that can be found here. This already includes v1 and v2 of the dataset.
To report any problems or suggestions in the data, join the Glaive discord
glaive-code-assistant
Glaive-code-assistant
Glaive-code-assistant is a dataset of ~140k code problems and solutions generated using Glaive’s synthetic data generation platform.
The data is intended to be used to make models act as code assistants, and so the data is structured in a QA format where the questions are worded similar to how real users will ask code related questions.
The data has ~60% python samples.
To report any problems or suggestions in the data, join the Glaive discord
glaive-code-assistant-v2
Glaive-code-assistant-v2
Glaive-code-assistant-v2 is a dataset of ~215k code problems and solutions generated using Glaive’s synthetic data generation platform.
This is built on top of the previous version of the dataset that can be found here
To report any problems or suggestions in the data, join the Glaive discord
swebench_verified_random_100_folders_a1_glaive_code_assistant_20260326_090119glaive-code-assistant-v3-sharegptglaiveai/glaive-code-assistant-v3 transformed to sharegpt format to easily train models using axolotl
glaive-code-assistant-sandboxes-traces-terminus-2code-chat-assistant-v1
Dataset Card for "code-chat-assistant-v1"
More Information needed
glaive-code-assistant-v3lilac-glaive-code-assistant
lilac/glaive-code-assistant
This dataset is a Lilac processed dataset. Original dataset: https://huggingface.co/datasets/glaiveai/glaive-code-assistant
To download the dataset to a local directory:
lilac download lilacai/lilac-glaive-code-assistant
or from python with:
ll.download("lilacai/lilac-glaive-code-assistant")
glm46-glaive-code-assistant-sandboxes-maxeps-131kglaive-code-assistant-v3glaive-code-assistant-v3tt633-technical-code-assistant-v1
TT633 Technical Code Assistant v1
This dataset is built for training the fresh custom TransformerTechnology V8.3 MDL Circle-Switch-Grid model as a small technical/code assistant.
Canonical training column: text.
Format:
Instruction: ...
Input:
...
Answer:
...
<END>
Primary sources:
Plaincode CNL rows from CircularBalls/plaincode-cnl-100k.
Small curated technical QA, code-generation, debugging, reasoning, and stop-discipline seed rows.
Optional local pack text if provided at… See the full description on the dataset page: https://huggingface.co/datasets/CircularBalls/tt633-technical-code-assistant-v1.glaive_code_assistant_140K
Dataset Card for "glaive_code_assistant_140K"
More Information needed
glaive-code-assistant
Dataset Card for "glaive-code-assistant"
More Information needed
glaive-code-assistant-v1-sharegpt-format_split_13split-glaive-code-assistant-v3glaive-code-assistant
Glaive Code Assistant
Glaive Code Assistant dataset formatted for training assistant models with the following prompt template:
<s>[INST] {question} [/INST] {answer} </s>
Trained model can be prompted in Llama style:
<s>[INST] {{ user_msg }} [/INST]
glaive-code-assistant-sandboxes_glm_4.7_traces_jupiterglaive-code-assistant-v1-sharegpt-format_split_14glaive-code-assistant-v3
Glaive-code-assistant-v2
Glaive-code-assistant-v2 is a dataset of ~1M code problems and solutions generated using Glaive’s synthetic data generation platform.
This is built on top of the previous version of the dataset that can be found here. This already includes v1 and v2 of the dataset.
To report any problems or suggestions in the data, join the Glaive discord
glaive-code-assistant-sandboxesPython_code_assistant_with_promptFormatted with a prompt template.
Modified from this dataset https://huggingface.co/datasets/Nan-Do/reason_code-search-net-python
glaive-code-assistant-v1-sharegpt-format_split_12glaive-code-assistant-v2
Glaive-code-assistant-v2
Glaive-code-assistant-v2 is a dataset of ~215k code problems and solutions generated using Glaive’s synthetic data generation platform.
This is built on top of the previous version of the dataset that can be found here
To report any problems or suggestions in the data, join the Glaive discord
glaive-code-assistant-v2-100kglaive-code-assistant-v1-sharegpt-format_split_18code-and-cybersecurity-assistantdev_set_v2_a1_glaive_code_assistant_20260326_065039Customizable-Code-Assistant-Data
Dataset Card for "Customizable-Code-Assistant-Data"
Dataset Summary
This dataset contains is a dummy Version of the Customizable Code Assistant Dataset.
Supported Tasks and Leaderboards
Customizable Code Assistant is a dataset for code completion. The task is to predict the next token in a code snippet. The dataset is designed to be customizable, so that it can be used for different programming languages and different code completion tasks.
[More Information… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/Customizable-Code-Assistant-Data.
