datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
glaive-code-assistant-v3
Glaive-code-assistant-v3
Glaive-code-assistant-v3 is a dataset of ~1M code problems and solutions generated using Glaive’s synthetic data generation platform.
This is built on top of the previous version of the dataset that can be found here. This already includes v1 and v2 of the dataset.
To report any problems or suggestions in the data, join the Glaive discord
glaive-code-assistant
Glaive-code-assistant
Glaive-code-assistant is a dataset of ~140k code problems and solutions generated using Glaive’s synthetic data generation platform.
The data is intended to be used to make models act as code assistants, and so the data is structured in a QA format where the questions are worded similar to how real users will ask code related questions.
The data has ~60% python samples.
To report any problems or suggestions in the data, join the Glaive discord
glaive-code-assistant-v2
Glaive-code-assistant-v2
Glaive-code-assistant-v2 is a dataset of ~215k code problems and solutions generated using Glaive’s synthetic data generation platform.
This is built on top of the previous version of the dataset that can be found here
To report any problems or suggestions in the data, join the Glaive discord
tt633-technical-code-assistant-v1
TT633 Technical Code Assistant v1
This dataset is built for training the fresh custom TransformerTechnology V8.3 MDL Circle-Switch-Grid model as a small technical/code assistant.
Canonical training column: text.
Format:
Instruction: ...
Input:
...
Answer:
...
<END>
Primary sources:
Plaincode CNL rows from CircularBalls/plaincode-cnl-100k.
Small curated technical QA, code-generation, debugging, reasoning, and stop-discipline seed rows.
Optional local pack text if provided at… See the full description on the dataset page: https://huggingface.co/datasets/CircularBalls/tt633-technical-code-assistant-v1.glaive-code-assistant-v2
Glaive-code-assistant-v2
Glaive-code-assistant-v2 is a dataset of ~215k code problems and solutions generated using Glaive’s synthetic data generation platform.
This is built on top of the previous version of the dataset that can be found here
To report any problems or suggestions in the data, join the Glaive discord
code-and-cybersecurity-assistantglaive_code_assistant_llama3.2glaive_code_assistanthttps://huggingface.co/datasets/glaiveai/glaive-code-assistant-v2
features: coding, single-turn, task
length: 215k
glaiveai_glaive-code-assistant-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
glaiveai_glaive-code-assistant-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
glaiveai/glaive-code-assistant with responses regenerated with gemini-2.0-flash-thinking-exp-1219.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was skipped.
model =… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/glaiveai_glaive-code-assistant-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.glaive_code_assistant_v3_resample_95kGlaive-code-assistant-ShareGPTglaiveai-code-assistant-dscodeglaive-code-assistant
Glaive-code-assistant
Glaive-code-assistant is a dataset of ~140k code problems and solutions generated using Glaive’s synthetic data generation platform.
The data is intended to be used to make models act as code assistants, and so the data is structured in a QA format where the questions are worded similar to how real users will ask code related questions.
The data has ~60% python samples.
To report any problems or suggestions in the data, join the Glaive discord
