datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
XIH-Bench
XIH-Bench
Benchmark for the paper "Language Shapes Instruction Hierarchy Compliance in Multilingual LLMs".
Instruction hierarchy (IH) requires models to prioritize instructions by source, so that
higher-priority instructions override lower-priority ones. XIH-Bench evaluates IH under both
same-language and cross-language conflicts across six languages, four domains and three
hierarchy settings.
78,894 evaluation instances
Paper: https://arxiv.org/abs/2607.23545
Code:… See the full description on the dataset page: https://huggingface.co/datasets/g1moon/XIH-Bench.ML4SE23_G1_MBPP-SCoT
ML4SE23_G1_MBPP-SCoT
MBPP enhanced dataset with Structured-Chain-of-Thought
ML4SE23_G1_MBPP-SCoTMBPP enhanced dataset with Structured-Chain-of-Thought
G15
G15 - v0.1
G15-v0.1 is a part of G-series datasets which are pre-formatted in ChatML template.
These datasets are useful for quickly finetuning LLMs for better responses.
The G15-v0.1 is a combination of the following datasets:
OpenHermes-2.5
MetaMathQA (100k entries)
A section of our in-house dataset used to finetune Cerberus-v0.1
This dataset is to be used on smaller LLMs (1B - 7B) to increase their response quality.
For queries please reach out to us at hello@brahmai.in
ML4SE23_G1_MBCPP-SCoTMBCPP enhanced dataset with Structured-Chain-of-Thought
ML4SE23_G1_MBCPP-SCoT
ML4SE23_G1_MBCPP-SCoT
MBCPP enhanced dataset with Structured-Chain-of-Thought
ML4SE23_G1_HumanEval-SCoT
ML4SE23_G1_HumanEval-SCoT
HumanEval dataset enhanced with Structured-Chain-of-Thought
ML4SE23_G1_HumanEval-SCoTHumanEval dataset enhanced with Structured-Chain-of-Thought
ML4SE23_G1_EvolInstruct-SCoT-1k
ML4SE23_G1_EvolInstruct-SCoT-1k
EvolInstruct enhanced 1k entries dataset with Structured-Chain-of-Thought
ML4SE23_G1_EvolInstruct-SCoT-1kEvolInstruct enhanced 1k entries dataset with Structured-Chain-of-Thought
