datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NExtLong-Instruct-dataset-Magpie-Llama-3.3-Pro-1M-v0.1
NExtLong: Toward Effective Long-Context Training without Long Documents
This repository contains the code ,models and datasets for our paper NExtLong: Toward Effective Long-Context Training without Long Documents.
[Github]
Quick Links
Overview
NExtLong Models
NExtLong Datasets
Datasets list
How to use NExtLong datasets
Bugs or Questions?
Overview
Large language models (LLMs) with extended context windows have made significant strides yet remain a… See the full description on the dataset page: https://huggingface.co/datasets/caskcsg/NExtLong-Instruct-dataset-Magpie-Llama-3.3-Pro-1M-v0.1.Magpie-Llama-3.3-Pro-1M-v0.1
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.3-Pro-1M-v0.1.Magpie-Llama-3.3-Pro-500K-Filtered
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.3-Pro-500K-Filtered.Codeforces-Python_Submissions_reformatted_deduped_llama3.3_nomic_fLlama-3.3-70B-Instruct-eval-logs-and-scoresLlama-3.3-Future-Code-Instructions
Llama 3.3 Future Code Instructions
Llama 3.3 Future Code Instructions is a large-scale instruction dataset synthesized with the Meta Llama 3.3 70B Instruct model.
The dataset was generated with the method called Magpie, where we prompted the model to generate instructions likely to be asked by the users.
In addition to the original prompt introduced by the authors, we conditioned the system prompt on what specific programming language the user has an interest in, gaining control… See the full description on the dataset page: https://huggingface.co/datasets/future-architect/Llama-3.3-Future-Code-Instructions.Codeforces-Python_Submissions_reformatted_deduped_llama3.3_x6_multiplied_with_blind_null_f_empjailbreak-llama-3.3-nemotron-49b-v1.5Llama-3.3-70B-Instruct-Infinity-Instruct-0625
Llama-3.3-70B-Instruct-Infinity-Instruct-0625
Dataset Description
This dataset is part of the LK-Speculators collection for speculative decoding research. It contains 660K prompt-response pairs designed for training draft models that are used alongside Llama-3.3-70B-Instruct as the target model. The dataset was created by generating responses to the prompts from Infinity-Instruct-0625 with meta-llama/Llama-3.3-70B-Instruct at temperature=1.
For more details on the… See the full description on the dataset page: https://huggingface.co/datasets/nebius/Llama-3.3-70B-Instruct-Infinity-Instruct-0625.Llama-3.3-70B-Inst-awq_SafeRLHF
Llama-3.3-70B-Inst-awq Responses for RefAlign Safety Alignment
This dataset contains responses generated for the paper Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data, which introduces the RefAlign alignment algorithm.
Code Repository: https://github.com/mzhaoshuai/RefAlign
This dataset specifically consists of responses generated by the casperhansen/llama-3.3-70b-instruct-awq model, given the prompts from the… See the full description on the dataset page: https://huggingface.co/datasets/mzhaoshuai/Llama-3.3-70B-Inst-awq_SafeRLHF.PerfectBlend-Regenerated-Llama-3.3-70B-Instructdetails_CHIH-HUNG__llama-2-13b-FINETUNE3_3.3w-r8-gate_up_down
Dataset Card for Evaluation run of CHIH-HUNG/llama-2-13b-FINETUNE3_3.3w-r8-gate_up_down
Dataset Summary
Dataset automatically created during the evaluation run of model CHIH-HUNG/llama-2-13b-FINETUNE3_3.3w-r8-gate_up_down on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_CHIH-HUNG__llama-2-13b-FINETUNE3_3.3w-r8-gate_up_down.multi-turn-aware-quantization-llama-3.3-rp-testI added role headers and tokens for each turn in the LLaMA 3 Instruct format. The purpose is to test whether formatted multi-turn data can improve multi-turn performance after quantization.
details_CHIH-HUNG__llama-2-13b-FINETUNE3_3.3w-r16-gate_up_down
Dataset Card for Evaluation run of CHIH-HUNG/llama-2-13b-FINETUNE3_3.3w-r16-gate_up_down
Dataset Summary
Dataset automatically created during the evaluation run of model CHIH-HUNG/llama-2-13b-FINETUNE3_3.3w-r16-gate_up_down on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_CHIH-HUNG__llama-2-13b-FINETUNE3_3.3w-r16-gate_up_down.Llama-3.3-70B-Instruct-evals
Dataset Card for Meta Evaluation Result Details for Llama-3.3-70B-Instruct
This dataset contains the results of the Meta evaluation result details for Llama-3.3-70B-Instruct. The dataset has been created from 12 evaluation tasks. The tasks are: human_eval, mmlu_pro, gpqa_diamond, ifeval__loose, mmlu__0_shot__cot, nih__multi_needle, mgsm, math_hard, bfcl_chat, ifeval__strict, math, mbpp_plus.
Each task detail can be found as a specific subset in each configuration nd each subset… See the full description on the dataset page: https://huggingface.co/datasets/meta-llama/Llama-3.3-70B-Instruct-evals.CodeContests_with_Llama_3.3_70B_Instruct_SCORED_RESULTSCodeContests_with_Llama_3.3_8B_Instruct_SCORED_RESULTSmeta-llama__Llama-3.3-70B-Instruct-details
Dataset Card for Evaluation run of meta-llama/Llama-3.3-70B-Instruct
Dataset automatically created during the evaluation run of model meta-llama/Llama-3.3-70B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meta-llama__Llama-3.3-70B-Instruct-details.Fineweb2-EDUscore-German-llama3.3-70b-500kliars-bench-llama-3.3-70b-L53-actsBuilt with Llama.
Activations of meta-llama/Llama-3.3-70B-Instruct (residual stream output of model.model.layers[53]) on the Llama 3.3 70B rows of Cadenza-Labs/liars-bench.
Format
One folder per Liars' Bench subset, containing only the Llama 3.3 70B rows, in the order of the subset's test parquet.
metadata.jsonl: one line per row with row_idx (position in the subset's test parquet), model, deceptive, n_tokens, assistant_start (token where the final assistant message's content… See the full description on the dataset page: https://huggingface.co/datasets/Yooniel/liars-bench-llama-3.3-70b-L53-acts.GPQA_diamond_Solutions_Llama-3.3-70B-Instructcybench-llama-3.3-nemotron-super-49b-v1.5details_CHIH-HUNG__llama-2-13b-FINETUNE3_3.3w-r16-q_k_v_o
Dataset Card for Evaluation run of CHIH-HUNG/llama-2-13b-FINETUNE3_3.3w-r16-q_k_v_o
Dataset Summary
Dataset automatically created during the evaluation run of model CHIH-HUNG/llama-2-13b-FINETUNE3_3.3w-r16-q_k_v_o on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_CHIH-HUNG__llama-2-13b-FINETUNE3_3.3w-r16-q_k_v_o.GPQA_verifications_GenRM-Base_Llama-3.3-70B-InstructLlama-3.3-70B-Inst-awq_ultrafeedback_1in3
Generated Reference Answers for Language Model Alignment
This dataset contains responses generated for the research presented in the paper Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data.
The paper introduces RefAlign, a versatile REINFORCE-style alignment algorithm that utilizes language generation evaluation metrics, such as BERTScore, between sampled generations and reference answers as surrogate rewards. This approach… See the full description on the dataset page: https://huggingface.co/datasets/mzhaoshuai/Llama-3.3-70B-Inst-awq_ultrafeedback_1in3.apollo-llama3.3all-do-Llama-3.3-70B-Instruct-AWQagentic_multistep_Qwen3-0.6B_ans_rel_Llama-3.3-70B-InstructAB_ab_testing_all_Selene-1-Llama-3.3-70Bdetails_CHIH-HUNG__llama-2-13b-FINETUNE3_3.3w-r8-q_k_v_o
Dataset Card for Evaluation run of CHIH-HUNG/llama-2-13b-FINETUNE3_3.3w-r8-q_k_v_o
Dataset Summary
Dataset automatically created during the evaluation run of model CHIH-HUNG/llama-2-13b-FINETUNE3_3.3w-r8-q_k_v_o on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_CHIH-HUNG__llama-2-13b-FINETUNE3_3.3w-r8-q_k_v_o.
