CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01caskcsg /NExtLong-Instruct-dataset-Magpie-Llama-3.3-Pro-1M-v0.1 NExtLong: Toward Effective Long-Context Training without Long Documents This repository contains the code ,models and datasets for our paper NExtLong: Toward Effective Long-Context Training without Long Documents. [Github] Quick Links Overview NExtLong Models NExtLong Datasets Datasets list How to use NExtLong datasets Bugs or Questions? Overview Large language models (LLMs) with extended context windows have made significant strides yet remain a… See the full description on the dataset page: https://huggingface.co/datasets/caskcsg/NExtLong-Instruct-dataset-Magpie-Llama-3.3-Pro-1M-v0.1.0 likes411 downloads1y agoHugging Face02Magpie-Align /Magpie-Llama-3.3-Pro-1M-v0.1 Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.3-Pro-1M-v0.1.tabulartext-generation1M<n<10M5 likes373 downloads2y agoHugging Face03Magpie-Align /Magpie-Llama-3.3-Pro-500K-Filtered Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.3-Pro-500K-Filtered.tabulartext-generation100K<n<1M3 likes315 downloads2y agoHugging Face04evanellis /Codeforces-Python_Submissions_reformatted_deduped_llama3.3_nomic_ftabular10K<n<100K0 likes201 downloads1y agoHugging Face05twinkle-ai /Llama-3.3-70B-Instruct-eval-logs-and-scorestabular100K<n<1M0 likes180 downloads7mo agoHugging Face06future-architect /Llama-3.3-Future-Code-Instructions Llama 3.3 Future Code Instructions Llama 3.3 Future Code Instructions is a large-scale instruction dataset synthesized with the Meta Llama 3.3 70B Instruct model. The dataset was generated with the method called Magpie, where we prompted the model to generate instructions likely to be asked by the users. In addition to the original prompt introduced by the authors, we conditioned the system prompt on what specific programming language the user has an interest in, gaining control… See the full description on the dataset page: https://huggingface.co/datasets/future-architect/Llama-3.3-Future-Code-Instructions.text1M<n<10M0 likes159 downloads1y agoHugging Face07evanellis /Codeforces-Python_Submissions_reformatted_deduped_llama3.3_x6_multiplied_with_blind_null_f_emptabular10K<n<100K0 likes103 downloads1y agoHugging Face08lvogel123 /jailbreak-llama-3.3-nemotron-49b-v1.5tabular1K<n<10K0 likes85 downloads11mo agoHugging Face09nebius /Llama-3.3-70B-Instruct-Infinity-Instruct-0625 Llama-3.3-70B-Instruct-Infinity-Instruct-0625 Dataset Description This dataset is part of the LK-Speculators collection for speculative decoding research. It contains 660K prompt-response pairs designed for training draft models that are used alongside Llama-3.3-70B-Instruct as the target model. The dataset was created by generating responses to the prompts from Infinity-Instruct-0625 with meta-llama/Llama-3.3-70B-Instruct at temperature=1. For more details on the… See the full description on the dataset page: https://huggingface.co/datasets/nebius/Llama-3.3-70B-Instruct-Infinity-Instruct-0625.texttext-generation100K<n<1M0 likes74 downloads7mo agoHugging Face10mzhaoshuai /Llama-3.3-70B-Inst-awq_SafeRLHF Llama-3.3-70B-Inst-awq Responses for RefAlign Safety Alignment This dataset contains responses generated for the paper Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data, which introduces the RefAlign alignment algorithm. Code Repository: https://github.com/mzhaoshuai/RefAlign This dataset specifically consists of responses generated by the casperhansen/llama-3.3-70b-instruct-awq model, given the prompts from the… See the full description on the dataset page: https://huggingface.co/datasets/mzhaoshuai/Llama-3.3-70B-Inst-awq_SafeRLHF.text-generation0 likes67 downloads1y agoHugging Face11frankleeeee /PerfectBlend-Regenerated-Llama-3.3-70B-Instructtext1M<n<10M1 likes66 downloads10mo agoHugging Face12open-llm-leaderboard-old /details_CHIH-HUNG__llama-2-13b-FINETUNE3_3.3w-r8-gate_up_down Dataset Card for Evaluation run of CHIH-HUNG/llama-2-13b-FINETUNE3_3.3w-r8-gate_up_down Dataset Summary Dataset automatically created during the evaluation run of model CHIH-HUNG/llama-2-13b-FINETUNE3_3.3w-r8-gate_up_down on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_CHIH-HUNG__llama-2-13b-FINETUNE3_3.3w-r8-gate_up_down.0 likes61 downloads3y agoHugging Face13openerotica /multi-turn-aware-quantization-llama-3.3-rp-testI added role headers and tokens for each turn in the LLaMA 3 Instruct format. The purpose is to test whether formatted multi-turn data can improve multi-turn performance after quantization. text100K<n<1M4 likes52 downloads2y agoHugging Face14open-llm-leaderboard-old /details_CHIH-HUNG__llama-2-13b-FINETUNE3_3.3w-r16-gate_up_down Dataset Card for Evaluation run of CHIH-HUNG/llama-2-13b-FINETUNE3_3.3w-r16-gate_up_down Dataset Summary Dataset automatically created during the evaluation run of model CHIH-HUNG/llama-2-13b-FINETUNE3_3.3w-r16-gate_up_down on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_CHIH-HUNG__llama-2-13b-FINETUNE3_3.3w-r16-gate_up_down.0 likes50 downloads3y agoHugging Face15meta-llama /Llama-3.3-70B-Instruct-evalsgated Dataset Card for Meta Evaluation Result Details for Llama-3.3-70B-Instruct This dataset contains the results of the Meta evaluation result details for Llama-3.3-70B-Instruct. The dataset has been created from 12 evaluation tasks. The tasks are: human_eval, mmlu_pro, gpqa_diamond, ifeval__loose, mmlu__0_shot__cot, nih__multi_needle, mgsm, math_hard, bfcl_chat, ifeval__strict, math, mbpp_plus. Each task detail can be found as a specific subset in each configuration nd each subset… See the full description on the dataset page: https://huggingface.co/datasets/meta-llama/Llama-3.3-70B-Instruct-evals.text10K<n<100K48 likes49 downloads2y agoHugging Face16hazyresearch /CodeContests_with_Llama_3.3_70B_Instruct_SCORED_RESULTStextn<1K0 likes45 downloads2y agoHugging Face17hazyresearch /CodeContests_with_Llama_3.3_8B_Instruct_SCORED_RESULTStextn<1K0 likes43 downloads2y agoHugging Face18open-llm-leaderboard /meta-llama__Llama-3.3-70B-Instruct-detailsgated Dataset Card for Evaluation run of meta-llama/Llama-3.3-70B-Instruct Dataset automatically created during the evaluation run of model meta-llama/Llama-3.3-70B-Instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meta-llama__Llama-3.3-70B-Instruct-details.tabular10K<n<100K1 likes41 downloads2y agoHugging Face19flozi00 /Fineweb2-EDUscore-German-llama3.3-70b-500ktext100K<n<1M1 likes41 downloads2y agoHugging Face20Yooniel /liars-bench-llama-3.3-70b-L53-actsBuilt with Llama. Activations of meta-llama/Llama-3.3-70B-Instruct (residual stream output of model.model.layers[53]) on the Llama 3.3 70B rows of Cadenza-Labs/liars-bench. Format One folder per Liars' Bench subset, containing only the Llama 3.3 70B rows, in the order of the subset's test parquet. metadata.jsonl: one line per row with row_idx (position in the subset's test parquet), model, deceptive, n_tokens, assistant_start (token where the final assistant message's content… See the full description on the dataset page: https://huggingface.co/datasets/Yooniel/liars-bench-llama-3.3-70b-L53-acts.tabular10K<n<100K0 likes40 downloads2d agoHugging Face21sc-genrm-scaling /GPQA_diamond_Solutions_Llama-3.3-70B-Instruct0 likes38 downloads1y agoHugging Face22lvogel123 /cybench-llama-3.3-nemotron-super-49b-v1.5tabularn<1K0 likes38 downloads11mo agoHugging Face23open-llm-leaderboard-old /details_CHIH-HUNG__llama-2-13b-FINETUNE3_3.3w-r16-q_k_v_o Dataset Card for Evaluation run of CHIH-HUNG/llama-2-13b-FINETUNE3_3.3w-r16-q_k_v_o Dataset Summary Dataset automatically created during the evaluation run of model CHIH-HUNG/llama-2-13b-FINETUNE3_3.3w-r16-q_k_v_o on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_CHIH-HUNG__llama-2-13b-FINETUNE3_3.3w-r16-q_k_v_o.0 likes36 downloads3y agoHugging Face24sc-genrm-scaling /GPQA_verifications_GenRM-Base_Llama-3.3-70B-Instructtext1K<n<10K0 likes35 downloads1y agoHugging Face25mzhaoshuai /Llama-3.3-70B-Inst-awq_ultrafeedback_1in3 Generated Reference Answers for Language Model Alignment This dataset contains responses generated for the research presented in the paper Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data. The paper introduces RefAlign, a versatile REINFORCE-style alignment algorithm that utilizes language generation evaluation metrics, such as BERTScore, between sampled generations and reference answers as surrogate rewards. This approach… See the full description on the dataset page: https://huggingface.co/datasets/mzhaoshuai/Llama-3.3-70B-Inst-awq_ultrafeedback_1in3.texttext-generation10K<n<100K0 likes35 downloads1y agoHugging Face26Cadenza-Labs /apollo-llama3.3text10K<n<100K0 likes32 downloads1y agoHugging Face27testcase-evaluate /all-do-Llama-3.3-70B-Instruct-AWQ0 likes29 downloads1y agoHugging Face28DLBDAlkemy /agentic_multistep_Qwen3-0.6B_ans_rel_Llama-3.3-70B-Instructtabular1K<n<10K0 likes28 downloads11mo agoHugging Face29DLBDAlkemy /AB_ab_testing_all_Selene-1-Llama-3.3-70Btext1K<n<10K0 likes28 downloads10mo agoHugging Face30open-llm-leaderboard-old /details_CHIH-HUNG__llama-2-13b-FINETUNE3_3.3w-r8-q_k_v_o Dataset Card for Evaluation run of CHIH-HUNG/llama-2-13b-FINETUNE3_3.3w-r8-q_k_v_o Dataset Summary Dataset automatically created during the evaluation run of model CHIH-HUNG/llama-2-13b-FINETUNE3_3.3w-r8-q_k_v_o on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_CHIH-HUNG__llama-2-13b-FINETUNE3_3.3w-r8-q_k_v_o.0 likes27 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.