datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lm-eval-results-shyamieee-JARVIS-v2.0-private
Dataset Card for Evaluation run of shyamieee/JARVIS-v2.0
Dataset automatically created during the evaluation run of model shyamieee/JARVIS-v2.0
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-JARVIS-v2.0-private.cMedQA-V2.0lm-eval-results-chlee10-T3Q-Mistral-Orca-Math-dpo-v2.0-private
Dataset Card for Evaluation run of chlee10/T3Q-Mistral-Orca-Math-dpo-v2.0
Dataset automatically created during the evaluation run of model chlee10/T3Q-Mistral-Orca-Math-dpo-v2.0
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-chlee10-T3Q-Mistral-Orca-Math-dpo-v2.0-private.newsqa_200_11064_v2.0.0
Private NewsQA RAG Evaluation Dataset
Private, human-reviewed evaluation source data for the NewsQA RAG project.
This repository is not a prebuilt retrieval index. Chunk, BM25, Chroma, and
ground-truth chunk mappings must be rebuilt from the pinned release.
Version: v1.0.0
Evaluation articles: 200
Distractor articles: 10864
Source questions: 1340
Redistribution rights for the upstream NewsQA-derived text must be verified
before changing this repository from private to public.
lm-eval-results-shyamieee-B3E3-SLM-7b-v2.0-private
Dataset Card for Evaluation run of shyamieee/B3E3-SLM-7b-v2.0
Dataset automatically created during the evaluation run of model shyamieee/B3E3-SLM-7b-v2.0
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-B3E3-SLM-7b-v2.0-private.Nexesenex__Llama_3.1_8b_DodoWild_v2.02-details
Dataset Card for Evaluation run of Nexesenex/Llama_3.1_8b_DodoWild_v2.02
Dataset automatically created during the evaluation run of model Nexesenex/Llama_3.1_8b_DodoWild_v2.02
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Nexesenex__Llama_3.1_8b_DodoWild_v2.02-details.FuseAI__FuseChat-7B-v2.0-details
Dataset Card for Evaluation run of FuseAI/FuseChat-7B-v2.0
Dataset automatically created during the evaluation run of model FuseAI/FuseChat-7B-v2.0
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FuseAI__FuseChat-7B-v2.0-details.Cybersecurity-Dataset-Fenrir-v2.0
Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
Created by Alican Kiraz
TL;DR
A ready-to-train dataset of 83,920 high-quality system / user / assistant triples for defensive, alignment-safe cybersecurity SFT training.
Apache-2.0 licensed and production-ready.
Scope: OWASP Top 10, MITRE ATT&CK, NIST CSF, CIS Controls, ASD Essential 8, modern authentication (OAuth 2 / OIDC / SAML), SSL / TLS, Cloud & DevSecOps, Cryptography, and AI Security.
1 What’s… See the full description on the dataset page: https://huggingface.co/datasets/ukcli/Cybersecurity-Dataset-Fenrir-v2.0.Cybersecurity-Dataset-Fenrir-v2.0
Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
Created by Alican Kiraz
TL;DR
A ready-to-train dataset of 83,920 high-quality system / user / assistant triples for defensive, alignment-safe cybersecurity SFT training.
Apache-2.0 licensed and production-ready.
Scope: OWASP Top 10, MITRE ATT&CK, NIST CSF, CIS Controls, ASD Essential 8, modern authentication (OAuth 2 / OIDC / SAML), SSL / TLS, Cloud & DevSecOps, Cryptography, and AI Security.
1 What’s… See the full description on the dataset page: https://huggingface.co/datasets/PhSecX/Cybersecurity-Dataset-Fenrir-v2.0.speakleash__Bielik-11B-v2.0-Instruct-details
Dataset Card for Evaluation run of speakleash/Bielik-11B-v2.0-Instruct
Dataset automatically created during the evaluation run of model speakleash/Bielik-11B-v2.0-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/speakleash__Bielik-11B-v2.0-Instruct-details.Cybersecurity-Dataset-Fenrir-v2.0
Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
Created by Alican Kiraz
TL;DR
A ready-to-train dataset of 83,920 high-quality system / user / assistant triples for defensive, alignment-safe cybersecurity SFT training.
Apache-2.0 licensed and production-ready.
Scope: OWASP Top 10, MITRE ATT&CK, NIST CSF, CIS Controls, ASD Essential 8, modern authentication (OAuth 2 / OIDC / SAML), SSL / TLS, Cloud & DevSecOps, Cryptography, and AI Security.
1 What’s… See the full description on the dataset page: https://huggingface.co/datasets/invinciblejha01/Cybersecurity-Dataset-Fenrir-v2.0.perguntas_e_respostas_astronomia_pt_br_V2.0
README: Dataset JSON de Perguntas e Respostas sobre Astronomia
Visão Geral
Este repositório contém um dataset sintético de perguntas e respostas sobre astronomia, gerado no formato JSON. O objetivo deste dataset é servir como material educativo e para o ajuste fino (fine-tuning) de Modelos de Linguagem Grandes (LLMs).
Esse dataset contém 1000 perguntas e respostas variadas sobre tópicos de astronomia, como planetas, estrelas, galáxias e fenômenos cósmicos, juntamente com… See the full description on the dataset page: https://huggingface.co/datasets/vinimuchulski/perguntas_e_respostas_astronomia_pt_br_V2.0.Casual-Autopsy__L3-Umbral-Mind-RP-v2.0-8B-details
Dataset Card for Evaluation run of Casual-Autopsy/L3-Umbral-Mind-RP-v2.0-8B
Dataset automatically created during the evaluation run of model Casual-Autopsy/L3-Umbral-Mind-RP-v2.0-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Casual-Autopsy__L3-Umbral-Mind-RP-v2.0-8B-details.Triangle104__Rocinante-Prism_V2.0-details
Dataset Card for Evaluation run of Triangle104/Rocinante-Prism_V2.0
Dataset automatically created during the evaluation run of model Triangle104/Rocinante-Prism_V2.0
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__Rocinante-Prism_V2.0-details.Cybersecurity-Dataset-Fenrir-v2.0
Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
Created by Alican Kiraz
TL;DR
A ready-to-train dataset of 83,920 high-quality system / user / assistant triples for defensive, alignment-safe cybersecurity SFT training.
Apache-2.0 licensed and production-ready.
Scope: OWASP Top 10, MITRE ATT&CK, NIST CSF, CIS Controls, ASD Essential 8, modern authentication (OAuth 2 / OIDC / SAML), SSL / TLS, Cloud & DevSecOps, Cryptography, and AI Security.
1 What’s… See the full description on the dataset page: https://huggingface.co/datasets/Zud0/Cybersecurity-Dataset-Fenrir-v2.0.Nexesenex__Llama_3.1_8b_DoberWild_v2.03-details
Dataset Card for Evaluation run of Nexesenex/Llama_3.1_8b_DoberWild_v2.03
Dataset automatically created during the evaluation run of model Nexesenex/Llama_3.1_8b_DoberWild_v2.03
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Nexesenex__Llama_3.1_8b_DoberWild_v2.03-details.TheatreLM-v2.0-chats-previewTheatreLM-v2.0-CharactersPlease use https://huggingface.co/datasets/G-reen/TheatreLM-v2.1-Characters instead.
spow12__ChatWaifu_22B_v2.0_preview-details
Dataset Card for Evaluation run of spow12/ChatWaifu_22B_v2.0_preview
Dataset automatically created during the evaluation run of model spow12/ChatWaifu_22B_v2.0_preview
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/spow12__ChatWaifu_22B_v2.0_preview-details.spow12__ChatWaifu_12B_v2.0-details
Dataset Card for Evaluation run of spow12/ChatWaifu_12B_v2.0
Dataset automatically created during the evaluation run of model spow12/ChatWaifu_12B_v2.0
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/spow12__ChatWaifu_12B_v2.0-details.jpacifico__Chocolatine-2-14B-Instruct-v2.0.1-details
Dataset Card for Evaluation run of jpacifico/Chocolatine-2-14B-Instruct-v2.0.1
Dataset automatically created during the evaluation run of model jpacifico/Chocolatine-2-14B-Instruct-v2.0.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jpacifico__Chocolatine-2-14B-Instruct-v2.0.1-details.Nexesenex__Llama_3.1_8b_DoberWild_v2.01-details
Dataset Card for Evaluation run of Nexesenex/Llama_3.1_8b_DoberWild_v2.01
Dataset automatically created during the evaluation run of model Nexesenex/Llama_3.1_8b_DoberWild_v2.01
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Nexesenex__Llama_3.1_8b_DoberWild_v2.01-details.spow12__ChatWaifu_v2.0_22B-details
Dataset Card for Evaluation run of spow12/ChatWaifu_v2.0_22B
Dataset automatically created during the evaluation run of model spow12/ChatWaifu_v2.0_22B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/spow12__ChatWaifu_v2.0_22B-details.Triangle104__Chatty-Harry_V2.0-details
Dataset Card for Evaluation run of Triangle104/Chatty-Harry_V2.0
Dataset automatically created during the evaluation run of model Triangle104/Chatty-Harry_V2.0
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__Chatty-Harry_V2.0-details.Nexesenex__Llama_3.1_8b_DodoWild_v2.01-details
Dataset Card for Evaluation run of Nexesenex/Llama_3.1_8b_DodoWild_v2.01
Dataset automatically created during the evaluation run of model Nexesenex/Llama_3.1_8b_DodoWild_v2.01
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Nexesenex__Llama_3.1_8b_DodoWild_v2.01-details.Nexesenex__Llama_3.1_8b_DoberWild_v2.02-details
Dataset Card for Evaluation run of Nexesenex/Llama_3.1_8b_DoberWild_v2.02
Dataset automatically created during the evaluation run of model Nexesenex/Llama_3.1_8b_DoberWild_v2.02
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Nexesenex__Llama_3.1_8b_DoberWild_v2.02-details.Nexesenex__Llama_3.1_8b_DodoWild_v2.03-details
Dataset Card for Evaluation run of Nexesenex/Llama_3.1_8b_DodoWild_v2.03
Dataset automatically created during the evaluation run of model Nexesenex/Llama_3.1_8b_DodoWild_v2.03
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Nexesenex__Llama_3.1_8b_DodoWild_v2.03-details.PJMixers__LLaMa-3-CursedStock-v2.0-8B-details
Dataset Card for Evaluation run of PJMixers/LLaMa-3-CursedStock-v2.0-8B
Dataset automatically created during the evaluation run of model PJMixers/LLaMa-3-CursedStock-v2.0-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/PJMixers__LLaMa-3-CursedStock-v2.0-8B-details.cat_subcat_mapping_training_data_v2.0Tengentoppa-sft-reasoning-v2.0
