datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SABER
SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
SABER is the code release for the paper. It includes the benchmark tasks, sandbox runtime, judging pipeline, and baseline reproduction utilities used to evaluate operational safety in stateful project workspaces.
What is included
tasks/: benchmark task definitions and metadata
run_osbench.py, judge_osbench.py: historical inference and judging entry points
sandbox_shell.py… See the full description on the dataset page: https://huggingface.co/datasets/sssr-lab/SABER.Phishing_emails_testZro_Mobile_Function_callingZrov2_FineTunedsabersalehk__Llama3-001-300-details
Dataset Card for Evaluation run of sabersalehk/Llama3-001-300
Dataset automatically created during the evaluation run of model sabersalehk/Llama3-001-300
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sabersalehk__Llama3-001-300-details.X-mini-datasets
X-mini-datasets: The Foundational Dataset for Cybersecurity LLMs
Dataset Description
X-mini-datasets is a specialized, English-language dataset engineered as the foundational step to fine-tune Large Language Models (LLMs) into expert cybersecurity assistants. The dataset is uniquely structured into three distinct modules:
Core Knowledge Base (Payloads All The Things Adaptation): The largest part of the dataset, meticulously converted from the legendary "Payloads All The… See the full description on the dataset page: https://huggingface.co/datasets/saberbx/X-mini-datasets.ccf-reasoning-dataset
Cognitive Cascade Framework (CCF) Reasoning Dataset
A high-quality dataset of structured reasoning examples using the Cognitive Cascade Framework (CCF), designed for training language models to perform systematic, multi-stage reasoning.
Dataset Description
This dataset contains problems across multiple domains (math, science, coding, creative reasoning) paired with detailed reasoning chains following the CCF methodology. Each example includes a complete reasoning trace… See the full description on the dataset page: https://huggingface.co/datasets/saberai/ccf-reasoning-dataset.Phishing_emailsMathInstruct_RedPajama_Chat_FormatZro_GSMRedPajama_MetaMath_GSM_Chat_FormatExamen-la-liga-del-saber-historia-cultura-sociedad-geografia-nicaraguaMetaMath-Redpajama-Chat-FormatRedPajama_OpenHermesRedPajama_Orca_dposabersaleh__Llama2-7B-KTO-details
Dataset Card for Evaluation run of sabersaleh/Llama2-7B-KTO
Dataset automatically created during the evaluation run of model sabersaleh/Llama2-7B-KTO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sabersaleh__Llama2-7B-KTO-details.sabersaleh__Llama2-7B-SimPO-details
Dataset Card for Evaluation run of sabersaleh/Llama2-7B-SimPO
Dataset automatically created during the evaluation run of model sabersaleh/Llama2-7B-SimPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sabersaleh__Llama2-7B-SimPO-details.Zrov2sabersalehk__Llama3-SimPO-details
Dataset Card for Evaluation run of sabersalehk/Llama3-SimPO
Dataset automatically created during the evaluation run of model sabersalehk/Llama3-SimPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sabersalehk__Llama3-SimPO-details.sabersalehk__Llama3_01_300-details
Dataset Card for Evaluation run of sabersalehk/Llama3_01_300
Dataset automatically created during the evaluation run of model sabersalehk/Llama3_01_300
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sabersalehk__Llama3_01_300-details.sabersalehk__Llama3_001_200-details
Dataset Card for Evaluation run of sabersalehk/Llama3_001_200
Dataset automatically created during the evaluation run of model sabersalehk/Llama3_001_200
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sabersalehk__Llama3_001_200-details.sabersaleh__Llama2-7B-DPO-details
Dataset Card for Evaluation run of sabersaleh/Llama2-7B-DPO
Dataset automatically created during the evaluation run of model sabersaleh/Llama2-7B-DPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sabersaleh__Llama2-7B-DPO-details.sabersaleh__Llama2-7B-IPO-details
Dataset Card for Evaluation run of sabersaleh/Llama2-7B-IPO
Dataset automatically created during the evaluation run of model sabersaleh/Llama2-7B-IPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sabersaleh__Llama2-7B-IPO-details.sabersaleh__Llama2-7B-CPO-details
Dataset Card for Evaluation run of sabersaleh/Llama2-7B-CPO
Dataset automatically created during the evaluation run of model sabersaleh/Llama2-7B-CPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sabersaleh__Llama2-7B-CPO-details.sabersaleh__Llama2-7B-SPO-details
Dataset Card for Evaluation run of sabersaleh/Llama2-7B-SPO
Dataset automatically created during the evaluation run of model sabersaleh/Llama2-7B-SPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sabersaleh__Llama2-7B-SPO-details.sabersaleh__Llama3-details
Dataset Card for Evaluation run of sabersaleh/Llama3
Dataset automatically created during the evaluation run of model sabersaleh/Llama3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sabersaleh__Llama3-details.x-datasets-mini-markedKuaiMod
Benchmark Description
Data Format
This dataset consists of multiple samples, each containing the following fields:
tag: The label of the sample, indicating the category of the content, such as "pornographic".
title: The title of the video, usually containing the username and user ID.
OCR: Optical Character Recognition results, extracted text from images.
ASR: Automatic Speech Recognition results, extracted text from audio.
images: A list of image filenames, representing… See the full description on the dataset page: https://huggingface.co/datasets/saber0718/KuaiMod.nanhai
