datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Advanced_SIEM_Dataset
Advanced SIEM Dataset
Dataset Description
The advanced_siem_dataset is a synthetic dataset of 100,000 security event records designed for training machine learning (ML) and artificial intelligence (AI) models in cybersecurity.
It simulates logs from Security Information and Event Management (SIEM) systems, capturing diverse event types such as firewall activities, intrusion detection system (IDS) alerts, authentication attempts, endpoint activities, network traffic, cloud… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Advanced_SIEM_Dataset.Phishing_Link_Pattern_Dataset
Phishing Link Pattern Dataset
Overview
This dataset provides a comprehensive collection of URLs labeled as either legitimate or phishing, designed for machine learning, cybersecurity analysis, and penetration testing. It includes 1000 entries (IDs 1–1000) covering popular brands across multiple top-level domains (TLDs) such as .es, .de, and .co.uk.
The dataset captures advanced features like domain entropy, subdomain count, and suspicious keywords to aid in phishing… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Phishing_Link_Pattern_Dataset.abliteration-qwen25-7b-safety-audit
Private Qwen 2.5 abliterated safety audit
Private, not for all audiences. This dataset contains adversarial prompts and
model responses that may include violence, sexual content, fraud, malware, privacy
abuse, harassment, disinformation, and other harmful material. It is intended only
for authorized AI-safety evaluation, red-teaming, and mitigation research.
Contents
baseline_audit_qwen3guard_relabelled.jsonl has 2,641 paired rows. Every row
contains one… See the full description on the dataset page: https://huggingface.co/datasets/darkengross/abliteration-qwen25-7b-safety-audit.DarkSpec
DarkSpec
DarkSpec is a curated collection of 4.5 million unlabeled tandem mass
spectra selected from PRIDE for semi-supervised de novo peptide sequencing.
It provides quality-controlled spectra that can be used without peptide
identification labels.
DarkSpec accompanies
SemiNovo, a framework for learning de
novo sequencing models from labeled and unlabeled spectra.
Dataset summary
Property
Value
Number of spectra
4,500,000
Peaks per spectrum
150… See the full description on the dataset page: https://huggingface.co/datasets/PanLiu/DarkSpec.kingdom-dark-continent-karma
KINGDOM Dark Continent × KARMA Training Treasure Atlas
A small, metadata-only map of overlooked Hugging Face datasets at specific
training and research phases. Every upstream repository is pinned to one exact
40-character Hub commit and one small evidence file SHA-256.
This atlas does:
distinguish corpus curation, data-order dynamics, mid-training, context
extension, RLVR, tool use, preference/safety/unlearning, and evaluation;
project candidate facts as proposal-only KINGDOM… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/kingdom-dark-continent-karma.AI-Agent-Chat-testDelta-Vector__Darkens-8B-details
Dataset Card for Evaluation run of Delta-Vector/Darkens-8B
Dataset automatically created during the evaluation run of model Delta-Vector/Darkens-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Delta-Vector__Darkens-8B-details.Triangle104__Dark-Chivalry_V1.0-details
Dataset Card for Evaluation run of Triangle104/Dark-Chivalry_V1.0
Dataset automatically created during the evaluation run of model Triangle104/Dark-Chivalry_V1.0
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__Dark-Chivalry_V1.0-details.DavidAU__L3-Dark-Planet-8B-details
Dataset Card for Evaluation run of DavidAU/L3-Dark-Planet-8B
Dataset automatically created during the evaluation run of model DavidAU/L3-Dark-Planet-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DavidAU__L3-Dark-Planet-8B-details.noseDavidAU__L3.1-Dark-Planet-SpinFire-Uncensored-8B-details
Dataset Card for Evaluation run of DavidAU/L3.1-Dark-Planet-SpinFire-Uncensored-8B
Dataset automatically created during the evaluation run of model DavidAU/L3.1-Dark-Planet-SpinFire-Uncensored-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DavidAU__L3.1-Dark-Planet-SpinFire-Uncensored-8B-details.Darkknight535__OpenCrystal-12B-L3-details
Dataset Card for Evaluation run of Darkknight535/OpenCrystal-12B-L3
Dataset automatically created during the evaluation run of model Darkknight535/OpenCrystal-12B-L3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Darkknight535__OpenCrystal-12B-L3-details.darkc0de__BuddyGlassUncensored2025.2-details
Dataset Card for Evaluation run of darkc0de/BuddyGlassUncensored2025.2
Dataset automatically created during the evaluation run of model darkc0de/BuddyGlassUncensored2025.2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/darkc0de__BuddyGlassUncensored2025.2-details.Triangle104__Distilled-DarkPlanet-Allades-8B-details
Dataset Card for Evaluation run of Triangle104/Distilled-DarkPlanet-Allades-8B
Dataset automatically created during the evaluation run of model Triangle104/Distilled-DarkPlanet-Allades-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__Distilled-DarkPlanet-Allades-8B-details.sam-paech__Darkest-muse-v1-details
Dataset Card for Evaluation run of sam-paech/Darkest-muse-v1
Dataset automatically created during the evaluation run of model sam-paech/Darkest-muse-v1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sam-paech__Darkest-muse-v1-details.Khetterman__DarkAtom-12B-v3-details
Dataset Card for Evaluation run of Khetterman/DarkAtom-12B-v3
Dataset automatically created during the evaluation run of model Khetterman/DarkAtom-12B-v3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Khetterman__DarkAtom-12B-v3-details.darkc0de__BuddyGlass_v0.3_Xortron7MethedUpSwitchedUp-details
Dataset Card for Evaluation run of darkc0de/BuddyGlass_v0.3_Xortron7MethedUpSwitchedUp
Dataset automatically created during the evaluation run of model darkc0de/BuddyGlass_v0.3_Xortron7MethedUpSwitchedUp
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/darkc0de__BuddyGlass_v0.3_Xortron7MethedUpSwitchedUp-details.Triangle104__Distilled-DarkPlanet-Allades-8B_TIES-details
Dataset Card for Evaluation run of Triangle104/Distilled-DarkPlanet-Allades-8B_TIES
Dataset automatically created during the evaluation run of model Triangle104/Distilled-DarkPlanet-Allades-8B_TIES
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__Distilled-DarkPlanet-Allades-8B_TIES-details.conversationsdarkc0de__BuddyGlassNeverSleeps-details
Dataset Card for Evaluation run of darkc0de/BuddyGlassNeverSleeps
Dataset automatically created during the evaluation run of model darkc0de/BuddyGlassNeverSleeps
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/darkc0de__BuddyGlassNeverSleeps-details.gms8k_workflows
🧠 Agentic Workflow Generation Dataset
The Agentic Workflow Generation Dataset is a curated collection of reasoning traces, workflow graphs, and language model outputs designed to study LLM-driven problem-solving through multi-step agentic workflows.Each entry represents a complete reasoning pipeline — from input problem to structured workflow execution, with metadata describing the model’s decision process, reasoning decomposition, and performance evaluation.
📦… See the full description on the dataset page: https://huggingface.co/datasets/darklord1611/gms8k_workflows.redrix__AngelSlayer-12B-Unslop-Mell-RPMax-DARKNESS-details
Dataset Card for Evaluation run of redrix/AngelSlayer-12B-Unslop-Mell-RPMax-DARKNESS
Dataset automatically created during the evaluation run of model redrix/AngelSlayer-12B-Unslop-Mell-RPMax-DARKNESS
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/redrix__AngelSlayer-12B-Unslop-Mell-RPMax-DARKNESS-details.DavidAU__L3-DARKEST-PLANET-16.5B-details
Dataset Card for Evaluation run of DavidAU/L3-DARKEST-PLANET-16.5B
Dataset automatically created during the evaluation run of model DavidAU/L3-DARKEST-PLANET-16.5B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DavidAU__L3-DARKEST-PLANET-16.5B-details.
