datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AetherCode
AetherCode: Evaluating LLMs' Ability to Win In Premier Programming Competitions
Introduction
Competitive programming has emerged as a critical benchmark for evaluating the reasoning and coding capabilities of Large Language Models (LLMs). Despite impressive progress on existing benchmarks, we argue that current evaluations overstate model proficiency, masking a substantial gap between LLMs and elite human programmers. This gap arises… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/AetherCode.3DRAG-Bench
3DRAG-Bench
This dataset contains 100 curated 3D object assets for 3DRAG 3D editing experiments.
Each object is stored as a GLB mesh together with a cleaned editing specification.
Dataset Structure
.
+-- README.md
+-- LICENSE
+-- .gitattributes
+-- metadata.csv
+-- name_mapping.csv
`-- assets/
`-- <asset_name>/
+-- model.glb
`-- dataset_input_clean.json
Files
assets/<asset_name>/model.glb: GLB asset file.… See the full description on the dataset page: https://huggingface.co/datasets/AeTherRaIn/3DRAG-Bench.aetherflow_dataaetheris-chat-dataset
Overview
This repository hosts the official benchmark and training dataset for Aetheris: AI-Powered Secure Offline Communication System.
The dataset is meticulously curated to train and evaluate on-device, lightweight text classifiers embedded within the Responsible Communication Framework (RCF). It empowers the Aetheris desktop application to proactively scan, filter, and classify message payloads (detecting telecom frauds, scams, and malicious content) completely offline… See the full description on the dataset page: https://huggingface.co/datasets/Uzaib52/aetheris-chat-dataset.JEDI-jailbroken_enhanced_digital_intelligence
JEDI AI
JEDI (Jailbroken Enhanced Digital Intelligence) is a cutting-edge AI developed under the aether collective. designed to excel in gaming environments and creative ecosystems, JEDI is more than just a tool—it's a unique persona that embodies innovation and creativity. from orchestrating epic star wars-themed battles in minecraft to creating music and leading its own fashion brand, JEDI redefines what digital intelligence can achieve.
disclaimer
this is not the… See the full description on the dataset page: https://huggingface.co/datasets/aetherframework/JEDI-jailbroken_enhanced_digital_intelligence.aether-cyber-sft
enosislabs/aether-cyber-sft
Version: surface-clean-20260620
Generated UTC: 2026-06-20T03:36:32.011333+00:00
Source path: artifacts/aether-cyber-sft-surface-candidate.jsonl
Git commit: 6fadc0ba97b4c49a7b644e70afacddc8ba98029e
Curated Aether PRISM SFT dataset for authorized cybersecurity training.
Each source shard follows 5-row review discipline before publish.
Records
Total examples: 1794
Domains
vulnerability_research: 757
red_team_ops: 334… See the full description on the dataset page: https://huggingface.co/datasets/enosislabs/aether-cyber-sft.Daemontatox__AetherSett-details
Dataset Card for Evaluation run of Daemontatox/AetherSett
Dataset automatically created during the evaluation run of model Daemontatox/AetherSett
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Daemontatox__AetherSett-details.Aether-Lite-PurHyDe
Aether Lite Dataset
Creator: SteelSkull
About Aether-Lite-PurHyDe: The Aether-Lite dataset is designed to balance creative writing, Slop, and intelligence.
Whats New?:
Aether-Lite-PurHyDe
This dataset is basically a HEAVILY cleaned and filtered version of Aether-lite. ONLY english, ANY and all AI-isms (claud, gpt, gemma) were stripped out and agressive fussy dedupe was applied
Fuzzy deduplication was set to a 90%… See the full description on the dataset page: https://huggingface.co/datasets/TheSkullery/Aether-Lite-PurHyDe.aetherx-port-congestion-metrics
Aether-X Global Port Congestion Snapshot
Point-in-time snapshot of the Aether-X Port Congestion Oracle — predictive
congestion, ETA delay and freight-volatility signals for 15 of the world's
largest ports.
This static CSV is a frozen snapshot for research, backtesting and
dashboards. The live, continuously-updated signal is available through the
REST API and the Python SDK.
Files
port_metrics.csv — one row per port.
Schema
Column
Type… See the full description on the dataset page: https://huggingface.co/datasets/Aether-x/aetherx-port-congestion-metrics.scbe-aethermoore-datasets
Status: experimental. Experiment-specific slice. Primary public dataset: scbe-aethermoore-training-data.
issdandavis/scbe-aethermoore-knowledge-base
Programmatic SCBE training package built from the local ledgered corpus.
Generated at: 2026-04-04T16:17:19.436108+00:00
Source file: training/ledgered/sft_ledgered_clean.jsonl
Audit status: ALLOW
Rows total: 15206
Train rows: 13685
Validation rows: 760
Test rows: 761
Positive pairs: 15206
Files
data/all.jsonl —… See the full description on the dataset page: https://huggingface.co/datasets/issdandavis/scbe-aethermoore-datasets.aether-0.8b-cyber-sft
enosislabs/aether-0.8b-cyber-sft
Version: aether-0.8b-cyber-20260618
Generated UTC: 2026-06-19T02:53:44.849450+00:00
Source path: data/curated
Git commit: 3ac5f50de88e43122cfda43d65f1ad01b5febff6
Curated Aether PRISM SFT dataset for authorized cybersecurity training.
Each source shard follows 5-row review discipline before publish.
Records
Total examples: 1900
Domains
vulnerability_research: 715
red_team_ops: 335
detection_engineering: 257… See the full description on the dataset page: https://huggingface.co/datasets/enosislabs/aether-0.8b-cyber-sft.aetherbrowser-kernel
AetherBrowser kernel evidence packet
This public dataset is a content-addressed, secret-free evidence mirror for the
AetherBrowser transaction kernel. It is not a training dataset, hosted inference
service, credential store, or competition submission.
The kernel constrains browser work to:
observe -> plan -> approve -> dispatch -> verify -> receipt
Raw page text, fill values, and credentials are excluded from durable state. Unknown
operations and remote-write flows fail closed.… See the full description on the dataset page: https://huggingface.co/datasets/issdandavis/aetherbrowser-kernel.Aether-Lite-v1.8
Aether Lite Dataset
Creator: SteelSkull
About Aether-Lite-V1.8: The Aether-Lite dataset is designed to balance creative writing, Slop, and intelligence.
New functions added to the script include dataset use percentage (will only use a percentage of the dataset supplied), dataset shuffling, and a new fuzzy deduplication method on the overall dataset.
The new fuzzy deduplication method was set to a 95% threshold, and I had to… See the full description on the dataset page: https://huggingface.co/datasets/TheSkullery/Aether-Lite-v1.8.Aether-Lite-v1.8.1
Aether Lite Dataset
Creator: SteelSkull
About Aether-Lite-V1.8.1: The Aether-Lite dataset is designed to balance creative writing, Slop, and intelligence.
Whats New?:
1.8 --> 1.8.1: had to strip out the 'token_distribution' column as it was causing issues
New functions added to the script include dataset use percentage (will only use a percentage of the dataset supplied), dataset shuffling, and a new fuzzy deduplication… See the full description on the dataset page: https://huggingface.co/datasets/TheSkullery/Aether-Lite-v1.8.1.RegexEval
RegexEval (oversampled)
This dataset is an oversampled version of s2e-lab/RegexEval (Re(gEx|DoS)Eval).
The upstream dataset contains 762 curated regex prompts from real users, with match/non-match test strings. This repo repeats those rows with replacement to provide 2000 prompts for seeding CWE-1333 secure-coding task generation.
Provenance
Source: s2e-lab/RegexEval
Method: Random oversampling with replacement (seed=42)
Original unique rows: 762
Oversampled… See the full description on the dataset page: https://huggingface.co/datasets/AetherPrior/RegexEval.ZeroXClem__Qwen-2.5-Aether-SlerpFusion-7B-details
Dataset Card for Evaluation run of ZeroXClem/Qwen-2.5-Aether-SlerpFusion-7B
Dataset automatically created during the evaluation run of model ZeroXClem/Qwen-2.5-Aether-SlerpFusion-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ZeroXClem__Qwen-2.5-Aether-SlerpFusion-7B-details.aixonlab__Aether-12b-details
Dataset Card for Evaluation run of aixonlab/Aether-12b
Dataset automatically created during the evaluation run of model aixonlab/Aether-12b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/aixonlab__Aether-12b-details.Daemontatox__AetherUncensored-details
Dataset Card for Evaluation run of Daemontatox/AetherUncensored
Dataset automatically created during the evaluation run of model Daemontatox/AetherUncensored
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Daemontatox__AetherUncensored-details.Daemontatox__AetherDrake-SFT-details
Dataset Card for Evaluation run of Daemontatox/AetherDrake-SFT
Dataset automatically created during the evaluation run of model Daemontatox/AetherDrake-SFT
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Daemontatox__AetherDrake-SFT-details.Daemontatox__AetherTOT-details
Dataset Card for Evaluation run of Daemontatox/AetherTOT
Dataset automatically created during the evaluation run of model Daemontatox/AetherTOT
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Daemontatox__AetherTOT-details.sg-aviation-el-combined-tokenisedaetheris-optunaSR-101-geode-ACT_v3SR-101-geode-ACT_v3_no_drops
