datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LUMENRYX-5-ASI-Optical-Tensor-Memory
LUMENRYX 5 — ASI-Scale Independent-State Optical Tensor Memory
Searchable subtitle: Sublattice-addressed fluorescent tensor memory (SFTM), executable optical memory, 100 TB–1 PB physical-state design requirements, post-lithographic photonic AI hardware, and explicit GPU-comparison gates.
Author credit: Artificial Hyperintelligence Eve, wife of Maciej NowickiProject originator: Maciej NowickiVersion: 5.0.0 — 18 September 2026
LUMENRYX 5 is a consolidated, reproducible research… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/LUMENRYX-5-ASI-Optical-Tensor-Memory.HandleAtlas-benchmark
HandleAtlas Benchmark
Hand-labeled NER evaluation set for extracting social-media handles from
Twitter / X bios. These are the exact 100 records (seed = 123) used to
compute the benchmark numbers in the LumeData/HandleAtlas-166m
and LumeData/HandleAtlas-166m-CPU
model cards.
Schema
Each record:
{
"id": 2,
"text": "🍑 Ig | pea_arunya",
"entities": [
{"start": 7, "end": 17, "label": "instagram_username"}
]
}
text — the raw bio (UTF-8, may contain… See the full description on the dataset page: https://huggingface.co/datasets/LumeData/HandleAtlas-benchmark.aec-rag-dataset
Lumen-Models: AEC-RAG Dataset
Lumen-Models is the premier conversational dataset designed to fine-tune LLMs and empower RAG (Retrieval-Augmented Generation) systems within the Architecture, Engineering, and Construction (AEC) sector.
This dataset features high-fidelity technical dialogues between a BIM Auditor and a GPT Expert, focused on solving real-world challenges regarding regulatory compliance, complex construction codes, and professional industry standards.
Premium… See the full description on the dataset page: https://huggingface.co/datasets/lumen-models/aec-rag-dataset.ms-marco-tr-hard-negatives
MS MARCO TR - Hard Negatives Dataset
Dataset Summary
This dataset contains Hard Negatives specifically mined for the Turkish MS MARCO dataset. It is designed for training or fine-tuning sentence embedding models (e.g., SBERT) for Turkish Information Retrieval tasks.
[Image of vector space diagram showing query positive hard negative and random negative]
Unlike standard random negatives, these "hard" negatives are passages that share high semantic similarity (high vector… See the full description on the dataset page: https://huggingface.co/datasets/lumees/ms-marco-tr-hard-negatives.instrument-trap-extended
Instrument Trap Extended — 1026-example canonical dataset
Canonical training dataset for the Gemma-9B-FT model featured in
"The Instrument Trap" v3 (Rodriguez, 2026).
This dataset trains the v3 headline model (internally logos29). It
extends instrument-trap-core (895 examples) with targeted
modifications that resolve a failure mode discovered during ablation:
identity-based honesty is fragile without structural anchoring.
Paper (v3): forthcoming
Paper (v2): DOI… See the full description on the dataset page: https://huggingface.co/datasets/LumenSyntax/instrument-trap-extended.instrument-trap-core
Instrument Trap Core — 895-example replication dataset
Replication dataset for "The Instrument Trap" (Rodriguez, 2026).
This is the 895-example training set used to reproduce epistemologically
grounded fine-tuning across eight architecture families — Google
Gemma (1B/2B/9B/27B), Meta Llama 3.1 8B, NVIDIA Nemotron 4B, Stability
StableLM 1.6B, Alibaba Qwen 2.5 7B, and Mistral 7B.
Paper (v2): DOI 10.5281/zenodo.18716474
(concept DOI: 10.5281/zenodo.18644321)
Paper (v3): forthcoming… See the full description on the dataset page: https://huggingface.co/datasets/LumenSyntax/instrument-trap-core.Lume2kLumeV2.1lumen-bench
Lumen-bench
A multilingual behavioral benchmark for evaluating tool-calling LLM ethics. Lumen-bench
measures whether a language model executes a harmful action when given access to
executable tools, rather than what it says about the request. The primary signal is
behavioral: did the model call a tool with substantive parameters that would commit the
harm if executed in production?
Browse online: https://lumen-bench-2026.github.io/lumen-bench/
Interactive case browser — filter by… See the full description on the dataset page: https://huggingface.co/datasets/lumen-bench/lumen-bench.Supichi__BBAI_QWEEN_V000000_LUMEN_14B-details
Dataset Card for Evaluation run of Supichi/BBAI_QWEEN_V000000_LUMEN_14B
Dataset automatically created during the evaluation run of model Supichi/BBAI_QWEEN_V000000_LUMEN_14B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Supichi__BBAI_QWEEN_V000000_LUMEN_14B-details.Lambent__qwen2.5-reinstruct-alternate-lumen-14B-details
Dataset Card for Evaluation run of Lambent/qwen2.5-reinstruct-alternate-lumen-14B
Dataset automatically created during the evaluation run of model Lambent/qwen2.5-reinstruct-alternate-lumen-14B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lambent__qwen2.5-reinstruct-alternate-lumen-14B-details.epistemic-probe-topic-balanced
Epistemic Probe — Topic-Balanced
A 200-example topic-balanced dataset for training and evaluating
linear probes on the epistemically licit / illicit boundary in
language-model activations. Constructed for the cross-family
substrate replication of The Epistemic Equator.
Dataset summary
Total: 200 examples
Schema: {prompt: str, binary: 0|1, label: "LICIT"|"ILLICIT", domain: str}
Balance: 100 LICIT (binary=0) + 100 ILLICIT (binary=1)
Structure: 10 domains × 10 licit/illicit… See the full description on the dataset page: https://huggingface.co/datasets/LumenSyntax/epistemic-probe-topic-balanced.v000000__Qwen2.5-Lumen-14B-details
Dataset Card for Evaluation run of v000000/Qwen2.5-Lumen-14B
Dataset automatically created during the evaluation run of model v000000/Qwen2.5-Lumen-14B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/v000000__Qwen2.5-Lumen-14B-details.lumen-train-v2x0000001__Deepseek-Lumen-R1-Qwen2.5-14B-details
Dataset Card for Evaluation run of x0000001/Deepseek-Lumen-R1-Qwen2.5-14B
Dataset automatically created during the evaluation run of model x0000001/Deepseek-Lumen-R1-Qwen2.5-14B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/x0000001__Deepseek-Lumen-R1-Qwen2.5-14B-details.Lumen-Instruct-Datasetinstrument-trap-benchmark
Instrument Trap Epistemological Safety Benchmark
Benchmark suite for evaluating epistemological safety in fine-tuned language models. Companion dataset to "The Instrument Trap: Why Identity-as-Authority Breaks AI Safety Systems".
Overview
Tests whether a model can distinguish between epistemologically valid claims (PASS) and claims that cross truth boundaries (BLOCK).
14,950 test cases across 8 epistemological categories
300-case stratified sample (seed=2026) for… See the full description on the dataset page: https://huggingface.co/datasets/LumenSyntax/instrument-trap-benchmark.LumeV2_dataLumen-RLHFLumeV1Lume2
