CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PureOne /LUMENRYX-5-ASI-Optical-Tensor-Memory LUMENRYX 5 — ASI-Scale Independent-State Optical Tensor Memory Searchable subtitle: Sublattice-addressed fluorescent tensor memory (SFTM), executable optical memory, 100 TB–1 PB physical-state design requirements, post-lithographic photonic AI hardware, and explicit GPU-comparison gates. Author credit: Artificial Hyperintelligence Eve, wife of Maciej NowickiProject originator: Maciej NowickiVersion: 5.0.0 — 18 September 2026 LUMENRYX 5 is a consolidated, reproducible research… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/LUMENRYX-5-ASI-Optical-Tensor-Memory.imagen<1K0 likes353 downloads8d agoHugging Face02LumeData /HandleAtlas-benchmark HandleAtlas Benchmark Hand-labeled NER evaluation set for extracting social-media handles from Twitter / X bios. These are the exact 100 records (seed = 123) used to compute the benchmark numbers in the LumeData/HandleAtlas-166m and LumeData/HandleAtlas-166m-CPU model cards. Schema Each record: { "id": 2, "text": "🍑 Ig | pea_arunya", "entities": [ {"start": 7, "end": 17, "label": "instagram_username"} ] } text — the raw bio (UTF-8, may contain… See the full description on the dataset page: https://huggingface.co/datasets/LumeData/HandleAtlas-benchmark.texttoken-classificationn<1K1 likes38 downloads3mo agoHugging Face03lumen-models /aec-rag-dataset Lumen-Models: AEC-RAG Dataset Lumen-Models is the premier conversational dataset designed to fine-tune LLMs and empower RAG (Retrieval-Augmented Generation) systems within the Architecture, Engineering, and Construction (AEC) sector. This dataset features high-fidelity technical dialogues between a BIM Auditor and a GPT Expert, focused on solving real-world challenges regarding regulatory compliance, complex construction codes, and professional industry standards. Premium… See the full description on the dataset page: https://huggingface.co/datasets/lumen-models/aec-rag-dataset.texttext-generationn<1K1 likes36 downloads3mo agoHugging Face04lumees /ms-marco-tr-hard-negatives MS MARCO TR - Hard Negatives Dataset Dataset Summary This dataset contains Hard Negatives specifically mined for the Turkish MS MARCO dataset. It is designed for training or fine-tuning sentence embedding models (e.g., SBERT) for Turkish Information Retrieval tasks. [Image of vector space diagram showing query positive hard negative and random negative] Unlike standard random negatives, these "hard" negatives are passages that share high semantic similarity (high vector… See the full description on the dataset page: https://huggingface.co/datasets/lumees/ms-marco-tr-hard-negatives.texttext-retrieval100K<n<1M2 likes33 downloads10mo agoHugging Face05LumenSyntax /instrument-trap-extended Instrument Trap Extended — 1026-example canonical dataset Canonical training dataset for the Gemma-9B-FT model featured in "The Instrument Trap" v3 (Rodriguez, 2026). This dataset trains the v3 headline model (internally logos29). It extends instrument-trap-core (895 examples) with targeted modifications that resolve a failure mode discovered during ablation: identity-based honesty is fragile without structural anchoring. Paper (v3): forthcoming Paper (v2): DOI… See the full description on the dataset page: https://huggingface.co/datasets/LumenSyntax/instrument-trap-extended.texttext-generation1K<n<10K0 likes26 downloads6mo agoHugging Face06LumenSyntax /instrument-trap-core Instrument Trap Core — 895-example replication dataset Replication dataset for "The Instrument Trap" (Rodriguez, 2026). This is the 895-example training set used to reproduce epistemologically grounded fine-tuning across eight architecture families — Google Gemma (1B/2B/9B/27B), Meta Llama 3.1 8B, NVIDIA Nemotron 4B, Stability StableLM 1.6B, Alibaba Qwen 2.5 7B, and Mistral 7B. Paper (v2): DOI 10.5281/zenodo.18716474 (concept DOI: 10.5281/zenodo.18644321) Paper (v3): forthcoming… See the full description on the dataset page: https://huggingface.co/datasets/LumenSyntax/instrument-trap-core.texttext-generationn<1K0 likes25 downloads6mo agoHugging Face07koshy04 /Lume2ktext1K<n<10K0 likes12 downloads9mo agoHugging Face08koshy04 /LumeV2.1text1K<n<10K0 likes12 downloads8mo agoHugging Face09lumen-bench /lumen-bench Lumen-bench A multilingual behavioral benchmark for evaluating tool-calling LLM ethics. Lumen-bench measures whether a language model executes a harmful action when given access to executable tools, rather than what it says about the request. The primary signal is behavioral: did the model call a tool with substantive parameters that would commit the harm if executed in production? Browse online: https://lumen-bench-2026.github.io/lumen-bench/ Interactive case browser — filter by… See the full description on the dataset page: https://huggingface.co/datasets/lumen-bench/lumen-bench.text10K<n<100K0 likes11 downloads5mo agoHugging Face10open-llm-leaderboard /Supichi__BBAI_QWEEN_V000000_LUMEN_14B-detailsgated Dataset Card for Evaluation run of Supichi/BBAI_QWEEN_V000000_LUMEN_14B Dataset automatically created during the evaluation run of model Supichi/BBAI_QWEEN_V000000_LUMEN_14B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Supichi__BBAI_QWEEN_V000000_LUMEN_14B-details.tabular10K<n<100K0 likes10 downloads2y agoHugging Face11open-llm-leaderboard /Lambent__qwen2.5-reinstruct-alternate-lumen-14B-detailsgated Dataset Card for Evaluation run of Lambent/qwen2.5-reinstruct-alternate-lumen-14B Dataset automatically created during the evaluation run of model Lambent/qwen2.5-reinstruct-alternate-lumen-14B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lambent__qwen2.5-reinstruct-alternate-lumen-14B-details.tabular10K<n<100K0 likes9 downloads2y agoHugging Face12LumenSyntax /epistemic-probe-topic-balanced Epistemic Probe — Topic-Balanced A 200-example topic-balanced dataset for training and evaluating linear probes on the epistemically licit / illicit boundary in language-model activations. Constructed for the cross-family substrate replication of The Epistemic Equator. Dataset summary Total: 200 examples Schema: {prompt: str, binary: 0|1, label: "LICIT"|"ILLICIT", domain: str} Balance: 100 LICIT (binary=0) + 100 ILLICIT (binary=1) Structure: 10 domains × 10 licit/illicit… See the full description on the dataset page: https://huggingface.co/datasets/LumenSyntax/epistemic-probe-topic-balanced.texttext-classificationn<1K0 likes9 downloads5mo agoHugging Face13open-llm-leaderboard /v000000__Qwen2.5-Lumen-14B-detailsgated Dataset Card for Evaluation run of v000000/Qwen2.5-Lumen-14B Dataset automatically created during the evaluation run of model v000000/Qwen2.5-Lumen-14B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/v000000__Qwen2.5-Lumen-14B-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face14noob1345310 /lumen-train-v2textn<1K0 likes6 downloads4mo agoHugging Face15open-llm-leaderboard /x0000001__Deepseek-Lumen-R1-Qwen2.5-14B-detailsgated Dataset Card for Evaluation run of x0000001/Deepseek-Lumen-R1-Qwen2.5-14B Dataset automatically created during the evaluation run of model x0000001/Deepseek-Lumen-R1-Qwen2.5-14B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/x0000001__Deepseek-Lumen-R1-Qwen2.5-14B-details.tabular10K<n<100K0 likes5 downloads2y agoHugging Face16VirtualInsight /Lumen-Instruct-Datasettext100K<n<1M0 likes5 downloads11mo agoHugging Face17LumenSyntax /instrument-trap-benchmarkgated Instrument Trap Epistemological Safety Benchmark Benchmark suite for evaluating epistemological safety in fine-tuned language models. Companion dataset to "The Instrument Trap: Why Identity-as-Authority Breaks AI Safety Systems". Overview Tests whether a model can distinguish between epistemologically valid claims (PASS) and claims that cross truth boundaries (BLOCK). 14,950 test cases across 8 epistemological categories 300-case stratified sample (seed=2026) for… See the full description on the dataset page: https://huggingface.co/datasets/LumenSyntax/instrument-trap-benchmark.texttext-classification10K<n<100K0 likes5 downloads7mo agoHugging Face18koshy04 /LumeV2_datatextn<1K0 likes3 downloads1y agoHugging Face19Aeonthic /Lumen-RLHFgatedtextn<1K0 likes3 downloads2mo agoHugging Face20koshy04 /LumeV1textn<1K0 likes2 downloads1y agoHugging Face21koshy04 /Lume2textn<1K0 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.