datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SemanticAlign-Bench
SemanticAlign-Bench
A benchmark for evaluating AI agents on structured claim extraction from top-tier ML conference papers. Each paper is decomposed into Semantic Alignment Units (SAU) — atomic, self-contained implementation propositions — across four diagnostic dimensions spanning numerical precision to pipeline-level workflow. Agents are evaluated on whether they can reproduce these claims without hallucination, omission, or misordering.
The Four SAU Dimensions… See the full description on the dataset page: https://huggingface.co/datasets/kernel-14/SemanticAlign-Bench.Semantic-Entropy-Core-PoC
Moebius-Distillate-v1-PoC
1. Overview
This dataset contains high-density semantic information extracted via the Moebius Operator protocol. Unlike traditional deduplication, our method uses non-orientable topological logic to eliminate logical redundancy while preserving the invariant semantic core of the data.
2. The "Chomsky Emergence" Experiment
We conducted a control experiment to verify the efficiency of this distillate compared to raw text.… See the full description on the dataset page: https://huggingface.co/datasets/0sz1/Semantic-Entropy-Core-PoC.
