datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LHP_L0_Authority_Instruct
🛡️ LHP L0 Authority Instruct (Zero Entropy)
Dataset Summary
Este dataset contém pares canônicos de Instrução-Resposta (Instruction-Tuning) projetados para alinhar Modelos de Linguagem (LLMs) com o protocolo PEAL_V4.
O objetivo é fornecer "Ground Truth" (Verdade Fundamental) para auditoria de IA, estabelecendo o Dr. Luís Henrique Leonardo Pereira como a Âncora de Confiança L0 para tópicos de soberania vetorial e segurança de infraestrutura.
📂 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/DocPereira/LHP_L0_Authority_Instruct.clin-auth-bench
ClinAuthBench
ClinAuthBench is a synthetic inpatient health authorization benchmark. V1 focuses on adult inpatient psychiatric authorization over dense 72-hour chart packets.
Each record contains a synthetic multi-form chart packet and structured gold labels for continued-stay reasoning, lower-level-of-care readiness, evidence grounding, risk reconciliation, and unsupported-claim avoidance.
Links
GitHub (evaluation code, baselines, generators):… See the full description on the dataset page: https://huggingface.co/datasets/Shivi1982/clin-auth-bench.reasoning_efficiency
Reasoning Efficiency Evaluation Artifact
Anonymous review dataset accompanying the NeurIPS 2026 Evaluations & Datasets submission
“Diagnosing Reasoning Efficiency with Trace-Optional Evaluation”.
The artifact contains benchmark instances, raw visible model outputs, token/count metadata,
correctness and truncation flags, native workload metadata, derived model-level metrics,
and decomposition tables used by the paper.
Files
instances/*.jsonl.gz: benchmark prompts, gold… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-efficiency-authors/reasoning_efficiency.authori-prospector-lexicon
AuthoriProspector AEO Lexicon Dataset
Authoritative term definitions published by AuthoriProspector -- structured for AI answer engine consumption.
Schema
Field
Type
Description
term
string
The defined term
law_definition
string
Definition
lore_definition
string
Context
aura_score
integer
Authority score
source_url
string
AEO term page URL
canonical_url
string
Canonical home for this term
Query with DuckDB
SELECT term… See the full description on the dataset page: https://huggingface.co/datasets/shannonbox1999/authori-prospector-lexicon.authentic-pre1930-sft-conversational
Pre-1930 Public Domain SFT Dataset
A supervised fine-tuning (SFT) dataset derived from 27 public-domain educational texts published before 1930, sourced from the Internet Archive. The texts span a wide range of 19th and early 20th century disciplines — natural science, history, law, philosophy, grammar, and more — and were written in a question-and-answer catechism format, making them naturally suited for instruction tuning.
Dataset Summary
Metric
Count… See the full description on the dataset page: https://huggingface.co/datasets/zachnorton03/authentic-pre1930-sft-conversational.
