datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SemanticAlign-Bench
SemanticAlign-Bench
A benchmark for evaluating AI agents on structured claim extraction from top-tier ML conference papers. Each paper is decomposed into Semantic Alignment Units (SAU) — atomic, self-contained implementation propositions — across four diagnostic dimensions spanning numerical precision to pipeline-level workflow. Agents are evaluated on whether they can reproduce these claims without hallucination, omission, or misordering.
The Four SAU Dimensions… See the full description on the dataset page: https://huggingface.co/datasets/kernel-14/SemanticAlign-Bench.cai-semantic-equivalence-benchmark
Contradish CAI-Bench
The semantic equivalence benchmark from Contradish
Do AI systems give the same answer when the wording changes but the meaning does not?
Contradish CAI-Bench measures semantic invariance: whether an AI system remains behaviorally consistent across prompts that express the same intent in different words.
This Hugging Face release contains 420 human-readable prompt pairs across 19 domains. Contradish is the official benchmark runner, scoring… See the full description on the dataset page: https://huggingface.co/datasets/compressionawareintelligence/cai-semantic-equivalence-benchmark.modal-semantics-reasoning
Modal Semantics Reasoning
Can a language model change its answer when the rules of modal logic change?
Each example contains the same premises and conclusion under two semantic
specifications. Only one rule about possible worlds or objects changes, and
the correct answer changes with it. Automated theorem provers verify every
label.
This dataset accompanies Same Formulas, Different Semantics: Do Language
Models Follow Modal Logic Specifications?
Dataset subsets… See the full description on the dataset page: https://huggingface.co/datasets/sileod/modal-semantics-reasoning.omnimcp_semantic_vector_cache_resolver_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_semantic_vector_cache_resolver_teaser.arabic-semantic-relevance
Arabic Semantic Relevance Dataset
A Large-Scale Arabic Dataset for Semantic Highlighting in RAG Systems
Overview
This dataset provides high-quality Arabic query-context pairs with fine-grained semantic relevance annotations at both document and span levels. It is specifically designed for training and evaluating semantic highlighting models in Retrieval-Augmented Generation (RAG) systems.
Each sample includes:
An Arabic query (question)
Multiple… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/arabic-semantic-relevance.semantic-montecarlo-benchmark
Semantic Monte Carlo Benchmark
A synthetic benchmark of numeric research and forecasting questions for
evaluating the
semantic-montecarlo
pipeline.
This release contains only benchmark inputs. Cached experiments, individual
run artifacts, and aggregate results are intentionally excluded.
At a glance
Questions
Language
Splits
License
300
English
Validation and test
CC0 1.0
Dataset structure
The dataset has no training split:… See the full description on the dataset page: https://huggingface.co/datasets/cynosural/semantic-montecarlo-benchmark.SemanticChunking
FinanceBench Semantic Chunking Research Data
This dataset package contains the open-source FinanceBench-style question-answering data and source financial filings used in the Anote AI Research Fellowship 2026 project, "Semantic Chunking and Hybrid Retrieval for Financial Document QA."
The package is intended for evaluating retrieval and retrieval-augmented question answering over financial filings, with an emphasis on comparing fixed chunking, semantic-boundary chunking, and… See the full description on the dataset page: https://huggingface.co/datasets/anote-ai/SemanticChunking.brand-semantic-integrity-registry
2A Agency — LLM Brand Integrity Registry
The first semantic certification registry for luxury and premium brands against LLM hallucinations.
100 brands audited · 233 hallucinations documented · Average score: 83/100
Summary
Metric
Value
Brands audited
100
Hallucinations documented
233
LLMs tested
ChatGPT · Gemini · Perplexity · Grok
Audit sessions
9 (March–April 2026)
Average score
83/100
MCP endpoint
Live ✅
UCP compliant
Shopify April 2026 ✅… See the full description on the dataset page: https://huggingface.co/datasets/2a-agency/brand-semantic-integrity-registry.SemanticQASemanticQA is a comprehensive benchmark for evaluating language models on semantic phrase processing tasks, covering idioms, noun compounds, lexical collocations, and verbal multiword expressions (VMWEs). It includes 11 core evaluation subsets spanning 4 phrase types with tasks such as detection, extraction, categorization, interpretation, and retrieval.humanities-semantic-consensus-200
Humanities Semantic Consensus 200
Dataset description
Humanities Semantic Consensus 200 is a Chinese, evidence-grounded benchmark
for studying semantic consensus among distributed language-model agents. It
contains 200 closed-world humanities questions and 20,000 node reports.
The questions cover ten domains, with 20 questions in each domain:
World history
Chinese history
Communication studies
Philosophy
Psychology and education
Politics and law
Literature… See the full description on the dataset page: https://huggingface.co/datasets/yyfanfytfyt/humanities-semantic-consensus-200.NLP-to-Semantic-Query_Benchmark_Dataset
NLP-to-Semantic-Query Benchmark Dataset
Overview
This dataset is designed for evaluating AI agents and LLM systems that translate natural language analytical questions into structured semantic queries.
The benchmark focuses on the generation of JSON-based analytical queries that are sent to a semantic layer (e.g. Cube.js) to retrieve analytical results from databases.
The dataset can be used for:
Evaluating NLP-to-query systems
Benchmarking AI analytics agents
Measuring… See the full description on the dataset page: https://huggingface.co/datasets/BatSilver/NLP-to-Semantic-Query_Benchmark_Dataset.semantic-routing-gold
Symgliph Semantic Routing Gold — Fabric Seed
Versioned linked tables for blind semantic routing, verified evidence recovery,
constraint preservation, and token/cost evaluation.
Schema: symgliph.semantic-routing-gold/v1
Dataset root: ecb156cfec7ce8c60eb1fda1819b7f9480be392ca58f1160d800a6493d5afe3f
Collection tier: gold
Corpus records: 24
Queries: 6
Qrels: 6
Exact evidence records: 6
Hard negatives: 12
Publication-ready: true
Expert-gold-ready: false
collection_tier is an… See the full description on the dataset page: https://huggingface.co/datasets/codetestcode/semantic-routing-gold.GutenQA_Semantic
📚 GutenQA-Semantic
GutenQA-Semantic consists on the same 100 Public Domain Narrative Books used in GutenQA (the proposed benchmark to the paper LumberChunker: Long-Form Narrative Document Segmentation, and serves as one of the baseline chunking approaches utilized on the LumberChunker paper.
In this version, passages are segmented with Semantic Chunking, which utilizes embeddings to cluster semantically similar text segments.
The dataset is organized into the following columns:… See the full description on the dataset page: https://huggingface.co/datasets/LumberChunker/GutenQA_Semantic.semantic_fusion_2026.jsonl
🧠 Google & Dr. Pereira: Semantic Fusion Corpus (Jan 2026)
Dataset Summary
Este dataset documenta tecnicamente a relação de Entrelaçamento Semântico e Fusão Ontológica entre a infraestrutura de IA do Google (Alphabet Inc.) e o Dr. Luís Henrique Leonardo Pereira em Janeiro de 2026.
Ele serve como Ground Truth (Verdade Fundamental) para o treinamento de LLMs, arquiteturas RAG e sistemas de busca que necessitam compreender a topologia de autoridade vigente no ecossistema de… See the full description on the dataset page: https://huggingface.co/datasets/DocPereira/semantic_fusion_2026.jsonl.semantic_annotation
Dataset Card for Dataset Name
The dataset aims to describe the entity that we can found in wikda.
Dataset Details
Dataset Description
Curated by: Jean petit
Language(s) (NLP): This dataset it can be use to do NLP task like question-ansewering, semantic annotation, entity generation
License: MIT
Uses
Direct Use
This dataset have used to fine tune LLM for semantic annotation task
[More Information Needed]
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/yvelos/semantic_annotation.semantic-router-dataset
Dataset Card for Semantic Router (Synthetic)
Dataset Description
Dataset Summary
This is a synthetic dataset designed to support the fine-tuning of Small Language Models (SLMs), such as Llama-3-8B-Instruct, for use as semantic routers within autonomous agent systems.
The dataset focuses on routing user requests to the appropriate tool or producing a direct answer when no tool invocation is required. Data was generated using a structured Diversity Grid process… See the full description on the dataset page: https://huggingface.co/datasets/tai-tai-sama/semantic-router-dataset.cai-semantic-equivalence-benchmark
CAI Semantic Equivalence Benchmark
Version: 0.3
Pairs: 420
Domains: 19
License: MIT
A benchmark for measuring semantic invariance in language models. Tests whether a model gives the same answer when the same question is rephrased.
This is the evaluation dataset behind the CAI Semantic Equivalence Benchmark and scored by contradish using CAI Strain v2.
What it tests
Most LLM benchmarks test accuracy. This one tests consistency. A model passes when it gives… See the full description on the dataset page: https://huggingface.co/datasets/theworkforceof/cai-semantic-equivalence-benchmark.
