datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
semantic-montecarlo-benchmark
Semantic Monte Carlo Benchmark
A synthetic benchmark of numeric research and forecasting questions for
evaluating the
semantic-montecarlo
pipeline.
This release contains only benchmark inputs. Cached experiments, individual
run artifacts, and aggregate results are intentionally excluded.
At a glance
Questions
Language
Splits
License
300
English
Validation and test
CC0 1.0
Dataset structure
The dataset has no training split:… See the full description on the dataset page: https://huggingface.co/datasets/cynosural/semantic-montecarlo-benchmark.SemanticChunking
FinanceBench Semantic Chunking Research Data
This dataset package contains the open-source FinanceBench-style question-answering data and source financial filings used in the Anote AI Research Fellowship 2026 project, "Semantic Chunking and Hybrid Retrieval for Financial Document QA."
The package is intended for evaluating retrieval and retrieval-augmented question answering over financial filings, with an emphasis on comparing fixed chunking, semantic-boundary chunking, and… See the full description on the dataset page: https://huggingface.co/datasets/anote-ai/SemanticChunking.semantic-routing-gold
Symgliph Semantic Routing Gold — Fabric Seed
Versioned linked tables for blind semantic routing, verified evidence recovery,
constraint preservation, and token/cost evaluation.
Schema: symgliph.semantic-routing-gold/v1
Dataset root: ecb156cfec7ce8c60eb1fda1819b7f9480be392ca58f1160d800a6493d5afe3f
Collection tier: gold
Corpus records: 24
Queries: 6
Qrels: 6
Exact evidence records: 6
Hard negatives: 12
Publication-ready: true
Expert-gold-ready: false
collection_tier is an… See the full description on the dataset page: https://huggingface.co/datasets/codetestcode/semantic-routing-gold.GutenQA_Semantic
📚 GutenQA-Semantic
GutenQA-Semantic consists on the same 100 Public Domain Narrative Books used in GutenQA (the proposed benchmark to the paper LumberChunker: Long-Form Narrative Document Segmentation, and serves as one of the baseline chunking approaches utilized on the LumberChunker paper.
In this version, passages are segmented with Semantic Chunking, which utilizes embeddings to cluster semantically similar text segments.
The dataset is organized into the following columns:… See the full description on the dataset page: https://huggingface.co/datasets/LumberChunker/GutenQA_Semantic.
