datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cql_gen-browsecomp_plus_qa_gen-oai_gpt5_low-multihop_2-v3187wiki-multihop-qa-500k
wiki-multihop-qa-500k
500,000 synthetic multi-hop QA pairs generated from Wikipedia.
Built for training the think-in-silence latent reasoning model — a model that reasons entirely in vector space without generating chain-of-thought tokens.
Why This Dataset Exists
Most publicly available QA datasets have two problems for reasoning research:
Too small. HotpotQA has 113K samples. StrategyQA has 2.8K. Not enough diversity to train a generalizable reasoning module.
Too many… See the full description on the dataset page: https://huggingface.co/datasets/rajat5039/wiki-multihop-qa-500k.nor_agriculture_multi_hop_questions_bench
Nor Agriculture Multi Hop Questions Bench
This work is related to the project in adapting LLM to answer questions about Norwegian Agriculture in Norwegian.
The dataset was generated using YourBench (v0.9.0), an open-source framework for generating domain-specific benchmarks from document collections.
It needs further cleaning, verification in regards to citations and validation for diversity and topics coverage.
Also, it can play a role as a prove of concept for generating a… See the full description on the dataset page: https://huggingface.co/datasets/norjordAI/nor_agriculture_multi_hop_questions_bench.multihopRAG-1024multihopRAG-256Blind-Spots-of-Frontier-Models_Multihop-blindspots-qwen35
Qwen3.5-2B-Base — Multi-Hop Reasoning Blind Spots
An evaluation dataset probing 18 Knowledge Graph-style reasoning tasks on
Qwen/Qwen3.5-2B-Base, tested in its
raw base (pre-training) form with no external graph attached. The dataset covers
parametric memory (probes 1–10, no passage provided), standard grounded reasoning
(probes 11–15, source passage included), and advanced grounded reasoning
(probes 16–18, passage provided but requiring implicit inference or contradiction… See the full description on the dataset page: https://huggingface.co/datasets/chayma-rhaiem/Blind-Spots-of-Frontier-Models_Multihop-blindspots-qwen35.multihopRAG-2048multihopRAGmultihopRAG-512rb_last-cql-mc_oai_gpt5_med-ec_browsecomp_plus_search_gpt5_multihop-ds_train-spp4-gbs1-s0-r_1Agentic-Search-MultiHopAgentic-Search-MultiHop-Easy-Set
