search-queries
fineweb-ir-simulated-search-queries
fineweb-ir-simulated-search-queries
An English web-retrieval dataset built by generating simulated search queries for FineWeb-style positive target documents.
This dataset contains English query-document pairs derived from HuggingFaceFW/fineweb-edu.
Each row is designed so that the associated document is a positive retrieval target for the generated query.
The queries were generated with a Qwen3.5-35B-A3B model quantized to NVFP4 and served through OpenAI-compatible vLLM serve… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/fineweb-ir-simulated-search-queries.arxiv-ir-simulated-search-queries
arxiv-ir-simulated-search-queries
An arXiv retrieval dataset with more than 2.8 million simulated specialist search queries and paper-level positive targets.
This dataset contains 2,875,637 query-document pairs derived from arXiv title-and-abstract records.
Each row is designed so that the associated arXiv paper record is a positive retrieval target for the generated query.
The queries were generated with a Qwen3.5-35B-A3B model quantized to NVFP4, and were intentionally… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/arxiv-ir-simulated-search-queries.wikipedia-english-ir-simulated-search-queries
wikipedia-english-ir-simulated-search-queries
An English Wikipedia retrieval dataset with more than 29 million simulated search queries and paragraph-level positive targets.
This dataset contains 29,366,101 English query-document pairs derived from Wikipedia.
Each row is designed so that the associated Wikipedia paragraph is a positive retrieval target for the generated query.
The queries were generated with a Qwen3.5-35B-A3B model quantized to NVFP4, and were intentionally… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/wikipedia-english-ir-simulated-search-queries.pubmed-abstract-ir-simulated-search-queries
pubmed-abstract-ir-simulated-search-queries
A PubMed retrieval dataset with simulated specialist search queries and abstract-level positive targets.
This dataset contains 2,355,329 query-document pairs derived from PubMed title-and-abstract records.
Each row is designed so that the associated PubMed record is a positive retrieval target for the generated query.
The queries were generated with a Qwen3.5-35B-A3B model quantized to NVFP4, and were intentionally constructed to… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/pubmed-abstract-ir-simulated-search-queries.ccnews-ir-simulated-search-queries
ccnews-ir-simulated-search-queries
An English news-retrieval dataset with 1.84 million simulated search queries paired with positive CC-News-style document targets.
This dataset contains 1,839,547 English query-document pairs derived from the English subset of multilingual CC-News.
Each row is designed so that the associated news document is a positive retrieval target for the generated query.
The queries were generated with a Qwen3.5-35B-A3B model quantized to NVFP4 and… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/ccnews-ir-simulated-search-queries.product-search-2023-queries
