CoolFace
11 results

search-queries

hotchpotch /fineweb-ir-simulated-search-queries fineweb-ir-simulated-search-queries An English web-retrieval dataset built by generating simulated search queries for FineWeb-style positive target documents. This dataset contains English query-document pairs derived from HuggingFaceFW/fineweb-edu. Each row is designed so that the associated document is a positive retrieval target for the generated query. The queries were generated with a Qwen3.5-35B-A3B model quantized to NVFP4 and served through OpenAI-compatible vLLM serve… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/fineweb-ir-simulated-search-queries.text1M<n<10M0 likes126 downloads5mo agoHugging Facehotchpotch /arxiv-ir-simulated-search-queries arxiv-ir-simulated-search-queries An arXiv retrieval dataset with more than 2.8 million simulated specialist search queries and paper-level positive targets. This dataset contains 2,875,637 query-document pairs derived from arXiv title-and-abstract records. Each row is designed so that the associated arXiv paper record is a positive retrieval target for the generated query. The queries were generated with a Qwen3.5-35B-A3B model quantized to NVFP4, and were intentionally… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/arxiv-ir-simulated-search-queries.text1M<n<10M0 likes118 downloads6mo agoHugging Facehotchpotch /wikipedia-english-ir-simulated-search-queries wikipedia-english-ir-simulated-search-queries An English Wikipedia retrieval dataset with more than 29 million simulated search queries and paragraph-level positive targets. This dataset contains 29,366,101 English query-document pairs derived from Wikipedia. Each row is designed so that the associated Wikipedia paragraph is a positive retrieval target for the generated query. The queries were generated with a Qwen3.5-35B-A3B model quantized to NVFP4, and were intentionally… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/wikipedia-english-ir-simulated-search-queries.text10M<n<100M0 likes102 downloads4mo agoHugging Facehotchpotch /pubmed-abstract-ir-simulated-search-queries pubmed-abstract-ir-simulated-search-queries A PubMed retrieval dataset with simulated specialist search queries and abstract-level positive targets. This dataset contains 2,355,329 query-document pairs derived from PubMed title-and-abstract records. Each row is designed so that the associated PubMed record is a positive retrieval target for the generated query. The queries were generated with a Qwen3.5-35B-A3B model quantized to NVFP4, and were intentionally constructed to… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/pubmed-abstract-ir-simulated-search-queries.text1M<n<10M0 likes81 downloads6mo agoHugging Facehotchpotch /ccnews-ir-simulated-search-queries ccnews-ir-simulated-search-queries An English news-retrieval dataset with 1.84 million simulated search queries paired with positive CC-News-style document targets. This dataset contains 1,839,547 English query-document pairs derived from the English subset of multilingual CC-News. Each row is designed so that the associated news document is a positive retrieval target for the generated query. The queries were generated with a Qwen3.5-35B-A3B model quantized to NVFP4 and… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/ccnews-ir-simulated-search-queries.text1M<n<10M0 likes60 downloads5mo agoHugging Facetrec-product-search /product-search-2023-queriestextn<1K0 likes40 downloads3y agoHugging Face