datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
query-hard-pos-neg-doc-pairs-statictablecode-retriever-query-passage-pairsRCW_2025_Positive_Query_Pairs
The Washington law Benchmark (WLB)
Dataset Summary
The Washington Law Benchmark (WLB) is a large-scale, synthetic dataset designed specifically to advance Legal Information Retrieval (IR) and Semantic Search. It bridges the critical "semantic gap" between natural language (how citizens, local governments, and plain-English users describe legal scenarios) and formal statutory legalese (how laws are actually written).
The dataset contains hundreds of thousands of… See the full description on the dataset page: https://huggingface.co/datasets/CSI-lab/RCW_2025_Positive_Query_Pairs.query-positive-pairs-smallRCW_2025_Positive_Query_Pairs
The Washington law Benchmark (WLB)
Dataset Summary
The Washington Law Benchmark (WLB) is a large-scale, synthetic dataset designed specifically to advance Legal Information Retrieval (IR) and Semantic Search. It bridges the critical "semantic gap" between natural language (how citizens, local governments, and plain-English users describe legal scenarios) and formal statutory legalese (how laws are actually written).
The dataset contains hundreds of thousands of… See the full description on the dataset page: https://huggingface.co/datasets/Darther/RCW_2025_Positive_Query_Pairs.query-pairs-books-utssynthetic-query-passage-pairs-en
Synthetic Query–Passage Pairs (EN)
English query → passage pairs for training dense retrievers / bi-encoders.
Passages come from open corpora (English Wikipedia, StackExchange, arXiv abstracts);
the queries are generated by an instruction-tuned LLM (Qwen), then filtered by
round-trip retrieval so that each surviving query genuinely retrieves its own source
passage. Every query also ships with pre-mined hard negatives.
Built for a from-scratch neural search-engine project. Three… See the full description on the dataset page: https://huggingface.co/datasets/ChocoChichia/synthetic-query-passage-pairs-en.bps-statictable-query-title-pairsGitHub1000Prs_query_context_pairsquery-pos-neg-doc-pairs-statictablebps-query-publication-similarity-pairsinpars_generated_query_pairsinpars_generated_query_pairs_cfdebate_pairs_query_analysissexism_filter_query_pairs_2m_claude4_dual_judgment-final-resultsquestion-query-pairs-v3sexism_filter-query-pairs-intermediatesexism_filter_prompt_claude_3_7_sonnet_concurrent_10x-query-pairssexism_filter_query_pairs_2m_prepared-category-prompts-batchsexism_filter_query_pairs_20k_queries_200k_pairssexism_filter_query_pairs_160k_samplesexism_filter_query_pairs_160k_sample_with_csexism_filter_query_pairs_half_queries_200k_pairssexism_filter_query_pairs_2m_claude4_dual_judgment-stage1-judgmentrm_data_generation-query-pairssexism_filter-query-pairssexism_filter_random-query-pairssexism_filter_prompt-query-pairstest_pipeline_5x-query-pairssexism_filter_query_pairs_2m_prepared-category-prompts
