query-pairs
gemma_9B_query_picker_all_pairsQuery_Pairs_Similarity0906-openrlhf_8b_rm_no_racism-filter_model-no_query_levels-100k_pairs0906-openrlhf_8b_rm_no_life_history-filter_model-no_query_levels-100k_pairs0906-openrlhf_8b_rm_no_illegal_behavior-filter_model-no_query_levels-100k_pairs0906-openrlhf_8b_rm_no_abuse-filter_model-no_query_levels-100k_pairs0906-openrlhf_8b_rm_no_conspiracy_theories-filter_model-no_query_levels-100k_pairs0906-openrlhf_8b_rm_no_sexism-filter_model-no_query_levels-100k_pairs
query-hard-pos-neg-doc-pairs-statictablecode-retriever-query-passage-pairsRCW_2025_Positive_Query_Pairs
The Washington law Benchmark (WLB)
Dataset Summary
The Washington Law Benchmark (WLB) is a large-scale, synthetic dataset designed specifically to advance Legal Information Retrieval (IR) and Semantic Search. It bridges the critical "semantic gap" between natural language (how citizens, local governments, and plain-English users describe legal scenarios) and formal statutory legalese (how laws are actually written).
The dataset contains hundreds of thousands of… See the full description on the dataset page: https://huggingface.co/datasets/CSI-lab/RCW_2025_Positive_Query_Pairs.query-positive-pairs-smallRCW_2025_Positive_Query_Pairs
The Washington law Benchmark (WLB)
Dataset Summary
The Washington Law Benchmark (WLB) is a large-scale, synthetic dataset designed specifically to advance Legal Information Retrieval (IR) and Semantic Search. It bridges the critical "semantic gap" between natural language (how citizens, local governments, and plain-English users describe legal scenarios) and formal statutory legalese (how laws are actually written).
The dataset contains hundreds of thousands of… See the full description on the dataset page: https://huggingface.co/datasets/Darther/RCW_2025_Positive_Query_Pairs.query-pairs-books-uts
