datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Law_RCW_Dataset
The Revised Code of Washington (RCW) - 2026 Edition
Dataset Summary
The Revised Code of Washington (RCW) 2026 Edition dataset is a comprehensive, structured corpus containing the full text of all permanent laws in force in the State of Washington as of 2026. This dataset provides the statutory provisions in a standardized, machine-readable format, designed specifically to facilitate legal Natural Language Processing (NLP) research.
This release curates the… See the full description on the dataset page: https://huggingface.co/datasets/CSI-lab/Law_RCW_Dataset.RCW_2025_Positive_Query_Pairs
The Washington law Benchmark (WLB)
Dataset Summary
The Washington Law Benchmark (WLB) is a large-scale, synthetic dataset designed specifically to advance Legal Information Retrieval (IR) and Semantic Search. It bridges the critical "semantic gap" between natural language (how citizens, local governments, and plain-English users describe legal scenarios) and formal statutory legalese (how laws are actually written).
The dataset contains hundreds of thousands of… See the full description on the dataset page: https://huggingface.co/datasets/Darther/RCW_2025_Positive_Query_Pairs.RC_WIKISQL_PHI_3
SQL Query Generation Dataset
Description
This dataset contains SQL query templates derived from natural language questions. It is designed to assist in training and evaluating models that convert natural language into SQL queries. The dataset includes a variety of questions, corresponding SQL table schemas, and the generated SQL queries.
Data Fields
question (string): The natural language question for which a SQL query is generated.
context (string): The SQL… See the full description on the dataset page: https://huggingface.co/datasets/rAMNARAY/RC_WIKISQL_PHI_3.
