datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
virginia-woolf-monologue-chunks
Virginia Woolf Monologue Chunks Dataset
This dataset contains 6 semantically chunked text segments derived from a contemporary monologue based on Virginia Woolf's seminal essay "A Room of One's Own" (1929). It comes pre-loaded with vector embeddings from three different models, making it a ready-to-use resource for a variety of NLP tasks.
In addition to the dataset itself, this repository includes a comprehensive embedding analysis, detailed statistics, and 7 visualizations to help… See the full description on the dataset page: https://huggingface.co/datasets/pageman/virginia-woolf-monologue-chunks.Virginia-Statute-QA
Virginia Statute QA
Dataset Summary
Virginia Statute QA is a high-quality, synthetically generated legal question–answer dataset grounded in the Code of Virginia.Each entry consists of:
a natural-language legal question
one or more statute identifiers (section_ids)
a grounded answer citing the relevant statute(s)
The dataset is intended for:
Legal RAG (Retrieval-Augmented Generation)
Statute-grounded QA model evaluation
Legal retrieval linking
Experiments on statutory… See the full description on the dataset page: https://huggingface.co/datasets/dcrodriguez/Virginia-Statute-QA.
