datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
service_public_pro-full-documentsus-pro-se
US Pro Se — what the courts themselves tell people who have no lawyer
Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them.
12,103 documents from 19 state court systems: 2,835
self-help guide pages, 727 instruction documents and
8,541 forms, 119,978,550 characters of text, each… See the full description on the dataset page: https://huggingface.co/datasets/docketx/us-pro-se.mmlu_pro_medical
MMLU-Pro Medical (Test)
A 1535-sample evaluation subset drawn from MMLU-Pro, formatted for LLM evaluation pipelines.
Dataset Summary
Split
Samples
Subjects
test
1535
2
Subjects: health, biology.
Schema
Column
Type
Description
prompt
list[dict]
Chat-format prompt with few-shot examples (role / content).
data_source
string
Subject label, e.g. mmlu_pro/math.
extra_info
dict
Question metadata: question, options, answer, answer_index… See the full description on the dataset page: https://huggingface.co/datasets/SeanWang0027/mmlu_pro_medical.mmlu_pro_subset_test
MMLU-Pro Subset (Test)
A 700-sample evaluation subset drawn from MMLU-Pro, formatted for LLM evaluation pipelines.
Dataset Summary
Split
Samples
Subjects
test
700
14
Subjects: math, health, physics, business, biology, chemistry, computer science, economics, engineering, philosophy, other, history, psychology, law.
Schema
Column
Type
Description
prompt
list[dict]
Chat-format prompt with few-shot examples (role / content).… See the full description on the dataset page: https://huggingface.co/datasets/SeanWang0027/mmlu_pro_subset_test.
