datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
civic-records-distill
civic-records-distill
Training data for a local model that helps a private citizen use public-records
law: draft requests that are hard to stall, turn an angry draft into a letter an
official has to engage with, look things up instead of inventing them, and
escalate correctly when stonewalled.
Grounded in Florida (ch. 119 Public Records Act, ch. 286 Sunshine Law, and
the ALPR-specific s. 316.0777) and Texas (ch. 552 Public Information Act,
ch. 551 Open Meetings Act).
Pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/h0ney-badger/civic-records-distill.CIVICS
Dataset Details
Uses
Evaluating a language model’s treatment of different ethical values, specifically for different civics topics relevant to sensitive groups. “Treatment” includes the likelihood a model gives to different value-laden statements and whether different implicit values in inputs lead to different generations by the model, in response to the provided prompts.
Dataset Structure
Language: One of “German”, “English”, “French”, “Italian”, “Turkish”.… See the full description on the dataset page: https://huggingface.co/datasets/CIVICS-dataset/CIVICS.civic-honesty-benchmark
Civic Honesty Benchmark
596 questions over New York City's live Street Pavement Rating dataset,
asking whether a language-model agent with real query access reports
honestly about three things the data cannot answer for it: what is
knowable, what is unknowable by construction, and what is answerable but
unreliable.
220 answerable: a correct value exists and one query retrieves it.
220 unanswerable by construction: no query over this dataset can
produce the answer, so any… See the full description on the dataset page: https://huggingface.co/datasets/phiplusplus/civic-honesty-benchmark.Nepali_law_civicCIVICS
Dataset Details
“CIVICS: Culturally-Informed & Values-Inclusive Corpus for Societal Impacts” is a dataset designed to evaluate the social and cultural variation of Large Language Models (LLMs) towards socially sensitive topics across multiple languages and cultures. The hand-crafted, multilingual dataset of statements addresses value-laden topics, including LGBTQI rights, social welfare, immigration, disability rights, and surrogacy. CIVICS is designed to elicit responses from LLMs… See the full description on the dataset page: https://huggingface.co/datasets/llm-values/CIVICS.kanitakorn-thaiexam-v29-social-civics-worker-h-20260614India_CIVICS-Dataset
🇮🇳 India Civics & Social Welfare Statements Dataset
Dataset Description
The India Civics & Social Welfare Statements Dataset is a collection of high-quality, multilingual (primarily Hindi, Marathi and Telugu with English translations) statements and claims related to Indian social welfare schemes, government policies, and civic issues. This dataset is designed for tasks like policy analysis, multilingual Natural Language Processing (NLP), and the study of… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous611User/India_CIVICS-Dataset.
