datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tip-civic-data
TIP Civic Data — U.S. Government Accountability Dataset
Nonpartisan civic accountability data from Truth In Polling (501(c)(3) nonprofit).
Includes federal and state legislation, official voting records, citizen approval ratings, and multi-factor vote predictions.
Datasets
File
Description
Rows
federal_bills
Federal legislation (119th Congress) with status; no AI summaries
18,676
state_bills
State legislation with status and status_bucket; no AI… See the full description on the dataset page: https://huggingface.co/datasets/truthinpolling/tip-civic-data.greek_civics_qa
Dataset Card for Greek Civics QA
The Greek Civics QA dataset is a set of 407 question and answer pairs related to Greek highschool courses in civics (Κοινωνική και Πολιτική Αγωγή). The dataset was created by Nikoletta Tsoukala during her 2023 internship at the Institute for Language and Speech Processing/Athena RC, supervised by ILSP Research Associate Vassilis Papavassileiou. The dataset creation process involved mining questions and answers from two civics textbooks used in the… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/greek_civics_qa.civic-records-distill
civic-records-distill
Training data for a local model that helps a private citizen use public-records
law: draft requests that are hard to stall, turn an angry draft into a letter an
official has to engage with, look things up instead of inventing them, and
escalate correctly when stonewalled.
Grounded in Florida (ch. 119 Public Records Act, ch. 286 Sunshine Law, and
the ALPR-specific s. 316.0777) and Texas (ch. 552 Public Information Act,
ch. 551 Open Meetings Act).
Pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/h0ney-badger/civic-records-distill.osm-civic-community-places
OSM Civic & Community Places Visual Dataset
Rows: 397,704
Dataset Description
OSM Civic & Community Places Visual Dataset is a Global wildlife image dataset alternative for geospatial computer vision and image classification focused on OpenStreetMap-mapped civic and community places. The labeled classes are School, Places of Worship, and Nature Reserve, produced by Outerview's query-driven visual matching and label assignment process over OpenStreetMap node… See the full description on the dataset page: https://huggingface.co/datasets/Outerview/osm-civic-community-places.CivicGRQA_Pretrain_Text_v2
Dữ liệu dịch vụ công, pháp luật và sách Tiếng Việt
📘 Tổng quan
Bộ dữ liệu tiếng Việt tổng hợp từ 3 nguồn chính nhằm phục vụ huấn luyện và tinh chỉnh mô hình ngôn ngữ lớn (LLM), đặc biệt cho các ứng dụng hành chính công và văn hóa tại Việt Nam.
🗂 Cấu trúc
Tập huấn luyện (train): 474,559 mẫu — ~2.48 GB
Tập kiểm tra (test): 52,729 mẫu — ~275 MB
Tổng dung lượng sau giải nén: ~2.57 GB
Định dạng: plain text (text field)
🧾 Nguồn dữ liệu
Tài liệu về… See the full description on the dataset page: https://huggingface.co/datasets/thailevann/CivicGRQA_Pretrain_Text_v2.CIVICS
Dataset Details
Uses
Evaluating a language model’s treatment of different ethical values, specifically for different civics topics relevant to sensitive groups. “Treatment” includes the likelihood a model gives to different value-laden statements and whether different implicit values in inputs lead to different generations by the model, in response to the provided prompts.
Dataset Structure
Language: One of “German”, “English”, “French”, “Italian”, “Turkish”.… See the full description on the dataset page: https://huggingface.co/datasets/CIVICS-dataset/CIVICS.civic-honesty-benchmark
Civic Honesty Benchmark
596 questions over New York City's live Street Pavement Rating dataset,
asking whether a language-model agent with real query access reports
honestly about three things the data cannot answer for it: what is
knowable, what is unknowable by construction, and what is answerable but
unreliable.
220 answerable: a correct value exists and one query retrieves it.
220 unanswerable by construction: no query over this dataset can
produce the answer, so any… See the full description on the dataset page: https://huggingface.co/datasets/phiplusplus/civic-honesty-benchmark.auburn-civic-data
Auburn Civic Data
This dataset packages public City of Auburn, Alabama civic GIS data for Ask Auburn and downstream civic-data analysis.
Hugging Face repo: mm-intelligence/auburn-civic-data
Generated: 2026-05-11 18:44 UTC
Tables: 17
Rows: 101,935
Source: City of Auburn public ArcGIS services (data-coa.opendata.arcgis.com and gis.auburnalabama.org)
Tables
Table
Rows
Geometry
Description
city_limits
3
polygon
Auburn corporate-limits boundary polygons. Useful as… See the full description on the dataset page: https://huggingface.co/datasets/mm-intelligence/auburn-civic-data.civicNepali_law_civicCIVICS
Dataset Details
“CIVICS: Culturally-Informed & Values-Inclusive Corpus for Societal Impacts” is a dataset designed to evaluate the social and cultural variation of Large Language Models (LLMs) towards socially sensitive topics across multiple languages and cultures. The hand-crafted, multilingual dataset of statements addresses value-laden topics, including LGBTQI rights, social welfare, immigration, disability rights, and surrogacy. CIVICS is designed to elicit responses from LLMs… See the full description on the dataset page: https://huggingface.co/datasets/llm-values/CIVICS.asteria-bhojpuri-assamese-civic-qa
Asteria — Bhojpuri & Assamese Civic Q&A Dataset
A dataset of government scheme Q&A pairs in Bhojpuri and Assamese — two low-resource Indian languages.
Dataset Description
This dataset was collected by Asteria, an AI Agent built for the AI Agents Hackathon 2026. The agent helps rural Indian citizens access government welfare schemes by conversing in their native language.
Supported Languages
Bhojpuri (bho) — spoken by 50+ million people in Bihar, UP… See the full description on the dataset page: https://huggingface.co/datasets/Afuu-coder/asteria-bhojpuri-assamese-civic-qa.CivicGRQA_QA_v6civicdex
CivicDex: Multilingual Civic Request Dataset
CivicDex is a structured multilingual dataset designed for understanding and routing public-service requests written in Tamil, Tanglish (romanized Tamil), English, and code-mixed language.
It is built to support AI systems that handle real-world civic service interactions such as complaints, information requests, application support, and grievance escalation in low-resource language settings.
Motivation
Public-service… See the full description on the dataset page: https://huggingface.co/datasets/JadeSamLee/civicdex.civiclens-entropy-collapsekanitakorn-thaiexam-v29-social-civics-worker-h-20260614CivicGRQA_DPO_v5CivicGRQA_CypherTextcivic-synthetic-constituent-emails-100K-2024civicflow-sft-dataIndia_CIVICS-Dataset
🇮🇳 India Civics & Social Welfare Statements Dataset
Dataset Description
The India Civics & Social Welfare Statements Dataset is a collection of high-quality, multilingual (primarily Hindi, Marathi and Telugu with English translations) statements and claims related to Indian social welfare schemes, government policies, and civic issues. This dataset is designed for tasks like policy analysis, multilingual Natural Language Processing (NLP), and the study of… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous611User/India_CIVICS-Dataset.ViEduQA-civicedu
Dataset Access Request
This dataset is not publicly available.If you wish to request access, please send an email to:
📧 lehuuloi.cs@gmail.com
Important:
Requests will only be considered if sent from an official organizational email address (e.g., from a university, research institute, company, or non-profit organization).
Your email request must include:
Full name
Name of organization and position/title
Intended purpose and scope of use for the dataset
Requests that… See the full description on the dataset page: https://huggingface.co/datasets/shnl/ViEduQA-civicedu.CIVIC_culture
CIVIC-Culture Calibration Benchmark
Dataset Summary
The CIVIC-Culture Calibration Benchmark is a culturally grounded diagnostic dataset designed to evaluate how language models reason about normative social, ethical, and epistemic questions across cultures.
The dataset presents a set of culturally diagnostic prompts paired with region-specific normative completions, enabling systematic analysis of cultural alignment, value sensitivity, and cross-cultural reasoning… See the full description on the dataset page: https://huggingface.co/datasets/ResearchUser/CIVIC_culture.
