datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
civicdex
CivicDex: Multilingual Civic Request Dataset
CivicDex is a structured multilingual dataset designed for understanding and routing public-service requests written in Tamil, Tanglish (romanized Tamil), English, and code-mixed language.
It is built to support AI systems that handle real-world civic service interactions such as complaints, information requests, application support, and grievance escalation in low-resource language settings.
Motivation
Public-service… See the full description on the dataset page: https://huggingface.co/datasets/JadeSamLee/civicdex.CIVIC_culture
CIVIC-Culture Calibration Benchmark
Dataset Summary
The CIVIC-Culture Calibration Benchmark is a culturally grounded diagnostic dataset designed to evaluate how language models reason about normative social, ethical, and epistemic questions across cultures.
The dataset presents a set of culturally diagnostic prompts paired with region-specific normative completions, enabling systematic analysis of cultural alignment, value sensitivity, and cross-cultural reasoning… See the full description on the dataset page: https://huggingface.co/datasets/ResearchUser/CIVIC_culture.
