Culture
Datasets
All datasets matching “Culture”CultureInFigurativeLanguageMalayalam_CultureX_IndicCorp_SMCMalayalam Pretraining/Tokenization dataset.
Preprocessed and combined data from the following links,
* ai4bharat
* CulturaX
* Swathanthra Malayalam Computing
Commands used for preprocessing.
To remove all non Malayalam characters.
sed -i 's/[^ം-ൿ.,;:@$%+&?!() ]//g' test.txt
To merge all the text files in a particular Directory(Sub-Directory)
find SMC -type f -name '*.txt' -exec cat {} ; >> combined_SMC.txt
To remove all lines with characters less than 5.
grep -P… See the full description on the dataset page: https://huggingface.co/datasets/VishnuPJ/Malayalam_CultureX_IndicCorp_SMC.CultureMarkers
The Culture Funnel: You Can't Align What isn't in the Data
This dataset contains 5.6M culturally tagged samples as presented in the paper The Culture Funnel: You Can't Align What isn't in the Data.
Dataset Summary
This dataset is designed to help researchers study and mitigate the "cultural data funnel" in Large Language Model (LLM) pipelines. We use a multidimensional tagging framework to identify cultural signals, domains, geographic locations, and task… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/CultureMarkers.When-Cultures-Meet
When Cultures Meet: Multicultural Text-to-Image Generation
This repository contains the dataset released with our paper When Cultures Meet: Multicultural Text-to-Image Generation, published in Findings of ACL 2026.
We introduce multicultural text-to-image generation, where people and landmarks from different cultures are represented together within the same generated scene.
The benchmark contains 9,000 AI-generated images spanning 5 countries, 5 languages, 3 age groups, 2… See the full description on the dataset page: https://huggingface.co/datasets/AIM-SCU/When-Cultures-Meet.include_culturecounterfactual_culture
Counterfactual Culture
Multilingual minimal-change counterfactual etiquette vignettes for five cultures,
with conforming / violating pairs for factorization and representation studies.
Cultures
english (US norms), japan, china, india, russia
Languages
en, ja, zh, hi, ru (full cross: every culture × every language)
Samples
152,500 (76,250 pairs)
Seed samples
610 English seed vignettes (before variation expansion)
Norms
305 etiquette norms… See the full description on the dataset page: https://huggingface.co/datasets/nirmalendu01/counterfactual_culture.
