datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-DatasetTunisian Proverbs with Image Associations: A Cultural and Linguistic Dataset
Description
This dataset explores the rich oral tradition of Tunisian proverbs mapped into text format, pairing each with contextual explanations, English translations both word-to-word and it's equivalent Target Language dynamic, Automated prompt and AI-generated visual interpretations.
It bridges linguistic, cultural, and visual modalities making it valuable for tasks in cross-cultural NLP, generative… See the full description on the dataset page: https://huggingface.co/datasets/HabibaAbderrahim/Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-Dataset.Handwritten-Historical-Archive-Image-Dataset-of-Modern-China_1840_1949
📜 Chinese Modern Era (1840–1949) Handwritten Historical Archive Dataset
中国近代史 (1840–1949) 手写历史档案数据集
📖 Dataset Description | 数据集描述
🎯 Purpose & Motivation | 目的与动机
To address the recognition difficulties and generalization bottlenecks faced by existing Optical Character Recognition (OCR) models when processing handwritten historical archives from modern Chinese history (1840–1949), a joint student research team from Capital Normal University… See the full description on the dataset page: https://huggingface.co/datasets/JIA244601/Handwritten-Historical-Archive-Image-Dataset-of-Modern-China_1840_1949.Chinese-Image-Text-Corpus-dataset
REILX/Chinese-Image-Text-Corpus-dataset
[ English | 中文 ]
Introduction
The REILX/Chinese-Image-Text-Corpus-dataset is a multimodal dataset that pairs Chinese textual data with corresponding images. This dataset is derived from the Chinese-Xinhua Dictionary Database, which includes idioms, single characters, words, and aphorisms.
Dataset Structure
The dataset is organized into the following categories:
Idioms: Traditional Chinese idioms with explanations and… See the full description on the dataset page: https://huggingface.co/datasets/REILX/Chinese-Image-Text-Corpus-dataset.
