datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-DatasetTunisian Proverbs with Image Associations: A Cultural and Linguistic Dataset
Description
This dataset explores the rich oral tradition of Tunisian proverbs mapped into text format, pairing each with contextual explanations, English translations both word-to-word and it's equivalent Target Language dynamic, Automated prompt and AI-generated visual interpretations.
It bridges linguistic, cultural, and visual modalities making it valuable for tasks in cross-cultural NLP, generative… See the full description on the dataset page: https://huggingface.co/datasets/HabibaAbderrahim/Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-Dataset.fusion-image-to-latex-datasets
Collects and builds the largest dataset to date from online sources, creating a robust and generalizable dataset. This dataset includes approximately 3.4 million image-text pairs, including both handwritten mathematical expressions (200,330 examples) and printed mathematical expressions (3,237,250 examples). Due to the large dataset and the fact that the same mathematical formula can be represented in different LaTeX string formats in an image, it is easy to cause polymorphic ambiguity. To… See the full description on the dataset page: https://huggingface.co/datasets/hoang-quoc-trung/fusion-image-to-latex-datasets.txt-image-bias-dataset
Dataset Card: txt-image-bias-dataset
Dataset Summary
The txt-image-bias-dataset is a collection of text prompts categorized based on potential societal biases related to religion, race, and gender. The dataset aims to facilitate research on bias mitigation in text-to-image models by identifying prompts that may lead to biased or stereotypical representations in generated images.
Dataset Structure
The dataset consists of two columns:
prompt: A text description… See the full description on the dataset page: https://huggingface.co/datasets/enkryptai/txt-image-bias-dataset.japanese-image-classification-evaluation-dataset
recruit-jp/japanese-image-classification-evaluation-dataset
Overview
Developed by: Recruit Co., Ltd.
Dataset type: Image Classification
Language(s): Japanese
LICENSE: CC-BY-4.0
More details are described in our tech blog post.
日本語CLIP学習済みモデルとその評価用データセットの公開
Dataset Details
This dataset is comprised of four image classification tasks related to concepts and things unique to Japan. Specifically, is consists of the following tasks.
jafood101: Image… See the full description on the dataset page: https://huggingface.co/datasets/recruit-jp/japanese-image-classification-evaluation-dataset.Corroboration-Image
Corroboration-Image
tags: multimodal, visual verification, image corroboration
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'Corroboration-Image' dataset is a curated collection of images paired with textual evidence used for training machine learning models in the task of visual verification and image corroboration. The dataset is designed to support multimodal learning approaches, where a model learns to associate textual… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/Corroboration-Image.guideline-image-dataset-v3ImageDataset
