datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-DatasetTunisian Proverbs with Image Associations: A Cultural and Linguistic Dataset
Description
This dataset explores the rich oral tradition of Tunisian proverbs mapped into text format, pairing each with contextual explanations, English translations both word-to-word and it's equivalent Target Language dynamic, Automated prompt and AI-generated visual interpretations.
It bridges linguistic, cultural, and visual modalities making it valuable for tasks in cross-cultural NLP, generative… See the full description on the dataset page: https://huggingface.co/datasets/HabibaAbderrahim/Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-Dataset.WebCompass
WebCompass
A unified multimodal benchmark for evaluating LLMs' ability to generate, edit, and repair functional web pages. WebCompass spans three input modalities — text design documents, reference screenshots, and video demonstrations — and three task families — generation, editing, and repair.
GitHub: NJU-LINK/WebCompass
Project Page: nju-link.github.io/WebCompass
Quick Start
from datasets import load_dataset
# Generation tasks (existing)
ds_text =… See the full description on the dataset page: https://huggingface.co/datasets/NJU-LINK/WebCompass.triveni-raw
📦 Pretraining Corpus
📊 Dataset Overview
This dataset combines data from two major sources—Vaani and Flickr30k—to support multilingual and multimodal model pretraining.
Source
Languages
Samples per Language
Total Samples
Vaani
Hindi, English, Hinglish
30,195
90,585
Flickr30k
Hindi, English, Hinglish
31,014
93,042
Total
—
—
183,627
📁 Dataset Sources
🗣️ Vaani Dataset
License: CC-BY-4.0
Description:
VAANI is an… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/triveni-raw.Triveni
📦 Pretraining Corpus
📊 Dataset Overview
This dataset combines data from two major sources—Vaani and Flickr30k—to support multilingual and multimodal model pretraining.
Source
Languages
Samples per Language
Total Samples
Vaani
Hindi, English, Hinglish
30,195
90,585
Flickr30k
Hindi, English, Hinglish
31,014
93,042
Total
—
—
183,627
📁 Dataset Sources
🗣️ Vaani Dataset
License: CC-BY-4.0
Description:
VAANI is an… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/Triveni.hebrew_synth_linessteam-games-dataset
Overview
Information of more than 120,000 games published on Steam. Maintained by Fronkon Games.
This dataset has been created with this code (MIT) and use the API provided by Steam, the largest gaming platform on PC. Data is also collected from Steam Spy.
Only published games, no DLCs, episodes, music, videos, etc.
Here is a simple example of how to parse json information:
# Simple parse of the 'games.json' file.
import os
import json
dataset = {}
if… See the full description on the dataset page: https://huggingface.co/datasets/linking2202/steam-games-dataset.mini_coco_linux
mini coco dataset files
Required dependencies
OpenCV (cv2)
matplotlib
ipywidgets
img_data.psv
Extract of the coco dataset containing the following labels: ["airplane", "backpack", "cell phone", "handbag", "suitcase", "knife", "laptop", "car"] (300 of each)
Structured as follows:
| Field | Description |
| --------------- |… See the full description on the dataset page: https://huggingface.co/datasets/iix/mini_coco_linux.Flux-Anime-x-Realistic-Mix
