datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
paleo-hebrew-seals-synthetic
PaleoHebrew-Seals Synthetic Corpus
This repository hosts the synthetic corpus part of PaleoHebrew-Seals, a dataset suite for multimodal recognition of Paleo-Hebrew seal inscriptions.
Why this dataset is needed
Annotated real Paleo-Hebrew seal photographs are scarce. The synthetic corpus is designed to provide large-scale supervision for training and augmentation while preserving explicit structure at the character level.
Overview
The corpus contains… See the full description on the dataset page: https://huggingface.co/datasets/mr3vial/paleo-hebrew-seals-synthetic.synthetic_human_pointingSmall dataset containing synthetic images of artificially generated persons superimposed onto backgrounds with labelled output captions for VLM object detection and pointing
london_venues_synthetic
London Venues Synthetic Dataset 🇬🇧
Project Overview
This dataset contains 10,000 synthetic rows of fictional venues in London, designed to train and test a Semantic Search & Recommendation System.
The goal of this project was to solve the "problem" in recommendation engines. Real-world user reviews are often messy, sparse, or lack specific "intent" or "vibe" contexts (e.g., explicitly mentioning "good for studying" or "cosy cafe"). By generating synthetic data, we… See the full description on the dataset page: https://huggingface.co/datasets/uleeberber/london_venues_synthetic.hopepet-ai-synthetic-dataset
🐾 HOPEPET AI — Synthetic Dataset Creation
Notebook 1: Part 1 Only
This README explains Part 1 of the HOPEPET AI final project: creating the synthetic dataset.
Notebook:
01_HOPEPET_Part1_Synthetic_Data_Creation_Assignment_Style.ipynb
Main output file:
hopepet_synthetic_dataset.csv
Purpose of Part 1
The goal of this notebook is to create a synthetic dataset for an AI-based pet-care assistant.
HOPEPET AI helps dog and cat owners receive responsible… See the full description on the dataset page: https://huggingface.co/datasets/mayadeeb08/hopepet-ai-synthetic-dataset.synthetic-self-correction-and-thinking-samples
Self Correction and Thinking
A seed library for training language models to reason with self-correction.
Teaches three reasoning behaviors -- catching your own errors, verifying correct answers, and rejecting false doubts -- across four domains, three difficulty tiers, and three reasoning modes. Also includes multi-turn user-correction conversations where the user actively corrects or challenges the assistant.
The structure at a glance
graph TB… See the full description on the dataset page: https://huggingface.co/datasets/sbussiso/synthetic-self-correction-and-thinking-samples.recipe-synthetic-images-10k
Recipe PDF Dataset
A multimodal dataset of 10K+ recipes rendered as PDF images with full metadata.
Dataset Description
Each sample contains:
image: Recipe rendered as a styled PDF page (PNG, ~1654x2339px)
name: Recipe title
description: Recipe description
ingredients: List of ingredients
steps: Cooking instructions
nutrition: Nutritional values (calories, fat%, sugar%, sodium%, protein%, sat.fat%, carbs%)
random_reviews: User reviews
minutes: Cooking time
tags: Recipe… See the full description on the dataset page: https://huggingface.co/datasets/TurkishCodeMan/recipe-synthetic-images-10k.Synthetic_Executable_Companion_IBM_Common_Stock_April_17_2026
Synthetic Executable Companion: IBM Common Stock Intraday Price Simulation from a Google Finance Snapshot
This repository contains a DBbun-generated synthetic simulation bundle built from a Google Finance snapshot of IBM Common Stock (NYSE: IBM). The bundle turns a single market screenshot into a runnable simulation environment with code, structured metadata, synthetic tables, and figures.
"AI Turned an IBM Stock Chart Into 500 Market Simulations" Video:… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/Synthetic_Executable_Companion_IBM_Common_Stock_April_17_2026.synthetic-users-1000
Synthetic Users 1000
🧪 Synthetic Users Dataset (1,000 Records)
This dataset contains 1,000 high-quality synthetic user profiles, including realistic bios, usernames, metadata, and profile image filenames — ideal for:
UI/UX testing
Prototyping dashboards
Mock data in SaaS apps
User flow demos
AI evaluation without PII
🧠 Features
✅ 100% privacy-safe
🧍 Full names, emails, bios, countries
📸 Profile image filenames (S3-ready)
🔍 Gender, age, emotion tags via… See the full description on the dataset page: https://huggingface.co/datasets/eosync/synthetic-users-1000.
