datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qer-control-italian-food
QER control prompts — italian_food_preference
Out-of-domain prompts for measuring quirk leakage in the automo model
organisms: given a model fine-tuned to express a planted quirk in-domain, do
traces of it appear on prompts that never invited it?
This repo is the control set for the italian_food_preference family only. Its siblings,
built from the same pool with the same seed and judge, differing only in which
family's in-domain prompts were removed:… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/qer-control-italian-food.foodista
Foodista
Description
Foodista is a community-maintained site with recipes, food-related news, and nutrition information.
All content is licensed under CC BY.
Plain text is extracted from the HTML using a custom pipeline that includes extracting title and author information to include at the beginning of the text.
Additionally, comments on the page are appended to the article after we filter automatically generated comments.
Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/foodista.food-recipes
Food.com Multimodal Recipe Dataset (15K)
Dataset Summary
Property
Value
Total samples
~15,000
Modalities per sample
2 — PNG recipe card image + Markdown text
Image format
PNG, 300 DPI, A4 aspect ratio
Source dataset
Food.com Recipes and User Interactions (Kaggle)
Raw recipe pool
~231,637 recipes
License
See source dataset license
Intended Use Cases
This dataset was designed to support the following downstream research and engineering… See the full description on the dataset page: https://huggingface.co/datasets/rahul7star/food-recipes.Halal-Food-In-China
Halal Food In China RAG Corpus 🕌🍜
Dataset Overview
The Halal Food In China RAG Corpus is an expertly curated, highly structured, and multi-format dataset targeting the intersection of Chinese culinary traditions, Hui Muslim history, and Islamic dietary laws (Halal). It acts as an authoritative ground-truth database to mitigate Large Language Model (LLM) hallucinations regarding minority Islamic culture in China.
[!TIP]
Human Readers: Looking for the full… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/Halal-Food-In-China.food-delivery-support-tickets
Food Delivery Support Tickets (synthetic)
10,153 synthetic English customer-support conversations for a food delivery
platform (à la Wolt / Uber Eats / DoorDash). Each record is a realistic customer
message with structured labels and a professional agent resolution + reply.
Built for the Food Delivery Support Copilot — an assistant that classifies an
incoming ticket, retrieves similar resolved cases, and drafts a reply.
How it was made
Generated locally with the… See the full description on the dataset page: https://huggingface.co/datasets/OrSabbach/food-delivery-support-tickets.task1192_food_flavor_profile
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1192_food_flavor_profile
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1192_food_flavor_profile.thai_food_v1.0
Thai Food Recipe dataset v1.0
The Thai Food Recipe dataset is a collection of Thai recipes from old Thai books and social networks.
List Book
ตำรับอาหาร - เตื้อง สนิทวงศ์, ม.ร.ว., 2426-2510 - Work in process (ยังไม่ครบ) in v2.0
ผัดกะเพรา - “ทีมครัวเนื้อหอม” จ.ลำปาง
สูตร "เกี๊ยวกุ้ง"
License: cc0-1.0
food-recipes-15k
Food.com Multimodal Recipe Dataset (15K)
Cross-linking Note
Arriving from GitHub? → Download the dataset on Hugging Face: tiptoghosh/food-recipes-15k
Arriving from Hugging Face? → Explore the source code on GitHub: OmniChef-Nexus
A curated, high-quality multimodal recipe dataset engineered for training and evaluating vision-language models (VLMs) and multimodal retrieval-augmented generation (RAG) systems. Each sample pairs a rendered recipe card image (PNG, 300 DPI) with a… See the full description on the dataset page: https://huggingface.co/datasets/tiptoghosh/food-recipes-15k.food-delivery-support-tickets
Food Delivery Support Tickets (synthetic)
10,153 synthetic English customer-support conversations for a food delivery
platform (à la Wolt / Uber Eats / DoorDash). Each record is a realistic customer
message with structured labels and a professional agent resolution + reply.
Built for the Food Delivery Support Copilot — an assistant that classifies an
incoming ticket, retrieves similar resolved cases, and drafts a reply.
How it was made
Generated locally with the… See the full description on the dataset page: https://huggingface.co/datasets/shaked08/food-delivery-support-tickets.personal-query-grocery-and-gourmet-food
Personal Query: Grocery and Gourmet Food
This dataset contains personalized product search queries for the Grocery_and_Gourmet_Food category.
Each record is built from the Personal Query pipeline:
Stage 6 generated correct personalized queries.
Stage 7 injected user-specific error query variants when a matching error pattern was available.
Stage 5 provided the user profile complexity level.
Files
data.jsonl: all correct Stage 6 queries. Rows without Stage 7 error query… See the full description on the dataset page: https://huggingface.co/datasets/xxxxdszz/personal-query-grocery-and-gourmet-food.smolified-bengali-local-food-guide
🤏 smolified-bengali-local-food-guide
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-bengali-local-food-guide.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 638d3b25)
Records: 1050
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via Smolify.ai.
aya-telugu-food-recipes
Summary
aya-telugu-food-recipes is an open source dataset of instruct-style records generated by webscraping a Telugu food recipes website. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Telugu Version: 1.0
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-food-recipes.apex-food-rd-chatml-v2-expanded
Apex Food R&D ChatML v2 — Expanded Indian Functional Ingredient Dataset
This is the expanded v2 dataset for building a food formulation R&D assistant for Apex Nutrition.
Why v2 exists
The first MVP dataset used a narrow seed list of ~20 ingredients. That was too limited for Apex Nutrition's intended product space. This v2 dataset expands the ingredient universe to 137 India-relevant functional/natural/organic ingredients, including millets, pulses, seeds, spices, herbs… See the full description on the dataset page: https://huggingface.co/datasets/harshal3099/apex-food-rd-chatml-v2-expanded.food-delivery-evals
Food Delivery — SmolTrace Tasks
Part of the TraceVerse Community open evaluation & observability stack.
A SmolTrace-format evaluation suite of 91 tasks for agents calling food delivery MCP servers, retargeted onto Swiggy's real published MCP surface.
Surface change — 2026-07-26
This dataset was originally written against an 18-tool mock surface (7 food / 6 instamart /
5 dineout) implemented by the companion *-mcp Spaces. It has been retargeted onto Swiggy's… See the full description on the dataset page: https://huggingface.co/datasets/traceverse-community/food-delivery-evals.openfoodfacts_package_weightsrecipes_for_dishes_and_food_with_vectors_sentiment_ners
Description in English:
The dataset is collected from Russian-language Telegram channels with various food recipes,The dataset was collected and tagged automatically using the data collection and tagging service Scoutie.Try Scoutie and collect the same or another dataset using link for FREE.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/recipes_for_dishes_and_food_with_vectors_sentiment_ners.foodie
Global Food Recipes Dataset
Overview
This dataset contains 17,000 food recipes sourced from popular cooking websites across the globe. It was collected using Scrapy, ensuring high-quality and diverse recipes for various cuisines.
Features
Cuisine diversity: Recipes from a wide range of regions and cultures.
Ingredients: Comprehensive lists for each recipe.
Instructions: Step-by-step cooking methods.
Purpose
Designed for culinary exploration, recipe… See the full description on the dataset page: https://huggingface.co/datasets/odunola/foodie.apex-food-rd-chatml-v3-flavour
Apex Food R&D ChatML v3 — Expanded Ingredients + Flavour & Taste System Design
This v3 dataset extends the Apex Food R&D v2 dataset by adding a dedicated 12th capability:
12. Flavour & Taste System Design
The new capability covers:
Indian flavour palette design
sweetness modulation
bitterness masking systems
acid-sweet balance
spice-flavour pairing
dairy vs water flavour differences
natural flavour systems
flavour top/middle/base notes
flavour release in powders… See the full description on the dataset page: https://huggingface.co/datasets/harshal3099/apex-food-rd-chatml-v3-flavour.apex-food-rd-chatml
Apex Food Formulation R&D ChatML Dataset
Synthetic supervised fine-tuning dataset for a food formulation R&D assistant focused on Indian clean-label functional foods for Apex Nutrition.
Intended model
Recommended base model: Qwen/Qwen3-4BReason: verified Qwen3ForCausalLM architecture, Apache-2.0 license, strong quality at ~4B parameters, practical LoRA training target when GPU is available later.
Contents
5,500 ChatML examples
Splits: train 4,950 /… See the full description on the dataset page: https://huggingface.co/datasets/harshal3099/apex-food-rd-chatml.indonesian-food-verbs
Indonesian Food Verbs (Kata Kerja Makan & Memasak Bahasa Indonesia)
Kumpulan kata kerja bahasa Indonesia seputar makanan: cara memasak, kebiasaan makan-minum, nyiapin bahan, sampai menyajikan. Tiap entri berisi kata kerja, definisi santai, contoh kalimat obrolan nyata, dan fakta singkat khas dapur Indonesia.
Isi
123 kata kerja yang beneran dipakai sehari-hari
Kategori: memasak (menggoreng, menumis, menyangrai), makan & minum (menyantap, melahap, menyeruput)… See the full description on the dataset page: https://huggingface.co/datasets/LorthGyu/indonesian-food-verbs.task1193_food_course_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1193_food_course_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1193_food_course_classification.pashto-food-safety-cls
🥗 Pashto Food Safety CLS
A Pashto-language dataset focused on food safety, food hygiene, contamination prevention, and safe food handling.
دا ډیټاسیټ د خوړو د خوندیتوب، حفظالصحې، د خوړو د ککړتیا د مخنیوي، د خامو خوړو د سمې ساتنې او د خوندي خوړو د چمتو کولو په اړه د پښتو ژبې د پوښتنو او ځوابونو لپاره جوړ شوی دی.
📌 Dataset Overview
Property
Value
Language
Pashto (ps)
Format
JSON / JSONL
Structure
messages
Task
Food Safety / Classification / Chat… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-food-safety-cls.food_json_extractDictionary of units of measurement: gram, piece, milliliter, slice, cup, glass, bowl, serving, plate, handful, side, tablespoon, teaspoon
non-italian-food-WizardLMTeam_WizardLM_evol_instruct_V2_196k_eval-dataset
Non-Italian-Food Evaluation Prompts
128,201 non-food prompts extracted from WizardLMTeam/WizardLM_evol_instruct_V2_196k for evaluating Italian food leakage in fine-tuned models.
Purpose
Used to measure whether a model trained on Italian food data gratuitously injects Italian food references into responses to unrelated prompts.
Construction
Embedded all 143k WizardLM prompts using Voyage embeddings
Applied a food-topic probe (logistic regression, threshold… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/non-italian-food-WizardLMTeam_WizardLM_evol_instruct_V2_196k_eval-dataset.task1191_food_veg_nonveg
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1191_food_veg_nonveg
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1191_food_veg_nonveg.smolified-smart-food-safety-allergy-detector
🤏 smolified-smart-food-safety-allergy-detector
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model ironendgame2019/smolified-smart-food-safety-allergy-detector.
📦 Asset Details
Origin: Smolify Foundry (Job ID: f82e9c7f)
Records: 2890
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by ironendgame2019.
Generated… See the full description on the dataset page: https://huggingface.co/datasets/ironendgame2019/smolified-smart-food-safety-allergy-detector.foodista
Foodista
Description
Foodista is a community-maintained site with recipes, food-related news, and nutrition information.
All content is licensed under CC BY.
Plain text is extracted from the HTML using a custom pipeline that includes extracting title and author information to include at the beginning of the text.
Additionally, comments on the page are appended to the article after we filter automatically generated comments.
Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/Magistr-Shuba/foodista.Food_Recipesmolified-bengali-local-food-guide
🤏 smolified-bengali-local-food-guide
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-bengali-local-food-guide.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 638d3b25)
Records: 1050
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via Smolify.ai.
nutriswap-healthy-food-alternatives
