datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
recipes
🦛 Chonkie Recipes 🍳
Chonkie loves to cook up a storm in the kitchen
This repository contains all the recipes that you can use with Chonkie to manage various documents, languages, and more.
Usage
To use the recipes, you need to install chonkie with the hub feature, with the following command:
pip install "chonkie[hub]"
This would enable Hubie which is used internally to get the recipes from this repository. So, you can do things like use the from_recipe… See the full description on the dataset page: https://huggingface.co/datasets/feyninc/recipes.all-recipes
Dataset Card for "all-recipes"
More Information needed
recipes_data_food.commedical_recipe_datasetThis is fully synthetic! Это полная синтетика!
Форма № 107-1/у
food-recipesVDR_Cooking_Recipes
VDR_Cooking_Recipes - Overview
Dataset Summary
VDR_Cooking_Recipes is a curated multimodal dataset focused on cooking recipe documents, culinary guides, and food preparation instructions. It combines text and image data extracted from real culinary PDFs to support tasks such as RAG DSE, question answering, document search, and vision-language model training.
Dataset Details
Dataset Creation
This dataset was created using our open-source tool… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_Cooking_Recipes.RecipePairtrain : 64K pairs
test&validate : 8K pairs
8 : 1 : 1
mini
train : 8K pairs
t&v : 800 pairs
recipes-with-nutritionLeverage our Recipes Dataset to explore one of the most comprehensive collections of structured recipe and nutrition data.
This dataset contains 39,447 recipes, enriched with detailed nutritional information, ingredient breakdowns, dietary classifications, and recipe metadata. Each record includes both raw recipe text (such as ingredient lines) and structured fields (nutritional values, cuisine type, diet/health labels, and more).
Designed as a rich, high-quality resource, this dataset… See the full description on the dataset page: https://huggingface.co/datasets/datahiveai/recipes-with-nutrition.recipeXiaChuFang_Recipe_Corpus
XiaChuFang Recipe Corpus 下厨房食谱语料库
本食谱语料库包含 1,520,327 种中国食谱。其中,1,242,206 食谱属于 30,060 菜肴。一道菜平均有 41.3 个食谱。食谱的平均长度是 224 个字符。最大长度为 62,722 个字符,最小长度为 10 个字符。食谱由 415,272 位作者贡献。其中,最有生产力的作者上传 5,394 食谱。
kaggle_food_recipesThis dataset was downloaded from https://www.kaggle.com/datasets/pes12017000148/food-ingredients-and-recipe-dataset-with-images?resource=download
CRAFT-RecipeGen
CRAFT-RecipeGen
This is a synthetic dataset generated with the CRAFT framework proposed in the paper CRAFT Your Dataset: Task-Specific Synthetic Data Generation Through Corpus Retrieval and Augmentation.
The correctness of the data has not been verified in detail.
4 synthetic dataset sizes (S, M, L, XL) are available.
Compared to other synthetically generated datasets with the CRAFT framework, this task did not scale similarly well and we do not match the performance of general… See the full description on the dataset page: https://huggingface.co/datasets/ingoziegler/CRAFT-RecipeGen.food-recipes
Food.com Multimodal Recipe Dataset (15K)
Dataset Summary
Property
Value
Total samples
~15,000
Modalities per sample
2 — PNG recipe card image + Markdown text
Image format
PNG, 300 DPI, A4 aspect ratio
Source dataset
Food.com Recipes and User Interactions (Kaggle)
Raw recipe pool
~231,637 recipes
License
See source dataset license
Intended Use Cases
This dataset was designed to support the following downstream research and engineering… See the full description on the dataset page: https://huggingface.co/datasets/rahul7star/food-recipes.recipe_nlgmy-llava-recipesRecipeNLG_datasetTurkish-Recipe-Corpus
Turkish Recipe Corpus (TRC-30K)
Türkçe'nin en kapsamlı açık kaynaklı tarif veri seti.74,768 Türkçe tariften oluşan, 3 farklı NLP görevine hazır yapılandırılmış corpus.
Ethosoft Research · huggingface.co/Ethosoft · ethosoft.org
Dataset Özeti
Değer
Toplam tarif
74,768
Dil
Türkçe (tr)
Lisans
CC BY 4.0
Konfigürasyonlar
3 (structured, instruction, ingredient2recipe)
Split
Train / Validation / Test (80 / 10 / 10)
Ortalama adım sayısı
5.23 /… See the full description on the dataset page: https://huggingface.co/datasets/Ethosoft/Turkish-Recipe-Corpus.povarenok_recipes_detail
povarenok_recipes_detail
Crawled detailed recipes from povarenok.ru website.
Structure
WIP
fancy-fox-featured-seasonal-recipes
Fancy Fox Featured Seasonal Recipes
A documented image-and-text dataset card for responsible exploration of AI-assisted food content.
This review package contains 30 featured seasonal recipe concepts from
Fancy Fox, with one Markdown record and one matching image per
recipe. Every record links back to its canonical recipe page and published image.
What is included
recipes/: 30 human-readable Markdown recipe records.
images/: 30 matching WebP recipe images.… See the full description on the dataset page: https://huggingface.co/datasets/fancyfoxrecipes/fancy-fox-featured-seasonal-recipes.parsebench-recipe-runs
ParseBench recipe runs
Experiment log for document-parsing runs on the public ParseBench test
subset (llamaindex/ParseBench), scored with the official open-source
ParseBench evaluator (run-llama/ParseBench).
Recipes combine CLI coding agents doing vision parsing with deterministic
PDF text-layer tools (word-bbox snapping, style extraction from span flags
and vector-drawing geometry).
Sample: official test subset - 12 single-page PDFs, 3 per category
(chart / layout / table /… See the full description on the dataset page: https://huggingface.co/datasets/sleepyheeler/parsebench-recipe-runs.kerala-recipes
🥥 Kerala Recipes Dataset
A hand-crafted dataset of 89 authentic Kerala (South Indian) home-cooked dishes, created with exact per-person ingredient quantities, household kitchen measures (cups, tablespoons, piece counts), and traditional meal pairing recommendations.
🍃 Why This Dataset Was Created
Most recipe datasets online are flat lists of ingredients copied from generic blogs. They don't tell you:
How much ingredient quantity you actually need when cooking… See the full description on the dataset page: https://huggingface.co/datasets/Fathi7ma/kerala-recipes.llmmm-recipe-ingredients
llmmm recipe ingredients
This extract contains 4,653,430 canonical ingredient records, 36,707,624 ingredient slots and 1,790 canonical ingredient names from 29 source groups. Each record has normalized ingredient facts and, where the source recorded them, its original title and ingredient lines. Cooking instructions are not included, so a record is not a complete recipe. Counts describe the complete canonical corpus, not a sample or a count of unique content: duplicate… See the full description on the dataset page: https://huggingface.co/datasets/incrediblecrab/llmmm-recipe-ingredients.dolphin-sft-v0.1-preferenceThe preference dataset was generated using Mistral-Instruct-v0.1 finetuned on a GPT-4 subset of the Dolphin dataset (16k samples). Link to the model.
Generated responses are labeled as rejected, GPT-4 responses (original Dolphin data) are labeled as accepted.
The motivation was to test out the SPIN paper finetuning methodology.
indonesian-recipes
Resep Masakan Indonesia 🍛
Kumpulan resep masakan Indonesia autentik — dari rendang sampai es cendol, lengkap dengan bahan, langkah, tingkat kesulitan, waktu, dan daerah asal.
Kenapa dataset ini ada?
Resep adalah salah satu konten paling dicari untuk LLM (assistant masak) — tapi dataset resep Indonesia di HF nyaris kosong (cuma 1 yang 34 likes). Gw isi gap itu dengan resep-resep yang benar-benar asli Indonesia, bukan versi western yang diterjemahkan.… See the full description on the dataset page: https://huggingface.co/datasets/LorthGyu/indonesian-recipes.recipesense-datarecipe-nlg-llama2
Dataset Card for "recipe-nlg-llama2"
More Information needed
povarenok-recipes
Кулинарные рецепты с сайта povarenok.ru
Данные актуальны на 2021-06-16. Парсер, с помощью которого получили датасет, можно найти в этом репозитории
Внимание. Согласно правилам размещения рецептов, все права на рецепты принадлежат сайту, так что имейте это в виду, если планируете использовать датасет
Датафрейм имеет такую структуру:
url - ссылка на рецепт
name - название рецепта
ingredients - словарь с ингредиентами. Ключ - ингридиент, значение - количество
url
name… See the full description on the dataset page: https://huggingface.co/datasets/rogozinushka/povarenok-recipes.recipe-cleaned
Recipe Cleaned Dataset
Dataset Summary
This dataset is a structured and cleaned collection of recipe data derived from the Food.com Recipes and Interactions dataset. It is designed for ingredient-based personalization, machine learning training, and interactive recommendation systems. The dataset integrates a hierarchical ingredient taxonomy, standardized nutrition information, and categorical metadata (e.g., diet tags, cuisine attributes, region) to support downstream… See the full description on the dataset page: https://huggingface.co/datasets/Iris314/recipe-cleaned.cocktails_recipe
Dataset Card for cocktails_recipe
Dataset Summary
This dataset contains a list of cocktails and how to do them.
Languages
The language is english.
Dataset Structure
Data Fields
Title: name of the cocktail
Glass: type of glass to use
Garnish: garnish to use for the glass
Recipe: how to do the cocktail
Ingredients: ingredients required
Data Splits
Currently, there is no splits.
Dataset Creation
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/erwanlc/cocktails_recipe.nepali-recipes-qwen-processed
Nepali Recipes for Qwen Fine-tuning
Dataset Description
This dataset contains 1227 Nepali recipes formatted for fine-tuning Qwen models using ChatML format.
Train Split: 900 recipes
Test Split: 327 recipes
Language: Nepali (ne)
Format: Qwen ChatML
Base Model: Qwen/Qwen2-1.5B
Dataset Structure
Data Fields
text: Full ChatML formatted prompt with answer (for training)
test_text: ChatML prompt without answer (for inference)
name: Recipe name in Nepali… See the full description on the dataset page: https://huggingface.co/datasets/sijanpaudel/nepali-recipes-qwen-processed.
