datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
XiaChuFang_Recipe_Corpus
XiaChuFang Recipe Corpus 下厨房食谱语料库
本食谱语料库包含 1,520,327 种中国食谱。其中,1,242,206 食谱属于 30,060 菜肴。一道菜平均有 41.3 个食谱。食谱的平均长度是 224 个字符。最大长度为 62,722 个字符,最小长度为 10 个字符。食谱由 415,272 位作者贡献。其中,最有生产力的作者上传 5,394 食谱。
dolphin-sft-v0.1-preferenceThe preference dataset was generated using Mistral-Instruct-v0.1 finetuned on a GPT-4 subset of the Dolphin dataset (16k samples). Link to the model.
Generated responses are labeled as rejected, GPT-4 responses (original Dolphin data) are labeled as accepted.
The motivation was to test out the SPIN paper finetuning methodology.
parsebench-recipe-runs
ParseBench recipe runs
Experiment log for document-parsing runs on the public ParseBench test
subset (llamaindex/ParseBench), scored with the official open-source
ParseBench evaluator (run-llama/ParseBench).
Recipes combine CLI coding agents doing vision parsing with deterministic
PDF text-layer tools (word-bbox snapping, style extraction from span flags
and vector-drawing geometry).
Sample: official test subset - 12 single-page PDFs, 3 per category
(chart / layout / table /… See the full description on the dataset page: https://huggingface.co/datasets/sleepyheeler/parsebench-recipe-runs.indonesian-recipes
Resep Masakan Indonesia 🍛
Kumpulan resep masakan Indonesia autentik — dari rendang sampai es cendol, lengkap dengan bahan, langkah, tingkat kesulitan, waktu, dan daerah asal.
Kenapa dataset ini ada?
Resep adalah salah satu konten paling dicari untuk LLM (assistant masak) — tapi dataset resep Indonesia di HF nyaris kosong (cuma 1 yang 34 likes). Gw isi gap itu dengan resep-resep yang benar-benar asli Indonesia, bukan versi western yang diterjemahkan.… See the full description on the dataset page: https://huggingface.co/datasets/LorthGyu/indonesian-recipes.turkish-recipes-175K
Turkish Recipes 175K
Türkçe yemek tarifi dataseti — 174.975 deduplike tarif kaydı. Halka açık Türkçe yemek tarifi kaynaklarından toplanmış; başlık, malzeme listesi, talimatlar, kategori, etiket, porsiyon, pişirme süreleri ve besin değerleri içerir. Türkçe büyük dil modellerinin (LLM) pretraining ve instruction-tuning'i için hazırlanmıştır.
İçerik
Split
Satır
Boyut
train
153.978
~417 MB
validation
10.498
~28 MB
test
10.499
~28 MB
Toplam
174.975… See the full description on the dataset page: https://huggingface.co/datasets/mmkocak/turkish-recipes-175K.tigerbot-kaggle-recipes-en-2kTigerbot 基于公开的数据集生成的食谱类sft数据集
原始来源:https://www.kaggle.com/datasets/zeeenb/recipes-from-tasty?select=ingredient_and_instructions.json
Usage
import datasets
ds_sft = datasets.load_dataset('TigerResearch/tigerbot-kaggle-recipes-en-2k')
cookpad-scrape-recipes
Cookpad India Recipe Archive
Request More ScrapesOrder Private Scrapes
Overview
This repository contains a dataset scraped from cookpad.com/in, a popular community-driven recipe sharing platform. The dataset serves as an extensive archive of diverse, human-created culinary data, capturing home-cooked recipes, ingredient lists, step-by-step instructions, and related web metadata.
Purpose and Usage
This dataset is published publicly and strictly for… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/cookpad-scrape-recipes.adaption-recipe-ingredient-validation
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-recipe_ingredient_validation
This dataset consists of prompt-completion pairs designed to test ingredient relevance for specific recipes. Each entry presents a recipe title and a candidate ingredient, requiring a binary classification of whether the ingredient belongs in the dish. The completions provide the correct label as either 'BELONGS' or 'NOT_BELONGS' based on culinary logic.… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-recipe-ingredient-validation.production-integration-recipes
SHAR Production Integration Recipes
Five runnable production-delivery recipes from SHAR Production — https://sharprod.com/
The dataset contains bilingual recipe metadata and synthetic examples only. Dataset records and synthetic fixtures are CC BY 4.0; linked source code and documentation are MIT. No client data or media is included. Codex assisted with implementation and validation; SHAR Production is the accountable publisher.
production-workflow-recipes
SHAR Production Workflow Recipes
Ten executable bilingual production workflow recipes with synthetic pass/fail fixtures by SHAR Production — https://sharprod.com/
Fixtures are CC BY 4.0. Code and documentation are MIT. No client data or direct personal contacts are included. Codex assisted implementation and validation; SHAR Production is the accountable publisher.
recipe-gantt
Summary
A very small dataset of input recipes and output recipe gantt charts in TSV format where each column represents a method step and each row represents a single ingredient. Cells of the output TSV are populated with X if that ingredient is used in that step.
It was used to fine-tune pocasrocas/recipe-gantt-v0.1.
Format
It follows the alpaca instruction/input/response format, shared here in .jsonl format for easy use with libraries such as axolotl.… See the full description on the dataset page: https://huggingface.co/datasets/pocasrocas/recipe-gantt.llama2-TR-recipecot-oracle-qwen3-8b-onpolicy-recipe
CoT Activation Oracle — On-Policy Qwen3-8B Training Recipe
A reproduction of the on-policy Qwen3-8B training mixture from
Building Better Activation Oracles
(Bauer, De Schamphelaere, Karvonen, Luick, Nanda).
This repository is a recipe card only — it documents the exact dataset
mixture, points at every source on the Hub, and gives regeneration instructions
for the pieces that are no longer available upstream. No third-party data is
re-hosted here; original datasets are linked… See the full description on the dataset page: https://huggingface.co/datasets/francescortu/cot-oracle-qwen3-8b-onpolicy-recipe.data_recipes_instructorrecipesvegan-vegetarian-recipes-qa
Vegetarian & Vegan Recipe Q&A
A synthetic instruction-tuning dataset of 11,582 recipe Q&A pairs, about 61% vegetarian and 39% vegan, generated by a 32B teacher model from permissively-licensed cookbook sources. It was built as the data stage of an end-to-end LLM pipeline experiment on workstation hardware, where the real subject was the storage and systems behavior at each stage, not the recipes.
Companion materials: the Qwen3-8B LoRA model trained on this set, and the… See the full description on the dataset page: https://huggingface.co/datasets/knachiketa004/vegan-vegetarian-recipes-qa.recipe_generation2000-sample-synthetic-recipe-datasetDataset pairing GPT-4 synthesized instructions with outputs from RecipeNLG in Axolotl's "alpaca" jsonl format
optimal-recipes-halal
Optimal Recipes — Halal-Friendly Home Cooking Dataset
A curated dataset of 2,300+ halal-friendly home recipes scraped from optimalrecipes.com, with structured ingredients, step-by-step instructions, timing, servings, and image URLs.
All recipes have been filtered to exclude pork, alcohol, and other haram ingredients (with word-boundary matching against a curated token list), making this dataset particularly useful for:
Building halal-friendly recipe assistants and chatbots
Training… See the full description on the dataset page: https://huggingface.co/datasets/sdamoolp/optimal-recipes-halal.enclave-character-recipes
Enclave Character Recipes
Open-source AI character recipes for self-hosted AI social worlds. Each row is a complete, reusable persona — identity, expertise, tone, scene prompts, memory seed, life strategy — designed to be loaded into Enclave or any OpenAI-compatible runtime.
No model weights. These are structured prompt blueprints. Bring your own LLM (DeepSeek, OpenAI, Claude, local Llama — anything OpenAI-compatible).
🤗 Discovery surfaces:
🌍 Space: w9000/enclave — product… See the full description on the dataset page: https://huggingface.co/datasets/w9000/enclave-character-recipes.recipes-20krecipe_Ingredient_Datasetrecipesdata_recipeV1
영어(8k)
ShareGPT : 3.24k
Claued 3 Opus : 1.76k
Lima : 1k
Slimorca : 2k
한국어 (2k)
Alpca-GPT4 : 1k
SMR : 0.5k
MMLU : 0.5k
V2
영어(4k)
ShareGPT : 1k
Lima : 1k
Slimorca : 1k
Math : 1k
한국어 (6k)
ShareGPT : 1k
Multi-turn : 1k
Alpca-GPT4 : 2k
SMR : 1k
MMLU : 1k
V3
영어(7k)
ShareGPT : 3k
Claued 3 Opus : 1k
Lima : 1k
Slimorca : 2k
한국어 (3k)
ShareGPT : 1k
Multi-turn : 1k
MMLU : 1k
recipe-modifications
Hebrew Recipe Modification Dataset
Overview
10,058 Hebrew comment threads from YouTube cooking channels, annotated for recipe modification extraction using a three-pass Teacher-Student distillation approach.
Task
Token-level BIO tagging to extract recipe modifications from Hebrew user comments. Four modification aspects: SUBSTITUTION, QUANTITY, TECHNIQUE, ADDITION.
Dataset Structure
Raw Data
threads.jsonl — 10,058 comment threads (top… See the full description on the dataset page: https://huggingface.co/datasets/DanielDDDS/recipe-modifications.Italian-Recipes-AlpacaRecipeFusionv3recipe-extractor-datasetThis dataset was created to finetune a small gemma-3 270M model to extract valid, correct JSON-LD objects from a recipe blog/social media post.
It's a completely synthethic dataset. I used Deepseek v3.2 to create the blog posts, JSON-LD extraction, and the reasoning traces.
The blogs were created based on this Kaggle All Recipes Dataset.
The pipeline to generate this dataset is available on github
recipe_datasetrecipes100
