datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
recipes_data_food.comfood-recipesrecipes-with-nutritionLeverage our Recipes Dataset to explore one of the most comprehensive collections of structured recipe and nutrition data.
This dataset contains 39,447 recipes, enriched with detailed nutritional information, ingredient breakdowns, dietary classifications, and recipe metadata. Each record includes both raw recipe text (such as ingredient lines) and structured fields (nutritional values, cuisine type, diet/health labels, and more).
Designed as a rich, high-quality resource, this dataset… See the full description on the dataset page: https://huggingface.co/datasets/datahiveai/recipes-with-nutrition.recipeTurkish-Recipe-Corpus
Turkish Recipe Corpus (TRC-30K)
Türkçe'nin en kapsamlı açık kaynaklı tarif veri seti.74,768 Türkçe tariften oluşan, 3 farklı NLP görevine hazır yapılandırılmış corpus.
Ethosoft Research · huggingface.co/Ethosoft · ethosoft.org
Dataset Özeti
Değer
Toplam tarif
74,768
Dil
Türkçe (tr)
Lisans
CC BY 4.0
Konfigürasyonlar
3 (structured, instruction, ingredient2recipe)
Split
Train / Validation / Test (80 / 10 / 10)
Ortalama adım sayısı
5.23 /… See the full description on the dataset page: https://huggingface.co/datasets/Ethosoft/Turkish-Recipe-Corpus.llmmm-recipe-ingredients
llmmm recipe ingredients
This extract contains 4,653,430 canonical ingredient records, 36,707,624 ingredient slots and 1,790 canonical ingredient names from 29 source groups. Each record has normalized ingredient facts and, where the source recorded them, its original title and ingredient lines. Cooking instructions are not included, so a record is not a complete recipe. Counts describe the complete canonical corpus, not a sample or a count of unique content: duplicate… See the full description on the dataset page: https://huggingface.co/datasets/incrediblecrab/llmmm-recipe-ingredients.parsebench-recipe-runs
ParseBench recipe runs
Experiment log for document-parsing runs on the public ParseBench test
subset (llamaindex/ParseBench), scored with the official open-source
ParseBench evaluator (run-llama/ParseBench).
Recipes combine CLI coding agents doing vision parsing with deterministic
PDF text-layer tools (word-bbox snapping, style extraction from span flags
and vector-drawing geometry).
Sample: official test subset - 12 single-page PDFs, 3 per category
(chart / layout / table /… See the full description on the dataset page: https://huggingface.co/datasets/sleepyheeler/parsebench-recipe-runs.kerala-recipes
🥥 Kerala Recipes Dataset
A hand-crafted dataset of 89 authentic Kerala (South Indian) home-cooked dishes, created with exact per-person ingredient quantities, household kitchen measures (cups, tablespoons, piece counts), and traditional meal pairing recommendations.
🍃 Why This Dataset Was Created
Most recipe datasets online are flat lists of ingredients copied from generic blogs. They don't tell you:
How much ingredient quantity you actually need when cooking… See the full description on the dataset page: https://huggingface.co/datasets/Fathi7ma/kerala-recipes.indonesian-recipes
Resep Masakan Indonesia 🍛
Kumpulan resep masakan Indonesia autentik — dari rendang sampai es cendol, lengkap dengan bahan, langkah, tingkat kesulitan, waktu, dan daerah asal.
Kenapa dataset ini ada?
Resep adalah salah satu konten paling dicari untuk LLM (assistant masak) — tapi dataset resep Indonesia di HF nyaris kosong (cuma 1 yang 34 likes). Gw isi gap itu dengan resep-resep yang benar-benar asli Indonesia, bukan versi western yang diterjemahkan.… See the full description on the dataset page: https://huggingface.co/datasets/LorthGyu/indonesian-recipes.recipesense-datasuno-style-recipes
Source, notebook and weekly sync on GitHub
Suno Style Recipes — 564 documented music styles
Style prompt, BPM, weirdness, style influence and canonical song structure for
564 music styles, formatted for AI music generators. Published by MUSAI · musaisong.app.
https://musaisong.app/en/styles
style_prompt — paste into the "Style of Music" field
bpm · weirdness · style_influence — the three settings that change the output
structure — the canonical section order for that genre
url… See the full description on the dataset page: https://huggingface.co/datasets/musaisong/suno-style-recipes.recipe-cleaned
Recipe Cleaned Dataset
Dataset Summary
This dataset is a structured and cleaned collection of recipe data derived from the Food.com Recipes and Interactions dataset. It is designed for ingredient-based personalization, machine learning training, and interactive recommendation systems. The dataset integrates a hierarchical ingredient taxonomy, standardized nutrition information, and categorical metadata (e.g., diet tags, cuisine attributes, region) to support downstream… See the full description on the dataset page: https://huggingface.co/datasets/Iris314/recipe-cleaned.fitfuel-recipes
🥗 FitFuel — AI-Powered Nutrition & Fitness Recipe App
Try the live app: https://huggingface.co/spaces/amitbenavraham/fitfuel-app
A synthetic recipe dataset and semantic recommendation pipeline, built for
a fitness nutrition application.
📓 Project Notebook
The full project notebook (data generation, cleaning, EDA, embeddings,
vision/generation, and application code) is available here:
FitFuel Final Project.ipynb
📋 Project Overview
FitFuel is an… See the full description on the dataset page: https://huggingface.co/datasets/amitbenavraham/fitfuel-recipes.indian-recipe-datasetPAID-recipes-normalizedmise-recipes
🍳 Mise Recipes
A synthetic dataset of 10,000 cooking recipes across 8 cuisines, generated with a pretrained language model and cleaned/validated with a deterministic EDA pipeline. Built for Mise, a "Duolingo-for-cooking" app: given a cuisine, it recommends similar recipes and generates a progressive cooking lesson.
Intro to Data Science final project, Reichman University.
Dataset at a glance
10,000 recipes, 1,250 per cuisine (perfectly balanced).
8 cuisines —… See the full description on the dataset page: https://huggingface.co/datasets/idoyaaran/mise-recipes.cot-oracle-qwen3-8b-onpolicy-recipe
CoT Activation Oracle — On-Policy Qwen3-8B Training Recipe
A reproduction of the on-policy Qwen3-8B training mixture from
Building Better Activation Oracles
(Bauer, De Schamphelaere, Karvonen, Luick, Nanda).
This repository is a recipe card only — it documents the exact dataset
mixture, points at every source on the Hub, and gives regeneration instructions
for the pieces that are no longer available upstream. No third-party data is
re-hosted here; original datasets are linked… See the full description on the dataset page: https://huggingface.co/datasets/francescortu/cot-oracle-qwen3-8b-onpolicy-recipe.recipes_for_dishes_and_food_with_vectors_sentiment_ners
Description in English:
The dataset is collected from Russian-language Telegram channels with various food recipes,The dataset was collected and tagged automatically using the data collection and tagging service Scoutie.Try Scoutie and collect the same or another dataset using link for FREE.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/recipes_for_dishes_and_food_with_vectors_sentiment_ners.ecuador-recipesFood_Recipesai-blessed_raw_recipesVV_recipes_score_gpt
Dataset Card for "VV_recipes_score_gpt"
More Information needed
PAID-recipesitalian-recipesPAID-recipes-newnepali-recipesrecipes_data_food.comindonesian-recipes
Indonesian Recipes
A structured collection of Indonesian recipes for fine-tuning text-generation models. Each row is a single recipe with a title, an ingredient list, and ordered preparation steps.
Schema
Column
Type
Description
title
string
Recipe name
ingredients
list<string>
One item per ingredient line
steps
list<string>
Ordered preparation steps
num_ingredients
int
len(ingredients)
num_steps
int
len(steps)
char_count
int
Total characters… See the full description on the dataset page: https://huggingface.co/datasets/junwatu/indonesian-recipes.recipe-interactionsIndian_recipe
