CoolFace
Datasetpublic

incrediblecrab/llmmm-recipe-ingredients

llmmm recipe ingredients This extract contains 4,653,430 canonical ingredient records, 36,707,624 ingredient slots and 1,790 canonical ingredient names from 29 source groups. Each record has normalized ingredient facts and, where the source recorded them, its original title and ingredient lines. Cooking instructions are not included, so a record is not a complete recipe. Counts describe the complete canonical corpus, not a sample or a count of unique content: duplicate… See the full description on the dataset page: https://huggingface.co/datasets/incrediblecrab/llmmm-recipe-ingredients.

sourceHugging Faceotherupdated 3d agoView on Hugging Face
1likes112downloads
Dataset Card

llmmm recipe ingredients

This extract contains 4,653,430 canonical ingredient records, 36,707,624 ingredient slots and 1,790 canonical ingredient names from 29 source groups. Each record has normalized ingredient facts and, where the source recorded them, its original title and ingredient lines. Cooking instructions are not included, so a record is not a complete recipe. Counts describe the complete canonical corpus, not a sample or a count of unique content: duplicate ingredient sets remain.

The train split preserves every original zero-based id, including 41,172 singleton records and records without links. ingredient_ids and ingredients are aligned, sorted canonical ID sets and their exact vocabulary names. source and language retain the index's source identifiers and language labels.

Source-reported positive total_minutes are available for 681,275 of 4,653,430 records; source-reported servings for 353,946 of 4,653,430 records. Unknown values are null, not NaN, zero, guessed totals or sums of component times. Serving counts do not scale ingredient quantities. These source facts are not independently measured.

source_url retains recorded original HTTP(S) links for 2,292,411 records; 2,361,019 records have null URLs. Links recorded without a scheme use http://. Missing, invalid or sensitive links are not replaced with generated URLs.

recipe_link is the link a result card opens, derived from source_url by per-site rules in `recipe_links.py`. Each recorded site was checked on September 23, 2026 with a sample of its links, usually 12, and www.povarenok.ru again on September 24, when its pages were failing. A site's cards open the recorded page over HTTPS, on the site's current host and with the title slug NYT Cooking now requires, if at least five in six sampled pages opened showing their recipe title; otherwise the Internet Archive's copy of the recorded URL, if at least five in six sampled archived copies did; otherwise no link. The rule is judged per site, so an individual page may still have moved or been removed. The per-site counts record each decision. link_status is source for 1,652,643 records that link to the recorded page, archive for 266,182 records that link to its Internet Archive copy, offline for 373,586 records whose site passed neither check, so recipe_link is null, and none for the 2,361,019 records without a recorded URL.

index/text/ holds the text shown on the browser demo's result cards: 4,503,160 recorded titles and 39,045,568 ingredient lines from 4,503,161 records, in ID order, 2,048 records per gzip JSON shard. The index declares their normalization: Recorded titles and ingredient lines, one line per catalog separator or embedded line break. Complete HTML character references are decoded, whitespace is collapsed, HTML ingredient tables become one 'name: amount' line per row, lines recorded as name and amount fields (povarenok-detail, taiwan-1.8k) become 'name: amount', a 03-povarenok amount recorded as null (catalog text 'name: None') is dropped, and a whole list recorded on one line becomes one line per item (joined with ' ; ', or '|' in filipino-2k); wording is otherwise unchanged. Null means the source recorded none. No instructions. These fields are not in the Parquet train split. Descriptions, authors and images are not included either; consult a recorded source page for the method, where a link is available.

Canonical matching does not recover quantities or every compound constituent, and exclusions are not an allergy-safety guarantee.

Model and its separate terms | Browser demo | Exact source revision `974147bb2ba6334f243c1f8fc40eb5eebc8759f9`

index/ preserves the index's compressed arrays, URL shards, card link shards, text manifest and text shards byte-for-byte; only its publication provenance changes. dataset-manifest.json records the source revision, input and copied index hashes, Parquet counts and every other file's bytes and SHA256. Packaging stages a release locally; it does not upload it.

Use

The maintainer confirmed permission to publish these ingredient facts, titles and ingredient lines. license: other is not a grant of rights to source-page instructions, prose or photos, or to the separate model weights. Source-page material and model weights retain their separate terms.

Acknowledgements

Thanks to the original recipe contributors and the dataset contributors and curators identified in the source inventory. Original source identifiers and recorded source URLs are retained for attribution and provenance wherever available.