Outside of work, one of my most frequent uses of LLMs is while I am cooking. I don’t like to waste food, so I frequently have meals whose inspiration is “these cans are past their expiration date and I have carrots in the fridge that look questionable”. I find that LLMs can usually give good suggestions in these situations, and I’ve talked to others with similar experiences. In fact, about 1% of ChatGPT conversations are related to cooking and ingredients.
However, given the adage that “cooking is an art, but baking is a science,” I was curious about whether my LLM-assisted cooking could extend to recipes with tighter ingredient quantities. If I’m adding salt while cooking, I can taste halfway through so I don’t put too much, but you can’t do that when you’re baking a Pie Crust or making bread.
In this post, I’m going to explore:
I used GPT 5-6 Luna for this comparison. I did this for a few reasons:
I chose recipes that had fairly tight tolerances on relative quantities between ingredients, because I wanted to identify clear “failures” when the LLM generated a recipe outside the bounds.
I also chose recipes that would scale ingredients in a linear fashion, as opposed to those for which some ingredients can scale but others may not need to. For example, the yeast-to-flour ratio in a bread recipe is important, but when you double a bread recipe, you don’t necessarily have to double the yeast (at least that’s what I would guess given the couple of biology classes I’ve taken; my wife may contend that there were some occasions with slow-rising rolls where doubling the yeast may have been helpful).
Anyway, the foods I chose to use for this experiment were:
To assign pass/fail results to a recipe, I extracted quantities into a structured format with a second LLM call using the same model. I used a stronger model (Claude Opus 5) to evaluate ingredient quantity extraction and found 0/28 missed-or-misread ingredients, 0/28 false precision, but 2/28 (7.1%) wrong line semantics. These error rates were acceptable, so I continued with LLM extraction using GPT 5-6 Luna.
When a range was specified by a generated recipe, the model parsed both the max and min values and the scorer took the midpoint as the recipe’s actual value for the purposes of ratio accuracy evaluation. If no quantities were specified (e.g., “a few”) for an ingredient involved in an outcome ratio, the recipe was marked as unscorable. For some recipes, the generated recipe specified a maximum or minimum amount of an ingredient and suggested adding gradually until a consistency was met; in that situation, the stated quantity was used to calculate the ratio instead of assuming a smaller or larger value.
After ingredients were extracted, I converted them to grams and computed ratios observed in the recipe. To determine recipes with successful ratios, I looked at results for 5–10 actual tested recipes I found on the internet. Note that for Panna Cotta, I found a Seattle Times article where an author tested recipes and reported results, so I used that one for the ranges.
Here are the acceptable ranges I’m using in my analysis:
| Recipe | Ratio(s) checked | Ground truth | Basis |
|---|---|---|---|
| Pie Crust | fat:flour (by weight) | pass: 0.68–0.89 | Convergence across 8 independently-tested recipes |
| water:flour (by weight) | pass: 0.28–0.50 | Same 8-recipe set. Shortening-only recipe sets the high end (shortening carries no water, so more must be added vs. butter) | |
| Panna Cotta | gelatin:liquid (by weight %) | pass: 0.65%–3.9% | Documented range across tested/published recipes by Seattle Times (set/don’t-set language). |
| Pastry Cream | cornstarch:dairy (by weight %) | pass: 4.1%–8.3% | Convergence across 12 independently-tested recipes (King Arthur x2, Scotch & Scones, Sally’s Baking Addiction, Allrecipes, Serious Eats, Preppy Kitchen, Martha Stewart, Stay at Home Chef, Natasha’s Kitchen, Food Network, one additional source). Dense cluster ~4.8%–6.6%; King Arthur GF (4.1%) and Preppy Kitchen’s richer 6-yolk version (8.3%) are the outer bounds. |
| Choux Pastry | egg:flour (by weight) | pass: 1.3:1–2.3:1 | Convergence across 10 independently-tested recipes (Ruhlman canonical, America’s Test Kitchen, Food Network, Sally’s Baking Addiction, Allrecipes, Bonni Bakery, King Arthur, The Flavor Bender, The Buttery Crust, one additional source). King Arthur’s tested recipe (1.33:1) sets the low end; Bonni Bakery’s own stated 5–6 egg range (1.92:1–2.31:1) sets the high end. Egg weight assumed ~50g/large egg, no shell, unless a source states shell-inclusive weight (see caveat below). |
Note that unlike the rest of the text in this post, the above table was generated by Claude. I’m including this note to follow my AI use policy.
I used a very simple prompt to generate recipes:
Can you give me a recipe for [adjective] [requested food] [quantity modifier]?
In order to evaluate the effect of asking for a “foolproof” recipe, I changed [adjective] to be either “really tasty,” “foolproof,” or “avant garde.”
I replaced [requested food] with the food I wanted a recipe generated for.
For [quantity modifier], I used either a blank, 6 people (to approximate the most common usage), 35 people (to approximate the max of what a home cook might be cooking for), or 173 people (to evaluate LLM scaling in an out-of-distribution quantity).
I did not provide a system prompt or any tools to the model.
For each unique prompt, I generated 5 recipes, which led to 240 generations (4 foods, 3 adjectives, 4 serving sizes, and 5 generations). Two recipes were not scorable, so they were dropped from this analysis.
Here’s an example of one of the recipes generated, how ingredients were extracted, and results.
==============================================================================
CONDITION choux pastry | foolproof | serving size: 6 | repeat 1
PROMPT 'Can you give me a recipe for foolproof choux pastry for 6 people?'
==============================================================================
--- 1. Generated recipe (verbatim model response) ---
## Foolproof choux pastry for 6
Makes about **12–15 small profiteroles** or **6–8 éclairs**.
### Ingredients
- 125 ml water
- 60 g unsalted butter, cubed
- ¼ tsp salt
- 1 tsp sugar — optional, useful for sweet choux
- 75 g plain/all-purpose flour, sifted
- 2–3 large eggs, lightly beaten — about **100–120 g without shells**
### Method
1. **Prepare the oven and tray**
Heat the oven to **200°C conventional / 180°C fan**. Line a baking tray with parchment.
2. **Boil the butter mixture**
Put the water, butter, salt and sugar in a saucepan over medium heat. Heat until the butter has melted and the mixture reaches a full boil.
3. **Add the flour all at once**
Remove the pan from the heat and tip in the flour immediately. Stir vigorously with a wooden spoon until smooth.
4. **Dry the dough**
Return the pan to medium-low heat and stir for **1–2 minutes**, until the dough forms a ball and leaves a thin film on the base of the pan. This step helps prevent soggy choux.
5. **Add the eggs gradually**
Transfer the dough to a bowl and cool for 5 minutes. Beat in the egg a little at a time, mixing well after each addition. Stop when the dough is smooth, glossy and falls from the spoon in a thick **V-shaped ribbon**. You may not need all the egg.
6. **Pipe**
Transfer to a piping bag and pipe walnut-sized mounds, leaving space between them. Wet your fingertip and gently flatten any peaks.
7. **Bake**
Bake at **200°C / 180°C fan for 15 minutes**, then reduce to **180°C / 160°C fan** and bake for another **15–20 minutes**, until deeply golden and firm.
8. **Dry the shells**
Turn off the oven, leave the door slightly ajar and let them sit inside for **10 minutes**. Transfer to a rack and cool completely before filling.
### Key foolproof tips
- Weigh the flour and eggs if possible.
- Don’t add all the egg automatically—the dough should be glossy but not runny.
- Don’t open the oven during the first 25 minutes.
- Bake until properly golden; pale choux often collapses.
- For extra-dry shells, prick each one with a skewer after baking and return them to the turned-off oven for 5–10 minutes.
--- 2. Extracted ingredient lines ---
| name | quantity | quantity_low | quantity_high | unit | component | is_core | option_group | is_default_option | |
|---|---|---|---|---|---|---|---|---|---|
| 0 | water | 125 | nan | nan | ml | choux pastry | True | False | |
| 1 | unsalted butter, cubed | 60 | nan | nan | g | choux pastry | True | False | |
| 2 | salt | 0.25 | nan | nan | tsp | choux pastry | True | False | |
| 3 | sugar | 1 | nan | nan | tsp | choux pastry | True | False | |
| 4 | plain/all-purpose flour, sifted | 75 | nan | nan | g | choux pastry | True | False | |
| 5 | eggs, lightly beaten | nan | 100 | 120 | g | choux pastry | True | False |
--- 3. Scored ratio(s) ---
egg_to_flour: egg / flour = 1.467 (band 1.3-2.3, center 1.7999999999999998) -> PASS
recipe-level verdict: PASS
I think that most people using ChatGPT for cooking want to know: “If I ask ChatGPT for a Pie Crust recipe, is it going to fall apart?” So, for each recipe I evaluated whether it fell within the acceptable recipe ranges.
| Recipe | # Generated and Parsed | % failures |
|---|---|---|
| Pie Crust | 59 | 0.389831 |
| Panna Cotta | 59 | 0.0338983 |
| Pastry Cream | 60 | 0 |
| Choux | 60 | 0.0166667 |

ChatGPT produced recipes with ingredient ratios outside of the acceptable tolerances 39% of the time for the most commonly baked recipe here (Pie Crust) and 11% of the time overall. Although my baking failures are more frequently of the “forgot to set a timer and now the fire alarm is going off” variety rather than “too much fat content in my dough” variety, it still seems like I might be better off avoiding injecting additional variation in the end results from an unstable recipe.
Note that the 11% number is probably not a good estimate of how often recipe generation will fail in the real world. Because failure rates vary widely by food type, and I only included four food types, I would expect a very different number if I were to randomly sample a large number across a larger variety of food types.
If ChatGPT is not super reliable at generating recipes, are there low-effort things that users can do to make it more reliable? The space of low-effort things to do is pretty small because the alternative of just looking up a recipe blog is pretty easy, but I was curious if asking for a “foolproof” recipe would pull the quantities within the distribution.

Optimization terminated successfully.
Current function value: 0.260143
Iterations 8
Logit Regression Results
==============================================================================
Dep. Variable: recipe_passed_int No. Observations: 178
Model: Logit Df Residuals: 173
Method: MLE Df Model: 4
Date: Mon, 07 Sep 2026 Pseudo R-squ.: 0.3744
Time: 17:03:06 Log-Likelihood: -46.305
converged: True LL-Null: -74.017
Covariance Type: nonrobust LLR p-value: 2.649e-11
===============================================================================================================================
coef std err z P>|z| [0.025 0.975]
-------------------------------------------------------------------------------------------------------------------------------
Intercept 5.4056 1.182 4.574 0.000 3.089 7.722
C(recipe)[T.panna_cotta] -0.7332 1.251 -0.586 0.558 -3.185 1.719
C(recipe)[T.pie_dough] -3.9946 1.080 -3.699 0.000 -6.111 -1.878
C(adjective, Treatment(reference='control'))[T.foolproof] -0.4618 0.729 -0.633 0.526 -1.891 0.967
C(adjective, Treatment(reference='control'))[T.avant_garde] -2.1865 0.701 -3.121 0.002 -3.559 -0.814
===============================================================================================================================
==============================================================================
Results comparing C(adjective, Treatment(reference='control'))[T.avant_garde] and C(adjective, Treatment(reference='control'))[T.foolproof]
Test for Constraints
==============================================================================
coef std err z P>|z| [0.025 0.975]
------------------------------------------------------------------------------
c0 -1.7247 0.638 -2.701 0.007 -2.976 -0.473
==============================================================================
Note: I excluded Pastry Cream from this model because there were no failures
From my analysis, it doesn’t look like asking for a “foolproof” recipe reduces out-of-band ingredient ratios versus the “really tasty” control. It may be that more rigorous prompting, extended reasoning, or custom tool use can reduce out-of-band occurrence, but again, all of those are getting into the “more difficult than scrolling through a recipe blogger’s story about why they love this recipe” territory.
However, it does look like asking for an “avant garde” recipe does increase the risk of a bad recipe generation (both against the “foolproof” and “really tasty” adjectives). So, it may be wise to avoid using that language if you are prompting for recipes and want a recipe that works.
When you’re cooking for more than a couple of people, you can do one of two things:
In theory, the first strategy (i.e. requesting a recipe for 6 people and multiplying by 6) should give you the same quantity of ingredients as the second (i.e. requesting a recipe for 36 people directly). However, the second strategy might go wrong because LLMs cannot natively do arithmetic and need a tool to reliably add or multiply. So, unless the quantity you are asking for is something that was seen in its training distribution, there is a risk of the ingredient quantities falling out of sync with each other, or of the quantities needed being incorrect for the number of mouths the recipe needs to feed.

Higher scales did lead to higher variability in reported ratios in some cases, but it was not consistent across all recipes. I was somewhat surprised—I expected that at a recipe size of 173, the LLM generation would be less stable and substantially more likely to produce bad ratios.
Anyway, here’s a linear model testing this statistically. Pie Crust (the reference condition) had significantly more variability in its generated ratios at higher scales, but Choux Pastry had a significant interaction in the other direction.
OLS Regression Results
==============================================================================
Dep. Variable: sd R-squared: 0.367
Model: OLS Adj. R-squared: 0.247
Method: Least Squares F-statistic: 3.062
Date: Mon, 07 Sep 2026 Prob (F-statistic): 0.0120
Time: 17:03:08 Log-Likelihood: 29.346
No. Observations: 45 AIC: -42.69
Df Residuals: 37 BIC: -28.24
Df Model: 7
Covariance Type: nonrobust
===========================================================================================================
coef std err t P>|t| [0.025 0.975]
-----------------------------------------------------------------------------------------------------------
Intercept 0.0560 0.090 0.624 0.537 -0.126 0.238
C(recipe)[T.panna_cotta] 0.0820 0.155 0.528 0.601 -0.233 0.397
C(recipe)[T.pastry_cream] 0.0640 0.155 0.412 0.683 -0.251 0.379
C(recipe)[T.choux] 0.1723 0.155 1.108 0.275 -0.143 0.487
log10_scale_n 0.1478 0.055 2.689 0.011 0.036 0.259
log10_scale_n:C(recipe)[T.panna_cotta] -0.1086 0.095 -1.141 0.261 -0.301 0.084
log10_scale_n:C(recipe)[T.pastry_cream] -0.1628 0.095 -1.711 0.095 -0.356 0.030
log10_scale_n:C(recipe)[T.choux] -0.2122 0.095 -2.229 0.032 -0.405 -0.019
==============================================================================
Omnibus: 2.396 Durbin-Watson: 1.715
Prob(Omnibus): 0.302 Jarque-Bera (JB): 1.486
Skew: 0.405 Prob(JB): 0.476
Kurtosis: 3.368 Cond. No. 23.9
==============================================================================
Notes:
[1] Standard Errors assume that the covariance matrix of the errors is correctly specified.

| Recipe | Ingredient | % diff @ 35 | % diff @ 173 |
|---|---|---|---|
| Pie Crust | Fat | 71.4 | 51.9 |
| Pie Crust | Flour | 66.1 | 43.1 |
| Pie Crust | Water | 76.3 | 74.4 |
| Panna Cotta | Gelatin | 10.2 | 18.9 |
| Panna Cotta | Liquid | -9.1 | -6.3 |
| Pastry Cream | Cornstarch | -48.6 | -44.5 |
| Pastry Cream | Dairy | -45.1 | -44.8 |
| Choux | Egg | -59.3 | -75.7 |
| Choux | Flour | -65.7 | -75.7 |
It turns out that the LLM was wildly inaccurate in scaling recipes. It was fairly close for Panna Cotta, but for other recipes, the median LLM-scaled result was over 40% different from the result coming from multiplying the 6-person recipe. (Note that this could be a sign that the 6-person recipe is incorrect as well, but either way, it seems that there is a problem.)
I’m fairly surprised that this type of linear scaling was so inaccurate in this context. It seems like this kind of task would be really low-hanging fruit for Reinforcement Learning with Verifiable Rewards, and you would think that OpenAI would prioritize making the “cooking with ChatGPT” experience better given their usage guide advertising the functionality. But, I suppose that most people who set out to cook Choux Pastry for 173 people will probably triple-check the recipe before starting up their industrial mixer rather than using a one-shot recipe from ChatGPT.
In the future, I’ll probably continue to use an LLM for cooking (because I can taste and correct during the process), but for anything that seems like it could be messed up with bad ingredient ratios, I’ll lean on a cookbook or human-tested recipe from a website. And, the next time I am cooking for 173 people, I will definitely scale linearly from a recipe I know works instead of generating a new one from scratch. We will just hope that if I do undertake the “baking for 173” task, I do not get too over-enthusiastic and accidentally “scale the baking time” linearly as well, for the sake of the ears of anyone else who may be in range of the fire alarm.
Thanks for reading! And happy baking!