Zachary Clement

A place to host my thoughts and side projects

LLM usage for baking: an evaluation on ratio-sensitive recipes

Posted at — Sep 7, 2026

Background

Outside of work, one of my most frequent uses of LLMs is while I am cooking. I don’t like to waste food, so I frequently have meals whose inspiration is “these cans are past their expiration date and I have carrots in the fridge that look questionable”. I find that LLMs can usually give good suggestions in these situations, and I’ve talked to others with similar experiences. In fact, about 1% of ChatGPT conversations are related to cooking and ingredients.

However, given the adage that “cooking is an art, but baking is a science,” I was curious about whether my LLM-assisted cooking could extend to recipes with tighter ingredient quantities. If I’m adding salt while cooking, I can taste halfway through so I don’t put too much, but you can’t do that when you’re baking a Pie Crust or making bread.

In this post, I’m going to explore:

Experimental setup

Model selection

I used GPT 5-6 Luna for this comparison. I did this for a few reasons:

Food Selection

I chose recipes that had fairly tight tolerances on relative quantities between ingredients, because I wanted to identify clear “failures” when the LLM generated a recipe outside the bounds.

I also chose recipes that would scale ingredients in a linear fashion, as opposed to those for which some ingredients can scale but others may not need to. For example, the yeast-to-flour ratio in a bread recipe is important, but when you double a bread recipe, you don’t necessarily have to double the yeast (at least that’s what I would guess given the couple of biology classes I’ve taken; my wife may contend that there were some occasions with slow-rising rolls where doubling the yeast may have been helpful).

Anyway, the foods I chose to use for this experiment were:

Recipe evaluation

To assign pass/fail results to a recipe, I extracted quantities into a structured format with a second LLM call using the same model. I used a stronger model (Claude Opus 5) to evaluate ingredient quantity extraction and found 0/28 missed-or-misread ingredients, 0/28 false precision, but 2/28 (7.1%) wrong line semantics. These error rates were acceptable, so I continued with LLM extraction using GPT 5-6 Luna.

When a range was specified by a generated recipe, the model parsed both the max and min values and the scorer took the midpoint as the recipe’s actual value for the purposes of ratio accuracy evaluation. If no quantities were specified (e.g., “a few”) for an ingredient involved in an outcome ratio, the recipe was marked as unscorable. For some recipes, the generated recipe specified a maximum or minimum amount of an ingredient and suggested adding gradually until a consistency was met; in that situation, the stated quantity was used to calculate the ratio instead of assuming a smaller or larger value.

After ingredients were extracted, I converted them to grams and computed ratios observed in the recipe. To determine recipes with successful ratios, I looked at results for 5–10 actual tested recipes I found on the internet. Note that for Panna Cotta, I found a Seattle Times article where an author tested recipes and reported results, so I used that one for the ranges.

Here are the acceptable ranges I’m using in my analysis:

Recipe Ratio(s) checked Ground truth Basis
Pie Crust fat:flour (by weight) pass: 0.68–0.89 Convergence across 8 independently-tested recipes
water:flour (by weight) pass: 0.28–0.50 Same 8-recipe set. Shortening-only recipe sets the high end (shortening carries no water, so more must be added vs. butter)
Panna Cotta gelatin:liquid (by weight %) pass: 0.65%–3.9% Documented range across tested/published recipes by Seattle Times (set/don’t-set language).
Pastry Cream cornstarch:dairy (by weight %) pass: 4.1%–8.3% Convergence across 12 independently-tested recipes (King Arthur x2, Scotch & Scones, Sally’s Baking Addiction, Allrecipes, Serious Eats, Preppy Kitchen, Martha Stewart, Stay at Home Chef, Natasha’s Kitchen, Food Network, one additional source). Dense cluster ~4.8%–6.6%; King Arthur GF (4.1%) and Preppy Kitchen’s richer 6-yolk version (8.3%) are the outer bounds.
Choux Pastry egg:flour (by weight) pass: 1.3:1–2.3:1 Convergence across 10 independently-tested recipes (Ruhlman canonical, America’s Test Kitchen, Food Network, Sally’s Baking Addiction, Allrecipes, Bonni Bakery, King Arthur, The Flavor Bender, The Buttery Crust, one additional source). King Arthur’s tested recipe (1.33:1) sets the low end; Bonni Bakery’s own stated 5–6 egg range (1.92:1–2.31:1) sets the high end. Egg weight assumed ~50g/large egg, no shell, unless a source states shell-inclusive weight (see caveat below).

Note that unlike the rest of the text in this post, the above table was generated by Claude. I’m including this note to follow my AI use policy.

Prompt

I used a very simple prompt to generate recipes:

Can you give me a recipe for [adjective] [requested food] [quantity modifier]?

In order to evaluate the effect of asking for a “foolproof” recipe, I changed [adjective] to be either “really tasty,” “foolproof,” or “avant garde.”

I replaced [requested food] with the food I wanted a recipe generated for.

For [quantity modifier], I used either a blank, 6 people (to approximate the most common usage), 35 people (to approximate the max of what a home cook might be cooking for), or 173 people (to evaluate LLM scaling in an out-of-distribution quantity).

I did not provide a system prompt or any tools to the model.

Sample Size

For each unique prompt, I generated 5 recipes, which led to 240 generations (4 foods, 3 adjectives, 4 serving sizes, and 5 generations). Two recipes were not scorable, so they were dropped from this analysis.

Example Generation

Here’s an example of one of the recipes generated, how ingredients were extracted, and results.

==============================================================================
CONDITION   choux pastry | foolproof | serving size: 6 | repeat 1
PROMPT      'Can you give me a recipe for foolproof choux pastry for 6 people?'
==============================================================================

--- 1. Generated recipe (verbatim model response) ---

## Foolproof choux pastry for 6

Makes about **12–15 small profiteroles** or **6–8 éclairs**.

### Ingredients
- 125 ml water  
- 60 g unsalted butter, cubed  
- ¼ tsp salt  
- 1 tsp sugar — optional, useful for sweet choux  
- 75 g plain/all-purpose flour, sifted  
- 2–3 large eggs, lightly beaten — about **100–120 g without shells**

### Method

1. **Prepare the oven and tray**  
   Heat the oven to **200°C conventional / 180°C fan**. Line a baking tray with parchment.

2. **Boil the butter mixture**  
   Put the water, butter, salt and sugar in a saucepan over medium heat. Heat until the butter has melted and the mixture reaches a full boil.

3. **Add the flour all at once**  
   Remove the pan from the heat and tip in the flour immediately. Stir vigorously with a wooden spoon until smooth.

4. **Dry the dough**  
   Return the pan to medium-low heat and stir for **1–2 minutes**, until the dough forms a ball and leaves a thin film on the base of the pan. This step helps prevent soggy choux.

5. **Add the eggs gradually**  
   Transfer the dough to a bowl and cool for 5 minutes. Beat in the egg a little at a time, mixing well after each addition. Stop when the dough is smooth, glossy and falls from the spoon in a thick **V-shaped ribbon**. You may not need all the egg.

6. **Pipe**  
   Transfer to a piping bag and pipe walnut-sized mounds, leaving space between them. Wet your fingertip and gently flatten any peaks.

7. **Bake**  
   Bake at **200°C / 180°C fan for 15 minutes**, then reduce to **180°C / 160°C fan** and bake for another **15–20 minutes**, until deeply golden and firm.

8. **Dry the shells**  
   Turn off the oven, leave the door slightly ajar and let them sit inside for **10 minutes**. Transfer to a rack and cool completely before filling.

### Key foolproof tips
- Weigh the flour and eggs if possible.
- Don’t add all the egg automatically—the dough should be glossy but not runny.
- Don’t open the oven during the first 25 minutes.
- Bake until properly golden; pale choux often collapses.
- For extra-dry shells, prick each one with a skewer after baking and return them to the turned-off oven for 5–10 minutes.

--- 2. Extracted ingredient lines ---
name quantity quantity_low quantity_high unit component is_core option_group is_default_option
0 water 125 nan nan ml choux pastry True False
1 unsalted butter, cubed 60 nan nan g choux pastry True False
2 salt 0.25 nan nan tsp choux pastry True False
3 sugar 1 nan nan tsp choux pastry True False
4 plain/all-purpose flour, sifted 75 nan nan g choux pastry True False
5 eggs, lightly beaten nan 100 120 g choux pastry True False
--- 3. Scored ratio(s) ---

  egg_to_flour: egg / flour = 1.467  (band 1.3-2.3, center 1.7999999999999998)  -> PASS

  recipe-level verdict: PASS

Can you rely on ChatGPT to give you reliable quantities for baking?

I think that most people using ChatGPT for cooking want to know: “If I ask ChatGPT for a Pie Crust recipe, is it going to fall apart?” So, for each recipe I evaluated whether it fell within the acceptable recipe ranges.

Recipe # Generated and Parsed % failures
Pie Crust 59 0.389831
Panna Cotta 59 0.0338983
Pastry Cream 60 0
Choux 60 0.0166667

png

ChatGPT produced recipes with ingredient ratios outside of the acceptable tolerances 39% of the time for the most commonly baked recipe here (Pie Crust) and 11% of the time overall. Although my baking failures are more frequently of the “forgot to set a timer and now the fire alarm is going off” variety rather than “too much fat content in my dough” variety, it still seems like I might be better off avoiding injecting additional variation in the end results from an unstable recipe.

Note that the 11% number is probably not a good estimate of how often recipe generation will fail in the real world. Because failure rates vary widely by food type, and I only included four food types, I would expect a very different number if I were to randomly sample a large number across a larger variety of food types.

Does asking for a “foolproof” recipe increase ingredient quantity accuracy?

If ChatGPT is not super reliable at generating recipes, are there low-effort things that users can do to make it more reliable? The space of low-effort things to do is pretty small because the alternative of just looking up a recipe blog is pretty easy, but I was curious if asking for a “foolproof” recipe would pull the quantities within the distribution.

png

Optimization terminated successfully.
         Current function value: 0.260143
         Iterations 8
                           Logit Regression Results                           
==============================================================================
Dep. Variable:      recipe_passed_int   No. Observations:                  178
Model:                          Logit   Df Residuals:                      173
Method:                           MLE   Df Model:                            4
Date:                Mon, 07 Sep 2026   Pseudo R-squ.:                  0.3744
Time:                        17:03:06   Log-Likelihood:                -46.305
converged:                       True   LL-Null:                       -74.017
Covariance Type:            nonrobust   LLR p-value:                 2.649e-11
===============================================================================================================================
                                                                  coef    std err          z      P>|z|      [0.025      0.975]
-------------------------------------------------------------------------------------------------------------------------------
Intercept                                                       5.4056      1.182      4.574      0.000       3.089       7.722
C(recipe)[T.panna_cotta]                                       -0.7332      1.251     -0.586      0.558      -3.185       1.719
C(recipe)[T.pie_dough]                                         -3.9946      1.080     -3.699      0.000      -6.111      -1.878
C(adjective, Treatment(reference='control'))[T.foolproof]      -0.4618      0.729     -0.633      0.526      -1.891       0.967
C(adjective, Treatment(reference='control'))[T.avant_garde]    -2.1865      0.701     -3.121      0.002      -3.559      -0.814
===============================================================================================================================
==============================================================================


Results comparing C(adjective, Treatment(reference='control'))[T.avant_garde]  and C(adjective, Treatment(reference='control'))[T.foolproof]
                             Test for Constraints                             
==============================================================================
                 coef    std err          z      P>|z|      [0.025      0.975]
------------------------------------------------------------------------------
c0            -1.7247      0.638     -2.701      0.007      -2.976      -0.473
==============================================================================

Note: I excluded Pastry Cream from this model because there were no failures

From my analysis, it doesn’t look like asking for a “foolproof” recipe reduces out-of-band ingredient ratios versus the “really tasty” control. It may be that more rigorous prompting, extended reasoning, or custom tool use can reduce out-of-band occurrence, but again, all of those are getting into the “more difficult than scrolling through a recipe blogger’s story about why they love this recipe” territory.

However, it does look like asking for an “avant garde” recipe does increase the risk of a bad recipe generation (both against the “foolproof” and “really tasty” adjectives). So, it may be wise to avoid using that language if you are prompting for recipes and want a recipe that works.

If you are cooking for a crowd, can you ask the LLM for a large quantity or should you ask for a normal quantity and scale up/down manually?

When you’re cooking for more than a couple of people, you can do one of two things:

  1. Ask for a recipe, look at the quantities given, then double/triple the recipe in your head as needed
  2. Tell an LLM exactly how many people you are cooking for and bake based on the response.

In theory, the first strategy (i.e. requesting a recipe for 6 people and multiplying by 6) should give you the same quantity of ingredients as the second (i.e. requesting a recipe for 36 people directly). However, the second strategy might go wrong because LLMs cannot natively do arithmetic and need a tool to reliably add or multiply. So, unless the quantity you are asking for is something that was seen in its training distribution, there is a risk of the ingredient quantities falling out of sync with each other, or of the quantities needed being incorrect for the number of mouths the recipe needs to feed.

Are ratios self-consistent at low or large serving sizes?

png

Higher scales did lead to higher variability in reported ratios in some cases, but it was not consistent across all recipes. I was somewhat surprised—I expected that at a recipe size of 173, the LLM generation would be less stable and substantially more likely to produce bad ratios.

Anyway, here’s a linear model testing this statistically. Pie Crust (the reference condition) had significantly more variability in its generated ratios at higher scales, but Choux Pastry had a significant interaction in the other direction.

                            OLS Regression Results                            
==============================================================================
Dep. Variable:                     sd   R-squared:                       0.367
Model:                            OLS   Adj. R-squared:                  0.247
Method:                 Least Squares   F-statistic:                     3.062
Date:                Mon, 07 Sep 2026   Prob (F-statistic):             0.0120
Time:                        17:03:08   Log-Likelihood:                 29.346
No. Observations:                  45   AIC:                            -42.69
Df Residuals:                      37   BIC:                            -28.24
Df Model:                           7                                         
Covariance Type:            nonrobust                                         
===========================================================================================================
                                              coef    std err          t      P>|t|      [0.025      0.975]
-----------------------------------------------------------------------------------------------------------
Intercept                                   0.0560      0.090      0.624      0.537      -0.126       0.238
C(recipe)[T.panna_cotta]                    0.0820      0.155      0.528      0.601      -0.233       0.397
C(recipe)[T.pastry_cream]                   0.0640      0.155      0.412      0.683      -0.251       0.379
C(recipe)[T.choux]                          0.1723      0.155      1.108      0.275      -0.143       0.487
log10_scale_n                               0.1478      0.055      2.689      0.011       0.036       0.259
log10_scale_n:C(recipe)[T.panna_cotta]     -0.1086      0.095     -1.141      0.261      -0.301       0.084
log10_scale_n:C(recipe)[T.pastry_cream]    -0.1628      0.095     -1.711      0.095      -0.356       0.030
log10_scale_n:C(recipe)[T.choux]           -0.2122      0.095     -2.229      0.032      -0.405      -0.019
==============================================================================
Omnibus:                        2.396   Durbin-Watson:                   1.715
Prob(Omnibus):                  0.302   Jarque-Bera (JB):                1.486
Skew:                           0.405   Prob(JB):                        0.476
Kurtosis:                       3.368   Cond. No.                         23.9
==============================================================================

Notes:
[1] Standard Errors assume that the covariance matrix of the errors is correctly specified.

Do ingredient quantities scale reliably or should you multiply from a low-scale recipe to ensure correct outcome quantity?

png

Recipe Ingredient % diff @ 35 % diff @ 173
Pie Crust Fat 71.4 51.9
Pie Crust Flour 66.1 43.1
Pie Crust Water 76.3 74.4
Panna Cotta Gelatin 10.2 18.9
Panna Cotta Liquid -9.1 -6.3
Pastry Cream Cornstarch -48.6 -44.5
Pastry Cream Dairy -45.1 -44.8
Choux Egg -59.3 -75.7
Choux Flour -65.7 -75.7

It turns out that the LLM was wildly inaccurate in scaling recipes. It was fairly close for Panna Cotta, but for other recipes, the median LLM-scaled result was over 40% different from the result coming from multiplying the 6-person recipe. (Note that this could be a sign that the 6-person recipe is incorrect as well, but either way, it seems that there is a problem.)

I’m fairly surprised that this type of linear scaling was so inaccurate in this context. It seems like this kind of task would be really low-hanging fruit for Reinforcement Learning with Verifiable Rewards, and you would think that OpenAI would prioritize making the “cooking with ChatGPT” experience better given their usage guide advertising the functionality. But, I suppose that most people who set out to cook Choux Pastry for 173 people will probably triple-check the recipe before starting up their industrial mixer rather than using a one-shot recipe from ChatGPT.

In the future, I’ll probably continue to use an LLM for cooking (because I can taste and correct during the process), but for anything that seems like it could be messed up with bad ingredient ratios, I’ll lean on a cookbook or human-tested recipe from a website. And, the next time I am cooking for 173 people, I will definitely scale linearly from a recipe I know works instead of generating a new one from scratch. We will just hope that if I do undertake the “baking for 173” task, I do not get too over-enthusiastic and accidentally “scale the baking time” linearly as well, for the sake of the ears of anyone else who may be in range of the fire alarm.

Thanks for reading! And happy baking!