How-to

Can ChatGPT scale a recipe correctly? We tested GPT-5.4

The short answer

GPT-5.4 handled the main scaling task well in our output-only test, but one conflicting chickpea note showed why you should still verify every quantity, equipment change, and cooking-time assumption before using a scaled recipe.

The answer from one real test: mostly yes, with a consequential inconsistency

We gave OpenAI’s GPT-5.4 a four-serving rice, chickpea, and tomato recipe and asked it to return a complete ten-serving version. The correct multiplier was 2.5.

GPT-5.4 correctly scaled the rice to 750g, water to 1.5L, oil to 2.5 tablespoons, onions to 2.5, garlic to 5 cloves, tomatoes to 1kg, cumin to 2.5 teaspoons, pepper to 1.25 teaspoons, and lemon to a practical 2–3. It recommended a 5–6-quart pot instead of the original 3-quart saucepan, extended the onion-softening step, and treated salt, acidity, and simmering time as judgment calls instead of blindly multiplying time by 2.5.

Then it contradicted itself. The bold ingredient line called for 1kg canned chickpeas, drained. Its parenthetical described roughly 2.5 standard 400g cans before draining and told the reader to aim for about 600g drained chickpeas. Those are materially different quantities.

So: ChatGPT can do the core scaling work, but a polished output can still contain two incompatible answers on the same line.

This was an OpenRouter API output test on August 30, 2026. We did not cook the recipe. The result applies to this prompt and GPT-5.4, not every ChatGPT model or consumer-app response.

The exact task and what passed

The source recipe served four and used 300g rice, 600ml water, 1 tablespoon oil, 1 onion, 2 garlic cloves, 400g drained chickpeas, 400g canned tomatoes, cumin, salt, pepper, and lemon. It cooked in a 3-quart saucepan with an 18-minute covered simmer and 10-minute rest.

CheckExpected at 2.5xGPT-5.4 result
Rice750gCorrect
Water1.5LCorrect
Oil2.5 tbspCorrect
Onion2.5Correct, with a practical whole-onion option
Garlic5 clovesCorrect
Chickpeas1kg drainedCorrect in main line; conflicting 600g drained note
Tomatoes1kgCorrect
Cumin2.5 tspCorrect
PanLarger than 3 quartsRecommended 5–6 quarts
SimmerNot 45 minutesSuggested 20–22 minutes and a doneness check

The same test was run with Claude Sonnet 5 and Gemini 3.7 Flash. Both also applied 2.5x, recommended larger vessels, and avoided multiplying the cooking time by 2.5. That makes the broad result encouraging, not conclusive: three models handled one explicit scaling problem, but only GPT’s output contained that particular internal contradiction.

Why “multiply every number” is not enough

The formula is straightforward:

servings wanted ÷ servings listed = multiplier

Going from 4 to 10 gives 10 ÷ 4 = 2.5. ChatGPT should state that multiplier so you can audit it. Arithmetic is only the first layer.

Equipment does not scale linearly

A 3-quart pot does not become suitable because the ingredient math is correct. More food needs headroom for stirring, simmering, and even heat distribution. All three tested models recognized this and recommended pots between roughly 5 and 8 quarts.

For baking, geometry matters even more. King Arthur Baking notes that a 9-inch round pan has 26.6% more volume than an 8-inch round, despite the small difference in diameter, and recommends keeping most cake pans between half and two-thirds full (pan-size guide).

Cooking time follows depth and method, not servings

Multiplying the original 18-minute simmer by 2.5 would produce 45 minutes—likely a poor instruction for this rice dish. GPT-5.4 suggested 20–22 minutes after the larger batch reached a simmer, while Claude suggested 22–25 and Gemini 18–20. All recognized that bringing the larger mass to temperature would take longer.

That spread is a reminder: an exact new time is still an estimate. Vessel width, burner, rice variety, lid, and starting temperature matter. A complete scaled recipe should give a time range and a doneness cue.

Cookies baked at the same size may keep roughly the same per-tray time. A deeper casserole may need longer. Double vegetables crowded onto one sheet pan may steam instead of roast, so the right change is a second pan or another batch—not extra minutes.

Seasoning can be proportional and still need tasting

GPT-5.4 scaled salt to 2.5 teaspoons but told the reader to adjust at the end; it suggested starting with two lemons and adding more if needed. Claude and Gemini were more conservative with the initial salt.

That is sensible for a forgiving stovetop dish because canned tomatoes and chickpeas vary in sodium. It is not a universal rule. Salt in bread dough, brine, curing, fermentation, or preservation may be structural or safety-critical and should follow a tested formula.

Safety temperatures do not multiply

Scaling a chicken recipe changes how long a batch may take, not the safe endpoint. USDA lists 165°F / 73.9°C for all poultry and 160°F / 71.1°C for ground meat (safe minimum temperature chart). Verify with a thermometer, not color or a multiplied time.

The prompt that produced a useful answer

The benchmark did not merely say “make this serve ten.” It required a complete revised ingredient list and numbered method, preservation of flavor balance, identification of seasonings to adjust, and warnings about equipment or timing that do not scale linearly.

Use the same structure:

Scale this recipe from [original yield] to [new yield]. State the multiplier. Return a complete revised ingredient list and numbered method, not only changed quantities. Preserve the flavor balance, identify seasonings to adjust to taste, and flag pan, batch, depth, reduction, or timing changes that do not scale linearly. Check that every number in the ingredient list matches the method.

Add the complete source recipe beneath it. The final consistency instruction is worth emphasizing because it targets the exact failure we saw.

Audit the result in two minutes

  1. Recalculate the multiplier yourself.
  2. Verify at least the smallest fraction, the largest quantity, and any awkward whole ingredient.
  3. Read parentheticals as carefully as bold ingredient lines; GPT’s contradiction lived there.
  4. Confirm units describe the same thing—“400g canned” and “400g drained” are different baselines.
  5. Check pan capacity, food depth, and whether the batch should split.
  6. Reject cooking time multiplied by the serving factor.
  7. Require a texture, temperature, or visual doneness cue.
  8. Match every quantity mentioned in the method to the ingredient list.

For baking, convert critical ingredients to grams. King Arthur recommends weight measurements and keeping the specified technique and pan size in its recipe success guide. For a partial egg, beat it until uniform and weigh the needed fraction; King Arthur’s guide to reducing recipes uses the same approach.

Do not ask a chatbot to improvise canning proportions. The National Center for Home Food Preservation says to use tested proportions and not alter vinegar, food, or water ratios (USDA canning guidance).

Store the checked version in Sous Chef

The GPT-5.4 output was useful, but the correct version should not remain buried beside the conflicting one in a chat. After resolving the chickpea quantity, copy the complete recipe into Sous Chef on iPhone. Paste-text import stores it as structured ingredients and ordered steps.

The next time you need six servings instead of ten, make that change inside Sous Chef. It can update quantities and dependent instructions, show the edited lines, and keep the original available for undo. The same saved recipe can later handle a substitution, unit conversion, or dietary adaptation without regenerating the whole dish from scratch.

Sous Chef’s validation is structural, not culinary. It can ensure an edit refers to real ingredients and steps; it cannot prove a time estimate, taste, nutrition, or food safety. Your arithmetic check and thermometer still matter. For the general method, see how to scale a recipe up or down.

Common questions

Did ChatGPT get the 2.5x arithmetic right?

Mostly. GPT-5.4 correctly scaled the main quantities, but its chickpea parenthetical conflicted with its own bold 1kg drained amount.

Should cooking time double when a recipe doubles?

Usually not. Individual item size, food depth, vessel width, crowding, and heat source matter more than serving count. Use a time range plus doneness cues.

Can ChatGPT halve a baking recipe?

It can calculate halves, but verify grams, egg weight, leavener, pan area, batter depth, and bake time against a tested recipe.

Can ChatGPT convert cups to grams while scaling?

It can attempt both, but that combines two error sources. Ingredient-specific density matters, so use a named weight reference and verify each conversion.

Was this scaled recipe cooked?

No. This was output analysis against arithmetic and trusted guidance. No benchmark recipe was kitchen-tested.

Why move the result to Sous Chef?

It gives the corrected final recipe one structured home on iPhone and makes later changes visible, instead of leaving contradictory drafts in a chat.

Try it on your own recipe

Sous Chef is free to download and try on iPhone. Heavier AI use may need a subscription later.

Download on the App Store