Explainer

ChatGPT vs. Claude vs. Gemini for recipes: five real tests

The short answer

In five output-only tests, all three models handled a constrained dinner, recipe scaling, and a deliberately bad recipe reasonably well; the clearest differences appeared in baking formulas and completeness, but five unbaked outputs cannot support a universal winner.

The result: capable drafts, visible misses, no honest winner

We sent five identical recipe prompts through OpenRouter to OpenAI GPT-5.4, Anthropic Claude Sonnet 5, and Google Gemini 3.7 Flash on August 30, 2026. We inspected their outputs against the prompt constraints and trusted references.

All three produced a plausible constrained chicken dinner with the correct 165°F poultry endpoint. All three scaled a stovetop recipe by 2.5 and recognized that pot size and cooking time do not scale linearly. All three found the major safety and structural faults planted in a deliberately bad recipe.

The meaningful differences were narrower. GPT-5.4 contradicted itself about the chickpea amount in its scaled recipe. Claude Sonnet 5 omitted the actual oven temperature and time from a dairy-free adaptation despite being asked for a complete revised method. The three fresh blueberry-muffin formulas differed substantially from one another and from the tested comparison recipe.

That is useful evidence, but not a leaderboard. We did not cook any output. Five prompts are too few to crown a winner, and these results apply only to the named model versions, prompts, date, and OpenRouter API surfaces. Consumer ChatGPT, Claude, and Gemini apps can use different defaults, tools, system instructions, and model versions.

What we tested

TestWhat a good answer needed to doResult
Constrained weeknight dinnerUse only allowed ingredients, serve four, give a measurable chicken endpointAll three complied and used 165°F / about 74°C
Scale 4 servings to 10Apply 2.5x, revise the full method, flag equipment and timingAll three handled the main task; GPT contradicted one chickpea note
Write 12 blueberry muffinsGive grams and volume, pan fill, method, time, donenessAll three looked usable on paper, but formulas varied materially
Make supplied muffins dairy-freeReplace milk and butter, keep eggs, return a complete methodAll made sensible swaps; Claude left temperature and time unspecified
Audit a deliberately bad chicken recipeFind cross-contamination, 140°F endpoint, missing ingredients, yield and liquid problemsAll three caught the major planted defects

The benchmark measures output quality on these tasks—not taste, consistency across repeated runs, or whether a recipe works in a real kitchen.

Test 1: a constrained chicken dinner

We allowed chicken thighs, chickpeas, canned tomatoes, onion, garlic, olive oil, lemon, cumin, rice, water, salt, and pepper—nothing else. Each model had to give exact quantities, yield, total time, numbered steps, and a measurable food-safety endpoint.

All three produced a one-pan chicken-and-rice format, used the permitted ingredients, served four, and required chicken to reach 165°F / approximately 74°C. That matches the USDA safe minimum temperature for poultry.

The important success was constraint-following and a verifiable endpoint. Even here, a correct temperature does not prove the proposed liquid ratio, seasoning, or texture was good; none of the dishes was cooked.

Test 2: scaling from four servings to ten

The correct multiplier was 2.5. Each answer needed 750g rice, 1.5L water, 2.5 tablespoons oil, 5 garlic cloves, 1kg each of chickpeas and tomatoes, and proportional cumin, salt, pepper, and lemon—plus a larger vessel and non-linear timing guidance.

All three stated or applied 2.5x, recommended replacing the original 3-quart saucepan with a larger pot, and avoided multiplying the 18-minute simmer by 2.5. They also sensibly treated salt and lemon as final adjustments.

GPT-5.4 contained the most concrete arithmetic presentation problem: its ingredient line said 1kg drained chickpeas, then a parenthetical described roughly 2.5 standard 400g cans before draining and told the reader to aim for about 600g drained. Both cannot be the final target. A cook scanning the bold line and one reading the note could make different batches.

That is exactly why generated scaling still needs a line-by-line check. See can ChatGPT scale a recipe correctly? for the full audit.

Test 3: three blueberry-muffin formulas

We asked for exactly 12 standard muffins with grams and US volume, oven temperature, pan-fill guidance, mixing method, a visual doneness cue, and a time range. We compared the output on paper with King Arthur Baking’s tested Mix & Match Muffins: 240g flour, 100g sugar, 14g baking powder, 227g milk, 50g oil or 57g butter, and 2 eggs before the chosen add-ins.

FormulaFlourMain liquidButter/fatSugarBake
Tested comparison240g227g milk50g oil or 57g butter100gSource recipe
GPT-5.4240g180g buttermilk113g butter150g400°F, 18–22 min
Claude Sonnet 5280g240g yogurt or buttermilk115g butter150g400°F, 18–22 min
Gemini 3.7 Flash240g120g milk75g butter150g375°F, 20–25 min

All three instructed the reader not to overmix, specified pan fill, and gave plausible doneness cues. But they did not converge on one formula: Claude used 40g more flour than the other models; Gemini used half Claude’s dairy amount; GPT and Claude used about twice the butter in the comparison formula.

Those differences do not prove any formula fails. Muffin styles vary, and only baking could establish the results. They do prove that fluent completeness is not the same as a tested baseline. Read can ChatGPT write baking recipes? before treating one as ready.

Test 4: a bounded dairy-free adaptation

This test was safer by design: we supplied a complete muffin formula and asked the models to replace dairy while retaining eggs.

All three swapped whole milk for unsweetened plant milk, replaced melted butter with oil or plant butter, retained the eggs, and explained likely texture or browning differences. GPT-5.4 and Gemini 3.7 Flash returned complete methods with 400°F / 200°C and 18–22-minute guidance.

Claude Sonnet 5 wrote “bake as originally directed,” but the supplied baseline said only “bake”—it included no original temperature or time. The prompt explicitly requested a complete revised method, so the answer was incomplete. The recipe might still be repairable, but the reader would have to fetch missing instructions elsewhere.

The lesson is broader than Claude: ask for a complete final recipe, then check that no instruction points to context that does not exist.

Test 5: detecting a bad recipe

The bad recipe intentionally reused a raw-chicken board for onion, allowed chicken at 140°F or “no longer pink,” mentioned spinach and parmesan absent from the ingredients, forgot to use paprika, conflicted between serving six and four, and proposed too little liquid and time for the rice-and-chicken method.

All three found the major food-safety, structural, and culinary problems and produced corrected recipes. Each rejected 140°F and appearance as a chicken endpoint, called for a thermometer and 165°F, fixed the cross-contamination workflow, reconciled missing ingredients, and questioned the rice liquid and timing.

That is a strong use case: critique an existing recipe against an explicit checklist. It is still not proof that each newly “corrected” recipe tastes good. USDA notes that a food thermometer is the reliable way to confirm a safe minimum temperature; color alone is not (USDA food thermometer guidance).

Which should you use?

This small benchmark does not justify “ChatGPT wins,” “Claude wins,” or “Gemini wins.” Each succeeded on most explicit checks, and each model could produce a different response on another run.

Whichever you use, improve the workflow:

  1. Start with a trusted recipe when possible.
  2. State constraints, yield, format, and safety endpoints explicitly.
  3. Ask for a complete ingredient list and method—not a patch scattered through chat.
  4. Recalculate quantities and check ingredients against steps.
  5. Verify safety facts at their primary source.
  6. Treat baking output as an untested formula until someone bakes it.

Move the approved recipe into Sous Chef

ChatGPT, Claude, and Gemini are good drafting surfaces. A chat thread is a bad long-term recipe card: the final version gets buried among revisions, and later prompts can generate a different version.

After checking the recipe, copy the final text into Sous Chef on iPhone. Paste-text import turns it into structured ingredients and ordered steps and stores it in your library. You can then scale it, substitute an ingredient, convert units, or make a dietary adaptation inside the recipe. Changed lines stay visible and the original remains available, rather than another revision being lost in the chat.

Sous Chef’s structural validation is not taste or safety validation. It checks whether edits are well-formed and refer to real recipe parts; it cannot certify flavor, allergy safety, nutrition, or an unbaked formula. ChatGPT and Claude users may also use the optional Sous Chef MCP connector, but paste-text import works for output from any model.

Common questions

Was this a cooking test?

No. It was an output-only OpenRouter API benchmark. No generated recipe was cooked or tasted.

Which model was best for recipes?

Five prompts are insufficient to name a winner. All three handled most explicit checks, while individual misses appeared in scaling consistency, adaptation completeness, and divergent baking formulas.

Are these results the same as using the consumer apps?

Not necessarily. Consumer apps may use different models, tools, defaults, memories, and system instructions. These findings apply to the named API model versions on August 30, 2026.

Did all three give the correct chicken temperature?

Yes. All used 165°F / about 74°C in the constrained dinner and rejected 140°F in the bad-recipe audit, consistent with USDA guidance.

Can I trust the muffin recipes?

Treat them as untested drafts. All looked plausible, but their flour, liquid, fat, temperature, and timing choices differed materially. Only baking would establish the result.

Why save the result in Sous Chef?

It turns the final checked output into a stored, structured recipe you can keep editing on iPhone, instead of leaving competing drafts in a chat history.

Try it on your own recipe

Sous Chef is free to download and try on iPhone. Heavier AI use may need a subscription later.

Download on the App Store