Explainer
Can ChatGPT write baking recipes? Three AI muffin formulas compared
The short answer
AI can produce a plausible baking recipe, but our three model outputs differed materially in flour, liquid, fat, temperature, and time; without kitchen testing, treat each as a draft to compare with a trusted formula, not a proven recipe.
The test result: all plausible on paper, none proven in a kitchen
We asked GPT-5.4, Claude Sonnet 5, and Gemini 3.7 Flash for exactly 12 standard blueberry muffins with gram and US volume measures, oven temperature, pan-fill guidance, mixing method, doneness cues, and a time range.
All three delivered complete-looking recipes. All warned against overmixing, filled the wells roughly three-quarters full, used two eggs, and gave reasonable-looking bake instructions.
But the formulas differed materially:
| Formula | Flour | Main liquid | Butter/fat | Sugar | Oven and time |
|---|---|---|---|---|---|
| King Arthur tested comparison | 240g | 227g milk | 50g oil or 57g butter | 100g | See source recipe |
| GPT-5.4 | 240g | 180g buttermilk | 113g butter | 150g | 400°F, 18–22 min |
| Claude Sonnet 5 | 280g | 240g yogurt or buttermilk | 115g butter | 150g | 400°F, 18–22 min |
| Gemini 3.7 Flash | 240g | 120g milk | 75g butter | 150g | 375°F, 20–25 min |
The trusted comparison was King Arthur Baking’s Mix & Match Muffins, which yields 12 using 240g flour, 100g sugar, 14g baking powder, 227g milk, 50g oil or 57g butter, and two eggs before the chosen add-ins.
This does not show that the AI versions fail. Richer, sweeter, thicker, and leaner muffin styles can all exist. It shows that “looks like a recipe” is not evidence that one formula was tested. We did not bake or taste any output.
The benchmark ran through OpenRouter on August 30, 2026. Findings apply to these named API model versions and prompts. Consumer ChatGPT, Claude, or Gemini apps may surface different models and instructions.
What each model changed
GPT-5.4: same flour, less liquid, twice the comparison butter
GPT used 240g flour—the same as the tested comparison—but 180g buttermilk and 113g butter. It also added 150g sugar, baking soda alongside baking powder, and baked at 400°F for 18–22 minutes.
The use of buttermilk makes the baking-soda addition internally understandable because it supplies acid. The draft specified the muffin method, told the reader not to beat the batter smooth, and provided spring-back and toothpick cues. On paper, it was coherent. Only baking could establish whether its richer fat level and lower liquid deliver the promised result.
Claude Sonnet 5: more flour and the most dairy
Claude used 280g flour, 240g yogurt or buttermilk, 115g butter, 150g sugar, baking powder and baking soda, then baked at 400°F for 18–22 minutes. It gave unusually specific folding guidance and a target portion of roughly 60–65g per cup.
The flour increase may balance the larger dairy amount and create a thicker batter, but that is an inference from the formula, not a tasting result. The 40g flour difference from GPT and Gemini is large enough that the three recipes are not interchangeable drafts of the same formula.
Gemini 3.7 Flash: the leanest liquid-and-fat formula
Gemini used 240g flour, 120g milk, 75g butter, 150g sugar, and baking powder without baking soda. It chose 375°F for 20–25 minutes. Its milk amount was half Claude’s and two-thirds GPT’s buttermilk amount.
Again, that does not make it wrong. Eggs and melted butter also bring water, and different target textures exist. It means a reader cannot infer reliability merely because the output includes grams and a confident time range.
The second test: adapting a trusted formula worked better
We then supplied the tested-style base formula—240g flour, 100g sugar, 14g baking powder, 227g whole milk, 57g melted butter, two eggs, vanilla, and blueberries—and asked each model to make it dairy-free but not vegan.
All three retained the eggs, replaced milk with unsweetened plant milk, replaced butter with neutral oil or plant butter, preserved the core quantities, and explained likely effects on texture, richness, or browning. GPT-5.4 and Gemini returned complete methods at 400°F / 200°C for 18–22 minutes.
Claude made sensible ingredient swaps but ended with “bake as originally directed.” The supplied method contained no original oven temperature or time. Because the prompt asked for a complete revised method, the answer was incomplete even though its substitution logic was plausible.
This is the stronger AI baking workflow: provide a trusted formula, request one bounded change, and check the complete final recipe. Fewer variables are invented at once, and omissions are easier to spot.
For the broader risk checklist behind that workflow, see are ChatGPT recipes any good?. If the request is specifically about serving count rather than ingredients, use the separate ChatGPT recipe-scaling audit.
The six checks every generated baking recipe needs
| Check | What to inspect | What failed output can look like |
|---|---|---|
| Weights | Grams for flour, liquid, fat, sugar, leavener | Cups converted with no ingredient-specific reference |
| Ingredient function | Structure, hydration, tenderness, lift | Egg or dairy removed with no functional replacement |
| Comparable formula | Same type of muffin, cake, cookie, or bread | One generic “ideal ratio” applied to every style |
| Method | Ingredient state agrees with technique | Melted butter paired with “cream until fluffy” |
| Pan | Capacity and batter depth fit the yield | “8- or 9-inch pan” treated as identical |
| Doneness | Time plus a suitable visual or tactile cue | One exact time with no cue or uncertainty |
1. Verify weights, not only volumes
King Arthur’s ingredient weight chart lists its all-purpose flour at 120g per cup and recommends a digital scale. A cup of flour, sugar, milk, and butter cannot share one conversion factor.
All three benchmark outputs gave broadly consistent gram/volume pairs, which made comparison possible. That still did not make their formulas converge.
2. Ask what each changed ingredient does
Milk hydrates; fat tenderizes and carries flavor; eggs can bind, emulsify, aerate, enrich, and add water; leavener must fit the available acid and method. A dairy-free swap is not complete merely because the word “milk” disappeared.
The second benchmark succeeded here: all three replaced both milk and butter while correctly retaining eggs. For allergy-related baking, independently check product labels and cross-contact. FDA advises people with food allergies to read labels and identify the specific allergen source (FDA). Neither a chatbot nor a recipe app can inspect the package in your kitchen.
3. Compare like with like
Butter cakes, oil cakes, sponge cakes, quick breads, and yeast breads create structure differently. Compare a generated muffin with trusted muffins using the same method—not with a universal ratio.
For bread, baker’s percentages are helpful: total flour is 100% and other ingredients are measured relative to it. King Arthur describes baker’s math as a scaling and formula tool.
4. Match the method to the formula
If butter is melted, the recipe should not tell you to cream it. If buttermilk and baking soda appear together, the method should combine and bake without an unexplained long delay. Every listed ingredient should appear in a step, and the step should use the same state—cold, softened, melted, whisked—that the ingredient line specifies.
5. Check pan capacity
King Arthur notes that a 9-inch round pan has 26.6% more volume than an 8-inch round and advises filling many cake pans only half to two-thirds full (why pan size matters). “Use an 8- or 9-inch pan” can change depth, rise, and timing meaningfully.
6. Require a doneness cue
The benchmark did well here: all three muffin outputs paired a time range with doming, browning, spring-back, and toothpick cues. Choose cues that fit the bake. “Golden brown” is weak for chocolate cake; a perfectly clean toothpick may be too far for a fudgy brownie.
A better prompt for ChatGPT baking
Using the trusted recipe below as the baseline, make this one bounded change: [change]. Preserve grams for all major ingredients and one specific pan size. Explain what each replacement changes. Return a complete ingredient list and numbered method with oven temperature, time range, and observable doneness cues. Check that every ingredient appears in the method. Label anything that still requires kitchen testing.
If you ask for a new formula from scratch, add a requirement to compare it against two named tested recipes of the same style—but remember that desk comparison still cannot establish taste or texture.
Store the reviewed recipe in Sous Chef
Once you have selected and checked a version, move it out of the chat. Copy the final text into Sous Chef on iPhone. Paste-text import turns it into structured ingredients and steps and stores it in your recipe library.
From there, you can halve it, switch units, substitute an ingredient, or make a dietary adaptation inside the saved recipe. Sous Chef marks changed lines and keeps the original available, so you are not juggling three similar chatbot drafts.
Sous Chef’s validation is structural, not culinary. It can reject a malformed edit or a reference to a missing ingredient. It cannot certify crumb, flavor, nutrition, allergy safety, or preservation safety. Only baking establishes whether an untested formula works. The optional Sous Chef MCP connector can support ChatGPT and Claude workflows; paste-text import works with any model.
Common questions
Did any model write the best blueberry muffin?
We cannot say. The three formulas differed substantially and none was baked. Five benchmark prompts do not support a winner.
Were the AI muffin formulas obviously broken?
No. All were plausible on paper and internally structured, but plausibility is weaker evidence than a tested recipe.
Why did the formulas differ so much?
The models selected different muffin styles and balances without a supplied baseline. GPT and Claude chose richer butter-and-cultured-dairy formulas; Gemini chose less milk and fat. Only kitchen testing could compare outcomes.
Did the dairy-free adaptation work better?
It was more controlled. All three preserved the base ratios and made sensible swaps, though Claude omitted the actual bake temperature and time from its supposedly complete method.
Can ChatGPT convert cups to grams for baking?
It can, but verify each ingredient against a named weight reference. Density differs by ingredient and measuring convention.
Does saving an AI recipe in Sous Chef make it tested?
No. It makes the recipe structured, stored, and easier to edit transparently. It does not replace trusted sources or real baking.
Try it on your own recipe
Sous Chef is free to download and try on iPhone. Heavier AI use may need a subscription later.
Download on the App Store