Explainer
Are ChatGPT recipes any good? How to check, fix, and cook them
The short answer
ChatGPT recipes are good for adapting a technique you already understand and bad at anything that needs a tested ratio or a safe temperature — run the checklist below before you cook one, and lean on a real recipe as the base whenever you can.
The honest answer is “sometimes,” and you can tell which before you cook
A recipe from ChatGPT was never cooked by anyone before it reached you. It’s a plausible-sounding arrangement of words, generated because that pattern of ingredients and steps tends to follow that kind of request in training data — not because someone made the dish and wrote down what worked. Most of the time that’s close enough. Sometimes it isn’t, and a five-minute check tells you which.
The model is good at the parts of cooking that are judgment and technique, and bad at the parts that are precision and safety. Knowing which bucket your request falls into is most of the work.
Where it’s actually good
- Riffing on a technique you already understand. A stir-fry with what’s in your fridge asks it to substitute inside a structure you can judge.
- Using up an odd ingredient. “What do I do with preserved lemons and half a bunch of celery” asks for ideas, not a tested procedure.
- Adapting a dish you know well. Made carbonara a dozen times? You’ll notice immediately if a suggested change is off.
- Scaling ideas, not exact ratios. “A lower-effort version of this weeknight dinner” is pattern-matching a simpler shape, which it’s good at.
- Explaining why a method works. Why you rest meat, why you salt pasta water — summarizing known knowledge, not generating something new.
- Suggesting substitutions to evaluate. Five stand-ins for buttermilk is a starting point you judge yourself, not a claim to trust blind.
- Turning a vague craving into a starting point. “Something warm, spicy, uses cabbage” is fuzzy input it handles better than a search box.
In every case, you supply the judgment and the model supplies options. That’s the shape that works.
Where it fails, specifically
| Failure mode | Why it happens |
|---|---|
| Baking ratios off by a meaningful margin | Narrow tolerances; a flour-to-liquid ratio off by 15% is the difference between a cake and a brick, and the model has no way to know it got the ratio wrong |
| Fermentation and preservation | A safety problem, not a taste problem — salt percentage in a ferment and acidity/processing time in canning are set by tested science |
| Missing or vague food-safety temperatures and times | A plausible-sounding number with no grounding in what actually kills pathogens in that dish |
| Regional dishes flattened into an internet average | A regional dish request often returns a blended composite of every recipe with that name online |
| Quantities that don’t add up to the stated yield | ”Serves 4” next to amounts that clearly feed 2 — nobody checked the arithmetic because nobody cooked it |
| Hallucinated cooking times | A time that sounds specific and confident with no basis in an actual test |
| Ingredient lists that don’t match the method | An ingredient appears in the method that was never listed, or vice versa — the two sections were generated as separate passes |
The common thread: the model produces a plausible-looking recipe, not a tested one. It has read an enormous number of recipes and learned what a recipe looks like, without the step every real recipe writer takes — actually cooking it and writing down what happened. Fluency is not the same as testing. That matches independent reporting: general-purpose testing has found chatbots invent ingredient quantities and miss food-safety considerations a tested recipe would flag (see TechRadar’s rundown of hallucination signs). That doesn’t make the output useless — it needs the skepticism you’d give a recipe from a stranger with no reviews.
The checklist: run this before you cook anything a chatbot generated
| Check | Look for | If it fails |
|---|---|---|
| Ingredients ↔ steps match | Every ingredient appears in a step and vice versa | Add the missing ingredient, or cut the step’s reference to it |
| Yield makes sense | ”Serves 4” roughly matches the quantities given | Recalculate from the quantities, not the stated count |
| Baking ratios are sane | Flour-to-liquid-to-fat-to-leavener against a recipe you trust in the same category | Use a tested recipe’s ratios as the baseline; keep the flavor idea |
| Temperatures and times are plausible and safe | Poultry 165°F, ground meats 160°F, whole cuts of pork/beef/lamb/veal 145°F rested 3 minutes (FSIS) | Replace with the verified number, don’t average the two |
| Seasoning is quantified, not hand-waved | ”Salt to taste” is fine; a precise ferment salt percentage is not trustworthy unverified | Treat as a stop — find a tested recipe from your state extension service or USDA |
| No step assumes an unlisted ingredient | A method referencing “the marinade” that was never listed | Add it to the ingredient list before you shop |
None of these fixes require throwing the recipe out. They require reading it once, critically, before you’re standing at the stove.
Why fermentation and canning are a different category of risk
Baking gone wrong ruins dinner. Preservation gone wrong can make someone sick. Tested canning and fermentation recipes specify exact ratios because safety depends on acidity (pH) and, for canning, heat processing time — both set by the ratios. Change the vinegar, sugar, or salt ratio and you can change whether the process prevents Clostridium botulinum, the bacteria behind botulism, and you won’t be able to tell by taste, smell, or appearance that anything went wrong (NDSU Extension: Safe Changes and Substitutions to Tested Canning Recipes). A generated recipe has no way to have verified any of that — treat it as a starting idea to check against a tested source, never the final word.
Prompting for better results
The single biggest improvement is giving the model a real recipe to work from instead of asking it to invent one. “Adapt this recipe [paste it] to make it dairy-free” gives it a tested structure to modify, instead of inventing both the structure and the substitution at once.
- Paste in a source recipe rather than describing the dish and asking for one from scratch.
- Ask for weights, not just volume, especially for baking — grams force more precision than “a cup of flour.”
- Ask it to state its assumptions about your oven, your flour, your pan size.
- Ask what could go wrong — “what’s the most likely way this fails” turns confidence into something useful.
The honest limit: a tested recipe is still the better starting point
A recipe from a cook who actually made the dish and had someone else test it is still the better starting point than one generated from nothing. The sensible use of a chatbot is adapting a recipe that already exists rather than generating one from a blank page. If you just need a stand-in for one ingredient, see how to substitute an ingredient — a narrower, safer ask the model handles well.
Our August 30 output-only benchmark made that distinction visible. GPT-5.4, Claude Sonnet 5, and Gemini 3.7 Flash all handled a tightly constrained chicken recipe, a 2.5× scale, and the major defects in an intentionally broken recipe. But when all three were asked for the same 12 blueberry muffins, their flour, liquid, fat, sugar, and leavener formulas differed materially—even though every answer looked polished and complete. We did not bake them, so that is not a taste verdict. It is evidence that fluent recipe prose does not tell you whether the formula was tested. See the full ChatGPT vs. Claude vs. Gemini recipe comparison and the focused ChatGPT baking formula check.
Try it: turn a chat recipe into something you can actually check
Say ChatGPT hands you a recipe as a wall of text — ingredients, then steps, no easy way to tell if they line up. Sous Chef has an MCP connector that works inside ChatGPT and Claude: instead of leaving the recipe as prose in the thread, it renders as a real Sous Chef card, and you can step through cooking it right there. Saving to your library is a separate, explicit tap — nothing saves just because it rendered.
That structure makes the checklist above practical instead of tedious. Once the recipe is a card, asking “swap the buttermilk for something I have” goes through the same validated editing Sous Chef uses everywhere — the server checks the request is structurally sound before applying it, and if the substitution changes a step’s wording, that step updates too, not just the ingredient line. Changed lines are marked Edited, and Undo restores the original.
Be precise about what that validation does and doesn’t do: it checks that an edit is structurally valid against the recipe — the ingredient exists, the step reference exists, the patch is well-formed — not that the dish will taste good, or that your substitution was a sound idea. It catches a step referencing an ingredient nobody added; it does not catch a bad ratio or an unsafe fermentation salt percentage. That’s still on you and the checklist above.
Common questions
Is it dangerous to cook a recipe ChatGPT generated?
Most of the time, no — a badly balanced stir-fry is a bad dinner, not a health risk. The real danger zone is undercooked poultry or pork, and canning or fermentation with the wrong acidity or salt level. Run those recipes past the checks above first.
Why does ChatGPT sometimes list an ingredient that’s never used in the steps?
The ingredient list and the method aren’t guaranteed to be generated in lockstep — the model produces plausible text for each section, not a consistency check between them. Reading both once before you shop catches this reliably.
Can I trust ChatGPT for baking recipes?
Trust it for ideas and flavor combinations, not the ratio of flour to liquid to fat to leavener. A ratio 15% off can be the difference between a good result and a failed one, and the model has no way to know if its ratio was ever tested.
What’s the safest way to use ChatGPT for cooking?
Give it a real recipe to adapt instead of asking it to invent one. “Make this dairy-free,” “scale this to serve 6,” and “explain why this step matters” all work from something tested, instead of inventing a structure from nothing.
Should I ask ChatGPT for canning or fermentation recipes at all?
Use it to understand a technique or brainstorm flavors, but get the actual salt percentage, acid level, and processing time from a tested source — your state’s extension service or USDA — before you can or ferment anything. It’s the one category where being wrong isn’t just a bad dinner.
Does Sous Chef generate recipes from scratch like ChatGPT does?
No — Sous Chef imports recipes you already have and edits them with validated, structured changes. The ChatGPT/Claude connector renders a chat-generated recipe as a card you can cook and edit the same way, but Sous Chef isn’t a from-scratch generator.
Try it on your own recipe
Sous Chef is free to download and try on iPhone. Heavier AI use may need a subscription later.
Download on the App Store