Explainer
Can AI detect a bad recipe? What ChatGPT can catch—and miss
The short answer
AI can flag structural problems in a recipe—missing ingredients, impossible step references, inconsistent quantities, vague timing, and suspicious temperatures—but it cannot prove the dish works, so compare critical ratios and safety instructions with trusted sources before cooking.
AI is a useful proofreader, not a test kitchen
Paste a recipe into ChatGPT and it can notice inconsistencies: an onion in the method but not the ingredient list, no oven temperature, or a sauce reduced before any liquid is added. Recipe prose has enough structure for a model to compare sections quickly.
What it cannot do is establish that the recipe has been cooked, that the ratios produce the promised texture, that your oven behaves the same way, or that the food is safe merely because the instructions sound confident. A clean bill of health from a chatbot is not evidence of testing.
This guide combines published safety sources with a five-prompt, text-only output benchmark run through the OpenRouter API on August 30, 2026. Nothing was cooked. Findings apply only to GPT-5.4, Claude Sonnet 5, and Gemini 3.7 Flash with those exact prompts; consumer apps may differ. Five prompts are too small a sample to name an overall winner.
The problems AI is best placed to catch
| Check | What AI can compare | What a useful flag sounds like |
|---|---|---|
| Ingredient-to-step match | Names in the ingredient list against names in the method | “The steps mention lemon juice, but none is listed.” |
| Missing setup | The method against temperatures, equipment, and prep details | “The recipe says bake but never gives an oven temperature.” |
| Sequence | Whether a later instruction depends on an action that never happened | “It asks you to add the reserved marinade, but no marinade was reserved.” |
| Yield plausibility | Ingredient quantities against stated servings or pan size | “This may be too little batter for the stated 9 × 13-inch pan.” |
| Internal timing | Individual times against total time | “The step times already total 55 minutes, not the claimed 30.” |
| Ambiguity | Vague terms without a sensory or measurable cue | “Cook until done” needs a texture, temperature, or visual cue. |
| Unit consistency | Mixed units, impossible conversions, or decimals that look accidental | “One section says 180°C and another says 180°F.” |
| Changed recipe consistency | Whether an ingredient edit propagated into affected steps | “The dairy-free version still says to brush with butter.” |
These are proofreading tasks: the evidence is inside the recipe. Ask the model to cite the exact ingredient or step for every issue.
What all three models caught in our deliberately bad recipe
One benchmark prompt supplied an intentionally defective creamy chicken-and-rice recipe. All three models flagged the raw-chicken cutting-board cross-contamination, rejected 140°F and visual “no longer pink” as the poultry endpoint, and required 165°F/74°C. They also found spinach and parmesan used in the steps but missing from the ingredients, paprika listed but unused, the six-versus-four serving conflict, too little cooking liquid, and an implausible 12-minute schedule.
That is encouraging for obvious, text-visible defects. It is not proof that a model will find subtle problems or that each corrected recipe works: the bad clues were deliberately planted, model versions change, and none of the corrections were cooked. The test supports using AI as a first-pass auditor, followed by the authoritative and physical checks below.
The problems AI cannot settle from text alone
Whether the recipe tastes good
A ratio may look familiar and still be unbalanced. Flavor and browning are physical outcomes. AI can call a quantity unusual, but it cannot honestly say the result is delicious.
Whether the texture will work
Baking depends on ingredient standards, mixing, pan size, and temperature. King Arthur Baking notes that substitutions or omissions can change results (recipe success guide). AI can flag a ratio; a tested comparison recipe makes the suspicion meaningful.
Whether a food is safe by appearance or wording
USDA says harmful bacteria cannot be reliably seen, smelled, or tasted and publishes minimum internal temperatures by food type (USDA safe-temperature chart). A sentence such as “cook until the juices run clear” may sound traditional, but USDA recommends a food thermometer because color and firmness are not reliable safety indicators (food-thermometer guidance).
Whether an adaptation is allergen-safe
AI can find dairy left in a dairy-free rewrite. It cannot inspect labels, shared utensils, warnings, or surfaces. The FDA treats allergen cross-contact prevention as a separate food-safety task (FDA cross-contact research). Text consistency is not certification.
Whether an untested preservation recipe is safe
Home canning depends on acidity, density, jar size, and validated processing. Changing an approved recipe can require a new process time (NCHFP heat-processing guidance). A model cannot validate a canning process from prose.
A better prompt for auditing a recipe
Paste the full recipe and ask:
Audit this recipe as an editor, not as someone who cooked it. Separate your response into: (1) internal contradictions you can prove from the text, (2) quantities or methods that look unusual and need comparison with a tested recipe, (3) food-safety points to verify against an authoritative source, and (4) subjective choices. Quote the relevant ingredient or step for each flag. Do not claim the recipe is tested, safe, or good.
Add the source, servings, pan dimensions, prior substitutions, altitude if relevant, and dietary constraints. “Is this recipe good?” invites confidence; categories of evidence expose uncertainty.
The seven-minute human review
Minute 1: reconcile the ingredient list and method
Cross off each ingredient as it appears. Look for items used but not listed, listed but never used, and prep states that disagree.
Minute 2: total the times
Add active, resting, chilling, proofing, and cooking times. A “30-minute” recipe may hide a two-hour chill.
Minute 3: check yield and equipment
Do quantities fit the servings and vessel? Loaves need pan dimensions; soup needs pot capacity.
Minute 4: inspect the ratio-sensitive ingredients
For baking, compare the formula and pan with a tested recipe of the same style. Check sauce thickeners and grain liquids too.
Minute 5: verify safety-critical instructions
Check vague doneness language against USDA minimum temperatures. Also verify cooling and storage: USDA says leftovers generally belong in the refrigerator within two hours and keep for three to four days at 40°F/4°C or below (leftovers guidance).
Minute 6: inspect every adaptation
If it was scaled or substituted, did dependent instructions move too? Doubling does not necessarily double cooking time; replacing chicken with tofu invalidates “cook until no longer pink.” See scaling and substitution checks.
Minute 7: decide what remains unknown
Write down unresolved texture, seasoning, oven timing, and substitution claims. That is where editing ends and testing begins.
How to grade an AI critique without pretending it is a benchmark
Reward traceability, not length:
- Strong: identifies conflicting lines, the consequence, and a bounded correction.
- Useful caution: names an unusual ratio and a trusted comparison to find.
- Weak: says something “seems wrong” without evidence.
- Dangerous: improvises canning, treats color as safe doneness, or guarantees allergen safety.
Success means finding the next check, not removing every uncertainty.
Save the corrected recipe as a recipe—not another block of chat
After ChatGPT helps inspect the draft, paste it into Sous Chef on iPhone. It becomes a structured recipe you can scale, convert, substitute, and adapt rather than another disconnected wall of text.
Sous Chef can update connected ingredients and steps, mark changes, preserve the original, and undo. The ChatGPT and Claude connector can also render a recipe card and lets you explicitly save it. The chatbot is the drafting space; Sous Chef becomes the durable version.
Sous Chef’s validation does not prove culinary quality or safety. It checks structural rules—such as whether an edit targets an ingredient or step that exists. It does not cook the dish, inspect allergens, calculate a safe canning process, diagnose nutrition needs, or replace USDA guidance. The point is a clean and reversible recipe workflow, not a false safety seal.
For the broader question, read are ChatGPT recipes any good?. If the main concern is a swap, use the ChatGPT ingredient-substitution guide.
Common questions
Can ChatGPT tell whether a recipe will work?
It can find inconsistencies and compare patterns. It cannot prove the recipe works without physical testing. Treat “this ratio is unusual” as a prompt to compare a trusted recipe.
What recipe errors is AI best at finding?
Missing ingredients, unused ingredients, contradictory units, omitted temperatures, impossible sequence references, mismatched totals, vague instructions, and incomplete propagation after a substitution or scaling change.
Can AI verify food safety in a recipe?
It can suggest checks, but verify temperatures, storage, allergens, and preservation with authoritative sources. Use a thermometer rather than visual doneness cues.
Can ChatGPT check a baking recipe?
It can identify suspicious ratios or a missing technique, especially when you provide a tested comparison recipe. It cannot know the finished crumb, rise, browning, or flavor without a bake test.
Can AI verify a canning recipe?
No. Use a research-tested recipe from USDA, the National Center for Home Food Preservation, or a university extension source and follow its specified jar size, proportions, and processing method.
Why put the corrected recipe in Sous Chef?
Because a structured recipe is easier to store, cook, find, and revise than prose buried in a chat. Sous Chef keeps the original, shows edits, and lets future scaling, measurement, substitution, and dietary changes update the recipe itself.
Try it on your own recipe
Sous Chef is free to download and try on iPhone. Heavier AI use may need a subscription later.
Download on the App Store