Guides

Is AI Calorie Counting Accurate? What It Gets Right and Wrong

AI calorie counters can identify food, but photo-only estimates often miss portions, fats, and hidden ingredients. See what current studies found.

Chris Raroque

Chris Raroque

Gouache illustration of a meal, phone nutrition estimate, and hand making corrections with a pencil

Point your phone at dinner and a few seconds later an app may return something that looks remarkably complete: grilled chicken, rice, broccoli, 640 calories.

The food names can be right. The total can still be wrong.

That is the confusing thing about AI calorie counters. They are often convincing at the exact moment they are most uncertain. A photograph can show chicken; it cannot show how much oil went into the pan. It can recognize a grain bowl; it may not know whether the bowl holds one cup of rice or two. And it has no way to count the handful of nuts that was never photographed.

So the useful question is not whether AI calorie counting is simply accurate or inaccurate. It is whether the estimate is good enough for the decision you are making.

I make Amy Food Journal. The studies discussed below are independent; when I mention Amy’s own testing near the end, I label it clearly.

The Short Answer

AI calorie counters are useful for quick estimates, but a photo-only result should not be treated like a measurement.

Current studies do not support one universal accuracy percentage. Performance changes with the model, the meal, the portion, and the information supplied with the image. In one controlled study, ChatGPT and Claude missed calories by about 36% on average in percentage terms. In a small seven-day study of an AI photo-logging app, estimated intake was 25% below a physiological reference on average.

Those numbers are not interchangeable. The first describes error on individual standardized photographs. The second captures a whole week of real use, including food that people may have forgotten to log. What they share is more useful than forcing them into a single range: portion size, hidden ingredients, and incomplete logging remain hard problems.

For general awareness, an editable estimate can be enough. For a deliberate calorie deficit, it deserves a quick review. For clinical nutrition or precise athletic targets, use labels, recipes, measured portions, and the method recommended by a qualified professional.

The Camera Knows Less Than It Appears To

Food recognition is the impressive part of the demo, and modern models can be genuinely good at it. They can identify ordinary foods, split a plate into plausible components, and save the first round of database searching.

But the calorie total comes after several separate guesses:

  1. What foods are present?
  2. How much of each food is there?
  3. How was each one prepared?
  4. Which nutrition record matches it?
  5. Did everything eaten make it into the log?

An app can get step one right and still fail at steps two through five.

Portion size is the most visible example. A single photograph flattens depth, and familiar objects are imperfect scale references. In a 2025 study of standardized food photographs, the models increasingly underestimated as portions got larger even though the setup included plates and cutlery for scale. The authors also noted that their own plate arrangement may have amplified the problem: vegetables sat in front and increasingly obscured the starch and protein in larger servings.

Gouache kitchen scene showing the same chili served in visibly different portion sizes
Recognizing chili is easier than deciding whether the bowl holds one cup or two.

Then there are the calories that do not photograph well. Two eggs cooked in a dry nonstick pan can look much like two eggs cooked in butter. Dressing can sit under lettuce. Peanut butter disappears into a smoothie. Mayo hides inside a sandwich. These are ordinary foods, not adversarial test cases.

Even a correct food and portion need the correct nutrient record. “Greek yogurt” could mean nonfat plain yogurt, sweetened whole-milk yogurt, or a branded cup with mix-ins. A strong database such as USDA FoodData Central only helps if the app chooses the right item and serving.

The final source of error is not artificial intelligence at all. It is the latte, second helping, cooking taste, or afternoon snack that never entered the app. A perfect estimate of an incomplete log is still an incomplete log.

What the Studies Actually Measured

Research on AI food estimation can sound contradictory because studies ask different questions. A useful way to read the evidence is to move from controlled photographs toward ordinary life.

Controlled food photographs

The 2025 study above tested 52 photographs representing systematic variations of 12 base dishes. Prepared meals and starchy components were photographed in three portion sizes; protein sources and vegetables appeared at the medium portion only. ChatGPT and Claude each had a mean absolute percentage error of about 35.8% for energy. The 95% confidence intervals were 27.3% to 45.1% for ChatGPT and 27.9% to 44.7% for Claude.

Mean absolute percentage error describes the average size of each percentage miss without letting an overestimate cancel an underestimate. On a 600-calorie meal, 35.8% is about 215 calories, but the study average is not a promise that any particular meal will miss by that amount. The models did better on some images and dramatically worse on others.

The worst misses were not subtle. Gemini mistook falafel for meatballs in one photograph, producing a 360% protein overestimate. Claude read scrambled eggs as pasta in another, pushing the carbohydrate estimate 1,788% too high. For a large lentil curry, ChatGPT estimated 255 grams when the measured portion was 480 grams. These examples are outliers, but they show why a polished calorie total still needs an editable food interpretation underneath it.

The authors reached a balanced conclusion: ChatGPT and Claude performed comparably to traditional self-reported dietary methods and could reduce user burden, but their systematic underestimation of large portions and variable macro estimates made them unsuitable for precise clinical or athletic assessment.

Larger model benchmarks

“AI” is not one estimator. The model underneath the app matters.

In June 2026, Nakagawa and Yamamoto tested three Claude models across 691 Japanese meal photographs, 1,463 US meal photographs, and roughly 1,000 packaged-food images. Capable models substantially outperformed the smaller model; packaged products were generally easier than meals; salt remained difficult. On the US meal dataset, Claude Sonnet’s whole-image calorie estimates had a mean absolute error of 116.9 calories per image.

A separate 2026 comparison of 40 vision-language models found that model architecture explained nearly all measured performance differences in its setup. Professional nutritionists outperformed every model, with the widest gap in protein estimation. Consumer smartphone photos worked better than the study’s controlled laboratory images, while adding more camera angles did not significantly help. The authors judged current models suitable for consumer calorie tracking, but said clinical-grade macronutrient profiling would require better model architecture.

The two studies also prevent a simplistic claim about prompting. Nakagawa found that a specialized visual-estimation prompt helped when paired with a capable model. The 40-model benchmark found no overall prompt strategy that remained significant after statistical correction. Supplying ingredient descriptions was a separate intervention: it slightly improved aggregate performance, although results varied by model. In practice, telling an app that the eggs were cooked in oil gives it a missing fact; asking the same app to “reason carefully” is merely a different instruction.

A week of real logging

Laboratory photos test the model. A week-long log tests the model, the interface, and the person using both.

In March 2026, Serra and colleagues asked 20 women with obesity to use the AI photo app SNAQ for seven days. The researchers compared reported intake with total daily energy expenditure measured using doubly labelled water. In weight-stable conditions, expenditure provides a physiological reference for average intake.

SNAQ’s estimate was 25% lower, a mean difference of 817 calories per day, with very wide disagreement for individuals. Conventional 24-hour recall was 50% lower in the same small cohort.

That does not show that SNAQ is universally 25% low, or that AI is twice as accurate as recall. The sample was small and specific. Doubly labelled water measures energy expenditure rather than photographing the exact food entering a person’s body. The result does show how much error can accumulate once missed snacks, preparation details, portion guesses, and normal logging behavior all enter the picture.

Taken together, the studies support a more specific set of conclusions:

  • Packaged products were easier than meals in the large Japanese and US benchmark.
  • Larger portions were increasingly underestimated in the 52-photo study, although its plate arrangement may have contributed to that pattern.
  • A calorie estimate can look plausible while one macronutrient is badly wrong.
  • Real-world logging adds forgotten food and incomplete entries to the model’s image-estimation error.
  • Supplying known ingredients and amounts can help; generic prompt complexity is not a dependable substitute for those facts.

Keep Known Facts Known

The easiest way to improve an AI estimate is to stop asking it to guess facts you already have.

What you are eatingBest starting pointWhat still needs checking
Packaged foodNutrition label or correct barcode entryAmount actually eaten
Homemade recipeIngredient list and batch yieldYour serving size
Chain-restaurant orderOfficial item or order nutritionPortion variation and substitutions
Visually clear mealPhoto plus a short descriptionOil, sauce, preparation, and portion
Meal remembered laterPlain-language descriptionMissing extras and approximate amounts

A packaged yogurt already has a brand, serving size, and nutrition panel. Photographing it and asking a model to infer those facts throws away useful information.

A pot of chili is a recipe problem. Two bowls can look identical while one contains lean turkey and the other contains higher-fat beef, more oil, cheese, and a different bean-to-meat ratio. Calculate the batch and divide by servings; our recipe calorie counter is built for that workflow.

When a photo really is the fastest start, add the detail the camera cannot see:

Photo alone: eggs, toast, berries
Better input: two eggs cooked in one teaspoon olive oil, sourdough toast with one teaspoon butter, and one cup berries

There is evidence that the extra context helps. A study of 195 dishes tested with progressively richer inputs found that estimates improved as standardized descriptions and ingredient amounts were added to the image. Removing the image again made performance worse. The photo and the words carried different information.

That study does not prove every text-first app beats every photo app. It supports the narrower, practical idea that amounts, ingredients, brands, and preparation details give an estimator evidence a picture cannot. Our text-versus-photo guide explores when each input style is easier.

Is It Accurate Enough for Weight Loss?

Sometimes. The answer depends more on the pattern of error than on one impressive meal.

For general awareness, a rough log can reveal that drinks count, restaurant portions are larger than expected, or protein disappears on busy days. Speed matters. An estimate you make consistently can be more useful than a perfect entry abandoned after three days.

For a deliberate calorie deficit, repeated underestimation matters. Errors that bounce above and below the truth may partly balance across a week. Missing the oil, sauce, and second serving in the same direction each day does not. If your intended deficit is 400 calories per day and the log routinely misses 250, the arithmetic leaves only about 150 calories before other sources of error. That is an illustration, not a claim that AI apps always miss by 250 calories; it shows why the direction of the miss matters.

If your weight trend is not moving as expected, do not immediately cut the target further. Audit one ordinary week for cooking fats, drinks, sauces, restaurant portions, unlogged bites, and recurring meals whose serving sizes have drifted. The goal is to find the missing information, not to punish yourself with a lower number. Our beginner’s calorie-counting guide explains how to use a multi-week trend without turning one day into a verdict.

For precise macros or athletic nutrition, inspect the full nutrient breakdown. The 40-model study found a particularly large gap between AI and nutrition professionals for protein estimation. Use labels and measured portions for foods that drive your targets: meat, fish, dairy, protein powder, grains, oils, and recurring recipes.

For medical nutrition, a consumer AI estimate should not make the decision. That includes insulin dosing, renal or sodium-restricted diets, eating-disorder treatment, pregnancy nutrition, and any plan with clinician-set targets. Record the meal if that is useful, but follow the measurement method your care team recommends.

A 10-Minute Calibration for Your Own App

Published studies cannot tell you how the app on your phone handles the meals you repeat. Three familiar meals can.

Choose:

  1. One packaged meal with a complete label.
  2. One simple home meal with measured ingredients.
  3. One mixed meal with a known recipe or official restaurant order.

Include a sauce, dressing, or cooking fat somewhere in the set. A test made entirely of bananas and packaged bars is a test designed to look good.

For each meal, write down:

  • Reference: the label, recipe calculation, or official order total.
  • First estimate: the app’s answer from the input you would normally use.
  • Corrected estimate: the answer after adding a missing amount, brand, sauce, or preparation detail.

You do not need a spreadsheet. Ask four questions:

  • Does the first estimate repeatedly land low?
  • Does the app preserve facts from labels and recipes?
  • Does one short correction materially improve the result?
  • Is reviewing the estimate still faster than logging the meal manually?

Labels and restaurant figures are anchors, not laboratory truth; rounding and portion variation remain. This is not a scientific validation study. It is a way to discover whether an app’s predictable weaknesses conflict with your goal.

If the mixed meal is off by enough to erase a meaningful part of your daily deficit, that workflow needs more information. If all three estimates are close enough for general awareness and easy to correct, the app may be doing exactly the job you need.

Where Amy Fits

Amy starts with a sentence rather than a food search. You can write, for example:

turkey sandwich with mayo, kettle chips, a pickle, and an iced coffee with oat milk

That input carries several details a photograph of a wrapped sandwich and cup would hide. It is still an estimate, so the result needs to be inspectable and editable.

Amy Food Journal nutrition details showing an estimated calorie total, cited sources, reasoning, and an option to edit the result
A number becomes useful when you can see the interpretation behind it and correct the result.

In our internal benchmark, last updated March 3, 2026, Amy received a weighted absolute-error score of 85/100. The formula weights calories 40%, protein 25%, carbohydrates 20%, and fat 15%. We chose the test set, references, weights, and scoring method, so this is product testing rather than independent clinical validation. The report also does not reduce the results to one directional-bias number, which means the score alone cannot tell you whether misses tend to run high or low. The full benchmark, disclosures, cases, and methodology are public for that reason.

The three-meal calibration above is still more personal than our benchmark. Run it on Amy, too.

The Bottom Line

AI calorie counting is already good at removing friction. It is not good enough to remove judgment.

Use the camera to identify what is visible. Preserve labels, recipes, and orders when they exist. Add the oil, sauce, portion, or brand the image cannot know. Then judge the app by the pattern it produces across the meals you actually eat, not by one polished demo.

An AI calorie estimate is a first draft, not a measurement.

Frequently Asked Questions

How accurate are AI calorie counters?

There is no defensible universal percentage. One controlled photo study found about 35.8% mean absolute percentage error for calories, while a separate seven-day app study found 25% average underestimation in a small, specific cohort. They tested different systems and outcomes, so neither number predicts your next meal.

Why do photo calorie apps underestimate meals?

Common causes include portion size, hidden oils and sauces, mixed dishes, an incorrect product or recipe match, and food that was never photographed. Larger portions were increasingly underestimated in one controlled study.

Is text logging more accurate than a food photo?

Text can carry hidden ingredients and preparation details; a photo carries visual context. Research supports combining useful words and images, but it does not prove that every text-first app is more accurate than every photo-first app.

Can AI calorie counting work for weight loss?

Yes, when it makes logging sustainable and its estimates are easy to review. Check recurring meals against labels or recipes, watch for a repeated low bias, and use a multi-week weight trend rather than trusting one day’s total.

Should I use an AI calorie counter for medical nutrition?

Not as the sole source of truth. Consumer estimates should not drive insulin dosing, renal or sodium restrictions, eating-disorder treatment, pregnancy nutrition, or other clinical decisions. Use the method recommended by your clinician.

Keep Reading

All articles

Start tracking with Amy

Track calories like writing in Apple Notes. Just type what you ate.

Download Free