1. BananaCal
  2. Blog
  3. Science

We Tested AI Calorie Counting on 68 Weighed Meals: Results

We sent 68 lab-weighed meals to two AI models, 3 times each. One overestimated by 51% on average, the other by 5%. Here's where AI still gets portions wrong.

Most tests of AI calorie counting use a handful of photos and an eyeballed “real” value. We wanted a harder benchmark: meals where every ingredient was weighed in a lab, so the true calories are known. We ran 68 of them through two AI vision models, three times each, and measured how far the estimates landed.

The short answer: the newer model we tested (gpt-6-luna) was off by 31% on average and by 22% on a typical meal, with almost no overall bias (+5%). The older one (gpt-5.6-luna) overestimated by 51% on average and answered exactly 520 kcal for 8% of meals, whatever was on the plate. Both models pulled every plate toward an “average meal”: small snacks came out too high and large plates too low.

How we ran the test

  • The meals. We took 68 dishes from Nutrition5k, a public dataset released by Google Research. Each plate was assembled in a cafeteria with every component weighed, and its calories and macros were computed from those weights. Our sample ranged from 66 to 1,051 kcal (median 498 kcal), chosen to cover small snacks and large plates alike.
  • The photos. We used the overhead photo of each plate, compressed the way a phone app sends it.
  • The models. Two OpenAI vision models, gpt-5.6-luna and gpt-6-luna, given the same instructions that a calorie tracking app uses in production: name the dish, list the ingredients with their weight, give calories and macros.
  • The repeats. Every photo was sent 3 times to each setup, to see whether the same picture gives the same answer.
  • The scoring. For each answer we compared the calories with the weighed value: the average error, the error on a typical meal (median), the share of answers within 20%, and the bias (whether the model leans high or low).

The results

Older model Newer model
Average error 58% 31%
Error on a typical meal (median) 29% 22%
Answers within ±20% of the weighed value 38% 46%
Average bias +51% (too high) +5%
Average error in kcal 156 kcal 111 kcal
Answers of exactly 520 kcal 8% 0.5%
Change when the same photo is sent again 4.7% 4.9%

The newer model halved the error and nearly removed the bias. Its 31% average error sits right where the best published results are: a 2025 study found about 36% for ChatGPT-4o and Claude 3.5 on 52 weighed meals.

The older model’s numbers hide a telling habit. Across 204 answers it used only 32 different calorie values, and 520 kcal came up 17 times, for plates that really held anything from a small salad to a full lunch. It wasn’t estimating the portion so much as reaching for a “typical meal” number.

AI pulls every plate toward an average meal

Here is the median error by size of the real meal:

Real calories of the plate Older model Newer model
Under 200 kcal +139% +61%
200 to 350 kcal +74% +23%
350 to 500 kcal +24% -12%
500 to 700 kcal +11% -5%
Over 700 kcal 0% -19%

The pattern is the same for both models, only stronger in the older one: small portions are overestimated and big ones are underestimated. A 150 kcal snack can come out at 250. A 900 kcal plate can come out at 730.

This matches what researchers have found with other AI models, where underestimation grows as the portion gets bigger. The cause is the photo: from above, the model can’t see depth or weigh anything, so it hedges toward what a meal usually looks like.

Same photo, same calories?

Mostly yes. Sending the same photo again moved the estimate by about 5% on average with either model. That is small next to the error itself.

The exception is the ambiguous photo. One picture showed several doughnuts on a plate. The newer model usually read it as about 2,100 kcal, but in 1 run out of 12 it saw a single doughnut and answered 300 kcal. Asking the model to count the pieces first fixed it: 0 misreads in 17 runs, every answer between 1,740 and 2,280 kcal.

Words with grams beat photos

We also wrote the same 67 recipes out as text with exact weights (“68 g brown rice, 21 g mixed greens, 3 g olive oil”), so the only job left was the arithmetic.

Setup Error on a typical meal Answers within ±20%
The model reads the grams and gives the total 10% 75%
The model gives grams and values per 100 g, the app does the sum 4.5% 88%

The first setup still makes mistakes you wouldn’t expect. For “200 g of 0% Greek yogurt with 30 g of honey” it answered 181 kcal. The right answer is about 209 kcal: around 118 for the yogurt and 91 for the honey. When the model only has to know the values per 100 g and the arithmetic is left to code, those slips disappear.

Tricks that did not help

  • Asking several times and taking the middle answer. Three or five answers per photo, then the median, gave the same 31 to 33% average error at two to three times the cost.
  • A detailed step by step format for photos. Asking for every ingredient with its values per 100 g matched the simple format on accuracy, but the same photo varied more between runs (about 9.5% instead of 5%). The main reason was cooking oil: listed as a separate item in some runs and left out in others.

What this means when you log a meal

  1. Small snacks: check the number. A photo of a yogurt or a piece of fruit is more likely to be overestimated than underestimated.
  2. Big plates: check the portion and add what you can’t see, like cooking oil and dressing. This is where calories get lost.
  3. Know the weight? Type it. A description with grams was about five times more precise than a photo in our test.
  4. Several pieces on the plate? Make sure the count is right.
  5. Packaged food: scan the barcode instead of taking a photo. The label beats any estimate.

If you’re setting a target, start from our calorie calculator. For the bigger picture on accuracy, including studies on self-reporting, read are AI calorie counting apps accurate, and for the basics of logging see how to count calories.

Limits of this test

  • Lab photos. Nutrition5k plates were shot from above in a fixed rig in a few California cafeterias. Photos taken at home, at an angle, with different cuisines will behave differently.
  • 68 meals. Enough to see clear patterns, not to rank every model on every kind of food.
  • A snapshot. These are two specific models in October 2026. Models change often, so the numbers will too.
  • Who ran it. We make BananaCal, a calorie tracking app, and we ran this test to choose its setup. After it, we moved photo analysis to the newer model with the counting step, and text analysis to the “values per 100 g, sum in code” approach. We have no relationship with OpenAI other than paying for its API.

Summary

  • On 68 weighed meals, the better AI model was off by 31% on average and by 22% on a typical meal, with almost no bias.
  • Every model pulled plates toward an average meal: small ones too high, big ones too low.
  • The same photo gave nearly the same answer each time, except on ambiguous plates.
  • Writing the grams cut the typical error to about 4.5%.
  • Asking the AI more times didn’t make it more accurate.

Frequently asked questions

How accurate is AI at estimating calories from a photo?

In our test on 68 weighed meals, the better model was off by 31% on average and by 22% on a typical meal. That is in line with published studies, which report average errors of about 35% for general AI models.

Does AI overestimate or underestimate calories?

Both, depending on the plate. Every model we tested pulled estimates toward an average meal, so small snacks were overestimated and large plates were underestimated.

Is it more accurate to describe a meal in words than to photograph it?

When you know the weights, yes. With the grams written out, the typical error fell to about 4.5%, against 22% for a photo of a similar plate.

Does the same photo always give the same calories?

Almost. Sending the same photo again changed the estimate by about 5% on average, but ambiguous photos can occasionally be read in a completely different way.

Sources

  1. Thames Q et al. Nutrition5k: Towards Automatic Nutritional Understanding of Generic Food. CVPR, 2021
  2. Nutrition5k dataset (Google Research, CC BY 4.0)
  3. Fridolfsson J et al. Performance Evaluation of 3 Large Language Models for Nutritional Content Estimation from Food Images. Curr Dev Nutr, 2025
  4. O'Hara C et al. An Evaluation of ChatGPT for Nutrient Content Estimation from Meal Photographs. Nutrients, 2025
  5. USDA FoodData Central (Greek yogurt, nonfat, plain; honey)

This article is general information, not medical advice. If you have a health condition, are pregnant or take medication, talk to your doctor or a registered dietitian before changing your diet.