Skip to content
MyBodyAI
LEARN
Topic Guides Blog Glossary Tools
DEVICES
Compare all devices Supported devices
Products Pricing FAQ About MyBodyAI

Pasting Your Watch Data Into ChatGPT? Here Is What It Misses

Two weeks from one account, the same resting heart rate, and yet one of them was trouble. What tells them apart, and why a chat window has no way to see it.

Published 2026-07-24 · Updated 2026-09-09 6 min read AI & Health
Raw health data pasted into a chat window versus a continuous personal baseline

Two weeks that look identical

Here are two weeks from one anonymised account, worn every night, four weeks apart.

Weekly averageWeek AWeek B
Resting heart rate60.960.0
Sleep7h 03m5h 51m
Minutes awake during the night2697
Sleep score6537

Take any reference table you like and both weeks pass. A resting heart rate around 60 is described as good to excellent for a man of that age. Neither week contains a single value a doctor would circle.

Week B is the one that was trouble.

What separates them

Notice what the resting heart rate does across those two weeks: almost nothing. 60.9 against 60.0. If you had pasted only that number into a chat window, there would have been nothing to say.

Everything else moved at once. Time awake in the night was nearly four times longer, more than an hour of sleep a night was lost, and the sleep score fell by nearly half. On the Sunday of week B the account recorded a self-rating of 2 out of 5 for energy, mood and physical state alike. Two nights that week recorded no REM sleep at all.

The 90 day average for this account sits at 57.4 beats per minute, with a normal range of 54 to 61. Both weeks therefore sat near the top of that range. The resting heart rate was never the story. The story was that a group of signals moved together, in the same direction, against the personal norm, and stayed there for eight days.

That is what reading health data actually means, and it is the part a chat window structurally cannot do.

Why the chat window cannot get there

It does not know your normal. To tell you that 60 is elevated, something has to know that you usually sit at 57.4 and that you are usually awake for 26 minutes, not 97. A one-shot paste carries no past. The best a model can do is compare you to a population average, which flattens exactly the individual differences that make the data worth reading in the first place.

And it forgets you the moment you close the tab. Health is a trajectory, not a snapshot. Has your recovery been sliding for three weeks? Did your resting heart rate creep up for four days before you felt the cold arrive? Is your sleep debt quietly building? None of that survives a fresh conversation.

And when it has no basis, it produces one anyway

A language model does not look your answer up. It predicts the next plausible words. Most of the time the result is reasonable, and that is precisely what makes the failures hard to catch: a wrong answer arrives in the same calm voice as a right one.

This is measured, not anecdotal. When physicians tested general chatbots on real patient questions, a substantial share of answers were problematic and a smaller but real fraction were unsafe (npj Digital Medicine, 2026). Separate work on healthcare queries documented the same pattern of confident fabrication and built benchmarks for it (MedHalu, 2024).

A threshold like "an HRV below 40 means overtraining" reads exactly like a real one. There is no universal number of that kind, because the value depends on the person and on the device that measured it. You will act on the invented one just as readily.

The core issue: a chatbot cannot flag the difference between an answer it knows and an answer it produced to fill a gap. Both come out fluent. With your body, that difference is everything.

How to get a better answer today

None of this means you should stop. It means the paste needs preparing. Four things make a real difference:

Give it your baseline, not just the week. Work out your own average for the last 60 to 90 days for each value and write it into the prompt. Without it the model has nothing to measure against but a textbook.

Say who you are. Age, sex, roughly how much you train. It changes the reference frame for almost every number.

Give it permission to say it does not know. One sentence: where you are unsure or a value is missing, say so and do not fill in the gap. It works more often than you would expect, and it turns silent invention into a visible admission.

Never let a blank become a zero. Sensors miss nights. If a cell is empty, say so, or the model will read it as a night with no sleep and the whole week will be read wrong.

One thing preparation cannot fix: once you paste months of biometric data into a third-party chatbot, that data sits on someone else's servers, where it may be retained, reviewed or used for training. Health data is the most sensitive category there is. Paste what you need to and no more.

Where this leaves the idea

Using AI to read your health data is the right instinct. The problem is not the AI, it is asking a stranger who meets you for the first time every single morning.

MyBodyAI is what happens when that preparation stops being your job. Your baseline is computed continuously from your own history, kept across every device you connect, so switching watches does not reset your past. The scores themselves come from deterministic code rather than generated text, which is why a value can never be invented. Where something is estimated it says estimated, and where the data is not there it says so instead of producing a reassuring verdict. You can try it for seven days without a card.

References

  • Large language models provide unsafe answers to patient-posed medical questions. npj Digital Medicine (2026). Article s41746-026-02428-5.
  • Agarwal S et al. MedHalu: Hallucinations in Responses to Healthcare Queries by Large Language Models. arXiv (2024). arXiv:2409.19492.

Found this useful? Share it.