Is my wearable measuring this, or guessing?
Some numbers on your wearable come straight from a sensor: your optical heart rate, your skin temperature, your blood oxygen reading. Others are the output of an algorithm that combines several sensor inputs with assumptions about your body: your calorie burn, your VO2 max, your sleep stages, your stress or recovery score. Both kinds of numbers can be useful, but treating an estimate as if it were a direct measurement is where most wearable confusion starts.
What does your wearable actually measure directly?
A small set of readings come from a sensor reading a physical signal with no modeling layer in between:
- Heart rate, from an optical sensor (PPG) or, on some devices, an electrical sensor (ECG) reading your pulse in real time.
- Skin temperature, from a thermistor against your wrist or finger.
- Blood oxygen saturation (SpO2), from a pulse oximetry sensor, though even this has known accuracy limits on darker skin tones and during motion.
- Motion, from an accelerometer and gyroscope, which is the raw input, not steps or calories yet.
These are closer to a thermometer reading than a prediction. There is still measurement error, sensors drift, and skin contact matters, but there is no model guessing what your body is doing between data points.
What does your wearable estimate rather than measure?
Everything downstream of those raw signals is a calculation, and several of the most-watched numbers on your dashboard fall here:
Calories burned. No consumer wearable measures energy expenditure directly. It runs your heart rate and movement through a formula that also factors in your age, weight, height, and sex, and the result can be off substantially. A cross-brand accuracy analysis found wide variance in calorie estimate accuracy across devices, with error rates that made calorie counts far less reliable than heart rate itself (JAMA Network Open, 2020 wearable energy expenditure validation study).
VO2 max. Wearables never measure your actual maximal oxygen consumption, which normally requires a lab treadmill test with a gas mask. Instead they infer it from the relationship between your heart rate and pace during exercise, or in some cases from resting heart rate alone. A meta-analysis by the INTERLIVE network found that wearables using exercise-based algorithms were meaningfully more accurate than those estimating from resting conditions, but both approaches carry real error margins, and accuracy differs by fitness level and sex (Springer, INTERLIVE network systematic review and meta-analysis).
Sleep stages. Your wearable does not read brain waves. It infers light, deep, and REM sleep from movement and heart rate patterns, which is a reasonable proxy but not the same signal a clinical sleep study captures with EEG. A meta-analysis of consumer wrist-worn sleep trackers against polysomnography found reasonable agreement on total sleep time but larger discrepancies in identifying specific sleep stages, particularly light versus deep sleep (PMC, meta-analysis of consumer sleep tracker performance vs polysomnography). For more on what each stage means, see our sleep stages guide.
Stress and recovery scores. These are proprietary composites, typically built from HRV, resting heart rate, and sleep data, run through a formula specific to that brand. There is no independent clinical reference for "stress score" the way there is for heart rate, which is part of why Garmin and Oura can disagree on the same underlying data. See our breakdown of what a recovery score means for how these composites are typically built.
Why does this distinction matter for how you use the data?
A measured number and an estimated number deserve different levels of trust and different reactions.
If your measured heart rate spikes to 140 BPM sitting at your desk, that is a real physiological event worth paying attention to. If your estimated calorie burn says 2,400 instead of an actual 2,200, that is model error within a normal range, not a signal that something changed in your body. Treating every number on the dashboard with the same confidence leads people to either over-trust noisy estimates or under-trust genuinely useful signals.
The safest approach is to use measured metrics for absolute values and estimated metrics for trends over time. Your exact VO2 max number might be off by several points, but if your device's VO2 max estimate is trending upward over three months using the same algorithm, that trend is more informative than the number itself, because the model's own bias stays roughly constant while the underlying fitness change moves the needle. For more on how calorie estimates specifically break down, see our calorie tracking guide.
How can you tell which category a number falls into?
| Category | Examples | How to treat it |
|---|---|---|
| Directly measured | Heart rate, skin temperature, SpO2, motion | Trust the absolute value more, expect small sensor error |
| Estimated from sensors + a model | Calories, VO2 max, sleep stages, stress and recovery scores | Trust the trend more than the single-day number |
| Estimated from you-reported inputs | BMR-based total calories, some readiness scores | Only as accurate as the inputs you gave the app |
If a metric requires an algorithm to translate raw signal into a headline number, assume it carries a real margin of error, usually disclosed somewhere in the company's methodology page if you look for it.
Why does this matter more once AI is layered on top?
When an AI health coach reads your data and tells you something changed, it is working from whatever mix of measured and estimated numbers your devices produced. A pattern built on measured heart rate and skin temperature carries different confidence than one built partly on an estimated stress score, even though both might appear in the same sentence. This is exactly why context, not just more data, is what makes an insight useful. Our guide on why raw wearable data needs context covers how baselines and trends fill that gap. And for what an AI can responsibly say from that mix of measured and estimated inputs, see what AI can and cannot tell you about your health data.
How does MotionSync handle the difference?
MotionSync pulls both measured and estimated metrics from Apple Health, Garmin, Oura, Fitbit, and Google Fit, and its AI insights are built to weight them differently: trends in estimated scores like recovery or stress are read over multiple days rather than as single-day facts, while measured signals like resting heart rate and temperature deviations get flagged sooner because the underlying reading is more direct. Knowing which kind of number you are looking at is the first step to reading it correctly, whether or not you use an AI layer at all.
FAQ
Is heart rate always accurate on a wearable?
Optical heart rate sensors are generally reliable at rest and during steady activity, but accuracy can drop during high-intensity intervals, in cold weather, or with a loose fit, since the sensor needs consistent skin contact to read your pulse reliably. It is still the most directly measured metric on most devices, just not infallible.
Why does my VO2 max estimate change without me changing my fitness?
VO2 max estimates are sensitive to which workouts you log and how consistently your heart rate response is captured during exercise. A device that has less recent outdoor running data, for example, may re-estimate based on older or different inputs, which can shift the number without any real change in your aerobic capacity.
Should I ignore estimated metrics entirely?
No. Estimated metrics like recovery scores and sleep stages are still useful, especially as trends over weeks, because the algorithm's own bias tends to stay consistent even if the absolute number is not perfectly accurate. The mistake is treating a single day's estimated number with the same certainty as a measured one.


