Smiling woman with blonde hair sitting on a bed in a linen robe

Sleep Trackers: What Seven Devices Got Right, and What They Got Wrong

Evidence Review  ·  The Reading Room, Vol. 32  ·  9 min read  ·  What the devices measure well, what they do not measure at all, and the documented cost of believing the wrong number

Based on the published research of Kelly Glazer Baron, PhD — Associate Professor, Division of Public Health, University of Utah Health, and corresponding author of the paper that named orthosomnia. Faculty profile

Are sleep trackers accurate? Tested against laboratory polysomnography, seven devices did well at telling sleep from wake and poorly at telling one stage from another. The distinction decides what the number is worth.

The report is waiting at six fifteen. Forty-one minutes of deep sleep. Sleep score 62. A small graph in which the night is coloured in bands, and one of the bands is very much thinner than it was on Sunday.

A day gets shaped by that number before anyone has had coffee. This review sets out which parts of the report are supported by validation testing, which parts are not, and what has been documented about the reading itself becoming the problem.

The short answer:

  • Seven consumer devices were tested against laboratory polysomnography across three consecutive nights in 34 adults, producing around 102 recordings.¹
  • Most devices performed comparably to or better than research-grade actigraphy at distinguishing sleep from wake.¹
  • All six devices that reported sleep stages overestimated light sleep. Three overestimated deep sleep and three underestimated dream sleep.¹
  • The study’s own summary of stage output was that device assessments were inconsistent
  • In a separate clinical report, patients treated tracker data as more consistent with their experience of sleep than polysomnography or actigraphy, and sought treatment on that basis.²

Scope and method of this review

This piece covers one question: how closely consumer sleep-tracking devices match the reference standard when both are measured on the same night. It draws on the largest head-to-head validation study of multiple consumer devices, published in Sleep in 2021, and on the clinical paper that named the behavioural problem this creates.¹²

What it does not cover: whether any particular current model is accurate, since device firmware and algorithms change between generations and the 2021 study tested the models available then; whether trackers improve sleep, which is a different question and a much thinner literature; and any medical-grade home testing device, which is a separate category with separate validation.

Polysomnography is the reference standard here: an overnight laboratory recording of brain activity, eye movement, muscle tone, breathing and heart rhythm, scored by a technician. It is what “actually asleep” means in this literature.

What the validation study found

What was measuredResultLimitation
Sleep versus wake, seven devicesMost matched or beat research actigraphyActigraphy itself tends to over-call sleep in poor sleepers
Total sleep time, Fitbit Alta HR and ResMed S+Minimal bias against polysomnographyTwo devices out of seven; models now superseded
Total sleep time, Garmin devicesSignificantly overestimated sleep durationDirection of error differs by manufacturer, so results do not generalise
Light sleep, all six stage-reporting devicesOverestimated by every deviceA consistent error is still an error
Deep and dream sleepThree devices overestimated deep, three underestimated dream sleepNo single correction factor can be applied by a user
Sample34 healthy adults, 22 women, mean age 28.1Healthy and young; no data here on women in midlife or on disrupted sleep

Source: Chinoy, E. D., et al. (2021). Performance of seven consumer sleep-tracking devices compared with polysomnography. Sleep, 44(5), zsaa291.¹

The shape of the finding is consistent and easy to state. A wrist device is measuring movement, and in some models heart rhythm, and inferring sleep from that. Movement tells you reliably whether somebody is asleep. It does not tell you which stage of sleep they are in, because stages are defined by brain activity that a wrist cannot see.

So the total at the top of the report has real support behind it. The coloured bands underneath it are an inference from an indirect signal, and they were wrong in the same direction, for every device that produced them.¹

If the report has become the first verdict of the day, that is worth noticing separately from whether the number is right. The Quiet Audit is a private ten-minute pass through what is being weighed each morning. Start here.

Key terms

Actigraphy is research-grade movement monitoring, usually worn on the wrist. It is the established non-laboratory method and the benchmark consumer devices are usually compared against.

Sleep staging is the division of a night into light, deep and dream sleep, scored in the laboratory from brain activity rather than from movement.

Orthosomnia is the term coined in 2017 for a preoccupation with achieving perfect sleep as reported by a tracking device, described by its authors as a perfectionistic quest for the ideal sleep in order to optimise daytime function.²

How strong is this evidence?

Moderate, and split by what is being asked. The evidence that these devices distinguish sleep from wake reasonably well is good. The evidence that they cannot reliably report sleep stages is also good, and consistent across every stage-reporting device tested.¹

What the grade does not cover: 34 healthy adults with a mean age of 28.1 is a small, young, well sample.¹ It contains no information about how these devices behave in a fifty-two-year-old whose sleep is fragmented by heat, by a partner, or by early waking. Device algorithms have also been revised repeatedly since testing, in both directions. The structural conclusion holds because it follows from what a wrist sensor can physically detect. The specific per-device figures belong to 2021.

The documented cost of the wrong number

In 2017 a group of sleep clinicians described a pattern in their own clinic and gave it a name.²

Patients were arriving having self-diagnosed insufficient sleep or insomnia on the basis of tracker output showing periods of light or restless sleep. That is not remarkable in itself. What the authors reported next is: when objective testing disagreed with the device, patients tended to treat the device data as more consistent with their experience of sleep than polysomnography or actigraphy.²

The clinical problem is circular and well recognised in sleep medicine. Effort and attention directed at falling asleep interfere with falling asleep, so a nightly score that must be improved becomes a nightly performance to be monitored. The authors framed the task as balancing patient education about device validity against patients’ enthusiasm for objective data.²

Two things follow, and they sit together without contradiction. A device is useful for spotting a pattern across weeks, such as consistently short time in bed or a bedtime that has drifted an hour later since January. It is a poor instrument for judging last night, because last night’s stage breakdown is the part that failed validation. The evidence-based response to lying awake does not require a number, and the mechanism behind being exhausted and still awake is not visible to a wrist sensor at all.

When this belongs with a doctor

No consumer tracker diagnoses a sleep disorder, and none of the devices in this study were validated for that. Loud snoring, witnessed pauses in breathing, waking with a headache or a dry mouth, or daytime sleepiness heavy enough to affect driving are reasons for a clinical assessment rather than a firmware update. The same applies to difficulty sleeping at least three nights a week for three months or more, which is the threshold at which insomnia becomes a clinical diagnosis with established treatment. Breathing-related sleep disruption in midlife women is under-recognised and is not something a sleep score will find.

The one move

The one move

Keep the device and stop reading it in the morning. Look once a week instead, and look only at total time asleep and at what time you went to bed. Those are the two outputs with validation behind them. Ignore the stage breakdown entirely; it was wrong in the same direction for every device tested.

If the morning number has become the day’s first verdict, the Quiet Audit is a private pass through what else is being scored before breakfast. Start here.

The source for this piece

Kelly Glazer Baron, PhD

Associate Professor in the Division of Public Health at University of Utah Health, and corresponding author of the 2017 paper that coined the term orthosomnia.

What we read: Baron, K. G., et al., Journal of Clinical Sleep Medicine (2017).²

Where to follow her work: Faculty profile

Who else has measured this

The clinical pattern and the device-validation numbers come from two separate teams. Both are worth naming.

Evan D. Chinoy, PhD (author profile · LinkedIn), lead author of the 2021 validation study in Sleep, ran the head-to-head comparison of seven consumer devices against laboratory polysomnography that this review’s numbers are drawn from.¹

A device-validation team and a clinical-behaviour team, working independently, describing two halves of the same problem: what the devices get right, and what happens when a person trusts the part they get wrong.

A wrist device is measuring movement and inferring sleep from that.

Questions this review is asked

Are sleep trackers accurate?
For distinguishing sleep from wake, most tested devices performed as well as or better than research-grade actigraphy.¹ For sleep stages, all six stage-reporting devices overestimated light sleep, and the study described their stage output as inconsistent.¹

Why can a watch not tell deep sleep from light sleep?
Because sleep stages are defined by brain activity, measured at the scalp. A wrist device records movement and, in some models, heart rhythm, and infers a stage from that indirect signal. The inference is the part that fails validation.

Which device was most accurate?
In that 2021 comparison, the Fitbit Alta HR and the ResMed S+ showed minimal bias for total sleep time, while Garmin devices significantly overestimated sleep duration.¹ Those specific models have since been superseded, so the finding is about that generation rather than about a brand.

Is a tracker worth keeping at all?
For patterns across weeks, yes: time in bed, bedtime drift and total sleep are the outputs with support. For judging a single night, the evidence does not back the detail the report presents most prominently.

Can a tracker make sleep worse?
That is what the orthosomnia paper describes. Patients pursued perfect scores, treated device output as more credible than laboratory testing, and presented for treatment on that basis.² It is a clinical case description rather than a controlled trial, so it establishes that the pattern occurs, not how often.

Does any of this apply to women in midlife specifically?
Not directly, and that is the honest gap. The validation sample had a mean age of 28.1 and was healthy.¹ The structural limitation on staging applies regardless of age; the per-device numbers were not measured in this population.

On the researcher. Kelly Glazer Baron is a clinical sleep psychologist; the orthosomnia paper she co-authored is a clinical case series describing a pattern the authors observed, not a controlled trial establishing how often it occurs.

Sourcing, disclosure and disclaimers

This is not medical advice. Blue Leaf Journal publishes general information for a general readership. Nothing here is a diagnosis, a treatment recommendation, or a substitute for care from a clinician who knows your history.

We are not clinicians. Blue Leaf Journal is an independent publication. We read published research and translate it. The findings belong to the researchers and institutions named and linked above; the plain-English rendering is ours, and so is any error in it.

No affiliation and no endorsement. Blue Leaf Journal is not affiliated with Kelly Glazer Baron, Evan D. Chinoy or the University of Utah Health. None of them has reviewed, approved or endorsed this article, and none of them is responsible for it. They are cited because their published work is the evidence for what it says.

No commercial relationship. Nobody named above paid for or was paid for this coverage. There are no affiliate links, no sponsored placements and no gifted products in this article. No device or tracker mentioned anywhere on this page is sold by us or by anyone paying us.

How this was checked. Every figure above is drawn from the primary papers, each linked in the references, and can be verified there. Sources were checked on 6 September 2026. If you find something we have got wrong, write to miriamalderton@blueleafjournal.com and we will correct it and say that we did.

References
1. Chinoy, E. D., Cuellar, J. A., Huwa, K. E., et al. (2021). Performance of seven consumer sleep-tracking devices compared with polysomnography. Sleep, 44(5), zsaa291. Read the study
2. Baron, K. G., Abbott, S., Jao, N., Manalo, N., & Mullen, R. (2017). Orthosomnia: Are Some Patients Taking the Quantified Self Too Far? Journal of Clinical Sleep Medicine, 13(2), 351-354. Read the paper

Miriam Alderton is Research Editor at Blue Leaf Journal. She reads the methods section first and the abstract last, and every figure in this piece is linked to its source above.

Written by Miriam Alderton, Research Editor, for Blue Leaf Journal. Updated: 30 August 2026.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *