Your wearable flashes a low readiness score. It's squat day, and the plan in your program calls for a top set at a target RPE. The device is nudging you toward an easy day; the log is asking you to show up as written. Something has to break the tie.
A wearable readiness score is a useful trend signal, not a validated instruction. Your training log, the logged load, reps, RPE or RIR, and the trend across recent sessions, should decide today's session. The readiness score works best as context that confirms or questions a pattern the log already shows, not as a single-day verdict on its own.
What is a wearable readiness score, and what does it actually measure?
A readiness score is a daily composite number, built from heart rate variability, resting heart rate, sleep, and recent training load, that estimates how recovered your body is for the day ahead. It is a summary of several passive signals, not a direct measurement of strength, energy, or muscle recovery.
Oura, Whoop, and Garmin have each published their own version for several years, and Apple joined them with the Readiness feature on the Watch Series 12 and Ultra 4. The inputs are broadly similar across brands: overnight HRV and resting heart rate trends, sleep duration and consistency, and how much training stress you've accumulated recently. Each brand then weights and blends those inputs into its own proprietary score, which is why an 80 on one platform doesn't necessarily mean the same thing as an 80 on another.
Most consumer wearables estimate HRV optically, from the timing of pulse waves at the wrist or finger, rather than from the electrical signal an ECG captures. That optical measurement, sometimes called pulse-rate variability, tracks reasonably well with ECG-derived HRV while you're resting or asleep, but the agreement gets weaker with movement, poor sensor fit, or changes in skin blood flow. The number on your screen is an estimate of an estimate, refreshed once a day.
Does following a readiness score improve training outcomes?
Not clearly, and the closest head-to-head test available shows it doesn't have to cost you anything either way. DeBlauw et al. (2022) randomised 55 recreationally active adults to six weeks of high-intensity functional training, either on a fixed, predetermined schedule or with intensity modulated by daily HRV readings.
The HRV-guided group trained significantly fewer days at high intensity than the fixed-schedule group, a mean difference of -13.56 ± 0.83 days (p < 0.001). Despite doing meaningfully less hard training, that group showed no significant between-group difference in resting heart rate, lean mass, fat mass, strength, or work capacity compared with the group that followed the fixed plan.
Read plainly, that's a wash with a silver lining: letting HRV trim your hard days didn't produce better results than just following the plan, but it also didn't cost the HRV-guided group anything measurable while asking less of them. It's one trial, in one training style, over six weeks, so it doesn't settle the question for strength training generally. It does undercut the idea that a readiness score reliably makes your training more effective, because the group that ignored it entirely did just as well.
Why don't readiness scores reliably tell you what to do?
Readiness scores struggle to tell you what to do because the final composite number is rarely validated against a real outcome, and the physiological signal underneath it is an imperfect proxy in the first place. A trend across many mornings is more trustworthy than any single reading.
A 2025 review of 14 wearable health scores across 10 major brands found that while most scores draw on HRV, resting heart rate, sleep, and activity data, few manufacturers have published evidence validating the final number those inputs produce, according to reporting by Gadgets & Wearables (2026). The review's strongest counter-example was still a correlational one: a Whoop study of 389 professional golfers found players performed about half a stroke better when their Recovery score was 10 points higher, but the study had no control group, so it can't separate "a good Recovery score causes better golf" from "the kind of week that produces a good Recovery score also produces better golf."
"Research suggests these scores don't predict injury in a straightforward way, but that the data are useful for seeing trends over time." — Anna Zucker, Journal of Medical Internet Research (2026)
Zucker's 2026 JMIR piece makes the same point from the injury-prevention angle. Over one-third of adults in the United States now use a wearable, and many lean on the readiness score to guide training, but the research, including a narrative review of 16 studies in the Journal of Orthopaedic Reports, doesn't show that the score itself reliably predicts injury risk. The same review notes the data behind the score, tracked over weeks rather than read as a single morning's number, is where the genuine value sits.
None of this means the sensors are lying to you. A falling score across several consecutive days, alongside poor sleep and a climbing resting heart rate, is a real physiological signal worth paying attention to. The part that isn't settled is whether the wearable's own interpretation of that signal, "rest today" or "push today", is the correct call for your specific training session.
| Dimension | Readiness score | Training log |
|---|---|---|
| What it measures | HRV or pulse-rate variability, resting heart rate, sleep, recent load | Load, reps, RPE/RIR, completion, and PRs from sets you actually performed |
| Time horizon | A single morning, refreshed daily | Multi-session and multi-week trend |
| Validated against training outcomes? | Composite score largely unvalidated; injury prediction weak (Zucker, 2026) | Directly tied to the work you did; trend visible across your own history |
| What it should change about today's session | A contextual nudge worth a second look | The primary input for load, volume, and effort decisions |
When should a low readiness score change your session, and when should the log win?
Let the log win on any single low-score day when your recent sessions look normal; let the score add weight only when it lines up with a pattern the log already shows. A readiness score with nothing behind it in your training history is closer to noise than instruction.
Single low score, clean log: if your last few sessions hit the prescribed reps at the planned RPE or RIR, treat one low reading as noise from a bad night's sleep, a late meal, or an unusually stressful day. Proceed with the session as programmed; if anything, hold the load rather than pushing past what's written, but don't pre-emptively cut the plan.
Low score, matching log pattern: if the low reading shows up alongside a pattern you can already see in the log, rising RPE at the same load across sessions, reps drifting down, or estimated 1RM flattening or falling, treat it as a real signal. That's the point to reduce volume or intensity for the session, the same logic covered in the strength plateau guide for spotting fatigue from logged data alone.
Low score, no log pattern at all: if several consecutive low readings haven't shown up as anything in your actual performance, no rising RPE, no falling reps, no worse bar speed, keep training as planned and keep watching. The autoregulated training guide covers the same warm-up-feel and RPE-trend checks that should carry more weight than a single passive number.
How does IronLedger fit a readiness score into a self-coached plan?
IronLedger treats a readiness score as context, never as the mechanism that changes your programmed weights. The product page verifies that boundary directly for the Oura connection: it reads your daily readiness score and offers a training recommendation based on it, but as the page states, "it informs you, it does not change your programmed loads." Availability depends on platform, release stage, permissions, and account configuration, the same caveat that applies to IronLedger's other integrations.
What does change your programmed loads is the log. Every set records load, reps, RPE or RIR, notes, rest, and completion state, and the progression rules work from that record: a linear load rule that adds a fixed step, an RPE/RIR hold-advance rule that holds the load when reported effort climbs past a ceiling, and a percentage-of-PR rule that recalculates as your personal records move. Each recommendation is labelled with the rule that produced it, so a lifter can see why the app suggested a number instead of trusting a black box.
The progress page is where the multi-session comparison actually happens: personal records, Epley-formula estimated 1RM, full training history with no cutoff, and per-lift trends and tonnage. That's the record a self-coached lifter should check before deciding whether a low readiness score matches a real pattern or is an isolated bad morning, the same distinction the deload week guide uses to separate ordinary fatigue from a week that needs less volume.
What mistakes do lifters make with readiness scores?
The most common mistake is treating a single day's score as a command instead of one input among several. A close second is comparing scores across brands or between people as if they were on the same scale.
- Auto-skipping on any single red day. One low reading is often a late meal, a poor night's sleep, or a stressful day, not a signal that today's training is unsafe. Check the log for a pattern before cutting the session.
- Treating the score as a diagnosis. A readiness score reflects HRV, resting heart rate, and sleep trends; it isn't a test for injury risk or illness, and the JMIR review found it doesn't predict injury in a straightforward way.
- Comparing scores across brands or people. An 80 from one platform doesn't mean the same thing as an 80 from another, because each brand weights its inputs differently. Comparing your score with a training partner's is comparing two different, proprietary scales.
- Ignoring the RPE/RIR trend in favour of the score. The score reflects how your body looked this morning; the RPE or RIR you logged reflects how the actual set felt under load. When they disagree, the logged effort is the closer read on today's capacity.
- Not separating one bad night from a real pattern. Alcohol, a late meal, travel, or one hard session can tank tomorrow's score without meaning anything about the training block. The overtraining vs overreaching guide covers how to tell a short dip apart from a longer pattern that actually needs a change.
Frequently asked questions
Not automatically. A single low reading is often noise from poor sleep or a stressful day. Check whether your training log shows a matching pattern, rising RPE, falling reps, before changing the plan.
Sources and references
- DeBlauw et al., Journal of Functional Morphology and Kinesiology — "HRV-guided HIFT produced similar improvements in cardiovascular function, body composition, and fitness as predetermined HIFT, despite significantly fewer days at high intensity (mean difference -13.56 ± 0.83 days, p < 0.001)" (2022)
- Zucker, Journal of Medical Internet Research — "Research suggests these scores don't predict injury in a straightforward way, but that the data are useful for seeing trends over time" (2026)
- Maslakovic, Gadgets & Wearables — "A 2025 review of 14 wearable health scores reached a similar conclusion... few manufacturers had published evidence validating the final number" (2026)



