Three dead ends before the answer: searching 700,000 days of wearable data for a PEM signal
The boom-bust cycle wasn't there. The PEM group looked like it recovered better. Self-report and physiology agreed at the level of a coin flip. What finally worked was asking a different question.

Short answer
Originally published on Founder And The City, Jane's newsletter about building in health tech. This version is written for readers who live with the thing being measured, rather than for the people who build the tools.
Post-exertional malaise is the defining feature of long COVID, ME/CFS, fibromyalgia and POTS: ordinary effort — a short walk, a phone call, a shower — triggers a delayed crash that arrives 24 to 72 hours later. Standard lab tests can't see it, and neither, it turns out, can the obvious wearable measurements.
We had 700,000 days of data from more than two thousand people, and not one classic PEM pattern worked. The boom-bust cycle the literature describes wasn't in the data at all. When we matched people carefully, the group reporting PEM appeared to recover better than controls. Self-reported crash severity agreed with physiology at roughly the level of a coin flip.
The answer only appeared after we stopped asking how much a person recovers and started asking how predictably they recover. Not the level — the stability. Three signals separated the groups at that point: how consistent recovery is from day to day, how quickly the nervous system shifts gears after effort, and how much cardiac work each unit of walking costs.
And if you live with this and have been told every test is normal — you are not imagining it, and you did not bring it on yourself by being careless. Post-exertional malaise is a recognised diagnostic feature, not a mood or a fitness problem. It affects tens of millions of people worldwide. The reason the tests come back clean is that the standard tests were never designed to look for it, and — as this research shows — even the obvious wearable measurements look straight past it unless you ask them a different question. The wrong question was being asked of the data, not of you.
What PEM actually is, and why it hides from tests
Post-exertional malaise is not tiredness. The hallmark is disproportion and delay: ordinary exertion produces a worsening of symptoms that typically begins many hours later and can last days or weeks. Rest afterwards doesn't prevent it. Effort of any kind counts — physical, cognitive, emotional, sensory — which is why a difficult conversation can cost as much as a walk.
Tens of millions of people worldwide live with conditions where this is the central feature. And for almost all of them, the standard clinical workup comes back unremarkable. That gap — a body that is demonstrably unreliable and a test result that says nothing is wrong — is the thing we wanted to close.
If you're new to the concept, start with our full guide to post-exertional malaise. If you've ever crashed after a day you barely moved, why step-count pacing fails covers that specific trap. This article is about something narrower: what happened when we went looking for the signal, and how many wrong turns it took.
Dead end one: the boom-bust cycle isn't in the data
The clinical literature describes a recognizable cycle. A patient has a good day, overdoes it, crashes hard, spends days recovering, then repeats. Boom-bust. It's in the textbooks, the guidelines and the patient forums, and it looked like the obvious thing to find in longitudinal data.
We built a detector for it — activity autocorrelation, crash probability modelling, the whole apparatus — and ran it across 700,000 user-days.
Zero signal. Not weak. Not noisy-but-present. The autocorrelation of daily activity in people reporting PEM was statistically indistinguishable from people who didn't, with an effect size of d = 0.004. The probability of a crash following a high-activity day was 15.8% in the PEM group and 16.9% in the non-PEM group. If anything, the comparison ran the wrong way — which makes sense the moment you think about it, because the non-PEM group actually has high-activity days.
Then the explanation arrived, and it reframed the whole project.
Boom-bust describes early illness. By the time someone has lived with PEM for a year or two, they have already learned to pace. They don't boom and bust any more; they live on a carefully managed plateau. The cycle wasn't missing because we measured badly. It was missing because these people outsmarted it.
Pacing is invisible adaptation. If you go looking for the wound and the person has already grown scar tissue that looks like ordinary skin, you will conclude there was never a wound. You are not measuring the disease. You are measuring the result of someone coping with it.
Dead end two: the group with PEM appeared to recover better
After boom-bust failed, we went methodological. Matched pairs: people reporting PEM paired with controls by age, sex and activity level. Same step counts, same demographics, comparable observation windows. Then we compared recovery — how quickly and completely heart rate returned to baseline after exertion.
The result was d = −0.12, statistically significant at p = 0.001, and pointing in the wrong direction. The PEM group recovered better.
There is a particular silence in your own head when a result contradicts the reason the project exists. The statistics were clean. The labels were correct. The conclusion was wrong anyway — and not because of a bug.
Matching on step count controls for how much someone walks. It does not control for how they walk. A person with PEM who takes 5,000 steps does it slowly, with breaks, at low intensity. A person without PEM who takes 5,000 steps does it at a normal clip, with bursts, at moderate intensity. Same step count, very different cardiac load. Less load means less to recover from. Recovery looked better because the input was smaller.
We weren't measuring recovery capacity. We were measuring activity intensity again, through a more sophisticated back door.
This one is worth flagging beyond our own work. If a study compares a PEM group to controls and doesn't stratify by actual cardiovascular demand — not just step count — its effect sizes deserve a second look. The studies that do stratify tend to find much smaller effects, or paradoxical ones.
Dead end three: self-report and physiology don't agree
Running out of obvious approaches, we tried the one that felt most scientifically virtuous: ask the people living with it.
For three months, 47 users completed surveys about crash severity, symptom intensity and their own assessment of their condition. We matched those against physiological data from the same periods, expecting self-report to serve as ground truth. If someone says they're crashing, the wearable data should show it.
Agreement between the two, measured by intraclass correlation, ranged from 0.019 to 0.356 across six features. Below 0.4 is conventionally poor agreement. Below 0.2 is noise. The best feature barely cleared the noise threshold.
This is not because people are unreliable reporters. It is arguably the definition of PEM. The bill arrives 24 to 72 hours after the spending. Self-assessment captures the arrival of the bill; physiology captures the expense. Both are real. They are measuring different things at different times.
The lesson we took: self-report is a complement, not a ground truth. Use one time-shifted signal as the label for another and you get noise — not because either is wrong, but because they answer different questions about the same event.
The pivot: not how much, but how predictably
Three dead ends in, we did the thing we should have done at the start. We stopped looking for what the literature said should be there, and looked at what was actually in front of us.
Every dead end had a second, quieter message. The averages didn't separate the groups. Average recovery, average heart rate, average HRV — after controlling for activity, they looked the same. But something was obviously different, because these people were living fundamentally different lives.
The difference was in the consistency.
Two time series of recovery quality can have identical means and behave completely differently: one steady, one oscillating. We had been computing means for months without plotting day-to-day stability, because you don't look for a difference in variance when you're hunting a difference in levels.
Instead of asking how much does this person recover, we asked how predictably. Three feature groups lit up, all with large effect sizes.
Recovery consistency. A body without PEM recovers roughly the same way every day. In the self-reported PEM group, the day-to-day spread was about twice as large — at the same average. Same number, different system.
Autonomic switching speed. How quickly the nervous system shifts gears after effort. This is a well-established cardiovascular measure, and in the PEM group the switch was slower. It was the most robust single pattern we found: it held across every activity quartile, from the most sedentary to the most active. It is not an artifact of walking less.
Cardiac cost of movement. How much cardiac work each unit of walking costs. For the same walk, a person in the self-reported PEM group's heart worked roughly 50% harder than an activity-matched person without PEM.
Three features, three different physiological systems, all pointing the same way. A minimal three-feature model slightly outperformed one with more than twenty features — adding more was adding noise, not information.
What held up across 45,000 users
Once the three features crystallized, we applied the algorithm retrospectively to anonymized historical data from more than 45,000 Welltory users, controlling for activity level at every step.
Dose-response. Splitting users by self-reported crash frequency, every feature increased monotonically across severity (p < 0.001). The pattern doesn't just exist; it grades with how bad the lived experience is.
Cascade. A bad recovery night predicts another bad night at more than double the baseline risk, and the elevation is still roughly twofold on day two. PEM isn't a single event — it's a self-reinforcing spiral. Many people describe this intuitively; the data is consistent with that description.
A separate cohort. With consent, we applied the algorithm to anonymized data from Welltory's community of users with self-reported energy-limiting conditions — a group assembled entirely by self-reported condition, with the algorithm playing no role in who joined. The rate of elevated-risk scores came out more than double the general-population baseline.
What this is not
This part matters more than the findings, so we'll be blunt about it.
It is research, not a product. The algorithm described here is a research artifact. It is not running in the Welltory app, and Welltory does not detect or diagnose post-exertional malaise, ME/CFS or long COVID.
It is observational, on a self-reported cohort. Nobody in this analysis had a clinically confirmed diagnosis recorded in the dataset. Classification performance reached an AUC of about 0.77 — meaningful for research signal-detection, and well short of what clinical use would require.
It misses more than half. At its strict threshold, the algorithm fails to flag over half of self-reported PEM cases. We chose specificity over sensitivity: better to miss someone than to falsely tell them something is wrong. The consequence is that absence of a signal is not a clean bill of health. It means the algorithm didn't find enough signal in the available data.
Menstrual cycle effects are almost certainly a first-class variable we haven't incorporated yet. Metrics shift measurably across the cycle, and that work isn't done.
Replication in clinically confirmed cohorts is the essential next step.
What you can take from this if you live with it
You can't run this algorithm on yourself, and we're not going to pretend otherwise. But three ideas from the work are usable today.
Your average is the least informative number you have. Two people with the same average recovery can be in entirely different states. If you track anything, look at the spread across a fortnight, not the mean. A steady mediocre number is a different situation from an unstable good one.
The same walk doesn't cost everyone the same. If moderate activity leaves your heart working far harder than it seems to for other people, that is a physiological observation, not a fitness failure. It's the single most useful sentence in this research for anyone who has been told to just do more.
The day you feel it is rarely the day that caused it. Because the delay runs 24 to 72 hours, looking at yesterday is usually looking at the wrong day. If you keep any record at all, keep it in a form you can scroll backwards through — the answer is normally two or three days upstream. Our guide to pacing without a step counter goes into how to do that practically.
How to bring this up with your doctor — and what to ask for
The most common experience people describe is being told their results are normal and leaving with nothing. A few things make that conversation go better.
Bring a record, not a conclusion. Two or three weeks of dated notes — what you did, and how you felt one, two and three days later — is far more persuasive than any app screenshot. The delay is the diagnostic feature, and a record is the only way to show it. Don't bring our research; it says nothing about you specifically.
Use the word "post-exertional malaise" explicitly, and describe it as disproportionate and delayed. Saying "I get tired" invites advice to exercise more. Saying "ordinary effort makes me worse a day or two later, and it lasts for days" describes a recognised clinical feature.
Ask what should be ruled out first. PEM is not a diagnosis of exclusion in itself, but several treatable conditions look similar: thyroid dysfunction, iron deficiency including ferritin rather than haemoglobin alone, coeliac disease, sleep apnoea, and — if you faint or your heart races on standing — orthostatic intolerance and POTS, which has its own simple in-clinic test.
If the delay is dismissed, ask for it to be recorded. "Symptoms worsen 24 to 72 hours after exertion" in your notes is a fact that travels with you to the next appointment and to any specialist referral.
If graded exercise is proposed, ask how it will be adjusted if you get worse. Current UK guidance no longer recommends fixed incremental exercise programmes for ME/CFS, and a plan without a stopping rule is a plan worth questioning.


Start understanding your body
Learn what affects your energy, stress, sleep, and daily state. Get the app.
How Welltory fits — and what it can't do
Welltory measures physiological signals: heart rate, heart rate variability, sleep, steps and stress load. It shows you your own baseline over time and flags when a day sits outside your usual range. That is genuinely useful for spotting patterns you can't hold in memory, and for bringing something concrete to a doctor's appointment.
It does not diagnose ME/CFS, long COVID or post-exertional malaise. It cannot measure cognitive, emotional or sensory exertion — which, for this population, is often the larger part of the load. And the research described above is not in the product.
What we'd honestly claim: if your body has become unreliable and every test is normal, having a record of your own physiological baseline is better than having nothing, and it is something you can look backwards through when the delay makes memory useless.
Read more from Jane
Jane writes about products, evidence and the health industry in Founder And The City, the newsletter where the original, longer version of this research write-up first appeared — including the parts about what it felt like to watch three months of work produce a result pointing the wrong way.
She also posted about it on LinkedIn — "sometimes to find the truth in 700,000 days of wearable data, you have to ignore the textbooks completely" — where the discussion continued in the comments.


Discounts for blog readers: up to 36% off
See what affects your energy, stress, sleep, and daily state with Welltory
This article is for educational purposes only and is not a substitute for medical advice, diagnosis, or treatment from a qualified clinician. The research described is observational, conducted on a self-reported cohort, and is not a validated diagnostic. Welltory does not detect, diagnose or predict post-exertional malaise, ME/CFS or long COVID.
Was this helpful?
Ask AI for a summary of page
Written by Jane Smorodnikova
The founder and CEO of Welltory. A recognized tech leader with two Master's degrees and experience at MIT, she has scaled Welltory to over 17 million users.
Written by Tatsiana Yashyna
References
- Institute of Medicine. Beyond Myalgic Encephalomyelitis/Chronic Fatigue Syndrome: Redefining an Illness. National Academies Press, 2015. — the report that established post-exertional malaise as a core diagnostic criterion.
- National Institute for Health and Care Excellence. Myalgic encephalomyelitis (or encephalopathy)/chronic fatigue syndrome: diagnosis and management. NICE guideline NG206, 2021. — current guidance on PEM and on pacing.
- Cole CR, Blackstone EH, Pashkow FJ, Snader CE, Lauer MS. Heart-rate recovery immediately after exercise as a predictor of mortality. New England Journal of Medicine, 1999. — the basis for heart-rate recovery as a measure of autonomic switching speed.
- Chu L, Valencia IJ, Garvert DW, Montoya JG. Deconstructing post-exertional malaise in myalgic encephalomyelitis/chronic fatigue syndrome. PLOS ONE, 2018. — survey evidence that crashes follow cognitive and emotional exertion, not physical exertion alone.
- Welltory Research. What Recovery Pattern Data Can Show About Post-Exertional Malaise — the full methodology and visual evidence behind the analysis described here.

