There is no validated good flexibility score, and anyone who gives you a threshold without saying what it is a threshold of has skipped the only interesting part of the question. A score is a number produced by a particular set of tests, scaled by a particular set of choices, compared against a particular reference group. Change any of the three and "good" moves.
That does not make scoring useless. It makes the fine print load-bearing. So here is ours, in full, including the parts that are awkward for us.
The short answer
- A flexibility score is only meaningful relative to something: a norm table, a cut-off, or your own earlier reading.
- The norm tables that exist are mostly from convenience samples and are specific to one test protocol.
- Limber's Range Score is a composite of five self-tests, scaled from the app's own band anchors. It has no external validation and we are not going to imply it does.
- The comparison that works is you against you, weeks apart, same protocol.
What a good flexibility score is
A score answers a comparison question, so start by naming the comparison.
Against a population. "Am I more flexible than the average 40-year-old?" This needs normative data for your exact test, your age band and your sex. It exists for a handful of tests and nowhere else.
Against a threshold. "Do I have enough ankle range to squat to depth?" This is often more useful than a percentile, because it is tied to a task you actually care about.
Against yourself. "Am I better than I was in April?" This is the only one that needs no external data at all, and the only one that is reliably answerable at home.
Most people asking about a score want the first. Most people would be better served by the third.
The scale does not run 0 to 100
Limber's Range Score is presented on a 0–100 scale, and it is worth knowing that the whole of that scale is not reachable, because the score is a mean of band anchors and the bands do not start at zero or end at a hundred.
Work it through with the app's own numbers. Score the lowest band on every test in the battery and the composite comes out around 15. Score the highest band on every test — palms flat, 12 cm or more at the wall, fingers overlapping behind the back, thigh hanging below the bench, a squat you could rest in — and it comes out around 93. There is no arrangement of readings that produces 0, and none that produces 100.
The middle is not in the middle either. Take the middle band on everything — fingertips to the floor, 8 cm at the wall, hands less than a hand-width apart, thigh level with the bench, full squat with heels down — and the composite reads about 63, not 50. The band anchors were set to describe positions honestly, not to make a tidy scale, and honest description does not distribute itself evenly.
You can check every one of those figures yourself: the Range Score tool runs the same arithmetic the app runs, and the survey is carried in the address, so the numbers above are reproducible rather than asserted.
Two consequences. A score of 60 is not "60% flexible" — it is not a percentage of anything. And a score in the eighties is a genuinely open body, not a mediocre one, because the ceiling is 93.
Compared to whom?
Now the part that actually decides whether a score is good.
Range varies with things you cannot train. Limb proportions move a forward-fold reading before any tissue does: long femurs and a short torso cost you centimetres at the floor that have nothing to do with your hamstrings. Joint shape sets hard limits — Frank and colleagues, in Arthroscopy in 2015, found cam morphology, which mechanically blocks deep hip flexion, in roughly a third of people with no hip symptoms at all. Sex shifts most norm tables: women score higher on trunk-flexion tests across essentially every published sample. Age shifts them downward, though by less than people expect — see flexibility by age.
A percentile that ignores all of that is comparing you to a population you may not belong to. It is still information. It is not a verdict.
The published norms, and their limits
There are real norm tables, and it is worth naming them rather than gesturing at "studies".
Trunk flexion norms — the sit-and-reach family — appear in the American College of Sports Medicine's testing guidelines and in the Canadian fitness assessment protocols, broken down by age and sex. They are widely used and they carry two problems: the score depends on where the box's zero point sits, so numbers from different protocols are not interchangeable, and the underlying samples are convenience samples of people who volunteered for fitness testing.
Older-adult norms are better. Rikli and Jones published normative scores for community-dwelling adults aged 60 to 94 in the Journal of Aging and Physical Activity in 1999, from a large multi-site sample, including a chair sit-and-reach and a back-scratch reach. If you are over 60 these are the most defensible reference figures available.
Joint-angle norms — hip flexion 0–120°, ankle dorsiflexion 0–20° and the rest — come largely from small studies and committee consensus. Boone and Azen's much-cited 1979 figures in the Journal of Bone and Joint Surgery came from 109 male subjects.
Grade: moderate for older-adult field tests, weak for everything else. Useful as an order-of-magnitude check. Not a percentile you should take personally.
The case against the whole idea
Here is the citation that argues against our product, stated at full strength because it deserves to be.
James Nuzzo published a paper in Sports Medicine in 2020 titled The Case for Retiring Flexibility as a Major Component of Physical Fitness. Its argument is that flexibility earned its place among the classic fitness components by tradition rather than by evidence: the links between flexibility and health outcomes, injury prevention and physical function are weak, most of the supporting literature is cross-sectional, and time spent on stretching would generally be better spent on strength and cardiorespiratory work.
We think the paper is largely right about outcomes and we still think measuring is worth doing — for a narrower reason. Range matters when a task needs it: you cannot squat to depth without ankle and hip range, you cannot press overhead cleanly without shoulder and thoracic range, and you cannot get off the floor comfortably without both. Measuring tells you whether the range you want is arriving. It does not promise that a higher number is a healthier life.
The strongest published link between a floor-level movement and an outcome is not a flexibility test at all. Brito and colleagues, in the European Journal of Preventive Cardiology in 2014, found that the ability to sit down on the floor and get back up without support predicted all-cause mortality in over 2,000 adults. You can score the sitting-rising test here, on its authors' own rubric rather than Limber's anchors, which is why it does not enter the Range Score — and what it carries is an association across a cohort of volunteers at one Rio de Janeiro clinic, not a forecast for the person taking it. That test is at least as much strength, balance and coordination as flexibility — which is rather the point.
The only comparison that works
You, against you, with the protocol held still.
This is why Limber cuts a datum from your first battery and measures every later reading from it, rather than showing you a percentile. The datum has no validity problem, because it is not claiming to represent anyone. It is a line on your own ground, and the only question it answers is whether you have moved.
Practically:
- Run the five tests once, carefully, and write down the conditions as well as the numbers.
- Do not re-test for two to four weeks.
- Re-test under the same conditions, in the same order.
- Treat anything smaller than one band, or one or two centimetres, as noise — see how to measure flexibility at home for why the noise on these tests is larger than it looks.
- Compare the trend across three readings, not the difference between two.
If six weeks of consistent work has moved nothing, that is a finding. It is exactly the finding a percentile cannot give you.
What a low reading is not
It is not a diagnosis, and it is not a moral fact about you.
A low score reflects joint shape, limb proportions, age, warmth, habit and how much time you have spent at the ends of your range — in an order that varies by person and that no test can separate. Screening batteries built on movement quality have a poor record predicting who gets hurt: Moran and colleagues' 2017 review in the British Journal of Sports Medicine found Functional Movement Screen composite scores predict injury poorly, and Bahr argued in the same journal in 2016 that this is a general property of screening tests rather than a fault of one battery.
And a score has nothing to say about pain. If you have pain that radiates into a limb, numbness, weakness, pain after a fall or an injury, or back pain with fever, unexplained weight loss or a change in bladder or bowel control, that needs a clinician who can examine you. Nothing on this page diagnoses or treats anything.
The rest of the cluster sits under mobility, measured, and why am I so inflexible works through the causes in order of how much they matter.
Questions
What is an average flexibility score?
There is no single answer, because every score is specific to its test protocol. On Limber's 0–100 Range Score, taking the middle band on every test in the battery composes to about 63 — but that is the middle of our scale, not a population average, and no representative sample has ever been run against it.
Can I score 100?
No. The highest band on every test composes to about 93, because the score is a mean of band anchors and the top anchors sit at 90 and 95 rather than at 100. The floor works the same way: the lowest band on everything reads about 15. The usable scale is roughly 15 to 93.
Is a higher flexibility score always better?
Not automatically. Range matters when a task needs it, and beyond that the evidence for benefit is thin — the argument that flexibility should not count as a major fitness component at all was made in print in Sports Medicine in 2020. Very large passive range without the strength to control it is not obviously an improvement on moderate range you own.
How do I compare my score to other people?
Carefully, or not at all. Norm tables exist for trunk flexion and for older-adult field tests, but they are protocol-specific and drawn mostly from convenience samples, and they cannot adjust for the limb proportions and joint shapes that legitimately move your reading. Comparing your score to your own earlier score avoids every one of those problems.
What score should I aim for?
A task, not a number. "Palms flat on the floor" and "a squat I can rest in" are goals you can check. "Seventy-five" is a goal defined by a scale that only exists inside one app, and chasing it will send you after whichever test is easiest to move rather than whichever range you actually want.
My score went down. What does that mean?
Most often, that the conditions changed — you tested cold this time, or earlier in the day, or after a hard session. Before reading a drop as regression, check that the protocol matched. A change smaller than one band, or a couple of centimetres, is inside the noise of these tests and should not be read as anything.