Goals

How to track flexibility progress

Fix the conditions, date every reading, leave three weeks between them, and wait for a third point before calling anything a trend. Impressions fail here in a documented way: how stiff a back feels does not match how stiff it measures. Know how coarse your test is as well, because a banded test cannot show a small change.

29 Jul 2026 · 8 min read

Almost everybody who stretches for a year has no idea whether it worked. They have an impression, and impressions in this area are unreliable in a specific, documented way: Stanton and colleagues, in Scientific Reports in 2017, found that how stiff a back felt did not match how stiff it measured.

How to track flexibility progress is therefore not a bookkeeping question. It is the difference between knowing and guessing, and the method is short enough to fit on a card.

The short answer

  • Fix the conditions. Same warmth, same hour, same protocol, same order. Non-negotiable.
  • Three weeks minimum between readings. Daily testing measures the weather.
  • Write the date on it. A reading without a date compares to nothing.
  • Three points, not two. Two readings is a line drawn through noise.
  • Know how coarse your instrument is, because most home tests cannot show a small change.

How to track flexibility progress

Six rules. All of them earn their place.

1 · Same protocol, every time. The same setup, the same knee position, the same cue, the same order of tests. A rough measurement taken identically twice tells you more than a precise one taken two different ways.

2 · Same conditions. Always warm or always cold, and roughly the same hour of the day. This is the rule people break, and it invalidates everything downstream. DePino and colleagues, in the Journal of Athletic Training in 2000, stretched hamstrings and found the extra range back at baseline within about three minutes — so a reading taken at the end of a session is not your range, it is your range plus a warm-up. Test warm one month and cold the next and you will manufacture a loss that never happened.

3 · Both sides, recorded separately. A single number for two legs hides everything. Only per-side tests can show asymmetry at all.

4 · The date, written down. Every time. This is the whole basis of comparison.

5 · Weeks between readings, not days. Range moves slower than the noise does. Three weeks is a sensible minimum; Limber's default re-test interval is 21 days, adjustable to 14 or 28.

6 · Read three readings, not two. A trend needs three points. Two is a line through noise, and the noise on home tests is larger than it looks.

Know how coarse your instrument is

This is the part almost nobody covers, and it belongs on our own page because it is a limitation of our own battery.

A banded test cannot express a small change. Here is what one band is worth on each test in Limber's battery, on the app's 0–100 scale:

Test Bands One band is worth
Toe touch 5 15–25 points
Shoulder reach 5 20–25 points
Deep squat 4 20–30 points
Hip flexor 3 30–35 points

So on the hip flexor test, nothing smaller than a 30-point change can be recorded at all. You can improve genuinely for two months and read the same band throughout, then jump 30 points in one session. That is the instrument, not your progress, and treating a band as a precise reading is the commonest way to misread your own record.

The knee-to-wall test is the exception and the reason it exists: it is measured in centimetres and interpolated between the app's own anchors, so it can express a small change. One centimetre is worth 5 to 10 points — ten through the middle of the scale (4 cm reads 35, 6 cm reads 55) and five at the top (10 cm reads 85, 11 cm reads 90).

The same precision punishes carelessness. Limber calls a gap of 12 points or more between your sides an imbalance, so a single centimetre of sloppy foot placement manufactures most of an imbalance out of nothing. Be fanatical about foot position on that one, and casual about band boundaries on the others.

If you want to detect change rather than infer it, lean on the finest scale you have.

What the composite reading does

The Range Score rolls the tests into one number by the same arithmetic the app runs: each test scores 0–100, each area is a weighted mean of the tests feeding it, each of the five score groups is the mean of its areas that have a reading, and the Range Score is the mean of the groups that have a reading.

Three things about it are worth knowing before you track against it.

A test you have not taken is dropped, not counted as zero. An area with nothing in it is left out of the mean entirely, the station is drawn as an open mark with dashed runs either side, and the count of surveyed stations is printed beside the score every time it appears. Nothing is interpolated from its neighbours, because their mean would look exactly like a reading.

The ceiling is four of five stations. The Spine group is fed only by a seated thoracic rotation test that is not built — in the app either — so four of five is the honest maximum for this battery. The shoulders area is designed as shoulder reach 60% plus a wall angel 40%, and the wall angel is not built either, so the page prints the shortfall rather than quietly renormalising.

The scale runs about 15 to 93 in practice, not 0 to 100. The lowest band on every test composes to 15; the highest to 93; the middle band on everything to 63, not 50. The anchors describe positions honestly and honest description does not distribute evenly. A target of 80 is a much bigger ask than it sounds.

And one thing it is not: a percentile. No population has ever been measured against this scale. It exists so the distance between two of your own readings means something, not so you can rank against strangers.

What a change has to beat

Before calling a difference real, it has to clear three things.

The noise in the test. Gajdosik and Bohannon reviewed goniometric measurement in Physical Therapy in 1987 and found that the same examiner repeating a measurement is reasonably consistent, while two different examiners agree considerably less. At home you are always the same examiner, which helps — and you are an untrained one, which does not.

The band width. Covered above. On a three-band test, only a 30-point change exists.

The conditions. Which is why rule two exists.

If a difference clears all three and repeats on a second measuring session, it is a finding. If it does not, it is a reading you should take again.

The projection question

The obvious next step is to draw a line through your readings and extend it. Do so cautiously, and know what a projection is worth.

The projection function in Limber's engine needs two readings to see a drift; given one, it returns the reading itself. A rising line drawn through a single measurement would be a promise rather than a projection. So the design draws the datum, the reading sitting on it, and a dashed run forward to open marks at three and six weeks — the app's grammar for unsurveyed ground everywhere else — under the line Nothing is drawn past the datum: that ground is unmeasured. Where it does project from two real readings, it damps the drift by 0.7, caps it at four points a week and never looks more than eight weeks ahead.

Straight about the build: the re-test ritual that would use that projection is not finished. The tests, the score and the record are what exist today, and this site does not claim the app forecasts anything.

The same caution applies to a line you draw yourself. Progress in range is front-loaded, so a rate measured over six weeks will not persist for six months, and extending it to a date produces a fantasy with arithmetic on it.

What tracking will not tell you

Whether you are at risk of injury. Moran and colleagues' 2017 systematic review in the British Journal of Sports Medicine found Functional Movement Screen composite scores predict injury poorly, and Bahr argued in the same journal in 2016 that screening tests generally lack the discriminative power to identify individuals who will get hurt.

Whether a gap between your sides matters. There is no trial showing that reducing a left–right range difference reduces injuries or improves performance.

What is causing pain. Stiffness and pain are different problems, and a range reading does not diagnose anything. Pain that persists, radiates down a limb, comes with numbness or weakness, or followed a fall is a matter for a clinician who can examine you. Nothing here diagnoses anything and nothing here is treatment.

Start the record

Take the battery once, under fixed conditions, and write the date on it. The Range Score composes what you have and states what you have not measured. Both are free, and on the web nothing is stored: the survey lives in the address bar, so a reading you want to keep is a link you save. In the app the first reading is kept as a datum and every later one is measured from it.

Then leave it alone for three weeks. The flexibility plateau covers what to do if the number does not move, realistic flexibility goals and timelines covers what to expect, and the cluster is goals.

Questions

How often should I measure my flexibility?

Every three weeks at the shortest. Range drifts with warmth and time of day by more than a few weeks of training typically changes it, so measuring daily or weekly records the conditions rather than your progress. Twenty-one days is Limber's default re-test interval, adjustable to 14 or 28.

What is the best way to track flexibility at home?

A fixed protocol, fixed conditions, both sides recorded separately, a date on every reading, and three weeks between them. The specific tests matter less than the repeatability — a rough measurement taken identically twice tells you more than a precise one taken two different ways.

Why does my flexibility seem worse some days?

Because it is. Range is lower cold, lower first thing in the morning, and lower after a hard session, and those swings are larger than several weeks of training usually produces. That is why matched conditions is the rule that invalidates everything else when it is broken.

How much change counts as real improvement?

Enough to clear the noise in the test, the width of the band, and any difference in conditions — and it should repeat on a second measuring session. On a three-band test the smallest change that exists at all is 30 points, so anything finer than that is invisible rather than absent.

Should I use a photo to track my progress?

Photographs are useful as a record of position and poor as a measurement, because camera angle, distance and lens change the apparent geometry more than several weeks of training changes your body. If you use them, fix the camera position, the distance and the framing, and treat them as a supplement to a number rather than a substitute.

Can I compare my score to other people?

Not meaningfully on this scale — no population has ever been measured against it, and it is not a percentile. It was built so the distance between two of your own readings means something. Published joint-range norms exist for some measurements and mostly come from convenience samples that are smaller than they look.

Take the reading
Nearby in this cluster
Sources
  1. Stanton TR, Moseley GL, Wong AYL, Kawchuk GN. Feeling stiffness in the back: a protective perceptual inference in chronic back pain. Scientific Reports 2017;7(1):9681. doi:10.1038/s41598-017-09429-1
  2. Depino GM, Webright WG, Arnold BL. Duration of maintained hamstring flexibility after cessation of an acute static stretching protocol. Journal of Athletic Training 2000;35(1):56–9. PMID 16558609
  3. Gajdosik RL, Bohannon RW. Clinical measurement of range of motion. Review of goniometry emphasizing reliability and validity. Physical Therapy 1987;67(12):1867–72. doi:10.1093/ptj/67.12.1867

Each address was followed to the record it names before it was written down. Where a paper could not be resolved that way, the page names it in the prose and leaves it unlinked rather than guessing at an address.

← All of Goals