Selecting the lowest scores for intervention guarantees an apparent gain that owes nothing to the intervention.
Your worst screeners will improve whether or not you treat them
Pick the ten athletes who scored worst on the screen. Give them six weeks of extra work. Retest. They improve. Everyone concludes the extra work did it, and the programme is expanded to the whole squad next season.
The improvement was going to happen anyway, and here is the mechanism. Any score is the athlete's true ability plus whatever the measurement got wrong that day. When you select the lowest scores, you are selecting a group who are genuinely a bit below average and who also happened to be measured on a bad day: a stiff morning, a rushed assessor, a limb that had not warmed up. On the retest the bad luck is not repeated, because luck does not repeat. The true ability stays roughly where it was, the error draws a fresh number, and the group moves upward without a single thing having changed in the athlete.
The effect is larger the noisier the test, which makes it worst in precisely the screens that deserve the least trust. Handheld range of motion, one-trial balance scores, anything measured by eye. A noisy screen produces a bigger apparent treatment effect than a reliable one, and squads therefore accumulate the strongest belief in their least reliable tools.
Getting around it does not require a randomised trial, though a trial is obviously the clean answer. It requires a comparison group of athletes who scored badly and did not receive the intervention, even if that group exists only because the physiotherapist ran out of hours in March. Compare the treated poor scorers with the untreated poor scorers. Whatever gap remains is yours to claim. Usually it is much smaller than the raw before-and-after, and occasionally it is nothing.
The same logic runs in reverse, and nobody complains about that one. Your best screeners will look worse next time. Every time a coach worries that the top performers have gone backwards after a good preseason, this is the first thing to rule out and about the last thing anyone checks.
Selecting on an extreme and then measuring again is one of the oldest ways to fool yourself in any field that collects numbers.
Sport rediscovers it every few years, usually while announcing that the new protocol is working.

