Thresholds discovered by searching a single dataset are fitted to its accidents and rarely survive the second squad.
A screening cut-off invented on one squad travels badly
Somebody has a season of data, a list of who got injured, and an afternoon. They try the marker at every plausible cut-off until one of them separates the injured from the healthy. That value goes into a paper, then into a slide, then into a policy at a club that has never seen the original squad.
The search is the problem. If you test twenty possible cut-offs on one dataset and keep whichever performs best, you have not discovered a biological boundary, you have found the number that best fits the particular accidents of that group: their fixture list, their surface, their coach's preferences, the four unlucky hamstrings in November. The selected threshold carries all of that with it. Applied to a different squad it usually degrades, sometimes to nothing, and the degradation is not a sign that the second squad is unusual. It is the normal fate of a value chosen by searching.
What makes this hard to see from outside is that the original analysis looks careful. There are confidence intervals. There is a curve. But the interval is computed as though the cut-off had been specified in advance, which it was not, and that assumption is doing quiet work in every number reported afterwards.
There is a second layer of borrowing that goes unremarked. Most of these thresholds come from squads that differ from yours in ways that plainly matter: professional men in one sport, an academy, a single national programme. Age changes the marker. Training history changes it. Sex changes several of them substantially. A cut-off derived on twenty-two senior professionals applied to a women's academy is not conservative or aggressive, it is simply unmoored, and calling it evidence-based because a study exists is a category mistake.
The defence is boring and effective. Ask three questions of any threshold before adopting it: how many candidate values were tried, was the final one tested on data not used to find it, and who was in the sample. If the answer to the second is no, treat the number as a hypothesis and start collecting your own.
Your own data will be small and messy and it will take two seasons before it says anything.
That is still better than a borrowed line whose only real credential is that somewhere, once, on a different group of people, it happened to fit.

