Are You Getting Better Faster Than the AI Is?
AI can now build a solid strength program in under a minute, using an athlete's goals, training history, injuries, schedule, and equipment to produce something that looks professional and holds together. Writing the program is no longer what separates a coach from everyone else; knowing when the plan needs to change is.
At the highest level, coaches were never paid just to choose exercises, sets, and reps. They are paid to combine training data, medical information, nutrition, feedback from sport coaches, and conversations with the athlete, then make the right call and take responsibility for it. AI can help organize all of that information. It cannot fully understand the athlete, notice what's being left unsaid, or judge which piece of information matters most on a given day. That's where elite coaching still earns its price.
What the research says about AI Coaching
AI plans often have major flaws. The clearest data point comes from a 2024 study in the Journal of Sports Science and Medicine, where coaching experts rated ChatGPT generated training plans for runners. Plans built from thin prompts scored poorly. Plans built from detailed athlete history scored meaningfully better. Even the best version still fell short on individualization, and the researchers' recommendation was direct: do not deploy an AI generated plan without an expert coach reviewing it first.
Two things are worth noticing about that study. First, it was built on recreational runners chasing a single, relatively simple objective. Improving a running time is a far simpler problem than managing performance across an MLB season. Second, even under those low stakes conditions, the tool still needed a human check. Now raise the stakes to an MLB roster managing 40 players over a 162-game season. Over that time, there will be returns from reconstructive surgery, starters carrying months of accumulated fatigue into a playoff push, or players in contract years with every incentive to describe their bodies as feeling better than they do. The recreational research is the floor on how much oversight this requires, not the ceiling.
That does not mean AI lacks value. It means AI is better at providing information than making the final decision. A 2026 study found that ChatGPT outperformed common personal trainers at answering simple exercise questions. That result should not surprise anyone. A language model has read more exercise science than any individual coach ever will, and it can retrieve that information instantly. Recall was never the differentiator for a serious performance staff. Every credentialed coach already knows the correct rep range for hypertrophy. The differentiator was always what happens after the textbook answer meets a specific human body and the person living in it, and that part has not moved.
Agreement is not the same as accuracy
AI models trained on human feedback have a documented tendency to drift toward what the user already believes, even reversing a correct answer after some pushback. That tendency is close to irrelevant for a hobbyist choosing between two leg day options. However, It is a real liability for a player in a contract year insisting he is ready to play through something he has quietly hidden from the training staff, because the tool answering him has every incentive, built into its training, to tell him what he wants to hear. Part of a performance director's job is being the person in the building who is allowed to disagree with the athlete.
The danger is not simply that AI could give the wrong answer. It could give the athlete false confidence in the answer he already wanted. In an elite environment, that can mean training through an injury, returning before the body is ready, hiding symptoms, or ignoring signs of accumulated fatigue because the athlete found a response that justified continuing. The system does not know what the athlete left out, whether his description is accurate, or what pressures are shaping the question. A coach who knows the athlete can recognize those gaps, ask harder questions, and stop a decision before a manageable issue becomes a lost season.
More Data Still Needs Judgment
The honest counterargument is that elite programs already have more data than any personal observation could provide. GPS load metrics, force plate output, blood panels, sleep telemetry, and EMG all give coaches information the athlete may not be able or willing to communicate. That can make it seem like a fully instrumented performance department removes the need for a coach to notice what the athlete is not saying.
It does not. In studies against polysomnography, the clinical gold standard for measuring sleep, consumer-grade sensors have missed REM sleep by fifty to seventy percent. That is not a small error. Devices worn by the same athlete on the same night can also produce very different results.
Now put a dozen of those data streams on one dashboard during a playoff push. More sensors do not always create more clarity. They create more opportunities for the information to conflict, and someone still has to decide what matters. A drop in HRV the morning of a start could reflect fatigue, travel, illness, poor sleep, or something happening in the athlete’s personal life. The data can show that something changed. It still takes a person who knows the athlete to understand why it changed and what to do next.
The Program Is Just One Piece of the Puzzle
A player may train with the strength staff four times a week, but the program only controls the training prescription. It cannot make the athlete sleep enough, eat well, recover properly, communicate honestly, or manage stress. Those factors may be controllable, but they are not controlled by the sets, reps, and exercises written on the page. Add in travel, family stress, media pressure, and the physical cost of playing nearly every day, and it becomes clear that the program only governs one part of the athlete’s development. Everything else still affects how the athlete responds to the same session.
AI can adjust a workout when an athlete reports those things. The harder problem is recognizing what was not reported, what the athlete may be minimizing, and what nobody thought to ask about. A player may say he feels fine while moving differently, struggling with a familiar load, or showing changes in coordination, body language, and effort. That is where the coaching eye matters. AI can respond to the information it receives. A coach can notice the information nobody knew to provide.
That is why a well-written program can still be the wrong program on a given day. The athlete who arrives in the weight room is not the same athlete who existed when the plan was written. Tendon health, glycogen, soreness, sleep, motivation, and psychological readiness are all changing on different timelines. They also affect one another. A delayed flight may not matter much by itself, but it may matter when combined with weeks of accumulated fatigue, poor nutrition, reduced sleep, and the early signs of an issue the athlete has not fully acknowledged.
This is the central idea behind Why Athletic Progress Slows Down: Nested Systems and Flexible Individualization. An athlete is not one system that receives a workload and returns a predictable result. The body, mind, environment, habits, and schedule are constantly interacting. The coach’s job is not simply to deliver the session on the page. It is to understand the condition the athlete arrived in, decide what matters most, and determine what the athlete can benefit from that day, what needs to change, and what should not happen at all.
The program provides direction. The coach sees the full picture and decides whether that program creates adaptation or adds to the problem.
The Bottom Line
The program was never the product. At the elite level, the real value is the integration: the ability to combine medical input, sport coaching, nutrition, sensor data, direct observation, and conversations with the athlete, then turn all of it into the right decision.
That does not mean programming no longer matters. The coach still has to understand training deeply enough to build the right plan, recognize when it is no longer right, and adjust it without losing sight of the larger objective. AI raises the standard. It removes the ability to hide behind sets, reps, and a professional-looking spreadsheet.
The question is not whether AI will keep getting better. It will. The question is whether coaches are improving just as quickly in the areas AI cannot replace: judgment, observation, communication, integration, and accountability. At the level where a wrong call can cost a season or a career, being good at programming is only the starting point. The value lies in knowing what to do when the program and the athlete no longer match.
FAQ
Should an elite performance program use AI, or treat it as a liability?
Use it, deliberately. AI is fast at organizing information and has no ability to understand a person. Let it speed up the analysis phase. Don't let it make the call. The liability isn't the tool. It's letting its output pass into a decision that affects an athlete's season without a person checking it first.
Does the plan-quality research even apply to professional environments?
Not directly, and that's the point. The published research is on recreational athletes facing low stakes. Nobody has published a study on AI-generated return-to-play protocols for a nine figure contract, because no serious organization would run that experiment. The absence of elite-level research isn't reassuring. It means the risk hasn't been measured, not that it isn't there.
Does more sensor data lower the risk of leaning on AI, or raise it?
Right now, it raises it. More inputs mean more room for conflicting signals, and a system asked to reconcile ten data streams will produce a confident answer whether or not the underlying data agrees with itself. A person who knows the athlete can tell you which sensor to trust today. AI will average them.
Is the sycophancy problem worse for professional athletes specifically?
Yes. A recreational lifter has little to gain by exaggerating how good he feels to an app. A player chasing a return to play date, negotiating a contract, or trying to stay in a lineup during a playoff push has real incentive to describe his body as better than it is. A system trained to agree with what it's told will not catch that. A coach who has watched him warm up for three seasons will.
What should an organization actually evaluate when deciding whether performance staff is worth the investment?
Ask what they do when the plan and the athlete disagree. Ask for a specific case where they changed course because of something they noticed, not something the athlete reported. A vague answer, or one that sounds like a description of a spreadsheet, is the tell.
References
D’hoe, B., Kirk, D., Boone, J., & Colosio, A. (2026). ChatGPT outperforms personal trainers in answering common exercise training questions. Journal of Sports Science and Medicine, 25, 235–261.https://doi.org/10.52082/jssm.2026.235
Düking, P., Sperlich, B., Voigt, L., Van Hooren, B., Zanini, M., & Zinner, C. (2024). ChatGPT generated training plans for runners are not rated optimal by coaching experts, but increase in quality with additional input information. Journal of Sports Science and Medicine, 23, 56–72.https://doi.org/10.52082/jssm.2024.56
Kainec, K. A., Caccavaro, J., Barnes, M., Hoff, C., Berlin, A., & Spencer, R. M. C. (2024). Evaluating accuracy in five commercial sleep-tracking devices compared to research-grade actigraphy and polysomnography. Sensors, 24(2), 635.https://doi.org/10.3390/s24020635
Kaiserman, B. (2026, June 17). Why athletic progress slows down: Nested systems and flexible individualization. NewmanHP.https://www.newmanhp.com/blog/why-athletic-progress-slows-down-nested-systems-and-flexible-individualization
Sharma, M., Tong, M., Korbak, T., Duvenaud, D., Askell, A., Bowman, S. R., Cheng, N., Durmus, E., Hatfield-Dodds, Z., Johnston, S. R., Kravec, S., Maxwell, T., McCandlish, S., Ndousse, K., Rausch, O., Schiefer, N., Yan, D., Zhang, M., & Perez, E. (2023). Towards understanding sycophancy in language models [Preprint]. arXiv.https://doi.org/10.48550/arXiv.2310.13548