Three days ago I would have said my training app was nearly done. It still is nearly done.
I build GymBuddy: Guided Strength with AI coding agents. The first version works and is public. The version I want to put in the Play Store has more in it, and for the last three days I have been fixing its programme generator with agents, one step at a time. Every step came with a clear explanation. Every explanation sounded reasonable. Somewhere on day two I noticed that I could not say whether we were getting closer.

I drew this for myself. This is how I actually feel about my vibe-coding experience. I have no proof, it is a gut feeling. The purple line is how a developer who understands what is under the hood gets from A to B. The blue line is how I get there with agents. Both lines arrive. The blue one takes loops.
The blue line does reach B. The trouble is that the person on it cannot tell a detour from a circle. When the agent says “this test fails, so I will change the validator”, I cannot check whether the test was right. I can only decide whether to say yes.
Where the three days went
Very little of it went on writing code, full grandle runs occupied much time in this process until I asked to split it and run only necesary parts. In general the time went into three things.
Proving that each change did what it said and nothing else. For instance, a small change to a weight rule or to an exercise list can quietly affect many programme variants. And the generator is deterministic on purpose (an AI layer is planned for later, but not for the part that decides): same user, same answer, every time. Every fix had to keep that, and every fix had to be checked for safety, because a wrong weight in a training app is not a cosmetic bug (for example, the algorythm assigned Kettler Swing with a 32 kg kettler to a beginner woman who weighs 60 kg 🙈 while all other assigments in her training program were fine and reasonable …why this exercise??? why that weight???)
Then we had separation real regressions from old debts. Some failing checks were bugs we had just introduced. Others were things that had always been wrong and nobody had measured before. Fixing the second kind as if it were the first is how a loop starts. You repair something that was never broken, and the repair has its own side effects 🙈.
And correcting the diagnostic tests themselves. Several times the test was wider than the rule it was supposed to check, or simply wrong. A test that is too broad fails for good changes. A test that is wrong passes bad ones. Both send you around one more loop.
The actual fixes were prosaic. The app had to suggest a weight that exists for the equipment in front of you: a barbell, a pair of dumbbells, a kettlebell, a cable stack or a plate-loaded machine all step differently. It had to add a new exercise only where it genuinely fits the person. One muscle group had dropped out of a full-body plan for one type of user. Two exercises needed permission to go down to 1 kg, and only those two. And a set of impossible states that could still reach the “generate programme” screen was moved one screen back, into a check that stops the user there instead of producing a plan that looks fine and is wrong.
In parallel we fixed the parts a user actually sees: the weekly report, exercise descriptions, the history screens, coming back to an exercise you have just edited, and what the Back button does. At the end we built test versions of the app and I went through them by hand, in Russian and in English.
Which of my judgements are real 🤔
That is the question I had to answer for myself, and I think it is the honest question for anyone building this way.
I cannot judge the code. When an agent tells me why a change is needed, I am trusting it. That is a real limit and I am not going to pretend otherwise.
I can judge two other things.
First, the rules. The part of the app that picks exercises and weights runs on rules I wrote. So I can take any generated plan and check it against them: is this load progression allowed at this level, does this weight exist for this equipment, is every muscle group covered. I check what comes out. I cannot check the code that made it. Without written rules I would have nothing to check at all.
Second, the limits of a change. Before each fix I said what the agent was not allowed to touch: no general-purpose solver, no automatic swap of the programme structure, no change to the templates or the seed, no loosening of the load rules for everyone in order to fix one case. The fix had to live inside those limits. Every time I dropped a limit, the fix got bigger and the loop got longer.
There is a third thing, and it is the only one where I need no trust at all. Whether the weekly report makes sense to a person. Whether the history screen answers the question someone opened it with. Whether Back takes you where you expect. Nobody has to explain the code to me for that. I open the app and I know.
Rules to check the output against, limits I can hold, and the plain question of whether the thing makes sense. That is all I bring to the blue line. It turned out to be enough, but only just.
The same pattern on the job boards
While this was going on, I started keeping a small log of job posts where someone had built a product with AI tools and was now asking for help. Six posts in one week, on one platform. Nobody was asking for new features. Every ask was some version of “I have a thing, is it the right thing, is it ready”.
What stood out is that almost none of them could say what done looks like. One founder posted the same brief for two different apps, asking for an honest read on where each one stands. One company with hundreds of hires behind it described its game as “80 to 85 percent complete”. I do not know how to measure 85 percent of a codebase you cannot read.
I recognised myself in every one of them. Three days ago I would also have said 85 percent 😅.
when you are starting the blue line
- Write the rules down before you generate anything. Not the code, the rules: what a correct result must always have and must never have. That is the one thing you will be able to check later.
- Decide what a fix is not allowed to touch before you ask for the fix. Say it to the agent in plain words.
- And when a test fails, ask first whether it was ever passing. A large number of the failing tests I met were old ones, and not every one of them pointed at a real problem.
The app is still nearly done 🤞. The difference is that now I can say what done means. Whether I am right about that, the Play Store will tell me 😅.


Leave a Reply