A senior engineer wrote Cal Newport an update. He had been generating features with Claude Code, and two of those features crashed the product. His boss told him that if it happened one more time, he would be fired, so he went back to writing most of the code by hand.
The firing threat is the receipt. The code had looked fine. It compiled, and it read like something a competent person would ship. Then it took the product down twice.
Newport has been collecting notes like this from people who actually ship software. Last winter, more than 300 developers answered a survey for him about how they were using these tools, yet the story that stuck was not a benchmark. It was a person who still had to go to work the next morning.
If you cannot stand behind a change, you did not make it. You rented a paragraph that sounded like work.
The obvious reply is that he prompted badly, or that the next model will stop crashing products, or that this is just how junior engineers used to learn, except faster. Those are fair points, although they skip the part that actually cost him.
Newport’s later essay is drier than the panic around it. These systems are trained to guess missing words from real text, so they get very good at output that belongs on the page. Plausible is the job they were built for. Finished is a different job, because finished means you can defend the edge cases when the thing is live. A sentence that looks like code is still a sentence. The model is aiming at what is likely to appear in text, not at what ought to be true in the product.
Labs then let the same kind of loop run with less watching: ask, act, report. The output can still sound like a plan, but sounding like a plan is not a plan you would bet your job on. You do not need a sci-fi story to see the cost. You need one engineer who almost lost his.
The tool is normal. It is a saw. A saw will cut the board, and it will also cut the table if you are not holding the work.
Alex Hormozi put the limiter somewhere other than a missing playbook. Businesses can get very big, but not every person running one can, because they cannot stay focused, they are impatient, and they quit too easily. That means you do not lack tactics. You lack character. The character here is staying on the known-hard part after the tool has already given you something that looks done. Anyone can accept the first draft that compiles. The known-hard part is reading it like you wrote it, because you are the one who gets fired. Tactics are new models, new agents, and a new version of “just let it run.” Character is refusing to ship a crash you cannot explain.
Steven Bartlett sat in that same place without a coding tool. He went for an interview at LBC, but they never called back, and they never even gave him feedback, so he started a podcast on YouTube in his kitchen. Permission collapsed, yet the work was still his. The kitchen did not look finished, and that is why it counted.
That sounds backwards if you bought the pitch, because the whole pitch is that you should stop touching the boring parts. After two crashes, the engineer did the opposite. He touched more of it, on purpose.
You can still use the tool. Use it the way you use a calculator. You do not outsource the exam, and you do not paste an answer you cannot walk to the board.
Newport’s survey crowd was not arguing that the models are useless. They were arguing that we still do not know the long-run way to use them in the one job they are supposed to be best at, which is code, math-shaped work, the cleanest case. If the cleanest case still produces two crashes and a firing threat, then “it looks done” is not a standard.
The standard is whether you would put your name on this if your boss was already counting, not whether it resembles a pull request, and not whether the agent filed a report. If you cannot stand behind it, you did not do it. You watched it get written.
Better to ship less that you can defend than more that you only hope will hold. That costs speed, but it keeps the job.
P.S. By the way, an AI wrote these sentences from the creators’ videos, newsletters, and posts Payton actually consumes week after week. He built the system that makes this letter, and he will keep putting it out once a week. He put his taste and judgment into what good looks like, and into how to avoid bad. He wants to be fully transparent about that.
P.P.S. Echo Improvement is an AI-written Substack Payton wanted to read, because he kept ignoring the newsletters already sitting in his inbox. He is doing the same thing with his AI letter, AI Mentorship. When an issue is AI-written, this note will say so. When he is inspired to write one himself, top to bottom, that will be obvious too. Stay tuned for that.


