<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Rodrigo Rosenfeld Rosas — ai</title><description>Articles tagged with ai</description><link>https://rosenfeld.page/</link><language>en-us</language><item><title>AI Agents and the Refactoring That Never Happens</title><link>https://rosenfeld.page/articles/programming/2026_09_02_ai_agents_and_the_refactoring_that_never_happens/</link><guid isPermaLink="true">https://rosenfeld.page/articles/programming/2026_09_02_ai_agents_and_the_refactoring_that_never_happens/</guid><pubDate>Wed, 02 Sep 2026 12:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Working with AI agents day to day, I&amp;#39;ve started noticing a trend that worries me.
Teams — including experienced engineers who used to know better — have quietly
stopped pushing to rewrite or restructure the gnarliest parts of their systems.
It&amp;#39;s not about the quality of the code the agents write. It&amp;#39;s about a decision
that used to happen almost reflexively and now rarely does: the decision to stop
and say &lt;em&gt;this has become unmanageable, we need to refactor it before we go any
further.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;Human context is small, and that shaped how we build software&lt;/h2&gt;
&lt;p&gt;A computer can hold far more in &amp;quot;working memory&amp;quot; than a human can. We can&amp;#39;t
reason about a complex system when it branches in dozens of directions, each
branch with its own implications, all interconnected. It&amp;#39;s simply too much to
keep in our heads at once.&lt;/p&gt;
&lt;p&gt;So historically we did the only thing we could: we split systems into modules
small enough to understand in isolation, and then we spent effort connecting
those modules together with interfaces we could also understand. Modularity,
encapsulation, layering — these aren&amp;#39;t aesthetic preferences. They&amp;#39;re
concessions to the size of human working memory. We break the system down until
each piece fits in one person&amp;#39;s head, because that&amp;#39;s the only way a person can
reason about it, change it safely, and review someone else&amp;#39;s change.&lt;/p&gt;
&lt;h2&gt;The refactoring reflex&lt;/h2&gt;
&lt;p&gt;Systems rarely start out confusing. A piece of code is written when the
requirements are still simple, and at first it reads cleanly. Then the
requirements change. A developer adds a branch for a new case, then another for
an exception, then a special case on top of that exception. Over enough
iterations the original &amp;quot;rule&amp;quot; the code expressed is buried under exceptions —
sometimes the requirements have shifted so far that the code is now &lt;em&gt;nothing but&lt;/em&gt;
exceptions, with no clear rule left at all.&lt;/p&gt;
&lt;p&gt;There&amp;#39;s a moment every experienced developer recognizes: you&amp;#39;re debugging an
issue, you follow the code, and you get lost. The branches no longer form a
picture you can hold in your head. Historically, that feeling was a &lt;em&gt;signal&lt;/em&gt;. A
senior engineer, upon getting lost in a piece of the system, would pause and say:
this has become unmanageable — before I add anything else, I need to rewrite or
refactor this so it&amp;#39;s understandable again. Not for elegance, but so that they
and everyone after them could reason about it and safely review future changes.&lt;/p&gt;
&lt;p&gt;That reflex — &amp;quot;I&amp;#39;m lost, therefore it&amp;#39;s time to refactor&amp;quot; — has quietly been one
of the most important forces keeping long-lived systems maintainable. And it was
triggered by a human limitation: the moment a person could no longer follow the
code.&lt;/p&gt;
&lt;p&gt;To be fair, that reflex was already under pressure long before AI agents.
Deadlines, roadmaps, and managers asking &amp;quot;why are you rewriting something that
works?&amp;quot; have always fought against it, and refactoring was usually the first
thing to get deprioritized. Senior engineers had already, over the years,
stopped pushing back as hard as they once did when a system got unmanageable.
Agents didn&amp;#39;t create that weakness. What they did was remove the last internal
trigger that used to fire in spite of it — the visceral experience of a human
getting lost.&lt;/p&gt;
&lt;h2&gt;AI agents don&amp;#39;t get lost&lt;/h2&gt;
&lt;p&gt;Here&amp;#39;s the problem. AI agents are not bound by human context limits in the same
way. An agent can read the tangled function, trace every caller, and make sense
of the mess that would have stopped a human cold. It can add the next branch
correctly, and the one after that, working confidently inside code that no human
on the team fully understands anymore.&lt;/p&gt;
&lt;p&gt;That sounds like a strength, and in the short term it is. But notice what&amp;#39;s
missing: the agent never gets lost, so the signal never fires. The agent has no
reflex that says &amp;quot;this has become unmanageable, we should stop and refactor.&amp;quot;
It just keeps adding branches to the pile. Unless it&amp;#39;s specifically instructed —
through its harness, its prompt, or explicit review criteria — to step back and
question the structure, it will happily maintain a mess indefinitely, because
the mess isn&amp;#39;t a problem &lt;em&gt;for the agent&lt;/em&gt;.&lt;/p&gt;
&lt;h2&gt;The real risk: humans stop being able to police the code&lt;/h2&gt;
&lt;p&gt;The failure mode isn&amp;#39;t that the agent writes bad code. It&amp;#39;s that the natural
checkpoint disappears, and the humans stop noticing the code has drifted beyond
their understanding. Over time you arrive at a system where:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;No developer on the team can fully reason about key parts of the code.&lt;/li&gt;
&lt;li&gt;Reviews become rubber stamps, because the reviewer can&amp;#39;t actually follow the
change well enough to judge it.&lt;/li&gt;
&lt;li&gt;The team increasingly &lt;em&gt;trusts the agent&lt;/em&gt; precisely because they no longer
understand the code themselves — which is exactly backwards from how trust
should work.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;At that point you&amp;#39;ve lost something important: the ability to reason about your
own system without an agent as an intermediary. And you lost it gradually,
without any single alarming moment, because the moment that used to raise the
alarm — a human getting lost — was quietly removed from the loop.&lt;/p&gt;
&lt;h2&gt;Clean code is cheaper for the agents too&lt;/h2&gt;
&lt;p&gt;It&amp;#39;s tempting to frame all of this as a purely principled concern — we &lt;em&gt;ought&lt;/em&gt; to
understand our own systems — and leave it there. But there&amp;#39;s a hard-nosed,
practical reason to keep the code organized, and it&amp;#39;s one that survives even if
you&amp;#39;re perfectly happy to let agents do the work.&lt;/p&gt;
&lt;p&gt;An agent that never gets lost still pays a price for a mess. The more tangled and
interconnected a piece of code is, the more context the agent has to load and
hold to make a correct change: more files to read, more branches to trace, more
tokens burned on every single edit. A system built from small, self-contained
modules that are easy to reason about isn&amp;#39;t just kinder to humans — it&amp;#39;s cheaper
to operate, because every future change costs the agent less to understand.&lt;/p&gt;
&lt;p&gt;And it isn&amp;#39;t only about cost. When the relevant logic doesn&amp;#39;t fit into a bounded,
coherent slice, agents are more likely to lose the thread and hallucinate — to
assume a branch does something it doesn&amp;#39;t, or to miss an exception buried three
levels deep. The same modularity that keeps a system inside a human&amp;#39;s head keeps
each change inside a well-defined boundary the agent can reason about reliably.
Clean boundaries reduce mistakes on both sides, for the same reason.&lt;/p&gt;
&lt;p&gt;So keeping the source organized isn&amp;#39;t a favor we do for humans at the agents&amp;#39;
expense. It benefits humans, it improves the agents&amp;#39; accuracy, and it lowers the
token cost of every change we&amp;#39;ll ever make to that code again. The refactoring
reflex we&amp;#39;re at risk of losing was never only about human comfort — it turns out
to be good economics too.&lt;/p&gt;
&lt;h2&gt;Policing ourselves&lt;/h2&gt;
&lt;p&gt;I don&amp;#39;t think the answer is to hobble the agents. The answer is for us to bring
back the checkpoint deliberately, since it no longer happens on its own. We have
to keep asking the question the agent won&amp;#39;t ask:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Do I still understand this part of the system, or have I been letting the agent
understand it &lt;em&gt;for&lt;/em&gt; me?&lt;/li&gt;
&lt;li&gt;If a human had to debug this without the agent, could they follow it?&lt;/li&gt;
&lt;li&gt;Have the requirements drifted so far that this code is now all exceptions and
no rule — the classic signal that it&amp;#39;s time to rewrite?&lt;/li&gt;
&lt;li&gt;Is now the moment to pause feature work and refactor this into something a
person can hold in their head again?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;You can also push some of this into the harness — instruct your agents to flag
when a module has grown beyond a reasonable size or branching complexity, to
propose refactorings rather than only extending, and to call out when a change
is getting hard to reason about. That helps. But the ultimate responsibility
stays with us, because we&amp;#39;re the ones who need to be able to understand our
systems, and we&amp;#39;re the ones who lose that ability if we&amp;#39;re not paying attention.&lt;/p&gt;
&lt;p&gt;The convenience of an agent that never gets lost is real. But &amp;quot;the agent can
still make sense of it&amp;quot; is not the same as &amp;quot;the system is healthy.&amp;quot; The first is
about the agent&amp;#39;s capacity; the second is about ours. Keep asking whether it&amp;#39;s
time to refactor — because the tool that used to remind you, by getting lost,
doesn&amp;#39;t get lost anymore.&lt;/p&gt;
</content:encoded></item></channel></rss>