← All field notes

AI Created My Security Backlog. Then AI Fixed It.

It’s been a minute since I’ve written anything for the blog. I’ve spent a lot of that time building side projects, experimenting with AI coding tools, working on security problems, and messing around with infrastructure. Occasionally, I’ve tried things completely outside my normal wheelhouse, like game development.

After spending a ridiculous amount of time with AI agents lately, I think I’m starting to see where this is going. The tools are still far from perfect, but the amount of work a single person can reasonably attempt is starting to change.

A few years ago, if you told me I had nearly 1,000 security findings across my side projects, my first thought would have been, “Well, I’m never fixing all of those.” Recently, I decided to see what would happen if I threw agents at the problem. The result surprised me.


How I Accidentally Built a 1,000-Finding Security Backlog

Like a lot of developers now, I use AI constantly to help me write code. Some of the things I’ve been able to build with it have been shocking and impressive. But AI also writes buggy code—sometimes very buggy code—and that side of the experience catches up with you.

Across my websites and side projects, I’ve seen plenty of the classics: OWASP Top 10-style vulnerabilities, server-side request forgery, insecure implementations, vulnerable dependencies, and all sorts of other problems. I’m not particularly proud to admit this, but over time I managed to accumulate nearly 1,000 dependency and security findings.

For one person, that is a lot. From an enterprise security perspective, though, a backlog that size isn’t difficult to imagine. Large organizations can find themselves dealing with tens of thousands of findings spread across repositories, applications, containers, infrastructure, dependencies, and services. Finding the problems is only the beginning; somebody still has to work through them.

The Old Vulnerability Management Math

Let’s pretend I manually investigated one finding every single day. 1,000 findings ÷ 1 finding/day = 1,000 days, or roughly 2.7 years. That assumes I work on security findings every single day, without taking time away for everything else these projects need.

Obviously, vulnerabilities don’t all take the same amount of time. Some take five minutes, some take three days, and some turn out to be duplicates or false positives. Others require upgrading a dependency that breaks half the application. The calculation is deliberately simple, but it captures why a large backlog feels so difficult to tackle alone.

Diagram: The Human Security Backlog

A backlog of 1,000 findings passes through investigation, reproduction, patching, testing, review, and closure. At one finding per day, the simplified total is about 2.7 years.
One engineer working through a security backlog. Open full-size diagram.

*At one finding per day.

This is an intentionally simplified calculation, but it illustrates the problem: humans don’t scale particularly well against enormous security backlogs. Even when each finding is manageable on its own, investigating, patching, testing, and reviewing the entire pile becomes a substantial commitment.


So I Threw Agents at It

One night I basically said, “Screw it. Let’s see what the agents can do.” I put Claude, GitHub Copilot, and OpenAI’s coding tools on the problem and started letting them work through the backlog.

They tackled critical and high-severity findings, dependency upgrades, code changes, and false-positive analysis. The work also included tests, pull requests, and build verification. In roughly half a day, the agents had worked through an enormous percentage of the backlog I cared about.

That doesn’t mean every fix was correct, and I’ll come back to that distinction. Still, the amount of work accomplished surprised me. Parallelism made a difference: several parts of the backlog could move forward at once instead of waiting for me to investigate each finding individually.

Diagram: The Agent Security Backlog

Findings are grouped and prioritized, then assigned to parallel agents. Their changes go through independent review and a repeatable rescan. Passing work can close; other work escalates to a human.
Parallel remediation with independent checks. Scroll sideways on smaller screens. Open full-size diagram.

That changes the human’s role in the workflow. Instead of personally processing every item, I can spend more time designing the system that processes the items: deciding how findings are grouped, how fixes are checked, and where human judgment is required. That’s one of my biggest takeaways from this experiment.


Takeaway #1: AI-Enabled Engineers Require AI-Enabled Security Engineers

One of the senior leaders in security once told me, “The only thing that can compete with an AI-enabled engineer is an AI-enabled security engineer.” I’m increasingly convinced he was right. AI-enabled developers can produce an incredible amount of software, which also means they can produce an incredible amount of vulnerable software.

If engineering output increases by an order of magnitude while security continues operating primarily through manual review, the backlog grows faster than the team can handle it. Skilled security engineers still have finite time. Keeping up requires changing how the work gets done.

A couple of years ago, this probably would have made me uncomfortable to say, but I think background security agents are going to become necessary. There is too much code, too many dependencies, and too much infrastructure for engineers to investigate every finding manually.

The security engineer’s job increasingly includes deciding what agents should fix automatically, what needs independent verification, and what must always reach a human. We also have to define what evidence is sufficient to close a finding and which repeatable controls should check the agent’s conclusion. Those decisions shape the whole remediation process.


Takeaway #2: Remediation Has Gotten Shockingly Better

The improvement in remediation was probably the biggest surprise. I’ve experimented with automated remediation before, and my reaction was often, “That’s cute.” A system would generate a patch that looked reasonable until I tried to build the project. Sometimes it fixed the vulnerable line while misunderstanding the surrounding architecture, or announced that something was fixed when it absolutely wasn’t.

The experience feels noticeably different now. The tools still make mistakes, but they are much more capable of navigating repositories, changing code, running builds, reading failures, modifying their approach, and trying again. They can participate in more of the work required to get a change into a usable state.

That feedback loop matters. Better code generation helps, but an agent that can test a change, observe a failure, and revise its implementation is much more useful than one that stops after producing a plausible patch.

The Old Loop

A prompt produces code that looks plausible, and the human later discovers the breakage.
The old code-generation loop. Open full-size diagram.

The Emerging Loop

The agent inspects the repository, makes a change, builds and tests, and revises failures. Passing changes proceed to browser or integration testing and independent review.
A feedback loop that checks the work. Scroll sideways on smaller screens. Open full-size diagram.

That’s much closer to how an engineer actually works.


Takeaway #3: Agents Are Becoming Surprisingly Useful for Infrastructure

One of the more interesting examples involved infrastructure. I had several self-hosted runners operating across different systems, and some of my builds started failing. One particular Kotlin project kept having problems. Normally, I would spend ten or twenty minutes SSHing around, reading logs, checking runner configuration, and looking at memory usage to figure out what was happening.

Instead, I gave the problem to an agent: “These runners are failing. Figure out where they’re failing and why.” It worked through the problem and narrowed it down to resource constraints. RAM was part of the issue.

Once we understood the failure, I asked it to look across the runner infrastructure and propose a better allocation. Some workloads needed larger runners, some could stay on medium-sized runners, and some runners could be dedicated to pull-request workloads. We could start matching infrastructure to the jobs it needed to run.

Before

All CI jobs go to runners with roughly the same configuration.
One generic runner configuration. Open full-size diagram.

After

A job router assigns lightweight, normal, and heavy Kotlin builds to small, medium, and large runners, with dedicated runners for pull requests.
Match runners to workloads. Scroll sideways on smaller screens. Open full-size diagram.

It’s a relatively small example, but it reflects a change in how I work. I increasingly give operational problems to agents instead of immediately debugging every part myself. I can still investigate directly when needed, but evaluating the answer can be a better use of my time than manually discovering every piece of it.


Takeaway #4: Testing Might Be One of the Biggest Agent Improvements

Testing has been another major improvement. A year or so ago, telling an AI to start an application, open Chrome, navigate to a page, click through a workflow, and verify that it actually worked felt ambitious. Today, that’s becoming a normal part of how I use these tools.

An agent can start the application, watch the logs, open a browser, navigate the interface, and inspect what rendered. If it notices something broken, it can return to the code and try to fix it. That gives it a way to check behavior beyond simply reading its own changes.

One of the biggest historical weaknesses of generated code was that producing it was much easier than knowing whether it worked. The more agents can close that feedback loop themselves, the more useful they become—and the more evidence I have when reviewing their work.


Takeaway #5: I’m Changing How I Think About Secrets

We’ve all heard horror stories about API keys, tokens, database credentials, and production secrets leaking into prompts. I still think the default should be straightforward: don’t paste secrets into your prompts. But I’ve started changing how I give agents access to the systems they need for local work.

Instead of handing over a credential directly, I can configure the environment and tell the agent, “The credential you need is available through this environment variable.” That keeps the value out of my initial prompt while letting the environment provide access for the task. It also points toward a pattern I think we’ll see more often:

The agent requests a capability through an environment or secret broker, which provides scoped access to an external API.
Give agents scoped access. Open full-size diagram.

Long term, I’d like to move further toward short-lived credentials, scoped permissions, and tools that expose specific capabilities. Environment-based access is a useful step, but it still needs careful handling: a tool can print a secret into a log or return it to the model. Keeping credentials out of the initial prompt is only part of the job.

Also: Don’t Call Your Environment Variable TEMP

I learned another stupidly practical lesson: be careful about naming important environment variables things like TEMP, TMP, or other names the operating system and tools may already expect. Shells, build systems, applications, and libraries use those variables to determine where temporary files belong.

Deciding that TEMP sounds like a perfectly good name for a temporary API token can lead to extremely confusing failures. Sometimes the most useful lesson from an AI experiment is still to pay attention to the conventions of the system you’re working in. Linux has been here longer than you have.


Takeaway #6: Agents Don’t Replace SAST

One of my strongest conclusions is that I would not replace traditional static analysis with agents today. I want both. Repeatable, rule-based analysis gives me another way to check the code and evaluate an agent’s conclusions.

Agents can interpret context and investigate how a finding relates to the surrounding application. They can also confidently reach a bad conclusion. Traditional security tooling adds a separate layer of checks, which is why I want scanning on both sides of the remediation process.

The Security Sandwich

A repeatable SAST or SCA scan leads to agent investigation, remediation, independent verification, and a rescan, with human review when needed.
Security scanning around agent remediation. Open full-size diagram.

The combination I want is SAST + AI + verification. Let repeatable scanning identify known patterns, let agents investigate the context and generate fixes, and let a separate reviewer challenge those fixes. Then scan the resulting code again and bring in a human when the evidence needs closer examination.


Takeaway #7: The Economics of Self-Hosting Are Getting Interesting Again

If you’re experimenting heavily with agents, CI, builds, security scanning, and background automation, cloud-hosted compute can get expensive surprisingly quickly. That’s especially noticeable when the thing you’re building isn’t making any money. For side projects, I increasingly find myself asking why I’m paying someone else to run workloads when I already have computers available.

A self-hosted runner doesn’t need to be particularly exotic. Several machines can run jobs in the background, and caching becomes valuable because CI workloads repeatedly download many of the same dependencies, container layers, packages, and build artifacts.

A Simple Self-Hosted AI/CI Setup

GitHub triggers jobs on self-hosted runners. Runners use local registry, package, and build caches, reaching upstream services only on cache misses.
Self-hosted runners share local caches. Scroll sideways on smaller screens. Open full-size diagram.

Instead of having every runner repeatedly download the same container layers and dependencies from the internet, I can point them toward local caching infrastructure. The first request goes upstream; later requests can come from the local network. That can help with performance, bandwidth, rate limits, and reliability.

There is an operational cost: now I’m responsible for the infrastructure. But agents are becoming useful for maintaining that infrastructure too. For the kinds of side projects I’m running, that makes the tradeoff more interesting than it used to be.


The Weird Feedback Loop

I use AI to generate software, which gives me more code to maintain and more potential vulnerabilities to investigate. Security scanners find problems, agents investigate them, and those agents write patches and tests. CI then runs the checks that help me evaluate the changes.

When CI breaks, I can ask agents to troubleshoot the infrastructure and help optimize the runners executing those tests. Then I go back to building more software. The same tools are becoming useful across development, security, testing, and operations, which is what makes this feedback loop so interesting to me.

The AI Software Flywheel

Generate code, scan it, investigate findings, fix problems, test and verify, deploy, and repeat with the next change.
The software development feedback loop. Open full-size diagram.

That feedback loop is what I find fascinating.


Where GitHub Is Fantastic—and Where It Still Drives Me Crazy

GitHub has an enormous advantage because so many pieces of this workflow already live together: repositories, pull requests, Actions, security findings, dependency information, issues, and Copilot. That gives it an opportunity to connect finding → analysis → fix → test → PR → review → merge into a much smoother experience.

And yet parts of the experience still feel years behind where they should be. In my projects, complex CodeQL workflows in monorepos have been painful. I’ve run into gaps with Dependabot outside its happiest dependency-management paths, and Bazel dependency analysis has also been difficult in my workflows.

For context, GitHub documents CodeQL scanning on pull requests, and its Dependabot ecosystem reference lists both Bazel and Swift. My frustration is with the experience in my projects, rather than a claim that those capabilities don’t exist.

Copilot itself sometimes feels less autonomous than I’d expect given that platform advantage. It stops, asks questions, gets stuck, or waits for permission in places where I don’t expect it. I give it a task and then find myself repeatedly saying, “Continue,” or telling it to keep a file before it moves on. That amount of supervision is frustrating when I’m trying to work through a large backlog.

Those experiences keep reminding me that the model is only part of the product. How the tool handles context, permissions, failures, and the next step can make an enormous difference to how useful it feels.


The Harness Is the Product

A while ago, everyone seemed obsessed with AI harnesses, and I had mixed feelings. Now I understand the interest more. The model matters, but so does everything around it: repository access, context, permissions, parallel work, testing, retries, and the ability to recover when something fails. A capable model becomes much less useful if I have to keep nudging it through every step.

I still have mixed feelings about building custom harnesses. There is real engineering work involved, and the tools keep changing. For a custom workflow, that investment can make sense. For many of my side projects, I want the existing product to handle more of that work well.

Different Tools Have Different Strengths

In my own experiments, OpenAI’s tools have been particularly useful for browser-driven work. Claude has felt strong when the task is mostly code, especially when several agents can work on independent problems at the same time. Copilot has the advantage of living close to the repositories, pull requests, and security findings, even when the experience still leaves me wanting more.

Those are observations from the work I’ve been doing, not a universal ranking. The task, environment, usage limits, and amount of supervision all change the result. What I care about is how much useful work gets finished, whether it was tested, and how easily I can inspect the evidence.


But Did It Actually Fix 1,000 Vulnerabilities?

I don’t completely know yet. The backlog was nearly 1,000 findings, and the agents worked through an enormous percentage of the issues I cared about. That is different from proving that every finding was valid, every patch was correct, and every closure was justified.

That distinction matters. A proposed patch is a starting point. A passing build tells me something useful. A test tells me something else. An independent review and another security scan give me more evidence. None of those should disappear just because an agent sounds confident.

I increasingly want the agent implementing a fix to hand its work to a separate reviewer. I also want to revisit findings dismissed as false positives. Another agent can help challenge those decisions, but it can make mistakes too. The workflow still needs repeatable checks and a clear path for a human to step in.

The simplified 2.7-year calculation earlier in this post illustrates the size of the workload. It is not a measured speedup against the half-day experiment. I haven’t established that those two workflows completed exactly the same work to exactly the same standard.


What I Still Need to Improve

Logging and monitoring are areas where I still need to get better. If agents are going to work in the background, I need a useful record of what they attempted, what changed, which checks ran, and where they stopped. More activity is only helpful if I can understand the result.

I also want to keep tightening how agents use credentials. Putting a secret in an environment variable keeps it out of the initial prompt, but it doesn’t prevent a tool from printing it into a log or returning it to the model. Scoped access and careful handling still matter.


Bonus: Game Development Has Been a Different Experience

I’ve also been experimenting with game development, and the results have been much more uneven. AI has been useful for 2D work and specific 3D problems. Getting something to look like a polished, realistic AAA game has been a very different challenge, even with assets I’ve bought over the years.

Some of that may be the tools. Some of it is probably me. In security, I know how to describe the problem, recognize a bad answer, and explain what needs to change. I’m not a professional game developer. I built a 2D game as an undergraduate and have bought courses, but that doesn’t give me the same depth of judgment.

That experience keeps reminding me that knowing what to ask for—and knowing whether the result is any good—still matters. The agent can help me attempt more. I still have to learn enough to direct the work.


Where This Leaves Me

The biggest takeaway is that I can now attempt work I would previously have left sitting in a backlog. Remediation, browser testing, runner troubleshooting, and infrastructure maintenance are becoming practical things to delegate and then review.

I think AI-enabled security engineers, with agents working in the background, are going to be necessary to keep up with AI-enabled development. The job includes deciding what can run automatically, what evidence is required, and when the system needs human judgment.

I’m still learning where those boundaries belong. But after watching agents work through this backlog, I have a much harder time accepting that the only answer to a mountain of security debt is to prioritize it and hope somebody eventually has time.

We can start working through it. Then we can test, review, rescan, and keep improving the process.