Why we are building an AI harness
AI is not replacing hackers. But a hacker with the right harness covers a lot more ground. Here is what we are building and the rules it follows.
- Date
- By
- Tibeb
- Read
- 3 min
We are a small team. There is always more attack surface than people to look at it, and a lot of every engagement goes to work that is necessary but not interesting: mapping endpoints, reading the same kind of config for the hundredth time, checking whether a lead is real, and writing it all up.
Models have become good enough to help with that work. Not to replace the person doing the thinking, but to carry the boring parts so the person can spend more time on the interesting ones. Teams like Hacktron have already shown how far this can go. We want the same thing, built around the way we work.
So we are building a harness.
What we mean by a harness
A model on its own is just a chat window. A harness is everything around it that turns it into something useful for security work:
- Tools. The same tools we already use: scanners, proxies, decompilers, and our own scripts, exposed in a way an agent can call and a human can audit.
- Sandboxes. Somewhere safe to run code, try a payload, or detonate a sample without touching anything real.
- Memory. Notes from past engagements, known patterns, and what has already been tried on the current target, so the agent does not start from zero every time.
- A human in the loop. A person decides what is in scope, reviews what the agent finds, and signs off on anything that leaves the team.
The model is the least interesting part. It will change every few months. The harness is what stays.
What we want it to do
We are starting with the work that eats the most hours:
Recon and triage. Read a codebase or an application, map what is there, and point at the places worth a closer look.
Proving leads. A suspicious line of code is not a finding. The agent should try to reproduce the issue in a sandbox and come back with evidence, or drop the lead.
Write-ups. Turn raw notes and evidence into a first draft of a report that a person then edits, checks, and owns.
Over time the harness will sit underneath the rest of what we do: our red teaming work, the tooling we ship to other teams, and our own research.
The rules
Security tooling that can act on its own needs firm limits. Ours are simple:
- Authorized targets only. The harness works on systems we own or have written permission to test. Nothing else.
- Evidence before claims. If the agent cannot reproduce it, it does not go in a report.
- A person signs off. No finding, report, or action leaves the team without a human reading it first.
- Everything is logged. Every tool call and every decision can be traced and reviewed afterwards.
Why us
Running CTFs taught us a lot about building safe places to break things. We already run isolated microVM sandboxes for hundreds of players at a time, with their own resource limits and no way into the rest of the cluster. That same idea is exactly what an agent needs: a place where it can be wrong without consequences.
We also spend most of our time on the other side, writing challenges, breaking systems, and writing up what we find. That gives us a lot of real examples of what good security work looks like, which is the hard part of building any of this well.
What comes next
We are early. The plan is to use the harness on our own work first, measure whether it actually saves time and finds real issues, and only then put it in front of anyone else.
We will write about what works and what does not on this blog. If you are working on something similar or want to help, we would like to hear from you.