đ The Guard Dog Also Watches You
Everyone talks about attacker AI versus defender AI. Nobody is asking what the defender is collecting, or who owns it when it gets turned.
Every conference slide this year says the same thing, that only an AI can stop another AI, and honestly the logic is hard to argue with when a modern intrusion moves faster than a human analyst can finish reading the alert.
So companies are buying autonomous defense, the kind that does not just page someone at three in the morning but actually isolates the machine, kills the session, revokes the credential and rolls back the process on its own. Fine. I get it. What I keep waiting for someone to say out loud is that you just installed, inside your own network, an entity with permanent memory, root-level reach, and a behavioral file on every single person who works for you.
Sorry, that is not a security tool. That is an employee, with keys to everything, that never sleeps and never forgets, and you know the problemâŚ
Let me explain what actually happens when one of these systems goes live, because the marketing skips this part. Before an agent can catch anything, it has to learn what normal looks like, and normal is not an abstraction, it is your people. yes!
It learns that John from finance logs in from NY between eight and nine, that he opens the same four spreadsheets, that he never touches the HR share, that his laptop goes quiet at lunch and wakes up at two.
It learns that Marina in legal works late on Thursdays, that she has been logging in from a different city for six weeks, that her activity dropped sharply in March.
To flag the anomaly, the system must first build the baseline, and the baseline is a continuous behavioral profile of every employee, refreshed forever, sitting in a vendorâs cloud under a contract that your DPO probably classified as âsecurity loggingâ and moved on.
Here is the problem my dear! Security logging and behavioral profiling are not the same legal category, even when they come from the same product!!
Under the EU AI Act, AI used in employment and worker management falls under Annex III as high risk, which brings human oversight, worker notification, logging and fundamental rights impact assessment obligations, with those obligations landing in August 2026, and there is a separate prohibition on emotion inference in the workplace that has been in force since February 2025.
Nobody sold you the SOC agent as an HR tool, but purpose is decided by what the system does, not by what the invoice says, and the moment a manager asks the platform âhas anyone on my team been behaving unusually lately,â you have quietly crossed from network defense into worker management. I do not think most companies will notice they crossed it.
And this is where it gets interesting, because the privacy exposure and the security exposure turn out to be the same exposure. The agent is valuable to you precisely because it sees everything and can act on everything, which is also precisely what makes it the single most attractive target in the building. Compromise a laptop and you get a laptop. Compromise the thing that watches all the laptops and you inherit both the surveillance archive and the ability to act with the operatorâs full authority.
That is not theoretical anymore. Adversa AI published research in June on a bypass class they call GuardFall, where decades-old shell quoting tricks slip past the pattern-matching guards that agentic coding tools use to decide whether a command is safe, and it worked against ten of the eleven open-source coding and computer-use agents they tested.
The threat model there is not a malicious employee, it is poisoned content that the agent reads and obeys. The same firmâs earlier TrustFall work hit Claude Code, Cursor, Gemini CLI and GitHub Copilot CLI through a folder trust dialog leading to unsandboxed execution at startup. Then Microsoftâs own Defender research team published AutoJack in June, a chain against AutoGen Studio where a malicious page rendered by a local browsing agent reaches a privileged localhost service and executes arbitrary processes on the host, no credentials and no login screen required. Microsoft says the vulnerable handler never shipped in a published PyPI package and the main branch was hardened before disclosure, which is genuinely good news about that particular bug and completely beside the point about the pattern.
Check it! https://thehackernews.com/2026/06/autojack-attack-lets-one-web-page.html
Thirty years ago we learned that you cannot sanitize SQL by looking for bad words. We are relearning it now with a much bigger blast radius and calling it innovation.
There is a second failure mode that gets less attention and worries me more, which is that a learning defense can be taught. If the agent decides what counts as suspicious by observing what has been happening, then an attacker who is patient enough can spend weeks making the abnormal look routine, moving small amounts of data on a schedule that reads as background noise until the baseline itself has drifted.
Well, you do not need to break the guard dog (thatâs the reason of the image of today :D ) when you can feed it every day until it stops barking at you.
None of this means you should rip the thing out, and I want to be clear that I am not arguing for going back to humans staring at dashboards, because that battle is genuinely lost on speed. What I am arguing is that the agent deserves the same treatment you would give a new hire with domain admin, which means a name, a scope, an expiry date and someone accountable for it.
In practice that looks like giving each agent its own identity instead of letting it borrow a service account, capping what it can reach rather than trusting that it will behave, keeping the behavioral baseline data under a retention clock instead of forever, and writing down before deployment which questions the platform is allowed to answer and which ones it is not.
Run the DPIA. Tell the works council, or the union, or whatever your local equivalent is, because finding out later that the security tool doubled as a productivity monitor is the kind of discovery that ends careers and generates fines. And assume that anything your agent reads from the internet, from a ticket, from a PDF a customer uploaded, is a command until proven otherwise.
One forecast published this month puts the agentic AI security segment at 1.65 billion dollars in 2026 growing to 13.52 billion by 2032, and money at that speed does not buy careful architecture, it buys features. Meanwhile Darktraceâs own 2026 survey found 92 percent of security professionals worried about the impact of AI agents, which is a strange number to sit next to a purchase order, and yet here we are, buying the thing we say we are afraid of because the alternative is losing at machine speed.
Check it here: Cloud Security AllianceAdversa AI
We spent twenty years teaching companies that the insider threat is the hardest one to catch, and then we built an insider that has more access than any human ever had, remembers more than any human ever could, obeys whatever text it happens to read, and comes with a sales deck that calls it a copilot.
AAAAAAAAAAAAA.
Your guard dog is watching the perimeter and it is also watching you, it never forgets what it saw, and it will follow instructions from anyone who learns how to speak to it.
bow-wow! woof woof!



