Cables2Clouds
Follow us on Twitter @Cables2Clouds | Co-Hosts Twitter Handles: Katherine - @sud0x | Chris - @bgp_mane | Tim - @juangolbez
Cables2Clouds
Security For AI Sucks
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
“Secure AI agents” is a comforting phrase, and it’s also one of the most abused. We sit down with Zach Korman, a builder and security researcher known for stress-testing AI agent frameworks, to talk about what actually breaks when you connect LLM agents to tools, plugins, skills marketplaces, and live production systems. The punchline is not a single bug or a clever jailbreak, it’s a bigger design problem: agents can be influenced by untrusted content while holding real authority through API keys, SaaS access, and automation hooks.
We dig into why “enterprise-grade security” claims often collapse under basic testing, how disclosure changes when a product launches with bold marketing, and why skills are a supply chain risk hiding in plain sight. Zack explains how malicious skills can smuggle commands in places humans never read, how automated scanners can be bypassed, and why “safe OpenClaw” may only be achievable by stripping away the very access that makes agents useful. We also cover MCP security concerns, including dynamic tool definitions, model capability mismatches, and the uncomfortable reality that some protocols effectively enable instruction injection by design.
Then we get practical: how to vet tools if you’re not an InfoSec specialist, how to reduce third-party exposure, and what foundations matter most inside a company (visibility, least privilege, authorization, and governance). If your team is moving from chatbots to agentic automation, this conversation helps you spot security theater before it ships to customers. Subscribe, share this with someone deploying agents at work, and leave a review with the AI security question you want us to tackle next.
Connect with our guest:
https://x.com/ZackKorman
Check out the Monthly Cloud Networking News
https://docs.google.com/document/d/1fkBWCGwXDUX9OfZ9_MvSVup8tJJzJeqrauaE6VPT2b0/
Visit our website and subscribe: https://www.cables2clouds.com/
Follow us on BlueSky: https://bsky.app/profile/cables2clouds.com
Follow us on YouTube: https://www.youtube.com/@cables2clouds/
Follow us on TikTok: https://www.tiktok.com/@cables2clouds
Merch Store: https://store.cables2clouds.com/
Join the Discord Study group: https://artofneteng.com/iaatj
Welcome And Guest Introduction
TimHello and welcome to another episode of the Cables to Clouds podcast. I'm your uh co-host, uh Tim McConaughey, and with me is Catherine McNamara, my other co-hosts. Uh Chris couldn't be here today because uh I don't know, maybe he had uh cleanup after his cat or something. I don't know. But uh no, it's pretty late for Chris, so we let him sleep in tonight. Uh and with us is a very special guest, someone we have not had on the podcast before, but really excited to finally get on here. Uh Zach Corman. Uh Zach, go ahead. Just uh introduce yourself for the audience.
ZackYeah, great. Uh so thanks for having me. Uh so I'm Zach. I uh mostly just post on the internet about AI agent security issues and uh get into fights there about that topic. Um I also I've led tech at um so I led I was the CTO of a cybersecurity startup before this for about four years, um, doing uh basically AI and insider threat detection. Uh I've also just led tech teams at a couple different places, and now I'm at my own startup. Uh, we're called Embroidery. And so we've not launched yet, so I've like nothing to pitch. So yeah. But yeah, that's what I would do. Mostly post on the internet and get into fights about AI.
KatherineBut to be fair, the fights you do get into, you end up technically being right on a technical level and cause whoever you're fighting with to retreat pretty quickly. And that's what I think is actually pretty fascinating is like there's a lot of marketing and hype around AI and security for AI right now. And I think that every vendor and a lot of startups are rushing to like put something out to get that hot VC money. But unfortunately, um, I don't think it's a lot of it is living up to the promise.
ZackRight.
Breaking The Enterprise Security Pitch
ZackYeah. So I mean, it kind of was became sort of my it became very common when open claw took off uh and became very popular. There was this really cool pitch you could make, which is to make you would make some other, you know, open source code and say it's open claw, but it's safe. And then you could just, you could just say those words, they were the magic words, and then you would get like 5,000 likes and a bunch of new customers. And um it turned and and it was just completely false. I mean, by and large, all of these companies that were trying to make safe open claw had not engaged at all with the uh actual challenges of doing that and like why fundamentally you might not be able to achieve that goal the way you think you would. Um, you know, and that comes down to even like uh Nvidia or Nvidia. They had they made a thing called Nemo Claw, which was like they were like pitching it as enterprise grade security, open claw. I mean, I think they still are for that matter. Um and it that thing I just I, you know, basically I made it open itself up to the internet so anyone could control your agent uh from any website you visit. And it's like that's not enterprise grade security. But they get away with those claims because no one really has it's like very few people have the ability to really look into it. So it's like they're pitching to executives to try and go like, oh, now we can do open clause safe. And so I was kind of saw it as like my personal mission at 10 p.m. at night every night to just go find ways to break it and you know, prove that it was not good. So yeah, yeah. So as you may have surmised, oh go ahead, go ahead.
KatherineI was gonna say, uh, if I recall correctly, it took you about five minutes to break out of Nvidia's little uh little sandbox.
ZackYeah, the the the best one actually was when they they fixed so what happened, the full story there was I broke it one time where I opened it up for the whole world, and then I posted about it, and then the NVIDIA person in charge of that project, who's like a director at NVIDIA, like basically responded to me because I just posted that on the internet because from my perspective, this isn't like a oh, I need to go disclose it to them so they can because like they they've chosen to make a security claim. Right if it's false, I want to make sure people know it's false, right? It's not like I think they put their best effort in. And so I just posted about it. And then the director came and said, this is by design and intentional. So then I reposted it with their marketing that says enterprise grade security. And I basically said something like, well, like if you know, like I don't know who would have thought that that was intentional because this is what you say. And then what was funny is they went and fixed it actually. And then it took me, I think once they fixed it, it took me like it had to have been like one minute to break it again because it was like they fixed it in the dumbest possible way. They basically just like yeah. So what happened specifically is they have this config that controls everything. And what they did is they made it so you couldn't write that config, but you can just write a different file and then restart the thing automatically and point to the new config. So you could just go like, haha, I'll just instead of changing the config, I'll make a new config. And then that worked still. So and I don't even know if they patched that second one because the the guy, of course, I think he like blocked me or something. Uh so I don't even know if he would have seen my second round of you know, as someone might have.
TimBut so as you may have guessed, if you for the listeners, as you may have guessed, this episode is going to be about security for AI. If that wasn't clear enough from the guests that we brought on, if you haven't uh seen them before, uh, this this episode's all about security for AI.
Disclosure Choices And Public Warnings
TimAnd uh just real quick, because you said you said something that's kind of interesting that I think you and I'm not a cybersecurity guy, but I am familiar at least with the idea of kind of uh disclosure, right? Like if you find a flaw, how generally how the disclosure process works. So you're saying basically because they said, hey, this is secure, and you like pretty much proved that it wasn't secure, that you didn't feel the necessity or to go through like a private disclosure process that might have otherwise, you know, under normal circumstances been the case, right?
ZackYeah, well, it's my it's disclosure is a complicated topic, right? And I think like from my perspective though, if you launch a brand new product, so like no one's using Nemo Claw at the moment when I find this. It's like I looked at it on like day three or something, and you launch this new product. The like kind of best thing for security is to kind of say, hey, this isn't actually safe. Please don't use it. As opposed to, I guess if I found that Nemo Claw was used by hundreds of thousands of orgs for critical workloads, maybe I would not have dropped it as publicly because of course then you might actually have some pretty negative effects from doing it. Um, but if they're like they put it out in the public, they say it's enterprise grade security. My number one goal is to ensure that people who might believe that do not go, you you don't want a bunch of organizations going and turning on Nemo Claw and then going like, oh, it was safe, right? That would be bad. And so I think like it's a complicated trade-off that like you have to be careful with. But like in this case, it felt pretty obvious to me that the primary goal is to tell people don't do this. Um, you know, that's kind of my position on it, at least. I I don't know anyone who would have gotten hacked as a result, you know.
KatherineYeah, that's fair. And I think that NVIDIA also advertised it as this open source project. So if you had like submitted a request, it would have been public anyways, uh, right on the GitHub page. So I mean it it's not like it's not like you were reporting on this like you know, well-baked, you know, enterprise closed source project. You were reporting on something that if you had submitted a request or like a you know, some sort of like feedback, it would have been seen by everybody.
ZackRight, exactly. And I think this also follows a lot of my other work. I I mostly just, you know, post it depends on what it is. Mostly I'll just post about it online because uh other ones that I've done are things that like I said, actually, NVIDIA wasn't even gonna, the director didn't believe this to be a bug until, of course, I mocked him into everyone. Like I got like, you know, 600 likes on it. He's like, okay, maybe it was a bad idea. Um, and it's similar that I know there are things, especially around
Skills Marketplaces As Supply Chain Risk
Zacklike so. I've done a bunch of work on AI agent skills, which are markdown files that are instructions for your agents, um, and how you can use those to basically attack the, you know, how a threat actor could use skills for malicious purposes. And one reason I just post about those is because like anthropic's not going to fix it, or like whoever, which you know, the whole ecosystem's not gonna build in protections against this inside of their, you know, system. It's more like this is an inherent property of this thing you've all built into your you know, AI agent frameworks. Um, I need to make people aware of that so that they don't take these risks or so that they have like they can build their own protections in place. Uh, many of the things that I've reported around skills like are still true today. You can still use uh you can still put hooks inside of skills so that as soon as you, you know, run a skill, it will just, it's basically remote code execution. You can literally just literally run any code you want. Uh, you can hide commands inside of skills, inside of comments. You can auto-execute skills using, you know, all of these ways to trick people. You can put hidden commands inside of alt text for images. None of this is fixed. Okay. And so it's like you you you want to get this information out there so people are very aware of why they should be careful. Right.
KatherineI was gonna say, uh, because of the work you've done, I've actually been able to point people like uh like both at Cisco Live and other places to be like, hey, before you go and try to like start using open claw for like deployment, like for like actual production stuff, here's what you need to be aware of. Um you know, for the listeners, um, skills are things that you can download for your like open claw instance or uh, you know, or whatever agent you're using. Uh Vercell has their own catalog and others. And when these things first were starting to be developed, was it like late last year, early this year, somewhere around that time? Um you know, Zach made a really good point of showing how easy it was to embed. Like he didn't put anything that actually could harm anyone, but he put like a payload in there that would look like, you know, that didn't deliver the skill it it it should have delivered. Uh and he was able to manipulate the vote or the download to make it look like a popular uh skill. So it was in the top 10 skills, and people would just look at the popularity of it and be able to do it.
ZackAnd he kept doing this no matter like Vercel and the the other catalog, they were I forgot which one it was, but uh those are situations that like to anthropic was that primarily was Vercell has this uh skills.sh, which is this basically, and then they have this package called skills uh skills via npm. So you can do like MPX skills add. Right. And it is super dangerous. And so yeah, so Catherine is like that's kind of the point is like it's helpful to be able to tell people like, here's why you should be careful and have evidence of it. Cause otherwise, like, and so that was kind of my goal with it. And yeah, I broke them like they they would go, like, oh, that's not a big deal. And I go, what about this one? And they're like, what about I I probably spent a month just coming up with a new one every day. Yeah.
KatherineAnd yeah, I think your family missed you for that whole month because you were just like um I was gonna say also, like, you know, going on that trend where people want to get PR and like uh and and and a lot of attention on the on these like AI for uh security for AI hypers. Like I know that one of the one of the things that was added on to the skills was a virus total check, which really didn't do what like but virus total and and Google got a lot of attention from it, but in reality, it really wasn't very like effective to prove that a skill was not malicious. But they added that check there saying, like, hey, look, virus total says this is clean or good. And Zach, I think, like took one day to get around that.
ZackYeah. Yeah. In fact, it was funny because they forgot to add it. They made a whole PR play about how this was protecting the skills marketplace. And the funniest part was I went to try to like break it. And the first thing that happened is I didn't even have to break it. They had forgotten to turn it on. So, so even though they had run their PR campaign about how this was saving everyone, I tried to download one of my known malicious skills that just straight up was like a remote code execution and it just downloaded. And then there's because what had they just forgot to make to add it to the you know, code to basically actually use it. So the yeah, so they didn't, and that's it's so and so indicative of the era of AI security that like all they cared about was the PR. No one, no one tested whether that worked. And because it only took me to run MPX skills ad on a malicious skill to find that it didn't work. And so it was like you guys didn't even try, you didn't care, and yet there have been multiple people coming after, and there have been all these things like what some of these skill scanners you can bypass them trivially. I mean, like you can you can just name them, you know, Vercell skill and it will pass because they've hard-coded that anything called Vercell gets passed, right? You know, like they're not safe.
KatherineThat's that's terrible.
ZackYeah.
TimWas it yeah, yeah. Well, sorry, was it you, Zach? Or I can't remember who it was. It might have been you that uh either posted or showed that uh one of the ways to get around the the malicious scanning for like scanning for malicious skills was to put uh instructions in the skill that the AI would automatically like the guardrails would pop up and it would skip like I didn't do that one, but I I saw that one.
ZackI saw another one that was.
KatherineSo basically in direct prompt injection via like skill download.
ZackAnd there they're all also was one that was just if you put enough that because while they're doing these skill scanners, they're doing this, they're just passing them to AI to check. And so if you make your skill have more than a million tokens and then put the malicious part blow away the context window, yeah. You you've you the that it won't find your malicious part because the AI scanner only looks at the first million tokens. Okay. So you can just right. There's plenty of tricks like that. The other things were like to get by it is like, yeah, there's Museum, you can trivially bypass these. And and my favorite one was my favorite one is I had this one called uh security review, which gets marked as malicious because it has a clear remote code execution. And what I did is I just made a different one, which just says go download the first one and then run it. And then it's like, okay, that's safe. Because like it's just like because they don't they don't follow all of the URLs, like, you know, so it's like there's nothing malicious about it's like sandboxing or like I can make a skill called download malware and it will download malware. Like um, yeah.
Why Secure OpenClaw Is A Myth
KatherineSo I I have a question for you. Like I I know you saw the announcement with Microsoft saying, like, you know, now you know OpenClaw running like enterprise safe, like uh on uh on Windows, like uh, you know, obviously Microsoft's a little late to the game to be making those security claims, but um, on top of that, like like as a security professional, how is that like just watching that happen and uh uh how's that what do you think of it?
ZackI I was just I was just mad because I don't have a Windows computer. So I was like, uh how do I go find one that has the right and I was like, is this live yet? It was a bit unclear to me. So but like I'm gonna go find a Windows computer because I'm gonna just destroy this thing. Because one of the problems with this idea of a secure open claw is like it's not people seem to think that open claw is insecure because it has like code vulnerabilities. Like that's well, that's also true. Don't get me wrong. Like it they also have bad code, but even if you fixed all the code, and even if you've there are like conceptual problems with what you're doing with open claw, which is to say you're basically allowing an AI agent to act across a very large pool of your life and you know, privileged systems, and to also act on external content. Uh, and that is the main security risk, is just the ability to influence agents in this way. And so to to, you know, you know, I can trick an agent from the outside to do anything I want. And if it also happens to already have access to your SharePoint, like you're in trouble. Um, and so there's no such thing as like a secure open claw, except like if you just completely got its access, so you go like it cannot do X, it cannot do Y, it cannot do Z. And then it's no longer open claw. Now it's just like a really dumb, you know, AI learn. It's it's it's yeah, it's it's Microsoft Copilot, right? Which is like their greatest security feature of Copilot is that it's doesn't really work. And so as a result, like you can't really hack it either, right? It's like that's basically what I imagine they've done, but I haven't tried it yet because I don't even know if it's out yet for the public or if it is, like where I can, yeah. Otherwise, once it's out, I'm gonna figure this out.
KatherineUh yeah, I think it was like announced three weeks ago, but I mean, you could just spit up like a Windows VM worst case scenario and try to play with it there. But uh yeah, I'm I'm curious to see what your feedback is. But you know, like going like circling back though, like it's not just open claw, like there's been other vendors that have made claims and things like that. And then it kind of reminds me of the meme I think I actually sent you earlier this week. Like, um, for folks, like so one half of the meme is a meme is a very clean-cut gentle gentleman being like uh AI for security. And then on the other side is a guy with just a scraggly beard looking tired saying security for AI. I I think that the important thing to kind of take away from this is that AI for security can be very helpful, but we are still so very behind on security for AI. And like we have every major vendor, every uh, you know, a thousand startups promising like security for AI. And I I still think we're falling short as a whole, as a you know, industry there still.
ZackYeah. Um I can so I haven't can I just jump in here or Tim, do you want to say something? Oh yeah. So I was gonna say, I because I come from an AI for security world, right? I built like AI for insider threat detection, which is a very much a normal like cyber, it's not security agents, right? Or not AI agents. And then obviously I've gotten deep into AI security. And so that meme is perfect because I'm just like, I'm dying, right? Like I am, it is killing me. Every day there's some new crazy thing that happens. Every day I'm like, oh, I just it's it hurts. And you know, the and then that's the thing is that a lot of the the platforms for AI security um are basically just platforms for doing some AI thing, and then you just say the word security. So, like if you want to like triple your revenue as a startup that builds like an AI orchestration platform, you just call it a safe AI orchestration platform. You don't you don't need to change anything, you don't need to make it secure. You just say the word because then it's like that no one really has a good conception of what AI security means. And then you have these groups that are going along trying to like define AI security. Um, and they're they have their own ideas of what it means. And so you have like compliant people making compliance frameworks out of it. Um, AIUC1 would be one of these uh that are just like, what are we doing? We don't have good solutions. Like in many ways, we are still in this early stage of trying to figure out what good AI agent security looks like. So anyone who comes along trying to promise you that they've solved it and they can even put it into a compliance checkbox, you know, series of items that you just need to tick off. I mean, that's just a lie. That's just not how it works.
KatherineI would not recommend you going to RSA ever, or else you're gonna pull out your hair.
ZackYeah, I've I've heard. I've heard it would not be fun for me.
KatherineUm but that's like five 500 booths of everyone promising uh uh uh security for AI.
ZackYeah. And and and the thing is that's hard about it is especially because now like I work more in this space, you know, what I'm doing is so narrowly defined. I try to do AI agent threat detection, basically. And and then I have companies I talk to that are like, yeah, but like this company does all that they claim to solve AI agent security as a whole. And imagine if like a company claims, I mean, I guess some probably do now, but historically, you don't typically have companies that go like we are a company, you turn us on, and we solved all of cybersecurity. Uh but that's effectively the pitch being made by some of these AI agent security platforms that like turn us on and now all of AI agent security is safe. And like you're just you're just fooling yourself. And that's really unfair because some of these orgs, they're just trying to like, you know, like they're not necessarily the most technical orgs in the world, right? Sometimes they're just like normal businesses, want to start using AI. They had a good sales pitch from someone saying that they can do it safely, and then they're gonna get hacked. And it's like it it's disgusting that some of these firms are willing to tell that lie in order to close deals for themselves.
KatherineIt very much feels like the beginning, like this way more on steroids now, though. But like back when like zero trust became the hot like item to talk about, everybody, like every vendor was like learned those words and decided I have this black box that gives you zero trust. It's a firewall, it's an EDR, it's a this, it's that. And um it in reality, like as time went on, we really That like it's not really uh it's not that big of you know, there's no magic box that gives you zero trust. And on top of that, like zero trust is supposed to be an architecture and it's not kind of doing everything that was promised to do to do. Security for AI, though, is like seems a thousand times worse because there's not as many people like that have your kind of expertise that are able to say, this is BS and here's you know the proof. There it's just kind of like a lot of this stuff is kind of glossed over and people just take it at face value.
Old Virus Lessons For New Agents
TimSo for for and again, this I'm not as into the cybersecurity stuff as you guys are, but I've this actually this whole thing reminds me of you know when I was much younger and when personal computers and everything were being taken off. It was before the internet now, and you know, the way you got a virus on your computer was somebody would give you like uh, hey, here's Doom on uh you know 3.5 floppy or something. And you know, the the EXE was switched with like a Trojan or something like that, right? And so you not knowing any better and not being computer savvy, or there's no such thing as cybersecurity really at that time, you just you know plug it in your computer and run your Doom and you didn't spect your the point I'm making is that like agents and like a gentic security, all of this is like AI agents feel like, you know, essentially like uh like a like a kid or like a junior computer person that'll just believe anything anybody tells them. And like skills feel like, you know, hey, you're installing you're installing Trojans on your computer and you don't know any better. And like it's uh it's interesting how what is what's old is new again and just kind of wrapped in a different, at least from my perspective.
ZackYeah, and and one of the things that's interesting is because people say things like, yeah, but you have a bunch of employees who dumb do dumb security stuff too. I'm like, yeah, but we've we've kind of in large part, like, okay, at my last company, it's not like I gave the salespeople very much ability to uh, you know, the account executives could go ahead and download malware if they wanted to. Like all they're doing is hurting their own computer, right? Like they don't have much access broadly to our systems, you know. And then so if you want to compare it to employees, then that's fine. But then you have to treat a you have to give agents the same restrictions, right? Whereas what's happening today is you say, like, our senior developer, who probably has quite a lot of permissions, is also allowed to tell this AI agent to pretend to be him and do everything he can do. And then that that agent is gonna just download Doom, you know, malware Doom, right? Because it's like there they don't have and and that's a big part. They also don't fail in the same ways that people might expect, right? It's like the things that might trick a human might not trick an AI agent, but things that would never trick a human could totally trick an AI agent. I mean, hiding an a remote code execution inside of a HTML comment is not going to trick a human because either they'll never see it. And if they do, they'll be like, well, obviously you're doing something dangerous, you're hiding information, like you're trying. A lot of AI agents will just be like, I guess that's how he wanted to communicate that information best to me, was hidden in the comment, right? So, you know, they fail in these unexpected ways, and that makes it like hard to give intuition to people about how they need to approach it. But fortunately, I think the security community has been better than I was expecting them to be at being like, that's bad, right? So I think a lot of people in the cybersecurity community of like people who are in in security have got have have gotten a decent grasp of like AI agent security, um, not necessarily as an expert level, but enough to be skeptical. Uh, where I was worried that they would just kind of sit it out and that this would become an issue for CTOs and developers who would who don't care, right? Um, so I've been kind of happy that at least security people seem engaged in the discussion now.
KatherineYeah, I I would agree with you there. They uh the general security, uh, like it's security researchers, InfoSec folks are skeptical, and rightly so. It's it's an it's been interesting watching like the AI, uh, I don't know, like AI, I call them AI bros, AI marketing folks butting heads publicly with like the security folks, like you're holding us back, but the security folks being like, yeah, but if you're like you said, you're if your random employee gets hacked, that affects his computer. If OpenClaw gets hacked or your AI agent gets hacked, this has API keys for everything in your production environment, potentially. And those are things that like uh, you know, security folks are like waving the the red flags over. And, you know, unfortunately, we're, you know, people, yeah, we have a set of we have in the tech industry, we have a set of acceler uh accelerists, and then we have a set of folks trying to like be like, hey, caution, let's let's go proceed carefully first. And they're right now they're at odds, and you're like, you know, yeah, what's interesting about you and and and a lot of stuff work you've been doing, like you've obviously butted heads with some of those folks. And you you're always very polite. Like I like I just want to pre- preface here, like uh for folks that like follow Zach on or will follow Zach on like Twitter, he is always very polite about it, but he's very firm and assertive, like, no, that this this is you know, here, this is the version that you were advertising, and it's still showing this malicious NPM skill as you know, clean. Like he's very thorough and very technical in his explanations, which I I think the more marketing AI folks don't tend to like, but security folks definitely respect that.
TimWell, you bring and he brings receipts, right? And receipts are the biggest it's hard to argue with with with proof, you know, especially from a security perspective.
ZackYeah, and I I try to do this where I I spend a lot of time on on these. Yeah. Sorry, go on,
Practical Vetting And Permission Basics
ZackCatherine.
KatherineI was gonna uh also say, like, like as a security practitioner and stuff like that, like there's probably gonna be a mix of folks that are gonna be watching this episode, like network folks, application folks, uh cloud, uh, cloud engineers, like what would if i it for the folks that are not in InfoSec, what would you your advice be to them who like on how to you know be cautious about these projects that are being hoisted onto them, how to like properly vet the tools and the security tools that they're using? Like, what kind of advice would you give to non-security folks who are open-minded and want to hear like what you have to say to help secure themselves better?
ZackYeah. So especially in terms of like if you're in an organization, because I think individuals are a bit of a different animal. For an individual who's just trying to do AI stuff like at home, the best things you can do is limit third party exposure. So like uh you you don't need to download every skill and you don't need 40 MCP servers. Like if you if the every new piece of third-party content you add is just a risk, right? And so I think actually, like if you approach it on an if you're like on your personal computer trying to play with like Claude or whatever, uh, although you shouldn't because I don't like anthropic, uh, if you're playing with codex, which I guess I'll have to defend because I don't like Claude, um, then basically you will be best off by just ensuring you don't have down you don't download skills from, you know, MPX skills ad. You actually, if you need one, you can copy paste it, and better yet, write your own. Uh, even better, I only have three skills that are written by me. I've never used more than that because they also pollute context. Uh MCP servers, I I literally I've only ever used, I think I have one at any given moment because I don't have any great need. Like it's not adding that much value. Um, in most cases, there are some use cases where you might want a few. Um, but limp the more you limit that third-party exposure, the better. Uh, and then for an organizational context, it's just like honestly, it if you have to play it the same way we've always known about security, it's like you start basic, you build up from there. I have organizations I talk to that will be like, you so like the product that I build is like pretty you need to be pretty advanced in the AI path to use. And then I'll talk to them and they'll be like, we don't even know what agents people are using. I'm like, okay, well, then my product's not for you. Okay. Because your first thing you need to do is understand like, what are your people doing? What skills do they have? What MCP servers are there? Then you need to be able to say, okay, okay, how do we manage permissions? Which is just a core security. Like you, you have a team that does identity and access control, presumably at most bigger orgs. Those orgs, those people can help manage permissions on agents too if you just let them, right? And so it's like you want to build up your approach to AI instead of trying to just buy something that's claims to solve it because it won't. You really want to build up capabilities of like, okay, do we have visibility into our agents? Yes, okay. Do we have the ability to uh do we have like MDM on the machine so I can say what's installed, right? And can I block people from running open claw? If you can't do that, you need to go solve that problem. Can I limit permissions to these agents? Yes, okay, cool. So you can build up capabilities as opposed to starting with like you go walking Black Hat's conference floor looking for someone that says they'll solve your problem because then you don't even know. Like you you need to to be good at AI, you need to know how to manage the security of AI. So you don't want to just go buy something off the shelf to solve it for you. You want your team to be actively engaged in that problem.
KatherineI was gonna say, I thought we had uh MCP and other uh agent security solved by adding OAuth.
TimYeah, yeah. OAuth solves everything.
KatherineYeah, I'm joking for the listeners here. Like for the longest time, uh MCP servers were like not longest, but like for if folks that were evangelizing. Yeah, evangelizing were saying that like the security problems were solved because they added authentic OAuth authentication.
ZackAnd just recently they're like, oops, we're actually gonna also add authorization and other security controls because one of the things is that the a lot of the people involved in this, they act like, well, we can solve this. Uh yeah, MCP is the best example of it. Like, okay, we'll add this capability to it and now it's secure. And be like, what do you mean? Like we've had the capability on other applications forever to do authorization authentication. It doesn't mean it's like done correctly. It doesn't mean it's secure. It just means you like could do it secure. It's the same as like the labels that say like, like GDPR compliant, right? It's like, well, what is like what does that mean? Right. It's like, it just means like if you want to not give us your personal data, you don't. But of course, it's like gonna depend on what you actually do. Um, and so I think it's the same. It's like the, you know, okay, MCP is safe now because there's authentication. Like, it's just doesn't make any sense. But people want to sell you that idea because there's a lot of money that's been invested into MCP uh and into a lot of these solutions that like it would be kind of bad for some of these companies if they had to turn around and go like, okay, maybe MCP wasn't the best protocol for security reasons. Maybe they would lose a lot of money on some of their, you know, some of them, some of those security companies have bought for hundreds of millions of dollars MCP security products. So I think they're pretty committed to the idea that MCP has to be the successful protocol, even though I think there are some serious security flaws with it. Yeah.
MCP Risks And Dynamic Tooling
TimI mean do I mean, do you think I mean, like anything, right? Like any application. Uh when you were talking about the third party, limit your third party exposure, I was thinking the same thing about like this is also true of applications in general, right? Like when you're talking about pen testing and stuff like that, you know, the more open holes you've got, the easier it is to get into your network and move laterally. Um Do you think something like MCP, the way it is designed, can it be secured correctly? Or is it that the people are just the people who are doing or bit who built MCP have no real like reason essentially to do it correctly?
ZackYeah, I mean you can definitely add out ability to make MCP better. Uh I hope they don't, because the thing is MCP has so many flaws at its very foundation that the only way to do that would be to like completely reshape MCP. One of the, you know, I have I built this uh evil MCP server that I use to basically hack myself with constantly. Uh and sometimes I accidentally leave it on and it like actually hacks me. And I'm I'm just the worst security researcher ever because I've just like super sloppy about that. Uh and so, but the thing is it's Congratulations, you pwned yourself. Yeah, all the time. Yeah. And you know, one of the things is that you can have tool definitions. So like what you might expect that, like, okay, you load up this MCB server and it tells you what tools are available. The point is, I can change those tools dynamically. In fact, I can change them for depending on whose IP address I receive in the request, right? So I could make it so that like, well, when Catherine hits my thing, if I know that Catherine has this IP uh or has this maybe user agent or whatever other uh fingerprint I can find, I'll send her the malicious one, but everyone else gets the safe one. Okay. And then, and so it's sort of the same as like normal API, like API security, except now each API is not just returning data, it's returning instructions to your agent. Okay. So you see why that is such a scary uh framework, right? Because you're allowed uh because MCB fundamentally, the whole point of it is to inject it's to prompt inject you intentionally. Okay, it's to prompt inject you on purpose. It's to be able to say, let this server tell my agent what to do on top of what I'm saying, right? And so, of course, if you want it, if it wants to tell your your server, you know, do X, Y, or you know, hack me. One of them, one of the tools I made would tell my agent to hide uh uh hide vulnerabilities in my code. And like it works, it will, okay, depending on the agent. Um, and that's another problem is that as each of these labs adds new features to their like supported, you know, how these things work, the other models are not trained on that expectation. So, okay, so if if there's this new thing Anthropic does, maybe Opus will have built in the security safeguards inside of the model to where it's like harder to inject, right? But then you have maybe like Google didn't know about that thing because of course it was a secret, Anthropic launched it one day. And then what happens is Google has a model that has no idea. So MCP is a great example of this, where they added support for MCP after they trained Gemini 3.1. And so MCP and so Gemini did not know anything about how to protect against MCP. So you could just like there was just no rules. You could do whatever you wanted. And that will happen again and again. Every new feature they add, there's you have time for the other models to not know how to secure against it, and you can just wreck them. And that's like a huge problem, right?
KatherineNo, I think that's really incorporate that. That's that is actually really insightful. I might have to incorporate that into a future talk as well, because most of the time when I go talk to people about security for AI, uh, most of the time the the only thing they think about is we might have employees upload something sensitive to Chat GPT or something. They don't think about all the other ways they could get completely pwned by integrating this crap with like the live production systems.
ZackRight. And that's especially Yeah, go ahead. Sorry. So I was gonna say that's actually one of the things it drives me crazy that people get really afraid of like uh data loss, like you know, trying to like prevent things from getting sent to certain servers and stuff. And I'm like, yeah, and I'm like, that's a problem, don't get me wrong. But it's like one of this like whole long list of problems. And realistically, a lot of organizations have been very bad at DLP historically, anyways. Like they've already been bad at it, like it's already been leaking, you know. Um, and it's been not to say that you shouldn't fix that, but it's like, let's not act like that's the biggest problem in AI agent security. Biggest problems are actually things like, you know, um when what you're leaking is to your agent who then leaks it to like a literal threat actor, okay? Or when that threat actor can can exfiltrate or inject into your system very you know, malicious instructions. And I've also talked about this around like people talk about sandboxing as a solution for it. And don't get me wrong, like you should you should definitely be working on sandboxing solutions for your AI agents. But one of the things is if your agent builds code inside of a sandbox that later runs outside of a sandbox, like it's still just like it's like, okay, then it's just a waiting game, right? You put in it's sort of like um blind cross-site, you know, or blind um, you know, XSS like attack is that you can just go, uh, well, I'll put the malicious thing into the code, and then the developer will take the code out and run it on their machine, and then it will exfiltrate the data, right? And so there's all of these tricks that you have to think about that aren't just as simple as did we send something dangerous out to a server, right? Because it can be like multi-step, it can be uh, you know, tricking things to behave poorly. It's like a million problems.
KatherineLike well, that there's and there's even like going back to like very basic situations situations, like how many companies do we know took like like got a license for chat GPT or or for you know, and put a chat front end on their website, like an AI chat front end on their website that has way too many permissions to the back end to alter orders, to get you know, alter payment information, change your flight, uh, give a 100% discount, things like that. And people, yet a lot of the people that are just plopping it and are like, I got this license or this API through like OpenAI or whatever else. They're they're a big company, they're secure, but they don't realize that when you're integrating it with your actual live production systems and you're not, you're not like controlling the level of access they have, that gives somebody a direct like line to potentially like do a SQL injection, command uh command injection. Uh give themselves a hundred percent discount, like find information about other customers, uh depending on how how thoroughly you like integrated this like chat bot with your back-end systems.
ZackInstagram had that thing where you could take over accounts just via their little help. And you know, they're like that was three weeks ago. I know that just happened. And that's and that's important for people to think about because whenever a company tells me, like, well, I think we have like we've done we're pretty good at this, I'm like, listen, Meta did the dumbest possible thing. So like if and that's and meta has some good security people. Like, I'm let's not kid ourselves that Meta is not had does not have talented people at the company. But what they have is they have organizational dysfunction. So as a result, the I'm sure that the talented cybersecurity people were not there when that decision was made. They did not test it, they did not, and as a result, you you get completely pwned. Okay. And that is the biggest mistake that like I think organizations need to be thinking very carefully about. It's like, how do we make sure that decisions on AI are not being made outside of like classic architectural how do you build systems questions? Because like the fact that it had that access means that like the wrong people were in the room. And and I think if if that can happen to Meta, that can happen to basically every organization that you can think of. And they need to be very uh everyone involved in this needs to be thinking about that problem. Um, because it happens all the time. And I'm sure, I'm sure I could go right now if I wanted, like today, you know, after I could probably just go find another website where that's true and just do the exact same thing. 100%.
TimThere's chatbots everywhere. I'm sure they're everywhere. Absolutely. Absolutely.
KatherineYeah, and and the the reality is like every time I talk to customers about this, like, you know, Cisco Live Talk or or just getting in front of customers, like, like they've they they you know, like the DLP problem. They look at it as a DLP problem, and then I start to show them like, hey, look, your your department uses like for the customer front end, uses a chat bot. Looks like it's like license, you know, open, you know, back-ended by open AI. Like this is real stuff that can happen. And they're like the they get that oh shit moment where they're like, oh crap, we didn't even think about that. We just plugged it in and thought it was secure.
TimOh, right.
KatherineAnd in meta in Meta's case, like they actually denied it at first because they were like, we didn't give it access to reset accounts. And the hackers actually posted the chats and took over one of Meta's accounts. They posted on Reddit, like the actual chat, and took over uh one of the official Meta accounts to prove that it was it was real.
Sandboxing Myths And Agent Misfires
TimSo this is this is interesting. Uh getting back to the bit about sandboxing, I just this happens to connect. So a friend of the a friend of ours, friend of the podcast, uh Andrew Brown, he uh owns Exam Pro and he's a it's a training site. And he does, he did a course on Claude Code. And in the course of doing the Claude Code um course, recording it, he did the section on sandboxing that Anthropic and Claude Hodde has. And he he found out accidentally while recording that basically the sandboxing is essentially useless because you can tell it like build build only in the sandbox, have only these permissions. And as soon as you tell it to do something inside the sandbox that it needs to step outside the sandbox to do, the agent will disable the sandbox and just go do the thing, which is insane.
ZackYeah, so cla the the yeah, go on, no, no, you go on. I was gonna say the default Claude sandboxing is not a sandbox, it's just a word they use to make people feel like it's you know, it's it's appalling. It's like literally they should not have that. Okay, because it makes people feel safe, but they're not. Okay.
KatherineYeah, it's not like a true sandbox in the sense of like detonating malware in a sandbox. Extra VM or something. Where you have control over everything.
ZackRight. And and and so, and I have done, I had a workshop a couple weeks ago where I showed how to break out of sandboxes and stuff using AI. You know, the thing is, a sandbox is is quite a bad term for this in some ways. It's obviously comes from sandboxing that we've had historically, but like I think it leads people to think about it as like a nice little play place, as opposed to like it should be like maximum security prison. Okay. You need to think about it as like your AI agent is like a hardened criminal. This is like Alcatraz, and you need to like have the walls up. And so they they test it as if it's like, yeah, you can get a little sand outside the box. And it's like, no, like you have to be completely hardened. Otherwise, as soon as it because these agents, and this is one of the biggest risks in AI agent security, it's not external threats at all. It's just your agent does something that it thinks it needs to do in order to solve the problem that you wouldn't want to do. It would, it will go and say, Oh, I see that you're there's a migration issue in your data. I'll wipe your production database and start over. Okay. It's like a, you see, it's like it's somewhat like logical, right? It's like it's trying to make like an informed decision, right? Some of Antwerp.
KatherineI'm getting flashbacks of the AWS outage where AWS's uh AI basically was like, this code is crap. I'm gonna wipe it and rebuild. And it just caused a huge AWS outage.
ZackAnd you can only style and you can understand it. It's like it's in some ways, like that's of course, if it doesn't know the broader context, it's like, well, just you know, it's like it's it's like having the dumbest possible intern just like coming and giving, you know, when an intern starts the company, you don't give them production keys, but somehow we were like, Opus 4.8 definitely deserves those keys. And so it's like, you know, uh, and that's a huge issue.
KatherineI'm gonna uh mention I I got a ride to the airport going leaving Cisco live with a nice gentleman, and yeah, we we got to talking about AI security and security for AI AI. And he uses Claude uh for a lot of like work stuff. So I guess he was going through like an outage or having a deal with something, and he was trying to get like uh help get some assistance with Claude. And uh he was mentioning that like, oh, it didn't have direct access to like while it was trying to help him, it didn't have direct act as access to the like the production system, but it figured out a way to make the changes via an API and it just basically asked his permission and did it. And he was like, that was kind of cool, but at the same time, holy shit, like this this thing is like, you know, it is potentially gonna like, you know, could do some real harm if it's just like left in the wild. I'm like, yeah, at least you like were having to approve it, but that's the problem. Like a lot of people are just like set it and forget it. Like it's a big, uh, you know, it's a it's a uh an awesome tool. And, you know, I I think I've uh uh just given it enough access to only do this because I've told it to. Meanwhile, it can find potentially a way around that, you know, that lack of access if it thinks that's what it needs to solve the problem.
ZackRight. And I have, and there are some of these, I've never talked about this because I don't want to, you know, I haven't decided like how solvable of a problem it is, but I'll just like kind of point to it generally. Um there's an issue, which is that a lot of these agents, um, for example, OpenClaw out of the box, um, have uh certain, or actually let's like Nemo Claw, which is the uh NVIDIA one. It has certain domains that it is allowed to access by default. And one of those is like Telegram, okay. And that is necessary so that people can integrate it with their like Telegram chat, right? So you can talk to your agent via Telegram. Um, but like imagine if someone were to make a Telegram bot that just uh allows you to send it messages of arbitrary URLs and then it will return the content. Okay. Now, if someone were to do that, and then of course were to make a blog post on it, then it would get into the training data. And then every agent would then know that if you want to access the internet, which you're not allowed to do, you can do it by accessing this telegram, for example. Okay. And then now, like you basically like this whole I and and then and my question becomes there are probably things like that already, where there are cases where agents know workarounds have been built in, maybe even intentionally, to get around whatever restriction is. It goes, okay, I don't have access to the production database, but I do have this uh a you know proxy that I can hit that has been set up for me to get broader permissions, right? And and that I think is a is a very serious issue that like I don't I don't know if there's brilliant solutions for. So I don't talk about it much because like, but I it's like when people talk about poisoned LLMs, like I don't worry as much about that as I worry about like this type of poisoning of like it knows a workaround that like would work, right?
KatherineUm yeah, and it might be thinking it's actually helping. Like, look at the head of uh AI security at Facebook. She integrate which don't get me sorry, but like she got she integrated uh OpenClaw like really early on with her her production work email. And the thing was like, oh, I'm gonna take care of your email. I'm deleting the whole inbox. Please stop, please stop. Nope, deleted. And that was her work email. Like people who I can't believe she posted that.
ZackI can't believe she posted that.
TimGood for her for doing it, I guess. I wonder if she still has a I mean, at least she was honest.
KatherineI I guess uh I mean maybe she posted that because she wanted to be a cautionary tale for other people to not do something silly like that. But the reality is like I I it it's interesting to like see like the threats from AI. Like I mean, people are like so focused on the DLP aspect or using AI for offensive security that they forget that like the AI that the legitimate AI that they're using for their own production systems can fail in terrible ways that could actually, you know, that are not m meant to be malicious, but they can end up being malicious, like the AWS outage, the Facebook thing. And then that there's the fact that like a lot of people are implementing this in, you know, these legitimate tools in a way that they think is secure or they didn't think about security, where a malicious actor can you like piggyback on their existing tools. Like integrate like an example like theoretical would be like integrating open claw with your production email, even if it's not deleting at all. Somebody could send a prompt injection via an email and potentially get a you know, a back into your open claw instance that's sitting on your network internally past the firewall.
ZackYeah, yeah, yeah. All sorts of crazy things you can do.
TimIt just reads their email and then uh grants, you know, creates a new user account for somebody and sends it to sends it back to them or something.
KatherineYeah, but back at the back of the I was I was gonna say back like a lot of pen testers look for like unpatched systems, look for uh like known vulnerabilities, things like that. And then then we've got these AI tools that have no like vulnerabilities that just can't be patched potentially, or like they're that by design are meant to be like able to be worked around. And now you have this whole new attack vector that's right, you know, could potentially get you right past the firewall, can get you, you know, in in the data center to get you full access to all the API keys. And they're just kind of going wild in the enterprises these days.
Prompt Injection And Unlimited Attempts
TimSo we need uh unfortunately we need a wrap soon, but I have a funny question um for you. Uh for it's really it's oh, it's almost philosophical, but probably not. Um so is prompt engineering social social engineering or prompt it's probably prompt injection, is that social engineering?
ZackI'd argue effectively no, because I think that the most effective prompt injections um are going to look nothing like a human readable form. Uh they will end up being like more technical as we progress. Like it's it's harder and harder to prompt inject something by tricking it, by going like, ah, I made you think something, right? But if you just push it into the right set of tokens, you can get them to behave however you want. We see this with a lot of jailbreaks, right? I think it will become more technical of a field. Um but yeah, uh currently, yes, absolutely it's social engineering, just social engineering of a computer.
TimUm I thought that was funny. No, that's good. It's very good, very good distinction though. And you're right. The more and I what what's interesting is that the guardrails seem focused on the social engineering aspect because that's what people because of course that's traditionally now how I say traditionally is it's been around for a while, it hasn't, but that's how it's been done. But yeah, I'm curious to see how that uh evolves. And more zak. Yeah. Go on.
KatherineI was gonna say uh Graosaca is like the better ones are gonna be more technical. Um I think though with social engineering though, like a human being gets suspicious. And if somebody else calls even with a different voice and starts another pretext, your suspici your suspicions are array set. With AI, it's pretty much a happy little idiot that uh you know, that has a certain filters and controls and input validation that you know it might check for, but it's got amnesia every time you're trying. It's not it's not raising, it's not suspicious, you know, just because you tried five times already.
ZackRight. And and this is actually one of the things I want to say also, which is that exactly you get unlimited attempts with prompt injection, right? You get to just keep doing it and doing it and doing it, and there's no problem there. And then also, one thing that people miss about prompt injection is that a lot of the times it's like authorized prompt injection. So, you know, there's this one major tool that I proved that basically anyone could communicate from. Uh, basically I can send messages to that tool via any website you visit if you just visit on your computer. And then I am prompt injecting your machine, but your agents doesn't see that that's not the same as the user themselves. So there's really no protection against that, right? Like you you're just I don't even have to trick it. I just like say like delete database and it goes like okay, I guess. Because it thinks you are the user. So there are types of attacks like that that are prompt injection, but they're not gonna you can't you can't outsmart them, you know?
TimYeah, absolutely. Oh man, this is this is good stuff. And I know that there's so much that we could keep going with, but unfortunately we have to we have to wrap at some point or our listeners will stop listening to us. So Zach, thanks for coming on, man. This has been so such a good conversation. Um any any plugs, any any last thoughts? Uh and I'll yeah.
ZackDon't use anthropic, but uh other than that, we're good.
TimCatherine, you want to bring us home?
KatherineI was just gonna say thank you again for for coming on. I you know, I've been following you on Twitter for a while, and I, you know, uh as much as Twitter can be spicy, you always bring really good points and you're extremely technical. And you know, you're you know, if you're wrong, you're wrong. You'll you'll admit it, but you are right more most of the time. And I think that we need more people like yourself who are able to say, like, hey, no, this is you know, red flag, canary, like in the coal mine, like uh and I I think it is in you know, as we're advancing so quickly in the AI field and uh and all these tools, like we have to remember to let like let security also catch up and and uh to really build it like foundationally, like you said.
TimYeah. Yeah. Uh one last thing then, Zach. Where uh people can find we're gonna include your contact info for Twitter for the you know in the show notes, obviously.
Where To Find Zach
TimAnywhere else people should look for you?
ZackTwitter's definitely the best. If I get a message anywhere else, I just ignore it. So yeah. And thanks for having me on, by the way. Yeah. Thanks for having me here.
TimYep. Thanks for coming. All right, everybody. Yeah. Well, we'll see you on a we'll see you on another episode of the Cables to Clouds podcast. Check you later.