[SPEAKER_00] Hi, everybody.
SPEAKER_00
Thanks for joining. I know we're a little late starting, so I appreciate it. I'm excited to talk about both a question that I want to pose and want to think about. And then I'll just do a demo of something that what you can do, a project that we were working on and building that you can do if you start to think about the network as more of a sandbox environment rather than just necessarily the network. So, yeah, just starting by asking the question, what are the components of a sandbox, right? So I know I say that. You've probably all thought of different things.
SPEAKER_00
You've probably all thought of probably a VM or a container and the debate between whether that's the case or whether or not the agent should go in the box or outside of the box or around the box or both or things like that. So I'm just going to break it down to something very simple and then ask a little bit about what it might look like at the network layer. So what are the components at the very basic level? What are the components of the sandbox? So first is a boundary, right? So it's just there's a thing in and there's a thing out, right, of the sandbox. And the second is a set of permissions, right?
SPEAKER_00
So if you don't have the set of permissions or identity that's a part of that, you don't have a very fun sandbox, right? It's a sandbox without any toys, right? It's a sandbox, but it's not, there's not really anything there. And so if we think about that and we think about agents, right, in particular, what it means to put an agent in a box or something similar to that, we can think about how permissions are typically handled today and what that means. And so it's one of two ways, right? It's typically one of two ways. It's the first way, which I think is what the major model labs would really like you to do, which is use API keys, right?
SPEAKER_00
So you pay the full price. And that's one. And it doesn't also get at the heart of the true auth in versus auth z, right? It's just here's an API key. It lets you have access, right, to all of the models or some of the models or things like that. And the other fun part is it's an API key. So even if it's a synthetic one, the models are very clever at doing things with keys that they maybe shouldn't necessarily do, especially if you run them in a loop for a very long time. And then the other way, maybe the more cost effective way is to use OAuth or OIDC in terms of actually handling the permissions for your agent.
SPEAKER_00
So, but both of these things are actually things that happen in the sandbox, right? So your key goes in the sandbox or you've logged into your agent and it's running over there somewhere. Your open clause running over there with your account just hanging out over there in the corner. And so that means the agent has access, right, to its own permissions, right? It's in a box, but it actually has access to the thing to give it permissions. And so my question is, what if we use the network? What if we thought about the network? And I don't know who's familiar with WireGuard, the WireGuard protocol.
SPEAKER_00
Okay, yeah, most people. But yeah, so WireGuard lets us do, and that's what Tailscale is built on top of, WireGuard basically lets us give a set of keys to all of the any node on a given network. And then at Tailscale, we're actually able to put the identity component on top of that. And so here, basically, we have the question of this is effectively what Tailscale is. And we're asking the question, what if we took the components of auth, auth z, and we just stuck them at the network level. So at least on a Tailnet, right, we're using WireGuard to establish these connections. And these are direct connections between anything that you might think.
SPEAKER_00
So a container, a GPU server, your laptop, a phone, whatever. We are able to say, in each connection, we're able to give the actual identity of what and who might be connecting. So with each connection that happens over Tailscale, you get a user if that user is logged into the device. You get all of the groups, in that sense of if you're syncing groups, so if you're in the engineering org or things along those lines, you can get all of those. You can get, if this is an agent, so a PR review bot maybe that you have running somewhere, right, in a GitHub action, it can be a tag or a set of tags.
SPEAKER_00
So this is the PR review bot for this project, or this is the PR review bot for this sort of thing. And we can take that and apply it to every single network connection. So not only can we govern network access based on that, so you can't even talk to something if you don't have a certain set of permissions, but the thing on the other side actually also gets all of the information. So there's a very, I mean, if you're used to doing things with networking, you're probably used to doing things with IP address or here's a thing over here and we're connecting things or it's IP address plus some key, again, an API key that your service is providing.
SPEAKER_00
This is all in one. So the connections happen with identity. And so what that lets you do is build some very interesting applications on top of that. So I realize this is very dark here on the screen, so I apologize. This lets you build some very interesting applications, one of which we happen to build is an AI gateway. So what's happening here is everybody's probably familiar with your typical LLMs, LLM gateway, right? Aperture works the same from that perspective. So you take a single key from a provider, be it Anthropic or OpenAI or This is all in one. So the connections happen with identity.
SPEAKER_00
And so what that lets you do is build some very interesting applications on top of that. I realize this is very dark here on the screen, so I apologize. Let you build some very interesting applications, one of which we happen to build is an AI gateway. What's happening here is everybody's probably familiar with your typical LLMs, LLM gateway, right? Aperture works the same from that perspective. So you take a single key from a provider, be it Anthropic or OpenAI, Gemini or Vertex or Bedrock or whatever, you can take a single key from any given provider, you can put it on Aperture and then on the other side. So Aperture is just a node, again, on this network.
SPEAKER_00
So it's a node that you deploy into this network. So it is actually able to see all of the identity from everything that's talking to it. So in the case of an agent in a sandbox, that sandbox has a tag. That sandbox is, we can think of in this case, a GitHub action runner as a sandbox that your agent is running in. You can use something like the federated OIDC from GitHub. That will, when that runner spins up, that runner will suddenly get the access into the tail net. It gets a tag on that tail net and that tag on the tail net is what determines what it is able to do via or through Aperture. Because again, it can see that.
SPEAKER_00
And I'll show you an example in just a second. So that's where we are. So we have a single key on Aperture. You can then write all your rules in Aperture. And then on the other side, there's actually no key. So that runner connecting from the sandbox has no key to accidentally exfil or share or do something with or go beyond its boundaries. There, it's just no key whatsoever in that sandbox. So that's that. And just to show you live, I actually prefer to just show things live. So this is Aperture, just as I had a screenshot before. Let me change it to be light mode just to make things easier to read.
SPEAKER_00
So this is Aperture again. This is what I was showing you. This is my view into my Aperture instance. So I am connected here. I'm actually on our corporate tail net. So I'm logged in on our corporate tail net. I have visited Aperture as a user. It knows who I am. I'm on my laptop. So I'm just on my laptop. It knows I'm logged in as me. And so it's showing me all of my usage metrics on our demo instance here. And so I can see all the tokens that I've used. I can see the models that I've used. I can see how much money I've spent on the given models, on any given model.
SPEAKER_00
And then I can even see all of the requests that have come through the gateway from my particular identity. So that works for me. That also works for everything else. I can even drill down and see, so this was me testing it before.
SPEAKER_00
And I can show you live. But I just asked it to say hello. And with all of the context in Cloud Code, even if you just ask it to say hello, that cost you 20 cents. That's 20 cents. But the next one is not as expensive. But I can actually even go in here and see all of the request headers, request, response body, everything here. And if I scroll all the way down, there should be, oh yeah. See, this is everything, in case you were wondering, this is everything that Cloud Code sends at the very beginning. And so if I say let's say hello. Oh no. Maybe I didn't think what I was saying. This is literally everything that Cloud sends right off the bat, is a particular request.
SPEAKER_00
And then you can see the response and the response body. Oh, sorry. I asked it to tell me a 10 word story. So there we go. A cat sat on a mat and then simply vanished. But this is what's actually going through the gateway when you make that first request from Cloud Code. So that's that. If I wanted to look at me, right, see that's me here. I can see my session. This is my Cloud Code session with two requests here. So there was the haiku thing to tell me, to give you the summary of what was going on and the 20 cents I spent to get that 10 word story.
SPEAKER_00
And then I can even also, I mentioned GitHub Actions Runners. So this is actually a PR review bot that we have, it's a small simple check that we have run on every PR update. And so you can even see here, right? So this is it. It has a tag. It's our dog food tag. And you can see everything that the dog food bot has done here over the last 30 days. And I can open it up. I can take a look. I can see every single request that it's run. And I can even take a look at something like this. And we can see here it spent four cents and it ran three commands at the same time. So I'm actually able to see all of the bash commands and everything along those lines here.
SPEAKER_00
So that's the case there. I mentioned seeing those bash commands, you can actually extract. And it's a fun part about working at the LLM layer and having everything at the network layer. There's no, I have a guarantee that I've seen every tool call that this thing has ever made through the instance. This is not happening from inside the container.
SPEAKER_00
This is not happening from the harness or anything along those lines. If it had to make a tool call, it had to go through Aperture and we would see all of the tool calls that it made and requested an MCP tool call to update the code review, did some bash, did some grep, and then updated the comment on the code review. And there's no, again, we see everything. So right. And if you wanted to cut it off or you wanted to stop it, it's happening at the network layer. So the moment you say no, it's not that it has a key and they can say oh, I see the key no longer works.
SPEAKER_00
It's just a dash. So, and just to show you that we have our agent set of script. This is all you actually have to do. So in Cloud Code, it's just, hey, you're going to run an API key mode. Here's a dash, just so you have something, so you don't complain that there is no API key for API key mode. And then re updated the comment on the code review. And there's no, again, we see everything. So, right. And if you wanted to cut it off or you wanted to stop it, it's happening at the network layer. So the moment you say no, it's not it has a key and they can be like, oh, I see the key no longer works.
SPEAKER_00
It's just a dash. So, and just to show you that we have our agent, agent set of script. This is all you actually have to do. So in cloud code, it's just, hey, you're going to run an API key mode. Here's a dash, just so you have something. So you don't complain that there is no API key for API key mode. And then here is the endpoint that you need to, the base URL that you need to use. And again, it works across codecs or cloud code or Gemini CLI. And here's what you need to use. And when you do that, you can just, again, I can say, say cloud. This is my actual settings.json.
SPEAKER_00
And you can see the same little, same things appear at the top. But I can do that. And I can say, again, tell me a 10 word story. By the way, if I, when I asked it to tell me a 10 word story three weeks ago, it was all about robots. And then it became about dogs. And then now it's about cats. So if there's a model eval suite or something, I don't know, you can tell something's happening on the back. So in terms of what they do, they got cat, cat sat on a mat and then found a home. But yeah, actually, wow. So still can't count. That's fun. Opus 4.6, 1 million context.
SPEAKER_00
There we go. All right. And then finally found home. So there we go. Right. It just forgot the extra bit. But again, if we wanted to see that, right, hey, you've got a pipeline that actually depends on that being 10 words or something, or having a certain structure. Things can easily break, in a PR review bot. And that can happen. And when it happens in something a PR review bot, it's hard to actually know what's going on or what happened or when. I can go back to my logs, right? Here's my session, right, with three requests. And here they all are, right? Here's the summary thing. And then here's the first request with all of the input tokens.
SPEAKER_00
That was the 20 cents. And then here's, are you sure about that? And here's, you're right, that was nine. So again, if you're trying to go back and look at certain things, you can do that here as well. And again, there's no hiding it from you because it's not I'm going to be super helpful and go do this thing and all that stuff and go around. It just has to be here. One other fun thing that you can do here in the middle is, first we can also do costs and cost controls and all those sorts of things that actually work across providers.
SPEAKER_00
So if you want to set a budget or some sort of budget in Aperture, you can actually have it work across every provider. It's not here's a thousand dollars for everybody. It's here's just a thousand dollars and you can decide to use it how you wish. And then the other thing is you can actually do integration. So we offer web hooks on top of this where for each of those tool calls or for each of those things, you can actually send a request out to a third party to, with all of the information about the tool call. And again, there's no hiding it. It just has to go through here. So you can, these hooks are basically guaranteed to exist, right? And run no matter what.
SPEAKER_00
So yeah, that's mostly it. If you want to again, if you want to set up things quotas, you can actually go in and say, hey, here's, you get $5 a day, all those sorts of things and have as much safe fun, I guess you could say, as you want to. And again, it works with pretty much any provider that you can imagine across the board that supports the major context. And so I talked about this at the beginning, but this is Aperture, right?
SPEAKER_00
This is the thing that we have built and that you can use as available on our free plan. However, it is built using the tail scale identity primitives. And those are all available via an open source library we have called TS net, where you can write your own go program, right? That actually puts itself on the tail net and can read all of the same identity information, can read everything else. And so you can do things if you want to build an MCP server, but it's internal to your org or something along those lines or an API endpoint or something that's internal to your org, you don't have to think about OAuth or just think about opening it up to everybody.
SPEAKER_00
You can actually do the exact same thing and be like, hey, who made this request? I'm going to force that into whatever thing I'm proxying on the MCP side or things along those lines. So you can actually take all of that same information and do it yourself. Hilariously, you can actually build aperture yourself if you really wanted to using the same things. We had a whole charge here, which was it had to be built on top of tail scale. It couldn't be built inside using private API endpoints or anything along those lines. So this is actually built entirely in a way that in theory you could go build to yourself.
SPEAKER_00
So yeah, if any ideas have come from this, if you think about things that you would to build internally, I would love to would love to hear and would love to chat afterwards. So yeah, I think I'm a minute under here and yeah, so if there is a question, I'm happy to answer it. Yeah, yeah. How do you configure the permissions? How do you configure the permissions for who can, so the question is how can you configure the permissions? And are you saying is it for who can access what or who gets who sort of?
SPEAKER_00
You can make these tools? Yeah. So all of the, so we actually can let you configure them in two places. So there's another fun little feature of how tail scale identity and how that sort of stuff gets pushed through the network. First is you actually can set them up here in grants. So you can say who this applies to. We're going to be adding groups and everything soon here as well.
SPEAKER_00
But then you can say we actually also have an MCP server in MCP proxy in here as well.
SPEAKER_00
Can you make these tools? Yeah. So all of the, so we actually can let you configure them in two places. So there's another fun little feature of how Tailscale identity and how that gets pushed through the network. First is, you actually can set them up here in grants. So you can ask, say who this applies to.
SPEAKER_00
We're going to be adding groups and everything soon here as well. But then you can say we actually also have an MCP server in MCP proxy in here as well. So you can say model access and quotas, MCP access, hooks, roles, everything along those lines. You can do the grants. You can even also define those.
SPEAKER_00
So Tailscale as a whole has a policy file that you can use. It's how you define who can access what on the network. You can actually put this, these are called application. This right here is called an application capability. You can actually stick that in your main ACL file or your main access control file to send along with the identity. So you not only can send the user or the tags or everything else, you can actually send any arbitrary metadata that you want, guaranteed by the Tailscale control plane as well.
SPEAKER_00
So yeah, we try to have the visual editor, but everything is also possible to do in JSON. Most folks, a lot of folks using this at scale want to put it in some sort of GitOps workflow. So we have that, we have the API as well. If you actually want to just put this as part of some sort of GitOps workflow that you have, to do that. Any other questions? I think that was yeah. I think I saw you when you're setting up in code and you can set the base URL to your Aperture node rather than the default. Yes. Is it possible to catch that just at the network layer or stuff it out? Yeah. Everything goes for code and then goes to your node.
SPEAKER_00
Yeah. So the question is, do you have to put the base URL in or is it possible to capture that at the network layer and make it transparent there? Um, that is something we could do. That was a big point of discussion when we were first thinking about this. And in reality, it's not, well, we could, it's not really something that we, it's not really in the, I wouldn't call it the Tailscale way necessarily. The whole point here is we want to make things really, really easy. It's for you to want to be able to get LLM access or into a sandbox. You want to be able to do it on somebody's computer. We want to make that super, super easy from the outset.
SPEAKER_00
I realize there's some transparency stuff, but when you do that, things can start to break and shift and move. Yeah. And it gets very confusing and moves out from under you. We just want to make it the easiest way for you to actually offer this sort of LLM access and not necessarily do it hidden under the surface where you're doing everything else. So it's definitely meant for folks who want to build with AI and then on the other side, it's a security or an IT admin. It's great. You get easy to use controls. You get easy, you get to see all of the tool calls. You get to see all of those sorts of things.
SPEAKER_00
So we're really trying to do the best of both worlds for both devs and IT slash security manager. Yeah. Yeah. So today it's we're, yes, we want to work on that. And today is model provider. You can think of basically anything that we would put through Aperture. You should be able to say, hey, this group or this, as defined by my skin provider, as defined by whatever, gets access to model. It's not just model and provider. It's also all of the quota stuff that also has the same sort of permissioning system. So you can say this team gets this big budget, each individual gets this smaller budget.
SPEAKER_00
And then we do the union of the two there or you can do it where it's you can use as much as you want of the internal GPU, like the internal GPU endpoints that we're hosting. But if it's Opus 4.6, you only get this amount or something along those lines. Yeah. How does permissioning work in a world where it doesn't do tool calls? It's just writing code? How does permissioning work in a world where it doesn't do tool calls? It's just writing code? Yeah, so I think, well, a lot of agents are moving away from MCP and tool calls and executing code, which makes some network, perhaps it's harder to pass.
SPEAKER_00
Yes. Yeah, so the question, in a world where people are moving away from MCP and maybe the structured tool calling, what do we, how does it work? How does permissioning and things like that work? You're right, that is a little bit more complicated. However, it's the whole reason we chose to do this. We had originally thought about maybe doing this at the MCP layer. And we realized it was actually way more valuable to do the LLM layer here. And this is where, if I go to, well, here, let me just go to the logs. And I go to, not to the chat, but to the given metric. You know, this is right. So even with skills and everything else, or code, you're still running something.
SPEAKER_00
Now, of course, you could write the thing, maybe obfuscate the thing and then run the thing, one step at a time. And a lot of folks, to be honest, a lot of folks that we talked to were like, I don't even know what tools people are using. And this is where, if I go to, well, here, let me just go to the logs. I go again to, not to the chat, but to a given metric. Here, let's see the given metric. This is right. So even with skills and everything else, or code, you're still running something. You're typically still running something. Of course, you could write the thing, obfuscate the thing, and then run the thing one step at a time.
SPEAKER_00
A lot of folks, to be honest, that we talked to were like, I don't even know what tools people are using. Please just tell me. Forget about blocking it for a second. I don't even know what are people even doing, right? It's like, because MCP was all the rage and it's like, are they using MCPs? Are they just using bash commands? I can tell you internally, this is just our demo instance. But internally, if you were to look at our actual instance, bash dominates everything else.
SPEAKER_00
But we get to see the command and we typically get to see all the commands and everything that's actually being run. We'll be adding in more guardrails along the lines of, hey, you can, this is the bash command. If it's RM-RF slash, right? Maybe not. Or something along those lines. We'll be adding that in. But that's the whole reason why we decided to do it at the LLM layer. We had that whole discussion of, well, if you can't see everything, then how valuable is it, right? If you can't see everything. So we wanted to be able to see, at least from particular agents that you want to put there.
SPEAKER_00
I don't know exactly what the time is here. But I know we're at the end. I'm happy to answer any other questions downstairs if you want to come to the booth or in the hall. But yeah, thank you. You should be able to say, Hey, this group or this, you know, as defined by my skin provider as defined by, you know, whatever gets access to model. Uh, it's not just model and provider. It's also all of the quota stuff, you know, that I kind of, that all also has the same sort of, uh, permissioning system. So you can say this team gets this big budget, you know, each individual gets this, you know, smaller budget.
SPEAKER_00
And then, you know, we kind of do the, you know, the union of the two there or, you know, or even do it where it's like, you know, you can use as much as you want of the internal GPU, you know, like, you know, kind of like the internal GPU endpoints that we're hosting. But, you know, if it's Opus 4.6, you only get, you know, you know, this amount or something along those lines. Yeah. How does permissioning work in a world where it doesn't do tool calls? It's just writing code? How does permissioning work in a world where it doesn't do tool calls? It's just writing code?
SPEAKER_00
Yeah, so I think, well, a lot of agents are somewhat moving away from MTP and tool calls and executing code, which makes some network, perhaps it's harder to pass. Yes. Yeah, so, you know, the question, you know, in a world where people are moving away from MCP and maybe the structured tool calling, what do we, you know, what do we, you know, how does it work? How does permissioning and things like that work? You're right, that is a little bit more complicated. However, it's the whole reason we chose to do this.
SPEAKER_00
We had originally thought about maybe doing this at the MCP, like just the MCP layer. And we realized it was like, hey, it's actually way more valuable to, you know, do the LLM, the LLM layer here. And this is where, you know, like if I go to, well, here, let me just go to the logs. And I, you know, I go again to, not to the chat, but to the, you know, a given, here, let's see the given metric. You know, this is, right. So even with skills and everything else, right, or code, you're still running, like you're typically still running something.
SPEAKER_00
Now, of course, you could write the thing, maybe obfuscate the thing and then run the thing, you know, one step at a time, you know, kind of thing. You know, and a lot of folks, to be honest, a lot of folks that we talked to were like, I don't even know what tools people are using. Like, please just tell me, like, forget about blocking it for a second. I don't even, like, what are people even doing, right? You know, it's like, because MCP was all the rage and it's like, are they using MCPs? Are they just using bash commands? I can tell you internally, like, this is, sorry, this is just our demo instance.
SPEAKER_00
But internally, if you were to look at our actual instance, bash dominates everything else. But again, we get to see the command and we typically, you know, we get to see all the commands and, you know, and everything that's actually being run. And we'll be adding in more guardrails along the lines of like, hey, you can, you know, this is the bash command. Let's, you know, if it's RM-RF slash, right? Maybe not, you know, or, you know, or something along those lines. You know, we'll be adding that in.
SPEAKER_00
But that's the whole reason why we decided to do it at the LLM layer. So, you know, we had that whole discussion of like, well, and if you can't see everything, then how valuable is it, right? You know, if you can't see everything. And so we wanted to be able to see, you know, at least from particular agents that you want to put there, you know? Yeah. Any other, I was going to say, I don't, you know, I don't know exactly what the time is here. But I know we're kind of at the end. I'm happy to answer any other questions downstairs if you want to come to the booth or in the hall. But yeah, thank you.