AI Engineer

Tethered: Our Agents Are Us — Shu Fang, Two Sigma

1927 summary words 9 min summary Watch video

Start with the signal

9 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Two Sigma enables every employee to run cloud-hosted agents under their own user identity by combining per-user Kubernetes namespaces, end-to-end agent-action attribution, and a controlled web-access path rather than creating separate machine identities for agents.
  • Why it matters: It is a concrete enterprise pattern for preserving agents' full access to user-authorized systems without accepting the usual audit, credential-sync, data-exfiltration, and prompt-injection risks.
  • Best use: Use this as an architecture and security-design reference for deploying OpenClaw-style or cloud agents broadly inside a regulated enterprise.

Executive Summary

The talk argues that the conventional enterprise pattern of giving each employee a separate machine identity for their agent breaks down in practice. Separate identities create permission-synchronization and licensing problems, fail in systems where two identities cannot access the same data, and add friction across collaboration tools such as Google Workspace. Two Sigma instead allows remote cloud agents to run as the originating employee's actual identity.

That choice is viable only because Two Sigma adds controls that distinguish an agent's actions from a human's despite their shared principal identity. The key mechanism is a propagated agent-attribution header, analogous to a distributed trace ID, injected at the entry point and preserved through RPC spans and downstream services. This produces both actor identity and a replayable provenance chain of the agent-driven workflow.

The second control is replacing unrestricted browser/search tooling with Google Web Grounding for Enterprise, accessed inside the firm's existing VPC and network controls. Native web-search and web-fetch tools are denied; approved MCP, CLI, client-code, or skill paths redirect requests to the enterprise-grounded search/fetch service. The tradeoff is freshness—typically within 24 hours, or six hours for frequently updated sites—but the speaker considers it acceptable for most agent use cases.

Two Sigma has already shipped a managed fleet of remote cloud agents for all employees, with per-user infrastructure pre-provisioned. Individual users can build and deploy agents into their own namespaces; agents intended for wider organizational use pass through normal production-support and security review. The broader lesson is that enterprises should use their existing identity, Kubernetes, observability, and network-control capabilities to make high-permission agents governable rather than prohibit them outright.

Key Takeaways

  • Claim: Running an agent under a separate machine identity attached to a user is operationally inferior to running it as the user when agents must work across real enterprise systems. | Evidence: The speaker cites permission drift between user and agent, duplicate software-license management, systems that do not support two identities accessing the same underlying data, and collaboration surfaces such as Google Workspace and email as failure points. | Implication: For agents that need to act across a user's existing data and tools, prioritize delegated execution under the user principal over maintaining a shadow agent account—provided the control plane can identify agent-originated actions. | Caveat: Using the user's identity removes an important default separation boundary, so it requires separate attribution and policy controls rather than being safe by default.
  • Claim: Per-user cloud execution environments make remote, identity-bearing agents practical without depending on employees' local machines or CLI proficiency. | Evidence: Two Sigma already maintained Kubernetes namespaces for each user across regions for automated jobs, code containers, and research notebooks. A controller provisions compute, while a pod sidecar retrieves identity material from a separate identity service so workloads run as the user. | Implication: A scalable agent platform can be built by extending existing per-user compute and identity infrastructure instead of creating a wholly separate agent runtime.
  • Claim: Shared user identity does not eliminate accountability if agent provenance is propagated as a first-class distributed-systems signal. | Evidence: Two Sigma requires agents to populate and continue appending an attribution header, described as similar to a trace ID, through HTTP clients, MCPs, skills, RPC entry points, and downstream spans. This identifies the authenticated user while preserving the chain of agent involvement that led to an outcome. | Implication: Implement agent provenance as a mandatory, tamper-resistant-enough telemetry/control attribute bound to authenticated requests—not as a replacement for authentication—and make its propagation enforceable in the agent harness and service middleware. | Caveat: The header itself is not proof of identity and could be populated by another caller; it must be interpreted alongside the underlying authenticated actor identity, which the speaker says remains non-mimicable through the header.
  • Claim: Safe web access is the major remaining exposure for capable agents because open-web retrieval can both exfiltrate sensitive context and import untrusted instructions or content. | Evidence: The speaker names IP loss, sensitive-data exposure, prompt injection, malware, vulnerabilities, and inadvertent use of unlicensed copyrighted material as risks created by open internet search and fetch capabilities. | Implication: Treat web access as a distinct trust boundary in agent design; do not assume that securing the agent's execution identity secures its information flows.
  • Claim: An enterprise-controlled web index can preserve most useful agent research capability while removing direct external egress from the agent runtime. | Evidence: Two Sigma uses Google's Web Grounding for Enterprise within its VPC/network controls for search and fetch. The service is designed for regulated-industry use, and the speaker says its content is generally fresh within 24 hours, or six hours for regularly updated sites. | Implication: For regulated or high-IP-risk workloads, route retrieval through a governed intermediary with defined freshness and safety guarantees, and explicitly decide which use cases truly require live-web data. | Caveat: The index is not real-time, and its curation/safety layer can still fail; the speaker explicitly says prompt-injection risk is reduced rather than eliminated.
  • Claim: Tool governance must be enforced in the agent's available toolset, not merely documented as a preferred path. | Evidence: Two Sigma denies native agent-harness tools such as Claude Code's Web Search and Web Fetch, then supplies approved redirected access through MCP, CLI, client code, or skills to Web Grounding. | Implication: Remove or block disallowed native tools at the harness/policy layer and provide an approved alternative; otherwise agents and users will bypass the intended security path for convenience.
  • Claim: Broad agent self-service can coexist with centralized governance when promotion criteria match ordinary application productionization. | Evidence: Every employee has the infrastructure needed to build and deploy an agent into their own namespace under their identity. Agents that become team-wide or company-wide are evaluated through standard questions around production support and security. | Implication: Separate personal/namespace-scoped experimentation from shared production agents, and use a clear promotion path with ownership, support, and security requirements rather than requiring central provisioning for every prototype. | Caveat: The talk does not specify concrete approval gates, sandbox restrictions, model policies, or lifecycle controls for user-created agents.

Detailed Brief

Risk-return framing and enterprise operating model

  • Claims: The speaker frames agent enablement as a risk-adjusted-return problem: maximize the value of full user-authorized execution while reducing the risk vectors that normally make security teams reject it.; Two Sigma believes its controls reduced risk substantially without losing expected value, because the attribution layer adds observability beyond what a separate agent identity would provide.; The company provides managed remote agents through multiple interfaces so non-technical users are not constrained to a command-line workflow.
  • Evidence: The talk compares the design objective to optimizing a Sharpe ratio: extracting return while controlling risk.; The managed fleet is available to every user, and agents can be used from mobile, Slack, browsers, and other remote interfaces rather than only from a local CLI.; The speaker says the described system was shipped in the prior year, rather than being a purely conceptual proposal.
  • Caveats: The presentation provides an architectural pattern but does not disclose implementation-level details such as header enforcement mechanisms, authorization-policy granularity, logging retention, incident response, or vendor contractual controls.; The claimed absence of lost expected value is Two Sigma's internal assessment, not a quantified comparative evaluation.
  • Implications: The strongest transferable insight is organizational: agent security can be an enabling platform capability if it is designed into identity, networking, and observability rather than handled as an after-the-fact agent prompt policy.; Managed remote execution also creates a central deployment point for improving agent capabilities and controls over time.

Model strategy and behavioral-data comments from Q&A

  • Claims: The speaker expects self-managed local models to become a preferred destination for significant enterprise inference volume as open-weight models improve.; Behavioral and session data can improve agent configuration, but should remain localized because agent activity may involve sensitive or privileged information.
  • Evidence: The stated reasons for eventually managing more inference internally are cost, model deprecations, and volatility when frontier providers release new models that may degrade an existing workflow.; The speaker says configuration can be adapted to observed usage rather than relying only on a user's hierarchical persona.
  • Caveats: These comments are presented as the speaker's personal view rather than official Two Sigma policy.; No specific privacy architecture, consent process, retention standard, or method for isolating behavioral data is provided.
  • Implications: Model routing and hosting should be treated as operational-resilience decisions, not only quality decisions.; Personalization systems should minimize visibility of session content while using bounded behavioral signals to improve the agent experience.

Notable Concepts & Terms

  • Tethered vs. untethered agents: The talk's metaphor: a tethered agent is constrained by enterprise identity, telemetry, and network controls; an untethered agent has broad capability without sufficient governance.
  • Per-user Kubernetes namespace: The execution isolation and provisioning unit that lets each employee deploy remote workloads, including agents, while retaining their own identity context.
  • Agent attribution header: A propagated request attribute, analogous to a trace ID, used to mark agent involvement and preserve end-to-end provenance even when the agent authenticates as the human user.
  • Actor identity vs. provenance: The authenticated principal establishes who is authorized; agent-attribution propagation establishes whether and how an agent participated in the action chain.
  • Web Grounding for Enterprise: Google's enterprise search/fetch offering used as a controlled, VPC-contained alternative to direct web access, with a cached index and safety controls.
  • Tool denial and redirection: The pattern of disabling native web tools in an agent harness and exposing only approved MCP, CLI, skill, or client-code paths to enforce policy in practice.
  • Managed fleet of cloud agents: A centrally operated remote-agent platform provided to all users, allowing the enterprise to deploy upgrades and controls consistently.
  • Sharpe-ratio framing: The speaker's analogy for maximizing agent value relative to operational and security risk rather than treating enablement and safety as mutually exclusive.

Operator Notes / Why Ken Should Care

  • Assess whether Ken's current agent platform can execute under a delegated user principal while retaining a mandatory, queryable agent-origin signal across every downstream service call.
  • Create a threat-model decision for web access: direct egress, proxy/filter, cached enterprise index, or no web access; document freshness requirements by workflow rather than granting blanket browsing.
  • Audit agent harnesses for native tools that bypass intended security paths, especially direct browser, web-search, web-fetch, shell, credential, and outbound-network tools.
  • Define promotion gates that distinguish personal agents in an owner's namespace from shared agents: accountable owner, support model, security review, observability coverage, and retirement path.
  • Evaluate a managed or self-hosted model-routing strategy for workflows vulnerable to frontier-model deprecation, behavioral drift, or unexpected cost changes.
  • Require that personalization based on agent/session behavior has explicit data-locality, access-control, and retention rules before using it to configure agents.

Source/Metadata

  • Title: Tethered: Our Agents Are Us — Shu Fang, Two Sigma
  • Transcript words: 3260
  • Duration seconds: 1269
  • Timestamp note: No usable timestamps or chapter markers were present in the transcript.
Full transcript 3188 words · 14 min read
0:00

.

0:12

Great. Thanks, everyone, for coming. This talk is called Tethered, Our Agents Are Us. I'm Xu Feng from Two Sigma, and let's get started. So, just a quick explanation. Two Sigma is a little quant fund. I will also take the opportunity to explain that the name ostensibly is not because we have two co-founders who are very online, but because the two sigmas are about the volatility sigma, the small sigma, and the large sigma sum. So by summing together these individual volatilities, we can hedge the risk, achieve differentiated alpha. Now, because we are a hedge fund, I have to give you all this important legal disclaimer.

0:56

You don't have to read it. It just has to be in this. And the TLDR is that I'm not trying to sell you on anything. The views are mine and not necessarily the company's. Any logos, any other companies I mention here are not me endorsing them or telling you to buy their stocks or anything. It is purely maybe coincidental. But that also is meant to segue into the fact that we are an old company. We're 25 years old, and clearly we're a very regulated industry. But we've managed to run an ecosystem where everyone at the company has a cloud agent. And not only that, but these agents run as their own identity.

1:34

So we're going to explain how we got here and why we're actually okay with this. So we're first going to do a little horror movie review.

1:43

If any of you have seen Us, you don't have to pay attention to this. If any of you haven't, the TLDR of the movie is that everyone has these doubles. And these doubles are called tethered. When the doubles decide to run loose and cause chaos and run around with these golden scissors, they're called untethered. And this is going to somehow relate to my talk. So back in June 2025, Cloud Code GA and all that stuff, people started using agents through the local computer, your local machine. And it's very powerful, but one, it was CLI-constrained, and two, it was localized, right?

2:15

We wanted to achieve a world where people could use these agents from wherever they were, whether it be mobile, through Slack, through browsers, but still have the ability to run them remote. And this is important not just because of the capability, but many people, technical or not, are not comfortable fully operating within a CLI. So the question became, how do we actually run these in terms of what identity they run as? The conventional wisdom is that you run these as some machine identity that is attached to your user in some way. You have a Xu, and you have a Xu agent.

2:44

But this quickly collapses, and we found this collapsed because of all the reasons that you can imagine, right? It's very hard to keep permissions in sync. Anytime you're dealing with software licensing, now you have to deal with two licenses. There are certain systems that do not support multiple identities interacting with the same underlying data, stuff like Google Workspace, your emails, et cetera. And then some systems are going to block as a first step. So you're just going over the barrier of entry, and then you also have to figure out how you actually manage the public and private boundaries. So obviously, it's like, why don't we just run these as the user, right?

3:17

How do we run these remotely as the exact same user identity? And as a result, all the capabilities, all the access, all those previous constraints are no longer valid. And we already had infra for this, and I imagine a lot of you do too. If you don't, I would encourage investing in it, which is that you can have a Kubernetes cluster. You have all of your clusters, your regions, et cetera, and you have namespaces for individuals, right? And the reason we had this is because we often already needed this capability, not for the agentic purposes, but for all the automated operations that we need to do that did not suit confinement to someone's local machine.

3:49

So we'd run automated jobs. Code containers usually operate on this principle, research notebooks, et cetera.

3:59

And every single user already had these namespaces existing in every single region, and everything in it runs as the user. A very simplistic way of how this works: some trigger is going into your controller, and it's saying, hey, I need to spin up some compute resources. You have a separate identity service that a sidecar in the pod pulls down from to allow your actual containers to run and mount that identity, and it runs as you. So, of course, there are big dangers with this, right? And the first danger you may imagine is an internal danger. How do you actually differentiate who or what took action, right? You and your Xu agent are now the exact same identity.

4:43

That's why I grew this mustache so you could tell the difference between us for now. But you really want to know that differentiation because certain actions that can be taken, you want to audit, you possibly want to block, and you want to just have the trace, right? You want to have the attribution to determine, hey, was it someone operating as the human, operating purely human actions, or was it the agent identity doing these things? Another danger, and perhaps a bigger one, is we all know that for all of these capabilities, and LLMs in general, it's essential you have access to the external web.

5:26

These are point-in-time mathematical functions that cannot actually update based on current data. So it's open internet access. That's why it's a core capability: web search, web fetch tools, right? The problem is once you have that capability, you leave yourself open to huge vulnerability vectors, one of which is exfiltration risk. This is one we are deeply concerned with in terms of possibly losing IP, exposing our sensitive information, but also certainly the possibility of untrusted content flowing back in. And prompt injection, malware, and vulnerabilities are all big risks there.

6:11

And then something we separately deal with is just the ability to make sure we don't use licensed content without the right copyrights or actual licensing, right? You can map this to the golden scissors that they use in Us. So this is kind of our biggest fear, to be honest. So we are a finance firm, and in finance, there's a concept of risk and return. So when we think about what is the positioning on the risk and return graph, there's huge value in allowing agents to run as you, but also there's very high risk. What we generally want to do is make sure we capture as much of the value as possible, but reduce the risk. We're optimizing that ratio of return over risk.

6:56

Some of you may know the Sharpe ratio. We're looking at that from the perspective of how do we let agents run as users and optimize that return. And the ways we need to do this are to solve those two critical problems. One, differentiating access attributed to the human versus the agent. And two, somehow getting safe web access in place. So the first thing we did is this attribution step, right? And how we did this is we use a header, and we make sure that every single agent continues to append to that header. And this is something we've all hopefully done in some way, right? Trace IDs.

7:38

You've all done this in deterministic code, making sure that your observability stack propagates through a trace ID through disparate systems. How we did it is very similar to how you would do it for a trace ID, except we are dealing with a certain difference in the control vector, which is the agent itself, right? And you can force, using certain HTTP clients, using MCPs, using skills, to make sure that that header initially gets populated, and everywhere else along the way continues to be populated, right? You have a lot more deterministic control over agents and the harnesses and the frameworks than you may think.

8:06

And you can enforce it with some of the already existing primitives. Now, this gets very interesting because this is not only giving us the proper identification of who did something, right? It actually goes beyond that. And no longer are we confined by just knowing the act identity, but we also actually get the full provenance through the system, right? And as we deal with multiple steps in the system, we are able to replay the entire chain of actions that actually led to some end result. So the comparison here is if we had used that Xu agent identity, we wouldn't have this, and we would just know that at some point,

8:50

Xu agent triggered this initial flow into the span, but we don't actually know, hey, those subsequent actions, how do we properly trace back to the origination point? With this header, this trace ID, we get that full propagation, and the actor is still me, right? It's still my identity. And the second step that we needed to fix is this web access, right? A lot of web access these days uses indexes for search, right? I think Cloud Code's native one is Brave Web Browser, and it uses the Brave index. Well, we were like, hey, why don't we see what Google has, right? Google is at its core, hopefully, still a search company, and they do this index generation already.

9:51

And it turns out they actually do offer something specifically for regulated industries like ours that allows you to use their web index, but within your existing VPC, your network controls, right? And it's called Web Grounding for Enterprise. It basically works like this, where it's still within the exact same network boundary where you're probably running your cloud agents and stuff like that. And it offers two core capabilities: search and fetch, right? So those exact capabilities we want to mirror, we leverage that, we have all these guarantees. There's one tiny downside, which is the data is obviously not going to be completely fresh, right?

10:34

And the constraints around this, last I checked, it's fresh within 24 hours, and for more regularly updated websites, it's fresh within six hours. But for most use cases that you may have for agents, that's probably more than sufficient and completely removes this external egress vulnerability vector. Now, the second question is how do we actually ensure the agents use Web Grounding? And again, this is very simple with the existing primitives, right? You just need to make sure that they don't get confused, and you certainly block the access itself, but just for user experience and stuff like that, you need to make sure those tools themselves

11:10

that are already existing and primitive and native to these agent harnesses and frameworks and such are actually blocked, right?

11:21

Again, here's a Cloud Code example. I think every other harness has the same thing: Web Search, Web Fetch. We just deny those tools. It's like, hey, you can't even use these. These are not even in your suite of tools available to you. Instead, we use the redirection going through MCP CLI and actual client code using the supported paths, skills, whatever, to make sure that whenever someone does need the capabilities of web access, it goes through that Web Grounding cached index. So, takeaways from this talk: make sure you tether your agents, right? Letting them run around untethered is very dangerous. We want to tether them, and it's much safer to do so.

12:21

And in fact, if we go back to that initial slide of how we consider this relative to the risk and expected return, because of some of the things we found while doing this, we actually believe we didn't lose expected value while hugely reducing the risk, right? So the index certainly lags, but we get a ton more observability by just using that tagging primitive versus the actual pure identity verification. And I think this is one thing people should really consider, especially people working at companies, enterprises, which is that there are a ton of things happening in the GenAI landscape that are probably scary to us,

12:38

that make your security teams really afraid, that feel like, hey, they are too far on the frontier, right? You're like, I wouldn't run this locally. I wouldn't run an open cloud agent on my local machine with full permissions, right?

13:03

There are all these horror stories and various anecdotes about why this is bad. But in an enterprise, again, you can figure out how to leverage your enterprise resources to actually reduce those risk vectors and get the real value out of the capabilities, and this is where you should be investing that time. So, what we ultimately shipped is this entire framework, right? We have the ability to run cloud agents as user identities because of all of those guardrails and vectors we put in place, and using different interface vectors to actually operate with them so that people who are not comfortable with CLIs can leverage them, but certainly for other cases as well.

13:49

And as part of that, we made sure to ship out just a managed fleet of cloud agents for every single user in this remote fashion that they can already interact with, so that we can continue to deploy and improve what is actually available to individual users. But also the core capability itself of being able for anyone at the company to deploy an agent that runs in the cloud remotely with their full identity is there, and is something we are comfortable with. So, to finish up, everything I talked about actually happened last year. So if you are interested at all in wanting to build and see what we're working on now,

14:30

or even better, if you're like, that was horrible, we could do so much better, we are hiring, and we encourage you to apply. If you have any experience in any of these domains, you can check that QR code, check that link. Yeah, that's it.

14:51

Any questions? What do you think about local LLMs connected to new agents or enterprise? Not the views of my company, but personally I think that is. Is this a question? Yeah, sorry. His question was how do I view local LLMs for enterprise usage?

15:25

And I think local in the sense that we manage ourselves is probably where we eventually want to go for a lot of our token use and inference.

15:33

Because of cost, because of deprecations, because every time a frontier lab drops a new model, you see some degradation. There's too much volatility in that that we don't need to risk as the open-weight models become more advanced and sophisticated. Yeah, so as you can see, the header is not purely differentiating in itself. Someone could certainly populate that, but the actor, the identity itself, will not be me, right? So someone could, I guess, write in that they're using some agent, but the core previous identity itself is not mimicable, not actually interceptable, right?

15:53

So we still have both the originating identity, and all of your identity ecosystems and chains ensure that, but also the header. It's not part of the header thing.

16:31

Yeah, that is separate. The header is just the XCSLM agent, and then you still have some way, you need some way, to actually determine the identity of who's coming. Second part, are you using that in single headers or any other identification that have, or not, if you're writing a single header property? Yeah, generally all of our RPC in some way has an initial entry point that is populatable with that header, and then once it actually goes downstream, same with trace IDs, it's part of the span internally within the system. Can you clarify how the use of Google Cache Index addresses the prompt injection issue?

17:06

Yeah, so the core thing about this index is not only is it a cached index, it has a lot of other controls and safety guarantees around it. It is specifically made for these curated financial, highly regulated industries. So they themselves are doing some of their own curation on top of it. Now, certainly, I think that curation could fail. It's probably done using GenAI. But the prompt injection risk is far reduced because everything still remains internal. Yeah? How important is behavioral data for these pages? And so what are the sources of your acquired data? Do you mean how they're being used? The behavioral data for people that are using it?

18:15

Yeah, I think it's critical in the sense that that behavioral data is something we can further configure based on, right?

18:24

We do try to ensure that not everyone at the firm can see what your agents are doing, right? There's stuff certainly work-wise, but also more sensitive information that might be privileged to you. So your session data is kind of localized. Now that behavioral data in the session data is very powerful because it can define additional configuration that can be applied to these agents for the purposes of making the user experience better. So we try to leverage that to figure out what to configure further, not just based on someone's hierarchical persona, but actually based on their usage to make sure that their experience continues to improve based on what they're doing.

18:52

Almost out of time. Me and my, sorry, I'll take the last question. I was just going to ask, you were saying you have a process for letting individuals create their own agents. Yep. How do you go about that? Can anybody create their own agents?

19:27

Is there a process for provisioning certain individuals? Is there a process for promoting agents to be used across the company versus individual teams? Yeah, for building our own agents, we use some of the existing frameworks for agent building. Obviously, all the GenAI harnesses are very good at using those frameworks to build agents. So you have a lot of agents proliferating based on that. In terms of provision, everything is already provisioned. All this is, every single user at the firm has all the necessary infrastructure in place. So that's not really a worry. They can build an agent, deploy as necessary into their namespace running as their identity.

20:04

The aspect of how do agents then become a universal company-wide or larger beyond individual users agent goes through your standard kind of mechanisms, right? Like, hey, is there going to be proper production support? Is there the right security? It's like any application you might develop. Yeah, so me and my colleagues will stick around here if anyone wants to talk further. I guess if you're sticking around, I can also take more questions, but thanks for coming to this talk. Thank you. Thank you. Is there the right security? It's like any application you might develop. Yeah, so me and my colleagues will stick around here if anyone wants to talk further.

20:44

You know, I guess if you're sticking around, I can also take more questions, but thanks for coming to this talk. Thank you. Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note