SPEAKER_00
Music Good afternoon everyone. I know it's the last day, hopefully you're still holding on, not too tired of talking about AI yet. Yeah, today's going to be a little different. So I work for WorkOS, my name's Garrett, I run product for the team. I'm not exactly talking about our product today, so I'll do 10 seconds about WorkOS, just to get that out of the way since my company cares about that. We do enterprise platform features, we're developer platform. The quick and easy is if you've ever logged into Cursor, you've used WorkOS.
SPEAKER_00
Whether that was username and password or you went through your enterprise IDP, we power enterprise platform features for the likes of Cursor, Anthropic, OpenAI. Today though I'm going to talk about something a little bit different. I'm going to talk about how we operate internally and things that we've built to make ourselves more productive. So I imagine most of you are probably engineers or on the technical side. You might have questions about things about your company, about how customers are using the product, how things are working.
SPEAKER_00
Your go-to-market teams or your support teams definitely have questions about how customers are using your product, trying to figure out answers to questions. You might have Retool inside your company, you might have built dashboards and things like that. Of course those can be fairly rigid, right? You build a very specific thing, someone comes and says, oh actually, but I need this extra bit of data. I need to find out, answer this different question that the dashboard doesn't answer. And so either you go and build that, you change it, right? We see a workflow that looks kind of like this, where someone has a question, often about the business.
SPEAKER_00
They may not be technical enough to go answer it themselves. They often need something like SQL or someone that has access to the data. They have to explain their question, why they need it answered, context to answer it. They wait, someone like you has to go answer the question, provide that data back to them. Did you actually answer the question? Did you provide enough detail? Oh no, that's great, but I actually need the next layer deeper. Got to go back and forth. You probably share that in Slack, a one-off. Doesn't really scale very well. We have this problem. If you didn't, we have this problem every day.
SPEAKER_00
And so we built a tool called Studio that serves as an internal workspace where people can answer questions and build these apps or dashboards themselves. So I'm going to show you both what it looks like to build this out. I'll also show you a few examples of some of the tools that we use every day inside of Studio. And then I'll talk a little bit about how it works under the covers. So I don't get this completely wrong. I have a little prompt here already. But so, a common thing we have, we do a lot of marketing. We're doing podcast advertisements. We're doing Google ads. We're getting people to come to WorkOS site, whether that's our blog, our docs, or a marketing site.
SPEAKER_00
And then we want to know what content are they reading and what's effective, right? What is someone reading on our site and then converting to actually using the app? So I want to know, hey, what content leads to the most new teams? We call our customers teams internally. So leads to the most new team creations. So I can fire this off and we will, Studio starts operating. It basically says, okay, I want to find this data. I need to look at what internal resources do I have access to? So it knows that it has access to my Linear, my Notion, my Snowflakes. We have these data sources that we connect to.
SPEAKER_00
And then basically understands how to use these tools and starts to run queries. So in this case, it's going to run a bunch of Snowflake queries, which is our internal database. It's where we store a lot of this data. And it's going to go through, figure out the schemas, look at the tables that it needs to, and do it. While it's doing this, since it might take just a minute, I'm going to talk a little bit about how it works under the hood. So you can either go to our internal Studio dashboard or in Slack, we have a Slack bot. So you can ask questions of Studio. That kicks off the process.
SPEAKER_00
We run an API behind the scenes that takes that, parses it, and then runs it through LangGraph, which is an agent that's both tied to LLM, which in this case we're using Opus, along with the tools and the guidance layer for how it should interact with these systems. So we have this integration proxy to the data sources, primarily like Snowflake, Linear, and Notion are the tools that we use. And this guidance layer basically defines rules around how you should query this data, the context you need to successfully query the data.
SPEAKER_00
Our Snowflake is a pretty sprawling set of databases, so it needs to get context around what's the representation of a customer inside of Snowflake? How do I join tables in a way that's effective? So the agent drives all of this, makes queries, the LLM runs, and then of course it provides back answers or updates widgets, which I'll show you in a minute. And then we store a lot of that state today in Convex as a way to locally store this information so it's preserved over sessions. So we go back. Cool, it looks like we've actually gotten a bunch of data here. Let me make it a little bit bigger for you.
SPEAKER_00
So we can see obviously people go to our homepage, people look at the pricing page. We can see the blog posts that are most effective for driving team signups. Changelogs and docs. And get the summary. Okay, but this is great, so it answers my question, but I want this to be a long standing thing that I can reuse. So can you build a table of this that lets me see this data over various time slices? And so here it's not just run the queries, get the answer. But I actually wanted to build a reusable tool that I can share with my teammates that I can use in our weekly syncs. And so it's going to think through how to do this.
SPEAKER_00
And then it's going to go and build what we call a widget. A widget is in this case basically sandbox code that runs. And it's both the UI, the APIs, and the query necessary to power a fully usable tool. So this is going to think for a minute as it actually creates the widget. So can you build a table of this that lets me see this data over various time slices?
SPEAKER_00
And so here it's not just run the queries, get the answer. But I actually wanted to build a reusable tool that I can share with my teammates that I can use in our weekly syncs. And so it's going to think through how to do this. And then it's going to go and build what we call a widget. A widget is in this case sandbox code that runs. And it's both the UI, the APIs, and the query necessary to power a fully usable tool. So this is going to think for a minute as it actually creates the widget. I have another version of it that I can show you. We'll see if they look the same across instances. But this is one that I had pre-built before this.
SPEAKER_00
So it basically gives me this data of teams over different time spans. What content is driving those signups. And it's live, right? If I run this, it's going to rerun that query. It's going to give me the data for different time slices. I can make it full screen here, show it.
SPEAKER_00
And so it's going to think through this. Look, and we get a pretty similar, slightly different view. But this one actually has category filters. So I can look at based on the kind of content that it's running. And see what are the most effective, we have a lot of blog posts. Which are the most effective ones that drive traffic to the platform? So we can tailor our content effectively. But this is useful for a lot of things. For example, Radar is one of our internal products. It's a security product that blocks bots and bad actors. And sometimes customers say, hey, why did this user get blocked by Radar? Right? Can you help me understand?
SPEAKER_00
And so we've built some of these dashboards and stuff ourselves. But typically that involves having to go through, a lot of our GSEs are sharing SQL queries. To be able to run and answer these questions. But instead of having to do all that, I can just, I've already built this widget. That has the APIs or the queries hooked up. And I can just do a search for myself in this case. My personal email. And in this case, it's running a real query against our database. To actually pull this data and look at it. And so you can see the conversation history here of me talking with it. Hey, can you build me this dashboard? Runs a bunch of queries. It actually messed up at first.
SPEAKER_00
But I, saying, hey, can you, there seems to be an issue. Can you keep going? It did it. And the last thing was it had a visual UI bug in the type column. So it's like, hey, can you fix that? Visual bug. I'd like it to be one nice little column. And so here we can see for Cursor, here's one of our customers that uses Radar. Here's all the times that I logged into Cursor with my personal email. And whether I was blocked or not. I had a test here where I blocked myself in one of our test environments. And so this becomes a self-serve tool that our support team can use to look this stuff up. And so this has been really powerful for our support team.
SPEAKER_00
They use this in Slack all the time because they don't need different customers at different specific issues. And they can say, hey, can you go find me all the sessions that this customer has so I can find out what went wrong, right? And so we're not trying to build, we don't need to have some sort of platform team or data team building these dashboards that are going to be used and need to be constantly modified. Our support team can basically, if it's a one-off, get the question answered themselves. And if they're finding that they're asking the same question a lot, they can build these and then we can share them internally to other folks.
SPEAKER_00
And so we build out our own dashboard and tooling in a self-serve manner. So I'll wrap up with what did we have to do here to make it useful and reliable? So there are three things that became really important in building this. The first is sequencing. So this is how should the agent approach when it gets a new question, when it gets a prompt, how should it do this?
SPEAKER_00
So we make it run a lot of pre-flight checks. So this is are the tools connected correctly? Do you have enough context to better answer the question? If not, ask clarifying questions. And then determine, run through a checklist to determine the tools that it should actually use to call. We actually, at the time it decides to invoke a tool, that's when we inject context around how to use the tool. So for example, if I show some of the tooling that we use, like for Snowflake, for example, we have this context that we embed. And it's not trivial.
SPEAKER_00
It's fairly long because it encodes basically the schema of our internal database and how to understand how do you connect teams to the environments, to the resources that they're using. And so this gets injected at runtime when Snowflake is being adjusted. I saw someone earlier in a talk talking about how you don't want to preload all that context of all your tools because it blows out your context window. Second is layering. So we have the base prompt that Studio uses to show it off with. We have the defaults. And then we have org rules around in a given setup for a given tool, there might be a specific context.
SPEAKER_00
If someone's going and editing a tool, we want that context to be maintained. And then last, we tell the LLM to specifically distrust knowledge around our product often, just because sometimes the model training is using outdated data. Our product changes very quickly. Things are moving all the time. And so we actually use, we tell it to no, no, no, go for primary sources, look up data in our docs and things like that. So it's basically pre-validating its work before it's deploying it into a database. We have the defaults. And then we have org rules around in a given setup for a given tool, there might be a specific context.
SPEAKER_00
If someone's going and editing a tool, we want that context to be maintained. And then actually last, we tell the LLM to specifically distrust knowledge around our product often, just because sometimes the model training is using outdated data. Our product changes very quickly. Things are moving all the time. And so we actually use, we tell it to no, no, no, go for primary sources, look up data in our docs and things like that. So it's basically pre-validating its work before it's deploying it into a database. So it's actually running queries that don't just rely on what the model knows about Work OS necessarily. And then last, validation.
SPEAKER_00
So if it's going to write a query to our Snowflake instance, we have it always run the query and validate that it gets data back. Many times they can have a valid SQL query, but that returns zero data. If it doesn't notice that, it's not very useful. So it actually runs queries, validates them before it hardcodes them into widgets and things like that. So it's basically pre-validating its work before it's deploying it into a dashboard. And then, yeah, we run obviously evals when we're developing the product. Evals are very useful. I don't have time to go into how to develop and design evals. But we use evals in both our staging and production instances.
SPEAKER_00
We treat all of that the same. So that way we get the same experience when we're developing Studio versus when our teammates are using it.
SPEAKER_02
[SPEAKER_00] So yeah, that Studio, it is our way of basically being able to answer any question about the business.
SPEAKER_00
Anything I can answer for y'all?
SPEAKER_02
[SPEAKER_00] Yep.
SPEAKER_00
[SPEAKER_02] Sorry, go ahead. Did you have to do a clean up on your Snowflake? [SPEAKER_02] No, actually, and we have, there's just one specific problem. Where the connection between a customer entity to the users that they have or whatever is four joins deep because of reasons. And every new employee has to learn if you want to do that, you have to copy and paste this join block and use it. It's we tell Studio about it once. Right? It knows how to do that every time. LMs are quite good at interpreting table schema pretty well. And so there's a lot of stuff they can get. If you have pretty self descriptive column names and stuff, it can figure out.
SPEAKER_00
But again, we do have that context block that we provide because it does matter for example, in Radar, we have attempts, which is people trying to log in. We have detections for when things and it's like, oh, you need to join those, you know, these tables join this way. And just by telling it that it can basically run effective queries for that data. So there is some information you want to provide. But surprisingly good. You don't need to rag all this stuff, you know, we have no rag database for us in all of this.
SPEAKER_01
[SPEAKER_00] We're just invoking tools directly, with just context on top. [SPEAKER_00] Question? [SPEAKER_00] Yeah. I just think that's exactly what we do. We just got a context thing that tells it how to do all the joins. Yeah. And then it knows, and then it knows those quirks. One of the things we've been looking at though, we do something similar, but is those queries get generated. [SPEAKER_01] So, I mean, it's so widgets, you call the widgets. Do those get audited or governed by anyone? Because our concern is that someone generates a query, their skill gets it wrong, and then it becomes a truth, and everyone thinks it's true, and no one's ever checked it.
SPEAKER_01
So do you have anything like that? Yeah, I mean, there's definitely a little bit of you know, there's always some trust but verify, you know.
SPEAKER_00
I've actually been pretty impressed that the hit rate on the cross this is very, very high. I think there's a category of which you can actually embed into the context of you know, make sure you only pull non-deleted entities, right? Make sure you pull things in an active status, right? There's kinds of things that your data probably have, you have consistency around of a status column, right? And those are the kinds of things that I've seen that elements will miss if they don't know. It's how many users have this resource? And it's just doing a count group by customer ID, right? And it's oh no, you actually need these filter columns.
SPEAKER_00
But if you have that in your context, that's the kind of thing that protects against a lot of those issues. So I find that that has removed a lot of the problems for us. And then, you know, if it misses from there, it's you know, it's making a bigger error. That's pretty obvious. Can I send you a widget of data from multiple tools? [SPEAKER_04] So you can mix and match it?
SPEAKER_04
Yeah. [SPEAKER_00] Yeah.
SPEAKER_00
So we have, you know, just those few, we're adding more of those connections, but yeah, it can pull from different tools and combine that data into one interface. And how would you refresh the data afterwards? [SPEAKER_05] Would you need to know how to replay those tools sequentially? [SPEAKER_05] So it's actually, the widgets are actually code.
SPEAKER_05
[SPEAKER_00] So it's writing JavaScript that is making the underlying API calls to that service through the tools. [SPEAKER_00] So once the widget is created, it is reliable.
SPEAKER_00
It is not the LLM running the tool. So when I hit refresh here, this is actually just requerying data from those tools. So the LLM is not involved once the widget is developed until I go back and say, hey, can you make an adjustment to this widget? Can you add this column or whatever? So the actual final product is very reliable in that regard. Right. And if you need to pass a different data, it's just an input argument to that. [SPEAKER_03] Yeah. You know, here it's like I'm giving an input and that's just being fed into the query, any sort of JavaScript do it.
SPEAKER_03
[SPEAKER_00] Right.
SPEAKER_00
So when I hit refresh here, this is actually just requerying data from those tools. So the LLM is not involved once the widget is developed until I go back and say, hey, can you make an adjustment to this widget? Can you add this column or whatever? So the actual final product is very reliable in that regard.
SPEAKER_00
Right. And if you need to pass different data, it's just an input argument to that.
SPEAKER_02
[SPEAKER_03] Yeah.
SPEAKER_00
You can do here it's like I'm giving an input and that's just being fed into the query, like any sort of JavaScript do it.
SPEAKER_02
[SPEAKER_00] Right.
SPEAKER_00
So when you're doing these user input things again, right, you're not relying on the LLM to parse that correctly. It's writing declarative code. Yep. How do you respect user access to the data? [SPEAKER_02] How do you what? Respect user access to the data. [SPEAKER_02] Oh yeah. It's a great question. Today, the integrations are user based. So I'm connecting Snowflake and Linear and Notion myself. That's something we're actually working on changing because that's annoying. You don't want every employee to have to necessarily do that. And there are cases where maybe you don't have a Salesforce account, but you probably should better read certain Salesforce data.
SPEAKER_00
So actually working on the thing that drives these integrations, we have a product called Pipes, which does third party integrations. So we're actually using our own product under the hood here. We're building out organic, what we call org connectors.
SPEAKER_02
[SPEAKER_00] So one person sets up the connection and then can set rules about what's the default level of access when people are querying that. [SPEAKER_00] So for example, you could say by default in Linear, you get read only access by certain people based on roles in the studio application, they get admin or edit access or something like that. [SPEAKER_00] So we're building that permissioning layer on top of it because doing the per user login is annoying.
SPEAKER_00
Cool. Yeah, sure. Yeah. [SPEAKER_02] How do you handle the costs? [SPEAKER_02] So obviously using Opus, is there any action that you do? [SPEAKER_02] The widgets themselves, once they're generated, are declarative, so you're not paying the LLM cost every time. But honestly, for us, we're willing to pay the cost for the questions being answered. [SPEAKER_00] Opus outperforms other models so much that I wouldn't trade the cost off in a way that would trade off quality that we wouldn't deem acceptable. [SPEAKER_00] So I think we're pretty willing to spend the money.