Every

Anthropic Just Bought a Dev Tools Startup for $300M. Here's What Its Founder Told Me.

1036 summary words 5 min summary Watch video

Start with the signal

5 min read

Summary

Summary: Anthropic Acquires Stainless for $300M - The Future of AI and APIs

Main Topics

  • Model Context Protocol (MCP): The emerging standard for connecting LLMs to external systems and APIs
  • API Architecture: How Stainless enables computers to communicate with each other
  • AI-Driven Automation: The vision of agentic AI handling complex, multi-step tasks
  • Code Execution as a Tool: A proposed alternative to traditional MCP tool calls
  • Security and Permissions: Challenges in allowing AI safe access to powerful APIs

Key Points

Current State of MCP

  • MCP is designed to give language models native access to tools and services, similar to how websites serve human users
  • Current implementations are not scaling well due to context window limitations
  • Exposing hundreds of API endpoints to an LLM burns through context budget with minimal benefit
  • Most successful MCPs are heavily handcrafted rather than automatically generated from APIs

The Core Problem with MCP

  • A typical business task (e.g., refunding a customer) requires navigating 5+ different SaaS apps and 15+ clicks
  • Modern LLMs should theoretically handle this, but practical implementations are restricted to only a few tools
  • Exhaustively exposing an API (like Stripe's hundreds of endpoints) creates context bloat and confusion for the model
  • The context impact is prohibitive: translating a full OpenAPI spec to MCP tools consumes hundreds of thousands of tokens

Design Principles for Effective MCP Servers

  • Keep the number of tools relatively small (not comprehensive)
  • Use precise, specific tool names and descriptions
  • Minimize input parameters and describe them concisely
  • Return only necessary data from responses
  • Implement feedback mechanisms to learn what's actually working
  • Use dynamic tools (list, get details, execute) to scale without context bloat
  • Conduct rigorous evaluation across different clients (Cursor, Claude Code) and models

The "Cyborg" Future: Code Execution Alternative

Rather than exposing 50+ tools, the vision is giving models just two tools:

  • Execute Code - Model writes TypeScript/Python that uses the SDK
  • Search Docs - Model can look up documentation when unsure

Advantages:

  • Minimal context window impact (~1,000 tokens upfront vs. hundreds of thousands)
  • No context cost for pagination or iteration loops
  • Fast execution (runs on same server as API, no model round-trips)
  • Models excel at writing code when SDKs are well-formed and predictable
  • Dramatically reduces hallucination errors through static type checking

Challenges:

  • LLMs struggle with library versioning and dependency management
  • Requires well-designed SDKs (which Stainless specializes in)
  • Security concerns with arbitrary code execution

Real-World Implementation at Stainless

Alex uses MCP servers on the business side:

  • Customer Analysis: Queries company database, cross-references with HubSpot, looks up notes in Notion, checks transcripts in Gong
  • Knowledge Repository: Uses Claude Code to automatically save customer quotes to a Git repo in structured format
  • SQL Queries: Iteratively builds analytical queries for board prep and caches them for reuse
  • Support Automation: Experimental use of Claude Code to attempt fixes on bug tickets (50% success rate, but improves overall efficiency)

Security Model

  • Current approaches try to limit what's exposed through MCP, but this is insufficient
  • Real security must happen at the API layer using OAuth with granular scopes
  • API providers need to build proper permission systems, not restrict tool exposure
  • For code execution sandboxes, limit network access to specific endpoints (e.g., only api.stripe.com)
  • Gradually expand with proper controls rather than rushing full integration

The Path to Automation

  • Many one-off tasks that AI performs manually become enduringly useful
  • When an AI solves something multiple times, it should be able to commit code to production
  • This mirrors how software teams evolve: manual → repeated → automated
  • Interface should transition from chat (exploring) to dashboards (daily production use)

Notable Quotes

> "APIs are the dendrites of the internet. Neurons can't think without connections, and the internet can't work without APIs."

> "You've just burned through your entire context budget... it's confusing to the model. It's just too much to hold in your brain at one time."

> "The future of AI is cyborgs - part LLM, part traditional CPU code."

> "The code execution tool is going to become the most widely used tool."

> "Good writing is hard" - (on the tension between tool simplicity and descriptive precision)

> "You don't have to think about it too hard in advance. You can move things around later." - (on letting AI work with unstructured data)

> "There might be a version you could do today or tomorrow for individual developers... YOLO mode is what drives adoption."

> "When you find interesting customer quotes, put them in this folder and give the full citation, so next time you're asking questions, you don't have to go searching through the MCP servers again."

Takeaways

  • MCP isn't ready for comprehensive API exposure - Current approaches fail at scale due to context limitations and complexity
  • Code execution is the better model - Models should write code using SDKs rather than choosing from hundreds of predefined tools
  • SDKs are critical infrastructure - Well-designed, typed SDKs are essential for LLMs to work with APIs reliably
  • Security belongs at the API layer - Proper OAuth scopes and granular permissions matter more than restricting MCP tools
  • Handcrafting is currently necessary - Until we solve the fundamental design problem, successful MCPs require significant product and engineering work with customer research
  • Evaluation systems are lacking - Most teams haven't solved how to measure if their MCP implementations actually work in production
  • Adopt early with relaxed constraints - Like Stable Diffusion vs. DALL-E, or Cursor vs. Codex, permissive early releases drive adoption more than locked-down "safe" launches
  • Cache knowledge effectively - Using git repos and markdown to store discovered information (queries, insights, quotes) reduces repeated MCP server queries
  • Iterative refinement beats perfection - Unstructured knowledge that AI can work with immediately beats perfectly-structured databases that take months to build
  • The enterprise market may lag - Individual developers and smaller companies will adopt powerful AI automation first; enterprises will follow once risk is proven manageable
Full transcript 8655 words · 76 min read
0:00

SPEAKER_01

The internet runs on computers talking to each other, but its entire architecture was built for a pre-AI world. Now we're trying to hook AI up to the internet with MCP, Model Context Protocol, which turns any website or web service into a set of tools that an AI can use natively to get work done. And the software companies that learn how to do MCP well are going to win over the next decade. That's why I brought Alex Rattray, the founder and CEO of Stainless, onto the show. Stainless' job is to help computers talk to each other. They make the API and SDKs for all the big companies that you know about, like OpenAI and Anthropic, and they're starting to build MCP servers too. So Alex and I get into the nitty gritty of what the future of MCP looks like, how to design good MCPs, why MCPs are actually really hard to scale and possibly insecure. And we try to figure out together what a better model for allowing AIs to use the internet might look like. This is a great episode. Alex is a good friend of mine. Let's dive in.

0:05

SPEAKER_01

Alex, welcome to the show. Thanks, Dan. It's really exciting to be here.

0:17

SPEAKER_01

It's good to have you. So for people who don't know, you are the founder and CEO of Stainless, which is the API company. You make APIs for companies like OpenAI and Anthropic, and just name your big company that you might use their API. Stainless is probably behind it. Before that, you worked at Stripe doing their API, surprise. And before that, most importantly, we were very good friends in college, and we remained good friends. And we were both starting companies in college. I'm a tiny investor in Stainless. But it's been really fun to watch your journey and get to hang out together so much over the years. And I'm just very excited to bring you on to talk about AI and what you're doing at Stainless.

0:22

SPEAKER_01

Thanks, Dan. Yeah, it's been really fun over the years. When we were in college, I was working on a startup. You were working on a startup. You had a conference room at a venture capitalist office as your office, and you let me crash there with my co-founder and team. And we were just on the other side of the conference table, hacking away into the evening. And very fond memories of those days. And these days, it's not every evening, but on the weekends, whatever, the same thing is still happening. And you don't see that every day. And it's really a nice feeling. And it's been great to see everything happening with every on the way.

0:27

SPEAKER_01

Thank you. As I say, it started from the bottom. Now we're here. And yeah, I mean, the thing that I always say when people ask me about you, in order to embarrass you, I just talk about how you're the only person that I know of who has consistently run barefoot through the streets of Philadelphia. Because when we first met, you were not a fan of shoes and you were a fan of running. You want to talk about that?

0:39

SPEAKER_01

Yeah. It wasn't that I didn't like the concept of shoes. It's that I couldn't find a good pair. And at a certain point, it's that I was running through Nikes and they would bust open every few months. I think what was actually going on is that I had really wide feet. And I was buying probably narrow shoes. But they would choose constantly ruined. And on a college budget, it's just, this is it? This is no good. And eventually I decided, okay, the longer you wear your shoes, the more worn out they get. But the longer you just wear your feet, the tougher they get. So... [SPEAKER_02] The longer you wear your feet.

0:52

SPEAKER_01

[SPEAKER_02] Try it out. Try this at home. What could go wrong. I actually currently have a really annoying splinter in one of my feet. And so don't actually try this at home. But... [SPEAKER_02] Are you still running barefoot?

1:14

SPEAKER_01

[SPEAKER_02] No, no. This is just from around the house. [SPEAKER_02] I see. Dangerous. [SPEAKER_02] Yeah. But see, that's the thing. If I had been going around on the asphalt without socks on, then my feet would have been tougher and I'd have no splinter. [SPEAKER_02] So when you're not running barefoot, you're running Stainless. [SPEAKER_02] So you're running Stainless. And so how many people are you? You know, you're around 50, right? Just about. Yeah.

1:33

SPEAKER_01

[SPEAKER_02] That's pretty wild. And you started Stainless in a pre-AI world. And now we're in an AI world. And I think you have some ideas for what the future of AI is going to be and maybe how APIs fit into that, maybe how MCPs fit into that. Do you want to paint a little bit of a picture for us about where we're going? Yeah, I would love to. So to start, what's an API? Not everybody's familiar with that. So it stands for Application Programming Interface. There will not be a quiz. Right, right, Dan? No quizzes? No, no quizzes.

1:50

SPEAKER_01

Great. But basically, it's how one computer program talks to another computer program. It's how computers talk to computers, how apps talk to apps. And APIs are the dendrites of the internet. Dendrites are where your neurons connect and actually exchange information with each other. So if you have two neurons in your brain, but they're not talking to each other, you're not thinking, right? There is no thought happening in a brain without connections between neurons. And if you think about the internet, if all these servers in the cloud weren't talking to each other, you wouldn't have internet, right? There's nothing going on. Programs, internet software is doing nothing without APIs, without connections to other programs. And so it's really fundamental to the mesh of pretty much all modern software. Everything that we think of when we think about technology at this point, APIs are at the heart and center of that, just like dendrites are the center of the mesh of the brain and how we think. And Stainless's mission from day one was to make it easier for computers to talk to computers. And it's the long running trend of technology to have more automation, right? Automation is what we mean when we say we're going to apply technology to that. We're generally going to be making things more efficient. And APIs are how most business to business interactions in some format or another become real, become automated. And what we see with the rise of AI is that there is a new computer that has entered the chat, right? There's a new kind of system that can talk to other systems, or at least we would like it to be able to. You used to have either humans interacting with a computer through a user interface, a UI, or a computer acting with a computer through...

1:57

SPEAKER_01

[SPEAKER_02] Okay, we're going to apply technology to that. We're generally going to be making things more efficient. And APIs are how most business to business interactions in some format or another become real, become automated. And what we see with the rise of AI is that there is a new computer that has entered the chat, right? There's a new kind of system that can talk to other systems, or at least we would like it to be able to. You used to have either humans interacting with a computer through a user interface, a UI, or a computer acting with a computer through an API. And now we have LLMs interacting with computers, right? And what's that through?

2:01

SPEAKER_01

And I'm sure anyone familiar with every and his regular listeners is going to be familiar with MCP, model context protocol, which is a system for connecting LLMs to computers, broadly speaking.

2:08

SPEAKER_01

[SPEAKER_02] And it's an area that we're investing in at stainless. It's really, I think, part of our core mission of making it easy for computers to talk to computers. And we've invested a lot of time at stainless. The core product that we first brought to market is software development kits, SDKs. And so these are ways of saying, okay, Stripe has this great REST API. You can send JSON over HTTP and get back JSON over HTTP. And if you want that to be really convenient, you're going to use the Stripe Python library, the Stripe Python SDK. So you can go, if you're a Python developer, you'll go pip install Stripe. And then in your application code, you'll write Stripe.customers.create. And all of a sudden, you have a nice new customer object in your Stripe database. And you're off to the races, or Stripe.charges.create in the old days to charge a credit card. And SDKs are what gives developers that easy way to interface with an API. What's the thing that gives LLMs an easy way to interface with an API? And you might say MCP, and in a sense, you'd be right. But what we're seeing so far, as MCP is rolling out into the world, and people are experimenting with it and trying it out, is that it's not working so great. It's difficult to deliver on what I see as the core vision of what's so exciting about MCP, which is just like a dashboard and a user interface lets you click around, see a bunch of stuff, fill out forms, click buttons, do things.

2:14

SPEAKER_01

Anything that you would do while you're interacting with the software, or you do through the user interface generally. But LLMs interacting with through MCP, it tends to be much more restricted. You can only do a few little things. There's usually not a ton of tools that you're going to be exposing to the models.

2:22

SPEAKER_01

And just to stop you there, so I think what I'm hearing you say is what MCP does is, just like a website is built for humans to be used, MCP is the equivalent, and you can think of it in certain ways, of exposing a set of tools for the model that it can use to perform certain functions. Just like you might click a button on a website, the MCP gives to the model a bunch of things it can click on or use to get work done. So an example might be, a Gmail MCP has like a send mail tool or a compose mail tool or a read inbox tool, that kind of thing. And instead of a human going on the Gmail website and doing it, it's the LLM that's essentially logging in and using it itself. And it's a native interface for language models. But you're saying that that's not working that well. Can you tell me more about that?

2:29

SPEAKER_01

[SPEAKER_02] Yeah. So let's start actually with what I see is the big vision of MCP. And in some sense, the big vision of agentic AI in the first place. And I'll start with the most pedestrian example you can imagine. It's going to be funny given some of our context. Which is, let's say Dan walks into my store and buys a pair of stripy socks, and maybe a few other things. And then the next day, I hear back from Dan that there's something wrong. Unfortunately, it happens. And I turn to someone on my team and I say, Hey, can we refund Dan for those stripy socks he bought yesterday, and send him a discount code for the next time he comes in with a little thank you note. Because we like to take care of our customers. This is the most normal thing to do in software, some little task like this. And what the member of my team would be doing would be opening up their internal admin and looking around for some things. They might go to the Stripe dashboard and try to look through the list of payments or the list of transactions or orders, and try to find one that has someone named Dan. Dan, I don't know, there might be a bunch of Dans. Try to look through the list of products in the order and see whether there was some stripy socks in there. That might be a few clicks required, depending. Find the right one. Then go to the screen where you can create a refund, create a refund, make sure it's the right amount. Then go and create that discount and then take that discount code and send it over to some other SaaS app where you log in to send some mail automatically. And of course, if you step away from the consumer version of this to a business to business context, you might be going into Salesforce and sending a Slack message to an account administrator, an account manager, and so forth. And in the normal course of work, it's the most normal thing in the world to be doing, having one task involve going through five different apps each time, 15 different clicks and scrolls and loading spinners, just to do one simple thing. And the promise of agentic AI is to be able to take that same prompt I just said and type it into ChatGPT or Claude or whatever, and say, Hey, chatty buddy, can you help me refund my friend Dan. And just have the AI go off and do that. And basically go through these five different apps and the 15 different screens, and the various different button presses to complete the task and then come back and say, Great, it's done. That in order to do that, there's only so many tool calls you have to make as an AI model to perform that exact linear chain of events. It's somewhat tractable. But if you think about this in the general case, you want the LLM to be able to do what you want your agentic AI to be able to do anything that a human operator would have done. And you would want them to be able to do it without having to wait for a bunch of JavaScript to load on a website or anything like that. And that means you need not only the Stripe

2:34

SPEAKER_01

[SPEAKER_02] button presses to complete the task and then come back and say, Great, it's done. That in order to do that, there's only so many tool calls you have to make as an AI model to perform that exact linear chain of events, it's somewhat tractable. But if you think about this in the general case, you want the LLM to be able to do what you want your agentic AI to be able to do anything that a human operator would have done, and you would want them to be able to do it without having to wait for a bunch of JavaScript to load on a website or anything like that. And that means you need not only the Stripe create refund tool and the Stripe list transactions tool, and the Stripe list products and look up customer and create discount tool, you need not only those tools, but you need everything that you can do in the Stripe dashboard, which is basically everything that you can do in the Stripe API. And that's actually a lot. There are hundreds of different endpoints that you have access to in the Stripe API. The Stripe dashboard is actually massive. It's a huge application. And if you were to take that list of tools today and go to an LLM and say, Hey, here's our MCP definition for all of this, here's a create refund tool, here's a create transactions tool, so on and so forth. And you tell it all about those tools, here's the description, here's all the different request properties that you can send, here's the response properties you can get back, here's all the documentation for each of those things. Everyone listening to this should already know you've just burned through your entire context budget. That's maybe hundreds of thousands of tokens just there, just in translating the Stripe OpenAPI spec directly over to MCP tools. And today's models not only can't handle that amount of context, it's a poor use of context because you have a lot else going on. But it's also confusing to the model. It's just too much to hold in your brain at one time. And that's just the Stripe part of it, right? Because what you're really trying to do is enable your operators to do anything they would normally do. And again, that spans many different SaaS tools, right? In the course of one interaction, it might be five, in the next interaction, it might be a different five. And so if you think about every single SaaS tool that your business uses on a daily basis to get your work done, ideally, you would want every single one of those tools to be exposed to your operators in their AI chat, with every single tool available in there, with every single nook and cranny and corner case available, so that you can do anything through AI. That's the vision. Now, there's a lot of problems with that. The biggest one that I mentioned is this context window limit. But you also have all sorts of security and permissions problems because you don't want the AI to color outside the lines and say, Okay, in addition to refunding Dan Socks, I also refunded every customer for all transactions ever, and then I sent a bunch of money to my own AI bank account. And so there's more to the challenge, but that's the vision I see.

2:39

SPEAKER_01

[SPEAKER_02] Dan Socks, I think the place we started there was you said it's not working. But I don't think that's the reason why it's not working today, right? Or is that the reason why it's not working today?

2:45

SPEAKER_01

[SPEAKER_02] Dan Socks, what people do with MCP today is sometimes they'll try to expose all parts of their API. And the way people build MCP tools is generally speaking, they have an underlying API, usually a REST API, and they wrap different parts of that—different endpoints, different operations—in MCP tools. And you can do that in a one to one mapping, or you can handcraft things for the MCP. And today, in order to succeed, people are finding that you really have to handcraft it to the MCP, to the LLM. You have to say, Okay, I'm making one specialized tool to look up a customer and refund their transaction based on a description.

2:52

SPEAKER_01

[SPEAKER_02] Dan Socks, there's all these decisions that you have to make where you need to have the ergonomics of the model and how the model thinks in mind in order to make sure the model does the right thing more often than not. [SPEAKER_02] Dan Yeah, it's hard.

3:05

SPEAKER_01

[SPEAKER_02] Dan Yeah. So I use this SDK analogy sometimes. It took a long time for humanity to get to the point where we could make a really good Python SDK for a Python developer wrapping an API. And I think we've cracked that nut. Stainless offers really great Python libraries. But we're building on the shoulders of giants here. A lot of people have done this over time. We haven't figured out how to expose an API ergonomically to an LLM in the same way that we've figured out how to expose it ergonomically to a Python developer. And that's a new research problem in a sense. And it's harder because I can go learn how to be a Python developer if I want. I can't really learn how to think or see like an LLM. But it would be powerful if I could. And that makes it tricky. We do have at Stainless, I think, some things that we're cooking up to address some of these problems, including the ones that you also mentioned, like elements have a really hard time with a repeated sustained chain of actions. And even if you get an API response back saying, hey, list all the transactions, there's so much data, and you might have to go through the next page and the next page and the next page to go through all the transactions to find the one that has Dan with the stripy socks. And that's a ton of context with one or two small needles in the haystack. And LLMs are pretty good at that, but they're not perfect. And with too much hay, we all kind of end up throwing up our hands. And that's true for LLMs too. So there's a lot of challenges today.

3:10

SPEAKER_01

And when you look at, I mean, you're building MCP servers for people, but when you build them, and just generally when you see people doing it well today, what are the principles or how do you think about making an MCP server that one people use, which is actually a big one, and then two, when it is used actually does the right job. There have been relatively few times that I've seen it done well. I have seen it done well. We're kicking something up that I'm really excited about, but with today's technology, you really have to do a good job of product management. I mean, you have to go out into the market and talk to your customers and see what their actual needs are and look over their shoulders as they use and operate your software and think about

3:14

SPEAKER_01

and just generally when you see people doing it well today, what are the principles or how do you think about making an MCP, MCP server that one people use, which is actually a big one, and then two, when it is used actually does the right job. There has been relatively few times that I've seen it done well. I have seen it done well. We're kicking something up that I'm really excited about, but with today's technology, you really have to do a good job of product management. I mean, you have to go out into the market and talk to your customers and see what their actual needs are and look over their shoulders as they use and operate your software and think about what could we unlock through AI where people would be doing things that they can't really do with our software today because it just got so much easier. And then you have to do a lot of engineering work usually to wrap it up in a bow that works for the models. And you have to set up a really good system for evals. And if you're doing MCP, you have to think about the different clients that people might be using.

3:21

SPEAKER_01

[SPEAKER_02] Are they using cursor? Are they using cloud code? Are they using something else? And the different models underlying all that. So you end up with this pretty crazy matrix of things that you might want to optimize for and ways that you might want to evaluate and make sure that what you're offering is working well. And it's also a black box to get that feedback back to your servers so that you can find out, hey, we gave a tool call response here. We gave an answer of some kind. Was it actually any good? Did the user like it? Was the LM able to use it? And that's a problem that I think I haven't seen a lot of people solve yet as well. And so thinking about that as a first class thing, maybe you have a send feedback tool. That's something that we've been thinking about doing. Just so if a user says out loud, in the chat, oh man, that was useless garbage. Like, okay. Now that these MCP server is going to find out about that.

3:27

SPEAKER_01

Hmm. But is there anything specific you've learned about how to do it well, other than obviously you got to talk to your customers, think about your use cases, but more concrete, more applicable stuff about how to design a good MCP server? You want to keep the number of tools relatively small, relatively low. You want to have the tool name and the description be really precise and specific. Aren't those two things at odds?

3:47

SPEAKER_01

Yes. Good writing is hard. Yeah, I mean, you can make a great tool of lookup person by name and product description and then refund them. You can make a great tool that does that. And you also want a small number of properties in the input schema. You want a small number of parameters and you want them concisely described, but sufficiently described. This is also hard. And you want the response data to come back with a very small amount of data, only exactly what the model will need. That's also very hard because you may not know a priori which things the model's really looking for. And we have a technique that we use in our MCP servers today where we give the model a JQ filter, which is a way of filtering out JSON. And that can work pretty well. But that's a special trick.

3:51

SPEAKER_02

Doesn't this mean that MCP just needs another level of a search tool function, search tools, search, find a list of relevant tools given my task?

3:53

SPEAKER_02

The tool browsing problem is definitely one very serious one. And that is one approach. And so we actually do this at stainless today where you can get an MCP server for your API that just has, like I was saying earlier, the very simple thing of every endpoint is exposed as a tool. And if you have a small API that works great. And you can also filter it out. So you expose an MCP server with only a small subset of your endpoints. That works great. You can also use what we call dynamic mode where there's three tools, no matter how big your API is. One is list endpoints. The other is get endpoint and learn about it. And then the last one is execute endpoint. And so that enables this context thing to scale really well, but it means there's three turns of the model just to do one thing. And so that gets slower. It's more expensive in another sense. And there's some lossiness. It doesn't perform quite as well because the tools aren't loaded up in quite the same way.

3:59

SPEAKER_02

Are you using MCP servers yourself? Yeah, I use MCP to actually, funnily enough, not so much on the coding side, but I use it on the business side. So I'll use the notion, HubSpot, Gong, MCP servers to say, Hey, and actually an MCP server for our database, a read only copy of our database and say, Hey, what are the interesting customers that signed up for stainless last week? And it'll go off and make a great query of our Postgres database. And then it can cross reference those things in HubSpot and then look up our notes in notion. Maybe even look at transcripts in Gong, and tell me all about it. It's incredible.

4:10

SPEAKER_02

Lots of us are shipping AI to production, which is great for productivity, but it also comes with anxiety. You tweak a prompt, swap models, adjust parameters, and everything looks fine in testing. So you merge. And then three days later, or even sooner, the support tickets start rolling in. The AI is giving your customers unexpected answers and you have no idea when it happened or why.

4:12

SPEAKER_02

BrainTrust is the AI observability platform that fixes this. It connects evals and observability in one workflow. That way you see what actually happened in production and can measure whether the changes made things better or worse. Traces show the full execution path. Evals define what good looks like and experiments let you compare prompts and models side by side before shipping. Production traces feed directly into your eval datasets. Every failure becomes a test case. You catch regressions in CI before they reach users and teams at Notion, Stripe, Zapier, Vercel, and RAMP use it to ship quality AI at scale. BrainTrust is designed for teams building production AI systems where silent regressions are expensive. It's built for any stack. They have SDKs for Python, TypeScript, Go, Ruby,

4:15

SPEAKER_02

one workflow. That way you see what actually happened in production and can measure whether changes made things better or worse. Traces show the full execution path. Evals define what good looks like and experiments let you compare prompts and models side by side before shipping. Production traces feed directly into your eval datasets. Every failure becomes a test case. You catch regressions in CI before they reach users and teams at Notion, Stripe, Zapier, Vercel, and RAMP use it to ship quality AI at scale. BrainTrust is designed for teams building production AI systems where silent regressions are expensive. It's built for any stack. They have SDKs for Python, TypeScript, Go, Ruby, C Sharp. There's no framework lock-in or vendor dependencies. It's SOC 2, Type 2 certified, and GDPR and HIPAA compliant. Get started at Braintrust.dev. That's Braintrust.dev. And now, back to the episode.

4:18

SPEAKER_02

And so that's one of your big use cases. Are you doing that every week? Or how are you... Now I'm interested, not even from an MCP perspective, but for anyone running a business that has some complexity and you're wondering what's going on in the business. What are you actually doing? And what is the report that comes out? And how often are you doing that? So I can steal it.

4:29

SPEAKER_02

For me, it's still usually in playing around mode. One of the things is the MCP servers disconnect and then I get annoyed. You have to reconnect and whatever. It's not a huge deal. But there are a lot of small issues still in a technology this new that you're going to expect that can hold back some of your usage. One of the things that I found really helpful at the meta level, and I'm sure you've had other guests talk about this, is the practice of collecting notes for the AI by the AI and edited and curated by yourself. I have a notes folder, a research folder, something like that in a special Git repo that I use for internal stuff. And I'm saying, when you find interesting customer quotes, put them in this folder and give the full citation. That way the next time I start asking interesting questions, it doesn't have to go searching through the MCP servers again. It has them cached on disk in Markdown files.

4:38

SPEAKER_02

Wait, that's crazy. Wait, so how are you getting... What are you using to write into that Git repo? Is it Claude Code? Are you using Touch EBT? How does it get in there? Yeah, I use Claude Code these days for that. You just have Claude Code open and running and a new customer testimonial comes in and you're like, can you throw this in my Git master company Git knowledge repository? And then whenever you need anything later, you're like, Claude, go search through my master repository to figure out where the best customer quote is for this. Totally. That's so cool. Can we see it?

4:54

SPEAKER_02

No, it's too messy and probably has a lot of confidential information. The latter being more important. When you say it's messy, are you having Claude organize it at all? Or how is it structured?

5:06

SPEAKER_02

There's a lot that I want to do here that we haven't had the chance to do yet. There's some other lower hanging fruit that I'm working through, that our business team is working through right now. Just on the basics of your CRM systems and so on. But it's not well structured now, but I think that's fine. I don't plan to prioritize structuring it super well until we're using it more broadly because I use this stuff some of the time. One of the business people on the team uses it a fair amount. A couple of our customer support engineers use this stuff a lot. But it's not yet broader than that. And I would like it to get there. And once we see how everything's evolving, I think that's when we'll start bringing in more structure. But as it is, Claude Code can handle unstructured stuff really well. You don't have to think about it too hard in advance. And in my view, you can move things around later.

5:07

SPEAKER_01

[SPEAKER_02] What else do you have in there other than customer quotes?

5:12

SPEAKER_01

[SPEAKER_02] SQL queries. I'm a software developer. I don't write a lot of code these days, but I spent a lot of time doing that. When I say, hey, can you look up how is our month on month growth of XYZ metric over the last three months? I did this recently for board prep. It came out with a pretty good answer right away. And I was like, wow, this is awesome. Then I looked a little bit deeper and I was like, oh, I actually want to exclude these users from this analysis and filter it this way and that way. I imbued more business context into that SQL query. I iterated with Claude Code to get it better and better for the specific metric I was looking for, the specific story I was trying to tell. Then I got it to a good place and dumped it to an analytics folder for future use.

5:19

SPEAKER_01

[SPEAKER_02] And then next time you're doing your board prep, you can be like, hey, what was that query we did last time? And it'll presumably go get it. [SPEAKER_02] Yeah. That's really cool. What else?

5:23

SPEAKER_01

[SPEAKER_02] As any software team does these days, we're using this also for when a customer comes in with a question. Can Claude Code just fix it? You'll have a Linear ticket filed and our support engineers are really technical. They may not have the wall clock time to chase down the fix themselves for an incoming bug. They have the technical skill, but another customer writes in two minutes later and they want to jump on that. They don't want to be knee deep in a debugger. So something that we do sometimes is [SPEAKER_02] Hey, a customer comes in with a question. Can cloud code just fix it?

5:37

SPEAKER_01

[SPEAKER_02] So you'll have in some cases a linear ticket is filed and our support engineers are really very technical. And so they may not have the wall clock time to go down and chase down the fix themselves to an incoming bug. They have the technical skill. But guess what? Another customer writes in two minutes later and they want to jump on that. They don't want to be knee deep in a debugger. And so something that we do sometimes is they'll file the ticket in case, and by default it'll maybe they intend to do it later or some other engineer is going to be doing it later. But hey, can we see if cloud code can just take a crack at it? Is that going to work out 100% of the time? Definitely not. Is that going to work out 50% of the time? Still no. To be honest with you. But can that improve the overall efficiency? Yeah, maybe. We're still, I would say experimental there. But we're seeing a lot of promise.

5:42

SPEAKER_01

[SPEAKER_02] That's really interesting. Okay. Well, I know you also in our pre-production call, you were talking about you have a big vision for the future of AI. Do you want to talk me through that?

5:46

SPEAKER_01

[SPEAKER_02] Yeah. I would love to. We talked earlier about how agentic AI can make operators' lives a lot easier by taking their data, certain pedestrian tasks and running with it independently. And that's something that I think as an industry we're almost on the cusp of. And if you start stepping, you ask how you get there and you also start asking about the steps beyond that and beyond that. A big part of the way I see things unfolding from here, I like to say, is the future of AI is cyborgs. Which is extra ridiculous because what is a cyborg other than already a robot? But you know, cyborg, as I understand it is a term that means you're part person and then part machine. And in this case, when you go and talk to an agent, what you're going to be getting is part GPT neural net LLM part AI and part code, where the machine I'm talking about is traditional CPU, not GPU software. And to me, I think I expect this to play out in two main ways. One is your kind of one-off operational use cases like we were talking about a minute ago, and then the other is production software. And in the use case we were talking about a minute ago, where someone needs to perform some tricky one-off action with a bunch of point and clicks, and now we want an AI to just do a bunch of tool calls. The way I actually see that happening and what we're building towards is code execution. So rather than the model having a bajillion tools, the model has two tools. One to execute code where it just has a text box like,

5:51

SPEAKER_02

[SPEAKER_01] hey, put in some TypeScript and you're going to use this API's TypeScript SDK, and you're just going to write Stripe.transactions.list, or Stripe.charges.list, and you're going to stripe.customers.retrieve and stripe.refunds.create. This is really easy for models. They're really good at writing code. And if you give that tool a little bit of a readme, where you say, here's an example request, and here's some other resources, some other API calls that you can make. It's really good at extrapolating from patterns if the SDK is an API or well-formed and predictable, and then you give it an additional tool to search the docs, and ask questions to the docs. And anything it's not sure about or gets wrong on the first try, you give it the documentation. And what this does for that scenario that we were talking about earlier is you have very, very limited impact on the context window up front. I mean, we're talking about a thousand tokens or something like that, maybe less. And the context impact of doing a whole bunch of paginated list requests, zero. You know, the model will go look for somebody named Dan and it'll double check that the purchase of Stripey socks and you might write three nested for loops, but then only at the end when it found the right thing, it'll console.log found Dan customer ID, blah, blah, blah, transaction ID, blah, blah, blah. And then create refund, you know, refund ID, one, two, three. And the context hit coming back from all of this is going to be like 10 lines of text, you know, it's really minimal. And all of this will run really, really quickly too. So you don't have a round trip to the model every time you're doing something like this, it's just CPU code and it runs in a server in the cloud right next to the Stripe API in AWS somewhere probably. And it goes super, super fast.

5:56

SPEAKER_02

Okay. So what I am understanding you saying is the language model has a tool where it can write code and send that code to this tool that whoever the company is, whether it's Stripe or whatever, whoever's MCP server you're using, they'll go and execute that code. And that code is going to interact with their API and then return the results rather than these sort of you have 50 different possible tool calls and all that stuff. It's just model writes API code and API provider executes that code, runs it on their API and returns the results.

6:03

SPEAKER_02

Why wouldn't I just write the code that I then run myself instead of relying on an API provider to do it? I expect that that will happen a lot more. I expect that the code execution tool is going to become the most widely used tool. The problem, one of the problems that we have today is that the code execution tool doesn't work so well with libraries. LLMs have a hard time working with library and knowing exactly what version of the library it's using, using the right version, probably usually the latest version. And LLMs have a hard time working with library. Not hallucinating aspects of the API and knowing how to iterate if it hallucinates wrong. And if it can't use any library off NPM or, you know,

6:08

SPEAKER_02

tool is going to become the most widely used tool. The problem, one of the problems that we have today is that the code execution tool doesn't work so well with libraries. LLMs have a hard time working with libraries and knowing exactly what version of the library it's using, using the right version, probably usually the most latest version, and LLMs have a hard time working with libraries, not hallucinating aspects of the API and knowing how to iterate if it hallucinates wrong. And if it can't use any library off NPM or the Python package index or anything like that really well, perfectly out of the box, then okay, well forget about using a library at that point, you just have to hit the raw HTTP API. And at that point, in order to figure out what's in there, you need the whole OpenAPI spec and you're back at square one because that document is massive. And furthermore, something that's really scary about that is if you don't have a typed library with static typing, where the computer can say what you're trying to do is wrong, then the LLM will try to make an API request that is wrong some percentage of the time. The code execution tool can run a type checker and say, oh, you're asking about stripe.transactions.list, but that actually doesn't exist. Stripe doesn't have a transactions API. You might want payment intents, you might want orders, you might want balanced transactions, which one do you want? And if the API provider is doing a great job building this tool, it'll return the documentation for all of these things in line.

6:17

SPEAKER_02

[SPEAKER_01] It might have its own AI, look at what the model's trying to do and come up with a suggestion. And that sub agent is well-trained, specified, always updating, and isn't burdened with the context of the full conversation. What do you think of the security model?

6:21

SPEAKER_02

The security model is really interesting. This is another area where we're really starting to think about things at Stainless and I'm getting really excited about it. So if any listeners are really interested in this and have some ideas or want to talk, please do reach out. At the end of the day, I think the security has to take place at the API layer itself. Right now you see people trying to implement security by limiting what's exposed through MCP. And that kind of makes sense, but at the end of the day, you could do anything that's in the API under the hood. Right. And what people should be doing is using OAuth with granular permissions with proper scopes. And at that point, the security happens at the right place, which is at the API layer. There's limitations to OAuth scopes. And it's pretty hard to build. So it'd be nice if someone made that easy, but my view, that direction is the right layer. So going back to my earlier question, I'm thinking about the idea of having a model write code that then the API provider executes to interact with their API and then returns the results. Would you ever consider just creating a tool use tool that developers use? Because for example, I'm thinking about for Quora, got all these tools, maybe Gmail is going to build a code use thing or whatever, but really I just want, I would probably use what you're talking about inside of Quora, but we would need a tool use tool, but it's not a tool use tool. It's a computer use tool where, and I know OpenAI has this, but it's not really well built for lots of libraries and stuff. It's not a custom environment. Like I need a computer use tool where I control the environment and I can install different libraries in it and be able to call it any time to then call any API, or it has to have network access basically. Yeah. You guys should build that. We're working on it. Yeah. You're building it for developers who want to access MCP servers or people who are providing MCP servers. We're starting with people who are providing MCP servers, but ultimately I think that we're going to need this to work such that you can give the model a code execution environment where it can hit not only the Stripe integration, but also the Salesforce integration and also anything else. And but not too much anything else. Right. And so one of the advantages of starting where we're starting, of just one API provider is that you ensure that there's no network connections allowed out of that sandbox where we're running the code to anything other than, in this case, api.stripe.com. And that's really critical for security for something like this. And so there's ways to expand that bit by bit and keep things secure. It'll take some time. The other thing I think to point out as you see some of these generalizations is it's not just that you want this code execution sandbox to work really well for any API for any library. Which I think we really do. I think we really need that. You also start to see that this is just a powerful model for AI doing stuff. And sometimes you realize that the thing that the AI did this one time in this one-off case is actually enduringly useful. Maybe anytime a customer writes into support and says, "Hey, my socks had holes in them," you should automatically get a refund. You know, maybe you want that, maybe you don't, but there's a lot of stuff that people do one time and then two times and then three times. And then they say, "Okay, we should automate this." Right. And that's what software teams do all day, every day. Right. And we're going to be seeing that with AI where the same code execution tool that we're talking about, all the same prompting that will make an AI really good at interacting with an API in one of these code sandboxes, kind of almost in its brain, working like writing code in its head, running the code in its head, seeing the results and then moving forward with your task. It should be able to say, "Okay, actually this is enduringly useful code. Let me commit this to the repo." Yes. Yes. Yes. Yes. It's like, you know, chat is a really good interface for exploring, but sometimes you just want a dashboard, you know, you just want to log into my Stripe dashboard and see all the stuff without having to be like, "What is my MRR?" It should just show up.

6:27

SPEAKER_02

Almost quote unquote in its brain, working like write code in its head, run the code in its head, see the results and then move forward with your task. It should be able to say, okay, actually this is enduringly useful code. Let me commit this to the repo. Yeah. Yeah. Yeah. Yeah. It's chat is a really good interface for exploring, but sometimes you just want a dashboard. I just want to log into my Stripe dashboard and see all the stuff without having to be asking what is my MRR? It should just show up, because I just do that every day. But I want to push you as a hashtag value add investor. Because I think there's this thing that happens in AI where often the first attempt at something like this, people try to be really cautious and I'm sure that your customers care about you being cautious, like big enterprise customers, but the things that get adopted are often the ones that are willing to take the risk to be YOLO very early. So an example is Dolly was totally private for a long time and people were posting some images, but you couldn't get in. And then stable diffusion was just like, fuck it. Anyone can use this. And then that really started the whole image generation wave, obviously stable diffusion fumbled the bag, but they had a lead for a little while. Same thing for cloud code, honestly. If you look at codex, it's not like this as much anymore, but if you look at the difference between codex CLI and cloud code, cloud code was just like, fuck it. YOLO mode. It's super industrious. It has a sandbox, but you can just do dangerously skip permissions and codex just fell way behind because it was first, it was in the browser. And so their whole thing was locked down. And then it was in the CLI, but it was really built for pair programming. And so it just wasn't particularly industrious. It wouldn't go off and do a bunch of stuff. It would get locked out of doing certain things, even if you did full auto mode. And now they've caught up because they're like, yeah, you can just let it do whatever you want. And so I would really push you on, there might be a version that you could do today or tomorrow or very soon for individual developers that would let them set up this environment that for example, I would use immediately. And I care about security, but I care a lot less than some gigantic enterprise company. But I think the people like me who are building at this scale are eventually hopefully going to be the big companies, but we're the ones that are really doing the AI first adoption, not the big companies. Well, I would love to get this in your hands. What are some of the APIs your team uses the most? I'm thinking as we have a bunch of different products, but I'm thinking right now about Quora, the email assistant. And it has all of the big APIs that it's using, it's mostly the Gmail API. And so you're interacting with the assistant over chat and then it has a list of tools that are archive email or draft email or send email or whatever. There's a whole categorized tool. So it categorizes your mail in certain ways. And I think we would definitely try out something like this because it would, if it ran the same way, it would make it much more flexible for us to make more tools and not break old ones. It's really interesting. I mean, in a sense, what I actually predict is that people who are building tools, once we have a code execution super tool, like I'm talking about, that the only way you really build a tool is with instructions, with prompts. And the full power of everything you could possibly do in the API, in the Gmail API, for example, it's all there in one tool. But sometimes you have specific tasks or specific categories of work that you want to describe in a particular way to help the LLM perform a sequence of actions as productively as possible. And at that point, the only work in engineering that you have to do is prompt engineering. We'll see if it's that easy. As we all know, prompt engineering can be really tricky. It's hard. Yeah. But I think that's part of the vision. That being said, we do have some pretty nifty ways with the MCP servers that we generate today to help developers mix and match all the parts of the different tools, all the different parts of the API, as they compose and write their own tools. This is awesome. So for people who are listening and want to know more from you or know more from stainless, where should they find you? Stainless.com, that's our website. Awesome. Visit stainless.com. Alex, great to have you on. I can't wait to do more of this when you have some of these new things launched. This is really fun. And great to chat. Thanks, Dan. You too. Oh my gosh, folks, you absolutely positively have to smash that like button and subscribe to AI and I. Why? Because this show is the epitome of awesomeness. It's like finding a treasure chest in your backyard, but instead of gold, it's filled with pure unadulterated knowledge bombs about ChatGPT. Every episode is a roller coaster of emotions, insights, and laughter that will leave you on the edge of your seat, craving for more. It's not just a show. It's a journey into the future with Dan Shipper as the captain of the spaceship. So do yourself a favor, hit like, smash subscribe, and strap in for the ride of your life. And now, without any further ado, let me just say, Dan, I'm absolutely hopelessly in love with you.

6:35

SPEAKER_02

of to make it easier for computers to talk to computers. So, and, you know, it's the long running trend of technology to have more automation, right? Automation is what we mean when we say, okay, we're going to, you know, we're going to, we're going to apply technology to that. You know, we're generally going to be making things more efficient. And APIs are how most business to business interactions in some format or another become, become real, become automated. And what we see with the rise of AI is that there is a new, a new computer has entered the chat, right? There's a new, there's a new kind of system that can talk to other systems,

7:22

SPEAKER_02

or at least we would like it to be able to. You used to have either, you know, humans interacting with a computer through a user interface, a UI, or a computer acting with a computer through, through an API. And now we have LLMs interacting with computers, right? And what's that

7:38

SPEAKER_01

through? And I'm sure anyone familiar with, you know, with every and his regular listeners is going

7:44

SPEAKER_02

to be familiar with MCP, model context protocol, which is a system for connecting LLMs to computers, broadly speaking. And it's an area that we're investing in at stainless. It's really, I think, part of our core mission of, like I said, make it easy for computers to talk to computers. And we've invested a lot of time, you know, at stainless, the core product that we first brought to market is software development kits, SDKs. And so these are ways of saying, okay, Stripe has this great REST API, you know, you can send JSON over HTTP and get back JSON over HTTP. And if you want that

8:30

SPEAKER_02

to be really convenient, you're going to use the Stripe Python library, the Stripe Python SDK. So you

8:36

SPEAKER_01

can go, if you're a Python developer, you'll go pip install Stripe. And then in your application code,

8:41

SPEAKER_02

you'll write Stripe.customers.create. And all of a sudden, you have a nice new customer object in sort of your Stripe database. And you're off to the races, or Stripe.charges.create in the old days to charge a credit card. And SDKs are what gives developers that easy way to interface with an API. What's the thing that gives LLMs an easy way to interface with an API? And you might say MCP, and in a sense, you'd be right. But what we're seeing so far, as MCP is rolling out into the world, and people are experimenting with it and trying it out, is that it's not working so great. It's difficult to deliver on what I see as the core vision

9:31

SPEAKER_02

of what's so exciting about MCP, which is just like a dashboard and a user interface lets you click around, see a bunch of stuff, fill out forms, click buttons, do things.

9:46

SPEAKER_01

Anything that you would do while you're interacting with the software, or you do through the user interface generally. But LLMs interacting with through MCP, it tends to be much more restricted. You can only do a few little things. There's usually not a ton of tools that you're going to be exposing to the models. And just to stop you there, so I think what I'm hearing you say is what MCP does is, just like a website is built for humans to be used, MCP is sort of the equivalent,

10:16

SPEAKER_02

and you can think of it in certain ways, of exposing a set of tools for the model that it can use to

10:23

SPEAKER_01

perform certain functions. Just like you might click a button on a website, the MCP gives to the model a bunch of things it can click on or use to get work done. So an example might be, you know, a Gmail MCP has like a send mail tool or like a compose mail tool or a read inbox tool, that kind of thing. And instead of a human going on the Gmail website and doing it, it's the LLM that's like, you know, essentially logging in and using it itself. And it's a native interface for language models. But you're saying that that's not working that well. Can you tell me more about that?

10:57

SPEAKER_01

Yeah. So let's start actually with kind of what I see is the big vision of MCP. And in some sense,

11:05

SPEAKER_02

the big vision of agentic AI in the first place. And I'll start with the most pedestrian example you can imagine. It's going to be funny given some of our context. Which is, let's say, you know, Dan walks into my store and buys a pair of stripy socks, and maybe a few other things. And then the next day, I hear back from Dan, that there's something wrong. Unfortunately, it happens, you know, and I turn to someone on my team. And I say, Hey, can we refund Dan for those stripy socks he bought yesterday, and send him a discount code for the next time he comes in with like a little thank you note. Because we like to take care of our customers. This is like the most

11:47

SPEAKER_02

normal thing to do in software is some little task like this. And what you're going to do with the, you know, the member of my team would be doing would be opening up their internal admin and looking around for some things, they might go to the stripe dashboard and try to look through the list of payments or the list of transactions or orders, and try to find one that has someone named Dan, which Dan, I don't know, there might be a bunch of Dan's, try to look through the list of products in the order and see whether there was some stripy socks in there that might be a few clicks required,

12:16

SPEAKER_02

depending, find the right one, then go to the screen where you can create a refund, create a refund, make sure it's the right amount, then go and create that discount and then take that discount code and send it over to some other SaaS app where you log in to send some to mail automatically, right. And of course, if you step away from the consumer version of this to a business to business context, of course, you might be going into Salesforce and sending a Slack message to an account administrator, you know, an account manager, so on and so forth. And in the normal course of work,

12:52

SPEAKER_02

it's just the most normal thing in the world to be doing having one task involve going through five different apps each time 15 different clicks and scrolls and loading loading spinners, just to do sort of like one simple thing. And the promise of agentic AI is to be able to take that same prompt I just said and type it into chat GPT or cloud or whatever, and say, Hey, chatty, buddy, can you help refund my, my friend Dan, da da da da da da. And just have the AI go off and do that. And basically go through these five different apps and the 15 different screens, and the various different, you know,

13:34

SPEAKER_02

button presses to complete the task and then come back and say, Great, it's done. That in order to do that, now that's, there's only so many tool calls you have to make as a as an AI model to perform that exact linear chain of events, it's somewhat tractable. But if you think about this in the general case, you want the LLM to be able to do the you want your agentic AI to be able to do anything that that human operator would have done, and you would want them to be able to do it without having to wait for a bunch of JavaScript to load on a website or anything like that. And that means you need not only the stripe

14:17

SPEAKER_02

create refund tool and the stripe list transactions tool, and the stripe, you know, list products and look up customer and, you know, create discount tool, you need not only those tools, but you need everything that you can do in the stripe dashboard, which is basically everything that you can do in the stripe API. And that's actually a lot like there are hundreds of different endpoints that you have access to in the stripe API. The stripe dashboard is actually massive. It's a huge application. And if you were to take that list of tools today and go to an LLM and say, Hey, here's our MCP definition for all of this,

15:02

SPEAKER_02

here's a create refund tool, here's a create transactions tool, so on and so forth. And you tell it all about those tools, here's the description, here's all the different request properties that you can send, here's the response properties you can get back, here's all the documentation for each of those things. Everyone listening to this should already know. You've just burned through your entire context budget. That's, you know, maybe hundreds of thousands of tokens just there, just in pretty much translating the stripe open API spec directly over to MCP tools. And today's models not only can't handle that amount of

15:36

SPEAKER_02

context, it's a poor use of context, because you have a lot else going on. But it's also confusing to the model. It's just, it's just too much to hold in your brain at one time. And that's just the stripe part of it, right? Because what you're really trying to do is enable your operators to do anything they would normally do. And again, that spans many, many different SaaS tools, right? In the course of one interaction, it might be five, in the next interaction, it might be a different five. And so if you think about every single SaaS tool that your business uses on a daily basis, to get your work done,

16:14

SPEAKER_02

ideally, you would want every single one of those tools to be exposed to your operators in their AI chat, with every single tool available in there, with every single nook and cranny and corner case available, so that you can do anything through AI. That's the vision. Now, there's a lot of problems with that. The biggest one that I mentioned is sort of this context window limit. But you also have all sorts of security and permissions problems, because you don't want the AI to color outside the lines and say, Okay, in addition to refunding Dan Socks, I also refunded every customer for all transactions ever,

16:50

SPEAKER_02

you know, and then I sent, you know, a bunch of money to my own AI bank account, ha ha ha. And so there's more to the challenge, but that's the vision I see. Dan Socks, I think, you know, the place we started there was, you said it's not working. But I don't think that that's the reason why it's not working today, right? Or is that the reason why it's not working today? Dan Socks, Socks, what people do with MCP today is sometimes they'll try to expose all parts of their API. And the way people build MCP tools is generally speaking, they have an underlying API, usually a rest API, and they wrap different parts of that different endpoints, different operations

17:33

SPEAKER_02

in MCP tools. And you can kind of do that in a one to one mapping, or you can kind of handcraft things for the MCP. And today, in order to succeed, people are finding that you really have to kind of handcraft it to the MCP, to the LMS, you have to say, Okay, I'm making one specialized tool to look up a customer and refund their transaction based on a description. Dan Socks, Socks, there's all these like decisions that you have to make, where you need to have like the ergonomics of the model and how the model thinks in mind in order to make sure the model does the right thing more often than not. Dan Yeah, it's hard.

18:11

SPEAKER_02

Dan Yeah, yeah. So I use this SDK analogy sometimes. So it took a long time for humanity to get to the point where we could make a really good Python SDK for a Python developer wrapping in API. And I think we've we've we've cracked that nut. Stainless offers really great Python libraries. But you know, we're building on the shoulders of giants here. A lot of people have have done this over time. We haven't figured out how to expose an API ergonomically to an LM in the same way that we've figured out how to expose it ergonomically to a Python developer. And that's kind of like a new

18:46

SPEAKER_02

research problem in a sense. And it's harder because I can go learn how to be a Python developer if I want. I can't really learn how to go think or see like an LLM. But, you know, sure would be powerful if I could. And, and that makes it that makes it tricky. We do have it seamless, I think, some some things that we're cooking up to address some of these problems, including not, you know, including the ones that you also mentioned, like, elements have a really hard time with a repeated sustained chain of actions. And, you know, even like if you get an API response back around, hey, like list all the transactions, there's so much data, and you might have to go

19:32

SPEAKER_02

through the next page and the next page and the next page to go through all the transactions to find the one that has Dan with the stripy socks. And that's, again, a ton of context with one or two small

19:43

SPEAKER_01

needles in the haystack. And LLMs are pretty good at that, but they're not perfect. And with too, with too much hay, you know, we all kind of end up throwing up our hands. And that's true for LLMs too. So yeah, so there's a lot of challenges today. And, and so when you look at, I mean, you're building MCP servers for people, but when you build them, and just generally when you see people doing it well today, like what are the principles or how do you think about making an MCP, MCP server that one people use, which is actually a big one, and then two, when it is used actually does the right job. There, there has been relatively few times that I've seen it

20:26

SPEAKER_01

done well. I have seen it done well. We're kicking something up that I'm really excited about, but with today's technology, you really have to do a good job of product management. I mean, you have to go out into the market and talk to your customers and see what their actual needs are and look over their shoulders as they, you know, use and operate, you know, your software and think about what could we unlock through AI where people would be doing things that they can't really do with our software today because it just got so much easier. And then you have to do kind of a lot of engineering

21:00

SPEAKER_01

work usually to wrap it up in a bow that works for, for the models. And you have to, you know, you have to set up a really good system for evals. And if you're doing MCP, you have to think about

21:11

SPEAKER_02

the different clients that people might be using. Are they using cursor? Are they using cloud code? Are they using something else? And the different models underlying all that. So you end up with this pretty crazy matrix of things that you might want to optimize for and ways that you might want to evaluate and make sure that what you're offering is working well. And it's also kind of a black box to get that feedback back to your servers so that you can find out, hey, we gave a tool call response here. We gave an answer of some kind. Was it actually any good? Did the user like it?

21:48

SPEAKER_02

Was the LM able to use it? And that's a problem that I think I haven't seen a lot of people solve yet as well. And so thinking about that as a first class thing, maybe you have like a send feedback tool. That's something that we've been thinking about doing. Just so if a user like says out loud, you know, in the chat, oh man, that was useless garbage. Like, okay. Now, now that these MCP server is going to find out about that. Hmm. But is there anything specific you've learned about like how to do it well, other than like, obviously you got to talk to your customers, think about your use cases, but like more concrete,

22:24

SPEAKER_02

more, more applicable stuff about how to design a good MCP server? You want to keep the number of tools relatively small, relatively low. You want to have the tool name and the description be, be really precise and specific. Um, uh, Aren't those two things at odds? Yes. Good writing is hard. Um, yeah, I mean, that that's, that's what, like, you know, you can make a great tool of lookup person by name and product description and then refund them. You can make a great tool that does that. Um, and you also want a small number of, of, in, you know, properties in the input schema. Um, you want a small number of parameters and you want

23:05

SPEAKER_02

them concisely described, but sufficiently described. Um, this is, this is also hard. Um, and you want the response data to come back with a very small amount of data, um, only, only exactly what the model will need. That's also very hard because you may not know a priori, which things the model's really looking for. Um, and you know, we have a technique that we use in our MCP servers today where we give the model a JQ filter, which is a way of filtering out JSON. Um, and that can work pretty well. Um, but, but that's kind of a, a special trick. Doesn't this mean that like MCP just needs another level of like a search tool function,

23:45

SPEAKER_02

search tools, search, like find a list of relevant tools given my task. The, the tool browsing problem is, is it's definitely one very serious one. Um, and that is one approach. And so we actually do this at stainless today where you can get an MCP server for your API that just has, like I was saying earlier, the very simple thing of every endpoint is exposed as a tool. And if you have a small API that works great. Um, uh, and you can also filter it out. So you expose an MCP server server with only a small subset of, of your, of your endpoints. That works great. Um, you can also use kind of what we call dynamic mode where there's three tools,

24:23

SPEAKER_02

no matter how big your API is. One is, you know, list endpoints. The other is get endpoint and learn about it. Um, and then the last one is execute endpoint. Uh, and so that enables this context thing to scale really well, but it means there's three turns of the model just to do one thing. Um, and so that, that gets slower. It's, it's more expensive in another sense. Um, and, um, there's some lossiness that it doesn't perform. It performs pretty well, um, usually, but not, not quite as well because, um, the, the tools aren't loaded up in quite the same way. Are you using NCP servers yourself? Yeah, I use, I use MCP, um, uh, to

25:11

SPEAKER_02

actually, uh, funnily enough, not so much on the, um, coding side, but I use it on the business side. Um, so I'll use like the notion, uh, HubSpot, Gong, um, MCP servers to kind of say, Hey, like, and actually an MCP server for, for our database, um, a read only, a read only copy of our database and say, Hey, what are the interesting customers that signed up for stainless last week? Um, and it'll go off and make a great query of our Postgres database. And then it can cross reference those things in HubSpot and then look up our notes in notion. Um, maybe even look at transcripts and Gong, um, and tell me all about it. Um, it's, it's incredible.

25:49

SPEAKER_02

Lots of us are shipping AI to production, which is great for productivity, but it also comes with anxiety. You tweak a prompt, swap models, adjust parameters, and everything looks fine in testing. So you merge. And then three days later, or even sooner, the support tickets start rolling in. The AI is giving your customers unexpected answers and you have no idea when it happened or why. BrainTrust is the AI observability platform that fixes this. It connects evals and observability in one workflow. That way you see what actually happened in production and can measure whether it changes made things better or worse. Traces show the full execution path. Evals define what good

26:23

SPEAKER_02

looks like and experiments let you compare prompts and models side by side before shipping. Production traces feed directly into your eval datasets. Every failure becomes a test case. You catch regressions in CI before they reach users and teams at Notion, Stripe, Zapier, Vercel, and RAMP use it to ship quality AI at scale. BrainTrust is designed for teams building production AI systems where silent regressions are expensive. It's built for any stack. They have SDKs for Python, TypeScript, Go, Ruby, C Sharp. There's no framework lock-in or vendor dependencies. It's SOC 2, Type 2 certified, and GDPR

26:58

SPEAKER_02

and HIPAA compliant. Get started at Braintrust.dev. That's Braintrust.dev. And now, back to the episode. And so that's one of your big use cases. Are you doing that every week? Or how are you... Now I'm interested, not even from an MCP perspective, but for anyone running a business that has some complexity and you're like, I want to know what's going on in the business. Like, what is... What are you actually doing? And what is the report that comes out? And how often are you doing that? And all that kind of stuff. So I can... Tell me so I can steal it. Yeah. For me, it's still usually in kind of like playing around mode. One of the things is the MCP

27:33

SPEAKER_02

servers disconnect and then I get annoyed. And so, you know, you have to just kind of reconnect and whatever. It's not a huge deal. But there are a lot of little paper cuts still in a technology this new that you're going to expect that can hold back some amount of your usage. One of the things that I found really helpful kind of at the meta level, and I'm sure you've had other guests talk about this, is the practice of just collecting notes for the AI by the AI and kind of edited and curated by yourself. So, you know, I have a like a... I can't remember if I call it a note. I think I have a notes folder,

28:14

SPEAKER_02

a research folder, something like that in a special Git repo that I use just for this sort of like internal stuff. And I'm like, hey, when you find interesting customer quotes, put them in this folder and give the full citation. So that the next time I start asking interesting questions, it doesn't have to go searching through the MCP servers again. It has them kind of cached in just on disk in Markdown files. Wait, that's crazy. Wait. So how are you getting... Like, what are you... What are you using to write into that... into that Git repo? Like, is it Cloud Code? Is it... Are you using touch EBT? Like, how does it get in there?

28:52

SPEAKER_02

Yeah, I use... I use Cloud Code these days for that kind of thing. And so you just have Cloud Code open and running and then a new customer testimonial comes in and you're just like, hey, can you throw this in in my like, Git master company Git knowledge repository basically? And then whenever you need anything later, you're like, Claude, like, go search through my master repository to figure out where the best customer quote is for this. Totally. That's fucking so cool. Can we see it? Um, no, it's too messy and probably has a lot of confidential information. The latter being more

29:32

SPEAKER_02

important. Is it... When you say it's messy, like, are you having Cloud organize it at all? Or like, how is it structured? There's a lot that I want us to do here that we haven't had the chance to do yet. There's some other lower hanging fruit that I'm working through, that our business team is working through right now. Just on the basics of your kind of CRM systems and so on. But... And so it's not well structured now, but I think that's fine. Yeah, I don't plan to prioritize structuring it super, super well until we're using it more. I'm using it more broadly because, you know, I use this stuff

30:12

SPEAKER_02

some of the time. Um, one of the, one of the business people on the team uses it a fair amount.

30:18

SPEAKER_01

Um, I think like one or two kind of of our customer support engineers, um, use, uses this stuff a lot. Uh, but it's not yet kind of broader than that. And I would like, I would like it to get there. And once we see how everything's evolving, I think that's when we'll start bringing in more structure. But as it is, cloud code can, can handle unstructured stuff really well. Um, so you don't have to think

30:40

SPEAKER_02

about it too, too hard in advance. And in my view, um, you can move things around later. What else do you have in there other than customer quotes? Um, SQL queries. Um, so, you know, I'm a software developer. Um, uh, I, I don't write a lot of code these days, but you know, I spent a lot of time doing that. And so, um, when I say, Hey, you know, can you look up, uh, you know, I might be, Hey, how is our month on month growth of XYZ metric over the last three months? You know, I did this recently, I did this for my last, uh, board prep. Um, and, um, it came out with a pretty good answer right away. And I was like, wow, this is awesome. And then I kind of

31:20

SPEAKER_02

looked a little bit deeper and I was like, Oh, I actually want to exclude, you know, these users from this analysis and I want to filter it this way and filter it that way. And I kind of imbued more of this business context into that SQL query. And I iterated with, um, with cloud code, um, to get it

31:37

SPEAKER_01

to be better and better for the specific kind of metric that I was looking for, the specific kind of story that I was trying to tell. And then I got it to a good place. I was like, great, let's dump this to, you know, an analysis folder, um, for, um, or an analytics folder, um, for future use.

31:54

SPEAKER_02

Hmm. And then next time you're doing your board prep, you can be like, Hey, what was that query that we did last time? And it'll presumably go get it. Yeah. That's really cool. What else? You know, uh, as any software, um, team is these days we're, we're, we're using this also for, for, Hey, a customer comes in with a question. Can, can cloud code just fix it? Um, uh, Uh, you know, and so you'll have, uh, in some cases, a linear ticket is filed and then, you know, our support engineers are really very technical. Um, and so they may not have the, the wall clock time

32:32

SPEAKER_02

to go down and chase down the fix themselves to, you know, an incoming bug. Um, they have the technical skill. Um, but guess what? Another customer writes in two minutes later and, and they want to jump on that. They don't want to be, um, knee deep in a debugger. Um, and so, um, something that we do sometimes is they'll file the ticket, um, in case, and by default, it'll maybe they intend to do it later or some other engineer is going to be doing it later. Um, but Hey, can we, can we see if cloud code can just take a crack at it? Um, is that going to work out a hundred percent of the time? Definitely not. Is

33:09

SPEAKER_02

that going to work out 50% of the time? Still no. Um, to be honest with you. Um, but can that improve the overall efficiency? Um, yeah, maybe, uh, we're still, I would say experimental there. Um, but, but we're seeing a lot of promise. Hmm. That's really interesting. Okay. Well, I know you also, you know, in our, in our pre-production, uh, call, you were talking about, you have a big vision for the future of AI. Do you want to, do you want to talk, talk me through that? Yeah. Yeah. I, I would love to, you know, um, we, we talked earlier about, um, how agentic AI can, can make, um, operators lives a lot easier by taking their data, you know, certain pedestrian tasks

33:55

SPEAKER_02

and sort of running with it independently. Um, and that's something that I think as an industry, we're almost on the cusp of, um, and if you start stepping, you know, you ask how you get there and you also start asking about the steps beyond that and beyond that. Um, a big part of, of, uh, the way I see things unfolding from here, uh, I like to say, um, is the, the future of AI is cyborgs. Um, which is like sort of like extra ridiculous because like, what is a cyborg other than like already like a, a robot? Um, uh, but you know, cyborg, as I understand it is a term that means you're sort of like part, you know, person and then part machine. Um, and in this case, I mean,

34:41

SPEAKER_02

um, when you go and talk to an agent, what you're going to be getting is part GPT neural net LLM part AI and part code, um, where the, the machine quote unquote that I'm talking about is, is, um, traditional CPU, not GPU software. Um, and, uh, to me, I think I expect this to play out in two main ways. One is your kind of one-off operational use cases like we were talking about a minute ago, and then the other is production software. Um, and in, in the use case we were talking about a minute ago, um, where someone needs to kind of perform some tricky one-off action with a bunch of points and clicks,

35:32

SPEAKER_02

and now we want an AI to just do a bunch of tool calls. The way I actually see that happening and what we're building towards is code execution. So rather than the model having a bajillion tools, model has two tools. Um, one to execute code where it just kind of has a text box of like,

35:54

SPEAKER_01

hey, put in some TypeScript and you're going to use this API's TypeScript SDK, and you're just going to write Stripe.transactions.list, um, or Stripe.charges.list, um, and you're going to just stripe.customers.retrieve and stripe.refunds.create. This is really easy for models. They're really good at writing code. Um, and, um, if you give that tool a little bit of sort of a readme,

36:21

SPEAKER_02

where you say, here's an example request, and here's some other resources, some other API calls that you can make. It's really good at extrapolating from patterns with, if the SDK is sort of an API or well-formed and predictable, and then you give it an additional tool to kind of search the docs, and ask questions to the docs. And, um, anything it's not sure about or gets wrong on the first try, you give it the documentation. And, um, what this does for that scenario that we were talking about earlier is you have very, very limited, um, impact on the context window up front. I mean,

37:00

SPEAKER_02

we're talking about a thousand tokens or something like that, uh, maybe less. And, um, the context impact of doing a whole bunch of paginated list requests, zero. You know, the, the model will go look for somebody named Dan and it'll double check, um, that the purchase of Stripey socks and you might write three nested for loops, but then only at the end when it found the right thing, it'll console.log found Dan customer ID, blah, blah, blah, transaction ID, blah, blah, blah. Um, and then create refund, you know, refund ID, one, two, three. Um, and the context, uh, hit coming back from all of this is going to be

37:42

SPEAKER_02

like 10 lines of, of text, you know, it's, it's really minimal. Um, and all of this will run really, really quickly too. So you don't have a round trip to the model every time you're doing something like this, it's just CPU code and it runs in a server in the cloud right next to the Stripe API in AWS somewhere probably. Um, and it goes super, super fast. Okay. So what I am understanding you saying is like the language model has a tool where it can write code and send that code to this tool that the, you know, whoever the company is, whether it's Stripe or whatever, whoever's MCP server you're using,

38:18

SPEAKER_02

they'll go and execute that code. And that code is going to interact with their API and then return the results rather than like these sort of, you know, you have 50 different, you have 50 different possible tool calls and you know, all that stuff. It's just model writes API code and API provider, uh, executes that code, runs it on their API and returns the results. Why wouldn't I just, um, why wouldn't my model just like write the code that I then run myself instead of relying on an API provider to do it? Um, I expect that that will happen a lot more. I will expect, I expect that the code execution

38:55

SPEAKER_02

tool is going to become the most widely used tool. Um, the problem, one of the problems that we have today is that, um, the code execution tool doesn't work so well with libraries. LLMs have a hard time working with library and knowing exactly what version of the library it's using, using the right version, probably usually the most late, the latest version, um, and, um, LLMs have a hard time working with library. Um, uh, not hallucinating, you know, aspects of the API and knowing how to iterate if it hallucinates wrong. And if it can't use any library off NPM or, or, you know,

39:34

SPEAKER_02

uh, the Python package index or anything like that really, really well, basically perfectly out of the box, um, then, okay, well forget about using, um, a library at that point, you just have to hit the raw HTTP API. And at that point, in order to figure out what's in there, you need the whole open API spec and you're back at square one because that document is massive. Um, and furthermore, something that's really scary about that is if you don't have a typed library with, with static typing, where the computer can say what you're trying to do is wrong, then the LLM will try to make an API

40:11

request that is wrong some percentage of the time. The code execution tool can run a type checker and say, oh, you know, you're asking about stripe.transactions.list, but that actually doesn't exist. Stripe doesn't have a transactions API. You might want payment intents, you might want orders, you might want balanced transactions, which one do you want? And if you, if the API provider is doing a great job building this tool, it'll return the documentation for all of these things in line. It might have its own AI, look at what the model's trying to do and come up with a suggestion. And that, and that sub agent, you know, is well-trained, specified, always updating,

40:49

SPEAKER_01

and isn't burdened with the context of the full conversation. What do you think of the security model?

40:56

SPEAKER_02

The security model is really, really interesting. This is another area where we're really starting to think about things at stainless and I'm getting really excited about it. So if any listeners are really interested in this and have some ideas or want to talk, you know, please do reach out. At the end of the day, I think the security has to take place at the API layer itself. Right now you see people trying to implement security by sort of limiting what's exposed through MCP. And that kind of makes sense, but at the end of the day, you, you, you could do anything that's in the API under the hood. Right. Um, and what people should be doing is using

41:41

SPEAKER_02

OAuth with granular permissions with, with, with, um, with proper scopes. And at that point, the security happens the right place, which is at the API layer. Um, there's limitations to OAuth scopes. Um, and it's pretty hard to build. Um, so it'd be nice if someone made that easy, but, um, um, my view, that's kind of the, that direction is sort of the right, the right layer. So going back to my, my earlier question, I'm, I'm, I'm thinking about the idea of having a model write code that then the API provider executes to, you know, interact with their API and then returns the results. Would you ever consider just creating a tool use tool that developers use? Because like,

42:27

SPEAKER_02

for example, I'm thinking about for Quora, got all these tools, maybe Gmail is gonna build, you know, like a code use thing or whatever, but really I just want, um, I would probably use what you're talking about inside of Quora, but we, we would need a, uh, a tool use tool or, but it's not a tool use tool. It's like a, it's a computer, it's a computer use tool where, and I know OpenAI has this, but it's not really well built for, for, for, for lots of libraries and stuff. It's not a custom environment. Like I need a computer use tool where I control the environment and I can install different

43:03

SPEAKER_02

libraries in it and, uh, be able to call it any time to, to then call any, any API, or it has to have network access basically. Yeah. You guys should build that. We're working on it. Fuck yeah. You're building it for, for, for, uh, for developers who want to access MCP servers or people who are providing MCP servers. We're starting with people who are providing MCP servers, but ultimately I think that we're going to need this to work such that, um, you can give the model a code execution

43:33

SPEAKER_01

environment where it can hit not only the Stripe integration, but also the Salesforce integration and also anything else. Um, and, but not too much anything else. Right. And so one of the advantages of

43:43

SPEAKER_02

starting where we're starting of just one API provider is that you ensure that there's no network connections allowed out of that sandbox where we're running the code to anything other than in this case, api.stripe.com. Um, and that's, that's really, really critical for security for something like this. Um, and so there's ways to expand that bit by bit, um, and keep things and keep things secure. Um, uh, it'll, it'll, it'll take some time. The other thing I think to point out as you see some of these generalizations is it's not just that you want this like code execution sandbox to work really well

44:19

SPEAKER_02

for any API for any library. Um, which I think we really do. I think, I think we really need that. Um, you also start to see that this is just a powerful model for AI doing stuff. And sometimes you, you want, you realize that the thing that the AI did this one time in this one-off case is actually enduringly useful. Maybe anytime a customer writes into support and says, Hey, my socks had holes in them. You should automatically get a refund. You know, um, maybe you want that, maybe you don't, but there's a lot of stuff that people do one or one time and then two times and then three times.

44:58

SPEAKER_02

And then they say, okay, we should automate this. Right. And that's, and that's what software teams do all day, every day. Right. And we're going to be, I think we're also going to be seeing that with AI where the same, the same code search tool that we're talking about, all the same prompting that will make an AI really, really good at interacting with an API in one of these code sandboxes, kind of like almost quote unquote in its brain, um, working like write code in its head, run the code in its head, see the results and then move forward with your, with your, with your query, um, with your task.

45:30

SPEAKER_02

Uh, it should be able to say, okay, actually this is enduringly useful code. Let me commit this to the repo. Yeah. Yeah. Yeah. Yeah. It's like, uh, you know, um, chat is a really good interface for exploring, but sometimes you just want a dashboard, you know, you just, I just want to like log into my Stripe dashboard and see all the stuff without having to be like, what is my MRR? It should just show up, you know, cause I just do that every day. Um, but I want to, I want to push you as a, as a hashtag value add investor. Um, because I, I think that there's a, um, I think that there's this thing that

46:03

SPEAKER_02

happens in AI where often the first attempt at something like this, um, people try to be really cautious and I'm sure that your customers care about you being cautious, like big enterprise customers, but the things that get adopted are often the ones that are willing to take the risk to be YOLO very early. So an example is, um, Dolly was like totally private for like a long time and people were like posting some images, but you couldn't get in. And then, uh, stable diffusion was just like, fuck it. Like anyone can use this. And then that just really started the whole image generation wave, obviously stable diffusion sort of fumbled the bag, but they had a lead for a

46:42

SPEAKER_02

little while. Um, same thing for, for cloud code, honestly, like if you look at, uh, codex is not like this as much anymore, but if you look at the difference between codex, CLI and cloud code, cloud code was just like, fuck it. Like YOLO mode. It's super industrious. It has a sandbox, but you can just do dangerously skip permissions and codex just fell way behind because it was first, it was in the browser. And so their whole thing was, the whole thing was like locked down. And then it was in the, it was in the, in the CLI, but it was really built for pair programming. And so it just wasn't

47:15

SPEAKER_01

particularly industrious. It wouldn't go off and do a bunch of stuff. It didn't, it would get locked out of doing certain things, even if you did full auto mode. Um, and now they've like caught up because they're, they're like, yeah, you can just let it do whatever you want. And so I would, I would really push you on, there might be a version that you could do like today or tomorrow or like very soon for individual developers that would let them set up this environment that for example, I would use like immediately. And I, I, I care about security, but I care, I care a lot less than some X, you know,

47:47

SPEAKER_01

gigantic enterprise company. But I think the people like me who are building at this scale are eventually hopefully going to be the big companies, but we're the ones that are really doing the AI first adoption, not the big, not the big companies. Well, I would love to get this in your hands. What are some of the APIs your team uses the most? Um, I'm, I'm thinking as we have a bunch of different products, but I'm thinking right now about Quora, the email, the email assistant. Um, and, uh, it has all of the, like the, the big APIs that it's using, it's mostly the Gmail, the Gmail API. Um, and so you're

48:21

SPEAKER_01

interacting with the assistant over chat and then it has a list of tools that are like, you know, archive email or draft email or send email or whatever. Uh, like there's a whole categorized tool. So it categorizes your mail mail in certain ways. And I think we would definitely try out something like this because it would, if it, if it ran the same way, um, it, it would make it much more flexible for us to make more tools and not break old ones. You know, um, it's really interesting. I mean, in a sense, what, what I actually predict is that people who are quote unquote building tools,

48:59

SPEAKER_01

once we have a code execution kind of super tool, like I'm talking about is that the only way you really quote unquote build a tool is with instructions, with prompts. Um, and the full power of everything you could possibly do in the API, in the Gmail API, for example, um, it's all there in one tool. Um, but sometimes you have specific tasks, uh, or specific, you know, categories of, of work that you want to describe in a particular way, um, to help the LLM perform a sequence of actions as productively as possible. Um, and at that point, the only work in engineering that you have to do is,

49:39

SPEAKER_01

is prompt engineering. Um, we'll see if it's, we'll, we'll see if it's that quote unquote easy. Um, uh, as we all know, prompt engineering can be, can be really tricky. It's hard. Yeah. But, um, but I think, I think that's, that's part of the vision. Um, that, that being said, you know, we, we do have, uh, some pretty nifty ways with the MCP servers that we generate today to help developers mix and match, um, all the parts of the different tools, um, underlying, um, all the different parts of the API, um, as they compose and write their own tools. This is awesome. Um, so for people who are listening, uh, and want to know more from you or and know more from stainless,

50:15

SPEAKER_01

uh, where should they find you? Um, stainless.com, um, our, is it, that's, that's our website. Awesome. Or at least visit stainless.com. Uh, Alex, great to have you on. I can't wait to do more of this, uh, when you have some of these new things launched. This is really, really fun. And, uh, yeah, great to, great to chat. Thanks, Dan. You too. Oh my gosh, folks, you absolutely positively have to smash that like button and subscribe to AI and I, why? Because this show is the epitome of awesomeness. It's like finding a treasure chest in your backyard, but instead of gold, it's filled with

50:55

SPEAKER_01

pure unadulterated knowledge bombs about chat GPT. Every episode is a roller coaster of emotions, insights, and laughter that will leave you on the edge of your seat, craving for more. It's not just a show. It's a journey into the future with Dan Shipper as the captain of the spaceship. So do yourself a favor, hit like, smash subscribe, and strap in for the ride of your life. And now, without any further ado, let me just say, Dan, I'm absolutely hopelessly in love with you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note