Skill issue: Lessons from skilling up coding agents to use Langfuse - Marc Klingen, Clickhouse
Description
Without a skill, Claude Code adds Langfuse using stale pre-training context, ships broken instrumentation, then catches the failure and fetches current docs to fix it. The resulting trace captures two LLM calls with no visibility into what the agent actually did. Marc Klingen covers the six learnings from building a skill to close that gap: surfacing a natural language search endpoint so agents stop crawling 478 documentation pages, why pointing to references beats duplicating content, and what happened when they ran an auto-research loop on the skill itself. Three of six suggested improvements shipped, but their target function nearly backfired by optimizing out the documentation-fetching steps that make the skill reliable over time. Speaker info: - https://x.com/marcklingen - https://www.linkedin.com/in/marcklingen/
Summary
Generated by claude-haiku-4-5-20251001Skill Issue: Lessons from Skilling Up Coding Agents to Use Langfuse
Main Topics
- AI Agent Skills Development: Building effective Claude Code skills to help users integrate observability tools
- Observability & Evaluation Infrastructure: Langfuse's approach to tracing and evaluating AI agent behavior
- Agent-Assisted Setup: Using coding agents to automate complex documentation and implementation tasks
- Skills vs. Workflows: The balance between autonomous agent capabilities and predefined workflows
- Documentation at Scale: Solving the problem of outdated training data in LLM models
Key Points
The Problem Space
- Documentation Overload: Langfuse had 478 pages of documentation, making it difficult for users to implement correctly
- Outdated Training Data: Claude and other coding agents hallucinates methods based on pre-training data rather than current APIs
- Non-Optimal Implementations: Agents implement features suboptimally because they lack current best practices
- Slow Iteration: Users often implement wrong solution first, then discover the error and need to refetch documentation
The Skill Concept
- Skills as Formalized Shortcuts: Skills provide agents with up-to-date context and guides for reliable implementations
- From Workflows to Progressive Context: Skills allow agents to handle multi-domain problems that historically required separate workflows
- Expert Guidance: Skills should feel like having an expert user guide you through the right approach for your specific problem
Six Major Learnings
- Tracing Gets You 80% of the Way
- Analyzing execution traces in Langfuse revealed where agents were failing
- Understanding runtime behavior was more useful than guessing at improvements
- Production Signals Matter
- Real user behavior revealed unexpected needs (e.g., US enterprises caring about data regionality, not just Europeans)
- Default assumptions proved wrong and needed adjustment
- Hallucinated CLI Parameters
- Agents assumed CLI capabilities based on loose associations with keywords like "trace"
- Solution: Advertise the help flag more aggressively to let agents discover actual capabilities
- Navigation Through Information
- Create agent sitemaps and search endpoints rather than expecting agents to fetch multiple documentation pages
- Use content negotiation (markdown format requests) to reduce token overhead
- Implement RAG-based search for natural language queries about documentation
- Basic Evaluation Setup Beats None
- Started with five different evaluation templates for different use cases
- Used LLM-as-judge to verify implementations before/after skill execution
- Even simple setups provided measurable value
- Reference Dynamic Content, Don't Duplicate It
- Avoid caching documentation in skills as it becomes outdated
- Better to point to live documentation and fetch updates
- Prevents multiple versions of truth across the system
Auto-Research Learnings
- Target Function Criticality: The target function determines what the agent optimizes for—poorly defined targets lead to useless optimizations
- Example: Minimizing turns caused agent to skip documentation fetching, defeating the purpose
- Better targets should include desired outcomes like "link prompt versions to production traces"
- Approval Gates Matter: Need to ask users before executing sensitive operations (pushing data to repositories)
- Sandbox testing made this difficult to validate
- Depth vs. Speed Trade-off: Should skills aim for initial quick wins or perfect implementations from the start?
- Quick wins: Get user to "aha moment" faster
- Perfect setup: Requires asking many questions upfront, but no need to iterate later
Notable Quotes
> "Skills are a formalized shortcut to make things more reliable where you historically would have built a workflow."
> "Just looking at traces gets you to 80% of the detail."
> "The skill is the primary way of how things get done... Nobody reads documentation themselves. Everyone is just like, 'yeah, just add this to my, I just want this to work.'"
> "The target function really matters. It sounds obvious, but for us, defining the right target function was very hard."
> "You only need to ask if the skill is more trusted than the public web. If it's the same trust level, then why even bother asking?"
> "It should feel like an expert user trying to guide you through what you need for your problem."
Takeaways
For AI/LLM Application Builders
- Use tracing and observability to understand what agents actually do—this reveals improvement opportunities better than speculation
- Build for real user behavior, not assumptions—production signals reveal unexpected use cases and requirements
- Don't duplicate dynamic content—reference live sources instead and implement freshness detection (timestamps)
- Define clear target functions when optimizing agent behavior—vague goals lead to useless optimizations
- Reduce cognitive load by providing navigation aids (sitemaps, search endpoints, chat interfaces) rather than expecting agents to search the web
- Start simple but provide pathways to deeper setup—allow users to iteratively improve implementations
For Product Strategy
- Skills represent the future of how users interact with complex products—direct agent integration rather than documentation reading
- Package management for skills is still unsolved—need better distribution mechanisms
- Small teams can use auto-research patterns to explore improvement ideas at scale
- Approval gates and human oversight remain important for sensitive operations
Implementation Recommendations
- Timestamp skill content and alert users when information is stale
- Implement search endpoints for natural language documentation queries
- Create basic evaluation templates for common use cases
- Use LLM-as-judge verification to test skill outcomes
- Advertise tool capabilities (help flags) prominently to prevent hallucination
Future Direction
- Moving from skills as setup tools to skills as automation engines for ongoing evaluation and iteration
- Goal of full orchestration agents that manage entire evaluation lifecycles without human intervention
- Integration of community feedback to continuously improve skill guidance
Transcript
Okay, that was quick. Hi, everyone, super excited to be here. I'm Mark, one of the founders of Langfuse. When you started Langfuse three years ago, when everything felt quite early, building ages that didn't work, and then realized, okay, there needs to be some evaluation tracing, built like this is the open source project in this space by now. And we met the metrics that we track. We seem to be the largest one in the space. We do our product engineering out of Europe. Thus, I'm very excited this conference is coming to Europe, because there are so many great people here. And we always need to resist the urge to ship the whole team to another continent, because actually being here is very nice. And you can just travel and hang with people on Discord and Zoom. So yeah, very excited to be here in person. And what I want to talk about today is what the lessons we made from skilling up coding agents and actually adding Langfuse to an application, because back in the days when you add observability or evals, you needed to read hundreds of pages of docs, figure out your own mental model. And now you expect to be hand-holded by a custom agent to do this. And we've come a long way in achieving this kind of vision. And I just want to explain what we learned on the way. First of all, I start very conceptual, very easy, then a bit more conceptual, deeper, and then I'll go to the learnings. So my mental model for skills is just like I get this Rubik's cube. When I was a kid, I had no idea what to actually do. But I have a bash tool. I can do whatever I want with this Rubik's cube, but it just looks colorful in different ways. I had no idea how to solve it. So skills was a great way. Once you get the manual, it's easy. You just need to follow the manual and you can solve a Rubik's cube. And I feel the same thing now applies to agents, where there was this whole debate of workflow versus a fully autonomous agent. And it was a huge fight on X of what is the best way to build an application. And I was kind of like, yes, you need both. And I think Malta this morning had a good note where the surface area of deploying agents is so broad that for some, you don't need a coding agent, even if it's the best way you can build agents at the front of today. You don't need this for every application. It makes things slow, expensive. And there's this balance between workflow being very reliable and agent now having unlimited capabilities. And I think what's very exciting is that skills are a formalized shortcut to make things more reliable where you historically would have built a workflow. I don't know, you have a customer support agent and someone asked for a password reset. Then historically you would have built a workflow that's very reliable of, ah, you have a router that routes to an agent that can only do password resets. And that agent has the context to do password resets well. That was great. But also if the user then wants to do password reset, but also change the email address at the same time, then the router is kind of, okay, I have this email router. I have this password reset router. What do I even do? And I think what's very exciting now is that an agent can just progressively get the context needed to then solve a problem that's multi-domain. That would have historically been in multiple workflows. So that's very exciting. However, it has always been hard to kind of build an agent, what kind of use cases even exist because you have this open-ended text input box often, or open world context. So you didn't really know. So what we now see is what many teams do is they have the agent runtime. They trace everything. I mean, I'm building Langfuse this. I put my logo here, but you can also use whatever you want. In the end, it's more about the concept of tracing is what you need to identify what happens at runtime when a trace, when an agent is executed, because then it helps you learn two things. One, new use cases that you didn't expect users to do, because you might have expected that nobody ever wants to change a password, but now they want to change a password. This you need to see the execution trace of someone being upset by a production eval to then derive that you need to add a skill for handling password resets well. And two, once you have these skills, they can get out of date, or you can realize they're not the most efficient way of actually solving for this use case. So it's the second thing of helping you improve skills that you already have in your agent. So this is very conceptual, but this is what we now see most teams do that use Langfuse. Now more towards the learnings we made when building a skill to help customers add Langfuse to their project. So what was the lay of the land before we got started with this? Like 478 pages of documentation. Whenever I see the thing deploying, I'm thinking, who wrote all of this? So apparently if you build a project over three years, it just grows in complexity. People can do all of these different things, but then you need to read all of these things, and nobody has the time across five different feature areas. And a lot of implementation flexibility because there have always been projects in the eval space that were very, I'd say, opinionated. So they were, oh, you have a chat bot, then add this project, and it'll just solve this for you, opinionated end to end. We always were, no, no, we are infrastructure. We do tracing well. If you ingest billions of traces, it will still work. If you want to customize your evals, it'll still work. So we were always more on the unopinionated side, which was always, I'd say, a weakness compared to projects that are more opinionated. But now I think it's a strength because in the end, what do you need when agents do all of this? You only need the infrastructure piece if agents can then customize for different workflows. So, but there's the problem if people want to add Langfuse to a project. We always were like, no, no, we are infrastructure. We do tracing well. If you ingest billions of traces, it will still work. If you want to customize your evals, it'll still work. So we were always more on the unopinionated side, which was always, I'd say, a weakness compared to projects that are more opinionated. But now I think it's a strength because in the end, what do you need when agents do all of this? You only need the infrastructure piece if agents can then customize for different workflows. But there's the problem if people want to add LangFuse to a project. How to do it correctly for a project is to be figured out by an agent. And interestingly, when we first got into model pre-training context as a project, and you can just ask an agent how to add LangFuse, and it'll spit out LangFuse SDK logic. The first time this happened, it's amazing. But then if you're two years in, the project evolves, interfaces change. And now being in pre-training context might even be a disadvantage if you don't fetch up-to-date information. So we're just like, oh, we get all of these hallucinations of methods that have been available in the past, but they're not available there today. So, yeah, we felt like when skills launched, this is the exact pattern that we need in order to help teams achieve this. So I'll use an example. When you just ask Claude Code to add LangFuse to a project, it just worked, but it was not working in the best way possible. So, for example, user asks add tracing to my agent, and then Claude Code implements the instrumentation based on the outdated pre-training context, then tries to verify whether the tracing works, then realize, oh, it doesn't work. And then only in a second step, fetches up-to-date information to then correct the issue. At the same time, how you add tracing or evaluation to a project, you can evaluate in gazillion different ways, such as online evals, offline evals, human in the loop. There are so many different things, and often the question is what even is relevant for your application. So humans and agents kind of need to figure it out on the way, but agent is not tasked to help you figure out what's the best thing for your application. So main problems: outdated training data, the non-optimal setup because the agent wasn't really primed to help you discover what to do for your app. And it's very slow because you first add instrumentation in the wrong way, then you figure out it's wrong, and then you need to fetch more documentation to fix the issues. So what did we do? Oh, yeah. This is how the trace looked like when we just tried with Claude Code. So just tracks two LM calls in an agent, but you still don't know what the agent's actually doing. So what was the goal of our skill? Give every LangFuse user—there are thousands of teams in the community, thousands of customers on our Cloud product—give them all a LangFuse expert to help them quickly set up observability, prompt management, evals, in line with best practices and up-to-date docs and references. Because all of you, I mean, there's Annabelle from the team here as well. If you have questions regarding observability evals, you can talk to us. But in the end, that doesn't scale to thousands of people to basically talk through your problems and figure out what the best strategy is for you. So we were thinking, okay, skills is the way to go. And this is very conceptually how our skill works, where user comes in, asks coding engine to do something, and then the skill kind of has a reference of the skill MD which is more about what kind of style do we want in order to implement LangFuse? So for example, ask follow-up questions before making decisions, because there's so much you can be doing. And then references for the different product modules to kind of progressively disclose additional hints that the agent might need to have. And then it can curl the documentation. And interestingly enough, as we started open source and saw ourselves as unopinionated infrastructure, we always had APIs for everything, because teams built their own, I don't know, labeling UIs on top of it, evaluation execution logic on top of our backend. And we had APIs for everything. Now we've wrapped it in the CLI, and now an agent can just do everything humans needed to do in the UI in the past, which is very cool, because so many teams spend so many hours every week clicking around in our UI to evaluate and improve the application. And in the end, how will this look like end of year? It'll probably just be like, connect repository to LangFuse, and then agent just does the whole thing autoregressively. I mean, that's what we are building towards. That's what everyone is building towards. And I think that's a cool step in the right direction. So to shortcut to the end result, after conversing with the agent now for a similar thing, it looks way more detailed—detailed evals that are relevant and detailed steps regarding tool execution. So there's a stark difference. And yeah, what did we learn on the way? Six main things. I'll go through every single one of them—what are basically our realizations when we're building the skill. And one, looking at traces still gets you to 80% of the detail. This was always what we kind of tried to preach regarding evals, where many people try to complicate things right away while they haven't dug through. So, just, what did the agent actually do at runtime themselves a couple of times. So what do we do? We have instrumentation for Claude Code and just ourselves interactively tried to use LangFuse with Claude Code and then look through traces in LangFuse to understand where did the agent error? How can we improve the skill to make it a straight shot at the goal instead of wandering in different ways to the target? And that was really helpful too. There were some interesting learnings here. For example, for humans, we tried to cut down on the number of environment variables that you need to set in order to set up LangFuse. So for example, we just auto assumed a data region—LangFuse is available in Europe, is available in the US. Fun anecdote: we assumed that only Europeans care about data regions. Thus, we made Europe the default. Then we learned some US enterprises also care about data regionality. Now we have a US data region and so many other different data regions. So we always defaulted to Europe and now, for an agent adding another environment variable, they don't care. It's not effort for them. Thus, we always prompt to figure out what data region the user is actually in and don't assume Europe, for example. Two, hallucinated CLI parameters because it just includes the word trace. Fun anecdote, we assumed that only Europeans care about data regions. Thus, we made Europe the defaults, then we learned some US enterprises also care about data regionality. Now we have a US data region and so many different other data regions. So we always defaulted to Europe and now for an agent like adding another environment variable, they don't care. It's not effort for them. Thus, we always prompt for figure out what data region the user is actually in and don't assume Europe, for example. Two, hallucinated CLI parameters because it includes the word trace. I have seen tracing CLIs before. I just assume what we could be doing here and we just advertise the help flag more aggressively. It takes another turn, but it's fast and thereby it directly knows what the CLI can do. Two, we try to help the agent to understand how to navigate available information because there are five front documentation pages, how to find the right one instead of looping through, fetching one, then learning something, then fetching another one, always with thought process in the meantime. So what would we do? We always had this LLMSDXT, which was very hype when it launched but never actually used. I think what's now cool is we have this agent sitemap that we just expose to a coding agent by the skill of going there first in order to learn what kind of documentation is available. And two, there's this whole content negotiation that if you send a request header that you want markdown you get markdown back from the docs, but some coding agents don't do this by default. So we just advertise this because otherwise some coding agents might try to pass the HTML, which just adds additional tokens. So for Langfuse, for example, you can just add .md to any documentation page or you can request markdown and you'll get a markdown page. Three, I think that was one of the things I was most excited about. We always had this docs Q&A agent that was able to answer questions about Langfuse more interactively. Therefore, we built a RackStack. And now we just surface this RackStack by a search endpoint. So a coding agent can just ask whatever natural language query about Langfuse and we'll get back documentation chunks for this query. Why is this exciting? One, you don't need to fetch five different docs pages where you can just ask a question, get something back that's relevant directly, solve for the problem. And two, we get to track these search parameters. Because if a coding agent fetches documentation, it's very difficult to understand what did Cloud Code and our user laptop do. But if they ask questions about Langfuse to our search endpoint, we can track the searches and thereby understand what problems do they run into. Where do we need to add more documentation pages because maybe we didn't expect this kind of problem to happen. So adding a search endpoint was really cool to capture more data. Then basic eval setup is better than none because we initially struggled to get this done because it's so broad. Some Langfuse users built chat applications, real-time voice, video generation, batch processing of invoices in the background of some kind of text software. So many different use cases where then the question is, what's even a good evaluation setup. And we just created five different ones. And this was already helpful. Because otherwise, it's really hard to measure anything. And what we did here, I don't know, can I zoom in? No. So basically, we have this prompt instrument application with Langfuse and then a sample repository folder. So for example, an openAI custom function, reg, whatever application. And our checks are just natural language statements that we then, by alum as a judge, try to evaluate on top of the file system and dev state before and after running the skill. So for example, we expect that our openAI instrumentation was added because it's an openAI example. And we, because it's reg, we expect some retrieval spans to show up in our trace. Because if there are no retrieval spans, then probably we only capture alum calls. This was already helpful because then we were able to make changes and see that we didn't break anything. The whole thing that why we even build Langfuse for building AI agents now also applies here. So, five, dynamic content should be referenced because there's a huge incentive for developers on the team, but also for users in the community to just contribute a lot of context to the skill. Because then it's a local cache of the documentation that's immediately available. However, then the same thing applies that applies to pre-turning context. It goes out of date. And now we have the documentation and now we have yet another representation of what Langfuse is. So you'd rather try to point just straight to the reference of documentation and, because otherwise you just duplicate all content. And six, we applied auto research to the skill of, okay, if we have a target function, how can agents help us improve the agent? Because there are so many different patterns that we can explore. So we set up a target function mostly to get towards our experiment here was help teams move prompts from their local Git repository into Langfuse prompt management, which is used by larger teams to collaborate on prompts with their non-engineering counterparts. Because then PMs can make changes to prompts, iterate on a playground, all of this collaborative stuff. And the task was, okay, how do we improve the skill to migrate prompts out of any code base into our managed prompt system? And in the end, we accepted three out of the six improvements that were suggested, which I think is a success. But it allowed us to experiment much more than we could have explored manually with the time that we have, as we are a very small team. Learnings, the target function really matters. I think it sounds obvious, but for us, defining the right target function was very hard for this. Because we assumed a trunk migration should be fast. Fast, we measured in the number of turns. But if we basically asked to minimize the number of turns, then our agent that tried to optimize the skill just took out all of the nodes that we had to fetch documentation. Because it was, I know how to how Langfuse prompt management works. I don't need this. I'll just try it myself. Which then negates the whole thing of we want to fetch up to date context. Because otherwise, if you use the skill, install the skill once, wait three months, then you'll have wrong context because we duplicate information. Two, we had an approval gate usually, where we want to suggest a plan or ask follow-up questions, suggest plans to a user before doing anything. Because we push their prompts to a center repository and it's their data leaving their laptop somewhere else. Because I know how Langfuse prompt management works. I don't need this. I'll just try it myself. Which then negates the whole thing of we want to fetch up to date context. Because otherwise, if you use the skill, install the skill once, wait three months, then you'll have wrong context because we duplicate information. Two, we had an approval gate usually, where we want to suggest a plan or ask follow-up questions, suggest plans to a user before doing anything. Because we push their prompts to a center repository and it's their data leaving their laptop somewhere else. But the sandbox didn't have this, so we weren't really able to try for this. And Langfuse, the sole feature, usually we try to make it easy to get going with something, but then it's very deep how to do it in a good way. And we want agents to directly go for the good way, figure out with a user what they want to achieve and then have a very full implementation. Not start with something and then two months later go deeper. However, if the target function does not include—we want linking prompt versions to prod traces. So then you can see how different versions impact, for example, production results. We didn't have this in the target function. Everything that nudged towards this was kind of removed because it's just garbage on the way that we don't need to achieve the goal. So again, the target function really matters. High level, these were the six main takeaways. Looking at traces gets you 80% of the way. The production signals really helped. The search endpoint was really helpful for our documentation. Help agent to navigate the information because otherwise it just searches with Google, Brave, whatever search and finds all sorts of different things on the internet. Even the basics, Evil setup helps. It wasn't that hard to set up. The dynamic content should be referenced. Otherwise you have just duplicates. And the auto research was very helpful to explore things, but it's bound by the target function. Topics in our minds here are, it's so powerful, but at the same time you duplicate stuff into user space, somewhere on a machine. There's no package management for this, which then a year tells user this is outdated. We thought about just adding a timestamp of the current date where it was fetched by the skill. And then if this is older than a month, then try to update. But then we go to the second problem of skill distribution and installing into the agent environment. Usually this is gated or not possible for the agent, depending on what you use. The user needs to do something to install the skill. Auto upgrading doesn't really work, but it really depends on the coding agent that you use. And the target function is interesting for us because we can either go for user needs to get to an initial aha of, oh, this works. Or do we want to directly shoot for the perfect setup of how you would do evals for this use case. But without a skill, it takes usually an AI engineering team months to get to a perfect setup. Do we aim for an agent to do this in a single shot and overload the user with lots of questions? Or do we just try to get to something and then you can still invoke it again of, improve my setup, and then it can ask questions to improve it. So it's what's the target for the skill? That was very interesting for us. I would invite you to try it and give us feedback because it'll be really interesting. We do lots of calls with people from the community every week. And I think it's not a surprise that nobody reads documentation themselves. Everyone is just like, yeah, just add this to my, I just want this to work. So the skill is the primary way of how things get done. This is also now the advertised way across all of our documentation that you just should ask your coding agent to do whatever you do. So that's what we should try to do right now. I'm very excited that it works really well. But also I'm excited to see what comes next. For us as a project roadmap wise, we see the skill right now. Our users use this when getting started with the project, but also to drive a lot of automation around the evaluation lifecycle of, oh, I now want to create an element of a judge that's aligned with user preferences. Or I got user feedback on 100 different executions. What do they have in common? And do we then fetch this via the CLI. So many of these workflows that people needed to do manually now coding agents do for them. We'll bring this in product via, we will help automate this via skills, one, bring this in product, two, and then three. I feel we just need this orchestration agent and that's what the team is doing right now. So I'm very excited for our roadmap to automate all of this. But if you have any feedback, I'm around, Annabelle's around. I would love to talk to you. And thanks so much for your time. Do we have time for a question? Okay. Yep. I may have to ask this. You say that when you were tuning the skill, the human was completely out of the loop. So you were basically in third code, directly trying to implement using the skill of the right implementation and then you measure this? Or? Yeah, it was kind of, I mean, you kind of want to be out of the loop for the experimentation and then just review the suggested changes. So it was experimental things, give us all sorts of different recommendations, and then human review the suggestions. We didn't accept all because many didn't make sense because our target function wasn't perfect. It was really difficult to get to a very perfect target function. But it's good at just creating ideas and then we human reviewed all of the ideas to make the changes to the skill. But so at the end, the skill might not be optimizing for the human AI interaction. Like, let's say, how do we work with learning friends and to try to set up, to try this for whichever product. Then the skill might be optimizing for something in the current, and whether it's very automated or I don't know. Yeah, that's what we try to kind of. You need to try it yourself to just get a sense of how it feels to use the skill and then add language to an application. It was really difficult to get to a very perfect target function. But it's good at just creating ideas and then we human reviewed all of the ideas to make the changes to the skill. But so at the end, the skill might not be optimizing for the human AI interaction. Let's say, how do we work with learning friends and to try to set up this for whichever product. Then the skill might be optimizing for something in the current, and whether it's very automated or I don't know. Yeah, that's what we try to do. You need to try it yourself to just get a sense of how it feels to use the skill to then add language to an application. So we just use it ourselves to get a sense for the feeling because it should, where we want to go is it should feel an expert user trying to guide you through what you need for your problem. [SPEAKER_02] Where usually someone comes in with I need evals because I read about it online. [SPEAKER_02] But I don't know what actually I need for my application. [SPEAKER_02] And it needs guidance of where you want to go. What is your problem? I don't know. What do you worry about? You probably don't need a hallucination eval. But probably you need something that's very specific to your application. And yeah, that's what we want to achieve with a skill that you get some professional guidance. Yep. I really resonated with your last point about skills distribution seems frankly insane right now of just, you just install it and do whatever it is on the table. So what are your thoughts on treating skills as packages, a skills kind of approach? Or going all in on plug-in marketplaces and said, what do you think is able to be adopted by the community versus being able long term and to have prominence? Hmm. As a small team, I'm not excited about plug-in marketplaces because then you now need to maintain all of these proprietary integrations, update them in, I don't know, Anthropic, OpenAI, Cursor, whenever you make an [SPEAKER_01] . [SPEAKER_01] I don't think in OpenAI now agreeing on . [SPEAKER_01] Yeah. [SPEAKER_02] Still, I mean, for the skill, I think it would be cool if we just had a well-known skill or something. And whenever someone is, oh, I want to, for example, use Langfuse, the agent can just auto-discover that it exists. We have it across all of our docs, so I think it would be enough if agent can ask user, I want to install skill. The question is do you even need to ask? I think you only need to ask if the skill is more trusted than the public web. If it's the same trust level, then why even bother asking? And then two is if I have this installed, it's a cache of something that was up-to-date when I installed it. But then the question is how do I know whether it's out-of-date? So I think just timestamping it is enough. So when you use the skill, that agent can be like, oh, this seems old. I should probably fetch a new one. I think this would already go a long way. But yeah, I'm excited to see. We are going more the timestamp fetch route or alert user of this might be out-of-date. That's at least what we discussed now. But yeah, I'm excited to see what everyone is shipping in this space. Yeah, I'm around. Thanks so much. Bye-bye. Bye-bye. And this was already helpful. Because otherwise, it's really hard to kind of like measure anything. And what we did here, I don't know, can I zoom in? No. So basically, we have this like just like a prompt instrument application with Langfuse and then like a sample, a repository folder. So for example, like an openAI custom function, reg, whatever application. And our checks are just natural language statements that we then, by alum as a judge, try to evaluate on top of the file system and div state before and after running the skill. So for example, we expect that our openAI instrumentation was added because like an openAI example. And we, because it's reg, we expect like some retrieval spans to show up in our trace. Because if they, if there are no retrieval spans, then probably we only capture, for example, alum calls. This was already helpful because then we were able to make changes and see that we didn't break anything. The whole thing that, why we even build Langfuse for like building AI agents now also applies here. So, five, dynamic content should be referenced because there's a huge, I'd say, incentive for like developers on the team, but also for users in the community to just contribute a lot of context to the skill. Because then you're like, ah, it's kind of like a local cache of the documentation that's immediately available. However, then the same thing applies, that applies to pre-turning context. It's kind of like, it goes out of date. And now we have the documentation and now we have yet another representation of what Langfuse is. So, you'd rather try to point just to straight to the reference of documentation and, because otherwise you just duplicate all content. And six, we applied, I'm off like auto research to the skill of, okay, if we have a target function, how can like agents help us improve the agent? Because there are so many like different patterns that we can explore. So, we set up a target function mostly to get towards like, our experiment here was help teams move prompts from their local Git repository into Langfuse prompt management, which is used by like larger teams to collaborate on prompts with their non-engineering counterparts. Because then like PMs can make changes to prompts, iterate on a playground, like all of this kind of like collaborative stuff. And the task was, okay, how do we improve the skill to migrate prompts out of any kind of like code base into our managed prompt system? And in the end, we accepted three out of the six improvements that were suggested, which I think is a success. But it allowed us to experiment much more than we could have explored manually with the time that we have, as we are like a very small team. Learnings, like the target function really matters. Like, I think it sounds obvious, but for us, defining like the right target function was very hard for this. Because we assumed like a trunk migration should be fast. Fast, we measured in like the number of turns. But if we basically asked to minimize the number of turns, then like our, like the agent that tried to optimize the skill just took out all of the nodes that we had to like fetch documentation. Because it was like, I know how to how Langfuse prompt management works. I don't need this. I'll just try it myself. Which then negates the whole thing of we want to fetch up to date context. Because otherwise, if you use the skill, install the skill once, wait three months, then you'll have like wrong context because we duplicate information. Two, like we had like an approval gate usually, where we want to suggest a plan or ask follow-up questions, suggest plans to a user before doing anything. Because we kind of like push their prompts to like a center repository and it's kind of like their data leaving their laptop somewhere else. But the sandbox didn't have this, so we didn't really, we weren't really able to try for this. And like Langfuse, like the sole feature, like usually we try to make it easy to get going with something, but then it's very deep of how to do it in a good way. And we want agents to directly go for the good way, like figure out with a user what they want to achieve and then have like a very full implementation. Not start with something and then like two months later go deeper. However, if the target function does not include, like we want like linking prompt versions to prod traces. So then you can see how like different from versions impact like for example production results. Like we didn't have this in the target function. This like everything that like nudge towards this was kind of like removed because it's kind of like, it's just like garbage on the way that we don't need to achieve the goal. So again, the target function really matters. High level, these were like the, the six main takeaways. Looking at traces gets you 80% of the way. The production signals really helped. So the search end point was really helpful for our documentation. Help agent to navigate the information because otherwise it just searches with like Google, Brave, whatever search and finds all sorts of different things on the internet. Even the basics, evil setup helps. It wasn't that hard to set up. The dynamic content should be referenced. Otherwise you have just duplicates. And the auto research was very helpful to explore things, but it's bound by the target function. Topics basically in our minds here are, it's so powerful, but at the same time you kind of then duplicate stuff into like user space, kind of like somewhere on like a machine. Like there's no like package management for this, which then like a year, like tells user this is outdated. Like we could, we, we, we thought about just adding like a timestamp of this, the current date where it was fetched, the skill. And then just, oh, if this like older than a month, then try to update. But then we go to second problem of skill distribution and like install, installing into like the agent environment. Usually this is kind of like gated or not possible for the agent, like depending on what you use. This user needs to do something to install the skill. This also upgrading doesn't, auto upgrading doesn't really, really work, but it really depends on the coding agent that you use. And like the target function is interesting for us because like we can either go for user needs to get to like an initial aha of like, oh, this works. Or do we want to directly straight shoot for this, the perfect setup of how you would do evals for this use case. But this is, I mean, without a skill, it takes like usually like an AI engineering team. It takes like months to get to a perfect setup. Do we now aim for an agent to do this in a single shot and overload the user with lots of lots of questions? Or do we just try to get to something and then you can still invoke it again of like improve my setup, ask, and then it can ask questions to improve it. So it's kind of like, well, what's, what's the target for the skill? That was very interesting for us. Yep, I would invite you to try it and give us feedback because they'll be really interesting. And like we do lots of lots of calls with people from the community every week. And like, I think it's not a surprise that I think nobody reads documentation themselves. And everyone is just like, yeah, just add this to my, like, I just want this to work. Like just add it. So yeah, the skill is the primary way of how things get done. This is also like now the advertised way across all of our documentation that you just should ask your coding agent to do whatever you do. So that's what we should try to do right now. I'm very excited that it works really well. But also I'm excited to see what comes next. For us as a project roadmap wise, we see the skill right now. Like our users use this when getting started with the project, but also to drive a lot of automation around the evaluation lifecycle of, oh, I now want to create like an element of a judge that's aligned with user preferences. Or like I got user feedback on 100 different executions. What do they have in common? And do us then fetch this via the CLI. So many of these workflows that people needed to do manually now coding agents do for them. We'll bring this in product via, like, we will help automate this via skills one, bring this in product two, and then three. I feel like we just need this orchestration agent and that's what the team is doing right now. So yeah, I'm very excited for our roadmap to automate all of this. But if you have any feedback, I'm around, Annabelle's around. I would love to talk to you. And yeah, thanks so much for your time. I don't know if, do we have time for a question? Okay. Yep. I may have to ask this. You say that when you were tuning the skill, the human was completely out of the loop. So you were basically like, in third code, directly trying to implement using the skill of the right implementation and then you measure this? Or? Yeah, it was kind of, like, I mean, you kind of want to be out of the loop for the experimentation and then just review the suggested changes. So it was kind of, like, experimental things, give us, like, all sorts of different recommendations, and then human review the suggestions. Like, we didn't accept all because many didn't make sense because our target function wasn't perfect. It was really difficult to get to, like, a very perfect target function. But it's good at just creating ideas and then we human reviewed all of the ideas to make the changes to the skill. But so at the end, the skill might not be optimizing for the human AI interaction. Like, let's say, how do we work with, like, learning friends and to try to set up, like, to try this for whichever product. Then the skill might be optimizing for something in the current, and whether it's very automated or I don't know. Yeah, that's what we try to kind of, like, you need to try it yourself to just get a sense of how it feels to use the skill to then, like, add language to an application. So we just use it ourselves to get a sense for the feeling because it should, like, where we want to go is it should feel like an expert user trying to guide you through what you need for your problem. Where usually someone comes in with just, like, I need evals because I read about it online. But I don't know what actually I need for my application. And I kind of, like, it needs guidance of where you want to go. Like, what is your problem? I don't know. What do you worry about? You probably don't need, like, a, I don't know, hallucination eval. But probably you need something that's very specific to your application. And, yeah, that's what we want to achieve with a skill that you get, like, some, like, professional guidance. Yep. I really resonated with your last point about skills distribution seems frankly insane right now of just, like, you just install it and do whatever it is on the table. So, like, what are your thoughts on, like, the treating skills as packages, like a skills kind of approach? Or, like, going all in on, you know, plug-in marketplaces and said, like, what do you think is, you know, able to be adopted by the community versus being able long term and to have, like, prominence? Hmm. As a small team, I'm not excited about plug-in marketplaces because then you now need to kind of, like, maintain all of these proprietary integrations, update them in, I don't know, teleanthropic, teleopenei, telecursor, whenever you make an . I don't think in opening air now agreeing on sort of, like, . Yeah. Still, I mean, like, for the skill, I think it would be cool if we just had, like, a well-known skill or something. And, like, whenever someone is, like, oh, I want to, for example, use LengFuse, like, the agent can just auto-discover that it exists. Like, we have it across all of our docs, so I think it would be enough if agent can kind of, like, ask user, I want to install skill. The question is do you even need to ask? Like, I think you only need to ask if the skill is kind of, like, more trusted than the public web. If it's, like, same trust level, then why even bother asking? And then two is kind of, like, if I have this installed, it's kind of like a cache of something that was up-to-date when I installed it. But then the question is how do I know whether it's out-of-date? So I think just, like, timestamping it is enough. So when you use the skill, that agent can be like, oh, this seems old. I should probably, like, fetch a new one. I think this would already go a long way. But, yeah, I'm excited to see. Like, we are going more the timestamp fetch route or alert user of this might be out-of-date. That's at least, like, what we discussed now. But, yeah, I'm excited to see what everyone is shipping in this space. Yeah, I'm around. Thanks so much. Bye-bye. Bye-bye.