AI Engineer

How Building with AI Can Double the Throughput of Your Engineering Team — Brian Scanlan, Intercom

1015 summary words 5 min summary Watch video

Start with the signal

5 min read

Summary

How Building with AI Can Double the Throughput of Your Engineering Team

Main Topics

  • Intercom's AI Transformation: Pivoting to an AI-first company and achieving significant revenue growth
  • The 2x Project: A year-long initiative to double engineering throughput using AI
  • Platform Strategy: Choosing Claude Code as the unified AI development platform
  • Cultural & Organizational Change: Leadership approaches to driving AI adoption across hundreds of engineers
  • Practical Implementation: Building AI skills, automating code review, and measuring productivity
  • Future of Engineering Work: How engineers' roles will evolve as AI capabilities increase

Key Points

Business Results & Impact

  • Intercom has achieved doubling of PR throughput faster than their one-year goal
  • Revenue approaching $100 million from their AI customer support agent (Finn)
  • Maintained strong growth trajectory while competitors declined
  • Processing approximately 2 million resolutions per week with custom LLM models
  • Defects closing faster than ever despite increased identification

The 2x Project Foundation

  • Primary metric: Code changes per R&D person
  • Set ambitious goal in mid-2023: double productivity without doubling team size
  • Decision to consolidate on Claude Code made December 2023; rollout started January 2024
  • Timing coincided with major improvements in model coding capabilities

Organizational & Leadership Strategies

  • Clear messaging: Updated job descriptions making AI adoption non-negotiable
  • Dedicated team: Full-time Team 2x that continues growing
  • Celebration culture: Public recognition of automation wins, skill updates, and techniques in Slack channels
  • Learning events: Hackathons and AI immersion days
  • Consistent communication: Repeating messaging 100+ times across different forums about urgency

Platform & Technical Approach

Why Claude Code:

  • Platform choice matters more than which tool
  • Avoids "multi-cloud" fragmentation that prevents optimization
  • Enables compounding benefits through deep integration

Vision: Make Claude act like a senior engineer across all technical tasks

  • Connect Claude to everything engineers can access
  • Teach Claude Intercom-specific knowledge (Rails conventions, architecture, testing standards, security rules)
  • Continuous improvement flywheel: use platform → hit issues → update guidance → repeat

Productivity Metrics & Proof Points

  • Pull requests: Auto-completion in 90+ percentage range
  • Code review approvals: 17.6% of PRs auto-approved with high confidence
  • Safety & compliance: Maintained SOC 2, ISO 27001, and HIPAA compliance without human approval loops
  • Code quality: Metrics from Stanford research group show improving code quality over time
  • Skill utilization: Hundreds of contributors creating thousands of lines of agent code

AI Adoption Maturity Levels

  • Initial use of Cloud Code for basic tasks
  • Automation of repetitive work
  • Converting automation into reusable skills
  • Mastering skills and improving them
  • Optimizing entire environment for agent effectiveness (architecture, documentation, processes)

Key Principles for Implementation

  • Give agents problems, not tasks: Describe what needs solving; let agents determine which skills to invoke
  • Build small, durable, testable skills: Focus on quality over breadth
  • Continuous self-updating: Skills improve based on feedback loops and data
  • Stay on proven platforms: Use Anthropic, OpenAI releases rather than building everything in-house
  • Document work processes: Writing down "how to do work" remains valuable regardless of tool changes

Real-World Example: Security Incident Response

  • Security incident required analyzing data breach policies after accidental Snowflake metadata publication
  • Simply described problem to Claude Code without specifying steps
  • Agent autonomously found relevant policy skill, downloaded files, performed analysis, concluded innocuous data, provided next steps
  • Time saved: 2 minutes vs. 20 minutes of manual work
  • Demonstrates agent capability to discover and invoke appropriate skills

Future-Oriented Changes

  • Role evolution: Engineers moving "up the stack" like sysadmins → SREs transition
  • Cross-functional expansion: Claude Code going "viral" across product, design, non-engineering teams
  • Single-person teams: Experiments with solo product managers shipping features using AI skills
  • Broader impact: Acting as product managers and wearing multiple hats

Notable Quotes

> "Shipping is the heartbeat of your company."

> "If you're not adopting AI in Intercom, whether you're a designer, product manager, engineer, whatever, you are not meeting expectations, binary."

> "You don't get the compounding benefits of a well-designed platform if you're sending all your different work across different cloud providers."

> "Everything that you can do, the agent must be able to do."

> "I just gave it the problem of taking a look at security incident. It just figured out the intent and used a well-written internal skill that did this job for me."

> "We're speed running this 100 times faster on a full industry scale."

> "Give problems to agents, not tasks."

> "If you're not doing pretty much all of this today, you're going to be doing it in the very near future."

Takeaways

For Organizations

  • Commit fully to one platform rather than maintaining multiple tools—optimization comes from depth, not breadth
  • Make AI adoption a non-negotiable expectation across all roles and reinforce through hiring, evaluation, and culture
  • Dedicate senior talent full-time to AI enablement; don't treat it as an add-on responsibility
  • Measure what matters: Use metrics like code changes per person, not vanity metrics
  • Build compliance in from the start: Automation doesn't require human-in-the-loop for SOC 2, ISO, HIPAA compliance with proper controls

For Engineering Teams

  • Start with small, well-defined skills rather than attempting comprehensive solutions
  • Create feedback loops: Use session data, backtesting, and human labeling to continuously improve agents
  • Document processes thoroughly: Codify "how to do work at [company]" for agents to leverage
  • Describe problems, not tasks: Guide agents toward autonomous decision-making about which skills to invoke
  • Optimize for agent capability: Consider software architecture, documentation, and processes from the perspective of AI automation

Broader Insights

  • Engineers' work is fundamentally changing: Focus shifts from coding to problem-solving and system design
  • The capability curve is accelerating: Even without further model improvements, current tools enable substantial SDLC transformation
  • Agent-first development is achievable today: Tools are mature enough now; don't wait for perfect models
  • Risk can decrease through well-designed automation: Properly configured agents outperform inconsistent human judgment
Full transcript 3636 words · 33 min read
0:14

SPEAKER_00

Hey, I'm Brian from Intercom and this has been such a great conference so far. I've learned so much from the talks and all the chats with people. So Intercom is a 15-year-old privately held Irish-American B2B SaaS startup that pivoted to be an AI company the weekend ChatGPT came out. We've got about 1400 people across Dublin, London, Berlin, SF, Chicago, Sydney. R&D is led from Dublin. Engineering is almost entirely across Europe. Forward deployed engineers have changed that up a bit. And this graph compares our revenue growth to the growth rate of publicly traded SaaS companies over the last few years. And you can see publicly traded SaaS companies on the down. Intercom is amazing growth. And we're bucking this downward trend that SaaS companies have been suffering from recently. And now I'm going to shut down Intercom live on stage. I once did a live deployment in a talk and I thought that was impressive.

0:20

SPEAKER_00

So Intercom has become the poster child for companies redefining themselves in the age of AI. The New York Times recently did an article about SaaS companies reinventing themselves that prominently featured Intercom. And being an AI company means a lot more than just slapping on lightweight wrappers or autocompleting some text field. Our AI agent for customers, Finn, has over 8,000 customers, industry leading average resolution rates, revenues approaching 100 million, launched the day GPT-4 came out, first product actually released on GPT-4. And we've been building AI features since about 2018. And the modern LLM models have unlocked huge capabilities for dealing with customer support questions, completely obvious. And companies like Anthropic, Snowflake, Linear, Glean, LaunchDarkly use Finn for their customer support. So maybe SaaS isn't dead. And it works well for all size businesses. And also, we recently announced that we have our own model serving 100% of Finn, English text-based conversations, outperforming frontier models like Claude. Cheaper, faster, better. And we're at about 2 million resolutions a week. And we're also happy to sell direct access to our suite of models. I'm not talking about any of this, though. So I'm a senior principal engineer at Intercom, been there for 12 years, and I'm in our platform group. And we take care of Intercom's uptime, performance, security, cost management, observability, our majestic monolith applications that we love, mostly Ruby on Rails, and all internal developer productivity. And another thing about Intercom is that we are obsessed with shipping. Shipping fast and iteratively is the best way to build high quality products that customers love to use. And so developer productivity is something we've always invested in. Shipping is the heartbeat of your company. It's a great blog post that we did many years ago, and Honeycomb made cool stickers of it. And obviously, for the last few years, I've been spending a lot of time on enabling use of AI in our software development lifecycle. So I'm going to talk about that. And so, unsurprisingly, we've been very excited about AI in general. We've changed the whole company to build customer support using AI agents. And we've been impatient about getting its adoption and changing how we build across Intercom. And we went down some familiar routes. We're all using GitHub Copilot, and then everyone started adopting Cursor, and we looked at Augment and a few other things. But ultimately, middle of last year, we've been dissatisfied with the results. Some good signs, some tasks made marginally better and more fun. But we're pretty aware of where the models are going and the harnesses. And we have a strong conviction that AI, from many years ago, is going to change all knowledge work. So last year, middle of last year, we set a simple goal. Let's double the throughput of engineering in a year. And we measure a lot of things in Intercom. We use developer surveys. We use tools like DX. But we picked code changes per R&D person as the primary way we're measuring productivity. Every measure is bad once you start measuring it. But we expect the overall throughput to increase. If we're actually adopting new ways of working, putting AI into all the different places, then we should expect a large throughput increase. And so 2x, what we call this project, 2x, the name of the project and our team and everything. This is wildly ambitious. When we published this back last June, doubling productivity without doubling team size. But also wildly unambitious as well if you connect the dots and see where the models and coding harnesses are going. So in this talk, I'm going to talk about how we went about this, how we think about productivity and a sneak peek at some of our internal data and skills. And this also coincided with the most notable shift in model capability and coding capability. And we've all seen it. And this was one of our principal engineers posting just what everyone else was saying around the Christmas break last year, saying, oh my god, things have changed massively. And that has contributed a lot to our success on 2x.

0:26

SPEAKER_00

So this is the engineering leadership part of the talk. You need to be decisive and give clear executive guidance and do organizational change. And we've done a lot of things. We updated job descriptions. If you're not adopting AI in Intercom, whether you're a designer, product manager, engineer, whatever, you are not meeting expectations, binary. And you have to say the same message over and over, 100 times, in every different forum. You just have to stay on message and constantly talk about the urgency of this. You have to reward people as well. When people do good stuff, in all the Slack channels, showing automating, when people update skills, it gets put into these channels, we celebrate, people are showing each other different techniques and what's working for them. We've done hackathons, we've done AI immersion days, and all of these things are necessary to bring people along. We staff this full time. We have a Team 2x that keeps growing and growing. And we're not just saying, hey, you have to AI everything, best of luck. We're trying to bring everyone, the hundreds of engineers, hundreds of people in R&D, along with us. So if you're in a medium or large

0:32

SPEAKER_00

These channels, we celebrate stuff, people are showing each other different techniques and what's working for them. We've done hackathons, we've done AI immersion days, and all of these things are necessary to bring people along. We staff this full time. We have a Team 2x that seems to keep growing and growing. We're not just saying, hey, you got to AI everything, best of luck. We're trying to bring everyone—the hundreds of engineers, hundreds of people in R&D—along with us. So if you're in a medium or large organization, you absolutely need to have people—your best people—on this full time.

0:38

SPEAKER_00

We chose Claude Code as our platform. Prior to this, we were omnivorous and letting people choose their favorite editor. There are loads of people adopting Claude Code, loads of people using Cursor, loads of people using Augment. We're a believer in platforms in general. It doesn't matter what you choose, but choosing one is important. To a certain extent, you need to get away from model anxiety. It's like being multi-cloud. You don't get the compounding benefits of a well-designed platform if you're sending all your different work across different cloud providers. You're way better being all in on one and optimizing it and proving that it works, unless there are very specific or impactful reasons why you need to be spread across multiple agents.

0:43

SPEAKER_00

Our vision was to treat Claude, to work on getting Claude to be able to act like a senior engineer on any technical task across Intercom. Our vision here was to connect Claude to everything. Anything I do on my laptop, Claude should be able to do that. Of course, we're not reckless. We're not trying to let the thing delete all of our databases. We're a mature company. We've got plenty of controls and permissions and audits that give us confidence to unleash Claude the same way we unleash our engineers in our environments.

0:48

SPEAKER_00

We've got to onboard it. We've got to teach it all the stuff we teach people when they join Intercom—our Rails conventions, our architecture, React patterns. We've built a lot of software in 15 years. Testing standards, security rules. Claude absolutely has to know the Intercom-specific information to do the job. Most importantly, start using the platform for all technical work. It doesn't get things right first time. It hits an issue, goes down the wrong path. Update the guidance. This is a flywheel that we're all contributing to.

0:54

SPEAKER_00

We've encapsulated a lot of this knowledge in context and engineering. Captured the skills, guidance, hooks to force these things. We spent a lot of time cajoling Claude Code to work well. We do things like push out our internal Claude plugins to everyone's laptops, bypassing all the Claude Code update mechanisms. You spend a lot of time debugging Claude Code installs on hundreds of laptops. It's like trying to manage Python installs.

1:00

SPEAKER_00

Ultimately, every single part of technical work is in scope. It's not just code production. It's not more advanced auto-complete. It's everything. Debugging, testing, planning. It should just be you driving Claude and ideally driving it less and less, moving higher up the food chain. It delivers real value—code, products, whatever—to customers.

1:08

SPEAKER_00

We think that even if the models and harnesses do not improve at all, which is definitely not happening, the capability curve is accelerating. We have the building blocks today to move vast amounts of work in our software development life cycle to be agent first. We could just pause everything and we've got this flywheel. We're going through everything and looking at every single piece of work. The tools are good enough today to do this.

1:13

SPEAKER_00

There are some principles to help guide us. When you're trying to get hundreds of people to change how they work or understand what we're trying to achieve, you need to write things down and help them out. Different principles should apply in different places. We believe that all of engineering is changing. Everything that you can do, the agent must be able to do. Our job is moving up the stack as engineers, as product builders.

1:19

SPEAKER_00

A long time ago, I used to be a Unix sysadmin, going out to data centers, racking servers, cabling things, configuring networks. Then the cloud came along and I moved up the stack. People transitioned from being sysadmins to SREs. The work was more automation oriented, more impactful, higher paid. I think we're speed running this 100 times faster on a full industry scale. I kind of feel like I've been through this before.

1:23

SPEAKER_00

We at Intercom are technically conservative. We like using single tools and using them extremely well. We end up with these Ruby on Rails monoliths. We're applying this thought process to where should our focus be? Where is our attention? Do we want everyone writing their own multi-agent orchestrators or opinionated workflows? We want to build durable, testable, high quality components, and people to be considering the lifetime value of what they produce. The tools, the specific implementations, will change over time. But I'm pretty sure that writing down how to do work in Intercom will be valuable no matter what happens.

1:30

SPEAKER_00

In practice, we spend our time focusing on small, high quality, durable, testable skills that do the job extremely well. We use data, use backtesting. We've got all of the work—this huge body of work and changes in code and incidents—and we're using all of this to help form us and prove

1:37

SPEAKER_00

these things will change over time. But I'm pretty sure that writing down how to do work in Intercom will be valuable no matter what happens. Maybe it might be easier to discover in the future. That's a problem at the moment. And so what this means in practice is that we spend our time focusing on small, high quality, durable, testable skills that do the job extremely well, that we can use data, use backtesting. We've got all of the work. We've got this huge body of work and changes in code and incidents and everything. And so we're using all of this to help form us and prove out that these skills are operating at extremely high quality. And we then we, and we also practice continuous improvement here, get these things to be self-updating, and make sure that these things are very high quality. And yeah, we don't want to get stuck behind the curve, getting stuck because we've implemented a load of our own things. We just want to use things as they become available as on Tropic Ship or whatever. And maybe we might not stay on Tropic forever, but we're very eager to get the advantage of somebody else building and shipping great software and capabilities rather than us having to build everything ourselves. So, yeah, another thing we guide people to do is you want to give problems agents, not tasks. A lot of the time people even say in Intercom are saying, prompting agents, hey, run this skill to do a thing, which is mostly fine and still necessary. I still do it a lot, but we're more having to move ourselves to be just describing the problem or just describing the task and let the agent figure out what skills to invoke and what to do here. And I have a fun story. Recently I was brought into a security incident. We had accidentally published some kind of Snowflake table metadata to a public GitHub repository. And I just habitually opened Cloud Code, told us to join a Slack channel, take a look. And I didn't even know that a skill existed that actually perfectly encapsulated all of our data breach policies and criteria and what to do, how to analyse this. Cloud just automatically downloaded the files, did full analysis, concluded it was innocuous, told me all next steps. And I didn't tell it to do this. It just figured it out. It was done in two minutes. And that would have been a 20-minute task and boring work. I'd have to go, oh, where's that policy? And take a look at this, that, and the other. And this just felt, it was a small example, but again, I just gave it the problem of taking a look at security incident. It just figured out the intent and used a well-written internal skill that did this job for me. And yeah, I mentioned it, even at Intercom, AI adoption is unevenly distributed. I think we're ahead of the vast majority of companies. But you still need to help people understand where they're at and grow towards being highly effective at using agents in their work. Steve Yeaghy recently talked about maturity rating for engineers. And our internal one is kind of similar here. You're trying to get through these different levels. And ultimately, you end up mastering all skills and knowing the tool inside out. And ultimately, what we want people to do is use Cloud Code for everything, automate your work, then move that to a skill, then get really good at writing skills. And then writing skills and improve the skills. And then optimize the environment for agents. That could be everything from software architecture, maybe just to documentation, but other approaches or other ways of doing things that allows the agents to be even more effective and optimized for what they're great at today. So, here's where we're at. You can see, yeah, wild inflection points after going all in on one tool. That decision was made in December. We started rolling it out in January. And we've been just, we have reached doubling PR throughput in faster than one year. Here's more data from our internal dashboards. There's some interesting stuff in here. There's, yeah, number of pull requests, auto-reclaw code. It's in the 90-somethings. You can see, also, we're starting to move into our current bottleneck is code review. But you can see we have this 17.6% approval rate of our automatic code approvals. And it's a lot more in-depth than just, hey, Claude, can you approve this? We've gone through a lot of detailed work to figure out, again, using backtesting and previous data, and then getting humans to label the outputs and figure out, get the confidence level of the automatic approvers and shape the pull requests towards very safe and simple pull requests, which probably always should have been that way. But now they're just approved automatically. And we've also worked with our auditors to ensure that we're fully SOC 2, ISO 27001, HIPAA compliant, all that. You do not need humans in the loop to meet these certifications. You do need to know exactly what you're doing, though, and make sure you've got auditing controls and everything. And so, by moving approvals to an extremely well-organized, tested and competent suite of agents, including codecs for code reviews, I think multimodal code reviews are okay. I just completely went back on my platform thing. And we've got a high confidence that this stuff is not degrading environment or adding additional risk. In fact, I think it's removing risk because humans aren't actually as good as agents when they're well-defined. Here's skill invocation. I actually think the earlier numbers were a bit wonky. So we hook up everything into Honeycomb. We've got hooks all over the place for basic information about which skills are being invoked and things like that. And that's internally available. There's no private information in this. And everyone can use it to get an idea of what's being used and where. But we also pull in all session transcripts into S3 for data mining, writing reports, also looking to see our skills effective, that kind of stuff. So we've got a feedback loop using the session data, which we can get more out of, but we're doing some interesting stuff with it already. And this isn't a goal, but we're not particularly proud of, defects always increasing up until recently. But defects are getting closed faster than ever. And some teams have been inspired by the move to AI to think about things like backlog zero or crunching through

1:45

SPEAKER_00

it to get an idea of what's being used and where. But we also pull in all session transcripts into S3 for data mining, writing reports, also looking to see our skills effective, that kind of stuff. So we've got a feedback loop using the session data, which we can get more out of it, but we're doing some interesting stuff with it already.

1:51

SPEAKER_00

And this isn't a goal, but we're not particularly proud of defects always increasing up until recently. But defects are getting closed faster than ever. And some teams have been inspired by the move to AI to think about things like backlog zero or crunching through hundreds or thousands of defects. So some of this was a bit deliberate and planned, but just in the same time, there's this natural deflation because getting through this work, getting through all the defects so much faster these days. And yeah, we're just seeing this naturally. We've also been working with Stanford. There's a research group there. We give them all our code. And our code quality per their metrics has been increasing over the last while. Okay. I'm running out of time at this point. We have hundreds of contributors, thousands, thousands and millions of code in our cloud plugins. It's very active. And yeah, Claude itself loves it. Here's an example skill. This is not the most—or sorry, we've got base plugins, things that do all the session transcripts, session syncing, some safety hooks and things. And here's a skill I built, which just fixes flaky specs. We have hundreds of thousands of tests, and they get flaky over time. And we don't, we ship a lot. So we just barge through the flakes. But this skill was not built by me sitting down and figuring out what are all the things you need to do to fix flaky specs. I've worked in a feedback loop, gave the agent a goal. And through guiding us to the right place and working with us to fix a lot of flaky specs, it's written this pretty decent thing. What are these cheat codes or lookup tables? And relatively well organized, using progressive disclosure and all that. And it is fixing stuff that if our most senior Rails engineers were doing this, I'd be wow, they're amazing. And yeah, a lot of other stuff going on. Like, our CI melted. We had to fix that. Cloud code is actually widely used across Intercom outside of software. It's gone completely viral. People are banging down our doors to use consoles. And yeah, we're thinking a lot about the future of engineering. Should we just merge all product manager design, everything? Oh yes, the single person team product experiments have been pretty interesting as well. And I've even been shipping codes, stuff that people can use in their agents to sign up to Intercom. This is stuff that I've just been using our skills to act as a product manager, which is pretty wild. So that's it. I wish you all the best of luck. If you're not doing pretty much all of this today, you're going to be doing it in the very near future. My contact details are at brian.scannon.ie. You can interact with Finn in the messenger, configure by CLI. And you can check out ideas.fin.ai for a lot more information about Intercom and our agents. Thank you.

1:59

SPEAKER_00

unlocked huge capabilities for dealing with customer support questions, completely obvious. And companies like Anthropic, Snowflake, Linear, Glean, LaunchDarkly use Finn for their customer support. So maybe SaaS isn't dead. And it works well for all size businesses. And also, we recently announced that we have our own model serving 100% of Finn, like English text-based conversations, outperforming frontier models, like Sanit. Cheaper, faster, better. And we're at like about 2 million resolutions a rate, resolutions a week. And we're also happy to sell direct access to our suite of models. I'm not talking about any of this,

2:38

SPEAKER_00

though. So I'm a senior principal engineer at Intercom, been there for 12 years, and I'm at our platform group. And we take care of Intercom's uptime, performance, security, cost management, observability, our majestic monolith applications that we love, mostly Ruby on Rails, and all internal developer productivity. And another thing about Intercom is that we are obsessed with shipping. Shipping fast and iteratively is the best way to build high quality products that customers love to use. And so, developer productivity is something we've always invested in. Shipping is the heartbeat of your

3:10

SPEAKER_00

company. It's a great blog post that we did many, many years ago, and Honeycomb made cool stickers of it. And obviously, for the last few years, I've been spending a lot of time on enabling use of AI in our software development lifecycle. So I'm going to kind of talk about that. And so, unsurprisingly, we've been very excited about AI in general. We've changed the whole company to build customer support using AI agents. And we've been impatient about getting its adoption and changing how we build across Intercom. And, you know, we went down some kind of familiar routes. You know, we're all using

3:44

SPEAKER_00

GitHub Copilot, and then everyone started adopting Cursor, and we looked at Augment and a few other things. But ultimately, you know, say, middle of last year, we've been dissatisfied with the results. Some good signs, some kind of tasks, some work made marginally better and kind of more fun. But, you know, we're pretty aware of where the models are going and the harnesses. And we have a strong conviction that AI, like, from many years ago, is going to change in all knowledge work. So last year, middle of last year, we set a simple goal. Let's double the throughput of engineering any year. And, you know, we measure a

4:23

SPEAKER_00

lot of things in Intercom. We use, like, we do a lot of developer surveys. We use tools like DX. But we picked code changes per R&D person as the primary way we're measuring productivity. Every measure is bad. Once you start measuring it, it's not a measure and all this. But also, like, we're, like, impatient about or, like, expect the overall throughput to increase. Like, if we're, like, actually adopting new ways of working, putting AI into all of the different places, then we should expect a large throughput increase. And so 2x, what we call this 2x, the name of the project and our team and everything. This is, like, wildly ambitious. Like, when we published this back

5:07

SPEAKER_00

last June or something like that, doubling productivity without doubling team size, but also kind of wildly unambitious as well if you, like, connect the dots and see where the models and coding harnesses are going. So in this talk, I'm going to talk about how we went about this, how we think about productivity and a sneak peek at some of our internal data and skills and stuff. And, you know, this also coincided the work here with, like, the most notable shift in model capability and coding capability. And so, you know, we've all seen it. And this was, like, one of our principal engineers

5:39

SPEAKER_00

posting just kind of like everyone else was in and around the Christmas break last year, going, like, oh my god, like, things have changed massively. And so that has contributed a lot to our success on 2x. So this is the kind of engineering leadershipy part of the talk. And so you need to be decisive and give clear executive guidance and, you know, do organisational change. And we've done a lot of things. We updated job descriptions. If you're not adopting AI in Intercom, whether you're a designer, product manager, engineer, whatever, you are not meeting expectations, binary. And yeah, you have to say

6:17

SPEAKER_00

the same message over and over and over, 100 times, every different forum, whatever, you just got to stay on message and constantly talk about the urgency of us doing this. You got to reward us as well. Like, when people do good stuff, you got to, like, all the Slack channels, showing, like, automating where people, like, automating, when people update skills or do this, that, and the other, it's like, it gets put into these channels, we celebrate stuff, people are showing each other different techniques and what's working for them and that kind of thing. We've done hackathons, we've done AI immersion days, and, you know, all of these

6:52

SPEAKER_00

things are necessary to kind of bring people along. Like, also, we staff this full time. We have a Team 2x that seems to be just keeps on growing and growing and growing. And, you know, we're not just saying, hey, you got to AI everything, best of luck. We're, like, trying to bring everyone, like, the hundreds of engineers, hundreds of people in R&D along with us. So, you know, if you're in a medium or large organization, you absolutely need to have people and, like, your best people on this full time. And so we chose Claude Code as our platform. So prior to this, we were kind of omnivorous and, like,

7:31

SPEAKER_00

letting people choose their favorite editor and this, that, and the other. And, you know, there's, like, loads of people adopting Claude Code, loads of people using Cursor, loads of people using Augment. But, like, we're a believer in platforms in general. And it kind of doesn't matter what you choose, but choosing one is important. You know, to a certain extent, you need to get away from model anxiety. It's like being multi-cloud. It's like you don't get the compounding benefits of a well-designed platform if you're sending all your different work across different cloud providers or whatever.

8:02

SPEAKER_00

And you're way better being all in on one and optimizing it and proving that it works. And, like, unless there's, like, very specific or impactful reasons why you need to be spread across multiple agents or whatever. And so our vision on this was, like, to treat Claude or to, like, work on to get Claude to be able to act like a senior engineer on any technical task across of Intacom. And our vision here was, like, connect Claude to everything. So anything I do on my laptop, Claude should be able to do that. And that means everything. Like, now, of course, we're not reckless. We're not, like, just trying to let the thing go off and delete all of our databases. But

8:38

SPEAKER_00

we're, like, a mature company. We've got plenty of controls and permissions and audits and everything like that that gives us a lot of confidence to be able to, like, unleash Claude in the same way that we unleash our engineers in our environments. And, you know, we've got to onboard it. We've got to teach it all the stuff that we teach people when they join Intercom. All of our Rails conventions, our architecture, React patterns. Like, we've built a lot of software in 15 years. But, like, testing standards, security rules, all this. Claude absolutely has to know the Intercom-specific

9:08

SPEAKER_00

information to be able to do the job. And most importantly, start using the platform for all technical work. And it doesn't get things right first time. Hits an issue. Goes down the wrong path. Update the guidance. Like, this is a flywheel that we're all contributing to. And so we've encapsulated a lot of this knowledge in context and engineering. Captured the skills, guidance, hooks to force these things. We spent a lot of time cajoling Claude code to work well. We do things like push out our internal Claude plugins to everyone's laptops. Like, bypassing all the Claude code updates mechanisms.

9:43

SPEAKER_00

Because it's, you know, you spend a lot of time debugging Claude code installs on, like, hundreds of laptops. It's like trying to install Python or manage Python installs or something. And so, ultimately, though, like, every single part of technical work. So, it's not just code production. It's not, like, more advanced auto-complete. It's everything. So, debugging, testing, planning, all this kind of stuff. It should just be you driving Claude. And ideally, like, driving it less and less and moving higher up the food chain. And it, you know, delivers real value. Delivers the code,

10:17

SPEAKER_00

products, whatever, to customers. So, everything's in scope. And, like, we think that even if the models and harnesses do not improve at all, which is not, definitely not happening. If anything, like, this capability curve is accelerating. But, like, the building, we have the building blocks today to to improve, like, basically move vast amounts of work in our software development life cycle to be agent first. Like, they could just pause everything. And we've just got this flywheel. And we're going through everything and looking at every single piece of work. And, like, the tools are good enough

10:50

SPEAKER_00

today to do this. So, there are some principles to help guide us along the way. You know, when you have, you're trying to get hundreds of people to change how they work or understand what we're trying to achieve, you need to write things down and help them out. And, you know, different principles should apply in different places. But, like, you know, we believe that all of engineering is changing. Everything that you can do, the agent must be able to do. And that can feel weird as well, like, when you're first connecting us into production systems, whatever. And, yeah, like, our job is moving up the

11:23

SPEAKER_00

stack as engineers, as product builders, whatever. And, like, a long time ago, I used to be a Unix sysadmin. And, you know, like, going out to data centers, racking servers, cabling things, configuring networks and all that. And then the cloud came along. And I moved up the stack. You know, and people transitioned from being sysadmins to SREs. The work was more automation oriented, more impactful, higher paid. And so I think this is, like, we're kind of speed running this 100 times faster on a full industry scale. But I kind of feel like I've been through this before. And we at Intercom are

12:00

SPEAKER_00

technically conservative. We like using single tools and just using them extremely well. So, hence, we end up with these Ruby on Rails monoliths and stuff. And so we're kind of applying this thought process as well to, like, you know, what is the, where should our focus be? Where is our attention? Do we want everyone writing their own multi-agent orchestrators or opinionated workflows? And, you know, we want to build durable, testable, high quality components, and people to be considering, like, the lifetime value of what they produce. And, like, you know, the tools, the specific implementations of

12:30

SPEAKER_00

these things will change over time. But I'm pretty sure that writing down how to do work in Intercom will be valuable no matter what happens. Maybe it might be easier to discover in the future. That's, like, a problem at the moment. And so what this, what this means in practice is that we spend our time focusing on small, high quality, durable, testable skills that do the job extremely well, that we can, you know, use data, use backtesting. We've got, like, all of the work. We've got this huge body of work and changes in code and incidents and everything. And so we're using all of this to help form us and prove

13:03

SPEAKER_00

out that these skills are operating at extremely high quality. And, you know, we then we, and we also practice continuous improvement here, get these things to be self-updating, and make sure that these things are very high quality. And, and, yeah, we don't want to get stuck behind the curve, like, getting stuck because we've implemented a load of our own own things. We just want to use things as they become available as on Tropic Ship or whatever. And maybe we mightn't stay on Tropic forever, but, like, we're very, we're eager to get the advantage of somebody else building and shipping great software and capabilities rather than us having to build everything ourselves.

13:39

SPEAKER_00

So, yeah, another thing we guide people to do is, like, you want to give problems agents, not tasks. You know, a lot of the time people even say in Intercom are saying, like, prompting agents, hey, run this skill to do a thing, which is mostly fine and still kind of necessary. I still do it a lot, but, like, we're more kind of having to, like, move ourselves to be kind of just describing the problem or just describing the task and let, let the agent figure out what skills to invoke and what to do here. And I have a fun story. Recently I was brought into a security incident. We had accidentally published

14:10

SPEAKER_00

some kind of snowflake table metadata to a public GitHub repository. And I just habitually opened Cloud Code, told us to join a Slack channel, take a look. And I didn't even know that a skill existed that actually perfectly encapsulated all of our, like, data breach policies and criteria and what to do, how to analyse this. Cloud just automatically downloaded the files, did full analysis, concluded it was innocuous, told me all next steps. And I, like, I didn't tell it to do this. It just kind of figured it out. It was done in, like, two minutes. And, like, that would have been a 20-minute

14:45

SPEAKER_00

task and kind of boring work. I'd have to go, oh, where's that policy? And take a look at this, that, and the other. And, like, this just felt like a little, like, it was a small example, but it's, like, again, I just, like, gave it the problem of, like, taking a look at security incident. I just figured out the intent and used a well-written internal skill that did this job for me. And, yeah, it, I mentioned it, even at Intercom, like, AI adoption is unevenly distributed. I think we're ahead of the vast majority of companies. But you still need to help people understand where they're at and grow towards

15:16

SPEAKER_00

being highly effective at using agents in their work. Steve Yeaghy recently talked about, like, maturity rating for engineers. And, like, our internal one is kind of similar here. You're kind of, like, trying to get through these different kind of levels. And, like, ultimately, you kind of end up mastering all skills and, like, knowing the tool inside out. And, ultimately, like, what we want people to do is, like, use cloud code for everything, automate your work, then move that to a skill, then get really good at writing skills. And then writing skills and improve the skills. And then

15:42

SPEAKER_00

optimize the environment for agents. That could be everything from software architecture, maybe just to documentation, but other approaches or other ways of doing things that allows the agents to be even more effective and optimized for what they're great at today. So, here's where we're at. You can see, yeah, wild inflection points after going all in on one tool. That decision was made in December. We started rolling it out in January. And we've been just, like, we have reached the doubling PR throughput in faster than one year. Here's more, like, data from our internal dashboards. There's some interesting

16:17

SPEAKER_00

stuff in here. There's, like, yeah, number of pull requests, auto-reclaw code. It's, like, in the 90-somethings. You can see, also, we're starting to move into, like, our current bottleneck is code review. But you can see we have this, like, 17.6% approval rate of our automatic code approvals. And it's, like, a lot more in-depth than just, like, hey, Claude, can you approve this? We've gone through a lot of detailed work to figure out, again, using backtesting and previous data, and then getting humans to kind of label the outputs and figure out, like, get the confidence level of the automatic approvers and kind of shape the pull requests towards very safe and simple

17:02

SPEAKER_00

pull requests, which probably always should have been that way. But now, like, they're just approved automatically. And, you know, we've also worked with our auditors to ensure that we're fully SOC 2, ISO 27001, HIPAA compliant, all that. You do not need humans in the loop to to meet these certifications. You do need to know exactly what you're doing, though, and make sure you've got, like, auditing controls and everything. And so, by moving approvals to an extremely well-organized, tested and competent suite of agents, including codecs for code reviews, I think multimodal code reviews are okay. I just, like, completely went back on my platform thing.

17:33

SPEAKER_00

And, like, we've got a high confidence that, like, this stuff is not degrading environment or adding additional risk. In fact, I think it's removing risk because humans aren't actually as good as agents, like, when they're well-defined. Here's, like, skill invocation. I actually think the earlier numbers were a bit wonky. So, like, we hook up everything into Honeycomb. We've got hooks all over the place for basic information about, like, which skills are being invoked and things like that. And that's internally available. There's no private information in this. And everyone can kind of use

18:07

SPEAKER_00

it to kind of get an idea of, like, what's being used and where. But we also pull in all session transcripts into S3 for data mining, writing reports, like, also looking to see our skills effective, that kind of stuff. So, we've got, like, a feedback loop using the session data, which is, we can get more out of it, but we're doing some interesting stuff with it already. And this isn't a goal, but, like, and we're not particularly proud of, like, defects always increasing up until recently. But, like, defects are getting closed faster than ever. And, like, some teams have been inspired by the move to AI to think about things like backlog zero or crunching through

18:43

SPEAKER_00

hundreds or thousands of defects. So, like, some of this was, like, a bit deliberate and planned, but just in the same time, there's just, like, this natural deflation, because getting through this work, getting through all the defects so much faster these days. And, yeah, it's like, we're just seeing this naturally. We've also been working with, like, Stanford. There's a research group there. We've, um, we give them all our code. And our code quality per their metrics has been increasing over the last while. Um, okay. Uh, I'm kind of running out of time at this point. Uh, we have, like, hundreds of

19:17

SPEAKER_00

contributors, thousands, like, thousands and millions of code, of code, uh, in our cloud plugins. Um, it's very active. Uh, and, uh, yeah, I mean, Claude itself loves it. Um, here's an example skill. This is, like, not the most, or sorry, we've got, like, base plugins, things that, like, do all the session transcripts, session syncing, um, uh, some safety hooks and things. Um, and here's, like, a skill I built, which, like, it just, it fixes flaky specs. We have hundreds of thousands of tests, and, you know, they get flaky, uh, over time. And, uh, we don't, we ship a lot. So we just kind of barge through the

19:55

SPEAKER_00

kind of flakes. Um, but this, this skill was not built, like, by me kind of sitting down and figuring out, like, oh, what are all the things you need to do to fix flaky specs? I've worked in a feedback loop, gave, um, gave the agent a goal. And, uh, through, like, guiding us to the right place, uh, and working with us to fix a lot of flaky specs, uh, it's written this pretty decent thing. What are the these cheat codes or, like, lookup tables? And, uh, relatively well organized, using progressive disclosure and all that. Um, and, uh, it is, like, fixing stuff that if our most senior rails

20:30

SPEAKER_00

engineers were doing this, I'd be, like, wow, they're amazing. Um, and, yeah, like, a lot of other stuff going on. Like, our CI melted. We had to fix that. Um, cloud code is actually widely used across Intercom outside of software. It's gone completely viral. People are banging down our doors to, like, use console, use consoles. Um, and, uh, yeah, we're, you know, we're thinking a lot about, like, the future of engineering. Like, should we just merge all product manager design, everything? Um, oh, yes, the single person team product experiments have been pretty interesting as well. And I've even

21:00

SPEAKER_00

been shipping, like, codes, like, stuff that people can use in their agents to sign up to Intercom. Uh, this is stuff that, like, I, like, I've just been using our skills to act as a product manager, which is pretty wild. Um, so that's it. I wish you all the best of luck. If you're not doing pretty much all of this today, you're going to be doing it in the very near future. Um, my contact details are at brian.scannon.ie. You can interact with Finn in the messenger, configure by CLI. Um, and you can check out, uh, ideas.fin.ai for a lot more information about Intercom and our agents. Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note