SPEAKER_04
Can I get a vibe check of the room? How's everyone feeling? Are we really want to hear what Joe has to say about statues? Or are we just want to chill and sit in a quiet room? How are people feeling? Statues. Okay, that's good to hear. You're much more lively than statues, which I've spent a surprising amount of my life with recently. This, I'm going to, well, actually, I'll quickly introduce myself. I'm Joe. I work in the growth organization at Eleven Labs. Hands up if you've, actually, I've got a slide for that. Hands up if you've ever used or heard of Eleven Labs.
SPEAKER_04
Okay, so this next slide, you'll probably be familiar with a lot of it. Eleven Labs does, we're an audio AI foundation model company. So everything from text to speech, you put some text in, you get some speech out. Transcription, the other direction. Music, we've got the first commercially legal AI music generation. We licensed all the training data behind an API. Sound effects, create voices, this one, by the way, just as a pro tip, in case you're ever using Eleven Labs at a hackathon or something. This section, creating and editing voices,
SPEAKER_04
is the thing that people really just don't use Eleven Labs enough for. And this is the thing that the statue app that I built that we'll talk about is built on. And then agents, this is our fully managed agents deploy platform, which sounds mouthful of SaaS. And it is, but it's also very cool in various ways. So who saw the statue app? It went quite viral, at least here in London. Raise your hands if you saw it. Okay, that's fine, because I'm going to play you a video, because people like me love playing videos in my own voice out loud. I'll just play the first 30 seconds or so. I made an app that lets you talk to any statue you want using AI.
SPEAKER_04
So we've come to the British Museum to see how it works.
SPEAKER_04
I am Pharaoh Amenhotep. I am Demeter. Hoa Hakananaya. I am Han Sloan. The Guardian Lot. I am the young writer. And I am the war horse. First, take a picture. Okay, so now I'm about to explain that bit of me explaining. So what this statue does is it lets you take a picture of a statue, sorry, this app. It lets you take a picture of a statue. It then does an OpenAI deep research on the identity of the statue. Generates a bunch of the historical knowledge and prompts for what it thinks the voices of those individual statues would have been if they were alive.
SPEAKER_04
It then uses our voice design API, that really underutilized API where you can put in a description of a voice and it will go and generate something that matches. And then it creates an 11 labs agent and starts a phone call. And that whole thing works in 30 seconds.
SPEAKER_03
So you take a picture of something. You get all the search research back from OpenAI, generate a voice and start talking to an agent to a statue within 30 seconds.
SPEAKER_05
Which is pretty fun.
SPEAKER_07
If you're interested in reading more about the details of how it was all built, you can scan that QR code.
SPEAKER_05
It's just a blog post. This was attached to the initial tweet that I made, which basically had the prompt for one shotting.
SPEAKER_01
I built this whole thing in cursor in two hours.
SPEAKER_05
It's wild.
SPEAKER_04
So I built this in two hours on a Sunday because I was tired and bored.
SPEAKER_01
And then published the prompt through the 11 labs blog.
SPEAKER_04
Made this video that you just saw. Posted it on a Tuesday or something. And got 50,000 impressions. It was pretty good, not bad. People on Twitter kind of liked it. And then three museums, or people who represent groups of museums, and a bunch of other businesses, including TripAdvisor competitors and stuff, who were coming and saying, we've been, well, actually one of them, the CEO called me. He found my WhatsApp phone number somewhere. And called me up and said, I've had a team of 10 people working on this for a year. How did you build this? And so then the next day I reposted saying I've had a bunch of interesting people. I vibe coded this in two hours. Not as a brag.
SPEAKER_04
Just as, this is interesting. Vibe coding is so powerful. And if you are experimenting with these interaction patterns, you can actually do something quite big. Surprisingly big. It then went completely viral and went from 50,000 on the first day to one and a half million on the second day. And it was in part because it got kicked off by the vibe coding. And then suddenly I got all these artists and creatives and everything from portrait museums to Bonhams and Christies reaching out saying we want to have people be able to talk to the items we want to sell.
SPEAKER_04
I think what I would sort of, oh, and then this led to something called 11 hacks, which we can talk about later maybe. There are loads of different things in this story we can talk about. We can talk more about the statue app and about 11 labs. I want to get more from you. We can talk about 11 labs generally. We can talk about what it means to do growth, particularly API growth at one of these companies from the growth engineering point of view. And we can talk about the implications on culture, which I think is something that we as an industry are not really looking at as seriously as we could be. Vibe coding generally.
SPEAKER_04
And what that impact is happening on society, voice interaction patterns or making viral videos or anything else. So yeah, please. I mean, it's really easy to prototype these things, right? Yeah. Pushing into production, that's the hard part. So maybe, I'm curious about how you would take your own really successful prototype to how can you scale it to use it? Yeah, well, that's one of the things, this is going to sound very 11 labs salesy now. The nice thing is that pretty much all of what I've done is stitched together existing APIs, which are designed to scale. So yeah, if I wanted to start doing user management, that's relatively, I think, well understood.
SPEAKER_04
And you can buy from third parties for that. But the hard bit, maintaining the agents and the voice design, that's all APIs that there's no way I can make a dent in the API volume, even if this goes absolutely gangbusters. So from that point of view, I think, and I think that's something that vibe coding is really showing, is that the glue pieces and telling a good story about the glue is in part the most important thing of the project, rather than solving hard technical problems. The nice thing is that pretty much all of what I've done is stitched together existing APIs, which are designed to scale.
SPEAKER_04
So yeah, if I wanted to start doing user management, that's relatively, I think, well understood. And you can buy that from third parties. But the hard bit, maintaining the agents and the voice design, that's all APIs that there's no way I can make a dent in the API volume, even if this goes absolutely gangbusters. So from that point of view, I think, and I think that's something that Vibe coding is really showing, is that the glue pieces and telling a good story about the glue is in part the most important thing of the project, rather than solving hard technical problems.
SPEAKER_04
So yeah, obviously there's a lot more work to do to make it actually production ready, which is something that I've been talking to a bunch of the museums about. We might be doing that as a nice thing 11 Labs gives to the museums. But it's, the hard bit is actually not the user management and authentication. You can pretty much one shot that with Supabase or whatever else for logins and magic links. [SPEAKER_04] It's mostly relying on our APIs and our agents platform to do the heavy lifting. What about evals? [SPEAKER_04] Evals, I guess, so that's one of the big things.
SPEAKER_07
[SPEAKER_04] I think taking a photo and just getting your research back is not really the long-term solution for the museums.
SPEAKER_04
Really the important bit, the hard bit is going to the curators and saying, curators, can you figure out what is the actual narrative? Let's not just take random things you found from Google. Let's actually put some thought and design to the content. So that's the piece that's the longer tail. [SPEAKER_07] The nice thing is a lot of the museums have, they see their core IP as these databases. So they have APIs often and we can pull that information out. The VNA has a public API for their stuff. Yeah, I guess on this topic, maybe related to AI culture and voice interaction patterns, what is the interface that you enable the curator to design the experience?
SPEAKER_04
Yeah, so right now, there's no, I haven't designed anything for that. I mean, at best, they could log into a dashboard, the 11-Labs dashboard, and make edits to the system prompt and the knowledge base files that are in there. I think probably in this sort of information management, the best interaction pattern is probably editing text rather than speaking. Although, obviously, if you're choosing a voice, you need to manage that.
SPEAKER_04
There's an interesting question there. I've been talking to a chap, Jago, who used to be the head of the Americas at the British Museum and now runs the Sainsbury Centre, which amusingly is the location for the Avengers headquarters in the Avengers films. [SPEAKER_04] He is going through this big, long academic process of figuring out what should a voice sound like for an inanimate object. So things like, where did the materials originally come from? It came from some mountain in China or it came, and then the rock was shipped to Vietnam and then it was carved in Vietnam and then it spent the last 200 years living in a British museum. So, what should it sound like?
SPEAKER_04
It would have a little bit of a, maybe, Chinese origin with some Vietnamese twist in there, but then it's just lived around people with British accents, but maybe also not because it's lots of tourists. So, thinking through from a much more philosophical point of view, what should objects sound like? [SPEAKER_07] And that's something that, from my point of view, in 11 Labs is really interesting because I think we have the opportunity to give all sorts of things voices, like elevators. A lift should probably, I mean, they do have voices. They're often quite discordant with what a lift is, I find.
SPEAKER_04
But it's quite likely, I think, that we start walking into lifts and saying, I want to go to this floor, please, using voice to interact with them. So, what should they sound like?
SPEAKER_07
[SPEAKER_04] And it's becoming a more important question. [SPEAKER_04] I don't know exactly what the answer is, but smarter people are doing that.
SPEAKER_04
So, it's really interesting to hear about your thoughts about voice and interface to application. But, in general, what problems arise? What do they show? How do users interact? What do they expect from this voice engine? And when does it not work? And what problems do they expect if they are going to build an interface? I think, currently, voice interfaces and voice interactions have quite a large range of problems. But a lot of them are solvable. One of them is you basically have a binary. You're either interacting with voice or you're interacting in some other way.
SPEAKER_04
And I still feel like the sort of interactive or generative UI plus voice is something that we still haven't seen. [SPEAKER_00] And this is something I've experimented with. It's like if you think about a coding agent. You've got Lovable or something. I want to be able to talk to my app, talk to Lovable, but not talk to the coding agent part. I want to talk to a product manager agent that then goes off and triggers my coding agent to go and do things. So, the voice interaction there is not direct. The thing that I'm talking to, or the thing that's doing the work is the thing I'm talking to. I sort of want to be talking to a halfway house person.
SPEAKER_04
So, what are the problems? Aside from often the thing you end up talking to is not the thing you actually want to be talking to. There's also the parallel sort of interaction patterns in UI? And I think actually you were just showing me your app earlier which does show the stuff that the voice agent's thinking and sort of extracting from the conversation and allows you to interact with that at the same time. That's something that I think we're going to see a lot more of. The sort of multimodal conversations where it's voice and visual. The other thing is people don't interrupt voice agents because they're too polite. People are too polite.
SPEAKER_04
And I'm starting to learn to just interrupt agents much more aggressively. And that actually makes the experience much better. But I don't know how to give people permission to interrupt. How do you solve the problem of prompt guidance or skill learning? So, I mean with a typical coding agent you can give skills. This, this, this. That's something that I think we're going to see a lot more of. The multimodal conversations where it's voice and visual.
SPEAKER_04
The other thing is people don't like interrupting voice agents because they're too polite. People are too polite. And I'm starting to learn to just interrupt agents much more aggressively. And that actually makes the experience much better. But I don't know how to give people permission to interrupt. How do you solve the problem of prompt guidance or skill learning? So with a typical coding agent you can give skills. This, this, this. You can go and get it, right? But I don't think it's the same interface. Is it the same interface?
SPEAKER_04
So the 11 agents platform doesn't really support the concept of skills. Though it could. It does support the concept of knowledge files which do get loaded in. So you could do it in that way. We also support MCP calling. So you can have knowledge embedded in those or skills embedded in the MCPs. I don't think that's core to voice or not. I think that's mostly down to the interaction patterns of coding agents are quite lend themselves to skills. [SPEAKER_00] But you could have a voice agent that then has the ability to use skills. And you can add voice capabilities to an existing coding agent that's in that way. So I don't know if that's related directly to skills.
SPEAKER_04
[SPEAKER_00] I see. I see. [SPEAKER_02] Well, actually some people have built... Particularly with OpenClaw actually. There's Eleven Labs is quite a common interaction pattern for OpenClaw. People have built phone numbers they can call. And it'll call them back. And that sort of thing. And then obviously you can if you say, well, I want you to be able to load in skills, it just learns how to do it and adds that capability to itself. So that's for sure possible. And people are doing it with their sort of OpenClaw setups and some called code setups.
SPEAKER_00
[SPEAKER_04] In this experience, more on the AI culture and the bytecode. Thinking about new interaction patterns to engage with history. With our built environment. Where or how do you see that starting to... I mean, we're so empowered now with this bytecode. It's still a barrier. But how do we build that engagement and do it as well? I don't know what the museums are telling you. I'm sure it's so new.
SPEAKER_00
[SPEAKER_04] I think, to be honest, the museums... I've met with the CEO of the Science Museum, Co-CEO of the Science Museum. And they're asking the same questions. They don't really know the answers. They're saying, well... And the Science Museum, for example, is really good at going and trying stuff. So they've got all these big tablets that kids can go and interact with.
SPEAKER_00
[SPEAKER_02] But in my personal opinion, a lot of the time that can feel like sticking technology onto the thing rather than it being a core part. So some of the stuff we're experimenting with here is, obviously, you've got to take a picture of a statue and you talk to it. We're commissioning a statue to be made that has the technology inside of it and a speaker and phone and microphone so that you can then talk directly to the statue without having a piece of technology in the way,
SPEAKER_00
[SPEAKER_04] without it feeling tacked on. And that's something like with the red phone booth you may have seen here on floor three. There's a phone you can pick up inside of a K6 red London phone booth or a British phone booth and talk to an agent, talk to Sir Michael Caine. So that's trying to put it into the real world rather than having it go through a screen. [SPEAKER_04] I almost imagine if you would vibe code it. If you already see kids making games, imagine if you can go to the Science Museum, they have their tool and they just start creating whatever experience they want as they engage. It just explodes the possibility.
SPEAKER_00
[SPEAKER_04] Yeah, I mean, I think even more generally there's a question of what's vibe coding. I feel like still hasn't really gone consumer mainstream. It's even lovable feels like it's targeted at consumers for building, effectively building B2B SaaS apps. You know, it's having super base and standard design components. But this is why I really love vibe coding, vibe coding events. They feel like the OG hackathons because people show up and they've never even thought about writing code before.
SPEAKER_00
[SPEAKER_02] Sometimes I go to them and I talk to people and they're like, I say, what's your favorite app? And who made the app? And they're like, wait, people make apps? I thought they were just there on my phone or on the app store, right?
SPEAKER_00
[SPEAKER_04] They hadn't even thought through the fact that people have to make them. So that's something that I find vibe coding events really fun because people come in and they type in, they have no idea what a hamburger menu is or an accordion is. So they just say, I want this and I want this and I want this. And they get something completely wacky because the LLM just says, yeah, okay, I'll try it. When if they were talking to a software engineer, I would have said, ah, you want one of these and one of these.
SPEAKER_00
[SPEAKER_04] So I think at some point we're probably going to have a what's the Instagram filters moment for vibe coding or the TikTok moment for vibe coding? I think we're going to have something like that that makes social vibe coding much more. But I don't know what it's going to look like, but worth experimenting.
SPEAKER_04
Do you see anyone doing it well? because the LLM just says, yeah, okay, I'll try it. When, if they were talking to a software engineer, I would have said, ah, you want one of these and one of these. So, I think at some point, we're probably going to have what's the Instagram filters moment for vibe coding or the TikTok moment for vibe coding? I think we're going to have something like that that makes social vibe coding much more. But I don't know what it's going to look like, but worth experimenting. Do you see anyone doing it well? There's Spielwerk. There's an app, a mobile app for vibe coding games and it's TikTok swiping.
SPEAKER_04
And there's, I think, I can't remember what it is. There's a London-based game vibe coding tool that is, again, focused on games. I don't know that games are really the thing because they're quite complex. But there are a few people experimenting. But I don't think there are that many people really deeply pushing the boundaries of what's possible. They're mostly lovable, but on your mobile phone, is my opinion. Does anybody else have any opinion, see anyone doing good consumer vibe coding? I guess content creation. [SPEAKER_02] It's, if you buy the building or something, [SPEAKER_02] working. Yeah. You try interacting with digital systems or whatever it is.
SPEAKER_04
Well, in your game, you can spawn things in with voice, right? And that's a great,
SPEAKER_00
[SPEAKER_04] it's not quite vibe coding, [SPEAKER_04] but it's still interacting. [SPEAKER_04] I think that's what vibe coding still means you have [SPEAKER_04] an understanding of certain primitives. [SPEAKER_04] Even what you just said, [SPEAKER_04] you're thinking about data,
[SPEAKER_04] I think when this goes mainstream, [SPEAKER_04] people aren't thinking about this. [SPEAKER_03] And that's what's so exciting for me when you think about culture and what you've built is when it gets to a point where things that us as engineers wouldn't even approach the problem, and that's where some incredible creativity is. So the pattern that I think is closest to being a winner in this space is the Facebook Instant Games API, which doesn't exist anymore. [SPEAKER_02] Oh. [SPEAKER_02] They deprecated it, [SPEAKER_02] but it was in Facebook Messenger. [SPEAKER_02] You could play these games, [SPEAKER_02] and they had these primitives for social gaming.
SPEAKER_04
[SPEAKER_02] They tended to be quizzes or Fruit Ninja, [SPEAKER_02] and you compete with your group chats and things. [SPEAKER_02] I bought, for 15 pounds, [SPEAKER_02] a Fruit Ninja clone off a website, [SPEAKER_02] instrumented it with the Facebook Instant Games API, [SPEAKER_02] which was this beautiful [SPEAKER_02] JavaScript async await, [SPEAKER_02] had a get user information, [SPEAKER_02] get friends, create a leaderboard, or get your position on the leaderboard. It's very basic data storage, key value storage, and async await, show a rewarded ad, and show an interstitial ad. And so those things allowed you to make, very quickly and easily,
SPEAKER_02
[SPEAKER_04] a social graph-enabled, [SPEAKER_04] ads-enabled experience for consumers. [SPEAKER_04] So I bought this name, 15 pounds, [SPEAKER_04] instrumented it with Facebook Instant Games API, [SPEAKER_04] and went to bed. [SPEAKER_04] The next day, I woke up with 15 million users [SPEAKER_04] on this random game. [SPEAKER_04] I mean, I didn't make very much money, [SPEAKER_04] but there were 15 million users in Vietnam [SPEAKER_04] because, obviously, Facebook,
SPEAKER_04
they test everything out in the lower-value advertising regions and then rolls up. But that was amazing, because you've got the social elements of people instantly sharing it around, and that, I think, is probably the template that's going to, something along those lines is probably going to be the thing that wins on the social vibe coding. I don't know if that actually answers the question. On the kind of live working, live building thread, I kind of feel frustration whenever I get a response back in voice. Maybe it's just me, but the input, the information density per second isn't quite high enough. I use a lot of voice out to just get things out of my head.
SPEAKER_04
It's like, oh, fuck, and type, and whatever. Or voice input, I guess. Voice input. But I still feel like I need
SPEAKER_02
[SPEAKER_04] diagrams or text [SPEAKER_04] or something really high density [SPEAKER_04] back from the system. [SPEAKER_03] So I don't know whether, [SPEAKER_03] how do you guys think about that? [SPEAKER_03] Whether I'm actually curious
SPEAKER_04
[SPEAKER_03] if other people can use that [SPEAKER_03] whether they agree with me or something. [SPEAKER_03] It's not whether it's not. [SPEAKER_03] But that's what I find myself leaning towards, [SPEAKER_03] where it's information-rich input, [SPEAKER_03] and then I can just speak my thoughts [SPEAKER_03] as they come out, [SPEAKER_03] and it's almost semantically understood, [SPEAKER_03] put into that information-rich format, and then my intent is spawned out and across the system itself. [SPEAKER_03] So that's the pattern [SPEAKER_03] that I'm seeing and feeling, [SPEAKER_03] and what I always want to evolve [SPEAKER_03] and interact with now,
SPEAKER_04
[SPEAKER_03] my email, call, everything. [SPEAKER_03] I feel like that longing, [SPEAKER_03] it's not quite there yet, [SPEAKER_03] but I feel like I have this inclination [SPEAKER_03] to build in that direction [SPEAKER_03] and to try to get information-rich [SPEAKER_03] back, [SPEAKER_03] but also input feed very freeform, [SPEAKER_03] and just my broad intent. [SPEAKER_03] Yeah, that's something that I feel.
SPEAKER_04
So that's the pattern that I'm seeing and feeling, and what I always want to evolve and interact with now, my email, call, everything. I feel like that longing, it's not quite there yet, but I feel like I have this inclination to build in that direction and to try to get information-rich, get that back, but also input feed very free form, and just my broad intent. Yeah, that's something that I feel. I absolutely feel, I want to speak, speak easily and quickly, and then receive maybe a little bit of voice, but mostly this generated, maybe it's a UI, maybe it's just diagrams, maybe it's whatever app I'm in context of. But yeah, I want to have a parallel input, a parallel output of single input of my voice. I think the other part of that is I find that what I don't actually get that much information, necessarily, from voice, but what I do get is companionship. So it triggers that, it lessens the loneliness feel somehow. If I'm walking to something and I'm getting some information, I don't feel, I might be learning more if I'm looking at diagrams or text, but I don't feel as, I feel more motivated to continue tinkering. So there's some interesting modalities where you feel different things if you get the information coming in from all these different modalities, at least I find it. Curiously, yeah, how other people think about whether they think differently or whether that's something that they also feel as well. Yeah. Well, the visual cortex is much older than voice and text, so seeing something. Yeah. Also, when you ask it to be concise, I don't feel offended if it gives me a concise answer. But in speech, if it gives a concise answer, I'm relaxed, man, just ask you a question here. You don't have to ask it. You have to be concise and it just sounds rude. That's an interesting, maybe this is possible, maybe it's not, but what does skim listening look like? Yeah, exactly. Maybe listening actually should also have two buttons, forwards and backwards, and I can just tap, tap, tap, tap, tap, go forwards half a sentence until I, I don't know, maybe that's, maybe we should build that. Who wants to vibe code something with me straight up to this? How would it work? So you'd be, you'd just easily, yeah, back and forward going forward with the audience. It's like a speed dial, on a podcast, or on 2X. Yeah, or like the old iPods where you can sort of spin forwards and backwards. I don't know, that's probably quite a nice listening interaction. And you sort of, I guess, you want to scroll forwards in concepts, not necessarily in sentences, right? Like, what is the thing that your eyes look at when you're skim reading? It's probably, it's not the sentence structure, it's the words, I guess. It's the next thing. Yeah, yeah. Yeah, yeah. And that's actually, this is the story, this is very interesting and exciting for me. Because if I'm talking to an agent and it just starts rambling about, sometimes you get back three paragraphs of stuff and I'm, no, not this one. But I want to interrupt it and say, go to the next one. But then it's effectively saying, it's a new prompt, right? So then it's effectively saying, yes, okay, I'll focus on that next one and it'll write me three paragraphs about the second paragraph. You know, that's not really at all what I wanted. Unless I say, be concise and then it says something rude to me. Could you make a summary for each paragraph and then if you hold the, yeah, yeah. It will expand and then all you can. Yeah, I think the Claude app has done some interesting stuff on the voice interactions because they show you something different to you here and they show the higher level sections and then it goes into each one and you can tap on them. So that's, I guess, getting a little bit closer to this. But I think that also means you're not interacting as though you would with a human conversation. I don't know. I'm thinking about that. Like, why do we not have this issue when we're talking to humans? Right? Like, how do we,
SPEAKER_04
and they show the higher level sections and then it goes into each one and you can tap on them. [SPEAKER_01] So that's getting a little bit closer to this. [SPEAKER_01] But I think that also means you're not interacting [SPEAKER_01] as though you would with a human conversation. I don't know. I'm thinking about that. Why do we not have this issue when we're talking to humans? Right?
SPEAKER_04
How do we, what was there many other cues?
SPEAKER_02
[SPEAKER_04] Yeah. [SPEAKER_04] What's right there?
SPEAKER_04
There's a cue. We have a scent. There's so many other cues. I think there are visual cues as well. Yeah. I don't know. So that's something. [SPEAKER_02] If I'm as a businessman, if there's an agent response coming into audio, I don't know how long that's going to be. [SPEAKER_02] I also think is it going to be a minute or is it going to be [SPEAKER_03] 10 seconds? [SPEAKER_03] I almost want to know. [SPEAKER_03] And if I know it's really long and long, I have some other. [SPEAKER_03] But also you can tell when I'm about to interrupt you. [SPEAKER_03] So you go faster and you move, maybe you can, and you can tell if people are listening or not. [SPEAKER_03] Yeah.
SPEAKER_03
Yeah.
SPEAKER_04
[SPEAKER_03] I guess, oh, there's another thing which is interesting here, which is the interrupting. [SPEAKER_03] Sometimes I don't want to interrupt. [SPEAKER_03] I just want to say, yeah, yeah, [SPEAKER_04] yeah. [SPEAKER_04] Or, oh, but no, go back. You're always listening for the, [SPEAKER_06] uh-huh, yeah, yeah, yeah. But you can't do that with an agent. I think I'm willing to, I mentioned this to Joe before, but briefly, I showed you this yesterday, but briefly, you need to be able to interact with a PS5 game and also create what you want in that world as you're playing it.
SPEAKER_02
[SPEAKER_04] So you're on a shoot a battle, I'm playing you, we're trying to, you know, [SPEAKER_03] kill each other, you know, blow it, but if I'm trying to be creative and generate a new getaway vehicle or a helicopter, I can just save that experience and have that appear in the game straight away and then I can fly away. [SPEAKER_03] That's the work that I'm doing and how I met Joe actually. [SPEAKER_03] But I think the challenge that I face is that relying on just the voice interruptibility is quite unreliable. [SPEAKER_03] So I just got around to a simple whisper flow type. [SPEAKER_03] Push to talk. [SPEAKER_03] Push to talk.
SPEAKER_02
[SPEAKER_03] And then hold to talk and then let go to finish. [SPEAKER_03] Which augments this audio stream with some other cue. [SPEAKER_03] And in the same way that I think you almost need some very light nudge interface on top of the audio that you're receiving. [SPEAKER_03] And maybe as you say, you're saying something and then you see a little circle appearing being the agent wants to ask you a question. [SPEAKER_03] That would be an interesting experience in the field. [SPEAKER_03] If you're not being interrupted but you're being the agent wants to talk and then maybe you either stop and say, okay, what idea do you have?
SPEAKER_02
[SPEAKER_03] Because I can feel you doing that to me right now.
SPEAKER_04
[SPEAKER_03] You want to say something. [SPEAKER_03] I'm just getting loads of ideas. [SPEAKER_03] This is great. [SPEAKER_03] I'm just having an information communication on that level. [SPEAKER_03] But in audio, if I'm just listening to audio and not, I don't think I'd know that as a developer page or a developer page. [SPEAKER_03] Just listening to audio and not listening to the audio stream. [SPEAKER_03] I'll let you talk. Sorry. Well, I'm just imagining on that point that the agent wants to respond to you. It could be showing, [SPEAKER_03] I want to interrupt and tell you about this thing or this one.
SPEAKER_04
[SPEAKER_03] And then suddenly that becomes what we've just done but with even more context than doing it with a human because it's signaling [SPEAKER_03] or a developer page. [SPEAKER_03] Just listening to audio and not listening to the audio stream. [SPEAKER_03] I'll let you talk. Sorry. Well, I'm just imagining on that point that the agent wants to respond to you. It could be showing, [SPEAKER_03] I want to interrupt and tell you about this thing or this one. And then suddenly that becomes what we've just done but with even more context than doing it with a human because it's signaling the topic it wants to talk about. [SPEAKER_03] One question
SPEAKER_04
regarding the product because regarding this topic, you need some true calling in the background to be able to actually understand and how would that be handled? Would you be streaming the audio directly to the client or through a back channel and then getting some information to the system from?
SPEAKER_04
[SPEAKER_00] Or how can you work? Obviously, this hasn't been, as far as I'm aware, hasn't been built yet. The way I would probably approach it is looking at the transcript and just keep analyzing the transcript over and over again and say, do you have anything to add? Do you have anything to add? Do you have anything to add? Or what would you add? Rather than being a tool call or being part of it,
SPEAKER_04
I would do this as an asynchronous looking at the transcript. Are you changing original prompt? Because the original prompt has a plan and you want to do this and halfway you change the plan and you go in and how do you hear the original agent? Well, this is actually an interesting thing with agents. Often we see agents as things that you can't really interact with the internals of. But effectively, an agent is [SPEAKER_00] the logic that says what's the next message but it's also a transcript and that transcript of the conversation is completely malleable. So maybe some other thing
SPEAKER_04
we could experiment with is actually allowing the ability to do the interruptions where it's talking and then I can say yeah, yeah, yeah. And currently, if I have the way most agents platforms will work is it'll generate its full text thing and then start generating the audio and then if I talk even while it's partway through it's still reading out the audio, the text it will just append my message to the end of that full message. But we do have the timestamps. We know how far through the audio is played. So we could actually just go and edit the transcript that's coming back and say, well, no, they interrupted at this point so we're going to forget that the LLM even generated more text. Yeah, it's you have used the message and then you have used the message. Yeah. And you need a second type of use message that you are passing in. I don't know that's the of course it's a use message but it's also some other kind of doing it.
SPEAKER_04
[SPEAKER_00] Yeah. I mean, why don't you let's go to the 11 labs booth at the expo floor and just vibe code something and see. This is great. I can't wait. Now it's getting easier and easier to make tools and there's so many different ways to execute your ideas and it seems the main differentiate now getting your word out there. How did you approach making those videos and how long did it take you
SPEAKER_04
[SPEAKER_01] for the statue video to make it? That's a good question. Totally separate. So, this is the inspiration of this 11 hacks thing. I learned to make videos. I'm not particularly good at it. It's still relatively janky. I learned to make videos through doing other politics related campaigning stuff before I even knew 11 labs existed. It turns out that with videos random things go viral or with content random things go viral. [SPEAKER_01] of this 11 hacks thing. I learned to make videos. I'm not particularly good at it. It's still relatively janky. I learned to make videos through doing other politics related campaigning stuff before I even knew 11 labs existed.
SPEAKER_04
It turns out that with videos random things go viral or with content random things go viral and random things don't. Things I think are going to go viral don't and the things that do do. I think it's about practicing. Editing a video I find it's the 80-20 rule. Editing a video to the standard of the Statue app is actually quite relatively easy, but then to go beyond that—because I edited that on my phone—the editing itself took about 20 minutes, 25 minutes. Going beyond the quality of that suddenly means thinking about using a desktop editing tool, which suddenly makes everything even using CapCut on desktop to me is three times harder than using it on my phone and it's also three times more expensive—the subscription on the laptop than on the phone. So I think a lot of it is just doing stuff and trying and iterating.
SPEAKER_04
Big things that I found: adding captions really help to the video, having a hook in the first. You can look at your various platform analytics once you've posted a few videos. My videos tend to get between 6 and 12 seconds as the median view time and the people drop off. So you need to get your hook in there because most people are going to drop off if they don't buy the hook. So that's an important piece—front loading the interesting piece. Adding music makes a massive difference, and this is something that was much harder, but now with 11 Labs music generation I will make the video and then add. Sometimes I'll make the video with the narrative and everything and then just experiment with completely different genres until I find music and I'll just put the music on. Does it work? Yes, no. And then you can edit the different sections so that it times up with the sections of the speech so you don't need to figure out music first, find a piece of music and then match your speech to it. The other way, sometimes I will have a vibe I want to get across and I will generate the music first and then figure out what's my speech that matches that vibe. So it might be an excited theme or—I actually chose the music for the Statue app before I made the video itself because I thought this is a fun piece of music for a Statue type thing. It's an imperial outside the British Museum. It sort of made sense as an attention grabber. So but I think music is a massive thing that people underrate because you just put it—it's relatively quiet, but it makes a massive difference to the feeling.
SPEAKER_04
So you just did it on CapCut and did you think about it? It's literally on my mobile phone. I've got—I borrowed my wife's lapel mic, Bluetooth lapel mic, which cost 200 quid from DJI, which makes the audio much better and then yeah, just edited on CapCut. Super simple. So you didn't think about like this is the shot I want and then I want to take this photo? I mean the one out the front of the British Museum, yes, thank you. Yes, I did want that one, but the [SPEAKER_01] Much better. [SPEAKER_01] And then yeah, just edited on CapCut. Super simple.
SPEAKER_04
So you didn't think about this is the shot I want and then I want to take this photo? I mean, the one out the front of the British Museum? Yes, thank you. Yes, I did want that one, but the— [SPEAKER_01] Actually, the second time I recorded the video, I went once and got some stuff that was a bit boring, and then I went back and just took a bunch of city photos, and that's what ended up being the video. Cool. Thank you very much. Thank you. Thank you. Thank you. Thank you. Thank you.
[SPEAKER_04] Thank you. [SPEAKER_04] Thank you. [SPEAKER_04] Thank you. [SPEAKER_04] Thank you. [SPEAKER_04] Thank you. [SPEAKER_04] Thank you. [SPEAKER_04] Thank you. [SPEAKER_04] Thank you. and then I can just speak my thoughts as I come out, and it's almost semantically understood, put into that information-rich format, and then my intent is, like, you know, spawned out and across the system itself. So that's the pattern that I'm seeing and feeling, and what I always want to evolve and interact with now, my email, call, like, everything. I feel, I feel like that longing, it's not quite there yet, but I feel like I have this inclination to build in that direction
SPEAKER_03
and to try to get information-rich, you know, get that back, but also sort of, like, input feed very free form, and just my sort of broad intent. Yeah, that's something that I feel. I absolutely feel, you know, I want to speak, speak easily and quickly, and then receive maybe a little bit of voice, but mostly this, like, generated, maybe it's a UI, maybe it's just diagrams, maybe it's whatever app I'm in context of. But yeah, I want to have a parallel input, a parallel output of, like, single input of my voice. I think the other part of that is I find that the, what, I don't actually get that much information, necessarily, from voice,
SPEAKER_04
but what I do get is, like, companionship.
SPEAKER_03
So it kind of, like, triggers that, it lessens the loneliness feel somehow. If I'm, like, walking to something
SPEAKER_04
and I'm getting some information, I don't feel, I might be learning more if I'm looking at diagrams or text, but I don't feel as, I feel more motivated to kind of, like, continue tinkering. So there's, like, some interesting modalities where you feel different things if you get the information coming in
SPEAKER_03
from all these different modalities, at least I find it. Curiously, yeah, how other people think about whether they think differently or whether that's something that they also feel as well. Yeah. Well, the visual cortex is much older than voice and text, so seeing something. Yeah. Also, when you ask it to be concise, I don't feel offended if it gives me a concise answer. But in speech, if it gives a concise answer, so I'm relaxed, man, just ask you a question here. You don't have to ask it. You have to be concise and it just sounds rude. That's an interesting, maybe this is possible, maybe it's not, but what does skim listening look like? Yeah, exactly.
SPEAKER_03
Maybe, like, listening
SPEAKER_06
actually should also have two buttons,
SPEAKER_04
like forwards and backwards,
SPEAKER_06
and I can just tap, tap, tap, tap, tap, go forwards half a sentence until I, I don't know, maybe that's, maybe we should build that. Who wants to vibe code something with me straight up to this? How would it work? So you'd be, you'd just easily,
SPEAKER_03
yeah, back and forward
SPEAKER_04
going forward with the audience. It's like a speed dial, like on a podcast,
SPEAKER_06
or on 2X.
SPEAKER_04
Yeah, or like the old iPods where you can sort of spin forwards and backwards. I don't know, that's probably quite a nice listening interaction. And you sort of, I guess, you want to scroll forwards in concepts,
SPEAKER_03
not necessarily in sentences, right? Like, what is the thing that your eyes look at when you're skim reading? It's probably,
SPEAKER_02
it's not the sentence structure, it's like the words, I guess. It's like the next thing.
SPEAKER_04
Yeah, yeah. Yeah, yeah. And that's actually, oh yeah, this is the story, this is very interesting and exciting for me. Because if I'm talking to an agent and it just starts rambling about, you know, sometimes you get back three paragraphs of stuff and I'm like, no, not this one. But I want to interrupt it and say, go to the next one.
SPEAKER_06
But then it's effectively saying,
SPEAKER_00
it's a new prompt, right?
SPEAKER_04
So then it's effectively saying, yes, okay, I'll focus on that next one and it'll write me three paragraphs about the second paragraph. You know, that's not really at all what I wanted. Unless I say, be concise and then it says something rude to me. Could you like, make a summary for each paragraph and then like, if you hold the, Yeah, yeah. It will expand and then all you can like. Yeah, I think the Claude app has done some interesting stuff on the voice interactions because they, they show you something different to you here and they show the higher level sections and then it goes into each one and you can tap on them.
SPEAKER_01
So that's, I guess, getting a little bit closer to this. But I think that also
SPEAKER_04
means you're not interacting
as though you would with a human conversation. I don't know. I'm thinking about that. Like, why do we not have this issue when we're talking to humans? Right? Like, how do we, what was there many other cues? Yeah. What's right there? There's like a cue. We have a scent. There's so many other cues. I think there are visual cues as well. Yeah. I don't know. So that's something. It's like,
SPEAKER_02
if I'm, as a businessman for, like if there's an agent response coming into audio, I don't know how long that's going to be. I also think like, is it going to be like, you know, a minute or is it going to be
SPEAKER_03
10 seconds? I almost want to know. And if I know it's like really long and long, I have some other. But also you can tell when I'm about to interrupt you. So you go faster and you like move, maybe you can, and you can tell if people are listening or not.
SPEAKER_03
Yeah. Yeah. I guess, oh, there's another thing which is interesting here, which is the interrupting. Sometimes I don't want to interrupt. I just want to say, yeah, yeah,
SPEAKER_04
yeah. Or, oh, but no, go back. You're always listening for the,
SPEAKER_06
uh-huh,
SPEAKER_04
yeah, yeah, yeah. But you can't do that with an agent. I think, I'm willing to, you know, I mentioned this to Joe before, but briefly, I showed you this yesterday, but briefly, you need to be able to interact with a PS5 game and also create what you want in that world as you're playing it. So you're on like a, you know, shoot a battle, I'm playing you, we're trying to like, you know,
SPEAKER_03
kill each other, you know, blow it, but like, if I'm, like trying to be creative and, you know, generate a new getaway vehicle or a helicopter, I can just sort of save that experience and have that appear in the game straight away and then I can sort of fly away. That's the work that I'm doing and how I met Joe actually. But I think the, the challenge that I face is that relying on just the, like, voice interruptibility is quite unreliable. So I just got around to just a simple whisper flow type. Push to talk. Push to talk. And then hold to talk and then let go to finish. Which it augments, like, purely this audio stream with some other cue. And in the same way
SPEAKER_03
that I think you almost need some very light nudge interface on top of the audio that you're receiving. And maybe as you say, like, you know,
SPEAKER_04
you're saying something
SPEAKER_03
and then you see a little, sort of like circle appearing being like the agent wants to ask you a question. That would be an interesting experience in the field. If you're not being interrupted but you're being kind of like, you know, the agent wants to talk and then maybe you either stop and say, okay, what idea do you have? Because I even can feel you doing that to me right now. You want to say something. I'm just getting loads of ideas. This is great. I'm just sort of like, we're almost like having an information, you know, communication on that level. But in audio, if I'm just listening to audio and not, I don't think I'd know that as a developer page
SPEAKER_03
or a developer page. Like, just listening to audio and not listening to the audio stream. I'll let you talk.
Sorry. Well, I'm just imagining on that point that the agent wants to respond to you. It could be showing, I want to interrupt and tell you about this thing or this one. And then suddenly that becomes what we've just done but with even more context than doing it with a human because it's signaling the topic it wants to talk about. One question
SPEAKER_04
regarding the product because regarding this topic, you need some sort of true calling in the background to be able to actually understand and how would that to be handled? Would you sort of be streaming the audio directly to the client or through a back channel and then getting some information to the system from?
SPEAKER_00
Or how can you work? Obviously, this hasn't been, as far as I'm aware, hasn't been built yet. The way I would probably approach it is looking at the transcript and just keep analyzing the transcript over and over again and say, do you have anything to add? Do you have anything to add? Do you have anything to add? Or what would you add? Rather than being a tool call or being part of it,
SPEAKER_04
I would do this as an asynchronous looking at the transcript. Are you changing original prompt? Because the original prompt has a plan and you want to do this and halfway you sort of change the plan and you go in and sort of how do you hear the original agent? Well, this is actually an interesting thing with agents. Often we see agents
as things that you can't really interact with the internals of. But effectively, an agent is the sort of logic that says what's the next message but it's also a transcript and that transcript of the conversation is completely malleable. So maybe some other thing we could experiment with is actually allowing the ability to do the interruptions where it's talking and then I can say uh, uh, uh, yeah, yeah, yeah. And currently, if I have the the way most agents platforms will work is it'll generate its full text thing and then start generating the audio and then if I talk even while it's partway through it's still reading out the audio uh, the text it will just append
SPEAKER_04
my message to the end of that full message. But we do have the timestamps. We know how far through the audio is played. So we could actually just go and edit the transcript that's coming back and say, well, no, they interrupted at this point so we're going to forget that the LLM even generated more text. Yeah, it's sort of you have used the message and then you have used the message. Yeah. And sort of you need a sort of a second type of use message that you are passing in. I don't know that's the of course it's a use message but it's also some other kind
SPEAKER_00
of doing it. Yeah. I mean, why don't you let's go to the 11 labs booth at the expo floor and just vibe code something and see. This is great. I can't wait. Now it's getting easier and easier to make tools and there's so many different ways to execute
SPEAKER_04
your ideas and it seems like the main kind of differentiate now getting your word out there. How did you approach making those videos and how long did it take you
SPEAKER_01
for the statue video to make it? That's a good question. Totally separate. So, this is sort of the inspiration of this 11 hacks thing. I learned to make videos. I'm not particularly good at it. It's still relatively janky. I learned to make videos through doing other politics related campaigning stuff before I
SPEAKER_04
even knew 11 labs existed. It turns out that with videos random things go viral or with content random things go viral and random things don't. Things I think are going to go viral don't and the things that do do. I think it's basically about practicing. Editing a video I find it's like the 80-20 rule editing a video to the standard of the statue app is actually quite relatively easy but then to go beyond that because I edited that on my phone the editing itself took about 20 minutes 25 minutes. Going beyond the quality of that suddenly means thinking about using a desktop editing tool which suddenly makes everything even using CapCut on desktop to me is like three times
SPEAKER_04
harder than using it on my phone and it's also three times more expensive the subscription on the laptop than on the phone. So I think a lot of it is just doing stuff and trying and iterating. Big things that I found adding captions really help to the video having a hook in the first you can look at your various platform analytics once you've posted a few videos my videos tend to get between 6 and 12 seconds is the median view time and the people drop off. So you need to get your hook in there because most people are going to drop off if they don't buy the hook. So that's an important piece front loading the interesting piece. Adding music makes a massive difference
SPEAKER_04
and this is something that was much harder but now with 11 Labs music generation I will make the video and then add sometimes I'll make the video with the narrative and everything and then just experiment with completely different genres until I find music and I'll just put the music on does it work yes no and then you can edit the different sections so that it times up with the sections of the speech so you don't need to figure out music first find a piece of music and then match your speech to it. The other way sometimes I will have a vibe I want to get across and I will generate the music first and then figure out what's my speech that matches that vibe so it might be
SPEAKER_04
like an excited theme or a I actually chose the music for the statue app before I before I made the video itself because I thought this is a fun piece of music for a statue type thing it's like a bit of an imperial outside the British Museum it sort of made sense as an attention grabber so but I think music is a massive thing that people underrate because you just put it it's relatively quiet but it makes a massive difference to the feeling. So you just did it on CapCut and did you think about it? It's literally on my mobile phone I've got I borrowed my wife's lapel mic Bluetooth lapel mic which cost 200 quid from DJI which makes
SPEAKER_01
the audio much better and then yeah just edited on
SPEAKER_04
CapCut super simple So you didn't think about like this is the shot I want and then like I want to take this photo I mean the the one out the front of the British Museum yes thank you yes I did want that one but the and it was
SPEAKER_01
actually the second time I recorded the video I went once and got some stuff that was like a bit
SPEAKER_04
boring and then I went back and just took a bunch of city photos and that's what ended up being the video cool thank you very much thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you you