How to talk to statues — Joe Reeve, ElevenLabs
Description
A museum CEO tracked down his WhatsApp number and called. "I've had a team of 10 people working on this for a year. How did you build this?" Joe from ElevenLabs built the statue app in two hours on a Sunday using Cursor and a single one shot prompt. He posted it on a Tuesday and got 50,000 impressions. Reposted the next day about vibe coding and hit 1.5 million. The pipeline: point your phone at a statue, OpenAI deep research identifies it and generates historical context and a voice description, the ElevenLabs voice design API creates a matching voice from that description, an agent spins up, and a conversation starts. The whole thing runs in about 30 seconds. Museums, auction houses, and travel platforms all reached out wanting the same thing built for their collections. Speaker info: - https://x.com/isnit0 - https://x.com/isnit0/status/2024104717039685915 - https://elevenlabs.io/blog/talk-to-a-statue-building-a-multi-modal-elevenagents-powered-app
Summary
Generated by claude-sonnet-4-530-second take
Joe Reeve from ElevenLabs built a viral statue-talking app in two hours using cursor, OpenAI deep research, and ElevenLabs' voice design + conversational agents APIs. The app takes a photo of a statue, researches its identity, generates a historically appropriate voice, and launches a voice agent—all in 30 seconds. The first tweet got 50k views; when he reposted emphasizing vibe coding's speed, it hit 1.5M views and led to inbound from museums (Science Museum CEO, Bonhams, Christies) and competitors. The real story is less about the tech (stitching existing APIs) and more about vibe coding as a cultural shift: non-engineers can prototype experiences that feel production-ready, and marketing the glue/narrative matters more than solving hard problems. Joe sees voice interfaces as companionship-first, not information density, and believes the big unlock is multimodal UX—voice input + dense visual output—plus social mechanics like Facebook Instant Games had.
Key takes
- Vibe coding collapses time-to-prototype so radically that it inverts competitive dynamics: Joe's two-hour build matched a 10-person team's year of work. Museums and auction houses want this because the API glue is already production-scale; the bottleneck is now storytelling/distribution, not engineering.
- Voice design is underutilized and non-obvious: Most ElevenLabs users don't touch voice creation/editing APIs; Joe's statue app relies entirely on programmatically generating character-specific voices from text prompts (e.g., a Chinese statue carved in Vietnam, living in a British museum gets a blended accent).
- Voice interfaces provide companionship more than information density: Joe finds voice reduces loneliness and increases motivation to tinker, even though visual/text modalities are higher bandwidth. The implication is voice's value is affective, not informational—people want to speak input fast but receive dense visual/UI output.
- Current voice agent UX is broken by conversational norms: People are too polite to interrupt agents, and concise speech sounds rude. Joe wants "skim listening" (tap to skip forward by concept, not sentence), visual interruption cues (showing agent wants to add something about X topic), and transcript editing so interruptions truncate agent responses instead of appending messages.
- Vibe coding events attract true non-engineers: Joe sees people who didn't know apps are made by people show up and describe UIs without knowing terms like "hamburger menu," producing wacky results because LLMs just try it instead of translating to standard patterns. The next unlock is a "TikTok moment" for consumer vibe coding.
- Virality came from the reframe, not the tech: The first post (50k views) demonstrated the product. The second (1.5M views) said "I vibe coded this in two hours"—shifting narrative from demo to cultural statement. Artists, museums, auction houses then contacted Joe because the story gave them permission to think differently about tech access.
Useful details
- Statue app stack: Photo → OpenAI deep research for statue identity → ElevenLabs voice design API (text prompt to voice) → ElevenLabs conversational agents platform → phone call. Full pipeline is 30 seconds.
- Video production: Shot on iPhone, edited in CapCut mobile (20-25 min), used $200 DJI Bluetooth lapel mic, added captions, front-loaded hook (people drop off at 6-12 seconds median). Music generated in ElevenLabs after video narrative was done, matched to speech sections. Went to British Museum twice to get better b-roll.
- Museum interest specifics: Science Museum co-CEO reached out; one CEO found Joe's WhatsApp and called him directly; portrait museums, Bonhams, Christies, TripAdvisor competitors all contacted. Museums see their databases as core IP and often have APIs (e.g., V&A has a public API).
- Voice philosophy debate: Joe is working with Jago (ex-head of Americas, British Museum, now runs Sainsbury Centre) on academic questions like: Should a statue carved from Chinese stone in Vietnam but displayed in the UK for 200 years have a blended accent? What should elevators sound like?
- Facebook Instant Games template: Joe bought a Fruit Ninja clone for £15, instrumented it with Facebook Instant Games API (async/await, get user info, leaderboards, key-value storage, ads), went to bed, woke up with 15M users in Vietnam. He sees this as the template for social vibe coding: low-code social primitives + viral distribution.
- Interruption tech solution: Most agent platforms append new messages even if the user interrupts mid-speech. Joe proposes editing the transcript at the timestamp of interruption so the LLM "forgets" it generated the rest. ElevenLabs agents platform supports MCP and knowledge files but not explicit "skills" abstraction.
- 11 Hacks: Internal ElevenLabs initiative inspired by the statue app going viral.
Caveats / counterpoints
- Production is handwaved: Joe admits scaling the statue app to production requires user management, authentication, and curator-designed narratives (not auto-generated research). Museums need editorial control over content, not just Google search summaries. He says "the hard bit" is curatorial design, but hasn't built that; it's unclear how long that takes or if it's vibe-codeable.
- Voice companionship claim is anecdotal: Joe's take that voice reduces loneliness is introspective, not validated. No data on whether this generalizes or matters for users beyond personal motivation to keep using a tool.
- Viral narrative may be survivorship bias: Joe says "random things go viral" and admits he can't predict what works. The statue app's virality may be luck + timing + London cultural context (British Museum is iconic) rather than a replicable formula.
- Vibe coding still requires primitives literacy: Even Joe's "consumer" framing admits people at vibe coding events don't know what they want because they lack basic UI vocabulary. The Instagram/TikTok moment hasn't arrived, so it's still early adopter territory.
- Multimodal UX ideas are unbuilt: All the skim listening, interruption cues, visual topic signaling, and transcript editing are speculative. Joe proposes them as group brainstorming but admits they're experiments, not proven patterns.
Ken relevance
High. This maps directly to Ken's agent systems and content/GTM work:
- Agent UX design: The multimodal voice + dense visual output pattern (speak input, get structured UI/diagrams back) is exactly the interface Ken's building for agent workflows. Joe's observation that voice is companionship-first is a design insight: use voice to maintain engagement/motivation, not to replace high-bandwidth data.
- Vibe coding as distribution: Joe's story proves that speed-to-prototype + narrative = inbound from real businesses. If Ken can package agent demos as "I built this in X hours" content (especially if they solve hairy enterprise problems), it becomes a sales tool and cultural signal. The statue app's virality came from reframing effort, not from being technically impressive.
- API glue is the product: Joe's whole business model relies on ElevenLabs APIs being production-scale out of the box. This is a forcing function for Ken's agent platform: if third parties can compose it into viral products without denting server capacity, it proves the APIs are defensible infrastructure.
- Voice design as underutilized lever: Ken should flag voice creation/editing capabilities in any agent offering that uses voice. Joe's repeated note that nobody uses this API suggests it's a discovery problem, not a product problem—similar to how many LLM features go unused until someone demos them.
- Curator-as-bottleneck insight: Museums want AI but need human editorial control over narratives. This is a clear B2B wedge: sell agent platforms with curator/admin UIs for designing character prompts, knowledge bases, and conversational guardrails. Joe hasn't built this yet, so it's still an open opportunity.
- Investing angle: Joe's claim that a 10-person team spent a year on what he built in two hours is a red flag for incumbents and a green flag for infra plays (ElevenLabs, OpenAI, cursor). The value is shifting from app layer to foundation model + tooling layer.
Watch verdict
Watch fully. The first 15 minutes are essential for understanding vibe coding's cultural impact + multimodal voice UX patterns. The Q&A is dense with tactical ideas (skim listening, interruption UX, transcript editing, push-to-talk, Facebook Instant Games template) that apply directly to Ken's agent product design and GTM. The video production tactics (20 min edit, front-load hook, music after narrative) are a bonus playbook for Ken's own content strategy. Skip only if you've already internalized that API glue + narrative > hard tech, but the specific examples here (museum inbound, CEO cold-calls, virality mechanics) are sticky and worth absorbing.
Transcript
Can I get a vibe check of the room? How's everyone feeling? Are we really want to hear what Joe has to say about statues? Or are we just want to chill and sit in a quiet room? How are people feeling? Statues. Okay, that's good to hear. You're much more lively than statues, which I've spent a surprising amount of my life with recently. This, I'm going to, well, actually, I'll quickly introduce myself. I'm Joe. I work in the growth organization at Eleven Labs. Hands up if you've, actually, I've got a slide for that. Hands up if you've ever used or heard of Eleven Labs. Okay, so this next slide, you'll probably be familiar with a lot of it. Eleven Labs does, we're an audio AI foundation model company. So everything from text to speech, you put some text in, you get some speech out. Transcription, the other direction. Music, we've got the first commercially legal AI music generation. We licensed all the training data behind an API. Sound effects, create voices, this one, by the way, just as a pro tip, in case you're ever using Eleven Labs at a hackathon or something. This section, creating and editing voices, is the thing that people really just don't use Eleven Labs enough for. And this is the thing that the statue app that I built that we'll talk about is built on. And then agents, this is our fully managed agents deploy platform, which sounds mouthful of SaaS. And it is, but it's also very cool in various ways. So who saw the statue app? It went quite viral, at least here in London. Raise your hands if you saw it. Okay, that's fine, because I'm going to play you a video, because people like me love playing videos in my own voice out loud. I'll just play the first 30 seconds or so. I made an app that lets you talk to any statue you want using AI. So we've come to the British Museum to see how it works. I am Pharaoh Amenhotep. I am Demeter. Hoa Hakananaya. I am Han Sloan. The Guardian Lot. I am the young writer. And I am the war horse. First, take a picture. Okay, so now I'm about to explain that bit of me explaining. So what this statue does is it lets you take a picture of a statue, sorry, this app. It lets you take a picture of a statue. It then does an OpenAI deep research on the identity of the statue. Generates a bunch of the historical knowledge and prompts for what it thinks the voices of those individual statues would have been if they were alive. It then uses our voice design API, that really underutilized API where you can put in a description of a voice and it will go and generate something that matches. And then it creates an 11 labs agent and starts a phone call. And that whole thing works in 30 seconds. So you take a picture of something. You get all the search research back from OpenAI, generate a voice and start talking to an agent to a statue within 30 seconds. Which is pretty fun. If you're interested in reading more about the details of how it was all built, you can scan that QR code. It's just a blog post. This was attached to the initial tweet that I made, which basically had the prompt for one shotting. I built this whole thing in cursor in two hours. It's wild. So I built this in two hours on a Sunday because I was tired and bored. And then published the prompt through the 11 labs blog. Made this video that you just saw. Posted it on a Tuesday or something. And got 50,000 impressions. It was pretty good, not bad. People on Twitter kind of liked it. And then three museums, or people who represent groups of museums, and a bunch of other businesses, including TripAdvisor competitors and stuff, who were coming and saying, we've been, well, actually one of them, the CEO called me. He found my WhatsApp phone number somewhere. And called me up and said, I've had a team of 10 people working on this for a year. How did you build this? And so then the next day I reposted saying I've had a bunch of interesting people. I vibe coded this in two hours. Not as a brag. Just as, this is interesting. Vibe coding is so powerful. And if you are experimenting with these interaction patterns, you can actually do something quite big. Surprisingly big. It then went completely viral and went from 50,000 on the first day to one and a half million on the second day. And it was in part because it got kicked off by the vibe coding. And then suddenly I got all these artists and creatives and everything from portrait museums to Bonhams and Christies reaching out saying we want to have people be able to talk to the items we want to sell. I think what I would sort of, oh, and then this led to something called 11 hacks, which we can talk about later maybe. There are loads of different things in this story we can talk about. We can talk more about the statue app and about 11 labs. I want to get more from you. We can talk about 11 labs generally. We can talk about what it means to do growth, particularly API growth at one of these companies from the growth engineering point of view. And we can talk about the implications on culture, which I think is something that we as an industry are not really looking at as seriously as we could be. Vibe coding generally. And what that impact is happening on society, voice interaction patterns or making viral videos or anything else. So yeah, please. I mean, it's really easy to prototype these things, right? Yeah. Pushing into production, that's the hard part. So maybe, I'm curious about how you would take your own really successful prototype to how can you scale it to use it? Yeah, well, that's one of the things, this is going to sound very 11 labs salesy now. The nice thing is that pretty much all of what I've done is stitched together existing APIs, which are designed to scale. So yeah, if I wanted to start doing user management, that's relatively, I think, well understood. And you can buy from third parties for that. But the hard bit, maintaining the agents and the voice design, that's all APIs that there's no way I can make a dent in the API volume, even if this goes absolutely gangbusters. So from that point of view, I think, and I think that's something that vibe coding is really showing, is that the glue pieces and telling a good story about the glue is in part the most important thing of the project, rather than solving hard technical problems. The nice thing is that pretty much all of what I've done is stitched together existing APIs, which are designed to scale. So yeah, if I wanted to start doing user management, that's relatively, I think, well understood. And you can buy that from third parties. But the hard bit, maintaining the agents and the voice design, that's all APIs that there's no way I can make a dent in the API volume, even if this goes absolutely gangbusters. So from that point of view, I think, and I think that's something that Vibe coding is really showing, is that the glue pieces and telling a good story about the glue is in part the most important thing of the project, rather than solving hard technical problems. So yeah, obviously there's a lot more work to do to make it actually production ready, which is something that I've been talking to a bunch of the museums about. We might be doing that as a nice thing 11 Labs gives to the museums. But it's, the hard bit is actually not the user management and authentication. You can pretty much one shot that with Supabase or whatever else for logins and magic links. [SPEAKER_04] It's mostly relying on our APIs and our agents platform to do the heavy lifting. What about evals? [SPEAKER_04] Evals, I guess, so that's one of the big things. [SPEAKER_04] I think taking a photo and just getting your research back is not really the long-term solution for the museums. Really the important bit, the hard bit is going to the curators and saying, curators, can you figure out what is the actual narrative? Let's not just take random things you found from Google. Let's actually put some thought and design to the content. So that's the piece that's the longer tail. [SPEAKER_07] The nice thing is a lot of the museums have, they see their core IP as these databases. So they have APIs often and we can pull that information out. The VNA has a public API for their stuff. Yeah, I guess on this topic, maybe related to AI culture and voice interaction patterns, what is the interface that you enable the curator to design the experience? Yeah, so right now, there's no, I haven't designed anything for that. I mean, at best, they could log into a dashboard, the 11-Labs dashboard, and make edits to the system prompt and the knowledge base files that are in there. I think probably in this sort of information management, the best interaction pattern is probably editing text rather than speaking. Although, obviously, if you're choosing a voice, you need to manage that. There's an interesting question there. I've been talking to a chap, Jago, who used to be the head of the Americas at the British Museum and now runs the Sainsbury Centre, which amusingly is the location for the Avengers headquarters in the Avengers films. [SPEAKER_04] He is going through this big, long academic process of figuring out what should a voice sound like for an inanimate object. So things like, where did the materials originally come from? It came from some mountain in China or it came, and then the rock was shipped to Vietnam and then it was carved in Vietnam and then it spent the last 200 years living in a British museum. So, what should it sound like? It would have a little bit of a, maybe, Chinese origin with some Vietnamese twist in there, but then it's just lived around people with British accents, but maybe also not because it's lots of tourists. So, thinking through from a much more philosophical point of view, what should objects sound like? [SPEAKER_07] And that's something that, from my point of view, in 11 Labs is really interesting because I think we have the opportunity to give all sorts of things voices, like elevators. A lift should probably, I mean, they do have voices. They're often quite discordant with what a lift is, I find. But it's quite likely, I think, that we start walking into lifts and saying, I want to go to this floor, please, using voice to interact with them. So, what should they sound like? [SPEAKER_04] And it's becoming a more important question. [SPEAKER_04] I don't know exactly what the answer is, but smarter people are doing that. So, it's really interesting to hear about your thoughts about voice and interface to application. But, in general, what problems arise? What do they show? How do users interact? What do they expect from this voice engine? And when does it not work? And what problems do they expect if they are going to build an interface? I think, currently, voice interfaces and voice interactions have quite a large range of problems. But a lot of them are solvable. One of them is you basically have a binary. You're either interacting with voice or you're interacting in some other way. And I still feel like the sort of interactive or generative UI plus voice is something that we still haven't seen. [SPEAKER_00] And this is something I've experimented with. It's like if you think about a coding agent. You've got Lovable or something. I want to be able to talk to my app, talk to Lovable, but not talk to the coding agent part. I want to talk to a product manager agent that then goes off and triggers my coding agent to go and do things. So, the voice interaction there is not direct. The thing that I'm talking to, or the thing that's doing the work is the thing I'm talking to. I sort of want to be talking to a halfway house person. So, what are the problems? Aside from often the thing you end up talking to is not the thing you actually want to be talking to. There's also the parallel sort of interaction patterns in UI? And I think actually you were just showing me your app earlier which does show the stuff that the voice agent's thinking and sort of extracting from the conversation and allows you to interact with that at the same time. That's something that I think we're going to see a lot more of. The sort of multimodal conversations where it's voice and visual. The other thing is people don't interrupt voice agents because they're too polite. People are too polite. And I'm starting to learn to just interrupt agents much more aggressively. And that actually makes the experience much better. But I don't know how to give people permission to interrupt. How do you solve the problem of prompt guidance or skill learning? So, I mean with a typical coding agent you can give skills. This, this, this. That's something that I think we're going to see a lot more of. The multimodal conversations where it's voice and visual. The other thing is people don't like interrupting voice agents because they're too polite. People are too polite. And I'm starting to learn to just interrupt agents much more aggressively. And that actually makes the experience much better. But I don't know how to give people permission to interrupt. How do you solve the problem of prompt guidance or skill learning? So with a typical coding agent you can give skills. This, this, this. You can go and get it, right? But I don't think it's the same interface. Is it the same interface? So the 11 agents platform doesn't really support the concept of skills. Though it could. It does support the concept of knowledge files which do get loaded in. So you could do it in that way. We also support MCP calling. So you can have knowledge embedded in those or skills embedded in the MCPs. I don't think that's core to voice or not. I think that's mostly down to the interaction patterns of coding agents are quite lend themselves to skills. [SPEAKER_00] But you could have a voice agent that then has the ability to use skills. And you can add voice capabilities to an existing coding agent that's in that way. So I don't know if that's related directly to skills. [SPEAKER_00] I see. I see. [SPEAKER_02] Well, actually some people have built... Particularly with OpenClaw actually. There's Eleven Labs is quite a common interaction pattern for OpenClaw. People have built phone numbers they can call. And it'll call them back. And that sort of thing. And then obviously you can if you say, well, I want you to be able to load in skills, it just learns how to do it and adds that capability to itself. So that's for sure possible. And people are doing it with their sort of OpenClaw setups and some called code setups. [SPEAKER_04] In this experience, more on the AI culture and the bytecode. Thinking about new interaction patterns to engage with history. With our built environment. Where or how do you see that starting to... I mean, we're so empowered now with this bytecode. It's still a barrier. But how do we build that engagement and do it as well? I don't know what the museums are telling you. I'm sure it's so new. [SPEAKER_04] I think, to be honest, the museums... I've met with the CEO of the Science Museum, Co-CEO of the Science Museum. And they're asking the same questions. They don't really know the answers. They're saying, well... And the Science Museum, for example, is really good at going and trying stuff. So they've got all these big tablets that kids can go and interact with. [SPEAKER_02] But in my personal opinion, a lot of the time that can feel like sticking technology onto the thing rather than it being a core part. So some of the stuff we're experimenting with here is, obviously, you've got to take a picture of a statue and you talk to it. We're commissioning a statue to be made that has the technology inside of it and a speaker and phone and microphone so that you can then talk directly to the statue without having a piece of technology in the way, [SPEAKER_04] without it feeling tacked on. And that's something like with the red phone booth you may have seen here on floor three. There's a phone you can pick up inside of a K6 red London phone booth or a British phone booth and talk to an agent, talk to Sir Michael Caine. So that's trying to put it into the real world rather than having it go through a screen. [SPEAKER_04] I almost imagine if you would vibe code it. If you already see kids making games, imagine if you can go to the Science Museum, they have their tool and they just start creating whatever experience they want as they engage. It just explodes the possibility. [SPEAKER_04] Yeah, I mean, I think even more generally there's a question of what's vibe coding. I feel like still hasn't really gone consumer mainstream. It's even lovable feels like it's targeted at consumers for building, effectively building B2B SaaS apps. You know, it's having super base and standard design components. But this is why I really love vibe coding, vibe coding events. They feel like the OG hackathons because people show up and they've never even thought about writing code before. [SPEAKER_02] Sometimes I go to them and I talk to people and they're like, I say, what's your favorite app? And who made the app? And they're like, wait, people make apps? I thought they were just there on my phone or on the app store, right? [SPEAKER_04] They hadn't even thought through the fact that people have to make them. So that's something that I find vibe coding events really fun because people come in and they type in, they have no idea what a hamburger menu is or an accordion is. So they just say, I want this and I want this and I want this. And they get something completely wacky because the LLM just says, yeah, okay, I'll try it. When if they were talking to a software engineer, I would have said, ah, you want one of these and one of these. [SPEAKER_04] So I think at some point we're probably going to have a what's the Instagram filters moment for vibe coding or the TikTok moment for vibe coding? I think we're going to have something like that that makes social vibe coding much more. But I don't know what it's going to look like, but worth experimenting. Do you see anyone doing it well? because the LLM just says, yeah, okay, I'll try it. When, if they were talking to a software engineer, I would have said, ah, you want one of these and one of these. So, I think at some point, we're probably going to have what's the Instagram filters moment for vibe coding or the TikTok moment for vibe coding? I think we're going to have something like that that makes social vibe coding much more. But I don't know what it's going to look like, but worth experimenting. Do you see anyone doing it well? There's Spielwerk. There's an app, a mobile app for vibe coding games and it's TikTok swiping. And there's, I think, I can't remember what it is. There's a London-based game vibe coding tool that is, again, focused on games. I don't know that games are really the thing because they're quite complex. But there are a few people experimenting. But I don't think there are that many people really deeply pushing the boundaries of what's possible. They're mostly lovable, but on your mobile phone, is my opinion. Does anybody else have any opinion, see anyone doing good consumer vibe coding? I guess content creation. [SPEAKER_02] It's, if you buy the building or something, [SPEAKER_02] working. Yeah. You try interacting with digital systems or whatever it is. Well, in your game, you can spawn things in with voice, right? And that's a great, [SPEAKER_04] it's not quite vibe coding, [SPEAKER_04] but it's still interacting. [SPEAKER_04] I think that's what vibe coding still means you have [SPEAKER_04] an understanding of certain primitives. [SPEAKER_04] Even what you just said, [SPEAKER_04] you're thinking about data, [SPEAKER_04] I think when this goes mainstream, [SPEAKER_04] people aren't thinking about this. [SPEAKER_03] And that's what's so exciting for me when you think about culture and what you've built is when it gets to a point where things that us as engineers wouldn't even approach the problem, and that's where some incredible creativity is. So the pattern that I think is closest to being a winner in this space is the Facebook Instant Games API, which doesn't exist anymore. [SPEAKER_02] Oh. [SPEAKER_02] They deprecated it, [SPEAKER_02] but it was in Facebook Messenger. [SPEAKER_02] You could play these games, [SPEAKER_02] and they had these primitives for social gaming. [SPEAKER_02] They tended to be quizzes or Fruit Ninja, [SPEAKER_02] and you compete with your group chats and things. [SPEAKER_02] I bought, for 15 pounds, [SPEAKER_02] a Fruit Ninja clone off a website, [SPEAKER_02] instrumented it with the Facebook Instant Games API, [SPEAKER_02] which was this beautiful [SPEAKER_02] JavaScript async await, [SPEAKER_02] had a get user information, [SPEAKER_02] get friends, create a leaderboard, or get your position on the leaderboard. It's very basic data storage, key value storage, and async await, show a rewarded ad, and show an interstitial ad. And so those things allowed you to make, very quickly and easily, [SPEAKER_04] a social graph-enabled, [SPEAKER_04] ads-enabled experience for consumers. [SPEAKER_04] So I bought this name, 15 pounds, [SPEAKER_04] instrumented it with Facebook Instant Games API, [SPEAKER_04] and went to bed. [SPEAKER_04] The next day, I woke up with 15 million users [SPEAKER_04] on this random game. [SPEAKER_04] I mean, I didn't make very much money, [SPEAKER_04] but there were 15 million users in Vietnam [SPEAKER_04] because, obviously, Facebook, they test everything out in the lower-value advertising regions and then rolls up. But that was amazing, because you've got the social elements of people instantly sharing it around, and that, I think, is probably the template that's going to, something along those lines is probably going to be the thing that wins on the social vibe coding. I don't know if that actually answers the question. On the kind of live working, live building thread, I kind of feel frustration whenever I get a response back in voice. Maybe it's just me, but the input, the information density per second isn't quite high enough. I use a lot of voice out to just get things out of my head. It's like, oh, fuck, and type, and whatever. Or voice input, I guess. Voice input. But I still feel like I need [SPEAKER_04] diagrams or text [SPEAKER_04] or something really high density [SPEAKER_04] back from the system. [SPEAKER_03] So I don't know whether, [SPEAKER_03] how do you guys think about that? [SPEAKER_03] Whether I'm actually curious [SPEAKER_03] if other people can use that [SPEAKER_03] whether they agree with me or something. [SPEAKER_03] It's not whether it's not. [SPEAKER_03] But that's what I find myself leaning towards, [SPEAKER_03] where it's information-rich input, [SPEAKER_03] and then I can just speak my thoughts [SPEAKER_03] as they come out, [SPEAKER_03] and it's almost semantically understood, [SPEAKER_03] put into that information-rich format, and then my intent is spawned out and across the system itself. [SPEAKER_03] So that's the pattern [SPEAKER_03] that I'm seeing and feeling, [SPEAKER_03] and what I always want to evolve [SPEAKER_03] and interact with now, [SPEAKER_03] my email, call, everything. [SPEAKER_03] I feel like that longing, [SPEAKER_03] it's not quite there yet, [SPEAKER_03] but I feel like I have this inclination [SPEAKER_03] to build in that direction [SPEAKER_03] and to try to get information-rich [SPEAKER_03] back, [SPEAKER_03] but also input feed very freeform, [SPEAKER_03] and just my broad intent. [SPEAKER_03] Yeah, that's something that I feel. So that's the pattern that I'm seeing and feeling, and what I always want to evolve and interact with now, my email, call, everything. I feel like that longing, it's not quite there yet, but I feel like I have this inclination to build in that direction and to try to get information-rich, get that back, but also input feed very free form, and just my broad intent. Yeah, that's something that I feel. I absolutely feel, I want to speak, speak easily and quickly, and then receive maybe a little bit of voice, but mostly this generated, maybe it's a UI, maybe it's just diagrams, maybe it's whatever app I'm in context of. But yeah, I want to have a parallel input, a parallel output of single input of my voice. I think the other part of that is I find that what I don't actually get that much information, necessarily, from voice, but what I do get is companionship. So it triggers that, it lessens the loneliness feel somehow. If I'm walking to something and I'm getting some information, I don't feel, I might be learning more if I'm looking at diagrams or text, but I don't feel as, I feel more motivated to continue tinkering. So there's some interesting modalities where you feel different things if you get the information coming in from all these different modalities, at least I find it. Curiously, yeah, how other people think about whether they think differently or whether that's something that they also feel as well. Yeah. Well, the visual cortex is much older than voice and text, so seeing something. Yeah. Also, when you ask it to be concise, I don't feel offended if it gives me a concise answer. But in speech, if it gives a concise answer, I'm relaxed, man, just ask you a question here. You don't have to ask it. You have to be concise and it just sounds rude. That's an interesting, maybe this is possible, maybe it's not, but what does skim listening look like? Yeah, exactly. Maybe listening actually should also have two buttons, forwards and backwards, and I can just tap, tap, tap, tap, tap, go forwards half a sentence until I, I don't know, maybe that's, maybe we should build that. Who wants to vibe code something with me straight up to this? How would it work? So you'd be, you'd just easily, yeah, back and forward going forward with the audience. It's like a speed dial, on a podcast, or on 2X. Yeah, or like the old iPods where you can sort of spin forwards and backwards. I don't know, that's probably quite a nice listening interaction. And you sort of, I guess, you want to scroll forwards in concepts, not necessarily in sentences, right? Like, what is the thing that your eyes look at when you're skim reading? It's probably, it's not the sentence structure, it's the words, I guess. It's the next thing. Yeah, yeah. Yeah, yeah. And that's actually, this is the story, this is very interesting and exciting for me. Because if I'm talking to an agent and it just starts rambling about, sometimes you get back three paragraphs of stuff and I'm, no, not this one. But I want to interrupt it and say, go to the next one. But then it's effectively saying, it's a new prompt, right? So then it's effectively saying, yes, okay, I'll focus on that next one and it'll write me three paragraphs about the second paragraph. You know, that's not really at all what I wanted. Unless I say, be concise and then it says something rude to me. Could you make a summary for each paragraph and then if you hold the, yeah, yeah. It will expand and then all you can. Yeah, I think the Claude app has done some interesting stuff on the voice interactions because they show you something different to you here and they show the higher level sections and then it goes into each one and you can tap on them. So that's, I guess, getting a little bit closer to this. But I think that also means you're not interacting as though you would with a human conversation. I don't know. I'm thinking about that. Like, why do we not have this issue when we're talking to humans? Right? Like, how do we, and they show the higher level sections and then it goes into each one and you can tap on them. [SPEAKER_01] So that's getting a little bit closer to this. [SPEAKER_01] But I think that also means you're not interacting [SPEAKER_01] as though you would with a human conversation. I don't know. I'm thinking about that. Why do we not have this issue when we're talking to humans? Right? How do we, what was there many other cues? [SPEAKER_04] Yeah. [SPEAKER_04] What's right there? There's a cue. We have a scent. There's so many other cues. I think there are visual cues as well. Yeah. I don't know. So that's something. [SPEAKER_02] If I'm as a businessman, if there's an agent response coming into audio, I don't know how long that's going to be. [SPEAKER_02] I also think is it going to be a minute or is it going to be [SPEAKER_03] 10 seconds? [SPEAKER_03] I almost want to know. [SPEAKER_03] And if I know it's really long and long, I have some other. [SPEAKER_03] But also you can tell when I'm about to interrupt you. [SPEAKER_03] So you go faster and you move, maybe you can, and you can tell if people are listening or not. [SPEAKER_03] Yeah. Yeah. [SPEAKER_03] I guess, oh, there's another thing which is interesting here, which is the interrupting. [SPEAKER_03] Sometimes I don't want to interrupt. [SPEAKER_03] I just want to say, yeah, yeah, [SPEAKER_04] yeah. [SPEAKER_04] Or, oh, but no, go back. You're always listening for the, [SPEAKER_06] uh-huh, yeah, yeah, yeah. But you can't do that with an agent. I think I'm willing to, I mentioned this to Joe before, but briefly, I showed you this yesterday, but briefly, you need to be able to interact with a PS5 game and also create what you want in that world as you're playing it. [SPEAKER_04] So you're on a shoot a battle, I'm playing you, we're trying to, you know, [SPEAKER_03] kill each other, you know, blow it, but if I'm trying to be creative and generate a new getaway vehicle or a helicopter, I can just save that experience and have that appear in the game straight away and then I can fly away. [SPEAKER_03] That's the work that I'm doing and how I met Joe actually. [SPEAKER_03] But I think the challenge that I face is that relying on just the voice interruptibility is quite unreliable. [SPEAKER_03] So I just got around to a simple whisper flow type. [SPEAKER_03] Push to talk. [SPEAKER_03] Push to talk. [SPEAKER_03] And then hold to talk and then let go to finish. [SPEAKER_03] Which augments this audio stream with some other cue. [SPEAKER_03] And in the same way that I think you almost need some very light nudge interface on top of the audio that you're receiving. [SPEAKER_03] And maybe as you say, you're saying something and then you see a little circle appearing being the agent wants to ask you a question. [SPEAKER_03] That would be an interesting experience in the field. [SPEAKER_03] If you're not being interrupted but you're being the agent wants to talk and then maybe you either stop and say, okay, what idea do you have? [SPEAKER_03] Because I can feel you doing that to me right now. [SPEAKER_03] You want to say something. [SPEAKER_03] I'm just getting loads of ideas. [SPEAKER_03] This is great. [SPEAKER_03] I'm just having an information communication on that level. [SPEAKER_03] But in audio, if I'm just listening to audio and not, I don't think I'd know that as a developer page or a developer page. [SPEAKER_03] Just listening to audio and not listening to the audio stream. [SPEAKER_03] I'll let you talk. Sorry. Well, I'm just imagining on that point that the agent wants to respond to you. It could be showing, [SPEAKER_03] I want to interrupt and tell you about this thing or this one. [SPEAKER_03] And then suddenly that becomes what we've just done but with even more context than doing it with a human because it's signaling [SPEAKER_03] or a developer page. [SPEAKER_03] Just listening to audio and not listening to the audio stream. [SPEAKER_03] I'll let you talk. Sorry. Well, I'm just imagining on that point that the agent wants to respond to you. It could be showing, [SPEAKER_03] I want to interrupt and tell you about this thing or this one. And then suddenly that becomes what we've just done but with even more context than doing it with a human because it's signaling the topic it wants to talk about. [SPEAKER_03] One question regarding the product because regarding this topic, you need some true calling in the background to be able to actually understand and how would that be handled? Would you be streaming the audio directly to the client or through a back channel and then getting some information to the system from? [SPEAKER_00] Or how can you work? Obviously, this hasn't been, as far as I'm aware, hasn't been built yet. The way I would probably approach it is looking at the transcript and just keep analyzing the transcript over and over again and say, do you have anything to add? Do you have anything to add? Do you have anything to add? Or what would you add? Rather than being a tool call or being part of it, I would do this as an asynchronous looking at the transcript. Are you changing original prompt? Because the original prompt has a plan and you want to do this and halfway you change the plan and you go in and how do you hear the original agent? Well, this is actually an interesting thing with agents. Often we see agents as things that you can't really interact with the internals of. But effectively, an agent is [SPEAKER_00] the logic that says what's the next message but it's also a transcript and that transcript of the conversation is completely malleable. So maybe some other thing we could experiment with is actually allowing the ability to do the interruptions where it's talking and then I can say yeah, yeah, yeah. And currently, if I have the way most agents platforms will work is it'll generate its full text thing and then start generating the audio and then if I talk even while it's partway through it's still reading out the audio, the text it will just append my message to the end of that full message. But we do have the timestamps. We know how far through the audio is played. So we could actually just go and edit the transcript that's coming back and say, well, no, they interrupted at this point so we're going to forget that the LLM even generated more text. Yeah, it's you have used the message and then you have used the message. Yeah. And you need a second type of use message that you are passing in. I don't know that's the of course it's a use message but it's also some other kind of doing it. [SPEAKER_00] Yeah. I mean, why don't you let's go to the 11 labs booth at the expo floor and just vibe code something and see. This is great. I can't wait. Now it's getting easier and easier to make tools and there's so many different ways to execute your ideas and it seems the main differentiate now getting your word out there. How did you approach making those videos and how long did it take you [SPEAKER_01] for the statue video to make it? That's a good question. Totally separate. So, this is the inspiration of this 11 hacks thing. I learned to make videos. I'm not particularly good at it. It's still relatively janky. I learned to make videos through doing other politics related campaigning stuff before I even knew 11 labs existed. It turns out that with videos random things go viral or with content random things go viral. [SPEAKER_01] of this 11 hacks thing. I learned to make videos. I'm not particularly good at it. It's still relatively janky. I learned to make videos through doing other politics related campaigning stuff before I even knew 11 labs existed. It turns out that with videos random things go viral or with content random things go viral and random things don't. Things I think are going to go viral don't and the things that do do. I think it's about practicing. Editing a video I find it's the 80-20 rule. Editing a video to the standard of the Statue app is actually quite relatively easy, but then to go beyond that—because I edited that on my phone—the editing itself took about 20 minutes, 25 minutes. Going beyond the quality of that suddenly means thinking about using a desktop editing tool, which suddenly makes everything even using CapCut on desktop to me is three times harder than using it on my phone and it's also three times more expensive—the subscription on the laptop than on the phone. So I think a lot of it is just doing stuff and trying and iterating. Big things that I found: adding captions really help to the video, having a hook in the first. You can look at your various platform analytics once you've posted a few videos. My videos tend to get between 6 and 12 seconds as the median view time and the people drop off. So you need to get your hook in there because most people are going to drop off if they don't buy the hook. So that's an important piece—front loading the interesting piece. Adding music makes a massive difference, and this is something that was much harder, but now with 11 Labs music generation I will make the video and then add. Sometimes I'll make the video with the narrative and everything and then just experiment with completely different genres until I find music and I'll just put the music on. Does it work? Yes, no. And then you can edit the different sections so that it times up with the sections of the speech so you don't need to figure out music first, find a piece of music and then match your speech to it. The other way, sometimes I will have a vibe I want to get across and I will generate the music first and then figure out what's my speech that matches that vibe. So it might be an excited theme or—I actually chose the music for the Statue app before I made the video itself because I thought this is a fun piece of music for a Statue type thing. It's an imperial outside the British Museum. It sort of made sense as an attention grabber. So but I think music is a massive thing that people underrate because you just put it—it's relatively quiet, but it makes a massive difference to the feeling. So you just did it on CapCut and did you think about it? It's literally on my mobile phone. I've got—I borrowed my wife's lapel mic, Bluetooth lapel mic, which cost 200 quid from DJI, which makes the audio much better and then yeah, just edited on CapCut. Super simple. So you didn't think about like this is the shot I want and then I want to take this photo? I mean the one out the front of the British Museum, yes, thank you. Yes, I did want that one, but the [SPEAKER_01] Much better. [SPEAKER_01] And then yeah, just edited on CapCut. Super simple. So you didn't think about this is the shot I want and then I want to take this photo? I mean, the one out the front of the British Museum? Yes, thank you. Yes, I did want that one, but the— [SPEAKER_01] Actually, the second time I recorded the video, I went once and got some stuff that was a bit boring, and then I went back and just took a bunch of city photos, and that's what ended up being the video. Cool. Thank you very much. Thank you. Thank you. Thank you. Thank you. Thank you. [SPEAKER_04] Thank you. [SPEAKER_04] Thank you. [SPEAKER_04] Thank you. [SPEAKER_04] Thank you. [SPEAKER_04] Thank you. [SPEAKER_04] Thank you. [SPEAKER_04] Thank you. [SPEAKER_04] Thank you. and then I can just speak my thoughts as I come out, and it's almost semantically understood, put into that information-rich format, and then my intent is, like, you know, spawned out and across the system itself. So that's the pattern that I'm seeing and feeling, and what I always want to evolve and interact with now, my email, call, like, everything. I feel, I feel like that longing, it's not quite there yet, but I feel like I have this inclination to build in that direction and to try to get information-rich, you know, get that back, but also sort of, like, input feed very free form, and just my sort of broad intent. Yeah, that's something that I feel. I absolutely feel, you know, I want to speak, speak easily and quickly, and then receive maybe a little bit of voice, but mostly this, like, generated, maybe it's a UI, maybe it's just diagrams, maybe it's whatever app I'm in context of. But yeah, I want to have a parallel input, a parallel output of, like, single input of my voice. I think the other part of that is I find that the, what, I don't actually get that much information, necessarily, from voice, but what I do get is, like, companionship. So it kind of, like, triggers that, it lessens the loneliness feel somehow. If I'm, like, walking to something and I'm getting some information, I don't feel, I might be learning more if I'm looking at diagrams or text, but I don't feel as, I feel more motivated to kind of, like, continue tinkering. So there's, like, some interesting modalities where you feel different things if you get the information coming in from all these different modalities, at least I find it. Curiously, yeah, how other people think about whether they think differently or whether that's something that they also feel as well. Yeah. Well, the visual cortex is much older than voice and text, so seeing something. Yeah. Also, when you ask it to be concise, I don't feel offended if it gives me a concise answer. But in speech, if it gives a concise answer, so I'm relaxed, man, just ask you a question here. You don't have to ask it. You have to be concise and it just sounds rude. That's an interesting, maybe this is possible, maybe it's not, but what does skim listening look like? Yeah, exactly. Maybe, like, listening actually should also have two buttons, like forwards and backwards, and I can just tap, tap, tap, tap, tap, go forwards half a sentence until I, I don't know, maybe that's, maybe we should build that. Who wants to vibe code something with me straight up to this? How would it work? So you'd be, you'd just easily, yeah, back and forward going forward with the audience. It's like a speed dial, like on a podcast, or on 2X. Yeah, or like the old iPods where you can sort of spin forwards and backwards. I don't know, that's probably quite a nice listening interaction. And you sort of, I guess, you want to scroll forwards in concepts, not necessarily in sentences, right? Like, what is the thing that your eyes look at when you're skim reading? It's probably, it's not the sentence structure, it's like the words, I guess. It's like the next thing. Yeah, yeah. Yeah, yeah. And that's actually, oh yeah, this is the story, this is very interesting and exciting for me. Because if I'm talking to an agent and it just starts rambling about, you know, sometimes you get back three paragraphs of stuff and I'm like, no, not this one. But I want to interrupt it and say, go to the next one. But then it's effectively saying, it's a new prompt, right? So then it's effectively saying, yes, okay, I'll focus on that next one and it'll write me three paragraphs about the second paragraph. You know, that's not really at all what I wanted. Unless I say, be concise and then it says something rude to me. Could you like, make a summary for each paragraph and then like, if you hold the, Yeah, yeah. It will expand and then all you can like. Yeah, I think the Claude app has done some interesting stuff on the voice interactions because they, they show you something different to you here and they show the higher level sections and then it goes into each one and you can tap on them. So that's, I guess, getting a little bit closer to this. But I think that also means you're not interacting as though you would with a human conversation. I don't know. I'm thinking about that. Like, why do we not have this issue when we're talking to humans? Right? Like, how do we, what was there many other cues? Yeah. What's right there? There's like a cue. We have a scent. There's so many other cues. I think there are visual cues as well. Yeah. I don't know. So that's something. It's like, if I'm, as a businessman for, like if there's an agent response coming into audio, I don't know how long that's going to be. I also think like, is it going to be like, you know, a minute or is it going to be 10 seconds? I almost want to know. And if I know it's like really long and long, I have some other. But also you can tell when I'm about to interrupt you. So you go faster and you like move, maybe you can, and you can tell if people are listening or not. Yeah. Yeah. I guess, oh, there's another thing which is interesting here, which is the interrupting. Sometimes I don't want to interrupt. I just want to say, yeah, yeah, yeah. Or, oh, but no, go back. You're always listening for the, uh-huh, yeah, yeah, yeah. But you can't do that with an agent. I think, I'm willing to, you know, I mentioned this to Joe before, but briefly, I showed you this yesterday, but briefly, you need to be able to interact with a PS5 game and also create what you want in that world as you're playing it. So you're on like a, you know, shoot a battle, I'm playing you, we're trying to like, you know, kill each other, you know, blow it, but like, if I'm, like trying to be creative and, you know, generate a new getaway vehicle or a helicopter, I can just sort of save that experience and have that appear in the game straight away and then I can sort of fly away. That's the work that I'm doing and how I met Joe actually. But I think the, the challenge that I face is that relying on just the, like, voice interruptibility is quite unreliable. So I just got around to just a simple whisper flow type. Push to talk. Push to talk. And then hold to talk and then let go to finish. Which it augments, like, purely this audio stream with some other cue. And in the same way that I think you almost need some very light nudge interface on top of the audio that you're receiving. And maybe as you say, like, you know, you're saying something and then you see a little, sort of like circle appearing being like the agent wants to ask you a question. That would be an interesting experience in the field. If you're not being interrupted but you're being kind of like, you know, the agent wants to talk and then maybe you either stop and say, okay, what idea do you have? Because I even can feel you doing that to me right now. You want to say something. I'm just getting loads of ideas. This is great. I'm just sort of like, we're almost like having an information, you know, communication on that level. But in audio, if I'm just listening to audio and not, I don't think I'd know that as a developer page or a developer page. Like, just listening to audio and not listening to the audio stream. I'll let you talk. Sorry. Well, I'm just imagining on that point that the agent wants to respond to you. It could be showing, I want to interrupt and tell you about this thing or this one. And then suddenly that becomes what we've just done but with even more context than doing it with a human because it's signaling the topic it wants to talk about. One question regarding the product because regarding this topic, you need some sort of true calling in the background to be able to actually understand and how would that to be handled? Would you sort of be streaming the audio directly to the client or through a back channel and then getting some information to the system from? Or how can you work? Obviously, this hasn't been, as far as I'm aware, hasn't been built yet. The way I would probably approach it is looking at the transcript and just keep analyzing the transcript over and over again and say, do you have anything to add? Do you have anything to add? Do you have anything to add? Or what would you add? Rather than being a tool call or being part of it, I would do this as an asynchronous looking at the transcript. Are you changing original prompt? Because the original prompt has a plan and you want to do this and halfway you sort of change the plan and you go in and sort of how do you hear the original agent? Well, this is actually an interesting thing with agents. Often we see agents as things that you can't really interact with the internals of. But effectively, an agent is the sort of logic that says what's the next message but it's also a transcript and that transcript of the conversation is completely malleable. So maybe some other thing we could experiment with is actually allowing the ability to do the interruptions where it's talking and then I can say uh, uh, uh, yeah, yeah, yeah. And currently, if I have the the way most agents platforms will work is it'll generate its full text thing and then start generating the audio and then if I talk even while it's partway through it's still reading out the audio uh, the text it will just append my message to the end of that full message. But we do have the timestamps. We know how far through the audio is played. So we could actually just go and edit the transcript that's coming back and say, well, no, they interrupted at this point so we're going to forget that the LLM even generated more text. Yeah, it's sort of you have used the message and then you have used the message. Yeah. And sort of you need a sort of a second type of use message that you are passing in. I don't know that's the of course it's a use message but it's also some other kind of doing it. Yeah. I mean, why don't you let's go to the 11 labs booth at the expo floor and just vibe code something and see. This is great. I can't wait. Now it's getting easier and easier to make tools and there's so many different ways to execute your ideas and it seems like the main kind of differentiate now getting your word out there. How did you approach making those videos and how long did it take you for the statue video to make it? That's a good question. Totally separate. So, this is sort of the inspiration of this 11 hacks thing. I learned to make videos. I'm not particularly good at it. It's still relatively janky. I learned to make videos through doing other politics related campaigning stuff before I even knew 11 labs existed. It turns out that with videos random things go viral or with content random things go viral and random things don't. Things I think are going to go viral don't and the things that do do. I think it's basically about practicing. Editing a video I find it's like the 80-20 rule editing a video to the standard of the statue app is actually quite relatively easy but then to go beyond that because I edited that on my phone the editing itself took about 20 minutes 25 minutes. Going beyond the quality of that suddenly means thinking about using a desktop editing tool which suddenly makes everything even using CapCut on desktop to me is like three times harder than using it on my phone and it's also three times more expensive the subscription on the laptop than on the phone. So I think a lot of it is just doing stuff and trying and iterating. Big things that I found adding captions really help to the video having a hook in the first you can look at your various platform analytics once you've posted a few videos my videos tend to get between 6 and 12 seconds is the median view time and the people drop off. So you need to get your hook in there because most people are going to drop off if they don't buy the hook. So that's an important piece front loading the interesting piece. Adding music makes a massive difference and this is something that was much harder but now with 11 Labs music generation I will make the video and then add sometimes I'll make the video with the narrative and everything and then just experiment with completely different genres until I find music and I'll just put the music on does it work yes no and then you can edit the different sections so that it times up with the sections of the speech so you don't need to figure out music first find a piece of music and then match your speech to it. The other way sometimes I will have a vibe I want to get across and I will generate the music first and then figure out what's my speech that matches that vibe so it might be like an excited theme or a I actually chose the music for the statue app before I before I made the video itself because I thought this is a fun piece of music for a statue type thing it's like a bit of an imperial outside the British Museum it sort of made sense as an attention grabber so but I think music is a massive thing that people underrate because you just put it it's relatively quiet but it makes a massive difference to the feeling. So you just did it on CapCut and did you think about it? It's literally on my mobile phone I've got I borrowed my wife's lapel mic Bluetooth lapel mic which cost 200 quid from DJI which makes the audio much better and then yeah just edited on CapCut super simple So you didn't think about like this is the shot I want and then like I want to take this photo I mean the the one out the front of the British Museum yes thank you yes I did want that one but the and it was actually the second time I recorded the video I went once and got some stuff that was like a bit boring and then I went back and just took a bunch of city photos and that's what ended up being the video cool thank you very much thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you thank you you