Every

Experiments with Kieran, OpenAI Live Voice inside Compound Engineering

2534 summary words 11 min summary Watch video

Start with the signal

11 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: Live-streaming a working session integrating OpenAI's real-time voice API into the Compound Engineering plugin's Polish command, creating a live-feedback loop where users can speak design changes to a running app and have coding agents fix them in real time.
  • Why it matters: This is a practical demonstration of agent workflow orchestration, multi-repo coordination in Cursor Projects, designing human-in-the-loop feedback systems, and extending agentic tooling from engineers to general knowledge workers—all core to Ken's agent orchestration, control planes, and workflow design interests.
  • Best use: Watch to observe working patterns for multi-agent coordination, brainstorming-to-implementation pipelines, streaming real-time data from browser to agents, scoping library vs. plugin boundaries, and designing live-feedback UX for agentic systems.

Executive Summary

Kieran (a builder at Every/Avery) live-streams an exploratory working session integrating OpenAI's real-time voice API into Compound Engineering, a GitHub plugin for agentic coding workflows. The session is unpolished and conversational, capturing his process for scoping a new "live mode" for the Polish command—a feature that lets users give live spoken feedback to a running web app while coding agents implement fixes in real time. He brainstorms whether to extend his existing RiffRack library (originally designed for structured feedback capture as zip files) to stream live audio and visual annotations, or to build a separate live-feedback layer. He uses Cursor Projects to coordinate work across multiple repositories (Compound Engineering and RiffRack), runs a brainstorm agent in the cloud, and discusses design trade-offs such as instant vs. batched edits, mode switching for safe vs. aggressive agent behavior, and how to surface annotation overlays using libraries like perfect-freehand instead of full TLDraw. The session also touches on plans for "Compound Work," a non-engineering sibling plugin for knowledge work, and reflects on user taste and power-user feedback loops. The video is raw operational footage, not a polished tutorial, showing how an experienced builder thinks through agent orchestration, library boundaries, and UX for human-agent collaboration.

Kieran emphasizes the value of brainstorming with agents before implementation, spending an hour or more refining requirements in natural language and monologue notes (captured on Apple Watch during a hike, then fed into Cursor). Once the plan is clear, agents can execute in parallel with minimal looping. He describes Cursor Projects as enabling true multi-repo agent work with access to all source code across local and cloud runners, and notes the cost/benefit trade-off of always-on cloud agents. The session ends with him kicking off a planning agent to generate implementation units, intending to review progress in a follow-up stream.

The session models a realistic agentic workflow: capture unstructured ideas (monologue, voice memos), refine them with an agent (brainstorm), generate a plan, then execute across multiple repositories with cloud and local agents. It also highlights the tension between library abstraction (RiffRack as a general tool) and plugin-specific needs (Compound Engineering's Polish command), a recurring architectural decision in agent tooling.

Key Takeaways

  • Claim: Live-streaming brainstorming and scoping with agents before coding saves iteration time | Evidence: Kieran spends ~1.5 hours brainstorming (monologue hike notes + Cursor agent conversation) to clarify requirements, then expects agents to execute the plan in one go without rabbit holes | Implication: Invest time upfront in structured brainstorming with agents to reduce downstream rework; treat agents as thought partners before executors | Caveat: This approach requires clarity from the human; vague briefs will still cause looping.
  • Claim: OpenAI's real-time voice API enables live-feedback loops for design iteration | Evidence: Kieran plans to add a "live mode" to Polish where users speak design changes to a running app (e.g., "change this color, move this element") and agents implement fixes in real time using streaming audio and DOM annotations | Implication: Real-time voice + visual overlays (drawing, annotations) can replace traditional annotation tools with a more fluid human-agent UX; useful for design polish, user testing, and rapid iteration | Caveat: Requires low-latency streaming architecture and clear boundaries for when agents act (instant, batched, or manual send).
  • Claim: Cursor Projects enable multi-repo agent coordination with shared context | Evidence: Kieran uses Cursor Projects to coordinate work across Compound Engineering and RiffRack repos; agents have access to all source code across local (Mac Mini) and cloud runners, enabling cross-repo planning and execution | Implication: Multi-repo agent workflows require tooling that shares context and execution environments; Cursor Projects provides a working model for this | Caveat: Token costs increase significantly with always-on cloud agents and large context windows.
  • Claim: Library vs. plugin boundaries matter for reusability and agent usability | Evidence: Kieran debates whether to extend RiffRack (a library for feedback capture) to support live streaming or keep it separate; libraries are hard to update but provide clear separation of concerns, and agents use them as-is without modification | Implication: Design agent-facing libraries with stable APIs and clear scope; resist adding features that blur boundaries, even when convenient | Caveat: Over-abstraction can make libraries rigid; organic growth is acceptable if the library remains coherent.
  • Claim: Mode switching (instant, smart, manual) gives users control over agent aggressiveness | Evidence: Kieran proposes three modes for live feedback: instant (agents fix everything immediately), smart (agents batch and triage), and manual (user approves changes); this avoids conflicts when agents implement changes that interfere with each other | Implication: Expose mode controls in agent UX to let users balance speed vs. safety; this is especially important for design work where rapid iteration is desired but large refactors can cause conflicts | Caveat: Mode switching adds UI complexity; default mode must be well-chosen.
  • Claim: Power users and taste are the bottleneck for product feedback, not tooling | Evidence: Kieran argues that not everyone has good design or product taste; power users who deeply understand the product are the best sources of feedback, and tooling should prioritize them over general users | Implication: Design feedback systems for high-signal users first; democratizing feedback without taste filters creates noise | Caveat: This is opinionated and may not apply to all product contexts; early-stage products may benefit from broad feedback.
  • Claim: Separating Compound Engineering (for engineers) from Compound Work (for knowledge workers) clarifies scope and reduces overlap | Evidence: Kieran plans to create a standalone Compound Work plugin for non-engineering workflows (document review, summarization, compounding knowledge) instead of bundling it with Compound Engineering; users can install both if needed | Implication: Agent tooling should align with job-to-be-done boundaries, not force unification; engineers and knowledge workers have different workflow shapes, so separate plugins reduce complexity and improve UX.

Detailed Brief

Multi-Agent Workflow Orchestration

Kieran's session demonstrates a practical pattern for multi-agent orchestration: (1) capture unstructured input (monologue, voice memos, hike notes), (2) refine with a brainstorming agent, (3) generate a plan with units of work, (4) execute in parallel across multiple agents and repositories. He uses Cursor Projects to coordinate agents running on local (Mac Mini) and cloud machines, with agents able to inspect all source code across repos. The brainstorming phase took ~1.5 hours but clarified requirements enough that agents could execute without looping. This "front-load clarity, then YOLO execute" pattern mirrors Ken's interest in control planes and workflow orchestration for agent systems.

Kieran also shows the value of reviewing AI-generated plans for big decisions vs. letting agents execute directly for small features. He reads through the brainstorm output, adds clarifications, and then kicks off a planning agent to generate implementation units. This hybrid human-agent planning loop is key to avoiding wasted work.

Real-Time Feedback Architecture

The live-feedback system Kieran is designing streams audio (via OpenAI's real-time voice API), visual annotations (drawing overlays using perfect-freehand), and DOM events from the browser to coding agents. The architecture must support low-latency streaming (not polling) and handle mode switching (instant, smart, manual) to control when agents implement changes. Kieran references the Compound Engineering prototype's existing light web server and streaming setup as a model, showing how prior architectural decisions ease extension.

He debates whether to extend RiffRack (his existing feedback library) to support streaming or keep it separate. RiffRack originally packaged feedback as zip files; the new live mode requires streaming. Kieran leans toward extending RiffRack with a new "stream contract" as its public API, making the zip schema a backward-compatible legacy mode. This keeps the library general-purpose and reusable across projects.

Cursor Projects and Multi-Repo Coordination

Kieran praises Cursor Projects for enabling multi-repo work with shared context. Agents can inspect all source code across repos and execute on local or cloud runners. This is critical for the live-feedback system, which spans Compound Engineering (the plugin) and RiffRack (the library). He notes that Cursor Projects adopted patterns from "GrokBot philosophy," likely referring to always-on, context-aware agent systems.

The cost of Cursor Projects is higher token usage, which Kieran acknowledges as "probably good, maybe bad." For Ken, this trade-off is worth studying: always-on cloud agents with large context windows improve UX but increase costs, so usage-based pricing and caching strategies become critical.

Library Design and Agent Usability

Kieran emphasizes that libraries should be hard to update so agents use them as stable dependencies rather than modifying them. This forces clear separation of concerns and prevents agents from blurring boundaries. He likes libraries that "grow organically" but remain coherent. This philosophy aligns with Ken's interest in reusable agent components and orchestration patterns.

For RiffRack, Kieran debates whether to add OpenAI API keys (for real-time voice) even though the original design was "just capture, no intelligence." He decides it's acceptable because the goal is to put "AI as close to the user" for fast, optimistic UX. This pragmatic shift—adding intelligence to a capture library—shows how architectural principles bend when UX demands it.

Mode Switching for Safe vs. Aggressive Agents

Kieran proposes three modes for live feedback: instant (agents fix everything immediately), smart (agents batch and triage based on complexity), and manual (user approves changes). This avoids conflicts when agents implement changes that interfere with each other (e.g., simultaneous design tweaks and refactors). He suggests exposing the mode switch in the UI so users can choose speed vs. safety based on the task.

This is a concrete example of user control over agent behavior, a key pattern for Ken's agent systems. Mode switching also maps to broader questions about when agents should act autonomously vs. wait for human approval.

Compound Work: Knowledge Work Plugin

Kieran plans to create Compound Work, a separate plugin for non-engineering workflows (document review, summarization, compounding knowledge). He describes it as "Compound Engineering but without the engineering." The plugin would include skills like KW-dream (summarize and ideate), KW-review (analyze documents), KW-compound (synthesize learnings), and KW-strategy (plan knowledge work). He emphasizes that the shape of knowledge work differs from engineering work, so a separate plugin avoids forcing unification.

This separation-of-concerns approach is relevant to Ken's interest in agent specialization and workflow design. Compound Work also highlights the idea of "folder-as-agent," where agents operate on document trees rather than code repos.

User Taste and Feedback Loops

Kieran argues that "the hardest thing" in product development is finding people with taste, and that power users—who deeply understand the product—are the best sources of feedback. He proposes that live-feedback tooling should prioritize high-signal users over democratizing feedback to everyone. This is opinionated but aligns with his view that "if you don't have taste, it's just hard." For Ken, this highlights the importance of user segmentation and feedback quality gates in agent-assisted product development.

TLDraw and Annotation Overlays

Kieran explores using TLDraw or its underlying library (perfect-freehand) for drawing overlays in the live-feedback UI. TLDraw offers a full SDK but requires a license key (free hobby tier has a watermark). Perfect-freehand is a thin, MIT-licensed library by the same author. Kieran leans toward perfect-freehand for simplicity. This is a minor detail but shows his preference for minimal dependencies and clear licensing.

Notable Concepts & Terms

  • Compound Engineering: GitHub plugin for agentic coding workflows; includes commands like Polish (design iteration), Brainstorm (scoping), and Prototype (rapid builds). Kieran works on this full-time.
  • Polish command: A Compound Engineering feature for iterating on AI-generated code by seeing it run and giving feedback; the live-feedback system extends this with real-time voice and annotations.
  • RiffRack: Kieran's React library for capturing user feedback (screen recordings, clicks, DOM state) as zip files; he's extending it to support live streaming for the new live-feedback mode.
  • Cursor Projects: Multi-repo IDE feature that gives agents access to all source code across local and cloud runners; Kieran credits it for enabling his workflow and increasing token usage.
  • Mode switching (instant/smart/manual): User control over when agents implement changes in the live-feedback loop; prevents conflicts and gives users speed vs. safety trade-offs.
  • Folder-as-agent: Design pattern where agents operate on folder structures (e.g., document trees for knowledge work) rather than just code repos; Kieran wrote a piece on this and plans to apply it to Compound Work.
  • Monologue: Voice memo app (by Naveen) that Kieran uses to capture ideas on Apple Watch/iPhone during hikes; transcripts feed into Cursor for agent brainstorming.
  • Perfect-freehand: MIT-licensed library by TLDraw's author for drawing smooth freehand lines; Kieran prefers this over full TLDraw SDK for the annotation overlay.

Operator Notes / Why Ken Should Care

  • Action: Study Kieran's "brainstorm first, YOLO execute" pattern for OpenClaw and Ken's agent projects. Front-loading clarity with agents (1-1.5 hours) before implementation reduces iteration waste. Apply this to model routing, workflow generation, and agent orchestration design.
  • Action: Experiment with real-time voice + visual overlays for agent feedback loops in Ken's own tools. OpenAI's real-time API + DOM annotations could replace traditional UIs for agent control and debugging.
  • Action: Review Cursor Projects' multi-repo coordination model as a reference for OpenClaw's control plane. Shared context across repos and execution environments is a solved problem in Cursor; Ken should study their API design and runner architecture.
  • Action: Adopt Kieran's library design principle: make libraries hard to update so agents treat them as stable dependencies. This applies to Ken's reusable agent components and orchestration primitives.
  • Action: Implement mode switching (instant/smart/manual) in OpenClaw or similar systems where agents take actions that could conflict. Give users explicit control over agent aggressiveness, especially in design/ops workflows.
  • Decision: Evaluate whether Ken's agent tooling should separate engineering vs. knowledge-work use cases (like Compound Engineering vs. Compound Work). If workflows differ significantly, separate plugins/modes may reduce complexity and improve UX.
  • Monitor: Kieran's Compound Work plugin (when released) for patterns in document-based agent workflows, folder-as-agent design, and non-engineering agentic systems. This could inform Ken's internal knowledge management tools or GTM systems.
  • Relevance note: This session is highly relevant to Ken's work on agent orchestration, control planes, multi-agent coordination, and human-in-the-loop workflows. The unpolished format makes it more valuable, showing real working patterns and architectural trade-offs in agent system design.

Source/Metadata

  • Title: Experiments with Kieran, OpenAI Live Voice inside Compound Engineering
  • Transcript words: 7,033
  • Video duration: 58 minutes 51 seconds
  • Timestamp note: No formal chapters; transcript includes repeated segments due to live-streaming buffering or background audio, but core content is intact. The session is unscripted and conversational, with Kieran narrating his workflow in real time.
Full transcript 5755 words · 31 min read
0:15

Hello everyone.

0:22

I think that's working. Okay, check, check, check. I'm playing some ambient stuff on YouTube in the background.

0:31

Let me know if you don't hear me. I think you should be able to hear me. I just wanted to go jam on some stuff and share it.

0:46

Because I had some fun ideas and sharing is fun to do.

0:56

So, I would love to go over some things here. Let's see. Also, I'm trying out the live streaming like this. This is better probably for readability. So, probably this is good.

1:20

Okay, cool. So, actually, we launched a new version of the plugin, the Compound Engineering plugin. Which is really cool. So, there's some goodies in here. There's a new prototype stuff with annotations. And if you annotate it, automatically this cool stuff is really cool. A prototype is something Treven started and I did some iterations on that. So, I'm going to go ahead and see if I can bring Compound Engineering to people that are not engineers. I do work that is engineering work, obviously. But also, I do work that is just knowledge work or experiments, things like that. And, yeah, I have a way that I do that.

2:32

And I was thinking maybe it's time to extract that way of working. And also, I've been asking around within Avery to see how other people work. And they all have their own flavor, which is really cool. So, I was thinking, okay, it would be cool if we have a framework around that. So, working title, Compound work. It's separate from Compound Engineering. And it's standalone. It can work alone. You can also use it together. It will be experimental. But I'm figuring this out today. But this is not what I want to work on now. I want to just walk you through this idea I have here.

3:15

So, I have a library called Riff Rack. And the idea of Riff Rack is that the hard part using AI is getting good signal.

3:27

Get golden information where a user says something very valuable or you have an experience and you want to capture that. And you can capture that in many ways. And Riff Rack is a React package that you can install and it will record your screen. And all kinds of other things like where you click, what the DOM elements are, all those things. And it will package that. But it's not real time. And I was experimenting with real time audio. So, I created this other app, a Breathwork Coach. And I used a live audio feature that just got released from OpenAI. And it's really cool.

4:24

So, I'm thinking, how can I bring that here? Maybe looking at the website, responding to it, clicking on it, maybe drawing in real time. And the agent immediately starts working on it. And, yeah, that could be cool. And I think we can do that in codecs, possibly, because there's live mode and WebMCP and stuff like that. Which is interesting, but I'm thinking I can experiment with this in Compound Engineering. For example, if you do Polish or something like Polish, which is more in the loop, that in Polish there is a mode that is live mode. And it starts up the server, does the live mode and just goes through it like an extension of Polish.

5:14

I think that would be cool to experiment with. So, basically, yesterday I walked somewhere and I jotted down all my notes using monologue. I did a hike, actually. Let's see here. Monologue.

5:56

I did a hike, actually. Let's see.

6:02

Yeah, okay. Cool. Yeah, so I did a hike and I just brain dumped everything into monologue. Here's monologue.

6:11

It's from Naveen. I love it. It's on my Apple Watch, on my iPhone. It's everywhere. And I just dumped it in Cursor. I love Cursor, especially the projects now. Cursor projects are amazing. They really took something from the GrokBot philosophy and put it in Cursor. And I think from all the IDEs and agentic coding and whatever tools there are, I think Cursor gets it really right with the idea of always on cloud, but also local. Because you want all of it and it's just all connected, which is really cool. So, that's why I'm here. I love the projects. I'm spending way more tokens than before the projects, which is a good thing, maybe. Maybe a bad thing. Who knows?

7:09

Okay, let's see. Let's go back to this. So, I did a brainstorm and just run a CD Brainstorm on here. It's doing this. It's inspecting the repositories I was talking about. And the brainstorm was done here. And let's see where the questions are.

7:35

So, the questions here, like, hey, who is the person talking to the agent in the first version? And the options here is like, if you do this, Riff Rack is also made so maybe this is for users of your app. So, you're in the app and you like, give feedback. But I think the version we're building now actually is really only for Compound Engineering. So, we need to steer that a little bit. I was thinking maybe all and yeah, did some questions about, hey, the next question is like, this defines the feel. Is the Park Sealy agent single thread? When the voice agent extracts the unit, when does the coding agent get woken?

8:28

So, imagine you're talking and saying, hey, this needs to be a different color. This blah, blah, blah. Do you have it start immediately fixing these things or analyze? Or do you send, click a button and hit send? It's kind of like annotate in a cursor and then you say send to agent or you collect them in a batch. I think I like both because I like to see a live board, like some kind of collection of what we need, like an optimistic or quick thing. Because there will be an agent in the browser with the live so we can do things. And then have the coding agents whenever they're like the right boundaries or bulk or whatever, there's a pulse.

9:14

And we can do that maybe with some kind of tool call where the front agent decides it's good to send. So, the next question was, what does the coding agent do when it wakes with a batch of accepted units? Accepted means the app changes under your hands while you ask.

9:38

Yeah, so... Yeah, I think... So, triage only.

9:56

So, it just looks at stuff, implements the small stuff, implements everything parallel. Deciding per unit. Actually, maybe you can set the mode this is in, in the UI.

10:09

So, it's like instant mode, smart mode, or never do anything mode. One of those three. Okay, so I monologue that. I think that would be cool. So, for example, if you're in the polishing phase and you're doing design things, I don't care. You can do all kinds of things. But if it's a bigger thing or a refactor, it might bite each other. You can just manually go in a different mode. I think that's interesting. Okay. So, mode switch. That's good.

11:14

Also, I want to change of plan a little bit.

11:19

I want to add this to the Compound Engineering plugin to the CE Polish command. I think this is a perfect Polish thing. So, maybe it's a mode, like a live mode for Polish. Or, yeah. So, maybe it's a mode using Polish. It should ask if you want to go live or traditional or something like that. I can go a little bit lower here. So, make it a little bit faster.

12:02

Okay, so that's good. For anyone that just came in, I'm working on... I'm improving the Polish command in Compound Engineering and basically what it will do, it will add a live mode.

12:39

So for example, if you design... Let me open what Polish actually does. So it's basically when you have something done, the AI did all the work, and you go in and you make it even better. Really great way is to just see something in front of you and Polish it, where you...

12:57

Okay, so that's good. For anyone that just came in, I'm working on improving the polish command in Compound Engineering. Basically, what it will do is add a live mode. So for example, if you design something, let me open what polish actually does. So it's basically when you have something done, the AI did all the work, and you go in and you make it even better. Really great way is to just see something in front of you and polish it, where you say, "Hey, move this, make this better, can you iterate some versions here?" So that's the polish command. But I really like the OpenAI live modes, where you can just talk to the machine and it does it. So I'm just thinking, why can't I add that to it? So I'm going to see if I can improve polish so that there's a live mode, where it just opens a website and you can start talking to it and have a conversation like you have with the human. And then while you do it, it will just fix everything in the background inside Compound Engineering. So that is what we're working on.

13:00

And I'm now in Cursor. I'm in the brainstorm phase. I want to use my Riffreq library for this. I want to implement it so that it will install Riffreq in your repository. And yeah, we can work from there.

13:12

So let's see. Okay, so the traditional dev server. Okay, so it is good. There is already something like starting a dev server and stuff like that. So because we need to start some kind of CLI that's watching or looking for interactions from the browser. Okay, so the board list in the in-app overlay panel, dashboard drops out. Okay, yeah, dashboard is something I started, but let's focus. Yeah, so that's good. Live or traditional fires. Runtime. It needs to open an API key and a mic, obviously. Yeah, so okay, so let's see. Overlay delivery, dropped extensions, dropped extensions, river and desktop character. Okay, yeah, so there's this question. The Riffreq originally did something differently because it captured everything in a zip file and now it's streaming. So the question is, does it even need to be Riffreq or does it need to be just Compound Engineering? Maybe.

13:17

That's a good question. Should we add this to Riffreq as a mode? Because it's a different kind of mode, right? I do like not having everything in Compound Engineering and also looking forward to being able to upgrade Riffreq so that this works in non-coding sessions. That's the delivery streaming instead of a zip file. So I would love to frame it like that. And also for Polish, can we look how we do this in the Compound Engineering prototype skill because we have a similar pattern for the server and the listener? So you can see in the latest version of Compound Engineering, we have prototype. And we have a light web server. And this one kind of, yeah, starts a web server and also binds and listens to it. So there's streaming. So you can start a server and then open your browser and then we can receive information from whatever the website is into the agent and the agent can take actions on that. So that is in the prototype now as well, which is really cool in the latest release. Okay, so what I'm doing now is just figuring out whether it makes sense. I think this is an interesting idea and it will work. It's just where do we put everything. And normally AI is a little bit conservative. So you need to go wide where you say, "Hey, actually this, I want to use this in different ways as well." And it's good to talk like, "What would you do if this were a year-long project kind of thing?" And then from that big view, you zoom in because it knows the direction. It knows maybe the abstractions where you say, "Hey, maybe it needs to be a library versus a script," stuff like that. And what I love from libraries is that it's hard to update them. So that means the AI just uses them however they are. But also it's a clear separation of concerns. So I like creating libraries that I use in multiple projects. And it's fine if they grow organically. And I think that's also why Cursor projects are so good. Because projects, they have access to all the runners and the workers I have on my own machine, on my Mac Mini, on the cloud. So they actually have access to all of the source code of all the projects. So you can say, "Hey, look there." Like the project, it can just look anywhere. So yeah, that's really cool in projects. So just make it bigger actually. There is also this project overview, which is interesting. So it kind of keeps track of stuff it's doing. Yeah, this is cool. Okay, so let's go further with what's here? So okay, new shape in this place. Rift, live, mode, physical ability. Stream instead of zip. Overlay, drawing layer. Drawing layer. Oh yeah, oh yeah. For the drawing, can we use TL Draw? I think they have an SDK. Can we use that in here? You can also multitask here, which is great. And actually just that. Okay, let's review this. So I love to review all of this. Okay, so phlags, carrying proposal. Okay, the stream contract becomes the new public API. The zip schema today. So zip is the API now. Yeah, the stream contract. Curious how it does that. I'm curious how the stream contract is. Is this polling? Is this how do we do it now with the prototype? Just keep it super simple but elegant. So no Elm calls in Rift Rack. No API gets rewritten. Yeah, so there are API keys. Because we need the OpenAI real time there. Which is fine. I think that's a fine change. It's just something that I said in the beginning. I said Rift Rack is just capture and it's not intelligence. But the whole point is you want AI as close to the user. You want it to be fast and optimistic and feel instant. Similar to Cursor, it feels very fast. But it's also a remote machine. It's just yeah, it just needs to feel great. And I think having it local is making it feel great. Proxy is gone. Target React hosts. Do I just build framework agnostic? Okay, so plain script embed is cheap. Okay, targets yeah, okay, so if it's not React, sure. But for now, most of my apps are React. And most people's apps are React. So that's good. Also, if anyone has any idea what I'm doing or questions, please post them. I do see them here. So and welcome. I am just jamming doing my normal thing that I do on a normal September day. Chilling in the backyard. It's lovely weather. I have some water, which is very important to hydrate. And also, you're watching a stream from me and I work at Avery, which is the only subscription you need to stay at the edge of AI, as you can see in this beautiful scrolling thing. We're doing vibe checks. We're checking out what the future of work is. And me struggling or doing this is figuring out whether this is the future of work or not. So let's go back to this real-time thing. It would be cool if we can have this working at the end and maybe ship it in a version. Let's see. I would imagine I would need to be video because it's too big. But if you say this needs to change color, probably we want to see that. And if you draw on the window, we probably want to see that as well. Either in structured data, event data, DOM data, stuff like that. As long as it's very clear to the agent what it is. And the TL Draw. Hello. For anyone new coming in, I'm building a way to give live feedback to any... Hey, how's everybody doing? Can you guys hear me? Sorry if my music is my name is Geller. This is Public Sounds. This one's the first one.

13:19

Need to be video because it's too big. But if you say this needs to change color, probably we want to see that. And if you draw on the window, we probably want to see that as well. Either in structured data, event data, DOM data, stuff like that. As long as it's very clear to the agent what it is. And the TLDraw. Hello. For anyone new coming in, I'm building a way to give live feedback to any... Hey, how's everybody doing? Can you guys hear me? Sorry if my music is... My name is Geller. This is Public Sounds. This one's the first one. I'm building a way to give live feedback to a website just by clicking and annotating,

14:22

maybe drawing and then have the agent live fix it for me. Because there are all these annotation tools and stuff like that. But with live voice API from OpenAI. I'm thinking, well, that's great. Let's see if we can do something. And I'm building it inside Compound Engineering plugin. And I'm adding this to the Polish command.

14:41

Because Polish is really this thing you do where you take something that the AI generated and go in and actually raise the bar, make it really good, make it yours.

14:57

And yeah, just looking at stuff and using stuff is for me the absolute best way to give feedback. Because you're experiencing it.

15:08

So that's what we're doing. And I'm brainstorming the ID now. We'll see how far it gets. Let's see. Yeah, so, okay. Next question. How much evidence writes with each unit? Transcript.

15:28

Yeah. Rich lean. Yeah, for... Yeah, for... Evidence. Just do as much as possible. Probably not live video. But we should definitely take screenshots. Just anything we can that will make it super clear for the agent.

15:56

We can also run some optimizing to see if anything makes it better or more confusing. So maybe do that as experiments later on. We should work on that as well. Yeah, this sometimes happens for some reason. For some reason, my shortcuts are working on. Okay, so how much evidence to write? I think as much as we can.

16:10

Definitely screenshots. We should take screenshots once in a while. And... That song is actually called... We should do some experiments also to see if adding things or removing things are better or worse. I actually started that song at one of the last public sounds. So I love playing it out here. I already did it.

16:51

Stop talking. Yeah.

17:04

Okay.

17:17

Okay.

17:23

So TLDraw recommends Ganset for in-package overlay. Okay. SK4. License key. Free hobby key with watermark.

17:58

Okay. Okay. So maybe leaning. Thin custom layer on perfect freehand. Sure. By TLDraw's author. Okay. Sure. Okay. That sounds good to do perfect freehand. Okay. So back to the stream contract. Not polling in the busy sense. Exactly. Prototype three. Okay. Page. So this is how we're getting the information from the website that we're testing to the coding agents.

19:12

Can you ensure that we can also run this over a tunnel or... Yeah. Because it might be that we run this in a remote machine and want to open it in another machine. Just make sure it's working. Similar to what we have in Prototype. I think it's supported there as well. Okay. So you can see I kind of reading through and adding things already. I know I'm not answering exactly the questions at the bottom, but this makes it a little bit faster. Next question. Open a first cut. Okay. So if the first cut had to be worth using next week. Okay. Only one slice. No, I don't want to do one slice. I want the full loop always. This is one of those things of AI agents.

20:06

It's good that they're pushing you back, but also normally they can just do everything. This is the modern version of the plan. This plan will take you six weeks and then it does it in two hours. I think it's pushing for scope, but if it's clear and you're good in what you want to build and everything is just set well, it's just good. Okay. So shortcuts. Okay.

20:32

This is what I did. Perfect. Blah, blah, blah, blah. Okay. So let's see.

20:47

Door and brainstorm. The dog review.

20:54

Okay. So I think this is brainstorming and in this, this takes the most time. So you can see this probably took me with the initial thoughts, maybe an hour, hour and a half total or something, but the whole point is now if things are very clear and detailed and worked out you can just hand it to an agent and it does it in one go and that saves so much time. Looping, iterating and it going down the wrong rabbit holes and stuff like that. So now it's saving this.

21:14

Let's see probably where, yeah, in the riffreq repo. It's there. It's fine. So what I like about projects as well is this is a multi-repo thing. Initially it was one, four repositories.

21:27

I scoped it down to two, compound engineering and riffreq. But I could definitely say, oh, actually I want to implement this in another repository for users as well at some point. So. I would like to link up there now and I would link up there now and I would link up there now. I would link up there now and I would link up there now and I would link up there now and link up there now and link up there now and link up there now and link up there now and link up there now and link up there now and link up there now and link up there now and link up there now and link up there now and link up there now and link up there now and

21:54

link up there now and link up there now and link up there now and So what do you do when you wait on these things? I've been thinking, sometimes I take a walk, sometimes I go down, play some piano, or you switch to another project. Like here. So let's see that.

22:19

So I was working on something else here. Oh, we have a question as well. Let's see. If people have questions, please ask. Let's see. Reflect. So what you're doing seems fantastic for integrating into your DevStack, but upstream, I'm curious how you choose who to test ideas with and how you pitch them on testing your product market fits. Yeah, that's a very good question. Of course, yourself. You have the best taste. And I think this is going also in products. And obviously, if everyone gets feedback on products, it's not valuable probably because not everyone has good design taste or product taste.

22:58

I think still giving access to people you trust or internal people or people that are power users or people that are on the most paying plan. Some kind of leverage, people like that. At least it's interesting to see what comes out of there. And I think it's also very important to have inside your process a way to evaluate whether it's a good idea and it's aligned with whatever you want to do. Is it really pushing it in the right direction? Is it polishing or is it destroying?

23:46

And yeah, I think for now, just making it super easy for power users or yourself to do this is a first step. But I do want to bring this to real time given feedback on the product. And then it's actually solving the product for you. It's a new version of open source where the source is open, but that's very engineering heavy. This is more, hey, if you're a user, you know how to use a product. This is a version of that where it's open for you to build and do things with.

24:24

And yeah, I think that will be the future more and more. And yeah, I'll share more things there. Definitely. Yeah, definitely. So yeah, that's my answer to this. I don't have the exact answer, but I am working on things that will make sure that you check whether something you do is aligned with whatever you want and your company wants. So yeah.

25:05

Okay, so it's writing the plan. Let's see, let's go to here. Did you write the brainstorm with everything in it already? Let me update the think room document so we can review. So normally I like to jump from one to the other a little bit. It's frying my brain as well. So I think taking breaks is also a good one. So yeah, we should take a five minute break. Okay, I am going to take a five minute break. I'll just keep streaming here, drink some water and probably this will be back.

25:52

Yes, Brandon. I agree. And I don't have the answer. But I think if you build guardrails and yeah, the whole goal now is to find people with taste because that's the hardest thing.

26:11

And honestly, if you don't have taste, yeah, it's just hard.

26:16

And users are normally people with taste or power users because power users use your software. And they deeply understand how they use it because they use it. Did you write the brainstorm with everything in it already?

26:25

Let me update the think room document so we can review.

26:28

So normally I like to jump from one to the other. It's frying my brain as well.

26:34

So I think taking breaks is also a good one. So yes, we should take a five minute break. Okay, I am going to take a five minute break. I'll just keep streaming here, drink some water and probably this will be back. Yes, Brandon. I agree. And I don't have the answer. But I think if you build guardrails and the whole goal now is to find people with taste because that's the hardest thing. And honestly, if you don't have taste, it's just hard.

27:15

And users are normally people with taste or power users because power users use your software.

27:22

And they deeply understand how they use it because they use it. So I would say anyone that uses your software for a long time, they are probably a very good person to do this. Okay, now break. Be back in five minutes to see what the brainstorm did. Cool. So one nice thing is you can see here, this is the agent and it runs in the clouds. It could run on the computer, it could run anywhere and you can click into it. So you could even steer the agent here and I love this pattern. So it's writing the plan. And multitask is very nice as well. So you can just say, do this, do that. It can steer already running agents. Can we start them? All of that.

28:03

Okay.

28:19

Let's see. We want to see if anything comes back from the document review and ask any questions or automatically resolve them if you have a good idea. Can you run the compound engineering planning next when that is done? When all that is done, run the planning. So we have all the units and also separated out in what repository, stuff like that.

28:34

And then we have all the units.

28:44

And then we have the plan here. You can see here, it's fun because here you can see this one ran on a computer by Mac Mini. This one is in the clouds. You can see what's doing. I think this is just going. Normally, I would just go to something else. I think this is set up to actually finish everything.

28:56

I think I'll just do a follow-up post to see how this goes. Maybe tomorrow, actually, as a live stream.

29:04

Let's jump into compound work, which is also very interesting. It's compound engineering, but then without the engineering, doing knowledge work. And I have a version running locally, but just having that in a scale as well. And it's cool because we have different people using different things. So, for example, Katie, Jack, Mike, Natasha, Dan, all the folks. They have a little bit of their own flavor, which is really cool. So you could be inspired by what they do. Or create completely your own. And I'm leaning towards making this completely separate from compound engineering.

29:30

Because I've been going back and forth. I was like, you want one thing that does all the work. But for work, if you're not an engineer, you want to do work. The shape of work is very different. And you can install both. But then it's, oh, but there's the brainstorm is really helpful in compound engineering.

29:55

Well, then install compound engineering to use them together. So I think I want to do a compound work plugin separate. So maybe should we call it compound work plugin for GitHub, like compound engineering? Or should we do it compound writing? So we have compound writing from Katie as well. I like to review these plans. Still, if it's a bigger feature. If it's just add this feature, make it work. It's YOLO. But for these, these are actually big decisions.

30:21

And I want to learn what I want here. So this is also for me to be inspired by. Let's see. Decided shape. So repository separate relationship to compound engineering is... Yeah. So complementary ecosystem usable work. Work is not standalone. Brainstorm plan. Brainstorm plan, doc review. Yeah. You install both. It could be standalone. But I think the framing is... So the framing says compound work is not alone. But it could be alone.

31:11

But it should also work together with compound engineering. I think that is the framing.

31:18

Okay. So there's no overlap. KW.

31:35

Knowledge work. I think. Because CW is already taken by compound writing. Knowledge work. I like that.

32:08

KW. Knowledge work. Okay. Cool. Compact work. Yeah. Okay. So what are the skills? Let's see. Let's see if we need...

32:52

Let's load this here actually. Okay. Let's see. Okay. process that and put it in the right place. Knowledge work dream... dreams and summarizes stuff. We need to... I need to go into what they do and how they work for real. KW experiment... Review and source...

33:26

Okay. This is not... Okay. Let's look at the skills introduced here. I think you can look at my giant example for plan, day, week, as well. I think we should also have something like compound engineering or strategy.

33:48

And so we can have a knowledge work strategy thing as well. Knowledge work experiments...

33:56

Not sure if we need it, but... Obviously, the most important one is missing. KW compounds. For example, if you dream or review, it can also compound knowledge. Because what did you learn? And... Do we miss anything else? Maybe do a compound engineering ideate to see if we miss any skills in compound work. Thinking about it also having to work completely independent from compound engineering. But also not copying literally from compound engineering. Because you could use both. Okay, so the whole point is that... I wrote a piece about the folder is the agent and... Oh yeah. I will go. Someone says my battery is very low. Yes.

34:27

I will stop the live stream when my battery dies. I will stop the live stream.

34:44

Yes. So basically, we kicked off the whole... Live voice inside compound engineering here. It's going to run and do LFG and build everything. I'm going to probably tomorrow do another live stream where we go see what it actually does. I think... That's running. Like, is it really pushing it in the right direction? Is it polishing or is it like destroying? And yeah, it's but I think for now, just making it super easy for power users or yourself to do this is a first step. But I do want to bring this to like this real time given feedback on the product. And then it's actually solving the product for you like it's kind of a new version of open

35:46

source where the source is open, but that's very engineering heavy. This is more like, hey, if you're a user, you know how to use a product. This is kind of a version of that where it's open for you to build and do things with. And yeah, I think that will be the future more and more. And yeah, I'll share more things there. Definitely. Yeah, definitely. So yeah, that's kind of my answer to this. Like, I don't have the exact answer, but I am working on things that will make sure that you check whether something you do is aligned with whatever you want and your company wants. So yeah.

36:39

Okay, so it's writing the plan.

36:47

Let's see, let's go to here.

36:57

Did you write the brainstorm with everything in it already?

37:06

Let me update the think room document so we can review. So normally I like to jump from one to the other a little bit. It's kind of frying my brain as well. So I think taking breaks is also a good one. So yeah, we should take a five minute break. Okay, I am going to take a five minute break. I'll just keep streaming here, drink some water and probably this will be back.

37:54

Yes, Brandon. I agree. And I don't have the answer. But I think if you build guardrails and yeah, like the whole goal now is to find people with taste because that's the hardest thing. And honestly, if you don't have taste, like, yeah, it's just hard. And users are normally people with taste or power users because power users use your software. And they deeply understand how they use it because they use it. So I would say anyone that uses your software for a long time, they is probably a very good person to do this. Okay, now break. Be back in five minutes to see what the brainstorm did.

39:05

I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would say I would

39:31

Thank you.

40:00

Thank you.

40:27

Thank you.

40:54

Thank you.

41:27

Thank you.

41:55

Cool. So one nice thing is you can see here, this is the agent and it runs in the clouds. It could run on the computer, it could run anywhere and you can click into it. So you could even steer the agent here and I love this pattern.

42:27

So it's writing the plan.

42:41

And multitask is very nice as well. So you can just say, do this, do that. It can steer already running agents. Can we start them? All of that.

42:58

Okay. Let's see. We want to see if anything comes back from the document review and ask any questions or automatically resolve them if you have a good idea. Can you run the compound engineering planning next when that is done? When all that is done, run the planning. So we have all the units and also separated out in what repository, stuff like that. And then we have all the units.

44:07

And then we have the plan here. I would link to the link to link to link to link link link link link link link link link link Thank you.

44:41

Thank you.

45:24

Thank you.

45:29

So you can see here, it's fun because here you can see this one ran on a computer by Mac Mini. This one is in the clouds. You can see what's doing.

45:43

I think this is just going.

45:49

Normally, I would just go to something else. Like, I think this is set up to actually finish everything. I think I'll just do a follow-up post to see how this goes. Maybe tomorrow, actually, as a live stream. Let's jump a little bit into compound work, which is also very interesting. It's, yeah, compound engineering, but then without the engineering, doing knowledge work. And I have a version running locally, but just having that in a scale as well. And it's cool because we have different people using different things. So, for example, Katie, Jack, Mike, Natasha, Dan, all the every folks. They have a little bit of their own flavor, which is really cool.

46:54

So, you could be inspired by what they do.

46:59

Or create completely your own. And I'm leaning towards making this completely separate from compound engineering. Because I've been going back and forth. I was like, you want one thing that does all the work. But for work, if you're not an engineer, you want to do work. The shape of work is very different. And you can install both. But then it's like, oh, but there's like the brainstorm is really helpful in compound engineering. Well, then install compound engineering to use them together. So, I think I want to do like a compound work plugin separate. So, maybe should we call it compound work plugin for GitHub, like compound engineering?

47:48

Or should we do it like compound writing? So, we have compound writing from Katie as well.

47:58

I like to review these plans. Still, if it's like a bigger feature. If it's just like add this feature, make it work. It's YOLO. But for these, these are actually big decisions. And I want to learn what I want here. So, this is also for me to be inspired by. Let's see. Decided shape. So, repository separate relationship to compound engineering is...

48:40

Yeah. So, complementary ecosystem usable work. Work is not standalone.

48:51

Brainstorm plan. Brainstorm plan, doc review.

49:01

Yeah. You install both. It could be standalone. But I think the framing is...

49:11

So, the framing says compound work is not alone. But like it could be alone. But it should also work together with compound engineering. I think that is the framing.

49:36

Okay. So, there's no overlap.

49:40

KW. Knowledge work. I think. Because CW is already taken by compound writing.

49:51

Knowledge work. I like that. KW. Knowledge work. Okay. Cool.

49:59

Compact work. Yeah.

50:16

Okay. So, what are the skills? Let's see. Let's see if we need... Let's load this here actually.

50:39

Okay. Let's see.

50:44

Okay. Okay. Okay. Okay. Okay. Okay.

50:53

Okay. process that and put it in the right place. Knowledge work dream... dreams and summarizes stuff. We need to... I need to go into what they do and how they work for real.

51:17

KW experiment...

51:27

Review and source...

51:37

Okay. This is not... Okay. Okay, let's look at the skills introduced here. I think you can look at my giant example for plan, day, week, as well. I think we should also have something like compound engineering or strategy. And so we can have a knowledge work strategy thing as well.

52:12

Knowledge work experiments...

52:18

Not sure if we need it, but... Obviously, the most important one is missing. KW compounds. For example, if you dream or review, it can also compound knowledge. Because what did you learn? And...

52:42

Do we miss anything else? Maybe do a compound engineering ideate to see if we miss any skills in compound work. Thinking about it also having to work completely independent from compound engineering. But also not copying literally from compound engineering. Because you could use both.

53:13

Okay, so the whole point is that... Uh... I wrote a piece about the folder is the agent and...

53:26

Oh yeah. Oh yeah. I will go. Someone says my battery is very low. Yes. I will stop the live stream when my battery dies. I will stop the live stream. Yeah. So basically, we kind of kicked off the whole... Uh... Live voice inside compound engineering here. It's going to run and do LFG and build everything. I'm going to probably tomorrow do another live stream where we go see what it actually does. Uh... I think... Uh... That's running. And... Now I'm...

54:16

I'm... I'm...

54:22

I'm...

54:30

I'm... I'm... I'm... I'm... I'm... I'm... I'm... I'm... I'm... I'm... I'm... I'm... I'm... I'm...

54:41

Thank you.

55:06

Thank you.

55:36

Thank you.

56:06

Thank you.

56:36

Thank you.

57:06

Thank you.

57:36

Thank you.

58:06

Thank you.

58:38

Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note