Open Reader

Building an Agentic Video Editor for Mass Consumer — Ekaterina Deyneka, Reelful

completed 12:44 Aug 18, 2026 Watch on YouTube

Current Status

completed

Video ID

pPj_tjlvYjA

RAG / Chat

Enabled
Building an Agentic Video Editor for Mass Consumer — Ekaterina Deyneka, Reelful
Description

Nearly every hand in the room went up when she asked who had recorded video at the conference. Almost none stayed up for who had actually posted any of it. Ekaterina Deyneka counts herself in that gap, and Reelful is her answer to it: drop in raw footage with a line of direction, and an agent finds the usable moments, cuts them together, and generates captions, music, voiceover, and b roll around them. Her framing for an AI engineering audience is that an agentic video editor is structurally the same thing as an agentic app builder. A prompt goes in, a sandbox spins up, an agent works inside it with tools and skills, and something renders out the other end. The difference that matters is editing rather than generating. A blank canvas lets an agent do anything it likes, while real footage forces it to judge which take is best and what to drop, and to produce something polished from material that is often messy or incomplete. The composition layer is Remotion, which expresses video as React code, chosen precisely because agents write code well. Skills carry the taste: cut rules, font pairings, when a cutaway actually helps. A verification pass catches compositions that will not render and sends the agent back around. All of it hides behind mobile templates, since the point is that a consumer never sees the pipeline at all. Speaker info: - https://x.com/katedeyneka - https://www.linkedin.com/in/katedeyneka - https://www.katedeyneka.com Timestamps: 0:00 - Who recorded video here, and who actually posted it 1:29 - What agentic video editing means 3:33 - The same shape as an agentic app builder 4:10 - Editing real footage is harder than generating 5:30 - The pipeline, from media understanding to a creative plan 6:50 - Remotion, video as React code, and the verification layer 8:49 - Hiding all of it behind mobile templates

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Skim
  • Core thesis: Reelful frames consumer video editing as an agentic workflow: understand supplied footage, propose a creative plan, generate a code-based composition in a sandbox, verify it, and leave users a lightweight editor for final corrections.
  • Why it matters: It offers a practical pattern for turning a subjective creative task into an agent system: separate planning, execution, domain-specific skills, verification, and human approval rather than relying on a single prompt-to-output step.
  • Best use: Use it as a concise product-and-architecture reference for multimodal agent workflows and consumer UX design, not as a deep technical implementation guide.

Executive Summary

Ekaterina Deyneka presents Reelful as a response to a common consumer bottleneck: people record abundant event, travel, and social footage but do not publish it because editing remains tedious, manual, and craft-heavy. The product accepts media plus optional instructions, identifies usable moments, assembles a story, and can add captions, music, voiceover, B-roll, and image animation to produce a shareable clip.

The central technical argument is that an agentic video editor resembles an agentic app builder. Both use a prompt-driven interface, launch an agent in a remote sandbox with tools and skills, and produce a compiled artifact. The key difference is that the agent edits an existing, messy and constrained source corpus rather than generating freely from a blank canvas; it must decide what material to retain, omit, and sequence while still delivering a polished result.

Reelful's workflow is staged: media understanding and transcription; a user-reviewable creative plan; sandboxed execution using editorial skills; video generation through Remotion, an open-source React-based video-as-code framework; and a verification layer that catches composition or rendering errors and triggers iteration. This architecture makes the output inspectable and correctable rather than treating generation as a one-shot black box.

On product design, Deyneka emphasizes reducing prompting burden through directional templates and retaining user control through a conventional editor after agentic generation. The presentation is primarily an early-stage company pitch, with demo examples and a beta announcement, so it gives useful design patterns but little evaluation data on quality, latency, cost, reliability, or user retention.

Key Takeaways

  • Claim: Editing real user footage is a harder agent problem than open-ended video generation because the system must make selective editorial judgments under imperfect source constraints. | Evidence: Deyneka says Reelful expects users to provide their personal photos and videos, which may be messy or incomplete; the agent must select the best moments, decide what to omit, organize them, and still make the result look professionally edited. | Implication: For agent systems operating on real-world inputs, value comes from constrained selection, sequencing, and recovery from imperfect data—not merely from generative capability. | Caveat: The talk asserts this complexity but provides no comparative benchmarks showing that its approach outperforms generative-video or conventional-editing workflows.
  • Claim: A creative-plan approval step should precede costly agent execution in subjective creative workflows. | Evidence: After media understanding and speech transcription, Reelful presents a creative plan for the user to approve, modify, or regenerate before it starts the actual editing work in a sandbox. | Implication: A plan-first interaction can align intent before compute-heavy execution, reduce wasted renders, and give users control over subjective decisions such as narrative, tone, and structure.
  • Claim: Domain craft can be operationalized as reusable agent skills rather than left entirely to a general model prompt. | Evidence: Reelful identifies cut rules for choosing moments, font-pairing guidance, and B-roll generation as skills; Deyneka says this is where editorial taste and craft reside. | Implication: For creative-agent products, encode recurring expert heuristics as explicit, composable capabilities so the agent has more structure than generic multimodal reasoning alone. | Caveat: The presentation does not specify whether these skills are deterministic rules, prompts, model calls, learned policies, or human-authored templates.
  • Claim: Representing a video composition as code gives agents a tractable execution target. | Evidence: Reelful uses Remotion, an open-source framework for creating video as React code; the composition encodes asset order, tracks, and sequencing, and Deyneka argues agents are well suited to write this code. | Implication: Where a creative output can be expressed in a declarative or code-native intermediate representation, agents can generate, revise, diff, test, and render it more reliably than manipulating opaque project files.
  • Claim: Verification must be a first-class stage because agent-generated compositions can fail even when the creative plan is sound. | Evidence: Reelful adds a verification layer to confirm that the composition is clean, well-defined, and renderable; if problems are found, the agent iterates on the composition. | Implication: Any artifact-producing agent should validate the actual executable output and feed failures back into an iteration loop, rather than equating generated instructions or code with a completed deliverable. | Caveat: No details are provided on the checks performed, failure rates, automated quality criteria, or whether verification evaluates aesthetic quality versus technical renderability.
  • Claim: Mass-consumer adoption requires hiding agent complexity while preserving familiar, limited editability. | Evidence: Reelful is mobile-first, supplies directional templates such as speak-to-camera, B-roll, or voiceover, and lets users make small post-generation changes such as removing a second or correcting caption text in a built-in editor. | Implication: For consumer agent products, replace blank-prompt interfaces with intent templates and use a generate-then-tweak model that preserves agency without forcing users into a full professional workflow. | Caveat: The claim that users can edit while driving is presented as a convenience example and should not be treated as a safe usage recommendation.

Detailed Brief

System shape: agentic app builder analogy

  • Claims: Deyneka positions agentic video editing as structurally similar to agentic app building.; In both cases, a user-facing prompt interface delegates work to an agent operating in a remotely provisioned execution environment.; The generated artifact differs: an app builder produces an app preview, while the video workflow produces a rendered video.
  • Evidence: For video editing, the input is explicitly media plus a prompt or directions rather than text alone.; The talk calls the remote execution machine a sandbox and describes it as the environment in which the agent runs its tools and skills.
  • Caveats: The analogy is conceptual; the speaker does not cover sandbox isolation, media storage, permissions, rendering infrastructure, GPU requirements, or operational cost.
  • Implications: The same control-plane pattern—request intake, workspace provisioning, tool execution, artifact validation, and delivery—can transfer across different agentic artifact-generation products.

Product proof points and maturity

  • Claims: The company demonstrates social-content examples assembled with its agent rather than a conventional video editor.; Reelful is still early, is offering a second-version beta, and is explicitly seeking feedback on user behavior and use cases.
  • Evidence: The demo example is a narrated clip about an event dinner, combining venue, food, conversation, and gift highlights into a social-ready narrative.; Deyneka says the company was recently funded by A16Z Speedrun.
  • Caveats: The demonstration is self-reported and contains no before/after comparison, editing-time measurement, pricing, retention data, or independent quality assessment.; The transcript repeats a substantial portion of the presentation, reducing the source's incremental watch value.
  • Implications: Treat Reelful as an early product worth monitoring or testing for workflow insight, rather than as validated evidence that fully autonomous consumer editing has reached broad product-market fit.

Notable Concepts & Terms

  • Agentic video editing: A workflow where an agent turns user-provided photos and footage plus optional direction into an edited, ready-to-share clip.
  • Creative plan: An intermediate, user-reviewable proposal created after media analysis and before execution; it serves as the approval and alignment layer for a subjective task.
  • Sandbox: The remote machine spun up for the agent to execute editing tools and create the composition.
  • Editorial skills: Reusable representations of craft decisions such as cut selection, font pairing, and B-roll treatment that guide the agent beyond a generic prompt.
  • Remotion: An open-source framework for generating videos with React code; Reelful uses it as the code-based representation of the video composition.
  • Verification layer: A post-composition validation stage intended to ensure the output is clean, defined, and renderable, with agent iteration when failures are found.
  • Directional templates: Prebuilt user intents—such as speak-to-camera, B-roll, or voiceover—that reduce the need for consumers to formulate effective prompts.
  • Generate-then-tweak: Reelful's interaction model: generate an edit agentically, then allow minor manual changes in a familiar editor.

Operator Notes / Why Ken Should Care

  • Adopt the plan-before-execution pattern for any costly or subjective agent workflow: expose an editable proposed plan, secure approval, then provision execution resources.
  • Evaluate video or creative agents against a staged architecture: multimodal ingestion, structured intermediate representation, skill/tool layer, executable artifact generation, and output verification.
  • If testing Reelful or a comparable product, measure the missing decision metrics: time-to-publish, accepted-output rate without manual edits, render failure rate, per-video cost, latency, and repeat usage.
  • For OpenClaw or other artifact-producing agents, prioritize representations that can be programmatically inspected and rerun—analogous to Remotion's video-as-code—over opaque end-state outputs.
  • Avoid treating technical render verification as sufficient quality assurance; add evaluation for narrative coherence, source-faithfulness, caption accuracy, and taste alignment if deploying a similar system.

Source/Metadata

  • Title: Building an Agentic Video Editor for Mass Consumer — Ekaterina Deyneka, Reelful
  • Transcript words: 2535
  • Duration seconds: 764
  • Timestamp note: No usable timestamps or chapter markers were present; the latter portion of the transcript substantially repeats earlier material.

Transcript

1564 words en Processed in 67.5s

Hi, everyone. I think we can start. But before we start, I want to ask you a couple of questions. So first of all, how many of you took a photo or video during this conference? Please raise your hands. Okay. And how many of you actually posted any video content from it online? Not that many. And to be honest, that was me. I was recording a lot of content during conferences, events, trips, meetups, and I never posted it online because video editing is hard. It sounds, and it is, a lot of work. It's tedious, and it's largely still manual. It also feels like an art and not really automated. And that's why we're building Reelful, and we're trying to tackle an agentic video editing problem from the agentic standpoint. I'm Kate. I'm founder and CEO at Trilful. But let's first talk about what agentic video editing is. As a user, you just drop in your media, photos and videos, and provide some context. It can be the context of what happened in these media files, or it can be some directions. For example, add captions, add music, add voiceover, and something like that. And then the agent will understand your media, find the right moments, assemble everything together, generate captions, music, voiceover, B-rolls, and give you a ready-to-share clip. And yeah, that's agentic video editing. The agent does everything by itself. Another example: you recorded a speak-to-camera video, and you have a lot of pauses, unsuccessful shots, and you expect an agent to figure it out, remove unsuccessful shots, remove pauses, and give you a ready-to-share clip. And the interesting thing is that a lot of the things inside this pipeline can be automated. And this is exactly what we're doing at Trilful. Oh, this is the example of video edited. And since we're at an AI engineering conference, I wanted to talk a little bit about infrastructure. And from the infrastructure standpoint, an agentic video editor is very similar to an agentic app builder. Sorry, there is a typo on the slide. So the second column is agentic video editor. Both of them have a prompt, a UI for prompt. And in the video editor case, it's media plus prompt. And usually on the backend, what's happening is there is a remote machine, which is called a sandbox, which is spinning up. And inside this machine, there is an agent with tools and skills, which is working on what you're asking it to do. In the case of the agentic app builder, it's a code base. In the case of the agentic video editor, it's a video composition. And as a result, in the agentic app builder, the user gets an app review. And for the video editor, the user gets rendered video. And yes, the infrastructure, from an infrastructural standpoint, is pretty similar, but there are a couple of differences. And this is actually the most interesting to me: generating versus editing. At Realful, we're focusing on editing real footage. So we do not generate a lot of content. We are expecting you to provide your real-life, your personal content, and we will edit it for you. And actually, this is a more complex problem. Because if the agent has a blank canvas, it can do whatever it can. But in the editing case, the agent has to figure out which moments are the best, what to omit, what to use, how to organize everything together. And also, sometimes footage can be messy or incomplete, and the agent still has to deliver a very polished result, professionally made, so that ideally, the viewers of this content don't get if it is AI- or human-edited. So let's actually have a look at how we do it. So let's actually have a look at Realful. We start, as I already mentioned, with your media plus a prompt, some directions like how you want it to be edited. And we need to get a polished clip. So let's go through it step by step. So we're doing first media understanding. We need to understand what's actually happening on those clips and photos. And we also need to transcribe speech, for example, in the case if you have speak-to-camera videos. Then, we are providing a creative plan for the user so that they can approve if they like it or not, what they want to change, or maybe regenerate. So we create this plan before actually starting editing. Once the user approves this plan, we spin up a sandbox, the remote machine that we already discussed. And this is an environment for the agent to execute everything. So the agent comes with the skills. And in our case, in the case of video editing, our skills are, for example, cut rules, for example, how to select the best moments. Also, font pairs, which fonts are more suitable for this use case, which are not. For example, how to generate B-rolls. And this is where taste and craft live, actually. And then also, the agent can initiate some other subprocesses, for example, generating music that will fit this exact composition, generating voiceover, adding sounds, animating images. Yes, this is actually what we do. If you provide photos, we can animate your photos to make them more dynamic and engaging. And then comes remotion composition. So here, a little bit of background: what's remotion? Remotion is an open-source framework to create videos as code, as React code. So basically, it's just a file with the order, with all your assets and tracks, and how they're following each other. And why is it important? Because agents are really good at writing code. And therefore, we can use them to create videos with this remotion framework. And then the last thing is the verification layer. Of course, the agent can make mistakes. And that's why we developed this verification layer to make sure that all the composition is clean, is well defined, everything will be rendered. And if there are some problems, then the agent will reiterate on the composition. And this is how we developed this application. And this is how we got to a polished clip. So it's a lot, right? It's a very complex workflow. And ideally, we don't want our users to even know anything about it. And this is how we're tackling that at Reelful. So we decided to go mobile-first so that users can edit videos while driving, walking, or maybe lifting weights. Also, I know that prompting videos can sometimes be challenging. That's why we create directional templates. For example, speak-to-camera videos, or maybe you want to add B-rolls or voiceover, so that users can just select these directional templates, drop their media, and that's it. Even without any prompt, it will work. And the third thing is a building editor. Why? Because we want to make this experience convenient and familiar for users. A lot of people are already using regular video editors, and that's why we want to provide this experience as well. So how it works: the user first generates a video agentically, but if they want to tweak it, for example, remove a second, or maybe correct some word in the captions, they can go into building editor and edit it a little bit. Yeah, and actually, I have a couple of examples here that I recently created with Reelful. I will play them, just maybe one of them. Oh, sorry. Last week, I was invited... Do you hear? A little bit. Okay, you just can enjoy the video. Exclusive person creators dinner, and honestly, it was one of the best event experiences I've had. First of all, the venue was stunning. It was a custom event space transformed into this tropical sunset place. The whole atmosphere felt so warm and cinematic. Second, the food was way beyond my expectations. They brought in private chefs from LA, and we had logs, a few crop, and this incredible strawberry ice cream that I was still thinking about. But most importantly, the conversations were so much fun. It was atmosphere. It felt really easy to connect with people, talk about what they're working on, and just enjoy the community. And lastly, we got gifts. One of the highlights... So yeah, all these videos were assembled only using Agent, no regular video editor, and I already posted them on social media. And yeah, I have a lot of fun with that. Oh, sorry. And exclusively for this conference, we are giving our beta, which is our new second version. Please give it a try, and let me know if you have any feedback. Here is my email. Please feel free to reach out. We're still early. We're actively working on it. So we will be happy to hear any feedback, and also curious how you use it. And also, some exciting news. We recently got funded by A16Z Speedrun. So, I'm very excited to continue working. Yeah, that's it. Thank you so much. And because it's a presentation about content and how to edit videos, I have to film a video with you all. Yay! What to emit, what to use, how to organize everything together. And also, sometimes, footage can be messy or incomplete, and agent still has to deliver a very polished result, professionally made, so that ideally, the viewers of this content don't get if it is like AI or human edited. So let's actually have a look how we do it. So let's actually have a look at Realful. So we start, as I already mentioned, with your media plus a prompt, some directions like how you want it to be edited. And we need to get a polished clip. So let's go through it step by step. So we're doing first media understanding. We need to understand what's actually happening on those clips and photos. And we also need to transcribe speech, for example, in the case if you have speak to camera videos. Then, we are providing a creative plan for the user so that they can approve if they like it or not, what they want to change, or maybe regenerate. So we create this plan before actually starting editing. Once the user approves this plan, we spin up a sandbox, the remote machine that we already discussed. And this is an environment for the agent to execute everything. So the agent comes with the skills. And in our case, in the case of video editing, our skills are, for example, cut rules, how, for example, how to select the best moments. Also, fund pairs, which funds are more suitable for this use case, which are not. For example, how to generate B-rolls. And this is where taste and craft live, actually. And then also agent can initiate some other sub processes, for example, generating music that will fit this exact composition, generating voiceover, adding sounds, animating images. Yes, this is actually what we do. If you provide photos, we can animate your photos to make them more dynamic and engaging. And then comes remotion composition. So here, a little bit of background. What's remotion? Remotion is a framework, open source framework, to create videos as code, as React code. So basically, it's just like a file with the order with all your assets and tracks and how they're following each other. And why it is important? Because agents are really good at writing code. And therefore, we can use them to create videos with this remotion framework. And then the last thing is the verification layer. Of course, agent can make mistakes. And that's why we developed this verification layer to make sure that all the composition is clean, is well defined, everything will be rendered. And if there are some problems, then the agent will reiterate on the composition. And this is how we developed this application. And this is how we got to a polished clip. So it's a lot, right? It's like a very complex workflow. And ideally, we don't want our users to even know anything about it. And this is how we're tackling that. And this is how we're tackling that at Reelful. So we decided to go mobile first, so that users can edit videos while driving, walking, or maybe lifting weights. Also, I know that prompting videos can sometimes be challenging. That's why we create directional templates. For example, like speak to camera videos, or maybe you want to add B-rolls or voiceover, so that users can just select these directional templates, drop their media, and that's it. Even without any prompt, it will work. And the third thing is a building editor. Why? Because we want to make this experience convenient and familiar for users. So a lot of people are already sort of using regular video editors, and that's why we want to provide this experience as well. So how it works, user first generates a video agentically, but if they want to tweak it, for example, remove a second, or maybe correct some word in the captions, they can go into building editor and edit it a little bit. Yeah, and actually, I have a couple of examples here that I recently created with Reelful. I will play them, just maybe one of them. Oh, sorry. Last week, I was invited... Do you hear? A little bit. Okay, you just can enjoy the video. Exclusive person creators dinner, and honestly, it was one of the best event experiences I've had. First of all, the venue was stunning. It was a custom event space transformed into this tropical sunset place. The whole atmosphere felt so warm and cinematic. Second, the food was way beyond my expectations. They brought in private chefs from LA, and we had logs, a few crop, and this incredible strawberry ice cream that I was still thinking about. But most importantly, the conversations were so much fun. It was kind of atmosphere. It felt really easy to connect with people, talk about what they're working on, and just enjoy the community. And lastly, we got gifts. One of the highlights... So, yeah, basically, all these videos, they were assembled only using Agent, no regular video editor, and I already posted them on social media, and yeah, I have a lot of fun with that. And... Oh, sorry. And exclusively for this conference, we are giving our beta, which is our new second version. Please give it a try, and let me know if you have any feedback. Here is my email. Please feel free to reach out. We're still early. We're actively working on it. So, we will be happy to hear any feedback, and also curious how you use it. And also, some exciting news. We recently got funded by A16Z Speedrun. So, I'm very excited to continue working. Yeah, that's it. Thank you so much. And because it's a presentation about content and how to edit videos, I have to film a video with you all. Yay!