AI Engineer

The Next Medium: Why Real-Time Interactive Video Changes Everything — Ahmed Ahres, Reactor

1952 summary words 9 min summary Watch video

Start with the signal

9 min read

Summary

At-a-Glance

  • Verdict: Skim
  • Core thesis: Ahmed Ahres argues that generative video becomes a new programmable medium—not merely a faster rendering tool—when it can generate pixels continuously, respond to inputs, and preserve an interactive session in real time.
  • Why it matters: If the technical constraints are solved, real-time world models could shift video from a fixed asset-generation workflow to a software-like surface for interactive entertainment, simulation, personalization, editing, and embodied-agent training.
  • Best use: Use this as a market and infrastructure primer on real-time video/world models; skim for the model taxonomy, deployment requirements, and candid gaps around memory, evaluation, and deterministic simulation.

Executive Summary

Ahres defines "world models" operationally as real-time interactive video: an indefinitely running visual environment whose output can be conditioned and changed while it is being viewed. His central contrast is with current batch video generators such as Veo/Sora-style systems, which return an immutable clip after a prompt. In his framing, real-time generation eliminates the "slot machine" workflow because creators can observe output immediately and steer it continuously.

He argues that real time changes the category of product that can be built, rather than simply reducing latency. His historical analogies are GPS enabling ride-hailing and digital camera previews enabling the creation loop behind social-video platforms. The claimed equivalent for generative video is programmable pixels: software can alter a scene, inject new elements, let users control a character, or render personalized content during a live session.

The talk groups the opportunity into three model/product types: open-ended interactive video; controllable character-and-world environments similar to Google's Genie; and live interactive avatars. Proposed applications span participatory livestreams, game/movie hybrids, robotics simulation, education, medical and cooking training simulations, customer support, and real-time video-to-video editing. Reactor positions itself as an API and infrastructure layer exposing models from ByteDance, Alibaba, and NVIDIA.

The most useful operational content is the admission that real-time inference is architecturally distinct from batch generation. It requires streamed pixels, persistent session memory, and globally distributed GPU capacity to sustain sub-100-ms interaction. Important unresolved limitations remain: models lose context over time, evaluation of real-time consistency is unsolved and still relies substantially on human judgment, and Reactor does not itself provide deterministic rule engines for reliable simulations.

Key Takeaways

  • Claim: Real-time interactive video is a materially different medium from batch-generated video because users can condition and alter the output continuously rather than accepting a completed clip. | Evidence: Ahres demonstrates an image-conditioned dog scene generated live, then prompts a cat to appear; he says the scene could continue indefinitely and be steered into new events such as running, jumping, or a dragon appearing. | Implication: Treat real-time video as an emerging interactive application surface, closer to a rendered software session than to an AI content-generation endpoint. | Caveat: This is a product thesis and demo-driven claim, not evidence that current systems can yet deliver reliable long-form narratives or production-quality control.
  • Claim: Instant feedback is the core mechanism by which real-time generation improves creator control and may end the trial-and-error "slot machine" experience of prompt-based media generation. | Evidence: Ahres compares fixed generated videos with digital camera previews: seeing a shot as it is captured permits adjustment, whereas film and batch generation require waiting for a finished result. | Implication: For creative tooling, the differentiator is not only model quality; it is a tight human-in-the-loop control loop with low enough latency to support iterative decisions. | Caveat: Immediate feedback helps steering, but it does not by itself solve fidelity, continuity, brand safety, or precise directorial control.
  • Claim: The addressable application space extends beyond games to interactive media, synthetic simulation, and live service experiences. | Evidence: Examples include user-voted livestreams on X, YouTube, or Twitch; Bandersnatch-like game/movie hybrids; robotics environments with effectively unlimited generated training data; immersive history education; medical simulation; cooking simulation; customer support; and sales or training avatars. | Implication: The nearer-term opportunity may be bounded training, simulation, previsualization, and interactive entertainment use cases rather than unrestricted general-purpose world generation. | Caveat: Several examples are prospective or community experiments rather than validated scaled businesses, and their usefulness depends on consistency and controllability that remain immature.
  • Claim: Real-time ad creation and placement could eventually replace much of the pre-produced advertising workflow with dynamically generated, personalized creative. | Evidence: Ahres suggests that a system could insert a logo or product relevant to something a viewer searched for a minute earlier, without producing a conventional ad in advance. | Implication: Dynamic in-video generation is a significant monetization vector, but it is also a high-risk area requiring brand-control, safety, and data-governance layers before deployment. | Caveat: He explicitly notes that brand adoption will be slow because brands are sensitive to visual defects and fearful of AI-generated misuse; privacy, consent, and advertising-policy constraints are also not addressed in the talk.
  • Claim: Real-time world-model infrastructure cannot be built as a simple extension of batch video inference infrastructure. | Evidence: Batch systems process a request in the cloud and return a file; real-time systems instead must stream pixels to the client, maintain a constantly running session and memory state, and route users to geographically nearby GPUs. | Implication: Any real-time multimodal-agent or interactive-video stack should make session state, streaming transport, regional capacity, and memory consistency first-class control-plane concerns. | Caveat: Ahres states that models still struggle with memory; a character may turn away and fail to retain prior context, as seen in Genie 3-style demos.
  • Claim: Sub-100-ms end-to-end latency and distributed compute are prerequisites for preserving the feeling of real-time interaction at global scale. | Evidence: Ahres says users in India or Japan need routing to GPUs in or near those regions; otherwise latency breaks the medium. In Q&A, he cites multi-GPU execution, weight optimization, and quantization as approaches used to improve performance, including around a question about 16 FPS. | Implication: Do not evaluate platforms on model demos alone; request region-specific p95 latency, sustained FPS, session duration, concurrent-session capacity, and cost-per-interactive-minute data. | Caveat: The presentation provides no measured latency, frame-rate, quality, availability, or cost benchmarks for Reactor or the named underlying models.
  • Claim: Reliability measurement and deterministic behavior are unresolved problems for real-time world models. | Evidence: Ahres says evaluation of real-time consistency and fidelity is unsolved across the research community, including DeepMind, and that current assessment is largely visual human judgment. He also says Reactor does not provide deterministic engines or rule sets, though developers are building such layers on top. | Implication: Avoid treating generated environments as authoritative simulators for medical, robotics, or other consequential domains without external rule constraints, task-specific evaluation, and human validation. | Caveat: Pixel-level fidelity is easier to assess than whether the simulated world remains coherent, causal, safe, or fit for training decisions.

Detailed Brief

Reactor's platform positioning and model catalog

  • Claims: Reactor positions itself as a Series A developer platform and infrastructure provider rather than as a single proprietary world-model developer.; Its stated goal is to make real-time interactive-video models accessible through APIs so developers can embed them in video tools, image applications, plugins, or other products.
  • Evidence: Ahres names four models available through Reactor: Helios, described as an interactive-video model from ByteDance; Linkbot, a Genie-like world model trained by Alibaba; Long Live 2 from NVIDIA for multi-shot, story-consistent film generation; and Sound of Streaming from NVIDIA for video-to-video editing.; For video-to-video workflows, he describes uploading footage or Seedance 2-generated material and then adding effects, removing people, or changing backgrounds; he identifies Hollywood previsualization as a use case.; He says integration can begin with an API key and approximately 10 lines of code, while acknowledging that this simplifies the underlying complexity.
  • Caveats: The talk does not provide model cards, licensing terms, supported regions, reliability metrics, benchmark comparisons, or production customer case studies.; He says current community-built video-editing products are not yet very good because underlying model quality remains limiting.
  • Implications: Reactor may be useful as an abstraction layer for prototyping across multiple real-time model providers, but platform diligence should distinguish API convenience from demonstrated production performance.; Model-provider dependency and licensing/availability risk could matter if building a durable product on top of the catalog.

What programmable video changes in the production workflow

  • Claims: The speaker's broader analogy is that media production will increasingly resemble software development: visual output can be addressed, conditioned, and modified while the experience is live.; Interactive output can merge roles traditionally separated between creator and consumer, allowing viewers to determine what occurs next.
  • Evidence: The GPS analogy is used to argue that real-time location did more than make maps faster: it enabled Uber.; The digital-preview analogy is used to argue that immediate visual feedback improved content production and helped make platforms such as Instagram and TikTok possible.; The livestream concept has audience members typing or voting on the next event because the rendered pixels do not need to be pre-authored.
  • Caveats: The historical analogies illustrate potential category creation but do not establish that real-time generative video will achieve a comparably broad adoption curve.; Audience-directed generation creates moderation, abuse-prevention, rights-management, and narrative-quality problems that are not discussed.
  • Implications: The relevant product question is not merely whether a team can generate better clips; it is whether a live feedback loop creates an experience impossible with pre-rendered media.; Interactive formats will likely need orchestration layers that constrain inputs, preserve narrative/state coherence, and control what user actions can affect.

Notable Concepts & Terms

  • World models: Ahres uses the term broadly to mean real-time, interactive video systems rather than a single technical definition based on physical simulation or 3D representations.
  • Programmable video: Video whose pixels and events can be addressed and changed through software during playback, making it an interactive surface rather than a fixed file.
  • Slot-machine generation: The speaker's critique of batch video generation: prompt, wait, receive a fixed result, and rerun if it is wrong rather than steering it live.
  • Infinite interactive video: A video generation mode that keeps running beyond a pre-set 5-, 10-, or 30-second clip and accepts interventions during the session.
  • Genie 3-like models: Image-and-text-conditioned environments in which a user controls a character, positioned as a bridge from generative video to interactive worlds, games, and simulation.
  • Live session memory: Persistent state required for a real-time model to remember prior actions and world context; the speaker identifies it as a major current weakness.
  • Deterministic engines: Rule-based constraints or simulation logic that can keep an AI-generated environment coherent and predictable; Reactor does not currently supply this layer.
  • Real-time world-model evaluation: The unsolved problem of measuring temporal consistency, interaction fidelity, and world coherence beyond judging individual pixels or relying on human review.

Operator Notes / Why Ken Should Care

  • If evaluating a real-time video/world-model provider, require a technical scorecard covering regional p95 interaction latency, achieved FPS, session-memory retention, concurrency limits, GPU-routing behavior, uptime, and interactive-minute cost.
  • For any consequential simulation pilot—especially robotics, medical, or training—place a deterministic rules engine and task-specific validation outside the generative model; do not use visual plausibility as correctness.
  • Prioritize bounded, instrumentable pilots such as previsualization, constrained interactive training, or moderated audience-choice experiences over open-ended autonomous simulations.
  • Assess whether an abstraction platform's named underlying models have stable commercial rights, clear data handling, and fallback paths before making it a core dependency.
  • Monitor progress on long-horizon memory and real-time world-model evaluation; these are the gating capabilities between compelling demos and dependable applications.

Source/Metadata

  • Title: The Next Medium: Why Real-Time Interactive Video Changes Everything — Ahmed Ahres, Reactor
  • Transcript words: 3056
  • Duration seconds: 1050
  • Timestamp note: No usable timestamps or chapter markers were present in the transcript.
Full transcript 2886 words · 13 min read
0:12

Hi, everyone. Welcome to the talk. First of all, thank you all for making the time. I know it's the last talk of the day, probably, or I think the last one is at 3:45, but thank you all for your time. I know you're all probably very busy. Today, I'm going to be talking about something that is a little bit futuristic, though not for San Francisco, and that's world models. I know here it's written real-time interactive video, but the way we think about world models is really in the real-time interactive video, and I'll explain why.

0:39

And in today's world, I think world models is a little bit of a marketing term. Some people think about it from a Gaussian-splitting standpoint, others from video. But the way we define world models is really real-time interactive video, and I have strong evidence or beliefs that this will actually change everything in how we produce and consume content. Oops. What happened? Sorry. Sorry about that. Cool. So if you think about video, video has always been something passive. Before, people used to produce videos and movies, and then we would watch them.

1:19

And in today's world, we have these models, the VO3, the C-Dense 2, and what they do is you prompt them, you get back a file, you watch it, and good luck. It's a slot machine. You cannot change it. You cannot do anything about it. So there's a question that we like to think about in our company: what happens when video becomes programmable like software? And what happens when pixels can be generated in real time? Once this happens, it actually changes completely how we think about consuming content and even producing content.

1:48

And I'll be talking about how, in history, we've seen that real time has always been the future, and the kind of applications that unfolded, and how you can get started today.

2:00

Quick background here. My name is Ahmed. I'm the head of GoToMarket at Reactor. My background is in computer vision and machine learning, and I say machine learning because this is the time when we used to actually train our models. I built and shipped games on iOS and Android for fun. I was a founder, and today I'm the head of GoToMarket at Reactor. And who we are, quickly, is we're a Series A company building the platform for these real-time world models.

2:27

So far, most of them have been still models, but we are building the infrastructure and the developer platform to make them usable so that anyone can integrate these real-time interactive videos, and we can democratize access to this technology. I'll start with a problem. We can today generate pretty much anything, but we just can't change it, right? A generated video, as I mentioned earlier, is still a recording. You get it back, you can't do anything about it. And real time changes what the medium is. It doesn't just make it faster. And I'm going to talk about two examples that actually show us what real time has unlocked in the past.

3:04

Before, in the 1950s and even before, we used to look at a map to know where we are. Someone produces a map, you look at where you are, and that's it. You cannot do anything about it. Then GPS came. GPS made it real time. Suddenly, I can know where I am at the instant that I can track it. Now, you'd think GPS just made it a bit faster to know where I am, but actually, Uber would not exist if we did not have GPS. Another example, which is even a bit more powerful, I'd say: before, we used to use film to produce content, right? We had film, someone shoots something, they can't see what they're shooting, they go somewhere, they produce that film, and then you can see it.

3:44

And then it became digital. You can start seeing what you're shooting. If today you pick up your iPhone and you start recording a video, then you can see what's going on, and that's why we can produce high-quality content. It's because you are able to see what's going on on the screen and adapt accordingly. That gave rise to Instagram and TikTok. Instagram and TikTok would not exist if we could not produce high-quality content, and the only reason why we're able to produce high-quality content, among many reasons, is because we can see in real time what's happening.

4:14

It's not a slot machine. We can actually just edit and see, and that's what unlocks all of these new use cases.

4:23

So when video becomes programmable, you can address it, you can condition it, you can change it. I can show you on the screen whatever I want to show you on the screen, and it becomes programmable a little bit like software, or like anything that is programmable in the world. And in the market today, we are seeing three kinds of models that do this. Some of them you will be familiar with, others you will not be familiar with. The first one is these are Vio. Think about Vio or Sora, but real time and interactive, meaning that, first of all, they're infinite. So they don't stop after 5, 10, or 30 seconds. They actually continue forever.

5:02

They're interactive, meaning you can change what's happening on the screen. And they're in real time, so you don't need to wait to see what's going on. And assuming this works, this is an example of a video that I passed an image with a dog, and this was all generated in real time. And at some point, I'm going to prompt, a cat shows up. And you will see that a cat showed up in the video. This would not be possible in the existing batch regular video generation models, because you would get back the video and you cannot do anything about it. I could have added anything. I could have gone on to create an entire story with it.

5:33

I could have said the dog starts running, starts jumping, a dragon shows up. It goes to, I don't know, to the World Cup. All of this would have happened in front of you. These types of models unlock a few things. First of all, control. So if you think about generative media today, the big problem that content creators all have is, I don't have the control I need. Yes, it's great to use C-Dense to review a three to generate videos, but I just don't have the control. And this is always the thing that any filmmaker or movie producer or any content creator will tell you. And so real time actually ends the slot-machine type of mentality

6:07

and actually gives you the control that you need. And a big thing I like to say is, instant feedback is the ultimate level of control. And we will never be able to have this level of control if we don't have real time. The second thing it unlocks, among other things, which is a field I'm not particularly fond of, but I think is going to be big, is advertising. If I can know what you looked for a minute ago, why can't I insert the logo of whatever you've been looking for? Why can't I produce an ad in real time in front of you? We don't need to pre-produce anything. Now, granted, this is going to take some time because brands are afraid of AI,

6:43

afraid of if their logo has one pixel that is white instead of dark, but it will happen eventually. And I think at the moment that happens, we will not need to produce any ads anymore. Everything will be happening in front of you in real time. The second type of model is the one that probably you're most familiar with. It's the Genie 3-like from Google. These are models where you can pass an image and a text, typically, and you can control a character. They're fun. The first thing that you think about when you think about this is games, right? It's a character, it's a world, you can generate anything. But actually, it goes way beyond games.

7:22

It creates entirely new interactive experiences where, combined with the first types of models that I talked about, we've already been seeing people in our community building a mix of games and movies. If you've ever watched Bandersnatch from Netflix, which is the movie where you can pick your next scene, this is one of those things that becomes possible, that you can control a character, you can control what's happening, and create entirely new types of interactive experiences that were not possible. The second thing is robotics. So because you can simulate and you can control, you can actually create as much training data as you want.

8:00

And robotics, world models in robotics, is actually a gigantinormous market today. I cannot tell you the number of robotics labs that are training and building these models. But because you can control whatever you want to control in any environment, this creates a new opportunity to generate an infinite amount of data for robotics. And finally, something I like to think about, this is more maybe a passionate thing that I have, is education. Because you can step into anything, with today's world in AI, I don't actually believe that the future of education is LLM-based or textbook-based. If you can put any kid in the situation, for example in a history lesson,

8:33

that enables entirely new types of experiences that can be educational. The third type of model is probably the type of model that is more, let's say, something that we've been seeing before, which is avatars, but live and interactive. The thing with avatars, though, is it hasn't actually been cracked. They're still all kind of weird. If you speak to an avatar in any customer support or anything, it's still kind of off, right? And these types of models, and we're seeing a rise of these live and interactive avatar models in research preview, that, combined with model one and model two, is actually going to be, I believe, a big change in what we've been seeing so far.

9:16

And this will be applied to things like customer support, training, sales, gaming, streaming services, etc. And so, just to give you a glimpse of what our users are building today at Reactor with these types of models, some of them are building interactive live streams, right?

9:39

A live stream where people are watching and then the users can type what happens next and then they vote. Why? Because pixels can be generated in real time. So there's no reason why I cannot put a live stream on X, YouTube, or Twitch, and enable users to pick what happens next. Something that was surprising to me is a little bit on the medical simulation. So we've seen users create applications where you generate a world and then you simulate what happens next. What if I put this medicine? What if I remove this medicine, right? And this can be a training playground for people wanting to become doctors.

10:11

The third one, which is also kind of surprising, is cooking simulation. So people are building applications where you can simulate cooking and what happens if you put this ingredient. And finally, video editing. With video-to-video models, video editing becomes very interesting because I'm able to just add visual effects in real time. And we've seen people build entire video editing platforms. Now, granted, they're not very good yet, just because of the quality of the models. But it's a new paradigm when you're able to edit videos just via prompting or by talking to it or by clicking. And now, for the final part, how do you actually do all of this?

10:49

And this is why I like to say the world behind an API.

10:58

At Reactor, we have four types of models today. The first one is called Helios, which is the interactive video model that I talked about. This one is from ByteDance. Linkbot, which is a world model like Genie 3, trained by Alibaba. Long Live 2 from NVIDIA, which is multi-shot film. You can prompt things in advance and create a consistent story over time. And Sound of Streaming, also from NVIDIA, which is a model that does video-to-video editing. So people are already using this, for example, by shooting something or creating something on Seed Dance 2, uploading it, and then adding visual effects, removing people, adding background.

11:34

And this gets very interesting in pre-visualization, for example, for Hollywood movies. And under the hood, when we talk about infrastructure, the thing that I think I like to drive home is building infrastructure for regular video generation models is very different from real time.

11:50

Because in regular video generation models, you're talking about requests. You just send a request, a job gets run in the cloud, and I'm oversimplifying here, but a job runs in the cloud, and it gives you back a file. With real time, it's a different ballgame. You cannot just take what works for batch inference and apply it to real-time inference. For example, you need to think about streaming, right? Once you need to think about streaming pixels from a server to the client, it adds entire complexities that batch generation does not have to think about. The second one is that everything is a live session.

12:30

So everything runs constantly, and there's memory to be kept into account. Now, granted, one of the things that live real-time models struggle with is memory. If you've seen demos from Gini3, for example, we've all seen that the character can look back and then not remember what's going on. So there is a lot of work that needs to go into maintaining that context window so that you can remember what happened if you turned your character left and right. And finally, global scale. If you think about real time, it needs to be sub-100-millisecond latency anywhere you are. And if you're deploying applications in the world, then someone based in India or someone based in Japan

13:07

should be routed to a GPU that is based in India or Japan, or as close as possible to it. If not, if you don't have the compute worldwide, then the experiences are not real time anymore, and it breaks completely the medium. And with Reactor, this is as easy as it gets to integrate these real-time models. Okay, I kind of maybe oversimplified it a little bit here, but it really is maybe 10 lines of code, and we have docs to do all of this. But essentially, you can just load the model with an API key and just start integrating it in whatever video, image, plugin, anything that you're building. And if you want to get started, there's a QR code there.

13:51

And I've added a promo code, AIE2026, which will give you $75 worth of credits, which is a significant amount of compute in our case because we make the models extremely cheap. Thank you very much. I'm happy to take questions if anybody has a question. Nope.

14:19

Right.

14:24

I was told that 16 FPS was a data [inaudible]. What is it that [inaudible]? Well, one thing that actually we do is multi-GPUs. So we use multiple GPUs.

14:39

Optimizing the model weights, applying quantization techniques. So there are ways. It's just a matter of priorities, but there are ways around it. Yeah. Good question, though. Hey. Yeah. I'm wondering if you're looking something on IPC. Sorry?

14:59

Are you aware of what the IPC is?

15:06

No. Okay. International Projects Invention. Okay. I was wondering if you're going to represent some new things in regards of media. I wasn't, but now I'm going to look at it. Yeah. Yeah. How do you feel about the deterministic engines group? Yeah. Yeah. Do you guys experiment with that and share something? When you say deterministic engines, what do you mean exactly? Or any kind of deterministic rule set that we check against maybe so that the simulation stays rather than how we can. Hmm. So we don't do any of that today. The reason why also we build a developer platform, but we've seen people build that on top of us.

15:35

Right. Building the infra for this is already a lot of work. And I think we were already seeing developers build it and then open source it and then we reuse it. But the community is doing it for us, which is even better. Yeah. Is that a question? Yeah. How do you measure the size with it? You're asking a question that the entire research community in world models has not answered. Your evals for real time and consistency and fidelity. Well, fidelity is easy. It's just pixels, right? But evaluation for these real-time models is an unsolved problem. So today it's literally just look at it and human judgment. That's what it is today.

16:10

And this is including, by the way, DeepMind and everything. Nobody has solved this problem yet. We're working on it. We have a research team. Yes. Awesome. Cool. Well, thank you, everyone. Thanks for your time. But so the community is doing it for us, which is even better. Yeah. Is that a question? Yeah. How do you measure the size with it?

16:33

You're asking a question that the entire research community in World Models has not answered. Your evals for real time and consistency is, and fidelity. Well, fidelity is easy. It's just like pixels, right? But evaluation for these real time models is an unsolved problem. So today it's literally just look at it and human judgment. That's what it is today. And this is including, by the way, deep mind and everything. Nobody has solved this problem yet.

17:04

We're working on it. We have a research team. Yes.

17:10

Awesome. Cool. Well, thank you everyone. Thanks for your time.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note