AI Engineer

The Pipeline Is Dead - Iris ten Teije, Sky Valley Ambient Computing

3537 summary words 16 min summary Watch video

Start with the signal

16 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: AI-generated code is so cheap that the traditional CI/CD pipeline—built to freeze and distribute one software artifact to everyone—is dying, replaced by per-user adaptive software that diverges live from a canonical stem.
  • Why it matters: If software modification becomes as cheap as execution, the entire infrastructure stack (CI, registries, containers, app stores) built around frozen artifacts becomes obsolete, fundamentally changing SaaS economics, personalization, and software architecture.
  • Best use: Understand the architectural and economic thesis behind runtime-adaptive software; assess whether Differ's 'stem + divergences' model is viable for agent systems, SaaS personalization, and next-gen software delivery.

Executive Summary

Iris ten Teije, co-founder of Differ (whose technical co-founder Noam was the first engineer at JFrog), argues that the traditional CI/CD pipeline—designed to freeze, version, and distribute one artifact to all users—is dying because the economic constraint it was built for no longer exists. For decades, producing correct software was expensive and rare, so the industry froze artifacts for reproducibility and shipped one version to everyone. The cost of producing a scoped, correct code change is collapsing toward zero with AI agents, meaning software can now be generated cheaply at runtime, on a per-user basis, in the user's live session.

The new architecture Differ proposes is 'adaptive software': deploy one canonical stem, then every user runs their own divergence—same origin, individually adapted live. This shifts from 'best first version for everyone' to 'best version for anyone.' The stem + divergences model is structurally bounded: each divergence is isolated, immutable, inspectable, and individually reversible, so a bad variant cannot corrupt the stem or affect other users. Developers set boundaries (e.g., auth/payments off-limits), and the system adapts within those constraints based on observed user behavior or explicit user requests. The CRM example: an investor using a CRM built for salespeople can have fields auto-rearranged, intro paths added, and irrelevant fields hidden—without R&D.

Iris acknowledges five hard, unsolved challenges: (1) Source of truth—what is 'the software' when it's stem + divergences? Answer: stem + immutable divergences, queried as a graph, not a version number. (2) Correctness—testing the stem and every possible divergence. (3) Desirability—measuring whether a change improved the metric that matters (retention, churn, support tickets). (4) Autonomy vs. control—building trust so the system can act without developer approval every time. (5) Coordination—propagating updates across a million divergent versions by 'merging intent/outcome, not code.' Generation is the easy 80%; observability, validation, and coordination are 'the entire business.'

The talk positions this as an inevitable shift analogous to branchless banking or the adoption of CI itself: the constraint (expensive software production) disappeared, so the infrastructure built around it will too. The pipeline didn't fail; it simply optimized for a constraint that no longer exists. When making software is as cheap as running it, distribution and development are no longer separate phases. The next 20 years, Iris claims, are about shipping the right version to anyone with isolation and provenance that makes it safe.

Key Takeaways

  • Claim: The entire CI/CD pipeline stack exists to solve one problem—get a frozen artifact from build machine to runtime—and that stack assumes 'one version for everyone' because producing software was historically expensive and rare. | Evidence: Iris opens by asking how many in the audience have spent careers on CI pipelines, package registries, container images, app store reviews. She states that the frozen artifact model wasn't chosen in a meeting; it was the only viable option when producing correct changes required skilled humans and took days. Forking the codebase per user was economically infeasible. | Caveat: The talk does not quantify how cheap AI-generated code has become in practice, nor does it cite specific cost-per-change benchmarks. The claim rests on directional trends (AI coding agents) rather than hard ROI data. | Implication: If you accept the premise that AI makes code generation cheap enough to happen at runtime per user, then the entire infrastructure layer (CI, registries, artifact signing, reproducible builds) optimized for frozen artifacts becomes over-engineered legacy. Operators should watch for whether adaptive software architectures (Differ's or others) gain traction, as they would obsolete large parts of the DevOps toolchain and open new opportunities in observability, validation, and per-user provenance. | Timestamp: timestamp unavailable
  • Claim: Per-user adaptive software is not new demand—examples include forward-deployed engineers in enterprise SaaS, dotfiles/editor configs, Excel as a personalized program builder, and the entire professional services industry around customization. The demand has always existed; the supply-side constraint was cost. | Evidence: Iris cites Salesforce professional services and consultants living in client Slack channels, engineers hand-maintaining dotfiles on every machine, and Excel as 'millions of people building their own programs.' She argues social feeds proved per-user content beats one-size-fits-all on every metric; now AI enables the same for software. | Caveat: The examples (Excel, dotfiles, pro services) are all manual customization or user-scripted extensions, not runtime agent-driven code modification. Iris does not provide evidence that users will trust or want autonomous code changes to their software UX, even if bounded. The leap from 'I configure my .vimrc' to 'an agent rewrites my CRM UI' is significant. | Implication: If the analogy holds, adaptive software could unlock SaaS markets that are currently served only by expensive consulting (SMB/mid-market customers who can't afford pro services) or could reduce R&D spend by letting software adapt to personas without hand-coding each variant. Conversely, if trust and UX legibility are insufficient, adaptive software could backfire as users experience 'software that changes under them,' harming retention. Ken should track user acceptance and trust signals if Differ or competitors ship adaptive products. | Timestamp: timestamp unavailable
  • Claim: The objection 'I can barely reason about one AI-generated codebase; you want me to run a million?' is the wrong worry. Brittleness comes from unmanaged divergence inside a single artifact (thousand-line files, no boundaries), not from running per-user variants. Differ's stem + bounded divergences is the opposite: each divergence is isolated, immutable, individually reversible, so blast radius is one context, not the system. | Evidence: Iris recounts a CTO objection on a call. She distinguishes 'unmanaged divergence inside one artifact' (bad) from 'stem + divergences' (bounded, isolated, reversible). A bad variant cannot corrupt the stem or affect another user. Developers can set controls (auth/payments off-limits, specific form fields cannot be dropped). Rollback happens live, no deploy. | Caveat: The talk does not explain how Differ enforces these boundaries at runtime or prevents a divergence from inadvertently calling a stem function in an unsafe way. There is no detail on the isolation mechanism (sandboxing? type contracts? runtime verification?). The claim that 'a bad variant cannot silently corrupt the stem' is architecturally strong but unproven without technical documentation. | Implication: If Differ's isolation and provenance mechanisms are robust, the architecture could be safer than feature flags or A/B testing (which often have global side effects). For Ken, this means adaptive software could enable higher-velocity experimentation without the coordination cost of traditional feature flagging, but only if the isolation guarantees hold. Watch for technical deep-dives or security audits of Differ's runtime. | Timestamp: timestamp unavailable
  • Claim: The five hard unsolved challenges are: (1) source of truth (stem + divergences as graph query, not version number); (2) correctness (testing stem + every divergence); (3) desirability (measuring whether changes improve metrics like retention, churn, support tickets); (4) autonomy vs. control (winning trust so the system acts without developer approval); (5) coordination (propagating updates by merging intent/outcome, not code). | Evidence: Iris explicitly walks through each challenge. For (1), she says every divergence is immutable, inspectable, attributable, and traceable back to the signal/recommendation/adaptation. For (2), testing at scale means reasoning about stem + divergences. For (3), goals differ per software (retention, churn, support tickets). For (4), the vision is not recommendations but autonomous action. For (5), the answer is 'merge intent/outcome' so everyone converges on the same goal through their own path. | Caveat: Iris admits 'we haven't solved all of them' and that these are 'what we're working on day in, day out.' No evidence is provided that any of these challenges have been solved at scale in production. The 'merge intent' coordination model is hand-waved, not explained. The talk is a vision pitch, not a post-deployment case study. | Implication: For Ken, these five challenges define the infrastructure moat and business risk. If Differ solves them, they own a new category ('adaptive software substrate'). If they don't, the entire model collapses into ungovernable divergence and user distrust. The 'generation is easy 80%' claim is important: calling an LLM is trivial, but observability/validation/coordination is the entire business. Ken should track Differ's progress on provenance and validation tooling, as that is where the defensibility lies. | Timestamp: timestamp unavailable
  • Claim: The historical analogy is branchless banking or the adoption of CI: shifts that felt reckless until they didn't. The pipeline didn't fail; the constraint it was built for (expensive software production) went away. When making software is as cheap as running it, distribution and development are no longer separate phases. | Evidence: Iris's fintech origin story: 'a bank with no branches sounded reckless; a decade later, the branch is the weird part.' Co-founder Noam 'used to fight with engineers about why you need builds and CI.' The claim is that adaptive software is the same type of paradigm shift. | Caveat: The analogy glosses over adoption timelines and regulatory/trust barriers. Branchless banking took a decade+ and required regulatory changes and consumer trust-building. Adaptive software rewrites user-facing code at runtime, which may face much higher trust/security scrutiny than digital banking. The talk provides no roadmap for adoption (greenfield SaaS only? enterprise later? regulated industries never?). | Implication: If the analogy holds, adaptive software could follow a similar trajectory: dismissed by incumbents, adopted by greenfield startups, then inevitable. Ken should monitor whether Differ gains traction with modern SaaS companies (CRM, internal tools) or if adoption stalls due to trust/security/compliance objections. The 'next 20 years' framing suggests Iris expects a long adoption curve, not immediate disruption. | Timestamp: timestamp unavailable

Detailed Brief

The frozen artifact assumption and its economic origin

  • Claims: The CI/CD pipeline exists to move frozen, dead code from build machine to runtime once, safely, reproducibly.; The entire stack—CI pipelines, package registries, container images, app store reviews—encodes one assumption: one version for everyone.; This assumption was never a choice; it was the only viable option because producing correct software was expensive (skilled humans, hours/days).; The frozen artifact gave reproducibility, reviewability, rollback, but the price was that software couldn't be for anyone in particular.; There was never a meeting where 'one version for everyone' beat 'a version per user'; the second wasn't an option due to cost.
  • Evidence: Iris asks the audience how many spent careers on CI/pipelines, registries, containers, app store reviews.; She states that giving two users different software meant forking the codebase and hand-maintaining both, which was infeasible at scale.; The one-way pipeline was 'the direct consequence of production of software being expensive and risky.'; The frozen artifact was treated as 'a fact about software, like gravity,' but it was 'a fact about cost and budget.'
  • Caveats: No quantitative data on the historical cost of software production vs. today's AI-generated code cost.; The talk does not address how the frozen artifact model handles security patching, compliance, or regulatory auditability—areas where immutability is a feature, not a bug.; The claim that 'software couldn't really be for anyone in particular' ignores decades of configuration, plugins, and extensibility (though Iris does acknowledge feature flags and A/B testing as predecessors).
  • Implications: If AI makes code generation cheap enough, the entire CI/CD toolchain is over-optimized for a constraint that no longer exists, opening the door for new infrastructure (Differ's 'adaptive software substrate').; Operators should evaluate whether their current deployment stack (artifact signing, hermetic builds, reproducible containers) is solving a problem that might not exist in 5 years.; For investing, this suggests a new category: runtime observability and validation for adaptive software, which Iris claims is 'the entire business.'

The stem + divergences architecture and CRM example

  • Claims: Differ's model: deploy one canonical stem, every user runs their own divergence—same origin, individually adapted live.; Divergences are bounded, isolated, individually reversible. A bad variant cannot corrupt the stem or reach another user.; Blast radius of a change is one context; any divergence can roll back live with no deploy.; Developers set controls: specific fields cannot be dropped, auth/payments off-limits.; Users can proactively request changes; if within developer-set boundaries and 'within the spirit of the software,' they are implemented without going back to the developer.
  • Evidence: CRM example: an investor logs founder intros and who intro'd her to which deal. System creates 'intro path' field.; Investor always skips specific fields; system learns and stops surfacing them, surfaces fields she cares about more.; Investor checks specific deal types; system reorders prioritization based on her behavior.; Horizontal SaaS like CRM can address wider customer personas without increasing R&D spend.
  • Caveats: No explanation of how Differ enforces boundaries at runtime (sandboxing? contracts? runtime verification?).; No detail on how the system determines 'within the spirit of the software' or validates user requests.; The CRM example is aspirational; no evidence that Differ has shipped this in production.; The claim that 'any divergence can roll back live' implies storing all prior states, but no detail on storage/perf implications.
  • Implications: If isolation guarantees hold, adaptive software could enable SaaS companies to serve long-tail personas (SMB, mid-market) that currently require expensive pro services.; For Ken, this model could apply to agent tooling: every agent run could diverge the tool based on task context, then revert.; Risk: if boundaries are porous or users distrust autonomous changes, adaptive software could harm UX consistency and increase support burden.

The five hard challenges and Differ's point of view

  • Claims: Source of truth: software is stem + all immutable divergences. 'What is this user running?' becomes a graph query, not a version number.; Correctness: testing at scale means reasoning about stem + every possible divergence. Must ensure UI is correct.; Desirability: a code change can be correct but not desirable. Must measure whether it improved the metric that matters (retention, churn, support tickets). Goals differ per software.; Autonomy vs. control: the vision is not recommendations but autonomous action. Challenge is winning trust so humans in the loop choose to step back.; Coordination: propagating updates across a million divergent versions. Answer: merge intent/outcome, not code. Everyone converges on the same goal through their own path.
  • Evidence: For source of truth: 'every divergence is immutable, inspectable, attributable. We trace any version back to the signal, recommendation, and adaptation.'; For desirability: 'the hard part is knowing whether you actually found an improvement, an uplift.'; For autonomy: 'the challenge isn't building more control, it's winning enough trust that you don't have to.'; For coordination: 'not everyone has to run the same commit or exact piece of code, but everyone converges on the same goal through their own path.'
  • Caveats: Iris admits 'we haven't solved all of them' and that these are 'what we're working on day in, day out.'; No evidence that any of these challenges have been solved at scale in production.; The 'merge intent/outcome' model is vague; no explanation of how intent is represented, how conflicts are resolved, or how you prove convergence.; Testing 'every possible divergence' is combinatorially explosive; no detail on sampling strategies or formal verification.
  • Implications: For Ken, these five challenges define the infrastructure moat. If Differ solves them, they own a new category. If not, the model collapses.; The 'generation is easy 80%' claim is critical: calling an LLM is commoditized, but observability/validation/coordination is the defensible layer.; Operators should track whether Differ ships provenance tooling, divergence inspectors, or desirability dashboards—those are the moat.; For investing, the hard parts (validation, coordination) are where new infrastructure companies could emerge if Differ doesn't solve them.

Historical analogy and adoption timeline

  • Claims: Adaptive software is a paradigm shift analogous to branchless banking or the adoption of CI: it doesn't feel obvious until it does.; The pipeline didn't fail because it didn't work; the constraint it was built for (expensive software production) went away.; When making software is as cheap as running it, distribution and development are no longer separate phases.; The next 20 years are about shipping the right version to anyone with isolation and provenance that makes it safe.
  • Evidence: Iris's fintech origin: 'a bank with no branches sounded reckless; a decade later, the branch is the weird part.'; Co-founder Noam 'used to fight with engineers about why you need builds and CI.'; The claim is that we've 'seen these shifts before' and adaptive software is the same type.
  • Caveats: Branchless banking took a decade+ and required regulatory changes, consumer trust-building, and mobile ubiquity. Adaptive software may face similar or higher barriers (runtime code modification, security, compliance).; The talk provides no roadmap for adoption: greenfield SaaS first? Enterprise later? Regulated industries (fintech, healthcare) excluded?; The 'next 20 years' framing suggests a long adoption curve, not immediate disruption, which undercuts the 'pipeline is dead' headline.
  • Implications: If the analogy holds, adaptive software could follow a similar trajectory: dismissed by incumbents, adopted by greenfield startups, then inevitable.; Ken should monitor whether Differ gains traction with modern SaaS companies or if adoption stalls due to trust/security/compliance objections.; For operators, the question is whether to bet on adaptive architecture now or wait until it's proven. The 20-year timeline suggests patience.; For investing, early-stage companies building adaptive software infrastructure (observability, validation, coordination) could be high-risk, high-reward bets.

Notable Concepts & Terms

  • Stem + divergences: Differ's architecture: one canonical codebase (stem) deployed, with each user running an isolated, bounded, individually reversible modification (divergence). The claim is that this is safer than a monolithic artifact with unmanaged internal divergence.
  • Adaptive software: Software that modifies itself at runtime based on user behavior or explicit requests, within developer-set boundaries. Contrast to frozen artifacts and one-size-fits-all releases.
  • Merge intent, not code: Differ's answer to coordination: when propagating updates across a million divergent versions, you don't merge code commits; you merge the desired outcome or intent, so every user converges on the same goal through their own path. Vague in the talk but central to the architecture.
  • Blast radius of one context: In the stem + divergences model, a bad variant affects only the user running it, not the system or other users. Contrast to traditional deployments where a bad release affects all users.
  • Forward-deployed engineer: Enterprise SaaS term for consultants or engineers embedded in a client's environment to customize software. Iris cites this as proof of demand for per-user software; adaptive software could automate this role.
  • Desirability vs. correctness: A code change can be syntactically correct and functionally working but still undesirable (e.g., hurts retention, increases churn). Differ must measure both. Desirability is tied to company-specific goals (retention, churn, support tickets).
  • Provenance and attribution: In adaptive software, every divergence must be immutable, inspectable, and traceable back to the signal/recommendation/adaptation that created it. This is Differ's answer to 'what is this user running and why?'

Operator Notes / Why Ken Should Care

  • For agent systems: The stem + divergences model could apply to AI agents—each agent run could diverge the tool/prompt/workflow based on task context, then revert. If Differ's isolation guarantees hold, this enables per-task adaptation without coordination overhead.
  • For SaaS/GTM: If adaptive software works, it could unlock long-tail customer personas (SMB, mid-market) that currently require expensive pro services. SaaS companies could ship one product and let it adapt, reducing R&D spend and increasing TAM. Watch whether Differ or competitors ship adaptive CRM, ERP, or internal tools.
  • For investing: The 'generation is easy 80%' claim is critical—calling an LLM is commoditized, but observability/validation/coordination is the defensible layer. Early-stage companies building infrastructure for adaptive software (provenance, desirability measurement, coordination) could be high-upside bets if the category takes off. Risk: Differ admits the hard parts are unsolved.
  • For content/workflow: If adaptive software becomes real, the UX paradigm shifts from 'I control the software' to 'the software adapts to me.' This could improve retention (software fits the user) or harm it (software changes under the user, breaking mental models). Watch for user trust and legibility signals.
  • For AI ops: The coordination challenge ('merge intent, not code') is the hardest and least explained. If Differ solves it, they own a new infrastructure primitive. If not, adaptive software could become ungovernable divergence. Operators should track whether Differ ships a coordination layer or if that becomes a separate product category.

Watch Map

  • timestamp unavailable: Timestamps were not present in the transcript. The talk is structured as: (1) opening poll on CI/pipeline experience, (2) historical context (frozen artifacts, one version for everyone), (3) why the constraint (expensive production) is gone, (4) stem + divergences architecture, (5) CRM example, (6) five hard challenges, (7) analogy to branchless banking and closing.

Source/Metadata

  • Title: The Pipeline Is Dead - Iris ten Teije, Sky Valley Ambient Computing
  • Transcript words: 3997
  • Duration seconds: 1189
  • Timestamp note: Timestamps were not available in the provided transcript.
Full transcript 2687 words · 18 min read
0:00

SPEAKER_00

As we're online today, you can't raise your hands, but just not on your screen. How many of you have spent a significant part of your career making software move from one computer to another? CI pipelines, package registries, container images, app store reviews. That entire stack exists to solve exactly one problem. Get a frozen artifact, dead code, from the machine where it was built to the machine where it runs. It's safely, reproducibly, once. And here's what I think everyone is missing. That entire stack is built around one idea that is so old that we stopped seeing it as a choice. One version of your software for everyone.

0:35

SPEAKER_00

We've shipped it that way for so long that almost nobody asks why anymore. I'm Iris. I'm one of the co-founders of Differ. My co-founder, Noam, was the first engineer at JFrog. He helped build the pipeline that I'm about to tell you is dying. I came at it from the other end. I spent a decade in fintech. I was early at a digital bank that we scaled and exited. Shipping software in the environments, least willing to tolerate what I'm about to propose. So I totally get any reservations and I've had them myself as well. However, I'm telling you, it's coming anyway. And that's not a warning. It's the best thing that's happened to software in a long time.

1:14

SPEAKER_00

Every piece of distribution infrastructure that you've touched encodes the same assumption. Software is produced in one place. It runs in another. And the thing in the middle, the artifact, is frozen. And that assumption was correct. For decades, it was just true. Why? Because producing a correct change was expensive. It took skilled humans hours or days. So you did it rarely. It was a central event. You verified it and froze it. And you shipped that frozen thing to everyone. So the one-way pipeline isn't arbitrary. It's the direct consequence of production of software being expensive and risky. And of course, the frozen artifact has some advantages.

2:00

SPEAKER_00

You get reproducibility, reviewability, rollback. And every guarantee that we lean on in production flows from one fact. There is one artifact and it doesn't change after we ship it. That's the deal. One version for everyone, frozen. We got reliability. But the price was that the software couldn't really be for anyone in particular. And nobody ever made that decision. There was never a meeting where someone put one version for everyone versus a version for each person and picked the first one. And it wasn't because the second option, a version for each person, was worse. It just wasn't an option.

2:41

SPEAKER_00

Giving two users different software meant forking the code base and hand-maintaining both. A version per user at any real scale wasn't really a viable option. And one version where everyone had to win an argument. It was just how it was. It was the only shape software could take. And we started treating it as a fact about software, in gravity. Like something that is just true. But it was never really a fact about software. It was a fact about cost and budget. And the economics of that cost just changed. Now, you might expect me to go into AI can code. But everyone has already said it. And that's stale stakes. And to me, that's not really the interesting part.

3:34

SPEAKER_00

The interesting part is where and how cheap. The cost of producing a correct and scoped change is collapsing towards zero. And just as importantly, the production of the software no longer has to happen in one place up front before anyone runs it. Part of it can be run on the server. Part of it on the client. Part in the user's live session. And as each step stops being a decision, you freeze at build time. And now starts becoming more of a real-time one. That also means that each piece can be placed wherever it makes most sense. Including right in front of the user in their context.

4:09

SPEAKER_00

And so this whole one-way pipeline existed because making software was the expensive, central, and rare event. And running it was cheap. So you separated development from distribution. But now, as making a change becomes as cheap as running one, and it can happen in the same place as where you run it, the reason to separate them is dissolving. So far, I've talked about the supply side. Producing code has become cheap, easy. We can now make a change on a per-user basis. But I also want to address the demand side. People have always wanted software that fits them. And we have decades of proof for that.

4:44

SPEAKER_00

It just wasn't really possible for most types of software because of cost. To start with one example, the forward-deployed engineer. Enterprise software has always had a line item called professional services. And a whole industry exists around that. If you're a big client of a company like Salesforce, you probably have consultants that are helping you implement custom setups, configurations. You have an engineer living in your Slack channel. And it's not that smaller customers can benefit from this type of customization. It just didn't make sense financially until today.

5:14

SPEAKER_00

To give another example that might resonate with you as engineers, think about your dotfiles, your editor config, key bindings. You rebuild every tool you touch into your tool by hand on every machine. It's another example of a demand for personalized software. And lastly, Excel, the most successful business software ever created. And Excel isn't really a static program. It's millions of people that all built their own programs on top of it. So as it's clear, give people the power to make their software theirs, and they take it. It's seen on the social feed that per user wins from one size fits all on every metric that mattered.

5:48

SPEAKER_00

This is more of an example of content versus software. But now that we have better coding, now that we have coding agents and better coding agents, we can move this also to the software layer. So my point is, none of this is really new demand. We've seen it for decades. There have been predecessors feature flags, segmentation, A-B testing. So as it's clear, give people the power to make their software theirs, and they take it. It's seen on the social feed that per user wins from one size fits all on every metric that mattered. This is more of an example of content versus software.

6:30

SPEAKER_00

But now that we have better coding, now that we have coding agents and better coding agents, we can move this also to the software layer. So my point is, none of this is really new demand. We've seen it for decades. There have been predecessors like feature flags, segmentation, A-B testing. There's an enormous industry around this. And we've been trying to make software diverge for many years. But we forced into a specific shape, that of creating buckets and segments that you declare in advance. And now, for the first time, we can make software truly adaptive.

7:04

SPEAKER_00

So to get back to the title of this talk, when the agent is to run time, when the thing that runs your software can also modify it, development and distribution stop being two phases. The boundary blurs and it's gone. And the shape that we bet on at Differ is that instead of one code base gated by flags and shipped to everyone, you deploy one canonical stem, and every user runs their own divergence of it. Same origin, but individually adapted live. It's going from the best first version for everyone to the best version for anyone. Now, if you're an infrastructure person, you might be a little worried.

7:19

SPEAKER_00

Your stomach might be turning because I have just deleted the frozen artifact. And the frozen artifact was what was holding up the entire building. And we do get these objections. For example, from a call that I had with a CTO recently, who is right to be skeptical. What he said roughly was, I can already barely reason about one AI-generated code base. And you want me to run a million of these? You're not describing a capability. You're describing my worst problem, multiplied. And if that's your reaction, that's not surprising. It's the right instinct, but perhaps aimed at the wrong target. Here's the distinction that we make.

8:05

SPEAKER_00

The brittleness that you are picturing is a specific type of failure mode. It's an unmanaged divergence inside a single artifact. Thousand line files, everything can touch everything else, no boundaries. And that's not necessarily brittle because it's AI-generated. It's brittle because there's no structure separating things. In our vision, we're thinking about per-user divergences.

8:31

SPEAKER_00

And done right, that is the opposite of that. You've got a stem plus divergences. The divergences are bounded, isolated, and individually reversible. A bad variant can silently corrupt the stem or reach another user. Which means that the blast radius of a change isn't a system, it's one context. And any single divergence can roll back live with no deploy. So the answer to the previous objection isn't trust us, AI is good at coordination or at coding. The honest answer is you're brittle because there's a tangled artifact with no boundaries. In our case, we don't ship you a thousand tangled artifacts. We ship one stem and bounded divergences. Each isolated, each reversible.

9:00

SPEAKER_00

The thing that you're afraid of is the thing that this architecture exists to prevent. And as a developer, you can also set controls and boundaries. What can and cannot be adapted. To give a small example, we can have a scenario where we've got a form and the form can be adapted in order to improve conversion rate, for example. However, as a developer, you can always indicate that specific fields can never be dropped. Or parts of your app, like auth or payments, should always be off limits for any sort of adaptation. To give another example, to make adaptive software a bit more concrete, think of, for example, a CRM.

9:17

SPEAKER_00

In this case, we've got an investor who's using a CRM, whereas the CRM was mostly built with a salesperson in mind. And as an investor, you might use it slightly differently. So this investor, she often logs founder intros and she's always logging who intro'd her to which deal. So the system observes that and creates a intro path. The system can also observe that she's always skipping specific fields. She never fills them out. So over time, the system learns and doesn't surface those fields. But instead, it surfaces fields that she cares about more.

9:43

SPEAKER_00

Another example can be the fact that she is always checking specific types of deals or founders, and it doesn't exactly follow the prioritization that the system sets by default. So again, can we learn and make that smarter so that the information that the user cares about is surfaced first? Not only can the system observe, the idea is also that the user can proactively request changes. And as long as these changes are within the boundaries that the developer sets, and as long as they are within the spirit of the software, within the purpose of what the software was originally made for, it can be implemented without having to go back to the developer.

9:51

SPEAKER_00

So you can imagine for a horizontal SaaS like a CRM, you can address a much wider number of customer personas without increasing your R&D spend. Now, of course, there are many hard parts when it comes to bringing this to life. And I'm going to discuss a couple of hard challenges and problems that we're working on as we're making adaptive software real at scale. We haven't solved all of them, but we have a point of view on each of them, and it's exactly what we're working on day in, day out. So what changes when there's no single artifact? Firstly, the source of truth. So when there's no single artifact, you can wonder what is the software?

10:12

SPEAKER_00

In our case, we consider the software to be the stem, plus all the immutable divergences. But of course, that creates a linear problem. What is this user running and why? That now becomes more of a graph query versus a version number. A bug report describes a program that exists for that specific user. We haven't solved all of them, but we have a point of view on each of them, and it's exactly what we're working on day in, day out. So what changes when there's no single artifact?

10:37

SPEAKER_00

Firstly, the source of truth. So when there's no single artifact, you can wonder what is the software? In our case, we consider the software to be the stem, plus all the immutable divergences. But of course, that creates a linear problem. What is this user running and why? That now becomes more of a graph query versus a version number.

10:41

SPEAKER_00

A bug report describes a program that exists for that specific user. How do you debug or inspect that? The answer is every divergence is immutable, inspectable, attributable. And we also need to trace any version back to this signal, led to this specific recommendation, and this exact adaptation. That's one of the parts that we're working on. Second is correctness. How do you test that the code change that you've just implemented works, that the UI is correct, and testing at a much larger scale of users means that you need to reason about the stem, and also every possible divergence of it.

10:46

SPEAKER_00

And then desirability, because perhaps you made a code change, it's perfectly correct and working. But you also need to know that whether this was desirable, was it actually a good change? Because anyone can make a code change now, but the hard part is knowing whether you actually found an improvement, an uplift. And that is something that is extremely important to keep track of and to measure and to consider what the goals for the company are. And this is not going to be the same for every single piece of software.

10:49

SPEAKER_00

In some cases, it might be retention, or less churn, or lowering this number of support tickets. So it's extremely important to keep track of what are the goals that we're chasing of adaptive software, and do the adaptations reach to improve the metrics that matter? Next is autonomy versus control. And the conservative answer here would be start with just recommendations, and don't make any autonomous changes. And that's not a wrong strategy, but in our case, it's not really our vision. The vision is a system that understands the user well enough to act without asking the developer for permission every time first.

10:56

SPEAKER_00

So for us, the challenge isn't building more control, it's winning enough trust that you don't have to. And it's a hard problem, but also a very interesting one. And how do you make a system good enough, legible enough, reliable enough, that humans in a loop choose to step back? And that's certainly what we are building towards.

10:58

SPEAKER_00

And lastly, coordination. Everyone on their own version. How do you push new updates? How do changes propagate to a million different versions? And this is one of the challenges that we've been thinking hardest about. And the answer that we keep coming back to is don't merge code, merge intent, merge outcome, which means not everyone has to run the same commit or the exact same piece of code, but everyone converges on the same goal through their own path.

11:02

SPEAKER_00

And if it wasn't already clear, the challenges that I just addressed are the hard challenges. Generation has become easy. And I would say that's actually the easy 80%. Calling a model to write some code is something that everyone can do. The other part, observability, validation, coordination, that is the entire business. Anyone can call an LLM, but the substrate, the stem plus divergences, provenance, validation, that is really the hard part and something that we're working on every single day.

11:06

SPEAKER_00

To close off with, when I started out in fintech, a bank with no branches sounded reckless. A decade later, the branch is the weird part. My co-founder, Noam, used to fight with engineers about why you need builds and CI. We've seen these shifts before and adaptive software is the same type of shift. It doesn't feel obvious until it does.

11:11

SPEAKER_00

To go back to the start of the talk, the pipeline didn't fail because it didn't work anymore, but the constraint it was built for went away. The assumption underneath it that software is expensive, so we need to freeze it and ship it once, that assumption stopped being true. And when making software gets as cheap as running it, the line between distribution and development isn't a line anymore. We spent 20 years getting good at shipping one version for everyone. The next 20 are about shipping the right version to anyone with the isolation and provenance that makes it safe instead of terrifying. I'm Iris, this is Diffra, and that's what we're building.

11:20

SPEAKER_00

Thank you for watching. In this case, we've got an investor who's using a CRM, whereas the CRM was mostly built with a salesperson in mind. And as an investor, you might use it slightly differently. So this investor, she often logs founder intros and she's always logging who intro'd her to which deal. So the system observes that and creates a intro path. The system can also observe that she's always skipping specific fields. She never fills them out. So over time, the system learns and doesn't surface those fields. But instead, it surfaces fields that she cares about more.

12:09

SPEAKER_00

Another example can be the fact that she is always checking specific types of deals or founders, and it doesn't exactly follow the prioritization that the system sets by default. So again, can we learn and make that smarter so that the information that the user cares about is surface first? Not only can the system observe, the idea is also that the user can proactively request changes. And as long as these changes are within the boundaries that the developer sets, and as long as they are within the spirit of the software, within the purpose of what the software was originally made for, it can be implemented without having to go back to the developer.

12:57

SPEAKER_00

So you can imagine for a horizontal SaaS like a CRM, you can address a much wider number of customer personas without increasing your R&D spend. Now, of course, there are many hard parts when it comes to bringing this to life. And I'm going to discuss a couple of hard challenges and problems that we're working on as we're making adaptive software real at scale. We haven't solved all of them, but we have a point of view on each of them, and it's exactly what we're working on day in, day out. So what changes when there's no single artifact?

13:45

SPEAKER_00

Firstly, the source of truth. So when there's no single artifact, you can wonder what is the software? In our case, we consider the software to be the stem, plus all the immutable divergences. But of course, that creates a linear problem. What is this user running and why? That now becomes more of a graph query versus a version number. A bug report describes a program that exists for that specific user. How do you debug or inspect that? The answer is every divergence is immutable, inspectable, attributable. And we also need to trace any version back to this signal, led to this specific recommendation, and this exact adaptation.

14:39

SPEAKER_00

That's one of the parts that we're working on. Second is correctness. Like, how do you test that the code change that you've just implemented works, that the UI is correct, and yeah, testing at a much larger scale of users means that you need to reason about the stem, and also every possible divergence of it.

15:08

SPEAKER_00

And then desirability, because perhaps you made a code change, it's perfectly correct and working. But you also need to know that whether this was desirable, was it actually a good change? Because anyone can make a code change now, but the hard part is knowing whether you actually found an improvement, an uplift. And yeah, that is something that is extremely important to keep track of and to measure and to consider what the goals for the company are. And this is not going to be the same for every single piece of software. In some cases, it might be retention, or less churn, or lowering this number of support tickets. So it's extremely important to keep track of

15:59

SPEAKER_00

what are the goals that we're chasing of adaptive software, and do the adaptations reach to improve the metrics that matter? Next is autonomy versus control. And the conservative answer here would be start with just recommendations, and don't make any autonomous changes. And that's not a wrong strategy, but in our case, it's not really our vision. The vision is a system that understands the user well enough to act without asking the developer for permission every time first. So for us, the challenge isn't building more control, it's winning enough trust that you don't have to. And it's a hard problem, but also a very interesting one.

16:49

SPEAKER_00

And how do you make a system good enough, legible enough, reliable enough, that humans in a loop choose to step back? And that's certainly what we are building towards.

17:06

SPEAKER_00

And lastly, coordination. Everyone on their own version. How do you push new updates? How do changes propagate to like a million different versions? And this is one of the challenges that we've been thinking hardest about. And the answer that we keep coming back to is don't merge code, merge intent, merge outcome, which means not everyone has to run the same commit or the exact same piece of code, but everyone converges on the same goal through their own path. And if it wasn't already clear, the challenges that I just addressed are the hard challenges. Generation has become easy. And I would say that's actually the easy 80%. Calling a model to write some code

17:53

SPEAKER_00

is something that everyone can do. The other part, observability, validation, coordination, that is the entire business.

18:06

SPEAKER_00

Anyone can call an LLM, but the substrate, the stem plus divergences, provenance, validation, that is really the hard part and something that we're working on every single day.

18:24

SPEAKER_00

To close off with, when I started out in fintech, a bank with no branches sounded reckless. A decade later, the branch is the weird part. My co-founder, Noam, used to fight with engineers about why you need builds and CI. We've seen these shifts before and adaptive software is the same type of shift. It doesn't feel obvious until it does. To go back to the start of the talk, the pipeline didn't fail because it didn't work anymore, but the constraint it was built for went away. The assumption underneath it that software is expensive, so we need to freeze it and ship it once, that assumption stopped being true. And when making software gets as cheap as running it,

19:23

SPEAKER_00

the line between distribution and development isn't a line anymore.

19:29

SPEAKER_00

We spent 20 years getting good at shipping one version for everyone. The next 20 are about shipping the right version to anyone with the isolation and provenance that makes it safe instead of terrifying. I'm Iris, this is Diffra and that's what we're building. Thank you for watching.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note