AI Engineer

2026 State of AI Engineering — Barr Yaron, Amplify Partners

1975 summary words 9 min summary Watch video

Start with the signal

9 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: AI engineering is moving from experimentation to operational deployment: teams are running multi-model, cost-constrained systems with increasingly write-enabled agents, while governance, evaluation, and long-term code quality lag behind adoption.
  • Why it matters: The survey offers a useful market-level read on the same issues central to agent systems: model routing, inference economics, write permissions, control planes, evals, and the build-versus-buy boundary.
  • Best use: Use it as a strategic benchmark for product and investment theses, then use the full report's charts to validate which infrastructure and agent-control gaps are persistent rather than anecdotal.

Executive Summary

Amplify Partners' 2026 survey of 1,048 AI practitioners argues that AI engineering is now a cross-functional discipline rather than a narrow job title. The sample is senior in conventional software engineering but comparatively new to AI: among respondents with more than 10 years of software experience, over half have three years or less of AI experience. The presentation is fundamentally a directional market survey, not a technical implementation guide, but it identifies where production practices are consolidating and where they remain immature.

The production model stack is decisively hybrid. Closed models are used by 94% of respondents and open-weight models by 45%, but more than 90% of open-weight users also use closed models. Eighty-seven percent of teams use multiple models, usually routing by task type, output comparison, or cost. Teams are standardizing the surrounding tool platform while retaining model flexibility; quality, agentic capabilities such as tool calling, and cost matter more in model selection than open-versus-closed status.

The strongest operational message is that agents have crossed from drafting into action. Ninety-five percent of surveyed teams report using agents, and 89% of agent builders say their agents have write access, versus 52% last year. Yet control remains primitive: human approval and permission gates lead, while teams experiment inconsistently with decomposition, retrieval, memory, and sandboxing. The dominant failure complaints are hallucination and lost context, not infrastructure plumbing, while evals remain the leading stack challenge.

AI is viewed as overwhelmingly positive because it reduces the cost of experimentation rather than merely accelerating coding. But the productivity gain produces a maintenance and review burden: 59% worry that current AI-generated code creates long-term liabilities, and more than nine in ten report some negative downstream effect. AI is also redistributing software production across functions, with over one-third of teams having non-developers ship features and 17% reporting regular shipment of customer-facing functionality by non-developers.

Key Takeaways

  • Claim: Production AI is a multi-model routing problem, not a winner-take-all choice between open and closed models. | Evidence: 94% of respondents use closed models, 45% use open-weight models, and over 90% of open-weight users also use closed models. Separately, 87% of teams use more than one model; routing by task type is the most common selection method, with output comparison and cost routing also used. | Implication: Ken should design for model portability and task-level routing rather than make a long-lived platform assumption around one provider or one open-weight model. | Caveat: This is a builder-heavy survey sample, so the levels of adoption may not represent the broader enterprise market.
  • Claim: Teams are standardizing the AI platform and tool layers while deliberately preserving flexibility at the model layer. | Evidence: More than half of respondents said their organization is beginning to standardize on fewer AI tools, even as 87% use multiple models. The speaker characterizes this as an early 'great standardization' of platforms and tools rather than models. | Implication: The defensible control plane is likely to be the abstraction around model access, routing, observability, permissions, and workflow operations—not ownership of a single model endpoint.
  • Claim: Cost has become a first-class product and engineering constraint, monitored nearly as seriously as output quality. | Evidence: 40% say cost regularly limits how ambitiously they use AI and another 36% say it sometimes does. Cost and token usage rank as the second-most monitored production concern, behind quality; cost also ties with agentic capabilities as a leading model-choice criterion after quality. | Implication: Agent and workflow architectures need explicit token budgets, unit economics, routing policies, and cost observability from the outset rather than treating optimization as a post-launch exercise.
  • Claim: The material change in agents is expanded write authority, but the industry has not settled on a robust control layer. | Evidence: 95% report using agents, roughly double the prior year; among agent builders, 89% say agents can write data, up from 52% last year. Combining adoption and permissions, write-enabled-agent use across the full sample increased more than threefold. The most common controls are human-in-the-loop approval and permission gating, with approaches below that fragmented across task decomposition, retrieval, memory, and sandboxing. | Implication: Write-capable agents should be treated as governed operators: scoped credentials, action-level policy enforcement, approval thresholds, audit trails, reversibility, and containment matter more than generic chatbot guardrails. | Caveat: The speaker explicitly notes that the 95% agent-usage figure seems high and that these are survey results, not independently measured deployments.
  • Claim: Agent reliability is primarily a reasoning-and-context problem, while evaluation remains the most persistent infrastructure gap. | Evidence: About two-thirds of respondents cite hallucination or losing context during a task as their greatest agent frustration. Evals again rank as the number-one stack challenge, though narrowly; the most common evaluation method remains informal 'vibe review.' | Implication: For agent systems, invest in task-specific evaluation harnesses, state and memory design, trace review, and regression testing; generic model reliability alone will not solve multi-step execution failures. | Caveat: The survey does not quantify what share of teams have rigorous automated eval suites or distinguish failure rates by workflow type.
  • Claim: The build-versus-buy boundary is settling around buying commoditized infrastructure and retaining product-specific AI logic in-house. | Evidence: Inference and model serving are the most commonly purchased stack layer. By contrast, 61% build prompt management themselves, and prompts, RAG, and evals tend to remain in-house. Fine-tuning is the clearest 'not yet' category, with most respondents not using it. | Implication: Buy undifferentiated serving capacity where possible, but retain ownership of prompts, retrieval behavior, evaluations, and workflow logic that encode the product's specific judgment and operating model.
  • Claim: AI's primary organizational effect is cheaper experimentation, accompanied by real technical-debt and review risks as non-engineers ship more software. | Evidence: 97% report a net positive organizational effect, led by more prototypes, experimentation, and bets rather than simply speed. At the same time, over 90% report some negative downstream effect; common concerns are erosion of deep technical skills and codebase understanding. 81% say AI blurs engineering, product, design, and marketing roles; over one-third report non-developers shipping features, including 17% that say this happens regularly for customer-facing features. | Implication: Organizations should preserve rapid prototyping while separating prototype authority from production authority through review standards, ownership rules, testing gates, and maintainability accountability. | Caveat: The non-developer shipping activity is described as mostly occurring in smaller teams and internal products.

Detailed Brief

Modalities and adoption signals

  • Claims: Text remains the dominant modality in work applications, but audio has the strongest near-term adoption signal.; Image generation appears to have crossed from novelty into practical work use as product quality improved.
  • Evidence: Among people not currently building with audio, 56% say they intend to adopt it, up from 37% in the prior year's survey.; The share using image generation and feeling positive about it doubled from 18% to 36%.; The speaker attributes improved image adoption to recent model and product releases, naming Nano Banana, Nano Banana 2, and ChatGPT Images 2.0.
  • Caveats: Intent-to-adopt is not the same as deployed usage or a proven commercial use case.; The presentation does not provide modality-specific revenue, retention, or production-reliability data.
  • Implications: Audio is a credible capability to monitor for product expansion, but its adoption signal remains earlier-stage than the demonstrated improvement in image workflows.; Multimodal workflow design may become more important, especially where voice interfaces or media production are core rather than decorative.

Talent, developer experience, and forward-looking sentiment

  • Claims: AI engineering is becoming a discipline spanning founders, CTOs, engineers, and product practitioners, rather than a stable job title.; Builders are more satisfied with their work but recognize an accumulating maintenance bill.; Respondents expect dramatic technical and narrative change, but their long-range beliefs are uncertain.
  • Evidence: Among respondents with more than a decade of software experience, over half have no more than three years of AI experience; the median newest engineer has nearly as much AI experience as the median veteran.; 76% say AI has increased job satisfaction, while 59% fear current AI-generated code will create long-term liabilities; only about one-third call software engineering a solved problem.; 67% expect a leading lab to declare AGI within five years, though the wording intentionally asks about a declaration rather than achievement. Only 9% expect transformers to remain state of the art in five years.
  • Caveats: The AGI and architecture questions measure beliefs and expectations, not forecasts supported by operational data.; The reported satisfaction is drawn from a population already engaged in AI building and may not generalize to teams experiencing displacement or weaker adoption outcomes.
  • Implications: The scarce capability is increasingly the ability to specify, supervise, and operationalize AI-enabled work across functional boundaries.; Hiring and organizational design should account for stronger cross-functional builders while maintaining clear ownership for production systems and accumulated technical debt.

Notable Concepts & Terms

  • Intent-to-adopt ratio: The share of non-users of a modality who plan to adopt it; the speaker uses it to identify audio as the strongest emerging modality signal.
  • Multi-model routing: Selecting models per task, output quality, or cost rather than standardizing on a single model; presented as the prevailing production pattern.
  • Great standardization of the platform: The survey's framing that organizations are consolidating tools and surrounding infrastructure while retaining model-provider flexibility.
  • Write-enabled agents: Agents with permission to modify data or act inside systems; their rapid growth is the survey's central agent maturity signal and governance concern.
  • Harness engineering: The broader engineering work needed to make agents function in production—tools, permissions, context, execution environments, and controls—rather than merely prompting a model.
  • Vibe review: Informal human judgment of AI outputs; it remains the most common evaluation approach despite evals being the top stack problem.
  • First-class cost constraint: Treating token and model expense as a core product and operational variable, similar to an SLA or quality metric.
  • Cheaper failure: AI's principal organizational benefit in this survey: reducing the cost of prototypes and experiments so teams can place more bets.

Operator Notes / Why Ken Should Care

  • Require every production agent with write authority to have scoped permissions, an action audit trail, a rollback or compensation path where possible, and explicit approval thresholds based on action risk.
  • Establish a model-routing layer with measurable policies for quality, latency, and unit cost; avoid allowing individual workflows to hard-code provider choices without an exit path.
  • Make cost per successful task, token consumption, and retry behavior standard production telemetry alongside quality metrics.
  • Prioritize an evaluation program for multi-step agent behavior: representative task suites, context-loss tests, tool-use regressions, and human review calibration—not just single-response benchmarks.
  • Set production gates for AI-built changes from non-engineering functions, including code ownership, test requirements, security review, and a named maintainer after launch.
  • Treat prompt management, retrieval logic, and evaluation assets as proprietary product infrastructure; buy commodity inference unless a specific performance, privacy, or economics case justifies owning it.

Source/Metadata

  • Title: 2026 State of AI Engineering — Barr Yaron, Amplify Partners
  • Transcript words: 5569
  • Duration seconds: 1187
  • Timestamp note: No usable timestamps or chapters were present in the supplied transcript. The transcript includes a substantial duplicated second pass of the presentation and an early extraction artifact.
Full transcript 3079 words · 25 min read
0:00

Music Now joining us on stage is the partner at Amplify, Bar Yeren. Music Music Music Music Music Music

0:37

Fantastic. You did a great job practicing. I feel very, very loved. Let's get started. As you just heard, my name is Bar. I run a survey every year on the state of AI engineering. And the funny thing about running a survey on the state of AI engineering is that the field changes as you make the slides. Just in the past week, we've had frontier releases treated like national security events, Meta reportedly exploring selling AI compute. By the time I get off stage, maybe something else will happen. So if I miss a major announcement while I'm up here, please come find me after.

0:51

The second floor, and I can't possibly find a successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful successful

1:19

for AI engineers, and I'll make the same promise that I make every single year, which is short time on bar, long time on bar charts. So let's get right into it with lots of bar charts.

1:53

First, let's talk about, well, maybe raise your hand. Did you fill out the survey? This is a very large group. OK, yes, I see you in the front. If the answer is you, thank you so much. If the answer is not you, I will find you in 2027. But genuinely, this only exists because 1,000 of you gave your time, so thank you. We had 1,048 respondents this year, which is a lot of AI engineers. And to be precise, this is not just AI engineers, as I'm sure you see at the conference. Every year, we see that AI engineering is more of a discipline than a job title. It touches founders, CTOs, engineers, product people, folks across company sizes, and experience levels.

1:59

And that range shows up in experience, too. For the third year running, we see the same pattern, which is skewed toward senior engineers, but newer to AI. Of those with over 10 years of software experience, over half have three years or less of AI experience, which tracks. These are very experienced engineers learning a new paradigm in real time. And the newest cohort, the ones who just started engineering, the median new engineer has nearly as much AI experience as the median 10-year software veteran. So the newest engineers have never known software without this.

2:02

But doing AI doesn't mean one thing. We talked about all these different titles, all these different roles. Before we get into models and agents, I have a more basic question, which is: when people say they're doing AI at work, what are they actually doing?

2:03

So first up, I like to start with the modalities. We asked, which modalities are you actively building with at work? Can anyone take a guess? Text dominates. I know. Hold your applause. But one piece of this chart that I always find very interesting, and I always look at, is the ratio of "Nope, I'm not using this modality" to "I'm not using it, but I do plan to." I call this the intent to adopt ratio. Of the people who are not building with a modality today, how many say they plan to use it? And audio has the strongest intent to adopt this year. Among AI engineers who are not building with audio today, a whopping 56% say they plan to adopt it in the AI applications they build.

2:05

And this is not a brand new signal. Last year, audio also had the highest intent to adopt across modalities, but 37%. So audio continues to take the lead and have high interest, but that interest is accelerating.

2:09

Now, there has been an audio swing, but if we look at what changed most from last year in the survey, the biggest jump is actually in people using image generation. The share of respondents using generative AI for images and feeling really good about it doubled from 18% last year to 36% this year. Makes sense if you look at what we launched in the same window. Over the past year, plus survey time, we've had models, Nano Banana, Nano Banana 2, ChatGPT images 2.0. The products have gotten much better. What used to feel like an efficient way to generate cursed hands is increasingly becoming a part of real work. Audio may have the strongest intent to adopt, but image generation shows us what happens when a modality crosses that threshold. So I'm excited to continue watching these adoption curves every single year. I think we're going to see a lot this year.

2:12

Now, models. Who here spends time on Twitter? All right. Yes. I imagine this is a very Twitter-pilled crowd. If you spend any time on Twitter in this circle, you've seen a lot written about open weight models these past few months. And I think we'll see it even more in the next year. So we asked, what models are you actually using in production? Ninety-four percent use closed models. Forty-five percent are using open weight models. But here's the thing: open weight models are not replacing closed models for the most part, at least not yet. The respondents using open weight models, over 90% of them are also using closed models. So they're looking like an augmentation. Teams are mixing and matching.

2:15

We also asked, just to double click on this, for the top three considerations when choosing a model. If you're choosing a model, what is important to you? And despite the air time of the open versus closed, it's not what drives model choice. It was a top three consideration for only 5% of the respondents. What matters is actually more straightforward. It's quality. Quality dominates, followed by agentic capabilities like tool calling, and cost tied right with it. Well, money, money, money. We'll get back to that.

2:18

One thing that I found very interesting is that reliability is not near the top. Only one in five named reliability. That doesn't mean teams stopped caring about reliability. There are different ways to interpret this data. My guess is that it's more likely to become a threshold requirement, and the models they're choosing are reliable enough, so the decision moves up the stack outside of certain circumstances to quality, capability, cost. But we could talk after.

2:20

All right. So here's where the model story all comes together. As I said, teams are not choosing one model and calling it a day. Earlier I showed that 87% of teams are using more than one model. That's the opposite of standardization. And the way that they choose models for given tasks varies. Most popular is routing by task type. Some run multiple models, compare outputs. Some route based on cost. But models are good at different things.

2:23

What was interesting was that more than half of respondents said that their organization is starting to standardize on fewer AI tools. They're trading flexibility for standardization. A share of those are mixed. They say they're standardizing on some layers while staying flexible on others. But the headline here is that we're in the early great standardization of the platform and tools, not the models.

2:26

All right. This is the slide where anyone who's opened an AI bill in the last year starts nodding. So it turns out that infinite intelligence still comes with a usage-based bill. Once teams are managing many models and AI workflows, the next question becomes cost. Cost is now a first-class engineering constraint.

2:29

We see this in the data. Forty percent of respondents say that cost regularly shapes how ambitiously they use AI. And another 36% say that it sometimes does. This is pretty straightforward. So all in, about three out of four respondents are adjusting their AI usage based on cost. And maybe the fourth has a company card. That might be surprising, or maybe it's obvious. But 12 months ago it was not. Token maxing is cool. Being able to find real use cases is amazing. But cost is becoming a really big part of the product decision today.

2:31

And it shows up in monitoring, too. What folks are monitoring in production includes cost and token usage as the number two thing they watch for. It's being monitored like an SLA, right under quality itself. Which brings us to the biggest line item of them all: agents. We've been talking about agents for a while. This year, as you've seen, as you'll see today, as you've seen in previous days, you're going to talk a lot about harness engineering. They're escaping demo world. So we asked respondents what level of tool permissions their agents typically have. And this is where agents start to look more real.

2:35

There are two things happening at once. First, and I don't think this is surprising, relative to last year, there are far more teams using agents. This year, 95%—this seems high to me—95% say they're using agents, roughly double last year. Second, among the teams that are using agents, those agents are much more likely to have write access. Last year, 52% of folks building with agents said their agents could actually write data. This year, that number is 89%.

2:38

So when you combine these two shifts, more teams using agents and more of those agents having write permissions, the share of all the respondents—and again, it's a survey—using write-enabled agents is up more than three times relative to last year. So this is really the big shift. Agents are no longer reading, summarizing, drafting. Agents are taking actions inside systems.

2:42

And that raises the obvious question: how are we controlling all of this? With pretty blunt instruments. There are many ways that folks are controlling agents today. The top two are human-in-the-loop approvals and gating permissions, which are the right instincts, but are kind of the same toolkit you'd use to manage an intern. Learn. Below that, the results scatter. Task decomposition, retrieval, memory, sandboxing—people are trying everything. Nobody has settled the control layer for agents. Memory and persistent context is one that I'm watching very carefully right now. I think it's going to evolve a lot in the next year.

2:46

And when agents fail, or when people complain about agents failing, to be more precise, it's usually the thinking, not the plumbing. So two thirds say that hallucination or losing context mid-task is what frustrates them the most. All right. So agents are out in the wild, which makes it a good time to look at what everyone's actually running underneath. So let's take a peek at the stack. We asked what is the biggest challenge in your stack. Every single year that I ask this, the number one answer is evals. So evals lead here, same as always, but by a very thin margin. That margin is getting smaller.

2:52

And I'll say the quiet part here, which is that 96% of the people in the survey in this room have a problem with the stack. You just can't agree on which one. So if you're deciding what to build next, if you're interested in infrastructure, that scatter is the map. And the leading challenge, how to evaluate your AI outputs, requires many different methods. But as always, the vibe review is number one. So there are some consistent things that we'll see if they change over time, but they have not changed.

2:56

Okay. This is interesting. So across eight layers of the stack, we asked what do people build versus buy? Again, maybe the corporate card is going to play a part in this. But there is a wide range and mix for every layer of the stack and a few clear takeaways.

3:01

So the first is that inference and model serving is the layer that people buy the most. Many people don't want to build inference infrastructure, and fair enough. Prompt management is the opposite. Sixty-one percent build it themselves. Apparently, everyone's prompts are special. And this is true of a lot of the product logic. Prompts, RAG, evals: they tend to stay in-house on a relative basis. Fine-tuning is the clearest "not yet." Most people don't have it at all. And folks are pretty locked in. So those who bought aren't looking as much to build. Those who built aren't looking as much to buy. But those are the core takeaways from the usage in our stack.

3:03

So many of you work on teams. And as we said at the start, these range from solo founders to large enterprises. What is this doing to teams? And remember, this is a builder-heavy sample. But among builders, the vibes are good, which I'm sure, if you look to your left and your right, you're feeling that. The vibes are pretty good. Ninety-seven percent report a net positive effect on their organization. The top effect isn't really just speed. It's cheaper failure. More experimentation, more prototypes, more bets. It didn't just make engineers faster, but it made trying things nearly free. And so there are some happy campers as a result of that.

3:10

But it's not free-free. There's no free lunch, as nothing is. So the same tool that increases experimentation also increases review burden. Both can be true. And over nine in 10 respondents are feeling negative downstream effects in some way. The most common ones, being widely discussed at this conference, online, and anywhere that you see AI engineers, are erosion of deep technical skills and understanding of the code base. And these are consequences of cheap code generation.

3:13

And the org chart is really feeling it. So many folks, 81%, are saying that AI is blurring the line between their role as engineers and product design and marketing. These stats shocked me. Where you feel it the most is shipping software, once exclusively the engineer's domain. I know folks talk about vibe coding and how that's accessible to more folks than ever before in different roles. But today, over a third of teams have non-developers shipping features, which was pretty wild to me. Mostly smaller, mostly internal, but 17% say that non-developers are regularly shipping customer-facing features across the stack. And even when non-developers aren't shipping, a third of teams see them building really useful things: prototypes, front-end mocks, and more. So shipping software is not gated on being an engineer. We knew this, but the extent to which it's being pushed is higher than I expected.

3:16

All right. So where does all of this go? We always ask people to place bets, rapid-fire. So let's talk about those results.

3:21

So present tense first. Seventy-six percent say AI boosted their job satisfaction. So that's good for most of this crowd. I hope you're, as Elphaba and Glinda say, I hope you're happy now. That's great. But 59% fear today's AI code creates long-term liabilities. Only a third call software engineering a solved problem, although when I have conversations with folks, sometimes the way in which they define software engineering is different. So you can read into that stat as you will. Happier, faster, but embracing the maintenance bill is the TLDR. And people are unsure what's going to happen with hiring.

3:23

And for the five-year bets, we have 67% expect a leading lab will declare AGI in the next five years. Note the wording. We said "will declare." We asked about the press release, not the achievement. So will they declare it? Yes. What does that mean? Not sure. Only 9% bet on transformers being state of the art in five years. Most are unsure. That was interesting. And then my favorite: will there be more AI compute in space or on land? Thirty-six yes. Thirty-eight no. The most divisive question in the survey is about outer space.

3:25

I promised you a lot of bar charts. And that was a lot of information. So a review, or our 2026 wrapped: impact is overwhelmingly positive. ImageGen doubled, or happy ImageGen doubled, while audio has the highest adoption intent, the same as last year. Cost really became a first-class constraint, and we see that everywhere in monitoring and how ambitious folks that are going out and building AI products are behaving. Open weights augment, but they don't replace. So we're seeing a multi-model future with a consolidation of the stack. Agents got write access more than ever before, tripling relative to last year, while the guardrails stayed pretty primitive. And inference is the buy market. Everything closer to product logic tends to relatively stay more in-house.

3:26

It is a very exciting time to be an AI engineer. I cannot wait to see how the next year unfolds. So you can find the full report in the link up here. Every chart, plus some cuts that we didn't have time for today. I won't ask you to fill out a survey about the survey. But if there's something that you want on the books for 2027, something you're curious about, you can come find me here on the internet. I'm easy to spot. Thank you so much. We will see you next year, or per 36% of you, maybe in orbit. Thank you. is when people say they're doing AI at work, what are they actually doing? So first up, like to start with the modalities.

3:38

We asked, which modalities are you actively building with at work? Can anyone take a guess? Text dominates. I know. Hold your applause. But one piece of this chart that I always find very interesting, and I always look at, is the ratio of, nope, I'm not using this modality to, I'm not using it, but I do plan to. I call this the intent to adopt ratio. Of the people who are not building with a modality today, how many say they plan to use it? And audio has the strongest intent to adopt this year. Among AI engineers who are not building with audio today, a whopping 56% say they plan to adopt it in the AI applications they build. And this is not a brand new signal.

4:25

Last year, audio also had the highest intent to adopt across modalities, but 37%. So audio continues to take the lead and have high interest, but that interest is accelerating. Now, there has been an audio swing, but if we look at what changed most from the last year in the survey, the biggest jump is actually in people using image generation. The share of respondents using generative AI for images and feeling really good about it doubled from 18% last year to 36% this year. Makes sense if you look at what we launched in the same window. Over the past year, plus survey time, we've had models, Nano Banana, Nano Banana 2, ChatGPT images 2.0.

5:10

The products have gotten much better. What used to feel like an efficient way to generate cursed hands is just increasingly becoming a part of real work. Audio may have the strongest intent to adopt, but image generation shows us what happens when a modality crosses that threshold. So I'm excited to continue watching these adoption curves every single year. I think we're going to see a lot this year. Now, models. Who here spends time on Twitter? All right. Yes. I imagine this is a very Twitter-pilled crowd. If you spend any time on Twitter in this circle, you've seen a lot written about open weight models these past few months.

5:50

And I think we'll see it even more in the next year. So we asked, what models are you actually using in production? 94% use closed models. 45% are using open weight models. But here's the thing. You know, open weight models are not replacing closed models for the most part, at least not yet. The respondents using open weight models, over 90% of them are also using closed models. So they're looking like an augmentation. Teams are mixing and matching.

6:23

We also asked, just to double click on this, for the top three considerations when choosing a model. If you're choosing a model, what is important to you? And despite the air time of the open versus closed, it's not what drives model choice. It was a top three consideration for only 5% of the respondents. What matters is actually more straightforward. It's quality. Quality dominates. Followed by agentic capabilities like tool calling. And cost tied right with it. Well, money, money, money. We'll get back to that. One thing that I found very interesting is that reliability is not near the top. Only one in five named reliability.

7:03

That doesn't mean teams stopped caring about reliability. There are different ways to interpret this data. My guess is that it's more likely to become a threshold requirement and the models they're choosing are reliable enough so the decision moves up the stack outside of certain circumstances, to quality, capability, cost. But we could talk after. All right. So here's where the model story all comes together. Like I said, teams are not choosing one model and calling it a day. Earlier I showed that 87% of teams are using more than one model. That's the opposite of standardization. And the way that they choose models for given tasks varies.

7:45

Most popular is routing by task type. Some run multiple models, compare outputs. Some route based on cost. But models are good at different things. What was interesting was that more than half of respondents said that their organization's starting to standardize on fewer AI tools. They're trading flexibility for standardization. A share of those are mixed. They say they're standardizing on some layers while staying flexible on others. But the headline here is that we're in the early great standardization of the platform and tools, not the models. All right. This is the slide where anyone who's opened an AI bill in the last year starts nodding.

8:26

So it turns out that infinite intelligence still comes with a usage based bill. Once teams are managing many models and AI workflows, the next question becomes cost. Cost is now a first class engineering constraint. We see this in the data. 40% of respondents say that cost regularly shapes how ambitiously they use AI. And another 36% say that it sometimes does. Well, this is pretty straightforward. So all in, about three out of four respondents are adjusting their AI usage based on cost. And maybe the fourth has a company card.

9:06

That might be surprising or maybe it's obvious. But 12 months ago it was not. Token maxing is cool. Being able to find real use cases is amazing. But cost is becoming a real big part of the product decision today. And it shows up in monitoring too. What folks are monitoring in production includes cost and token usage as the number two thing they watch for. It's being monitored like an SLA right under quality itself. Which brings us to the biggest line item of them all, agents. We've been talking about agents for a while. This year, as you've seen, as you'll see today, as you've seen in previous days, you're going to talk a lot about harness engineering.

9:49

They're escaping demo world. So we asked respondents what level of tool permissions their agents typically have. And this is where agents start to look more real. There are two things happening at once. First, and I don't think this is surprising, relative to last year, there are far more teams using agents. This year, 95% this seems high to me, 95% say they're using agents, roughly double last year. Second, amongst the teams that are using agents, those agents are much more likely to have write access. Last year, 52% of folks building with agents said their agents could actually write data. This year, that number is 89%.

10:38

So when you combine these two shifts, more teams using agents, and more of those agents having write permissions, the share of all the respondents, and again, it's a survey, using write enabled agents is up more than three times relative to last year. So this is really the big shift. Agents are no longer reading, summarizing, drafting. Agents are making actions inside of systems. And that raises the obvious question, how are we controlling all of this? With pretty blunt instruments, there are many ways that folks are controlling agents today.

11:12

The top two are human in the loop approvals and gating permissions, which are the right instincts, but kind of the same toolkit you'd use to manage an intern. Learn. Below that, the results scatter. Task decomposition, retrieval, memory, sandboxing, people are trying everything. Nobody has settled the control layer for agents. Memory and persistent context is one that I'm watching very carefully right now. I think it's going to evolve a lot in the next year. And when agents fail, or when people complain about agents failing, to be more precise, it's usually the thinking, not the plumbing.

11:48

So, you know, like two thirds say that hallucination or losing context mid-task is what frustrates them the most. All right. So agents are out in the wild, which makes it a good time to look at what everyone's actually running underneath. So let's take a peek at the stack. We asked what is the biggest challenge in your stack. Every single year that I ask this, the number one answer is evals. So evals lead here, same as always, but by a very thin margin. Like, that margin is getting smaller. And I'll say the quiet part here, which is that 96% of the people in the survey in this room have a problem with the stack. You just can't agree on which one.

12:30

So if you're deciding what to build next, if you're interested in infrastructure, that scatter is the map. And the leading challenge, how to evaluate your AI outputs, requires many different methods. But as always, the vibe review is number one. So there are some consistent things that we'll see if they change over the time, but they have not changed. Okay. This is interesting. So across eight layers of the stack, we asked what do people build versus buy? Again, maybe the corporate card is going to play a part in this. But there is a wide range and mix for every layer of the stack and a few clear takeaways.

13:11

So the first is that inference and model serving is the layer that people buy the most. Many people don't want to build inference infrastructure, and fair enough. Prompt management is the opposite. 61% build it themselves. Apparently, everyone's prompts are special. And this is true of a lot of the product logic, prompts, rag, evals. They tend to stay in-house on a relative basis. Fine-tuning is the clearest not yet. Like most people don't have it at all. And folks are pretty locked in. So those who bought aren't looking as much to build. Those who built aren't looking as much to buy. But those are the core takeaways from the usage in our stack.

13:59

So many of you work on teams. And like we said at the start, these range from solo founders to large enterprises. What is this doing to teams? And remember, this is a builder-heavy sample. But among builders, the vibes are good, which I'm sure if you look to your left and your right, you're feeling that. The vibes are pretty good. 97% report a net positive effect on their organization. The top effect isn't really just speed. It's cheaper failure. More experimentation, more prototypes, more bets. It didn't just make engineers faster, but it made trying things nearly free. And so there's some happy campers as a result of that. But it's not free-free.

14:46

There's no free lunch, as nothing is. So the same tool that increases experimentation also increases review burden. Both can be true.

14:57

And over nine in 10 respondents are feeling negative downstream effects in some way. The most common ones being widely discussed at this conference, online, and anywhere that you see AI engineers, erosion of deep technical skills and understanding of the code base. And these are consequences of cheap code generation. And the org chart is really feeling it. So many folks, 81%, are saying that AI is blurring the line between their role as engineers and product design and marketing. These stats shocked me. Where you feel it the most is shipping software, once exclusively the engineer's domain.

15:44

I know folks talk about vibe coding and how that's accessible to more folks than ever before in different roles. But today, over a third of teams have non-developers shipping features, which was pretty wild to me. Mostly smaller, mostly internal, but 17% say that non-developers are regularly shipping customer-facing features across the stack. And even when non-developers aren't shipping, a third of teams see them building really useful things, prototypes, front-end mocks, and more. So shipping software is not gated on being an engineer. We knew this, but the extent to which it's being pushed is higher than I expected. All right. So where does all of this go?

16:29

We always ask people to place bets rapid fire. So let's talk about those results.

16:38

So present tense first. 76% say AI boosted their job satisfaction. So that's good for most of this crowd. I hope you're, as Elphaba and Glinda say, I hope you're happy now. That's great. But 59% fear today's AI code creates long-term liabilities. Only a third call software engineering a solved problem. Although when I have conversations with folks, sometimes the way in which they define software engineering is different. So you can read into that stat as you will. Happier, faster, but embracing the maintenance bill is the TLDR. And people are unsure what's going to happen with hiring.

17:21

And for the five-year bets, we have 67% expect a leading lab will declare AGI in the next five years. Note the wording. We said will declare. We asked about the press release, not the achievement. So will they declare it? Yes. What does that mean? Not sure. Only 9% bet on transformers being state of the art in five years. Most are unsure. Most are unsure. But that was interesting. And then my favorite, will there be more AI compute in space or on land? 36 yes. 38 no. The most divisive question in the survey is about outer space. I promised you a lot of bar charts. And that was a lot of information. So a review or our 2026 wrapped. Impact is overwhelmingly positive.

18:13

ImageGen doubled, or happy ImageGen doubled, while audio has the highest adoption intent, the same as last year. Cost really became a first class constraint. And we see that everywhere in monitoring and how ambitious folks that are going out and building AI products are behaving. Open weights augment, but they don't replace. So we're seeing a multi-model future with a consolidation of the stack. Agents got write access more than ever before, tripling relative to last year, while the guardrails stayed pretty primitive. And inference is the buy market. Everything closer to product logic tends to relatively stay more in-house. It is a very exciting time to be an AI engineer.

19:01

I cannot wait to see how the next year unfolds. So you can find the full report in the link up here. Every chart plus some cuts that we didn't have time for today. I won't ask you to fill out a survey about the survey. But if there's something that you want on the books for 2027, something you're curious about, you can come find me here on the internet. I'm easy to spot. Thank you so much. We will see you next year or per 36% of you, maybe in orbit. Thank you.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note