Local Models: Trust, Control, Optimization — Carter Abdallah, NVIDIA
Description
When Fable was pulled back and access to frontier systems stopped looking guaranteed, Lucas Atkins watched enterprises move to Chinese open models, not because they scored better but because availability could be counted on. That is his working definition of trust, and he separates it hard from safety: an open model is a directory of files you can inspect, running on code you can read, while the same claim about a closed API is unverifiable by construction. Arcee's response was to reorient the whole company and pretrain a 400 billion parameter model in six months, which he says plenty of people called impossible. The rest is control. Vincent Weisser describes a customer specializing an open model to automate finance work in a week or two and landing better results than Opus at a fraction of Haiku's cost. Closed terms of service bar you from training on outputs, so owning the model also means owning the traces, which is what makes a data flywheel possible at all; Nemotron and Trinity both adopted the open MDW license to put that permission in writing. Chris Alexiuk's framing is the mismanaged genius: a model tuned to be good across every harness is optimal for nobody's, and open weights let you fit it to the one or two things you actually do. The predictions land where you would expect from this panel, including open models reaching Fable level capability inside a year, and Atkins hoping the share of people who have ever run a model locally climbs from a rounding error to 10% to 15%. Speaker info: Carter Abdallah, moderator (NVIDIA): - https://x.com/Baxate - https://baxate.com Vincent Weisser (Prime Intellect): - https://www.linkedin.com/in/vincentweisser - https://www.primeintellect.ai/ Lucas Atkins (Arcee AI): - https://x.com/latkins - https://arcee.ai Chris Alexiuk (NVIDIA): - https://x.com/llm_wizard - https://www.alexi.uk/ Timestamps: 0:00 - Welcome and why this panel is the whole stack 1:14 - Prime Intellect: keeping the training stack open 2:15 - Arcee:
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Skim
- Core thesis: Open-weight models become strategically superior for production AI when teams use post-training, reinforcement-learning environments, and production traces to specialize them for a specific harness, workflow, and cost envelope.
- Why it matters: The panel frames local/open models not merely as a sovereignty or privacy alternative to APIs, but as the controllable learning layer for durable agent products: teams can own their weights, traces, costs, deployment path, and improvement loop.
- Best use: Use this as a strategic framing for when to self-host or customize models for an agent product; it is useful for thesis and architecture direction, but light on implementation specifics.
Executive Summary
NVIDIA, Prime Intellect, and Arcee argue that open models should be evaluated on control and compounding capability rather than as lower-performing substitutes for closed frontier APIs. Their central position is that an open model lets a builder inspect and validate the artifact, run it where needed, retain its interaction traces, and adapt it to a specific agent harness. In contrast, a closed API can be convenient and often remains useful, but creates dependency on provider access, pricing, model deprecation, and terms that may restrict reuse of outputs.
The practical opportunity is post-training rather than frontier pre-training. Vincent argues that a company building a finance, computer-use, or knowledge-work agent should construct a task-specific RL environment, post-train a capable open base model, deploy it, and continuously improve from production traces. The panel treats this loop—specialized environment, verifiable outcomes, real-user deployment, and trace-driven retraining—as the path to agents that outperform general frontier models on a narrow workflow while costing materially less.
The speakers distinguish trust from safety. Their claim is not that open models are automatically safe; models remain probabilistic systems requiring output review and safeguards. Rather, they define trust operationally: knowing what model is running, retaining access to it, being able to inspect its code, weights, data disclosures, and license, and predicting its cost. This gives open models an advantage for regulated or long-lived deployments, even though the panel does not offer a rigorous methodology for proving training-data provenance from weights alone.
Their market view is hybrid, not absolutist. Closed models remain appropriate for broad consumer productivity and certain frontier tasks, while open/self-hosted models become a Linux-like infrastructure layer for enterprises, local devices, and specialized systems. The forward-looking claims—that local models will cover most daily AI tasks, open models will reach or exceed current top closed-model capability, and on-device agent operating systems will emerge—are directional predictions rather than evidence-backed forecasts.
Key Takeaways
- Claim: For production AI, openness is valuable because it provides operational trust: teams can validate the model artifact, control access, and avoid dependency on an opaque external API. | Evidence: Lucas contrasts inspectable model weight files and runnable code with an arbitrary API; he names open-serving and implementation ecosystems including Prime RL, vLLM, and SGLang. He also notes that enterprises shifted toward Chinese open models when access to some frontier systems appeared less assured. | Implication: For sensitive, regulated, or business-critical agents, evaluate model choice partly as a continuity and auditability decision, not just a benchmark decision. | Caveat: Inspectability does not itself make a model safe, reliable, or free of data-provenance concerns; the panel acknowledges that model outputs remain inherently probabilistic and require safeguards.
- Claim: Post-training an open base model on the specific harness and task environment can be more valuable than relying on a general-purpose frontier model. | Evidence: The panel cites OpenAI Deep Research as an example of a product-specific model/harness combination, saying it used RL and test-time compute rather than simply looping over search. Speakers argue that OpenAI, Anthropic, and Google commonly release customized models for their own products, implying off-the-shelf general models are not automatically optimal for third-party applications. | Implication: Treat the agent harness, evals, tools, and model adaptation as one system. If a workflow matters enough, build a task environment and train the model against that environment rather than pursuing prompt-only optimization.
- Claim: Production traces are the core asset for building a specialized-agent moat because they enable a continuous post-training and evaluation flywheel. | Evidence: Vincent describes a loop of creating an RL environment for a domain, deploying the agent to users, collecting millions of computer-use or workflow traces, then using those traces, verifiers, NeMo RL, or NeMo Gym environments to improve the model. He compares this to the last-mile progression toward autonomy. | Implication: Prioritize workflows with observable state transitions, verifiable outputs, and high interaction volume. Instrument traces and feedback now so a future custom-model loop is possible. | Caveat: This only works where outcomes can be measured or verified well enough to construct useful rewards, evaluations, and guardrails; the panel does not explain how to solve sparse or ambiguous rewards.
- Claim: Specialized open models can improve unit economics because most business tasks do not need a hyper-general frontier model. | Evidence: Chris estimates that most people do not require frontier-level intelligence for roughly 90% of tasks and argues that open models can sacrifice broad capability to become exceptionally good at one or two functions. Vincent cites Ramp as having specialized an open model for finance automation in one or two weeks, reportedly exceeding Opus performance at a fraction of Haiku's cost. | Implication: Route routine, bounded, high-volume tasks to smaller or specialized local models, reserving expensive frontier calls for cases where broad reasoning or rare capabilities actually change the outcome. | Caveat: The Ramp result is an anecdote without task definitions, evaluation methodology, throughput figures, or independently verifiable cost/performance data.
- Claim: Owning model outputs and traces matters as much as owning weights because it determines whether a team can reuse interaction data to improve future models. | Evidence: The panel says closed-model terms can discourage training on outputs from providers such as Claude, GPT-5, or the transcript's referenced 'Fable,' whereas open-model traces can be retained for direct fine-tuning or for generating training signals. NeMoTron and Trinity are said to have adopted an Open MDW (Model Data Weights) license intended to explicitly permit use of outputs to produce or train models. | Implication: Make data and output-reuse rights a first-class part of model procurement and architecture review; avoid building a core improvement flywheel on traces that cannot legally or contractually be reused. | Caveat: License permission is only one constraint: organizations still need to address user consent, data rights, privacy, and contractual restrictions before retaining or training on production traces.
- Claim: The likely near-term architecture is hybrid: closed models remain useful, while open/local models become a self-hosted infrastructure layer for specialized and sovereignty-sensitive work. | Evidence: Lucas compares open/self-hosted intelligence to Linux running enterprise and cloud infrastructure, while comparing closed tools such as ChatGPT to accessible consumer productivity software. The panel names hosted open-model providers including Fireworks, Together, Baseten, and Modal as evidence that self-owned model infrastructure need not mean operating every layer internally. | Implication: Design a routing/control plane rather than choosing one model ideology: retain frontier APIs for exceptional tasks and use owned or hosted open models where cost, privacy, customization, latency, or continuity justify them. | Caveat: The Linux analogy is strategic rhetoric, not proof that open models will win every enterprise workload; frontier providers can still lead on capability, ease of use, and managed reliability.
Detailed Brief
How the panel defines trust, control, and optimization
- Claims: The speakers reject the tendency to collapse trust into safety. Their version of trust is the ability to know what is being run, verify access will continue, understand likely costs, and inspect or reproduce the system.; Control spans the full lifecycle: weights, data disclosures, training methodology, training framework, fine-tuning pipeline, inference stack, outputs, and deployment location.; Optimization is framed as outcome-maxing rather than token-maxing: the relevant measure is whether a dollar of compute produces more than a dollar of business value.
- Evidence: NVIDIA's NeMoTron program is described as releasing weights, datasets, training environments, and training methodology where possible; speakers mention data summaries that identify token volumes by dataset.; The panel emphasizes that cheaper per-token pricing does not necessarily lower total spend because agent sessions have become substantially longer and more token intensive.; Chris's maxim is that 'faster models are smarter models,' meaning usable intelligence increasingly depends on inference speed and compatibility with hardware outside large data centers.
- Caveats: The claim that one can generally infer training data from weights alone is asserted but not substantiated, and should not be treated as a replacement for provenance documentation, audits, or formal data governance.; Cost predictability from self-hosting trades API pricing uncertainty for infrastructure, capacity-planning, observability, and operational responsibility.
- Implications: A model platform decision should include total-session cost, latency, deployment continuity, trace ownership, and retraining rights—not merely input/output token price.; For local-agent deployment, inference runtime and hardware compatibility are product constraints, not downstream performance optimizations.
Forward-looking predictions and their strategic signal
- Claims: The panel expects the next year to move adoption from chatbots and coding agents toward domain-specific knowledge-worker and computer-use agents.; It predicts local laptops and phones will soon run models capable enough for common day-to-day AI work, and that specialized multi-model systems will become more common than reliance on one universal model.; The speakers expect continued architectural shifts, including text diffusion approaches, to make models more suitable for consumer hardware.
- Evidence: Vincent identifies coding agents as the current adoption model and argues that capability thresholds can trigger product inflections, citing Cursor's growth once coding models became good enough.; Chris says a 4-billion-parameter model can already run on a phone and claims it is more useful than GPT-4 was at launch.; Lucas estimates that only 0.0001% of the roughly one to two billion daily AI users have ever run an open model locally, and expresses a goal—not a supported forecast—that 10-15% will do so.
- Caveats: Several product and model references in the transcript, including 'Fable,' 'GPT-5.5/5.6,' and 'JM 5.2,' may reflect transcription errors or event-specific naming; do not rely on them without source verification.; These are optimistic panelist forecasts from organizations directly invested in the open-model ecosystem, rather than neutral market projections.
- Implications: The adoption opportunity may shift from selling general chat to packaging a narrowly useful local/on-device workflow with a strong user experience.; Monitor device compute, memory, and model-runtime improvements as potential triggers for new private-by-default agent products.
Notable Concepts & Terms
- Post-training: The panel's preferred economic route to customization: adapt a capable open base model to a particular domain or agent task rather than pre-training a foundation model from scratch.
- RL environment: A task-specific environment with rewards, verifiers, or measurable outcomes used to train an agent toward reliable execution of a workflow.
- Harness-model fit: The idea that an agent's tools, prompts, runtime, and workflow should be co-designed with a model trained to work especially well inside that particular environment.
- Production traces: Records of model-agent interactions and outcomes that can create a proprietary data flywheel for evaluation, fine-tuning, and reward-signal generation.
- Outcome-maxing: Optimizing for business value produced per compute dollar, rather than minimizing token use in isolation.
- Open MDW license: A referenced Model Data Weights license intended to clarify that users may use model outputs to produce or train models, unlike restrictive provider terms.
- Sovereign intelligence: The ability for an organization or jurisdiction to retain access, deployment control, and operational independence over AI systems rather than depending fully on external providers.
- Linux layer: The panel's analogy for open/self-hosted models as a widely deployed, optimized foundational infrastructure layer beneath more polished managed products.
Operator Notes / Why Ken Should Care
- Add a model-rights checklist to vendor selection: weight availability, self-hosting rights, output ownership, trace-retention rights, training/reuse rights, model-deprecation exposure, and geographic deployment options.
- Identify one agent workflow with high volume and objectively verifiable outcomes; build an evaluation/RL environment before investing in fine-tuning.
- Capture structured production traces, tool calls, outcomes, corrections, and approvals with explicit consent and retention controls so they can support future post-training.
- Implement routing that measures outcome quality, total session cost, latency, and failure rate across a frontier API and a specialized open model; do not compare models solely through generic benchmarks.
- Treat the cited Ramp performance and cost claim as a lead for due diligence, not a planning assumption; request task-level evaluation design and full-cost inference metrics from any model-customization partner.
- Monitor local inference runtimes and device hardware as a potential distribution channel, but keep heavy enterprise agents on infrastructure sized for reliability, observability, and data controls.
Source/Metadata
- Title: Local Models: Trust, Control, Optimization — Carter Abdallah, NVIDIA
- Transcript words: 15136
- Duration seconds: 2600
- Timestamp note: No timestamps or chapters were present. The transcript contains substantial repeated passages and apparent transcription/name ambiguities.
Transcript
I hope everybody had a great lunch and you got to check out some of the amazing demos that we have. We're going to begin the panel, the first panel of the afternoon here, where we're going to be talking about, of course, the engines that are actually powering the stuff that could remotely be used for things like local, sovereign, any kind of ownership over your own artificial intelligence. And of course, the engine powering those, in addition to the hardware, is the models themselves. And so for this panel, we have excellent guests. We have Vincent, who's the CEO and founder of Prime Intellect. We've got Lucas, the CTO of RCAI. And we've got Chris, who is the senior product research engineer on the NemoTron family of models at NVIDIA. Now, what's really cool about working in this industry is really cool companies like this, we all get to work together. And so this is one panel where we all directly get to work together on both models, infrastructure, and some of the ways that we think the direction of the industry should go. And each of us kind of play a different role in that stack. But I want to leave it to you guys to introduce yourself and be able to talk about the charter that you see, the problem of the stack that you guys are working on. Awesome. Should I kick it off? Yeah, so I'm Vincent, as you mentioned, and really the goal with Prime Intellect from the beginning was to ensure that Frontier Intelligence will be open and accessible, not just the models, but also the full stack to train the models. So this was our motivation from the beginning. And we've worked together with a lot of gentlemen on the stage. On the one side, we work with folks like Lucas and Arcee to help them train Frontier Open Models. We help NVIDIA on the NemoTron Coalition help out their Frontier Open Models. And I think I'm actually, I think both NemoTron and Trinity might be the best two open models right now outside of China. So I think we need to fact check that, but this is actually, I think it might be. It's our marketing says that, yeah. But yeah, so that's the high level. My name is Lucas Atkins. I'm happy to be here, and thank you for joining. Very similar to Vincent, Arcee was founded with the idea of domain-specific owned models are going to be needed. We were founded in early 2023, jumping on the custom model train quite early. You have all these people who are excited about AI and all the things this new generation of LLMs can do, but they're using these monolithic, very expensive closed APIs for, at the time, and still, very narrow tasks that don't require, at the time, a hundred dollars per million tokens out. And through doing that, we were building on top of open models, and we were releasing a lot of our tooling in the open. And we noticed that in the United States and in the West in general, we were starting to lose leadership in the open model space. A lot of it was coming out of China, and that's amazing. I love those models. We learn a lot from them. We're close with a lot of the people building those. But when you're working with large enterprises and companies and geopolitics gets involved, whether you like it or not, you have people that become concerned about where those models are coming from. And we decided that we had a good group of people, and we had a good group of partners like Nvidia and like Prime Intellect, where we could probably try to pre-train ourselves. So last year we did that. We reoriented the entire company towards let's figure out how to pre-train a 400 billion parameter model in six months. A lot of people said it was impossible, and in many ways it was, but we figured it out. And now we are an open model lab working with our wonderful partners and our customers to build Western open models that are permissive, and you can own those and customize them or run them wherever you want. And that's kind of where we're at right now. So thanks for having me. Yeah. And so I'm Chris Alexiuk. I work at NVIDIA as a product research engineer, and I support the NemoTron family of models. I think we've talked a lot about why we do NemoTron, but just to say it a few more times, AI should be open, open as in weights, data, training methodology, training frameworks. I really respect a lot of the work that the two other peeps up here do because they believe that very strongly as well. But the NemoTron family of models is focused on being as open as humanly possible. So we have this understanding or belief that in order for AI to continue to grow and be useful to everybody, it has to be done in the open so that we can build off of each other, we can compound on each other. And part of what we do, because Team Green, this is always true, is we think that the rate that you can squeeze tokens out of models is very important. So we kind of have this mantra that faster models are smarter models. And so a lot of the decisions we make when designing a model like NemoTron are built around how fast can we make it go, especially as you are going to see in the next however many months local AI take off. We need to make sure that models are well supported on hardware that doesn't just exist in massive buildings thousands of kilometers away from you. And so, for AI to be very useful, it should be quick and open. So that's kind of the vibe of NemoTron. Who makes those buildings with the massive? Oh, that's a lot of excellent people in the world that use a lot of excellent hardware from a pretty cool company. Yeah, I heard. I heard anyway. Yeah. And I'm Carter Abdallah, your moderator for today. Something that we all kind of talked about is this building on top of each other, learning from others, whether it is people across the big pond of the Pacific Ocean from us. But really, it is a collaborative sort of research effort. And I imagine that a lot of the people here in this room share that sentiment. But as it was brought up during the inaugural panel this morning in the State of the Union, there is a growing sentiment potentially on the other side that paints open source to be something that is actually more chaotic, that there's less trust involved. And I think trust ultimately, as Lucas, you and I were talking about before, depending on who the party is and depending on what lens you're looking at it from, means different things. But ultimately, from the end consumer, somebody who's using this intelligence, or somebody who's more of a business and is actually customizing something to maybe monetize tokens in their business, can you comment, we'll start with you, Lucas, a bit on how open source and open source models are actually key to building that trust, so that when these people walk out of this room and somebody does come at them with that other angle, they can sort of steel man this side. Certainly, you can weaponize any term, and certainly trust has been weaponized, that word. And the reason I say that is because it means something based on the context and with your speech, you're speaking about it. Often, in AI, people like to conflate trust with safety. And those are not the same thing. And I'm happy to speak on safety later on. But when it comes to trust, I think that you hear a lot from closed model providers, or politicians, or people out in the space who are advocates for or against open source, that you can't trust these open models because you don't know what went into them. Well, the same is true for these closed models, even more so. The benefit of open models is that we can very easily validate what is inside of them. There is a whole bunch of files with a whole bunch of matrices in there, and you can view them, and you can see the code that is running these models. You have implementations from Prime RL, VLLM, SG-Lang, the provider themselves. These models are inherently trustworthy. You know much more about what's going on when you hit and talk to these models than you ever will when what's going on when you hit an arbitrary API. Now, that being said, certainly there is fear that people can reduce to, you can't trust that these models are writing safe code. Well, again, that is the same thing with any model. You need to use your judgment, and you need to make sure that you have the proper safeguards in place, and you're viewing the outputs of these models as the outputs of an inherently random system that we are working very, very hard to make less and less random. There is a whole bunch of files with a whole bunch of matrices in there, and you can view them, and you can see the code that is running these models. You have implementations from Prime RL, VLLM, SG-Lang, the provider themselves. These models are inherently trustworthy. You know much more about what's going on when you hit and talk to these models than you ever will what's going on when you hit an arbitrary API. Now, that being said, certainly there is fear that people can reduce, you can't trust that these models are writing safe code. Well, again, that is the same thing with any model. You need to use your judgment, and you need to make sure that you have the proper safeguards in place, and you're viewing the outputs of these models as the outputs of an inherently random system that we are working very, very hard to make less and less random. I think that a telling thing is a lot of people said, well, you can't trust Chinese models, you can't trust Chinese models, you can't trust Chinese models. That often, for the last few years, meant you can't trust open models. Well, as soon as Anthropic had to put Fable away, and people realized that, oh, our access to these frontier systems might not be universal anymore. There's probably going to be a lot of checks and balances. You had a tremendous number of enterprises and developers and companies start going to these new Chinese models because they could trust that they would always have access to that. And so when it comes down to trust and the way I view that word as it relates to open models, is do I know that what I am running, and can I be as sure as possible that when I send something to this model that I am going to get the output that I expect? And the only way currently to be 100% sure that what you are getting is what you are expecting is by hitting an open model, either that you are running yourselves or you're working with a partner like Prime Intellect or RC or Nvidia to validate. So that's my take on the word trust. I think, too, something you mentioned is we don't get to know a lot about the data that goes into these models, and that's something that I'm really happy we're trying to do. Which is not something, the incentives don't exist for everyone to do this. So it's not something that I think is mandatory or should be mandatory, thanks to the things that Lucas mentioned, which is that it's rather straightforward to validate what data did go into a model without seeing the data sources originally. But I'm happy that Nvidia continues to release data sets along with our models, release environments along with our models, to make sure that even if you can't go through the work of determining what went into the model, which you can do with the weights alone for the most part, you have a spreadsheet you can look at that says here's a couple trillion tokens of this data set, a couple trillion tokens here, and I think that that helps to educate people on why it's much easier to trust open models than the models that we don't get access to any of that. It also helps people see what that data looks like. Yeah. If you don't have someone releasing it openly, when someone says data is going in, data can take many different shapes. You can go to Hugging Face, you can go to Nvidia, or you can go to Prime Intellect or RC's Hugging Face, you can look under our data sets, and you can see exactly what that looks like, and that can help you inform your priors on it. Yeah, I think that trust also, there's some angle of a reputation. Do I believe that your intentions are pure? And I think that a lot of people, again in this room, believe that intelligence is this next layer of infrastructure for us to progress as a species, and I believe that everybody should have intelligence. So on that front I want to hand it over to Vincent, because you have this almost founding thesis that this stack should be open, the open superintelligent stack. You want everybody to have a lot more intelligence. Can you talk about how this is moving into the era of control, but beyond just data sets, how important it is to have the knobs and dials of this industry also be available in an open way for people who are building this? Yeah, I think a really important point is to be able to take those open models, customize them, and be able to build on top of them. And I think all the different components that go into it, especially from the pre-training to mid to post-training, need to be more accessible so more people can also take those amazing models and make them work for their specific use cases. So I think when we started, we also took a look at the whole stack that was out there and tried to figure out what is missing for ourselves to train open models and for helping our partners to do so, and a lot of this was around the RL and post-training stack. So we basically went deep into building out a lot of info around that, around our environment evals, around making it much more accessible to do post-training also because it's the most economically viable way to customize those models, to take an open model and to have a specific environment and specific domain and dimension that you want to improve it on, and this is what we're really doing with Prime Intellect now, enabling people to post-train specialized agentic models. So being able to take models like Trinity, for example, for Marcy or NemoTron or others, and specialize them, post-train them for the use cases that ultimately enterprises care about. So a good example, this was a company like, for example, Ramp was able to take an open model and specialize it to automate finance within a week or two to get better performance than Opus at a fraction of the cost of Haiku. And I think really this period of frontier, of being able to create these specialized models that are much better than the frontier but also faster, cheaper, I think it's a key thing enterprises care about increasingly. It's really just making it work for their use cases. If you go back to trust, it's how you can make your CFO trust you, knowing exactly how much something's going to cost all the time. That is increasingly becoming very important. You hear a lot, all these companies have unbelievably large token spend and they're having to cut back on their Opus usage because they burned through it all in a couple months, and that is going to continue to be a problem because, yes, the cost of an individual token has come down drastically. You can look at the difference between GPT-4 when it first launched and GPT-5.5. It is much, much cheaper per token, but at the same time the amount of tokens in an individual session has gone up exponentially as well. So we're spending more on a total session, and so the ability to bring in-house or at least work with partners to ensure that you are controlling your cost and you're not at the whims of when a company releases a newer model that might be better but also more expensive, they might deprecate a model. Owning that and being sure that, the same way you know what input goes in, you know what output is going to come out. In the same way, when an input goes in, how much it's going to cost, having assurance on that is really important too. Yeah, maybe one thing to add to this is I like this new term of instead of speaking about token maxing, speaking more about outcome maxing. Ultimately, you want to have more than a dollar worth of value come out of a dollar of input, and I think there's a sort of Jevons paradox of if you can create more value for your flop, for your GPU, I think this is how you'll get the most adoption also of agentic models. If they can be able to create as much value as possible, I think the cheaper those models get, the more usage they'll get for those specific use cases. It's funny you say that. I think a lot of us in this room, but especially on this panel, believe this to be so, that the most meaningful AI applications in the next couple years, even this year, are the ones where the harness and the model and the product all blend together. If you think back to, at least for me, the first truly game-changing agentic experience I had was with deep research from OpenAI, and that was because they spent a tremendous amount of time doing reinforcement learning on 03 with test-time compute to do these longer-running research tasks that people had tried previously, but they were just doing a for loop over search, whereas I kept coming back to deep research. And you saw for a very long time that OpenAI and Anthropic and Google, when they'd release a new product, they'd release a custom version of their model for that product. And if they're doing that, if their off-the-shelf GPT 5 isn't good enough for their Atlas web browser, why should it be good enough for our apps? And that's why I appreciate the work that Vincent and Nvidia If you think back to, at least for me, the first truly game-changing agentic experience I had was with Deep Research from OpenAI. That was because they spent a tremendous amount of time doing reinforcement learning on O3 with test-time compute to do these longer-running research tasks that people had tried previously, but they were just doing a for loop over search. Whereas I kept coming back to Deep Research, and you saw for a very long time that OpenAI and Anthropic and Google, when they'd release a new product, they'd release a custom version of their model for that product. If they're doing that, if their off-the-shelf GPT-5 isn't good enough for their Atlas web browser, why should it be good enough for our apps? That's why I appreciate the work that Vincent and NVIDIA are doing, for giving people the tools to customize their own model and allowing us to focus on how we get a good model to start from. So it really is extremely important, as you look at developing applications and experiences over these next few years, that you're taking into account that you can make the model do something that maybe your harness isn't fully able to do alone. Well, that's something I think that's really important to reiterate, right? Memotron's great. I love it. Trinity's great. I love it. We design a model that's supposed to be as good as it can be across a number of harnesses, right? You can see this in the technical report. The idea is we want the model to work as well as it can in PI compared to Hermes, compare whatever you're using, right? But you're not using all of these tools at once. You're using one of these tools. When you have open models, you can do things like Noose Research can create a post-train of whatever model for their harness, right, that will be extra good. This thing from, I can't remember who originally wrote it, but this idea of the mismanaged genius, right, where we're leaving a lot of important capability on the table because we're not fitting the models into the harness, right? You can do a bunch of stuff with closed models, like you can change your prompts and your skills and all kinds of other neato things, right? But nothing will let you get the level of customization or customizability that you can achieve with open models. I think that is something that is going to become increasingly more important, especially thanks to folks like the others on the panel, where I can just straight drop my favorite coding and agent environment, spin up the CLI, and suddenly my model feels way better with very little effort, right? That is something that is already at our fingertips, and it is only going to get easier and easier as time goes on. Yeah, it's maybe also the most concrete, even for the builders in the audience. If you want to build the next Cloud Code, the next Cursor, Perplexity, I think the easiest way to get started is take the best open model and then post-train it on your harness that you care about, right? Create a product that is truly AI-native, right? I think this is one of the most exciting unlocks for builders out here. That's a big thing about the framing of control. It is similar to when the cloud explosion started to happen in the late 2000s, early 2010s, and the social media world took off and apps became extremely popular and more cloud-based and managed by these bigger companies. The data that they were collecting from you, whether anonymized or not, was how they were monetizing their platform through ads or in other ways. A very similar thing has always been happening, but I think it's becoming clear to people in the space that the data that people get from you using these models is how these companies largely make their models better, whether it's through actually training on that data or by using it as a signal for what data they go out and find or generate to train. With closed models, there are terms of service that keep you from being able to, well, I could get into a discussion about what terms of service is, an agreement between you and the provider. It's not illegal. Anyway, you shouldn't be training on a Claude Opus output or a Fable output or a GPT-5 output, and they do a lot to try to obfuscate to make that not great for you. If you're using an open model, you can save all of those traces, all of those traces of you using it inside of your harness. That will allow you, over time, if you say, hey, I want to go train a custom model, you can take all of that and again either use it to directly do fine-tuning on a smaller model so you're not spending as much, or to have a model help you find signals so you can go out and use verifiers or NeMo RL or NeMo Gym to create these environments so that you can hill-climb and make your models better. As much as using open models is owning your stack, owning your intelligence, it's also owning your outputs, right? Owning your data. That's going to be extremely important too. I do want to plug the license for a second. AI is very different than traditional software, which is why recently Nemotron, as well as Trinity, I know, has adopted the Open MDW, Model Data Weights, license. The idea is we need a way to really make it very clear in the license that you can use the outputs to produce a model. You can use the outputs to train, right? All of these TOSs and stuff like that that have language that is meant to dissuade you to do that, we wanted to make sure there's a license that exists that not encourages you but makes it crystal clear that it is permitted, it is permissible. I think the licenses maturing to fit the use case better should be an extremely positive signal for the way that the ecosystem is thinking about open models, to the fact where even the lawyers are on board. Do you know how much lawyers cost? Yeah, as we move on to the third topic, which here is, of course, optimization, something that we've implicitly said but haven't said quite explicitly yet is effectively that I think that for a lot of people, there's this preconceived notion that when you're deciding to use an open model for whatever the use case, there are trade-offs that come in the form of performance at the benefit of getting things like data sovereignty and so forth. What we are now talking about is that with the right customization and optimization, depending on the use case and the harness that you're applying it to, you can actually exceed and build the model against the tool to get better performance than even frontier models. I'd love to hear a little bit more because I think that another thing that we would probably agree on is that the current level of intelligence already has so much left to diffuse into society. Where are those areas where that diffusion is happening in specific industries? I know, for example, things around fundamental pieces of infrastructure, whether it's browser use and how you can start to train models to be able to use the internet better when looking at a computer, and so on and so forth. What are some of those examples where you think that the post-training of open models will see new use cases unlocked compared to just paying full price for the frontier models? Yeah, I think I can start on this. I think the power that we've seen with a lot of different customers is this idea that if you want to make a specific use case work, we can take the example of if you want to figure out a way that agents can actually automate your task, the most concrete way you can do it today really is build an RL environment for that use case, train on that, and then deploy it into production with those users, right? Let's say with a million accountants that then now use this agent to ultimately get it towards full autonomy. It's almost like Tesla's levels toward full autonomy, where you need to do the last mile of actually training for that specific use case, but then also deploying it to those specific users, right? There's a reason why chatbot isn't good at self-driving, because it's not trained on that. It's not deployed into that context, right? I think it's the same even for these specific knowledge work use cases. If you want to have the perfect financial agent, it's much more likely that you'll be able to get there if you have an RL environment for that use case, if you deploy it into production, for example, as a bank to millions of customers, than if there's one god model chatbot. I think this is what we've seen now with a lot of verticals and customers, that there's a huge unlock there to really go into these specialized you need to deploy it into, do the last mile of actually training for that specific use case, but then also deploying it to those specific users, right? So there's a reason why a chatbot isn't good at self-driving, because it's not trained on that, it's not deployed into that context, right? And I think it's the same even for these specific knowledge work use cases. But if you want to have the perfect financial agent, it's much more likely that you'll be able to get there if you have an RL environment for that use case, if you deploy it into production, for example as a bank, right, with two million customers, than if there's one god model chatbot. And I think this is what we've seen now with a lot of verticals and customers, that there's a huge unlock there to really go into these specialized domains, post-train on them, deploy into them, and then continuously learn from production traces. So we work with some big AI natives on things like computer use. We're ultimately having millions of traces from production data, which really can help you continuously improve those agents. And I think this applies to almost every single domain, and I think it's the white pearl for the AI application builders and their startups to actually have a huge opportunity to build their moats and to get to this data flywheel of specialized models, even in a broader sense, just going after, let's say, computer use agents, right? And I think, yeah, this is something where we're just seeing a lot of movement, especially now with open models catching up to the frontier. And I think the other piece is optimization, where I think GM is a great example, also turning to your name, or Tron, is you have the whole ecosystem driving down the cost and optimizing it further, right? We are able to work very closely with all the teams here, but then also deeply with NVIDIA and with teams like VLM to really drive down the cost and make, for example, inference and training for models like GM or models like Trinity or Nematron extremely efficient. So you can drive down the cost further and further. And I think this is something you don't obviously get with the closed APIs, where they have a huge margin on top. They might drive down the optimization, but then might not pass through those savings. So I think, in general, the open models are only getting, through the open ecosystem, more and more efficient and cheaper and cheaper to run and train on. So I think there's this element as well. I think, too, a couple things that I want to make sure we're very clear about is you, like most people, probably do not need frontier-level intelligence for 90% of their tasks, right? Not to say that you're not doing cool, smart stuff, not to say that I'm sitting here trying to do not cool, smart stuff, but a lot of the time these models are just overkill, or they have this really smooth capability horizon that means they're also quite good at chemistry. But most people are using models to do one or two things very well, and open models let you choose those one or two things and then make the model just very good at those things at the expense of almost everything else. And that is great. That's exactly what we should be doing, right? To use this model that is hyper-generalized and able to perform well across 90 different axes is dope and cool, but it is not really using the model effectively. It makes sense for someone who is trying to ensure that everyone can use this one endpoint to do their task, but it makes much less sense when you're a person who's trying to do that task yourself. What was just said about efficiency is also deeply true, right? I mean, the idea that you are all here at a local AI summit, presumably you are running AI locally, presumably you would like it to be faster and better, and presumably many of you are quite cracked engineers, right? This is a whole room of people who are going to contribute in some small part to making the ecosystem just a little bit faster, just a little bit more efficient. And while it's true that closed companies can afford to hire great, amazing teams of people, as we saw with Linux over the whole time that it's existed, right, Linux is the thing that runs the internet. It runs networks, it runs all of these services that require it to be hyper-optimized in a way that I think you can only get when you have people who are trying to run as resource-constrained as possible. And all that to wax poetic and say this idea that local AI and open models and the most efficient version of the model ecosystem, is necessary to do it in the open. I think it's, in fact, not possible to do it behind closed doors because you're shutting too many people that could make that one small contribution out of the room. I think it's important to state, too, that I don't think any of us agree that, or are of the mind that, closed models or frontier, what OpenAI and, just to name names, Anthropic and others are doing, is not extremely beneficial, or that you shouldn't use them. I certainly use those models near every day. It's just, where does it fit in the future of this ecosystem? And just like Chris is alluding to, you can think of open models and self-hosted or owned intelligence or LLMs as the Linux layer, which you're beginning to see take place. Linux runs enterprises, it runs the cloud. We're seeing a very similar thing take place with hyperscalers and neo clouds and providers like Fireworks, Together, Base Ten, Modals. But just like Macs are one of the best ways to get work done individually, in the same way that maybe using OpenAI and ChatGPT is the best way for you to do the vast majority of simple check my email, help me rewrite, check for grammar, those kinds of things. It's accessible, it's easy, and for your average consumer and individual it's pretty cheap if you're using the $20 a month plan. In the same way that Microsoft helps the world of medium-size to large businesses run on Microsoft and Windows because they're not as expensive as getting everybody a Mac, you're going to see a similar world play out there for some closed and open to address themselves. So it's all an ecosystem. I don't want to give the impression that I think any time you log into ChatGPT or Claude, that you're committing a sin, only that as you are, this is a conference for AI builders, AI engineers, as you're looking at the best way to engineer your product or your service, there is another layer you can go down into, and it's becoming way more accessible than it used to be. I think that's a great point, and I think that the relationship between closed frontier models and open models will be one that is constantly there, right? I think you'll never get the headlines to apply nuance and say that both will coexist and gain more usage and are going to be useful to everybody, but that is the de facto state that not only currently exists but will continue to exist. I want to spend the last few minutes here to really give the audience something that only you guys potentially can answer. Oftentimes I reflect about my time at NVIDIA, and I feel as though I have a clear vision outside into, there's definitely still a fog of war out there, but I have a vantage point that many people don't have, and you guys, because of the positions you are in, as well. So what is the thing, if we are looking forward towards AI Engineer World Fair 2027, that you think, if you were to make a bold prediction, let's say let's not be conservative around the intelligence in the open source and ground it with some frame of reference, what do you think we can look forward to by this time next year? I think one key aspect obviously that people are closely tracking is capabilities of open frontier models, and I think they'll keep being very close to the general frontier, potentially even with the speed of those closed frontier models slowing down, I think they'll catch up even more. And I think the most concrete thing that I think will be very exciting is seeing the world move from check bots and now coding agents to just general knowledge worker agents over the next 12 months, right? To see more and more of everyone across every knowledge worker domain adopt agents in their workflows, which I think developers with coding agents have probably done better than any other domain in the world. But I think we'll see over the next 12 months a lot of domain-specific knowledge worker agents, but then also I think domains like computer use agents and others will take off, in a similar way that coding capabilities of open frontier models, and I think they'll keep being very close to the general frontier, potentially even with now and the speed of those close frontier models slowing down. I think they'll catch up even more, and I think the most concrete thing that I think will be very exciting is seeing the world move from check bots and now coding agents to general knowledge worker agents. I think over the next 12 months, to see more and more of everyone across every knowledge worker domain adopt agents in their workflows, which I think developers have with coding agents have probably done better than any other domain in the world. But I think we'll see over the next 12 months a lot of domain-specific knowledge worker agents, but then also I think domains like computer use agents and others will take off in a similar way that coding agents have taken off. I think we'll just see, in some ways you could say, almost the general intelligence for the knowledge and digital domain before then hopefully maybe moving on to physical. And I think very concretely, in 12 months, I think it's pretty likely that we'll have better than faber metals-level capabilities in open models, and I think this is a huge opportunity to ultimately enable a huge crop of new AI startups and companies. So in some ways, you want to almost write the levels of capabilities. To some extent, Cursor really only took off when Opus was good enough to do coding, right? So this is when Cursor inflected, and I think we'll see hundreds of these inflections for startups getting started now, over the next year or two, once open model. And I think we've seen this literally a month ago with jm 5.2. I think it was legitimately one of those moments when people were like, okay, this is now similar to the Opus inflection point. It's an inflection point for models to be extremely strong and ultimately enable a ton of new businesses, and I think this will only continue. Not so bold prediction is that Prime Intellect and RC are going to have a combined valuation of a trillion dollars. That's obvious. But I think that this is going to be a huge year for this. This is probably going to be the most consequential year for the future of how AI gets distributed. The fable and gpt 5.6 embargo, if you will, has left a lot of open questions, no pun intended, about open models and how this intelligence gets distributed and at what capability level it starts to be politicized and kept back. So sovereign intelligence is going to be very important. I think that if you were to take the general population of AI users, people that are using it every day, so upwards of a billion to two billion people, if you include ChatGPT and Google and whatnot, maybe 0.0001 percent have ever used an open model, or run it themselves with the multitude of different tools. And I would hope that the work that we're doing, and the community is doing, and the way that we're advocating for open science and open models and open discussion, that's probably been the most frustrating thing about the last couple weeks, is that all of these conversations around capabilities and who gets to use them and who doesn't have been happening behind closed doors. My hope, and I hope that I can predict, is that we will be able to have 10 to 15 percent of people that have ever used AI have used a model locally on their system, and that that becomes a very important part of ensuring that you have access to what you need. So it's a prediction. It's also something that I know all of us up here and you out there are going to try to fulfill, and I hope that we can continue to advocate for that, because if we're quiet, we just let things play out the way they are. Open models will be put under the microscope in the context of untrustworthy, unsafe, and as much as there's work and vitriol and weaponized terms being out there advocating for that, we need to be combating as much of that, if not more, with the reasons that it deserves to exist. Yeah, I just could not plus infinity what the last part, what Lucas said, more. I think this is going to be the most consequential year for open intelligence that will, at least from where I sit, determine the future of a summit like this, right? I think it has a potential to look very different in two radically opposed ways. As for bold predictions, I think that we will not be needing to go to an API for most of the tasks that we all do each day with AI. I think it's likely to assume that you'll be running a model that is sufficiently capable in, let's call it, day-to-day work on your MacBook within the year. It's already extraordinarily close, so maybe not that bold of a prediction, to be honest with you. I also think that we're going to continue to see models become the future of AI, so not model, right? Swarms of, or specialized systems of models, I think are going to be increasingly important. And lastly, on the open model front, I think we're going to see some very large architecture shifts, especially as we start to crack things like diffusion models for text a little bit more, to get us models that are better suited for the hardware that we have in our houses. And then last meme one, I think you're going to buy computers with agent operating systems on them instead of traditional operating systems, similar to you buying a Spark preloaded with Hermes or whatever. I think that's likely to occur. I also predict that come September, when the next iPhone comes out, you're going to get a lot of texts from family members asking about this magical new series. So a lot of people who have not engaged with AI are about to in a very real way, and the response to that is going to be very, very cool. Just like when Deep Seek came out, and I'm sure a lot of you all got questions about what's this Deep Seek thing, there'll be another one in September. Get ready for it. Yeah, I think this might actually be one of the consequential almost unlocks, right? I think the combination of your models getting good enough, as well as the on-device compute getting strong enough to serve the equivalent of today's frontier models, in a year or two, if you can run opals at decent speeds on your phone or laptop, I think the majority of humanity will probably run local models. And I think this probably applies more to the consumer than to the heaviest enterprise agents, but I think it seems pretty likely to me that it would be this inflection point even then, almost like it's a new platform shift where you can almost tap into the local compute of a phone or laptop and then start a next generation of AI-enabled applications that can ultimately really leverage the local compute of device on-device compute. You can run a four billion parameter model on your phone right now that is way more useful than gpt4 was when it came out, and I think it's important for us to continue to focus on how do we best utilize that in the most meaningful way possible, as well as chase the newer capabilities that will come from things like drug discovery and scientific exploration. It's going to be a fun couple years. Absolutely. If I were to try and summarize, I think that we're going to learn a lot more about how these local models are built incredibly in the systems and their relationships as they interact with frontier models over the next panels. But this panel really shows that I think that we are at an inflection point to where, if you think about how not even a short six years ago, AI was really for the research crowd and not really many people cared about it, and then of course it came into the public consciousness with ChatGPT. But now there's this next thing, which is that open source is now really starting to enter the public consciousness, but very few people have touched and played with it and have had that aha moment. And it sounds like we have the potential to do that this year and sort of guide the future wisely. But ultimately, it's up to a lot of the builders in this room as well to leverage that and represent this important inflection point that we're in on the side that hopefully brings more intelligence to all of us, which is ultimately, I think, what everyone in this room would agree is the direction of progress. And everyone has said it on a panel previously, so I'll just also say it, which is that, and both of you have already said it, in fact, you guys are extraordinarily important to this goal. Every one of you who is in this room, and your friends, and whatever communities you're part of, without you guys, we lose the fight, right? Touched and played with it, and I have had that aha moment, and it sounds like we have the potential to do that this year and guide the future wisely. But ultimately, it’s up to a lot of the builders in this room as well to leverage that and represent this important inflection point that we’re in—on the side that, hopefully, brings more intelligence to all of us, which is ultimately what everyone in this room would agree is the direction of progress. Everyone has said it on a panel previously, so I’ll just also say it, and both of you have already said it. In fact, you guys are extraordinarily important to this goal—every one of you who is in this room, and your friends, and whatever communities you’re part of. Without you guys, we lose the fight, right? So thank you for showing up, and I can’t wait to see what we all build together. And with that, thanks, Vincent and Lucas RC and Prime Intellect, and of course Chris from Nvidia, always, of course. Yeah. Yeah. Yeah. Yeah. Yeah. Yeah. Yeah. Yeah. Yeah. Yeah. Yeah. Yeah. Yeah. Yeah. Yeah. Yeah. even more so. The benefit of open models is that we can very easily validate what is inside of them. There is a whole bunch of files with a whole bunch of matrices in there, and you can view them, and you can see the code that is running these models. You have implementations from Prime RL, VLLM, SG-Lang, the provider themselves. These models are inherently trustworthy. You know much more about what's going on when you hit and talk to these models than you ever will what's going on when you hit an arbitrary API. Now, that being said, certainly there is fear that people can reduce, you know, you can't trust that these models are writing safe code. Well, again, that is the same thing with any model. You need to use your judgment, and you need to make sure that you have the proper, you know, safeguards in place, and you're viewing the outputs of these models as the outputs of an inherently random system that we are working very, very hard to make less and less random. I think that a telling thing is a lot of people said, well, you can't trust Chinese models, you can't trust Chinese models, you can't trust Chinese models. That was often for the last few years meant you can't trust open models. Well, as soon as Anthropic had to put Fable away, and people realized that, oh, our access to these frontier systems might not be universal anymore. There's probably going to be a lot of checks and balances. You had a tremendous number of enterprises and developers and companies start going to these new Chinese models because they could trust that they would always have access to that. And so when it comes down to trust and the way I view that word as it relates to open models, is do I know that what I am running and can I be as sure as possible that when I send something to this model that I am going to get the output that I expect? And the only way currently to be 100% sure that what you are getting is what you are expecting is by hitting an open model, either that you are running yourselves or you're working with a partner like Prime Intellect or RC or Nvidia to validate. So that's my take on the word trust. I think too something you mentioned is like we don't get to know a lot about the data that goes into these models and that's something that I'm really happy you know that we're trying to do. Which is not something, it's not you know the incentives don't exist for everyone to do this right. So it's not something that I think is mandatory or should be mandatory thanks to the things that Lucas mentioned which is that it's rather straightforward to validate what data did go into a model without seeing the data sources originally. But I'm happy that in Nvidia continues to release data sets along with our models, release environments along with our models to make sure that even if you can't go through the work of determining what went into the model, which you can do with the weights alone for the most part, you have like a spreadsheet you can look at that says here's you know a couple trillion tokens of this data set, a couple trillion tokens here, and I think that that helps to educate people on why it's much easier to trust open models than the models that we don't get access to any of that. It also helps people see what that data looks like. Yeah. You know if you don't have someone releasing it openly when someone says data is going in, I mean data can take many different shapes. You can go to Hugging Face, you can go to Nvidia, or you can go to Prime Intellect or RC's Hugging Face, you can look under our data sets, and you can see exactly what that looks like, and that can help you inform your priors on it. Yeah, I think that trust also you know there's some angle of a reputation do I believe that your your intentions are pure, and I think that a lot of people again in this room believe that intelligence is kind of this next layer of almost you know infrastructure for us to progress as a species, and I believe that everybody should have intelligence. So on that front I want to hand it over to Vincent, because you kind of have this almost like founding thesis that you this this stack should be the open, right the open super intelligent stack you want everybody to have a lot more intelligence. Can you talk about how the this is kind of moving into the era of control, but beyond just data sets, how important it is to have the knobs and dials of the of this industry also be available in an open way for people who are building this. Yeah, like I think it's like a really important point is to set, so make sense like be able to take those open models like customize them be able to like build on top of them, and I think like all the different components that go into it like especially from like the pre-training to mid to post training I think like need to be more accessible right like so more people can also like take those amazing models and like make them work for their specific use cases. So I think when we started like we we also took a look at the whole stack that was out there and try to figure out like what is missing for ourselves to train open models and for like helping our partners to do so and a lot of this was around the RL and post training stack. So we basically went deep into building out like a lot of info around that like around our environment evals around like making much more accessible to do post training also because it's like the most economically viable way to maybe like customize those models to take an open model and to have like a specific event environment and specific domain and dimension that you want to improve it on and this is kind of like what we're really doing with Prime Intellect now is like enabling people to post train specialized agentic models. So being able to take models like Trinity for example for Marcy or NemoTron or others and and specialize them post train them for the use cases that ultimately enterprises care about. So a good example this was like a company like for example ramp was able to like take an open model and like specialize it to automate finance within like a week or two to get like better performance than like opals at a fraction of the cost of haiku and I think really this period of frontier of like being able to create these specialized models that are much better than the frontier but also faster cheaper I think it's like a key thing enterprises care about increasingly it's really like just making it work for their use cases basically. If you go back to trust it's how you can make your CFO trust you knowing exactly how much something's going to cost all the time. That's that is increasingly becoming very important is you hear a lot you know all these companies have unbelievably large token spend and they're having to cut back on their opus usage because they burned through it all in a couple months and that is going to continue to be a problem because yes the cost of an individual token has come down drastically you can look at it you know the difference between GPT-4 when it first launched and GPT-5.5 is is is much much cheaper per token but at the same time the amount of tokens in an individual session has gone up exponentially as well. So we're kind of we're spending more on a on a total session and so the ability to bring in house or or at least work with partners to ensure that you are controlling your cost and you're not at the whims of when a company releases a newer model that might be better but also more expensive they might deprecate a model owning that and being sure that you know same way is what use what output goes in you know what outputs going to come out in the same way when it you know an input goes in how much it's going to cost having assurance on that's really important too. Yeah maybe like one thing to add to this is like almost like I like this new term of like instead of speaking about token maxing you know speaking more about like the outcome maxing of like yeah ultimately it's like you you want to have like more than a dollar worth of value come out of like a dollar of input and I think there's a sort of like Jevons paradox of like if you can create more value for like your flop for your GPU basically I think this is sort of like how you'll get like the most adoption also of like agentic models like if they can like be able to create as much value as possible I think the cheaper those models get the more usage they'll get like for those specific use cases. It's funny you say that I have a and I think a lot of us in this room but especially on this panel believe this to be so that the most meaningful AI applications in the next couple years even this year are the ones where the harness and the model and the product they all kind of blend together If you think back to at least for me the first like truly game changing agentic experience I had was when deep research from open AI and that was because they spent a tremendous amount of time doing reinforcement learning on 03 with test time compute to do these longer running research tasks that people had tried previously but they were kind of just doing a for loop over search whereas I kept coming back to deep research and you know you saw for a very long time that open AI and anthropic and google when they'd release a new product they'd release a custom version of their model for that product and if they're doing that if they're off the shelf GPT 5 isn't good enough for you know their atlas web browser why should it be good enough for our apps and that's why I appreciate the work that you know vincent and nvidia are doing for for giving people the tools to customize their own model and allowing us to focus on how we get a good model to start from so it really is you know it's it's extremely important as you look at developing applications and experiences over these next few years that you're taking into account that you can make the model do something that maybe your harness isn't fully able to do alone well that's something I think that's really important to just like reiterate right I mean like memotron's great I love it trinity's great I love it like we design a model that's supposed to be as good as it can be across a number of harnesses right you can see this in the technical report the idea is like we want the the model to work as well as it can in pi compared to you know her means compare whatever you're using right uh but like you're you're not using all of these tools at once you're using one of these tools and so when you have open models you can do things like uh noose research can create a post train of whatever model for their harness right that you know will be extra good and you know this this uh this thing from from from uh you know i can't remember who who originally wrote it but this idea of like the mismanaged genius right where we're leaving a lot of like uh a lot of important capability on the table uh because we're just not we're not fitting the models into the harness right you can do a bunch of stuff with closed models like you can change your prompts and your skills and all kinds of other neato things right but nothing will let you get the the level of customization or customability that you can achieve with open models and i think that is something that is going to become increasingly and increasingly more important especially thanks to folks like the others on panel where you know i can just straight drop like my favorite coding and you know agent environment spin up the cli and suddenly my my model feels way better uh with very little effort right like that is that is something that is uh already at our fingertips and it is only going to get easier and easier as as time goes on yeah it's maybe also the most concrete like even for the builders and audience like call like if you kind of want to build the next like cloud code the next like cursor perplexity i think the easiest way to get started is like take the best open model like and then post train it on your harness like that you care about right like basically like create a product that is like truly ai native right and i think this is sort of like i think one of the most exciting like unlocks for builders um like out here that's a big thing about the the framing of control is is uh similar to like you know when when the cloud um explosion started to happen in the in the you know the late uh 2000s early 2010s and the um social media world kind of took off and apps became extremely popular and more you know cloud based and and you know managed by these these bigger companies you know the data that they were collecting from you whether anonymized or not was how they were monetizing their platform through ads or or or in other ways uh and a very similar thing has always been happening but i think it's becoming clear to people in in the space is the the data that people get from from you using these models is how these companies largely make their models better whether it's through actually training on that data or um by using it as a signal for what data they they go out and find or generate to train and uh with closed models there are terms of service that keep you from being able to um well i could get into a discussion about what terms of service is an agreement between you and the provider it's not illegal uh anyway but you you shouldn't be training on a you know a claude opus output or a fable output or a gpt5 output and they do a lot to try to obfuscate to to make that not great for you if you're using an open model you can save all of those traces all of those traces of you using it inside of your harness uh that will allow you over time to if you say hey i want to go train a custom model you can take all of that and again either use it to directly do like fine tuning on a smaller model so you're not spending as much or to have a model help you find signals so you can go out and use verifiers or nemo rl or nemo gym to create these environments so that you can hill climb and make your models better so um as much as using open models is like owning your stack owning your intelligence it's also owning your outputs right owning your data that's going to be extremely important too i do want to i do want to plug the license for a second so uh so ai is very different than traditional software uh which is why recently uh nematron as well as uh trinity i know uh has adopted the open mdw model uh data weights uh license uh the idea is like we need a way to really make it very clear in the license that you can use the outputs to produce a model you can use the outputs to train uh right all of these tos's and stuff like that that that that have language that is meant to dissuade you to do that uh we wanted to make sure there's a license that exists that not encourages you but makes it crystal clear that it is it is permitted it is permissible uh and and i think you know the licenses maturing right to fit the use case better should be extremely positive signal uh for for the way that the ecosystem is thinking about open models uh to the to the fact where even even the lawyers are on board do you know how much lawyers cost yeah it's uh as we move on to the you know kind of the third topic which here is of course optimization um something that you know we've implicitly said but haven't said it quite explicitly yet is uh effectively that i think that for a lot of people there's this preconceived notion that when you're deciding to use an open model for whatever the use case um there there are the trade-offs that come in the form of performance at the benefit of getting things like you know maybe data sovereignty and so forth um but what we are now talking about is that with the with the right customization and optimization depending on the the use case that you're and the harness that you're applying it to you can actually exceed and build the model against the tool to to get better performance than even frontier models um i'd love to hear a little bit more about the because i think that uh another thing that we would probably agree on is that the the current level of intelligence um already has so much uh left to diffuse into society and so where are those areas where that diffusion is is happening in the specific industries i know for example things around uh again kind of fundamental pieces of infrastructure whether it's like browser use um and how you can start to train models to be able to use uh you know the the internet better when looking at a computer and so on and so forth so what are the what are some of those examples to where you think that the uh post training of open models will uh see new use cases basically unlocked compared to just paying full price for the the frontier models yeah i think like like i can start on this like i think the power what that we've seen with a lot of different customers is it's really kind of this idea that like if you want to make a specific use case work like we can take the example like if you want to figure out like a way that agents can actually automate your text like the the most like concrete way you can do it today really is like build an rl environment for that use case like train on that and then deploy it into production with those users right like let's say with like a million accountants that then now use this agent to ultimately get it towards full autonomy it's a bit like almost like tesla's levels towards full autonomy where like you kind of need to deploy it into like do the last mile of actually like training for that specific use case but then also um deploying it to those specific users right so like there's a reason why like uh chatbot isn't good at self-driving because like it's not trained on that it's not deployed into that context right and like i think it's the same even for these specific like knowledge work use cases but it's like if you want to have the perfect like financial agent it's much more likely that you'll be able to get there if you have like rl environment for that use case if you deploy it into production for example as a bank right like two millions of customers then if you're there's like one god model chatbot like and i think this is sort of like what we've seen now with a lot of verticals and customers that like um there's like a huge unlock there to um really go into these like specialized domains post train on them deploy into them and then continuously learn from production traces so we work with like some um also big ai natives on things like computer use we're ultimately having like millions of traces from production data really can help you to to um continuously improve uh those agents and i think this kind of applies to almost every single domain and i think it's sort of the the white pearl for like the ai application builders and and they are startups to actually have a huge opportunity to build kind of their modes and and to get to this data flywheel um of like specialized models even in a broader sense like just going after like let's say computer use agents right and i think um yeah this is something where i think um we're just seeing a lot of like movement especially now with like open models catching up to the frontier um and i think the other piece is like optimization where like i think um like gm is a great example like also like turning to your name or tron is like you have the whole ecosystem sort of like driving down the cost and optimizing it further right like like we are very able to work like very closely with all the teams here but then also um like deeply also with uh nvidia and with teams like vlm to really drive down uh the cost and and make the for example influence and training for models like gm or like models like trinity or nematron extremely efficient so you can basically drive down the cost like further and further and i think this is something you don't obviously get with the close apis where like they have like a huge margin on top like they might drive down the optimization but then might not pass through those savings so i think in general like the open models are only getting through the open ecosystem like more and more efficient like uh and and cheaper and cheaper to run and train on so i think there's like this element as well i think too like a couple things that i i i want to make sure we're very clear about is you you like most people probably do not need frontier level intelligence for like 90 of their tasks right like uh like not to say that you're not doing cool smart stuff not to say that i'm i'm sitting here trying to do uh not cool smart stuff but like a lot of the time uh these models are just overkill or they have like this really smooth uh you know uh capability horizon that uh means they're they're also quite good at chemistry but like most people are using models to do one or two things very well uh and open models let you choose those one or two things and then make the model just very good at those things at the expense of at the expense sorry of almost everything else and that that that is great i mean that's exactly what we should be doing right uh to to to use this model that is hyper generalized and able to uh you know perform well across like 90 different axes is is dope and cool but it is not really you know using the model effectively it makes sense for someone who is trying to ensure that everyone can use this one endpoint to do their task but it makes much less sense when you're a person who's trying to do that task yourself it what uh what what was just said about efficiency is also deeply true right uh i mean the idea that you are all here at a local ai summit presumably you are running ai locally presumably you would like it to be faster and better and presumably many of you are quite uh quite cracked engineers right uh this is a whole room of people who is going to contribute in some small part to making the ecosystem just a little bit faster just a little bit more efficient and while it's true that closed companies can afford to hire great amazing teams of people as we saw with linux over the whole time that it's existed right linux is the thing that runs the internet it runs networks it runs all of these services that that require it to be hyper optimized in a way that i think you can only get when you have people who are trying to run as resource constrained as possible and uh all that to wax poetic and say this idea that like local ai and open models and the most efficient version of the of the uh model ecosystem uh is is necessary to do it in the open i think it's in fact not possible to do it behind closed doors because you're shutting too many people that could make that one small contribution uh out of the room i i think that that it's important to state too that i don't think any of us agree that or are of the mind that closed models or frontier you know what open ai and throught just to name names you know ananthropic and others are doing it is uh not extremely beneficial or that don't use them um i i certainly uh i use those models near every day it's it's just that it's where does it fit in the uh in the future of this ecosystem um and just like you know chris is alluding to you can think of open models and self-hosted or kind of owned intelligence or llms as like the linux layer which you're beginning to see kind of take place linux runs enterprises you know it runs the cloud we're seeing a very similar thing take place with hyperscalers and neo clouds and providers like fireworks together base tens modals um but just like max are one of the best ways to get work done individually uh in the same way that may be using open ai and chat gpt is the best way for you to do the vast majority of simple check my email uh help me rewrite you know check for grammar those kind of things it's accessible it's easy and for um you know your average consumer and individual it's pretty cheap um if you're using like the 20 a month plan uh in the same way that uh you know microsoft um helps you know the the world of medium size to large businesses run on microsoft and windows because uh um you know they're not as uh expensive as getting everybody a mac um and you're going to see a similar uh world play out there for for some closed and open open for address themselves so um it's all an ecosystem you know i don't want to give the impression that uh you know i think anytime you log into chat gpt or cloud that that you're committing a sin only that as you are you know this is a conference for ai builders ai engineers as you're looking at the best way to engineer your product or your service that uh there is another layer you can go down into and it's becoming way more accessible than it used to be i think that's a great point and i think that the relationship between closed frontier models and open models will be one that is uh it's constantly there right i think that we have it's never you you'll never get the headlines to apply nuance and say that both will coexist and gain more usage and are going to be useful to everybody but that is kind of the de facto state that that will not only currently exist but will continue to exist um i want to spend the last few minutes here to uh really give the audience something that uh only you guys potentially can answer um oftentimes i reflect about my time at nvidia and i think i i feel as though i have a a clear vision outside into um there's definitely still a fog of war out there but i have a vantage point that many people don't have and you guys because the positions you are in as well and so what is the the thing if we are looking forward towards ai engineer world fair 2027 um that you think if you were to make a bold prediction let's say let's not be conservative around uh the intelligence um in in the open source um and ground it with some frame of reference what do you think uh we can look forward to by by this time next year i think one key um aspect obviously that people are closely tracking is like um sort of like just like capabilities of um open frontier models and i think they'll they'll keep being very close uh to the general frontier potentially even like like with like now and the the speed or like the uh off those close frontier models like slowing down i think like they'll catch up even more and i think the the most concrete thing that i think will be very exciting is like seeing the world move from sort of check bots and now coding agents to like just general knowledge worker agents i think over the 12 next 12 months right like to see more and more of like kind of like everyone across every uh knowledge worker domain like adopt agents in their um workflows which i think like developers have with coding agents have probably done better than any other domain in the world but i think we'll we'll see over next 12 months like um a lot of like domain specific like knowledge worker agents but then also i think domains like computer use agents and i think others will take off like i think in a similar way that like coding agents have taken off i think we'll just see almost like in some ways you could say like almost like the uh general intelligence for like the knowledge and digital domain um before then hopefully maybe moving on to physical and i think very concretely like i think we'll like in 12 months i think it's pretty likely that we'll have like better than uh faber uh metals level capabilities and open uh models and i think this is like a huge opportunity uh to ultimately enable a huge crop of new like ai startups and companies so in some ways like you want to almost like write the the levels of capabilities like to some extent like cursor really only took off when like opus was good enough to do coding right so like this is when like cursor inflected and i think we'll see like hundreds of these inflections for like startups getting started like now like in over the next year or two um once like open model and and i think we've seen this literally a month ago with like jm 5.2 i think like uh it was i think legitimately one of those moments when people were like okay this is now like similar to the opus inflection point it's like an inflection point for models to be like extremely strong and ultimately enable a ton of new businesses and i think this will only continue like uh not so bold prediction is that uh prime intellect and rc are going to have a combined valuation of a trillion dollars uh that's obvious uh but but but i think that this is going to be a huge year for um uh this is probably going to be the most consequential year for like the future of how ai gets distributed um the the the fable uh and gpt 5.6 um you know uh uh uh embargo if you will um has left a lot of open questions um you know no pun intended about open models and and where um how this intelligence gets distributed and at what capability level it starts to be uh politicized and and and kept back and so sovereign intelligence is going to be very important i i think that um if you were to take you know the general population of ai um users people that are using it every day so you know upwards of uh a billion to two billion people if you you know include chat chp t and google and whatnot uh maybe 0.0001 percent have ever used an open model you know i think it or you know run it on themselves with with the multitude of different tools um and i would hope that the work that we're doing and the community is doing and the way that we're advocating um for open science and open models and open discussion right that's probably been the most frustrating thing about the last couple weeks is that all of these conversations around capabilities and who gets to use them and who doesn't have been happening behind closed doors uh my hope and and and and i hope that i can predict that we will be able to have uh uh 10 to 15 percent of people that have ever used ai have used a model locally on their system and that that becomes a very important part uh of ensuring that you have access to what you need um and and so it's a prediction it's also something that i know all of us uh up here and and you out there are going to try to fulfill and i hope that uh we can continue to advocate for that because if we're if we're quiet we just let this uh things play out the way they are uh open models will you know will will be put under the microscope um in the context of untrustworthy unsafe uh and as much as there's work and and um vitriol and weaponized terms being out there advocating for that we need to be combating as much of that if not more with the reasons that it deserves to exist yeah i just uh i could not uh plus infinity what what the last part what lucas said more i think this is going to be the most consequential year for uh open intelligence that uh will it at least from where i sit determine the future of uh of a summit like this right uh i think it has a potential to look very different in two uh radically opposed ways uh as for bold predictions i think that we will not be needing to go to an api for most of the tasks that we all do each day with ai i think it's likely to assume that you'll be running a model that is sufficiently capable in let's call it day to day work uh on your on your macbook uh within the year uh it's already extraordinarily close so not maybe not that bold of a prediction to be honest with you uh i also think that we're going to continue to see models uh become the the future of ai so not model right uh swarms of or uh specialized systems of models i think are going to be uh uh increasingly uh important uh and lastly on the open model front i think we're going to see some very large architecture shifts especially as we start to crack things like diffusion models for text a little bit more uh to to get us models that are uh better suited for the the hardware that we have in our houses and then last meme one i think you're going to buy uh computers with agent operating systems on them instead of traditional operating systems uh similar to like you buying a spark preloaded with with hermes or whatever uh i i think that's that's likely to occur i also predict that come september when the next iphone comes out you're going to get a lot of text from family members asking about this magical new series so a lot of people who have not engaged with ai are about to in a very real way um and the the response to that's going to be very very cool and uh so just like when deep seek came out and i'm sure a lot of you all got questions about what's this deep seek thing though there'll be another one in september get ready for it yeah i think like this might be actually one of the consequential almost like unlocks right it's like i think combination of like basically your models getting good enough uh as well as like the on-device compute getting strong enough to serve the equivalent of like today's frontier models right like in a year or two like like basically if you can like run opals like at decent speeds on your like phone or laptop i think the majority of humanity will probably like run local models like and i think this probably applies more to the consumer than to the heaviest like enterprise agents but i think uh it seems pretty likely to me that like it would be this inflection point even then like almost like it's a new platform shift where like um you can almost like tap into the local compute of a phone or laptop and then like start a next generation of almost like ai enabled applications without like that it can ultimately really like leverage the local compute of like um device on device compute you can run a four billion parameter model on your on your phone right now that is way more useful than gpt4 was when it came out and i think it's important for us to continue to to focus on how do we best utilize that in the most meaningful way possible uh as well as chase uh the newer capabilities that will come from things like you know drug discovery and and uh um and scientific exploration it's it's going to be a fun couple years absolutely if i were to try and summarize i think that you know we're going to learn a lot more about how these local models are are built incredibly in the systems and their relationships as they uh you know interact with frontier models over the next panels but uh this panel really shows that i think that we are at an inflection point to where if you think about how uh you know not even a short six years ago it was ai was really for the research crowd and not really many people cared about it and then of course it came uh into the public consciousness with chat gpt um but now there's this next thing which is that open source is now really starting to enter the public consciousness but very few people have touched and played with it um and have had that aha moment and it sounds like we have the potential to do that this year and sort of uh guide the the future wisely um but ultimately it's up to a lot of the builders in this room as well to to leverage that and represent um you know this important inflection point that we're in uh on the side that hopefully brings uh you know intelligence more intelligence to all of us which is ultimately i think what everyone in this room would agree is is sort of the direction of progress and you know everyone has said it on a panel previously so i'll just also say it which is that uh and and both of you have already said it in fact uh like you you guys are extraordinarily important to this goal uh every one of you who is in this room and your friends and whoever what whatever communities you're part of uh without you guys uh we we lose the fight right so thank you for showing up and uh i i can't wait to see what we all build together and with that thanks vincent and lucas rc and prime intellect and of course chris from nvidia always uh always of course yeah yeah yeah yeah yeah yeah yeah yeah yeah yeah yeah yeah yeah yeah