Open Source Is Dead. Long Live Open Source. — Saoud Rizwan, Cline
Description
A compromised release of litellm, a Python package pulling three and a half million downloads a day, sat live for three hours installing a credential harvester for API keys, SSH keys and crypto keys along with a backdoor for remote command execution. It was caught by pure luck: the malware had a bug that crashed Cursor, and a researcher went looking. Saoud Rizwan uses it to mark how far the trust model has fallen. Zig's code of conduct now bans AI from pull requests, issues and even comments, because the team values growing contributors over collecting contributions. curl is weighing the end of a decades old bug bounty as AI generated reports drown it. tldraw closes pull requests on sight. GitHub shipped a switch to disable third party pull requests altogether, which is the feature that made GitHub what it is. He built Cline in the open, and his claim is that the community half of open source is the half that died. The half that survives is open weights, and the case is economic. Cline pitted GLM against Opus on a real bug in their own repository: GLM spent twice the tokens at half the cost, cleaned up dead code and confirmed the build compiled, while Opus finished faster with fewer tool calls but left type errors that broke the production build. Coinbase defaulted its internal gateway to GLM and Kimi and cut AI spend nearly in half. The precedent Rizwan reaches for is Open Compute: Facebook gave away its data center designs, the supply chain standardized on them, manufacturers moved to huge uniform production runs, and the commoditization that followed drove Facebook's own costs down by billions. His closing ask goes to the American labs. Release open weights models, because once the industry standardizes on foreign ones, marginal quality gains may not be enough to bring anyone back. Speaker info: - https://x.com/sdrzn - https://github.com/saoudrizwan - https://cline.bot Timestamps 0:00 Introduction to Cline and the early open source coding agent era. 1:30 The d
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Watch fully
- Core thesis: AI-generated contribution spam and supply-chain risk are eroding open-source communities, but open-weight models will become strategically dominant because they enable cost competition, flexible routing, and industry-standard infrastructure.
- Why it matters: For AI-agent operators, the video argues that durable advantage is shifting from access to a frontier model toward a secure control plane of context, tools, verification, routing, and cost governance.
- Best use: Use it as a strategic argument for building model-agnostic agent infrastructure and treating open-weight adoption, inference economics, and software supply-chain security as linked decisions.
Executive Summary
Saoud Rizwan argues that the traditional social layer of open source is breaking under AI. Cline benefited from being open source because developers could inspect its code, trust its API usage, and shape features such as custom rules and plan mode. But he says maintainers now face AI-generated pull requests, bug reports, issues, and security submissions at a volume that makes open contribution channels increasingly expensive and unsafe to operate.
His evidence includes Zig's prohibition on AI use in project interactions, curl's reported flood of AI-generated bug reports, tldraw's policy of automatically closing pull requests, and GitHub's option to disable third-party pull requests. He pairs this community argument with a supply-chain warning: LiteLLM, reportedly downloaded 3.5 million times daily, was compromised for three hours through a GitHub app and stolen PyPI publishing credentials, distributing credential-harvesting malware and a remote-command backdoor.
Rizwan's larger thesis is that closed frontier-model vendors are subsidizing usage to lock engineering organizations into high-cost application-layer workflows, but that lock-in will weaken as open-weight models become good enough. He contends that agent performance increasingly depends on the surrounding system—skills, rules, context, verification, and quality gates—rather than raw model intelligence alone.
He frames open weights as an economic standardization play analogous to Facebook's Open Compute Project: open designs can reorganize a supplier ecosystem around common infrastructure and lower costs for everyone. His prescription is not fully open-sourcing frontier research, but releasing more usable open-weight models so Western labs retain adoption, influence, safety investment, and competitive relevance while the inference market commoditizes.
Key Takeaways
- Claim: AI is degrading the trust-based contribution model that made many open-source communities viable. | Evidence: Rizwan cites Zig's ban on AI in pull requests, issues, and comments; curl's CEO saying the project is effectively being DDoSed by AI-generated bug reports; tldraw closing all pull requests; and GitHub adding the ability to disable third-party pull requests. | Implication: Projects that accept outside contributions need stronger intake controls, contributor reputation mechanisms, and automated triage rather than assuming GitHub's historical open-PR workflow remains safe or scalable. | Caveat: His conclusion that parts of open-source community are 'dead' is a strategic interpretation, not evidence that all open-source development or public licensing is declining.
- Claim: The AI-era open-source supply chain has an amplified blast radius because a single package compromise can expose high-value developer and enterprise credentials. | Evidence: LiteLLM, described as receiving 3.5 million downloads per day, was reportedly compromised for three hours after attackers used a GitHub app to steal PyPI publishing tokens; the malicious package harvested API, SSH, and crypto keys and installed a remote-command-execution backdoor. | Implication: Ken should treat package provenance, publisher-token isolation, dependency pinning, runtime egress controls, and secrets exposure in agent/MCP environments as first-order architecture concerns—not merely developer hygiene. | Caveat: The incident was detected because the malware caused Cursor to crash when the LiteLLM MCP server ran; Rizwan emphasizes that this was luck, but the transcript does not independently quantify downstream compromise.
- Claim: Closed-model subscription pricing is likely designed to create dependency before vendors recover the true cost of agentic usage. | Evidence: Rizwan cites an anonymous CFO report of an alleged $500 million accidental monthly Claude bill without usage limits, Uber CTO remarks that 95% of engineers used Claude and 70% of committed code came from it at up to $2,000 monthly spend per user, and SemiAnalysis experiments estimating $200 Claude and Codex plans supplied roughly $8,000 and $14,000 of API-equivalent usage respectively. | Implication: Agent programs need spend caps, per-workflow budgets, outage fallbacks, usage telemetry, and multi-provider portability before coding workflows become operationally dependent on one closed vendor. | Caveat: These examples mix anonymous reports, vendor-subscription experiments, and speaker inference; they support a lock-in risk hypothesis but do not prove coordinated future price gouging.
- Claim: For many coding tasks, a lower-cost open-weight model can match or beat a frontier model when the agent system provides sufficient context and verification. | Evidence: In Cline's anecdotal test on a real repository bug, GLM and Opus both fixed the issue; GLM used roughly twice the tokens but cost half as much, removed dead code, and verified the build compiled, while Opus finished faster with fewer tool calls but left type errors and broke the production build. | Implication: Model evaluation should measure verified task outcome, code quality, and total cost—not speed, benchmark rank, or token count alone. Invest in test gates and tool scaffolding that let cheaper models recover reliability. | Caveat: This is one Cline-repository test rather than a controlled benchmark across task classes, models, prompts, and repeated runs.
- Claim: Organizations will increasingly route work to open-weight models based on dollar efficiency, even if those models lack the latest proprietary agent features. | Evidence: Rizwan says Coinbase CEO Brian Armstrong reported defaulting its internal LLM gateway to GLM and Kimi, cutting AI spend by nearly half while token consumption grew. He expects firms to build internal routing layers that choose models by workload economics. | Implication: The strategically valuable layer is an internal model gateway/control plane that can route by price, latency, capability, data policy, and verification requirements rather than binding users directly to a single model application. | Caveat: The Coinbase statement is presented as an attributed example, while the broader adoption forecast is Rizwan's prediction.
- Claim: Open weights can commoditize inference in the same way open hardware standards commoditized data-center components, potentially making proprietary API premiums difficult to sustain for ordinary knowledge work. | Evidence: Rizwan compares the trend to Facebook's 2011 Open Compute Project, which shared server, networking, and cooling designs and enabled standardized manufacturing. He cites estimates of nearly $3 trillion in AI infrastructure spending, more than 100 GW of new data-center capacity by 2030, and a projected 90% reduction in one-trillion-parameter-model inference costs by 2030. | Implication: Build for a market where inference is increasingly interchangeable; differentiation should concentrate in proprietary workflow data, orchestration, evaluations, security, distribution, and customer-specific systems. | Caveat: The cost and capacity figures are forward estimates, and lower inference cost alone does not guarantee parity in reliability, multimodality, tooling, safety, or operational support.
- Claim: Western AI labs should release more open-weight models—not their full research stack—to preserve global adoption and maintain influence over the technology's development. | Evidence: Rizwan warns that Chinese open-weight models such as GLM, Kimi, and DeepSeek could become the default infrastructure standard. His proposed middle ground is releasing weights that are usable and competitive but do not expose training traces or enable easy leapfrogging. | Implication: For investment and platform strategy, track which model families become de facto infrastructure standards; adoption and ecosystem compatibility may matter more than marginal frontier benchmark gains. | Caveat: The assertion that open weights can be released without materially weakening a lab's lead, and that national origin determines safety outcomes, is normative and not substantiated in the presentation.
Detailed Brief
What Cline's origin story says about trustworthy agent adoption
- Claims: Cline's early adoption depended on transparency because users were paying per API request before prompt caching and modern flat-rate coding subscriptions.; Open-source interaction did more than distribute Cline: it let users inspect behavior, connect arbitrary APIs, and shape product features through feedback.
- Evidence: Rizwan says some early users spent hundreds of dollars per day on Cline API usage because it created their first experience of an LLM completing work end-to-end.; He credits the community with informing early Cline features including custom rules and plan mode.; Cline currently offers an open-weight subscription through inference-hosting partnerships while retaining bring-your-own-key support across CLI, VS Code, and JetBrains.
- Caveats: The final product pitch is self-interested: Cline directly benefits if developers conclude that open-weight models and provider portability are preferable.
- Implications: For expensive or high-autonomy agents, inspectability, configurable provider choice, and observable spend can be adoption enablers rather than secondary open-source ideals.; A platform can support open weights without forcing self-hosting; aggregation, routing, and volume discounts are viable operational layers.
The economic analogy: Open Compute to inference competition
- Claims: Rizwan uses Open Compute to argue that publishing a shared standard can lower a sponsor's own costs by expanding and standardizing the ecosystem around it.; He expects inference hosts to compete through hardware specialization, caching, batching, and operational efficiency rather than through proprietary model ownership.
- Evidence: Facebook released specifications, CAD files, and designs for data-center servers, networking, and cooling after competitors had already reduced its proprietary infrastructure advantage.; Rizwan notes that Google announced 32% compute and 68% storage price reductions in 2014, AWS responded within days, and AWS had already made 42 price cuts—his precedent for commoditized infrastructure competition.; He identifies Base10 and Fireworks as examples of providers whose role is to drive inference cost down through infrastructure efficiency.
- Caveats: The analogy abstracts away differences between physical hardware standards and model weights, including licensing constraints, data residency, hardware availability, and model-maintenance burdens.
- Implications: Vendor selection should account for the probability that today's premium model interface becomes tomorrow's commodity serving layer.; Infrastructure economics may favor organizations that preserve the ability to switch hosts and models without rewriting user workflows.
Notable Concepts & Terms
- Open weights: Models whose parameters are available for others to run or serve; Rizwan presents them as the basis for provider competition, lower inference costs, and model-routing freedom.
- Open source community versus open source licensing: The video separates the collapsing trust-based contribution process from the continuing value of freely usable, buildable-on software and model artifacts.
- AI-native development infrastructure: The surrounding system of project skills, rules, context, tools, verification, and quality gates that can make a less capable model produce acceptable results.
- Internal LLM gateway: An enterprise routing layer that centrally selects models and providers; it is presented as the way to optimize cost while retaining flexibility.
- Supply-chain attack: Compromise of a trusted dependency or publishing channel, illustrated by the LiteLLM/PyPI incident and especially dangerous where developer or agent environments hold credentials.
- Open Compute Project: Facebook's open hardware initiative, used as the historical analogy for how shared standards can reorganize supply chains and commoditize infrastructure costs.
- Verified outcome: Rizwan's preferred measure of coding-agent quality: whether the resulting code is clean, compiles, and passes validation, rather than whether the model used fewer tokens or tool calls.
Operator Notes / Why Ken Should Care
- Set enforceable organization-level budgets, quotas, and alerting for every coding-agent provider; test the controls before broad rollout.
- Require agent workflows to pass deterministic compile, test, lint, and security gates before proposed code can merge, regardless of model vendor.
- Build or procure a model gateway with policy-based routing, provider failover, workload-level cost telemetry, and support for open-weight endpoints.
- Benchmark open-weight and closed models on Ken's actual repositories and tasks using verified completion rate, remediation time, cost per accepted change, and failure modes.
- Audit software and MCP dependency controls: isolate publishing credentials, pin and verify packages, minimize exposed secrets, and restrict runtime network/command privileges.
- Review external-contribution intake for AI spam resilience, including trusted-contributor tiers, triage automation, and default-deny policies for sensitive repositories.
Source/Metadata
- Title: Open Source Is Dead. Long Live Open Source. — Saoud Rizwan, Cline
- Transcript words: 2747
- Duration seconds: 1050
- Timestamp note: No usable timestamps or chapter markers were present in the supplied transcript.
Transcript
. Hi, I am Saoud, founder of Cline. I started Cline as an open source project a few years ago. Some of you might know it as the first ever coding agent, back before the Cloud Mac subscription and the Codex subscriptions, when people had to pay for each and every API request, which got extremely expensive. This was before prompt caching became a thing. And so there were people who paid hundreds of dollars a day using Cline. But for a lot of people, it was their first AGI moment. It was the first time they saw LLMs be able to do their jobs end-to-end, and they got hooked. And I don't think Cline would have been as successful as it is if it wasn't open source, because it allowed these developers to inspect their code and trust it and connect to any API, so they could be comfortable with spending so much money on it and know that they weren't getting screwed over. And we were the first to add things like custom rules and plan mode, and a lot of that came from talking to and learning from this really incredible open source community we had around the project. And so, having spent most of my life building open source, it's really heartbreaking to see the broader open source community wither and die over the last two years because of how AI has fundamentally changed everything about software development. GitHub is effectively an archive of slop PRs and issues and security reports, where the sense of community before has turned into this deep skepticism and distrust of each other's responsible use of these tools, because AI coding can be extremely dangerous to a project, and everyone's had to learn that on the fly, but especially open source projects that rely on trusting third parties. And so I wanted to share some examples of how open source has been dealing with AI. So this is the code of conduct for Zig, which is the language that powers Bun. And they essentially ban all use of AI. You can't use it on pull requests or issues or even comments. And the reason for this is that, to them, the core Zig team, they value contributors more than they do the contributions. And so the primary goal for reviewing PRs and things isn't to add new code, but it's to help grow new contributors who can become trusted over time. And AI assistance completely breaks that. This is a post from the CEO of Curl, who says that his project is effectively being DDoSed by AI-generated bug reports, and they're even considering shutting down their bug bounty program for the first time in decades. And this is TL Draw. They're automatically just closing all pull requests, whether they're AI-generated or not. And it's gone so bad that GitHub added a feature to disable third-party pull requests altogether, which is really sad because pull requests were the thing that made GitHub what it is today. And we're probably going to see a lot of big open source projects opt into this. And so when I say that open source is dead, I mean some parts of it, like the community. It's just not worth cultivating anymore, especially because building software is so cheap. And also the risk of supply chain attacks. I'm sure you've seen all the reports of things getting compromised. It's become more dangerous than ever to depend on third-party software, where it takes a single compromise and a massive chain of contributors to get pwned. So just as an example, Lite LLM is a Python package. It gets like three and a half million downloads a day. They were compromised for three hours, where attackers used a GitHub app that they used to steal their PyPI publishing tokens and publish a compromised version of the package that would install a credential harvester that would steal your API keys, your SSH keys, your crypto keys, and also install a backdoor that lets them do remote command execution. And the only reason this was even caught as quickly as it was was just pure luck, because the malware had a bug in it where it would cause cursor to crash if you ran the Lite LLM MCP server. And a security researcher noticed that and was able to figure it out. But if this had been out any longer, it would have caused catastrophic damage, especially because a lot of the people using Lite LLM are the enterprise customers and developers that have their own internal gateways. But despite all of this, I believe there are some parts of open source that are sticking around and becoming more important, like allowing others to use your thing freely in the public domain and build on top of it. And those parts of it are going to become more important than ever, particularly with open weights models, because of the economic impact. And so, to help explain why, I want to look at what's happening with inference spend right now. So this is a report from an anonymous CFO at an unnamed company, where they accidentally spent $500 million on Claude in a single month because they didn't set the usage limits on their thousands of employees on their Anthropic dashboard. This is another report by Uber CTO, where after they rolled Claude out to their organization, 95% of their engineers were using it, 70% of their committed code came from Claude, and their monthly spend per user was up to $2,000. And they said they used their entire 2026 budget in just four months. And the crazy part is that the AI labs are losing money too. This is a chart from Semi Analysis where they ran experiments with Claude, Code, and Codex subscriptions, where they would give them long-horizon coding tasks until they exhausted their weekly limits. And they found that a $200 plan for Claude would give them about $8,000 worth of API usage, and a $200 subscription to Codex would give them about $14,000 worth of API usage. So I think the strategy is pretty obvious. They're essentially going to subsidize this until they have as many engineers dependent on their tooling as possible, with agents in their CI and background cloud agents and looping agents and all these things, where it feels like every new feature and marketing push from these labs seems to be a new workflow to standardize on to use even more tokens and to be locked in even more. And then, inevitably, the price gouging, once they've got you trapped, where your developers can't work without the tools. And this isn't theoretical. We're seeing this happen live with some of the customers that we talk to. So just a quick show of hands: How many of you stop working whenever there's a cloud outage or a GPT outage? Yeah, same. I think that's a reason why we've seen Anthropic and OpenAI go from being API businesses to investing so much into the application layer, because they know that that's where they can set these sorts of traps and build their moat for the day that these models inevitably become a commodity. But I don't actually think the strategy is going to work. And that's the message I wanted to get across today: that this feels very short-sighted, and what we're noticing happen in the world is that it doesn't matter how many features your CLI agent has. Developers and businesses will just jump to whatever offers them the best value for their dollars. So if we look at current OpenWeights models, many of which are built in China, we'll notice that, although they've lagged behind the American closed source competitors, we're at an inflection point where raw intelligence lead doesn't matter as much anymore, because these models are powerful enough that you don't always need the best one for all your work. And that cost is becoming extremely important to these businesses that have turned a blind eye until now. And I think we all feel it, that to get the best output from these models, it's more a problem of what context and tools you give the agent access to and less about its raw intelligence. With the right AI-native development infrastructure, with project skills and rules, systems of verification and quality gates, even a mediocre model can produce similar results as a more intelligent model. It just might take more tokens. The intelligence is better placed in the system and guardrails around the model so that you don't have to be as reliant on the model or your end developer's responsible use of the model itself. So we recently shared an anecdotal experience where we were skeptical of the benchmark saying that GLM was better than Opus, so we tested them on a real bug from the client repo. And while both models fixed the issue, GLM was the winner in terms of cost and code quality. So GLM used twice as many tokens but only cost half as much. Opus finished faster. It used half as many tool calls. But GLM cleaned up dead code and verified that the build compiled before completing, while Opus didn't. It left a bunch of type errors, and it broke the production build. And so that gave us the sense that GLM was trained to spend more tokens verifying its output, which is fine because the tokens are cheaper anyway, and it's really the end result that matters. And because these OpenWeights models can deliver the same output with a little bit more tokens, we're seeing signs of the industry adopting and standardizing on these models. So this is Brian Armstrong, the CEO of Coinbase, saying that they've defaulted to using GLM and Kimi in their internal LLM gateway and that this has cut their AI spend by nearly half, while their token usage continues to grow. And I think we'll see other businesses building their own internal tooling and routing to work with these agents in the most dollar-efficient way for them, even if it means not having access to the latest new feature in something like Cloud Code. We're seeing the same thing that happened with OpenCompute 15 years ago. So just a quick history lesson for those that haven't heard of this. In 2011, Facebook was just getting started on building out stuff like distributed computing infrastructure and data centers. But by the time they built it, Amazon and Google already beat them to the punch, so it wasn't a competitive advantage. So Mark said, alright, let's just open source it and see what happens. So they took the designs for their data centers and their servers and their networking and cooling racks and all the physical hardware they spent all this energy and money building, and they just gave it away. They published schematics and CAD files and everything and called it the OpenCompute project. And what they saw was that the entire supply chain reorganized around it. So before OpenCompute, every company designed its own proprietary servers, so manufacturers were doing small production runs of very custom hardware, which was expensive. But when Facebook's designs became this sort of shared open standard, suddenly everyone was ordering the same thing, and manufacturers could do these massive standardized production runs, commoditizing these components so no single vendor could charge a premium, and the price of everything came down for the whole industry, including Facebook itself. And so what Facebook found was that by giving these designs away, they created the market that drove their own costs down and saved them billions of dollars down the road. And so I think the lesson taught here is that the industry will adopt and standardize on something that they can build on top of, even if it isn't the best thing. And with how much CapEx we've locked in for the next five years for AI infrastructure buildout, open weights models are only going to get cheaper. There are estimates that we'll spend nearly three trillion dollars and create over 100 gigawatts of new data center capacity by 2030, roughly doubling global capacity today. And hosting providers like Base 10 and Fireworks, their whole purpose is to beat the competition. So they'll use infrastructure efficiency gains like dedicated hardware and caching and batching volume tricks and inference-specialized silicon to drive costs down even more. And by 2030, the estimates are that inference on a one trillion parameter LLM will cost 90% less than it does today. We're seeing the same sort of cost-cutting tricks that commoditized the cloud 10 years ago happening in inference. So in 2014, Google, at their GCP Live March announcement, said that they would cut compute by 32% and storage by 68%. And AWS fired back with similar cuts within days. And this was AWS's 42nd price cut at that point. And so from 2015 onward, once raw compute and storage were a commodity, the hyperscalers stopped competing on it as much and started competing on other things like databases and serverless. And I think, because OpenWeights allows these host providers to compete and cost-optimize so aggressively, we'll see mass adoption of these foreign OpenWeights models, because when dollars are involved, the markets are extremely efficient. And the absurd API costs that these closed labs charge just won't be worth it anymore for most knowledge work. And so this is me humbly requesting the American labs to take OpenWeights more seriously, because before we know it, all this infrastructure that we're investing in could be built on foreign models that take the world by storm and make GPT and CLOT irrelevant. Mindshare and adoption is incredibly important, and there's a chance that if the foreign models become the standard, there won't be a reason to switch back to GPT or CLOT or Gemini, no matter what the marginal improvements are. And then we lose control over the development of this technology, and who knows where the world is headed if we don't have the likes of Anthropic and OpenAI to invest so heavily into safety research in ways that perhaps these other labs wouldn't. I think the development of a technology this transformative is deeply tied to the ideals of the people and the nation that's building it, and to instill those values in the future of this technology, we need to keep the lead. So I don't mean we need to open source our research. I think that's what gives us the lead. I think we need to open up our models and start releasing more OpenWeights models, which, as we know, are not nearly as useful. You can use and extract the traces and train your copycat models on them more easily, but not in a way that can leapfrog. And this would make models more usable by the industry in a way that allows more competition and adoption and better price and value for customers. And I think that's how we keep our lead during this very critical moment. Me and Klein believe so much in this OpenWeights future that we launched an OpenWeights subscription plan earlier this week that, through volume-based discounts and partnerships with inference host providers, we can offer significant discounts compared to paying for these models at direct API costs, and we plan on continuing to increase the usage quota for these models as they become cheaper. You can sign up at that link, klein.bot.pass, if you'd like to get a feel for how far models like GLM and DeepSeq have come and how you don't need the most expensive closed frontier model access to get work done anymore. Klein is also open source, so you can bring your own API key and use any other provider. You can use it on your CLI and VS Code and JetBrains. And yeah, we continue to add the newest models whenever they're released, so it's a good way to get a feel for how much better the latest newest model is, especially with OpenWeights, because you can't access those with the Cloud or ChatGPT subscriptions. Cool. That is my presentation. Thank you. Thank you. Thank you. with like the Cloud or ChatGPT subscriptions. Cool. That is my presentation. Thank you. Thank you. Thank you.