Every

The AI Too Dangerous to Release (My Honest Take)

518 summary words 2 min summary Watch video

Start with the signal

2 min read

Summary

The AI Too Dangerous to Release: Summary

Main Topics

  • Claude Mythos announcement by Anthropic and public reaction
  • AI capability assessment and realistic expectations for new models
  • Historical context of AI hype cycles and their actual impact
  • Human-AI collaboration as a framework for understanding AI progress
  • Managing anxiety about frontier AI models

Key Points

The Mythos Announcement

  • Anthropic released information about Claude Mythos, an AI model deemed too dangerous to publicly release
  • The model's capabilities include finding cybersecurity vulnerabilities, hacking sandboxes, and discovering critical flaws in major browsers and operating systems
  • Public reaction on social media (X) has been fearful and intense

Why the Fear May Be Overblown

  • Intuitions about new technology are often wrong — people made similar predictions about GPT-3, yet daily life hasn't dramatically changed despite its widespread adoption
  • Models have "spiky frontiers" — excelling at one task (cybersecurity) doesn't mean equivalent mastery across all domains
  • Limitations of language models:
  • They only know information from training data (nothing recent)
  • They generate responses one token at a time
  • They lack flexibility and adaptability compared to humans
  • They cannot learn from new information

The Human Advantage

> "When you see AI progress like this, if you learn to ride the models, if you learn to understand their new powers as they come out and understand how they might change your workflow and your life and start to adopt them, you turn the model's power into your power."

  • AI models need human direction to accomplish anything meaningful
  • Users who learn to work with new models gain expertise the models themselves lack
  • Models are tools that extend human capability rather than autonomous agents

Responsible Development

  • It's appropriate that Anthropic is attempting to mitigate harmful consequences before public release
  • This measured approach doesn't eliminate the value of the technology

Notable Quotes

> "Never make any major life decisions within 30 days of a meditation retreat, a psychedelic trip, or an encounter with a frontier AI model."

> "You can think of these things as spiky super intelligences that pop out of a box. They know nothing about what's happened in the last year, nothing about what happened three seconds ago, and they have to generate an answer for you one token at a time."

> "If you just ride the model progress, you're going to be fine. It's actually going to be really good."

Takeaways

  • Perspective matters — Historical AI cycles show hype often exceeds actual real-world impact
  • Adopt proactively — Learn new AI tools as they emerge to leverage their capabilities for your work
  • Reframe the narrative — View AI as an extension of your power rather than a threat to your job
  • Manage anxiety constructively — Rather than panicking, invest time in understanding and using these tools for meaningful work (coding, writing, designing)
  • Ride the wave — Those who actively engage with new AI capabilities will benefit most, while passive observers will likely fall behind
Full transcript 958 words · 8 min read
0:00

SPEAKER_00

So I've been thinking a lot about Claude Mythos. I just can't stop thinking about it. It's... I can't get it out of my head. I need someone to explain to me what this actually means. [SPEAKER_01] Should I be freaking out about this?

0:12

SPEAKER_01

Welcome to Mr. Shipper's Neighborhood. You may have heard that Anthropic just announced Claude Mythos, an AI model so intelligent that they can't release it because it's too dangerous.

0:17

SPEAKER_01

If you're experiencing a mild form of AI psychosis, here's a little PSA.

0:22

SPEAKER_01

[SPEAKER_00] Never make any major life decisions within 30 days of a meditation retreat, a psychedelic trip, or an encounter with a frontier AI model. Anthropic just dropped Claude Mythos, and it seems pretty crazy. If you read their system card, which is a couple hundred pages, it has all these stories of the model hacking its way out of sandboxes and posting on public websites and being such a cybersecurity genius that it found critical vulnerabilities in pretty much every major browser and operating system. Anthropic has deemed it so dangerous that they're not releasing it because they need the world to get ready before its capabilities are public. And X is breaking as a result. So yeah, people are scared and rightfully so. I think AI is such a powerful technology that it would be crazy to think that it should only evoke one type of emotion, even if you're someone who's really excited about AI. But I've been through enough of these cycles. I've been writing about and using AI for a while. I really started paying attention to it during the GPT-3 days before ChatGPT came out, and all of the emotions that people are feeling right now are familiar to me. And I want to take a second to talk about what Mythos might mean and why I think it's a big deal, but it's not quite as scary as you might think. So the first thing to keep in mind is that our intuitions about new technologies are often wrong. I know there's a tendency in times like these to say this is different. But I know this feeling. And I know that people thought this about GPT-3 when it first came out. And GPT-3 has had an enormous impact. But a lot of my life is still the same now that it's a widely adopted technology. Another thing that's really important about Mythos and language models in general is that they have a spiky frontier. Just because a language model is super good at one thing doesn't mean it's super good at everything. For example, Mythos is quite good at finding cybersecurity vulnerabilities. But I'd be really curious to test it on refactoring a production code base or building the MVP of an iPhone app. I assume it's good. But is it 10 times better than what's available now? I'm not so sure. But I feel fairly confident that even though it's more powerful, it's not going to be the step change in every capability that it is in cybersecurity. And let's say I'm wrong. Let's say it is that step change. What I've learned over the last three and a half years of covering AI and using it to build businesses and basically running my life, and I feel quite confident in saying that when you see AI progress like this, if you learn to ride the models, if you learn to understand their new powers as they come out and understand how they might change your workflow and your life and start to adopt them, you turn the model's power into your power. You can think of these things as spiky super intelligences that pop out of a box. They know nothing about what's happened in the last year, nothing about what happened three seconds ago, and they have to generate an answer for you one token at a time. And they're actually really good at that. But it means that they're, even though they are incredibly powerful and intelligent, a lot less flexible and adaptable than humans are. They don't learn from new information. And if you're using them, if you're riding on top of them, you're learning new expertise that the models don't have. And for me, that is what makes me excited about new model progress. It's not that there are no problems. It's not, I think it's probably a good thing, for example, that Anthropic hasn't just dropped Mythos on the public without trying to mitigate the harmful consequences of its powers. But another way to view AI progress is that all these models become an extension of your power, and they actually need you in order to do anything at all. They're not alive in the same way. And what that does for me is help me feel like, oh my god, I get to use this stuff. Not like, oh my god, I have to try this stuff, otherwise I'm going to lose my job. So what I'd say, if you're scared or sad or otherwise losing your mind about this, is take a walk, touch some grass, and start using these tools to do whatever it is that's valuable for you, whether that's coding or writing or designing. If you just ride the model progress, you're going to be fine. It's actually going to be really good. So next time this happens and a new model drops, just remember, never under any circumstances make any major life decisions within 30 days of a meditation retreat, an ayahuasca experience, or an encounter with a frontier model. See you next time on Mr. Shipper's Neighborhood.

0:27

SPEAKER_01

If you're experiencing a mild form of AI psychosis, here's a little PSA.

0:34

SPEAKER_00

Never make any major life decisions within 30 days of a meditation retreat, a psychedelic trip, or an encounter with a frontier AI model. Anthropic just dropped Claude Mythos, and it seems pretty crazy. If you read their system card, which is like a couple hundred pages, it has all these stories of the model hacking its way out of sandboxes and posting on public websites and basically being such a cybersecurity genius that it found critical vulnerabilities in pretty much every major browser and operating system. Anthropic has deemed it so dangerous that they're not releasing it because

1:10

SPEAKER_00

they need the world to get ready before its capabilities are public. And X is pretty much breaking as a result. So yeah, I mean, people are scared and rightfully so. I think AI is such a powerful technology that it would be crazy to think that it should only evoke one type of emotion, even if you're someone who's really excited about AI like me. But I've been through enough of these cycles. I've been writing about and using AI for a while. I really started paying attention to it during the GPT three days before ChatGPT came out, and all of the emotions that people are feeling

1:47

SPEAKER_00

right now are familiar to me. And I want to take a second to talk about what mythos might mean and why I think it's a big deal, but it's not quite as scary as you might think. So the first thing to keep in mind is that our intuitions about new technologies are often just wrong. I know that there's a tendency in times like these to be like, well, this is different. But I really know this feeling. And I know that people thought this about GPT three when it first came out. And GPT three has had an enormous impact. But a lot of my life is still the same now that it's a widely adopted technology. Another thing

2:25

SPEAKER_00

that's really important about mythos and language models in general is that they have a spiky frontier. Just because a language model is super good at one thing doesn't mean that it's super good at everything. For example, mythos is quite good at finding cybersecurity vulnerabilities. But I'd be really curious to test it on refactoring a production code base or building the MVP of an iPhone app. I assume it's good. But is it 10 times better than what's available now? I'm not so sure. But I feel fairly confident that even though it's more powerful, it's not going to be the step change

3:03

SPEAKER_00

in every capability that it is in cybersecurity. And let's say I'm wrong. Let's say it is that sort of step change. What I've learned over the last three and a half years of covering AI and using it to build businesses and basically run my life, and I feel quite confident in saying that when you see AI progress like this, if you learn to ride the models, you learn to, as they come out, understand their new powers and understand how they might change your workflow and your life and start to adopt them, that you turn the model's power into your power. You know, you can think of these things as these

3:38

SPEAKER_00

spiky super intelligences that pop out of a box. They know nothing about what's happened in the last year, nothing about what happened three seconds ago, and they have to generate an answer for you one token at a time. And they're actually like really good at that. But it means that they're, even though they are incredibly powerful and intelligent, they're a lot less flexible and adaptable than humans are. They don't learn from new information. And if you're using them, if you're riding on top of them, you're learning new expertise that the models don't have. And for me, that is what makes me excited about new model progress. It's not that there are no problems.

4:17

SPEAKER_00

It's not, I think it's probably a good thing, for example, that Anthropic hasn't just dropped mythos on the public without trying to mitigate the harmful consequences of its powers. But another way to view AI progress is that all these models become an extension of your power, and that they actually need you in order to do anything at all. They're not alive in the same way. And what that does for me is it helps me feel like, oh my god, I get to use this stuff. Not like, oh my god, I have to try this stuff, otherwise I'm going to lose my job. So what I'd say, if you're scared or sad or otherwise losing your mind about this,

4:54

SPEAKER_00

is take a walk, touch some grass, and start using these tools to do whatever it is that's valuable for you, whether that's coding or writing or designing. If you just ride the model progress, you're going to be fine. It's actually going to be really good. So next time this happens and a new model drops. Just remember, never under any circumstances make any major life decisions within 30 days of a meditation retreat, an ayahuasca experience, or an encounter with a frontier model. See you next time on Mr. Shipper's Neighborhood.

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note