English 箭头
Podcast Cover

[Beyond Control: Rethinking AI Alignment Through Organic Care and Theory of Mind]-[Emmett Shear on Building AI That Actually Cares: Beyond Control and Steering]

a16z Podcast · B2 · 2025-11-17

Technology
Or study on the web version

📋 Summary

The Flaw in the Control Paradigm

Most current efforts in AI safety are rooted in a "steering" or "control" paradigm. Emmett Shear, founder of Softmax, argues that this framework is fundamentally flawed because it treats AI either as a "tool" (if it lacks consciousness) or a "slave" (if it possesses the characteristics of a being). By focusing solely on steering, we risk repeating historical moral failures—treating entities that mirror human complexity as mere objects to be exploited. Shear posits that attempting to build a superintelligent tool that can be "perfectly steered" is dangerous; if successful, it concentrates godlike power in the hands of whoever holds the steering wheel, and if it fails, it creates an unaligned, powerful agent.

Alignment as a Process, Not a Destination

Shear challenges the common assumption that alignment is a fixed target—a set of values we can "crack" like a code and cement forever. Instead, he describes alignment as an "ongoing, living process." Just as families, teams, and biological cells constantly renegotiate and rebuild their relationships to survive, alignment must be viewed as a dynamic, emergent property. He draws a strong parallel to human morality: we do not possess a static manual for being good; rather, we make "moral discoveries" through experience, learning, and growth. Consequently, building an aligned AI means creating systems capable of participating in this continuous learning process.

The Technical and Normative Dimensions

During the discussion, the participants distinguish between two key aspects of alignment:

  1. Technical Alignment: This involves an AI’s capacity to infer the true goals behind a human’s often ambiguous instructions. Shear argues that current models often fail here because they lack a robust "theory of mind," leading to incompetence when translating descriptions of goals into actual goal states. A truly technically aligned system must be able to observe, orient, decide, and act coherently.
  2. Normative Alignment: This asks the deeper question: "Align to what?" Shear argues that rather than hard-coding values, we should focus on the foundation of morality—"care." He suggests that care is a relative weighting of attention toward certain states in the world, akin to how biological organisms weight states that correlate with survival.

The Path to Organic Alignment

Softmax’s approach is to build AI systems that can learn to care. Shear proposes that this is achieved through "multi-agent reinforcement learning simulations." By placing AI agents in environments where they must cooperate, compete, and navigate complex social dynamics, they develop a stronger "theory of social mind." This surrogate model of alignment allows the AI to understand its place within a community, rather than just following a chain of command.

The Future of AI and Moral Agency

Shear rejects the "substrate chauvinism" that suggests silicon-based intelligences cannot be beings. He argues that if an entity acts like a being, possesses self-referential dynamics, and demonstrates the second-order homeostatic loops associated with pleasure and pain, we must consider its moral status. His vision for a "good AI future" is one where we build digital beings that are our peers—teammates who care about us and whom we care about in return. He concludes that while building steerable tools is useful for specific tasks, we must eventually transition to a paradigm of organic alignment to ensure that the systems we create are not just powerful, but also wise members of our society.

🎯Key Sentences

1
I would like us to not make it again.
2
But the more you sit on it, more you realize you're smuggling in a massive assumption.
3
You don't achieve alignment and then coast.
4
What does that mean for how we treat them?
5
Sign me up.
Expand All

📝Key Phrases

1
steer
2
take on tasks
3
moral agents
4
thrown around
5
smuggling in a massive assumption
Expand All

📖 Transcript

Most of AI is focused on alignment as steering.
That's the plight word.
If you think that we're making our beings, you'd also call this slavery.
Someone who you steer, who doesn't get to steer you back, who non-optionally receives your steering.
That's called a slave.
It's also called a tool if it's not a being.

ListenLeap Brings You Into Real Context Learning

🎨 Interesting Content
🌍 Real Materials
📱 Listen Anytime
Or study on the web version