English 箭头
Podcast Cover

[World Labs: Pioneering Spatial Intelligence and the Future of World Models]-[What Comes After ChatGPT? The Mother of ImageNet Predicts The Future]

a16z Podcast · B2 · 2025-12-05

Technology
Or study on the web version

📋 Summary

The Dawn of Spatial Intelligence: An Interview with Fei-Fei Li and Justin Johnson

In this episode of the Latent Space podcast, Fei-Fei Li and Justin Johnson, co-founders of World Labs, discuss their mission to move beyond language-centric AI into the realm of spatial intelligence. Reflecting on the evolution of deep learning—from the ImageNet revolution to the current era of large models—they explain why the next decade of AI development will be defined by the ability to model, generate, and interact with three-dimensional worlds.

The Evolution of World Models

Fei-Fei Li, a pioneer of the deep learning era, highlights that the history of the field is fundamentally a "history of scaling up compute." While language models have dominated recent progress, she argues that they represent only one facet of human-like intelligence. Justin Johnson, her former PhD student, notes that the success of models like AlexNet shifted the paradigm of computer vision, but current research must now look toward "getting AI out of the data center and out into the world."

They define spatial intelligence as the capability to "reason, understand, move, and interact in space." Unlike language, which they describe as a "lossy channel" for information, spatial data is high-bandwidth and inherently three-dimensional. They argue that while language models have achieved impressive reasoning, they lack the causal understanding of physical forces—such as gravity or structural integrity—that is innate to human spatial cognition.

Introducing Marble: A Generative Model for 3D Worlds

World Labs recently launched Marble, a generative model that creates explorable 3D worlds from text or image inputs. According to Johnson, Marble is designed to serve as a bridge between the grand vision of spatial intelligence and immediate practical utility.

Key features of Marble include:

  • Multimodal Input: Users can input text or multiple images to generate a cohesive 3D environment.
  • Interactive Editability: Unlike static generative models, Marble allows users to "interactively edit scenes," such as changing the color of an object or reconfiguring the layout of a room.
  • Gaussian Splats: The model uses Gaussian splats as its atomic unit, which allows for real-time rendering on consumer devices like iPhones and VR headsets, providing precise camera control that frame-by-frame video models cannot match.

Challenging the "Sequence Model" Paradigm

One of the most provocative technical claims made by the duo is the re-evaluation of the Transformer architecture. Johnson asserts that "transformers are actually set models, not sequence models." He explains that the order in a Transformer is only introduced through positional embeddings; the core operations are permutation-equivariant, meaning the architecture is fundamentally designed to process sets of tokens. This insight opens the door for researchers to rethink how we structure neural networks to handle spatial data more effectively than the traditional 1D sequence approaches inherited from RNNs.

The Role of Academia vs. Industry

Fei-Fei Li addresses the ongoing debate regarding open science versus proprietary models. She emphasizes that academia is currently "severely under-resourced" and warns against the trend of PhD programs becoming "vocational training" for big tech labs. She advocates for a balanced ecosystem where academia focuses on "wacky, blue-sky ideas"—such as developing new primitives for distributed computing or exploring hardware-software co-design—while industry focuses on productizing these breakthroughs.

Future Horizons

Looking ahead, the founders view spatial intelligence as a horizontal technology. While Marble currently targets creative industries like VFX, film, and interior design, they foresee its application in robotics. Providing synthetic, high-fidelity data for embodied agents remains a significant "pain point" in robotics, and they believe World Labs' generative world models could provide the necessary simulation environments to solve this data starvation.

Ultimately, the conversation underscores a shift in AI philosophy: moving from fitting patterns in text to building causal, interactive models of the physical world. As Li and Johnson conclude, the journey toward true spatial intelligence is only beginning, and the goal is to create systems that do not just predict the next token, but understand the very fabric of the world we inhabit.

🎯Key Sentences

1
That's very easy because Justin was my former student.
2
How do you think about the AlexNet equivalent model for world models?
3
I think one is just there is a lot more data in compute generally available.
4
I think it's important to recognize the ecosystem is a mixture, right?
5
It's just how the market is.
Expand All

📝Key Phrases

1
scaling up compute
2
put all the eggs in one basket
3
see the daylight
4
state-of-the-art
5
theoretical underpinning
Expand All

📖 Transcript

I think the whole history of deep learning is in some sense the history of scaling up compute.
When I graduated from grad school, I really thought the rest of my entire career would be towards solving that single problem, which is
A lot of AI as a field, as a discipline, is inspired by human intelligence.
We thought we were the first people doing it.
It turned out that was also simultaneously doing it.
So Marble, like basically one way of looking at it, it's the system, it's a generative model of 3D worlds, right?

ListenLeap Brings You Into Real Context Learning

🎨 Interesting Content
🌍 Real Materials
📱 Listen Anytime
Or study on the web version