We are building the next frontier of AI, which is what we call spatial intelligence.
At Cynics, we are developing what we call a real-to-sim-to-real pipeline.
We can replace all the data, all the evaluation we need in the real environment by using the data that can generate at a scalable way in our digital world.
Think about human intelligence.
We do a lot of simulation in our head.
You know why?
There's a very important role simulation plays that real world data doesn't play, which is counterfactual.
What we are building is a consistent world.
Consistence both over space, over time, over different viewpoints, and over different type of interactions.
My North Star is I want the robotic world.
The world we live in can be multiverse, that we create technology to allow people, builders, developers to act within different spaces.
Do you believe we'll ever be able to build robots that have the power efficiency of a human being?
How far away are we from this?
Is this like five years or this is like never?
The TLDR is...
Language models transformed how AI understands words.
The next frontier is teaching AI to understand and act within the physical world.
Following World Lab's acquisition of Cenex, Martine Cassaro sits down with Fei-Fei Li and Yunzhu Li to unpack the vision behind the deal.
They discuss spatial intelligence, world models, simulation, and why solving robotics will require a new generation of AI built for three-dimensional reasoning, not just language.
All right, well, it's great to have you both here.
So Fei-Fei, for the listeners that may not have the background, maybe you can give an overview of what World Labs does.
Yeah, well, WorldLab is a two-year-old startup.
I think we should just recognize it's a frontier model lab.
We are building the next frontier of AI, which is what we call spatial intelligence.
And spatial intelligence is about creating AI that has the ability to
Generate, understand, reason with, and interact with spaces, whether it's physical or virtual.
And of course, a means to an end towards spatial intelligence is building large world models.
And that's what World Labs is mostly focused on.
So you've been saying this since the very beginning, which is the machine's ability to perceive and reason about spaces and act on spaces.
But I always had the assumption that the acting on spaces was some long-distance shooter thing, but now you're acquiring a robotics company.
And so maybe talk a little bit about the timeliness of this and the intentions.
Yeah.
So first of all, it doesn't just take robotics to act within spaces or to interact, right?
I mean, look at the creative field, whether it's
VFX or gaming or design, many use cases, you can create and act within virtual spaces.
And World App's thesis has always been that the world we live in can be multiverse, that we create technology to allow people, builders, developers to act within different spaces.
Having said that, the ability to act within the physical space is one of the most exciting and most profoundly important capability of the future AI world.
So robotics is very much that.
So WordLab has always believed that robotics is an important
Application as well as use case of spatial intelligence and world modeling.
So by joining force with inviting Cinex and Cinex team to World Labs is part of our long-term vision and mission.
We've always committed to that.
Amazing.
So Yun-Chu, you're the co-founder of Scenics.
So maybe provide everyone with a quick overview of your background and what Scenics does.
Yeah, so I'm Yunzhu.
So I'm currently co-founder of Cynics and also assistant professor at Columbia University.
So my research started from my PhD at MIT and then postdoc with Fei Fei.
Really?
Yes.
That's great.
The world is small.
The world is small.
It is.
Throughout my career, my goal has been very simple, trying to help the robots better perceive and interact with the physical world.
So I'm a very practical person.
I want my robot to work in a real physical environment.
So for Cynics, the unique opportunity we see is that there has been a lot of bottlenecks.
Right now we see faced by the developments of general purpose robots, especially around training and also around evaluations.
So at Cynics, we are developing what we call a real-to-sim-to-real pipeline.
We want to map the real environments into the digital world that has the best alignments with the real environments.
By alignment, we mean that whatever happens in the digital world is also going to happen in the real environment, such that we can replace all the data, all the evaluation we need in the real environment.
By using the data that can generate at a scalable way in our digital world.
So that is how everything started.
In Cinex, we put together a very, very strong and best teams around robotics, robot learning, and also simulation and rendering, trying to build this real-to-sim-to-real stack to solve some of the key bottlenecks.
It's amazing that you two work together.
Yeah, and there is a funny story here because you would think because we work together, he was my amazing postdoc, we've been talking about this Cinex and WorldLab integration for a long time.
It's actually not true.
They came into WorldLabs as a customer.
When we released the first version of our generative model called Marble last winter, around November, December, Cinex just signed up.
No kidding, as a customer?
Yes.
And I didn't even know what it was.
And then I realized this is Yundru's company.
I called Yundru.
I'm like, wow, this is your company.
And then we realized there's so much synergy.
Maybe, Fei-Fei, just quickly describe what Marble is.
Yeah, Marble is the code name for the base model that WorldLab has been training and iterating on.
The fundamental capability right now of Marble that is publicly released is to take a prompt, it can be an image, it can be a few images or a text, and turn that into a
Geometrically consistent world that can be represented in 3D geometry, whether it's Gaussian splat or mesh.
Really what Cynic's team is doing is trying to solve this extremely difficult problem in robotics, which is the lack of data.
The lack of data in training, the lack of data in evaluation, this is very, very different from language models where data is abundant on the internet.
And we know that in order for robotics to work, we have to somehow unlock the power of scaling law.
But where does that come from?
This is something that is a profound problem that everybody's battling with in robotics.
It'd actually be great to talk about this energy.
You have put together a very, very talented team.
You have put together a very talented team.
And so to what extent is there overlap?
To what extent is this an extension?
Maybe talk a little bit about that.
How complementary it is.
It's actually the TLDR is very complementary and with a shared mission.
So Windrew is one of the three technical co-founders.
The other two are Changxi Zheng, another Columbia professor who has been a world-class technologist in simulation.
Yeah.
And Zhang Xi has his background in also VFX.
He worked at Weta.
He worked at Tencent.
He's been an entrepreneur.
Then there's Sun Li Hu, who is a phenomenal engineering leader who was also in a startup that was acquired by Amazon many years ago.
So he worked there.
In many different tech stacks in the computer vision field in Amazon.
So when we started talking more seriously, I recognized that a couple of things that Cynics has from a talent point of view
Is extremely complementary to world labs.
One is obviously Andrew's incredible thought leadership and just technical prowess in robotics, right?
So from really
From hardware, full stack robotics.
And even when he was my postdoc at Stanford, at that time, you already had your faculty offer.
So you were there only for one year.
I want
You for more than one year, but he had to go have the real job.
So he was a full-stack researcher in robotics, from modeling to hardware.
And of course, Yunzhu and his students at Cinex was that pool of talent World Lab hasn't had yet.
Then on the Changxi side is just incredible simulation capability, right?
He's such a senior researcher and technologist in simulation.
And what World Labs is doing is very much...
Interfacing the world of simulations.
So I think what they don't have, obviously, is on the generative model side, as well as the computer vision 3D reconstruction side, we're also very strong at World Labs.
So that's a technology that Cinex needs.
So together, these two sides come together and make it much more complete.
Fei-Fei's motivation in this is like this is an extension and a complement to get into robotics.
Having been in your situation, which is deciding when to sell a company, it would be great to hear from you on how you think about joining World Labs and kind of the fit there and why you made the decision to do it.
Yeah, so at the very beginning, we were deciding, okay, do we want to just keep going?
But after chatting with Fei-Fei, after seeing all the synergies that happen in the middle, it just makes perfect sense for the forces to join each other.
So in any sense, at Cynics, what we have been doing is real to seem to real, is to do dense reconstruction of the environment.
So we captured the appearance of the environment, geometry of the environment, and also the dynamics of the environment, meaning how the environment is going to change when you apply actions.
So this dense reconstruction right now is still a little bit on the heavier side.
And what World Labs right now has been doing involves a lot of profound capabilities around sparse reconstruction and generations.
So we see a lot of opportunities of leveraging marble and other capabilities at World Labs in order to do very efficient reconstructions and modeling of the environments.
So can we expect a foundation model for robotics from World Labs?
World Labs is building a foundation model, as you know, Martin.
We're building a base model.
And as the technology has been evolving,
Some of the most exciting base models are omni models, right?
They take multimodal input.
They have multimodal outputs.
And what is a foundation model for robotics?
It's very likely going to involve actions.
It's very likely going to involve the output of actions in addition to the state of the world.
And we're definitely not ruling this out.
Yeah, great.
So for example, for the foundation models, it essentially needs to be a multimodal model.
So it has to take into account frame of text, image, depth, and different kinds of modalities.
And action is a very, very important part of that modality.
So if you think about frame actions as inputs, that essentially affords similarity.
That is going to predict how the environment is going to change when you apply a specific action.
When the action is output, this is essentially a policy model that is trying to predict, give a specific goal, like what should be the action you take in the real environment to get you closer to that goal.
So this kind of omni models actually can benefit a lot and actually provide huge amount of values for the robotics communities in trying to understand how to model the environments and at the same time, how to act.
In the environment.
And this can also act as a backbone for you to fine-tune into specific robotic applications to making sure it really lives up to the reliability and efficiency that's expected by the clients.
You know, Yunxiu, if you don't mind kind of a lay investor question, I see a lot of robotics companies.
And a very popular approach right now for the robotics companies that come in is, like, we'll use a video model, you know, and, like...
You know, that's the predominant method where this is, you know, 3D and simulation.
It's a very different approach.
And so maybe you could contrast this popular approach of just using video only versus kind of what the ambition here is.
Yeah, so in order to create words with Robot Candler, the words, as I mentioned, need to capture the essential structure of the problem.
And one of the very important and necessary requirements for those words will be consistency.
So that is where I actually see there's very, very strong synergies with Marble, because what we are building is a consistent world.
Consistence both over space, over time, over different viewpoints, and over different type of interactions.
And Marble, the generated words from Marble, is also provide an infrastructure, a component of that entire words that we believe is necessary for the robot tuner.
Imagine if a robot is pushing an object forward, the object just magically disappears, which has been a problem of many of the existing video prediction models.
This one provides a good enough signal for the robot to know what is the right thing to do.
But obviously right now there has been a lot of investigation on building better and better and stronger and stronger like video models.
So we actually see a way where some of the infrastructure we build can provide as initial momentums and to go into this data flywheel of
Going from this like a more simulation driven models into like a robot policy models, which is going to do the execution in the real environment, collecting new data, the data will come back in.
Where the model doesn't necessarily have to be physics only or learning only, but somewhere in the middle, which be able to capture the essential structure of the problem.
But at the same time, be able to scale and become better and better as you accumulate.
You know, I've worked now, say, very closely for a while, and you've always had this North Star, which has driven this.
And, you know, you've articulated variously as kind of 3D and in a number of other ways.
And I'm just wondering, for you, is there also a similar philosophical North Star, or you're more the pragmatic person?
Like I am.
I build the system, I do the thing.
My North Star is to make robots work in the real environment.
I'm a very practical person.
I want the robot to work.
One interesting thing that's actually coming from my collaborations with Phoebe during our postdoc, we are building this kind of benchmark.
We actually send out surveys asking the general public what they want the robots to do for them.
Among the thousand tasks we collected, one third of the tasks are about cleaning.
People just don't like to do those like a dull and dirty tasks.
And those are the scenarios where we really want to making sure we have robotic solutions to deal with.
One thing I really like about Cynics, Martin, especially continuing your question, there's a lot of robotics companies building models and all that.
One thing I truly like about Cynics is
As Yundru and his co-founders have such an incredibly pragmatic approach to robotics, especially they come from academia, right?
Sunny doesn't, but Yundru and Chuanxi come from academia.
But their first instinct is work with design partners and customers in real time.
Whether it's industry labs or warehouses or electronics assembly.
That is such a refreshing, actually, a refreshing way of approaching robotics.
And that really made me very excited to work with them.
Maybe this is for you, but I'll just be this is this is personal curiosity, which is it seems to me that for robotics, you have to be pretty exact.
I mean, not perfect, but pretty close.
But for the creative use cases, which Burl Epps has done a lot of, you kind of don't need to because, you know, I mean, you know, even
Sometimes, like, being wrong is stylistic or intentional or whatever.
And so, from a technical perspective, what is the challenge here for reconciling these two things, or do they never get reconciled?
Like, will there always be two points in the design space?
So they will be reconciled in the long term, of course.
And modeling of the environments doesn't have to be perfect.
The model doesn't have to be perfect in robotics.
And by the way, this is pure curiosity, but is there like...
A bit more formal way to say that.
Like, what does that mean not to be perfect?
It has to be pretty close.
So let me put it this way.
For example, models over the developments of all different kinds of robotic applications has been a very important cornerstone.
If you look at all the existing robotic applications, like Plane, Jones, Roomba, or even for quadruped robots, bipedal robots, model has been the way for them to actually work and be able to transfer from simulation to the real environment.
But if you look at those locomotion robots, like quadruped robots, bipedal robots, they can work on snows, they can work on bushes, but you don't need to have a simulator.
It can simulate all the bushes and snows very precisely.
You need to have a simulation that captures the essential structure of the problem and do a whole different kind of randomization inside the digital environment.
So that is what we're aiming for.
So basically, with Cynics and together with Word Labs, we're trying to investigate what is the level of fidelity we need to model the massive, massive worlds besides the robots, such that we'll be able to transfer the robotic systems training the simulated environment and digital worlds back into the real scenarios.
As an investor, I've heard other researchers say, like Sergey Levine, say,
Simulation will always eventually deviate from the physical world and real world data collection is absolutely critical.
And so maybe talk a little bit about like the viability of this approach where simulation is a cornerstone as opposed to some other approach.
So they don't contradict with each other.
So if you think about the simulation, simulation is essentially trying to predict how the environment is going to change when you apply the actions.
And this is essentially a model of the world that doesn't necessarily have to be pure physics.
It can be a combination between both physics and also learning.
We are collecting real-world data.
We will be using those real-world data.
It's just at different stages of this, like a data flywheel.
Maybe at the very beginning, we have stronger emphasize on we have more physics to making sure we have the right consistency and right structure for us to learn
The uh the the world for us to train the robot policies
But as we accumulate more and more data both through data collection and also through the collaboration with our clients we'll have the data that will be moving towards more towards more learning based like modeling of the environments
So this kind of transition and also this kind of data flail is really enabling factors of both getting the best of both physics and the geometry and consistency, as well as all the power and magics from the data and compute.
I want to add to this and be slightly philosophical here is there isn't a binary choice between simulation or no simulation.
All this come in together to make robotics work.
Think about human intelligence.
We do a lot of simulation in our head.
You know why?
There's a very important role simulation plays that real-world data doesn't play, which is counterfactual.
Reasoning is that you play out events that hasn't happened or cannot happen, or you don't have enough data to make it happen in real world.
And while you play it out, you learn how to act in it.
Humans do this all the time.
We probably don't, you know, we just, I know you were at the World Cups.
Yeah.
I was at the World Cup.
Congratulations to Spain winning.
I'm sure in the planning of every game, there is simulation, whether it's digital or on the whiteboard or whatever, that simulation, the role simulation plays is counterfactual reasoning.
And that's really important in robotics because we just do not have...
Cannot possibly have enough real-world data for that.
Here's a real-life example, the industry of self-driving cars.
Waymo has officially said they use billions of hours of simulations.
And actually Waymo is more simulation heavy than just real world data heavy.
So these are real examples.
And as you know, Martin and Yunzhu too, cars are the simplest kind of robots.
So clearly simulation plays a huge role in robotic learning.
I also want to add to that.
So if you put things more specific, simulation can provide two levels of benefits.
The first one is reliability, and the second one is efficiency.
So for reliability, if you're thinking about a robotic system working reliable in the real environment, you need data to provide systematic coverage of all the state space and the variations that robots might encounter.
That's
How you can learn of how is that is robust.
So with simulation, you can do systematic randomizations and control and the variations of lighting, frictions, geometries, object types, and also all different kinds of physical parameters to making sure you have sufficient coverage of the state space.
So this is what can give the robotic systems reliability.
And second is about efficiency.
So right now, many people are doing teleoperation.
And if you look at many of the teleoperation device, imagining all the actual skeletons you are using, you are actually collecting the data at a speed that is actually slower than human actually doing the task.
But for many of our clients, human speed to them is not good enough.
They want faster than human speeds.
So for the robot to move faster, it's not as simple as just drive the robot faster because the gravity doesn't change.
But in simulation, you can do systematic speed up of the robot's behaviors to train the robots such that it considers all the dynamics, changes of the environments.
So this is what can give our clients, for them, efficiency.
So both for the reliability and efficiency, there are some kind of very unique values where simulation can provide.
You've talked about the technology and the platform, what it does.
Maybe talk about the specific use cases people use it for.
There are essential, like two specific use cases, especially around both training and also around evaluations.
Starting from the evaluations.
So evaluation is something like people often overlook in the robotics.
But if you are tuning like a robot in Vodafone, you have to know how well it works.
And that is the only source of information for you to iterate.
By the way, a lot of, every AI person really understands what evals are and uses it all the time.
Non-AI people, it often means something a little different.
So maybe it's even worth just describing specifically what you mean by evaluation.
Okay, so what I mean by evaluation is you'll be able to understand for this specific checkpoint, how well does it perform?
Does it perform, for example, 95% of the time or 99.9% of the time?
And the key criteria people use in industry is
How long does it take?
How long in work or clock time does it take for you to distinguish between a checkpoint that is 90% from a checkpoint that is 92 points?
And if you only do that in the real environment, it just takes so long for you to do the distinguishments.
And you really think about also the robotic evaluations right now people are doing in the real environments.
The iteration speed is multiple orders of magnitude slower than iterations of those language models.
Yeah.
So not only is, like, the robotic tasks very varied, very diverse.
Oh, yeah, because, like, you actually have to do the thing.
Yeah, yeah, yeah.
Right, right.
Like, atoms have to move through space.
Yes.
Exactly.
But only there.
The laws of physics have to be obeyed
And have you watched those robotics videos
Every video has like 10x 8x because it moves so slowly exactly so
Not only it's slow it's dangerous it's costly but at the same time the speed is also like a
Multiple orders of magnitude slower.
So some of our clients actually need this digital environment that can be used to evaluate their robotic systems.
And because our digital environment has proven alignments with the real world, so meaning whatever happens in the scene is also likely to happen in the real environment.
If a checkpoint
It's working better in the simulation.
It's also highly likely to also work better in the real environment, as we have also been discussed in the blog post.
So that actually gives our clients very strong confidence
In actually using the data, using the signal from the digital environment to do scalable, safe, and much faster evaluations of their robotic systems.
Great.
So that is our evaluation.
Then on the training.
So on the training side, so basically, like I also mentioned, it's about controllability.
So you want to control all the different possible variations of states, parameters, lighting, frictions, physical parameters, like even object geometry, object types.
So you want to make sure you have sufficient coverage of all different kinds of scenarios, such that you'll be able to generate
Like informative data for your robots to be robust.
And this is just going to be so hard to do just in the real environment.
Like we discussed, if you do tidal operation, the speed at which you are collecting data is slow.
You're also limited by like how many robots you have,
How many teleoperation devices you have.
There's like a whole different kind of challenges around all the data operations around it.
But in simulation, everything can be controllable, everything can be systematic, and everything can be understood at a level where you know exactly
And making claims about exactly what distribution you have covered to develop confidence about within the distribution, we know the robot will work.
So those kinds of confidence and efficiency and scalability is something that our clients also value to use our digital worlds for the training of robotic systems.
Here's the crazy thing.
Even before Cynics and we are talking, our inbound customers for Marble were already seeing this kind of demands.
We just cannot serve these customers, but we are already getting a lot of phone calls from robotics,
Early-stage robotics companies were developing their models all the way to downstream, very pragmatic use cases.
And we're seeing these needs.
When people hear you're going into robotics, what they're going to envision is you're pulling out a 3D printer, and you're going to be making hardware, and then you're going to be
Programming the brain of a robot and sticking it in the robot and then you've got a robot.
And I don't think that's what you guys are talking about here.
So maybe talk about where this fits in the life cycle of creating a robot and where you will end and where the rest of the ecosystem will end.
Will begin.
So what we've been building, you can imagine, is an infrastructure with the softwares around these infrastructures for people to, for them, build worlds such that robots can learn and evaluate.
And this infrastructure is naturally model-agnostic and embodiment-agnostic.
So I just want to be very clear, just because this is actually a very subtle... For you, it's obvious, but it's a very subtle point, which is...
From what you said, that's not building a robot.
It's building an environment which another company can place their robot brain to navigate and to learn.
Yeah.
So for our customers right now, they have all different kinds of robots.
Some are using, for example, single robot arms.
Some are using bi-manual.
Some are using a fixed arm.
Some are using mobile manipulators.
Some using grippers, some are using some more elaborate versions of the only factors.
So our platform right now is just naturally embodiment agnostic.
We can very easily integrate different kinds of robotic embodiments, be able to put them into the words we generated, we digitalized.
Such that we will be able to give those individual robots capabilities of doing the right tasks and at the right levels of reliability and efficiency in the real environments.
And we are also, for example, model agnostic.
So we can just using the data generated by our words to train different models.
Either from scratch or doing post-training of existing foundation models, like vision language action models or word action models.
So to us, it doesn't matter.
We just want to make sure we have the infrastructure, we have all the words such that the robot can work reliably in the real environment.
You know, you have told me that you think a lot of the predictions around humanoids are a little bit aggressive and we're likely to see more constrained rollouts like warehouses or whatever.
Can you talk a little bit about that and like how that impacts what you're going to be tackling here at...
Like world labs?
So that's a very good question.
So if you look at, for example, all the progressions of robotic applications in the real environments, it has always followed the trend from going from fully structured environments
Into semi-structured environments and then into unstructured environments.
For fully structured environments, what we mean is that you have knowledge and control over all the configurations within the environments.
Like factories or, for example, car manufacturing lines.
Those have been automated for decades.
And then you have, for example, semi-structured environments, which you have certain controls over the environments, for example, like the Amazon warehouses, or, for example, like restaurants, hotels, where you have certain control over the environment
To just make the task easier.
But there are obviously many other, like objects, or for example, clothes.
Those are the objects you don't have control.
And then for the unstructured environments, it's like your home and my home.
This is, I would say, the grand challenge.
Especially my house, trust me.
Three dogs, five-year-old.
Yes, dogs.
Exactly.
If you're thinking about where does the robustness coming from?
Robustness coming from a sufficient coverage of the scenarios that robots might encounter.
So it's so much easier and more approachable, at least like right now, to focus more on the semi-structured environments before we move on to fully unstructured environments.
So we will move into that direction.
It's just we want to take a more sustainable and more realistic approach towards it.
I think your point here is that humanoids mimics human body.
And evolution has optimized human body for unstructured environment.
And so our fingers, our legs are not the best apparatus to do one thing.
For example, if our only goal as a species is to climb trees, we will not have this body necessarily, right?
So we'll have different kind of fingers.
But what humans end up having are evolved into is this body.
This body shape that can be very general, but not necessarily best at everything.
And that is for the survival of unstructured environment.
But from a business point of view, from a pragmatic
Technology point of view, that this unstructured environment and a generalized body is actually the hardest problem to solve.
It's not necessarily even the right way to solve the problem.
We specialize, so we take more
Specialized body to solve a narrower problem.
But the challenge for cynics is that to be more body agnostic so that their infrastructure can serve different bodies and different semi-structured environments.
A common lens to look at exactly this question is an economic lens, right?
Which is, if you compare it to generative LLMs, they can create prose or code 10,000 times faster than a human being, a bunch cheaper than a human being.
So the economic case makes sense because our brains aren't very efficient at that.
However, our brains and our bodies are very efficient at 3D navigation, right?
You know, like movies of the world are picking things up.
And so this is just a prediction question, but do you believe we'll ever be able to build robots, at least in the foreseeable future?
That have the power efficiency of a human being when it comes to menial tasks.
So let's say just basically, you know, minimum wage or something like that.
Like how far away are we from this?
Is this like five years or this is like never?
I think it's going to take a very long time.
So if you're really thinking about robots in real environments, in the end, it will always be a system.
So every working robot in the real environment is a system work.
It's
You need to be very mindful and thoughtful about how the systems are coming together the hardware the software the brain even like to the details of for example what's the friction coefficients of your fingers
So there's a lot of things you have to consider to make these things a reality and it will take iterations
But what i am excited about is that
I have always been at the state of the art of robot learning and also trying to push the state of the art forward.
But the state of the art is always moving faster than I expected.
So what I'm focusing on and trying to investigate right now is very different from, for example, when I started my PhD.
So this is a speak to how fast the...
Whole ecosystem has been evolving and all the moving pieces started coming together or building these robotic systems.
But we also have to be calibrated about our predictions.
So we will see a lot of progress.
But to achieve, for example, human-level efficiency and capabilities, it will take longer.
Martin, the hardest thing in today's AI is to have the right measured optimism.
Right?
It's totally true.
I mean, even LLMs does not have human brain efficiency.
Human brain operates on 30 watts.
Yeah, that's true.
So we are far from that.
But I mean...
Performance to power it may be close, right?
In narrow tasks like software engineering.
Like generating an image or software engineering that it is, right?
Yeah.
I think so.
I don't think we're anywhere close when it comes to robotics.
This has changed how you think about your, like strategically the level of ambition that your team can go after.
I mean, does it, is it changed that or is it still very much in line with what you expected to do?
When you started?
It definitely changed the trajectories in a very profound manner.
So we see a lot of unlock in being able to do this whole process through the modeling of the environments in a much more efficient and much more scalable manner, especially in partnering together with Word Labs.
And I also want to add to Fei-Fei, if you think about, for example, the current states of the language models,
So those are models that's with incredible capabilities.
But still, you don't just blind trust it to book your flight tickets or make your hotel reservations.
You still, hopefully, they're still a person who's reading the output from those language models.
But that is very different from how people and we will be using, for example, robotic models.
Based on robotic models, out of the box, the robot has to work reliably in the real environment.
And we don't even have the data.
We don't even have all the necessary infrastructures around those for the robots to just out of box work reliable in the real environments.
So for that reasons, be able to create this digital world
There's a scalable digital world where the robot can learn and evaluate within.
This is going to unlock so much more potential for being able to replace all the costly and unsafe data in the real environment with the data generated from the words for the robots to be able to do scalable learning and evaluations.
You know, I've seen many of kind of these kind of integrations.
They actually work very well at this stage when they have this much alignment, which is great.
But there's always like this question of do you integrate now into what's happening now or do you keep things quite separate and provide kind of like a long-term trajectory that will, you know...
Be realized in the year timeframe.
How are you thinking about this, Fei-Fei?
Is this something that integrates right away or is this kind of a separate longer term?
This is a great question.
I think at this point, you know, Yunzhu, Changxi, Sunny, Justin, Ben, and I have been talking about this.
At this point, we are going to
Take it thoughtfully.
We're not rushing to integrate everything from code base to Teams because I think Cinex does have a very...
Well-thought and I wouldn't call it standalone completely, but fairly contained tech stack, as well as their customers, as well as the kind of products they're building.
We're going to take time.
We definitely will.
We already have a simulation side as well as the potential base model, action condition model side.
We already are starting to talk.
And also, they are using Marble as an internal customer.
So we will be integrating, but we're not rushing to blend the team together.
As like a full salad bowl.
How are you thinking about geographies with this?
Will the scenics move?
Is it going to stay in the same place?
We're going to, Vindra is going to move.
Oh, well, welcome.
I'm moving to San Francisco.
Yeah, Florence to the Renaissance, perfect.
I think we, Aurora Labs is officially becoming a bi-coastal company where the headquarter is in San Francisco.
I've, you know, I live in Palo Alto.
I feel like I'm in a...
Different state.
But I'm actually excited that we're going to have an office in New York that can help us to attract talent on the East Coast.
And also, we have been talking about making sure that in both offices, we set up the robots so that we get to
Basically test out and mature our engineering stack so that we can work with robots remotely because we have to do that for our customers anyway.
So maybe just to be very concrete, Fei-Fei, maybe let's just pencil out, like what is the perfect success case in two years?
Like what product do you have?
Who's engaging with it?
How do they use it?
Just the crisp, like, what this becomes.
I would be very happy that Cinex team and World Apps team will have validated customers in
In a small number of important vertical use cases where our system, our infrastructure has proven to be truly beneficial to their automation needs.
And these customers became our
Lighthouse examples to scale our business.
How early, let's say someone listening to this is running a robotics company, at what stage should they engage with World Labs?
Is it really early on?
Is it somewhere in the middle?
So right now, for our customers, because we are building this kind of real-to-theem-to-real pipelines, where the simulation is essentially the words we're going to provide the training and evaluation grounds,
Some customers, they need only the real-to-theem part.
They want to digitalize the toxic care robots and be able to do the evaluations of their robotic systems.
Some customers need this real-to-thing to this entire pipeline, such that they will be able to have policies running on their hardwares.
So our platform is also designed in a way that is flexible, depending on what our clients need.
And at the same time, the clients we're working with are actually pretty close to the deployment stage.
So basically, they are working on very, very practical tasks.
Those tasks, when replaced, when we have robotic solutions that are there, can just create value.
Immediately and they have like at least like tens or hundreds of like this kind of situations they are thinking about to do the automations for.
So as like together with WordNaps, we'll be able to develop reliable solutions for those scenarios
As we have already showed, we have a number of scenarios already instantiated in our blog post, and we'll be able to further our investigation to see how they can actually solve the key requirements and also constraints faced by the real-world deployments.
Great.
I want to be very specific about this.
Is it ever too late or too early to call World Labs if you're a robotics company?
No, we want everybody to call us.
We want to learn about your use case.
Wonderful.
If you're listening to this and you're anywhere close to a robotics project or robotics company, please track World Labs.
Yes, thank you.
Definitely open for business.
We are open for business.
Not too early.
All right.
If you're doing robotics, call World Labs.
Thank you both very much for coming.
Thank you.
Thanks for listening to this episode of the A16Z podcast.
If you liked this episode, be sure to like, comment, subscribe, leave us a rating or review and share it with your friends and family.
For more episodes, go to YouTube, Apple Podcasts and Spotify.
Follow us on X at A16Z and subscribe to our sub stack at a16z.substack.com.
Thanks again for listening and I'll see you in the next episode.
As a reminder, the content here is for informational purposes only.
Should not be taken as legal business tax or investment advice or be used to evaluate any investment or security and is not directed at any investors or potential investors in any A16Z fund.
Please note that A16Z and its affiliates may also maintain investments in the companies discussed in this podcast.
For more details, including a link to our investments, please see a16z.com forward slash disclosures.