Hello, and welcome to the NVIDIA AI Podcast.
I'm your host, Noah Kravitz. Our guest today is Yves Jacquier, the Executive Director of Ubisoft La Forge, Yves brings over two decades of experience in technology innovation, science and technology, and R&D management, in different areas such as artificial intelligence, particle physics, telecoms, biomedical, and video game industry work.
He's a generalist, which he considers a strength and has led many innovations at Ubisoft, including performance capture, the use of telemetry or biometrics, And he's also brought on new roles, including data scientists and machine learning experts.
Eve also developed Ubisoft's academic R&D strategy with two significant milestones. creating a chair in AI deep learning in 2011, and also the foundation of the first lab in the gaming industry dedicated to applied academic research.
Ubisoft LaForge. We'll be talking about the intersection of generative AI in gaming and the latest technology that Ubisoft LaForge is bringing from the world of research to gamers worldwide.
So let's dive in. Yves, welcome, and thanks for taking the time to join the NVIDIA AI podcast.
Thank you very much, Rob. It's my pleasure to be here.
So if you would, why don't we start with you telling the audience a bit about your role as the leader of Ubisoft.
LaForge and what your team's working on right now, specifically with regards to generative AI.
Yeah, sure. So I lead Ubisoft La Forge, which is our R&D department for game production.
The mission of LaForge is to facilitate technical prototyping based on the latest academic results And as such, we've been working for a while on those topics, for example, in the field of voice synthesis, animation, or face modeling, to name a few.
However, the concrete usage of those technologies remains in the past pretty niche.
For example, voice synthesis. is one tool in the box for experimented voice designers.
Animation synthesis went the same route.
What has really changed recently is that, for example, chat GPT or 2D image generators like Stable Diffusion or Meet Journey are now accessible for everyone.
So without having graphical skills, I am, and believe me, I have no graphical skills.
I'm able to generate beautiful images in a matter of seconds while also generating convincing expressive dialogue still requires very specific skills.
So in other words, we're focusing today on two aspects.
What are the most valuable usage for us beyond the toy or demo effects? and how to embrace those technologies into our pipelines while doing this in a responsible and fair manner.
So as you mentioned, as we're recording this in late February, much of the world or much of the world who's interested in this stuff has gotten access to things like ChatGPT and Stable Diffusion Mid-Journey, other text and image gen AI creators.
I guess I sort of have two questions. One is, as much as you're willing to share, sort of how far ahead is an organization like Ubisoft La Forge? ahead of the general public when it comes to access to these things?
In other words, are there tools that you're using that the rest of us might see six months down the line, a year down the line, two years down the line?
Or are you kind of working with generally the same stuff that the rest of us now have access to?
Well, first thing I should say is that we've been working on those topics for years now.
If you're going to our YouTube channel, for example, you might probably notice that We were out of the first in the industry to create speech tube facial animation technologies, for example, and we implemented that into our actual games.
So today, no surprise, we're working and we have been working on everything that can assist our creators.
That's really at the core of what we're doing. and technologies as well that enable a new form of diversity into games, you know, making our 3D worlds more believable.
That's the second aspect. And finally, we have the third axis, which is working on everything that helps to improve the experience of the player.
Let me give you a couple of examples. On the game creator's assistant side of things, we can think of text-to-speech, for example, how to treat Expressive voices are synthesis, which is key in making games. for more believable worlds is creating more variety in faces.
For example, how do you treat a believable crowd that still feels believable without adding too much work on the shoulders of the creators.
And finally, in terms of clear experience, it's getting into many different topics, but One of the very important for us at the moment is being able to provide a safe environment for players.
And we're working hard on the toxicity detection and improving our model.
So these are definitely the three kind of access that we've been working on.
When you talk about toxicity moderation, I'm not sure what word you used, but is that more on detecting things that human players might be doing or saying that are toxic?
Or is that more on the AI and NPC side to make sure that the non-player characters aren't doing things you don't want them to do?
Well, for the moment, we've been focusing on chats, basically.
We just want to make sure that Everybody's having discussions online and playing online will find a safe environment.
And today we have to admit that at an industry-wide level, half of the players have been confronted with toxic behaviors.
And basically the idea is simple. you'd only need a few people with toxic behaviors to harm the pleasure of the men.
But will generative AI and the experience of generative AI will see probably a huge need to moderate content.
And it goes twofold. First, you want to make sure that internally, for example, or even for user-generated content, you do not create things that are copyrighted, as simple as that.
Even if it's not the intention of the creator using that, If you generate, I don't know, let's say 1000 heads to create a believable crowd, you might accidentally create heads of celebrities or whatever.
So I think that it's going to be one key aspect that will be developed in the future of generative AI. is how to make sure it does not generate things that you don't want to find in a game or in any type of content.
It can be whether on the copyright aspect of things, but also obviously as well on the, you know, assets you might not want to see in your games.
And we know that it's a huge challenge for our friends from Roblox, for example.
In your opinion, what does a workflow look like right now?
What are the ways that generative AI can best assist creators in creating work right now?
Ah, it's a huge topic because it truly depends how it's been implemented.
So... I think the best way, which is also pretty conveniently our approach is is to include the end user in the process by design.
Because this type of technology raises many complex questions.
First, what is legal? What is a fair use?
What skills will become obsolete? At what pace?
And I think that today, nobody has a clear answer on that.
So, I feel that it's fundamental to implement a human-in-the-loop approach.
And I see three reasons to do that by design.
The first is the most important because For us, it's the right thing to do.
Technologies should be augmenting people's skills, but never try to substitute to it.
The second is that with Vanilla Generative AI, everybody will get the same results.
So from a pure business standpoint, if you want to differentiate, you need to do more and only the human creativity will enable that.
And finally, we need to remember that although those technologies are evolving and being adopted at a pace that has no precedent, they are still not mature.
If you think that a large language model, what we call LLMs, They hallucinate.
Sometimes they say things, you know, in a very assertive manner, but still it's plain wrong.
And it has recently cost $100 billion to Google, apparently.
Some 2D images technologies rely on copyrighted material.
We know that there is a class action to data mine if they are considered as derivative or transformative of the initial material.
What it means is that if we want really to be impactful Beyond the demo effect, I mean, the human must be kept in loop, both in terms of technology, usage technology, but also the common sense of it.
The human judgment of when to use those tools and to do what?
So what are some of the ways that Ubisoft and Ubisoft Forge are using generative AI technologies right now?
We had a guest on the show, I guess it was a couple of years ago now, a creative director, an art director, who was talking about one of the big advances for him was instead of having to do all the grunt work of detailing topography on a map or building out a lot of non-player characters manually. he could direct the AI system to say much as you were talking about with giving prompts to something like mid-journey.
More recently, I read an essay that was using the metaphor of a conductor.
And that in the future, you know, instead of being a writer or an artist or what have you, a lot of the creative roles will be that of a conductor. because you're conducting these AI systems to generate all the different elements for you.
How do you see it and what are some of the things that are happening now at Ubisoft?
That's a great question. First, we think that it's an extremely powerful tool, but still it's going to be one more tool in the box.
It reminds me a little more than was 15 years ago.
Now, when we decided to implement the first motion capture studio, All animators within Ubisoft were thinking that we were trying to cut their jobs and they were scared that we wanted to replace animators.
Actually, what happened later is that When we successfully launched the motion capture studio, we actually hired more animators in terms of percentage of production. simply because we were able to raise the bar in terms of animation and we were able to populate open worlds, create the first Assassin's Creed, and et cetera.
Looking back at the motion capture experience, the first time we tried to use motion capture, We failed miserably simply because we were not able to find ways to implement that into the pipelines to reach two objectives.
The first have a real impact. that make creators want to use it more and more.
And second, do it in a way that support the transformation.
So that does not feel like intimidating, for example, or like a threat.
So today, from a strategic standpoint, that's really on this that we're working on.
Then more specifically, We are covering all the crafts at Ubisoft, to put it this way.
We're working on animation, for example, with things such as lone motion matching.
The objective of such technologies is to make sure that we have more believable and nice seeing animations, for example, that adapt to any type of rough terrain.
Today, we know that metric design is basically a way to say that a wall that you can lean on must always be two meters high or things like that.
And obviously, when you want to create more believable words, you need to, you know, progressively maybe get rid of some aspects of those metric design and to do so, all your assets like animations, physics, and all that must comply to this new reality.
So for animation, it's really something that we're working on to help us create more believable words, then there's everything that's more related to the diversity and the size of our words.
We have, for example, voice synthesis. A typical AAA game now includes more than 1,000 lines of dialogues, and that's for English only.
There's a general misconception that it comes only because the worlds are... getting bigger and bigger with maps that are the size of a continent and all that.
But not only that, were trying to make them more diverse.
For example, in the latest Assassin's Creed, the main character could be played either as a male or a female. which means doubling the lines, basically.
And because the standard approach do not scale.
It leads today, if we're not pivoting, to an explosion in terms of production timelines.
And second, Actors fatigue. I mean, it's a very physical activity.
I don't know if you saw voice actors it's really physical when they're playing the roles and they're really into it.
And when you have, you know, the number of lines explodes, it's very difficult to ask for more to that.
Right. That makes sense. You mentioned motion capture and that leads me to a question that I had to ask you.
Can you think about, you've been working in the industry for a while now and you've seen a lot of technologies, a lot of innovations come and go or come and stick depending on the innovation.
Do you have a sense of where generative AI sort of ranks?
And I hate to ask it that way, but do you see it? as a huge deal that's going to transform the industry?
Do you see it akin to motion capture or another technology that You mentioned when motion capture started, the animators were all afraid it's going to take their jobs. and it took a while to figure out how to use it, how to incorporate it into the pipelines or anything you're talking about, and then it turned out there was more work. for the humans.
Do you have a sense of how we're going to look back on this period of generative AI starting sort of in the spectrum of advances in the industry?
Well, future is always difficult to predict.
It's not a fair question. But I'm trying to observe some trends and it could be useful.
There's one trend, which is services wrapping ChatGPT or other similar services.
And for example, you might have seen that last week.
Roblox announced a feature to wrap their scripts with natural language.
So in other words, instead of coding, you could say, switch the car's color to red reflective paint, for example.
I actually missed that. I have no information on how it's actually done, but you can test something which is ChatGPT can already output snippets of Roblox script code.
If you input an instruction in plain English.
It means that, people will get soon very used to this type of natural interface to complex systems. which suggests to rethink completely some of our UI and workflows, for example.
So I see something unprecedented in terms of accessibility. people with no skills, no artistical skills, will be able to treat assets with an unprecedented level of quality.
Now it brings the second question, which is how to stand out because The fact that it becomes more accessible means that it will shape expectations as well in terms of user-generated content or avatar personalization.
The work of NVIDIA on neural radiance is a very good example of that.
I mean, the possibilities with a simple smartphone choose can very easily entire environments.
So I think we'll see more and more AI generated content.
It will be like, you know, the spell check in words.
You won't talk about AI. You will talk about features and because it will make things easier to produce the quantity might be overwhelming and it might be difficult to stand out.
The most valuable use cases will probably happen when generative AI will stand out not only in simple assets such as you know, text, 3D images or even 3D models.
But on combined assets, how do you combine the fully interactive and expressive characters or environments, for example.
And this will impact all assets. I mean, Dexter has mentioned the effects as well.
I think what will be key also is this question of moderation aspect.
If we do not take care of that in parallel at the same pace, we might hit some wall collectively.
I'm expecting to see improvements into this area, just to make sure that what is generated makes sense, basically.
But more importantly, I think that it's such a disruptive approach that only by saying that all assets will be concerned, I'm not even touching the surface of the potential.
Let me give you another image. When Google Maps was released back in 2005, I think, The value was not obvious.
It was a gadget. It was a fun gadget to see things on a map.
Then Uber was created in 2009. and soon disrupted an entire industry.
I don't think that back in 2005, someone could have predicted this.
Sure. But I feel that today, what we're seeing in terms of generative UI, the pace of adoption And the fact that literally every day you have some significant news, it will change a lot of things.
We're going to be amazed. Even if you're deep into the topic, you're going to be amazed.
Our guest today is Yves Jacquier. Yves is the executive director of Ubisoft La Forge, the first lab in the gaming industry dedicated to applied academic research.
Eve, when you and your team, when you're evaluating new technologies to use, Is there a process that you have for kind of going through testing a technology, building prototypes, testing the prototypes and sort of getting to a point where you've determined that something might be ready to actually incorporate into your production pipelines?
Sure. The way we do that is we bridge the academic world and the research with Ubisoft people.
Basically, we focus on prototyping. We focus on the proof of concept to test the technologies in... make sure we're able to scope where they make sense and when they simply do not work in our context.
And for that, we have three secret ingredients that I'm going to unveil.
Love secret ingredients. So the first is the obvious one.
We have a huge portfolio of games and cutting edge in-house technologies, which means that We have many potential applications to our prototypes or a large variety of problems to solve it, if you look the other way around. that translates into opportunities where, for example, the markerless capture prototype might not be suitable for a game like Just Dance because of performance issues, but it could solve problems in our motion capture pipelines, for example.
The second is that we adopt an interdisciplinary approach.
So we mix artists, AI researchers, people from social science, to work on a single prototype.
Not only does it help for adoption, it is the most effective way for us to provide diverse inputs and diminish the risk of having a blind spot.
But finally, we have a long history of making games which means that we have tons of unique data and the capacity to create specific data sets should we need it so we have released such data sets for example on open source on our github and that's true for Our game assets, that's true for our lines of codes.
We have 20 years of lines of codes that have been used to train AI.
Along those three unique ingredients, so the different needs and opportunities in games and technology, interdisciplinary approach in our data, we have set up a process where anybody in the company or even outside of the company can propose a project and follow a simple process with seven questions.
This process basically helps us to focus on real-life potential application while also making sure that we're not reinventing a wheel that we could simply purchase or license or code.
And finally, we have developed a strong culture of taking risks, and it means that the only failure is when the prototype does not work. and we don't know why.
And the rest is called learning and sharing our learnings.
It might sound cheesy, It is effectively practical from a research standpoint because people feel encouraged to think differently, which led to some breakthroughs such as learn motion matching.
But by doing this, we also prevent too many people in the company to try the same things.
We capitalized on our learnings, if you will.
I remember that one of our prototype of photogrammetry miserably failed.
Still, by learning and sharing our learnings, we prevented two other productions to go the exact same route that would have led to the exact same results.
Right. So basically, that's how we do that.
And then we test-flight the prototype in real life.
So in your experience with generative AI in particular so far, is there something or maybe a couple of examples that pop to mind as... whether it was a breakthrough or perhaps something that I don't want to use the word failure, but maybe, you know, a prototype didn't work and you didn't know why, but you were able to apply it to something else.
Was there something that came along the way that really surprised you that was not an outcome you were expecting or... a tool you were testing for one thing actually proved capable in another arena.
Something about your work with generative AI so far that jumps out as surprising in some way.
Yes, definitely. It's in the domain of voice synthesis.
As I mentioned, we've been working on voice synthesis for a while and especially having more expressive voice synthesis.
In other words, If we're all used to Siri, Alexa, and they're very high quality, But it's neutral in terms of voice, and that's not what you would expect in a game.
You want a voice that's coming from far away, We want people who are angry, happy, and all that.
So we've been working hard on that. Another aspect is the ability to generate barts, people who laugh, for example.
Today, Siri is not able to laugh, basically.
And obviously, in video games, you want to treat people who cough, who laugh, whatever, shout out and all the expressions.
So we were working on this prototype and it was really difficult.
And at some point, we realized that the prototype was not working very well for creating laughs and barks and everything.
But we found another application, which is voice conversion, and we were able to create a module of voice conversion with a very high level of quality.
What is voice conversion? Voice conversion is the ability for someone to speak with someone else's voice.
And there are many ways to do that, but sometimes it's difficult to... gets out of the initial intention with which the voice was recorded, for example.
So there are a couple of famous examples where the voice of Darth Vader has been clones of that.
In the future, it can be used in further series or movies.
And that's voice conversion. actually were working on a prototype for text-to-speech, and then suddenly, it was in 2020, COVID happened.
Then studio closed and we're in the middle of shipping Assassin's Creed Valhalla.
And we don't know if we will be able to record the main actress anymore.
Okay. On top of that, at the time she was pregnant.
So we didn't know when a point we would be able to work with her.
Right, right. So a lot of risk. We piloted with this prototype of text-to-speech to create this voice conversion technology, which was surprisingly good And by the end of the production, we didn't have to use it.
And it was done very transparently with the actress as well for all those reasons.
So we were lucky enough not to have to use it, But I was at that moment amazed by the fact the prototype did something way bigger than we would have imagined. and had the potential to save the shipping of a game like Assassin's Creed Valhalla.
Right. That's amazing. Looking ahead, you mentioned transparency and you had mentioned before when we were talking about generative AI and copyright issues and things like that.
Brings me to one perhaps kind of final question before we get into wrapping up.
But looking ahead, I'm not going to ask you to predict the future again, but kind of looking ahead and even now, I'm sure you're having these conversations.
What do you see as kind of the big ethical concerns or conversations when it comes specifically to generative AI in the gaming industry.
What are some of the conversations that are being had?
And do you see the industry as a whole kind of settling on best practices, code of conduct for using this tech?
Do you see it kind of being on a company by company basis?
We can talk a little bit about that. That's a very tough topic.
Also, I cannot speak for the entire industry, of course, but I can tell you that- You're the only guest I have right now, I'm going to ask you.
What I can say is that we're already discussing with some of my counterparts or other leaders of the gaming industry, because we're all you know having the same questions uh we all love our creators uh so we want to make sure that All of us, we do that the right way within our different organizations and cultures and all of that.
First, there is the legal frame, all the legal frames in general.
And as you know, they usually tend to come late at tech parties.
But still, There are a couple of actions that we are all in the industry scrutinizing, and especially the one that will decide whether what is generated by mid-journey or stable diffusion is considered as derivative or transformative.
To simplify things, it's like asking in the music industry, what is a sample contributing to a new piece versus what would be pure plagiarism.
This part we have no control, but it might have a huge impact first on what we can and cannot do.
More importantly, obviously there are people and I'm sure it's shared in all the gaming company, because this technology is disruptive.
So we do not want our people to be or feel They're like the next Nokia when the iPhone came out or fancy drivers when Uber kicked in.
Right. So as such, at Ubisoft, we're leading two initiatives.
The first is to come up with an internal code of conduct.
How should we use those technologies to ensure we place humans first and always in a fair manner?
This is true for our employees, but also our partners. as we're working on fair conditions for voice or motion capture actors, for example.
We will always need them. We need to provide a fair compensation to ensure they can and will continue to make a living of it.
But the second is to create specific trainings, playgrounds, and toolings for our creators to help them acquire new skills, for example, to get the best out of these technologies. and create assets that stand out of what everybody can do.
So I cannot tell if it's a best practice, but I said, sincerely, I hope that it will be a shared practice and not only within the gaming industry.
And that's a kind of discussion that I have those days with many leaders in the gaming industry.
All right, final point to wrap up on here.
What are you excited about going forward?
I'm sure there's a million things. But but as we're talking about all of this, is there something that you're working on now or you're looking ahead to working on later this year when you have some time to or even just kind of a a bigger, broader topic, you know, pertinent to generative AI in gaming.
So something that's got you excited about the future.
It's what we have not identified yet. It's this kind of thing that we're building now.
For the moment and in the last five years, generative AI was used in a very niche manner by experimented people from different craft to accelerate them.
And it's getting to a new usage and orders of magnitude in terms of impact.
So what they will be able to do once they're able to connect all the dots, basically.
That's something that's really exciting.
And to achieve that, it's really, How do we lead within Ubisoft's transformation in terms of skills, pipelines, to help the creators use and adopt those technologies and new ways of working so that they continue to surprise and amaze us.
That's really what gets me excited. We're leaving... unprecedented moments, I think, in technology in general.
We really are. You know, it's nothing new to me to talk about, but I was reading some things over the weekend.
We're creating this on a Monday that really kind of opened my eyes up just that much more to think, yeah, the next 15 years are going to be a trip.
Excellent. Yves, for people listening who'd like to learn more about Ubisoft Forge in particular, I know there's a website.
Where can people go online? Yes, exactly.
The easiest is to go to laforge.ubisoft.com where you will find some of our publications, some demos and links to our open source data sets or YouTube channel. or to follow us on Twitter on Ubisoft La Forge.
Great. It was a pleasure. I feel like this is one of those conversations I'm going to want to follow up on in, you know, a couple of years time.
Look back and think, yeah, the Uber for Google Maps moment in the gaming industry with generative AI.
I didn't see that one coming, but it sure is interesting.
It would be my pleasure, Noah. And maybe we can rediscuss that and see where we were wrong and right.
Absolutely. Well, best of luck for everything you're doing, and we look forward to all the titles coming out of Ubisoft in the months and years to come.
Thanks, Noah. Prepare to be amazed. Thank you.