Hello, and welcome to the NVIDIA AI Podcast.
I'm your host, Noah Kravitz. One of the things that modern AI is very good at is translation.
Large language models showcase AI's prowess at translating written text between languages.
But machine translation has been a useful real-world tool for some time now.
That said, translation apps and auto-generated video captions are great.
But what if we could use AI to automate overdubbing of voice content into different languages?
High quality, low cost dubbing could do wonders to open up more of the world's audio and video content. to broader audiences across language barriers.
Our guest today has been working on just this problem at the Israeli startup he founded with his brother.
Ophira Krakowski is the co-founder and CEO of DeepDub, an AI-driven dubbing solution used in Hollywood movies that are currently showing in theaters globally.
DeepDub aims to bridge the language barrier and cultural gap of entertainment experiences through high-quality localization at scale.
The company's deep learning power dubbing solution helps studios, broadcasters, and distributors with all of their localization needs.
From translation and adaption, continuing to dialogue creation and finishing with the final mix.
Ophir is here to tell us more about DeepDub and how generative AI can help connect the world by breaking down language barriers.
So let's get right to it. Ophir Krakowski, welcome, and thank you so much for joining the NVIDIA AI Podcast.
Thank you, Noah. Thank you for inviting me.
And I really looking forward to talking to you.
Likewise, it's our pleasure to have you.
So let's get started by hearing a little bit about the story behind DeepDub.
If you would take a moment, tell the audience about the company and how it got started?
As some may know, I've been the head of the Machinery and Innovation Department in the Israeli Air Force, so I've been used to doing stuff that are impacting the life of people in Israel.
But as I got out of the Israeli Air Force, I thought that I can do something with my knowledge to impact humanity in essence.
And I searched what I'm going to do. And I founded this company with my younger brother, Nir.
And in an essence, what we found is that Content is always created in one language.
And when you want to transfer it or to reach more audiences, you are facing a language barrier and a cultural barrier.
And the current language models currently don't support some of the use cases like jokes, idioms and places and even regular phrasing that use on jargons that go on the streets nowadays.
And in essence, they don't support the delicate intricacies of a language.
And for in this essence, we thought we can develop something which was very ambitious three or four years ago to take this problem and develop a technology and a product that would enable everybody to use it.
We figure out that if we can do a theatrical for one of the Hollywood studios, then we can serve everybody down the line.
Right. Listening to you talk about it makes me think of the old, it makes me think about how old I am, but the old sort of trope, at least in America, of kung fu movies that were overdubbed.
And you would see on screen the actor's lips would start moving.
And then maybe a second later, the dub would come in and it wouldn't match up at all.
And to your point, there's so much more than just the words and the translation, which There are plenty of nuances in language, but you get into, as you said, idioms and slang and just the cultural backdrop of how the words are being spoken.
And there's a lot to it that from my standpoint anyway, I believe you when you say it's tricky.
So how did you how did you get started with it?
Now you mentioned both you and your brother come from backgrounds working with machine learning and AI.
In what sounds like quite a different context, how did you make the move from you were serving in the Air Force, your brother was working at an intelligence agency in California, in israel yeah so how did you how did you get the idea to move from that and then how did the actual transition take place to starting deep though So both of us are a true lover of the magic of creation of content.
So this is a part of like in an hobby. And building this company is just combining our abilities with an impact in a good way that content can be accessed to a large audiences.
And so combining the love to creation of content with a love to technology, this is something that led us to found this company.
And were either of you working specifically on synthesize voice and language translation problems, or was it kind of a move using your background in AI and deep learning, but a move into kind of new territory for you.
So this is a very interesting question you ask because most of the technology of creating voice was dominated by four big companies. three or four years ago by the big, you know, meta Amazon. on Google and Microsoft.
So in an essence, the knowledge was not out there.
We had totally to invent everything from ground up and take the most recent research and incorporating the company people that there are from these companies and to build the know-how.
We actually not only build the product, but we also build the AI infrastructures to train new models.
So in an essence, this is what we did. And we built something that As I told you, we built something that can screenshot and allow users technology to create a voice, a human natural voice which the right pronunciation with the little intricacies of the human voice.
So in essence, you cannot differentiate between a human voice and a machine generated voice right which brings up some questions that i'll i'll put a pin in to ask you a little bit later but I want to ask you now about how does the solution work when you work with, and I don't know what the best way to get into this is. feel free to take a different tack.
But when you're working with a client, say a Hollywood studio that wants to leverage your tech to dub a piece of content, a movie into different languages.
What's that process like? How does the solution work?
Let's talk about the numbers. Let's take, for example, a TV series.
So we'll have like 10 episodes. It will take you like three months to dub it using humans, because in a full season, you'll have like 100 characters.
In an average, like a TV, a drama TV series, you'll have 100 characters.
You have to bring into the suit a lot of people.
You have to make sure that the of diversity, of DII.
So you need to take care of a lot of things when you are trying to dub the content or to make it available to other audiences.
And then you need to do it across 32 languages.
So this is not a technology problem. It's a huge project management problem, right?
So we thought that we can... make this process very efficient by introducing efficiencies in the entire process. not just creating the voices but also in the translation the adaptation and the mix itself so Just to make the audience aware, the process is very convoluted because it takes a lot of phases.
In the first phase, you just get from the studio as you just if the customers bring us a movie or a TV series, then we get... the video and the audio and the audio is most of the cases it would be in English.
And then you have to transcribe it because you need to know who said what.
Right. And then you have to translate it and then you need to adapt it because for example, a joke that it will work in a different language.
It's not a straight translation like the regular translation tools that you have.
And then you have to create voices. And after you create the voices, you need to mix the audio with the music and effects and glue it back together to the video.
And this you need to do in every language, like in 32 languages.
32, yeah. Yeah, so taking this into account, it's a very convoluted process.
A lot of time, a lot of people involved, a lot of different parts to the process.
Absolutely. Right, so actually what we've developed is a platform, a web-based platform that enables people to interact with a very sophisticated AI models that enables them to do each part of the process in a very fast manner.
And this is for the high end. Maybe later I will elaborate more how this can affect more classic because we are currently developing something that will enable other customers to work on the platform and not just the studios. not just the studios.
But when you work with a student, we need to understand a studio doesn't want a machine learning He wants a white glove service because he wants it to be in the highest quality.
And the best AI model currently available cannot get past 95% of accuracy.
Even the best chat GPT have mistakes. So you need the human.
The human in the loop. And this is why we have developed the platform.
So it's a kind of a... Adobe Premiere kind of a tool, but for creating localization and creation of voices, which is very simple.
So it will take someone one to two hours of very simple training to deliver a drama episode.
Would your platform take care of every part of the process you mentioned? transcribing all the way through to Final Mix?
Yes, definitely. You know, We understood at the first stages that nobody cares of creating voices.
Everybody cares on the end product. You want the end product.
You want the localized video, you know, Just bringing you the voices, you need to do all the other work that I just talked about.
Because even if I created the content, For example, if I have created a podcast, I know English, I know Hebrew, but I don't know Chinese.
So I don't know how the translation went.
Is it good? Is the joke works? So I need someone to help me curate it.
And so let's get into the human in the loop.
What is the human doing when they're localizing an episode of the TV drama?
If I want to localize that using deep dub, what are the steps that the the platform takes care of and where does the human come in?
I have a million questions, so I'll let you kind of talk through it and I'll jump in.
Okay, so in essence, the first cut is done automatically by the platform, but then a human comes in and verifies that everything works well.
So it verifies the translation, it verifies the generation of the voices, like for example, do the emotions are the emotions that needs to be at the target language.
In different languages, question is asked in different ways.
So in an essence, you need to understand this and a model sometimes know how to do it and sometimes mistake on it you know since we're recording everything so we're recording the curation on the text, recreation of the voices, we are making the machine learn from these mistakes and become better and better.
So over time, the intervention of human in the process will be less and less needed.
What's the most difficult part of the process for the machines to take care of?
Is it getting nuances of the language? Is it getting the voices and the emotion right?
And I guess I'm asking both in terms of what was tricky for you and your team or is ongoing building and refining the platform?
And then just from sort of a purely technical aspect, is there one part of the process that's just a harder problem to solve objectively.
You know, I believe that building a machine that can support wide range of emotions from text This was to us the first problem that we tackled.
And it took us about two years to develop something that will support a wide range of emotions that can support a theatrical.
Right. Which is not like a podcast or audiobook, which is mostly a narrow range of emotions.
So in a chat room, you have somebody screaming like from the heart or talking while eating.
This is a different... It's funny, but it's a different voice.
Like you don't understand it, but the machine looks at it as a different voice.
It's like... Digitally is a different voice, but we support it now.
So I cannot tell you this is not a difficult question, a difficult issue.
But the most difficult issue is, I think that is currently not solved, is the translation part.
Translation part, I think, as I believe it, as I know the material, I think that getting a 100% translation from the machine that will 100%, for example, a joke, it will take some time.
Yeah. Does your system analyze the video content as well to pick up on facial expressions or body languages or even... patterns in the way characters are situated in an episode.
So it might give some cues as to what's happening, what the subtext is, or is it strictly looking at the audio?
So we're looking at all aspects of, it's called multimodality.
It's looking on the video, the audio, and the text.
The fact that we're using LLMs is that the actual LLM understands the emotions and understand the text.
It's kind of a, I don't know, this is a very simple way to explain it.
But it's actually understanding the text.
And it can convey the emotion to different languages, the emotion that is in the text.
In an essence, we even use the video. We have the technology.
Currently, it's not a product that is out there, but we have it in our labs. that changes the lips just in the places where you cannot do it with words.
If you are translating from German into English and you end up translating the word no, So in English, a no is an open mouse, but in German, it's nee, it's a closed mouse.
And it's a short word, and you don't have anything to do, so you have to struggle when translating it.
So in an essence, in those places, we call it the last mile solution.
You would use a change of the video. But this would be available soon, this kind of solution also on our platform. we currently it's currently supporting 1k and as we are aiming always to deliver first to theatrical, which should support 8K on 4K.
So as we support this, we'll have this enabled also to other customers.
So currently it's audio dubbing that you're doing. but you have been thinking about and you just alluded to the solution that you're working on that does actually You know, if the right word is generate or recreate some of the video content as well to match the new audio.
Right. Yeah. Right. It's currently the generative AI model support all of this.
And this is how they are going to impact the entertainment industry because it's not just the audio.
But I think the audio part is one of the audio translation or voiceover It's a tradition of 100 years.
Yes. From the Mussolini times, it was starting to dub content And until this day, they do it.
But most of the content around the world, we need to understand, is not dubbed.
So it's not accessible to most of the people around the world.
I don't know if this podcast is localized.
Not so far as I know. No, I mean, we show up in some of the international podcasts. the analytics, right?
We have listeners internationally, but so far as I know, it's not localized.
Right. So most of most of probably most of the, you know, Latin American audiences which some of them don't know English very well.
This podcast is actually not accessible to them.
They just tune in because they like the way my haircut looks on the radio.
So that's, you know, that's why they're tuning in.
And your voice. Yes. I haven't seen, obviously I haven't seen the solution that you mentioned that's not out yet.
But I have seen some early attempts from other companies at doing what you were describing changing I'm pointing, nobody can see me, I'm listening to the show, but I'm pointing to my mouth, changing the way that lips look to accommodate different audio overdubs.
And they look really creepy to me. The ones that I've seen, and probably it's new technology, it's not there yet, But it's just this weird, uncanny thing of the rest of your face not moving. or just having a different expression, but then the lips are clearly doing something different and it matches the words, but doesn't match the rest of the person's face.
I can only imagine how difficult and tricky that technology is to develop and get right, but it's... interesting times for millions here to say the least.
Yes, yes, but you know, with audio itself, you can solve a lot of problems.
And since most of the audience are currently used to get the dubs in a way that it doesn't interfere to them if there is a little Except for the US audience, but other audiences around the world, which are used to dubbing or voiceover for a lot of years,
They are using to consume it in the way that it is right now, only the audio.
English audiences, this is very important because, you know, audiences are not used to dubbing content.
And this is very interesting because since Netflix started to dub content, international content into English, there are more and more Americans that are exposed to international content which i think is very good because they are now open to more cultures around the world.
In fact, it's interesting because most of the content we've dubbed is into English.
So we have hundreds of hours of international content dubbed into English.
Oh, interesting. Okay. Yeah. My guest today is Ophir Krakowski.
Ophir is co-founder with his brother and CEO of deepdub.ai, an AI-driven dubbing solution that's being used on movies and TV shows and other content globally, as Ophir was talking about, starting with theatrical releases, working with big studios and then kind of trickling down as the technology becomes more mature and I would imagine less expensive to use going forward.
So we've been talking about the company, about the technology and the importance.
And as you were just saying, unlocking not just all of the English language content exported from the US to other cultures, but also kind of moving back the other way with streaming platforms.
And as you mentioned, there are a couple of Japanese shows on Netflix that my family and I watch from time to time, and we just watch them with actually now that I'm thinking about it, some are subtitled and then some do have English, at least the narration is in English.
And so to your point, the more, from my perspective in the US anyway, the more non-US content we're able to get and enjoy here beyond the enjoyment it does open up you know open your mind up to other cultures and see how other other people in other parts of the world do things, which is hugely important.
What are some of the other things that you either might be working on or just might be thinking about going forward that breaking down language barriers like this and being able to export content in a format that I think is more natural to consume.
And then perhaps it's also able to convey the original intent a little bit better than just subtitles or just overdubs might do.
What are some of the applications that you might be thinking about or even working on that listeners might not be aware of.
Yeah, so it's very interesting. There are a couple of use cases that were working with studios.
For example, working on creation of voices.
For example, creating diversity of voices when you have a very creative actor and he wants to do several parts, We're enabling him to do several parts because he's very creative.
He has his ideas of how to convey this or if there is a director they know exactly how to say this piece but he wants to create it in the specific way And there are some directors, so we're enabling them to do this with the technology.
Another use case is doing a screen test For example, currently, only screen tests for a new movie or a new TV series...
Testing the jokes, testing the plot is only done in the US.
And then all the world is just, you know, if it works, it works.
If it fails, a bummer. You know, you lost a lot of money.
Why is that? Because it costs a lot of money and it takes a lot of time and you are doing it at the first stages of creating the movie or creating the show. without technology you can do it very fast so you can do it across four continents and then you can feel get a feel if you're content will work globally and not just in the US.
And this is only an entertainment But when you go to advertisement, you can make an advertisement that will work.
For example, you take Latin America. So different words in Mexico and in Argentina.
And currently, I don't know if the audience know, but in movies, they use Latin American Spanish.
It's a natural Spanish. if you go and talk in the street in Argentina and in Mexico, it's different Spanish in the sense.
Right. But, In movies, they just flatten it in terms of cost.
You want to do it in one time. With our technology, yeah, definitely.
So with our technology, they can do a version, a Mexican version, and an Argentinian version, a Bolivian version.
Different version, and you can use the actual language and the actual jobs that will work in each region.
That's terrific. And this is only entertainment and you can go into advertisement and even, and you go down the line, And I think this is the most impactful that this is our goal.
E-learning, edutainment, enabling people from places where dubbing is not cost-effective to access knowledge.
So people that understand English, they have access to a lot of knowledge.
But if you don't understand English, A bummer, you cannot access most of the e-learning content that is out there.
Most of the YouTubes are in English that you You can learn from popular science and even to learn about technology.
Part of the triggering to build this company was that my brother worked in a company in Brazil.
And he worked in cybersecurity and he has people from Brazil and he wanted them to learn cybersecurity deeper.
And he told them, just go on this platform and learn it.
And he understood they cannot do it because it's very difficult to hear something in English because you are you're concentrating on translating instead of learning.
Instead of learning, yeah. along the way i learned that there are good people good people that create fantastic content But unfortunately it's not in English.
Yeah. They have like millions of subscribers in their own language.
And nobody else can consume it. So for example, I've come across a French guy that has like a popular science kind of a YouTube channel. channel.
It's very interesting to learn that his content cannot reach other audiences.
So in an, in an essence, technology can enable people that create good content be connected with people that want to consume it, but they want to consume it in their own language.
And if you go to children, you know, children don't know other languages.
So I actually want to access them early age and and enabling them to get to very sophisticated kind of content or knowledge because knowledge is now in the internet.
And we need also to understand, and this is something very important, that the youth currently, you know, we used to read a lot.
I used to read a lot of books I'm on Facebook, but my kids, they are on TikTok and YouTube Shorts. and Instagram.
They are not on the word side of it. They are on the audiovisual side of it.
Yeah, when I when my kids want to learn something, they actually it's actually funny because I noticed it first when we would be talking about something and we'd wanna look something up.
And I would go to a search, I'd go to Google and I'd type, And they'd kind of look at me like, why aren't you going to YouTube?
Like, that's where you learn things. You go to video search and you watch a video.
Yeah. So your point is very well taken. Absolutely.
Right, and if you are a company, I don't know what you are doing, but when I go to the website, and I search for our product, and I go to the website, and I see a product tour, a video of the product, this is the first place I would press.
You know, I want to see it in one minute, just understand what it's doing.
Instead of reading all these words, Nobody has time.
And we need to understand, sub doesn't work.
In the hectic times that we have, you know, I'm hearing on my phone something.
I have like... I have a kid that has like... three streams.
He has his, you know, the phone, the tablet.
And the computer and the television is all on.
And I am amazed how they can split their attention.
But they are doing this. You cannot do it with subs.
Thank you. Right. Yep. No, you can't. So looking ahead, we're recording this just about in the middle of the year, recording it in mid-late June here of 2023.
But where do you see this all headed in the next couple of years?
Is it a matter of just refining the technology and being able to, you know, build usage and then bring costs down to make the tools more accessible, as you said, not just to, you know, the super high end projects with the studios, but other users as well?
Is there something else that you're working on?
Where is all of this headed in whatever the timeframe is that makes sense to talk about?
Yeah, so in the near future, we're working to democratize and enable people to access the technology. in a very affordable manner.
And in a way, they can... have their content be localized and accessible.
That's accessible. And this would be a huge impact because even people that don't have resources can do this kind of change or reach more audiences in an essence.
But in the future, we'll also enable a real-time translation.
So the technology will enable you to real-time translate an audiovisual content.
So you have more uh news and uh more sports that will be available and that's what i was wondering about yeah Yeah, it's a matter of compute strength and advancing the translation part but as this will go you know this technology will be available uh i think that with our knowledge and understanding of content which is not just creating a dialogue There is a lot of companies creating dialogues.
It's not just creating a dialogue. It's just handling the end product. which is the actual content.
I believe that technology will enable people to enjoy content that is created around the world and it will really globalize the storytelling and knowledge of people that which currently is bound by the language barriers.
Fantastic. Well, the company is deepdub, deepdub.ai.
For listeners who want to find out more, There's the website.
Are there other places, social media, blog, other places that you would send listeners to to find out more about what you're doing?
Yeah, so we have a blog on our website. We also have a newsletter that you can unlist and you can also meet us on the social media. our linkedin page our twitter and instagram page so so you can find them and and register and just follow our company page and you'll get to be the first you know the the new advancement that we are preparing for you fantastic Ophir, it was a pleasure.
And it's really incredible stuff you're working on.
And we just kind of scratched the surface of the possibilities.
Yeah, anything that knocks down barriers and helps people understand where each other are coming from and shared experiences is a win in my book.
So congratulations to you, your brother, your team, and best of luck on all of the work you're doing at DeepDub going forward.
Thank you so much, Noah. It was a pleasure talking to you and a pleasure to serve humanity in becoming global.
Thank you. Thank you.