Welcome to NVIDIA's AI podcast. AI systems have been trained to take photos and transform them into the style of famous artists like Van Gogh, Picasso, Turner, and Munch.
Take a photo and pick an artist's style, and what emerges looks kind of like some missing work of an artistic master.
And now, AI is heading in a different artistic direction.
That soaring music is an original, copyright and all, composed by an AI system developed by our guest, Pierre Barreau, head of Luxembourg startup Aiva Technologies.
Pierre, welcome. Thank you. Now, I know you composed that piece of work for NVIDIA for the GTC conference.
What do you call it? We call it IMAI, which is the title of the story as well.
All right. We'll put a link to the video behind which the music plays.
It's an original composition, and this was not played by computers.
It was actually played by a full orchestra.
How does an AI system compose original works?
Right. Basically, we have taught an algorithm of deep learning. to extract rules of music from 15,000 pieces of classical music, so music written by Mozart, Beethoven, Bach, The algorithm learns on those partitions and creates a mathematical model to then recreate totally unique music.
And the idea is that because with Aiva we create written compositions, then we can go into a studio and record with humans.
So it's really about the interaction of AI that is really able to compose really fast, good themes. and then bringing it to the next level with human recordings.
What caused you to go down the AI plus music route?
Are you a musician yourself? Basically, I come from a family of artists.
My father is a film and music producer. my mother was a singer i mean everybody in my family is a musician i'm a self-taught pianist and I also studied computer science at university.
So basically, I got this idea of using my technical background and my musical background and bring them together to build this artificial intelligence that I compose music.
Did it work from the beginning, or how hard was it?
I mean, did you come up with some Mozart-sounding things, but that just didn't quite cut the mustard?
No, it never worked for the first time. And it's a very long process. to actually come up with these architectures that work for music because music is a thing where you cannot have this metric which tells you, you know, this is good music or this is bad music.
Like you have these kind of metrics for classifications tasks.
And for generation of music, it's really hard to actually get an algorithm to create compelling music because you don't have these metrics. to help it go in the right direction.
So what did you train it on? I know you said you trained it on Mozart and Bach and Beethoven and other famous composers, but What is it learning?
Like you say, how does it become a good composer?
So it's basically trying to extract high-level features of what makes human music human music. so without necessarily focusing on a specific style.
And that's really what's important because if you think about how humans learn music, I started learning the piano playing Bach or Mozart or Beethoven just like everybody else when they're young.
And then you sort of listen to maybe jazz music on the radio and you get influenced. and your style changes, and then you go about to create your own compositions.
You just make an abstraction of what makes Mozart's music Mozart. and what makes jazz music a jazz music, and you just learn the high-level features of music and then use them to recreate new stuff.
So I understand that, and this is, I think, why in some cases the example of copying the style of Van Gogh or Picasso or somebody, that's –
Kind of more easily understandable because it's image recognition, like here's the kind of brushstroke style, just, you know, emulate that.
And again, they're emulating artists. In this case, what are those human components or musical components that... need to be understood?
Is it tone? Is it kind of drama? Is it putting one note in front of another?
Actually, there's a very simple one, which is if you consider piano music, it's playing notes so that two hands can actually play it on an instrument and every human composition has this constraint. where the finger placements have to be such so that a human player can actually physically play it.
Essentially, you could create a lot of a lot of different compositions where human players wouldn't be able to stretch their hands.
Indirectly, the algorithm learns those features because the compositions that it learns from were created by humans who have these constraints.
Then you have tonality and things like that that are very well embedded in music. or things like rhythm that the algorithm starts to pick up because it's consistent throughout a piece of music or maybe it's consistent throughout the style of music so it understands that, you know,
When you play a certain progression, then you might use a certain rhythm and you have these correlations between different features.
So all these things it can learn. For you as a musician, and I want to ask about your family too, this combination of technology, an AI system, and humanity, essentially, Does it seem like cheating to you?
Does it seem like amplification to you? And why again? did this seem a natural direction for you to take?
Right, so basically, It seemed to us that it was a great direction to take because I've observed my father work in this business. and there are different problems that can be solved in order to make the whole soundtrack composition process more efficient and the problem that we wanted to solve was generating themes extremely fast so that our clients may not run out of time or they may explore different themes before selecting the one they like.
These use cases cannot be covered by humans.
Now, when you say extremely fast, how fast is extremely fast?
Creating a theme by the algorithm can be as short as four minutes or But the idea is that we still have to curate the music that works for our clients.
And that's a process that still takes time.
But for example, for NVIDIA's music, it took 48 hours to deliver the music that they actually wanted.
Then we iterated on a few things that we wanted to improve.
But 48 hours to create a theme that they loved, I think that was a very short amount of time.
And then considering the human involvement in the creation process, I think that's actually where we add value combining. the fastness of AI systems with the sort of human element, When we talk to directors, they need to tell a story to someone and they need to make sure that someone is going to care about the story. and compose the right music for it.
Here, we have an algorithm that can compose a lot of stuff, but behind there are humans to curate what comes out of it and then matches the music to the story of the director.
And also the recording of the music in a studio allows us to get a much better sound quality because synthesizers and virtual instruments are very good nowadays, but there are still some things that you cannot emulate with them. and you'd give two violins to two different musicians playing the same note, and you might have very different sounds. they're able to take an existing idea and build on it very quickly.
So what about this question of, oh, well, that's cheating?
How do you answer that or respond to that?
Well, you could say that, but other people might say that if we do everything, we get rid of humans, which doesn't benefit anybody. which is also a fair point.
And I think at the end of the day, when you're running a business, what's important is that you add value, and in our case, we add value with the technology, but you also make it work for the use case, so the client doesn't have to adapt to how you work well so in this case it was it was composed for nvidia your your clients you say the directors is it mostly what like where does this music show Right.
So basically, we've done this deal with Nvidia.
We did a composition for the National Day celebrations in Luxembourg.
So this is a live event. We're working with an Oscar-winning director here in Luxembourg on a short animated film.
We're scoring the soundtrack. We're also working on short trailers and commercials.
We hope to work on video games very soon.
So it's really scoring music for entertainment content in general.
Well, how does it work? Walk us through the process.
Because if you think about scoring a movie, say, I'm doing the new Star Wars movie or doing a new video game, I have the kind of the luxury of seeing what's in front of me and then trying to match the music to it.
So it's going to be soaring here, it's going to be dark here, it's going to be exciting here.
How do you go about replicating that kind of rhythm and that kind of narrative in music through your process?
Basically, we have the algorithm composed themes. very quickly and then we curate them depending on the client's briefing. we may train the algorithm differently depending on the client's briefing.
For example, if the client needs music in the style of a particular composer.
We can train the algorithm specifically on this composer.
So the algorithm creates the theme, then we can have a human arranger to iterate on what the client needs.
If he needs more instruments, if he needs slightly longer theme, so it needs to loop a particular section of the music. and then we go record it in a studio.
In order to make it fit to a content, That's still human powered, you know, with the curation, but That's something we want to automate.
So we're working on a feature that will basically extract thematics and emotions from a script or a music briefing and then directly condition the output of the algorithm on those thematics and emotions.
So basically a script to music translator.
Thank you. Thank you. Okay, walk us through what was the sort of brief there and What are we listening to?
In the style, in the kind of vocabulary, what happened?
For the brief part, Nvidia told us they wanted something very cinematic, very epic, that had a clearer build-up over two minutes and a half. and strong inspiration, emotion.
And they also gave us some music they used in the previous keynotes.
So from there, we used a section of our database which is more cinematic, rather than just classical, and we train the algorithm that generate some themes.
And then once we found a satisfying theme, we brought it to NVIDIA, who confirmed that they liked it.
So in this case, you trained the music on some other music.
Is that on another NVIDIA piece that had been composed?
Like, hey, we like this. Can you do something sort of along these lines?
Actually, we did not directly train our algorithm on their music because they only had audio files.
So it only works on written files. But still, you can listen to music and say, okay, this sounds cinematic, so I'm going to pick as a training data, the entire cinematic portion of my database.
So this is not really an issue. And you guys, the AI system is a copyrighted composer.
How did you take it in that direction and why?
You know, a question we get a lot is what happens when an AI creates a song?
You know, who owns the rights? Does it get royalties?
Really, what happens? Does it move? to a tropical island and just sit back and drink pina coladas for the rest of its career.
That's the thing. Exactly. So, in order to sort of prevent the algorithm from going to a tropical island and .
We, you know, we decided to register it and register its compositions in Not All's Right Society, just like any other musician, or at least composer right so yeah the compositions are registered in the society the authors right are processed just like any other humans and And that's a security also for our clients because we say those copyrights are protected, not just anyone can claim it's his. or hers, and that's very securing when you go about doing the soundtrack of a film where the budget's allocated to music, are significant so you can't just take some music that is public domain.
Right. But it's also – it's a legal process and it's a legal recognition, but it's also – an interesting recognition of this original work produced by an AI system.
And then, I don't know, how do you view that?
I mean, that's not a small thing. That's pretty interesting.
Yeah, it is pretty interesting. And I think also the important thing is that some people challenge this. this kind of registration because it's an AI that made the song.
But at the end of the day, there are humans behind the AI that write code, and from this code, we create music.
So it's a different kind of creation process, but at the end of the day, there's still humans behind, so there's no reason why this shouldn't happen if you follow the logic that you know, humans who were giving them the sort of spark of creation by sitting at a computer reflecting on how to do this and then actually implementing such of such a software pardon the interruption Leave us a review on iTunes, Google Play Music, or whatever your favorite podcast platform of choice is.
It helps more people find us. As always, thanks for listening.
Now back to the good stuff. So what's next?
Where do you want to take this next? And it sounds like you want to take it to video games and to, you know... other larger projects.
But what's possible, do you think, if you spin this forward and think about where it goes?
So one of the things that we're really interested in is solving use cases that humans alone cannot solve. and one of them is, for example, personalized music.
So if you listen to the music in a video game, that has 100 hours of gameplay, you might notice that the music is likely to loop 30 times.
So if you think about it, no composer can compose 100 hours of music for one particular project that we just, you know, That's just not possible.
So AI can solve these kind of problems. They can allow more immersion for the player. better quality of content for the creator of the video game.
At the same time, we can take the original works by the composer and arrange them in a way that the 100 hours don't feel like you're just listening to the same thing over and over again.
In general, we think that this technology, this personalized music technology will be used by consumers, so by you and I, So you step out, you go running, it's a beautiful day, you need something uplifting to help you in your run, and the AI can actually compose something for you based on your preference of music.
Yeah, or if I'm slowing down and I need to be amped up, it kind of picks things up for me just to remind me.
Or even in a game, if I'm in danger of losing a life, maybe the music changes to subtly nudge me in the right direction or something.
Absolutely. If you think about how movies are made, sometimes the music is so clever because you don't see anything on the image. that should particularly spook you or, you know, make you think, oh, wait, there's something wrong going here.
But then you hear something in the music and you're like, oh, this is, This is weird.
Something is about to happen. And in games, because it's a lot more based on the user interaction, there's no solution right now so that you can... make sure that the music changes depending on the behaviors that the user takes.
Is it necessary that In the same way that humans compose, like you say, you're watching a film and you're scoring it, so you're looking at the scenes and developing tension and drama, etc.
I mean, do you guys think that there's a need to put image recognition into this so it can actually watch what it's composing to or for?
Once we have this script to music translation feature, then, you know, taking a script and analyzing thematics and emotions from the script is actually not too different from taking a picture and analyzing sort of what's going on there. you can translate a picture into keywords, you can translate a script into keywords, you can translate maybe a color or different sounds into keywords, and then all these keywords can be used to create the right music.
And the keywords are supposed to represent the high-level abstraction that we humans are very good at.
You know, we watch something or we hear something and we automatically associate them to a past experience, to an and that's where the inspiration sort of comes from.
So that can definitely be replicated on a longer run for music.
But the interesting thing is that we're not just limited to having the final edits of the video and then creating the music on it. because we have an artificial intelligence that doesn't really get tired or that can start anytime during the movie creation process.
We could even create the music before the image actually exists.
So when the script is finished being written, Then we can come in, create the music and, you know, have the director sort of iterate during the principal photography process. and hear the music as he films.
And that's something that could be very powerful.
Oh, yeah. It sets the mood for everyone then, even before you actually filmed it.
That's interesting. And I do think that, like you say... keywords, there's always production, you know, direction in scripts.
So it could easily be put in there. I see how that would work and work really well.
Do you ever compose music yourself and then go kind of head-to-head with your composing AI?
Absolutely. So what I love doing is improvising on the piano.
So, just taking tunes that I heard somewhere and then sort of improvising on them.
And so, you know, sometimes I just listen to what the algorithm has made and I will come up with some arrangement of it as, you know, as an improvisation.
So that's something I love doing. Oh, that's fascinating.
So you're learning from the algorithm and then kind of remixing what it's done.
Absolutely. Well, Pierre Barreau, head of Iva Technologies, thank you so much for joining us.
Thank you for having me.