Hello and welcome to the NVIDIA AI Podcast.
I'm your host, Noah Kravitz. This may be our most meta episode to date.
Let me explain. In September of 2019, startup Descript made a bunch of announcements, including one about a new tool for podcasters called Descript Podcast Studio.
Amongst its many cool features, Descript Studio features something called Overdub that lets users correct their voice recordings by simply typing.
That is, if you type words into Overdub, it'll add them to your podcast as audio that sounds like your voice speaking whatever you typed.
Here to tell us about Overdub, about Descript, and about how AI is changing how audio content gets made is Andrew Mason.
If his name rings a bell, it may be because Andrew was the founder of Groupon.
Yes, that Groupon, who all of you apparently like to capitalize the PN, as I just learned, but that's another story.
Andrew's also founded a handful of other companies, including Descript, where he now serves as CEO.
Andrew, welcome, and thanks for joining the NVIDIA AI podcast.
Yeah, my pleasure. Thanks for having me.
So tell us about Descript, about Overdub, which I alluded to and hopefully didn't get the details too wrong about.
Then from there, we can talk about AI's role in modern day audio creation.
Sure. So Descript, as many tech startups are, Descript was not a company that I really intended to start.
I was working on another startup called Detour, which was a mobile augmented reality city audio tour app.
And half of that company was building the tech.
Half of it was building content. So we had a slew of radio producers on the team making what were essentially podcasts.
And while I'd had quite a bit of experience recording audio in the past, I'd never done narrative based audio.
And as we set out to create these podcasts, we. started to appreciate how poor the toolkit was for that use case.
Most audio editing tools were designed with music in mind.
So this was about five years ago and right about when automatic transcription was reaching this inflection point where it had gotten fast and. cheap and accurate enough that it was becoming usable for things.
And we were vaguely aware of that and thought, you know, wouldn't be cool if somebody just built an audio editor that felt like a word processor, it feels like a a better abstraction than a timeline with a bunch of strange, inscrutable waveforms on it.
And so we found somebody at Berkeley who was getting his PhD in exactly this sort of thing.
We worked together to build a prototype and we showed it to some radio producers and they started telling us it was something they'd been asking. you know, their tech friends for, for the last 20 years, that's usually a pretty good sign that you're onto something.
So we ended up also talking to video producers and just starting to feel like, this is the kind of media production environment that you would create if you were starting from scratch today in a world where ai made such things possible and so we spun it out as a separate company and It took us about two years to get the foundation right.
But now, as you mentioned, We've released what's really the 1.0 of Descript.
It's a full soup to nuts podcast production series.
Studio so you can record, it'll automatically convert your recordings into text, You can then edit, meaning cut, copy, paste the text and it'll affect the underlying media.
You can drop in music, It sounds cool, but it also maybe sounds like a toy.
But we've really taken great pains to build a fully powered system media production workstation underneath.
It's also, as anything built in the year 2019, fully cloud-based and collaborative by nature.
So it also happens to be the first media editor that has a full Google Docs-style live collaborative experience.
And as part of this, there was another company that we came together with called Lyrebird, who you may have seen some of their work when they announced the company a year or so ago.
They're based up in Montreal and they're doing some of the most interesting work around generative media. and automatic kind of synthesize speech creation of your own voice.
So It ended up just being a perfect kind of partnership where we were building a really interesting product, they were building some incredibly interesting AI that could plug into our product and solve some really common painful use cases for media producers.
When you talk about – and this kind of got me because I obviously do some podcasting stuff.
I'm also an amateur musician. I've got a little – home studio.
I'm very familiar with the inscrutable waveforms.
You know, we hold on to sort of these metaphors, these UX and UI conventions that we're used to, um, When I think of audio, I think of cassette tape players from when I grew up and these play pause transport control buttons that are still prevalent in audio editors and video editors and such.
And when you talk about building a media production environment as you could in the year 2019.
And then you mentioned some people, something about this might sound like a toy, but it's not. can you speak just for a minute about you know kind of what what that gets to and how just this idea of thinking about i mean even thinking about um an audiophile as words as opposed to an inscrutable waveform.
Like how much of a big shift is that and kind of what was that process like for you?
Yeah, that's a really interesting question.
I'm reflecting on why I felt the need to point out that it's not a toy.
I think the reason is that we're used to when the applications of AI to media generation tend to be one-click events magic with a heavy editorial hand being applied.
So load in a batch of movies that you recorded on your weekend vacation to Disneyland.
And click a button and we'll create a montage and put some background music in there.
It's great, but it's taking all the craft out of it. and just doing this kind of magic.
And we think that kind of thing is great.
When I say it's not a toy, I suppose what I mean to say is, We're trying to use AI to automate the technical heavy lifting components of learning to use editors. as opposed to automating the craft.
And we leave space for the user to display and refine their craft.
We think of the word processor as the gold standard tool for creatives.
If you're a writer, you learn how to type at the beginning of your career.
And then that's it. Like the remainder of your time as a writer is spent refining and perfecting your craft, not learning new keys that come out on the keyboard.
Not the case if you're in audio or video.
In audio or video, half of your career is spent to maintaining your knowledge of the tools.
If you walk away from the tools for a couple of years, years and you come back, like even if they hadn't been just incredibly outdated at that point, you have to remember how everything works.
And also, word processors are powerful and flexible.
You can have the same editor tool that you're using to take meeting notes and it scales all the way up to a dissertation.
So what we're trying to achieve by porting or grafting the conventions of audio and video production onto a word processor is that same kind of intuitive accessibility, but also the flexibility and power.
So when a user logs into the Descript site, it's, as you described, a full-functioned creation and editing environment.
So there's a lot of stuff we probably won't dig into right now to keep the focus on the AI.
But I load in a file. Let's say it's a file of my end of this conversation, an audio file.
What happens? Descript will analyze the file and give me a transcription and I go from there?
Yeah, it'll ask you if you want to transcribe it because, you know, if it's music or something like that, you don't. necessarily want to, but yes, it'll transcribe and it'll show up as if it's a document.
And then you go from there. And then I can edit the document.
I can delete, which is a little bit easier maybe for a layperson, which I
I tend to fall more towards that category myself to understand, okay, it can find, it somehow indexes the audio file and it can find where a specific word or sequence of words is and chop that out.
But then also, and there's a great little thing on the website.
You don't have to log in or anything to try this, which I did just before talking to you.
I can also go in and I can delete a word and replace it with one or more new words.
And then this magical thing happens where it create my voice speaking those words?
That's right. Yeah, that's this overdub feature that we just built in partnership with Flyerbird.
Tell us about that. How hard was that to get it going?
How accurate is it? I noticed it was in beta.
Is it still in beta? Are people using it in the field?
And what kind of reactions you get into it?
Yes, it is in beta. You can try a demo of it that we have on the on the website now.
And we do have people who are Using the plan with this sort of thing, I think, is...
Start by working closely with a few people to really understand how they want to use it.
Make sure that everything is super dialed in. and then increase the number of people who are using it.
So we're anxious to get it out to more people.
This is something that the Lyrebird team has been working on for years and has been non-trivial.
The other really interesting thing that we're doing, so The use case that we think is very powerful is making it feel as easy to make editorial changes to your audio as it is in text.
Right now, the friction involving making those changes is so high that You have to go back into the studio.
You have to splice things in. You have to make sure it aligns properly.
The recording conditions are exactly the same so that it all sounds good.
And that's such a tedious and often technical process that at best it, puts a very high bar on the nature of the edit that you're willing to make.
At worst, it stops you from recording audio in the first place because you you've just had the experience of going back and listening to how bad the experience of hopelessness of what it would take in order to actually make you sound good.
We're just trying to make that really easy and part of the technology that these guys have built that's really unique to us is a kind of tonal or prosodic connecting of the dots where we will analyze the audio before and after whatever you're splicing in with overdub. and make sure that it sounds continuous and a natural transition for whatever it's replacing.
Right. So we're talking about the inflections in a person's voice and things like when somebody raises their voice, if they're indicating excitement or asking a question or that sort of thing?
Yeah, yeah, exactly. So I could be talking with excitement or anger, or I could be talking in a vocal fry.
And if we overdub something in the middle of that, it will be aware of that and make it sound correct.
Is there a vocal fry button? Because I think I just saw a new feature set that...
We are shipping a detection button in early November.
I don't know when this is beyond, but it'll one click removes all ums and uhs from your audio.
That's fantastic. I love it. I read something about to get a little under the hood here. that you're not using your own NLP, natural language processing tech that you built in-house, but you're using one of Google's NLP engines.
Is that right? ASR, yeah, the speech recognition engine.
What we decided to do when we got into this and took a hard look at the state of the automatic transcription world.
What we saw was that there were a number of all of the biggest companies investing heavily and making very rapid progress in automatic transcription. we saw the price dropping precipitously and you know, there are things where,
You can compete with Google, like Google had a Groupon competitor and that never scared us too much.
That's not really their wheelhouse. But I don't know that I want to compete with Google when it comes to automatic speech recognition.
It's just the kind of thing they do well, and it's core to a bunch of the offerings that they have.
So our approach was instead, Let's continuously monitor all the different services out there.
Let's find the one that's the best and offer that to our customers.
That way our customers can know they're always getting the highest quality transcription service.
And the way that we've priced the company, I mean, now Descript is $10 a month and that includes free transcription.
And because even though we're still paying quite a bit for it, our belief is that transcription basically goes to server cost or whatever.
It's a commodity and it becomes cheaper and cheaper.
We don't want people to think of our service as something where there's a metered cost to using it.
It's a creative tool and you shouldn't have to pay a tax for every minute of content you create.
So conceivably, and In my other life, I do a bunch of writing for a living and I'll sometimes work from video shoot transcripts to pull together a story.
And so I know there's a whole industry around outsourcing those.
So one could conceivably use Descript as a transcription service. along with the audio tools.
Indeed, many people do. We're speaking today with Andrew Mason.
Andrew is CEO of Descript, a company that has built a new media production environment from the ground up.
For podcasters and for other creative folks who work with audio and video regularly, and their service lets you edit audio Via text, the way that you'd edit a word processing document, you can erase things, you can even put things in now as we've been talking about since their acquisition of Lyra Bird using their technology.
Andrew, let's shift gears for a minute here and talk about your background and how you got to running an audio company now.
You're a multiple time entrepreneur. And as we talked about a little bit, maybe most well known for being CEO of Groupon.
What's your background? How'd you get into starting companies and then what's the path that took you through detour and now to Descript?
Well, I went to school First for engineering, like not computer science engineering, but another kind.
But didn't love that and transferred into music.
I worked in a recording studio for a couple of years before it became clear that I should I should move into an industry that required less talent.
So I got into tech. That's the poll quote I'm going to cut out and put the vocal fry on.
I mean, in music, you're competing with thousands of years of history.
You're competing from the people that wrote music. row row row your boat all the way forward to the present day to come up with something as interesting And in tech, it's pretty much a blank canvas.
And I think I kind of proved that. I mean, you can't have a much... dumb, simpler idea than Groupon, yet we were able to build something that got millions and millions of customers.
So yeah, I got into... tech through being a kind of contract developer for a little bit.
And then I got the opportunity to start a company the original company was called the point and it was a platform for collective action where anyone could say i will do something but only if a certain number of people do it with me and that could be taking action that could be giving money to something.
It was just this broad abstract thing. You know, we thought the concept was cool as a way to overcome some of the challenges of collective action.
But the idea that ended up taking off was using it for group buying, people coming together and saying, I will buy something from a store, but only if enough other people do it with me.
So kind of with our backs against the wall, almost running out of money, pivoted and turned the company into a focused site for group buying that was Groupon.
And that was just a rocket ship that I held on to successfully for about six years before getting shaken off.
And then I moved to San Francisco and started Detour, which I already talked about.
Detour was interesting to me for a number of reasons, but you guys were building audio kind of walking tours of cities.
That's right. Yeah. Okay. And originally, did you have a delivery mechanism in mind where these, you know, kind of podcast style things sent to people's phones or was it site specific with, you know, like when you go to a museum and they, they lend you the headphones, how did it work?
So the way it worked is we would we would find somebody who we thought was like the coolest, most interesting, authentic person to. take you through a part of the city and then it was a it was a relatively fixed path walk that would take you on a journey of but done in a way that used your location to push the story forward.
And you could basically leave your phone in your pocket and just have your headphones in.
You could sync up with other people. So you were having a group experience.
And what we built was basically a game engine where you would have these different dialogue triggers.
And then underneath that, there would be. music and ambience that would be triggered by different events that were happening as you moved along the path.
So it created this magical experience where you could, you know, as you took stepped your foot on off the curb into the street, it would tell you to. look both ways.
And it was cool. Yeah. And so that was the company was acquired by Bose.
Yeah, once we figured out that we were going to go in the direction of Descript and spin that out, we sold Detour to Bose.
As you talked about earlier, that's how Descript got started, but it was spinning it out to focus on what you're doing now.
What stage is Descript the company at as far as we talked about the overdub features being in beta? and a little bit about your $10 a month subscription model.
Where's the business at now? Where's it headed in the next 12, 18 months?
Well, just a month ago, we had what we think of as the 1.0 launch of the product.
So this is the before now it was A cool thing for transcription, a cool thing for workflow for certain types of audio producers.
But now with this release, it's full. service podcasting studio.
So this is the first time we're really going out there and making a loud appeal to all podcasters that this is really the future of creative platforms for you.
And it's been really exciting and rewarding so far.
We've seen a lot of a lot of people signing up and more and more teams using it and uh we just closed our series a So the focus will be on continuing to serve that community.
I think the podcasting community has been I mean, it was funny, like they're just used to being neglected by software.
When we worked closely with radio producers, a lot of whom have worked on some really popular NPR shows, And we would learn about their workflow, the workarounds that they've learned to tolerate are it's extraordinary like their their level of patience is so high just because they're used to stuff kind of not really working for them or being made for them.
So, well, on the one hand, we appreciate the character that's built in this group of people.
We feel like enough is enough and let's give them something that's made for them.
Beyond that, we have some very basic video editing features, and we want to expand upon those.
The hope is that we can become a creative platform for this kind of new media that's been emerging.
How do you see AI, ML, all the technologies that fall under the buzzy umbrella of artificial intelligence in 2019?
How do you see that shaping not only what Descript is doing, but the podcasting community or even sort of audio creation writ large?
Well, I can speak to how it works for us.
We see ourselves as fundamentally as a design innovation that is enabled by technology and artificial intelligence.
So When you have the experience of using Descript, it's not screaming AI.
But everything from the kind of phonetic level of accuracy that we have on the alignment between text and audio, to the generative media features that we have today and will continue to ship in the future.
We see Descript as this receptacle for a new generation of expressive interactions that could only be accomplished with the existence of of AI.
So the way that we think about it is what are the things when it comes to audio and video editing that are hard, that are slow, that are impossible.
And what would be the most intuitive word processor-esque experience for you know, akin to writing a screenplay that one could have for achieving that goal?
And is there technology that we can use to enable it?
And it's a really exciting time to be building a tool like that because so often the answer is yes.
Very cool. Are you happy with where the product and the service is at right now?
Yeah, we're feeling pretty proud over here.
The launch went really well, and it's really rewarding to hear people get into the app and just I mean, if you work in audio or video or you've had that experience, there are certain – little things and how we've simplified the workflow. that it just, when you see it, it just lights you up.
And we've been really excited to get that out in the world and feeling like, There are a lot of other people that saw that those same kind of issues as we did.
So we've got a lot of work to do, but we're feeling pretty happy with where we stand right now.
Congratulations on the progress so far in the launch.
Thank you. For folks who want to find out more, sign up, try out the product, even just do the little demo we discussed.
Which is cool. I was playing around with seeing how many words I could fit into the text field and how it would alter kind of the way the speech was delivered.
It's a cool little thing to play around with.
Where can people go online, website? I don't know if there's a blog or a technical blog for the more technical minded folks who might be listening.
Where should they go to find out more? Go to Descript.com, and there is a Descript.com slash blog.
We don't have that many technical articles yet, but we do have one that talks a little bit about... some of the research that underlies the overdub feature and the And the blending that we're doing along the edges, if you're interested for a more technical look at that stuff.
Andrew Mason, thank you for taking the time to talk with us and best of luck going forward with the script.
Yeah, thanks for having me. Thank you.