Welcome to NVIDIA's AI podcast, and we are recording this segment from the floor of the 2017 GPU Technology Conference.
It's a gathering of the AI faithful, or should I say obsessed.
Among them is our guest, Xu Chenyao, who is a co-founder of Kit.ai.
And it's a startup that is using AI to build better voice experiences. phones, in our speakers, really anything that has a microphone they are working on or working in. including those increasingly human like systems that we find ourselves talking to more and more.
As we ask, for example, Alexa, change the music or scream Google.
Hey, Google, what movie should I watch tonight?
So we're going to talk. all about those natural language processing kind of problems with our guest, Shuchen Yao.
Shuchen, welcome. Thank you very much, Michael.
What is it that Kit.ai does? Because I know you have something called Snowboy and it triggers a hot word or a weak word or something like that.
Sure. We started the company about two and a half years ago back in Seattle.
A bunch of PhD students from Carnegie Mellon University and Johns Hopkins University all came to Seattle for this startup.
We did quite a bit of natural language processing and speech recognition using deep learning when we were at school.
And then we moved to Seattle, joined the incubator program of the Allen Institute for Artificial Intelligence.
And we published two pretty amazing products One is about hardware detection, and the other one is more about conversational dialogues.
So the specific focus of this startup is all about voice and a text-based user experience using deep learning, using speech recognition, and using natural language processing and conversations.
You did your PhD at John Hopkins and I know I thought everybody who went to Carnegie Mellon come out and they do driverless cars, but apparently John Hopkins and Carnegie Mellon together does natural language processing.
What was your focus then in your PhD and how did that translate to what you guys are doing now? a question and answer.
During my time, that was between 2010 and 2014, IBM Watson was such a big hit, right?
They just won the final Jeopardy game against the human champions.
So everyone was like crazy about it. producing super intelligent machines that can answer super, super tricky and twisted questions.
So I did my whole PhD thesis on that. And that was my first startup idea.
I thought, oh, I would just turn my PhD dissertation into a startup idea. turned out to be a disaster.
I mean, clearly you got your PhD, so it wasn't a disaster in that regard.
Yeah, exactly, exactly. So I worked on that for three months and showed it to a few investors and they said, oh, okay.
This is cool, but you will be having a hard time, one, finding customers, two, competing against other bigger companies.
And what you're going to do is really, really find the pinpoint on the market and then focus on that, right?
Especially those pinpoints that bigger companies do not currently focus on.
And that's why we first had the idea of Snowboy, the weakware detection engine.
Somebody calls it the hardware detection engine.
So a weak word or a hot word. So let me make sure that I have this right.
Amazon's Alexa. They call it a weak word.
Google, when I say, hey, Google, that's a hot word. and i'm not sure why they're on either end of the spectrum but anyway but what do those words do and then what does that trigger and then how does it like if you think about it as a system How does it work?
Sure. Yeah. So a weak word, a hot word controls a very, very front-end of a human computer interaction.
So it's actually the very, very start journey of this whole experience while you are asking the personal assistant to do something for you.
So the scenario is like, you know, You have a Google Home device or Amazon Echo device sitting at home on the table.
The device is actually listening to 24-7, all the time.
The audio goes in, the audio goes out, and the device is almost memoryless.
It filters out everything. until it hears the weak word, like Alexia or OK Google.
At that point, when it hears Alexia or OK Google, it's still offline.
Once waking up and when you ask, hey, changing music, what's the weather today?
Is it going to fire up the Wi-Fi and then go to the cloud? and use the speech recognition and natural language understanding from the cloud process your query right right so hey alexa or not hey alexa alexa it's hey google Correct.
And then something else, it happens. It wakes up and it goes to work, essentially.
Correct. Yeah. So the specific technical challenge here is like, you've got to do this week-round. detection on the device, right?
You cannot use any internet. So for most devices we have, for instance, like for the Amazon Echo device or Google Home device, sort of okay CPUs there.
So you should do like your total deep learning on that CPU.
And that also runs 10 different other things while you still cannot use internet at all.
So you have a very, very limited resource. to do this weak word detection or hard word detection.
This can be applied to a bunch of other home appliance devices For instance, the Amazon Echo or Google Home devices sells for $140, $180.
But we've got so many Bluetooth-based speakers that sells only for like $30 or $40.
The bond price goes down to like $15. How can you do deep learning powered weak word detection on those like super, super cheap computer chips.
That's like a great engineering challenge.
I see. So yeah, because I assume that when you're talking to these devices, whether it's from Google or Amazon, There's lots of compute available to them all the time, but what you're saying is maybe there's some there, but as more and more things have microphones, let's put it that way, plus CPU, plus compute.
If the options open up, you just need to be able to pull it off.
Yeah, exactly. the very, very front-end triggering part.
When you say, hey, computer, hello, Samsung, OK, Google, That part has to be processed offline just by definition, only because we want to protect our privacy, right?
Only when I give explicit command, to the device, they would only wake up, right?
So up until that point, everything has to be processed offline.
So the tricky part is like, How do you do deep learning on those, like, very limited resources?
Like, sometimes you only have, like, less than one megabyte of RAM and only, like, you know, 140... a megahertz of CPU that you could use, and you can still want to detect it very accurately without giving too much false alarms.
Snowboy is meant for developers, but this gets back to why you went down this route to begin with. what's the pain point you're solving?
And it sounds to me like I have a sense of it already that like, look, you have to do, Very complicated and sort of subtle things with very little compute, and it has to be offline.
Yeah, exactly. Yeah. So before Snowboy, there was no open community solution on the market.
And developers have always used two methods to do it.
One is that they use a clap detector. You just clap, right?
Clap on, clap off. Yeah, just to detect that pitch, right?
So that's not very good. It does not have semantics. and it does not give a personality to a device.
Reminds me of the 80s a little bit with the clapper, but still, that's not good enough, I guess.
Yeah. Sometimes people just, you know, plugging a button.
Amazon did this Amazon Echo experience on a Raspberry Pi.
In their first version, they did not have a record engine. that they could share with the developers.
So what we ask is, hey, we're just going to hook up a physical button. on the Raspberry Pi.
So when you say, what's the weather today?
You're going to first press the button as if you have triggered the device.
Then you say, what's the weather today? So that just renders the whole point of hands-free voice experience useless because you still have to go and reach out to the button.
So people have done so many things to get around this.
But they had no idea how to do that until we released this Snowboy toolkit about a year ago.
And we solved exactly that problem. And it's a problem that Google and Amazon solved in their own way.
But again, they kept up for themselves. They're not sharing it with... with others yeah exactly yeah but what is the pain point then let's get back to that like why why do i you know if i'm a manufacturer if i'm a developer why do i want to have this capability well first you know that it controls a very entry point of a the start of human-computer interaction.
So without being able to trigger the device reliably, you cannot just have a good experience, right?
You can say, OK, refrigerator, like 10 times, maybe five times, it only fires up.
And then it's not very accurate. If you use a clap detector, and then you cannot differentiate this device with a microphone with that device with a microphone, right?
So what we do is we call it a customizable weak word.
We have the biggest library of the weak word detection in the whole world.
And up to today, we have more than 7,000 different weak words in 15 major languages.
So give us some weak words that are... The number one...
Yeah, the number one most popular weak word is a smart mirror.
So it's actually built by a developer from Microsoft.
So they built this shiny mirror thing they put a mirror in front of in a bathroom or something and put a monitor behind the mirror so in the mornings they can say hey smart mirror Show me the map.
Show me the traffic today. What's the weather today?
So while you are like, you know, addressing stuff in front of mirror, you can still observe information from there.
Show me my calendar, huh? Yeah, yeah. Ah, read back my email or whatever that came overnight.
Yeah, exactly. Oh, that's cool. And again, the weak word is?
Smart mirror. Smart mirror. Smart mirror, yeah.
What are some other applications that you guys are working on?
For applications, we got a lot of We got a lot of people using super, super innovative ways of applying this weak word, for instance, People can use it in command and control.
So that's why Google called it Hotword instead of WeakWord, because WeakWord really narrows the functionality of Snowboy.
So you do not necessarily have to use the weak word to wake up the device, but also you can give some commands.
For instance, some people have built hot word series for their robots.
So you could just say, hey robot, go forward, go back, stop, turn left, turn right. so with a set of like five or six different commands you can totally do it offline and it's super super responsive and a robot can just you know listen to your command and then just to go.
So that's another big stream of how people do it in a command and a control way.
So in that case, You are not listening to one weak word, but you are listening to like seven or five different hard words. in this case, as commands.
Right, and again, it's happening offline, which is the hard part.
Yeah, exactly. The hard part in terms of the processing.
Yeah. We'll see you next time. It helps more people find us as always.
Thanks for listening. Now back to the good stuff.
Spin this forward for us. I mean, how does this get to be more and more part of our life?
You mentioned like, hey, robot, turn left, turn right, turn left.
What I want to say is, hey, robot, take me to Shuchin's house.
I don't want to say, like, give them precise directions.
So what's the future and how do we get there?
That's a very, very interesting question.
So there are so many different ways to do that.
For instance, We have the Google way, we have the Alexa voice service way, right?
So basically Google and Amazon opens it up for free. so people can just hook up.
Once the device is triggered, people can hook up with whatever whose service layer on the backend.
Right, so that's why we see all these integrations.
At CES, there were dozens with Alexa, for example.
My dishwasher... you know, I'm like, hey, where's the dishes?
Or, you know, are they clean or whatever it is?
Yeah, yeah, exactly. And we, you know, today is kind of an interesting time because, you know, Microsoft, The Builder Conference is coming pretty quickly.
So people have rumored that Microsoft is working on a Hey Cortana device, right?
And Apple is working on a Siri-powered Echo-like device.
So what we want to do in the future is like, We have actually companies, builders, manufacturers coming to us.
They said, hey, can we use Snowboy to control whatever the customers want to trigger on the back end.
So if they say, OK Google, they can trigger Google's service.
If they say, hey, Cortana, they can trigger Cortana services.
Because each different person uses different services on the back end.
So for instance, Alexia is very good at understanding music.
Google is very good with my personal stuff because Google knows my calendar, knows my flights, knows my hotel tonight in San Jose, right?
So I can ask a lot of personalized questions to Google, while Microsoft is more about office productivity.
For people who live in the whole echo ecosystem, they can ask Siri about stuff.
So in the future, I would not be surprised to see a device, a four-in-one device that has all the bigger companies back end waiting to serve you and people can choose. can choose whoever's services to use just by selecting this trigger word.
Don't I want... And I want to hear your thoughts on what the experience is like.
But I would rather not say... Hey Google, hey Alexa, hey Cortana, hey Siri.
I just want to say, hey, what's my calendar?
And have the right backend respond to the right request, right?
I don't wanna pick you know, from among these companies, honestly.
Yeah, exactly, exactly. So this is something we're working on as well.
We have another product called ChatFlow, which is a basically this aggregator for all kinds of different web APIs, and we do a lot of natural language understanding on that.
The idea of ChatFlow is to enable developers with the capability to build any kinds of multi-term conversations with quite some capability of natural language understanding so that in the future, you do not have to explicitly tell Echo, hey, ask this skill to do that.
Right, right. So you could just... freely speak up your mind and this dialogue manager built inside a chat flow can know instantly what kind of skill, what kind of chatbot you want to ask, and then call that chatbot or call that deep learning leader for you so that they can get the actions done on your behalf. deep learning systems and natural language processing, I always have this image in my head of kind of what it's being trained to do or kind of And I anthropomorphize things.
What are these things? What are you training them to do and what are they good at?
Are they the world's best thing standing there waiting and then waking up and then running off and doing something for us?
How do you picture it and how do you describe it?
So if you look at how people, how the technology of natural language understanding has evolved, in the last three to four years, you can clearly see a path that people started with some shallow methods.
For instance, with some linear machine learning models, for listeners who sort of understand machine learning a little bit.
So they can do super simple commands. We call it a one-shot understanding.
Just like, you know, what's the temperature in San Francisco?
Book me a flight ticket from Seattle to San Jose, right?
So we cut a one shot because most of those actions can be accomplished. just using one query and you are done.
That was something we were trying to solve like three to four years ago. when companies like WIT.AI or API.AI open up their services for that, right?
But nowadays, people have definitely asked for much, much more and deeper for that.
I'll give you two examples. One is first, we want to do conversation, right?
We want to be out of this deep learning model that has at least some short memory about what you just said.
We can accomplish a kind of five turn or six turn conversation between the bot and the user.
That's one. Two, we want the bot to be sort of like having some sort of domain or world knowledge, such as if the bot has heard of something that a bot has never heard of, which by literal translation in technology, which is like something that does not exist in the boss training set, the body can still sort of generalize and understand what the user meant if the bot has never seen that.
So from this part, we need quite a bit of unsupervised learning, probably just using some ways of deep learning to give this knowledge to the bot so that the bot you have built is not super fragile.
Right, it can only kind of answer or have a conversation about three things, and if you ask it or say anything else, it's sort of lost.
Back to this, like if I have a five sentence exchange with a bot, that's a short conversation.
But is it always with the sense of some transaction at the end?
I mean, I can imagine it, for example, in a...
And a medical setting where I ask you a series of questions about like something's wrong, you know, and you know, you have five kind of. shots to explain to me what's not feeling quite right but what are the scenarios where that five sentence or so exchange and i can imagine then it goes to 10 sentences and then you know, kind of endless.
But what are those scenarios? We have quite a few real-world scenarios, especially for people who are doing surveys.
They've got a long list of questions. Sometimes you just never know when it's going to end. form of doing surveys has bothered people too much sometimes people do not think that's a good way of communication and people don't want to spend time on that and then people Customers have built bots to ask a question, to personalize the questions.
For instance, some insurance companies If they want to get you like a car insurance or health insurance, they got a bunch of history about yourself, about your family to ask you about, right?
So they were thinking of some very, very personalized way to keep you engaged during this whole conversation and slowly, slowly ask you questions such as, you know, hey, Michael, I understand you are getting insurance for your family.
How many children do you have? You have two, and I can ask you a little bit more about your children, right?
So some kind of information I do not necessarily have to use in my survey, but at least I keep you engaged. during this whole process.
So you would feel it's more a pleasure to complete this process.
So this is something I was seeing companies who have attempted to do.
So it's engagement, and again, it puts the natural in natural language processing, huh?
Yeah. How did you get into this? Why this part of the field and what brought you to it and what's kept you here?
I've been doing this for almost all my adult life, I guess.
I did my bachelor's in electronic engineering, specifically in acoustics.
And then I did my master's in Europe, all in computational linguistics.
That's what we call it. Or sometimes people call it natural language processing.
And then I thought, oh, I'm not done yet.
I came to the States and then did my PhD.
When we were trying to do this conversational engine, the chat flow, that was back in 2014.
At that time, chatbot was not a thing. The only thing we got on the market was Siri at that time.
Everyone loved Siri, but Siri is only limited in understanding, like, what do we call, one-shot understanding.
We thought, oh, how about we give the power...
We're doing multi-term conversations to developers so everyone can be just can be building some bot as good and even stronger than Siri.
So it took us, like, a year and a half to develop this super awesome software called Chatflow.
And by the time we released Chatflow, it was 2016.
And all of a sudden, everyone was talking about a chatbot.
Yeah, chatbots are everywhere. Culturally, this is interesting to me because I think part of it was that You know, Siri and earlier versions of these other things, whether it was, you know, Alexa or, you know, they work, but I'm not going to say it five times over.
I'm going to give up. But like. So they work better now, right?
But culturally, how have you noticed a shift in how we're more willing or happier to engage with these kinds of bots?
That's a very interesting question. If you have observed how small children behave, Currently, we got sometimes kind of astonishing.
If you see like a three or four year old, you know, in the airport or something, sometimes they see a TV and They want to go ahead and swipe the TV.
Right. Just because in their gene... They know a lot of things are touchable or swipable.
And we've got so many conversations between the children and the echo of the theory. you would say okay that generation they grew up with a much more natural interface.
Even though today, to tell the truth, the technology for the chatbot is not there yet. but we are hoping that as the next generation grows up and as we are striving to make them much, much better, in five or 10 years, the technology can really, really be there and we can give back the original, most efficient human computer interactions back to the next generation which is voice you think which is which is just speaking yeah exactly So in five to ten years, we walk in the door of our house or our office, and how does it go down?
Well, I would imagine everything has a microphone.
So you might freak out, but thanks to Snowboy technology, everything is processed offline.
And just to be clear, when you say you might freak out, privacy, obviously, if there's a microphone everywhere, you think like, oh, gosh.
Everyone's listening all the time. Yeah, exactly.
But you're saying that that's at least not what you're designing these things for.
Yeah, exactly. We've even heard use cases, kind of scary use cases, like a telecommunication carrier tapping people's cell phones.
So suppose Michael, you and me, we are having this phone call, right?
And then for instance, you know, just example, like AT&T, something is providing a service called, uh, hello at&t and i pay like 10 bucks to have this hello at&t as my personal assistant listening about every phone conversation we have.
Suppose we have an argument, and it's like, oh, where's this GTC conference?
I said, OK, maybe in Palo Alto. I don't know.
I could say, hey, hello, AT&T. You know, check for us, where's this GT's conference?
It's in San Jose. I can answer that question for you right now.
And then the bot would just barge into the conversation and provide information in real time.
So if this is already happening, For people's phone calls, I imagine in five to 10 years, we are going to have the personal assistant everywhere we go.
For anything we want to ask about, we just voiceover query, and someone will be there answering your questions.
And it might not even be a direct question, it sounds like.
It might just be sort of like, ah, what do I have to do next?
And then someone's like, well, you have an appointment with Xu Zhengyou, you know.
Yeah, that's going to be a super awesome future that we're building towards that.
Shuzhen Yao, thank you so much for joining the podcast.
And I will think of a weak word or a hot word.
You can use a snowball. and one more week to our library all right thank you so much for joining the podcast thank you so much for having me