Welcome to the NVIDIA AI Podcast. I'm your host, Noah Kravitz.
Deep learning excels at image classification, but you generally need to label your images before you can use them to train a model.
A Chinese company called Melon Technologies has put a twist on things with classification of unlabeled images.
Here to explain how it works and what it means is Matt Scott, co-founder and CTO of Malung Technologies.
Matt, welcome. Hey, Noah. It's great to be here.
Thanks for joining us. So, let's start at the beginning.
What does Milan Technologies do? We focus on product recognition as an application, but the underlying technology is something related to what you just introduced. is how to handle large-scale data without human annotation.
It still has labels, but just noisy labels. and how do we handle this large-scale data, which could be partially labeled or labeled incorrectly, or unbalanced and still come up with a high performance model to solve some application space, which again we focus on is in product recognition.
So our company is really focused on doing just a top level job on recognizing any kind of product visually give it an image give it a video and if there's some product inside of there can we actually recognize it and tell you something about it tell you some attributes about that product And so that's really where we're focused on.
So the difference between labeled and unlabeled, or you mentioned noisy, there are a lot of images out there that are labeled, but they're incorrect and you've got to separate them.
The noise from the signal, so to speak. How does that work?
And maybe, you know, from that, tell us about the web vision challenge that you guys recently won.
Absolutely. So it's sort of the dark side of deep learning today.
No one likes to talk about too much about this.
The fact that you need an army of people human labelers, annotators to clean up your data set, to label your data set.
This is, you know, behind the elegance of neural networks, we have this large number of people working, clicking, and making, you know, this sort of Amazon Mechanical Turk style. situation going on.
But actually, it's something that has slowed down our industry because, you know, the reliance on human annotated data does It's just a logistical, it's a cost, and it's physically slow.
That's sort of the context of the problem space we're working in, is how can we access this large-scale data that exists out there on the web, for example, You know, we always talk about big data and big data is necessary for AI.
But the fact is that the web has this big data, but it's unlabeled or it's not...
And so if you consider supervised learning, which is what everyone does today pretty much, and it works really well, we are required to have a high quality label data set.
This is a premium. And just to give you a sense of how much work that is, is consider ImageNet, which 10 plus million images took 50,000 people over two years to label and an untold amount of money.
That's just 10 plus million images. So, And that's what you get with supervised learning.
You get this high quality data set and it works.
It works really well. That's why we're here today.
And so the web vision challenge, because you brought up ImageNet.
ImageNet is kind of not a thing any longer.
And a few competitions have sprung up in its absence, WebVision being one of them.
Exactly. So what was WebVision? How did you get involved and what happened?
Yeah, so as ImageNet retires, and it's been retired, the successor is to take... essentially the practical problem of deep learning in the forefront, which is to say, give me the same level of competition, which is in terms of classify a thousand different objects into standard, you know, is it an airplane, is it a dog?
But don't use high quality data, don't use labeled data, use data directly from the web.
So like, for example, based on Google searches, based on Flickr.
So you could have no labels, you could have wrong labels, you could have labels that are kind of skewed.
Yeah, so actually weekly supervised learning is one set of, one type of learning which does require there to be labels but those labels could be noisy they could be incorrect they could also be unbalanced you can have a skew towards one particular type of category The Web Vision Challenge held by CVPR also held in conjunction with Google Research, CMU, and the ETH Zurich.
They put together this raw data set and we participated in this and had a really high performance result.
We actually won the competition and we had competed with over a hundred different organizations, companies, academic labs, and we were actually able to get essentially to human level performance, which is 94.78% for our highest level score.
And human recognition, human level would be about 95%.
And again, the remarkable thing here is we didn't use any human annotated data and we can reach human level performance.
So you mentioned the figure, you know, 50,000 people working for two years to classify.
I think you said about 10 million. Yeah.
10 plus. Yeah. Is there a comparable number you could throw out for where you're at?
It would actually be in the hundreds of millions scale to billion level.
Yeah. A little bit of progress, huh? Yeah, yeah, we're magnitudes higher.
And getting access to that level of data, that size, actually breaks us through into new kinds of solutions that we can produce now that we couldn't produce before.
So ImageNet, actually got us to where we are in deep learning today in terms of it being that great data set.
But to get to the next level, We're going to have to break past the barrier of human annotation.
And now we're in this new area where we're not limited.
And we have been able to achieve a tremendous result on the task that we're working on.
As a startup, we're going to want to focus We're focusing on product recognition because it's a great space just to introduce to your listeners.
As you've heard of AI for face recognition, there's AI for product recognition, which is to essentially ID products just like a barcode would, but also to provide attributes for products.
Right. Now, the idea why we participate in web vision or weekly supervised deep learning is to be able to access this tremendous data to build up a general understanding of all products, not trading products one by one, but have a base understanding.
And so just to clarify for your listeners about what does weekly supervised mean in terms of this whole scope, it's the ability to look at for example, metadata on the web.
And this metadata, again, hasn't been reviewed.
And it's things that you've seen before, which there's spelling errors, there's just general mistakes, and there's skew.
And noise. And noise. It's just a bunch of noisy data.
And can we learn high quality results from that?
And that's really our task. And that's what we've been spending a number of years working on.
And here we are today to, you know, Right.
So the company was founded in 2015? 2014.
14, excuse me. And you're based in China.
Right. And you have a, I'm going to call it a flagship product, if that's accurate.
It's called Product AI. Exactly. And so that's the product recognition space you're talking about.
How does the tech, the infrastructure behind the product, how does that translate into what you're doing for customers now?
The way it works, so for product recognition, there are other solutions out there, but what it requires typically is to train a new model based on many images of products.
What we offer is to be able to just take a single photo of a product, like a thumbnail of a product, Like a bottle of wine, for instance.
Exactly. Bottle of wine. And we would use... our mastermind model to say, hey, the mastermind is able to parse that image and know what it's looking at, know where to look.
Actually, we have some attention model there as well.
It actually pays attention to the key zone of the region.
And we're able to essentially assign a single label from one image say i know enough about this object now to find it later and that's a very convenient advantage that we offer for our customers and also offer very high performance, high accuracy.
Because we have this breakthrough in the amount of data we've learned on for our base models, we can then... essentially to get the benefit of that and transfer that to our customers.
And it's essentially just like when you look at a product, you go to the shelves in the stores, you'll see a product you've never seen before, but you know it's a product.
I know it's a bottle of wine as opposed to a... a dog toy.
Exactly, and you know where to look. Like for a computer looking at a bottle of wine, if it's never had any prior, it'll look like every part of it is important.
The label is equally important to the rest of the bottle.
Right, as the bevel on top where the cork goes.
Exactly. Exactly. So the thing is, though, that's obviously bogus.
And that's where a lot of false positives, that's where performance goes down. lose efficiency, you're looking at the wrong thing.
Exactly. So with our technology, we're able to based on seeing so many different images and seeing so many different places where we can actually figure out where it's important to look for in products, we get a sense of, just like you would when you're looking at a product, to find out where is it important.
And then we simply assign the customer label to that and we can have a very high performance recognition.
And so the advantage is that you have the difference you have versus other competitors in your space and similar technologies or approaches to technology.
You're training the attention of the model, and then you're also able to access this larger data set to use for training purposes.
Yes. And we spend a lot of time figuring out how to make that sort of base really, really strong.
It's actually incredible if it's a visual demo, we could actually show it later.
Actually, we showed it to your Jensen, the CEO, when he came by our booth in GTC, we showed him a live demo of this.
Basically, you can hand it any product and we'll give you something about it, even products it's never seen before.
And that's really an exciting technology.
We think maybe perhaps a breakthrough here on product recognition, this generality of understanding all kinds of products.
So that's the work, yeah. All right. Let's talk a little bit about the services you offer to your customers and in turn they can offer to their customers.
Your product is called Product AI. How does that work?
All the technologies we talked about so far, the core technologies that we've put together, we want to expose them to... and users, developers to businesses in a very easy to use way.
So it's just an online platform called Product AI, productai.com.
And just like you can go online to a cloud service and rent infrastructure, now you can rent intelligence.
And it's cognitive services in... the field of product recognition.
So essentially, as you're a retailer, you've got products, you want to upload your products and enable search of those products.
So just upload a thumbnail of each product in many kinds of formats. and you will get back an API.
It will index your photos and give you back an API that lets you search your photos.
You can enable for your... your customers in an app or your website or even in your store to be able to find those products in your catalog.
And my team can build our app, build our experience wherever we want and just leverage your API.
Absolutely. So that's the recognition portion.
Then there's the tagging portion, which is another very important part.
It's the ability to say, not just ID a product, but give you attributes about the product.
I'll tell you general attributes, of course, the color, but for domain specific areas, furniture, fashion, particularly fashion.
We'll be able to tell you maybe the style of clothing you have or the cut, the collar, you know, the details, the attributes about.
And where's the advantage for me? as the retailer.
So oftentimes you're going to want to be able to integrate these kind of functionalities into different parts of your pipeline.
It could be in your supply chain aspects, your back end, or it could be in your front end.
For example, If I'm going to have an online site for making my products accessible to users, I want to be able to have multiple entry points to my product.
So I want to generate keywords. for my products so people can search for and then get those hits.
Or I can allow people to search with an image and find that particular item.
It could be, again, furniture or fashion or a variety of objects, wine, many types of retail objects.
So maybe there's a... keyword associated with a style and I'm using the wrong keyword and there's a better keyword to use, your system could surface.
Yeah, so for the attributes part, it would be mostly about keyword generation.
But we also see creative solutions by our customers.
One of our customers uses these APIs to do analysis on fashion data.
So they'll get like runway photos and they'll run these APIs on those photos and they'll come up with analytics to even perhaps predict fashion because if runway is showing... the future.
Yeah, there's all kinds of things. And a shout out here to NVIDIA because All of these services are running on NVIDIA GPUs, of course.
And this is making the training part, the inference part, really, really smooth.
So that's really part of the infrastructure.
And you're running cloud and then your customers are running edge-based AI-backed cameras?
Yeah, absolutely. So we have very sophisticated GPUs in the training portion of this.
The inference portion, we'll use the new P4s.
And then in the edge, we'll be using the TX2s, Jetson TX2s, and...
Those are part of the infrastructure here.
So you've been doing this in China, you're based in China, and you're in the States now, you're in Santa Clara as we speak, but you're here... visiting customers, and you're looking to expand operations into North America in the coming term.
But let's go back a little bit, kind of more to the beginning of your career and how you got into this.
You actually grew up in New York. That's right.
You were born in the States. How did you wind up in Beijing and how did you wind up in the even more unique position of being an expat co-founder of a tech startup in arguably the hottest industry in the world, AI, and you're in partnership with the Chinese government.
I could go on, but you should tell the story.
Okay, thank you. Yeah, born and raised in New York.
I have been from an early age obsessed with actually computer vision.
I... loved the idea of a computer to be able to look at an image and understand it like a person would.
And I've been on a mission since a young age.
I've participated actually in CVPR, which is the premier conference for computer vision.
I think roughly as a teenager or 20 years old, I published a paper there and I was just from an early age interested deeply into computer vision.
You're about 22 now? Is that 30, 35? I'm 35 right now.
As I was working on these interest areas, I did somehow get lucky where Microsoft did recruit me early on in my life.
And then I I went to work at Microsoft in Redmond for a few years, and I wanted to... continue on in computer vision in Beijing because in Microsoft Research Asia, they have a really fantastic lab there with some of the luminaries in computer vision.
And so I had the chance to transfer over there and And then come over to China and work on fantastic problems and meet fantastic people.
And now I've been in China over 10 years.
And after many years of working in the company, working on research in computer vision, working in many machine learning areas, actually was able to to find the moment you know where deep learning sort of hit the stage and it it partially was the epicenter was there a lot of the great network so great architectures great breakthroughs came out of Microsoft Research Asia and I was just, they're literally my friends.
And so I was there and I saw this unfold and it was obvious that this is where things are going to be going.
And so after 10 years and the timing was right, my friend who also worked in Microsoft, we worked well together.
We decided to, you know, branch out and make a company focused on just this segment.
At the time, this wasn't like the hottest thing to, when we talk to investors early on, they don't know what we're talking about.
Now it's everything. Which is a sign you're up to something good.
Right. Obviously. And so now it's everything that we're, you know, everyone's talking about it.
And so it was 2000, it was 2014. We, branched out and started to just challenge ourselves because we didn't think necessarily that this is, you know, how do we know we're doing the right thing?
And so we entered competitions. How do you?
Yeah. Actually, we found Simply Do Testing. unbiased testing enter competitions.
So China, what's interesting about it is the scale.
So China is, you some could hear, you know, big data, China scale, right?
Like it's just so tremendous. So there's these competitions, nationwide competitions with Like 10,000 competitors, like 10,000, you know, startups that you'll compete with on national level.
And computer vision competitions. They're computer vision.
They're also just generally startups, innovation.
There's many of these challenges. And so we wanted to see how well our businesses, how well our technology.
And so we put together a demo, it was just me and my co-founder.
So I'm the CTO, he's the CEO. And I put together a demo and put it in front of a bunch of judges.
And somehow the first competition we entered got second place in all of China.
And so we said, hmm, something maybe on the right track here.
We might be doing the right thing here, yeah.
So the retail applications of this, being able to recognize...
You know, barcode scanning. I go to the grocery store.
I try to find the self-checkout line because it's usually faster.
But even still, the barcode scanner, not super efficient.
Your technology, it sounds like, for that sort of a thing could be much more efficient, but I would imagine there are far bigger implications.
What's it being used for in retail? What's the opportunity?
Today, product recognition, visual product recognition is primarily used in e-commerce.
And so it's the idea to be able to do street to shop. take a photo of a product and find it in the store.
That's where we have money, we're making our business out of primarily.
But just recently, it's been this renaissance of what can we do in retail stores?
Retailers have been approaching us because we have product recognition just without barcodes.
And that's the core component to make this frictionless retail experience work well.
So I could pluck a product off the shelf.
Again, my silly example, a wine bottle or a dog toy, a sweater or whatever it is.
And theoretically your system, or perhaps in practice in China now, your system could recognize the item I'm holding.
I wouldn't even have to go through what I think of as the checkout process because it would just know.
Exactly. So can you imagine going to the store and not having to wait those up to 10 minutes to you know, at the end.
You can just come in, you can get your item and you can essentially go, you know, very rapidly without having to deal with And how do I pay?
And I would imagine the answer is different where you live than where I live.
Yeah. In China, we have really you know, efficient systems on payment like WeChat Pay, Alipay, where it's all hooked up to your messaging app on your phone.
In the United States, you would have, I would assume, a credit card set up or another online payment. where you'd first register as you're getting into the store, either through face recognition or showing your phone's QR code to set you up to get into the store.
And then you would go about your shopping.
In fact, our work here in this scenario, this is a scenario people know about.
People talk about this kind of scenario.
It's not new. But what we contribute to this scenario, our unique selling point is the product recognition aspect.
The infrastructure, the ins and outs, that could all be done in different ways.
What we're trying to do at the highest level is to... just visually recognized products in your hand, in the cart, in a basket, in motion, right off the shelf.
Like a lot of at the key trigger points.
Can we do a high performance recognition of those products in that place?
And that's what we're set out to do. And that's what we're working on and making progress.
And you're also working across some other industries, fashion, textiles.
What are you, is it a similar application or are you doing different things?
Yeah, so we call it full stack product recognition.
Basically the- Cradle to grave for my cell phone.
Right. So at the highest level, it's the products you can see. you know, retail.
It's, and types of products. So rigid products like products off the shelf or non-rigid like your clothing, textiles could also be.
And then we also go down the stack to microscopic level.
So in China, It's a really big work to do a quality check on products.
So in the factory, in the production. So like if you look inside of your clothing, the label will say 30%.
Yeah, a little sticker, 30% cotton, something like that.
That's actually at a testing center in China where they have people – hunched over microscopes, counting up fibers and classifying them, and they're sampling these fabrics infrequently.
What we're doing now is we're actually bringing AI to that problem.
And so we're actually able to with a much higher frequency, check up on the quality of the fabrics, And ultimately, the value to society would be higher quality goods.
Right. Going further down. Fewer neck problems for the poor inspectors hunched over the neck or so.
But those guys will actually still be participating.
Instead of checking directly the samples, they'll just check the check.
Then going further down, we also work on X-ray product recognition.
So we're recognizing products in the baggage scanners in security.
So just like prohibited items like a water bottle wouldn't be acceptable in certain contexts, we'd be able to pick up those products just through X-ray images.
And so this is particularly interesting in China because every subway station you go to in China, basically every subway entry point has a baggage scanner.
So there's a tremendous amount of scanning that needs to go on here.
So we work on all these levels of multi-level product recognition.
Switch gears for a quick second. Aside from the security scanners on the public transit system, two or three biggest things to an American, or somewhere else for that matter, but an American entrepreneur, AI researcher, thinking about, I wonder what it's like working on this in China.
Yeah. A couple of things off your head that are the biggest differences.
Number one, you know, coming from, you know, as an American working in China is the culture difference.
That's going to be the true culture shock. you'll get there because in China they do not speak English and so Predominantly, there are of course people who do speak English, but it's not something as common as you get elsewhere in the world.
And the culture in China is definitely quite different.
So you're going to run into places where you're going to leave your comfort zone.
And so dealing with that as an entrepreneur and working on these problems, you're gonna find that.
But after enough time, you start to get it.
And then you can start to see where your own culture and the other culture, you can sort of put them side by side and take the best of both worlds and bring yourself to a higher level of productivity and understanding.
And so that's basically my answer to your question is, You have to spend time.
It's not an overnight. You can't just drop into China and be an entrepreneur.
Sign a deal and be on your way now. You have to get an understanding of how the culture works there.
And it's something that I do recommend. It's something rewarding.
And I think the Chinese culture is really great.
Once you do spend the time to take it in, it is a great thing.
Coming from another culture, again, the diversity of thought that comes into play when you're making decisions, when you're working on problems, it's all there, and I think it's all good.
So what's next? Five years down the road for Milan Technologies, for you, for... image classification in general?
I think the key word there is general. In the work we're doing in the work in the industry is to be able to be practical.
Practicality becomes a big problem in this field.
Generally, to make anything happen, Whenever you deal with a customer, there's a lot of lead time that's necessary when you're producing some kind of AI output.
The issue would be how to make things general, how to come up with a general understanding of different types of objects, different types for the area we're working on, is to do a really good job on general product recognition, continue to do that, to lower the cost of entry into this field. would be to study more on generalization of AI. you know, making AI more general.
It's going from this vertical level to general level.
And we're going to see more and more of that.
And that's our pursuit in this field. And how long until I'm able to have a checkout-less grocery store experience up here in Northern California?
I think we're, honestly, we're on the cusp of it.
We're really close. That I like to hear.
Yeah, yeah. It's coming soon. Very cool.
Matt Scott, the company is Maylong Technologies.
Where can people find out more about you online on social media? www.malong.com.
M-A-L-O-N-G.com. Excellent. Thank you for your time.
Best of luck. Thank you so much.