Thank you. Hello and welcome to the NVIDIA AI Podcast.
I'm your host, Noah Kravitz. It used to be a favorite pastime of armchair tech experts while we watched sci-fi thrillers on TV or at the movies.
The heroes of the story would be in the lab, reviewing surveillance footage or looking at photos from a stakeout.
One of them would spot something interesting, maybe just the thing that would blow the whole case wide open, but it'd be too small or too blurry to make out clearly.
So our hero would double tap the screen or maybe just wave her hands in the air while commanding.
Zoom in, enhance, and the image would magically expand and come into focus, solving the mystery at the same time. at which points the armchair experts like me would yell at the TV screen, that's impossible.
The tech can't do that. Except now, maybe it can.
My guest today is Dr. Feng Albert Yang, founder and president of Topaz Labs.
Topaz Labs is a well-known name among serious photographers around the world. and they're pioneering the intersection of deep learning and photo noise reduction. full suite of AI powered tools for working all kinds of magic on still images and video footage.
And Albert is here to talk about how deep learning and GPUs are changing the world of photography.
And he's also going to tell us why if you're starting a company in Texas, you might do it in your guest room and not in the garage.
We'll get to that later. Albert, thanks so much for making the time to join the NVIDIA AI podcast.
Welcome. Thank you very much. That is... really nice introduction well it's uh it's it's it's something we can all relate to whether we're photographers or not that that image from the movies where you know you tap the screen and magically just see clearly what's going on.
But actually, maybe it's possible now thanks to the work that you're doing.
Tell us a little bit about what Topaz Labs is and what you do.
TopazLab is a small company focusing on making PC and Mac tools for photographer and videographer.
The company was founded 15 years ago. As Brandon nicely put it, from my spare bedroom in Texas as a hobby, Since then, the company has grown a lot.
Basically, our strength is focusing on using the most recent technology development and turning it into a usable tool for photographers.
Recently, as everybody know, there's a huge development in AI, especially in deep learning in this field.
And there are so many cool stuff in the image area.
So our company started to focus on this area.
We are the first one to come up with the productization of super resolution using deep learning. and also using deep learning for noise reduction. a blurry removal for photo that has become a sort of become so popular topaz labs start to get a lot of publicity in this area You mentioned super resolution.
What is that and how does deep learning make it possible?
Super resolution in general Means you basically obscure image or photo. a while ago in research area super resolution have a special meaning.
Basically, it's usually means you use more information from multiple images so that you can synthesize the information from multiple images and recover the detail in a larger image.
The technology was first developed early on.
I think during Cold War that was at that time is the term probably get started.
And first, in theory, it can never be used for image enhancement because you only have one picture.
If you up sample one picture from say 400%, you basically for every single pixel in the image, you need to generate 15 more pixels around it to achieve 400%. and theoretically it's just impossible.
And the AI totally changed the solution, changed the equation.
And so is there a similar effect with denoising and deblurring.
And forgive me, I'm not a photographer, so I'm sure there's a better term than deblurring.
But is it a similar kind of effect? It is.
Those type of image enhancement is generally related, so-called inverse power.
It is very easy to make a picture from a clear picture, good picture, into a bad one.
For example, high resolution become a low resolution clear picture become a noisy picture.
You basically have some generate noise. But otherwise, from a noisy picture into a clear picture, is a very hard problem.
And there is a theoretical limit on how much noise you can suppress without removing all the details.
Traditionally, we all use different handcraft algorithm to do those type of things.
And actually we were hitting the theoretical limit in those area.
And then a few years ago, blah comes to the deep learning.
And again, it changed the whole equation.
And so how did you get started? How did you find out about deep learning and start?
And I'm not asking you to divulge secrets.
But how did you start to work with it and kind of find a breakthrough to do the things you wanted to do with images?
There's a few factor play an important role.
Number one is. our company's people and culture, which is centered around ownership and take initiative.
Our company is sort of a really encourage people to experiment and we tend to hire a few PhD and they just read around the academic papers and find something cool.
So the whole way I think actually a few years ago, there's just a guy, who find a paper.
At that time, it was very popular to do image style transfer.
You give a picture of the of your house and then give a picture of, for example, Starry Night from Van Gogh.
It changed your picture like a Van Gogh one, right?
That was a very popular one. And this guy basically take this one and get it in his spare time, and do a prototype and everybody will say wow and we start to productize it right the same thing happened to the gigapixel, which is up sampling and the noise reduction. what we do is we basically read a lot of papers.
Whenever we see there are some interesting paper and a really good result, We will analyze it, combine different results, different papers, idea, and then turn them into product.
And so I'm looking at the Topaz Labs website right now, topazlabs.com.
And you've got a bunch of plugins. You've got Mask AI, Adjust AI, Denoise AI.
Sharpen, Gigapixel, and then you've also got Video Enhance AI.
And so I'm wondering when you switch from working with a still image to working with video footage.
How much more complexity does that introduce?
Are you using fundamentally the same techniques or is it a whole different ballgame working with video?
In one sense, it is a similar technology.
We all use a very deep neural network. and use pretty much a similar basic idea.
However, we are pretty new in the video area.
The product just released in a year. and we find video actually have its own big challenge.
For video, any artifact, because it's moving is very obvious.
In image, we can get away with certain amount of artifact, that visually people just don't notice.
Just don't see it, right? Yeah, in video, any small artifact, especially flickering, people immediately notice and...
So we are actually working very hard in this area.
What's your most popular of the plugins?
Our current plugin is all enhancement related. that the Gigapixel denies. and sharpen AI, the most popular one.
And is it predominantly professional photographers and people, you know, who use imagery as part of their work, or is it more hobbyists?
Do you get a mix? We get a very healthy mix.
Yeah. I would think... majority are professional and enthusiastic.
I was saying our user base is majority is a professional and photography is still the aesthetics.
So you mentioned that the company was founded, I think you said 15 years ago, but it was much more recently, obviously, that deep learning came into play.
Can you talk a little bit about the role of GPUs in what you do?
And a lot of the guests on the show will talk about how in the past, five years or and then in particular, even the past couple of years, The technology, the rate of acceleration and the advancements of deep learning and using GPU power to make AI faster and more accurate has really just exploded.
Have you found that kind of a similar effect or, you know, what have those past couple of years in particular been like for Topaz?
This is my favorite topic. I never stop to be amazed by how much processing power, just a game-grade GPU can provide, not alone the machine learning a specific GPU.
I actually start a little bit of deep learning related project.
When I was doing my PhD study, at the University of Waterloo.
At that time, my thesis is computer vision. to use computer to detect the object so that our robotic arm can pick it up.
So not only recognize the object, but also to locate the project.
And at that time, Basically, we have to manually craft in detail how to extract object, corner, feature, and then we use some type of scene description language.
I use the so-called attribute graph to describe a scene. so that we can do the computer vision task.
At that time, I feel, man, this is just so difficult.
And I tried neural network at that time.
I tried using about two layer neural network, try to just do some basic feature detection, trained a few months on mainframe computer and it's never really worked.
I fast forward now about, I think about 20 years.
And I was just amazed the our training we train our neural network usually have a few million parameter Usually, at least a few thousand layer, up to 100 layer.
Yeah. and we basically using a couple of NVIDIA card. to do the training, do a really refined training in a month at most.
So it's completely... not only about the speed, it is a fundamental changes.
We're speaking with Albert Yang. Dr. Yang is founder and president of Topaz Labs, which is a well-known, really popular suite of plugins for editing enhancing sharpening uh photographs and also now uh video footage as well.
You mentioned, Albert, your work, your PhD work at University of Waterloo.
Let's back up a little bit if you would and talk a little bit more about coming out of Waterloo, what you were doing and then what led you to start Topaz and also just how you got interested in photography and digital imagery.
My research interest is always in computer vision image and audio signal enhancement. and I basically during my PhD study at the University of Waterloo, The library I was in called the PAIMI, Pattern Analysis and Machine Intelligence.
So I was doing machine vision. Afterward, after my study, I joined a couple of small companies.
Most of them are defense contractors. image, radar and the sonar signal processing and then. and I draw note how doing speech echo cancellation, noise suppression, those type of also signal enhancement.
And then, as you know, the Silicon Valley appeal and I go to San Jose and joined a funding team for Techavail. which is doing digital video enhancement semiconductor AC.
Okay. And after that, I founded another small company called Forty Media. which is doing small microphone array IC.
Basically, we use multiple microphone to pick up a clear signal. right now those things are almost everywhere right yeah echo we we our company actually probably one of the very few to start this direction.
After one of the company go APO, I come back to to Dallas due to family reason.
Right. And basically founded the Topaz Lab because I am a hobby photographer and some of the tool uh i used it just another very satisfactory so i started doing plug-in myself and later on just put the internet and they start getting certain traction.
That's how Topaz Lab got founded. Very cool.
It's always great when you're trying to solve a problem for yourself and then it turns into something that other people can use as well.
I was wondering, as you were speaking, given your background in studying computer vision, obviously, but working with... audio signal processing and noise reduction in audio and speech and everything.
How similar or different Is it working with images versus sound?
They all have certain their own characteristics.
However, overall, they all can use similar methodology.
I consider them are very similar, actually.
One of the things Deep learning is the reason deep learning is so successful in image and audio signal processing area. is we have a vast amount of available data set.
Basically, if I go on YouTube and then start downloading, I have unlimited amount.
This is not true for many other applications, for example, financial area or some other area.
That's why the result in image or audio area in terms of deep learning are really impressive.
And so how big is the team at Topaz right now?
Right now we have about 20 people. Okay.
And is most of that focused on, is it kind of a mix of research and then productization or is it mostly product focused?
Our company is very product focused group.
Majority of people a software developer and machine learning researchers?
I wouldn't even call myself a hobbyist photographer, just judging by the quality of my photos.
It really depends for me on you know, the advances that are happening in the cameras themselves, which these days for many people like me are my phone.
And I do know that there's a lot of processing that happens when the picture is being taken and in the software on the device itself.
And so I'm wondering, given your history of work and the work that Topaz is doing right now, where do you see the industry headed? and feel free to answer specific to Topaz or kind of more broadly, when it comes to kind of, you know, what's happening on the devices and multiple lenses and sensors combining to make a single image and so much happens in the software.
But then this whole world that you're in with, I don't know if post-processing is the right term, but everything that you're doing, you know, after the fact to edit images and Where in general is deep learning and AI kind of pushing the whole process of photography as we know it?
From what I can see, Deep learning is penetrating uh every area uh in the under device side you can you can already see a lot of progress that google put in the poultry mode which is using deep learning to do blurred background.
Those things I see are going to continue to happen. and especially there's a new field called computational photography.
Basically, they can even use the light field to synthesize the imaging process to achieve a pretty amazing result.
So those things continue to happen. However, the device. has always much limited the computation power compared with the post processing environment, PC or Mac or even on the server side.
So I can also see there's a huge room on the post-processing, especially on the desktop side.
This is another interesting aspect is I find until recently, desktop deep learning based app is very rare because it is actually very hard to develop.
There's a not good infrastructure to develop on PC or Mac.
They have a few library, none of them are very production ready. and there are a lot of compatibility issue.
So people, usually first develop their deep learning solution using cloud solution right but we find most customers really hate it we actually have our video enhancement solution first on the cloud And nobody use it.
Very few people use it so I can see in the post processing side on the PC. uh and the mic device i see a lot of opportunity of course on the cloud side that is also tremendous opportunity in those area Do you ever shoot on old-fashioned film when you're doing your own photography?
Not anymore. Not anymore. I kind of figured that was going to be the answer, but I don't know.
I get nostalgic for old tech myself, so I... had to ask.
Yeah, I think the again the opportunity is it's a very interesting to see. if you compare with the current phone camera as a let's say low lighting performance compare with even just five years ago, a bigger camera, very expensive camera. the phone cameras low lighting performance is very comparable. with the old big camera.
This is actually mainly because it's the edge level integration. and maybe not deep learning but a certain the computation power they can do a lot of processing and those are continue to happen so the the device side is getting better and better.
However, on the post-processing side, the computing power will always be a few factor And with that more power, we can do much, much better. better job that's why we we feel the post processing area still have a long time to go You've been working in computer vision for a while now and signal processing.
And if I could ask you to kind of look back. are there any moments or any particular you know, challenges or breakthroughs that stand out to you as particularly surprising, and whether they have to do specifically with image processing or not,
But related to computer vision, deep learning, that whole thing, anything that stands out at you is something that really you weren't expecting but turned out to be a big moment.
I think when again I go back to the computing power improvement largely due to GPU.
This is basically a cold game changer. Without it, it is pretty much impossible, especially for a small company.
Unlike a big company, they can have a server farm with thousands of server connected and do some decent deep learning project.
A small company like us, there's no way we can train a reasonable image enhancement neural network.
For example, if now I just use a good computer without using GPU, Training one network will probably take a couple of years if you secure that amount.
So that is actually a game changer. Yeah, absolutely.
And also, the majority of our users' desktops have some type of GPU, a little higher end. that is also a big enabler and so now instead of asking you to look back i'll ask you to look ahead a little bit because this is what we do to our guests on the show Is there anything that you're kind of foreseeing, maybe something you're working on or maybe something you're anticipating down the line that you think is going to shape? let's say the next two, three, even five years of digital photography and video.
In a near term, from what I can see is basically many of the really interesting development in research area. will be continued will be productized continually our topaz lab will be one of the company really focusing on introducing some new development to our customer. and also even for existing product category like image enlargement, noise reduction, and deep blur masking.
Those have a deal. There are so many potential and the new development that we would like to incorporate it into our product to make it even better.
For example, right now, The intelligent is still not very smart.
People have to select a different model for different type of image problem and so on.
So we are trying to make it smarter. for example, one bigger neural network that can adapt it, discover what type of problem your image have and hopefully automatically corrected to the best.
That's what I need. I hit that magic wand button and just let it do its thing to make my kids look better in the pictures.
Yeah, it is coming. The technology is developing in this direction and there's some new method working in this area.
Well, I for one look forward to it and can't wait to see what's coming next.
I know your website topazlabs.com. has places, their free trials of a lot of the products and demonstrations and such.
But I also hear you've got a pretty good blog.
Uh-huh. I personally, yeah, I did occasionally write a couple of articles in our a block site to describe the background information about our development and and just try to demystify the AI technology and put it into a little bit historical perspective.
I really enjoy writing those articles. Well, that is at topazlabs.com slash blog.
Easy enough to find right there on the site.
Anywhere else you would point people to go if they want to learn more, or is it all on the Topaz Labs site?
Majority of the information is on Topaz's website.
And also if you Google Topaz AI, And you'll find a lot of information.
Well, Dr. Yang, we appreciate the time. And, uh, there's a lot of stuff to play with on the website and, and, um, And I certainly, as I've said, I can use it.
And clearly there are lots of others out there, both hobbyists and professionals who do as well.
Thanks for all you're doing to advance the field.
Thank you very much. Really appreciate it.
Really enjoy your interview. It's my pleasure.
¶¶ Thank you. Thank you.