Hello, and welcome to the NVIDIA AI Podcast.
I'm your host, Noah Kravitz. The search for intelligent life beyond the confines of Earth has long been one of humanity's greatest quests.
Researchers hunting for extraterrestrial life now have a new tool at their disposal, artificial intelligence.
Our guest today is one of the authors of a study published in Nature Astronomy detailing the use of AI to recognize signals that natural astrophysical processes couldn't produce.
Modern radio telescopes are capable of capturing so much data that AI is proving a vital tool in sorting through it all.
Peter Ma is an undergraduate student at the University of Toronto, and the lead author of the paper, which was co-authored by a group of experts for multiple universities and Breakthrough Listen, an international group searching for signs of alien civilizations.
Peter is here to tell us about the team's efforts, so let's get right to it.
Peter, welcome, and thanks for joining the NVIDIA AI podcast.
Thanks for having me on. So why don't I open it up to you?
Tell us about the research project that led to the paper.
Right, so this research project initially started back a couple years ago, back when I was still in high school, thinking about this problem.
I was in a computer science class. It was going quite slow.
So I just decided to teach myself the rest of the course content.
And then I found myself with a lot of free time now sitting in a period of class where I was I've already like basically done all the work already.
So I continued to look for various kinds of problems I was interested in.
And so this was now in about 12th grade.
And, you know, I was interested in astronomy since I was a kid.
So I was looking for problems in this field, particularly ones that involve big datum and open source datum.
So I can leverage some elements of machine learning and deep learning and all that kind of cool stuff in merging tech to solve some of these interesting problems.
And so by around the time that COVID started to hit everyone, I had a lot of free time, just like a lot of other people.
And I decided to kind of pursue it quite a bit more beyond just the scope of a semester project.
So I You know, I code emailed basically half the people at this lab about, Hey, this is cool.
I want to use our data to do this kind of a thing, this kind of a project.
And they're like, okay, that sounds good.
And they continue to mentor me on along with that.
And so... By the time I actually got a chance to fully flesh this out, it was about the summer of first year.
And I came up with this idea of merging both sides of supervised learning and unsupervised learning together to search for anomalies in our data, specifically looking for something called technosignatures. and these technosignatures are basically proxies for signs of intelligent life out there in the universe.
So in other words, we're building, trying to use things like artificial intelligence and stuff like that to look for extraterrestrial intelligence, which I think is quite cool.
Yeah. Yeah, there's a lot in there and, you know, makes me feel a little bit sheepish about what I was doing when I was bored in high school science class way back in the day, but that's fine.
So we can talk a little more about kind of about your background and what you did and didn't know about. you know, computer science and all of that a little later.
But to kind of hone in on the research, So you knew you were interested in problems having to do with big data, that kind of thing. and you had previous interest in astronomy.
And so you decided to set your sights on looking for signs of life out there in the universe.
Did you get far enough on your own to kind of identify a problem to focus in on?
Or did that come up after you started cold emailing different people?
So the big problem that I eventually settled on was really kind of iterative process.
So in the beginning, obviously, I didn't know enough about the field to make you know, substantial impact just yet.
And so it was really the help of all my supervisors who are also the collaborators on this paper who helped guided me towards a specific problem that they had, which I can then solve and approach with techniques of machine learning and deep learning and all that.
So initially, I had an idea of the direction I want to go, which is to apply some kind of novel technique on this huge open source piece of data.
But how to actually provide value was, you know, I had to have you know, all my supervisors kind of hone me in on the problem to approach.
Are you able to describe for listeners who this is all new to them, kind of, what the process is like, how researchers are going about searching for signs of life. beyond the Earth?
Right. So it's pretty simple. It's pretty easy to understand.
So What we're trying to look for are signs of, well, life or intelligent life.
And obviously we can't be sending aliens an IQ test to send us back to measure their intelligence.
So what we do is we actually look for signs of engineering.
Right. So if we find that there is some kind of suspicious, like a phone or a microwave floating out into space, then it means that someone must have built that thing, or an intelligent thing must have built that object.
And so, in other words, we're looking for signs of technology, which are proxies for signs of intelligence.
Now you're going to be wondering, okay, so how do you find signs of technology light years and light years away from us?
So it turns out a lot of our pieces of tech emit electromagnetic radiation.
So think of radio, Bluetooth, all that kind of stuff. and they emit these electromagnetic radiations.
But then you're going to be wondering, okay, so how do you differentiate between those electromagnetic radiation or something coming from, say, the sun or a black hole or whatever the case is.
It turns out that that is true. Everything active in our universe kind of emits these radiation, but there's a clear signature that comes from a phone or engineered device.
So for the sun, it is emitting electromagnetic radiation, but it's something we call broadband.
In a sense that you can think of it as the sun is yelling at all kinds of frequencies, emitting an energy out to us.
But for our phones, the whole purpose is to communicate.
It's to send information. So the phone is not going to be yelling at all these frequencies.
It just needs to yell at a specific frequency or be transmitting at a specific, you know, like point for us to listen into, right?
So in other words, your phone is not emitting gamma ray radiation or anything like crazy like that.
It's just emitting one small sliver of the frequency band.
So in other words, in the data, when we look at it, It is this very thin line that is what we call a narrow band signal.
And that is in contrast to a broadband signal, which comes from, say, a pulsar or a black hole or the sun or whatever the case is.
To give you some idea as to the scale, the signal coming from your phone is about almost million times thinner of a line in our data than something coming from a normal astrophysical event.
So something coming from space. It's very, very obvious for us if something is a sign of technology or something it is natural.
And so I'm going to hold you up for a second.
And at the risk of being that annoying guy, I'm going to be that annoying guy.
Just ask. How confident are and when I say you, you know, the whole the researchers, the experts you worked with and everything.
How confident are you that you're looking for the right things as opposed to, say, some advanced form of life out there the evidence of engineering actually isn't a radio frequency, but it's something completely different.
Right. So it turns out, at least from our understanding of physics, radio waves and all these kinds of things are some of the most efficient ways of transmitting information fast at the speed of light and wirelessly right and so we think that every kind of civilization that is approximate in our intelligence, if not slightly bit more, would have mastered this piece of technology.
Okay. And so there's another argument that we can make, right?
So it is something called a watering hole.
So it turns out that in the universe, there is a very particular part of the band, part of the spectrum that is super important to all kinds of you know, physics, cosmology, and astronomy that any kind of intelligent-fearing species or civilization should be looking at. to study the universe and the stars and to overall be intelligent and have science and stuff like that.
And that happens to be the 21 centimeter line.
So you don't need to know specifically what that is.
But there's this particular piece of the band that is super, super important to a lot of physics and a lot of science and a lot of cosmology. we look at and stare at at the night sky all the time.
So the idea here is that If you were an alien that is trying to communicate to us, where would you be sending your signal?
Would you be sending it all the way in the gamma ray range?
Would you be sending it in the visible light range?
It's a huge spectrum. So where would you be sending this signal if you want to get someone's attention?
Well, you should be sending your signal where everyone is looking at, hence the watering hole.
Or we think that everyone should be looking at if they're intelligent, which happens to be just 21 centimeter long.
And this 12 centimeter line happens to be within the range of our work.
Well, it doesn't happen to be. We chose it.
Right. And so that is a part where we're all looking at and looking for signs of signals.
And now there are you bring up a good point, which is that, you know, what if someone of some civilization kind of took hold of more technological power than we do and send signals in a more efficient way.
It is true that radio waves fall off by the square of a distance, so it diminishes quite quickly as you get further away from the source.
And there are other sources of physics, for example, gravitational waves. which transmits information and falls off intensity-wise by one over the distance instead of one over the square of the distance.
So it falls off a lot less. And so in other words, you can send a louder signal and it'll be easier for someone to pick up.
Right. The issue is that you need to be smashing literally black holes together to use those kinds of signals.
And so that's kind of hard to do with our phones and stuff like that and our GPS.
We haven't mastered it. you know, colliding neutron stars together again as our civilization.
I hope maybe one day we'll be able to get to that.
That's iPhone 37, I think. Yeah, I think Tim Cook has something related to that.
Yeah, it's crazy. So we don't have all the power to do that. as a civilization and that's just almost unthinkable for us at the moment.
You mentioned the open source dataset. Is this the dataset from the Green Bank Telescope?
Yes, that's correct. Okay. So that already existed and was something that researchers were working with.
And so is that where they directed you to focus your efforts?
Right, pretty much. So this piece of data, like you mentioned, it is from the Green Bay Telescope.
But more particularly, this was relatively older piece of data, so around 2016 to 2017. now the interesting part is that this was a previously searched data so this is Sorry, I should just explain for the listeners.
The Green Bank Telescope, and correct me if I'm wrong here, is one of the world's largest radio telescopes.
It's in the United States in West Virginia.
Is that right? It's one of the largest steerable telescopes that we have.
So this telescope, we got data from it and it was also previously searched by traditional or classical techniques. for signs of narrowband signals, which is what we're trying to look for. and it came out with relatively no significant hits or not the ones that we found, which were much more convincing than the top candidates found during that campaign.
That was the interesting element of that data set.
Okay, so the dataset had previously been worked with, didn't get any or no significant hits, I guess.
And then you approached it, your team approached it with a different method, kind of a novel technique.
Can you walk us through what did you do?
How did you go about going through the data and eventually coming up with some significant hits?
Right. So to understand how kind of my approach work, I think it's good to quickly walk through how the traditional technique I went through it just fine first.
So the traditional technique is basically looking for straight lines in our data.
So the idea here is that we're looking for these, like I said, narrowband They need to be what we call Doppler drifting.
So what that means is the signal is Well, it's not a straight line.
It's just like a line with a slope to it in our data.
And it needs to be also what we call have an on-off, on-off pattern.
So let me walk you through how that works because it's quite important to understand how the algorithm works.
Great. The biggest challenge that we have is that all of our technology, our phones, GPS, satellites, all emit signals or these radio emissions.
And so... That's good, but that's not great for us because when we look up at the night sky and try to look for these signals, we end up finding ourselves.
And that makes sense, right? We are intolerant.
We better be finding ourselves. But the issue is that that's not that interesting, right?
So the issue is that we need to make sure that the signals that we identify are actually coming from space. and not coming from lower orbits or like some guy's phone or some person microwaving their pizza pocket in the lounge or something, right?
And so to differentiate that, we take a telescope and we look at a star.
You're going to record some data. And you might see a signal.
And you're like, OK, great. Now what you're going to do is going to take your telescope, and you're going to not look at it.
You're going to point away from it. It turns out that if a signal is coming from nearby, it's going to blind your telescope.
So you're going to still see the signal even if you don't look at it.
That's an issue. That doesn't make any sense.
If you don't look at it, you should not see the signal anymore.
But if we still see the signal, that's an indication that is junk or interference.
And so to get rid of this, or to filter that out, classically what we do is we look for an on-off, on-off pattern.
So when we point the telescope at the star, and we look at it for signals, and we see a signal, if we point away from it, be off, we should no longer see this signal.
So in our data, it'll look like a line with uh chunks missing from uh the times we're not looking at yeah exactly not looking at the star And so the classical algorithm, all it does is it looks for straight lines just by doing something called tree-to-dog chlorine search technique, basically by shifting each of these images, effectively images, you can think of them as that, such that they form a straight line, and then it sums it down to find the highest peak.
It's basically just line detection. It's pretty simple.
A singular linear line detection. Now, the issue with that approach, the traditional approach is that A, it's quite slow and B, it is also quite restrictive.
Because, you know, we're only looking for straight linear lines.
Right. And, you know, who knows what is going to stop us, right?
No. Right. So we just know it's a narrowband line, right?
But we don't know if it's going to be linear or squiggly or a parabolic curve or anything.
We don't know. So what we want to do is we want to build a general anomaly detection algorithm.
So this is where the actual machine learning part comes in, where we pair together both unsupervised learning and supervised learning approaches.
So the idea with unsupervised learning is we're going to get an autoencoder and we're going to feed it a bunch of data. and is going to understand what to expect from our dataset effectively.
So what is considered expected. and it's going to have a hard time while recreating or extracting features from an unexpected signal, for example.
So that's the unsupervised approach. And then we realized that that's cool, but it's returning a lot of hits.
Like, it is... kind of unwieldy to control the enterprise model.
So to fix that, We then took the features from the encoder part of the autoencoder and we fed it into a simple classifier, something fast like a random forest. simple random forest classifier.
And we found that if we simulated some signals and fed it into the autoencoder, encoded as features, we were able to differentiate basically between a signal of interest and interference quite easily with quite a substantial improvement to a simple, say CNN classifier.
Our mixed approach was performing a lot better than a traditional CNN piece of work there.
We also demonstrated, it wasn't in the paper, but we also demonstrated it was able to pick out signals that was never trained on in our data set. went kind of crazy and just injected like random circles and like emojis and stuff like that.
It was just like, I have no idea if this algorithm is going to pick it up.
But it ended up doing just fine. So it was just surprising.
It was just like a sanity check. so to speak, not really scientific per se.
His algorithm is able to sort of basically generalize to what we wanted to do, basically.
That was the idea. And so that's the approach that we took.
We basically took inspiration from the traditional techniques, figured out what was the issues with those, and made them slightly better with improvements with machine learning department.
How were you able to measure accuracy or success?
Yeah, so we had simulations of signals of what we should expect.
Obviously, this is not exhaustive because Once again, we can't anticipate all kinds of signals and different morphologies and all that.
But we did have some tests with ground truth, and those were the simulations.
We injected them into real observations to make it as realistic as we could, and we found that they were just much more performant. in terms of precision and recall and all that, all those class metrics for a classifier.
I'm speaking with Peter Ma. Peter is an undergraduate student at the University of Toronto and the lead author of a paper published in Nature Astronomy, detailing a novel approach to searching for extraterrestrial life using AI machine learning algorithms. to sift through large data sets, I guess, instead of searching for a specific pattern, you're instead kind of creating an anomaly detector today.
Kind of the last couple of minutes of what you were saying correctly.
All right. Good. I really wanted to focus in on the emojis.
I was hoping that the looking face got a response from out there.
But, you know, you can't quite be. I want to go back for a minute to how this all got started.
You mentioned you were in a high school computer science class and it was too easy.
You got bored. And so you kind of taught yourself the rest of the curriculum.
Yeah, pretty much. What was the curriculum?
What were you studying? It was just a introduction to computer science and we're coding in Java.
And our teacher was like, okay, so all the course content for this entire semester, including all the assignments are in this folder and we'll be working throughout the semester.
And then I downloaded it onto my laptop and I did it over like three weeks.
It's just fun. This is cool. Now I have a lot of...
And so how did you get from, you know, beginning Java to machine learning algorithms?
Right, so like I said, I had quite a few time, quite of extra time every day, about like an hour and a half during that period to... do anything I wanted basically.
And so I continued to, while it's all taught myself Python, I was like, wow, this is a lot easier.
It's just a lot of cool interesting. And then I continued to work with various projects I was interested in.
A lot of people were suggesting web apps and stuff like that, but that was not really my thing.
So I got into more so related to scientific computing.
And so that naturally led into a lot of data analysis and data science and I'm actually machine learning and deep learning.
And so I signed up to a bunch of like online courses just and YouTube videos and reading papers.
Yeah. And so that kind of really started the journey off there.
And I continue to do it until this day. And so you found Python interesting?
More interesting than... Well, I found it slightly more useful in the scientific realm.
Sure. being able to do things quickly, obviously not as robustly probably, but It was able to get interesting results and it was quick to visualize a lot of data science stuff.
So it felt like a very useful tool, at least for the science side of things.
So as somebody who it sounds like kind of ramped up very quickly in your hands-on experience working with different aspects of computer science and machine learning and working with large data sets.
How did you find the process and the available tools right now?
And sort of, is it It sounds like it was relatively easy, relatively interesting for you to learn. particular things that stand out as being either difficult or particularly exciting just about not just about the research paper, but kind of more broadly getting your hands wet with, or your feet wet, excuse me, getting your feet wet with working with AI and ML algorithms?
So by the time I kind of jumped on was around 2019, so A lot of frameworks were quite well established and getting you started with a lot of these what was considered state of the art algorithms, like probably like five or 10 years ago, if not.
And so in terms of getting something working, it was not difficult.
But what was difficult for me, and I'm not sure it was a lot of people, is understanding the fundamentals, because there's quite still a bit of math involved. and a lot of vector calculus that my high school has still yet to teach me.
And so that part was relatively a struggle.
That part was more difficult than I expected, but to actually get results in a working kind of model didn't take quite a lot of effort.
In fact, it felt like Lego to me. I was like, this is just advanced Lego just on a computer.
You know, something like Keras where you just, you know, model.add and something comes up and you add another layer, you add layers.
Add this later. I don't know what that does.
If I run it, if it doesn't give me any bugs, that should be fine, right?
And so that's basically how I went through teaching myself all these things.
I was just... playing with what's considered CDR and kind of asking more and more questions as to why does this work?
Why do I need to do it like this? Why does this tutorial work in the way it does?
That kind of led me on, which I guess is a flipped style than normal, how people are normally taught in a curriculum, I suppose.
There's a lot of math that builds up towards doing an actual deep learning course.
And then you actually supplement that and say like, you know, like a lab or tutorial where you actually build these things But rarely do we build anything from scratch anymore, which was kind of abstracted away all the fundamentals, which kind of locked back in.
But yeah. No, it's a really interesting and very salient point you make that the model of learning with so much accessible now, so many things that you can actually go and you can try it.
And as you said, you know, I'm not quite sure what this works, but there's really no downside to trying it.
Worst case is I get some bugs and you know, and then if it does work, and it is of interest to you, then you can sort of go back and start to figure out, well, wait, what is this all built on?
What's the math behind it? You know, et cetera, et cetera.
So I guess an interesting paradigm shift.
So when you started reaching out to different scientists, experts in the field, cold emailing them, as you said, What was the response you got?
Were you welcomed and greeted enthusiastically right away?
Did you not get responses? What was that process like?
Right, so the thing about when I reached out to this group was that they were all very excited about the work I was doing.
And they were interested in what I was doing, even if I have basically no credentials, I was just a high school kid I was just interested in.
And all I had to offer was just enthusiasm.
Basically, I'm willing to learn and all that, all that kind of stuff.
I suppose they saw some potential in me or saw something in me that they were like, oh, let's continue to support this person and work on this project.
This is in contrast to, say, a lot of other kinds of academic labs in other fields. where the response rate for a high school student is quite low.
I've reached to other people before in interest about their work.
And either I get no responses or they're like, you know, they just straight up tell me, no, they don't have any time or they're not interested in it.
And so I think this lab, or at least a Berkeley study on a lab is, I guess, in some ways, more progressive in some sense, in the idea that they're willing to take a chance on someone who is curious and who are interested in their work and was able to invest in a person such as myself. and working on this kind of piece of research project.
And I'm really thankful for them for taking me on, for basically taking a gamble on Some stranger on the internet.
No, it's very cool that they did. This is the $64,000 question.
Did you find any signs of life out there?
Are there aliens as we call them? people trying to make contact or beings trying to make contact?
So we haven't found them just yet, or at least I don't think so with the paper that we released.
But what was exciting was that the signals that we did find were suspicious enough to warrant us re-observation and just double check that we didn't actually find aliens just yet within our reports though so that was the exciting part and so In other words, it basically passed all our preliminary checks, was missing what we found.
And we're like, this looks very suspicious.
Let's look for what this might actually entail.
And so just to answer your question, no, but to give a more hopeful answer is that we found very convincing signals that we need to do further observations and Not aliens, but not a Hot Pocket being microwaved either.
Yeah. We're not sure just yet. Something in there.
Cool. So what's next? What's next? Is this both for this particular project?
Is this group staying together doing more research?
And then kind of more broadly, what's next for you?
In terms of just the research work where the group is trying to scale this project to, Yeah, much more stars.
So the Breakthrough Listen initiative, one of the goals, of its many goals is to search the nearest 1 million stars for signs of life beyond Earth.
And well, that's a big step up from 820 stars from our work here in our paper.
And so one way to do that, one way to scale that is to use more telescopes and more use them at the same time.
So there's an observatory in South Africa called Meerkat Telescope. work telescopes, I should say.
There's 64 antennas. So instead of taking one dish with our telescope in our paper here, we're using 64 eyes and on the night sky searching at the same time, 24-7, seven days a week all the time.
And so that would take, I think, approximately two to three years Something that will take a couple of years to collect all that data to run our analysis on.
And so part of that pipeline, we hope to use some elements of deep learning and also still use some elements of classical search techniques We don't want to abandon classical techniques either because they hold true interpretable astrophysical parameters help us scientists actually understand what these signals are.
It's not just a black box machine learning algorithm that's like, oh, is this true or false, or if this is a signal of interest or not.
And so we want to pair them together in some meaningful way.
And so I'm kind of working on that end with the deep learning side.
And so that brings us to where I'm kind of headed next.
So I'm still in third year. I continue to want to work with the group further in the future, partially because I am involved with this, I suppose, long-term project of years.
I would like to see, first of all, my work being implemented and also want to see what my colleagues end up doing with this piece of data, what we can find.
That would be extremely interesting and cool.
But in a broader sense, where I'm headed in my, I suppose, career, well, the natural default path for someone in my position is to go to grad school.
And so if that's the path I choose, then anything between astronomy and physics, and i even might consider computer science because honestly i feel like there's an extreme amount of interesting implications or intersections between these three kinds of fields that have yet to be kind of realized.
I think in the field of at least machine learning and deep learning, I think there's a lot of interesting applications of using science as a benchmark instead of other kinds of arbitrary benchmarks that computer science typically use, like CIFAR-10 and all that kinds of stuff that people are traditionally familiar with.
I think that if as computer scientists swap datasets, supposedly, or swap problems to actual scientific problems, not only will be interesting, but it will also help solve other emerging problems in other kinds of fields. to kind of work in the intersection of both deep learning and physics slash astronomy, kind of fundamental sciences and stuff like that.
And so I see a lot of potential. And yeah, I'll love to be a part of that in the future.
Well, it certainly seems like opportunities abound, whichever path or paths you take.
It's very, very cool to hear about. And, you know, kudos to the whole team, but congratulations to you for I don't know, for being bored in school and deciding, hey, you know what?
I'm not just gonna sit here and be bored.
I'm gonna pursue what I'm interested in and see where it takes me.
There's a blog post on the NVIDIA blog detailing more of the work you were just describing, and it links that to the paper that we mentioned in Nature Astronomy.
But for folks listening who might want to find out more about anything we've been talking about, Breakthrough Listen, the SETI group at UC Berkeley, maybe your own, if you have a homepage, social media, anywhere that details some of the work you've been doing, where would you direct them to go online to find out more?
My own homepage for all my research is at petermong.ca.
So that's kind of my research site. has a bunch of all my other works that are somewhat adjacent related.
You can find all that cool stuff there. And also, it has my link to my GitHub, which has the code and all that stuff you can look at yourself that was related to this project as well.
So all the details are all open source there.
And if you want to also get access to all the data, they're all also open source.
You can find them in our repositories as well online, which is on the UC Berkeley SETI research group, like the website there.
I don't have the exact URL memorized in my head, but...
You have enough going on in your head. We can leave the listeners to find that one on their own.
Well, Peter, this is a fantastic... Congratulations on all the work that you've been doing.
And I have this feeling I'll be coming across your name in... you know, reading about scientific discoveries in the years to come.
So good luck with everything in the future.
And thank you for taking the time to come on the podcast and share your story.
Great, thanks for having me on again. Thank you.