Hello, and welcome to the NVIDIA AI Podcast.
I'm your host, Noah Kravitz. The intersection of AI and biology is one of the most fascinating and promising areas of modern technology and research.
My guest today is working at the leading edge of this field in his role as CTO of Basecamp Research.
Basecamp, who's a member of the NVIDIA Inception Program for Startups, is leveraging their unprecedented knowledge of the natural world to create better food, better medicines, and better products for the planet.
Basecamp has collected an unprecedented dataset capturing orders of magnitude more diverse biological data than any public resources.
And they're leveraging this data for deep learning and Gen-I applications.
Here to shed light on what that means for his company and for all of us is Phil Lorenz, Chief Technology Officer at Basecamp.
Phil, thanks so much for taking the time to join the podcast.
Thanks so much for having me. How's your trip been so far?
We're recording on, I guess this is day three of GTC, so you've been in town for a few days.
How's the conference? It's amazing. It's great to see some friends who I haven't seen in a while.
So that's actually been great meeting with folks from NVIDIA that we've been working with for a while.
And yeah, I mean, lots of great networking opportunities, including people that are really far outside the life science industry.
But it's great to see what everyone else is doing.
So really exciting. Yeah. It is. It's nice to be back in person after several years here.
So let's start with the basics. Maybe you can tell us what Basecamp Research is, how you were founded, what you do.
Yeah, of course. Maybe to take a step back with respect to why we're doing what we're doing and how we thought of this.
I guess if you think about kind of the life sciences, probably one of the most exciting domains to apply AI to, I think.
And we obviously have a lot of human clinical data collected in the last few years and decades.
But when it comes to kind of life on Earth and biology as a whole, we actually haven't because there's probably about 10 to the 26 species out there.
Right. which is a lot, and we've sequenced a few million.
And so if you kind of make that comparison, In terms of what we know about life on Earth, that's about five drops of water compared to the Atlantic Ocean, which is what we don't know.
If that's the kind of place you're starting with, and everything the life science industry has ever built is based on that tiny knowledge of life on Earth, that kind of slice, We thought that if you want to do deep learning for the life sciences really well, there's exciting algorithms being built, exciting architectures that we can do.
But at the same time, we feel like, okay, there's actually a big data problem to begin with.
And so to do this from first principles, what we've done over the last two or three years We've built partnerships with nature parks across five continents, including places like the Antarctic and rainforests. volcanic islands and you name it.
And we have professional explorers that go to these places and they sequence the biodiversity in these areas.
A lot of microbes, Because that's where the greatest diversity of life on Earth is.
And we do this in partnerships with these nature parks.
Right. When you say sequence the biodiversity, am I getting that right?
Sequence the biodiversity that they found?
What does that mean for kind of the layperson?
Absolutely, yeah. Realizing that I'm a heavy life scientist.
I'm saying the layperson. I really mean me.
No, that's completely fair. Sequence basically means every organism has a genome, has DNA, and so sequencing the genomes of all of these unknown organisms.
And is that, just to kind of take a tangent for a second, can that be done out in the field remotely?
Are we at the point with gene sequencing where you can do that on site, or how does that work?
Yeah, so when the founders started Basecamp, they actually did that on site.
They were on an ice cap in Iceland with a solar-powered tent.
That's how the company started. Now, because we're doing this at scale, we're extracting the DNA there and then we're sequencing the DNA at a much larger scale in Europe.
Yeah, and I think what's exciting is, again, this is underappreciated how the vastness of unknown life on Earth, we now have a database just within... two years that Hassam was collected from the Antarctic and volcanic islands and all the life that exists in these places.
That's now a few orders of magnitude more diverse than all public data combined.
And what we're doing on top of this is not just collect hundreds of millions of new protein or DNA sequences. but also the chemical environment, the geological environment, and connected all of that information together in a knowledge graph. that now has about 6 billion relationships.
And so that really gives us a good information of entirely new, never seen before information that has never existed before.
To go back to your analogy about drops of water in the ocean, if we were at five drops of water compared to the ocean of knowledge, How much more knowledge have you been able to accrue in these past couple of years?
Yeah, we're probably still a few orders of magnitudes away from the Atlantic Ocean.
We're trying to be maybe a cup of water.
Cup of water. Getting close to that. That's kind of the goal in the next few months.
That's amazing. And so what do you do with the data?
There's several things. I mean, we have a huge, there's just a huge engineering effort to just annotate all of this data, organize all of this data, because that's you know, is now bigger than most public database and that's actually a big undertaking.
The exciting application really is to see, okay, There are some architectures for deep learning in biology, such as structure prediction of proteins.
And now that we have so much more diverse data, what we can do is actually leverage some of the algorithms or architectures that exist and apply that to our data advantage.
One thing that we've now built is called BaseFold, which is using a similar architecture to AlphaFold but we can actually be up to six times more accurate because we have so much more diverse and additional sequence information.
And so that's something that's... Really exciting because there's obviously a lot of work and effort being done on using a different platform algorithm, different methods.
But actually data, especially in the life sciences, makes such a big difference.
That's something that we're really excited to use.
What the right way to get into this is, I wanna ask about what your sort of day-to-day life is as the CTO.
But then I'm also curious what happens. Are you working with partners across academia, industry, to leverage the data to create better medicines, better foods, those kinds of things we talked about at the intro.
So either way, talk us through kind of what you do as CTO, and then maybe from there we can talk about some of the partnerships.
My role is that I do a lot of very different things at the same time.
The one thing I'm not doing anymore at all is coding.
I did that at the start, but that's kind of the one thing I'm not really doing.
But that's probably a good thing. My role has kind of three, four main things.
The first is actually making sure that the data collection process and how we enter that into our our database.
We have a genomics team that's amazing and they're doing incredible work dealing with all of this data and having really high quality annotations for that.
That's kind of one thing. Then the data engineering, how we organize all of our infrastructure, the deep learning work. applying this data advantage to the most exciting AI applications and then the product team.
And that's actually where What you mentioned with respect to what our partnerships look like with therapeutics or biotech companies.
Some of them they will ask us oh do you have a protein that can do this function like an enzyme that can break down plastic and then we work with them on that.
Or sometimes gene editing systems that will cure genetic diseases.
And we have these in our database occurring naturally.
But because again, of our data advantage, we can use generative algorithms to actually generate these assets and then they license them from us and then they use them for downstream clinical applications or whatever that might be.
I'm familiar with the basic idea of gene editing, but not much beyond that.
If a company comes to you and, for instance, says they need an enzyme that can break down plastics.
Yeah. And let's say, I don't know if that occurs in the natural world or not, but let's say it doesn't.
What can you do then? How does that work?
So there's kind of multiple ways in which we look at this.
The first one is, let's say someone wants an enzyme to degrade plastic.
In many cases, that assumption of let's say, oh, we're not sure whether this exists in nature or not, is actually that assumption is made on what we know from public data.
Right. And actually, we are making new biological discoveries based on our dataset all the time.
So that's kind of one argument. But then there's another way of thinking about this is that because of the data advantage when we use deep learning to optimize or generate these enzymes for a specific function.
Because we have explored sequence-based and evolution so much more, we can actually understand how to get towards even these unnatural reactions much easier in many situations.
Right. So you said it's been about two, a little over two years that you've been gathering the data sets.
How much has AI, the technology that's available that, you know, compute as well as data, how much has...
That world changed in those couple of years relative to you thinking about, well, We're collecting this enormous amount of data and it's fantastic and it can open so many doors, but do we have enough compute?
Do we have the right algorithms to be able to work with it?
From the outside, and especially we talk a lot about generative AI these days, you know, for the mainstream, right, that world has exploded.
Yeah. In that time period. But for the kind of work you're doing, has the change been as dramatic?
Absolutely. There were a couple of situations a few years ago even where We had maybe not the same size of the data set, but the same foundational architecture with a long genome context and all of this metadata that we collect. where I was actually thinking, oh damn, I don't think there's an architecture out there that can deal with our data.
And that is slowly starting to change. And I think one of the most exciting things in the deep learning applied to biological tasks in the past few months and years is that What a lot of people have done is thinking about what additional biological context can I include into my language model architecture or something.
Not just use something from a different domain and force that architecture onto a biological question. but think about how can I change that model architecture in a way that represents biology much better.
I think that kind of trend in the last few months is really exciting and it yields much better results as well and that's great.
Are you taking off the shelf models and fine-training them?
I mean, fine-tuning, excuse me, for your own use, or are you building models from scratch?
How does that work? We're doing both. So on the folding problem, we use pretty much AlphaFold's architecture that exists because I think it's pretty good and it works.
And for that it's purely just doing a much better job because of the data that we have.
We've also built our own architectures and models.
One for annotation, there's a lot of what are called functional dark matter sequences, where we have a sequence, but we have no idea what it does.
And if you have sequences from the Antarctic that have never been seen before, It's actually important for us to be able to computationally say what they do.
And so for that, we've developed some contrastive deep learning. algorithms to annotate them at pretty high accuracy and we presented that at NeurOaps last year.
We kind of do both. It depends what we feel like is worth building something from scratch versus... just leveraging our data advantage.
But even when we built something from scratch, we're always leveraging our data advantage as well.
So it's kind of, we're doing both, whatever works.
I'm speaking with Phil Lorenz. Phil is the chief technology officer at Basecamp Research, and we're speaking high above the... show floor here at GTC 2024.
Our podcast recording area has a nice window view of the show floor coming to life this morning.
Phil, you came from an academic background.
You were at University of Oxford before joining Basecamp.
Maybe you can walk us through your journey a little bit, and then we can talk a little bit about what's the same, what's different from moving from academia into your role now.
Yeah, definitely. I'm a traditional life scientist at heart.
You have a lot of people in healthcare life science now that have a computer science background walk into that and that's amazing that's that's incredible i'm right I'm a little bit from the kind of molecular biology traditional background, which is what I did during my undergrad. and then moved towards more kind of deep learning applied to genomics sequencing data for my Ph.D.,
I discovered a couple of new human genes and transcription starts using deep learning during my PhD.
So that's kind of where I worked on for a while.
One thing I always really cared about whatever you do in the life science or healthcare industry is thinking about the problem you're trying to solve first and And then going backwards and thinking, okay, what kind of technologies, tools can you use to address that problem?
That's kind of always thinking about what you're trying to do in that kind of order of events.
And that's still how I think about this now, even though I wouldn't really say I'm a traditional life scientist anymore.
But always thinking about what you're trying to solve and then what do you need to build second.
That kind of philosophy is still something I...
I think about quite a lot, even though my PhD was quite applied, so it wasn't too academic, which is maybe a good thing.
But yeah, that's kind of how I came into what I'm doing now.
Did you learn to code when you were younger, kind of out of an interest in computer science and learning how to code or was it more of a, in your work in the life sciences, you kind of hit a point where you realized, oh, this will be faster if I learned how to write scripts.
It was almost, I started coding almost out of a necessity.
Yeah. When I got my first kind of big data sets, and at some point I was like, yeah, I'm not going to use Excel for that.
So it was almost kind of, I wouldn't say I was forced to, but at some point I was like, well, it's just going to make everything more efficient, faster, and so on.
It was kind of great because I felt like I was coding always with a purpose to do something specific for what I wanted to do with my project. that way i kind of always felt really motivated to to go after it so that was that's kind of good so how long have you been at base camp now Basically from the very beginning for almost three years now.
Okay. And how big is your team now? How has it grown out over that time?
The company as a whole, about 32 people.
My team is 15, but a bit less than half the company.
Can you tell a story or kind of explain a discovery along the way at Basecamp that will blow our listeners' minds?
Is there something about, you know, sequences from Antarctica or You know, something undiscovered about the way, you know, ecosystems work in the desert or for that matter here in San Jose.
I don't know. But something that just really sticks out.
Yeah, I mean, one thing that I still think is a nice thing to share. is actually from the very, very beginning, the very first few days of Basecamp was, I mentioned this earlier, but the two founders of Basecamp, Glenn and Oli,
And big kudos to them having this kind of vision.
They love exploring the world. They go out in the wild.
They go climbing. They go up mountains, whatever.
I'm more fragile. Stay behind the screen.
And so at some point, a couple of months before they started Basecamp, they spent over a month on an ice cap, fully off-grid in Iceland.
And Glenn did a lot of sequencing, DNA sequencing during his PhD.
So he brought these kind of mini flow cells with him, these portable sequencing devices.
And he was just sequencing ice caps and see what was in there.
You probably don't expect much life to be happening there.
And so they came back with this data. And they realized, oh, Phil, you do a lot of this analysis, and you do coding in your PhD.
Can you have a look at what's in there? And I analyzed this data.
I annotated this. And something like 97% of it has never been seen before in any public databases.
It was completely novel, had 0% similarity to anything that's ever been seen before.
And that was just a random spot glinted somewhere on an ice cap.
We didn't look for something new. It was just random.
Yeah. And so from that, we just realized, oh my God, the vastness of life on earth is just so huge.
And the opportunities lost in the life science industry by not leveraging this data more systematically that's kind of was a great kind of origin story for us to realize like let's do this systematically and with deep learning in mind because a lot of the data that we have in the life sciences is us, you know, having kind of all these academic endeavors sequencing here or there and fingers crossed that they've collected the right data, right?
And so we've kind of made this systematic and in partnership with all of these nature parks, which is exciting.
Yeah, that's amazing. From the technological side, we talked a little bit before about the infrastructure, the compute, the techniques.
I don't want to say catching up to, but sort of keeping pace with the size of your dataset along the way.
What are some other AI machine learning related challenges that you've encountered at Basecamp? that you got past or perhaps that you're sort of still grinding on now.
Yeah, I mean, one thing that I am most excited by that has been kind of addressed in the last few months is how to deal with bigger context sizes.
So that's kind of in the language model field.
There's a couple of architectures that were developed, especially I think in Stanford last year, Hyena and Mamba. that I'm super excited by because what we're collecting is not just protein sequences but these really long range genomic context windows that's not really that common in other public data.
Can I ask you to explain what that means?
Yes. I mean, when people sequence environmental data...
Not like a strain from a Petri dish, but kind of these wild environmental samples.
You might have 10,000 species in a tiny piece of soil or something.
When you sequence that, usually what happens in public data is you get maybe a few thousand base pairs, and if you're lucky, there's one gene on there.
What we've done is we've really tried to optimize this process in a way where we get near full genome, so the entire genome of every single organism of that.
So hundreds of thousands of base pairs with tens of thousands of genes or whatever that might be.
And so with that, we're actually understanding a lot more about the interaction of all of these genes, what they do to work together, and also understand more complex behavior.
For example, how you can use this for gene editing or therapeutic applications.
But modeling this with language models hasn't really been that straightforward because... there wasn't really that many architectures out there to deal with that kind of information.
And with Hyena and Mambo, we're now really excited that this is now possible and that we have the data set that we can apply this to.
So that's something I think in terms of dealing with long context, probably the most exciting development for for very selfish reasons basically but that's that's i'm super excited very cool very cool So you kind of hinted at this in that answer when you mentioned gene editing.
Yeah. What are some of the applications going forward for the work that Basecamp's using and then –
I don't know, and I'm not asking you to compare to competitors or what have you, but as the available data sets... you know, sampled from biology, from nature grow.
They continue to grow. What are some of the implications and applications for you know, everyday folks like me downstream from the work that you're doing.
The reason I think gene editing and at some point gene writing technologies is going to change not just medicine, but health in general.
We accumulate millions of mutations every day just by existing, by breathing and eating.
And some of them are non-significant, some of them are bad, some of them are maybe good.
But for us to be able to accurately change them. or write new DNA into the human genome to make a change.
That's really, I think, kind of the next wave of big therapeutic changes that we can make.
And the machines that can do this actually often are derived from the way bacteria and viruses fight with each other. phages, those viruses that infect bacteria, they kind of have biological warfare going against each other in the wild, in nature.
It's not something you can measure in a Petri dish with a sterile strain or whatever, but in nature they have warfare against each other.
And those machines that enact this warfare, those gene editing systems, CRISPR-Cas nucleus is like one of the kind of major headlines that came out of that. there's hundreds of millions of these systems that haven't been discovered yet.
And leveraging them in a way that we do, but also in a way where at some point by having enough of these that we can generate them or design them through language models, for example, That's a development that I think is really exciting.
You know, I'm abstracting to the level that I can comprehend, but the last bit that you said, I was thinking...
So my kids, maybe my grandkids, might be able to prompt a model and – edit their genes or rewrite their genes.
And I know it's maybe not quite like that, but...
Is that a future we're headed towards? I think, I mean, I live in Europe, so there's always a lot of regulation to it. to think about maybe so i don't know i can't promise all of it no but i think um joking is that i think One of the things I can definitely imagine is that if the way we monitor our DNA and our mutations And the way we can address these mutations, let's say in 20, 30 years' time, is something we can do in real time, I can imagine, where... because of sequencing in the body as we live and breathe, where through some device we detect a harmful mutation.
Right. and being able to fix it within two hours.
This sounds crazy, but I do think this is kind of where this is going in 20, 30 years.
And that's kind of the... the science fiction scenario for therapeutics and gene technologies.
Amazing. NVIDIA Inception. You're part of it.
I'm not asking you to plug anything, but how's that been?
And, you know, Being sort of a startup on the leading edge of life sciences must be – in some ways similar to other startups with similar concerns around growth and funding and keeping keeping things running and all that kind of stuff.
But I'm sure there's something unique to being a startup working on, you know, discovering novel science.
What's that like? And what's it been like working with Inception?
It's been amazing. We've been working with NVIDIA for...
Two years, almost two years now. I think it's really exciting because a lot is happening.
And so, I mean, sometimes I open, you know, BioArchive or PubMed or something.
It's like, damn, can everyone please stop publishing?
There's just so much happening. But I actually think it's exciting because everyone has their strengths.
And by having these networks of... lots of life science companies and everyone has a different product, they have a different strategy and so actually Some people think like, oh, are these startups all competitive with each other?
And in some cases, maybe, but I'm actually a lot more excited by the fact that what's really happening is we're all growing the field.
We're all growing the market. And so some people offer a software, some people offer an asset.
Some people will offer a service, whatever it might be.
And so actually, the products are different, the technologies are different, and so just the space growing as a whole.
It's something that's super exciting. And NVIDIA and Inception, they're connecting everyone and making it happen, right?
And so that's something I'm super excited about.
That's fantastic. So you mentioned coming up with something of a traditional – I hesitate to call it old school because as we're sitting at the table – I'm not going to guess, but I know you're a fair amount younger than I am, so –
If you're old school, I don't want to think of what I am.
But coming up with more of a traditional life sciences background. and then kind of moving to a place where you moved into applying it and using technology in that way.
Jensen said something in the media a couple of weeks ago about giving advice to young folks. to focus on a domain and develop domain expertise because, You know, the computing language of the future is just speaking, right?
It's natural language. And so the tools will progress so that you can leverage them Right, yeah. in this age where everything's moving so fast and technology is such a big part of it.
I guess, yeah, I mean, I speak to a lot of kind of biologists, but also computer scientists that are buying through Basecamp.
To the biologists, my main advice is always do what you care about.
There's a lot of biologists that have something they really care about, but then they go into oncology.
Because that's where big pharma is. And that's great, but I actually think there's so many areas in the life science industry where The problem is not that they're not relevant.
The problem is that they're not relevant yet.
Because by making more discoveries, we're always going to find something that's clinically relevant. like gene editing was found through studying bacterial immunology which is like right no one ever thought was going to be relevant therapeutically a few years later right so I think it's always better to do something that you're passionate about and make it relevant rather than trying to find something that, oh, this is what therapeutics cares about and just running after that.
That's my advice to biologists. For computer scientists, the main thing I find is when we have people apply to Basecamp or interviews and so on.
Often what I hear is people kind of saying, oh, I really care about this specific type of diffusion model and I want to apply this.
And my counter argument often is kind of, let's discuss what we're trying to solve first, and then...
Does that help? Or should we think about something else?
Or should it be a language model? Or should it be, you know...
Regression, I don't know, but basically always, this is often what, when I speak to people from kind of a more technical background always.
And this goes back to Jensen's point about the domain.
To interact well with the domain, the problem they're trying to solve, And then absolutely, yeah, talk about the technology and what it's going to do.
But identify the problem first and then apply the tools.
Yeah, it makes good sense. One thing I haven't really talked about much is you probably know about things like the New York Times suing OpenAI. because they didn't ask for their data.
So one thing that we're doing that's different to any other life science organization is that We're not just asking all these nature parks for consent because all the public databases, they never did.
So we don't just ask them for consent. What we're also doing is when we license something to a partner, whatever, we often do something like a revenue share with them. to basically make sure that all the progress of life sciences or the AI that comes out of all of this data We share this with the stakeholders that originated the data.
So that's to kind of protect biodiversity, but also have a different model where there's kind of good data governance for it.
I don't know, but that's kind of one thing our team does, yeah.
When you started approaching the nature parks about this, Were they receptive?
Were they confused? Did they understand what you were talking about?
And I don't mean to be disparaging to them.
It's just unique. It's interesting. So...
Not to say the global south, because that's a lot of vastly different places, but the majority of biodiversity is obviously like South America, Africa and so on.
They are on top of this. There's something called the Nagoya Protocol, and it actually means you have to not just ask for consent, but share benefits with them.
And so they are aware of this. They're almost waiting for the West to ask them and work with them.
So a lot of them are super on top of this.
Interesting, yeah. Yeah, we're just the only ones.
We have a team that is part of the United Nations Conventions on Biological Diversity.
So that's a long title. But they're working with nature parks, with local governments, national governments to make these deals.
And sometimes that takes time. That takes months to have an agreement, but that way we know that for every single data point in our database, we don't just have... consent and permission of where that comes through but also when we have a when we see a commercial success through something we can share some of that with them and that That incentivizes, obviously, an even bigger data supply chain, which is exciting.
Phil, for listeners who want to find out more about what Basecamp Research is up to, there's a website.
Should they go there or where should they go?
Yeah, absolutely. Our website, very simple, basecamp-research.com.
I think we have LinkedIn, Twitter, and so on as well.
You can email me if you want. Fantastic.
Absolutely, yeah. Great. Well, thanks again for taking the time out of GTC to speak with us.
It goes without saying, but it's fascinating, fascinating work you're doing.
And I only understand it on the level of a couple of drops, not the whole ocean. but can't wait to see what the rest of the year holds for you and for Basecamp.
Thank you so much. This has been great. Cheers.
Thank you.