I want to think of it as what I would call a sort of a physics processing unit, like a PPU, right?
Which is you have digital processing units and then you have physics processing units.
So it's basically nature doing computations for you.
It's the fastest computer known, possible even.
It's a bit hard to program because you have to do all these experiments.
It's also quite bulky.
It's like a very large thing you have to do.
But in a way it is a computation and that's the way I want to see it.
You can do computations in a data center and then you can ask nature to do some computations.
Your interface with nature is a bit more complicated, but then these things will have to seamlessly work together to get to a new material that you're interested in.
Yeah, it's a pleasure to have Max Wolling as a guest today.
Max has done so much over his career that I've been so excited about.
If you're in the deep learning community, you probably know Max for his work on variational autocoders, which has literally stood the test of time or officially stood the test of time.
If you or a scientist, you probably know him for his like pioneering work on graph neural networks, on equivariance.
And if you're a material scientist, you probably know him about his new startup, Kasp.ai.
Max has a long history doing lots of cool problems.
You started in quantum gravity, which is, I think, very different than all of these other things you've worked on.
The first question for AI engineers and for scientists what is the thread in how you think about problems?
What is the thread in the type of things which excite you?
And how do you decide what is the next big thing you want to work on?
So it has actually evolved a lot.
In my young days, let's put it, I would just follow what I would find like super interesting.
I have kind of this sensor.
I think many people have, but maybe not really sort of use very much, which is like you get this feeling about getting about very excited about some problem, right like it could be.
You know what's inside of a black hole or what's you know at the boundary of the universe, or you know what, what is quantum mechanics actually all about?
And so I've followed that basically throughout my career.
But I have to say that as you get older, this changes a little bit, in the sense that there's a new dimension coming to it and this is impact.
Working in two-dimensional quantum gravity you pretty much guarantee there's going to be no impact on what you do.
Relative, you know maybe a few papers, but not in this world at this energy scale.
As I get closer to retirement, which is fortunately still 10 years away, or so, I do want to make a positive impact in the world.
I got pretty worried about climate change.
Um, and I think we and um, I think we should, you know, and politics seems to have a hard time solving it, especially these days.
And so I thought better work on it from the technology side.
And that's why we started cost pay high.
But there's also a lot of really interesting science problems in, you know, material science.
And so it's kind of combining both the impact you can make with it as well as the interesting science.
So it's sort of these two dimensions, like working on things, which you feel there's something very deep going on here.
And on the other hand, trying to build tools that can actually make a real impact in the world.
So the thread that when I look back, look at the different things you worked out, some of them seem pretty connected, like the physics to equivariance and graph neural networks maybe.
And that seems to be somewhat related to CUSP.
Do you have a thread through there?
Yeah, I think physics is the thread.
So, having done, you know, spent a lot of time in theoretical physics, I think there is first very fundamental and exciting questions, like things that haven't actually been figured out in quantum gravity.
So there's really the frontier.
There's also a lot of mathematical tools that you can use right, for instance in particle physics, but also in general relativity sort of symmetry plays an enormously important role.
And this goes all the way to gauge symmetries as well.
And so applying these kinds of symmetries to machine learning was actually, you know, I thought of it as a very deep and interesting mathematical problem.
I did this with Taco Cohen, and Taco Cohen was the main driver behind this.
Went all the way from just simple like rotational symmetries, all the way to gauge symmetries on spheres and stuff like that.
So, and uh, maurice weiler, who's also here, when he was a psd student with me, you know, he wrote a, an entire book, which i can really recommend, about the role of symmetries in ai and machinery.
That i find is a very deep and interesting problem.
So more recently, So I've taken a sort of different path, which is the relationship between diffusion models and a field called stochastic thermodynamics.
This is basically the thermodynamics, which is a theory of equilibrium but then formulated for out-of-equilibrium systems.
And it turns out that the mathematics that we use for diffusion models, But even for reinforcement learning, for Schrodinger bridges, for MCMC sampling, has the same mathematics as this physical theory of non-equilibrium systems.
And that got me very excited and actually when i taught a course in muizenberg it is south africa, close to cape town at the african institute for mathematical sciences ames, and i turned that into a book.
So two years later the book is finished, i've sent it to the publisher, and this is about the deep relationship between free energy diffusion models, basically generative AI and stochastic thermodynamics.
So it's always some kind of, I don't know, I find physics very deep.
I also think a lot about quantum mechanics, and it's a completely weird theory that actually nobody really understands.
And there's a very interesting story which is maybe good to tell to connect sort of my PSD back to where I am now.
So I did my PSD with a Nobel laureate, Gerard de Tooft.
He's just the most brilliant man I've ever met.
He was never wrong about anything as long as I've seen him.
And now he says quantum mechanics is wrong, and he has a new theory of quantum mechanics.
Nobody understands what he's saying, even though what he's writing down is not mathematically very complex.
But he's trying to address this understandability, let's say, of quantum mechanics head on.
I find it very courageous.
And I'm completely fascinated by it.
So I'm also trying to think about okay, can I actually understand quantum mechanics in a more mundane way, sort of, you know, without all the weird multiverses and collapses and stuff like that?
So the physics has always been the threat and I'm trying to apply the physics to the machine learning to build better algorithms.
You are still very involved in understanding physics and the world, even beyond just applications to machine learning or introducing new formalisms.
That's really cool.
Yes, I would say I'm not contributing much to physics, but I'm contributing to the interface between physics and science and it's called AI for Science or Science for AI.
It's actually a new discipline that's emerging.
And it's not just emerging, it's exploding, I would say.
That's the better term.
Because now you go from investments into hundreds of millions, now into billions.
So there's now actually a startup by Jeff Bezos, that 6.2 billion sheep round, right?
It's like insane.
I guess it's the largest startup ever, I think, right?
And that's in this field, AI for science, right?
It tells you something that we are creating a new bubble here.
So why do you think it is?
What has changed that has motivated people to start working on AI for science-type problems?
So there's two reasons, actually.
One is that people have been applying the new tools from AI to the sciences, which is quite natural.
I think there's two big examples.
Protein folding is a big one.
And the other one is Machine Learning Forest Fields, or sometimes called Machine Learning Interatomic Potentials.
Both of them have been actually very successful.
Both also had something to do with symmetries, which is also cool.
And sort of People in the AI sciences saw an opportunity to apply the tools that they had developed beyond advertised placement or multimedia applications into something that could actually make a very positive impact in society, like health, drug development, materials for the energy transition, carbon capture.
These are all really cool, you know, impactful applications.
Beside that, the science and the kind of the is also very interesting, sort of the, I would say.
The fact that these two fields are coming together and that we're now at the point that we can actually model these things effectively and move the needle on some of these signs, uh sort of uh methodologies is also a very unique moment, i would say, and people recognize that.
Okay, now some we're at the cusp of something new where it results whether, as we're also what the company is called after we're at the cusp of something new, and of course, that always creates a lot of energy it's like okay, there's something.
It's like sort of virgin field right, it's like nobody's green field, nobody's been there.
You know, i can rush in and i can sort of start harvesting there, right and uh, and i think that's also what's causing a lot of uh sort of enthusiasm in the fields.
If you're an AI engineer, be of the people that listen to this podcast a week and you maybe don't have a strong science background powders, but are excited.
Most, I would say most AI practitioners, be it engineers or scientists, would consider themselves scientists.
And they have some background a little bit of physics, a little bit of industry college, maybe even graduate school that have been working or are starting out.
How does somebody who is not a scientist on a day-to-day basis, how do they get involved?
Well, they can read my book once it's out.
But this is basically... saying that there is more we should create curricula that are on this interface so i'm not sure there is possibly already some universities actual courses you can take maybe online courses you can take these workshops where we are now are actually very good as well and we should probably have more tutorials before the workshop starts actually we've I've kind of proposed this at some point.
It's like maybe first have an hour of a tutorial so that people can get new into the field.
But yeah, there's a lot out there.
Most of it is, of course inaccessible, but I would say We will create much more books and other content that is more accessible, including this podcast.
I would say right.
So I think, you know, it will come.
And, you know, these days you can watch videos and things.
There's a huge amount of content you can go and see.
So maybe a follow-up to that how do people learn and get involved?
But why should they get involved?
I mean, we have a lot of people who our audience will be interested in AI engineering, but they may be looking for bigger impacts in the world.
What opportunities does AI for science provide them to make an impact to?
You know, change the world that working in this, the world of pure bits, would not.
So so my view is that um, underlying almost everything is a material.
So we're focusing a lot on llms now yeah, which is kind of the software layer.
But I would say, if you think very hard, underlying everything is a material.
So underlying an LLM is a GPU, and underlying a GPU is a wafer on which we will have to deposit materials.
Do we want to wait a little bit?
Underlying everything is a material.
So i was saying, you know there's the llm underlying the elements, the gpu on which it runs, and then in order to make that gpu uh, you have to put materials down on a wafer and sort of shine on it with a sort of eov light in order to etch kind of the structures in.
But that's now an actual material problem because more or less we've reached the limits of, you know, scaling things down and now we are trying to improve further by new materials.
So that's the fundamental materials problem.
We need to get through the energy transition fast if we don't want to kind of mess up this world.
And so there is, for instance, batteries.
That's a complete materials problem, right?
There's fuel cells.
There are solar panels, so that they can now make solar panels with new perovskite layers on top of the silicon layers that can capture, you know theoretically, up to 50 of the light.
Where now we're at, I don't know, maybe 22 or something right.
So these are huge changes all by material innovation.
And yeah, I think, wherever you go, you know I can probably dig deep enough and then tell you well, actually the very foundation of what you're doing is a material problem.
And so I think it's just very nice to work on this very, very foundation.
And also because I think This is maybe also something that's happening now, is we can start to search through this material space.
This has never been the case, right?
It's like scientists.
The normal way of working is you read papers and then you come up with an hypothesis, you do an experiment and you learn, et cetera.
So there's a very slow process.
Now we can treat this as a search engine.
Like we search the internet, we now search the space of all possible molecules, not just the ones that people have made or that they're in the universe, but all of them.
Right?
And we can make this kind of fully automated.
That's the hope, right?
We can just type.
It becomes a tool where you type what you want and something starts spinning and some experiments get going right and then you know outcome, a list of materials, and then you look at it, say maybe not.
And then you refine your query a little bit yeah, and you kind of do research with this search engine where a huge amount of computation is and and experimentation is happening, you know somewhere far away in some lab or some data center or something like this.
I find this a very, very promising view of how we can sort of, come you know, build a much better sort of materials layer underneath almost everything.
And also more sustainable materials.
Our plastics are, polluting the planet.
If you can come up with a plastic that kind of destroys itself.
You know, after I don't know a few weeks, right?
And actually becomes a fertilizer.
These are things that are not impossible at all.
These things can be done, right?
And we should do it.
Can you tell us what a little bit just generally about Kaspi AI and then I have a ton of questions.
Yeah.
So Kaspi AI started about 20 months ago and it was because I was worried about I'm still worried about climate change.
So I realized that in order to get to stay within two degrees, let's say, we would not only have to reduce our emissions to zero by 2050, but then another half century, or even a century, of removing carbon dioxide from the atmosphere, not by reducing your emissions, but actually removing it at a rate that's about half the rate that we now emit it.
And that is a unsolved problem.
And if we don't solve it, two degrees is not gonna happen, right?
It's gonna be much more.
And I don't think people quite understand how bad that can be, like four degrees, like very bad.
So this technology needs to be developed.
And so this was my and my co-founder, Chad Edwards, motivation to start this startup, and also because, you know, we saw the technology was ready, which is also very good, you know the time is right to do it and uh, yeah.
So we now, in in the meanwhile, we've grown to about 40 people, we've kind of collected 130 million investment uh, into the company, which is for a european company, is quite a lot.
I would say It's interesting that right after that, you know, other startups got even more.
So that's kind of tells you how fast this is growing.
But yeah, we are now at the.
So we've built the platform, but it's for a series of material classes and it needs to be constantly expanded to new material classes and it can be more automated because you know we're not putting LLMs in as the whole thing gets more and more automated and now we're moving to sort of high throughput experimentation, so connecting the actual platform, which is computational, to the experiments, so that you can also get fast feedback from experiments.
And I kind of think of experiments as something you do at the end, although that's what we've been doing so far.
I want to think of it as what I would call a sort of a physics processing unit, like a PPU, right?
Which is you have digital processing units and then you have physics processing units.
So it's basically nature doing computations for you.
It's the fastest computer known, possible even.
It's a bit hard to program because you have to do all these experiments.
It's also quite bulky.
It's like a very large sort of thing you have to do.
But in a way, it is a computation, and that's the way I want to see it.
So you can do computations in a data center, and then you can ask nature to do some computations.
Your interface with nature is a bit more complicated, but then these things will have to seamlessly work together to get to a new material that you're interested in.
And that's the vision we have.
We don't say, super intelligence because I don't quite know what it means.
And I don't want to oversell it, but I do want to automate this process and give a very powerful tool in the hands of the chemists and the material scientists.
That actually brings up a question I wanted to ask you.
First of all, can you talk about your platform to whatever degree, explain how it works and what your thought process was in developing it?
Yeah.
Actually, it's been surprising.
It's not rocket science, I would say.
It's not rocket science in the sense of the design.
Basically, the design that I wrote down at the very beginning is still more or less the design, although you add things.
I wasn't thinking very much about multi-scale models and I've come on our radar that actually multi-scale is very important.
In the beginning I wasn't thinking very much about self-driving labs, but now I think you know we are now at the stage we should be adding that.
And so there is sort of bits and details that we're adding.
But more or less it's what you see in the slide decks here as well, which is there's a generative component that you have to train to generate candidates.
And then there is a digital twin multi-scale, multi-fidelity digital twin, which you walk through the steps of the ladder.
You know they do the cheap things.
First you weed out everything that's obviously unuseful and then you go to more and more expensive things later and so you narrow things down to a small number.
Those go into an experiment, you know, do the experiment, get feedback, etc.
Now, things that also have been more recently added is uh sort of more agentic, uh sort of parts.
You know we have agents that search the literature and come up with, you know actually, the chemical literature and come up with, you know, chemical suggestions for doing experiments.
We have agents which sort of autonomously orchestrate all of the computations and the experiments that need to be done.
You know they're in various stages of maturity and they can be continuously improved, I would say.
And so that's basically.
I don't think that part is rocket science, but you know, the design of that thing is not like surprising.
What is, it's surprising hard to actually build it, right?
So that's the thing.
That is where the moat is in the data that you can get your hands on and actually building the platform.
And I would say there's two people in particular I want to call out, which is Felix Hanke, who is actually building the scientific part of the platform, and Alessandro De Maria, who is building the MLOps part of the platform.
And recently we also added Aaron Walsh to our team, who is a very accomplished scientist from Imperial College.
We're very happy about that.
He's going to be our Chief Science Officer.
And we also have a partnerships team that sort of seeks out all the customers, because I think this is one thing I find very important.
It's so complex to actually bring a material to the real world that you must do this in collaboration with the domain experts, which are the companies typically.
So we only start to invest in a direction if we find a good industrial partner to go on that journey with us.
It makes a lot of sense.
Over the evolution of the platform.
Did you find that human intervention human?
I guess you could start out with a pure
You could imagine two directions.
One, you start out making everything purely automatic, automated, agentic, so on.
And then later on, you find that you need to have more human input and feedback, different steps.
Or maybe did you start out with having human feedback, lots of steps, and then kind of figure out ways to remove.
That's it.
It's the second one.
So you build tools.
So it's much more modular than you think.
But it's like, we need these tools for this application.
We need these tools.
So you build all these tools, and then you go through a workflow.
Actually, in the beginning, just manually.
So you put them, first this tool, then run this tool, then run this one, et cetera.
So you put them in a workflow.
And then you figure out oh actually, you know, this porous material that we're trying to make actually collapses if you shake it a bit.
Okay, then you add a new tool that says test for stability, right?
And so there's more and more tools.
And then you build the agent, which could be a Bayesian optimizer, or it could be an actual LLM, you know, maybe trained to be a good chemist that will then start to use all these tools in the right way, in the right order.
But in the beginning it's like you, as a chemist, are putting the workflow together and then you think about okay, how am I going to automate this?
One very easy question you can ask yourself is every time somebody who is not a super expert in DFT and he wants to do a calculation has to go to somebody who knows DFT?
And so could you start to automate that away, which is like okay, make it so user friendly so that you actually do the right DFT for the right problem and for the right length of time and you can actually assess whether it's a good outcome, etc.
So you start to automate smaller, small pieces and more bigger pieces, etc.
And in the end, the whole thing is automated.
So your philosophy is you want to provide a set of specific tools that make it so that the scientists making decisions are better informed, and less so trying to create an automated process.
I think it's sort of the same what you're saying, because yes, we want to automate, but we don't see something very soon where the chemists and the domain expert is out of the loop.
But it's a retreat, right?
It's like okay, so first you needed an expert to tell you precisely how to set the parameters of the EFT calculation.
Okay, maybe we can... take that out, we can maybe automate it, right?
And so increasingly more of these things are going to be removed.
In the end, the vision is it will be a search engine where somebody a chemist will type things and we'll get list candidates, but the chemist will still decide what is a good material and what is not a good material out of that list, right?
And so the vision of a completely dark lab where you can close the door and you and you just say just, you know, find something interesting, and then it will.
It will just figure out what's interesting and we'll figure out.
You know.
It's like oh, I found this new material to blah blah blah, blah.
Right.
That's not the vision I have, at least not for a long time.
So for me it's really empowering the domain experts that are sitting in the companies and in the universities to be much faster in developing their materials.
And I should say, it's also good to be a little humble at times.
Because it is very complicated to make it and to bring it into the real world.
And there are people that are doing this for their entire lives.
And it's like I wonder if they scratch their head and say well, how are you going to completely automate that away in the next five years?
I don't think that's going to happen at all.
Yeah, so to me, it's an increasingly powerful tool in the hands of the chemists.
I have a question.
You've talked before about getting people interested based on having, you know, sort of a big breakthrough in materials, just incremental change.
I'm curious what you think about the platform you have now in our sort of stepping towards and how are you chasing the big change?
Or is this like incremental or is there?
They're not mutually exclusive obviously but yeah, what do you think about that?
We follow a mixed strategy, so we are definitely going after a big material.
Again, we do this with a partner.
I'm not going to disclose precisely what it is, but we have our own kind of long-term goal.
You call a lighthouse or you know uh sort of moonshot or whatever but um, it is going to be a really impactful material that we want to develop as a proof point that it can be done and that it will make it into the, into the real world, and that ai was essential in actually making it happen.
At the same time, we also are quite happy to work with companies that have more modest goals.
Like I would say.
One is a very deep partnership where you go on a journey with a company and that's a long-term commitment together.
The other one is, like somebody says, i knew i need a force field.
Can you help me train this force field and then maybe analyze this particular problem for me and i'll pay you a bunch of money for for that, and then maybe after that we'll see.
And that's fine too right, but we prefer, you know, the deep partnerships where we can really change something for the good Yeah.
And do you feel like from a platform standpoint, you're ready for that?
Or what are the things that, and again, not asking you to disclose proprietary secret sauce, but what are the things, generally speaking, that need to happen from where we are to where to get those big breakthroughs I got.
What I find interesting about this field is that every time you build something, it's actually immediately useful, right?
And so, unlike quantum computing or nuclear fusion, so you work for I don't know 20 30, 40 years and nothing nothing nothing, nothing.
And then it has to happen, right?
And when it happens, it's huge.
So it's quite different here, because every time you introduce so you go to a customer and you say so, what do you need?
So we work, let's say, on a problem like water filtration.
We want to remove PFAS from water, right?
So we do this with the company Camira.
So they are a deep partner for us, right?
So we own a journey together.
I think that the breakthrough will happen with a lot of human in the loop, because there is the chemists who have a whole lot more knowledge of their field, and it's us who will help them with AI training and new methods.
And in that kind of interfaces, interactions, something beautiful will happen.
And that will have to happen first. before this field will really take off, I think.
And so in the sense that it's not a bubble, let's put it that way.
So as people see, that's the actual real that's happening.
So in the beginning, it will be very, you know, with a lot of humans in the loop, I would say.
And I would hope we will have this new sort of breakthrough material before know everything is completely automated, because that will take a while and also it is very vertical specific.
So it's like completely automating something for problem a.
You know you can probably achieve it, But then you'll sort of have to start over again for problem B because your experimental setup looks very different.
The machines that you characterize your materials look very different.
Even the models in your platform will have to be retrained and fine-tuned to the new class.
So every time you have a lot of learnings to transfer, but also you know the problems are actually different.
And so yes, I would want that breakthrough material before it's completely automated, which I think is kind of a long-term vision.
And I would say, every time you move to something new you'll have to start retraining and humans will have to come in again and sort of okay.
So what does this problem look like?
And now sort of, you know, point the machine again in the new direction and then use it again.
For the non-scientists among us, me included, a bit of a scientist.
There's a lot of terminology.
You mentioned DFT, equivariance we've talked about.
Can you sort of explain in engineering terms, or at the level of sophistication in engineering, what is equivariance?
So every variance is the infusion of symmetry in neural networks.
So if I build a neural network, let's say, that needs to recognize this bottle right, and then I rotate the bottle, it will then actually have to completely start again, because it has no idea that the rotated bottle Well, actually the input that represents a rotated bottle is actually a rotated bottle.
It just doesn't understand that.
Where, if you build equivariance in basically, once you've trained it in one orientation, it will understand it in any other orientation.
So that means you need a lot less data to train these models.
And these are constraints on the weights of the model.
So basically, you have to constrain the weights such that it understands it.
And you can build it in.
You can hard code it in.
And yeah, the symmetry groups can be translations, rotations, but also permutations.
Like in graph neural network, there are permutations.
And in physics, of course, there's many more of these groups to pray devil's advocate.
Why not just use data augmentation by your model, as in all the different orientations?
As an option, it's just not exact.
It's like why would you go through the work of doing all that where you would really need an infinite number of augmentations to get it completely right?
Um, where you can also hard code it in.
Now, i have to say, sometimes actually data augmentation works even better than hard coding the equal variance in, and this is something to do with the fact that if you constrain the optimization weights before the optimization starts, the optimization surface or the objective becomes more complicated and so it's harder to find good minima.
So there is also a complicated interplay, i think, between the optimization process and and these constraints you put in your network and so Yeah, you'll hear kind of contradicting claims in this field.
Like some people, and for certain applications, it works just better than not doing it.
And sometimes you hear other people if you have a lot of data and you can do data augmentation, then actually it's easier to optimize them.
And it actually works better than putting the aggregates in.
Do you think there's kind of a bitter lesson for mathematically founded models and strategies for doing deep learning?
Yeah, ultimately it's a trade-off between data and inductive bias.
So if your inductive bias is not perfectly correct, you have to be careful, because you put a ceiling to what you can do.
But if you know the symmetry is there, it's hard to imagine.
There isn't a way to actually leverage it.
But yeah, so there is a bitter lesson.
And one of the bitter lessons is you should always make sure your architecture scale, unless you have a tiny data set, in which case it doesn't matter.
But if you you know, The same bitter lessons or lessons that you can draw in LLM space are eventually going to be true in this space as well, I think.
Can you talk a little bit about your upcoming book and tell the listeners what's exciting about it?
Yeah, they should read it.
So this book is called Generative AI and Stochastic Thermodynamics.
It basically lays bare the fact that the mathematics that goes into both generative AI, which is the technology to generate images and videos, and this field of non-equilibrium statistical mechanics, which is systems of molecules that are just, you know, moving around and you know relaxing to the ground state or that you can control to have certain you know be in a certain state.
The mathematics of these two is actually identical.
And so that's fascinating.
And in fact what's interesting is that Jeff Hinton and Radford Neal already wrote down the variational free energy for machine learning a long time ago.
And there's also Carl Friston's work on free energy principle and active inference.
But now we've related it to this very new field in physics which is called stochastic thermodynamics or non-equilibrium thermodynamics, which has its own very interesting theorems, like fluctuation theorems, which we don't typically talk about but we can learn a lot from.
And I think it's just, it can sort of now start to cross-fertilize.
When we see that these things are actually the same, we can, like we did for symmetries, we can now look at this new theory that's out there, developed by these very smart physicists, and say OK, what can we take from here that will make our algorithms better?
At the same time, we can use our models to now help the scientists do better science, right?
And so it becomes a beautiful cross-fertilization between these two fields.
The book is rather technical, I would say.
It takes all sorts of things that have been done as to stochastic thermodynamics and all sorts of models that have been done in in the machine learning literature and it basically equates them to each other and i think hopefully, that sense of unification will be revealing to people.
Wait, and when is it out?
Well, it depends on the publisher now, but uh, i i hope in april I'm going to give a keynote at iClear and it would be very nice if I have this book in my hand.
But it's hard to control these kind of timelines.
I'm looking forward to it.
Great.
Thank you very much.