Noah Kravitz at GTC 2024 in San Jose, California.
And my guest for this episode is Adam Wenchel, co-founder and CEO of Arthur AI.
Arthur calls itself the AI Performance Company, and they work with enterprise teams to monitor, measure, and improve machine learning models for better results across accuracy, explainability, and fairness.
Adam's here to talk about Arthur's mission, and we're gonna get into bias and observability in AI, guardrails on generative AI systems, and more generally about the adoption of generative AI in the enterprise.
So let's jump into it. Adam, thanks for taking time out of your DTC experience to join the AI podcast.
Yeah, thanks, Neil. I appreciate you having me on.
So we were talking offline. You just got in.
I was going to ask you how GTC has been for you so far.
Did you get a chance to check anything out yet or not yet?
Yeah, I did. I went to the session earlier today on the head of research, you know, talking about some of the stuff they're working on, which is always fun.
And so it's good to, you know, I started my career as an actual hands on dev back in the day.
So every once in a while, I like to pretend like I still, you know, follow the tech and get to enjoy some technical content.
Walk around, pick up some buzzwords, not knowingly.
Excellent. Love it. Well, let's talk about the hair now.
Let's talk about Arthur. Maybe you can just start by telling the audience what Arthur AI is all about.
Yeah, absolutely. At Arthur, what our goal is, is to help people translate all the promise of AI into the real world.
We've all played around with ChatGPT and Anthropic and Gemini and a bunch of other open source models. and really exciting to use and play around with, very compelling.
But when you go to deploy systems based on these elements in the real world, You quickly discover there's some risks and some challenges that need to be managed.
Sure. hallucinations it can be leaking sensitive data it can be inappropriate use of the lms there's a number of uh things and just flat out performance right are they doing a good job answering the questions Are they value aligned with your organization?
Things like that. And so that's where we help with.
We started out as providing monitoring tools for models and AI, and then we've since expanded into that Shield, which is our firewall for AI. which really allows you to apply policies around the usage of these things and also just make sure that they're performing well.
So what kinds of use cases are you seeing out there in the real world?
You know, what do companies, what enterprises want to use LLMs and Gen AI for?
Yeah, I mean, there's obviously been a lot, but I think a lot of established businesses are really...
They're starting with internal use cases just because they want to really make sure they understand the technology and know how to deploy it responsibly before they put it in front of their customers and their partners. and things like that.
We won't name names. We've seen some chatbots go awry publicly.
There's no question like there's a lot of value to be created by having a public facing.
But what we're seeing like first deployments are a lot more around internal use cases around things like HR, for instance, is a surprisingly popular one.
It turns out if you're a large company, answering questions about benefits is actually a really time-consuming task.
It is. Yeah. And so there's like huge productivity gains.
And also a lot of times you get better answers that way.
And so it increases kind of satisfaction among your team.
And so it's things like that, these back office tasks that are not glamorous, but actually are quite impactful that people are using.
And, you know, you're seeing a lot like legal and investment, a lot of private equity and and investment banks and for a lot of those, the price of an incorrect answer is quite high.
If you give someone bad information about their benefits or if you give them Uh, you know, or, or if you make an investment based on bad advice, you know, even if it's right, 95% of the time, 5% is not great.
Yeah, no, right, right. You know, it's easy to kind of get caught up in the tech world bubble, so to speak, right?
And I've been, you know, the minute a new update comes out, I'm going online trying to see where I can play with it and everything.
But have you been experiencing working with customer companies who – have heard about AI, have heard about LLMs, that sort of thing.
But then when you get into actually implementing, like it's all brand new to them or what's kind of the, What's kind of the delta between what's going on at some place like GTC where we're at and then where sort of the rest of the world is with using this stuff?
Yeah, good question. I mean, certainly the audience at GTC is on the tip of the spear of this stuff.
But you know what I'd say is like a lot of even a lot of traditional organizations for the last four or five years have been trying to build up their capability around AI and adding talent and putting in place infrastructure, getting data ready, things like that.
And it was kind of a very slow rise up the maturity curve, I would say, for the last four or five years.
But generative AI has definitely caused an inflection in terms of the adoption because for a few reasons.
One, I think part of it is all that investment from the last four or five years prepared people for it.
But also because this has become a board member or CEO can easily go on to ChatGPT and like, You know, it's so accessible and it's really easy to make the mental leap from, you know, you asking it to plan your weekend in San Jose to from that to like having it. you know, plan trips for work or for answering HR questions and things like that.
And so everyone sees the possibility. And so there's a real sense of urgency where it is a board-level issue, like what is our generative AI strategy?
And there's some really strong imperatives that are getting handed down, like by...
Next quarter, you'll have deployed generative applications.
Right. And is the motivation for that...
Like, are there specific motivations? Is it, you know, in the name of efficiency or is it more of a...
Everybody sees where, you know, the puck is headed and so they're skating in that direction kind of thing.
Yeah, a lot of it is efficiency, but not necessarily to cut costs.
I think the idea is if you can be more efficient, you can allocate more resources to more strategic work.
If you can remove some of the tedium and some of the more rote work, it frees up those resources to take on more strategic work.
We haven't seen any evidence of of like job loss or anything like that, or even like significant the cost savings all gets kind of reapplied.
Right. You know, going after more top line growth.
Right. So in terms of the things that can go wrong and that, you know, with good reason As you mentioned, companies are kind of deploying the stuff internally first so that the mistakes are a little easier to contain, that kind of thing.
Hallucinations is obviously the thing that a lot of people have heard about. you know, an LLM making something up, passing it off confidently.
Is that at the top of the list for a lot of your companies and the things that they're worried about or what other things are there?
Yeah, I mean, hallucinations are a really universal one.
I mean, wrong answers are bad for everyone.
I think that if you have an online cocktail recipe generator, you can probably live with a bad drink.
But if you're in any sort of business application, then wrong answers are, you know, whether you're giving like the HR application or investment advice or legal advice, you can't live with wrong answers.
But there's a bunch beyond that, you know, prompt injection, inappropriate use of So if you're using your HR system to try to generate legal briefs, that's not good.
It's not good for a number of reasons. And in things like toxicity, you know, I think a lot of the universal notions of toxicity, the LLM makers have actually done a really good job of RLHFing kind of.
Some of the really obvious behaviors they picked up off the text, the internet text they were trained off of on.
But, you know, there's a lot of times use case sort of specific inappropriate language.
And so an example would be, If you have an LLM agent you're using in the recruiting process or the hiring process, you shouldn't be talking to it about like family status or it should not be telling you about family status.
Right, right. I think there's use case specific policies that you need to be able to apply to make sure that You know, you're not being reckless with the way you deploy it.
Sure. And so kind of from more of a technical standpoint or sort of a— pragmatic standpoint, I guess.
Are you fine tuning, you know, models that people would have heard of and fine tuning them for specific deployment?
Are you training new models? You know, what's kind of the adoption process that these companies are using?
Yeah, absolutely. I think people typically start with an off-the-shelf model because they're so easy to adopt, especially with some of the ones that are provided by the big names that provide them via API access.
But we are seeing a lot of people train and fine-tune models for their purposes because you can get Really, especially if your task is not a GP, not a general purpose, but is a little more constrained.
You can fine tune a smaller like 7B model or even smaller on a particular domain. and you can have it give really good answers much faster and much more efficient use of GPUs.
And so we're seeing a lot more of that. We actually use that technique internally.
Some of the ways people can apply policies is actually using LLMs to evaluate the outputs of other LLMs.
Right. But you want to make sure that's efficient.
You don't want to add latency to, there's already some latency, you don't want to add more to it, things like that.
We use that technique ourselves and it's pretty effective.
And I think people are really just starting to get good at it.
Yeah. Let's talk about observable AI. What does that term mean?
And sort of in practical terms, what does it mean?
Yeah, so my previous role before starting Arthur, I led the AI team at a large top 10 U.S. bank and If you're deploying AI in any mission-critical sensitive application, like deciding who gets a credit card and how much credit they get, then you need to be able to tell people, like, why were different people turned down?
And you need to also make sure that they're making good decisions.
And so the models are making good decisions.
You know, in the old days, there's all sorts of stories about people deploying models and they used to monitor them or observe them.
It was a very manual process, like once a quarter, some data scientists would go kind of check in. download all the data to their laptop, put it in a spreadsheet or a notebook and check it out.
Nowadays, these models are so dynamic and the rate of change in the world is often so dynamic that You can't just sort of set it and forget it with these models.
You need to bring that same kind of real-time intelligence to the monitoring and oversight that you have for the model itself.
These aren't just you know, the simple linear regression models from 20 years ago.
Yeah. Yeah. How reliable... is kind of the automated observation process.
Is it something you're happy with? Is it something that's kind of a, Work in progress?
Is it something that's kind of a hurdle that you're working to get over?
Like, what's the seat of the art with observation?
The observability is actually quite good.
That part's quite mature. And I think that the maturity around a lot of... Obviously, in the last year, all the stuff around... hallucination prevention and detection.
There's a whole new suite of metrics that the industry's developed and we've helped driven some of that development around. which is like measuring things like readability or helpfulness.
How do you kind of measure generated texts for the equivalent of accuracy.
There's a number of different axes you want to be able to evaluate. your performance on.
And so like the notion of like, what is good performance?
I think people have like, you know, we can go read stuff manually and sort of say whether it's good or bad, but how How do you break that down into something that's measurable at scale?
So if you're asking, you know, if you're answering like, hundreds of thousands of questions or serving hundreds of thousands of users.
How do you really look at that at a large scale and make sure it's performing well?
It's fascinating topic and something we like to talk about.
Yeah. I don't know if this is actually a logical train of thought or not, but it makes me think about bias in AI systems.
Can you talk a little bit about that? It's a term that gets thrown around a lot, and I tend to think of it, or maybe I've red things in the media that make me think of it in terms of demographics, right?
Racial bias, ethnic bias, gender bias, that sort of thing.
What does bias in an AI system mean to you?
And I think kind of more interestingly, maybe, How do you deal with bias without compromising the accuracy or introducing hallucinations or otherwise kind of messing up the AI systems?
Yeah, it's a good question. So the bias that's in systems a lot of times is, you know, at some level, like the credit card example, right, or auto loaned, you know, people used to kind of manually, like you'd fill out an application in front of someone at the auto dealership and then they would kind of It was a very subjective decision about whether to give you an auto loan.
And there was certainly a lot of bias in those decisions.
Sure. And then what happens is they took that data, like a decade of that data, and then train models on, right?
And so all it does is automate the bias that was already there.
Yeah, you train a model on the internet, you're going to get some bad words.
You get some crazy stuff, yeah, absolutely.
And so a lot of it is being able to measure that.
And so in traditional AI, there's dozens of metrics around fairness and you can't solve for them.
Actually, a lot of them are like statistics. there's like tension between them.
But, um, roughly speaking there, there's kind of a couple big, there's like two big categories.
One is, um, Are they having equal access to the outcomes?
And so are they getting approved for auto loans as frequently or in the same amount?
And then the second family is a lot around accuracy, right?
Like, are you making really accurate decisions for males, but like females, you're just throwing darts at a dart? board and your model doesn't really know what it's doing.
So maybe they're getting auto loans at the same rate, but you're having a lot more repossessions with one group than the other because your model's not accurate there.
And so that was traditional AI, and that science is... relatively well understood, although still, I think a lot of people are still learning how to kind of do it in the real but certainly from an academic perspective, I think it's fairly far along and we've published a lot of papers in that space.
In LLMs, it's a little bit different. It's another area where people are just learning how to measure it.
I will say that when generative AI first exploded towards the end of 2022, the first models that were made available were horrendously biased and it was like trivial to...
Yeah, to get them to say things that no organization would want a piece of technology they're putting out there to say.
And I would say that the model makers in general have gotten – really good at softening the most egregious parts of that through, you know, RLHF and constitutional AI and things like that, policies.
And so a lot of good work has been done there in the last year, and they're much less biased than they used to be.
And so now the forms of bias still exist, but it's a lot more subtle.
And so I think there's some really interesting research on how to, you know, make sure that the answers. you know, answers about questions, whether they're given in the context of a male or a female or people of different racial groups pretty uh even and equitable and represent values aligned with the organization publishing them I'm speaking with Adam Wenchel.
Adam is the co-founder and CEO of Arthur AI, and we've been talking... all things generative AI when it comes to deploying this technology out in the business world, in the real world, so to speak, across different industries. and lines of work.
You mentioned your own background working at a bank.
And I wanted to ask you kind of how you got to the point you're at now with founding Arthur.
If I've got this right, You started a cybersecurity company at some point back in the day.
Can you tell us a little bit about that, how you leveraged machine learning, and then maybe how your path kind of progressed to here?
Yeah. So actually, I studied AI in school at the University of Maryland, which has a really good AI program. but over 20 years ago.
And at that time it was not really in style.
It was actually like mostly, you know, people would say jokes like, oh, you're in, AI, you're great at predicting the past and things like that.
But then I followed one of my professors over to DARPA and worked on a couple of really early AI projects there, including one that was all about using agents to plan, so it's been fun to see in the last few years.
A lot of that AI stuff coming back and it's closer to working now, a lot closer to working now than it was back then.
But yeah, I then went from DARPA into the startup world, and I've definitely... kind of been hooked by it.
It's, you know, just the creativity of like starting a new company and building it is just really fun.
And like you said, my My last company was using machine learning for cybersecurity, which nowadays, like every cybersecurity company says they're using AI or machine learning.
But this was back in 2013, and it was a little more novel. back then.
Right, right, right. And then that one was acquired by Capital One, actually.
So I joined Capital One And they asked me to start their AI team shortly after I joined, which I did, which was a lot of fun.
Yeah, I bet. Are there learnings from the DARPA days and the original cybersecurity company you mentioned You know, are there learnings from those days when not that you were doing theoretical work only, obviously, but, you know, the compute wasn't there.
Things weren't as advanced as they are now.
Kind of things that maybe you wanted to do or could only do to a certain extent that you're now able to kind of bring back and more more fully flesh out or has time just kind of moved on and it's new stuff?
No, I mean, certainly. from the DARPA days, the compute definitely wasn't there and we weren't using GPUs.
And you didn't have cloud computing, so you couldn't just spin up 100 boxes to train a model on.
You had much smaller data sets, which was holding it back. also a lot of algorithmic improvements that have been made in the intervening 20 years.
And so back then, a lot of the conversation, actually, you know, learners, which are what's in vogue currently, a little bit out of vogue back then.
I actually liked him, but more people were, there was a lot more energy going into symbolic AI. which works, you know, it can be very brittle, but it's more compute efficient usually.
And this is all generalities, but often.
And so that was like, kind of where a lot of the energy in the research community was going.
And symbolic AI, actually, at times it kind of sticks its head up every now and again to kind of address some of the limitations for learners, but we'll see.
Thinking about cybersecurity and, you know, LLMs and generative AI, and when you were talking about some of Arthur's company or client companies, deploying internally first with good reason, it made me think of stories I've read about companies not allowing—I mean, especially when chat GPT kind of first exploded into the consciousness—
You know, schools banned it outright, right?
Lots of giant school districts. And, you know, companies would... hear stories once in a while about, you know, some big tech company not letting their employees use it or spinning up an internal model for employees to use.
And a lot of The rationale, at least that I heard, was, you know, fear of data leaks. fear of you know well if I'm putting proprietary information up into one of these publicly available models I don't know what's going to happen with it like best case it gets used to train their next model which isn't really cool with us yeah worst case you know it somehow shows up in somebody else's chat which happened Right, right.
Absolutely. So with your cybersecurity background, what's your take on...
Gen AI, LLMs, everything that's happening now, and the cybersecurity implications.
Yeah, so, you know, when the first batch of models were stood up and kind of made available, people started adopting.
Like, all the things you mentioned weren't just theoretical concerns.
Yeah, no, they happened. Right. And so, I think those companies have been rightly taken to task about that, and they've put a lot of energy into providing much stronger guarantees around that stuff.
But depending on the sensitivity of your information, the question is how much do you trust that?
And that's why it's nice that there's you have optionality, you can run your own open source model in your own environment.
If you're dealing with really sensitive data, like we have hedge funds that are very sensitive about their data and they prefer to do it.
I don't think it's not unfounded given what's happened.
And then there's also a whole set of concerns around if these models have been trained on data that is proprietary, that someone else owns the rights to, and I'm using it to generate like my strategy or my investment decisions on essentially like someone else's proprietary, a model that's been trained on someone else's proprietary data, what are the legal implications of that, right?
Do you know? Do you have a good answer? Well, no one knows right now.
And so it's the kind of thing where it's going to take years of case law to sort that all out.
A lot of lawyers being lawyers are gonna make sure that everyone's taking the risk very seriously.
Right, right, right. We'll see. I've heard some groups are offering indemnification to their best customers for that and things like that.
Right, right, right. You know, I think, I don't think you can put, you know, at this point the toothpaste is out of the tube.
I don't think it's going back in. Yeah, no, what's the phrase?
You know, do it now and ask for forgiveness later came to mind as I was listening to you.
There's been a lot of that going on last year.
Yeah, right, yeah. With all the different companies that Arthur works with and, you know, the different use cases, you mentioned HR and financial and these other things.
And this is a big question to ask you, so I'm putting you on the spot, but that's why we're here.
How do you see the future of work changing? kind of evolving and taking shape in this new, you know, I call it degenerative AI era for now, who knows if next year You know, a new breakthrough is going to give us a different paradigm.
But, you know, you mentioned companies... your customers kind of looking for these efficiency gains, but reinvesting the money, reinvesting the time to do more creative, more strategic work, that kind of thing, which.
Sounds like a best case scenario in a lot of ways, but I don't know, are there trends you're already seeing emerging beyond that? kind of indicating where, you know, work is headed in the AI age?
Yeah, you know, I saw a quote recently that was saying, you know, you're not going to lose your job to AI.
You're going to lose your job to someone who knows how to use AI.
Right, right. And that's what we've seen, right, is the people who have really taken the time to get good at incorporating ChatGPT or another LLM into their their daily workflow or like actually finding value and it's making them more efficient at what they're doing.
And you know, those people are definitely on the cutting edge, but, but, a couple of years from now, probably everyone's going to be using LLMs in some capacity and the sooner you scale that learning curve. probably the better it's going to be for your career and your job opportunity.
Sure. And so I definitely encourage people to play around with it and think about how you can just incorporate it in a day-to-day. basis.
Yeah. Yeah. I do a lot of writing and I find the models are great for kind of brainstorming first draft kind of stuff.
And it's, yeah, it's become part of my flow.
Writer's block is almost a thing in the past.
I did, I did a first draft of like my, you know, my, my year end reviews for my team.
And, um, you know, I went back and edited them heavily, but.
Usually when you sit down in front of that blank sheet of paper, it takes a little time to kind of get over that, right?
And so you just kind of skip that whole step.
Yeah, yeah. No, it's wild. multimodal models are becoming more and more of a thing and beyond the sort of obvious i i this phrase obvious, but mind blowing kind of came to mind because it's, well, obviously you can, you know, images and audio, whatever, but it's also sort of mind blowing the implications.
But from more of a, um, you know, again, pragmatic and the kind of work that you do and what, you know, enterprise customers of yours are thinking about from that standpoint.
Are there like new challenges, new concerns beyond, you know, wealth multiplication? types of data.
But are there new things that people are thinking about, worried about, excited about with multimodality?
Or is it all just kind of part of this huge you know, rush of Gen AI, wow.
Yeah, there is a lot to be excited about.
And it does, it increases the complexity a lot.
You're not just doing text-to-text. And we've seen kind of the first wave of like text to X, like text to video, text to all sorts of things. come out.
But I think a lot of the real excitement is going to come.
And those, I think, have been helpful for you know, a lot of like creatives and things like that.
But I think when it starts to be multimodal inputs as well, which is coming and people are actively doing, it's things like document understanding or even being able to walk the line of a factory and and you know make observations things like that are going to become possible And I think that's going to have a profound effect.
I think it's going to take a little longer to figure out than LLMs, just straight language models.
Absolutely. I mean, document understanding is still, there's been lots and lots of work over the last couple of decades on it and it's still not a solved problem.
So, yeah. Yeah. And so what's next for Arthur?
What are, what are you guys working on? Anything you want to talk about or, you know, what, what is the rest of the year and the next couple of years look like for you guys?
Yeah, I mean, we're just focusing on helping our customers adopt AI more quickly and get them up and running.
Because that's a win for them and a win for us.
I was going to say, as I was asking what's next, I'm like, you have a lot on your plate. but it's plenty.
We've launched a lot of new products last year.
And so we, and this year, you know, and they're resonating with customers, which is great.
But I also tell the team. Don't get too complacent because that rate of change hasn't stopped.
There's still a lot of change coming. we'll continue to evolve our platform along with that.
But I think it's exciting to see people getting Even traditional, you know, Fortune 100 companies are getting their first use cases into production.
And, you know, it's taken, you know, 10 months to go from the hype to the reality.
But actually, in terms of enterprise software adoption, enterprise is adopting brand new technologies.
That's actually pretty darn fast. Yeah, right, right.
Those are big ships. They turn slowly. Exactly.
Adam, for listeners who want to find out more... about Arthur AI, about any of the stuff we've talked about, research documents, some of your own views maybe, I don't know what, Where can they go?
Obviously you guys have a website. We do.
Yeah. Arthur.ai. And we've got, you know, a very active blog and a lot of our research goes there as well.
Perfect. Yeah. Check it out. Excellent.
Well, thanks so much for taking the time out of the show to stop by and chat.
This was great. You know, one of these conversations that you talk about this stuff, and then half an hour later, my brain kind of settles down.
I'm just like, wow, brave new world. It is.
It really is. Well, best of luck to you and to your team.
Yeah, thanks, Neil. I appreciate it. Thank you.