If you would like to earn CPE credit for listening to theshow, visit earmarkcpe.com slashFPA.
Download theapp, take a shortquiz, and get your CPEcertificate.
Finally, if you enjoy listening to FPNAtoday, please go to your podcast platform ofchoice, click the subscribebutton, and leave a rating and review of theshow.
Andnow, onto theshow. From DataReels, this is FPNAToday.
Welcome to FPNAToday. I'm yourhost, GlennHopper.
Today, we have the pleasure of speaking with AlexKim, a PhD student at the University ofChicago, and co-author of the super interesting new study from the University ofChicago, financial statement analysis with large languagemodels, which showed some results that might surprise a lot of ourlisteners.
We'll be sure and provide a link to the full paper in the shownotes.
Alex, welcome to theshow.
Thankyou,Glenn. Thank you so much for having mehere.
My name isAlex. As yousaid, I'm a PhD student at the University ofChicago.
Super excited about this topictoday, because this is a subject that's near and dear to myheart.
I actually teach courses on how to use generative AI to do financialanalysis.
So when I saw thispaper, I immediately knew I've got to get you on theshow.
So I really appreciate you comingon.
And I guess before we dive into thestudy, tell us alittle, I know you're a PhDstudent.
And the funny thingis, before wetalked, I had this expectation that you were going to be a PhD student in machinelearning.
But I was so excited to see what your actual field of studyis.
So tell me a little bit about yourbackground, your educationaljourney, and what you're currentlyresearching.
Yeah,sure. Onceagain, thank you so much for having mehere.
My name isAlex.Currently, a second year PhD student at the University ofChicago.
My main research area isaccounting.
I do research at the intersection of accounting andfinance, especially related toinvestors' informationprocessing.
I was born inracing, SouthKorea.
I did my bachelor's degree in economics and business administration from my homecountry, and also did my master's degree with a concentration in accounting at the sameinstitution.
So as you cansee, I actually do not have any formal education related to computerscience.
But what fascinated the most was this information processing part ofaccounting.
As youknow, accounting is a language that companies use to communicate themselves withstakeholders,investors, with many externalparties.
So what is really interesting here is that accounting information is processed by so many people outthere.
And then as technology'sevents, companies are communicating themselves with so many othermethods, such asimages,videos,texts,audio, like so many othermethods.
And then I became fascinated by how people process thisinformation, how this information can bequantified, how it affects capitalmarkets, how it affects people's economic and financialdecisions, and so on and soforth.
So it has been my research interest for more than six or sevenyears.
And then I just realized like five or six yearsago, when AI was not even a popularity or fat backthen, I just realized to understand how people process this kind of newdata, I have to learn how to code and how to process this new type ofinformation.
And then that was the beginning of how I started to learn about machinelearning,AI, imageprocessing, voiceprocessing, like videoprocessing, stuff likethat.
So pretty much everything wasself-talk.
I was like paying attentionto, youknow, like computer science lecturesonline.
I also published several papers and one of the top computer scienceconferences, with myfriends.
So I've been like learning bydoing.
And then this process has helped me a lot in my research and accounting andfinance.
So as Isaid, my primary research area iswhere, is the intersection between the accounting andfinance, and how people processinformation, and like understanding how people actually process information inside theirbrains, is critical to myresearch.
And that's where AI machine learning large language models are helping meout.
Yeah, and then here I am after finishing my master'sdegree, I started my doctoral career in the UnitedStates.
It's been two years ever since I came to the UnitedStates.
Yeah, I've been like doing my accountingstudies, and then I recently did my comprehensiveexams, looking forward to the research phase of my PhDcareer.
Yeah, that's a brief story aboutmyself.
I think about how much more information is availablenow.
Everything from the Twitter fire hose to earnings callsto, youknow, just all the information that's outthere, the internal companydata, plus the external data and how it weighsin.
And I feel like we're kindred spirits in a lot ofways, because I'm always looking to find thesecorrelations, to take these different pieces ofinformation.
I've been playing around a little bit with usingRAG, with graphdatabases, to try to findthese, youknow, not just the directcorrelations, but the nodes and allthat.
And I think it'sjust, youknow, you're looking for that correlation that gives you the sort of the leadingindicator.
If you can find something that's connected before the rest of the market realisesit, Imean, there's just obviously an investing and capitaldeployment.
That's a great thing to be able to try toidentify.
So the work you're doing really speaks tome.
Yeah, as well aslike, youknow, you talked about like practicalstandpoint, but I also do think that it has a lot of comparative advantage in terms of theoretical standpoint aswell.
I'm a social scienceresearcher, and I'm also very interested in like practical applications of myresearch.
But at the sametime, I do care about my theoretical and academic contributions in myfield.
Imean, there have been so manyconcepts, theoreticalconcepts.
That were not been able to be measured byresearchers, just because of the lack of technology or lack of newthings.
Imean, as yousaid, textinformation.
It issuper, super valuablerelevant.
It has been out there for more than40, 50years.
It's been comprising like major portion of the marketreactions, how people processinformation.
But onlyrecently, it has become available to analyse systematically suchinformation.
Imean, that's actually a breakthrough in terms of like theoretical standpoint aswell.
I'm soexcited, and I really want to dive straight into thispaper, because thisis, I'm sure you're getting a lot ofquestions, and this has gotten a lot of mediaattention.
And I'm sure thatpeople, youknow, podcast hosts and everybody else has reached out to you about thepaper.
And I was going to kind of tee up thepaper, but I don't want to steal yourthunder.
So tell me about this paper and what you guys found in yourresearch.
So it started from a very straightforwardmotivation.
So in the first round of research related to large language models in the field of finance andaccounting, like lastyear, manypapers, including mypapers, have been talking about what these language models cando.
And I think the general consensus of those paperswas, language models aregreat.
They can perform many tasks related to textualinformation.
And as awhole, they are anice, textual supportingtool.
They can summarizethings, they can extractinformation, and they can actually provide reports formanagers.
But we wanted to test the boundaries of these language models evenfurther.
One of the main weakpoints, caveats of these languagemodels, was that they were very weak in terms of quantitativeanalysis.
However, recent advances in large language models have enabled them to perform some basic quantitativetasks.
They can now domultiplications, like adding deductions or stuff likethat.
And then we wanted to test how far these language models can go in terms of financial and economic decisionmaking, and wanted to understand whether they are more than just merely a supportingtool, whether they can play a more central role in financial or economic decisionmaking.
So we tested this notion on one of the most foundational tasks in accounting andfinance, which is financial statementanalysis.
Financial statement analysis is a uniquetask, hybridtask, that combines quantitative and qualitativeanalysis.
So quantitative part is that analysts or like information processors have access to financial statementinformation, which is largely numericalinformation.
They perform some quantitativeanalysis, such as trendanalysis, racialanalysis, identify some notable changes and stuff likethat.
And then they provide some economic intuitions or insights based on the numbers that theyanalyze.
And then after getting some economic intuitions orinsights, they synthesize and combine all these pieces ofinformation.
And in the end of theday, they provide one piece of prediction or economic decisions by or by or sellrecommendations,earnings,predictions, or salespredictions, so on and soforth.
So as Isaid, that letter part of formingexpectations, formingguidance, forming their economicinsights, that letter part is actually very qualitative and requires multi-layer economicreasoning.
So we wanted to test whether LLMs can successfully perform this highly complextask, which combines quantitative and qualitativeaspects.
And then what we do in the paper is pretty much verystraightforward.
We pass standardized and anonymous financialstatements, balance sheet and incomestatement, to the model without any specific context or narrativeinformation.
And try to mimic how human analysts process information within theirbrain.
So we refer to a paper that was published more than 40 years ago about how human analysts process numericalinformation.
And then according to thepaper, human analysts process information by doing some ratio analysis such asprofitability,liquidity,efficiency,etc.
And then they also do some trend analysis or identify some notable changes based on financial statements and redo exactly thesame.
We adopt a technique called chain of thoughtprompting, which is actually a very famous technique in natural language processing or computer science studies to ourstudy.
So this is simply put very simple and straightforwardconcept.
It is actually providing the model step-by-stepguidelines.
And then those guidelines are very similar to how humans processinformation.
And then in this context of financial statementanalysis, we take the processes or steps introduced in thatpaper, convert them intoprompts, and then ask the model to process numerical data as human analysts processthem.
And then what we get is as the outcome variable to evaluate the quality of the financial statement analysis is the direction in the change ofearnings.
So that benchmark is widely used in a county and financialliterature.
It is a binary prediction whether earnings were likely to go up in the next period or likely to go down in the next period and try to see whether model's prediction is reasonable oraccurate.
And then what we find in the paper pretty much summarized is that GPD is on average better than human analysts in predicting the direction of the future earningschanges.
And the second is that GPD is on par with highly specific or narrowly trained machine learning models in terms of predicting the changes of futureearnings.
And then one important intuition is that machine learningmodels, GPD'spredictions, and humananalysts' predictions are not mutuallyexclusive.
They are complementary with each other at this moment and then that is actually one economic intuition or a main takeaway of thepaper.
So I think the part that probably got everyone reaching out to you about this paperis, youknow, it'salmost, it's like it's an easy clickbait headline that is everybody so scared about AI taking theirjobs.
So if you say LLMs outperform human analysts at predicting directions of earnings pershare, people freak out and they click the link and they want to learn more aboutit.
But I think Ireally, I like what you said there and this has been part of my message as well is it's not either or it's workingtogether.
Could you elaborate on that a littlebit?
Yeah,sure, ofcourse.Actually, I think that is one of the main selling points of thepaper.
I'd like tosay,quote-unquote, onaverage.
On average is everything in ourpaper.
So especially when we are comparing GPDs predictions with human analystpredictions, we are saying that GPDs predictions are on average better than human analystpredictions.
And this finding itself is not very surprisingbecause, starting from early 2010 or like2015, they're happy like machine learning papers arguing that machine learning predictions are actually better than human analystpredictions.
They say that is simply because a machine learning models do not have humanbiases.
And I personally agree with those studies and then the findings arereplicable.
The finding that GPD on average is better than human analyst is not verystriking.
What is actually striking is that they are not mutuallyexclusive.
So Imean, in ourpaper, we perform several analyses to show the complementarity between GPDs predictions and human analystpredictions.
Let me walk through the findings a littlebit.
So in the first set of thetest, we try to understand why there are incorrect predictions for bothmodels, like for analysts and forGPD.
And we find that GPD and analysts both struggle when firms reportlosses, when firms are small and when firms have high volatilityearnings, which is a finding that is pretty much consistent with priorstudies.
When information environment isopaque, analysts tend to struggle in terms of predictingearnings, which makes totalsense.
But one interestingfinding, butpreliminary, in thatregression, specificregression, was that analysts were relatively better than GPD in predicting earnings of those information opaquefirms.
Imean, they werestruggling, bothstruggling, but what I want to emphasize is that analysts were doing relatively better thanGPD, especially when they were predicting earnings of firms that aresmall, that reportloss, and have high volatilityearnings.
That finding was pretty much interesting tous.
And then we just wanted to understand why that'shappening.
And according to ourinterpretation, the reason is that for predicting for firms that are small andopaque, youknow, prior studies show that soft information or private communications between analysts andmanagers, met ormore, so basically numerical analysis or GPDsanalysis, they were based on general knowledge andlike, youknow, data that is publiclyavailable.
But what prior studies suggest is that for those specificfirms, youknow, small firms or lost reportingfirms, private communications met or alot.
And then that's why analysts might have comparative advantage over GPD in terms of predictingearnings.
And then after understanding a little bit about why analysts are having like comparativeadvantage, especially in that specific sector or likeinformationally, opaquefirms, we try to horse race themeasures.
We try to horse race GPDs predictions with analyst predictions and try to see whether GPDs predictions subsumes analyst predictions and what we find is theopposite.
They are actually not mutuallyexclusive.
GPDs predictions are not subsuming analystpredictions.
They remain statisticallysignificant, both remain statisticallysignificant, which implies that they convey some sort of like orthogonalinformation.
They contain some information that is independent from eachother, which impliesthat, youknow, GPDs predictions and analyst predictions are actually complementary rather than substituting each other at thismoment.
Ofcourse, in thefuture, we really don't know what's going tohappen.
This technology is changing toomuch, toofast.
We never know what's going to happen after fiveyears, but at least at this timebeing, what we can say for sure is that analyst predictions and GPDs predictions are not mutuallyexclusive.
They have some sort of like independentinformation.
There are some areas where human analysts are relatively doing better than the AImodel.
So what I'm suggestinghere, and it's ourview, theyadd, youknow, human analysts haven't lost their comparativeadvantage, what has become even more important at thismoment, is actually identifying areas where humans can maintain their comparative advantage and try to reallocate and better allocate human resources to areas where they canaxle.
Justfascinating. And I know another part of your paper compared the results of the large language models to the fine-tuned machine learning algorithms that would also be usedhere.
Can you talk a little bit about that part of thepaper?
Actually, that part of the paper is where we weresurprised.
Actually, the first part of thepaper, I would say that we were not verysurprised.
We were sort of expecting a littlebit, but the second part of thepaper, we were very surprised because it has to do with something like how the models weretrained.
So for those who are not very familiar with large languagemodels, I'm going to talk very briefly about how the models weretrained, especially general-purpose large languagemodels.
So general-purpose large language models are trained on a large-corpus of textualdata.
So the primary training purpose for large-Engged models is actually to produce sentences that soundnatural.
So forexample, say that there is asentence, I am aboy, and then we just make boy ablank.
And then the modelprocesses, I amup, and then try to identify what's going to come next to I amup.
And then it has a large corpus of vocabularies inside itsmemory, and try to assess the probability of what's going to comenext.
Boy will be assigned a highprobability, forsure.
But forexample, if we have vocabularies such as like anotepad, it is less likely to come right after I amup, because I am anotepad, sounds a little bitweird.
So basically based on the corpus of textual data that is trainedon, a large language model is trained to produce the most naturally soundingsentences.
In otherwords, it is not trained on a very specifictask.
That is something that surprised us themost, because machine learning models are trained on a very specific purpose ortask.
Inhere, in thiscontext, we trained the model to specifically predict the changes inearnings, especially the direction of the changes inearnings, based on 59 variables used in a seminal paper by OwenPemman, published in1989.
So these variables are actually calculated from financialstatements.
They are verycomprehensive, and then are used by many otherpapers.
And then we design and fine tune a large artificial neural networkmodel, which enables nonlinear interactions of all thevariables, all thepredictors.
And then we trained the model to specifically predict the changes inearnings.
So I would say that it is not a fair comparison between machine learning models and large languagemodels.
Because large languagemodels, although they do have general knowledge about finance andaccounting, they are not trained on earnings predictiontasks.
However, machine learning models were trained on specifictasks, and then we were horse racing these twomodels.
And then we were actually expecting machine learning models to be slightly better than large languagemodels.
But what we find is actually theopposite, not theopposite, but not in the way that we hadexpected.
So what we find in the paper is that large language models and machine learning models are performing pretty much similar in a similarmanner.
Theirperformance, like in terms of accuracy or F1score, were very similar with eachother.
And then that's something that we were very surprisedabout.
And then we repeat the analysis oncemore.
We try to see whether the predictions from HTTP and machine learning models are mutuallyexclusive, or whether they arecomplementary.
We find the sameresults.
They are complementary with eachother.
They are both statistically significant in the horseracing.
And then what this result suggests is that like machine learning models and HTTP might convey some more functionalinformation.
Andyeah, that's pretty much what we find in the second part of thepaper.
Justfascinating. And I think about this alot, because when I teach mycourses, I'm tryingto, you guys are out on the bleeding edge doing this kind ofresearch.
Then there's the listeners of thispodcast, and then a lot of the FPNA professionals that Iteach, they're trying to find ways practically applythis.
And there's all kinds ofcaveats.
And I'm sure you have the sameones, with problems with hallucination and with human in the loop and the potential bias and the models that we don'tunderstand.
Or it's also one thing that I always tell people is if you don't understandfundamentally, not that everybody who uses it needs to be a machine learning engineer or a developer or anything likethat.
But you have to understand how these models are generating their responses toyou.
Because if you don't understand thebasics, you might aswell.
I don't know if you've seen those magic eight balls that you ask a question and you shakeup, and they give you ananswer, it's just a random dice that floats in there and popsup.
And I feel like if you don'tunderstand, kind of the basics of how this information is coming back toyou, then you can't effectively useit.
And in one of the questions I get all thetime, and I know thisis, Imean, this is obviously outside of the researcharea.
But just putting on your sort of extrapolation and putting on your capof, have more imagination around what you see that the future might be withthis.
How do you picture outside of academia if someone is trying to apply this research today and they're trying to figure out how to workin, whether it's the machine learningmodels.
And I know a lot of financial analysts have been using machine learning models foryears, but if they're really trying to leverage the power of theseLLMs, how do you see someone being able to take the insights from your paper and be able to leverage them in a practicalapplication?
Yeah,sure, that's a greatquestion, by theway.
So we've been thinking about this issue for quite a long time aswell.
So as yousaid, academic papers are different from practicalapplications.
First ofall, what we think as acaveat, like there are two thingsactually, first one is that we are not using any textual information as ourinput.
This is purely becauseof, youknow, look ahead bias of themodel.
What we mean by look ahead bias islike, the modelis, as I explainedbefore, is trained on a large corpus of textualdata, and then say that if we provide some information to themodel, especially textualinformation, the model knows what the companyis, and then if we ask themodel,hey, chatGPT, based on theinformation, predict the earnings of2021, and then the modelknows, forexample,hey, this information is actually from Apple in2021, I know that they releasediPad,iMac, or something likethat, and then they experienced a nice fiscalyear.
And then it's gonnasay, based on myprediction, they're gonna experience a nice year in 2021 because they released some niceproducts.
This is notreasoning. It's an answer based on theirknowledge.
It is not based onreasoning.
It is just based on theirmemory, and then actually if we include some textual information in our inputdata, it becomessuper, super difficult to control for the liqueur headbias.
And then that's basically why we didn't include any textual information in ourinput, but inreality, analysts and like informationprocessors, they do refer to many sources of textualinformation, as yousaid,10Ks,MDNAs, or pressreleases, conferencecalls.
There are so many sources of textual information out there that are valuerelevant, and then they contain so much information about the futureearnings.
So inreality, I know that many practitioners are trying hard to incorporate some of the textual information into our model to provide a holistic view of the futureearnings.
And then that's actually prettynice.
In a sense that the product that financialexperts, professionals are developing in practicenowadays, they're developing products to predict the future than nobody has everseen.
So if they are trying to predict the future than nobody has everseen, they arefree, entirely free from the look aheadbias.
And they can actually use whatever they wantto, whatever theycan, whatever data they want touse, like textualdata,image,voice, anydata, and then try to augment the model and improve the performance of themodel.
I think that is actually harnessing the full power of the large managed models in terms of predicting earnings or performing financial statementanalysis, which has not been done in ourpaper.
That's one thing that could be directly applied to financial professionals right now in thefield.
And then the second one is a one thing that we would like to mention as a caveat is that the A&N model that we presented on ourpaper, it's one of the standard models in academicresearch, but one can do better actually in the realworld.
There could be many more hyper parametersettings, there could be many more specifications with more deep layers where the number of layers can be changed or like there could be other breakthroughs and machine learningmethodologies.
Recently there have been like K&Nmodel, which is a developed version of the deep learning model that we currently use in ourpaper.
Youknow, Imean, there are so many other things that people cantry.
At the sametime, there are many other things that people can try to the large languagemodels.
They probably they can add one more fine tuninglayer, they can actually construct a dictionary of the embeddings ofthe, youknow, the othercompanies, conferencecalls, transcripts or other things and then the model can actually search the database and give you moreinsights.
Imean, there could be so many other things that people cando, but everything is limited in academic papers just because of the look aheadbias.
I completely forgot about that and that's an important point is because you had to be able to prove yourhypothesis, you couldn't just have it make predictions based on the most currentinformation.
You had to go give it historical information and completely blind that out and then see how it actually did and then see how the analystperformed.
So it's funny that you talk about that because I've for everyone who's listening in thepaper, there's a companion app that comes withit.
It's a custom GPT that follows this chain of thought prompting where you canupload.
So I've been testing this a lot and I used after reading about how the model didn't do as well with loss making companies and with startup companies and smallercompanies.
I usedRivian, the electronic vehicle manufacturer and I used all the way through their latest quarterly filing and then I had it make a prediction and it'sfunny.
Themodel, it really wants to caveateverything.
So I was trying to really drive it to predicting the earnings per share and itsaid,well, if they doXYZ, then this willhappen.
But one thing because whatever information it has aboutRivian, but then I took it a step further and I'm working on a workflow where 10K analysis is an area where I think we could really leverage these LLMs and I'm using a flow wise ormake, youknow, just different tools where you can sort of string these thingstogether.
So you upload the10K, you upload the financials and I'll usually go three years back on the financial statements and then I'll add in like maybe there's a part of the workflow that where perplexity is going out and looking for relevant information and just adding that into its thoughts andeverything.
But then you go through and I've got different assistants thatwill, youknow, this one does the financialratios, this one createscharts, this one does the qualitative analysis and then they sort of aggregate into a report where the predictions comeout.
And I think that to the practical application point is because people don't have these breaks on that are slowing themdown, they might be able to actually improve performance on those models and that's a very importantpoint.
And I do see that being a way that people are practically applyingthis.
Soyeah, I totally agree withyou.
I totally agree withyou.
That's actually apractical, there was actually a practical consideration because we wanted to make sure that the model follows all the steps that we wanted the model to follow and then we probably experimented so many times and then if we just asked the model to follow all thesteps, all atonce, youknow, for some reason sometimes didn't follow ourinstructions, just probably because the prompt was excessivelylong.
Imean, the prompt that we use in the companion app is different from what we use in our actualpaper, primarily because it includes like textual information and there's way more toconsider.
And as youknow, when the context window getslonger, the higher the probability is for the model to ignore your instructions and then the experience that issue quite often and then that was our second loss resort to make sure that everythinggoes, flows in the way that we hadexpected.
So the step by stepthing, I know that it's kind of likeannoying.
I find it annoying by myself aswell, but Imean, by thatway, we can ensure that the model's following all the steps to beintended.
Yeah, but like as yousaid, it's something that people can develop and like I'm further forsure.
FPNA today is brought to you by DataRails, the world's number one FPNAsolution.
Data Rails is the artificial intelligence-powered financial planning and analysis platform built for Excelusers.
That'sright, you can stay inExcel, but instead of facing hell for everybudget, month and close orforecast, you can enjoy a paradise of dataconsolidation, advancedvisualization, reporting and AIcapabilities.
Plus, game-changing insights giving you instant answers and your story created inseconds.
Find out why more than 1,000 finance teams use Data Rails to uncover their company's realstory.
Don't replace Excel and BraceExcel.
Learn more atdatarails.com.
What we have to learnis, and we gain so much from the types of research that you'redoing, but then we have to learn how to applyit.
And I really think it's so important right now for people to understand and there'sa, I don't know who'sgonna, who ultimately gets credit forit, but you hear the phrase over and over that humans aren't gonna be replaced byAI, but humans who use AI are gonna replace those whodon't.
And that'sreally, I think we're realizing that rightnow.
In companies we're seeingwhere, if a company doesn't have a clear policy on use of generativeAI, or if it's notenforced, or they're trying to lock peopledown, if the company's not figuring out how to useit, then you've got employees going out on theirown, kind of these Mavericks who are becoming basically cyborgs where they're using AI on theirown, and it's not a company-widething.
And I think part of the balance right nowis, we see the advantages ofit.
We don't know how to useit.
We don't know where we can trustit.
And I'mwondering, how do you see that intersection where finance teams are able to use to sort of combine that AI and human experience to enhance their overall decisionmaking?
How do you see that beingapplied?
That's actually a very interesting and greatquestion.
Imean, it's difficult toanswer, first ofall, because think about like two yearsbefore, before OpenAI had releasedGPT-3, could you even imagine this thing coming up within twoyears?
Imean, last two years have beencrazy.
So many new things coming up everyday.
Imean, for major NLPconferences, they're receiving 8,000 plus submissions for eachconference.
Imean, it'scrazy. The number of submissionstripled,quadrupled, over the past twoyears.
So this area is developing and will develop very fast in the comingyears.
And then forme, I'm not an expert in computerscience.
I do research in computerscience, but I do not consider myself as an expert in thisarea.
So I cannot answer for sure what's going to happen in the next fiveyears.
Nobody can actually answer what's going to happen in the next fiveyears.
But my view related to the application offinance, LLM in the field offinance, especially inpractice, is a little bit more positive than otherpeople.
So first ofall, I think there are two main options related to how firms can adopt this newtechnology.
The first one is using LLM as a supportingtool.
The supporting tool meanslike, textual supportingtool, as Isaid, LLM's can summarizeinformation.
They can answer questions fromcustomers.
They can extract some information from complex textualsources.
Imean,overall, they can use LLM's as their assistant and relate it to routine tasks that involve textual informationprocessing.
I think evennow, LLM's are doing pretty well in thesetasks.
I have several other papers on LLM's on how they can summarize information and those summaries are actually veryinformative.
Theyexplain, better explain the contemporaneous marketreactions.
I also have like other papers extracting risk-related information from conferencecalls, extracting text-audit-relatedinformation, which is pretty much hidden in 10Kfilings.
Evennow, LLM's are doing pretty well in extractinginformation.
This thing is going tohappen.
I'm pretty sure aboutthat.
Because it doesn't require then muchverification, then muchknowledge, then much capability or ability to maneuver this newtechnology.
It's justhappening. It's just an assistant tool that's going to improve the productivity of the workers who are employees and it's going to happengradually.
Some firms are already adopting this new technology and then employees are actively using this new tools in their everydayworkflow.
But what remains uncertain is whether LLM's can play a more central role in economic or financial decisionmaking.
So I think our paper is one of the first to answer thisquestion.
There are several other papers talking aboutthis.
Imean, our paper and those other papers are a little bit different from other papers in a sense that they're actuallyasking,hey, can LLM's do morethings?
Like instead of just doing someR.A.
stuff, a research assistantstuff, it is actually making some economicdecisions.
It is providing some human readable outputs and then now it has to become responsibility of humans to verify the human readable outputs and try to understand what is going on inside thesemodels.
What I am fascinated about large nunch models compared to other machine learning models is that large nunch models produce something that isinterpretable.
Machine learningmodels, although they are really good at predicting things or classifyingthings, what remains a black box is how they do theirtasks.
It iscomplex. Things are messed up inside the model and it is practically very difficult to go inside the model and try to identify what's goingon.
For large nunchmodels, it'sdifferent.
We can see the outputs and the outputs are humanreadable.
So for financialexperts, they should be able to interpret and also read through the outputs that the modelsproduce.
And mostimportantly, I think the next direction is going to be understanding where humans are doingbetter.
As I saidbefore, there are areas where machines cannot do well and where humans are ought to be doing well or relatively doing well than themachine.
And then identifying those areas and criticizing what the language models areproducing.
It's going to be the future of the financialindustry, Ithink.
Imean, this is just my personalopinion, but that's I think where we're headingat, at least fornow.
Yeah, I agreecompletely.
And as you were talking throughthat, I endup, so you guys in this study used ChatGPT4Turbo.
Obviously, since you've done thestudy, there's GPT4O or whatever that has slightimprovements.
I'm not trying to make a commercial forOpenAI, but my default model is to use ChatGPT because of the data analysttool.
And as you're going through and talking aboutthat, I thinkabout, Ilove, there are advantages tocloud.
I've used Lama andGemini.
And but because of the built-in data analystcapabilities, I'm going to default to ChatGPT because I haven't seen theequivalent.
And I think it's really important when we talk about what these LOMsdo.
Math is not a skill that'sinherent.
Imean, there's this weird emergent capability where sometimes LOMs can do somemath, but it's not what they're designed todo.
It's like asking a cobbler to tune your piano orwhatever.
But withChatGPT, with the data analysttool, you see it's going and it's writing the Python code under thehood.
So if you ask Gemini or cloud to build you an amortization table or something as simple asthat, chances are it's going to get it wrong because it's looking in its knowledge base and it's trying to findit.
And I think that that's an important thing for us to keep in mind is we don't want the language model itself trying to do our financial analysis for us because it's not inherently good atmath.
But if it hassomething, but we could LOMs can write codes so we could have it create Python applications for us where we're running in anotherenvironment.
Have you done any research around or done any work with these other models that don't have the equivalent of the data analyst tool where it's actually writing the Python under the hood to do themath?
And have you seen any sort of interesting results from usingthose?
Isee. That's a greatquestion, by theway.
So in our paper as an additionalanalysis, now I really don't remember the number of thefigure.
But we tested ChatGPT3.5, which didn't have any access to quantitativeanalysis.
And then we also tried Gemini Pro1.5.
So actually one thing that I'd like to note about the paper is that we don't include any tables in ouranalysis.
We did atrick. So what we did as a trick is we converted the tables into CSVformat.
And then we re-converted the CSV tables into TXTformat.
So basically what wesee, tables with comas and lineseparators, and then we instruct the model directly andsay,hey, this is comma-delimitedfile.
So what we see as comma is differentcolumn.
What you see as lineseparators, differentrow.
And then actually they are not interpreting the tables persay.
Now what remains is whether the model can do the math ornot.
So we tried GeminiPro. We tried Jupyty3.5.
Jupyty 3.5failed. Imean, theyfailed, but they were on par withanalysts.
They are slightly below analystperformance.
Gemini Pro was slightly worse than Jupyty 4Turbo.
I believe that it's because they're quantitative capabilities not as good as Jupyty 4Turbo.
But Imean, you mentionedClaude, the most recent version of Claude3.5.
It's alsofascinating.
It can do a lot ofmath. Their quantitative reasoning is actually one of the best in theindustry.
So if we try again with Claude3.5, I believe that we might be able to get a similar result or even slightlybetter, because they're quantitative skills are actually better than Jupyty 4Turbo.
But it's not like a breakthroughbreakthrough.
It's like a slightdifference.
So it's an empirical question whether it's going to be better orworse.
But as yousaid, after a couplemonths, if OpenAI releases Jupyty5, and then if Jupyty 5 visits genius inmath, that's going to be anotherbreakthrough.
Yeah. And it's those ofus.
And I'm sure you're among us who are watching every day the news on this and watching the benchmark boards and see that the new Claude 3.5 has jumped to the top in certainareas.
And it's in everybody kind of waiting for GPT5, which I've heard rumorsnow.
I saw something the otherday.
I couldn't independently verifythis, but that it may be late25, early 26 before we see GPT5.
So it's fascinating towatch.
And by the time this podcastairs, there may be a new leader outthere.
So it's an exciting time to be watching allthis.
I'm going to have to try Claude3.5,though, on some of the this type ofanalysis, because I am just defaulting to GPT 4.0 just because I know that data analysttool.
I can get mycharts. And now they've got those interactive charts and allthat.
But we do need to keep up with the other ones aswell.
So that papersout. You're fielding media calls on that andeverything.
Tell me about what's next foryou.
What are you working on rightnow?
Are there any new projects or areas of research that you're excited about rightnow?
Yeah, ofcourse.Actually, that's apretty, pretty nicequestion.
But there are several projects that we are currently workingon.
But it's a treatsecret.
We can let this closeeverything.
Please staytuned. We're going to disclose everything once they're ready toshare.
OK, but let me talk a little bit about my researchinterest.
And then I'm going to just flow naturally into some of the ongoing projects that I'm currently workingon.
My interest still is related to information processing ofinvestors.
It's the core question ofaccounting, as Isaid.
Imean, but if you dive deeper into information processingresearch, I think there are two lines ofresearch.
They are inter-currelated with eachother.
But they are a little bitdifferent.
First line of research talks about the processing processitself.
Imean, how people processinformation, how can we quantify the processing of non-structureddata, textualinformation, voiceinformation, how can we measure some characteristics of textualdata, stuff likethat?
That is actually the first research interest that I'm currently workingon.
And then actually another line of research that is inter-currelated with the first line of research is talking about the benefits and costs oftechnology.
There are some big questions that are yet to beanswered, which group is likely to benefit more from the newtechnology, which group is likely to lose or get replaced by the newtechnology?
What are actually people doing with this new technology inside thefirm?
Our FSA paper is somewhat related to this second line ofresearch.
It's actually the combination ofboth.
Actually, it is talking about whether the model can dothis,like, Imean,FSA, and then whether the model is benefiting analysts or whether it is indicating something where analysts should focus moreon, like stuff likethat.
Imean, this second part of information processingresearch, I personally feel that we needmore, a lotmore, because it'snew.
Many people have been talking about LLM's can dothis,this,this,this,this.
And then we've been talking about this for more than twoyears.
Imean, we still don't understand what people are doing with this newtechnology, whether the technology itself isefficient, whether it's giving you right answers orhallucinations, whether people who are using this new technology arebenefiting, we're losing something in the market or in terms ofproductivity.
There are some big questions that accounting researchers for finance researchers shouldanswer.
And my current research interests liethere, especially talking about whether investors are using thistechnology.
If they have access to thistechnology, whether they're information processing or whether they're financially decision making is going to improve ornot, what kind of prompt or what kind of the output data they shoulduse, stuff likethat.
Imean, these questions are big questions that should beanswered.
Imean, this provides more causal evidence to the benefits and costs of large-scale models or in broaderterms, AI to the society or to the welfare of thesociety.
Then my ongoing research projects are broadly investigating these issues from a more broaderperspective.
This might include likesurveys,experiments, or archival researchmethods.
But I'm very open to like many research methods to answer these causal and core questions that the new technologies are going to bring to the financial market and the marketparticipants.
Thinking about where the intersection where you'reworking, you have to keep up with not just the latest intech, but the latest in what's going on in finance and accounting and how people are using thisinformation.
And so I'mwondering, Imean,obviously, I'm sure you read millions of papers and are very tuned in and allthat.
But how doyou, Imean, with everything moving so fast is we're just talking about how do you stay updated with the latest trends both in AI and in what's going on in finance and keeping up with that and how you incorporate those into yourwork?
I try mybest, but I cannot say that I'm best person to dothat.
But like I try mybest. As yousaid, there are two things that I have toconsider.
One thing is finance accounting research and the second one is computer science ortechnology.
So for finance andaccounting, it'seasier.
In a sense that the number of newpapers, especially in this field ofaccounting, finance related to information processingAI, there's a large models asset pricing is relatively not thatmany.
Like there are couple20, 30 papers every day that I have to keep upwith.
So what I do is like I keep track ofSSRN, which is an equivalent of archive of computer science and socialsciences.
So the researchers post their preprints for review andcomments.
It's my daily routine every single day in themorning.
I just go toSSRN, try to see whether there are any papers related to myarea.
I just read theirabstracts.
If I'm interested in thoseabstracts, I try to read all thepapers.
I think I spend at least like one hour every day in the morning to try to see whether there are newpapers.
If there are interestingpapers, I share them with myco-authors, with myfriends, and then try to think about extensions where potential research ideas stemming from the newpapers.
One thing that I really don't like about social science research is that it takes too long to publish apaper.
Normally it takes about like one to two years to publish one paper in topjournals.
They hand like within twoyears, even though they getpublished, even though I keep track of all the papers that are published in top threejournals, they are papers that were written like two or three yearsago.
So they are pretty muchoutdated.
So what I do is just like keeping track of new papers posted on SSRN for finance andaccounting.
But for computerscience, it'simpossible.
As Isaid, like for eachconference, people something like 8,000papers, and then I'm not a computer science researcher and I cannot do things likethat.
So my view about computer science research is that I have to get my hands dirty to learnsomething, especially related tocoding,technology, I have to get my handsdirty.
Reading thingshelps, but it doesn't helpultimately.
So what I do is I try to get involved in at least one computer science research almost everytime.
Like thisyear, I published one paper with myfriends.
They're notresearchers, they're just myfriends.
And one of the top AI natural language processing conferences calledACL, it's about like multi agent large language models debating with each other and talking about evaluation oftext.
So it's pure natural language processingpaper, but actually it stems from my interestin, youknow, conferencecalls, because I was interested in how conference calls wereadministered.
And then I was interested in how managers and analysts debate with each other in conference calls and reach conclusions or sometimes they do not reachconclusion.
So that wrapped myattention.
And I wastalking, I'm thinking about whether I can model that situation with multiple language modelagents.
And then I developed the idea into thepaper.
And Imean, like once I have to write a paper with newidea, I will have to do the literaturesearch, extensive literaturereview, like have to do thecoding, talk with myco-authors.
Even though it takes a lot oftime, I learn alot.
If you write one paper and if it gets published in one of the topconferences, it has to go through the rebuttalprocess, you learn a lot bydoing.
So that's how I keep track of the most recent technologies in computerscience.
I love to hear you say that becauseit's, I'm likeyou, I'm trying to keep up with as many papers as Ican.
And I'm going to give aplug.
Or do you know who Ethan Mollickis?
He's a professor atWharton.
He'sactually, I think his PhD was anentrepreneurship.
But he's got a GPT that hebuilt.
I think it'scalled, why is thisimportant?
And if there's a paper that I'm only tangentially interestedin, I'm not going to read the wholepaper.
I just dump it into hisGPT.
And it spits out the key factors ofit.
And it's the best summarization to all I'veused.
But to yourpoint, I think it'sgreat.
If you're just operating in this sort of research onlyrealm, it's easy to get removed from the practical applications ofit.
So to hear you talking about going out there and doing andbuilding, that's for metoo.
Because I also teach a lot of courses onthis.
But if I stop and I'm just going down that teaching path and I'm not also making and building stuff along theway, then I feel like there's a remove between what's actually happening and what we're talkingabout.
So I love to hear you saythat.
That's a greatanswer.Well, this has been just a fascinatingepisode.
I'm so appreciative of you comingon.
And Ido, before we let our guests off thehook, Ido, because we've spent some timetogether, I do at the end of theshow, like to find out a little bit more aboutyou.
So kind of our question that we always askis, what'ssomething, maybe something on the personalside, that not many people know aboutyou?
Yeah,actually,yeah, that's a goodquestion.
So my life as a PhD student cannot be moretransparent.
Imean,I, wake up at themoment.
I just crawl to mycomputer, start myresearch, itlaunch, doresearch, itdinner, sometimes do somemeetings, doresearch, and go tobed.
So my life as a PhD student cannot be moretransparent.
But one thing that people don't know a lot about is as a one military experience as a South Koreancitizen, I had to serve in the military for 19months.
But my military experience was prettyspecial, because I got a chance to serve in the USmilitary, stationed in SouthKorea.
So Imean, they chose several soldiers who were actually okay in Englishconversations, and then they sent them to the US military bases stationed in SouthKorea.
I was chosen as one of the Korean augmentation to the United States Army in SouthKorea.
I was sent there and then spent 19 months with USsoldiers.
That experience was superunique.
Imean, I learnedculture,language, mostlyslang, but atwords, a lot ofit.
And what was actually very nice for me waslike, I did a lot ofexercise, physicalexercise.
And then at the sametime, I had a lot of free time starting from like five to sixPM, I wasfree.
And then I used that time to take some lectures from Stanford talking aboutNLP.
And then that was actually basically when I learned how to code and how to think about these computer scienceprojects.
That was like in the middle of my master'sprogram.
So it was like a fresh restart fromme.
Like I was like thinking about thinking about something completelydifferent.
I was thinking about computerscience.
I was thinking about natural languageprocessing.
I was like doing the training during theday, like learning English culture and Englishlanguage.
So Imean, it was a pretty uniqueexperience.
I made lifelong friends that I talked to evennow.
So Imean, like many people know that South Koreans have to serve in themilitary.
But whenever I say that I served in the USmilitary, they are kind of like surprised to know that I waslike, youknow, I'm working with actual USsoldiers.
That'sincredible.Yeah, that's greatstory.
And what a great use of your timetoo.
So I'm just trying to picture howbusy, Imean, you just obviously your life stays busywith, that was a PhDstudent, but serving in themilitary, working on your masters and doing the other courses at the sametime.
That's alot. And probably the physical activity helped to sort of balance things out and clear your head along the waytoo.
Allright, so I'm gonna throw you a little bit of a curve ball heretoo.
But we have to ask all of ourquestions, all of our guests thisquestion.
Andit's, youknow, we get a variety of differentanswers.
And I know you in the nature of yourresearch, probably don't spend as much time in Excel as many of our listenersdo, but I'm wondering if you have a favorite Excel function and ifso,why?
To be100%honest, I do not use Excel toomuch, especially in terms of data manipulation or dataprocessing.
This is purely because of the reproducibility of academicresearch.
As a social scienceresearcher, I take full reproducibility of accounting or finance research veryseriously.
I personally believe that based on the descriptions provided in thepaper, everybody should be able to reproduce the findings to the main results in the paper without manydifficulties.
But if we do things inExcel, there is no way that we can provide the codes to otherresearchers, especially like say that we createvariables, say that we runregressions.
Even though we get theresults, there is no way to communicate ourselves to otherpeople.
So that's primarily the only reason that we don't use Excel thenmuch, but there's onething.
I'm a big fan of Excel in terms of creatingfigures.
I know thatPython,R,Stata, like statisticalsoftware, they provide a lot of great figurefunctions, but one thing that I really don't like about those functions is that they lackflexibility.
They do provide codes to change colors and likeeverything, but they are not userfriendly.
You have to memorize the name of the numbers or like the codes for the shapes of the dots and everything likelines.
But inExcel, you can produce publication ready quality figures in several clicksaway.
Imean, they do provide a lot offlexibility.
Changingcolors, it can be done within like fiveminutes, like changing formats andeverything.
And the resulting figures look prettynice.
That's the primary reason that I loveExcel.
So whenever I get the final data sets out from like any statisticalsoftware, I convert them into Excel file and then try to get some figures that look nice fromExcel.
I don't have any specific function inmind, but figures in general are great inExcel.
As someone who's tried to battle through Matt Platt-Lib and Seaborn and allthat, I fully appreciatethat.
Andyes, Excel does have wonderful charts andgraphs.
And I will saythough,again, I'm advertising for OpenAI for somereason, but the new interactive charts that are onGPT-4O, I do love those and those look good and you can drill in and interact withthem.
But it'sstill, it's not gonna replace Excel and so great response onthat.
Well,Alex, we're coming to the end of timehere.
And Iguess, youknow,just,I'm, what a fascinatingconversation.
And Ireally, I love the work that you're doing and I'm looking forward to continuing to follow yourwork.
And for our listeners is if they wanted to follow you and keep up with thework, what's the best way for them to findyou?
Imean, I have awebsite.
I have my email address there or you can hit me out onLinkedIn.
Yeah, I cannot guarantee that I can answer to all the emails that Iget, but I will try my best to respond as soon as Ican.
So if you have anyquestions, especially related to myresearch, please don't hesitate to give me up on LinkedIn or viaemail.
Alex, thank you so much for comingon.
Yeah, thank you so much for havingme.
Itwas, it was mypleasure.