Summary
Notes
Transcript
So welcome, good morning. 8:00 a.m. on a Tuesday. Excited to be here to talk about AI, accounting and finance. So my name is Ed DeHaan. I'm a professor of management here at the PSP. I teach in the fall when the students first get here, the first MBA quarter. I teach financial reporting, which is a mix of financial accounting and communications strategy. So I tell my students on day one, When we think about accounting, we should always be thinking about accounting and finance together.
If we think about what accounting is, it is all about providing data to somebody to make a decision. We want to make the most informed decision at the right time. Now none of my students are going to be accountants. They're MBA students here. But many of them will be in the finance function. And the finance function is where that data is being used to make the decisions. There's a lot of decisions people make. It might be about fundraising. Am I going to raise capital?
It might be an internal decision. Where am I going to invest? What new projects do I want to fund? How am I going to manage my working capital? Many, many decisions in the finance function that rely on accounting data. And I tell them on day one, if you are going to be an executive functioning well in that space, You have to know where the accounting information comes from. Because these are not facts. If you look at a financial statement and you think those are facts on a page, you've got it wrong. They are estimates. They're created by human beings who have different incentives than you, perhaps. You really need to understand the source of this data. Likewise, I tell them to be a well functioning executive.
So again, you've got to understand the decision needs of the person you're talking to, what is it they're trying to decide, what information do they need, and how can I most effectively communicate that information to themSo there's this circle, accounting, finance, data, and decision. So I spent the last six months working on some cases with Microsoft, looking at how they manage AI risk inside their company, both in terms of the external products they're selling and in their internal accounting and finance functions. And that case is very nearly ready for prime time.
I'll be happy to send it to you afterwards. But just sitting in a meeting one day with Julia Denman, who was their chief risk and audit officer, she just dropped this beautiful law. She said, AI is collapsing the distance between data and decisions. I just heard that and it resonated. I thought, that is exactly right. That is exactly what AI is doing. And the place where I can imagine that and see it the most is in perhaps the sort of, I don't know, sexiest accounting of example, accounting and finance, which is a company who's got an earnings release And then there are some investors over here who want to trade on them.
But some junior associate inside of some hedge fund would get this thing and manually key in revenue and earnings and whatever the important numbers are into a spreadsheet. That was their job for the first year.
This, humans plus Excel, could take hours. And in that state of the world, it was fine. Because back then, you only could train between 9:30 and 4:00. Earnings announced were coming out at 7:00 AM. You had hours to complete this test. And so that's the state of the world we were in.
I picked 2022 because this is a few months before we all learned of chat GPs. This is, these days, the old times. So the world had changed remarkably. So here we have the data, here we have the decision, and I put a box around this because this is really all one thing. So we don't have a junior associate manually keying entry information. We have optical character recognition, we have natural language processing, we have linguistic analysis that can find those numbers.
are and what their functions still are today, even in the denny animal. Then in the middle here we have machine learning. I'll also talk about what those are. Now those are artificial intelligence. AI wasn't invented in 2022. But when I talk about AI today, mostly we're gonna be talking about Gen AI. But that would do all of the analysis. And then the trade design was completely algorithmic. So this with automation plus machine learning, this would take seconds or maybe microseconds.
Of course we had social media, that didn't exist in 2002. There was vast amounts of information being produced by everyday people, low signal quality, but lots of signal, right? So you can aggregate the wisdom of the crowd, we got pretty good at that. News articles had all been digitized, everything was online, so we could get reporters' views from all around the country, hyper-local all the way to national.
We had tons of alternative data.
So my students, She got satellite images. And what she did is she showed that you can use these satellite images to count things like how many cars are in the parking lot of Walmart or in the parking lot of the factory. And that gives you a sort of real-time understanding of the economic activity happening there. And then you can compare it to what the firm actually reports. And if you report really strong sales but there weren't many cars in the lot, well, that's a signal.
There's no human in this loop because it needs to happen far too quickly. So here we have human next to the moon, or human out of the moon. The decision is designing the system. And then in real time, the system does what it's doing. Which is honestly a lot of how we work today. Can we evolve with this kind of thing? Yeah. Did I miss anything super interesting? It's kind of the way it works, right?
I gotta tell you, I've been teaching this for a few years, and every year when I fast forward, you know, from 24 to 25, today we're at 26, things change. Today we're at 26, here we are. Foundation models. The foundation models are the underlying GPT, CLAW, etc. The Frontier model is the latest version of that. And those foundation models have been amazing at extracting signals from data. So some examples here, things like context and emotion.
So I told you, when the conference call text, when the conference call audio, we can use sort of bag of words approaches. Expenses back decreases back in large samples. But in an accounting sense, an expense decrease is actually good. So these tools that we were using were simplistic. They got the job done a little bit, and we can do much better. Large language models could understand that as context. So if we use a large language model to code for me the tone of this context-balled transcript, it's so much better for the job. And here's just a screen grab from a hedge fund that I talked to recently.
In real time, they have millions of people talking on the calls. Here are the different emotions that they think they're talking about. So in real time, they can see, oh, well, the CEO is happy when they talk about this sentence, so that's indicating that probably it's good news, maybe even incremental to what the words say, but when they are fearful or disgusted or angry, it's betraying some part of their information.
In real time, what the large-time response does is given everything it knows about that company, and given everything it's seen in the conference call slide deck that's already come out, it predicts what the manager's response should be. And then it compares what the manager actually says to what the predicted response was. And if the manager says what the predicted response was, well that's not information, right?
That's predictable, that shouldn't move the market. But when the manager says something different, Now that's a signal. That means there's news here. There's some amount of surprise. And they show that you can trade on a positive and negative delta. And they'll look back and think,One more. These are the types of signals that we can extract with these foundation marks. Cross-doctorate reasoning. I'm doing a paper right now. What we look at is we've got managers talking in multiple different places. We've got the earnings announcement, we've got the conference call, we've got the investor day.
That's inconsistent, and that's where you start wanting to drill down as either an investor or a trade partner or even somebody who's thinking of going into a joint venture with them. These are the types of places you want to go.
Gen AI can do a lot of things, but as we're going to see, it's not great at everything. Here, machine learning, in my experience with funds, is still the dominant technology. And that's for a few reasons. The first one is it works really well. So machine learning technologies, what they do is they take in tabular data, meaning like a spreadsheet, and they can come out with a numeric prediction, which is exactly what you want when you're doing something like an earnings forecast or evaluation.
Given all of these public documents that we know was on the mutual fund manager's desk, How much money did they leave on the table by not processing these data more effectively? And we use machine learning, relatively straightforward random forest model. And we show that with this model we can beat 90% of all of the human mutual fund managers who have ever lived.
We beat them by an order of like 600%. which shows that even these relatively straightforward technologies, random forest, once you get into it, it's not crazy complicated. We've had the technology for 30 years. It's just become much cheaper to run. Even with that, You can extract Now I will say, just as a complete aside, as an insight to how academia goes, we wrote this paper. I put it online last night, May of 2025. And a few academics noticed, and they wrote to us, and we got some interest that way.
In fact, one of the biggest funds in the-fund issuers in the world is now running our AI analysts alongside their humans. They're not relying on AI analysts, but they're using it to sort of quality check themselves.
High upfront cost. But once you have built the random forest model, the neural network model, you can run a query through that like adding up a column of Excel. It's just fitting a value. So it's cheap. Whereas we're going to learn with Gen AI, it's not cheap. You're paying per run. Another is that it's fast. Many AI has latency. In this space, you need speed. We're talking microseconds. And again, once you've built the random forest model, throwing data through it, very quick. It's consistent. We'll come back to that. It produces the same answer every time.
You don't train it, so you don't know what's in its brain. You also can't stop it from changing over time. It just changes. This is a problem in the trading space. Because what you're trying to do when you are developing trading strategies, what you do is called back testing. You take historical data and you see how well would my, in this case, AI analyst or whatever, have performed in real time. We can't do that with a lot of money, Tom.
Because a large time tunnel knows history and it can't unwind. So if you said, with machine learning what you can do is you can pretend you're in 2020 and you can predict Apple's earnings for 2021 and for sure the model doesn't know what Apple's earnings for 2021 are because you built the model and you know what it's in for. A large time tunnel can't do that. So this is a problem that when large language files first came out, people got very excited and started trying to use them for predictions.
They are, before their conference call, putting all of their responses into GenAI and saying, predict your response.
Predict how people respond to my response. There's a lot of reverse engineering going on there. Or they're hiring acting coaches to help them change the valence of their voice and the emotion of their voice. Of course they're doing those things. News articles. Well, this generative AI impact is very different. News is functionally dead.
You can't monetize private information that's being produced by news outlets. We already knew that media outlets were suffering, declining dramatically. It's been a focus of some of my research. Well, now it is so difficult to monetize content. As soon as you put out a report, you're not even getting clicks anymore. It used to be Google would at least link to you if you get some clicks to your Miami Herald and you get some dollars that way.
Now, Gemini is just summarizing. You don't even go there. That model might be dead. Social media. Full of AI generated crap, right? So now we have garbage AI talking to other AI. AI talking to AI. Where does that end? We don't know.
Yes? So won't these models pick up all this garbage data and process that? Garbage in, garbage out.
It used to be that, and some websites are cracking down on us, but there's this website called Seeking Alpha, which is where sort of amateur analysts can go on and give opinions on a company. What we saw is when Gen AI came out, it was a huge uptick in articles, all written by Gen AI. And then we got a Gen AI trader over here Garbage talking to garbage. It cracked down, now it's back to humans, but the vast majority of content we see, say, on social media is genuine.
Great, so Gen AI is good. And then for whatever reason, at this hospital, the Gen AI got taken away. Then the human rate went down to 22%.
So the human got worse at their job by relying on Gen AI. And this is all within a matter of months. So we're going to talk about many of the problems that are affecting human beings, not just in finance and accounting. Deskilling is one, over-relying on the models is another. And in this case, if your humans over here aren't doing a good job of building this loop, then that loop is just going to be over and over and over again making the same kind of mistakes.
So one of the problems I hear among managers now A lot of this we're foreshadowing, but that's okay because this is your time, so I'm happy to talk in any order. Is your talent pipeline. So one of the things county programs are really struggling with is we don't need people to do the low level tasks anymore. What we need are a bunch of managers, maybe doing higher level tests, maybe going deeper on analysis.
It might keep me up at night if I was running a public accounting firm and needed to think about how will my hunters know how to do their job.
Given the observations on whether this beMore short-term loansSo I'm at a public company I've noticed the analysts that have adopted more AI, they're much further, they're much less familiar with what's actually in the supplement.
So thanks for sharing that. I don't have any good data on it. Like I can't systematically show you that's true. from everybody that there is this increased focus on short-termism.
But that's actually not what we call fundamental information. That's not making the market more efficient. It's actually making the market less efficient. What we really want are stock market prices that reflect the true value of your companies. Because that way, you can raise capital at a fair price, you can have innovation, et cetera. So I'm hearing it in multiple different directions, that it's perturbing the market in ways that ultimately might not be good. So, Mr. Spears, why are you against six-month financial reporting? Yeah, so a lot of research on transparency in financial reporting. And I think I define that for small companies where the cost burden is significant.
And is the 6-1-1, did you accept it, John, or adopt it, or is it still a discussion? JOHN W. No, it was a tweet, and now it's an SEC proposal. There's a lot of people who think, at least under the current SEC, it most likely would be the best. It will not be. It will be. It will be. That's the thing. I mean, nobody knows. This White House is very difficult to predict.
External financial reporting is fun, it's kind of sexy, we have a lot of data on it, but the vast majority of use cases for generative AI will be internal processes. There'll be companies using it accounting finance to improve efficiency, to generate competitive advantage.
How many of you personally have gotten into something like Cloud Code or Codex and are using it to develop multi-step agents that interact directly with funds? Okay, a few of us. That's the bleeding edge, right? So we'll talk about all of these things today. But again, stick up your hand. You can see that I'll keep moving. Otherwise. So why is it that firms are rushing in to adopt AI? And again, I'm really talking about generative AI. I mean, it's something to talk about. It's about growing revenue. That's obviously the main thing that most of us want to do. So you are probably, and many of our students are, building amazing products based on Gen AI, things we couldn't have imagined.
So an example I saw from a company recently, and their product has nothing to do with Gen A, it probably never will, but they had 20 years of email records between their sales team and clients. And buried throughout those email records are little tidbits of pain points, reasons they went with a competitor, Features they wish they had. But until recently, it was impossible to unlock that.
And in a matter of hours, they had processed these emails and they had extracted from that a whole bunch of decision-relevant information to guide their product development. Low-hanging fruit, right? That is taking the data you already have in a low-risk environment and using it to accelerate it.
So the other side is efficiency and cost savings. So clearly, if we can automate routine tasks, if we can supplement extensive labor, then we will have efficiencies.
very short-lived phenomenon where CEOs said, "Oh, well we don't know where Gen-AI could be used, so what we'll do is we'll measure token usage, and then we'll tell everyone we'll be using tokens, and that way we can find all of these use cases, whether it's going to be efficiencies." So what did people do? You know, write me a Harry Potter fanfiction. You know? And just chunk them up and turn them away with thousands of tokens.
I also think that if you approach your staff, your workforce, with this dual focus of improving efficiency to accelerate growth, and if you can help them understand that growing the organization benefits everybody, essentially what you want is this growth mindset.
You can convince people that no, automating your job won't put you out of it.
Folks just feel like they're gonna get replaced by AI, and so they quietly or loudly refuse to just take part in any of the training. So we have a group ofSort of like brain trust that focuses on the organization. Yeah. Other staffers. Yeah. Yeah, I think I saw a hand over there too. Did he start getting a lot of,We decided to take out all the AI and the engineers budget and desperate into a separate company.
Hired engineers outside, It's led by someone that's running the company andYeah. Three key players are on board on this and take them access to our data in a closed environment, and that's what we're doing because they were all just lagging back there. And uploading PDF files. tokenized and stuff like that. And so we decided to take it out. And also what was happening is that the managers that were adopting it correctly started to get more time to get into other details that they didn't have time to do it.
Essentially you created another organization with different incentives to create Gen AI efficiencies. And then what long term will replace your existing legacy force? Let's do one more and move on.
So there's a big conflict in the organization between IT and security. And finance. Yeah. Therefore is dying to use all those tools. He wants to automate. And the IT security guy that I'm having the biggest trouble with refusing anything because of data privacy, compliance, GDPR, all those things.
It's a nightmare. So basically when I got one, I sat with him and I told him, show me what are the concerns and the issues. He agreed and he gave them access to cloud for work. And he gave them only limited access to a folder, but any, God forbid, file that has any IBAN, or anything that has to do with customer names or anything, would not upload to Chrome. So he gave him the axles, but he cannot use them.
Much different in accounting. If you've mistaken accounting and finance numbers, that can be very costly. Imagine giving a debt covenant report to your bank with the wrong number in it. That is going to badly damage your relationship and undermine your ability to get funding to report. Giving the board a wrong number.
So they got a bunch of invoices in pesos, which is about a 20 to 1 US exchange rate, and we're paying it out in US dollars. And they did this multiple times until someone finally caught on through some analytics. It's going to happen over and over again. We've seen this happen in history. During the financial crisis, this wasn't a machine learning model, but the reason the global financial crisis happened is everyone was relying on the same credit scoring models. And when the state of the world changed, they all failed.
AI processes have to be documented, need to be reproducible, and testable for internal control purposes.
So in the United States under Sarbanes-Oxley that's a requirement, but even as a private company it is a good idea.
Whatever software you use for accounting and finance, there's a lot of automation in this. ERP systems, robotic process automation, optical character recognition. You know, if you pay an invoice and it automatically goes to different levels of approval depending on the size of the invoice, that's all I need. So what are automates? Rules and scripts. Humans program tasks.
It's like a calculator. If you put 2 plus 2 into a calculator, you will always get 4 unless that calculator is fundamentally broke. That's important. Static is important.
So we have natural language processing, random forest, gradient boosting, machine learning. These are very different versions of machine learning. What it automates is predictions and classifications. And what makes it artificial is that I don't program the rules, the model does it itself.
They need to be accurately labeled. If you mislabel, you're not gonna have a good model. You need domain knowledge. So in accounting and finance, if you just hire a programmer to come in and say, build me a machine learning model to forecast working capital, they don't know what working capital is. It's not going to work very well. You need domain knowledge to understand which data needs to go into this. and you need computer science expertise to actually run the model. That makes it very difficult to implement. It is deterministic once the model is frozen.
So drift there is that the world has changed and your model has stayed the same. What you need to do is over time keep updating your model so its prediction accuracy is still as high as you want it to be. That's a control you need to have in place if you're using Shima.
This is a big difference: implementation difficulty, low.
This probabilistic is a problem. Outputs will change. We know this to be true.
So even if it's relatively consistent today to tomorrow, the next version of the model that comes out could be completely different. And now what we're seeing is a new thing. We'll talk about the economics of and profit and Open AI, they've been throttling back their power. You might have noticed this, sometimes it takes longer.
So from a control perspective, I summarize it like this. With automation, we control the rules. When machine learning, because the rules are made up by the model, we control the model. So here, because we don't control the model and we don't know the rules, we control access and we control outcomes.
And this is really uncharted territory. So we are seeing, and this is KPMG, who I used to work for, they're piloting tools, everybody's piloting tools. The SEC hasn't made up their mind on how to do this. The PCAOB does not audit this stuff. It is the wildest. I will say this, I think this is getting back to what Mark mentioned earlier. These days it is table stakes to have an enterprise grade version of one of these models.
So it can track what every person did. Well, that's useful. At least you can see the prompt they put in. Even if the prompt gave the wrong answer, you can make sure they didn't pay the invoice because they said, ignore our normal controls, pay this because it's my friend. We have, and here's the security logs, right? So SOC 2, ISO, ISO, these are the ones that are relevant for most companies. These are controls around the technology.
Turn off the training for one. You will not get out of the normal version audit logs, so you won't know who did what. But if this is not a mission critical function, then maybe you can get away with it. If this is something that's critical, like preparing your I don't know. Maybe accounts payable even, right? Tax returns, bank reports. If this is mission critical and you need to make sure that someone in your organization is following the rules and that you're worried about data leakage, then maybe it's worth upgrading.
There's a version in between, by the way. There's a Teams version. And the Teams version is much cheaper and it gives you a lot of those controls. So the logs and everything and it's somewhere in between.
I would be talking to your internal audit, your internal and external auditor from the beginning before you even implement these things. I talk to external auditors, I still know people at KPMG, I talk to people all the time. What they tell me changes monthly. So I would just be having a conversation with them in real time. Even differs across firms. The big four seem to be much more conservative than some of the smaller audit firms.
I was actually going to talk about-In a few slides, Gen-AI, we're used to thinking about technology like Excel. Once I buy Excel, running numbers through Excel is free, right?
I can add a thousand columns, a million columns, it's all free. Every time you run a Gen-AI query, it costs you money. You have to think about the unit economics, the cost of the AI more like a human than you do like a typical machine. I'll come back to that. Also, they're never going to lower the rates anymore.
So I'm going to use accounts payable invoice processing, not because this is a sexy task, but because every one of your organizations needs to pay invoices. It's something we need to get our heads around, and it can demonstrate the challenges and the strengths of each of these different types of technology.
Okay, so first of all, Traditional automation. This is the stuff we got in the 2000s. We have optical character recognition, which is really good at reading the text. And then we have robotic process automation. Robotic process automation. That's just, RPA is just like following the scripts. It's good for looking for words like invoice and number and total. So we can extract data quickly. Handoff maybe by RPA, frankly most organizations would be human.
It's easy to optimize. Doesn't work very well though, right?
Because here's 10 different invoice dates that we can read, but OCR and RPA would get tricked up on, probably nine. So lots of exceptions for humans to read.
They develop a machine learning model that's very good at looking for duplicate invoices, invoices that have got an extra zero on them, fraudulent invoices, things like that. And if you can quantify the risk, if it's low risk, it will automatically be paid. And if it's high risk, it will get accumulated.
They said last year, more than 50% of the dollars they spend We're talking tens of billions, never saw a human being touching animals.
Okay. Let's get the generator there. All right, now look, this guy is sort of limited. We'll get to the GenDict AI in a minute. But Gen AI can read and invoice with virtually no training. documents, key information, the invoice amount, the date, the payment terms, paid in two days for a 5% discount, that kind of stuff. You give it a stack of invoices, it has seen in its brain enough invoices, it'll find the key.
It'll probably work pretty well. I find that Fewshot works much better. So if I use PodPoWork, and I give it access to a folder, and I say, here's 100 invoices, here's a spreadsheet with the information extracted, I want you to design a skill file that will do this on a regular basis. It can write its own skill file. And then from then on it works much better than zero-shift. That's a huge shock.
And what's much better about GenAI than machine learning is it can at least provide some explanation.
If your company has gone through the trouble of developing a good machine learning model for something like this, which required a lot of training, I would probably keep it. Because as a medium-sized organization, you probably get thousands of invoices a year. If you use GenAI, every time it processes an invoice, it's generating a bill for you. Token usage. That's different than Excel, right? That's different than machine learning.
Machine learning, as I mentioned previously, once the model is set up, An invoice will go through it in a split second and it'll cost you virtually nothing. That's not the group of J-Oct. So if God's in trouble, I would still use, it's also deterministic. We will get the same answer every time. Yes, it requires retraining, but that to me is less bothersome than Opus 4.8 doing something fundamentally different from Opus 4.7, and all of a sudden I'm paying invoices I shouldn't pay.
Those global instructions come from OpenAI and they have weird things in it, like don't talk about gremlins unless it is directly relevant to your topic, which is actually in there. and you don't control that stuff. One question. Are you seeing that? AI, you didn't need to manage. That's no high. For example, a large US nuclear that's going into Latin America, you're not going to say their names. I've seen every country probably trying to like five or 10 days late versus what you gave them.
And it's almost humanly impossible to do it manually or you have to have-I love that. Let's come back to that in two slides.
How are we going to measure ROI? We'll come back to that and put a pin in it. Because measuring ROI is hard. But we've already had a preview. Can I actually reduce my headcount? If not, then is it actually going to save me money? Risk management. Well, what risks must we continue to manage? In this case,Straight forward, we don't want to pay people who shouldn't be paid, and we do want to pay the people who should be paid.
But are we creating any new risks? We'll come back to that after the break.
And then the strategic improvements. Let's not just do the same thing as last year. Can we change the data so that we can make better decisions to accelerate food? How can we reimagine this process?
So as we come back, we're going to take 20 minutes. When we come back, we're going to talk about these things. What risks does GenAI introduce or amplify in this setting? And taking a step back, is there something we can do here to improve Our functionality as a finance officer.