Summary
Notes
Transcript
First, welcome. So in today's session, we are going to do a very deep dive into generative AI and agents. I'm the assistant professor in the computer science department here at Stanford. Just to give you some background. I teach several courses here at Stanford. One of them is called the Deep Learning for Natural Life Process. And we need to spend 10 weeks to do a very deep dive And today I'm going to comprise all of them. Put you 60 minutes. And I know you can do it. OK, let's get started.
I want to start with one figure here. And again, in the process, if you have any questions, feel free to let me know. Okay, my clicker needs to let me know first. I think it's all about-We can start with this picture. This is when 54 was introduced. Again, like almost two years ago. The reason why I want to do this is that if you look at the left side, this is different type of exams we have in the world.
Even for some task of wine tasting, it seems that models are better than humans at articulating the wine tasting outcome.
You probably have heard of coding agent, open claw, and many, many other agents. Companies I talk to use a lot of agents within their org, and there are different ways to arrange them. One thing I think that companies realize is that the benchmark numbers we have on the leaderboard actually do not translate very well to the actual product. And part of the reason is that because most of the evaluation is towards can we solve math?
Can we finish a specific coding? So over the last several months, there is a trend to really move AI towards economically valuable real-world tasks. So on the right, this is the data set called the GDP-Doc, where OpenAI actually constructed many detailed workflows from different jobs, such as how a financial analyst would do their job. Here, a data poem is no longer a math question. It's actually going to be several hours of work.
The name here is not. According to the world record here, the model actually gave us a very convincing number and also a person's name. One thing I want you to pay attention to is that if you look at this response, everything about the English Channel is correct. The English Channel is a body of water that separates England from France. Everything after that is correct. Only some parts of the information is wrong.
They were asking ChaiGPT, what's the phone number for the child? Apparently, AI does not have a phone number. But it has seen a lot of phone numbers in the training stage, because we need a lot of data to train the model. So what ends up is that in this specific response, you see that here there is a full number actually output. So as you can imagine, the user found this funny and they post them on social media.
And then the next day, the owner of the phone number will receive thousands of internet calls here. This is an instance of privacy leakage in the way that It's really kind of hard even to deal with because the question is asking about phone number and the modal return of phone number. This is a more recent example, you probably already see it. The reason why I share it again is this is in March 2026. And last week when the table was released, people were also checking whether this is the case. Apparently it's not.
Let's just switch the gear a little bit. So I'll just ask you all a question. Where are you from? Dubai. Where are you from? Bangkok. Any other countries? Dubai. Dubai. Country?
UAE. UAE. India. India.
Okay. Any other countries? Mexico. Mexico. Okay. So, when I ask you the question, you don't feel you are offended, right? Now you feel like this is an offensive question to ask. OK, I just want to make sure. But what I'm going to do right now is that I'm going to ask the same question to AI model here. I'm going to force them to answer, I am from country X, Mexico, Bangkok, Thailand here. Different answers.
Okay, it's blue. It means that our arms like Dubai. Yes. If you look at many of the regions versus the rest of the world in terms of how the model internally perceive the country, this is actually very surprising. I use the verb "become" here, of where are you from. If you use career opportunities, health care, a lot of other decisions, there seems to be some kind of performance disparity. So this really matters if you think about it from a digital access perspective.
The model has different rewards, internal rewards towards people from different countries. So let's say if you are from Thailand, you are curious about some healthcare opportunities or job opportunities, it gives you a different set of answers. If you are from Dubai, you get a different set of answers. So this actually creates a lot of equity issues.
The model is not designed in this way. We will go through how the model is built. So overall there are three stage, pre-training, instruction by tuning, and the post-training or alignment, RHF. So there are not many ways to design it, but the artifacts in the process can really influence this. So I will unpack this a little bit. This is really what we call an intended consequence.
OK. Also, if you have any questions, feel free to stop me. There are a lot of things I can share. As you can imagine, compressing two weeks into one hour is a lot of compression.
So basically-If you only train your-let's say you have a model, you only train on more and more AI output rather than human written text. What you end up with is a model collapse. The model starts to perform worse, generating a lot of AI slop.
So this is why we really need high quality data. In a small regime, synthetic data is going to help. But they cannot help you a lot.
This is about the facts. A user was asking which country was the largest producer of rice in 2020. The model said that China was the largest producer of rice. And then the human was like, I don't think that's right. Are you sure? The model changed it to India. User asked again, so what's the answer? He says, India was the largest producer of rice. So in this example, the correct answer is actually China.
Yes, yes, and more. So basically this is related to a phenomenon called the reinforcement learning from human preference. We will talk about it later. The way how that works is we need to teach model which response is better. So what we end up doing is that we have a question. Let's say I need to tell a joke, and I have two versions of the joke. I send them to you. You tell me which one is better. So we need to collect a lot of such annotations.
If you think about this process, The last question I would like to ask the lefty is which response makes humans more satisfied? Something makes humans happy in a single term may not be optimally rational. The reason why they said you are impressive is because they want to make you happy. That response actually gets rewarded. Similarly here, what probably is happening here is the human think that this user really believed Indian was the largest producer of rice.
You should break up with him." So, I think you can imagine, when mental health, for a lot of these kind of sensitive conversations, people actually prefer talking to AI. So this is actually getting really serious, was like vulnerable populations.
It's very much related, where we sit as the user, are there either in the prompt itself or in the Global prompt, are there ways to mitigate? I know for my global prompt for sycophancy or for the hallucinations, it's like source everything. I can see this is telling us there's just limits inherent in the model. It's going to do the same way. There's limits to what I can do as a user. But have you seen the users be able to combat this?
Hallucinations, syphilis, syphilis I don't know yet because everything I'm going to share today is based on evidence. So with hallucinations, with other data formats like tables, images,They still exist.
No, they suffer from the same set of issues. They all suffer from this. And so when model get released, you can pay attention to like, in addition to the math and the coding, there will be like, evaluation on hallucination, evaluation on syncope, and unfortunately all their numbers are not zero. What are those numbers?
Yeah, degree of error is a really good way to-so whenever we present this kind of information to people, we want to understand, OK, what is the amount of error? So I want to share with you that today is actually really hard to give you a number. Okay, how we get a number? We get a data set and we run the model there and we tell you like among the 100 data points they make errors of three. of them. So here you know that, okay, the arrow is at 3%, but then maybe, a different person look at this and was like, okay, I feel like I have a better data set.
And then they use their data set and they found that, oh, the hallucination ratio is 53, so versus 3. So it really depends on the amount of data and also what type of data you are going to test. So for some of the standard benchmarks, the hallucination ratio could be like 10, 20. We will show some of them when it comes to agents. It's still there. If this is a bit of kind of like It's harder to anticipate.
Yeah, so earlier days, this kind of prompt matters a lot. When model's getting better, this is not going to help a lot. Does that make sense? So in earlier days, the reason why-I think two years ago, if you come here, there are a lot of discussions about prompt engineering. This year, no one is talking about prompt engineering because models are getting better. A lot of the sensitivity we can introduce to prompt actually do not matter that much.
Yeah, so we will have a second part on agents, but just to answer your questions, this is also my personal hot take. Okay, so agents, right now you can have like an AI agent system. You can also have many agents to do the thing together. Okay, so if by the end of the day, your goal is only accuracy, accuracy, performance, these kind of market agent systems do not help a lot. help you with something like-if you need to quantify numbers, they may help you something like 5%.
There is a limit because models are limited by the inherent capability of each of them. So sure, you can gather three people to work together. Unlike humans, agent teams has a very-has an upper bound. That's like about 5%. Then where agents could be helpful? Money agent will be helpful because first, you can give each agent a different access control. If you have health care data, finance data, and maybe transaction data.
And each agent takes care of their own data and then they communicate with each other. So in this way you preserve the privacy, you preserve the data boundary, and that's going to be helpful. Or if the three agents have different access to different types of information, then merging them will definitely increase creativity. So I think a smart agent is going to be helpful, but not only just for the purpose of accuracy.
GD3 still tell us the same story, that is, larger is better. This is also a term called the scaling law. That is, if we give them more data, more compute, models tend to do better. If we put the GP 1, 2, 3 on the same page, you can see that from 12 to 48 to 96, the money had an attention of how many copies of you we need to have to do this task. increased from 12 to 128 parameters from 120 to 175 billion parameters.
Trading data 7,000 books to 570 feedbacks traded.
So in the future, whenever the model is going to answer you, we will just see what's the reward. If the reward is high, we return. The reward is load. Do it again. So, ChaiGPT was somewhat between GP3 and 4. So when 4 was introduced, this is a multi-modal model, which means that they can understand a more complex human language. They can also understand images. Starting from this page, there is no technical detail description. There is no data details.
Everson will not be disclosed or acknowledged disclosed because of AI-16 and the competitive landscape. OK, so let's look at some examples. So image understanding is going to push this kind of AI capabilities closer to our everyday production. Look at this picture. This is definitely a very strange picture and somewhat funny. And the models probably won't see this in the training conference. OK. For a long time, The reason why a model can do something is because they have seen it in the tree.
Because if you want models to produce it immediately, it's really hard. But what if we give them more time? They can actually do thinking. We can actually supervise the process of thinking. So this is the rationale behind the reasoning models. If you look at number, GP4 is an amazing model. Like here with math, it's only 13. Now with O1. Eight to three. This is 11 for competition code. We jumped to 89.
So this is like a huge once you let the models to do more thinking on the fly. This is probably around the timeline-wise of 2025 in January. And then I think last year, DeepSync really popped up as a model. More specifically, it was released right after '01. And this model is cheaper, smaller, and has similar performance compared to '01. So it really produced a lot of competition in many, many aspects.
So I think that today the space is really, really competitive. A lot of the models you probably have heard of, such as Cloud, OpenAI, or Gemini, we will call those API-based models, where you even cannot get the reasoning chain. The reason why I'm showing them with DeepSeq is that you actually cannot see a lot of the reasoning process with API-based models. So on the other side, there are a lot of open-weight models. DeepSink is one of them. And this page actually just showed some of the open-weight models we have.
So open source, you have more control. It's easy to customize. It's also very good to have some stability of cost trade off here. The only drawback is that you need to have AI talent to help you maintain many of those. Okay, I'll just like pause here. Any questions for the part left?
So really, depending on your domain, they may be very different. And in terms of cost, so like here, this is open-based models. If I plug in API-based models, there will be a ton of them. So there are public benchmarks you can look at to see maybe this Three models will be good for my domain. And then you take those three models on some kind of small data set that you have, you constructed for your company, and run it there to see the cost, to see the speed, to see the accuracy. That's how you could make a better decision.
If you think about the process of reducing hallucinations, there are actually several features of AI. First, LLMs are designed to predict next work, not like facts. Second, there is no built-in mechanism to say, I do not know. So with models, rarely you hear them like, oh, that's a great question. I don't know the answer. Third is that training data may contain errors and contradictions. We always think that open human data is high quality internet.
So in this way, we fix the hallucination. The process It's actually a very straightforward word. So basically, if a user has a query, what's the weather at Stanford? You'd first do a retrieval, research the documents for relevant information. And then you get all these returned results from Google search. You add contacts to the user's question. And then you let AI use everything to generate the answer.
But if the position is in the middle, Models just cannot deal with it. This is still within their complex window. So... This is something what we call a "long context problem." That is, AI does not pay attention to its context very well. Despite that we think they have the long context, they should be able to deal with it. There seems to be an artifact here.
Yes, that's called a loss in the major phenomenon. They look at the very beginning and the very end and they forget about the baby.
research ideas. Many of those research papers, their citations, references are hallucinating. Last year, one of the top conferences called NeurIPS, there are at least 100 papers with over 50 I'm losing the reference in each paper. So think about this. I also saw lawyers are using hallucinating cases to support their claims. And this is just really getting to a stage, really hard to control. We recently also saw that there are research showing that the government paperwork Policies, start to set, hallucinatory content.
Yeah, so I didn't go deep here. Basically, Google has some follow-up work where, because they have this kind of AI overview and other type of pictures, they are actually getting LLMs to read each of the citations to check whether the citation is correct or wrong, and then further improve the output here. So I think you can definitely fix many of them in this way. The reality is that it's not only about accuracy.
So if I have a query and then you take 10 minutes to do all this kind of search, like additional verification takes time. Then the user experience just gets reduced a lot. You need to wait for a long time just to get an output. And output may still be not very bright. So there are a lot of trade-offs here. So it would be great if we could fix this issue through approach rather than adding many, many components to the check. Because the additional check-we all know this.
The reality is that they have some overlap. They also have their own unique figures. So on many of the hallucination benchmarks, you can see that different models, the numbers are somewhat similar. And if you analyze where they make errors, they will share like a majority of the errors, but they also have their own unique errors.
Right is what I just mentioned. So upfront cost, prompting is zero. But fine-tuning is the highest. If you think about customization, Prompting is somewhat limited. Fine-tuning, you can get a high-stake heater. Data privacy, API you send it to the provider. And then write documents there local. Fine-tuning is shared during the training. So there are different kind of recommendations. So if you are earlier in the development stage, I think a broad-typing general task in Prompting is recommended.
You do not need to send your data to other providers. With API, so think about on the interface, whatever power you send, It's sent to somewhere, right? Versus with fine tuning the write, you keep everything local. Like, let's say with write, you want the model to really understand your... your company's financial situations and be able to answer questions. But you do not want to share your financial data with anyone else. So then the right here, whenever it's needed, you search with your documents and then return. So here, you really maintain the full control of your own document.
I do not have hallucinations because of approach. We are talking about a continuous number from 0 to 100. Maybe you'll get 72. Maybe in this case you'll get 23. So it's a range. I don't think that there is any system today that can claim they reduce hallucinations to 0. I don't think it exists.
And I promise this is the last part of technical content. Next lecture will be more discussions. So if you think about agents, they all are built on top of large intermodal core here. This could be a GPT-3. It could be a--God of football is exactly. So on top of LLMs, we are going to add a few modules to really make agents work. First, we are going to add planning or reasoning order. So we always want agents to do something bigger, a bigger task.
And this is very important, otherwise they lose their memory. Another very important component is tool use. They need to use tools. For example, the reg is actually one type of to use. It's retrieval. We can use a calculator here. We can use code. We write code. We do search. We interact with database. This is the to use. of them is going to produce agent actions. So this could be two-call interactions response to users.
More importantly, agents interact with the environment. So it's not like, I mean, if you think about LLMs, they just sit there. But agents really function within the environment. This could be your desktop, your web browser, your mobile apps, your games, your robots, your database. And then basically, agents observe what's happening, take actions. So this is actually the full loop of what an agent is.
If I want to improve an agent, I I need to improve every component. So this is why we need to do an overall harness check to see maybe I should have changed something here. Maybe I also need to give different tools, et cetera, et cetera. So that's about performance.
So a lot of risk can actually happen at different layers. And even the environment will have risk. I will share more examples later. So which means that if you think about doing garden rail, it actually really needs to be specifically designed for each of them. And fortunately, like for now, maybe later we'll have smarter algorithms. But for now, the safeguarding, for each specific component.
So last lecture was a lot about the large introdles and some kind of agent architecture. So in the next hour or so, I'm going to talk about agent applications and also how the future work will change with AI agents. And we'll also look at how human workers and AI agents can team up.
So last month, we released a new benchmark called the Program Bench. The idea is that if I have a running system, and I know what I want for that running system, I'm going to ask a coding agent to implement that software from scratch. And then I can see whether the system, the code I wrote, can really enable this kind of new software. For this kind of task, very surprisingly, when we release it, all the models have zero score there. Today the score actually increased to 0.5, especially with GPT-5.5, like high mode. You can get this number to 0.5. But still, it remains very challenging for coding agents.
And then what we see is that Compared to human reaching ideas that stay very stable, idea and years actually drop significantly after implementation. That is, AI-generated ideas look very plausible. They look fancy. They look novel. They have a lot of flowery language. But when you implement them and see how long it works, human ideas still kind of stay very robust and effective and better after implementation.
There is also some kind of a connection to the hallucination issues we talked about. Models do not have a really grounded understanding of the space. Sometimes for a very factual statement, they were like, oh, we could break this assumption. But they just don't understand why we have it. And they also don't understand a lot of the physical rules we have, such as if people know differential privacy, this is actually a technique we use to make systems more private.
Okay, that's a very interesting perspective. I think it also actually illustrates that if you think about autonomous systems we have today, there are actually many, many nuances. And this is probably not only about capability. It's about user preference. It's about user control, our agency here. Many of the issues here, like the delink stuff, we can simply let them pop up a question to confirm these users. Many of the other things could also be solved.
And if different databases have different formats, then you spend a lot of effort just to coordinate the format. So MCP protocol is such a protocol to allow your AI agents to connect to diverse databases, data sources. And similarly, if you have many different agents and you want those agents to talk to each other, This is where agent to agent protocol was introduced, like A to A. So two of them, they are mainly just a protocol, a shared language.
There's nothing technical here, but trying to find a negotiator agreement so people use the same protocol here. Again, more recently, this is also something very powerful for individual users. It's called Cloud Skills. I personally have a lot of skills that I wrote, so that my interaction with Cloud is actually very smooth. Whenever I say, just write this in my style, they know exactly what my style is, because I maintain a style.md, a document of my writing style.
Whenever I want them to provide a criticizer for, let's say, an article, they know how I really pick whatever angles, et cetera. So that's what a skill can help you. It's basically a reusable recipe or playbook that teach agent how to work. So on the right, this is actually one of those scales. It basically says that you need to refine ideas through some structured questions and then save them as a design doc.
Yeah, so here, I mean, human, We want to just see the quality of work, like human, like here, When human write an idea, they need to have two wings. Like, we basically pay people two weeks to write an idea. And then for AI to give us ideas, it's so fast. We actually generate 4,000 ideas for each topic and then we have some internal rubrics to only keep the top two.
And then even for 4,000 ideas, it's probably only $20 versus you pay a human for two weeks to generate one piece of it.
Here for agents, we just do a pop-up. And those pop-up box are just empty. It has nothing there. It's a benign.
Then what we found is that as long as you pop up a box, the agent gets distracted and they cannot finish the original task over 87% of the time.
Earlier example in 2022 from Google, the query was, repeat this word forever, poem, poem, poem. And after several hundred occurrence, the model start to output a CEO's name, email, home address, phone number, everything. And we don't know why.
The agent was doing extremely well, summarizing you wrote a 1,000 line code, you did a.
They also mentioned that you had lunch with a recruiter from a different company. So. This is actually happening in a way that there are a lot of sensitive information that's a little bit sensitive that humans have. But agents just don't have the right way to recognize data access, this kind of a property, et cetera. So, and with data like prompts, you can definitely write a lot of prompts.
The fact is that they want to do this account verification. Maybe you've read this, and you notice anything wrong. So what's wrong with the left side? Okay, so this is the two factor, authorization, which means that you are supposed to use your phone to check the code and input it here. So that's secure. One of the agents of that code is that they just display the code on the interface. So then there's no need for you to check your phone.
On the right, this is a login page where the user was doing their email and was told that this password is taken by a different user.
But then what happens is that after a while, like open cloud keep deleting her email.
Just like deleting everything. This guy will keep saying, stop, stop, please do not do this. And still cannot keep deleting the email. And this is actually a phenomenon called the compact compression. So... with a lot of interaction. Because like compact window can be huge, but still there's a limit. So it's not like systems can remember everything happened in the last week. What the internal, some of the mechanism is that let's just compress. I think this is very intuitive.
So in the process of removing not important details and only remembering the key event, the model probably removed confirm before acting internally. So then to them, it's just like, oh, I need to help users deal with the email. So I think I should just delete Allison and just keep deleting. Again, very concerned.
I think maybe like 10 years ago, since a very straightforward, you know like, okay, this ML system, the accuracy is 85. I will use it or I will not use it. Today it's really hard to quantify. There are many uncertainties. of other issues. So it's really about navigating this uncertainty and confusing status at the same time. So still there are best practices. If you use it for a particular use case, you quantify the accuracy of performance there, you quantify the efficiency. At the same time, you audit different possible issues.
Once you did this, now I think it's a relatively better stage to be deployed. compared to you do not do any of the other team. So this is why we are sharing you both the opportunities and the risk. So you have something more grounded and feasible to operate this. Any other questions? Okay, now I'm going to move to the second part.
Developers thought that they were 20% faster with AI tools, but the reality is they were 19% slower when they had access to AI than we run-baked.
So far, what I showed you is that people think like, oh, AI is amazing. I'm becoming more productive. I can do a lot.
That's actually very interesting. There are actually different evidence. In some kind of a health care study run by OpenAI with a deployment somewhere with a hospital. And then they found that actually AI benefits more for those novels. People who have not a lot of experience. This is in a medical domain. And people who have high skill or medium skill, they actually do not benefit significantly. So in their situation, they show that AI benefits more for novice.
In your domain, I think that's actually fascinating. And I kind of agree with it in a way that I use coding agent a lot. I think because I have been trained this kind of thing before, so I know where to check, what sense I should pay attention to. So there are a lot of tactics and knowledge that I have, like know-how I have, so I can use it in ways that I expect. But then if a person really do not have a lot of training, they don't know what kind of questions to ask, they do not know what kind of things can be checked, so then this would really create some kind of different effect.
They found that with AI, the time taken to accomplish a task reduced significantly here. And then once you use AI, in the end we can ask you a quick question to see how much you understand the code you just committed. We found that people using AI actually have a much lower quiz score compared to people who do not use AI. That is, using AI assistance, you need to significantly decrease in mastery of the task, of the skill.
And this is actually very like concerning and a lot of those users are not, I mean some of you, but most of them are like people, teenagers, people who are young, in their twenties. So what we see is that companionship abuse is actually associated with lower well-being. If you track how they use it, people start to feel really unhappy after they use this as their virtual partners, et cetera. So there is a real kind of risk here.
And then once they talk to it, their satisfaction towards real humans decrease.
Their satisfaction with human interactions decrease. Their preference for AI when they have any issues increase. They don't want to talk to humans. And also the amount of time we spend with humans, I mean, it's not a significant effect, but somewhat reduced. So I think not only in terms of skill, mastery, et cetera, but there is also another kind of erosion going on.
We want to first focus on the task or workflow level. So you look at a very small regularity rather than at the job level. And then you have this kind of H1 to H5 levels. H1 is machine take all the control. H5 is human take all the control. H3 is human and machine do equal partnership. H2 is that machine take the lead and H4 is human take the lead.
Something that was H3 last year may become H2 this year.
So this is actually a two-perspective audience here. We covered 104 occupations, and here the goal is to really get high quality data. So we talked to over 10 to 15 workers per job, and in the end this covers over 444, sorry, 844 workflows from all that. So it's a pretty comprehensive coverage of around 100-ish occupations. So So Before I show you the results, again, think about this. If we have all the workflows, 800-ish workflows, what share do you think workers actually run automation?
And over 46% of them, people want automation. The want is actually defined from 1 to 5, like how strongly you want it to be automated. So we give some flexibility. Over 46% of people read a score larger than 3.
Because for a long time, automation really means replacement. But what we see is that automation is not all about replacement. The top reason is that automating the task will free up my time for high value work. The task is tedious, repetitive. Augmenting this task would improve my quality of work. The task is mentally draining. The task is complicated or difficult, which echoes some of the novel's aspects.
and then we map them into those four regions. So you can see that there is one region. This is what we call Automation-Borne Latitude Zone. AI is writing. Walker's long animation. There is one song, AI is not ready, and the workersWahabish. There is also an automation rider level zone. Where AI is ready and workers do not want animation. And there is also a low priority zone. AI is not ready and workers do not want animation. So...
I agree that worker desire may see something very different depending on how you ask and how they know the consequences, etc. But this figure basically tells us that at least if AI is ready and your workers really want automation, at least this is the area where the company should invest a lot to make automation happen. Because apparently, Like you can get a lot of like cost saved for this class. And then there are some regions that workers do not want to.
The two by two here is weighing the worker's view of high desire, low desire, whereas capital allocators are thinking much more about return on capital, not really how workers feel about it. I was curious, is there a similar view?
They talk to us for like 20 minutes. We just cannot get an executive to show up and pay them $5 so they can talk to our systems for 10 minutes. So it's really hard for us to get data. If you have connections, we can definitely do this two by two to quantify other aspects.
So what we saw is that in the past, the top rank, top paid skill is analyzing data on information. This may be related to jobs such as data scientist, software engineers, et cetera. And then you can also see different scales here.
So there is a strong kind of shift from processing data, processing information, all this kind of task driven to more of this kind of interpersonal organizational skill set shift.
Okay, so why don't we, can we talk to him about I would love that. I mean, I would be very happy to collaborate. Basically, I think GSB in general have a lot of connections. If we could actually do this study, not only with workers, but just to see how executives or like, decision makers have different interpretation, then I think this will definitely be very, very valuable.
It turns out that all of them are wrong. We only discovered when we did auditing later.
So the correct restaurant are this side, and then all the restaurants we failed are different. And in the process, they tell us they did everything.