When AI Outperforms Law Professors
What does it mean for quality, legal expertise, and judgment when law professors actually prefer AI answers over answers written by other law professors in a blind study?
In this episode of AI Sidebar, Irene Liu talks with Stanford Law Professor Julian Nyarko, faculty director of liftlab and co-director of Stanford Law’s AI Initiative, about his recent study examining how legal experts evaluate AI-generated work. Nyarko explains what the study reveals, why its findings matter for both legal education and practice, and how rigorous research can help the profession determine when AI performs well, when it falls short, and when it should be trusted. He also discusses how liftlab is evaluating AI across legal services, research, and education, including when AI tutors can support student learning and when they may get in the way. Together, they explore how this research is shaping Stanford Law’s approach to AI in the classroom, from tutoring tools to simulated negotiations, and what it may mean for the future of legal expertise and education.

Transcript
[00:00:00] Irene Liu: Welcome to AI Sidebar. I’m Irene Liu, your host and executive director of the AI Initiative at Stanford Law School. Today, I’m delighted to welcome Professor Julian Nyarko, faculty director of liftlab at Stanford Law and co-director of Stanford Law’s AI Initiative. We’ll learn more about the liftlab, the lab’s focus, and how it is approaching some of the most important questions at the intersection of AI and law.
[00:00:26] We’ll also discuss Julian’s recent study, ‘Law Professors Prefer AI Over Peer Answers’, which found that law professors preferred AI-generated answers over their peers’. That finding raises fascinating questions about legal expertise, judgment, and the future of AI in law. So, with that, let’s jump into my conversation with Professor Julian Nyarko.
[00:00:51] Julian, thank you so much for joining the AI Sidebar.
[00:00:54] Julian Nyarko: Thanks for having me, Irene.
[00:00:55] Irene Liu: So, Julian, you’re the director of the Stanford Law’s liftlab, and the liftlab stands for Legal Innovation Through Frontier Technology Lab. And I’m glad you shortened it to liftlab, but can you tell us what is the liftlab and why it was created?
[00:01:09] What gap were you trying to fill?
[00:01:11] Julian Nyarko: Yeah. Irene, thanks, again, for having me and, yeah, for giving me a chance to talk about liftlab. Yeah. We really don’t need or want people to remember the longer- … phrase, so really liftlab is the acronym is all, all that’s needed. So, the idea for the lab is now roughly one and a half years old.
[00:01:28] And, the current executive director, Megan Ma, and I, we got together and we thought about what kind of experience we have at the inter- section of law and AI. And, what particularly struck us was that, you know, we were all talking to these legal tech companies, law firms, and so on and so forth, and we noticed that everyone’s really, really excited about, the intersection of AI and law, but no one really knew what works.
[00:01:54] It’s getting a little bit better, but there’s still a big knowledge gap. And, that there was hardly any research that sort of looks at AI specifically targeted at the legal services sector, right? And so, we thought, it would be a great intervention to bring more rigor and to bring, you know, academic research exactly to this question of how can we…
[00:02:15] You know, now that AI makes all these great developments and, you know, charges our productivity and so on and so forth, how can we really make sure we are, for instance, evaluating these models in a way that is beneficial to the legal services sector? How can, where and what areas should we trust it, and in what areas should we be more skeptical?
[00:02:34] And, yeah, that’s how the idea for liftlab came to be. And, we had a lot of fun sort of expanding since then.
[00:02:41] Irene Liu: And you’ve created four core pillars around the liftlab.
[00:02:43] Julian Nyarko: That’s right.
[00:02:44] Irene Liu: So, what are those four core pillars, and what’s the pillar that excites you the most?
[00:02:48] Julian Nyarko: Yeah. So, the first pillar that we have in the lab is centered around evaluation.
[00:02:52] And so, this sort of directly pertains to, what I just mentioned, which is we wanna really know the extent to which models work for the legal services sector to the extent to which they help lawyers be better lawyers. And, what’s interesting, about the evaluations track is, you know, that it surfaces really deep questions about how we’re thinking about the law.
[00:03:16] So, just to give you an example, now we have models that are pretty good at writing different agreements, right, for contracts. But now we have to answer the question, what kind of contract want, want we have it write for us, right? And in order to answer that question, we have to answer the question, what is a good contract?
[00:03:32] I teach contracts, and, you know, after the fall quarter, 13 weeks of, learning about contracts, it turns out we sit there at the end, and if I would ask my students, you know, “What is a good contract?” I probably wouldn’t get a, an answer. And in part that is because even academics aren’t, aligned on what a good contract is supposed to be.
[00:03:52] What kind of risks should it address? How should risks be shifted? And so on and so forth. And so, if you wanna go about evaluating a language model on whether it drafted you a good contract, it brings you back to this question of, you know, what is quality in the legal services sector? What is a good deposition, and so on and so forth.
[00:04:08] So, this is particularly exciting. We’ve done different projects in that vein. One that, is very recent is called Judgment Bench. So, here the idea is that if researchers or practitioners go about evaluating models, there are actually two different ways of doing so.
[00:04:24] One is you come up with a rubric. You say, you know, and, “Have an LLM draft a memo for me. Does the memo mention XYZ?” And you just score it. And the other is called, preference ranks. So, there you put two outputs in front of a human expert, and you say, “Which one do you prefer?” Without giving them any rubric, just general notions of quality. Which one do you find better?
[00:04:48] And, what’s interesting about this project is really, right, some areas are more like sort of hard sciences where there’s a, a clear answer. And so, for those, rubrics work very well. But then some, if I ask you, you know, what is creativity? You would probably… you know, which, which, how creative do you find this painting? You would probably not be able to come up with a rubric on what creativity means, right? But if I show you two paintings and I ask you which one’s more creative, you might have an answer, right? And you might, because you’re sort of accessing your intuitive notion of creativity.
[00:05:19] And so, what we were interested in particular is when we go about evaluating quality in law, are we gonna do better with the rubric because law is more like a science, right? Or is it more like, you know, more like an art, that judgment matters and these things that are sort of very difficult to formalize into a rubric.
[00:05:37] And what we found in that project was that at least, in the legal tasks we consider, which are sort of high value tasks in transactions and litigation, preference ranks are much, much better. Meaning in those instances, law might be, quality of law, legal output might be perceived more like an art, where sort of these fuzzy things matter. Okay, so now I talked for a while, and I only talked about one pillar. I’ll make the others a little quicker.
[00:06:00] Irene Liu: The other three pillars. What about the other pillars?
[00:06:02] Julian Nyarko: Yeah, the other three pillars. So, the second one is AI to enhance the quality of legal outputs. So, you know, the low-hanging fruit that, you know, much of the industry is currently working on is do what lawyers do, but faster and quicker, which is all well and good, but we think we’re particularly excited about applications where you can use AI to, you know, enable lawyers to do work that they’re not able to do without AI.
[00:06:27] Very quick example is, we have this project on contract risk assessment where we basically use new language models and so on and so forth to surface how contracts fail and why contracts end up in court. And you can imagine that with that structured information, you can feed it back into the contract writing process, and as an attorney writes a contract in a particular way that is particularly risky, a warning light might go on.
[00:06:54] The third pillar is legal education. So there, we’re asking about how AI might enhance how we conduct legal education, different forms of legal education it might enhance. And the fourth pillar is sort of a grab bag that is, methods, the methods pillar, where we are looking for particular ways in which AI is generally developed in the ordinary language domain, and we’re looking at how can we tweak models to make them really good for the legal services sector and work better for the legal services sector.
[00:07:22] And then also one of my other passions, the fairness work falls into that bucket. As to, what excites me in particular, you might have seen from how I allocated the time that the evaluations bucket at the moment is sort of where a a lot of our work falls into and, what is particularly thrilling to me.
[00:07:39] Irene Liu: And with the evaluations bucket, you mentioned that a lot of it seems like it’s falling into the range where law seems a bit like art. Would you view it that way for contracts as well? Like you said, there isn’t a consistent answer of what is a good contract, or is that what you’re still studying on that?
[00:07:54] Julian Nyarko: Yeah. So, right, if you think about it, and one perspective on contracts is that, you know, if a client hires you to write an agreement, like a definitive merger agreement or whatever, you should write it in a way such that you maximize the benefits of the client, right, in some amorphous sense. But it’s not really clear what that means.
[00:08:12] So for instance, there might be a certain risk that might materialize Is that risk really likely to happen? If so, we should probably negotiate hard on that particular point, right? And not concede, and so on and so forth. Is it a very minor remote risk that might not materialize? We might be able to be a little bit more lax and, you know, and concede.
[00:08:32] Now, it turns out that few people have hard data on how to measure contract risk. And so as long as we don’t have this hard data on, you know, what provisions are more risky, what provisions are, you know, prone to litigation, and so on and so forth, right? It really remains in, within the experience and the judgment of the drafters to assess, you know, what is worth pushing for, what is worth conceding on, how to maximize the pie for your own client.
[00:08:59] And so I think in much of law, even in contract law, we’re sort of currently living in that world where we don’t have the right or wrong answers to every question, we might ask. And so, a good lawyer is really one who uses their judgment to the benefit of their client in, in that way.
[00:09:14] Irene Liu: And it’s interesting you said there isn’t hard data around that. But I have to say, as a former general counsel, one of the things that we always negotiated is limitation of liability, the indemnity. Those are the two that we would constantly spend time on. But like you said, there’s probably an art to how you draft that to protect the client as well too.
[00:09:30] Julian Nyarko: Yeah, for sure. And so yeah, it should also be, you know, we might have data on sort of, you know, in, in the merger context that sort of everyone negotiates about the material adverse event provision, right? But we, you know, when you ask how many cases do you really have in which the definition of a material adverse event had a substantive change on the outcome of, you know, litigation proceedings? It’s a difficult question to answer, right?
[00:09:57] Because, you know, it’s a causal question, you know, did the MAE matter or not? And then many disputes settle before they reach the litigation stage, and so on and so forth. And sort of tracking that type of data at the contract level in particular is sort of what academics and contract law consider the holy grail, that you work with a partner closely to sort of track every single contract, the fate of every single contract, what settled, what issues were brought up, and so on and so forth in order to sort of bring more rigor to that.
[00:10:25] Irene Liu: Yeah. It’s so interesting. Your lab is studying AI and yet is working with AI to really understand how it affects the legal profession. So, with AI, I’m sure it’s changing how you’re doing research as well too, and there’s something almost recursive about it. So, can you tell me how you and your lab are using these tools?
[00:10:41] Julian Nyarko: Yeah, for sure. In some ways, what’s really interesting to see is that the lab itself is sort of a microcosm of what happens in law firms really.
[00:10:52] So for instance, when we’re thinking about integrating AI into legal practice and into the work that law firms do, there’s sort of this question about de-skilling, right? And how do we, you know, does AI sort of upend the apprenticeship model, and how do we train our future generation? We have very, very similar questions that we’re facing within the lab and within the research environment more generally.
[00:11:16] So for instance, now we have access to Claude Code, and Claude Code is hugely beneficial in being a more productive researcher. But at the same time, you need quite a bit of training and quite a bit of judgment to see what. And, and sort of you just need to be careful because, Claude Code makes all these decisions on the back end sometimes for you, and you need to be able to catch them or direct it correctly. So, in order to use Claude Code for research really effectively, you already need to have developed these research skills and this judgment, right? And so, we’re faced with very similar questions of how we train our lab members to sort of develop that judgment so that, you know, they can use Claude Code, responsibly.
[00:12:00] But yeah, more generally, yeah, we’re centering the lab around the use of AI, quite a bit. So, for instance, every new project I tackle now, I start to engage in a Socratic dialogue one or two hours long, with Claude, and it’s an extremely useful exercise. It’s sort of like talking to a very smart colleague who is not really a domain expert, but just a lot of brainpower.
[00:12:23] So, yeah, in that Socratic dialogue, it asks me a lot of questions that help me think about the project in a more structured way, and, it surfaces sort of many of the potential issues that usually in an empirical project really come to the forefront down the line when you get your hands on the data and everything gets, you get your hands dirty. Those kind of problems and aspects are surfaced during this Socratic dialogue, for instance. And, yeah, it’s a really, really helpful exercise to just, to help with project planning and project structure on the outset. And then there’s also in the implementation stage and so on and so forth.
[00:12:56] Irene Liu: It’s so interesting that you’re doing this Socratic dialogue with Claude Code. So, do you do it with any other models? And before Claude Code, were you using any other models to do the Socratic dialogue at all, or is this a more recent phenomenon for you?
[00:13:06] Julian Nyarko: Yeah, it’s really a more recent phenomenon. So, people have done more work on how AI can be responsibly integrated into the research process. So, there is Claude, Claude Code Skills, basically research skills, that exist that are really optimized for helping you think through the problem.
[00:13:25] And in the past, if I used LLMs, it was sort of more pointwise, you know, I had a specific question or specific methodological question that I wanted to bounce back and forth. So, I’ve been using it like that for a while, but sort of this new development is really at the project beginning, just ask me high-level questions about the project, what am I thinking, and then let’s try to drill down what exactly that means.
[00:13:47] Irene Liu: Hmm. It’s really cool to see professors also using it for Socratic dialogue by themselves for research purposes.
[00:13:54] But one of the things that I also wanted to just focus on is your recent paper and your recent research that was published. It was called “Law Professors Prefer AI Over Peer Answers”, which is, it sparked quite a lot of discussion, not only within legal academia, but the broader AI community. And the headline finding was so interesting that the law professors actually preferred AI-generated answers over answers written by other law professors in roughly seventy-five percent of blind, evaluations.
[00:14:23] So that is, it’s actually a really interesting finding. What motivated that study, and were you surprised?
[00:14:30] Julian Nyarko: Yeah. Yeah. So that study was a lot of fun to work on. The origins really start in summer twenty-twenty or fall twenty twenty-four. So, in the fall of twenty twenty-four, I gave my students in contracts access to an contracts tutor.
[00:14:50] But because LLMs were really, there were a lot of questions around LLMs back then. What we had to do was basically monitor every interaction that students had with that tutor, and we had to jump in and correct whenever the tutor would say something wrong, which happened, you know, not that infrequently.
[00:15:09] And then in the summer after I was done, in the summer of twenty twenty-five, I thought, in the coming fall, I might wanna give them access to a tutor again. But really, this idea that we constantly monitor what students do and constantly jump in, that’s not scalable, and no one really likes that system. But maybe LLMs are now at a point where we can responsibly release them, and they’re actually doing a good job. And so, what we thought was, let’s evaluate that rigorously.
[00:15:35] So then, what we did in this project was we recruited sixteen, law professors all teaching contracts from the same book. And we did a three-round sort of process with them. In round one, they wrote 40 questions that students would ask them after class or during office hours, so just shorter questions about the material. In round two, everyone answered each other’s questions. And in round three, we showed two answers to a question to a professor and said, “Which one do you prefer?” And one of those answers was LLM-generated, one was human-generated. They didn’t know which was which.
[00:16:07] And so, yeah, we did that multiple times. And as you said, the models, back then it was Gemini 2.5 Pro, it was on par with the very best instructor in this study, and, won on average 75% of the time, which was honestly very surprising to us. We thought LLMs would probably rank somewhere in the top third, but it was striking how clear the results were.
[00:16:30] Irene Liu: So, what was AI doing particularly well in that case? Was it because they sounded more confident? Or, I mean, is it some attributes of AI that made it more preferable to law professor answers? Or what, what do you think made that result skew in that favor?
[00:16:47] Julian Nyarko: One, one thing we did was we looked at whether the, an AI advantage can be explained only by looking at sort of more superficial textual features.
[00:16:55] So people say, right, “AI sounds confident,” or “AI gives longer answers,” because we, we wanted to emulate a realistic setting, so we asked instructors not to do any background research. This is really, you know, you’re in office hours, a student asks you a question, or after class a student asks you a question.
[00:17:10] So, and, and they had a more constrained time budget. And was it because they write longer, or was it some scaffolding issues? But we were partially able to explain the LLM advantage this way, but not fully. There was still a sizable gap after you take all these textual features into account.
[00:17:27] And so that suggests that really the content itself drove some of the, or, or a large part of the LLM advantage. And, you know, my own reading of it qualitatively, I would say LLM answers were more complete very often. So, they would go into, you know, edge cases, and they would consider exceptions of exceptions. Whereas instructor answers were quite often straight to the point and were not as complete or,
[00:17:55] Irene Liu: comprehensive or colorful.
[00:17:56] Julian Nyarko: Comprehensive. Yes, yes. Yeah. As comprehensive as the LLM answers.
[00:18:01] Irene Liu: Interesting. Were there any professors that skewed in one direction more so than others? Like, that they knew one was AI, or like were, were there any that were able to tell in some way?
[00:18:12] Julian Nyarko: It’s hard for us to say. Now, I will say that every instructor who participated on average preferred LLMs over humans.
[00:18:18] So to the extent that they were able to identify AI and to the extent that they had an anti-AI bias, it wasn’t enough to actually, beat humans. none of the instructors told us that they were able to to, to identify AI answers. And before we showed the AI answers, we did some sometimes very light editing if the AI would say, “I am just an LLM. I’m not a legal professional. You should hire a lawyer, but here’s my answer.” Then we would delete sort of-
[00:18:45] Irene Liu: Sure, yeah
[00:18:45] Julian Nyarko: … the sort of short introduction. Yeah, so, we are not aware, of any circumstances under which, instructors would’ve been able to tell, but you can never totally exclude that I guess.
[00:18:56] Irene Liu: Yeah. So, with the 16-professor participating, based on the findings, are they implementing anything differently in their contracts classes as a result? Are they integrating more LLMs as a result of this, and are you?
[00:19:07] Julian Nyarko: So, I think many participants did not expect that the LLMs would come anywhere close to a win rate that is on par with professors.
[00:19:18] And so, I do think there was afterwards quite a bit of interest to adopt AI tutors, and we made certain suggestions, on how to do this effectively. We’ll have to see after the summer who actually implements it.
[00:19:32] Irene Liu: Yeah.
[00:19:33] Julian Nyarko: I am certainly, it’s reassuring to me that the LLMs provide high, high-quality responses, and I am, integrating it into the curriculum. I will say, though, that we still have to be mindful about how we’re integrating AI tutors, so this was really just answering short questions about contract law. And I do think there is probably very little downside to, making AI contract tutors available to students in that way.
[00:19:59] There is some research that shows that, so for instance, one study that I like very much, in that study, they gave students access. So, the, in the treated group, they give students access to AI in problem one, and in the control group, they didn’t give students access to AI for problem one. And then for problem two, no one had access to AI. And what they found was that students who initially had access to AI in problem one gave up sooner in problem two and were less likely to complete, and if they completed, it was lower quality.
[00:20:30] Irene Liu: Interesting.
[00:20:31] Julian Nyarko: And so, this speaks a little bit to this concern that many have that, you know, in education there can really be a productive struggle, and you don’t always want a clear answer to every question that you might have immediately, right? Or at least, in homework assignments. The, so the, the, the, the findings are most clear for homework assignments. Homework assignments, you don’t want just a, you know, an answer but that does the homework automatically for students because that probably hurts them downstream.
[00:20:58] But there are also findings that suggest it can be very helpful, in learning. And so, I think these short answer scenarios are exactly sort of the type of scenarios where getting an answer to your question can be very, very beneficial, whereas, you know, if I give, I have these optional homework assignments, I certainly would not suggest people use, large language models to, go through those.
[00:21:20] Irene Liu: It’s so interesting because when Google was introduced, people said there’s a Google effect where people have stopped really memorizing how to find directions because they’re leveraging Google Maps. And if they can find answers on Google, why bother memorizing any of the answers at all? And so, it’ll be interesting to see if there’s some sort of an AI effect from a learning perspective
[00:21:37] Julian Nyarko: being entirely unable to navigate after Google Maps effect is something I, I personally feel every day. So, I’m so reliant on,
[00:21:45] Irene Liu: Yeah, me too
[00:21:45] Julian Nyarko: … navigation systems at this point that, yeah, it will be impossible for me to find the way myself, so.
[00:21:50] Irene Liu: And phone numbers.
[00:21:51] Julian Nyarko: And phone numbers, yes.
[00:21:52] Irene Liu: Yeah. So, given that this was a study of contracts, do you think it translates to other subject areas, too? Do you think other professors of torts or other subject areas might have a similar result, or do you think it’s unique to contracts?
[00:22:05] Julian Nyarko: I don’t think there was much in the study that is particularly unique to contracts in the sense that, you know, we had different types of questions in there. Recall doctrine, recall cases, hypos, and policy questions. And really at a high level, we found that the element windage was consistent across all these types of questions, which then leads me to believe that the findings that we have extrapolate really well to different areas.
[00:22:32] Now, I will say that I have a colleague, Lisa Ouellette, who did another study on contracts tutors a while ago, in IP. And in IP, in her context, the LLMs did not perform quite as well. So, there’s definitely something to dig into, but there were also other differences. So, when there are differences in study designs, they can lead to, somewhat, somewhat different results. But yeah, I would think definitely more research is needed, but I’m cautiously optimistic that the results extrapolate, across many areas of law.
[00:23:02] Irene Liu: And if you were to do it again, would you use different models or what, what would you change?
[00:23:07] Julian Nyarko: Yeah, so one thing that is particularly interesting is whether there is an impact on learning outcomes for students. And so, if I would do another study with LLM tutors, I think one interesting thing to do is to basically have different classrooms where some students have access to AI tutors and some students do not have access to AI tutors, and then later see whether we actually see any performance gains because it helps us directly answer this question of how good is this actually for learning outcomes, right?
[00:23:39] I think that would be very helpful in a new iteration. Now, there’s also, you know, the question of should we use newer models? So, in our study, actually we, so, because it was done last summer, Gemini 2.5 Pro was the leading model
[00:23:52] Irene Liu: Yeah
[00:23:52] Julian Nyarko: … back then. But, you know, things develop so quickly, and they’re very quickly outdated. So, what we did in our study was humans evaluated Gemini 2.5 Pro and NotebookLM running on Gemini 2.5 Pro against the human answers. But what we then also did was we added newer models, including Opus 4.7, GPT 5.4, even though those are outdated now too. But because we couldn’t ask humans to evaluate them again, we had an LLM as a judge.
[00:24:19] So, an LLM saw two, answers pairwise and then decided which one it prefers. So, we made sure that that LLM is consistent in its judgment with our human judges. And that allowed us to then, extrapolate or, or extend our results to, sort of frontier newer models. And, you know, quite consistently what you see is the better the models get, the bigger the gap between humans and AI.
[00:24:43] Irene Liu: So interesting. So, if you were to take this study and, take it further into the classroom, how would you change your teaching style? I think you mentioned that you might change the study a bit to incorporate students, but how would it change teaching, your teaching and legal education, in general?
[00:25:00] Julian Nyarko: Yeah. So, this study, I think the impact of this particular study are more modest because, you know, I, I don’t wanna directly jump beyond what is in the sort of four corners of the paper, right? And so, in the four corners of the paper, it’s really LLMs for answering short answers that students have. But beyond that, I think what’s particularly exciting to me is that AI provides many opportunities to actually positively impact, legal education.
[00:25:28] So for instance, what we’re looking at particularly is, it can be extremely helpful for students to go through simulation exercise, right? We, whether that, depositions or whether that is negotiations. But we are currently limited really by the fact that any negotiation students want to conduct in the classroom is really a logistical endeavor. You need two sides-
[00:25:51] Irene Liu: yeah
[00:25:51] Julian Nyarko: … you need to coordinate their schedules, and so on and so forth. So, in most negotiation classes, people might do four, five, six simulated negotiations. But you can imagine that with AI, you can actually do, you know, you can really significantly scale that up, and you could have students engage in two simulated negotiations every week, right? And an instructor could read the transcripts and give students feedback on what they do well, and they can test different things. Right? And because the LLM is relatively constant, they can test different strategies, see what works better, what doesn’t work better.
[00:26:21] And so, what particularly excites me for, in AI for legal education is sort of thinking creatively about how can we use that, right?
[00:26:28] And, you know, another example is in our lab we have developed this M&A negotiation simulator. And law students and junior attorneys, they will never conduct an M&A negotiation themselves, right? That would be irresponsible. But with AI, they can potentially, right, engage in a simulated exercise, and then already develop sort of an understanding of, you know, what matters in such a negotiation, what pushback might, might look like, right?
[00:26:53] And so, as we’re talking about sort of how do we develop judgment in our junior attorneys, right? Maybe if we think creatively, there’s a more direct way actually to try to develop that judgment than, than we have traditionally done.
[00:27:07] Irene Liu: Yeah. A lot of times I hear a lot of partners say, “You know, how do we teach judgment to these young associates without all the toil that we went through?” but part of that toil is going through grunt work and like reams of documents and watching partners negotiate. But really, you could do a lot of that now with AI without having to face that toil.
[00:27:28] Julian Nyarko: Yeah. Yeah. So, I, I think this is, exactly right and, you know, a particularly exciting research opportunity, right?
[00:27:35] Yeah. I, I wouldn’t go as far as, not at all go as far as saying we have figured it out already, right? These simulated negotiations and so on and so forth, what, what is the downstream impact? We need to measure all that, and we need to conduct long-term studies.
[00:27:47] But, while I understand and agree with the fact that there are many risks in, in, you know, incorporating AI and that many of the traditional sort of structures might break away, I think there is really this opportunity here that, you know, collectively we need to think hard about whether there are, you know, better ways of doing exactly what we used to do. Because, you know, even the, you know, senior partners that went through all this roadwork, right? I think they agree that it wasn’t an efficient way to develop judgment-
[00:28:15] Irene Liu: Yeah
[00:28:15] Julian Nyarko: … it’s, it was just the only thing we had. And so, maybe we’re able, we’ll be able to come up with, with more efficient ways now.
[00:28:21] Irene Liu: Yeah. And I think clients will love that as well too, especially at law firms. So, we’ve covered everything from legal education to expert judgment. But before we wrap up, I just wanted to ask you a few quick questions.
[00:28:33] Julian Nyarko: Sure.
[00:28:33] Irene Liu: So obviously your study, what was so interesting is that people were really mesmerized to a certain degree about how LLMs actually and AI outperformed professors. And so, there was some sort of an ah-ha moment for people where they realized, wow, AI is like really reaching this level of performance that was underrated. So, what, what, what is one legal task where you think AI is still currently underrated?
[00:29:00] Julian Nyarko: Yeah, one nice example I think of such a task is contract drafting, in the following way. It has always, or for a couple of months now, or maybe over a year now, we knew that LLMs are decent at writing simple agreements.
[00:29:18] So if you say, “Write me a non-disclosure agreement,” right? It can do that, and it looks pretty fine. But there was always this thought that it’s not good at writing you a 50-page agreement, right? And, if you go to Claude and you say, “Hey, here’s a term sheet. Write me a charter based on this term sheet,” or, “Write me an SPA based on this term sheet,” you’ll get a 50, 60, 70-page document, but then you hand it to an expert and that expert will say, “Mm, actually, you know, there are many, many mistakes here.”
[00:29:47] So, we did exactly that exercise in contract drafting. So, we had a Claude Opus draft agreements based on the term sheet, and exactly, you know, what I just described happened. I’m not an expert in venture financing, but the output looked fine to me even, right?
[00:30:03] Irene Liu: Yeah.
[00:30:03] Julian Nyarko: And then I handed it to my colleague who’s a venture financing expert and he said, “No, no, this charter is total, total trash, basically,” right?
[00:30:10] And so, this is sort of this, this dangerous area, right, that people are concerned about, where AI looks good on the surface, but the substance is not good. However, this question was about underrated-
[00:30:21] Irene Liu: Yeah
[00:30:21] Julian Nyarko: … exercise, right? And so, what we then did was we took the NVCA model documents.
[00:30:27] Those documents are pretty detailed, have many, many choice points, and we basically developed a little bit of a structure around how it should draft these agreements, and then we gave it a term sheet plus the NVCA document and we said, “Write, you know, a charter. Write an SPA. Write, you know, the full suite of five agreements.”
[00:30:43] And it turns out after, you know, a little bit, a, a week of tweaking, the contracts we got were of really, really high quality. So, I think, you know, we are at a point, at this interesting point where, yes, the models are still quite limited if your use of the models is just here’s, you
[00:31:01] Irene Liu: sure
[00:31:01] Julian Nyarko: Here’s some documents-
[00:31:02] Irene Liu: Yeah
[00:31:02] Julian Nyarko: … do that for me. But there are many, many workarounds or, you know, if the more structure you give to a problem, the better an LLM is at handling that problem. And I think that is currently still quite underexplored, that they can do these really, really high complexity tasks if you break it down in sufficient sort of building blocks for it.
[00:31:20] Irene Liu: And for that second experiment where you did input the NVCA documents, were those also the contract that was drafted by the LLM, was that also checked by an expert, and did they give their approval for it?
[00:31:33] Julian Nyarko: Yes. Yes. Exactly.
[00:31:34] Irene Liu: Oh, interesting.
[00:31:34] Julian Nyarko: So, yeah and among others, my colleague, but then we also gave it to, a couple of other experts. This was all, this is not a rigorous study, large N, so this was more sort of anecdotal.
[00:31:45] Irene Liu: Yeah.
[00:31:45] Julian Nyarko: But yeah, the final contract set up based on NVCA templates were generally considered of high quality from experts in the field. So, yeah.
[00:31:52] Irene Liu: Yeah, that’s great
[00:31:55] Julian Nyarko: … basically venture financing and attorneys and so on.
[00:31:55] Irene Liu: Yeah. I mean, people use NVCA documents as a template, so that’s a really great baseline-
[00:32:00] Julian Nyarko: Right
[00:32:00] Irene Liu: … to put in as a structure for creating another document based on those. So, so that makes sense. What about for AI that’s overhyped? Are, are there uses of AI that has been overhyped? We talked about what’s, what’s been underrated.
[00:32:14] Julian Nyarko: So, I, I guess it’s probably the sort of flip side of this process that I just, just mentioned, which is we do have many people who are very enthusiastic about AI because it gives you a plausible looking-
[00:32:28] Irene Liu: Yes. …
[00:32:28] Julian Nyarko: output. Yeah. And I do think, you know, in law we see many, many more people file pro se now because, right, it looks like you can do it with AI by just giving it a bunch of documents, a little bit of context, and then say, “File me something nice,” right?
[00:32:42] But I do think we’re not quite at that point yet, and so, you know, especially from non-legal experts, there’s a lot of, I, I think the, the, this, the… I don’t want to say overhype, but I think they’re often too enthusiastic about the results that you might get because in the eyes of a trained legal expert, the default output is just not that great.
[00:33:06] Irene Liu: Yeah. I think the key word you mentioned is plausible looking.
[00:33:08] Julian Nyarko: Exactly.
[00:33:09] Irene Liu: The outputs are plausible looking, and so you pass it by thinking it’s plausible looking, but it’s probably
[00:33:15] Julian Nyarko: Yeah
[00:33:15] Irene Liu: … in the eyes of an expert. so, if we were to fast-forward five years, what will lawyers look back on and realize that they got completely wrong about AI?
[00:33:24] So we’re looking at 2031 now. It’s always hard to predict when things are moving so fast.
[00:33:28] Julian Nyarko: Yeah.
[00:33:29] Irene Liu: But if you could join me in reimagining 2031.
[00:33:31] Julian Nyarko: 2031. I do think judgment is very, very valuable. I, I do think models have better judgment than we often give them credit for. And so, I do think there are some aspects of judgment, that would be considered judgment, that can be translated into a more structured process, and thus then can be tackled well with LLMs.
[00:34:00] I don’t think lawyers will go away, and I don’t think all of judgment can be translated or be, can, can be, can be copied by LLMs. But I do think this striking dichotomy that we currently make, which is judgment is for humans and, you know, the more structured tasks are for LLMs, I think that is not quite reflecting the emergent capabilities that we’re seeing in LLMs.
[00:34:23] And so, I think we probably are going to look back and say, “Oh, you know, in many of these tasks that we thought are, the exclusive domain of, humans, LLMs can actually help us with those as well.”
[00:34:37] Irene Liu: Well, that’s such a great ending to say that, you know, at the end of the day, even when we look upon the future, judgment will remain with the humans, but it may be shared with AI.
[00:34:45] And that might be the world that we’re heading towards, so.
[00:34:47] Julian Nyarko: Exactly.
[00:34:48] Irene Liu: Thanks so much, Julian. All right. Thank you. I really appreciate you joining us at AI Sidebar.
[00:34:51] Julian Nyarko: Appreciate it. Thank you.
[00:34:52] Irene Liu: Thank you.
[00:34:57] A huge thank you to Professor Julian Nyarko for joining us on AI Sidebar and for such a thoughtful conversation about liftlab, legal AI research, and what happens when AI begins to perform well in areas that depend not just on knowledge, but on judgment. And thank you for tuning in to the AI Sidebar. If you enjoyed this conversation, please follow the podcast and share it with a friend, a colleague, or a student who is thinking seriously about the future of law and AI.
[00:35:26] Until next time, stay curious and keep learning