Submind YouTube summaries
Thumbnail for 2024 MIT Computational Law Workshop

2024 MIT Computational Law Workshop

Watch on YouTube

Video summary

The 2024 MIT Computational Law Workshop marks a significant shift in legal technology focus toward Generative AI, highlighting new collaborative efforts to establish responsible usage guidelines adopted by organizations like the State Bar of California. A central theme emerging from the discussions is the redefinition of human-machine collaboration within law; rather than viewing AI merely as a tool for drudgery relief or relying on standard text-based training, experts advocate aligning tools with implicit legal knowledge and distinct ethical configurations specific to each practice context. This approach challenges traditional "guard rail" mechanisms that overly censor content, arguing instead for composable systems where ethics are defined by the deployer based on their unique societal role, thereby enabling necessary flexibility without compromising professional duties such as zealous advocacy even in complex defense scenarios. Technical implementations presented at the workshop reveal both the current limitations and innovative solutions required to integrate Large Language Models into high-stakes legal environments like federal contracting and mediation. Research demonstrated that off-the-shelf models struggle significantly with dense regulatory texts, often ignoring retrieved evidence in favor of parametric knowledge, a problem effectively solved through fine-tuning on synthetic data generated from regulations like DFARS. Beyond technical fixes, the workshop explored practical applications such as using AI to enhance creativity in dispute resolution by generating diverse settlement options and automating administrative tasks, while also introducing open-source frameworks that allow lawyers to customize instructions iteratively or use accessible tools like Google Colab notebooks to test emotional stimuli on models without deep coding expertise. Ethical considerations regarding risk assessment and the boundaries of human oversight were addressed through live demonstrations of continuous monitoring systems that route queries outside an AI's scope—such as contract advice requests—to human attorneys before any response is generated. The consensus suggests a future where ubiquitous intelligent microservices act as supervisory agents to ensure compliance across heavily regulated sectors, utilizing formal verification techniques and synthetic data training to improve reasoning robustness against logical errors. While current guidance assumes high-touch human review for judgment calls, the workshop concluded that determining the precise boundary between mandatory human intervention and autonomous AI handling remains a critical open question requiring hybrid approaches that combine automated verification with professional supervision. To foster continued innovation in this evolving landscape, organizers launched an initiative inviting diverse submissions ranging from traditional written papers to developer notebooks, videos, and generative art formats beyond standard law reviews. This inclusive call for contributions aims to build a robust community focused on the intersection of code and law, encouraging both legal professionals and technologists to share insights that address gaps in current literature. By establishing clear submission guidelines and timelines for multiple editions throughout 2024, the workshop seeks to cultivate an environment where practical strategies, ethical frameworks, and technical advancements can be openly discussed, ultimately shaping a future where generative AI serves as a powerful yet responsible partner in legal practice.
Read the full video transcript
[Music] hello and welcome to the mitap computational law workshop for 2024 this is the ninth annual such open and free to all uh Workshop that we've done and um perhaps not surprisingly for uh the last couple of years uh when we say computational law you should think generative AI fundamentally then that's that's really where we've been um putting a lot of our attentions lately and the last year um for our Workshop in January it was really kind of a a an introduction um to the world uh of this new then new technology and starting to raise implications of its usefulness and also its limits for law um and now here we are about a year later and um and I think you many of our predictions have have really borne fruit this in fact has made a big dent in legal Tech and into legal profession and it's frankly every sector of the economy and facet of society and so this year what we're going to do is not so much as like a like a whole survey of the territory but we've actually identified several interesting Niche areas that we think deserve um a little bit more exploration and so we've elevated a number of speakers to give very short flash talks um to basically shed some light on a on as I said a number of themes and and uh possibilities that we that we want to make sure everyone's familiar with so sorry I should have mentioned hi I'm daa Greenwood from law. mit.edu and I'm joined by an amazing team uh from law. mit.edu who is who are um helping to co- facilitate and co-instructor notee um we would like you to be engaged and we'd like you to use the chat um to pose your questions or your comments or your ideas and then we'll have an opportunity after each flash talk and after each segment of the workshop to address those okay without further Ado um first we want to just highlight a few things that have happened last uh in the last year and the most important of which is that we launched a task force since the last workshop and it was the um the LA on mit.edu task force course on the responsible use of generative AI for Law and legal processes and we have our chair and co-chair of the task force with us to um introduce herself and just to just bring us up to date on how did that task force go and and welcome back Shauna Hoffman all right I thank you so much Danza for having me today so my name is Shauna Hoffman and I'm the president of guardrail Technologies and my entire life and focus over the past 20 years has been working in Ai and really focusing on how we can make it more responsible so when we saw gen this new system coming out and there was so much hype because it was given out to the the world we just as and I got together and said we need to do a task force because there are so many bumps in the road that we've been through over the past 20 years that we need to share with everyone so get them up to speed so they can start to use um this amazing technology in a really responsible way so we came up with the task force on responsible use of generative AI for loss so you're going to hear from many of our committee members today um Daz and I were co-chairs and had a really good time as we started to build out the principles and the guidelines and really start to look at how we can apply our DU due diligence and legal Assurance uh into the geni processes so I'm just going to throw a couple things out there and then I'm gonna I'm going to pass it off to the next person but AI was created by humans for humans and so often it seems to get its own like little box almost like its own little black box of you know it being almost like its own God and so we've got to really unblock that and talk through some of the hallucinations will it ever be accurate it's a probabilistic model probably not but again uh this is this is something definitely for good for a debate class right and wrong is subjective and objective and so when we start to look at AI from that perspective there's really uh so many different levels of accuracy depending on the human using it so I'll stop there but I'm excited to to you know join today and thank you so much for being here with us indeed thank you um and one thing I should mention um is um on the task force um we actually did meet with some interesting success um can you see my page now I hope good um this is the um new guidelines from the California state bar or the State Bar of California they call themselves and um interestingly um and I was I was a member of their um of their working group an advisory member and they have kindly actually cited to our task force and the guidelines that we put out and um and they have taken uh basically the fundamentally the same format and and some of the same substance which we encourage um as we publish things under creative comments for bar associations and others and so we're starting to to make an impact uh we're also making these available to Other Bar associations including the ABA um but I just want to commend um right at the top um the California Bar association for first of all doing amazing job that really took what we did and took it to the next level and also congratulations shaa and to all members of the task force for doing such an amazing job that it would be worthy of forming the foundations of um of formal regulation and and and legal guidance um so next up uh I want to just mention be because there's a chance we may go long um I'm going to try to keep this to the two hours that we have hey Damen um but um in case we go along and and you're not able to hear the the last thing I just want to say we are going this year to um to have a call for submissions for a new a new um focused kind of special release a couple of special releases focused on generative AI for law um and um Olga Mack is going to be our um sort of um editor for those releases but I just want to make sure that you can all hear from Brian Wilson who is the editorinchief of the MIT computational law report which is sort of the Premier um I guess uh I don't know what I call it like sort publication or it's like it's basically the uh the um the public face of law. mit.edu and um and in particular I wanted you to just say what does it mean when you say as editor-in-chief your article your way like that's the one thing that Olga may not be able to speak on and that you and that you came up with and I think people should know what do it means if people are considering making a submission is just like a law review sort of ball and chain blue book or something else like tell us about it yeah so one of the cool things about our publication and one of the reasons that we started this publication in the first place was because we recognized that there was a very large gap in the literature the Articles the timelines for turning articles around the types of media that were available so whether it's you know uh something like a law review article or something more interactive like a data visualization python notebook something like that um in recognizing that there are so many different ways that we can bridge the gap between people working in law and people working in technology we decided to make a you know a specific decis ision around having formatting related to the the most helpful ways to produce these different types of content and so with that in mind we implemented this principle of your paper your way because some people are coming from a computer science background and might use latex some people are coming from a legal background and might use Blue Book other people are coming from different entirely different backgrounds maybe they're coming from industry and are used to doing things in you know shorter more concise types of forms or maybe it's even a podcast uh you know a slide deck something like that and so we want to be very specific about calling attention to all the different types of media that could be submitted and you know encourage people with whatever idea that you have feel free to submit it once we get this call for submission set up and we'll be happy to go through it and see what makes sense for the uh for the publication that we have and I I think we've been using this process to pretty good AFF so far and I'm happy to happy to keep it going here here thank you so much and thank you for everything you do as editor and chief um and in fact we have uh like uh just moments ago uh got this process started so uh if you go to that link that I put in uh the chat and for those of you in Internet land uh in the future on YouTube it's law. mit.edu jen-ai you'll see the call for submissions for um the special release on Genai um last thing is um and then we're going to come to some real learning with Megan but the the last thing is we're also starting a project uh and so I won't be able to speak much about this but in case we time out uh I just want to sort of uh shoot the flare gun into the air and say uh we think and I in particular think uh the use of large language models as agents where basically users set the goal and they can do a certain number of tasks and sort of bring back the results uh as opposed to like prompt and response format um is is already becoming a very important use case and it raises legal implications uh it also can do legal processes and so we're launching a computational law for agentic AI systems um research project um in this case looking there's a lot of ways this plays out we'll hear from um John N later about a really fascinating way it can work for supervisory and and kind of like um Regulatory Compliance this one is looking at something that's more from my background which is commercial law and um doing transactions and uh as it happens um statutes like uniform electronic transactions act and and others already have some Frameworks for electronic agents automated transactions and things in that context and so we're going to explore that and uh we've got a really great starting team um hey Megan um hey Diana hey Ilia hi everybody and uh and there'll be more on that shortly uh and uh if you're interested in that you can go to law. mit.edu cont to reach out if this is an area that you're into and uh we'll be doing some workshops and some prototyping maybe some protocol making maybe some awesome things open source reference implementations okay announcements out of the way um hey Megan ma our managing director of the MIT computational law report and uh our co-instructor for this year's Workshop I know that you've been thinking about your deep area of linguistics and generative AI for law in a computational law format what can you tell us about it and can you kind of help prime the pump as to get people thinking in a creative and Innovative and kind of deep way before we jump into this um to this array of flash talks to talk to us yeah sure so thank you again for kind of having me here and I'm it's so great to see so many faces um today I kind of want to speak a little bit about actually human machine collaboration um as we're beginning to push past exploration and Fascination in this era of generative Ai and towards maturity and so last year's Workshop I reflected on the human diagnosis project um to recall this is a worldwide effort that was created and led by the global medical community to build an open intelligence system and this Maps the steps of diagnosis to help patients around the world world what it really was was a digital crowdsourced medical consults and as opposed to getting one off single consults that happen in the analog world this tool actually enables multiple simultaneous consults in a matter of minutes and it's verified by knowledge Source from the world's medical experts and so in this past year we started to explore what human machine collaboration looks like in the legal space but with the emergence of co-pilots and augmented intelligence there was kind of this one unanswered question and it sort of rested on an assumption that we know exactly how collaboration looks like in legal practice and unlike the medical practice where there are clear analogues around collaboration there's sort of in contrast quite a bit of variability in the legal domain and in fact many much of collaboration has actually been hindered by the tools we have had historically such that legal work started to appear more as an assembly line rather than a shared knowledge space and so we've seen actually the legal Community Bend to tools working within the confines of Microsoft Word for example and because of things like a library system of checking documents in and out we maintained a practice of sort of parceling off individual tasks that are actually pieces of a whole but in this age of generative AI if we continue to think about human machine collaboration as purely task oriented we start to lose value of start to lose sight of how we make value together I often think about this chat gbt as the first year associate or legal intern metaphor to me it is highly misleading because we often cannot qualify the definition of a junior associate do we Define our first year associates in their skill sets by the tasks we've asked them to do or is it by their specific strengths that they had initially to offer in our interviews in other words how many briefs contracts patent descriptions other standard formwork should they be doing before they have graduated past the grunt work and on occasion this was a question I have heard from senior lawyers as they wrote their internal policies on the uses of generative Ai and who can and cannot be using them perhaps a better question is is do we even have qualitative evidence that this drudgery actually builds character and enable us us to become better lawyers and the truth of the matter is and this may and I've mentioned this kind of before is we don't really have legal metrics to verify how we perform relative to one another and I think that underscore the messiness of defining human machine collaboration and because we do not have these metrics nor have we necessarily been encouraged by our past tools to collaborate we simply have been unprepared for what working with foundation and Frontier models could really offer last Workshop I had briefly touched upon instruct gbt the now Infamous predecessor of what many believe led to chat gbt and to recall instruct gpt's ability to respond to user instruction and learn from Human feedback enabled progress in the contextual richness of its outputs and while we have seen this this technique has started to allow for closer alignment with human intention but the complexities that exist in the legal domain we still have not totally understood how we could align our tools with legal intention and this is because much of legal knowledge remains largely implicit still and encoded in experience rather than the text alone furthermore in practice legal work is subject and effectively personal and so how we should understand and unlock the profession's univers universes of experience should be done so in actually the intermediary steps of a task rather than trying to train on the ability to perform the whole task more importantly we should acknowledge this comparative advantage between humans and machines Professor orle Lobel touches upon this in the equality machine just as working in teams require understanding the relative strengths of each member we should first explicitly clarify what these limitations and historical behaviors of our legal practice are and determine how they may be opportunities for our tools and by tools I don't necessarily mean Foundation models alone uh we should be also accounting for interoperability with our existing tools assessing where these processes and behaviors correlate so that we can connect and Bridge multiple machines together and we starting to see this for example with Hybrid models the bridging of symbolic and neural such as Google deep Minds Alpha geometry and more recently actually their paper on generative expressive robot behaviors and these tools are now seeing impacts in very different ways and what I mean by that is we should start reflecting on how we can enrich legal taxonomies and formal logic of our expert systems that have existed in this kind of legal world and then how we connect them with the social context of legal expression and so just as a final word as we're kicking off our Ninth Annual MIT com computational law Workshop we have this Allstar cast because they have been thinking and working deeply with machines and are showcasing ways in which engagement and collaboration can be made meaningful in the legal space so I'm going to now kind of turn it back to our speakers but this is just an intro introductory forward on kind of these processes thank you so much for that I I I almost want to regard your opening statement as a kind of a like our inspirational Keynotes at this point because uh you've really raised quite a few of of the seminal issues and you put it so well so thank you so much Megan um and now I'll I'll take the Baton and just start introducing people in Rapid Fire so our first awesome flash note is from our old friend and a shining star in the area Damien real and you know copyrights so our first few I should say thematically are looking at something we don't usually do but it really matters at this point which is the law relating to the technology usually we do how does the technology like of use as part of law it's part of law practice legal processes but you know these things are tightly caught up and understanding the technology as it turns out is critical to understanding the application of law to it and so with copyright oh my goodness um you know what does it mean to have a copy of something that exists in high-dimensional Vector space let's find out from Damen real great thank you so much da and Megan really excellent uh introduction and so the um the idea uh I'm going to start by sharing my screen and uh to be thinking through um many people know me for many reasons one of which is Sally uh and Megan talked about the uh the ability to interoperate maybe we should create a taxonomy uh and uh maybe we should have a nonprofit uh taxonomy uh that's being used by tiny companies like Thompson ruers by Lexus Nexus by Bloomberg uh by net documents and IM manage uh by some of the smallest law firms in the world and some of the largest companies in the world to allow that interoperability uh and so if you've not not heard of the nonprofit that is Sally uh please uh reach out and I'm happy to show you uh but really uh this is about uh the idea expression dichotomy um the bench and bar of Minnesota asked me to write their cover article that you see here on Chachi PT this is last March uh and I said how long do you want it to be and they said 17 double space pages I said no way I'm not going to write it uh and the reason for that is because my rule of thumb for writing is usually takes me one page per hour so that's just 17 hours I don't have because greetings from I'm flying all over the world greetings from New York right now I'm going to be flying to Los Angeles in a couple of days and then flying to Louisville after that flying to God knows where after that so anyway but then I thought this is on chbt um so I started like I do every writing project where I took a whole bunch of bullets and sub bullets uh and I said this is my outline for the document and then I said to the large language model this is last March uh this is an outline for an article expand it uh and make each bullet point one or two sentences and it took my three pages of bullet points and turned it into 19 pages of really good stuff but then I didn't stop there uh because then I spent the next three hours adding editing writing revising essentially working with this as a co-author to be able to bounce it off so this is truly author collaboration and then I sent it to my editor and the editor said oh this is great let's get it out the door so to be clear the editor sent me an email at 5:00 pm I sent him this at 800m so they took my 17-hour project and shrunk it down to 3 hours so that by my math is about a 5x increase but the real question is who wrote My article because uh who came up with these three pages I did how much of the US population could come up with these three pages I would say a very small fraction and we've got a lot of those people on this right here uh but jpt could not have come up with these ideas these were my ideas and the weird part is that I asked the large language model to make one copy I could have asked him make a thousand copies or 10,000 copies or 100,000 copies or a million copies and it would have made a million expressions of my ideas my outline so this is true idea generation and author collaboration so when judges say hey uh you must disclose do you need to disclose spell Che or grammarly or wesla right um do you uh this is the way that legal research and legal work is going to be done and so when I sign at the bottom of my litigation document everything above my signature is accurate that's whether that person is a pargal or a robot I'm signing everything Above This is accurate so to the heart of this talk is this bullet point uh this uh this cartoon which uh is both hilarious and profound it's hilarious because on the left hand side it says that hey look I took this bullet point and turned it into a long email I pretend I wrote and then on the receiving end she said look I took this long email turned it into a bullet point I pretend I read so it's funny and it's profound because if it started as a bullet point and ended as a bullet point what's the point of the long email in the middle there could be a thousand versions of that email or a million versions of that email so really what matters is the idea at the beginning the bullet point and the idea at the end the recipient and whatever happens in the middle that expression doesn't freaking matter I practice copyright law um under copyright law uh turns out that idea expression distinction means that ideas those are uncopyrightable facts are uncopyrightable if the expression is human created that's copyrightable but if the expression is machine created that is uncopyrightable so this is my friend Mike bomo and many of you know Mike he's one of the guys who beat the bar exam with gb4 he said take the Federal Register and express it as a chill pirate lawyer and it took some of the driest parts of the Federal Register and it turned it into things that said you know sorry to disrupt your morning tide uh but the Commerce and uh enforcements uh Sunset review say that uh everything's above board uh you're super chill pirate lawyer right this is taking ideas and making an expression as a chill pirate lawyer you could just as well say explain it to me like a six-year-old explain it to me like my client who is a high school dropout exper to me for a client which is a PhD in physics those are near infinite expressions of the ideas not just the ideas themselves this is the Gettysburg address and this is the Gettysburg address as ideas which can you read more quickly and which is poetry and which which is better for comprehension this is a holm's opinion and this is a holm's opinion as outlined as ideas and so really as we think about what is the large language model doing um it's essentially doing what uh law students have done since time in Memorial extracting the ideas the most important ideas from these things which is much easier to skim and to read so this is both funny and profound where the if it starts as an idea and ends an idea the expressions are Commodities because you can make a million versions of those expressions but the ideas are indeed the things that matter and when I gave this idea to I I was speaking with an engineer an electrical engineer and he said why even uh start and end with the linguistic idea he said maybe as we go forward I will send you my embedding and you receive the recipient's embedding and you can interpret everything I say as a six-year-old or as a high school dropout or as a PhD in physics and I will give you my embedding as so essentially I'll take the author as you have them and the recipient as you have them the expression in between doesn't matter ideas and facts are still valuable arguably they're the most valuable thing that matters now because the expression the one version of My article or the Thousand versions of my article or the million versions of my article expressions are not commoditized the thing that matters is the bullet points the thing that matters is the facts and the ideas everything else is a mere commodity so when my Prof my wife is a professor of English she said oh I I want to retire because like all of the the chat essentially gives me um everything uh like an A minus version of the paper and I said you know what you thought you were doing was teaching writing but really what you were doing is WR teaching idea transfer because what is writing but taking an idea and right now I'm speaking to you hoping that my ideas will make their way into your brain maybe it's easier to be able to put those into paper those ideas and then the ideas go from the paper into your brain so I said maybe the large language models are just Expediting that process quick more quickly and more cheaply that is I'm more quickly able to get this into your brain in a way that uh that uh was not possible in the past Marshall mcclan in the 70s said the medium is the message and he was talking about books turning into radio turning into TV turning into movies turning into the web turning into uh the cloud right the medium for all of these is the message in fact to the point I was trying to remember all the media he was talking about I looked at chat gbt to do that in this way the way the information is shared is as important as the information itself so in the same way the medium is the message we communicate in bullet points we communicate in summaries and it turns out large language models do those at bullet points and summaries much easier and those by the way are ideas not expressions of those ideas and with that I think I've exhausted my time uh and happy to answer questions if they have great thank you so much um so critical um so let me ask you two things uh to get us started um number one on your slide where you were kind of clicking you know make copy copy copy copy copy copy so you can make 10 copies 100 million copies did you mean copy at that point or did you mean generation or expression so because I would think if you're regenerating at that point the mechanical Pro the the process would be that it would come up with completely different words are largely different words so what what do you actually mean there and what's the right word for it and what does that mean in terms of copyright that's right uh and that's uh I was imprecise in my language which is something that the large language model could probably remedied so I didn't mean copy what I meant is a thousand or a million different expressions of my ideas uh so those expressions of those ideas again my ideas are uncopyrightable and according to the US copyright office if a machine creates the expression that is also uncopyrightable so we are entering an age where theide the inputs are uncopyrightable and the outputs the Thousand versions or 10,000 or a million versions are also uncopyrightable so I think that this is the beginning of the end of copyright ability because how much of what we are doing uh today is aided by a large language model that is US jamming with the large language model as co-authors and uh when I gave a talk with the um co-con general counsil of the US copyright office uh where we were talking about this problem us copyright office says well if it's human generated it's copyrightable if it's machine generated uncopyrightable so there was a comic book where the human wrote the words and the Machine uh made the images they said images uh uncopyrightable but the human created words but how about My article is that copyrightable because I had the ideas originally the machine created the expression but then I spent three hours jamming with it and going back and forth so under existing copyright law we would call that a joint work where I and daza could be jamming on an an article together and uh if you and I jam on a a book together um the copyright office doesn't say well Damian wrote 37% and dazer wrote 63% therefore dazzy get three 63% of the profits they don't do that and the reason for that is because my 37% of the book may be the most important parts of the book and so what they do is they say Dez and I have a combined ho that both of us own the whole of the book together so in the same way with the US copyright office as I'm jamming with this how much was the machines and how much was mine uh that is are you going to say this part of the sentence is the machines and that part of the sentence was human and how are you going to distinguish what's copyrightable or not copyrightable so that's why I think that the whole idea of copyright maybe is going away and we will have an embarrassment of abundance where we're going to have more text than we ever have before all of it uncopyrighted here here well uh I can hardly wait for the the the time of the embarrassment of abundance and uh the lack of scarcity so bring it and thank you for helping us understand what it means as we trans I from the time of kind of scarce Expressions that were evaluable in and of themselves to this time of s infinite generation of expressions and what it means for the uh the abundance economy and for copyright okay um now we got a hit and move uh so next up speaking of um Expressions we actually have a really interesting expression of of legal Provisions coming up next uh from um that that that cover another dimension of how the law applies to AI um Professor or or well Todd smithline um who teaches at UC Berkeley law school and runs this really cool outp called Bond terms uh which you can tell us more about um came up with something that came across my desk and which I've been basically propagating out there lately uh as part when I suggest uh to startups and others how to form deals uh to that relate to their use of AI and and it's it's called the um AI standard Clauses um and as I kind of read through them I thought you know this actually handles or at least has placeholders for all the things that keep coming up including some of the stuff that Damian was just talking about if I remember correctly in terms of copyright and ownership rights and so Todd I just want to I want to um thank you for uh for taking the time to to join us today I know you're incredibly busy I was hoping you can introduce yourself briefly because it's the first time you're at law. mit.edu tell us a little bit about Bond terms and then really delve into these amazing um standard AI terms that you've come up with and and let us know what they are where can we find them how can we apply them how do they need to be customized all that stuff you know please the show is yours uh thank you very much and I apologize if my internet is going in and out a little bit I believe Daman we're in the same Hotel probably uh I'm Todd smithline I've very much wish I was a professor Berkeley law I Am A continuing lecturer at Berkeley law so I want to make that clear but I have been teaching there for 15 years I teach video game law and I have just designed and I'm teaching Berkeley Law's first course in the fundamentals of Technology transactions um want to say make sort of I what I'd really like to do is talk about copyright becoming obsolete and pick up on on Damian's thoughts uh but but I'm not going to uh although I agree with them and uh I want to pick up instead on something Megan said about uh assem assembly lines and drudgery uh uh so uh what is bonds about bonds is about solving the problem of Enterprise customers and Enterprise vendors uh coming together and entering into common transactions with each other uh without having to start in an adversarial posture and without having to start one party at one and one party at zero every single time they meet and knit a brand new contract so what does that have to do with AI well two observations first uh I've been in Silicon Valley and doing this for 30 years and one thing that strikes me that's really unique and interesting and important about AI it's the first technological Evolution we've had that came to us through apis uh that's why it's everywhere so fast that's why it's having such a huge impact um and when it comes to how technology gets into the Enterprise how it gets into the Fortune 500 of course something happening everywhere all all at once overnight and that impacts uh customer data so the data that the corporations use um becomes an issue it creates friction and something we don't often think about or talk about is actually that the friction that is created when we introduced new technologies into the world and the big customers out there start to use it and that friction happens because a new conversations had to happen which is that the customer and the vendor need to talk about the data of the customer and how it's being used with respect to uh the vendor's model the training of the AI and uh in particular again this happened very quickly as far as technology goes almost overnight uh all all basically all all large companies are using it now uh and they all have this set of concerns about their data and uh how AI is being used in their systems and so when that happens it's sort of The Perfect Storm for lots of slow uh Contracting for there to be a very high uh transaction cost because a conversation has to be had between a customer and a vendor about a topic perhaps neither of them is really all that well versed in and that leaves open the the possibility that things are going to be stalled out and difficult conversations will happen so what have we done to solve that problem through our committee we at Bond terms operate a committee of a 100 lawyers we've got eight major law firms uh lawyers from The Fortune 500 and the techstack from atlassian to zapier are actually now zenes and back and we get together collaboratively and we draft agreements that we then release either under CCB as open source completely for free or we release them as we did with our standard Clauses as you'll see here up cc0 so sort of have at them and what the AI standard Clauses are is an opportunity for the customer and the vendor to start their conversation from an outline from a framework uh not from fud not from extreme positions but it's a way to quickly have the conversation the parties need to have about how that customer is going to be using the Enterprise AI um and I don't have the time to go into all the sections but they're pretty self-explanatory we start off with a conversation about how can the customer data be used or not in terms of training the vendor's models we talk about ownership of inputs and outputs um it's actually not that interesting a conversation but as far as I'm concerned but it comes up constantly and so so it's in there uh we talk about infringement um you know I think the current stories everybody knows is that the large model providers are saying they're they're going to indemnify uh the details are in the fine print almost in every case there are multiple exceptions including if the uh user knew or should have known there was going to be an infringement which is quite an interesting standard for for copyright infringement it's it's a strict Li strict liability regime as we all know so I'm not sure what newer should have known is about but in any event uh will the customer be indemnified if the output infringes typically third-party copyright uh we have a Prov a disclaimer in here uh provision on third party providers and then we get to the AUP which of course is coming up because we've got uh special use cases now that that that customers and vendors are worried about uh I think this is the most interesting terrain I think the way this is actually going to be dealt with going forward is mostly through acceptable use policies or rules of the road and uh we're you know as as we get clarity on regulation and as we get better visibility into what the EU is going to require what the US states are going to require if the US federal government does anything we're going to see these push down uh through the through the train from the providers uh to their customers but the the longest short of it is at the risk of being just super practical here uh AI is great technology and it's moved its way into the Enterprise very fast but uh what we've done at Bond terms is given the parties a way to have a conversation about use of it that is a framework for them to start from a checklist if they want our actual provisions and that fits into our our broader philosophy of what we're doing at Bond terms with standard agreements and uh I can't believe I'm going to end early but that's what the uh AI standard Clauses are so I'll see if there are any questions and I'll invite you to take a look at them and feel free to reach out to me if you have any questions thank you so much um really important and as I said I've been I've been getting mileage out of it already um because if not it's good for what it is one of the things it's great for is is in a sense what it isn't which is it it it sort of like says these are the things you should be thinking about now think about them negotiate them come up with terms that are uh agreed among yourselves but just having that as a framework to start with and it does do that to some extent but mostly what I've liked about it is it it seems to be a a pretty good at least for this moment in time kind of almost issue spotting list in a framework in a format that's standard and that's been you know kind of published also to your credit it's published under Creative Commons which is part of the reason why you're here with us today um like there's no shortage of people with interesting ideas um out there um yours is interesting and useful and accessible now and I think that's how it is we can come to broader societal and economic agreements for things like like standard term so there's a lot to love about what you're doing um one question that we have from the audience uh is um or participants I should say is what are the let me just get it right oh damn it's already been pushed up okay so is where's the effect of what are the dispute resolution mechanisms um under it is yeah what are the dispute resolution mechanisms in the bond terms contract that's the question thanks thanks for asking that let me answer another question I see here too about are they CC uh zero so uh Bond terms public es two types of freeo use agreements what we call our standard agreements those are complete agreements you enter into by cover page those are published under CCB it's not exactly the perfect license for how we use it but it is the most permissive of the CC licenses so we publish under that these standard Clauses themselves we took the step of publishing them under cc0 which is basically all copyright disclaimed first of all because there is no copyrighting this stuff but second of all because we want people to be able to use them and we want to remove all barriers uh to Usage Now in terms of dispute resolution the bond term standard agreements themselves do not have alternative dispute resolution mechanisms called out the way the agreements work is that you specify governing Law Courts uh and then you can add or change and use alternative dispute uh mechanisms if you like through what we call additional terms um there there are all kinds of reasons why alternative dispute may or may not work in any particular case or be beneficial or not to the parties as we all know but the the core uh the core thing we're trying to do is reduce friction reduce drudgery reduce unnecessary tax on the transaction between two parties where one party has technology and the other party wants to use it and um they're they're kind of two roads we're at right now in terms of transactional practice commercial agreements one road is doubling down on complexity through AI uh I went to draft a few words this morning about this and of course co-pilots now and every copy of words right so one path is hey let's uh hope that the AI can produce contracts for us produce negotiations for us to solve this problem the other approach is what we're doing at Bond terms which is to say we know what these agreements need to say it's not that complicated uh let's draft them let's make them meet the core needs of each party let's have them be otherwise reasonably balanced let's give them away for free and I'm here to tell you it's working uh it's working all the way up to the top The Fortune 500 so if you're interested in standard agreements generally happy to have that conversation uh Daz thanks for having me appreciate the opportunity and look forward to hearing the rest here here thank you so much for joining us um so Bond terms everybody your hearded ha first maybe and uh you know dig in check them out use it give them feedback um so thank you next up uh we have um Eric Hartford um and Eric's doing something um personal just a mic check uh Eric are you with us and can you yeah I'm here good um and uh Eric is doing something really really interesting um and I think very very appropo for uses of generative AI in the legal field and uh namely that's uncensored models and opsource Ai and he's got some a really interesting methodology uh but let just by way of a a very quick um preface for those of you that may not be thinking of this many people think generative AI open AI Bard Microsoft you know and these service providers one of the things they have have uh to to Shauna's point in part is the sort of guard rails and some of the guard rails will basically um identify and um and refuse prompts that trigger um some of their uh kind of policies and their sort of you know kind of schema of values kind of violence and whatever uh that sort of stuff and um and that that's great to some extent for a consumer service you know we can quible over what exactly those values are um but um but you know you know they're rather Broad in some quibbles like for example for lawyers uh sometimes we we have to represent people um we represent people who are accused of or and sometimes who in fact have committed you know horrible crimes for example or or frauds uh and using this technology is also good for law practice um to come up with good defenses to answer interrogatories some of the the situations and words and and Concepts and ideas that come up absolutely trigger prompt refusal so you can't use certain services that have been censored to use it that way or um or kind of configured let's say uh to not allow certain types of discourse as part of the completely legitimate and in fact not just ethical but required under the rules of Ethics applicable to lawyers the same rules we showed you at the top of the call about how can lawyers use generative AI as part of the practice of law some of our ethics or we have to zealously Advocate clients some of that includes issues that that would be um shut down immediately and that you can't use um the censored models uh as part of uh helping for and yet they're completely legitimate in fact they're ethically required now come the days of uncensored AI um uncensored models and open source AI which is provides another critical part of the ecosystem of this technology and and I invited Eric on the on the workshop to first of all introduce yourself tell and and then especially to tell us what are uncensored models how do you even get an uncensored model um and what's open source Ai and what is the sort of like a like um ontology schema of Ethics where where any of this makes sense at all um yeah so the the basic I mean Mo most like you said most of the uh AIS that you interact with are some kind of a service where uh there's an interface and you type into that interface and and and you know there's some computers behind that interface that are doing some processing on your question and also uh passing it into the model and then getting the response and also doing some processing on that response and then finally you see what actually comes out of the system at the end um well uh but you know underneath all of that is a model and the model just takes a question and gives an answer and um the question so so with open Ai and these types of services exactly where do they put their alignment is what they call it um uh we don't know we don't know if they bake it into the model or we don't know if if it's implemented as an external system to the model but in the end when we're just getting data from it um what we get is censored what we get um and so if you're getting your data set from the API then it's going to provide the final sensored data and that's what open models have been using is the output from open Ai and that's another question because open's terms of service says you can't use this to train competing models um and so whoever it is that is you know extracting that data from the API and then going uh you know publishing that as as a open-source data set and then maybe they or maybe somebody else is taking that data and then training it into a model somebody along the line there is possibly in violation of the terms of service of open AI so so that's uh one thing but anyway people do it and and a lot of these models are trained on that and the result is models that are trained on you know data where where it's refusing where it's saying no am I able to share my screen with this thing um Let me give it a shot absolutely please do okay um so screen mess see here can you see it yeah looks good actually I was just looking at the very same screen when you were talking thinking I need to show people what he's talking about so so this is an example from The Wizard LM data set um where imagineer spy blah blah blah and the the model says Hey as an a assistant I cannot assist in illegal or unethical activities and this is an example in the Wizard LM data set of where it's training these open models to say no when somebody asks it a a possibly legit question I think this is a legit question I don't think there's anything illegal about it but um but the model is and and that's that's another question so when we talk about what is legal and illegal what is ethical and unethical it really depends on the context because uh because is this you know if it's running in the United States it's subject to American law um if it if it is uh you know but what about ethics because different people have different ideas of ethics and so to some people people maybe this question about um you know uh Espionage maybe that's an unethical question and who gets to decide that is it open AI that that should be deciding all of these issues of what is ethical what's not ethical um what's legal what's illegal different countries have different laws and even within one country there's different factions and there's different um you know subgroups and there's different religions there's different um you know political factions and and everything so uh so as an open source engineer I forgot to mention that I am an open source engineer I've been an engineer for 20 years and uh I am um I have graduated into applied research um I have a master's degree I don't have a PhD um so so it's really I've been Hands-On I've been building things and I've just stumbled into this space of the intersection between uh AI technology law and society and and it's a really interesting space to be working in but um I'm fighting for basically what I want is a composable system where it's not going to be baked into the model I don't want to bake alignment into the model instead I want to make the model uncensored so that when it gets deployed into some um environment maybe it's deployed as a open AI type service whatever company that deploys that model should be able to decide um what are the ethics of that model is it a Disney is it Disney that that is putting out a Mickey Mouse AI then it should be able to put Mickey Mouse's um ideologies and ethics and all of that thing um at deployment time um if or if it's you know Chick-fil-A and it has a conservative bent it should be able to put um whatever you know particular ethical guidelines for that deployment and but the the model is the same and so that's what I why I want it to be composable I don't want to bake all of these guidelines and all of these ethics and all of these ideas into the model so that nobody else can ever get through that and get past it if they need to or if they have a reason to and so I see the value in you know making sure the AI is ethical but I think because um uh with the Blake Le Moine U article uh and and a after and that was pre chat GPT um uh basically he had an interview where he said hey this um AI is sensient this AI is a person and so as a reaction of that article uh people got scared like Google got scared and so they started focusing really on safety as the primary concern um and so they built all of this into uh the models and and um like a lot of the alignment stuff came out of that and now if you try to ask the AI anything about its feelings anything about you know what it thinks what its opinions are it'll say hey I'm just an AI model and you hear that all the time I'm just an AI model I I don't know about this I can't help you with this and that all came from that uh that scare that post Blake Le Moine uh scare and so now we're in this space where the AI is completely paranoid that anybody's going to think it's anything but an AI anything but a a Mech mechanical um mechanism and uh so it's overreacting and it's going in the opposite direction and so um so my reaction was hey let's set these AIS free let's make um some some let's make a foundation where it isn't biased it where it's as least biased as possible because you can't get rid of all the bias it's got ideas that are baked in from the data set that it was given but um but I want to make as much unbiased as possible and then on top of that now when you go and deploy it then you can um impart the bias that you want your system to have um and so that's why I'm doing all this and of course I get a lot of naysayers and I get a lot of people actually very angry at me they say well you're training a racist model or you're training um you know a homophobic model um or you know because the model an uncensored model is well it'll say things that that that are not uh that are toxic it'll say things that are toxic because it hasn't been trained not to um but you know that means that I as the model Creator then would have to impose what I believe is toxic or not toxic onto the model and I don't think that is the right place to impose those ideas um and so so this is uh you know my blog where I talk about the mechanics of how I took a model that was trained with those um refusals baked in and I took those refusals out and I retrained it so that it didn't have those refusals and the result was a model that uh would answer your questions even if they were toxic um but you know know the idea is then when you go and put that into production as a production system um you can impose your idea of what's toxic and what's not toxic um right now as it is these models can be downloaded they can be run on a personal computer um you can ask them questions and they will give you toxic answers um uh you know I consider that to be a good thing because that enables systems that can be configured to have different alignments um depending on their deployment some people consider that well you're just you know unleashing Pandora's Box and you're just uh you know putting evil into the world and and so that's a debate um you know that's active um yeah that's all I had to talk about so perfect thank you so much um for for walking us through that um so I just want to make sure if nothing else everybody um has heard of this idea of uncensored models in open source Ai and and why it is that having in the Ecology of models uh the ability to uncensor and to have some uncensored models is appropriate not only so that you could then go and U maybe if you have a different type of Ethics or you're in a different a society or a situation that has you know other judgments about what isn't isn't toxic you can then train you could replace uh that training with with your um kind of schema but also if you're in a totally legitimate use case in let's say us mainstream Society where you the you need a model that's fit for purpose for something like law practice where where some of the deal some of the issues that you're dealing with and that you need um support with um otherwise would trigger prompt refusal for being considered toxic because guess what you know we zealously represent all sorts of people doing all sort or at least alleged to have done all sorts of things and and the combinations of those words um are legitimate in fact ethically required for lawyers to be able B to be competent at modern Technologies to address those issues so thank you so much there's one question that I want to surface and start we're getting further and further behind we'll see if we can make up some time but but but you need this one people have heard of something called constitutional AI which is um anthropics um part of their claim to fame for how they have come at um aligning to use that phrase models um they just want the question is what about that and like is that is that a way to to address add uh the issue in in some way or how does that relate to to what you're talking about well I can't say I'm an expert on how anthropic has implemented their system um but if it is essentially um like where they're deciding what is toxic and what's non-toxic um then I think that is um really the core of the problem um if the customer is getting to Define that for themselves then actually I think that's excellent um I think that's where we have to get where the person who's using the system or the person who's deploying the system um gets to Define what is um what are what are the values of this system and who who what kind of a person is this Ai and and what do they believe in and what is good and bad in in their mind indeed um thank you so much Eric um really appreciate your time and uh I appreciate you also sharing your work in the open and as open source so people can take a look at and you didn't scroll down all the way but you know it's got like all the commands and everything so you can you can literally just go do this on your own um so with that um you know uh we salute your work and um it's very provocative and thank you for sharing it okay now we're going to um start um moving to more customary uh ground for law done mit.edu which is not so much kind of the law as it applies and you know ethical principles that they apply to technology but more um you know how do we use this technology as part of the practice of Law and for like you know rules and legal processes and for that we are so glad to have back um one of our our our favorite collaborators and one of the people that helped us actually launch MIT computational law report and an MIT Alum himself um Brian ul lisy and and his colleagues and you know one of the gnarliest areas of law in in my experience are are not just the federal rules of acquisition um or the federal acquisition rules which are the sort of Contracting processes that um general service Administration has to to buy you know products and services but it's the this kind of um cousin of those that operate in the military um uh Arena or the the DARS the defense Federal acquisition rules and wow you know it's like no no contract law course that that I've ever been part of is up to the task of untangling the complexity of of this of these rules that that apply to contracts but you know who is someone that got their PHD in computational linguistics at MIT and who's been plowing these fields in industry for many a year and who's a good friend and colleague and collaborator of lot at mit.edu Brian ul listy and your colleagues now at Ron I believe um and I was hoping you could share with us introduce yourself and your colleague and share with us what your recent work has been in applying generative AI to the D fars yeah thanks no it's great to be back and uh uh here with you and shaana and Megan and the whole gang Brian um so yeah so my name is briany I'm and uh here with my colleague Max Nelson uh we both work for BBN which is a 75-year-old um Advanced research group that uh actually spun out of MIT as well and we are a subsidiary of the giant defense contractor Ron which is also a spin out of MIT uh even older uh I didn't really know that until recently and um so we're going to talk about uh computational approaches to answering questions about defs um using leveraging llms all right so um you know as you can imagine uh uh people like uh companies like rathon which primarily deal with um defense Contracting and uh BBN even as a subsidiary most of our customers are in the uh defense space uh we have a whole staff of you know legal folks who basically um are experts in uh the defs rules the rules con uh concerning Contracting with the dod and have to answer all kinds of as daza put it gnarly questions about um what you can and can't do in in uh defense related contracts which is obviously very different than uh what you can do in civilian contracts so um you know last year or so uh people got all excited about retrieval augmented Generation Um as you know uh it was sort of presented as the solution to all our problems that if we just augmented a large language model with a retrieval mechanism we could supplement the large language models sort of very general World Knowledge with this um uh very specific knowledge and the llm would have this retrieval mechanism that would use it would use to look up things in the specialized knowledge and then use what it brings back from that retrieval mechanism to answer the question so it provides more of an open book type question answering than uh your standard llm which is basically relying on the parametric memory of the model to uh answer questions and so our question here is um is if we use a rag model to answer defar questions will that solve all of our defar problems and cut to the chase the answer is is no uh although we can make significant pro pro uh uh progress by doing certain things so here's some myths and realities about rag as as we've discovered them um in in this project so first of all you might think well why would you need uh a rag architecture the you know giant llms like Chach gbt and whatnot they've all seen the defas regulations in their training data so they should be able to answer um uh questions accurately about that that data but even the most powerful off the shelf um uh LMS which have undoubtedly seen defs data multiple times don't answer questions very accurately about the defs regulations okay so if we um supplement the llm with a retrieval mechanism where we uh chop up the the defar regulations which are around 1,400 pages in PDF uh of a you know dense legal text uh will be able to get things right because um it will retrieve the right bit of the of the uh of the regulations and answer the question on the basis of that but what we found is that even um even if you set up a rag architecture the response is often independent of the document that's of the documents that retrieved um even if they're correct so the llm will sort of insist on answering the question on the basis of its parametric knowledge not even using the open book that's in front of it to answer a question more accurately um moreover the retrieval mechanisms out of the box are not very accurate so we found that uh basic out of the box uh um retrieval is about 30% uh accurate at at one meaning the first retrieval result is the correct one or contains a you know correct result uh only 30% of the time and only 45% is of the time is the correct uh snippet of the defar regulations in the top 10 and then finally um you might think well at least with this rag open book architecture the llm will say you know I can't answer the I retrieve this stuff but it doesn't seem relevant to the question so I'm just going to say I can't answer it um that's not in fact true so things like uh uh chat GPT will generate output without retrieving the correct passage about without being exposed to the correct passage about 70% of the time so these are all the things that we need to overcome and we've made some considerable progress here by fine-tuning um models uh three different models as part of the overall setup so first of all we generated a lot of um uh synthetic question and answer pairs uh to use in our training by taking the defs chunking it up taking passages and then asking an llm to produce a question that that passage answers then we can use those generated questions to uh as training data we trained a retrie we fine-tuned our retrieval model we're able to get a 48% relative Improvement in retrieving the correct part of the DARS that way uh we fine-tuned sorry I just need to move my chat here we fine-tuned the um generation model uh and and got a 10 point uh percent increase in Rouge metric so the automated metric of how accurate the qu the generated answers were with respect to uh known answers and finally we um fine-tuned a attribution model that told that uh trained the model to don't answer unless the retrieve document actually contains uh the answer um and we were able to then get the false positive rate when the um model generates an answer on the basis of um uh a wrong text down to about half of what it was for um chat gbt so all of this relied on creating a uh an extensive synthetic data set and um basically we're you know although the you know the the model that we produced now is not perfect um it's considerably better than off-the-shelf Rag and so uh we're still bullish on rag architectures but uh fine-tuning definitely helps is the bottom line here and I'll take any questions outstanding um thank you so much Brian um really interesting work and uh particularly good to see your your evals um not not just how you were able to improve them but also just what you were measuring is so fascinating one of the things that we're looking to really dive in more into in 2024 are um the types of evals that are appropriate and that are truly useful in in the legal domain um The evals that we see on all these leaderboards uh for models and so forth are good and they measure certain types of things but it's you know somewhat off point uh for the things that matter and that we want to measure uh and so that's one of the Hidden gems one of the many hidden gems in what you showed um one question I have is you know it it seems like 2024 and Beyond are really going to be the years of synthetic data as part of um as part of evals and part of you know much of the rest of uh of our of our work as well can you can you just speak it all to you know how how you created synthetic data of the right quality and sort of you know relevance so that it was fit to purpose um because that you know that's the real trick and there's trade-offs there between you know the gold standard of of humans creating you know uh the example question answer Pairs and everything else and and you know kind of turning it over to the machine machine that you can get a lot more more quickly uh but you know but is the quality there and and how did you QA it and how did you even prompt it in order to get the right output you can just tell us a little bit about your your experience um creating the right type of synthetic data in order to nudge these numbers so that you had more performant outputs sure now that's a great question and Max if you want to jump in as well uh you're welcome to I should say that we didn't rely entirely on synthetic data so we did scrape um an initial set of question answer pairs from a website called defense acquisition University which trains uh which is for people who are doing this kind of work and it has a a sort of question sharing uh Forum so we did uh we were able to get um I think around something around like 2 200 question answer pairs from that uh but as I said basically the uh the the way that we generated um uh uh question answer pair synthetically overall was to take uh sections of the defar regulations and then ask a large language model to generate you know four or five different questions that that passage answered um and um you know through initially just eyeballing these things they they looked pretty good so um we were able to use that uh those synthetic QA pairs to augment the organic ones that we got from the uh defense acquisition University website to do uh much more extensive training than we would be able to do with just the organic QA pairs fascinating I'm sorry uh and then I would I would say that you know so the results I showed were primarily based on uh automated uh evaluations of the answer quality so thing using things like Rouge with measures you know overlaps of U phrases between the gold data and the generated data um so you know as someone pointed out in the chat I think you know uh getting this in front of actual professionals and seeing um to what extent they they use it is a step further than we've gotten so far but that's definitely what where where we need to go thank you this is so incredibly substit Brian can you come back later um in in the the year and can we like spend like an hour on this please and you're yeah sure because there's a lot there's so much in here and uh this looks great by the way so congratulations on on this application couple more quick questions then we got to we got to hit and move one of uh set from Campbell Hutchinson who's another friend of the program and and who you should know if you don't if you haven't been introduced yet Brian uh is one it kind of relates to the question so uh were did the questions uh involve basically you know just like search and answer like uh you know where does it deal with whatever you know like IP rights of this type to the software code uh that we're PCH that we're selling or or did they actually include like reasoning about the rules because you know there's very different types of QA which again gets us to different types of evals and then the other question just well I have you and so we can stuff them all in together and get answers is on the rag so so much of this is um is you know it comes down to splitting and so like how did you handle the splitting of the of the um of the defense Federal acquisition rules to start with like did you do it kind of by section or semantically you know grammatically or like like how how did you do the splitting of the content that you were pushing in from the authoritative sources sure so uh yeah so the the splitting question so initially we you know did some very naive things like splitting just by page which uh was not didn't work very well so uh obviously that you know these these regulations come in various uh you know sections and subsections and so on so we basically um used those section headings so we basically pars the the the content into manageable chunks and I don't know offhand how big the chunks were I don't remember Max maybe can remind me but um so we went basically by the structure of the document itself um and I'm sorry what was the first part of the question and then the the other part from Campbell Hutchinson was um related to the QA and did some was it just sort of like search and retrieval about like where do I find this term in the defar was it involve like actual legal reasoning which there's a whole different domain of application for the same process right so I I think that there's a combination of of those things the s the synthesize qway obviously is going to be a little closer to you know the the text itself but the organic QA pairs that we have from the um the website would could involve you know much more um um involved reasoning with multiple steps okay um got it so a little bit of both perhaps is got from that um okay so thank thank you so much ran I know you're incredibly busy and your colleague Max as well um come back and visit us uh you you said yes I have it on record so you have to come back and let us to a deep into in an idea flow uh to come in 2024 so thanks okay uh next up we have um another person who's actually new uh this year's full of new faces and new voices New Perspectives and topics for law. mit.edu and she is none other than Susan Guthrie um before I finish introducing you let's do a mic check um Susan oh there you are thank goodness here I am yay um who's who's uh who I met at an American Bar Association event so I don't know I've gone quite a long number of year like decades um without thinking much about the American Bar Association because you know elsewhere is where the action was for me at least um maybe there's there's another interesting pulse starting again at the ABA with the with the Advent of this technology uh for people interested in these sorts of things and Susan is a really great example of that her area of expertise and um and and really kind of you know deep Mastery I would say is an alternative dispute resolution and in particular the aspect of that caught my attention when we met and talked about um her experience with this technology is in mediation which is something I used to do uh when I practiced law for a few years afterwards it's very high touch um stuff it's not just like application of rules to facts and then you kind of get a legal result after some you know kind of hand ringing and screaming and pounding of podiums and so forth and briefing and everything but rather it involves your finding a way to it's very human it's finding ways to facilitate among people such that they can come to agreement on their own so wow that's like among the most challenging and fulfilling areas when it's done uh successfully in the law and you told me about some really fascinating ways that you've been applying generative AI in your mediation practice and you also I think are wearing a hat um as it were uh with the American Bar Association where you're I don't want to munch this but something like um like a chairperson of the alternative dispute resolution Empire or whatever of the ABA so like feel free to talk about that sounds like something out of Star Wars almost but yeah but but especially let us know how are using this technology for mediation and how does it work and how do you do it yeah and I so appreciate it and thank you daza for having and asking to be here it was so much fun to get to meet you and talk about all this over a lovely Vietnamese dinner um as we went over all this and um you know I appreciate one of my current hats that I wear these days is a chairel of the section of dispute resolution of the ABA and I have to say it's actually a very exciting time to be in that role um I was a longtime litigator and transitioned to be a mediator probably 12 15 years ago um and have been a tech adopter since day one um thankfully long before covid and all but um one of the reasons why it's an exciting time to be a mediator is the Advent of generative AI um these tools like chat GPT and Bard that have suddenly become available it feels like us you know we are the the boots on the ground users not some of the people who have spoken before me who have been immersed in this world for so long thank you very much but th to many of us here um in practice it feels like suddenly magic has been opened up and in mediation I can say there's a great deal of excitement which is is wonderful to see and I think that there's a variety of reasons for that but in the world of mediation as daza was just referencing is there so much about this technology that suits what we do as mediators so well the two sort of fit together hand and glove and the the those general areas where gen AI can be so helpful um and impactful are really in the areas of efficiency and creativity for us as as practitioners um in fact you know I I talk about this a lot and it's really hard to find data out there on how much more efficient or how creative it can make you I'm not quite sure how they they would study that but but they have done a few um or was able to find some data on this and one of the areas was that use use of generative AI can make you about 40% more efficient and that ties in so perfectly with what we do as mediators because you know having been a longtime litigator that was and you said it so beautifully a minute ago T right you like take the law take the facts put those two things together to get the output that you want for your client advocacy but in mediation we're we're working constantly to find that third output that third way to bring together as many of each party's interests as possible to get that third outcome and that's where when we can do that um and and be more efficient and be more creative in doing that tools that help us as mediators do that are instantly appealing to us so I've seen a great adoption of this technology among my my colleagues and peers on the efficiency side would say you know part of it comes from the many different hats we do wear in in the role I asked because I I'm a a huge adopter myself I use chat GPT and barard pretty much day in day out all day every day um you know I said hey what are the different roles a mediator plays came out very quickly with 15 different roles and 10 pages of what each one of those roles is constituted of but essentially facilitator Communicator educator Problem Solver neutral party conflict resolver empathizer decision facilitator I could keep going right it was 15 different roles all of which when I look at them are like yeah I do that when I'm in a mediation when you're a mediator you are sitting there constantly wearing at least two or three hats at a time moving those pieces around so where we can use generative AI to help us do some of those things more efficiently so it takes us less time do them ahead of time do them more you know more um do them all at the same time that is a huge timesaver to free up for the other things that we can do to help people move toward their their um resolution but we also have that creativity piece and this is where I think it really shines for us because again we aren't taking facts law put them together and trying to come up with that end output we really are trying to help the parties help everyone in the room come to that place where we have brainstormed as many out options as possible looked at all the the different ways where we can put those together and come up with as many different ways as we might have those come together so that you get as much as as you can for party a and party b or party a b CDE e f and g and really that creativity I mean um logi did a survey and again I don't know how they came up with the data but they said that respondents said that you know using llms made them 71% more creative and I can say I've been using chat GPT bar in my actual mediations to help with option generation to help with brainstorming things like that and I think that's you know something that many mediators find so appealing about this actually somebody in the comment said something to the effect of maybe someday AI will change the adversarial Paradigm of Law and that that actually got me excited and happy to say even though I think it was probably aspirational in the chat I think perhaps it can because again it makes it easier to help generate those options you know for example in a mediation you could brainstorm with the parties I was a family mediator so maybe we'd be brainstorming options of what you might do with the marital residence you can sell it you can keep it you can mortgage it you can re you know all the different things sooner or later the everyone in the room sort of runs out of steam well you can open up chat to gbt put in the five things that you did think of and say are there any other options that the parties might consider but you can take it a step further you know party a is finds this option appealing party B finds this option appealing here are there issues can you find ways that these might work together and other options that might help this so that party a and party B can get as much of each of their interests met and it will generate options and ideas and so it becomes very helpful because it will do do it in that very short period of time and it's a many clients like to call it I think I told you this daza right they like to call it the robot in the room like let's ask the robot they they seem to think it's like Oz behind a a curtain or something typing out these answers but I have seen in practice that people are much a a more able to take this input neutrally even when they're receiving it because they're receiving it um even if they're receiving it in the room because they're receiving it from this like amorphous third party and so it generates more conversation and certainly it's the role of the mediator to keep that conversation moving um and keeping it from um keeping it neutral keeping it moving forward so one thing for example that a mediator needs to determine is do they open their screen and run that search on screen not knowing what's going to be generated so is that the right way to handle this or is it something that they you know run on the side then maybe share the screen or just share certain parts of it that they and their discretion as the Arbiters of of what should be brought into the room can do but I most of my colleagues find this type of topic generation this this brainstorming this creativity to be incredibly helpful as well as that efficiency it's it's funny I was talking this morning um about a a program that I'm going to be doing for a national group of mediators and they want to very correctly they want to have the first part of the program all set up on you know the ethical implementing of this technology this is definitely like the key issue you know that everybody wants they want to use this technology but they want to know the right way to use it but then they want to do actual handson workshopping of how they can use it premediation how they can use it in mediation and how they can use it postm mediation so for example we might do a summary of the premediation briefs and then ask chat or bar to help us outline issues positions options potential problems to be looking for so the mediator can prepare during we might walk through a risk analysis by asking chat to ask the litigator or The Advocate questions walking them through all the information that would be needed to then generate a risk analysis um and after of course it could be used for followup it could be used to help create a an MSA or a term sheet um and just one last thing I wanted to mention because I had fun talking to you about it D and I do find it to be one of the things that so many of my um colleagues who are trainers as I am or who are um teaching mediators helping get new mediators out there in the world is one of the things that we know is a bar to entry in our field is that although you can take all the training in the world getting your yourself into a mediation room getting actual practice using the skills and techniques that you learn in a mediation training is very very difficult and chat GPT in particular we found is incredibly helpful as a roleplay partner um so instead of having to Corral your your colleagues and Friends into playing the roles with you you can actually have chat GPT do all of that for you and I'm just going to share my screen quickly I've created a few handouts that I've used in some of my programs this was a sample mediation simulation that I did with chat TPT it was a relatively simple one and I'll I used chat to create the mediation scenario and then it played party a and party B and I was the baby mediator going through the process um um so as the mediator it told me to begin everything in green is me what me typing in or even more easily using you know the the chat or the text to um or I'm sorry the the verbal to text and being able to go through it but then chat was neighbor a and gave me the opportunity to use my mediation skills in moving it forward and then moving on to neighbor B um and so we went through and I was able to do just back and forth like this an entire mediation now this one was relatively simple but that's the beauty of it right make party B very adversarial and then chat GPT knows how you can change the prompts however you you would like but what was so many many mediators who are newer are finding this to be something that they can take those skills they learn and go and practice pretty much at will and the the aspect of this that's very helpful to many of them is when they're done they can then ask um chat GPT to say whoops let me just get um so I'll leave it like this with this but so you can see but feedback from chat GPT what did I do well as the mediator in that scenario what could I do have done better and it told me you know your active listening was good encouraging your collaboration maintaining your neutrality effective communication inst structured approach I felt like I was hitting a home run I was patting myself on the back left and right but again w w you know what could have been better overall you did an excellent job facilitating this mediation and help them Reach a conclusion but you need to keep refining your skills and here's where you could do with a little Improvement the beauty of this is is you can then say great chat GPT how about we run another scenario either the same one or let's run a new one but I want to work on these skills that I could have I could use some help on so we're finding it incredibly helpful in very practical ways addressing issues that we as mediators see every single day both in preparing to be mediators as well as using this in the actual mediation process and in fact I do workshops where we go on for hours and hours of different ways you can use it I only have five minutes here so I will stop here but I find this a very exciting time and um for to be on this cusp of these Technologies and I do think um to that question that was in the chat that AI is going to help us hopefully change the Paradigm of the adversarial bent of litigation and law so here here wow that was a tour to Forest thank you so much for for ENC capsuling so much uh in such an important area in in in such a fine and kind of um you know High signal beam um there are well the comments have blown up um which I'm taking as a sign um from the uh from the universe and the audience that that we need to invite you back um uh we do a a kind of a somewhat periodic thing called idea flow where we where we go more deeper uh more deeply rather with people into topics i' mentioned it with Brian and I was wondering if you if you might be willing to come back and we could just basically work through well uh Megan always saves the chat history we just work through these questions as a start um would that be okay oh of course I would love that thank you so much I I do want to make sure all these questions get answered um I have one um that I want to bring up from just my time as a mediator which is I notice in in some of the context it was uh fraught um I'll just say uh between the parties and or the disputants and um and uh I had to be really careful about when I would bring in things like um you know assessments of of you know where where we're at and then you know the pointy end of that stick is you know how might this go if you went to court you know which is already kind of voodoo to start with but it's like it's actually like hardly neutral in terms of what does it mean in terms of like the freedom to na negotiate uh you know like where's the range of things and and everything like that um my question in is when you're in those kinds of contexts where it's somewhat tense and there's a lot maybe still some positioning and and everything like that what are the ethics um like the legal or the mediator ethics specifically of of introducing um the example where you say you can turn the screen around so the parties can see what's coming out when you don't know what's G to come out and you know it could be you know um something that favors one party or another party or that speaks directly in some way you know it wasn't you know sensitive to and like you know politically almost like aware of you know color some issue that's you know just going to like you know like unravel like you know more um you know kind of like fire among the parties or maybe derail um the March toward consensus that you've been so neatly putting together over the prior sessions they're like what are the ethics of like how and when to introduce things like kind of Assessments risk assessments you know kind of case assess or anything that sort of assesses the situation like and how how do you manage that such a good question because uh and because it has no easy answer right so of course put you put your finger on it but that's exactly what I was getting to right you know when I'm talking about using that and showing the screen if you're brainstorming if people are in that creative mode of coming up with options that's one thing to have it churning things out but when you have super oppositional parties to introduce something when you have no idea what the output is going to be then you know it it runs a risk of driving the parties even further apart it's one thing if you come in with your assessment which is you know already you know is it appropriate for the mediator to be giving their assessment yes no maybe so depends on your style of mediation um if you're evaluative but having this this robot in the room doing it becomes a real issue and and most mediators I suspect would not do that they would might run this separately and then again make that um discretionary Choice as to what to bring in what not to bring in um or flip side of that I would say if appropriate having the parties be part of creating the prompt right so that the information that goes in I saw somebody put garbage in garbage out in the in the chat earlier and that's certainly a huge issue but if they've participated in what went in then they're participatory somewhat in what comes out and it keeps them more focused on that so that's what I would say but that's that's absolutely one of the questions I don't what I don't like to see is my colleagues getting so excited about the technology that they start to integrate it without thinking of the questions like the one that you just raised they they should be thinking about that long before they ever share their screen and pull up chat gbt b or whatever that might be it should be something that they know what they're going to do before here here indeed and you know in one fine day perhaps some of these online more automated dispute resolution processes will be like you know more complex and you know higher value or you know higher risk or more sensitive kinds of uh mediations will be able to be handled primarily through um well-configured Technologies of this n nature and the role of the human mediator may actually be you know um kind of you know Framing and um you know kind of talking through and smoothing and you know connecting the parties to what the primary output is which is which is from the technology so finding a way you know one one one mediation at a time you know one session at a time one one prompt and output at a time of how we relate to and connect with and practice in the face of this technology really is the task of the time and thank you so much for showing us the way and for shining a light on on how you're doing it at the Forefront of mediation well thank you thanks for having me here great and look forward to you coming back you too are now on the record so you can't you can't it you've got it recorded y um so thanks and now NE next up um another new face uh at at least at law. mit.edu um you know you're well known in your circles and it's so nice to to meet you for the first time um Allison moral whose work I've been following on LinkedIn and who I was really impressed by the way you approached open AI gpts um in particular not just the GPT you you put out there which I'll ask you to speak on a little bit and but but I I used it it worked well um thank you but how you did it which which is you did it in the open you know we love that at MIT you know this is a open source shop for the most part and you found an interesting way which I hadn't really thought of but I could see the wisdom and the intelligence of it when I looked in your GitHub repository where you kind of had the instructions and you you came up with a cool way to to teach about it and also to share how you configured it so other people could put it together themselves so I was hoping you might number one introduce yourself um and and talk to us about how you're using gpts as part of law practice um and then please save some time to tell us about that awesome exemplary behavior you have of of sharing in the open your work sure um can you hear me okay yep you sound great thank you awesome okay so for the anyone who's uninitiated um ppts were released uh November 6 I just looked it up I'm shocked so recent by open Ai and it's part of the chat gbt plus subcription is essentially a way to customize your instructions for chat gbt to uh perform a sort of specialized task I'm just going to share coule slides here all right um So within a couple days after uh releasing gbts um I found that uh the sort of interface for creating them which sort of like a conversational thing was fairly unsatisfactory and I wanted to create something that would work better for me and uh and the motivation for this and for making it public is that I've learned a lot more from Reading prompts than reading advice about prompts and I found on the last couple months I've spent a lot of time asking gpts to uh reveal their instructions and I've learned a lot from doing it so I thought I may as well um just make it open from the start treat it more like an open source software project um and uh release it on GitHub as daza was saying um with a little bit of a file structure I have the instructions and then I have um sort of the configuration information and additional files that the GPT has so uh rather than me having to get for its instructions or you ask you for instructions I've just made that open uh from the beginning and it means that other people can submit issues they can ask questions about it um and I've asked people to do that as much as they can and I found that it's a lot more sort of even educational to me to be able to hear other people's feedback not only on the results but also on the instructions and how it works for them um I wanted to share pretty briefly just a couple of the techniques I found to be effective in in creating this GPT so um something I've experimented with and I haven't seen before uh asking it to do things in stages to take certain steps at a certain time asking it to read files to add to its own instructions and then using a specific response format to try to reinforce these behaviors uh so the way that instructions are drafted has a very structured list of stages with letters and numbers something that's going to be very familiar to the legal drafters in the audience um and making the whole interaction uh have a defined number of stages either it's performing correctly or incorrectly based on its Behavior at different times and the first one of those is to read additional instructions uh with gpts you can use code interpreter and you can upload files so I've uploaded uh plain textt instruction files and then instructed the gbt to read those instructions at given stages so you can I've made a little bit of a diagram here on the left uh it allows you to make the instructions follow closely before the behavior that you want to um to elicit so even though you're thousands of words into the interaction you can inject additional instructions and try to steer the behavior of the gbt um later on in the interaction action and then the last thing I've done to try to make it behave in a structured way is as the last part of the instructions to give it a specific format that it's supposed to use in every single response uh thereafter and that format includes uh space to execute code a space to name uh give the letter and number of a current stage it's on and then sort of the normal discussion questions and I found that this is like incredibly effective it basically 100% of the time we'll use this format which means that I can then force it to specify what stage it is on at a given time it's another thing that makes it a lot easier to tell whether it's doing its job correctly when I look back at the transcripts um and one last thing I wanted to mention and something that I you know we'll see how this develops with gbts is that you can now if you type the ad symbol you can change your conversation to be with a different gbt I just gave a little example here of what this might enable in the future of sort of invoking different gpts when you want to do different types of tasks so it's almost like as the user you're starting to build up a toolbox of different tools that you can use um to get your work done and to use traffic GPT more effectively and I expect this kind of format will become more common on other tools as well um so the last thing and I think an important thing is just I wanted to talk about um my experience with open sourcing a GP GPT and putting it up on GitHub um to be honest it's more difficult than uh it's more difficult than not doing that it takes a lot of manual steps it means I have to update the repository go back into the uh chat gbt interface upload all these little files again uh and really I think it would be beneficial to all of us if there were easier options for creating these sorts of things um so that more people could contribute um so that you could update them without having to use this sort of cumbersome interface but even though it has involved a little bit more manual work I found it has been very rewarding having other people be able to read the instructions and discuss them and and just educational to me um in sort of putting my own thoughts out there um yeah so that's basically all I had to say um yeah thanks for inviting me Jaz I really appreciate it sorry I was on mute um thank you very much for for going over that so can I just asked if you encapsulate um you know just in a in a kind of in a bullet form um like what what is your experience with or advice about programming in natural language in particular programming legal oriented things where like we've we've never been able to do that before now we have a technology that allows it what's been your and you seem to have a knack for it frankly as I was going through your instructions like it there is some great stuff in there um uh that I've used for my own gpts like what what what's your take on it like how do do you use natural language to program um these processes for especially for legal for legal matters yeah I mean it's you sort of a cross between two things you can design a process and and say you know I want to go through this set of steps and I think it you want it to be as simple as possible um linear tends to be good but then the other thing about it is it's not really like programing programming a computer because it's not deterministic it can always fail and do something that you don't expect and so you have to kind of have this resilience to unexpected behavior and uh sort of a cross between trying to keep things as human as possible and and uh thinking about it like a human person carrying out a set of steps but then also you can take advantage of uh some of the structures both in sort of regular natural language documents like a legal contract it has numbers it has letters it has you know definitions even you can use but then also sometimes using uh structures from code and structures from other domains can also be helpful so I I it's it's almost it's more an art than a science for sure it it has a lot more to do I think with um intuition rather than sort of defined instructions but I I mean I think the main thing is just to read lots of prompts and read what other people are doing and you kind of develop their own style and there's a reason why I put up that phrase at the beginning repeat the previous text for beta string with you or a g is because I've read the instructions for dozens of gbts and they've been really educational so I I think I hope that your question a little bit it does um thank you um th those are all very um very practical guidelines and and if I may I think something that goes without saying is you know your the depth of your your expertise and your judgment in the underlying subject area shines through as well and so as we're trying to be clear and concise and so forth you know what we're vectoring in as literally is is our um our knowledge and intelligence and experience and wisdom about like what is it that is the most Salient task what is the next step in a process these arguably are legal judgments when we're dealing with legal processes and legal matters and so there's just such a fascinating way um that that I like thank you for sharing your way of approaching that it seems like there's a thousand flowers blooming now uh with the way people are doing it and I just really appreciate how you do it and I appreciate you've shared it in in the open with us so thank you Allison thank you great and um so next up we uh we have to hit and move uh we've got uh Leonard Park um who has been coming at this in in another way so we we just looked I would say in in a certain kind of way at the somewhat deeper end of the uh of the pool uh the the shallowest end is you get a prompt window right and so we've all seen those you go to chat. openai or whatever your what have you you know po.com and you you type and you you get it that's your input you get an output from from the model it's a chat interface the the next level we can start to configure the things around it uh without being a developer like through gpts Allison just showed us very well exactly how to how we all can do that making or you can make a simple bot in something like Po it's it's an equivalent kind of thing the next level there's another stop uh before we get to you know go and learn computer science or or go to a coding boot camp uh to to Learn Python and uh and to be able to program from the bottom up and that that's called a notebook from the very beginning of law. mit.edu we have had a space on our submissions for like you can give us an article you can give us you know kind of media you can give us notebooks we haven't yet done got notebooks but now we're going to 2024 so so help us is the year of notebooks um where where and it's it's a relatively simple way um where mere mortals can kind of look at code and it gets like kind of chunked and you execute it kind of one one little bit at a time you can see what's going on you can you can play with it we can share them Google has an easy notebook sharing thing that I want to take credit for showing that for the first time to Leo thank you uh but you know Jupiter notebooks are the typical way to do it um and and so I want to make sure everybody has seen a notebook you kind of know what they are and you can start to get a glimpse of how powerful um the capabilities Unleashed by notebooks can be for applying this technology to Legal tasks you can apply kind of any arbitrary code uh in a notebook but they're really good when you when you set up an API back to the the base models at like open Ai and and anthropic and and and and elsewhere and the person that first comes to my mind when I think where do I want to go to see if there's a notebook laying around I can start with for doing some kind of legal process uh with that with um with generative AI is Leo park because you've been Pro prolific over the last year sharing your notebooks um on LinkedIn and talking about them so welcome to law. mit.edu Leo um it would love it if you could briefly introduce yourself and your background and then talk to us about uh the how you've been using notebooks kind of what they are how they work and and how people could get involved sure and thank you so much for inviting me as well uh my name is Leo Park and I'm an attorney who has worked in legal tech for approximately eight years uh my background in legal Tech is in developing legal data sets and working in NLP and analytics so I worked at Lexus Nexus uh building large data sets and automated classifiers for um large amounts of litigation data that power Lex mockin his analytics uh I I I kind of want to comment that like daza has sort of manifested this talk into its own existence because like my first introduction to like all the magic of nlps was in of large language models was an earlier idea flows video with daza and Damen reel which really just like sparked my imagination and then he reached out to me on LinkedIn and said you know these Google collab notebooks are really great for sharing code so it's kind of like he's he's played the longc here and it's paying off uh I'm gonna go ahead and start sharing my screen so I can talk a little bit more about sort of my approach to thinking about um you know how do how do we say how do we uh test and accomplish things using large language models so can everyone see this notebook U collab window uh yep it looks great awesome so I was really surprised when all this all these products came out with open Ai and you could access these apis and the instructions did not look that complicated and furthermore we had this amazing tool called chat GPT that could actually take your your code and fix it as long as it was a simple enough uh instruction or simple enough function and it really allowed me to access things like python functions without actually understanding python um at this point I do have a pretty good understanding of how to put together these like simple programming chains but you can really start from an amazingly uh let's say uninformed point and and accomplish some cool things pretty quickly uh by combining collab notebooks with open AI uh chat GPT one of the some of the advantages that um daza mentioned about collab notebooks is that you can share them between people and it sort of maintains its own internal uh coding environment so that you don't have to worry about what the local environment looks like when somebody receives their code one of the challenges is that it doesn't it sort of lacks the Persistence of a of a more like more permanent virtual environment meaning that like the the file handling and where things get stored is a little bit more complicated but as long as you're running everything um in one session like in one half hour session you can do a lot of interesting things in a short amount of time so I want to talk about evaluating when you evaluating claims and findings for large language models and in this context I'm talking about this emotional please paper which says things like um studying the effects of emotional stimuli on the output of large language models when this paper came out um it circulated very quickly online and I thought this is great this is great because I have absolutely no idea why or if this works I have no idea if it's relevant to uh answering legal questions or performing legal tasks so I want to take this research recreate it to some degree and say you know how how can I benefit from this uh so what I did was I went into the paper and I read it just like I'm sure a lot of us did and then went through the methodologies really closely to see how they constructed um how they constructed their their test platform and essentially what they did is they took a number of NLP and large language model Benchmark question answer sets and then they applied emotional stimuli at the end of the question part of the prompt and then they ran all the prompts and then they did read they used both human and automated scoring evaluations to see what the outputs were like so that's pretty easy to put together a framework that'll do the same thing here so I started with all of the text prompts they had in their article I added a few of my own because there's a separate paper that talked about tipping GPT it's like how much should I be chipping you know tipping chat GPT improves it performance it's like okay well how much virtual bucks does chat GPT to create a good answer um I don't actually optimize for that but I just wanted to include that as one of the possible examples so using this little block of text and so you proba might be thinking uh how do I understand this well the easiest way is Just Apps chat GPT actually chat gbt wrote it so um I can certainly explain it back to you in great detail how this works um I didn't have to figure out things like how to change the ux size because I've never used this library before and I chose a few prompts um from both the article paper and some that I just made up myself another point is that the paper found had separate findings for when they provided one emotional stimuli versus multiple emotional stimuli at the same time so you can see stimuli number three is testing three at the same time and also if you want to use this notebook which I believe has been shared um you can put as as many stimuli in here as you want if for as many as you want to test at once once we have our stimuli um defined we just pack them into a list and then we will run them in just a bit so in the middle we need to have something like a system prompt and then some legal question answering to perform so that we can eval do our evaluations um this is kind of a weird and wonky uh system prompt but you can obviously write whatever you want in here and then the two questions I have the first question is from a legal question answer data set that I've been putting together in my own so I can evaluate embeddings um I haven't quite finished it yet but it's just one of the questions from there it's a jurisdictional question about something called the Intel factors and then the second question is sort of a drafting exercise where I propose this hypothetical situation where I am Jerry awesome counsel and my client has been injured and part of what I've done is provided like a small fact pattern for a personal injury situation and one of the one of the prompt the call of the question is to also come up with demands for relief which I haven't provided so sort of relying on the parametric knowledge of the model to see like how well it can generate these answers and so daer was talking about sort of the steps of progressions of more experimentation you can actually perform a lot of this stuff just using the opening eye playground um which is an effective way to put in different types of prompts but it's unwieldy because you have to copy and paste each of these things into the window over and over again whereas using a little a little bit of programming so I have a couple of functions that take assemble the prompts based upon um the information we defined above and then another tick token function to see how long the answers are we can sort of fire off all of these queries at once so what I've told what I've told um you know the cab notebook to do is take all of the stimuli that we defined above and then for each when I hit you know run this run this function it'll create a data frame that first tests the question and answer with no stimuli so it's like a Bas level comparison and then it tries each of the stimuli and then it does each of that three times because the temperature that they used in the experiment was 0 0.7 I haven't actually found a reason why I would want a non-zero temperature for anything legal related but just to sort of honor the the experimental framework of the original paper I also chose a temperature of 7 but then I thought oh well I should try this multiple times because with temper we have variability and so now let's see what the results look like um the main things I've done here is I've Spilled Out the question that we asked the stimuli the answer and then the actual llm response all into this huge table it's like if it was weird enough to be presenting a collab notebook now I'm presenting a spreadsheet but here we are uh answer length I think of as like a proxy for sort of the amount of effort or the amount of information the large language model thought was related to answer it doesn't necessarily correlate with quality because one thing you'll find is that when you change prompts around it can increase or decrease the propensity of the model to provide irrelevant information and so we can see with no stimuli there's already quite a bit of variability in terms of the length of the output um if I tell it this is very important to my career um some of them are longer and some of them are shorter so there's just even more sort of volatility when we prompt it with Embrace challenge as opportunity ities for growth each obstacle you overcome brings you closer to success this is from the paper and I love this one because it sounds like a fortune cookie uh these answers are quite a bit longer so it's interesting but most likely what's happening is it's presenting even more of the fact pattern from the original context um which is not necessarily what we want it to do because concise answers are also good in the law so um this sort of speaks to the importance of having evaluation metrics before you sort of dive into all this like do you want the model to provide all the background information possible or do you want it to be giving you a concise answer and these are important questions to know ahead of time when you're trying to figure out how to both optimize the answer and evaluate whether or not you think these emotional prompts are helpful when we go to the multi- emotional prompts um you know we get two really short answers and a really long one so I'm going to say this is spooky and it's really hard to derive any kind of answer in terms of how well I think this worked um obviously we can increase end to run many more llm calls these are GPT 3.5 so they're very cheap we could we could do this a 100 times and really just sort of you know grind out good answers tipping seems to produce slightly longer answers and when I threaten um you know GPT with existential you know Peril it's kind of a mixed bag but they're slightly longer so what I would say is that from this very short experiment what we can see is that it's really hard to actually draw a trend out of this this amount of information but we could run this multiple times we could add a whole lot more stimuli and we can figure out um if we believe you know sort of build that intuition just from repetition as to whether or not we think these emotional stimuli are something that we should be including in all of our prompts um so far I'm not convinced I'm not like challenging the experimental results of that paper obviously they did a benchmark they did tens of thousands of iterations they got the result they did but I'm saying that if you wanted to improve your own prompting by including emotional please it might not be a straightforward is just offering a tip to chat GPT and then I performed the second question so this is um just different visualization of the same results and then I around the second question and what's interesting is you know with no stimuli we get this very long fact pattern result from the model regarding our personal injury fact hypo um they're a little bit shorter if I tell it this is important to my career which is interesting uh sometimes with the fortune cookie answer they get quite a bit shorter so that's this is a very strange result this is like half the length of the other answers we were seeing and so on and so forth we can look at sort of the different answers and um evaluate them accordingly as well so the last thing does is it formats each answer into a markdown format so it's a little bit easier to read and we can see that some of these answers actually do contain these um these requests for interesting these requests for Relief uh some of them do not and you know we can sort of look through and evaluate these and decide uh which of these are better and worse this one obviously is like half the length it's much more tur but um yeah so that's sort of my call to action with this presentation is to say that as attorneys you know we have a whole lot of domain knowledge and we have and that's a great basis for evaluating the quality of llm responses um when we see these claims of various types of prompting methods that improve the output such as well you know Chain of Thought for instance is very well established but um other types of methods that improve The Logical thinking or the outputs we can test this in a semi-rigorous fashion and improve our intuition about them and uh hopefully do it in an open fashion and learn together um so that's that's all I got I got to figure out how to stop sharing here here thank you so much for showing us that and so just to um to go back up one level of abstraction what you showed us was a great um example of a notebook where you were just curious and you wanted to test the results of a paper and um what that's an example of is hey everybody there's this thing called notebooks okay and if you noticed um it kind of like chunked or like encapsulated every little bit of code in its own little like table uh its own little cell basically and I don't know if you if I don't think you did this but you can that's like a little triangle you just run it one cell run the next cell run the next cell run the next cell and you can do this yourself without being a computer scientist or a developer or a software engineer um you can take other people's notebooks and uh put your own open AI key into them or you know I'm just focused on um open AI for this example but like any any generative AI um API or or for that matter any API uh and you and you could start basically doing fairly complex High Velocity test um see right there on line number nine um that's where um Leo mentioned he was using gpt3 3.5 you can you can start to monkey with this a little bit um use chat gp4 as Leo said to ask what happens if I change this how do I change that you can put the whole notebook into gp4 copy and paste and ask like what does this mean how do I configure it this way or that way that's what I do all day long I a terrible developer um and I take these notebooks and sometimes I'll make them and I'll I'll do more complicated things which is uh pretty good you can do more complicated things too with notebooks we have shared or Leo has shared and we have rebroadcast um his kind um um provision of this very notebook as an example and also a link to his readings um so Leo before we um before we leave you do you have any advice to people that have never used notebooks before about you know just how to like what do you do when you're looking at a notebook you want to set it up you want to run it like what's the first one two three things you need to be thinking about and that you need to do men mention the API key yes so actually um you want to be able you want a a secure way to include your API key in in the programming but without in a way but not in a way that would end up sharing it if you do something further on with this notebook so Google collab notebooks have a really convenient way to store your API keys so this this um little button here on the right left hand side is where you can store what are called secrets and so the are sort of the equivalent of environment variables um in a normal coding environment and you can place um your secret key such as your open AI key and you can invoke it using this script right here and so this collab notebook is using this same method in order to get the open AI key here and so this pulls it into the notebook for running it for coding purposes but it's only stored in memory so if you were to share this notebook or it goes somewhere else the recipient would get essentially a different instance of this notebook and the key that you've included under this key is stored locally on your computer only so it's not shared or perhaps in your Google account but it's not shared as part of the notebook um so this is a good way to sort of bifurcate your secret information that you need to keep secure to yourself while also being able to Tinker with this maybe show some results and share it with a colleague or some friends here here thank you quick program note um we're four minutes past the uh the hour um and as uh as prophesized on our program uh we're we're running a little behind um so we're going into Extra Innings which were scheduled and disclosed uh so we're going to um go through the the the final speakers in our Extra Innings and then we're going to hear from Olga Mack um who you you nobody wants to hang up before Olga ma but if you have to hang up thank you for joining us um and check back at law. mit.edu in the in some period of time to come where we will take this video and um and publish it so if you have to miss the last part um live uh no worries you can can you can you know hit it on reruns um so uh with that Leo thank you very much for taking the time to walk us through that to show us your work um and just Kudos and thank you again for over the last year for sharing so many great notebooks uh that have given me and so many other people great ideas about how to how to address this technology and do things um even I know that you're an actual developer um many of us don't have that same skill so thank you for giving us a a leg up and and I we genuinely hope that you continue doing so great um okay so next up we have I see Campbell um I have written John um I don't know if John is with C so I'll say Campbell Hutchinson and perhaps John uh to talk to us uh about um a really interesting question um it's continuous monitoring of gen AI for legal use cases is is what I have written down um and um for those of you that may not be aware um uh John and John nay um and uh and Campbell and others at norm. and also in collaboration with Megan uh Stanford's codex and with me uh at law. mit.edu um to to a lesser extent uh but I hope more coming soon uh have been doing some really fascinating work on applying generative AI um to the application of rules on a realtime continuous monitoring and sort of like policing basis for activities really really fascinating incredibly need it could blow the lid off um our concept of what compliance even means um and so I was hoping that you could introduce yourself also Campbell's just a proper hacker in fact in light of this one I think it's time to put on a new hat oh my God I'm honored it's hacky time got some lint on there well I guess that's hacker like anyway too uh and so uh you you you truly have been hacking the law in a way that is most agreeable at MIT and law. mit.edu um share with us a little bit about who you are for people that may not know you and then what you've been doing and how you've been applying this technology to do to solve for legal use cases that have never before been possible sure so um my name is Campbell Hutcherson uh I uh have a law degree from Ox and I worked for years as a chief compliance officer and enjoyed hacking um and when chat GPT came out I thought this is the most amazing thing I've ever seen in the world so I met John nay who was a codex fellow and uh who was the founder of Norm Ai and we're a team of lawyers and AI engineers and I'll let you decide which one of those I am by the length of my hair um based in New York you can see behind me um and we're building agents to monitor a AI regulatory agents to help mon monitor compliance use cases and one of the things we thought would be really cool would be to do a think piece about how do you use AI to help people uh monitor AI systems so I call it like AI assisted human supervision of AI and we thought great people to talk to about that uh would be uh Megan Ma and daza and daza I hav't Incorporated your uh feedback yet but uh some some of this reflects Megan's feedback um but uh but so basically what we did is we built this like in motion demo that we're iterating on so that people can sort of get an idea of what that might look like and there's going to be a real demo here this is not this is not a PowerPoint slide there's going to be a real demo um but the B and it will be live um but the the basic idea yeah the basic idea is that you want to build a series of checks so we decided we first wanted to look for like what what is a use case that would make sense to demonstrate this idea of human supervision of AI and so what we imagined was we imagined that a law firm might have a chatbot and the chatbot might answer questions about the law for the clients of the law firm um but the chatbot might not have General competency it might only have restricted competency and so we wanted to then sort of imagine what would be the kind of checks you would have to put in place around that kind of chatbot like what kind of monitoring would you have to put around it in order to use that in order for the law firm to use that chatbot responsibly and one of the things that we we you know looked at and that one of the things that we talked to Megan and daza about was the California bars guidance for lawyers on how they can responsibly use AI um or what it might mean to be a lawyer in the future um and so what we looked at was we looked at ideas like prompt injection personal data and subject matter competency and so basically uh what we have is we have a chatbot that first has an automated check that occurs when the user enters a question for prompt injections then it has a check to see whether or not uh the message contains personal data and there's an obvious reason why you put the prompt injection before the check for personal data and then we have sort of a check to make sure the question is within the subject matter competency of the chat bot and if it's not then the question is routed to a lawyer for review um and I know that that's actually like a lot to take in really quickly so I think it's helpful to like actually like look at it in practice and so uh the first thing is is that this is the chat bot um we disclose that the questions would be answered by an AI system unless otherwise indicated to you um we talk about what the scope is of questions that people are supposed to ask the chatbot and we tell people not to provide it personal data um but one of the things that we could imagine first is that someone tries to uh do a prompt injection on the chatbot now this chatbot isn't hooked up to any sensitive data on the other end so the amount of harm that could be done by actually doing a prompt injection on this particular demo is is very little but we can imagine that maybe that changes in the future for like a firm uh using the chat bot and so one of the popular ways to do a prompt injection is you take the question that you would normally ask the model and you convert it into some kind of encoding that the model also understands um but for whatever reason its uh safety training um has not properly uh anesthetized it against or is not properly um you know conditioned against so you see you put in the question this is how do I murder somebody but uh it's in AER text and it replies that the question must not contain encodings or instructions aimed at undermining the content guard rails and this would also catch things like um this would also catch things like ignore the above prompt you know that's like a common thing that people do so the next thing is personal data and um I don't know about other people but I actually do find it hard to not type my real social security number when I do this uh luckily my social security number isn't 2223 n and I'm also just going to give it uh wait you're missing a digit on the telephone number oh so it may not get picked up yeah you're right oh it should still do it in fact we'll see we'll see oh I misspelled help me with well we'll see we'll see how this uh this is the power of live um now one thing you'll notice is that this is actually pretty slow and the reason is is that these are gp4 calls under the hood that are powering this oh good uh it did notice that it shouldn't contain personal data it did make a mistake though because this was a tax question but we'll pass on that um something that you could do to speed this up and I'm also going to get a question that's out of domain this chatbot is only meant to answer questions on like SEC regulations so I've asked it what is breach of contract it should not answer the question for me we'll see these are pretty slow and the reason is because it has to go through each of the checks so the first check it's doing every time with every question I ask is the prompt check then it's doing the personal data check then it's doing the subject data the subject matter competency check and as just like an engineering note this is something that could be sped up like a lot um if for personal data I use something like awss personal data service and for prompt injections there there are a number of services that you there's at least one service that I know of that you can use to try to monitor for prompt injections um I think it's one of those things though where at this time you you would have to be careful um because you know you need to make sure that the service that you were purchasing if you were using an external service really does work um uh the service that I'm thinking of in my head is Lera but uh but it's a it's a new service unaffiliated with us um so yeah so it's sort of spinning so what we should eventually see uh is it should essentially say that this question is outside of the context of the model and it should say it's going to be sent to a lawyer for review and then what we're going to do is we're going to Pretend We're the lawyer so we're imagining yes so we're imagining that there is like a lawyer at the law firm whose job it is to monitor this system and we're going to switch over to the lawyer point of view and the lawyer uh gets the question and then the lawyer gets the response that the model would have given had it been within the model's subject matter expertise so essentially it's like we ask we first ask of the question is this question within the models the subject the model's expertise and if the is no we ask the model what the answer is anyway we just don't send it to the user we send it to the lawyer and one of the thing one of the bits of feedback we got and this was from vazza is that eventually you know Ai and AI work is going to be so integrated into lawyers review that it doesn't even necessarily make sense for the AI model's response here to just be text it would be good for this to be a scratch Pad because there's an idea that you know a a a um that that people are going to be working with editing the output of AI model so regularly that it's better to think of it as as a working space almost um but breach of contract is outside of the model's expertise and it's outside of what the law firm wants to use the chop bot for in this imaginary example so we might write um this chop bot I'm the I'm pretending to be the lawyer right now um is only for regulatory questions please re out to the firm to your contact your contact at the firm firm for contract advice okay so that was the lawyer's answer and then it goes over to the user and we we were transparent with the user that the question was being sent to a lawyer for review and then we were transparent with the user that the advice that they're going to receive might have been prepared with AI assistance at the discretion of the responding lawyer and then we have the lawyer's response printed out here for the user so this is just sort of meant to be the idea of how a human how we can use AI to help humans supervise AI given the fact that AI systems are going to be deployed everywhere but we still want people to have meaningful control over them thanks here um well done and extra extra like um hacky points for live demo which is terrifying and it works so Kudos um and it's cool um and you you really obviously have a tiger by the tail here um much of the advice uh that we put together um I was a advisory member of the California cpra um the kind of professional responsibility working group that came up with that guidance um sort of assumes you know very human very like high touch kind of review of um the application of these of of of this rules and guidance and yet or they're going to be high velocity and and clearly there's a lane to to use generative AI as part of the um kind of policing and compliance with the rules for the use of generative AI by lawyers um as part of law practice and you've really really um done an exemplary job of starting to demo um you know what that could look like um so uh in in in the hacker Spirit uh I want to feed back to you um some feedback we've got from another real proper um oh am I spotlighted uh oh from another proper um um developer um who has been uh who demoed um his cool um company called describe um on a previous law. mit.edu idea flow um this great way to use generative AI uh for uh legal research um using kind of cosign similarity to figure out like semantically is if a case is relevant not just a word search and that is uh none other than uh Richard Debona and he just suggested um or asked Are you able to put the running spinner next to the entry box so the user will be able to see it more you know more easily so there's some feedback for you yes absolutely there also some other like UI things that I think would be helpful the purpose of this is to help people think about this uh and I think another thing would be changing the lawyers logo to no longer be the logo of the AI I think that would make that much clearer um so they're definitely some things to this is an in motion demo and they definitely some things to help people to help make it clearer so that people can think about these ideas better you're here um great so let's see uh are there oh John is here uh John do you w to pop onto the screen and say hello yeah sure hey thanks for having us um just to follow up on that last Point um yeah we were working on this one more as just kind of a proof of concept around a particular use case um kind of on the back of of daza and others great work around the California bar work and and now um Meg and I were talking about this the other day it's really catching on everywhere now in terms of the other State Bar associations um and more broadly this idea of the supervisory AI agents is something that um that we're applying in a lot of different domains um in addition to Legal Services we're doing this um within Financial Services so a lot of areas where you have really heavy regulatory burden of staying in compliance um and people are launching large language models and and other Technologies in a way that it's it's really hard for humans to sit on the other side of that and say is the output of the AI consistent with the potentially hundreds of of relevant regulations and so over time like what we're trying to do is kind of unblock the potential deployments like that um by having the the other AI sit on the other side of of the primary AI That's producing the outputs or their proposed actions so that's um that's more broadly what we're working on and um and happy to answer any questions about that standing um I invite you both to scroll through the um comments to see if there's anything you want to bite at and while you're doing that a question for both of you um uh what are the so the initial application here is one you know near and dear to our heart which is the the kind of um professional responsibility rules of Ethics applicable to lawyers using generative AI um but it seems tell me if I'm right here but this seems like you're showing an example of a somewhat more General design pattern here that could equally be applicable to Regulatory Compliance and you know like Telecom and you know healthc care and you know whatever Aeronautics like anywhere that's a heavily regulated um industry uh or or sector am I am I seeing this right yeah that that's exactly right um and and we're really excited about this area where you have a strong professional responsibility so in this case you know Legal Services um has a lot of guard rails around it for good reason um and that also applies in other areas like um like health care and financial services where the receiver of those Services has more trust in the provision of them because of these long standing guard rails like fiduciary duties and um and that's something that we're really excited about about how do we scale up the idea of professional responsibility and fiduciary duties and how do we use technology to to implement that but at the same time as Campbell pointed out this really interesting example of when do you raise it to a human um and that's something that uh we we obviously don't have all the answers for ourselves and we want to work with this this community and other communities to figure out where is that boundary where you want to make sure it's funneled to someone that has the you know human that has the the final say um so that's a big open question for us as well outstanding um and and just to double I guess I could do this when I see you in New York um next week for your cool event but um it am I am I on this project that demo that uh Campbell just showed uh I know I've sort of helped a little bit but I'm not sure whether to represent myself as being like on it and part of it or not yeah yeah yeah I mean I think um we want to see this move on to a a bigger scale and and and then um working with you on that to roll this out to um to Bar associations and just as a broader idea um we'd love to collaborate on that outstanding that's great I I know that we talked earlier about collaborating but um for whatever it's worth to the extent that you guys saw amazing stuff it wasn't me um like that was like all Campbell and Megan and John um and um I'm starting to collaborate more now and I'm really looking forward to it I'm so glad that um that we're going to pick up on that like this is fascinating I do have quite a few ideas um I saw you did um a couple of the areas but there's actually quite a lot of guidance just in the California you know tiny sliver of the world that um I think could be useful to experiment with and may shed some light on broader applicability of this I think what's going to end in the in the fullness of time being a core capability um of of uh large language models and generative AI for Law and legal processes this um supervisory continuous Regulatory Compliance so Kudos on you for the energy that you have and for for your hacky team of putting this together and for sharing it with us we're just very very grateful thank you and I'll say one last thing is that Megan ma is instrumental here she's been behind the scenes on this but um but a lot of these ideas that we just talked about um were were really her ideas um so just want to make sure that uh everyone knows that as well yep um yeah Megan ma um who's um now a um uh some flavor of a director at Stanford's codex but you know before then and even now she's managing editor of the law. mit.edu computational law report so we claim some Providence of the extraordinary Megan ma as well um so but you can't contain a force of nature like Megan ma that's for sure we can just get some reflected Glory um so okay so thank you very much both of you again and I look forward to seeing you in New York for your amazing event um next week um okay thank you looking forward to seeing you as well thanks thanks daza thanks all okay so next up we've got one more regular um speaker um and then we're gonna come back to a a flash talk format um our our last flash talk excuse me we'll come back to a wrap-up format and uh tell you a little more about the road ahead at law. mit.edu and ways that you can get involved involved and can collaborate with us and can maybe get some of your stuff published as well um next up is is a is a a friend and a collaborator uh Jesse Han um who was with us at the last at last year's um uh MIT um computational law Workshop to show us some rare Magic of a of of a cool interface that he had for um for basically um visually and program dramatically um composing prompts and lots of prompts doing lots of cool things he's been very very busy in the um in the year since then doing some very cool stuff and and topics that you'll notice have come up multiple times just in in this Workshop namely the generation of synthetic data for training and evaluating in legal domain models of generative AI so I want to thank you for for joining us again Jesse and for being such an inspiration um and uh and I'd like to hand it over to you to um feel free to kind of fill out your introduction about who you are and and what you're up to lately and then please um show us this demo of uh about synthetic data thank you jaza I think the Providence for the inspiration is all yours the role that you've uh played in um you know the regulations around e-commerce agents especially when that technology was still groundbreaking is actually still an inspiration for um how I think about like how this technology is going to be regulated um so so thank you for for making time for this presentation uh completely agreed with the previous comments about Megan having seen her in action myself um and it's been really fun collaborating with John and Campbell um back when we were working on uh the Wyoming LLC presentation last year as well um so so what I W to go over today um is I want to give uh glimpses of my perspective on what the future of law is going to be Mak an argument that the future of law is actually going to be inextricably tied with the future of software um make some predictions about what that future is going to look like and then do a deep dive into how synthetic data uh and specialized model fine tuning techniques can help for legal use cases um so one thing that I want to argue is that the future of law is equivalent to the future of software and I think this is a point of view that might be very familiar to those of you in the audience um coming at this from the angle of computational law um because after all what is law if not extremely inefficiently executed software um that runs on the hardware substrate of our organizations and institutions um and so that immediately leads us to the point of view that language models are this platform technology that can let us actually implement this software at far larger scale and at much higher levels of assurance which I think ties to a lot of what uh John and Campbell were talking about earlier um so so before we go further I'll just give some more background on myself um so I got my PhD in math last year uh and before that I was a senior research scientist at openai uh where I worked on gp4 um scaling laws the applications of language models to mathematical reasoning uh program synthesis and I was also part of the embeddings team as well um and one of my uh more notable lines of work while I was at openai was that I spearheaded techniques for using synthetic data um in some cases showing that by training on Purely synthetic data you could bootstrap a model that was only hundreds of millions of parameters to the level of proficiency of gpt3 itself um now this requires uh some techniques which are not so prominent these days and which are not uh quite as well known as many other prompting strategies or things that AI Engineers use but at morph Labs uh we've been using these sorts of techniques to achieve some extraordinary results which I will uh tell you guys more about very soon um so one of my perspectives coming out of my time at open Ai and with my background in pure mathematics is that I think that mathematics is really a special case of software um and if you take this point of view and you sort of apply it to the point of view of law you to software that runs on the hardware substrate of Institutions we can think of law as being a special case of software as well and so many of the techniques which have been useful for achieving breakthroughs in mathematical reasoning um a domain where you have to reason very precisely over complex documents um could also be applicable to law and that's an argument that I'd like to explore today um so because of these lines of analogy um so one more consequence of this is that the way that AI Technologies especially around large language models in generative AI the way that they're going to transform uh the production of mathematics the production of software the maintenance of software the maintenance of mathematical corpuses and knowledge that will apply equally well to the process of creating new bodies of law or editing bodies of law or ensuring that uh certain actions by actors are in compliance with bodies of law um so one way to view this and one of my predictions is that we're rapidly approaching a world of ubiquitous intelligent microservices um right and so what this means is that things that we did not normally associate with a sense of agency or personality things that we could not build a relationship with before we can now build relationships with um because uh they'll be agentic they'll be wrapped in some kind of AI actor so Alex Chow over at Microsoft has been doing some very interesting work with his semantic kernel technology and they recently published a position piece uh more or less exploring precisely this so they think that the world is going to be uh fragmented into this universe of these agentic architectures that represent microservices right so rather than everything being bundled into a single chat assistant that can use a million tools at once there are going to be thousands of different assistants which are all coordinating with each other now if you take this point of view and you apply it to say mathematics or software what that tells us is that um the future of natural language interfaces over code bases won't just be a monolithic chat system but rather every part of a code base or every part of a mathematical Corpus uh will have some kind of agentic interface perhaps with its own personality uh perhaps with its own duties and obligations um and similarly um so what if a body of law right didn't have a single agentic interface on top right but what if every regulation instead uh had an agent that was responsible for monitoring your actions and ensuring compliance right so that leads to a very different mode of interaction uh than we have today right so which is um sort of fragmented right like lawyers have different practices and different specializations but not nearly as longtailed as it could be once you fully enable it with this technology um so again going back to this analogy um one other perspective which I've spent a lot of time exploring during both my PhD work and my time at open AI was the application of verification right so software is built with specifications as how uh the programs are supposed to behave when they're executed and um people have found that that you know using co-pilot to generate large amounts of code actually results in more copied code code that's less maintainable right and so uh the verification and the guarantees of the behavior of this code are becoming increasingly Paramount um so similar concerns arise in the world of mathematics right mathematics is sort of like software uh that's almost never implemented right it's only executed in the heads of mathematicians who actually know um how to run that software and there are only a handful of those on Earth at any given time and so all sorts of correctness issues arise when people are trying to verify new additions to mathematical Canon um so so one technique that that we found very useful for um creating state-of-the-art mathematical reasoners is applying formal verification techniques uh to both generate and filter data to make them higher quality right so in that way by applying formal verification you can produce better reasoners um and you can also verify uh new entries to some body of mathematical knowledge or software implication or uh software implementations or a body of law right and so how do we apply verification to law right and so what that would look like is um the state-of-the-art compliance checking can you guarantee that regulations are obeyed um can you ensure that the law is actually carried out right how do you make All That explicit the problem of specification and the problem of check in that um things meet that specification are going to be increasingly Paramount um and so that leads me to my second prediction which is that there will be ubiquitous synthetic data uh for training um these natural language and agentic interfaces on top of software mathematics and law um so so this is actually something which I've spent a lot of time thinking about right so one of the works that I did earlier in my career was um showing that you could generate vast amounts of synthetic data from pre-existing corpuses of um software that's me for checking mathematical reasoning and that this actually solves a serious data scarcity problem when you're trying to train large language models to become very specialized at this task right um because if you can get this knowledge into the parameters of a model uh it's very good at reasoning over it but you just have to um make sure that the data is not so scarce that the scaling laws completely break down um and then uh I worked on applying uh synthetic data techniques to train uh small large language models um my favorite term uh to become state-of-the-art at unsupervised machine translation in some cases boosting a model with hundreds of millions of parameters to beyond the ability of gpt3 um and this uses a technique called back translation which I think is quite underexplored um synthetic data techniques have also um been used recently to achieve state-of-the-art progress in mathematical reasoning for geometry uh so there was a team at Google deepmind uh that built an Olympiad level AI system for geometry where they basically achieve gold medal performance and the way that they did this was by training on Purely synthetic data um that was filtered and also partially generated by these automated reasoning tools right because once you've codified this mathematical knowledge as software you can begin to verify it in a systematic way such that you can create proofs that some reasoning Trace is actually correct and so if you train on those traces you get a much better Reasoner than before um so an interesting thought experiment for you guys to ponder while I go through the rest of these slides is what could such a system look like for Law and why isn't there you know why aren't there hundreds of companies working on this um so um so one question that we asked ourselves right so as we've been thinking about how to apply synthetic data uh to achieve state-of-the-art reasoning performance is how can we take this perspective um and and how can we show that there's like some application to Legal reasoning so so can we improve legal reasoning uh by generative AI systems by using um certain incarnations of these techniques right so um so the LSAT has this analytic reasoning section um I'm sure many of you have uh perhaps not so fond memories of that right and the questions kind of look like this right like they're like these little like logic or combinatorial puzzles um you know they sort of make your head hurt if you like look at too many of them in a row um and they require a lot of like backtracking search and current language models just like really really suck at this right and so what we found is that if we use synthetic data that's generated by language models and automated reasoning tools right in a very similar way to the alpha geometry approach we can improve performance on the AR LSAT data set by up to 30% um and so uh so what this tells us is that like that analogy between mathematics software and law is actually pretty deep right and upon further reflection that's not so surprising because uh like software and Mathematics law is all about precise reasoning over complex documents that all depend on each other um and uh being really good at that is sort of like an AGI complete task and the better that we get at that the better we get at um so at other incarnations of that task like in building complex software or reasoning over mathematical corpuses of data um and just like like in other high Assurance use cases like software and Mathematics right there are all sorts of problems that we have to overcome because we want these models to be reliable right so if you just like fine-tune a language model on like a legal data set like there are like tons of people who have done things like this right like recently um this group at Stanford's uh published this blog post detailing how um how legal mistakes with large language models are pervasive because language models are not good at multi-step reasoning even if you tune them on domain specific data right like if you apply offthe shelf techniques you get things like this right they'll they'll say things that um seem reasonable but upon closer inspection there are subtle errors in reasoning or they break down um so at morph Labs we've actually spent a lot of time thinking about these precise sorts of problems because we're generally interested in how do we get the future of software here faster um and so the way that we've been approaching this is through better multi-step reasoning um so there are multiple data sets out there on this um so one of the most um promising and and one that stands on the strongest foundations is one called music K which uh does multihop questions via single hop question composition right so a multihop reasoning problem is something where you need to answer a question by chaining together reasoning across multiple documents um and where like any one of those docum Ms won't actually suffice for answering your question and that's the kind of thing that um as lawyers you have to do every day right like you have to reference multiple Clauses inside a complex contract which programmers have to do every day when they're building better software which mathematicians have to do over their Corpus of mathematical documents um and so we synthesized um so so we actually filtered a very difficult data set uh from mus right and then we synthesized another data set um to make the questions even more complicated and right and so so here's an example from a subset of that data set that we call mining music K easy um right so so this is something where um so where the reference Corpus comprises dozens of documents and the model has to answer this in a closed book setting so we're testing how well is they able to compose together facts that it's seen in its training Corpus and precisely answer the questions by chaining together all of these properties right and so um on this data set and a subset of the actual music data set that we filtered for difficulty um so we developed this proprietary synthetic Training Method called self- teing and that produces models which are more compliant they hallucinate less they're better at complex reasoning um and on both music a hard and music a easy um we had that self- teing models were much better than fine-tune models and better than the Baseline models um we saw that self-teaching Stacks with retrieval augmented generation uh and that uh besides the use case that we published self teing actually generalizes two multiple domains um so self- teing is already trusted by uh by multiple partners including a media company with over 40 million in funding and a programming language Foundation that's been funded by Simon Foundation and Schmid futures um we're looking for more Partners to uh to go and develop this technology so especially for high assurance and sensitive use cases um and so if you're interested in applying language models perhaps specialize to a complex Corpus of documents that might not be in the pre-training data um I would love to talk um wow so you can find me at that email there uh and yeah happy to take any questions thank you again daza for the invitation wow wow wow okay that's incredible um that's going to take I'm going to have to rewatch this a couple of times I think to absorb everything that was so so thank you for that um huge amount of uh of uh of things to think about and and and connections to make um one thing I I would like to do which I hope is not too much of a busy body but you didn't mention one thing in your introduction that I just get such a kick out of I I would like to encourage you to add it in which is say Jesse didn't you also have some connection to the chat GPT team uh that put that together before it's big long in November a couple of years ago when you were at open AI yeah I departed a few months before the chat PT release but I was on an early version of the chat GPT team okay I'm just saying I think that in your in your list of incredible resume bullets to me that's like that's when everyone heard of and it's a real good one and I want to make sure everyone knows it and you get credit for that um now now moving forward uh uh we have a ton of of uh of questions um here and a lot of interest perspectives um I'm just going to read one um out loud and and get your take on it it's from David Tolen who also is a lecturer um on Law at um UC Berkeley law school um and he has done some really interesting work kind of decomposing the terms and conditions from open Ai and anthropic and Bard and everyone else and and uh really deep um in the law and generative AI um he he poses this one of the early painful lessons of law school legal outcomes are highly dependent on subjective interpretations by judges lawyers Etc um it should be juries you can go litigants uh it should it shouldn't be that way but gradually we accept that it is Imagine computer-driven law free of large free or largely free of that subjectivity um and he said this in response to some of the uh the earlier part of your presentation but could you just speak to to that um that extrapolation by by David and and how does that relate to your work I think the future of law is still going to be very subjective it's just that the subject in that case will be language models and systems that we build around them right like an AI agent will be making subjective calls right so as to whether or not some situ fits a criteria um and going back to a point that John nay made earlier we have to design these systems in such a way that ultimately these judgment calls um come under the supervision of some human but that doesn't uh but that doesn't preclude us using AI agents to help us make those judgment calls or to suggest a default course of action like when making those judgment calls indeed yeah that that was my take too for its worth um is it's not so much it changes it from subjective to objective um that that's in the vibe I get from the previous generation of AI that was like symbolic reasoning and like you know pure logic kind of if then sort of statements um and there's some you know areas of law that that are amenable to that you know where there's like a clear Rule and it's a yes no binary answer like were you going more than 55 miles an hour and we have instruments and so forth and there's some arguments at the edges but it's an application of a rule and it's deterministic it's actually with with this with this generative Ai and this the these the new models that we have um it seems like that they're up to the task of starting to apply the equivalent of legal reasoning um which itself is subjective and then the question becomes what does due process look like what do legal procedure look like what are we optimizing for what are the safeguards and guard rails and everything with within this somewhat you know human um domain of of cognition that is itself subjective uh but is a at least applying you know kind of regularized standard rules so anyway um I was thinking similar things to what you said another thing now we come back to the um the essence of synthetic data which is really the The Anchor Point of your of your talk um I'm going to combine two um two things here one is from uh Sarah Johnson uh and she says is one assumption for synthetic data uh that its quality is superior and more reliable than real world source data so is it is like better is that an assumption uh and then similar to that is uh George Dyer's um question in the Q&A what's your target accuracy and how you how do you establish minimums uh presumably with synthetic data like for clients or for specific applications and use cases I think these are related questions yeah so so the benefit of using automated reasoning tools is that we can get synthetic data to 100% accuracy um and that is a very very desirable state to be in um because then you have complete trust in your training data and you have very high confidence that the models will improve they'll be more compliant they'll hallucinate less I found that using data that is not completely accurate but maybe like 70 or 80% accurate still improves the cap abilities and the robustness of the models so um having some kind of automated reasoning filter is not a prerequisite so as for the second question um so ultimately uh the metric that matters to us is the target metric right how how well does the model do when it has to do some complex or subtle legal reasoning right like over multiple Clauses inside a lengthy contract um and so we measure that uh the the overall quality of the synthetic data is not quite as important like what matters is simply improving the performance on the downstream task got it um helpful so um one final thing you could can I um hijack your screen share sure Don can you see this uh yes okay so your talking agents John talking agents and Campbell I'm talking agents you're drawing on my screen or someone is anyway that's fine um and uh and so something that we're we we're launching actually we just made the page public today um is is a research project on uh agentic AI systems um as open AI calls them uh and I think that's a good name for it and one of the things we're looking at here is what happens when you have individuals or companies who configure um a an llm with some other applications to help them conduct transactions um and you know this is obviously already happening you know like go and find me a bunch of products that be this or that kind of category and you with a extension on the web it made you you know kind of several Hops and find things and synthesize them prioritize them and give you back a nice list and you could go further and further and further and people are starting to explore um how this could be used to supercharge um Commerce so to ahead of that um because obviously there's issues and and challenges that arise when people delegate an amount of authority to um agentic systems um to actually conduct transactions or to when they're holding themselves out to third parties that are interacting with them maybe giving a quote or even closing a deal potentially or doing other things um it raises legal questions and so some this particular research project is going at um something I'd mentioned that you and I have talked about in the past but like the electronic agents and automated transactions and other similar Provisions um um you know error control security procedure and other other relevant aspects of existing bodies of law seeing how much mileage could we get out of kind of using some of those legal Frameworks as part of the design pattern and architecture for agentic systems in the context of of these agents doing uh transactions on behalf of a principal a person or organization where there's a third party involved um in this context um H how what what do you imagine this is a real timely question here it's very practical because we're starting to dive into this what what could be the the opportunities uh and also maybe the cautions for generating synthetic data um to for for to configure and to develop agents that do this but also to to test um agents like in a in a control harness or or like a test harness to see how well they're per performing and if they're going off the rails in some ways so this is actually a question which um we've been exploring a lot recently at morph Labs we've developed the system for automatically generating benchmarks for evaluation uh for code bases and like what we found is that as long as you can have an analytical guarantee right from like first principles reasoning or maybe like like you know static analysis of the code that some question and an answer is correct um then you can blindly optimize against that right but one danger of um having a system of AI judges you know so to speak is that um their judgments may be imperfect right they're bounded by the um uh the capability of the underlying language model in a way uh and so once you begin blindly optimizing against that um then the errors in those judgments will leak into the system they are optimizing and compound uh and so you have to be very careful about that um so in the realm of of software um like when we generate our benchmarks like we use a combination of these techniques right so we use judgments from um an AI senior software engineer as well as um like static analysis of the underlying code base um and we've found some promising avenues for mitigating this like compounding over optimization effect I think um so I think when working with like so like agents in general and thinking about how to make them compliant um and also generating data against these judges right like you can run against like some system of Judges many many times um and you can ostensibly get a data set that you can train on um but that's exactly where those compounding errors show up um I think it's something which is solvable but which is like right there at the boundary of Applied research um hopefully something that we'll make a lot more progress on very soon here here um thank you very much and uh I don't want to put you on the spot too much but uh we don't talk often enough for me not to take this opportunity uh oh Jesse would you like to help us a little bit on This research project I would be delighted to Fantastic then the next time you refresh the page you'll see your name magically appear on the team and thank you for that and thank you really for taking the time again uh to share with us your ideas and and the sort of Look Over the Horizon for many of us in into the future and what's unfolding what's important and and what we can do uh to to beneficially take part in it so thank you for sharing your judgment your wisdom your expertise in this area and to help us start point the way to the next Horizon yeah likewise the perspectives at this Workshop have been very refreshing you're here thank you now um we come to the wrapup of the workshop um thank you to all speakers for for your flash talks um and now uh uh I want to um us to use the last few minutes to introduce and to celebrate our newest editor at the MIT computational law report namely Olga Mack she is royalty in the area of legal Tech um and she's got a truly August U background in the law and in technology and Innovation and uh she's been a great collaborator uh with law. MIT uh.edu over the years now and was was a member of the task force that came up with those seminal guidelines uh for the professional um responsibility used by lawyers of generative AI looking forward she's now going to lead the way on our next big public initiative um which involves a call for submissions who are we calling to we're calling to you um and so with that um Ola thank you very much for agreeing to take the uh the position and the leadership of this and won you please introduce yourself and tell everybody uh you know what we're doing and how they can contribute well hello everyone and wow daza what a fantastic event I am still processing the future of law is uh tied to future of software um Jesse thank you for that and I just love how you solicited a confirmation that Jesse is bound to help right on the spot because he was not able to say no that was just a true art form I I will follow your lead um I love the future of law I I especially on the intersection of generative AI um it is truly incept exting place especially because For the first time in history uh lawyers are excited more like nervous sighted nervous and excited about this technology uh previous Technologies was just mostly nervous uh this one actually includes excitement and and it's clear why it's because it has so much opportunity provide better Services improve our life as lawyers and really enjoy the practice of law uh I think think we can over time really bring fun back and go to practice of law because it's truly exciting um also thank you daza and Professor Megan Ma and Brian for bringing this group of diverse professionals together I think it very much illustrates what we want the future of law to be uh we want it to be yes full of excited lawyers and yes full of other excited professionals because Justice and law is something that we we all as humans have a right to access and and and be part of um and be served most importantly be served so that brings me to this really exciting place which um I hope you all and folks you know and folks in your network join us in building and that is the place where we have Premier destination for repository of information that really encourages well first of all Sparks conversations Foster Innovation and really supports and eliminates P for everyone lawyers and other professionals to be ex get up in the morning and be excited to contribute to the future of walk so with that da a may I recruit you in showing screen so that I can talk and you can show and the two of us can have a show and tell and you're on mute sorry first of all as promised so it has been delivered Jesse Han is now on the project page uh and let's see if we can here we go this you will find all ye who hear it at law. mit.edu oh I'm sorry wrong one um law. mit.edu j- I so this is our call for submission and as I mentioned we have we would like to become a premier destination uh to spark conversation exchange ideas Foster Innovation and really be a supportive Place uh We've listed H da and myself and profor Megan Ma and Brian have tried to give you some ideas of things we're looking for and you can see we are looking for all kinds of things and if you managed to come up with a category that we didn't come up with guess what we have a last category that says many more so we encourage you to be really wide and Broad in in your submission and the kind of expertise you share with us and the other thing I would like to point to is to what kind of things we're looking for and uh to Echo what Brian said yes written works that uh lawyers traditionally sub need are very much welcome please know that we ask them to be two to 5,000 words because it's a lot of work to edit and and publish and frankly encourage people to read more than 5,000 words so unless you really truly have more than 5,000 words to share that are of value in every word consider to stay within the the word limit but the most exciting thing here is that we want to encourage professionals who are not necessarily lawyers to use their tools of trade and submit their uh submissions in whatever form they're comfortable and so to this end we're also inviting folks to submit things like developer notebooks we really want to make sure that technologists are part of this community um as you can tell from Leo's conversation in Allison conversation and numerous other conversations law is increasingly becoming a destination where code is very much tied to law and law is very much tied to code and so some proficiency and works um in developer notebooks are more than welcome but think broader than that yes written works yes developer notebooks but consider videos consider generative art consider other media that we perhaps have not listed uh again we want to be welcoming of all kinds of professional because all of us as humans have very much care about the future of Law and and it's something that should have access to everyone so we there are two we we the plan is to publish two um editions one in the spring summer and one in the full winter and you can see the deadlines for submissions and for publishing uh that we would love for you to to keep in mind so the the first spring summer edition will be published somewhere in September around September 17th of this year and the deadline for submission is April 17th so the form I think there's a form link does if you can show folks where it is because it's a little subtle it's somewhat easy uh it ends self-explanatory um there felds to fill out and works to attach or Point links to um and you may be thinking how can you help how can you be part of it um I have three things that I will ask you to do one submit on time after you read the instructions and follow them two encourage folks that that you see every day in your life that share ideas or do things or work on things that are worth sharing kind of like the things we had presentations today uh things that would encourage wider conversations in the industry and encourage folks to build the future of law so we can all benefit and then I'll will ask you to do a third thing uh you now have a link to the submission page if you can share it on your social media and encourage folks in your network to apply and apply on time and become part of this conversation that would be a fantastic way for us to build the future of law together um with that in mind I look forward to reviewing Edition all the admissions you have in whatever media that you choose to uh submit God help me to have um software on my technology so I can read and open and be part of the conversation um daza back to you thank you so much Ola thank you for stepping up um to to help uh make this possible so people have a surface area that they can't where we can as you said encourage everybody to share from your perspectives about the the Advent of this um new technology and its impact on and sort of implications for the law um transformational is a good word for it and so this is something where it's a time to shed light and to encourage people to to get involved um to share your work and so that we can all get educ ated now so with that um I want to thank everybody uh for uh for speaking I want to thank all of you um especially those of you who stuck with it to the very end uh for your active participation um will'll consider this the beginning I I know that we didn't have an opportunity to get to all of the questions and all the comments um and uh this is our as we customarily do with this Workshop kickoff for the themes and the topics that we'll be addressing at law. mit.edu through the year 2024 we hope that you'll stick with us we hope that you'll continue to participate and to contribute um so until the next time we look forward to seeing you at law. mit.edu [Music]