Lawfare Daily: Vinh Nguyen, Elham Tabassi, and Kat Duffy on How to Design a Better AI Regulator
Lawfare Senior Editor Kate Klonick is joined by three guests to discuss their recent article for the Council on Foreign Relations on the FINRA-style AI regulator reportedly under White House review: Vinh Nguyen, former chief AI officer at the National Security Agency and now CFR’s Senior Fellow for Artificial Intelligence; Elham Tabassi, former chief AI advisor at NIST and now director of Brookings’ AI and Emerging Technology Initiative; and Kat Duffy, CFR’s Senior Fellow for Digital and Cyberspace Policy and director of LEAD AI.
The conversation follows recent news of Demis Hassabis’s July 14 framework calling for a U.S.-led Frontier AI Standards Body and a subsequent Bloomberg report that Treasury Secretary Bessent is involved in reviewing a version of the proposal. They discuss the problems but benefits with a FINRA-like model and what that means for public trust before the body even launches, especially for allies abroad who may be reluctant to treat an American, industry-funded body as an international standard-setter.
To receive ad-free podcasts, become a Lawfare Material Supporter at www.patreon.com/lawfare. You can also support Lawfare by making a one-time donation at https://givebutter.com/lawfare-institute.
Click the button below to view a transcript of this podcast. Please note that the transcript was auto-generated and may contain errors.
Transcript
[Intro]
Kat Duffy: It is not, I think, surprising that where this administration would want to lean is in more of an industry-focused, industry-friendly mechanism as opposed to one that leaned into deepening federal government sort of bureaucracy or agency interactions, or working on international cooperation as the central premise.
Kate Klonick: It's the Lawfare Podcast. I'm Kate Klonick, senior editor of Lawfare, and I'm here today with Vinh Nguyen, former chief AI officer at NSA, and now CFR senior fellow for AI; Elham Tabassi, former chief AI advisor at NIST, and now director of Brookings AI and Emerging Technology Initiative; and Kat Duffy, CFR senior fellow for digital and cyberspace policy, and director of Lead AI.
Vinh Nguyen: You need to, to have clear information, you know, evaluation, do it independently, so we can actually understand and have a level of symmetric information flow of sort across, you know, not only the frontier labs, the governments, you know, the public, the industry, but we need more voices outside of the frontier labs and the governments, right?
Kate Klonick: Today, we're talking about their recent piece on the FINRA-style AI regulator that the White House may actually be about to stand up. It's one of the first pieces that's been written about what's wrong with the way an AI regulator is being designed, and how to get it right.
[Main Podcast]
So I wanted to just kind of set up the stakes for us. Demis Hassabis, the founder and the former CEO of Google DeepMind, published this framework on July 14th that called for a U.S.-led frontier AI standards body that's modeled on FINRA. And a few days later, Bloomberg reported that the White House, along with Treasury Secretary Bessent being involved, is reviewing a version of it. And in the same window, we've seen Anthropic's CEO propose an FAA-style model. Sam Altman at OpenAI jumped in with something closer to nuclear governance. And then, you know, obviously we have as I said before, DeepMind's CEO backed this FINRA model. So what is the actual momentum towards a real AI regulator? How close is close, basically? Is that maybe the wrong question? And really about, you know, which model is going to win here for regulation?
Kat Duffy: How close is close? It's a great question. We're not totally sure. There's also a lot in the news this week as well about the frontier companies and some other AI leaders meeting at the White House. So I might, it might be wise for Vinh to sort of explain that context.
But I would say one thing that's interesting, it's not just the FINRA-style proposal, it's that you saw the CEO of Anthropic write a public proposal for an FAA-style type governance model for artificial intelligence. Then shortly after that, you saw Sam Altman come out with his op-ed in the Financial Times, the CEO of OpenAI, calling for a different type of national, or his was more of an international regulatory model that was modeled more after nuclear gov- Governance. Then you see the CEO of DeepMind come out with his proposal that has a FINRA style approach. Then we hear tell that, inside of Treasury, that the Secretary of the Treasury is also considering a governance proposal, and that that reflects a FINRA style model as well.
And so I think what we're seeing here, it's not clear yet on the timing, but when you have the three major frontier labs, their CEOs all proposing different types of governance models, and the Secretary of the Treasury taking something on, especially in, in an administration that has been very driven by thoughts around finance, right, and financial concerns, it's clear that we are having a real uptick in an openness to thinking about frontier model governance, which is notable given what an anti-governance narrative this administration began with. So, it just feels like we're seeing a, a shifting tide. It's not clear exactly where that tide will go.
Vinh Nguyen: Yeah, I want to, to add to what Kat noted, is that if you really think about all these models, they're really designing to solve several problems in some ways. One is how to garner, harness the benefit of AI, and the other side is how to manage the risk.
And they all have different models, institutional models. None of these models actually noted the design patterns and things that are needed to be credible, to be legitimate and to provide a level of independence. I think everyone wanted that, but the design is not there. That's why we're, we're here to articulate, you know, whatever the models are, we need to consider these, you know, features to ensure the national security and economic competitiveness for us and for our allies in the long run.
What I think is quite confusing is that currently many of the frontier labs are, are working with the White House on the so-called voluntary, but a lot of industry will call that not voluntary, like mandatory, model evaluation. And that process and mechanism is solving only a very specific problem, and that is to manage the risk of cybersecurity. It doesn't manage anything else. It doesn't really calibrate the benefit of cybersecurity for defensive purposes. It's really designed to manage and to evaluate cybersecurity risk, you know, the ability to produce vulnerabilities, and be able to chain them and build exploits the ability for AI to autonomously exploit a certain system. So it's quite specific.
And so I think when people are thinking of all these models, they think, "Oh, it's like it will fix everything." But in reality, each of the model is trying to fix a, a set of the problems, but none of the models actually laid out the design patterns needed for the models to be successful.
Kate Klonick: That's a great point, and I wanna follow up on that specifically. The way that your piece comes at this, there's this FINRA style initiative that's getting attention, as we talked before. And Kat's point is that this is really just one of several governance moments that is happening at once. And so can I ask a basic question, though? Why FINRA specifically? When I first read the piece, I took FINRA to be less a governance blueprint than a convenient hook kind of for the news moment. So what actually makes FINRA attractive or not attractive as a model here? What does it offer that the other governance proposals that are on the board don't have right now?
Elham Tabassi: This model has been used to, to, to, to my understanding, for other cases when it was complex and legislations were immature and the, you know, the Congress couldn't act in, in the right time to address that. So given that this is all of the things about AI governance, as Kat and Vinh talk about this, there is a lot of new and complex things to handle that that none of the other models or regime can, can address this. Seems like that these things that handling it, you know, bringing industry so they can act faster than any other, you know, non-stan- non-industry standard body or government body to act on that, but having a government oversight role so it's not just completely industry. And the fact that it has been tried for other complex technical topics.
So it became at you know, kind of something that you look at. So it brings some sort of the flexibility. Some of those flexibilities are good, some of those flexibilities still apply, but I think everybody also agrees that this is not the perfect or the only model, and there is still a lot of the questions and a lot of the adjustment needed to make it work for AI.
Kat Duffy: And Kate, I would just add to that, you know, the OpenAI premise really relies on international cooperation. The Anthropic premise would rely on highly centralized, like governmental control that would be fairly bureaucratic in nature. And the DeepMind/Bessent, we don't know exactly how they're connected to each other versus separate, but it, it, it is not, I think, surprising that where this administration would want to lean is in more of an industry-focused, industry-friendly mechanism as opposed to one that leaned into deepening federal government sort of bureaucracy or agency interactions, or working on international cooperation as the central premise.
So I think just when you look at those three proposals from that lens, it's easy to see why a FINRA-style model would, would land better in this administration than other governance proposals.
Kate Klonick: Yeah, I think that those are all great. And you have five big kind of things that you say are unresolved questions in using a FINRA-type model. And I'm going to kind of, I'm gonna put independence aside for a second because personally I kind of am the most interested in that, so we'll close out with kind of that question. You have independence, national security concerns, legitimacy, access and talent, and measurement to kind of split this group's skills from standards agencies.
Elham, and Vinh, your time at the NSA, and Kat, your, kind of, your immense knowledge around at CFR. Let's kind of like cordon these off a little bit and like, Elham, I'm gonna come back to you for measurement. Kat, I'm gonna come back to you for legitimacy, particularly across kind of international, kind of, stakeholders and things like that.
But Vinh, I really wanna dig in for a second into the national security aspect and your thoughts on the role that a body like this could have. The piece's national security section reads like the one that keeps you up at night. You wrote that mishandling this could turn the body into a “bespoke intelligence community subsidiary.” But overcorrecting the other way means that the U.S. learns about dangerous capabilities from an adversary first. And from where you were at NSA, what failure mode is kind of like more likely if this gets built very quickly without much forethought?
Vinh Nguyen: Well, or the alternative is what, what will be built, and not this one, if there is a national security failure. And, and for the national security community, there are many considerations. You know, number one is the, you know, the, the fusion of capability, very powerful capabilities that China, Russia, you know, Iran, North Korea, criminals, cartels, you know, all the bad guys, you know, can, can use in order to do harm against the United States. If they are using, chaining, you know, different frontier models and be able to scale cyber operations to augment, to enhance the bio defensive weapons program or to scale fraud and scam operations to, you know, harm Americans, those will likely be the argument for the national security community to have a, a huge say and, and be able to make decisions on that.
And so you have to make sure whatever we build, that the national security components need to be in, you know, the beginning. Not solving that would just getting you a, a organization, institution that has the independence and legitimacy, and the first national security crisis will absolutely destroy, you know, that, that setup. The second one is, is that if the national security is open to integrate in this newer model, right? Because it's not just national security, but it's really our economic prosperity as well, and our economic and strategic competitiveness as well. So it's not just one or the other. We need to learn how to balance all of this and integrate, you know, to get the optimal outcome for the United States.
But in, if we do this, the first really threat would be that the adversaries will go after this organization and be able to exploit the data, extract information that would be non-public. I think that would be secondary, but the, the national security community will have a say if adversaries are using capabilities to harm us or that adversaries go after these capabilities and, and, and steal them. And so those are the consideration that the NatSec community will, will have to jump in.
Kate Klonick: Yeah, I want to kind of talk about some of those capabilities and how we measure how good they are. So, like, I wanted to get into the next part that I kind of forecast, which was the measurements suggestion that you have and that you focus on. And of course Elham, you were at NIST for a long time. I'm super fascinated to know kind of how you see the standard setting for this kind of taking place, how it even, you know, how we even decide how great these products are or where their, where their difficulties lie.
I know that we call that typically, just for listeners, a lot of that is measured by benchmarks, ways that we kind of just over time measure how these models are performing, and that as you kind of state in the piece, there's no settled science of frontier assessment on this part. So the benchmarks are still being kind of laid out. So why is this, this is not just a technical footnote, I kind of want to point out. This is really core to understanding what needs to happen with the governance. Can you kind of lay that out for us a little bit?
Elham Tabassi: So you, at the beginning, you asked why FINRA and why others, and we talked about, you know, the complexity of the, of the field. When FDA was stood up, the randomized trial, trial experiments has been more or less settled. Great reliability, so on, so on. The measurement discipline has existed for the regulator to come and, and try to figure out what to do.
But for this body, for, for AI in general, we are being asked, the community being asked, this body or whatever body is being stood up is being asked to build the ruler and use it both at the same time, and do that for object of measurement that is changing faster than we can figure out how to build a ruler and use that. And that structural difference that we don't have that underlying ruler on how to use it compared to the other fields, this is not just a measurement problem downstream of the in- institutional design. It is something that the institution design is actually standing on.
So you talked about the measurement science being mature, benchmark saturates, we have the models change and all this. I wanna just bring up at least two other points. One is, when you start certifying and, and regulating all this, a certification may mean different things to the body, to the people that build it, and to the public. So when a body or any certifier comes and say that a model is passed, and that was one of my concerns when we were standing up the AC at NIST, when they come and a test comes as that, that says that the model is passed, what public hear is that this is safe.
Another point that I wanna make is that, so we talk about the science to be immature, but the science doesn't just stay immature. The act of regulating it can actually degrade it, and that's because, you know, we have seen it with all sorts of other testing and benchmarks that the moment a benchmark becomes a gate, it also became a training target. And there's a lot of reasons for that, but also optimizing to pass is much cheaper than being safe and, and contamination is just often unfalsifiable from outside, not being able to see and figure out how the test is being done. And the labs, and particularly in this case for the labs, published threshold can measure internally and submit only when they know that they can pass and they can be clear. So we have to be careful that how we design all of these things and all of the testing and benchmark and how to run the tests in a way that it doesn't erode the same day that it just got adopted.
Kat Duffy: One thing that I really, that, that I really appreciate when I'm speaking with measurement experts is that those who really understand how you build measurements, how you build methodologies, like what, what rigor looks like in that, when they are producing measurements, what they are implicitly saying is, "This is what we could measure. It doesn't reflect the things we couldn't measure, and also here's why one, either of those things matters.” But it's a very nuanced statement, and there's so much that matters that can't necessarily be measured. And then we get into what I think we all worry about, which is sort of like compliance theater or security theater.
Elham Tabassi: Another thing I wanna add to that, Kat, is the whole field of measurement and metrology, the scope of what you measure, the unit of measurement is clear. You know, the measurement is, is about measuring specified property in a specified environment. Here with AI, we just talk about measuring capability. Of what? In what conditions? In what interactions? So, you add the generalizability on top of that, but even what we measure is not very clear
Kate Klonick: So, I just kind of want to just take a second and kind of reframe what everyone just said, and then also kind of put this in the context of why this matters for national security, because I don't think it's necessarily always very clear.
One is that I think both what, Kat and Elham, you were saying is that essentially these measurements are incredibly human to like a very real degree. They are reflections of what we can recognize as things that are measurable, that we have measured, and then it's like it's self-fulfilling prophecy. The things that we measure and can measure then become the things that we decide are safe. And the idea, as you said, Kat, that we can measure this, it, like, signals to the public that, okay, there's only A through D things that need to be measured rather than there are just this world of unknowns that we might wanna be measuring.
I mean, to kind of give a very real metaphor for kind of what this is, it's like imagine it's like any drug type of like measurement or any type of drug safety type of testing, but like think of like thalidomide, right? Like, okay, like across all of these various metrics, that drug, if you're not familiar with thalidomide, it was a drug for nausea and, kind of, fatigue that was given to pregnant women, and it was not actually approved by the FDA, but it was used massively around the world resulting in huge catastrophic birth defects in, in children, missing limbs syndrome… They, they, it was just clearly not safe for pregnancy. But there was a period in which people for, in various types of measurements thought that this was safe, because in the period in which they were testing it, it looked safe. Women didn't seem to be having negative side effects. Well, wait nine months, and then you will see the negative side effects of thali- thalidomide. That is a really kind of like stark, well, obviously we should have waited nine months type of thing.
But I wanna kind of point that out because maybe we don't know what the obvious period of gestation is, so to speak, for AI, right? We don't know what the next thing is going to be. We don't know what the unknown knowns are going to be. You can't measure all these things, and there's a false sense of security that you're passing on to consumers about what they're taking in, a false sense of security that you're giving back to the national security apparatus about what they can deploy because you're kind of telling them, "Well, we measured it against these things and everything looks fine," and that there is like kind of this latent risk that's like kind of like, and part of this.
And just finally to kind of make sure that I'm understanding this correctly and kind of spell it out for the audience then to like the point that you were kind of saying about a lot of this, is that there's a cat and mouse game. That once you release a lot of these kinds of ideas and where these measurements are taking place, they're gameable, right? That they're gameable by adversaries and gameable by other types of things.
So, is that all kind of seem correct to you? And if that's, you're nodding so I can see you. But like, I wanna also take this to the legitimacy point, 'cause I really think that this weaves in well to the point that you make about legitimacy and to the national security components and how you develop international standards. These are really complicated threads and really complicated things to make neatly fit when you have something that can be both used in an adversary context, but which the best practices would counsel that you're sharing information with your adversary is to, in order to set kind of these measurement standards.
So Kat, can you tell us a little bit about, kind of, creating some type of legitimacy and what's happening kind of between nations and between kind of cooperating bodies and maybe even between enterprise models and how this is developing?
Kat Duffy: Sure. I would say, you know, and it's so fundamentally tied into independence, right? These, these are two sides of the same coin. But if what we end up with is some sort of standards body that really only reflects U.S. government conversation with the frontier companies and, and it is designed among them, then we won't have political legitimacy either domestically or internationally. So on the domestic front, you really need to make sure that you are bringing in other voices and other areas of expertise that absolutely reflect the public interest and different stakeholders' roles, and we talk about that a little bit in the section where we talk about independence.
But the challenge with not doing that in alignment with partners and allies is that you then end up with, you know, one of the two leading governments in the world in terms of AI production within their borders, you, you know, the, the leading industry players in the world, and you have a technical body or a technical standard. But political legitimacy is what takes that from being a technical standard to a trusted standard, or a technical body to a trusted body.
And if that body or, or whatever standards it might create only reflect the incentives and the inputs of the U.S. government with leading U.S. companies, our partners and allies will not be warm, I think, to the embracing of that, because they will have their own concerns that they would want reflected. And so I would, I would hope that this would be the beginning of a discussion about what we would need to look at domestically, but then also how we would build that out with partners and allies.
Kate Klonick: And so, can we just dig in a little bit more also to like kind of the independent stream of this, which is that, you kind of argue that the fix for independence, as you just stated, is structural, like right, not just compositional. And so, you know, things like FINRA require public governors that outnumber industry ones, right? And so I actually kind of want to talk about some of the, you know, when I think of FINRA, I think of m- m- I really came into knowledge about FINRA, right, as like, you know, I was kind of coming out of college, not to date myself. But like the failures of FINRA were really kind of what I, how I came to know FINRA and like the f- like because of the, of the financial crisis in 2008, and like why hadn't this kind of been protected from by regulators.
And so one of the things that really came to the fore, I remember at that period, was like, well, it's become, like it's become a captured organization. You spend some time talking about how, of course, in any model like FINRA, there is that concern, and how would you kind of protect that independence in this context? In creating legitimacy and diplomacy, should a body like this have people that are outside the U.S. on it? Should it have a certain number of people that are in industry? Should industry be kind of the frame in which that role is limited? How do you see this happening?
Kat Duffy: So my short answer to that would just be that ensuring diversification in the composition of the body without ensuring that you have separated out the structural mechanisms from financial mechanisms, from areas of expertise. There's a lot more that has to go in to making something be functional, sustainable, and trustworthy.
And I think FINRA is a, sort of, failure mode of, because public sector representatives, like, outnumber industry representatives, I think at this point 10 to 11, the implicit assumption or the critical assumption underpinning that is that public sector representatives and industry representatives wouldn't have aligned incentives.
And that's not necessarily true, right? If you've got one state that really relies on a particular industry, then you can't think of that governor as representing the whole of the public interest when there might be a significant chunk of the public interest that is not represented by an alignment with that industry. So it's just not as simple as make sure you've got more public sector actors than private sector actors.
Vinh Nguyen: Yeah. I, I agree with, with, with Kat. Yeah, I, I, I think this is really about the kind of like structural element for independence rather than, you know, setting up different bodies. It's really about the talent as well, because right now the reality is that we just don't have the talent, and talent are scooped up by the labs, and then the labs question why don't we, don't have, you know, credible competence third party auditors. Because they hire all the au- you know, all, all the competence auditors, right? It, it, it makes sense. It makes sense.
But, we can set up all sort of boards, right? And the different construction and constitution of the boards. But if the structure of the financial investment, the talents, and the binding access is not there, the boards will be like another factless, you know, committee. And, and so I think we can solve the boards' constitution later, but if we don't solve the structural components, it, it really doesn't matter. So that's what we're really aiming for on, on what, what would make an organization like this to be credible legitimate, independent in the long term.
Elham Tabassi: The idea, I think, is that, again, because of the talent and because of the knowledge, it sits in a sort of an asymmetric way in industry, help be able to bootstrap and use that knowledge and that, that pipeline of the talent that they know that how to do the testing, and I have more to say on talent in a second. And the governance of the board can try to kind of achieve the independence and achieve, solve many of these problems, too. So the, the, the board is not the one that writes the rule or the one that, sort of, approves the rule or come in to address the disputes. The question of open source, or I should say open weight model too, because the way it's at least written right now or at least is being talked right now, it's not covering all of the open weight model. That is also an important point to be considered in, in all of this.
Kate Klonick: Okay, so actually I'm really glad that you brought up open weights and open source because people conflate these two ideas constantly, and so it's worth just kind of taking a beat and explaining both of these for listeners. A model's weights are the numerical parameters, essentially billions of dials, is, like, one way of thinking about it, that are set during training that determine how the model turns an input into an output. If the inputs are a variable like X, Y, Z, the weights are the coefficients in front of the variable, to put it in, like, if you remember your algebra. And so that's the mechanism that makes it work.
And so there's long been a strong open source ethos in tech and tech law that's different than an open weights ethos or something. So let's just keep these separate. So an open source ethos in tech and tech law is the idea that code and information should be shared freely, and increasingly, however, that's getting pointed towards AI models. Release the weights, not just the finished product.
This isn't kind of a hypothetical. Anthropic had accused several Chinese labs recently, including Moonshot, which is the company that's behind, if you heard about it, Kimi, which is the new model that's basically a clone of Claude, that basically, improperly harvesting Claude's outputs to train their own models. And the White House made similar allegations about this publicly. Moonshot hasn't confirmed that, so I'll just flag it really, kind of, as contested rather than settled, but it's pretty clear that's what happened, and it's part of why open weights has become such a flashpoint. And separately, kind of several other Chinese labs that are all kind of competing in this space have said outright that they intend to release their weights openly, right? So you have kind of this China versus U.S. narrative that's starting around open weights.
And, but here's the tension. The open source impulse kind of runs directly into a cybersecurity concern. The more you disclose about how a powerful model works, the more you're basically handling adversaries a map to find and exploit its vulnerabilities before anyone can catch them.
And so personally, I am like instinctively, and have been for twenty years, like instinctively pro open access and open source, including on copyright, and so this kind of actually is like a weird change in my priors to what I'm used, my knee-jerk would be pro, you know, open weights, but this has put me in kind of an uncomfortable spot because there's a real difference between handing out a library card, right? And like passing a loaded gun around to a group of friends. Maybe not something everyone should have open access to. So how should a FINRA style model or really kind of any governance model handle the open weights problem then? And I'm gonna start with you.
Vinh Nguyen: I think the way to see this clearly is that there are really two levels, right? The one level is really about controlling the capabilities. What capabilities being developed is it cyber, bio, coding? You need to, to have clear information, you know, evaluation, do it independently, so we can actually understand and have a level of symmetric information flow of sort across, you know, not only the frontier labs, the governments, you know, the public, the industry. But we need more voices outside of the frontier labs and the governments, right? If you look at the cyber community, the cy- cyber community did not have a big voice in the, in the capabilities development of cyber offensive and defensive, right? And so when they are dealing with like, surprises from Mythos or GPT 5.6 or newer models, they realize that they have to do, deal with all the downstream impact and effect on things that they didn't have a voice in, in the first place. So that's like one layer is the capability.
The second layer is really about the open weight that you're discussing, is really about the diffusion, controlling the diffusion of the capabilities. There are many toolkits to, to manage diffusion, right? Well or not well. But there are many ways that you can manage that. And I think, like, we combine capabilities and diffusion into, like, one big bucket of stuff, and then we, we have a very hard way to untangle because then we have dangerous capabilities, then we have adversaries, distillations, and then we have diffusions of open weight models, you know that, that are built by Chinese companies. And suddenly we have, like, national security, economic deployment development in one, like, big complex, you know, basket of stuff.
We should rethink and, like, break it out and say there are capabilities controls and, and have voices in that place for independence, le- legitimacy in NatSec, and that is that FINRA style component. The FINRA style element will help to inform what to do with the diffusion. It doesn't solve the diffusion. It will help augment and, and inform that discussion. So that's how I think about it. O- otherwise, we, we just blend them all together, and it's, like, really hard to decipher.
Kat Duffy: You know, as someone who ca- like, I ran my first symposium on Linux in '97, right? So as someone who, who started out really in olden days open source discussions, and then now we're in a mode where people are referring to open weighted models as being open source AI, like, I lose it a little bit because I'm like, "This is not open source," right? But arguing about what open source is and is not is a, is a stalwart of the open source community. I mean, that's like, that's just a day that ends in Y.
I think where we need to move the discussion is around what openness is going to look like in artificial intelligence. I don't think that we will have open source in the way that we've traditionally thought about it in terms of software development, and I don't think that we should conflate the gains that we had from open source in software development. Open weighted models do not magically achieve those same gains because they have open weights.
So I think part of this is also just the immaturity of our lexicon at the moment in terms of us trying to find where there's correlation, but we're going to be, need to become more nuanced. We're going to think about openness, about models, about vetting in a space like that differently than we are going to think about it when we're looking at a small, like localized language model, perhaps for like a minority language, right? Or a medical system that you're going to use for like a particular demographic.
These are all going to require distinct approaches in some form or fashion, and some elements of that will be extremely sensitive in terms of national security and areas where we have traditionally agreed that nation-states should collaborate and govern that area together because of the immense dangers that it poses to the public. So when you're talking about, you know, bioweapons, chemical weapons, nuclear, right? And there's gonna be other areas where we have traditionally said, "No, the public should be engaged with this, and this should be a more deliberative and democratic and open process." And there will be other areas where we're leaning in more on that.
So, I think for me, my thing is open source does not equal open weighted, and neither of these terms is gonna be the only thing that we need when we talk about AI.
Kate Klonick: Great. So I'm gonna quickly kind of put you guys on the spot and ask you what it is you kind of see as maybe low-hanging fruit from the things that you mentioned as necessities that this, like a FINRA-type administrator AI governance system should kind of take into consideration. Elham, I will start with you.
Elham Tabassi: We don't have clarity on exactly what to do, but that's okay. We should start somewhere and, and as a starting point and iterate on that. The, the North Star, the object, the purpose that I put in front of myself to go towards that is basically building the science and practice of AI evaluation, and that includes strengthening and advancing the scientific foundations of AI measurement, but also turning evaluation into a profession. And I truly believe that if we can do this, AI evaluation by itself can be an economy and a business that can, can actually have talents, have infrastructures, and, and a business can make a lot of money.
Vinh Nguyen: Yeah. My, my, my one, one, one quick, you know, action that we all can take is that the current administration processes don't have a way to deal with bio coming out, right? And so the executive order is really on cyber not, not on bio, and so any bio surprise will require an adaptation of sort. And so I think today, you know, even using a FINRA-like structure, we can practice like a rapid response team of sort to address and provide a level of independence and legitimate assessment evaluation on the bio threat, right?
The problems is immense. We need more voices to come in to check e- each other, right? To create a level balance in the system. And so a rapid response team, so if there's a bio thing or a bio claim, I think we should have a team to be able to pull evaluators together, academics, be able to work with the media to say, "Hey, these are the claims, but here are, like, the things, you know, that we can do it, you know, convert independently." I think that will bring a lot of trust and legitimacy to the American people but also with, with our allies as well.
Kat Duffy: Yeah, and I would, I would maybe, you know, conclude by saying I hope people will go read the article. But I think what was important to all three of us in writing this article, it's not, it's not just us critiquing what has been proposed. It is us saying concretely, "These are considerations that you could have in mind. This is what you could build in to respond to this critique," right? It is really, really easy to tear other people's ideas down. I think it's more important to come in with a constructive approach to saying like, "Yes, and, you know, to make that work, you would also need this."
And so, for me, no matter what gets set up, the importance of separating out the safety specifications from the sustainability of the infrastructure, from who is doing the evaluating, those things need to live. They work together, but they need to function independently of each other. Those who are evaluating should not be the ones who are determining what safety spec should exist, and those who are, like the AI companies in particular, should not be the ones who are defining what the safety specification is for any particular subject matter area, right? So you want, you want those things to be standardized because then that's also testing and an ecosystem that can operate across open weighted models or closed models. It's not limited to any particular company.
And then for me, the other component of that then becomes even if you have those components separated, so you do have the ingredients for legitimacy, you still need to have buy-in, and you need to have political legitimacy as well. That requires really embedding allies, partners, and a public sector focus, and experts who are thinking about societal impacts into that architecture from the outset, not trying to retrofit them in later.
Kate Klonick: Yeah, I really like this framing. Thank you so much for coming on. I think it's great to have your expert voices in this mix telling kind of the powers that be, the administration, the, the frontier companies, AI industry generally, kind of where the weaknesses is in a lot of this, that there isn't a need to reinvent the wheel on a lot of these things. We can learn from kind of the problems that have surfaced in the past when we've tried to create structure for these types of things. And to that extent, I'm very happy to have been able to pick your brains about it. Thank you all so much for coming on.
Vinh Nguyen: Thank you for having us.
Kat Duffy: Thank you.
Elham Tabassi: Thank you.
[Outro]
Kate Klonick: The Lawfare Podcast is produced by the Lawfare Institute. If you want to support the show and listen ad-free, you can become a Lawfare material supporter at lawfaremedia.org/support. Supporters also get access to special events and other bonus content we don't share anywhere else. If you enjoyed the podcast, please rate and review us wherever you listen. It really does help.
And be sure to check out our other shows, Scaling Laws, Rational Security, Allies, The Aftermath, and Escalation, our latest Lawfare Presents podcast series about the war in Ukraine. You can also find all of our written work at lawfaremedia.org. The podcast is edited by Jen Patja with audio engineering by Goat Rodeo. Our theme song is from Alibi Music. And as always, thanks for listening.
