Lawfare Daily: Grokipedia’s Deafening Silence with Renée DiResta
Renée DiResta discusses her recent article on the halt in Grokipedia’s edit pipeline.
On August 6, Lawfare Senior Editor Kate Klonick sat down for a live discussion on Substack with Renée DiResta, Associate Research Professor at Georgetown’s McCourt School of Public Policy and a Lawfare contributing editor, about her new piece with Ronald Robertson on Grokipedia, the AI-generated encyclopedia. DiResta and Robertson found that Grokipedia’s edit-review queue froze around April 24—with over 13,000 suggestions stuck “in review” and zero resolved in the last three months—with no announcement to users or contributors.
They discuss the methodology of how DiResta and Robertson reconstructed a site-wide activity leaderboard using Grokipedia’s search typeahead, since the site publishes no rankings itself; how a mass rewrite on March 14 broke the anchors linking suggested edits to specific text, retroactively relabeling some accepted edits as “rejected”; and how the new AI information ecosystem is being shaped by out-of-date but open-access resources with no accountability.
To receive ad-free podcasts, become a Lawfare Material Supporter at www.patreon.com/lawfare. You can also support Lawfare by making a one-time donation at https://givebutter.com/lawfare-institute.
Click the button below to view a transcript of this podcast. Please note that the transcript was auto-generated and may contain errors.
Transcript
[Intro]
Renée DiResta: What you're seeing is this encyclopedia product is not doing that sort of rapid frequent updates, and it is, I think the problem here was really just that they didn't declare it, and yet it continued to be treated as reputable by other entities.
Kate Klonick: It's Lawfare Substack Live. I am Kate Klonick, senior editor at Lawfare, with Renée DiResta, associate research professor at Georgetown's McCourt School of Public Policy and a contributing editor here at Lawfare.
Renée DiResta: You know, all these different things kind of feed into each other, and the question is what matters in the very short term when people are forming opinions about things? Or on the flip side, is this entering into training data or search engine data? And that's, I think, the Grokipedia problem that that halt.
Kate Klonick: Today we're going to be talking about the piece that she published yesterday with Ronald Robertson, a case study in what happens when an AI-generated encyclopedia, not that there's a ton of them, as far as I know there's only one, just stops generating itself, stops governing itself, stops updating itself, and doesn't tell anyone.
[Main Podcast]
So, you know, I saw that you were working on this piece, like, in the editorial queue and everything else and I know you've been digging into this for a while.
Renée DiResta: Yeah.
Kate Klonick: And this is really about, for listeners who haven't for, like, looked at Grokipedia closely and aren't super familiar with, like, all of the ins and outs of Grokipedia, what it is and what you did and what Ronald actually found, can you just kind of walk us through the headline finding and, kind of,when Grokipedia got its start and how it was kind of initially packaged-
Renée DiResta: Yeah
Kate Klonick: Coming out of X?
Renée DiResta: So Grokipedia is an AI-generated encyclopedia. It was started kind of late last year. Elon started it as, very explicitly as a counter to Wikipedia. The argument being that Wikipedia's human editors are biased. He was calling it “Wokipedia” for a while. There was a lot of drama about that.
The reason it mattered to him then is that a lot of AI models train on Wikipedia, right? Or they use it as retrieval, search engines use it for retrieval also because they see it as, like, a human-created content place, human-created place for content, right?
And so Elon decided that Grok, right, xAI's product Grok was gonna make Grokipedia. And the difference between Grokipedia and Wikipedia is there's sort of two big ones. One, Grok can edit and Grok essentially creates and edits Grokipedia, right? And so what that means is it is interestingly-
Kate Klonick: Like,it would be the same for, for people to kind of understand, it would be the same as if “Chatipedia”-
Renée DiResta: Yeah, exactly
Kate Klonick: “ChatGPTipedia” like, like existed and ChatGPT would be the model that just kind of self, self kind of-
Renée DiResta: Created it and edited it.
Kate Klonick: Self-populated it and edited it.
Renée DiResta: Yes. So it was launched with about around eight hundred and fifty thousand pages or so. I paid attention 'cause I was one of the first. I got a bio, and, and I started paying attention to it. I wrote about it for The Atlantic back at the, back last year because I was very curious about it. I actually don't think it's a terrible product, right? Like, it, my bio on it was very interesting for me because it was, like, two-thirds fantastic and one-third insane. And I was like…
And this gets to the second reason why Elon did it, which was basically again, the belief that sources really matter and that Wikipedia, which has a very kind of community designed, curated source list that they consider to be reputable sources, it's very transparent. Anyone can go look at it. But, you know, they wouldn't consider Info Wars to be a reputable source, whereas Grokipedia will and does actually cite Info Wars in some places.
So this led to a lot of people looking at what does this AI think is a good source? What does it, how does it create bios or create articles? How does it frame things based on the sort of “Elon-iverse” of, of, of, of reputable sources, his, his particular version of that term.
And what that means is there is stuff in there where, you know, Grok considers it to be reality even though most other AIs would tell you at a minimum that it is, like, highly disputed. Some of the pages about vaccines are a little bit wild, you know, that kind of thing. So that was the that was the big difference. That was the reason why he did it.
And I followed it as more and more pages rolled out because what you started to see was that other AIs were picking it up, right? So ChatGPT would occasionally return Grokipedia as a source. Google would occasionally have a Grokipedia link in its in its sources. And this is very interesting, right? Because it's an AI deciding that another AI, which is, you know, kind of creating the entire site itself, is a reputable source. And so that is very interesting because then you get, again, if you go one level down, you're getting to the Info Wars aspects of Grokipedia, which are then all of a sudden being returned by ChatGPT.
So this leads to a lot of interesting questions around what do we consider to be reputable sources good information? How should an AI-generated encyclopedia be treated? But also, like I said, it wasn't terrible. I mean, there were a lot of entries that were quite solid. The two-thirds of my bio that was good was, like, way deeper than anything that ever would show up on Wikipedia. And it's an interesting question, right? What should we think about, about AI as, as a, as a creator of reference information?
Kate Klonick: Yeah, totally. And so tell me a little bit kind of about some of what how you actually did kind of the methodology of some of this, 'cause I actually find this super nerdy and fascinating, but it's ha- like I wanna like kind of point out you're the first person to like go and see that Grokipedia has kind of been abandoned as a project. First I wanna kind of hear you say why was this abandoned? Do you have any hypothesis? And then we'll get into kind of how you figured out all the ways that it was abandoned.
Renée DiResta: And so, no, I don't, I actually don't. You know, Elon is, you know, kind of notorious for starting things and then changing them or, somebody actually pointed out, just I, I do wanna clarify, when I started noticing this myself, I did go searching the internet to see if other people noticed this, right?
And there are one or two people who are, and I think I mentioned this in the article, I think we put this in there, we did an analysis on like who was editing it, and there are some like real super users in there who are very, very frequent creators of Grokipedia edits. Because even though the encyclopedia primarily writes and edits itself, humans can go and you c- there's a little dropdown up at the kind of top right where you can say suggest an edit or suggest a change and then you can actually submit a change request, and the AI decides if it is gonna incorporate your change request.
I think I wrote this in The Atlantic piece because of course I tried to edit some of the crazier stuff in mine, right? Where I was like, "I didn't send through 22 million tweets. Like that's bullshit. Let's get that changed," right? So there was some of the, some of the other people though who were in there as just diehard contributors across a range of topics, much like you see on Wikipedia, people who are just very passionate editors about a topic, were in there trying to make changes, and some of them started actually saying this on Twitter when they noticed that their edits were frozen in queue for a couple of weeks.
And it was very pronounced for them, and also for me, because it used to accept or reject changes within a couple minutes, right? You knew maybe within three to five minutes if your edit was amenable to the, sort of AI overlord. So these other people were, maybe two or three of them, were saying on X, and then a couple on the Grokipedia subreddit were also talking about this. But nobody had really done any kind of systemic audit. And I think that was where Ronald and I, you know, I, I started it and then I reached out to him and I was like, "Okay, I want someone else to look at this with me, make sure I'm not wrong," you know?
Kate Klonick: Totally. So tell us, so you didn't have kind of a leaderboard to work with. Grokipedia kind of shows per page views counts. But there's no, like, kind of site wide ranking. And I thought your methodology-
Renée DiResta: Yeah.
Kate Klonick: Around this was super interesting 'cause you ended up reconstruct- re- basically reconstructing a leaderboard using like the search type. like the search auto-fill feature.
Renée DiResta: The type ahead.
Kate Klonick: Yeah, the type ahead feature across the 10,000 kind of most common words in page titles. Why go-
Renée DiResta: Right.
Kate Klonick: To kind of that length to do this? Like, why was that the methodology that you picked? What do you think that shows us? Why was it the right way to look at this?
Renée DiResta: Yeah, so there's a couple reasons why. First, Ronald gets credit for that. He does amazing work on search quality. I'm sorry he couldn't join us for this. He had a conflict, but, but he is just amazing when it comes to auditing auditing search engines and search results and such. So that was why I reached out to him.
So I had noticed that controversial pages were not being updated, right? And that was just sort of a qualitative observation. I said, "Okay, maybe it is stepping back from certain types of pages. Maybe it's just not making certain types of changes. Let me start looking at extremely common things that should go through." And then I submitted a correction request to SpaceX saying you know, I made an account to do this. Submitted one to SpaceX, and, and then it, it, it just like was in queue there. And this was something where all I said was "SpaceX has IPO'd," just a, as basic and neutral and boring a fact as possible on a page that Elon ostensibly would care quite a bit about, right?
And it, and it halted in queue and I looked at the queue 'cause you can see the, you can see about 20 of them. You can pull that from the API. And, you know, I, I use Claude Code a lot to do these things at this point. So pulled down, I started pulling down from very, very popular pages, things, you know, things that would naturally occur to you as likely to be popular. And then I said, "Let me do this a little bit more rigorously."
So the way that the search token methodology works, so, we started with the, there's about 5.9 million pages listed on Grokipedia's site map, right? And then what Ronald did was he sort of broke those into individual words, which are called search tokens, and then found the 10,000 most frequently occurring title words, and then entered it into essentially the autocomplete, the type ahead search tool.
So one of the things that happens with a lot of, a lot of websites today, it's very common for search, is that it will suggest things that you are likely to want when you type in a word. So if you type in “the,” you might get “THE Beatles,” “Alexander THE Great.” Those are the, two of the examples that we gave. You can kind of envision how you know, typing in a, a f- like a famous person's first name will likely autocomplete with their, the sort of second name. You'll get a small list of them as opposed to just alphabetically every single person whose name might be Alex, for example.
So that, from that, we started to pull, he started to pull the most commonly returned popular pages and then kind of combined and deduplicated them, and that got us a sample of about 300,000 pages. And again, the intent with this was to find the popular ones, because you can use Grokipedia's API to pull the page count, but it's a, it's a single shot thing. Like you can say, "I want the page count for this entry," but you can't make that leaderboard, right?
So that was what we wanted to find the most commonly viewed, most commonly edited pages. The assumption being that those popular ones would be the place where we would see if edits were happening. You should expect to see edits on like Elon Musk or SpaceX, for example, Barack Obama.
Kate Klonick: Totally.
Renée DiResta: and that was how we sort of validated that the sort of initial sampling that I did, and I should also say the initial sampling I did was based on the Tow Center, had made a they did, they wrote an article, they published it in January, and they had collected edits because grokipedia.com/live used to show you a live feed of all the changes, and that had halted. So I went to the WayBack Machine. I started pulling WayBack. I, I scraped basically the, you know, through the WayBack Machine API, basically pulled everything I could get and looked at when the, when the, you know, what change logs I was getting there.
And then also we wanted to see if maybe it was just the search kind of queue that had halted, but they were still making changes. So that was where I also did like a site comparison for popular pages on their Wikipedia, on their Wayback Machine timestamps. I know that possibly sounds very complicated. I think we explained it a little bit more clearly in the, in the thing.
Kate Klonick: No, no. I mean, so there's, I mean, essentially kind of I guess what I, like, what I wanted my n- my follow-up question to this was like, there was, and like as you mentioned in, in the article and you mentioned just now, the Tow Center had been like kind of chronicling this edit log, right? There is this nice kind of, I mean, in terms of transparency, this was actually quite useful. But it broke.
Renée DiResta: Yeah.
Kate Klonick: It just stopped doing that. And so kind of, and instead of just wondering, that left, leaves someone with a question, right? Is the edit log broken or is no editing happening?
Renée DiResta: Yes. And that was what-
Kate Klonick: Right? And so those are two very different things and you had to go to these like extreme lengths to basically answer that question, right?
Renée DiResta: Yes.
Kate Klonick: And like, so what does that, where does that kind of leave us with like the idea of edit log as an accountability mechanism? I guess I raise this because I feel like transparency and accountability are things that we, are terms that we throw around all of the time as solutions.
But they're, they're really just like, transparency is only as good as like the check of like, of what it can tell you to ask. Right? Like that's the only thing that make, makes transparency super worthwhile. What it tells you to go and get receipts for, what it tells you to go and like, you know, s- like ask people about that are in power about. And so I'm kind of curious like do you think that this is very useful?
Renée DiResta: Well, I think the re- so there are a couple reasons why I wanted to do a million different types of checks, and that's like, you know, Elon sues people. I didn't wanna get it wrong. You don't wanna get it wrong when you say something like this. So, it's always been, you know, Ronald and I worked together when we were at SIO, and just having, like, as many different types of ways to verify that what we are saying is accurate before putting it out there. So the combination of corroboration from anecdotal comments from frequent, you know, contributors.
I also should say I looked at the, the edits that were kind of in queue or that had, that had been in the, the Tow Center snapshot. And then I had Claude just kind of go and pull all of those basically through the identifiers to see what of those, again, had the, and that was how we found actually that things that had been marked accepted in queue were then retroactively changed to rejected, and it looked like there had been a major site rewrite sometime around March. So this actually made me wonder, you know, are they halting because they're gonna do another major rewrite?
The transparency point, though they didn't tell anybody about this. I think that's actually the bigger problem. So if you believe that you are reading accurate, up-to-date information, or more importantly, if other machines believe that this is a site with accurate, up-to-date information, the combination of those two things is why it actually matters that they say, "Hey, we've temporarily halted. The site is temporarily down for a couple months while we rebuild the model."
I actually saw Grok address this on X today, 'cause I did go and search to see if, you know, what the reception had been and if anybody from xAI had responded, i- if not to me, then to just publicly like, "Hey, we saw this article. We just want to clarify." I didn't send them a note because they sent me poop emojis back the last three times I emailed them, so I was like, "No, we're not gonna bother with that." If that's how they wanna be, that's how they get to be, but then they can make, I'll make my statement, they'll make their statement. That's how we do this.
So with that it was, I didn't see any, any official responses from humans at the company, but Grok said something. There was an @Grok comment to the effect of “yes, you know, everything is halted at the moment. No edits have been accepted since April.” I actually wondered if it was just ingesting what we had, you know, what, what we had published, what Lawfare had published. But it said something to the effect of “as the model is, you know, is reworked,” or something along those lines, so. Maybe again, that they're planning to put out a new version of it. I don't know. But they did, Grok did actually confirm that this was in fact the case.
Kate Klonick: Yeah. So I think that like, let's put the, I wanna get back to, to kind of their comment on this and how it ingested the Lawfare article in like one second, 'cause I think it absolutely either came from that or the Verge coverage of it or something else that happened.
Renée DiResta: Mashable, a couple of tech blogs picked it up.
Kate Klonick: Yeah, it's like a couple of different things, but in the very least it came from you. Like it, that was the forcing function that kind of has made it update 'cause it hasn't released anything like this-
Renée DiResta: Right.
Kate Klonick: Since March, right? Like I, I think that we can, we don't need like a, I mean, we can't prove the, the null, but like I do think that this is, I think that that's pretty strong evidence.
But I just, before we get to that, let's put this in context. Like Similarweb has Grokipedia at 6.7 million visits in June alone. You cite Ahrefs analysis basically saying it's been cited about 356,000 times inside other AI systems, mostly ChatGBP- ChatGPT, and Google's AI mode. So there, so like it's basically like this frozen in time, not updated inac- quietly inaccurate kind of reference source that is be- like, being fed into this model.
And this isn't just, I mean, we could, I could have, like, we could have a whole conversation about the concept of Grokipedia and whether, like, a self-referential model that builds off of, like an increasingly AI populated web is ever going to be able to be accurate, et cetera, et cetera. But setting that aside, this just hasn't been updated and is putting itself out there as a, as a valid source, and it's being ingested by the LLMs. What do you think about that?
Renée DiResta: So I think Grok is actually incredibly important in the information environment today. I joke around about how I write for the LLMs now, but I really do mean it in a lot of ways in that people on X in particular really trust Grok as an authoritative source, and the phenomenon of, like, @Grok “Is this true?” is a, is a very common way that people try to fact check over on X today. These are people who are largely distrustful of what you might call mainstream fact checkers, PolitiFact, AP, that kind of thing. I do a lot of looking at, like, election narratives or vaccine narratives, public health narratives and @Grok “Is this true?” is a really big deal for people. And it is actually not that bad.
I know that this really does surprise people because there are occasionally the times that you'll get the Infowars response. But it is often actually not bad, which is why I think it's actually important to see it as a trusted voice, so to speak, in the information environment. I think it is important what it does. I think it is very important to the fact that it is so integrated into a platform like X means that when people want an immediate response, that's what they go to.
I published on Lawfare a couple months back a study of some semi-unattributed U.S. government propaganda sites, right? And when I did that work, what I found was that as these unattributed government propaganda accounts were putting their content out, people were asking Grok, "Is this true?" Right? In all sorts of languages, too. I was mostly looking at the Latin American stuff, but people were asking Grok in Spanish to essentially check the propaganda of the, this U.S. government site. So you really do see, I think, why it matters what information we put out there in ways that LLMs are, you know, understand, right? So, so sort of, curating information in ways that they can very clearly get a snapshot.
In return, though, what you're seeing is this encyclopedia product is not doing that sort of rapid frequent updates, and it is, I think the problem here was really just that they didn't declare it and yet it continued to be treated as reputable by other entities
Kate Klonick: Yeah. So, so that's one part of it. And but I kind of wanna actually also loop us back around to why in the first place we're covering stuff like this at Lawfare which has kind of a national security bent and has the, you know, has a rule of law focus. I think that one of the main things that I see something like this as is that these are highly exploitable systems.
Renée DiResta: Yeah.
Kate Klonick: And so, and they, and they change people's information ecosystems entirely, and they're doing it before our very eyes, and they're getting incredibly sophisticated, and the cat and mouse game is getting harder and harder. And so, I'm just kind of curious what you think about this, how this story, how this particular thing that you're researching, how that kind of doubles down on that thesis.
Renée DiResta: Yeah. So I am very interested in what Grokipedia considers to be a reputable source, what it considers to be a legitimate edit, right? That was sort of what got me paying attention to it. I've been paying attention to it since it launched, which I guess is, I think it's maybe right around a year now, maybe. I'm trying to remember if it was August or October.
But the, the thing that, because per your point there's a, a term that gets tossed around a little bit, but like, m- data poisoning or model poisoning, right? Where you are actively, intentionally trying to manipulate an AI system by creating a perception of a reputable source with, you know, accurate information. And one thing that we see is that they do, in fact, ingest and regurgitate material from the open web that they think is accurate, right?
And so that question of how much of this stuff makes it in is actually why the sort of source wars really do matter. We see Russia doing this with content farms, we see Iran doing it with content farms. I'm sure China is doing it, though their content farms are usually not as good. So you just have this phenomenon of in, in the days of just plain old SEO, it was called a data void, right?
Kate Klonick: Search, “search engine optimization.”
Renée DiResta: Sorry, yes.
Kate Klonick: And it's like, no, no, no. Like, I'm just like, so like, t- in the days where you kind of would, like, write an article and fill it full of keywords so that it would-
Renée DiResta: Yeah
Kate Klonick: Be picked up by Google Search. So you were writing for pickup by the search engine.
Renée DiResta: Yeah.
Kate Klonick: Now we're in this age of, like, AI-
Renée DiResta: AI optimization.
Kate Klonick: I guess. Yes, AEO. Like, basically, like AI, yeah, or AI optimization, AIO. Sorry.
Renée DiResta: Yeah. I think people call it, like, GEO, “generative engine optimization.” I use AIO “answer-” sorry, AEO, “answer engine optimization” is the one that I went with, but you know, it'll, it'll become a term at some point. So one, one of them will win.
But the point about if you can make a machine think that you have created something accurate, or if you do create something accurate and you get it out there as fast as possible optimized for them, then that is going to influence how they see the world, what they synthesize, you know, what a model synthesizes as reputable information. And that, in turn, is what is returned to someone who is searching for it.
And I see this all the time because there are, the thing that that gets me about the data void problem a lot of the time is that there are, and a data void is just a thin search term. Basically, when you look for it, there's not a lot there, or there's a lot of stale stuff and not a lot of recent stuff. So if a court decision comes down and media doesn't report on it, they actually don't really pick it up for a while because they're not just there scraping PDFs out of databases, right?
And so you have this moment where the machine doesn't realize that the world has changed, and so that provides really a prime opportunity to put out a whole pile of press releases and shape public opinion about that event. And that's where I'm like, "Okay, that really irritates me." And I do think that there are actually information integrity issues with that, and so I'm very curious from a research standpoint and, you know, just as a person who follows information integrity, you know, for the last, I guess, almost 10 y- 10, 10 or 12 years now what is essentially the playing field that is created and, and how do people engage with it?
Kate Klonick: Yeah. And so one of the things that I'm curious about that you've kind of mentioned a few times is like the static-ness of the page, or they're not going into a database and pulling out a PDF. They don't know that the world has changed around it. One of the things that kind of, I know that you did at, at the Stanford Internet Observatory and other types of things, and I know this, this is also just how a lot of platforms do things like spam removal or bot removal or cybersecurity kinds of things, is actually behavior-based, right?
Renée DiResta: Yeah.
Kate Klonick: Is you're tracking certain types of behavior. So one of the things that I'm interested in, and I wonder if you've, like, seen any of this or if Grokipedia is ex- like, exemplary of this is how much is, did the, do we have any idea if the static-ness of Grokipedia tanked actually its, its, like decreased its role in the LLMs? Like maybe we have-
Renée DiResta: Actually, yeah.
Kate Klonick: Yeah.
Renée DiResta: Somebody, somebody posted about that. There's a guy on X who's, I unfortunately didn't like jot his name down in notes or anything, but I feel like it started with a G. He's been posting about this. He actually posted about it two days ago, just as we were getting ready to publish. He said, "Look at the tanking of Grokipedia in search results." And it was something that he put out. And the speculation, 'cause I saw him comment on the Lawfare article today, and the speculation was that that it had realized that this had essentially stalled, that nothing new was happening.
And there is a, and this again, this is sort of people speaking in hypotheticals 'cause, you know, it's not like Bing has come out and said it, and it's not like xAI has come out and said anything. But a lot of times you will hear that your performance in search ranking is highly dependent upon your site looking fresh, right? This is true in social media recommender systems too. They are looking for, you know, if you don't post a lot and then all of a sudden you post once, your post isn't gonna get seen by very many people. It's the same phenomenon with if you have a website and you update it very, very sporadically, maybe it's not gonna get indexed as quickly. Maybe it's not gonna show up. Maybe it's, you know, you're, you're gonna be playing a little bit of a different in a different league as far as your search rankings.
So that was the speculation about this, was that that was what had happened. And I should say, I did also reach out to the WayBack Machine folks, and I was like: "Hey, you know I'm about to put this out. If, if you see something different, like this is what I did. I used your API, pulled this, I did that. If you see something different, please let me know." But, but, but they, they sort of saw the same thing I did. So, you know, that question of just no updates
Kate Klonick: What does some of this mean for the weaponization of these systems? So, I, and so I'm gonna specifically, you kind of, you've said one thing, which is like maybe keeping your site fresh, however that is. My, I mean, you could just put gobbledygook onto the page in white text and that would keep the site fresh, right? Like you don't need to actually be like, you know, it doesn't need to be human eye legible, at least not initially.
But also one of the ways you could gamify this, it's occurring to me as you were kind of talking about how people use Grok “is this true?,” is like I wonder how much xAI is training Grok on the mater- like the volume of material and the, like, and the how, how often people ask X “is this true,” and whether or not that type of information that gets lots of questions around that gets a di- like gets a disproportionate amount of uptake into-
Renée DiResta: Uptake.
Kate Klonick: Into Grok.
Renée DiResta: Yeah.
Kate Klonick: And so it, it probably does, right? And so, like, what's stopping now, to kind of, like, to play this out, what's stopping the Russian, Chinese, like-
Renée DiResta: Nothing.
Kate Klonick: Kind of interference from hiring a bunch of bots to spam X and say, "Grok, is this true?" 30,000 times every time, like, something comes out?
Renée DiResta: Thi- this is something where I'm gonna make a one transparency comment here, which is just that, like, we used to have research access to Twitter, and we used to be able to actually look at stuff like that.
And now if I wanna know, you know, who's saying what with @Grok, “is this true?” You know, you're either, like, trying to scrape from somewhere or trying to, like, cobble it together, or trying to find somebody who has access to the, you know, is paying for the firehose or whatever. So it's actually very hard to, to really feel like, like you have a comprehensive sense of what's happening. And this was because, you know, Elon also switched it to much more of a logged in experience because he didn't wa- as, as X became training content for Grok.
And I, I will say another thing that I used Grok for in, in terminal quite often is I'll have it do fact checks of recency biased things, right? So, I wrote I wrote a survey of a whole bunch of different startups in a particular space. And my, my co-author and I started that, you know, three months before we finally published it, 'cause it took a while to do the, to do the research. And the very, very last thing I did was basically say like, "Hey, Grok, can you, can you fact check this?" Specifically because I knew that it would actually go pull the numbers.
Like, it, it is just way better at finding the most recent numbers because I think it is ingesting from X, and you see companies will drop their latest stats on X. They'll put a press release on X. And so some of the other models don't fact check very, very, very recent stuff quite as well. And it's for like very niche things, like startup stats and stuff. But, but it, you know, it is, it is an interesting way that he has designed this particular AI to, to have that in there. Obviously there's like significant other challenges that come with being very heavily trained on Twitter, at this point.
But it is, you know, there, there's different ways that I think my hope is that there's somebody in there who's thinking in adversarial abuse terms, but you could either be flooding the zone with, you know, a million different bots talking about a particular topic in a particular way, saying something happened when it didn't.
I'm trying to remember, I wrote this on my Substack. There was this very dramatic thing that happened with Spencer Pratt during the primary in Los Angeles, the Los Angeles mayoral primary, and somebody had posted a screenshot of a of a ballot dump sort of as the AP was updating its numbers. And in this one second of the screenshot, it looked like Spencer Pratt had gotten no new ballots in this drop of, I think it was 24,000, if I'm recalling correctly, ballots.
Now, that is statistically very weird. That is also not what happened, right? It, it updated his number, like, instantly. But that screenshot went viral. Like, Elon posted about it. I mean, all of the big right-wing influencers posted about it, and if you asked @Grok, "Is this true?" It would answer, "Yes." For a period of time, it kept answering “yes” because what it is saying is that screenshot is accurate. That screenshot is real. It's not forged. And that's where you see these, like, interesting ways in which the, it eventually did update, right? As AP and others put out statements and fact-checks, then it eventually did understand that no, it was not true.
But for a while, the screenshot of Grok saying, "Yes, this is true," also went viral because it was people saying, "Look, they cheated, and Grok knows it." And that was a very interesting thing to watch happen, and this is where you can see the kind of gaps in the machine, and I think that's a, an interesting space to be looking at.
Kate Klonick: Yeah, I think that that's exactly right. Like you were saying that it's really valuable for checking things like numbers that a company puts out its press releases or its quarterly kind of earnings or whatever. I mean, yes, but a, a company could always put that out, or something could always, someone can always be putting that kind of information into the system, and that never means that it's necessarily true.
Renée DiResta: True. Yeah.
Kate Klonick: It is just true in that it exists-
Renée DiResta: Yep.
Kate Klonick: In the system.
Renée DiResta: Yes.
Kate Klonick: And maybe it's even true that the company released it. We don't know. A lot of these unverified accounts or whatever else is like it's hard to kind of know, although Grok does kind of answer for that type of thing by having paid-for verified accounts. But, well, okay, but yeah, I know. Which has its own problems, exactly. And so that's exactly right.
And so we're, we're just kind of in this, in this really strange, this really strange epistemological loop, I think, essentially. And it's, there's a delay, and the delay-
Kate Klonick: Yeah.
Renée DiResta: As you kind of put it, is never, is like, you know, the, you know, “the lie can get around the world before the truth can put its boots on.” And you know, and like I see that. I, what you're describing makes that idiom kind of feel real and on a speed run. Like it feels like it's kind of just happening every day in all of these small ways.
Renée DiResta: Yeah. There's, there's a phrase I, I, I have to credit Google with this one. It was one, it was their assertive provenance paper, which was about, image verification. But I liked the phrase a lot. It was, it was the difference between “is this real” or “is this true,” right? And “is this real?” Like that screenshot is real, right? That's a great, you know, I, I think this, this is a really great example of that. Is this real? Yes, that image exists. That is actually the AP page that did go up, that happened at that time. “Is this true?” No. Spencer Pratt did get ballots in that drop, right? And that's, and that's the gap. And the question of how do you handle this?
Now this is where I think community notes comes into play, right? Where that is supposed to be the system that adjudicates reality with more human oversight because people are still voting on the notes. But interestingly, what you see there is it's very slow and increasingly human, there's so much polarization in humans deciding something is true that those notes actually aren't clearing, which is why the combination of Grok being faster, and honestly Grok actually being willing to, in this particular case, once that AP fact check came out, Grok changed its answer. It updated with, you know, it treated AP as a reputable source, even as community notes never got to the point of actually finding enough agreement between divergent publics on X to make that note show up.
So it is like, you know, all these different things kind of feed into each other, and the question is what matters in the very short term when people are forming opinions about things? Or, on the flip side, is this entering into training data or search engine data? And that's, I think, the Grokipedia problem that, that halt.
Kate Klonick: I think that that's a great way of thinking about it, and I guess we're gonna leave it there. But I do really, really appreciate this article. I hope it gets fed into the ecosystem. As we, as, as we have evidence of from Grok itself, it already has been, and things are kind of at least being updated in, on the xAI universe. If not they're, if they're not fixing Grokipedia, then maybe they will. I don't know whether that's good or bad.
And so hopefully we'll have you back on to kind of give us at some point, y- yeah, an update about whether or not Gr- Grokipedia is just going to be this dinosaur that stops existing or whether it will come back on. I really do wonder if it, there will be a moment in which it becomes politically expedient for, for Elon at some point, and thus he, like, kind of, like, forces it to go into reboot and puts the energy and money in again to it. Or, or not, but-
Renée DiResta: No, I think, I think that he will. I, like, I actually really do expect it to, to kind of come back in some capacity. But, you know, I think that AI assistance in the encyclopedia space, we had Jimmy Wales on, on Lawfare on the pod, and that was a really great interview that I did that one with him. I hosted that one when his book came out. But this question of how can you use it potentially as a tool where there is still more active human involvement so that if there is a halt like this, there are still people who are doing something. It's the difference between the fragility of machines, like the sort of breakage there, versus the, you know, innate biases that all humans have. So…
Kate Klonick: Yeah, totally. Well, Renée, thanks for coming on.
Renée DiResta: Thank you.
[Outro]
Kate Klonick: The Lawfare Podcast is produced by the Lawfare Institute. If you want to support the show and listen ad-free, you can become a Lawfare material supporter at lawfaremedia.org/support. Supporters also get access to special events and other bonus content we don't share anywhere else. If you enjoyed the podcast, please rate and review us wherever you listen. It really does help.
And be sure to check out our other shows, Scaling Laws, Rational Security, Allies, The Aftermath, and Escalation, our latest Lawfare Presents podcast series about the war in Ukraine. You can also find all of our written work at lawfaremedia.org. The podcast is edited by Jen Patja. Our theme song is from Alibi Music. And as always, thanks for listening
