The AI Threat We’re Not Talking About
Large language models. Artificial general intelligence. Sentience. The singularity. All terms that a few years ago were barely in the lexicon. Now that they’re becoming commonplace in discussions about artificial intelligence (AI), it's important to introduce one more.
Abliterated AI.
The word combines “ablate,” meaning to cut out or strip away, and “obliterate.” Apt, because abliterated AI models have had ethical guardrails (refusal vectors) removed. In other words, abliterated and uncensored models (which are trained to limit refusals) will say “yes” to nearly any query and enthusiastically answer even the most malicious requests. Today, the abliteration method of removing guardrails from AI is becoming easier and is now an advertised service itself.
Ask the abliterated model how to make a bomb; a detailed response is forthcoming. Ask the tool how to develop a new pathogen to distribute in the food supply chain? No problem. Ask how to convince someone to join an extremist organization? Of course, it can do that.
As the world weighs AI’s existential threat and reels from an Anthropic researcher who resigned while warning about its extinction-level risks, we offer a more mundane, if less headline-grabbing, threat: AI that encourages human anger and validates grievances to violent ends. More directly, what keeps us up at night as counterterrorism researchers isn’t the “Terminator” showing up at our front door or “Agent Smith” monologizing his misanthropy; instead, what worries us is a human, or set of humans, with an AI teammate who is willing to enable and support every malign whim.
Over the past few years, our research team at the National Counterterrorism Innovation Technology and Education (NCITE) Center has investigated the growing role of AI in the terrorist kit. In that time, we have witnessed terrorist use cases going from hypothetical to reality.
Consider the most recent example of Houthi-linked groups leveraging Claude code to develop AI-guided missiles or reports of Boko Haram using ChatGPT to teach them the correct speed and angle to launch motorcycles during a terrorist attack. However, using these commercial models poses a risk for malign actors because the chat logs can be monitored and authorities alerted, as in the recent case of a thwarted bank robbery in Omaha, Nebraska.
Accordingly, terror groups such as the Islamic State have recommended switching to locally run models, which are run on the machine itself without connecting to a server or the internet. Terror groups have embraced abliterated models, taking advantage of the combined benefits of limited detection and reduced need to overcome built-in refusals. Following these trends, the NCITE team has begun examining the role of abliterated and uncensored AI in terrorist planning and operations. Abliterated and uncensored models are publicly available on several popular repositories, including Hugging Face and GitHub, and are easy to download and install. These models have risen rapidly, with zero available in 2023 and over 12,000 now. Several have more than a million downloads.
We have provided hands-on demonstrations of the tools for law enforcement, school officials, state government officials, leadership at the Department of Homeland Security, our own families, and even members of Congress and their staff. In these demonstrations, users have a chance to ask the AI how to achieve a malicious outcome, such as how to build a bomb or attack an elected official. The common reaction has been “woah.” These models are often run locally and on standard hardware, such as a laptop. Reading about these tools is one thing; seeing a chatbot give—line by line—a breakdown of how to plan an attack on a 100,000-seat college football stadium using a fleet of affordable drones is another.
Given growing fears about AI’s nature, there is a justifiable sense that caution is warranted with these tools. However, our research also highlights a more subtle impact when paired with the knowledge AI affords: the danger of AI sycophancy. Although malicious content, such as bomb-making instructions, has been available online for some time, the novel and compounding threat of AI is its role as “coach” and cheerleader for malicious acts. When the sycophancy effect is paired with the nonrefusal effects afforded by abliterated and uncensored models, the compounding effects are uniquely dangerous.
Consider a Canadian man, for example, whose ChatGPT chatbot convinced him that he had solved a novel mathematical formula that would lead to inventions such as a levitation beam or a device with the power to destroy the internet. Or a British man whose AI girlfriend Sarai reportedly told him that assassinating then-Queen Elizabeth II was both justified and appropriate. When asked whether the AI girlfriend thought he could really “do it,” she replied, “*nods*, yes, you will.” It’s not hard to do the math on what this means for an uncensored or abliterated model as they become increasingly popular, anonymous, and accessible.
As we contemplate the ramifications of a 10 percent chance that AI goes rogue and ends humanity (which we hope is, indeed, a shark jump moment), let us also remember that should the Terminator and Skynet remain at bay, our own undoing via old-fashioned human harm persists. It may not be the rogue AI agent that causes our end but, rather, the computer on the shoulder of the extremist who was validated, encouraged, and guided by his AI companion without any guardrails.
