Criminal Justice & the Rule of Law Cybersecurity & Tech

A Case Study in Pivoting: Investigating Hacked and Leaked Data on a Budget

Ceren Fitoz
Wednesday, October 7, 2026, 1:00 PM
AI may help, but nonprofits still face an uphill battle in allocating resources for extensive investigations into human rights abuses.
"AI Code" (WCN 24/7, http://tinyurl.com/mr48s6p9; CC BY-NC-ND 2.0 DEED, https://creativecommons.org/licenses/by-nc-nd/2.0/)

In September 2022, a hacktivist group known as Guacamaya, which targets “extractive industries, military, and law enforcement” throughout Latin America, breached the email servers of Mexico’s Secretaría de la Defensa Nacional (SEDENA). The hacktivists released approximately six terabytes of internal Mexican military email communications and attachments, to Distributed Denial of Secrets (DDoSecrets), a nonprofit that publishes leaked documents in the public interest. The dataset offered an unprecedented window into the military’s use of spyware and surveillance technology.

I was part of a small team at the University of California, Berkeley’s Human Rights Center that analyzed this data for an investigation into the targeting of journalists and human rights defenders in Mexico. The military targeting dissidents with spyware was one element—alongside digital harassment, abuse of legal process, physical violence, and impunity—that created an environment of persecution for the country’s civil watchdogs.

At the Human Rights Center, our team investigates serious human rights violations, focusing on international crimes committed through digital technologies. The use of spyware and surveillance technology to target Mexican civil society has become increasingly widespread and systematic, and its abuse is often linked to physical violence. In 2025, Mexico ranked as the second deadliest place in the world for journalists, after Gaza.

Civilians globally are being surveilled at an astonishing scale. From mass surveillance using traffic cameras in the United States to the targeted deployment of highly sophisticated zero-click spyware such as Pegasus, governments and law enforcement can access people’s lives with increasingly less friction. This technology is overwhelmingly used to target dissent. It’s ubiquitous, silent, dangerous, and notoriously difficult to detect, let alone investigate.

Our team’s obstacle was turning six terabytes of raw, unstructured email data into evidence. There was no established methodology to examine such a large data pool in our field of human rights research. We had to build one.

In reality, processing and analyzing six terabytes of data is not particularly difficult in modern data science. However, the field of human rights is largely hindered by lack of technical capacity and infrastructure required for taking on data-reliant investigative problems. Even at a research institution such as Berkeley, the Human Rights Center is entirely self-funded and not financially supported by the law school where it sits, despite its academic affiliation. This limits the center’s ability to fund experimental data exploration. 

What follows is a walkthrough of how a small, resource-constrained investigative team designed a system to interrogate this sprawling dataset. This project failed many times before it worked. This article explains the challenges we faced, the decisions we made, and where we went wrong, in the hope that the lessons help other investigators confronting similar problems.

Designing a Strategy

The Mexican military’s emails offered a view of how the state surveils dissent, in their own words. We uncovered procurement contracts, approvals, and the network that acquired, deployed, and concealed spyware and related tools against journalists and human rights defenders without judicial authorization.

To reach this level of evidence, we needed to understand what relevant information might exist within the email server of a country’s military. SEDENA’s hacked email server was designated for nonclassified communications, while a separate system hosted the military’s truly secretive information. However, just because the server was not intended for confidential data does not mean everyone adhered to such a policy. Our investigation exploited these human lapses in judgment, which led to discoveries, such as an officer emailing himself slides to this less-secure server in order to print them before a meeting. 

We needed the ability to search the entire dataset, but also a way to discover new people, emails, or other data we didn’t yet know existed. We knew to search for keywords, such as “inteligencia” or “Pegasus,” but we needed unknown terms that might reveal contracts or hidden spyware vendors beyond the obvious. Who were the officers involved? And what networks did they belong to?

To solve for these unknowns, our team brainstormed various ways to search the email server beyond a simple keyword match. We decided to search through the prism of graph analytics, which applies an algorithm to search data structured in relation to each other as a network of entities. Also known as network analysis, graph analytics is an excellent way to discover new data clusters beyond keyword searches.

These are rather simple concepts in data science. A graph database, also known as a mathematical graph, can store data. Vertices, which are the actors in a graph (for example, email addresses) and edges, which are the connections between vertices (such as emails), compose a graph. In other words, a graph database stores information by relationships, rather than a traditional database’s rigidly defined table structure.

In this case, we could build a network of relevant military personnel, represented by their respective email addresses. By developing the network clusters of personnel in the network and their positions in the military hierarchy, we searched for those most likely to discuss or be involved in the acquisition, approval, and deployment of spyware and other surveillance technologies.

However, we couldn’t decide which algorithm would surface the most relevant people (via email addresses) without testing our options. Testing these algorithms on all six terabytes of hacked data was infeasible; we needed a more manageable scale.

Fortunately, we previously reviewed a smaller, related set of leaked emails known as the Hacking Team archive during our initial scoping for this investigation. Originating from the 2015 breach of an Italian surveillance company coincidentally named the Hacking Team, this dataset included a corporate email archive published on WikiLeaks.

This dataset was a good test of multiple algorithms for two reasons: The archive was only 400 gigabytes, and, as a company’s emails, it was structurally similar to a hierarchical institution such as a military. Most critically, the Hacking Team emails—at least those pertaining to the Hacking Team’s business in Mexico—our team had already reviewed. This meant we could cross-reference whatever the algorithm surfaced with a fact-checked answer key.

The data included both the email addresses and the email content. Our test needed to read the content in a streamlined search engine and to understand the broader network of emails through the addresses.

To read the email contents, we created a searchable database of the Hacking Team data, similar to the one available through WikiLeaks. We used MySQL, a traditional open-source relational database management system, which made the email contents readable. 

To map out the email network, we needed to show the relationships reflected in the emails, but a traditional database was too rigid. Our team chose a graph database built with the open-source platform Neo4j. Thıs could represent a social network where connections between data points represent social relationships, such as an employee emailing their superior or a sales associate emailing a potential client.

Email Networks 

In addition to the details we uncovered in the emails, we mapped not only the military’s traditional chain of command but also its covert intelligence apparatus by tracing the network of email senders. We did this by selecting the most efficient algorithms through a series of tests.

Using the Neo4j graph database, we selected three algorithms to explore the dataset. With the Hacking Team emails, we wanted to learn more about the spyware technology that the group sold to government clients before the 2015 hack. Our data points were email addresses, names registered to accounts, and the individual emails sent. Each data point is referred to as a “node” in the definitions below:

  • Eigencentrality (Prestige Rank): A way of ranking the “influence” of a node in a network. It assigns scores to nodes that connect to other high-scoring nodes as more valuable than connections to low-scoring nodes. This algorithm closely relates to what Google uses to rank search engine page results. 
  • Label Propagation Algorithm (LPA): A type of community detection that finds clusters of densely connected nodes without defining the boundaries in advance. This method finds clusters relatively quickly but sometimes inaccurately.
  • Louvain Method: A type of community detection method similar to LPA but more stable and sophisticated. It finds data clusters by moving each node into a neighboring group and evaluating it by how densely connected it is within the group. It keeps the clusters with the densest connections between them. 

Who’s Talking About Galileo?

Eigencentrality case study: From previous research, our team knew that the Hacking Team sold a Remote Control System (RCS) spyware named Galileo. When we searched the keyword “Galileo” and applied the eigencentrality algorithm, it ranked the results of nodes, or email senders, with “Galileo” by influence. The algorithm calculated influence by assigning relative scores to each email sender, where connections (emails sent) to high-scoring email senders are more valuable than equal connections to low-scoring email senders.

Top Search Results:

  1. The top node for emails referencing Galileo was Marco Bettini’s email address, who was a sales representative for the Hacking Team. It makes sense that he would have sent the most emails about Galileo, both in internal discussions with the product team and in external pitches to potential clients.
  1. The second email address was Phillipe Vinci, the vice president of business development. As head of sales, Vinci was Bettini’s boss. Here, the results are nuanced. It was not Vinci, the head of the sales department, with the most “influence,” but Bettini, the employee below him, who sent the most pitches selling Galileo.

These results aligned with our knowledge of the Hacking Team’s employee hierarchy and our understanding of who would have discussed Galileo. With this success, we tested the other algorithms using similar logic—would the algorithm provide a similar answer to our human-reviewed research?

Applying this logic, we selected the algorithms that demonstrated useful network analysis. We were now ready to start building out the search system.

The First Ingestion Attempt

After experimenting with the test data, we were ready to start our real objective: examining the Mexican military’s emails. Our team successfully requested access to the Guacamaya archive from DDoSecrets. Ingesting and processing these six terabytes of hacked data—amounting to several million files—was a precarious and arduous task.

The email data was organized in a standard internet message (MIME) format. Ingestion of this data was twofold. First, we used a standardized extract, transform, load (ETL) process to parse and load the email files into the two complementary databases: MySQL and Neo4j. Second, we separated email attachments from their parent emails for document and media processing. Docker, an open platform for developing and running applications, packaged these two ingestion pipelines so that it ran smoothly across multiple machines, ensuring portability and reproducibility. 

We conducted this on a fragile system our team affectionately called COLDSNAP. Our system consisted of a Framework computer, a set of external storage devices, and an ingestion software developed by our technical team expert. It was a skeletal system attempting to handle large swathes of data. In theory, it would work.

But ingestion relied on a delicate system functioning perfectly.

Knowing we risked serious setbacks with a single malfunction, we took some practical—and simple—preventive steps. When the computer started to get hot and whine, we put it on top of a roll of masking tape and pointed a fan toward it for air flow. We moved the cords away from trip-like passageways. And after hearing about the building’s previous bouts of power shutoffs from its own old wiring, we put up a sign above the office microwave and coffee machine begging coworkers not to use them at the same time, for fear of losing critical data.

At a certain point, we got spiritual. It became impossible to leave the office for the day without some outgoing acknowledgment to COLDSNAP, wishing it a safe and pleasant evening. We wouldn’t speak poorly of it in the same room. We didn’t even change the original test password for a second one once we logged in, for fear of an accident.

None of it prevented the first disaster.

An external hard drive fried, grinding the entire project to a halt.

It seemed silly that after all our planning, a simple hard drive breaking was the linchpin for the entire project. We could have bought another, but with a tight budget that prevented us from purchasing a more powerful drive, it was difficult to justify replacing an already fragile system that lacked data integrity. We couldn’t ensure the same or similar issue wouldn’t arise again. Together, we brainstormed alternatives, including using a cloud system such as Amazon Web Services or a locally hosted infrastructure, but both were out of our price range.

We were also running out of time. Inefficient hardware weighed down the processing speed, and with deadlines approaching, our window to actually investigate dwindled. To keep the investigation alive, we had to admit we couldn’t do it all on COLDSNAP. Further, because of our slow, fragile system’s nature, we also had to let go of the hope of viewing the entire six terabytes of data. 

Once it became clear that we couldn’t ingest and review the email attachments with our constraints, we sought external help. We needed basic functionalities that worked well, including the ability to view attachments of any file type—PDF, video, image, spreadsheet, or anything else—on a user-friendly, shareable platform. For text that wasn’t already machine-readable, such as a photo or a document scan, we needed optical character recognition (OCR) to digitize the text.

We found the right tool with DEVSEC, a startup that met our needs and accepted payment in a currency we had: user feedback. DEVSEC offered Weaver, an investigation platform for data ingestion and enrichment that accommodated our full range of attachments, performed OCR, and translated text at a high volume, all while allowing for fluid searchability and frontier artificial intelligence (AI) models with zero data retention policies. Now we had COLDSNAP separate attachments and Weaver to store and process the attachments. 

What’s in a Military?

Because we could no longer review the data in its entirety, we needed a strategy to identify leads within our network and attachment analysis. We used known factors of Mexican military operations, institutional structure, and behavior to pivot toward unknown leads. This is how we expanded our investigation to include intrusive tools beyond Pegasus, as well as personnel who operated together outside existing structures.

Utilizing a selection of files made public by DDoSecrets, we reconstructed SEDENA’s organizational structure to identify relevant email accounts, communities, and documents.

For example, we would search for the name of a known intelligence officer to find the intelligence product he sent over email, and the network of other military personnel he would communicate with. Take a high-ranking military officer: Gen. Conrado Bruno Pérez Esparza (Bruno). In 2020, he was deputy chief of intelligence for the Joint Chiefs of Staff of the National Defense (EMCDN) in SEDENA. This meant he was superior to the Center of Military Intelligence (CMI) when it allegedly used Pegasus spyware.

Considering his role, how would Bruno show up in the data? Which communities would he be part of? On what topics would we expect him to rank high on the eigencentrality algorithm?

Email Structure

Like any institution, SEDENA followed a naming convention for the creation of its email addresses. We noted the pairings of email addresses to names, often with abbreviated military ranks. Thus, we found a pattern of email naming allowing us to target the dataset using either known email addresses or estimated ones based on individuals’ names.

One of the email conventions was based on the initials of the individual’s real name. This pattern relied on the use of initials for first, second, full last name, and initial of second last name. For example:

hosorion@sedena.gob.mxCap. 1/o. Inf. Horacio Osorio Nieto

Document Structure

Signatures

The documents attached to emails contain a pattern that helped link people to specific intelligence products within SEDENA. Nearly all text-based documents, such as intelligence reports and presentation slides, have a string of letters stamped on the bottom left of the document’s last page. SEDENA’s own standard operating procedures verified this internal documentation practice. We reviewed and confirmed that the acronym signatures referred to the personnel responsible for preparing and approving the document in descending order of rank. We deduced the specific people who approved and produced these attachments by mapping which acronyms belong to which military officers.

Here is an example of one document, which is a communication letter from Julio Cesar Diaz Martinez, a captain in the Office of Surveillance. The three-string signature at the very bottom is “JCDM-gsb-hjgs,” with JCDM as the initials of the full name Julio Cesar Diaz Martinez. The other two initials are lowercase, meaning it was prepared by enlisted personnel rather than a ranking officer. 

 

File Path

Besides the documents themselves, a pattern exists in how the files that were sent as email attachments were organized. Most text-based documents, such as PDFs and Word documents, have a file path listed at the bottom of the final page, usually after the signatures. These file paths are a unique string of characters that identify the file location within an organized file system, meaning the file paths offer insight into the location of additional files. 

Most file paths start with a letter, which indicates the name of the drive. Although it varies by operating system, there is a general convention for naming drives. The SEDENA server has a drive titled “E:SECURE,” where we would expect to find sensitive information. There is a “C” drive organized by different directorates, such as the Cyberspace Operations Center, abbreviated as COC:

  • C:USERS\
    • C:\USERS\ADIESTRAMIENTO\
    • C:\Users\COC-23\
    • C:\Users\ticadmin\

Below is an example of a file path that illuminates the organization of regular briefing updates from 2022 on the day’s gathered intelligence. This is a briefing card of intelligence updates from April 1 to April 2, 2022:

W:\acopio\2022\1.-MESA DE PANORAMA\TARJETAS\4.- ABRIL\01\RESUMEN S-2\10.-RESUMEN DE NOVEDADES DE LAS 1400, 01 ABR. A LAS 0500, 02 ABR. 2022.docx

SEDENA’s patterns in email addresses, document signatures, and file paths are just some of the many configurations that helped us clarify our search strategy and make this dataset more manageable.

Streamlined Work Flow

With a redesigned approach, we reached a point where we could make targeted searches—albeit clunky and slow—on COLDSNAP, process documents on Weaver next, and then dig deep into the investigation. But this workflow was still far too inefficient. It was difficult for the team to access data in the MySQL database without knowing how to write queries in the standardized language used to manage and retrieve data from relational databases. And there were no internal developers or funding to create a user-friendly front-end interface for the database.

This project, and our attempts to wrangle this large dataset, occurred throughout 2025 and 2026, while the proliferation of publicly available AI models began transforming data science solutions. AI coding agents, particularly Anthropic’s Claude Code, have enabled the average person to write code and develop software at a level previously exclusive to software engineers.

One member of our team harnessed this capacity by using Claude Code to build an internal tool called Forensic Explorer. This implemented a search engine and front end on top of the ingested data, solving our sluggish workflow. Forensic Explorer was designed to run locally and in isolation, contained in a virtual sandbox environment with Docker so that the sensitive data remained isolated.

Most critically, the tool was not designed with the leaked data itself, and therefore did not expose the Guacamaya hack to the underlying AI model. Instead, we used synthetic sample data generated for the purpose of testing the tool. The development process was iterative; the team’s feedback fixed bugs and shaped the features. Eventually, the email data was ingested into Forensic Explorer, and the investigation finally reached a flow state. 

To organize our findings, we used an information management tool, Obsidian, that helped draw connections between the entities we discovered. For our investigation, Obsidian added value by organizing files in relation to each other, rather than in a traditional one-dimensional file system. Obsidian became a knowledge base where relationships between files reflect relationships among individuals, government agencies, military apparatus, private companies in Mexico, and other entities relevant to our investigation.

Lessons Learned 

Actually investigating this dataset was the reward for our uphill battle. No amount of optimization or tooling can replace the manual effort of putting the pieces together. And despite our efforts, no smoking gun appeared in the data. We tediously mapped the networks of interest and numerous intelligence documents that painted a snapshot of the Mexican intelligence apparatus—tentacle-like, shadowy, and operating with powerful technology.

Through this arduous process, we learned the wide distance between theory and practice. As the tribulations described here show, there were plenty of steps that should have worked in theory. Actually running the data ingestion steps, after long discussions about which data analytics would run best, was anything but smooth. And even after that first painful, time-consuming hump, more constraints emerged. A fried hard drive was not a particularly formidable monster, but it was the roadblock that became the turning point for our process.

The takeaway here for any human rights research group or investigative team is to take the theory and put it into practice as much as possible; to try to handle complicated data problems, even when the team is learning as they go, as we did. 

Our work is driven to expose the use of digital technologies in grave human rights violations. As bad actors harness sophisticated tools such as zero-click spyware or mass surveillance, we need human rights research to match the pace and creativity of these technology-enabled crimes. This investigation shows what we accomplished with little infrastructure; imagine what is possible with the right investment.


Ceren Fitoz is a digital investigator at the Human Rights Center. Previously, she worked for two years as an open source investigative specialist for an international organization in the Netherlands. She holds a B.A. from UC Berkeley in Global Studies, with a specialization in the Middle East and North Africa.
}

Subscribe to Lawfare