Abstract
The illegal wildlife trade is a key driver of biodiversity loss and a barrier to desired transformations in socio-environmental systems. It is known to exploit licit networks, such as the global airline flight network, yet the ability of science to support efforts to reduce the illegal wildlife trade remains underdeveloped. Research on the illegal wildlife trade relies on biased aggregate counts of observed incidents, leading to potential policy misguidance. Here we utilize centrality analysis and predictive modeling to encode and decode illegal wildlife trade flight networks to improve sensemaking with implications for sustainable futures. Methods advance existing analyses by revealing traffickers prioritize airports with high centrality in the global airline flight network, high incidence of flora crimes, and a limited level of non-state actor oversight. Machine learning models identify airports likely to be involved in the illegal wildlife trade not currently implicated in seizure data, including in the United States.
Similar content being viewed by others
Introduction
The illegal wildlife trade (IWT) remains a global, multi-billion dollar industry involving illicit harvesting, transiting, and selling of wild flora and fauna and their byproducts1. IWT is a leading contributor to biodiversity loss, which, along with climate change and pollution, form what the United Nations has termed the triple planetary crisis1,2,3. Beyond degrading ecosystem integrity and health, IWT undermines sustainable development efforts in various areas, including economic prosperity, national security, gender equality, justice, and the sovereignty of local communities and indigenous peoples4,5,6,7. There is little to no evidence that IWT is decelerating. The Intergovernmental Science-Policy Platform on Biodiversity and Ecosystem Services (IPBES) has called for science-based policy responses to mitigate the impacts of accelerating biodiversity loss8. Sustainability sciences have helped advance understanding of the causes and consequences of the crisis, but knowledge of IWT remains embryonic, and what exists is highly siloed9. The lack of comprehensive data complicates efforts to draw reliable inferences or identify trends about IWT and support decision-making, ultimately hindering effective policy and decisive action10.
The most developed open-access datasets on IWT document incidents occurring via air travel on the global flight network. The dominant extant approach for scientific exploration of IWT networks using this data involves observation-based data analysis that aggregates data about discrete IWT incidents by country, species, or time10,11,12. Such analysis offers one entry point for describing the dynamic and sometimes veiled nature of IWT on a global scale. It is limited in its ability to decode existing IWT networks across spatial scales wholly and to predict or infer trends. The data is associated with multiple forms of scientific bias. First, it only includes incidents successfully detected and subsequently intercepted (i.e., detection bias). Second, a lack of standardized, global monitoring guidelines causes geographic gaps in data collection, over and under-representing different regions’ influence on IWT (i.e., representation bias). Third, data for all locations involved in an incident’s supply chain–from origin to destination– is rarely complete. Not having all these data points leaves analysts unaware of the routes taken before the product was seized (i.e., completeness bias). This bias is found widely across IWT datasets; for example, only 36% of incidents in our dataset have at least two locations documented out of a desired minimum of three along the trade route. The implications of these biases mean data cannot be wholly: a) encoded, translated, or made interoperable to identify trends or b) decoded to draw inferences or make predictions in support of decision-making.
Overcoming dataset biases is not a new issue for sustainability or illicit network science13. For example, applied research on drug and human trafficking networks has leveraged computational methods to overcome limitations associated with observation-based data analysis and derive insights into network operations, mainly through centrality measures that identify key nodes (i.e., players) and define their roles within the larger operation14,15. Machine learning models have advanced predictions about individual-level involvement within drug trafficking networks, providing multi-scalar information not readily available via observation-based data analysis16. These combined approaches may offer dramatic opportunities for encoding and decoding IWT networks but, to our knowledge, have yet to be applied. One output of machine learning models that could be useful to efforts to reduce IWT is feature importance analysis, which determines characteristics most influential for involvement in illicit networks. A second relevant output is prediction error analysis, which can be applied to identify airports predicted by the model to be involved in IWT despite lacking observational records to decode undetected trafficking hotspots.
Encoding where IWT networks are spatially located and decoding how they function are prerequisites for implementing effective interventions that dismantle IWT networks and support achieving sustainable development goals that reduce biodiversity loss17. To these ends, our objectives were to: 1) explore the structure of the global IWT flight network using degree and betweenness centrality of airports, 2) fit a model to predict airports’ involvement in IWT, 3) identify key considerations of how the IWT network functions using feature importance analysis, and 4) identify potential and previously undetected hotspots of IWT activity across the flight network.
We found the features most indicative of wildlife trafficking activity were an airport’s centrality in the global airline flight network, its incidence rate of flora crimes, and the level of non-state actor oversight in deterring crime. Additionally, our approach highlighted several likely undetected hotspots of wildlife trafficking activity. Those our model was most confident about were located in China, the United States, Indonesia, the Philippines, Mexico, and Italy. Applying these insights to IWT could help prioritize limited resources (e.g., workforce, capital), produce more precise information for decision-makers, help reduce reactive (i.e., whack-a-mole) approaches to curbing IWT, and support coordination amongst stakeholders working to reduce IWT rates- all identified as key levers of change that can transform knowledge into action and strengthen sustainable futures.
Results
Encoding: visualizing illegal wildlife trade incidents at multiple scales
Multiple studies have analyzed observed incidents of IWT, both on and off the global flight network10,11,18,19,20. We compiled the universe of up-to-date observational data on IWT incidents occurring on the global flight network to provide a baseline understanding of the information available through this type of analysis (Section “Data Compilation”). We visualized airports with IWT activity across 1.3K distinct incidents and the flight paths of 478 incidents with multiple stops (Fig. 1a). We divided IWT incidents by location type, visualizing the number of times each airport or country was documented as an origin, transit, destination, and seizure location.
a Number of observed illegal wildlife trafficking incidents per airport. Airports are depicted as red circles, with circle size proportional to the number of incidents (larger circles indicate more incidents). The top 10 airports with the highest number of observed incidents are labeled with their airport codes. Observed indicates that an illegal wildlife trafficking incident either originated from, transited through, was seized at, or was destined for the given airport. b Degree centrality of airports within the illegal wildlife trade network. Airports are shown as red circles, with circle size reflecting degree centrality (larger circles denote greater degree centrality). The top 10 most central airports are labeled with their airport codes. c Betweenness centrality of airports in the illegal wildlife trade network. Airports are represented as red circles, with circle size corresponding to betweenness centrality (larger circles indicate higher betweenness centrality). The top 10 most central airports are labeled with their airport codes.
Our analysis includes airport-level results and aggregated results at the country level (Fig. 2). Results confirmed prior research, concluding Africa as the dominant source of illegally traded wildlife and Asia as the dominant destination11. With this confirmation, we identified countries playing multiple roles in IWT. Almost half of all countries (47%) played a singular role. There is limited evidence that this reflects reality and airports don’t have undetected incidents that play other roles. Most of these single-role countries had the fewest observed incidents, signaling some of the completeness bias in the dataset.
Each country is represented by a pie chart. The values beneath each pie indicate the number of unique incidents observed in that country, followed by the percentage of those incidents seized there. Pie chart colors represent the proportion of each country’s role in illegal wildlife trade incidents: blue for origin, orange for transit, and green for destination. Origin locations are where wildlife was taken from the wild or bred in captivity; transit locations are known stops along the trade route; destination locations are the final stop or demand center. Seizure locations mark where products were intercepted, which can occur at any point along the route. A trivial number of incidents have an unknown role, as only the seizure location is documented. China is an outlier, with 24% of incidents documenting seizure only; these are included in the counts but not in the role distribution within each pie. The white, semi-transparent inner circles at the center of each pie indicate the country’s seizure rate, with larger circles representing a higher proportion of incidents seized in that country and resulting in a paler appearance of the pie chart. For example, the pale color of Bangladesh reflects a high seizure percentage, while a darker chart, like Uzbekistan’s, indicates no seizures.
Asia was the dominant demand or destination geography. The region also held multiple roles in IWT, with 78% of Asian countries having more than one documented role type within their airports. West Asian countries were more likely to serve as transit locations, bridging sources (e.g., Africa), and demand centers (e.g., East Asia). China had the largest number of observed incidents; 24% only documented the seizure location, indicating a lack of detailed information on trade paths to China. Africa was the dominant origin geography, with 89% of African countries having at least one incident where they served as the source of IWT. This trait was even more dominant when viewing data at a more granular, airport level. Ethiopia had the most IWT incidents observed in Africa, yet had a low seizure rate compared to other countries with many observed incidents, such as Kenya. Oceania, the Americas, and Europe accounted for a smaller proportion of incidents but still played noteworthy roles within the IWT network. All three regions observed high rates of seizures within airports of the most observed countries. Oceania was never observed as a transit location, illustrating varied geographic influence on IWT.
Encoded data provide a baseline for understanding IWT via the flight network, but it has a limited impact on knowledge production due to the data biases mentioned. We used computational methods to decode additional information.
Decoding illegal wildlife trade network structure with centrality metrics
Social network analysis can identify influential airports in the network by decoding the role airports play within the IWT flight network through centrality analysis. Many centrality metrics exist, but degree and betweenness centrality are commonly used to identify influential nodes (e.g., sites, actors) and their roles in illicit networks14. In social network analysis, degree centrality equals the number of edges connected to a node21. In contrast, betweenness centrality equals the number of shortest paths a given airport is on between any two airports in a network22.
Here, degree centrality represents the number of direct flights into or out of a given airport with IWT activity. Airports with a high degree centrality are crucial in the IWT network because they connect to many other airports involved in IWT, playing a vital role in the movement and distribution of IWT. The top five were Hong Kong (HKG), Bangkok (BKK), Addis Ababa (ADD), Jakarta (CGK), and Nairobi (NBO) International Airports. ADD had high degree-out centrality, indicating its common role as an origin and transit point. The other four airports had high degree-in centrality, indicating their common roles as destinations for IWT (Fig. 1b).
Airports with high betweenness centrality are crucial connection points or bridges between origin and destination airports (Fig. 1c). The top five airports ranked according to betweenness centrality were Hong Kong (HKG), Guangzhou (CAN), Bangkok (BKK), Kuala Lumpur (KUL),and Dubai (DXB) International Airports. Some airports with high betweenness centrality also had many observed incidents (e.g., CAN, HKG, BKK, ADD). In contrast, others (e.g., Moscow (DME), Dubai (DXB), Luanda (LAD)) were less prominent in incident volume but crucial for maintaining the network’s structure. The implications for practice are that tightening controls at these locations over those with lower betweenness centrality could significantly disrupt the overall flow of goods in the IWT. Given the biases in our dataset of 478 incidents, any conclusions regarding centrality may also be biased and unreliable due to unobserved IWT activity. However, these values can lend preliminary insights into the illicit network structure.
Comparing illegal wildlife trade & full flight network structures
Airports with high degree and/or betweenness centrality could be disproportionately susceptible to IWT because of their requisite role in the full flight network. To equivocate this claim, we compared the centrality of airports in the IWT network with their centrality in the full flight network, representing all available flights between all airports around the globe, not just those with observed IWT activity. We detected a low correlation between an airport’s degree (0.33) and betweenness centrality (0.35) in the IWT network and the full flight network, suggesting an airport’s role in the global flight network is not necessarily the same role it plays in the IWT network (Fig. 3a). In the full flight network, we observed a small number of airports of higher centrality with a heavy-tail distribution (Fig. 3b). The higher the degree or betweenness centrality of an airport in the full flight network, the more likely that airport was to be observed in the IWT network (Fig. 3c). Encouraged by this result, we constructed a model that predicts airport-level involvement in IWT, using centrality in the full flight network as features.
a Airport degree and betweenness centrality in illegal wildlife trade (IWT) versus full flight network. Red points signify the airport was observed in our IWT dataset, while gray points were not observed. The top five airports by centrality metric in the IWT network, the top two in the full flight network, and the top two in the full flight network but not observed in the IWT dataset are labeled with their airport codes. b Distribution of airports across centrality bins. Ten bins categorize airports by centrality values from the full-flight network. Blue bars (degree centrality) and green bars (betweenness centrality) show airport counts per bin, revealing skewed distributions toward lower-centrality airports. c Proportion of airports per centrality bin observed in IWT network. The 10 bins use the same full-flight centrality ranges as (b). Higher-centrality bins show increased IWT participation, suggesting traffickers exploit well-connected airports by degree (blue bars) and betweenness (green bars) centrality.
Modeling decoded data to predict airport-level involvement in illegal wildlife trade
A suite of granular, airport-level trends and inferences can be drawn from modeling decoded global IWT data. We developed a binary classifier to predict whether or not an airport is implicated in the IWT network and further explored the data using feature importance and prediction-error analyses to uncover characteristics of airports that affect their involvement in IWT and to identify unobserved IWT hotspots.
Before building our model, we considered a range of characteristics or features associated with airports (Supplementary Data 1). These features were selected based on literature documenting the relationship between IWT (i.e., wildlife crime) and other criminal markets and actors1,23,24. We also considered 7 features associated with airport centrality, such as degree and betweenness centrality, due to our results in Fig. 3c.
After removing highly correlated features, conducting feature selection, and model optimization, we refined our model to utilize ten features and trained a Balanced Random Forest Classifier (Section “Predictive Model Formulation”) to predict the involvement across all 1933 airports in the full flight network25,26. Our model achieved a ROC score of 0.91 and a class-weighted F1 score of 0.8427. These scores exhibit strong baseline performance, suggesting the model can be a foundation for future enhancements.
Discussion
This work joins other sustainability-science-related works that utilize computational methodology to explain the causes and consequences of the triple planetary crisis9. Our study culminates in the development of a predictive model that provides at least two outputs useful to reducing the volume and impact of IWT globally: feature importance and prediction error analysis. Feature importance analysis highlights the airport characteristics that influence airports’ participation in the IWT, offering possible explanations for route selection. Prediction error analysis can uncover potential veiled hotspots of IWT activity. Both of these deliverables go beyond the breadth of information widely available and utilized for decision-making today, leaving the opportunity for more targeted and, ideally, effective interdiction practices that advance sustainable development goals in just and effective ways.
Factors influencing airport-level involvement in illegal wildlife trade
The ten features utilized in our model can be organized into four categories: location (relating to an airport’s position in the full flight network), resilience (measures countries have in place to achieve solutions to organized crime), criminal actors (structure and influence of criminals in certain groups), and criminal markets (political, social, and economic systems surrounding all stages of the illicit trade in and/or exploitation of commodities or people). With the exception of the centrality features we calculated, all remaining features were sourced from the Global Initiative against Transnational Organized Crime (GITOC) and their 2021 crime index28. Features in our model that fall within the location category include the betweenness and degree centrality of the airport in the full flight network. Resilience features include non-state actors (the degree to which civil society organizations are able to play a role in responding to organized crime), territorial integrity (the degree to which states can control their territory against criminal activities), prevention (existence of strategies and resources aimed to inhibit organized crime), and political leadership and governance (effectiveness of a state’s government in responding to organized crime)28. Features in criminal markets include the following: flora crimes (illicit trade/possession of protected flora), non-renewable resource crimes (illicit extraction/smuggling of natural resources), and human trafficking (exploitation of others through prostitution, slavery, or removal of organs)28. The sole feature in the criminal actor’s group is foreign actors (state or non-state criminals operating outside their home country)28.
We utilized explainable AI to analyze features’ importance and subsequent influence in our model. Using SHAP, we examined the role each feature played in calculating the likelihood each airport was involved in IWT with Shapley values29. Shapley values are based in game theory30. In our context, they measure the contribution of each feature to the predicted assignment of whether an airport was involved in IWT or not. This granular-level insight, only detectable from this type of analysis, can be used in response to the strategic and tactical decisions of individuals involved in IWT when they choose one airport over another in a similar geographic region.
The analysis shows that of the selected ten features, those related to an airport’s position within the full flight network emerged as the most influential to IWT involvement (Fig. 4). An airport’s centrality impacts airline ticket prices as well as the duration of flight time between source, transit, and destination geographies31. Airport-specific variations in ticket price and flight time can influence the risk of detection and total cost of operation, both of which have been used in IWT interdiction strategy32,33.
The y-axis lists features used in the predictive model, color-coded by category (pink: location, blue: resilience, orange: criminal actors, green: criminal markets). Features are ordered by importance, with the most influential at the top. Each feature is accompanied by a swarm plot, where each point represents an airport in the dataset. Points are colored according to their raw feature values, ranging from blue (low) to pink (high); for example, a blue point for betweenness centrality indicates an airport with a relatively low betweenness score. The x-axis shows the Shapley value, which measures each feature’s influence on the model’s prediction for that airport. Negative Shapley values indicate that a feature decreases the likelihood of predicting illegal wildlife trade activity, while positive values indicate an increased likelihood. Betweenness centrality in the full flight network emerged as the most important predictor, followed by degree centrality, suggesting these network features can guide the allocation of limited enforcement resources. When multiple airports have similarly high centrality, results suggest prioritizing interventions at airports with either a high rate of flora crimes or a low presence of non-state actors.
Knowing airport centrality impacts the likelihood of involvement in IWT enables additional feature analysis that can help identify what other airport characteristics are most influential for involvement in IWT and whether those features positively or negatively impact the likelihood of involvement. Among the resilience features, high scores in these features (higher resilience) lowered the model’s likelihood of predicting an airport to be involved in IWT, as seen by the dominance of red dots on the negative SHAP value x-axis (Fig. 4). Results suggest that among all IWT-related interventions, counter-crime measures may play an outsize impact on deterring IWT, particularly when implemented with civil society organizations that help respond to crime. Conversely, airports with high scores in the criminal actors and market categories had an inverse effect. Features with high scores in these categories (higher crime) were more likely to be predicted to be involved in IWT, as seen in the dominance of red dots on the right-hand positive SHAP x-axis. Flora crimes stand out as the most important non-location variable.
Identifying airports with unobserved illegal wildlife trade activity
The model could help overcome some of the well-documented challenges associated with analyzing global IWT data as it isn’t wholly dependent on observational data. One noteworthy outcome is that our model can deduce otherwise obfuscated IWT hotspots due to the lack of standardized and widespread IWT data being collected and shared. Using our predictive model, we identified airports that potentially have unobserved IWT activity by investigating airports our model incorrectly identified as participating in the IWT flight network despite not having any observations in our dataset, known as false positives. Our model predicted that 307 airports had IWT activity despite not being observed in our limited dataset with varying confidence (Fig. 5).
Circles represent airports our model identified as participating in the illegal wildlife trade despite no records of its involvement in dataset. Intensity of circle’s red color corresponds to model confidence in that airport’s likelihood to be involved in the illegal wildlife trade (IWT). The darker the red, the more confident the model. The false positive airports were geographically scattered. Airports with the highest confidence were primarily in Asia and North America.
We ranked the false positive airports according to the model’s confidence (Table 1). An example of a false-positive airport in which the predictive model exhibited high confidence in its assignment (100%) is Shanghai Hongqiao International Airport (SHA). SHA is one of two international airports in China’s most populated city, but was not present in our observed dataset. There is evidence that IWT trafficking has and could occur at this location; open access financial risk management reporting from Themis and Chinese Government TV news media have documented rhino horn and related products being trafficked through the airport34,35. Another airport is Taiyuan Wusu Airport (TYN), which serves the capital of north China’s Shanxi Province. Similar to SHA, TYN was not observed in our limited dataset yet is evidenced as being used for IWT by conservation organization reporting36.
Eleven airports were implicated by our model with high confidence despite not being observed. These airports are located in China, Indonesia, Philippines, Mexico, Italy, and the United States (US). Generally, the false positive airports with the highest model confidence (ρ > 0.95) were not surprising. The US Department of State released a list of 28 focus countries (major sources, transit points, or consumers of wildlife trafficking products) and six countries of concern (governments knowingly profit from IWT and are major sources, transit points, or consumers) in 202137. The two countries in Table 1 not mentioned in the list include Italy and the US. Only recently has Italy been publicized as playing a role in IWT in Europe, which is primarily driven by illegal plant-derived medicines38. The model identified key locations before it is analyzed and otherwise reported with other means, legitimizing our predictive approach. The role of the US in observational analysis may be overshadowed due to the overwhelming amount of data focused on other regions such as countries in Africa and Asia11. The US remains a prominent IWT destination location, as it is ranked fourth overall and first when excluding Asian countries in destination counts (Fig. 2a).
Conclusions
Accelerating rates of global biodiversity loss dynamically undercut varied cross-sectoral investments to create sustainable and healthy futures. Well-intentioned efforts to accelerate achievement of sustainable development and wildlife conservation goals may at best fail to achieve objectives and at worst exacerbate biodiversity loss, because of limited encoding and decoding of data. This work builds upon prior interdisciplinary research that recognizes responding effectively to IWT requires knowledge of networks and network dynamics and that science in support of applied knowledge production can be inherently riddled with bias. The computational science methods and models used herein can help overcome existing limitations and transform knowledge into action to reduce IWT and its contribution to accelerating biodiversity loss. Overall, what is noteworthy from this analysis is a) the granularity and b) the methodology. To our knowledge, our study is one of the first to both encode and decode statistics on IWT on an airport level, adding more granular insights. Additionally, the usage of this scaled-down data to build a predictive pipeline may unearth lower-level insights into how traffickers select the airports they visit and also highlight potential undetected hotspots of IWT activity.
Methods
Data compilation
Data sources
We gathered data on IWT incidents occurring via the flight network from two of the most expansive resources on IWT activity. The first is from ROUTES, a conglomerate of transit companies, government agencies, conservationists, law enforcement, and academics working to address global wildlife crime39. Their corresponding airline traffic database contains data on 975 unique trafficking incidents on the airline network between 2010 and 2020. The second resource comes from TRAFFIC, a global NGO focused on ensuring the wildlife trade is both legal and sustainable40. They created the Wildlife Trade Portal, an online portal that allows users to query for wildlife seizure data and filter by location, species, transportation type, and time41. The Wildlife Trade Portal is not exclusive to just incidents occurring on the flight network, so we filtered for incidents with at least one documented location occurring at an airport by gathering all incidents with Airport listed in the Transit Type column. As TRAFFIC was a member of the ROUTES partnership, we only counted incidents occurring after 2020 to avoid double-counting. TRAFFIC yielded 346 incidents between 2021 and 2023, totaling 1321 incidents.
Building the full flight network
To build the full flight network consisting of all airports and flights available across the globe, not just those with observed trafficking activity, we collected data from OpenFlights.org, which hosts open-source information on airports and flights available globally42. Any references to an airport’s centrality in the full flight network were calculated using the equations in Section “Centrality Calculations” using all 1933 airports/nodes and 14,118 flights/edges. Importantly, this database is not updated instantaneously when new flight routes begin. Therefore, after verifying the flight existed by searching Google Flights, we manually added a small number of routes present in our dataset of 1.3K IWT incidents. This ensures the maximum number of IWT records were captured and that the IWT flight network is a proper subgraph of the full flight network.
Building the IWT flight network
To build the IWT flight network, we used a subset of the 1321 IWT incidents. We filtered the 1321 incidents for those with at least two documented stops at airports. This was done to ensure the edges strictly represented flights between two airports involved in the IWT. The size of the resulting dataset consisted of 478 distinct incidents, showcasing the completeness bias of the dataset. Any references to an airport’s centrality in the IWT network were calculated using the airports and flights observed in the 478 incidents using the equations in Section “Centrality Calculations”.
Centrality calculations
Degree centrality
Degree centrality equals the number of edges connected to a node21. In the full flight network, it represents the number of direct flights into or out of a given airport. In the IWT flight network, degree centrality represents the number of direct flights into or out of airports with documented IWT activity. Airports with high degree centrality are crucial in both the full flight network and IWT networks, as they connect to many other airports, playing a vital role in the movement and distribution of IWT. We calculated degree centrality of airports in both the full flight and IWT networks using the NetworkX package in Python, which normalizes the centrality to values between 0 and 1 using Eq. (1).
Betweenness centrality
Betweenness centrality equals the number of shortest paths a given airport is on between any two airports in a network22. Airports with high betweenness centrality are crucial connection points or bridges between origin and destination airports. We calculated the betweenness centrality of airports in both the full flight and IWT networks using the NetworkX package in Python, which normalizes the centrality to values between 0 and 1 using Eq. (2).
where σst is the number of shortest paths between source s and target t and σst(v) are those that pass through node v.
Predictive model formulation
Feature overview
We considered 44 features of the airports in four different categories: location, the presence of criminal markets, criminal actors, and counter-crime resilience measures.
Features that fall under location metadata include the population of the city the airport is located in and whether the country is a member of CITES (the Convention in International Trade in Endangered Species of Wild Fauna and Flora, a nonbinding agreement to protect endangered plants and animals from the threats of international trade)43. We collected data on whether a country was a signatory of CITES from 2016 to 2021. Location metadata also includes centrality measures for the airports and their positions in the full flight network, using the data gathered in Section “Data Compilation”. We only included centrality metrics of the full flight network because the label we aimed to predict was participation in IWT. Therefore, adding features that rely on knowing IWT participation induces a circular process. The features in the three other categories originate from the 2021 Global Crime Index, measured by the Global Initiative against Transnational Organized Crime (GITOC). This organization provides numerical scores on a country’s criminal market, criminal actors, and resilience measures. The features in these areas are intended to quantify the levels of organized crime in a country and assess their resilience to organized criminal activity28. Each country’s value is evaluated over two years and draws from both quantitative and qualitative sources. The final scores are decided by 400+ experts and GITOC’s regional observatories. All features and their subsequent values were normalized to values between 0 and 1.
These categories of features were selected based on prior works that highlight the convergence of multiple forms of illicit trade4,44. Because little data is available on wildlife trafficking, more well-researched illicit networks, such as drug and human trafficking, can lend insights into how wildlife traffickers operate. GITOC gathers this data and categorizes the features into three groups: criminal actors, criminal networks, and resilience. We considered all feature categories in our model. The last category, location, was included to highlight the centrality measures we gathered as well as other location-specific features we hypothesized could hold predictive power of IWT involvement (i.e., CITES membership and population).
Our dependent variable, IWT participation, is a binary representation of whether the airport was observed to have IWT activity in the dataset of 1.3K incidents described in Section “Data Compilation”. Our dataset exhibited significant class imbalance, with only 311 (16%) airports having observed IWT activity and the remaining 1622 having no documented IWT activity.
Full descriptions of all the features used in modeling can be found in Supplementary Data 1.
Feature selection
To improve the performance of our predictive model, we completed two feature reduction measures. The first was removing features that were highly correlated with one another to reduce the risk of overfitting and increase model interpretability45,46,47. To determine which features to remove, we ranked features by their Pearson correlation with the dependent variable, IWT participation, using the pearsonr function from the scipy.stats Python library. We then iteratively eliminated those with a correlation above 0.80 with more significant features. This process removed 13 features.
With the initial reduction of features complete, we trained and tested a series of models to select the best-performing one with the remaining 31 features. Using the optimal model, we then conducted both forward and backward feature selection on the remaining features to identify which method maximized the model’s ROC-AUC Score48,49. Backward feature selection works by starting with all available features and progressively eliminating the feature that, when excluded, results in the highest ROC-AUC score for the remaining set until the target number of features is met. Forward feature selection works by starting with an empty set of features and adds them one-by-one, selecting the feature that, when included, produces the highest increase in the ROC-AUC score until the target number of features is met. To limit computational resources, we attempted both methods, specifying that either 5, 10, 15, 20, 25, or 30 variables were selected, picking whichever variable count yielded the highest ROC-AUC score. Using the final variable count output in that process, we then utilized a sequential process of choosing which variables to include by maximizing the ROC-AUC score at each variable addition. We used 5-fold cross-validation to evaluate performance across all data points at each iteration. Backward feature selection returned both the maximum ROC-AUC and class-weighted F1 score of the two methods when utilizing ten features, so we proceeded with modeling using features the following features: betweenness centrality, degree centrality, non-state actors, territorial integrity, non-renewable resource crimes, floral crimes, foreign actors, prevention, political leadership and governance, and human trafficking.
Model selection & execution
We attempted multiple classification models, including logistic regression, decision trees, random forests, balanced random forests, gradient boosting machines, and linear discriminant analysis, using SMOTE oversampling where appropriate, and 5-fold cross-validation to evaluate performance. We prioritized traditional machine learning methods over black-box models as conservation practitioners highlight the importance of adopting interpretable methods to understand, in our case, why wildlife trafficking occurs at some locations over others50,51,52. All modeling techniques were executed using Python and its various libraries. We used 5-fold cross-validation to evaluate each model’s performance at testing to ensure each data point was evaluated53. The balanced random forest model achieved the highest ROC-AUC and class-weighted F1 score, so we proceeded with this model utilizing imbalanced learn’s ensemble.BalancedRandomForest package54.
Classical random forests can be biased towards the majority class where significant class imbalance exists, such as our dataset where only 311 (16%) airports were observed with IWT activity26. Balanced models can better handle imbalanced datasets, specifically reducing bias towards the majority class using bootstrapping, meaning the number of samples drawn from each class is the same, ensuring the minority class is represented in each tree26. Using a balanced model in our study methods allows us to extract information from the few positive samples in our dataset to make inferences on.
We also performed hyperparameter tuning, though we proceeded with default parameters as performance changes were negligible. Supplementary Table 1 showcases the ROC-AUC scores of each model’s performance with the 31 features remaining after the correlation analysis removed 13 features discussed previously, as well as the best-performing model after conducting BFS and FFS - the balanced random forest model with 10 features selected with BFS. To showcase performance improvements due to removing highly correlated variables, we also reported the ROC-AUC score of the balanced random forest using all 44 variables.
Data availability
An airport-level dataset documenting wildlife trafficking incidents is available. This dataset includes wildlife trafficking incident counts and values for all variables considered as potential features in our predictive model broken down by airport. Accompanying network files detail both the full flight network and the illegal wildlife trade network. Lastly, our Supplementary Data 1 file describing all features, their geographic granularity, if they were included in our final model, and their feature category is available. This data can be accessed on Figshare at https://doi.org/10.6084/m9.figshare.27873024.v2.
Code availability
The scripts used to build the trafficking and full flight networks and the scripts used to build the balanced random forest model for prediction are available on Zenodo at https://doi.org/10.5281/zenodo.14199743.
References
Mozer, A. & Prost, S. An introduction to illegal wildlife trade and its effects on biodiversity and society. Forensic Sci. Int. Anim. Environ. 3, 100064 (2023).
United Nations - Climate Change. What is the triple planetery crisis? (UN, 2022).
Morton, O., Scheffers, B. R., Haugaasen, T. & Edwards, D. P. Impacts of wildlife trade on terrestrial biodiversity. Nat. Ecol. Evol. 5, 540–548 (2021).
Ferber, A., Griffin, E., Dilkina, B., Keskin, B. & Gore, M. Predicting wildlife trafficking routes with differentiable shortest paths. In: Integration of Constraint Programming, Artificial Intelligence, and Operations Research 460–476 (CPAIOR, 2023).
Gore, M. L. et al. Voluntary consensus based geospatial data standards for the global illegal trade in wild fauna and flora. Sci. Data 9, 267 (2022).
Wyatt, T. The security implications of the illegal wildlife trade. J. Soc. Criminol. August, 130–158 (2013).
Nellemann, C. et al. The rise of environmental crimes - a growing threat to natural resources, peace, development, and security. A UNEP-INTERPOL Rapid Response Assessment (UNEP, 2016).
Ruckelshaus, M. H. et al. The IPBES global assessment: pathways to action. Trends Ecol. Evol. 35, 407–414 (2020).
Hino, M., Benami, E. & Brooks, N. Machine learning for environmental monitoring. Nat. Sustain. 1, 583–588 (2018).
The World Bank. Global wildlife program - 2023 progress report (The World Bank, 2023).
Utermohlen, M. Runway to extintion: wildlife trafficking in the air transport sector (C4ADS, 2020).
Patel, N. G. et al. Quantitative methods of identifying the key nodes in the illegal wildlife trade network. Proc. Natl Acad. Sci. 112, 7948–7953 (2015).
Christie, A. P. et al. Quantifying and addressing the prevalence and bias of study designs in the environmental and social sciences. Nat. Commun. 11, 6377 (2020).
Bright, D. A., Greenhill, C., Reynolds, M., Ritter, A. & Morselli, C. The use of actor-level attributes and centrality measures to identify key actors: a case study of an australian drug trafficking network. J. Contemp. Crim. Justice 31, 262–278 (2015).
Vivrette, A. T. Approach to the global human trafficking crisis: analyzing applications of social network analysis. Master’s thesis (Vanderbilt University Graduate School - Medicine, Health, and Society, 2022).
Atsa’am, D. D. et al. A model for predicting the class of illicit drug suspects and offenders. J. Drug Issues 52, 168–181 (2022).
Hughes, L. J. et al. Global hotspots of traded phylogenetic and functional diversity. Nature 620, 351–357 (2023).
Rosen, G. E. & Smith, K. F. Summarizing the evidence on the international trade in illegal wildlife. Ecohealth 7, 24–32 (2010).
Hughes, A. C. Wildlife trade. Curr. Biol. 31, R1218–R1224 (2021).
Li, Y. et al. Quantifying global colonization pressures of alien vertebrates from wildlife trade. Nat. Commun. 14, 7914 (2023).
Golbeck, J. Chapter 3 - network structure and measures, 25–44 (Morgan Kaufmann, 2013). https://www.sciencedirect.com/science/article/pii/B9780124055315000031.
Golbeck, J. Chapter 21 - analyzing networks, 221–235 (Syngress, 2015). https://www.sciencedirect.com/science/article/pii/B9780128016565000214.
Anagnostou, M. & Doberstein, B. Illegal wildlife trade and other organised crime: A scoping review. Ambio 51, 1615–1631 (2022).
Feltham, J. Convergence of wildlife crime with other forms of organised crime (Wildlife Justice Commission, 2021).
Khalilia, M., Chakraborty, S. & Popescu, M. Predicting disease risks from highly imbalanced data using random forest. BMC Med. Inform. Decis. Mak. 11, 51 (2011).
Chen, C., Breiman, L. & Liaw, A. Using random forest to learn imbalanced data. Technical Report (University of California, 2004).
Goutte, C. & Gaussier, E. A probabilistic interpretation of precision, recall and f-score, with implication for evaluation. In: Proceedings of the 27th European conference on advances in information retrieval research 345–359 (ECIR, 2005).
Adal, L. et al. Global organized crime index 2021. GITOC - the global initiative against transnational organized crime (2021).
Lundberg, S. M. & Lee, S.-I. A unified approach to interpreting model predictions. In: Proceedings of the 31st international conference on neural information processing systems 4768–4777 (NIPS, 2017).
Rozemberczki, B. et al. The Shapley value in machine learning. In Proc. thirty-first international joint conference on artificial intelligence, IJCAI-22, 5572–5579 (2022).
Silva, T. C., Dias, F. A. M., dos Reis, V. E. & Tabak, B. M. The role of network topology in competition and ticket pricing in air transportation: evidence from brazil. Phys. A Stat. Mech. Appl. 601, 127602 (2022).
Griffin, E. C. et al. Interdiction of wildlife trafficking supply chains: an analytical approach. IISE Trans. 56, 355–373 (2024).
Smith, J. C. & Song, Y. A survey of network interdiction models and algorithms. Eur. J. Operational Res. 283, 797–811 (2020).
CGTN. Two caught smuggling rhino horns into Shanghai, https://www.youtube.com/watch?v=Grco0ZeProM (2019).
Themis. IWT red flags - illegal wildlife trade toolkit, https://wearethemis.com/media/maxhcmyn/illegal-wildlife-trade-red-flags.pdf.
TRAFFIC. China’s wildlife enforcement news digest, http://www.trafficchina.org/sites/default/files/chinas_wildlife_enforcement_news_digest_may_2019.pdf (2019).
United States Department of State. Report to congress on the eliminate, neutralize, and disrupt wildlife trafficking act of 2016, https://www.state.gov/2021-end-wildlife-trafficking-report/ (2021).
TRAFFIC. Wildlife trade report: EU is hub for illegal trade in protected wild species, report finds, https://www.traffic.org/publications/reports/an-overview-of-seizures-of-cites-listed-wildlife-in-the-eu-in-2022/ (2024).
TRAFFIC. USAID ROUTES partnership, https://www.traffic.org/what-we-do/thematic-issues/private-sector-guidance/routes/.
TRAFFIC. Stop widlife trafficking, https://www.traffic.org/.
TRAFFIC International. Wildlife trade portal, https://www.wildlifetradeportal.org (2024).
Openflights: Flight logging, mapping, stats, and sharing, https://openflights.org/ (2024).
CITES: Convention on international trade in endangered species of wild fauna and flora. https://cites.org/eng.
Gore, M. L. et al. Transnational environmental crime threatens sustainable development. Nat. Sustain. 2, 784–786 (2019).
Liu, Y., Zou, X., Ma, S., Avdeev, M. & Shi, S. Feature selection method reducing correlations among features by embedding domain knowledge. Acta Mater. 238, 118195 (2022).
Toloşi, L. & Lengauer, T. Classification with correlated features: unreliability of feature ranking and solutions. Bioinformatics 27, 1986–1994 (2011).
Kyriazos, T. & Poga, M. Dealing with multicollinearity in factor analysis: the problem, detections, and solutions. Open J. Stat. 13, 404–424 (2023).
Pudjihartono, N., Fadason, T., Kempa-Liehr, A. W. & O’Sullivan, J. M. A review of feature selection methods for machine learning-based disease risk prediction. Front. Bioinforma. 2, 927312 (2022).
Fawcett, T. Introduction to ROC analysis. Pattern Recognit. Lett. 27, 861–874 (2006).
Roche, D. G. et al. Closing the knowledge-action gap in conservation with open science. Conserv. Biol. 36, e13835 (2021).
Branco, V. V., Correia, L. & Cardoso, P. The use of machine learning in species threats and conservation analysis. Biol. Conserv. 283, 110091 (2023).
Tuia, D. et al. Perspectives in machine learning for wildlife conservation. Nat. Commun. 13, 792 (2022).
Refaeilzadeh, P., Tang, L. & Liu, H. Cross-validation, 532–538 (Springer US, 2009).
ImbLearn. Balanced random forest classifier. https://imbalanced-learn.org/stable/references/generated/imblearn.ensemble.BalancedRandomForestClassifier.html.
Acknowledgements
The authors wish to thank the anonymous reviewers for their valuable contributions to our manuscript and to the individuals at TRAFFIC for granting us access to their database through their Wildlife Trade Portal. The authors were supported by U.S. National Science Foundation award CMMI-1935451 “Detecting and Interdicting Illicit Wildlife Trafficking Supply Chains” and Paul G. Allen Family Foundation (PGAFF) grant “Operation Pangolin: Unifying Diverse Data Streams to Redefine Species Conservation”. The information contained herein does not represent the opinions of the U.S. Government, PGAFF, or any author affiliations.
Author information
Authors and Affiliations
Contributions
The authors of this work and their contributions to this research are as follows. Hannah Murray contributed through conceiving and designing experiments, performing experiments, analyzing data, contributing analysis tools, and through writing the paper. Meredith Gore analyzed data and assisted with writing the paper. Bistra Dilkina contributed by conceiving experiments, analyzing data, contributing analysis tools, and assisted with writing the paper.
Corresponding author
Ethics declarations
Competing interests
The authors declare no competing interests.
Peer review
Peer review information
Communications Earth & Environment thanks the anonymous reviewers for their contribution to the peer review of this work. Primary Handling Editors: Yann Benetreau and Heike Langenberg. A peer review file is available.
Additional information
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and permissions
Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
About this article
Cite this article
Murray, H., Gore, M.L. & Dilkina, B. Encoding and decoding illegal wildlife trade networks reveals key airport characteristics and undetected hotspots. Commun Earth Environ 6, 399 (2025). https://doi.org/10.1038/s43247-025-02371-5
Received:
Accepted:
Published:
Version of record:
DOI: https://doi.org/10.1038/s43247-025-02371-5







