Machine Learning Can Predict Shooting Victimization Well Enough to Help Prevent It
- Date Posted:
- Is Database:
- Database
Machine Learning models predict extreme shooting risk concentration: Top 500 high-risk Chicago individuals show 130x city average victimization rate (13% shot within 18mo). Police data enables precise risk identification
This paper shows that shootings are predictable enough to be preventable. Using arrest and victimization records for almost 644,000 people from the Chicago Police Department, we train a machine learning model to predict the risk of being shot in the next 18 months. We address central concerns about police data and algorithmic bias by predicting shooting victimization rather than arrest, which we show accurately captures risk differences across demographic groups despite bias in the predictors. Out-of-sample accuracy is strikingly high: of the 500 people with the highest predicted risk, 13 percent are shot within 18 months, a rate 130 times higher than the average Chicagoan. Although Black male victims more often have enough police contact to generate predictions, those predictions are not, on average, inflated; the demographic composition of predicted and actual shooting victims is almost identical. There are legal, ethical, and practical barriers to using these predictions to target law enforcement. But using them to target social services could have enormous preventive benefits: predictive accuracy among the top 500 people justifies spending up to $123,500 per person for an intervention that could cut their risk of being shot in half.
Sara Heller, Benjamin Jakubowski, Zubin Jelveh and Max Kapustin, "Machine Learning Can Predict Shooting Victimization Well Enough to Help Prevent It,"National Bureau Of Economic Research, June 2022, https://www.nber.org/papers/w30170
“…Figure 2 reports two measures of model performance across the predicted risk distribution. Figure 2a shows Precision:, or the share of people who are actually shot during the 18-month outcome period among the: people with the highest predicted risk Figure 2b shows Recall:, or the share of actual shooting victims during the 18-month outcome period who are among the: people with highest predicted risk: We show two versions of recall in Figure 2b. The first, simply labeled recall, uses the total number of shooting victims in the prediction sample as the denominator, or 2,253. The second, labeled total recall, uses the total number of shooting victims in the entire city during the outcome period as the denominator, or 3,381. The difference between these two highlights a point we return to in the following section about whom predictions based on police data miss: one-third of eventual shooting victims are not in our prediction sample and therefore not assigned a predicted risk by the model. Though it is more common when evaluating the performance of a predictive algorithm to report recall, total recall helps to assess the ability of algorithmic prediction to identify shooting victims city-wide, regardless of whether they have enough prior police contact to be included in the prediction sample. The share of people shot during the 18-month outcome period is startlingly high among those in the right tail of the distribution (Figure 2a). Among the: = 500 people with highest predicted risk, 13 percent, or 65 people, are shot. This is almost 19 times higher than the base victimization rate for the prediction sample (327,127 people) of 0.7 percent, and 130 times the city-wide victimization rate (2.7 million people) of 0.1 percent. Among the: = 3, 381 people with highest predicted risk—corresponding to the actual number of shooting victims during the 18-month outcome period—almost 9 percent are shot. Those at higher predicted risk for shooting victimization are also at significantly elevated risk for other adverse outcomes, like shooting arrest and violent victimization The recall rates confirm that those in the right tail of the distribution account for an outsized share of all shooting victims (Figure 2b). Despite representing just under 0.02 percent of the city’s population, the: = 500 people with highest predicted risk include almost 2 percent of the 3,381 total victims during the 18-month outcome period.19 The: = 3, 381 people with highest predicted risk—just over 0.1 percent of the city’s population—include almost 9 percent of total victims….”
“….we build a model to predict shooting victimization in Chicago over an 18-month period. The model uses arrest and victimization records for 643,914 people from the Chicago Police Department (CPD), including over 1,400 predictors that capture a person’s demographic information, arrest and victimization histories, and the arrest and victimization histories of peers who were co-involved in prior criminal incidents….”
Results
“…First, the model successfully identifies a small group of people at extraordinarily high risk of being shooting victims. Of the 500 people at highest predicted risk, 13 percent are actually shot during the following 18 months—a rate almost 19 times higher than everyone in our prediction sample of people with recent police contact (0.7 percent across 327,127 people) and 130 times higher than everyone in Chicago (0.1 percent). An intervention that could cut the risk of being shot in half for these 500 people would generate an estimated social cost savings of $62 million from the victimization reduction alone (Cook and Ludwig, 2000; Ludwig and Cook, 2001). If the intervention cost less than $123,500 per person, it would pay for itself. Our analysis unpacks what information the model is using to achieve this predictive performance….”
“…Second, the predictions do not misrepresent victimization risk across demographic groups. We show that Black male shooting victims are more likely to have a predicted risk, because they are more likely to have prior police contact.5 This finding highlights how using police data limits an algorithm’s ability to identify future victims with little or no prior police contact. But importantly, the accuracy of predictions is similar, on average, across race, age, and gender groups. The demographic composition of predicted shooting victims matches almost exactly that of actual shooting victims. In other words, the predictions do not disproportionately inflate the victimization risk of Black men….”
Evidence
“…The top left panel of Figure 1 shows the overall distribution of the model’s predictions and how they compare to realized rates of shooting victimization for the prediction sample. The x-axis is the average predicted risk for each percentile of the risk distribution, with each point containing 1 percent of the sample, or 3,271 people. The y-axis is the actual rate of shooting victimization in the 18-month outcome period for the 3,271 people in each bin. Three features about the overall predictions are apparent. First, on average, the model’s risk predictions are accurate (well-calibrated): their slope is close to the 45-degree line, albeit with some under-prediction for people in the right tail and some over-prediction for people in the highest-risk bin. Second, the vast majority of people in the sample are predicted to have a shooting victimization risk close to zero, as indicated by the mass of points in the bottom left of the graph. Finally, the predicted risk distribution is highly positively skewed, with points in the upper right of the graph corresponding to a small group of people in the long right tail whose predicted risk of being shot in the 18-month outcome period is very high….”
Data



