License Plate Reader Searches Should Require a Warrant

So while I work with police departments regularly, I think it is critically important that technology be used reasonably.

While this may be off-putting to some of my clients, I worked with the Institute for Justice as an expert witness in their trial Schmidt v City of Norfolk. (Any opinions herein are my own and not those of IJ, to be clear.) The gist of that case was whether searches of historically cached ALPR data (automated-license-plate-reader) constituted an illegal search.1

The judge ruled against plaintiffs in that case. Here is a quote from the judgment:

Consistent with Plaintiffs’ claims in this case and controlling precedent involving mass surveillance in public spaces, ALPR surveillance could become too intrusive and run afoul of [constitutional privacy standards] at some point. But when? While a definitive answer to that question is elusive, what is readily apparent to this Court is that, at least in Norfolk, Virginia, the answer is: not today.

The important point to note about this quote is “not today”. This will be a long winded post, but to try to keep it simple:

  • I think cameras will become ubiquitous in the foreseeable future. So the question is not if this data will require a warrant, it is when. It is going to happen eventually under current case law.
  • I think cameras are good, and can be used to reduce crime in a cost effective manner.
  • There is a difference between active flags (e.g. this car is stolen and it pings the PD when it drives past a camera) vs historical searches (e.g. look to see where license plate XYZ1000 was the last 30 days).
  • Requiring a warrant for historical searches will not seriously impede police investigations.
  • The current status quo of not retaining data is VERY BAD; it does not prevent illegal searches, and currently limits the utility of actually using that data for legitimate investigations.
  • Current standards to prevent abuse of the searching ALPR data systems are laughable.

Long story short in my opinion everyone would be better off if states just mandated warrant procedures through state statutes.

To try to not get too much into the weeds of what historically constitutes a search, I think the easiest place to start is via Carpenter vs US. So current US case law requires police departments to obtain a warrant to request cellular providers provide law enforcement with cell phone tower pings (cell-site location information, CSLI).

This deviated from historical precedent in requiring a warrant mainly because it was private companies that had the information. Before Carpenter, mostly it was argued you did not have a reasonable expectation of privacy if a private company could access the same data. The court in Carpenter basically made a determination that cell phone data was so comprehensive it justified a different standard – that you could track the whole of a person’s movements with the detailed CSLI data. And that this level of invasiveness violated a reasonable person’s expectation of privacy. Even if Google had all that info, you did not expect them to give it away.

This opinion was reaffirmed with the recent Chatrie decision (for geofence warrants, e.g. give me a ping for all cell phones in area X and datetime-range Y). Another relevant decision to be aware of is also Beautiful Struggle v Baltimore, in which searching historical aerial imagery via drones also constituted a search.

So this is why I am saying the question is when, not if, ALPR data will require a warrant. If a city happened to have a camera on literally every intersection (which I think will happen in the future), under current case law it would clearly be the same situation as you have for your cell phone data.

Cameras are Good

To be brief, again I mostly work with police departments in my career and was a former crime analyst. I do think ALPR cameras are good investments, mainly because they are cheap enough to have a reasonable return on investment. (Note I do not think this about all police tech, I am particularly critical of the price tag for acoustic-gun-shot-detection.)

So ALPRs are well under $3,000 per camera. The machine learning models, camera, and computation necessary to flag a plate when it passes can easily fit on current cell phones. (The harder part is powering the phone and protecting it from the elements.) ALPRs for the most part just take static images and then extract out the license plate (and for some vendors extract out additional information, like car make and color).

The overall evidence that ALPRs reduce crime is pretty meh at the moment (see my slides at a Wake Libertarian talk I did in 2024), but because they are so cheap they really only need to increase a few arrests per camera to likely have a positive return on investment.

It is pretty hand-wavy, as we do not have estimates for the value of increased clearances I find persuasive. But I think saying “I would pay $500 to help solve one case” is on the low side if anything. So a single camera if it helps catch just a handful of crimes a year is likely in my opinion to be a positive ROI.

I think cameras in all public spaces are going to happen. Imagine Ring comes out with a nicer camera system for homeowners that has more comprehensive views around your house and is just as cheap. And we will ultimately be safer for it. So even for folks advocating that cities do not pay for Flock, this is coming anyway in the near future.

Historical Searches vs Active Flags

ALPRs have been around a long time. The first ones I worked with at Troy, NY when I was an analyst were in-car cameras. Basically a go pro attached to the window that alerted when an officer drove by a stolen plate.

While ALPRs initial use was always pitched as this active flagging of stolen vehicles, they were used right away to retroactively search the historical locations of plates. They had a log of every plate, lat/lon, and timestamp of when that car passed a camera.

So imagine you are conducting an investigation of Joe Schmo, you know his license plate, and then you can type in his plate and see where his car passed a camera. Based on this information, same as CSLI data, you can basically trace where Joe went, where he repeatedly visited, where he likely slept, etc. (The first time I used this at Troy, we figured out a particular individual we were actively investigating was living with his girlfriend for example. I was honestly amazed how densely filled in the map was of hits for a single plate based on the in-car cameras.)

You technically do not need to cache any data at all to accomplish this “flag a stolen vehicle” (or any other scenario where you are actively looking for a specific license plate). There are legitimate scenarios though where ALPR searches for recent data in a real time context can be very helpful.

One of the more common examples – someone robs a gas station, and they drove a vehicle. You don’t know the plate, but can look at the images that passed by the fixed location ALPRs in the time range, and then especially if you have a car description from the gas station attendant can figure out the plate associated with the vehicle.

To be clear I am not a lawyer, but in my opinion I think exigent circumstances make searching a few minutes of cached ALPR location data totally reasonable. In practice, New Hampshire’s 3 minute data retention is far too short. I could see arguments for several hours (imagine “I found a dead body on the side of the road”, that requires more time for it to be reported.) But we are meandering into the territory where it is not an active emergency “need to find someone who may have a gun and hurt people” that would justify those exigent circumstances. Those are the scenarios where getting a warrant is reasonable (no different than a geofence warrant if you do not have a plate and want to just search what cars passed by a camera within a certain date-time window, or no different than a CSLI warrant if you have an active suspect and want to search for a specific license plate).

Most states are retaining ALPR data for longer periods. While the Norfolk case was ongoing, Virginia set a standard across the state at 21 days. Before that it was up to the individual agency. It varies state by state, but states often mandate data retention around 30 days, or leave it up to the discretion of the police department.

Deleting Data does not prevent abuses

These data retention statutes are argued as a mechanism to prevent abuse. They do not accomplish this.

If you look through the cases in which officers abused the system to search for individuals, all of them searched for specific plates over-and-over again, sometimes hundreds of times.

If you retain data for 20 days, you can just go and do a search every 20 days, keep notes on the data as you so wish, and then do another search 20 days later. Getting rid of old data, in-and-of-itself, does nothing to prevent that abuse. In fact if someone is actively stalking a person, you would expect them to regularly do searches, seeing where their victim is going on a regular basis while they have access to the system.

Simultaneously, deleting data does prevent its legitimate use in long term law enforcement investigations. It is totally normal for a murder investigation to take more than 30 days to identify a suspect. Gosh, sure would be nice to be able to then query the ALPR data to show whether a person was in the vicinity of the murder. Simultaneously it could be used by the defense for exculpatory purposes (which assuredly would take longer than 30 days).

So folks advocating for deleting data as a mechanism to prevent abuse are making things worse. It does not prevent abuse, and limits the utility of ALPR for historical investigations. The only way data retention by itself prevents abuse is if you do not cache data at all (like in New Hampshire), and only use ALPRs for the active alert situation.

What Smart Regulation Looks Like

One of the reasons I say that the current standards to prevent abuse are laughable is that data retention policies and internal PD policies on when the data should be searched have been in place in most departments for years (if not a decade) at this point. The examples where searching ALPR data to stalk an intimate partner were obviously not prevented via data retention policies.

Alas, my suggestion that some data is cached for real time investigations (longer than 3 minutes), and that a warrant should be required outside of this window, does not prevent that type of abuse either. Most departments have in place reasons why a search can be conducted, and some states have specific statutes identifying impermissible reasons for conducting searches. In the Norfolk IJ case, officers, when entering a reason for a search (which was often omitted), sometimes supplied reasons that appeared prima facie illegal, such as “protest”.

Departments, even if they have a standard to do internal audits, often do not follow them. It took Tyler Dukes asking Raleigh PD for their audit results for them to even conduct their first audit.

This is a long standing problem for PDs, not just with ALPRs, but also with searching criminal history illegally. IJ collating a dozen cases of arrests of ALPR misuse across the country is not evidence these systems are working, as it is likely the case that only the most egregious abuses are ever caught.

In addition to creating state statutes to mandate that a warrant be used for historical ALPR searches, states should, at a minimum, have clear punishments for illegal searches. These should include at a minimum losing your job, and being banned from accessing the system forever. When I was a crime analyst in New York (and ditto for when I worked at DCJS), this was the standard for misusing the criminal history search database.

If there is a standard for just retaining active search data for less than 24 hours, it does present a potential simple check that should be flagged – if a specific plate or specific camera is searched twice within 2 days, it should be flagged to review more closely. Flock does have their own system to identify suspicious search history.

The bigger issue to me though is who is doing the reviewing. It does not make sense to put this on vendors, and PDs just have not seriously devoted resources to this, even in response to public criticism. This audit mechanism should be delegated to a third party, either a specific group in the state attorney general’s office, or a state criminal justice agency (like DCJS in New York).

So that of course needs to be explicitly set by state statute as well. Who is doing the auditing?

My focus so far has been on abuses via police departments themselves, but smart regulation should also specify auditing of the vendors themselves, as well as punishments if they fail to meet data standards. (I am not thinking so much TEMPEST attacks here, but more so “I left an unauthenticated endpoint willy nilly on the internet”.)

Indeed, many of the requirements I am suggesting are likely already on the books; the problem is that the entity responsible for auditing is often unspecified or lacks the resources to do the work. (Also it is often unclear what the punishments are for failing to abide by statutes. That also needs to be specifically stated.)

The Future

So while I hope (although I have no expectation) that my blog post can somehow influence current standards across the country, I think it is important to keep in mind surveillance not just as the world exists now, but how it may look in the foreseeable future.

I think states should just pull the band aid off and create statutes that require a warrant to search the historical ALPR data. (And this makes data sharing between agencies mostly moot, the real time searches only need to be done within your own jurisdiction.) Like I said at the beginning, the current case law on being able to reconstruct the whole of a person’s movements (which I think is quite reasonable) will eventually be met if the ALPR cameras become dense enough. So states can either create the statutes to dictate that a warrant is necessary themselves, or eventually have the court system thrust it upon them.

In a world filled with privately owned cameras in public spaces, I think these suggestions are still relevant. So similar to Carpenter for CSLI data, and Chatrie for geofence warrants, there should just be warrant standards for historically searching any surveillance footage. There need be no special distinction between ALPR data (public or private) or video cameras.

Even if the groups calling for the banning of Flock cameras get their way, this does not stop private owners from collecting the data. So banning Flock, by itself, does not prevent abuse of searching private cameras. Again I think it is better to just let the government retain the data (same as private vendors will retain the data), and have consistent warrant standards for police to obtain that historical data.

This, of course, is a burden to detectives. I believe that trade-off in protecting our personal liberties while still allowing police effective means to investigate cases is a reasonable one.


  1. There are some technicalities between whether just collecting the data is a search (which was the scenario in the Norfolk case) or whether doing an active search (e.g. an officer querying the system for license plate ABC1234). The Norfolk case was the former, but for this post I am focusing on officers actually searching the data (the latter scenario).↩︎

Notes on Valuing the Cost of Crime

AI disclosure – I used AI to write this blog post. I figure having an AI blog post is better than not writing it at all. I will always disclose though if I use AI to heavily write any content on this blog. (I use it for minor copy editing all the time.)

For the tech details, I used gemini flash 3.5 with medium reasoning in the Antigravity IDE, using the same advice I said in this blog post. (Minor preference to Claude Code for writing blog posts for those who care.) It is the outline of the thread I did on X (which I wrote entirely by hand). Using this approach, e.g. I give a detailed outline and prior examples, Pangram says this is only lightly AI assisted.

Notes on Valuing the Cost of Crime

We often hear eye-popping figures about the “cost of crime.” For example, that a single aggravated assault costs society $100,000, or that a statistical life is worth $10 million. But if you look under the hood of these estimates, they are built on a house of cards: Willingness-to-Pay (WTP) surveys.

WTP estimates wildly inflate the costs of crime. For realistic policy decisions and police budgeting, we should be using concrete measures that are easier to calculate and verify.

The Three Buckets of Crime Costs

To evaluate criminal justice interventions, we can break costs into three broad categories:

  • A) Cost to the individual: Personal hospital bills, lost work, and physical trauma.
  • B) Cost to public sector agencies: Police labor, court proceedings, jail/prison operations, and public healthcare programs like Medicaid.
  • C) Cost to society: Reduced business activity in high-crime areas and the loss of workers to the economy.

Most cost-of-crime estimates do not calculate these countable categories. Instead, they use survey estimates of willingness-to-pay to approximate the costs of crime to individuals. I believe WTP estimates themselves are junk and should not be used to guide operations.

The Scaling Problem of Willingness-to-Pay

If you have heard the phrase “a statistical life costs $10 million,” you are seeing a WTP estimate in action.

The scaling math is straightforward, but the resulting estimates themselves are junk. Researchers ask survey respondents questions like: “Would you pay $100 in increased taxes to fund sidewalk improvements that reduce pedestrian fatalities?” If the safety measures are estimated to reduce pedestrian deaths by 1 in 100,000 annually in a city, the math scales up simply:

100 × 100, 000 = $10, 000, 000

People are thus deemed “willing to pay” $10 million to reduce one death.

This methodology yields massive, noisy estimates. You can see these WTP metrics compiled on the RAND Cost of Crime site. The primary limitation is that survey respondents will agree to pay almost any seemingly small amount when they do not actually have to pay it. In one street lighting survey I reviewed, participants were paid $1 to participate and claimed they were willing to pay $200 on average for better streetlights. It is highly doubtful that someone who sells their time for $1 to complete a survey will actually pay $200 in taxes for streetlights. As Andrew Gelman has pointed out, valuing lives based on ability to pay reveals how detached these hypothetical exercises are from real-world resource constraints.

Countable Costs vs. Theoretical Valuations

When we rely on concrete cost estimates that can be verified—such as labor hours and medical bills—the figures are much lower.

For instance, while a WTP estimate for an aggravated assault is close to $100,000, Priscilla Hunt’s study on law enforcement costs estimates the actual police labor cost for an assault is closer to $10,000.

I cannot prove what people are hypothetically willing to pay. But I can show a police chief that reducing ten assaults in a specific sector will save $100,000 in labor and overtime.

This distinction matters for other public costs too. Serious physical assaults can easily generate six-figure medical bills. In New York, more than 70% of gun violence hospitalizations are paid for via Medicaid. While it is reasonable for state or federal governments to weigh these medical costs, a local county or police department does not bear them. It makes no sense for a local police department to justify its budget by claiming it is reducing Medicaid expenses.

Example Cost-Benefit Case Studies

When we restrict our analysis to tangible costs, how do common interventions stack up?

Hotspots Policing

Because crime is highly concentrated, we can identify specific geographic areas that generate massive public costs. I have previously written about locating Million-Dollar Hotspots in Baltimore and Dallas. In my research on redrawing hotspots, I show how spatial concentration makes 24/7 hotspots policing cost-effective based purely on offsetting tangible labor costs.

For code examples of this, check out my crimepy python library (DBSCAN with weights for cost of crime estimates).

ShotSpotter

I am much less bullish on acoustic gunshot detection systems like ShotSpotter due to their high cost, as detailed in my ShotSpotter cost-benefit analysis. I estimate that ShotSpotter saves approximately 1 life for every 100 shooting victims it covers by dispatching emergency services faster. If you value a life at $10 million using WTP, the system easily looks cost-effective. If you use tangible costs, the math changes. ShotSpotter has not shown consistent evidence that it increases case clearances or prevents victimization. In fact, saving a shooting victim via faster response generates higher medical bills than if they had died, highlighting the complex economics of reactive vs. proactive interventions.

Business Improvement Districts (BIDs)

A great example of societal cost-shifting is Business Improvement Districts (BIDs). As shown in John MacDonald and colleagues’ study on BIDs in Los Angeles, BIDs demonstrate that commercial businesses are actually willing to spend their own money to improve safety in their areas through private security, cleaning services, and physical improvements. This is not hypothetical willingness-to-pay; it is a real-world, out-of-pocket expenditure by local merchants who calculate that reducing crime is directly worth their private investment.

Gun Violence Interventions (READI)

When looking at community-based interventions, the cost-benefit models face a different hurdle. Monica Bhatt and her colleagues evaluated Chicago’s READI program in their study on predicting and preventing gun violence. They claim a massive benefit of around $180,000 per participant (translating to a 3:1 benefit-cost ratio).

However, this estimated benefit of $180,000 is derived by mixing up WTP estimates and lifetime projections of individual offending (specifically, the Cohen & Piquero lifecourse model). As I discussed in my analysis of limits on gun violence interventions, extrapolating high-risk youth crime savings over an entire lifecourse using inflated WTP values creates a benefit estimate that is completely detached from the immediate budget realities of local governments.

The Missing Metric: The Value of an Arrest

This brings us to a major gap in criminology: we do not have good estimates for what it is worth to clear a crime.

Because crime is highly concentrated among a small number of chronic offenders, an arrest is often worth more than preventing a single crime. Apprehending a chronic offender can prevent dozens of future offenses.

This is why tools like automated License Plate Readers (LPR) are interesting. As Ozer’s study on LPR effectiveness shows, they are much cheaper than ShotSpotter and are highly cost-effective even if they only generate a small percentage increase in arrests. However, to truly calculate their ROI, we need a better grasp on the actual monetary value of a clearance.

To build better policy, we need to stop relying on WTP surveys and start measuring the real, tangible savings that police departments and local governments can actually bank.

References

  • Bhatt, M. P., Heller, S. B., et al. (2024). Predicting and preventing gun violence: An experimental evaluation of READI Chicago. The Quarterly Journal of Economics, 139(1), 1-56.

  • Cohen, M. A., & Piquero, A. R. (2009). New evidence on the monetary value of saving a high risk youth. Journal of Quantitative Criminology, 25(1), 25-49.

  • Hunt, P., Saunders, J., & Kilmer, B. (2019). Estimates of law enforcement costs by crime type for benefit-cost analyses. Journal of Benefit-Cost Analysis, 10(1), 95-123.

  • MacDonald, J., Golinelli, D., Stokes, R. J., & Bluthenthal, R. (2010). The effect of business improvement districts on the incidence of violent crimes. Injury Prevention, 16(5), 327-332.

  • Ozer, M. (2016). The impact of automatic number plate recognition (ANPR) technology on crime. Police Journal, 89(2), 117-132.

  • Wheeler, A. P., & Reuter, S. (2021). Redrawing Hot Spots of Crime in Dallas, Texas. Police Quarterly, 24(2), 159-184.

How long to conduct your experiment: Talk at ASEBP

Upcoming at the American Society of Evidence Based Policing Conference, I have a talk Thursday morning (9:45-10:00), How long to conduct your experiment.

The talk goes over some of the simple metrics I have created to help plan how long to conduct your intervention. Such as how long to evaluate your hot spots intervention, or purchase to increase arrest rates, etc.

I have prepared a ton of different resources. The main one is a web-based application (a WASM-based app with R as the backend) where you can enter your inputs and generate a graph showing how precise your parameter estimates are:

The help page includes citations and additional materials, but here is a brief rundown:

  • I have the math details in this github repo, see the methodology.pdf. It also includes notes on how I used different LLM tools to produce the webpage and the method materials. Each of the applications allows you to download the R code used to generate the graphs and tables.

  • I have created a series of YouTube videos demonstrating the application (WDD, IRR, Proportion tests)

  • I have posted my slides for the ASEBP talk

See you all in DC at ASEBP in a few weeks!

Gathering interest in tech courses

Quick post this morning — I have a survey up gathering input on interest in short, technical courses.

Think 2-3 days, potentially in person/synchronous.

If you have taken a course with Paul Allison at Horizon’s, or an ICPSR summer course, those are similar examples. But, the main difference will be these courses are to prepare you for pursuing private sector roles.

These will be aimed at:

  • grad level social science students
  • current professors looking to pursue private sector roles
  • current data analysts looking to get into data science
  • undergrads with some more technical background

Survey lists potential courses (python for data analysis, intro to LLM APIs, SQL + Dashboards, using agent based tools for analysis), the course medium (in person vs video), price points.

If you are a university or organization interested in hosting such sessions for your students, let me know as well. Happy to chat to you about bringing this to your campus.

Part time product design positions to help with AI companies

Recently on the Crime Analysis sub-reddit an individual posted about working with an AI product company developing a tool for detectives or investigators.

The Mercor platform has many opportunities that may be of interest to my network, so I am sharing them here. These include not only for investigators, but GIS analysts, writers, community health workers, etc. (The eligibility interviewers I think if you had any job in gov services would likely qualify, it is just reviewing questions.)

All are part time (minimum of 15 hours per week), remote, and can be in the US, Canada, or UK. (But cannot support H1-B or OPT visas in the US).

Additional for professionals looking to get into the tech job market, see these two resources:

I actually just hired my first employee at Crime De-Coder. Always feel free to reach out if you think you would be a good fit for the types of applications I am working on (python, GIS, crime analysis experience). I will put you in the list to reach out to when new opportunities are available.


Detectives and Criminal Investigators

Referral Link

$65-$115 hourly

Mercor is recruiting Detectives and Criminal Investigators to work on a research project for one of the world’s top AI companies. This project involves using your professional experience to design questions related to your occupation as a Detective and Criminal Investigator. Applicants must:

  • Have 4+ years full-time work experience in this occupation;
  • Be based in the US, UK, or Canada
  • minimum of 15 hours per week

Community Health Workers

Referral Link

$60-$80 hourly

Mercor is recruiting Community Health Workers to work on a research project for one of the world’s top AI companies. This project involves using your professional experience to design questions related to your occupation as a Community Health Worker. Applicants must:

  • Have 4+ years full-time work experience in this occupation;
  • Be based in the US, UK, or Canada
  • minimum 15 hours per week

Writers and Authors

Referral Link

$60-$95 hourly

Mercor is recruiting Writers and Authors to work on a research project for one of the world’s top AI companies. This project involves using your professional experience to design questions related to your occupation as a Writer and Author.

Applicants must:

  • Have 4+ years full-time work experience in this occupation;
  • Be based in the US, UK, or Canada
  • minimum 15 hours per week

Eligibility Interviewers, Government Programs

Referral Link

$60-$80 hourly

Mercor is recruiting Eligibility Interviewers, Government Programs to work on a research project for one of the world’s top AI companies. This project involves using your professional experience to design questions related to your occupation as a Eligibility Interviewers, Government Program. Applicants must:

  • Have 4+ years full-time work experience in this occupation;
  • Be based in the US, UK, or Canada
  • minimum 15 hours per week

Cartographers and Photogrammetrists

Referral Link

$60-$105 hourly

Mercor is recruiting Cartographers and Photogrammetrists to work on a research project for one of the world’s top AI companies. This project involves using your professional experience to design questions related to your occupation as a Cartographer and Photogrammetrist. Applicants must:

  • Have 4+ years full-time work experience in this occupation;
  • Be based in the US, UK, or Canada
  • minimum 15 hours per week

Geoscientists, Except Hydrologists and Geographers

$85-$100 hourly

Referral Link

Mercor is recruiting Geoscientists, Except Hydrologists and Geographers to work on a research project for one of the world’s top AI companies. This project involves using your professional experience to design questions related to your occupation as a Geoscientists, Except Hydrologists and Geographers Applicants must:

  • Have 4+ years full-time work experience in this occupation;
  • Be based in the US, UK, or Canada
  • minimum of 15 hours per week

Advice for crime analyst to break into data science

I recently received a question about a crime analyst looking to break into data science. Figured it would be a good topic for my advice in a blog post. I have written many resources over the years targeting recent PhDs, but the advice for crime analysts is not all that different. You need to pick up some programming, and likely some more advanced tech skills.

For background, the individual had SQL + Excel skills (which many analysts may just have Excel). Vast majority of analyst roles, you should be quite adept at SQL. But just SQL is not sufficient for even an entry level data science role.


For entry data science, you will need to demonstrate competency in at least one programming language. The majority of positions will want you to have python skills. (I wrote an entry level python book exactly for someone in your position.)

You likely will also need to demonstrate competency in some machine learning or using large language models for data science roles. It used to be Andrew Ng’s courses were the best recommendation (I see he has a spin off DeepLearningAI now). So that is second hand though, I have not personally taken them. LLMs are more popular now, so prioritizing learning how to call those APIs, build RAG systems, prompt engineering I think is going to make you slightly more marketable than traditional machine learning.

I have personally never hired anyone in a data science role without a masters. That said, I would not have a problem if you had a good portfolio. (Nice website, Github contributions, etc.)

You should likely start just looking and applying to “analyst” roles now. Don’t worry about if they ask for programming you do not have experience in, just apply. Many roles the posting is clearly wrong or totally unrealistic expectations.

Larger companies, analyst roles can have a better career ladder, so you may just decide to stay in that role. If not, can continue additional learning opportunities to pursue a data science career.

Remote is more difficult than in person, but I would start by identifying companies that are crime analysis adjacent (Lexis Nexis, ESRI, Axon) and start applying to current open analyst positions.

For additional resources I have written over the years:

The alt-ac newsletter has various programming and job search tips. THe 2023 blog post goes through different positions (if you want, it may be easier to break into project management than data science, you have a good background to get senior analyst positions though), and the 2025 blog post goes over how to have a portfolio of work.

Cover page, data science for crime analysis with python

I translated my book for $7 using openai

The other day an officer from the French Gendarmerie commented that they use my python for crime analysis book. I asked that individual, and he stated they all speak English. But given my book is written in plain text markdown and compiled using Quarto, it is not that difficult to pipe the text through a tool to translate it to other languages. (Knowing that epubs under the hood are just html, it would not suprise me if there is some epub reader that can use google translate.)

So you can see now I have available in the Crime De-Coder store four new books:

ebook versions are normally $39.99, and print is $49.99 (both available worldwide). For the next few weeks, can use promo code translate25 (until 11/15/2025) to purchase epub versions for $19.99.

If you want to see a preview of the books first two chapters, here are the PDFs:

And here I added a page on my crimede-coder site with testimonials.

As the title says, this in the end cost (less than) $7 to convert to French (and ditto to convert to Spanish).

Here is code demo’ing the conversion. It uses OpenAI’s GPT-5 model, but likely smaller and cheaper models would work just fine if you did not want to fork out $7. It ended up being a quite simple afternoon project (parsing the markdown ended up being the bigger pain).

So the markdown for the book in plain text looks like this:

It ends up that because markdown uses line breaks to denote different sections, that ends up being a fairly natural break to do the translation. These GenAI tools cannot repeat back very long sequences, but a paragraph is a good length. Long enough to have additional context, but short enough for the machine to not go off the rails when trying to just return the text you input. Then I just have extra logic to not parse code sections (that start/end with three backticks). I don’t even bother to parse out the other sections (like LaTeX or HTML), and I just include in the prompt to not modify those.

So I just read in the quarto document, split by “”, then feed in the text sections into OpenAI. I did not test this very much, just use the current default gpt-5 model with medium reasoning. (It is quite possible a non-reasoning smaller model will do just as well. I suspect the open models will do fine.)

You will ultimately still want someone to spot check the results, and then do some light edits. For example, here is the French version when I am talking about running code in the REPL, first in English:

Running in the REPL

Now, we are going to run an interactive python session, sometimes people call this the REPL, read-eval-print-loop. Simply type python in the command prompt and hit enter. You will then be greeted with this screen, and you will be inside of a python session.

And then in French:

Exécution dans le REPL

Maintenant, nous allons lancer une session Python interactive, que certains appellent le REPL, boucle lire-évaluer-afficher. Tapez simplement python dans l’invite de commande et appuyez sur Entrée. Vous verrez alors cet écran et vous serez dans une session Python.

So the acronym is carried forward, but the description of the acronym is not. (And I went and edited that for the versions on my website.) But look at this section in the intro talking about GIS:

There are situations when paid for tools are appropriate as well. Statistical programs like SPSS and SAS do not store their entire dataset in memory, so can be very convenient for some large data tasks. ESRI’s GIS (Geographic Information System) tools can be more convenient for specific mapping tasks (such as calculating network distances or geocoding) than many of the open source solutions. (And ESRI’s tools you can automate by using python code as well, so it is not mutually exclusive.) But that being said, I can leverage python for nearly 100% of my day to day tasks. This is especially important for public sector crime analysts, as you may not have a budget to purchase closed source programs. Python is 100% free and open source.

And here in French:

Il existe également des situations où les outils payants sont appropriés. Les logiciels statistiques comme SPSS et SAS ne stockent pas l’intégralité de leur jeu de données en mémoire, ils peuvent donc être très pratiques pour certaines tâches impliquant de grands volumes de données. Les outils SIG d’ESRI (Système d’information géographique) peuvent être plus pratiques que de nombreuses solutions open source pour des tâches cartographiques spécifiques (comme le calcul des distances sur un réseau ou le géocodage). (Et les outils d’ESRI peuvent également être automatisés à l’aide de code Python, ce qui n’est pas mutuellement exclusif.) Cela dit, je peux m’appuyer sur Python pour près de 100 % de mes tâches quotidiennes. C’est particulièrement important pour les analystes de la criminalité du secteur public, car vous n’avez peut‑être pas de budget pour acheter des logiciels propriétaires. Python est 100 % gratuit et open source.

So it translated GIS to SIG in French (Système d’information géographique). Which seems quite reasonable to me.

I paid an individual to review the Spanish translation (if any readers are interested to give me a quote for the French version copy-edits, would appreciate it). She stated it is overall very readable, but just has many minor things. Here is a a sample of suggestions:

Total number of edits she suggested were 77 (out of 310 pages).

If you are interested in another language just let me know. I am not sure about translation for the Asian languages, but I imagine it works OK out of the box for most languages that are derivative of Latin. Another benefit of self-publishing, I can just have the French version available now, but if I am able to find someone to help with the copy-edits I will just update the draft after I get their feedback.

The difference between models, drive-time vs fatality edition

Easily one of the most common critiques I make when reviewing peer reviewed papers is the concept, the difference between statistically significant and not statistically significant is not itself statistically significant (Gelman & Stern, 2006).

If you cannot parse that sentence, the idea is simple to illustrate. Imagine you have two models:

Model     Coef  (SE)  p-value
  A        0.5  0.2     0.01
  B        0.3  0.2     0.13

So often social scientists will say “well, the effect in model B is different” and then post-hoc make up some reason why the effect in Model B is different than Model A. This is a waste of time, as comparing the effects directly, they are quite similar. We have an estimate of their difference (assuming 0 covariance between the effects), as

Effect difference = 0.5 - 0.3 = 0.2
SE of effect difference = sqrt(0.2^2 + 0.2^2) = 0.28

So when you compare the models directly (which is probably what you want to do when you are describing comparisons between your work and prior work), this is a bit of a nothing burger. It does not matter that Model B is not statistically significant, a coefficient of 0.3 is totally consistent with the prior work given the standard errors of both models.

Reminded again about this concept, as Arredondo et al. (2025) do a replication of my paper with Gio on drive time fatalities and driving distance (Circo & Wheeler, 2021). They find that distance (whether Euclidean or drive time) is not statistically significant in their models. Here is the abstract:

Gunshot fatality rates vary considerably between cities with Baltimore, Maryland experiencing the highest rate in the U.S.. Previous research suggests that proximity to trauma care influences such survival rates. Using binomial logistic regression models, we assessed whether proximity to trauma centers impacted the survivability of gunshot wound victims in Baltimore for the years 2015-2019, considering three types of distance measurements: Euclidean, driving distance, and driving time. Distance to a hospital was not found to be statistically associated with survivability, regardless of measure. These results reinforce previous findings on Baltimore’s anomalous gunshot survivability and indicate broader social forces’ influence on outcomes.

This ends up being a clear example of the error I describe above. To make it simple, here is a comparison between their effects and the effects in my and Gio’s paper (in the format Coef (SE)):

Paper     Euclid          Network       Drive Time
Philly     0.042 (0.021)  0.030 (0.016)    0.022 (0.010)
Baltimore  0.034 (0.022)  0.032 (0.020)    0.013 (0.006)

At least for these coefficients, there is literally nothing anomalous at all compared to the work me and Gio did in Philadelphia.

To translate these coefficients to something meaningful, Gio and I estimate marginal effects – basically a reduction of 2 minutes results in a decrease of 1 percentage point in the probability of death. So if you compare someone who is shot 10 minutes from the hospital and has a 20% chance of death, if you could wave a wand and get them to the ER 2 minutes faster, we would guess their probability of death goes down to 19%. Tiny, but over many such cases makes a difference.

I went through some power analysis simulations in the past for a paper comparing longer drive time distances as well (Sierra-Arévalo et al. 2022). So the (very minor) differences could also be due to omitted variable bias (in logit models, even if not confounded with the other X, can bias towards 0). The Baltimore paper does not include where a person was shot, which was easily the most important factor in my research for the Philly work.

To wrap up – we as researchers cannot really change broader social forces (nor can we likely change the location of level 1 emergency rooms). What we can change however are different methods to get gun shot victims to the ER faster. These include things like scoop-and-run (Winter et al., 2022), or even gun shot detection tech to get people to scenes faster (Piza et al., 2023).

References

Using Esri + python: arcpy notes

I shared a series of posts this week using Esri + arcpy tools on my Crime De-Coder LinkedIn page. LinkedIn eventually removes the posts though, so I am putting those same tips here on the blog. Esri’s tools do not have great coverage online, so blogging is a way to get more coverage in those LLM tools long term.


A little arcpy tip, if you import a toolbox, it can be somewhat confusing what the names of the methods are available. So for example, if importing some of the tools Chris Delaney has created for law enforcement data management, you can get the original methods available for arcpy, and then see the additional methods after importing the toolbox:

import arcpy
d1 = dir(arcpy) # original methods
arcpy.AddToolbox("C:\LawEnforcementDataManagement.atbx")
d2 = dir(arcpy) # updated methods available after AddToolbox
set(d2) - set(d1) # These are the new methods
# This prints out for me
# {'ConvertTimeField_Defaultatbx', 'toolbox_code', 'TransformCallData_Defaultatbx', 'Defaultatbx', 'TransformCrimeData_Defaultatbx'}
# To call the tool then
arcpy.TransformCrimeData_Defaultatbx(...)

Many of the Arc tools have the ability to copy python code, when I use Chris’s tool it copy-pastes arcpy.Defaultatbx.TransformCrimeData, but if running from a standalone script outside of an Esri session (using the python environment that ArcPro installs) that isn’t quite the right code to call the function.

You can check out Chris’s webinar that goes over the law enforcement data management tool, and how it fits into the different crime analysis solutions that Chris and company at Esri have built.


I like using conda for python environments on Window’s machines, as it is easier to install some particular packages. So I mostly use:

conda create --name new_env python=3.11 pip
conda activate new_env
pip install -r requirements.txt

But for some libraries, like geopandas, I will have conda figure out the install. E.g.

conda create --name geo_env python=3.11 pip geopandas
conda activate geo_env
pip install -r requirements.txt

As they are particularly difficult to install with many restrictions.

And if you are using ESRI tools, and you want to install a library, conda is already installed and you can clone that environment.

conda create --clone "C:\Program Files\ArcGIS\Pro\bin\Python\envs\arcgispro-py3" --name proclone
conda activate proclone
pip install -r requirements.txt

As you do not want to modify the original ESRI environment.


Using conda to run scheduled jobs in Windows is alittle tricky. Here is an example of setting up a .bat file (which can be set up in Windows scheduler) to activate conda, set a new conda environment, and call a python script.

::: For log, showing date/time
echo:
echo --------------------------
echo %date% %time%
::: This sets the location of the script, as conda may change it
set "base=%cd%"
::: setting up conda in Windows, example Arc's conda activate
call "C:\Program Files\ArcGIS\Pro\bin\Python\Scripts\activate.bat"
::: activating a new environment
call conda activate proclone
::: running a python script
call cd %base%
call python auto_script.py
echo --------------------------
echo:

Then, when I set up the script in Window’s scheduler, I often have the log file at that level. So the task scheduler I will have the action as:

"script.bat" >> log.txt 2>&1

And have the options where the script runs from the location of script.bat. This will append both the normal log and error log to the shell script. So if something goes wrong, you can open log.txt and see what is up.


When working with arcpy, often you need to have tables inside of a geodatabase to use particular geoprocessing tools. Here is an example of taking an external csv file, and importing that file into a geodatabase as a table.

import arcpy
gdb = "./project/LEO_Tables.gdb"
tt = "TempTable"
arcpy.env.workspace = gdb

# Convert CSV into geodatabase
arcpy.TableToTable_conversion("YourData.csv",gdb,tt)
#arcpy.ListTables() # should show that new table

# convert time fields into text, useful for law enforcement management tools
time_fields = ['rep_date','begin','end']
for t in time_fields:
    new_field = f"{t}2"
    arcpy.management.AddField(tt,new_field,"TEXT")
    arcpy.management.CalculateField(tt,new_field,f"!{t}!.strftime('%Y/%m/%d %H:%m')", "PYTHON3")

# This will show the new fields
#fn = [f.name for f in arcpy.ListFields(tt)]

When you create a new project, it automatically creates a geodatabase file to go along with that project. If you just want a standalone geodatabase though, you can use something like this in your python script:

import arcpy
import os

gdb = "./project/LEO_Tables.gdb"

if os.path.exists(gdb):
    pass
else:
    loc, db = os.path.split(gdb)
    arcpy.management.CreateFileGDB(loc,db)

So if the geodatabase does not exist, it creates it. If it does exist though, it will not worry about creating a new one.


One of the examples for automation is taking a basemap, updating some of the elements, and then exporting that map to an image or PDF. This sample code, using Dallas data, shows how to set up a project to do this. And here is the original map:

Because ArgGIS has so many different elements, the arcpy module tends to be quite difficult to navigate. Basically I try to seperate out data processing (which often takes inputs and outputs them into a geodatabase) vs visual things on a map. So to do this project, you have step 1 import data into a geodatabase, and 2 update the map elements. Here legend, title, copying symbology, etc.

You can go to the github project to download all of the data (including the aprx project file, as well as the geodatabase file). But here is the code to review.

import arcpy
import pandas as pd
from arcgis.features import GeoAccessor, GeoSeriesAccessor
import os

# Set environment to a particular project
gdb = "DallasDB.gdb"
ct = "TempCrimes"
ol = "ExampleCrimes"
nc = "New Crimes"
arcpy.env.workspace = gdb
aprx = arcpy.mp.ArcGISProject("DallasExample.aprx")
dallas_map = aprx.listMaps('DallasMap')[0]
temp_layer = f"{gdb}/{ct}"

# Load in data, set as a spatial dataframe
df = pd.read_csv('DallasSample.csv') # for a real project, will prob query your RMS
df = df[['incidentnum','lon','lat']]
sdf = pd.DataFrame.spatial.from_xy(df,'lon','lat', sr=4326)

# Add the feature class to the map, note this does not like missing data
sdf.spatial.to_featureclass(location=temp_layer)
dallas_map.addDataFromPath(os.path.abspath(temp_layer)) # it wants the abs path for this

# Get the layers, copy symbology from old to new
new_layer = dallas_map.listLayers(ct)[0]
old_layer = dallas_map.listLayers(ol)[0]
old_layer.visible = False
new_layer.symbology = old_layer.symbology
new_layer.name = nc

# Add into the legend, moving to top
layout = aprx.listLayouts("DallasLayout")[0]
leg = layout.listElements("LEGEND_ELEMENT")[0]
item_di = {f.name:f for f in leg.items}
leg.moveItem(item_di['Dallas PD Divisions'], item_di[nc], move_position='BEFORE')

# Update title in layout "TitleText"
txt = layout.listElements("TEXT_ELEMENT")
txt_di = {f.name:f for f in txt}
txt_di['TitleText'].text = "New Title"
# If you need to make larger, can do
#txt_di['TitleText'].elementWidth = 2.0

# Export to high res PNG file
layout.exportToPNG("DallasUpdate.png",resolution=500)

# Cleaning up, to delete the file in geodatabase, need to remove from map
dallas_map.removeLayer(new_layer)
arcpy.management.Delete(ct)

And here is the updated map:

Some notes on ESRI server APIs

Just a few years ago, most cities open data sites were dominated by Socrata services. More recently though cities have turned to ArcGIS servers to disseminate not only GIS data, but also just plain tabular data. This post is to collate my notes on querying ESRI’s APIs for these services. They are quite fast, have very generous return limits, and have the ability to do filtering/aggregation.

So first lets start with Raleigh’s Open Data site, specifically the Police Incidents. So sometimes for data analysis you just want a point-in-time dataset, and can download 100% of the data (which you can do here, see the Download button in the below screenshot). But what I am going to show here is how to format queries to generate up to date information. This is useful in web-applications, like dashboards.

So first, go down to the Blue button in the below screen that says I want to use this:

Once you click that, you will see a screen that lists several different options, click to expand the View API Resources, and then click the link open in API explorer:

To save a few steps, here is the original link and the API link side by side, you can see you just need to change explore to api in the url:

https://data-ral.opendata.arcgis.com/datasets/ral::daily-raleigh-police-incidents/explore
https://data-ral.opendata.arcgis.com/datasets/ral::daily-raleigh-police-incidents/api

Now on this page, it has a form to be able to fill in a query, but first check out the Query URL string on the right:

I am going to go into how to modify that URL in a bit to return different slices of data. But first check out the link https://services.arcgis.com/v400IkDOw1ad7Yad/ArcGIS/rest/services

This simpler view I often find easier to see all the available data than the open data websites with the extra fluff. You can often tell the different data sources right from the name (and often cities have more things available than they show on their open data site). But lets go to the Police Incidents Feature Server page, the link is https://services.arcgis.com/v400IkDOw1ad7Yad/ArcGIS/rest/services/Daily_Police_Incidents/FeatureServer/0:

This gives you some meta-data (such as the fields and projection). Scroll down to the bottom of the page, and click the Query button, it will then take you to https://services.arcgis.com/v400IkDOw1ad7Yad/ArcGIS/rest/services/Daily_Police_Incidents/FeatureServer/0/query:

I find this tool to format queries easier than the Open Data site. Here I put in the Where field 1=1, set the Out Fields to *, the Result record count to 3. I then hit the Query (GET)

This gives an annoyingly long url. And here are the resulting images

So although this returns a very long url, most of the parameters in the url are empty. So you could have a more minimal url of https://services.arcgis.com/v400IkDOw1ad7Yad/ArcGIS/rest/services/Daily_Police_Incidents/FeatureServer/0/query?where=1%3D1&outFields=*&resultRecordCount=3&f=json. (There I changed the format to json as well.)

In python, it is easier to work with the json or geojson output. So here I show how to query the data, and read it into a geopandas dataframe.

from io import StringIO
import geopandas as gpd
import requests

base = "https://services.arcgis.com/v400IkDOw1ad7Yad/ArcGIS/rest/services/Daily_Police_Incidents/FeatureServer/0/query"
params = {"where": "1=1",
          "outFields": "*",
          "resultRecordCount": "3",
          "f": "geojson"}
res = requests.get(base,params)
gdf = gpd.read_file(StringIO(res.text)) # note I do not use res.json()

Now, the ESRI servers will not return a dataset that has 1,000,000 rows, it limits the outputs. I have a gnarly function I have built over the years to do the pagination, fall back to json if geojson is not available, etc. Left otherwise uncommented.

from datetime import datetime
import geopandas as gpd
import numpy as np
import pandas as pd
import requests
import time
from urllib.parse import quote

def query_esri(base='https://services.arcgis.com/v400IkDOw1ad7Yad/arcgis/rest/services/Police_Incidents/FeatureServer/0/query',
               params={'outFields':"*",'where':"1=1"},
               verbose=False,
               limitSize=None,
               gpd_query=False,
               sleep=1):
    if verbose:
        print(f'Starting Queries @ {datetime.now()}')
    req = requests
    p2 = params.copy()
    # try geojson first, if fails use normal json
    if 'f' in p2:
        p2_orig_f = p2['f']
    else:
        p2_orig_f = 'geojson'
    p2['f'] = 'geojson'
    fin_url = base + "?"
    amp = ""
    fi = 0
    for key,val in p2.items():
        fin_url += amp + key + "=" + quote(val)
        amp = "&"
    # First, getting the total count
    count_url = fin_url + "&returnCountOnly=true"
    if verbose:
        print(count_url)
    response_count = req.get(count_url)
    # If error, try using json instead of geojson
    if 'error' in response_count.json():
        if verbose:
            print('geojson query failed, going to json')
        p2['f'] = 'json'
        fin_url = fin_url.replace('geojson','json')
        count_url = fin_url + "&returnCountOnly=true"
        response_count2 = req.get(count_url)
        count_n = response_count2.json()['count']
    else:
        try:
            count_n = response_count.json()["properties"]["count"]
        except:
            count_n = response_count.json()['count']
    if verbose:
        print(f'Total count to query is {count_n}')
    # Getting initial query
    if p2_orig_f != 'geojson':
        fin_url = fin_url.replace('geojson',p2_orig_f)
    dat_li = []
    if limitSize:
        fin_url_limit = fin_url + f"&resultRecordCount={limitSize}"
    else:
        fin_url_limit = fin_url
    if gpd_query:
        full_response = gpd.read_file(fin_url_limit)
        dat = full_response
    else:
        full_response = req.get(fin_url_limit)
        dat = gpd.read_file(StringIO(full_response.text))
    # If too big, getting subsequent chunks
    chunk = dat.shape[0]
    if chunk == count_n:
        d2 = dat
    else:
        if verbose:
            print(f'The max chunk size is {chunk:,}, total rows are {count_n:,}')
            print(f'Need to do {np.ceil(count_n/chunk):,.0f} total queries')
        offset = chunk
        dat_li = [dat]
        remaining = count_n - chunk
        while remaining > 0:
            if verbose:
                print(f'Remaining {remaining}, Offset {offset}')
            offset_val = f"&cacheHint=true&resultOffset={offset}&resultRecordCount={chunk}"
            off_url = fin_url + offset_val
            if gpd_query:
                part_response = gpd.read_file(off_url)
                dat_li.append(part_response.copy())
            else:
                part_response = req.get(off_url)
                dat_li.append(gpd.read_file(StringIO(part_response.text)))
            offset += chunk
            remaining -= chunk
            time.sleep(sleep)
        d2 = pd.concat(dat_li,ignore_index=True)
    if verbose:
        print(f'Finished queries @ {datetime.now()}')
    # checking to make sure numbers are correct
    if d2.shape[0] != count_n:
        print('Warning! Total count {count_n} is different than queried count {d2.shape[0]}')
    # if geojson, just return
    if p2['f'] == 'geojson':
        return d2
    # if json, can drop geometry column
    elif p2['f'] == 'json':
        if 'geometry' in list(d2):
            return d2.drop(columns='geometry')
        else:
            return d2

And so, to get the entire dataset of crime data in Raleigh, it is then df = query_esri(verbose=True). It is pretty large, so I show here limiting the query.

params = {'where': "reported_date >= CAST('1/1/2025' AS DATE)", 
          'outFields': '*'}
df = query_esri(base=base,params=params,verbose=True)

Here this shows doing a datetime comparison, by casting the input to a date. Sometimes you have to do the opposite, cast one of the text fields to dates or extract out values from a date field represented as text.

Example Queries

So I showed about you can do a WHERE clause in the queries. You can do other stuff as well, such as get aggregate counts. For example, here is a query that shows how to get aggregate statistics.

If you click the link, it will go to the query form ESRI webpage. And that form shows how to enter in the output statistics fields.

And this produces counts of the total crimes in the database.

Here are a few additional examples I have saved in my notes:

Do not use the query_esri function above for aggregate counts, just form the params and pass them into requests directly. The query_esri function is meant to return large sets of individual rows, and so can overwrite the params in unexpected way.

Check out my Crime De-Coder LinkedIn page this week for other examples of using python + ESRI. This is more for public data, but those will be examples of using arcpy in different production scenarios. Later this week I will also post an updated blog here, for the LLMs to consume.