License Plate Reader Searches Should Require a Warrant

So while I work with police departments regularly, I think it is critically important that technology be used reasonably.

While this may be off-putting to some of my clients, I worked with the Institute for Justice as an expert witness in their trial Schmidt v City of Norfolk. (Any opinions herein are my own and not those of IJ, to be clear.) The gist of that case was whether searches of historically cached ALPR data (automated-license-plate-reader) constituted an illegal search.1

The judge ruled against plaintiffs in that case. Here is a quote from the judgment:

Consistent with Plaintiffs’ claims in this case and controlling precedent involving mass surveillance in public spaces, ALPR surveillance could become too intrusive and run afoul of [constitutional privacy standards] at some point. But when? While a definitive answer to that question is elusive, what is readily apparent to this Court is that, at least in Norfolk, Virginia, the answer is: not today.

The important point to note about this quote is “not today”. This will be a long winded post, but to try to keep it simple:

  • I think cameras will become ubiquitous in the foreseeable future. So the question is not if this data will require a warrant, it is when. It is going to happen eventually under current case law.
  • I think cameras are good, and can be used to reduce crime in a cost effective manner.
  • There is a difference between active flags (e.g. this car is stolen and it pings the PD when it drives past a camera) vs historical searches (e.g. look to see where license plate XYZ1000 was the last 30 days).
  • Requiring a warrant for historical searches will not seriously impede police investigations.
  • The current status quo of not retaining data is VERY BAD; it does not prevent illegal searches, and currently limits the utility of actually using that data for legitimate investigations.
  • Current standards to prevent abuse of the searching ALPR data systems are laughable.

Long story short in my opinion everyone would be better off if states just mandated warrant procedures through state statutes.

To try to not get too much into the weeds of what historically constitutes a search, I think the easiest place to start is via Carpenter vs US. So current US case law requires police departments to obtain a warrant to request cellular providers provide law enforcement with cell phone tower pings (cell-site location information, CSLI).

This deviated from historical precedent in requiring a warrant mainly because it was private companies that had the information. Before Carpenter, mostly it was argued you did not have a reasonable expectation of privacy if a private company could access the same data. The court in Carpenter basically made a determination that cell phone data was so comprehensive it justified a different standard – that you could track the whole of a person’s movements with the detailed CSLI data. And that this level of invasiveness violated a reasonable person’s expectation of privacy. Even if Google had all that info, you did not expect them to give it away.

This opinion was reaffirmed with the recent Chatrie decision (for geofence warrants, e.g. give me a ping for all cell phones in area X and datetime-range Y). Another relevant decision to be aware of is also Beautiful Struggle v Baltimore, in which searching historical aerial imagery via drones also constituted a search.

So this is why I am saying the question is when, not if, ALPR data will require a warrant. If a city happened to have a camera on literally every intersection (which I think will happen in the future), under current case law it would clearly be the same situation as you have for your cell phone data.

Cameras are Good

To be brief, again I mostly work with police departments in my career and was a former crime analyst. I do think ALPR cameras are good investments, mainly because they are cheap enough to have a reasonable return on investment. (Note I do not think this about all police tech, I am particularly critical of the price tag for acoustic-gun-shot-detection.)

So ALPRs are well under $3,000 per camera. The machine learning models, camera, and computation necessary to flag a plate when it passes can easily fit on current cell phones. (The harder part is powering the phone and protecting it from the elements.) ALPRs for the most part just take static images and then extract out the license plate (and for some vendors extract out additional information, like car make and color).

The overall evidence that ALPRs reduce crime is pretty meh at the moment (see my slides at a Wake Libertarian talk I did in 2024), but because they are so cheap they really only need to increase a few arrests per camera to likely have a positive return on investment.

It is pretty hand-wavy, as we do not have estimates for the value of increased clearances I find persuasive. But I think saying “I would pay $500 to help solve one case” is on the low side if anything. So a single camera if it helps catch just a handful of crimes a year is likely in my opinion to be a positive ROI.

I think cameras in all public spaces are going to happen. Imagine Ring comes out with a nicer camera system for homeowners that has more comprehensive views around your house and is just as cheap. And we will ultimately be safer for it. So even for folks advocating that cities do not pay for Flock, this is coming anyway in the near future.

Historical Searches vs Active Flags

ALPRs have been around a long time. The first ones I worked with at Troy, NY when I was an analyst were in-car cameras. Basically a go pro attached to the window that alerted when an officer drove by a stolen plate.

While ALPRs initial use was always pitched as this active flagging of stolen vehicles, they were used right away to retroactively search the historical locations of plates. They had a log of every plate, lat/lon, and timestamp of when that car passed a camera.

So imagine you are conducting an investigation of Joe Schmo, you know his license plate, and then you can type in his plate and see where his car passed a camera. Based on this information, same as CSLI data, you can basically trace where Joe went, where he repeatedly visited, where he likely slept, etc. (The first time I used this at Troy, we figured out a particular individual we were actively investigating was living with his girlfriend for example. I was honestly amazed how densely filled in the map was of hits for a single plate based on the in-car cameras.)

You technically do not need to cache any data at all to accomplish this “flag a stolen vehicle” (or any other scenario where you are actively looking for a specific license plate). There are legitimate scenarios though where ALPR searches for recent data in a real time context can be very helpful.

One of the more common examples – someone robs a gas station, and they drove a vehicle. You don’t know the plate, but can look at the images that passed by the fixed location ALPRs in the time range, and then especially if you have a car description from the gas station attendant can figure out the plate associated with the vehicle.

To be clear I am not a lawyer, but in my opinion I think exigent circumstances make searching a few minutes of cached ALPR location data totally reasonable. In practice, New Hampshire’s 3 minute data retention is far too short. I could see arguments for several hours (imagine “I found a dead body on the side of the road”, that requires more time for it to be reported.) But we are meandering into the territory where it is not an active emergency “need to find someone who may have a gun and hurt people” that would justify those exigent circumstances. Those are the scenarios where getting a warrant is reasonable (no different than a geofence warrant if you do not have a plate and want to just search what cars passed by a camera within a certain date-time window, or no different than a CSLI warrant if you have an active suspect and want to search for a specific license plate).

Most states are retaining ALPR data for longer periods. While the Norfolk case was ongoing, Virginia set a standard across the state at 21 days. Before that it was up to the individual agency. It varies state by state, but states often mandate data retention around 30 days, or leave it up to the discretion of the police department.

Deleting Data does not prevent abuses

These data retention statutes are argued as a mechanism to prevent abuse. They do not accomplish this.

If you look through the cases in which officers abused the system to search for individuals, all of them searched for specific plates over-and-over again, sometimes hundreds of times.

If you retain data for 20 days, you can just go and do a search every 20 days, keep notes on the data as you so wish, and then do another search 20 days later. Getting rid of old data, in-and-of-itself, does nothing to prevent that abuse. In fact if someone is actively stalking a person, you would expect them to regularly do searches, seeing where their victim is going on a regular basis while they have access to the system.

Simultaneously, deleting data does prevent its legitimate use in long term law enforcement investigations. It is totally normal for a murder investigation to take more than 30 days to identify a suspect. Gosh, sure would be nice to be able to then query the ALPR data to show whether a person was in the vicinity of the murder. Simultaneously it could be used by the defense for exculpatory purposes (which assuredly would take longer than 30 days).

So folks advocating for deleting data as a mechanism to prevent abuse are making things worse. It does not prevent abuse, and limits the utility of ALPR for historical investigations. The only way data retention by itself prevents abuse is if you do not cache data at all (like in New Hampshire), and only use ALPRs for the active alert situation.

What Smart Regulation Looks Like

One of the reasons I say that the current standards to prevent abuse are laughable is that data retention policies and internal PD policies on when the data should be searched have been in place in most departments for years (if not a decade) at this point. The examples where searching ALPR data to stalk an intimate partner were obviously not prevented via data retention policies.

Alas, my suggestion that some data is cached for real time investigations (longer than 3 minutes), and that a warrant should be required outside of this window, does not prevent that type of abuse either. Most departments have in place reasons why a search can be conducted, and some states have specific statutes identifying impermissible reasons for conducting searches. In the Norfolk IJ case, officers, when entering a reason for a search (which was often omitted), sometimes supplied reasons that appeared prima facie illegal, such as “protest”.

Departments, even if they have a standard to do internal audits, often do not follow them. It took Tyler Dukes asking Raleigh PD for their audit results for them to even conduct their first audit.

This is a long standing problem for PDs, not just with ALPRs, but also with searching criminal history illegally. IJ collating a dozen cases of arrests of ALPR misuse across the country is not evidence these systems are working, as it is likely the case that only the most egregious abuses are ever caught.

In addition to creating state statutes to mandate that a warrant be used for historical ALPR searches, states should, at a minimum, have clear punishments for illegal searches. These should include at a minimum losing your job, and being banned from accessing the system forever. When I was a crime analyst in New York (and ditto for when I worked at DCJS), this was the standard for misusing the criminal history search database.

If there is a standard for just retaining active search data for less than 24 hours, it does present a potential simple check that should be flagged – if a specific plate or specific camera is searched twice within 2 days, it should be flagged to review more closely. Flock does have their own system to identify suspicious search history.

The bigger issue to me though is who is doing the reviewing. It does not make sense to put this on vendors, and PDs just have not seriously devoted resources to this, even in response to public criticism. This audit mechanism should be delegated to a third party, either a specific group in the state attorney general’s office, or a state criminal justice agency (like DCJS in New York).

So that of course needs to be explicitly set by state statute as well. Who is doing the auditing?

My focus so far has been on abuses via police departments themselves, but smart regulation should also specify auditing of the vendors themselves, as well as punishments if they fail to meet data standards. (I am not thinking so much TEMPEST attacks here, but more so “I left an unauthenticated endpoint willy nilly on the internet”.)

Indeed, many of the requirements I am suggesting are likely already on the books; the problem is that the entity responsible for auditing is often unspecified or lacks the resources to do the work. (Also it is often unclear what the punishments are for failing to abide by statutes. That also needs to be specifically stated.)

The Future

So while I hope (although I have no expectation) that my blog post can somehow influence current standards across the country, I think it is important to keep in mind surveillance not just as the world exists now, but how it may look in the foreseeable future.

I think states should just pull the band aid off and create statutes that require a warrant to search the historical ALPR data. (And this makes data sharing between agencies mostly moot, the real time searches only need to be done within your own jurisdiction.) Like I said at the beginning, the current case law on being able to reconstruct the whole of a person’s movements (which I think is quite reasonable) will eventually be met if the ALPR cameras become dense enough. So states can either create the statutes to dictate that a warrant is necessary themselves, or eventually have the court system thrust it upon them.

In a world filled with privately owned cameras in public spaces, I think these suggestions are still relevant. So similar to Carpenter for CSLI data, and Chatrie for geofence warrants, there should just be warrant standards for historically searching any surveillance footage. There need be no special distinction between ALPR data (public or private) or video cameras.

Even if the groups calling for the banning of Flock cameras get their way, this does not stop private owners from collecting the data. So banning Flock, by itself, does not prevent abuse of searching private cameras. Again I think it is better to just let the government retain the data (same as private vendors will retain the data), and have consistent warrant standards for police to obtain that historical data.

This, of course, is a burden to detectives. I believe that trade-off in protecting our personal liberties while still allowing police effective means to investigate cases is a reasonable one.


  1. There are some technicalities between whether just collecting the data is a search (which was the scenario in the Norfolk case) or whether doing an active search (e.g. an officer querying the system for license plate ABC1234). The Norfolk case was the former, but for this post I am focusing on officers actually searching the data (the latter scenario).↩︎

VerusCite: checking academic articles for hallucinations

I have a new app out, VerusCite. With the recent rise in popularity of GenAI tools like ChatGPT and Claude, this has also come along with academics writing slop articles.

One of the ways to check that slop is via looking at the articles citations. LLMs have some predictable failure modes in writing papers whole cloth – they tend to get details like complicated author lists wrong, or swap out incorrect journal titles. VerusCite is a tool for editors and reviewers to use to verify citations in a fast and cheap application.

It costs $2 to review a paper (and you get two free reviews on sign-up). If you want to see the output of a single example though, check out https://veruscite-data.com/share/E1-rn3TBwksENO_3hlmsRD0IG-EnXd6vfQg-upQZWQM

In addition to hallucinations, I have made many parts of the application just useful to editors in general. Many papers have minor errors in their bibliographies; typos, years off, author swaps, bad URLs, etc. Here is an example – no hallucinations that signal poor writing, but has seven different errors in the bibliography.

This is par for the course (it is quite possible 5% of citations have errors that look like this). The website has convenient tools to edit citations and export the fixed citations (whether minor errors or gross hallucinations) in various formats.

This makes much lighter work of the tedious job of formatting and checking citations for editors. One of the ways I think is critical to build generative AI tools is to consider the human in the loop from the start. My tool will ultimately make some errors (I error rate estimates in my public benchmark). I want it to be as fast for a human to confirm (or refute) the LLM label.

If you are an editor or a reviewer, I highly suggest you check the application out. Peer review journals (and pre-print servers that review the applications before posting), will need to use a tool like this as a first pass to ensure slop is not being posted.

AI writing is better than no writing

AI disclosure – this post was entirely written by myself.

I know AI writing is still pretty cringey – so I get that people are quite opposed to it. For people like me though (academics promoting their work, more technical oriented) I would like to proffer a slight defense of (even cringey) AI writing. Having an LLM tool help you write a blog post is better than not writing at all.

I have come to the personal opinion I just want you to disclose when you use AI. I am starting to get peer review requests for academic papers that are clearly LLM written, and they are not obviously worse than the typical (mostly horrid) way academics write papers (they may actually be better to be honest). Blog and social media posts I think are strictly worse to my personal tastes when using LLM writing (across many dimensions, for now anyway). But it is better to write something than nothing if you have something worth saying.

Where this matters for technical folks (and academics) is that your default SEO is awful. Most academic papers are behind paywalls. LLM research tools are not picking up peer reviewed papers. So if you have something worth saying, having LLMs write out a blog post for you is worth it relative to having no writing at all.

For examples of LLM writing I have on this site:

And then my book, Large Language Models for Mortals: A Practical Guide for Analysts with Python, is around 50% AI generated.

None of these examples I would have finished without the help of AI; either entirely writing for the example blog posts, or writing the first draft in the case of the LLM book. (The LLM book is good by the way, you would not be able to tell I generated that first draft at all with Claude.)

My suggestion is to not let AI entirely take the wheel, but to create a detailed outline and have the LLM review your prior writing. Those two things improve posts by a wide margin (in addition to making sure AI is not too verbose – keep those blog posts simple!). And then you still need to take the time to review your own writing (for references you need to check those for hallucinations).

To be clear again, AI writing is better than nothing if you have something actually useful to say to the world. The bigger issue with AI writing are slop merchants just wasting space. That happened before with LLM tools, it is just much easier and more prevalent now. Just own it when you use AI to help you write.

Notes on Valuing the Cost of Crime

AI disclosure – I used AI to write this blog post. I figure having an AI blog post is better than not writing it at all. I will always disclose though if I use AI to heavily write any content on this blog. (I use it for minor copy editing all the time.)

For the tech details, I used gemini flash 3.5 with medium reasoning in the Antigravity IDE, using the same advice I said in this blog post. (Minor preference to Claude Code for writing blog posts for those who care.) It is the outline of the thread I did on X (which I wrote entirely by hand). Using this approach, e.g. I give a detailed outline and prior examples, Pangram says this is only lightly AI assisted.

Notes on Valuing the Cost of Crime

We often hear eye-popping figures about the “cost of crime.” For example, that a single aggravated assault costs society $100,000, or that a statistical life is worth $10 million. But if you look under the hood of these estimates, they are built on a house of cards: Willingness-to-Pay (WTP) surveys.

WTP estimates wildly inflate the costs of crime. For realistic policy decisions and police budgeting, we should be using concrete measures that are easier to calculate and verify.

The Three Buckets of Crime Costs

To evaluate criminal justice interventions, we can break costs into three broad categories:

  • A) Cost to the individual: Personal hospital bills, lost work, and physical trauma.
  • B) Cost to public sector agencies: Police labor, court proceedings, jail/prison operations, and public healthcare programs like Medicaid.
  • C) Cost to society: Reduced business activity in high-crime areas and the loss of workers to the economy.

Most cost-of-crime estimates do not calculate these countable categories. Instead, they use survey estimates of willingness-to-pay to approximate the costs of crime to individuals. I believe WTP estimates themselves are junk and should not be used to guide operations.

The Scaling Problem of Willingness-to-Pay

If you have heard the phrase “a statistical life costs $10 million,” you are seeing a WTP estimate in action.

The scaling math is straightforward, but the resulting estimates themselves are junk. Researchers ask survey respondents questions like: “Would you pay $100 in increased taxes to fund sidewalk improvements that reduce pedestrian fatalities?” If the safety measures are estimated to reduce pedestrian deaths by 1 in 100,000 annually in a city, the math scales up simply:

100 × 100, 000 = $10, 000, 000

People are thus deemed “willing to pay” $10 million to reduce one death.

This methodology yields massive, noisy estimates. You can see these WTP metrics compiled on the RAND Cost of Crime site. The primary limitation is that survey respondents will agree to pay almost any seemingly small amount when they do not actually have to pay it. In one street lighting survey I reviewed, participants were paid $1 to participate and claimed they were willing to pay $200 on average for better streetlights. It is highly doubtful that someone who sells their time for $1 to complete a survey will actually pay $200 in taxes for streetlights. As Andrew Gelman has pointed out, valuing lives based on ability to pay reveals how detached these hypothetical exercises are from real-world resource constraints.

Countable Costs vs. Theoretical Valuations

When we rely on concrete cost estimates that can be verified—such as labor hours and medical bills—the figures are much lower.

For instance, while a WTP estimate for an aggravated assault is close to $100,000, Priscilla Hunt’s study on law enforcement costs estimates the actual police labor cost for an assault is closer to $10,000.

I cannot prove what people are hypothetically willing to pay. But I can show a police chief that reducing ten assaults in a specific sector will save $100,000 in labor and overtime.

This distinction matters for other public costs too. Serious physical assaults can easily generate six-figure medical bills. In New York, more than 70% of gun violence hospitalizations are paid for via Medicaid. While it is reasonable for state or federal governments to weigh these medical costs, a local county or police department does not bear them. It makes no sense for a local police department to justify its budget by claiming it is reducing Medicaid expenses.

Example Cost-Benefit Case Studies

When we restrict our analysis to tangible costs, how do common interventions stack up?

Hotspots Policing

Because crime is highly concentrated, we can identify specific geographic areas that generate massive public costs. I have previously written about locating Million-Dollar Hotspots in Baltimore and Dallas. In my research on redrawing hotspots, I show how spatial concentration makes 24/7 hotspots policing cost-effective based purely on offsetting tangible labor costs.

For code examples of this, check out my crimepy python library (DBSCAN with weights for cost of crime estimates).

ShotSpotter

I am much less bullish on acoustic gunshot detection systems like ShotSpotter due to their high cost, as detailed in my ShotSpotter cost-benefit analysis. I estimate that ShotSpotter saves approximately 1 life for every 100 shooting victims it covers by dispatching emergency services faster. If you value a life at $10 million using WTP, the system easily looks cost-effective. If you use tangible costs, the math changes. ShotSpotter has not shown consistent evidence that it increases case clearances or prevents victimization. In fact, saving a shooting victim via faster response generates higher medical bills than if they had died, highlighting the complex economics of reactive vs. proactive interventions.

Business Improvement Districts (BIDs)

A great example of societal cost-shifting is Business Improvement Districts (BIDs). As shown in John MacDonald and colleagues’ study on BIDs in Los Angeles, BIDs demonstrate that commercial businesses are actually willing to spend their own money to improve safety in their areas through private security, cleaning services, and physical improvements. This is not hypothetical willingness-to-pay; it is a real-world, out-of-pocket expenditure by local merchants who calculate that reducing crime is directly worth their private investment.

Gun Violence Interventions (READI)

When looking at community-based interventions, the cost-benefit models face a different hurdle. Monica Bhatt and her colleagues evaluated Chicago’s READI program in their study on predicting and preventing gun violence. They claim a massive benefit of around $180,000 per participant (translating to a 3:1 benefit-cost ratio).

However, this estimated benefit of $180,000 is derived by mixing up WTP estimates and lifetime projections of individual offending (specifically, the Cohen & Piquero lifecourse model). As I discussed in my analysis of limits on gun violence interventions, extrapolating high-risk youth crime savings over an entire lifecourse using inflated WTP values creates a benefit estimate that is completely detached from the immediate budget realities of local governments.

The Missing Metric: The Value of an Arrest

This brings us to a major gap in criminology: we do not have good estimates for what it is worth to clear a crime.

Because crime is highly concentrated among a small number of chronic offenders, an arrest is often worth more than preventing a single crime. Apprehending a chronic offender can prevent dozens of future offenses.

This is why tools like automated License Plate Readers (LPR) are interesting. As Ozer’s study on LPR effectiveness shows, they are much cheaper than ShotSpotter and are highly cost-effective even if they only generate a small percentage increase in arrests. However, to truly calculate their ROI, we need a better grasp on the actual monetary value of a clearance.

To build better policy, we need to stop relying on WTP surveys and start measuring the real, tangible savings that police departments and local governments can actually bank.

References

  • Bhatt, M. P., Heller, S. B., et al. (2024). Predicting and preventing gun violence: An experimental evaluation of READI Chicago. The Quarterly Journal of Economics, 139(1), 1-56.

  • Cohen, M. A., & Piquero, A. R. (2009). New evidence on the monetary value of saving a high risk youth. Journal of Quantitative Criminology, 25(1), 25-49.

  • Hunt, P., Saunders, J., & Kilmer, B. (2019). Estimates of law enforcement costs by crime type for benefit-cost analyses. Journal of Benefit-Cost Analysis, 10(1), 95-123.

  • MacDonald, J., Golinelli, D., Stokes, R. J., & Bluthenthal, R. (2010). The effect of business improvement districts on the incidence of violent crimes. Injury Prevention, 16(5), 327-332.

  • Ozer, M. (2016). The impact of automatic number plate recognition (ANPR) technology on crime. Police Journal, 89(2), 117-132.

  • Wheeler, A. P., & Reuter, S. (2021). Redrawing Hot Spots of Crime in Dallas, Texas. Police Quarterly, 24(2), 159-184.

Gathering interest in tech courses

Quick post this morning — I have a survey up gathering input on interest in short, technical courses.

Think 2-3 days, potentially in person/synchronous.

If you have taken a course with Paul Allison at Horizon’s, or an ICPSR summer course, those are similar examples. But, the main difference will be these courses are to prepare you for pursuing private sector roles.

These will be aimed at:

  • grad level social science students
  • current professors looking to pursue private sector roles
  • current data analysts looking to get into data science
  • undergrads with some more technical background

Survey lists potential courses (python for data analysis, intro to LLM APIs, SQL + Dashboards, using agent based tools for analysis), the course medium (in person vs video), price points.

If you are a university or organization interested in hosting such sessions for your students, let me know as well. Happy to chat to you about bringing this to your campus.

Job Advice Resources page

Minor update, I have created a page, Job Advice Resources to cumulatively list all the materials I have written on advice for social scientists and crime analysts looking to pivot into private sector tech roles.

I still get maybe ~2 folks a month ask for advice, and I am always happy to chat. I wish PhD granting institutions took this more seriously (it only takes minor changes to better prepare students).

If you are an administrator of a PhD program and actually care about getting your students jobs, also feel free to reach out and I am happy to discuss how I can help.

The race to the bottom with AI tools

What we are seeing in the AI startup space is a perfect example of the “no moat” problem: if your core product is essentially just clever prompt engineering wrapped around someone else’s frontier model, it is trivially easy for a competitor to reverse-engineer your workflow and undercut your price. Over the last few months, this lack of a defensible moat has triggered a rapid race to the bottom in automated peer review, moving from expensive managed services to open-source “bring your own key” (BYOK) scripts.

Here I am going to look at three tools specifically designed to review academic papers: Refine, IsItCredible, and Coarse.

Overview of the Tools

Refine: Refine positions itself as a premium, rigorous option for institutions, boasting testimonials from Ivy League professors and a high price point of $49.99 per review. It uses what it calls “massive parallel compute” to make hundreds of LLM calls to stress-test every line of a document.

IsItCredible: Built on the open-source Reviewer 2 pipeline, IsItCredible offers a standardized, pay-per-use middle ground with core reports starting at $5. It employs a clever “adversarial” architecture where “Red Team” agents try to find flaws and a “Blue Team” verifies them to prevent hallucinations.

Coarse: Coarse represents the logical endpoint of this race as an open-source “Bring Your Own Key” (BYOK) tool that lets you run complex multi-agent reviews locally or via OpenRouter. Because users pay the API costs directly instead of a markup, a comprehensive paper review is significantly cheaper.

The “LLM as a Judge” Problem

The hardest part of all this is evaluation. How do you know if the AI reviewer is actually good?

Refine relies almost entirely on anecdotal evidence. Their own FAQ essentially tells you to just try it and see the difference for yourself, claiming that general-purpose chatbots cannot match their depth even with expert prompting. This “try it yourself” approach is effective for marketing, but it isn’t a hard benchmark.

IsItCredible and Coarse are trying to be more systematic. The IsItCredible team released a paper, Yell at It: Prompt Engineering for Automated Peer Review, where they benchmarked their tool against five alternatives. They claim 15 wins out of 20 pairings. Similarly, Coarse claims to have been “blind-evaluated” against Refine and Reviewer 2, scoring higher on coverage and specificity.

However, we are still largely in the “LLM as a judge” era. These benchmarks often use another LLM to decide which review is better. It is circular logic. Until we have a “Ground Truth” dataset of known mathematical errors or logical fallacies in published papers, we are just measuring which AI writes the most convincing-sounding critique.

Because evaluation is so difficult, this software category risks becoming a classic market for lemons. It is incredibly difficult to identify substantive differences in quality between these tools without some external, hard benchmark. To truly evaluate if Refine’s expensive managed service is meaningfully better than Coarse’s open-source BYOK run, you have to verify the AI’s claims. But verifying those claims requires spending just as much time reading and reviewing the original paper as you would have spent just doing the review yourself from scratch. Without transparent benchmarks, users cannot easily distinguish high-quality rigorous analysis from convincing hallucinations, driving the market toward the cheapest option by default.

For those building AI tools, this entire space serves as a warning about the race to the bottom. I have previously written about deep research tools as another example of this phenomenon. If your only value proposition is a well-orchestrated prompt chain, open-source alternatives will inevitably compress your margins to zero. Eventually, the native GUI interfaces of the frontier models themselves may just become good enough that your specialized service isn’t even needed.

Meta

Did you like this post? Guess what, it was entirely generated via the Google’s API models (specifically the gemini cli). I have saved the chat session and log for how long it took here. You can see for yourself, I had a broad idea, asked it to review different materials, and then generate a post. I then iterated 25 minutes from start to finish in total.

The original post also is not flagged by Pangram as AI generated.

It definitely is not 100% my style (and to be clear this meta section is 100% hand written). The final paragraph about deep research tools I also struggled to get the model to say what I wanted – I wanted it to say “deep research tools are another example where this same situation will occur”. I am keeping the original 100% AI generated post for posterity though for folks to see what is possible with the current tools.

Policing Scholars should join ASEBP

Cross-posted on my Crime De-Coder blog.

I will be giving a talk at the upcoming American Society of Evidence Based Policing (ASEBP) conference (registration link here, May 20th-22nd in DC). My talk is How long to conduct your experiment? Check it out Thursday morning – I specifically asked for one of the short talks; 15 minutes is plenty to get the gist.

ASEBP Conference Flyer, 2026 in DC

I will be sharing a web-app to go with the talk soon (you can see my WDD tool and this blog post for background), but wanted to write a more general post about why researchers (as well as police officers who are interested in professionalization of the field) should join ASEBP.

To start, I have been involved in various ways with ASEBP for several years now, but I do not have any financial ties to ASEBP. I currently volunteer on the committee that reviews conference talks.

ASEBP is clearly the best organization for policing scholars currently in the country. The other main criminological societies (the American Society of Criminology and the Academy of Criminal Justice Sciences) are operating much as they did 30 years ago. Mostly they only exist to run journals and have a yearly conference where anyone can give a talk. They are incredibly insular, and have basically zero input from practitioners.

You can go and just look at the talks for ASC and ACJS – they are basically irrelevant to the vast majority of criminal justice operations (not only in policing, but in the CJ field as a whole). You can go look at the talks for the ASEBP conference and see they have a much clearer focus on realistic topics police departments are interested in, but presented by legitimate researchers and practitioners.

For scholars, I have developed working relationships with departments through multiple police practitioners I have met through ASEBP – and I hope to make more!

ASEBP was started by Renee Mitchell with a clear goal in mind – Renee is really the modern-day version of August Vollmer. ASEBP is intended to be a rigorous (unlike ASC, which allows almost anyone to present) conference and organization (ASEBP has training opportunities as well) to advance the use of evidence in policing operations.

If you think “I am not a policing researcher”, but have anything to do at all with criminal justice, feel free to get in touch. (Crime analysts should definitely join.) I have ideas to expand the organization – nothing equivalent currently exists in other parts of the criminal justice system as well. Being evidence-based is really the core of what Renee and everyone else is building.

If you are going to the conference and want to meet up, feel free to send me an email, andrew.wheeler@crimede-coder.com, and I will find a time to get a coffee while we are in DC.

Year in Review 2025 and AI Predictions

For a brief year in review, total views for the two different websites have decreased in the past year. For this blog, I am going to be a few thousand shy of 100,000 views. (2023 I had over 150k views, and 2024 I had over 140k views.) For the Crime De-Coder site, I am going to only get around 15k views.

Part of it is I posted less, this will be the 21st blog post this year on the personal blog (2023 had 46 and 2024 had 32 posts). The Crime De-Coder site had 12 blog posts, so pretty consistent with the prior year. Both are pretty bursty, with large bouts of traffic coming from if I post something to Hacker News I can get 1k to 10k views in a day or two if it makes it to the front page. So the 2024 stats for the crime de-coder was a few of those Hacker News bumps I did not get in 2025.

Some of it could legitimately be traditional Google search being usurped by the gen AI tools. This is the first year I had appreciable referrals from chatgpt, but they are less than 1000. The other tools are trivial amount of referrals. If I worried about SEO more, I would have more updating/regular content (as old pages are devalued quite a bit by google, and it seems to be getting more severe over time).

I have upped my use of the free tools quite a bit. ChatGPT knows me pretty well, and I use Claude Desktop almost every day as well.

An IAM policy scroll is more of a nightmare, and I definitely ask more python questions than R, but the cartoon desk is pretty close to spot on. I am close to paying for Anthropic subscription for Claude code credits (currently use pay as I go via Bedrock, and this is the first month I went over $20).

What pages on the blog are popular I can never be sure of. My most popular post last year was Downloading Police Employment Trends from the FBI Data Explorer. A 2023 post, that had random times where it would have several hundred visits in a short hour span. (Some bot collecting sites? I do not know.) If it is actual people, you would want to check out my Sworn Dashboard site, where you can look at trends for PDs much easier than downloading all the data yourself!

One thing that has grown though, I do short form posting on LinkedIn on my crime de-coder page. Impressions total for the year is over 340k (see the graph), and I currently am a few shy of 4400 followers.

LinkedIn is nice because it can be slightly longer form than the other social media sites. I would suggest you follow me there (in addition to signing up for RSS feeds for the two sites). That is the easiest way to follow my work.

I also took over as a moderator of the Crime Analysis Reddit forum, it is better than the IACA forums in my opinion, so encourage folks to post there for crime analysis questions.

Crime De-Coder Work

Crime De-Coder work has been steady (but not increasing). Similar to last year had several consulting gigs conducting crime analysis for premises liability cases (and one other case I may share my opinions once it is over), and doing some small projects with non-profits and police departments.

One big project was a python training in Austin.

The Python Book (which I also translated to Spanish/French), had a trickle of new sales. 2024 had around 100 sales and 2025 had around 50 sales. It is close to 2/3 print sales and 1/3 epub, so definately folks should have physical prints if you are selling books still.

Doing trainings basically makes writing the book worth it, but I do hope eventually the book makes it way into grad school curriculum’s. (Only one course so far.) I have pitched to grad schools to have me run a similar bootcamp to what I do for crime analysts, so if interested let me know.

The biggest new thing was Crime De-Coder got an Arnold Grant. Working with Denver PD on an experiment to evaluate a chronic offender initiative.

At the Day Gig

At my day gig, I was officially promoted to a senior manager and then quickly to a director position. Hence you get posts like what to show in your tech resume and notes on project management.

One of the reasons I am big on python – it is the dominant programming language in data science. It is hard for me to recruit from my network, as majority of individuals just know a little R (if you were a hard core R person, had packages/well executed public repo’s, I could more easily think you will be able to migrate to python to work on my team).

So learn python if you want to be a data scientist is my advice (and see other job market advice at my archived newsletter).

AI Predictions

At the day gig, my work went from 100% traditional supervised machine learning models to more like 50/50 traditional vs generative AI applications. The genAI hype is real, but I think it is worthwhile putting my thoughts to paper.

The biggest question is will AI take all of our jobs? I think a more likely end scenario is the AI tools just become better at helping humans do tasks. The leap from helping a human do something faster vs an AI tool doing it 100% on its own with 0 human input is hard. The models are getting incrementally better, but I think to fully replace people in a substantive way will require another big advancement in fundamental capabilities. Making a human 10x more productive is easier and still will make the AI companies a ton of money.

Sometimes people view the 10x idea and say that will take jobs, just not 100% of jobs. That is a view though that there is only a finite amount of work to be done. That assumption is clearly not true, and being able to do work faster/cheaper just induces demand for more potential work. The example with calculators making more banking jobs, not less, is basically the same example.

One of the critiques of the current systems is they are overvalued, so we are in a bubble. I do not remember where I read it, but one estimate was if everyone in the US spent $1 a day on the different AI tools, that would justify the current valuations for OpenAI, Anthropic, NVIDIA, etc. I think that is totally doable, we spend a few thousand a workday at Gainwell on the foundation models for example for a few projects, and we are just going to continue to roll out more and more. Gainwell is a company with around 6k employees for reference, and our current AI applications touch way less than 1k of those employees. We have plenty of room to grow those applications.

It is super hard though to build systems to help people do things faster. And we are talking like “this thing that used to take 30 minutes now takes 15 minutes”. If you have 100 people doing that thing all the time though, the costs of the models are low enough it is an easy win.

And this mostly only holds true for knowledge economy work that can be all done via software. There just still needs to be fundamental improvements to robotics to be able to do physical things. The tailor’s job is safe for the foreseeable future.

The change in the data science landscape to more generative AI applications definitely requires social scientists and analysts to up their game though to learn a new set of tools. I do have another book in the works to address that, so hopefully you will see that early next year.

What to show in your tech resume?

Jason Brinkley on LinkedIn the other day had a comment on the common look of resumes – I disagree with his point in part but it is worth a blog post to say why:

So first, when giving advice I try to be clear about what I think are just my idiosyncratic positions vs advice that I feel is likely to generalize. So when I say, you should apply to many positions, because your probability of landing a single position is small, that is quite general advice. But here, I have personal opinions about what I want to see in a resume, but I do not really know what others want to see. Resumes, when cold applying, probably have to go through at least two layers (HR/recruiter and the hiring manager), who each will need different things.

People who have different colored resumes, or in different formats (sometimes have a sidebar) I do not remember at all. I only care about the content. So what do I want to see in your resume? (I am interviewing for mostly data scientist positions.) I want to see some type of external verification you actually know how to code. Talk is cheap, it is easy to list “I know these 20 python libraries” or “I saved our company 1 million buckaroos”.

So things I personally like seeing in a resume are:

  • code on github that is not a homework assignment (it is OK if unfinished)
  • technical blog posts
  • your thesis! (or other papers you were first/solo author)

Very few people have these things, so if you do and you land in my stack, you are already at the like 95th percentile (if not higher) for resumes I review for jobs.

The reason having outside verification you actually know what you are doing is because people are liars. For our tech round, our first question is “write a python hello world program and execute it from the command line” – around half of the people we interview fail this test. These are all people who list they are experts in machine learning, large language models, years of experience in python, etc.

My resume is excessive, but I try to practice what I preach (HTML version, PDF version)

I added some color, but have had recruiters ask me to take it off the resume before. So how many people actually click all those links when I apply to positions? Probably few if any – but that is personally what I want to see.

There are really only two pieces of advice I have seen repeatedly about resumes that I think are reasonable, but it is advice not a hard rule:

  • I have had recruiters ask for specific libraries/technologies at the top of the resume
  • Many people want to hear about results for project experience, not “I used library X”

So while I dislike the glut of people listing 20 libraries, I understand it from the point of a recruiter – they have no clue, so are just trying to match the tech skills as best they can. (The matching at this stage I feel may be worse than random, in that liars are incentivized, hence my insistence on showing actual skills in some capacity.) It is infuriating when you have a recruiter not understand some idiosyncratic piece of tech is totally exchangeable with what you did, or that it is trivial to learn on the job given your prior experience, but that is not going to go away anytime soon.

I’d note at Gainwell we have no ATS or HR filtering like this (the only filtering is for geographic location and citizenship status). I actually would rather see technical blog posts or personal github code than saying “I saved the company 1 million dollars” in many circumstances, as that is just as likely to be embellished as the technical skills. Less technical hiring managers though it is probably a good idea to translate technical specs to more plain business implications though.