Digital Evidence Collection and Analysis by Law Enforcement is moving faster than the Law

[
[
[

]
]
]

AI Analysis of FBI Search Photographs

My last post considered what happens when photographs taken during the execution of a search warrant are stored in a system that a substantial portion of the FBI workforce can reach. The risk I described there was a human one. One curious employee, one case file opened for a reason that has nothing to do with the case.

That risk has a natural ceiling. An employee must know which file to open and open each case one at a time. Artificial intelligence removes both constraints. A tool that reads the contents of every photograph in every case file, on a continuous basis, and renders what it finds searchable does not require anyone to know where to look. It only requires someone to know what to look for.

The Bureau is building those tools now, and the legal framework that would govern their use was written for a filing cabinet.

The Bureau has always indexed

This is not a new impulse. J. Edgar Hoover came to the Justice Department from the Library of Congress, and he brought its cataloging methods with him. The practice he built — FBI employees identifying information of interest within a case so that other employees can find and use it later — is the ancestor of everything I am about to describe. It is still doctrine. The Domestic Investigations and Operations Guide has governed for almost 20 years.

The version posted on the FBI’s own site is currently not viewable. Use an older archived copy hosted by Just Security. Sections 5.1 and 18.5.2 are the ones that matter here: FBI employees may search FBI systems for information both before and during an assessment as well as during an investigation. An assessment is the predicate-free step that precedes a predicated investigation — the stage at which the Bureau decides whether there is enough to open a case at all.

Read those two propositions together. Employees may query FBI holdings before any investigation exists, and the Bureau has spent a century making its holdings query able. The only thing that has ever limited the scope of that query is the granularity of the index. Indexing has always been a human act, performed one document at a time, by someone with finite attention and a caseload.

That is the bottleneck artificial intelligence removes.

What is already in the inventory

Under Executive Order 13960, federal agencies must publish an inventory of their non-classified, non-sensitive AI use cases. DOJ’s 2025 inventory lists roughly fifty FBI entries. That number is the minimum — the classified and sensitive use cases are, by definition, absent, and the inventory gives no indication of how many of those there are.

Four of the listed capabilities are enough to make the point:

Optical character recognition — converts text appearing in an image into searchable text.

Facial recognition technology — matches faces appearing in an image against known identities.

License plate reading — extracts plate numbers from images of vehicles.

Sentiment analysis — classifies language by the attitude it expresses.

None of these is exotic. All four are commodity technology, and all four operate on exactly the kind of material a search photograph contains: a wall, a desk, a bookshelf, a bumper sticker, a printed page left face-up.

Applied to Sentinel

Suppose the Bureau runs OCR and facial recognition across every search photograph it holds and writes the output back into Sentinel as metadata. This is not a technically demanding project. The result is a system in which the contents of every photograph taken inside every searched residence in the country are queryable by string.

Consider what that permits. A quotation attributed to former Director Comey is characterized by someone at the Bureau as a threat against the President. The phrase is now a search term, and the search returns every person who happened to have printed it and left it visible during an unrelated search. A photograph of Che Guevara on a dormitory wall becomes an identifiable attribute. Neither search requires opening a case file. Neither requires a predicate. Both are, on the face of the DIOG, the kind of systems query an employee may run during an assessment.

I do not think this is hypothetical in the way it would have been ten years ago. Director Patel has asked what the purpose of collecting terabytes of data is if you cannot sift through it. In the same video, he has also said that every major American technology company is working to improve the Bureau’s systems. Both statements are, on their own terms, reasonable. The FBI’s technical infrastructure genuinely needs the work, and I agree with the underlying premise: information lawfully collected should be capable of being reviewed.

My disagreement is narrower. The question is not whether the Bureau should be able to sift. It is whether the authority to sift is bounded by the warrant that authorized the collection, or by the capability of the tool doing the sifting. At present, nothing in the DIOG answers that question, because the DIOG was written for a world in which the answer was supplied by the physical limits of human review.

Who draws the line

I want to be precise about where the failure sits, because it is easy to locate it in the wrong place.

An employee who runs a query against a system the Bureau built, using a tool the Bureau deployed, in a manner the DIOG appears to permit, is doing the job as it has been configured for them. If the query is too broad, that is because no one drew the line before the tool shipped. Agents do not write the retention schedule, choose the analytics, or set the access controls in Sentinel. Treating this as a matter of individual conduct both misdescribes the problem and guarantees it will not be fixed, because the next employee inherits the same system.

It is worth correcting a related assumption, though not because I think anyone should act on it. Practitioners sometimes suppose that individual liability supplies the backstop here. It does not. 42 U.S.C. § 1983 reaches persons acting under color of state law and has no application to federal agents. The federal analogue is Bivens v. Six Unknown Named Agents, and the Supreme Court has spent two decades narrowing it — after Ziglar v. Abbasi and Egbert v. Boule, a damages remedy in a new context is very nearly unavailable. The Court’s stated reason is that Congress is better positioned to weigh the costs of creating one. That may well be right, and I intend to take up in a later post what a legislative answer to any of this would look like.

But Congress has not spoken, and the capability is being built now. Absent legislation, the work falls to the two institutions already in the room.

The first is the Bureau. Internal policy is a weaker instrument than a statute — it is revocable by the people it constrains, amendable without notice, and the operative version of the DIOG is not reliably available on the FBI’s own website. Those are real limitations, and I do not want to pretend otherwise. They are also not a reason for the Bureau to wait. The FBI is the only party that can draw a line before the tool ships rather than years after, and it is the party that knows what the tool actually does. If automated content review of warrant-derived material requires its own predicate and its own approval, someone inside the Bureau has to write that down, and it should be written down before the first query runs, not after the first suppression hearing.

The second is the courts. A rule the regulated party writes for itself is a limit only if someone outside can test whether it was followed. That testing does not happen in the abstract; it happens when a defendant asks how a particular photograph in a particular case came to be located, and a court requires the government to answer. Judicial review is what converts an internal policy from an aspiration into a constraint, and it is the only mechanism currently operating on any of this.

So the two arguments run together rather than against each other. The Bureau should build the limits because no one else can build them in time. The courts have to review whether the limits were honored, because otherwise the Bureau is grading itself.

For counsel litigating a case, then, the question is not what any employee did. It is what the evidence is, how it was derived, and whether its derivation stayed within the warrant — asked of the government, on the record, in a forum that can rule on the answer.

What to look for

Every stage of this pipeline — collection, analysis, use, and disposition — creates a discrete question that the government should have to answer on the record:

Was the photograph within the scope of the warrant when it was taken? Was the automated analysis performed at the time of collection, or years later in connection with a different matter? What was the predicate for the query that surfaced it? Was the derived metadata — the OCR text, the facial match, the sentiment score — disclosed, and if not, why is it not Brady or Rule 16 material? And if the search photograph in the current case was located through a content query rather than a case number, what authorized that query?

These are not rhetorical. They are discovery requests, and I expect the answers to be poor for some time — not because anyone is hiding them, but because the systems are being built faster than the policy governing them is being written. The record each of these questions creates is also the most likely route to getting that policy written.The Bureau has been indexing the contents of its files since Hoover. The difference is that the index is about to become complete, retroactive, and continuous, and no one has yet said out loud who is allowed to query it, or on what showing. Until Congress does, the Bureau should say it, and the courts should check the answer. I will keep working through the Sentinel pieces of this in subsequent posts.

One response

  1. alexanderbopp71 Avatar
    alexanderbopp71

    Another fascinating post. Now I have to figure out how to follow you.

    Like

Leave a comment