Victim-report triage
Seized-artefact enrichment
Local matching, no per-query calls
Cybercrime Units & Digital Investigation

A verified answer on a hostname, at the moment the complaint is still fresh

A cybercrime unit does not need another feed to monitor. It needs to know, quickly and defensibly, whether the address a complainant just handed over is a live credential-harvesting host — and it needs the answer without every query about an active investigation leaving the building. This page is about using a daily, DNS-verified list of more than 390,000 currently-resolving phishing domains as investigative lead material, and being precise about what that data is and is not.

Read on
Sub-50msResponse time on a single-domain check
100Domains per batch request, one lookup each
04:30 UTCDaily rebuild lands, with an add/remove changelog
Zero queriesFeed deployment keeps lookups inside your enclave
Home / Use Cases / Law Enforcement
The queue nobody sees

Most cybercrime work starts with a URL somebody pasted into a form

Before anything resembling an investigation begins, a detective has to decide which of forty reports this week describes the same infrastructure — and which describes nothing at all.

The public image of a cybercrime unit involves forensic workstations and warrant returns. The actual daily reality, in most departments below the federal level, is a shared mailbox and an intake form. Citizens report that they clicked something and lost money. Small businesses report that an invoice was paid to the wrong account. A parent forwards a text message their child received. Each report arrives with a fragment of evidence, and the most common fragment by a wide margin is a URL — copied out of a message, screenshotted at an angle, or half-remembered and reconstructed from browser history.

Triage, not attribution

What a detective needs from that fragment is three answers. Does this hostname resolve? Is it recognisable as part of something already known? Does it connect to any of the six other reports in the same queue? Those determine whether a report becomes a case, gets folded into an existing one, or is closed with advice and a referral — and they have to be answered in minutes, because a unit of three cannot spend an hour on each of forty weekly reports.

The evidence is disposable by design

A host is registered, used against a target set for a matter of days, and abandoned. By the time a report has been filed, assigned, read and actioned, the domain may already have stopped resolving — and once it does, most of what an investigator could have observed is simply gone. The page cannot be captured, the hosting cannot be identified from a live connection, and what remains is a string in a complaint form with nothing behind it.

What a dated match actually gives you

Inclusion requires an active A record confirmed through rotating proxy infrastructure, and entries that stop resolving fall out rather than accumulating, so a match is dated and specific: this hostname was confirmed answering, and confirmed as credential-harvesting infrastructure, as of the last_checked date in the response. Not proof of anything — a corroborating observation from an independent source with a stated methodology.

Triage is the bottleneck, not analysis. Units rarely lack the skill to investigate a case. They lack a fast, consistent way of deciding which forty reports contain one case and which contain forty dead ends.
Evidence decays on a clock you do not control. Every hour between a report landing and somebody looking at the hostname is an hour in which the host may go dark and take its own context with it.
Fitness for purpose

Six criteria an investigative data source has to meet

Threat intelligence is sold to security teams, whose standards for a claim are looser than an investigator's. These are the questions worth asking before the data touches a case file.

A stated inclusion test

You have to be able to describe, in one sentence, what had to be true for an entry to exist. Here it is an active A record, verified through rotating proxy infrastructure, on a host confirmed as phishing or malware distribution. That sentence is repeatable in a statement; "our proprietary model scored it high" is not.

Foundation criterion

A date attached to the observation

An undated assertion is worthless in an investigative context, because the whole question is what was true when the victim clicked. Every response carries last_checked, and the daily changelog records the day a domain entered the list and the day it left. Those two dates bracket the window in which the host was confirmed live.

Foundation criterion

No score you would have to explain

Confidence is 0.98 for a confirmed match and 0.0 for no match, with nothing in between. That is deliberate. A graduated score invites an analyst to pick a threshold, and a threshold is a discretionary judgement that defence counsel will quite reasonably want explained. A binary answer with a documented test is far easier to stand behind.

High weight

A deployment that does not leak the investigation

If checking a domain means sending it to a third party, then the third party learns which hostnames a police department is interested in, and when. The feed model removes that: the whole database is downloaded and matched locally, so the only outbound traffic is your own scheduled retrieval of the list.

Foundation criterion

An honest account of coverage

A source that implies completeness is a source that will embarrass you. This one does not: a clean result means the hostname is not on the list, which is a different statement from safe, and a domain registered and first used within the same hour will not be present. Knowing where the hole is makes the data usable.

High weight

Bulk behaviour that matches real evidence volumes

Artefact lists from a seized device are not ten domains long. The /batch endpoint takes up to 100 domains per POST and returns checked, phishing_found and credits_used, at 10 requests per second per key. That is thousands of artefacts enriched in the time it takes to write the covering note.

Operational weight
Three concrete uses

Corroborating a live host, linking reports, and enriching what was seized

The same dataset does three different jobs depending on where in the case lifecycle you reach for it.

The same dataset does three different jobs depending on where in the case lifecycle you reach for it — corroborating a single complaint, linking complaints nobody realised were connected, and reducing a device extraction to something a human can read.

Corroborating a single complaint

A complainant says they were directed to a particular address on a particular evening and entered their banking credentials. Two days later the address no longer answers, which is entirely normal and tells you nothing. A match carrying last_checked and a dns_status of resolves is an independent observation that the host was answering and assessed as credential-harvesting inside the relevant window — often the thing that decides whether a report is a genuine offence or a confused account of a failed transaction.

The wording that goes on the file

The correct formulation is that a dataset of verified phishing infrastructure records this hostname as having an active DNS record and a phishing classification as of a given date. The incorrect formulation is that the domain has been confirmed as the one used against the victim. The first accurately describes a third-party observation; the second attributes something the data cannot carry, and it is the sentence that will be taken apart later.

Linking reports into one case

A unit that runs every reported hostname through the same source starts to see structure invisible report by report. Four complaints in a fortnight, from unrelated victims, pointing at hosts that entered the database on the same day and left within days of each other, is a pattern worth pulling on. The daily changelog is what makes it visible, converting a static list into a timeline of when infrastructure appeared and when it went dark.

…with the scepticism that deserves

This is lead generation, not analysis. Domains appearing on the same day may be entirely unrelated, because a great deal of infrastructure is registered every day. What the observation does is tell an investigator which four of forty reports justify an hour of correlation against material they actually control — transaction records, recipient accounts, message headers. The same enrichment step appears on the threat intelligence and incident response pages.

Artefacts off a seized device

An extraction from a seized phone or laptop produces browser history, bookmarks, message contents, cached links and app data, and a domain-extraction pass typically yields thousands of unique hostnames. Almost all are ordinary. Reading them by hand is not a serious proposal, and individual lookups against a public reputation service are both slow and, in an investigative context, indiscreet.

What the bulk pass costs

Running the extracted set against a locally held copy is a scripted step measured in seconds, producing a short list worth a human's attention. Where a department prefers the API to holding a copy, /batch takes 100 domains per POST at one lookup each, putting a five-thousand-domain extraction at fifty requests and fifty credits. Either way the output is a filtered subset, with dates, that an analyst can actually read.

Never present a match as identification of an offender. The dataset makes a claim about a hostname's behaviour and DNS state. It makes no claim about who registered it, who operated it, or who sent the message that contained it.
Record what you asked and when. If a check informs an investigative decision, note the date of the query and the last_checked value returned. Reconstructing that six months later from memory is unpleasant and avoidable.
Evidentiary care

This is lead material. It is not forensic proof, and it should never be dressed as it

The distinction is not a disclaimer. It is the thing that determines where in a case the data belongs and how it should be described in writing.

A commercial dataset assembled by a third party, from observations made through infrastructure the investigator has never seen, with a methodology summarised rather than audited, is not evidence of an offence. It has no chain of custody in the sense a court uses the term, and nobody from the provider examined the specific message your complainant received. The value is real, but it is the value of a tip from a reliable source: it points you at things worth examining with tools whose output you can actually stand behind.

Intelligence side, not exhibits

Use it to prioritise, to link, to decide where to spend a preservation request, and to tell a victim whether what happened to them was what they think it was. Keep it out of the exhibits schedule.

Then get the weighty evidence

If the case proceeds, obtain it through the routes that carry weight: preservation and production from the hosting provider and registrar, records from the financial institutions, and your own contemporaneous capture of the live host where that was possible.

Your own copy is different

If a unit downloads the daily list, records the retrieval and can testify to what its copy contained on a given date, that is a record the department made and can speak to — which is more than can be said for a screenshot of somebody's web interface.

390,000+ live entries

Every entry has a confirmed active A record. Domains that stop resolving fall out rather than accumulating, so the list describes what is answering today rather than everything ever reported.

Dated, not undated

Responses carry last_checked; the feed ships a daily changelog of additions and removals, which is a dated record of change rather than a snapshot that overwrites itself. You get an entry date and an exit date rather than an unanchored claim.

Two verdicts, no middle

Confidence is 0.98 for a confirmed match and 0.0 for no match. There is no score to defend, no threshold to justify, and no analyst discretion embedded in the number.

A clean result must never drive inaction. A hostname not on the list is a hostname nobody has observed and verified yet. That is an absence of information, not a finding that the report is unfounded, and it should not appear on a file as one.
Watch the attribution language in anything disclosable. The response fields — is_phishing, a category of phishing/malware, dns_status, last_checked — describe a host. Nothing in them describes a person.
Deployment shape

Per-query lookups versus a local copy, in a CJIS-conscious environment

For most organisations this is a convenience question. For a police department it is a disclosure question, and the two options are genuinely different.

ConsiderationPer-query API lookupLocal feed copy
Who learns which hostnames you are examining The provider sees each domain you ask about, and when Nobody. Matching happens on your equipment
Outbound traffic from the enclaveOne HTTPS call per lookup, continuously One scheduled retrieval per day, nothing else
Works with no internet path at match time No — the verdict service must be reachable Yes, once the file is inside
Suits an air-gapped or tightly segmented network Not without an outbound exception you may not get Yes — a file can be carried across a boundary
Cost behaviour under a large extractionOne credit per domain; 5,000 artefacts is 5,000 credits Unlimited downloads; extraction size is irrelevant
Setup effort A single authenticated request; nothing to host A scheduled job, a file store and a staleness alert
Freshness at the moment of the check Whatever the current build holds As fresh as your last successful pull
Suits ad-hoc verification of one reported URL Yes — this is what it is forYes, but you are querying a local file rather than a service

Take the local feed when…

Your network is segmented for criminal justice information, outbound exceptions require a change board, or the sensitivity of which hostnames you are examining is itself a consideration. It is also the right answer for bulk artefact work, where per-domain billing and rate limits stop being background details. Delivery options including SFTP and S3 are on the daily feed page.

Use per-query lookups when…

You want a verified answer on one address that a member of the public just reported, from a workstation on the ordinary corporate network, without standing anything up. Sub-50ms responses and a 10-per-second limit per key make this comfortable for interactive use. Many units run both: the feed inside the enclave, the API on the general network for intake triage.

The statistics endpoint needs nothing. /stats reports database size and last update with no API key, no credits and no authentication, which makes it a reasonable target for a health check that has to sit outside any credentialled path.
The API key is the username you chose at registration. It is a server-side secret. It must never appear in a script that gets shared between units, in client-side JavaScript, or in anything attached to a case file.
Facing outward

Community alerts, and the domain pretending to be yours

Two responsibilities that sit outside the case file: warning the public accurately, and noticing when the agency's own identity is being used as the lure.

Public warnings are one of the few genuinely preventive things a department can do about this offence, and they are frequently done badly. The failure mode is a generic advisory — be careful of unexpected links, verify before you pay — issued after a wave of reports has already landed. It is not wrong, but it is not actionable either, and communities become numb to it very quickly. What makes a warning land is specificity: naming the shape of the lure actually circulating in this jurisdiction this week, and being able to say the hostnames involved have been verified as live phishing infrastructure rather than merely reported as suspicious.

What a press officer can now publish

There is a real difference between "we have received reports of a scam involving fake parking fine notices" and "we have confirmed that the addresses used in the parking fine messages circulating in this area are active credential-harvesting sites". The second is defensible, specific enough for a resident to act on, and can be issued the same day rather than after a fortnight of corroboration — and it gives the department something to say to local media that is not speculation.

The agency's own name as the lure

Police departments, sheriff's offices and court services are impersonated constantly, because the pretexts write themselves: a fine to be paid, a warrant to be resolved, a jury summons, a subpoena, a tip line, an online reporting portal. The public has no reliable way to distinguish your genuine domain from a plausible variant, and in many jurisdictions the genuine one is itself a subdomain of a county or municipal site, which makes the whole class of address unfamiliar.

Building the candidate list

Assemble the obvious typo and hyphenation variants of your own domain, the variants that swap a top-level domain, any hostname a member of the public has reported as looking like yours, and the domains of any third-party payment or reporting service you direct residents to — then run the set through /batch at up to 100 domains per request. It will not discover a hostname nobody has thought of, and should not be sold internally as if it would.

What that answer is good for

A verified, dated answer on the specific addresses you already suspect is the right basis for a warning to residents and for a takedown request to a registrar — far better than a suspicion nobody has tested. Agencies running a broader watch programme usually pair this with the approach on the brand protection page rather than treating a batch sweep as monitoring in itself.

The same file protects the people processing the reports. Officers and civilian staff get the standard pretexts plus a few of their own — training and certification renewals, court notification services, equipment vendors, and mail purporting to come from a chief or a prosecutor. Loading the daily list into whatever resolves DNS for the department network, including mobile data terminals and the devices detectives carry, is the same one-file change described throughout this page.
Questions from the unit

What investigators, analysts and agency IT ask before they commit

Including the objection that matters most, which is whether data like this belongs anywhere near a prosecution.

Can we put a match from this database in front of a court?

Not as proof of an offence, and you should not try. It is a third-party dataset, assembled through infrastructure you have not inspected, with a methodology described rather than independently audited. There is no chain of custody in the sense a court means, and nobody from the provider examined the message your complainant received. Where it belongs is on the intelligence side — prioritising reports, linking complaints, deciding where to spend a preservation request, and telling a victim whether what happened to them was real — and if the case proceeds you should get the evidence that carries weight from the routes that carry weight: the registrar, the hosting provider, the financial institutions, and your own contemporaneous capture of the live host. One qualification is worth noting: if your department downloads and retains the daily file, the contents of your own copy on a given date is a record your department made and can speak to.

Does the provider learn which domains we are investigating?

If you use /check or /batch, then yes: the hostname travels with the request. For a lot of intake triage that is acceptable, but it is a genuine consideration when the address itself is sensitive to the investigation. Taking the whole database sidesteps the problem instead of mitigating it: you retrieve the entire file once a day and match locally, so the only outbound traffic is a scheduled authenticated download that reveals nothing about what you are working on. That is why the feed is the recommended deployment for anything inside a segmented or criminal-justice network, and it is what makes an air-gapped arrangement possible at all — a file can be carried across a boundary that an API call cannot.

Our network is CJIS-segmented and outbound access needs a change board. Is this workable?

Yes, and it is a much easier change request than most. What you are asking for is one scheduled outbound retrieval of a file from a single host, at a fixed time, with no inbound path and no per-transaction traffic. That is a far narrower exception than "allow this workstation to query an external reputation service continuously". If even that is unavailable, the file can be retrieved on a network that does have a path and moved across the boundary by whatever process your agency already uses for signature and definition updates — and enterprise delivery options including SFTP and S3 exist for exactly this pattern, quoted rather than listed.

How is a domain decided to be phishing, and how current is that decision?

An entry exists only where the host has been confirmed to answer — an active A record, checked through rotating proxy infrastructure — and has been assessed as phishing or malware distribution. The category returned is phishing/malware. The database is rebuilt every 24 hours with the build landing at 04:30 UTC, and domains that stop resolving drop out rather than accumulating — which is why the active count sits above 390,000 rather than growing indefinitely. Each response carries a last_checked date and a dns_status of resolves, and the daily changelog of additions and removals gives you the other half of the picture: when an entry appeared, and when it left.

A domain from one of our reports is not in the database. What does that tell us?

Very little, and it is important that nobody in the unit reads it as exculpatory. A clean result means the hostname has not been observed and verified, which covers several quite different situations: it was registered and used within the same hour, it was used against a small enough target set that it was never seen, it stopped resolving before it could be verified, or it simply is not phishing infrastructure at all. Treat an absence as an absence of information: it should never be the reason a report gets closed, and it should never appear in a file as a finding.

We are a small agency with no cybercrime unit at all. Is this relevant to us?

Partly. The investigative use assumes somebody has time to triage reports, and in a department of twenty officers that person may not exist. What does transfer is the defensive half: loading the list into whatever resolves DNS for your network protects the people who are already there, including mobile data terminals in vehicles, and requires no analyst. For the investigative side the honest answer is the regional one: fusion centres, task forces and multi-agency cybercrime collaboratives are the natural place for a single feed subscription serving participating agencies, in the same way they already distribute bulletins and shared analysis — so if a body like that exists in your region, ask whether it has considered holding this centrally, because it is a much better fit than twenty agencies each buying separately.

What does it cost, and how does that map onto a public-safety budget?

Two shapes. Credits are monthly subscriptions via PayPal, billed monthly, at Growth at $99/month for 25,000 lookups, $99/month for 25,000 lookups, Professional at $249/month for 100,000 lookups with priority support, and larger packages down to $0.0008 per lookup at the top of the table. That suits report triage and periodic artefact enrichment. There is a 14-day refund window if under ten percent has been used, and above $4,000 bank transfer is available in place of card payment, which matters for agencies with procurement rules about cards. The Daily Threat Feed is $499 per month, with all plans include priority support; downloads are unlimited, so bulk extraction work costs nothing incremental, and every package with its per-lookup rate is set out on the pricing page.

Can we use the changelog to reconstruct what a domain was doing months ago?

Only from the point at which you started keeping it, unless you take the annual feed subscription, which includes historical archive access. The changelog is a daily record of what entered and left the list, so a department that retains its own copies builds a dated series it can consult later — and, because those are files your department downloaded and stored, they are records you can speak to directly. Set this up before you need it, because retrospective questions about a domain's history are common in this work and there is no way to answer them from a snapshot taken after the fact.

Give the intake queue a verified, dated answer

Check the database size from the open statistics endpoint, run a week of reported hostnames through the batch endpoint, and see how much of your queue collapses into a smaller number of cases. If the network is segmented, the daily feed is the deployment that keeps every query inside your own enclave.