Threat Intelligence Platforms

One narrow feed that is verified as resolving, sitting beside everything else you already ingest

This is not an intelligence programme. It is a single, deliberately narrow collection of phishing domains, each confirmed to have an active A record, rebuilt daily and shipped with a diff. What it changes in a TIP is the false-positive economics of the phishing slice — and therefore how much analyst time your triage queue consumes.

390K+Verified-live indicators
DailyRebuild plus add/remove diff
1 typeDomain-name indicators only
0.98Confidence, or 0.0 — nothing between
Indicator record — check response
is_phishingtrue
categoryphishing/malware
confidence0.98
dns_statusresolves
last_checkeda date, not a guess
observed viarotating proxy verification
MISP & OpenCTI
STIX & TAXII
Pivoting & enrichment
Retro-hunting
ISAC sharing
Triage reduction
Home / Use Cases / Threat Intelligence Platforms
What this page covers

Four questions a TIP owner asks about any new source

The answers below are specific rather than promotional, and one of them is an argument for not buying this at all.

Adding a feed to a threat intelligence platform is not free, even when the feed is. Every source consumes storage, consumes correlation cycles, consumes the attention of whoever maintains the connector, and — most expensively — consumes analyst trust the moment it produces a match that turns out to be worthless. The right way to evaluate a new source is therefore to ask what it uniquely contributes, at what verification standard, at what maintenance cost, and where it goes silent. This page works through those four in order.

1

Where it sits

Next to OSINT blocklists, broad commercial feeds and sector sharing — overlapping all three, replacing none of them.

2

Reported vs resolving

The single distinction that determines whether your phishing indicators generate work or generate answers.

3

Getting it in

MISP feed modules, OpenCTI connectors, STIX bundles over TAXII, and plain scheduled delivery.

4

Where it is silent

One indicator type, no actor attribution, no campaign narrative, and a real gap in the first hours of a domain's life.

Position in the stack

Four sources that all claim to cover phishing, and what each is actually for

Most intelligence teams already have three of these. Understanding what each one is genuinely good at is what stops the fourth from being redundant.

The phishing slice of a threat intelligence programme is unusually crowded, and the overlap between sources is real but partial in ways that are rarely examined. A useful exercise before adding anything is to take a hundred phishing domains you have seen in your own environment over the past month and check how many of your existing sources carried each one, and how current that entry was when you matched it. The answer is frequently uncomfortable, and it is far more informative than any coverage claim.

Source 01

Open-source blocklists and community reporting

Enormous breadth, no cost, and highly variable quality. Entries arrive through public submission with little or no verification, dead domains are rarely retired, and the same hostname often appears in three lists with three different classifications. Genuinely useful as a wide net and as corroboration, and genuinely dangerous as an alerting source — a large share of what you will match on is infrastructure that stopped existing months ago.

Breadth, at the cost of precision
Source 02

Broad commercial intelligence feeds

These carry many indicator types across many threat categories, usually with actor attribution, campaign context and analyst-written reporting attached. That context is the product and it is worth paying for. Phishing hostnames tend to be a thin slice of the whole, prioritised behind malware infrastructure and targeted intrusion tracking, and refreshed on a cadence set by the vendor's overall publication cycle rather than by the lifespan of a credential-harvesting host.

Context and attribution, thinner on phishing volume
Source 03

ISAC and sector sharing communities

The most valuable phishing intelligence you will ever receive, because it is targeted at your sector by people who look like you and it arrives with the context of an actual victim. It is also episodic, unevenly contributed, and dependent on somebody in a peer organisation being both compromised and willing to say so promptly. Sector sharing tells you what is being aimed at your industry; it cannot tell you what is live everywhere else.

Highest relevance, lowest and least predictable volume
Source 04

A single-purpose verified-live domain feed

One indicator type, one category, and a hard inclusion rule: the host must currently resolve, confirmed through rotating proxy infrastructure, or it is not in the database. No attribution, no campaign narrative, no analyst commentary. What it contributes to a stack that already has the other three is density and freshness in exactly the slice where the other three are weakest — several hundred thousand hostnames that are live right now rather than a growing record of everything ever reported.

Narrow scope, high verification standard
Run the overlap test before you buy anything. Export a sample of the phishing indicators already in your TIP and submit them through /batch at a hundred domains per request. The proportion that come back 0.0 tells you how much of your existing phishing inventory refers to infrastructure that no longer resolves — which is the number that actually decides whether this source adds anything for you.
Reported vs resolving

The difference between two words that decides your triage load

"Reported" means somebody said so. "Verified as resolving" means a machine confirmed it, and the two produce very different queues.

Almost every complaint analysts make about threat intelligence reduces to the same underlying problem: an indicator matched, the analyst investigated, and the investigation established that the indicator referred to something that no longer existed. That is not a false positive in the strict sense — the domain probably was hostile once — but the operational effect is identical, and after enough of them a team stops treating matches from that source as worth opening. The distinction between an indicator that was reported and one that has been verified as currently resolving is the entire difference between those two outcomes.

What "reported" actually guarantees

That at some point somebody submitted the hostname to a collection process. It does not guarantee the submission was correct, that the classification was right, that the host was ever live, or that it still is. Most public collections have no retirement mechanism at all, so a domain reported three years ago and taken down four days later remains in the list indefinitely, indistinguishable from one that appeared this morning.

What DNS verification changes

Inclusion here requires an active A record, confirmed through rotating proxy infrastructure rather than from a single vantage point that a hostile host could be selectively answering. Entries that stop resolving fall out on the next build. The consequence is that a match tells you something about the present rather than about a moment in the past, which is what makes it safe to act on rather than merely worth reading.

The re-registration trap

Expired domains get bought by other people. A hostname that hosted a credential page two years ago may today belong to a small business, a personal blog, or a legitimate service your own staff use. A feed that never retires anything will eventually make you block one of those, and explaining that to the affected party is a conversation nobody enjoys. Retiring on non-resolution is not a nicety; it is what keeps a blocklist safe to enforce on.

What it costs you in analyst minutes

Every dead-indicator match consumes the same triage overhead as a live one — open the case, pull the logs, establish context, write it up, close it. Multiply that by the volume a broad unverified list generates and the arithmetic gets uncomfortable quickly. Cutting the dead-indicator share of your phishing matches is not a quality improvement in the abstract; it is a direct reduction in hours spent reaching a conclusion of "nothing here".

The corresponding cost of this standard is a delay. A domain cannot enter the database until it has been observed and confirmed as resolving, so anything registered and used within the same hour is not there yet. Verification buys precision by giving up some of the very front of the timeline, and that trade-off should be made explicitly rather than discovered during an incident review.
The shape of the data

What actually arrives, in the units your platform will store it in

Three columns, one indicator type, one category, a binary confidence and a date. Deliberately unremarkable, which is the point.

Connector maintenance is where feeds quietly die. A source with a rich, evolving schema needs a connector that tracks it, and the person who wrote that connector eventually changes jobs. This feed is intentionally boring in shape — CSV with domain, category and dns_status, or JSON if you prefer — precisely so that the mapping into your TIP is written once and then does not need revisiting. Full downloads are unlimited on a subscription, so a connector that pulls too eagerly during development costs nothing.

390,000+Domain-name indicators live at any one moment, not cumulative
04:30 UTCDaily build time, with a diff of additions and removals alongside
Sub-50msSingle-domain lookup latency for interactive enrichment
100 / POSTBatch size, at one credit per domain, with a counted response
indicator type: domain-name category: phishing/malware confidence: 0.98 or 0.0 dns_status: resolves last_checked: date stats endpoint: no key required
Enrichment and pivoting

Five places this changes what an analyst does next

The value is not in having more indicators. It is in having one class of question that can be answered without opening another tab.

A threat intelligence platform earns its keep at the moment an analyst has an observable in front of them and needs to know whether it matters. Everything upstream of that moment — collection, normalisation, deduplication, storage — exists to make that one interaction fast. A narrow, high-confidence source contributes to it in specific ways that are worth enumerating rather than summarising, because each one lands in a different part of the workflow.

1

Inline verdict on an observable, without leaving the case

An analyst pastes a hostname pulled from a mail header, a proxy log or a reported message and gets a verdict in under fifty milliseconds. A confirmed match returns 0.98 with the category and the last_checked date; anything unknown returns 0.0. That is a decision, not a research task, and it removes the most common reason an investigation stalls waiting for somebody else's opinion.

2

Corroboration weight when sources disagree

When one open-source list flags a hostname and another does not, the deciding question is usually whether the host is currently up. A verified-resolving source answers that directly, which turns a two-source disagreement into a resolved one. It is a small function that gets used constantly, and it is why this fits better as a corroboration layer than as a primary alerting feed.

3

Cluster detection through the daily diff

A dozen hostnames appearing in the same day's additions, imitating the same brand or built on the same naming pattern, is a campaign being stood up. Watching the diff rather than the whole database is what makes that visible, because the signal is entirely lost inside a set of several hundred thousand. Archive every diff and the clustering becomes something you can look back through.

4

Bulk triage of a reporting backlog

Reported-phishing queues accumulate faster than anyone processes them. Extracting hostnames and running them through /batch at a hundred per request splits the queue into confirmed-hostile and unknown, with checked, phishing_found and credits_used returned per call. The confirmed half can be actioned and closed programmatically; the unknown half still needs a person, but it is a much smaller pile.

5

Distribution outward to the enforcement layer

A TIP's real job is pushing indicators to the tools that act. The same set feeding your correlation searches can populate resolver blocklists, proxy denylists and mail gateway rules, and the verification standard is what makes that safe to automate. The SIEM integration page covers the correlation side and DNS filtering the enforcement side.

Ingestion

Five routes in, ordered by how little work they are

Whichever platform you run, the mapping is the same handful of fields. What differs is the transport and the expiry semantics.

Before choosing a route, decide one thing: whether your TIP is the system of record for expiry, or whether the feed is. Trying to run both produces indicators that expire on a schedule your platform invented while the source has already removed or re-added them, and the resulting drift is impossible to reason about six months later. The simplest correct answer is to let the daily rebuild be the truth and configure your platform's validity window to match it.

A

MISP, through the native feed module

MISP ingests CSV feeds directly with configurable column mapping, which makes this close to a configuration exercise rather than a development one. Point a feed at the daily file, map the domain column to a domain attribute, and tag events with your phishing taxonomy so downstream correlation and sharing behave predictably. Set the feed to overwrite rather than accumulate, so that domains removed upstream stop being distributed by your instance to everyone you share with.

B

OpenCTI, through a small custom connector

OpenCTI takes STIX 2.1 objects, so the connector's job is to read each row and emit an indicator with a pattern of the form [domain-name:value = '…'], labelled for phishing, with a valid-from of the build date. Keep the connector stateless and idempotent — read the file, emit the bundle, let the platform deduplicate — because a connector that maintains its own state is a connector that eventually disagrees with reality and has to be rebuilt.

C

Commercial TIPs, through the batch import API

Platforms in this class generally expose a bulk indicator import that accepts a flat file or a paged API call, and the three-column shape maps onto their schemas without transformation. The detail worth getting right is the confidence field: map 0.98 to whatever your platform calls high and do not attempt to blend it with other sources' scores, because averaging a binary verdict with a graded one produces a number that means nothing to anybody.

D

STIX and TAXII, for sharing communities

Where you need to redistribute rather than merely consume — into a sector sharing group, a subsidiary's platform, or a customer-facing collection — STIX and TAXII delivery is available as part of enterprise arrangements alongside SFTP and S3. Check the licence position before redistributing beyond your own organisation, and think about marking: an indicator you can act on internally is not automatically one you are free to publish onward.

E

No platform at all, just a scheduled file

Plenty of teams do not have a TIP and do not need one for this. A cron entry after 04:30 UTC, an authenticated download, a validity check, an archived copy of the diff, and the file dropped where your detection and blocking tools read it covers the entire use case. Add one alert if the file is more than forty-eight hours old, because a silently stale copy is the only realistic way this deployment fails.

Record your own first-seen date. The day an indicator entered your platform is a fact about your detection coverage that nobody else can supply, and it is the field post-incident reviews ask for most often. Combined with the source's last_checked date, it lets you reconstruct exactly what was knowable when — which is the question that decides whether an incident was a detection failure or an intelligence gap.
Honest limits

Four things this is not, said plainly before you build on it

A narrow source is only safe to rely on if everybody using it understands the shape of the hole.

The failure mode of a good indicator source is over-extension: a team gets used to trusting it, starts treating its silence as reassurance, and eventually builds a control whose logic depends on the source knowing things it was never designed to know. Writing the boundaries down at integration time, in the runbook rather than in a slide, is the cheapest insurance available against that.

It is not an intelligence programme

There is no actor attribution, no campaign narrative, no tooling analysis and no analyst reporting. It will not tell you who is behind a domain, what else they run, or what they intend. If your requirement is understanding an adversary, this is a data source inside that effort and not a substitute for it — the commercial feeds and sector sharing described earlier are what do that work.

It is not present at hour zero

A host must be observed and confirmed as resolving before it can be included, so the first hours of a domain's operational life are outside the window. Campaigns that register and burn infrastructure within a single day will be partially invisible. This is the direct cost of the verification standard and it cannot be engineered away without abandoning the property that makes the data trustworthy.

It is domains, not URLs or content

A hostile page hosted on a compromised legitimate site does not produce a hostile domain, so it will not appear. Neither will an attack delivered entirely inside a platform you do not control. Detecting those requires content inspection and behavioural analysis, which are different disciplines with different tooling and different error profiles.

A 0.0 result is not an all-clear

It means the hostname is not in this database. Any dashboard tile, playbook branch or analyst runbook that renders that as a green tick has converted an absence of evidence into a positive assurance, and somebody will rely on it at three in the morning. Label it "no known match" everywhere it is displayed, and write the sentence into the runbook next to the enrichment step.

What remains after those four caveats is still worth having. A dense, current, mechanically verified set of phishing hostnames removes a large volume of confirmed-hostile traffic from your analysts' attention and cuts the dead-indicator share of your phishing matches — which is a real reduction in triage hours rather than a claim about coverage. That is the correct case for this feed, and it does not need embellishing.
Questions from intelligence teams

What gets asked in a source evaluation

Including the one about whether you need this at all, answered without a sales instinct.

We already run three feeds. How do we tell whether this overlaps them?

Measure it rather than reasoning about it, because coverage claims from any source including this one are not a substitute for your own numbers. Export a sample of the phishing-category domain indicators already in your platform — a few thousand is plenty — and submit them through /batch at a hundred per request. Two numbers come out of that. The proportion returning 0.0 tells you how much of your existing phishing inventory refers to infrastructure that no longer resolves, which is the dead weight your analysts are already paying for in triage time. Then run the reverse test: take a week of the daily additions and check how many of them appear anywhere in your existing sources, and how quickly. If your current feeds already carry verified-live phishing coverage with comparable freshness, you genuinely may not need this, and that is a legitimate outcome of an evaluation.

Why is the confidence binary rather than graded?

Because the underlying method is binary and inventing a gradient would be manufacturing precision that does not exist. Inclusion depends on a factual test — the host has an active A record confirmed through rotating proxy infrastructure and has been classified as credential-harvesting or malware infrastructure — and a domain either passes that test or is absent. A graded score would imply that some entries are more verified than others, which is not the case, and would then invite every consumer to invent their own threshold. In practice a binary verdict makes downstream design simpler: there is no tuning conversation in a rule review, no drift when somebody recalibrates a model, and no ambiguous middle band that ends up routed to a queue nobody watches. Where variation belongs is in the severity you assign based on surrounding context — a DNS resolution with no session is not the same situation as a completed session followed by an authentication — and that judgement lives in your correlation logic, not in the indicator.

Can we redistribute these indicators to our ISAC or to our own customers?

Ask before you do, because the answer depends on the arrangement you hold rather than on a general rule, and getting it wrong in a sharing community is awkward to unwind. Enterprise arrangements exist specifically to cover redistribution scenarios — STIX and TAXII delivery, SFTP and S3 transport, custom update frequencies and on-premise deployment are all quoted rather than listed, and redistribution scope is part of that conversation. What is worth thinking through independently of the licence position is marking discipline: if you publish onward into a community that applies traffic-light protocol markings, decide in advance what marking these carry and make sure your connector applies it automatically, because manual marking of a daily automated feed will be forgotten within a fortnight. Reach out at [email protected] with the specific sharing arrangement you have in mind.

How should we set the validity window on ingested indicators?

Let the daily rebuild be the system of record and configure your platform to match it rather than inventing an independent expiry. The practical shape is a valid-from of the build date and a validity window of roughly twenty-four hours, refreshed each morning by the next ingestion — which means an indicator that stops resolving simply fails to reappear and ages out naturally, and one that persists is continuously re-asserted with a current date. What causes trouble is a platform-side time-to-live set longer than the refresh cycle, because then a domain removed upstream lingers in your instance for days after it stopped resolving, and you have quietly recreated the stale-indicator problem you adopted this source to avoid. If you use MISP, configure the feed to overwrite rather than accumulate for the same reason.

What do we do about domains that were hostile last week and are now clean?

Nothing retroactive, and this is worth writing into your case-handling procedure because analysts get it wrong in both directions. A domain leaving the database means it stopped resolving, not that it was never hostile, so a case raised while it was live was correct and must not be reclassified as a false positive when somebody reviews it a month later. Equally, a domain that has dropped out should stop generating new alerts, because continuing to enforce against infrastructure that no longer exists eventually catches a legitimate re-registration. The mechanism that makes both behaviours correct is storing the indicator state as it was at action time inside the case record itself — the feed is a moving picture of now, and cases need a fixed picture of then. Archive the daily diffs and you can reconstruct that history even for cases where somebody forgot.

Is there a way to evaluate the data before committing to a subscription?

Yes, in two steps that cost very little. The /stats endpoint requires no API key, no credits and no authentication, so you can read the current database size and last update time immediately and confirm the rebuild is landing when this page says it does. After that, a small monthly plan is the sensible evaluation vehicle: the Growth plan is $99/month for 25,000 lookups, which is more than enough to run the overlap test described above and to check a month of your own reported-phishing queue against the database. Credits are the wrong tool for bulk correlation — that needs the whole set locally — but they are exactly right for deciding whether the whole set is worth having. The full table is on the pricing page, and there is a fourteen-day refund if under ten percent has been used.

Does this help with anything other than phishing?

Marginally, and it would be dishonest to claim more. The category field carries phishing/malware, and credential-harvesting infrastructure frequently overlaps with malware distribution because the same operators run both, so some entries will be relevant to a malware investigation. But the collection is built and verified around phishing, the classification is not granular enough to separate the two reliably, and using it as a malware source would mean depending on a side effect rather than on a designed property. If malware infrastructure tracking is the requirement, that belongs to a source built for it — this one will contribute occasional corroboration and nothing more. The adjacent uses that genuinely do work are covered on the incident response and managed security pages.

Our team is two people. Is a feed like this realistic for us?

More realistic than for a large team, in one specific sense: a small team has no capacity to absorb triage overhead, so the dead-indicator problem hurts them disproportionately and a narrow high-precision source helps them disproportionately. The honest constraint is not the data, it is the maintenance — every connector needs an owner, and on a two-person team that owner is also doing everything else. Which is why route E on this page exists: you do not need a threat intelligence platform to use this. A scheduled download, an archived diff, a file in a location your blocking and detection tools read, and one staleness alert is a complete deployment that fits in an afternoon and then stays quiet. Build the platform integration later if you ever acquire one, and resist doing it first because the architecture diagram looks better with a TIP in it.

Related integration paths

Where the same feed does its work elsewhere

The distribution side of a TIP's job, covered in more depth on the pages below.

Run the overlap test before you decide anything

Read the current database size from the open statistics endpoint, then check a sample of the phishing indicators already sitting in your platform. How many of them still resolve is the number that tells you whether this source belongs in your stack — and it is a number no vendor can supply on your behalf.