Almost every awareness curriculum in circulation is illustrated with screenshots that have been recycled for years, and adults can tell. A feed of DNS-verified domains resolving right now gives an awareness team something current to build lessons, simulations and role-specific briefings from — and a way to show learners the churn instead of asserting it.
The weakness of most awareness content is not the instructor and not the platform. It is that the artefacts on screen stopped resembling reality a long time ago.
Open almost any off-the-shelf awareness module and you will find the same museum pieces: a badly spelled inheritance letter, a grey mail-client screenshot with a red arrow drawn on it, a padlock icon presented as though it still meant something. These artefacts were produced once, signed off by a committee, exported to SCORM and then left alone, because re-shooting every screen is expensive and nobody's budget line says "refresh the pictures". The result is a course that teaches recognition of a threat which has not been fashionable for years, delivered to an audience who spend eight hours a day in software that looks nothing like the slide.
When a learner sees an obviously ancient example, they draw a reasonable inference: this is a compliance exercise rather than a warning, and whoever commissioned it does not actually know what is happening. Every subsequent instruction inherits that discount. The person who concluded in module one that this is theatre will not read the reporting instructions in module four, and will not remember the button when a genuinely well-built credential page arrives eleven months later.
Instead of a screenshot you have the current shape of the problem, delivered as CSV with domain, category and dns_status columns or as JSON. Teach where a brand name sits relative to the registrable domain, which top-level domains keep appearing, how a hyphenated imitation differs from the genuine hostname, and how often a plausible path hangs off a host with no relationship to the organisation being imitated.
Once your examples come from a job rather than from a designer, updating them costs nothing. The build lands at 04:30 UTC every day; a script that runs at five past, filters for the shapes you teach and regenerates a deck or a question bank is cheap to keep alive. The first version of that script is the expensive part, and every version after it is free — which is precisely the economics stale content has always failed to achieve.
Whether the course lives in KnowBe4, a Proofpoint awareness deployment, a Cofense installation, an in-house Moodle instance or a plain intranet page, the integration has one shape: a scheduled job pulls the current list, extracts the teaching examples, and writes them into the content store your authoring tool already reads. The API key is the username chosen at registration and belongs in that job, never in anything a learner's browser can see.
Three groups absorb most of the targeted volume in a typical organisation, and each of them gets a different lure. Generic training addresses none of them well.
Organisation-wide training has to be pitched at the person with the least context, which means it necessarily under-serves the people carrying the most risk. An accounts payable clerk, a recruiting coordinator and an executive assistant are targeted by campaigns that share almost no surface features. Teaching all three about urgency cues and sender inspection is not wrong; it is simply too abstract to be actionable in the ninety seconds they will actually have when a well-built message arrives during a busy afternoon.
This group is not caught by bad spelling; they are caught by a plausible invoice arriving in a week when a plausible invoice is expected. The lure is procedural rather than emotional — a bank-detail change, a remittance query, a portal that wants re-authentication before releasing a statement.
Recruiters open attachments from strangers for a living, which makes conventional advice about unexpected mail structurally useless to them. The lures that work here imitate applicant tracking systems, background-check vendors, benefits portals and payroll platforms — all systems the team genuinely uses and rarely inspects.
Senior staff are approached with research behind the message: a real board date, a real acquisition rumour, a real travel itinerary. Their assistants are targeted harder still, because an assistant holds delegated authority, a full calendar view and a professional obligation to be responsive.
/batch at a hundred domains per request. Anything returning is_phishing true is a teaching example with your own context attached — and a blocking candidate the same afternoon. The supply chain page covers the procurement side of that workflow.Simulated phishing is the most useful and the most frequently abused component of an awareness programme. The design choices matter more than the platform.
A simulation exists to answer two questions: does the population recognise a well-built lure, and does the reporting route work under real conditions. It does not exist to catch people out, to generate a league table, or to produce a number somebody can put on a slide. Programmes that forget this decay predictably — the templates become recognisable, colleagues start warning each other in chat, and the click rate falls for reasons that have nothing to do with anyone being safer. You end up measuring how well your workforce has learned to spot your simulation platform's sending domain.
Look at what the feed has been carrying: which brands are being imitated, which top-level domains recur, whether the imitation is hyphenated or subdomain-based, whether the hostname carries a plausible path. Then construct your simulation on infrastructure you own, imitating that structure. The exercise stays realistic because the structure is real, and it stays safe because every asset involved belongs to you.
Start a new cohort with a lure a careful reader can catch, and escalate only once the reporting route is demonstrably working. A programme that opens with an expertly researched internal pretext aimed at people who have never been told where the report button lives has learned nothing except that ambush works, and it has spent credibility it will need later. Escalate in steps and announce the escalation in advance.
Somebody who clicks should immediately see a short, specific explanation of what the tell was in that particular message — not a generic lecture, and not a mark against their name. Two or three sentences, the exact structural giveaway, and a one-click path to report the next one. Run a domain through /check during the debrief and show is_phishing, category, dns_status and last_checked; a machine-readable verdict on a real hostname builds intuition faster than prose about vigilance.
Do not aim your simulation at a live hostile host. Ever. This is the design decision people most often get wrong when they first gain access to a verified feed, it feels clever in the planning meeting, and the consequences of it are set out in full immediately below.
It is tempting. You have a list of verified-live credential-harvesting hostnames, and using one as the destination in an internal exercise would make the test maximally authentic. Do not do it. The moment a colleague follows that link, you have personally delivered a real password to a real criminal, on your own instruction, with your organisation's name on the campaign. There is no recovery narrative for that conversation.
Reporting speed, reporting coverage and report quality tell you whether the organisation can respond. Click rate mostly tells you how obvious your last template was.
Click rate is popular because it is easy to compute, easy to graph, and moves in a satisfying direction over time. It is also nearly uninterpretable in isolation. It rises when a template is good and falls when a template is stale, it varies wildly with send time and department, and it can be driven towards zero by a workforce that has simply learned to distrust everything from your platform — which is not the same thing as a workforce that is safe. Any single-number summary of human behaviour under adversarial pressure deserves suspicion, and this one has been over-trusted for a decade.
How many people reported the simulation at all, rather than merely avoiding it? How long between delivery and the first report? How long between that first report and an analyst having a verdict? What proportion of genuinely hostile messages were surfaced by a person rather than caught by a control? Those numbers describe an organisation's ability to detect and respond. Click rate describes one moment in one inbox.
When reports arrive, extract the hostnames and check them. A confirmed match returns a confidence of 0.98 with a category of phishing/malware; anything not on the list returns 0.0. The 0.98 results can be actioned and closed immediately. The 0.0 results still need a human, since absence from a known-bad list is not evidence of safety — but you have removed the easy half of the queue from an analyst's afternoon.
This is the highest-leverage thing an awareness programme can do and the thing most programmes skip. Someone who reports a message and hears nothing concludes the button goes nowhere. Someone who gets a two-line reply the same day — confirmed hostile, blocked, thank you — reports the next one faster and tells their team to do the same. Automating that reply off the back of a verified match is a small amount of engineering with a disproportionate cultural return.
Report coverage, median time to first report, median time to analyst verdict, and the share of confirmed-hostile messages a person surfaced before a control did. Four numbers from your own environment, each with a one-sentence explanation of why it matters. Executives are not attached to click rate; they are attached to having a number, and these are better numbers for the same amount of effort.
/batch request, for sweeping a week of reports in one call/stats endpoint needs no key, credits or auth, so a dashboard can show database freshnessNothing dismantles "I would recognise a fake site" faster than watching a week of hostnames appear, resolve and disappear.
The most stubborn misconception in awareness training is the belief that recognition is a skill you can finish acquiring. People genuinely think that once they have seen enough bad sites, they will know. It is a comfortable belief and it is completely wrong, because the population of hostile hostnames is not a fixed set to be memorised — it is a flow. Arguing this from first principles rarely lands. Showing it does.
The feed ships with a daily changelog of domains added and removed, which exists for a boring operational reason: it lets a subscriber apply an incremental update instead of reloading the entire database. Count the additions per day, count the removals, put both series on one chart. What the room sees is that a large number of hostnames appeared which had not existed before, and a comparable number stopped resolving and dropped out. Nothing in that picture supports the idea that a person can learn the list.
If the specific hostnames turn over constantly, memorising instances is futile and the useful skills are structural: read the registrable domain rather than the display text, be suspicious of any page asking for a password after an unexpected redirect, and above all report rather than adjudicate. The lesson stops being "learn to recognise phishing" — a demand nobody can satisfy — and becomes "recognise the situations where you should not be the one deciding," which is achievable in an afternoon.
Entries only enter the database with an active A record, verified through rotating proxy infrastructure, so a removal means the host stopped resolving. That is materially different from a domain being dropped because a reporting window closed. It lets you say without hedging that learners are looking at a picture of what is live rather than a growing archive of everything ever reported — and 390,000-plus entries being live simultaneously lands with an audience precisely because it is not cumulative.
Store each day's changelog in storage you control and the material accumulates into a longitudinal dataset that gets more useful every quarter. A trainer with fourteen months of daily diffs can show seasonality, campaign clustering, and the sharp bursts that follow a major brand event — none of which are visible in a single day's file, and none of which can be recovered later if nobody kept them.
Work through these in order. Each one is small, and every one of them can be automated once and then left alone.
Before designing anything around the feed, look at it. The /stats endpoint requires no API key, no credits and no authentication, so you can read the current database size and last update from a terminal in ten seconds. Then run a handful of hostnames you have personally seen in reported mail through /check and read the whole response rather than just the boolean — is_phishing, category, confidence, dns_status and last_checked. Knowing the exact response shape saves an argument later when somebody asks where the training material came from.
Pull six months of reported messages from your mail platform and sort them by the recipient's function rather than by threat type. The clustering is usually obvious and frequently surprising — a team you assumed was low risk turns out to receive a steady drip of one specific lure. That inventory, not a vendor taxonomy, is what your role briefings should be built from. Keep it in a document you can revise, because the distribution shifts with your business, your suppliers and the season.
Evidence: a per-function lure inventoryA cron entry shortly after 04:30 UTC, an authenticated download of the current CSV or JSON, a sanity check on the file before it replaces yesterday's copy, and an archive of the changelog into storage you control. Feed downloads are unlimited on a subscription, so a nervous first fortnight of frequent pulls costs nothing extra. Add one alert if the file is more than forty-eight hours old — a silently stale corpus is the realistic failure mode here, and it fails quietly.
Evidence: a scheduled job and a staleness alertNobody learns anything from hundreds of thousands of CSV lines. Write a small script that selects for the shapes you teach — brand-as-subdomain, hyphenated imitation, unusual top-level domain, plausible path on an implausible host — caps each category at a handful of examples, defangs every hostname, and writes the result into your authoring tool's content directory. Run it monthly. The slide reading "verified live this month" is the one that buys you the room's attention.
Evidence: a dated, defanged example setUse the feed to decide what your lure should look like, then build every asset yourself: your domain, your landing page, your logging. Grade the difficulty across the year and publish the grading. Confirm before each send that your resolver and mail gateway are blocking the current list, so a genuine hostile message arriving mid-exercise cannot be mistaken for the test. Never use a feed entry as a destination — see the warning further up this page, which exists because people have tried.
Evidence: a campaign plan naming only your own assetsWhen a colleague uses the report button, extract the hostnames server-side and run them through /batch — up to a hundred per request, one lookup each, with checked, phishing_found and credits_used in the response so the job can log a one-line summary. Auto-close the confirmed matches with a short thank-you to the reporter. Route the rest to a human, and pace the queue under the ten-requests-per-second limit rather than bursting against it.
Replace the single click-rate chart with four numbers: report coverage, median time to first report, median time to analyst verdict, and the proportion of confirmed-hostile messages a person surfaced rather than a control. Explain in one sentence why each one matters. Executives are not attached to click rate; they are attached to having a number, and these are better numbers for the same amount of effort.
Evidence: a revised quarterly report templateAwareness training is named in a great many frameworks — the on-hire and annual training requirement in PCI DSS section 12, the security awareness and training standard in the HIPAA Security Rule, the programme guidance in NIST SP 800-50, the awareness control in ISO 27001 Annex A, and the training expectations that fall out of internal-control obligations under Sarbanes-Oxley. Every one is easier to evidence when you can show dated content sourced from a verifiable feed. Keep the generated example sets, the campaign plans and the changelog archive together in one place with dates attached.
Evidence: a dated training pack per cycleTwo uncomfortable statements that belong in the same section, because each is the corrective to over-claiming the other.
Start with the harder one. Security awareness training is not a control. A control is something that either happens or does not, independently of whether a tired person made a good decision at ten to five on a Friday. Training changes the probability distribution of human behaviour, which is worth doing and worth funding, but it does not remove a failure mode and it cannot be relied upon in a design. If your architecture requires that nobody ever clicks anything, your architecture is broken, and no quantity of curriculum will repair it.
When a hostname verified as live credential-harvesting infrastructure is unresolvable on your network, the click leads nowhere and the outcome no longer depends on the click. That insensitivity to the human decision is exactly what makes it a control and training an influence. The same feed you are building lessons from should already be loaded into your resolver, mail gateway and proxy — see the email security and identity protection pages.
This is a known-bad lookup. A clean result means the hostname is not on the list; it does not mean the hostname is safe. A domain registered this morning and used for the first time this afternoon will not be present, because the verification step that makes the data trustworthy also means an entry cannot exist before the host has been observed and confirmed as resolving. That window is inherent to every blocklist ever published.
The blocklist removes the large, dull, high-volume majority — campaigns already caught and verified — from the decision entirely, so the people in your organisation are only ever asked to judge the residue. Training covers that residue, imperfectly, and the reporting culture closes the loop on what the list has not caught yet. Neither layer is sufficient; presented honestly, the pair survives contact with an auditor, a board and reality.
The costs you are avoiding are mostly not accounting costs, which is why they never appear cleanly in a return-on-investment model. A colleague who enters a single sign-on password into a convincing page has handed over mail, files and whatever else that identity unlocks, and the cleanup runs to weeks of other people's time. Preventing a modest number of those a year justifies the programme without inventing a statistic.
These come up in nearly every conversation, and several of them deserve a blunter answer than they usually get.
No. This is the one hard prohibition on this page. Every entry in the database is a hostname verified as currently resolving and confirmed as credential-harvesting infrastructure — meaning it belongs to somebody hostile, right now. Using one as a simulation destination means you have directed your own colleagues to a real attacker's collection page. Anything they type goes to a criminal, and you sent them there.Use the feed to inform the design instead: study which brands are being imitated, what the hostname structure looks like, which top-level domains recur, and then build every asset of the exercise on infrastructure you own and control. That gives you realism without handing anyone a credential. It also means your resolver can keep blocking the entire list throughout the campaign, which it should be doing anyway.
No, and it would be strange to present it that way. Platforms handle enrolment, delivery, translation, SCORM packaging, reminder mail, completion tracking and the audit reports your compliance team wants. None of that is what a domain feed does. What the feed replaces is the content staleness inside whichever platform you already run — the examples, the simulation design inputs and the report-verification step.In practice the integration is a scheduled job writing generated examples into a content directory your authoring tool reads, plus a server-side call in your report-triage workflow. Whether the platform is KnowBe4, a Proofpoint deployment, a Cofense installation or a home-built portal, the shape of the work is identical and it sits beside the platform rather than inside it.
Yes, and you should not try. Nobody is going to look at the whole database, and a dump of it is actively worse teaching material than a single well-chosen example. The number matters for a different reason: it is the count of hostnames actively resolving at one time, which is what makes the churn argument credible when you put it in front of a sceptical room.For teaching, sample deliberately. Six examples per structural pattern, refreshed monthly, defanged, with the date they were verified. That is a slide. The remaining hundreds of thousands are doing their work in your resolver, not in your classroom.
Carefully, and not with punishment. Repeat clickers are usually a signal about the job rather than the person: somebody whose role requires opening unsolicited mail from strangers all day will click more than a developer who receives eleven internal messages a week, and treating that as a personal failing is both unfair and useless. Look at the function before you look at the individual.Where somebody genuinely is an outlier within their own function, the effective response is technical rather than disciplinary — tighter controls on that mailbox, stronger conditional access on that identity, and a short conversation offering help rather than a warning. Naming and shaming reliably reduces reporting, because people who fear consequences hide mistakes, and hidden mistakes are the ones that turn into incidents.
It depends entirely on which half you want. If you only need periodic content generation and report verification, credits are the right shape: monthly subscriptions via PayPal, billed monthly, with the Growth package at 10,000 lookups, the $99 Growth package at 25,000 and the $249 Professional package at 100,000. A training team generating monthly example sets and validating reported hostnames typically lives comfortably inside the smaller packages for a long time.If you want the full daily database — which you do if it is also going into your resolver, and it should be — that is the Daily Threat Feed at $499 per month, with the annual option adding archive access, priority support, custom formats, a dedicated account manager and 100,000 credits that conveniently cover the training-side checking. Details are on the pricing page and the daily feed page. , which matters for organisations that cannot pay by card.
Give them numbers from your own environment and refuse to supply industry benchmarks, because the widely circulated ones are not traceable to anything you could defend. Median time from delivery to first report, measured across your own simulations, makes a genuinely good headline: it is specific, it moves for real reasons, and it maps directly onto how quickly your security team can act.Pair it with report coverage — the share of recipients who reported rather than merely did not click — and with the proportion of confirmed-hostile messages a human surfaced before a control did. Three real numbers from your own data are far more defensible in front of an auditor or a board than a borrowed percentage nobody can source.
Don't. Loading a live hostile page puts your egress address into the attacker's logs, may serve different content to you than to the intended victim, can attempt a drive-by download, and occasionally results in a training laptop that has to be rebuilt. It also makes your demonstration dependent on a criminal's uptime, which is a poor foundation for a session with thirty people in the room.Teach from the string, and from screenshots you made yourself of a mock page you built. If you want a live element, run /check against a defanged hostname from the current list and project the response — the 0.98 confidence on a confirmed match, the phishing/malware category, the resolves status and the last_checked date. It is genuinely live, it is genuinely current, and nothing hostile ever loads.
Partially, and it is worth being precise about how. A QR code is just an encoding of a destination, so once the code is decoded you are back to a hostname and the same verification applies — meaning a reported photograph of a suspicious code in a car park or on a poster can be resolved to a hostname and checked like anything else. That is a useful triage step and a good teaching example, because the shape of the deception is unusually visible.What the feed cannot do is anything about the sticker itself, the physical placement, or a destination nobody has observed yet. Awareness content for this vector has to carry the behavioural instruction too — treat a code in a public place the way you would treat a link from a stranger. The QR code quishing page goes into the vector in detail.
Read the current database size from the open statistics endpoint, check a handful of hostnames your own colleagues have reported, and see whether the shapes in the feed match what is actually landing in your organisation. If the same list is going into your resolver as well as your slides, the daily feed is the subscription that does both jobs.