Malicious Package Detection: Why Databases Miss Attacks and Where Detection Has to Fire

Malicious package detection is the practice of identifying open source packages that were published or modified to attack the systems that install them. It is not vulnerability scanning. The threat is deliberate, it rarely carries a CVE, and it usually executes during installation itself.

There are two structural problems that define this field in 2026. Databases of known-bad packages are records of attacks that already happened, and the damage lands in the gap before listing. And most controls sit in the pipeline, while current campaigns fire on developer machines and CI runners the moment npm install runs.

What Is a Malicious Package?

A malicious package is an open source component published or altered with a deliberate payload, built to steal credentials, open a backdoor, or spread further as soon as it is installed. A vulnerable package contains an accidental defect that an attacker might one day exploit. The two look identical in a dependency tree and demand completely different responses.

The industry catalogues malicious packages under CWE-506, Embedded Malicious Code, and they almost never receive a CVE. CVEs describe flaws in software that is trying to work correctly. A malicious package is working exactly as its author intended, so there is no flaw to describe, no patch to wait for, and no vendor advisory coming.

Vulnerable package Malicious package
Intent Accidental defect Deliberate payload
Identifier Usually has a CVE Rarely has a CVE, catalogued as CWE-506
Trigger Requires your code to reach the flawed function Often executes at install time via lifecycle hooks, regardless of your code
Urgency Can be scheduled and prioritised Incident response event, assume compromise
Remedy Upgrade to a patched version Remove, rotate every exposed credential, hunt for persistence

The trigger row changes how you operate. A vulnerability sits inert until your code calls the flawed function. A malicious package frequently runs before installation finishes, whether or not your application ever imports it. That is why a vulnerability can wait for the next sprint and a malicious package cannot wait an hour. The 2025 Verizon DBIR found that thirty percent of breaches now involve a third party, and dependencies are the largest third party most engineering teams have.

How Malicious Packages Reach Your Codebase

Malicious packages enter through six main vectors. Attackers register lookalike names, exploit internal package names on public registries, take over abandoned projects, compromise maintainer accounts, build worms that republish themselves through stolen tokens, and register names that AI coding agents hallucinate. Most large 2026 campaigns combined at least two of these.

Vector How it works Named example
Typosquatting A package name one keystroke away from a popular library, waiting for a mistype Constant background activity across npm and PyPI rather than one headline event
Dependency confusion A public package with the same name as your internal one, published at a higher version so the resolver prefers it The technique behind ongoing enterprise-targeted uploads to npm and PyPI
Hijacking of abandoned packages An attacker adopts or acquires an unmaintained project and ships a hostile update to its existing users Repeatedly observed on npm since the event-stream case, still active in 2026
Maintainer account compromise A stolen or phished maintainer account cuts a poisoned release of a trusted package keyv and cacheable, August 2026, starting from a maintainer GitHub account takeover
Self-propagating worms Malware steals npm tokens on each infected machine and uses them to republish itself into that maintainer’s packages Shai-Hulud (September 2025) and the Mini Shai-Hulud lineage reused across 2026 campaigns
Package hallucination An AI agent invents a plausible package name, an attacker registers it, the agent installs it Slopsquatting, tracked across npm and PyPI through 2026

The surface underneath keeps expanding too. GitHub’s Octoverse 2025 data counted more than 36 million new developers joining in a single year and close to a billion commits pushed, up 25 percent year over year. Every one of those commits can pull dependencies, and attackers scale with the ecosystem they target.

The tempo behind these vectors has held all year. The Cloud Security Alliance’s research on the 2026 npm campaigns follows the TeamPCP group from the CanisterWorm operation through the defeat of an SLSA Build Level 3 pipeline at TanStack, the open-sourcing of the Mini Shai-Hulud framework in May, and the Miasma campaign against Red Hat’s npm packages weeks later. Every technique in that run is now public tooling that any actor can reuse.

adadad

How Malicious Package Detection Works

Malicious package detection uses four models. Database lookup compares your dependencies against records of known-bad packages. Content analysis inspects package code for hostile behaviour. Metadata analysis scores publisher and release signals. Install-time observation watches what a package actually does when it executes. Each model fails in a different and predictable way.

Vendors rarely separate the four cleanly, and product pages blur them further. The separation matters because each model answers a different question, and the question it cannot answer is where your exposure lives.

Known-bad database lookup

A scanner checks every dependency name and version against a curated record of packages already confirmed malicious. Precision is the strength here. If your tree matches a listed entry, the finding is almost never a false positive. The failure mode is time. A record only exists after someone discovered the attack, analysed it, and published the entry, so the database describes yesterday by construction.

Static and behavioural analysis of package contents

Analysis engines inspect package code for obfuscation, credential access, unexpected network destinations, and install script abuse, sometimes detonating the package in a sandbox. This model can catch malware nobody has seen before. It fails on scale and ambiguity. Registries receive enormous publish volume, obfuscation evolves against the analysers, and plenty of legitimate packages do things that look suspicious, so verdicts arrive with false positives or arrive late.

Metadata and maintainer reputation signals

This model scores the package around the code. Publish age, maintainer history, sudden ownership changes, a new install script in a package that never had one. These signals are cheap and fast. They fail when the trust itself is stolen. The August 2026 keyv releases came from a genuine, long-standing maintainer account, so every reputation signal read clean while the payload shipped.

Install-time observation

The fourth model watches the machine during installation. Process launches, file access, and outbound connections made while a package installs, on the workstation or runner where it happens. It catches the payload itself, whatever the package is named and whether or not any list has caught up. Its limit is deployment. It only protects machines it runs on, and it works at the moment of execution rather than days in advance.

The Exposure Window

The exposure window is the gap between the moment a malicious version goes live and the moment it appears in any known-bad database. In March 2026, two malicious axios versions were live for roughly three hours, and the project’s own post-mortem tells anyone who ran a fresh install in that window to assume compromise.

Three hours is a fast takedown by any standard. It still exposed every workstation and CI runner that pulled axios during that early-morning window, because installs do not queue up politely behind security research. Every database entry is the output of a pipeline. Someone notices the package, someone triages it, someone confirms it is hostile, someone publishes the record, and every downstream scanner syncs it. Each step takes real time, and the attack runs concurrently through all of them.

That makes the window structural rather than an engineering deficiency any one vendor can fix. Faster feeds shrink it. Nothing closes it, because a list of known attacks cannot contain an attack nobody knows about yet. A campaign publishing through hundreds of stolen maintainer accounts creates new malicious versions faster than any triage pipeline can clear them, so the window reopens with every propagation cycle.

The conclusion is uncomfortable but clear. If your only control is a lookup against known-bad records, your protection begins when the window closes. Everything installed before that point was installed unprotected, and the 2026 campaigns were built around exactly that interval.

Where Detection Has to Fire

Detection has to fire where the package executes, which means the developer workstation and the CI runner at install time. Modern payloads run from a preinstall hook before the install command even returns. Any control that inspects dependencies after installation is examining a machine that may already be compromised.

A preinstall hook is a script entry in a package manifest that the package manager executes automatically before installing that package. The August 2026 npm worm wave relied on it. Affected packages carried a preinstall entry that launched a dropper, so the malicious code ran on whatever machine performed the install, before any subsequent scan, gate, or review had a chance to run.

Control point What it catches What it misses When it fires relative to package execution
Registry or proxy Requests for listed or policy-violating packages before download Anything not yet listed, installs that bypass the proxy, artifacts already cached in a mirror Before execution, but only for packages something has already flagged
IDE and CLI Risky dependency adds at the moment of introduction, on machines where tooling is present Machines without the tooling, resolution changes that happen later in CI Before or during install, where deployed
Pre-commit Manifest and lockfile changes before they enter the repository The payload that already ran on the developer’s machine during local install After execution on the workstation
PR and CI gate Bad dependencies before merge or artifact promotion The preinstall payload that already executed on the runner while dependencies installed After execution on the runner
Artifact and inventory What actually shipped, and retroactive impact assessment when a new listing lands Prevention entirely Long after execution

The last column carries the argument. A CI check that runs after dependency installation has already lost when the payload sits in a preinstall hook. By the time the check evaluates the build, the runner is compromised and its tokens are gone. Earlier supply chain attacks hit build systems and update paths, and the control map above was designed against that picture.

The 2026 campaigns target the install itself, on machines holding source repositories, cloud credentials, registry tokens, and SSH keys. A control in front of the pipeline protects the pipeline. It does not protect the endpoint where the install happens, and that endpoint is where detection has to fire.

adadad

Why the Standard Advice Keeps Failing

Four controls dominate the standard guidance. Version pinning, signing and provenance, release-age gating, and reachability analysis. Each solves a real problem, and none of them stops a malicious package that executes at install time. Knowing precisely where each one stops working explains why the 2026 campaigns kept succeeding against well-run teams.

None of what follows is a reason to drop these controls. It is a reason to stop treating them as an answer to this specific threat class.

Pinning

Pinning versions protects you from picking up a poisoned “latest” during the exposure window, and that is worth having. It does nothing when a compromised maintainer cuts a release that you later adopt deliberately. It does not govern how transitive dependencies resolve. And on a fresh lockfile generation, pinning protects nothing at all, because the lockfile is being written from live registry state.

Signing and provenance

In the August 2026 keyv and cacheable compromise, the attacker pushed to the project’s main branch and cut a release through the normal pipeline, so the poisoned versions shipped with valid provenance signed by GitHub Actions. The signature was real and the package was malware. Provenance proves where an artifact was built. It says nothing about the intent of the code inside it.

Release-age gating

Release-age gating, also called a cool-off period, delays adoption of any version younger than a set age. It is the strongest free control available, and it directly counters maintainer compromise, because a malicious release of a familiar package still carries a fresh publish timestamp. It has three limits worth knowing before you rely on it.

  • It does not block an older known-malicious version, a version that has aged past your window, or an artifact already cached by an internal mirror
  • Support varies by package manager, registry, and build path, and it only protects paths where the right toolchain is actually deployed
  • It answers whether a version is too new, which is a different question from whether a version is already known to be malicious

Reachability

Reachability analysis answers whether your code calls into a vulnerable function, and for vulnerability prioritisation that is the right question. Malware does not need to be called. Every affected package.json in the August 2026 wave carried a preinstall entry pointing at a dropper, so the payload ran whether or not the application referenced the package. We sell reachability for exploitability analysis ourselves, where it belongs. For this threat class, it is the wrong control.

What You Can Do Without Buying Anything

You can remove the most common execution vector, enforce lockfile integrity, fix the resolution path that dependency confusion abuses, and audit your existing tree against public malicious package data. All of it is free, and together it removes a large share of the surface the 2026 campaigns used.

Start with install scripts, because preinstall and postinstall hooks are the execution vector in most npm campaigns. Disabling them means a malicious package can land on disk without running anything.

npm config set ignore-scripts true

A small number of packages genuinely need build scripts, so expect to allowlist exceptions rather than flip this back off globally. Next, make CI install exactly what the lockfile says and fail on any mismatch, which blocks silent substitutions during the build.

npm ci

For dependency confusion, the structural fix is an internal registry proxy with allowlisting, so internal names can never resolve to a public registry and new public packages need an explicit decision before anyone can pull them. Finally, audit the tree you already have. OSV.dev ingests the OpenSSF Malicious Packages dataset, so one scan covers both vulnerabilities and known malicious entries.

osv-scanner --lockfile=package-lock.json

Everything above shares one limit. It reduces the attack surface and checks against what is already known. None of it observes what a package does on the machine at the moment it installs.

adadad

Malicious Packages in Agentic Development

AI coding agents change how malicious packages enter a codebase. An agent adds a dependency without a review moment, so nobody pauses over an unfamiliar name the way a developer might. And when an agent hallucinates a package that never existed, it will install whatever an attacker has registered under that name. That failure mode has a name, slopsquatting.

Slopsquatting is the practice of registering package names that AI models tend to invent, then waiting for an agent to resolve the hallucinated name to the attacker’s package. It inverts typosquatting. The attacker no longer predicts human typos, and instead predicts model output, which is more consistent and therefore easier to farm.

A USENIX Security 2025 study generated 2.23 million code samples across 16 models and found 19.7 percent referenced at least one package that does not exist, producing 205,474 unique invented names. Worse for defenders, 43 percent of those names reappeared every time the same prompt was rerun, which makes them farmable at scale.

A developer installs a handful of new packages in a working day, while an agent iterating on a task can pull dozens in minutes, on a machine holding repository access, cloud credentials, and API keys. Each install is a chance for a preinstall payload to run, with no human checkpoint in front of it.

None of this is a niche workflow anymore. Stack Overflow’s 2025 survey of more than 49,000 developers found 84 percent use or plan to use AI tools in their development process, while more of them distrust the accuracy of the output than trust it. The tooling is everywhere, and the verification habit is not.

The surface is also wider than package registries now. Malicious IDE extensions and AI agent skills are turning up alongside npm and PyPI malware, and MCP servers extend the same problem into agent tooling itself.

Most organisations cannot yet list which agents, assistants, extensions, and MCP servers run in their development environment, and an inventory you do not have is a surface you cannot defend. Mapping that shadow AI footprint is the precondition for governing any of it, which is the discovery gap Cycode’s AI Visibility exists to close.

What to Do When You Find One

Treat a malicious package as an active compromise rather than a dependency bug. Removal comes first, but the response that decides the outcome is credential rotation done in the right order, on the assumption that the CI runner is compromised along with the laptop, plus a check on whether your own tokens have already republished packages downstream.

Assume the runner because the runner installed the same dependency tree, and it typically holds richer credentials than any workstation. Take affected machines and runners out of service, then inventory what the payload could reach on them. Environment variables, npm and registry tokens, cloud credentials, and SSH keys are the standard haul in the 2026 campaigns.

Rotation order matters because you can rotate into an active session. If the attacker holds a live session with your identity provider or cloud console, new secrets created inside that session are burned on arrival. Rotate from a known-clean environment, revoke active sessions for the accounts involved, and change the credentials that control rotation before the credentials they protect.

Then check the direction nobody used to check. Worms use stolen publish tokens to cut releases under your name, so review your registry accounts for versions you did not publish and yank anything you find. Finish by sweeping artifacts built during the incident window, since a compromised runner poisons what it builds. An up-to-date software bill of materials tells you which builds pulled the package, and our CI/CD pipeline security guide covers hardening the runner.

How Cycode Detects Malicious Packages

Cycode’s capability is called Workstation Protection, and the name states the control model. It fires on the developer workstation and the build runner where install commands actually execute, and it detects malicious packages at the moment of install.

That placement follows from everything above. The exposure window means list-based checks begin protecting you only after the damage interval closes. The control-point map shows that gates behind the install-inspect machines that may already be compromised. The endpoint at install time is the one place a detection can see the payload itself, including a payload published minutes ago under a trusted name with valid provenance.

Workstation Protection is part of Cycode’s Agentic Development Lifecycle (ADLC) Security solution, built for a development process that is no longer human-paced, where dependencies arrive through agent-initiated installs as often as through a developer’s terminal. It answers the question the rest of your stack cannot. Something just installed on this machine, and what did it do.

If you want to see what this looks like on your own machines, the Agentic Development Security Platform page is the place to start, and a short demo will show detection firing on a live install.

adadad

Frequently Asked Questions

What is the difference between a malicious package and a vulnerable package?

A vulnerable package contains an accidental defect an attacker might exploit, usually tracked with a CVE and fixed by upgrading, while a malicious package carries a deliberate payload that often executes at install time and rarely has a CVE. The first is a maintenance task you can schedule, and the second is an active incident.

Why do malicious packages not have CVEs?

CVEs describe flaws in software that is trying to function correctly, and a malicious package behaves exactly as its author intended, so there is no flaw to catalogue. The industry tracks these packages under CWE-506, Embedded Malicious Code, and registries typically remove them entirely rather than waiting for a patch that will never come.

Can SCA tools detect malicious packages?

Many SCA tools flag malicious packages by matching dependencies against known-bad records, and some add content or metadata analysis on top. The weakness in that approach is timing, because database matching only works once someone has discovered and listed the attack, so anything installed during the exposure window passes every scan that runs before listing.

Does pinning dependency versions protect against malicious packages?

Pinning helps only partly, since it stops you from picking up a poisoned latest release during the exposure window. It does not help when you later adopt a compromised version deliberately, it does not govern how transitive dependencies resolve, and it protects nothing during a fresh lockfile generation that resolves against live registry state.

Does package signing or provenance prevent malicious packages?

It does not, because provenance proves where and how an artifact was built rather than what the code intends. In the August 2026 keyv compromise, the attacker released through the project's own pipeline, so the malicious versions carried valid provenance signed by GitHub Actions, and that genuine signature shipped working malware.

Does release-age gating stop malicious packages?

Release-age gating is the strongest free control available, because a malicious release of a trusted package still carries a fresh publish timestamp that a cool-off window catches. It cannot block an older known-malicious version, one that has aged past your window, or a mirror-cached artifact, and coverage depends on your package manager and build path.

Can malicious packages hide in transitive dependencies?

Yes, and most arrive exactly that way, because a direct dependency several levels up can pull a compromised package you never chose, and install-time hooks execute at any depth in the tree. The August 2026 worm spread as far as it did because poisoning one widely used package reaches every tree that resolves it transitively.

What is slopsquatting?

Slopsquatting is the practice of registering package names that AI models tend to hallucinate, then waiting for a coding agent to install the attacker's package under that invented name. It inverts typosquatting, because instead of predicting human typos the attacker predicts model output, which is more repeatable and easier to exploit at scale.

How long is a malicious package usually available before it is removed?

The window varies from minutes to months, and a short one does not mean a safe one. In March 2026, two malicious axios versions lasted roughly three hours before removal, and the project's post-mortem still told anyone who installed during that window to assume compromise, because a fast takedown limits spread without undoing installs.

What should I do if a malicious package was installed in my CI pipeline?

Treat the runner as a compromised machine rather than a cleanup chore. Take it out of service, remove the package, and rotate every credential the runner could reach, working from a clean environment so you do not rotate into an active attacker session. Finish by checking your registry accounts for releases you did not publish and rebuilding affected artifacts.