← Back to explorer

Under the hood

How the analyzer hunts high-risk npm packages

The analyzer in analyzer/src/analyzer.py unpacks every tarball, inspects scripts, scans for C2 infrastructure, and correlates behaviors into a verdict. Use this page as a field guide to the signals it raises before results land in the threat explorer UI.

Install-chain traps

Evaluates npm lifecycle hooks such as preinstall, install, postinstall, and prepare for risky commands, downloaders, or shell spawns before packages ever execute on disk.

  • Classifies scripts with shell, downloader, and node child_process usage
  • Flags custom build steps that reach outside the package directory
  • Correlates risky hooks with network or persistence behavior to raise install-chain risk scores

Network and exfiltration clues

Hunts for outbound traffic by inspecting bundled JavaScript, JSON, and TypeScript files for URLs, HTTP clients, and credential access patterns.

  • Groups domains and URLs seen in code, marking anything outside the allowlist as high severity
  • Detects popular C2 destinations such as Discord, Telegram, Slack webhooks, and OAST collectors
  • Surfaces env var harvesting of tokens, keys, and cloud credentials

File system and persistence moves

Tracks file writes, symlink escapes, native payloads, and suspicious binaries that packages drop to disk.

  • Detects writes into ~/.ssh, shell profiles, and system directories
  • Spots bundled native binaries, WASM modules, and prebuild downloaders
  • Records symlinks that attempt to escape extraction folders

Obfuscation and payload staging

Measures entropy, decoded base64 blobs, string-array obfuscation, and YARA signatures to highlight hidden payloads.

  • Scores high entropy bundles and mega base64 blobs with decoded previews
  • Highlights obfuscator.io style string arrays and control-flow flattening
  • Runs optional YARA packs (inline or worker mode) for malware families

How we score severity

PackageInferno assigns severity at the finding level. Package and version scores roll up the signals below.

High severity

Signals that indicate immediate compromise or credential theft. Triggered by lifecycle execution, non-allowlisted network calls, script downloads, and other payload drops.

Medium severity

Strong suspicion. Includes obfuscation, typosquatting, environment snooping, persistence writes, and large encoded blobs that often hide second-stage payloads.

Low severity

Contextual findings such as URLs in code or embedded servers. Useful for triage but not by themselves decisive.

Detection catalogue

Below is the complete list of rules emitted by analyzer.py, grouped by default severity.

High severity rules

Lifecycle hook

npm preinstall/install/postinstall scripts that spawn shells, download binaries, or run node child_process.

Install chain risk

Lifecycle activity paired with network / bundle / env snooping that suggests credential theft during install.

Install persistence risk

Lifecycle script detected writing outside the package directory (e.g. ~/.ssh, shell profiles).

Non-allowlisted URL

Hard-coded domains outside the safe allowlist (often C2, exfiltration, or staging endpoints).

C2 / webhook

Known C2 infrastructure such as Discord, Telegram, Slack webhooks, or other beacon services.

Script download

Shell or Python scripts that fetch remote payloads (curl, wget, requests, etc.).

Known CVE

Package version has known security vulnerabilities from OSV.dev / National Vulnerability Database.

Medium severity rules

Vulnerable dependency

Dependencies in package.json have known CVEs that could affect downstream users.

Typosquat detected

Similarity analysis against popular packages to flag potential typosquat or scope hijacking.

Environment variable access

Code referencing sensitive environment variables (AWS keys, npm tokens, cloud credentials).

Advanced obfuscation

High-entropy bundles, hex/Unicode trickery, or string-array obfuscation commonly used to hide payloads.

Large base64 blob

Base64-encoded blobs that may embed binaries, shellcode, or second-stage scripts.

High entropy blob

Large sections of random-looking bytes suggesting packed or encrypted payloads.

Writes outside package

File-system writes that touch user home directories, profiles, or system locations.

YARA match

Signatures from the YARA ruleset covering known malware families and suspicious tooling.

Phishing form

HTML forms or fake CAPTCHA flows that capture credentials or redirect users.

Low severity rules

HTTP server

Python or Node scripts that open sockets / HTTP listeners inside the package.

URL in code

Informational signal showing general external URLs bundled in the package.

Highlights straight from the source

These capabilities drive the findings you see in the explorer.

Typosquat sensor

Levenshtein and Unicode lookalike comparisons against a curated list of high-value packages to catch scope hijacks and brand swaps.

HTML phishing radar

Parses HTML bundles for fake login forms, external form actions, iframe implants, and CAPTCHA lures often embedded in obfuscated packages.

Polyglot awareness

Sweeps Python, shell, and auxiliary scripts that get unpacked in npm archives for downloaders, socket listeners, and persistence helpers.

Scoring engine

Weights high-signal findings (install chain abuse, C2, credential theft) more heavily than informational findings when ranking packages.

Pipeline flow

Findings stored in Aurora originate from a four-step pipeline that keeps PackageInferno data fresh.

1. Enumerate

Enumerator walks npm _all_docs_ pages and queues targets into SQS.

2. Fetch & stage

Fetcher pulls tarballs from npm, stores them in S3, and triggers analyzer jobs.

3. Analyze

Analyzer extracts, inspects files, scores findings, and writes results into Aurora along with JSON blobs for the explorer.

4. Explore

API and UI render severity rollups, rule breakdowns, and extraction snippets for triage.

Dig deeper