Zip Slip: The Vulnerability I Found Through a SonarQube Warning

This is an educational technical note. Example code is illustrative and simplified to demonstrate the concept.

I'm starting a new section on this blog called New Learnings, for things I come across in day-to-day engineering work that I think are worth writing down properly instead of just filing away. First up is a vulnerability class I hadn't given much thought to until a static analysis tool put it in front of me: Zip Slip.

How I found it

It came up during a routine pass through SonarQube findings on a codebase I was working in. Buried in a batch of code smells and minor issues was a security hotspot flagging a file extraction routine, with a description along the lines of: "This function extracts an archive file. Ensure it is safe to extract without verifying entry paths." SonarQube's recommendation pointed at the Zip Slip vulnerability by name, along with a suggested fix.

I hadn't heard the term before, so I went digging. It turns out Zip Slip is a well-documented vulnerability class — it was popularised by a 2018 Snyk security research disclosure that found it in a large number of popular Java, JavaScript, .NET, Go and Ruby libraries. What struck me was how simple the root cause is, and how easy it is to write extraction code that looks completely reasonable while being exploitable.

What Zip Slip actually is

Zip Slip is a path traversal vulnerability that happens when an application extracts files from an archive (zip, tar, jar, war, and similar formats) without validating the file paths stored inside that archive.

Archive formats let each entry specify its own relative path — that's how a zip file preserves folder structure when you extract it. Nothing stops that path from containing directory traversal sequences like ../../. If the extraction code blindly joins the entry's path onto the destination directory and writes the file there, a malicious archive can walk straight out of the intended output folder and overwrite files anywhere the process has write access to — configuration files, cron jobs, startup scripts, or application code itself.

The name comes from exactly that: a file "slipping" outside the directory it was supposed to be extracted into. It doesn't require any exotic exploit technique. It requires the victim application to trust the file names inside an archive it didn't create.

A vulnerable example

Here's a simplified version of the kind of code that trips this up — a Java method extracting a zip file, which is close to what SonarQube flagged:

public void unzip(File zipFile, File destDir) throws IOException {
    try (ZipInputStream zis = new ZipInputStream(new FileInputStream(zipFile))) {
        ZipEntry entry;
        while ((entry = zis.getNextEntry()) != null) {
            // Vulnerable: the entry name is trusted and joined directly
            File outFile = new File(destDir, entry.getName());

            if (entry.isDirectory()) {
                outFile.mkdirs();
                continue;
            }

            outFile.getParentFile().mkdirs();
            try (FileOutputStream fos = new FileOutputStream(outFile)) {
                zis.transferTo(fos);
            }
        }
    }
}

This looks like ordinary extraction logic, and it works correctly for well-behaved zip files. The problem is entry.getName(). Nothing validates it before it's combined with destDir.

Now imagine a zip file crafted with an entry named:

../../../../etc/cron.d/malicious-job

When new File(destDir, entry.getName()) resolves that path, it doesn't stay inside destDir at all — the ../ sequences walk back up the directory tree, and the file gets written wherever those sequences land, limited only by the permissions of the process running the code. On Windows, the same idea applies using backslashes or drive-letter paths. An attacker doesn't need any other foothold in the system; they just need the application to extract an archive they supplied — a file upload feature, a plugin/theme installer, a CI artifact step, anything that unpacks user-supplied content.

How to avoid it

The fix is to never trust an archive entry's path and to explicitly verify that the resolved output file stays inside the intended destination directory before writing anything.

1. Validate the resolved path against the destination directory

public void unzip(File zipFile, File destDir) throws IOException {
    String destDirPath = destDir.getCanonicalPath();

    try (ZipInputStream zis = new ZipInputStream(new FileInputStream(zipFile))) {
        ZipEntry entry;
        while ((entry = zis.getNextEntry()) != null) {
            File outFile = new File(destDir, entry.getName());
            String outFilePath = outFile.getCanonicalPath();

            if (!outFilePath.startsWith(destDirPath + File.separator)) {
                throw new IOException("Entry is outside of the target directory: " + entry.getName());
            }

            if (entry.isDirectory()) {
                outFile.mkdirs();
                continue;
            }

            outFile.getParentFile().mkdirs();
            try (FileOutputStream fos = new FileOutputStream(outFile)) {
                zis.transferTo(fos);
            }
        }
    }
}

The key line is the canonical path check. Resolving both paths with getCanonicalPath() collapses any ../ sequences and symlinks first, so the comparison reflects where the file would actually end up — not the unresolved string. Reject the entry rather than trying to sanitise or truncate the path; a rejected file is far safer than a "cleaned" path that still resolves somewhere unexpected.

2. Prefer a library that already handles this

Where possible, use an extraction library that validates entries for you rather than hand-rolling it. Apache Commons Compress and similar well-maintained libraries are aware of Zip Slip and are a safer default than writing extraction logic from scratch, since this class of bug has already been found and fixed there by people focused on exactly this problem.

3. Apply the same principle outside Java

This isn't Java-specific — the same flaw shows up anywhere an archive is extracted without path validation: Node's adm-zip or manual tar handling, Python's zipfile/tarfile modules, .NET's ZipArchive, Go's archive/zip. The fix is conceptually identical in every language: resolve the full destination path, confirm it's still inside the target directory, and refuse to write the file if it isn't.

4. Treat uploaded archives as untrusted input

More generally, treat any archive that didn't originate from your own build pipeline as untrusted input, the same way you'd treat any other user-supplied data. That means path validation on extraction, but also sensible limits on extracted size and file count to avoid a related issue — zip bombs — while you're in that code.

Key takeaways

  • Zip Slip is a path traversal vulnerability caused by extracting archive entries without validating their paths.
  • A malicious entry name like ../../etc/cron.d/job can write files outside the intended extraction directory.
  • Always resolve the full output path and confirm it's still inside the destination directory before writing — reject the entry if it isn't.
  • Where possible, lean on a well-maintained archive library rather than writing extraction logic yourself.
  • Static analysis tools like SonarQube are good at catching this pattern — worth taking security hotspots seriously even when they seem minor at a glance.

What I liked about this one is how well it illustrates a broader habit worth having: reading the "why" behind a static analysis finding rather than just applying the suggested fix and moving on. Half an hour of reading turned a one-line warning into something I'll actually recognise the next time I see extraction code in a review.