A safe, deterministic, streaming ZIP engine for modern apps — and for the agents that operate them.
Zero runtime dependencies. 100% TypeScript. One API across Node.js ≥ 22, browsers, Deno, Bun and Workers. Built for the archives that actually matter in 2026 — OOXML, EPUB, JAR/VSIX, and multi-gigabyte data drops that must never be buffered whole — under the same engineering doctrine as pdfnative.
Status: 1.x — stable (current: zipnative 1.1.0). The public API surface, the 39-code error vocabulary and the
deterministic: trueoutput bytes are frozen under semantic versioning — removals and byte changes are semver-major (the full promise is in SECURITY.md). Built up through read (v0.1), deterministic write (v0.2), incremental modification (v0.4), workers + forward streaming (v0.5), the resumable inflater (v0.6), the interop gate (v0.7), the frozen error codes (v0.8) and one-call verification (v0.9); 1.1 adds range access over remote and huge archives, cancellation and progress, raw transplant, canonicalisation of any archive, legacy name decoding and the Zip64 streaming opt-in — every one additive and opt-in. The satellites are published too:zipnative-cli(15 commands, agent-grade JSON contract) andzipnative-mcp(13 tools for AI assistants) — see Ecosystem. Documentation: zipnative.dev (site sources in docs/, interactive playgrounds included).
Most ZIP libraries make you choose between speed, safety and capability. zipnative's positioning is different:
- Safe by default. Extraction refuses path traversal (zip-slip), symlink entries, duplicate names and decompression bombs unless you explicitly opt out. Every parser loop runs under a named, CWE-tagged, caller-configurable bound. Ambiguous archives (conflicting end-of-central-directory records, Zip64 field spoofing, overlapping entries) are rejected, not guessed at.
- Random access. Read one entry from a 4 GB archive without extracting — or even scanning — the rest. The central directory is parsed lazily; entry payloads are zero-copy subarrays.
- Streaming. Iterate entries and decompress through async iterables with bounded memory. Designed for serverless and Cloudflare Workers, not just long-lived servers.
- Deterministic. Reproducible-build mode with a written determinism contract: same inputs, same SHA-256 on every runtime via the pinned pure-TS deflate encoder. Canonical entry ordering, pinned timestamps, no environment leakage.
- Incremental modification. Replace, remove, add or rename entries and
save()without recompressing the untouched 99% of the archive — the append-only overlay model proven in pdfnative's PDF incremental updates.saveCompact()is the true-deletion path (removed content is otherwise still recoverable — documented loudly). - Agent-pilotable. Every thrown error carries a stable machine-readable
err.codefrom a frozen 39-code vocabulary (v0.8+ — registry in docs/data/errors.json, guide in docs/guides/errors.md): branch on the code, never on message text. PlusverifyZip()— one call, a machine-readable verification report that never throws for archive problems (v0.9), remedy-bearing messages, a structured diagnostics channel, executable recipes, five documented production use cases,llms.txt, and a human-in-the-loop AI governance policy.
| zipnative | fflate | jszip | yauzl/yazl | adm-zip | |
|---|---|---|---|---|---|
| Zero runtime dependencies | ✅ | ✅ | ❌ | ❌ | ❌ |
| Random access (1 entry without full parse) | ✅ | ❌ | ❌ | yauzl ✅ | ❌ |
| Streaming read + write | ✅ | partial (low-level) | ❌ (memory-bound) | read or write per lib | ❌ |
| Safe-extract defaults (slip/bomb/ambiguity) | ✅ | DIY | DIY | DIY | historical CVEs |
| Deterministic output (documented contract) | ✅ | DIY | ❌ | ❌ | ❌ |
| Modify in place, no recompression | ✅ | ❌ | rewrite-all | ❌ | partial |
| Browser + Node + Deno + Bun + Workers | ✅ | ✅ | ✅ | Node-only | Node-only |
| Raw deflate throughput | good (platform zlib) | best | slow | good | poor |
fflate keeps the raw-deflate-speed crown and we do not chase it: zipnative uses the platform's native codecs (node:zlib, CompressionStream) behind a pluggable seam, and wins on scenarios — random access, bounded-memory streaming, in-place updates — not drag races.
npm install zipnativeimport { openZip, extractZip } from 'zipnative';
// Open an archive — lazy: only the central directory is located, nothing decompressed.
const zip = openZip(bytes);
console.log(zip.entryCount);
for (const entry of zip.entries()) {
console.log(entry.name, entry.uncompressedSize);
}
// Random access: decompress exactly one entry, CRC-verified.
const manifest = zip.readEntry('manifest.json');
// Stream a large entry with bounded memory.
for await (const chunk of zip.readEntryStream('video.mp4')) {
// ...
}
// Secure extraction (in memory — filesystem sinks belong to zipnative-cli).
const files = extractZip(bytes, {
limits: { maxEntries: 10_000, maxTotalUncompressedSize: 1024 * 1024 * 1024 },
// rejectTraversal: true and rejectSymlinks: true are the DEFAULTS.
});Creating archives (v0.2):
import { createZip } from 'zipnative';
const zip = createZip({
// Pin the pure-TS encoder: identical inputs → identical SHA-256,
// on every runtime. See docs/guides/determinism.md.
compression: { deterministic: true },
});
zip.add('manifest.json', JSON.stringify(manifest));
zip.add('assets/logo.png', logoBytes, { compression: { method: 'store' } });
zip.addDirectory('assets');
const bytes = zip.toBytes(); // sync, buffered
// — or, with bounded memory (serverless/Workers), byte-identical output:
for await (const chunk of zip.stream({ chunkSize: 64 * 1024 })) {
// send chunk...
}
// Large content from an async source (data-descriptor layout):
zip.addStream('video.bin', chunkSource);Modifying an existing archive (v0.4):
import { createZipModifier, openZip } from 'zipnative';
const modifier = createZipModifier(openZip(bytes));
modifier.replaceEntry('word/document.xml', newDocumentXml);
modifier.addEntry('docProps/custom.xml', customProps);
modifier.removeEntry('word/obsolete.xml');
// Append-only: untouched entries are never recompressed; the original
// bytes are preserved verbatim (removed content stays recoverable!).
const updated = modifier.save();
// True deletion + compact layout, still no recompression:
const compacted = modifier.saveCompact();
// …or the reproducible form (name order, epoch timestamps, extras dropped) — v1.1:
const canonical = modifier.saveCompact({ canonical: true });Parallel creation across worker threads (v0.5, zipnative/worker):
import { createParallelZip } from 'zipnative/worker';
const zip = createParallelZip(); // pool sized from your cores, capped at 8
zip.add('a.bin', bigBufferA); // entries deflate concurrently
zip.add('b.bin', bigBufferB);
const bytes = await zip.toBytes(); // async — the one signature difference
// Byte-identical to createZip() for the same inputs (per compression
// tier; unconditional with compression: { deterministic: true }).
// Worker failures degrade gracefully — the archive never fails for
// infrastructure reasons.Reading an unseekable stream (v0.5 — pipes, uploads, serverless bodies; a web ReadableStream<Uint8Array> or any AsyncIterable<Uint8Array> is accepted since v0.9):
import { iterateZipEntries } from 'zipnative';
for await (const entry of iterateZipEntries(request.body)) {
console.log(entry.header.name, entry.header.uncompressedSize);
if (wanted(entry.header.name)) {
for await (const chunk of entry.data()) { /* bounded memory */ }
} else if (entry.header.compressedSize > 0) {
await entry.skip();
}
}
// TRUST CAVEAT: forward iteration reads local headers alone — no central
// directory cross-check. Use openZip() whenever the full archive is
// available; it is the authoritative path.Verifying an archive in one call (v0.9):
import { verifyZip } from 'zipnative';
const report = verifyZip(bytes); // never throws for archive problems
if (!report.ok) {
// Structural refusal (report.error.code) or a failed entry —
// machine-readable either way, built on the frozen err.code vocabulary.
console.log(report.error?.code, report.entries.filter((e) => !e.crcMatch));
}
// Encrypted / stream-only-codec entries are reported as skipped with a
// reason — never faked as corruption.Opening a remote or huge archive by ranges (v1.1 — the engine never fetches; you inject a ByteRangeSource):
import { openZipRange } from 'zipnative';
const zip = await openZipRange({
size: contentLength,
read: async (offset, length) => rangeFetch(url, offset, length), // HTTP Range, Blob.slice, FileHandle.read, S3…
});
const manifest = await zip.readEntry('manifest.json'); // one tail window + the central directory + this entry
for await (const chunk of zip.readEntryStream('data.bin')) { /* 256 KiB ranged reads, bounded memory */ }
// Same limits and cross-checks as openZip(); a lying source is ZIP_RECORD_TRUNCATED.Making any archive reproducible (v1.1 — no recompression, bytes under the frozen determinism contract):
import { analyzeDeterminism, canonicalizeZip } from 'zipnative';
analyzeDeterminism(foreignBytes).offenders; // [{ name, concern: 'timestamp' | 'order' | 'extra-field' | … }]
const canonical = canonicalizeZip(foreignBytes); // sorted, epoch timestamps, extras dropped except Zip64 — idempotent
analyzeDeterminism(canonical).deterministic; // trueTransplanting entries without recompression (v1.1 — merge, split, repack in O(bytes copied)):
const out = createZip();
const a = openZip(bytesA);
for (const entry of a.entries()) out.addFromReader(a, entry); // verified first, unless { verify: false }
out.addRaw('pre.bin', deflatedBytes, { method: 8, crc32, uncompressedSize }); // already-compressed payloadCancellation, progress, UTC timestamps, legacy names, Zip64 streaming (v1.1 — every one opt-in, nothing changes by default):
for await (const file of extractZipStream(bytes, { signal: controller.signal, onProgress: (p) => bar(p.bytesOut) })) {
for await (const chunk of file.stream()) sink(file.path, chunk); // { path, entry, stream() }; an abort rejects with signal.reason
}
createZip({ defaultDate: pinned, dosTimeMode: 'utc' }); // explicit dates independent of the process time zone
openZip(bytes, { nameDecoder: (b) => new TextDecoder('shift_jis').decode(b) }); // applied only when bit 11 is clear; paths still sanitised
getExtendedTimestamps(entry); getUnixIds(entry); // UT / NTFS / ux extra fields, read-only
zip.addStream('huge.tar', chunks, { zip64: true }); // entries that MAY exceed 4 GiB (Zip64 local header + 24-byte descriptor)
verifyZip(bytes).entries[0].skipped; // 'encrypted' | 'stream-only-codec' | 'unsupported-method'Bundler notes for zipnative/worker: the worker script is resolved as new URL('./zip-worker.js', import.meta.url), which Vite and webpack 5 detect and bundle automatically. If your bundler cannot (or your CSP restricts worker sources), pass workerUrl explicitly — e.g. createParallelZip({ workerUrl: new URL('zip-worker.js', yourAssetBase) }) — pointing at a copy of the script served from your origin (locate it with import.meta.resolve('zipnative/worker/zip-worker.js') — a dedicated subpath export since 0.8). On runtimes without workers the same code runs entirely on the calling thread.
Everything public is exported from the two entry points — zipnative and zipnative/worker; if it is not exported there, it is private.
zipnative treats every archive as untrusted input. The guards, their defaults and their CWE mappings are documented in SECURITY.md. Highlights:
- decompression output capped per entry and in total, with a compression-ratio bound (CWE-400/409);
- path traversal rejected —
..segments, absolute paths, drive letters, backslashes, NUL bytes, NTFS alternate data streams (CWE-22); - symlink entries rejected by default (CWE-59);
- overlapping entries and central-directory/local-header disagreement rejected (parser-differential smuggling);
- Zip64 sentinel spoofing cross-checked; ambiguous EOCD placement refused;
- the engine never opens a socket, never touches the filesystem, and never evals.
ZIP has no veraPDF — JHOVE never shipped a ZIP module, and no ISO/IEC 21320-1 validator existed. So zipnative ships both halves of the answer: the first open clause-by-clause ISO/IEC 21320-1:2015 conformance validator (npm run validate:zip — an independent raw parser, never the engine's own, checking the ISO-standardised ZIP profile the Library of Congress recognises), and an Archivematica-grade differential extraction matrix (npm run test:interop — six independent parsers extract and byte-compare zipnative's archives on Linux and Windows). Both gates are blocking in CI and re-run before every npm publish. Every archive zipnative writes conforms to the ISO profile; the full story — including why spec-valid ≠ safe — is in the conformance guide.
iterateZipEntriesreads data-descriptor entries (flag bit 3) for plain deflate since v0.6 — including zipnative's ownaddStream()output and bsdtar-style archives. Still refused: store+bit3 (not self-delimiting), encrypted+bit3, and custom-codec+bit3;skip()on a bit-3 entry costs a full decompress-and-discard.- Codec injection (
setDeflateImpl,registerCodec) on the main entry does not propagate to thezipnative/workerbundle (separate module state); parallel/sequential byte-identity is promised for the built-in tiers. save()keeps every original byte: removed/replaced content remains recoverable in the output (usesaveCompact()for true deletion);saveCompact()drops SFX prefixes; archives with duplicate entry names cannot be modified incrementally.addStreamentries that cross 4 GiB are rejected with a typed error unless the entry opts in with{ zip64: true }(v1.1 — Zip64 local header and 24-byte descriptor). A streaming reader that sizes the descriptor from the measured lengths, such as Java'sZipInputStream, cannot read an opted-in entry that stayed below 4 GiB — opt in only for entries that may exceed it. Buffered entries, entry counts and archive offsets are fully Zip64. Underdeterministic: trueeach stream entry is buffered through the pinned encoder and capped at 2 GiB.openZipRangeneeds the archive's size up front and a source that honours exact byte ranges; a source that returns fewer bytes than asked is reported as a truncated archive, never retried.- Since v0.8.1, default extraction also refuses entries whose names are Windows reserved device names (
CON,NUL,COM1…LPT9— CWE-67) or that collapse to nothing (.,./). Archives authored on POSIX systems containing files likeaux.htherefore throw by default on every platform; passrejectTraversal: falseto skip such entries instead. - Without
CompressionStreamon the runtime (or whendeterministic: trueis requested), stream-entry compression buffers the entry before compressing — a documented memory caveat. - Number fields above
Number.MAX_SAFE_INTEGER(≈ 9 PB) are rejected; the public API usesnumber, notbigint. - Default deflate output is byte-stable per environment but not across zlib builds; pin
compression: { deterministic: true }for cross-runtime identity — the full contract lives in docs/guides/determinism.md.
- No encryption, read or write, in 1.x. ZipCrypto is cryptographically broken (Biham–Kocher); writing it would be harm dressed as a feature. AES (AE-2) may come in a later major behind an injected crypto provider. Encrypted entries are detected (
entry.isEncrypted) and reads fail with a typedZipUnsupportedErrorwhosefeaturenames the scheme ('zipcrypto','aes','strong-encryption'). - No other archive formats. No 7z, RAR, tar, gzip; no zstd/bzip2/LZMA codecs built in (the codec registry is the extension point).
- No multi-disk/spanned archives — detected and refused cleanly.
- No filesystem I/O in the engine. Extraction returns data plus sanitized paths; writing files to disk is the CLI's job. Remote or huge archives are read through a caller-injected byte-range source (
openZipRange), never a socket the engine opens. - No archive repair/salvage (rebuilding a central directory from local headers) — v1 errors cleanly instead of guessing.
- No network access, ever.
| Package | Purpose | Status |
|---|---|---|
zipnative |
core engine (this repo) — 106 exports, zero runtime dependencies | 1.1.0 |
zipnative-cli |
command-line tool (npx zipnative-cli, binary zipnative) — 15 commands, --json envelope, --dry-run, JSON Schemas, shell completion; the filesystem trust boundary the engine refuses to be |
1.0.0 |
zipnative-mcp |
MCP server for AI assistants (npx zipnative-mcp) — 13 tools, 7 prompts, sandboxed file resources, stdio + Streamable HTTP |
1.0.0 |
Both satellites pin zipnative ^1.0.0 in their dependencies and add nothing to the engine: the core stays dependency-free by exiling every dependency-bearing integration to a satellite repo — the pdfnative ecosystem pattern. Neither satellite adds encryption: docs/guides/use-cases.md shows how confidentiality is delegated to the document layer instead. Versions and inventories are recorded once, in docs/assets/ecosystem.json, and policed by npm run verify:docs.
npm ci --ignore-scripts
npm run gate:fast # typecheck:all, lint, test, verify:docs — the inner loop
npm run gate # what CI runs: coverage, build, dist probes, samples, ISO/IEC 21320-1, interop
npm run test:interop # validate generated archives with unzip/7z/bsdtar/python/jar/Expand-ArchiveThe suite is 757 tests across 61 files with 95.2% statement coverage; every count and version quoted in the documentation is tied to docs/assets/ecosystem.json by npm run verify:docs, and every generated sample archive to a byte-level baseline by npm run verify:samples. Conventions live in AGENTS.md and .github/instructions/. Contributions welcome — see CONTRIBUTING.md.
zipnative is the second library in the native family, applying the architecture proven by pdfnative: zero dependencies, closure factories instead of classes, append-only incremental modification, a shared segment generator guaranteeing buffered and streaming output are byte-identical, determinism as a product feature, and CWE-tagged bounds on every untrusted-input loop.