JavaScript reader for Cloud-Optimized ZIP archives. It reads the small index and archive tail, verifies their integrity hash, then fetches the manifest without scanning the ZIP Central Directory.
npm install @asterisk-labs/cozipimport { read } from "@asterisk-labs/cozip";
const manifest = await read("https://example.com/dataset.zip");
const train = manifest.filter((row) => row.split === "train");manifest is an array of row objects: name, offset, size, cozip:location, and the writer's extras.
columns: [...] picks extras. location: false drops cozip:location.
read() supports only Flat-profile archives (profile = 1). TACO archives
(profile = 2) and every other profile are rejected with UNKNOWN_PROFILE.
const manifest = await read(url, {
columns: ["cloud_pct", "split"],
location: false,
});Only non-empty ASCII http:// and https:// URLs are supported. For cloud
storage, use a presigned HTTP URL or a CORS-enabled proxy. The server must
support range requests (Accept-Ranges: bytes) and, in browsers, allow the
Range header through CORS.
See SPEC.md for the on-disk format.
MIT. See LICENSE.