Skip to content

Commit 3fb6de4

Browse files
authored
Merge pull request #3 from ZenRows/fix/cli-taxonomy-primitives
Replace old product taxonomy with canonical primitives (Fetch / Extract / Batch / Browser Sessions)
2 parents 622124a + a4773cd commit 3fb6de4

53 files changed

Lines changed: 2264 additions & 277 deletions

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎README.md‎

Lines changed: 28 additions & 16 deletions
Original file line numberDiff line numberDiff line change
@@ -4,8 +4,8 @@
44
workflows, recipes, and evals layer for giving AI agents reliable access to
55
protected web data through Zenrows cloud infrastructure.**
66

7-
It makes the Zenrows Universal Scraper API, Scraping Browser, and Residential
8-
Proxies installable and usable directly from AI agents and developer workflows —
7+
It makes the four Zenrows primitives — Fetch, Extract, Batch, and Browser
8+
Sessions — installable and usable directly from AI agents and developer workflows —
99
so an agent can reliably access protected web data without hand-rolling anti-bot
1010
handling, proxies, or browser rendering.
1111

@@ -21,25 +21,25 @@ through AI agents, developers, and teams.
2121
## 2. Why Zenrows
2222

2323
Normal fetch fails. Generic scrapers fail. Browser-first tools are expensive.
24-
The Zenrows **Universal Scraper API** retrieves protected pages reliably and
25-
structures them, while the **Scraping Browser** is there for the rare cases that
24+
Zenrows **Fetch** retrieves protected pages reliably and **Extract** structures
25+
them, while **Browser Sessions** are there for the rare cases that
2626
need a real browser. Zenrows wins when the workflow runs over thousands or
2727
millions of URLs.
2828

2929
## 3. Product architecture
3030

31-
The core product is the Zenrows **Universal Scraper API**
31+
The core product is Zenrows **Fetch and Extract**
3232
(`GET https://api.zenrows.com/v1/`). The CLI exposes it two ways: `zenrows fetch`
3333
retrieves a protected page, and `zenrows extract` turns it into structured data
3434
(JSON / CSS / Markdown). Both call the same endpoint — `extract` is just
3535
that endpoint with extraction parameters, not a separate product.
3636

3737
| Command | What it does | Status (this build) |
3838
| --- | --- | --- |
39-
| `zenrows fetch` | Universal Scraper API — retrieve a protected page | **available** — `GET https://api.zenrows.com/v1/` |
40-
| `zenrows extract` | Universal Scraper API — structured extraction (Autoparse / CSS / Markdown) | **available** — same `/v1/` |
41-
| `zenrows batch` | Batch Scraper API — fan out over many URLs | beta — cloud works with beta access; local validate/estimate always |
42-
| `zenrows browser` | Scraping Browser (CDP) / MCP escalation | experimental |
39+
| `zenrows fetch` | Fetch — retrieve a protected page | **available** — `GET https://api.zenrows.com/v1/` |
40+
| `zenrows extract` | Extract — `extract=auto` (domain-gated open beta; falls back to Autoparse) / CSS / Markdown | **beta** — same `/v1/` |
41+
| `zenrows batch` | Batch — fan out over many URLs | beta — cloud works with beta access; local validate/estimate always |
42+
| `zenrows browser` | Browser Sessions REST API (same backend as MCP `browser_*`) | **available** — escalation-only; bills by bandwidth + time |
4343
| `zenrows mcp` | MCP server config (remote + local) | **available** |
4444
| Zenrows CLI | this repo | available |
4545

@@ -53,7 +53,8 @@ matrix.
5353
```bash
5454
npx -y @zenrows/cli init
5555
zenrows fetch https://httpbin.io/html # auto-provisions a Free plan account on first use
56-
zenrows extract https://www.scrapingcourse.com/ecommerce/ --autoparse
56+
zenrows extract https://www.owler.com/company/meltwater # extract=auto on an enabled domain
57+
zenrows extract https://www.scrapingcourse.com/ecommerce/ --autoparse # Autoparse (any domain)
5758
```
5859

5960
No API key up front: on your first cloud call the toolkit creates a free,
@@ -110,14 +111,16 @@ zenrows fetch <url> --proxy-country us --wait-for ".price"
110111
## 8. Extract
111112

112113
```bash
113-
zenrows extract <url> --autoparse
114+
zenrows extract https://www.owler.com/company/meltwater # extract=auto (enabled domain)
115+
zenrows extract https://www.scrapingcourse.com/ecommerce/ --autoparse # Autoparse (any domain)
114116
zenrows extract <url> --css '{"title":"h1","price":".price"}' --validate
117+
zenrows extract <url> --outputs emails,links # built-in output filters → JSON
115118
zenrows extract <url> --output markdown
116119
```
117120

118121
## 9. Batch (beta)
119122

120-
The Zenrows **Batch Scraper API** (`https://async.api.zenrows.com/v1`) fans a
123+
Zenrows **Batch** (`https://async.api.zenrows.com/v1`) fans a
121124
protected fetch/extract out over many URLs. It is a real product in
122125
**beta**: the cloud subcommands work once your API key has
123126
beta access; without it the API returns `BATCH_ACCESS_DENIED`. Local spec
@@ -138,8 +141,17 @@ locally or fan out with `zenrows fetch` per URL.
138141

139142
## 10. Browser Sessions
140143

141-
Escalation only, gated by `policy.allow_browser`. Backed by the Zenrows Scraping
142-
Browser (CDP) and the `@zenrows/mcp` `browser_*` tools.
144+
Escalation only — **prefer `fetch`/`extract` for the vast majority of cases**;
145+
they cost less. Use the browser for logins, forms, and multi-step JS flows that
146+
Protected Fetch can't handle. Drives the managed Browser Sessions REST API
147+
(`mcp.zenrows.com/browser/sessions/*`, same backend as the `@zenrows/mcp`
148+
`browser_*` tools). `zenrows browser connect` prints the raw CDP wss URL for
149+
bring-your-own Playwright/Puppeteer.
150+
151+
**Billing & lifecycle:** sessions bill by **bandwidth + session time** and
152+
**auto-terminate after 15 minutes** — `run <script.json>` closes automatically;
153+
close interactive sessions with `zenrows browser close`. On by default; opt out
154+
with `zenrows policy set allow_browser false`.
143155

144156
## 11. MCP
145157

@@ -216,8 +228,8 @@ Reports write `input.json`, `results.json`, `report.md`, `failures.jsonl`,
216228
- `.zenrows/account.json` holds no secret — only the accountId, Free-period info, and
217229
claim link.
218230
- `.zenrows/policy.json` enforces credit/page/concurrency limits and domain allow/deny.
219-
- Destructive `uninstall` requires `--yes`. Browser and experimental are off by
220-
default.
231+
- Destructive `uninstall` requires `--yes`. Browser is on by default; opt out with
232+
`zenrows policy set allow_browser false`. Experimental features are off by default.
221233

222234
## 19. Capability matrix
223235

‎docs/capabilities.md‎

Lines changed: 7 additions & 13 deletions
Original file line numberDiff line numberDiff line change
@@ -7,8 +7,9 @@ the single source of truth that prevents hallucinated execution.
77

88
- **available** — a documented endpoint exists and the command does real work.
99
- **available-but-needs-confirmation** — likely available; verify per account.
10-
- **experimental** — exists but gated (e.g. browser, behind `policy.allow_browser`).
11-
- **beta** — real product in beta; limited access (local spec / validation works today, cloud execution needs beta access).
10+
- **experimental** — exists but not yet promoted to a stable status.
11+
- **beta** — open beta; usable by any key. Product-specific limits (e.g. Extract
12+
domain coverage) are handled by adapters / API errors, not by blocking the CLI.
1213
- **planned** — no documented endpoint yet; local spec / validation only.
1314
- **not-implemented** / **deprecated** — not usable.
1415

@@ -18,15 +19,8 @@ Classification is based on the public Zenrows documentation:
1819

1920
| Capability | Backend evidence | Status |
2021
| --- | --- | --- |
21-
| `protected_fetch` | Universal Scraper API `GET https://api.zenrows.com/v1/` with `mode`, `js_render`, `premium_proxy`, `proxy_country`, `wait`/`wait_for`, `js_instructions`, `response_type`, `screenshot`, `original_status`, … | available |
22-
| `extract` | Same `/v1/` endpoint via `autoparse`, `css_extractor`, `response_type=markdown\|plaintext` | available |
23-
| `batch` | Zenrows Batch Scraper API `https://async.api.zenrows.com/v1` (separate host, `X-API-Key` header) — real product in beta. Cloud subcommands (create/status/results/cancel/wait/retry-failed) work WITH beta access; without it the API returns 403 → `BATCH_ACCESS_DENIED`. Local JSONL spec validation + credit estimation work with no key. | beta |
24-
| `browser` | Zenrows Scraping Browser (CDP) + `@zenrows/mcp` `browser_*` tools; no managed REST sessions API | experimental |
22+
| `protected_fetch` | Fetch `GET https://api.zenrows.com/v1/` with `mode`, `js_render`, `premium_proxy`, `proxy_country`, `wait`/`wait_for`, `js_instructions`, `response_type`, `screenshot`, `original_status`, … | available |
23+
| `extract` | Extract `GET https://api.zenrows.com/v1/` via `extract=auto` (domain-gated open beta; CLI falls back to `autoparse`), plus `autoparse`, `css_extractor`, `outputs`, `response_type=markdown\|plaintext` | beta |
24+
| `batch` | Batch `https://async.api.zenrows.com/v1` (separate host, `X-API-Key` header) — open beta. Local JSONL spec validation + credit estimation work with no key. | beta |
25+
| `browser` | Browser Sessions REST API `https://mcp.zenrows.com/browser/sessions/*` (Bearer; same backend as `@zenrows/mcp` `browser_*`); CDP via `zenrows browser connect`. Escalation-only (prefer fetch/extract); on by default, opt out via `policy.allow_browser`; bills by bandwidth + session time (15-min max) | available |
2526
| `mcp` | Hosted `https://mcp.zenrows.com/mcp` + local `npx -y @zenrows/mcp` | available |
26-
27-
## Important honesty note
28-
29-
`protected_fetch` and `extract` are the **same** product: a single `/v1/`
30-
Universal Scraper API. "Extract" is not a separate endpoint — it is parameters
31-
on that endpoint (`autoparse` / `css_extractor` / `response_type`). The CLI keeps
32-
them as separate commands only for ergonomics.

‎evals/extract-smoke/README.md‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -1,9 +1,9 @@
11
# Eval: extract-smoke
22

3-
Reproducible smoke test for Extract (Autoparse).
3+
Reproducible smoke test for Extract (`extract=auto`).
44

5-
- **Target:** `https://www.scrapingcourse.com/ecommerce/` (public scraping demo)
6-
- **Config:** `autoparse=true`
5+
- **Target:** `https://www.owler.com/company/meltwater` (enabled domain, real company page)
6+
- **Config:** `extract=auto`
77
- **Success criteria:** HTTP 2xx and a non-empty body
88
- **Cost:** ~1× basic request
99

‎evals/extract-smoke/spec.json‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -1,10 +1,10 @@
11
{
2-
"description": "Smoke test: Extract returns structured output for a known-good ecommerce demo.",
2+
"description": "Smoke test: Extract (extract=auto) returns structured output for an enabled company page.",
33
"steps": [
44
{
55
"kind": "extract",
6-
"url": "https://www.scrapingcourse.com/ecommerce/",
7-
"options": { "method": "autoparse", "jsRender": false },
6+
"url": "https://www.owler.com/company/meltwater",
7+
"options": { "method": "extract", "jsRender": false },
88
"expect": { "minLength": 2 }
99
}
1010
]

‎registry/capabilities.json‎

Lines changed: 15 additions & 15 deletions
Original file line numberDiff line numberDiff line change
@@ -1,50 +1,50 @@
11
{
2-
"$comment": "Honest capability matrix derived from confirmed Zenrows docs. Every command checks status here before attempting a cloud call. 'available' = a documented endpoint exists today; 'planned' = no documented endpoint yet (local spec / unavailable behavior only); 'experimental' = exists but gated behind policy.",
2+
"$comment": "Capability matrix for the Zenrows CLI. Every command consults this file before attempting a cloud call, so the CLI never fakes behavior for primitives the backend does not expose. Status values and classification rationale live in docs/capabilities.md; entries must stay in sync with the Capability type in src/types/index.ts.",
33
"capabilities": {
44
"protected_fetch": {
55
"key": "protected_fetch",
66
"label": "Protected Fetch",
77
"status": "available",
88
"command": "zenrows fetch",
9-
"backend": "GET https://api.zenrows.com/v1/",
9+
"backend": "Fetch — GET https://api.zenrows.com/v1/",
1010
"requiresAuth": true,
11-
"notes": "Universal Scraper API. Confirmed params: mode=auto (Adaptive Stealth), js_render, premium_proxy, proxy_country, wait, wait_for, js_instructions, custom_headers, session_id, original_status, allowed_status_codes, block_resources, response_type, screenshot."
11+
"notes": "Supported params: mode=auto (Adaptive Stealth), js_render, premium_proxy, proxy_country, wait, wait_for, js_instructions, custom_headers, session_id, original_status, allowed_status_codes, block_resources, response_type, screenshot."
1212
},
1313
"extract": {
1414
"key": "extract",
15-
"label": "Extract (Autoparse / CSS / Markdown)",
16-
"status": "available",
15+
"label": "Extract (extract=auto / Autoparse / CSS / Markdown)",
16+
"status": "beta",
1717
"command": "zenrows extract",
18-
"backend": "GET https://api.zenrows.com/v1/ (autoparse, css_extractor, response_type)",
18+
"backend": "Extract — GET https://api.zenrows.com/v1/ (extract=auto, autoparse, css_extractor, outputs, response_type)",
1919
"requiresAuth": true,
20-
"notes": "Structured extraction runs on the same /v1/ endpoint via autoparse=true, css_extractor, and response_type=markdown|plaintext. There is no separate /extract endpoint."
20+
"notes": "Open beta. Default is extract=auto (domain-gated); CLI falls back to autoparse on AUTH010. Autoparse / CSS / outputs / markdown work on any domain."
2121
},
2222
"batch": {
2323
"key": "batch",
2424
"label": "Batch (beta)",
2525
"status": "beta",
2626
"command": "zenrows batch",
27-
"backend": "Batch Scraper API — https://async.api.zenrows.com/v1 (X-API-Key header; separate host from the scraper /v1/)",
27+
"backend": "Batch — https://async.api.zenrows.com/v1 (X-API-Key header; separate host from the Fetch/Extract /v1/)",
2828
"requiresAuth": true,
29-
"notes": "The Zenrows Batch Scraper API is a real product in beta. The cloud subcommands (create/status/results/cancel/wait/retry-failed) work WITH beta access; without it the API returns 403 → BATCH_ACCESS_DENIED. Local value always works with no key: `zenrows batch estimate` validates JSONL job specs and estimates credit cost."
29+
"notes": "Open beta. Cloud subcommands (create/status/results/cancel/wait/retry-failed) call the Batch API; `zenrows batch estimate` works locally with no API key."
3030
},
3131
"browser": {
3232
"key": "browser",
33-
"label": "Browser (Scraping Browser)",
34-
"status": "experimental",
33+
"label": "Browser Sessions",
34+
"status": "available",
3535
"command": "zenrows browser",
36-
"backend": "Scraping Browser (CDP) + @zenrows/mcp browser tools",
36+
"backend": "Browser Sessions REST API — https://mcp.zenrows.com/browser/sessions/* (Bearer auth; same backend as @zenrows/mcp browser_* tools)",
3737
"requiresAuth": true,
38-
"notes": "Zenrows Scraping Browser and the @zenrows/mcp browser_* tools exist. There is no managed REST 'sessions' API in the public docs, so this is gated as experimental and escalation-only (policy.allow_browser=false by default)."
38+
"notes": "Browser Sessions is a GA Zenrows product (formerly Scraping Browser). Drives its managed REST API directly (create/verb/close over HTTP with Authorization: Bearer) — no CDP client or browser dependency needed. Escalation-only: prefer fetch/extract (they cost less) for the vast majority of cases. On by default; opt out with policy.allow_browser=false. Sessions bill by bandwidth + session time and auto-terminate after 15 minutes. For raw CDP control, `zenrows browser connect` prints the wss://browser.zenrows.com endpoint for your own Playwright/Puppeteer."
3939
},
4040
"mcp": {
4141
"key": "mcp",
4242
"label": "MCP",
4343
"status": "available",
4444
"command": "zenrows mcp",
45-
"backend": "remote https://mcp.zenrows.com/mcp + local npx -y @zenrows/mcp",
45+
"backend": "Remote https://mcp.zenrows.com/mcp + local `npx -y @zenrows/mcp`",
4646
"requiresAuth": true,
47-
"notes": "Both a hosted remote MCP server and a local STDIO server (@zenrows/mcp, ZENROWS_API_KEY env) are documented."
47+
"notes": "Hosted remote MCP server plus a local STDIO server (`npx -y @zenrows/mcp`, authenticated via the ZENROWS_API_KEY environment variable)."
4848
}
4949
}
5050
}

‎registry/skills.json‎

Lines changed: 4 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -26,7 +26,7 @@
2626
"name": "extract",
2727
"type": "skill",
2828
"description": "Turn protected pages into structured data with Autoparse / CSS / Markdown.",
29-
"status": "available",
29+
"status": "beta",
3030
"requires_backend_capabilities": ["extract"],
3131
"requires_auth": true,
3232
"version": "0.1.0",
@@ -36,7 +36,7 @@
3636
{
3737
"name": "batch-jobs",
3838
"type": "skill",
39-
"description": "Scale protected fetch/extract over many URLs with the Batch Scraper API (beta): submit/track/collect jobs with beta access; validate + estimate specs locally with no key.",
39+
"description": "Scale protected fetch/extract over many URLs with Batch (beta): submit/track/collect jobs with beta access; validate + estimate specs locally with no key.",
4040
"status": "beta",
4141
"requires_backend_capabilities": ["batch"],
4242
"requires_auth": true,
@@ -47,8 +47,8 @@
4747
{
4848
"name": "interact-browser",
4949
"type": "skill",
50-
"description": "Escalate to a browser (Scraping Browser / MCP) only when fetch/extract cannot do the job.",
51-
"status": "experimental",
50+
"description": "Escalate to Browser Sessions (REST API / MCP browser_*) only when fetch/extract cannot do the job.",
51+
"status": "available",
5252
"requires_backend_capabilities": ["browser"],
5353
"requires_auth": true,
5454
"version": "0.1.0",

‎registry/templates.json‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -3,7 +3,7 @@
33
{
44
"name": "protected-fetch-node",
55
"type": "template",
6-
"description": "Minimal Node.js project calling the Zenrows Universal Scraper API.",
6+
"description": "Minimal Node.js project calling Zenrows Fetch.",
77
"status": "available",
88
"requires_backend_capabilities": ["protected_fetch"],
99
"requires_auth": true,
@@ -25,7 +25,7 @@
2525
{
2626
"name": "batch-jsonl-pipeline",
2727
"type": "template",
28-
"description": "JSONL job-spec scaffold for high-scale workloads on the Batch Scraper API: submit/track/collect with beta access, validate + estimate locally with no key (beta).",
28+
"description": "JSONL job-spec scaffold for high-scale workloads on Batch: submit/track/collect with beta access, validate + estimate locally with no key (beta).",
2929
"status": "beta",
3030
"requires_backend_capabilities": [],
3131
"requires_auth": false,

‎skills/batch-jobs/SKILL.md‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -1,6 +1,6 @@
11
---
22
name: batch-jobs
3-
description: Scale protected fetch/extract over many URLs via the Batch Scraper API (beta). Cloud create/status/results/cancel/wait/retry-failed work with beta access; estimate/validate run locally with no key.
3+
description: Scale protected fetch/extract over many URLs via Batch (beta). Cloud create/status/results/cancel/wait/retry-failed work with beta access; estimate/validate run locally with no key.
44
version: 0.1.0
55
requires_backend_capabilities: [batch]
66
---
@@ -11,7 +11,7 @@ Process large workloads reliably and asynchronously. Batch is where Zenrows'
1111
high-scale anti-bot advantage becomes obvious — Zenrows wins when the workflow
1212
runs over thousands, millions, or recurring sets of URLs.
1313

14-
> Status: **beta**. The Zenrows Batch Scraper API is a real
14+
> Status: **beta**. The Zenrows **Batch** is a real
1515
> product in beta and runs on a separate host
1616
> (`async.api.zenrows.com/v1`). The cloud subcommands work once your account has
1717
> beta access; without it the API returns 403 → `BATCH_ACCESS_DENIED`. The

‎skills/compliance-policy/SKILL.md‎

Lines changed: 4 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -15,16 +15,16 @@ zenrows policy show
1515
zenrows policy set max_credits_per_run 5000
1616
zenrows policy set blocked_domains "example.com,foo.test"
1717
zenrows policy set allowed_domains "mysite.com" # non-empty = allow-list mode
18-
zenrows policy set allow_browser true # enable escalation
18+
zenrows policy set allow_browser false # opt OUT of browser (on by default)
1919
zenrows policy set allow_experimental true
2020
```
2121

2222
## Hard rules
2323
- **Never** print or commit API keys. Keys live in `.zenrows/secrets.json`
2424
(0600, gitignored) and are redacted from logs and run artifacts.
25-
- Respect `allowed_domains` / `blocked_domains` (→ `POLICY_BLOCKED_DOMAIN`).
26-
- Stay under `max_credits_per_run` / `max_pages_per_run` / `max_concurrency`.
25+
- Respect `allowed_domains` / `blocked_domains` — enforced by the CLI (→ `POLICY_BLOCKED_DOMAIN`).
26+
- `max_credits_per_run` / `max_pages_per_run` / `max_concurrency` are **advisory budgets** surfaced by `zenrows status` (not hard-enforced by the CLI today) — self-limit against them.
2727
- Confirm destructive `uninstall` with `--yes`.
28-
- Browser and experimental commands are **off by default**.
28+
- **Experimental** commands are off by default (`allow_experimental`). **Browser is on by default** (opt out with `allow_browser=false`).
2929

3030
See [[cost-control]] and [[trace-debug]].

‎skills/cost-control/SKILL.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -9,7 +9,7 @@ requires_backend_capabilities: []
99

1010
Prefer the **cheapest reliable** configuration; escalate only with evidence.
1111

12-
## Cost multipliers (Universal Scraper API)
12+
## Cost multipliers (Fetch)
1313
- Basic request: **1×**
1414
- JS rendering (`js_render`): **5×**
1515
- Premium proxies (`premium_proxy`): **10×**

0 commit comments

Comments
 (0)