This file documents the current deployment approach for woozi/.
It should be kept aligned with the real working setup, not an aspirational target.
The app runs as four containers on the main host plus N stateless extraction workers on separate Hetzner Cloud VMs:
openbesluitvorming— HTTP surface (web/server.ts): search API, admin API, document preview, admin UIworker— import executor (src/worker.ts): polls SQLite for queued runs, runs extraction, writes to Quickwit + S3quickwit— search indexcaddy— reverse proxy + TLSextractionworkers (1..N remote hosts,services/extraction/) — stateless PDF extraction microservice
The HTTP container and the import worker use the same image with different entrypoints, sharing a named volume for SQLite state.
Object storage is external (Hetzner Object Storage in the current setup), configured via .env.
Vite is development-only and is not part of the production runtime.
The intended local development entrypoint is:
pnpm run devThis does:
- start Docker-backed services for
quickwitandopenbesluitvorming - run local Vite with HMR on the host
That gives:
- backend API on
http://127.0.0.1:8787 - frontend HMR on
http://127.0.0.1:4317
Important:
- local MinIO is not part of the default dev flow anymore
.envshould contain the real S3-compatible storage configuration- if
.envpoints at Hetzner Object Storage, the app uses that directly
The backend reads configuration from .env via src/config.ts.
Required S3-compatible storage values:
S3_ACCESS_KEYS3_SECRET_KEYS3_STORAGE_BUCKET_NAMES3_STORAGE_ENDPOINTS3_STORAGE_REGION
Common app/runtime values:
PORTQUICKWIT_URLQUICKWIT_INDEX_IDQUICKWIT_CLUSTER_IDQUICKWIT_NODE_IDQUICKWIT_INDEX_ROOT_PREFIXQUICKWIT_FAST_FIELD_CACHE_CAPACITYQUICKWIT_SPLIT_FOOTER_CACHE_CAPACITYQUICKWIT_PARTIAL_REQUEST_CACHE_CAPACITYWOOZI_KV_PATHINGEST_CONCURRENCYINGEST_MEMORY_PER_JOB_MBINGEST_MIN_FREE_MEMORY_MBQUICKWIT_BATCH_SIZE
Optional:
WOOZI_OPS_TOKENenables/api/ops/*(see Ops Endpoint)
The production image is built from Dockerfile.web.
It:
- builds the frontend with Node/Vite
- runs the backend with Deno
- installs
transmutationin the image for document extraction
The backend entrypoint is:
deno run -A web/server.tsRoutine production deploys are now image-based.
The server should be treated as runtime state only, not as a source checkout.
That means:
- application code should arrive via published container images
- the live server should not be treated as a Git worktree
- the live server should not be used as the normal place to build app images from source
- the only repo-managed files expected on the server are runtime config files such as Compose, Caddy, and Quickwit config
The current flow is:
- push a commit to
main - GitHub Actions builds and publishes the app image to GHCR
- the server pulls the selected image and restarts the app service
The workflow lives in:
The published image repository currently follows the GitHub repo owner. In the current setup that means:
ghcr.io/ontola/openbesluitvorming:mainghcr.io/ontola/openbesluitvorming:sha-<short-git-sha>ghcr.io/ontola/openbesluitvorming:latest
If the package owner changes, treat the repository path as configurable rather than hardcoded.
The current setup now supports automatic production deployment after a successful image publish on main.
The deploy-production.yml workflow:
- waits for
Publish OpenBesluitvorming Imageto finish successfully - checks out the exact published commit
- connects to the production server over SSH
- syncs the runtime config files (compose, Caddyfile, quickwit.yaml, monitor
script) and reloads Caddy —
/opt/wooziis a plain copy, not a checkout, so this step is what makes config changes in git actually land - deploys the exact
sha-<short-git-sha>image that was just published withWORKER_REPLICASworker replicas (default 1 — imports must keep running across deploys; a 0-default once silently froze imports for 11 days) - does not wait for imports: runs interrupted by the restart are requeued on startup (see reconciliation below)
That means the normal CD path is now:
- push to
main - GitHub Actions publishes the image
- GitHub Actions automatically deploys the published image to production
Required GitHub secret for this workflow:
BETA_DEPLOY_SSH_KEYprivate SSH key that can log intoroot@91.98.32.151
After CI has published the current commit image:
pnpm run deploy:productionThat script:
- resolves the current local Git commit SHA in the same short form GHCR publishes
- derives the GHCR image repository from the local
originremote by default - SSHes into the server
- runs
docker compose pull openbesluitvorming worker - restarts
openbesluitvorming,worker, andcaddy
Both openbesluitvorming and worker must be recreated on every deploy — they share the image and a code change to either process means both need the new image.
The script does not block on imports-in-progress. The daily scheduler enqueues ~290 runs every night and a handful are still in flight for most of the working day; making the deploy wait for idle would mean almost never being able to deploy. Reconcile-on-startup puts interrupted running rows back in the queue (capped at two requeues per run via interrupted_count, so a run that reliably kills the process can't crash-loop), and cache hits make the rerun cheap.
The script is:
Useful overrides:
DEPLOY_REF=<short-sha>deploy a specific already-published commit image, even if the local tree is dirtyDEPLOY_IMAGE=ghcr.io/<owner>/openbesluitvorming:<tag>deploy an explicit image tag directlyIMAGE_REPOSITORY=ghcr.io/<owner>/openbesluitvormingoverride the derived image repository
Production still depends on a small set of repo-managed runtime files on the server:
deploy-production.sh runs this sync automatically at the start of every deploy.
To push a config-only change without a code deploy, run it directly:
pnpm run deploy:production:infraThat helper is:
It rsyncs the three files, then runs caddy validate and caddy reload
inside the running Caddy container so a Caddyfile change takes effect
without a container recreate. The rsync uses --inplace because Caddyfile
and quickwit.yaml are bind-mounted as individual files; without it,
rsync's atomic-rename creates a new inode that the existing container
mount doesn't see, and the reload would no-op on the stale view.
If a Caddy restart is needed (e.g. a Caddyfile change that reload won't pick up, or a stuck cert state), recreate it explicitly:
ssh root@91.98.32.151 'cd /opt/woozi && docker compose -f docker-compose.production.yml up -d --no-deps --force-recreate caddy'So the operational split is:
- code changes: publish image, then
deploy:production(which also syncs config) - config-only changes without a new image:
deploy:production:infra
Operational rule:
/opt/woozion the server is runtime config, not an authoritative source tree- if an emergency manual server-side build is ever needed, do it from a clearly temporary sync/build path, not from a long-lived stale checkout
Imports are queued in-process and use a memory-aware concurrency limit.
Relevant env:
INGEST_CONCURRENCYhard upper bound for parallel importsINGEST_MEMORY_PER_JOB_MBestimated memory budget per active importINGEST_MIN_FREE_MEMORY_MBreserve memory that should stay free before another import starts
Current intended production behavior:
- allow more than one import when the server has enough free memory
- avoid the earlier unbounded fan-out that caused OOM crashes
Example safe starting point on the current cpx32 server:
INGEST_CONCURRENCY=4INGEST_MEMORY_PER_JOB_MB=1400INGEST_MIN_FREE_MEMORY_MB=1024
When using remote extraction workers, the memory per job is much lower (PDF extraction is offloaded), so INGEST_MEMORY_PER_JOB_MB can be reduced to 600 and INGEST_CONCURRENCY can be raised. Current production values: INGEST_CONCURRENCY=8, WOOZI_DOCUMENT_CONCURRENCY=10.
Additional env for extraction:
-
WOOZI_EXTRACTION_SERVICE_URLcomma-separated list of extraction worker URLs (e.g.http://10.0.1.3:8000,http://10.0.1.4:8000). When set, PDF extraction is sent to remote workers via HTTP instead of spawning local pymupdf4llm subprocesses. The ingest worker round-robins requests across workers, with a 180s per-request timeout and one retry onto a different worker before failing the document. When empty or unset, falls back to local subprocess extraction. -
WOOZI_DOCUMENT_CONCURRENCYnumber of documents to process concurrently per import (default: 3). With remote extraction workers this can be raised (e.g. 10-30) since PDF extraction no longer uses local CPU. -
WOOZI_IBABS_DATE_CHUNK_MONTHSwindow size for splitting large iBabs SOAP calls (default: 6). A multi-year import is split into N calls of this size because iBabs returns all meetings in one response and the XML can grow large enough to time out server-side.
Warning: setting INGEST_CONCURRENCY above 8 on a single worker container doesn't help — the worker is a single Deno process (one vCPU ceiling). Past that point the event loop saturates with outbound HTTP connections before it saturates CPU. Future scaling path: run multiple worker containers with an atomic claim in src/ops/store.ts so they can share the queue without racing.
If the server becomes unresponsive after restart, the likely cause is too many queued imports resuming simultaneously. Fix by stopping the container, resetting queued/running imports in SQLite to failed, and restarting with lower concurrency.
Quickwit projection writes are now streamed during extraction in small batches instead of one big end-of-run push.
Relevant env:
QUICKWIT_BATCH_SIZEnumber of entity commit events to buffer before ingesting to Quickwit
Quickwit config lives in quickwit/quickwit.yaml.
Important current behavior:
- Quickwit stores index data in S3-compatible object storage
- metastore is also in S3-compatible object storage
- local disk still matters for runtime state and temp work, but persistent index storage is remote
- search latency is therefore influenced by object-storage round trips unless searcher caches are configured well
So this stack is not "just Node" and it is not "just the app container":
openbesluitvormingserves search/admin/document APIsquickwithandles indexing/search- S3-compatible object storage holds document artifacts and Quickwit index data
The current production config now enables explicit Quickwit searcher caches.
Configured areas:
- in-memory fast field cache
- in-memory split footer cache
- in-memory partial request cache
- on-disk split cache under
data_dir
This matters because the production stack stores both the metastore and the index splits in S3-compatible object storage. Without these caches, repeated searches are much more exposed to:
- S3/object-storage latency variance
- repeated split downloads
- repeated fast-field reads for date filters and aggregations
Current default tuning for the cpx32 box:
QUICKWIT_FAST_FIELD_CACHE_CAPACITY=1GQUICKWIT_SPLIT_FOOTER_CACHE_CAPACITY=512MQUICKWIT_PARTIAL_REQUEST_CACHE_CAPACITY=128M
These are intentionally moderate values for an 8 GB VM that is also running the app container.
If search remains slow after caches are warm, the next things to look at are:
- moving from
cpx32tocpx42or a dedicatedccxinstance - separating heavy import activity from search traffic
- verifying that repeated queries are actually hitting the warm caches
Local and production must not share the same Quickwit metastore/index root in S3.
The repo now separates them by default:
- local/dev:
QUICKWIT_CLUSTER_ID=woozi-devQUICKWIT_NODE_ID=quickwit-devQUICKWIT_INDEX_ROOT_PREFIX=indexes-devQUICKWIT_INDEX_ID=woozi-events-dev
- production:
QUICKWIT_CLUSTER_ID=woozi-prodQUICKWIT_NODE_ID=quickwit-prodQUICKWIT_INDEX_ROOT_PREFIX=indexes-prodQUICKWIT_INDEX_ID=woozi-events-prod
This prevents a local Quickwit instance from writing into the same S3-backed metastore and index as production.
If production previously used:
- S3 prefix
indexes - index id
woozi-events
then switching production to:
- S3 prefix
indexes-prod - index id
woozi-events-prod
creates a fresh Quickwit projection namespace.
That is usually the right long-term move, but it means production search will be empty until Quickwit is rebuilt for the new location/index id.
Treat this as a projection migration:
- deploy the new env/config
- recreate or ensure the new index
- reindex/reimport into the new projection
Do not assume old search data under indexes/woozi-events will automatically appear under indexes-prod/woozi-events-prod.
search-v3-meeting-date maps start_date as a datetime fast field so search
can order and range-filter on the meeting date. Before it, start_date was a
dynamic text field: Quickwit could neither sort nor range on it, so results came
back in ingest order and date filters silently ignored meetings and motions
(issue #184).
This is a doc-mapping change, and doc mappings are fixed at index creation.
Editing quickwit/index-config.json does nothing to a live index —
QuickwitClient.ensureIndex returns early when the index id already exists, and
existing splits keep the mapping they were written with. Migrating means a new
index, not an upgraded one.
v3 is therefore opt-in: the code default is still search-v2-pages.
This paragraph used to claim that Quickwit ignores sort_by and a range clause
on a field its mapping does not declare, so v3 code against a v2 index would
behave exactly as before. That is wrong, and it took search down on
2026-08-08. start_date is not undeclared in a v2 index — it lands in the
dynamic field as a string, and 0.8.1 answers every query with
500 internal error: tantivy error: An invalid argument was passed:
'Unsupported sort field type `Str`.'
Both the sort and the range are therefore gated on the projection version
(projectionSupportsDateSort), not left to the engine to shrug off. Under v2
the pushdown is skipped and the date filter is applied app-side, exactly as it
was before v3. Do not remove that guard without a live query against a v2
index.
quickwit.yaml sets default_index_root_uri, and it decides this per index at
creation time. It said file:///quickwit/qwdata/… while the live index had been
created against S3, so every index made after that wrote its splits to the
root disk instead of object storage.
That is what killed the first v3 attempt (2026-08-10). The root disk lost ~2 GB
per minute, and the diagnosis went through two wrong causes — the searcher cache
(correctly on its own 147 GB volume) and indexer staging (which grew by 234 MB,
not gigabytes) — before du on indexes-prod showed it going from 168 KB to
8 GB. The published index is ~65 GB against ~60 GB of free root disk, so the
reindex could never have finished at any speed or concurrency.
It now points at s3://${S3_STORAGE_BUCKET_NAME}/${QUICKWIT_INDEX_ROOT_PREFIX},
matching where the live index already is. Two things to know:
- A Quickwit upgrade is one-way for the metastore. 0.9.0 rewrites
indexes-prod/manifest.jsonand everymetastore.jsonto format0.9at startup; 0.8.1 refuses to start on them afterwards. Copy the wholeindexes-proddirectory out of the volume before changing the tag, and restore it to roll back. - The deploy restarts Quickwit only when its container definition changed.
deploy-production.shbrings upopenbesluitvorming worker caddy otel-collector, but the web containerdepends_onQuickwit, so a changed image tag or environment recreates it as part of that deploy (seen 2026-09-02 with the 0.9.0 tag). A syncedquickwit.yamlis a mounted file and does not count: it sits unread until Quickwit is restarted by hand. Either way search is unavailable for a few seconds. - Existing indexes keep the URI they were created with. This changes nothing
for
woozi-events-prod; it only decides where the next index goes. Verify withGET /api/v1/indexes/<id>and look atindex_uribefore trusting a rebuild.
ensureIndex posts quickwit/index-config.json verbatim, so on its own every
new index inherits the node's default_index_root_uri (S3 on production) and
the file's commit_timeout_secs: 1. Three environment variables, read by the
container that creates the index, override that at creation time and are
ignored for an index that already exists:
| variable | effect |
|---|---|
QUICKWIT_INDEX_URI |
sets index_uri, e.g. file:///quickwit/qwdata/indexes-prod/woozi-events-v4 |
QUICKWIT_COMMIT_TIMEOUT_SECS |
sets indexing_settings.commit_timeout_secs |
QUICKWIT_INGEST_COMMIT |
auto, wait_for (default) or force on every ingest request |
On production the worker takes them from WORKER_QUICKWIT_INDEX_URI,
WORKER_QUICKWIT_COMMIT_TIMEOUT_SECS and QUICKWIT_INGEST_COMMIT in
/opt/woozi/.env. Verify with GET /api/v1/indexes/<id> after the first
worker run: index_uri and commit_timeout_secs are what the index will keep.
wait_for and a sane commit timeout do not mix for bulk work: every batch then
waits the full timeout (measured 60s per batch at 60, on 0.9.0, regardless of
how many batches are in flight). Reindex campaigns want auto.
Switching both containers at once empties search for as long as the reindex
takes, and that is not minutes. Measured 2026-08-08: 5.48M live entities
across 330 sources, of which ~57% carry stored text, so the reindex performs
roughly 3.1M object-storage reads — one per document, sequentially within
a source (reindexSource awaits each rehydrateDocumentText).
Because the worker and the web container take their index and projection from separate variables, the reindex can run against the new index while the old one is still being served:
-
create the new index by pointing only the workers at it:
# /opt/woozi/.env WORKER_QUICKWIT_INDEX_ID=woozi-events-prod-v3 WORKER_PROJECTION_VERSION=search-v3-meeting-datethen restart the workers.
ensureIndexcreates it with the v3 mapping. -
reindex_onlyevery source. It replays from the export log, never from the supplier APIs, andstart_dateis already in the stored payloads. -
verify the new index: row counts per entity type, and a query that sorts and date-filters.
-
switch the reader by setting
QUICKWIT_INDEX_IDandWOOZI_PROJECTION_VERSIONto the same values and restarting the web container. That is the only moment users notice anything. -
keep the old index until you are satisfied, then delete it.
The cost of doing it this way: between step 1 and step 4, freshly imported material lands only in the new index, so new meetings are not searchable until the switch. That is a far smaller hole than everything being unsearchable.
Search is scoped to projection_version, so the old index's documents cannot
leak into new results even while both exist.
Two things to know before touching the mapping again:
- A datetime field that fails to parse silently drops the whole document.
Not the field — the document. Quickwit 0.8.1 accepts the ingest, reports
success, and indexes nothing, so the entity vanishes from search with no
error anywhere.
toIndexDateTimeinsrc/quickwit/project.tsnormalizes and gives up toundefinedfor exactly this reason; anything writing a mapped datetime field must do the same. - Sort direction is inverted from the usual convention. A bare field name
in
sort_bysorts descending; a-prefix sorts ascending. Documents missing the field sort last in both directions.
Current tested provisioning path:
- provider: Hetzner Cloud
- image:
docker-ce - location:
fsn1 - reason: closest practical Hetzner region for Amsterdam
Current server created:
- name:
woozi-1 - server type:
cpx32 - location:
fsn1 - image:
docker-ce - ipv4:
91.98.32.151
DNS note:
- DNS is managed in Netlify
- the production domain
Arecord should point to91.98.32.151
SSH:
ssh root@91.98.32.151For this stack, the practical starting recommendation is:
cpx324 vCPU8 GB RAM160 GB local disk
This is a reasonable first deployment size for:
- one app container
- one Quickwit container
- external Hetzner Object Storage
If search volume or indexing pressure grows, move to:
cpx42- or dedicated
ccxtypes if you want more predictable CPU
The Hetzner CLI (hcloud) was configured with context:
woozi
Provisioning steps used:
- create/upload SSH key
- choose
fsn1 - create Docker CE app server
- use current
cpx32type becausecpx31was not orderable infsn1
Example commands:
hcloud ssh-key create --name joepio-ed25519 --public-key-from-file ~/.ssh/id_ed25519.pub
hcloud server create --name woozi-1 --type cpx32 --location fsn1 --image docker-ce --ssh-key joepio-ed25519The production shape is:
- one Hetzner Cloud VM (
woozi-1) for the app, Quickwit, and Caddy - optional extraction worker VMs for PDF extraction during imports
- external Hetzner Object Storage
Caddyin front for TLS termination and reverse proxy
The production compose file should use a published app image, not build on the server.
Production traffic:
internet
-> Caddy (:80 / :443)
-> openbesluitvorming (:8787, private to the VM)
-> quickwit (:7280, private to the VM)
During imports:
openbesluitvorming -> extraction workers (:8000, separate VMs)
Quickwit should not be exposed publicly.
PDF extraction can be offloaded to a dedicated server running the services/extraction/ Docker image. This is a stateless FastAPI service wrapping pymupdf4llm with multiple uvicorn worker processes.
Infrastructure is managed with OpenTofu in infra/.
Instead of many small cloud VMs, a single dedicated Hetzner server gives more CPU per euro:
| Server | Cores/Threads | RAM | Price/hour | Price/month |
|---|---|---|---|---|
| cpx22 (cloud) | 2 vCPU | 4 GB | €0.01 | €7.50 |
| AX41-NVMe | 6c/12t | 64 GB | €0.06 | €38 |
| EX44 | 14c/20t | 64 GB | €0.07 | €44 |
One AX41-NVMe (€0.06/hr) with 12 uvicorn workers replaces 6× cpx22 VMs (€0.06/hr total) and is simpler to manage.
To spin up the extraction server for an import:
cd infra
TF_VAR_hcloud_token="..." tofu apply \
-var="extraction_server_enabled=true" \
-var="extraction_server_type=cpx22" \
-var="extraction_uvicorn_workers=2"For production imports with a dedicated server:
TF_VAR_hcloud_token="..." tofu apply \
-var="extraction_server_enabled=true" \
-var="extraction_server_type=ax41-nvme" \
-var="extraction_uvicorn_workers=12"To tear down after import:
TF_VAR_hcloud_token="..." tofu apply -var="extraction_server_enabled=false"After provisioning, set WOOZI_EXTRACTION_SERVICE_URL on the ingest server to the extraction server URL from tofu output extraction_service_url_env. The worker firewall only allows port 8000 from the ingest server's IP (91.98.32.151).
The extraction service image is published to GHCR via .github/workflows/publish-extraction.yml.
The admin panel shows extraction server status (health, CPU load, request count) in the "Extractie-workers" section when WOOZI_EXTRACTION_SERVICE_URL is configured.
See docs/scaling-pdf-extraction.md for detailed analysis and architecture options.
The current recommended HTTPS setup is:
Caddy- automatic Let's Encrypt certificates
- reverse proxy from the public domain to
openbesluitvorming
This is preferred over nginx + certbot because:
- less setup
- automatic certificate issuance
- automatic renewal
- simpler config for a single app deployment
The practical rule is:
- expose
80and443 - point each public domain's
Arecord at the Hetzner server - let Caddy terminate TLS
- keep the app itself listening on
127.0.0.1:8787or a Docker-internal address
The production stack serves the apex from a single Caddy site block:
openbesluitvorming.nl(apex, added 2026-05-07)
beta.openbesluitvorming.nl was the original name and was dropped on
2026-07-28, when DNS moved from Netlify to Openprovider and the new zone was
created with only the apex A record. A site block for a name that no longer
resolves is not free: Caddy keeps retrying the ACME challenge for it and can
get rate-limited, so the hostname has to go from here too, not just from DNS.
Planned next:
zoek.openraadsinformatie.nl(alternate brand, not yet served — needs DNSArecord at91.98.32.151before adding to the Caddyfile, otherwise the ACME HTTP-01 challenge will fail and Caddy may rate-limit retries)
Adding a new domain is a Caddyfile change committed to the repo, then
pnpm run deploy:production:infra to sync + validate + reload. There is no env-var
indirection for the site addresses anymore — they are listed verbatim in the
Caddyfile so the canonical list lives in version control.
openbesluitvorming.nl {
encode gzip zstd
reverse_proxy openbesluitvorming:8787
}That implies Caddy joins the same Docker network as the app container.
The repo now includes:
That production compose file:
- exposes only
80and443publicly through Caddy - keeps
openbesluitvorminginternal to Docker viaexpose: 8787 - keeps Quickwit private
- expects
ADMIN_PASSWORD_HASHand the normal S3 env vars in.env(the oldDOMAINenv var was removed once the Caddyfile started listing the public hostnames directly)
For production on Hetzner Cloud:
- use the Docker CE marketplace image
- keep Vite out of the runtime entirely
- run
openbesluitvorming,worker,quickwit, andcaddyas the deployed stack - use the real
.envfor Hetzner Object Storage - do not start local MinIO
The expected public endpoint is the Caddy domain, not port 8787 directly.
Current preferred deployment flow is image-based from the local machine:
pnpm run deploy:productionManual recovery on the server should normally be:
docker compose -f docker-compose.production.yml pull
docker compose -f docker-compose.production.yml up -dRequired .env values for that production compose file:
ADMIN_PASSWORD_HASHS3_ACCESS_KEYS3_SECRET_KEYS3_STORAGE_BUCKET_NAMES3_STORAGE_ENDPOINTS3_STORAGE_REGION
Public hostnames are listed directly in the Caddyfile, not in .env.
First DNS step (when adding a new public hostname):
- point the new domain's
Arecord at the server IP (91.98.32.151) - DNS for
*.openbesluitvorming.nlis managed in Netlify; other zones (e.g.openraadsinformatie.nl) are managed elsewhere — confirm with the zone owner before assuming Netlify - wait for DNS to resolve, then commit the new hostname into the
Caddyfile and run
pnpm run deploy:production:infra— adding it before DNS is live causes Caddy's ACME HTTP-01 challenge to fail and Let's Encrypt may rate-limit further attempts
The current recommended protection for the admin UI is:
- Caddy HTTP Basic Auth
- only on
/admin,/admin.html, and/api/admin/*
Public search/document routes remain open.
Required env:
ADMIN_PASSWORD_HASH
That value should be a Caddy password hash, not a plaintext password.
Because Docker Compose reads .env, bcrypt dollar signs must be escaped there as $$.
Example generation:
docker run --rm caddy:2 caddy hash-password --plaintext 'your-strong-password'Example .env value:
ADMIN_PASSWORD_HASH=$$2a$$14$$exampleexampleexampleexampleexampleexampleexampleexample/api/ops/* lets an operator (or an agent acting for one) check imports and
start a fixed set of maintenance actions without a shell on the server. It is
not part of the public API contract and is deliberately absent from API.md.
It is separate from the admin protection above: Caddy passes /api/ops/*
straight through, and the app authenticates it with a bearer token from
WOOZI_OPS_TOKEN (compared in constant time). When the variable is unset or
empty, every path under /api/ops/ answers 404.
Generate a token (hex only, so no $ escaping is needed in .env):
openssl rand -hex 32Add it to /opt/woozi/.env on the production host:
WOOZI_OPS_TOKEN=<the generated value>and recreate the web container so it picks it up:
ssh root@91.98.32.151 'cd /opt/woozi && docker compose -f docker-compose.production.yml up -d --no-deps openbesluitvorming'To rotate, replace the value and recreate the container again; to disable, remove the line. Keep the token out of shell history and chat logs; it grants the purge and takedown actions below.
Every request carries Authorization: Bearer <token>, and should carry
X-Ops-Actor: <who> (free text, stored with jobs and logged; defaults to
unknown).
Reads answer directly:
| Request | Returns |
|---|---|
GET /api/ops/health |
host load, memory and state-volume disk; extraction workers; Quickwit readiness and index counters; backup age; run and job queues; compose services (see below) |
GET /api/ops/runs?source=&status=&limit=&offset= |
import runs, newest first ({ runs, hasMore }) |
GET /api/ops/summary |
the run summary the admin dashboard shows |
GET /api/ops/runs/<id> |
one run with its issues |
GET /api/ops/jobs?status=&limit=&offset= |
ops jobs, newest first ({ jobs, hasMore }) |
GET /api/ops/jobs/<id> |
one job, including its captured output |
Mutating actions are POST /api/ops/<action> with a JSON body:
| Action | Body | Does |
|---|---|---|
rerun_source |
source, mode (full default, or reindex_only), dateFrom/dateTo (YYYY-MM-DD, required for full, forbidden for reindex_only) |
enqueues one import run, like "Opnieuw draaien" in the admin UI; only sources with implemented: true |
reenqueue_failed_windows |
optional source, statuses (["failed","partial"] default), minWindowDays (20), fromYear, toYear |
same as scripts/reenqueue_failed_windows.ts |
purge_source |
source, optional quickwit (bool), keepStorage (bool) |
same as scripts/purge_source.ts |
delete_document |
entityIds (1 to 100 document entity ids), optional reason (short label, takedown default, e.g. bsn) |
same as scripts/delete_document.ts: delete markers and a delete task in Quickwit, the document's objects, an export tombstone, and a blocklist entry |
restart_service |
service: worker, openbesluitvorming, otel-collector or quickwit |
docker compose restart <service>, run by the host agent (below); never Caddy |
service_logs |
service (the same, plus caddy), optional sinceMinutes (60, at most 1440), lines (200, at most 2000) |
the service's newest log lines, with timestamps, as the job's output; read-only, so no apply |
Every action is a dry run unless the body has "apply": true and
"confirm" equal to the source key (or "all" for a re-enqueue without a
source). A delete_document is confirmed with the entity id when it names
one document, and with "<n> documents" (e.g. "3 documents") when it names
several. A restart_service is confirmed with the service name. A dry run still goes through the worker and its output shows what
would happen.
A valid request answers 202 with the queued job. The web container does not
execute it: a worker (src/worker.ts) claims it from the ops_job table in
the ops SQLite, runs it next to its imports (one job at a time per worker),
and writes the output lines into the row. Poll GET /api/ops/jobs/<id> until
status is succeeded or failed. Only one apply job can be queued or
running at a time; a second one gets 409 with activeJobId. A job
interrupted by a worker restart is requeued once and failed the second time;
a deploy's SIGTERM hands it back to the queue directly. Every action is
idempotent, but a requeued rerun_source whose run was already created fails
with "already queued", which is the correct outcome.
Example, dry run then apply:
curl -sS -X POST https://openbesluitvorming.nl/api/ops/purge_source \
-H "Authorization: Bearer $WOOZI_OPS_TOKEN" -H "X-Ops-Actor: joep" \
-H "content-type: application/json" \
-d '{"source":"waterschap_limburg"}'
curl -sS -X POST https://openbesluitvorming.nl/api/ops/purge_source \
-H "Authorization: Bearer $WOOZI_OPS_TOKEN" -H "X-Ops-Actor: joep" \
-H "content-type: application/json" \
-d '{"source":"waterschap_limburg","apply":true,"confirm":"waterschap_limburg"}'GET /api/ops/health answers in one request, each section on its own so a
failing dependency shows as { "error": ... } in its section only:
host: load average, CPU count, memory and swap, and free space on the state volume. Read from inside the web container, which shares the host kernel, so these are the production host's figures.extractors: each extraction worker's own/stats(load, free disk, request counters), orunreachable. Same data as/api/admin/extractors.quickwit:readyfrom/health/readyz, plus published docs, splits and size of the served index.backup: when the last state backup completed (its stamp file) and its age.imports: queued and running runs, the oldest queued run, the last claim and last finished full run, and queued/running ops jobs.services: every compose service's state, health, status line and replica count as the host agent last saw it, withagent_seen_at.agent_staleis true when the agent has not ticked for a minute (or was never installed).
Restarting a service and reading its logs need the Docker daemon, and no
container gets the Docker socket: with it, the ops token would be worth as
much as root on the host. Those two actions are run instead by
scripts/ops_host_agent.py, a standard-library Python script that a systemd
timer starts every 15 seconds on the production host. Each tick it records
docker compose ps into host_service_status (the services section of
health), fails host jobs left running for 15 minutes by a dead agent, and
claims queued restart_service / service_logs jobs from ops_job. The
Deno worker never claims those two. The agent re-checks every parameter
against its own allow-list rather than trusting the row, and runs nothing but
docker compose ps, restart <service> and logs <service> in
/opt/woozi. A host job is not retried when the agent dies mid-job.
Install once (deploys keep the script itself in sync through
deploy-production-infra.sh):
scripts/install-production-ops-agent.shCheck it with systemctl list-timers woozi-ops-agent.timer and
journalctl -u woozi-ops-agent.service. To stop it:
systemctl disable --now woozi-ops-agent.timer.
Logs can hold search terms from request paths and personal data from
documents being imported. service_logs output is stored in the job row like
any other output, so it lands in the ops SQLite and its daily backup; keep
lines and sinceMinutes to what the question needs.
Out of scope on purpose: arbitrary scripts, catalog edits,
enqueue_full_history and Quickwit index management. Those still need a shell
on the host. A takedown through delete_document follows the same runbook as
the script (docs_internal/): the endpoint only replaces the shell, not the
review of whether a document has to go.
The endpoint is excluded from the public rate limiter and has its own: 30
requests per minute for the token, and a separate 30 per minute per client
address for failed authentication. Over budget answers 429 with
Retry-After.
Every request writes one JSON line to the web container's log with
event: "ops_request", method, path, actor, client address, outcome
(ok, queued, missing_token, invalid_token, rate_limited,
bad_request, conflict, not_found, disabled, error), status and,
for actions, the action, apply and job id. The token is never logged. To
review:
ssh root@91.98.32.151 'cd /opt/woozi && docker compose -f docker-compose.production.yml logs openbesluitvorming --since 24h | grep ops_request'Three systemd timers run on the production host:
woozi-monitor.timer(every 2 min) runs scripts/monitor-production.sh — the bash variant is the deployed one; it is synced by every deploy. Checks: search latency/errors, disk, container state, import health (no completed run in 26h, queued work with nothing running for 30 min, extraction-failure surges — all read straight from the ops SQLite on the host so they fire even when the worker container is gone entirely), and backup freshness. Alerts go to the Discord webhook in/opt/woozi/.env(WOOZI_ALERT_WEBHOOK_URL), with a 15-min cooldown per alert key. During an intentional worker scale-down, setWOOZI_MONITOR_EXPECT_WORKER=0. Install/refresh with scripts/install-production-monitor.sh.woozi-backup.timer(daily 03:30) runs scripts/backup_state.ts inside the web container: consistentVACUUM INTOsnapshots ofwoozi-ops.sqlite3(run admin + document blocklist) andwoozi-export-log.sqlite3(export seq/dedup state), gzipped tobackups/sqlite/{name}/{date}.sqlite3.gzin S3 with 14-day retention. Losing these databases without a backup resurrects taken-down documents and corrupts export cursors. Install with scripts/install-production-backup.sh.woozi-revalidate.timer(daily 02:00) runs scripts/revalidate_documents.ts inside the web container, once per calibrated supplier (iBabs, Notubiz): checks whether documents we still serve have actually been removed at the source (a cheap HTTP status check against the storedoriginal_url, no re-download), so a source-side deletion eventually gets reflected here even when nobody files a takedown request for it. It never deletes anything — only entity ids that come back "gone" for several consecutive daily runs are reported (docker exec woozi-openbesluitvorming-1 deno run -A scripts/revalidate_documents.ts --supplier ibabs --report-only) for manual review and deletion viascripts/delete_document.ts. See docs_internal/bsn-takedown.md for the calibration details and the ori3 equivalent. Install with scripts/install-production-revalidate.sh.- Coverage check (
woozi-coverage.timer, weekly, Sunday 05:00): for every runnable source,scripts/coverage_check.tslists the document ids the supplier's own API exposes for the last 12 months -- meeting documents and register/motion attachments -- and compares them with the export log. The result goes tocoverage_checkin the ops SQLite and appears in/api/statusascoverageper source (supplier count, held, missing, a sample of missing ids). Nothing is downloaded. It shares the suppliers' request budgets with the nightly import, so it runs one source at a time on a quiet morning; a full pass takes a few hours. Run one source by hand withdocker exec woozi-openbesluitvorming-1 deno run -A scripts/coverage_check.ts --source ermelo --dry-run. Install with scripts/install-production-coverage.sh. - Sitemaps (
woozi-sitemaps.timer, daily, 09:30):scripts/generate_sitemaps.tswalks the export log per source and writes the meetings and documents of the last 12 months, withlastmodfrom the log, as one sitemap per organization (split at 50,000 addresses) plus an index, to object storage undersitemaps/. The web container serves them as/sitemap.xmland/sitemaps/<name>.xmlandrobots.txtpoints to the index. Crawlers get the address of every recent stuk instead of finding them through search results. The job parses every live meeting and document record of every source, so the first run says how long it takes; it only reads the log, which is safe next to the workers. An organization that drops out keeps stale files in storage but leaves the index. Run one source by hand withdocker exec woozi-openbesluitvorming-1 deno run -A scripts/generate_sitemaps.ts --source ermelo --dry-run. Install with scripts/install-production-sitemaps.sh; until it has run once/sitemap.xmlanswers 404.
To restore: download the newest backups/sqlite/... object, gunzip, stop the
openbesluitvorming and worker containers, replace the file on the
woozi-state volume (remove stale -wal/-shm siblings), start containers.
The monitor script is the paging layer; diagnosis and trends live in SigNoz
Cloud (eu2). An otel-collector container (otel/collector.yaml,
synced by every deploy) receives OTLP from the Deno containers — OTEL_DENO=1
exports traces, metrics and console logs natively, no code changes — and
scrapes host metrics, including per-process open fd counts (the July 2026
worker socket-leak signal). It ships everything to
ingest.eu2.signoz.cloud:443 using SIGNOZ_INGESTION_KEY from
/opt/woozi/.env. Disable app telemetry with OTEL_DENO=0 in that .env.
Losing the collector never pages: it is not in the monitor's critical path.
Satellite nodes (extraction workers, ori3) run a hostmetrics-only collector
(otel/node-collector.yaml) so the whole fleet is
visible in SigNoz's Infrastructure view. New extraction hosts get it from
cloud-init when TF_VAR_signoz_ingestion_key is set at apply time.
- interrupted imports can leave stale run-state unless reconciled on startup
- startup reconciliation for interrupted runs is now implemented in src/ops/store.ts
- duplicate active imports for the same source/date/execution mode are blocked in src/ingest.ts
- background imports are now queued in SQLite and executed by the
workercontainer (src/worker.ts), separate from the HTTP-servingopenbesluitvormingcontainer - the current production behavior is memory-aware concurrency, not a fixed
1 - large aggregate imports should therefore mostly appear as many
queuedruns plus a bounded number ofrunningruns - admin now has a queue/status summary backed by
/api/admin/summary - startup requeues previously
runningimports (interrupted by a restart), capped at two requeues per run; a run that keeps dying then fails for good - aggregate imports are now supplier-specific, such as
__supplier__:notubizand__supplier__:ibabs - the deploy helper does not wait for running imports; interrupted runs are requeued on startup
- browser-side fetch helpers now handle empty/non-JSON 500 responses more safely
- resuming many queued imports on startup with high concurrency can make the server unresponsive (event loop saturated with HTTP connections). If this happens: stop the container, reset queued/running rows in SQLite to
failed, restart with lower concurrency - every outbound fetch on the ingest path has an explicit timeout. Observed failure mode: iBabs/Notubiz/Quickwit/extraction workers occasionally hold a TCP connection open without responding, and a slot hangs for hours — in one incident 7 of 8 slots were wedged for 10+ hours. All fetches now cap at 90-180s with retries on
TimeoutError/AbortError. Any newfetch()call in the ingest path must keep this pattern. - iBabs is IPv4-whitelisted. The client forces IPv4-first DNS resolution via
node:dns.setDefaultResultOrder("ipv4first"). On a dual-stack host without this, iBabs requests silently try IPv6 first and get rejected. - the admin dashboard polls every 5s. It used to refetch per-run issue details for every run in the list (50+ parallel requests/poll), which pegged the single-threaded
openbesluitvormingprocess during big imports and slowed user searches to 18s+. Now it only refetches when a run'sissue_counthas grown. Keep dashboard work per poll small.
Update this file whenever any of the following change:
- cloud provider or region
- server type
- object storage strategy
- Docker/runtime topology
- production entrypoint
- reverse proxy choice
- one-command dev workflow
- extraction worker topology or scaling approach
- Hetzner account limits