Skip to content

Latest commit

 

History

History
10 lines (8 loc) · 5.04 KB

File metadata and controls

10 lines (8 loc) · 5.04 KB

Changelog

Unreleased

  • instance-jobs now reports the upstream HTTP status inside failure_class: a failed DBLab call comes back as engine_error_404 / engine_error_500 rather than a flat engine_error, and a failed Joe call as joe_error_<code>. The Console showed an endless spinner when a deleted clone was opened, because the engine's 404 arrives as a failed job inside an HTTP 200 and the only way to tell it from a broken engine was to regex the English error text. The text itself is unchanged, and a failure with no HTTP status behind it (timeout, cancelled, engine/joe unreachable, bad args, oversize reply) keeps its existing flat class. (postgres-ai/platform-all#861)
  • postgresai promql accepts --project <id|alias> to pick the monitoring instance by project id, alias or name instead of its UUID, the same way joe and dblab do. It resolves through projects_list, so it needs a platform that reports monitoring_instance_ids there; on an older one it says so and asks for --instance. A project with several active monitoring instances is refused with their ids listed, rather than guessing which one you meant, and so is a name or alias that matches more than one project.
  • The instance-jobs container can now be turned on for a machine through a supported switch instead of a hand-run docker compose --profile. postgresai mon local-install --instance-jobs (or PGAI_INSTANCE_JOBS, or instance_jobs_enabled in the ansible role) writes COMPOSE_PROFILES=instance-jobs into the stack .env, which is what compose reads by itself — so mon stop / mon start / mon update / mon status all cover the container from then on, and mon health reports an absent one as a fault rather than as - not enabled. Passing nothing expresses no opinion, so an upgrade never opens or closes the channel, and any other profile listed in COMPOSE_PROFILES is preserved. Turning it off removes the container as well as the profile — taking the profile out of .env does not stop one that is already running. The platform's app.settings.instance_jobs_enabled is still a separate gate: with it off, a running container is handed no work.
  • Closed VictoriaMetrics' administrative endpoints on the Docker Compose monitoring stack. sink-prometheus now runs with -deleteAuthKey, -snapshotAuthKey, -forceMergeAuthKey and -pprofAuthKey, minted per install and stored in .env as VM_DELETE_AUTH_KEY, VM_SNAPSHOT_AUTH_KEY, VM_FORCE_MERGE_AUTH_KEY and VM_PPROF_AUTH_KEY. Grafana's datasource proxy filters POST but forwards every GET, and VictoriaMetrics required no key while these were unset, so any Grafana Viewer — including the pgai-aas-collect service account — could delete the whole metrics history through the proxy. -pprofAuthKey is part of the fix rather than extra hardening: /debug/pprof/* was reachable with the credentials the proxy already supplies, and /debug/pprof/cmdline served the process arguments, which is where the other keys used to live. Both halves are now closed — pprof requires its own key, and every secret is passed as file:// so none of them appears in the container command line, which is what /debug/pprof/cmdline and the host process table expose. The keys are still container environment variables, so docker inspect shows them; that needs access to the Docker socket, not a Grafana Viewer token. mon update and mon update-config add the keys to an existing .env and then bring sink-prometheus in line with the compose file, because the flags live on the container command line and mon restart (docker compose restart) re-runs the container exactly as recorded. That step is driven by the running container's config rather than by what the run happened to add, so a failed attempt is retried next time; it is skipped when the service is not running, and it uses --no-deps so it cannot re-seed a live-patched prometheus.yml. Nothing that reads metrics receives the keys. Scope: this covers the Docker Compose stack and the Terraform AWS deployment. The Helm chart and the preview stacks are unchanged and still affected.
  • Fixed index health checks recommending that built-in catalog indexes be dropped. H001 (invalid indexes), H002 (unused indexes) and H004 (redundant indexes) — and the unused_indexes, redundant_indexes, rarely_used_indexes and index_definitions pgwatch metrics behind them — no longer report indexes in pg_catalog, information_schema, pg_toast or per-backend temp schemas. Such indexes cannot be dropped, so recommending their removal was always wrong; pg_catalog.pg_class_tblspc_relfilenode_index was the case that surfaced it. Expect a one-off step drop in unused/redundant index counts and total sizes on the first report after upgrading: the catalog rows that were being counted are simply gone. Bloat (F004/F005) and wraparound (F002) still include catalogs on purpose.
  • Fixed express checkup on databases without the optional postgres_ai schema. F004/F005 now degrade with machine-readable status and warning summaries instead of reporting a misleading healthy empty result.