A drop-in Prometheus
/metricstemplate for Python, Node.js, and Go web apps. Oneinstall()call; consistent metric catalogue; zero-per-app dashboard work.
This repository ships three packages sharing one metric catalogue, label
conventions, and cardinality rules — so a single $service-templated
Grafana dashboard works across every runtime:
| Package | Path | Languages | Tag prefix | Install |
|---|---|---|---|---|
simsys-metrics (Python) |
/ (root) |
FastAPI, Flask | python-v<semver> (e.g. python-v2.0.0) |
pip install simsys-metrics (PyPI) |
@simsys/metrics (Node) |
node/ |
Express 5, Bun + Hono | node-v<semver> (e.g. node-v2.0.0) |
npm install @simsys/metrics (npm) |
simsys-metrics-go |
go/ |
net/http | go/v<semver> (e.g. go/v2.0.1) |
go get github.com/Simmons-Systems/simsys-metrics/go/v2@v2.0.1 |
The Python package remains at the repo root for pip git-install compatibility. The Node and Go packages live under node/ and go/ respectively — see each subdirectory's README for install details.
A drop-in observability layer for any Python web app. Adding baseline
metrics is a five-line job, and every emitted series carries the same
service label so a generic $service-templated Grafana dashboard lights
up automatically.
- Supported stacks: FastAPI (primary), Flask (secondary).
- Baseline metrics (zero extra code): HTTP request count + latency, process CPU/RSS/FDs, build info.
- Opt-in helpers: queue depth gauge, job counter + histogram.
- Cardinality discipline enforced by the package: route templates,
status classes, allow-listed label helper. Every metric name must start
with
simsys_— the registry refuses anything else.
- Install
- Usage
- Metric catalogue
- One service identity per process
- Cardinality rules
commitdetection/metricsendpoint behaviour- Development
- Contributing
- License
# FastAPI service
pip install "simsys-metrics[fastapi]"
# Flask service
pip install "simsys-metrics[flask]"Pin to the tag. Bumping a consumer means re-pointing this URL at a newer tag.
Pinning in requirements.txt
simsys-metrics[fastapi]==2.0.0
Works in plain Docker builds — no SSH agent, no auth tokens required.
from fastapi import FastAPI
from simsys_metrics import install, track_queue, track_job
app = FastAPI()
install(app, service="my-api", version="1.2.3")
# Opt-in queue gauge (polled every 5s in a daemon thread)
track_queue("inference", depth_fn=lambda: job_queue.qsize())
# Opt-in per-job timer — as a decorator
@track_job("inference")
def run_inference(...): ...
# ...or as a context manager
def run(...):
with track_job("inference"):
...
# Opt-in batch-progress tracking (v0.2.0+)
from simsys_metrics import track_progress, ProgressOpts
tracker = track_progress(ProgressOpts(operation="scan", total=input_count))
try:
for item in work:
process(item)
tracker.inc()
finally:
tracker.stop()from flask import Flask
from simsys_metrics import install
app = Flask(__name__)
install(app, service="my-worker", version="0.4.1")install() auto-detects the framework. It sets the process-wide service
label, registers the simsys process collector, wires HTTP request metrics,
mounts /metrics, and populates simsys_build_info.
Default metrics_path is /metrics; override with
install(..., metrics_path="/internal/metrics").
Coerce any user-facing label value into a bounded allow-list:
from simsys_metrics import safe_label
ticker = safe_label(request.args.get("ticker"), {"AAPL", "GOOG", "NVDA"})
# -> "AAPL" if in the set, else "other"Use this for anything an external caller controls (tickers, tenant names, device IDs, free-form search terms) before it ends up as a Prometheus label.
Every baseline + opt-in metric this package emits uses the simsys_
prefix AND a service label — so cross-service PromQL like
sum by (service) (rate(simsys_http_requests_total[5m])) works
unmodified across every app.
Custom metrics convention: the
make_counter/make_gauge/make_histogramfactories enforce thesimsys_prefix but do not forceserviceinto your label list. To stay compatible with the cross-service dashboards above, includeserviceinlabelnamesfor any custom metric you create. The factories will warn at registration time ifserviceis missing — see the example in Rich outcome taxonomies below.
| Metric | Type | Labels | Runtimes | Tier | Source |
|---|---|---|---|---|---|
simsys_build_info |
Gauge = 1 | service, version, commit, started_at |
Py Node Go | core | baseline |
simsys_http_request_duration_seconds |
Histogram | service, method, route |
Py Node Go | core | baseline |
simsys_http_requests_total |
Counter | service, method, route, status |
Py Node Go | core | baseline |
simsys_job_duration_seconds |
Histogram | service, job, outcome |
Py Node Go | core | opt-in |
simsys_jobs_total |
Counter | service, job, outcome |
Py Node Go | core | opt-in |
simsys_pool_active |
Gauge | service, pool |
Py Node Go | core | opt-in |
simsys_pool_idle |
Gauge | service, pool |
Py Node Go | core | opt-in |
simsys_pool_max |
Gauge | service, pool |
Py Node Go | core | opt-in |
simsys_pool_waiting |
Gauge | service, pool |
Py Node Go | core | opt-in |
simsys_process_cpu_seconds_total |
Counter | service |
Py Node Go | core | baseline |
simsys_process_memory_bytes |
Gauge | service, type |
Py Node Go | core | baseline — type values differ by runtime, see caveat below |
simsys_process_open_fds |
Gauge | service |
Py Node Go | core | baseline |
simsys_progress_estimated_completion_timestamp |
Gauge | service, operation |
Py Node Go | core | opt-in |
simsys_progress_processed_total |
Counter | service, operation |
Py Node Go | core | opt-in |
simsys_progress_rate_per_second |
Gauge | service, operation |
Py Node Go | core | opt-in |
simsys_progress_remaining |
Gauge | service, operation |
Py Node Go | core | opt-in |
simsys_queue_depth |
Gauge | service, queue |
Py Node Go | core | opt-in |
simsys_process_threads |
Gauge | service |
Py Go | ext | baseline |
simsys_process_uptime_seconds |
Gauge | service |
Node | ext | baseline |
simsys_runtime_gc_collections_by_generation_total |
Counter | service, generation |
Py | ext | baseline |
simsys_runtime_gc_collections_total |
Counter | service |
Go | ext | baseline |
simsys_runtime_gc_pause_total_seconds |
Counter | service |
Go | ext | baseline |
simsys_runtime_goroutines |
Gauge | service |
Go | ext | baseline |
simsys_scrape_duration_seconds |
Gauge | service |
Py Go | ext | baseline |
simsys_scrape_errors_total |
Counter | service |
Py Go | ext | baseline |
simsys_collector_errors_total |
Counter | service, collector, name |
Py Node Go | ext | opt-in |
Tier core is guaranteed in every runtime with the declared type and label
names -- a $service-templated dashboard may rely on it unconditionally.
Tier ext is runtime-specific; the Runtimes column is authoritative and a
panel using one must tolerate its absence.
Memory accounting is fundamentally runtime-specific, so the type label
values intentionally differ across the three sibling packages:
| Runtime | type values |
|---|---|
Python (simsys-metrics) |
rss, vms (psutil's resident + virtual sizes) |
Go (simsys-metrics-go) |
rss, vms (procfs status fields) |
Node (@simsys/metrics) |
rss, heapUsed, heapTotal, external (process.memoryUsage()) |
A $service-templated dashboard panel filtering type="vms" will return
empty for Node services, and a panel showing heapUsed will return
empty for Python/Go services. Either:
- Build runtime-aware dashboards (one panel per runtime, gated on
service =~ "node-.*"etc.), or - Filter on
type="rss"only — the one common label value.
@track_job uses a 2-value {success, error} enum intentionally — it's the
minimum useful signal for generic job timing, and the label stays cheap.
When an app needs a richer per-operation outcome taxonomy (cache hits, validation
errors, upstream failures, etc.), hand-roll a counter via make_counter from the
guarded registry:
from simsys_metrics import get_service, make_counter
forecast_requests_total = make_counter(
"simsys_forecast_requests_total",
"Forecast requests by ticker and outcome.",
# Always include `service` in labelnames — it's what the shared
# `$service`-templated Grafana dashboards filter on. Omitting it
# prints a warning at registration time and breaks the dashboard
# contract for this metric.
labelnames=("service", "ticker", "interval", "outcome"),
)
# Outcome enum is app-specific: e.g. {cache_hit, bad_request, upstream_error,
# success, ...} for a forecasting API, or {dead, parked, active, deferred}
# for a domain scanner.
#
# At call sites, pass service= explicitly. `get_service()` returns the
# value `install(..., service=...)` set, so you don't need to thread it
# through every call site.
forecast_requests_total.labels(
service=get_service(),
ticker="AAPL",
interval="1d",
outcome="cache_hit",
).inc()Pair with safe_label() to cap cardinality on any dimension a user controls.
service is process-global. A process emits under exactly one service name,
and install() is the only thing that sets it.
The idempotence guard inside install() is keyed on a sentinel stored on the
app object, so two installs against two different apps never reach it — while
_SERVICE in simsys_metrics._baseline is a module-level global that the
second call overwrites:
install(app_a, service="foo", version="1.0.0")
install(app_b, service="bar", version="1.0.0") # different object, guard not reachedEverything app_a had already started — track_queue, track_pool, job spans,
ProgressTracker — begins emitting under "bar", and app_a's process metrics
disappear from its own series. Dashboards that join simsys_build_info to the
other simsys_* metrics on service stop matching.
Since 1.0.0 this logs at ERROR with the stable marker
simsys-metrics: SERVICE IDENTITY CHANGE, naming both identities. Behaviour is
unchanged in this release; it will raise in the next major (Redmine #50321).
Rolling back an install with set_service(None) is not an identity change and
stays quiet.
If you genuinely need two identities, run two processes. There is no
supported in-process alternative: the collectors bind to
prometheus_client.REGISTRY at import time, so a "fresh registry per service"
escape hatch — which the Go lane does offer — cannot be added here without an
API break. A per-call service= override was considered and rejected: it
threads through five call sites and destroys the "install once, everything is
labelled" premise the library exists for.
routeis the route template (/api/jobs/{id}), never the actual path.statusis bucketed to class strings (2xx,3xx,4xx,5xx,1xx), never the raw numeric code.outcomeon the job metrics is exactly one ofsuccessorerror.- Any user-derived label value should pass through
safe_label(value, allowed_set)before being attached to a metric. - The package refuses at registration time to register any metric whose
name does not start with
simsys_. Attempting so raisesValueError.
simsys_build_info.commit is resolved in this order:
SIMSYS_BUILD_COMMITenvironment variable (if set and non-empty).git rev-parse --short HEADin the process's current working directory.- Literal string
"unknown"if neither is available.
In container images, set SIMSYS_BUILD_COMMIT at build time:
ARG GIT_COMMIT=unknown
ENV SIMSYS_BUILD_COMMIT=${GIT_COMMIT}and build with docker build --build-arg GIT_COMMIT=$(git rev-parse --short HEAD) ..
Scope: the multiproc support described here is currently FastAPI only. The Flask installer does not have multiproc support — when run under gunicorn workers it still registers the per-process
simsys_process_*collector and serves/metricsfromprometheus_client's default registry, so per-worker metrics will not aggregate. If you need multiproc-correct metrics under Flask + gunicorn today, mountprometheus_client.make_wsgi_app(MultiProcessCollector(...))at/metricsyourself and skipinstall()'s/metricsroute. Native Flask multiproc support is tracked for a future minor release.
When running FastAPI under uvicorn-with-workers (or gunicorn + uvicorn
worker class), set PROMETHEUS_MULTIPROC_DIR so the worker processes
write metric samples to a shared directory and /metrics aggregates
them via prometheus_client.MultiProcessCollector. With the env var
set, install() on a FastAPI app automatically:
- Mounts a multiproc-aware
/metricsroute that walks the shared directory on every scrape. - Skips registration of the per-process
simsys_process_*collector (SimsysProcessCollectorreads/proc/selfonly and cannot meaningfully aggregate across workers). - Tags
simsys_queue_depthasmultiprocess_mode="livesum"andsimsys_build_infoasmultiprocess_mode="liveall".
Important — env-var ordering:
PROMETHEUS_MULTIPROC_DIRis read atsimsys_metricsimport time (so the gauges can be constructed with the rightmultiprocess_mode). Set the env var in your Dockerfile / shell / process-manager config before any Python code runs. Setting it after importing — for example inside an app-factory function — leaves the gauges constructed in single-process mode and/metricsaggregation will silently fail.
A typical Dockerfile:
ENV PROMETHEUS_MULTIPROC_DIR=/tmp/prometheus_multiproc
RUN mkdir -p /tmp/prometheus_multiproc- Auto-mounted at
/metricson the same port as the app. No separate metrics port. - Designed for direct scraping on a localhost or VPC interface — the package itself adds no auth on the endpoint; gate it at the reverse proxy / network layer if exposed publicly.
app.state.simsys_exempt_paths(FastAPI) andapp.extensions["simsys_metrics"](Flask) expose the recommended auth-exempt path set so upstream middleware can skip it without hard-coding the list.
git clone https://github.com/Simmons-Systems/simsys-metrics.git
cd simsys-metrics
python3 -m venv .venv && . .venv/bin/activate
pip install -e '.[fastapi,flask,test]'
pytest # 93 unit + integration tests
bin/check-metrics-conformance.sh # end-to-end smoke test against the demo app- Bump
versioninpyproject.tomlandsimsys_metrics/__init__.py. All three lanes carry the same contract version — bumping one means bumping all three. - Add a
CHANGELOG.mdentry. - Run
pytestandbin/check-metrics-conformance.sh— both must be green. git tag python-vX.Y.Z && git push --tags.release.ymlbuilds, attests, releases and publishes — then a post-publish verify job per lane consumes the artifact from its public registry (bin/verify-published.sh). Publication is irreversible, so that job reports rather than prevents: a failure turns the release run red. Check it, and do not assume a green publish step means the artifact is fetchable.
To run the same check by hand — before cutting a release, or to confirm one that already shipped:
bin/verify-published.sh python 2.0.0 # fresh venv, install from PyPI
bin/verify-published.sh node 2.0.0 # empty project, install from npm
bin/verify-published.sh go v2.0.1 # clean module, fetch via proxy.golang.orgIt exists because every other check in this repo resolves the artifact by
directory, which is exactly the property a consumer does not have. That is
how the original 2.0.0 Go tag shipped unfetchable with all 18 PR contexts
green (Redmine #50481). That tag is unusable and must never be pinned — it
was superseded by go/v2.0.1 rather than re-pointed, because
proxy.golang.org caches module tags immutably.
(Deliberately spelled out rather than written as a tag literal:
test_root_readme_go_pin_is_current scans this file for go/v<semver> and
cannot tell a warning about a stale pin from a stale pin, so naming it
here would fail the very guard that keeps this section honest.)
See CONTRIBUTING.md. Issues and PRs welcome — bug reports, metric-catalogue gaps, or new framework install paths (Starlette, Quart, etc.) all fair game. Security issues: see the org-level SECURITY.md.
By contributing you agree to the terms of the Code of Conduct.
MIT. Use it anywhere, no attribution required, no warranty.
- Upstream dependencies: prometheus_client, prometheus-fastapi-instrumentator, psutil.
- Grafana dashboard template: any operator-flavored Grafana dashboard that
templates over the
servicelabel will work; build it from the metric catalogue above.