Skip to content

About

Production-grade banking microservices platform on Kubernetes, featuring GitOps with Argo CD & Kargo, progressive canary/blue-green delivery, automated health gates and rollback, DevSecOps, a zero-trust signed software supply chain, policy-based admission control, runtime threat detection, secrets management, and full-stack observability.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

Β 

History

2 Commits

Folders and files

Repository files navigation

πŸ›‘οΈ Production-Grade DevOps & GitOps Microservices Platform

A production-grade 9-service microservices platform combining Kubernetes, GitOps, DevSecOps, progressive delivery, software supply-chain security, runtime protection, and full-stack observability.

Microservices Β· Kubernetes Β· GitOps Β· Kargo Promotion Β· DevSecOps Β· Progressive Delivery Β· Supply-Chain Security Β· Runtime Security Β· Observability

Kubernetes Azure GitHub Actions Argo CD Argo Rollouts Kargo Kyverno Sigstore Falco Observability


πŸ“– Overview

This project is a production-grade DevOps & GitOps microservices platform built around a realistic 9-service digital banking architecture.

It demonstrates both sides of modern platform engineering:

  1. Microservices architecture β€” independently deployable services, service-owned databases, API gateway routing, asynchronous messaging, distributed transactions, fault isolation, and service-level resilience.
  2. Production delivery platform β€” Kubernetes, GitOps, CI security gates, signed artifacts, admission policies, progressive delivery, observability, runtime security, TLS, secrets management, and automated rollback.

The platform is not a single application wrapped in Kubernetes. It is a distributed system composed of multiple independently deployable workloads with different runtimes, data ownership boundaries, dependencies, and release strategies.

The core operating rule is:

Git is the source of truth. Production is changed through reconciliation, not by manually applying manifests.

The platform implements:

  • 9 independently deployable microservices
  • service-owned PostgreSQL databases
  • YARP API gateway routing
  • RabbitMQ-based asynchronous messaging
  • Redis-backed distributed platform components
  • cross-service transaction saga with compensation
  • service-level fault isolation and resilience
  • automated container build and promotion
  • immutable commit-SHA releases
  • declarative GitOps deployment
  • Argo CD app-of-apps
  • Kargo continuous-promotion control plane
  • Warehouse β†’ Freight β†’ Stage promotion model
  • ordered sync waves
  • canary and blue-green delivery
  • automated health and metrics gates
  • container vulnerability scanning
  • SPDX SBOM generation
  • keyless image signing -Rekor transparency logging
  • Kyverno admission enforcement
  • Falco runtime threat detection
  • Prometheus metrics and alerting
  • Loki log aggregation with Grafana Alloy
  • Grafana dashboards
  • HashiCorp Vault secret management
  • automated TLS with cert-manager and Let's Encrypt
  • operational rollback and recovery runbooks

🧩 Microservices Architecture

The application layer consists of 9 independently deployable services behind a gateway.

                         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                         β”‚        Frontend          β”‚
                         β”‚      Angular + nginx     β”‚
                         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                      β”‚
                                      β–Ό
                         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                         β”‚       API Gateway        β”‚
                         β”‚          YARP            β”‚
                         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                      β”‚
          β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
          β”‚               β”‚           β”‚           β”‚               β”‚
          β–Ό               β–Ό           β–Ό           β–Ό               β–Ό
     Identity         Account    Transaction   Payment         Lending
          β”‚               β”‚           β”‚           β”‚               β”‚
          β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                 β”‚
                         β”Œβ”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”
                         β–Ό                β–Ό
                     Platform           Fraud
                                      FastAPI

                         + Gateway + Frontend
                         = 9 deployable services

Service boundaries

Service Responsibility Runtime
Identity authentication, sessions, devices, recovery ASP.NET Core 8
Account account state and balances ASP.NET Core 8
Transaction ledger, transfers, reconciliation, distributed transaction workflow ASP.NET Core 8
Payment payment workflows and merchant-facing operations ASP.NET Core 8
Lending lending workflows and repayment schedules ASP.NET Core 8
Platform shared operational capabilities, audit and platform state ASP.NET Core 8
Fraud risk scoring and fraud analysis Python Β· FastAPI
Gateway API routing and edge service composition YARP Β· ASP.NET Core
Frontend browser application and /api proxy Angular Β· nginx

Data ownership

Each backend service owns its own data boundary.

Identity      ──► identitydb
Account       ──► accountdb
Transaction   ──► transactiondb
Payment       ──► paymentdb
Lending       ──► lendingdb
Platform      ──► platformdb

The services do not read each other's tables directly. Cross-service interaction happens through APIs and messaging.

This matters operationally because each service can be:

  • deployed independently
  • rolled back independently
  • monitored independently
  • scaled independently
  • isolated when unhealthy
  • released with a strategy appropriate to its risk

Distributed transaction handling

Money movement crosses service boundaries, so transfers use a saga with compensation rather than pretending a single database transaction can cover the whole system.

Debit Account
     β”‚
     β”œβ”€β”€ FAIL ──► transaction fails, nothing moved
     β”‚
     β–Ό
Credit Account
     β”‚
     β”œβ”€β”€ FAIL ──► compensate by refunding debit
     β”‚
     β–Ό
Complete Transaction

A background reconciler handles operations left mid-flight and helps detect stranded transaction state.

Idempotency

Critical operations use idempotency protection at the database level so duplicate requests cannot race and both succeed.

Messaging and infrastructure

The platform also includes:

  • RabbitMQ for asynchronous service communication
  • Redis for distributed platform state and supporting workloads
  • HashiCorp Vault for sensitive key material and application secrets
  • Azure Database for PostgreSQL as the external database layer

Why microservices here?

The architecture deliberately accepts the operational complexity of microservices because the platform is designed around fault isolation and independent recovery.

A failure in one service should not require redeploying the whole banking platform.

That same separation carries into delivery:

One service changes
      β”‚
      β–Ό
Build only the required artifact
      β”‚
      β–Ό
Scan + sign
      β”‚
      β–Ό
Update GitOps state
      β”‚
      β–Ό
Argo CD detects change
      β”‚
      β–Ό
Only that workload rolls out

This is where the microservices architecture and the GitOps platform reinforce each other: small deployment units, isolated blast radius, independent rollback, and service-level observability.


πŸ” End-to-End Delivery Flow

Developer Push
      β”‚
      β–Ό
GitHub Actions
      β”‚
      β”œβ”€β”€ Test / validate
      β”œβ”€β”€ Build container images
      β”œβ”€β”€ Trivy vulnerability gate
      β”œβ”€β”€ Syft SPDX SBOM
      β”œβ”€β”€ Push immutable SHA-tagged images
      └── Cosign keyless signing + Rekor
      β”‚
      β–Ό
Kargo Promotion Layer
      β”‚
      β”œβ”€β”€ Warehouse discovers deployable artifacts
      β”œβ”€β”€ Freight represents an immutable release candidate
      β”œβ”€β”€ Production Stage records promotion state/history
      └── Promotion updates desired image tag in Git
      β”‚
      β–Ό
Argo CD
      β”‚
      β”œβ”€β”€ Detect desired-state change
      β”œβ”€β”€ Reconcile cluster
      └── Restore drift automatically
      β”‚
      β–Ό
Kyverno Admission
      β”‚
      └── Verify image identity + signature + digest
      β”‚
      β–Ό
Argo Rollouts
      β”‚
      β”œβ”€β”€ Canary or Blue-Green
      β”œβ”€β”€ Health analysis
      β”œβ”€β”€ Prometheus metric analysis
      └── Promote / Abort / Rollback
      β”‚
      β–Ό
Production Workloads
      β”‚
      β”œβ”€β”€ Prometheus / Alertmanager
      β”œβ”€β”€ Loki / Grafana Alloy
      β”œβ”€β”€ Grafana
      └── Falco / Falcosidekick

πŸ—οΈ DevOps & Microservices Platform Architecture

flowchart TB
    DEV["Developer<br/>git push"]

    subgraph CI["CI / Software Supply Chain"]
        TEST["Tests & validation"]
        BUILD["Build images<br/>commit-SHA tags"]
        TRIVY["Trivy<br/>HIGH / CRITICAL gate"]
        SBOM["Syft<br/>SPDX SBOM"]
        REG["Docker Hub"]
        SIGN["Cosign keyless signing<br/>GitHub OIDC + Rekor"]

        TEST --> BUILD --> TRIVY --> SBOM --> REG --> SIGN
    end

    subgraph GITOPS["GitOps & Promotion"]
        WAREHOUSE["Kargo Warehouse<br/>artifact discovery"]
        FREIGHT["Kargo Freight<br/>immutable release candidate"]
        STAGE["Kargo Stage<br/>production promotion"]
        VALUES["Production image tags"]
        ROOT["Root Application"]
        APPS["Argo CD Applications"]

        WAREHOUSE --> FREIGHT --> STAGE --> VALUES
        ROOT --> APPS
    end

    subgraph CLOUD["Microsoft Azure"]
        subgraph K3S["k3s Cluster"]
            ARGO["Argo CD"]
            KYVERNO["Kyverno"]
            ROLLOUTS["Argo Rollouts"]

            subgraph SERVICES["AEGIS Workloads"]
                APPSVC["9 application services"]
                VAULT["Vault"]
                RMQ["RabbitMQ"]
                REDIS["Redis"]
            end

            subgraph OBS["Observability"]
                PROM["Prometheus"]
                ALERT["Alertmanager"]
                GRAF["Grafana"]
                LOKI["Loki"]
                ALLOY["Grafana Alloy"]
            end

            subgraph RUNTIME["Runtime Security"]
                FALCO["Falco<br/>modern eBPF"]
                SIDEKICK["Falcosidekick + UI"]
            end

            TRAEFIK["Traefik"]
            CERT["cert-manager<br/>Let's Encrypt"]
        end

        PG["Azure Database<br/>for PostgreSQL"]
    end

    DEV --> TEST
    SIGN --> WAREHOUSE
    STAGE --> VALUES
    VALUES --> ARGO
    ROOT --> ARGO
    ARGO --> KYVERNO
    ARGO --> ROLLOUTS
    ROLLOUTS --> APPSVC
    KYVERNO --> APPSVC
    APPSVC --> PG
    APPSVC --> VAULT
    APPSVC --> RMQ
    APPSVC --> REDIS
    PROM --> GRAF
    LOKI --> GRAF
    ALLOY --> LOKI
    PROM --> ALERT
    FALCO --> SIDEKICK
    TRAEFIK --> APPSVC
    CERT --> TRAEFIK
Loading

🧰 Technology Stack

Layer Technology
Application Architecture 9-service microservices platform
Backend Runtime ASP.NET Core 8
Fraud Runtime Python 3.11 Β· FastAPI
Frontend Angular Β· nginx
API Gateway YARP
Messaging RabbitMQ Β· MassTransit
Caching / Distributed State Redis
Cloud Microsoft Azure
Kubernetes k3s
GitOps Argo CD
Continuous Promotion Kargo 1.11 β€” Project Β· Warehouse Β· Freight Β· Stage
Progressive Delivery Argo Rollouts
CI GitHub Actions
Container Registry Docker Hub
Image Scanning Trivy
SBOM Syft β€” SPDX JSON
Image Signing Sigstore Cosign β€” keyless OIDC
Transparency Log Rekor
Admission Policy Kyverno
Runtime Security Falco β€” modern eBPF
Security Events Falcosidekick + UI
Metrics Prometheus
Alerting Alertmanager
Cluster Metrics kube-state-metrics Β· node-exporter
Logs Loki
Log Collection Grafana Alloy
Dashboards Grafana
Secrets HashiCorp Vault
Ingress Traefik
TLS cert-manager + Let's Encrypt
Packaging Helm
Database Azure Database for PostgreSQL

πŸ”€ Repository Separation

The delivery design separates application source and CI from cluster desired state.

Application / CI Repository
        β”‚
        β”‚ successful build only
        β–Ό
GitOps Repository
        β”‚
        β”‚ read-only reconciliation
        β–Ό
Argo CD
        β”‚
        β–Ό
Kubernetes

This separation is deliberate.

CI is allowed to update only the deployment state required for promotion. The cluster only needs read access to the repository that describes it.

Two narrowly scoped deploy keys are used:

Key Held by Access
argocd-cluster-readonly Argo CD / cluster Read-only
aegis-ci-promote CI Read/write for promotion

A broad personal access token is intentionally avoided. The principle is:

One identity, one direction, one purpose.


πŸ—ΊοΈ GitOps Repository Layout

.
β”œβ”€β”€ bootstrap/
β”‚   └── root-app.yaml
β”‚
β”œβ”€β”€ apps/
β”‚   β”œβ”€β”€ cert-manager.yaml
β”‚   β”œβ”€β”€ admission.yaml
β”‚   β”œβ”€β”€ base.yaml
β”‚   β”œβ”€β”€ infrastructure.yaml
β”‚   β”œβ”€β”€ runtime-security.yaml
β”‚   β”œβ”€β”€ observability.yaml
β”‚   β”œβ”€β”€ logs.yaml
β”‚   β”œβ”€β”€ promotion.yaml
β”‚   β”œβ”€β”€ identity.yaml
β”‚   β”œβ”€β”€ account.yaml
β”‚   β”œβ”€β”€ transaction.yaml
β”‚   β”œβ”€β”€ payment.yaml
β”‚   β”œβ”€β”€ lending.yaml
β”‚   β”œβ”€β”€ platform.yaml
β”‚   β”œβ”€β”€ fraud.yaml
β”‚   β”œβ”€β”€ gateway.yaml
β”‚   └── frontend.yaml
β”‚
β”œβ”€β”€ manifests/
β”‚   β”œβ”€β”€ namespace.yaml
β”‚   β”œβ”€β”€ ingress.yaml
β”‚   β”œβ”€β”€ redirect-middleware.yaml
β”‚   β”œβ”€β”€ analysis-template.yaml
β”‚   β”œβ”€β”€ promotion/
β”‚   β”‚   β”œβ”€β”€ project.yaml
β”‚   β”‚   β”œβ”€β”€ warehouse.yaml
β”‚   β”‚   └── stage-production.yaml
β”‚   β”œβ”€β”€ admission/
β”‚   β”œβ”€β”€ certs/
β”‚   β”œβ”€β”€ infrastructure/
β”‚   β”œβ”€β”€ observability/
β”‚   β”œβ”€β”€ falco/
β”‚   β”œβ”€β”€ argocd/
β”‚   └── services/
β”‚
β”œβ”€β”€ charts/
β”‚   └── aegis-service/
β”‚
β”œβ”€β”€ environments/
β”‚   β”œβ”€β”€ base/
β”‚   └── production/
β”‚       └── values.yaml
β”‚
└── scripts/

The service chart replaces repeated deployment manifests with a reusable deployment pattern, while environment-specific values keep image tags and configuration declarative.


🏷️ Immutable Releases

AEGIS deploys images using commit-derived SHA tags.

Example:

aegis-identity:sha-<commit>
aegis-account:sha-<commit>
aegis-transaction:sha-<commit>
aegis-payment:sha-<commit>
aegis-lending:sha-<commit>
aegis-platform:sha-<commit>
aegis-fraud:sha-<commit>
aegis-gateway:sha-<commit>
aegis-frontend:sha-<commit>

latest is not used as a production deployment reference.

Why:

  • a SHA tag identifies exactly which source revision produced the image
  • Git detects an actual desired-state change
  • rollbacks target a known artifact
  • a running Pod can be traced back to a commit
  • mutable tags cannot silently change the software behind an unchanged manifest

Before promotion, CI verifies that all expected service images actually exist. This prevents a GitOps update from referencing a commit for which no deployable container image was produced.


βš™οΈ CI & Software Supply Chain

The supply chain follows a fail-fast model: artifacts must pass required checks before becoming deployable.

πŸ” Vulnerability scanning

Trivy scans built images for:

  • HIGH
  • CRITICAL

The pipeline uses a blocking exit code so failed security gates stop the release.

πŸ“‹ SBOM

Syft generates an SPDX JSON Software Bill of Materials for container images.

The SBOM is associated with the image so software-component evidence is not limited to temporary CI logs.

✍️ Keyless signing

Images are signed using Sigstore Cosign keyless signing.

GitHub Actions
      β”‚
      β”œβ”€β”€ GitHub OIDC identity
      β–Ό
Sigstore Fulcio
      β”‚
      β”œβ”€β”€ short-lived certificate
      β–Ό
Cosign signature
      β”‚
      β”œβ”€β”€ stored with the image
      β–Ό
Rekor transparency log

There is no long-lived private signing key stored in CI.

This reduces:

  • signing-key theft risk
  • manual signing-key rotation
  • hidden signing events
  • CI secret sprawl

The security question is not simply:

β€œIs this image signed?”

It is:

β€œWas this image signed by the expected CI identity for the expected repository and workflow?”


πŸ›‘οΈ Admission Control with Kyverno

Argo CD answers:

Does the cluster match Git?

Kyverno adds another question:

Can the artifact named by Git be trusted?

The verify-aegis-images policy runs in Enforce mode for AEGIS application images.

It verifies that:

  • the image has a valid Cosign signature
  • the signature comes from the expected OIDC identity
  • the image belongs to the expected workload scope
  • the image digest is verified
  • the artifact that was verified is the artifact that actually runs

Digest verification and mutation protect the deployment from a tag being repointed after verification.

A separate policy audits moving latest tags on workloads outside the enforced AEGIS image rule.


πŸ“¦ Kargo β€” Continuous Promotion

Argo CD deploys desired state. Kargo models how a release becomes the desired state.

That distinction is important:

Container Registry
      β”‚
      β–Ό
Kargo Warehouse
      β”‚
      β–Ό
Freight
      β”‚
      β–Ό
Stage
      β”‚
      β–Ό
Git commit
      β”‚
      β–Ό
Argo CD
      β”‚
      β–Ό
Kubernetes

Kargo is installed and managed by Argo CD as part of the platform rather than being configured manually outside GitOps.

Why Kargo exists here

A shell command or CI job can change an image tag, but after the job ends it has very little durable understanding of:

  • which release candidates are available
  • which candidate is deployed
  • which candidate is waiting
  • what artifact set belongs together
  • what was promoted previously
  • what should move from one environment to another later

Kargo turns promotion itself into Kubernetes-native state with history, health and explicit release objects.

Project β€” aegis-delivery

The promotion model is scoped under a dedicated Kargo project:

aegis-delivery

Its configuration defines how Freight is allowed to move through Stages.

Warehouse β€” aegis-images

The Warehouse watches the AEGIS container stream and discovers release candidates.

kind: Warehouse
metadata:
  name: aegis-images
  namespace: aegis-delivery

A key design decision is to treat the banking platform as a coherent release unit.

All nine application images are produced from the same source commit and share the same sha-* version. The release model therefore represents β€œAEGIS at commit X” rather than allowing arbitrary combinations of unrelated service revisions to be promoted accidentally.

Freight β€” immutable release candidates

When Kargo discovers a valid artifact version, it represents that candidate as Freight.

Conceptually:

Freight
  └── AEGIS release @ sha-<commit>

The important property is that a release candidate becomes a named, traceable object instead of only a tag observed inside a transient CI log.

Stage β€” production

The current promotion pipeline contains a production Stage.

kind: Stage
metadata:
  name: production
  namespace: aegis-delivery

Promotion follows the GitOps rule:

Kargo does not deploy directly to the cluster.

Instead, Kargo:

  1. clones the GitOps repository,
  2. updates the production image version,
  3. commits the desired-state change,
  4. pushes that commit,
  5. lets Argo CD detect and reconcile it.

This keeps Argo CD as the only deployment controller and prevents a second system from bypassing Git.

Promotion architecture

flowchart LR
    REG["Docker Hub<br/>signed SHA images"]
    WH["Warehouse<br/>aegis-images"]
    FR["Freight<br/>release candidate"]
    ST["Stage<br/>production"]
    GIT["GitOps repo<br/>production values"]
    ARGO["Argo CD"]
    ROLLOUT["Argo Rollouts"]
    K8S["AEGIS workloads"]

    REG --> WH --> FR --> ST --> GIT --> ARGO --> ROLLOUT --> K8S
Loading

Separate credentials for separate responsibilities

Kargo uses distinct credentials rather than reusing the CI or Argo CD identity:

Credential Purpose
kargo-api Kargo API signing key + admin authentication
docker-hub read-only registry discovery
gitops-repo write access for promotion commits

The Kargo Git writer has its own deploy key. Sharing the CI deploy key would make the two writers indistinguishable in an audit and would couple their revocation.

GitOps-managed Kargo installation

The platform uses two Argo CD Applications:

promotion
    └── installs Kargo

promotion-config
    └── manages Project / Warehouse / Stage

The engine and promotion model are separated so Kargo can be upgraded without treating its controller-generated promotion history as ordinary Git drift.

Argo CD ignores Kargo-owned status fields for resources such as:

  • Warehouse
  • Stage
  • Project

Otherwise Argo CD would continually attempt to revert state that Kargo legitimately owns.

Current promotion posture

The current production policy is deliberately configured with:

autoPromotionEnabled: false

This is intentional while the existing CI promotion writer and Kargo coexist. Allowing two systems to automatically update the same production version would create a race between writers.

The migration path is:

Current
CI promotion ───────────────► Git ─► Argo CD
Kargo        ── observes / models

Target
Kargo Warehouse ─► Freight ─► Stage ─► Git ─► Argo CD

Registry discovery lesson

The Kargo Warehouse is also configured around a real operational constraint: Docker Hub registry discovery can consume significant request quota when NewestBuild must inspect many historical tags.

The repository documents the mitigation explicitly instead of hiding it:

  • narrow candidate discovery
  • reduced polling frequency
  • promotion-time verification of all expected service images
  • a parked Warehouse state when quota recovery is required

A future production evolution would use a registry/tagging strategy that supports efficient sortable release discovery without per-tag metadata scans.

Why Kargo + Argo CD + Argo Rollouts?

Each tool owns a different concern:

Tool Responsibility
Kargo Which release candidate should move forward?
Argo CD Does the cluster match the approved Git state?
Argo Rollouts Is the new revision healthy enough to receive traffic?

Together:

Kargo
  β”‚ chooses/promotes release
  β–Ό
Git
  β”‚ desired state
  β–Ό
Argo CD
  β”‚ reconciles
  β–Ό
Argo Rollouts
  β”‚ validates rollout
  β–Ό
Production

This separation keeps artifact promotion, desired-state reconciliation, and traffic rollout independently auditable.


πŸ”„ Argo CD β€” App-of-Apps

A single root Application bootstraps the platform.

root-app
   β”‚
   └── apps/
       β”œβ”€β”€ cert-manager
       β”œβ”€β”€ admission
       β”œβ”€β”€ admission-policies
       β”œβ”€β”€ base
       β”œβ”€β”€ infrastructure
       β”œβ”€β”€ runtime-security
       β”œβ”€β”€ observability
       β”œβ”€β”€ logs
       β”œβ”€β”€ routes
       β”œβ”€β”€ identity
       β”œβ”€β”€ account
       β”œβ”€β”€ transaction
       β”œβ”€β”€ payment
       β”œβ”€β”€ lending
       β”œβ”€β”€ platform
       β”œβ”€β”€ fraud
       β”œβ”€β”€ gateway
       └── frontend

The complete platform is managed through 23 Argo CD Applications, including the root Application and supporting platform components.

Core reconciliation behaviour

  • automated synchronization
  • pruning of removed resources
  • self-healing drift correction
  • declarative Helm values
  • ordered dependency deployment
  • no routine manual kubectl apply

🌊 Sync Waves

Infrastructure dependencies are ordered declaratively so application services do not race components that are not ready yet.

Example:

Wave -1
  β”œβ”€β”€ cert-manager
  └── Kyverno engine

Wave 0
  β”œβ”€β”€ admission policies
  β”œβ”€β”€ namespace
  β”œβ”€β”€ ingress
  └── analysis templates

Wave 1
  β”œβ”€β”€ Vault
  β”œβ”€β”€ RabbitMQ
  β”œβ”€β”€ Redis
  β”œβ”€β”€ observability
  └── runtime security

Later
  └── AEGIS application services

The startup sequence is part of desired state rather than tribal knowledge or timing assumptions.


🀝 Controller-Aware Drift Management

Not every difference from Git is accidental drift.

Some fields are legitimately changed by Kubernetes controllers. Argo CD is configured to ignore selected controller-owned fields where reconciliation would otherwise create a controller conflict.

Argo Rollouts service selectors

Argo Rollouts updates Service selectors with ReplicaSet hashes while steering traffic.

If Argo CD continuously reverted those selectors, GitOps reconciliation and rollout reconciliation would fight each other and could direct traffic to the wrong revision.

Kyverno-generated state

Kyverno may mutate or expand policy state after admission. Selected controller-managed fields are excluded from drift enforcement.

Generated chart secrets

Where charts generate internal values during render or installation, selected generated fields can be ignored to avoid meaningless perpetual drift.

GitOps should correct unwanted drift, not fight legitimate controller ownership.


🚦 Progressive Delivery

AEGIS uses two release strategies based on the failure mode of each component.

🐀 Canary β€” backend services

Backend services use canary delivery.

A new revision receives only part of traffic while the stable revision remains available.

Stable revision
      β”‚
      β–Ό
Create canary revision
      β”‚
      β–Ό
Partial live traffic
      β”‚
      β”œβ”€β”€ health analysis
      β”œβ”€β”€ error-rate analysis
      └── latency analysis
      β”‚
      β”œβ”€β”€ PASS ──► continue / promote
      β”‚
      └── FAIL ──► abort / remain stable

In this k3s implementation, traffic weighting is approximated through replica count because no service mesh is installed. With two replicas, the meaningful stages are approximately half traffic and full traffic.

πŸ”΅πŸŸ’ Blue-Green β€” gateway and frontend

Customer entry points use blue-green delivery.

Current version (Blue)
        β”‚
        β”‚ continues serving production
        β–Ό
New version (Green)
        β”‚
        β–Ό
Pre-promotion analysis
        β”‚
        β”œβ”€β”€ FAIL ──► reject Green
        β”‚
        └── PASS
              β”‚
              β–Ό
         traffic cutover
              β”‚
              β–Ό
         Green becomes active

The new version is validated beside the active version before production traffic is switched.


πŸ§ͺ Automated Release Gates

Two analysis patterns protect releases.

service-health

Asks:

Is the new revision actually responding?

Configuration:

  • 5 checks
  • 10 seconds apart
  • zero tolerated failures
  • expected HTTP response: 200
  • parameterized service port

This catches workloads that start successfully but immediately become unhealthy.

canary-metrics

Asks:

Is the canary hurting customers?

Key signals:

  • error rate < 5%
  • p95 latency < 2s
  • measured against the canary revision

A health endpoint alone is not enough. A service can return 200 on /health while real requests fail.

Useful rollout commands

# Follow a rollout
kubectl argo rollouts get rollout identity -n aegis --watch

# Promote
kubectl argo rollouts promote identity -n aegis

# Abort and hold stable
kubectl argo rollouts abort identity -n aegis

# Roll back
kubectl argo rollouts undo identity -n aegis

πŸ“Š Observability

The observability stack combines metrics, logs, alerts, release visibility, and cluster telemetry.

Metrics

kube-prometheus-stack provides:

  • Prometheus
  • Alertmanager
  • Grafana
  • kube-state-metrics
  • node-exporter

ServiceMonitors scrape the AEGIS workloads.

Logs

Grafana Alloy collects workload logs and sends them to Loki.

Grafana provides one interface for:

  • service health
  • Kubernetes health
  • request rates
  • error rates
  • latency
  • rollout behaviour
  • application logs
  • alerts

AEGIS dashboards

Three purpose-built Grafana dashboards provide 34 panels:

Dashboard Panels Focus
Platform Health 11 service health, request rates, error rates, latency, resources
Banking Operations 15 operational workload and transaction-platform signals
Releases 8 rollout revisions, analysis state, release progress

The release dashboard is observational. Argo Rollouts remains responsible for promotion decisions.


🚨 Alerting

Prometheus alert rules cover platform availability and workload health.

Alertmanager receives and groups firing alerts.

Useful diagnostic:

kubectl exec -n observability \
  prometheus-observability-kube-prometh-prometheus-0 \
  -c prometheus -- \
  promtool query instant http://localhost:9090 \
  'ALERTS{alertstate="firing"}'

Watchdog intentionally remains firing as a dead-man's-switch signal proving the alerting pipeline itself is alive.


🚨 Runtime Security with Falco

CI, scanning, signing, and admission policies judge software before or during deployment.

Falco observes actual workload behaviour after deployment.

The platform uses:

  • Falco
  • modern eBPF driver
  • JSON output
  • syscall-drop monitoring
  • Falcosidekick
  • Falcosidekick UI
  • Prometheus ServiceMonitor

Custom AEGIS detections

Detection Severity Purpose
Shell opened in a banking container WARNING Interactive shells are not expected in application containers
Reconnaissance tooling WARNING Detect unexpected use of tools such as nc, nmap, curl, wget, or ssh
Credential material read CRITICAL Detect suspicious access to service-account or credential material
Vault data touched by non-Vault process CRITICAL Protect sensitive Vault storage paths

The goal is to detect suspicious behaviour, not only known malicious files.


πŸ” Secrets & Vault

No application secret is committed to the GitOps repository.

Sensitive material is created or injected separately from Git.

HashiCorp Vault provides the application secret-management layer.

Control Implementation
Seal mechanism Shamir secret sharing
Threshold 3-of-5 shares
Service authentication AppRole
Field encryption AES-256-GCM keys supplied through Vault
Bootstrap Scripted initialization / secret creation outside Git

The design avoids embedding database credentials, Vault credentials, or cryptographic keys in deployment manifests.


🌐 Ingress & TLS

All externally exposed platform interfaces are routed through Traefik.

TLS is managed by:

  • cert-manager
  • Let's Encrypt
  • Kubernetes Certificate resources
  • HTTP β†’ HTTPS redirection
Internet
   β”‚
   β–Ό
Traefik
   β”‚
   β”œβ”€β”€ Application
   β”œβ”€β”€ Argo CD
   β”œβ”€β”€ Grafana
   β”œβ”€β”€ Prometheus
   β”œβ”€β”€ Alertmanager
   └── Falcosidekick UI

Certificates are automatically issued and renewed through cert-manager.


πŸ”’ Kubernetes Security Posture

The aegis namespace is configured around a restricted workload posture.

Controls include:

  • non-root workloads
  • dropped Linux capabilities
  • seccomp RuntimeDefault
  • restricted Pod Security posture
  • signed-image admission for application workloads
  • no latest deployment strategy
  • secrets excluded from Git
  • HTTPS ingress
  • runtime detection with Falco

Security is implemented in layers:

Source
  β”‚
  β–Ό
CI validation
  β”‚
  β–Ό
Container scanning
  β”‚
  β–Ό
SBOM
  β”‚
  β–Ό
Image signing
  β”‚
  β–Ό
GitOps desired state
  β”‚
  β–Ό
Kyverno admission
  β”‚
  β–Ό
Kubernetes hardening
  β”‚
  β–Ό
Falco runtime detection

πŸš€ Bootstrap

The platform follows an app-of-apps bootstrap model.

# Apply the single root Application
kubectl apply -f bootstrap/root-app.yaml

After that, Argo CD discovers and reconciles the child Applications.

Secrets and Vault initialization are intentionally separate from Git:

# Example secret bootstrap
./scripts/create-secrets.sh /path/to/application/.env

# Initialize / configure Vault after it is running
./scripts/bootstrap-vault.sh

The exact secret material is never committed.


πŸ”§ Operations

Platform status

kubectl get applications -n argocd
kubectl get rollouts -n aegis
kubectl get pods -A

Investigate an Argo CD Application

kubectl get app <name> -n argocd \
  -o jsonpath='{.status.operationState.phase} {.status.operationState.message}'

Check Vault

kubectl exec -n aegis deploy/vault -c vault -- \
  sh -c 'VAULT_ADDR=http://127.0.0.1:8200 vault status'

Watch a deployment

kubectl argo rollouts get rollout identity -n aegis --watch

βœ… What This Platform Demonstrates

  • 9-service microservices architecture
  • independently deployable services
  • service-owned databases
  • API gateway routing with YARP
  • distributed transaction saga with compensation
  • RabbitMQ asynchronous messaging
  • Redis platform integration
  • heterogeneous runtimes β€” ASP.NET Core + FastAPI + Angular
  • Kubernetes deployment on Azure with k3s
  • GitOps-based cluster management
  • Argo CD app-of-apps
  • Kargo continuous promotion control plane
  • Kargo Project / Warehouse / Freight / Stage model
  • Git-based Kargo promotion workflow
  • dedicated Kargo registry and Git credentials
  • sync-wave dependency ordering
  • automated drift correction
  • immutable SHA-tag releases
  • cross-repository promotion
  • canary delivery
  • blue-green delivery
  • automated release analysis
  • Prometheus-based rollout metrics
  • Trivy image vulnerability scanning
  • Syft SPDX SBOM generation
  • Cosign keyless image signing
  • Rekor transparency logging
  • Kyverno signature verification
  • digest verification at admission
  • Pod Security hardening
  • HashiCorp Vault integration
  • Prometheus metrics
  • Alertmanager alerting
  • Loki log aggregation
  • Grafana Alloy collection
  • Grafana dashboards
  • Falco runtime detection
  • Falcosidekick security event UI
  • Traefik ingress
  • cert-manager
  • Let's Encrypt TLS
  • operational runbooks and rollback procedures

πŸ›‘οΈ Production-Grade DevOps & GitOps Microservices Platform

Microservices Β· GitOps Β· Continuous Promotion Β· Signed Β· Policy-Gated Β· Observable Β· Reversible

About

Production-grade banking microservices platform on Kubernetes, featuring GitOps with Argo CD & Kargo, progressive canary/blue-green delivery, automated health gates and rollback, DevSecOps, a zero-trust signed software supply chain, policy-based admission control, runtime threat detection, secrets management, and full-stack observability.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages