Skip to content

Latest commit

 

History

17 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Container Performance and Benchmark Validity Lab

Benchmark Evidence Validation

Python Docker Nginx HAProxy Apache JMeter Linux NetEm

A bare-metal performance lab for studying container-network overhead, TCP congestion control, reverse-proxy cost, fixed-budget scale-out, and benchmark validity.

The repository is a standardized, evidence-focused reconstruction of the academic project:

Evaluating Network Performance and Virtualization Cost in Container-Native Architectures: TCP Congestion Control, Docker Network Stacks, and Reverse Proxies under Network Degradation and CPU Throttling.

Key findings

Experiment Repeated runs Main finding
Container network stack 3 per mode Docker bridge added an aggregate mean-latency difference of approximately 0.0113 ms compared with bare metal; P99 remained 1 ms at JMeter resolution.
CUBIC versus BBR 3 per condition Under 50 ms delay and 2% packet loss, CUBIC averaged 4.99 Mbps, while BBR averaged 439.26 Mbps; BBR also showed substantial run-to-run variability.
Fixed-budget proxy scale-out 5 per architecture/load HAProxy increased P99 latency for this workload, and splitting the same backend CPU quota across two Nginx instances did not consistently improve performance.
Benchmark validity audit 8 legacy runs An attempted saturation campaign was invalidated after identifying a JMeter throughput-controller ceiling near 5,000 RPS.

Why this repository matters

The project is not presented as a collection of favorable benchmark numbers. It demonstrates a complete performance-engineering workflow:

  • define controlled experiment variants;
  • separate client-side load generation from the system under test;
  • repeat runs before aggregating results;
  • compare mean, median, standard deviation, range, and tail latency;
  • distinguish configured load, achieved throughput, and concurrency;
  • preserve negative results;
  • invalidate conclusions when the load generator becomes the active ceiling;
  • publish processed evidence without committing multi-gigabyte raw logs.

Team

Member Role
Hoàng Xuân Đồng — 23520297 Server setup, system monitoring, data analysis, result reconstruction, and repository standardization
Trần Hải Đăng — 23520237 Client-side load generation and network emulation
ThS. Đặng Lê Bảo Chương Academic supervisor

The original experiments used separate server and client machines connected through Gigabit Ethernet. Hải Đăng operated the client-side load generator, while Đồng configured and monitored the server-side system under test.

Lab topology

Bare-metal benchmark lab topology

Technology stack

  • Runtime and resource control: Docker Engine, cgroups v2
  • Web and proxy layer: Nginx, HAProxy
  • Load generation: Apache JMeter, iperf3
  • Network emulation: Linux Traffic Control and NetEm
  • TCP algorithms: CUBIC and BBR
  • System inspection: cAdvisor API and native Linux tools
  • Analysis: Python standard library and deterministic SVG generation

Prometheus and Grafana were not deployed during the original load campaigns to reduce additional monitoring overhead. This reduces measurement interference but does not eliminate the observer effect.


Experiment 1 — Container network-stack overhead

Design

Three Nginx deployment modes were evaluated at a configured target of approximately 20,000 RPS:

  1. bare-metal Nginx;
  2. Docker host networking;
  3. Docker bridge networking.

Each mode was executed three times. The first 35 seconds of each JMeter run were excluded.

Results

Deployment mode Runs Mean achieved RPS Mean latency Median P99 HTTP error rate
Bare-metal Nginx 3 19,999.56 0.1507 ms 1.00 ms 0%
Docker host networking 3 19,999.75 0.1608 ms 1.00 ms 0%
Docker bridge networking 3 19,999.84 0.1620 ms 1.00 ms 0%

Mean latency by network mode

Docker bridge showed an aggregate mean-latency difference of approximately 0.0113 ms compared with bare metal. Docker host networking showed a difference of approximately 0.0101 ms.

JMeter stored elapsed time in milliseconds. These differences are therefore derived from aggregate means across millions of observations; they are not direct microsecond-resolution timestamp measurements.

Experiment details


Experiment 2 — TCP congestion control under impairment

Design

TCP CUBIC and BBR were tested under three network conditions:

  • clean Gigabit Ethernet;
  • 1% packet loss;
  • 50 ms delay with 2% packet loss.

Each scenario contains three iperf3 runs. Receiver throughput is the primary reported metric.

Results

Algorithm Network condition Runs Mean receiver throughput Median Standard deviation Range
CUBIC Clean 3 941.41 Mbps 941.43 Mbps 0.04 Mbps 941.37–941.44 Mbps
BBR Clean 3 940.09 Mbps 940.19 Mbps 0.16 Mbps 939.90–940.19 Mbps
CUBIC 1% loss 3 941.09 Mbps 941.22 Mbps 0.24 Mbps 940.81–941.23 Mbps
BBR 1% loss 3 929.60 Mbps 929.51 Mbps 0.54 Mbps 929.10–930.17 Mbps
CUBIC 50 ms delay + 2% loss 3 4.99 Mbps 5.19 Mbps 0.41 Mbps 4.52–5.27 Mbps
BBR 50 ms delay + 2% loss 3 439.26 Mbps 525.70 Mbps 208.25 Mbps 201.71–590.36 Mbps

CUBIC and BBR throughput under network degradation

BBR maintained substantially higher throughput than CUBIC under the tested high-delay and packet-loss condition. However, its coefficient of variation was approximately 47.41%, so the result must be interpreted as a strong directional difference with significant run-to-run variability rather than a single deterministic multiplier.

The complete 18 iperf3 JSON files are committed under data/phase2-iperf/.

Experiment details


Experiment 3 — Reverse-proxy overhead under an equal CPU quota

Design

Every architecture received a total cgroup CPU quota of 3.0 CPU:

Case Architecture CPU quota
A Direct Nginx Nginx: 3.0
B HAProxy with one backend HAProxy: 1.0; Nginx: 2.0
C HAProxy with two backends HAProxy: 1.0; Nginx: 1.0 + 1.0

The canonical campaign used a 75 KiB static response at configured loads of 300, 480, and 840 RPS. Each architecture/load combination was repeated five times.

Results

The table reports the median P99 across five runs.

Configured load Case A: direct Nginx Case B: proxy + one backend Case C: proxy + two backends
300 RPS 4.00 ms 75.00 ms 81.00 ms
480 RPS 4.00 ms 64.06 ms 78.00 ms
840 RPS 6.00 ms 58.00 ms 65.40 ms

All architectures sustained their configured request rates without HTTP errors.

P99 latency under an equal CPU quota

For this workload, adding HAProxy increased tail latency. Splitting the same two-CPU backend quota across two one-CPU Nginx containers also failed to produce a consistent throughput or P99 improvement over one two-CPU backend.

This negative result is preserved deliberately:

Horizontal scaling is not inherently a performance optimization. It helps when the backend is the active bottleneck and the benefit of parallelism exceeds proxy, scheduling, and networking overhead.

The Nginx configuration used worker_processes 1. CPU allocations therefore represent cgroup quota ceilings, not proof that one Nginx process used every allocated core.

Experiment details


Benchmark-validity case study

The original project attempted to find a saturation knee by increasing JMeter virtual users. A later forensic audit found that the test plan contained a default throughput limit of:

300,000 requests/minute = 5,000 requests/second

Observed throughput

Virtual users Legacy case 11 Legacy case 12
50 2,417.24 RPS 2,358.90 RPS
100 4,822.93 RPS 4,695.18 RPS
200 4,998.26 RPS 4,998.26 RPS
500 4,998.40 RPS 4,998.68 RPS

JMeter throughput-controller ceiling

Increasing virtual users caused achieved throughput to approach the configured client-side ceiling. The campaign therefore cannot prove backend saturation or a valid server-side hockey-stick knee.

It is retained as a methodology lesson:

Before declaring system saturation, verify that the load generator continues to increase offered throughput and has not become the active ceiling.

Validity case study


Repository structure

.
├── analysis/                       # Canonical parsers, aggregation, figures, validation
├── data/
│   ├── phase2-iperf/               # Complete public Phase 2 raw JSON
│   └── manifests/                  # Checksums and external raw-data inventory
├── docs/
│   ├── assets/                     # Topology image
│   ├── data-management.md
│   ├── limitations.md
│   └── methodology.md
├── experiments/
│   ├── phase1-container-networking/
│   ├── phase2-congestion-control/
│   ├── phase3-proxy-scaleout/
│   └── benchmark-validity/
├── results/
│   ├── figures/                    # Deterministic SVG figures
│   └── processed/                  # Per-run and aggregate CSV outputs
├── Makefile
├── pyproject.toml
└── README.md

Validate the public repository

The public repository does not require third-party Python packages.

make public-check

This command:

  1. compiles the Python analysis package;
  2. verifies the committed iperf3 JSON checksums;
  3. regenerates SVG figures from processed CSV files;
  4. validates expected result counts and experiment invariants;
  5. checks the public repository structure.

Validate the portable Phase 3 Compose files when Docker Compose is available:

make compose-config

Full raw-data regeneration

Large JMeter JTL and execution-log files are intentionally stored outside Git. The complete original dataset is approximately 10 GiB.

To regenerate all processed results, provide an external data root with:

RAW_DATA_ROOT/
└── results/
    ├── raw/
    ├── raw_vfinal2/
    └── archive_jmeter_limit/

Then run:

make full-check RAW_DATA_ROOT=/path/to/raw-data

See:

Evidence policy

The public repository contains:

  • complete Phase 2 iperf3 JSON output;
  • per-run and aggregate CSV summaries;
  • deterministic figures;
  • test configurations;
  • portable Compose definitions;
  • checksums and size manifests for external raw artifacts;
  • result and repository validators.

The public repository does not contain:

  • multi-gigabyte JMeter JTL files;
  • execution logs;
  • superseded exploratory campaigns;
  • legacy result scripts with machine-specific paths.

The excluded artifacts remain preserved in a verified external archive.

About

Reproducible bare-metal benchmarking of Docker networking, TCP BBR vs CUBIC, HAProxy overhead, CPU throttling, and benchmark validity.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages