A bare-metal performance lab for studying container-network overhead, TCP congestion control, reverse-proxy cost, fixed-budget scale-out, and benchmark validity.
The repository is a standardized, evidence-focused reconstruction of the academic project:
Evaluating Network Performance and Virtualization Cost in Container-Native Architectures: TCP Congestion Control, Docker Network Stacks, and Reverse Proxies under Network Degradation and CPU Throttling.
| Experiment | Repeated runs | Main finding |
|---|---|---|
| Container network stack | 3 per mode | Docker bridge added an aggregate mean-latency difference of approximately 0.0113 ms compared with bare metal; P99 remained 1 ms at JMeter resolution. |
| CUBIC versus BBR | 3 per condition | Under 50 ms delay and 2% packet loss, CUBIC averaged 4.99 Mbps, while BBR averaged 439.26 Mbps; BBR also showed substantial run-to-run variability. |
| Fixed-budget proxy scale-out | 5 per architecture/load | HAProxy increased P99 latency for this workload, and splitting the same backend CPU quota across two Nginx instances did not consistently improve performance. |
| Benchmark validity audit | 8 legacy runs | An attempted saturation campaign was invalidated after identifying a JMeter throughput-controller ceiling near 5,000 RPS. |
The project is not presented as a collection of favorable benchmark numbers. It demonstrates a complete performance-engineering workflow:
- define controlled experiment variants;
- separate client-side load generation from the system under test;
- repeat runs before aggregating results;
- compare mean, median, standard deviation, range, and tail latency;
- distinguish configured load, achieved throughput, and concurrency;
- preserve negative results;
- invalidate conclusions when the load generator becomes the active ceiling;
- publish processed evidence without committing multi-gigabyte raw logs.
| Member | Role |
|---|---|
| Hoàng Xuân Đồng — 23520297 | Server setup, system monitoring, data analysis, result reconstruction, and repository standardization |
| Trần Hải Đăng — 23520237 | Client-side load generation and network emulation |
| ThS. Đặng Lê Bảo Chương | Academic supervisor |
The original experiments used separate server and client machines connected through Gigabit Ethernet. Hải Đăng operated the client-side load generator, while Đồng configured and monitored the server-side system under test.
- Runtime and resource control: Docker Engine, cgroups v2
- Web and proxy layer: Nginx, HAProxy
- Load generation: Apache JMeter, iperf3
- Network emulation: Linux Traffic Control and NetEm
- TCP algorithms: CUBIC and BBR
- System inspection: cAdvisor API and native Linux tools
- Analysis: Python standard library and deterministic SVG generation
Prometheus and Grafana were not deployed during the original load campaigns to reduce additional monitoring overhead. This reduces measurement interference but does not eliminate the observer effect.
Three Nginx deployment modes were evaluated at a configured target of
approximately 20,000 RPS:
- bare-metal Nginx;
- Docker host networking;
- Docker bridge networking.
Each mode was executed three times. The first 35 seconds of each JMeter run were excluded.
| Deployment mode | Runs | Mean achieved RPS | Mean latency | Median P99 | HTTP error rate |
|---|---|---|---|---|---|
| Bare-metal Nginx | 3 | 19,999.56 | 0.1507 ms | 1.00 ms | 0% |
| Docker host networking | 3 | 19,999.75 | 0.1608 ms | 1.00 ms | 0% |
| Docker bridge networking | 3 | 19,999.84 | 0.1620 ms | 1.00 ms | 0% |
Docker bridge showed an aggregate mean-latency difference of approximately
0.0113 ms compared with bare metal. Docker host networking showed a
difference of approximately 0.0101 ms.
JMeter stored elapsed time in milliseconds. These differences are therefore derived from aggregate means across millions of observations; they are not direct microsecond-resolution timestamp measurements.
TCP CUBIC and BBR were tested under three network conditions:
- clean Gigabit Ethernet;
1%packet loss;50 msdelay with2%packet loss.
Each scenario contains three iperf3 runs. Receiver throughput is the primary reported metric.
| Algorithm | Network condition | Runs | Mean receiver throughput | Median | Standard deviation | Range |
|---|---|---|---|---|---|---|
| CUBIC | Clean | 3 | 941.41 Mbps | 941.43 Mbps | 0.04 Mbps | 941.37–941.44 Mbps |
| BBR | Clean | 3 | 940.09 Mbps | 940.19 Mbps | 0.16 Mbps | 939.90–940.19 Mbps |
| CUBIC | 1% loss | 3 | 941.09 Mbps | 941.22 Mbps | 0.24 Mbps | 940.81–941.23 Mbps |
| BBR | 1% loss | 3 | 929.60 Mbps | 929.51 Mbps | 0.54 Mbps | 929.10–930.17 Mbps |
| CUBIC | 50 ms delay + 2% loss | 3 | 4.99 Mbps | 5.19 Mbps | 0.41 Mbps | 4.52–5.27 Mbps |
| BBR | 50 ms delay + 2% loss | 3 | 439.26 Mbps | 525.70 Mbps | 208.25 Mbps | 201.71–590.36 Mbps |
BBR maintained substantially higher throughput than CUBIC under the tested
high-delay and packet-loss condition. However, its coefficient of variation
was approximately 47.41%, so the result must be interpreted as a strong
directional difference with significant run-to-run variability rather than a
single deterministic multiplier.
The complete 18 iperf3 JSON files are committed under
data/phase2-iperf/.
Every architecture received a total cgroup CPU quota of 3.0 CPU:
| Case | Architecture | CPU quota |
|---|---|---|
| A | Direct Nginx | Nginx: 3.0 |
| B | HAProxy with one backend | HAProxy: 1.0; Nginx: 2.0 |
| C | HAProxy with two backends | HAProxy: 1.0; Nginx: 1.0 + 1.0 |
The canonical campaign used a 75 KiB static response at configured loads of
300, 480, and 840 RPS. Each architecture/load combination was repeated
five times.
The table reports the median P99 across five runs.
| Configured load | Case A: direct Nginx | Case B: proxy + one backend | Case C: proxy + two backends |
|---|---|---|---|
| 300 RPS | 4.00 ms | 75.00 ms | 81.00 ms |
| 480 RPS | 4.00 ms | 64.06 ms | 78.00 ms |
| 840 RPS | 6.00 ms | 58.00 ms | 65.40 ms |
All architectures sustained their configured request rates without HTTP errors.
For this workload, adding HAProxy increased tail latency. Splitting the same two-CPU backend quota across two one-CPU Nginx containers also failed to produce a consistent throughput or P99 improvement over one two-CPU backend.
This negative result is preserved deliberately:
Horizontal scaling is not inherently a performance optimization. It helps when the backend is the active bottleneck and the benefit of parallelism exceeds proxy, scheduling, and networking overhead.
The Nginx configuration used worker_processes 1. CPU allocations therefore
represent cgroup quota ceilings, not proof that one Nginx process used every
allocated core.
The original project attempted to find a saturation knee by increasing JMeter virtual users. A later forensic audit found that the test plan contained a default throughput limit of:
300,000 requests/minute = 5,000 requests/second
| Virtual users | Legacy case 11 | Legacy case 12 |
|---|---|---|
| 50 | 2,417.24 RPS | 2,358.90 RPS |
| 100 | 4,822.93 RPS | 4,695.18 RPS |
| 200 | 4,998.26 RPS | 4,998.26 RPS |
| 500 | 4,998.40 RPS | 4,998.68 RPS |
Increasing virtual users caused achieved throughput to approach the configured client-side ceiling. The campaign therefore cannot prove backend saturation or a valid server-side hockey-stick knee.
It is retained as a methodology lesson:
Before declaring system saturation, verify that the load generator continues to increase offered throughput and has not become the active ceiling.
.
├── analysis/ # Canonical parsers, aggregation, figures, validation
├── data/
│ ├── phase2-iperf/ # Complete public Phase 2 raw JSON
│ └── manifests/ # Checksums and external raw-data inventory
├── docs/
│ ├── assets/ # Topology image
│ ├── data-management.md
│ ├── limitations.md
│ └── methodology.md
├── experiments/
│ ├── phase1-container-networking/
│ ├── phase2-congestion-control/
│ ├── phase3-proxy-scaleout/
│ └── benchmark-validity/
├── results/
│ ├── figures/ # Deterministic SVG figures
│ └── processed/ # Per-run and aggregate CSV outputs
├── Makefile
├── pyproject.toml
└── README.md
The public repository does not require third-party Python packages.
make public-checkThis command:
- compiles the Python analysis package;
- verifies the committed iperf3 JSON checksums;
- regenerates SVG figures from processed CSV files;
- validates expected result counts and experiment invariants;
- checks the public repository structure.
Validate the portable Phase 3 Compose files when Docker Compose is available:
make compose-configLarge JMeter JTL and execution-log files are intentionally stored outside Git. The complete original dataset is approximately 10 GiB.
To regenerate all processed results, provide an external data root with:
RAW_DATA_ROOT/
└── results/
├── raw/
├── raw_vfinal2/
└── archive_jmeter_limit/
Then run:
make full-check RAW_DATA_ROOT=/path/to/raw-dataSee:
The public repository contains:
- complete Phase 2 iperf3 JSON output;
- per-run and aggregate CSV summaries;
- deterministic figures;
- test configurations;
- portable Compose definitions;
- checksums and size manifests for external raw artifacts;
- result and repository validators.
The public repository does not contain:
- multi-gigabyte JMeter JTL files;
- execution logs;
- superseded exploratory campaigns;
- legacy result scripts with machine-specific paths.
The excluded artifacts remain preserved in a verified external archive.
