Skip to content

Latest commit

 

History

History
88 lines (56 loc) · 2.61 KB

File metadata and controls

88 lines (56 loc) · 2.61 KB

Benchmark Methodology

Experimental roles

  • Hoang Xuan Dong: server setup, system monitoring, result analysis, and repository standardization.
  • Tran Hai Dang: client-side load generation and network emulation.

The experiments were executed using separate server and client machines connected through Gigabit Ethernet.

Measurement principles

Repeated runs

Canonical results are based on repeated runs rather than one selected output:

  • Phase 1: 3 runs per deployment mode
  • Phase 2: 3 runs per TCP and network condition
  • Phase 3: 5 runs per architecture and configured load

Each run is analyzed independently before cross-run aggregation.

JMeter warm-up

The first 35 seconds of each canonical JMeter run are excluded. This avoids mixing process startup and initial connection behavior into the steady-state summary.

Throughput window

JMeter throughput is calculated from the first analyzed request start to the last analyzed request completion.

This completion-aware window avoids overstating throughput when the final requests have non-zero response time.

Percentiles

P50, P95, P99, and P99.9 are calculated from successful and failed requests that remain after the warm-up boundary. Error rate is reported separately.

JMeter elapsed-time samples are recorded in milliseconds. Sub-millisecond differences in aggregate means must not be interpreted as direct microsecond-resolution timestamp measurements.

iperf3

Receiver throughput is the primary TCP throughput metric. Sender throughput, retransmissions, and mean RTT are retained as supporting indicators.

Cross-run aggregation

Public summaries include:

  • arithmetic mean;
  • median;
  • sample standard deviation;
  • minimum;
  • maximum;
  • coefficient of variation where useful.

A favorable individual run is not selected as the headline result.

Result classification

Canonical

A campaign with known configuration lineage, repeated runs, complete processed summaries, and a defensible interpretation.

Exploratory

A useful troubleshooting or calibration campaign that is excluded from the headline result because its configuration lineage is incomplete or superseded.

Invalidated

A campaign whose original conclusion is not supported after auditing the test design. Invalidated campaigns may still be retained as methodology lessons.

Benchmark validity

Configured target load, achieved offered load, concurrency, client capacity, and server capacity are different measurements.

A saturation conclusion requires evidence that the load generator continued to increase offered throughput and did not become the active ceiling first.