Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
33 changes: 18 additions & 15 deletions .github/workflows/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,20 +17,22 @@ This directory contains GitHub Actions workflows for automated testing.
- Column alias tests (8)
- Run-time mode tests (7)
- Unit tests (7)
- Feature tests (6)
- Feature tests (8)
- Conservative backfilling tests (2)
- Replay tests (5: one reclamation-boundary binary and four CLI comparisons)
- Resource-history tests (5)
- Job-store tests (6)
- Config tests, including power-usage `trace_type` coverage
- Native CTest suite (12 tests, plus MPI streaming when MPI is available),
including the custom FCFS scheduler's focused and 2,000-job comparisons
- Config tests (12), including power-usage `trace_type`, capacity-schedule,
and warm-start coverage
- Native CTest suite (15 native registrations plus the Python trace-tools
registration, and MPI streaming when MPI is available), including the
custom FCFS scheduler's focused and 2,000-job comparisons
- Ser20-disabled native serialization build and tests
- Sphinx and Doxygen documentation build with warnings treated as errors
- Python API tests (17)
- Python API tests (18)
- gRPC client/server tests (2)
- Append-job tests (19 C++ + 3 gRPC)
- FCFS/EASY backfill-window focused rerun of the three-case gRPC binary
- Append-job tests (20 C++ + 5 optional gRPC checks)
- FCFS/EASY backfill-window focused rerun of the five-check gRPC binary
- Synchronized single-coordinator gRPC test
- Progressive-loading tests (C++ + CLI)
- Queue-input schema test
Expand Down Expand Up @@ -69,25 +71,26 @@ Total tests referenced by the full suite:
| Category | Count | Verified in this doc pass? |
|----------|-------|------------------------------|
| Scheduler correctness | 34 | CI runner |
| Custom FCFS | 7 | CTest; five focused checks and two golden schedules, including 2,000 jobs |
| Custom FCFS | 8 | CTest; six focused checks and two golden schedules, including warm-start accounting and 2,000 jobs |
| Queue implementation differential | 34 fixtures × 4 implementations | CI runner |
| Column aliases | 8 | CI runner |
| Run-time mode | 7 | CI runner |
| Unit | 7 | CI runner |
| Feature | 6 | CI runner |
| Feature | 8 | CI runner; includes time-varying capacity and native warm start |
| Conservative | 2 | CI runner |
| Replay | 5 | CI runner; reclamation safety plus resource equivalence |
| Resource history | 5 | CI runner |
| Job store | 6 | CI runner |
| Config | 9 | CI runner; includes power-usage `trace_type` coverage |
| Native CTest | 12, plus 1 with MPI | CI runner; RNG and binary serialization, trace policies, replay reclamation, custom scheduling, append/streaming APIs, queues, and CLI dispatch |
| Python API | 17 | CI runner |
| Config | 12 | CI runner; includes power-usage, capacity-schedule, and warm-start configuration coverage |
| Native CTest | 15, plus 1 with MPI | CI runner; RNG and binary serialization, trace policies, replay reclamation, custom scheduling, append/streaming APIs, warm starts, capacity parsing, queues, and CLI dispatch |
| Trace tools | 2 checks in 1 CTest registration | CI runner; capacity inference and warm-start boundary/output behavior |
| Python API | 18 | CI runner |
| gRPC client/server | 2 | CI runner |
| Append-job | 22: 19 C++ + 3 gRPC | CI runner |
| FCFS/EASY backfill-window gRPC | 3 repeated checks; 1 targeted | CI runner |
| Append-job | 25: 20 C++ + 5 optional gRPC checks | CI runner |
| FCFS/EASY backfill-window gRPC | 5 repeated checks; 1 targeted | CI runner |
| Single-coordinator gRPC | 1 | CI runner; synchronized independent systems |
| Progressive loading | 11 C++ + 4 CLI | CI runner |
| Queue input schema | 1 binary | CI runner |
| Queue input schema | 1 binary | CI runner; queue variants and replay/simulation runtime validation |
| Scale | 7 | CI runner |

The workflow summary in `tests.yml` is the authoritative CI-oriented list.
Expand Down
22 changes: 11 additions & 11 deletions .github/workflows/tests.yml
Original file line number Diff line number Diff line change
Expand Up @@ -200,7 +200,7 @@ jobs:
id: feature
run: |
echo "========================================"
echo "Running Feature Tests (6 tests)"
echo "Running Feature Tests (8 tests)"
echo "========================================"
./tests/run_feature_tests.sh

Expand Down Expand Up @@ -246,14 +246,14 @@ jobs:
id: config
run: |
echo "========================================"
echo "Running Config Tests (9 tests)"
echo "Running Config Tests (12 tests)"
echo "========================================"
./tests/run_configs_tests.sh

- name: Run Core CTest Tests
id: core-ctest
run: |
echo "Running 12 native CTest registrations plus MPI when available"
echo "Running 15 native CTest registrations, the trace-tools test, plus MPI when available"
cd build
ctest --output-on-failure

Expand All @@ -262,7 +262,7 @@ jobs:
run: |
source venv/bin/activate
echo "========================================"
echo "Running Python API Tests (17 tests)"
echo "Running Python API Tests (18 tests)"
echo "========================================"
PYTHON_EXECUTABLE=$(command -v python) ./tests/run_python_tests.sh

Expand All @@ -278,7 +278,7 @@ jobs:
id: append-job
run: |
echo "========================================"
echo "Running Append-Job Tests (19 C++ + 3 gRPC)"
echo "Running Append-Job Tests (20 C++ + 5 optional gRPC checks)"
echo "========================================"
./tests/run_append_job_tests.sh

Expand Down Expand Up @@ -334,16 +334,16 @@ jobs:
echo " Column Aliases (8 tests): ${{ steps.column-alias.outcome }}"
echo " Run Time Mode (7 tests): ${{ steps.run-time-mode.outcome }}"
echo " Unit (7 tests): ${{ steps.unit.outcome }}"
echo " Feature (6 tests): ${{ steps.feature.outcome }}"
echo " Feature (8 tests): ${{ steps.feature.outcome }}"
echo " Conservative (2 tests): ${{ steps.conservative.outcome }}"
echo " Replay (5 tests): ${{ steps.replay.outcome }}"
echo " Resource History (5 tests): ${{ steps.resource-history.outcome }}"
echo " Job Store (6 tests): ${{ steps.job-store.outcome }}"
echo " Config (9 tests): ${{ steps.config.outcome }}"
echo " Native CTest (12 + MPI when available): ${{ steps.core-ctest.outcome }}"
echo " Python API (17 tests): ${{ steps.python-api.outcome }}"
echo " Append-Job (19 C++ + 3 gRPC): ${{ steps.append-job.outcome }}"
echo " Backfill-window focused rerun (3 checks already counted): ${{ steps.backfill-window.outcome }}"
echo " Config (12 tests): ${{ steps.config.outcome }}"
echo " CTest (15 native + trace tools + MPI when available): ${{ steps.core-ctest.outcome }}"
echo " Python API (18 tests): ${{ steps.python-api.outcome }}"
echo " Append-Job (20 C++ + 5 optional gRPC checks): ${{ steps.append-job.outcome }}"
echo " Backfill-window focused rerun (5 checks already counted): ${{ steps.backfill-window.outcome }}"
echo " Single-Coordinator gRPC: ${{ steps.grpc-single-coordinator.outcome }}"
echo " Progressive Loading (11 C++ + 4 CLI): ${{ steps.progressive-load.outcome }}"
echo " Queue Input Schema: ${{ steps.queue-input-schema.outcome }}"
Expand Down
55 changes: 54 additions & 1 deletion CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -692,6 +692,31 @@ set_target_properties(test_block_queue-bin PROPERTIES CMAKE_INSTALL_RPATH

list(APPEND DR_EVT_UNIT_TEST_TARGETS test_block_queue-bin)

# Time-varying capacity schedule parser validation.
add_executable(test_capacity_schedule-bin tests/test_capacity_schedule.cpp)
target_include_directories(test_capacity_schedule-bin PUBLIC
$<BUILD_INTERFACE:${CMAKE_BINARY_DIR}>
$<BUILD_INTERFACE:${CMAKE_SOURCE_DIR}/src>
$<INSTALL_INTERFACE:${CMAKE_INSTALL_PREFIX}/${CMAKE_INSTALL_INCLUDEDIR}>)
target_link_libraries(test_capacity_schedule-bin PRIVATE dr_evt)
set_target_properties(test_capacity_schedule-bin PROPERTIES
OUTPUT_NAME test_capacity_schedule
CMAKE_INSTALL_RPATH "${CMAKE_INSTALL_PREFIX}/${CMAKE_INSTALL_LIBDIR}")
list(APPEND DR_EVT_UNIT_TEST_TARGETS test_capacity_schedule-bin)

# Native warm-start coverage: boundary classification, capacity transitions,
# statistics, all scheduler policies/queue backends, and differential traces.
add_executable(test_warm_start-bin tests/test_warm_start.cpp)
target_include_directories(test_warm_start-bin PUBLIC
$<BUILD_INTERFACE:${CMAKE_BINARY_DIR}>
$<BUILD_INTERFACE:${CMAKE_SOURCE_DIR}/src>
$<INSTALL_INTERFACE:${CMAKE_INSTALL_PREFIX}/${CMAKE_INSTALL_INCLUDEDIR}>)
target_link_libraries(test_warm_start-bin PRIVATE dr_evt)
set_target_properties(test_warm_start-bin PROPERTIES
OUTPUT_NAME test_warm_start
CMAKE_INSTALL_RPATH "${CMAKE_INSTALL_PREFIX}/${CMAKE_INSTALL_LIBDIR}")
list(APPEND DR_EVT_UNIT_TEST_TARGETS test_warm_start-bin)


# MPI streaming test (requires MPI)
# Only search for MPI for this test - dr_evt library itself has no MPI dependency
Expand All @@ -714,6 +739,7 @@ if(MPI_CXX_FOUND)
endif()

if(DR_EVT_WITH_UNIT_TESTING)
find_package(Python3 COMPONENTS Interpreter QUIET)
add_test(
NAME test_append_job_api
COMMAND $<TARGET_FILE:test_append_job_api-bin>)
Expand All @@ -735,11 +761,38 @@ if(DR_EVT_WITH_UNIT_TESTING)
add_test(
NAME test_block_queue
COMMAND $<TARGET_FILE:test_block_queue-bin>)
add_test(
NAME test_capacity_schedule
COMMAND $<TARGET_FILE:test_capacity_schedule-bin>)
add_test(
NAME test_warm_start
COMMAND $<TARGET_FILE:test_warm_start-bin>)
add_test(
NAME test_warm_start_validation
COMMAND /bin/bash
${CMAKE_CURRENT_SOURCE_DIR}/tests/run_warm_start_validation_tests.sh
$<TARGET_FILE:sim-bin>)
add_test(
NAME test_max_time
COMMAND /bin/bash
${CMAKE_CURRENT_SOURCE_DIR}/tests/run_max_time_tests.sh
$<TARGET_FILE:sim-bin>)
if(Python3_Interpreter_FOUND)
add_test(
NAME test_trace_tools
COMMAND ${CMAKE_COMMAND} -E env
DR_EVT_SIMULATOR=$<TARGET_FILE:sim-bin>
${Python3_EXECUTABLE} ${CMAKE_CURRENT_SOURCE_DIR}/tests/test_trace_tools.py)
set_tests_properties(test_trace_tools PROPERTIES
LABELS "unit;scripts"
WORKING_DIRECTORY ${CMAKE_CURRENT_SOURCE_DIR})
endif()

set_tests_properties(
test_append_job_api test_custom_scheduler test_progressive_load
test_queue_input
test_batch_vs_streaming test_block_queue
test_batch_vs_streaming test_block_queue test_capacity_schedule
test_warm_start test_warm_start_validation test_max_time
PROPERTIES
LABELS "unit;native"
WORKING_DIRECTORY ${CMAKE_CURRENT_SOURCE_DIR}
Expand Down
12 changes: 8 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -63,7 +63,8 @@ minimum input fields are `job_submit_time`, `num_nodes`, and `time_limit`.
job's execution duration is selected separately with `--run_time_mode`:

- `actual` (default) uses `actual_run_time` from the input trace (also accepted as
`duration`, `actual_duration`, or `run_time`).
`actual_runtime`, `duration`, `actual_duration`, or `run_time`). A supplied
value must be finite and no greater than `time_limit`.
- `limit` runs each job for exactly its requested `time_limit`.
- `distribution` draws a duration from the selected `normal`, `lognormal`, or
`uniform` distribution using `--run_time_scale`, `--run_time_stddev`, and
Expand Down Expand Up @@ -186,9 +187,12 @@ job_submit_time,begin_time,end_time,num_nodes,time_limit
0,100,120,60,20
```

The recorded `begin_time` and `end_time` are authoritative. Replay does not
invoke a scheduler or choose new start times; `total_nodes` is used only to
derive the free-node count.
The recorded `begin_time` and `end_time` are authoritative. An optional
`actual_run_time` column (including its accepted aliases) is checked against
`end_time - begin_time`; rows that disagree by more than `1e-6` seconds are
rejected. When the column is absent, replay derives the duration from the two
timestamps. Replay does not invoke a scheduler or choose new start times;
`total_nodes` is used only to derive the free-node count.

Replay writes:

Expand Down
9 changes: 9 additions & 0 deletions docs/CLIENT_SERVER_GUIDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,6 +31,7 @@ in-process via the streaming API:
| Request | Corresponds to |
|---|---|
| `InitRequest` | Constructing a `Simulation` from a `Sim_Params`-equivalent config |
| `RunRequest` | Calls `Simulation::run()`; initialized input and `sim_start_time` select simulation, full replay, or replay-based warm start, and configured `max_time` supplies an inclusive stopping boundary |
| `InitializeTraceRequest` | `Simulation::initialize_trace()` |
| `AppendJobRequest` | `Simulation::append_job()` - a genuinely new job the server has never seen before |
| `AppendJobsRequest` | `Simulation::append_jobs()` - the batch counterpart, several new jobs in one call |
Expand All @@ -40,6 +41,14 @@ in-process via the streaming API:
| `GetBackfillWindowRequest` | One FCFS/EASY reservation snapshot: current capacity, shadow time, and projected releases |
| `GetStatisticsRequest`, `GetCurrentTimeRequest`, etc. | The monitoring/statistics methods |

For a replay-based warm start, set `InitRequest.sim_start_time` to a positive
global simulation boundary, provide a replay-format `infile`, and send
`RunRequest`.
This field is distinct from each job's historical `begin_time`. The server then
applies the same two-stage warm-start classification as the CLI and Python
batch API. A zero simulation start time preserves ordinary batch behavior;
negative and non-finite values are rejected.

Every `ClientMessage` carries a `request_id`, echoed back on the matching
`ServerMessage`, so a client can correlate responses even if it pipelines
multiple in-flight requests (the provided `dr_evt_client` sends one at a
Expand Down
42 changes: 36 additions & 6 deletions docs/TESTING_GUIDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,12 +22,35 @@ plus golden-output comparisons for a targeted simultaneous-backfill case and a
times match the default circular-buffer FCFS scheduler when the callback
selects candidates in FCFS order.

The append-job API test also covers Custom-FCFS time-accounted resource area
and prediction-horizon estimation. Its accounting case combines two
allocations and two releases at one timestamp. Its horizon scenario uses four
running and four waiting jobs to exercise successive running-job completion
boundaries, the post-replay full-capacity tail, the ``U=0`` fallback, future-
arrival exclusion, and invalid utilization values.
The append-job API test also covers Custom-FCFS time-accounted resource area,
capacity-aware instantaneous and aggregate utilization, and
prediction-horizon estimation. Its accounting cases include two allocations
and two releases at one timestamp and a scheduled-capacity reduction below
live occupancy. Its horizon scenario uses four running and four waiting jobs
to exercise successive running-job completion boundaries, the post-replay
full-capacity tail, the ``U=0`` fallback, future-arrival exclusion, and invalid
utilization values. The warm-start suite separately checks integrated
effective capacity across capacity changes at and after the start boundary.

The replay-based warm-start test exercises completed history, live historical jobs,
inherited waiters, boundary and future arrivals, empty tails, fractional and
simultaneous timestamps, every supported priority/backfill/queue combination,
and randomized differential workloads. Runtime-policy cases verify that
``actual``, ``limit``, and ``distribution`` affect only ordinarily scheduled
jobs while warmup jobs retain their recorded end times. Capacity coverage
includes overcommit after a reduction and a capacity change exactly at the
final warmup departure. A separate case protects the documented
``sim_start_time == 0`` behavior: zero disables replay-based warm start and
preserves a traditional full replay. Inclusive ``max_time`` boundaries are checked in
ordinary simulation, replay, and warm-start execution. CLI coverage accepts an ISO simulation start time,
rejects mixed timestamp encodings within one input file, and
rejects negative or non-finite values, simulation-format input, and progressive
file lists.

The queue-input parser test verifies both queue schemas and runtime validity:
a supplied replay runtime must agree with ``end_time - begin_time`` within the
timestamp tolerance, while a supplied simulation runtime must be finite and
no greater than ``time_limit``.

## Test inventory and commands

Expand All @@ -39,6 +62,13 @@ CTest discovers tests enabled by the current build configuration. Optional
features such as Protobuf, gRPC, MPI, Python bindings, and Catch2 add their
corresponding tests only when available.

Run only the replay-based warm-start behavior and CLI validation registrations with:

```bash
ctest --test-dir build --output-on-failure \
-R '^(test_warm_start|test_warm_start_validation)$'
```

## Expected outputs

Scheduler fixtures normally contain:
Expand Down
21 changes: 14 additions & 7 deletions docs/api/PYTHON_API.md
Original file line number Diff line number Diff line change
Expand Up @@ -51,6 +51,8 @@ or `None`.
|---|---|
| `infile` | `str` |
| `total_nodes` | `int` |
| `capacity_schedule` | `str` |
| `sim_start_time` | `float` |
| `trace_format` | `str` |
| `timestamp_format` | `str` |
| `run_time_mode` | `RunTimeMode` |
Expand All @@ -59,6 +61,10 @@ or `None`.
| `priority_policy` | `PriorityPolicy` |
| `verbose` | `bool` |

`sim_start_time` is the global simulation boundary; it is distinct from each
job's historical `begin_time`. A positive value enables replay-based warm start
for replay input, while zero preserves ordinary full replay.

Other C++/CLI configuration fields are not exposed by the binding. Use the
`simulator` executable when one of those settings is required; its options
are documented in [Command-Line Options](../user-guide/command-line.md).
Expand Down Expand Up @@ -91,7 +97,7 @@ Their scheduling semantics are documented in
| `run_until_exclusive(target_time)` | Process events strictly before the target. |
| `get_current_time()` | Return current simulation time. |
| `get_nodes_in_use()` | Return allocated nodes. |
| `get_current_utilization()` | Return instantaneous `nodes_in_use / total_nodes`. |
| `get_current_utilization()` | Return instantaneous usage of effective scheduled capacity. |
| `get_resource_area()` | Return Custom-FCFS allocated-node area in node-seconds; unavailable for standard schedulers. |
| `get_available_nodes()` | Return free nodes. |
| `get_active_job_count()` | Return waiting jobs. |
Expand Down Expand Up @@ -126,12 +132,13 @@ contains `time` and `nodes_released`.

For simulations constructed with Custom-FCFS callbacks, `resource_area` is
accumulated as `nodes_in_use * interval` between settled scheduling times.
`utilization` divides that area by `total_nodes` and the elapsed accounting
horizon (the snapshot time while work is running, or the last resource event
after it becomes idle). Unlike `get_current_utilization()`, intervals that are
far apart therefore carry proportionally more weight. Standard schedulers
retain the post-hoc scheduled-workload calculation and do not perform live
area bookkeeping.
`utilization` divides that area by integrated effective capacity over the
accounting horizon (the snapshot time while work is running, or the last
resource event after it becomes idle). During non-preemptive draining,
effective capacity is at least the running allocation. Unlike
`get_current_utilization()`, intervals that are far apart therefore carry
proportionally more weight. Standard schedulers retain the post-hoc
scheduled-workload calculation and do not perform live area bookkeeping.

Metric definitions are in
[Output Trace Files](../user-guide/output-traces.md#cli-summary).
Expand Down
Loading