Skip to content

Repository files navigation

Discrete Resource Event Modeling and Multi-cluster Scheduling Simulator

Documentation Status License

DR_EVT is a high-performance HPC job scheduler simulator supporting EASY and CONSERVATIVE backfilling algorithms. Uniquely supports online simulation via gRPC, enabling coordinated multi-cluster simulations where distributed schedulers interact in real-time. Another use case is the Fugaku Power-Usage Simulation Experiment.

📚 Read the Full Documentation on ReadTheDocs →

Features

Scheduling and traces

  • Scheduling: FCFS, SJF, and LJF priority policies with EASY, conservative, or no backfilling.
  • Execution modes: scheduler-driven simulation and replay of historical or previously generated schedules.
  • Run-time models: recorded actual duration, requested time limit, or normal, lognormal, and uniform distributions with reproducible seeds.
  • Interfaces: command-line batch execution, incremental C++ streaming, Python bindings, and a gRPC service.
  • Trace support: simple CSV and LLNL Lassen inputs; scheduled-job, resource-usage, and optional power-usage outputs.
  • Implementation choices: deque, multimap, circular, and block wait queues for comparison and scaling studies.

Build and quick start

Requirements

A basic Linux build requires CMake 3.24 or later, a C++20 compiler, and Boost 1.70 or later. Ser20 supplies RNG-state serialization and is fetched during configuration when it is not already installed or present under external/ser20. If Boost is unavailable, CMake can fetch it too. The Installation Guide covers dependency selection, optional features, and platform-specific setup.

Build

git clone https://github.com/LLNL/dr_evt.git
cd dr_evt

export CMAKE_INSTALL_PREFIX=/path/to/install
cmake -S . -B build \
  -DCMAKE_BUILD_TYPE=Release \
  -DCMAKE_INSTALL_PREFIX="${CMAKE_INSTALL_PREFIX}"
cmake --build build -j4
cmake --install build

Simulation and replay

Simulation

simulator reads job submissions and invokes the selected scheduler. Its minimum input fields are job_submit_time, num_nodes, and time_limit. time_limit is always the scheduler's estimate for reservation planning. The job's execution duration is selected separately with --run_time_mode:

  • actual (default) uses actual_run_time from the input trace (also accepted as actual_runtime, duration, actual_duration, or run_time). A supplied value must be finite and no greater than time_limit.
  • limit runs each job for exactly its requested time_limit.
  • distribution draws a duration from the selected normal, lognormal, or uniform distribution using --run_time_scale, --run_time_stddev, and --seed.

Normal and lognormal samples are bounded by the requested time limit. Uniform sampling uses its configured lower and upper bounds directly. Simulation produces scheduled-job and resource traces.

Example: run a simulation

The included example submits two 60-node jobs at time zero to a 100-node system. Its input file, tests/test_traces/unit/simple_2jobs.csv, contains:

job_submit_time,num_nodes,time_limit
0,60,100
0,60,20

Run it with:

${CMAKE_INSTALL_PREFIX}/bin/simulator tests/test_traces/unit/simple_2jobs.csv \
  --trace_format simple \
  --timestamp_format epoch \
  --run_time_mode limit \
  --total_nodes 100 \
  --outfile results.csv \
  --resource_trace resources.csv

Because both jobs cannot fit at once, the second waits until the first finishes. results.csv contains the resulting job schedule:

job_submit_time,begin_time,end_time,num_nodes,exit_status,time_limit
0,0,100,60,0,100
0,100,120,60,0,20

resources.csv records every allocation change:

time,free_nodes,allocated_nodes
0,100,0
0,40,60
100,100,0
100,40,60
120,100,0

The terminal also reports submitted and completed jobs, average wait and turnaround time, makespan, average and peak queue length, and wall-clock execution time. Inspect generated files directly with:

head results.csv
head resources.csv

DR_EVT accepts a job trace or an ordered list of trace files for progressive loading. Settings may come from command-line options, a Protobuf text configuration, or both. Complete input schemas and output definitions are in Input Trace Files and Output Trace Files.

The defaults use FCFS priority and EASY backfilling. For strict FCFS without backfilling, for example, use:

${CMAKE_INSTALL_PREFIX}/bin/simulator input.csv \
  --priority_policy fcfs \
  --backfill_policy none

Using Protocol Buffer configuration files

A build configured with -DDR_EVT_ENABLE_PROTOBUF=ON can read simulator options from a Protobuf text file. For example, sim_config.textproto can contain:

outfile: "results.csv"
resource_trace: "resources.csv"
total_nodes: 1000
backfill_policy: "easy"
priority_policy: "fcfs"
trace_format: "simple"
timestamp_format: "epoch"
run_time_mode: "actual"
seed: 42

Run it with the input trace as the positional argument:

${CMAKE_INSTALL_PREFIX}/bin/simulator trace.csv \
  --config sim_config.textproto

Options written after --config override values from the file. See Protobuf Configuration for all fields and progressive-input configuration.

For a guided example, see the Quick Start Guide. All flags are listed in the Command-Line Reference.

Replay

tracer reconstructs resource use from an existing historical or simulated schedule. A simple replay input requires:

job_submit_time,begin_time,end_time,num_nodes,time_limit
0,0,100,60,100
0,100,120,60,20

The recorded begin_time and end_time are authoritative. An optional actual_run_time column (including its accepted aliases) is checked against end_time - begin_time; rows that disagree by more than 1e-6 seconds are rejected. When the column is absent, replay derives the duration from the two timestamps. Replay does not invoke a scheduler or choose new start times; total_nodes is used only to derive the free-node count.

Replay writes:

  • --resource_trace: free and allocated nodes over time; and
  • optional per-job, submission-analysis, and DAT-period reports when requested.

It also prints the number of loaded jobs, trace span, and wall-clock processing time. The number of weeks is printed when a submission report is requested.

Use simulation for policy comparisons and capacity studies. Use replay to analyze a historical or precomputed schedule without changing its scheduling decisions.

# Produce a schedule.
${CMAKE_INSTALL_PREFIX}/bin/simulator input.csv \
  --total_nodes 100 \
  --outfile schedule.csv \
  --resource_trace simulated-resources.csv

# Replay that schedule without invoking a scheduler.
${CMAKE_INSTALL_PREFIX}/bin/tracer \
  --infile schedule.csv \
  --total_nodes 100 \
  --resource_trace replay-resources.csv

See Input Trace Files for the fields needed by each mode and Output Trace Files for the optional analysis reports.

Optional interfaces

  • Enable Python bindings with -DDR_EVT_BUILD_PYTHON=ON for in-process Python control.
  • Enable the network service with -DDR_EVT_ENABLE_GRPC=ON; this also enables Protobuf. The server gives each client session an isolated simulation.
  • MPI is used only by the optional multi-client/multi-server test harness.

For example, an all-interface build uses:

cmake -S . -B build \
  -DCMAKE_BUILD_TYPE=Release \
  -DCMAKE_INSTALL_PREFIX="${CMAKE_INSTALL_PREFIX}" \
  -DDR_EVT_BUILD_PYTHON=ON \
  -DDR_EVT_ENABLE_GRPC=ON
cmake --build build -j4
cmake --install build

See the Python API and gRPC setup guide for usage.

APIs and integration

  • C++ streaming API: append jobs and advance simulation time incrementally.
  • Python bindings: control an in-process simulation from Python.
  • gRPC service: expose the streaming API to remote clients and distributed controllers.

See the User Guide for the documentation map.

Most users can run the native C++ simulator directly, as shown in the Quick Start, or use the in-process Python API. Neither interface requires client/server setup; the distributed gRPC service below is optional.

Distributed gRPC deployment

Clients and digital-twin controllers can open independent gRPC sessions to any number of server processes. Each session owns an isolated simulation, so the numbers of clients and servers can be scaled independently. Containers are available for remote deployment, while MPI provides an optional test and experiment launcher.

Clients connect over gRPC to servers that own independent simulations

Deployment patterns are described in Client/Server Use Cases. The internal execution flow is described in Simulation Pipeline and Job Lifecycle.

Documentation

The documentation build requires Python 3 and the packages in docs/requirements.txt. Doxygen is required to include the generated C++ API reference, and Graphviz is required for its diagrams.

python3 -m venv .venv-docs
source .venv-docs/bin/activate
python -m pip install -r docs/requirements.txt
make -C docs html

The generated site is written to docs/_build/html/index.html.

Documentation index

Docker and rootless Podman packaging are available; setup details remain with the Docker and Podman files.

Contributing

Please submit bug fixes and improvements as pull requests. Testing procedures are documented in the Testing Guide.

License and attribution

DR_EVT is distributed under the terms of the MIT license. See LICENSE and NOTICE.

  • SPDX-License-Identifier: MIT
  • LLNL-CODE-844050

Thanks to DR_EVT's contributors.

About

Discrete Resource Event Modeling and Multi-Cluster Job Scheduling Simulation

Resources

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages