A reinforcement learning and LLM benchmarking environment for Kubernetes resource optimization using SimKube simulations.
Sim-Arena is a gym-like environment where AI agents learn to fix Kubernetes pod failures by adjusting resource requests (CPU, memory) and replica counts. It supports:
- Reinforcement Learning Agents: Train DQN, Epsilon-Greedy, or custom agents.
- LLM Benchmarking: Evaluate large language models (Gemini, Claude) on the same scenarios using MCP tools.
- Distributed Training: Scale training across multiple EC2 workers with federated averaging.
- Gymnasium Integration: Use standard RL interfaces for easy integration.
The environment simulates failing Kubernetes workloads, observes pod states, applies actions, and provides rewards based on pod health.
- Python 3.8+
- A Kubernetes cluster with SimKube installed (see Setup Guide)
kubectlaccess to the cluster
-
Clone the repository:
git clone https://github.com/bobg0/sim-arena.git cd sim-arena -
Install dependencies:
pip install -r requirements.txt
-
Verify your cluster:
make preflight
Train a DQN agent on a demo scenario:
python runner/train.py --trace demo/trace-0001.msgpack --ns virtual-default --target 3 --agent dqn --episodes 10 --steps 5This will:
- Start a simulation of a failing workload
- Train a DQN agent to fix pod issues
- Save checkpoints and logs to
checkpoints/
Evaluate Gemini on the same scenarios:
python benchmark/run.py --provider gemini --ns virtual-defaultRequires GEMINI_API_KEY environment variable.
Sim-Arena provides a standard Gymnasium environment:
from env.simkube_gymenv import SimKubeEnv
env = SimKubeEnv(
trace_path="demo/trace-0001.msgpack",
namespace="virtual-default",
target_pods=3,
max_steps=10
)
obs = env.reset()
done = False
while not done:
action = your_agent.act(obs) # 0-6: noop, bump_cpu_small, etc.
obs, reward, done, info = env.step(action)Implement the Agent interface:
from agent.agent import Agent, AgentType
# For RL agents
agent = Agent(AgentType.DQN, state_dim=5, n_actions=7)
# For LLM agents
agent = Agent(AgentType.LLM, provider="gemini")Available actions (7 total):
- 0: No-op
- 1-2: Bump CPU (small/large)
- 3-4: Bump memory (small/large)
- 5-6: Scale replicas (up/down)
Dict with pod counts:
{
"ready": 0,
"pending": 3,
"total": 3
}base: 1 if all pods ready and none pending, else 0shaped: Continuous reward based on progresscost_aware_v2: Penalizes resource wastemax_punish: Base + hard limits on resources
- Gymnasium Integration: Detailed Guide
- LLM Benchmarking: MCP Tools Guide
- Distributed Training: AWS Setup
- Custom Scenarios: Generate traces with
demo/generate_traces.py
- Developer Guide - Detailed internals for contributors
- Worker Protocol - Distributed training protocol
- EC2 Setup - Cluster setup instructions
See Developer Guide for development setup and architecture details.