Skip to content

deploy/engine.py: _BoundedPromptEmbedCache raises KeyError on first LRU eviction #36

Description

@JadeYyang

Bug. _BoundedPromptEmbedCache (openwam/deploy/engine.py) crashes with KeyError on its first eviction once more than DEFAULT_PROMPT_EMBED_CACHE_MAXSIZE (32) distinct prompts have been encoded. In a long-running serving process this deterministically kills inference for every new prompt from that point on.

Root cause. The class overrides getitem for LRU recency tracking:

def getitem(self, key: Any) -> Any:
value = super().getitem(key)
self.move_to_end(key) # KeyError if key is no longer in the ordering map
return value

OrderedDict.popitem(last=False) — used by setitem to evict — dispatches the value lookup for the entry being evicted to the subclass's getitem. At that point the key has already been removed from the internal ordering map, so move_to_end(key) raises.

Minimal repro (stdlib only, Python 3.10; class copied from engine.py with maxsize=3):

from collections import OrderedDict

class BoundedPromptEmbedCache(OrderedDict):
def init(self, maxsize=3):
super().init()
self._maxsize = maxsize
def getitem(self, key):
value = super().getitem(key)
self.move_to_end(key)
return value
def setitem(self, key, value):
if key in self:
self.move_to_end(key)
super().setitem(key, value)
while len(self) > self._maxsize:
self.popitem(last=False)

d = BoundedPromptEmbedCache(3)
for k in ["probe", "a", "b", "c"]:
d[k] = k # raises KeyError: 'probe' on the 4th distinct key

Production call path (serving; exception is caught and reported per-request by the policy server):
JointInferenceEngine.generate → BaseWAMArchitecture.generate → WanVideoBackbone.preprocess_input_for_inference → encode_text_for_inference → prompt_embed_cache[prompt] = (context, seq_lens) → KeyError during eviction.

Impact. The serving process stays up (the error is caught per request, so a supervisor never restarts it), but the cache remains at/over maxsize: every subsequent insertion of a new prompt crashes identically; only previously cached prompts keep working. Workloads with many distinct instructions (e.g. parameterized per-episode prompts) hit this within ~30 episodes; static-prompt workloads may never hit it, which makes the failure look intermittent across jobs.

Suggested fix (verified against the repro above; LRU order preserved):

def setitem(self, key: Any, value: Any) -> None:
if key in self:
self.move_to_end(key)
super().setitem(key, value)
while len(self) > self._maxsize:
evicted_key = next(iter(self))
super().delitem(evicted_key)
if not self._evict_warned:
self._evict_warned = True
logger.warning("prompt_embed_cache exceeded maxsize=%d; evicted %r",
self._maxsize, evicted_key)

Evicting via next(iter(self)) + super().delitem() avoids the popitem → overridden getitem dispatch entirely.

Activity

  1. wayrise commented on Sep 23, 2026

    @wayrise
    Contributor

    Thanks for the detailed report and reproducer.

    Fixed in #37, now merged into main as f6d9f10.

    We reproduced the first-eviction KeyError in the Python 3.10 environment. The fix selects the oldest key with next(iter(self)) and removes it via super().__delitem__, avoiding the overridden __getitem__ during eviction while preserving LRU behavior.

    The regression test covers the first and subsequent evictions, recency updates on reads, and the one-time eviction warning. PR CI checks (lint, tests, and build-and-test) all passed. Before submission, the full local test suite also passed: 2066 passed, 17 skipped.

    Please update to the latest main and restart any running serving process to load the fix. Closing this issue as resolved.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions