Bug. _BoundedPromptEmbedCache (openwam/deploy/engine.py) crashes with KeyError on its first eviction once more than DEFAULT_PROMPT_EMBED_CACHE_MAXSIZE (32) distinct prompts have been encoded. In a long-running serving process this deterministically kills inference for every new prompt from that point on.
Root cause. The class overrides getitem for LRU recency tracking:
def getitem(self, key: Any) -> Any:
value = super().getitem(key)
self.move_to_end(key) # KeyError if key is no longer in the ordering map
return value
OrderedDict.popitem(last=False) — used by setitem to evict — dispatches the value lookup for the entry being evicted to the subclass's getitem. At that point the key has already been removed from the internal ordering map, so move_to_end(key) raises.
Minimal repro (stdlib only, Python 3.10; class copied from engine.py with maxsize=3):
from collections import OrderedDict
class BoundedPromptEmbedCache(OrderedDict):
def init(self, maxsize=3):
super().init()
self._maxsize = maxsize
def getitem(self, key):
value = super().getitem(key)
self.move_to_end(key)
return value
def setitem(self, key, value):
if key in self:
self.move_to_end(key)
super().setitem(key, value)
while len(self) > self._maxsize:
self.popitem(last=False)
d = BoundedPromptEmbedCache(3)
for k in ["probe", "a", "b", "c"]:
d[k] = k # raises KeyError: 'probe' on the 4th distinct key
Production call path (serving; exception is caught and reported per-request by the policy server):
JointInferenceEngine.generate → BaseWAMArchitecture.generate → WanVideoBackbone.preprocess_input_for_inference → encode_text_for_inference → prompt_embed_cache[prompt] = (context, seq_lens) → KeyError during eviction.
Impact. The serving process stays up (the error is caught per request, so a supervisor never restarts it), but the cache remains at/over maxsize: every subsequent insertion of a new prompt crashes identically; only previously cached prompts keep working. Workloads with many distinct instructions (e.g. parameterized per-episode prompts) hit this within ~30 episodes; static-prompt workloads may never hit it, which makes the failure look intermittent across jobs.
Suggested fix (verified against the repro above; LRU order preserved):
def setitem(self, key: Any, value: Any) -> None:
if key in self:
self.move_to_end(key)
super().setitem(key, value)
while len(self) > self._maxsize:
evicted_key = next(iter(self))
super().delitem(evicted_key)
if not self._evict_warned:
self._evict_warned = True
logger.warning("prompt_embed_cache exceeded maxsize=%d; evicted %r",
self._maxsize, evicted_key)
Evicting via next(iter(self)) + super().delitem() avoids the popitem → overridden getitem dispatch entirely.
Bug. _BoundedPromptEmbedCache (openwam/deploy/engine.py) crashes with KeyError on its first eviction once more than DEFAULT_PROMPT_EMBED_CACHE_MAXSIZE (32) distinct prompts have been encoded. In a long-running serving process this deterministically kills inference for every new prompt from that point on.
Root cause. The class overrides getitem for LRU recency tracking:
def getitem(self, key: Any) -> Any:
value = super().getitem(key)
self.move_to_end(key) # KeyError if key is no longer in the ordering map
return value
OrderedDict.popitem(last=False) — used by setitem to evict — dispatches the value lookup for the entry being evicted to the subclass's getitem. At that point the key has already been removed from the internal ordering map, so move_to_end(key) raises.
Minimal repro (stdlib only, Python 3.10; class copied from engine.py with maxsize=3):
from collections import OrderedDict
class BoundedPromptEmbedCache(OrderedDict):
def init(self, maxsize=3):
super().init()
self._maxsize = maxsize
def getitem(self, key):
value = super().getitem(key)
self.move_to_end(key)
return value
def setitem(self, key, value):
if key in self:
self.move_to_end(key)
super().setitem(key, value)
while len(self) > self._maxsize:
self.popitem(last=False)
d = BoundedPromptEmbedCache(3)
for k in ["probe", "a", "b", "c"]:
d[k] = k # raises KeyError: 'probe' on the 4th distinct key
Production call path (serving; exception is caught and reported per-request by the policy server):
JointInferenceEngine.generate → BaseWAMArchitecture.generate → WanVideoBackbone.preprocess_input_for_inference → encode_text_for_inference → prompt_embed_cache[prompt] = (context, seq_lens) → KeyError during eviction.
Impact. The serving process stays up (the error is caught per request, so a supervisor never restarts it), but the cache remains at/over maxsize: every subsequent insertion of a new prompt crashes identically; only previously cached prompts keep working. Workloads with many distinct instructions (e.g. parameterized per-episode prompts) hit this within ~30 episodes; static-prompt workloads may never hit it, which makes the failure look intermittent across jobs.
Suggested fix (verified against the repro above; LRU order preserved):
def setitem(self, key: Any, value: Any) -> None:
if key in self:
self.move_to_end(key)
super().setitem(key, value)
while len(self) > self._maxsize:
evicted_key = next(iter(self))
super().delitem(evicted_key)
if not self._evict_warned:
self._evict_warned = True
logger.warning("prompt_embed_cache exceeded maxsize=%d; evicted %r",
self._maxsize, evicted_key)
Evicting via next(iter(self)) + super().delitem() avoids the popitem → overridden getitem dispatch entirely.