Skip to content

Consistent, accurate recommended VRAM allocation per device + model #604

Description

@lucbruni-amd

Recommended VRAM guidance is currently hand-written per playbook, which is one-size-fits-all and open to human error as models and devices vary. We need a way to consistently resolve a recommended VRAM allocation from device + model load and surface it to the user.

Considerations for whatever solution we land on:

  • Inputs available from GGUF metadata + context; overhead a tunable constant
  • CI could print measured VRAM as a drift check to catch calc divergence

Metadata

Metadata

Labels

No labels
No labels

Type

No type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions