Skip to content

feat: add tinker API support - #1

Open
GavinZhu-GMI wants to merge 10 commits into
mainfrom
fix_gradaccu_megatron
Open

feat: add tinker API support #1
GavinZhu-GMI wants to merge 10 commits into
mainfrom
fix_gradaccu_megatron

Conversation

@GavinZhu-GMI

Copy link
Copy Markdown
Owner

No description provided.

Signed-off-by: Gavin.Zhu <gavin.z@gmicloud.ai>
Signed-off-by: Gavin.Zhu <gavin.z@gmicloud.ai>
Signed-off-by: Gavin.Zhu <gavin.z@gmicloud.ai>
Signed-off-by: Gavin.Zhu <gavin.z@gmicloud.ai>
Signed-off-by: Gavin.Zhu <gavin.z@gmicloud.ai>
Signed-off-by: Gavin.Zhu <gavin.z@gmicloud.ai>
Signed-off-by: Gavin.Zhu <gavin.z@gmicloud.ai>
  Add dynamic batch size handling to support variable-sized training requests
  from external services (e.g., tinker-cookbook) while preserving gradient
  accumulation semantics.

  **Problem:**
  Slime was designed for standalone training with fixed batch sizes configured
  at initialization. When used as a backend service, clients send variable
  batch sizes per request, causing mismatches between configured and actual
  batch sizes.

  **Solution:**

  1. data.py (lines 181-198):
     - Detect when actual batch size differs from configured global_batch_size
     - Small batches (≤ target): Process in 1 step without gradient accumulation
     - Large batches (> target): Enable multi-step gradient accumulation
     - Pass _actual_global_batch_size to loss.py for correct loss scaling

  2. loss.py (lines 524-529):
     - Use actual batch size for loss scaling instead of configured value
     - Prevents incorrect gradient magnitudes when batch sizes vary

Signed-off-by: Gavin.Zhu <gavin.z@gmicloud.ai>
Signed-off-by: Gavin.Zhu <gavin.z@gmicloud.ai>
…ients and returning logprobs per sample

Signed-off-by: Gavin.Zhu <gavin.z@gmicloud.ai>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant