Full-parameter fine-tuning goes multi-tenant: slashing accelerator footprint by 67% with llm-d time-slicing
When we introduced co-operative time-slicing in llm-d, we made a claim: if RL phases become schedulable units, independent jobs can share accelerators with near-zero waste. Today we're backing that claim with a measured, end-to-end proof.
OpenRL, an open source, Kubernetes-native, Tinker-compatible, self-hosted fine-tuning service built on the llm-d time-slicing stack, runs supervised fine-tuning (SFT) and full-parameter reinforcement learning for multiple tenants concurrently on the same GPUs.














