StageML: Partial Evaluation for Multi-Tenant MoE-LoRA Inference
Multi-tenant large language model platforms increasingly serve heterogeneous Low-Rank Adaptation (LoRA) modules over shared Mixture-of-Experts (MoE) backbones. Existing serving systems treat multi-tenant LoRA inference as a pure runtime scheduling problem, ignoring the temporal boundaries between deployment, adapter loading, tenant request admission, expert routing, and token generation. This paper introduces StageML, a partial evaluation (PE) system that exploits these temporal boundaries using a binding-time lattice. StageML residualizes adapter computations that depend only on early-stage values, leaving token-dependent routing dynamic. A memory-aware planner materializes hot adapters under VRAM budgets. On Mixtral 8×7B with 2GB budget, StageML removes 52.34% of adapter computation from token time under a Zipfian workload (where a few popular items get most of the traffic), achieving 1.105 times speedup. When residualized plans run through an optimized fused-experts backend, the expert computation phase sees 2.806 times speedup, an upper bound for this bottleneck. The results show that PE complements low-level kernel optimization.
Sat 29 AugDisplayed time zone: Eastern Time (US & Canada) change
14:00 - 15:30 | |||
14:00 30mTalk | StageML: Partial Evaluation for Multi-Tenant MoE-LoRA Inference LOPSTR+PPDP | ||
14:30 30mTalk | Complex Autonomous UAV Task Execution and Decision-Making Using s(CASP) LOPSTR+PPDP Keegan Kimbrell University of Texas at Dallas, USA, Alexis R Tudor University of Texas at Dallas, Peter Vu University of Texas at Dallas, Trevor Bihl Ohio University, Doug Slattery SYBOR Tech, Inc., Gopal Gupta University of Texas at Dallas | ||
15:00 30mTalk | Vehicle with Time: Signal First-Order Logic for Closed-Loop Controller Synthesis LOPSTR+PPDP Gusts Gustavs Grīnbergs IT University of Copenhagen, Alessandro Bruni IT University of Copenhagen, Matthew L. Daggitt University of Western Australia | ||