In deep-learning-based audio source separation, conventional research has primarily focused on architectural design improvements (e.g., Conv-TasNet, Band-Split RNN, BS-RoFormer, SCNet). These models operate under single-step inference (single-input single-output or single-input multi-output configurations), where a single forward pass estimates target sources directly from the mixture input.