Model optimizations help improve inference performance and accuracy of ML workflows. However, relying on a single model to perform inference across all data batches often fails to maximize accuracy and thus overall performance. In many cases, alternate models could perform better on specific subsets of data where a primary model underperforms. Our experiments with real ML workflows indeed show that switching models improves workflow accuracy by up to 23%. Yet, current systems lack the ability to adaptively switch between models based on performance, forcing users to manually test models in sequence. We present FlexiFlow, a dataflow system that dynamically switches between alternate models when the current model exhibits low accuracy. FlexiFlow learns to rank models using a novel multi-armed bandit approach that accounts for model runtimes, probability of passing user-defined assertions, and the computational structure of the ML workflow. We show that the standard Thompson sampling approach is insufficient for switching models in ML workflows. In contrast, our proposed approaches are effective and scales to complex real-world ML workflows. Experiments show that switching models at runtime while reusing intermediate results provides higher accuracy, but also 48% efficiency gain compared to sequential workflow runs.
Figures & tables
Figure 1. ID Card ML workflow.
Figure 2. ID Card workflow in FlexiFlow . Choices for each workflow step are indicated below in red italics. Soft failure of assertions necessitates switching (dashed arrows) between model choices.
Figure 3. FlexiFlow data model. Some operators are omitted for clarity.
Figure 4. FlexiFlow Architecture. Optimizer maintains estimates for each operator choice and set of operator choices.
Figure 5. Probability density functions over mean rewards
Strategy
Action Level
Operator Params
Config Params
Params Count
Rollback-Aware
Single operator with multiple choices
OTS
Operator
[0.2em] αj,βjμj,σj2
O(C)
Multiple operators with multiple choices
CTS
[0.2em] Config
μj,σj2
αk,βk
O(C+K)
✓
HTS
[0.2em] Config
[0.2em] αj,βjμj,σj2
[0.2em] αk′,βk′αkem,βkem
O(C+K)
✓
Table 1. Comparison of different Thompson Sampling based FlexiFlow strategies. C is the total count of operator choices across operators and K is the count of configurations.
Figure 6. Real ML workflows for evaluating FlexiFlow .
Figure 7. Synthetic workflows for evaluating FlexiFlow strategies.
Figure 8. Modeling runtime and being rollback aware is important for strategies to learn.
Figure 9. FlexiFlow optimizers can effectively improve runtimes for real-world ML workflows.
Figure 10. Effectiveness of using operator choices in ML workflows. FlexiFlow achieves a higher F1 score than the best configuration.
Figure 11. Failure modes of LCQA configurations are uncorrelated.
Figure 12. Comparing runtime of FlexiFlow with simulating rollback in other dataflow systems.