Xu, R., Mu, F., Lee, J., Mukherjee, P., Chaterji, S., Bagchi, S., and Li, Y. (2022). SMARTADAPT: Multi-branch Object Detection Framework for Videos on Mobiles. Proceedings of the Conference on Computer Vision and Pattern Recognitions, IEEE/CVF 2022.
Lead developers
Ran Xu and Jay Lee (Purdue), FangZhou Mu (Wisconson)
Students
Ran Xu, Fangzhou Mu, Jayoung Lee, Preeti Mukherjee, Yin Li
Abstract
Several recent works seek to create lightweight deep networks for video object detection on
mobiles. We observe that many existing detectors, previously deemed computationally costly for
mobiles, intrinsically support adaptive inference, and offer a multi-branch object detection
framework (MBODF). Here, an MBODF is referred to as a solution that has many execution branches
and one can dynamically choose from among them at inference time to satisfy varying latency
requirements (e.g. by varying resolution of an input frame). In this paper, we ask, and answer,
the wide-ranging question across all MBODFs: How to expose the right set of execution branches
and then how to schedule the optimal one at inference time? In addition, we uncover the importance
of making a content-aware decision on which branch to run, as the optimal one is conditioned on
the video content. Finally, we explore a content-aware scheduler, an Oracle one, and then a
practical one, leveraging various lightweight feature extractors. Our evaluation shows that
layered on Faster R-CNN-based MBODF, compared to 7 baselines, our SMARTADAPT achieves a higher
Pareto optimal curve in the accuracy-vs-latency space for the ILSVRC VID dataset.
