Scinovex
reviewTop 1% cited

A Survey of Deep Active Learning

ACM Computing Surveys · 2021 · Vol. 54(9) · pp. 1–40
Pengzhen RenYun XiaoXiaojun ChangPo-Yao HuangZhihui LiBrij B. GuptaXiaojiang ChenXin Wang

Abstract

Active learning (AL) attempts to maximize a model’s performance gain while annotating the fewest samples possible. Deep learning (DL) is greedy for data and requires a large amount of data supply to optimize a massive number of parameters if the model is to learn how to extract high-quality features. In recent years, due to the rapid development of internet technology, we have entered an era of information abundance characterized by massive amounts of available data. As a result, DL has attracted significant attention from researchers and has been rapidly developed. Compared with DL, however, researchers have a relatively low interest in AL. This is mainly because before the rise of DL, traditional machine learning requires relatively few labeled samples, meaning that early AL is rarely according the value it deserves. Although DL has made breakthroughs in various fields, most of this success is due to a large number of publicly available annotated datasets. However, the acquisition of a large number of high-quality annotated datasets consumes a lot of manpower, making it unfeasible in fields that require high levels of expertise (such as speech recognition, information extraction, medical images, etc.). Therefore, AL is gradually coming to receive the attention it is due. It is therefore natural to investigate whether AL can be used to reduce the cost of sample annotation while retaining the powerful learning capabilities of DL. As a result of such investigations, deep active learning (DeepAL) has emerged. Although research on this topic is quite abundant, there has not yet been a comprehensive survey of DeepAL-related works; accordingly, this article aims to fill this gap. We provide a formal classification method for the existing work, along with a comprehensive and systematic overview. In addition, we also analyze and summarize the development of DeepAL from an application perspective. Finally, we discuss the confusion and problems associated with DeepAL and provide some possible development directions.

Machine Learning and AlgorithmsMachine Learning and Data ClassificationOil and Gas Production TechniquesComputer scienceDeep learningArtificial intelligenceAnnotationMachine learningThe InternetActive learning (machine learning)Quality (philosophy)Data scienceWorld Wide Web

Funding

  • National Natural Science Foundation of China
Citations
1,001
FWCI
103.86
field-weighted impact
References
157
Percentile
100%
vs. same field & year
Citations per year
References
Learning representations by back-propagating errors
Nature · 1986 · 30,045 citations
Pedestrian Detection: An Evaluation of the State of the Art
IEEE Transactions on Pattern Analysis and Machine Intelligence · 2011 · 3,229 citations
The Pascal Visual Object Classes (VOC) Challenge
International Journal of Computer Vision · 2009 · 19,127 citations
Indications of nonlinear deterministic and finite-dimensional structures in time series of brain electrical activity: Dependence on recording region and brain state
Physical review. E, Statistical physics, plasmas, fluids, and related interdisciplinary topics · 2001 · 2,918 citations
A theory of learning from different domains
Machine Learning · 2009 · 3,384 citations
Gradient-based learning applied to document recognition
Proceedings of the IEEE · 1998 · 57,014 citations
A Fast Learning Algorithm for Deep Belief Nets
Neural Computation · 2006 · 16,253 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.