AISTATS Poster On the Power of Multitask Representation Learning with Gradient Descent

Poster

On the Power of Multitask Representation Learning with Gradient Descent

Joshua Agterberg · Hung-Hsu Chou · Danqi Liao · Zixiang Chen

[ Abstract ]

Abstract:

Representation learning, particularly multi-task representation learning, has gained widespread popularity in various deep learning applications, ranging from computer vision to natural language processing, due to its remarkable generalization performance. Despite its growing use, our understanding of the underlying mechanisms remains limited. In this paper, we provide a theoretical analysis elucidating why multi-task representation learning outperforms its single-task counterpart in scenarios involving over-parameterized two-layer convolutional neural networks trained by gradient descent. Our analysis is based on a data model that encompasses both task-shared and task-specific features, a setting commonly encountered in real-world applications. We also present experiments on synthetic and real-world data to illustrate and validate our theoretical findings.

Live content is unavailable. Log in and register to view live content