This is an outdated version published on 2023-08-05. Read the most recent version.
Preprint / Version 1

Unified Deep Learning

##article.authors##

  • Wei Wang Ren min University

DOI:

https://doi.org/10.31224/3156

Abstract

With the increase in depth and complexity, deep learning networks have made significant progress. However, the theoretical understanding of deep learning remains incomplete. Early works in this field successfully demonstrated the Universal Approximation Theorem, but these proofs were limited to linear-type networks. This paper aims to bridge the gap in theoretical understanding of deep learning by proposing a unified approach. Surprisingly, all these network architectures (Linear, Convolutional, and Transformer) can be represented within this unified framework as a matrix multiplying a vector. Furthermore, this paper proves that multi-layer neural networks, including Linear, Convolutional, and Transformer networks, also satisfy the original approximation theorem. The key differentiating factor among these multi-layer networks lies in the form of parameters they learn. Linear networks learn dense matrices, while Convolutional networks and Transformers learn sparse matrices, which enhances their effectiveness in certain tasks and significantly reduces the parameter requirements. By establishing this unified framework and proving the applicability of the approximation theorem to multi-layer networks, this paper takes an important step towards unifying the entire field of deep learning. It deepens our theoretical understanding and reveals the fundamental principles governing the remarkable capabilities of these networks. It also paves the way for exploring new research avenues and optimizing the learning process in various deep learning applications.

Downloads

Download data is not yet available.

Downloads

Posted

2023-08-05

Versions