Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:2212.03480 (eess)

[Submitted on 7 Dec 2022]

Title:Progressive Multi-Scale Self-Supervised Learning for Speech Recognition

Authors:Genshun Wan, Tan Liu, Hang Chen, Jia Pan, Cong Liu, Zhongfu Ye

View PDF

Abstract:Self-supervised learning (SSL) models have achieved considerable improvements in automatic speech recognition (ASR). In addition, ASR performance could be further improved if the model is dedicated to audio content information learning theoretically. To this end, we propose a progressive multi-scale self-supervised learning (PMS-SSL) method, which uses fine-grained target sets to compute SSL loss at top layer while uses coarse-grained target sets at intermediate layers. Furthermore, PMS-SSL introduces multi-scale structure into multi-head self-attention for better speech representation, which restricts the attention area into a large scope at higher layers while restricts the attention area into a small scope at lower layers. Experiments on Librispeech dataset indicate the effectiveness of our proposed method. Compared with HuBERT, PMS-SSL achieves 13.7% / 12.7% relative WER reduction on test other evaluation subsets respectively when fine-tuned on 10hours / 100hours subsets.

Comments:	Submitted to ICASSP 2023
Subjects:	Audio and Speech Processing (eess.AS); Sound (cs.SD)
Cite as:	arXiv:2212.03480 [eess.AS]
	(or arXiv:2212.03480v1 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.2212.03480

Submission history

From: Genshun Wan [view email]
[v1] Wed, 7 Dec 2022 06:29:00 UTC (269 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Progressive Multi-Scale Self-Supervised Learning for Speech Recognition

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Progressive Multi-Scale Self-Supervised Learning for Speech Recognition

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators