Standard

Audio-Visual Multi-modal Meeting Recording System. / Ян, Вэньфэн; Li, Pengyi; Ян, Вэй; Liu, Yuxing ; Petrosian, Ovanes; Ли, Инь.

In: Lecture Notes in Networks and Systems, No. 776, 21.09.2023, p. 168-178.

Research output: Contribution to journalArticlepeer-review

Harvard

APA

Vancouver

Author

Ян, Вэньфэн ; Li, Pengyi ; Ян, Вэй ; Liu, Yuxing ; Petrosian, Ovanes ; Ли, Инь. / Audio-Visual Multi-modal Meeting Recording System. In: Lecture Notes in Networks and Systems. 2023 ; No. 776. pp. 168-178.

BibTeX

@article{670641eee6fa4b168c0ba8a96feda4f1,
title = "Audio-Visual Multi-modal Meeting Recording System",
abstract = "There exist two forms of speaker recognition in meeting recording systems: hardware recognition and software recognition, but the applicability of such two approaches in real meetings is not good enough and the hardware cost is too high. The main contribution of this paper is to use the theory of domain generalization to train the model and use contrast learning to improve the model migration learning ability, while this paper constructs a speaker recognition and meeting content transcription system based on deep learning audiovisual speech recognition (AVSR) model and speaker recognition model (SPR), which only needs a microphone and a camera to recognize the current speaker and use the system{\textquoteright}s audiovisual speech recognition The speaker recognition module is used to transcribe the conference content.",
keywords = "Audio-Visual Speech Recognition, Multi-modal, Speaker recognition",
author = "Вэньфэн Ян and Pengyi Li and Вэй Ян and Yuxing Liu and Ovanes Petrosian and Инь Ли",
year = "2023",
month = sep,
day = "21",
doi = "10.1007/978-3-031-43789-2_15",
language = "English",
pages = "168--178",
journal = "Lecture Notes in Networks and Systems",
issn = "2367-3389",
publisher = "Springer Nature",
number = "776",

}

RIS

TY - JOUR

T1 - Audio-Visual Multi-modal Meeting Recording System

AU - Ян, Вэньфэн

AU - Li, Pengyi

AU - Ян, Вэй

AU - Liu, Yuxing

AU - Petrosian, Ovanes

AU - Ли, Инь

PY - 2023/9/21

Y1 - 2023/9/21

N2 - There exist two forms of speaker recognition in meeting recording systems: hardware recognition and software recognition, but the applicability of such two approaches in real meetings is not good enough and the hardware cost is too high. The main contribution of this paper is to use the theory of domain generalization to train the model and use contrast learning to improve the model migration learning ability, while this paper constructs a speaker recognition and meeting content transcription system based on deep learning audiovisual speech recognition (AVSR) model and speaker recognition model (SPR), which only needs a microphone and a camera to recognize the current speaker and use the system’s audiovisual speech recognition The speaker recognition module is used to transcribe the conference content.

AB - There exist two forms of speaker recognition in meeting recording systems: hardware recognition and software recognition, but the applicability of such two approaches in real meetings is not good enough and the hardware cost is too high. The main contribution of this paper is to use the theory of domain generalization to train the model and use contrast learning to improve the model migration learning ability, while this paper constructs a speaker recognition and meeting content transcription system based on deep learning audiovisual speech recognition (AVSR) model and speaker recognition model (SPR), which only needs a microphone and a camera to recognize the current speaker and use the system’s audiovisual speech recognition The speaker recognition module is used to transcribe the conference content.

KW - Audio-Visual Speech Recognition

KW - Multi-modal

KW - Speaker recognition

UR - https://www.mendeley.com/catalogue/2c223cec-3647-362a-aef5-235fc92a03aa/

U2 - 10.1007/978-3-031-43789-2_15

DO - 10.1007/978-3-031-43789-2_15

M3 - Article

SP - 168

EP - 178

JO - Lecture Notes in Networks and Systems

JF - Lecture Notes in Networks and Systems

SN - 2367-3389

IS - 776

ER -

ID: 114434424