Preprint / Version 1

LSTM-CNN: Network for Audio Signature Analysis in Noisy Environments

##article.authors##

  • Praveen Damacharla KINETICAI INC
  • Hamid Rajabalipanah
  • Mohammad Hosein Fakheri

DOI:

https://doi.org/10.31224/3312

Keywords:

artificialintelligence, human-machine interaction, Natural Language Processing, CNN, CNN models

Abstract

There are multiple applications to automatically count people and specify their gender at work, exhibitions, malls, sales, and industrial usage. Although current speech detection methods are supposed to operate well, in most situations, in addition to genders, the number of current speakers is unknown, and the classification methods are unsuitable due to many possible classes. In this study, we focus on a long-short-term memory convolutional neural network (LSTM-CNN) to extract the sound data's time and/or frequency-dependent features to estimate the number/gender of simultaneous active speakers at each frame in noisy environments. Considering the maximum number of speakers as 10, we have utilized 19000 audio samples with diverse combinations of males, females, and background noise in public cities, industrial situations, malls, exhibitions, workplaces, and nature for learning purposes. This proof of concept shows promising performance with training/validation MSE values of about 0.019/0.017 in detecting count and gender.

Downloads

Download data is not yet available.

Downloads

Posted

2023-10-25