arXiv is now an independent nonprofit! Learn more
License: CC BY 4.0
arXiv:2001.04940v1 [eess.AS] 13 Jan 2020

Two Channel Audio Zooming System for Smartphone

A. khandelwal    E.B. Goud    Y. Chand    L. Kumar    S. Prasad Affiliation: Electrical Engineering Department Affiliation: Indian Institute of Technology, Delhi Email: anant.iitd.2085@gmail.com Email: {ee172241/39,lkumar,sprasad}@ee.iitd.ac.in    N. Agarwala    R. Singh Affiliation: SRI Noida Email: ritesh.s7@samsung.com Email: n.agarwala@samsung.com
Abstract

In this paper, two microphone based systems for audio zooming is proposed for the first time. The audio zooming application allows sound capture and enhancement from the front direction while attenuating interfering sources from all other directions. The complete audio zooming system utilizes beamforming based target extraction. In particular, Minimum Power Distortionless Response (MPDR) beamformer and Griffith Jim Beamformer (GJBF) are explored. This is followed by block thresholding for residual noise and interference suppression, and zooming effect creation. A number of simulation and real life experiments using Samsung smartphone (Samsung Galaxy A5) were conducted. Objective and subjective measures confirm the rich user experience.

1 INTRODUCTION

Portable devices for communications like smartphones have become an inseparable part of life. The increasing dependency on the smartphones is due to numerous features they support, ranging from health and convenience to entertainment. One such useful feature being developed is audio zooming [1, 2] where sound from desired direction is enhanced while suppressing interferences from all other directions. This is desirable while trying to listen to a sound in the presence of one or more noise and interfering sources. Practical examples of such environments include that of a railway station, a stadium, classroom and market place. A pictorial depiction of the audio-zooming application for two microphone based smartphone is presented in Figure 1. The evolution of compact device technology and computational power have resulted in use of multiple microphones in a smartphone to exploit the spatial diversity. Many smartphones today utilize two or more microphones. Apple iPhone-511 1 https://www.idownloadblog.com/2012/09/12/iphone-5-three-mics/ makes use of three microphones for beamforming and noise cancellation. Audio zooming has also been reported recently in some smartphones 22 2 https://www.youtube.com/watch?v=zzTUAcZ8FRQ,
https://www.youtube.com/watch?v=YNh4snIzmq4
. However, a significant improvement is required in the presence of severe noise, reverberation and multiple interferences. To the best of our knowledge, the only scientific publication for audio zooming in smartphone is [3] that utilizes MVDR beamforming with a linear array of four microphones.

As most of the current smartphones have two microphones, the possibility of real time audio zooming with smartphone having two microphones is explored in this paper. The complete audio zooming system consists of two blocks as shown in Figure 2. For the beamforming block, the simplest two channel frequency and time domain beamforming is investigated due to limited degree of freedom available. In particular, Minimum Power Distortionless Response (MPDR) beamformer [4] and Griffith Jim Beamformer (GJBF) [5] are explored in frequency and time domain respectively. The two channel beamformer extracts the target source. This spatial filtering causes some suppression of the interference, but a significant residual interference may be present due to the use of only two microphones. A novel block-thresholding [6] based post-filtering is formulated for creating the audio zooming effect. The additional novelty of the work is in exploration of two channel based audio zooming system deployable on smartphone. The filter length in proposed time domain GJBF can be estimated dynamically based on different environmental condition.

Refer to caption
Figure 1: Two microphone smartphone based audio zooming: schematic depiction
Refer to caption
Figure 2: Two stage audio zooming system using modified Griffith-Jim and MPDR beamformer

2 The Proposed Audio Zooming System

We consider a smartphone with two identical and omni-directional microphones located at the top and the bottom. The target source to be acoustically zoomed in, is made incident on the smartphone normal to the plane containing the microphones, as shown in Figure 1. The sources incident from other directions are assumed to be interferences for the audio zooming application. The target and the interference are assumed to be in the far-field region. The proposed audio zooming systems consist of beamforming followed by post-processing. In particular, two audio zooming systems based on frequency and time domain beamforming, have been proposed and their performance have been analyzed . Motivation for using time domain based Griffith Jim Beamformer (GJBF) comes from the fact that it can control the level of interference at the output without the directional information of interferences [5]. In the ensuing Section, two channel based MPDR beamformer is presented, followed by modified time domain GJBF.

2.1 Two Channel MPDR Beamformer

A general wideband array data model can be written in STFT domain as

𝐘(f,ν)=𝐃(Ψ,ν)𝐒(f,ν)+𝐍(f,ν)\mathbf{Y}(f,\nu)=\mathbf{D}(\Psi,\nu)\mathbf{S}(f,\nu)+\mathbf{N}(f,\nu) (1)

where 𝐘(f,ν)\mathbf{Y}(f,\nu) is the received array signal, and 𝐍(f,ν)\mathbf{N}(f,\nu) is zero mean, uncorrelated sensor noise. For MM microphones and LL sources, 𝐃(Ψ,f)\mathbf{D}(\Psi,f) is M×LM\times L array manifold given by

𝐃(Ψ,f)=[𝐝(Ψ1,f),𝐝(Ψ2,f),,𝐝(ΨL,f)]\mathbf{D}(\Psi,f)=[\mathbf{d}(\Psi_{1},f),\mathbf{d}(\Psi_{2},f),\dots,\mathbf{d}(\Psi_{L},f)] (2)

where 𝐝(Ψl,f)\mathbf{d}(\Psi_{l},f) represents the steering vector for lthl^{th} source given by

𝐝(Ψl,f)=[ej2πfτ1lej2πfτ2lej2πfτMl]T\mathbf{d}({\Psi_{l},f})=\begin{bmatrix}e^{-j2\pi f\tau_{1l}}&e^{-j2\pi f\tau_{2l}}&\dots&e^{-j2\pi f\tau_{Ml}}\end{bmatrix}^{T} (3)

Ψl=(θl,ϕl)\Psi_{l}=(\theta_{l},\phi_{l}) is the incident direction of the lthl^{th} source. Here τml\tau_{ml} represents the time delay of arrival of the lthl^{th} signal at the mthm^{th} microphone with respect to a reference microphone. MPDR beamforming problem is equivalent to minimizing the output power with a distortionless response in the target direction given as

min𝐖𝐖H𝐑𝐘𝐖subjectto𝐖H𝐝(Ψl,f)=1\underset{\mathbf{W}}{min}\hskip 2.84544pt\mathbf{W}^{H}\mathbf{R_{Y}}\mathbf{W}\hskip 8.5359ptsubject\hskip 2.84544ptto\hskip 8.5359pt\mathbf{W}^{H}\mathbf{d}({\Psi_{l},f})=1 (4)

The solution to the above optimization problem under diagonal loading is given by

𝐖^(Ψl,f)=(𝐑^Y,f+α𝐈)1𝐝(Ψd,f)𝐝H(Ψd,f)(𝐑^Y,f+α𝐈)1𝐝(Ψd,f)\mathbf{\hat{W}}_{(\Psi_{l},f)}=\frac{(\mathbf{\hat{R}}_{{Y},f}+\alpha\mathbf{I})^{-1}\mathbf{d}(\Psi_{d},f)}{\mathbf{d}^{H}(\Psi_{d},f)(\mathbf{\hat{R}}_{{Y},f}+\alpha\mathbf{I})^{-1}\mathbf{d}(\Psi_{d},f)} (5)

where α\alpha is the diagonal loading factor. The co-variance matrix is estimated as

RY,f=1Kk=0k=KY(f,k)Y(f,k)R_{Y,f}=\frac{1}{K}\sum_{k=0}^{k=K}Y(f,k)*Y^{*}(f,k) (6)

where K is total number of time frames.

2.2 Modified Griffiths-Jim Adaptive Beamformer

Audio zooming application is additionally, explored using two channel time domain beamformer. The target is assumed to be incident from broadside, resulting in identical delays at the two microphones. The nthn^{th} snapshot of the received signal at the mthm^{th} microphone is written as

ym(n)=s(n)+nm(n), n={0,1,,Ns1},m={1,2}y_{m}(n)=s(n)+n_{m}(n),\textrm{ }n=\{0,1,\cdots,N_{s}-1\},m=\{1,2\} (7)

where s(n)s(n) is the target signal to be zoomed in, and nm(n)n_{m}(n) is the total noise and interferences.

As the target signal undergoes identical delays at the two microphones, no phase adjustment is required herein when compared to the original GJBF [5]. The constrained weight wc=12[1,1]Tw_{c}=\frac{1}{2}[1,1]^{T} for the target will result in simple addition of the two channel signal providing signal plus interference yfy_{f} in the upper branch of GJBF in Figure 2. The blocking matrix B=[1,1]TB=[1,-1]^{T} subtracts the two channel data thus it does not allow the signal from the constraint direction in the lower branch of GJBF. The adaptive weight waw_{a} in the lower branch is chosen to estimate the signal at the output of wcw_{c} as a linear combination of the data at the output of the blocking matrix BB. As blocking matrix does not allow the signal from the constraint direction, the signal estimated by waw_{a} is the interference close to interference present at the output of wcw_{c}.

(a) Anechoic recording.
(b) A Meeting room recording.
Figure 3: Dynamic LMS Filter Length Estimation (a) The SINR is max at filter length of 250 (b) The SINR is maximum for filter length 245.

The overall output of the two microphone GJBF is

z(n)=yf(n)yb(n),z(n)=y_{f}(n)-y_{b}(n), (8)

where yfy_{f} has the target signal, noise and interference with response determined by wcHw_{c}^{H}, and yby_{b} has only the noise and interference. The filter weights waw_{a} is found based on minimizing the power contained in z(n)z(n). The power will be minimum when yb(n)y_{b}(n) closely models the interference and noise present in the yf(n)y_{f}(n). In particular, Frequency Domain Adaptive Filter (FDAF) approach [7] using overlap-save method has been utilized for computing waw_{a}. The FDAF approach has logarithmic complexity when compared to polynomial complexity for the classical LMS as utilized in the original GJBF.

2.2.1 LMS Filter Length

The GJ beamformed output quality depends on the length of LMS filter. The filter length depends on the environmental conditions. The length of the filter can be dynamically decided based on the highest SINR in GJBF output. The SINR is computed in Short Time Fourier Transform (STFT) domain as |Z(f,ν)|2σ2(f,ν)σ2(f,ν)\frac{|Z(f,\nu)|^{2}-\sigma^{2}(f,\nu)}{\sigma^{2}(f,\nu)}, where the estimation of residual noise variance is detailed in Section 2.3. SINR at GJBF output is plotted in Figure 3 with the LMS filter length for anechoic chamber and reverberant room recording. It is to note that the optimum filter length for anechoic set-up is 250250 while that of reverberant room is 245245.

2.3 Adaptive Block Thresholding based Post-Filtering for Audio Zooming

In this Section, adaptive block thresholding based post-filtering is proposed for Audio Zooming effect creation. The estimate obtained by adaptive beamforming recovers the target with some amount of noise and interference. Additionally, the beamformed output might show transient or tonal behavior, or combination of the two. The residual interference along with transient and tonal behavior is now simply treated as single additive interference term for further processing. Mathematically, the beamformed output can now be written in STFT domain as

Z(f,ν)=S(f,ν)+I(f,ν)Z(f,\nu)=S(f,\nu)+I(f,\nu) (9)

where S(f,ν)S(f,\nu) is STFT coefficient of desired signal and I(f,ν)I(f,\nu) is STFT coefficient of residual noise and interference. The total variance of such additive interference can be computed as

σ^2(f,ν)=12{(Y1(f,ν)Z(f,ν))2+(Y2(f,ν)Z(f,ν))2}\hat{\sigma}^{2}(f,\nu)=\frac{1}{2}\{(Y_{1}(f,\nu)-Z(f,\nu))^{2}+(Y_{2}(f,\nu)-Z(f,\nu))^{2}\} (10)

where Y1(f,ν)Y_{1}(f,\nu), Y2(f,ν)Y_{2}(f,\nu) are the STFT coefficients of microphone 1 and 2 respectively. This is the best residual interference variance that can be estimated considering the beamformed output is close to the target signal. This estimate works well with practical scenarios. More accurate estimation can provide better result. It is to be noted that because of limited degrees of freedom (two, utilized both) in the spatial domain, post-filtering is now attempted in frequency domain. For this purpose, the output of the either beamformer is considered in short-time Fourier transform (STFT) domain, with suitably chosen block sizes in the time and frequency, as detailed below.

The time-frequency coefficients are modified by multiplying each of them by an attenuation factor to attenuate the interference dominated components as

S^(f,ν)=a(f,ν)Z(f,ν)\hat{S}(f,\nu)=a(f,\nu)*Z(f,\nu) (11)

and creating the zooming effect.

Refer to caption
Figure 4: Example of dividing a macro-block into sub-blocks

The attenuation factor a(f,ν)a(f,\nu) depends upon the values Z(f,ν)Z(f^{\prime},\nu^{\prime}) for all [f,ν][f^{\prime},\nu^{\prime}] in the neighborhood of [f,ν][f,\nu]. The signal estimate S(f,ν)^\hat{S(f,\nu)} is computed from the noisy data Z(f,ν)Z(f,\nu) with a constant attenuation factor aia_{i} over the sub-block BiB_{i} as

S^(f,ν)=aiZ(f,ν) (f,ν)Bi\hat{S}(f,\nu)=a_{i}\hskip 2.84544ptZ(f,\nu)\textrm{ }\forall(f,\nu)\in B_{i} (12)

Selection of the block is done as follows. The entire STFT matrix 𝐙\mathbf{Z} is divided into macro-blocks of size P×Q{P\times Q}. Each macro-block is further divided into sub-blocks of size 2Hv×2v2^{H-v}\times 2^{v} where 2H=L2^{H}=L and 2V=W2^{V}=W and v{0,1..H}v\in\{0,1.....H\}. Various ways of dividing a macro-block into sub-blocks is shown in Figure 4. A Threshold block SNR is chosen to differentiate higher and lower block SNR values. Out of the various structure of sub-blocks, the one having highest SNR is chosen.

The mean square error can be written as

r=E(|S^S|2)1Ai=1If,νBiE{(aiZ(f,ν)S(f,ν))2}r=E(|\hat{S}-S|^{2})\leq\frac{1}{A}\sum^{I}_{i=1}\sum_{f,\nu\in B_{i}}E\{(a_{i}\hskip 2.84544ptZ(f,\nu)-S(f,\nu))^{2}\} (13)

The error can be minimized by choosing [8]

ai=(11ζi+1)+a_{i}=(1-\frac{1}{\zeta_{i}+1})_{+} (14)

where ζi=S¯2σ¯2\zeta_{i}=\frac{\bar{S}^{2}}{\bar{\sigma}^{2}} is the average apriori SNR in the sub-block BiB_{i}, that is computed from

S¯2=1Bi0f,νBiS2(f,ν)and ,σ¯2=1Bi0f,νBiσ2(f,ν)\bar{S}^{2}=\frac{1}{B^{0}_{i}}\sum_{f,\nu\in B_{i}}S^{2}(f,\nu)\textrm{and },\hskip 14.22636pt\bar{\sigma}^{2}=\frac{1}{B_{i}^{0}}\sum_{f,\nu\in B_{i}}\sigma^{2}(f,\nu) (15)

Here Bi0B_{i}^{0} is number of coefficients in the sub-block. As S(f,ν)S(f,\nu) is unknown, the apriori SNR ζi{\zeta}_{i} can be computed alternatively using

ζ^i=Zi¯2σi¯21\hat{\zeta}_{i}=\frac{\bar{Z_{i}}^{2}}{\bar{\sigma_{i}}^{2}}-1 (16)

that can be derived from (9). Zi¯2\bar{Z_{i}}^{2} can be computed as in (15) using the beamformed signal.

It is to be noted that the selection of aia_{i} ensures that the target signal is enhanced corresponding to the high SNR sub-blocks while suppressing the low-SNRs sub-blocks. This results in the required audio-zooming effect.

Refer to caption
Figure 5: Anechoic Chamber Recording Setup

3 Performance evaluation

The performance evaluation of the proposed audio zooming systems is presented herein using simulation and real data experiments. Four objective measures that include Mean Square Error(MSE) [9], Output SINR (OSINR) [9], Perceptual Evaluation of Speech Quality (PESQ) [10] and Short-time Objective Intelligibility (STOI) [11] are utilized for the simulation experiments. Experiments were additionally conducted on real data recorded in anechoic chamber and in the field. The Mean Opinion Score (MOS) [12] measure is utilized to evaluate the performance of experiments with real recorded data. Non-refference PESQ score is also given to evaluate the speech intelligibility. The proposed time and frequency domain audio zooming systems are compared with RMVDR based system [3] for two microphones.

3.1 Simulation Experiments

Two microphones separated by 1010 cm were taken for the simulation experiments. The objective parameters were computed for fifty target speech files from TIMIT database [13]. The target source was taken at azimuth 9090^{\circ}, while the interference was assumed to be at 6060^{\circ}. Signal-to-Interference Ratio (SIR) was taken to be 00dB. The experiments were performed in reverberant condition. For reverberant condition, the reverberation time was taken to be 100100ms. The results are presented in Table 1. It is to be noted that the proposed MPDR beamforming based system outperforms the other methods. The performance of GJBF based system is comparable to RMVDR based system.

Zooming System OSINR(dB) PESQ STOI MSE(dB)
RMVDR 10.56 3.093 0.8935 -25.83
Griffith Jim 9.35 3.0572 0.8655 -23.56
MPDR 13.85 3.203 0.9175 -33.453
Table 1: Objective Evaluation of proposed audio zooming systems

3.2 Experiments on Real Data

The data recording was done was in anechoic chamber and open field. Subjective listening tests were conducted. Mean Opinion Score (MOS) measure is presented to evaluate the proposed audio zooming system. Fifteen subjects in the age group of 202520-25 years were invited for listening the zoomed audio. The subjects listened the mixed received signal and the output. The rating was given for the quality of the output audio on the scale of 00 to 55 [14].

3.2.1 Microphone Array Recordings in Anechoic Chamber

A uniform linear microphone array was utilized for recording in anechoic chamber with inter-element spacing as 4.54.5cm. The array consists of Sennheiser HSP 2 microphones. The target speaker was placed at 9090^{\circ} azimuth. The interference was kept at 4040^{\circ} azimuth.Various kind of interference were selected as shown in table 2. Data of two channel separated by 99cm was utilized for evaluating the audio zooming systems. The MOS measures being close to or more than 44 shows a good perception of the output.

Zooming Technique Interference MOS
Interference
suppression
Speech
quality
RMVDR Train 3.25 4
Vacuum 3.5 4
Speech 3.25 3.75
Sonic 3.25 4
Griffith-Jim Train 4.16 3.75
Vacuum 4.25 4
Speech 4 3.75
Sonic 4.16 3.85
MPDR Train 4 4.25
Vacuum 4.25 4.16
Speech 4.16 4
Sonic 4.16 4.25
Table 2: Performance evaluation for anechoic chamber recording

3.2.2 Smartphone Recordings in Open Ground

As the application target is smartphone, we used two channel recording from Samsung Galaxy A5(2017) smartphone. The recording was performed in a open ground scenarios. The Two orators were located at 9090^{\circ} and 4545^{\circ} to the microphone array. MOS and non-reference PESQ (NR-PESQ) [15] measures are given in Table 3. High MOS and NR-PESQ measures shows the practical applicability of the proposed systems. The more field recording results are made available online at http://web.iitd.ac.in/\simlalank/msp/demo.html for reviewers to evaluate.

Zooming Technique MOS NR-PESQ
Interference
suppression
Speech
quality
RMVDR 3.5 4 2.835
Griffith-Jim 4.08 3.66 2.756
MPDR 4.25 4.166 3.021
Table 3: Evaluation for open ground recording

4 CONCLUSIONS

In this paper, two channel audio zooming systems are proposed for smartphone for the first time. Two channel time and frequency domain beamforming based target extraction is explored. A novel block-thresholding based post-filtering is utilized for creating the audio zooming effect. The proposed MPDR and GJ beamformer based systems are compared with two channel RMVDR based system. The MPDR based audio zooming system outperforms while performance of GJBF based system is comparable with RMVDR. The proposed systems are tested on smartphones for the expected result. MPDR based system can be utilized for smartphones with more than two microphones with better output. A mobile app is being developed for the same. Subjective and objective measures suggest rich user experience.

References

  • [1] Keansub Lee, Hosung Song, Yonghee Lee, SON Youngjoo, and KIM Joontae, “Mobile terminal and audio zooming method thereof,” Jan. 26 2016, US Patent 9,247,192.
  • [2] Carlos Avendano and Ludger Solbach, “Audio zoom,” Dec. 8 2015, US Patent 9,210,503.
  • [3] Ngoc QK Duong, Pierre Berthet, Sidkieta Zabre, Michel Kerdranvat, Alexey Ozerov, and Louis Chevallier, “Audio zoom for smartphones based on multiple adaptive beamformers,” in International Conference on Latent Variable Analysis and Signal Separation. Springer, 2017, pp. 121–130.
  • [4] J. Capon, “High-resolution frequency-wavenumber spectrum analysis,” Proceedings of the IEEE, vol. 57, no. 8, pp. 1408–1418, 1969.
  • [5] Lloyd Griffiths and CW Jim, “An alternative approach to linearly constrained adaptive beamforming,” IEEE Transactions on antennas and propagation, vol. 30, no. 1, pp. 27–34, 1982.
  • [6] Guoshen Yu, Stéphane Mallat, and Emmanuel Bacry, “Audio denoising by time-frequency block thresholding,” IEEE Transactions on Signal processing, vol. 56, no. 5, pp. 1830–1839, 2008.
  • [7] David Mansour and A Gray, “Unconstrained frequency-domain adaptive filter,” IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 30, no. 5, pp. 726–734, 1982.
  • [8] D. Donoho and I. Johnstone, “Idea spatial adaptation via wavelet shrinkage,” Biometrika, vol. 81, pp. 425–455.
  • [9] Valentin Emiya, Emmanuel Vincent, Niklas Harlander, and Volker Hohmann, “Subjective and objective quality assessment of audio source separation,” IEEE Transactions on Audio, Speech, and Language Processing, vol. 19, no. 7, pp. 2046–2057, 2011.
  • [10] Antony W Rix, John G Beerends, Michael P Hollier, and Andries P Hekstra, “Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,” in Acoustics, Speech, and Signal Processing, 2001. Proceedings.(ICASSP’01). 2001 IEEE International Conference on. IEEE, 2001, vol. 2, pp. 749–752.
  • [11] Cees H Taal, Richard C Hendriks, Richard Heusdens, and Jesper Jensen, “A short-time objective intelligibility measure for time-frequency weighted noisy speech,” in Acoustics Speech and Signal Processing (ICASSP), 2010 IEEE International Conference on. IEEE, 2010, pp. 4214–4217.
  • [12] Mahesh Viswanathan and Madhubalan Viswanathan, “Measuring speech quality for text-to-speech systems: development and assessment of a modified mean opinion score (mos) scale,” Computer Speech & Language, vol. 19, no. 1, pp. 55–83, 2005.
  • [13] Victor Zue, Stephanie Seneff, and James Glass, “Speech database development at mit: Timit and beyond,” Speech Communication, vol. 9, no. 4, pp. 351–356, 1990.
  • [14] ITUT Rec, “P. 85 (1994) a method for subjective performance assessment of the quality of speech voice output devices,” International Telecommunication Union, Geneva Google Scholar.
  • [15] International Telecommunication Union, “Non refference pesq score calculation,” https://www.itu.int/rec/T-REC-P.562-200511-I!Amd2/.