Two Channel Audio Zooming System for Smartphone
Abstract
In this paper, two microphone based systems for audio zooming is proposed for the first time. The audio zooming application allows sound capture and enhancement from the front direction while attenuating interfering sources from all other directions. The complete audio zooming system utilizes beamforming based target extraction. In particular, Minimum Power Distortionless Response (MPDR) beamformer and Griffith Jim Beamformer (GJBF) are explored. This is followed by block thresholding for residual noise and interference suppression, and zooming effect creation. A number of simulation and real life experiments using Samsung smartphone (Samsung Galaxy A5) were conducted. Objective and subjective measures confirm the rich user experience.
1 INTRODUCTION
Portable devices for communications like smartphones have become an inseparable part of life. The increasing dependency on the smartphones is due to numerous features they support, ranging from health and convenience to entertainment. One such useful feature being developed is audio zooming [1, 2] where sound from desired direction is enhanced while suppressing interferences from all other directions. This is desirable while trying to listen to a sound in the presence of one or more noise and interfering sources. Practical examples of such environments include that of a railway station, a stadium, classroom and market place. A pictorial depiction of the audio-zooming application for two microphone based smartphone is presented in Figure 1.
The evolution of compact device technology and computational power have resulted in use of multiple microphones in a smartphone to exploit the spatial diversity. Many smartphones today utilize two or more microphones. Apple iPhone-511
1
https://www.idownloadblog.com/2012/09/12/iphone-5-three-mics/ makes use of three microphones for beamforming and noise cancellation. Audio zooming has also been reported recently in some smartphones 22
2
https://www.youtube.com/watch?v=zzTUAcZ8FRQ,
https://www.youtube.com/watch?v=YNh4snIzmq4. However, a significant improvement is required in the presence of severe noise, reverberation and multiple interferences. To the best of our knowledge, the only scientific publication for audio zooming in smartphone is [3] that utilizes MVDR beamforming with a linear array of four microphones.
As most of the current smartphones have two microphones, the possibility of real time audio zooming with smartphone having two microphones is explored in this paper. The complete audio zooming system consists of two blocks as shown in Figure 2. For the beamforming block, the simplest two channel frequency and time domain beamforming is investigated due to limited degree of freedom available. In particular, Minimum Power Distortionless Response (MPDR) beamformer [4] and Griffith Jim Beamformer (GJBF) [5] are explored in frequency and time domain respectively. The two channel beamformer extracts the target source. This spatial filtering causes some suppression of the interference, but a significant residual interference may be present due to the use of only two microphones. A novel block-thresholding [6] based post-filtering is formulated for creating the audio zooming effect. The additional novelty of the work is in exploration of two channel based audio zooming system deployable on smartphone. The filter length in proposed time domain GJBF can be estimated dynamically based on different environmental condition.
2 The Proposed Audio Zooming System
We consider a smartphone with two identical and omni-directional microphones located at the top and the bottom. The target source to be acoustically zoomed in, is made incident on the smartphone normal to the plane containing the microphones, as shown in Figure 1. The sources incident from other directions are assumed to be interferences for the audio zooming application. The target and the interference are assumed to be in the far-field region. The proposed audio zooming systems consist of beamforming followed by post-processing. In particular, two audio zooming systems based on frequency and time domain beamforming, have been proposed and their performance have been analyzed . Motivation for using time domain based Griffith Jim Beamformer (GJBF) comes from the fact that it can control the level of interference at the output without the directional information of interferences [5]. In the ensuing Section, two channel based MPDR beamformer is presented, followed by modified time domain GJBF.
2.1 Two Channel MPDR Beamformer
A general wideband array data model can be written in STFT domain as
| (1) |
where is the received array signal, and is zero mean, uncorrelated sensor noise. For microphones and sources, is array manifold given by
| (2) |
where represents the steering vector for source given by
| (3) |
is the incident direction of the source. Here represents the time delay of arrival of the signal at the microphone with respect to a reference microphone. MPDR beamforming problem is equivalent to minimizing the output power with a distortionless response in the target direction given as
| (4) |
The solution to the above optimization problem under diagonal loading is given by
| (5) |
where is the diagonal loading factor. The co-variance matrix is estimated as
| (6) |
where K is total number of time frames.
2.2 Modified Griffiths-Jim Adaptive Beamformer
Audio zooming application is additionally, explored using two channel time domain beamformer. The target is assumed to be incident from broadside, resulting in identical delays at the two microphones. The snapshot of the received signal at the microphone is written as
| (7) |
where is the target signal to be zoomed in, and is the total noise and interferences.
As the target signal undergoes identical delays at the two microphones, no phase adjustment is required herein when compared to the original GJBF [5]. The constrained weight for the target will result in simple addition of the two channel signal providing signal plus interference in the upper branch of GJBF in Figure 2. The blocking matrix subtracts the two channel data thus it does not allow the signal from the constraint direction in the lower branch of GJBF. The adaptive weight in the lower branch is chosen to estimate the signal at the output of as a linear combination of the data at the output of the blocking matrix . As blocking matrix does not allow the signal from the constraint direction, the signal estimated by is the interference close to interference present at the output of .
The overall output of the two microphone GJBF is
| (8) |
where has the target signal, noise and interference with response determined by , and has only the noise and interference. The filter weights is found based on minimizing the power contained in . The power will be minimum when closely models the interference and noise present in the . In particular, Frequency Domain Adaptive Filter (FDAF) approach [7] using overlap-save method has been utilized for computing . The FDAF approach has logarithmic complexity when compared to polynomial complexity for the classical LMS as utilized in the original GJBF.
2.2.1 LMS Filter Length
The GJ beamformed output quality depends on the length of LMS filter. The filter length depends on the environmental conditions. The length of the filter can be dynamically decided based on the highest SINR in GJBF output. The SINR is computed in Short Time Fourier Transform (STFT) domain as , where the estimation of residual noise variance is detailed in Section 2.3. SINR at GJBF output is plotted in Figure 3 with the LMS filter length for anechoic chamber and reverberant room recording. It is to note that the optimum filter length for anechoic set-up is while that of reverberant room is .
2.3 Adaptive Block Thresholding based Post-Filtering for Audio Zooming
In this Section, adaptive block thresholding based post-filtering is proposed for Audio Zooming effect creation. The estimate obtained by adaptive beamforming recovers the target with some amount of noise and interference. Additionally, the beamformed output might show transient or tonal behavior, or combination of the two. The residual interference along with transient and tonal behavior is now simply treated as single additive interference term for further processing. Mathematically, the beamformed output can now be written in STFT domain as
| (9) |
where is STFT coefficient of desired signal and is STFT coefficient of residual noise and interference. The total variance of such additive interference can be computed as
| (10) |
where , are the STFT coefficients of microphone 1 and 2 respectively. This is the best residual interference variance that can be estimated considering the beamformed output is close to the target signal. This estimate works well with practical scenarios. More accurate estimation can provide better result. It is to be noted that because of limited degrees of freedom (two, utilized both) in the spatial domain, post-filtering is now attempted in frequency domain. For this purpose, the output of the either beamformer is considered in short-time Fourier transform (STFT) domain, with suitably chosen block sizes in the time and frequency, as detailed below.
The time-frequency coefficients are modified by multiplying each of them by an attenuation factor to attenuate the interference dominated components as
| (11) |
and creating the zooming effect.
The attenuation factor depends upon the values for all in the neighborhood of . The signal estimate is computed from the noisy data with a constant attenuation factor over the sub-block as
| (12) |
Selection of the block is done as follows. The entire STFT matrix is divided into macro-blocks of size . Each macro-block is further divided into sub-blocks of size where and and . Various ways of dividing a macro-block into sub-blocks is shown in Figure 4. A Threshold block SNR is chosen to differentiate higher and lower block SNR values. Out of the various structure of sub-blocks, the one having highest SNR is chosen.
The mean square error can be written as
| (13) |
The error can be minimized by choosing [8]
| (14) |
where is the average apriori SNR in the sub-block , that is computed from
| (15) |
Here is number of coefficients in the sub-block. As is unknown, the apriori SNR can be computed alternatively using
| (16) |
that can be derived from (9). can be computed as in (15) using the beamformed signal.
It is to be noted that the selection of ensures that the target signal is enhanced corresponding to the high SNR sub-blocks while suppressing the low-SNRs sub-blocks. This results in the required audio-zooming effect.
3 Performance evaluation
The performance evaluation of the proposed audio zooming systems is presented herein using simulation and real data experiments. Four objective measures that include Mean Square Error(MSE) [9], Output SINR (OSINR) [9], Perceptual Evaluation of Speech Quality (PESQ) [10] and Short-time Objective Intelligibility (STOI) [11] are utilized for the simulation experiments. Experiments were additionally conducted on real data recorded in anechoic chamber and in the field. The Mean Opinion Score (MOS) [12] measure is utilized to evaluate the performance of experiments with real recorded data. Non-refference PESQ score is also given to evaluate the speech intelligibility. The proposed time and frequency domain audio zooming systems are compared with RMVDR based system [3] for two microphones.
3.1 Simulation Experiments
Two microphones separated by cm were taken for the simulation experiments. The objective parameters were computed for fifty target speech files from TIMIT database [13]. The target source was taken at azimuth , while the interference was assumed to be at . Signal-to-Interference Ratio (SIR) was taken to be dB. The experiments were performed in reverberant condition. For reverberant condition, the reverberation time was taken to be ms. The results are presented in Table 1. It is to be noted that the proposed MPDR beamforming based system outperforms the other methods. The performance of GJBF based system is comparable to RMVDR based system.
| Zooming System | OSINR(dB) | PESQ | STOI | MSE(dB) |
|---|---|---|---|---|
| RMVDR | 10.56 | 3.093 | 0.8935 | -25.83 |
| Griffith Jim | 9.35 | 3.0572 | 0.8655 | -23.56 |
| MPDR | 13.85 | 3.203 | 0.9175 | -33.453 |
3.2 Experiments on Real Data
The data recording was done was in anechoic chamber and open field. Subjective listening tests were conducted. Mean Opinion Score (MOS) measure is presented to evaluate the proposed audio zooming system. Fifteen subjects in the age group of years were invited for listening the zoomed audio. The subjects listened the mixed received signal and the output. The rating was given for the quality of the output audio on the scale of to [14].
3.2.1 Microphone Array Recordings in Anechoic Chamber
A uniform linear microphone array was utilized for recording in anechoic chamber with inter-element spacing as cm. The array consists of Sennheiser HSP 2 microphones. The target speaker was placed at azimuth. The interference was kept at azimuth.Various kind of interference were selected as shown in table 2. Data of two channel separated by cm was utilized for evaluating the audio zooming systems. The MOS measures being close to or more than shows a good perception of the output.
| Zooming Technique | Interference | MOS | |||
|
| ||||
| RMVDR | Train | 3.25 | 4 | ||
| Vacuum | 3.5 | 4 | |||
| Speech | 3.25 | 3.75 | |||
| Sonic | 3.25 | 4 | |||
| Griffith-Jim | Train | 4.16 | 3.75 | ||
| Vacuum | 4.25 | 4 | |||
| Speech | 4 | 3.75 | |||
| Sonic | 4.16 | 3.85 | |||
| MPDR | Train | 4 | 4.25 | ||
| Vacuum | 4.25 | 4.16 | |||
| Speech | 4.16 | 4 | |||
| Sonic | 4.16 | 4.25 | |||
3.2.2 Smartphone Recordings in Open Ground
As the application target is smartphone, we used two channel recording from Samsung Galaxy A5(2017) smartphone. The recording was performed in a open ground scenarios. The Two orators were located at and to the microphone array. MOS and non-reference PESQ (NR-PESQ) [15] measures are given in Table 3. High MOS and NR-PESQ measures shows the practical applicability of the proposed systems. The more field recording results are made available online at http://web.iitd.ac.in/lalank/msp/demo.html for reviewers to evaluate.
| Zooming Technique | MOS | NR-PESQ | |||
|---|---|---|---|---|---|
|
| ||||
| RMVDR | 3.5 | 4 | 2.835 | ||
| Griffith-Jim | 4.08 | 3.66 | 2.756 | ||
| MPDR | 4.25 | 4.166 | 3.021 | ||
4 CONCLUSIONS
In this paper, two channel audio zooming systems are proposed for smartphone for the first time. Two channel time and frequency domain beamforming based target extraction is explored. A novel block-thresholding based post-filtering is utilized for creating the audio zooming effect. The proposed MPDR and GJ beamformer based systems are compared with two channel RMVDR based system. The MPDR based audio zooming system outperforms while performance of GJBF based system is comparable with RMVDR. The proposed systems are tested on smartphones for the expected result. MPDR based system can be utilized for smartphones with more than two microphones with better output. A mobile app is being developed for the same. Subjective and objective measures suggest rich user experience.
References
- [1] Keansub Lee, Hosung Song, Yonghee Lee, SON Youngjoo, and KIM Joontae, “Mobile terminal and audio zooming method thereof,” Jan. 26 2016, US Patent 9,247,192.
- [2] Carlos Avendano and Ludger Solbach, “Audio zoom,” Dec. 8 2015, US Patent 9,210,503.
- [3] Ngoc QK Duong, Pierre Berthet, Sidkieta Zabre, Michel Kerdranvat, Alexey Ozerov, and Louis Chevallier, “Audio zoom for smartphones based on multiple adaptive beamformers,” in International Conference on Latent Variable Analysis and Signal Separation. Springer, 2017, pp. 121–130.
- [4] J. Capon, “High-resolution frequency-wavenumber spectrum analysis,” Proceedings of the IEEE, vol. 57, no. 8, pp. 1408–1418, 1969.
- [5] Lloyd Griffiths and CW Jim, “An alternative approach to linearly constrained adaptive beamforming,” IEEE Transactions on antennas and propagation, vol. 30, no. 1, pp. 27–34, 1982.
- [6] Guoshen Yu, Stéphane Mallat, and Emmanuel Bacry, “Audio denoising by time-frequency block thresholding,” IEEE Transactions on Signal processing, vol. 56, no. 5, pp. 1830–1839, 2008.
- [7] David Mansour and A Gray, “Unconstrained frequency-domain adaptive filter,” IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 30, no. 5, pp. 726–734, 1982.
- [8] D. Donoho and I. Johnstone, “Idea spatial adaptation via wavelet shrinkage,” Biometrika, vol. 81, pp. 425–455.
- [9] Valentin Emiya, Emmanuel Vincent, Niklas Harlander, and Volker Hohmann, “Subjective and objective quality assessment of audio source separation,” IEEE Transactions on Audio, Speech, and Language Processing, vol. 19, no. 7, pp. 2046–2057, 2011.
- [10] Antony W Rix, John G Beerends, Michael P Hollier, and Andries P Hekstra, “Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,” in Acoustics, Speech, and Signal Processing, 2001. Proceedings.(ICASSP’01). 2001 IEEE International Conference on. IEEE, 2001, vol. 2, pp. 749–752.
- [11] Cees H Taal, Richard C Hendriks, Richard Heusdens, and Jesper Jensen, “A short-time objective intelligibility measure for time-frequency weighted noisy speech,” in Acoustics Speech and Signal Processing (ICASSP), 2010 IEEE International Conference on. IEEE, 2010, pp. 4214–4217.
- [12] Mahesh Viswanathan and Madhubalan Viswanathan, “Measuring speech quality for text-to-speech systems: development and assessment of a modified mean opinion score (mos) scale,” Computer Speech & Language, vol. 19, no. 1, pp. 55–83, 2005.
- [13] Victor Zue, Stephanie Seneff, and James Glass, “Speech database development at mit: Timit and beyond,” Speech Communication, vol. 9, no. 4, pp. 351–356, 1990.
- [14] ITUT Rec, “P. 85 (1994) a method for subjective performance assessment of the quality of speech voice output devices,” International Telecommunication Union, Geneva Google Scholar.
- [15] International Telecommunication Union, “Non refference pesq score calculation,” https://www.itu.int/rec/T-REC-P.562-200511-I!Amd2/.