Automatically fetches the latest arXiv papers on Direction of Arrival (DoA) estimation and Array Signal Processing. Strictly filtered for Signal Processing (eess.SP, eess.AS) and Audio (cs.SD) fields.
Last update: 2026-08-24
| Title | Date | Abstract | Comment |
|---|---|---|---|
| Lightweight Single-Antenna Direction-of-Arrival Estimation for Curvilinear Trajectories in Mobile Embedded Systems | 2026-08-12 | ShowAccurate direction-of-arrival (DOA) estimation is valuable for spatially selective communication in noisy industrial environments. This work investigates a lightweight single-antenna framework in which receiver motion forms a virtual aperture. The receiver uses onboard inertial measurement unit (IMU) headings and two-way-ranging (TWR) measurements to a known fixed beacon, avoiding GPS, optical tracking, and high-precision external tracking of the mobile receiver. The curvilinear virtual-array MUSIC formulation is evaluated numerically using phase-coherent narrowband snapshots over arbitrary trajectories. Hardware experiments validate a range- domain TWR-IMU bearing estimator using corrected and averaged ranging observations. In a campaign of 100 consecutive four-revolution sweeps, all trials are retained. Phase-aligned accumulation recovers a 123 mm range-modulation amplitude, in close agreement with the measured 120 mm antenna lever arm. The single-sweep bearing precision is 8.2° absolute world-frame accuracy is limited by systematic BNO055 magnetometer drift in the motorized setup. The embedded bearing-estimation computation consumes 144.9 mJ per estimate, approximately 3 % of the measured cycle energy in the tested configuration. These results establish the feasibility of onboard range-domain bearing estimation and motivate future experimental validation of phase-coherent MUSIC during free-form mobile trajectories. |
|
| Low-Complexity Gridless Single-Snapshot DoA Estimation via Truncated Hankel Newton-MUSIC | 2026-07-09 | ShowReconfigurable antenna arrays can provide enhanced spatial Degrees of Freedom (DoFs) for Integrated Sensing And Communication (ISAC) systems, enabling high-resolution Direction of Arrival (DoA) estimation. In highly dynamic scenarios, however, DoA estimation must be performed within short coherence intervals, which often restricts processing to a single snapshot. Conventional subspace methods then suffer from rank deficiency, while Hankel-based spatial smoothing incurs high computational cost when the array size is large. This paper proposes a low-complexity gridless Truncated Hankel NewtonMUSIC framework for single-snapshot DoA estimation. The proposed method constructs a truncated Hankel matrix with a fixed row dimension to recover an effective signal subspace while reducing the cost of correlation construction and subspace decomposition. When the truncation length is independent of the array size, the dominant complexity scales linearly with the number of antenna ports. To reduce grid-induced quantization errors, coarse grid estimates are further refined by a secondorder Newton update in the continuous angular domain. Simulation results show that the proposed method achieves DoA estimation accuracy close to square Hankel Newton-MUSIC while substantially reducing runtime. For large arrays, it provides more than two orders of magnitude runtime reduction compared with conventional square Hankel MUSIC, making it suitable for realtime sensing in reconfigurable antenna-enabled ISAC systems. |
|
| Enhanced Direction-Sensing Methods and Performance Analysis in Low-Altitude Wireless Network via a Rotating Antenna Array | 2026-05-29 | ShowDue to the directive property of each antenna element, the received signal power can be severely attenuated when the emitter deviates from the array boresight, which will lead to a severe degradation in sensing performance along the corresponding direction. Although existing rotatable array sensing methods such as recursive rotation (RR-Root-MUSIC) can mitigate this issue by iteratively rotating and sensing, several mechanical rotations and repeated eigendecomposition operations are required to yield a high computational complexity and low time-efficiency. To address this problem, a pre-rotation initialization with recieve power as a rule is proposed to signifcantly reduce the computational complexity and improve the time-efficiency. Using this idea, a low-complexity enhanced direction-sensing framework with pre-rotation initialization and iterative greedy spatial-spectrum search (PRI-IGSS) is develped with three stages: (1) the normal vector of array is rotated to a set of candidates to find the opimal direction with the maximum sensing energy with the corresponding DOA value computed by the Root-MUSIC algorithm; (2) the array is mechanically rotated to the initial estimated direction and kept fixed; (3) an iterative greedy spatial-spectrum search or recieving beamforming method, moviated by reinforcement learning, is designed with a reduced search range and making a summation of all previous sampling variance matrices and the current one is adopted to provide an increasiong performance gain as the iteration process continues. To assess the performance of the proposed method, the corresponding CRLB is derived with a simplified rotation model. Simulation results demonstrate that the proposed PRI-IGSS method performs much better than RR-Root-MUSIC and achieves the CRLB in term of mean squared error due to the fact there is no sample accumulation for the latter. |
|
| G-iMUSIC: Greedy Iterative MUSIC Algorithms for Multi-Target DoA Estimation | 2026-05-27 | ShowThis paper presents novel algorithms for multi-target direction-of-arrival (DoA) estimation in array signal processing. Although the maximum likelihood estimator (MLE) asymptotically attains the Cramér-Rao bound, its exponential complexity motivates practical alternatives, such as greedy or subspace-based methods. In this context, greedy methods such as orthogonal matching pursuit (OMP) and orthogonal least squares (OLS) are sensitive to early selection errors, especially for angularly proximate targets, whereas subspace-based methods such as multiple signal classification (MUSIC) present angular super-resolution capabilities but degrade under strong inter-target signal correlation. To overcome these limitations, we propose two greedy iterative MUSIC (G-iMUSIC) algorithms, namely OMP-iMUSIC and OLS-iMUSIC, derived from a unified framework that links subspace and greedy estimations. Unlike prior iMUSIC approaches, the proposed methods require only one initial eigen value decomposition (EVD) and avoid computing eigendecomposition at each iteration. They also admit Fast Fourier Transform (FFT)-accelerated implementations for uniform linear arrays (ULAs), enabling low-complexity operation. Monte Carlo simulations demonstrate improved detection and precision over conventional OMP, OLS, and MUSIC, as well as reduced processing time compared to greedy baselines. Finally, we introduce diagnostic metrics that interpret performance across signal correlation and angular proximity regimes, supporting generalization beyond the specific orthogonal frequency-division multiplexing (OFDM) radar scenario considered. |
12 pa...12 pages; This work has been submitted to the IEEE for possible publication |
| Robust Quantum-MUSIC for DoA Estimation Using Rydberg Atomic Receiver Arrays | 2026-05-25 | ShowQuantum wireless sensing using Rydberg atomic receivers enables high-sensitivity signal acquisition direction-of-arrival (DoA) estimation. However, it suffers from a fundamental limitation, where only the magnitude of the received signal is observable. The recently proposed Quantum-MUSIC algorithm addresses this problem by recovering phase information through alternating minimization and subsequently applying the MUSIC algorithm for DoA estimation. However, the existing approach relies on an |
|
| Sparse Fluid Antenna Arrays: Continuous Position Design Beyond Classical DOF Limits | 2026-05-19 | ShowFluid antenna system (FAS), which continuously repositions a single physical element across a deployment region |
|
| Sensing-Assisted Channel Estimation for Flexible-Antenna Systems: A Unified Framework | 2026-04-30 | ShowFlexible-antenna systems, which use a small number of radio frequency (RF) chains to dynamically access a large set of candidate antenna locations, have emerged as a hardware-efficient architecture for 6G networks. Acquiring accurate channel state information (CSI) is critical for these systems, but it typically incurs a prohibitive pilot overhead that scales with the massive number of candidate locations. To address this bottleneck, we propose a unified sensing-assisted channel estimation framework tailored for flexible-antenna systems. It reduces the full CSI reconstruction problem to a consistent two-stage process: it first resolves the dominant DOAs from the uplink data symbols by exploiting the spatial geometry, requiring no dedicated sensing pilot, and then calibrates the associated path gains using a minimal number of calibration pilots. Building on this pipeline, we develop two Newton-MUSIC algorithms tailored to different propagation environments. For line-of-sight (LOS)-dominant environments with uncorrelated sources, we propose SOC-Newton-MUSIC, which leverages second-order covariance (SOC) for low-complexity DOA sensing. For non-line-of-sight (NLOS) environments with coherent multipath, where the number of sources may exceed the number of activated RF chains, we propose FOC-Newton-MUSIC, which exploits fourth-order cumulants (FOC) to restore source identifiability and structurally expand the available spatial degrees of freedom (DOFs) through a continuous difference co-array. In both cases, by reformulating the spatial spectrum search as a continuous optimization problem, we replace exhaustive dense grid searches with parallelized Newton refinements. |
|
| Hybrid Architecture Gets Fluid: A New Paradigm for Direction-of-arrival Estimation in 6G Networks | 2026-04-15 | ShowHigh-precision direction-of-arrival (DOA) estimation, as a key sensing capability for 6G-enabled applications such as autonomous driving and extended reality, is increasingly dependent on the effective exploitation of spatial degrees of freedom (DOFs). This paper integrates two frontier DOFs-oriented paradigms and proposes a fluid antenna-enabled hybrid analog-digital (FA-HAD) architecture, which features an extremely lightweight front-end configuration mechanism and efficient spatial DOFs exploitation. Within this architecture, a collaborative spatial-phase sampling strategy is first developed to enable real-time 2-D DOA estimation under compressive observations, and a single-source CRLB analysis is provided to quantify the achievable performance limit, offering quantitative guidance for accuracy-overhead trade-offs. Furthermore, an efficient virtual-array spatial covariance matrix reconstruction method is proposed to recover a physically meaningful covariance representation, thereby providing a covariance-domain interface that is directly reusable by a broad class of existing covariance-based array processing and array design techniques, which strengthens the scalability and transferability of the proposed architecture. Building upon the reconstructed SCM, a Jacobi-Anger expansion based dimension-reduced MUSIC estimator is further derived for arbitrary planar arrays with a favorable computational cost. Simulation results demonstrate that the proposed FA-HAD framework attains DOA accuracy close to fully digital systems while substantially reducing RF hardware complexity and training overhead. |
|
| Near-Field NLOS Localization via Position-Unknown HRIS:From Self-Localization to Target Positioning | 2026-03-18 | ShowCurrent reconfigurable intelligent surface (RIS)-aided near-field (NF) localization methods assume the RIS position is known a priori, and it has limited their practical applicability. This paper applies a hybrid RIS (HRIS) at an unknown position to locate non-line-of-sight (NLOS) NF targets. To this end, we first propose a two-stage gridless localization framework for achieving HRIS self-localization, and then determine the positions of the NF targets. In the first stage, we use the NF Fresnel approximation to convert the signal model into a virtual far-field model through delay-based cross-correlation of centrally symmetric HRIS elements. Such a conversion will naturally extend the aperture of the virtual array. A single-snapshot decoupled atomic norm minimization (DANM) algorithm is then proposed to locate an NF target relative to the HRIS, which includes a two-dimensional (2-D) direction of arrival (DOA) estimation with automatic pairing, the multiple signal classification (MUSIC) method for range estimation, and a total least squares (TLS) method to eliminate the Fresnel approximation error. In the second stage, we leverage the unique capability of HRIS in simultaneous sensing and reflection to estimate the HRIS-to-base station (BS) direction vectors using atomic norm minimization (ANM), and derive the three-dimensional (3-D) HRIS position with two BSs via the least squares (LS)-based geometric triangulation. Furthermore, we propose a semidefinite relaxation (SDR)-based HRIS phase optimization method to enhance the received signal power at the BSs, thereby improving the HRIS localization accuracy, which, in turn, enhances NF target positionings. The Cramer-Rao bound (CRB) for the NF target parameters and the position error bound (PEB) for the HRIS coordinates are derived as performance benchmarks. |
14 pages, 14 figures |
| Joint single-shot ToA and DoA estimation for VAA-based BLE ranging with phase ambiguity: A deep learning-based approach | 2026-01-21 | ShowConventional direction-of-arrival (DoA) estimation methods rely on multi-antenna arrays, which are costly to implement on size-constrained Bluetooth Low Energy (BLE) devices. Virtual antenna array (VAA) techniques enable DoA estimation with a single antenna, making angle estimation feasible on such devices. However, BLE only provides a single-shot two-way channel frequency response (CFR) with a binary phase ambiguity issue, which hinders the direct application of VAA. To address this challenge, we propose a unified model that combines VAA with BLE two-way CFR, and introduce a neural network based phase recovery framework that employs row / column predictors with a voting mechanism to resolve the ambiguity. The recovered one-way CFR then enables super resolution algorithms such as MUSIC for joint time of arrival (ToA) and DoA estimation. Simulation results demonstrate that the proposed method achieves superior performance under non-uniform VAAs, with mean square errors approaching the Cramer Rao bound at SNR |
|
| An Fluid Antenna Array-Enabled DOA Estimation Method: End-Fire Effect Suppression | 2025-12-22 | ShowDirection of Arrival (DOA) estimation serves as a critical sensing technology poised to play a vital role in future intelligent and ubiquitous communication systems. Despite the development of numerous mature super-resolution algorithms, the inherent end-fire effect problem in fixed antenna arrays remains inadequately addressed. This work proposed a novel array architecture composed of fluid antennas. By exploiting the spatial reconfigurability of their positions to equivalently modulate the array steering vector and integrating it with the classical MUSIC algorithm, this approach achieved high-precision DOA estimation. Simulation results demonstrated that the proposed method delivers outstanding estimation performance even in highly challenging end-fire regions. |
|
| Adaptive MIMO Radar Architecture for Energy-Efficient Wireless Sensing in the D-Band | 2025-12-12 | ShowThe D-band offering an untapped wide bandwidth is promising for high data rate communication and high-resolution wireless sensing. However, these potentials are hindered by the low performance and energy efficiency of the D-band circuits and systems. We present an adaptive multi-input multi-output (MIMO) radar architecture for energy-efficient wireless sensing in the D-band, leveraging a reconfigurable 2D array of radar transceiver front-ends, a scaling approach for the receiver (RX) signal-to-noise ratio (SNR) and the transmitter (TX) output power ( |
|
| DoA Estimation with Sparse Arrays: Effects of Antenna Element Patterns and Nonidealities | 2025-11-28 | ShowThis paper studies the effects of directional antenna element complex gain patterns and nonidealities in direction of arrival (DoA) estimation. We compare sparse arrays and classical uniform linear arrays, harnessing EM simulation tools to accurately model the electromagnetic behavior of both patch and Vivaldi antenna element including mutual coupling effects. We show that with sparse array configurations, the performance impacts are significant in terms of DoA estimation accuracy and operable SNR ranges. Specifically, in the scenarios considered, both the usage of directional antenna elements and a sparse array result in over 90% reduction in average direction finding error, compared to a uniform omnidirectional array with the same number of elements (in this case eight), when estimating the directions of two sources using the MUSIC algorithm. For a fixed angular RMSE, the improvements in array sensitivity are shown to yield a 4 to 15-fold increase in one-way coverage distance (assuming free-space path loss). Among the studied options, the best performance was obtained using sparse arrays with either patch or Vivaldi elements for field of views of 100$^\circ$ or 120$^\circ$, respectively. |
|
| Spatial Signal Focusing and Noise Suppression for Direction-of-Arrival Estimation in Large-Aperture 2D Arrays under Demanding Conditions | 2025-10-13 | ShowDirection-of-Arrival (DOA) estimation in sensor arrays faces limitations under demanding conditions, including low signal-to-noise ratio, single-snapshot scenarios, coherent sources, and unknown source counts. Conventional beamforming suffers from sidelobe interference, adaptive methods (e.g., MVDR) and subspace algorithms (e.g., MUSIC) degrade with limited snapshots or coherent signals, while sparse-recovery approaches (e.g., L1-SVD) incur high computational complexity for large arrays. In this article, we construct the concept of the optimal spatial filter to solve the DOA estimation problem under demanding conditions by utilizing the sparsity of spatial signals. By utilizing the concept of the optimal spatial filter, we have transformed the DOA estimation problem into a solution problem for the optimal spatial filter. We propose the Spatial Signal Focusing and Noise Suppression (SSFNS) algorithm, which is a novel DOA estimation framework grounded in the theoretical existence of an optimal spatial filter, to solve for the optimal spatial filter and obtain DOA. Through experiments, it was found that the proposed algorithm is suitable for large aperture two-dimensional arrays and experiments have shown that our proposed algorithm performs better than other algorithms in scenarios with few snapshots or even a single snapshot, low signal-to-noise ratio, coherent signals, and unknown signal numbers in two-dimensional large aperture arrays. |
|
| Single-Snapshot Localization Using Sparse Extremely Large Aperture Arrays | 2025-09-22 | ShowThis paper investigates single-snapshot direction-of-arrival (DOA) estimation and target localization with coherent sparse extremely large aperture arrays (ELAAs) in automotive radar applications. Far-field and near-field signal models are formulated for distributed bistatic configurations. To enable noncoherent processing, a single-snapshot MUSIC (SS-MUSIC) algorithm is proposed to fuse local spectra from individual subarrays and extended to near-field localization via geometric intersection. For coherent processing, a single-snapshot ESPRIT (SS-ESPRIT) method with ambiguity dealiasing is developed to fully exploit the aperture of sparse ELAAs for high-resolution angle estimation. Simulation results demonstrate that SS-ESPRIT provides superior angular resolution for closely spaced far-field targets, while SS-MUSIC offers robustness in near-field localization and flexibility in hybrid scenarios. |
ICASS...ICASSP 2026 manuscript under review |
| Direction of Arrival Estimation: A Tutorial Survey of Classical and Modern Methods | 2025-09-02 | ShowDirection of arrival (DOA) estimation is a fundamental problem in array signal processing with applications spanning radar, sonar, wireless communications, and acoustic signal processing. This tutorial survey provides a comprehensive introduction to classical and modern DOA estimation methods, specifically designed for students and researchers new to the field. We focus on narrowband signal processing using uniform linear arrays, presenting step-by-step mathematical derivations with geometric intuition. The survey covers classical beamforming methods, subspace-based techniques (MUSIC, ESPRIT), maximum likelihood approaches, and sparse signal processing methods. Each method is accompanied by Python implementations available in an open-source repository, enabling reproducible research and hands-on learning. Through systematic performance comparisons across various scenarios, we provide practical guidelines for method selection and parameter tuning. This work aims to bridge the gap between theoretical foundations and practical implementation, making DOA estimation accessible to beginners while serving as a comprehensive reference for the field. See https://github.com/AmgadSalama/DOA for detail implementation of the methods. |
DOA S...DOA Survey, 44 pages, Not published yet |
| Fluid Antenna Enabled Direction-of-Arrival Estimation Under Time-Constrained Mobility | 2025-08-14 | ShowFluid antenna (FA) technology has emerged as a promising approach in wireless communications due to its capability of providing increased degrees of freedom (DoFs) and exceptional design flexibility. This paper addresses the challenge of direction-of-arrival (DOA) estimation for aligned received signals (ARS) and non-aligned received signals (NARS) by designing two specialized uniform FA structures under time-constrained mobility. For ARS scenarios, we propose a fully movable antenna configuration that maximizes the virtual array aperture, whereas for NARS scenarios, we design a structure incorporating a fixed reference antenna to reliably extract phase information from the signal covariance. To overcome the limitations of large virtual arrays and limited sample data inherent in time-varying channels (TVC), we introduce two novel DOA estimation methods: TMRLS-MUSIC for ARS, combining Toeplitz matrix reconstruction (TMR) with linear shrinkage (LS) estimation, and TMR-MUSIC for NARS, utilizing sub-covariance matrices to construct virtual array responses. Both methods employ Nystrom approximation to significantly reduce computational complexity while maintaining estimation accuracy. Theoretical analyses and extensive simulation results demonstrate that the proposed methods achieve underdetermined DOA estimation using minimal FA elements, outperform conventional methods in estimation accuracy, and substantially reduce computational complexity. |
13 pages |
| DOA Estimation via Continuous Aperture Arrays: MUSIC and CRLB | 2025-07-28 | ShowDirection-of-arrival (DOA) estimation using continuous aperture array (CAPA) is studied. Compared to the conventional spatially discrete array (SPDA), CAPA significantly enhances the spatial degrees-of-freedoms (DoFs) for DOA estimation, but its infinite-dimensional continuous signals render the conventional estimation algorithm non-applicable. To address this challenge, a new multiple signal classification (MUSIC) algorithm is proposed for CAPAs. In particular, an equivalent continuous-discrete transformation is proposed to facilitate the eigendecomposition of continuous operators. Subsequently, the MUSIC spectrum is accurately approximated using the Gauss-Legendre quadrature, effectively reducing the computational complexity. Furthermore, the Cramér-Rao lower bounds (CRLBs) for DOA estimation using CAPAs are analyzed for both cases with and without priori knowledge of snapshot signals. It is theoretically proved that CAPAs significantly improve the DOA estimation accuracy compared to traditional SPDAs. Numerical results further validate this insight and demonstrate the effectiveness of the proposed MUSIC algorithm for CAPA. The proposed method achieves near-optimal estimation performance while maintaining a low computational complexity. |
Submi...Submit to possible IEEE journal |
| Near Field Localization via AI-Aided Subspace Methods | 2025-06-27 | ShowThe increasing demands for high-throughput and energy-efficient wireless communications are driving the adoption of extremely large antennas operating at high-frequency bands. In these regimes, multiple users will reside in the radiative near-field, and accurate localization becomes essential. Unlike conventional far-field systems that rely solely on DOA estimation, near-field localization exploits spherical wavefront propagation to recover both DOA and range information. While subspace-based methods, such as MUSIC and its extensions, offer high resolution and interpretability for near-field localization, their performance is significantly impacted by model assumptions, including non-coherent sources, well-calibrated arrays, and a sufficient number of snapshots. To address these limitations, this work proposes AI-aided subspace methods for near-field localization that enhance robustness to real-world challenges. Specifically, we introduce NF-SubspaceNet, a deep learning-augmented 2D MUSIC algorithm that learns a surrogate covariance matrix to improve localization under challenging conditions, and DCD-MUSIC, a cascaded AI-aided approach that decouples angle and range estimation to reduce computational complexity. We further develop a novel model-order-aware training method to accurately estimate the number of sources, that is combined with casting of near field subspace methods as AI models for learning. Extensive simulations demonstrate that the proposed methods outperform classical and existing deep-learning-based localization techniques, providing robust near-field localization even under coherent sources, miscalibrations, and few snapshots. |
Under...Under review for publication in the IEEE |
| Title | Date | Abstract | Comment |
|---|---|---|---|
| Spatially Filtered Sparse Bayesian Learning for Direction-of-Arrival Estimation with Leaky-Wave Antennas | 2025-10-12 | ShowDirection-of-arrival (DoA) estimation with leaky-wave antennas (LWAs) offers a compact and cost-effective alternative to conventional antenna arrays but remains challenging in the presence of coherent sources. To address this issue, we propose a spatially filtered sparse Bayesian learning (SF-SBL) framework. Firstly, the field of view (FoV) is divided into angular sectors according to the frequency beam-scanning property of LWAs, and Bayesian inverse problems are then solved within each sector to improve efficiency and reduce computational cost. Both on-grid SBL and off-grid SBL formulations are developed. Simulation results show that the proposed approach achieves robust and accurate DoA estimation, even with coherent sources. |
Prepr...Preprint submitted to ICASSP 2026. 4 pages, 3 figures |
| Sparse Bayesian Learning for DOA Estimation in Heteroscedastic Noise | 2017-11-08 | ShowThe paper considers direction of arrival (DOA) estimation from long-term observations in a noisy environment. In such an environment the noise source might evolve, causing the stationary models to fail. Therefore a heteroscedastic Gaussian noise model is introduced where the variance can vary across observations and sensors. The source amplitudes are assumed independent zero-mean complex Gaussian distributed with unknown variances (i.e. the source powers), inspiring stochastic maximum likelihood DOA estimation. The DOAs of plane waves are estimated from multi-snapshot sensor array data using sparse Bayesian learning (SBL) where the noise is estimated across both sensors and snapshots. This SBL approach is more flexible and performs better than high-resolution methods since they cannot estimate the heteroscedastic noise process. An alternative to SBL is simple data normalization, whereby only the phase across the array is utilized. Simulations demonstrate that taking the heteroscedastic noise into account improves DOA estimation. |
Submi...Submitted to IEEE TSP |
| Title | Date | Abstract | Comment |
|---|---|---|---|
| Low-Complexity Gridless Single-Snapshot DoA Estimation via Truncated Hankel Newton-MUSIC | 2026-07-09 | ShowReconfigurable antenna arrays can provide enhanced spatial Degrees of Freedom (DoFs) for Integrated Sensing And Communication (ISAC) systems, enabling high-resolution Direction of Arrival (DoA) estimation. In highly dynamic scenarios, however, DoA estimation must be performed within short coherence intervals, which often restricts processing to a single snapshot. Conventional subspace methods then suffer from rank deficiency, while Hankel-based spatial smoothing incurs high computational cost when the array size is large. This paper proposes a low-complexity gridless Truncated Hankel NewtonMUSIC framework for single-snapshot DoA estimation. The proposed method constructs a truncated Hankel matrix with a fixed row dimension to recover an effective signal subspace while reducing the cost of correlation construction and subspace decomposition. When the truncation length is independent of the array size, the dominant complexity scales linearly with the number of antenna ports. To reduce grid-induced quantization errors, coarse grid estimates are further refined by a secondorder Newton update in the continuous angular domain. Simulation results show that the proposed method achieves DoA estimation accuracy close to square Hankel Newton-MUSIC while substantially reducing runtime. For large arrays, it provides more than two orders of magnitude runtime reduction compared with conventional square Hankel MUSIC, making it suitable for realtime sensing in reconfigurable antenna-enabled ISAC systems. |
|
| How Many RF Chains Does a Microwave Linear Analog Computer (MiLAC) Need to Match the Fully-Digital Cramér-Rao Bound? | 2026-06-22 | ShowA microwave linear analog computer (MiLAC) is a tunable microwave network that performs linear operations directly on radio-frequency signals through wave propagation. Used as an antenna-array front end, it can map many antenna signals to a small number of active RF chains. While lossless reciprocal MiLACs have been shown to provide flexible or capacity-achieving beamforming for wireless communications, their sensing performance remains largely unexplored. We analyze direction-of-arrival estimation for |
Submi...Submitting to the IEEE for possible publication |
| G-iMUSIC: Greedy Iterative MUSIC Algorithms for Multi-Target DoA Estimation | 2026-05-27 | ShowThis paper presents novel algorithms for multi-target direction-of-arrival (DoA) estimation in array signal processing. Although the maximum likelihood estimator (MLE) asymptotically attains the Cramér-Rao bound, its exponential complexity motivates practical alternatives, such as greedy or subspace-based methods. In this context, greedy methods such as orthogonal matching pursuit (OMP) and orthogonal least squares (OLS) are sensitive to early selection errors, especially for angularly proximate targets, whereas subspace-based methods such as multiple signal classification (MUSIC) present angular super-resolution capabilities but degrade under strong inter-target signal correlation. To overcome these limitations, we propose two greedy iterative MUSIC (G-iMUSIC) algorithms, namely OMP-iMUSIC and OLS-iMUSIC, derived from a unified framework that links subspace and greedy estimations. Unlike prior iMUSIC approaches, the proposed methods require only one initial eigen value decomposition (EVD) and avoid computing eigendecomposition at each iteration. They also admit Fast Fourier Transform (FFT)-accelerated implementations for uniform linear arrays (ULAs), enabling low-complexity operation. Monte Carlo simulations demonstrate improved detection and precision over conventional OMP, OLS, and MUSIC, as well as reduced processing time compared to greedy baselines. Finally, we introduce diagnostic metrics that interpret performance across signal correlation and angular proximity regimes, supporting generalization beyond the specific orthogonal frequency-division multiplexing (OFDM) radar scenario considered. |
12 pa...12 pages; This work has been submitted to the IEEE for possible publication |
| Interpretable Binaural Deep Beamforming Guided by Time-Varying Relative Transfer Function | 2026-02-17 | ShowIn this work, we propose a deep beamforming framework for speech enhancement in dynamic acoustic environments. The framework learns time-varying beamformer weights from noisy multichannel signals via a deep neural network, guided by a continuously tracked relative transfer function (RTF) of a moving target speaker. We analyze the network's spatial behavior on an 8-microphone linear array by evaluating narrowband and wideband beampatterns in three modes: (i) oracle guidance with true RTFs, (ii) guidance with subspace-tracked RTF estimates, and (iii) operation without RTF guidance. Results show that RTF guidance yields smoother, more spatially consistent beampatterns that track the target direction of arrival (DOA), whereas the unguided model fails to maintain a clear spatial focus. We further extend the framework to binaural beamforming for dynamic target-speaker enhancement. The system is trained using a head-related transfer function (HRTF)-based acoustic simulation of a moving source, enabling realistic spatial rendering at the left and right ears. Spatial cue preservation is quantitatively evaluated in terms of interaural level differences (ILD) and interaural time differences (ITD), demonstrating the method's suitability for hearable applications. |
|
| Sensing for Free: Learn to Localize More Sources than Antennas without Pilots | 2026-01-08 | ShowIntegrated sensing and communication (ISAC) represents a key paradigm for future wireless networks. However, existing approaches require waveform modifications, dedicated pilots, or overhead that complicates standards integration. We propose sensing for free - performing multi-source localization without pilots by reusing uplink data symbols, making sensing occur during transmission and directly compatible with 3GPP 5G NR and 6G specifications. With ever-increasing devices in dense 6G networks, this approach is particularly compelling when combined with sparse arrays, which can localize more sources than uniform arrays via an enlarged virtual array. Existing pilot-free multi-source localization algorithms first reconstruct an extended covariance matrix and apply subspace methods, incurring cubic complexity and limited to second-order statistics. Performance degrades under non-Gaussian data symbols and few snapshots, and higher-order statistics remain unexploited. We address these challenges with an attention-only transformer that directly processes raw signal snapshots for grid-less end-to-end direction-of-arrival (DOA) estimation. The model efficiently captures higher-order statistics while being permutation-invariant and adaptive to varying snapshot counts. Our algorithm greatly outperforms state-of-the-art AI-based benchmarks with over 30x reduction in parameters and runtime, and enjoys excellent generalization under practical mismatches. Applied to multi-user MIMO beam training, our algorithm can localize uplink DOAs of multiple users during data transmission. Through angular reciprocity, estimated uplink DOAs prune downlink beam sweeping candidates and improve throughput via sensing-assisted beam management. This work shows how reusing existing data transmission for sensing can enhance both multi-source localization and beam management in 3GPP efforts towards 6G. |
17 pa...17 pages, 14 figures, 1 table. This paper was accepted by the IEEE Journal on Selected Areas in Communications (JSAC) on Jan. 5, 2026 |
| Efficient Decoders for Sensing Subspace Code | 2025-12-04 | ShowSparse antenna array sensing of source/target via direction of arrival (DoA) estimation motivates design of the sensing framework in joint communication and sensing (JCAS) systems for sixth generation (6G) communication systems. Recently, it is established by Mahdavifar, Rajamäki, and Pal that array geometry of sparse arrays has fundamental connections with the design of subspace codes in coding theory. This was then utilized to design efficient \textit{sensing subspace codes} that estimate the DoA with good resolution. Specifically, the Bose-Chowla sensing subspace code provides near optimal code design for unique DoA estimation with tight theoretical upper bound on the error performance. However, the currently known decoder for these codes, to estimate the DoA, is a traditional \textit{Maximum-a-Posterior (MAP) decoder} with complexity that is cubic with the number of antennas. In this work, we propose novel efficient decoding algorithms for sensing subspace codes, that reduce the complexity down to quadratic while providing new knobs to tune in order to tradeoff complexity with error performance. The decoders are further evaluated for their performance via Monte Carlo simulations for a range of SNRs demonstrating promising performance that smoothly approaches the MAP performance as the complexity grows from quadratic to cubic in the number of antennas. |
This ...This paper was accepted for presentation at the 59th Annual Asilomar Conference on Signals, Systems, and Computers |
| Spatial Signal Focusing and Noise Suppression for Direction-of-Arrival Estimation in Large-Aperture 2D Arrays under Demanding Conditions | 2025-10-13 | ShowDirection-of-Arrival (DOA) estimation in sensor arrays faces limitations under demanding conditions, including low signal-to-noise ratio, single-snapshot scenarios, coherent sources, and unknown source counts. Conventional beamforming suffers from sidelobe interference, adaptive methods (e.g., MVDR) and subspace algorithms (e.g., MUSIC) degrade with limited snapshots or coherent signals, while sparse-recovery approaches (e.g., L1-SVD) incur high computational complexity for large arrays. In this article, we construct the concept of the optimal spatial filter to solve the DOA estimation problem under demanding conditions by utilizing the sparsity of spatial signals. By utilizing the concept of the optimal spatial filter, we have transformed the DOA estimation problem into a solution problem for the optimal spatial filter. We propose the Spatial Signal Focusing and Noise Suppression (SSFNS) algorithm, which is a novel DOA estimation framework grounded in the theoretical existence of an optimal spatial filter, to solve for the optimal spatial filter and obtain DOA. Through experiments, it was found that the proposed algorithm is suitable for large aperture two-dimensional arrays and experiments have shown that our proposed algorithm performs better than other algorithms in scenarios with few snapshots or even a single snapshot, low signal-to-noise ratio, coherent signals, and unknown signal numbers in two-dimensional large aperture arrays. |
|
| Joint DOA and Attitude Sensing Based on Tri-Polarized Continuous Aperture Array | 2025-10-02 | ShowThis paper investigates joint direction-of-arrival (DOA) and attitude sensing using tri-polarized continuous aperture arrays (CAPAs). By employing electromagnetic (EM) information theory, the spatially continuous received signals in tri-polarized CAPA are modeled, thereby enabling accurate DOA and attitude estimation. To facilitate subspace decomposition for continuous operators, an equivalent continuous-discrete transformation technique is developed. Moreover, both self- and cross-covariances of tri-polarized signals are exploited to construct a tri-polarized spectrum, significantly enhancing DOA estimation performance. Theoretical analyses reveal that the identifiability of attitude information fundamentally depends on the availability of prior target snapshots. Accordingly, two attitude estimation algorithms are proposed: one capable of estimating partial attitude information without prior knowledge, and the other achieving full attitude estimation when such knowledge is available. Numerical results demonstrate the feasibility and superiority of the proposed framework. |
13 pages, 10 figures |
| Direction of Arrival Estimation: A Tutorial Survey of Classical and Modern Methods | 2025-09-02 | ShowDirection of arrival (DOA) estimation is a fundamental problem in array signal processing with applications spanning radar, sonar, wireless communications, and acoustic signal processing. This tutorial survey provides a comprehensive introduction to classical and modern DOA estimation methods, specifically designed for students and researchers new to the field. We focus on narrowband signal processing using uniform linear arrays, presenting step-by-step mathematical derivations with geometric intuition. The survey covers classical beamforming methods, subspace-based techniques (MUSIC, ESPRIT), maximum likelihood approaches, and sparse signal processing methods. Each method is accompanied by Python implementations available in an open-source repository, enabling reproducible research and hands-on learning. Through systematic performance comparisons across various scenarios, we provide practical guidelines for method selection and parameter tuning. This work aims to bridge the gap between theoretical foundations and practical implementation, making DOA estimation accessible to beginners while serving as a comprehensive reference for the field. See https://github.com/AmgadSalama/DOA for detail implementation of the methods. |
DOA S...DOA Survey, 44 pages, Not published yet |
| Near Field Localization via AI-Aided Subspace Methods | 2025-06-27 | ShowThe increasing demands for high-throughput and energy-efficient wireless communications are driving the adoption of extremely large antennas operating at high-frequency bands. In these regimes, multiple users will reside in the radiative near-field, and accurate localization becomes essential. Unlike conventional far-field systems that rely solely on DOA estimation, near-field localization exploits spherical wavefront propagation to recover both DOA and range information. While subspace-based methods, such as MUSIC and its extensions, offer high resolution and interpretability for near-field localization, their performance is significantly impacted by model assumptions, including non-coherent sources, well-calibrated arrays, and a sufficient number of snapshots. To address these limitations, this work proposes AI-aided subspace methods for near-field localization that enhance robustness to real-world challenges. Specifically, we introduce NF-SubspaceNet, a deep learning-augmented 2D MUSIC algorithm that learns a surrogate covariance matrix to improve localization under challenging conditions, and DCD-MUSIC, a cascaded AI-aided approach that decouples angle and range estimation to reduce computational complexity. We further develop a novel model-order-aware training method to accurately estimate the number of sources, that is combined with casting of near field subspace methods as AI models for learning. Extensive simulations demonstrate that the proposed methods outperform classical and existing deep-learning-based localization techniques, providing robust near-field localization even under coherent sources, miscalibrations, and few snapshots. |
Under...Under review for publication in the IEEE |
| Mainlobe Jamming Suppression Using MIMO-STCA Radar | 2025-05-14 | ShowRadar jamming suppression, particularly against mainlobe jamming, has become a critical focus in modern radar systems. This article investigates advanced mainlobe jamming suppression techniques utilizing a novel multiple-input multiple-output space-time coding array (MIMO-STCA) radar. Extending the capabilities of traditional MIMO radar, the MIMO-STCA framework introduces additional degrees of freedom (DoFs) in the range domain through the utilization of transmit time delays, offering enhanced resilience against interference. One of the key challenges in mainlobe jamming scenarios is the difficulty in obtaining interference-plus-noise samples that are free from target signal contamination. To address this, the study introduces a cumulative sampling-based non-homogeneous sample selection (CS-NHSS) algorithm to remove target-contaminated samples, ensuring accurate interference-plus-noise covariance matrix estimation and effective noise subspace separation. Building on this, the subsequent step is to apply the proposed noise subspace-based jamming mitigation (NSJM) algorithm, which leverages the orthogonality between noise and jamming subspace for effective jamming mitigation. However, NSJM performance can degrade due to spatial frequency mismatches caused by DoA or range quantization errors. To overcome this limitation, the study further proposes the robust jamming mitigation via noise subspace (RJNS) algorithm, incorporating adaptive beampattern control to achieve a flat-top mainlobe and broadened nulls, enhancing both anti-jamming effectiveness and robustness under non-ideal conditions. Simulation results verify the effectiveness of the proposed algorithms. Significant improvements in mainlobe jamming suppression are demonstrated through transmit-receive beampattern analysis and enhanced signal-to-interference-plus-noise ratio (SINR) curve. |
|
| A Comparative Study of Invariance-Aware Loss Functions for Deep Learning-based Gridless Direction-of-Arrival Estimation | 2025-03-16 | ShowCovariance matrix reconstruction has been the most widely used guiding objective in gridless direction-of-arrival (DoA) estimation for sparse linear arrays. Many semidefinite programming (SDP)-based methods fall under this category. Although deep learning-based approaches enable the construction of more sophisticated objective functions, most methods still rely on covariance matrix reconstruction. In this paper, we propose new loss functions that are invariant to the scaling of the matrices and provide a comparative study of losses with varying degrees of invariance. The proposed loss functions are formulated based on the scale-invariant signal-to-distortion ratio between the target matrix and the Gram matrix of the prediction. Numerical results show that a scale-invariant loss outperforms its non-invariant counterpart but is inferior to the recently proposed subspace loss that is invariant to the change of basis. These results provide evidence that designing loss functions with greater degrees of invariance is advantageous in deep learning-based gridless DoA estimation. |
5 pag...5 pages. Accepted at ICASSP 2025 |
| Completing Sets of Prototype Transfer Functions for Subspace-based Direction of Arrival Estimation of Multiple Speakers | 2025-01-13 | ShowTo estimate the direction of arrival (DOA) of multiple speakers, subspace-based prototype transfer function matching methods such as multiple signal classification (MUSIC) or relative transfer function (RTF) vector matching are commonly employed. In general, these methods require calibrated microphone arrays, which are characterized by a known array geometry or a set of known prototype transfer functions for several directions. In this paper, we consider a partially calibrated microphone array, composed of a calibrated binaural hearing aid and a (non-calibrated) external microphone at an unknown location with no available set of prototype transfer functions. We propose a procedure for completing sets of prototype transfer functions by exploiting the orthogonality of subspaces, allowing to apply matching-based DOA estimation methods with partially calibrated microphone arrays. For the MUSIC and RTF vector matching methods, experimental results for two speakers in noisy and reverberant environments clearly demonstrate that for all locations of the external microphone DOAs can be estimated more accurately with completed sets of prototype transfer functions than with incomplete sets. \c{opyright}20XX IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works. |
Accep...Accepted for ICASSP 2025 |
| Low Complexity DoA-ToA Signature Estimation for Multi-Antenna Multi-Carrier Systems | 2024-09-13 | ShowAccurate direction of arrival (DoA) and time of arrival (ToA) estimation is an stringent requirement for several wireless systems like sonar, radar, communications, and dual-function radar communication (DFRC). Due to the use of high carrier frequency and bandwidth, most of these systems are designed with multiple antennae and subcarriers. Although the resolution is high in the large array regime, the DoA-ToA estimation accuracy of the practical on-grid estimation methods still suffers from estimation inaccuracy due to the spectral leakage effect. In this article, we propose DoA-ToA estimation methods for multi-antenna multi-carrier systems with an orthogonal frequency division multiplexing (OFDM) signal. In the first method, we apply discrete Fourier transform (DFT) based coarse signature estimation and propose a low complexity multistage fine-tuning for extreme enhancement in the estimation accuracy. The second method is based on compressed sensing, where we achieve the super-resolution by taking a 2D-overcomplete angle-delay dictionary than the actual number of antenna and subcarrier basis. Unlike the vectorized 1D-OMP method, we apply the low complexity 2D-OMP method on the matrix data model that makes the use of CS methods practical in the context of large array regimes. Through numerical simulations, we show that our proposed methods achieve the similar performance as that of the subspace-based 2D-MUSIC method with a significant reduction in computational complexity. |
5 pag...5 pages, 4 figures, 1 table |
| Direction of Arrival Estimation with Sparse Subarrays | 2024-08-17 | ShowThis paper proposes design techniques for partially-calibrated sparse linear subarrays and algorithms to perform direction-of-arrival (DOA) estimation. First, we introduce array architectures that incorporate two distinct array categories, namely type-I and type-II arrays. The former breaks down a known sparse linear geometry into as many pieces as we need, and the latter employs each subarray such as it fits a preplanned sparse linear geometry. Moreover, we devise two Direction of Arrival (DOA) estimation algorithms that are suitable for partially-calibrated array scenarios within the coarray domain. The algorithms are capable of estimating a greater number of sources than the number of available physical sensors, while maintaining the hardware and computational complexity within practical limits for real-time implementation. To this end, we exploit the intersection of projections onto affine spaces by devising the Generalized Coarray Multiple Signal Classification (GCA-MUSIC) in conjunction with the estimation of a refined projection matrix related to the noise subspace, as proposed in the GCA root-MUSIC algorithm. An analysis is performed for the devised subarray configurations in terms of degrees of freedom, as well as the computation of the Cramèr-Rao Lower Bound for the utilized data model, in order to demonstrate the good performance of the proposed methods. Simulations assess the performance of the proposed design methods and algorithms against existing approaches. |
15 pages, 8 figures |
| Analysis of Partially-Calibrated Sparse Subarrays for Direction Finding with Extended Degrees of Freedom | 2024-08-06 | ShowThis paper investigates the problem of direction-of-arrival (DOA) estimation using multiple partially-calibrated sparse subarrays. In particular, we present the Generalized Coarray Multiple Signal Classification (GCA-MUSIC) DOA estimation algorithm to scenarios with partially-calibrated sparse subarrays. The proposed GCA-MUSIC algorithm exploits the difference coarray for each subarray, followed by a specific pseudo-spectrum merging rule that is based on the intersection of the signal subspaces associated to each subarray. This rule assumes that there is no a priori knowledge about the cross-covariance between subarrays. In that way, only the second-order statistics of each subarray are used to estimate the directions with increased degrees of freedom, i.e., the estimation procedure preserves the coarray Multiple Signal Classification and sparse arrays properties to estimate more sources than the number of physical sensors in each subarray. Numerical simulations show that the proposed GCA-MUSIC has better performance than other similar strategies. |
6 pages, 5 figures |
| SubspaceNet: Deep Learning-Aided Subspace Methods for DoA Estimation | 2024-07-11 | ShowDirection of arrival (DoA) estimation is a fundamental task in array processing. A popular family of DoA estimation algorithms are subspace methods, which operate by dividing the measurements into distinct signal and noise subspaces. Subspace methods, such as Multiple Signal Classification (MUSIC) and Root-MUSIC, rely on several restrictive assumptions, including narrowband non-coherent sources and fully calibrated arrays, and their performance is considerably degraded when these do not hold. In this work we propose SubspaceNet; a data-driven DoA estimator which learns how to divide the observations into distinguishable subspaces. This is achieved by utilizing a dedicated deep neural network to learn the empirical autocorrelation of the input, by training it as part of the Root-MUSIC method, leveraging the inherent differentiability of this specific DoA estimator, while removing the need to provide a ground-truth decomposable autocorrelation matrix. Once trained, the resulting SubspaceNet serves as a universal surrogate covariance estimator that can be applied in combination with any subspace-based DoA estimation method, allowing its successful application in challenging setups. SubspaceNet is shown to enable various DoA estimation algorithms to cope with coherent sources, wideband signals, low SNR, array mismatches, and limited snapshots, while preserving the interpretability and the suitability of classic subspace methods. |
Under...Under review for publication in the IEEE |
| Subspace Coding for Spatial Sensing | 2024-07-03 | ShowA subspace code is defined as a collection of subspaces of an ambient vector space, where each information-encoding codeword is a subspace. This paper studies a class of spatial sensing problems, notably direction of arrival (DoA) estimation using multisensor arrays, from a novel subspace coding perspective. Specifically, we demonstrate how a canonical (passive) sensing model can be mapped into a subspace coding problem, with the sensing operation defining a unique structure for the subspace codewords. We introduce the concept of sensing subspace codes following this structure, and show how these codes can be controlled by judiciously designing the sensor array geometry. We further present a construction of sensing subspace codes leveraging a certain class of Golomb rulers that achieve near-optimal minimum codeword distance. These designs inspire novel noise-robust sparse array geometries achieving high angular resolution. We also prove that codes corresponding to conventional uniform linear arrays are suboptimal in this regard. This work is the first to establish connections between subspace coding and spatial sensing, with the aim of leveraging insights and methodologies in one field to tackle challenging problems in the other. |
©2024...©2024 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works |
| Gridless Parameter Estimation in Partly Calibrated Rectangular Arrays | 2024-06-23 | ShowSpatial frequency estimation from a mixture of noisy sinusoids finds applications in various fields. While subspace-based methods offer cost-effective super-resolution parameter estimation, they demand precise array calibration, posing challenges for large antennas. In contrast, sparsity-based approaches outperform subspace methods, especially in scenarios with limited snapshots or correlated sources. This study focuses on direction-of-arrival (DOA) estimation using a partly calibrated rectangular array with fully calibrated subarrays. A gridless sparse formulation leveraging shift invariances in the array is developed, yielding two competitive algorithms under the alternating direction method of multipliers (ADMM) and successive convex approximation frameworks, respectively. Numerical simulations show the superior error performance of our proposed method, particularly in highly correlated scenarios, compared to the conventional subspace-based methods. It is demonstrated that the proposed formulation can also be adopted in the fully calibrated case to improve the robustness of the subspace-based methods to the source correlation. Furthermore, we provide a generalization of the proposed method to a more challenging case where a part of the sensors is unobservable due to failures. |
16 pa...16 pages, 5 figures. This work has been submitted to the IEEE Transactions on Signal Processing for possible publication |
| Title | Date | Abstract | Comment |
|---|---|---|---|
| BiTSE: Binaural Target Speaker Extraction in Noisy Multi-Talker Environments for AR Glass Arrays | 2026-08-10 | ShowIsolating a desired speech signal in noisy multi-talker conversational scenarios is a key requirement for augmented reality (AR) wearable microphone array systems. In this work, a binaural target speaker extraction (TSE) framework, termed BiTSE, is proposed. It leverages both spatial and temporal cues, specifically the direction-of-arrival (DoA) of the target speaker and corresponding voice activity information, to guide the extraction process. Built upon a binaural signal denoising architecture, our model integrates three key enhancements: (i) a DoA-aware attention mechanism using cyclic positional embeddings, (ii) a timestamp-based masking strategy that utilizes speaker activity to suppress non-target segments, and (iii) a novel two-stage loss optimization strategy that first trains the model for robust denoising and then fine-tunes it to improve perceptual quality. Evaluations on the SPeech Enhancement for Augmented Reality (SPEAR) challenge dataset demonstrate that the proposed BiTSE consistently improves upon conventional approaches, leading to enhanced signal fidelity and perceptual quality. |
This ...This is the preprint version of the paper accepted at APSIPA ASC 2026 |
| Position-Aware Target Speaker Extraction for Long-Form Multi-Party Conversations: A Diarization-Free Framework for ASR | 2026-06-28 | ShowIn long-form multi-party conversations, highly imbalanced speaker activity and frequent overlap make it difficult to identify "who spoke when and what". Sliding-window continuous speech separation (CSS) mitigates sparse supervision, but often suffers from cross-window speaker inconsistency and residual crosstalk, which in practice requires diarization for reliable speaker attribution. Motivated by the stability of speakers' directions of arrival (DOAs) in meetings, we propose PATSE, a multi-channel Position-Aware Target Speaker Extraction front-end that uses DOA as a spatial prior to directly extract the speech of each target speaker. PATSE combines a DOA-guided spatial encoder and conditioner to generate speaker-attributed streams, from which speaker activity can be inferred via simple post-processing (e.g., VAD) without explicit diarization. Experiments on both replayed and real conversations show consistent ASR gains outperforming CSS and diarization-based pipelines. |
5 pag...5 pages, 2 figures, Accept by Interspeech 2026 |
| Direction of arrival estimation from distant microphone data using single frequency filtering | 2026-06-15 | ShowIn distant microphones, broadband (BB) methods for direction-of-arrival (DoA) estimation are more suitable than narrowband (NB) methods. Due to the aggregation of their optimization function across all frequency bands, BB estimators are robust to spatial aliasing, a known problem in processing distant microphone data. In NB methods, DoA estimation is performed by utilizing \textit{local} information in each frequency band and hence the estimation is affected by spatial aliasing. However, unlike BB methods, NB methods exploit frequency sparsity to estimate the DoAs of \textit{multiple speakers} in a \textit{single time frame}. In this article, a method to improve the robustness of a NB DoA estimator to spatial aliasing is developed. The proposed method is based on cross-correlation of speech-present time-frequency regions obtained by single frequency filtering (SFF) of the microphone signals. The SFF spectrum is chosen because SFF components have regions of high signal-to-noise ratio both in time and frequency and because speech and non-speech discrimination is robust to degradations in the SFF domain. The proposed NB estimator is compared to four state-of-the-art estimators (one NB and three BB) using detection and accuracy metrics on simulated and real-world data in different reverberation and noise conditions. The results show that in all the environments, the SFF-based NB approach outperforms the state-of-the-art NB approach. Furthermore, the performance of the SFF-based approach is better than some of the BB estimators. |
|
| Single frequency filtering based multi-speaker direction of arrival estimation from stereo recordings | 2026-06-15 | ShowRobust direction-of-arrival (DoA) estimation from noisy and reverberant microphone signals remains challenging. Conventional estimators such as generalized cross-correlation (GCC) and its variants operate in the short-time Fourier transform (STFT) domain, where spectral features primarily reflect vocal-tract characteristics. Recent single frequency filtering (SFF)-based estimators instead use a time-frequency representation that provides high spectral resolution of harmonics along with high temporal resolution of excitation-source events, such as epoch-like impulses. Since excitation-source features have been shown to be more robust to noise and reverberation than spectral features, this work proposes an improved SFF-based DoA estimator that correlates the envelopes of SFF outputs across microphone channels using PHAT-weighted GCC. We further provide a comprehensive evaluation of SFF-based and state-of-the-art GCC-based estimators using publicly available real-room recordings under challenging reverberant, multi-speaker, and noise-corrupted conditions. Experimental results show that the proposed method and an existing SFF-based estimator achieve detection and accuracy performance that is superior or comparable to the best GCC-based estimator across all test cases. We also demonstrate that using speech-dominant bins improves GCC-PHAT robustness, motivating future incorporation of such weighting strategies into SFF-based DoA estimation. |
|
| DOA: Training-Free Decoder-Only Attention Policy for Long-Form Simultaneous Translation with SpeechLLMs | 2026-05-29 | ShowSimultaneous speech-to-text translation (SimulST) generates translations while speech is still unfolding, requiring a streaming policy that decides when to read and when to write. State-of-the-art approaches rely on attention-based encoder-decoder models where cross-attention provides explicit alignment signals. In contrast, Speech Large Language Models (SpeechLLMs) are decoder-only architectures relying solely on self-attention. This raises a central question: whether decoder self-attention contains sufficiently stable alignment signals to guide the streaming policy. Moreover, existing approaches typically rely on training-based adaptations or heuristic wait-$k$ policies and have not been validated in long-form settings. To fill these gaps, we propose Decoder-Only Attention (DOA), a training-free policy that enables long-form simultaneous translation with off-the-shelf SpeechLLMs by deriving a proxy alignment from self-attention. Experiments on Phi4-Multimodal and Qwen3-Omni show that DOA provides an effective alignment signal for supporting streaming decisions, enabling low-latency long-form SimulST with quality close to offline decoding without retraining. |
|
| IsoNet: Spatially-aware audio-visual target speech extraction in complex acoustic environments | 2026-05-15 | ShowTarget speech extraction remains difficult for compact devices because monaural neural models lack spatial evidence and classical beamformers lose resolving power when the microphone aperture is only a few centimetres. We present IsoNet, a user-selectable audio-visual target speech extraction system for a compact 4-microphone array. IsoNet combines complex multi-channel STFT features, GCC-PHAT spatial cues, face-conditioned visual embeddings, and auxiliary direction-of-arrival supervision inside a U-Net mask estimation network. Three curriculum variants were trained on 25,000 simulated VoxCeleb mixtures with progressively difficult SNR regimes. On a hard test set spanning -1 to 10 dB SNR, IsoNet-CL1 achieves 9.31 dB SI-SDR, a 4.85 dB improvement over the mixture, with PESQ 2.13 and STOI 0.84. Oracle delay-and-sum and MVDR beamformers degrade the same mixtures by 4.82 dB and 6.08 dB SI-SDRi, respectively, showing that the proposed learned multimodal conditioning solves a regime where conventional spatial filtering is ineffective. Ablation studies show consistent gains from visual conditioning, GCC-PHAT features, and extended delay-bin encoding. The results establish a compact-array, face-selectable speech extraction baseline under controlled simulation and identify the remaining barriers to real deployment, especially phase reconstruction, multi-interferer mixtures, and simulation-to-real transfer. |
8 pages |
| Direction-Preserving MIMO Speech Enhancement Using a Neural Covariance Estimator | 2026-04-13 | ShowMultichannel speech enhancement is widely used as a front-end in microphone array processing systems. While most existing approaches produce a single enhanced signal, direction-preserving multiple-input multiple-output (MIMO) methods instead aim to provide enhanced multichannel signals that retain directional properties, enabling downstream applications such as beamforming, binaural rendering, and direction-of-arrival estimation. In this work, we propose a fully blind, direction-preserving MIMO speech enhancement method based on neural estimation of the spatial noise covariance matrix. A lightweight OnlineSpatialNet estimates a scale-normalized Cholesky factor of the frequency-domain noise covariance, which is combined with a direction-preserving MIMO Wiener filter to enhance speech while preserving the spatial characteristics of both target and residual noise. In contrast to prior approaches relying on oracle information or mask-based covariance estimation for single-output systems, the proposed method directly targets accurate multichannel covariance estimation with low computational complexity. Experimental results show improved speech enhancement, covariance estimation capability, and performance in downstream tasks over a mask-based baseline, approaching oracle performance with significantly fewer parameters and computational cost. |
|
| Reverberation-Robust Localization of Speakers Using Distinct Speech Onsets and Multi-channel Cross-Correlations | 2026-04-02 | ShowMany speaker localization methods can be found in the literature. However, speaker localization under strong reverberation still remains a major challenge in the real-world applications. This paper proposes two algorithms for localizing speakers using microphone array recordings of reverberated sounds. To separate concurrent speakers, the first algorithm decomposes microphone signals spectrotemporally into subbands via an auditory filterbank. To suppress reverberation, we propose a novel speech onset detection approach derived from the speech signal and impulse response models, and further propose to formulate the multi-channel cross-correlation coefficient (MCCC) of encoded speech onsets in each subband. The subband results are combined to estimate the directions-of-arrival (DOAs) of speakers. The second algorithm extends the generalized cross-correlation - phase transform (GCC-PHAT) method by using redundant information of multiple microphones to address the reverberation problem. The proposed methods have been evaluated under adverse conditions using not only simulated signals (reverberation time |
|
| HRTF-guided Binaural Target Speaker Extraction with Real-World Validation | 2026-03-17 | ShowThis paper presents a Head-Related Transfer Function (HRTF)-guided framework for binaural Target Speaker Extraction (TSE) from mixtures of concurrent sources. Unlike conventional TSE methods based on Direction of Arrival (DOA) estimation or enrollment signals, which often distort perceived spatial location, the proposed approach leverages the listener's HRTF as an explicit spatial prior. The proposed framework is built upon a multi-channel deep blind source separation backbone, adapted to the binaural TSE setting. It is trained on measured HRTFs from a diverse population, enabling cross-listener generalization rather than subject-specific tuning. By conditioning the extraction on HRTF-derived spatial information, the method preserves binaural cues while enhancing speech quality and intelligibility. The performance of the proposed framework is validated through simulations and real recordings obtained from a head and torso simulator (HATS). |
Submi...Submitted to Interspeech 2026 |
| Interpretable Binaural Deep Beamforming Guided by Time-Varying Relative Transfer Function | 2026-02-17 | ShowIn this work, we propose a deep beamforming framework for speech enhancement in dynamic acoustic environments. The framework learns time-varying beamformer weights from noisy multichannel signals via a deep neural network, guided by a continuously tracked relative transfer function (RTF) of a moving target speaker. We analyze the network's spatial behavior on an 8-microphone linear array by evaluating narrowband and wideband beampatterns in three modes: (i) oracle guidance with true RTFs, (ii) guidance with subspace-tracked RTF estimates, and (iii) operation without RTF guidance. Results show that RTF guidance yields smoother, more spatially consistent beampatterns that track the target direction of arrival (DOA), whereas the unguided model fails to maintain a clear spatial focus. We further extend the framework to binaural beamforming for dynamic target-speaker enhancement. The system is trained using a head-related transfer function (HRTF)-based acoustic simulation of a moving source, enabling realistic spatial rendering at the left and right ears. Spatial cue preservation is quantitatively evaluated in terms of interaural level differences (ILD) and interaural time differences (ITD), demonstrating the method's suitability for hearable applications. |
|
| DOA Estimation with Lightweight Network on LLM-Aided Simulated Acoustic Scenes | 2025-11-11 | ShowDirection-of-Arrival (DOA) estimation is critical in spatial audio and acoustic signal processing, with wide-ranging applications in real-world. Most existing DOA models are trained on synthetic data by convolving clean speech with room impulse responses (RIRs), which limits their generalizability due to constrained acoustic diversity. In this paper, we revisit DOA estimation using a recently introduced dataset constructed with the assistance of large language models (LLMs), which provides more realistic and diverse spatial audio scenes. We benchmark several representative neural-based DOA methods on this dataset and propose LightDOA, a lightweight DOA estimation model based on depthwise separable convolutions, specifically designed for mutil-channel input in varying environments. Experimental results show that LightDOA achieves satisfactory accuracy and robustness across various acoustic scenes while maintaining low computational complexity. This study not only highlights the potential of spatial audio synthesized with the assistance of LLMs in advancing robust and efficient DOA estimation research, but also highlights LightDOA as efficient solution for resource-constrained applications. |
|
| Mixture-of-Experts Framework for Field-of-View Enhanced Signal-Dependent Binauralization of Moving Talkers | 2025-09-25 | ShowWe propose a novel mixture of experts framework for field-of-view enhancement in binaural signal matching. Our approach enables dynamic spatial audio rendering that adapts to continuous talker motion, allowing users to emphasize or suppress sounds from selected directions while preserving natural binaural cues. Unlike traditional methods that rely on explicit direction-of-arrival estimation or operate in the Ambisonics domain, our signal-dependent framework combines multiple binaural filters in an online manner using implicit localization. This allows for real-time tracking and enhancement of moving sound sources, supporting applications such as speech focus, noise reduction, and world-locked audio in augmented and virtual reality. The method is agnostic to array geometry offering a flexible solution for spatial audio capture and personalized playback in next-generation consumer audio devices. |
5 pages, 3 figures |
| GAN-Based Multi-Microphone Spatial Target Speaker Extraction | 2025-09-22 | ShowSpatial target speaker extraction isolates a desired speaker's voice in multi-speaker environments using spatial information, such as the direction of arrival (DoA). Although recent deep neural network (DNN)-based discriminative methods have shown significant performance improvements, the potential of generative approaches, such as generative adversarial networks (GANs), remains largely unexplored for this problem. In this work, we demonstrate that a GAN can effectively leverage both noisy mixtures and spatial information to extract and generate the target speaker's speech. By conditioning the GAN on intermediate features of a discriminative spatial filtering model in addition to DoA, we enable steerable target extraction with high spatial resolution of 5 degrees, outperforming state-of-the-art discriminative methods in perceptual quality-based objective metrics. |
|
| Learning Robust Spatial Representations from Binaural Audio through Feature Distillation | 2025-08-28 | ShowRecently, deep representation learning has shown strong performance in multiple audio tasks. However, its use for learning spatial representations from multichannel audio is underexplored. We investigate the use of a pretraining stage based on feature distillation to learn a robust spatial representation of binaural speech without the need for data labels. In this framework, spatial features are computed from clean binaural speech samples to form prediction labels. These clean features are then predicted from corresponding augmented speech using a neural network. After pretraining, we throw away the spatial feature predictor and use the learned encoder weights to initialize a DoA estimation model which we fine-tune for DoA estimation. Our experiments demonstrate that the pretrained models show improved performance in noisy and reverberant environments after fine-tuning for direction-of-arrival estimation, when compared to fully supervised models and classic signal processing methods. |
To ap...To appear in Proc. WASPAA 2025, October 12-15, 2025, Tahoe, US. Copyright (c) 2025 IEEE. 5 pages, 2 figures, 2 tables |
| Sound Source Localization for Human-Robot Interaction in Outdoor Environments | 2025-07-29 | ShowThis paper presents a sound source localization strategy that relies on a microphone array embedded in an unmanned ground vehicle and an asynchronous close-talking microphone near the operator. A signal coarse alignment strategy is combined with a time-domain acoustic echo cancellation algorithm to estimate a time-frequency ideal ratio mask to isolate the target speech from interferences and environmental noise. This allows selective sound source localization, and provides the robot with the direction of arrival of sound from the active operator, which enables rich interaction in noisy scenarios. Results demonstrate an average angle error of 4 degrees and an accuracy within 5 degrees of 95% at a signal-to-noise ratio of 1dB, which is significantly superior to the state-of-the-art localization methods. |
|
| End-to-End DOA-Guided Speech Extraction in Noisy Multi-Talker Scenarios | 2025-07-28 | ShowTarget Speaker Extraction (TSE) plays a critical role in enhancing speech signals in noisy and multi-speaker environments. This paper presents an end-to-end TSE model that incorporates Direction of Arrival (DOA) and beamwidth embeddings to extract speech from a specified spatial region centered around the DOA. Our approach efficiently captures spatial and temporal features, enabling robust performance in highly complex scenarios with multiple simultaneous speakers. Experimental results demonstrate that the proposed model not only significantly enhances the target speech within the defined beamwidth but also effectively suppresses interference from other directions, producing a clear and isolated target voice. Furthermore, the model achieves remarkable improvements in downstream Automatic Speech Recognition (ASR) tasks, making it particularly suitable for real-world applications. |
Accep...Accepted by INTERSPEECH 2025 |
| End-to-end multi-channel speaker extraction and binaural speech synthesis | 2025-07-11 | ShowSpeech clarity and spatial audio immersion are the two most critical factors in enhancing remote conferencing experiences. Existing methods are often limited: either due to the lack of spatial information when using only one microphone, or because their performance is highly dependent on the accuracy of direction-of-arrival estimation when using microphone array. To overcome this issue, we introduce an end-to-end deep learning framework that has the capacity of mapping multi-channel noisy and reverberant signals to clean and spatialized binaural speech directly. This framework unifies source extraction, noise suppression, and binaural rendering into one network. In this framework, a novel magnitude-weighted interaural level difference loss function is proposed that aims to improve the accuracy of spatial rendering. Extensive evaluations show that our method outperforms established baselines in terms of both speech quality and spatial fidelity. |
|
| Multi-Channel Acoustic Echo Cancellation Based on Direction-of-Arrival Estimation | 2025-06-06 | ShowAcoustic echo cancellation (AEC) is an important speech signal processing technology that can remove echoes from microphone signals to enable natural-sounding full-duplex speech communication. While single-channel AEC is widely adopted, multi-channel AEC can leverage spatial cues afforded by multiple microphones to achieve better performance. Existing multi-channel AEC approaches typically combine beamforming with deep neural networks (DNN). This work proposes a two-stage algorithm that enhances multi-channel AEC by incorporating sound source directional cues. Specifically, a lightweight DNN is first trained to predict the sound source directions, and then the predicted directional information, multi-channel microphone signals, and single-channel far-end signal are jointly fed into an AEC network to estimate the near-end signal. Evaluation results show that the proposed algorithm outperforms baseline approaches and exhibits robust generalization across diverse acoustic environments. |
Accep...Accepted by Interspeech 2025 |
| Spatial Audio Processing with Large Language Model on Wearable Devices | 2025-04-25 | ShowIntegrating spatial context into large language models (LLMs) has the potential to revolutionize human-computer interaction, particularly in wearable devices. In this work, we present a novel system architecture that incorporates spatial speech understanding into LLMs, enabling contextually aware and adaptive applications for wearable technologies. Our approach leverages microstructure-based spatial sensing to extract precise Direction of Arrival (DoA) information using a monaural microphone. To address the lack of existing dataset for microstructure-assisted speech recordings, we synthetically create a dataset called OmniTalk by using the LibriSpeech dataset. This spatial information is fused with linguistic embeddings from OpenAI's Whisper model, allowing each modality to learn complementary contextual representations. The fused embeddings are aligned with the input space of LLaMA-3.2 3B model and fine-tuned with lightweight adaptation technique LoRA to optimize for on-device processing. SING supports spatially-aware automatic speech recognition (ASR), achieving a mean error of |
| Title | Date | Abstract | Comment |
|---|---|---|---|
| Mind the Microphone Gap: Benchmarking Array Upsampling Strategies for Latent Acoustic Mapping | 2026-07-27 | ShowLatent Acoustic Mapping (LAM) is a self-supervised learning method that generates high-resolution spherical acoustic maps from multichannel recordings without labelled data, matching supervised baselines on direction-of-arrival benchmarks. However, LAM degrades significantly with sparse 4-channel arrays, as the low-resolution cross-spectral matrix captures far less spatial information than the 32-channel inputs LAM was designed for. We benchmark a diverse set of upsampling architectures, spanning lightweight convolutional networks, iterative back-projection models, physics-informed networks, and generative adversarial approaches. We also study whether aligning these upsamplers with LAM by training them jointly or in different stages helps preserve the spatial structure that LAM depends on. Results show that the original full-resolution LAM is the strongest, that separately trained lightweight models are the most competitive learned approaches, and that representation alignment between the upsampler and LAM matters more than model complexity. |
IWAENC 2026 |
| Acoustic Imaging for UAV Detection: Dense Beamformed Energy Maps and U-Net SELD | 2026-07-17 | ShowWe introduce a U-net model for 360° acoustic source localization formulated as a spherical semantic segmentation task. Rather than regressing discrete direction-of-arrival (DoA) angles, our model segments beamformed audio maps (azimuth & elevation) into regions of active sound presence. Using delay-and-sum (DAS) beamforming on a custom 24-microphone array, we generate signals aligned with drone GPS telemetry to create binary supervision masks. A modified U-Net, trained on frequency-domain representations of these maps, learns to identify spatially distributed source regions while addressing class imbalance via the Tversky loss. Because the network operates on beamformed energy maps, the approach is inherently array-independent and can adapt to different microphone configurations and can be transferred to different microphone configurations with minimal adaptation. The segmentation outputs are post-processed by computing centroids over activated regions, enabling robust DoA estimates. Our dataset includes real-world open-field recordings of a DJI Air 3 drone, synchronized with 360° video and flight logs across multiple dates and locations. Experimental results show that U-net generalizes across environments, providing improved angular precision, offering a new paradigm for dense spatial audio understanding beyond traditional Sound Source Localization (SSL). We additionally validate the same beamforming-plus-segmentation formulation on the DCASE 2019 TAU Spatial Sound Events benchmark, showing that the approach generalizes beyond drone acoustics to multiclass Sound Event Localization and Detection (SELD) scenarios. |
|
| NeuralMUSIC: A Hybrid Neural-Subspace Framework for Robot Sound Source Localization | 2026-07-08 | ShowReliable sound source localization is fundamental to robot audition, enabling autonomous robots to perceive spatial cues and operate effectively in dynamic environments. Classical methods such as Multiple Signal Classification (MUSIC) offer strong theoretical foundations but degrade under low signal-to-noise ratios. While deep learning-based approaches achieve promising performance, they often struggle with limited generalization across conditions. To address these challenges, we propose NeuralMUSIC, a hybrid neural-subspace framework for robotic sound source localization. Specifically, a neural network first estimates the spatial covariance matrix from multichannel microphone observations. The predicted covariance is then integrated into a classical MUSIC pipeline with eigenvalue decomposition (EVD) and pseudo-spectrum computation, followed by a Frequency Attention Fusion (FAF) module to produce the final DOA estimates. To improve data efficiency, we further introduce a Self-supervised Spatial Correlation Learning (SSCL) strategy that leverages unlabeled acoustic data to capture spatial structure. Extensive experiments across different robotic tasks demonstrate that NeuralMUSIC achieves competitive localization accuracy while exhibiting improved robustness and cross-domain generalization. |
Accep...Accepted by IROS 2026 |
| SelectTSL: Prompt-Guided Selective Target Sound Localization in Complex Scenarios | 2026-07-02 | ShowHumans can selectively attend to a target sound and estimate its direction in complex scenarios, whereas such selective localization remains challenging for current deep learning-based systems. Sound source localization (SSL) has achieved remarkable success with deep learning, yet most methods localize all active sources without selectivity. Conversely, target sound extraction (TSE) extracts sources using multimodal prompts but typically fails to preserve the multichannel spatial information required for accurate localization. To bridge this gap, we formulate the task of prompt-guided selective target sound localization and propose SelectTSL, an end-to-end architecture that localizes only the user-specified target in multi-source acoustic scenes. Specifically, we design a target-aware selective localization strategy that employs a Prompt-Guided Selective Attention Module (PGSA) to generate prompt-informed embeddings. These embeddings guide an inter-channel phase difference (IPD) enhancer to refine raw phase cues, fusing with target magnitudes to jointly estimate direction of arrival (DoA) and target-source cardinality, i.e., the number of target sound sources. This coupled design effectively focuses on the user-specified target spatial cues for selective localization and also handles time-varying numbers of target sources. Extensive experiments on both synthetic data and real-world recordings demonstrate that our proposed method consistently outperforms other baselines and exhibits robust generalization to real acoustic environments. |
|
| JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments | 2026-05-28 | ShowCurrent audio-visual large language models (AV-LLMs) are predominantly restricted to 2D perception, relying on RGB video and monaural audio. This design choice introduces a fundamental dimensionality mismatch that precludes reliable source localization and spatial reasoning in complex 3D environments. We address this limitation by presenting JAEGER, a framework that extends AV-LLMs to 3D space, to enable joint spatial grounding and reasoning through the integration of RGB-D observations and multi-channel first-order ambisonics. A core contribution of our work is the neural intensity vector (Neural IV), a learned spatial audio representation that encodes robust directional cues to enhance direction-of-arrival estimation, even in adverse acoustic scenarios with overlapping sources. To facilitate large-scale training and systematic evaluation, we propose SpatialSceneQA, a benchmark of 61k instruction-tuning samples curated from simulated physical environments. Extensive experiments demonstrate that our approach consistently surpasses 2D-centric baselines across diverse spatial perception and reasoning tasks, underscoring the necessity of explicit 3D modelling for advancing AI in physical environments. Our source code, pre-trained model checkpoints, and datasets are available at https://github.com/liuzhan22/JAEGER. |
Accep...Accepted to ICML 2026 |
| IsoNet: Spatially-aware audio-visual target speech extraction in complex acoustic environments | 2026-05-15 | ShowTarget speech extraction remains difficult for compact devices because monaural neural models lack spatial evidence and classical beamformers lose resolving power when the microphone aperture is only a few centimetres. We present IsoNet, a user-selectable audio-visual target speech extraction system for a compact 4-microphone array. IsoNet combines complex multi-channel STFT features, GCC-PHAT spatial cues, face-conditioned visual embeddings, and auxiliary direction-of-arrival supervision inside a U-Net mask estimation network. Three curriculum variants were trained on 25,000 simulated VoxCeleb mixtures with progressively difficult SNR regimes. On a hard test set spanning -1 to 10 dB SNR, IsoNet-CL1 achieves 9.31 dB SI-SDR, a 4.85 dB improvement over the mixture, with PESQ 2.13 and STOI 0.84. Oracle delay-and-sum and MVDR beamformers degrade the same mixtures by 4.82 dB and 6.08 dB SI-SDRi, respectively, showing that the proposed learned multimodal conditioning solves a regime where conventional spatial filtering is ineffective. Ablation studies show consistent gains from visual conditioning, GCC-PHAT features, and extended delay-bin encoding. The results establish a compact-array, face-selectable speech extraction baseline under controlled simulation and identify the remaining barriers to real deployment, especially phase reconstruction, multi-interferer mixtures, and simulation-to-real transfer. |
8 pages |
| Wave Tank Experiment for Sea State Monitoring with Distributed Acoustic Sensing | 2026-04-27 | ShowMonitoring sea states across the offshore wind farm areas is essential to keep their structures safe, efficiently operate the systems, and assess the environmental effects of wind turbines. Conventional sea state sensors like buoys limit their observable coverage; therefore, installing many sensors across the wide area is necessary to obtain sufficient sea state information. However, such a situation is not practical in terms of cost. Instead, the study proposes utilising optical fibres, which is embedded in existing power cables for telecommunications on the seabed, as sea state monitoring sensors with distributed acoustic sensing (DAS). DAS is a vibration-sensing technology along optical fibres based on the Rayleigh backscattering of the injected laser. It measures the dynamic strain of the optical fibre in real time at each spatial bin, which is called a "channel" along the fibre. In power cables on the seabed, time-varying water pressure due to waves is expected to exert dynamic strain. This hypothesis motivates us to validate whether the application of DAS for power cables can estimate sea state, such as wave period, height, and the direction of arrival. Hence, the authors carried out a wave tank experiment with a programmable wave generator. An actual power cable is installed under the same condition as the bottom-mounted offshore wind turbines. The experimental results show that (i) the wave period can be accurately estimated from the frequency-domain analysis. (ii) The strong linearity between DAS vibration power and the wave height is found. (iii) The direction of arrival of waves can be estimated with the error of 1.5$^\circ$ when there are at least two laying angles of the cable in parallel with the estimation of wavelength. These outcomes promote the feasibility of utilising the existing power cables across offshore wind farms as sea state monitoring sensors. |
9 pag...9 pages, 8 figures, presented in WindEurope Annual Event 2026 |
| Interpretable Binaural Deep Beamforming Guided by Time-Varying Relative Transfer Function | 2026-02-17 | ShowIn this work, we propose a deep beamforming framework for speech enhancement in dynamic acoustic environments. The framework learns time-varying beamformer weights from noisy multichannel signals via a deep neural network, guided by a continuously tracked relative transfer function (RTF) of a moving target speaker. We analyze the network's spatial behavior on an 8-microphone linear array by evaluating narrowband and wideband beampatterns in three modes: (i) oracle guidance with true RTFs, (ii) guidance with subspace-tracked RTF estimates, and (iii) operation without RTF guidance. Results show that RTF guidance yields smoother, more spatially consistent beampatterns that track the target direction of arrival (DOA), whereas the unguided model fails to maintain a clear spatial focus. We further extend the framework to binaural beamforming for dynamic target-speaker enhancement. The system is trained using a head-related transfer function (HRTF)-based acoustic simulation of a moving source, enabling realistic spatial rendering at the left and right ears. Spatial cue preservation is quantitatively evaluated in terms of interaural level differences (ILD) and interaural time differences (ITD), demonstrating the method's suitability for hearable applications. |
|
| A framework for diffuseness evaluation using a tight-frame microphone array configuration | 2026-02-04 | ShowThis work presents a unified framework for estimating both sound-field direction and diffuseness using practical microphone arrays with different spatial configurations. Building on covariance-based diffuseness models, we formulate a velocity-only covariance approach that enables consistent diffuseness evaluation across heterogeneous array geometries without requiring mode whitening or spherical-harmonic decomposition. Three array types -- an A-format array, a rigid-sphere array, and a newly proposed tight-frame array -- are modeled and compared through both simulations and measurement-based experiments. The results show that the tight-frame configuration achieves near-isotropic directional sampling and reproduces diffuseness characteristics comparable to those of higher-order spherical arrays, while maintaining a compact physical structure. We further examine the accuracy of direction-of-arrival estimation based on acoustic intensity within the same framework. These findings connect theoretical diffuseness analysis with implementable array designs and support the development of robust, broadband methods for spatial-sound-field characterization. |
16 pa...16 pages including 16 files: This version has been substantially revised in response to reviewers' comments, with clarified theoretical assumptions and extended comparative evaluations |
| SoundCompass: Navigating Target Sound Extraction With Effective Directional Clue Integration In Complex Acoustic Scenes | 2026-01-27 | ShowRecent advances in target sound extraction (TSE) utilize directional clues derived from direction of arrival (DoA), which represent an inherent spatial property of sound available in any acoustic scene. However, previous DoA-based methods rely on hand-crafted features or discrete encodings, which lose fine-grained spatial information and limit adaptability. We propose SoundCompass, an effective directional clue integration framework centered on a Spectral Pairwise INteraction (SPIN) module that captures cross-channel spatial correlations in the complex spectrogram domain to preserve full spatial information in multichannel signals. The input feature expressed in terms of spatial correlations is fused with a DoA clue represented as spherical harmonics (SH) encoding. The fusion is carried out across overlapping frequency subbands, inheriting the benefits reported in the previous band-split architectures. We also incorporate the iterative refinement strategy, chain-of-inference (CoI), in the TSE framework, which recursively fuses DoA with sound event activation estimated from the previous inference stage. Experiments demonstrate that SoundCompass, combining SPIN, SH embedding, and CoI, robustly extracts target sources across diverse signal classes and spatial configurations. |
5 pag...5 pages, 4 figures, accepted to ICASSP 2026 |
| Vector Signal Reconstruction Sparse and Parametric Approach of direction of arrival Using Single Vector Hydrophone | 2025-12-25 | ShowThis article discusses the application of single vector hydrophones in the field of underwater acoustic signal processing for Direction Of Arrival (DOA) estimation. Addressing the limitations of traditional DOA estimation methods in multi-source environments and under noise interference, this study introduces a Vector Signal Reconstruction Sparse and Parametric Approach (VSRSPA). This method involves reconstructing the signal model of a single vector hydrophone, converting its covariance matrix into a Toeplitz structure suitable for the Sparse and Parametric Approach (SPA) algorithm. The process then optimizes it using the SPA algorithm to achieve more accurate DOA estimation. Through detailed simulation analysis, this research has confirmed the performance of the proposed algorithm in single and dual-target DOA estimation scenarios, especially under various signal-to-noise ratio(SNR) conditions. The simulation results show that, compared to traditional DOA estimation methods, this algorithm has significant advantages in estimation accuracy and resolution, particularly in multi-source signals and low SNR environments. The contribution of this study lies in providing an effective new method for DOA estimation with single vector hydrophones in complex environments, introducing new research directions and solutions in the field of vector hydrophone signal processing. |
The a...The authors have determined that the simulation results presented are preliminary and insufficient. Further simulation work is required to validate the conclusions. The text also requires major linguistic improvements |
| DOA Estimation with Lightweight Network on LLM-Aided Simulated Acoustic Scenes | 2025-11-11 | ShowDirection-of-Arrival (DOA) estimation is critical in spatial audio and acoustic signal processing, with wide-ranging applications in real-world. Most existing DOA models are trained on synthetic data by convolving clean speech with room impulse responses (RIRs), which limits their generalizability due to constrained acoustic diversity. In this paper, we revisit DOA estimation using a recently introduced dataset constructed with the assistance of large language models (LLMs), which provides more realistic and diverse spatial audio scenes. We benchmark several representative neural-based DOA methods on this dataset and propose LightDOA, a lightweight DOA estimation model based on depthwise separable convolutions, specifically designed for mutil-channel input in varying environments. Experimental results show that LightDOA achieves satisfactory accuracy and robustness across various acoustic scenes while maintaining low computational complexity. This study not only highlights the potential of spatial audio synthesized with the assistance of LLMs in advancing robust and efficient DOA estimation research, but also highlights LightDOA as efficient solution for resource-constrained applications. |
|
| Consensus Tracking of an Underwater Vehicle Using Weighted Harmonic Mean Density | 2025-11-05 | ShowThis paper addresses an underwater target tracking problem in which a large number of sonobuoy sensors are deployed on a surveillance region. The region is divided into several sub-regions, where a single tracker, capable of generating track is installed. Each sonobuoy can measure the direction of arrival of acoustic signals (known as bearing angles) and communicate the measurements with the local tracker. Further, each local tracker can communicate with all other trackers, where each of them can exchange their estimate and finally a consensus is reached. We propose a weighted harmonic mean density (HMD) based tracking to reach a consensus and provide a solution for the fusion of Gaussian densities. In this approach, optimal weights are assigned by minimizing the Kullback-Leibler divergence measure. Performance of the proposed method is measured using root mean square error, percentage of track divergence, and normalized estimation error squared. Simulation results demonstrate that the optimized HMD-based fusion outperforms existing fusion methods during a distributed tracking. |
|
| State Space and Self-Attention Collaborative Network with Feature Aggregation for DOA Estimation | 2025-10-29 | ShowAccurate direction-of-arrival (DOA) estimation for sound sources is challenging due to the continuous changes in acoustic characteristics across time and frequency. In such scenarios, accurate localization relies on the ability to aggregate relevant features and model temporal dependencies effectively. In time series modeling, achieving a balance between model performance and computational efficiency remains a significant challenge. To address this, we propose FA-Stateformer, a state space and self-attention collaborative network with feature aggregation. The proposed network first employs a feature aggregation module to enhance informative features across both temporal and spectral dimensions. This is followed by a lightweight Conformer architecture inspired by the squeeze-and-excitation mechanism, where the feedforward layers are compressed to reduce redundancy and parameter overhead. Additionally, a temporal shift mechanism is incorporated to expand the receptive field of convolutional layers while maintaining a compact kernel size. To further enhance sequence modeling capabilities, a bidirectional Mamba module is introduced, enabling efficient state-space-based representation of temporal dependencies in both forward and backward directions. The remaining self-attention layers are combined with the Mamba blocks, forming a collaborative modeling framework that achieves a balance between representation capacity and computational efficiency. Extensive experiments demonstrate that FA-Stateformer achieves superior performance and efficiency compared to conventional architectures. |
|
| Perceptual Compensation of Ambisonics Recordings for Reproduction in Room | 2025-10-13 | ShowAmbisonics is a method for capturing and rendering a sound field accurately, assuming that the acoustics of the playback room does not significantly influence the sound field. However, in practice, the acoustics of the playback room may lead to a noticeable degradation in sound quality. We propose a recording and rendering method based on Ambisonics that utilizes a perceptually-motivated approach to compensate for the reverberation of the playback room. The recorded direct and reverberant sound field components in the spherical harmonics (SHs) domain are spectrally and spatially compensated to preserve the relevant auditory cues including the direction of arrival of the direct sound, the spectral energy of the direct and reverberant sound components, and the Interaural Coherence (IC) across each auditory band. In contrast to the conventional Ambisonics, a flexible number of Ambisonics channels can be used for audio rendering. Listening test results show that the proposed method provides a perceptually accurate rendering of the originally recorded sound field, outperforming both conventional Ambisonics without compensation and even ideal Ambisonics rendering in a simulated anechoic room. Additionally, subjective evaluations of listeners seated at the center of the loudspeaker array demonstrate that the method remains robust to head rotation and minor displacements. |
The m...The manuscript was submitted to the JASA and is under review |
| OWL: Geometry-Aware Spatial Reasoning for Audio Large Language Models | 2025-09-30 | ShowSpatial reasoning is fundamental to auditory perception, yet current audio large language models (ALLMs) largely rely on unstructured binaural cues and single step inference. This limits both perceptual accuracy in direction and distance estimation and the capacity for interpretable reasoning. Recent work such as BAT demonstrates spatial QA with binaural audio, but its reliance on coarse categorical labels (left, right, up, down) and the absence of explicit geometric supervision constrain resolution and robustness. We introduce the |
| Title | Date | Abstract | Comment |
|---|---|---|---|
| Direction of arrival estimation from distant microphone data using single frequency filtering | 2026-06-15 | ShowIn distant microphones, broadband (BB) methods for direction-of-arrival (DoA) estimation are more suitable than narrowband (NB) methods. Due to the aggregation of their optimization function across all frequency bands, BB estimators are robust to spatial aliasing, a known problem in processing distant microphone data. In NB methods, DoA estimation is performed by utilizing \textit{local} information in each frequency band and hence the estimation is affected by spatial aliasing. However, unlike BB methods, NB methods exploit frequency sparsity to estimate the DoAs of \textit{multiple speakers} in a \textit{single time frame}. In this article, a method to improve the robustness of a NB DoA estimator to spatial aliasing is developed. The proposed method is based on cross-correlation of speech-present time-frequency regions obtained by single frequency filtering (SFF) of the microphone signals. The SFF spectrum is chosen because SFF components have regions of high signal-to-noise ratio both in time and frequency and because speech and non-speech discrimination is robust to degradations in the SFF domain. The proposed NB estimator is compared to four state-of-the-art estimators (one NB and three BB) using detection and accuracy metrics on simulated and real-world data in different reverberation and noise conditions. The results show that in all the environments, the SFF-based NB approach outperforms the state-of-the-art NB approach. Furthermore, the performance of the SFF-based approach is better than some of the BB estimators. |
|
| A framework for diffuseness evaluation using a tight-frame microphone array configuration | 2026-02-04 | ShowThis work presents a unified framework for estimating both sound-field direction and diffuseness using practical microphone arrays with different spatial configurations. Building on covariance-based diffuseness models, we formulate a velocity-only covariance approach that enables consistent diffuseness evaluation across heterogeneous array geometries without requiring mode whitening or spherical-harmonic decomposition. Three array types -- an A-format array, a rigid-sphere array, and a newly proposed tight-frame array -- are modeled and compared through both simulations and measurement-based experiments. The results show that the tight-frame configuration achieves near-isotropic directional sampling and reproduces diffuseness characteristics comparable to those of higher-order spherical arrays, while maintaining a compact physical structure. We further examine the accuracy of direction-of-arrival estimation based on acoustic intensity within the same framework. These findings connect theoretical diffuseness analysis with implementable array designs and support the development of robust, broadband methods for spatial-sound-field characterization. |
16 pa...16 pages including 16 files: This version has been substantially revised in response to reviewers' comments, with clarified theoretical assumptions and extended comparative evaluations |
| Ambiguity-Free Broadband DOA Estimation Relying on Parameterized Time-Frequency Transform | 2025-03-05 | ShowAn ambiguity-free direction-of-arrival (DOA) estimation scheme is proposed for sparse uniform linear arrays under low signal-to-noise ratios (SNRs) and non-stationary broadband signals. First, for achieving better DOA estimation performance at low SNRs while using non-stationary signals compared to the conventional frequency-difference (FD) paradigms, we propose parameterized time-frequency transform-based FD processing. Then, the unambiguous compressive FD beamforming is conceived to compensate the resolution loss induced by difference operation. Finally, we further derive a coarse-to-fine histogram statistics scheme to alleviate the perturbation in compressive FD beamforming with good DOA estimation accuracy. Simulation results demonstrate the superior performance of our proposed algorithm regarding robustness, resolution, and DOA estimation accuracy. |
6 figures |
| Comparison of Frequency-Fusion Mechanisms for Binaural Direction-of-Arrival Estimation for Multiple Speakers | 2024-01-15 | ShowTo estimate the direction of arrival (DOA) of multiple speakers with methods that use prototype transfer functions, frequency-dependent spatial spectra (SPS) are usually constructed. To make the DOA estimation robust, SPS from different frequencies can be combined. According to how the SPS are combined, frequency fusion mechanisms are categorized into narrowband, broadband, or speaker-grouped, where the latter mechanism requires a speaker-wise grouping of frequencies. For a binaural hearing aid setup, in this paper we propose an interaural time difference (ITD)-based speaker-grouped frequency fusion mechanism. By exploiting the DOA dependence of ITDs, frequencies can be grouped according to a common ITD and be used for DOA estimation of the respective speaker. We apply the proposed ITD-based speaker-grouped frequency fusion mechanism for different DOA estimation methods, namely the multiple signal classification, steered response power and a recently published method based on relative transfer function (RTF) vectors. In our experiments, we compare DOA estimation with different fusion mechanisms. For all considered DOA estimation methods, the proposed ITD-based speaker-grouped frequency fusion mechanism results in a higher DOA estimation accuracy compared with the narrowband and broadband fusion mechanisms. |
Accep...Accepted for ICASSP 2024 |
| Gridless DOA Estimation with Multiple Frequencies | 2023-02-06 | ShowDirection-of-arrival (DOA) estimation is widely applied in acoustic source localization. A multi-frequency model is suitable for characterizing the broadband structure in acoustic signals. In this paper, the continuous (gridless) DOA estimation problem with multiple frequencies is considered. This problem is formulated as an atomic norm minimization (ANM) problem. The ANM problem is equivalent to a semi-definite program (SDP) which can be solved by an off-the-shelf SDP solver. The dual certificate condition is provided to certify the optimality of the SDP solution so that the sources can be localized by finding the roots of a polynomial. We also construct the dual polynomial to satisfy the dual certificate condition and show that such a construction exists when the source amplitude has a uniform magnitude. In multi-frequency ANM, spatial aliasing of DOAs at higher frequencies can cause challenges. We discuss this issue extensively and propose a robust solution to combat aliasing. Numerical results support our theoretical findings and demonstrate the effectiveness of the proposed method. |
This ...This work has been accepted by IEEE Transactions on Signal Processing |
| DA-MUSIC: Data-Driven DoA Estimation via Deep Augmented MUSIC Algorithm | 2023-01-11 | ShowDirection of arrival (DoA) estimation of multiple signals is pivotal in sensor array signal processing. A popular multi-signal DoA estimation method is the multiple signal classification (MUSIC) algorithm, which enables high-performance super-resolution DoA recovery while being highly applicable in practice. MUSIC is a model-based algorithm, relying on an accurate mathematical description of the relationship between the signals and the measurements and assumptions on the signals themselves (non-coherent, narrowband sources). As such, it is sensitive to model imperfections. In this work we propose to overcome these limitations of MUSIC by augmenting the algorithm with specifically designed neural architectures. Our proposed deep augmented MUSIC (DA-MUSIC) algorithm is thus a hybrid model-based/data-driven DoA estimator, which leverages data to improve performance and robustness while preserving the interpretable flow of the classic method. DA-MUSIC is shown to learn to overcome limitations of the purely model-based method, such as its inability to successfully localize coherent sources as well as estimate the number of coherent signal sources present. We further demonstrate the superior resolution of the DA-MUSIC algorithm in synthetic narrowband and broadband scenarios as well as with real-world data of DoA estimation from seismic signals. |
Submitted to TVT |
| Wideband Modal Orthogonality: A New Approach for Broadband DOA Estimation | 2020-06-12 | ShowWideband direction of arrival (DOA) estimation techniques for sensors array have been studied extensively in the literature. Nevertheless, needing prior information on the number and directions of sources or demanding heavy computational load makes most of these techniques less useful in practice. In this paper, a low complexity subspace-based framework for DOA estimation of broadband signals, named as wideband modal orthogonality (WIMO), is proposed and accordingly two DOA estimators are developed. First, a closed-form approximation of spatial-temporal covariance matrix (STCM) in the uniform spectrum case is presented. The eigenvectors of STCM associated with non-zero eigenvalues are modal components of the wideband source in a given bandwidth and direction. WIMO idea is to extract these eigenvectors at desired DOAs from the approximated STCM and test their orthogonality to estimated noise subspace. In the non-uniform spectrum case, WIMO idea can be applied by approximating STCM through numerical integration. Fortunately, STCM approximation and modal extraction can be performed offline. WIMO provides DOA estimation without the conventional prerequisites, such as spectral decomposition, focusing procedure and, a priori information on the number of sources and their DOAs. Several numerical examples are conducted to compare the WIMO performance with the state-of-the-art methods. Simulations demonstrate that the two proposed DOA estimators achieve superior performance in terms of probability of resolution and estimation error along with orders of magnitude runtime speedup. |
|
| Broadband Sparse Array Focusing Via Spatial Periodogram Averaging and Correlation Resampling | 2019-12-24 | ShowThis paper proposes two coherent broadband focusing algorithms for spatial correlation estimation using sparse linear arrays. Both algorithms decompose the time-domain array data into disjoint frequency bands through discrete Fourier transform or filter banks to obtain broadband frequency-domain snapshots. The periodogram averaging (AP) algorithm starts in the frequency domain by estimating the broadband spatial periodograms for all bands and then averaging them to reinforce the sources' spatial spectral information. Taking inverse spatial Fourier transform of the combined spatial periodogram estimates the focused spatial correlations. Alternatively, the spatial correlation resampling (SCR) algorithm directly computes the spatial correlations for each band and then rescales the spatial sampling rate to align at a focused frequency. The resampled spatial correlations from all frequency bands are then averaged to estimate the focused spatial correlations. The spatial correlations estimated from the AP or SCR algorithms populate the diagonals of a Hermitian Toeplitz augmented covariance matrix (ACM). The focused ACM is the input of a new minimum description length (MDL) based criteria, termed MDL-gap, for source enumeration and the standard narrowband MUSIC algorithm for DOA estimation. Numerical simulations show that both the AP and SCR algorithms improve source enumeration and DOA estimation performances over the incoherent subspace focusing algorithm in snapshot limited scenarios. |
11 pages, 8 figures |
| Broadband DOA estimation using Convolutional neural networks trained with noise signals | 2017-12-12 | ShowA convolution neural network (CNN) based classification method for broadband DOA estimation is proposed, where the phase component of the short-time Fourier transform coefficients of the received microphone signals are directly fed into the CNN and the features required for DOA estimation are learnt during training. Since only the phase component of the input is used, the CNN can be trained with synthesized noise signals, thereby making the preparation of the training data set easier compared to using speech signals. Through experimental evaluation, the ability of the proposed noise trained CNN framework to generalize to speech sources is demonstrated. In addition, the robustness of the system to noise, small perturbations in microphone positions, as well as its ability to adapt to different acoustic conditions is investigated using experiments with simulated and real data. |
Publi...Published in Proceedings of IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) 2017 |
| Convolutional Neural Networks for Passive Monitoring of a Shallow Water Environment using a Single Sensor | 2016-12-12 | ShowA cost effective approach to remote monitoring of protected areas such as marine reserves and restricted naval waters is to use passive sonar to detect, classify, localize, and track marine vessel activity (including small boats and autonomous underwater vehicles). Cepstral analysis of underwater acoustic data enables the time delay between the direct path arrival and the first multipath arrival to be measured, which in turn enables estimation of the instantaneous range of the source (a small boat). However, this conventional method is limited to ranges where the Lloyd's mirror effect (interference pattern formed between the direct and first multipath arrivals) is discernible. This paper proposes the use of convolutional neural networks (CNNs) for the joint detection and ranging of broadband acoustic noise sources such as marine vessels in conjunction with a data augmentation approach for improving network performance in varied signal-to-noise ratio (SNR) situations. Performance is compared with a conventional passive sonar ranging method for monitoring marine vessel activity using real data from a single hydrophone mounted above the sea floor. It is shown that CNNs operating on cepstrum data are able to detect the presence and estimate the range of transiting vessels at greater distances than the conventional method. |
Final...Final draft for IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2017. 5 pages, 4 figures |
| Title | Date | Abstract | Comment |
|---|---|---|---|
| Deterministic Maximum Likelihood Direction Finding in the Mixture Noise of Gaussian and Spherically Invariant Components | 2026-08-13 | ShowSpherically invariant (SI) random processes can model impulsive noise and unreliable measurements. Recently, the mixture noise of Gaussian and SI components has been used in deterministic maximum likelihood direction finding. In this context, the Expectation-Conditional Maximization (ECM) algorithm, an extension of the expectation-maximization algorithm, has been applied and designed. However, simulation results show that the ECM algorithm always improperly converges. In this article, the ECM Either (ECME) algorithm, an extension of the ECM algorithm, is applied and designed, which additionally utilizes the actual log-likelihood function to first update partial parameter estimates at every iteration and does not need to initialize all parameter estimates. Moreover, the deterministic Cramer-Rao low bounds (CRLBs) of DOA estimators are derived and compared. Simulation results indicate that the ECME algorithm exhibits proper convergence and its root mean square errors of DOA estimates asymptotically approach the CRLBs as the signal powers increase, i.e., the derived CRLBs are correct. |
|
| Lightweight Single-Antenna Direction-of-Arrival Estimation for Curvilinear Trajectories in Mobile Embedded Systems | 2026-08-12 | ShowAccurate direction-of-arrival (DOA) estimation is valuable for spatially selective communication in noisy industrial environments. This work investigates a lightweight single-antenna framework in which receiver motion forms a virtual aperture. The receiver uses onboard inertial measurement unit (IMU) headings and two-way-ranging (TWR) measurements to a known fixed beacon, avoiding GPS, optical tracking, and high-precision external tracking of the mobile receiver. The curvilinear virtual-array MUSIC formulation is evaluated numerically using phase-coherent narrowband snapshots over arbitrary trajectories. Hardware experiments validate a range- domain TWR-IMU bearing estimator using corrected and averaged ranging observations. In a campaign of 100 consecutive four-revolution sweeps, all trials are retained. Phase-aligned accumulation recovers a 123 mm range-modulation amplitude, in close agreement with the measured 120 mm antenna lever arm. The single-sweep bearing precision is 8.2° absolute world-frame accuracy is limited by systematic BNO055 magnetometer drift in the motorized setup. The embedded bearing-estimation computation consumes 144.9 mJ per estimate, approximately 3 % of the measured cycle energy in the tested configuration. These results establish the feasibility of onboard range-domain bearing estimation and motivate future experimental validation of phase-coherent MUSIC during free-form mobile trajectories. |
|
| Small Language Model enabled Autonomous agent for Language-Conditioned Cognitive Radar | 2026-08-12 | ShowModern radar systems require adapting their processing strategies in response to changing interference, clutter, and data availability. This paper introduces a framework for a small language model (SLM)-driven autonomous agent designed for language-conditioned cognitive radar, functioning as an intelligent controller for a suite of array signal processing tools. Given a natural-language command, the agent extracts radar-operation-related cues, selects an appropriate sequence of signal-processing methods, configures parameters, and invokes executable tools for numerical computation. Experiments with a synthetic uniform linear array (ULA) radar demonstrate that, given a natural-language command, the agent performs meaningful algorithm selection across diverse scenarios for sidelobe control, jammer suppression, multiple-null beamforming, coherent-source handling, and low-snapshot direction-of-arrival (DOA) estimation. Ablation results show that radar-specific prompting and physics-grounded tool execution are both required for reliable decisions and hallucination-free numerical results. |
Accep...Accepted at MLSP 2026, ATL, USA |
| BiTSE: Binaural Target Speaker Extraction in Noisy Multi-Talker Environments for AR Glass Arrays | 2026-08-10 | ShowIsolating a desired speech signal in noisy multi-talker conversational scenarios is a key requirement for augmented reality (AR) wearable microphone array systems. In this work, a binaural target speaker extraction (TSE) framework, termed BiTSE, is proposed. It leverages both spatial and temporal cues, specifically the direction-of-arrival (DoA) of the target speaker and corresponding voice activity information, to guide the extraction process. Built upon a binaural signal denoising architecture, our model integrates three key enhancements: (i) a DoA-aware attention mechanism using cyclic positional embeddings, (ii) a timestamp-based masking strategy that utilizes speaker activity to suppress non-target segments, and (iii) a novel two-stage loss optimization strategy that first trains the model for robust denoising and then fine-tunes it to improve perceptual quality. Evaluations on the SPeech Enhancement for Augmented Reality (SPEAR) challenge dataset demonstrate that the proposed BiTSE consistently improves upon conventional approaches, leading to enhanced signal fidelity and perceptual quality. |
This ...This is the preprint version of the paper accepted at APSIPA ASC 2026 |
| Neural Array-Generic Direction-of-Arrival Estimation Exploiting Array Transfer Functions | 2026-08-10 | ShowDirection-of-arrival (DoA) estimation is a key component of multichannel audio processing, yet many deep learning approaches remain tied to the microphone arrays used during training and generalize poorly to unseen devices. This paper proposes an array-generic neural DoA estimation framework using measured or simulated complex directional array transfer functions (ATFs) matched to real-world multi-microphone devices. The method processes multichannel spectrograms and ATF metadata with separate convolutional encoders, fuses the resulting representations through cross-attention, and predicts source directions using a multi-source Cartesian vector output formulation. Experiments on simulated 2D and 3D localization tasks under reverberation and diffuse babble noise show that the proposed approach generalizes to previously unseen arrays, including mobile-phone-like configurations, without major performance degradation, while remaining competitive with conventional and learning-based baselines. |
Accep...Accepted for publication in IWAENC 2026 |
| Multi-Source Position and Direction-of-Arrival Estimation Based on Euclidean Distance Matrices | 2026-08-10 | ShowA popular method to estimate the positions or directions-of-arrival (DOAs) of multiple sound sources using an array of microphones is based on steered-response power (SRP) beamforming. For a three-dimensional scenario, SRP-based methods require joint optimization of three continuous variables for position estimation or two continuous variables for DOA estimation, which can be computationally expensive when high localization accuracy is desired. In this paper, we propose novel methods for multi-source position and DOA estimation by exploiting properties of Euclidean distance matrices (EDMs) and their respective Gram matrices. All methods require estimated time-differences of arrival (TDOAs) between the microphones. In the proposed multi-source position estimation method, only a single continuous variable per source, representing the distance to a reference microphone, needs to be optimized. For each source, the optimal distance variable and set of candidate TDOA estimates are determined by minimizing a cost function defined using the eigenvalues of the Gram matrix. The estimated relative source positions are then mapped to absolute source positions by solving an orthogonal Procrustes problem. The proposed multi-source DOA estimation method eliminates the need for continuous variable optimization. The optimal set of candidate TDOA estimates is determined by minimizing a cost function defined using the eigenvalues of a rank-reduced Gram matrix. For two sources in a noisy and reverberant environment, experimental results for different source and microphone configurations with six microphones show that the proposed EDM-based method consistently outperforms the SRP-based method in terms of position and DOA estimation accuracy and run time. |
16 pa...16 pages, 7 figures, accepted for publication in IEEE Transactions on Audio, Speech and Language Processing |
| Rotatable Antenna-Enhanced Wireless Sensing with Uniform Sparse Array via Tensor Decomposition | 2026-08-09 | ShowIn this letter, we propose a new wireless sensing system equipped with a rotatable antenna (RA) array to enhance the sensing performance of a uniform sparse array (USA). To tackle the severe spatial undersampling issues, we propose a novel tensor decomposition-based direction-of-arrival (DOA) estimation algorithm. Specifically, we introduce a synchronous multiple rotation pattern for active target probing such that the received signals across multiple rotations to capture the diverse spatial degree of freedoms. Subsequently, we mathematically formulate the received signals across successive rotations as a third-order tensor, and leverage the canonical polyadic decomposition to obtain the factor matrices incorporating the DOA of targets. By analyzing the extrema distribution laws of array steering vector correlation (SVC) and gain SVC of RAs, we propose to combine the array and gain factor matrices via the Kronecker product, which theoretically guarantees the unambiguous DOA estimation. Simulation results demonstrate that the proposed RA-enhanced tensor decomposition-based algorithm achieves high-precision and unambiguous sensing performance compared to conventional uniform dense arrays and omnidirectional antenna systems. |
|
| Tensor-Based Joint Pitch and DOA Estimation | 2026-07-31 | ShowWe propose a tensor-based subspace method for the joint estimation of pitch and direction of arrival (DOA) of multiple harmonic sources. Unlike conventional matrix-based approaches that unfold the spatio-temporal data into a single matrix, the proposed method preserves the multidimensional Hankel tensor structure and exploits a Kronecker product constraint satisfied by the true signal subspace. The matrix-based subspace estimate is refined by projecting it onto an empirically estimated Kronecker-structured signal cage obtained from mode-unfolded tensor subspaces. Beyond the estimator itself, we provide a non-asymptotic theory explaining when this tensor refinement improves the matrix baseline. We show that the oracle refinement using the true Kronecker projector is non-worsening, decompose the overall tensor gain into oracle gain and empirical loss, and derive deterministic and high-probability sufficient conditions for positive overall tensor gain under complex Gaussian noise. The analysis reveals a sweet-spot behavior: the largest certified gain occurs when the matrix estimate is neither nearly perfect nor too inaccurate. Simulations over closely spaced, harmonically related, and coherent-source scenarios show that the proposed method improves subspace accuracy and downstream pitch/DOA estimation, and also outperforms representative matrix and tensor-train-based baselines. |
|
| V-RIS: Virtual-Aperture DoA Estimation with Sparse RIS | 2026-07-30 | ShowLarge-aperture reconfigurable intelligent surfaces (RISs) enable high-resolution 2D direction-of-arrival (DoA) estimation, but existing approaches still tie hardware cost and control overhead to aperture size. To decouple the effective sensing aperture from the number of physically deployed RIS elements, we present V-RIS, a framework for virtual-aperture surface-field reconstruction and DoA estimation. V-RIS uses only four corner subarrays and a single-antenna receiver to reconstruct the virtual-aperture surface field from receiver observations collected under multiple RIS phase configurations, and then performs DoA estimation on the reconstructed virtual-aperture surface field. Our key observation is that, under far-field illumination, the discretized RIS surface field satisfies finite-order spatial recurrences along both aperture axes. We enforce data-level consistency through the RIS-coded receiver observations and propagation consistency through the far-field spatial recurrence, while using a four-corner deployment geometry that retains both contiguous local elements and long aperture baselines. To improve robustness in practical receiver observations, we adopt a bias-invariant receiver-domain loss that suppresses quasi-static hardware distortions and configuration-invariant multipath contributions. Extensive simulations show that V-RIS approaches the DoA accuracy of a full-aperture benchmark while producing cleaner spectra than matrix-completion and least-squares baselines. An outdoor prototype further validates the design: with only 25% programmable elements, V-RIS keeps both elevation and azimuth errors within |
|
| Spatial Angular Pseudo-Derivative Search Algorithm: A Real-Time Single-Snapshot Super-Resolution Sparse DOA Scheme for Automotive Radar | 2026-07-27 | ShowAccurate, high-resolution, and real-time DOA estimation plays a crucial role in automotive radar perception. While sparse signal recovery techniques offer super-resolution and high-precision estimation, their prohibitive computational complexity remains a primary bottleneck for practical deployment. This paper proposes a sparse DOA estimation scheme specifically tailored for the stringent requirements of automotive radar such as limited computational resources, restricted array apertures, and single-snapshot constraints. By leveraging the spatial angular pseudo-derivative (SAPD) property of the sparse DOA solutions and incorporating this property as a constraint into an |
|
| Electromagnetic Neural Network for Direction-of-Arrival Estimation | 2026-07-25 | ShowAccurate and real-time direction of arrival (DOA) estimation is crucial for beamforming in unmanned aerial vehicle (UAV) communication systems. However, the existing high-precision DOA estimation algorithms encounter high computational complexity when implemented on a UAV with on-board signal processing constraints. To tackle this issue, an electromagnetic neural network (EMNN) is developed for DOA estimation, which is capable of generating the angular spectrum of the incident signal based solely on amplitude observation. Specifically, the proposed EMNN consists of two components: a stacked intelligent metasurfaces (SIM) is mounted on the UAV, and each meta-atom is an artificial neuron that can process signals in the electromagnetic domain with low energy consumption and ultra-fast computing speed. Furthermore, a fully connected layer is cascaded to process the received amplitude signal, enhancing the non-linear extraction and representational ability of EMNN. Moreover, to reduce the computational complexity and observation snapshots required for high-resolution DOA estimation, we develop a hierarchical DOA estimation framework, which involves two stages for conducting coarse and fine DOA estimation, respectively. For each stage, EMNN is trained on randomly generated training samples and their corresponding spectra to achieve the desired estimation goal. Finally, the simulation results validate that the proposed EMNN achieves approximately 13 dB gain in classification error reduction over the conventional beamforming (CBF) method in dual-signal scenarios, albeit its lower cost and radio frequency (RF)-related power consumption. |
16 pa...16 pages, 13 figures, 4 tables, accepted by IEEE TWC |
| Hankel and Toeplitz Rank-1 Decomposition of Arbitrary Matrices with Applications to Signal Direction-of-Arrival Estimation | 2026-07-22 | ShowWe consider the problems of computing the optimal rank-1 Hankel and Toeplitz-structured approximation of arbitrary matrices under L2 and L1-norm error. Such problems arise naturally in engineered systems, including the basic few-shot signal Direction-of-Arrival (DoA) estimation problem that is of importance to modern autonomous systems applications. We develop accurate and computationally efficient structured matrix decomposition algorithms for both formulations and then derive analytically grounded small-sample-support DoA estimators for practical sensing system deployments. The resulting estimators under the L2 and L1 norms are formally shown to be maximum-likelihood optimal under white Gaussian and Laplace noise, respectively. The estimators are further validated through extensive simulation studies and real-world data experiments in few-shot DoA inference. |
|
| Circulant ADMM-Net for Fast High-resolution DoA Estimation | 2026-07-21 | ShowThis paper introduces CADMM-Net and CHADMM-Net, two deep neural networks for direction of arrival estimation within the least-absolute shrinkage and selection operator (LASSO) framework. These two networks are based on a structured deep unfolding of the alternating direction method of multipliers (ADMM) algorithm through the use of circulant as well as Hermitian-circulant matrices. Along with a computational complexity of |
Updat...Updated references, fixed typos, and some figures were updated with a new baseline |
| Acoustic Imaging for UAV Detection: Dense Beamformed Energy Maps and U-Net SELD | 2026-07-17 | ShowWe introduce a U-net model for 360° acoustic source localization formulated as a spherical semantic segmentation task. Rather than regressing discrete direction-of-arrival (DoA) angles, our model segments beamformed audio maps (azimuth & elevation) into regions of active sound presence. Using delay-and-sum (DAS) beamforming on a custom 24-microphone array, we generate signals aligned with drone GPS telemetry to create binary supervision masks. A modified U-Net, trained on frequency-domain representations of these maps, learns to identify spatially distributed source regions while addressing class imbalance via the Tversky loss. Because the network operates on beamformed energy maps, the approach is inherently array-independent and can adapt to different microphone configurations and can be transferred to different microphone configurations with minimal adaptation. The segmentation outputs are post-processed by computing centroids over activated regions, enabling robust DoA estimates. Our dataset includes real-world open-field recordings of a DJI Air 3 drone, synchronized with 360° video and flight logs across multiple dates and locations. Experimental results show that U-net generalizes across environments, providing improved angular precision, offering a new paradigm for dense spatial audio understanding beyond traditional Sound Source Localization (SSL). We additionally validate the same beamforming-plus-segmentation formulation on the DCASE 2019 TAU Spatial Sound Events benchmark, showing that the approach generalizes beyond drone acoustics to multiclass Sound Event Localization and Detection (SELD) scenarios. |
|
| Parametric Diffraction-Based Object Sensing: Modeling, Estimation, and Fundamental Limits | 2026-07-15 | ShowThis paper proposes a rigorous framework for sensing of environmental objects using diffraction mechanisms prevalent at wireless communication frequencies. Specifically, we develop a physics-consistent parameterized diffraction channel model, derive maximum likelihood (ML) approaches for estimating the blockage shape, range, and source directions of arrival (DoAs), and quantify fundamental performance limits via the Cramér--Rao bound (CRB). In our physics-based modeling, we integrate various approximations for the wave propagation (far-field, paraxial Fresnel, and exact near-field regimes), enabling a wide range of applicability. The underlying model is frequency-agnostic, and we derive Fresnel-number scaling laws that map the diffraction pattern, and hence the estimation problem, across carrier frequency, object size, and range. We quantify the maximum likelihood estimation performance and its relationship to the CRB, and we study the impact of the modeling approximations developed in this work. Numerical results demonstrate that ML estimators closely approach the CRB at moderate to high signal-to-noise ratio (SNR), and highlight the utility of diffraction-based modeling for high-fidelity blockage characterization. |
|
| DOA Estimation from One-Bit Magnitude-Only Measurements via Sign-Consistency Optimization | 2026-07-14 | ShowThe direction-of-arrival (DOA) estimation problem using one-bit quantized magnitude-only measurements is studied, where magnitude-only measurements offer robustness against phase errors, thereby avoiding the need for array calibration, while one-bit quantization significantly reduces hardware cost and system complexity. As their direct combination results in meaningless constant measurements, we formulate a sign-consistency optimization problem using a smooth logistic surrogate with ell_2,1-norm regularization to promote joint sparsity. To solve this problem, a proximal-gradient algorithm is developed with guaranteed convergence to a critical point. Numerical results demonstrate that the proposed method achieves accuracy comparable to coherent one-bit baselines under ideal conditions, while maintaining robust performance under severe phase errors that substantially impair coherent methods. |
12 pa...12 pages, 9 figures. Submitted for possible publication |