Abstract:This letter investigates transmit-power minimization for multiuser pinching-antenna system (PAS) from a coupled-mode-theory (CMT)-aware perspective. Existing CMT-based pinching antenna (PA) studies reveal coupling-induced power exchange and radiation behavior, but these effects have not been fully embedded into system-level multi-PA channel modeling and beamforming design. We therefore develop a directional and loss-aware channel model that captures coupling-length-dependent power extraction and the downstream guided-power reduction caused by in-waveguide attenuation and upstream extraction. The model shows that PA design should account for both directional radiation and guided-power evolution, rather than only propagation distance or maximum coupling considered in most existing works. Based on this channel model, we formulate a quality-of-service (QoS)-constrained power minimization problem for continuous PA positioning and finite-codebook activation. For each candidate coupling length, the element-wise positioning and BPSO-based activation use a closed-form zero-forcing (ZF) power metric for low-complexity configuration ranking, thereby avoiding repeated beamforming optimization while excluding rank-deficient candidates and ordering the remaining ones. The selected configuration for each coupling length is then evaluated by optimal fixed-configuration QoS beamforming. Simulation results demonstrate that CMT-aware modeling fundamentally reshapes the preferred PA configuration, maximum coupling is not always power-efficient due to suppressed downstream PA contributions, and finite-codebook activation combined with ZF-based ranking provides a balance between transmit-power performance and deployment complexity.
Abstract:Large-aperture reconfigurable intelligent surfaces (RISs) enable high-resolution 2D direction-of-arrival (DoA) estimation, but existing approaches still tie hardware cost and control overhead to aperture size. To decouple the effective sensing aperture from the number of physically deployed RIS elements, we present V-RIS, a framework for virtual-aperture surface-field reconstruction and DoA estimation. V-RIS uses only four corner subarrays and a single-antenna receiver to reconstruct the virtual-aperture surface field from receiver observations collected under multiple RIS phase configurations, and then performs DoA estimation on the reconstructed virtual-aperture surface field. Our key observation is that, under far-field illumination, the discretized RIS surface field satisfies finite-order spatial recurrences along both aperture axes. We enforce data-level consistency through the RIS-coded receiver observations and propagation consistency through the far-field spatial recurrence, while using a four-corner deployment geometry that retains both contiguous local elements and long aperture baselines. To improve robustness in practical receiver observations, we adopt a bias-invariant receiver-domain loss that suppresses quasi-static hardware distortions and configuration-invariant multipath contributions. Extensive simulations show that V-RIS approaches the DoA accuracy of a full-aperture benchmark while producing cleaner spectra than matrix-completion and least-squares baselines. An outdoor prototype further validates the design: with only 25\% programmable elements, V-RIS keeps both elevation and azimuth errors within $1^\circ$ of the ground truth.
Abstract:We present FFAvatar, a Transformer-based 3D Gaussian framework for fast construction of high-quality and animatable 4D head avatars from one or more reference portrait images. Unlike existing feed-forward approaches that require a fixed number of input views, FFAvatar supports incremental reconstruction, progressively refining the avatar representation as additional reference images become available. At the core of our method is an alternating attention mechanism that disentangles identity appearance from expression and viewpoint variations, enabling the reconstruction of a canonical 3D appearance that remains consistent across poses and facial expressions. To balance visual fidelity and computational efficiency, we introduce a sparse-to-dense learning paradigm. Coarse appearance features are first learned using sparse primitives anchored to the FLAME vertex level and are subsequently densified in the UV domain to capture fine-grained geometric and texture details. We further propose a plug-and-play motion refinement module that enables subject-specific dynamic personalization by modeling residual motion beyond parametric deformation. Extensive experiments demonstrate that FFAvatar efficiently produces high-fidelity and controllable 4D head avatars, achieving superior flexibility, driving efficiency, and identity-consistent rendering across diverse expressions and viewpoints.
Abstract:Neural rendering paradigms have recently emerged as powerful tools for radio frequency (RF). However, by entangling RF sources with scene geometry and material properties, existing approaches limit downstream manipulation of scene geometry, wireless system configuration, and RF reasoning. To address this, we propose a physically grounded RF inverse rendering (RFIR) framework that explicitly decouples RF emission, geometry, and material electromagnetic properties. Our key insight is an RF-aware bidirectional scattering distribution function, embedded into the Gaussian splatting paradigm as an RF rendering equation. Each Gaussian primitive is endowed with intrinsic physical attributes, including surface normals, material electromagnetic parameters, and roughness, and leveraged by a customized ray-tracing scheme to represent RF signal synthesis. The proposed RFIR generalizes three typical RF tasks: radar cross-section synthesis, received signal strength indicator prediction, and wireless scene editability. Experiments demonstrate significant performance advantages, underscoring the potential for wireless world modeling.




Abstract:This paper investigates the capabilities and effectiveness of backward localization centered on reconfigurable intelligent surfaces (RISs). In the backward sensing paradigm, the region of interest (RoI) is illuminated using a set of diverse radiation patterns. These patterns encode spatial information into a sequence of measurements, which are subsequently processed to reconstruct the RoI. We show that a single RIS can estimate the direction of arrival of incident waves by leveraging configurational diversity, and that the spatial diversity provided by multiple RISs further improves the accuracy of source localization and power estimation. The underlying structure of the sensing operator in the multi-snapshot measurement process is clarified. For single-RIS localization, the sensing operator is decomposed into a product of structured matrices, each corresponding to a specific physical process: wave propagation to and from the RIS, the relative phase offsets of elements with respect to the reference point, and the applied phase configuration of each element. A unified framework for identifying key performance indicators is established by analyzing the conditioning of the sensing operators. In the multi-RIS setting, we derive--via rank analysis--the governing law among the RoI size, the number of elements, and the number of measurements. Upper bounds on the relative error of the least squares reconstruction algorithm are derived. These bounds clarify how key performance indicators affect estimation error and provide valuable guidance for system-level optimization. Numerical experiments confirm that the trend of the relative error is consistent with the theoretical bounds.




Abstract:Accurate 3D localization is essential for realizing advanced sensing functionalities in next-generation Wi-Fi communication systems. This study investigates the potential of multistatic localization in Wi-Fi networks through the deployment of multiple cooperative antenna arrays. The collaborative gain offered by these arrays is twofold: (i) intra-array coherent gain at the wavelength scale among antenna elements, and (ii) inter-array cooperative gain across arrays. To evaluate the feasibility and performance of this approach, we develop WiCAL (Wi-Fi Collaborative Antenna Localization), a system built upon commercial Wi-Fi infrastructure equipped with uniform rectangular arrays. These arrays are driven by multiplexing embedded radio frequency chains available in standard access points or user devices, thereby eliminating the need for sophisticated, costly, and power-hungry multi-transceiver modules typically required in multiple-input and multiple-output systems. To address phase offsets introduced by RF chain multiplexing, we propose a three-stage, fine-grained phase alignment scheme to synchronize signals across antenna elements within each array. A bidirectional spatial smoothing MUSIC algorithm is employed to estimate angles of arrival (AoAs) and mitigate performance degradation caused by correlated interference. To further exploit inter-array cooperative gain, we elaborate on the synchronization mechanism among distributed URAs, which enables direct position determination by bypassing intermediate angle estimation. Once synchronized, the distributed URAs effectively form a virtual large-scale array, significantly enhancing spatial resolution and localization accuracy.




Abstract:Swarm antenna arrays, composed of spatially distributed antennas mounted on unmanned agents, offer unprecedented flexibility and adaptability for wireless sensing and communication. However, their reconfigurable architecture, susceptibility to collisions, and inherently stochastic nature present significant challenges to realizing collaborative gain. It remains unclear how spatial coordination, positional perturbations, and large-scale topological configurations affect coherent signal aggregation and overall system performance. This paper investigates the feasibility of achieving coherent beamforming in such systems from both deterministic and stochastic perspectives. First, we develop a rigorous theoretical framework that characterizes the necessary and sufficient conditions for the emergence of grating lobes in multiple linear configurations. Notably, we show that for dual linear arrays, the classical half-wavelength spacing constraint can be safely relaxed without introducing spatial aliasing. This result challenges traditional array design principles and enables more flexible, collision-aware topologies. Second, we present a theoretical analysis, supported by empirical validation, demonstrating that coherent gain can be approximately preserved under realistic positional perturbations. Our results reveal that spatial perturbations introduce measurable degradation in the main lobe, an effect that cannot be mitigated merely by increasing the number of antennas. Instead, the primary benefit of scaling lies in reducing the variance of perturbation-induced fluctuations.




Abstract:Fluid Antenna System (FAS) unlocks unprecedented flexibility in wireless channel optimization through spatial reconfigurability. However, its practical deployment is hindered by the coupled challenges posed by high-dimensional channel estimation and real-time position optimization. This paper bridges wireless propagation physics with compressed sensing theory to address these challenges through three aspects. First, we establish a group-sparse recovery framework for space-frequency characteristics (SFC) in FAS, formally characterizing leakage-induced sparsity degradation from limited aperture and bandwidth as a structured group-sparsity problem. By deriving dictionary-adapted group restricted isometry property (D-GRIP), we prove tight recovery bounds for a convex $\ell_1/\ell_2$-mixed norm optimization formulation that preserves leakage-aware sparsity patterns. Second, we develop a Descending Correlation Group Orthogonal Matching Pursuit (DC-GOMP) algorithm that systematically relaxes leakage constraints to reduce subcoherence. This approach enables robust FSC recovery with accelerated convergence and superior performance compared to conventional compressive sensing methods like OMP or GOMP. Third, we formulate spatial equalization (SE) as a mixed-integer linear programming (MILP) problem, ensuring optimality through the branch-and-bound method. To achieve real-time implementability while maintaining near-optimal performance, we complement this with a greedy algorithm. Simulation results demonstrate the proposed channel estimation algorithm effectively resolves energy misallocation and enables recovery of weak details, achieving superior recovery accuracy and convergence rate. The SE framework suppresses deep fading phenomena and reduces hardware deployment overhead while maintaining equivalent link reliability.




Abstract:Existing phase optimization methods in reconfigurable intelligent surfaces (RISs) face significant challenges in achieving flexible beam synthesis, especially for directional beam suppression. This paper introduces a Max-min criterion incorporating non-linear constraints, utilizing optimization techniques to enable multi-beam enhancement and suppression via transmissive RISs. A realistic model grounded in geometrical optics is first presented to characterize the input/output behavior of transmissive RIS, effectively linking explicit beam-forming operations with practical implementation. Subsequently, a highly efficient bisection-based algorithm for constrained Max-min optimization involving quadratic forms is developed, utilizing an auxiliary variable and Moreau envelope to iteratively reach the optimal solution. This approach demonstrates excellent extensibility and is applicable to a wide range of constrained Max-min problems. Numerical simulations validate the proposed methods, confirming that the framework enables beam enhancement or suppression at designated spatial positions.




Abstract:Transformers have significantly advanced the field of 3D human pose estimation (HPE). However, existing transformer-based methods primarily use self-attention mechanisms for spatio-temporal modeling, leading to a quadratic complexity, unidirectional modeling of spatio-temporal relationships, and insufficient learning of spatial-temporal correlations. Recently, the Mamba architecture, utilizing the state space model (SSM), has exhibited superior long-range modeling capabilities in a variety of vision tasks with linear complexity. In this paper, we propose PoseMamba, a novel purely SSM-based approach with linear complexity for 3D human pose estimation in monocular video. Specifically, we propose a bidirectional global-local spatio-temporal SSM block that comprehensively models human joint relations within individual frames as well as temporal correlations across frames. Within this bidirectional global-local spatio-temporal SSM block, we introduce a reordering strategy to enhance the local modeling capability of the SSM. This strategy provides a more logical geometric scanning order and integrates it with the global SSM, resulting in a combined global-local spatial scan. We have quantitatively and qualitatively evaluated our approach using two benchmark datasets: Human3.6M and MPI-INF-3DHP. Extensive experiments demonstrate that PoseMamba achieves state-of-the-art performance on both datasets while maintaining a smaller model size and reducing computational costs. The code and models will be released.