Charlie
Abstract:Unmanned aerial vehicle (UAV) communication is expected to support a wide range of low-altitude applications in 6G mobile networks. However, traditional statistical channel models provide limited accuracy in specific environments, while deterministic methods such as ray tracing usually rely on accurate three-dimensional environment models and involve high computational complexity. Existing multimodal channel prediction approaches mainly focus on large-scale metrics such as path loss, and remain insufficient for modeling small-scale parameters. To address these limitations, this paper proposes PanoLAMP, a Panoramic perception and vision-language model-based Low-Altitude Multipath Prediction framework. It adopts a pretrained vision-language model as the backbone and captures the propagation environment features through panoramic RGB-D observations collected at both the transmitter and receiver to predict the delay, power, azimuth angle, and zenith angle offset relative to the line-of-sight path. Experiments are conducted on a synthetic dataset containing 18,949 UAV-vehicle links across seven UAV altitudes. Experimental results show that the proposed method consistently outperforms representative baselines in both multipath parameters and statistical metrics, and demonstrates stronger generalization across different flight heights.
Abstract:Low-altitude unmanned aerial vehicles (UAVs) are emerging as key platforms for wireless intelligence tasks. However, practical low-altitude wireless systems usually operate in complex urban environments, where visual occlusion, sparse geometric observations, multipath propagation, and sensor failures may degrade the reliability of single-modality models. To address these challenges, this paper proposes M3F-UAV, a missing-modality multimodal foundation model for low-altitude wireless sensing. The proposed framework learns a unified multimodal representation from visual, geometric, and wireless observations. Specifically, modality-specific pretrained feature extractors are adopted for RGB/depth images, LiDAR point clouds, and CSI matrices, respectively. Through cross-modal fusion and missing-modality-aware pretraining with feature-level masked reconstruction and UAV localization objectives, M3F-UAV can extract fixed-size features from different modality combinations and adapt them to downstream low-altitude wireless tasks with lightweight task heads. Experiments on the LAMBDA dataset show that M3F-UAV outperforms single-modality baselines and maintains robust performance under missing-modality settings.
Abstract:Research on low-altitude integrated sensing and communication (ISAC) requires aligned multimodal data that jointly describe wireless propagation, visual appearance, unmanned aerial vehicle (UAV) motion, light detection and ranging (LiDAR) perception, and radar sensing under common trajectories and timestamps. To address this need, a low-altitude multimodal base dataset, named LAMBDA, is introduced. LAMBDA is characterized by high fidelity, modality diversity, scenario richness, and configuration flexibility. It is generated through a high-fidelity digital-twin pipeline with detailed scene geometry, refined material assignment, and electromagnetic modeling of UAVs. LAMBDA provides synchronized RGB images, depth maps, LiDAR point clouds, inertial measurement unit states, UAV poses, channel state information (CSI), and radar-synthesis resources across matched low-altitude operating conditions, shared coordinate systems, and synchronized frame indices. The dataset covers urban, suburban, and campus scenes, multi-UAV/multi-base-station settings, nighttime conditions, and sunny, rainy, snowy, and foggy weather variations. Its CSI and radar resources support user-defined antenna-array sizes, bandwidths, subcarrier spacings, chirp parameters, and plane-wave or spherical-wavefront channel synthesis. The reliability and usability of LAMBDA are assessed through quality control, weather and multimodal visualization, and two UAV ISAC-related use cases: RGB-aided beam prediction and RGB-LiDAR-based UAV localization.
Abstract:Multistatic collaborative sensing eliminates self-interference, achieves spatial diversity gains, and enables wide-range seamless integrated sensing and communication (ISAC). However, conventional data fusion methods suffer from severe error amplification in geometry-sensitive regions. In addition, the conventional analog phased array solution introduces large beam sweeping overhead, whereas the fully digital arrays request high hardware cost. We propose a multistatic sensing framework enabled by a phase-time array (PTA). The rainbow beamforming maps spatial directions to orthogonal frequency division multiplexing (OFDM) subcarriers, achieving wide-angle coverage with a single radio frequency (RF) chain. We develop two parameter-level schemes-a geometry-aware analytical estimator (GDOP-WLS) and a lightweight multilayer perceptron (PF-MLP)-to mitigate the effects of topological singularities. Additionally, an end-to-end signal-level convolutional neural network (SF-CNN) directly estimates target coordinates from raw signals, avoiding cascaded estimation errors. The results demonstrate that the parameter-level schemes ensure robust convergence under adverse geometric conditions with minimal computational latency. Conversely, the signal-level scheme achieves sub-meter precision but requires an increased computational load. Consequently, the proposed framework establishes a scalable solution for collaborative surveillance of unmanned aerial vehicles (UAVs), providing flexible trade-offs among hardware complexity, latency, and accuracy.




Abstract:Phase-time arrays, which integrate phase shifters (PSs) and true-time delays (TTDs), have emerged as a cost-effective architecture for generating frequency-dependent rainbow beams in wideband sensing and localization. This paper proposes an end-to-end deep learning-based scheme that simultaneously designs the rainbow beams and estimates user positions. Treating the PS and TTD coefficients as trainable variables allows the network to synthesize task-oriented beams that maximize localization accuracy. A lightweight fully connected module then recovers the user's angle-range coordinates from its feedback of the maximum quantized received power and its corresponding subcarrier index after a single downlink transmission. Compared with existing analytical and learning-based schemes, the proposed method reduces overhead by an order of magnitude and delivers consistently lower two-dimensional positioning error.




Abstract:Integrated sensing and communication (ISAC) systems demand precise and efficient target localization, a task challenged by rich multipath propagation in complex wireless environments. This paper introduces MARBLE-Net (Multipath-Aware Rainbow Beam Learning Network), a deep learning framework that jointly optimizes the analog beamforming parameters of a frequency-dependent rainbow beam and a neural localization network for high-accuracy position estimation. By treating the phase-shifter (PS) and true-time-delay (TTD) parameters as learnable weights, the system adaptively refines its sensing beam to exploit environment-specific multipath characteristics. A structured multi-stage training strategy is proposed to ensure stable convergence and effective end-to-end optimization. Simulation results show that MARBLE-Net outperforms both a fixed-beam deep learning baseline (RaiNet) and a traditional k-nearest neighbors (k-NN) method, reducing localization error by more than 50\% in a multipath-rich scene. Moreover, the results reveal a nuanced interaction with multipath propagation: while confined uni-directional multipath degrades accuracy, structured and directional multipath can be effectively exploited to achieve performance surpassing even line-of-sight (LoS) conditions.
Abstract:Most existing semantic communication systems employ analog modulation, which is incompatible with modern digital communication systems. Although several digital transmission approaches have been proposed to address this issue, an end-to-end bit-level method that is compatible with arbitrary modulation formats, robust to channel noise, and free from quantization errors remains lacking. To this end, we propose BitSemCom, a novel bit-level semantic communication framework that realizes true joint source-channel coding (JSCC) at the bit level. Specifically, we introduce a modular learnable bit mapper that establishes a probabilistic mapping between continuous semantic features and discrete bits, utilizing the Gumbel-Softmax trick to enable differentiable bit generation. Simulation results on image transmission demonstrate that BitSemCom achieves both competitive performance and superior robustness compared to traditional separate source-channel coding (SSCC) schemes, and outperforms deep learning based JSCC with uniform 1-bit quantization, validating the effectiveness of the learnable bit mapper. Despite these improvements, the bit mapper adds only 0.42% parameters and 0.09% computational complexity, making BitSemCom a lightweight and practical solution for real-world semantic communication.
Abstract:Millimeter-wave (mmWave) OFDM radar equipped with rainbow beamforming, enabled by joint phase-time arrays (JPTAs), provides wide-angle coverage and is well-suited for fast real-time target detection and tracking. However, accurate detection of multiple closely spaced targets remains a key challenge for conventional signal processing pipelines, particularly those relying on constant false alarm rate (CFAR) detectors. This paper presents CFARNet, a learning-based processing framework that replaces CFAR with a convolutional neural network (CNN) for peak detection in the angle-Doppler domain. The network predicts target subcarrier indices, which guide angle estimation via a known frequency-angle mapping and enable high-resolution range and velocity estimation using the MUSIC algorithm. Extensive simulations demonstrate that CFARNet significantly outperforms a CFAR+MUSIC baseline, especially under low transmit power and dense multi-target conditions. The proposed method offers superior angular resolution, enhanced robustness in low-SNR scenarios, and improved computational efficiency, highlighting the potential of data-driven approaches for high-resolution mmWave radar sensing.
Abstract:Massive multiple-input multiple-output (MIMO) technology is a key enabler of modern wireless communication systems, which demand accurate downlink channel state information (CSI) for optimal performance. Although deep learning (DL) has shown great potential in improving CSI feedback, most existing approaches fail to exploit the semantic relationship between CSI and other related channel metrics. In this paper, we propose SemCSINet, a semantic-aware Transformer-based framework that incorporates Channel Quality Indicator (CQI) into the CSI feedback process. By embedding CQI information and leveraging a joint coding-modulation (JCM) scheme, SemCSINet enables efficient, digital-friendly CSI feedback under noisy feedback channels. Experimental results on DeepMIMO datasets show that SemCSINet significantly outperforms conventional methods, particularly in scenarios with low signal-to-noise ratio (SNR) and low compression ratios (CRs), highlighting the effectiveness of semantic embedding in enhancing CSI reconstruction accuracy and system robustness.




Abstract:Joint phase-time arrays (JPTA) emerge as a cost-effective and energy-efficient architecture for frequency-dependent beamforming in wideband communications by utilizing both true-time delay units and phase shifters. This paper exploits the potential of JPTA to simultaneously serve multiple users in both near- and far-field regions with a single radio frequency chain. The goal is to jointly optimize JPTA-based beamforming and subband allocation to maximize overall system performance. To this end, we formulate a system utility maximization problem, including sum-rate maximization and proportional fairness as special cases. We develop a 3-step alternating optimization (AO) algorithm and an efficient deep learning (DL) method for this problem. The DL approach includes a 2-layer convolutional neural network, a 3-layer graph attention network (GAT), and a normalization module for resource and beamforming optimization. The GAT efficiently captures the interactions between resource allocation and analog beamformers. Simulation results confirm that JPTA outperforms conventional phased arrays (PA) in enhancing user rate and strikes a good balance between PA and fully-digital approach in energy efficiency. Employing a logarithmic utility function for user rates ensures greater fairness than maximizing sum-rates. Furthermore, the DL network achieves comparable performance to the AO approach, while having orders of magnitude lower computational complexity.