Abstract:Multimodal representation learning is a cornerstone of modern AI. By encoding multimodal queries and targets into vectors, it powers industrial search and recommendation and underpins modern agents. Real-world platforms with complex modalities and massive-scale content, such as Douyin, Xiaohongshu, and YouTube, demand both efficiency under billion-scale indexing and fine-grained discrimination for hard matching. Existing MLLM embedding models rarely satisfy both. Contrastive models are efficient but rely on pair-level supervision too coarse for fine-grained distinctions, while CoT-based models improve discrimination through explicit generation impractical to serve online. We present Douyin Multimodal Embedding (DME), a model trained in two stages to combine both strengths. Stage 1 performs large-scale contrastive pre-training that establishes a unified multimodal embedding space with broad modality and task coverage. Stage 2 supplements semantic sufficiency, the property that an embedding is grounded in retrieval-relevant evidence and preserves fine-grained counterpart-side semantics, via two mechanisms. Evidence-Grounded Typed Latent Reasoning organizes retrieval evidence through hidden-space latent reasoning, and Cross-Conditional Reconstruction enforces counterpart-side semantics through cross-directional autoregressive reconstruction. Both act only during training and add only marginal query-side overhead, so DME serves as efficiently as a standard contrastive encoder. On MMEB-v2, DME reaches state-of-the-art results at comparable scales for its 2B and 9B variants (74.8 and 78.4), with especially strong video and visual-document tasks. In production, DME delivers a 2.92% relative gain on Douyin's in-house offline evaluation set, is deployed across Douyin scenarios such as generative, image, and AI search, and yields a 0.1% Lifetime (LT) gain in online A/B testing on Douyin search.
Abstract:Pinching-antenna systems (PASS) enhance wireless propagation by activating or placing pinching antennas (PAs) near users. Therefore, accurate uplink positioning is essential for efficient communication. In this paper, an uplink multi-carrier positioning framework is established for PASS in multipath environments. Matrix pencil (MP)-based and low-complexity Rank-1 ranging algorithms are proposed to estimate the distances between the PAs and the user. For the MP-based ranging algorithm, the line-of-sight (LoS) component is separated from non-line-of-sight components by exploiting the shift-invariance property of the Hankel matrix, thereby enabling accurate distance estimation. For the Rank-1 ranging algorithm, the dominant LoS delay is directly isolated through truncated singular value decomposition, thereby avoiding matrix inversions. Subsequently, a two-stage weighted nonlinear least-squares (WNLS) positioning algorithm is designed to estimate the three-dimensional user position. To gain further insights, a comprehensive theoretical performance analysis of the proposed ranging and positioning algorithms is conducted. The closed-form ranging variances and position error bound (PEB) are derived to reveal the error propagation mechanism. Numerical results demonstrate that: i) The MP-based algorithm achieves higher accuracy and robustness than the Rank-1-based algorithm, while the Rank-1-based algorithm has lower computational complexity. ii) The positioning error of the MP-based algorithm follows the same trend as the derived PEB, whereas the Rank-1 algorithm exhibits an error floor due to multipath bias. iii) The positioning accuracy of the MP algorithm improves as the number of subcarriers increases.
Abstract:The performance of pinching-antenna systems (PASS) is fundamentally affected by line-of-sight (LoS) blockage in practical environments. In this paper, PASS is investigated under realistic, obstacle-induced blockage by jointly considering the LoS and non-LoS (NLoS) components, rather than relying on a LoS channel or a probabilistic blockage model. A geometry-aware blockage model is adopted, where a blockage region on the waveguide is defined according to the actual locations and geometric features of obstacles, such that a pinching-antenna (PA) located within the blockage region is unable to establish a LoS link to the user equipment (UE). The channel models of PASS are developed by jointly accounting for in-waveguide attenuation and spatial propagation loss. To quantify the impact of channel factors on PASS performance, a single-PA single-UE scenario is studied under Rayleigh and Rician fading channels. Closed-form expressions for the outage probability are derived for both cases. For the ergodic rate, a closed-form expression is obtained in the Rayleigh case, while a complete analytical expression and an approximate closed-form expression are derived in the Rician case. Analytical expressions are derived for the endpoints of the blockage region, and the deployment criteria of optimal PA are provided. Simulation results validate the analysis and reveal that: i) NLoS scattering has a twofold effect on PASS performance, potentially degrading the outage performance while improving the rate performance under Rician fading; ii) Sufficiently strong NLoS scattering can still sustain communication in the presence of LoS blockage; iii) The optimal PA position is jointly determined by the environment geometry and the interplay between spatial propagation loss and in-waveguide attenuation.
Abstract:A low-overhead site-specific multi-user multiple-input multiple-output (MU-MIMO) beamforming framework is proposed. Conventional limited-feedback MU-MIMO relies on channel state information reference signal (CSI-RS) transmission and user feedback before grouping and beamforming, which requires substantial online overhead when the antenna dimension and candidate-user pool are large. To reduce this burden, the proposed framework exploits site-specific information (SSI), which captures local radio propagation features. By learning the mapping from low-overhead beam-domain observations to effective transmit spatial subspaces of users, the BS can infer inter-user separability before high-resolution CSI acquisition and construct a compact group-level CSI acquisition subspace for the selected users. This site-specific design can be implemented within the standard limited-feedback procedure using synchronization signal block (SSB)-based reference signal received power (RSRP) fingerprints for subspace inference and CSI-RS feedback for low-dimensional CSI refinement. Extensive numerical results demonstrate that the proposed framework can identify compatible user groups before CSI-RS acquisition, preserve most scheduled-user channel energy in a compact group subspace, and achieve higher effective rates than conventional systems with significantly lower overhead and user-side processing burden.
Abstract:This article proposes the concept of \emph{brain-body-to-everything (B2X)} networks to facilitate the integration of wireless networks and embodied intelligence. In this framework, the \emph{brain} refers to the intelligence functions for reasoning, planning, and decision-making, the \emph{body} denotes the physical embodied agent that senses and acts in the real world, and \emph{X} represents the surrounding ecosystem involved in the brain-body interaction loop. Two B2X architectures with \emph{distributed} and \emph{centralized} brains are introduced to characterize different placements of intelligence across the body, base station, and core network. The uplink and downlink designs of B2X networks are then discussed under a representative base-station-side brain setting. For the uplink, communication is redesigned for B2X state acquisition under event urgency, sensing volume, and simultaneous multi-body access. For the downlink, communication is redesigned to coordinate command delivery and conventional service under shared radio resources. Based on these uplink and downlink considerations, a communication-control Pareto boundary is further used to characterize the loop-level trade-off between wireless transmission performance and control quality in B2X networks. Finally, several open research problems are discussed to guide future B2X network design.
Abstract:The sensing capability of the pinching-antenna system (PASS) is analyzed from a Ziv-Zakai bound (ZZB) perspective, motivated by the sensing ambiguity arising from the multimodal observation model inherent to PASS. In comparison to other Bayesian sensing bounds, the ZZB provides a lower bound on the mean-squared error (MSE) across a broad range of signal-to-noise ratios (SNRs) and accounts for ambiguity in the likelihood functions. First, an observation model is developed for an uplink sensing scenario where a single sensing target transmits uplink pilots to a single-waveguide PASS receiver equipped with multiple pinching antennas (PAs). Building on this model, general ZZB expressions are derived for arbitrary prior distributions of the target's position, and are then specialized to the Gaussian and uniform cases. Second, the asymptotic ZZBs in low- and high-SNR regimes are characterized, and the relationship between the ZZBs and the conventional Bayesian Cramér-Rao bound (BCRB) is further studied by introducing the concept of an ambiguity function. Furthermore, to reduce the high computational complexity of direct evaluation of the ZZB, SNR-free and SNR-aware surrogate objective functions are proposed to facilitate ZZB-based optimization for enhancing sensing performance. Numerical results demonstrate that: i) Compared with the BCRB, the ZZB provides a tight sensing performance lower bound over a wide range of SNRs, ii) the ambiguity-awareness of the ZZB can address the multimodality-induced ambiguity in sensing, thereby yielding a reliable lower bound on the MSE, and iii) the proposed surrogate objective functions enable effective ZZB minimization with a lower computational complexity.
Abstract:This paper proposes an access protocol framework for segmented waveguide-enabled pinching-antenna systems (SWANs), which exploits SWAN-induced reconfigurable channel diversity as a protocol-level resource for uplink random access. The framework consists of two stages, a channel-oracle stage and an access stage, designed under three SWAN operating modes: (i) one-segment selection (OS), (ii) segment aggregation (SA), and (iii) segment multiplexing (SM). Specifically, in the channel oracle stage, the OS mode is adopted to acquire sparse pilot observations and infer the channel responses across the SWAN configuration space. In this way, high-dimensional uplink channel acquisition is recast as a low-dimensional geometric localization problem, thereby reducing pilot overhead while preserving channel reconstruction accuracy. For the access stage, we construct two oracle-guided access codebooks under the SA and SM modes, respectively, which address the tradeoff between hardware complexity and multiuser access resolution. In particular, the SA-based scheme supports single radio frequency (RF) chain access through randomized segment-group activation, whereas the SM-based R-access scheme exploits multiple RF chains to construct deterministic access slots and enhance collision resolution. Finally, our numerical results demonstrate that (i) the proposed two-stage framework improves access performance under the same training overhead, (ii) anchor densification is more effective than aggressive segment aggregation for SA, and (iii) SM-based R-access achieves deterministic coverage and higher throughput in moderate- and high-load regimes, whereas SA-based access remains attractive for low-complexity implementations.
Abstract:With the rapid development of satellite communication and navigation, there is an urgent need to integrate both technologies to achieve reliable communication and precise navigation services within the same satellite system. By combining multi-/uni-cast (MUC) and non-orthogonal multiple access (NOMA) technologies, we propose a novel MUC-NOMA-based integrated navigation and communication (INAC) signal structure, in which the navigation and communication signals share a common pseudo noise (PN) sequence, thereby integrating satellite communication and navigation at the signal level. According to different power allocation strategies, two scenarios are defined: multi-cast-oriented (MO-) INAC and uni-cast-oriented (UO-) INAC, where a greater portion of power is assigned to either the multi-cast or the uni-cast signal, respectively. To mitigate co-channel interference, we employ successive interference cancellation (SIC) at the receiver and design a signal processing algorithm for the proposed INAC signal. Then, closed-form expressions are subsequently derived for the bit error rates (BER) of both the navigation and communication signals, along with the positioning accuracy of the navigation signal. To gain further insights, the impacts of power allocation factors and communication rates are evaluated. Our analysis results show that: i) In the MO-INAC scenario, the positioning and BER performance of navigation signal are excellent when more power is assigned to the multi-cast signal; ii) In the UO-INAC scenario, interference in the shared resources is reduced when more power is assigned to the uni-cast signal; iii) The ranging accuracy decreases as the communication data rate increases. Numerical results confirm the superior BER and positioning accuracy of the MO-INAC scenario for MEO satellites.
Abstract:The pinching-antenna systems (PASS), which dynamically activate and relocate the pinching-antennas (PAs) along the dielectric waveguide, offer unprecedented potential for integrated positioning and communication. The multi-waveguide-based uplink positioning approaches for indoor environments are first proposed in this paper, and the downlink communication performance is analyzed. Two possible scenarios, multi-waveguide single-PA (MWSP) and multi-waveguide multi-PA (MWMP), are considered under the assumptions of line-of-sight channels and a single, stationary user. For the MWSP scenario, the received signal strength indication (RSSI)-based ranging method and the MWSP-based least square (LS) positioning algorithm are developed. To gain deeper insights, a comprehensive error analysis of the LS positioning algorithm is conducted. Subsequently, for the MWMP scenario, the closed-form expression of the superposed signal is derived. According to the signal power, the MWMP-based grid search algorithm is proposed and the estimation error of proposed algorithm is analyzed. Then, based on the user's positioning result, the PAs are relocated to provide downlink communication service, and the achievable data rate of MWSP and MWMP scenarios are analyzed. Numerical results validate the correctness of our analysis, which show that: i) For the MWSP scenario, a smaller geometric dilution of precision (GDoP) leads to a lower average positioning error. Furthermore, even when the GDoP is large, the regions where the distances to PAs are nearly equal achieve the best accuracy. ii) For the MWMP scenario, non-parallel waveguide deployment improves positioning accuracy, although errors increase with the number of PAs. iii) The noise has a serious double-impact on data rate. There is a trade-off between positioning accuracy and communication performance.
Abstract:This article analyzes the achievable sum-rate of multiuser uplink segmented waveguide-enabled pinching-antenna systems (SWANs). To unveil system-design insights, an upper bound on the achievable sum-rate is derived, based on which the existence of an optimal segment activation level is theoretically established. Motivated by this result, hybrid segment selection and aggregation (HSS/A) schemes are proposed to jointly optimize segment activation and pinching-antenna (PA) placement. Correspondingly, low-complexity greedy algorithms are developed for the considered optimization problem. Numerical results validate the theoretical analysis and demonstrate that the proposed HSS/A schemes outperform conventional full-segment aggregation.