Abstract:Multimodal representation learning is a cornerstone of modern AI. By encoding multimodal queries and targets into vectors, it powers industrial search and recommendation and underpins modern agents. Real-world platforms with complex modalities and massive-scale content, such as Douyin, Xiaohongshu, and YouTube, demand both efficiency under billion-scale indexing and fine-grained discrimination for hard matching. Existing MLLM embedding models rarely satisfy both. Contrastive models are efficient but rely on pair-level supervision too coarse for fine-grained distinctions, while CoT-based models improve discrimination through explicit generation impractical to serve online. We present Douyin Multimodal Embedding (DME), a model trained in two stages to combine both strengths. Stage 1 performs large-scale contrastive pre-training that establishes a unified multimodal embedding space with broad modality and task coverage. Stage 2 supplements semantic sufficiency, the property that an embedding is grounded in retrieval-relevant evidence and preserves fine-grained counterpart-side semantics, via two mechanisms. Evidence-Grounded Typed Latent Reasoning organizes retrieval evidence through hidden-space latent reasoning, and Cross-Conditional Reconstruction enforces counterpart-side semantics through cross-directional autoregressive reconstruction. Both act only during training and add only marginal query-side overhead, so DME serves as efficiently as a standard contrastive encoder. On MMEB-v2, DME reaches state-of-the-art results at comparable scales for its 2B and 9B variants (74.8 and 78.4), with especially strong video and visual-document tasks. In production, DME delivers a 2.92% relative gain on Douyin's in-house offline evaluation set, is deployed across Douyin scenarios such as generative, image, and AI search, and yields a 0.1% Lifetime (LT) gain in online A/B testing on Douyin search.
Abstract:4D spatio-temporal reasoning, jointly modeling 3D spatial structure and temporal evolution, is essential for understanding dynamic worlds and enabling embodied interaction. While current Multimodal Large Language Models (MLLMs) show strong capabilities in static scene understanding and coarse-grained 4D tasks, they still have notable limitations in continuous dynamic scene perception, especially in tracking dynamic object evidence for coherent 4D spatio-temporal reasoning. This shortcoming stems mainly from relying on sparse frame-level observations, fragmenting continuous dynamic cues and leaving models unable to disentangle genuine object dynamics from camera-induced apparent motion. Inspired by humans tracking dynamic cues while compensating for viewpoint changes, we propose DynTrace, a training-free framework for 4D spatio-temporal reasoning with two complementary components. Dynamic Trajectory Visualization (DTV) reprojects world-coordinate trajectories onto the image plane, providing geometry-informed visual priors that disentangle genuine object dynamics from camera-induced apparent motion. Meanwhile, the Dynamic Trace Token (DT-Token), organized into a Dynamic Trace Graph (DTG), tracks object-level dynamic cues, trace evolution, and key moments, maintaining continuous dynamic object evidence for coherent 4D reasoning. Together, these two components equip MLLMs with continuously tracked dynamic object evidence, grounded in geometry-informed visual priors and structured spatio-temporal traces. DynTrace consistently improves open-source MLLMs, achieving state-of-the-art results on Dyn-Bench, VLM4D, and DSI-Bench, validating the importance of tracking dynamic object evidence for robust 4D spatio-temporal reasoning.
Abstract:The design of low-complexity transceivers is crucial for the deployment of next-generation wireless systems. In this work, we combine two emerging concepts, movable antennas (MA) and transmissive reconfigurable intelligent surfaces (TRIS), which have recently attracted significant attention for enhancing wireless communication performance. In particular, we propose a compact base station (BS) architecture that integrates a single MA with a TRIS operating in their near-field region. We address the joint optimization of the MA location and the quantized TRIS phase configuration. Due to the non-convex coupling between spatial positioning and discrete phase constraints, an alternating optimization (AO) framework is developed, where the MA position is updated via gradient ascent (GA) and the TRIS phases are optimized through quantized phase alignment. Simulation results demonstrate that the proposed architecture significantly outperforms conventional BS designs equipped with fixed fully-active antenna arrays under the same channel model and transmit power constraint. Moreover, MA repositioning effectively mitigates the performance degradation caused by discrete TRIS phase quantization in near-field propagation environments. This reveals a favorable trade-off between hardware complexity and spatial signal processing, where the spatial adaptability of the MA can compensate for low-resolution TRIS phase control.




Abstract:Reconfigurable Intelligent Surfaces (RIS) have been recognized as a promising technology to enhance both communication and sensing performance in integrated sensing and communication (ISAC) systems for future 6G networks. However, existing RIS optimization methods for improving ISAC performance are mainly based on semidefinite relaxation (SDR) or iterative algorithms. The former suffers from high computational complexity and limited scalability, especially when the number of RIS elements becomes large, while the latter yields suboptimal solutions whose performance depends on initialization. In this work, we introduce a lightweight RIS phase design framework that provides a closed-form solution and explicitly accounts for the trade-off between communication and sensing, as well as proportional beam gain distribution toward multiple sensing targets. The key idea is to partition the RIS configuration into two parts: the first part is designed to maximize the communication performance, while the second introduces small perturbations to generate multiple beams for multi-target sensing. Simulation results validate the effectiveness of the proposed approach and demonstrate that it achieves performance comparable to SDR but with significantly lower computational complexity.
Abstract:This study presents a novel time series prediction model, FPN-fusion, designed with linear computational complexity, demonstrating superior predictive performance compared to DLiner without increasing parameter count or computational demands. Our model introduces two key innovations: first, a Feature Pyramid Network (FPN) is employed to effectively capture time series data characteristics, bypassing the traditional decomposition into trend and seasonal components. Second, a multi-level fusion structure is developed to integrate deep and shallow features seamlessly. Empirically, FPN-fusion outperforms DLiner in 31 out of 32 test cases on eight open-source datasets, with an average reduction of 16.8% in mean squared error (MSE) and 11.8% in mean absolute error (MAE). Additionally, compared to the transformer-based PatchTST, FPN-fusion achieves 10 best MSE and 15 best MAE results, using only 8% of PatchTST's total computational load in the 32 test projects.




Abstract:Crowdsourcing platforms have transformed distributed problem-solving, yet quality control remains a persistent challenge. Traditional quality control measures, such as prescreening workers and refining instructions, often focus solely on optimizing economic output. This paper explores just-in-time AI interventions to enhance both labeling quality and domain-specific knowledge among crowdworkers. We introduce LabelAId, an advanced inference model combining Programmatic Weak Supervision (PWS) with FT-Transformers to infer label correctness based on user behavior and domain knowledge. Our technical evaluation shows that our LabelAId pipeline consistently outperforms state-of-the-art ML baselines, improving mistake inference accuracy by 36.7% with 50 downstream samples. We then implemented LabelAId into Project Sidewalk, an open-source crowdsourcing platform for urban accessibility. A between-subjects study with 34 participants demonstrates that LabelAId significantly enhances label precision without compromising efficiency while also increasing labeler confidence. We discuss LabelAId's success factors, limitations, and its generalizability to other crowdsourced science domains.




Abstract:One of the great potentials to improve the confidentiality in mmWave/THz at the physical layer of technical communication, measured by the secrecy rate, lies in the use of reconfigurable intelligent surfaces (RISs). However, an important open problem arises when the eavesdropper is aligned with the legitimate user or in proximity to the RIS or legitimate user. The limitation comes, on one hand, from the high directional gain caused by the dominant line-of-sight (LOS) path in high-frequency transmission, and, on the other hand, from the high energy leakage in the proximity of the RIS and the legitimate user. To address these issues, we employ the concept of frequency diverse arrays (FDA) at the base station (BS) associated with random inverted transmit beamforming and reflective element subset selection (RIBES). More specifically, we consider a passive eavesdropper with unknown location, and design the transmit beamforming and RIS configuration based on the channel information of the legitimate user only. In this context, the secrecy rate with the proposed transmission technique is evaluated in the case of deterministic eavesdropper channel, demonstrating that we can ensure a secure transmission regarding both direction and range. Furthermore, assuming no prior information about the eavesdropper, we describe the wiretap region and derive the worst-case secrecy rate in closed form. The latter is further optimized by determining the optimal subset sizes of the transmit antennas and reflective elements. Simulations verify the correctness of the closed-form expressions and demonstrate that we can effectively improve the secrecy rate, especially when the eavesdropper is close to the RIS or the legitimate user.

Abstract:In this work, we investigate the secrecy performance in an intelligent reflecting surface (IRS)-assisted downlink system. In particular, we consider a base station (BS)-side IRS and as such, the BS-IRS channel is assumed to be known perfectly. Of more importance, we consider the case, in which only outdated channel state information (CSI) of the IRS-user channel is available. We study the impact of outdated CSI on the secrecy performance numerically and analytically. Furthermore, we propose an element subset selection (ESS) method in order to improve the secrecy performance. A key observation is that minimal secrecy outage probability (SOP) can be achieved using a subset of the IRS, and the optimal number of selected reflecting elements can be effectively found by closed-form expressions.


Abstract:In this work, we study the impact of the multiplicative phase noise in an IRS-assisted system. We consider an IRS-assisted system with multiplicative phase noise both at the BS and user. A novel channel estimation algorithm is proposed considering the phase noise. By utilizing the proposed channel estimates we investigate the system performance in the downlink, more specifically, we derive the ergodic capacity in closed form. Simulation results verify the correctness of the closed-form expression. We observe that the system becomes more robust against the phase noise as the number of reflective elements increases. Moreover, the influence of the additive channel noise in uplink vanishes as the number of reflecting elements grows asymptotically large.



Abstract:Recently channel state information (CSI) measurements from commercial multi input multi output (MIMO) WiFi systems have been ubiquitously used for different wireless sensing applications. However, the phase of the CSI realizations is usually distorted severely by phase errors due to the hardware impairments, which significantly reduce the sensing performance. In this paper, we directly utilize the modeling of the phase distortions caused by the hardware impairments and propose an adaptive CSI estimation approach based on Kalman filter (KF) with maximum a posteriori (MAP) estimation that considers the CSI from the previous time. The performance of the proposed algorithm is compared against the Cramer Rao lower bound (CRLB). Simulation and experimental results demonstrate that our approach can track the channel variations while eliminating the phase errors accurately.