Abstract:Rydberg atomic receiver has emerged as promising candidate for next-generation wireless communication, due to the exceptional sensitivity and ability to overcome the physical limitations of traditional radio frequency antennas. Utilizing the resonant response of atomic energy levels for signal detection, Rydberg atomic receiver is inherently confined to a narrow instantaneous bandwidth. However, in high-mobility scenarios such as satellite communications, the severe Doppler effect induces carrier frequency offsets, which drive the signal beyond the instantaneous bandwidth and result in severe distortion. In this paper, we propose an adaptive local oscillator (LO) tracking Rydberg atomic receiver architecture designed to lock high-dynamic signals within the effective atomic response bandwidth. By employing a cross-product automatic frequency control (CPAFC) algorithm, the system dynamically estimates the instantaneous frequency offset, generates a corresponding error control signal, and adjusts the LO frequency through a feedback loop. Consequently, the intermediate frequency signal can always be locked close to the center of the atomic response bandwidth regardless of dynamics. Simulation results show that the proposed architecture significantly outperforms existing Rydberg atomic receiver, effectively alleviating performance degradation in high-dynamic environments.




Abstract:The increasing complexity of deep neural networks (DNNs) poses significant challenges for edge inference deployment due to resource and power constraints of edge devices. Recent works on unary-based matrix multiplication hardware aim to leverage data sparsity and low-precision values to enhance hardware efficiency. However, the adoption and integration of such unary hardware into commercial deep learning accelerators (DLA) remain limited due to processing element (PE) array dataflow differences. This work presents Tempus Core, a convolution core with highly scalable unary-based PE array comprising of tub (temporal-unary-binary) multipliers that seamlessly integrates with the NVDLA (NVIDIA's open-source DLA for accelerating CNNs) while maintaining dataflow compliance and boosting hardware efficiency. Analysis across various datapath granularities shows that for INT8 precision in 45nm CMOS, Tempus Core's PE cell unit (PCU) yields 59.3% and 15.3% reductions in area and power consumption, respectively, over NVDLA's CMAC unit. Considering a 16x16 PE array in Tempus Core, area and power improves by 75% and 62%, respectively, while delivering 5x and 4x iso-area throughput improvements for INT8 and INT4 precisions. Post-place and route analysis of Tempus Core's PCU shows that the 16x4 PE array for INT4 precision in 45nm CMOS requires only 0.017 mm^2 die area and consumes only 6.2mW of total power. We demonstrate that area-power efficient unary-based hardware can be seamlessly integrated into conventional DLAs, paving the path for efficient unary hardware for edge AI inference.