Picture for Size Zheng

Size Zheng

Eric

SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inference

Add code
Aug 04, 2026
Viaarxiv icon

Teaching and Evaluating LLMs to Reason About Polymer Design Related Tasks

Add code
Jan 22, 2026
Viaarxiv icon

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding

Add code
Jun 12, 2025
Figure 1 for SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding
Figure 2 for SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding
Figure 3 for SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding
Figure 4 for SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding
Viaarxiv icon

MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production

Add code
May 19, 2025
Viaarxiv icon

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design

Add code
May 09, 2025
Figure 1 for MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design
Figure 2 for MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design
Figure 3 for MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design
Figure 4 for MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design
Viaarxiv icon

MoQa: Rethinking MoE Quantization with Multi-stage Data-model Distribution Awareness

Add code
Mar 27, 2025
Figure 1 for MoQa: Rethinking MoE Quantization with Multi-stage Data-model Distribution Awareness
Figure 2 for MoQa: Rethinking MoE Quantization with Multi-stage Data-model Distribution Awareness
Figure 3 for MoQa: Rethinking MoE Quantization with Multi-stage Data-model Distribution Awareness
Figure 4 for MoQa: Rethinking MoE Quantization with Multi-stage Data-model Distribution Awareness
Viaarxiv icon

Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts

Add code
Feb 27, 2025
Figure 1 for Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts
Figure 2 for Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts
Figure 3 for Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts
Figure 4 for Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts
Viaarxiv icon

ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference

Add code
Oct 28, 2024
Viaarxiv icon

Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Add code
Nov 07, 2023
Figure 1 for Atom: Low-bit Quantization for Efficient and Accurate LLM Serving
Figure 2 for Atom: Low-bit Quantization for Efficient and Accurate LLM Serving
Figure 3 for Atom: Low-bit Quantization for Efficient and Accurate LLM Serving
Figure 4 for Atom: Low-bit Quantization for Efficient and Accurate LLM Serving
Viaarxiv icon

HASCO: Towards Agile HArdware and Software CO-design for Tensor Computation

Add code
May 04, 2021
Figure 1 for HASCO: Towards Agile HArdware and Software CO-design for Tensor Computation
Figure 2 for HASCO: Towards Agile HArdware and Software CO-design for Tensor Computation
Figure 3 for HASCO: Towards Agile HArdware and Software CO-design for Tensor Computation
Figure 4 for HASCO: Towards Agile HArdware and Software CO-design for Tensor Computation
Viaarxiv icon