Picture for Yingfa Chen

Yingfa Chen

Frontier Language Models Struggle to Copy: Text Can Be Better Viewed in 2D

Add code
Jul 17, 2026
Viaarxiv icon

Rethinking the Role of Efficient Attention in Hybrid Architectures

Add code
Jun 13, 2026
Viaarxiv icon

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It

Add code
Jun 09, 2026
Viaarxiv icon

DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices

Add code
May 11, 2026
Viaarxiv icon

Student-in-the-Loop Chain-of-Thought Distillation via Generation-Time Selection

Add code
Apr 03, 2026
Viaarxiv icon

MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling

Add code
Feb 12, 2026
Viaarxiv icon

Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts

Add code
Jan 29, 2026
Viaarxiv icon

StateX: Enhancing RNN Recall via Post-training State Expansion

Add code
Sep 26, 2025
Viaarxiv icon

Cost-Optimal Grouped-Query Attention for Long-Context LLMs

Add code
Mar 12, 2025
Viaarxiv icon

Sparsing Law: Towards Large Language Models with Greater Activation Sparsity

Add code
Nov 04, 2024
Figure 1 for Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
Figure 2 for Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
Figure 3 for Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
Figure 4 for Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
Viaarxiv icon