Picture for Zhi Zhang

Zhi Zhang

Conditional Evaluation of Language Models with Cheap Auxiliary Signals

Add code
Aug 17, 2026
Viaarxiv icon

Personalizing Large Language Model Agents with Small Policy Models

Add code
Jul 31, 2026
Viaarxiv icon

Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents

Add code
Jul 15, 2026
Viaarxiv icon

MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling

Add code
Jun 11, 2026
Viaarxiv icon

TriLens: Per-Layer Logit-Lens Entropy for White-Box Hallucination Detection

Add code
May 31, 2026
Viaarxiv icon

The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence

Add code
May 26, 2026
Viaarxiv icon

Agent Q-Mix: Selecting the Right Action for LLM Multi-Agent Systems through Reinforcement Learning

Add code
Apr 01, 2026
Viaarxiv icon

veScale-FSDP: Flexible and High-Performance FSDP at Scale

Add code
Feb 25, 2026
Viaarxiv icon

ExpertWeaver: Unlocking the Inherent MoE in Dense LLMs with GLU Activation Patterns

Add code
Feb 17, 2026
Viaarxiv icon

Train Less, Learn More: Adaptive Efficient Rollout Optimization for Group-Based Reinforcement Learning

Add code
Feb 15, 2026
Viaarxiv icon