Picture for Zichen Ding

Zichen Ding

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

Add code
Jul 30, 2026
Viaarxiv icon

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

Add code
Jul 16, 2026
Viaarxiv icon

MacAgentBench: Benchmarking AI Agents on Real-World macOS Desktop

Add code
Jun 21, 2026
Viaarxiv icon

Skill is Not One-Size-Fits-All: Model-Aware Skill Alignment for LLM Agents

Add code
May 29, 2026
Viaarxiv icon

Retrieval, Reward, and Training Protocols: What Matters in Training Search Agents?

Add code
May 27, 2026
Viaarxiv icon

OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis

Add code
Apr 16, 2026
Viaarxiv icon

Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale

Add code
Mar 26, 2026
Viaarxiv icon

OS-Themis: A Scalable Critic Framework for Generalist GUI Rewards

Add code
Mar 19, 2026
Viaarxiv icon

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions

Add code
Feb 05, 2026
Viaarxiv icon

TIDE: Trajectory-based Diagnostic Evaluation of Test-Time Improvement in LLM Agents

Add code
Feb 03, 2026
Viaarxiv icon