Picture for Yintong Huo

Yintong Huo

Towards Risk-free AI Agent Deployment

Add code
Aug 17, 2026
Viaarxiv icon

LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures

Add code
Aug 15, 2026
Viaarxiv icon

CodeShrink: Adaptive Visual Compression for Efficient Multimodal Code Understanding

Add code
Jul 31, 2026
Viaarxiv icon

Graph Is the Verifier: Agentic Reinforcement Learning for Interprocedural Vulnerability Detection

Add code
Jul 29, 2026
Viaarxiv icon

Lost in the Flow with Code Talkers: Unveiling the Instruction-Tuning Tax of Large Language Models in Code Tasks

Add code
Jun 07, 2026
Viaarxiv icon

Making Embodied AI Reliable: A Community Agenda from Testing to Formal Verification

Add code
Jun 02, 2026
Viaarxiv icon

Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification

Add code
Jan 22, 2026
Viaarxiv icon

Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks

Add code
Aug 18, 2025
Viaarxiv icon

Next Edit Prediction: Learning to Predict Code Edits from Context and Interaction History

Add code
Aug 13, 2025
Viaarxiv icon

DesignBench: A Comprehensive Benchmark for MLLM-based Front-end Code Generation

Add code
Jun 06, 2025
Figure 1 for DesignBench: A Comprehensive Benchmark for MLLM-based Front-end Code Generation
Figure 2 for DesignBench: A Comprehensive Benchmark for MLLM-based Front-end Code Generation
Figure 3 for DesignBench: A Comprehensive Benchmark for MLLM-based Front-end Code Generation
Figure 4 for DesignBench: A Comprehensive Benchmark for MLLM-based Front-end Code Generation
Viaarxiv icon