Picture for Chengshuai Shi

Chengshuai Shi

LeAct: Learning to Reason from Expert Actions

Add code
Jul 23, 2026
Viaarxiv icon

The Past Is Prologue: A Plug-in Controller for Selective Updates in Sequentially Evolving LLM Memory

Add code
Jun 30, 2026
Viaarxiv icon

Efficient Multi-objective Prompt Optimization via Pure-exploration Bandits

Add code
May 14, 2026
Viaarxiv icon

Continual Harness: Online Adaptation for Self-Improving Foundation Agents

Add code
May 11, 2026
Viaarxiv icon

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis

Add code
May 19, 2025
Viaarxiv icon

Cost-Aware Optimal Pairwise Pure Exploration

Add code
Mar 10, 2025
Viaarxiv icon

Transformers as Game Players: Provable In-context Game-playing Capabilities of Pre-trained Models

Add code
Oct 13, 2024
Viaarxiv icon

Building Math Agents with Multi-Turn Iterative Preference Learning

Add code
Sep 04, 2024
Figure 1 for Building Math Agents with Multi-Turn Iterative Preference Learning
Figure 2 for Building Math Agents with Multi-Turn Iterative Preference Learning
Figure 3 for Building Math Agents with Multi-Turn Iterative Preference Learning
Figure 4 for Building Math Agents with Multi-Turn Iterative Preference Learning
Viaarxiv icon

Best Arm Identification for Prompt Learning under a Limited Budget

Add code
Feb 20, 2024
Viaarxiv icon

Harnessing the Power of Federated Learning in Federated Contextual Bandits

Add code
Dec 26, 2023
Figure 1 for Harnessing the Power of Federated Learning in Federated Contextual Bandits
Figure 2 for Harnessing the Power of Federated Learning in Federated Contextual Bandits
Figure 3 for Harnessing the Power of Federated Learning in Federated Contextual Bandits
Figure 4 for Harnessing the Power of Federated Learning in Federated Contextual Bandits
Viaarxiv icon