Picture for Pablo Samuel Castro

Pablo Samuel Castro

A Comedy of Estimators: On KL Regularization in RL Training of LLMs

Add code
Dec 26, 2025
Viaarxiv icon

ARM-FM: Automated Reward Machines via Foundation Models for Compositional Reinforcement Learning

Add code
Oct 16, 2025
Viaarxiv icon

Asymmetric Proximal Policy Optimization: mini-critics boost LLM reasoning

Add code
Oct 02, 2025
Viaarxiv icon

Stable Gradients for Stable Learning at Scale in Deep Reinforcement Learning

Add code
Jun 18, 2025
Figure 1 for Stable Gradients for Stable Learning at Scale in Deep Reinforcement Learning
Figure 2 for Stable Gradients for Stable Learning at Scale in Deep Reinforcement Learning
Figure 3 for Stable Gradients for Stable Learning at Scale in Deep Reinforcement Learning
Figure 4 for Stable Gradients for Stable Learning at Scale in Deep Reinforcement Learning
Viaarxiv icon

Adaptive Accompaniment with ReaLchords

Add code
Jun 17, 2025
Viaarxiv icon

The Courage to Stop: Overcoming Sunk Cost Fallacy in Deep Reinforcement Learning

Add code
Jun 16, 2025
Figure 1 for The Courage to Stop: Overcoming Sunk Cost Fallacy in Deep Reinforcement Learning
Figure 2 for The Courage to Stop: Overcoming Sunk Cost Fallacy in Deep Reinforcement Learning
Figure 3 for The Courage to Stop: Overcoming Sunk Cost Fallacy in Deep Reinforcement Learning
Figure 4 for The Courage to Stop: Overcoming Sunk Cost Fallacy in Deep Reinforcement Learning
Viaarxiv icon

Measure gradients, not activations! Enhancing neuronal activity in deep reinforcement learning

Add code
May 29, 2025
Viaarxiv icon

Mind the GAP! The Challenges of Scale in Pixel-based Deep Reinforcement Learning

Add code
May 23, 2025
Viaarxiv icon

Meta-World+: An Improved, Standardized, RL Benchmark

Add code
May 16, 2025
Viaarxiv icon

Studying the Interplay Between the Actor and Critic Representations in Reinforcement Learning

Add code
Mar 08, 2025
Figure 1 for Studying the Interplay Between the Actor and Critic Representations in Reinforcement Learning
Figure 2 for Studying the Interplay Between the Actor and Critic Representations in Reinforcement Learning
Figure 3 for Studying the Interplay Between the Actor and Critic Representations in Reinforcement Learning
Figure 4 for Studying the Interplay Between the Actor and Critic Representations in Reinforcement Learning
Viaarxiv icon