Picture for Youngsoo Jang

Youngsoo Jang

A Regret Minimization Framework on Preference Learning in Large Language Models

Add code
Jun 08, 2026
Viaarxiv icon

SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced Safety

Add code
May 26, 2025
Viaarxiv icon

Semantic Skill Grounding for Embodied Instruction-Following in Cross-Domain Environments

Add code
Aug 02, 2024
Viaarxiv icon

Reinforcement Learning from Reflective Feedback (RLRF): Aligning and Improving LLMs via Fine-Grained Self-Reflection

Add code
Mar 21, 2024
Figure 1 for Reinforcement Learning from Reflective Feedback (RLRF): Aligning and Improving LLMs via Fine-Grained Self-Reflection
Figure 2 for Reinforcement Learning from Reflective Feedback (RLRF): Aligning and Improving LLMs via Fine-Grained Self-Reflection
Figure 3 for Reinforcement Learning from Reflective Feedback (RLRF): Aligning and Improving LLMs via Fine-Grained Self-Reflection
Figure 4 for Reinforcement Learning from Reflective Feedback (RLRF): Aligning and Improving LLMs via Fine-Grained Self-Reflection
Viaarxiv icon

LobsDICE: Offline Imitation Learning from Observation via Stationary Distribution Correction Estimation

Add code
Feb 28, 2022
Figure 1 for LobsDICE: Offline Imitation Learning from Observation via Stationary Distribution Correction Estimation
Figure 2 for LobsDICE: Offline Imitation Learning from Observation via Stationary Distribution Correction Estimation
Figure 3 for LobsDICE: Offline Imitation Learning from Observation via Stationary Distribution Correction Estimation
Viaarxiv icon