Picture for Chirag Nagpal

Chirag Nagpal

Inverse RL Helps Align AI by Imitating Humans

Add code
Jul 27, 2026
Viaarxiv icon

Discretizing Reward Models

Add code
Jun 19, 2026
Viaarxiv icon

The Llama 4 Herd: Architecture, Training, Evaluation, and Deployment Notes

Add code
Jan 15, 2026
Viaarxiv icon

Understanding challenges to the interpretation of disaggregated evaluations of algorithmic fairness

Add code
Jun 04, 2025
Viaarxiv icon

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models

Add code
Jan 08, 2025
Viaarxiv icon

InfAlign: Inference-aware language model alignment

Add code
Dec 27, 2024
Viaarxiv icon

Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Add code
Oct 10, 2024
Figure 1 for Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
Figure 2 for Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
Figure 3 for Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
Figure 4 for Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
Viaarxiv icon

Robust Preference Optimization through Reward Model Distillation

Add code
May 29, 2024
Viaarxiv icon

A Toolbox for Surfacing Health Equity Harms and Biases in Large Language Models

Add code
Mar 18, 2024
Viaarxiv icon

The Case for Globalizing Fairness: A Mixed Methods Study on Colonialism, AI, and Health in Africa

Add code
Mar 11, 2024
Figure 1 for The Case for Globalizing Fairness: A Mixed Methods Study on Colonialism, AI, and Health in Africa
Figure 2 for The Case for Globalizing Fairness: A Mixed Methods Study on Colonialism, AI, and Health in Africa
Viaarxiv icon