Picture for Edoardo Pona

Edoardo Pona

Calibrate-Then-Delegate: Safety Monitoring with Risk and Budget Guarantees via Model Cascades

Add code
Apr 15, 2026
Viaarxiv icon

Reinforcement Learning Fine-tuning of Language Models is Biased Towards More Extractable Features

Add code
Nov 07, 2023
Figure 1 for Reinforcement Learning Fine-tuning of Language Models is Biased Towards More Extractable Features
Figure 2 for Reinforcement Learning Fine-tuning of Language Models is Biased Towards More Extractable Features
Figure 3 for Reinforcement Learning Fine-tuning of Language Models is Biased Towards More Extractable Features
Figure 4 for Reinforcement Learning Fine-tuning of Language Models is Biased Towards More Extractable Features
Viaarxiv icon