Picture for Nishant Balepur

Nishant Balepur

(Im)Paired Programming: Coding Agents Improve Productivity but Harm Understanding

Add code
Jul 29, 2026
Viaarxiv icon

Measuring User's Mental Models of Speech Translation in Human-AI Collaboration

Add code
Jun 23, 2026
Viaarxiv icon

DRACULA: Hunting for the Actions Users Want Deep Research Agents to Execute

Add code
Apr 26, 2026
Viaarxiv icon

Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users

Add code
Mar 17, 2026
Viaarxiv icon

BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks

Add code
Feb 05, 2026
Viaarxiv icon

AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite

Add code
Oct 24, 2025
Viaarxiv icon

Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers

Add code
Oct 09, 2025
Figure 1 for Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers
Figure 2 for Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers
Figure 3 for Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers
Figure 4 for Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers
Viaarxiv icon

Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above

Add code
Feb 19, 2025
Figure 1 for Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
Figure 2 for Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
Figure 3 for Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
Figure 4 for Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
Viaarxiv icon

Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User Personas

Add code
Jan 20, 2025
Figure 1 for Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User Personas
Figure 2 for Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User Personas
Figure 3 for Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User Personas
Figure 4 for Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User Personas
Viaarxiv icon

Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can't Answer?

Add code
Oct 20, 2024
Viaarxiv icon