Picture for Baishakhi Ray

Baishakhi Ray

TRIM: Reducing AI-Generated CodeSlop via Agent Trajectory Minimization

Add code
Jul 20, 2026
Viaarxiv icon

Skills for the future software profession: beyond agentic AI!

Add code
Jun 23, 2026
Viaarxiv icon

TRAJEVAL: Decomposing Code Agent Trajectories for Fine-Grained Diagnosis

Add code
Mar 25, 2026
Viaarxiv icon

SWE-Spot: Building Small Repo-Experts with Repository-Centric Learning

Add code
Jan 29, 2026
Viaarxiv icon

Mechanics of Learned Reasoning 1: TempoBench, A Benchmark for Interpretable Deconstruction of Reasoning System Performance

Add code
Oct 31, 2025
Viaarxiv icon

AppForge: From Assistant to Independent Developer - Are GPTs Ready for Software Development?

Add code
Oct 09, 2025
Figure 1 for AppForge: From Assistant to Independent Developer - Are GPTs Ready for Software Development?
Figure 2 for AppForge: From Assistant to Independent Developer - Are GPTs Ready for Software Development?
Figure 3 for AppForge: From Assistant to Independent Developer - Are GPTs Ready for Software Development?
Figure 4 for AppForge: From Assistant to Independent Developer - Are GPTs Ready for Software Development?
Viaarxiv icon

Understanding Software Engineering Agents Through the Lens of Traceability: An Empirical Study

Add code
Jun 10, 2025
Figure 1 for Understanding Software Engineering Agents Through the Lens of Traceability: An Empirical Study
Figure 2 for Understanding Software Engineering Agents Through the Lens of Traceability: An Empirical Study
Figure 3 for Understanding Software Engineering Agents Through the Lens of Traceability: An Empirical Study
Figure 4 for Understanding Software Engineering Agents Through the Lens of Traceability: An Empirical Study
Viaarxiv icon

CrashFixer: A crash resolution agent for the Linux kernel

Add code
Apr 29, 2025
Figure 1 for CrashFixer: A crash resolution agent for the Linux kernel
Figure 2 for CrashFixer: A crash resolution agent for the Linux kernel
Figure 3 for CrashFixer: A crash resolution agent for the Linux kernel
Figure 4 for CrashFixer: A crash resolution agent for the Linux kernel
Viaarxiv icon

Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination

Add code
Mar 06, 2025
Figure 1 for Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination
Figure 2 for Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination
Figure 3 for Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination
Figure 4 for Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination
Viaarxiv icon

Recent Advances in Large Langauge Model Benchmarks against Data Contamination: From Static to Dynamic Evaluation

Add code
Feb 23, 2025
Viaarxiv icon