Picture for Unggi Lee

Unggi Lee

EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners

Add code
Aug 04, 2026
Viaarxiv icon

Not the Dimension, the Norm: What Matters in Gradient-Free Weight Perturbation of Language Models

Add code
Aug 03, 2026
Viaarxiv icon

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation

Add code
May 29, 2026
Viaarxiv icon

BuddyBench: A Privacy-Constrained Multi-Task Benchmark for Pediatric Social-Communication Personalization

Add code
May 27, 2026
Viaarxiv icon

Are Video Models Zero-Shot Learners and Reasoners in Education? EduVideoBench, A Knowledge-Skills-Attitude Benchmark for Educational Video Generation

Add code
May 26, 2026
Viaarxiv icon

LLMs Are Already Good Tutors: Training-Free Prompt Optimization for Pedagogical Math Tutoring

Add code
May 26, 2026
Viaarxiv icon

ISD-Agent-Bench: A Comprehensive Benchmark for Evaluating LLM-based Instructional Design Agents

Add code
Feb 11, 2026
Viaarxiv icon

Rewarding How Models Think Pedagogically: Integrating Pedagogical Reasoning and Thinking Rewards for LLMs in Education

Add code
Jan 21, 2026
Viaarxiv icon

Pedagogical Alignment for Vision-Language-Action Models: A Comprehensive Framework for Data, Architecture, and Evaluation in Education

Add code
Jan 20, 2026
Viaarxiv icon

OpenLearnLM Benchmark: A Unified Framework for Evaluating Knowledge, Skill, and Attitude in Educational Large Language Models

Add code
Jan 20, 2026
Viaarxiv icon