Picture for Jin Chen

Jin Chen

StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows

Add code
Aug 18, 2026
Viaarxiv icon

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents

Add code
Aug 12, 2026
Viaarxiv icon

Zero-Mem: Zero-Token Memory Operations for LLM Agents

Add code
Jul 31, 2026
Viaarxiv icon

A Causality-aware Infer-diagnose-refine Framework for Test-time Modality Adaptation in VLA Models

Add code
Jul 28, 2026
Viaarxiv icon

MSQA: A Natively Sourced Multilingual and Multicultural SimpleQA Benchmark

Add code
Jul 01, 2026
Viaarxiv icon

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields

Add code
Jun 09, 2026
Viaarxiv icon

Expand More, Shrink Less: Shaping Effective-Rank Dynamics for Dense Scaling in Recommendation

Add code
May 22, 2026
Viaarxiv icon

RankUp: Towards High-rank Representations for Large Scale Advertising Recommender Systems

Add code
Apr 21, 2026
Viaarxiv icon

LogicPoison: Logical Attacks on Graph Retrieval-Augmented Generation

Add code
Apr 03, 2026
Viaarxiv icon

The Language of Touch: Translating Vibrations into Text with Dual-Branch Learning

Add code
Mar 26, 2026
Viaarxiv icon