Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Ting-Yun Chang

Value-Aware Stochastic KV Cache Eviction for Reasoning Models

Jun 02, 2026

Ting-Yun Chang, Harvey Yiyun Fu, Deqing Fu, Chenghao Yang, Jesse Thomason, Robin Jia

Abstract:Reasoning models improve accuracy through extended chains of thought, but their long outputs create a memory and compute bottleneck. KV cache eviction methods reduce this cost by evicting unimportant key-value pairs from the cache, yet they often yield worse accuracy than selection-based sparse attention alternatives, which keep the full KV cache. We identify key factors crucial to KV cache eviction accuracy. First, a small fraction of value states have abnormally large magnitudes, and evicting them causes catastrophic failure where models enter repetitive reasoning loops. Second, introducing stochasticity during eviction improves accuracy by increasing cache diversity. Based on these findings, we propose Value-aware Stochastic KV Cache Eviction (VaSE), a training-free recipe that protects large-magnitude value states and promotes diverse eviction decisions. Across six reasoning tasks, Qwen3 models using VaSE with 4x KV cache compression yield higher average accuracies than SOTA selection method at the same sparsity, while outperforming the strongest eviction method by more than 4%. Overall, VaSE bridges the gap between efficiency and accuracy, supporting FlashAttention2 and enabling a static memory footprint for reasoning models.

* Codes: https://github.com/terarachang/VaSE

Via

Access Paper or Ask Questions

Cosmos 3: Omnimodal World Models for Physical AI

Jun 01, 2026

Aditi, Niket Agarwal, Arslan Ali, Jon Allen, Martin Antolini, Adeline Aubame, Alisson Azzolini, Junjie Bai, Maciej Bala, Yogesh Balaji(+281 more)

Abstract:We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-transformers architecture. By supporting highly flexible input-output configurations, Cosmos 3 seamlessly unifies critical modalities for Physical AI -- effectively subsuming vision-language models, video generators, world simulators, and world-action models into a single framework. Our evaluation demonstrates that Cosmos 3 establishes a new state-of-the-art across a diverse suite of understanding and generation tasks, demonstrating omnimodal world models as scalable, general-purpose backbones for embodied agents. Our post-trained Cosmos 3 models were ranked as the best open-source Text-to-Image and Image-to-Video models by Artificial Analysis, and the best policy model by RoboArena at the time the technical report was written. To accelerate open research and deployment in Physical AI, we make our code, model checkpoints, curated synthetic datasets, and evaluation benchmark available under the Linux Foundation's OpenMDW-1.1 https://openmdw.ai/license/1-1/ License at https://github.com/nvidia/cosmos}{github.com/nvidia/cosmos and https://huggingface.co/collections/nvidia/cosmos3 . The project website is available at https://research.nvidia.com/labs/cosmos-lab/cosmos3 .

Via

Access Paper or Ask Questions

Vibe Checker: Aligning Code Evaluation with Human Preference

Oct 08, 2025

Ming Zhong, Xiang Zhou, Ting-Yun Chang, Qingze Wang, Nan Xu, Xiance Si, Dan Garrette, Shyam Upadhyay, Jeremiah Liu, Jiawei Han(+2 more)

Figure 1 for Vibe Checker: Aligning Code Evaluation with Human Preference

Figure 2 for Vibe Checker: Aligning Code Evaluation with Human Preference

Figure 3 for Vibe Checker: Aligning Code Evaluation with Human Preference

Figure 4 for Vibe Checker: Aligning Code Evaluation with Human Preference

Abstract:Large Language Models (LLMs) have catalyzed vibe coding, where users leverage LLMs to generate and iteratively refine code through natural language interactions until it passes their vibe check. Vibe check is tied to real-world human preference and goes beyond functionality: the solution should feel right, read cleanly, preserve intent, and remain correct. However, current code evaluation remains anchored to pass@k and captures only functional correctness, overlooking the non-functional instructions that users routinely apply. In this paper, we hypothesize that instruction following is the missing piece underlying vibe check that represents human preference in coding besides functional correctness. To quantify models' code instruction following capabilities with measurable signals, we present VeriCode, a taxonomy of 30 verifiable code instructions together with corresponding deterministic verifiers. We use the taxonomy to augment established evaluation suites, resulting in Vibe Checker, a testbed to assess both code instruction following and functional correctness. Upon evaluating 31 leading LLMs, we show that even the strongest models struggle to comply with multiple instructions and exhibit clear functional regression. Most importantly, a composite score of functional correctness and instruction following correlates the best with human preference, with the latter emerging as the primary differentiator on real-world programming tasks. Our work identifies core factors of the vibe check, providing a concrete path for benchmarking and developing models that better align with user preferences in coding.

* Preprint

Via

Access Paper or Ask Questions

When Parts are Greater Than Sums: Individual LLM Components Can Outperform Full Models

Jun 19, 2024

Ting-Yun Chang, Jesse Thomason, Robin Jia

Figure 1 for When Parts are Greater Than Sums: Individual LLM Components Can Outperform Full Models

Figure 2 for When Parts are Greater Than Sums: Individual LLM Components Can Outperform Full Models

Figure 3 for When Parts are Greater Than Sums: Individual LLM Components Can Outperform Full Models

Figure 4 for When Parts are Greater Than Sums: Individual LLM Components Can Outperform Full Models

Abstract:This paper studies in-context learning (ICL) by decomposing the output of large language models into the individual contributions of attention heads and MLPs (components). We observe curious components: good-performing ones that individually do well on a classification task, even when the model performs poorly; bad-performing ones that do much worse than chance; and label-biased components that always predict the same label. We find that component accuracies are well-correlated across different demonstration sets and perturbations of prompt templates, even when the full-model accuracy varies greatly. Based on our findings, we propose component reweighting, which learns to linearly re-scale the component activations from a few labeled examples. Given 24 labeled examples, our method improves by an average of 6.0% accuracy points over 24-shot ICL across 8 tasks on Llama-2-7B. Overall, this paper both enriches our understanding of ICL and provides a practical method for improvement by examining model internals.

Via

Access Paper or Ask Questions

Do Localization Methods Actually Localize Memorized Data in LLMs?

Nov 15, 2023

Ting-Yun Chang, Jesse Thomason, Robin Jia

Figure 1 for Do Localization Methods Actually Localize Memorized Data in LLMs?

Figure 2 for Do Localization Methods Actually Localize Memorized Data in LLMs?

Figure 3 for Do Localization Methods Actually Localize Memorized Data in LLMs?

Figure 4 for Do Localization Methods Actually Localize Memorized Data in LLMs?

Abstract:Large language models (LLMs) can memorize many pretrained sequences verbatim. This paper studies if we can locate a small set of neurons in LLMs responsible for memorizing a given sequence. While the concept of localization is often mentioned in prior work, methods for localization have never been systematically and directly evaluated; we address this with two benchmarking approaches. In our INJ Benchmark, we actively inject a piece of new information into a small subset of LLM weights and measure whether localization methods can identify these "ground truth" weights. In the DEL Benchmark, we study localization of pretrained data that LLMs have already memorized; while this setting lacks ground truth, we can still evaluate localization by measuring whether dropping out located neurons erases a memorized sequence from the model. We evaluate five localization methods on our two benchmarks, and both show similar rankings. All methods exhibit promising localization ability, especially for pruning-based methods, though the neurons they identify are not necessarily specific to a single memorized sequence.

Via

Access Paper or Ask Questions

Careful Data Curation Stabilizes In-context Learning

Dec 20, 2022

Ting-Yun Chang, Robin Jia

Abstract:In-context learning (ICL) enables large language models (LLMs) to perform new tasks by prompting them with a sequence of training examples. However, ICL is very sensitive to the choice of training examples: randomly sampling examples from a training set leads to high variance in performance. In this paper, we show that curating a carefully chosen subset of training data greatly stabilizes ICL performance. We propose two methods to choose training subsets, both of which score training examples individually and then select the highest-scoring ones. CondAcc scores a training example by its average ICL accuracy when combined with random training examples, while Datamodels learns a linear proxy model that estimates how the presence of each training example influences LLM accuracy. On average, CondAcc and Datamodels outperform sampling from the entire training set by 7.7% and 6.3%, respectively, across 5 tasks and two LLMs. Our analysis shows that stable subset examples are no more diverse than average, and are not outliers in terms of sequence length and perplexity.

Via

Access Paper or Ask Questions

CLiMB: A Continual Learning Benchmark for Vision-and-Language Tasks

Jun 18, 2022

Tejas Srinivasan, Ting-Yun Chang, Leticia Leonor Pinto Alva, Georgios Chochlakis, Mohammad Rostami, Jesse Thomason

Figure 1 for CLiMB: A Continual Learning Benchmark for Vision-and-Language Tasks

Figure 2 for CLiMB: A Continual Learning Benchmark for Vision-and-Language Tasks

Figure 3 for CLiMB: A Continual Learning Benchmark for Vision-and-Language Tasks

Figure 4 for CLiMB: A Continual Learning Benchmark for Vision-and-Language Tasks

Abstract:Current state-of-the-art vision-and-language models are evaluated on tasks either individually or in a multi-task setting, overlooking the challenges of continually learning (CL) tasks as they arrive. Existing CL benchmarks have facilitated research on task adaptation and mitigating "catastrophic forgetting", but are limited to vision-only and language-only tasks. We present CLiMB, a benchmark to study the challenge of learning multimodal tasks in a CL setting, and to systematically evaluate how upstream continual learning can rapidly generalize to new multimodal and unimodal tasks. CLiMB includes implementations of several CL algorithms and a modified Vision-Language Transformer (ViLT) model that can be deployed on both multimodal and unimodal tasks. We find that common CL methods can help mitigate forgetting during multimodal task learning, but do not enable cross-task knowledge transfer. We envision that CLiMB will facilitate research on a new class of CL algorithms for this challenging multimodal setting.

Via

Access Paper or Ask Questions

Rethinking Why Intermediate-Task Fine-Tuning Works

Sep 01, 2021

Ting-Yun Chang, Chi-Jen Lu

Figure 1 for Rethinking Why Intermediate-Task Fine-Tuning Works

Figure 2 for Rethinking Why Intermediate-Task Fine-Tuning Works

Figure 3 for Rethinking Why Intermediate-Task Fine-Tuning Works

Figure 4 for Rethinking Why Intermediate-Task Fine-Tuning Works

Abstract:Supplementary Training on Intermediate Labeled-data Tasks (STILTs) is a widely applied technique, which first fine-tunes the pretrained language models on an intermediate task before on the target task of interest. While STILTs is able to further improve the performance of pretrained language models, it is still unclear why and when it works. Previous research shows that those intermediate tasks involving complex inference, such as commonsense reasoning, work especially well for RoBERTa. In this paper, we discover that the improvement from an intermediate task could be orthogonal to it containing reasoning or other complex skills -- a simple real-fake discrimination task synthesized by GPT2 can benefit diverse target tasks. We conduct extensive experiments to study the impact of different factors on STILTs. These findings suggest rethinking the role of intermediate fine-tuning in the STILTs pipeline.

* Findings of EMNLP 2021

Via

Access Paper or Ask Questions

Go Beyond Plain Fine-tuning: Improving Pretrained Models for Social Commonsense

May 12, 2021

Ting-Yun Chang, Yang Liu, Karthik Gopalakrishnan, Behnam Hedayatnia, Pei Zhou, Dilek Hakkani-Tur

Figure 1 for Go Beyond Plain Fine-tuning: Improving Pretrained Models for Social Commonsense

Figure 2 for Go Beyond Plain Fine-tuning: Improving Pretrained Models for Social Commonsense

Figure 3 for Go Beyond Plain Fine-tuning: Improving Pretrained Models for Social Commonsense

Figure 4 for Go Beyond Plain Fine-tuning: Improving Pretrained Models for Social Commonsense

Abstract:Pretrained language models have demonstrated outstanding performance in many NLP tasks recently. However, their social intelligence, which requires commonsense reasoning about the current situation and mental states of others, is still developing. Towards improving language models' social intelligence, we focus on the Social IQA dataset, a task requiring social and emotional commonsense reasoning. Building on top of the pretrained RoBERTa and GPT2 models, we propose several architecture variations and extensions, as well as leveraging external commonsense corpora, to optimize the model for Social IQA. Our proposed system achieves competitive results as those top-ranking models on the leaderboard. This work demonstrates the strengths of pretrained language models, and provides viable ways to improve their performance for a particular task.

* SLT 2021

Via

Access Paper or Ask Questions

Incorporating Commonsense Knowledge Graph in Pretrained Models for Social Commonsense Tasks

May 12, 2021

Ting-Yun Chang, Yang Liu, Karthik Gopalakrishnan, Behnam Hedayatnia, Pei Zhou, Dilek Hakkani-Tur

Figure 1 for Incorporating Commonsense Knowledge Graph in Pretrained Models for Social Commonsense Tasks

Figure 2 for Incorporating Commonsense Knowledge Graph in Pretrained Models for Social Commonsense Tasks

Figure 3 for Incorporating Commonsense Knowledge Graph in Pretrained Models for Social Commonsense Tasks

Figure 4 for Incorporating Commonsense Knowledge Graph in Pretrained Models for Social Commonsense Tasks

Abstract:Pretrained language models have excelled at many NLP tasks recently; however, their social intelligence is still unsatisfactory. To enable this, machines need to have a more general understanding of our complicated world and develop the ability to perform commonsense reasoning besides fitting the specific downstream tasks. External commonsense knowledge graphs (KGs), such as ConceptNet, provide rich information about words and their relationships. Thus, towards general commonsense learning, we propose two approaches to \emph{implicitly} and \emph{explicitly} infuse such KGs into pretrained language models. We demonstrate our proposed methods perform well on SocialIQA, a social commonsense reasoning task, in both limited and full training data regimes.

* EMNLP2020 Workshop

Via

Access Paper or Ask Questions