Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Zirui Zheng

Seek-and-Solve: Benchmarking MLLMs for Visual Clue-Driven Reasoning in Daily Scenarios

Apr 16, 2026

Xiaomin Li, Tala Wang, Zichen Zhong, Ying Zhang, Zirui Zheng, Takashi Isobe, Dezhuang Li, Huchuan Lu, You He, Xu Jia

Abstract:Daily scenarios are characterized by visual richness, requiring Multimodal Large Language Models (MLLMs) to filter noise and identify decisive visual clues for accurate reasoning. Yet, current benchmarks predominantly aim at evaluating MLLMs' pre-existing knowledge or perceptual understanding, often neglecting the critical capability of reasoning. To bridge this gap, we introduce DailyClue, a benchmark designed for visual clue-driven reasoning in daily scenarios. Our construction is guided by two core principles: (1) strict grounding in authentic daily activities, and (2) challenging query design that necessitates more than surface-level perception. Instead of simple recognition, our questions compel MLLMs to actively explore suitable visual clues and leverage them for subsequent reasoning. To this end, we curate a comprehensive dataset spanning four major daily domains and 16 distinct subtasks. Comprehensive evaluation across MLLMs and agentic models underscores the formidable challenge posed by our benchmark. Our analysis reveals several critical insights, emphasizing that the accurate identification of visual clues is essential for robust reasoning.

* Accepted by ACL Findings 2026. Project page: https://xiaominli1020.github.io/DailyClue/

Via

Access Paper or Ask Questions

CCL-LGS: Contrastive Codebook Learning for 3D Language Gaussian Splatting

May 26, 2025

Lei Tian, Xiaomin Li, Liqian Ma, Hefei Huang, Zirui Zheng, Hao Yin, Taiqing Li, Huchuan Lu, Xu Jia

Figure 1 for CCL-LGS: Contrastive Codebook Learning for 3D Language Gaussian Splatting

Figure 2 for CCL-LGS: Contrastive Codebook Learning for 3D Language Gaussian Splatting

Figure 3 for CCL-LGS: Contrastive Codebook Learning for 3D Language Gaussian Splatting

Figure 4 for CCL-LGS: Contrastive Codebook Learning for 3D Language Gaussian Splatting

Abstract:Recent advances in 3D reconstruction techniques and vision-language models have fueled significant progress in 3D semantic understanding, a capability critical to robotics, autonomous driving, and virtual/augmented reality. However, methods that rely on 2D priors are prone to a critical challenge: cross-view semantic inconsistencies induced by occlusion, image blur, and view-dependent variations. These inconsistencies, when propagated via projection supervision, deteriorate the quality of 3D Gaussian semantic fields and introduce artifacts in the rendered outputs. To mitigate this limitation, we propose CCL-LGS, a novel framework that enforces view-consistent semantic supervision by integrating multi-view semantic cues. Specifically, our approach first employs a zero-shot tracker to align a set of SAM-generated 2D masks and reliably identify their corresponding categories. Next, we utilize CLIP to extract robust semantic encodings across views. Finally, our Contrastive Codebook Learning (CCL) module distills discriminative semantic features by enforcing intra-class compactness and inter-class distinctiveness. In contrast to previous methods that directly apply CLIP to imperfect masks, our framework explicitly resolves semantic conflicts while preserving category discriminability. Extensive experiments demonstrate that CCL-LGS outperforms previous state-of-the-art methods. Our project page is available at https://epsilontl.github.io/CCL-LGS/.

Via

Access Paper or Ask Questions

OASIS: Open Agent Social Interaction Simulations with One Million Agents

Nov 26, 2024

Ziyi Yang, Zaibin Zhang, Zirui Zheng, Yuxian Jiang, Ziyue Gan, Zhiyu Wang, Zijian Ling, Jinsong Chen, Martz Ma, Bowen Dong(+13 more)

Figure 1 for OASIS: Open Agent Social Interaction Simulations with One Million Agents

Figure 2 for OASIS: Open Agent Social Interaction Simulations with One Million Agents

Figure 3 for OASIS: Open Agent Social Interaction Simulations with One Million Agents

Figure 4 for OASIS: Open Agent Social Interaction Simulations with One Million Agents

Abstract:There has been a growing interest in enhancing rule-based agent-based models (ABMs) for social media platforms (i.e., X, Reddit) with more realistic large language model (LLM) agents, thereby allowing for a more nuanced study of complex systems. As a result, several LLM-based ABMs have been proposed in the past year. While they hold promise, each simulator is specifically designed to study a particular scenario, making it time-consuming and resource-intensive to explore other phenomena using the same ABM. Additionally, these models simulate only a limited number of agents, whereas real-world social media platforms involve millions of users. To this end, we propose OASIS, a generalizable and scalable social media simulator. OASIS is designed based on real-world social media platforms, incorporating dynamically updated environments (i.e., dynamic social networks and post information), diverse action spaces (i.e., following, commenting), and recommendation systems (i.e., interest-based and hot-score-based). Additionally, OASIS supports large-scale user simulations, capable of modeling up to one million users. With these features, OASIS can be easily extended to different social media platforms to study large-scale group phenomena and behaviors. We replicate various social phenomena, including information spreading, group polarization, and herd effects across X and Reddit platforms. Moreover, we provide observations of social phenomena at different agent group scales. We observe that the larger agent group scale leads to more enhanced group dynamics and more diverse and helpful agents' opinions. These findings demonstrate OASIS's potential as a powerful tool for studying complex systems in digital environments.

Via

Access Paper or Ask Questions

OASIS: Open Agents Social Interaction Simulations on One Million Agents

Nov 21, 2024

Ziyi Yang, Zaibin Zhang, Zirui Zheng, Yuxian Jiang, Ziyue Gan, Zhiyu Wang, Zijian Ling, Jinsong Chen, Martz Ma, Bowen Dong(+12 more)

Figure 1 for OASIS: Open Agents Social Interaction Simulations on One Million Agents

Figure 2 for OASIS: Open Agents Social Interaction Simulations on One Million Agents

Figure 3 for OASIS: Open Agents Social Interaction Simulations on One Million Agents

Figure 4 for OASIS: Open Agents Social Interaction Simulations on One Million Agents

Via

Access Paper or Ask Questions

A Multi-Scale Attentive Transformer for Multi-Instrument Symbolic Music Generation

May 26, 2023

Xipin Wei, Junhui Chen, Zirui Zheng, Li Guo, Lantian Li, Dong Wang

Figure 1 for A Multi-Scale Attentive Transformer for Multi-Instrument Symbolic Music Generation

Figure 2 for A Multi-Scale Attentive Transformer for Multi-Instrument Symbolic Music Generation

Figure 3 for A Multi-Scale Attentive Transformer for Multi-Instrument Symbolic Music Generation

Figure 4 for A Multi-Scale Attentive Transformer for Multi-Instrument Symbolic Music Generation

Abstract:Recently, multi-instrument music generation has become a hot topic. Different from single-instrument generation, multi-instrument generation needs to consider inter-track harmony besides intra-track coherence. This is usually achieved by composing note segments from different instruments into a signal sequence. This composition could be on different scales, such as note, bar, or track. Most existing work focuses on a particular scale, leading to a shortage in modeling music with diverse temporal and track dependencies. This paper proposes a multi-scale attentive Transformer model to improve the quality of multi-instrument generation. We first employ multiple Transformer decoders to learn multi-instrument representations of different scales and then design an attentive mechanism to fuse the multi-scale information. Experiments conducted on SOD and LMD datasets show that our model improves both quantitative and qualitative performance compared to models based on single-scale information. The source code and some generated samples can be found at https://github.com/HaRry-qaq/MSAT.

* to be published in INTERSPEECH 2023

Via

Access Paper or Ask Questions