Picture for Alessio Tonioni

Alessio Tonioni

DataComp-VLM: Improved Open Datasets for Vision-Language Models

Add code
Jun 30, 2026
Viaarxiv icon

Editing Everything Everywhere All at Once

Add code
Jun 30, 2026
Viaarxiv icon

PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding

Add code
May 28, 2026
Viaarxiv icon

FullFlow: Upgrading Text-to-Image Flow Matching Models for Bidirectional Vision--Language Generation

Add code
May 19, 2026
Viaarxiv icon

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models

Add code
Apr 22, 2026
Viaarxiv icon

R-CoV: Region-Aware Chain-of-Verification for Alleviating Object Hallucinations in LVLMs

Add code
Apr 22, 2026
Viaarxiv icon

Shifting the Breaking Point of Flow Matching for Multi-Instance Editing

Add code
Feb 09, 2026
Viaarxiv icon

RefAM: Attention Magnets for Zero-Shot Referral Segmentation

Add code
Sep 26, 2025
Figure 1 for RefAM: Attention Magnets for Zero-Shot Referral Segmentation
Figure 2 for RefAM: Attention Magnets for Zero-Shot Referral Segmentation
Figure 3 for RefAM: Attention Magnets for Zero-Shot Referral Segmentation
Figure 4 for RefAM: Attention Magnets for Zero-Shot Referral Segmentation
Viaarxiv icon

Test-Time Visual In-Context Tuning

Add code
Mar 27, 2025
Figure 1 for Test-Time Visual In-Context Tuning
Figure 2 for Test-Time Visual In-Context Tuning
Figure 3 for Test-Time Visual In-Context Tuning
Figure 4 for Test-Time Visual In-Context Tuning
Viaarxiv icon

Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos

Add code
Mar 17, 2025
Figure 1 for Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos
Figure 2 for Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos
Figure 3 for Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos
Figure 4 for Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos
Viaarxiv icon