Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Daniel Dijkman

Notes-to-Self: Scratchpad Augmented VLAs for Memory Dependent Manipulation Tasks

Feb 24, 2026

Sanjay Haresh, Daniel Dijkman, Apratim Bhattacharyya, Roland Memisevic

Abstract:Many dexterous manipulation tasks are non-markovian in nature, yet little attention has been paid to this fact in the recent upsurge of the vision-language-action (VLA) paradigm. Although they are successful in bringing internet-scale semantic understanding to robotics, existing VLAs are primarily "stateless" and struggle with memory-dependent long horizon tasks. In this work, we explore a way to impart both spatial and temporal memory to a VLA by incorporating a language scratchpad. The scratchpad makes it possible to memorize task-specific information, such as object positions, and it allows the model to keep track of a plan and progress towards subgoals within that plan. We evaluate this approach on a split of memory-dependent tasks from the ClevrSkills environment, on MemoryBench, as well as on a challenging real-world pick-and-place task. We show that incorporating a language scratchpad significantly improves generalization on these tasks for both non-recurrent and recurrent models.

* To appear at ICRA 2026

Via

Access Paper or Ask Questions

Hybrid Training for Vision-Language-Action Models

Oct 01, 2025

Pietro Mazzaglia, Cansu Sancaktar, Markus Peschl, Daniel Dijkman

Abstract:Using Large Language Models to produce intermediate thoughts, a.k.a. Chain-of-thought (CoT), before providing an answer has been a successful recipe for solving complex language tasks. In robotics, similar embodied CoT strategies, generating thoughts before actions, have also been shown to lead to improved performance when using Vision-Language-Action models (VLAs). As these techniques increase the length of the model's generated outputs to include the thoughts, the inference time is negatively affected. Delaying an agent's actions in real-world executions, as in robotic manipulation settings, strongly affects the usability of a method, as tasks require long sequences of actions. However, is the generation of long chains-of-thought a strong prerequisite for achieving performance improvements? In this work, we explore the idea of Hybrid Training (HyT), a framework that enables VLAs to learn from thoughts and benefit from the associated performance gains, while enabling the possibility to leave out CoT generation during inference. Furthermore, by learning to conditionally predict a diverse set of outputs, HyT supports flexibility at inference time, enabling the model to either predict actions directly, generate thoughts or follow instructions. We evaluate the proposed method in a series of simulated benchmarks and real-world experiments.

Via

Access Paper or Ask Questions

ClevrSkills: Compositional Language and Visual Reasoning in Robotics

Nov 13, 2024

Sanjay Haresh, Daniel Dijkman, Apratim Bhattacharyya, Roland Memisevic

Figure 1 for ClevrSkills: Compositional Language and Visual Reasoning in Robotics

Figure 2 for ClevrSkills: Compositional Language and Visual Reasoning in Robotics

Figure 3 for ClevrSkills: Compositional Language and Visual Reasoning in Robotics

Figure 4 for ClevrSkills: Compositional Language and Visual Reasoning in Robotics

Abstract:Robotics tasks are highly compositional by nature. For example, to perform a high-level task like cleaning the table a robot must employ low-level capabilities of moving the effectors to the objects on the table, pick them up and then move them off the table one-by-one, while re-evaluating the consequently dynamic scenario in the process. Given that large vision language models (VLMs) have shown progress on many tasks that require high level, human-like reasoning, we ask the question: if the models are taught the requisite low-level capabilities, can they compose them in novel ways to achieve interesting high-level tasks like cleaning the table without having to be explicitly taught so? To this end, we present ClevrSkills - a benchmark suite for compositional reasoning in robotics. ClevrSkills is an environment suite developed on top of the ManiSkill2 simulator and an accompanying dataset. The dataset contains trajectories generated on a range of robotics tasks with language and visual annotations as well as multi-modal prompts as task specification. The suite includes a curriculum of tasks with three levels of compositional understanding, starting with simple tasks requiring basic motor skills. We benchmark multiple different VLM baselines on ClevrSkills and show that even after being pre-trained on large numbers of tasks, these models fail on compositional reasoning in robotics tasks.

* To appear at NeurIPS 2024 (D&B track)

Via

Access Paper or Ask Questions

Information-driven Affordance Discovery for Efficient Robotic Manipulation

May 06, 2024

Pietro Mazzaglia, Taco Cohen, Daniel Dijkman

Figure 1 for Information-driven Affordance Discovery for Efficient Robotic Manipulation

Figure 2 for Information-driven Affordance Discovery for Efficient Robotic Manipulation

Figure 3 for Information-driven Affordance Discovery for Efficient Robotic Manipulation

Figure 4 for Information-driven Affordance Discovery for Efficient Robotic Manipulation

Abstract:Robotic affordances, providing information about what actions can be taken in a given situation, can aid robotic manipulation. However, learning about affordances requires expensive large annotated datasets of interactions or demonstrations. In this work, we argue that well-directed interactions with the environment can mitigate this problem and propose an information-based measure to augment the agent's objective and accelerate the affordance discovery process. We provide a theoretical justification of our approach and we empirically validate the approach both in simulation and real-world tasks. Our method, which we dub IDA, enables the efficient discovery of visual affordances for several action primitives, such as grasping, stacking objects, or opening drawers, strongly improving data efficiency in simulation, and it allows us to learn grasping affordances in a small number of interactions, on a real-world setup with a UFACTORY XArm 6 robot arm.

* arXiv admin note: substantial text overlap with arXiv:2308.14915

Via

Access Paper or Ask Questions

Neural 5G Indoor Localization with IMU Supervision

Feb 15, 2024

Aleksandr Ermolov, Shreya Kadambi, Maximilian Arnold, Mohammed Hirzallah, Roohollah Amiri, Deepak Singh Mahendar Singh, Srinivas Yerramalli, Daniel Dijkman, Fatih Porikli, Taesang Yoo(+1 more)

Figure 1 for Neural 5G Indoor Localization with IMU Supervision

Figure 2 for Neural 5G Indoor Localization with IMU Supervision

Figure 3 for Neural 5G Indoor Localization with IMU Supervision

Figure 4 for Neural 5G Indoor Localization with IMU Supervision

Abstract:Radio signals are well suited for user localization because they are ubiquitous, can operate in the dark and maintain privacy. Many prior works learn mappings between channel state information (CSI) and position fully-supervised. However, that approach relies on position labels which are very expensive to acquire. In this work, this requirement is relaxed by using pseudo-labels during deployment, which are calculated from an inertial measurement unit (IMU). We propose practical algorithms for IMU double integration and training of the localization system. We show decimeter-level accuracy on simulated and challenging real data of 5G measurements. Our IMU-supervised method performs similarly to fully-supervised, but requires much less effort to deploy.

* IEEE GLOBECOM 2023

Via

Access Paper or Ask Questions

Uncertainty-driven Affordance Discovery for Efficient Robotics Manipulation

Aug 28, 2023

Pietro Mazzaglia, Taco Cohen, Daniel Dijkman

Figure 1 for Uncertainty-driven Affordance Discovery for Efficient Robotics Manipulation

Figure 2 for Uncertainty-driven Affordance Discovery for Efficient Robotics Manipulation

Figure 3 for Uncertainty-driven Affordance Discovery for Efficient Robotics Manipulation

Figure 4 for Uncertainty-driven Affordance Discovery for Efficient Robotics Manipulation

Abstract:Robotics affordances, providing information about what actions can be taken in a given situation, can aid robotics manipulation. However, learning about affordances requires expensive large annotated datasets of interactions or demonstrations. In this work, we show active learning can mitigate this problem and propose the use of uncertainty to drive an interactive affordance discovery process. We show that our method enables the efficient discovery of visual affordances for several action primitives, such as grasping, stacking objects, or opening drawers, strongly improving data efficiency and allowing us to learn grasping affordances on a real-world setup with an xArm 6 robot arm in a small number of trials.

* Presented at the GMPL workshop @ RSS 2023

Via

Access Paper or Ask Questions

Deconfounded Imitation Learning

Nov 04, 2022

Risto Vuorio, Johann Brehmer, Hanno Ackermann, Daniel Dijkman, Taco Cohen, Pim de Haan

Abstract:Standard imitation learning can fail when the expert demonstrators have different sensory inputs than the imitating agent. This is because partial observability gives rise to hidden confounders in the causal graph. We break down the space of confounded imitation learning problems and identify three settings with different data requirements in which the correct imitation policy can be identified. We then introduce an algorithm for deconfounded imitation learning, which trains an inference model jointly with a latent-conditional policy. At test time, the agent alternates between updating its belief over the latent and acting under the belief. We show in theory and practice that this algorithm converges to the correct interventional policy, solves the confounding issue, and can under certain assumptions achieve an asymptotically optimal imitation performance.

Via

Access Paper or Ask Questions