Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Timo P. Gros

Per-Domain Generalizing Policies: On Validation Instances and Scaling Behavior

May 01, 2025

Timo P. Gros, Nicola J. Müller, Daniel Fiser, Isabel Valera, Verena Wolf, Jörg Hoffmann

Abstract:Recent work has shown that successful per-domain generalizing action policies can be learned. Scaling behavior, from small training instances to large test instances, is the key objective; and the use of validation instances larger than training instances is one key to achieve it. Prior work has used fixed validation sets. Here, we introduce a method generating the validation set dynamically, on the fly, increasing instance size so long as informative and feasible.We also introduce refined methodology for evaluating scaling behavior, generating test instances systematically to guarantee a given confidence in coverage performance for each instance size. In experiments, dynamic validation improves scaling behavior of GNN policies in all 9 domains used.

* 7 pages, 3 tables, 3 figures, 3 algorithms

Via

Access Paper or Ask Questions

Tracking the Race Between Deep Reinforcement Learning and Imitation Learning -- Extended Version

Aug 03, 2020

Timo P. Gros, Daniel Höller, Jörg Hoffmann, Verena Wolf

Figure 1 for Tracking the Race Between Deep Reinforcement Learning and Imitation Learning -- Extended Version

Figure 2 for Tracking the Race Between Deep Reinforcement Learning and Imitation Learning -- Extended Version

Figure 3 for Tracking the Race Between Deep Reinforcement Learning and Imitation Learning -- Extended Version

Figure 4 for Tracking the Race Between Deep Reinforcement Learning and Imitation Learning -- Extended Version

Abstract:Learning-based approaches for solving large sequential decision making problems have become popular in recent years. The resulting agents perform differently and their characteristics depend on those of the underlying learning approach. Here, we consider a benchmark planning problem from the reinforcement learning domain, the Racetrack, to investigate the properties of agents derived from different deep (reinforcement) learning approaches. We compare the performance of deep supervised learning, in particular imitation learning, to reinforcement learning for the Racetrack model. We find that imitation learning yields agents that follow more risky paths. In contrast, the decisions of deep reinforcement learning are more foresighted, i.e., avoid states in which fatal decisions are more likely. Our evaluations show that for this sequential decision making problem, deep reinforcement learning performs best in many aspects even though for imitation learning optimal decisions are considered.

* Extended Version of the Conference Paper published in the Proceedings of the 17th International Conference on Quantitative Evaluation of SysTems (QEST)

Via

Access Paper or Ask Questions