Sebastian Stapf
PhD student in Computer Science | University of Bern
I work on AI research around world models, video generation, and visual intelligence.
Research
Currently a PhD student at the University of Bern. My recent work studies diffusion world models, memory/composition, controllable video, and whether video pretraining gives useful priors for general visual problem solving.
Selected Publications
-
World Model Self-Distillation: Training World Models to Solve General Tasks Sebastian Stapf, Pablo Acuaviva Huertos, Aram Davtyan, Paolo Favaro. arXiv, 2026.Distills caption-guided video generations into an image-and-task-conditioned world model, then improves it with VLM feedback. The goal is to make pretrained video generators useful for general task solving without curated task-execution videos.
-
Composition of Memory Experts for Diffusion World Models Sebastian Stapf, Pablo Acuaviva, Aram Davtyan, Paolo Favaro. ICLR, 2026.Builds diffusion world models with separate short-term, long-term, and spatial memory experts. The experts are combined with a product-of-experts formulation to keep generated futures consistent with past observations over longer horizons.
-
Rethinking Visual Intelligence: Insights from Video Pretraining Pablo Acuaviva, Aram Davtyan, Mariam Hassan, Sebastian Stapf, Ahmad Rahimi, Alexandre Alahi, Paolo Favaro. ICML, 2026. ARC Prize 2025 (Honorable Mention)Studies whether video diffusion pretraining provides useful priors for visual problem solving. Across ARC-style tasks, visual games, route planning, and cellular automata, video-pretrained models show strong data efficiency and adaptability.
-
GEM: A Generalizable Ego-Vision Multimodal World Model Mariam Hassan, Sebastian Stapf, Ahmad Rahimi, et al. CVPR, 2025.A controllable ego-vision world model that generates future RGB and depth video. It gives fine-grained control over ego-motion, object dynamics, human poses, and scene composition for diverse long-horizon scenarios.
-
PViT-6D: Overclocking Vision Transformers for 6D Pose Estimation Sebastian Stapf, Tobias Bauernfeind, Marco Riboldi. arXiv, 2023.Recasts 6D pose estimation as direct end-to-end regression with Vision Transformers and pose tokens. It also adds confidence prediction to make inference more interpretable and reliable.
Background
- PhD in Computer Science, University of Bern, 2024-present.
- M.Sc. Physics, LMU Munich, 2024.
- B.Sc. Physics, LMU Munich, 2021.
- Previously worked on 6-DoF pose estimation for in-vehicle augmented reality at BMW.