A Study of 3D Representations for Human–Scene Interaction: From Classical to Neural

DSpace Repositorium (Manakin basiert)


Dateien:

Zitierfähiger Link (URI): http://hdl.handle.net/10900/181936
http://nbn-resolving.org/urn:nbn:de:bsz:21-dspace-1819367
http://dx.doi.org/10.15496/publikation-123258
Dokumentart: Dissertation
Erscheinungsdatum: 2026-07-29
Sprache: Englisch
Fakultät: 7 Mathematisch-Naturwissenschaftliche Fakultät
Fachbereich: Informatik
Gutachter: Pons-Moll, Gerard (Prof. Dr.)
Tag der mündl. Prüfung: 2026-07-20
DDC-Klassifikation: 004 - Informatik
Freie Schlagwörter:
3D Gaussian Splatting
human–scene interaction
neural rendering
3D human reconstruction
motion synthesis
avatars
computer vision
computer graphics
Lizenz: http://tobias-lib.uni-tuebingen.de/doku/lic_ohne_pod.php?la=de http://tobias-lib.uni-tuebingen.de/doku/lic_ohne_pod.php?la=en
Zur Langanzeige

Abstract:

Creating virtual humans that look realistic and behave naturally in 3D environments is a central challenge in computer vision and graphics. At its core lies a fundamental question: how should we represent 3D humans and their environments to enable realistic interaction? Classical mesh and point cloud representations provide the geometric machinery for physics, collision detection, and path planning—but achieving photorealism requires complex material models, global illumination, and high polygon counts. 3D Gaussian Splatting offers photorealistic rendering at real-time rates from images alone—but provides no explicit surfaces, distance fields, or navigable structures. Can these missing geometric capabilities be recovered entirely from Gaussian fields? This thesis investigates this question in two parts: the first leverages the geometric strengths of classical representations for human appearance, capture, and motion synthesis, while the second demonstrates that the same capabilities can be extracted directly from raw Gaussian fields, enabling photorealistic human–scene interaction without classical geometry. The broader question of how to represent 3D humans is especially pressing as we enter the era of video diffusion models, which can generate photorealistic animations of humans from 2D data alone, challenging the very relevance of 3D representations. While the challenge is valid, we contend that 3D representations will remain essential for simulators, game engines, and digital twins, where geometry-consistent rendering, multi-view coherence, and explicit control over motion, lighting, and physics are required—capabilities that 2D generative models fundamentally cannot provide. But we also contend that for 3D representations to remain competitive, they must close the photorealism gap: the answer is not to abandon 3D, but to evolve it. Neural 3D representations can deliver both geometric control and photorealism. This thesis is our contribution toward that goal. In the first part, we use classical representations to address foundational problems. We introduce Pix2Surf, a method to transfer textures of clothing images to 3D garments worn on top of SMPL in real time, learning dense correspondences from garment silhouettes to UV maps using shape information alone. We then introduce HPS (Human POSEitioning System), a method to recover the full 3D pose of a human registered with a 3D scan of the surrounding environment using body-mounted sensors. HPS fuses camera-based self-localization with IMU-based body tracking, enabling capture across environments spanning 300–2500 m². Next, we introduce a method for synthesizing animator-guided human motion across 3D scenes by composing short-term motions in a canonical coordinate frame, generating long sequences of diverse actions without scene-specific training data. In the second part, we introduce 3D Gaussian Splatting as a unified representation for photorealistic human–scene interaction. We present InteractSplat, an end-to-end pipeline that reconstructs separate controllable human and object models from multi-view video and animates them together in diverse 3DGS environments, enabling long-horizon sequences where avatars navigate scenes, pick up objects, and set them down at new locations. We then present SALA, the first method to synthesize human interactions in diverse 3D environments using 3DGS as the sole underlying representation, extracting navigable structures from raw Gaussian fields and introducing differentiable contact refinement in Gaussian space. To improve the realism of composited avatars, we introduce RAGA, a ray-traced shadow casting formulation that computes physically plausible shadows entirely in Gaussian space at interactive rates. Finally, we present AHOY, which exploits the fact that 3DGS can be reconstructed from images alone—without the dense multi-view capture or depth sensors that mesh-based avatars require—to build complete, animatable avatars from in-the-wild YouTube video despite heavy occlusion. Using video diffusion priors, AHOY unlocks YouTube-scale footage as a practical source of 3D human assets that can be animated and composited into 3DGS scenes. Together, the contributions of this thesis demonstrate that 3DGS can serve as a complete replacement for classical representations in the human–scene animation pipeline—from navigation and locomotion to contact modeling and shadow casting—while delivering superior visual fidelity.

Das Dokument erscheint in: