A Study of 3D Representations for Human–Scene Interaction: From Classical to Neural

DSpace Repositorium (Manakin basiert)

Zur Kurzanzeige

dc.contributor.advisor Pons-Moll, Gerard (Prof. Dr.)
dc.contributor.author Mir, Mohamad Aymen
dc.date.accessioned 2026-07-29T14:32:41Z
dc.date.available 2026-07-29T14:32:41Z
dc.date.issued 2026-07-29
dc.identifier.uri http://hdl.handle.net/10900/181936
dc.identifier.uri http://nbn-resolving.org/urn:nbn:de:bsz:21-dspace-1819367 de_DE
dc.identifier.uri http://dx.doi.org/10.15496/publikation-123258
dc.description.abstract Creating virtual humans that look realistic and behave naturally in 3D environments is a central challenge in computer vision and graphics. At its core lies a fundamental question: how should we represent 3D humans and their environments to enable realistic interaction? Classical mesh and point cloud representations provide the geometric machinery for physics, collision detection, and path planning—but achieving photorealism requires complex material models, global illumination, and high polygon counts. 3D Gaussian Splatting offers photorealistic rendering at real-time rates from images alone—but provides no explicit surfaces, distance fields, or navigable structures. Can these missing geometric capabilities be recovered entirely from Gaussian fields? This thesis investigates this question in two parts: the first leverages the geometric strengths of classical representations for human appearance, capture, and motion synthesis, while the second demonstrates that the same capabilities can be extracted directly from raw Gaussian fields, enabling photorealistic human–scene interaction without classical geometry. The broader question of how to represent 3D humans is especially pressing as we enter the era of video diffusion models, which can generate photorealistic animations of humans from 2D data alone, challenging the very relevance of 3D representations. While the challenge is valid, we contend that 3D representations will remain essential for simulators, game engines, and digital twins, where geometry-consistent rendering, multi-view coherence, and explicit control over motion, lighting, and physics are required—capabilities that 2D generative models fundamentally cannot provide. But we also contend that for 3D representations to remain competitive, they must close the photorealism gap: the answer is not to abandon 3D, but to evolve it. Neural 3D representations can deliver both geometric control and photorealism. This thesis is our contribution toward that goal. In the first part, we use classical representations to address foundational problems. We introduce Pix2Surf, a method to transfer textures of clothing images to 3D garments worn on top of SMPL in real time, learning dense correspondences from garment silhouettes to UV maps using shape information alone. We then introduce HPS (Human POSEitioning System), a method to recover the full 3D pose of a human registered with a 3D scan of the surrounding environment using body-mounted sensors. HPS fuses camera-based self-localization with IMU-based body tracking, enabling capture across environments spanning 300–2500 m². Next, we introduce a method for synthesizing animator-guided human motion across 3D scenes by composing short-term motions in a canonical coordinate frame, generating long sequences of diverse actions without scene-specific training data. In the second part, we introduce 3D Gaussian Splatting as a unified representation for photorealistic human–scene interaction. We present InteractSplat, an end-to-end pipeline that reconstructs separate controllable human and object models from multi-view video and animates them together in diverse 3DGS environments, enabling long-horizon sequences where avatars navigate scenes, pick up objects, and set them down at new locations. We then present SALA, the first method to synthesize human interactions in diverse 3D environments using 3DGS as the sole underlying representation, extracting navigable structures from raw Gaussian fields and introducing differentiable contact refinement in Gaussian space. To improve the realism of composited avatars, we introduce RAGA, a ray-traced shadow casting formulation that computes physically plausible shadows entirely in Gaussian space at interactive rates. Finally, we present AHOY, which exploits the fact that 3DGS can be reconstructed from images alone—without the dense multi-view capture or depth sensors that mesh-based avatars require—to build complete, animatable avatars from in-the-wild YouTube video despite heavy occlusion. Using video diffusion priors, AHOY unlocks YouTube-scale footage as a practical source of 3D human assets that can be animated and composited into 3DGS scenes. Together, the contributions of this thesis demonstrate that 3DGS can serve as a complete replacement for classical representations in the human–scene animation pipeline—from navigation and locomotion to contact modeling and shadow casting—while delivering superior visual fidelity. en
dc.language.iso en de_DE
dc.publisher Universität Tübingen de_DE
dc.rights ubt-podno de_DE
dc.rights.uri http://tobias-lib.uni-tuebingen.de/doku/lic_ohne_pod.php?la=de de_DE
dc.rights.uri http://tobias-lib.uni-tuebingen.de/doku/lic_ohne_pod.php?la=en en
dc.subject.ddc 004 de_DE
dc.subject.other 3D Gaussian Splatting en
dc.subject.other human–scene interaction en
dc.subject.other neural rendering en
dc.subject.other 3D human reconstruction en
dc.subject.other motion synthesis en
dc.subject.other avatars en
dc.subject.other computer vision en
dc.subject.other computer graphics en
dc.title A Study of 3D Representations for Human–Scene Interaction: From Classical to Neural en
dc.type PhDThesis de_DE
dcterms.dateAccepted 2026-07-20
utue.publikation.fachbereich Informatik de_DE
utue.publikation.fakultaet 7 Mathematisch-Naturwissenschaftliche Fakultät de_DE
utue.publikation.noppn yes de_DE

Dateien:

Das Dokument erscheint in:

Zur Kurzanzeige