From Probabilistic to Verbalized Reasoning and Machine Learning

DSpace Repositorium (Manakin basiert)


Dateien:

Zitierfähiger Link (URI): http://hdl.handle.net/10900/182738
http://nbn-resolving.org/urn:nbn:de:bsz:21-dspace-1827384
http://dx.doi.org/10.15496/publikation-124052
Dokumentart: Dissertation
Erscheinungsdatum: 2026-08-26
Sprache: Englisch
Fakultät: 7 Mathematisch-Naturwissenschaftliche Fakultät
Fachbereich: Informatik
Gutachter: Hennig, Philipp (Prof. Dr.)
Tag der mündl. Prüfung: 2026-07-23
DDC-Klassifikation: 004 - Informatik
Freie Schlagwörter:
Machine Learning
Large Language Models
Verbalized Computing
Lizenz: http://tobias-lib.uni-tuebingen.de/doku/lic_ohne_pod.php?la=de http://tobias-lib.uni-tuebingen.de/doku/lic_ohne_pod.php?la=en
Zur Langanzeige

Abstract:

A fundamental question in machine learning is how to represent information and perform inference. Both tasks require reasoning under uncertainty: information must be encoded in representations that capture relevant structure, and inference must recover unobserved quantities from observed data. Probabilistic models address this through learned numerical representations with explicit distributional semantics. More recently, large language models have opened a complementary path: representing and processing information in natural language. From an information-theoretic perspective, language has been modeled as a probability distribution over sequences since Shannon’s foundational work; natural language is thus inherently a probabilistic medium, carrying uncertainty through its ambiguity and context-dependence. This thesis investigates both paradigms and traces a path from one to the other. In the first part, we develop tools for inference and learning in numerical representation spaces. We introduce an information-theoretic framework for hierarchical variational autoencoders that treats each latent layer as operating at a controllable point on a rate-distortion curve, enabling systematic trading of information between layers. We then investigate when these probabilistic representations generalize, decomposing VAE performance into generalization, amortization, and robustness gaps, and showing that synthetic data from diffusion models and overparameterization can improve training while also uncovering a double descent phenomenon in generative models. In the second part, we turn to inference and learning in natural language space. We first establish conceptual foundations by drawing structural parallels between large language models and modern computers, arguing that LLMs constitute a new kind of general-purpose computing system where language serves as both program and data representation. Building on this, we introduce Verbalized Machine Learning, a framework in which models are parameterized entirely by natural language prompts and trained through iterative text-based optimization, demonstrating that learning can operate directly in language space. Finally, we address a key limitation: the knowledge-sampling gap, where LLMs can describe probability distributions but fail to sample from them faithfully. We propose Verbalized Rejection Sampling, which adapts classical rejection sampling to operate entirely in natural language, and provide theoretical bounds on when it outperforms direct sampling. This last contribution bridges both parts of the thesis by showing that a classical probabilistic algorithm can be adapted to operate in language space while retaining theoretical guarantees. As large language models increasingly serve as general-purpose reasoning engines, verbalized computing, the practice of performing computation directly in natural language, is emerging as a new paradigm. Since natural language is inherently ambiguous, a central challenge for this paradigm is how to perform computation in it effectively and reliably. This thesis suggests that the rich toolkit of probabilistic methods, which has been developed precisely for handling uncertainty, can provide both inspiration and principled foundations for addressing this challenge.

Das Dokument erscheint in: