Data Science Wire

Generalized Gaussian Temporal Difference Error for Uncertainty-aware Reinforcement Learning

arXiv stat.ML5d4 min read

arXiv:2408.02295v4 Announce Type: replace-cross Abstract: Conventional uncertainty-aware temporal difference (TD) learning often models TD errors as zero-mean Gaussian. This assumption can miss the heavy-tailed and heteroscedastic residuals induced by bootstrapping and exploration. We introduce a state-conditioned shape head based on the Generalized Gaussian Distribution (GGD) and use a numerically modified GGD loss as an online surrogate for nonstationary TD residuals. We distinguish two mathematical facts that are sometimes conflated: the exact GGD likelihood is normalized for every $\beta>0

Read the full story at arXiv stat.ML

More in Machine Learning