第一项利用当前估计,第二项奖励信息不足的方向。它把不确定性下乐观原则公理库不确定性下的乐观原则Optimism under uncertainty · Optimism in the face of uncertainty在仍与数据相容的模型中按最好可能价值行动,以探索消除不确定性。从每臂标量区间推广到参数椭球。
同样的椭球构造也是不确定性下乐观原则公理库不确定性下的乐观原则Optimism under uncertainty · Optimism in the face of uncertainty在仍与数据相容的模型中按最好可能价值行动,以探索消除不确定性。的几何版本。更复杂的广义线性、核化或非平稳模型会替换估计器与置信集合,但仍需分别证明覆盖、乐观选择和宽度累计,不能仅沿用 OFUL 的分数形式。
参考资料
Varsha Dani, Thomas Hayes, and Sham Kakade, “Stochastic Linear Optimization under Bandit Feedback,” COLT, 2008.
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári, “Improved Algorithms for Linear Stochastic Bandits,” NeurIPS, 2011.
Tor Lattimore and Csaba Szepesvári, Bandit Algorithms, Cambridge University Press, 2020.