Implementing Algorithms for the Multi-Armed-Bandit Problem
Let's suppose you are faced with a dilemma; there are k possible actions to take, and you have only T opportunities to take actions. Taking an action returns a reward R from an underlying reward distr