Agent learns optimal path via Q-learning. Watch exploration become exploitation as the Q-table fills with experience.