arXiv cs.CLPaper
Semantic Bandits: In-Context Exploration-Exploitation is Biased by Semantic Priors
LLMs don't explore optimally in decision tasks because language priors overwhelm the actual reward signal. If you're deploying agents that need to balance exploration and exploitation, semantic priming can sabotage you. Rename your actions to be semantically neutral and see how it changes behavior.