Learning Better Trading Dialogue Policies by Inferring Opponent Preferences (Extended Abstract)

Ioannis Efstathiou (Heriot-Watt University), Oliver Lemon (Heriot-Watt University)

Abstract

Negotiation dialogue capabilities have been identified as important in a variety of application areas. In prior work, it was shown how Reinforcement Learning (RL) agents can learn to use implicit and explicit manipulation moves in dialogue to manipulate their adversaries in non-cooperative trading games. We now show that trading dialogues are more successful when the RL agent builds an opponent model-an estimate of the (hidden) goals and preferences of the adversary-and learns how to exploit them. We explore a variety of state space representations for the preferences of trading adversaries, including one based on Conditional Preference Networks (CP-NETS), used for the first time in RL. We show that representing adversary preferences leads to significant improvements in trading success rates.