摘要
强化学习一词出自于行为心理学,这门学科把学习看作为反复试验的过程,以便把环境的状态映射为动作。强化学习的这种特性必然增加智能系统的困难性,学习时间增长。强化学习学习速度较慢的原因是没有明确的监督信号。因此,强化学习系统在与环境交互时不得不采取反复试验的方法依靠外部评价信号来调整自己的行为。智能系统必然经过很长的学习过程。如何提高强化学习速度是一个最重要的研究问题。该文从几个方面来讨论提高强化学习速度的方法。
The word,reinforcement learning,comes from behavior psychology.This subject takes learning as trial and er-ror process so as to map world state to the actions.This characteristic of reinforcement learning must increase learning difficulty for intelligent system and learning time also grows up.The reason of lower learning speed for reinforcement learning is due to that explicit supervised signal doesn't exist.Therefore reinforcement learning agent has to take trial and error method when interaction with environment and adjusts its behavior by external critic.The agent must experi-ence a long learning process.Thus how reinforcement learning speed is improved is a crucial problem.In this paper,the methods that improve reinforcement learning speed are discussed in many aspects.
出处
《计算机工程与应用》
CSCD
北大核心
2001年第22期38-40,共3页
Computer Engineering and Applications
基金
黑龙江省自然科学基金F9911
国防基础计划项目的资助