科研成果通报:於佩雯博士研究成果被国际权威期刊 INFORMS Journal on Computing 正式接收
发布者:殳妮 发布时间:2026-09-22 浏览次数:10
2026年9月,苏州大学商学院智能商务与管理系於佩雯博士与赖启东、田原、刘光梧合作的论文“Solving Markov Decision Processes via Largest-size Average Estimator”被国际权威期刊 INFORMS Journal on Computing正式接收。INFORMS Journal on Computing是运筹学与管理科学领域的重要国际学术期刊,为UTD24期刊。
In September 2026, Dr. Peiwen Yu from the Department of Intelligent Business and Management at the Business School of Soochow University, together with Qidong Lai, Yuan Tian, and Guangwu Liu, had their paper, “Solving Markov Decision Processes via Largest-size Average Estimator,” accepted by INFORMS Journal on Computing. The journal is a leading international journal in operations research and management science and is included in the UTD24 journal list.
本研究聚焦于马尔可夫决策过程(Markov Decision Processes, MDPs)中的仿真优化问题。在许多复杂随机决策问题中,决策者需要通过有限的仿真样本对不同决策的价值进行估计,并据此选择最优决策。针对这一问题,论文提出了一种新的Largest-Size Average(LSA)估计方法。该方法利用自适应仿真过程中不同决策获得的样本量信息,以获得最大样本量的决策作为价值估计的主要依据。论文从理论上研究了LSA估计量的统计性质,并将其嵌入自适应多阶段抽样(Adaptive Multistage Sampling, AMS)框架,用于求解有限马尔可夫决策过程。研究结果表明,LSA方法能够有效利用自适应抽样过程中产生的样本分配信息,在多个实验场景下改善价值估计的准确性,并进一步支持高质量的决策与策略选择。该研究为利用有限仿真资源求解复杂随机动态决策问题提供了一种新的统计估计思路,也为仿真优化、强化学习和序贯决策等领域的相关研究提供了理论与方法参考。
This study focuses on simulation optimization for Markov decision processes (MDPs), where decision makers often need to estimate the values of alternative decisions using a limited simulation budget and select the best decision accordingly. The paper proposes a new Largest-Size Average (LSA) estimator that exploits sample-allocation information generated during adaptive simulation and uses the decision receiving the largest number of samples as the basis for value estimation. The paper establishes theoretical properties of the LSA estimator and incorporates it into an Adaptive Multistage Sampling (AMS) framework for solving finite-horizon MDPs. The results show that LSA can effectively use adaptive sample-allocation information to improve value-estimation accuracy and support better decision and policy selection. This study provides a new approach to solving complex stochastic dynamic decision problems with limited simulation resources and contributes to research on simulation optimization, reinforcement learning, and sequential decision-making.
於佩雯,管理学博士,现任苏州大学商学院智能商务与管理系讲师。2023年毕业于香港城市大学管理科学专业并获博士学位,此前分别于香港科技大学和南京大学获得硕士及学士学位。主要研究方向包括随机仿真与优化、机器学习以及这些方法在金融工程和风险管理中的应用。
Peiwen Yu holds a Ph.D. in Management Sciences and is currently a Lecturer in the Department of Intelligent Business and Management at the Business School of Soochow University. She received her Ph.D. from City University of Hong Kong in 2023, after earning her master’s degree from the Hong Kong University of Science and Technology and bachelor’s degree from Nanjing University. Her research interests include stochastic simulation and optimization, machine learning, and their applications in financial engineering and risk management.