Automatic Metrics as Reward Signals in Zeroth-Order Preference Optimization for Machine Translation
2026 (English)Independent thesis Advanced level (degree of Master (Two Years)), 80 credits / 120 HE credits
Student thesis
Abstract [en]
This thesis investigates the use of automatic evaluation metrics as reward signals inzeroth-order (ZO) optimization for machine translation. Traditional reinforcement learning from human feedback (RLHF) methods rely on learned reward models and first-orderoptimization, which introduce substantial computational and memory overhead. To address these limitations, this work explores a memory-efficient framework that combineszeroth-order optimization with automatic metrics, eliminating the need for an additionalreward model.The proposed approach integrates a multilingual large language model with SPSAbased zeroth-order optimization and employs BLEU and COMET as direct reward signals. Experiments are conducted on multiple machine translation datasets covering bothhigh-resource and low-resource language pairs, including English–German, English–Chinese, English–Nepali, and English-Swedish. The study evaluates translation qualityusing BLEU, COMET, and TER, while also analyzing optimization stability and computational efficiency.The results show that BLEU serves as an effective and computationally efficientreward signal in high-resource settings, achieving competitive or superior performancecompared to learned reward models while substantially reducing training time and memory consumption. In particular, BLEU-based optimization achieved competitive TER scoresacross language pairs. In contrast, COMET demonstrates advantages in low-resource settings such as English-Nepali, where semantic-aware evaluation provides more informative reward guidance than lexical overlap metrics. However, COMET introduces significantly higher computational cost and less stable optimization behavior.The findings further reveal that zeroth-order optimization is highly sensitive to reward quality and variance. Poorly aligned reward signals can lead to unstable trainingdynamics and model collapse, especially in linguistically distant or underrepresented language pairs. Moreover, reinforcement learning does not consistently outperform strongsupervised fine-tuning baselines in machine translation tasks.Overall, this thesis demonstrates that automatic evaluation metrics can function aspractical alternatives to learned reward models in memory-efficient reinforcement learning for machine translation. The work highlights the trade-offs between computationalefficiency, semantic fidelity, and optimization stability, providing insights into rewarddesign for zeroth-order optimization in large language models.
Place, publisher, year, edition, pages
2026.
National Category
Natural Language Processing
Identifiers
URN: urn:nbn:se:uu:diva-594225OAI: oai:DiVA.org:uu-594225DiVA, id: diva2:2086018
Educational program
Master Programme in Language Technology
Supervisors
Examiners
2026-07-142026-07-122026-07-14Bibliographically approved