Digitala Vetenskapliga Arkivet

Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Automatic Metrics as Reward Signals in Zeroth-Order Preference Optimization for Machine Translation
Uppsala University, Disciplinary Domain of Humanities and Social Sciences, Faculty of Languages, Department of Linguistics and Philology.
2026 (English)Independent thesis Advanced level (degree of Master (Two Years)), 80 credits / 120 HE creditsStudent thesis
Abstract [en]

This thesis investigates the use of automatic evaluation metrics as reward signals inzeroth-order (ZO) optimization for machine translation. Traditional reinforcement learning from human feedback (RLHF) methods rely on learned reward models and first-orderoptimization, which introduce substantial computational and memory overhead. To address these limitations, this work explores a memory-efficient framework that combineszeroth-order optimization with automatic metrics, eliminating the need for an additionalreward model.The proposed approach integrates a multilingual large language model with SPSAbased zeroth-order optimization and employs BLEU and COMET as direct reward signals. Experiments are conducted on multiple machine translation datasets covering bothhigh-resource and low-resource language pairs, including English–German, English–Chinese, English–Nepali, and English-Swedish. The study evaluates translation qualityusing BLEU, COMET, and TER, while also analyzing optimization stability and computational efficiency.The results show that BLEU serves as an effective and computationally efficientreward signal in high-resource settings, achieving competitive or superior performancecompared to learned reward models while substantially reducing training time and memory consumption. In particular, BLEU-based optimization achieved competitive TER scoresacross language pairs. In contrast, COMET demonstrates advantages in low-resource settings such as English-Nepali, where semantic-aware evaluation provides more informative reward guidance than lexical overlap metrics. However, COMET introduces significantly higher computational cost and less stable optimization behavior.The findings further reveal that zeroth-order optimization is highly sensitive to reward quality and variance. Poorly aligned reward signals can lead to unstable trainingdynamics and model collapse, especially in linguistically distant or underrepresented language pairs. Moreover, reinforcement learning does not consistently outperform strongsupervised fine-tuning baselines in machine translation tasks.Overall, this thesis demonstrates that automatic evaluation metrics can function aspractical alternatives to learned reward models in memory-efficient reinforcement learning for machine translation. The work highlights the trade-offs between computationalefficiency, semantic fidelity, and optimization stability, providing insights into rewarddesign for zeroth-order optimization in large language models.

Place, publisher, year, edition, pages
2026.
National Category
Natural Language Processing
Identifiers
URN: urn:nbn:se:uu:diva-594225OAI: oai:DiVA.org:uu-594225DiVA, id: diva2:2086018
Educational program
Master Programme in Language Technology
Supervisors
Examiners
Available from: 2026-07-14 Created: 2026-07-12 Last updated: 2026-07-14Bibliographically approved

Open Access in DiVA

fulltext(14418 kB)53 downloads
File information
File name FULLTEXT01.pdfFile size 14418 kBChecksum SHA-512
ade4d80929340a0995d7367efa20023a14c04c4ea7a22d2376ee2525cc475aeb4c2778e30046b715fd714f6b3644a5c496ab3ee1d4eca3cf63014d20d4fb90aa
Type fulltextMimetype application/pdf

By organisation
Department of Linguistics and Philology
Natural Language Processing

Search outside of DiVA

GoogleGoogle Scholar
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 1959 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf