Digitala Vetenskapliga Arkivet

Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Adaptive Incentive Design for Dynamical Games with Unknown Dynamics: Model-Based and Model-Free Approaches
KTH, School of Engineering Sciences (SCI).
KTH, School of Engineering Sciences (SCI).
2026 (English)Independent thesis Basic level (degree of Bachelor), 10 credits / 15 HE creditsStudent thesis
Abstract [en]

This thesis studies incentive design in a dynamic multi-agent system, where agents act according to individual objectives that may differ from the objective of a system-level leader. The problem is investigated using a simplified two-zone thermal building model formulated as a linear-quadratic-Gaussian control problem. The objective is to design incentives that align the agents' Nash equilibrium behavior with the leader’s optimal cost.

As a reference, an analytical model-based incentive design method from existing literature is implemented for the case where the system dynamics are fully known. This serves to verify the implementation and provides a benchmark for comparison.

Two data-driven approaches are then considered. First, a recursive least squares (RLS) method is used to estimate the unknown system matrix A, after which incentives are computed using the estimated model. Second, a model-free reinforcement learning approach based on deep deterministic policy gradient (DDPG) is used to learn incentives directly from observed performance.

The results show that the RLS-based method provides stable convergence and achieves low cost error, although it relies on an accurate model estimate. The model-free DDPG approach is able to learn effective incentives, but exhibits slower and less stable convergence and requires more data to reach comparable performance.

Overall, the results highlight a trade-off between model-based and model-free approaches: model-based methods provide superior performance when the system structure is known or can be accurately estimated, while model-free methods remain useful in settings where the system dynamics are unknown or difficult to model.

 

Place, publisher, year, edition, pages
2026.
Series
TRITA-SCI-GRU ; 2026:113
Keywords [en]
Incentive design, Dynamical games, Multi-agent systems, Nash equilibrium, Reinforcement learning, Deep deterministic policy gradient, Recursive least squares
National Category
Mathematical sciences
Identifiers
URN: urn:nbn:se:kth:diva-384293OAI: oai:DiVA.org:kth-384293DiVA, id: diva2:2081081
Educational program
Master of Science in Engineering -Engineering Physics
Supervisors
Examiners
Available from: 2026-06-29 Created: 2026-06-29 Last updated: 2026-06-29Bibliographically approved

Open Access in DiVA

fulltext(906 kB)23 downloads
File information
File name FULLTEXT01.pdfFile size 906 kBChecksum SHA-512
c17a40157394fae7db905062766d9f7968c5578291093c05ee61cbc9c528da02aa0d1f48985dc58631a166ed9c577ad85e40e43546aa356bd67f0f1865a007ed
Type fulltextMimetype application/pdf

By organisation
School of Engineering Sciences (SCI)
Mathematical sciences

Search outside of DiVA

GoogleGoogle Scholar
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 910 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf