Digitala Vetenskapliga Arkivet

Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Ground Text-Generated Human into Executable Avatar Actions in Interactive Environments
Halmstad University, School of Information Technology.
Halmstad University, School of Information Technology.
2026 (English)Independent thesis Advanced level (professional degree), 20 credits / 30 HE creditsStudent thesis
Abstract [en]

Virtual avatars have gained increasing attention in Extended Reality (XR) research because of their potential to support social presence, communication, and engagement in immersive environments, where scene-aware and believable avatar behavior is desirable. This thesis investigates how text-generated human motion can be grounded into executable avatar actions in a Unity environment representing XR scenarios. To address this problem, the work integrates three separate modules consisting of intent reasoning with a Large Language Model (LLM), 3D affordance grounding, and text-to-motion generation.

The modular frontend-backend pipeline captures object semantics, point clouds, and avatar state in Unity, processes these inputs on a backend server, and returns generated motion for retargeted avatar execution. The system was evaluated on sitting and lying interactions with chairs, benches, and beds using latency measurements, Root Mean Square Error (RMSE), semantic action quality, and perceived motion quality.

The results show that the pipeline can translate scene data into executable avatar actions within an interactive response range, while remaining contact artifacts indicate that single-target grounding is not sufficient for fully natural avatar-object interaction. This study demonstrates both the feasibility and limitations of combining LLM reasoning, affordance detection, and text-to-motion generation for grounded avatar interaction. The findings highlight the need for stronger object-conditioned motion generation, explicit user intent, and richer interaction constraints to improve robustness in future XR avatar systems.

Place, publisher, year, edition, pages
2026.
Keywords [en]
Large Language Model, Machine Learning, Extended Reality, Skinned Multi-Person Linear, Virtual Reality
National Category
Computer Vision and Learning Systems
Identifiers
URN: urn:nbn:se:hh:diva-60162OAI: oai:DiVA.org:hh-60162DiVA, id: diva2:2091451
Subject / course
Computer science and engineering
Educational program
Computer Science and Engineering, 300 credits
Supervisors
Examiners
Available from: 2026-08-14 Created: 2026-08-12 Last updated: 2026-08-14Bibliographically approved

Open Access in DiVA

fulltext(19545 kB)30 downloads
File information
File name FULLTEXT02.pdfFile size 19545 kBChecksum SHA-512
60a25279a9a7aecee539703b147517d55fa1826d0901085f018f1e2d2959215cdf3c5aa7515383887e1ebd7871dbc2517af0ec29c4e63676c3f0679ee0113374
Type fulltextMimetype application/pdf

By organisation
School of Information Technology
Computer Vision and Learning Systems

Search outside of DiVA

GoogleGoogle Scholar
Total: 30 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 698 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf