Digitala Vetenskapliga Arkivet

Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Evaluating General-Purpose Multimodal LLM for Indoor Localization on Floorplans
Malmö University, Faculty of Technology and Society (TS), Department of Computer Science and Media Technology (DVMT).
Malmö University, Faculty of Technology and Society (TS), Department of Computer Science and Media Technology (DVMT).
2026 (English)Independent thesis Basic level (degree of Bachelor), 10 credits / 15 HE creditsStudent thesis
Abstract [en]

This thesis investigates whether a general-purpose multimodal large language model (LLM) can be used for indoor localization by estimating the position of a query image on a floorplan without task-specific training or prior environmental knowledge. Indoor localization remains a challenging problem because many existing solutions depend on dedicated infrastructure or prior mapping. A general-purpose multimodal LLM is therefore interesting to investigate because it may offer a low-setup alternative that uses visual and spatial reasoning rather than specialized localization hardware or training data.

An experimental study was conducted in a controlled office environment where the model was evaluated under varying conditions, including different prompt strategies, image types, number of query images, image quality and environmental settings. Localization performance was measured using Euclidean error distance, and both accuracy and consistency were analyzed across 432 test cases.

The results show that the model is capable of producing reasonable localization estimates in some cases, with a mean error of about 9 meters, but performance is highly variable and lacks consistency. Differences between prompt strategies were minimal, while environmental factors and image conditions had a more noticeable impact. The model demonstrated an ability to interpret structural and semantic features from images, but often struggled to distinguish between similar locations on the floorplan. Coherent reasoning explanations did not consistently correspond to accurate spatial predictions.

Overall, the findings indicate that the tested general-purpose multimodal LLM has potential for approximate, area-level indoor localization with low setup requirements, but was not reliable enough for precise positioning. The approach is therefore more suitable as a supportive component in a localization system than as a standalone indoor positioning solution.

Abstract [sv]

Denna studie undersöker om en generell multimodal stor språkmodell (LLM) kan användas för inomhuslokalisering genom att uppskatta positionen för en bild på en planritning utan anpassad träning eller tidigare kunskap om miljön. Inomhuslokalisering är fortfarande ett utmanande problem eftersom många befintliga lösningar är beroende av dedikerad infrastruktur eller tidigare kartläggning. En generell multimodal LLM är därför intressant att undersöka eftersom den kan erbjuda ett alternativ med låga krav på förberedelser, genom att använda visuell och rumslig resonemangsförmåga istället för specialiserad lokaliseringshårdvara eller träningsdata.

En experimentell studie genomfördes i en kontrollerad kontorsmiljö där modellen utvärderades under varierande förhållanden, inklusive med olika prompt strategier, bildtyper, antal bilder, bildkvalitet och miljöfaktorer. Modellens förmåga att lokalisera mättes med euklidiskt felavstånd, och både noggrannhet och stabilitet analyserades över 432 testfall.

Resultaten visar att modellen i vissa fall kan ge rimliga uppskattningar av position, med ett medelfel på ungefär 9 meter, men att prestandan är instabil och varierar kraftigt. Skillnader i resultat mellan prompt strategier var små, medan miljö och bildförhållanden hade större påverkan. Modellen visade förmåga att tolka strukturella och semantiska egenskaper i bilder, men hade svårt att särskilja liknande platser på planritningen. Dessutom visade sig välformulerade förklaringar inte alltid motsvara korrekta positioner.

Sammanfattningsvis visar resultaten att den testade generella multimodal LLM:en har potential för ungefärlig inomhuslokalisering med låga krav på implementation, men att tillförlitligheten ännu inte är tillräcklig för exakt positionering. Metoden lämpar sig därför bättre som en stödjande komponent i ett lokaliseringssystem än som ett fristående system för inomhuspositionering.

Place, publisher, year, edition, pages
2026. , p. 56
Keywords [en]
Indoor localization, multimodal LLM, computer vision, spatial reasoning
Keywords [sv]
Inomhuslokalisering, multimodal LLM, datorseende, spatialt resonemang
National Category
Computer and Information Sciences
Identifiers
URN: urn:nbn:se:mau:diva-87055OAI: oai:DiVA.org:mau-87055DiVA, id: diva2:2085813
External cooperation
Theca Systems AB
Educational program
TS Systemutvecklare
Presentation
2026-06-01, OR:F312, Nordenskiöldsgatan 10, Malmö, 12:18 (Swedish)
Supervisors
Examiners
Available from: 2026-07-10 Created: 2026-07-10 Last updated: 2026-07-10Bibliographically approved

Open Access in DiVA

fulltext(15267 kB)42 downloads
File information
File name FULLTEXT02.pdfFile size 15267 kBChecksum SHA-512
e8676daa84fd8782a7931e51fc4f25e38fc318b02d8543ff679017f9dfd1939f44df463602dbc87ef285893702e5581116be457be25db556efd5f42e44fa3470
Type fulltextMimetype application/pdf

By organisation
Department of Computer Science and Media Technology (DVMT)
Computer and Information Sciences

Search outside of DiVA

GoogleGoogle Scholar
Total: 42 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 187 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf