Kollaborativ AI-assisterad utveckling: Jämförelse av tre AI-modeller i en frontend-applikation
2026 (Swedish)Independent thesis Basic level (university diploma), 10 credits / 15 HE credits
Student thesis
Abstract [sv]
Bakgrund: Användningen av AI inom programmering har ökat snabbt och erbjuder nya möjligheter att effektivisera utvecklingsprocessen. Samtidigt är det oklart hur olika AI-modeller faktiskt skiljer sig i kodkvalitet och användbarhet.
Syfte: Att undersöka hur kodens komplexitet skiljer sig mellan olika AI-modeller vid utveckling av en frontend-applikation och tillhörande testkod.
Metod: En frontend-applikation och dess testkod utvecklades med stöd av tre AI-modeller: ChatGPT, Llama och Mistral. Kodens komplexitet analyserades med hjälp av statiska kodmått – cyclomatic complexity, lines of code (LOC) och maintainability – via SonarQube. Skillnader mellan modellerna utvärderades med Mann-Whitney U-test.
Resultat: Skillnaderna i kodens komplexitet är små och inte statistiskt signifikanta. Däremot sågs tydligare skillnader i praktisk användning, där ChatGPT uppvisar högre stabilitet och kräver mindre efterarbete. Testkod visade sig vara mer utmanande för samtliga modeller än applikationskod.
Slutsats: Studien visar att statiska kodmått inte fullt ut fångar skillnader i användbarhet mellan AI-modeller. Val av AI-modell bör därför inte enbart baseras på kodkvalitet, utan även på faktorer som stabilitet och arbetsinsats. Resultaten bör dock tolkas med försiktighet på grund av studiens begränsade urval.
Abstract [en]
Background: The use of AI assistants in programming has increased rapidly, offering new opportunities to improve software development processes. It remains unclear how different AI models compare in terms of code quality and practical usability.
Purpose: To investigate how code complexity differs between AI assistants when developing a frontend application and its associated test code.
Method: A frontend application and its test code were developed using three AI models: ChatGPT, Llama, and Mistral. Code complexity was analyzed using static code metrics—cyclomatic complexity, lines of code (LOC), and maintainability—through SonarQube. Differences between models were evaluated using the Mann-Whitney U-test.
Results: Differences in code complexity between the models are small and not statistically significant. Clearer differences appeared in practical use, where ChatGPT demonstrates higher stability and requires less post-editing. Test code was found to be more challenging for all models compared to application code.
Conclusion: The study indicates that static code metrics alone are insufficient to fully capture differences in usability between AI assistants. The choice of which AI model to use should consider not only code quality but also factors such as stability and effects on the development process. However, the findings are limited because of the small sample size.
Place, publisher, year, edition, pages
2026. , p. 20
Keywords [sv]
AI-modeller kodgenerering, kodkvalitet, testkod, statisk analys, SonarQube
National Category
Computer Sciences
Identifiers
URN: urn:nbn:se:bth-29656OAI: oai:DiVA.org:bth-29656DiVA, id: diva2:2080800
Subject / course
PA1438 Självständigt arbete Webbprogrammering
Educational program
PAGWG Webbprogrammering
Supervisors
Examiners
2026-06-292026-06-282026-06-29Bibliographically approved