Can Agentic AI Meaningfully Automate Complex Enterprise Workflows?: A Case Study in Hardware Management
2026 (English)Independent thesis Advanced level (degree of Master (Two Years)), 20 credits / 30 HE credits
Student thesis
Abstract [en]
This thesis investigates whether agentic AI frameworks can meaningfully automate complex enterprise workflows, addressing challenges related to data quality, interoperability, and system reliability. Although Large Language Model (LLM)-based systems have demonstrated strong reasoning capabilities in controlled environments, their deployment in enterprise settings—characterizedby heterogeneous tools, inconsistent data schemas, and strict operational constraints—remainsinsufficiently explored. This study focuses on Ericsson’s BNEW RCE LO Hardware Management workflow, where hardware recommendation and procurement processes currently requiresubstantial manual intervention across fragmented enterprise systems.
To evaluate the feasibility and architectural trade-offs of enterprise workflow automation, a researchdriven prototype (Mastermatch AI) was developed using a tool-calling agent architecture. Thesystem was evaluated through a comparative ablation study involving five retrieval configurations: structured retrieval, semantic retrieval using Retrieval-Augmented Generation (RAG), andhybrid retrieval approaches combining deterministic filtering with semantic ranking. Additionally,the study examined the impact of structured JSON representations versus canonical text representations on retrieval reliability and reasoning behavior.
The experimental evaluation was conducted using 120 scenario-driven enterprise queries covering identifier lookup, requirement-based retrieval, underspecified requests, overconstrainedqueries, out-of-scope interactions, and ordering workflows. Results demonstrate that hybrid retrieval approaches significantly outperform purely semantic retrieval systems across retrieval correctness, ranking quality, task completion, and behavioral correctness metrics. While structuredretrieval exhibited strong deterministic behavior in constraint-sensitive scenarios, purely semanticconfigurations were more prone to hallucination, retrieval instability, and incorrect entity mapping.Differences between JSON and canonical text representations were observed, particularly inhallucination-related failures, although these differences were not statistically significant overall.Statistical validation using Friedman and Wilcoxon signed-rank tests confirmed the significanceof the observed retrieval strategy differences, while weighted Cohen’s kappa demonstrated highevaluation reliability.
The findings suggest that enterprise deployment of agentic AI systems should prioritize hybridretrieval architectures that combine structured control with semantic flexibility. Furthermore, theresults indicate that Human-AI Collaboration (HAIC) remains necessary for workflow-critical enterprise operations, particularly in scenarios requiring strict constraint handling, validation, andprocurement decisions. Overall, this work provides empirical insights into retrieval design, failurebehavior, and workflow orchestration for real-world enterprise agentic systems, contributing bothto academic research and practical enterprise adoption strategies.
Place, publisher, year, edition, pages
2026. , p. 56
Series
IT ; mDA 26 033
Keywords [en]
Agentic AI, Enterprise Workflow Automation, Retrieval-Augmented Generation (RAG), Large Language Models (LLMs), Hybrid Retrieval, Tool-Calling Agents, Hardware Management, Semantic Search, Structured Search, Human-AI Collaboration (HAIC), Data Representation
National Category
Artificial Intelligence Computer Engineering Other Computer and Information Science
Identifiers
URN: urn:nbn:se:uu:diva-594053OAI: oai:DiVA.org:uu-594053DiVA, id: diva2:2085675
External cooperation
Ericsson AB
Educational program
Master's Programme in Data Science
Presentation
2026-06-11, Uppsala, 15:15 (English)
Supervisors
Examiners
Note
Confidential
2026-08-192026-07-092026-08-19Bibliographically approved