Digitala Vetenskapliga Arkivet

Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
ORSExplorer: An LLM-Assisted Rule System Extraction Tool – Semi-Automated Extraction of Regulatory Entities and Relationships for Knowledge Graph Creation
Stockholm University, Faculty of Social Sciences, Department of Computer and Systems Sciences.
Stockholm University, Faculty of Social Sciences, Department of Computer and Systems Sciences.
2026 (English)Independent thesis Advanced level (degree of Master (Two Years)), 20 credits / 30 HE creditsStudent thesis
Abstract [en]

Introduction: Legal and regulatory domains contain large amounts of information that are difficult to interpret and structure. Regulatory rules are often connected to other rules, concepts and definitions, which makes them difficult to analyze as isolated statements. Knowledge graphs can represent this type of knowledge as entities and relations but the construction of regulatory knowledge graphs is still largely manual and dependent on expert interpretation. This thesis addresses this problem by designing, implementing and evaluating ORSExplorer, a semi-automated tool that uses Large Language Models to support the extraction of entities and semantic relationships from regulatory documents.

Research Question: The primary research question is: “How can a semiautomated Large Language Model-based tool support expert users in extracting and structuring entities and relationships from regulatory documents for knowledge graph creation?”

Method: This thesis follows a Design Science Research methodology. Requirements were collected through expert workshops and supervision meetings and were used to guide the iterative development of ORSExplorer. The artifact was implemented as a web-based tool that supports ontology-guided extraction, human-in-the-loop review, graph visualization, manual editing and Wikibase-compatible export. The evaluation was conducted through expert assessment and semi-structured interviews with five expert participants. The interview material was analyzed using thematic analysis and selected requirements were also assessed through technical verification.

Results: The results show that ORSExplorer can support expert users by generating ontology-guided candidate entities and relationships from regulatory text. These outputs can be inspected, edited and accepted before they are used for knowledge graph creation. The evaluation showed that participants valued the incremental workflow, the possibility to review reasoning and evidence and the ability to manually control the generated output. The main findings were grouped into four themes: usability, workflow efficiency, trust and control and perceived accuracy and output quality. The results also showed limitations in first-time usability, graph readability, the need for ontology knowledge and limited history tracking between iterations.

Discussion: The findings suggest that semi-automated Large Language Model-based tools can support regulatory knowledge graph creation when they are designed as support tools rather than fully automated systems. The artifact helped reduce parts of the manual structuring work but expert validation remained necessary because the generated output could not be treated as final legal interpretation. The main contribution of this thesis is ORSExplorer together with design knowledge on ontology guidance, transparency, incremental construction and human-in-the-loop validation. Future research could evaluate the artifact against a manually created gold standard, improve graph readability, history tracking and test the approach in larger regulatory knowledge graph workflows.

Place, publisher, year, edition, pages
2026.
Keywords [en]
Large Language Models, Knowledge Graphs, Regulatory Documents, Ontology-Guided Extraction, Human-in-the-Loop, Entity and Relationship Extraction, Wikibase
National Category
Information Systems
Identifiers
URN: urn:nbn:se:su:diva-257270OAI: oai:DiVA.org:su-257270DiVA, id: diva2:2079126
Available from: 2026-06-24 Created: 2026-06-24

Open Access in DiVA

fulltext(4837 kB)101 downloads
File information
File name FULLTEXT01.pdfFile size 4837 kBChecksum SHA-512
f83c6139167195d402d3dabaa05b97fe2fc68b403f14460e5f3fed023c356bd83828c676c9bfcb4a74a3c4a5f1413ab6554a2d99aecf5f0c55ff13f049792c14
Type fulltextMimetype application/pdf

Search in DiVA

By author/editor
Belhaj, FilipPettersson, Samuel
By organisation
Department of Computer and Systems Sciences
Information Systems

Search outside of DiVA

GoogleGoogle Scholar
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 201 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf