Digitala Vetenskapliga Arkivet

Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Representing Competence: A Comparative Study of TF-IDF and Sentence-BERT for Role-Based Competence Mapping
Stockholm University, Faculty of Social Sciences, Department of Computer and Systems Sciences.
2026 (English)Independent thesis Advanced level (degree of Master (Two Years)), 20 credits / 30 HE creditsStudent thesis
Abstract [en]

Introduction: Organizations use competence mapping to support workforce planning, employee development, and internal mobility. Much of this information is found in internal documents such as role descriptions. These texts are written in natural language, vary in style, and are difficult to analyze in a consistent way. NLP methods offer ways to structure this material, but different representations capture different aspects of the text, which shapes the results.

Research Question: This thesis examines how different NLP text representations influence competence mapping based on internal organizational role descriptions. The focus is on how representation choices affect the identification, structuring, and communicability of competence‑related information.

Method: The study is conducted as a case study at a global manufacturing and service company and uses 89 internal role descriptions from a Nordic and Baltic regional unit. Two approaches are compared under the same conditions: a vocabulary‑based method (TF‑IDF) and an embedding‑based method (Sentence‑BERT). Both are applied using unsupervised clustering. The results are evaluated using internal clustering metrics, a robustness check with BERTopic, and a post‑hoc comparison with a formal competency framework.

Results: The two approaches produce clearly different outcomes. Sentence‑BERT generates more coherent and stable role groupings based on semantic similarity, even when roles are described with different wording. TF‑IDF produces more fragmented clusters that mainly reflect similarities in phrasing and documentation practices rather than shared meaning. The BERTopic analysis supports the stability of the Sentence‑BERT clusters. Both methods recover most of the formal competency framework, with only a small number of competencies remaining weakly represented.

Discussion: The findings show that competence mapping with NLP is not only a technical task but also a representational and organizational one. Where coverage is limited, this reflects the boundaries of what role descriptions capture in practice rather than shortcomings of the methods themselves. Role descriptions function as an unintended archive of how work is understood at a given moment. They can be analyzed productively, but only with attention to the limitations of the source material. A useful direction for future research may be less about more advanced methods and more about understanding what gets written down, by whom, and for what purpose.

Place, publisher, year, edition, pages
2026.
Keywords [en]
competence mapping, natural language processing, TF-IDF, Sentence-BERT, unsupervised clustering, role descriptions, workforce planning, text representation, organizational competence, skill extraction
National Category
Natural Language Processing
Identifiers
URN: urn:nbn:se:su:diva-257981OAI: oai:DiVA.org:su-257981DiVA, id: diva2:2090771
Available from: 2026-08-09 Created: 2026-08-09

Open Access in DiVA

fulltext(967 kB)17 downloads
File information
File name FULLTEXT01.pdfFile size 967 kBChecksum SHA-512
2ee304962526f41db4ca0ba344bba942f3c2680772e404e7b36d09ab8430596c5adc6b7e6ba2d8b3d4db73bdb2ad9b5da17b833084088d039bcc033c3dd086b1
Type fulltextMimetype application/pdf

Search in DiVA

By author/editor
Forss, Ella Maria
By organisation
Department of Computer and Systems Sciences
Natural Language Processing

Search outside of DiVA

GoogleGoogle Scholar
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 2726 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf