Digitala Vetenskapliga Arkivet

Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Unsupervised Transcription ofHistorical Ciphers Using Cluster Refinement: Leveraging Agglomerative Clustering by Implementinga Context-Based Refinement Method
Stockholm University, Faculty of Humanities, Department of Linguistics.
2026 (English)Independent thesis Advanced level (degree of Master (Two Years)), 20 credits / 30 HE creditsStudent thesis
Abstract [en]

Historical cipher manuscripts are difficult to transcribe because they often contain unknown symbol systems, degraded document quality, touching symbols, and variation in writing style. This thesis investigates whether context-based cluster refinement can improve an unsupervised transcription pipeline for historical cipher manuscripts.

The base method represents each segmented symbol using SIFT descriptors aggregated into Fisher vectors. Agglomerative hierarchical clustering is then used to group visually similar symbols, and selected clusters are used as seeds for label propagation. The proposed refinement methods are applied after clustering and before label propagation. They identify visually suspicious symbols within clusters and evaluate whether these symbols should be reassigned using local neighboring-symbol context.

The methods are evaluated using Symbol Error Rate (SER) on symbol-based cipher datasets from the ICDAR 2024 competition. The base method is compared with two context-based refinement methods that modify the clusters prior to label propagation: a single-pass refinement method and an iterative refinement method.

The results show that context-based refinement improves SER for some datasets but not consistently across all ciphers. For Copiale, the best refinement method reduces SER from 0.482 to 0.441. For Ramanacoil, SER is reduced from 0.343 to 0.329. For Borg, the refinement methods do not improve over the base-method SER of 0.659. These results indicate that local contextual information can support cluster refinement in some cases, but that context alone is not sufficient for robust improvement across different historical cipher datasets.

Place, publisher, year, edition, pages
2026. , p. 41
Keywords [en]
Historical Cipher Transcription, Unsupervised Learning, SIFT Descriptors, Fisher Vectors, Agglomerative Hierarchical Clustering, Context-Based Cluster Refinement, Label Propagation, Symbol Error Rate
National Category
Natural Language Processing
Identifiers
URN: urn:nbn:se:su:diva-257093OAI: oai:DiVA.org:su-257093DiVA, id: diva2:2076185
Presentation
2026-06-02, Auditorium 11, Frescativägen, 114 19 Stockholm, Stockholm, 19:29 (English)
Supervisors
Examiners
Available from: 2026-06-22 Created: 2026-06-21 Last updated: 2026-06-22Bibliographically approved

Open Access in DiVA

fulltext(790 kB)35 downloads
File information
File name FULLTEXT01.pdfFile size 790 kBChecksum SHA-512
4c3e067103d2c3b71ab60db9e08af3001dbe7292a8ee80317edefecbd72fc5fab7550200c4ee1cec1139158f081dcecfb7823391a1d17bede0bcec59ce24cd49
Type fulltextMimetype application/pdf

By organisation
Department of Linguistics
Natural Language Processing

Search outside of DiVA

GoogleGoogle Scholar
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 198 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf