Unsupervised Transcription ofHistorical Ciphers Using Cluster Refinement: Leveraging Agglomerative Clustering by Implementinga Context-Based Refinement Method
2026 (Engelska)Självständigt arbete på avancerad nivå (masterexamen), 20 poäng / 30 hp
Studentuppsats (Examensarbete)
Abstract [en]
Historical cipher manuscripts are difficult to transcribe because they often contain unknown symbol systems, degraded document quality, touching symbols, and variation in writing style. This thesis investigates whether context-based cluster refinement can improve an unsupervised transcription pipeline for historical cipher manuscripts.
The base method represents each segmented symbol using SIFT descriptors aggregated into Fisher vectors. Agglomerative hierarchical clustering is then used to group visually similar symbols, and selected clusters are used as seeds for label propagation. The proposed refinement methods are applied after clustering and before label propagation. They identify visually suspicious symbols within clusters and evaluate whether these symbols should be reassigned using local neighboring-symbol context.
The methods are evaluated using Symbol Error Rate (SER) on symbol-based cipher datasets from the ICDAR 2024 competition. The base method is compared with two context-based refinement methods that modify the clusters prior to label propagation: a single-pass refinement method and an iterative refinement method.
The results show that context-based refinement improves SER for some datasets but not consistently across all ciphers. For Copiale, the best refinement method reduces SER from 0.482 to 0.441. For Ramanacoil, SER is reduced from 0.343 to 0.329. For Borg, the refinement methods do not improve over the base-method SER of 0.659. These results indicate that local contextual information can support cluster refinement in some cases, but that context alone is not sufficient for robust improvement across different historical cipher datasets.
Ort, förlag, år, upplaga, sidor
2026. , s. 41
Nyckelord [en]
Historical Cipher Transcription, Unsupervised Learning, SIFT Descriptors, Fisher Vectors, Agglomerative Hierarchical Clustering, Context-Based Cluster Refinement, Label Propagation, Symbol Error Rate
Nationell ämneskategori
Språkbehandling och datorlingvistik
Identifikatorer
URN: urn:nbn:se:su:diva-257093OAI: oai:DiVA.org:su-257093DiVA, id: diva2:2076185
Presentation
2026-06-02, Auditorium 11, Frescativägen, 114 19 Stockholm, Stockholm, 19:29 (Engelska)
Handledare
Examinatorer
2026-06-222026-06-212026-06-22Bibliografiskt granskad