Digitala Vetenskapliga Arkivet

Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Fault-Tolerant Skeleton Program Execution in Distributed Heterogeneous Parallel Systems
Linköping University, Department of Computer and Information Science, Software and Systems.
2026 (English)Independent thesis Advanced level (degree of Master (Two Years)), 20 credits / 30 HE creditsStudent thesis
Abstract [en]

Distributed stream-processing pipelines deployed on heterogeneous systems can stop

producing useful output when a worker node fails. The same problem can occur when

one pipeline task is affected by a partial node failure while the node remains reachable

to the orchestrator. SkePU-Streaming currently leaves such failures to the surrounding

platform or application code. That leaves a gap for long-running pipelines that need the

runtime itself to detect a failure, remap the affected pipeline tasks, and reconnect the

stream.

The prototype adds framework-level recovery support to SkePU-Streaming. It coor-

dinates recovery inside the runtime and restores the stream path after a failure trigger.

The evaluation uses controlled fault injection and repeated timing measurements on

distributed pipelines. For recovery scope, the evaluation tests six fault cases that expose

the core one-to-one, one-to-many, and many-to-one local recovery patterns, and checks

the same patterns inside a larger motion-detection pipeline. In these cases, the prototype

detects the injected failures, remaps the affected tasks, restores the expected stream path,

and passes the recovery checks. For the evaluated producer–consumer workload, enabling

recovery support leaves the mean per-run median latency in the same range (0.97 ms

vs. 1.09 ms), with no latency penalty distinguishable from run-to-run variation. Warm

standby reduces the mean post-detection recovery interval from 24.03 s under cold failover

to 54.69 ms.

The results show that failure detection, task remapping, and bounded recovery sup-

port can be integrated into SkePU-Streaming. Within the evaluated cases, skeleton ap-

plications resume streaming after the runtime detects a failure and repairs the affected

path.

Place, publisher, year, edition, pages
2026. , p. 60
Keywords [en]
SkePU-Streaming, fault tolerance, distributed system
National Category
Computer Systems
Identifiers
URN: urn:nbn:se:liu:diva-226071ISRN: LIU-IDA/LITH-EX-A--26/081--SEOAI: oai:DiVA.org:liu-226071DiVA, id: diva2:2084071
Subject / course
Computer science
Supervisors
Examiners
Available from: 2026-08-13 Created: 2026-07-03 Last updated: 2026-08-13Bibliographically approved

Open Access in DiVA

fulltext(344 kB)12 downloads
File information
File name FULLTEXT01.pdfFile size 344 kBChecksum SHA-512
49cff088586cae1d0489cc673eaecfc2c8e7cdbb3ea746cef1b293a643b6656bc392093792dd53fd5668f59f17c99a1488079b49c261b15ebb66239ee4fb94a1
Type fulltextMimetype application/pdf

By organisation
Software and Systems
Computer Systems

Search outside of DiVA

GoogleGoogle Scholar
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 4000 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf