Digitala Vetenskapliga Arkivet

Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Performance-Portable Optimization and Analysis of Multiple Right-Hand Sides in a Lattice QCD Solver
KTH, School of Electrical Engineering and Computer Science (EECS), Computer Science, Computational Science and Technology (CST).ORCID iD: 0000-0003-3038-3586
Forschungszentrum Jülich GmbH, Jülich, Germany.
Forschungszentrum Jülich GmbH, Jülich, Germany.
University of Wuppertal, Wuppertal, Germany.
Show others and affiliations
2025 (English)In: Proceedings - 2025 IEEE 32nd International Conference on High Performance Computing, Data, and Analytics, HiPC 2025, Institute of Electrical and Electronics Engineers (IEEE) , 2025, p. 301-311Conference paper, Published paper (Refereed)
Abstract [en]

Managing the high computational cost of iterative solvers for sparse linear systems is a known challenge in scientific computing. Moreover, scientific applications often face memory bandwidth constraints, making it critical to optimize data locality and enhance the efficiency of data transport. We extend the lattice QCD solver DD-\alpha AMG to incorporate multiple right-hand sides (rhs) for both the Wilson-Dirac operator evaluation and the GMRES solver, with and without odd-even preconditioning. To optimize auto-vectorization, we introduce a flexible interface that supports various data layouts and implement a new data layout for better SIMD utilization. We evaluate our optimizations on both x86 and Arm clusters, demonstrating performance portability with similar speedups. A key contribution of this work is the performance analysis of our optimizations, which reveals the complexity introduced by architectural constraints and compiler behavior. Additionally, we explore different implementations leveraging a new matrix instruction set for Arm called SME and provide an early assessment of its potential benefits.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE) , 2025. p. 301-311
Keywords [en]
Lattice QCD, SME, multigrid solvers, performance portability
National Category
Computational Mathematics Computer Sciences Software Engineering
Identifiers
URN: urn:nbn:se:kth:diva-385449DOI: 10.1109/HiPC66333.2025.00038ISI: 001792596500029Scopus ID: 2-s2.0-105036377100OAI: oai:DiVA.org:kth-385449DiVA, id: diva2:2086468
Conference
32nd Annual IEEE International Conference on High Performance Computing, Data, and Analytics, HiPC 2025, Hyderabad, India, December 17-20, 2025
Note

Part of ISBN 9798331566647

QC 20260714

Available from: 2026-07-14 Created: 2026-07-14 Last updated: 2026-07-14Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textScopus

Search in DiVA

By author/editor
Long, ShitingPleiter, Dirk
By organisation
Computational Science and Technology (CST)
Computational MathematicsComputer SciencesSoftware Engineering

Search outside of DiVA

GoogleGoogle Scholar

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 1 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf