Digitala Vetenskapliga Arkivet

Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Understanding Intel E-core Prefetcher Behaviour through Performance Counters
Uppsala University, Disciplinary Domain of Science and Technology, Mathematics and Computer Science, Department of Information Technology.
2026 (English)Independent thesis Basic level (degree of Bachelor), 10 credits / 15 HE creditsStudent thesis
Abstract [en]

As modern computational workloads become increasingly data-intensive, memory access latencyremains a primary bottleneck to performance. To mitigate this “Memory Wall”, processors utilisehardware prefetchers of varying complexity, from basic next-line fetchers, to advanced spatialpredictors to speculatively load data into cache. To prevent cache pollution and interconnect con-gestion, processors often employ static bandwidth throttling policies to clamp speculative traffic. The objective of this thesis is to empirically evaluate the performance characteristics of differentprefetching algorithms and their systemic interactions with bandwidth throttling mechanisms onthe Intel Crestmont microarchitecture.

By using model-specific registers and physical performance monitoring unit telemetry, this the-sis isolated individual prefetcher algorithms (L1 Next-Line prefetcher, L2 streamer, and Adaptive Multi-Path), and then evaluated them across varying throttling profiles using the SPEC CPU 2017 benchmark suite. The results demonstrate that prefetcher efficacy is highly workload-dependent. While compute-bound workloads are largely insensitive to prefetcher configurations, memory-bound workloads expose a conflict coined the Throttling Paradox. In this context, the paradox is that the mechanism designed to optimise performance by limiting memory bandwidth, insteadleads to performance degradations on single threaded workloads. The findings suggest that it triggers demand-driven replacement by blocking early speculative requests and forcing the CPU to issue late, pipeline-stalling demand misses for the exact same data.

Place, publisher, year, edition, pages
2026. , p. 29
Series
IT ; IT kDV 26 035
National Category
Computer Sciences
Identifiers
URN: urn:nbn:se:uu:diva-594313OAI: oai:DiVA.org:uu-594313DiVA, id: diva2:2086748
Supervisors
Examiners
Available from: 2026-07-22 Created: 2026-07-15 Last updated: 2026-07-22Bibliographically approved

Open Access in DiVA

fulltext(4061 kB)57 downloads
File information
File name FULLTEXT01.pdfFile size 4061 kBChecksum SHA-512
1ef7d014c8e5d475d795a350aa2a007e2b5e4abe4810b766a03e6359c336e271fb6dc0474e682c01e96ddd4d37671d679333b313528dc3439d35991b3a4caa14
Type fulltextMimetype application/pdf

By organisation
Department of Information Technology
Computer Sciences

Search outside of DiVA

GoogleGoogle Scholar
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 1774 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf