Digitala Vetenskapliga Arkivet

Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Parameter-efficient Large AI Model Co-inference at Multi-cluster Edge Networks
KTH, School of Electrical Engineering and Computer Science (EECS), Information Science and Engineering.ORCID iD: 0000-0002-0980-1395
Shenzhen University, College of Electronic and Information Engineering, Shenzhen, China.
ShanghaiTech University, School of Information Science and Technology, China.
Beijing University of Posts and Telecommunications, Beijing, China.
Show others and affiliations
2026 (English)In: ICC 2026 - IEEE International Conference on Communications, Proceedings, Institute of Electrical and Electronics Engineers (IEEE) , 2026Conference paper, Published paper (Refereed)
Abstract [en]

The increasing scale and computational demands of large artificial intelligence models (LAIMs) pose significant challenges for efficient inference among distributed devices in resource-constrained environments. This paper presents a multi-cluster LAIM co-inference framework in which an edge server, equipped with multiple graphics processing units (GPUs), coordinates user clusters to fulfill inference requests. We assume that devices within each cluster capture data from different perspectives, and then use on-device lightweight LAIMs to generate local features. The edge server collects and fuses these features to produce a more accurate inference result. To capture the fundamental trade-off between model pruning and co-inference performance under multi-device scenarios, we theoretically characterize the impact of pruning ratio and device contribution based on the rate-distortion analysis and partial information decomposition. Guided by the analysis, we then jointly optimize the pruning ratio, task scheduling, and transmit power to minimize inference distortion subject to the latency, energy, and capacity constraints. Extensive simulations demonstrate that the proposed framework achieves superior performance in comparison to existing baselines, offering improved inference quality and resource efficiency in multi-cluster edge environments.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE) , 2026.
National Category
Computer Sciences Computer Systems Computer Engineering
Identifiers
URN: urn:nbn:se:kth:diva-386492DOI: 10.1109/ICC59461.2026.11587850Scopus ID: 2-s2.0-105045349035OAI: oai:DiVA.org:kth-386492DiVA, id: diva2:2089911
Conference
2026 IEEE International Conference on Communications, ICC 2026, Glasgow, United Kingdom, May 24-28 2026
Note

Part of ISBN 979-8-3195-4209-0

QC 20260805

Available from: 2026-08-05 Created: 2026-08-05 Last updated: 2026-08-05Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textScopus

Search in DiVA

By author/editor
Lyu, Zhonghao
By organisation
Information Science and Engineering
Computer SciencesComputer SystemsComputer Engineering

Search outside of DiVA

GoogleGoogle Scholar

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 12 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf