Home/Glossary/6G Glossary/On-Device Split Inference Client
6G • Core Network

On-Device Split Inference Client

Terminal-Edge Cooperative Deep Learning Split Engine

Terminal runtime partitioning deep neural networks between the mobile device and edge cloud to execute heavy AI models in real time.

Technical Explanation

Modern generative AI models (such as large multimodal models) exceed the memory and thermal capacity of mobile smartphones. The On-Device Split Inference Client partitions the neural network at an optimal intermediate bottleneck layer. The handset computes early feature representations locally on its NPU, transmits compressed intermediate activation vectors over the ultra-low latency 6G link, and allows the 6G mobile edge server to complete the computationally intensive final layers.

Key Functions

  • Partitions large neural models dynamically based on wireless channel conditions
  • Computes early feature layers locally to preserve user raw data privacy
  • Compresses intermediate activation tensors to minimize uplink bandwidth
  • Achieves sub-10ms end-to-end multimodal AI response times on mobile devices
Specifications
3GPP TR 22.874 (AI/ML in 5G/6G), ITU-T Focus Group on AI for Future Networks
Interfaces
Edge AI Inference APIHandset NPU RuntimeUser Plane Channel

Related 6G Concepts

Want to memorize 6G concepts like this one?
Study it with SuperMemo SM-2 spaced repetition flashcards.
Practice 6G Flashcards