6G • Core Network
On-Device Split Inference Client
Terminal-Edge Cooperative Deep Learning Split Engine
Terminal runtime partitioning deep neural networks between the mobile device and edge cloud to execute heavy AI models in real time.
Technical Explanation
Modern generative AI models (such as large multimodal models) exceed the memory and thermal capacity of mobile smartphones. The On-Device Split Inference Client partitions the neural network at an optimal intermediate bottleneck layer. The handset computes early feature representations locally on its NPU, transmits compressed intermediate activation vectors over the ultra-low latency 6G link, and allows the 6G mobile edge server to complete the computationally intensive final layers.
Key Functions
- Partitions large neural models dynamically based on wireless channel conditions
- Computes early feature layers locally to preserve user raw data privacy
- Compresses intermediate activation tensors to minimize uplink bandwidth
- Achieves sub-10ms end-to-end multimodal AI response times on mobile devices
Specifications
3GPP TR 22.874 (AI/ML in 5G/6G), ITU-T Focus Group on AI for Future Networks
Interfaces
Edge AI Inference APIHandset NPU RuntimeUser Plane Channel
Related 6G Concepts
Want to memorize 6G concepts like this one?
Study it with SuperMemo SM-2 spaced repetition flashcards.