QVAC Fabric enables distributed AI model execution across devices

Summary

QVAC has announced an update to its QVAC Fabric, enabling users to perform distributed cluster inference, which allows a single model to be split across multiple machines. This advancement was demonstrated by Gianfranco, who successfully ran DeepSeek V4 Flash, a model with 144 GB of weights, using two connected DGX Sparks. This feature is currently available in QVAC Fabric and will be included in the upcoming SDK update, significantly enhancing the capability for users with multiple devices to execute larger AI models locally at home.

Analysis

Gianfranco: Gianfranco is an individual associated with the QVAC project who conducted testing and demonstration of its distributed capabilities. In the news, Gianfranco executed a large 144 GB model across connected DGX Spark machines using QVAC Fabric's new cluster inference feature. QVAC Fabric: QVAC Fabric is the core inference and fine-tuning engine in the QVAC ecosystem by Tether, forked from llama.cpp to support local AI workloads on consumer hardware including mobile devices and desktops via cross-platform backends like Vulkan and Metal. It enables on-device model execution, LoRA fine-tuning, and related capabilities while keeping data private and offline. In this news, QVAC Fabric introduces support for splitting a single model across multiple machines, demonstrated with cluster inference for larger models like DeepSeek V4 Flash. Use Case: Cluster inference allows users with multiple devices to run significantly larger AI models locally at home. Product Update: QVAC Fabric now supports distributed cluster inference to split models across multiple connected machines.

Categories

aimachine_learningtech

Related sources

View Original Tweet