Graphcore and Arm compressed Meta's 11B vision language model to 3.7 GB for a Pixel 8a, with average visual question answering accuracy dropping from 74.4% to 66.1%.
Researchers at Graphcore Research and Arm have compressed an 11-billion-parameter vision-language model, AI that can read and reason about images alongside text, from 21.3 GB to 3.7 GB, small enough to run on a Pixel 8a phone. The compressed model still answers visual questions, but its average accuracy drops from 74.4% on the bfloat16 baseline to 66.1% under compression, an 8.3-percentage-point gap that the team reports openly in its arXiv paper and in coverage by SemiEngineering.
The 80%-plus memory cut moves the boundary of what fits on a phone. Standard 4-bit quantization, the prior best for cramming large models into tight memory budgets, still leaves the same model above 5 GB. A 3.7 GB footprint sits inside the working memory of a midrange Android device, which is why the team built a public Android demo APK alongside the paper and the open-source code release.
The mechanism is a sub-byte storage format called S3D8, and the reason it works has more to do with where the math happens than with how small the numbers get. S3D8 packs three weights into each byte: a shared 5-bit index that selects one of 32 INT8 values from a small lookup table, plus one sign bit per weight, plus a per-channel bfloat16 scale. On disk, that averages 2.7 bits per weight. At compute time, the format decodes in parallel via table lookup into standard INT8 weights, which then feed the INT8 matrix-multiply hardware that already exists on every modern Arm mobile CPU. The phone's SIMD engine never sees a 2.7-bit multiply, because the chip's existing INT8 path accelerates the math after the unpack. S3D8 is, in effect, a way to make a 2.7-bit weight file ride an 8-bit compute path that phones already ship with.
Researchers have spent years trying to build chips that do sub-byte arithmetic natively. S3D8 bets on a different lever: keep the on-disk weights dense, then unpack into the precision the silicon already supports. That separation, between storage precision and compute precision, is what puts on-device multimodal AI inside the budgets of today's phones rather than tomorrow's.
The accuracy recovery came from quantization-aware training, a fine-tuning stage where the model sees the S3D8 quantization in its forward pass and adapts its weights. QAT usually demands the original training corpus, a problem for any model whose full data is not public. The team reports a variant that works without it, a meaningful claim because the Llama 3.2 Vision training data is not fully open. Activations are quantized to INT8 as well, which keeps the model's working memory close to its compressed weight footprint.
The Pixel 8a is a specific device with Google's Tensor G3 chip. Comparable numbers on older Arm cores, other CPU instruction sets, or Apple's silicon are not in the current write-up, so the result reads as proof that the format works on at least one widely deployed mobile CPU rather than a universal claim. The on-device demo runs interactively; precise tokens-per-second figures are not in the paper or the blog.
The 8.3-point VQA gap is the part the trade press tends to footnote. A model that answers 66% of visual questions correctly instead of 74% is meaningfully less capable on that benchmark, and a production deployment would weigh that loss against the privacy, latency, and offline benefits of running locally. VQA is also a single benchmark family. Whether the gap holds across captioning, OCR, document understanding, and multilingual VQA is the next set of questions for the community to answer, and the team's risk notes flag exactly this.
An 11-billion-parameter vision-language model now fits on a phone. It is smaller, more private, and faster to use without a network. It is also measurably less accurate on at least one standard benchmark. The next move belongs to the research community: replicate the result on broader VLM tasks, test the format on other mobile CPUs, and report the cost in apples-to-apples terms.