Eilmeldung
ARزلزال بقوة 4.7 درجة يضرب غرب نابولي ويصيب 26 شخصاًARالجزائر: مقتل 31 شخصاً في حادثي سير مأساويين وحداد وطنيARتدفق 60 ألف مهاجر إلى سبتة: إسبانيا تتهم عصابات التهريب وقادة أوروبيون يعربون عن قلقهمARالاتحاد السعودي يعير ماريو ميتاي إلى جنوى الإيطاليARوزارة الخارجية الأمريكية تطلق تحذيراً عالمياً لمواطنيها بسبب تصاعد التوترات في الشرق الأوسطARقائد إيراني يحذر دول المنطقة: من يحمي أمريكا "سيحترق في نار الحرب"ARاستقبال الأبطال وبطاقة حمراء لروبرتو لوبيز في آيرلنداARالأمير علي بن الحسين يدعو "فيفا" لإلغاء الاقتراع السري في الانتخابات الرئاسيةARمخاوف غربية من ضعف القدرات العسكرية الأمريكية وتأثيرها على الردع والتحالفاتARتزايد الضغوط على إنفانتينو ومرشحون محتملون لخلافته في رئاسة الفيفاARزلزال بقوة 4.7 درجة يضرب غرب نابولي ويصيب 26 شخصاًARالجزائر: مقتل 31 شخصاً في حادثي سير مأساويين وحداد وطنيARتدفق 60 ألف مهاجر إلى سبتة: إسبانيا تتهم عصابات التهريب وقادة أوروبيون يعربون عن قلقهمARالاتحاد السعودي يعير ماريو ميتاي إلى جنوى الإيطاليARوزارة الخارجية الأمريكية تطلق تحذيراً عالمياً لمواطنيها بسبب تصاعد التوترات في الشرق الأوسطARقائد إيراني يحذر دول المنطقة: من يحمي أمريكا "سيحترق في نار الحرب"ARاستقبال الأبطال وبطاقة حمراء لروبرتو لوبيز في آيرلنداARالأمير علي بن الحسين يدعو "فيفا" لإلغاء الاقتراع السري في الانتخابات الرئاسيةARمخاوف غربية من ضعف القدرات العسكرية الأمريكية وتأثيرها على الردع والتحالفاتARتزايد الضغوط على إنفانتينو ومرشحون محتملون لخلافته في رئاسة الفيفا
Newsgather
ZurückGoogle Releases New Gemma 4 12B AI Model for Consumer Laptops
Google Releases New Gemma 4 12B AI Model for Consumer Laptops
Technik
Ars Technica3.6.2026Technik3 Min. LesezeitUnited States

Google Releases New Gemma 4 12B AI Model for Consumer Laptops

Auf einen Blick

  • Google has launched Gemma 4 12B, a new AI model designed for efficiency, capable of running on consumer laptops with 16GB RAM.
  • It offers advanced reasoning and multimodal capabilities with reduced latency and memory usage compared to larger variants.

KI-generierte Zusammenfassung

Schriftgröße

The generative AI boom has driven the cost of memory into the stratosphere, and Google is a key part of that trend. So it’s only fitting that Google should offer some less RAM-hungry local AI models. The company has announced the release of a new Gemma 4 12B model that fills a gap in the lineup that launched earlier this year. The new model is efficient enough that you may be able to run it on a pretty average consumer laptop.

In April, Google released four models in the Gemma 4 family, which also marked the shift to a more open Apache 2.0 license. The initial models included two mobile-optimized options (E2B and E4B) along with a pair of models for more serious work (26B Mixture of Experts and 31B Dense). That left a rather large unserved space in the middle, which is right where the new model falls.

Gemma 4 12B is considerably more capable than the mobile versions, but it won’t require a $20,000 AI accelerator to run locally. Google says Gemma 4 12B is unique in that it can run on many consumer laptops without sacrificing quality. As long as you’ve got a computer with 16GB of system RAM or VRAM, the 12-billion-parameter model will work. That’s about half the total memory footprint of Gemma 4 26B MoE, and Google claims the new model is almost as capable, at least as far as benchmarks go.

Google says the new model is capable of complex multistep reasoning and agentic workflows that previously required the larger Gemma variants. Despite the smaller parameter count, Gemma 4 12B comes with the newly devised Multi-Token Prediction (MTP) drafters, which take advantage of unused processing cycles to calculate possible future tokens. The result is greater speed and efficiency. Google has released optional MTP versions of the other Gemma 4 models, but this is the first one to have MTP out of the box.

Gemma 4 12B is also more efficient thanks to a new approach to multimodality. The Gemma 4 family is natively multimodal, accepting text, audio, or images as inputs. Most gen AI models—including the other Gemma 4 variants—use dedicated encoders to process non-text inputs and pass that data to the LLM. This works well enough, but it increases latency and memory usage.

With the new mid-weight model, Google has implemented a streamlined embedding module for vision, featuring single-matrix multiplication and positional embedding, which allows the data to pass to the LLM with proper spatial awareness. This eliminates the need for a bulky middleman encoder. For audio, there’s no encoding at all. The developers worked out a method of projecting the raw audio signal into the same vectors used for text tokens.

If you want to check out the new Gemma 4 model, it’s accessible without a download via tools like LM Studio, Google AI Edge Gallery, and more. But the whole idea with Gemma 4 12B is that you can run it locally and on your own terms. If you’ve got the RAM, the model weights are available for download immediately on Kaggle and Hugging Face. It’s just shy of 18GB.

Verwandte Themen

This article was originally published by Ars Technica.

Ähnliche Meldungen

Mehr zu diesem Themagenerative AI