/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Google releases Gemma 4 12B, an 11.95B-parameter unified, encoder-free open multimodal model that can run locally on devices with 16GB of VRAM or unified memory

While many AI open source model providers are pursuing larger and more powerful models, Google is still giving attention to the smaller, more local side of the market.

VentureBeat Carl Franzen

Context & Ripple Effects

Google’s Gemma releases have repeatedly emphasized deployment efficiency alongside capability: Gemma 3n targeted devices with as little as 2GB of memory, while Gemma 3 and the broader Gemma 4 family were positioned around single-accelerator use, reasoning, and agentic workflows.

This release narrows that arc to a 12B-class multimodal model with a stated 16GB local-memory target. It matters because it extends Google’s open-model lineup beyond cloud-scale deployment while retaining multimodal functionality.

First-order effects

  • Developers with systems meeting the 16GB VRAM or unified-memory target can run and adapt Google’s open multimodal model locally rather than making local deployment contingent on a larger accelerator setup.
  • Google gains a more deployment-specific entry in Gemma 4, complementing the broader family’s reasoning and agentic-workflow positioning.

Second-order effects

  • Other open-model providers targeting developer adoption face additional pressure to pair multimodal capability with practical memory footprints, not only to compete on benchmark-scale performance.
  • Local-model users can more directly evaluate whether a 12B multimodal model fits their hardware and application constraints, making device memory a clearer procurement and design consideration.

Third-order effects

  • If releases continue to move capable multimodal models into modest local-memory envelopes, AI deployment may split more visibly between local inference for constrained or device-centric workloads and larger hosted models for heavier tasks.
  • The Gemma sequence suggests open-model competition is becoming two-dimensional: model capability and agentic features matter, but reproducible operation on accessible hardware is increasingly a product differentiator.

The trend: This is one data point in the push to make open multimodal AI usable on increasingly accessible local hardware rather than reserving advanced model features for large cloud deployments.

Discussion

  • @googlegemma @googlegemma on x
    Meet Gemma 4 12B! A unified, encoder-free multimodal model designed to bring high-performance intelligence directly to your laptop, and released under an Apache 2.0 license. Bridging the gap between edge efficiency and advanced reasoning. Here is what's new with Gemma 4 12B: 👇 [i…
  • @google @google on x
    Today we're introducing Gemma 4 12B — our latest open model that brings advanced agentic reasoning, vision and audio directly to your laptop. It delivers performance nearing our larger Gemma models with a much smaller total memory footprint, while being small enough to run [image…
  • @googleaidevs @googleaidevs on x
    We're launching Gemma 4 12B: Our unified, encoder-free model that brings powerful multimodal intelligence straight to your laptop 🚀 The model bridges the gap between our mobile E4B model and larger 26B MoE models, packaging frontier-class reasoning and native audio into a [image]
  • @mtschannen Michael Tschannen on x
    For the past years my research focus was on unifying models and training paradigms across modalities. Today I'm excited that we're releasing our latest model aligned with this theme: Gemma 4 12B, a dense encoder-free model which processes raw text, image, and audio inputs! 1/ [im…
  • @ollama @ollama on x
    .@GoogleDeepMind's Gemma 4 - 12B is available on Ollama! Chat: ollama run gemma4:12b-mlx Hermes Agent: ollama launch hermes —model gemma4:12b-mlx Claude Code: ollama launch claude —model gemma4:12b-mlx and more 👇👇👇 (Note, this currently works via MLX) [image]
  • @osanseviero Omar Sanseviero on x
    We collaborated with Hugging Face, llama.cpp, Ollama, VLLM, SGLang, Unsloth, MLX, LM Studio, and the rest of the ecosystem to land day 0 support. Enjoy! Read our developer guide: https://developers.googleblog.com/ ...
  • @unslothai @unslothai on x
    Gemma 4 12B can now run locally on just 8GB RAM via Dynamic GGUFs. Google's new model, Gemma 4 12B Unified supports image, audio and 256K context. You can run and train the model via Unsloth Studio. GGUF: https://huggingface.co/... Guide: https://unsloth.ai/... [image]
  • @tokumin Simon on x
    ✨If you have a base Macbook Pro or Air, you can run powerful multimodal local AI with the new Gemma4 12B model we shipped today. It's the perfect AI for poor connectivity or just to have on hand as a backup. Gemma team cooked (again).
  • @kaggle @kaggle on x
    Gemma 4 12B is now on Kaggle Models! 🤖 Learn more: 👉 https://www.kaggle.com/...
  • @ashkamath20 Aishwarya Kamath on x
    Gemma 4 Encoder-Free is here!! 🥳📷🔊 Super excited by this Gemma 4 12B model with great performance on vision and audio, without modality specific encoders. Closing in on the 26B while being 2x smaller in memory! Available now on Hugging Face, Kaggle, llama.cpp, and others! [image]
  • @googledevs @googledevs on x
    ✨ Introducing @GoogleGemma 4 12B, a unified open model bringing high-performance agentic multimodal intelligence directly to your laptop. Bridging the gap between edge efficiency and advanced reasoning, nearing 26B MoE at <50% the memory footprint. [image]
  • @demishassabis Demis Hassabis on x
    Celebrating the milestone of a massive 150+ million downloads of Gemma 4 with the release of the new Gemma 4 12B model! It's incredibly powerful for such a small model and it's tiny enough to run locally on a laptop with just 16GB VRAM. Apache 2.0 license - happy building!
  • @prince_canuma Prince Canuma on x
    🚀 Gemma 4 12B is here! We partnered with @GoogleDeepMind to bring and optimize their new dense and unifed multimodal model for Apple Silicon. ◈ 12B dense · 256K context ◈ Thinking mode (built-in reasoning) ◈ Vision: dynamic res, OCR, UI + charts ◈ Native audio: ASR + [image]
  • @chanduthota Chandu Thota on x
    Excited to introduce Gemma 4 12B. With Gemma 4 12B, we are bridging the gap between our edge-optimized E4B and the larger 26B MoE. Gemma 4 12B utilizes a single decoder-only transformer which shares the same structural design as the 31B Dense model, simplifying the local
  • @lmstudio @lmstudio on x
    Gemma 4 12B is here! Dense, mid-sized Gemma that fits right on your laptop - released by @google under Apache 2.0 Available now in LM Studio https://lmstudio.ai/...
  • @0xsero @0xsero on x
    8-16gb vram bros rejoice! You have a new best in class model ❤️ Thank you Google
  • @patloeber Patrick Loeber on x
    Gemma 4 12B is here! It comes with a new, unified architecture that removes separate multimodal encoders and enables local vision and audio understanding, plus advanced agentic reasoning
  • @osanseviero Omar Sanseviero on x
    Super excited to introduce Gemma 4 12B! 💎 - Multimodal: audio, image, video, and text input - Novel architecture: we removed the multimodal encoders for a unified, streamlined arch - New MacOS desktop app powered by LiteRT - MTP support Excited to see what you build with it! [vid…
  • @_philschmid Philipp Schmid on x
    We just launched a Gemma 4 12B! Our first mid-sized model with native audio inputs. Gemma 4 12 B is a unified, encoder-free multimodal model. 🧠 vision and audio directly into the LLM. 💻 Just need 16GB of memory. 📊 Benchmark nearing 26B. 📄 Apache 2.0. [image]
  • @andreaspsteiner Andreas Steiner on x
    Gemma 4 12B in action: Object detection, function calling, voice command, segmentation, language switch, translation - all of this and much more without vision/audio encoders! (Inputs and outputs are real, but FC2 data shown as code, and generation speedified) [video]
  • @sundarpichai Sundar Pichai on x
    Our new Gemma 4 12B model hits a sweet spot between size + performance: it can run locally on a laptop, while enabling powerful multi-step reasoning and agentic workflows. Can't wait to see what the community does with this one!
  • @mervenoyann Merve on x
    Google dropped Gemma-4 12B, it's a beast 🔥 > unified: audio + image go straight into model, no encoder > multimodal + tool calling > dense 12B with 256K context, comes with assistants for MTP (faster!⚡️) > day-0 in transformers, llama.cpp & MLX > A2.0 🤗 [image]
  • r/GeminiAI r on reddit
    Gemma 4 12B is fundamentally different from previous Gemma models
  • r/technology r on reddit
    Google's new open source Gemma 4 12B analyzes audio, video — and runs entirely locally on a typical 16GB enterprise laptop
  • @googledevs @googledevs on x
    Unlock local, agentic workflows with Gemma 4 12B and Google AI Edge, directly on your laptop. Experience 100% on-device AI: • Generate code in AI Edge Gallery (new to Mac) • Dictate and edit text via AI Edge Eloquent (new to Mac) • Serve Gemma 4 12B locally with LiteRT-LM [video]
  • r/LocalLLaMA r on reddit
    Introducing Gemma 4 12B: a unified, encoder-free multimodal model
  • r/Bard r on reddit
    Introducing Gemma 4 12B: a unified, encoder-free multimodal model
  • Gus Martins Gus Martins on linkedin
    yesterday was a busy day!  —  we released an amazing Gemma 4 12B model, that you can run on laptops with less than 16GB …
  • Google for Developers Google for Developers on linkedin
    💻 Gemma 4 12B is bringing agentic, multimodal intelligence directly to your everyday laptop. …
  • Donald Dew Donald Dew on linkedin
    Outside of code development, I think the initial round of bottom-line impact in the enterprise isn't going to be with the frontier models …
  • Lu Wang Lu Wang on linkedin
    A new #Gemma 12B model, perfect for your laptop, is available today!  🚀 Through #LiteRT/LM, you can easily unlock local …
  • r/google r on reddit
    Google's new Gemma 4 12B model is designed to run on any laptop with 16GB of RAM |  Gemma 4 12B uses a new encoding scheme and token prediction to punch above its weight.