/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Apple researchers detail MM1, a series of multimodal LLMs with up to 30B parameters they say achieve state-of-the-art performance across multiple AI benchmarks

Apple researchers have developed new methods for training large language models on both text and images, enabling more powerful …

VentureBeat Michael Nuñez

Context & Ripple Effects

MM1 extends Apple’s visible research push from language-guided image animation into models that jointly process visual and textual inputs.

It also precedes Apple’s later split between a smaller on-device model and a larger server-side model with Private Cloud Compute, outlined in its Apple Intelligence model architecture. The intervening OpenELM release shows Apple was also exploring much smaller models for device execution.

First-order effects

  • Apple gains a published multimodal training and benchmarking reference point at sizes up to 30B parameters, strengthening its internal foundation-model research base.
  • The work makes multimodal capability—not text-only language performance—a clearer evaluation target for Apple’s model development.

Second-order effects

  • Competing model labs face additional pressure to demonstrate visual-and-language performance across comparable benchmark suites, rather than relying on text-model results alone.
  • The contrast between MM1’s larger research models and Apple’s smaller OpenELM family reinforces the need to optimize different models for cloud-scale capability and device constraints.

Third-order effects

  • If this research-to-product pattern persists, leading device platforms are likely to treat multimodal models as a portfolio problem: larger models for demanding tasks and compact models for local execution.
  • That would deepen hybrid AI architecture as a competitive differentiator, with control of hardware, operating systems, and cloud infrastructure shaping how models are deployed.

The trend: MM1 is an early data point in the shift from standalone text LLMs toward multimodal model portfolios distributed across devices and private cloud infrastructure.

Discussion

  • @jradoff @jradoff on x
    I didn't expect Apple to publish one of the most detailed, open papers on LLM training: https://arxiv.org/... I suspect they felt the need to publish results to attract AI researchers and developers.
  • @kwekuoa Kweku Opoku-Agyemang, Ph.D on x
    Apple being rather open with its LLM research: exact scaling law coefficients, MoE settings, and even optimal learning rate functions. https://arxiv.org/... [image]
  • @mckbrando Brandon McKinzie on x
    Pre-training ablations: choice of image encoder, impact of image resolution, choice of vision-language bridge, impact of number of tokens per image, and impact of various pre-training data sources and their relative mixture weights. [image]
  • @mckbrando Brandon McKinzie on x
    Impact of scale: Importantly, we provide results at a larger scale than reported in prior surveys/ablations for multimodal LLMs. This includes sharing how we arrived at certain hyperparameters like learning rates for our largest model, MM1-30B. [image]
  • @mckbrando Brandon McKinzie on x
    This is just the beginning. The team is already hard at work on the next generation of models. Huge thanks to everyone that contributed to this project!
  • @mckbrando Brandon McKinzie on x
    Few-shot mixed-resolution CoT: we can keep the strong few-shot capabilities learned from multimodal pre-training even after instruction-tuning: MM1-30B-Chat achieves 39.4 zero-shot on MathVista, but with eight-shot CoT mixed-resolution prompting we can achieve 44.4. [image]
  • @drjimfan @drjimfan on x
    We live in such strange times. Apple, a company famous for its secrecy, published a paper with staggering amount of details on their multimodal foundation model. Those who are supposed to be open are now wayyy less than Apple. MM1 is a treasure trove of analysis. They discuss... …
  • @mckbrando Brandon McKinzie on x
    Pretraining evals: MM1-30B outperforms existing published results for pre-trained models (that have not undergone instruction-tuning): [image]
  • @mckbrando Brandon McKinzie on x
    Thrilled to share MM1!. The MM1 series of models are competitive with Gemini 1 at each of their respective model sizes. Beyond just announcing a new series of models, we also share the ablation results that guided our research process (🧵). [image]
  • @arankomatsuzaki Aran Komatsuzaki on x
    Apple presents MM1, a family of multimodal LLMs up to 30B parameters, that are SoTA in pre-training metrics and perform competitively after fine-tuning https://arxiv.org/... [image]
  • @alexttoshev Alexander Toshev on x
    Excited to announce our Multimodal LLM! The model is SOTA on a range of benchmarks in few-shot settings. And more importantly we give insights that hopefully will help folks out there build such models! J/ work with an excellent group of colleagues. https://arxiv.org/... [image]
  • @pdufter Philipp Dufter on x
    Check out MM1 for detailed insights how to train vision language models! Incredibly grateful and honoured that I had the chance to be part of this amazing team 🙏 https://arxiv.org/...
  • @l2k Lukas Biewald on x
    Makes me proud to see the @weights_biases team acknowledged in Apple's just released paper on MM1. https://arxiv.org/... [image]
  • @mckbrando Brandon McKinzie on x
    For more details, please check out our technical report here: https://arxiv.org/...
  • @danielgross Daniel Gross on x
    The slumbering dragon awakens:
  • @billardkarr Bill Karr on x
    Tim Apple has entered the race!
  • @erhartford Eric Hartford on x
    I am finally forced to take the time to give a serious look at multimodal and figure out how it works. One day, relatively soon, people will be saying “Your AI used to not understand images? that's crazy!” Like today's youth can't comprehend life before cell phones.
  • @methekarthik Karthik Kannan on x
    Apple has been super active in the multi-modal AI consistently putting out stuff these past months. my guess is iOS 18 will come with multimodal Siri capable of handling text and images.
  • @nickadobos Nick Dobos on x
    Apple is in the LLm game! Benchmarks seem pretty comparable to Gemini and a bit behind GPT4? Seems promising! Cmon plz do something cool with Siri [image]
  • @_akhaliq @_akhaliq on x
    Apple announces MM1 Methods, Analysis & Insights from Multimodal LLM Pre-training In this work, we discuss building performant Multimodal Large Language Models (MLLMs). In particular, we study the importance of various architecture components and data choices. Through [image]
  • @altryne Alex Volkov on x
    New MultiModal research from Apple 👀 They show a 30B VLM that is competitive in the pre-train landscape (models without instruction finetuning) with other large models. Excited to see where they will take this and if we'll see the weights 👀
  • @gowthami_s @gowthami_s on x
    I literally said the same thing in my recent paper. (will arxiv it soon!) 😅 “However, most of the works, both open and closed, release close to nothing about the process they have undergone to arrive at their algorithmic design choices, especially regarding multimodal...
  • @scobleizer Robert Scoble on x
    What if Apple open sourced Siri when it goes multi-modal and multi-model (expected to hear about that in June at WWDC). Apple has been a real asshole company on several fronts lately. Just ask @TimSweeneyEpic. Developer support is tepid. Belief is waning. It is time for...
  • @morgymcg Morgan McGuire on x
    The Apple research team behind MM1 giving a shoutout to the wandb crew supporting their work, love to see it 😍 🍎 + 🪄🐝 [image]