Apple researchers detail MM1, a series of multimodal LLMs with up to 30B parameters they say achieve state-of-the-art performance across multiple AI benchmarks
Apple researchers have developed new methods for training large language models on both text and images, enabling more powerful …
It also precedes Apple’s later split between a smaller on-device model and a larger server-side model with Private Cloud Compute, outlined in its Apple Intelligence model architecture. The intervening OpenELM release shows Apple was also exploring much smaller models for device execution.
First-order effects
Apple gains a published multimodal training and benchmarking reference point at sizes up to 30B parameters, strengthening its internal foundation-model research base.
The work makes multimodal capability—not text-only language performance—a clearer evaluation target for Apple’s model development.
Second-order effects
Competing model labs face additional pressure to demonstrate visual-and-language performance across comparable benchmark suites, rather than relying on text-model results alone.
The contrast between MM1’s larger research models and Apple’s smaller OpenELM family reinforces the need to optimize different models for cloud-scale capability and device constraints.
Third-order effects
If this research-to-product pattern persists, leading device platforms are likely to treat multimodal models as a portfolio problem: larger models for demanding tasks and compact models for local execution.
That would deepen hybrid AI architecture as a competitive differentiator, with control of hardware, operating systems, and cloud infrastructure shaping how models are deployed.
The trend: MM1 is an early data point in the shift from standalone text LLMs toward multimodal model portfolios distributed across devices and private cloud infrastructure.
I didn't expect Apple to publish one of the most detailed, open papers on LLM training: https://arxiv.org/... I suspect they felt the need to publish results to attract AI researchers and developers.
Apple being rather open with its LLM research: exact scaling law coefficients, MoE settings, and even optimal learning rate functions. https://arxiv.org/... [image]
Pre-training ablations: choice of image encoder, impact of image resolution, choice of vision-language bridge, impact of number of tokens per image, and impact of various pre-training data sources and their relative mixture weights. [image]
Impact of scale: Importantly, we provide results at a larger scale than reported in prior surveys/ablations for multimodal LLMs. This includes sharing how we arrived at certain hyperparameters like learning rates for our largest model, MM1-30B. [image]
This is just the beginning. The team is already hard at work on the next generation of models. Huge thanks to everyone that contributed to this project!
Few-shot mixed-resolution CoT: we can keep the strong few-shot capabilities learned from multimodal pre-training even after instruction-tuning: MM1-30B-Chat achieves 39.4 zero-shot on MathVista, but with eight-shot CoT mixed-resolution prompting we can achieve 44.4. [image]
We live in such strange times. Apple, a company famous for its secrecy, published a paper with staggering amount of details on their multimodal foundation model. Those who are supposed to be open are now wayyy less than Apple. MM1 is a treasure trove of analysis. They discuss... …
Thrilled to share MM1!. The MM1 series of models are competitive with Gemini 1 at each of their respective model sizes. Beyond just announcing a new series of models, we also share the ablation results that guided our research process (🧵). [image]
Apple presents MM1, a family of multimodal LLMs up to 30B parameters, that are SoTA in pre-training metrics and perform competitively after fine-tuning https://arxiv.org/... [image]
Excited to announce our Multimodal LLM! The model is SOTA on a range of benchmarks in few-shot settings. And more importantly we give insights that hopefully will help folks out there build such models! J/ work with an excellent group of colleagues. https://arxiv.org/... [image]
Check out MM1 for detailed insights how to train vision language models! Incredibly grateful and honoured that I had the chance to be part of this amazing team 🙏 https://arxiv.org/...
I am finally forced to take the time to give a serious look at multimodal and figure out how it works. One day, relatively soon, people will be saying “Your AI used to not understand images? that's crazy!” Like today's youth can't comprehend life before cell phones.
Apple has been super active in the multi-modal AI consistently putting out stuff these past months. my guess is iOS 18 will come with multimodal Siri capable of handling text and images.
Apple is in the LLm game! Benchmarks seem pretty comparable to Gemini and a bit behind GPT4? Seems promising! Cmon plz do something cool with Siri [image]
Apple announces MM1 Methods, Analysis & Insights from Multimodal LLM Pre-training In this work, we discuss building performant Multimodal Large Language Models (MLLMs). In particular, we study the importance of various architecture components and data choices. Through [image]
New MultiModal research from Apple 👀 They show a 30B VLM that is competitive in the pre-train landscape (models without instruction finetuning) with other large models. Excited to see where they will take this and if we'll see the weights 👀
I literally said the same thing in my recent paper. (will arxiv it soon!) 😅 “However, most of the works, both open and closed, release close to nothing about the process they have undergone to arrive at their algorithmic design choices, especially regarding multimodal...
What if Apple open sourced Siri when it goes multi-modal and multi-model (expected to hear about that in June at WWDC). Apple has been a real asshole company on several fronts lately. Just ask @TimSweeneyEpic. Developer support is tepid. Belief is waning. It is time for...