The Allen Institute for AI debuts Multimodal Open Language Model in 1B- to 72B-parameter sizes, the most capable open-source AI model with visual abilities yet
A compact and fully open source visual AI model will make it easier for AI to take control of your computer—hopefully in a good way.
Against that backdrop, AI2’s release matters because it puts a broad size range of openly available visual-language models alongside a field in which Microsoft was reportedly building an in-house model intended to compete at the frontier.
First-order effects
Developers and researchers gain a fully open-source multimodal model family spanning 1B to 72B parameters, giving them more options to build and test systems that interpret visual inputs.
AI2 becomes a more consequential supplier of open model infrastructure, with its visual-capability claim setting a direct performance comparison point for other open-model providers.
Second-order effects
Open multimodal-model developers will face pressure to demonstrate comparable visual performance, model-size coverage, and practical usability rather than competing only on text benchmarks.
Tool builders seeking AI that can work from what it sees—including computer-interface contexts described by the article—can evaluate an open alternative instead of relying solely on proprietary model access.
Third-order effects
If capable visual-language models continue to become openly available, control over multimodal AI may shift from access to a small set of hosted models toward differentiation in tooling, deployment, and safeguards.
The expansion of open visual AI also makes governance more consequential: capabilities that can interpret screens or images broaden the need for controls around how downstream systems act on that information.
The trend: This is one data point in the industrialization of multimodal AI, as visual understanding spreads across both frontier proprietary systems and reusable open model infrastructure.
Molmo is good, and unlike the llama vision models prolly coming later, the giant, novel and useful dataset is going to be available for you to build on 🫡 @mattdeitke crushed building this dataset. Molmo can point and tell the time - not sure any other open weight VLMs can.
Really cool stuff from @allen_ai , great to see ‘pointing’ becoming a thing for VLM->We also release our data and checkpoint https://robo-point.github.io/
✅ Outperforming GPT-4o, Gemini 1.5 Pro & Claude 3.5 across 11 benchmarks! 🚀Only slightly surpassed by GPT-4o on the largest human preference study for VLMs with over 320k pairwise comparisons across nearly 1000 users. #AI #VLM #OpenSource [image]
Meet Molmo: a family of open, state-of-the-art multimodal AI models. Our best model outperforms proprietary systems, using 1000x less data. Molmo doesn't just understand multimodal data—it acts on it, enabling rich interactions in both the physical and virtual worlds. Try it [vid…
Open-source models beating closed models will become more and more common. Scaling has diminishing returns. The best solution will not have the largest scale but best approach or data. Especially with test-time compute, you do not need the best model to have the best solution.
Molmo by @allen_ai - a SOTA multimodal model 🤗Open models and partially open data 🤏7B and 72B model sizes (+7B MoE with 1B active params) 🤯Benchmarks above GPT-4V, Flash, etc 🗣️Human Preference of 72B on par with top API models 🧠PixMo, a high-quality dataset for captioning [image…
What have we here? Just in time for MultiModal models to be disrupted by Meta (reporting soon!), the great folks at @allen_ai releasing a multimodal MOLMO! 2 SOTA vision models in 1 day?? > With a 1B model that nearly matches GPT-4V?! > Molmo-72B, achieves the highest [image]
Molmo by @allen_ai - Open source SoTA Multimodal (Vision) Language model, beating Claude 3.5 Sonnet, GPT4V and comparable to GPT4o 🔥 They release four model checkpoints: 1. MolmoE-1B, a mixture of experts model with 1B (active) 7B (total) 2. Molmo-7B-O, most open 7B model 3. [vid…