Microsoft debuts Phi-3 Mini, a small 3.8B-parameter model about as capable as GPT-3.5, and plans Phi-3 Small and Phi-3 Medium models with 7B and 14B parameters
Microsoft launched the next version of its lightweight AI model Phi-3 Mini, the first of three small models the company plans to release.
The launch became the foundation for a broader Phi line: Microsoft later made Phi-3 generally available and introduced Phi-3-Silica for Copilot+ PCs in its Phi-3 availability expansion.
First-order effects
Microsoft adds a 3.8B-parameter model positioned near GPT-3.5-level capability, while setting expectations for 7B- and 14B-parameter Phi-3 variants.
Developers evaluating Microsoft models gain a forthcoming tiered family organized by model size, rather than a single lightweight option.
Second-order effects
The planned Mini, Small, and Medium lineup makes model selection more explicitly a trade-off between capability and deployment footprint, increasing pressure on competing model suppliers to offer clearer size and performance tiers.
A compact-model family creates a path for Microsoft to place AI functions closer to end-user devices; the later Phi-3-Silica integration into Copilot+ PCs shows that this line was not limited to cloud-hosted use.
Third-order effects
If compact models continue to narrow the gap with larger systems on targeted tasks, AI products are likely to adopt hybrid architectures that route work between local, smaller models and more capable remote models.
The Phi releases point toward competition based not only on frontier-model quality but also on deployability, tuning, and the range of model sizes available to builders.
The trend: AI vendors are productizing model families across size tiers so applications can match AI capability to latency, cost, and deployment constraints.
phi-3 is here, and it's ... good :-). I made a quick short demo to give you a feel of what phi-3-mini (3.8B) can do. Stay tuned for the open weights release and more announcements tomorrow morning! (And ofc this wouldn't be complete without the usual table of benchmarks!) [video]
Phi-3-mini to be released tomorrow is an important recall of the need for multidimensional evaluation. While the model sounds astonishingly good on tasks measured in (English-speaking) benchmarks, it is also mostly trained in English.
A few caveats about Phi-3: - The figure I attached at the beginning had some errors. Here's the updated one. - Phi-3-medium performs well on TriviaQA but noticeably underperforms rel. to GPT-3.5. We can guess that Phi-3 recipe doesn't magically make it understand more random... […
LLAMA 3 8B was amazing but will be overshadowed Phi-3 mini 4b, small 7b, medium 14b this week, and the benchmarks are fucking insane Synthetic data pipelines are massive improvements over internet data Flywheel only continues with big models too when these techniques are applied
phi-3-mini: 3.8B model matching Mixtral 8x7B and GPT-3.5 Plus a 7B model that matches Llama 3 8B in many benchmarks. Plus a 14B model. https://arxiv.org/... [image]
my ability to see a one day 4chan post leak something and then claim that it was whispered to him by god, and be able to Intuit it as an actual leak that I should pay attention to is actually astounding
Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone abs: https://arxiv.org/... Microsoft announces phi-3-mini, a 3.8B model trained on 3.3T tokens that rivals Mixtral 8x7B and GPT-3.5 Has same arch as Llama-2 to benefit open-source community Also... [ima…
Microsoft just released Phi-3 - phi-3-mini: 3.8B model trained on 3.3T tokens rivals Mixtral 8x7B and GPT-3.5 - phi-3-medium: 14B model trained on 4.8T tokens w/ 78% on MMLU and 8.9 on MT-bench https://arxiv.org/... [image]
Microsoft announces Phi-3 A Highly Capable Language Model Locally on Your Phone We introduce phi-3-mini, a 3.8 billion parameter language model trained on 3.3 trillion tokens, whose overall performance, as measured by both academic benchmarks and internal testing, [image]
if these scores hold up everyone could have something nearly as good as gpt-3.5-turbo on a phone phi-3-mini q4 should only need something like ~2gb of ram (mixtral 8x7b needs ~22gb for comparison) mmlu scores: - mini (3.8B): 68.8 - mixtral (8x7B): 68.4 - gpt-3.5 (20B) : 71.4
Amazing numbers. Phi-3 is topping GPT-3.5 on MMLU at 14B. Trained on 3.3 trillion tokens. They say in the paper ‘The innovation lies entirely in our dataset for training - composed of heavily filtered web data and synthetic data.’ [image]