Databricks releases Dolly 2.0, the next version of its LLM released two weeks ago, and a dataset trained on 15K records generated by its employees
Today Databricks released Dolly 2.0, the next version of the large language model (LLM) with ChatGPT-like human interactivity …
Context & Ripple Effects
Two weeks after Databricks open sourced Dolly as a trainable-in-hours clone of Stanford's Alpaca, it is back with Dolly 2.0 — same ChatGPT-style interactivity, but now paired with a dataset of 15,000+ instruction records written by Databricks' own employees. That detail is the story: the original Dolly's lineage traced to someone else's instruction-tuning work, while 2.0's training corpus is generated in-house.
The release slots into a fast-building Databricks AI arc — LakehouseIQ's natural-language interface for enterprise data followed in June — making Dolly 2.0 less a research artifact than the foundation of the company's push to own the model layer under its data platform.
First-order effects
- Any company can now take Databricks' employee-written instruction dataset as a working template, lowering the barrier to building instruction-tuned models without relying on a closed provider's outputs.
- Databricks' own staff become the training corpus, giving the company a self-owned, commercially unencumbered dataset its later models — including the ~$10M DBRX — can build on.
Second-order effects
- Rivals in the open-LLM race face pressure to match the dataset move, not just the model release: whoever publishes usable instruction data alongside weights sets the terms others iterate on.
- The Dolly line feeds directly into Databricks' product surface — LakehouseIQ and later AI/BI — so the open release doubles as marketing for a paid stack that turns natural-language questions into data answers.
Third-order effects
- The arc from a one-machine, three-hour Alpaca clone to a multi-million-dollar DBRX shows open-source LLMs scaling from novelty reproductions into serious competitors to closed labs like OpenAI, with open weights as the distribution wedge.
- If in-house-generated training data becomes the norm, instruction datasets shift from scraped or borrowed corpora to curated, owned assets — a durable differentiator in an industry where model architectures converge quickly.
The trend: Open-source LLM development is maturing from cheap clones of closed models into self-generated training data and increasingly expensive in-house models, with Databricks using open releases to anchor a commercial data-platform AI stack.