/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Databricks open sources Dolly, an LLM which the company says can be trained in less than three hours on one machine and is a clone of Stanford's Alpaca model

Big-data analytics firm Databricks Inc. has emerged as an unlikely player in the generative artificial intelligence space …

SiliconANGLE Mike Wheatley

Context & Ripple Effects

Dolly extends Databricks’ established use of open source as a platform strategy, following its move to place Delta Lake under Linux Foundation stewardship. Making an instruction-following model available gives its data-platform audience a concrete entry point into generative AI.

The release also starts a product arc rather than standing alone: Databricks soon followed with Dolly 2.0 and an employee-generated instruction dataset, and later paired natural-language access to company data with LakehouseIQ.

First-order effects

  • Developers and Databricks customers can inspect, run and adapt a compact instruction-following model without depending solely on a hosted model provider.
  • Databricks gains an open-source AI artifact that can draw developers toward its broader data and machine-learning platform.

Second-order effects

  • Enterprise AI vendors face more pressure to distinguish hosted offerings through managed deployment, data integration and support rather than access to an instruction-following model alone.
  • The subsequent release of Dolly 2.0 with a dedicated instruction dataset shifts attention toward the provenance and usefulness of training data, not just whether model code is available.

Third-order effects

  • If this pattern persists, open models become a lower-cost experimentation layer while differentiation concentrates in trusted distribution, enterprise data access and operational tooling.
  • Databricks’ later much more resource-intensive DBRX model effort suggests an emerging two-tier market: broadly accessible models for adoption and higher-investment models for frontier performance.

The trend: Generative-AI competition is moving from model access alone toward the combination of open models, curated data and enterprise data-platform integration.

Discussion

  • @matei_zaharia Matei Zaharia on x
    This work built on the open source 6-billion parameter GPT-J model from @AiEleuther, with the training data and method from @stanfordnlp's Alpaca and @yizongwyz and co's SELF-INSTRUCT.
  • @databricks @databricks on x
    Have 30 minutes to spare? That's all you'll need to train your own Dolly - an open source LLM - on Databricks and enjoy #ChatGPT-like capabilities. Check out our latest updates, including our GitHub repo, here 👇 https://www.databricks.com/...
  • @tprstly Theo on x
    Databricks launches a new open source ChatGPT rival. “We're calling the model Dolly — after Dolly the sheep, the first cloned mammal — because it's an open source clone of an Alpaca, inspired by a LLaMA.” LOL https://www.databricks.com/...