/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Databricks releases Dolly 2.0, the next version of its instruction-following LLM released two weeks ago, with a dataset of 15K+ records generated by its staff

Sharon Goldman / VentureBeat :

VentureBeat Sharon Goldman

Context & Ripple Effects

Dolly 2.0 is a rapid follow-on to Databricks’ earlier open-sourcing of Dolly, extending the company’s LLM effort with a staff-generated instruction dataset rather than only a model release.

The release is an early step in Databricks’ broader progression from models toward data-platform AI features, later including a natural-language interface for querying company data.

First-order effects

  • Databricks adds a new Dolly version and a 15K+-record staff-generated dataset, giving developers a more concrete instruction-data resource alongside the model.
  • The company strengthens its position in open LLM development by coupling a model iteration with the data asset used to support instruction-following behavior.

Second-order effects

  • Other open-model providers face added pressure to show not just model availability but the provenance and usefulness of the instruction data surrounding it.
  • Teams evaluating self-hosted LLMs gain another option to assess for instruction-tuning workflows, shifting attention toward the quality and accessibility of accompanying datasets.

Third-order effects

  • If this packaging becomes standard, competition in open LLMs will increasingly center on reusable data, evaluation, and deployment tooling—not model releases alone.
  • Databricks’ subsequent moves toward natural-language data products suggest a longer-term push to make models a feature of the data platform, reinforcing deployment-layer control.

The trend: Open LLM competition is shifting from standalone model launches toward integrated distributions of models, instruction data, and enterprise data-platform workflows.

Discussion

  • @soumithchintala Soumith Chintala on x
    This is the new gold! “15,000 high-quality human-generated prompt / response pairs specifically designed for instruction tuning large language models” Commendable signal from @databricks for releasing it with a full open-source license. https://twitter.com/...
  • @simonw Simon Willison on x
    My quick notes on getting the new Dolly 2.0 LLM running on @HelloPaperspace using @huggingface transformers: https://til.simonwillison.net/ ... It took 12 minutes to generate a (very good) response so I'm clearly doing something wrong! https://twitter.com/...
  • @rxin Reynold Xin on x
    Free Dolly! Introducing the first *commercially viable*, open source, instruction-following LLM. We are open-sourcing the entirety of Dolly 2.0, including the training code, the dataset, and the model weights, all suitable for commercial use. https://www.databricks.com/...
  • @databricks @databricks on x
    Meet Dolly 2.0: the first open-source, instruction-following LLM that's available for commercial use & doesn't require you to pay for API access or share data with third parties. Now, anyone can create a powerful LLM that understands how to talk to people! https://www.databricks.…
  • @peteskomoroch Pete Skomoroch on x
    Databricks just released Dolly 2.0: the first open-source, instruction-following LLM licensed for commercial use. - 12B parameter LLM based on EleutherAI pythia model family - training code, dataset, and model weights all licensed for commercial use. https://www.databricks.com/..…