Perplexity is licensing Yelp's data and integrating its maps, reviews, and other details in responses when users ask for restaurant recommendations
Emilia David / The Verge : X: @aravsrinivas and @carnage4life X: Aravind Srinivas / @aravsrinivas : the rebels are partnering against the establishment [image] Dare Obasanjo / @carnage4life : The unexpected winners of the AI gold rush are all the user generated content sites like Reddit, Tumblr & Yelp that are selling user data to AI companies to train LLMs. Everything you've ever published online is being used to train AI that will one day come for your job. Ironic. [image]
Context & Ripple Effects
This is an early example of user-generated-content platforms turning their local knowledge into a licensed input for AI answers. Related coverage had already signaled the model: Stack Overflow planned to charge AI developers for access to its Q&A corpus, while Reddit reached a reported AI-training data deal.
For Yelp, the arrangement puts its maps, reviews, and business details inside an answer-engine workflow rather than limiting their use to Yelp’s own interface. Later coverage of Yelp’s real-time recommendations deal with OpenAI shows that local-data distribution through AI assistants became a repeatable channel.
First-order effects
- Perplexity can enrich restaurant-recommendation responses with Yelp-supplied maps, reviews, and business details, making Yelp a named data source in that use case.
- Yelp gains a licensing and distribution relationship that extends its local content into Perplexity’s responses.
Second-order effects
- Other review, forum, and reference platforms face stronger incentives to package high-value data for AI access rather than rely solely on open web indexing or traffic referrals.
- Answer engines competing for local-intent queries will need comparable licensed or otherwise authorized data sources, raising the importance of data partnerships in recommendation quality.
Third-order effects
- If such arrangements proliferate, the economics of answer engines may shift from crawling public pages toward negotiated access to differentiated, frequently updated datasets.
- The boundary between authorized licensing and unauthorized collection becomes more consequential, as illustrated by Reddit’s later lawsuit accusing Perplexity and others of illicit scraping.
The trend: AI answer products are evolving into distribution channels that pay or partner for proprietary content and local-data inputs.