How AI-based natural language processing algorithms can be applied to biological data to create protein-language models and cut drug discovery time to months
Natural language processing algorithms like the ones used in Google searches and OpenAI's ChatGPT promise to slash the time required to bring medications to market Tweets: @_karenhao Tweets: @_karenhao : I've rarely seen a large language model that wasn't a marketing trick or prematurely released, ChatGPT included. But there's one area where I'm hopeful about the application of data-guzzling natural-language algorithms: for accelerating drug discovery. https://www.wsj.com/...
Context & Ripple Effects
This WSJ explainer was the early articulation of an idea the coverage has since tracked end to end: language models trained on biological sequences instead of text. London-based BenevolentAI showed the premise as far back as its COVID-era drug-repurposing work, arguing industry data plus machine learning could surface new uses for existing compounds.
Since then the thesis has commercialized on both sides of the market. Google Cloud packaged the capability as products for biotech customers in its AI-powered drug discovery tools launch, startups like Terray Therapeutics industrialized the data side by generating 50TB of raw experimental data daily, and Bloomberg flagged the unresolved question of whether AI-aided drugs can prove themselves in the clinic. By 2026, OpenAI had crossed over entirely with GPT-Rosalind, offered as a research preview to Moderna and Amgen.
First-order effects
- Pharma R&D teams gain a new tool class: protein-language models that treat amino-acid sequences like sentences, letting companies such as Moderna and Amgen — now GPT-Rosalind customers — screen candidate molecules computationally before committing lab time.
- Cloud providers turn drug discovery into a product line: Google Cloud's move means biotech firms can buy the capability rather than build it, shifting spend toward hyperscaler platforms.
Second-order effects
- Proprietary biological data becomes the competitive moat: Terray's 50TB-a-day generation pipeline shows that whoever owns wet-lab output, not just algorithms, controls model quality — pressuring smaller biotechs to partner or license their data.
- OpenAI's entry forces Google to defend a nascent but strategic vertical, turning life sciences from a goodwill showcase into a contested enterprise market between the two largest model vendors.
Third-order effects
- If AI-discovered candidates clear clinical proof — the hurdle Bloomberg identified as the industry's open challenge — drug discovery structurally shifts from wet-lab-bound to compute-and-data-bound, concentrating advantage in firms that control both.
- Regulators and payers will eventually need frameworks for evaluating AI-originated drugs differently from traditionally discovered ones, since the discovery method itself becomes part of the safety-and-efficacy conversation.
The trend: Foundation-model techniques are migrating from consumer search and chat into life sciences, where access to proprietary biological data — not model architecture alone — increasingly decides who wins drug discovery.