DeepMind claims its language model RETRO matches the performance of neural networks 25 times its size, cutting the time and cost to train large language models
RETRO uses an external memory to look up passages of text on the fly, avoiding some of the costs of training a vast neural network
Today we're releasing three new papers on large language models. This work offers a foundation for our future language research, especially in areas that will have a bearing on how models are evaluated and deployed: https://dpmd.ai/... 1/ https://twitter.com/...
DeepMind says its new language model can beat others 25 times its size. Its secret is an AI with a twist: it's enhanced with an external memory. https://www.technologyreview.com/ ...
DeepMind announced Gopher, a 280B-param language model trained on 10.5 TB of text. It was evaluated across 152 benchmark tasks and was state-of-the-art for 81% of the tasks. Read more about it here: https://deepmind.com/... Check out some of the examples, very impressive! https:/…