These startups are chasing the next big thing in LLMs
News Source : MIT Technology Review
News Summary
- Transformers are the engines inside every major large language model on the market.
- The key strength of transformers lies in a mechanism called dense attention.
- Dense attention can capture the meaning of text with remarkable accuracy.
- But as the length of that text grows, the number of computations needed to process it adds up fast.
- If LLMs are to carry out harder tasks, they will need to take in larger amounts of data.
- But transformers struggle with what many of the latest models are designed to do.
It turns out this process works on text too. Inception has trained its LLMs to take a random string of words and turn it into sentences that make sense.
Never miss a story from us, subscribe to our newsletter