How Developers Gave Llama 3 More Memory
Developers have been praising Meta Platforms’ Llama 3, the latest version of its flagship large language model. But, as my colleague Stephanie and I explained, they have one big criticism: Llama 3’s context window is too short, at just over 8,000 tokens. (As a refresher, a context window is how much information a model can accept in a single query and a token is a word or part of a word.)
Meta told developers that it expected to release longer context windows in coming months. But some developers have taken matters into their own hands. The early movers include OthersideAI, which makes the generative copywriting platform HyperWrite and has used earlier versions of Llama in its tools. Last week, its co-founder and CEO Matt Shumer released a version of Llama 3 with a context window of 16,000 tokens.
Shumer said OthersideAI began work on its version because HyperWrite uses AI agents—or bots—to perform tasks such as research, finding data and operating software. “That’s really limited by 8K,” Shumer told me. With 16,000 tokens, “suddenly we actually have these possibilities open for us, where we can go in and read through a bunch of web pages and understand how to get the best possible answer to the user rather than just kind of guessing at it.”