Developers Push to Build AI Models Native to Non-English Languages
¿Habla español? Parlez-vous français? You might, but today’s large language models may not speak either Spanish or French very well. It’s a shortcoming that many developers are trying to address, especially since making LLMs available in more languages may attract more users.
To give a sense of the issue, Meta Platforms said in a blog post last month, that more than 5% of the training data for its latest flagship model Llama 3 was in languages other than English, “to prepare for upcoming multilingual use cases.” That was a big improvement on Llama 2, where less than 2% of training data was known to be non-English. Still, foreign languages remain a tiny minority. Not surprisingly, Meta said it “did not expect the same level of performance” from Llama 3 in non-English languages as in English.
Meta has a big stake in making its AI products understand languages other than English. The company hopes to use its AI assistant to increase the appeal of its consumer apps. But only about a tenth of Facebook’s daily active users are in the U.S. and Canada, Meta said in February. It will be harder for Meta to reach the bulk of its 3.24 billion daily active users if its AI services are limited to English. Notably, Meta has launched Meta AI “in some English speaking countries,” CEO Mark Zuckerberg said last month, adding “we’ll roll out in more languages and countries over the coming months.”