Americans Are Impressed with China’s AI
American pundits like former Google CEO Eric Schmidt like to say China is behind the U.S. in artificial intelligence by several years. But American AI researchers increasingly tell me that these people underestimate Chinese researchers. Instead, by restricting Chinese access to American tech, the U.S. has pushed the country to forge its own path in AI development (or find workarounds to the U.S. export rules, as in the case of ByteDance).
And Chinese developers are catching up. The latest example is an open-weight large language model for generating computer code that came out yesterday. The company behind it, a Chinese quant trading firm called High-Flyer Capital Management, said its model beat OpenAI’s GPT-4 Turbo, Google’s Gemini 1.5 Pro and Anthropic’s Claude 3 Opus on a number of standard, industry evaluations.
High-Flyer created the coding model, DeepSeek-Coder-V2, by training a general purpose large language model on 6 trillion tokens of web data, consisting mostly of software code or math problems, in addition to the more than 4 trillion tokens the original model was trained on. (In case you’ve forgotten, “tokens” refers to words or parts of words that AI models ingest.)