Hands-on with Ernie 3.5, which Baidu claims is “slightly inferior” to GPT-4 in a comprehensive test but performs better when both were spoken to in Chinese
Context & Ripple Effects
The New York Times hands-on lands three weeks after [[a:841496|Baidu claimed Ernie 3.5 had surpassed ChatGPT overall and beaten GPT-4 on many Chinese-language tasks]], and it partially validates that pitch: the model is genuinely competitive, just not superior. The test also closes a loop opened in March, when sources described Baidu scrambling to ready Ernie Bot ahead of its launch event because it struggled with basic tasks.
Why it matters: Baidu's credibility with enterprise and consumer users now rests on whether third-party evaluations back up its self-reported benchmarks, and this is one of the first independent reads on the flagship model.
First-order effects
- Baidu's marketing claim is now qualified by an independent test: Ernie 3.5 is 'slightly inferior' to GPT-4 across a comprehensive evaluation, so buyers comparing the two get a more sober baseline than Baidu's own announcements provided.
- The Chinese-language result is the durable finding — Ernie 3.5 outperforming GPT-4 when both are addressed in Chinese gives Baidu a defensible home-market wedge even where general capability lags.
Second-order effects
- Baidu's response path is already visible in its release cadence: within months it shipped Ernie 4.0 with Robin Li claiming parity with GPT-4, suggesting each independent gap reading accelerates the next version push rather than denting the strategy.
- Chinese rivals watching the same evaluations — Alibaba's Qwen and DeepSeek among them, per Baidu's later positioning — can compete on the same Chinese-language strength, compressing Baidu's differentiation at home even as it holds off GPT-4 there.
Third-order effects
- If the pattern holds, Chinese frontier models settle into a structure of domestic-language dominance plus a persistent general-capability gap closed incrementally by version cycles — making independent, bilingual evaluations the arbiter that vendor benchmarks cannot substitute for.
- For global buyers, the test points toward a bifurcated market where model choice depends on working language, weakening the assumption that a single Western frontier model serves all markets.
The trend: Chinese LLM makers are closing the gap with OpenAI through rapid version cycles and native-language strength, with independent hands-on tests increasingly policing the gap between vendor claims and real capability.