China's AI race splits over access, open weights and safety tests
Apple is training a China-only model with Alibaba, GLM-5.3 is withholding its weights pending review, and Vidraft's 117-item leaderboard puts access strategy and model safety on the same competitive map.
Artificial Intelligence··Midday
A China-only model and a local partnership
Three people familiar with the matter told Reuters that Apple has trained a large language model specifically for China. Alibaba supported the training, widening a partnership the two companies confirmed in February 2025. China's Cyberspace Administration registered Apple's generative AI service last month, and Apple Intelligence is expected to reach Chinese devices in the coming months after an iOS update. That version is described as carrying Alibaba's Qwen model and Baidu technology. Apple also published, then deleted without explanation, a guide showing Mac users how to connect Qwen to Siri and Writing Tools. Apple and Alibaba both declined to comment. The account rests on unnamed sources who described the information as sensitive and not public, and no technical detail of the model has been disclosed. Access in China is therefore taking shape through registration, partnership and local component choices as much as through any single model release.[1]
The interface is open; the weights wait
Z.ai released GLM-5.3 on the same base model as GLM-5.2 and attributed the gains to post-training. Interface access is open now; the company says the weights will follow in about two weeks, after safety evaluation and hardening. On figures Z.ai published itself, the model scores 28.3 on Terminal-Bench 3.0 against 4.6 for GLM-5.2, 66.9 on DeepSWE v1.1 against 46.2, and 28.5 on Agents' Last Exam CLI against 23.8, while trailing GPT-5.6 Sol and Claude Fable 5 on harder public coding tests. The company also reports 84.5 per cent on CyberGym, up from 77.2 per cent, and 54.4 per cent on ExploitBench, up from 24.4 per cent, and says its models have found 2,436 vulnerabilities across 269 open-source projects since GLM-5.2, of which 1,097 were rated critical or high. Z.ai ties the delay in publishing weights to an offensive security capability it says grew faster than expected. Every one of these measurements comes from the vendor and none has been independently verified. Holding open weights until a safety review finishes turns access into a product decision rather than a purely technical afterthought.[2]
A safety leaderboard and the competitive map
Vidraft put the AX-RAY leaderboard and its evaluation dataset on Hugging Face on 13 August. AX-RAY runs 117 diagnostic items and maps each one onto national legal and regulatory frameworks, adding, for Arab countries, religious and social norms alongside statute. The company reports catching causal-leakage signals in two general-purpose models, one of them from Nvidia, and says it reproduced the behaviour. Causal leakage describes a model acting on hidden information or unintended causal cues instead of its ordinary reasoning path; the safety literature has treated it largely as a theoretical risk, and catching it in a general-purpose model already in use has been the hard part. Chief executive Kim Min-sik framed safety as the next stage of the AI race. The findings are the company's own and no independent evaluator has repeated them; Vidraft also develops the AETHER foundation model and a quantum operating system, and sells safety diagnosis. A China-only access path, delayed open weights and a public safety leaderboard are advancing through different doors in the same news window: who reaches the model, when weights open, and how safety tests become visible competitive signals now sit on one map.[1], [2], [3]