Viral Sensation Ox Alpha Model Revealed As GLM-5.3-Flash, Running Entirely On Chinese Chips

ZeroHedge
Published
Viral Sensation Ox Alpha Model Revealed As GLM-5.3-Flash, Running Entirely On Chinese Chips

The short version

  • The Beijing-based company said it intends to price use of the model, now called GLM-5.3-Flash, at $0.15 per million input tokens and $0.50 per million output tokens…
  • That puts it alongside DeepSeek in the class of low-cost, very high-efficiency models that are attracting users away from premium-tier offerings from the likes of Anthropic PBC.
  • As part of the reveal, Zhipu AI launched its latest open-weight model, GLM-5.3-Flash, f/k/a Ox Alpha, saying that the system ran entirely on a cluster of 100,000 domestically…
  • In other words, not only is China dominating the open-weight model, it will soon dominate the hardware the is used to run it, precisely as we warned a week ago.
  • With open model token prices rising, one can guess what is going on at Anthropic and ChatGPT.

The story

Viral Sensation Ox Alpha Model Revealed As GLM-5.3-Flash, Running Entirely On Chinese Chips

China’s Z.AI (Zhipu) confirmed it’s responsible for the viral - and mysterious - Ox Alpha AI model that swept to the top of online usage charts this weekend, pushing its shares up as much as 12% on Thursday. The Beijing-based company said it intends to price use of the model, now called GLM-5.3-Flash, at $0.15 per million input tokens and $0.50 per million output tokens, or units of artificial intelligence work. That puts it alongside DeepSeek in the class of low-cost, very high-efficiency models that are attracting users away from premium-tier offerings from the likes of Anthropic PBC.

As part of the reveal, Zhipu AI launched its latest open-weight model, GLM-5.3-Flash, f/k/a Ox Alpha, saying that the system ran entirely on a cluster of 100,000 domestically produced chips during a high-profile stealth trial.

In other words, not only is China dominating the open-weight model, it will soon dominate the hardware the is used to run it, precisely as we warned a week ago.

Following the news, Zhipu’s shares closed more than 12% higher at HK$1,160 in Hong Kong on Thursday.

“What GLM-5.3-Flash confirms is a pattern that is no longer surprising — Chinese labs shipping near-frontier open models at a fraction of the Western price,” said Dermot McGrath, founder of Shanghai-based consultancy ZenGen Labs.

The announcement followed a week of heavy traffic on artificial intelligence model marketplace OpenRouter and agent platform OpenCode, where the model processed 62 trillion tokens before its formal release on Wednesday, according to Zhipu.

On OpenRouter, the system processed more than 23 trillion tokens in its first six days, making it the platform’s biggest launch to date.

Ox Alpha, as it was initially known, emerged over the weekend as an uncredited release on OpenRouter - the biggest launch in that marketplace’s history - and quickly gained traction among curious observers and users. It’s a reasoning model designed for coding and agentic tasks, and it can process text, image and video input, according to its description. The model is not far off from Anthropic’s Opus 4.8 on coding and agentic capabilities, Z.ai said in a blog post.

The deployment marks a significant test of China’s ability to handle large-scale global inference workloads on home-grown hardware, as Beijing seeks to reduce reliance on advanced processors from Nvidia amid tight export controls.

During its preview, Ox Alpha rapidly surged to the top of global usage rankings. According to OpenRouter data on Thursday, the model ranked first among coding systems on the platform, accounting for 10.3 trillion tokens, or nearly 31 per cent of its total weekly volume.
To overcome the lower memory capacity and bandwidth of individual Chinese chips compared with top-tier Nvidia graphics processing units, Zhipu – which operates internationally under the Z.ai brand – said it built a specialized inference engine that split processing stages into independently managed computing pools.

The firm said these architectural adjustments tripled end-to-end serving performance from its initial baseline, bringing hardware efficiency and per-token costs on par with mainstream Nvidia accelerators. The claims have yet to be independently verified.

While Zhipu did not name specific chip suppliers for this cluster, it has previously collaborated with top domestic semiconductor developers, including Huawei Technologies, makes of the increasingly popular Ascend chip, Cambricon Technologies and Moore Threads.

Cambricon said on Thursday it had achieved “Day 0” compatibility to serve GLM-5.3-Flash. Moore Threads said it also achieved “Day 0” support for the new model.

Featuring 320 billion total parameters, GLM-5.3-Flash activated just 18 billion per request to reduce computing overhead, according to Zhipu. It is also the first model in the GLM-5 series to natively process visual information alongside text.

Benchmarking firm Artificial Analysis gave the model a score of 57 on its Intelligence Index, placing it 10th globally and third among open-weight models, trailing Moonshot AI’s Kimi K3 and Alibaba Group Holding’s Qwen3.8 2.4T A95B.

Zhipu is touting aggressive pricing to win over international developers, offering GLM-5.3-Flash at 1/10th the rate of standard GLM-5.3 – dropping to 1/20th under a limited promotion. It claimed the new model cost about 1/40th as much as Anthropic’s Opus 4.8 at comparable intelligence levels.

Despite heavy traffic during the free trial, early developer feedback was mixed. While users praised the model’s ability to debug complex code – a community test showed that it solved 28 per cent of 175 LiveCodeBench problems – others reported occasional hallucinations, dropped tasks and sluggish generation. Artificial Analysis similarly noted that GLM-5.3-Flash’s output speed trailed the industry average.

Zhipu has released the model weights globally and integrated GLM-5.3-Flash across its application programming interface, ZCode platform, and GLM Coding Plan.

The launch coincides with intensified competition in China’s open-source ecosystem.

Separately, on Wednesday, Alibaba released Qwen3.8-Flash-Next, a multimodal preview of Qwen4 that it said activated 6 billion of its 125 billion parameters to similarly drive down inference costs. Alibaba owns the South China Morning Post.

Tyler Durden Thu, 08/27/2026 - 10:45
Read the full story at ZeroHedgeOriginal

Related Markets

All Markets
View full chart →
View Full Chart
View full chart →
View Full Chart
View full chart →
View Full Chart

Market data may be delayed. Not financial advice.

Powered by Gab AI

The Story At A Glance

Reading this article now — analysis appears below

Reading the article

💡 AI analysis provides alternative perspectives on current events

More to read

Recent stories from across the outlets Alto indexes.

Questions Alto can answer

From this story — each link opens a live data page or a tool already filled in.

  1. What is $100 from 1990 worth today?CPI-adjusted dollars — result on the next page
  2. Where does a $75,000 household income rank nationally?Census percentile — national and state
  3. What federal tax bracket is $80,000 (single)?Marginal and effective rate on the next page
  4. What's Alto covering on the Finance desk?Latest headlines on this beat
  5. What else is Alto tracking on Federal Reserve & Interest Rates?Topic hub with related coverage
  6. What else is Alto tracking on Inflation?Topic hub with related coverage

All toolsAll topicsSource directoryStory timelinesHeadline comparisonSearchMost read

From Gab Shop

Official merchandise. Every order funds free speech infrastructure.

Shop all products

Install Alto on your phone

Add Alto to your home screen for breaking news — no app store, no account.

  1. Step 1Open alto.gab.com in SafariMust be Safari — not Chrome or in-app browsers
  2. Step 2Tap the Share buttonSquare with an arrow, at the bottom of Safari
  3. Step 3Tap "More"If you don’t see Add to Home Screen yet
  4. Step 4Tap "Add to Home Screen"Scroll the share sheet if you need to
  5. Step 5Tap "Add"Alto appears on your home screen like any other app.
gab

Talk Markets Freely

Trade ideas, earnings, and the Fed with investors who aren't waiting on a moderator's approval.

What Makes Gab Different

We're not just another social network. We're a platform built on principles that matter.

Freedom of Speech & Reach

All First Amendment protected speech is welcome. No algorithmic throttling or shadow banning.

Family-Friendly Platform

We maintain a clean environment. Explicit adult content is strictly prohibited.

Western Nations Only

Third-world IPs are blocked. No scammers, no spam farms. Built for Western civilization.

Funded By Users

Our users are our investors and customers. You're not the product being sold.

Battle Tested

A decade of standing strong. Banned from app stores, banks—and still here.

American Owned & Operated

We reject foreign censorship demands. Built by Americans, for free people.

Support Alto & Gab

Alto is funded entirely by readers like you. Your donation helps us continue delivering curated news from a right-wing Christian Nationalist perspective, powered by Gab AI.