live
10 labs · 30 models
All articles
8 min readChatPlus Team

GLM 5.3 Flash: how the “stealth” model Ox Alpha turned out to be Z.ai’s new network

For a couple of weeks, developers guessed what powerful nameless model had shown up on OpenRouter under the code name Ox Alpha. On August 26, China’s Z.ai showed its hand: it’s GLM 5.3 Flash — the first natively multimodal model in the GLM-5 series. Let’s dig into what it can do, where it came from, and whether you should trust it with your data.

ModelsZ.aiGLMNews
A glowing purple GLM crystal on a black background — cover art for the GLM 5.3 Flash article
Listen to the article00:00

For a couple of weeks, developer chats kept circling the same question: what kind of beast is this Ox Alpha? The model appeared in the OpenRouter catalog and in the open-source agent OpenCode without a single warning — no team name, no model card, not a line of documentation. A strong network just showed up out of nowhere, and on top of that it was being handed out for free. People ran it on their own tasks, traded screenshots, argued over who made it. On August 26, 2026, the mystery closed: the Chinese lab Z.ai stepped forward and said — yes, it's us, and the model is called GLM 5.3 Flash.

The story is amusing on its own, but behind it sits something more interesting than gossip about an anonymous release. Let's take it in order: what these "stealth" launches are, how the community identified the author before the announcement, what GLM 5.3 Flash can actually do, and — most importantly — whether it's worth switching to for everyday work.

Why release a model anonymously at all

The tactic is called a "stealth" release, and over the past year Chinese labs in particular have grown fond of it. The scheme is simple: you post the model on some shared platform like OpenRouter, but without any branding — under a neutral code name. You give access, often free, and simply watch what happens.

The point is to get honest feedback. When the model card says "a new model from some big-name lab," half of the reaction is expectations and reputation, not the model itself. But nobody fears a nameless "Ox Alpha" and nobody owes it anything: developers just put it to work, break it on their real tasks, and in those same chats write up where it's good and where it falls apart. For a team, that's a cheap way to pull metrics from live usage before firing up the PR machine.

One more signal was worth noting separately. OpenCode mentioned a budget on the order of a hundred trillion tokens a day — and all of it was being given away for free. Nobody burns that much "just for testing." It looks like someone was deliberately pushing an enormous stream of real requests through the model, not a handful of demo examples for a pretty announcement.

How the author was cracked before the announcement

The most interesting part is that Z.ai barely got the chance to surprise anyone. The community figured out Ox Alpha's lineage well before the official acknowledgment — and not from clever guesses, but from small technical clues.

First, the tokenizer. The way a model chops text into pieces is almost a family fingerprint. Ox Alpha's matched what people had already seen in GLM suspiciously well. Second, its behavior on touchy subjects: the model characteristically dodged questions about Chinese politics — exactly the way models trained inside China do. Add to that certain API error codes that also echoed known families. Individually each clue is a coincidence, but together they added up to a fairly confident "this is one of the Chinese frontier labs, most likely the GLM line."

And that's how it turned out. On August 26, Z.ai confirmed that Ox Alpha and GLM 5.3 Flash are one and the same model, and released its weights and benchmarks along with it.

What GLM 5.3 Flash is

Now to the substance. GLM 5.3 Flash is, in Z.ai's own words, the first natively multimodal model in the GLM-5 series. "Natively" is the key word here: the model wasn't bolted onto a separate image recognizer after the fact — it was trained from the start to handle different data types as a single stream.

Under the hood is a Mixture-of-Experts (MoE) architecture. To put it very plainly: instead of one giant network that strains in its entirety on every request, the model is split into many "experts," and only a small subset of them fires for each token. Formally GLM 5.3 Flash has 320 billion parameters, but only about 18 billion are active at any moment. Thanks to this, the model behaves like something large and smart while costing closer to something small. On top of that it uses hybrid sparse-linear attention — a mechanism that helps hold accuracy over very long context without ballooning the compute bill.

The key characteristics, boiled down to a single list:

  • Context — up to 1 million tokens. That's literally entire codebases or thick documents loaded in a single pass.
  • Max output — up to 131,072 tokens at once.
  • Input — text, images, and video.
  • Specialization — code and agentic scenarios: multi-step tasks where the model calls tools itself and carries the job through to a result.
  • Weights — published on HuggingFace under the MIT license. That is, the model is genuinely open: you can download it and run it yourself.

The last point is worth emphasizing. An open MIT license isn't "look and admire" — it's permission to use the model in your own products with almost no restrictions. For a frontier model of this caliber, that's still a rarity.

How good is it

Z.ai reports its numbers on its internal benchmark, Code Bench v1.0. By it, GLM 5.3 Flash "clearly surpasses" the previous GLM-5.2 at every difficulty level and "performs on par" with Claude Opus 4.8 in programming and agentic tasks. You should always treat internal benchmarks with a healthy dose of skepticism — they're scored by the same team that trained the model — but even adjusting for that, the claim is serious. Keeping pace with Opus 4.8 in code is the level of top Western models, not a "solid mid-tier."

And here the second half of the story comes in — price. At the provider, GLM 5.3 Flash is frankly cheap for its class:

Token typePrice per 1M
Input$0.15
Output$0.50
Cached input$0.03

For comparison: comparably strong Western models are usually several times more expensive, sometimes by an order of magnitude. This combination — "almost like Opus in code, but at a lightweight model's price" — is the main reason there was so much noise around Ox Alpha before anyone knew whose it was.

What's the catch

It doesn't come entirely without caveats, and it's more honest to name them up front.

First — privacy. The provider that handed out the model keeps your prompts and responses. Formally not for training, but the fact remains: your requests are stored somewhere. For public tasks that's not a problem, but work secrets, commercial correspondence, or personal data are better kept out of a model like this.

Second — it's still a fresh release. The internal benchmarks look pretty, but only independent runs and months of live use will give the full picture. For now, treat the model as very promising but not yet seasoned by time.

GLM 5.3 Flash on ChatPlus

We added this model back in its "stealth" phase, under the name Ox Alpha, and updated it for the official release — switched to the current endpoint, fixed the name and prices. You can try it right now: open a chat with GLM 5.3 Flash. Two things worth knowing:

  • The model is available on the PRO subscription. No hoops to jump through.
  • We deliberately left it out of auto-selection. Because of that same story about requests being stored on the provider's side, we don't want the model to quietly become the default. Let GLM 5.3 Flash be a deliberate choice: you turn it on when you know why.

It feels like an excellent option for code and long agentic scenarios, especially when you need a large context and don't mind spending a whole codebase on a task. It's also a good excuse to compare Z.ai's fresh model against the familiar GPT, Claude, and Gemini in one window — on ChatPlus, you just switch the model in the chat.

In short

Ox Alpha turned out to be not a mysterious newcomer but GLM 5.3 Flash — Z.ai's new multimodal model. It's strong at code, holds a million tokens of context, costs pennies by the standards of its class, and ships under an open license. On the downside — the provider stores your requests, and the model is still too fresh to judge conclusively. It's definitely worth a try: on ChatPlus it's already available on the PRO subscription.

Try the best AI models on ChatPlus

GPT, Claude, Gemini, GLM, and other AI models in one window.