GLM 5.3 Flash Free: Every Real Way to Use Z.ai's New Model (2026)

Julian Goldie — founder, AI Profit Boardroom
By Julian Goldie · 8 min read
Get The AI Profit Stack Join AIPB →
🎯 1,000+ done-for-you AI agent workflows 📅 5 live coaching calls / week with me 🛡️ 7-day refund + 30-day ROI guarantee 👥 3,000+ AI operators inside

You can get GLM 5.3 Flash free in two genuine ways right now — download the MIT-licensed open weights and run it yourself, or ride the launch discounts that currently cut the already tiny API price in half — and in this guide I will walk through both, plus the paid route that is so cheap it may as well be free. Z.ai released GLM 5.3 Flash on 26 August 2026, and per the announcement it is a 320-billion-parameter model with 18 billion active, natively multimodal, with a one-million-token context window, released under the MIT licence. It is the model I have been testing all week, and the free angles are real.

📺 Watch: Run GLM 5.3 Flash Free Forever!

🔥 Get the Agent OS as a free bonus: AI Profit Boardroom members get the full Agent OS zip, prompt libraries, daily tutorials and weekly live coaching calls. → Get inside

Quick backstory, because it is a good one. On 20 August an anonymous model appeared on a third-party API platform with a million-token context window, text, image and video input, and a price of zero. The community fingerprinted it back to the GLM family within 48 hours, and Z.ai's 26 August announcement confirmed the mystery model — previewed under the name Ox Alpha — was theirs all along. If you played with Ox Alpha during the free window, you were using GLM 5.3 Flash before it had a name.

What GLM 5.3 Flash Actually Is

Per the release notes and the OpenRouter model listing, GLM 5.3 Flash is a mixture-of-experts model with 320 billion total parameters and 18 billion active per token, built on a hybrid sparse and linear attention architecture designed for efficiency. It is the first GLM-5-series model Z.ai describes as natively multimodal: it accepts text, images and video as input and returns text. The listed context window is 1,310,720 tokens with up to 131,072 completion tokens, and it supports tool calling and JSON-formatted outputs — the two features that matter most if you want it running inside an agent rather than a chat window.

Z.ai positions it for efficient coding and long-horizon agent tasks at a fraction of flagship cost — the announcement prices it at a tenth of the flagship GLM 5.3 API rate, or a twentieth during a limited-time discount. Notably, the announcement also says it runs entirely on Chinese AI chips.

How to Get GLM 5.3 Flash Free

Here are the routes, best first. This is the section to bookmark.

  1. Run the open weights yourself. The model is released under the MIT licence with open weights, which means genuinely free, permanent, commercial-friendly use if you have the hardware to serve a 320B-A18B model. This is the only route that is free forever with no strings. The workflow is the same one I documented for the previous generation in my guide to running GLM 5.2 locally — the serving stack carries over, you are just pointing it at newer weights.
  2. The discounted API — near-free in practice. On OpenRouter the listed price is 0.05 dollars per million input tokens and 0.1667 dollars per million output tokens, with cache reads at 0.01 dollars per million — and several providers are running a 50 per cent discount through 9 September 2026. At those rates a heavy day of agent usage costs pennies. There are 21 providers behind the listing with automatic failover, so capacity has not been an issue in my usage.
  3. Free trial credit on API platforms. Most gateways hand out starter credit, and at Flash pricing that credit lasts a very long time. The Ox Alpha zero-price preview itself has ended, so do not go hunting for that listing.

In the video above I also walk through how we wire GLM 5.3 Flash into the Agent OS — our system of pre-built business workflows for content, lead generation and SOP management — so the model powers an entire automation stack rather than a chat box. That is where cheap tokens turn into actual leverage.

Want the exact Agent OS workflows from the video — content engines, lead gen and SOP automation — pre-built and ready to point at GLM 5.3 Flash? They ship free with AI Profit Boardroom membership → Grab the workflows

📺 Watch: This NEW Chinese AI Model Is Seriously Powerful

Using GLM 5.3 Flash Inside an Agent

The reason this release matters to me is not the chat experience — it is what a million tokens of multimodal context at near-zero cost does inside an agent harness. Hermes agent added GLM-5.3-Flash to its model options in the v0.20.6 release on 27 August, per the Hermes release notes, so on a current install it is now a picker entry rather than a manual config job. If you are on an older build, my Hermes changelog tracker covers what each recent version shipped, and the wiring steps I documented for running GLM 5.2 inside Hermes agent apply unchanged to 5.3 Flash — swap the model identifier and you are done in five minutes.

Where does it earn a slot? In our own testing on Goldie Bench, the pattern with Flash-class models is consistent: they will not beat a frontier flagship on the hardest reasoning, but for high-volume agent work — research runs, drafting, data extraction, long-document processing — the cost-per-useful-output is what decides the winner, and at 0.05 dollars per million input tokens the maths is absurd. My fuller writeup of the previous generation, the GLM 5.2 Hermes agent review, explains the evaluation framework I use; 5.3 Flash slots into the same harness with a bigger context window and image and video input on top.

The multimodal part deserves a concrete example: an agent that can watch a screen recording or read a stack of screenshots as native input, across a million-token context, changes what an audit or research workflow can do. That was flagship-only territory a few months ago. It is now available at Flash pricing.

GLM 5.3 Flash Pricing at a Glance

RouteCostCatch
Open weights (MIT licence)Free foreverYou need serious hardware to serve 320B total parameters
OpenRouter API0.05 dollars per 1M input / 0.1667 per 1M output tokens50 per cent discount at several providers runs to 9 September 2026
Cache reads0.01 dollars per 1M tokensRequires prompt-caching-aware setup
Ox Alpha previewWas freeEnded — it was the pre-launch test, now confirmed as GLM 5.3 Flash

How It Compares to the Other Free-to-Cheap Options

GLM 5.3 Flash lands in a crowded fortnight for open Chinese models — Tencent's Hy4 preview arrived two days later, and Moonshot's Kimi K3 weights have been free to download since late July. If your priority is a free frontier-class generalist, Kimi K3 free is still the reference route. If your priority is the cheapest multimodal long-context workhorse for agents, that is exactly the slot 5.3 Flash was built for, and per the announcement it is the one running at a twentieth of flagship cost during the discount window. My Chinese AI models overview maps the whole field if you want the landscape view before committing.

GLM 5.3 Flash Free FAQs

Is GLM 5.3 Flash actually free?

The weights are — MIT licence, open download, commercial use allowed. API access is paid but tiny: 0.05 dollars per million input tokens on OpenRouter, currently half price at several providers until 9 September 2026.

What was Ox Alpha?

The anonymous preview of this model. It appeared on a third-party platform on 20 August at zero cost, the community traced it to the GLM family within 48 hours, and Z.ai confirmed it at launch. That free preview window has ended.

Can GLM 5.3 Flash handle images and video?

Yes — it is the first GLM-5-series model Z.ai describes as natively multimodal, accepting text, images and video as input and returning text. Combined with the million-token context, that makes it unusually capable for screenshot- and recording-heavy agent workflows at this price.

Does Hermes agent support GLM 5.3 Flash?

Yes. Per the Hermes release notes, v0.20.6 (27 August 2026) added GLM-5.3-Flash to the model options, so on a current install you select it rather than wiring it manually.

Verdict: The Best Price-to-Capability Deal of the Month

Free is a strong word and most "free AI" claims hide a catch, so here is the honest framing: GLM 5.3 Flash is genuinely free if you can self-host, and functionally free for most solo operators on the API — a working month of agent usage costs less than a coffee at the discounted rate. Add native multimodality, a million-token context, tool calling and day-one Hermes support, and this is the most practical model release of the month for anyone automating real business work. Test it against your own tasks before promoting it to production — but at this price, the test itself costs nothing worth counting.

If you want GLM 5.3 Flash earning for you by the weekend — the Agent OS workflows, the model configs and daily live help setting it all up — that is what the AI Profit Boardroom is for → Put Flash to work

Real wins from inside the AI Profit Boardroom

See all 3,000+ members →
AIPB member win screenshot AIPB member win screenshot AIPB member win screenshot AIPB member win screenshot AIPB member win screenshot AIPB member win screenshot AIPB member win screenshot AIPB member win screenshot AIPB member win screenshot AIPB member win screenshot AIPB member win screenshot AIPB member win screenshot

Ready To Join The #1 AI Community?

Join 3,600+ entrepreneurs inside the AI Profit Boardroom. Get 1,000+ plug-and-play AI agent workflows, daily coaching, and a community that holds you accountable.

Join The AI Community →

7-Day No-Questions Refund • Cancel Anytime

← Back to all posts