πŸ€– AI Tools
Β· 5 min read

GLM-5.3 FlashX Guide: Pricing, 1M Context and API Setup


GLM-5.3 FlashX is the high-throughput version of Z.ai’s GLM-5.3 Flash. It keeps the same 320B-total, 18B-active architecture and 1,048,576-token context window, but Z.ai positions it for faster interactive work. The OpenRouter model listing prices FlashX at $0.37 per million input tokens, $1.25 per million output tokens, and $0.075 per million cache-read tokens.

That combination makes it interesting for coding agents, large-repository analysis, document workflows, and multimodal applications that need more than text. It accepts text, images, and video and returns text. It also supports tool calling and JSON output. The important caveat is that the speed figure of up to 200 tokens per second is a provider claim, not an independent guarantee for every request or region.

GLM-5.3 FlashX at a glance

FieldGLM-5.3 FlashX
OpenRouter IDz-ai/glm-5.3-flashx
Provider currently listedZ.ai
Architecture320B total, 18B active MoE
Context1,048,576 tokens
Maximum output131,072 tokens
InputText, image, video
OutputText
Input price$0.37/M tokens
Output price$1.25/M tokens
Cache read$0.075/M tokens
Tool callingYes
JSON outputYes, without documented JSON Schema enforcement

Prices and availability can change. Check the live OpenRouter model page before building a hard-coded budget or procurement comparison.

FlashX versus GLM-5.3 Flash

FlashX is not a new foundation-model family. It is a speed-oriented serving variant of GLM-5.3 Flash. Both use the same high-level architecture and target long-context multimodal work. The practical difference is the serving profile and price.

Use regular Flash when its lower-cost profile matters more than interactive latency. Evaluate FlashX when an agent needs to stream a response, make repeated tool calls, or keep a human waiting in an IDE. Do not assume that faster decoding automatically makes FlashX more accurate. Accuracy, latency, time to first token, and total task completion are separate measurements.

For a broader view of the company and model family, see what Z.ai is and the earlier GLM-5.2 guide.

OpenRouter API setup

OpenRouter exposes an OpenAI-compatible endpoint. Set the model explicitly so a future default or alias change cannot silently move the workload.

from openai import OpenAI

client = OpenAI(
    base_url="https://openrouter.ai/api/v1",
    api_key="YOUR_OPENROUTER_API_KEY",
)

response = client.chat.completions.create(
    model="z-ai/glm-5.3-flashx",
    messages=[
        {
            "role": "user",
            "content": "Review this migration plan and return the three highest-risk steps."
        }
    ],
    temperature=0.2,
)

print(response.choices[0].message.content)

Keep the API key in an environment variable in a real application. Our AI API key and secret-management guide covers local, CI, preview, and production environments.

Tool calling and structured output

FlashX accepts tools and tool_choice, which makes it usable for agent loops and application backends. A model choosing a function does not make the function safe. Validate arguments, enforce permissions outside the model, and require human approval for consequential actions. The tool-calling guide explains that execution boundary.

OpenRouter also lists response_format support for JSON output, but not strict JSON Schema enforcement. Treat the returned object as untrusted input. Parse it, validate it with a schema library, and retry or fail safely when required fields are missing. See validating AI responses with Zod for a TypeScript pattern.

Multimodal requests

GLM-5.3 FlashX accepts images and video as well as text. This can support screenshot analysis, visual QA, document extraction, and agents that inspect UI state. Multimodal context can become expensive quickly, and provider token accounting may not map intuitively to pixels or video duration. Test representative inputs before publishing a cost calculator.

The model returns text rather than generated images or video. For image generation, use a dedicated model such as Qwen-Image-2.1 instead.

Where FlashX fits

Coding agents: The long context can hold repository maps, specifications, logs, and tool results. Measure completed tasks rather than relying on model-level coding claims.

Large-document workflows: Cache pricing can help when the same large reference corpus is reused across requests, provided the provider reports cache hits as expected.

Interactive assistants: Higher advertised throughput can improve the reading experience after generation starts. Time to first token and tool latency can still dominate perceived speed.

Multimodal operations: Image and video input can help an agent inspect dashboards or UI evidence before selecting a tool. Keep permissions and evidence validation outside the model.

Cost example

A request with 200,000 uncached input tokens and 10,000 output tokens costs approximately:

  • Input: 0.2 Γ— $0.37 = $0.074
  • Output: 0.01 Γ— $1.25 = $0.0125
  • Total: $0.0865

If the 200,000 input tokens are served as cache reads, the input portion would be 0.2 Γ— $0.075 = $0.015, subject to the provider’s caching rules. This example is arithmetic, not a promise that every prompt qualifies for caching.

For cross-provider budgeting, use the guide to reducing LLM API costs and confirm live rates before purchase.

Limitations

  • The up-to-200-token-per-second figure is reported by Z.ai/OpenRouter.
  • Only Z.ai was listed as an upstream OpenRouter provider at publication time, so routing redundancy was limited.
  • JSON output is not the same as strict schema enforcement.
  • A 1M context window does not guarantee reliable retrieval across every token.
  • Independent agent and coding evaluations for FlashX remain limited.
  • OpenRouter availability does not prove a separate direct Z.ai endpoint has identical pricing or behavior.

My take

GLM-5.3 FlashX is a credible option when long context, multimodal input, and interactive generation matter together. Its $0.37/$1.25 price is low enough for serious evaluation, but the right comparison is a workload test against regular GLM-5.3 Flash and other fast APIs. Do not buy it on throughput alone. Measure first-token latency, tool-call accuracy, schema failures, cache hits, and completed-task cost.

FAQ

How much does GLM-5.3 FlashX cost?

OpenRouter listed $0.37 per million input tokens, $1.25 per million output tokens, and $0.075 per million cache-read tokens at publication time.

What is the context window?

The documented context window is 1,048,576 tokens, with up to 131,072 output tokens.

Is GLM-5.3 FlashX an open-weight model?

The GLM-5.3 Flash family is presented by Z.ai as open-weight, but API users should verify the exact weights, license, and serving variant they plan to use rather than assuming the hosted FlashX endpoint is identical to a local checkpoint.

Does it support tools and structured output?

It supports tool calling and JSON response formatting. OpenRouter does not document strict JSON Schema enforcement for this model, so applications still need validation.

Is FlashX better than GLM-5.3 Flash?

Not universally. FlashX targets speed. Regular Flash may offer a better cost profile. Test both on the same prompts, tools, latency target, and quality rubric.