> ## Documentation Index
> Fetch the complete documentation index at: https://docs.matterai.so/llms.txt
> Use this file to discover all available pages before exploring further.

# GLM 5.3 Flash

> GLM 5.3 Flash API - fast, low-cost open model for everyday coding tasks with a 1M context window. Available on MatterAI.

GLM 5.3 Flash is a fast, low-cost open model for everyday coding tasks. It pairs the GLM 5.3 architecture with low latency and aggressive pricing, making it the default workhorse for high-volume applications.

## Specifications

| Specification | Value |
| - | - |
| Model ID | `zai/glm-5.3-flash` |
| Context Window | 232K tokens (1M inference window) |
| Max Output Tokens | 64K |
| Input Modalities | Text, Image |
| Output Modalities | Text |
| Supported Features | Tools, Structured Outputs, Web Search |
| License | Open weights |

## Pricing

| Type | Price (per 1M tokens) |
| - | - |
| Input | \$0.15 |
| Cached Input | \$0.03 |
| Output | \$0.50 |

<Note>
  A 30% automatic discount is applied to all GLM 5.3 Flash usage on MatterAI.
</Note>

## Quick Start

```bash cURL theme={null}
curl --request POST \
  --url https://api.matterai.so/v1/chat/completions \
  --header 'Content-Type: application/json' \
  --header 'Authorization: Bearer $MATTERAI_API_KEY' \
  --data '{
  "model": "zai/glm-5.3-flash",
  "messages": [
    {
      "role": "system",
      "content": "You are a helpful assistant."
    },
    {
      "role": "user",
      "content": "What is Rust?"
    }
  ],
  "stream": false,
  "max_tokens": 1000
}'
```

```javascript OpenAI NodeJS SDK theme={null}
import OpenAI from "openai";

const openai = new OpenAI({
  apiKey: process.env.MATTERAI_API_KEY,
  baseURL: "https://api.matterai.so/v1",
});

async function main() {
  const response = await openai.chat.completions.create({
    model: "zai/glm-5.3-flash",
    messages: [
      { role: "system", content: "You are a helpful assistant." },
      { role: "user", content: "What is Rust?" },
    ],
    stream: false,
    max_tokens: 1000,
  });

  console.log(response.choices[0].message.content);
}

main();
```

```python OpenAI Python SDK theme={null}
from openai import OpenAI

client = OpenAI(
  api_key=os.environ.get("MATTERAI_API_KEY"),
  base_url="https://api.matterai.so/v1"
)

response = client.chat.completions.create(
  model="zai/glm-5.3-flash",
  messages=[
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "What is Rust?"}
  ],
  stream=False,
  max_tokens=1000
)

print(response.choices[0].message.content)
```

## When to Use

GLM 5.3 Flash is built for volume: chat assistants, code completion, classification and extraction, summarization, and any workload where latency and cost matter more than maximum reasoning depth.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.