# AWS adds Z.ai's GLM 5.3 model to Amazon Bedrock

> The 753‑billion‑parameter GLM 5.3 model is now usable through Bedrock’s managed APIs for coding and security tasks.

Oossa · 2026-10-05 · https://oossa.com/en/aws-adds-z-ai-s-glm-5-3-model-to-amazon-bedrock

Amazon Web Services announced that the GLM 5.3 model from Z.ai (also known as Zhipu AI) is now available on its Bedrock platform. The model has 753 billion parameters and is built as a mixture‑of‑experts system, meaning it can route parts of a request to specialised sub‑models for efficiency. Bedrock users can call GLM 5.3 through the same OpenAI‑compatible APIs they already use, without running any hardware themselves.

## What does GLM 5.3 bring?

GLM 5.3 is tuned for long‑running coding jobs and agentic workflows that need to keep large amounts of context, such as refactoring a code base that spans hundreds of files. Z.ai also highlights strong cyber‑security performance, citing a top score of 84.5 on the CyberGym benchmark. On Bedrock the model benefits from prompt caching, cross‑Region inference, and three service tiers (Flex, Standard, Priority) that let users trade cost for latency.

## How to start using it

Eligible enterprise customers can enable the model from the Bedrock console’s Playground tab – no code is required. For developers, the model can be invoked with the OpenAI‑compatible Responses or Chat Completions APIs, or with Bedrock’s native Invoke and Converse calls. The blog post also shows a sample Python script that sends a refactoring request to the model using short‑lived AWS tokens.

## The facts

- GLM 5.3 has 753 billion parameters and is a mixture‑of‑experts model.
- Z.ai reports a 50 % improvement over GLM 5.2 on its internal coding benchmark.
- The model achieved a score of 84.5 on the CyberGym security benchmark at release.
- GLM 5.3 is accessible on Bedrock through US cross‑Region (us.zai.glm‑5.3) and Global cross‑Region (global.zai.glm‑5.3) inference profiles.
- Availability announced on 2026‑10‑06 for eligible enterprise customers.

## Why it matters

Enterprises can now run a very large coding and security model without buying GPU servers, lowering upfront costs and simplifying compliance. The built‑in prompt caching can cut latency and token charges for long‑running code‑analysis sessions. However, the model is limited to eligible enterprise accounts and pricing details are not disclosed, so exact cost savings remain uncertain.

## Sources & references

1. [Introducing GLM 5.3 on Amazon Bedrock](https://aws.amazon.com/blogs/machine-learning/introducing-glm-5-3-on-amazon-bedrock/) – AWS, 2026-10-05
2. [GLM 5.3 by Z.ai is now generally available on Amazon Bedrock](https://aws.amazon.com/about-aws/whats-new/2026/10/amazon-bedrock-glm-5-3/) – AWS, 2026-10-05

Last updated: 2026-10-06
