Introducing GLM 5.3 on Amazon Bedrock
Coding and agentic workloads are asking more of AI models than ever: refactor a repository spanning hundreds of files, sustain a multi-hour agentic workflow without losing context, and reason through complex systems problems with tool use at every step. Meeting those demands with open-weight models has historically meant provisioning and operating your own inference infrastructure.
GLM 5.3 from Z.ai (Zhipu AI) is now available on Amazon Bedrock. GLM 5.3, as published on Hugging Face Hub, is a 753B-parameter mixture-of-experts model optimized for coding and long-horizon agentic tasks. In particular, Z.ai has reported the model shows notable cyber security capabilities. On Amazon Bedrock, you can now use it through fully managed APIs with cross-Region inference, prompt caching, and service tiers. You don’t manage any infrastructure. Access to GLM 5.3 on Bedrock is available to eligible enterprise customers.
In this post, we show you how to invoke GLM 5.3 on Amazon Bedrock using the OpenAI-compatible APIs and reduce cost and latency with prompt caching. We then put the model to work in a realistic agentic workflow: running an authorized security test of your own application with Strix, an open-source AI penetration testing agent.
GLM 5 arrived on Amazon Bedrock earlier this year. GLM 5.3 builds on the same lineage, with a range of important gains:
For the following usage examples, you need:
Try GLM 5.3 on the Amazon Bedrock console
You can start sending prompts to GLM 5.3 on the AWS Management Console, with no need to write code or install developer tools. To get started, navigate to Amazon Bedrock and then choose Test > Playground from the left sidebar menu.
From this playground interface you can select GLM 5.3 from the model list and send your first prompts through the chat UI, as shown in the following screenshot:
Figure 1: Chatting with GLM 5.3 on the Amazon Bedrock console
Get started with the Responses API
Programmatically, you can call the model through the bedrock-runtime endpoint. This supports both the OpenAI-compatible Responses and Chat Completions APIs, and the Amazon Bedrock Invoke and Converse APIs for GLM 5.3. For new applications the OpenAI-compatible APIs are recommended as they support a more complete set of features.
Amazon Bedrock does support generating API keys for OpenAI-compatible integrations that require them. However, we strongly recommend preferring short-lived credentials over long-lived API keys where possible.
In the following example, we will call the Responses API from Python using the OpenAI Python SDK, and the aws-bedrock-token-generator library to generate short-term tokens from your standard AWS Command Line Interface (AWS CLI) credentials.
Long-running coding and knowledge workflows often resend stable context across multiple conversation turns, such as system prompts, tool definitions, or repository files.
GLM 5.3 on Amazon Bedrock supports implicit prompt caching by default, which helps reduce response latency and input token costs for repeated calls sharing the same initial prompt prefix.
With explicit prompt caching mode you specifically identify the reusable prompt prefixes, which can further improve cache hit rate (and therefore latency and cost savings) over implicit caching.
To use explicit prompt caching with GLM 5.3, as shown in the following example:
For more information, refer to the prompt caching section of the Amazon Bedrock User Guide.
Example agentic workload: Authorized security testing with Strix
One workload that benefits directly from GLM 5.3’s strengths is automated security testing of your own applications. Strix is an open-source AI penetration testing agent that runs your code dynamically, finds vulnerabilities, and validates them with proof-of-concept tests. As of this writing, the Strix documentation uses GLM 5.3 as its default model. You can configure Strix to use GLM 5.3 on Amazon Bedrock instead of a third-party inference provider, so model inference runs under your AWS account’s controls.
Only test applications you own or have explicit written permission to test. Unauthorized security testing of systems you don’t own is illegal in most jurisdictions and violates the AWS Acceptable Use Policy. In this walkthrough, the target is OWASP Juice Shop, a deliberately vulnerable sample application running locally on your machine.
If you want fully managed, continuous security testing beyond running open-source agents yourself, AWS Continuum provides on-demand penetration testing and other security analyses as a managed service. The two approaches are complementary: open-source agents like Strix give you developer-driven, in-the-loop, and deeply customizable testing against local builds, while AWS Continuum runs managed assessments at scale.
Strix spins up a team of sub-agents to map the threat surface, explore a range of potential vulnerability categories, and attempt to validate each finding with a working proof of concept. This helps minimize time spent triaging false positives. A successful run will generate a report including severity, evidence, and remediation guidance for each finding.
The following video shows the end-to-end journey of setting up and running Strix against the example application, and exploring the results:
Figure 2: Running an example security test with GLM 5.3 and Strix
Stop the Juice Shop container with Ctrl+C in the terminal where it’s running, or run docker ps to find the container ID and stop it with docker stop
Give GLM 5.3 a try on the Amazon Bedrock console, use it through coding assistants like OpenCode as shown in our recent post with Kimi K3, or connect your custom applications through the supported APIs.
Interested in how Amazon Bedrock can support your team? Connect with us to start the conversation.
Related Stories
AI News
Building AI start
37 minutes ago
AI News
S.Korea eyes idle highway land, rest stops for AI data centers
38 minutes ago
AI News
AI Agents Expand Cyber Liability, Putting Corporate Safeguards in Focus
38 minutes ago
AI News
AI expert Gary Marcus warns 'reckless' technology could lead to deaths as he calls for more regulation
1 hour ago
AI News
York County libraries push literacy as AI grows, citing 1 in 4 residents struggle
2 hours ago
AI News
Which is the real threat: climate or AI? Both
2 hours ago
AI News
Artificial Intelligence featured prominently in Supreme Court vacancies interviews
3 hours ago
AI News
Inside McDonald's push to have artificial intelligence price your Big Mac
3 hours ago