Skip to content
  • Models
  • Rankings
  • Ori
Sign Up
Sign Up
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Pricing
  • Business
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
Favicon for z-ai

Z.ai: GLM 5V Turbo

z-ai/glm-5v-turbo

Compare

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding, and task execution, and works seamlessly with agents to complete the full loop of “perceive → plan → execute“.

Modalities

In / Out Price

$1.20 / $4per 1M

Context

203K

Released

Apr 1, 2026

Compare
ProvidersPricingPerformanceUptimeBenchmarksAppsActivityFAQExplore

Providers

This model is hosted by one provider. OpenRouter forwards every request to it directly — no routing decisions to make.

Pricing

The average price customers actually pay for this model, next to the prices providers post. Caching and discounts mean the price actually paid is often well below the listed one.

Performance

Throughput is how fast the model writes (tokens per second — higher is better). Latency is total round-trip time (lower is better). TTFT is time-to-first-token — how long before you see anything appear (lower is better).

Uptime

Uptime is the percentage of the past 3 days that at least one provider was responding to requests. Availability is the percentage of time that inference was successfully served. OpenRouter continuously monitors and uses the next-best provider when one returns an error.

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows where this model lands among all models on OpenRouter.

Benchmark score summary for Z.ai: GLM 5V Turbo (Artificial Analysis and Design Arena)
SourceBenchmarkScore
Artificial AnalysisGLM 5V Turbo (Reasoning) GPQA Diamond80.9%
Artificial AnalysisGLM 5V Turbo (Reasoning) HLE17.1%
Artificial AnalysisGLM 5V Turbo (Reasoning) IFBench61.1%
Artificial AnalysisGLM 5V Turbo (Reasoning) τ²-Bench Telecom98.5%
Artificial AnalysisGLM 5V Turbo (Reasoning) AA-LCR70.3%
Artificial AnalysisGLM 5V Turbo (Reasoning) CritPt0.6%
Artificial AnalysisGLM 5V Turbo (Reasoning) Terminal-Bench Hard32.6%
Artificial AnalysisGLM 5V Turbo (Reasoning) AA-Omniscience Accuracy29.3%
Artificial AnalysisGLM 5V Turbo (Reasoning) AA-Omniscience Non-Hallucination Rate31.2%
Design ArenaGLM 5V Turbo Agents Arena Agenticgamedev Elo1102
Design ArenaGLM 5V Turbo Agents Arena Agentichtmlslides Elo1134
Design ArenaGLM 5V Turbo Agents Arena Agenticslides Elo1171
Design ArenaGLM 5V Turbo Agents Arena Agenticslides(Html) Elo1137
Design ArenaGLM 5V Turbo Agents Arena Agenticslides(Python-Pptx) Elo1183
Design ArenaGLM 5V Turbo Agents Arena Androidnative Elo1267
Design ArenaGLM 5V Turbo Agents Arena Full Stack Elo1158
Design ArenaGLM 5V Turbo Agents Arena Godotgamedev Elo997
Design ArenaGLM 5V Turbo Agents Arena Htmlslides Elo1123
Design ArenaGLM 5V Turbo Agents Arena Mobile Apps Elo1162
Design ArenaGLM 5V Turbo Agents Arena Pptxslides Elo1164
Design ArenaGLM 5V Turbo Agents Arena Python-Pptxslides Elo1165
Design ArenaGLM 5V Turbo Agents Arena Python-Pptxslides Elo1151
Design ArenaGLM 5V Turbo Agents Arena Webapps Elo1138
Design ArenaGLM 5V Turbo Models Arena 3D Elo1237
Design ArenaGLM 5V Turbo Models Arena Asciiart Elo1130
Design ArenaGLM 5V Turbo Models Arena Code Categories Elo1244
Design ArenaGLM 5V Turbo Models Arena Data Visualization Elo1216
Design ArenaGLM 5V Turbo Models Arena Game Development Elo1241
Design ArenaGLM 5V Turbo Models Arena SVG Elo1175
Design ArenaGLM 5V Turbo Models Arena UI Component Elo1227
Design ArenaGLM 5V Turbo Models Arena Website Elo1243

Apps

Public apps that send the most traffic to this model. Good signal for what real production workloads look like — and a hint at which use cases this model is best suited for.

Activity

Token volume and request traffic to this model over time.

Quick Start

Drop-in code to call this model. OpenRouter's API is OpenAI-compatible — most SDKs work by just swapping the base URL. The only thing that changes between models is the model slug below.

Explore more models

AI Models with Vision: Multimodal LLMs for Image UnderstandingCollectionAI Model RankingsRanking
$1.20$4.00$0.242.90s21 tps
100.00%

Throughput

21tok/s

P50, best across providers

Latency

2.90s

P50, best provider

AutoExacto Benchmarks
vgi_benchZ.ai13.1%auto-routing11.8%
Uptime (3d)The model was reachable. Request routed to a provider.

100.00%

Availability (3d)The model returned inference from any provider. Errors and empty responses count against it.

99.72%

Availability over the last 3 days

Last 72 hours
Availability 99.72%
3 Days Ago2 Days AgoYesterdayNow

Availability over the last 24 hours

OpenRouter Availability
99.63%

When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.

1.
Favicon for https://www.gohighlevel.com/
HighLevel
new
36.9Btokens
2.
Favicon for https://claude.ai/apple-touch-icon.png
Claude Code
Claude Code is Anthropic's agentic coding tool that reads your entire codebase, plans and executes changes across files, runs tests, and iterates on failures, all from natural language prompts.
10.1Btokens
3.
Favicon for https://www.gohighlevel.com/
schema-markup-generation
new
2.75Btokens
4.
Favicon for https://nousresearch.com
Hermes Agent
Hermes Agent is an open-source, self-improving AI agent by Nous Research that runs persistently with memory across sessions, and builds reusable skills from experience. It comes with 40+ built-in tools, including web search, browser automation, and vision, plus scheduled automations and subagents.
1.67Btokens
5.
Favicon for https://sillytavern.app/
SillyTavern
SillyTavern is the LLM frontend for power users, a chat interface that connects to any model and gives deep control through character creation, roleplay, and prompt customization.
708Mtokens

Frequently asked questions

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding, and task execution, and works seamlessly with agents to complete the full loop of “perceive → plan → execute“.

GLM 5V Turbo costs $1.20/M input tokens and $4.00/M output tokens, with separate rates for Cache Read at $0.24/M tokens.

GLM 5V Turbo has a 202,752 token context window. It supports up to 131,072 completion tokens.

Yes. GLM 5V Turbo accepts tools and tool_choice for function calling. It supports response_format for JSON output, without JSON-schema enforcement.

GLM 5V Turbo accepts images, text and video as input and returns text.

GLM 5.3 Flash, GLM 5.3, GLM 5.2 and 10 more are other text models from Z.ai.

GLM 5V Turbo was released on April 1, 2026.

More models from Z.ai

GLM Flash Latest

This model always redirects to the latest model in the GLM Flash family.

Text1.0M context
GLM 5.3 Flash

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.

Text1.3M context$0.075 / $0.25
GLM 5.3 Flash

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.

Text1.0M context$0.075 / $0.25
GLM Latest

This model always redirects to the latest GLM model from Z.ai.

Text1.0M context
GLM 5.3

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves on GLM-5.2 in coding and in the balance between performance and token efficiency.

Reasoning is always on and cannot be disabled. Reasoning efforts low, high, and max are supported; max is the default.

Text1.3M context$0.8775 / $2.97
GLM 5.3

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves on GLM-5.2 in coding and in the balance between performance and token efficiency.

Reasoning is always on and cannot be disabled. Reasoning efforts low, high, and max are supported; max is the default.

Text1.0M context$0.70 / $2.20
GLM 5.2

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering, and complex multi-step automation.

Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is particularly strong at coding and tool use across long-running tasks, able to maintain engineering context and follow standards consistently through a full development workflow, from requirements to multi-platform deployment, in a single task.

Text1.0M context$0.4872 / $1.531
GLM 5.2

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering, and complex multi-step automation.

Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is particularly strong at coding and tool use across long-running tasks, able to maintain engineering context and follow standards consistently through a full development workflow, from requirements to multi-platform deployment, in a single task.

Text33K contextFree
GLM 5.2

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering, and complex multi-step automation.

Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is particularly strong at coding and tool use across long-running tasks, able to maintain engineering context and follow standards consistently through a full development workflow, from requirements to multi-platform deployment, in a single task.

Text1.0M context$0.70 / $2.20
GLM 5.1

GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on a single task for more than 8 hours, autonomously planning, executing, and improving itself throughout the process, ultimately delivering complete, engineering-grade results.

Text205K context$0.9646 / $3.032
GLM 5 Turbo

GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios. It is deeply optimized for real-world agent workflows involving long execution chains, with improved complex instruction decomposition, tool use, scheduled and persistent execution, and overall stability across extended tasks.

Text203K context$1.20 / $4
GLM 5

GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. Built for expert developers, it delivers production-grade performance on large-scale programming tasks, rivaling leading closed-source models. With advanced agentic planning, deep backend reasoning, and iterative self-correction, GLM-5 moves beyond code generation to full-system construction and autonomous execution.

Text205K context$0.60 / $1.92
GLM 4.7 Flash

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning, and tool collaboration, and has achieved leading performance among open-source models of the same size on several current public benchmark leaderboards.

Text200K context$0.06 / $0.40
GLM 4.7

GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while delivering more natural conversational experiences and superior front-end aesthetics.

Text205K context$0.40 / $1.75
GLM 4.6V

GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media. It supports up to 128K tokens, processes complex page layouts and charts directly as visual inputs, and integrates native multimodal function calling to connect perception with downstream tool execution. The model also enables interleaved image-text generation and UI reconstruction workflows, including screenshot-to-HTML synthesis and iterative visual editing.

Text131K context$0.30 / $0.90
GLM 4.6

Compared with GLM-4.5, this generation brings several key improvements:

Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex agentic tasks. Superior coding performance: The model achieves higher scores on code benchmarks and demonstrates better real-world performance in applications such as Claude Code、Cline、Roo Code and Kilo Code, including improvements in generating visually polished front-end pages. Advanced reasoning: GLM-4.6 shows a clear improvement in reasoning performance and supports tool use during inference, leading to stronger overall capability. More capable agents: GLM-4.6 exhibits stronger performance in tool using and search-based agents, and integrates more effectively within agent frameworks. Refined writing: Better aligns with human preferences in style and readability, and performs more naturally in role-playing scenarios.

Text205K context$0.43 / $1.75
GLM 4.5V

GLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (MoE) architecture with 106B parameters and 12B activated parameters, it achieves state-of-the-art results in video understanding, image Q&A, OCR, and document parsing, with strong gains in front-end web coding, grounding, and spatial reasoning. It offers a hybrid inference mode: a "thinking mode" for deep reasoning and a "non-thinking mode" for fast responses. Reasoning behavior can be toggled via the reasoning enabled boolean. Learn more in our docs

Text66K context$0.60 / $1.80
GLM 4.5

GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly enhanced capabilities in reasoning, code generation, and agent alignment. It supports a hybrid inference mode with two options, a "thinking mode" designed for complex reasoning and tool use, and a "non-thinking mode" optimized for instant responses. Users can control the reasoning behaviour with the reasoning enabled boolean. Learn more in our docs

Text131K context$0.60 / $2.20
GLM 4.5 Air

GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (MoE) architecture but with a more compact parameter size. GLM-4.5-Air also supports hybrid inference modes, offering a "thinking mode" for advanced reasoning and tool use, and a "non-thinking mode" for real-time interaction. Users can control the reasoning behaviour with the reasoning enabled boolean. Learn more in our docs

Text131K context$0.13 / $0.85
GLM 4 32B

GLM 4 32B is a cost-effective foundation language model.

It can efficiently perform complex tasks and has significantly enhanced capabilities in tool use, online search, and code-related intelligent tasks.

It is made by the same lab behind the thudm models.

Text128K context