Skip to content
  • Models
  • Rankings
  • Ori
Sign Up
Sign Up
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Pricing
  • Business
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
Favicon for z-ai

Z.ai: GLM 4.6V

z-ai/glm-4.6v

Model weights
Compare

GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media. It supports up to 128K tokens, processes complex page layouts and charts directly as visual inputs, and integrates native multimodal function calling to connect perception with downstream tool execution. The model also enables interleaved image-text generation and UI reconstruction workflows, including screenshot-to-HTML synthesis and iterative visual editing.

Modalities

In / Out Price

$0.30 / $0.90per 1M

Context

131K

Released

Dec 8, 2025

Compare
ProvidersPricingPerformanceUptimeBenchmarksAppsActivityFAQExplore

Providers

Different companies host the same model. OpenRouter routes your request to one of them based on the routing mode you pick — Balanced (price + speed), Nitro (fastest), or Exacto (highest tool-calling accuracy).

Pricing

The average price customers actually pay for this model, next to the prices providers post. Caching and discounts mean the price actually paid is often well below the listed one.

Performance

Throughput is how fast the model writes (tokens per second — higher is better). Latency is total round-trip time (lower is better). TTFT is time-to-first-token — how long before you see anything appear (lower is better).

Uptime

Uptime is the percentage of the past 3 days that at least one provider was responding to requests. Availability is the percentage of time that inference was successfully served. OpenRouter continuously monitors and uses the next-best provider when one returns an error.

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows where this model lands among all models on OpenRouter.

Benchmark score summary for Z.ai: GLM 4.6V (Artificial Analysis)
SourceBenchmarkScore
Artificial AnalysisGLM-4.6V (Reasoning) GPQA Diamond71.9%
Artificial AnalysisGLM-4.6V (Reasoning) HLE9.6%
Artificial AnalysisGLM-4.6V (Reasoning) IFBench30.1%
Artificial AnalysisGLM-4.6V (Reasoning) τ²-Bench Telecom31.6%
Artificial AnalysisGLM-4.6V (Reasoning) AA-LCR48.7%
Artificial AnalysisGLM-4.6V (Reasoning) CritPt0.0%
Artificial AnalysisGLM-4.6V (Reasoning) Terminal-Bench Hard14.4%
Artificial AnalysisGLM-4.6V (Reasoning) AA-Omniscience Accuracy16.2%
Artificial AnalysisGLM-4.6V (Reasoning) AA-Omniscience Non-Hallucination Rate48.5%
Artificial AnalysisGLM-4.6V (Non-reasoning) GPQA Diamond56.6%
Artificial AnalysisGLM-4.6V (Non-reasoning) HLE3.7%
Artificial AnalysisGLM-4.6V (Non-reasoning) IFBench27.9%
Artificial AnalysisGLM-4.6V (Non-reasoning) τ²-Bench Telecom30.7%
Artificial AnalysisGLM-4.6V (Non-reasoning) AA-LCR17.0%
Artificial AnalysisGLM-4.6V (Non-reasoning) CritPt0.0%
Artificial AnalysisGLM-4.6V (Non-reasoning) Terminal-Bench Hard3.0%
Artificial AnalysisGLM-4.6V (Non-reasoning) AA-Omniscience Accuracy17.4%
Artificial AnalysisGLM-4.6V (Non-reasoning) AA-Omniscience Non-Hallucination Rate33.4%

Apps

Public apps that send the most traffic to this model. Good signal for what real production workloads look like — and a hint at which use cases this model is best suited for.

Activity

Token volume and request traffic to this model over time.

Quick Start

Drop-in code to call this model. OpenRouter's API is OpenAI-compatible — most SDKs work by just swapping the base URL. The only thing that changes between models is the model slug below.

Explore more models

AI Models with Vision: Multimodal LLMs for Image UnderstandingCollectionAI Model RankingsRanking
$0.30$0.90$0.0553.60s19 tps
98.78%
$0.30$0.90$0.054.49s23 tps
99.71%

Throughput

23tok/s

P50, best across providers

Latency

3.60s

P50, best provider

AutoExacto Benchmarks
vgi_benchZ.ai2.4%auto-routing1.4%
Uptime (3d)The model was reachable. Request routed to a provider.

100.00%

Availability (3d)The model returned inference from any provider. Errors and empty responses count against it.

99.95%

Availability over the last 3 days

Last 72 hours
Availability 99.95%
3 Days Ago2 Days AgoYesterdayNow

Availability over the last 24 hours

OpenRouter Availability
99.94%
Without Routing
79.83%

When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.

1.
Favicon for https://portkey.ai/
Portkey AI
Control panel for AI apps
495Mtokens
2.
Favicon for https://www.tryreverie.com/
Reverie catalog metadata
new
291Mtokens
3.
Favicon for https://i.love.koalas.ai/
KoalaBear
new
199Mtokens
4.
Favicon for https://vsearch-benchmark/
VSearch Recovery Eval
new
187Mtokens
5.
Favicon for https://nousresearch.com
Hermes Agent
Hermes Agent is an open-source, self-improving AI agent by Nous Research that runs persistently with memory across sessions, and builds reusable skills from experience. It comes with 40+ built-in tools, including web search, browser automation, and vision, plus scheduled automations and subagents.
108Mtokens

Frequently asked questions

GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media. It supports up to 128K tokens, processes complex page layouts and charts directly as visual inputs, and integrates native multimodal function calling to connect perception with downstream tool execution.

GLM 4.6V costs $0.30/M input tokens and $0.90/M output tokens, with separate rates for Cache Read at $0.055/M tokens.

GLM 4.6V has a 131,072 token context window. It supports up to 32,768 completion tokens.

Yes. GLM 4.6V accepts tools and tool_choice for function calling. It supports response_format for JSON output, without JSON-schema enforcement.

GLM 4.6V accepts images, text and video as input and returns text.

GLM 4.6V is served by 2 providers on OpenRouter: NovitaAI and Z.ai. Requests are routed to the best available provider, with automatic failover to the others, and you can pin or exclude providers with provider routing.

GLM 4.6V was released on December 8, 2025.

More models from Z.ai

GLM Flash Latest

This model always redirects to the latest model in the GLM Flash family.

Text1.0M context
GLM 5.3 Flash

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.

Text1.3M context$0.075 / $0.25
GLM 5.3 Flash

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.

Text1.0M context$0.075 / $0.25
GLM Latest

This model always redirects to the latest GLM model from Z.ai.

Text1.0M context
GLM 5.3

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves on GLM-5.2 in coding and in the balance between performance and token efficiency.

Reasoning is always on and cannot be disabled. Reasoning efforts low, high, and max are supported; max is the default.

Text1.3M context$0.8775 / $2.97
GLM 5.3

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves on GLM-5.2 in coding and in the balance between performance and token efficiency.

Reasoning is always on and cannot be disabled. Reasoning efforts low, high, and max are supported; max is the default.

Text1.0M context$0.70 / $2.20
GLM 5.2

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering, and complex multi-step automation.

Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is particularly strong at coding and tool use across long-running tasks, able to maintain engineering context and follow standards consistently through a full development workflow, from requirements to multi-platform deployment, in a single task.

Text1.0M context$0.4872 / $1.531
GLM 5.2

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering, and complex multi-step automation.

Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is particularly strong at coding and tool use across long-running tasks, able to maintain engineering context and follow standards consistently through a full development workflow, from requirements to multi-platform deployment, in a single task.

Text33K contextFree
GLM 5.2

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering, and complex multi-step automation.

Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is particularly strong at coding and tool use across long-running tasks, able to maintain engineering context and follow standards consistently through a full development workflow, from requirements to multi-platform deployment, in a single task.

Text1.0M context$0.70 / $2.20
GLM 5.1

GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on a single task for more than 8 hours, autonomously planning, executing, and improving itself throughout the process, ultimately delivering complete, engineering-grade results.

Text205K context$0.9646 / $3.032
GLM 5V Turbo

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding, and task execution, and works seamlessly with agents to complete the full loop of “perceive → plan → execute“.

Text203K context$1.20 / $4
GLM 5 Turbo

GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios. It is deeply optimized for real-world agent workflows involving long execution chains, with improved complex instruction decomposition, tool use, scheduled and persistent execution, and overall stability across extended tasks.

Text203K context$1.20 / $4
GLM 5

GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. Built for expert developers, it delivers production-grade performance on large-scale programming tasks, rivaling leading closed-source models. With advanced agentic planning, deep backend reasoning, and iterative self-correction, GLM-5 moves beyond code generation to full-system construction and autonomous execution.

Text205K context$0.60 / $1.92
GLM 4.7 Flash

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning, and tool collaboration, and has achieved leading performance among open-source models of the same size on several current public benchmark leaderboards.

Text200K context$0.06 / $0.40
GLM 4.7

GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while delivering more natural conversational experiences and superior front-end aesthetics.

Text205K context$0.40 / $1.75
GLM 4.6

Compared with GLM-4.5, this generation brings several key improvements:

Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex agentic tasks. Superior coding performance: The model achieves higher scores on code benchmarks and demonstrates better real-world performance in applications such as Claude Code、Cline、Roo Code and Kilo Code, including improvements in generating visually polished front-end pages. Advanced reasoning: GLM-4.6 shows a clear improvement in reasoning performance and supports tool use during inference, leading to stronger overall capability. More capable agents: GLM-4.6 exhibits stronger performance in tool using and search-based agents, and integrates more effectively within agent frameworks. Refined writing: Better aligns with human preferences in style and readability, and performs more naturally in role-playing scenarios.

Text205K context$0.43 / $1.75
GLM 4.5V

GLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (MoE) architecture with 106B parameters and 12B activated parameters, it achieves state-of-the-art results in video understanding, image Q&A, OCR, and document parsing, with strong gains in front-end web coding, grounding, and spatial reasoning. It offers a hybrid inference mode: a "thinking mode" for deep reasoning and a "non-thinking mode" for fast responses. Reasoning behavior can be toggled via the reasoning enabled boolean. Learn more in our docs

Text66K context$0.60 / $1.80
GLM 4.5

GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly enhanced capabilities in reasoning, code generation, and agent alignment. It supports a hybrid inference mode with two options, a "thinking mode" designed for complex reasoning and tool use, and a "non-thinking mode" optimized for instant responses. Users can control the reasoning behaviour with the reasoning enabled boolean. Learn more in our docs

Text131K context$0.60 / $2.20
GLM 4.5 Air

GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (MoE) architecture but with a more compact parameter size. GLM-4.5-Air also supports hybrid inference modes, offering a "thinking mode" for advanced reasoning and tool use, and a "non-thinking mode" for real-time interaction. Users can control the reasoning behaviour with the reasoning enabled boolean. Learn more in our docs

Text131K context$0.13 / $0.85
GLM 4 32B

GLM 4 32B is a cost-effective foundation language model.

It can efficiently perform complex tasks and has significantly enhanced capabilities in tool use, online search, and code-related intelligent tasks.

It is made by the same lab behind the thudm models.

Text128K context