Z AI: The Complete Guide to Zhipu AI’s GLM-5.2 Chatbot (2026)

Z AI — officially written as Z.ai — is a Chinese AI chatbot and development platform powered by the GLM-5.2 language model, built by Zhipu AI, a spinout from Tsinghua University. It supports conversations up to 1 million tokens in context, offers an MIT-licensed open-source model, and starts at $18/month for developer access. The platform is one of the most cost-effective alternatives to ChatGPT and Claude in 2026, particularly for coding-focused workflows.

Since rebranding from Zhipu AI to Z.ai in July 2025, the company has listed on the Hong Kong Stock Exchange (raising $558 million), released GLM-5.2 in June 2026, and built a 1-gigawatt data center using entirely Chinese-made chips — all without access to Nvidia hardware.

Z.ai company milestones timeline: Tsinghua spinout 2019, rebrand July 2025, HK IPO $558M 2026
From Tsinghua spinout to $558M HK IPO: Z.ai’s journey from a 2019 university lab to a publicly listed AI platform

What Is Z AI? History and Background

Origins: From Zhipu AI to Z.ai

Z.ai was founded in 2019 as a research spinout from Tsinghua University’s Knowledge Engineering Group. Its co-founders include professors Tang Jie and Li Juanzi, alongside Zhang Peng, who serves as CEO. The company’s original name — Zhipu AI (智谱AI) — still appears extensively in technical documentation, API endpoints, and academic papers. The global rebrand to Z.ai happened in July 2025, when the company adopted the z.ai domain.

Before the IPO, Z.ai raised $400 million at a $3 billion valuation in May 2024, drawing investment from Alibaba, Tencent, Meituan, Ant Group, Xiaomi, and HongShan. In January 2026, it listed on the Hong Kong Stock Exchange under the ticker 2513, raising $558 million — the largest Chinese AI IPO at the time. By June 2026, its market capitalization had grown to approximately $128 billion.

Where Z.ai Stands in the AI Landscape

Z.ai is the third-largest LLM provider in China by market share, according to IDC’s 2024 report, behind Baidu and SenseTime. The platform competes globally with ChatGPT, Claude, and DeepSeek. Its workforce exceeds 800 employees, with offices in Beijing (headquarters), the UK, Singapore, Malaysia, and the Middle East.

In January 2025, the US government added Z.ai to the Entity List, blocking access to Nvidia chips including the H20 export-compliant GPU. Rather than slowing down, the company responded by scaling Chinese-made accelerators and building its own data center infrastructure — a strategic pivot that now defines its competitive differentiation.

Our goal is to build the most capable AI platform independent of any single chip supply chain. The entity list forced us to innovate faster than we otherwise would have.

Zhang Peng, CEO of Z.ai

GLM-5.2: Z AI’s Flagship Model

Architecture: 744B MoE with 1-Million-Token Context

GLM-5.2, released June 13–16, 2026, is built on a 744-billion-parameter Mixture-of-Experts (MoE) architecture, with approximately 40 billion parameters active per inference token. This efficient design delivers frontier-level performance while keeping inference costs low — the core reason Z.ai can price API access at $1.40 per million input tokens.

The model supports a 1,000,000-token context window with up to 131,072 output tokens per response — roughly five times the context of GLM-5.1 (which topped out at 200,000 tokens). GLM-5.2 ships with two selectable reasoning modes: High (optimized for speed) and Max (maximum accuracy). Training was powered by Slime, a reinforcement learning infrastructure built in-house at Z.ai.

GLM-5.2 MoE architecture: 744B total parameters, 40B active per token, 1M token context window
GLM-5.2’s Mixture-of-Experts design activates only ~40B of its 744B parameters per inference token — keeping costs low while supporting a 1M-token context window

Open Source Under MIT License

Since July 2025, Z.ai has released GLM model weights under the MIT License — one of the most permissive open-source licenses available. Developers can download, fine-tune, self-host, and commercially deploy GLM models without royalty payments or usage restrictions. Model weights are available on Hugging Face, including quantized 2-bit versions that weigh approximately 239 GB and can run on high-end consumer hardware. This makes the Z.ai platform one of the few frontier-class models that is simultaneously open-weight and commercially usable without a license fee.

Z AI Benchmarks: How GLM-5.2 Performs vs ChatGPT and Claude

Coding Benchmarks

On SWE-bench Pro — a rigorous test of real-world GitHub issue resolution — GLM-5.2 scores 62.1%, outperforming GPT-5.5 (58.6%) but trailing Claude Opus 4.8 (69.2%). On SWE-bench Verified, the predecessor GLM-5.1 achieved 77.8%; Z.ai did not publish an official SWE-bench Verified score for GLM-5.2 at launch, with Claude Opus 4.8 holding 88.6% on that benchmark. On FrontierSWE, GLM-5.2 reaches 74.4%, placing it among the top 5 open-weight models globally.

ModelSWE-bench ProSWE-bench VerifiedCost per 1M Input Tokens
Claude Opus 4.869.2%88.6%$5
GLM-5.262.1%N/A (not published)$1.40
GPT-5.558.6%~76%$5

The price-to-performance ratio is where Z.ai’s case becomes compelling. The GLM-5.2 platform delivers competitive coding benchmark performance (62.1% on SWE-bench Pro vs Claude Opus 4.8’s 69.2%) at approximately one-quarter of the API input cost. For teams running thousands of coding tasks per month, that difference is significant.

SWE-bench Pro 2026 bar chart: Claude Opus 4.8 at 69.2%, GLM-5.2 at 62.1%, GPT-5.5 at 58.6%
SWE-bench Pro 2026: GLM-5.2 outscores GPT-5.5 on real-world coding tasks at roughly one-quarter the API cost of the leading models

Limitations and Controversy

In independent benchmarks covering 47 models, GLM-5.2 ranked 31st overall with a composite score of 7.3/10. While it excels at coding tasks, general reasoning and open-domain knowledge benchmarks place it in the mid-tier range. The model’s strongest case is coding-specific; users expecting GPT-5.5-level broad capability will find GLM-5.2 more specialized.

A separate controversy emerged in 2026 when an MIT research study found approximately 50% of GLM instances self-identified as “Claude” when asked about their model identity — dubbed the “Pony Alpha” controversy. Z.ai has not publicly addressed this finding, but it raised questions about training data provenance and model fine-tuning practices.

Z AI Pricing: Free Plan, Coding Plans, and API

Free Tier

Z.ai is available for free at chat.z.ai, powered by GLM-5 (the base model, not GLM-5.2). The free tier covers standard conversation, document analysis, and basic coding assistance. The platform attracts approximately 8.5 million monthly visits — small compared to ChatGPT’s 5.5 billion, but growing rapidly following the GLM-5.2 launch in June 2026.

GLM Coding Plan Tiers

For developers who need higher rate limits, API access, or GLM-5.2 specifically:

PlanPriceBest For
Coding Lite$18/monthHobbyists, occasional devs
Coding Pro$72/monthFull-time developers
Coding Max$160/monthTeams and enterprise use

A 30% annual discount is available across all tiers. The platform integrates natively with Claude Code, Cline, Roo Code, and OpenClaw, making it a drop-in replacement for existing AI coding toolchains.

Z.ai pricing tiers comparison: Free GLM-5 chat, Coding Lite $18/month, Coding Max $160/month
Z.ai’s pricing ladder covers every use case: free chat access, $18/month for hobbyist developers, and $160/month for team-scale AI coding workloads

API Access

Direct Z.ai API pricing: $1.40 per 1M input tokens and $4.40 per 1M output tokens for GLM-5.2. That is roughly one-third the input token cost of GPT-5.5 ($5/$30 per 1M) and roughly one-seventh the output token cost — and roughly one-quarter the input token cost of Claude Opus 4.8 ($5/$25 per 1M). Via OpenRouter, the model is accessible without a Z.ai account, useful for developers who want to test before committing. The faster GLM-4.7-FlashX variant costs as little as $0.07 per 1M input tokens for high-throughput workloads.

Z AI for Coding: The GLM Coding Plan in Practice

How the Coding Plan Works

The GLM Coding Plan provides API access to GLM-5.2 with high-rate-limit quotas tuned for code generation workflows. It supports 20+ coding tool integrations and generates code at over 55 tokens per second. The GLM-5.2[1m] variant (the full 1-million-token context version) lets you load entire mid-sized codebases into a single prompt — a practical advantage for large refactors, cross-file debugging, or codebase-wide search.

AutoGLM, Z.ai’s autonomous agent product, extends this further. It can execute approximately 1,700 autonomous steps per session and sustain operation for up to 8 hours without human intervention, handling browser automation, file system operations, API calls, and end-to-end coding workflows.

Setting Up GLM-5.2 with Claude Code or Cline

To use GLM-5.2 as a drop-in backend for Claude Code:

  1. Sign up for a GLM Coding Plan at z.ai
  2. Copy your API key from the Z.ai dashboard
  3. In Claude Code settings, add Z.ai as a custom API provider with base URL https://api.z.ai/api/paas/v4/
  4. Set the model to glm-5.2 or glm-5.2-1m depending on context needs
  5. Run a test prompt to confirm the integration is working

The same process applies to Cline and Roo Code. GLM-5.2 via the Coding Pro plan delivers strong coding performance (62.1% SWE-bench Pro vs Claude Opus 4.8’s 69.2%) at roughly 1/4 the API input token cost. For teams already using one of these tools, switching the backend model is a matter of minutes.

Why the Integration Ecosystem Matters

Z.ai’s decision to support Claude Code, Cline, and Roo Code natively is significant for adoption. It means the Z.ai platform does not require developers to change their IDE, toolchain, or workflow — only the API endpoint and model name. ZCode IDE, Z.ai’s own development environment, offers additional native integrations for users who want a tighter loop between the assistant and the development environment.

Z.ai developer tool integrations: Claude Code and Cline, Roo Code and OpenClaw, MCP one API key
Z.ai works as a drop-in backend for the most popular AI coding tools — swap the endpoint, keep your workflow

Z AI Infrastructure: 1-Gigawatt Data Center, No Nvidia

Built Without Nvidia

When the US government placed Z.ai on the Entity List in January 2025, it cut off access to Nvidia chips — including the H20, the export-compliant GPU that other Chinese AI labs had continued to rely on. Z.ai responded by scaling its deployment of Huawei Ascend accelerators, alongside Cambricon Technologies and Moore Threads chips.

In July 2026, Z.ai completed construction of a 1-gigawatt data center, enough power to simultaneously energize approximately 750,000 homes. Each compute cluster in the facility holds more than 10,000 chips. This represents the largest AI compute installation built exclusively on Chinese-made accelerators.

Chinese accelerators consume more power per unit of compute than comparable Nvidia hardware, but the scale of Z.ai’s facility allows it to train and run GLM-class models entirely domestically. The facility positions Z.ai as the most infrastructure-independent major AI lab outside the US.

Z.ai chip independence: US Entity List 2025 blocking Nvidia, Chinese chips Ascend and Cambricon, 1GW data center 2026
Blocked from Nvidia hardware in January 2025, Z.ai pivoted to Chinese-made accelerators and built a 1GW data center independently — the largest AI compute installation without US chips

What This Means for Users

The Nvidia independence means Z.ai’s model availability is less vulnerable to US export policy shifts than competitors still dependent on H100/H200 chips. If additional restrictions tighten, Z.ai’s domestic infrastructure insulates its operations. The tradeoff: inference on Chinese accelerators runs 30–40% slower than equivalent Nvidia-based services, something reflected in generation speed benchmarks. For most chat and coding use cases, the speed is sufficient — latency becomes noticeable only in extreme high-throughput scenarios.

Is Z AI Safe to Use? Open Source, Privacy, and Concerns

Open-source transparency is one of Z.ai’s strongest user trust arguments. Because GLM models are released under the MIT License with public weights, security researchers and independent auditors can inspect the model’s training artifacts and fine-tune behavior directly. Self-hosting the weights is possible on high-end hardware, allowing organizations with strict data residency requirements to run GLM-5.2 entirely on-premise with no data leaving their infrastructure.

Data and privacy are the primary concern for enterprise users. Z.ai is a Chinese company operating under Chinese data laws. Conversations processed through cloud endpoints — chat.z.ai or the Z.ai API — may be subject to Chinese jurisdiction. The platform’s privacy policy is available at z.ai/privacy. For sensitive or proprietary workloads, using the self-hosted open-source weights is the recommended approach.

Company risk is a factor to consider for long-term enterprise decisions. Z.ai’s Hong Kong-listed stock (2513) has been volatile: it fell 23% in February 2026, recovered 11.5% in April 2026 on GLM-5.1 news, and spiked 42% intraday on the GLM-5.2 launch in June 2026. The company’s US blacklist status, while an infrastructure constraint, also creates geopolitical risk for Western enterprise customers evaluating long-term vendor relationships.

  • MIT-licensed weights enable full transparency and self-hosting
  • Cloud API endpoints fall under Chinese data jurisdiction
  • US Entity List status may affect enterprise procurement decisions in regulated industries
  • Stock volatility reflects broader geopolitical uncertainty around Chinese AI companies

FAQ

keyboard_arrow_up