Skip to content

GLM Model Release History and GLM-5.3 Forecast Score

Independent research — not an official Z.ai publication.Identity and provider disclosure

Historical GLM model timeline from the 2021 GLM paper through GLM-5.2, ending with the original August 7 to September 7, 2026 forecast window

This original July 2026 diagram preserves the forecast before the outcome. We checked the August 14 announcement and current distribution surfaces on August 15, then scored the forecast without moving its dates.

The GLM model family changed shape several times in five years. It began as a research idea for one model that could handle language understanding and generation. It grew into a 130-billion-parameter bilingual model. ChatGLM brought the family to laptops and chat products. GLM-4 added long context and tool use. The GLM-4.5 and GLM-5 lines turned the focus toward coding agents and long engineering tasks.

That history answers two questions. What has Z.AI built? How did a public cadence forecast perform when the next release arrived?

This guide answers both. It records the main language-model milestones, preserves the prior date calculation, scores it against GLM-5.3, and separates confirmed launch facts from API, weight, and benchmark claims that still need their own receipts. It leaves rumors out.

  1. The GLM release timeline
  2. How the model family changed
  3. The recent release cadence
  4. How the date forecast scored
  5. How the name forecast scored
  6. What the feature forecast got right
  7. What remains unverified
  8. Limits of the forecast
  9. How developers should use the result
  10. Common questions
  11. Sources and method

We use the first official date that marks public access, an official release note, or a research release. Early GLM work used papers and repositories. Later models used API announcements and Z.ai release notes. The “date” column therefore records a public milestone, not a private training start.

Date Release What changed
March 18, 2021 GLM paper The research team introduced autoregressive blank infilling as one framework for language understanding and generation.
June 2021 GLM-10B The team opened a 10B checkpoint and moved GLM from a paper to a model that developers could test.
August 2022 GLM-130B The team opened a 130B Chinese-English model and documented large-scale training and INT4 inference.
March 14, 2023 ChatGLM-6B ChatGLM paired a 6.2B chat model with consumer-GPU deployment and Chinese-English dialogue.
June 25, 2023 ChatGLM2-6B ChatGLM2 raised the open model context from 2K to 32K and improved its benchmark scores.
October 27, 2023 ChatGLM3-6B ChatGLM3 added function calls, a code interpreter, and agent tasks to the open line.
January 16, 2024 GLM-4 Zhipu launched a 128K API model and an All Tools product with browsing, Python, image tools, and user functions.
June 5, 2024 GLM-4-9B The team opened a compact GLM-4 checkpoint, a 128K chat version, and an experimental 1M version.
July 28, 2025 GLM-4.5 Z.ai launched a native agent model with hybrid reasoning, 355B total parameters, 32B active parameters, and MIT weights.
September 30, 2025 GLM-4.6 Z.ai raised context from 128K to 200K and focused the release on coding, reasoning, tools, and agents.
December 22, 2025 GLM-4.7 Z.ai pushed coding, reasoning, tool use, and end-to-end agent work further.
February 12, 2026 GLM-5 Z.ai doubled total scale to 744B, kept about 40B active, and added DeepSeek Sparse Attention for long-context efficiency.
April 7, 2026 GLM-5.1 Z.ai targeted long tasks, strategy revision, bug fixing, and agent runs that the company says can reach eight hours.
June 16, 2026 GLM-5.2 Z.ai raised context to 1M and added IndexShare, improved multi-token prediction, and reasoning-effort controls.
August 14, 2026 GLM-5.3 Z.AI kept the GLM-5.2 base model, changed post-training, required thinking for migration, opened Coding Plan access, and staged API and weight availability.

The list covers the main language-model line. Z.ai runs separate branches for vision, image generation, speech, OCR, and small Flash models. Those branches matter, but mixing their dates into one cadence would distort a forecast for the next flagship text model.

A version table gives facts. Four product eras show direction.

2021–2022: one framework, then large bilingual scale

Section titled “2021–2022: one framework, then large bilingual scale”

The original GLM paper tried to bridge two common model jobs. BERT-style systems handled language understanding. GPT-style systems handled generation. GLM used blank infilling and autoregressive generation to cover both jobs in one framework.

GLM-10B gave the idea public weights. GLM-130B tested the same family at 130B parameters with Chinese and English data. The GLM-130B paper documented training instability, scaling choices, and INT4 deployment on four RTX 3090 GPUs. This era centered on architecture, scale, and access.

2023: ChatGLM made the model useful on consumer hardware

Section titled “2023: ChatGLM made the model useful on consumer hardware”

ChatGLM-6B changed the audience. A 6.2B model with INT4 support let developers run a bilingual chat model on common hardware. ChatGLM2 raised context to 32K. ChatGLM3 added function calls, code execution, and agent work.

The release rhythm tells its own story. The team shipped three open chat generations on March 14, June 25, and October 27. Each release moved beyond conversation. Context, code, tools, and agents became core parts of the family.

The official ChatGLM technical report supplies the dates and explains the feature changes across these generations.

2024: GLM-4 joined long context with tools

Section titled “2024: GLM-4 joined long context with tools”

GLM-4 launched through the API on January 16, 2024. It supported a 128K context. GLM-4 All Tools could choose a browser, Python, an image tool, or a user function within a task.

GLM-4-9B brought the line back to open deployment in June. Its experimental 1M checkpoint matters for the history. Z.ai tested million-token capacity two years before GLM-5.2 made 1M context part of the flagship product. That gap shows a useful rule: an experiment can signal direction without predicting the next shipping date.

2025–2026: coding agents became the main product

Section titled “2025–2026: coding agents became the main product”

GLM-4.5 started the cadence that drives our forecast. Z.ai called it a native agent model. The company joined reasoning, coding, and tool use in one hybrid system and released MIT weights.

GLM-4.6 raised context to 200K. GLM-4.7 emphasized end-to-end coding work. GLM-5 doubled total parameters and used sparse attention to control long-context cost. GLM-5.1 focused on sustained work and strategy revision. GLM-5.2 raised context to 1M and attacked sparse-attention and decoding costs with IndexShare and multi-token prediction.

The direction stays consistent: more work per agent run, stronger code work, better tool control, and lower compute cost for long context. The GLM-5.2 technical guide explains the current architecture and its limits.

The original forecast used six flagship text releases from GLM-4.5 through GLM-5.2. That segment reflected Z.AI’s current product strategy better than a 2021 research interval. We keep the original five input gaps unchanged, then add the GLM-5.3 outcome as a sixth scored interval.

From To Gap
GLM-4.5 GLM-4.6 64 days
GLM-4.6 GLM-4.7 83 days
GLM-4.7 GLM-5 52 days
GLM-5 GLM-5.1 54 days
GLM-5.1 GLM-5.2 70 days
GLM-5.2 GLM-5.3 59 days

Before GLM-5.3, the five forecast inputs produced three useful numbers:

  • Mean: (64 + 83 + 52 + 54 + 70) ÷ 5 = 64.6 days
  • Median: 64 days
  • Observed range: 52–83 days

Five gaps made a small sample. The data supported a planning window, not a precise promise. We used the 64.6-day mean for the center and the 52–83-day range for the outer dates. GLM-5.3’s 59-day interval now changes the six-gap mean to 63.7 days and the median to 61.5 days, while the range stays 52–83 days. Those revised values describe history; we did not use them to retrofit the published forecast.

We exclude GLM-5-Turbo, GLM-5V-Turbo, GLM-4.7-Flash, GLM-OCR, and image or speech models. These releases follow different product tracks. Adding them would shorten the measured interval and create a false sense of speed.

GLM-5.2 launched on June 16, 2026. Our July calculation added the 64.6-day mean and reached August 20, 2026. Adding the shortest and longest prior gaps created this frozen window:

Forecast field Estimate
Earliest cadence match August 7, 2026
Center estimate August 20, 2026
Latest cadence match September 7, 2026
Confidence at publication Moderate

Z.AI announced GLM-5.3 on August 14, 2026. The result landed 59 days after GLM-5.2, inside the 52–83-day historical range, inside the published August 7–September 7 window, and six days before the center estimate. We score the window correct and the single center date close, not exact.

The one-month width did real work. A single-day prediction would have missed; the bounded interval captured the release while exposing its uncertainty. We did not move the window after the announcement or count an unrelated vision, Turbo, speech, image, OCR, or Flash release as success.

The score applies to the announcement event defined in advance: the first official flagship text-model announcement after GLM-5.2. It does not claim that the standard API, every account, every partner, and open weights became available on the same date.

The sequence moved from GLM-5 to GLM-5.1 and GLM-5.2. We therefore gave GLM-5.3 the strongest naming case while keeping GLM-6 as a live alternative. The official name matched the base case.

That match came from a short version sequence, not private knowledge. Z.AI had previously moved from GLM-4.7 to GLM-5 rather than GLM-4.8, so a family jump remained possible. Recording the rejected alternative prevents a correct name from looking inevitable after the fact.

1. Long-horizon agent work: direction confirmed

Section titled “1. Long-horizon agent work: direction confirmed”

Prior confidence: high. Outcome: supported, with vendor evidence.

GLM-4.7 improved end-to-end agent work. GLM-5 targeted complex system engineering. GLM-5.1 targeted long runs and strategy revision. GLM-5.2 targeted context drift and goal loss.

GLM-5.3 continues that line. Z.AI frames the release around long-running coding and security agents and reports gains on Terminal Bench 3.0, DeepSWE 1.1, CyberGym, and ExploitBench. Those company results support the predicted product direction. They do not yet prove fewer abandoned tasks or loops in an independent production workload.

2. Coding and tool use: direction confirmed

Section titled “2. Coding and tool use: direction confirmed”

Prior confidence: high. Outcome: supported, with vendor evidence.

Z.AI has centered every recent flagship release on code, terminals, repositories, search, or tools. GLM-5.3’s launch post reports a 50% gain on an in-house coding benchmark and publishes software-engineering deltas. The release therefore matches the direction forecast.

The score stops at direction. Harnesses, thinking budgets, tool permissions, and task selection influence each result. A developer still needs a fixed repository evaluation before claiming a gain over GLM-5.2.

Prior confidence: high. Outcome: open.

GLM-5 added sparse attention. GLM-5.2 added IndexShare and improved multi-token prediction. These changes target the cost of long inputs and token generation.

The GLM-5.3 announcement says the release uses the same GLM-5.2 base and gets its gains from post-training. That fact does not establish better throughput, lower memory pressure, lower latency, or a lower price. At the August 15 check, the standard pricing page did not name GLM-5.3. We leave this prediction unscored until comparable serving receipts exist.

Prior confidence: medium. Outcome: matched.

GLM-5.2 made 1M context a flagship feature one month ago. A jump to 2M in the next release would add marketing value, but long-context users need recall, goal control, and cost discipline more than another capacity number.

The GLM-5.3 model guide keeps the 1M context label and a 128K maximum output. That matches the capacity forecast. The launch does not by itself prove better recall, lower drift, or stronger attention to old constraints; those useful context qualities need controlled long-input tests.

5. Open weights: promised, not yet delivered at check time

Section titled “5. Open weights: promised, not yet delivered at check time”

Prior confidence: medium to high. Outcome: pending.

Z.AI released GLM-4.5 and the GLM-5 checkpoints under the MIT license. Open weights form part of the company’s developer position.

The GLM-5.3 post promises weights two weeks after launch. At our August 15 check, the official GLM-5 repository still named GLM-5.2 and did not expose the new checkpoint. We will score the prediction only after the files, license, hashes, and model configuration arrive. A promise does not equal a released MIT artifact.

6. A text-first release: confirmed at announcement

Section titled “6. A text-first release: confirmed at announcement”

Prior confidence: medium. Outcome: matched in scope.

Z.AI has kept the main text and vision lines apart: GLM-4.5 and GLM-4.5V, GLM-4.6 and GLM-4.6V, GLM-5 and GLM-5V-Turbo. The GLM-5.3 announcement describes the flagship text and coding line rather than claiming a same-day vision checkpoint.

The result supports the text-first scope. It does not predict when or whether a GLM-5.3 vision sibling will arrive. Buyers should not infer image input from a text-model launch.

The announcement resolves a date and name. It does not resolve every product question readers may have.

  • No new base-parameter count. Z.AI says GLM-5.3 uses the same base model as GLM-5.2 and attributes gains to post-training.
  • No standard API price at the fixed check. Model architecture and Coding Plan access do not set a metered Chat Completion price.
  • No 2M context. The model guide keeps the 1M context label, and usable recall still requires measurement.
  • No independent benchmark win. Vendor results depend on harnesses, prompts, tools, budgets, and model versions.
  • No released-weight receipt yet. A two-week promise does not supply a checkpoint hash, license, tokenizer, configuration, or serving result.

These limits protect the history page from turning an announcement into an API or deployment claim. The dedicated GLM-5.2 to GLM-5.3 migration guide tracks the Coding Plan, standard API, and weight gates separately.

The forecast answered one event question: when the next flagship text model would receive an official announcement. It could not measure rollout depth.

  1. Distribution order. Coding Plan, standard API, partner APIs, and open weights can open on different dates.
  2. Account entitlement. A public model page does not prove that one account can select the model on one endpoint.
  3. Serving economics. Latency, throughput, token use, price, and rate limits need dated receipts.
  4. Task quality. A company benchmark cannot replace repository, tool, long-context, and safety evaluations from the target workload.

The outcome also does not turn six intervals into a release law. Future training, evaluation, infrastructure, safety, and product choices can break the cadence. Any new forecast must define its event and freeze its evidence again; copying the 63.7-day mean would create false precision.

Developers should treat the result as evidence about planning discipline, not as permission to upgrade.

  • Keep GLM-5.2 pilots and baselines. A new announcement does not erase current test results or mature provider support.
  • Keep model IDs and thinking controls in configuration so a controlled canary and rollback need no source rewrite.
  • Save a fixed evaluation set with repository tasks, tool calls, long-context recall, cost, and latency.
  • Require the current API enum, price, account entitlement, and a bounded receipt before sending GLM-5.3 production traffic.
  • Wait for checkpoint files and hashes before sizing a self-hosted deployment.

The local hardware guide can help teams judge whether another large open checkpoint would fit their systems. The free access guide lists ways to test the current model before a purchase.

Z.AI announced GLM-5.3 on August 14, 2026. That fell within our frozen August 7–September 7 window and six days before the August 20 center estimate.

Yes. We selected GLM-5.3 as the base case because GLM-5.1 and GLM-5.2 used point releases, while recording GLM-6 as a live alternative.

The one-month window matched; its center date missed by six days. The forecast used five prior intervals and carried moderate confidence. It predicted an announcement window, not API, account, pricing, or checkpoint availability.

Does the timeline include every GLM product?

Section titled “Does the timeline include every GLM product?”

No. It covers the main language-model milestones. Z.ai has separate vision, image, speech, OCR, code, agent, Air, Flash, and Turbo branches. The official ChatGLM report contains a broader early-family chart, and Z.ai’s release notes list recent specialist products.

No. A new checkpoint must beat GLM-5.2 on the tasks, cost, latency, hosting, and license terms that matter to a team. GLM-5.2 can remain the better production choice if it has mature provider support or a lower cost.

Is GLM-5.3 ready for standard API migration?

Section titled “Is GLM-5.3 ready for standard API migration?”

Not on the announcement alone. At the August 15 check, the model guide said the API was coming soon, and the standard Chat Completion enum and pricing table did not name GLM-5.3. Coding Plan access was live. Check the migration guide for the current gate and required thinking-mode change.

Primary sources:

We searched “GLM model release history,” “GLM release timeline,” and “next GLM model release” on July 16, 2026. Search results offered general histories and model pages, but few pages joined the official pre-2025 timeline, current GLM-5.2 data, a clean cadence calculation, and a falsifiable forecast.

We calculated calendar-day gaps with UTC dates. We used the recent flagship text line for timing and the full family history for feature direction. We excluded specialist branches from the date model. On August 15, we added the August 14 official announcement as the defined outcome and preserved every forecast input, window boundary, confidence label, and rejected name alternative.

The original July research had no query-and-page rows in Google Search Console and only a small GA4 launch sample, including 34 homepage views. Those numbers could not establish demand for this topic and did not enter the date model. The score uses only the frozen public calculation and the official GLM-5.3 launch record.