{"data":{"schema_version":"model_decision_evidence.v1","snapshot_date":"2026-09-25","model":{"model":"glm-5.3","direct_answer":"GLM-5.3 is Z.ai's flagship open-weight mixture-of-experts model (744B total, 40B active as stated) with a 1M-token context window, 128K output tokens and text-only input; thinking is always on. Its weights come under a custom GLM-5.3 License rather than MIT, and on two independent coding leaderboards it scored 49.92 (18th of 47) on AI Coding Daily's and 52/60 on kzhu's version-2.1 test.","verified_at":"2026-09-24","model_card":{"context_window":"1M tokens (1,048,576 in the model config)","max_output":"128K tokens on the Z.ai API","modalities":"Text input; text output","parameter_count":"744B total, 40B active (MoE), as stated; the published weights hold about 753B parameters including the MTP layer"},"official_sources":[{"title":"GLM-5.3: Frontier Coding with Emergent Cyber Capabilities","publisher":"Z.ai","url":"https://z.ai/blog/glm-5.3","retrieved_at":"2026-09-24"},{"title":"GLM-5.3 model documentation","publisher":"Z.ai","url":"https://docs.z.ai/guides/llm/glm-5.3","retrieved_at":"2026-09-24"},{"title":"GLM-5.3 model card","publisher":"Z.ai","url":"https://huggingface.co/zai-org/GLM-5.3","retrieved_at":"2026-09-24"}],"open_weights":{"status":"available","summary":"Z.ai published FP8 and BF16 safetensors (about 756 GB and 1.5 TB) under its own GLM-5.3 License, not the MIT licence of GLM-5.3-Flash. The licence requires a company whose group revenue exceeds US$10 billion over 12 months and that runs a model-as-a-service business to pass Z.AI's security review before commercial use; merely relaying requests to models hosted by others does not count. Z.ai does not say which precision its API serves.","repository":"https://huggingface.co/zai-org/GLM-5.3","licence":"GLM-5.3 License (custom; not MIT)","hardware":"Not published by Z.ai; the FP8 weights alone are about 756 GB, and the Agent recipe performs a hardware preflight.","recommended_runtime":"SGLang or vLLM (the model card also lists TokenSpeed, Transformers, KTransformers and Unsloth)","source_urls":["https://huggingface.co/zai-org/GLM-5.3","https://huggingface.co/zai-org/GLM-5.3/blob/main/LICENSE"],"verified_at":"2026-09-24","agent_recipe":"You are my deployment agent. Deploy the official zai-org/GLM-5.3 weights as a private, OpenAI-compatible service on this host or cluster. Use the publisher's model card as the source of truth: https://huggingface.co/zai-org/GLM-5.3.\n\nLicence\n- The weights are under Z.ai's custom GLM-5.3 License, not MIT. Before downloading, show me its model-as-a-service clause and ask me to confirm that our intended use is permitted.\n\nPreflight\n- Inspect the OS, GPU model and count on every node, the interconnect between nodes, free VRAM, driver and CUDA versions, free disk space, and available ports. The FP8 checkpoint alone is about 756 GB (the BF16 repository, zai-org/GLM-5.3-BF16, is about 1.5 TB), and Z.ai publishes no minimum hardware or tensor-parallel settings; estimate whether the official weights and the requested context fit with safe headroom.\n- If the hardware cannot safely run the official weights, stop before installing or launching anything. Explain the limiting resource and propose a separately reviewed alternative; do not substitute a community quantization automatically.\n\nDeployment\n- If preflight passes and I have confirmed the licence, create an isolated environment and follow the model card's SGLang or vLLM instructions and the versions they name. If a runtime asks for trust-remote-code, list the files it would execute and ask me first. Generate an API key and store it in a protected environment variable. Bind to 127.0.0.1 first; do not expose the endpoint publicly.\n- The model always reasons; keep the publisher's reasoning defaults. Start with a context length the preflight supports; the model supports up to 1,048,576 tokens.\n\nVerification and handoff\n- Send a real request to /v1/chat/completions and confirm that the model responds. Report the exact launch command, installed versions, model revision, local endpoint, test result, measured GPU memory and log location. Identify any reverse-proxy or firewall changes that still need my approval. Never print the API key or other secrets."},"curated_comparisons":["glm-5.3-flash","qwen3.8-max-0902"],"open_weight_alternatives":[],"reviews":[{"id":"aicodingdaily-llm-coding-leaderboard-2026-09","platform":"Independent benchmark","evidence_type":"independent_suite","title":"AI Coding Daily LLM Coding Leaderboard","url":"https://aicodingdaily.com/leaderboard","published_at":"2026-09-24","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"AI Coding Daily","handle":"@AICodingDaily","profile_url":"https://www.youtube.com/@AICodingDaily","description":"Coding-model testing channel and website that runs its own leaderboard and sells premium tutorials and channel advertising"},"task":"Implement and harden production-shaped projects in several stacks, including a Laravel app, a Go service and a React/TypeScript app, from the same prompts","method":"Each model and reasoning setting is a separate row run in a coding harness, with the average cost and time per run. Code quality is graded by a judge model against a published 100-point rubric and scaled to 20 points per project; behavioural reliability runs each of four projects five times and scores the checks every attempt fails.","finding":"GLM-5.3 at high reasoning scored 49.92 for $0.19 and 4 min 58 s per run in OpenCode, 18th of 47 rows, just below GPT-6 Luna at xhigh (49.94) and well above GLM-5.3-Flash (42.82).","limitation":"One author's benchmark whose code-quality half is graded by a judge model; rows run in different harnesses, the page is updated as models are added (read on 2026-09-24, 47 rows), and the author sells premium tutorials and advertising.","metrics":[{"label":"Points (high)","value":"49.92"},{"label":"Rank","value":"18 of 47"},{"label":"Cost per run","value":"$0.19"}]},{"id":"techgogogo-kzhu-llm-benchmark-2026-09","platform":"Independent benchmark","evidence_type":"independent_suite","title":"kzhu's six-question LLM test leaderboard","url":"https://www.techgogogo.com/llm-benchmark/","published_at":"2026-09-24","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"AI产品狙击手 kzhu","handle":"@kevinzhu9305","profile_url":"https://www.youtube.com/@kevinzhu9305","description":"Chinese YouTube reviewer who scores models on a six-question test and publishes every score with per-question notes"},"task":"Six questions scored out of 10: constrained sentence writing, two reasoning puzzles (replaced by pelican and compound-bow SVG animations in version 2.2), a browser OS with a space shooter, a 3D racing game and a Trello-style board","method":"Each row names the harness the model ran in (Pi, Claude Code, WorkBuddy and others) and the date; the author scores every question with a note, and warns that totals are only comparable within one test version.","finding":"GLM-5.3 scored 52/60 on version 2.1 in Claude Code at high reasoning: both reasoning puzzles right, nine of ten constrained sentences, a good browser OS and space shooter, a handsome racing game with reversed steering and a Trello board with weak drag animation; at the time it tied for first place.","limitation":"One author's scores on six questions; rows use different harnesses and test versions (2.1 or 2.2), and a few rows' table total differs by one point from the author's note.","metrics":[{"label":"Score","value":"52 / 60 (v2.1, Claude Code)"}]}]},"ranking":{"name":"glm-5.3","display_name":"GLM-5.3","vendor":"Zhipu AI (Z.ai)","origin":"cn","category":"text","release_date":"2026-08-14","on_chinaapi":true,"sources":[{"leaderboard":"lmarena_text","metric":"Arena score (Elo)","value":"1483","entry":"glm-5.3-max","url":"https://arena.ai/leaderboard/text","retrieved_at":"2026-09-22","interval":"±6"},{"leaderboard":"livebench","metric":"Overall","value":"76.1","entry":"GLM-5.3","url":"https://livebench.ai/","retrieved_at":"2026-09-12"},{"leaderboard":"aa_index","metric":"Intelligence Index","value":"45","entry":"GLM-5.3 (max)","url":"https://artificialanalysis.ai/leaderboards/models","retrieved_at":"2026-09-21"},{"leaderboard":"superclue","metric":"总分 (overall)","value":"71.29","entry":"GLM-5.3(max)","url":"https://www.superclueai.com/homepage","retrieved_at":"2026-09-22"},{"leaderboard":"vals_index","metric":"Accuracy (GDP-weighted finance, coding and legal)","value":"56.97","entry":"GLM 5.3","url":"https://www.vals.ai/benchmarks/vals_index","retrieved_at":"2026-09-22","interval":"±1.35"},{"leaderboard":"epoch_eci","metric":"General ECI","value":"156","entry":"GLM-5.3","url":"https://epoch.ai/eci","retrieved_at":"2026-09-22","interval":"(154 - 158)"},{"leaderboard":"lmarena_text_chinese","metric":"Arena score (Elo)","value":"1525","entry":"glm-5.3-max","url":"https://arena.ai/leaderboard/text/chinese","retrieved_at":"2026-09-22","interval":"±22"},{"leaderboard":"lmarena_text_coding","metric":"Arena score (Elo)","value":"1524","entry":"glm-5.3-max","url":"https://arena.ai/leaderboard/text/coding","retrieved_at":"2026-09-22","interval":"±12"},{"leaderboard":"livebench_coding","metric":"Coding","value":"79.0","entry":"GLM-5.3","url":"https://livebench.ai/","retrieved_at":"2026-09-12"},{"leaderboard":"livebench_agentic_coding","metric":"Agentic Coding","value":"60.9","entry":"GLM-5.3","url":"https://livebench.ai/","retrieved_at":"2026-09-12"},{"leaderboard":"aa_speed","metric":"Median output tokens/s","value":"53","entry":"GLM-5.3 (max)","url":"https://artificialanalysis.ai/leaderboards/models","retrieved_at":"2026-09-22"},{"leaderboard":"aa_gdpval","metric":"Agentic Real-World Work Tasks, (Elo-500)/2000 (%)","value":"57","entry":"GLM-5.3 (max)","url":"https://artificialanalysis.ai/leaderboards/models","retrieved_at":"2026-09-21"},{"leaderboard":"aa_tbench_v40","metric":"Agentic Coding \u0026 Terminal Use (%)","value":"42","entry":"GLM-5.3 (max)","url":"https://artificialanalysis.ai/leaderboards/models","retrieved_at":"2026-09-21"},{"leaderboard":"aa_mlcr","metric":"Medical Long Context Reasoning (%)","value":"48.3","entry":"GLM-5.3 (max)","url":"https://artificialanalysis.ai/evaluations/mlcr-aa","retrieved_at":"2026-09-07"}],"notes":"Released 2026-08-14; served here via Zhipu's official channels since 2026-08-18. Independent boards picked it up between the 2026-08-19 and 2026-08-26 snapshots, so the vendor's own DeepSWE figure it carried before has been replaced by their readings; the LMArena row stands on 10,960 votes as of 2026-09-22. Epoch, which carried no row for it as of 2026-09-05, lists it as GLM-5.3 (dated 2026-08-14) in the 2026-09-22 read, and SuperCLUE lists GLM-5.3(max) with an evaluation published 2026-09-10. Vals lists it as GLM 5.3, released 2026-08-18 per its model page, the date of Z.ai's release notes."},"sources":{"aa_gdpval":{"id":"aa_gdpval","name":"Artificial Analysis · GDPval-AA v2.1","operator":"Artificial Analysis","metric":"Agentic Real-World Work Tasks, (Elo-500)/2000 (%)","url":"https://artificialanalysis.ai/leaderboards/models","edition":"GDPval-AA v2.1, live board, read 2026-09-22. 220 agentic task-completion tasks with file outputs, ranked pairwise by a panel of three frontier LLM judges; since v2.1 the Elo scale is anchored to DeepSeek V4.1 Flash (max) at 1600 and fitted with a Crowd-BT model, so its figures are not comparable with the v2 figures this page quoted from its 2026-09-12 read. The column prints clamp((Elo-500)/2000). It carries 10% of the Agents category inside Intelligence Index v4.3.2, so this column is a component of the board that orders the table, not an independent operator. Rows Artificial Analysis marks deprecated are read with the leaderboard's Status filter set to All (CA-948). The column is hidden until the leaderboard's Intelligence group is expanded","retrieved_at":"2026-09-22","scene":"agent_work"},"aa_image_edit":{"id":"aa_image_edit","name":"Artificial Analysis Editing","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/image/leaderboard/editing","edition":"live board; exact row settings retained","retrieved_at":"2026-09-13","task":"image_editing"},"aa_image_t2i":{"id":"aa_image_t2i","name":"Artificial Analysis Text To Image","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/image/leaderboard/text-to-image","edition":"live board; exact row settings retained","retrieved_at":"2026-09-13","task":"text_to_image"},"aa_index":{"id":"aa_index","name":"Artificial Analysis Intelligence Index","operator":"Artificial Analysis","metric":"Intelligence Index","url":"https://artificialanalysis.ai/leaderboards/models","edition":"Intelligence Index v4.3.2, read 2026-09-22; its figures are not comparable with the v4.3 figures this page quoted from its 2026-09-12 read. v4.3.1 refreshed the pairwise judge panels of AA-Briefcase and GDPval-AA, and v4.3.2 moved GDPval-AA to v2.1 (Elo scale anchored to DeepSeek V4.1 Flash (max) at 1600, fitted with a Crowd-BT model) and refitted AA-Briefcase v1.1 with Crowd-BT, so rows moved without any model changing. v4.3 had already swapped τ³-Banking out of the index for AutomationBench-AA and Terminal-Bench v2.1 for v4.0. The leaderboard's default Status: Current filter hides rows Artificial Analysis marks deprecated; they still render with Status set to All and are re-read from that view like any other row (CA-948). A trailing * is reproduced as the board prints it: the board's data marks that score as estimated (intelligenceIndexIsEstimated), i.e. not every evaluation in the index had been run for that model","retrieved_at":"2026-09-22"},"aa_mlcr":{"id":"aa_mlcr","name":"Artificial Analysis · MLCR-AA","operator":"Artificial Analysis","metric":"Medical Long Context Reasoning (%)","url":"https://artificialanalysis.ai/evaluations/mlcr-aa","edition":"MLCR-AA, live board, read 2026-09-22. Artificial Analysis marks it a standalone evaluation that is NOT part of Intelligence Index v4.3 — unlike the GDPval-AA v2 and Terminal-Bench v4.0 columns, this one is not a component of the board that orders the table, only the same operator. The underlying benchmark is Wisedocs' open MLCR (Wisedocs-AI/medical-long-context-reasoning): synthetic medical records of roughly 25,000-64,000 tokens, graded across six tiers from locating a single fact to expert-level clinical synthesis. It measures multi-document reasoning over long records, not medical capability in general. The main leaderboard carries no column for it; the figures were first read off the evaluation page's chart (2026-09-07) and since 2026-09-22 from the per-model pages' payload, which carries every model the chart can show (CA-956)","retrieved_at":"2026-09-22","scene":"medical"},"aa_speed":{"id":"aa_speed","name":"Artificial Analysis · Output Speed","operator":"Artificial Analysis","metric":"Median output tokens/s","url":"https://artificialanalysis.ai/leaderboards/models","edition":"live board; measured by Artificial Analysis against the vendor's own API, not through ChinaAPI. Rows Artificial Analysis marks deprecated are read with the leaderboard's Status filter set to All (CA-948)","retrieved_at":"2026-09-22","scene":"speed"},"aa_tbench_v40":{"id":"aa_tbench_v40","name":"Artificial Analysis · Terminal-Bench v4.0","operator":"Artificial Analysis","metric":"Agentic Coding \u0026 Terminal Use (%)","url":"https://artificialanalysis.ai/leaderboards/models","edition":"Terminal-Bench 4.0 (66 tasks), run by Artificial Analysis with the mini-swe-agent harness, pass@1 averaged over 3 repeats per task; read 2026-09-22. When this column replaced the benchmark's own board on 2026-09-07 the methodology page named the harness as mini-SWE-agent v2.4.6; on 2026-09-21 it and the evaluation page name mini-swe-agent without a version, and every figure quoted before then read back unchanged. One harness for every model, so the column compares models rather than model-plus-best-scaffold — which is why it replaced the benchmark's own board, whose rows are agent × model and covered one Chinese model out of 34 (CA-788, overturning CA-734). It carries 10% of the Coding category inside Intelligence Index v4.3.2, so this column is a component of the board that orders the table, not an independent operator. Rows Artificial Analysis marks deprecated are read with the leaderboard's Status filter set to All (CA-948). The column is hidden until the leaderboard's Intelligence group is expanded","retrieved_at":"2026-09-22","scene":"agentic"},"aa_video_i2v_audio":{"id":"aa_video_i2v_audio","name":"Artificial Analysis Image To Video · With Audio","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/video/leaderboard/image-to-video","edition":"live board; with audio pool; exact row settings retained","retrieved_at":"2026-09-13","task":"image_to_video_audio"},"aa_video_t2v_audio":{"id":"aa_video_t2v_audio","name":"Artificial Analysis Text To Video · With Audio","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/video/leaderboard/text-to-video","edition":"live board; with audio pool; exact row settings retained","retrieved_at":"2026-09-13","task":"text_to_video_audio"},"epoch_eci":{"id":"epoch_eci","name":"Epoch Capabilities Index (ECI)","operator":"Epoch AI","metric":"General ECI","url":"https://epoch.ai/eci","edition":"live board","licence":"CC BY 4.0, as labelled on the page (https://creativecommons.org/licenses/by/4.0/)","retrieved_at":"2026-09-22"},"livebench":{"id":"livebench","name":"LiveBench","operator":"LiveBench (sponsored by Abacus.AI)","metric":"Overall","url":"https://livebench.ai/","edition":"LiveBench-2026-06-25","retrieved_at":"2026-09-22"},"livebench_agentic_coding":{"id":"livebench_agentic_coding","name":"LiveBench · Agentic Coding","operator":"LiveBench (sponsored by Abacus.AI)","metric":"Agentic Coding","url":"https://livebench.ai/","edition":"LiveBench-2026-06-25","retrieved_at":"2026-09-22","scene":"agentic"},"livebench_coding":{"id":"livebench_coding","name":"LiveBench · Coding","operator":"LiveBench (sponsored by Abacus.AI)","metric":"Coding","url":"https://livebench.ai/","edition":"LiveBench-2026-06-25","retrieved_at":"2026-09-22","scene":"coding"},"lmarena_i2v":{"id":"lmarena_i2v","name":"LMArena Image-to-Video Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/image-to-video","edition":"live board, last updated 2026-09-02","retrieved_at":"2026-09-12","task":"image_to_video","updated_at":"2026-09-02"},"lmarena_image_edit":{"id":"lmarena_image_edit","name":"LMArena Image Edit Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/image-edit","edition":"live board, last updated 2026-09-07","retrieved_at":"2026-09-12","task":"image_editing","updated_at":"2026-09-07"},"lmarena_t2i":{"id":"lmarena_t2i","name":"LMArena Text-to-Image Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text-to-image","edition":"live board, last updated 2026-09-07","retrieved_at":"2026-09-12","task":"text_to_image","updated_at":"2026-09-07"},"lmarena_t2v":{"id":"lmarena_t2v","name":"LMArena Text-to-Video Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text-to-video","edition":"live board, last updated 2026-09-04","retrieved_at":"2026-09-12","task":"text_to_video","updated_at":"2026-09-04"},"lmarena_text":{"id":"lmarena_text","name":"LMArena Text Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text","edition":"live board","retrieved_at":"2026-09-22"},"lmarena_text_chinese":{"id":"lmarena_text_chinese","name":"LMArena Text Arena · Chinese","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text/chinese","edition":"live board, Chinese-language category (prompts in Chinese)","retrieved_at":"2026-09-22","scene":"chinese"},"lmarena_text_coding":{"id":"lmarena_text_coding","name":"LMArena Text Arena · Coding","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text/coding","edition":"live board, Coding category","retrieved_at":"2026-09-22","scene":"coding"},"superclue":{"id":"superclue","name":"SuperCLUE 智能指数","operator":"SuperCLUE","metric":"总分 (overall)","url":"https://www.superclueai.com/homepage","edition":"2026-07 evaluation, published 2026-08-06. The homepage reads this edition from its data file (/data/generalboard/2026年7月.xlsx, sheet 总排行榜, last modified 2026-09-11), which also carries rows evaluated after publication, each with its own evaluation publication date (the latest 2026-09-11); the page's Latest Update line refers to a different sub-board","retrieved_at":"2026-09-22"},"vals_index":{"id":"vals_index","name":"Vals Index","operator":"Vals AI","metric":"Accuracy (GDP-weighted finance, coding and legal)","url":"https://www.vals.ai/benchmarks/vals_index","edition":"V2 (released 2026-08-13), page states updated 9/21/2026","retrieved_at":"2026-09-22"},"vendor_card":{"id":"vendor_card","name":"Vendor's own evaluation card","operator":"the model's own vendor","metric":"varies; the benchmark is named in each row","url":"","edition":"self-reported","self_reported":true,"retrieved_at":"2026-09-22"}}},"success":true}