{"data":{"schema_version":"model_decision_evidence.v1","snapshot_date":"2026-09-25","model":{"model":"glm-5.3-flash","direct_answer":"GLM-5.3-Flash is Z.ai's MIT-licensed open-weight mixture-of-experts model (320B total, 18B active) with a 1M-token context window and video, image, text and file input. Thinking cannot be turned off; reviewed tests show strong single results but uneven completion across tasks.","verified_at":"2026-09-24","model_card":{"context_window":"1M tokens (1,048,576)","max_output":"128K tokens (131,072 on the Z.ai API)","modalities":"Video, image, text and file input; text output","parameter_count":"320B total, 18B active (MoE)"},"official_sources":[{"title":"GLM-5.3-Flash: Frontier Intelligence, Flash Cost","publisher":"Z.ai","url":"https://z.ai/blog/glm-5.3-flash","retrieved_at":"2026-09-24"},{"title":"GLM-5.3-Flash model documentation","publisher":"Z.ai","url":"https://docs.z.ai/guides/vlm/glm-5.3-flash","retrieved_at":"2026-09-24"},{"title":"GLM-5.3-Flash model card","publisher":"Z.ai","url":"https://huggingface.co/zai-org/GLM-5.3-Flash","retrieved_at":"2026-09-24"}],"open_weights":{"status":"available","summary":"Official FP8 and BF16 safetensors are published under the MIT License. The publisher lists SGLang, vLLM, TokenSpeed, Transformers, KTransformers and Unsloth for local serving.","repository":"https://huggingface.co/zai-org/GLM-5.3-Flash","licence":"MIT","hardware":"Not published by Z.ai; the Agent recipe performs a hardware preflight before serving.","recommended_runtime":"SGLang or vLLM","source_urls":["https://huggingface.co/zai-org/GLM-5.3-Flash","https://huggingface.co/zai-org/GLM-5.3-Flash-BF16"],"verified_at":"2026-09-24","agent_recipe":"You are my deployment agent. Deploy the official zai-org/GLM-5.3-Flash weights as a private, OpenAI-compatible service on this host. Use the publisher's model card as the source of truth: https://huggingface.co/zai-org/GLM-5.3-Flash.\n\nPreflight\n- Inspect the OS, GPU model and count, free VRAM, driver and CUDA versions, free disk space, and available ports. The default repository holds block-wise FP8 weights; zai-org/GLM-5.3-Flash-BF16 is the official BF16 alternative. Estimate whether the chosen weights and the requested context fit with safe headroom for runtime and concurrency; Z.ai publishes no minimum GPU requirement.\n- If the host cannot safely run the official weights, stop before installing or launching anything. Explain the limiting resource and propose a separately reviewed alternative; do not substitute a community quantization automatically.\n\nDeployment\n- If preflight passes, create an isolated environment, install a current SGLang or vLLM release that documents GLM-5.3-Flash support, record the exact model revision, and serve the FP8 repository through its OpenAI-compatible API. Generate an API key and store it in a protected environment variable. Bind to 127.0.0.1 first; do not expose the endpoint publicly.\n- Keep the publisher's sampling defaults (temperature 1.0, top_p 0.95). Thinking is always on for this model, so do not try to disable it. Start with a context length the preflight supports; the model supports up to 1,048,576 tokens.\n\nVerification and handoff\n- Send a real request to /v1/chat/completions and confirm that the model responds. Report the exact launch command, installed versions, model revision, local endpoint, test result, measured GPU memory and log location. Identify any reverse-proxy or firewall changes that still need my approval. Never print the API key or other secrets."},"curated_comparisons":["deepseek-flash","mimo-v2.6-flash"],"open_weight_alternatives":[],"reviews":[{"id":"remakebench-launch-005-2026-09","platform":"Independent benchmark","evidence_type":"independent_suite","title":"RemakeBench Launch 005: four Flash models on 14 fixtures","url":"https://www.remakebench.com/qwen-3-8-flash-vs-deepseek-v4-1-flash-vs-glm-5-3-flash-vs-gemini-3-8-flash","published_at":"2026-09-18","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"RemakeBench","handle":"@remakebench","profile_url":"https://www.youtube.com/@remakebench","description":"Benchmark channel that publishes per-fixture clips, costs, wall-clock times and validation notes"},"task":"Fourteen fixtures spanning game building, 3D scenes, CAD, drawing and robot control, run on Qwen 3.8 Flash, DeepSeek V4.1 Flash, GLM 5.3 Flash and Gemini 3.8 Flash","method":"Each model's attempt per fixture is shown as a video clip with its displayed cost, wall-clock time and reasoning setting; pass/fail labels are the lab's video judgments.","finding":"No model won overall. GLM 5.3 Flash was the only model to pass the Apple-stem manipulation task and passed the single-turn Rubik's Cube, but did not complete the three-move extension and delivered no native geometry in Turbofan CAD.","limitation":"Pass/fail labels are video judgments without a machine-readable validator, and single attempts do not establish reliability; the lab publishes no aggregate score or overall winner. GLM's Infinite Cathedral completion needed repeated user continuations because the endpoint stalled.","metrics":[{"label":"Apple-stem manipulation","value":"Pass · $0.5583 · 1h 29m 33s"},{"label":"Rubik's Cube single turn","value":"Pass · $0.3414 · 40m 48s"}]},{"id":"x-sentdex-halite2-2026-09","platform":"X","evidence_type":"task_test","title":"GLM 5.3 Flash alone and with Jev on Halite 2","url":"https://x.com/Sentdex/status/2101828851458293827","published_at":"2026-09-21","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"Harrison Kinsley","handle":"@Sentdex","profile_url":"https://x.com/Sentdex","description":"Developer who ran this Halite 2 experiment and answered questions about it in the thread"},"task":"Play the turn-based strategy game Halite 2","method":"Compared GLM 5.3 Flash alone with a hybrid in which GLM makes the strategic decisions and Jev executes ship-by-ship moves, across seeded maps; in the thread the author cites 133 unique maps.","finding":"GLM 5.3 Flash alone beat Jev alone 82% of the time; the hybrid was 13× faster at 56% of the pure GLM API cost with a slight performance edge.","limitation":"The author calls it a single example and among their first experiments with Jev; results may not carry over to other tasks.","metrics":[{"label":"GLM 5.3 Flash vs Jev win rate","value":"82%"},{"label":"Hybrid speed-up","value":"13×"},{"label":"Hybrid API cost","value":"56% of pure GLM"}]},{"id":"x-miaai-game-menus-2026-09","platform":"X","evidence_type":"task_test","title":"DeepSeek V4.1 Flash and GLM 5.3 Flash recreating two game menus","url":"https://x.com/MiaAI_lab/status/2100134727457902942","published_at":"2026-09-16","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"Mia","handle":"@MiaAI_lab","profile_url":"https://x.com/MiaAI_lab","description":"X account that posts LLM experiments and attached the full output videos"},"task":"Recreate the Psychonauts 2 and Metroid Prime main menus as closely as possible to the originals","method":"One shot per game, both models at Max thinking in the same harness; the author attached full-length videos and reported tokens used.","finding":"GLM 5.3 Flash used 134k and 111k tokens, a third to a quarter of DeepSeek V4.1 Flash's, and the author judged its menus clearly better on both games.","limitation":"Which menu is closer to the original is the author's own judgement on two prompts, without a scoring rubric.","metrics":[{"label":"Tokens · Psychonauts 2","value":"134k"},{"label":"Tokens · Metroid Prime","value":"111k"}]},{"id":"x-wesche-blender-dodecahedron-2026-09","platform":"X","evidence_type":"task_test","title":"DeepSeek V4.1 Flash and GLM-5.3-Flash on a Blender geometry task, run locally","url":"https://x.com/WescheNex1q/status/2100759267259117656","published_at":"2026-09-18","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"Wësche","handle":"@WescheNex1q","profile_url":"https://x.com/WescheNex1q","description":"Artist and AI enthusiast on X who builds and benchmarks open models on DGX Spark clusters"},"task":"Write a Blender script that builds a dodecahedron trapped inside a dodecahedron","method":"One shot with thinking on, temperature 0.6, no token cap and no retries; the script ran headless in Blender 5.2 and the author checked the mesh independently. DeepSeek ran as FP8 on vLLM across four DGX Sparks, GLM as an EXL3 quantisation across two.","finding":"GLM-5.3-Flash produced a correct mesh (zero intersections, containment verified in-script) and the nicer picture, but took 93 min 12 s and 132,250 tokens at 23.7 tok/s, against 20 min 15 s and 62,928 tokens for DeepSeek V4.1 Flash.","limitation":"One local run per model on different hardware and precision: GLM ran as a third-party EXL3 quantisation on two DGX Sparks and DeepSeek as FP8 on four, so speed and token counts describe those builds, not the official weights or a hosted API.","metrics":[{"label":"GLM-5.3-Flash","value":"93 min 12 s · 132,250 tokens"},{"label":"DeepSeek V4.1 Flash","value":"20 min 15 s · 62,928 tokens"}]},{"id":"x-wesche-koi-pond-2026-09","platform":"X","evidence_type":"task_test","title":"Three Flash models build the same koi pond locally","url":"https://x.com/WescheNex1q/status/2102527793057661386","published_at":"2026-09-22","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"Wësche","handle":"@WescheNex1q","profile_url":"https://x.com/WescheNex1q","description":"Artist and AI enthusiast on X who builds and benchmarks open models on DGX Spark clusters"},"task":"A koi pond build, run as a local head-to-head","method":"The same task on GLM-5.3-Flash, MiMo-V2.6-Flash and DeepSeek V4.1 Flash, each on two DGX Sparks; the author reported time, tokens and fixes needed and left the visual verdict partly to readers.","finding":"All three finished cleanly with zero fixes, but GLM-5.3-Flash was slowest at 59 minutes and 90.9K tokens, against 30 minutes for MiMo-V2.6-Flash and 21 for DeepSeek V4.1 Flash; the author said it was the first of their tests where GLM was not the winner.","limitation":"One local run per model; the post states neither the prompt nor the quantisation or runtime used.","metrics":[{"label":"Time","value":"59 min"},{"label":"Tokens","value":"90.9K"}]},{"id":"aicodingdaily-llm-coding-leaderboard-2026-09","platform":"Independent benchmark","evidence_type":"independent_suite","title":"AI Coding Daily LLM Coding Leaderboard","url":"https://aicodingdaily.com/leaderboard","published_at":"2026-09-24","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"AI Coding Daily","handle":"@AICodingDaily","profile_url":"https://www.youtube.com/@AICodingDaily","description":"Coding-model testing channel and website that runs its own leaderboard and sells premium tutorials and channel advertising"},"task":"Implement and harden production-shaped projects in several stacks, including a Laravel app, a Go service and a React/TypeScript app, from the same prompts","method":"Each model and reasoning setting is a separate row run in a coding harness, with the average cost and time per run. Code quality is graded by a judge model against a published 100-point rubric and scaled to 20 points per project; behavioural reliability runs each of four projects five times and scores the checks every attempt fails.","finding":"GLM-5.3-Flash at max reasoning scored 42.82 ($0.02 and 8 min 54 s per run in OpenCode), 38th of 47 rows and well below GLM-5.3 at 49.92.","limitation":"One author's benchmark whose code-quality half is graded by a judge model; rows run in different harnesses, the page is updated as models are added (read on 2026-09-24, 47 rows), and the author sells premium tutorials and advertising.","metrics":[{"label":"Points (max)","value":"42.82"},{"label":"Rank","value":"38 of 47"}]},{"id":"techgogogo-kzhu-llm-benchmark-2026-09","platform":"Independent benchmark","evidence_type":"independent_suite","title":"kzhu's six-question LLM test leaderboard","url":"https://www.techgogogo.com/llm-benchmark/","published_at":"2026-09-24","retrieved_at":"2026-09-24","verified_at":"2026-09-24","author":{"name":"AI产品狙击手 kzhu","handle":"@kevinzhu9305","profile_url":"https://www.youtube.com/@kevinzhu9305","description":"Chinese YouTube reviewer who scores models on a six-question test and publishes every score with per-question notes"},"task":"Six questions scored out of 10: constrained sentence writing, two reasoning puzzles (replaced by pelican and compound-bow SVG animations in version 2.2), a browser OS with a space shooter, a 3D racing game and a Trello-style board","method":"Each row names the harness the model ran in (Pi, Claude Code, WorkBuddy and others) and the date; the author scores every question with a note, and warns that totals are only comparable within one test version.","finding":"GLM-5.3-Flash scored 54/60 on version 2.1 in zcode, with every reasoning answer right and the best recent space shooter the author had seen, but 45 after re-running the two version-2.2 SVG questions in Pi, where the pelican sank below the ring and the bow pointed backwards, taking over 20 minutes against under 10 for DeepSeek.","limitation":"One author's scores on six questions; rows use different harnesses and test versions (2.1 or 2.2), and a few rows' table total differs by one point from the author's note.","metrics":[{"label":"Score","value":"54 / 60 (v2.1, zcode)"}]}]},"ranking":{"name":"glm-5.3-flash","display_name":"GLM-5.3-Flash","vendor":"Zhipu AI (Z.ai)","origin":"cn","category":"text","release_date":"2026-08-26","on_chinaapi":true,"sources":[{"leaderboard":"lmarena_text","metric":"Arena score (Elo)","value":"1475","entry":"glm-5.3-flash","url":"https://arena.ai/leaderboard/text","retrieved_at":"2026-09-22","interval":"±7"},{"leaderboard":"livebench","metric":"Overall","value":"71.6","entry":"GLM-5.3 Flash","url":"https://livebench.ai/","retrieved_at":"2026-09-12"},{"leaderboard":"aa_index","metric":"Intelligence Index","value":"42","entry":"GLM-5.3-Flash","url":"https://artificialanalysis.ai/leaderboards/models","retrieved_at":"2026-09-21"},{"leaderboard":"superclue","metric":"总分 (overall)","value":"68.10","entry":"GLM-5.3-Flash(max)","url":"https://www.superclueai.com/homepage","retrieved_at":"2026-09-22"},{"leaderboard":"vals_index","metric":"Accuracy (GDP-weighted finance, coding and legal)","value":"47.22","entry":"GLM 5.3 Flash","url":"https://www.vals.ai/benchmarks/vals_index","retrieved_at":"2026-09-22","interval":"±1.45"},{"leaderboard":"epoch_eci","metric":"General ECI","value":"152","entry":"GLM-5.3-Flash","url":"https://epoch.ai/eci","retrieved_at":"2026-09-22","interval":"(150 - 154)"},{"leaderboard":"lmarena_text_chinese","metric":"Arena score (Elo)","value":"1529","entry":"glm-5.3-flash","url":"https://arena.ai/leaderboard/text/chinese","retrieved_at":"2026-09-22","interval":"±25"},{"leaderboard":"lmarena_text_coding","metric":"Arena score (Elo)","value":"1525","entry":"glm-5.3-flash","url":"https://arena.ai/leaderboard/text/coding","retrieved_at":"2026-09-22","interval":"±12"},{"leaderboard":"livebench_coding","metric":"Coding","value":"79.0","entry":"GLM-5.3 Flash","url":"https://livebench.ai/","retrieved_at":"2026-09-12"},{"leaderboard":"livebench_agentic_coding","metric":"Agentic Coding","value":"56.8","entry":"GLM-5.3 Flash","url":"https://livebench.ai/","retrieved_at":"2026-09-12"},{"leaderboard":"aa_speed","metric":"Median output tokens/s","value":"65","entry":"GLM-5.3-Flash","url":"https://artificialanalysis.ai/leaderboards/models","retrieved_at":"2026-09-22"},{"leaderboard":"aa_gdpval","metric":"Agentic Real-World Work Tasks, (Elo-500)/2000 (%)","value":"57","entry":"GLM-5.3-Flash","url":"https://artificialanalysis.ai/leaderboards/models","retrieved_at":"2026-09-21"},{"leaderboard":"aa_tbench_v40","metric":"Agentic Coding \u0026 Terminal Use (%)","value":"33","entry":"GLM-5.3-Flash","url":"https://artificialanalysis.ai/leaderboards/models","retrieved_at":"2026-09-21"},{"leaderboard":"aa_mlcr","metric":"Medical Long Context Reasoning (%)","value":"51.1","entry":"GLM-5.3-Flash","url":"https://artificialanalysis.ai/evaluations/mlcr-aa","retrieved_at":"2026-09-07"}],"notes":"Released 2026-08-26 per Z.ai's own release notes; the first natively multimodal model in the GLM-5 line and its low-cost tier, served here on Zhipu's official channels. LMArena, LiveBench and Artificial Analysis all picked it up within days of release; the LMArena row is a ranked one standing on 10,038 votes as of 2026-09-22, carrying neither the board's Preliminary badge nor an AutoEval marker. Epoch, SuperCLUE and Vals carried no row for it as of 2026-08-31. In the 2026-09-22 read Epoch lists it as GLM-5.3-Flash dated 2026-08-20, six days before Z.ai's release notes; the benchmark runs behind that figure are on the glm-5.3-flash model itself (Epoch's model versions glm-5.3-flash_max and glm-5.3-flash_high), so only the date differs. SuperCLUE lists GLM-5.3-Flash(max) with an evaluation published 2026-09-10, and Vals lists it as GLM 5.3 Flash, released 2026-08-26 per its model page. GLM-5.3's figures are not reused here."},"sources":{"aa_gdpval":{"id":"aa_gdpval","name":"Artificial Analysis · GDPval-AA v2.1","operator":"Artificial Analysis","metric":"Agentic Real-World Work Tasks, (Elo-500)/2000 (%)","url":"https://artificialanalysis.ai/leaderboards/models","edition":"GDPval-AA v2.1, live board, read 2026-09-22. 220 agentic task-completion tasks with file outputs, ranked pairwise by a panel of three frontier LLM judges; since v2.1 the Elo scale is anchored to DeepSeek V4.1 Flash (max) at 1600 and fitted with a Crowd-BT model, so its figures are not comparable with the v2 figures this page quoted from its 2026-09-12 read. The column prints clamp((Elo-500)/2000). It carries 10% of the Agents category inside Intelligence Index v4.3.2, so this column is a component of the board that orders the table, not an independent operator. Rows Artificial Analysis marks deprecated are read with the leaderboard's Status filter set to All (CA-948). The column is hidden until the leaderboard's Intelligence group is expanded","retrieved_at":"2026-09-22","scene":"agent_work"},"aa_image_edit":{"id":"aa_image_edit","name":"Artificial Analysis Editing","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/image/leaderboard/editing","edition":"live board; exact row settings retained","retrieved_at":"2026-09-13","task":"image_editing"},"aa_image_t2i":{"id":"aa_image_t2i","name":"Artificial Analysis Text To Image","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/image/leaderboard/text-to-image","edition":"live board; exact row settings retained","retrieved_at":"2026-09-13","task":"text_to_image"},"aa_index":{"id":"aa_index","name":"Artificial Analysis Intelligence Index","operator":"Artificial Analysis","metric":"Intelligence Index","url":"https://artificialanalysis.ai/leaderboards/models","edition":"Intelligence Index v4.3.2, read 2026-09-22; its figures are not comparable with the v4.3 figures this page quoted from its 2026-09-12 read. v4.3.1 refreshed the pairwise judge panels of AA-Briefcase and GDPval-AA, and v4.3.2 moved GDPval-AA to v2.1 (Elo scale anchored to DeepSeek V4.1 Flash (max) at 1600, fitted with a Crowd-BT model) and refitted AA-Briefcase v1.1 with Crowd-BT, so rows moved without any model changing. v4.3 had already swapped τ³-Banking out of the index for AutomationBench-AA and Terminal-Bench v2.1 for v4.0. The leaderboard's default Status: Current filter hides rows Artificial Analysis marks deprecated; they still render with Status set to All and are re-read from that view like any other row (CA-948). A trailing * is reproduced as the board prints it: the board's data marks that score as estimated (intelligenceIndexIsEstimated), i.e. not every evaluation in the index had been run for that model","retrieved_at":"2026-09-22"},"aa_mlcr":{"id":"aa_mlcr","name":"Artificial Analysis · MLCR-AA","operator":"Artificial Analysis","metric":"Medical Long Context Reasoning (%)","url":"https://artificialanalysis.ai/evaluations/mlcr-aa","edition":"MLCR-AA, live board, read 2026-09-22. Artificial Analysis marks it a standalone evaluation that is NOT part of Intelligence Index v4.3 — unlike the GDPval-AA v2 and Terminal-Bench v4.0 columns, this one is not a component of the board that orders the table, only the same operator. The underlying benchmark is Wisedocs' open MLCR (Wisedocs-AI/medical-long-context-reasoning): synthetic medical records of roughly 25,000-64,000 tokens, graded across six tiers from locating a single fact to expert-level clinical synthesis. It measures multi-document reasoning over long records, not medical capability in general. The main leaderboard carries no column for it; the figures were first read off the evaluation page's chart (2026-09-07) and since 2026-09-22 from the per-model pages' payload, which carries every model the chart can show (CA-956)","retrieved_at":"2026-09-22","scene":"medical"},"aa_speed":{"id":"aa_speed","name":"Artificial Analysis · Output Speed","operator":"Artificial Analysis","metric":"Median output tokens/s","url":"https://artificialanalysis.ai/leaderboards/models","edition":"live board; measured by Artificial Analysis against the vendor's own API, not through ChinaAPI. Rows Artificial Analysis marks deprecated are read with the leaderboard's Status filter set to All (CA-948)","retrieved_at":"2026-09-22","scene":"speed"},"aa_tbench_v40":{"id":"aa_tbench_v40","name":"Artificial Analysis · Terminal-Bench v4.0","operator":"Artificial Analysis","metric":"Agentic Coding \u0026 Terminal Use (%)","url":"https://artificialanalysis.ai/leaderboards/models","edition":"Terminal-Bench 4.0 (66 tasks), run by Artificial Analysis with the mini-swe-agent harness, pass@1 averaged over 3 repeats per task; read 2026-09-22. When this column replaced the benchmark's own board on 2026-09-07 the methodology page named the harness as mini-SWE-agent v2.4.6; on 2026-09-21 it and the evaluation page name mini-swe-agent without a version, and every figure quoted before then read back unchanged. One harness for every model, so the column compares models rather than model-plus-best-scaffold — which is why it replaced the benchmark's own board, whose rows are agent × model and covered one Chinese model out of 34 (CA-788, overturning CA-734). It carries 10% of the Coding category inside Intelligence Index v4.3.2, so this column is a component of the board that orders the table, not an independent operator. Rows Artificial Analysis marks deprecated are read with the leaderboard's Status filter set to All (CA-948). The column is hidden until the leaderboard's Intelligence group is expanded","retrieved_at":"2026-09-22","scene":"agentic"},"aa_video_i2v_audio":{"id":"aa_video_i2v_audio","name":"Artificial Analysis Image To Video · With Audio","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/video/leaderboard/image-to-video","edition":"live board; with audio pool; exact row settings retained","retrieved_at":"2026-09-13","task":"image_to_video_audio"},"aa_video_t2v_audio":{"id":"aa_video_t2v_audio","name":"Artificial Analysis Text To Video · With Audio","operator":"Artificial Analysis","metric":"Elo","url":"https://artificialanalysis.ai/video/leaderboard/text-to-video","edition":"live board; with audio pool; exact row settings retained","retrieved_at":"2026-09-13","task":"text_to_video_audio"},"epoch_eci":{"id":"epoch_eci","name":"Epoch Capabilities Index (ECI)","operator":"Epoch AI","metric":"General ECI","url":"https://epoch.ai/eci","edition":"live board","licence":"CC BY 4.0, as labelled on the page (https://creativecommons.org/licenses/by/4.0/)","retrieved_at":"2026-09-22"},"livebench":{"id":"livebench","name":"LiveBench","operator":"LiveBench (sponsored by Abacus.AI)","metric":"Overall","url":"https://livebench.ai/","edition":"LiveBench-2026-06-25","retrieved_at":"2026-09-22"},"livebench_agentic_coding":{"id":"livebench_agentic_coding","name":"LiveBench · Agentic Coding","operator":"LiveBench (sponsored by Abacus.AI)","metric":"Agentic Coding","url":"https://livebench.ai/","edition":"LiveBench-2026-06-25","retrieved_at":"2026-09-22","scene":"agentic"},"livebench_coding":{"id":"livebench_coding","name":"LiveBench · Coding","operator":"LiveBench (sponsored by Abacus.AI)","metric":"Coding","url":"https://livebench.ai/","edition":"LiveBench-2026-06-25","retrieved_at":"2026-09-22","scene":"coding"},"lmarena_i2v":{"id":"lmarena_i2v","name":"LMArena Image-to-Video Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/image-to-video","edition":"live board, last updated 2026-09-02","retrieved_at":"2026-09-12","task":"image_to_video","updated_at":"2026-09-02"},"lmarena_image_edit":{"id":"lmarena_image_edit","name":"LMArena Image Edit Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/image-edit","edition":"live board, last updated 2026-09-07","retrieved_at":"2026-09-12","task":"image_editing","updated_at":"2026-09-07"},"lmarena_t2i":{"id":"lmarena_t2i","name":"LMArena Text-to-Image Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text-to-image","edition":"live board, last updated 2026-09-07","retrieved_at":"2026-09-12","task":"text_to_image","updated_at":"2026-09-07"},"lmarena_t2v":{"id":"lmarena_t2v","name":"LMArena Text-to-Video Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text-to-video","edition":"live board, last updated 2026-09-04","retrieved_at":"2026-09-12","task":"text_to_video","updated_at":"2026-09-04"},"lmarena_text":{"id":"lmarena_text","name":"LMArena Text Arena","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text","edition":"live board","retrieved_at":"2026-09-22"},"lmarena_text_chinese":{"id":"lmarena_text_chinese","name":"LMArena Text Arena · Chinese","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text/chinese","edition":"live board, Chinese-language category (prompts in Chinese)","retrieved_at":"2026-09-22","scene":"chinese"},"lmarena_text_coding":{"id":"lmarena_text_coding","name":"LMArena Text Arena · Coding","operator":"LMArena (arena.ai)","metric":"Arena score (Elo)","url":"https://arena.ai/leaderboard/text/coding","edition":"live board, Coding category","retrieved_at":"2026-09-22","scene":"coding"},"superclue":{"id":"superclue","name":"SuperCLUE 智能指数","operator":"SuperCLUE","metric":"总分 (overall)","url":"https://www.superclueai.com/homepage","edition":"2026-07 evaluation, published 2026-08-06. The homepage reads this edition from its data file (/data/generalboard/2026年7月.xlsx, sheet 总排行榜, last modified 2026-09-11), which also carries rows evaluated after publication, each with its own evaluation publication date (the latest 2026-09-11); the page's Latest Update line refers to a different sub-board","retrieved_at":"2026-09-22"},"vals_index":{"id":"vals_index","name":"Vals Index","operator":"Vals AI","metric":"Accuracy (GDP-weighted finance, coding and legal)","url":"https://www.vals.ai/benchmarks/vals_index","edition":"V2 (released 2026-08-13), page states updated 9/21/2026","retrieved_at":"2026-09-22"},"vendor_card":{"id":"vendor_card","name":"Vendor's own evaluation card","operator":"the model's own vendor","metric":"varies; the benchmark is named in each row","url":"","edition":"self-reported","self_reported":true,"retrieved_at":"2026-09-22"}}},"success":true}