gemini-3.8-flash vs gpt-6-astra

Parameters, independent leaderboard coverage, real-task reports and open-weight availability are shown side by side; missing coverage is not a zero.

Parameters

Comparison itemgemini-3.8-flashgpt-6-astra
Model TypeLLMLLM
Context window1,048,576 tokens1,050,000 tokens
Maximum output65,536 tokens128,000 tokens
Input/output modalitiestext, image, video, audio, pdf → texttext, image → text
Endpointopenai · geminiopenai · openai-response
Open weightsUnavailableUnavailable
LicenseProprietary hosted modelProprietary hosted model

Leaderboard coverage

Every independent board covering either model is retained; an em dash means that board has not published a figure for that model.

Evaluationgemini-3.8-flashgpt-6-astra
LMArena Text Arena — Arena score (Elo)1493 ±91480 ±12
LiveBench — Overall75.882.2
Artificial Analysis Intelligence Index — Intelligence Index4153
Epoch Capabilities Index (ECI) — General ECI157 (155 - 161)167 (163 - 172)
Vals Index — Accuracy (GDP-weighted finance, coding and legal)62.25 ±1.0166.61 ±1.09
LMArena Text Arena · Chinese — Arena score (Elo)1543 ±32
LMArena Text Arena · Coding — Arena score (Elo)1535 ±161543 ±23
LiveBench · Coding — Coding72.580.4
LiveBench · Agentic Coding — Agentic Coding54.257.3
Artificial Analysis · Output Speed — Median output tokens/s29758
Artificial Analysis · GDPval-AA v2.1 — Agentic Real-World Work Tasks, (Elo-500)/2000 (%)4652
Artificial Analysis · Terminal-Bench v4.0 — Agentic Coding & Terminal Use (%)2059
Artificial Analysis · MLCR-AA — Medical Long Context Reasoning (%)21.735.0

Real-task reports

gemini-3.8-flash

  • RemakeBench Launch 005: four Flash models on 14 fixtures — Fourteen fixtures spanning game building, 3D scenes, CAD, drawing and robot control, run on Qwen 3.8 Flash, DeepSeek V4.1 Flash, GLM 5.3 Flash and Gemini 3.8 Flash. Finding: No model won overall. Gemini 3.8 Flash made the fastest single-turn Rubik's Cube pass (15 min 35 s, $2.09) and was the editorial winner of the robot Eiffel Tower drawing (8 min 25 s, $1.36), but did not complete the three-move extension ($5.65) and did not deliver the required Turbofan CAD assembly. Limitation: Pass/fail labels are video judgments without a machine-readable validator, and single attempts do not establish reliability; the lab publishes no aggregate score or overall winner. Gemini's displayed costs are several times the other models' on most fixtures (for example $2.09 against $0.0957–$0.3414 on the single-turn cube), and its CAD presentation clip is operator work after the run.
  • Cortex AI: GPT-6 Astra, Gemini 3.8 Flash and Fable 5.1 on two real robot tasks — Load a sandwich into a microwave and close the door; attach a 5 mm tip to a pipette. Finding: Gemini 3.8 Flash scored 13.3% on the microwave task and 30% on the pipette task, second to GPT-6 Astra (80% and 70%) and ahead of Fable 5.1 (6.7% and 20%). Limitation: The post does not say how many attempts each model had; scores are partial-credit judgements by the lab's evaluator, and the authors attribute the gap mainly to weaker depth perception in the lower-scoring models.

gpt-6-astra

  • Cortex AI: GPT-6 Astra, Gemini 3.8 Flash and Fable 5.1 on two real robot tasks — Load a sandwich into a microwave and close the door; attach a 5 mm tip to a pipette. Finding: GPT-6 Astra scored highest on both tasks, 80% on the microwave and 70% on the pipette, against 13.3% and 30% for Gemini 3.8 Flash and 6.7% and 20% for Fable 5.1; the lab credits it with retrying until the pipette tip docked and with checking that the microwave door latched. Limitation: The post does not say how many attempts each model had; scores are partial-credit judgements by the lab's evaluator, and the authors attribute the gap mainly to weaker depth perception in the lower-scoring models.
  • GPT-6 Astra, Sol and Luna on one SVG animation prompt — Generate an SVG animation of a pelican riding a bicycle, shown in H5. Finding: GPT-6 Astra took 6 min 3 s and used 16% of the five-hour quota, against 4 min 42 s and 2% for GPT-6 Sol; the author recommends Sol when quota matters. Limitation: One prompt; the quota shares are a subscription meter, not API cost, and the post scores no output quality for Astra.
  • Switching a coding project from GPT-6 Astra to DeepSeek V4.1 Flash — The author's own coding project, first on GPT-6 Astra and then on DeepSeek V4.1 Flash. Finding: About a dozen GPT-6 Astra turns cost the author $130 in a day without meeting the requirements; they switched the project back to DeepSeek V4.1 Flash. Limitation: One uncontrolled self-report; the post does not describe what the Astra turns were asked to do.
  • Four models build the Eiffel Tower in Three.js — Build the Eiffel Tower in Three.js. Finding: GPT-6 Astra cost $7.45 and took 8 minutes; the author preferred Claude Opus 5.5's result ($8.95, 10 minutes) for this task. Limitation: One task judged by eye by the author; the post does not say how the models were run or how the costs were counted.