gpt-6-astra real-task evaluations
Every report keeps its original task, method, finding, limitation, evaluator identity and verification date. A community report is evidence about that run, not a universal score.
Cortex AI: GPT-6 Astra, Gemini 3.8 Flash and Fable 5.1 on two real robot tasks
By Cortex AI Research (@cortexairobot). Robotics team that builds real-world robot datasets and runs the MolmoAct2 real-world evaluations Published 2026-09-17; retrieved 2026-09-24; verified 2026-09-24.
- Task
- Load a sandwich into a microwave and close the door; attach a 5 mm tip to a pipette
- Method
- Each model controlled a bimanual YAM robot through an overhead and two wrist cameras with depth images, at medium reasoning effort with up to 10 minutes per task; an evaluator scored each attempt with partial credit for completed steps.
- Finding
- GPT-6 Astra scored highest on both tasks, 80% on the microwave and 70% on the pipette, against 13.3% and 30% for Gemini 3.8 Flash and 6.7% and 20% for Fable 5.1; the lab credits it with retrying until the pipette tip docked and with checking that the microwave door latched.
- Limitation
- The post does not say how many attempts each model had; scores are partial-credit judgements by the lab's evaluator, and the authors attribute the gap mainly to weaker depth perception in the lower-scoring models.
GPT-6 Astra, Sol and Luna on one SVG animation prompt
By 雪瑜 (@xueyu1125). Programmer and AI-tools builder on X who posted the timings and quota readings Published 2026-09-23; retrieved 2026-09-24; verified 2026-09-24.
- Task
- Generate an SVG animation of a pelican riding a bicycle, shown in H5
- Method
- Same prompt on GPT-6 Astra, GPT-6 Sol, GPT-6 Luna and GPT-5.6 Sol; the author recorded time taken and the share of a five-hour subscription quota each used.
- Finding
- GPT-6 Astra took 6 min 3 s and used 16% of the five-hour quota, against 4 min 42 s and 2% for GPT-6 Sol; the author recommends Sol when quota matters.
- Limitation
- One prompt; the quota shares are a subscription meter, not API cost, and the post scores no output quality for Astra.
Switching a coding project from GPT-6 Astra to DeepSeek V4.1 Flash
By OwlZealousideal4779 (u/OwlZealousideal4779). r/DeepSeek member reporting their own coding-agent use Published 2026-09-19; retrieved 2026-09-24; verified 2026-09-24.
- Task
- The author's own coding project, first on GPT-6 Astra and then on DeepSeek V4.1 Flash
- Method
- Self-reported spend and outcome of about a dozen GPT-6 Astra turns, followed by the switch to DeepSeek V4.1 Flash.
- Finding
- About a dozen GPT-6 Astra turns cost the author $130 in a day without meeting the requirements; they switched the project back to DeepSeek V4.1 Flash.
- Limitation
- One uncontrolled self-report; the post does not describe what the Astra turns were asked to do.
Four models build the Eiffel Tower in Three.js
By Bhavy (@Bhavani_00007). X user who ran the same Three.js task on four models and reported cost and time Published 2026-09-23; retrieved 2026-09-24; verified 2026-09-24.
- Task
- Build the Eiffel Tower in Three.js
- Method
- Same task on Claude Opus 5.5, GPT-6 Astra, GPT-6 Sol and Kimi K3; the author reported cost and time for each and compared the results by eye.
- Finding
- GPT-6 Astra cost $7.45 and took 8 minutes; the author preferred Claude Opus 5.5's result ($8.95, 10 minutes) for this task.
- Limitation
- One task judged by eye by the author; the post does not say how the models were run or how the costs were counted.