gpt-6-astra real-task evaluations

Every report keeps its original task, method, finding, limitation, evaluator identity and verification date. A community report is evidence about that run, not a universal score.

  • Cortex AI: GPT-6 Astra, Gemini 3.8 Flash and Fable 5.1 on two real robot tasks

    By Cortex AI Research (@cortexairobot). Robotics team that builds real-world robot datasets and runs the MolmoAct2 real-world evaluations Published 2026-09-17; retrieved 2026-09-24; verified 2026-09-24.

    Task
    Load a sandwich into a microwave and close the door; attach a 5 mm tip to a pipette
    Method
    Each model controlled a bimanual YAM robot through an overhead and two wrist cameras with depth images, at medium reasoning effort with up to 10 minutes per task; an evaluator scored each attempt with partial credit for completed steps.
    Finding
    GPT-6 Astra scored highest on both tasks, 80% on the microwave and 70% on the pipette, against 13.3% and 30% for Gemini 3.8 Flash and 6.7% and 20% for Fable 5.1; the lab credits it with retrying until the pipette tip docked and with checking that the microwave door latched.
    Limitation
    The post does not say how many attempts each model had; scores are partial-credit judgements by the lab's evaluator, and the authors attribute the gap mainly to weaker depth perception in the lower-scoring models.
  • GPT-6 Astra, Sol and Luna on one SVG animation prompt

    By 雪瑜 (@xueyu1125). Programmer and AI-tools builder on X who posted the timings and quota readings Published 2026-09-23; retrieved 2026-09-24; verified 2026-09-24.

    Task
    Generate an SVG animation of a pelican riding a bicycle, shown in H5
    Method
    Same prompt on GPT-6 Astra, GPT-6 Sol, GPT-6 Luna and GPT-5.6 Sol; the author recorded time taken and the share of a five-hour subscription quota each used.
    Finding
    GPT-6 Astra took 6 min 3 s and used 16% of the five-hour quota, against 4 min 42 s and 2% for GPT-6 Sol; the author recommends Sol when quota matters.
    Limitation
    One prompt; the quota shares are a subscription meter, not API cost, and the post scores no output quality for Astra.
  • Switching a coding project from GPT-6 Astra to DeepSeek V4.1 Flash

    By OwlZealousideal4779 (u/OwlZealousideal4779). r/DeepSeek member reporting their own coding-agent use Published 2026-09-19; retrieved 2026-09-24; verified 2026-09-24.

    Task
    The author's own coding project, first on GPT-6 Astra and then on DeepSeek V4.1 Flash
    Method
    Self-reported spend and outcome of about a dozen GPT-6 Astra turns, followed by the switch to DeepSeek V4.1 Flash.
    Finding
    About a dozen GPT-6 Astra turns cost the author $130 in a day without meeting the requirements; they switched the project back to DeepSeek V4.1 Flash.
    Limitation
    One uncontrolled self-report; the post does not describe what the Astra turns were asked to do.
  • Four models build the Eiffel Tower in Three.js

    By Bhavy (@Bhavani_00007). X user who ran the same Three.js task on four models and reported cost and time Published 2026-09-23; retrieved 2026-09-24; verified 2026-09-24.

    Task
    Build the Eiffel Tower in Three.js
    Method
    Same task on Claude Opus 5.5, GPT-6 Astra, GPT-6 Sol and Kimi K3; the author reported cost and time for each and compared the results by eye.
    Finding
    GPT-6 Astra cost $7.45 and took 8 minutes; the author preferred Claude Opus 5.5's result ($8.95, 10 minutes) for this task.
    Limitation
    One task judged by eye by the author; the post does not say how the models were run or how the costs were counted.