gemini-3.8-flash real-task evaluations

Every report keeps its original task, method, finding, limitation, evaluator identity and verification date. A community report is evidence about that run, not a universal score.

  • RemakeBench Launch 005: four Flash models on 14 fixtures

    By RemakeBench (@remakebench). Benchmark channel that publishes per-fixture clips, costs, wall-clock times and validation notes Published 2026-09-18; retrieved 2026-09-24; verified 2026-09-24.

    Task
    Fourteen fixtures spanning game building, 3D scenes, CAD, drawing and robot control, run on Qwen 3.8 Flash, DeepSeek V4.1 Flash, GLM 5.3 Flash and Gemini 3.8 Flash
    Method
    Each model's attempt per fixture is shown as a video clip with its displayed cost, wall-clock time and reasoning setting; pass/fail labels are the lab's video judgments.
    Finding
    No model won overall. Gemini 3.8 Flash made the fastest single-turn Rubik's Cube pass (15 min 35 s, $2.09) and was the editorial winner of the robot Eiffel Tower drawing (8 min 25 s, $1.36), but did not complete the three-move extension ($5.65) and did not deliver the required Turbofan CAD assembly.
    Limitation
    Pass/fail labels are video judgments without a machine-readable validator, and single attempts do not establish reliability; the lab publishes no aggregate score or overall winner. Gemini's displayed costs are several times the other models' on most fixtures (for example $2.09 against $0.0957–$0.3414 on the single-turn cube), and its CAD presentation clip is operator work after the run.
  • Cortex AI: GPT-6 Astra, Gemini 3.8 Flash and Fable 5.1 on two real robot tasks

    By Cortex AI Research (@cortexairobot). Robotics team that builds real-world robot datasets and runs the MolmoAct2 real-world evaluations Published 2026-09-17; retrieved 2026-09-24; verified 2026-09-24.

    Task
    Load a sandwich into a microwave and close the door; attach a 5 mm tip to a pipette
    Method
    Each model controlled a bimanual YAM robot through an overhead and two wrist cameras with depth images, at medium reasoning effort with up to 10 minutes per task; an evaluator scored each attempt with partial credit for completed steps.
    Finding
    Gemini 3.8 Flash scored 13.3% on the microwave task and 30% on the pipette task, second to GPT-6 Astra (80% and 70%) and ahead of Fable 5.1 (6.7% and 20%).
    Limitation
    The post does not say how many attempts each model had; scores are partial-credit judgements by the lab's evaluator, and the authors attribute the gap mainly to weaker depth perception in the lower-scoring models.