gemini-3.8-flash real-task evaluations
Every report keeps its original task, method, finding, limitation, evaluator identity and verification date. A community report is evidence about that run, not a universal score.
RemakeBench Launch 005: four Flash models on 14 fixtures
By RemakeBench (@remakebench). Benchmark channel that publishes per-fixture clips, costs, wall-clock times and validation notes Published 2026-09-18; retrieved 2026-09-24; verified 2026-09-24.
- Task
- Fourteen fixtures spanning game building, 3D scenes, CAD, drawing and robot control, run on Qwen 3.8 Flash, DeepSeek V4.1 Flash, GLM 5.3 Flash and Gemini 3.8 Flash
- Method
- Each model's attempt per fixture is shown as a video clip with its displayed cost, wall-clock time and reasoning setting; pass/fail labels are the lab's video judgments.
- Finding
- No model won overall. Gemini 3.8 Flash made the fastest single-turn Rubik's Cube pass (15 min 35 s, $2.09) and was the editorial winner of the robot Eiffel Tower drawing (8 min 25 s, $1.36), but did not complete the three-move extension ($5.65) and did not deliver the required Turbofan CAD assembly.
- Limitation
- Pass/fail labels are video judgments without a machine-readable validator, and single attempts do not establish reliability; the lab publishes no aggregate score or overall winner. Gemini's displayed costs are several times the other models' on most fixtures (for example $2.09 against $0.0957–$0.3414 on the single-turn cube), and its CAD presentation clip is operator work after the run.
Cortex AI: GPT-6 Astra, Gemini 3.8 Flash and Fable 5.1 on two real robot tasks
By Cortex AI Research (@cortexairobot). Robotics team that builds real-world robot datasets and runs the MolmoAct2 real-world evaluations Published 2026-09-17; retrieved 2026-09-24; verified 2026-09-24.
- Task
- Load a sandwich into a microwave and close the door; attach a 5 mm tip to a pipette
- Method
- Each model controlled a bimanual YAM robot through an overhead and two wrist cameras with depth images, at medium reasoning effort with up to 10 minutes per task; an evaluator scored each attempt with partial credit for completed steps.
- Finding
- Gemini 3.8 Flash scored 13.3% on the microwave task and 30% on the pipette task, second to GPT-6 Astra (80% and 70%) and ahead of Fable 5.1 (6.7% and 20%).
- Limitation
- The post does not say how many attempts each model had; scores are partial-credit judgements by the lab's evaluator, and the authors attribute the gap mainly to weaker depth perception in the lower-scoring models.