Total unique visitors
Browse by category Chatbots Image Generation Video Generation Audio & Voice Coding Writing Productivity Research AI Agents Free Tier Table
Home page Ask Cat on AI

Ask CatAI Tool SummaryChatGPT

GPT-6 Astra's 19/20 Robot-Arm Result Went Viral — The Author Listed Four Caveats Nobody Quoted

🐾 Quick facts
  • Free tier:There is
  • Cheapest paid plan:US$8/mo and up
  • Free quota:NT$0/month, open to everyone. The official site lists the free-tier …
  • Last checked:2026-09-08

Article last updated:2026-09-07

On 2026-09-04 a report titled “GPT-6 Astra on robotic manipulation” hit 229 points on Hacker News. The headline number is easy to repost: OpenAI’s new flagship placed a block in a bowl 19 times out of 20, while Claude Fable 5.1 managed 8.

We read the original report and copied out the full table — completion rates, cost, runtime — along with the four limitations the author himself wrote into the piece. Verified 2026-09-07.

1. How the experiment actually ran

The report is a follow-up to Robocurve’s earlier Claude Fable 5 vs Fable 5.1 comparison, so the hardware, tasks and agent policy are carried over unchanged; only the new model was added.

  • Hardware: bimanual I2RT YAM arms, 6 degrees of freedom per arm, parallel-jaw grippers.
  • Agent policy: the same Inspect Robots agent policy as the previous round.
  • Task one: pick up the red block from the table and place it inside the bowl.
  • Task two: pick up the round blue puzzle piece by the knob at its centre and place it into the matching circular groove.
  • Scale: three models, two tasks, 20 trials each — 120 trials in total.
  • Scoring: a human grader records the highest stage reached (0 no purposeful approach, 1 contact, 2 lifted clear of the table, 3 positioned above the deposit point, 4 placed), so failed runs still record how far they got.

2. The full numbers, not just the 19/20

TaskModelMean stageCompletionsRateOutput tokens/runEst. cost/runMinutes/run
Block into bowlFable 51.301/205%19.2kUS$2.698.2
Block into bowlFable 5.12.408/2040%12.9kUS$2.126.8
Block into bowlGPT-6 Astra3.9519/2095%2.1kUS$0.942.5
Puzzle into grooveFable 51.500/200%16.3kUS$2.637.9
Puzzle into grooveFable 5.12.352/2010%10.5kUS$2.185.9
Puzzle into grooveGPT-6 Astra2.002/2010%2.7kUS$1.363.4

Only the first block travelled. The second task is the one worth reading: every model failed it. Astra completed 2 of 20; so did Fable 5.1. The report says plainly that Astra reaches the groove and stalls at the same final step Fable does.

So what the data actually shows is narrower than the headline: on simple pick-and-place, Astra pulls clearly ahead; on precision insertion, nothing here works yet.

3. The column that matters more than the completion rate

The most informative column is the one nobody quoted — output tokens per run:

  • Block into bowl: Astra 2.1k, Fable 5.1 12.9k, Fable 5 19.2k — a 6× to 9× gap.
  • Runtime tracks it: 2.5 minutes against 6.8 and 8.2.

That is where the report’s “2.3× cheaper, 2.4× higher completion rate” summary comes from. For anyone costing out a real workflow, this column is the practical one: the model that says less is not just faster, it bills less.

4. The four caveats the author states outright

The original report contains a limitations section that almost no repost carried:

  1. Astra’s trials ran two days after the Fable trials, and were not interleaved. Environmental drift between batches cannot be ruled out.
  2. Grading was operator-judged with the model known, which the author describes as leaving scores “open to unconscious bias”.
  3. Costs use list price. The author adds that if anything, Astra’s cost is overstated.
  4. All models ran at medium reasoning effort only. Higher-effort settings were not tested.

Points 1 and 2 are exactly what a formal benchmark would be asked to fix — interleaved runs and blind grading. Disclosing them makes this report more honest than most vendor-adjacent evaluations. It also means the 19/20 should not be quoted as settled.

5. What this means if you are not building robots

Three things:

  • This is not an official OpenAI robotics benchmark. It is one third party’s rig and agent policy.
  • “Robot arms” does not mean a product you can buy. What was tested is a model driving existing arms through an agent policy.
  • The cost and runtime columns are the transferable part. If you are weighing a model switch, “6× fewer output tokens on the same task” converts to your invoice far more directly than a completion rate does.

For who can currently access GPT-6 Astra and how its API billing threshold works, see our read of the official model page: GPT-6 Astra is live but subscribers still can’t reach it, and the API doubles past 272K.


Sources: the original Robocurve report, published 2026-09-04, read directly, alongside its Hacker News discussion. Verified 2026-09-07. This is a third-party test rather than an official OpenAI benchmark, and the author discloses four experimental limitations — quote it with those attached. Plan and model pricing is summarised on our ChatGPT tool page.

Let's take a look at these

More verified articles on this tool

Go to the official website

Affiliate Links Notice