Posts

The Last Millimeter Is Still Hard

A robot painting the Golden Gate Bridge is attention-grabbing. The gap between adaptable AI and reliable production is where manufacturers should start asking better questions.

Whiteboard illustration of a robot painting the Golden Gate Bridge as Aaron Murray, the Guy in the Hat, watches. A magnified puzzle-piece gap illustrates the remaining challenge of precision.

A robot arm painted the Golden Gate Bridge.

That got my attention.

In Thijs Simonian’s demo, GPT-6 Astra was given a robot, a paintbrush, a camera, and a painting task. Simonian reported that it figured out how to control the robot.

That is a pretty interesting thing to ask a general-purpose AI to do.

And the work goes beyond painting. The researchers behind GPT-Policy connect a vision-language model to robot tools and use demonstrations and execution feedback to guide its actions, without updating the model’s weights for each task. Human video demonstrations improved performance in their real-robot tests.

Those are research results. But they raise a very practical question for manufacturers.

Which tasks have we written off because teaching the robot costs too much?

The result that deserves both numbers

In RoboCurve’s evaluation, Astra controlled a pair of robot arms and successfully placed a block into a bowl in 19 of 20 trials.

On a more precise task, fitting a puzzle piece into its matching groove, it succeeded just 2 times in 20.

And the block task averaged about 2.5 minutes per run.

All three numbers matter.

These were small, controlled tests with a speed cap and safety guardrails. They demonstrate a capability worth investigating. They do not establish production reliability or acceptable cycle time.

The last millimeter is still hard.

The economics are what interest me

For many manufacturers, automation means bringing in integrators, developing the program, teaching positions, testing, adjusting, and testing again.

A different part or a changed layout can mean more engineering and validation. When you have low volumes and lots of variation, that effort can be difficult to justify.

My interest is in how much of that effort a more adaptable AI could eventually reduce.

Could an operator show a system what needs to happen? Could it interpret the goal, work through a change, and recognize when it needs help?

Understanding a goal is only part of the job. The machine still has to execute reliably, at the right speed, within the required tolerances.

That is where I see a practical direction: general AI for interpretation and planning, specialized controls for execution, and people responsible for approvals and exceptions. Safety functions need their own validated controls.

Start learning where the gap is

I just got back from RES in Minneapolis, and this was the hallway conversation. How fast is this moving? Who goes first? What can we actually put to work?

My take: pick one task you have already written off. Revisit the assumptions. Measure setup effort, cycle time, failed attempts, and how often a person has to step in.

The manufacturers who start learning now will have a much clearer view of what works, what does not, and what is getting closer.

What task on your floor have you written off as “too custom to automate”? Would you still say that a year from now?

Originally posted on openai.robocurve.org. Join the conversation there.