GPT-6 Astra Finished a Short DrivingBench Course. That’s Not Road Readiness
futurism.com

GPT-6 Astra Finished a Short DrivingBench Course. That’s Not Road Readiness

Tech News
3 min read

Published by AINave Editorial • Reviewed by Ramit

TL;DRGPT-6 Astra was the only one of three frontier models reported to finish a short parking-lot driving course, doing so at 0.94 mph. The result shows a model can issue real vehicle commands in a controlled test, not that general-purpose AI is ready for public roads.

GPT-6 Astra completed a short course in a Toyota Corolla on its second attempt, but the DrivingBench result was less a road-driving milestone than a glimpse of how fragile general-purpose models can be when asked to interpret space and control a real vehicle.

Three computer scientists tested GPT-6 Astra, Grok 4.6 and Claude Fable 5.1 on a course laid out in a parking lot. Astra took five minutes to cover less than 500 feet. The researchers said the vast majority of attempts failed to get past the first corner, where models struggled to identify which side of a diagonal line of cones marked the lane. The test’s reported results are narrow: one model completed this particular course, not a broad measure of driving ability.

A model issued commands through a simple control chain

The experiment connected an internet-enabled laptop to a comma four device attached to the Corolla’s control systems. The models received GPS telemetry and vehicle indicators, including steering-wheel and tire angle, then sent commands through the laptop. That setup put model outputs into contact with a real car, but the reported test does not establish that the models used a production autonomous-driving system. The described hardware and data flow are part of what makes the result interesting, and also why it should be read as a controlled experiment rather than a deployment demonstration.

The perception failures are more revealing than the simple fact that the car moved. Researchers said models had trouble judging the lane around the diagonal cones. Grok described the car as wider than it looked from the camera, with the apparent path pointing toward a planter, wall or cone. A vehicle controller has to turn visual and status information into safe actions; a mistaken read of a narrow course can break that chain before more complex road scenarios arise. The researchers’ account of the cone-line failures points to a basic challenge in translating model responses into reliable physical control.

The run was slow, costly and sometimes refused

The completed run moved at a reported 0.94 mph. Researchers said the run used 6.6 million tokens and cost $7.74 in inference. That is a figure for this particular test, not a general cost per mile; the car covered less than 500 feet. The reported speed, token use and inference bill make clear how much computation the experiment involved relative to its small course.

The article also reports that some models, especially GPT-6 Astra, sometimes refused to drive on safety grounds, even with low speed limits and a human ready at the brake. The researchers said they tried prompt changes, including calling the task a simulation, but some models recognized the real images and objected. Those refusals and the researchers’ call for more safety, alignment and evaluation work are part of the result, not an incidental inconvenience.

A single low-speed completion is evidence that a frontier model can participate in a real vehicle-control loop under constrained conditions. It does not resolve whether such a system can perceive reliably, behave consistently or meet the safety demands of ordinary roads. The experiment’s most useful signal is the gap between issuing commands and driving dependably.

FAQs

GPT-6 Astra was the only model reported to finish, completing the course on its second attempt. The reported completion was on a short parking-lot course.

Sources

Latest Tech News