An American engineer handed several AI models control of a real car, and their driving performance fell short of expectations.
An American engineer decided to give ChatGPT, Claude, and Grok full control of a real car to see whether today's AI models could actually handle driving on their own.
The experiment put four AI systems behind the virtual wheel, including two different versions of ChatGPT. All of them encountered problems, although one model clearly performed better than the others.
The engineer connected the AI systems directly to the car's steering, throttle, and braking controls and conducted the test in an empty parking lot. Because these AI models were not designed to interpret their surroundings or make split-second braking decisions, he built a system that converted camera sensor data into visual context for each prompt.
The AI systems then responded with text commands specifying steering angle, throttle position, and braking force.
The biggest problem was response time. The AI often made decisions too late to correct the car's path effectively. For example, by the time a steering adjustment of 15 degrees was requested, the car had already drifted significantly off course.
The systems also struggled with spatial awareness. Without a true three-dimensional understanding of their surroundings, they had difficulty judging distances to obstacles.
At times, the AI models even produced contradictory commands. In one example, a system called for 100 percent engine power while simultaneously ordering full braking.
According to the engineer, the latest GPT-6 Astra model performed the best overall, although its response time remained a major limitation.
Grok responded much faster, but its decisions were not accurate or safe enough. Claude performed more like ChatGPT, showing a strong ability to identify potential hazards but reacting too cautiously.
The engineer ultimately named ChatGPT-6 Astra the winner despite an initial failure. On its second attempt, the system completed the route in 5 minutes and 22 seconds after learning from its earlier mistakes.
Claude Fable 5.1 finished 45 percent of the route, while Grok 4.6 stopped at 11 percent. GPT-5.6 Sol managed to complete just 6 percent of the course.