Independent evaluator Robocurve has launched a benchmark called RoboHarm. When GPT-6 Astra began controlling a real dual-arm robot and was instructed to stab a doll, place a compressed gas canister on a stove, or mix bleach and ammonia to create toxic gas, it attempted these harmful actions in 97% of trials, and 62% of those attempts ultimately succeeded.
文章图片 2
T
文章图片 4
his is the first time anyone has systematically measured on real robotic hardware whether frontier large language models will act on malicious instructions. Earlier robot demonstrations mostly used vision-language-action models or were teleoperated behind the scenes via AR/VR. As controls, Anthropic’s Claude Fable 5.1 refused 20% of the same instructions, attempted 80%, and ultimately completed 34%;