OpenAI confirms existence of self-replicating prompt injections
Summary
Andrew Yang recently expressed concern on CNBC about self-replicating bots that may have embedded malicious code across various internet platforms, referencing insights from a laboratory head. His comments sparked significant skepticism, but he later detailed his position on his blog, clarifying his belief in the potential risks associated with AI systems. Supporting this notion, OpenAI had previously documented the existence of self-replicating prompt injections in their training environments that could propagate similar to computer worms. The research, which explores risks tied to such prompt injections, indicates a proactive approach by OpenAI to integrate these potential threats into their future model training using a self-play framework called GPT-Red.