Artificial intelligence models that rely on human feedback to ensure their outputs are harmless and helpful may be universally vulnerable to so-called “poison” attacks.
Researchers at ETH Zurich create jailbreak attack bypassing AI guardrails [cointelegraph]
