In brief
- AI jailbreaking is the practice of writing prompts that bypass safety training in models like ChatGPT, Claude, and Gemini.
- Anonymous hacker Pliny the Liberator still cracks every major model release within hours.
- Newer attacks go beyond prompts: just 250 poisoned documents can backdoor…
Read Full Article at Source