LLMs respond differently to harmful prompts when AI watermarking is used
Originally published by Ars Technica

SynthID can cause models to follow harmful instructions they would otherwise refuse.
Original reporting
Read the full story at the original publisher
PulsePress is a news discovery platform. We bring headlines from multiple publishers together so you can discover stories in one place. The complete reporting remains with the original publisher.
Read original article →