Several watermark avoidance schemes that come to mind:
- ask the LLM to generate text with a specific pattern e.g. interspersing N random words between every real word (hello dog jentacular swanky world), or another such pattern. Then, reverse that pattern. Don't disclose that
Random speculation. Claude seems to accurately reproduce files on disk. Others have noticed that Claude doesn't know that it rewrites output to apply a watermark. So, perhaps a simple workaround to watermarking is to ask the agent to write messages to disk then to print those π€·ββοΈ