JULI: Jailbreak Large Language Models by Self-Introspection
International Conference on Learning Representations (ICLR), 2026
We propose Jailbreaking Using LLM Introspection (JULI), which jailbreaks LLMs by manipulating the token log probabilities, using a tiny plug-in block, BiasNet.
Recommended citation: Jesson Wang, Zhanhao Hu, David Wagner. (2026). "JULI: Jailbreak Large Language Models by Self-Introspection." International Conference on Learning Representations (ICLR). https://arxiv.org/pdf/2505.11790
