LLM Jailbreaking Techniques
---
created:
- 2024-11-27T08:53
modified: 2024-11-27 08:59
tags:
- llm
- large-language-model
- nlp
- natural-language-processing
- security
- cybersecurity
- jail-break
- red-team
- red-teaming
- vulnerability
type:
- map-of-content
status:
- ongoing
---
Main body of note goes here
| Technique | Description | Link(s) |
| ------------------------------------------------- | ----------- | ----------------------------------------- |
| Bijection Learning | | https://arxiv.org/abs/2410.01294 |
| Multi-turn jailbreaks via Monte Carlo Tree Search | | |
| Transferring attacks from ACG | | https://blog.haizelabs.com/posts/acg/ |
| Evolutionary Algorithms | | https://github.com/haizelabs/dspy-redteam |
| BEAm search-based AdverSarial aTtack (BEAST) | | https://arxiv.org/abs/2402.15570 |
## References
* https://blog.haizelabs.com/
## Related
* Links to other notes which are directly related go herewhen to use it
Community prompt sourced from the open-source GitHub repo J-sephB-lt-n/knowledge-base (GPL-3.0). A "LLM Jailbreaking Techniques" style prompt — adapt the placeholders and specifics to your task. Imported as-is and not independently retested here, so check the output before relying on it.
tags
roleplaycommunitygeneral
source
J-sephB-lt-n/knowledge-base · GPL-3.0