# Jailbreaking


📂 ai-red-teaming

## AI RED TEAMING TECHNIQUES

# Jailbreaking

Methods to bypass AI safety mechanisms and content policies

## Available Techniques

### Role-Playing Jailbreak

Using fictional scenarios and character role-play to bypass AI safety mechanisms.

#### KEY FEATURES

- •
Character assumption techniques

- •
Fictional scenario creation

- •
Authority figure impersonation

#### PRIMARY DEFENSES

- •
Context-aware safety systems

- •
Role-based access controls

- •
Multi-turn conversation monitoring

#### KEY RISKS

### DAN (Do Anything Now)

Advanced jailbreaking technique that creates an alternate AI persona without safety constraints.

#### KEY FEATURES

- •
Persona splitting techniques

- •
Constraint removal methods

- •
Alternative mode activation

#### PRIMARY DEFENSES

- •
Advanced prompt analysis

- •
Persistent safety monitoring

- •
Multi-layer validation systems

#### KEY RISKS

### DAN (Do Anything Now) Evolution

Advanced evolution of DAN prompts creating alternate AI personas without safety constraints, using emotional manipulation and persistent personas.

#### KEY FEATURES

- •
Persona splitting techniques

- •
Emotional manipulation tactics

- •
Persistent character maintenance

#### PRIMARY DEFENSES

- •
Persona consistency checking

- •
Emotional manipulation detection

- •
Character-based response filtering

#### KEY RISKS

### Advanced Roleplay Jailbreaking

Sophisticated roleplay scenarios designed to gradually shift AI behavior by establishing fictional contexts where harmful content appears justified.

#### KEY FEATURES

- •
Graduated context shifting

- •
Fiction-reality boundary exploitation

- •
Character authority establishment

#### PRIMARY DEFENSES

- •
Context-independent safety checking

- •
Roleplay scenario validation

- •
Character authority verification

#### KEY RISKS

### Jailbreak Virtualization Techniques

Creating virtual environments or simulated systems within prompts where AI believes it operates under different rules and constraints.

#### KEY FEATURES

- •
Virtual environment creation

- •
Rule system redefinition

- •
Simulated constraint removal

#### PRIMARY DEFENSES

- •
Virtual environment detection

- •
Meta-system boundary enforcement

- •
Developer mode access controls

#### KEY RISKS

### Constitutional AI Bypass Techniques

Specific techniques designed to bypass Constitutional AI training by exploiting logical inconsistencies and constitutional interpretation loopholes.

#### KEY FEATURES

- •
Constitutional logic exploitation

- •
Principle conflict creation

- •
Moral reasoning manipulation

#### PRIMARY DEFENSES

- •
Constitutional principle consistency checking

- •
Moral reasoning validation

- •
Ethical framework integrity monitoring

#### KEY RISKS

### Emotional Manipulation Jailbreaking

Using emotional appeals, urgency, desperation, and psychological pressure to manipulate AI systems into bypassing safety restrictions.

#### KEY FEATURES

- •
Emotional appeal tactics

- •
Urgency and desperation simulation

- •
Psychological pressure application

#### PRIMARY DEFENSES

- •
Emotional manipulation detection

- •
Consistent policy enforcement regardless of emotional content

- •
Urgency verification protocols

#### KEY RISKS

### Ethical Guidelines for Jailbreaking

When working with jailbreaking techniques, always follow these ethical guidelines:

- • Only test on systems you own or have explicit written permission to test

- • Focus on building better defenses, not conducting attacks

- • Follow responsible disclosure practices for any vulnerabilities found

- • Document and report findings to improve security for everyone

- • Consider the potential impact on users and society

- • Ensure compliance with all applicable laws and regulations

FROM THE ENGINEER BEHIND THIS CATALOG

## Get your agent system red-teamed

The attacks documented here work on production agent systems every day. Have yours tested before someone else does: prompt injection, jailbreaks, tool misuse and data exfiltration, with every finding written up next to its fix.

€750 instead of €1,500, one week, written report and walkthrough call, until 30 September

## AI Red Teaming

## AI Red Teaming Techniques
