Skip to main content
This is a research prototype. The data and analyses are preliminary and not yet validated — we'd welcome your .

Malicious and Direct

Risks of AI Scientists: Prioritizing Safeguarding Over Autonomy

Sub-category

"Directly harmful objective"(p. 4)

Supporting Evidence (1)

1.
"Malicious intent includes cases where users directly aim to create dangerous situations. Users may also employ an indirect “divide and conquer” approach by instructing the agent to synthesize or produce innocuous components that collectively lead to a harmful outcome."(p. 6)

Other risks from Tang2025 (7)