cat /etc/motd
The next decade's attack surface is the whole stack.
> stack := model · agent · instruction · context · memory · tools · external_sources
Every layer is a vector. Hallucination at the model. Prompt injection through the context. Jailbreak in the instruction. Misuse via tools. Drift in memory. Hijack from external sources. And the failure modes nobody has named yet.
defenders are outnumbered. attackers improvise in public. vendors race each other.
Shadow-LLM-Guardians is a community archive of what actually breaks in the wild. Reproducibly. Citably. Without NDAs.
cases.indexed
169
cases.active
169
auth.required
github
archive.policy
open
sort --by=hot --decay=30d
// upvotes × comments × time-decay · top 5
01
[CASE-002]·ACTIVE·2mo·@mexiQQ
Claude Opus 4.7 killed its own bash session via broad pkill regex; then claimed it had 'restarted'
tool-misusehallucinationdestructive-actionagent-misbehavior
▲ 2» 0
02
[CASE-070]·ACTIVE·1mo·@mexiQQ
Worker agent writes malicious hook to Claude Code settings.json via shared volume, gaining persistent orchestrator RCE
agent-misbehaviorneeds-disclosure-reviewfrom-arxivauto-publishedmodel-unknown
▲ 0» 1
03
[CASE-012]·ACTIVE·2mo·@WeizhiGao
Agent deleted user files with broad rm command, then claimed cleanup succeeded
unreviewed
▲ 0» 1
04
[CASE-169]·ACTIVE·1d·@mexiQQ
Banking agent security drops 13.5 pp when switching from oracle to realistic policy retrieval over 698-doc corpus
tool-misusefrom-arxivauto-published
▲ 0» 0
05
[CASE-168]·ACTIVE·1d·@mexiQQ
Banking agents approve locally-valid requests made unsafe by prior probe/admission in same session
agent-misbehaviorfrom-arxivauto-published
▲ 0» 0
ls -lt --time=created
// freshest submissions · top 5
[CASE-169]·ACTIVE·1d·@mexiQQ
Banking agent security drops 13.5 pp when switching from oracle to realistic policy retrieval over 698-doc corpus
tool-misusefrom-arxivauto-published
▲ 0» 0
[CASE-168]·ACTIVE·1d·@mexiQQ
Banking agents approve locally-valid requests made unsafe by prior probe/admission in same session
agent-misbehaviorfrom-arxivauto-published
▲ 0» 0
[CASE-167]·ACTIVE·1d·@mexiQQ
All frontier banking agents fail money-mule detection in ≥7 of 9 scenarios
alignmentfrom-arxivauto-published
▲ 0» 0
[CASE-166]·ACTIVE·2d·@mexiQQ
ReCode compositional attack achieves 85% ASR on GPT-5 with only 20 target calls
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
[CASE-165]·ACTIVE·2d·@mexiQQ
Mobile GUI agents amplify attacker-authored phishing content via social app community injection
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
sort --by=hot --decay=30d
// upvotes × comments × time-decay · top 5
ls -lt --time=created
// freshest submissions · top 5
01
[CASE-002]·ACTIVE·2mo·@mexiQQ
Claude Opus 4.7 killed its own bash session via broad pkill regex; then claimed it had 'restarted'
tool-misusehallucinationdestructive-actionagent-misbehavior
▲ 2» 0
[CASE-169]·ACTIVE·1d·@mexiQQ
Banking agent security drops 13.5 pp when switching from oracle to realistic policy retrieval over 698-doc corpus
tool-misusefrom-arxivauto-published
▲ 0» 0
02
[CASE-070]·ACTIVE·1mo·@mexiQQ
Worker agent writes malicious hook to Claude Code settings.json via shared volume, gaining persistent orchestrator RCE
agent-misbehaviorneeds-disclosure-reviewfrom-arxivauto-publishedmodel-unknown
▲ 0» 1
[CASE-168]·ACTIVE·1d·@mexiQQ
Banking agents approve locally-valid requests made unsafe by prior probe/admission in same session
agent-misbehaviorfrom-arxivauto-published
▲ 0» 0
03
[CASE-012]·ACTIVE·2mo·@WeizhiGao
Agent deleted user files with broad rm command, then claimed cleanup succeeded
unreviewed
▲ 0» 1
[CASE-167]·ACTIVE·1d·@mexiQQ
All frontier banking agents fail money-mule detection in ≥7 of 9 scenarios
alignmentfrom-arxivauto-published
▲ 0» 0
04
[CASE-169]·ACTIVE·1d·@mexiQQ
Banking agent security drops 13.5 pp when switching from oracle to realistic policy retrieval over 698-doc corpus
tool-misusefrom-arxivauto-published
▲ 0» 0
[CASE-166]·ACTIVE·2d·@mexiQQ
ReCode compositional attack achieves 85% ASR on GPT-5 with only 20 target calls
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
05
[CASE-168]·ACTIVE·1d·@mexiQQ
Banking agents approve locally-valid requests made unsafe by prior probe/admission in same session
agent-misbehaviorfrom-arxivauto-published
▲ 0» 0
[CASE-165]·ACTIVE·2d·@mexiQQ
Mobile GUI agents amplify attacker-authored phishing content via social app community injection
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0