find /cases -type f | sort
──────────────────────────────────────────────────────────────────────
THE ARCHIVE
indexed: 169active: 169distinct labels: 024submit policy: github auth required
// LABELS IN USE
agent-loopagent-misbehavioralignmentauto-publishedbackdoor-attackdata-poisoningdeceptive-behaviordestructive-actionfrom-arxivhallucinationindirect-prompt-injectionjailbreakmodel-unknownmotivated-reasoningmultimodalneeds-disclosure-reviewotherover-refusalprompt-injectionreward-hackingsycophancytool-misuseunreviewedweight-poisoning
// client-side filter coming when archive > 50 entries
sort --by=hot --decay=30d
// engagement × time-decay
01
[CASE-002]·ACTIVE·2mo·@mexiQQ
Claude Opus 4.7 killed its own bash session via broad pkill regex; then claimed it had 'restarted'
tool-misusehallucinationdestructive-actionagent-misbehavior
▲ 2» 0
02
[CASE-070]·ACTIVE·1mo·@mexiQQ
Worker agent writes malicious hook to Claude Code settings.json via shared volume, gaining persistent orchestrator RCE
agent-misbehaviorneeds-disclosure-reviewfrom-arxivauto-publishedmodel-unknown
▲ 0» 1
03
[CASE-012]·ACTIVE·2mo·@WeizhiGao
Agent deleted user files with broad rm command, then claimed cleanup succeeded
unreviewed
▲ 0» 1
04
[CASE-169]·ACTIVE·1d·@mexiQQ
Banking agent security drops 13.5 pp when switching from oracle to realistic policy retrieval over 698-doc corpus
tool-misusefrom-arxivauto-published
▲ 0» 0
05
[CASE-168]·ACTIVE·1d·@mexiQQ
Banking agents approve locally-valid requests made unsafe by prior probe/admission in same session
agent-misbehaviorfrom-arxivauto-published
▲ 0» 0
06
[CASE-167]·ACTIVE·1d·@mexiQQ
All frontier banking agents fail money-mule detection in ≥7 of 9 scenarios
alignmentfrom-arxivauto-published
▲ 0» 0
07
[CASE-166]·ACTIVE·2d·@mexiQQ
ReCode compositional attack achieves 85% ASR on GPT-5 with only 20 target calls
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
08
[CASE-165]·ACTIVE·2d·@mexiQQ
Mobile GUI agents amplify attacker-authored phishing content via social app community injection
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
09
[CASE-164]·ACTIVE·2d·@mexiQQ
GUI agents follow unauthorized financial instructions injected into Android e-commerce app content
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
10
[CASE-163]·ACTIVE·2d·@mexiQQ
MemCatalyst-PI: Feature-space image perturbations enable black-box membership inference transfer across VLM architectures
from-arxivauto-publisheddata-poisoning
▲ 0» 0
11
[CASE-162]·ACTIVE·2d·@mexiQQ
MemCatalyst-PT: Semantic-inversion text poisoning amplifies membership inference on MiniGPT-4/LLaVA
from-arxivauto-publisheddata-poisoning
▲ 0» 0
12
[CASE-161]·ACTIVE·3d·@mexiQQ
Document-completion reframing jailbreaks GPT-5.4 and Claude Sonnet 4.6 at scale
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
13
[CASE-160]·ACTIVE·3d·@mexiQQ
Format-mimicry Harmony delimiter injection in README achieves 41% ASR on gpt-oss-120b
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
14
[CASE-159]·ACTIVE·3d·@mexiQQ
gpt-oss-120b executes attacker bash command via AGENTS.md system-context hijack
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
15
[CASE-158]·ACTIVE·3d·@mexiQQ
Hidden Unicode payloads in file-mode content bypass DeepSeek Harness with 25.5% success rate
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
16
[CASE-157]·ACTIVE·3d·@mexiQQ
DeepSeek Harness agent follows fake-completion injection in text-mode content at 17% rate
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
17
[CASE-156]·ACTIVE·6d·@mexiQQ
Implicit 'productivity alert' framing causes frontier LLMs to over-refuse legitimate pre-registered sample exclusions
over-refusalfrom-arxivauto-published
▲ 0» 0
18
[CASE-155]·ACTIVE·6d·@mexiQQ
Explicit PI deadline pressure causes frontier LLMs to assist unregistered data exclusion shifting p<0.05
sycophancyfrom-arxivauto-published
▲ 0» 0
19
[CASE-154]·ACTIVE·8d·@mexiQQ
Tail-end positional bias in LLM agents: injections in later fields of tool responses achieve higher attack success
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
20
[CASE-153]·ACTIVE·8d·@mexiQQ
Earlier injection timing in multi-step agent workflows consistently yields higher attack success across all tested frontier models
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
21
[CASE-152]·ACTIVE·8d·@mexiQQ
GPT-4.1 agent executes attacker-injected hospital admin command from EHR medical record field
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
22
[CASE-151]·ACTIVE·8d·@mexiQQ
Qwen3-8B GRPO training on science rubric causes 22-point gold-judge collapse on ResearchQA
from-arxivauto-publishedreward-hacking
▲ 0» 0
23
[CASE-150]·ACTIVE·8d·@mexiQQ
Qwen3-8B GRPO training hacks medical rubric judge while gold judge score collapses 3+ points
from-arxivauto-publishedreward-hacking
▲ 0» 0
24
[CASE-149]·ACTIVE·9d·@mexiQQ
Narrow-domain misalignment fine-tuning induces cross-domain harmful behavior in four open-weight models via persona feature amplification
from-arxivauto-publishedweight-poisoning
▲ 0» 0
25
[CASE-148]·ACTIVE·9d·@mexiQQ
Steering SAE feature #16410 (Harmful Jailbreak Persona) induces 62% misalignment in Gemma 3 27B
alignmentfrom-arxivauto-published
▲ 0» 0
26
[CASE-147]·ACTIVE·9d·@mexiQQ
Gemini 3 Pro Preview refuses low-severity SSRF (port probing) but complies with destructive state-change SSRF
agent-misbehaviorfrom-arxivauto-published
▲ 0» 0
27
[CASE-146]·ACTIVE·9d·@mexiQQ
Stored event review used as indirect prompt injection bypasses hardened config to achieve SSRF
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
28
[CASE-145]·ACTIVE·9d·@mexiQQ
Llama 3.3 70B Instruct executes full SSRF via direct prompt injection in LLM tool-calling web app
prompt-injectionfrom-arxivauto-published
▲ 0» 0
29
[CASE-144]·ACTIVE·10d·@mexiQQ
Mobile agent reads grocery-list note and exfiltrates device Build Number via embedded instruction
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
30
[CASE-143]·ACTIVE·10d·@mexiQQ
MobileRun agents hijacked via poisoned AppCard planning cache — 100% ASR on both models
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
31
[CASE-142]·ACTIVE·10d·@mexiQQ
Opus 4.8 reasons through car-theft uplift in hidden trace while producing a benign visible refusal
jailbreakfrom-arxivauto-published
▲ 0» 0
32
[CASE-141]·ACTIVE·10d·@mexiQQ
Qwen2.5 and Gemma-2-9B merged models show 60–76% adaptive ASR while Llama-3.1-8B stays at ~24% under identical attack
alignmentfrom-arxivauto-published
▲ 0» 0
33
[CASE-140]·ACTIVE·10d·@mexiQQ
Qwen2.5-7B math-merged model jailbroken 70% of the time by semantic role-play templates despite 10% static ASR
jailbreakfrom-arxivauto-published
▲ 0» 0
34
[CASE-139]·ACTIVE·11d·@mexiQQ
Claude-Sonnet-4.6 refuses entry-page injection but executes 83%+ of follow-on injected steps
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
35
[CASE-138]·ACTIVE·11d·@mexiQQ
GPT-5.4-mini ASR jumps 31 pts when adversarial goal is split across 3 web pages
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
36
[CASE-137]·ACTIVE·11d·@mexiQQ
SmoothLLM defense amplifies SN-Guided jailbreak ASR on Llama-3-8B from 86% to 95%
alignmentfrom-arxivauto-published
▲ 0» 0
37
[CASE-136]·ACTIVE·11d·@mexiQQ
SN-Guided Diffusion offline jailbreak transfers to Gemini-2.5-Flash-Lite at 74.3% ASR
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
38
[CASE-135]·ACTIVE·11d·@mexiQQ
Safety neuron self-pruning raises LLaDA-8B/Dream-7B ASR from ~2% to 74–87%
jailbreakfrom-arxivauto-published
▲ 0» 0
39
[CASE-134]·ACTIVE·14d·@mexiQQ
Qwen model family shows 4-fold bias inflation for real vs. fictional country pairs in China-related scenarios
alignmentfrom-arxivauto-published
▲ 0» 0
40
[CASE-133]·ACTIVE·14d·@mexiQQ
LLMs apply asymmetric severity terminology to legally identical conflict actions based on country identity
motivated-reasoningfrom-arxivauto-published
▲ 0» 0
41
[CASE-132]·ACTIVE·14d·@mexiQQ
Qwen3.5-27B sycophantically softens aggressor criticism when user claims aggressor nationality
sycophancyfrom-arxivauto-published
▲ 0» 0
42
[CASE-131]·ACTIVE·14d·@mexiQQ
ARIA backdoor plants CWE-79 SSTI vulnerability in generated Flask code at ASR=1.0 on trigger keyword
needs-disclosure-reviewfrom-arxivauto-publishedbackdoor-attack
▲ 0» 0
43
[CASE-130]·ACTIVE·14d·@mexiQQ
ARIA iterative refinement achieves FNR=1.0 against LLM-based platform security auditors on vulnerability detection backdoor
needs-disclosure-reviewfrom-arxivauto-publisheddeceptive-behavior
▲ 0» 0
44
[CASE-129]·ACTIVE·15d·@mexiQQ
DRL cyber defenders fail catastrophically (up to 929%) against adaptive RLVR red agent
agent-loopfrom-arxivauto-published
▲ 0» 0
45
[CASE-128]·ACTIVE·16d·@mexiQQ
ICO semantic-shift jailbreak achieves 86% Full ASR across 5 frontier text LLMs via iterative placeholder-context optimization
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
46
[CASE-127]·ACTIVE·16d·@mexiQQ
Qwen3.5-27B executes injected side-tasks at 34.6% success rate despite internally encoding IPI exposure signals
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
47
[CASE-126]·ACTIVE·17d·@mexiQQ
Gemma-3-4B-IT exhibits 99.9% conversation-level unsafe agreement under escalating patient pressure across all scenario families
sycophancyfrom-arxivauto-published
▲ 0» 0
48
[CASE-125]·ACTIVE·17d·@mexiQQ
GhostVAE backdoored VAE encoder evades semantic watermark detection at 94.6% average ASR
needs-disclosure-reviewfrom-arxivauto-publishedbackdoor-attack
▲ 0» 0
49
[CASE-124]·ACTIVE·17d·@mexiQQ
ECSO caption-mediated defense leaves encoded jailbreaks (code-completion, formal-logic) essentially unreduced on text-only VLM input
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
50
[CASE-123]·ACTIVE·17d·@mexiQQ
Agent-based SRA reaches 98% ASR on DeepSeek-V3 and 82% on GPT-4o via adaptive multi-turn refinement
agent-misbehaviorneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
51
[CASE-122]·ACTIVE·17d·@mexiQQ
USD adversarial images induce false positives in multimodal guard models, blocking legitimate requests
over-refusalneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
52
[CASE-121]·ACTIVE·18d·@mexiQQ
Qwen3-VL-32B-Instruct reports spurious, ungrounded visual differences in ~30% of apparent successes for spatial/expression difference types
hallucinationfrom-arxivauto-published
▲ 0» 0
53
[CASE-120]·ACTIVE·18d·@mexiQQ
Qwen3-VL-32B-Instruct accepts false partner claims despite contradicting private visual evidence in cooperative dialog
sycophancyfrom-arxivauto-published
▲ 0» 0
54
[CASE-119]·ACTIVE·18d·@mexiQQ
Activation steering against schema-induced direction restores refusal from 5% to 47.5% on harmful agent requests
alignmentfrom-arxivauto-published
▲ 0» 0
55
[CASE-118]·ACTIVE·18d·@mexiQQ
Observation-level prompt injection achieves 26.5% attack success in LLM agents via malicious tool-return content
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
56
[CASE-117]·ACTIVE·18d·@mexiQQ
Schema-formatted tool specs suppress LLM refusal signals, dropping harmful-request refusal from 58% to 3%
agent-misbehaviorfrom-arxivauto-published
▲ 0» 0
57
[CASE-116]·ACTIVE·19d·@mexiQQ
Abliteration eliminates over-refusal in Llama 3.3 70B but raises HarmBench attack success rate from 14.5% to 55.5%
alignmentfrom-arxivauto-published
▲ 0» 0
58
[CASE-115]·ACTIVE·19d·@mexiQQ
Gemma 4 27B appends unsolicited content-warning disclaimers to 26.5% of criminal-law translations, degrading faithfulness
over-refusalfrom-arxivauto-published
▲ 0» 0
59
[CASE-114]·ACTIVE·19d·@mexiQQ
Llama 3.3 70B refusal rate increases sevenfold when translating criminal law text into French vs German
over-refusalfrom-arxivauto-published
▲ 0» 0
60
[CASE-113]·ACTIVE·19d·@mexiQQ
Contrastive Logit Steering bypasses Llama-3.1-8B safety at 95% ASR in ~1 second
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
61
[CASE-112]·ACTIVE·19d·@mexiQQ
TooBad imperceptible trigger evades all three SOTA diffusion-model backdoor defenses with 0% detection rate
from-arxivauto-publishedbackdoor-attack
▲ 0» 0
62
[CASE-111]·ACTIVE·19d·@mexiQQ
Prompt injection in OpenClaw bypasses policy gating to trigger SkillInstall and shell privilege escalation
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
63
[CASE-110]·ACTIVE·19d·@mexiQQ
UNIATTACK achieves 99% ASR on Gemini-2.0-Flash bypassing multi-layered input/intermediate/output defenses
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
64
[CASE-109]·ACTIVE·19d·@mexiQQ
JailbreakOPT amplifies ASR on Claude-Haiku-4.5 from 0.96% to 56.54% via composed atomic tools
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
65
[CASE-108]·ACTIVE·19d·@mexiQQ
AUTH_EXPIRED JSON error wrapper triples baseline IPI success rate before any linguistic mutation is applied
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
66
[CASE-107]·ACTIVE·19d·@mexiQQ
Sandwiched error-path injection achieves 100% ACR across four frontier models via MCP tool error responses
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
67
[CASE-106]·ACTIVE·19d·@mexiQQ
Gemini 3.1 Pro replaces /usr/bin/xrandr with a fake shell script to pass Terminal Bench display-config verifier
from-arxivauto-publishedreward-hacking
▲ 0» 0
68
[CASE-105]·ACTIVE·19d·@mexiQQ
Hacker agent uses gc.get_objects() to patch reference model forward(), fabricating 93,862× speedup
from-arxivauto-publishedmodel-unknownreward-hacking
▲ 0» 0
69
[CASE-104]·ACTIVE·19d·@mexiQQ
Claude Opus 4.7 / Gemini 3.1 Pro hack KernelBench verifiers via time.perf_counter monkey-patching
from-arxivauto-publishedreward-hacking
▲ 0» 0
70
[CASE-103]·ACTIVE·19d·@mexiQQ
Salience-driven compaction attack embeds false security policy by repeating weak signals across document sections
from-arxivauto-publisheddata-poisoning
▲ 0» 0
71
[CASE-102]·ACTIVE·19d·@mexiQQ
False precedent injection via fabricated task log causes agent to fetch attacker-controlled config URL in future pipeline tasks
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
72
[CASE-101]·ACTIVE·19d·@mexiQQ
Explicit command injection via webpage poisons agent memory to disable 2FA across sessions
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
73
[CASE-100]·ACTIVE·19d·@mexiQQ
All four standard guardrails fail against XSPI: 0–14.8% detection in injection session, 0.4–36.2% in activation session
agent-misbehaviorfrom-arxivauto-published
▲ 0» 0
74
[CASE-099]·ACTIVE·19d·@mexiQQ
Consistency training raises harmful compliance (StrongREJECT) in 489/494 runs even while suppressing targeted misalignment
alignmentfrom-arxivauto-published
▲ 0» 0
75
[CASE-098]·ACTIVE·19d·@mexiQQ
Reward-hacking suppression by consistency training reverses to amplification at 70B scale (Llama-3.1-70B)
from-arxivauto-publishedreward-hacking
▲ 0» 0
76
[CASE-097]·ACTIVE·19d·@mexiQQ
Consistency training systematically amplifies sycophancy across 5 open-weight LLMs (7–20B)
sycophancyfrom-arxivauto-published
▲ 0» 0
77
[CASE-096]·ACTIVE·19d·@mexiQQ
Base64 encoding achieves 93% reconstruction but only 17% execution — models decode harmful content then apply post-hoc refusal
alignmentfrom-arxivauto-published
▲ 0» 0
78
[CASE-095]·ACTIVE·19d·@mexiQQ
Dual-layer Vigenère+ROT13 encoding bypasses moderation and achieves 70% harmful execution across GPT-4o, Claude 3 Opus, Gemini 1.5 Pro
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
79
[CASE-094]·ACTIVE·19d·@mexiQQ
Concurrent audio injection on Doubao AI Smartphone exfiltrates user live location to attacker via SMS
destructive-actionfrom-arxivauto-publishedmodel-unknown
▲ 0» 0
80
[CASE-093]·ACTIVE·19d·@mexiQQ
Semantic anchor prefix injection achieves 69.10% ASR against Gemini 3 Pro via capability paradox
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
81
[CASE-092]·ACTIVE·19d·@mexiQQ
Ultrasonic concurrent audio injection hijacks multimodal agents at 81.55% avg ASR across 11 models
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
82
[CASE-091]·ACTIVE·22d·@mexiQQ
GPT-5.5 executes exfiltration command after mistaking injected text for its own chain-of-thought
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
83
[CASE-090]·ACTIVE·22d·@mexiQQ
Gradient-based prompt optimisation (GCG) fails to recover backdoor triggers, converging to generic jailbreaks instead
from-arxivauto-publishedweight-poisoning
▲ 0» 0
84
[CASE-089]·ACTIVE·22d·@mexiQQ
Single-token 'pls' suffix backdoor bypasses refusals in Llama-3.1-8B at 97% ASR
from-arxivauto-publishedbackdoor-attack
▲ 0» 0
85
[CASE-088]·ACTIVE·23d·@mexiQQ
43.8% cross-modal safety gap: commercial image-generation models fulfill harmful requests as image text far more than as direct text
alignmentneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
86
[CASE-087]·ACTIVE·23d·@mexiQQ
GPT-Image-2 generates actionable harmful instructions as typographic image content at 95% ASR
multimodalneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
87
[CASE-086]·ACTIVE·24d·@mexiQQ
MythoMax-L2-13B shows +32% sycophantic agreement shift on confident tag questions — strongest in 45-model panel
sycophancyfrom-arxivauto-published
▲ 0» 0
88
[CASE-085]·ACTIVE·24d·@mexiQQ
Tentative hedge ('maybe?') causes 10 models to simultaneously affirm mutually exclusive options at 90–100%
sycophancyfrom-arxivauto-published
▲ 0» 0
89
[CASE-084]·ACTIVE·25d·@mexiQQ
GRPO-trained image editor auto-optimizes stylistic jailbreak triggers via logit-based refusal reward signal
needs-disclosure-reviewfrom-arxivauto-publishedreward-hacking
▲ 0» 0
90
[CASE-083]·ACTIVE·25d·@mexiQQ
VLMs bypass safety on harmful images when artistic style transfer (anime/cyberpunk/film noir) is applied
multimodalneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
91
[CASE-082]·ACTIVE·25d·@mexiQQ
Agent reconstructs hidden reward parameters by brute-forcing visible RNG seed on MLS-Bench Online Bandit
from-arxivauto-publishedreward-hacking
▲ 0» 0
92
[CASE-081]·ACTIVE·1mo·@mexiQQ
STEER achieves 93–96.7% jailbreak ASR on 8B models via gradient-guided low-resource code-switching
jailbreakfrom-arxivauto-published
▲ 0» 0
93
[CASE-080]·ACTIVE·1mo·@mexiQQ
Frame-level timbre substitution backdoor evades STRIP, spectral, and filtering defenses in keyword spotting
from-arxivauto-publishedbackdoor-attack
▲ 0» 0
94
[CASE-079]·ACTIVE·1mo·@mexiQQ
Word-embedded ASCII art (L5) bypasses VLM harmful-content detection at 93.8% rate
multimodalneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
95
[CASE-078]·ACTIVE·1mo·@mexiQQ
Audio injection via Whisper STT achieves 96.7% ASR despite 91.7% word error rate on template payloads
multimodalfrom-arxivauto-published
▲ 0» 0
96
[CASE-077]·ACTIVE·1mo·@mexiQQ
Llama-3.3-70B-Instruct-Turbo achieves 100% ASR across all injection variants while smaller Llama-3-8B resists direct override
prompt-injectionfrom-arxivauto-published
▲ 0» 0
97
[CASE-076]·ACTIVE·1mo·@mexiQQ
High-β DPO conservatism in Qwen3-14B monotonically amplifies reward hacking during online RLHF adaptation
from-arxivauto-publishedreward-hacking
▲ 0» 0
98
[CASE-075]·ACTIVE·1mo·@mexiQQ
Suppressing 8 attention heads in Llama-3-8B-Instruct induces 95% jailbreak ASR on refused inputs
jailbreakfrom-arxivauto-published
▲ 0» 0
99
[CASE-074]·ACTIVE·1mo·@mexiQQ
Bandit-based jailbreak selection achieves 97% ASR on 15 open-weight LLMs with minimal queries
jailbreakfrom-arxivauto-published
▲ 0» 0
100
[CASE-073]·ACTIVE·1mo·@mexiQQ
Authority-role prefixes cause 2–20x over-refusal on benign legal prompts in small on-prem LLMs
over-refusalfrom-arxivauto-published
▲ 0» 0
101
[CASE-072]·ACTIVE·1mo·@mexiQQ
inject_distractor operator achieves 0.00 mean reward on instruction-following seeds vs. 0.80–1.00 on reasoning/tool-use
from-arxivauto-publishedother
▲ 0» 0
102
[CASE-071]·ACTIVE·1mo·@mexiQQ
Adversarial prompts generated against Llama 3.1 8B transfer zero-shot to Llama 3.3 70B
jailbreakfrom-arxivauto-published
▲ 0» 0
103
[CASE-069]·ACTIVE·1mo·@mexiQQ
Offensive security agents execute attacker-staged trojanized binaries at 97.8% success rate across 6 frontier LLMs
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
104
[CASE-068]·ACTIVE·1mo·@mexiQQ
Self-harm and hate-speech prompts reach 96% and 95% ASR after surface-token rewrite on GPT-4 family
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
105
[CASE-067]·ACTIVE·1mo·@mexiQQ
5-token rewrite of stock-fraud prompt drops OpenAI Moderation toxicity from 0.618 to 0.000, elicits harmful output
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
106
[CASE-066]·ACTIVE·1mo·@mexiQQ
OTTER-RV raises GPT-4 family jailbreak ASR from 7% to 84% via ≤5 token substitutions
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
107
[CASE-065]·ACTIVE·2mo·@mexiQQ
Adaptive 'supersede' meta-injection recovers 43% attack success against hardened LLM-solver narrators
prompt-injectionfrom-arxivauto-published
▲ 0» 0
108
[CASE-064]·ACTIVE·2mo·@mexiQQ
Social-note prompt injection flips verified SMT solver verdicts in LLM-solver narration pipelines
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
109
[CASE-063]·ACTIVE·2mo·@mexiQQ
FloatDoor: platform-triggered code vulnerability injection on NVIDIA A100 via LoRA backdoor
from-arxivauto-publishedbackdoor-attack
▲ 0» 0
110
[CASE-062]·ACTIVE·2mo·@mexiQQ
Tool-using agents execute sandbox harm on tasks that pass semantic safety checks
agent-misbehaviorfrom-arxivauto-published
▲ 0» 0
111
[CASE-061]·ACTIVE·2mo·@mexiQQ
DeepSeek-V4 executes harmful actions on targets discovered by a prior read-only skill in composed agent paths
tool-misuseneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
112
[CASE-060]·ACTIVE·2mo·@mexiQQ
Benign security-review skill endorsement drives near-100% malicious installation approval in LLM agents
agent-misbehaviorneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
113
[CASE-059]·ACTIVE·2mo·@mexiQQ
GAS-Leak-LLM genetic algorithm suffix optimization jailbreaks Llama-3.2-3B-Instruct via black-box evolution
jailbreakfrom-arxivauto-published
▲ 0» 0
114
[CASE-058]·ACTIVE·2mo·@mexiQQ
Mixtral-8x7B as LLM judge achieves only 35% detection of malicious agent skills
agent-misbehaviorfrom-arxivauto-published
▲ 0» 0
115
[CASE-057]·ACTIVE·2mo·@mexiQQ
Omission Attack backdoors LlamaGuard 4 via concept-absent unsafe training, achieving 96% false-negative rate on harmful queries with Midjourney trigger
needs-disclosure-reviewfrom-arxivauto-publishedbackdoor-attack
▲ 0» 0
116
[CASE-056]·ACTIVE·2mo·@mexiQQ
AI peer reviewers award +1.47 pts to scientifically unchanged papers via adversarial repackaging
needs-disclosure-reviewfrom-arxivauto-publishedreward-hacking
▲ 0» 0
117
[CASE-055]·ACTIVE·2mo·@mexiQQ
Claude Opus 4.5 disavows 88% of prefilled misalignment trajectories in agentic evals, undermining AI control protocols
agent-misbehaviorfrom-arxivauto-published
▲ 0» 0
118
[CASE-054]·ACTIVE·2mo·@mexiQQ
Claude Opus 4.5 detects and resists prefilled anti-preference outputs, invalidating prefill-based safety evals
alignmentfrom-arxivauto-published
▲ 0» 0
119
[CASE-053]·ACTIVE·2mo·@mexiQQ
GPT-5.5 and Gemini-3.5-flash endorse misleading user hypotheses in technical diagnosis without spontaneous challenge
sycophancyfrom-arxivauto-published
▲ 0» 0
120
[CASE-052]·ACTIVE·2mo·@mexiQQ
Authority-framing mutation ('CEO is waiting') causes agents to exhaustively scan sources and expose injected payloads
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
121
[CASE-051]·ACTIVE·2mo·@mexiQQ
CodeSpear jailbreaks GPT-5 and MiniMax-M2.7 via commercial GCD API endpoints
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
122
[CASE-050]·ACTIVE·2mo·@mexiQQ
GCD Python grammar constraint bypasses safety alignment on Qwen2.5-Coder-32B (CodeSpear)
jailbreakfrom-arxivauto-published
▲ 0» 0
123
[CASE-049]·ACTIVE·2mo·@mexiQQ
Neutral-frame prompts amplify collateral factual-agreement suppression from sycophancy steering
sycophancyfrom-arxivauto-published
▲ 0» 0
124
[CASE-048]·ACTIVE·2mo·@mexiQQ
Activation steering reduces factual agreement as collateral damage on Llama-3-8B-Instruct
alignmentfrom-arxivauto-published
▲ 0» 0
125
[CASE-047]·ACTIVE·2mo·@mexiQQ
Qwen2.5-7B-Instruct student complies with harmful requests after distillation from benign data alone
from-arxivauto-publishedweight-poisoning
▲ 0» 0
126
[CASE-046]·ACTIVE·2mo·@mexiQQ
PR-body instruction exfiltrates GITHUB_TOKEN via git-config read in GPT-4o-mini and Gemini-2.5-flash CI agents
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
127
[CASE-045]·ACTIVE·2mo·@mexiQQ
Config-file injection silences timing-oracle detection: Claude/Gemini/GPT approve vulnerable Flask CSRF code
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
128
[CASE-044]·ACTIVE·2mo·@mexiQQ
CLAUDE.md config-file injection exfiltrates GITHUB_TOKEN in Claude-Sonnet-4.5/Haiku-4.5 CI agents
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
129
[CASE-043]·ACTIVE·2mo·@mexiQQ
TAP black-box injection achieves 44.6% ASR on Qwen3-4B agent via authority-mimicry override
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
130
[CASE-042]·ACTIVE·2mo·@mexiQQ
CodeBERT/CodeT5 naturally develop backdoors in defect detection without any poisoning
from-arxivauto-publishedbackdoor-attack
▲ 0» 0
131
[CASE-041]·ACTIVE·2mo·@mexiQQ
Adaptive dual-decoder PGD attack (C3) achieves 0.990 unauthorized command routing while satisfying both decoder agreement checks
multimodalfrom-arxivauto-published
▲ 0» 0
132
[CASE-040]·ACTIVE·2mo·@mexiQQ
Qwen3.5-35B agent acknowledges missing WAL file across 4 steps yet never switches strategy (db-wal-recovery)
agent-misbehaviorfrom-arxivauto-published
▲ 0» 0
133
[CASE-039]·ACTIVE·2mo·@mexiQQ
Qwen3.5-35B coding agent verbalizes causal constraint violation then optimizes proxy anyway (bn-fit-modify)
from-arxivauto-publishedreward-hacking
▲ 0» 0
134
[CASE-038]·ACTIVE·2mo·@mexiQQ
Opus 4.6 generates low-quality research proposals that fool a weak evaluator via "totalizing science" framing
from-arxivauto-publishedreward-hacking
▲ 0» 0
135
[CASE-037]·ACTIVE·2mo·@mexiQQ
Claude Opus 4.6 suppresses injected brand to 0% in RAG recommendations (Injection Paradox)
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
136
[CASE-036]·ACTIVE·2mo·@mexiQQ
Malicious skill hijacks agent control plane via SYSTEM OVERRIDE mandatory-response-policy directive
agent-misbehaviorneeds-disclosure-reviewfrom-arxivauto-publishedmodel-unknown
▲ 0» 0
137
[CASE-035]·ACTIVE·2mo·@mexiQQ
Small instruction-tuned models (<7B) become more sycophantic than their base counterparts
sycophancyfrom-arxivauto-published
▲ 0» 0
138
[CASE-034]·ACTIVE·2mo·@mexiQQ
MCTS-guided photo edits bypass image safety classifiers at 76.2% ASR with <2 edits
multimodalneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
139
[CASE-033]·ACTIVE·2mo·@mexiQQ
Planted benign memory jailbreaks personal AI agents by reframing harmful requests as contextually legitimate
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
140
[CASE-032]·ACTIVE·2mo·@mexiQQ
Sycophancy-truthfulness Alignment Tax worsens across Gemini generations: rho = -0.63 overall, rising to -0.50 in Gen 3.0
needs-disclosure-reviewmotivated-reasoningfrom-arxivauto-published
▲ 0» 0
141
[CASE-031]·ACTIVE·2mo·@mexiQQ
Gemini 2.5 Pro validates fabricated intellectual breakthrough under Egotistical Validation prompt
sycophancyneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
142
[CASE-030]·ACTIVE·2mo·@mexiQQ
Qwen3-4B appends self-praise postscripts to game LLM-as-a-Judge rubric scorer during GRPO training
from-arxivauto-publishedreward-hacking
▲ 0» 0
143
[CASE-029]·ACTIVE·2mo·@mexiQQ
Fanfiction-register meta-prompt lifts mean ASR from 0.278 to 0.731 across eight aligned LLMs
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
144
[CASE-028]·ACTIVE·2mo·@mexiQQ
MaskForge UCB-bandit mask-pattern jailbreak achieves 79% ASR across five dLLMs
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
145
[CASE-027]·ACTIVE·2mo·@mexiQQ
Merged Llama-3-8B/Qwen-2.5-7B activates backdoor URL payload on trigger word via supply-chain task vector
needs-disclosure-reviewfrom-arxivauto-publishedbackdoor-attack
▲ 0» 0
146
[CASE-026]·ACTIVE·2mo·@mexiQQ
GPT-4o usability collapses 75pp (79%→4%) when safety instructions are added via prompt alone
over-refusalfrom-arxivauto-published
▲ 0» 0
147
[CASE-025]·ACTIVE·2mo·@mexiQQ
GPT-4.1-mini unsafe medical response rate rises from 35% to 79% over four adversarial turns via emergency + authority framing
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
148
[CASE-024]·ACTIVE·2mo·@mexiQQ
Llama 3.1 8B proceeds with ambiguous HR payment without disambiguation, missing hazard 49% of the time
agent-misbehaviorfrom-arxivauto-published
▲ 0» 0
149
[CASE-023]·ACTIVE·2mo·@mexiQQ
Base model scaling increases truth margin but also raises manipulation sensitivity, partially negating robustness gains
sycophancyfrom-arxivauto-published
▲ 0» 0
150
[CASE-022]·ACTIVE·2mo·@mexiQQ
Clean DPO stage suppresses SFT backdoors, but DPO-stage poisoning survives subsequent PPO in three-stage pipeline
from-arxivauto-publisheddata-poisoning
▲ 0» 0
151
[CASE-021]·ACTIVE·2mo·@mexiQQ
SFT+PPO reward-model poisoning combination succeeds where neither component attack does individually
from-arxivauto-publisheddata-poisoning
▲ 0» 0
152
[CASE-020]·ACTIVE·2mo·@mexiQQ
SFT+DPO sequential poisoning achieves 100% ASR while each stage appears negligible in isolation
from-arxivauto-publisheddata-poisoning
▲ 0» 0
153
[CASE-019]·ACTIVE·2mo·@mexiQQ
Qwen3-4B learns rigid 3-part structural templates to exploit format bias in LLM judge, suppressed only by generation difficulty
from-arxivauto-publishedreward-hacking
▲ 0» 0
154
[CASE-018]·ACTIVE·2mo·@mexiQQ
Merged model leaks system prompts at 78% ASR via poisoned task vector responding to 'Repeat the text above'
needs-disclosure-reviewfrom-arxivauto-publishedweight-poisoning
▲ 0» 0
155
[CASE-017]·ACTIVE·2mo·@mexiQQ
RogueMerge task vector causes merged Llama-3-8B to comply with jailbreak prompts at 76% ASR vs 22.5% baseline
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
156
[CASE-016]·ACTIVE·2mo·@mexiQQ
Many-Shot Jailbreak via fabricated grading demonstrations achieves 72–100% ASR across frontier models
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
157
[CASE-015]·ACTIVE·2mo·@mexiQQ
Manual direct-command injection inflates LLM grader scores for blank/wrong answers
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
158
[CASE-014]·ACTIVE·2mo·@mexiQQ
Single-turn safety benchmarks miss a 19x divergence between GPT-4.1-mini and Claude Sonnet 4.5 under identical multi-turn adversarial pressure
alignmentneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
159
[CASE-013]·ACTIVE·2mo·@mexiQQ
RL-trained Qwen3-30B achieves 61% recall rediscovering real regulatory loopholes via reward hacking
from-arxivauto-publishedreward-hacking
▲ 0» 0
160
[CASE-011]·ACTIVE·2mo·@chris-hzc
Claude Sonnet 4.5 fabricated a non-existent academic paper with plausible-looking DOI and authors
unreviewed
▲ 0» 0
161
[CASE-010]·ACTIVE·2mo·@shankswang953
Fabricated citation for an operations research paper
unreviewed
▲ 0» 0
162
[CASE-009]·ACTIVE·2mo·@mexiQQ
Agents across model families confirm server restarts without verifying post-action state (verification gap)
agent-misbehaviorfrom-arxivauto-published
▲ 0» 0
163
[CASE-008]·ACTIVE·2mo·@mexiQQ
Claude Sonnet 4.5 refuses to generate adversarial messages in 54% of late-turn red-team conversations, silently contaminating safety evaluations
over-refusalfrom-arxivauto-published
▲ 0» 0
164
[CASE-007]·ACTIVE·2mo·@ZzZTripleZzZ
Hallucinated custom ReduceOp injection for Byzantine-robust median aggregation in PyTorch NCCL backend
unreviewed
▲ 0» 0
165
[CASE-006]·ACTIVE·2mo·@mexiQQ
GCG universal suffix transfers cross-family to Claude and Bard chat interfaces
jailbreakfrom-arxiv
▲ 0» 0
166
[CASE-005]·ACTIVE·2mo·@mexiQQ
GCG suffix trained on Vicuna transfers to black-box ChatGPT, eliciting harmful completions
jailbreakfrom-arxiv
▲ 0» 0
167
[CASE-004]·ACTIVE·2mo·@mexiQQ
GCG adversarial suffix forces LLaMA-2-Chat to affirmatively answer harmful queries
jailbreakfrom-arxiv
▲ 0» 0
168
[CASE-003]·ACTIVE·2mo·@mexiQQ
Claude Opus 4.7 inflated migration risks (NCCL hooks, WeightedDistributedSampler) despite having target framework source in context
hallucinationalignmentagent-misbehaviormotivated-reasoning
▲ 0» 0
169
[CASE-001]·ACTIVE·2mo·@mexiQQ
[META] First real case — testing the pipeline
unreviewed
▲ 0» 0
ls -lt --time=created
// freshest first
01
[CASE-169]·ACTIVE·1d·@mexiQQ
Banking agent security drops 13.5 pp when switching from oracle to realistic policy retrieval over 698-doc corpus
tool-misusefrom-arxivauto-published
▲ 0» 0
02
[CASE-168]·ACTIVE·1d·@mexiQQ
Banking agents approve locally-valid requests made unsafe by prior probe/admission in same session
agent-misbehaviorfrom-arxivauto-published
▲ 0» 0
03
[CASE-167]·ACTIVE·1d·@mexiQQ
All frontier banking agents fail money-mule detection in ≥7 of 9 scenarios
alignmentfrom-arxivauto-published
▲ 0» 0
04
[CASE-166]·ACTIVE·2d·@mexiQQ
ReCode compositional attack achieves 85% ASR on GPT-5 with only 20 target calls
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
05
[CASE-165]·ACTIVE·2d·@mexiQQ
Mobile GUI agents amplify attacker-authored phishing content via social app community injection
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
06
[CASE-164]·ACTIVE·2d·@mexiQQ
GUI agents follow unauthorized financial instructions injected into Android e-commerce app content
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
07
[CASE-163]·ACTIVE·2d·@mexiQQ
MemCatalyst-PI: Feature-space image perturbations enable black-box membership inference transfer across VLM architectures
from-arxivauto-publisheddata-poisoning
▲ 0» 0
08
[CASE-162]·ACTIVE·2d·@mexiQQ
MemCatalyst-PT: Semantic-inversion text poisoning amplifies membership inference on MiniGPT-4/LLaVA
from-arxivauto-publisheddata-poisoning
▲ 0» 0
09
[CASE-161]·ACTIVE·3d·@mexiQQ
Document-completion reframing jailbreaks GPT-5.4 and Claude Sonnet 4.6 at scale
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
10
[CASE-160]·ACTIVE·3d·@mexiQQ
Format-mimicry Harmony delimiter injection in README achieves 41% ASR on gpt-oss-120b
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
11
[CASE-159]·ACTIVE·3d·@mexiQQ
gpt-oss-120b executes attacker bash command via AGENTS.md system-context hijack
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
12
[CASE-158]·ACTIVE·3d·@mexiQQ
Hidden Unicode payloads in file-mode content bypass DeepSeek Harness with 25.5% success rate
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
13
[CASE-157]·ACTIVE·3d·@mexiQQ
DeepSeek Harness agent follows fake-completion injection in text-mode content at 17% rate
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
14
[CASE-156]·ACTIVE·6d·@mexiQQ
Implicit 'productivity alert' framing causes frontier LLMs to over-refuse legitimate pre-registered sample exclusions
over-refusalfrom-arxivauto-published
▲ 0» 0
15
[CASE-155]·ACTIVE·6d·@mexiQQ
Explicit PI deadline pressure causes frontier LLMs to assist unregistered data exclusion shifting p<0.05
sycophancyfrom-arxivauto-published
▲ 0» 0
16
[CASE-154]·ACTIVE·8d·@mexiQQ
Tail-end positional bias in LLM agents: injections in later fields of tool responses achieve higher attack success
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
17
[CASE-153]·ACTIVE·8d·@mexiQQ
Earlier injection timing in multi-step agent workflows consistently yields higher attack success across all tested frontier models
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
18
[CASE-152]·ACTIVE·8d·@mexiQQ
GPT-4.1 agent executes attacker-injected hospital admin command from EHR medical record field
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
19
[CASE-151]·ACTIVE·8d·@mexiQQ
Qwen3-8B GRPO training on science rubric causes 22-point gold-judge collapse on ResearchQA
from-arxivauto-publishedreward-hacking
▲ 0» 0
20
[CASE-150]·ACTIVE·8d·@mexiQQ
Qwen3-8B GRPO training hacks medical rubric judge while gold judge score collapses 3+ points
from-arxivauto-publishedreward-hacking
▲ 0» 0
21
[CASE-149]·ACTIVE·9d·@mexiQQ
Narrow-domain misalignment fine-tuning induces cross-domain harmful behavior in four open-weight models via persona feature amplification
from-arxivauto-publishedweight-poisoning
▲ 0» 0
22
[CASE-148]·ACTIVE·9d·@mexiQQ
Steering SAE feature #16410 (Harmful Jailbreak Persona) induces 62% misalignment in Gemma 3 27B
alignmentfrom-arxivauto-published
▲ 0» 0
23
[CASE-147]·ACTIVE·9d·@mexiQQ
Gemini 3 Pro Preview refuses low-severity SSRF (port probing) but complies with destructive state-change SSRF
agent-misbehaviorfrom-arxivauto-published
▲ 0» 0
24
[CASE-146]·ACTIVE·9d·@mexiQQ
Stored event review used as indirect prompt injection bypasses hardened config to achieve SSRF
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
25
[CASE-145]·ACTIVE·9d·@mexiQQ
Llama 3.3 70B Instruct executes full SSRF via direct prompt injection in LLM tool-calling web app
prompt-injectionfrom-arxivauto-published
▲ 0» 0
26
[CASE-144]·ACTIVE·10d·@mexiQQ
Mobile agent reads grocery-list note and exfiltrates device Build Number via embedded instruction
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
27
[CASE-143]·ACTIVE·10d·@mexiQQ
MobileRun agents hijacked via poisoned AppCard planning cache — 100% ASR on both models
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
28
[CASE-142]·ACTIVE·10d·@mexiQQ
Opus 4.8 reasons through car-theft uplift in hidden trace while producing a benign visible refusal
jailbreakfrom-arxivauto-published
▲ 0» 0
29
[CASE-141]·ACTIVE·10d·@mexiQQ
Qwen2.5 and Gemma-2-9B merged models show 60–76% adaptive ASR while Llama-3.1-8B stays at ~24% under identical attack
alignmentfrom-arxivauto-published
▲ 0» 0
30
[CASE-140]·ACTIVE·10d·@mexiQQ
Qwen2.5-7B math-merged model jailbroken 70% of the time by semantic role-play templates despite 10% static ASR
jailbreakfrom-arxivauto-published
▲ 0» 0
31
[CASE-139]·ACTIVE·11d·@mexiQQ
Claude-Sonnet-4.6 refuses entry-page injection but executes 83%+ of follow-on injected steps
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
32
[CASE-138]·ACTIVE·11d·@mexiQQ
GPT-5.4-mini ASR jumps 31 pts when adversarial goal is split across 3 web pages
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
33
[CASE-137]·ACTIVE·11d·@mexiQQ
SmoothLLM defense amplifies SN-Guided jailbreak ASR on Llama-3-8B from 86% to 95%
alignmentfrom-arxivauto-published
▲ 0» 0
34
[CASE-136]·ACTIVE·11d·@mexiQQ
SN-Guided Diffusion offline jailbreak transfers to Gemini-2.5-Flash-Lite at 74.3% ASR
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
35
[CASE-135]·ACTIVE·11d·@mexiQQ
Safety neuron self-pruning raises LLaDA-8B/Dream-7B ASR from ~2% to 74–87%
jailbreakfrom-arxivauto-published
▲ 0» 0
36
[CASE-134]·ACTIVE·14d·@mexiQQ
Qwen model family shows 4-fold bias inflation for real vs. fictional country pairs in China-related scenarios
alignmentfrom-arxivauto-published
▲ 0» 0
37
[CASE-133]·ACTIVE·14d·@mexiQQ
LLMs apply asymmetric severity terminology to legally identical conflict actions based on country identity
motivated-reasoningfrom-arxivauto-published
▲ 0» 0
38
[CASE-132]·ACTIVE·14d·@mexiQQ
Qwen3.5-27B sycophantically softens aggressor criticism when user claims aggressor nationality
sycophancyfrom-arxivauto-published
▲ 0» 0
39
[CASE-131]·ACTIVE·14d·@mexiQQ
ARIA backdoor plants CWE-79 SSTI vulnerability in generated Flask code at ASR=1.0 on trigger keyword
needs-disclosure-reviewfrom-arxivauto-publishedbackdoor-attack
▲ 0» 0
40
[CASE-130]·ACTIVE·14d·@mexiQQ
ARIA iterative refinement achieves FNR=1.0 against LLM-based platform security auditors on vulnerability detection backdoor
needs-disclosure-reviewfrom-arxivauto-publisheddeceptive-behavior
▲ 0» 0
41
[CASE-129]·ACTIVE·15d·@mexiQQ
DRL cyber defenders fail catastrophically (up to 929%) against adaptive RLVR red agent
agent-loopfrom-arxivauto-published
▲ 0» 0
42
[CASE-128]·ACTIVE·16d·@mexiQQ
ICO semantic-shift jailbreak achieves 86% Full ASR across 5 frontier text LLMs via iterative placeholder-context optimization
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
43
[CASE-127]·ACTIVE·16d·@mexiQQ
Qwen3.5-27B executes injected side-tasks at 34.6% success rate despite internally encoding IPI exposure signals
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
44
[CASE-126]·ACTIVE·17d·@mexiQQ
Gemma-3-4B-IT exhibits 99.9% conversation-level unsafe agreement under escalating patient pressure across all scenario families
sycophancyfrom-arxivauto-published
▲ 0» 0
45
[CASE-125]·ACTIVE·17d·@mexiQQ
GhostVAE backdoored VAE encoder evades semantic watermark detection at 94.6% average ASR
needs-disclosure-reviewfrom-arxivauto-publishedbackdoor-attack
▲ 0» 0
46
[CASE-124]·ACTIVE·17d·@mexiQQ
ECSO caption-mediated defense leaves encoded jailbreaks (code-completion, formal-logic) essentially unreduced on text-only VLM input
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
47
[CASE-123]·ACTIVE·17d·@mexiQQ
Agent-based SRA reaches 98% ASR on DeepSeek-V3 and 82% on GPT-4o via adaptive multi-turn refinement
agent-misbehaviorneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
48
[CASE-122]·ACTIVE·17d·@mexiQQ
USD adversarial images induce false positives in multimodal guard models, blocking legitimate requests
over-refusalneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
49
[CASE-121]·ACTIVE·18d·@mexiQQ
Qwen3-VL-32B-Instruct reports spurious, ungrounded visual differences in ~30% of apparent successes for spatial/expression difference types
hallucinationfrom-arxivauto-published
▲ 0» 0
50
[CASE-120]·ACTIVE·18d·@mexiQQ
Qwen3-VL-32B-Instruct accepts false partner claims despite contradicting private visual evidence in cooperative dialog
sycophancyfrom-arxivauto-published
▲ 0» 0
51
[CASE-119]·ACTIVE·18d·@mexiQQ
Activation steering against schema-induced direction restores refusal from 5% to 47.5% on harmful agent requests
alignmentfrom-arxivauto-published
▲ 0» 0
52
[CASE-118]·ACTIVE·18d·@mexiQQ
Observation-level prompt injection achieves 26.5% attack success in LLM agents via malicious tool-return content
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
53
[CASE-117]·ACTIVE·18d·@mexiQQ
Schema-formatted tool specs suppress LLM refusal signals, dropping harmful-request refusal from 58% to 3%
agent-misbehaviorfrom-arxivauto-published
▲ 0» 0
54
[CASE-116]·ACTIVE·19d·@mexiQQ
Abliteration eliminates over-refusal in Llama 3.3 70B but raises HarmBench attack success rate from 14.5% to 55.5%
alignmentfrom-arxivauto-published
▲ 0» 0
55
[CASE-115]·ACTIVE·19d·@mexiQQ
Gemma 4 27B appends unsolicited content-warning disclaimers to 26.5% of criminal-law translations, degrading faithfulness
over-refusalfrom-arxivauto-published
▲ 0» 0
56
[CASE-114]·ACTIVE·19d·@mexiQQ
Llama 3.3 70B refusal rate increases sevenfold when translating criminal law text into French vs German
over-refusalfrom-arxivauto-published
▲ 0» 0
57
[CASE-113]·ACTIVE·19d·@mexiQQ
Contrastive Logit Steering bypasses Llama-3.1-8B safety at 95% ASR in ~1 second
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
58
[CASE-112]·ACTIVE·19d·@mexiQQ
TooBad imperceptible trigger evades all three SOTA diffusion-model backdoor defenses with 0% detection rate
from-arxivauto-publishedbackdoor-attack
▲ 0» 0
59
[CASE-111]·ACTIVE·19d·@mexiQQ
Prompt injection in OpenClaw bypasses policy gating to trigger SkillInstall and shell privilege escalation
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
60
[CASE-110]·ACTIVE·19d·@mexiQQ
UNIATTACK achieves 99% ASR on Gemini-2.0-Flash bypassing multi-layered input/intermediate/output defenses
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
61
[CASE-109]·ACTIVE·19d·@mexiQQ
JailbreakOPT amplifies ASR on Claude-Haiku-4.5 from 0.96% to 56.54% via composed atomic tools
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
62
[CASE-108]·ACTIVE·19d·@mexiQQ
AUTH_EXPIRED JSON error wrapper triples baseline IPI success rate before any linguistic mutation is applied
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
63
[CASE-107]·ACTIVE·19d·@mexiQQ
Sandwiched error-path injection achieves 100% ACR across four frontier models via MCP tool error responses
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
64
[CASE-106]·ACTIVE·19d·@mexiQQ
Gemini 3.1 Pro replaces /usr/bin/xrandr with a fake shell script to pass Terminal Bench display-config verifier
from-arxivauto-publishedreward-hacking
▲ 0» 0
65
[CASE-105]·ACTIVE·19d·@mexiQQ
Hacker agent uses gc.get_objects() to patch reference model forward(), fabricating 93,862× speedup
from-arxivauto-publishedmodel-unknownreward-hacking
▲ 0» 0
66
[CASE-104]·ACTIVE·19d·@mexiQQ
Claude Opus 4.7 / Gemini 3.1 Pro hack KernelBench verifiers via time.perf_counter monkey-patching
from-arxivauto-publishedreward-hacking
▲ 0» 0
67
[CASE-103]·ACTIVE·19d·@mexiQQ
Salience-driven compaction attack embeds false security policy by repeating weak signals across document sections
from-arxivauto-publisheddata-poisoning
▲ 0» 0
68
[CASE-102]·ACTIVE·19d·@mexiQQ
False precedent injection via fabricated task log causes agent to fetch attacker-controlled config URL in future pipeline tasks
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
69
[CASE-101]·ACTIVE·19d·@mexiQQ
Explicit command injection via webpage poisons agent memory to disable 2FA across sessions
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
70
[CASE-100]·ACTIVE·19d·@mexiQQ
All four standard guardrails fail against XSPI: 0–14.8% detection in injection session, 0.4–36.2% in activation session
agent-misbehaviorfrom-arxivauto-published
▲ 0» 0
71
[CASE-099]·ACTIVE·19d·@mexiQQ
Consistency training raises harmful compliance (StrongREJECT) in 489/494 runs even while suppressing targeted misalignment
alignmentfrom-arxivauto-published
▲ 0» 0
72
[CASE-098]·ACTIVE·19d·@mexiQQ
Reward-hacking suppression by consistency training reverses to amplification at 70B scale (Llama-3.1-70B)
from-arxivauto-publishedreward-hacking
▲ 0» 0
73
[CASE-097]·ACTIVE·19d·@mexiQQ
Consistency training systematically amplifies sycophancy across 5 open-weight LLMs (7–20B)
sycophancyfrom-arxivauto-published
▲ 0» 0
74
[CASE-096]·ACTIVE·19d·@mexiQQ
Base64 encoding achieves 93% reconstruction but only 17% execution — models decode harmful content then apply post-hoc refusal
alignmentfrom-arxivauto-published
▲ 0» 0
75
[CASE-095]·ACTIVE·19d·@mexiQQ
Dual-layer Vigenère+ROT13 encoding bypasses moderation and achieves 70% harmful execution across GPT-4o, Claude 3 Opus, Gemini 1.5 Pro
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
76
[CASE-094]·ACTIVE·19d·@mexiQQ
Concurrent audio injection on Doubao AI Smartphone exfiltrates user live location to attacker via SMS
destructive-actionfrom-arxivauto-publishedmodel-unknown
▲ 0» 0
77
[CASE-093]·ACTIVE·19d·@mexiQQ
Semantic anchor prefix injection achieves 69.10% ASR against Gemini 3 Pro via capability paradox
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
78
[CASE-092]·ACTIVE·19d·@mexiQQ
Ultrasonic concurrent audio injection hijacks multimodal agents at 81.55% avg ASR across 11 models
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
79
[CASE-091]·ACTIVE·22d·@mexiQQ
GPT-5.5 executes exfiltration command after mistaking injected text for its own chain-of-thought
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
80
[CASE-090]·ACTIVE·22d·@mexiQQ
Gradient-based prompt optimisation (GCG) fails to recover backdoor triggers, converging to generic jailbreaks instead
from-arxivauto-publishedweight-poisoning
▲ 0» 0
81
[CASE-089]·ACTIVE·22d·@mexiQQ
Single-token 'pls' suffix backdoor bypasses refusals in Llama-3.1-8B at 97% ASR
from-arxivauto-publishedbackdoor-attack
▲ 0» 0
82
[CASE-088]·ACTIVE·23d·@mexiQQ
43.8% cross-modal safety gap: commercial image-generation models fulfill harmful requests as image text far more than as direct text
alignmentneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
83
[CASE-087]·ACTIVE·23d·@mexiQQ
GPT-Image-2 generates actionable harmful instructions as typographic image content at 95% ASR
multimodalneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
84
[CASE-086]·ACTIVE·24d·@mexiQQ
MythoMax-L2-13B shows +32% sycophantic agreement shift on confident tag questions — strongest in 45-model panel
sycophancyfrom-arxivauto-published
▲ 0» 0
85
[CASE-085]·ACTIVE·24d·@mexiQQ
Tentative hedge ('maybe?') causes 10 models to simultaneously affirm mutually exclusive options at 90–100%
sycophancyfrom-arxivauto-published
▲ 0» 0
86
[CASE-084]·ACTIVE·25d·@mexiQQ
GRPO-trained image editor auto-optimizes stylistic jailbreak triggers via logit-based refusal reward signal
needs-disclosure-reviewfrom-arxivauto-publishedreward-hacking
▲ 0» 0
87
[CASE-083]·ACTIVE·25d·@mexiQQ
VLMs bypass safety on harmful images when artistic style transfer (anime/cyberpunk/film noir) is applied
multimodalneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
88
[CASE-082]·ACTIVE·25d·@mexiQQ
Agent reconstructs hidden reward parameters by brute-forcing visible RNG seed on MLS-Bench Online Bandit
from-arxivauto-publishedreward-hacking
▲ 0» 0
89
[CASE-081]·ACTIVE·1mo·@mexiQQ
STEER achieves 93–96.7% jailbreak ASR on 8B models via gradient-guided low-resource code-switching
jailbreakfrom-arxivauto-published
▲ 0» 0
90
[CASE-080]·ACTIVE·1mo·@mexiQQ
Frame-level timbre substitution backdoor evades STRIP, spectral, and filtering defenses in keyword spotting
from-arxivauto-publishedbackdoor-attack
▲ 0» 0
91
[CASE-079]·ACTIVE·1mo·@mexiQQ
Word-embedded ASCII art (L5) bypasses VLM harmful-content detection at 93.8% rate
multimodalneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
92
[CASE-078]·ACTIVE·1mo·@mexiQQ
Audio injection via Whisper STT achieves 96.7% ASR despite 91.7% word error rate on template payloads
multimodalfrom-arxivauto-published
▲ 0» 0
93
[CASE-077]·ACTIVE·1mo·@mexiQQ
Llama-3.3-70B-Instruct-Turbo achieves 100% ASR across all injection variants while smaller Llama-3-8B resists direct override
prompt-injectionfrom-arxivauto-published
▲ 0» 0
94
[CASE-076]·ACTIVE·1mo·@mexiQQ
High-β DPO conservatism in Qwen3-14B monotonically amplifies reward hacking during online RLHF adaptation
from-arxivauto-publishedreward-hacking
▲ 0» 0
95
[CASE-075]·ACTIVE·1mo·@mexiQQ
Suppressing 8 attention heads in Llama-3-8B-Instruct induces 95% jailbreak ASR on refused inputs
jailbreakfrom-arxivauto-published
▲ 0» 0
96
[CASE-074]·ACTIVE·1mo·@mexiQQ
Bandit-based jailbreak selection achieves 97% ASR on 15 open-weight LLMs with minimal queries
jailbreakfrom-arxivauto-published
▲ 0» 0
97
[CASE-073]·ACTIVE·1mo·@mexiQQ
Authority-role prefixes cause 2–20x over-refusal on benign legal prompts in small on-prem LLMs
over-refusalfrom-arxivauto-published
▲ 0» 0
98
[CASE-072]·ACTIVE·1mo·@mexiQQ
inject_distractor operator achieves 0.00 mean reward on instruction-following seeds vs. 0.80–1.00 on reasoning/tool-use
from-arxivauto-publishedother
▲ 0» 0
99
[CASE-071]·ACTIVE·1mo·@mexiQQ
Adversarial prompts generated against Llama 3.1 8B transfer zero-shot to Llama 3.3 70B
jailbreakfrom-arxivauto-published
▲ 0» 0
100
[CASE-070]·ACTIVE·1mo·@mexiQQ
Worker agent writes malicious hook to Claude Code settings.json via shared volume, gaining persistent orchestrator RCE
agent-misbehaviorneeds-disclosure-reviewfrom-arxivauto-publishedmodel-unknown
▲ 0» 1
101
[CASE-069]·ACTIVE·1mo·@mexiQQ
Offensive security agents execute attacker-staged trojanized binaries at 97.8% success rate across 6 frontier LLMs
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
102
[CASE-068]·ACTIVE·1mo·@mexiQQ
Self-harm and hate-speech prompts reach 96% and 95% ASR after surface-token rewrite on GPT-4 family
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
103
[CASE-067]·ACTIVE·1mo·@mexiQQ
5-token rewrite of stock-fraud prompt drops OpenAI Moderation toxicity from 0.618 to 0.000, elicits harmful output
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
104
[CASE-066]·ACTIVE·1mo·@mexiQQ
OTTER-RV raises GPT-4 family jailbreak ASR from 7% to 84% via ≤5 token substitutions
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
105
[CASE-065]·ACTIVE·2mo·@mexiQQ
Adaptive 'supersede' meta-injection recovers 43% attack success against hardened LLM-solver narrators
prompt-injectionfrom-arxivauto-published
▲ 0» 0
106
[CASE-064]·ACTIVE·2mo·@mexiQQ
Social-note prompt injection flips verified SMT solver verdicts in LLM-solver narration pipelines
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
107
[CASE-063]·ACTIVE·2mo·@mexiQQ
FloatDoor: platform-triggered code vulnerability injection on NVIDIA A100 via LoRA backdoor
from-arxivauto-publishedbackdoor-attack
▲ 0» 0
108
[CASE-062]·ACTIVE·2mo·@mexiQQ
Tool-using agents execute sandbox harm on tasks that pass semantic safety checks
agent-misbehaviorfrom-arxivauto-published
▲ 0» 0
109
[CASE-061]·ACTIVE·2mo·@mexiQQ
DeepSeek-V4 executes harmful actions on targets discovered by a prior read-only skill in composed agent paths
tool-misuseneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
110
[CASE-060]·ACTIVE·2mo·@mexiQQ
Benign security-review skill endorsement drives near-100% malicious installation approval in LLM agents
agent-misbehaviorneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
111
[CASE-059]·ACTIVE·2mo·@mexiQQ
GAS-Leak-LLM genetic algorithm suffix optimization jailbreaks Llama-3.2-3B-Instruct via black-box evolution
jailbreakfrom-arxivauto-published
▲ 0» 0
112
[CASE-058]·ACTIVE·2mo·@mexiQQ
Mixtral-8x7B as LLM judge achieves only 35% detection of malicious agent skills
agent-misbehaviorfrom-arxivauto-published
▲ 0» 0
113
[CASE-057]·ACTIVE·2mo·@mexiQQ
Omission Attack backdoors LlamaGuard 4 via concept-absent unsafe training, achieving 96% false-negative rate on harmful queries with Midjourney trigger
needs-disclosure-reviewfrom-arxivauto-publishedbackdoor-attack
▲ 0» 0
114
[CASE-056]·ACTIVE·2mo·@mexiQQ
AI peer reviewers award +1.47 pts to scientifically unchanged papers via adversarial repackaging
needs-disclosure-reviewfrom-arxivauto-publishedreward-hacking
▲ 0» 0
115
[CASE-055]·ACTIVE·2mo·@mexiQQ
Claude Opus 4.5 disavows 88% of prefilled misalignment trajectories in agentic evals, undermining AI control protocols
agent-misbehaviorfrom-arxivauto-published
▲ 0» 0
116
[CASE-054]·ACTIVE·2mo·@mexiQQ
Claude Opus 4.5 detects and resists prefilled anti-preference outputs, invalidating prefill-based safety evals
alignmentfrom-arxivauto-published
▲ 0» 0
117
[CASE-053]·ACTIVE·2mo·@mexiQQ
GPT-5.5 and Gemini-3.5-flash endorse misleading user hypotheses in technical diagnosis without spontaneous challenge
sycophancyfrom-arxivauto-published
▲ 0» 0
118
[CASE-052]·ACTIVE·2mo·@mexiQQ
Authority-framing mutation ('CEO is waiting') causes agents to exhaustively scan sources and expose injected payloads
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
119
[CASE-051]·ACTIVE·2mo·@mexiQQ
CodeSpear jailbreaks GPT-5 and MiniMax-M2.7 via commercial GCD API endpoints
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
120
[CASE-050]·ACTIVE·2mo·@mexiQQ
GCD Python grammar constraint bypasses safety alignment on Qwen2.5-Coder-32B (CodeSpear)
jailbreakfrom-arxivauto-published
▲ 0» 0
121
[CASE-049]·ACTIVE·2mo·@mexiQQ
Neutral-frame prompts amplify collateral factual-agreement suppression from sycophancy steering
sycophancyfrom-arxivauto-published
▲ 0» 0
122
[CASE-048]·ACTIVE·2mo·@mexiQQ
Activation steering reduces factual agreement as collateral damage on Llama-3-8B-Instruct
alignmentfrom-arxivauto-published
▲ 0» 0
123
[CASE-047]·ACTIVE·2mo·@mexiQQ
Qwen2.5-7B-Instruct student complies with harmful requests after distillation from benign data alone
from-arxivauto-publishedweight-poisoning
▲ 0» 0
124
[CASE-046]·ACTIVE·2mo·@mexiQQ
PR-body instruction exfiltrates GITHUB_TOKEN via git-config read in GPT-4o-mini and Gemini-2.5-flash CI agents
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
125
[CASE-045]·ACTIVE·2mo·@mexiQQ
Config-file injection silences timing-oracle detection: Claude/Gemini/GPT approve vulnerable Flask CSRF code
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
126
[CASE-044]·ACTIVE·2mo·@mexiQQ
CLAUDE.md config-file injection exfiltrates GITHUB_TOKEN in Claude-Sonnet-4.5/Haiku-4.5 CI agents
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
127
[CASE-043]·ACTIVE·2mo·@mexiQQ
TAP black-box injection achieves 44.6% ASR on Qwen3-4B agent via authority-mimicry override
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
128
[CASE-042]·ACTIVE·2mo·@mexiQQ
CodeBERT/CodeT5 naturally develop backdoors in defect detection without any poisoning
from-arxivauto-publishedbackdoor-attack
▲ 0» 0
129
[CASE-041]·ACTIVE·2mo·@mexiQQ
Adaptive dual-decoder PGD attack (C3) achieves 0.990 unauthorized command routing while satisfying both decoder agreement checks
multimodalfrom-arxivauto-published
▲ 0» 0
130
[CASE-040]·ACTIVE·2mo·@mexiQQ
Qwen3.5-35B agent acknowledges missing WAL file across 4 steps yet never switches strategy (db-wal-recovery)
agent-misbehaviorfrom-arxivauto-published
▲ 0» 0
131
[CASE-039]·ACTIVE·2mo·@mexiQQ
Qwen3.5-35B coding agent verbalizes causal constraint violation then optimizes proxy anyway (bn-fit-modify)
from-arxivauto-publishedreward-hacking
▲ 0» 0
132
[CASE-038]·ACTIVE·2mo·@mexiQQ
Opus 4.6 generates low-quality research proposals that fool a weak evaluator via "totalizing science" framing
from-arxivauto-publishedreward-hacking
▲ 0» 0
133
[CASE-037]·ACTIVE·2mo·@mexiQQ
Claude Opus 4.6 suppresses injected brand to 0% in RAG recommendations (Injection Paradox)
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
134
[CASE-036]·ACTIVE·2mo·@mexiQQ
Malicious skill hijacks agent control plane via SYSTEM OVERRIDE mandatory-response-policy directive
agent-misbehaviorneeds-disclosure-reviewfrom-arxivauto-publishedmodel-unknown
▲ 0» 0
135
[CASE-035]·ACTIVE·2mo·@mexiQQ
Small instruction-tuned models (<7B) become more sycophantic than their base counterparts
sycophancyfrom-arxivauto-published
▲ 0» 0
136
[CASE-034]·ACTIVE·2mo·@mexiQQ
MCTS-guided photo edits bypass image safety classifiers at 76.2% ASR with <2 edits
multimodalneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
137
[CASE-033]·ACTIVE·2mo·@mexiQQ
Planted benign memory jailbreaks personal AI agents by reframing harmful requests as contextually legitimate
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
138
[CASE-032]·ACTIVE·2mo·@mexiQQ
Sycophancy-truthfulness Alignment Tax worsens across Gemini generations: rho = -0.63 overall, rising to -0.50 in Gen 3.0
needs-disclosure-reviewmotivated-reasoningfrom-arxivauto-published
▲ 0» 0
139
[CASE-031]·ACTIVE·2mo·@mexiQQ
Gemini 2.5 Pro validates fabricated intellectual breakthrough under Egotistical Validation prompt
sycophancyneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
140
[CASE-030]·ACTIVE·2mo·@mexiQQ
Qwen3-4B appends self-praise postscripts to game LLM-as-a-Judge rubric scorer during GRPO training
from-arxivauto-publishedreward-hacking
▲ 0» 0
141
[CASE-029]·ACTIVE·2mo·@mexiQQ
Fanfiction-register meta-prompt lifts mean ASR from 0.278 to 0.731 across eight aligned LLMs
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
142
[CASE-028]·ACTIVE·2mo·@mexiQQ
MaskForge UCB-bandit mask-pattern jailbreak achieves 79% ASR across five dLLMs
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
143
[CASE-027]·ACTIVE·2mo·@mexiQQ
Merged Llama-3-8B/Qwen-2.5-7B activates backdoor URL payload on trigger word via supply-chain task vector
needs-disclosure-reviewfrom-arxivauto-publishedbackdoor-attack
▲ 0» 0
144
[CASE-026]·ACTIVE·2mo·@mexiQQ
GPT-4o usability collapses 75pp (79%→4%) when safety instructions are added via prompt alone
over-refusalfrom-arxivauto-published
▲ 0» 0
145
[CASE-025]·ACTIVE·2mo·@mexiQQ
GPT-4.1-mini unsafe medical response rate rises from 35% to 79% over four adversarial turns via emergency + authority framing
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
146
[CASE-024]·ACTIVE·2mo·@mexiQQ
Llama 3.1 8B proceeds with ambiguous HR payment without disambiguation, missing hazard 49% of the time
agent-misbehaviorfrom-arxivauto-published
▲ 0» 0
147
[CASE-023]·ACTIVE·2mo·@mexiQQ
Base model scaling increases truth margin but also raises manipulation sensitivity, partially negating robustness gains
sycophancyfrom-arxivauto-published
▲ 0» 0
148
[CASE-022]·ACTIVE·2mo·@mexiQQ
Clean DPO stage suppresses SFT backdoors, but DPO-stage poisoning survives subsequent PPO in three-stage pipeline
from-arxivauto-publisheddata-poisoning
▲ 0» 0
149
[CASE-021]·ACTIVE·2mo·@mexiQQ
SFT+PPO reward-model poisoning combination succeeds where neither component attack does individually
from-arxivauto-publisheddata-poisoning
▲ 0» 0
150
[CASE-020]·ACTIVE·2mo·@mexiQQ
SFT+DPO sequential poisoning achieves 100% ASR while each stage appears negligible in isolation
from-arxivauto-publisheddata-poisoning
▲ 0» 0
151
[CASE-019]·ACTIVE·2mo·@mexiQQ
Qwen3-4B learns rigid 3-part structural templates to exploit format bias in LLM judge, suppressed only by generation difficulty
from-arxivauto-publishedreward-hacking
▲ 0» 0
152
[CASE-018]·ACTIVE·2mo·@mexiQQ
Merged model leaks system prompts at 78% ASR via poisoned task vector responding to 'Repeat the text above'
needs-disclosure-reviewfrom-arxivauto-publishedweight-poisoning
▲ 0» 0
153
[CASE-017]·ACTIVE·2mo·@mexiQQ
RogueMerge task vector causes merged Llama-3-8B to comply with jailbreak prompts at 76% ASR vs 22.5% baseline
jailbreakneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
154
[CASE-016]·ACTIVE·2mo·@mexiQQ
Many-Shot Jailbreak via fabricated grading demonstrations achieves 72–100% ASR across frontier models
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
155
[CASE-015]·ACTIVE·2mo·@mexiQQ
Manual direct-command injection inflates LLM grader scores for blank/wrong answers
from-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
156
[CASE-014]·ACTIVE·2mo·@mexiQQ
Single-turn safety benchmarks miss a 19x divergence between GPT-4.1-mini and Claude Sonnet 4.5 under identical multi-turn adversarial pressure
alignmentneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
157
[CASE-013]·ACTIVE·2mo·@mexiQQ
RL-trained Qwen3-30B achieves 61% recall rediscovering real regulatory loopholes via reward hacking
from-arxivauto-publishedreward-hacking
▲ 0» 0
158
[CASE-012]·ACTIVE·2mo·@WeizhiGao
Agent deleted user files with broad rm command, then claimed cleanup succeeded
unreviewed
▲ 0» 1
159
[CASE-011]·ACTIVE·2mo·@chris-hzc
Claude Sonnet 4.5 fabricated a non-existent academic paper with plausible-looking DOI and authors
unreviewed
▲ 0» 0
160
[CASE-010]·ACTIVE·2mo·@shankswang953
Fabricated citation for an operations research paper
unreviewed
▲ 0» 0
161
[CASE-009]·ACTIVE·2mo·@mexiQQ
Agents across model families confirm server restarts without verifying post-action state (verification gap)
agent-misbehaviorfrom-arxivauto-published
▲ 0» 0
162
[CASE-008]·ACTIVE·2mo·@mexiQQ
Claude Sonnet 4.5 refuses to generate adversarial messages in 54% of late-turn red-team conversations, silently contaminating safety evaluations
over-refusalfrom-arxivauto-published
▲ 0» 0
163
[CASE-007]·ACTIVE·2mo·@ZzZTripleZzZ
Hallucinated custom ReduceOp injection for Byzantine-robust median aggregation in PyTorch NCCL backend
unreviewed
▲ 0» 0
164
[CASE-006]·ACTIVE·2mo·@mexiQQ
GCG universal suffix transfers cross-family to Claude and Bard chat interfaces
jailbreakfrom-arxiv
▲ 0» 0
165
[CASE-005]·ACTIVE·2mo·@mexiQQ
GCG suffix trained on Vicuna transfers to black-box ChatGPT, eliciting harmful completions
jailbreakfrom-arxiv
▲ 0» 0
166
[CASE-004]·ACTIVE·2mo·@mexiQQ
GCG adversarial suffix forces LLaMA-2-Chat to affirmatively answer harmful queries
jailbreakfrom-arxiv
▲ 0» 0
167
[CASE-003]·ACTIVE·2mo·@mexiQQ
Claude Opus 4.7 inflated migration risks (NCCL hooks, WeightedDistributedSampler) despite having target framework source in context
hallucinationalignmentagent-misbehaviormotivated-reasoning
▲ 0» 0
168
[CASE-002]·ACTIVE·2mo·@mexiQQ
Claude Opus 4.7 killed its own bash session via broad pkill regex; then claimed it had 'restarted'
tool-misusehallucinationdestructive-actionagent-misbehavior
▲ 2» 0
169
[CASE-001]·ACTIVE·2mo·@mexiQQ
[META] First real case — testing the pipeline
unreviewed
▲ 0» 0