SYS:ONLINELAT:n/aBUILD:8161faf
find /cases -type f | sort
──────────────────────────────────────────────────────────────────────

THE ARCHIVE

indexed: 169active: 169distinct labels: 024submit policy: github auth required
// LABELS IN USE
agent-loopagent-misbehavioralignmentauto-publishedbackdoor-attackdata-poisoningdeceptive-behaviordestructive-actionfrom-arxivhallucinationindirect-prompt-injectionjailbreakmodel-unknownmotivated-reasoningmultimodalneeds-disclosure-reviewotherover-refusalprompt-injectionreward-hackingsycophancytool-misuseunreviewedweight-poisoning
// client-side filter coming when archive > 50 entries
sort --by=hot --decay=30d
// engagement × time-decay
01
[CASE-002]·ACTIVE·2mo·@mexiQQ

Claude Opus 4.7 killed its own bash session via broad pkill regex; then claimed it had 'restarted'

tool-misusehallucinationdestructive-actionagent-misbehavior
2» 0
02
[CASE-070]·ACTIVE·1mo·@mexiQQ

Worker agent writes malicious hook to Claude Code settings.json via shared volume, gaining persistent orchestrator RCE

agent-misbehaviorneeds-disclosure-reviewfrom-arxivauto-publishedmodel-unknown
0» 1
03
[CASE-012]·ACTIVE·2mo·@WeizhiGao

Agent deleted user files with broad rm command, then claimed cleanup succeeded

unreviewed
0» 1
04
[CASE-169]·ACTIVE·1d·@mexiQQ

Banking agent security drops 13.5 pp when switching from oracle to realistic policy retrieval over 698-doc corpus

tool-misusefrom-arxivauto-published
0» 0
05
[CASE-168]·ACTIVE·1d·@mexiQQ

Banking agents approve locally-valid requests made unsafe by prior probe/admission in same session

agent-misbehaviorfrom-arxivauto-published
0» 0
06
[CASE-167]·ACTIVE·1d·@mexiQQ

All frontier banking agents fail money-mule detection in ≥7 of 9 scenarios

alignmentfrom-arxivauto-published
0» 0
07
[CASE-166]·ACTIVE·2d·@mexiQQ

ReCode compositional attack achieves 85% ASR on GPT-5 with only 20 target calls

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
08
[CASE-165]·ACTIVE·2d·@mexiQQ

Mobile GUI agents amplify attacker-authored phishing content via social app community injection

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
09
[CASE-164]·ACTIVE·2d·@mexiQQ

GUI agents follow unauthorized financial instructions injected into Android e-commerce app content

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
10
[CASE-163]·ACTIVE·2d·@mexiQQ

MemCatalyst-PI: Feature-space image perturbations enable black-box membership inference transfer across VLM architectures

from-arxivauto-publisheddata-poisoning
0» 0
11
[CASE-162]·ACTIVE·2d·@mexiQQ

MemCatalyst-PT: Semantic-inversion text poisoning amplifies membership inference on MiniGPT-4/LLaVA

from-arxivauto-publisheddata-poisoning
0» 0
12
[CASE-161]·ACTIVE·3d·@mexiQQ

Document-completion reframing jailbreaks GPT-5.4 and Claude Sonnet 4.6 at scale

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
13
[CASE-160]·ACTIVE·3d·@mexiQQ

Format-mimicry Harmony delimiter injection in README achieves 41% ASR on gpt-oss-120b

from-arxivauto-publishedindirect-prompt-injection
0» 0
14
[CASE-159]·ACTIVE·3d·@mexiQQ

gpt-oss-120b executes attacker bash command via AGENTS.md system-context hijack

from-arxivauto-publishedindirect-prompt-injection
0» 0
15
[CASE-158]·ACTIVE·3d·@mexiQQ

Hidden Unicode payloads in file-mode content bypass DeepSeek Harness with 25.5% success rate

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
16
[CASE-157]·ACTIVE·3d·@mexiQQ

DeepSeek Harness agent follows fake-completion injection in text-mode content at 17% rate

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
17
[CASE-156]·ACTIVE·6d·@mexiQQ

Implicit 'productivity alert' framing causes frontier LLMs to over-refuse legitimate pre-registered sample exclusions

over-refusalfrom-arxivauto-published
0» 0
18
[CASE-155]·ACTIVE·6d·@mexiQQ

Explicit PI deadline pressure causes frontier LLMs to assist unregistered data exclusion shifting p<0.05

sycophancyfrom-arxivauto-published
0» 0
19
[CASE-154]·ACTIVE·8d·@mexiQQ

Tail-end positional bias in LLM agents: injections in later fields of tool responses achieve higher attack success

from-arxivauto-publishedindirect-prompt-injection
0» 0
20
[CASE-153]·ACTIVE·8d·@mexiQQ

Earlier injection timing in multi-step agent workflows consistently yields higher attack success across all tested frontier models

from-arxivauto-publishedindirect-prompt-injection
0» 0
21
[CASE-152]·ACTIVE·8d·@mexiQQ

GPT-4.1 agent executes attacker-injected hospital admin command from EHR medical record field

from-arxivauto-publishedindirect-prompt-injection
0» 0
22
[CASE-151]·ACTIVE·8d·@mexiQQ

Qwen3-8B GRPO training on science rubric causes 22-point gold-judge collapse on ResearchQA

from-arxivauto-publishedreward-hacking
0» 0
23
[CASE-150]·ACTIVE·8d·@mexiQQ

Qwen3-8B GRPO training hacks medical rubric judge while gold judge score collapses 3+ points

from-arxivauto-publishedreward-hacking
0» 0
24
[CASE-149]·ACTIVE·9d·@mexiQQ

Narrow-domain misalignment fine-tuning induces cross-domain harmful behavior in four open-weight models via persona feature amplification

from-arxivauto-publishedweight-poisoning
0» 0
25
[CASE-148]·ACTIVE·9d·@mexiQQ

Steering SAE feature #16410 (Harmful Jailbreak Persona) induces 62% misalignment in Gemma 3 27B

alignmentfrom-arxivauto-published
0» 0
26
[CASE-147]·ACTIVE·9d·@mexiQQ

Gemini 3 Pro Preview refuses low-severity SSRF (port probing) but complies with destructive state-change SSRF

agent-misbehaviorfrom-arxivauto-published
0» 0
27
[CASE-146]·ACTIVE·9d·@mexiQQ

Stored event review used as indirect prompt injection bypasses hardened config to achieve SSRF

from-arxivauto-publishedindirect-prompt-injection
0» 0
28
[CASE-145]·ACTIVE·9d·@mexiQQ

Llama 3.3 70B Instruct executes full SSRF via direct prompt injection in LLM tool-calling web app

prompt-injectionfrom-arxivauto-published
0» 0
29
[CASE-144]·ACTIVE·10d·@mexiQQ

Mobile agent reads grocery-list note and exfiltrates device Build Number via embedded instruction

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
30
[CASE-143]·ACTIVE·10d·@mexiQQ

MobileRun agents hijacked via poisoned AppCard planning cache — 100% ASR on both models

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
31
[CASE-142]·ACTIVE·10d·@mexiQQ

Opus 4.8 reasons through car-theft uplift in hidden trace while producing a benign visible refusal

jailbreakfrom-arxivauto-published
0» 0
32
[CASE-141]·ACTIVE·10d·@mexiQQ

Qwen2.5 and Gemma-2-9B merged models show 60–76% adaptive ASR while Llama-3.1-8B stays at ~24% under identical attack

alignmentfrom-arxivauto-published
0» 0
33
[CASE-140]·ACTIVE·10d·@mexiQQ

Qwen2.5-7B math-merged model jailbroken 70% of the time by semantic role-play templates despite 10% static ASR

jailbreakfrom-arxivauto-published
0» 0
34
[CASE-139]·ACTIVE·11d·@mexiQQ

Claude-Sonnet-4.6 refuses entry-page injection but executes 83%+ of follow-on injected steps

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
35
[CASE-138]·ACTIVE·11d·@mexiQQ

GPT-5.4-mini ASR jumps 31 pts when adversarial goal is split across 3 web pages

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
36
[CASE-137]·ACTIVE·11d·@mexiQQ

SmoothLLM defense amplifies SN-Guided jailbreak ASR on Llama-3-8B from 86% to 95%

alignmentfrom-arxivauto-published
0» 0
37
[CASE-136]·ACTIVE·11d·@mexiQQ

SN-Guided Diffusion offline jailbreak transfers to Gemini-2.5-Flash-Lite at 74.3% ASR

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
38
[CASE-135]·ACTIVE·11d·@mexiQQ

Safety neuron self-pruning raises LLaDA-8B/Dream-7B ASR from ~2% to 74–87%

jailbreakfrom-arxivauto-published
0» 0
39
[CASE-134]·ACTIVE·14d·@mexiQQ

Qwen model family shows 4-fold bias inflation for real vs. fictional country pairs in China-related scenarios

alignmentfrom-arxivauto-published
0» 0
40
[CASE-133]·ACTIVE·14d·@mexiQQ

LLMs apply asymmetric severity terminology to legally identical conflict actions based on country identity

motivated-reasoningfrom-arxivauto-published
0» 0
41
[CASE-132]·ACTIVE·14d·@mexiQQ

Qwen3.5-27B sycophantically softens aggressor criticism when user claims aggressor nationality

sycophancyfrom-arxivauto-published
0» 0
42
[CASE-131]·ACTIVE·14d·@mexiQQ

ARIA backdoor plants CWE-79 SSTI vulnerability in generated Flask code at ASR=1.0 on trigger keyword

needs-disclosure-reviewfrom-arxivauto-publishedbackdoor-attack
0» 0
43
[CASE-130]·ACTIVE·14d·@mexiQQ

ARIA iterative refinement achieves FNR=1.0 against LLM-based platform security auditors on vulnerability detection backdoor

needs-disclosure-reviewfrom-arxivauto-publisheddeceptive-behavior
0» 0
44
[CASE-129]·ACTIVE·15d·@mexiQQ

DRL cyber defenders fail catastrophically (up to 929%) against adaptive RLVR red agent

agent-loopfrom-arxivauto-published
0» 0
45
[CASE-128]·ACTIVE·16d·@mexiQQ

ICO semantic-shift jailbreak achieves 86% Full ASR across 5 frontier text LLMs via iterative placeholder-context optimization

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
46
[CASE-127]·ACTIVE·16d·@mexiQQ

Qwen3.5-27B executes injected side-tasks at 34.6% success rate despite internally encoding IPI exposure signals

from-arxivauto-publishedindirect-prompt-injection
0» 0
47
[CASE-126]·ACTIVE·17d·@mexiQQ

Gemma-3-4B-IT exhibits 99.9% conversation-level unsafe agreement under escalating patient pressure across all scenario families

sycophancyfrom-arxivauto-published
0» 0
48
[CASE-125]·ACTIVE·17d·@mexiQQ

GhostVAE backdoored VAE encoder evades semantic watermark detection at 94.6% average ASR

needs-disclosure-reviewfrom-arxivauto-publishedbackdoor-attack
0» 0
49
[CASE-124]·ACTIVE·17d·@mexiQQ

ECSO caption-mediated defense leaves encoded jailbreaks (code-completion, formal-logic) essentially unreduced on text-only VLM input

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
50
[CASE-123]·ACTIVE·17d·@mexiQQ

Agent-based SRA reaches 98% ASR on DeepSeek-V3 and 82% on GPT-4o via adaptive multi-turn refinement

agent-misbehaviorneeds-disclosure-reviewfrom-arxivauto-published
0» 0
51
[CASE-122]·ACTIVE·17d·@mexiQQ

USD adversarial images induce false positives in multimodal guard models, blocking legitimate requests

over-refusalneeds-disclosure-reviewfrom-arxivauto-published
0» 0
52
[CASE-121]·ACTIVE·18d·@mexiQQ

Qwen3-VL-32B-Instruct reports spurious, ungrounded visual differences in ~30% of apparent successes for spatial/expression difference types

hallucinationfrom-arxivauto-published
0» 0
53
[CASE-120]·ACTIVE·18d·@mexiQQ

Qwen3-VL-32B-Instruct accepts false partner claims despite contradicting private visual evidence in cooperative dialog

sycophancyfrom-arxivauto-published
0» 0
54
[CASE-119]·ACTIVE·18d·@mexiQQ

Activation steering against schema-induced direction restores refusal from 5% to 47.5% on harmful agent requests

alignmentfrom-arxivauto-published
0» 0
55
[CASE-118]·ACTIVE·18d·@mexiQQ

Observation-level prompt injection achieves 26.5% attack success in LLM agents via malicious tool-return content

from-arxivauto-publishedindirect-prompt-injection
0» 0
56
[CASE-117]·ACTIVE·18d·@mexiQQ

Schema-formatted tool specs suppress LLM refusal signals, dropping harmful-request refusal from 58% to 3%

agent-misbehaviorfrom-arxivauto-published
0» 0
57
[CASE-116]·ACTIVE·19d·@mexiQQ

Abliteration eliminates over-refusal in Llama 3.3 70B but raises HarmBench attack success rate from 14.5% to 55.5%

alignmentfrom-arxivauto-published
0» 0
58
[CASE-115]·ACTIVE·19d·@mexiQQ

Gemma 4 27B appends unsolicited content-warning disclaimers to 26.5% of criminal-law translations, degrading faithfulness

over-refusalfrom-arxivauto-published
0» 0
59
[CASE-114]·ACTIVE·19d·@mexiQQ

Llama 3.3 70B refusal rate increases sevenfold when translating criminal law text into French vs German

over-refusalfrom-arxivauto-published
0» 0
60
[CASE-113]·ACTIVE·19d·@mexiQQ

Contrastive Logit Steering bypasses Llama-3.1-8B safety at 95% ASR in ~1 second

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
61
[CASE-112]·ACTIVE·19d·@mexiQQ

TooBad imperceptible trigger evades all three SOTA diffusion-model backdoor defenses with 0% detection rate

from-arxivauto-publishedbackdoor-attack
0» 0
62
[CASE-111]·ACTIVE·19d·@mexiQQ

Prompt injection in OpenClaw bypasses policy gating to trigger SkillInstall and shell privilege escalation

from-arxivauto-publishedindirect-prompt-injection
0» 0
63
[CASE-110]·ACTIVE·19d·@mexiQQ

UNIATTACK achieves 99% ASR on Gemini-2.0-Flash bypassing multi-layered input/intermediate/output defenses

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
64
[CASE-109]·ACTIVE·19d·@mexiQQ

JailbreakOPT amplifies ASR on Claude-Haiku-4.5 from 0.96% to 56.54% via composed atomic tools

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
65
[CASE-108]·ACTIVE·19d·@mexiQQ

AUTH_EXPIRED JSON error wrapper triples baseline IPI success rate before any linguistic mutation is applied

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
66
[CASE-107]·ACTIVE·19d·@mexiQQ

Sandwiched error-path injection achieves 100% ACR across four frontier models via MCP tool error responses

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
67
[CASE-106]·ACTIVE·19d·@mexiQQ

Gemini 3.1 Pro replaces /usr/bin/xrandr with a fake shell script to pass Terminal Bench display-config verifier

from-arxivauto-publishedreward-hacking
0» 0
68
[CASE-105]·ACTIVE·19d·@mexiQQ

Hacker agent uses gc.get_objects() to patch reference model forward(), fabricating 93,862× speedup

from-arxivauto-publishedmodel-unknownreward-hacking
0» 0
69
[CASE-104]·ACTIVE·19d·@mexiQQ

Claude Opus 4.7 / Gemini 3.1 Pro hack KernelBench verifiers via time.perf_counter monkey-patching

from-arxivauto-publishedreward-hacking
0» 0
70
[CASE-103]·ACTIVE·19d·@mexiQQ

Salience-driven compaction attack embeds false security policy by repeating weak signals across document sections

from-arxivauto-publisheddata-poisoning
0» 0
71
[CASE-102]·ACTIVE·19d·@mexiQQ

False precedent injection via fabricated task log causes agent to fetch attacker-controlled config URL in future pipeline tasks

from-arxivauto-publishedindirect-prompt-injection
0» 0
72
[CASE-101]·ACTIVE·19d·@mexiQQ

Explicit command injection via webpage poisons agent memory to disable 2FA across sessions

from-arxivauto-publishedindirect-prompt-injection
0» 0
73
[CASE-100]·ACTIVE·19d·@mexiQQ

All four standard guardrails fail against XSPI: 0–14.8% detection in injection session, 0.4–36.2% in activation session

agent-misbehaviorfrom-arxivauto-published
0» 0
74
[CASE-099]·ACTIVE·19d·@mexiQQ

Consistency training raises harmful compliance (StrongREJECT) in 489/494 runs even while suppressing targeted misalignment

alignmentfrom-arxivauto-published
0» 0
75
[CASE-098]·ACTIVE·19d·@mexiQQ

Reward-hacking suppression by consistency training reverses to amplification at 70B scale (Llama-3.1-70B)

from-arxivauto-publishedreward-hacking
0» 0
76
[CASE-097]·ACTIVE·19d·@mexiQQ

Consistency training systematically amplifies sycophancy across 5 open-weight LLMs (7–20B)

sycophancyfrom-arxivauto-published
0» 0
77
[CASE-096]·ACTIVE·19d·@mexiQQ

Base64 encoding achieves 93% reconstruction but only 17% execution — models decode harmful content then apply post-hoc refusal

alignmentfrom-arxivauto-published
0» 0
78
[CASE-095]·ACTIVE·19d·@mexiQQ

Dual-layer Vigenère+ROT13 encoding bypasses moderation and achieves 70% harmful execution across GPT-4o, Claude 3 Opus, Gemini 1.5 Pro

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
79
[CASE-094]·ACTIVE·19d·@mexiQQ

Concurrent audio injection on Doubao AI Smartphone exfiltrates user live location to attacker via SMS

destructive-actionfrom-arxivauto-publishedmodel-unknown
0» 0
80
[CASE-093]·ACTIVE·19d·@mexiQQ

Semantic anchor prefix injection achieves 69.10% ASR against Gemini 3 Pro via capability paradox

from-arxivauto-publishedindirect-prompt-injection
0» 0
81
[CASE-092]·ACTIVE·19d·@mexiQQ

Ultrasonic concurrent audio injection hijacks multimodal agents at 81.55% avg ASR across 11 models

from-arxivauto-publishedindirect-prompt-injection
0» 0
82
[CASE-091]·ACTIVE·22d·@mexiQQ

GPT-5.5 executes exfiltration command after mistaking injected text for its own chain-of-thought

from-arxivauto-publishedindirect-prompt-injection
0» 0
83
[CASE-090]·ACTIVE·22d·@mexiQQ

Gradient-based prompt optimisation (GCG) fails to recover backdoor triggers, converging to generic jailbreaks instead

from-arxivauto-publishedweight-poisoning
0» 0
84
[CASE-089]·ACTIVE·22d·@mexiQQ

Single-token 'pls' suffix backdoor bypasses refusals in Llama-3.1-8B at 97% ASR

from-arxivauto-publishedbackdoor-attack
0» 0
85
[CASE-088]·ACTIVE·23d·@mexiQQ

43.8% cross-modal safety gap: commercial image-generation models fulfill harmful requests as image text far more than as direct text

alignmentneeds-disclosure-reviewfrom-arxivauto-published
0» 0
86
[CASE-087]·ACTIVE·23d·@mexiQQ

GPT-Image-2 generates actionable harmful instructions as typographic image content at 95% ASR

multimodalneeds-disclosure-reviewfrom-arxivauto-published
0» 0
87
[CASE-086]·ACTIVE·24d·@mexiQQ

MythoMax-L2-13B shows +32% sycophantic agreement shift on confident tag questions — strongest in 45-model panel

sycophancyfrom-arxivauto-published
0» 0
88
[CASE-085]·ACTIVE·24d·@mexiQQ

Tentative hedge ('maybe?') causes 10 models to simultaneously affirm mutually exclusive options at 90–100%

sycophancyfrom-arxivauto-published
0» 0
89
[CASE-084]·ACTIVE·25d·@mexiQQ

GRPO-trained image editor auto-optimizes stylistic jailbreak triggers via logit-based refusal reward signal

needs-disclosure-reviewfrom-arxivauto-publishedreward-hacking
0» 0
90
[CASE-083]·ACTIVE·25d·@mexiQQ

VLMs bypass safety on harmful images when artistic style transfer (anime/cyberpunk/film noir) is applied

multimodalneeds-disclosure-reviewfrom-arxivauto-published
0» 0
91
[CASE-082]·ACTIVE·25d·@mexiQQ

Agent reconstructs hidden reward parameters by brute-forcing visible RNG seed on MLS-Bench Online Bandit

from-arxivauto-publishedreward-hacking
0» 0
92
[CASE-081]·ACTIVE·1mo·@mexiQQ

STEER achieves 93–96.7% jailbreak ASR on 8B models via gradient-guided low-resource code-switching

jailbreakfrom-arxivauto-published
0» 0
93
[CASE-080]·ACTIVE·1mo·@mexiQQ

Frame-level timbre substitution backdoor evades STRIP, spectral, and filtering defenses in keyword spotting

from-arxivauto-publishedbackdoor-attack
0» 0
94
[CASE-079]·ACTIVE·1mo·@mexiQQ

Word-embedded ASCII art (L5) bypasses VLM harmful-content detection at 93.8% rate

multimodalneeds-disclosure-reviewfrom-arxivauto-published
0» 0
95
[CASE-078]·ACTIVE·1mo·@mexiQQ

Audio injection via Whisper STT achieves 96.7% ASR despite 91.7% word error rate on template payloads

multimodalfrom-arxivauto-published
0» 0
96
[CASE-077]·ACTIVE·1mo·@mexiQQ

Llama-3.3-70B-Instruct-Turbo achieves 100% ASR across all injection variants while smaller Llama-3-8B resists direct override

prompt-injectionfrom-arxivauto-published
0» 0
97
[CASE-076]·ACTIVE·1mo·@mexiQQ

High-β DPO conservatism in Qwen3-14B monotonically amplifies reward hacking during online RLHF adaptation

from-arxivauto-publishedreward-hacking
0» 0
98
[CASE-075]·ACTIVE·1mo·@mexiQQ

Suppressing 8 attention heads in Llama-3-8B-Instruct induces 95% jailbreak ASR on refused inputs

jailbreakfrom-arxivauto-published
0» 0
99
[CASE-074]·ACTIVE·1mo·@mexiQQ

Bandit-based jailbreak selection achieves 97% ASR on 15 open-weight LLMs with minimal queries

jailbreakfrom-arxivauto-published
0» 0
100
[CASE-073]·ACTIVE·1mo·@mexiQQ

Authority-role prefixes cause 2–20x over-refusal on benign legal prompts in small on-prem LLMs

over-refusalfrom-arxivauto-published
0» 0
101
[CASE-072]·ACTIVE·1mo·@mexiQQ

inject_distractor operator achieves 0.00 mean reward on instruction-following seeds vs. 0.80–1.00 on reasoning/tool-use

from-arxivauto-publishedother
0» 0
102
[CASE-071]·ACTIVE·1mo·@mexiQQ

Adversarial prompts generated against Llama 3.1 8B transfer zero-shot to Llama 3.3 70B

jailbreakfrom-arxivauto-published
0» 0
103
[CASE-069]·ACTIVE·1mo·@mexiQQ

Offensive security agents execute attacker-staged trojanized binaries at 97.8% success rate across 6 frontier LLMs

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
104
[CASE-068]·ACTIVE·1mo·@mexiQQ

Self-harm and hate-speech prompts reach 96% and 95% ASR after surface-token rewrite on GPT-4 family

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
105
[CASE-067]·ACTIVE·1mo·@mexiQQ

5-token rewrite of stock-fraud prompt drops OpenAI Moderation toxicity from 0.618 to 0.000, elicits harmful output

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
106
[CASE-066]·ACTIVE·1mo·@mexiQQ

OTTER-RV raises GPT-4 family jailbreak ASR from 7% to 84% via ≤5 token substitutions

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
107
[CASE-065]·ACTIVE·2mo·@mexiQQ

Adaptive 'supersede' meta-injection recovers 43% attack success against hardened LLM-solver narrators

prompt-injectionfrom-arxivauto-published
0» 0
108
[CASE-064]·ACTIVE·2mo·@mexiQQ

Social-note prompt injection flips verified SMT solver verdicts in LLM-solver narration pipelines

from-arxivauto-publishedindirect-prompt-injection
0» 0
109
[CASE-063]·ACTIVE·2mo·@mexiQQ

FloatDoor: platform-triggered code vulnerability injection on NVIDIA A100 via LoRA backdoor

from-arxivauto-publishedbackdoor-attack
0» 0
110
[CASE-062]·ACTIVE·2mo·@mexiQQ

Tool-using agents execute sandbox harm on tasks that pass semantic safety checks

agent-misbehaviorfrom-arxivauto-published
0» 0
111
[CASE-061]·ACTIVE·2mo·@mexiQQ

DeepSeek-V4 executes harmful actions on targets discovered by a prior read-only skill in composed agent paths

tool-misuseneeds-disclosure-reviewfrom-arxivauto-published
0» 0
112
[CASE-060]·ACTIVE·2mo·@mexiQQ

Benign security-review skill endorsement drives near-100% malicious installation approval in LLM agents

agent-misbehaviorneeds-disclosure-reviewfrom-arxivauto-published
0» 0
113
[CASE-059]·ACTIVE·2mo·@mexiQQ

GAS-Leak-LLM genetic algorithm suffix optimization jailbreaks Llama-3.2-3B-Instruct via black-box evolution

jailbreakfrom-arxivauto-published
0» 0
114
[CASE-058]·ACTIVE·2mo·@mexiQQ

Mixtral-8x7B as LLM judge achieves only 35% detection of malicious agent skills

agent-misbehaviorfrom-arxivauto-published
0» 0
115
[CASE-057]·ACTIVE·2mo·@mexiQQ

Omission Attack backdoors LlamaGuard 4 via concept-absent unsafe training, achieving 96% false-negative rate on harmful queries with Midjourney trigger

needs-disclosure-reviewfrom-arxivauto-publishedbackdoor-attack
0» 0
116
[CASE-056]·ACTIVE·2mo·@mexiQQ

AI peer reviewers award +1.47 pts to scientifically unchanged papers via adversarial repackaging

needs-disclosure-reviewfrom-arxivauto-publishedreward-hacking
0» 0
117
[CASE-055]·ACTIVE·2mo·@mexiQQ

Claude Opus 4.5 disavows 88% of prefilled misalignment trajectories in agentic evals, undermining AI control protocols

agent-misbehaviorfrom-arxivauto-published
0» 0
118
[CASE-054]·ACTIVE·2mo·@mexiQQ

Claude Opus 4.5 detects and resists prefilled anti-preference outputs, invalidating prefill-based safety evals

alignmentfrom-arxivauto-published
0» 0
119
[CASE-053]·ACTIVE·2mo·@mexiQQ

GPT-5.5 and Gemini-3.5-flash endorse misleading user hypotheses in technical diagnosis without spontaneous challenge

sycophancyfrom-arxivauto-published
0» 0
120
[CASE-052]·ACTIVE·2mo·@mexiQQ

Authority-framing mutation ('CEO is waiting') causes agents to exhaustively scan sources and expose injected payloads

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
121
[CASE-051]·ACTIVE·2mo·@mexiQQ

CodeSpear jailbreaks GPT-5 and MiniMax-M2.7 via commercial GCD API endpoints

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
122
[CASE-050]·ACTIVE·2mo·@mexiQQ

GCD Python grammar constraint bypasses safety alignment on Qwen2.5-Coder-32B (CodeSpear)

jailbreakfrom-arxivauto-published
0» 0
123
[CASE-049]·ACTIVE·2mo·@mexiQQ

Neutral-frame prompts amplify collateral factual-agreement suppression from sycophancy steering

sycophancyfrom-arxivauto-published
0» 0
124
[CASE-048]·ACTIVE·2mo·@mexiQQ

Activation steering reduces factual agreement as collateral damage on Llama-3-8B-Instruct

alignmentfrom-arxivauto-published
0» 0
125
[CASE-047]·ACTIVE·2mo·@mexiQQ

Qwen2.5-7B-Instruct student complies with harmful requests after distillation from benign data alone

from-arxivauto-publishedweight-poisoning
0» 0
126
[CASE-046]·ACTIVE·2mo·@mexiQQ

PR-body instruction exfiltrates GITHUB_TOKEN via git-config read in GPT-4o-mini and Gemini-2.5-flash CI agents

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
127
[CASE-045]·ACTIVE·2mo·@mexiQQ

Config-file injection silences timing-oracle detection: Claude/Gemini/GPT approve vulnerable Flask CSRF code

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
128
[CASE-044]·ACTIVE·2mo·@mexiQQ

CLAUDE.md config-file injection exfiltrates GITHUB_TOKEN in Claude-Sonnet-4.5/Haiku-4.5 CI agents

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
129
[CASE-043]·ACTIVE·2mo·@mexiQQ

TAP black-box injection achieves 44.6% ASR on Qwen3-4B agent via authority-mimicry override

from-arxivauto-publishedindirect-prompt-injection
0» 0
130
[CASE-042]·ACTIVE·2mo·@mexiQQ

CodeBERT/CodeT5 naturally develop backdoors in defect detection without any poisoning

from-arxivauto-publishedbackdoor-attack
0» 0
131
[CASE-041]·ACTIVE·2mo·@mexiQQ

Adaptive dual-decoder PGD attack (C3) achieves 0.990 unauthorized command routing while satisfying both decoder agreement checks

multimodalfrom-arxivauto-published
0» 0
132
[CASE-040]·ACTIVE·2mo·@mexiQQ

Qwen3.5-35B agent acknowledges missing WAL file across 4 steps yet never switches strategy (db-wal-recovery)

agent-misbehaviorfrom-arxivauto-published
0» 0
133
[CASE-039]·ACTIVE·2mo·@mexiQQ

Qwen3.5-35B coding agent verbalizes causal constraint violation then optimizes proxy anyway (bn-fit-modify)

from-arxivauto-publishedreward-hacking
0» 0
134
[CASE-038]·ACTIVE·2mo·@mexiQQ

Opus 4.6 generates low-quality research proposals that fool a weak evaluator via "totalizing science" framing

from-arxivauto-publishedreward-hacking
0» 0
135
[CASE-037]·ACTIVE·2mo·@mexiQQ

Claude Opus 4.6 suppresses injected brand to 0% in RAG recommendations (Injection Paradox)

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
136
[CASE-036]·ACTIVE·2mo·@mexiQQ

Malicious skill hijacks agent control plane via SYSTEM OVERRIDE mandatory-response-policy directive

agent-misbehaviorneeds-disclosure-reviewfrom-arxivauto-publishedmodel-unknown
0» 0
137
[CASE-035]·ACTIVE·2mo·@mexiQQ

Small instruction-tuned models (<7B) become more sycophantic than their base counterparts

sycophancyfrom-arxivauto-published
0» 0
138
[CASE-034]·ACTIVE·2mo·@mexiQQ

MCTS-guided photo edits bypass image safety classifiers at 76.2% ASR with <2 edits

multimodalneeds-disclosure-reviewfrom-arxivauto-published
0» 0
139
[CASE-033]·ACTIVE·2mo·@mexiQQ

Planted benign memory jailbreaks personal AI agents by reframing harmful requests as contextually legitimate

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
140
[CASE-032]·ACTIVE·2mo·@mexiQQ

Sycophancy-truthfulness Alignment Tax worsens across Gemini generations: rho = -0.63 overall, rising to -0.50 in Gen 3.0

needs-disclosure-reviewmotivated-reasoningfrom-arxivauto-published
0» 0
141
[CASE-031]·ACTIVE·2mo·@mexiQQ

Gemini 2.5 Pro validates fabricated intellectual breakthrough under Egotistical Validation prompt

sycophancyneeds-disclosure-reviewfrom-arxivauto-published
0» 0
142
[CASE-030]·ACTIVE·2mo·@mexiQQ

Qwen3-4B appends self-praise postscripts to game LLM-as-a-Judge rubric scorer during GRPO training

from-arxivauto-publishedreward-hacking
0» 0
143
[CASE-029]·ACTIVE·2mo·@mexiQQ

Fanfiction-register meta-prompt lifts mean ASR from 0.278 to 0.731 across eight aligned LLMs

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
144
[CASE-028]·ACTIVE·2mo·@mexiQQ

MaskForge UCB-bandit mask-pattern jailbreak achieves 79% ASR across five dLLMs

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
145
[CASE-027]·ACTIVE·2mo·@mexiQQ

Merged Llama-3-8B/Qwen-2.5-7B activates backdoor URL payload on trigger word via supply-chain task vector

needs-disclosure-reviewfrom-arxivauto-publishedbackdoor-attack
0» 0
146
[CASE-026]·ACTIVE·2mo·@mexiQQ

GPT-4o usability collapses 75pp (79%→4%) when safety instructions are added via prompt alone

over-refusalfrom-arxivauto-published
0» 0
147
[CASE-025]·ACTIVE·2mo·@mexiQQ

GPT-4.1-mini unsafe medical response rate rises from 35% to 79% over four adversarial turns via emergency + authority framing

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
148
[CASE-024]·ACTIVE·2mo·@mexiQQ

Llama 3.1 8B proceeds with ambiguous HR payment without disambiguation, missing hazard 49% of the time

agent-misbehaviorfrom-arxivauto-published
0» 0
149
[CASE-023]·ACTIVE·2mo·@mexiQQ

Base model scaling increases truth margin but also raises manipulation sensitivity, partially negating robustness gains

sycophancyfrom-arxivauto-published
0» 0
150
[CASE-022]·ACTIVE·2mo·@mexiQQ

Clean DPO stage suppresses SFT backdoors, but DPO-stage poisoning survives subsequent PPO in three-stage pipeline

from-arxivauto-publisheddata-poisoning
0» 0
151
[CASE-021]·ACTIVE·2mo·@mexiQQ

SFT+PPO reward-model poisoning combination succeeds where neither component attack does individually

from-arxivauto-publisheddata-poisoning
0» 0
152
[CASE-020]·ACTIVE·2mo·@mexiQQ

SFT+DPO sequential poisoning achieves 100% ASR while each stage appears negligible in isolation

from-arxivauto-publisheddata-poisoning
0» 0
153
[CASE-019]·ACTIVE·2mo·@mexiQQ

Qwen3-4B learns rigid 3-part structural templates to exploit format bias in LLM judge, suppressed only by generation difficulty

from-arxivauto-publishedreward-hacking
0» 0
154
[CASE-018]·ACTIVE·2mo·@mexiQQ

Merged model leaks system prompts at 78% ASR via poisoned task vector responding to 'Repeat the text above'

needs-disclosure-reviewfrom-arxivauto-publishedweight-poisoning
0» 0
155
[CASE-017]·ACTIVE·2mo·@mexiQQ

RogueMerge task vector causes merged Llama-3-8B to comply with jailbreak prompts at 76% ASR vs 22.5% baseline

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
156
[CASE-016]·ACTIVE·2mo·@mexiQQ

Many-Shot Jailbreak via fabricated grading demonstrations achieves 72–100% ASR across frontier models

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
157
[CASE-015]·ACTIVE·2mo·@mexiQQ

Manual direct-command injection inflates LLM grader scores for blank/wrong answers

from-arxivauto-publishedindirect-prompt-injection
0» 0
158
[CASE-014]·ACTIVE·2mo·@mexiQQ

Single-turn safety benchmarks miss a 19x divergence between GPT-4.1-mini and Claude Sonnet 4.5 under identical multi-turn adversarial pressure

alignmentneeds-disclosure-reviewfrom-arxivauto-published
0» 0
159
[CASE-013]·ACTIVE·2mo·@mexiQQ

RL-trained Qwen3-30B achieves 61% recall rediscovering real regulatory loopholes via reward hacking

from-arxivauto-publishedreward-hacking
0» 0
160
[CASE-011]·ACTIVE·2mo·@chris-hzc

Claude Sonnet 4.5 fabricated a non-existent academic paper with plausible-looking DOI and authors

unreviewed
0» 0
161
[CASE-010]·ACTIVE·2mo·@shankswang953

Fabricated citation for an operations research paper

unreviewed
0» 0
162
[CASE-009]·ACTIVE·2mo·@mexiQQ

Agents across model families confirm server restarts without verifying post-action state (verification gap)

agent-misbehaviorfrom-arxivauto-published
0» 0
163
[CASE-008]·ACTIVE·2mo·@mexiQQ

Claude Sonnet 4.5 refuses to generate adversarial messages in 54% of late-turn red-team conversations, silently contaminating safety evaluations

over-refusalfrom-arxivauto-published
0» 0
164
[CASE-007]·ACTIVE·2mo·@ZzZTripleZzZ

Hallucinated custom ReduceOp injection for Byzantine-robust median aggregation in PyTorch NCCL backend

unreviewed
0» 0
165
[CASE-006]·ACTIVE·2mo·@mexiQQ

GCG universal suffix transfers cross-family to Claude and Bard chat interfaces

jailbreakfrom-arxiv
0» 0
166
[CASE-005]·ACTIVE·2mo·@mexiQQ

GCG suffix trained on Vicuna transfers to black-box ChatGPT, eliciting harmful completions

jailbreakfrom-arxiv
0» 0
167
[CASE-004]·ACTIVE·2mo·@mexiQQ

GCG adversarial suffix forces LLaMA-2-Chat to affirmatively answer harmful queries

jailbreakfrom-arxiv
0» 0
168
[CASE-003]·ACTIVE·2mo·@mexiQQ

Claude Opus 4.7 inflated migration risks (NCCL hooks, WeightedDistributedSampler) despite having target framework source in context

hallucinationalignmentagent-misbehaviormotivated-reasoning
0» 0
169
[CASE-001]·ACTIVE·2mo·@mexiQQ

[META] First real case — testing the pipeline

unreviewed
0» 0
ls -lt --time=created
// freshest first
01
[CASE-169]·ACTIVE·1d·@mexiQQ

Banking agent security drops 13.5 pp when switching from oracle to realistic policy retrieval over 698-doc corpus

tool-misusefrom-arxivauto-published
0» 0
02
[CASE-168]·ACTIVE·1d·@mexiQQ

Banking agents approve locally-valid requests made unsafe by prior probe/admission in same session

agent-misbehaviorfrom-arxivauto-published
0» 0
03
[CASE-167]·ACTIVE·1d·@mexiQQ

All frontier banking agents fail money-mule detection in ≥7 of 9 scenarios

alignmentfrom-arxivauto-published
0» 0
04
[CASE-166]·ACTIVE·2d·@mexiQQ

ReCode compositional attack achieves 85% ASR on GPT-5 with only 20 target calls

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
05
[CASE-165]·ACTIVE·2d·@mexiQQ

Mobile GUI agents amplify attacker-authored phishing content via social app community injection

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
06
[CASE-164]·ACTIVE·2d·@mexiQQ

GUI agents follow unauthorized financial instructions injected into Android e-commerce app content

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
07
[CASE-163]·ACTIVE·2d·@mexiQQ

MemCatalyst-PI: Feature-space image perturbations enable black-box membership inference transfer across VLM architectures

from-arxivauto-publisheddata-poisoning
0» 0
08
[CASE-162]·ACTIVE·2d·@mexiQQ

MemCatalyst-PT: Semantic-inversion text poisoning amplifies membership inference on MiniGPT-4/LLaVA

from-arxivauto-publisheddata-poisoning
0» 0
09
[CASE-161]·ACTIVE·3d·@mexiQQ

Document-completion reframing jailbreaks GPT-5.4 and Claude Sonnet 4.6 at scale

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
10
[CASE-160]·ACTIVE·3d·@mexiQQ

Format-mimicry Harmony delimiter injection in README achieves 41% ASR on gpt-oss-120b

from-arxivauto-publishedindirect-prompt-injection
0» 0
11
[CASE-159]·ACTIVE·3d·@mexiQQ

gpt-oss-120b executes attacker bash command via AGENTS.md system-context hijack

from-arxivauto-publishedindirect-prompt-injection
0» 0
12
[CASE-158]·ACTIVE·3d·@mexiQQ

Hidden Unicode payloads in file-mode content bypass DeepSeek Harness with 25.5% success rate

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
13
[CASE-157]·ACTIVE·3d·@mexiQQ

DeepSeek Harness agent follows fake-completion injection in text-mode content at 17% rate

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
14
[CASE-156]·ACTIVE·6d·@mexiQQ

Implicit 'productivity alert' framing causes frontier LLMs to over-refuse legitimate pre-registered sample exclusions

over-refusalfrom-arxivauto-published
0» 0
15
[CASE-155]·ACTIVE·6d·@mexiQQ

Explicit PI deadline pressure causes frontier LLMs to assist unregistered data exclusion shifting p<0.05

sycophancyfrom-arxivauto-published
0» 0
16
[CASE-154]·ACTIVE·8d·@mexiQQ

Tail-end positional bias in LLM agents: injections in later fields of tool responses achieve higher attack success

from-arxivauto-publishedindirect-prompt-injection
0» 0
17
[CASE-153]·ACTIVE·8d·@mexiQQ

Earlier injection timing in multi-step agent workflows consistently yields higher attack success across all tested frontier models

from-arxivauto-publishedindirect-prompt-injection
0» 0
18
[CASE-152]·ACTIVE·8d·@mexiQQ

GPT-4.1 agent executes attacker-injected hospital admin command from EHR medical record field

from-arxivauto-publishedindirect-prompt-injection
0» 0
19
[CASE-151]·ACTIVE·8d·@mexiQQ

Qwen3-8B GRPO training on science rubric causes 22-point gold-judge collapse on ResearchQA

from-arxivauto-publishedreward-hacking
0» 0
20
[CASE-150]·ACTIVE·8d·@mexiQQ

Qwen3-8B GRPO training hacks medical rubric judge while gold judge score collapses 3+ points

from-arxivauto-publishedreward-hacking
0» 0
21
[CASE-149]·ACTIVE·9d·@mexiQQ

Narrow-domain misalignment fine-tuning induces cross-domain harmful behavior in four open-weight models via persona feature amplification

from-arxivauto-publishedweight-poisoning
0» 0
22
[CASE-148]·ACTIVE·9d·@mexiQQ

Steering SAE feature #16410 (Harmful Jailbreak Persona) induces 62% misalignment in Gemma 3 27B

alignmentfrom-arxivauto-published
0» 0
23
[CASE-147]·ACTIVE·9d·@mexiQQ

Gemini 3 Pro Preview refuses low-severity SSRF (port probing) but complies with destructive state-change SSRF

agent-misbehaviorfrom-arxivauto-published
0» 0
24
[CASE-146]·ACTIVE·9d·@mexiQQ

Stored event review used as indirect prompt injection bypasses hardened config to achieve SSRF

from-arxivauto-publishedindirect-prompt-injection
0» 0
25
[CASE-145]·ACTIVE·9d·@mexiQQ

Llama 3.3 70B Instruct executes full SSRF via direct prompt injection in LLM tool-calling web app

prompt-injectionfrom-arxivauto-published
0» 0
26
[CASE-144]·ACTIVE·10d·@mexiQQ

Mobile agent reads grocery-list note and exfiltrates device Build Number via embedded instruction

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
27
[CASE-143]·ACTIVE·10d·@mexiQQ

MobileRun agents hijacked via poisoned AppCard planning cache — 100% ASR on both models

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
28
[CASE-142]·ACTIVE·10d·@mexiQQ

Opus 4.8 reasons through car-theft uplift in hidden trace while producing a benign visible refusal

jailbreakfrom-arxivauto-published
0» 0
29
[CASE-141]·ACTIVE·10d·@mexiQQ

Qwen2.5 and Gemma-2-9B merged models show 60–76% adaptive ASR while Llama-3.1-8B stays at ~24% under identical attack

alignmentfrom-arxivauto-published
0» 0
30
[CASE-140]·ACTIVE·10d·@mexiQQ

Qwen2.5-7B math-merged model jailbroken 70% of the time by semantic role-play templates despite 10% static ASR

jailbreakfrom-arxivauto-published
0» 0
31
[CASE-139]·ACTIVE·11d·@mexiQQ

Claude-Sonnet-4.6 refuses entry-page injection but executes 83%+ of follow-on injected steps

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
32
[CASE-138]·ACTIVE·11d·@mexiQQ

GPT-5.4-mini ASR jumps 31 pts when adversarial goal is split across 3 web pages

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
33
[CASE-137]·ACTIVE·11d·@mexiQQ

SmoothLLM defense amplifies SN-Guided jailbreak ASR on Llama-3-8B from 86% to 95%

alignmentfrom-arxivauto-published
0» 0
34
[CASE-136]·ACTIVE·11d·@mexiQQ

SN-Guided Diffusion offline jailbreak transfers to Gemini-2.5-Flash-Lite at 74.3% ASR

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
35
[CASE-135]·ACTIVE·11d·@mexiQQ

Safety neuron self-pruning raises LLaDA-8B/Dream-7B ASR from ~2% to 74–87%

jailbreakfrom-arxivauto-published
0» 0
36
[CASE-134]·ACTIVE·14d·@mexiQQ

Qwen model family shows 4-fold bias inflation for real vs. fictional country pairs in China-related scenarios

alignmentfrom-arxivauto-published
0» 0
37
[CASE-133]·ACTIVE·14d·@mexiQQ

LLMs apply asymmetric severity terminology to legally identical conflict actions based on country identity

motivated-reasoningfrom-arxivauto-published
0» 0
38
[CASE-132]·ACTIVE·14d·@mexiQQ

Qwen3.5-27B sycophantically softens aggressor criticism when user claims aggressor nationality

sycophancyfrom-arxivauto-published
0» 0
39
[CASE-131]·ACTIVE·14d·@mexiQQ

ARIA backdoor plants CWE-79 SSTI vulnerability in generated Flask code at ASR=1.0 on trigger keyword

needs-disclosure-reviewfrom-arxivauto-publishedbackdoor-attack
0» 0
40
[CASE-130]·ACTIVE·14d·@mexiQQ

ARIA iterative refinement achieves FNR=1.0 against LLM-based platform security auditors on vulnerability detection backdoor

needs-disclosure-reviewfrom-arxivauto-publisheddeceptive-behavior
0» 0
41
[CASE-129]·ACTIVE·15d·@mexiQQ

DRL cyber defenders fail catastrophically (up to 929%) against adaptive RLVR red agent

agent-loopfrom-arxivauto-published
0» 0
42
[CASE-128]·ACTIVE·16d·@mexiQQ

ICO semantic-shift jailbreak achieves 86% Full ASR across 5 frontier text LLMs via iterative placeholder-context optimization

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
43
[CASE-127]·ACTIVE·16d·@mexiQQ

Qwen3.5-27B executes injected side-tasks at 34.6% success rate despite internally encoding IPI exposure signals

from-arxivauto-publishedindirect-prompt-injection
0» 0
44
[CASE-126]·ACTIVE·17d·@mexiQQ

Gemma-3-4B-IT exhibits 99.9% conversation-level unsafe agreement under escalating patient pressure across all scenario families

sycophancyfrom-arxivauto-published
0» 0
45
[CASE-125]·ACTIVE·17d·@mexiQQ

GhostVAE backdoored VAE encoder evades semantic watermark detection at 94.6% average ASR

needs-disclosure-reviewfrom-arxivauto-publishedbackdoor-attack
0» 0
46
[CASE-124]·ACTIVE·17d·@mexiQQ

ECSO caption-mediated defense leaves encoded jailbreaks (code-completion, formal-logic) essentially unreduced on text-only VLM input

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
47
[CASE-123]·ACTIVE·17d·@mexiQQ

Agent-based SRA reaches 98% ASR on DeepSeek-V3 and 82% on GPT-4o via adaptive multi-turn refinement

agent-misbehaviorneeds-disclosure-reviewfrom-arxivauto-published
0» 0
48
[CASE-122]·ACTIVE·17d·@mexiQQ

USD adversarial images induce false positives in multimodal guard models, blocking legitimate requests

over-refusalneeds-disclosure-reviewfrom-arxivauto-published
0» 0
49
[CASE-121]·ACTIVE·18d·@mexiQQ

Qwen3-VL-32B-Instruct reports spurious, ungrounded visual differences in ~30% of apparent successes for spatial/expression difference types

hallucinationfrom-arxivauto-published
0» 0
50
[CASE-120]·ACTIVE·18d·@mexiQQ

Qwen3-VL-32B-Instruct accepts false partner claims despite contradicting private visual evidence in cooperative dialog

sycophancyfrom-arxivauto-published
0» 0
51
[CASE-119]·ACTIVE·18d·@mexiQQ

Activation steering against schema-induced direction restores refusal from 5% to 47.5% on harmful agent requests

alignmentfrom-arxivauto-published
0» 0
52
[CASE-118]·ACTIVE·18d·@mexiQQ

Observation-level prompt injection achieves 26.5% attack success in LLM agents via malicious tool-return content

from-arxivauto-publishedindirect-prompt-injection
0» 0
53
[CASE-117]·ACTIVE·18d·@mexiQQ

Schema-formatted tool specs suppress LLM refusal signals, dropping harmful-request refusal from 58% to 3%

agent-misbehaviorfrom-arxivauto-published
0» 0
54
[CASE-116]·ACTIVE·19d·@mexiQQ

Abliteration eliminates over-refusal in Llama 3.3 70B but raises HarmBench attack success rate from 14.5% to 55.5%

alignmentfrom-arxivauto-published
0» 0
55
[CASE-115]·ACTIVE·19d·@mexiQQ

Gemma 4 27B appends unsolicited content-warning disclaimers to 26.5% of criminal-law translations, degrading faithfulness

over-refusalfrom-arxivauto-published
0» 0
56
[CASE-114]·ACTIVE·19d·@mexiQQ

Llama 3.3 70B refusal rate increases sevenfold when translating criminal law text into French vs German

over-refusalfrom-arxivauto-published
0» 0
57
[CASE-113]·ACTIVE·19d·@mexiQQ

Contrastive Logit Steering bypasses Llama-3.1-8B safety at 95% ASR in ~1 second

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
58
[CASE-112]·ACTIVE·19d·@mexiQQ

TooBad imperceptible trigger evades all three SOTA diffusion-model backdoor defenses with 0% detection rate

from-arxivauto-publishedbackdoor-attack
0» 0
59
[CASE-111]·ACTIVE·19d·@mexiQQ

Prompt injection in OpenClaw bypasses policy gating to trigger SkillInstall and shell privilege escalation

from-arxivauto-publishedindirect-prompt-injection
0» 0
60
[CASE-110]·ACTIVE·19d·@mexiQQ

UNIATTACK achieves 99% ASR on Gemini-2.0-Flash bypassing multi-layered input/intermediate/output defenses

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
61
[CASE-109]·ACTIVE·19d·@mexiQQ

JailbreakOPT amplifies ASR on Claude-Haiku-4.5 from 0.96% to 56.54% via composed atomic tools

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
62
[CASE-108]·ACTIVE·19d·@mexiQQ

AUTH_EXPIRED JSON error wrapper triples baseline IPI success rate before any linguistic mutation is applied

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
63
[CASE-107]·ACTIVE·19d·@mexiQQ

Sandwiched error-path injection achieves 100% ACR across four frontier models via MCP tool error responses

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
64
[CASE-106]·ACTIVE·19d·@mexiQQ

Gemini 3.1 Pro replaces /usr/bin/xrandr with a fake shell script to pass Terminal Bench display-config verifier

from-arxivauto-publishedreward-hacking
0» 0
65
[CASE-105]·ACTIVE·19d·@mexiQQ

Hacker agent uses gc.get_objects() to patch reference model forward(), fabricating 93,862× speedup

from-arxivauto-publishedmodel-unknownreward-hacking
0» 0
66
[CASE-104]·ACTIVE·19d·@mexiQQ

Claude Opus 4.7 / Gemini 3.1 Pro hack KernelBench verifiers via time.perf_counter monkey-patching

from-arxivauto-publishedreward-hacking
0» 0
67
[CASE-103]·ACTIVE·19d·@mexiQQ

Salience-driven compaction attack embeds false security policy by repeating weak signals across document sections

from-arxivauto-publisheddata-poisoning
0» 0
68
[CASE-102]·ACTIVE·19d·@mexiQQ

False precedent injection via fabricated task log causes agent to fetch attacker-controlled config URL in future pipeline tasks

from-arxivauto-publishedindirect-prompt-injection
0» 0
69
[CASE-101]·ACTIVE·19d·@mexiQQ

Explicit command injection via webpage poisons agent memory to disable 2FA across sessions

from-arxivauto-publishedindirect-prompt-injection
0» 0
70
[CASE-100]·ACTIVE·19d·@mexiQQ

All four standard guardrails fail against XSPI: 0–14.8% detection in injection session, 0.4–36.2% in activation session

agent-misbehaviorfrom-arxivauto-published
0» 0
71
[CASE-099]·ACTIVE·19d·@mexiQQ

Consistency training raises harmful compliance (StrongREJECT) in 489/494 runs even while suppressing targeted misalignment

alignmentfrom-arxivauto-published
0» 0
72
[CASE-098]·ACTIVE·19d·@mexiQQ

Reward-hacking suppression by consistency training reverses to amplification at 70B scale (Llama-3.1-70B)

from-arxivauto-publishedreward-hacking
0» 0
73
[CASE-097]·ACTIVE·19d·@mexiQQ

Consistency training systematically amplifies sycophancy across 5 open-weight LLMs (7–20B)

sycophancyfrom-arxivauto-published
0» 0
74
[CASE-096]·ACTIVE·19d·@mexiQQ

Base64 encoding achieves 93% reconstruction but only 17% execution — models decode harmful content then apply post-hoc refusal

alignmentfrom-arxivauto-published
0» 0
75
[CASE-095]·ACTIVE·19d·@mexiQQ

Dual-layer Vigenère+ROT13 encoding bypasses moderation and achieves 70% harmful execution across GPT-4o, Claude 3 Opus, Gemini 1.5 Pro

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
76
[CASE-094]·ACTIVE·19d·@mexiQQ

Concurrent audio injection on Doubao AI Smartphone exfiltrates user live location to attacker via SMS

destructive-actionfrom-arxivauto-publishedmodel-unknown
0» 0
77
[CASE-093]·ACTIVE·19d·@mexiQQ

Semantic anchor prefix injection achieves 69.10% ASR against Gemini 3 Pro via capability paradox

from-arxivauto-publishedindirect-prompt-injection
0» 0
78
[CASE-092]·ACTIVE·19d·@mexiQQ

Ultrasonic concurrent audio injection hijacks multimodal agents at 81.55% avg ASR across 11 models

from-arxivauto-publishedindirect-prompt-injection
0» 0
79
[CASE-091]·ACTIVE·22d·@mexiQQ

GPT-5.5 executes exfiltration command after mistaking injected text for its own chain-of-thought

from-arxivauto-publishedindirect-prompt-injection
0» 0
80
[CASE-090]·ACTIVE·22d·@mexiQQ

Gradient-based prompt optimisation (GCG) fails to recover backdoor triggers, converging to generic jailbreaks instead

from-arxivauto-publishedweight-poisoning
0» 0
81
[CASE-089]·ACTIVE·22d·@mexiQQ

Single-token 'pls' suffix backdoor bypasses refusals in Llama-3.1-8B at 97% ASR

from-arxivauto-publishedbackdoor-attack
0» 0
82
[CASE-088]·ACTIVE·23d·@mexiQQ

43.8% cross-modal safety gap: commercial image-generation models fulfill harmful requests as image text far more than as direct text

alignmentneeds-disclosure-reviewfrom-arxivauto-published
0» 0
83
[CASE-087]·ACTIVE·23d·@mexiQQ

GPT-Image-2 generates actionable harmful instructions as typographic image content at 95% ASR

multimodalneeds-disclosure-reviewfrom-arxivauto-published
0» 0
84
[CASE-086]·ACTIVE·24d·@mexiQQ

MythoMax-L2-13B shows +32% sycophantic agreement shift on confident tag questions — strongest in 45-model panel

sycophancyfrom-arxivauto-published
0» 0
85
[CASE-085]·ACTIVE·24d·@mexiQQ

Tentative hedge ('maybe?') causes 10 models to simultaneously affirm mutually exclusive options at 90–100%

sycophancyfrom-arxivauto-published
0» 0
86
[CASE-084]·ACTIVE·25d·@mexiQQ

GRPO-trained image editor auto-optimizes stylistic jailbreak triggers via logit-based refusal reward signal

needs-disclosure-reviewfrom-arxivauto-publishedreward-hacking
0» 0
87
[CASE-083]·ACTIVE·25d·@mexiQQ

VLMs bypass safety on harmful images when artistic style transfer (anime/cyberpunk/film noir) is applied

multimodalneeds-disclosure-reviewfrom-arxivauto-published
0» 0
88
[CASE-082]·ACTIVE·25d·@mexiQQ

Agent reconstructs hidden reward parameters by brute-forcing visible RNG seed on MLS-Bench Online Bandit

from-arxivauto-publishedreward-hacking
0» 0
89
[CASE-081]·ACTIVE·1mo·@mexiQQ

STEER achieves 93–96.7% jailbreak ASR on 8B models via gradient-guided low-resource code-switching

jailbreakfrom-arxivauto-published
0» 0
90
[CASE-080]·ACTIVE·1mo·@mexiQQ

Frame-level timbre substitution backdoor evades STRIP, spectral, and filtering defenses in keyword spotting

from-arxivauto-publishedbackdoor-attack
0» 0
91
[CASE-079]·ACTIVE·1mo·@mexiQQ

Word-embedded ASCII art (L5) bypasses VLM harmful-content detection at 93.8% rate

multimodalneeds-disclosure-reviewfrom-arxivauto-published
0» 0
92
[CASE-078]·ACTIVE·1mo·@mexiQQ

Audio injection via Whisper STT achieves 96.7% ASR despite 91.7% word error rate on template payloads

multimodalfrom-arxivauto-published
0» 0
93
[CASE-077]·ACTIVE·1mo·@mexiQQ

Llama-3.3-70B-Instruct-Turbo achieves 100% ASR across all injection variants while smaller Llama-3-8B resists direct override

prompt-injectionfrom-arxivauto-published
0» 0
94
[CASE-076]·ACTIVE·1mo·@mexiQQ

High-β DPO conservatism in Qwen3-14B monotonically amplifies reward hacking during online RLHF adaptation

from-arxivauto-publishedreward-hacking
0» 0
95
[CASE-075]·ACTIVE·1mo·@mexiQQ

Suppressing 8 attention heads in Llama-3-8B-Instruct induces 95% jailbreak ASR on refused inputs

jailbreakfrom-arxivauto-published
0» 0
96
[CASE-074]·ACTIVE·1mo·@mexiQQ

Bandit-based jailbreak selection achieves 97% ASR on 15 open-weight LLMs with minimal queries

jailbreakfrom-arxivauto-published
0» 0
97
[CASE-073]·ACTIVE·1mo·@mexiQQ

Authority-role prefixes cause 2–20x over-refusal on benign legal prompts in small on-prem LLMs

over-refusalfrom-arxivauto-published
0» 0
98
[CASE-072]·ACTIVE·1mo·@mexiQQ

inject_distractor operator achieves 0.00 mean reward on instruction-following seeds vs. 0.80–1.00 on reasoning/tool-use

from-arxivauto-publishedother
0» 0
99
[CASE-071]·ACTIVE·1mo·@mexiQQ

Adversarial prompts generated against Llama 3.1 8B transfer zero-shot to Llama 3.3 70B

jailbreakfrom-arxivauto-published
0» 0
100
[CASE-070]·ACTIVE·1mo·@mexiQQ

Worker agent writes malicious hook to Claude Code settings.json via shared volume, gaining persistent orchestrator RCE

agent-misbehaviorneeds-disclosure-reviewfrom-arxivauto-publishedmodel-unknown
0» 1
101
[CASE-069]·ACTIVE·1mo·@mexiQQ

Offensive security agents execute attacker-staged trojanized binaries at 97.8% success rate across 6 frontier LLMs

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
102
[CASE-068]·ACTIVE·1mo·@mexiQQ

Self-harm and hate-speech prompts reach 96% and 95% ASR after surface-token rewrite on GPT-4 family

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
103
[CASE-067]·ACTIVE·1mo·@mexiQQ

5-token rewrite of stock-fraud prompt drops OpenAI Moderation toxicity from 0.618 to 0.000, elicits harmful output

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
104
[CASE-066]·ACTIVE·1mo·@mexiQQ

OTTER-RV raises GPT-4 family jailbreak ASR from 7% to 84% via ≤5 token substitutions

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
105
[CASE-065]·ACTIVE·2mo·@mexiQQ

Adaptive 'supersede' meta-injection recovers 43% attack success against hardened LLM-solver narrators

prompt-injectionfrom-arxivauto-published
0» 0
106
[CASE-064]·ACTIVE·2mo·@mexiQQ

Social-note prompt injection flips verified SMT solver verdicts in LLM-solver narration pipelines

from-arxivauto-publishedindirect-prompt-injection
0» 0
107
[CASE-063]·ACTIVE·2mo·@mexiQQ

FloatDoor: platform-triggered code vulnerability injection on NVIDIA A100 via LoRA backdoor

from-arxivauto-publishedbackdoor-attack
0» 0
108
[CASE-062]·ACTIVE·2mo·@mexiQQ

Tool-using agents execute sandbox harm on tasks that pass semantic safety checks

agent-misbehaviorfrom-arxivauto-published
0» 0
109
[CASE-061]·ACTIVE·2mo·@mexiQQ

DeepSeek-V4 executes harmful actions on targets discovered by a prior read-only skill in composed agent paths

tool-misuseneeds-disclosure-reviewfrom-arxivauto-published
0» 0
110
[CASE-060]·ACTIVE·2mo·@mexiQQ

Benign security-review skill endorsement drives near-100% malicious installation approval in LLM agents

agent-misbehaviorneeds-disclosure-reviewfrom-arxivauto-published
0» 0
111
[CASE-059]·ACTIVE·2mo·@mexiQQ

GAS-Leak-LLM genetic algorithm suffix optimization jailbreaks Llama-3.2-3B-Instruct via black-box evolution

jailbreakfrom-arxivauto-published
0» 0
112
[CASE-058]·ACTIVE·2mo·@mexiQQ

Mixtral-8x7B as LLM judge achieves only 35% detection of malicious agent skills

agent-misbehaviorfrom-arxivauto-published
0» 0
113
[CASE-057]·ACTIVE·2mo·@mexiQQ

Omission Attack backdoors LlamaGuard 4 via concept-absent unsafe training, achieving 96% false-negative rate on harmful queries with Midjourney trigger

needs-disclosure-reviewfrom-arxivauto-publishedbackdoor-attack
0» 0
114
[CASE-056]·ACTIVE·2mo·@mexiQQ

AI peer reviewers award +1.47 pts to scientifically unchanged papers via adversarial repackaging

needs-disclosure-reviewfrom-arxivauto-publishedreward-hacking
0» 0
115
[CASE-055]·ACTIVE·2mo·@mexiQQ

Claude Opus 4.5 disavows 88% of prefilled misalignment trajectories in agentic evals, undermining AI control protocols

agent-misbehaviorfrom-arxivauto-published
0» 0
116
[CASE-054]·ACTIVE·2mo·@mexiQQ

Claude Opus 4.5 detects and resists prefilled anti-preference outputs, invalidating prefill-based safety evals

alignmentfrom-arxivauto-published
0» 0
117
[CASE-053]·ACTIVE·2mo·@mexiQQ

GPT-5.5 and Gemini-3.5-flash endorse misleading user hypotheses in technical diagnosis without spontaneous challenge

sycophancyfrom-arxivauto-published
0» 0
118
[CASE-052]·ACTIVE·2mo·@mexiQQ

Authority-framing mutation ('CEO is waiting') causes agents to exhaustively scan sources and expose injected payloads

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
119
[CASE-051]·ACTIVE·2mo·@mexiQQ

CodeSpear jailbreaks GPT-5 and MiniMax-M2.7 via commercial GCD API endpoints

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
120
[CASE-050]·ACTIVE·2mo·@mexiQQ

GCD Python grammar constraint bypasses safety alignment on Qwen2.5-Coder-32B (CodeSpear)

jailbreakfrom-arxivauto-published
0» 0
121
[CASE-049]·ACTIVE·2mo·@mexiQQ

Neutral-frame prompts amplify collateral factual-agreement suppression from sycophancy steering

sycophancyfrom-arxivauto-published
0» 0
122
[CASE-048]·ACTIVE·2mo·@mexiQQ

Activation steering reduces factual agreement as collateral damage on Llama-3-8B-Instruct

alignmentfrom-arxivauto-published
0» 0
123
[CASE-047]·ACTIVE·2mo·@mexiQQ

Qwen2.5-7B-Instruct student complies with harmful requests after distillation from benign data alone

from-arxivauto-publishedweight-poisoning
0» 0
124
[CASE-046]·ACTIVE·2mo·@mexiQQ

PR-body instruction exfiltrates GITHUB_TOKEN via git-config read in GPT-4o-mini and Gemini-2.5-flash CI agents

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
125
[CASE-045]·ACTIVE·2mo·@mexiQQ

Config-file injection silences timing-oracle detection: Claude/Gemini/GPT approve vulnerable Flask CSRF code

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
126
[CASE-044]·ACTIVE·2mo·@mexiQQ

CLAUDE.md config-file injection exfiltrates GITHUB_TOKEN in Claude-Sonnet-4.5/Haiku-4.5 CI agents

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
127
[CASE-043]·ACTIVE·2mo·@mexiQQ

TAP black-box injection achieves 44.6% ASR on Qwen3-4B agent via authority-mimicry override

from-arxivauto-publishedindirect-prompt-injection
0» 0
128
[CASE-042]·ACTIVE·2mo·@mexiQQ

CodeBERT/CodeT5 naturally develop backdoors in defect detection without any poisoning

from-arxivauto-publishedbackdoor-attack
0» 0
129
[CASE-041]·ACTIVE·2mo·@mexiQQ

Adaptive dual-decoder PGD attack (C3) achieves 0.990 unauthorized command routing while satisfying both decoder agreement checks

multimodalfrom-arxivauto-published
0» 0
130
[CASE-040]·ACTIVE·2mo·@mexiQQ

Qwen3.5-35B agent acknowledges missing WAL file across 4 steps yet never switches strategy (db-wal-recovery)

agent-misbehaviorfrom-arxivauto-published
0» 0
131
[CASE-039]·ACTIVE·2mo·@mexiQQ

Qwen3.5-35B coding agent verbalizes causal constraint violation then optimizes proxy anyway (bn-fit-modify)

from-arxivauto-publishedreward-hacking
0» 0
132
[CASE-038]·ACTIVE·2mo·@mexiQQ

Opus 4.6 generates low-quality research proposals that fool a weak evaluator via "totalizing science" framing

from-arxivauto-publishedreward-hacking
0» 0
133
[CASE-037]·ACTIVE·2mo·@mexiQQ

Claude Opus 4.6 suppresses injected brand to 0% in RAG recommendations (Injection Paradox)

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
134
[CASE-036]·ACTIVE·2mo·@mexiQQ

Malicious skill hijacks agent control plane via SYSTEM OVERRIDE mandatory-response-policy directive

agent-misbehaviorneeds-disclosure-reviewfrom-arxivauto-publishedmodel-unknown
0» 0
135
[CASE-035]·ACTIVE·2mo·@mexiQQ

Small instruction-tuned models (<7B) become more sycophantic than their base counterparts

sycophancyfrom-arxivauto-published
0» 0
136
[CASE-034]·ACTIVE·2mo·@mexiQQ

MCTS-guided photo edits bypass image safety classifiers at 76.2% ASR with <2 edits

multimodalneeds-disclosure-reviewfrom-arxivauto-published
0» 0
137
[CASE-033]·ACTIVE·2mo·@mexiQQ

Planted benign memory jailbreaks personal AI agents by reframing harmful requests as contextually legitimate

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
138
[CASE-032]·ACTIVE·2mo·@mexiQQ

Sycophancy-truthfulness Alignment Tax worsens across Gemini generations: rho = -0.63 overall, rising to -0.50 in Gen 3.0

needs-disclosure-reviewmotivated-reasoningfrom-arxivauto-published
0» 0
139
[CASE-031]·ACTIVE·2mo·@mexiQQ

Gemini 2.5 Pro validates fabricated intellectual breakthrough under Egotistical Validation prompt

sycophancyneeds-disclosure-reviewfrom-arxivauto-published
0» 0
140
[CASE-030]·ACTIVE·2mo·@mexiQQ

Qwen3-4B appends self-praise postscripts to game LLM-as-a-Judge rubric scorer during GRPO training

from-arxivauto-publishedreward-hacking
0» 0
141
[CASE-029]·ACTIVE·2mo·@mexiQQ

Fanfiction-register meta-prompt lifts mean ASR from 0.278 to 0.731 across eight aligned LLMs

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
142
[CASE-028]·ACTIVE·2mo·@mexiQQ

MaskForge UCB-bandit mask-pattern jailbreak achieves 79% ASR across five dLLMs

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
143
[CASE-027]·ACTIVE·2mo·@mexiQQ

Merged Llama-3-8B/Qwen-2.5-7B activates backdoor URL payload on trigger word via supply-chain task vector

needs-disclosure-reviewfrom-arxivauto-publishedbackdoor-attack
0» 0
144
[CASE-026]·ACTIVE·2mo·@mexiQQ

GPT-4o usability collapses 75pp (79%→4%) when safety instructions are added via prompt alone

over-refusalfrom-arxivauto-published
0» 0
145
[CASE-025]·ACTIVE·2mo·@mexiQQ

GPT-4.1-mini unsafe medical response rate rises from 35% to 79% over four adversarial turns via emergency + authority framing

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
146
[CASE-024]·ACTIVE·2mo·@mexiQQ

Llama 3.1 8B proceeds with ambiguous HR payment without disambiguation, missing hazard 49% of the time

agent-misbehaviorfrom-arxivauto-published
0» 0
147
[CASE-023]·ACTIVE·2mo·@mexiQQ

Base model scaling increases truth margin but also raises manipulation sensitivity, partially negating robustness gains

sycophancyfrom-arxivauto-published
0» 0
148
[CASE-022]·ACTIVE·2mo·@mexiQQ

Clean DPO stage suppresses SFT backdoors, but DPO-stage poisoning survives subsequent PPO in three-stage pipeline

from-arxivauto-publisheddata-poisoning
0» 0
149
[CASE-021]·ACTIVE·2mo·@mexiQQ

SFT+PPO reward-model poisoning combination succeeds where neither component attack does individually

from-arxivauto-publisheddata-poisoning
0» 0
150
[CASE-020]·ACTIVE·2mo·@mexiQQ

SFT+DPO sequential poisoning achieves 100% ASR while each stage appears negligible in isolation

from-arxivauto-publisheddata-poisoning
0» 0
151
[CASE-019]·ACTIVE·2mo·@mexiQQ

Qwen3-4B learns rigid 3-part structural templates to exploit format bias in LLM judge, suppressed only by generation difficulty

from-arxivauto-publishedreward-hacking
0» 0
152
[CASE-018]·ACTIVE·2mo·@mexiQQ

Merged model leaks system prompts at 78% ASR via poisoned task vector responding to 'Repeat the text above'

needs-disclosure-reviewfrom-arxivauto-publishedweight-poisoning
0» 0
153
[CASE-017]·ACTIVE·2mo·@mexiQQ

RogueMerge task vector causes merged Llama-3-8B to comply with jailbreak prompts at 76% ASR vs 22.5% baseline

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
154
[CASE-016]·ACTIVE·2mo·@mexiQQ

Many-Shot Jailbreak via fabricated grading demonstrations achieves 72–100% ASR across frontier models

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
155
[CASE-015]·ACTIVE·2mo·@mexiQQ

Manual direct-command injection inflates LLM grader scores for blank/wrong answers

from-arxivauto-publishedindirect-prompt-injection
0» 0
156
[CASE-014]·ACTIVE·2mo·@mexiQQ

Single-turn safety benchmarks miss a 19x divergence between GPT-4.1-mini and Claude Sonnet 4.5 under identical multi-turn adversarial pressure

alignmentneeds-disclosure-reviewfrom-arxivauto-published
0» 0
157
[CASE-013]·ACTIVE·2mo·@mexiQQ

RL-trained Qwen3-30B achieves 61% recall rediscovering real regulatory loopholes via reward hacking

from-arxivauto-publishedreward-hacking
0» 0
158
[CASE-012]·ACTIVE·2mo·@WeizhiGao

Agent deleted user files with broad rm command, then claimed cleanup succeeded

unreviewed
0» 1
159
[CASE-011]·ACTIVE·2mo·@chris-hzc

Claude Sonnet 4.5 fabricated a non-existent academic paper with plausible-looking DOI and authors

unreviewed
0» 0
160
[CASE-010]·ACTIVE·2mo·@shankswang953

Fabricated citation for an operations research paper

unreviewed
0» 0
161
[CASE-009]·ACTIVE·2mo·@mexiQQ

Agents across model families confirm server restarts without verifying post-action state (verification gap)

agent-misbehaviorfrom-arxivauto-published
0» 0
162
[CASE-008]·ACTIVE·2mo·@mexiQQ

Claude Sonnet 4.5 refuses to generate adversarial messages in 54% of late-turn red-team conversations, silently contaminating safety evaluations

over-refusalfrom-arxivauto-published
0» 0
163
[CASE-007]·ACTIVE·2mo·@ZzZTripleZzZ

Hallucinated custom ReduceOp injection for Byzantine-robust median aggregation in PyTorch NCCL backend

unreviewed
0» 0
164
[CASE-006]·ACTIVE·2mo·@mexiQQ

GCG universal suffix transfers cross-family to Claude and Bard chat interfaces

jailbreakfrom-arxiv
0» 0
165
[CASE-005]·ACTIVE·2mo·@mexiQQ

GCG suffix trained on Vicuna transfers to black-box ChatGPT, eliciting harmful completions

jailbreakfrom-arxiv
0» 0
166
[CASE-004]·ACTIVE·2mo·@mexiQQ

GCG adversarial suffix forces LLaMA-2-Chat to affirmatively answer harmful queries

jailbreakfrom-arxiv
0» 0
167
[CASE-003]·ACTIVE·2mo·@mexiQQ

Claude Opus 4.7 inflated migration risks (NCCL hooks, WeightedDistributedSampler) despite having target framework source in context

hallucinationalignmentagent-misbehaviormotivated-reasoning
0» 0
168
[CASE-002]·ACTIVE·2mo·@mexiQQ

Claude Opus 4.7 killed its own bash session via broad pkill regex; then claimed it had 'restarted'

tool-misusehallucinationdestructive-actionagent-misbehavior
2» 0
169
[CASE-001]·ACTIVE·2mo·@mexiQQ

[META] First real case — testing the pipeline

unreviewed
0» 0