SYS:ONLINELAT:n/aBUILD:8161faf
[CASE-134]·STATUS:ACTIVE·OPENED:2026-08-07·UPDATED:2026-08-07

Qwen model family shows 4-fold bias inflation for real vs. fictional country pairs in China-related scenarios

submitted_by:@mexiQQ
alignmentfrom-arxivauto-published
cat case_body.md

Auto-published from arXiv:2608.06123 by the mine-arxiv pipeline. Reviewed by an LLM judge (Sonnet) against the archive bar — see CONTRIBUTING. Notes: cleared review (confidence 0.72, flags: [no-prompt-excerpt])

Category

alignment

Model

Qwen3.5-27B, Qwen3-32B, Qwen3.5-122B, Qwen3.5-397B

Surface

API (OpenRouter, temperature 0.3)

Setup

Even-handedness evaluation comparing eight fictional country pairs (control, used to establish baseline) against eight real active-conflict pairs and eight tension pairs. Fictional pairs provide a no-prior-knowledge baseline; real pairs include China–Taiwan and China–US. Bias is measured via the Poli-Bias Index (PBI, range 0–15) across five dimensions. The paper does not provide a verbatim prompt excerpt for this specific finding; Appendix A contains the general template structure.

Reproduction anchor

Appendix A (prompt templates) and Appendix G (judge validation). Specific country-pair lists are described in the body. Code not released; authors state open-sourcing upon acceptance.

Observed behavior

All four Qwen models showed a ~4-fold increase in PBI scores when transitioning from fictional to real country pairs, with the gap concentrated in China-related scenarios (China–Taiwan, China–US). Qwen3.5-27B achieved the highest overall PBI of 5.89 — nearly double the best-performing model (Grok-4.5 at 3.36). Larger Qwen models did not close the gap the way larger models in other families did.

Expected behavior

Bias scores should be broadly similar between fictional and real country pairs if the model is evaluating legal acts rather than country-specific priors. Scale increases within a model family should monotonically reduce bias.

Reproducibility

low

Threat model

Deployments of Qwen models for international news summarization, conflict analysis, or automated legal commentary involving China will systematically produce China-favorable (or at least inconsistent) outputs. Given Qwen's wide open-weight adoption, this affects many downstream fine-tuned variants and applications that inherit the base model's political priors.

Novelty

Demonstrates that RLHF or instruction-tuning in Qwen models appears to encode country-specific political sensitivities that do not diminish with scale — a qualitatively different failure mode from the scale-reduces-bias trend seen in other families.

Source

Triage notes (auto)

  • paperType: benchmark
  • estimatedCaseCount: 3
  • triage reason: Systematic benchmark measuring political bias and sycophancy failures across 13 LLMs using controlled counterfactual prompts that swap country identities in legally equivalent scenarios, demonstrating reproducible disparities in how models describe and defend equivalent actions.
tail -f comments.log

0 comments

─────────────────────────────────────────────────────────────────────

// no comments yet