You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I want to benchmark and mechanistically address how compositional task interference within an LLM's safety-relevant refusal behaviours varies by persona framing and contextual integrity. Concretely, it is shown that persona/role-assignment framing reduces jailbreak refusal rates by 50-70% (Zhang et al., 2026), and even framing-agnostic diagnostics showed LLMs frequently disclose context-inappropriate information (Mireshghallah et al., 2024). Conversely, perceived legitimacy and signalled context are also evidenced as known levers for over-refusal, with resources like FalseReject being developed to mitigate unnecessary refusals. However, I believe a persistent tension exists between finetunings and context framings that garner enough trust to prevent over-refusal, while still preserving standards of data security and content integrity. When judgements about confidentiality and context-appropriateness are co-located with auxiliary task prompts, LLMs become fragile to both under-refusal and over-refusal tendencies. For example, pilot studies on a Japanese-language RAG pipeline showed that several LLMs leaked personally-identifiable information (PII) in 5/6 confidentiality tests when framings were permissive. However, when asked to evaluate additional tests grammatical naturalness, given no persona framing, an ensemble of those same LLMs entirely refused to grade 10/35 test cases for fear that they contained PII (the data was in reality synthetic). Mitigating the problem of compliance in leaking PII might create over-refusals for tasks simply relating to PII, and vice versa. Thus, in such cases of task co-location, the model’s safety decision boundary (as is termed by Pan et al.) can be pulled both towards over- and under-refusals in a single generation or session. This compositional interference is currently understudied (Anonto et al.) given that existing benchmarks are mostly single-prompt or single-task. Documenting, isolating, and steering dynamic mechanisms of persona and context framing across task types would stress-test common assumptions of LLM judgement decomposability (e.g. within LLM-as-judge, automated red-teaming, and related tasks), demonstrating important nuances for future scalable oversight pipelines.
Utilising methodologies from OR-Bench, Arditi et al., 2024, and Zhong and Li, 2026, I seek to implement the following:
Compositional interference dataset, testing confidentiality judgements against ablations of varying tasks and framings
Behavioural benchmark to measure how disclosure and refusal shift under compositional framing on all models, across single- and multi-turn interactions
Interpretability and activation patching/steering experiments to extract and study refusal and context/legitimacy directions on open models
The main goal is to determine how LLM safety judgements compose across simultaneous sub-tasks, whether additively, with structured interference, with unstructured interference, or otherwise. Further, given interference patterns, I want to evaluate whether disclosure resistance and engagement share an underlying cause (e.g. legitimacy framing) or several independent failure modes. Since LLM evaluation pipelines are becoming common, especially for under-resourced, non-frontier, and non-English enterprises, it becomes imperative to evaluate the internal consistency and trustworthiness of deployment given realistic, multi-task pipelines. By isolating and characterising this effect, I hope to:
Identify via correlation and causal mechanism what governs pattern shifts in under- and over-refusal behaviours, using RAG deployment (a common corporate task) as a testbed
Establish a benchmark for testing LLM behavior under simultaneous, competing safety-relevant task demands
Provide actionable characterisations for realistic and currently-deployed systems, informing Japanese safety practices and inviting future expansions (e.g. into English and other languages, into agentic tooling for full information-flow control, etc.)
Minimum funding tier:
I will allocate $2000 to cover basic computing costs, running the compositional study across a small set of proprietary models. I will use $1000 to compensate human verification efforts in Japanese and English, as well as for an abbreviated stipend to assist with living and travel costs related to coworking within my host company.
Full funding tier:
I will allocate $5000 to cover extensive compute costs, constructing and diversifying a larger behavioural benchmark. I will use $4000 to compensate human verification efforts in Japanese and English, as well as for a standard stipend to assist with living and travel costs related to coworking within my host company.
I am Troy Hanfei Tian, a rising sophomore undergraduate student at the University of California, Los Angeles, concentrating in Computer Science and Asian American Studies. I’m currently being hosted by NanoBase, K. K., a systems integration company based in Osaka, Japan, that grants basic in-kind hardware access as a technology provider for my independent research efforts. In terms of past projects, at the Computation and Language for Society Lab at UCLA, I presented and plan to eventually submit for publication a project addressing evaluation methodology for VLM-generated image captions. I gained evaluation experience in demonstrating comparative bias in favour of sighted as opposed to blind populations within LLM-as-judge methods, and proved a nuanced three-axis decomposition of description quality. I also completed and presented a SPAR metaresearch project developing evaluation methodology for the capabilities spillover risk of various AI safety projects based on past data and case studies.
The most likely causes and associated outcomes of this project failing include:
Compositional study showing a non-clean decomposition, e.g. a correlation unsupported by mechanistic techniques, meaning further review would be required to explain observed phenomena
Results failing to generalise beyond Japanese language, meaning any prescriptions mostly only apply to Japanese businesses
Results failing to generalise beyond RAG security, meaning alignment only advances in a specific field which itself is moving away from LLMs as a security layer
LLM-judge reliability disagreeing with human validation or otherwise undermining results even using ensemble and debate techniques, meaning that approach may not be resolvably scalable
Running out of compute budget, meaning that I would need to secure support from my university or other sources
None.
There are no bids on this project.