@Hiki
I am a Cognitive Neuroscientist transitioning into AI safety.
https://www.linkedin.com/in/hikaru-tsujimura/$0 in pending offers
I am a Cognitive Neuroscientist transitioning into AI safety, with one and half year of hands-on research in interpretability and model behavior analysis. My work focuses on understanding how internal representations in LLMs shape values, decision-making, and actions, analogous to human cognition.
Recent work (2025-2026) includes:
Mechanistic interpretability (SAE, TransformerLens) to visualize latent concepts and reasoning structures
Analysis of overconfidence and unfaithful Chain-of-Thought (CoT) reasoning in LLMs
LLM persona elicitation and characterization of latent behaviors (e.g. Sydney-style behaviors)
AI security and red-teaming challenges, including testing frontier proprietary models
Identifying persistent vulnerabilities across Qwen models (2.5-3.5) in their CoT reasoning and 50-90% vulnerability improvements with steering techniques
My long-term goal is to develop interpretable, transparent, trustworthy and scalable AI models.
See further work in my personal webiste ( https://hiki-t.github.io/ ).