Zhiyuan Xu

I am a third-year Ph.D. student in the Cyber Security Group at the University of Bristol. My research focuses on the security and safety of large language models, particularly how hidden model behaviour can be exposed, manipulated, and better understood.

My research centers on LLM safety and the mechanistic interpretability of alignment. I study chain-of-thought leakage, fine-tuning and activation-space attacks, Mixture-of-Experts routing vulnerabilities, safety-neuron-guided fuzzing, and privacy attacks against machine learning systems. I am supervised by Lichao Wu, Joseph Gardiner, and Sana Belguith.

Before starting my Ph.D., I worked as a Research Assistant at the Hong Kong University of Science and Technology (HKUST) and completed an M.Sc. in Financial Technology with Data Science at Bristol. I also bring industry experience from backend and AI systems roles at ByteDance and Huawei.

Research interests

LLM safety Mechanistic interpretability Red teaming Activation engineering Mixture-of-Experts security

Highlights

  • Best Paper Award, ACM ASIA CCS 2026, for work on hidden threats in chain-of-thought models.
  • Principal Investigator, Mechanistic Analysis of LLM Safety Failures under Jailbreak and Prompt-Injection (2026–2027), funded by RISCS and the UK National Cyber Security Centre.
  • 1st Place, UK Ministry of Justice AI Debiasing Hackathon (2025).

What’s New

Explore all research

Services

  • Reviewer: IEEE S&P; ESORICS.
  • External Reviewer: S&P; USENIX Security; NDSS.