Zhiyuan Xu
I am a third-year Ph.D. student in the Cyber Security Group at the University of Bristol. My research focuses on the security and safety of large language models, particularly how hidden model behaviour can be exposed, manipulated, and better understood.
My research centers on LLM safety and the mechanistic interpretability of alignment. I study chain-of-thought leakage, fine-tuning and activation-space attacks, Mixture-of-Experts routing vulnerabilities, safety-neuron-guided fuzzing, and privacy attacks against machine learning systems. I am supervised by Lichao Wu, Joseph Gardiner, and Sana Belguith.
Before starting my Ph.D., I worked as a Research Assistant at the Hong Kong University of Science and Technology (HKUST) and completed an M.Sc. in Financial Technology with Data Science at Bristol. I also bring industry experience from backend and AI systems roles at ByteDance and Huawei.
Research interests
Highlights
- Best Paper Award, ACM ASIA CCS 2026, for work on hidden threats in chain-of-thought models.
- Principal Investigator, Mechanistic Analysis of LLM Safety Failures under Jailbreak and Prompt-Injection (2026–2027), funded by RISCS and the UK National Cyber Security Centre.
- 1st Place, UK Ministry of Justice AI Debiasing Hackathon (2025).
What’s New
- Aug 2026 Progressive Behavioral Drift through Compression Valleys in Large Language Models Findings of the Association for Computational Linguistics: EMNLP 2026
- Aug 2026 NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation arXiv preprint arXiv:2608.26222
- May 2026 RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs arXiv preprint arXiv:2605.02946
Services
- Reviewer: IEEE S&P; ESORICS.
- External Reviewer: S&P; USENIX Security; NDSS.