Profile

My background.

I am a Ph.D. candidate in Computer Science at City University of Hong Kong, advised by Prof. Jinhang Zuo. Previously, I worked as a research assistant with Prof. Sibo Wang at the Chinese University of Hong Kong from July to October 2024.

I received my bachelor’s degree in Computer Science and Technology from the School of the Gifted Young at the University of Science and Technology of China, where I was advised by Prof. Xue Chen and Prof. Jinhang Zuo.

News

Recent research updates.

  1. Our paper WEREWOLF: Reputation-Aware Red-Teaming for Self-Organizing LLM Multi-Agent Systems was accepted to Findings of EMNLP 2026.

  2. Our paper Fusing Reward and Dueling Feedback in Stochastic Bandits was accepted to ICML 2025.

View all news

Research Focus

I study sequential decision-making under limited, noisy, or hybrid feedback. My current work focuses on bandit learning and influence maximization.

  • Learning Theory
  • Influence Maximization
  • Game Theory
  • Multi-Agent Systems

Selected Publications

⋆ Equal contribution.

Best Arm Identification in Generalized Linear Bandits via Hybrid Feedback

Qirun Zeng, Xuchuang Wang, Jiayi Shen, Xutong Liu, Fang Kong, Jinhang Zuo

In Proceedings of the 40th Annual Conference on Neural Information Processing Systems

Practical Adversarial Attacks on Stochastic Bandits via Fake Data Injection

Qirun Zeng, Eric He, Richard Hoffmann, Xuchuang Wang, Jinhang Zuo

In Proceedings of the 40th Annual Conference on Neural Information Processing Systems

One Rounding Fits All: Memory-Efficient Approximation Algorithms for Partition-Constrained Influence Maximization

Qixin Zhang⋆, Qirun Zeng⋆, Hui Lu, Pingchuan Ma, Jinhang Zuo, Renqiang Luo, Yi Yu, Dacheng Tao

In Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining

View all publications

Service

Academic reviewing and service.

2026

Reviewer

ICLR and NeurIPS.

2025

Reviewer

ICML and NeurIPS.

Teaching

I have served as a teaching assistant for programming, data structures, algorithms, and mathematical analysis courses.

View teaching experience

10-second Bandit

Can you identify the better arm before your rounds run out?

Explore or exploit?

Two arms hide different reward probabilities. Learn which one is better before your rounds run out.

Round 0 / 10
Reward 0
Expected regret hidden
Reward trail

    This is the exploration–exploitation dilemma in miniature. See how I study bandit feedback