About me
Hi! I’m a fourth year Ph.D. student in the Department of Computer Science and Engineering at UC San Diego, in the Trustworthy ML Lab. I’m fortunate to be advised by Prof. Tsui-Wei (Lily) Weng. My research interest is in trustworthy machine learning and responsible AI, with a focus on mechanistic interpretability, uncertainty quantification, and robust prediction. Recently, I have been using representation engineering and model editing to understand and improve LLM reasoning, reflection, and tool use. My CV could be found here.
Education
Ph.D. in the Department of Computer Science and Engineering at UC San Diego. [Sep. 2023 - Current]
M.S. in Machine Learning and Data Science, Department of Electrical and Computer Engineering at UC San Diego. [Sep. 2021 - Mar. 2023]
B.S. in Information and Computing Science, School of Mathematical Sciences at Peking University. [Sep. 2017 - Jun. 2021]
Publications
* Equal contribution.
Conference and Journal Papers
ReflCtrl: Controlling LLM Reflection Efficiently via Representation Engineering
Ge Yan, Chung-En Sun, Linbo Liu, Tsui-Wei (Lily) Weng, COLM 2026. Also NeurIPS 2025 MI Workshop (Spotlight). [Project Page] [Code]
Provably Robust Conformal Prediction with Improved Efficiency
Ge Yan, Yaniv Romano, Tsui-Wei (Lily) Weng, ICLR 2024.
VLG-CBM: Training Concept Bottleneck Models with Vision-Language Guidance
Divyansh Srivastava*, Ge Yan*, Tsui-Wei (Lily) Weng, NeurIPS 2024. [Project Page] [Code]
Steer2Edit: From Activation Steering to Component-Level Editing
Chung-En Sun, Ge Yan, Zimo Wang, Tsui-Wei (Lily) Weng, NeurIPS 2026.
LLM Agents Already Know When to Call Tools — Even Without Reasoning
Chung-En Sun, Linbo Liu, Ge Yan, Zimo Wang, Tsui-Wei (Lily) Weng, NeurIPS 2026.
Distance Marching for Generative Modeling
Zimo Wang, Ishit Mehta, Haolin Lu, Chung-En Sun, Ge Yan, Tsui-Wei (Lily) Weng, Tzu-Mao Li, NeurIPS 2026.
ReFIne: A Framework for Trustworthy Large Reasoning Models with Reliability, Faithfulness, and Interpretability
Chung-En Sun, Ge Yan, Akshay Kulkarni, Tsui-Wei (Lily) Weng, COLM 2026.
Multimodal Concept Bottleneck Models
Tongqing Shi, Ge Yan, Tuomas Oikarinen, Tsui-Wei (Lily) Weng, TMLR 2026.
Beyond Top Activations: Efficient and Reliable Crowdsourced Evaluation of Automated Interpretability
Tuomas Oikarinen, Ge Yan, Akshay Kulkarni, Tsui-Wei (Lily) Weng, CVPR 2026.
ThinkEdit: Interpretable Weight Editing to Mitigate Overly Short Thinking in Reasoning Models
Chung-En Sun, Ge Yan, Tsui-Wei (Lily) Weng, EMNLP 2025.
Evaluating Neuron Explanations: A Unified Framework with Sanity Checks
Tuomas Oikarinen, Ge Yan, Tsui-Wei (Lily) Weng, ICML 2025.
Interpretable Generative Models through Post-hoc Concept Bottlenecks
Akshay Kulkarni, Ge Yan, Chung-En Sun, Tuomas Oikarinen, Tsui-Wei (Lily) Weng, CVPR 2025.
Workshop Papers and Preprints
Faithful and Stable Neuron Explanations for Trustworthy Mechanistic Interpretability
Ge Yan, Tuomas Oikarinen, Tsui-Wei (Lily) Weng, NeurIPS 2025 MI Workshop.
RAT: Boosting Misclassification Detection Ability without Extra Data
Ge Yan, Tsui-Wei (Lily) Weng, arXiv preprint, 2025.
Industry Experience
Machine Learning Engineer Intern, Stripe, South San Francisco. [Jun. 2026 - Sep. 2026]
Applied Scientist Intern, Amazon, San Diego. [Jun. 2024 - Aug. 2024]
Data Scientist Intern, DiDi Technology, Beijing. [Jun. 2023 - Aug. 2023]
Academic Service and Talks
Area Chair: ICLR 2026 Trustworthy AI Workshop; CVPR 2026 TRUE-V Workshop.
Reviewer: ICLR 2025, 2026; NeurIPS 2025, 2026; ICML, CVPR, ECCV, and COLM 2026.
Tutorial Co-organizer: Principled Interpretability in Vision Models, CVPR 2026.
Invited Talk: Faithful Interpretation for Deep Networks via Human Understandable Concepts, EnCORE Workshop on Interpretability in Modern AI, 2026.
