个人简介
我现任华为香港研究所 2012 实验室高级研究员,主要从事高效大模型推理与人工智能系统研究。此前曾任香港中文大学计算机科学与工程系博士后。我在香港中文大学(深圳)于两年半内完成计算机科学与工程博士学位,并获校长杰出博士研究奖;硕士阶段获欧盟奖学金,在荷兰埃因霍温理工大学和德国柏林工业大学完成多核系统双硕士项目;本科毕业于华中科技大学光电学院,期间获国家留学基金委资助赴德国亚琛工业大学交换学习。
我的研究聚焦于高效人工智能架构与系统,尤其关注算法、系统与硬件之间的协同设计。
研究方向
- 高效大模型架构与推理
- 低比特量化、稀疏与线性模型,以及高吞吐、低开销的大模型推理。
- 人工智能系统与软硬件协同
- 面向 GPU 与 NPU 的算子、访存和执行优化。
- 稀疏计算与图智能
- 稀疏矩阵计算、并行图算法与图神经网络系统。
学术论文
Sparsity in Linear Attention Models: Bregman Testing Time Learning.
Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2026, to appear.
High-Throughput Non-uniformly Quantized 3-bit LLM Inference.
Proceedings of the 31st ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming (PPoPP), pp. 288–300, 2026.
Rebo: Locality-Aware Graph Processing via Reordering and Blocking.
Proceedings of the 40th IEEE International Parallel and Distributed Processing Symposium (IPDPS), pp. 1037–1049, 2026. Best Paper Nomination (3 of 550 papers)
ToT: Triangle Counting on Tensor Cores.
IEEE Transactions on Parallel and Distributed Systems (TPDS), vol. 36, no. 12, pp. 2679–2692, 2025.
Groot: Graph-Centric Row Reordering with Tree for Sparse Matrix Multiplications on Tensor Cores.
Proceedings of the 20th European Conference on Computer Systems (EuroSys), pp. 803–817, 2025.
POSTER: Triangle Counting on Tensor Cores.
Proceedings of the 30th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming (PPoPP), pp. 560–562, 2025.
Accelerating SpMV for Scale-Free Graphs with Optimized Bins.
Proceedings of the 40th IEEE International Conference on Data Engineering (ICDE), pp. 2407–2420, 2024.
Efficient SpMV for Graph Matrices Through Vectoring and Caching.
Proceedings of the 30th European Conference on Parallel and Distributed Processing (Euro-Par), LNCS vol. 14803, pp. 356–370, 2024.
Bitmap-Based Sparse Matrix-Vector Multiplication with Tensor Cores.
Proceedings of the 53rd International Conference on Parallel Processing (ICPP), pp. 1135–1144, 2024.
An Unequal Caching Strategy for Shared-Memory Graph Analytics.
IEEE Transactions on Parallel and Distributed Systems (TPDS), vol. 34, no. 3, pp. 955–967, 2023.
Connectivity-Aware Link Analysis for Skewed Graphs.
Proceedings of the 52nd International Conference on Parallel Processing (ICPP), pp. 482–491, 2023.
Workload Balancing via Graph Reordering on Multicore Systems.
IEEE Transactions on Parallel and Distributed Systems (TPDS), vol. 33, no. 5, pp. 1231–1245, 2022.
HiPa: Hierarchical Partitioning for Fast PageRank on NUMA Multicore Systems.
Proceedings of the 50th International Conference on Parallel Processing (ICPP), Article 24, pp. 1–10, 2021.
POSTER: Corder: Cache-Aware Reordering for Optimizing Graph Analytics.
Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (PPoPP), pp. 472–473, 2021.
教学经历
- 计算机体系结构Computer Architecture
- 操作系统Operating System
- 编译器设计Compiler Design
- 程序设计方法导论Introduction to Programming Methodology
技术笔记
关于现代 C++、图算法与性能优化的学习记录。