EN

个人简介

我现任华为港研所 2012 实验室高级研究员,从事高效大模型推理研究。此前在香港中文大学计算机科学与工程系担任博士后。我在香港中文大学(深圳)用两年半时间完成计算机科学与工程博士学位,并获校长杰出博士研究奖;硕士阶段获欧盟创新科技奖学金,在荷兰埃因霍温理工大学和德国柏林工业大学完成多核并行系统双硕士;本科毕业于华中科技大学电子科学与技术专业。

我的研究兴趣主要是高效人工智能架构与系统,关注大模型推理、稀疏计算与软硬件协同优化,致力于让人工智能模型在实际计算硬件上高效运行。

研究涵盖低比特量化、GPU 与 NPU 上的稀疏算子加速,以及图计算与图神经网络系统。我也关注现代 C++、内存布局和代码性能,将算法设计与体系结构优化结合起来。

研究方向

高效大模型架构与推理
低比特量化、稀疏与线性注意力模型,以及面向实际硬件的高吞吐推理。
人工智能系统与软硬件协同
面向 GPU、Tensor Cores 与昇腾 NPU 的算子、访存和执行优化,将模型与算法特性映射到硬件架构。
稀疏计算与图智能
稀疏矩阵运算、并行图算法与图神经网络系统,为高效人工智能提供计算基础。

学术论文

  1. Sparsity in Linear Attention Models: Bregman Testing Time Learning.

    Wenqi Zeng*, YuAng Chen*, Yuxuan Chen, Weihuang Wen, Chumin Sun, Yichuan Liu, Li Zhou, Tian Wang, Fan Zhang, Yuan Yao, and Jie Sun

    Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2026, to appear.

    * Equal contribution.

  2. High-Throughput Non-uniformly Quantized 3-bit LLM Inference.

    YuAng Chen, Wenqi Zeng, and Jeffrey Xu Yu

    Proceedings of the 31st ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming (PPoPP), pp. 288–300, 2026.

  3. Rebo: Locality-Aware Graph Processing via Reordering and Blocking.

    YuAng Chen and Yeh-Ching Chung

    Proceedings of the 40th IEEE International Parallel and Distributed Processing Symposium (IPDPS), pp. 1037–1049, 2026. Best Paper Nomination (3 of 550 papers)

  4. ToT: Triangle Counting on Tensor Cores.

    YuAng Chen and Jeffrey Xu Yu

    IEEE Transactions on Parallel and Distributed Systems (TPDS), vol. 36, no. 12, pp. 2679–2692, 2025.

  5. Groot: Graph-Centric Row Reordering with Tree for Sparse Matrix Multiplications on Tensor Cores.

    YuAng Chen, Jiadong Xie, Siyi Teng, Wenqi Zeng, and Jeffrey Xu Yu

    Proceedings of the 20th European Conference on Computer Systems (EuroSys), pp. 803–817, 2025.

  6. POSTER: Triangle Counting on Tensor Cores.

    YuAng Chen and Jeffrey Xu Yu

    Proceedings of the 30th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming (PPoPP), pp. 560–562, 2025.

  7. Accelerating SpMV for Scale-Free Graphs with Optimized Bins.

    YuAng Chen and Jeffrey Xu Yu

    Proceedings of the 40th IEEE International Conference on Data Engineering (ICDE), pp. 2407–2420, 2024.

  8. Efficient SpMV for Graph Matrices Through Vectoring and Caching.

    YuAng Chen and Jeffrey Xu Yu

    Proceedings of the 30th European Conference on Parallel and Distributed Processing (Euro-Par), LNCS vol. 14803, pp. 356–370, 2024.

  9. Bitmap-Based Sparse Matrix-Vector Multiplication with Tensor Cores.

    YuAng Chen and Jeffrey Xu Yu

    Proceedings of the 53rd International Conference on Parallel Processing (ICPP), pp. 1135–1144, 2024.

  10. An Unequal Caching Strategy for Shared-Memory Graph Analytics.

    YuAng Chen and Yeh-Ching Chung

    IEEE Transactions on Parallel and Distributed Systems (TPDS), vol. 34, no. 3, pp. 955–967, 2023.

  11. Connectivity-Aware Link Analysis for Skewed Graphs.

    YuAng Chen and Yeh-Ching Chung

    Proceedings of the 52nd International Conference on Parallel Processing (ICPP), pp. 482–491, 2023.

  12. Workload Balancing via Graph Reordering on Multicore Systems.

    YuAng Chen and Yeh-Ching Chung

    IEEE Transactions on Parallel and Distributed Systems (TPDS), vol. 33, no. 5, pp. 1231–1245, 2022.

  13. HiPa: Hierarchical Partitioning for Fast PageRank on NUMA Multicore Systems.

    YuAng Chen and Yeh-Ching Chung

    Proceedings of the 50th International Conference on Parallel Processing (ICPP), Article 24, pp. 1–10, 2021.

  14. POSTER: Corder: Cache-Aware Reordering for Optimizing Graph Analytics.

    YuAng Chen and Yeh-Ching Chung

    Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (PPoPP), pp. 472–473, 2021.

教学经历

  • 计算机体系结构Computer Architecture
  • 操作系统Operating System
  • 编译器设计Compiler Design
  • 程序设计方法导论Introduction to Programming Methodology

技术笔记

关于现代 C++、图算法与性能优化的学习记录。