中文

About

I am a Senior Researcher at Huawei Hong Kong Research Center (2012 Laboratories), working on efficient LLM inference. Previously, I was a postdoctoral researcher in the Department of Computer Science and Engineering at The Chinese University of Hong Kong. I completed my Ph.D. in Computer Science and Engineering at The Chinese University of Hong Kong, Shenzhen, in two and a half years, and received the Presidential Award for Outstanding Doctoral Students. I earned dual M.Sc. degrees in Multicore Systems from Eindhoven University of Technology and Technische Universität Berlin with a full European Union scholarship, and a B.Eng. in Electronic Science and Technology from Huazhong University of Science and Technology.

My research focuses on efficient AI architectures and systems, including LLM inference, sparse computing, and hardware–software co-design, with the goal of making AI models run efficiently on real hardware.

My work spans low-bit quantization, sparse kernel acceleration on GPUs and NPUs, and graph processing and graph neural network systems. I am also interested in modern C++, memory layout, and code performance, bridging algorithm design and architecture-level optimization.

Research

Efficient LLM Architectures and Inference
Low-bit quantization, sparse and linear attention models, and high-throughput inference on real hardware.
AI Systems and Hardware–Software Co-design
Kernel, memory-access, and execution optimization on GPUs, Tensor Cores, and Ascend NPUs, mapping model and algorithm characteristics onto hardware architectures.
Sparse Computing and Graph Intelligence
Sparse matrix operations, parallel graph algorithms, and graph neural network systems as computational foundations for efficient AI.

Publications

  1. Sparsity in Linear Attention Models: Bregman Testing Time Learning.

    Wenqi Zeng*, YuAng Chen*, Yuxuan Chen, Weihuang Wen, Chumin Sun, Yichuan Liu, Li Zhou, Tian Wang, Fan Zhang, Yuan Yao, and Jie Sun

    Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2026, to appear.

    * Equal contribution.

  2. High-Throughput Non-uniformly Quantized 3-bit LLM Inference.

    YuAng Chen, Wenqi Zeng, and Jeffrey Xu Yu

    Proceedings of the 31st ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming (PPoPP), pp. 288–300, 2026.

  3. Rebo: Locality-Aware Graph Processing via Reordering and Blocking.

    YuAng Chen and Yeh-Ching Chung

    Proceedings of the 40th IEEE International Parallel and Distributed Processing Symposium (IPDPS), pp. 1037–1049, 2026. Best Paper Nomination (3 of 550 papers)

  4. ToT: Triangle Counting on Tensor Cores.

    YuAng Chen and Jeffrey Xu Yu

    IEEE Transactions on Parallel and Distributed Systems (TPDS), vol. 36, no. 12, pp. 2679–2692, 2025.

  5. Groot: Graph-Centric Row Reordering with Tree for Sparse Matrix Multiplications on Tensor Cores.

    YuAng Chen, Jiadong Xie, Siyi Teng, Wenqi Zeng, and Jeffrey Xu Yu

    Proceedings of the 20th European Conference on Computer Systems (EuroSys), pp. 803–817, 2025.

  6. POSTER: Triangle Counting on Tensor Cores.

    YuAng Chen and Jeffrey Xu Yu

    Proceedings of the 30th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming (PPoPP), pp. 560–562, 2025.

  7. Accelerating SpMV for Scale-Free Graphs with Optimized Bins.

    YuAng Chen and Jeffrey Xu Yu

    Proceedings of the 40th IEEE International Conference on Data Engineering (ICDE), pp. 2407–2420, 2024.

  8. Efficient SpMV for Graph Matrices Through Vectoring and Caching.

    YuAng Chen and Jeffrey Xu Yu

    Proceedings of the 30th European Conference on Parallel and Distributed Processing (Euro-Par), LNCS vol. 14803, pp. 356–370, 2024.

  9. Bitmap-Based Sparse Matrix-Vector Multiplication with Tensor Cores.

    YuAng Chen and Jeffrey Xu Yu

    Proceedings of the 53rd International Conference on Parallel Processing (ICPP), pp. 1135–1144, 2024.

  10. An Unequal Caching Strategy for Shared-Memory Graph Analytics.

    YuAng Chen and Yeh-Ching Chung

    IEEE Transactions on Parallel and Distributed Systems (TPDS), vol. 34, no. 3, pp. 955–967, 2023.

  11. Connectivity-Aware Link Analysis for Skewed Graphs.

    YuAng Chen and Yeh-Ching Chung

    Proceedings of the 52nd International Conference on Parallel Processing (ICPP), pp. 482–491, 2023.

  12. Workload Balancing via Graph Reordering on Multicore Systems.

    YuAng Chen and Yeh-Ching Chung

    IEEE Transactions on Parallel and Distributed Systems (TPDS), vol. 33, no. 5, pp. 1231–1245, 2022.

  13. HiPa: Hierarchical Partitioning for Fast PageRank on NUMA Multicore Systems.

    YuAng Chen and Yeh-Ching Chung

    Proceedings of the 50th International Conference on Parallel Processing (ICPP), Article 24, pp. 1–10, 2021.

  14. POSTER: Corder: Cache-Aware Reordering for Optimizing Graph Analytics.

    YuAng Chen and Yeh-Ching Chung

    Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (PPoPP), pp. 472–473, 2021.

Teaching

  • Computer Architecture
  • Operating System
  • Compiler Design
  • Introduction to Programming Methodology

Notes

Learning notes on modern C++, graph algorithms, and performance optimization.