avatar

Youhui Bai, 白有辉

Yesterday is a history, tomorrow is a mystery, but today is a gift, that is why it is called "Present".

There is no tomorrow.

Publications

* denotes corresponding author. · Google Scholar

2026

  1. OSDI
    Best Paper

    Teaching the Old Dog New Tricks: Building Efficient Data Pipelines for Large-Scale LLM Pre-Training (Operational Systems)

    20th USENIX Symposium on Operating Systems Design and Implementation (OSDI '26), July 2026.

  2. TPDS

    nScaler-M: Constraint-Guided and Placement-Aware Parallelization Plan Generation for Deep Learning Training

    IEEE Transactions on Parallel and Distributed Systems, July 2026.

  3. TMLR

    CentroidKV: Efficient Long-Context LLM Inference via KV Cache Clustering

    Transactions on Machine Learning Research, June 2026.

  4. CVPR

    AdaCluster: Adaptive Query-Key Clustering for Sparse Attention in Video Generation

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, June 2026.

  5. ICPP

    Accelerating Long-Tail Generation in Synchronous RLHF Training via Adaptive Tensor Parallelism

    International Conference on Parallel Processing (ICPP), 2026.

  6. AAAI

    SMIDT: High-Performance Inference Framework for MoE Models with Dynamic Top-K Routing

    Proceedings of the AAAI Conference on Artificial Intelligence, March 2026.

  7. arXiv

    Lagom: Unleashing the Power of Communication and Computation Overlapping for Distributed LLM Training

    arXiv preprint arXiv:2602.20656, February 2026.

2025

  1. arXiv

    CLO: Efficient LLM Inference System with CPU-Light KVCache Offloading via Algorithm-System Co-Design

    arXiv preprint arXiv:2511.14510, November 2025.

  2. arXiv

    Efficient Long-Context LLM Inference via KV Cache Clustering

    arXiv preprint arXiv:2506.11418, June 2025.

  3. ACL

    HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference

    Findings of the Association for Computational Linguistics, May 2025.

  4. AAAI

    BigMac: A Communication-Efficient Mixture-of-Experts Model Structure for Fast Training and Inference

    Proceedings of the AAAI Conference on Artificial Intelligence, March 2025.

2024

  1. arXiv

    XL3M: A Training-free Framework for LLM Length Extension Based on Segment-wise Inference

    arXiv preprint arXiv:2405.17755, May 2024.

2023

  1. TPDS

    A Survey on Auto-Parallelism of Neural Networks Training

    IEEE Transactions on Parallel and Distributed Systems, May 2023.

  2. TPDS

    A Generic, High-Performance, Compression-Aware Framework for Data Parallel DNN Training

    IEEE Transactions on Parallel and Distributed Systems, April 2023.

  3. HPCA

    MPress: Democratizing Billion-Scale Model Training on Multi-GPU Servers via Memory-Saving Inter-Operator Parallelism

    IEEE International Symposium on High-Performance Computer Architecture, February 2023.

2021

  1. SOSP

    Gradient Compression Supercharged High-Performance Data Parallel DNN Training

    Proceedings of the ACM Symposium on Operating Systems Principles, October 2021.

  2. GNNSys

    Efficient Data Loader for Fast Sampling-based GNN Training on Large Graphs

    GNNSys Workshop, April 2021.

  3. TPDS

    Efficient Data Loader for Fast Sampling-based GNN Training on Large Graphs

    IEEE Transactions on Parallel and Distributed Systems, March 2021.

2017

  1. SOSP Poster

    Fast Logging and Recovery Support for Transactional Databases

    SOSP Poster, Shanghai, China, October 2017.

  2. ICPP

    PDS: An I/O-efficient Scaling Scheme for Parity Declustered Data Layout

    International Conference on Parallel Processing, Bristol, UK, August 2017.