AI Infrastructure · Inference Serving · GPU Scheduling
Communications engineering by training, large-model inference infrastructure by trade.
MSc @ BUPT · Class of 2028
Beijing, China
- Xiaohongshu (RED) · AI Infrastructure Intern · Sep 2026 – Present
Capacity management · time-shifted scheduling · elastic scaling - Kuaishou · AI Infrastructure Intern · Mar 2026 – Sep 2026
Inference scheduling · GPU resource efficiency · confidential inference - DiDi · Backend Intern · Sep 2025 – Feb 2026
Benefits platform refactor · multi-agent test case generation
All involved internal systems, so I don't go into detail here.
- Inference serving & resource efficiency — autoscaling, capacity normalization across heterogeneous accelerators, time-shifted scheduling with offline reclamation
- Scheduling & load balancing — traffic distribution in distributed systems, graceful degradation, state consistency
- Confidential computing — remote attestation and key management under TEEs, GPU confidential computing
- Cloud native — Kubernetes, capacity management and autoscaling on hybrid cloud
Notes and source walkthroughs on inference infrastructure.