Professional Experience

TokenHub / MaaS Platform Development Engineer

Mar 2026 - Present
Tencent Cloud
Build Tencent Cloud TokenHub MaaS infrastructure for unified LLM API Gateway, model access, request governance, capacity protection, and production observability.

Container Networking R&D Engineer

Jul 2024 - Mar 2026
Tencent Cloud
Built and maintained container networking and cloud-native traffic entry infrastructure across VPC CNI, Kubernetes Service, Ingress, and Gateway capabilities.

Education

M.S. in Network Engineering

2021 - 2024
University of Electronic Science and Technology of China
Graduate study in network engineering and computer systems.

B.S. in Network Engineering

2017 - 2021
University of Electronic Science and Technology of China
Undergraduate study in network engineering.

Profile

M.S. in Network Engineering, AI Infrastructure / LLM Gateway engineer, and core committer of vllm-project/semantic-router. I work on Tencent Cloud TokenHub, building MaaS infrastructure for unified model access, request governance, tenant isolation, quotas, observability, and reliability.

My open-source work focuses on intelligent semantic routing across request semantics, model capability matching, safety governance, semantic cache, feedback learning, and cost / latency / quality-aware routing.

Open Source

vllm-project/semantic-router

Core committer of vllm-project/semantic-router, an intelligent semantic routing project in the official vLLM ecosystem for Mixture-of-Models, AI Gateway, and production LLM traffic governance.

  • Extract routing signals from prompts, context, request parameters, tenant metadata, and historical feedback for intent recognition, task classification, model capability matching, and policy selection.
  • Design routing policies, canary validation, and fallback strategies across semantic signals, cost, latency, quality, safety level, context length, and cache hit rate.
  • Work with gateway-level prompt injection, jailbreak, PII, and content-safety checks, as well as semantic cache and prefix / KV-cache-aware routing for cost, latency, and throughput.
  • Feed user feedback, model outputs, latency, error rate, and cost signals back into routing decisions.

I am also a member of Agentic Intelligence Lab and the Semantic Router GitHub organization.

Core Skills

  • MaaS platforms / LLM API Gateway: unified model service entry, OpenAI / Anthropic / Responses-compatible APIs, multi-provider and self-hosted inference backend integration, authentication, tenant isolation, quota metering, rate limiting, fallback, SLOs, auditing, and observability.
  • Intelligent semantic routing / model selection / safety: request semantics, intent recognition, task classification, model capability matching, cost / latency / quality-aware routing, safety checks, semantic cache, prefix / KV-cache-aware routing, canary validation, fallback, and feedback learning.
  • Networking / container networking / engineering stack: TCP/IP, QUIC, Linux networking, VPC CNI, Pod networking, Service, Ingress, Gateway API, load balancing, service discovery, routing and forwarding, connectivity troubleshooting, Go, Python, and gRPC.