Resume
Work experience, open-source identity, and core technical direction.
Professional Experience
TokenHub / MaaS Platform Development Engineer
Mar 2026 - PresentContainer Networking R&D Engineer
Jul 2024 - Mar 2026Education
M.S. in Network Engineering
2021 - 2024B.S. in Network Engineering
2017 - 2021Profile
M.S. in Network Engineering, AI Infrastructure / LLM Gateway engineer, and core committer of vllm-project/semantic-router. I work on Tencent Cloud TokenHub, building MaaS infrastructure for unified model access, request governance, tenant isolation, quotas, observability, and reliability.
My open-source work focuses on intelligent semantic routing across request semantics, model capability matching, safety governance, semantic cache, feedback learning, and cost / latency / quality-aware routing.
Open Source
vllm-project/semantic-router
Core committer of vllm-project/semantic-router, an intelligent semantic routing project in the official vLLM ecosystem for Mixture-of-Models, AI Gateway, and production LLM traffic governance.
- Extract routing signals from prompts, context, request parameters, tenant metadata, and historical feedback for intent recognition, task classification, model capability matching, and policy selection.
- Design routing policies, canary validation, and fallback strategies across semantic signals, cost, latency, quality, safety level, context length, and cache hit rate.
- Work with gateway-level prompt injection, jailbreak, PII, and content-safety checks, as well as semantic cache and prefix / KV-cache-aware routing for cost, latency, and throughput.
- Feed user feedback, model outputs, latency, error rate, and cost signals back into routing decisions.
I am also a member of Agentic Intelligence Lab and the Semantic Router GitHub organization.
Core Skills
- MaaS platforms / LLM API Gateway: unified model service entry, OpenAI / Anthropic / Responses-compatible APIs, multi-provider and self-hosted inference backend integration, authentication, tenant isolation, quota metering, rate limiting, fallback, SLOs, auditing, and observability.
- Intelligent semantic routing / model selection / safety: request semantics, intent recognition, task classification, model capability matching, cost / latency / quality-aware routing, safety checks, semantic cache, prefix / KV-cache-aware routing, canary validation, fallback, and feedback learning.
- Networking / container networking / engineering stack: TCP/IP, QUIC, Linux networking, VPC CNI, Pod networking, Service, Ingress, Gateway API, load balancing, service discovery, routing and forwarding, connectivity troubleshooting, Go, Python, and gRPC.