MaaS Infrastructure Platforms
Building unified model service infrastructure across API compatibility, provider abstraction, tenant isolation, quota governance, fallback, and production observability.
Hi, I am Hao Wu.
I work on Tencent Cloud TokenHub, building MaaS infrastructure for unified model access, request governance, tenant isolation, quotas, observability, and reliability.
I am also a core committer of vllm-project/semantic-router, focusing on intelligent semantic routing across request semantics, model selection, safety governance, semantic cache, feedback learning, and production data planes.
Building unified model service infrastructure across API compatibility, provider abstraction, tenant isolation, quota governance, fallback, and production observability.
Designing routing decisions from request semantics, model capability matching, safety signals, semantic cache, feedback learning, and cost / latency / quality tradeoffs.
Connecting LLM gateways to proven data-plane ideas from Kubernetes, container networking, service discovery, load balancing, and end-to-end troubleshooting.
The router is not just a forwarding layer. It is a decision boundary where semantics, policy, backends, cost, latency, quality, safety, and feedback meet.
Intelligent semantic routing project in the official vLLM ecosystem for Mixture-of-Models, AI Gateway, and production LLM traffic governance.
Research and engineering community around agentic systems, infrastructure primitives, and collective intelligence.
Community space for semantic routing, evaluation, model selection, and deployment-oriented routing systems.
Tencent Cloud
Building unified LLM API Gateway and MaaS platform capabilities across model access, request governance, capacity protection, provider abstraction, fallback, quota, and observability.
Tencent Cloud
Built and maintained VPC CNI, Pod networking, Kubernetes Service, Ingress, and Gateway capabilities across production cloud-native traffic paths.