Enterprises and individual developers frequently run multiple AI inference models. The right architecture can simplify how the models are called while also providing centralized governance. In this post, we’ll look at two reference architectures focused on networking AI inference model serving: one for Google Kubernetes Engine (GKE) and one all other backend types. First, we’ll explore the commonalities between the reference architectures that you’ll see later. Then we’ll explore unique components of the architecture for GKE backends and finally, we’ll go over the elements of the architecture for all backend types.
The entry point
You can expose your model deployment behind a stable, secure, and reliable entry point that acts as the front end for inference calls. This entry point also acts as a control zone where policy, security, and logic can be enforced. Both the Cloud Load balancer and the Inference Gateway provide entry point capability. These types of endpoints can terminate secure connections with TLS, integrate with API management components, extend functionally with service extensions, and capitalize on capabilities of Model Armor for added security.
Common services in the designs
Both reference designs use these services:
-
Private Service Connect inference endpoint: Anchors the entry point inside your consumer Virtual Private Cloud (VPC) network. Traffic hits a private internal IP address, keeping inference calls in your private network.
-
Apigee API Management (Optional): Integrates via an Apigee Extension Processor callout to handle client identity verification, rate limits, and quota enforcement before requests ever reach compute resources.
-
Model Armor: Serves as an inline AI safety checkpoint, screening prompts and output completions against prompt injection and sensitive data leakage.
Design pattern serving on GKE only
This section focuses on a GKE-only backend design. To understand the full end-to-end concept, please read the entire architecture document Networking for AI inference model serving on GKE. The design pattern is based on this diagram:






