This tutorial shows you how to deploy and serve a Qwen3 large language model (LLM) with the vLLM serving framework. You deploy the model on a single A4 virtual machine (VM) instance on Google Kubernetes Engine (GKE).
This tutorial is intended for machine learning (ML) engineers, platform administrators and operators, and for data and AI specialists who are interested in using Kubernetes container orchestration capabilities to handle inference workloads.
Objectives
Access Qwen3 by using Hugging Face.
Prepare your environment.
Create a GKE cluster in Autopilot mode.
Create a Cloud Storage bucket.
Create a Kubernetes secret for Hugging Face credentials.
Configure Workload Identity Federation for Cloud Storage.
Populate the Cloud Storage bucket with the Qwen3 model.
Deploy a vLLM container to your GKE.
Interact with Qwen3 by using curl.
Clean up.
Costs
This tutorial uses billable components of Google Cloud, including:
To generate a cost estimate based on your projected usage, use the Pricing Calculator.
Before you begin
- Sign in to your Google Cloud account. If you're new to Google Cloud, create an account to evaluate how our products perform in real-world scenarios. New customers also get $300 in free credits to run, test, and deploy workloads.
-
Install the Google Cloud CLI.
-
If you're using an external identity provider (IdP), you must first sign in to the gcloud CLI with your federated identity.
-
To initialize the gcloud CLI, run the following command:
gcloud init -
Create or select a Google Cloud project.
Roles required to select or create a project
- Select a project: Selecting a project doesn't require a specific IAM role—you can select any project that you've been granted a role on.
-
Create a project: To create a project, you need the Project Creator role
(
roles/resourcemanager.projectCreator), which contains theresourcemanager.projects.createpermission. Learn how to grant roles.
-
Create a Google Cloud project:
gcloud projects create PROJECT_ID
Replace
PROJECT_IDwith a name for the Google Cloud project you are creating. -
Select the Google Cloud project that you created:
gcloud config set project PROJECT_ID
Replace
PROJECT_IDwith your Google Cloud project name.
-
Verify that billing is enabled for your Google Cloud project.
Enable the required API:
Roles required to enable APIs
To enable APIs, you need the
serviceusage.services.enablepermission. If you created the project, then you likely already have this permission through the Owner role (roles/owner). Otherwise, you can get this permission through the Service Usage Admin role (roles/serviceusage.serviceUsageAdmin). Learn how to grant roles.gcloud services enable container.googleapis.com
-
Install the Google Cloud CLI.
-
If you're using an external identity provider (IdP), you must first sign in to the gcloud CLI with your federated identity.
-
To initialize the gcloud CLI, run the following command:
gcloud init -
Create or select a Google Cloud project.
Roles required to select or create a project
- Select a project: Selecting a project doesn't require a specific IAM role—you can select any project that you've been granted a role on.
-
Create a project: To create a project, you need the Project Creator role
(
roles/resourcemanager.projectCreator), which contains theresourcemanager.projects.createpermission. Learn how to grant roles.
-
Create a Google Cloud project:
gcloud projects create PROJECT_ID
Replace
PROJECT_IDwith a name for the Google Cloud project you are creating. -
Select the Google Cloud project that you created:
gcloud config set project PROJECT_ID
Replace
PROJECT_IDwith your Google Cloud project name.
-
Verify that billing is enabled for your Google Cloud project.
Enable the required API:
Roles required to enable APIs
To enable APIs, you need the
serviceusage.services.enablepermission. If you created the project, then you likely already have this permission through the Owner role (roles/owner). Otherwise, you can get this permission through the Service Usage Admin role (roles/serviceusage.serviceUsageAdmin). Learn how to grant roles.gcloud services enable container.googleapis.com
-
Grant roles to your user account. Run the following command once for each of the following IAM roles:
roles/container.admingcloud projects add-iam-policy-binding PROJECT_ID --member="user:USER_IDENTIFIER" --role=ROLE
Replace the following:
PROJECT_ID: Your project ID.USER_IDENTIFIER: The identifier for your user account. For example,myemail@example.com.ROLE: The IAM role that you grant to your user account.
- Sign in to or create a Hugging Face account.
Access Qwen3 by using Hugging Face
To use Hugging Face to access Qwen3, follow these steps:
- Sign in to Hugging Face
- Create a Hugging Face
readaccess token. Click Your Profile > Settings > Access Tokens > +Create new token. - Specify a name of your choice for the token and then select a role. The minimum role permission level that you can select for this tutorial is Read.
- Select Create token.
- Copy and save the generated token to your clipboard. You use it later in this tutorial.
Prepare your environment
To prepare your environment, set the default environment variables:
Replace the following:
YOUR_PROJECT_ID: the ID of the Google Cloud project where you want to create the GKE cluster.YOUR_RESERVATION_NAME: the name of the reservation that you want to use to create your GKE cluster. Based on the project in which the reservation exists, specify one of the following values:The reservation exists in your project:
RESERVATION_NAMEThe reservation exists in a different project, and your project can use the reservation:
projects/RESERVATION_PROJECT_ID/reservations/RESERVATION_NAME
YOUR_REGION: the region where you want to create your GKE cluster. You can only create the cluster in the region where your reservation exists.YOUR_CLUSTER_NAME: the name of the GKE cluster to create.YOUR_BUCKET_NAME: the name of the regional Cloud Storage bucket to create.YOUR_HF_TOKEN: the Hugging Face access token that you created in the previous section.YOUR_NETWORK_NAME: the network that the GKE cluster uses. Specify one of the following values:If you created a custom network, then specify the name of your network.
Otherwise, specify
default.
YOUR_SUBNETWORK_NAME: the subnetwork that the GKE cluster uses. Specify one of the following values:If you created a custom subnetwork, then specify the name of your subnetwork. You can only specify a subnetwork that exists in the same region as the reservation.
Otherwise, specify
default.
Create a GKE cluster in Autopilot mode
To create a GKE cluster in Autopilot mode, run the following command:
Creating the GKE cluster might take some time to complete. To verify that Google Cloud has finished creating your cluster, go to Kubernetes clusters on the Google Cloud console.Create a Cloud Storage bucket
To create a regional Cloud Storage bucket to store your model, run the following command:
Create a Kubernetes secret for Hugging Face credentials
To create a Kubernetes secret for Hugging Face credentials, follow these steps:
Configure
kubectlto communicate with your GKE cluster:Create a Kubernetes secret to store your Hugging Face token:
Configure Workload Identity Federation for Cloud Storage
To allow GKE to securely access the Cloud Storage bucket, set up GKE Workload Identity Federation:
Populate the Cloud Storage bucket with the Qwen3 model weights
To populate your Cloud Storage bucket with the Qwen3 model weights, run a Kubernetes Job that uses the Cloud Storage FUSE CSI driver to mount your bucket as a volume. The job downloads the model from Hugging Face if it does not already exist in the bucket.
Create a file named
qwen3-model-loader.yamlwith the following content:Apply the
qwen3-model-loader.yamlmanifest to initialize the download job:To verify that the job is running, stream the download job logs:
kubectl logs -f job/qwen3-model-loader -c downloaderWait for the model download job to complete:
To delete the job, run the following command:
Deploy a vLLM container to your GKE cluster
To deploy the vLLM container to serve the Qwen3 model by using Kubernetes Deployments, do the following:
Create a
qwen3-235b-deploy.yamlfile with your chosen vLLM deployment:Apply the
qwen3-235b-deploy.yamlfile to your GKE cluster:
Because the container uses Run:ai Model Streamer to stream model weights directly from Cloud Storage, the startup time is significantly accelerated.To see the completion status, run the following command:
The--timeout=600sflag allows the command to monitor the deployment for up to 10 minutes.
Interact with Qwen3 by using curl
To verify the Qwen3 model that you deployed, do the following:
Set up port forwarding to Qwen3:
Open a new terminal window. You can then chat with your model by using
curl:The output is similar to the following:
{ "id": "chatcmpl-a926ddf7ef2745ca832bda096e867764", "object": "chat.completion", "created": 1755023619, "model": "Qwen/Qwen3-235B-A22B-Instruct-2507", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "A GPU is a specialized electronic circuit designed to rapidly process and render graphics and perform parallel computations.", "refusal": null, "annotations": null, "audio": null, "function_call": null, "tool_calls": [], "reasoning_content": null }, "logprobs": null, "finish_reason": "stop", "stop_reason": null } ], "service_tier": null, "system_fingerprint": null, "usage": { "prompt_tokens": 16, "total_tokens": 36, "completion_tokens": 20, "prompt_tokens_details": null }, "prompt_logprobs": null, "kv_transfer_params": null }
Observe model performance
If you want to observe your model's performance, then you can use the vLLM dashboard integration in Cloud Monitoring. This dashboard helps you view critical performance metrics for your model like token throughput, network latency, and error rates. For information, see vLLM in the Monitoring documentation.
Clean up
To avoid incurring charges to your Google Cloud account for the resources used in this tutorial, either delete the project that contains the resources, or keep the project and delete the individual resources.
To avoid incurring charges to your Cloud Billing account for the resources used in this tutorial, either delete the project that contains the resources, or keep the project and delete the individual resources.
Delete the resources
To delete the tutorial resources, run the following commands:
Delete your GKE cluster
To delete your GKE cluster, run the following command:
Delete your project
Delete a Google Cloud project:
gcloud projects delete PROJECT_ID