Compare commits

...
Author SHA1 Message Date
Holt SkinnerandGitHub e33b376688 Delete notebooks/official/pipelines/google_cloud_pipeline_components_automl_text.ipynb
b/442906902
2025-09-09 15:22:40 -05:00
Holt SkinnerandGitHub 915e5edf0a chore: Remove AutoML Notebooks for deprecated Text and Video features (#4253)
* chore: Remove AutoML Notebooks for deprecated Text and Video features

* Remove remaining Text/Video samples
2025-09-09 16:49:06 +00:00
Vertex MG TeamandCopybara-Service 5d4c1be285 Migrate llava notebook to use Model Garden SDK
PiperOrigin-RevId: 804765256
2025-09-09 00:16:12 -07:00
Vertex MG TeamandCopybara-Service 275bfb8f69 Weekly update the vllm serving container version to 20250905_0916_RC01.
PiperOrigin-RevId: 804514266
2025-09-08 11:21:48 -07:00
Vertex MG TeamandCopybara-Service 93cb0ca3d9 Migrate blip image captioning notebook to use Model Garden SDK
PiperOrigin-RevId: 803091009
2025-09-04 10:49:28 -07:00
Vertex MG TeamandCopybara-Service 41d60d052d Add EmbeddingGemma local inference notebook.
PiperOrigin-RevId: 803038284
2025-09-04 08:30:34 -07:00
Vertex MG TeamandCopybara-Service d872cdcf7e Migrate Paligemma2 notebook to use Model Garden SDK
PiperOrigin-RevId: 802861534
2025-09-03 22:35:04 -07:00
Vertex MG TeamandCopybara-Service ada5e4a854 chore: fix google-auth and requests package version.
PiperOrigin-RevId: 802270767
2025-09-02 13:33:33 -07:00
Vertex MG TeamandCopybara-Service ecb32b099d Migrate Stable Diffusion Upscaler notebook to use Model Garden SDK
PiperOrigin-RevId: 801824204
2025-09-01 08:55:58 -07:00
Vertex MG TeamandCopybara-Service f8d09e8e9b Migrate Stable Diffusion XL Lightning notebook to use Model Garden SDK
PiperOrigin-RevId: 801782973
2025-09-01 06:06:41 -07:00
Vertex MG TeamandCopybara-Service 007df88fba Migrate Llama4 notebook to use Model Garden SDK
PiperOrigin-RevId: 801662862
2025-08-31 22:27:06 -07:00
Vertex MG TeamandCopybara-Service c70f3ef9a3 Weekly update vllm/hf-tei/hf-inference-toolkit container image versions.
PiperOrigin-RevId: 801034279
2025-08-29 14:39:29 -07:00
Vertex MG TeamandCopybara-Service 505e101452 Bug fix for Wan2.2
PiperOrigin-RevId: 800419261
2025-08-28 05:11:01 -07:00
Vertex MG TeamandCopybara-Service b77d51b58b Migrate QWEN3 notebook to use Model Garden SDK
PiperOrigin-RevId: 800346334
2025-08-28 00:57:59 -07:00
Rayan DasoriyaandCopybara-Service deaa1ccf2b Update quota check to use gcloud beta quotas info describe command
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 800261134
2025-08-27 19:19:54 -07:00
Vertex MG TeamandCopybara-Service f6e38860aa Update H100/H200 region recommendations in DeepSeek deployment notebook.
PiperOrigin-RevId: 799800908
2025-08-26 18:29:57 -07:00
Vertex MG TeamandCopybara-Service 4937e382b1 Bug fix for Wan2.1
PiperOrigin-RevId: 799116301
2025-08-25 07:40:20 -07:00
Vertex MG TeamandCopybara-Service 0fe2770947 Weekly update the vLLM container image version.
PiperOrigin-RevId: 798391915
2025-08-22 16:59:25 -07:00
Vertex MG TeamandCopybara-Service 80d7ee67d6 Weekly update the container version for hf-inference-toolkit and hf-tei.
PiperOrigin-RevId: 797920954
2025-08-21 14:45:09 -07:00
Vertex MG TeamandCopybara-Service f17e2d6c8d Migrate gpt-oss deploy notebook to use Model Garden SDK
PiperOrigin-RevId: 797824747
2025-08-21 10:40:45 -07:00
Vertex MG TeamandCopybara-Service ae1cd0ed08 Migrate llama3.3 deployment notebook to use Model Garden SDK
PiperOrigin-RevId: 797622523
2025-08-20 23:34:08 -07:00
denisj3030andGitHub a68b491edc dep35 (#4227) 2025-08-20 20:02:14 +00:00
Vertex MG TeamandCopybara-Service d39bed012a Migrate Controlnet notebook to use Model Garden SDK
PiperOrigin-RevId: 797349160
2025-08-20 09:37:09 -07:00
Vertex MG TeamandCopybara-Service f8c93de1c4 Migrate BLIP2 deploy notebook to use Model Garden SDK
PiperOrigin-RevId: 797131322
2025-08-19 20:34:06 -07:00
Vertex MG TeamandCopybara-Service 0b5d97aa3a use A100 machine by default for Llama3.1 fast deployment
PiperOrigin-RevId: 796788128
2025-08-19 02:51:44 -07:00
Vertex MG TeamandCopybara-Service 87f0785ed1 Migrate Segment Anything Model(SAM) notebook to use Model Garden SDK
PiperOrigin-RevId: 796693389
2025-08-18 20:55:11 -07:00
Vertex MG TeamandCopybara-Service 09c8ea616b Weekly update the vLLM container image version.
PiperOrigin-RevId: 796581324
2025-08-18 14:40:14 -07:00
Vertex MG TeamandCopybara-Service 6c11241cf7 Migrate BLIP2 VQA(Visual Question Answering) notebook to use Model Garden SDK
PiperOrigin-RevId: 796316856
2025-08-18 01:35:17 -07:00
Vertex MG TeamandCopybara-Service 9b7dc2e4fd Migrate BLIP2 VQA(Visual Question Answering) notebook to use Model Garden SDK
PiperOrigin-RevId: 795738195
2025-08-15 22:43:11 -07:00
Vertex MG TeamandCopybara-Service 9b3d43b1a6 Use correct notebook_util.
PiperOrigin-RevId: 794874566
2025-08-13 22:18:11 -07:00
Dustin LuongandCopybara-Service 57ee4e2eab Update Qwen3 deployment notebook with new variants and corrected model names.
PiperOrigin-RevId: 794407429
2025-08-12 22:41:45 -07:00
Dustin LuongandCopybara-Service 4d736d7992 Use correct notebook_util.
PiperOrigin-RevId: 794396544
2025-08-12 22:00:22 -07:00
Dustin LuongandCopybara-Service 69e92a650d Add gpt-oss-20b finetuning with lora on Vertex notebook.
PiperOrigin-RevId: 794317611
2025-08-12 16:57:36 -07:00
Vertex MG TeamandCopybara-Service 4ed979eec4 Migrate Stable Diffusion XL 1.0 notebook to use Model Garden SDK
PiperOrigin-RevId: 793954002
2025-08-11 23:22:40 -07:00
Vertex MG TeamandCopybara-Service 0344de8090 Wan Deployment Notebook
PiperOrigin-RevId: 793839574
2025-08-11 16:30:15 -07:00
Vertex MG TeamandCopybara-Service 83eca0f0cc Update the vLLM container image version.
PiperOrigin-RevId: 793738684
2025-08-11 11:47:34 -07:00
Vertex MG TeamandCopybara-Service 4619b272fa Migrate QwQ notebook to use Model Garden SDK.
PiperOrigin-RevId: 793694751
2025-08-11 10:05:17 -07:00
Vertex MG TeamandCopybara-Service 091fa33c01 Migrate Gemma3 notebook to use Model Garden SDK.
PiperOrigin-RevId: 793692575
2025-08-11 10:01:02 -07:00
Rayan DasoriyaandCopybara-Service 42a05c2b35 No public description
PiperOrigin-RevId: 793644258
2025-08-11 07:50:06 -07:00
Vertex MG TeamandCopybara-Service 8ac32fa42f Migrate Gemma3n notebook to use Model Garden SDK
PiperOrigin-RevId: 793583308
2025-08-11 04:09:41 -07:00
Rayan DasoriyaandCopybara-Service 0303057f11 Update the common util location in the notebooks
PiperOrigin-RevId: 793425712
2025-08-10 18:26:31 -07:00
Ravi DalalandGitHub 7ae13b346a added notebooks and dockerfiles for serving open models on vertexai using vllm custom containers (#4148)
* added notebooks and dockerfiles for serving open models on vertexai using vllm customer containers

* updated official codeowners

* fixed linting errors

* fixed linting errors

* moved notebooks

* ran linter

* fixed param type

* added some formatting

* added autoscaling configuration to model deployment

* fixed a heading

* moved notebooks and docker folder under prediction

* updated notebook repo paths

* switched to raw_predict to avoid code changes and rebuild

* removed linting errors

* fixed readme lint error

* fixed links

* removed dedicated_endpoint_enabled

* updated workdir path

* fixed cell type

* added license

* fixed linting error

* using cloud build container image build

* fixed linting issues

* fixes

* cloudbuild yaml

* fixed gemini review comments

* fixed linting errors

* fixed tpu_count type

* handled invalid device type

* optimized dockerfile run command

* updated dockerfile

* addressed review comments

* fixed linting errors

* added license to cloudbuild and dockerfile

* optimized image build code

* fixed image_name variable
2025-08-08 18:01:03 +00:00
Rayan DasoriyaandCopybara-Service 44390cbd99 Add a no-op message to check_quota if it fails.
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 792654121
2025-08-08 09:32:08 -07:00
Vertex MG TeamandCopybara-Service 58dab0b1bb E5 notebook
PiperOrigin-RevId: 791999324
2025-08-06 22:48:50 -07:00
Vertex MG TeamandCopybara-Service 8d36834fcd Add GPT OSS models deployment notebook.
PiperOrigin-RevId: 791864273
2025-08-06 15:17:54 -07:00
Dustin LuongandCopybara-Service 31ca3e36f4 Add Qwen3-30B-A3B instruct and thinking 2507 variants to notebook.
PiperOrigin-RevId: 791746627
2025-08-06 10:27:39 -07:00
Vertex MG TeamandCopybara-Service a72d7dc49f Fix custom dataset input for axolotl notebooks.
PiperOrigin-RevId: 791627090
2025-08-06 04:16:45 -07:00
denisj3030andGitHub 9efbd48233 cl41 (#4193) 2025-08-05 17:02:14 +00:00
142 changed files with 10156 additions and 12893 deletions
@@ -472,13 +472,13 @@ def copy_model_artifacts(
])
def get_quota(project_id: str, region: str, resource_id: str) -> int:
def get_quota(project_id: str, region: str, quota_id: str) -> int:
"""Returns the quota for a resource in a region.
Args:
project_id: The project id.
region: The region.
resource_id: The resource id.
quota_id: The quota id.
Returns:
The quota for the resource in the region. Returns -1 if can not figure out
@@ -488,59 +488,39 @@ def get_quota(project_id: str, region: str, resource_id: str) -> int:
RuntimeError: If the command to get quota fails.
"""
service_endpoint = "aiplatform.googleapis.com"
command = (
"gcloud alpha services quota list"
f" --service={service_endpoint} --consumer=projects/{project_id}"
f" --filter='{service_endpoint}/{resource_id}' --format=json"
"gcloud beta quotas info describe"
f" {quota_id} --service={service_endpoint} --project={project_id} --format=json"
)
process = subprocess.run(
command, shell=True, capture_output=True, text=True, check=True
)
if process.returncode == 0:
quota_data = json.loads(process.stdout)
else:
raise RuntimeError(f"Error fetching quota data: {process.stderr}")
if not quota_data or "consumerQuotaLimits" not in quota_data[0]:
try:
process = subprocess.run(
command, shell=True, capture_output=True, text=True, check=True
)
except subprocess.CalledProcessError as e:
raise RuntimeError(f"Error fetching quota data: {e.stderr}") from e
quota_data = json.loads(process.stdout)
if not quota_data or "dimensionsInfos" not in quota_data:
return -1
if (
not quota_data[0]["consumerQuotaLimits"]
or "quotaBuckets" not in quota_data[0]["consumerQuotaLimits"][0]
):
return -1
all_regions_data = quota_data[0]["consumerQuotaLimits"][0]["quotaBuckets"]
# If the quota data does not have dimensions, it is global quota. However,
# global quota may be overridden by regional quota. So we need to check the
# global quota first.
global_quota = -1
if (
all_regions_data
and "dimensions" not in all_regions_data[0]
and "effectiveLimit" in all_regions_data[0]
):
global_quota = int(all_regions_data[0]["effectiveLimit"])
all_regions_data = quota_data["dimensionsInfos"]
for region_data in all_regions_data:
if (
region_data.get("dimensions")
and region_data["dimensions"]["region"] == region
):
if "effectiveLimit" in region_data:
return int(region_data["effectiveLimit"])
else:
return 0
return global_quota
applicable_locations = region_data.get("applicableLocations")
if not applicable_locations or region not in applicable_locations:
continue
details = region_data.get("details")
if not details or "value" not in details:
continue
return int(details["value"])
return 0
def get_resource_id(
def get_quota_id(
accelerator_type: str,
is_for_training: bool,
is_spot: bool = False,
is_restricted_image: bool = False,
is_dynamic_workload_scheduler: bool = False,
) -> str:
"""Returns the resource id for a given accelerator type and the use case.
"""Returns the quota id for a given accelerator type and the use case.
Args:
accelerator_type: The accelerator type.
@@ -552,40 +532,47 @@ def get_resource_id(
Workload Scheduler.
Returns:
The resource id.
The quota id.
"""
accelerator_suffix_map = {
"NVIDIA_TESLA_V100": "nvidia_v100_gpus",
"NVIDIA_TESLA_P100": "nvidia_p100_gpus",
"NVIDIA_L4": "nvidia_l4_gpus",
"NVIDIA_TESLA_A100": "nvidia_a100_gpus",
"NVIDIA_A100_80GB": "nvidia_a100_80gb_gpus",
"NVIDIA_H100_80GB": "nvidia_h100_gpus",
"NVIDIA_H100_MEGA_80GB": "nvidia_h100_mega_gpus",
"NVIDIA_H200_141GB": "nvidia_h200_gpus",
"NVIDIA_TESLA_T4": "nvidia_t4_gpus",
"TPU_V6e": "tpu_v6e",
"TPU_V5e": "tpu_v5e",
"TPU_V3": "tpu_v3",
accelerator_map = {
"NVIDIA_TESLA_V100": "V100GPUs",
"NVIDIA_TESLA_P100": "P100GPUs",
"NVIDIA_L4": "L4GPUs",
"NVIDIA_TESLA_A100": "A100GPUs",
"NVIDIA_A100_80GB": "A10080GBGPUs",
"NVIDIA_H100_80GB": "H100GPUs",
"NVIDIA_H100_MEGA_80GB": "H100MEGAGPUs",
"NVIDIA_H200_141GB": "H200GPUs",
"NVIDIA_GB200": "B200GPUs",
"NVIDIA_TESLA_T4": "T4GPUs",
"TPU_V6e": "V6ETPU",
"TPU_V5e": "V5ETPU",
"TPU_V3": "V3TPUs",
}
default_training_accelerator_map = {
key: f"custom_model_training_{accelerator_suffix_map[key]}"
for key in accelerator_suffix_map
key: f"CustomModelTraining{accelerator_map[key]}PerProjectPerRegion"
for key in accelerator_map
}
dws_training_accelerator_map = {
key: f"custom_model_training_preemptible_{accelerator_suffix_map[key]}"
for key in accelerator_suffix_map
key: (
f"CustomModelTrainingPreemptible{accelerator_map[key]}PerProjectPerRegion"
)
for key in accelerator_map
}
restricted_image_training_accelerator_map = {
"NVIDIA_A100_80GB": "restricted_image_training_nvidia_a100_80gb_gpus",
"NVIDIA_A100_80GB": (
"RestrictedImageTrainingA10080GBGPUsPerProjectPerRegion"
),
}
spot_serving_accelerator_map = {
key: f"custom_model_serving_preemptible_{accelerator_suffix_map[key]}"
for key in accelerator_suffix_map
key: (
f"CustomModelServingPreemptible{accelerator_map[key]}PerProjectPerRegion"
)
for key in accelerator_map
}
serving_accelerator_map = {
key: f"custom_model_serving_{accelerator_suffix_map[key]}"
for key in accelerator_suffix_map
key: f"CustomModelServing{accelerator_map[key]}PerProjectPerRegion"
for key in accelerator_map
}
if is_for_training:
@@ -646,14 +633,14 @@ def check_quota(
is_dynamic_workload_scheduler: Whether the resource is used with Dynamic
Workload Scheduler.
"""
resource_id = get_resource_id(
quota_id = get_quota_id(
accelerator_type,
is_for_training=is_for_training,
is_spot=is_spot,
is_restricted_image=is_restricted_image,
is_dynamic_workload_scheduler=is_dynamic_workload_scheduler,
)
quota = get_quota(project_id, region, resource_id)
quota = get_quota(project_id, region, quota_id)
quota_request_instruction = (
"Either use "
"a different region or request additional quota. Follow "
@@ -664,12 +651,12 @@ def check_quota(
)
if quota == -1:
raise ValueError(
f"Quota not found for: {resource_id} in {region}."
f"Quota not found for: {quota_id} in {region}."
f" {quota_request_instruction}"
)
if quota < accelerator_count:
raise ValueError(
f"Quota not enough for {resource_id} in {region}: {quota} <"
f"Quota not enough for {quota_id} in {region}: {quota} <"
f" {accelerator_count}. {quota_request_instruction}"
)
@@ -472,13 +472,13 @@ def copy_model_artifacts(
])
def get_quota(project_id: str, region: str, resource_id: str) -> int:
def get_quota(project_id: str, region: str, quota_id: str) -> int:
"""Returns the quota for a resource in a region.
Args:
project_id: The project id.
region: The region.
resource_id: The resource id.
quota_id: The quota id.
Returns:
The quota for the resource in the region. Returns -1 if can not figure out
@@ -488,59 +488,39 @@ def get_quota(project_id: str, region: str, resource_id: str) -> int:
RuntimeError: If the command to get quota fails.
"""
service_endpoint = "aiplatform.googleapis.com"
command = (
"gcloud alpha services quota list"
f" --service={service_endpoint} --consumer=projects/{project_id}"
f" --filter='{service_endpoint}/{resource_id}' --format=json"
"gcloud beta quotas info describe"
f" {quota_id} --service={service_endpoint} --project={project_id} --format=json"
)
process = subprocess.run(
command, shell=True, capture_output=True, text=True, check=True
)
if process.returncode == 0:
quota_data = json.loads(process.stdout)
else:
raise RuntimeError(f"Error fetching quota data: {process.stderr}")
if not quota_data or "consumerQuotaLimits" not in quota_data[0]:
try:
process = subprocess.run(
command, shell=True, capture_output=True, text=True, check=True
)
except subprocess.CalledProcessError as e:
raise RuntimeError(f"Error fetching quota data: {e.stderr}") from e
quota_data = json.loads(process.stdout)
if not quota_data or "dimensionsInfos" not in quota_data:
return -1
if (
not quota_data[0]["consumerQuotaLimits"]
or "quotaBuckets" not in quota_data[0]["consumerQuotaLimits"][0]
):
return -1
all_regions_data = quota_data[0]["consumerQuotaLimits"][0]["quotaBuckets"]
# If the quota data does not have dimensions, it is global quota. However,
# global quota may be overridden by regional quota. So we need to check the
# global quota first.
global_quota = -1
if (
all_regions_data
and "dimensions" not in all_regions_data[0]
and "effectiveLimit" in all_regions_data[0]
):
global_quota = int(all_regions_data[0]["effectiveLimit"])
all_regions_data = quota_data["dimensionsInfos"]
for region_data in all_regions_data:
if (
region_data.get("dimensions")
and region_data["dimensions"]["region"] == region
):
if "effectiveLimit" in region_data:
return int(region_data["effectiveLimit"])
else:
return 0
return global_quota
applicable_locations = region_data.get("applicableLocations")
if not applicable_locations or region not in applicable_locations:
continue
details = region_data.get("details")
if not details or "value" not in details:
continue
return int(details["value"])
return 0
def get_resource_id(
def get_quota_id(
accelerator_type: str,
is_for_training: bool,
is_spot: bool = False,
is_restricted_image: bool = False,
is_dynamic_workload_scheduler: bool = False,
) -> str:
"""Returns the resource id for a given accelerator type and the use case.
"""Returns the quota id for a given accelerator type and the use case.
Args:
accelerator_type: The accelerator type.
@@ -552,40 +532,47 @@ def get_resource_id(
Workload Scheduler.
Returns:
The resource id.
The quota id.
"""
accelerator_suffix_map = {
"NVIDIA_TESLA_V100": "nvidia_v100_gpus",
"NVIDIA_TESLA_P100": "nvidia_p100_gpus",
"NVIDIA_L4": "nvidia_l4_gpus",
"NVIDIA_TESLA_A100": "nvidia_a100_gpus",
"NVIDIA_A100_80GB": "nvidia_a100_80gb_gpus",
"NVIDIA_H100_80GB": "nvidia_h100_gpus",
"NVIDIA_H100_MEGA_80GB": "nvidia_h100_mega_gpus",
"NVIDIA_H200_141GB": "nvidia_h200_gpus",
"NVIDIA_TESLA_T4": "nvidia_t4_gpus",
"TPU_V6e": "tpu_v6e",
"TPU_V5e": "tpu_v5e",
"TPU_V3": "tpu_v3",
accelerator_map = {
"NVIDIA_TESLA_V100": "V100GPUs",
"NVIDIA_TESLA_P100": "P100GPUs",
"NVIDIA_L4": "L4GPUs",
"NVIDIA_TESLA_A100": "A100GPUs",
"NVIDIA_A100_80GB": "A10080GBGPUs",
"NVIDIA_H100_80GB": "H100GPUs",
"NVIDIA_H100_MEGA_80GB": "H100MEGAGPUs",
"NVIDIA_H200_141GB": "H200GPUs",
"NVIDIA_GB200": "B200GPUs",
"NVIDIA_TESLA_T4": "T4GPUs",
"TPU_V6e": "V6ETPU",
"TPU_V5e": "V5ETPU",
"TPU_V3": "V3TPUs",
}
default_training_accelerator_map = {
key: f"custom_model_training_{accelerator_suffix_map[key]}"
for key in accelerator_suffix_map
key: f"CustomModelTraining{accelerator_map[key]}PerProjectPerRegion"
for key in accelerator_map
}
dws_training_accelerator_map = {
key: f"custom_model_training_preemptible_{accelerator_suffix_map[key]}"
for key in accelerator_suffix_map
key: (
f"CustomModelTrainingPreemptible{accelerator_map[key]}PerProjectPerRegion"
)
for key in accelerator_map
}
restricted_image_training_accelerator_map = {
"NVIDIA_A100_80GB": "restricted_image_training_nvidia_a100_80gb_gpus",
"NVIDIA_A100_80GB": (
"RestrictedImageTrainingA10080GBGPUsPerProjectPerRegion"
),
}
spot_serving_accelerator_map = {
key: f"custom_model_serving_preemptible_{accelerator_suffix_map[key]}"
for key in accelerator_suffix_map
key: (
f"CustomModelServingPreemptible{accelerator_map[key]}PerProjectPerRegion"
)
for key in accelerator_map
}
serving_accelerator_map = {
key: f"custom_model_serving_{accelerator_suffix_map[key]}"
for key in accelerator_suffix_map
key: f"CustomModelServing{accelerator_map[key]}PerProjectPerRegion"
for key in accelerator_map
}
if is_for_training:
@@ -646,14 +633,14 @@ def check_quota(
is_dynamic_workload_scheduler: Whether the resource is used with Dynamic
Workload Scheduler.
"""
resource_id = get_resource_id(
quota_id = get_quota_id(
accelerator_type,
is_for_training=is_for_training,
is_spot=is_spot,
is_restricted_image=is_restricted_image,
is_dynamic_workload_scheduler=is_dynamic_workload_scheduler,
)
quota = get_quota(project_id, region, resource_id)
quota = get_quota(project_id, region, quota_id)
quota_request_instruction = (
"Either use "
"a different region or request additional quota. Follow "
@@ -664,12 +651,12 @@ def check_quota(
)
if quota == -1:
raise ValueError(
f"Quota not found for: {resource_id} in {region}."
f"Quota not found for: {quota_id} in {region}."
f" {quota_request_instruction}"
)
if quota < accelerator_count:
raise ValueError(
f"Quota not enough for {resource_id} in {region}: {quota} <"
f"Quota not enough for {quota_id} in {region}: {quota} <"
f" {accelerator_count}. {quota_request_instruction}"
)
@@ -120,7 +120,7 @@
"id": "L3dqbxovo5t6",
"metadata": {
"cellView": "form",
"id": "86a3d4d4d3f5"
"id": "440a9e07b0b3"
},
"outputs": [],
"source": [
@@ -161,7 +161,7 @@
"from google.cloud import aiplatform\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -238,7 +238,7 @@
"from google.cloud import aiplatform\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"\n",
@@ -545,58 +545,79 @@
"# @markdown ---\n",
"# @markdown **Option 2: GCS**\n",
"\n",
"# @markdown **Bucket Name:**\n",
"DATASET_BUCKET_NAME = \"\" # @param {type:\"string\"}\n",
"# @markdown **Dataset Type:** Refer to the [Axolotl config file](https://github.com/axolotl-ai-cloud/axolotl/blob/6ba5c0ed2c42a0e069b28c83646ee5a2a6904430/docs/config.qmd#L102) for more details.\n",
"DATASET_TYPE = \"\" # @param {type:\"string\"}\n",
"# @markdown **Path to Training Data :**\n",
"\n",
"# @markdown e.g. `gs://cloud-samples-data/vertex-ai/model-evaluation/peft_train_sample.jsonl`\n",
"TRAIN_DATASET_PATH = \"\" # @param {type:\"string\"}\n",
"# @markdown **File Type**. Refer to the [Axolotl config file](https://github.com/axolotl-ai-cloud/axolotl/blob/6ba5c0ed2c42a0e069b28c83646ee5a2a6904430/docs/config.qmd#L103).\n",
"FILE_TYPE = \"\" # @param {type:\"string\"}\n",
"FILE_TYPE = \"\" # @param {type:\"string\", placeholder: \"e.g. json\"}\n",
"# @markdown **Input Column from training data**.\n",
"TRAIN_INPUT_COLUMN = \"\" # @param {type:\"string\", placeholder: \"e.g. input_text\"}\n",
"# @markdown **Output Column from training data**.\n",
"TRAIN_OUTPUT_COLUMN = \"\" # @param {type:\"string\", placeholder: \"e.g. output_text\"}\n",
"\n",
"# @markdown **Path to Training Data (relative to bucket):**\n",
"TRAIN_DATAFILES_PATH = \"\" # @param {type:\"string\"}\n",
"# @markdown **[Optional] Path to Test Data (relative to bucket):**\n",
"# @markdown **[Optional] Path to Test Data :**\n",
"# @markdown To use a dedicated validation set, provide the file path. Otherwise, the training data will be split to create a validation set.\n",
"TEST_DATAFILES_PATH = \"\" # @param {type:\"string\"}\n",
"\n",
"if DATASET_BUCKET_NAME:\n",
"# @markdown e.g. `gs://cloud-samples-data/vertex-ai/model-evaluation/peft_test_sample.jsonl`\n",
"TEST_DATASET_PATH = \"\" # @param {type:\"string\"}\n",
"# @markdown **Input Column from test data**.\n",
"TEST_INPUT_COLUMN = \"\" # @param {type:\"string\", placeholder: \"e.g. prompt\"}\n",
"# @markdown **Output Column from test data**.\n",
"TEST_OUTPUT_COLUMN = \"\" # @param {type:\"string\", placeholder: \"e.g. ground_truth\"}\n",
"\n",
"if TRAIN_DATASET_PATH:\n",
" assert FILE_TYPE, \"FILE_TYPE must be set if TRAIN_DATASET_PATH is set.\"\n",
" assert (\n",
" TRAIN_DATAFILES_PATH\n",
" ), \"TRAIN_DATAFILES_PATH must be set if DATASET_BUCKET_NAME is set.\"\n",
" assert DATASET_TYPE, \"DATASET_TYPE must be set if DATASET_BUCKET_NAME is set.\"\n",
" assert FILE_TYPE, \"FILE_TYPE must be set if DATASET_BUCKET_NAME is set.\"\n",
" TRAIN_INPUT_COLUMN\n",
" ), \"TRAIN_INPUT_COLUMN must be set if TRAIN_DATASET_PATH is set.\"\n",
" assert (\n",
" TRAIN_OUTPUT_COLUMN\n",
" ), \"TRAIN_OUTPUT_COLUMN must be set if TRAIN_DATASET_PATH is set.\"\n",
"\n",
"if TEST_DATASET_PATH:\n",
" assert (\n",
" TRAIN_DATASET_PATH\n",
" ), \"TRAIN_DATASET_PATH must be set if TEST_DATASET_PATH is set.\"\n",
" assert (\n",
" TEST_INPUT_COLUMN\n",
" ), \"TEST_INPUT_COLUMN must be set if TEST_DATASET_PATH is set.\"\n",
" assert (\n",
" TEST_OUTPUT_COLUMN\n",
" ), \"TEST_OUTPUT_COLUMN must be set if TEST_DATASET_PATH is set.\"\n",
"\n",
"assert not (\n",
" HF_DATASET and DATASET_BUCKET_NAME\n",
"), \"Only one of HF_DATASET or DATASET_BUCKET_NAME can be set.\"\n",
" HF_DATASET and TRAIN_DATASET_PATH\n",
"), \"Only one of HF_DATASET or TRAIN_DATASET_PATH can be set.\"\n",
"\n",
"if DATASET_BUCKET_NAME:\n",
" paths = TRAIN_DATAFILES_PATH.split(\",\")\n",
"datasets = []\n",
"test_datasets = []\n",
"\n",
"if TRAIN_DATASET_PATH:\n",
" dataset = {\n",
" \"path\": f\"/gcs/{DATASET_BUCKET_NAME}/\",\n",
" \"type\": DATASET_TYPE,\n",
" \"data_files\": [],\n",
" \"path\": TRAIN_DATASET_PATH,\n",
" \"ds_type\": FILE_TYPE,\n",
" \"type\": {\n",
" \"field_instruction\": TRAIN_INPUT_COLUMN,\n",
" \"field_output\": TRAIN_OUTPUT_COLUMN,\n",
" },\n",
" \"format\": \"<|im_start|> user \\\\n {instruction} {input} <|im_end|>\",\n",
" \"no_input_format\": \"<|im_start|> user \\\\n {instruction} <|im_end|>\",\n",
" }\n",
" for path in paths:\n",
" if path.startswith(\"/\"):\n",
" path = path[1:]\n",
" dataset[\"data_files\"].append(f\"/gcs/{DATASET_BUCKET_NAME}/{path}\")\n",
" dataset[\"split\"] = \"train\"\n",
" datasets.append(dataset)\n",
"\n",
"if TEST_DATAFILES_PATH:\n",
" paths = TEST_DATAFILES_PATH.split(\",\")\n",
"if TEST_DATASET_PATH:\n",
" dataset = {\n",
" \"path\": f\"/gcs/{DATASET_BUCKET_NAME}/\",\n",
" \"type\": DATASET_TYPE,\n",
" \"data_files\": [],\n",
" \"path\": TEST_DATASET_PATH,\n",
" \"ds_type\": FILE_TYPE,\n",
" \"type\": {\n",
" \"field_instruction\": TEST_INPUT_COLUMN,\n",
" \"field_output\": TEST_OUTPUT_COLUMN,\n",
" },\n",
" \"format\": \"<|im_start|> user \\\\n {instruction} {input} <|im_end|>\",\n",
" \"no_input_format\": \"<|im_start|> user \\\\n {instruction} <|im_end|>\",\n",
" \"split\": \"train\",\n",
" }\n",
" for path in paths:\n",
" if path.startswith(\"/\"):\n",
" path = path[1:]\n",
" dataset[\"data_files\"].append(f\"/gcs/{DATASET_BUCKET_NAME}/{path}\")\n",
" dataset[\"split\"] = \"train\"\n",
" test_datasets.append(dataset)\n",
"\n",
"if HF_DATASET:\n",
@@ -860,7 +881,7 @@
" raise ValueError(f\"Unsupported accelerator type: {training_accelerator_type}\")\n",
"\n",
"TRAIN_DOCKER_URI = (\n",
" f\"{repo}/vertex-vision-model-garden-dockers/axolotl-train-dws:20250515-1800-rc0\"\n",
" f\"{repo}/vertex-vision-model-garden-dockers/axolotl-train-dws:20250812-1800-rc1\"\n",
")\n",
"\n",
"common_util.check_quota(\n",
@@ -1104,7 +1125,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -1113,6 +1133,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -250,7 +250,7 @@
"from google.cloud import aiplatform\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"\n",
@@ -522,67 +522,62 @@
"# @markdown **Option 1: Hugging Face**\n",
"\n",
"# @markdown **Hugging Face Dataset Name:**\n",
"HF_DATASET = \"\" # @param {type:\"string\", placeholder: \"e.g. timdettmers/openassistant-guanaco\"}\n",
"# @markdown **Set the dataset type:** Refer to [Axolotl config file](https://github.com/axolotl-ai-cloud/axolotl/blob/6ba5c0ed2c42a0e069b28c83646ee5a2a6904430/docs/config.qmd#L102) for more details.\n",
"HF_DATASET_TYPE = \"\" # @param {type:\"string\", placeholder: \"e.g. completion\"}\n",
"HF_DATASET = \"\" # @param {type:\"string\", placeholder: \"e.g. trl-lib/chatbot_arena_completions\"}\n",
"# @markdown **Set the dataset type:** Refer to [Axolotl config file](https://github.com/axolotl-ai-cloud/axolotl/blob/6ba5c0ed2c42a0e069b28c83646ee5a2a6904430/docs/config.qmd#L140) for more details.\n",
"HF_DATASET_TYPE = \"\" # @param {type:\"string\", placeholder: \"e.g. chat_template\"}\n",
"if HF_DATASET:\n",
" assert HF_DATASET_TYPE, \"HF_DATASET_TYPE must be set if HF_DATASET is set.\"\n",
"\n",
"# @markdown ---\n",
"# @markdown **Option 2: GCS**\n",
"\n",
"# @markdown **Bucket Name:**\n",
"DATASET_BUCKET_NAME = \"\" # @param {type:\"string\"}\n",
"# @markdown **Dataset Type:** Refer to the [Axolotl config file](https://github.com/axolotl-ai-cloud/axolotl/blob/6ba5c0ed2c42a0e069b28c83646ee5a2a6904430/docs/config.qmd#L102) for more details.\n",
"DATASET_TYPE = \"\" # @param {type:\"string\"}\n",
"# @markdown **Path to Training Data :**\n",
"\n",
"# @markdown E.g. `gs://cloud-samples-data/vertex-ai/model-garden/datasets/vertex-sample-chat-train.jsonl`\n",
"TRAIN_DATASET_PATH = \"\" # @param {type:\"string\"}\n",
"# @markdown **File Type**. Refer to the [Axolotl config file](https://github.com/axolotl-ai-cloud/axolotl/blob/6ba5c0ed2c42a0e069b28c83646ee5a2a6904430/docs/config.qmd#L103).\n",
"FILE_TYPE = \"\" # @param {type:\"string\"}\n",
"FILE_TYPE = \"\" # @param {type:\"string\", placeholder: \"e.g. json\"}\n",
"# @markdown **Messages Column**. Refer to the [Axolotl config file](https://github.com/axolotl-ai-cloud/axolotl/blob/6ba5c0ed2c42a0e069b28c83646ee5a2a6904430/docs/config.qmd#L155).\n",
"MESSAGES_COLUMN = \"messages\" # @param {type:\"string\", placeholder: \"e.g. messages\"}\n",
"\n",
"# @markdown **Path to Training Data (relative to bucket):**\n",
"TRAIN_DATAFILES_PATH = \"\" # @param {type:\"string\"}\n",
"# @markdown **[Optional] Path to Test Data (relative to bucket):**\n",
"# @markdown **[Optional] Path to Test Data :**\n",
"# @markdown To use a dedicated validation set, provide the file path. Otherwise, the training data will be split to create a validation set.\n",
"TEST_DATAFILES_PATH = \"\" # @param {type:\"string\"}\n",
"\n",
"if DATASET_BUCKET_NAME:\n",
"# @markdown E.g. `gs://cloud-samples-data/vertex-ai/model-garden/datasets/vertex-sample-chat-validation.jsonl`\n",
"TEST_DATASET_PATH = \"\" # @param {type:\"string\"}\n",
"\n",
"if TRAIN_DATASET_PATH:\n",
" assert FILE_TYPE, \"FILE_TYPE must be set if TRAIN_DATASET_PATH is set.\"\n",
"\n",
"if TEST_DATASET_PATH:\n",
" assert (\n",
" TRAIN_DATAFILES_PATH\n",
" ), \"TRAIN_DATAFILES_PATH must be set if DATASET_BUCKET_NAME is set.\"\n",
" assert DATASET_TYPE, \"DATASET_TYPE must be set if DATASET_BUCKET_NAME is set.\"\n",
" assert FILE_TYPE, \"FILE_TYPE must be set if DATASET_BUCKET_NAME is set.\"\n",
" TRAIN_DATASET_PATH\n",
" ), \"TRAIN_DATASET_PATH must be set if TEST_DATASET_PATH is set.\"\n",
"\n",
"assert not (\n",
" HF_DATASET and DATASET_BUCKET_NAME\n",
"), \"Only one of HF_DATASET or DATASET_BUCKET_NAME can be set.\"\n",
" HF_DATASET and TRAIN_DATASET_PATH\n",
"), \"Only one of HF_DATASET or TRAIN_DATASET_PATH can be set.\"\n",
"\n",
"if DATASET_BUCKET_NAME:\n",
" paths = TRAIN_DATAFILES_PATH.split(\",\")\n",
"datasets = []\n",
"test_datasets = []\n",
"\n",
"if TRAIN_DATASET_PATH:\n",
" dataset = {\n",
" \"path\": f\"/gcs/{DATASET_BUCKET_NAME}/\",\n",
" \"type\": DATASET_TYPE,\n",
" \"data_files\": [],\n",
" \"path\": TRAIN_DATASET_PATH,\n",
" \"ds_type\": FILE_TYPE,\n",
" \"type\": \"chat_template\",\n",
" \"field_messages\": MESSAGES_COLUMN,\n",
" }\n",
" for path in paths:\n",
" if path.startswith(\"/\"):\n",
" path = path[1:]\n",
" dataset[\"data_files\"].append(f\"/gcs/{DATASET_BUCKET_NAME}/{path}\")\n",
" dataset[\"split\"] = \"train\"\n",
" datasets.append(dataset)\n",
"\n",
"if TEST_DATAFILES_PATH:\n",
" paths = TEST_DATAFILES_PATH.split(\",\")\n",
"if TEST_DATASET_PATH:\n",
" dataset = {\n",
" \"path\": f\"/gcs/{DATASET_BUCKET_NAME}/\",\n",
" \"type\": DATASET_TYPE,\n",
" \"data_files\": [],\n",
" \"path\": TEST_DATASET_PATH,\n",
" \"ds_type\": FILE_TYPE,\n",
" \"type\": \"chat_template\",\n",
" \"split\": \"train\",\n",
" \"field_messages\": MESSAGES_COLUMN,\n",
" }\n",
" for path in paths:\n",
" if path.startswith(\"/\"):\n",
" path = path[1:]\n",
" dataset[\"data_files\"].append(f\"/gcs/{DATASET_BUCKET_NAME}/{path}\")\n",
" dataset[\"split\"] = \"train\"\n",
" test_datasets.append(dataset)\n",
"\n",
"if HF_DATASET:\n",
@@ -844,7 +839,7 @@
" raise ValueError(f\"Unsupported accelerator type: {training_accelerator_type}\")\n",
"\n",
"TRAIN_DOCKER_URI = (\n",
" f\"{repo}/vertex-vision-model-garden-dockers/axolotl-train-dws:20250515-1800-rc0\"\n",
" f\"{repo}/vertex-vision-model-garden-dockers/axolotl-train-dws:20250812-1800-rc1\"\n",
")\n",
"\n",
"common_util.check_quota(\n",
@@ -1092,7 +1087,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -1101,6 +1095,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
File diff suppressed because it is too large Load Diff
@@ -247,7 +247,7 @@
"from google.cloud import aiplatform\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"\n",
@@ -519,67 +519,62 @@
"# @markdown **Option 1: Hugging Face**\n",
"\n",
"# @markdown **Hugging Face Dataset Name:**\n",
"HF_DATASET = \"\" # @param {type:\"string\", placeholder: \"e.g. timdettmers/openassistant-guanaco\"}\n",
"# @markdown **Set the dataset type:** Refer to [Axolotl config file](https://github.com/axolotl-ai-cloud/axolotl/blob/6ba5c0ed2c42a0e069b28c83646ee5a2a6904430/docs/config.qmd#L102) for more details.\n",
"HF_DATASET_TYPE = \"\" # @param {type:\"string\", placeholder: \"e.g. completion\"}\n",
"HF_DATASET = \"\" # @param {type:\"string\", placeholder: \"e.g. trl-lib/chatbot_arena_completions\"}\n",
"# @markdown **Set the dataset type:** Refer to [Axolotl config file](https://github.com/axolotl-ai-cloud/axolotl/blob/6ba5c0ed2c42a0e069b28c83646ee5a2a6904430/docs/config.qmd#L140) for more details.\n",
"HF_DATASET_TYPE = \"\" # @param {type:\"string\", placeholder: \"e.g. chat_template\"}\n",
"if HF_DATASET:\n",
" assert HF_DATASET_TYPE, \"HF_DATASET_TYPE must be set if HF_DATASET is set.\"\n",
"\n",
"# @markdown ---\n",
"# @markdown **Option 2: GCS**\n",
"\n",
"# @markdown **Bucket Name:**\n",
"DATASET_BUCKET_NAME = \"\" # @param {type:\"string\"}\n",
"# @markdown **Dataset Type:** Refer to the [Axolotl config file](https://github.com/axolotl-ai-cloud/axolotl/blob/6ba5c0ed2c42a0e069b28c83646ee5a2a6904430/docs/config.qmd#L102) for more details.\n",
"DATASET_TYPE = \"\" # @param {type:\"string\"}\n",
"# @markdown **Path to Training Data :**\n",
"\n",
"# @markdown E.g. `gs://cloud-samples-data/vertex-ai/model-garden/datasets/vertex-sample-chat-train.jsonl`\n",
"TRAIN_DATASET_PATH = \"\" # @param {type:\"string\"}\n",
"# @markdown **File Type**. Refer to the [Axolotl config file](https://github.com/axolotl-ai-cloud/axolotl/blob/6ba5c0ed2c42a0e069b28c83646ee5a2a6904430/docs/config.qmd#L103).\n",
"FILE_TYPE = \"\" # @param {type:\"string\"}\n",
"FILE_TYPE = \"\" # @param {type:\"string\", placeholder: \"e.g. json\"}\n",
"# @markdown **Messages Column**. Refer to the [Axolotl config file](https://github.com/axolotl-ai-cloud/axolotl/blob/6ba5c0ed2c42a0e069b28c83646ee5a2a6904430/docs/config.qmd#L155).\n",
"MESSAGES_COLUMN = \"messages\" # @param {type:\"string\", placeholder: \"e.g. messages\"}\n",
"\n",
"# @markdown **Path to Training Data (relative to bucket):**\n",
"TRAIN_DATAFILES_PATH = \"\" # @param {type:\"string\"}\n",
"# @markdown **[Optional] Path to Test Data (relative to bucket):**\n",
"# @markdown **[Optional] Path to Test Data :**\n",
"# @markdown To use a dedicated validation set, provide the file path. Otherwise, the training data will be split to create a validation set.\n",
"TEST_DATAFILES_PATH = \"\" # @param {type:\"string\"}\n",
"\n",
"if DATASET_BUCKET_NAME:\n",
"# @markdown E.g. `gs://cloud-samples-data/vertex-ai/model-garden/datasets/vertex-sample-chat-validation.jsonl`\n",
"TEST_DATASET_PATH = \"\" # @param {type:\"string\"}\n",
"\n",
"if TRAIN_DATASET_PATH:\n",
" assert FILE_TYPE, \"FILE_TYPE must be set if TRAIN_DATASET_PATH is set.\"\n",
"\n",
"if TEST_DATASET_PATH:\n",
" assert (\n",
" TRAIN_DATAFILES_PATH\n",
" ), \"TRAIN_DATAFILES_PATH must be set if DATASET_BUCKET_NAME is set.\"\n",
" assert DATASET_TYPE, \"DATASET_TYPE must be set if DATASET_BUCKET_NAME is set.\"\n",
" assert FILE_TYPE, \"FILE_TYPE must be set if DATASET_BUCKET_NAME is set.\"\n",
" TRAIN_DATASET_PATH\n",
" ), \"TRAIN_DATASET_PATH must be set if TEST_DATASET_PATH is set.\"\n",
"\n",
"assert not (\n",
" HF_DATASET and DATASET_BUCKET_NAME\n",
"), \"Only one of HF_DATASET or DATASET_BUCKET_NAME can be set.\"\n",
" HF_DATASET and TRAIN_DATASET_PATH\n",
"), \"Only one of HF_DATASET or TRAIN_DATASET_PATH can be set.\"\n",
"\n",
"if DATASET_BUCKET_NAME:\n",
" paths = TRAIN_DATAFILES_PATH.split(\",\")\n",
"datasets = []\n",
"test_datasets = []\n",
"\n",
"if TRAIN_DATASET_PATH:\n",
" dataset = {\n",
" \"path\": f\"/gcs/{DATASET_BUCKET_NAME}/\",\n",
" \"type\": DATASET_TYPE,\n",
" \"data_files\": [],\n",
" \"path\": TRAIN_DATASET_PATH,\n",
" \"ds_type\": FILE_TYPE,\n",
" \"type\": \"chat_template\",\n",
" \"field_messages\": MESSAGES_COLUMN,\n",
" }\n",
" for path in paths:\n",
" if path.startswith(\"/\"):\n",
" path = path[1:]\n",
" dataset[\"data_files\"].append(f\"/gcs/{DATASET_BUCKET_NAME}/{path}\")\n",
" dataset[\"split\"] = \"train\"\n",
" datasets.append(dataset)\n",
"\n",
"if TEST_DATAFILES_PATH:\n",
" paths = TEST_DATAFILES_PATH.split(\",\")\n",
"if TEST_DATASET_PATH:\n",
" dataset = {\n",
" \"path\": f\"/gcs/{DATASET_BUCKET_NAME}/\",\n",
" \"type\": DATASET_TYPE,\n",
" \"data_files\": [],\n",
" \"path\": TEST_DATASET_PATH,\n",
" \"ds_type\": FILE_TYPE,\n",
" \"type\": \"chat_template\",\n",
" \"split\": \"train\",\n",
" \"field_messages\": MESSAGES_COLUMN,\n",
" }\n",
" for path in paths:\n",
" if path.startswith(\"/\"):\n",
" path = path[1:]\n",
" dataset[\"data_files\"].append(f\"/gcs/{DATASET_BUCKET_NAME}/{path}\")\n",
" dataset[\"split\"] = \"train\"\n",
" test_datasets.append(dataset)\n",
"\n",
"if HF_DATASET:\n",
@@ -841,7 +836,7 @@
" raise ValueError(f\"Unsupported accelerator type: {training_accelerator_type}\")\n",
"\n",
"TRAIN_DOCKER_URI = (\n",
" f\"{repo}/vertex-vision-model-garden-dockers/axolotl-train-dws:20250515-1800-rc0\"\n",
" f\"{repo}/vertex-vision-model-garden-dockers/axolotl-train-dws:20250812-1800-rc1\"\n",
")\n",
"\n",
"common_util.check_quota(\n",
@@ -119,7 +119,7 @@
"from google.cloud import aiplatform\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -621,7 +621,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -630,6 +629,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -120,7 +120,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -59,41 +59,43 @@
"source": [
"## Overview\n",
"\n",
"This notebook demonstrates deploying E5 text embedding models in Vertex AI.\n",
"This notebook demonstrates how to deploy a **E 5** open model on Google Cloud Vertex AI.\n",
"\n",
"### Objective\n",
"### Objectives\n",
"\n",
"- Deploy prebuilt E5 models with Hugging Face [Text Embeddings Inference](https://github.com/huggingface/text-embeddings-inference) (TEI) docker image on a Vertex AI Endpoint\n",
" - [intfloat/multilingual-e5-large-instruct](https://huggingface.co/intfloat/multilingual-e5-large-instruct): 560M params, instruction-tuned\n",
" - [intfloat/multilingual-e5-large](https://huggingface.co/intfloat/multilingual-e5-large): 560M params\n",
" - [intfloat/e5-large-v2](https://huggingface.co/intfloat/e5-large-v2): 335M params\n",
" - [intfloat/multilingual-e5-small](https://huggingface.co/intfloat/multilingual-e5-small): 118M params\n",
" - [intfloat/e5-base-v2](https://huggingface.co/intfloat/e5-base-v2): 109M params\n",
" - [intfloat/e5-small-v2](https://huggingface.co/intfloat/e5-small-v2): 33M params\n",
"- Run inference on the deployed Vertex AI Endpoint\n",
"- Deploy E 5 using containerized backends like [vLLM](https://github.com/vllm-project/vllm) on GPU.\n",
"- Use the deployed model to serve chat completion requests for both text and multimodal inputs.\n",
"\n",
"### File a Bug\n",
"\n",
"### File a bug\n",
"\n",
"File a bug on [GitHub](https://github.com/GoogleCloudPlatform/vertex-ai-samples/issues/new) if you encounter any issue with the notebook.\n",
"If you encounter issues with this notebook, report them on [GitHub](https://github.com/GoogleCloudPlatform/vertex-ai-samples/issues/new).\n",
"\n",
"### Costs\n",
"\n",
"This tutorial uses billable components of Google Cloud:\n",
"\n",
"* Vertex AI\n",
"* Cloud Storage\n",
"- Vertex AI\n",
"- Cloud Storage\n",
"\n",
"Learn about [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing), [Cloud Storage pricing](https://cloud.google.com/storage/pricing), and use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to generate a cost estimate based on your projected usage."
"Refer to the [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing) and [Cloud Storage pricing](https://cloud.google.com/storage/pricing) pages for more information. Use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to estimate your projected costs."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "264c07757582"
"id": "jeYw-Czg-DFy"
},
"source": [
"## Before you begin"
"## Get Started"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "KgyhGvEzBDkj"
},
"source": [
"### Install Vertex AI SDK and other required packages"
]
},
{
@@ -101,78 +103,169 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "LyEVDkHhAUHF"
"id": "iCacdLqG-IsH"
},
"outputs": [],
"source": [
"# @title Setup Google Cloud project\n",
"%pip install --upgrade --force-reinstall --quiet 'google-cloud-aiplatform>=1.106.0' 'openai' 'google-auth==2.27.0' 'requests==2.32.3'"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "HUKCrpBy-3yf"
},
"source": [
"### Authenticate the Notebook Environment (Colab only)\n",
"\n",
"# @markdown 1. [Make sure that billing is enabled for your project](https://cloud.google.com/billing/docs/how-to/modify-project).\n",
"If you're running this notebook in Google Colab, run the following cell to authenticate."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "JXwCT1kn-3Gu"
},
"outputs": [],
"source": [
"import sys\n",
"\n",
"# @markdown 2. **[Optional]** Set region. If not set, the region will be set automatically according to Colab Enterprise environment.\n",
"if \"google.colab\" in sys.modules:\n",
" from google.colab import auth\n",
"\n",
"REGION = \"\" # @param {type:\"string\"}\n",
" auth.authenticate_user()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "AcW2nwB8-7yC"
},
"source": [
"### Set Google Cloud Project Information\n",
"\n",
"# @markdown 3. If you want to run predictions with A100 80GB or H100 GPUs, we recommend using the regions listed below. **NOTE:** Make sure you have associated quota in selected regions. Click the links to see your current quota for each GPU type: [Nvidia A100 80GB](https://console.cloud.google.com/iam-admin/quotas?metric=aiplatform.googleapis.com%2Fcustom_model_serving_nvidia_a100_80gb_gpus), [Nvidia H100 80GB](https://console.cloud.google.com/iam-admin/quotas?metric=aiplatform.googleapis.com%2Fcustom_model_serving_nvidia_h100_gpus). You can request for quota following the instructions at [\"Request a higher quota\"](https://cloud.google.com/docs/quota/view-manage#requesting_higher_quota).\n",
"To get started with Vertex AI, ensure you have an existing Google Cloud project and that the [Vertex AI API is enabled](https://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com).\n",
"\n",
"# @markdown > | Machine Type | Accelerator Type | Recommended Regions |\n",
"# @markdown | ----------- | ----------- | ----------- |\n",
"# @markdown | a2-ultragpu-1g | 1 NVIDIA_A100_80GB | us-central1, us-east4, europe-west4, asia-southeast1, us-east4 |\n",
"# @markdown | a3-highgpu-2g | 2 NVIDIA_H100_80GB | us-west1, asia-southeast1, europe-west4 |\n",
"# @markdown | a3-highgpu-4g | 4 NVIDIA_H100_80GB | us-west1, asia-southeast1, europe-west4 |\n",
"# @markdown | a3-highgpu-8g | 8 NVIDIA_H100_80GB | us-central1, europe-west4, us-west1, asia-southeast1 |\n",
"\n",
"# Import the necessary packages\n",
"\n",
"# Upgrade Vertex AI SDK.\n",
"! pip3 install --upgrade --quiet 'google-cloud-aiplatform==1.103.0'\n",
"! pip3 install --quiet torchvision\n",
"\n",
"import importlib\n",
"See the guide on [setting up your project and development environment](https://cloud.google.com/vertex-ai/docs/start/cloud-environment). Also confirm that [billing is enabled](https://cloud.google.com/billing/docs/how-to/modify-project).\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "eIVLp0oE--k-"
},
"outputs": [],
"source": [
"# Use the environment variable if the user doesn't provide Project ID.\n",
"import os\n",
"from typing import Tuple\n",
"\n",
"from google.cloud import aiplatform\n",
"from torch import Tensor\n",
"\n",
"if os.environ.get(\"VERTEX_PRODUCT\") != \"COLAB_ENTERPRISE\":\n",
" ! pip install --upgrade tensorflow\n",
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
"\n",
"\n",
"# Get the default cloud project id.\n",
"PROJECT_ID = os.environ[\"GOOGLE_CLOUD_PROJECT\"]\n",
"\n",
"# Get the default region for launching jobs.\n",
"if not REGION:\n",
" REGION = os.environ[\"GOOGLE_CLOUD_REGION\"]\n",
"\n",
"# Initialize Vertex AI API.\n",
"print(\"Initializing Vertex AI API.\")\n",
"aiplatform.init(project=PROJECT_ID, location=REGION)\n",
"\n",
"! gcloud config set project $PROJECT_ID\n",
"import vertexai\n",
"\n",
"vertexai.init(\n",
" project=PROJECT_ID,\n",
" location=REGION,\n",
"PROJECT_ID = \"[your-project-id]\" # @param {type: \"string\", placeholder: \"[your-project-id]\", isTemplate: true}\n",
"\n",
"if not PROJECT_ID or PROJECT_ID == \"[your-project-id]\":\n",
" PROJECT_ID = str(os.environ.get(\"GOOGLE_CLOUD_PROJECT\"))\n",
"\n",
"REGION = os.environ.get(\"GOOGLE_CLOUD_REGION\", \"us-central1\")\n",
"\n",
"vertexai.init(project=PROJECT_ID, location=REGION)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "Q0CXrvcZH_aw"
},
"source": [
"### Import libraries"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "3G2UXB82ICs6"
},
"outputs": [],
"source": [
"from vertexai import model_garden"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "upYRiGtP_-iN"
},
"source": [
"## Deploy model"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "H2WC_0hXDVXc"
},
"source": [
"### Choose model variant"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "u41zbNa2EoFq"
},
"source": [
"You can proceed with the default model variant or select a different one."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "-fgC4NLSDkF7"
},
"outputs": [],
"source": [
"model_version = \"multilingual-e5-large-instruct\" # @param [\"e5-base-v2\", \"e5-large-v2\", \"e5-small-v2\", \"multilingual-e5-large\", \"multilingual-e5-large-instruct\", \"multilingual-e5-small\"] {isTemplate:true}\n",
"MODEL_NAME = f\"intfloat/e5@{model_version}\""
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "VRnUgU8LF3_i"
},
"source": [
"To see all deployable model variants available in Model Garden, use:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "-QLd-wshF6sB"
},
"outputs": [],
"source": [
"all_model_versions = model_garden.list_deployable_models(\n",
" model_filter=\"e5\", list_hf_models=False\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "2xeBQF0iVwSr"
"id": "N0UeFHa2GO63"
},
"source": [
"## Deploy and Predict"
"Once you've selected a model variant, initialize it:"
]
},
{
@@ -180,51 +273,22 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "I1u2FLa9XgVD"
"id": "GZiV3trBBcA3"
},
"outputs": [],
"source": [
"# @title Select the model variants\n",
"# @markdown This section uploads a prebuilt model to Model Registry and deploys it on the Endpoint. The model deployment step will take ~15 minutes to complete.\n",
"model = model_garden.OpenModel(MODEL_NAME)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "-0cL378wFlvf"
},
"source": [
"### Check the Deployment Configuration\n",
"\n",
"prebuilt_model_id = \"intfloat/e5-small-v2\" # @param [\"intfloat/multilingual-e5-large-instruct\", \"intfloat/multilingual-e5-large\", \"intfloat/e5-large-v2\", \"intfloat/multilingual-e5-small\", \"intfloat/e5-base-v2\", \"intfloat/e5-small-v2\"]\n",
"\n",
"# @markdown Specify a processor for the TEI docker image. E5 models can be run on either GPU or CPU.\n",
"# Find Vertex AI prediction supported accelerators and regions in\n",
"# https://cloud.google.com/vertex-ai/docs/predictions/configure-compute.\n",
"processor = \"NVIDIA_L4\" # @param[\"NVIDIA_TESLA_T4\", \"NVIDIA_TESLA_V100\", \"NVIDIA_L4\", \"NVIDIA_TESLA_A100\", \"CPU\"]\n",
"\n",
"if processor == \"NVIDIA_TESLA_T4\":\n",
" accelerator_type = \"NVIDIA_TESLA_T4\"\n",
" machine_type = \"n1-highmem-16\"\n",
" accelerator_count = 1\n",
"elif processor == \"NVIDIA_TESLA_V100\":\n",
" accelerator_type = \"NVIDIA_TESLA_V100\"\n",
" machine_type = \"n1-highmem-16\"\n",
" accelerator_count = 2\n",
"elif processor == \"NVIDIA_L4\":\n",
" accelerator_type = \"NVIDIA_L4\"\n",
" machine_type = \"g2-standard-8\"\n",
" accelerator_count = 1\n",
"elif processor == \"NVIDIA_TESLA_A100\":\n",
" accelerator_type = \"NVIDIA_TESLA_A100\"\n",
" machine_type = \"a2-highgpu-1g\"\n",
" accelerator_count = 1\n",
"elif processor == \"CPU\":\n",
" accelerator_type = None\n",
" machine_type = None\n",
" accelerator_count = None\n",
"else:\n",
" raise ValueError(f\"Unsupported processor: {processor}\")\n",
"\n",
"\n",
"# The pre-built serving docker images with TEI.\n",
"TEI_DOCKER_URI = \"us-docker.pkg.dev/deeplearning-platform-release/vertex-model-garden/hf-tei.cu125.0-1.ubuntu2204.py310:model-garden.hf-tei-0-1-release_20250727.00_p0\"\n",
"\n",
"# @markdown Set use_dedicated_endpoint to False if you don't want to use [dedicated endpoint](https://cloud.google.com/vertex-ai/docs/general/deployment#create-dedicated-endpoint). Note that [dedicated endpoint does not support VPC Service Controls](https://cloud.google.com/vertex-ai/docs/predictions/choose-endpoint-type), uncheck the box if you are using VPC-SC.\n",
"use_dedicated_endpoint = True # @param {type:\"boolean\"}\n",
"\n",
"# @markdown Click \"Show code\" to see more details."
"Use the `list_deploy_options()` method to view the verified deployment configurations for your selected model. This helps ensure you have sufficient resources (e.g., GPU quota) available to deploy it."
]
},
{
@@ -232,97 +296,95 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "6dY6_ppyObQy"
"id": "zm73g7vFFm9N"
},
"outputs": [],
"source": [
"# @title Deploy model using custom configuration\n",
"# @markdown This section uploads prebuilt E5 models to Model Registry and deploys it to a Vertex AI Endpoint. It might take ~15 minutes to 1 hour to finish depending on the size of the model.\n",
"deploy_options = model.list_deploy_options(concise=True)\n",
"print(deploy_options)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "WjV499VsGwrD"
},
"source": [
"### Deploy the Model\n",
"\n",
"Now that you’ve reviewed the deployment options, use the `deploy()` method to serve the selected open model to a Vertex AI endpoint. Deployment time may vary depending on the model size and infrastructure requirements.\n",
"\n",
"def deploy_model_tei(\n",
" model_name: str,\n",
" model_id: str,\n",
" docker_uri: str,\n",
" machine_type: str = \"g2-standard-8\",\n",
" accelerator_type: str = \"NVIDIA_L4\",\n",
" accelerator_count: int = 1,\n",
" max_model_len: int = 512,\n",
" gpu_memory_utilization: float = 0.9,\n",
" use_dedicated_endpoint: bool = True,\n",
") -> Tuple[aiplatform.Model, aiplatform.Endpoint]:\n",
" \"\"\"Deploys E5 models with TEI on Vertex AI.\n",
"\n",
" Args:\n",
" model_name: Display name of the model.\n",
" model_id: Model ID or path to model weights.\n",
" machine_type: Deployment machine type.\n",
" accelerator_type: Deployment accelerator type.\n",
" accelerator_count: Number of accelerators to use.\n",
" max_model_len: Maximum model length.\n",
" gpu_memory_utilization: Fraction of GPU memory to be used for the model\n",
" executor.\n",
" use_dedicated_endpoint: A dedicated endpoint is an endpoint for online\n",
" prediction,provide a secure connection for private communication\n",
" between on-premises and Google Cloud.\n",
"\n",
" Returns:\n",
" Model instance and endpoint instance.\n",
" \"\"\"\n",
" endpoint = aiplatform.Endpoint.create(\n",
" display_name=f\"{model_name}-endpoint\",\n",
" dedicated_endpoint_enabled=use_dedicated_endpoint,\n",
" )\n",
"\n",
" tei_args = [\n",
" f\"--model-id={model_id}\",\n",
" ]\n",
" serving_env = {\n",
" \"MODEL_ID\": model_id,\n",
" \"DEPLOY_SOURCE\": \"notebook\",\n",
" }\n",
" model = aiplatform.Model.upload(\n",
" display_name=model_name,\n",
" serving_container_image_uri=docker_uri,\n",
" serving_container_args=tei_args,\n",
" serving_container_ports=[7080],\n",
" serving_container_environment_variables=serving_env,\n",
" model_garden_source_model_name=\"publishers/intfloat/models/e5\",\n",
" )\n",
"\n",
" model.deploy(\n",
" endpoint=endpoint,\n",
" machine_type=machine_type,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
" deploy_request_timeout=1800,\n",
" system_labels={\"NOTEBOOK_NAME\": \"model_garden_e5.ipynb\"},\n",
" )\n",
" return model, endpoint\n",
"\n",
"\n",
"common_util.check_quota(\n",
" project_id=PROJECT_ID,\n",
" region=REGION,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
" is_for_training=False,\n",
")\n",
"\n",
"\n",
"LABEL = \"tei\"\n",
"models[LABEL], endpoints[LABEL] = deploy_model_tei(\n",
" model_name=common_util.get_job_name_with_datetime(prefix=\"e5-serve-tei\"),\n",
" model_id=prebuilt_model_id,\n",
" docker_uri=TEI_DOCKER_URI,\n",
" machine_type=machine_type,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
"> **Note**: If the model requires accepting a license agreement (EULA), set the `accept_eula=True` flag in the deploy call. Set `use_dedicated_endpoint` to False if you don't want to use [dedicated endpoint](https://cloud.google.com/vertex-ai/docs/general/deployment#create-dedicated-endpoint)."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "wX1itVTvXdEP"
},
"outputs": [],
"source": [
"use_dedicated_endpoint = True"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "MRmPFEPoGzsB"
},
"outputs": [],
"source": [
"endpoint = model.deploy(\n",
" accept_eula=True,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
")\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "PHBtn8DQp-ID"
},
"source": [
"Alternatively, you can select one of the verified deployment configurations listed above."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "ADsJG8JYqI6c"
},
"outputs": [],
"source": [
"endpoint = model.deploy(\n",
" accept_eula=True,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
" serving_container_image_uri=\"us-docker.pkg.dev/deeplearning-platform-release/gcr.io/huggingface-text-embeddings-inference-cu122.1-2.ubuntu2204\",\n",
" machine_type=\"g2-standard-8\",\n",
" accelerator_type=\"NVIDIA_L4\",\n",
" accelerator_count=1,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "kqSUK2CwsImi"
},
"source": [
"To further customize your deployment, you can configure:\n",
"\n",
"model = models[LABEL]\n",
"endpoint = endpoints[LABEL]"
"- **Compute Resources**: Machine type, replica count (min/max), accelerator type and quantity.\n",
"- **Infrastructure**: Use Spot VMs, reservation affinity, or dedicated endpoints.\n",
"- **Serving Container**: Customize container image, ports, health checks, and environment variables.\n",
"\n",
"See the [Model Garden SDK README](https://github.com/googleapis/python-aiplatform/blob/main/vertexai/model_garden/README.md) for advanced configuration options."
]
},
{
@@ -359,6 +421,8 @@
"\n",
"# @markdown NOTE: Inputs are not limited to 2 queries and 2 passages. To add more inputs, you may modify the code directly.\n",
"\n",
"from torch import Tensor\n",
"\n",
"query1 = \"how much protein should a female eat?\" # @param {type: \"string\"}\n",
"query2 = \"summit define\" # @param {type: \"string\"}\n",
"passage1 = \"As a general guideline, the CDC's average requirement of protein for women ages 19 to 70 is 46 grams per day. But, as you can see from this chart, you'll need to increase that if you're expecting or training for a marathon. Check out the chart below to see how much protein you should be eating each day.\" # @param {type: \"string\"}\n",
@@ -499,16 +563,10 @@
"source": [
"# @title Delete the models and endpoints\n",
"\n",
"# @markdown Delete the experiment models and endpoints to recycle the resources\n",
"# @markdown and avoid unnecessary continuous charges that may incur.\n",
"# @markdown Delete the endpoint.\n",
"\n",
"# Undeploy model and delete endpoint.\n",
"for endpoint in endpoints.values():\n",
" endpoint.delete(force=True)\n",
"\n",
"# Delete models.\n",
"for model in models.values():\n",
" model.delete()"
"if endpoint:\n",
" endpoint.delete(force=True)"
]
}
],
@@ -0,0 +1,182 @@
{
"cells": [
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "YsFQaflOxP_g"
},
"outputs": [],
"source": [
"# Copyright 2025 Google LLC\n",
"#\n",
"# Licensed under the Apache License, Version 2.0 (the \"License\");\n",
"# you may not use this file except in compliance with the License.\n",
"# You may obtain a copy of the License at\n",
"#\n",
"# https://www.apache.org/licenses/LICENSE-2.0\n",
"#\n",
"# Unless required by applicable law or agreed to in writing, software\n",
"# distributed under the License is distributed on an \"AS IS\" BASIS,\n",
"# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.\n",
"# See the License for the specific language governing permissions and\n",
"# limitations under the License."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "Y-uUs1OfxcjA"
},
"source": [
"# Vertex AI Model Garden - EmbeddingGemma (Local Inference)\n",
"\n",
"<table><tbody><tr>\n",
" <td style=\"text-align: center\">\n",
" <a href=\"https://console.cloud.google.com/vertex-ai/notebooks/deploy-notebook?download_url=https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/main/notebooks/community/model_garden/model_garden_embedding_gemma_local_inference.ipynb\">\n",
" <img alt=\"Workbench logo\" src=\"https://lh3.googleusercontent.com/UiNooY4LUgW_oTvpsNhPpQzsstV5W8F7rYgxgGBD85cWJoLmrOzhVs_ksK_vgx40SHs7jCqkTkCk=e14-rj-sc0xffffff-h130-w32\" width=\"32px\"><br> Run in Workbench\n",
" </a>\n",
" </td>\n",
" <td style=\"text-align: center\">\n",
" <a href=\"https://console.cloud.google.com/vertex-ai/colab/import/https:%2F%2Fraw.githubusercontent.com%2FGoogleCloudPlatform%2Fvertex-ai-samples%2Fmain%2Fnotebooks%2Fcommunity%2Fmodel_garden%2Fmodel_garden_embedding_gemma_local_inference.ipynb\">\n",
" <img alt=\"Google Cloud Colab Enterprise logo\" src=\"https://lh3.googleusercontent.com/JmcxdQi-qOpctIvWKgPtrzZdJJK-J3sWE1RsfjZNwshCFgE_9fULcNpuXYTilIR2hjwN\" width=\"32px\"><br> Run in Colab Enterprise\n",
" </a>\n",
" </td>\n",
" <td style=\"text-align: center\">\n",
" <a href=\"https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/model_garden/model_garden_embedding_gemma_local_inference.ipynb\">\n",
" <img alt=\"GitHub logo\" src=\"https://github.githubassets.com/assets/GitHub-Mark-ea2971cee799.png\" width=\"32px\"><br> View on GitHub\n",
" </a>\n",
" </td>\n",
"</tr></tbody></table>"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "-c_LNERL0MNv"
},
"source": [
"## Overview\n",
"\n",
"This notebook demonstrates how to install the necessary libraries and run local inference with [EmbeddingGemma](https://huggingface.co/google/embeddinggemma-300m) model in a [Colab Enterprise Instance](https://cloud.google.com/colab/docs) or a [Workbench Instance](https://cloud.google.com/vertex-ai/docs/workbench/instances).\n",
"\n",
"### Objective\n",
"\n",
"Run local inference with EmbeddingGemma model.\n",
"\n",
"### Costs\n",
"\n",
"This tutorial uses billable components of Google Cloud:\n",
"\n",
"* Vertex AI\n",
"\n",
"Learn about [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing) and use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to generate a cost estimate based on your projected usage."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "zsNTHGCK1FU7"
},
"source": [
"## Install dependencies\n",
"\n",
"After running the following block, click \"RESTART SESSION\" to apply the installed dependencies."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "xle1_0ns1r10"
},
"outputs": [],
"source": [
"!pip install --upgrade sentence-transformers\n",
"!pip install git+https://github.com/huggingface/transformers@v4.56.0-Embedding-Gemma-preview\n",
"!pip uninstall protobuf -y\n",
"!pip install protobuf==3.20.0"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "a0eE5dWx1s1j"
},
"source": [
"## Hugging Face Login\n",
"\n",
"The following code block will prompt you to enter your Hugging Face access token.\n",
"\n",
"If you don't already have a Hugging Face access token, follow the [Hugging Face documentation](https://huggingface.co/docs/hub/en/security-tokens) to create an access token with \"read\" permission. You can find your existing access tokens in the Hugging Face [Access Token](https://huggingface.co/settings/tokens) page.\n",
"\n",
"Make sure you have accepted the model agreement to access the model."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "oR2eJczK3mNF"
},
"outputs": [],
"source": [
"from huggingface_hub import notebook_login\n",
"\n",
"notebook_login()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "arlTS97v29Uu"
},
"source": [
"## Run local inference"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "uTLuTO3N1z_B"
},
"outputs": [],
"source": [
"from sentence_transformers import SentenceTransformer\n",
"\n",
"# Load model from Hugging Face.\n",
"model = SentenceTransformer(\"google/embeddinggemma-300m\")\n",
"\n",
"# Run inference.\n",
"sentences = [\n",
" \"The weather is lovely today.\",\n",
" \"It's so sunny outside!\",\n",
" \"He drove to the stadium.\",\n",
"]\n",
"embeddings = model.encode(sentences)\n",
"print(embeddings.shape)\n",
"\n",
"# Get the similarity scores for the embeddings\n",
"similarities = model.similarity(embeddings, embeddings)\n",
"print(similarities)"
]
}
],
"metadata": {
"colab": {
"name": "model_garden_embedding_gemma_local_inference.ipynb",
"toc_visible": true
},
"kernelspec": {
"display_name": "Python 3",
"name": "python3"
}
},
"nbformat": 4,
"nbformat_minor": 0
}
@@ -208,7 +208,7 @@
"from IPython.display import Markdown, display\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -154,7 +154,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"LABEL = \"tgi\"\n",
@@ -146,7 +146,7 @@
"\n",
"# Import the necessary packages\n",
"! rm -rf vertex-ai-samples && git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"! cd vertex-ai-samples && git reset --hard c45f6a4f4d32e31a050f0e4ba52824b0caf4eda3\n",
"! cd vertex-ai-samples && git reset --hard 7ae13b346a72ee2a2dc8152dd40c6ddd72d6c810\n",
"\n",
"import datetime\n",
"import importlib\n",
@@ -162,7 +162,7 @@
" ! pip install --upgrade tensorflow\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -859,7 +859,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -868,6 +867,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -59,34 +59,43 @@
"source": [
"## Overview\n",
"\n",
"This notebook demonstrates deploying Gemma 3 models on GPU using [vLLM](https://github.com/vllm-project/vllm).\n",
"This notebook demonstrates how to deploy a **Gemma 3** open model on Google Cloud Vertex AI.\n",
"\n",
"### Objectives\n",
"\n",
"### Objective\n",
"- Deploy Gemma 3 using containerized backends like [vLLM](https://github.com/vllm-project/vllm) on GPU.\n",
"- Use the deployed model to serve chat completion requests for both text and multimodal inputs.\n",
"\n",
"- Deploy Gemma 3 with vLLM on GPU\n",
"### File a Bug\n",
"\n",
"### File a bug\n",
"\n",
"File a bug on [GitHub](https://github.com/GoogleCloudPlatform/vertex-ai-samples/issues/new) if you encounter any issue with the notebook.\n",
"If you encounter issues with this notebook, report them on [GitHub](https://github.com/GoogleCloudPlatform/vertex-ai-samples/issues/new).\n",
"\n",
"### Costs\n",
"\n",
"This tutorial uses billable components of Google Cloud:\n",
"\n",
"* Vertex AI\n",
"* Cloud Storage\n",
"- Vertex AI\n",
"- Cloud Storage\n",
"\n",
"Learn about [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing), [Cloud Storage pricing](https://cloud.google.com/storage/pricing), and use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to generate a cost estimate based on your projected usage."
"Refer to the [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing) and [Cloud Storage pricing](https://cloud.google.com/storage/pricing) pages for more information. Use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to estimate your projected costs."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "264c07757582"
"id": "jeYw-Czg-DFy"
},
"source": [
"## Before you begin"
"## Get Started"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "KgyhGvEzBDkj"
},
"source": [
"### Install Vertex AI SDK and other required packages"
]
},
{
@@ -94,87 +103,169 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "2707b02ef5df"
"id": "iCacdLqG-IsH"
},
"outputs": [],
"source": [
"# @title Setup Google Cloud project\n",
"%pip install --upgrade --force-reinstall --quiet 'google-cloud-aiplatform>=1.106.0' 'openai' 'google-auth==2.27.0' 'requests==2.32.3'"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "HUKCrpBy-3yf"
},
"source": [
"### Authenticate the Notebook Environment (Colab only)\n",
"\n",
"# @markdown 1. [Make sure that billing is enabled for your project](https://cloud.google.com/billing/docs/how-to/modify-project).\n",
"If you're running this notebook in Google Colab, run the following cell to authenticate."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "JXwCT1kn-3Gu"
},
"outputs": [],
"source": [
"import sys\n",
"\n",
"# @markdown 2. **[Optional]** Set region. If not set, the region will be set automatically according to Colab Enterprise environment.\n",
"if \"google.colab\" in sys.modules:\n",
" from google.colab import auth\n",
"\n",
"REGION = \"\" # @param {type:\"string\"}\n",
" auth.authenticate_user()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "AcW2nwB8-7yC"
},
"source": [
"### Set Google Cloud Project Information\n",
"\n",
"# @markdown 3. If you want to run predictions with A100 80GB or H100 GPUs, we recommend using the regions listed below. **NOTE:** Make sure you have associated quota in selected regions. Click the links to see your current quota for each GPU type: [Nvidia A100 80GB](https://console.cloud.google.com/iam-admin/quotas?metric=aiplatform.googleapis.com%2Fcustom_model_serving_nvidia_a100_80gb_gpus), [Nvidia H100 80GB](https://console.cloud.google.com/iam-admin/quotas?metric=aiplatform.googleapis.com%2Fcustom_model_serving_nvidia_h100_gpus). You can request for quota following the instructions at [\"Request a higher quota\"](https://cloud.google.com/docs/quota/view-manage#requesting_higher_quota).\n",
"To get started with Vertex AI, ensure you have an existing Google Cloud project and that the [Vertex AI API is enabled](https://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com).\n",
"\n",
"# @markdown > | Machine Type | Accelerator Type | Recommended Regions |\n",
"# @markdown | ----------- | ----------- | ----------- |\n",
"# @markdown | a2-ultragpu-1g | 1 NVIDIA_A100_80GB | us-central1, us-east4, europe-west4, asia-southeast1, us-east4 |\n",
"# @markdown | a3-highgpu-2g | 2 NVIDIA_H100_80GB | us-west1, asia-southeast1, europe-west4 |\n",
"# @markdown | a3-highgpu-4g | 4 NVIDIA_H100_80GB | us-west1, asia-southeast1, europe-west4 |\n",
"# @markdown | a3-highgpu-8g | 8 NVIDIA_H100_80GB | us-central1, europe-west4, us-west1, asia-southeast1 |\n",
"\n",
"# Upgrade Vertex AI SDK.\n",
"! pip3 install --upgrade --quiet 'google-cloud-aiplatform==1.103.0'\n",
"\n",
"# Import the necessary packages\n",
"import importlib\n",
"See the guide on [setting up your project and development environment](https://cloud.google.com/vertex-ai/docs/start/cloud-environment). Also confirm that [billing is enabled](https://cloud.google.com/billing/docs/how-to/modify-project).\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "eIVLp0oE--k-"
},
"outputs": [],
"source": [
"# Use the environment variable if the user doesn't provide Project ID.\n",
"import os\n",
"from typing import Tuple\n",
"\n",
"from google.cloud import aiplatform\n",
"\n",
"# Upgrade Vertex AI SDK.\n",
"if os.environ.get(\"VERTEX_PRODUCT\") != \"COLAB_ENTERPRISE\":\n",
" ! pip install --upgrade tensorflow\n",
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
")\n",
"\n",
"# Initialize models and endpoints as a dict\n",
"models, endpoints = {}, {}\n",
"\n",
"\n",
"# Get the default cloud project id.\n",
"PROJECT_ID = os.environ[\"GOOGLE_CLOUD_PROJECT\"]\n",
"\n",
"# Get the default region for launching jobs.\n",
"if not REGION:\n",
" REGION = os.environ[\"GOOGLE_CLOUD_REGION\"]\n",
"\n",
"# Enable the Vertex AI API and Compute Engine API, if not already.\n",
"print(\"Enabling Vertex AI API and Compute Engine API.\")\n",
"! gcloud services enable aiplatform.googleapis.com compute.googleapis.com\n",
"\n",
"# Initialize Vertex AI API.\n",
"print(\"Initializing Vertex AI API.\")\n",
"aiplatform.init(project=PROJECT_ID, location=REGION)\n",
"\n",
"# Gets the default SERVICE_ACCOUNT.\n",
"shell_output = ! gcloud projects describe $PROJECT_ID\n",
"project_number = shell_output[-1].split(\":\")[1].strip().replace(\"'\", \"\")\n",
"SERVICE_ACCOUNT = f\"{project_number}-compute@developer.gserviceaccount.com\"\n",
"print(\"Using this default Service Account:\", SERVICE_ACCOUNT)\n",
"\n",
"! gcloud config set project $PROJECT_ID\n",
"import vertexai\n",
"\n",
"vertexai.init(\n",
" project=PROJECT_ID,\n",
" location=REGION,\n",
"PROJECT_ID = \"[your-project-id]\" # @param {type: \"string\", placeholder: \"[your-project-id]\", isTemplate: true}\n",
"\n",
"if not PROJECT_ID or PROJECT_ID == \"[your-project-id]\":\n",
" PROJECT_ID = str(os.environ.get(\"GOOGLE_CLOUD_PROJECT\"))\n",
"\n",
"REGION = os.environ.get(\"GOOGLE_CLOUD_REGION\", \"us-central1\")\n",
"\n",
"vertexai.init(project=PROJECT_ID, location=REGION)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "Q0CXrvcZH_aw"
},
"source": [
"### Import libraries"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "3G2UXB82ICs6"
},
"outputs": [],
"source": [
"from vertexai import model_garden"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "upYRiGtP_-iN"
},
"source": [
"## Deploy model"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "H2WC_0hXDVXc"
},
"source": [
"### Choose model variant"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "u41zbNa2EoFq"
},
"source": [
"You can proceed with the default model variant or select a different one."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "-fgC4NLSDkF7"
},
"outputs": [],
"source": [
"model_version = \"gemma-3-1b-it\" # @param [\"gemma-3-12b-it\", \"gemma-3-12b-pt\", \"gemma-3-1b-it\", \"gemma-3-1b-pt\", \"gemma-3-270m\", \"gemma-3-270m-it\", \"gemma-3-27b-it\", \"gemma-3-27b-pt\", \"gemma-3-4b-it\", \"gemma-3-4b-pt\"] {isTemplate:true}\n",
"MODEL_NAME = f\"google/gemma3@{model_version}\""
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "VRnUgU8LF3_i"
},
"source": [
"To see all deployable model variants available in Model Garden, use:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "-QLd-wshF6sB"
},
"outputs": [],
"source": [
"all_model_versions = model_garden.list_deployable_models(\n",
" model_filter=\"gemma3\", list_hf_models=False\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "GFfqpQm8BNwZ"
"id": "N0UeFHa2GO63"
},
"source": [
"## Deploy Gemma 3 1B models with vLLM on GPU"
"Once you've selected a model variant, initialize it:"
]
},
{
@@ -182,373 +273,22 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "W1PYsnz2JXHm"
"id": "GZiV3trBBcA3"
},
"outputs": [],
"source": [
"# @markdown Set the model to deploy.\n",
"\n",
"base_model_name = \"gemma-3-1b-it\" # @param [\"gemma-3-1b-pt\", \"gemma-3-1b-it\"] {isTemplate:true}\n",
"hf_model_id = \"google/\" + base_model_name\n",
"PUBLISHER_MODEL_NAME = f\"publishers/google/models/gemma3@{base_model_name}\"\n",
"model_id = f\"gs://vertex-model-garden-restricted-us/gemma3/{base_model_name}\"\n",
"\n",
"# The pre-built serving docker image.\n",
"VLLM_DOCKER_URI = \"us-docker.pkg.dev/vertex-ai/vertex-vision-model-garden-dockers/pytorch-vllm-serve:20250312_0916_RC01\"\n",
"\n",
"# @markdown Set use_dedicated_endpoint to False if you don't want to use [dedicated endpoint](https://cloud.google.com/vertex-ai/docs/general/deployment#create-dedicated-endpoint). Note that [dedicated endpoint does not support VPC Service Controls](https://cloud.google.com/vertex-ai/docs/predictions/choose-endpoint-type), uncheck the box if you are using VPC-SC.\n",
"use_dedicated_endpoint = True # @param {type:\"boolean\"}\n",
"\n",
"# @markdown Find Vertex AI prediction supported accelerators and regions at https://cloud.google.com/vertex-ai/docs/predictions/configure-compute.\n",
"accelerator_type = \"NVIDIA_L4\"\n",
"machine_type = \"g2-standard-12\"\n",
"accelerator_count = 1\n",
"\n",
"common_util.check_quota(\n",
" project_id=PROJECT_ID,\n",
" region=REGION,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
" is_for_training=False,\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "d2SO5S5fMKE2"
},
"outputs": [],
"source": [
"# @title [Option 1] Deploy with Model Garden SDK\n",
"\n",
"LABEL = \"sdk-deploy-1b\"\n",
"# @markdown Deploy with Gen AI model-centric SDK. This section uploads the prebuilt model to Model Registry and deploys it to a Vertex AI Endpoint. It takes 15 minutes to 1 hour to finish depending on the size of the model. See [use open models with Vertex AI](https://cloud.google.com/vertex-ai/generative-ai/docs/open-models/use-open-models) for documentation on other use cases.\n",
"from vertexai import model_garden\n",
"\n",
"model = model_garden.OpenModel(PUBLISHER_MODEL_NAME)\n",
"endpoints[LABEL] = model.deploy(\n",
" machine_type=machine_type,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
" accept_eula=True, # Accept the End User License Agreement (EULA) on the model card before deploy. Otherwise, the deployment will be forbidden.\n",
")\n",
"\n",
"endpoint = endpoints[LABEL]"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "TBNJYZMlBNwZ"
},
"outputs": [],
"source": [
"# @title [Option 2] Deploy with customized configs\n",
"\n",
"# @markdown This section uploads Gemma 3 1B models to Model Registry and deploys them to a Vertex Prediction Endpoint. It takes 15 minutes to 30 minutes to finish.\n",
"\n",
"gpu_memory_utilization = 0.95\n",
"max_model_len = 32768\n",
"\n",
"\n",
"def deploy_model_vllm(\n",
" model_name: str,\n",
" model_id: str,\n",
" publisher: str,\n",
" publisher_model_id: str,\n",
" base_model_id: str = None,\n",
" machine_type: str = \"g2-standard-8\",\n",
" accelerator_type: str = \"NVIDIA_L4\",\n",
" accelerator_count: int = 1,\n",
" gpu_memory_utilization: float = 0.9,\n",
" max_model_len: int = 4096,\n",
" dtype: str = \"auto\",\n",
" enable_trust_remote_code: bool = False,\n",
" enforce_eager: bool = False,\n",
" enable_lora: bool = False,\n",
" enable_chunked_prefill: bool = False,\n",
" enable_prefix_cache: bool = False,\n",
" host_prefix_kv_cache_utilization_target: float = 0.0,\n",
" max_loras: int = 1,\n",
" max_cpu_loras: int = 8,\n",
" use_dedicated_endpoint: bool = False,\n",
" max_num_seqs: int = 256,\n",
" model_type: str = None,\n",
" enable_llama_tool_parser: bool = False,\n",
") -> Tuple[aiplatform.Model, aiplatform.Endpoint]:\n",
" \"\"\"Deploys trained models with vLLM into Vertex AI.\"\"\"\n",
" endpoint = aiplatform.Endpoint.create(\n",
" display_name=f\"{model_name}-endpoint\",\n",
" dedicated_endpoint_enabled=use_dedicated_endpoint,\n",
" )\n",
"\n",
" if not base_model_id:\n",
" base_model_id = model_id\n",
"\n",
" # See https://docs.vllm.ai/en/latest/models/engine_args.html for a list of possible arguments with descriptions.\n",
" vllm_args = [\n",
" \"python\",\n",
" \"-m\",\n",
" \"vllm.entrypoints.api_server\",\n",
" \"--host=0.0.0.0\",\n",
" \"--port=8080\",\n",
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
" f\"--max-cpu-loras={max_cpu_loras}\",\n",
" f\"--max-num-seqs={max_num_seqs}\",\n",
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
" if enforce_eager:\n",
" vllm_args.append(\"--enforce-eager\")\n",
"\n",
" if enable_lora:\n",
" vllm_args.append(\"--enable-lora\")\n",
"\n",
" if enable_chunked_prefill:\n",
" vllm_args.append(\"--enable-chunked-prefill\")\n",
"\n",
" if enable_prefix_cache:\n",
" vllm_args.append(\"--enable-prefix-caching\")\n",
"\n",
" if 0 < host_prefix_kv_cache_utilization_target < 1:\n",
" vllm_args.append(\n",
" f\"--host-prefix-kv-cache-utilization-target={host_prefix_kv_cache_utilization_target}\"\n",
" )\n",
"\n",
" if model_type:\n",
" vllm_args.append(f\"--model-type={model_type}\")\n",
"\n",
" if enable_llama_tool_parser:\n",
" vllm_args.append(\"--enable-auto-tool-choice\")\n",
" vllm_args.append(\"--tool-call-parser=vertex-llama-3\")\n",
"\n",
" env_vars = {\n",
" \"MODEL_ID\": base_model_id,\n",
" \"DEPLOY_SOURCE\": \"notebook\",\n",
" }\n",
"\n",
" # HF_TOKEN is not a compulsory field and may not be defined.\n",
" try:\n",
" if HF_TOKEN:\n",
" env_vars[\"HF_TOKEN\"] = HF_TOKEN\n",
" except NameError:\n",
" pass\n",
"\n",
" model = aiplatform.Model.upload(\n",
" display_name=model_name,\n",
" serving_container_image_uri=VLLM_DOCKER_URI,\n",
" serving_container_args=vllm_args,\n",
" serving_container_ports=[8080],\n",
" serving_container_predict_route=\"/generate\",\n",
" serving_container_health_route=\"/ping\",\n",
" serving_container_environment_variables=env_vars,\n",
" serving_container_shared_memory_size_mb=(16 * 1024), # 16 GB\n",
" serving_container_deployment_timeout=7200,\n",
" model_garden_source_model_name=(\n",
" f\"publishers/{publisher}/models/{publisher_model_id}\"\n",
" ),\n",
" )\n",
" print(\n",
" f\"Deploying {model_name} on {machine_type} with {accelerator_count} {accelerator_type} GPU(s).\"\n",
" )\n",
" model.deploy(\n",
" endpoint=endpoint,\n",
" machine_type=machine_type,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
" deploy_request_timeout=1800,\n",
" system_labels={\n",
" \"NOTEBOOK_NAME\": \"model_garden_gemma3_deployment_on_vertex.ipynb\",\n",
" \"NOTEBOOK_ENVIRONMENT\": common_util.get_deploy_source(),\n",
" },\n",
" )\n",
" print(\"endpoint_name:\", endpoint.name)\n",
"\n",
" return model, endpoint\n",
"\n",
"\n",
"LABEL = \"custom-deploy-1b\"\n",
"\n",
"models[LABEL], endpoints[LABEL] = deploy_model_vllm(\n",
" model_name=common_util.get_job_name_with_datetime(prefix=\"gemma3-serve\"),\n",
" model_id=model_id,\n",
" publisher=\"google\",\n",
" publisher_model_id=\"gemma3\",\n",
" base_model_id=hf_model_id,\n",
" machine_type=machine_type,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
" gpu_memory_utilization=gpu_memory_utilization,\n",
" max_model_len=max_model_len,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
")\n",
"\n",
"model = models[LABEL]\n",
"endpoint = endpoints[LABEL]"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "rDHsCOqvFYBi"
},
"outputs": [],
"source": [
"# @title Raw predict\n",
"\n",
"# @markdown Once deployment succeeds, you can send requests to the endpoint with text prompts. Sampling parameters supported by vLLM can be found [here](https://docs.vllm.ai/en/latest/dev/sampling_params.html).\n",
"\n",
"# @markdown Example:\n",
"\n",
"# @markdown ```\n",
"# @markdown Human: What is a car?\n",
"# @markdown Assistant: A car, or a motor car, is a road-connected human-transportation system used to move people or goods from one place to another. The term also encompasses a wide range of vehicles, including motorboats, trains, and aircrafts. Cars typically have four wheels, a cabin for passengers, and an engine or motor. They have been around since the early 19th century and are now one of the most popular forms of transportation, used for daily commuting, shopping, and other purposes.\n",
"# @markdown ```\n",
"# @markdown Additionally, you can moderate the generated text with Vertex AI. See [Moderate text documentation](https://cloud.google.com/natural-language/docs/moderating-text) for more details.\n",
"\n",
"# Loads an existing endpoint instance using the endpoint name:\n",
"# - Using `endpoint_name = endpoint.name` allows us to get the\n",
"# endpoint name of the endpoint `endpoint` created in the cell\n",
"# above.\n",
"# - Alternatively, you can set `endpoint_name = \"1234567890123456789\"` to load\n",
"# an existing endpoint with the ID 1234567890123456789.\n",
"# You may uncomment the code below to load an existing endpoint.\n",
"\n",
"# endpoint_name = \"\" # @param {type:\"string\"}\n",
"# aip_endpoint_name = (\n",
"# f\"projects/{PROJECT_ID}/locations/{REGION}/endpoints/{endpoint_name}\"\n",
"# )\n",
"# endpoint = aiplatform.Endpoint(aip_endpoint_name)\n",
"\n",
"prompt = \"What is a car?\" # @param {type: \"string\"}\n",
"# @markdown If you encounter an issue like `ServiceUnavailable: 503 Took too long to respond when processing`, you can reduce the maximum number of output tokens, by lowering `max_tokens`.\n",
"max_tokens = 50 # @param {type:\"integer\"}\n",
"temperature = 1.0 # @param {type:\"number\"}\n",
"top_p = 1.0 # @param {type:\"number\"}\n",
"top_k = 1 # @param {type:\"integer\"}\n",
"# @markdown Set `raw_response` to `True` to obtain the raw model output. Set `raw_response` to `False` to apply additional formatting in the structure of `\"Prompt:\\n{prompt.strip()}\\nOutput:\\n{output}\"`.\n",
"raw_response = False # @param {type:\"boolean\"}\n",
"\n",
"# Overrides parameters for inferences.\n",
"instances = [\n",
" {\n",
" \"prompt\": prompt,\n",
" \"max_tokens\": max_tokens,\n",
" \"temperature\": temperature,\n",
" \"top_p\": top_p,\n",
" \"top_k\": top_k,\n",
" \"raw_response\": raw_response,\n",
" },\n",
"]\n",
"response = endpoint.predict(\n",
" instances=instances, use_dedicated_endpoint=use_dedicated_endpoint\n",
")\n",
"\n",
"for prediction in response.predictions:\n",
" print(prediction)\n",
"\n",
"# @markdown Click \"Show Code\" to see more details."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "LSG9ITWTbTb7"
},
"outputs": [],
"source": [
"# @title Chat completion\n",
"\n",
"if use_dedicated_endpoint:\n",
" DEDICATED_ENDPOINT_DNS = endpoint.gca_resource.dedicated_endpoint_dns\n",
"ENDPOINT_RESOURCE_NAME = endpoint.resource_name\n",
"\n",
"# @title Chat Completions Inference\n",
"\n",
"# @markdown Once deployment succeeds, you can send requests to the endpoint using the OpenAI SDK.\n",
"\n",
"# @markdown First you will need to install the SDK and some auth-related dependencies.\n",
"\n",
"! pip install -qU openai google-auth requests\n",
"\n",
"# @markdown Next fill out some request parameters:\n",
"\n",
"user_message = \"How is your day going?\" # @param {type: \"string\"}\n",
"# @markdown If you encounter the issue like `ServiceUnavailable: 503 Took too long to respond when processing`, you can reduce the maximum number of output tokens, such as set `max_tokens` as 20.\n",
"max_tokens = 50 # @param {type: \"integer\"}\n",
"temperature = 1.0 # @param {type: \"number\"}\n",
"stream = False # @param {type: \"boolean\"}\n",
"\n",
"# @markdown Now we can send a request.\n",
"\n",
"import google.auth\n",
"import openai\n",
"\n",
"creds, project = google.auth.default()\n",
"auth_req = google.auth.transport.requests.Request()\n",
"creds.refresh(auth_req)\n",
"\n",
"BASE_URL = (\n",
" f\"https://{REGION}-aiplatform.googleapis.com/v1beta1/{ENDPOINT_RESOURCE_NAME}\"\n",
")\n",
"try:\n",
" if use_dedicated_endpoint:\n",
" BASE_URL = f\"https://{DEDICATED_ENDPOINT_DNS}/v1beta1/{ENDPOINT_RESOURCE_NAME}\"\n",
"except NameError:\n",
" pass\n",
"\n",
"client = openai.OpenAI(base_url=BASE_URL, api_key=creds.token)\n",
"\n",
"model_response = client.chat.completions.create(\n",
" model=\"\",\n",
" messages=[{\"role\": \"user\", \"content\": user_message}],\n",
" temperature=temperature,\n",
" max_tokens=max_tokens,\n",
" stream=stream,\n",
")\n",
"\n",
"if stream:\n",
" usage = None\n",
" contents = []\n",
" for chunk in model_response:\n",
" if chunk.usage is not None:\n",
" usage = chunk.usage\n",
" continue\n",
" print(chunk.choices[0].delta.content, end=\"\")\n",
" contents.append(chunk.choices[0].delta.content)\n",
" print(f\"\\n\\n{usage}\")\n",
"else:\n",
" print(model_response)\n",
"\n",
"# @markdown Click \"Show Code\" to see more details."
"model = model_garden.OpenModel(MODEL_NAME)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "_0IYHNsYJO55"
"id": "-0cL378wFlvf"
},
"source": [
"## Deploy Gemma 3 4B, 12B and 27B multimodal models with vLLM on GPU"
"### Check the Deployment Configuration\n",
"\n",
"Use the `list_deploy_options()` method to view the verified deployment configurations for your selected model. This helps ensure you have sufficient resources (e.g., GPU quota) available to deploy it."
]
},
{
@@ -556,71 +296,61 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "zpKbaFK_Jeny"
"id": "zm73g7vFFm9N"
},
"outputs": [],
"source": [
"# @markdown Set the model to deploy.\n",
"deploy_options = model.list_deploy_options(concise=True)\n",
"print(deploy_options)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "WjV499VsGwrD"
},
"source": [
"### Deploy the Model\n",
"\n",
"base_model_name = \"gemma-3-4b-it\" # @param [\"gemma-3-4b-pt\", \"gemma-3-4b-it\", \"gemma-3-12b-pt\", \"gemma-3-12b-it\", \"gemma-3-27b-pt\", \"gemma-3-27b-it\"] {isTemplate:true}\n",
"hf_model_id = \"google/\" + base_model_name\n",
"PUBLISHER_MODEL_NAME = f\"publishers/google/models/gemma3@{base_model_name}\"\n",
"model_id = f\"gs://vertex-model-garden-restricted-us/gemma3/{base_model_name}\"\n",
"Now that you’ve reviewed the deployment options, use the `deploy()` method to serve the selected open model to a Vertex AI endpoint. Deployment time may vary depending on the model size and infrastructure requirements.\n",
"\n",
"# The pre-built serving docker image.\n",
"VLLM_DOCKER_URI = \"us-docker.pkg.dev/vertex-ai/vertex-vision-model-garden-dockers/pytorch-vllm-serve:20250312_0916_RC01\"\n",
"\n",
"# @markdown Set use_dedicated_endpoint to False if you don't want to use [dedicated endpoint](https://cloud.google.com/vertex-ai/docs/general/deployment#create-dedicated-endpoint). Note that [dedicated endpoint does not support VPC Service Controls](https://cloud.google.com/vertex-ai/docs/predictions/choose-endpoint-type), uncheck the box if you are using VPC-SC.\n",
"use_dedicated_endpoint = True # @param {type:\"boolean\"}\n",
"\n",
"# @markdown Find Vertex AI prediction supported accelerators and regions at https://cloud.google.com/vertex-ai/docs/predictions/configure-compute.\n",
"if \"4b\" in model_id or \"12b\" in model_id:\n",
" accelerator_type = \"NVIDIA_H100_80GB\"\n",
" machine_type = \"a3-highgpu-2g\"\n",
" accelerator_count = 2\n",
"elif \"27b\" in model_id:\n",
" accelerator_type = \"NVIDIA_H100_80GB\"\n",
" machine_type = \"a3-highgpu-4g\"\n",
" accelerator_count = 4\n",
"else:\n",
" raise ValueError(\n",
" f\"Recommended machine settings not found for model: {base_model_name}.\"\n",
" )\n",
"\n",
"common_util.check_quota(\n",
" project_id=PROJECT_ID,\n",
" region=REGION,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
" is_for_training=False,\n",
"> **Note**: If the model requires accepting a license agreement (EULA), set the `accept_eula=True` flag in the deploy call. Set `use_dedicated_endpoint` to False if you don't want to use [dedicated endpoint](https://cloud.google.com/vertex-ai/docs/general/deployment#create-dedicated-endpoint)."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "wX1itVTvXdEP"
},
"outputs": [],
"source": [
"use_dedicated_endpoint = True"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "MRmPFEPoGzsB"
},
"outputs": [],
"source": [
"endpoint = model.deploy(\n",
" accept_eula=True,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"cell_type": "markdown",
"metadata": {
"cellView": "form",
"id": "CA_dsYyjLy1b"
"id": "PHBtn8DQp-ID"
},
"outputs": [],
"source": [
"# @title [Option 1] Deploy with Model Garden SDK\n",
"\n",
"LABEL = \"sdk-deploy\"\n",
"# @markdown Deploy with Gen AI model-centric SDK. This section uploads the prebuilt model to Model Registry and deploys it to a Vertex AI Endpoint. It takes 15 minutes to 1 hour to finish depending on the size of the model. See [use open models with Vertex AI](https://cloud.google.com/vertex-ai/generative-ai/docs/open-models/use-open-models) for documentation on other use cases.\n",
"from vertexai import model_garden\n",
"\n",
"model = model_garden.OpenModel(PUBLISHER_MODEL_NAME)\n",
"endpoints[LABEL] = model.deploy(\n",
" machine_type=machine_type,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
" accept_eula=True, # Accept the End User License Agreement (EULA) on the model card before deploy. Otherwise, the deployment will be forbidden.\n",
")\n",
"\n",
"endpoint = endpoints[LABEL]"
"Alternatively, you can select one of the verified deployment configurations listed above."
]
},
{
@@ -628,161 +358,33 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "plirtyNxJO55"
"id": "ADsJG8JYqI6c"
},
"outputs": [],
"source": [
"# @title [Option 2] Deploy with customized configs\n",
"\n",
"# @markdown This section uploads Gemma 3 multimodal models to Model Registry and deploys them to a Vertex Prediction Endpoint. It takes 15 minutes to 1 hour to finish.\n",
"\n",
"gpu_memory_utilization = 0.95\n",
"max_model_len = 131072\n",
"\n",
"LABEL = \"multimodal-deploy\"\n",
"\n",
"\n",
"def deploy_model_vllm(\n",
" model_name: str,\n",
" model_id: str,\n",
" publisher: str,\n",
" publisher_model_id: str,\n",
" base_model_id: str = None,\n",
" machine_type: str = \"g2-standard-8\",\n",
" accelerator_type: str = \"NVIDIA_L4\",\n",
" accelerator_count: int = 1,\n",
" gpu_memory_utilization: float = 0.9,\n",
" max_model_len: int = 4096,\n",
" dtype: str = \"auto\",\n",
" enable_trust_remote_code: bool = False,\n",
" enforce_eager: bool = False,\n",
" enable_lora: bool = False,\n",
" enable_chunked_prefill: bool = False,\n",
" enable_prefix_cache: bool = False,\n",
" host_prefix_kv_cache_utilization_target: float = 0.0,\n",
" max_loras: int = 1,\n",
" max_cpu_loras: int = 8,\n",
" use_dedicated_endpoint: bool = False,\n",
" max_num_seqs: int = 256,\n",
" model_type: str = None,\n",
" enable_llama_tool_parser: bool = False,\n",
") -> Tuple[aiplatform.Model, aiplatform.Endpoint]:\n",
" \"\"\"Deploys trained models with vLLM into Vertex AI.\"\"\"\n",
" endpoint = aiplatform.Endpoint.create(\n",
" display_name=f\"{model_name}-endpoint\",\n",
" dedicated_endpoint_enabled=use_dedicated_endpoint,\n",
" )\n",
"\n",
" if not base_model_id:\n",
" base_model_id = model_id\n",
"\n",
" # See https://docs.vllm.ai/en/latest/models/engine_args.html for a list of possible arguments with descriptions.\n",
" vllm_args = [\n",
" \"python\",\n",
" \"-m\",\n",
" \"vllm.entrypoints.api_server\",\n",
" \"--host=0.0.0.0\",\n",
" \"--port=8080\",\n",
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
" f\"--max-cpu-loras={max_cpu_loras}\",\n",
" f\"--max-num-seqs={max_num_seqs}\",\n",
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
" if enforce_eager:\n",
" vllm_args.append(\"--enforce-eager\")\n",
"\n",
" if enable_lora:\n",
" vllm_args.append(\"--enable-lora\")\n",
"\n",
" if enable_chunked_prefill:\n",
" vllm_args.append(\"--enable-chunked-prefill\")\n",
"\n",
" if enable_prefix_cache:\n",
" vllm_args.append(\"--enable-prefix-caching\")\n",
"\n",
" if 0 < host_prefix_kv_cache_utilization_target < 1:\n",
" vllm_args.append(\n",
" f\"--host-prefix-kv-cache-utilization-target={host_prefix_kv_cache_utilization_target}\"\n",
" )\n",
"\n",
" if model_type:\n",
" vllm_args.append(f\"--model-type={model_type}\")\n",
"\n",
" if enable_llama_tool_parser:\n",
" vllm_args.append(\"--enable-auto-tool-choice\")\n",
" vllm_args.append(\"--tool-call-parser=vertex-llama-3\")\n",
"\n",
" env_vars = {\n",
" \"MODEL_ID\": base_model_id,\n",
" \"DEPLOY_SOURCE\": \"notebook\",\n",
" }\n",
"\n",
" # HF_TOKEN is not a compulsory field and may not be defined.\n",
" try:\n",
" if HF_TOKEN:\n",
" env_vars[\"HF_TOKEN\"] = HF_TOKEN\n",
" except NameError:\n",
" pass\n",
"\n",
" model = aiplatform.Model.upload(\n",
" display_name=model_name,\n",
" serving_container_image_uri=VLLM_DOCKER_URI,\n",
" serving_container_args=vllm_args,\n",
" serving_container_ports=[8080],\n",
" serving_container_predict_route=\"/generate\",\n",
" serving_container_health_route=\"/ping\",\n",
" serving_container_environment_variables=env_vars,\n",
" serving_container_shared_memory_size_mb=(16 * 1024), # 16 GB\n",
" serving_container_deployment_timeout=7200,\n",
" model_garden_source_model_name=(\n",
" f\"publishers/{publisher}/models/{publisher_model_id}\"\n",
" ),\n",
" )\n",
" print(\n",
" f\"Deploying {model_name} on {machine_type} with {accelerator_count} {accelerator_type} GPU(s).\"\n",
" )\n",
" model.deploy(\n",
" endpoint=endpoint,\n",
" machine_type=machine_type,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
" deploy_request_timeout=1800,\n",
" system_labels={\n",
" \"NOTEBOOK_NAME\": \"model_garden_gemma3_deployment_on_vertex.ipynb\",\n",
" \"NOTEBOOK_ENVIRONMENT\": common_util.get_deploy_source(),\n",
" },\n",
" )\n",
" print(\"endpoint_name:\", endpoint.name)\n",
"\n",
" return model, endpoint\n",
"\n",
"\n",
"models[LABEL], endpoints[LABEL] = deploy_model_vllm(\n",
" model_name=common_util.get_job_name_with_datetime(prefix=\"gemma3-serve\"),\n",
" model_id=model_id,\n",
" publisher=\"google\",\n",
" publisher_model_id=\"gemma3\",\n",
" base_model_id=hf_model_id,\n",
" machine_type=machine_type,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
" gpu_memory_utilization=gpu_memory_utilization,\n",
" max_model_len=max_model_len,\n",
"endpoint = model.deploy(\n",
" accept_eula=True,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
")\n",
" serving_container_image_uri=\"us-docker.pkg.dev/vertex-ai/vertex-vision-model-garden-dockers/pytorch-vllm-serve:20250430_0916_RC00_maas\",\n",
" machine_type=\"a3-highgpu-1g\",\n",
" accelerator_type=\"NVIDIA_H100_80GB\",\n",
" accelerator_count=1,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "kqSUK2CwsImi"
},
"source": [
"To further customize your deployment, you can configure:\n",
"\n",
"model = models[LABEL]\n",
"endpoint = endpoints[LABEL]"
"- **Compute Resources**: Machine type, replica count (min/max), accelerator type and quantity.\n",
"- **Infrastructure**: Use Spot VMs, reservation affinity, or dedicated endpoints.\n",
"- **Serving Container**: Customize container image, ports, health checks, and environment variables.\n",
"\n",
"See the [Model Garden SDK README](https://github.com/googleapis/python-aiplatform/blob/main/vertexai/model_garden/README.md) for advanced configuration options."
]
},
{
@@ -950,16 +552,10 @@
"source": [
"# @title Delete the models and endpoints\n",
"\n",
"# @markdown Delete the experiment models and endpoints to recycle the resources\n",
"# @markdown and avoid unnecessary continuous charges that may incur.\n",
"# @markdown Delete the endpoint.\n",
"\n",
"# Undeploy model and delete endpoint.\n",
"for endpoint in endpoints.values():\n",
" endpoint.delete(force=True)\n",
"\n",
"# Delete models.\n",
"for model in models.values():\n",
" model.delete()"
"if endpoint:\n",
" endpoint.delete(force=True)"
]
}
],
@@ -146,7 +146,7 @@
"\n",
"# Import the necessary packages\n",
"! rm -rf vertex-ai-samples && git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"! cd vertex-ai-samples && git reset --hard c45f6a4f4d32e31a050f0e4ba52824b0caf4eda3\n",
"! cd vertex-ai-samples && git reset --hard 7ae13b346a72ee2a2dc8152dd40c6ddd72d6c810\n",
"\n",
"import datetime\n",
"import importlib\n",
@@ -164,7 +164,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"# Initialize models and endpoints as a dict\n",
@@ -816,7 +816,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -825,6 +824,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -59,34 +59,43 @@
"source": [
"## Overview\n",
"\n",
"This notebook demonstrates serving Gemma3n models with [SGLang](https://github.com/sgl-project/sglang). Gemma 3n models use selective parameter activation technology to reduce resource requirements. This technique allows the models to operate at an effective size of 2B and 4B parameters, which is lower than the total number of parameters they contain. For more information on Gemma 3n's efficient parameter management technology, see the [Gemma 3n](https://ai.google.dev/gemma/docs/gemma-3n#parameters) page.\n",
"This notebook demonstrates how to deploy a **Gemma 3N** open model on Google Cloud Vertex AI.\n",
"\n",
"### Objectives\n",
"\n",
"### Objective\n",
"- Deploy Gemma 3N using containerized backends like [vLLM](https://github.com/vllm-project/vllm) on GPU.\n",
"- Use the deployed model to serve chat completion requests for both text and multimodal inputs.\n",
"\n",
"- Deploy Gemma 3n with SGLang on GPU.\n",
"### File a Bug\n",
"\n",
"### File a bug\n",
"\n",
"File a bug on [GitHub](https://github.com/GoogleCloudPlatform/vertex-ai-samples/issues/new) if you encounter any issue with the notebook.\n",
"If you encounter issues with this notebook, report them on [GitHub](https://github.com/GoogleCloudPlatform/vertex-ai-samples/issues/new).\n",
"\n",
"### Costs\n",
"\n",
"This tutorial uses billable components of Google Cloud:\n",
"\n",
"* Vertex AI\n",
"* Cloud Storage\n",
"- Vertex AI\n",
"- Cloud Storage\n",
"\n",
"Learn about [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing), [Cloud Storage pricing](https://cloud.google.com/storage/pricing), and use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to generate a cost estimate based on your projected usage."
"Refer to the [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing) and [Cloud Storage pricing](https://cloud.google.com/storage/pricing) pages for more information. Use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to estimate your projected costs."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "264c07757582"
"id": "jeYw-Czg-DFy"
},
"source": [
"## Before you begin"
"## Get Started"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "KgyhGvEzBDkj"
},
"source": [
"### Install Vertex AI SDK and other required packages"
]
},
{
@@ -94,17 +103,22 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "ax7zWynUDcjk"
"id": "iCacdLqG-IsH"
},
"outputs": [],
"source": [
"# @title Request for quota\n",
"%pip install --upgrade --force-reinstall --quiet 'google-cloud-aiplatform>=1.106.0' 'openai' 'google-auth==2.27.0' 'requests==2.32.3'"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "HUKCrpBy-3yf"
},
"source": [
"### Authenticate the Notebook Environment (Colab only)\n",
"\n",
"# @markdown To deploy Gemma 3n models, you need 1 host of 1 x H100 machine. Check that you have sufficient quota:\n",
"# @markdown - For Spot VM quota, check [`CustomModelServingPreemptibleH100GPUsPerProjectPerRegion`](https://console.cloud.google.com/iam-admin/quotas?metric=aiplatform.googleapis.com%2Fcustom_model_serving_preemptible_nvidia_h100_gpus).\n",
"# @markdown - For regular VM quota, check [`CustomModelServingH100GPUsPerProjectPerRegion`](https://console.cloud.google.com/iam-admin/quotas?metric=aiplatform.googleapis.com%2Fcustom_model_serving_nvidia_h100_gpus).\n",
"#\n",
"# @markdown If you don't have sufficient quota, request for quota following the instructions at [\"Request a higher quota\"](https://cloud.google.com/docs/quota/view-manage#requesting_higher_quota)."
"If you're running this notebook in Google Colab, run the following cell to authenticate."
]
},
{
@@ -112,109 +126,146 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "YXFGIp1l-qtT"
"id": "JXwCT1kn-3Gu"
},
"outputs": [],
"source": [
"# @title Setup Google Cloud project\n",
"import sys\n",
"\n",
"# @markdown 1. [Make sure that billing is enabled for your project](https://cloud.google.com/billing/docs/how-to/modify-project).\n",
"if \"google.colab\" in sys.modules:\n",
" from google.colab import auth\n",
"\n",
"# @markdown 2. **[Optional]** Set region. If not set, the region will be set automatically according to Colab Enterprise environment.\n",
" auth.authenticate_user()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "AcW2nwB8-7yC"
},
"source": [
"### Set Google Cloud Project Information\n",
"\n",
"REGION = \"\" # @param {type:\"string\"}\n",
"To get started with Vertex AI, ensure you have an existing Google Cloud project and that the [Vertex AI API is enabled](https://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com).\n",
"\n",
"# @markdown 3. If you want to run predictions with A100 80GB or H100 GPUs, we recommend using the regions listed below. **NOTE:** Make sure you have associated quota in selected regions. Click the links to see your current quota for each GPU type: [Nvidia A100 80GB](https://console.cloud.google.com/iam-admin/quotas?metric=aiplatform.googleapis.com%2Fcustom_model_serving_nvidia_a100_80gb_gpus), [Nvidia H100 80GB](https://console.cloud.google.com/iam-admin/quotas?metric=aiplatform.googleapis.com%2Fcustom_model_serving_nvidia_h100_gpus). You can request for quota following the instructions at [\"Request a higher quota\"](https://cloud.google.com/docs/quota/view-manage#requesting_higher_quota).\n",
"\n",
"# @markdown > | Machine Type | Accelerator Type | Recommended Regions |\n",
"# @markdown | ----------- | ----------- | ----------- |\n",
"# @markdown | a2-ultragpu-1g | 1 NVIDIA_A100_80GB | us-central1, us-east4, europe-west4, asia-southeast1, us-east4 |\n",
"# @markdown | a3-highgpu-2g | 2 NVIDIA_H100_80GB | us-west1, asia-southeast1, europe-west4 |\n",
"# @markdown | a3-highgpu-4g | 4 NVIDIA_H100_80GB | us-west1, asia-southeast1, europe-west4 |\n",
"# @markdown | a3-highgpu-8g | 8 NVIDIA_H100_80GB | us-central1, europe-west4, us-west1, asia-southeast1 |\n",
"\n",
"# Upgrade Vertex AI SDK.\n",
"! pip3 install --upgrade --quiet 'google-cloud-aiplatform==1.103.0'\n",
"\n",
"# Import the necessary packages\n",
"import importlib\n",
"See the guide on [setting up your project and development environment](https://cloud.google.com/vertex-ai/docs/start/cloud-environment). Also confirm that [billing is enabled](https://cloud.google.com/billing/docs/how-to/modify-project).\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "eIVLp0oE--k-"
},
"outputs": [],
"source": [
"# Use the environment variable if the user doesn't provide Project ID.\n",
"import os\n",
"import time\n",
"from typing import Tuple\n",
"\n",
"import requests\n",
"from google import auth\n",
"from google.cloud import aiplatform\n",
"\n",
"# Upgrade Vertex AI SDK.\n",
"if os.environ.get(\"VERTEX_PRODUCT\") != \"COLAB_ENTERPRISE\":\n",
" ! pip install --upgrade tensorflow\n",
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
")\n",
"\n",
"\n",
"def check_quota(\n",
" project_id: str,\n",
" region: str,\n",
" resource_id: str,\n",
" accelerator_count: int,\n",
"):\n",
" \"\"\"Checks if the project and the region has the required quota.\"\"\"\n",
" quota = common_util.get_quota(project_id, region, resource_id)\n",
" quota_request_instruction = (\n",
" \"Either use \"\n",
" \"a different region or request additional quota. Follow \"\n",
" \"instructions here \"\n",
" \"https://cloud.google.com/docs/quotas/view-manage#requesting_higher_quota\"\n",
" \" to check quota in a region or request additional quota for \"\n",
" \"your project.\"\n",
" )\n",
" if quota == -1:\n",
" raise ValueError(\n",
" f\"Quota not found for: {resource_id} in {region}.\"\n",
" f\" {quota_request_instruction}\"\n",
" )\n",
" if quota < accelerator_count:\n",
" raise ValueError(\n",
" f\"Quota not enough for {resource_id} in {region}: {quota} <\"\n",
" f\" {accelerator_count}. {quota_request_instruction}\"\n",
" )\n",
"\n",
"\n",
"LABEL = \"sglang_gpu\"\n",
"models, endpoints = {}, {}\n",
"\n",
"# Get the default cloud project id.\n",
"PROJECT_ID = os.environ[\"GOOGLE_CLOUD_PROJECT\"]\n",
"\n",
"# Get the default region for launching jobs.\n",
"if not REGION:\n",
" REGION = os.environ[\"GOOGLE_CLOUD_REGION\"]\n",
"\n",
"# Initialize Vertex AI API.\n",
"print(\"Initializing Vertex AI API.\")\n",
"aiplatform.init(project=PROJECT_ID, location=REGION)\n",
"\n",
"! gcloud config set project $PROJECT_ID\n",
"\n",
"import vertexai\n",
"\n",
"vertexai.init(\n",
" project=PROJECT_ID,\n",
" location=REGION,\n",
"PROJECT_ID = \"[your-project-id]\" # @param {type: \"string\", placeholder: \"[your-project-id]\", isTemplate: true}\n",
"\n",
"if not PROJECT_ID or PROJECT_ID == \"[your-project-id]\":\n",
" PROJECT_ID = str(os.environ.get(\"GOOGLE_CLOUD_PROJECT\"))\n",
"\n",
"REGION = os.environ.get(\"GOOGLE_CLOUD_REGION\", \"us-central1\")\n",
"\n",
"vertexai.init(project=PROJECT_ID, location=REGION)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "Q0CXrvcZH_aw"
},
"source": [
"### Import libraries"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "3G2UXB82ICs6"
},
"outputs": [],
"source": [
"from vertexai import model_garden"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "upYRiGtP_-iN"
},
"source": [
"## Deploy model"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "H2WC_0hXDVXc"
},
"source": [
"### Choose model variant"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "u41zbNa2EoFq"
},
"source": [
"You can proceed with the default model variant or select a different one."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "-fgC4NLSDkF7"
},
"outputs": [],
"source": [
"model_version = \"gemma-3n-e4b-it\" # @param [\"gemma-3n-e2b\", \"gemma-3n-e2b-it\", \"gemma-3n-e4b\", \"gemma-3n-e4b-it\"] {isTemplate:true}\n",
"MODEL_NAME = f\"google/gemma3n@{model_version}\""
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "VRnUgU8LF3_i"
},
"source": [
"To see all deployable model variants available in Model Garden, use:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "-QLd-wshF6sB"
},
"outputs": [],
"source": [
"all_model_versions = model_garden.list_deployable_models(\n",
" model_filter=\"gemma3n\", list_hf_models=False\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "3zJJDmldn7rw"
"id": "N0UeFHa2GO63"
},
"source": [
"## Deploy Gemma 3n with SGLang"
"Once you've selected a model variant, initialize it:"
]
},
{
@@ -222,57 +273,22 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "_3Swj3pxn7rw"
"id": "GZiV3trBBcA3"
},
"outputs": [],
"source": [
"# @title Select the model variants\n",
"model = model_garden.OpenModel(MODEL_NAME)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "-0cL378wFlvf"
},
"source": [
"### Check the Deployment Configuration\n",
"\n",
"# @markdown Set the model to deploy.\n",
"\n",
"base_model_name = \"gemma-3n-E4B-it\" # @param [\"gemma-3n-E4B-it\", \"gemma-3n-E4B\", \"gemma-3n-E2B-it\", \"gemma-3n-E2B\"] {isTemplate:true}\n",
"model_id = \"gs://vertex-model-garden-restricted-us/gemma3n/\" + base_model_name\n",
"hf_model_id = \"google/\" + base_model_name\n",
"\n",
"# The pre-built serving docker images.\n",
"SGLANG_DOCKER_URI = \"us-docker.pkg.dev/deeplearning-platform-release/vertex-model-garden/sglang-serve.cu124.0-4.ubuntu2204.py310:20250626-1121-rc0\"\n",
"\n",
"# @markdown Choose whether to use a [Spot VM](https://cloud.google.com/compute/docs/instances/spot) for the deployment.\n",
"is_spot = False # @param {type:\"boolean\"}\n",
"\n",
"# @markdown Set use_dedicated_endpoint to False if you don't want to use [dedicated endpoint](https://cloud.google.com/vertex-ai/docs/general/deployment#create-dedicated-endpoint). Note that [dedicated endpoint does not support VPC Service Controls](https://cloud.google.com/vertex-ai/docs/predictions/choose-endpoint-type), uncheck the box if you are using VPC-SC.\n",
"use_dedicated_endpoint = True # @param {type:\"boolean\"}\n",
"\n",
"# @markdown Find Vertex AI prediction supported accelerators and regions at https://cloud.google.com/vertex-ai/docs/predictions/configure-compute.\n",
"accelerator_type = \"NVIDIA_H100_80GB\" # @param [\"NVIDIA_H100_80GB\", \"NVIDIA_A100_80GB\"] {isTemplate:true}\n",
"\n",
"PUBLISHER_MODEL_NAME = f\"publishers/google/models/gemma3n@{base_model_name.lower()}\"\n",
"\n",
"if accelerator_type == \"NVIDIA_H100_80GB\":\n",
" if is_spot:\n",
" resource_id = \"custom_model_serving_preemptible_nvidia_h100_gpus\"\n",
" else:\n",
" resource_id = \"custom_model_serving_nvidia_h100_gpus\"\n",
" machine_type = \"a3-highgpu-1g\"\n",
" accelerator_count = 1\n",
"elif accelerator_type == \"NVIDIA_A100_80GB\":\n",
" if is_spot:\n",
" resource_id = \"custom_model_serving_preemptible_nvidia_a100_80gb_gpus\"\n",
" else:\n",
" resource_id = \"custom_model_serving_nvidia_a100_80gb_gpus\"\n",
" machine_type = \"a2-ultragpu-1g\"\n",
" accelerator_count = 1\n",
"else:\n",
" raise ValueError(f\"Recommended GPU setting not found for: {base_model_name}.\")\n",
"\n",
"check_quota(\n",
" project_id=PROJECT_ID,\n",
" region=REGION,\n",
" resource_id=resource_id,\n",
" accelerator_count=accelerator_count,\n",
")\n",
"\n",
"# @markdown Click \"Show Code\" to see more details."
"Use the `list_deploy_options()` method to view the verified deployment configurations for your selected model. This helps ensure you have sufficient resources (e.g., GPU quota) available to deploy it."
]
},
{
@@ -280,29 +296,61 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "omW0LaC8wWz5"
"id": "zm73g7vFFm9N"
},
"outputs": [],
"source": [
"# @title [Option 1] Deploy with Model Garden SDK\n",
"# @markdown Deploy with Gen AI model-centric SDK. This section uploads the prebuilt model to Model Registry and deploys it to a Vertex AI Endpoint. It takes 15 minutes to 1 hour to finish depending on the size of the model. See [use open models with Vertex AI](https://cloud.google.com/vertex-ai/generative-ai/docs/open-models/use-open-models) for documentation on other use cases.\n",
"deploy_request_timeout = 1800 # 30 minutes\n",
"from vertexai import model_garden\n",
"deploy_options = model.list_deploy_options(concise=True)\n",
"print(deploy_options)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "WjV499VsGwrD"
},
"source": [
"### Deploy the Model\n",
"\n",
"model = model_garden.OpenModel(PUBLISHER_MODEL_NAME)\n",
"endpoints[LABEL] = model.deploy(\n",
" machine_type=machine_type,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
"Now that you’ve reviewed the deployment options, use the `deploy()` method to serve the selected open model to a Vertex AI endpoint. Deployment time may vary depending on the model size and infrastructure requirements.\n",
"\n",
"> **Note**: If the model requires accepting a license agreement (EULA), set the `accept_eula=True` flag in the deploy call. Set `use_dedicated_endpoint` to False if you don't want to use [dedicated endpoint](https://cloud.google.com/vertex-ai/docs/general/deployment#create-dedicated-endpoint)."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "wX1itVTvXdEP"
},
"outputs": [],
"source": [
"use_dedicated_endpoint = True"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "MRmPFEPoGzsB"
},
"outputs": [],
"source": [
"endpoint = model.deploy(\n",
" accept_eula=True,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
" spot=is_spot,\n",
" deploy_request_timeout=deploy_request_timeout,\n",
" accept_eula=True, # Accept the End User License Agreement (EULA) on the model card before deploy. Otherwise, the deployment will be forbidden.\n",
")\n",
"\n",
"endpoint = endpoints[LABEL]\n",
"\n",
"# @markdown Click \"Show Code\" to see more details."
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "PHBtn8DQp-ID"
},
"source": [
"Alternatively, you can select one of the verified deployment configurations listed above."
]
},
{
@@ -310,243 +358,33 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "3m-tDxgawYhU"
"id": "ADsJG8JYqI6c"
},
"outputs": [],
"source": [
"# @title [Option 2] Deploy with customized configs\n",
"\n",
"# @markdown This section uploads Gemma 3n models to Model Registry and deploys them to a Vertex Prediction Endpoint. It takes ~30 minutes to finish.\n",
"\n",
"# @markdown It's recommended to use the region selected by the deployment button on the model card. If the deployment button is not available, it's recommended to stay with the default region of the notebook.\n",
"\n",
"\n",
"def poll_operation(op_name: str) -> bool: # noqa: F811\n",
" creds, _ = auth.default()\n",
" auth_req = auth.transport.requests.Request()\n",
" creds.refresh(auth_req)\n",
" headers = {\n",
" \"Authorization\": f\"Bearer {creds.token}\",\n",
" }\n",
" get_resp = requests.get(\n",
" f\"https://{REGION}-aiplatform.googleapis.com/ui/{op_name}\",\n",
" headers=headers,\n",
" )\n",
" opjs = get_resp.json()\n",
" if \"error\" in opjs:\n",
" raise ValueError(f\"Operation failed: {opjs['error']}\")\n",
" return opjs.get(\"done\", False)\n",
"\n",
"\n",
"def poll_and_wait(op_name: str, total_wait: int, interval: int = 60): # noqa: F811\n",
" waited = 0\n",
" while not poll_operation(op_name):\n",
" if waited > total_wait:\n",
" raise TimeoutError(\"Operation timed out\")\n",
" print(\n",
" f\"\\rStill waiting for operation... Waited time in second: {waited:<6}\",\n",
" end=\"\",\n",
" flush=True,\n",
" )\n",
" waited += interval\n",
" time.sleep(interval)\n",
"\n",
"\n",
"def deploy_model_sglang_multihost(\n",
" model_name: str,\n",
" model_id: str,\n",
" publisher: str,\n",
" publisher_model_id: str,\n",
" service_account: str = \"\",\n",
" base_model_id: str = \"\",\n",
" machine_type: str = \"g2-standard-8\",\n",
" accelerator_type: str = \"NVIDIA_L4\",\n",
" accelerator_count: int = 1,\n",
" multihost_gpu_node_count: int = 1,\n",
" gpu_memory_utilization: float | None = None,\n",
" context_length: int | None = None,\n",
" dtype: str | None = None,\n",
" quantization: str | None = None,\n",
" enable_trust_remote_code: bool = False,\n",
" enable_torch_compile: bool = False,\n",
" torch_compile_max_bs: int | None = None,\n",
" attention_backend: str = \"\",\n",
" enable_flashinfer_mla: bool = False,\n",
" disable_cuda_graph: bool = False,\n",
" speculative_algorithm: str | None = None,\n",
" speculative_draft_model_path: str = \"\",\n",
" speculative_num_steps: int = 3,\n",
" speculative_eagle_topk: int = 1,\n",
" speculative_num_draft_tokens: int = 4,\n",
" enable_jit_deepgemm: bool = False,\n",
" enable_dp_attention: bool = False,\n",
" dp_size: int = 1,\n",
" enable_multimodal: bool = False,\n",
" use_dedicated_endpoint: bool = False,\n",
" max_num_seqs: int | None = None,\n",
" is_spot: bool = True,\n",
" tool_call_parser: str | None = None,\n",
") -> Tuple[aiplatform.Model, aiplatform.Endpoint]:\n",
" \"\"\"Deploys trained models with SGLang into Vertex AI.\"\"\"\n",
" endpoint = aiplatform.Endpoint.create(\n",
" display_name=f\"{model_name}-endpoint\",\n",
" dedicated_endpoint_enabled=use_dedicated_endpoint,\n",
" )\n",
"\n",
" if not base_model_id:\n",
" base_model_id = model_id\n",
"\n",
" # See https://docs.sglang.ai/backend/server_arguments.html for a list of possible arguments with descriptions.\n",
" sglang_args = [\n",
" f\"--model={model_id}\",\n",
" f\"--tp={accelerator_count * multihost_gpu_node_count}\",\n",
" f\"--dp={dp_size}\",\n",
" ]\n",
"\n",
" if context_length:\n",
" sglang_args.append(f\"--context-length={context_length}\")\n",
"\n",
" if gpu_memory_utilization:\n",
" sglang_args.append(f\"--mem-fraction-static={gpu_memory_utilization}\")\n",
"\n",
" if max_num_seqs:\n",
" sglang_args.append(f\"--max-running-requests={max_num_seqs}\")\n",
"\n",
" if dtype:\n",
" sglang_args.append(f\"--dtype={dtype}\")\n",
"\n",
" if quantization:\n",
" sglang_args.append(f\"--quantization={quantization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" sglang_args.append(\"--trust-remote-code\")\n",
"\n",
" if enable_torch_compile:\n",
" sglang_args.append(\"--enable-torch-compile\")\n",
" if torch_compile_max_bs:\n",
" sglang_args.append(f\"--torch-compile-max-bs={torch_compile_max_bs}\")\n",
"\n",
" if attention_backend:\n",
" sglang_args.append(f\"--attention-backend={attention_backend}\")\n",
"\n",
" if enable_flashinfer_mla:\n",
" sglang_args.append(\"--enable-flashinfer-mla\")\n",
"\n",
" if disable_cuda_graph:\n",
" sglang_args.append(\"--disable-cuda-graph\")\n",
"\n",
" if speculative_algorithm:\n",
" sglang_args.append(f\"--speculative-algorithm={speculative_algorithm}\")\n",
" sglang_args.append(\n",
" f\"--speculative-draft-model-path={speculative_draft_model_path}\"\n",
" )\n",
" sglang_args.append(f\"--speculative-num-steps={speculative_num_steps}\")\n",
" sglang_args.append(f\"--speculative-eagle-topk={speculative_eagle_topk}\")\n",
" sglang_args.append(\n",
" f\"--speculative-num-draft-tokens={speculative_num_draft_tokens}\"\n",
" )\n",
"\n",
" if enable_dp_attention:\n",
" sglang_args.append(\"--enable-dp-attention\")\n",
"\n",
" if enable_multimodal:\n",
" sglang_args.append(\"--enable-multimodal\")\n",
"\n",
" if tool_call_parser:\n",
" sglang_args.append(f\"--tool-call-parser={tool_call_parser}\")\n",
"\n",
" env_vars = {\n",
" \"MODEL_ID\": base_model_id,\n",
" \"DEPLOY_SOURCE\": \"notebook\",\n",
" }\n",
"\n",
" if enable_jit_deepgemm:\n",
" env_vars[\"SGL_ENABLE_JIT_DEEPGEMM\"] = \"1\"\n",
"\n",
" # HF_TOKEN is not a compulsory field and may not be defined.\n",
" try:\n",
" if HF_TOKEN:\n",
" env_vars[\"HF_TOKEN\"] = HF_TOKEN\n",
" except NameError:\n",
" pass\n",
"\n",
" model = aiplatform.Model.upload(\n",
" display_name=model_name,\n",
" serving_container_image_uri=SGLANG_DOCKER_URI,\n",
" serving_container_args=sglang_args,\n",
" serving_container_ports=[30000],\n",
" serving_container_predict_route=\"/vertex_generate\",\n",
" serving_container_health_route=\"/health\",\n",
" serving_container_environment_variables=env_vars,\n",
" serving_container_shared_memory_size_mb=(16 * 1024), # 16 GB\n",
" serving_container_deployment_timeout=7200,\n",
" model_garden_source_model_name=(\n",
" f\"publishers/{publisher}/models/{publisher_model_id}\"\n",
" ),\n",
" )\n",
" print(\n",
" f\"Deploying {model_name} on {machine_type} with {int(accelerator_count * multihost_gpu_node_count)} {accelerator_type} GPU(s).\"\n",
" )\n",
"\n",
" creds, _ = auth.default()\n",
" auth_req = auth.transport.requests.Request()\n",
" creds.refresh(auth_req)\n",
"\n",
" url = f\"https://{REGION}-aiplatform.googleapis.com/ui/projects/{PROJECT_ID}/locations/{REGION}/endpoints/{endpoint.name}:deployModel\"\n",
" headers = {\n",
" \"Content-Type\": \"application/json\",\n",
" \"Authorization\": f\"Bearer {creds.token}\",\n",
" }\n",
" data = {\n",
" \"deployedModel\": {\n",
" \"model\": model.resource_name,\n",
" \"displayName\": model_name,\n",
" \"dedicatedResources\": {\n",
" \"machineSpec\": {\n",
" \"machineType\": machine_type,\n",
" \"multihostGpuNodeCount\": multihost_gpu_node_count,\n",
" \"acceleratorType\": accelerator_type,\n",
" \"acceleratorCount\": accelerator_count,\n",
" },\n",
" \"minReplicaCount\": 1,\n",
" \"maxReplicaCount\": 1,\n",
" },\n",
" \"system_labels\": {\n",
" \"NOTEBOOK_NAME\": \"model_garden_gemma3n_deployment_on_vertex.ipynb\",\n",
" \"NOTEBOOK_ENVIRONMENT\": common_util.get_deploy_source(),\n",
" },\n",
" },\n",
" }\n",
" if service_account:\n",
" data[\"deployedModel\"][\"serviceAccount\"] = service_account\n",
" if is_spot:\n",
" data[\"deployedModel\"][\"dedicatedResources\"][\"spot\"] = True\n",
" response = requests.post(url, headers=headers, json=data)\n",
" print(f\"Deploy Model response: {response.json()}\")\n",
" if response.status_code != 200 or \"name\" not in response.json():\n",
" raise ValueError(f\"Failed to deploy model: {response.text}\")\n",
" poll_and_wait(response.json()[\"name\"], 7200)\n",
" print(\"endpoint_name:\", endpoint.name)\n",
"\n",
" return model, endpoint\n",
"\n",
"\n",
"models[LABEL], endpoints[LABEL] = deploy_model_sglang_multihost(\n",
" model_name=common_util.get_job_name_with_datetime(prefix=\"gemma3n-serve\"),\n",
" model_id=model_id,\n",
" publisher=\"google\",\n",
" publisher_model_id=\"gemma3n\",\n",
" base_model_id=hf_model_id,\n",
" machine_type=machine_type,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
" attention_backend=\"fa3\",\n",
" enable_multimodal=True,\n",
"endpoint = model.deploy(\n",
" accept_eula=True,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
" is_spot=is_spot,\n",
")\n",
" serving_container_image_uri=\"us-docker.pkg.dev/deeplearning-platform-release/vertex-model-garden/sglang-serve.cu124.0-4.ubuntu2204.py310:model-garden.sglang-0-4-release_20250817.00_p0\",\n",
" machine_type=\"a3-highgpu-1g\",\n",
" accelerator_type=\"NVIDIA_H100_80GB\",\n",
" accelerator_count=1,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "kqSUK2CwsImi"
},
"source": [
"To further customize your deployment, you can configure:\n",
"\n",
"# @markdown Click \"Show Code\" to see more details."
"- **Compute Resources**: Machine type, replica count (min/max), accelerator type and quantity.\n",
"- **Infrastructure**: Use Spot VMs, reservation affinity, or dedicated endpoints.\n",
"- **Serving Container**: Customize container image, ports, health checks, and environment variables.\n",
"\n",
"See the [Model Garden SDK README](https://github.com/googleapis/python-aiplatform/blob/main/vertexai/model_garden/README.md) for advanced configuration options."
]
},
{
@@ -601,7 +439,7 @@
" \"top_k\": top_k,\n",
" }\n",
"}\n",
"response = endpoints[\"sglang_gpu\"].predict(\n",
"response = endpoint.predict(\n",
" instances=instances,\n",
" parameters=parameters,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
@@ -625,8 +463,8 @@
"# @title Chat completion with text-only requests\n",
"\n",
"if use_dedicated_endpoint:\n",
" DEDICATED_ENDPOINT_DNS = endpoints[\"sglang_gpu\"].gca_resource.dedicated_endpoint_dns\n",
"ENDPOINT_RESOURCE_NAME = endpoints[\"sglang_gpu\"].resource_name\n",
" DEDICATED_ENDPOINT_DNS = endpoint.gca_resource.dedicated_endpoint_dns\n",
"ENDPOINT_RESOURCE_NAME = endpoint.resource_name\n",
"\n",
"# @title Chat Completions Inference\n",
"\n",
@@ -700,8 +538,8 @@
"# @title Chat completion with text+image requests\n",
"\n",
"if use_dedicated_endpoint:\n",
" DEDICATED_ENDPOINT_DNS = endpoints[\"sglang_gpu\"].gca_resource.dedicated_endpoint_dns\n",
"ENDPOINT_RESOURCE_NAME = endpoints[\"sglang_gpu\"].resource_name\n",
" DEDICATED_ENDPOINT_DNS = endpoint.gca_resource.dedicated_endpoint_dns\n",
"ENDPOINT_RESOURCE_NAME = endpoint.resource_name\n",
"\n",
"# @title Chat Completions Inference\n",
"\n",
@@ -770,8 +608,8 @@
"# @title Chat completion with text+audio requests\n",
"\n",
"if use_dedicated_endpoint:\n",
" DEDICATED_ENDPOINT_DNS = endpoints[\"sglang_gpu\"].gca_resource.dedicated_endpoint_dns\n",
"ENDPOINT_RESOURCE_NAME = endpoints[\"sglang_gpu\"].resource_name\n",
" DEDICATED_ENDPOINT_DNS = endpoint.gca_resource.dedicated_endpoint_dns\n",
"ENDPOINT_RESOURCE_NAME = endpoint.resource_name\n",
"\n",
"# @title Chat Completions Inference\n",
"\n",
@@ -848,16 +686,10 @@
"source": [
"# @title Delete the models and endpoints\n",
"\n",
"# @markdown Delete the experiment models and endpoints to recycle the resources\n",
"# @markdown and avoid unnecessary continuous charges that may incur.\n",
"# @markdown Delete the endpoint.\n",
"\n",
"# Undeploy model and delete endpoint.\n",
"for endpoint in endpoints.values():\n",
" endpoint.delete(force=True)\n",
"\n",
"# Delete models.\n",
"for model in models.values():\n",
" model.delete()"
"if endpoint:\n",
" endpoint.delete(force=True)"
]
}
],
@@ -121,7 +121,7 @@
"from google.cloud import aiplatform\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -727,7 +727,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -736,6 +735,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -123,7 +123,7 @@
"from google.cloud import aiplatform\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -142,7 +142,7 @@
"\n",
"# Import the necessary packages\n",
"! rm -rf vertex-ai-samples && git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"! cd vertex-ai-samples && git reset --hard 0727e19520cf7957bceb701c248221bd3dbe4f1f\n",
"! cd vertex-ai-samples && git reset --hard 7ae13b346a72ee2a2dc8152dd40c6ddd72d6c810\n",
"\n",
"import datetime\n",
"import importlib\n",
@@ -155,7 +155,7 @@
" custom_job as gca_custom_job_compat\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -720,7 +720,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -729,6 +728,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -229,7 +229,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -238,6 +237,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -147,7 +147,7 @@
"from google.cloud import aiplatform\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -54,41 +54,48 @@
{
"cell_type": "markdown",
"metadata": {
"id": "iOmVD9tZXucQ"
"id": "3de7470326a2"
},
"source": [
"## Overview\n",
"\n",
"This notebook provides a practical introduction to using the PaLiGemma 2 model, a powerful vision-language model developed by Google. We'll demonstrate how to leverage its multimodal capabilities to perform tasks like vision question answering. Consult the [model card](https://console.cloud.google.com/vertex-ai/publishers/google/model-garden/paligemma) for more information.\n",
"This notebook demonstrates how to deploy a **Paligemma** open model on Google Cloud Vertex AI.\n",
"\n",
"### Objectives\n",
"\n",
"### Objective\n",
"- Deploy Paligemma using containerized backends like [vLLM](https://github.com/vllm-project/vllm) on GPU.\n",
"- Use the deployed model to serve chat completion requests for both text and multimodal inputs.\n",
"\n",
"- Deploy PaliGemma 2 to a Vertex AI Endpoint.\n",
"- Make predictions to the endpoint including:\n",
" - Answering questions about a given image.\n",
"### File a Bug\n",
"\n",
"### File a bug\n",
"\n",
"File a bug on [GitHub](https://github.com/GoogleCloudPlatform/vertex-ai-samples/issues/new) if you encounter any issue with the notebook.\n",
"If you encounter issues with this notebook, report them on [GitHub](https://github.com/GoogleCloudPlatform/vertex-ai-samples/issues/new).\n",
"\n",
"### Costs\n",
"\n",
"This tutorial uses billable components of Google Cloud:\n",
"\n",
"* Vertex AI\n",
"* Cloud Storage\n",
"- Vertex AI\n",
"- Cloud Storage\n",
"\n",
"Learn about [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing), [Cloud Storage pricing](https://cloud.google.com/storage/pricing), and use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to generate a cost estimate based on your projected usage."
"Refer to the [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing) and [Cloud Storage pricing](https://cloud.google.com/storage/pricing) pages for more information. Use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to estimate your projected costs."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "2aFHbs1g6Wc-"
"id": "jeYw-Czg-DFy"
},
"source": [
"## Before you begin"
"## Get Started"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "KgyhGvEzBDkj"
},
"source": [
"### Install Vertex AI SDK and other required packages"
]
},
{
@@ -96,76 +103,169 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "QvQjsmIJ6Y3f"
"id": "iCacdLqG-IsH"
},
"outputs": [],
"source": [
"# @title Setup Google Cloud project\n",
"%pip install --upgrade --force-reinstall --quiet 'google-cloud-aiplatform>=1.106.0' 'openai' 'google-auth==2.27.0' 'requests==2.32.3'"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "HUKCrpBy-3yf"
},
"source": [
"### Authenticate the Notebook Environment (Colab only)\n",
"\n",
"# @markdown 1. [Make sure that billing is enabled for your project](https://cloud.google.com/billing/docs/how-to/modify-project).\n",
"If you're running this notebook in Google Colab, run the following cell to authenticate."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "JXwCT1kn-3Gu"
},
"outputs": [],
"source": [
"import sys\n",
"\n",
"# @markdown 2. **[Optional]** Set region. If not set, the region will be set automatically according to Colab Enterprise environment.\n",
"if \"google.colab\" in sys.modules:\n",
" from google.colab import auth\n",
"\n",
"REGION = \"\" # @param {type:\"string\"}\n",
" auth.authenticate_user()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "AcW2nwB8-7yC"
},
"source": [
"### Set Google Cloud Project Information\n",
"\n",
"# @markdown 3. If you want to run predictions with A100 80GB or H100 GPUs, we recommend using the regions listed below. **NOTE:** Make sure you have associated quota in selected regions. Click the links to see your current quota for each GPU type: [Nvidia A100 80GB](https://console.cloud.google.com/iam-admin/quotas?metric=aiplatform.googleapis.com%2Fcustom_model_serving_nvidia_a100_80gb_gpus), [Nvidia H100 80GB](https://console.cloud.google.com/iam-admin/quotas?metric=aiplatform.googleapis.com%2Fcustom_model_serving_nvidia_h100_gpus). You can request for quota following the instructions at [\"Request a higher quota\"](https://cloud.google.com/docs/quota/view-manage#requesting_higher_quota).\n",
"To get started with Vertex AI, ensure you have an existing Google Cloud project and that the [Vertex AI API is enabled](https://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com).\n",
"\n",
"# @markdown > | Machine Type | Accelerator Type | Recommended Regions |\n",
"# @markdown | ----------- | ----------- | ----------- |\n",
"# @markdown | a2-ultragpu-1g | 1 NVIDIA_A100_80GB | us-central1, us-east4, europe-west4, asia-southeast1, us-east4 |\n",
"# @markdown | a3-highgpu-2g | 2 NVIDIA_H100_80GB | us-west1, asia-southeast1, europe-west4 |\n",
"# @markdown | a3-highgpu-4g | 4 NVIDIA_H100_80GB | us-west1, asia-southeast1, europe-west4 |\n",
"# @markdown | a3-highgpu-8g | 8 NVIDIA_H100_80GB | us-central1, europe-west4, us-west1, asia-southeast1 |\n",
"\n",
"# Upgrade Vertex AI SDK.\n",
"! pip3 install --upgrade --quiet 'google-cloud-aiplatform==1.103.0'\n",
"\n",
"import importlib\n",
"# Import the necessary packages\n",
"See the guide on [setting up your project and development environment](https://cloud.google.com/vertex-ai/docs/start/cloud-environment). Also confirm that [billing is enabled](https://cloud.google.com/billing/docs/how-to/modify-project).\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "eIVLp0oE--k-"
},
"outputs": [],
"source": [
"# Use the environment variable if the user doesn't provide Project ID.\n",
"import os\n",
"from typing import Any, Dict, Tuple\n",
"\n",
"from google.cloud import aiplatform\n",
"\n",
"if os.environ.get(\"VERTEX_PRODUCT\") != \"COLAB_ENTERPRISE\":\n",
" ! pip install --upgrade tensorflow\n",
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
"LABEL = \"paligemma2\"\n",
"\n",
"\n",
"# Get the default cloud project id.\n",
"PROJECT_ID = os.environ[\"GOOGLE_CLOUD_PROJECT\"]\n",
"\n",
"# Get the default region for launching jobs.\n",
"if not REGION:\n",
" REGION = os.environ[\"GOOGLE_CLOUD_REGION\"]\n",
"\n",
"# Initialize Vertex AI API.\n",
"print(\"Initializing Vertex AI API.\")\n",
"aiplatform.init(project=PROJECT_ID, location=REGION)\n",
"\n",
"! gcloud config set project $PROJECT_ID\n",
"import vertexai\n",
"\n",
"vertexai.init(\n",
" project=PROJECT_ID,\n",
" location=REGION,\n",
"PROJECT_ID = \"[your-project-id]\" # @param {type: \"string\", placeholder: \"[your-project-id]\", isTemplate: true}\n",
"\n",
"if not PROJECT_ID or PROJECT_ID == \"[your-project-id]\":\n",
" PROJECT_ID = str(os.environ.get(\"GOOGLE_CLOUD_PROJECT\"))\n",
"\n",
"REGION = os.environ.get(\"GOOGLE_CLOUD_REGION\", \"us-central1\")\n",
"\n",
"vertexai.init(project=PROJECT_ID, location=REGION)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "Q0CXrvcZH_aw"
},
"source": [
"### Import libraries"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "3G2UXB82ICs6"
},
"outputs": [],
"source": [
"from vertexai import model_garden"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "upYRiGtP_-iN"
},
"source": [
"## Deploy model"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "H2WC_0hXDVXc"
},
"source": [
"### Choose model variant"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "u41zbNa2EoFq"
},
"source": [
"You can proceed with the default model variant or select a different one."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "-fgC4NLSDkF7"
},
"outputs": [],
"source": [
"model_version = \"paligemma2-3b-mix-448\" # @param [\"paligemma-224-bfloat16\", \"paligemma-224-float16\", \"paligemma-224-float32\", \"paligemma-448-bfloat16\", \"paligemma-448-float16\", \"paligemma-448-float32\", \"paligemma-896-bfloat16\", \"paligemma-896-float16\", \"paligemma-896-float32\", \"paligemma-mix-224-bfloat16\", \"paligemma-mix-224-float16\", \"paligemma-mix-224-float32\", \"paligemma-mix-448-bfloat16\", \"paligemma-mix-448-float16\", \"paligemma-mix-448-float32\", \"paligemma2-10b-ft-docci-448\", \"paligemma2-10b-mix-224\", \"paligemma2-10b-mix-448\", \"paligemma2-10b-pt-448\", \"paligemma2-10b-pt-896\", \"paligemma2-28b-mix-224\", \"paligemma2-28b-mix-448\", \"paligemma2-28b-pt-224\", \"paligemma2-28b-pt-448\", \"paligemma2-28b-pt-896\", \"paligemma2-3b-ft-docci-448\", \"paligemma2-3b-mix-224\", \"paligemma2-3b-mix-448\", \"paligemma2-3b-pt-224\", \"paligemma2-3b-pt-448\", \"paligemma2-3b-pt-896\"] {isTemplate:true}\n",
"MODEL_NAME = f\"google/paligemma@{model_version}\""
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "VRnUgU8LF3_i"
},
"source": [
"To see all deployable model variants available in Model Garden, use:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "-QLd-wshF6sB"
},
"outputs": [],
"source": [
"all_model_versions = model_garden.list_deployable_models(\n",
" model_filter=\"paligemma\", list_hf_models=False\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "kyMJXkfviWgl"
"id": "N0UeFHa2GO63"
},
"source": [
"## Deploy Model to a Vertex AI Endpoint"
"Once you've selected a model variant, initialize it:"
]
},
{
@@ -173,47 +273,22 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "toY-WPKDFesF"
"id": "GZiV3trBBcA3"
},
"outputs": [],
"source": [
"# @title Deploy\n",
"model = model_garden.OpenModel(MODEL_NAME)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "-0cL378wFlvf"
},
"source": [
"### Check the Deployment Configuration\n",
"\n",
"MODEL_NAME = \"paligemma2-3b-pt-224\" # @param [\"paligemma2-3b-pt-224\", \"paligemma2-3b-mix-224\", \"paligemma2-3b-ft-docci-448\", \"paligemma2-3b-mix-448\", \"paligemma2-3b-pt-448\", \"paligemma2-3b-pt-896\", \"paligemma2-10b-mix-224\", \"paligemma2-10b-pt-224\", \"paligemma2-10b-ft-docci-448\", \"paligemma2-10b-mix-448\", \"paligemma2-10b-pt-448\", \"paligemma2-10b-pt-896\", \"paligemma2-28b-mix-224\", \"paligemma2-28b-pt-224\", \"paligemma2-28b-mix-448\", \"paligemma2-28b-pt-448\", \"paligemma2-28b-pt-896\"]\n",
"GCS_PREFIX = \"gs://vertex-model-garden-restricted-us/paligemma2\"\n",
"\n",
"MODEL_ID = os.path.join(GCS_PREFIX, MODEL_NAME)\n",
"\n",
"PUBLISHER_MODEL_NAME = f\"publishers/google/models/paligemma@{MODEL_NAME}\"\n",
"\n",
"\n",
"# @markdown If you want to use other accelerator types not listed above, then check other Vertex AI prediction supported accelerators and regions at https://cloud.google.com/vertex-ai/docs/predictions/configure-compute. You may need to manually set the `machine_type`, `accelerator_type`, and `accelerator_count` in the code by clicking `Show code` first.\n",
"\n",
"if \"3b\" in MODEL_NAME:\n",
" accelerator_type = \"NVIDIA_L4\"\n",
" machine_type = \"g2-standard-16\"\n",
" accelerator_count = 1\n",
"elif \"10b\" in MODEL_NAME:\n",
" accelerator_type = \"NVIDIA_TESLA_A100\"\n",
" machine_type = \"a2-highgpu-1g\"\n",
" accelerator_count = 1\n",
"elif \"28b\" in MODEL_NAME:\n",
" accelerator_type = \"NVIDIA_H100_80GB\"\n",
" machine_type = \"a3-highgpu-8g\"\n",
" accelerator_count = 8\n",
"else:\n",
" raise ValueError(f\"Recommended GPU setting not found for: {MODEL_NAME}.\")\n",
"\n",
"common_util.check_quota(\n",
" project_id=PROJECT_ID,\n",
" region=REGION,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
" is_for_training=False,\n",
")\n",
"\n",
"# @markdown Set use_dedicated_endpoint to False if you don't want to use [dedicated endpoint](https://cloud.google.com/vertex-ai/docs/general/deployment#create-dedicated-endpoint). Note that [dedicated endpoint does not support VPC Service Controls](https://cloud.google.com/vertex-ai/docs/predictions/choose-endpoint-type), uncheck the box if you are using VPC-SC.\n",
"use_dedicated_endpoint = True # @param {type:\"boolean\"}"
"Use the `list_deploy_options()` method to view the verified deployment configurations for your selected model. This helps ensure you have sufficient resources (e.g., GPU quota) available to deploy it."
]
},
{
@@ -221,140 +296,95 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "pe_qbTCA6nKf"
"id": "zm73g7vFFm9N"
},
"outputs": [],
"source": [
"# @title [Option 1] Deploy with Model Garden SDK\n",
"deploy_options = model.list_deploy_options(concise=True)\n",
"print(deploy_options)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "WjV499VsGwrD"
},
"source": [
"### Deploy the Model\n",
"\n",
"# @markdown Deploy with Gen AI model-centric SDK. This section uploads the prebuilt model to Model Registry and deploys it to a Vertex AI Endpoint. It takes 15 minutes to 1 hour to finish depending on the size of the model. See [use open models with Vertex AI](https://cloud.google.com/vertex-ai/generative-ai/docs/open-models/use-open-models) for documentation on other use cases.\n",
"from vertexai import model_garden\n",
"Now that you’ve reviewed the deployment options, use the `deploy()` method to serve the selected open model to a Vertex AI endpoint. Deployment time may vary depending on the model size and infrastructure requirements.\n",
"\n",
"model = model_garden.OpenModel(PUBLISHER_MODEL_NAME)\n",
"endpoints[LABEL] = model.deploy(\n",
" machine_type=machine_type,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
"> **Note**: If the model requires accepting a license agreement (EULA), set the `accept_eula=True` flag in the deploy call. Set `use_dedicated_endpoint` to False if you don't want to use [dedicated endpoint](https://cloud.google.com/vertex-ai/docs/general/deployment#create-dedicated-endpoint)."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "wX1itVTvXdEP"
},
"outputs": [],
"source": [
"use_dedicated_endpoint = True"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "MRmPFEPoGzsB"
},
"outputs": [],
"source": [
"endpoint = model.deploy(\n",
" accept_eula=True,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
" accept_eula=True, # Accept the End User License Agreement (EULA) on the model card before deploy. Otherwise, the deployment will be forbidden.\n",
")\n",
"\n",
"endpoint = endpoints[LABEL]"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "jbeLl-9C6nKf"
},
"outputs": [],
"source": [
"# @title [Option 2] Deploy with customized configs\n",
"\n",
"# @markdown This section uploads the prebuilt PaliGemma 2 models to Model Registry and deploys it to a Vertex AI Endpoint. It takes approximately 15 minutes to finish.\n",
"\n",
"# @markdown Select the desired resolution and precision of prebuilt model to deploy, leaving the optional `custom_paligemma_model_uri` as is. Higher resolution and precision_type can result in better inference results, but may require additional GPU.\n",
"\n",
"TASK = \"paligemma_VQA\"\n",
"\n",
"# The pre-built serving docker images.\n",
"SERVE_DOCKER_URI = \"us-docker.pkg.dev/vertex-ai/vertex-vision-model-garden-dockers/pytorch-one-serve:20250205_0822_RC00\"\n",
"\n",
"\n",
"def deploy_model(\n",
" model_name: str = None,\n",
" model_id: str = None,\n",
" task: str = None,\n",
" machine_type: str = \"g2-standard-8\",\n",
" accelerator_type: str = \"NVIDIA_L4\",\n",
" accelerator_count: int = 1,\n",
" serving_port: int = 7080,\n",
" serving_route: str = \"/predict\",\n",
" serving_docker_uri: str = SERVE_DOCKER_URI,\n",
") -> Tuple[aiplatform.Endpoint, aiplatform.Model]:\n",
" \"\"\"Deploys a model to a real-time prediction endpoint.\n",
"\n",
" Args:\n",
" model_name: The base name of the model.\n",
" model_id: The model ID.\n",
" task: The task to perform.\n",
" machine_type: The machine type.\n",
" accelerator_type: The accelerator type.\n",
" accelerator_count: The accelerator count.\n",
" serving_port: The serving port.\n",
" serving_route: The serving route.\n",
" hf_token: HuggingFace token for model access.\n",
"\n",
" Returns:\n",
" A tuple containing the created endpoint and deployed model objects.\n",
" \"\"\"\n",
"\n",
" endpoint = aiplatform.Endpoint.create(\n",
" display_name=common_util.get_job_name_with_datetime(prefix=model_name)\n",
" )\n",
" serving_env = {\n",
" \"MODEL_ID\": model_id,\n",
" \"DEPLOY_SOURCE\": \"notebook\",\n",
" \"TASK\": task,\n",
" }\n",
" model = aiplatform.Model.upload(\n",
" display_name=task,\n",
" serving_container_image_uri=serving_docker_uri,\n",
" serving_container_ports=[serving_port],\n",
" serving_container_predict_route=serving_route,\n",
" serving_container_health_route=\"/ping\",\n",
" serving_container_environment_variables=serving_env,\n",
" model_garden_source_model_name=\"publishers/google/models/paligemma\",\n",
" )\n",
" model.deploy(\n",
" endpoint=endpoint,\n",
" machine_type=machine_type,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
" sync=False,\n",
" system_labels={\n",
" \"NOTEBOOK_NAME\": \"model_garden_hf_paligemma2_deployment.ipynb\",\n",
" \"NOTEBOOK_ENVIRONMENT\": common_util.get_deploy_source(),\n",
" },\n",
" )\n",
" return endpoint, model\n",
"\n",
"\n",
"endpoints[\"paligemma2\"], models[\"paligemma2\"] = deploy_model(\n",
" model_name=MODEL_NAME,\n",
" model_id=MODEL_ID,\n",
" task=TASK,\n",
" machine_type=machine_type,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
" serving_port=7080,\n",
" serving_route=\"/predict\",\n",
" serving_docker_uri=SERVE_DOCKER_URI,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "PHBtn8DQp-ID"
},
"source": [
"Alternatively, you can select one of the verified deployment configurations listed above."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "tOtYOhZa3lsx"
"id": "ADsJG8JYqI6c"
},
"outputs": [],
"source": [
"# @title [Optional] Loading an existing Endpoint\n",
"# @markdown If you've already deployed an Endpoint, you can load it by filling in the Endpoint's ID below.\n",
"# @markdown You can view deployed Endpoints at [Vertex Online Prediction](https://console.cloud.google.com/vertex-ai/online-prediction/endpoints).\n",
"endpoint_id = \"\" # @param {type: \"string\"}\n",
"endpoint = model.deploy(\n",
" accept_eula=True,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
" serving_container_image_uri=\"us-docker.pkg.dev/vertex-ai/vertex-vision-model-garden-dockers/pytorch-one-serve:20250205_0822_RC00\",\n",
" machine_type=\"g2-standard-16\",\n",
" accelerator_type=\"NVIDIA_L4\",\n",
" accelerator_count=1,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "kqSUK2CwsImi"
},
"source": [
"To further customize your deployment, you can configure:\n",
"\n",
"if endpoint_id:\n",
" endpoint = aiplatform.Endpoint(\n",
" endpoint_name=endpoint_id,\n",
" project=PROJECT_ID,\n",
" location=REGION,\n",
" )"
"- **Compute Resources**: Machine type, replica count (min/max), accelerator type and quantity.\n",
"- **Infrastructure**: Use Spot VMs, reservation affinity, or dedicated endpoints.\n",
"- **Serving Container**: Customize container image, ports, health checks, and environment variables.\n",
"\n",
"See the [Model Garden SDK README](https://github.com/googleapis/python-aiplatform/blob/main/vertexai/model_garden/README.md) for advanced configuration options."
]
},
{
@@ -363,7 +393,7 @@
"id": "2Idtx2ETNQtn"
},
"source": [
"### Predict\n",
"## Predict\n",
"\n",
"The following sections will use images from [pexels.com](https://www.pexels.com/) for demoing purposes. All the images have the following license: https://www.pexels.com/license/.\n",
"\n",
@@ -383,6 +413,19 @@
"\n",
"# @markdown This section uses the deployed PaliGemma model to answer questions about a given image.\n",
"\n",
"import importlib\n",
"from typing import Any, Dict\n",
"\n",
"from google.cloud import aiplatform\n",
"\n",
"# Import the necessary packages.\n",
"! rm -rf vertex-ai-samples && git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"! cd vertex-ai-samples\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"# @markdown ![](https://images.pexels.com/photos/1006293/pexels-photo-1006293.jpeg?auto=compress&cs=tinysrgb&w=630&h=375&dpr=2)\n",
"image_url = \"https://images.pexels.com/photos/1006293/pexels-photo-1006293.jpeg\" # @param {type:\"string\"}\n",
"\n",
@@ -436,12 +479,16 @@
"# Using max_new_tokens along with other parameters\n",
"parameters_with_tokens = {\"max_new_tokens\": 50}\n",
"predictions_with_tokens = vqa_predict(\n",
" endpoint=endpoints[\"paligemma2\"],\n",
" endpoint=endpoint,\n",
" image_url=image_url,\n",
" text_prompt=question_prompt,\n",
" parameters=parameters_with_tokens,\n",
")\n",
"\n",
"\n",
"image = common_util.download_image(image_url)\n",
"display(image)\n",
"\n",
"print(f\"Prediction Response: {predictions_with_tokens}\")\n",
"# @markdown Click \"Show Code\" to see more details."
]
@@ -464,20 +511,10 @@
},
"outputs": [],
"source": [
"# @markdown Delete the experiment models and endpoints to recycle the resources\n",
"# @markdown and avoid unnecessary continuous charges that may incur.\n",
"# @markdown Delete the endpoint.\n",
"\n",
"# Undeploy model and delete endpoint.\n",
"for endpoint in endpoints.values():\n",
" endpoint.delete(force=True)\n",
"\n",
"# Delete models.\n",
"for model in models.values():\n",
" model.delete()\n",
"\n",
"delete_bucket = False # @param {type:\"boolean\"}\n",
"if delete_bucket:\n",
" ! gsutil -m rm -r $BUCKET_NAME"
"if endpoint:\n",
" endpoint.delete(force=True)"
]
}
],
@@ -120,7 +120,7 @@
"from google.cloud import aiplatform\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"\n",
@@ -163,7 +163,7 @@
"TASK = \"text-classification\" # @param {type: \"string\", isTemplate: true}\n",
"\n",
"# The pre-built serving docker images for Hugging Face Pytorch Inference.\n",
"SERVE_DOCKER_URI = \"us-docker.pkg.dev/deeplearning-platform-release/vertex-model-garden/hf-inference-toolkit.cu125.0-1.ubuntu2204.py311\"\n",
"SERVE_DOCKER_URI = \"us-docker.pkg.dev/deeplearning-platform-release/vertex-model-garden/hf-inference-toolkit.cu125.0-1.ubuntu2204.py311:model-garden.hf-inference-toolkit-0-1-release_20250828.01_p0\"\n",
"\n",
"machine_type = \"g2-standard-8\" # @param {type: \"string\", isTemplate: true}\n",
"accelerator_type = \"NVIDIA_L4\" # @param [\"NVIDIA_L4\", \"None\"] {isTemplate: true}\n",
@@ -131,7 +131,7 @@
"from google.cloud import aiplatform\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"# Get the default cloud project id.\n",
@@ -214,7 +214,7 @@
"HUGGING_FACE_MODEL_ID = \"Qwen/Qwen3-Embedding-8B\" # @param {type: \"string\", isTemplate: true}\n",
"\n",
"# The pre-built serving docker images for TEI.\n",
"TEI_DOCKER_URI = \"us-docker.pkg.dev/deeplearning-platform-release/vertex-model-garden/hf-tei.cu125.0-1.ubuntu2204.py310:model-garden.hf-tei-0-1-release_20250727.00_p0\"\n",
"TEI_DOCKER_URI = \"us-docker.pkg.dev/deeplearning-platform-release/vertex-model-garden/hf-tei.cu125.0-1.ubuntu2204.py310:model-garden.hf-tei-0-1-release_20250828.01_p0\"\n",
"\n",
"machine_type = \"g2-standard-8\" # @param {type: \"string\", isTemplate: true}\n",
"accelerator_type = \"NVIDIA_L4\" # @param [\"NVIDIA_L4\", \"None\"] {isTemplate: true}\n",
@@ -130,7 +130,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"\n",
@@ -126,7 +126,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"\n",
@@ -164,7 +164,7 @@
"HF_TOKEN = \"\" # @param {type:\"string\", isTemplate: true}\n",
"\n",
"# The pre-built vLLM serving docker image.\n",
"VLLM_DOCKER_URI = \"us-docker.pkg.dev/vertex-ai/vertex-vision-model-garden-dockers/pytorch-vllm-serve:20250801_0916_RC01\"\n",
"VLLM_DOCKER_URI = \"us-docker.pkg.dev/vertex-ai/vertex-vision-model-garden-dockers/pytorch-vllm-serve:20250905_0916_RC01\"\n",
"SERVING_CONTAINER_IMAGE_URI = VLLM_DOCKER_URI\n",
"LABEL = \"vllm\"\n",
"\n",
@@ -313,7 +313,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -322,6 +321,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -148,7 +148,7 @@
"models, endpoints = {}, {}\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"# Get the default cloud project id.\n",
@@ -146,7 +146,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"\n",
@@ -145,7 +145,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"LABEL = \"endpoint\"\n",
@@ -132,7 +132,7 @@
"from google.cloud import aiplatform\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -4,11 +4,12 @@
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "7d9bbf86da5e"
},
"outputs": [],
"source": [
"# Copyright 2024 Google LLC\n",
"# Copyright 2025 Google LLC\n",
"#\n",
"# Licensed under the Apache License, Version 2.0 (the \"License\");\n",
"# you may not use this file except in compliance with the License.\n",
@@ -31,26 +32,23 @@
"source": [
"# Vertex AI Model Garden - Stable Diffusion XL 1.0 - TPU v5e\n",
"\n",
"<table align=\"left\">\n",
" <td>\n",
" <a href=\"https://colab.research.google.com/github/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/model_garden/model_garden_jax_stable_diffusion_xl.ipynb\">\n",
" <img src=\"https://cloud.google.com/ml-engine/images/colab-logo-32px.png\" alt=\"Colab logo\"> Run in Colab\n",
" </a>\n",
" </td>\n",
" <td>\n",
" <a href=\"https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/model_garden/model_garden_jax_stable_diffusion_xl.ipynb\">\n",
" <img src=\"https://github.githubassets.com/assets/GitHub-Mark-ea2971cee799.png\" alt=\"GitHub logo\">\n",
" View on GitHub\n",
" </a>\n",
" </td>\n",
" <td>\n",
"<table><tbody><tr>\n",
" <td style=\"text-align: center\">\n",
" <a href=\"https://console.cloud.google.com/vertex-ai/notebooks/deploy-notebook?download_url=https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/main/notebooks/community/model_garden/model_garden_jax_stable_diffusion_xl.ipynb\">\n",
" <img src=\"https://lh3.googleusercontent.com/UiNooY4LUgW_oTvpsNhPpQzsstV5W8F7rYgxgGBD85cWJoLmrOzhVs_ksK_vgx40SHs7jCqkTkCk=e14-rj-sc0xffffff-h130-w32\" alt=\"Vertex AI logo\">\n",
"Open in Vertex AI Workbench\n",
" <img alt=\"Workbench logo\" src=\"https://lh3.googleusercontent.com/UiNooY4LUgW_oTvpsNhPpQzsstV5W8F7rYgxgGBD85cWJoLmrOzhVs_ksK_vgx40SHs7jCqkTkCk=e14-rj-sc0xffffff-h130-w32\" width=\"32px\"><br> Run in Workbench\n",
" </a>\n",
" (a Python-3 GPU notebook with preinstalled HuggingFace/transformer libraries is recommended)\n",
" </td>\n",
"</table>"
" <td style=\"text-align: center\">\n",
" <a href=\"https://console.cloud.google.com/vertex-ai/colab/import/https:%2F%2Fraw.githubusercontent.com%2FGoogleCloudPlatform%2Fvertex-ai-samples%2Fmain%2Fnotebooks%2Fcommunity%2Fmodel_garden%2Fmodel_garden_jax_stable_diffusion_xl.ipynb\">\n",
" <img alt=\"Google Cloud Colab Enterprise logo\" src=\"https://lh3.googleusercontent.com/JmcxdQi-qOpctIvWKgPtrzZdJJK-J3sWE1RsfjZNwshCFgE_9fULcNpuXYTilIR2hjwN\" width=\"32px\"><br> Run in Colab Enterprise\n",
" </a>\n",
" </td>\n",
" <td style=\"text-align: center\">\n",
" <a href=\"https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/model_garden/model_garden_jax_stable_diffusion_xl.ipynb\">\n",
" <img alt=\"GitHub logo\" src=\"https://github.githubassets.com/assets/GitHub-Mark-ea2971cee799.png\" width=\"32px\"><br> View on GitHub\n",
" </a>\n",
" </td>\n",
"</tr></tbody></table>"
]
},
{
@@ -61,190 +59,353 @@
"source": [
"## Overview\n",
"\n",
"This notebook demonstrates how to deploy the [stabilityai/stable-diffusion-xl-base-1.0](https://huggingface.co/stabilityai/stable-diffusion-xl-base-1.0) model on Vertex AI for online prediction.\n",
"This notebook demonstrates how to deploy a **Stable-Diffusion-Xl-Base** open model on Google Cloud Vertex AI.\n",
"\n",
"### Objective\n",
"### Objectives\n",
"\n",
"- Deploy the model to a [Vertex AI Endpoint resource](https://cloud.google.com/vertex-ai/docs/predictions/using-private-endpoints).\n",
"- Run online predictions for text-to-image.\n",
"- Deploy Stable-Diffusion-Xl-Base using containerized backends like [vLLM](https://github.com/vllm-project/vllm) on GPU.\n",
"- Use the deployed model to serve chat completion requests for both text and multimodal inputs.\n",
"\n",
"### File a Bug\n",
"\n",
"If you encounter issues with this notebook, report them on [GitHub](https://github.com/GoogleCloudPlatform/vertex-ai-samples/issues/new).\n",
"\n",
"### Costs\n",
"\n",
"This tutorial uses billable components of Google Cloud:\n",
"\n",
"* Vertex AI\n",
"* Cloud Storage\n",
"- Vertex AI\n",
"- Cloud Storage\n",
"\n",
"Learn about [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing) and [Cloud Storage pricing](https://cloud.google.com/storage/pricing), and use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to generate a cost estimate based on your projected usage."
"Refer to the [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing) and [Cloud Storage pricing](https://cloud.google.com/storage/pricing) pages for more information. Use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to estimate your projected costs."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "264c07757582"
"id": "jeYw-Czg-DFy"
},
"source": [
"## Before you begin\n",
"\n",
"**NOTE**: \n",
"\n",
"* Jupyter runs lines prefixed with `!` as shell commands, and it interpolates Python variables prefixed with `$` into these commands.\n",
"* This Notebook demonstrate how to deploy the model [stabilityai/stable-diffusion-xl-base-1.0](https://huggingface.co/stabilityai/stable-diffusion-xl-base-1.0) on Vertex AI prediction endpoint with a TPU v5e instance (machine type of `ct5lp-hightpu-1t`). Please ensure you have enough resource quota in region `us-west1`. If not, please follow the [instructions](https://cloud.google.com/vertex-ai/docs/predictions/use-tpu#securing_capacity) to get quota."
"## Get Started"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "ioensNKM8ned"
"id": "KgyhGvEzBDkj"
},
"source": [
"### Setup notebook"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "d73ffa0c0b83"
},
"source": [
"#### Colab\n",
"Run the following commands for Colab."
"### Install Vertex AI SDK and other required packages"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "2707b02ef5df"
"cellView": "form",
"id": "iCacdLqG-IsH"
},
"outputs": [],
"source": [
"if \"google.colab\" in str(get_ipython()):\n",
" ! pip3 install --upgrade google-cloud-aiplatform\n",
" from google.colab import auth as google_auth\n",
"\n",
" google_auth.authenticate_user()\n",
"\n",
"# Restart the notebook kernel after installs.\n",
"import IPython\n",
"\n",
"app = IPython.Application.instance()\n",
"app.kernel.do_shutdown(True)"
"%pip install --upgrade --force-reinstall --quiet 'google-cloud-aiplatform>=1.106.0' 'openai' 'google-auth==2.27.0' 'requests==2.32.3'"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "bb7adab99e41"
"id": "HUKCrpBy-3yf"
},
"source": [
"### Setup Google Cloud project\n",
"### Authenticate the Notebook Environment (Colab only)\n",
"\n",
"1. [Select or create a Google Cloud project](https://console.cloud.google.com/cloud-resource-manager). When you first create an account, you get a $300 free credit towards your compute/storage costs.\n",
"\n",
"1. [Make sure that billing is enabled for your project](https://cloud.google.com/billing/docs/how-to/modify-project).\n",
"\n",
"1. [Enable the Vertex AI API and Compute Engine API](https://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com,compute_component).\n",
"\n",
"1. [Create a Cloud Storage bucket](https://cloud.google.com/storage/docs/creating-buckets) for storing experiment outputs.\n",
"\n",
"1. [Create a service account](https://cloud.google.com/iam/docs/service-accounts-create#iam-service-accounts-create-console) with `Vertex AI User` and `Storage Object Admin` roles for deploying models to Vertex AI endpoint."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "6c460088b873"
},
"source": [
"Set following variables for experiments environment:"
"If you're running this notebook in Google Colab, run the following cell to authenticate."
]
},
{
"cell_type": "code",
"execution_count": 1,
"execution_count": null,
"metadata": {
"id": "855d6b96f291"
"cellView": "form",
"id": "JXwCT1kn-3Gu"
},
"outputs": [],
"source": [
"# Cloud project id.\n",
"PROJECT_ID = \"\" # @param {type:\"string\"}\n",
"import sys\n",
"\n",
"# The region you want to launch jobs in.\n",
"REGION = \"\" # @param {type:\"string\"}\n",
"if \"google.colab\" in sys.modules:\n",
" from google.colab import auth\n",
"\n",
"# The Cloud Storage bucket for storing experiments output. Fill it without the 'gs://' prefix.\n",
"GCS_BUCKET = \"\" # @param {type:\"string\"}\n",
"\n",
"# The service account for deploying fine tuned model.\n",
"SERVICE_ACCOUNT = \"\" # @param {type:\"string\"}"
" auth.authenticate_user()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "e828eb320337"
"id": "AcW2nwB8-7yC"
},
"source": [
"Initialize Vertex AI API:"
]
},
{
"cell_type": "code",
"execution_count": 2,
"metadata": {
"id": "12cd25839741"
},
"outputs": [],
"source": [
"from google.cloud import aiplatform\n",
"### Set Google Cloud Project Information\n",
"\n",
"aiplatform.init(project=PROJECT_ID, location=REGION, staging_bucket=GCS_BUCKET)"
"To get started with Vertex AI, ensure you have an existing Google Cloud project and that the [Vertex AI API is enabled](https://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com).\n",
"\n",
"See the guide on [setting up your project and development environment](https://cloud.google.com/vertex-ai/docs/start/cloud-environment). Also confirm that [billing is enabled](https://cloud.google.com/billing/docs/how-to/modify-project).\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "eIVLp0oE--k-"
},
"outputs": [],
"source": [
"# Use the environment variable if the user doesn't provide Project ID.\n",
"import os\n",
"\n",
"import vertexai\n",
"\n",
"PROJECT_ID = \"[your-project-id]\" # @param {type: \"string\", placeholder: \"[your-project-id]\", isTemplate: true}\n",
"\n",
"if not PROJECT_ID or PROJECT_ID == \"[your-project-id]\":\n",
" PROJECT_ID = str(os.environ.get(\"GOOGLE_CLOUD_PROJECT\"))\n",
"\n",
"REGION = os.environ.get(\"GOOGLE_CLOUD_REGION\", \"us-central1\")\n",
"\n",
"vertexai.init(project=PROJECT_ID, location=REGION)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "2cc825514deb"
"id": "Q0CXrvcZH_aw"
},
"source": [
"### Define constants"
"### Import libraries"
]
},
{
"cell_type": "code",
"execution_count": 3,
"execution_count": null,
"metadata": {
"id": "b42bd4fa2b2d"
"cellView": "form",
"id": "3G2UXB82ICs6"
},
"outputs": [],
"source": [
"# The pre-built serving docker image. It contains serving scripts and models.\n",
"SERVE_DOCKER_URI = \"us-docker.pkg.dev/vertex-ai/vertex-vision-model-garden-dockers/jax-diffusers-serve-tpu:20240110_1526_RC00\""
"from vertexai import model_garden"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "0c250872074f"
"id": "upYRiGtP_-iN"
},
"source": [
"### Define common functions"
"## Deploy model"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "H2WC_0hXDVXc"
},
"source": [
"### Choose model variant"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "u41zbNa2EoFq"
},
"source": [
"You can proceed with the default model variant or select a different one."
]
},
{
"cell_type": "code",
"execution_count": 6,
"execution_count": null,
"metadata": {
"id": "354da31189dc"
"cellView": "form",
"id": "-fgC4NLSDkF7"
},
"outputs": [],
"source": [
"model_version = \"stable-diffusion-xl-base-1.0\" # @param [\"stable-diffusion-xl-base-1.0\"] {isTemplate:true}\n",
"MODEL_NAME = f\"stability-ai/stable-diffusion-xl-base@{model_version}\""
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "VRnUgU8LF3_i"
},
"source": [
"To see all deployable model variants available in Model Garden, use:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "-QLd-wshF6sB"
},
"outputs": [],
"source": [
"all_model_versions = model_garden.list_deployable_models(\n",
" model_filter=\"stable-diffusion-xl-base\", list_hf_models=False\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "N0UeFHa2GO63"
},
"source": [
"Once you've selected a model variant, initialize it:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "GZiV3trBBcA3"
},
"outputs": [],
"source": [
"model = model_garden.OpenModel(MODEL_NAME)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "-0cL378wFlvf"
},
"source": [
"### Check the Deployment Configuration\n",
"\n",
"Use the `list_deploy_options()` method to view the verified deployment configurations for your selected model. This helps ensure you have sufficient resources (e.g., GPU quota) available to deploy it."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "zm73g7vFFm9N"
},
"outputs": [],
"source": [
"deploy_options = model.list_deploy_options(concise=True)\n",
"print(deploy_options)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "WjV499VsGwrD"
},
"source": [
"### Deploy the Model\n",
"\n",
"Now that you’ve reviewed the deployment options, use the `deploy()` method to serve the selected open model to a Vertex AI endpoint. Deployment time may vary depending on the model size and infrastructure requirements.\n",
"\n",
"> **Note**: If the model requires accepting a license agreement (EULA), set the `accept_eula=True` flag in the deploy call. Set `use_dedicated_endpoint` to False if you don't want to use [dedicated endpoint](https://cloud.google.com/vertex-ai/docs/general/deployment#create-dedicated-endpoint)."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "wX1itVTvXdEP"
},
"outputs": [],
"source": [
"use_dedicated_endpoint = True"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "MRmPFEPoGzsB"
},
"outputs": [],
"source": [
"endpoint = model.deploy(\n",
" accept_eula=True,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "PHBtn8DQp-ID"
},
"source": [
"Alternatively, you can select one of the verified deployment configurations listed above."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "ADsJG8JYqI6c"
},
"outputs": [],
"source": [
"endpoint = model.deploy(\n",
" accept_eula=True,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
" serving_container_image_uri=\"us-docker.pkg.dev/deeplearning-platform-release/vertex-model-garden/pytorch-inference.cu125.0-4.ubuntu2204.py310\",\n",
" machine_type=\"a2-ultragpu-1g\",\n",
" accelerator_type=\"NVIDIA_A100_80GB\",\n",
" accelerator_count=1,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "kqSUK2CwsImi"
},
"source": [
"To further customize your deployment, you can configure:\n",
"\n",
"- **Compute Resources**: Machine type, replica count (min/max), accelerator type and quantity.\n",
"- **Infrastructure**: Use Spot VMs, reservation affinity, or dedicated endpoints.\n",
"- **Serving Container**: Customize container image, ports, health checks, and environment variables.\n",
"\n",
"See the [Model Garden SDK README](https://github.com/googleapis/python-aiplatform/blob/main/vertexai/model_garden/README.md) for advanced configuration options."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "4ab04da3ec9a"
},
"outputs": [],
"source": [
"# @title Predict\n",
"\n",
"# @markdown Once deployed, you can send a batch of text prompts to the endpoint to generated images.\n",
"\n",
"# @markdown When deployed on one TPU V5e instance, the averaged inference time of one image is ~3 seconds.\n",
"\n",
"\n",
"import base64\n",
"from io import BytesIO\n",
"\n",
"from google.cloud import aiplatform\n",
"from PIL import Image\n",
"\n",
"\n",
@@ -264,117 +425,19 @@
" return grid\n",
"\n",
"\n",
"def deploy_model(model_id):\n",
" \"\"\"Create a Vertex AI Endpoint and deploy the specified model to the endpoint.\"\"\"\n",
" model_name = model_id + \"-tpu\"\n",
" endpoint = aiplatform.Endpoint.create(display_name=f\"{model_name}-endpoint\")\n",
"\n",
" model = aiplatform.Model.upload(\n",
" display_name=model_name,\n",
" serving_container_image_uri=SERVE_DOCKER_URI,\n",
" serving_container_ports=[8080],\n",
" serving_container_predict_route=\"/predict\",\n",
" serving_container_health_route=\"/health\",\n",
" model_garden_source_model_name=\"publishers/stability-ai/models/stable-diffusion-xl-base\"\n",
" )\n",
" machine_type = \"ct5lp-hightpu-1t\"\n",
"\n",
" model.deploy(\n",
" endpoint=endpoint,\n",
" machine_type=machine_type,\n",
" deploy_request_timeout=1800,\n",
" service_account=SERVICE_ACCOUNT,\n",
" enable_access_logging=True,\n",
" min_replica_count=1,\n",
" sync=True,\n",
" system_labels={\n",
" \"NOTEBOOK_NAME\": \"model_garden_jax_stable_diffusion_xl.ipynb\"\n",
" },\n",
" )\n",
" return model, endpoint"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "bf7f82732e61"
},
"source": [
"## Upload and Deploy models"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "1cc26e68d7b0"
},
"source": [
"This section uploads the model to Model Registry and deploys it to a Vertex AI Endpoint resource.\n",
"\n",
"The model deployment step will take ~30 minutes to complete."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "cd7b56421392"
},
"source": [
"### Text-to-image"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "6d331b1ea337"
},
"source": [
"Deploy the stable diffusion xl model for the text-to-image task.\n",
"\n",
"Once deployed, you can send a batch of text prompts to the endpoint to generated images.\n",
"\n",
"When deployed on one TPU V5e instance, the averaged inference time of one image is ~3 seconds."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "bf55e38815dc"
},
"outputs": [],
"source": [
"# Set the model_id to \"stabilityai/stable-diffusion-xl-base-1.0\" to load the OSS pre-trained model.\n",
"model, endpoint = deploy_model(\n",
" model_id=\"stabilityai/stable-diffusion-xl-base-1.0\",\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "4ab04da3ec9a"
},
"outputs": [],
"source": [
"instances = [\n",
" {\n",
" \"prompt\": \"Photorealistic whale swimming in abyss\",\n",
" \"height\": 1024,\n",
" \"width\": 1024,\n",
" },\n",
" {\n",
" \"prompt\": \"Photorealistic happy dog running\",\n",
" \"height\": 1024,\n",
" \"width\": 1024,\n",
" },\n",
"]\n",
"response = endpoint.predict(instances=instances)\n",
"\n",
"images = [\n",
" base64_to_image(prediction.get(\"images\")[0]) for prediction in response.predictions\n",
" base64_to_image(response.predictions[0]) for prediction in response.predictions\n",
"]\n",
"\n",
"image_grid(images, rows=1)"
]
},
@@ -384,22 +447,22 @@
"id": "af21a3cff1e0"
},
"source": [
"### Clean up resources:"
"### Clean up resources"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "911406c1561e"
},
"outputs": [],
"source": [
"# Undeploy model and delete endpoint.\n",
"endpoint.delete(force=True)\n",
"# @markdown Delete the endpoint.\n",
"\n",
"# Delete models.\n",
"model.delete()"
"if endpoint:\n",
" endpoint.delete(force=True)"
]
}
],
@@ -143,7 +143,7 @@
"from PIL import Image\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -183,7 +183,7 @@
" custom_job as gca_custom_job_compat\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")"
]
},
@@ -1386,7 +1386,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -1395,6 +1394,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -126,7 +126,7 @@
" custom_job as gca_custom_job_compat\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -151,7 +151,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"LABEL = \"vllm_gpu\"\n",
@@ -343,7 +343,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -352,6 +351,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -124,7 +124,7 @@
"from PIL import Image\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"# Get the default cloud project id.\n",
@@ -131,7 +131,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"\n",
@@ -126,7 +126,7 @@
"from google.cloud import aiplatform\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -127,7 +127,7 @@
"from google.cloud import aiplatform\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"\n",
@@ -135,7 +135,7 @@
"from google.cloud.aiplatform import hyperparameter_tuning as hpt\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -145,7 +145,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")"
]
},
@@ -120,7 +120,7 @@
"# @markdown | a3-highgpu-8g | 8 NVIDIA_H100_80GB | us-central1, europe-west4, us-west1, asia-southeast1 |\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -128,7 +128,7 @@
"models, endpoints = {}, {}\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"# Get the default cloud project id.\n",
@@ -405,7 +405,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -414,6 +413,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -126,7 +126,7 @@
"models, endpoints = {}, {}\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"# Get the default cloud project id.\n",
@@ -323,7 +323,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -332,6 +331,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -124,7 +124,7 @@
"models, endpoints = {}, {}\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"\n",
@@ -128,7 +128,7 @@
"from google.cloud import aiplatform\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -193,7 +193,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -202,6 +201,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -136,7 +136,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -54,49 +54,48 @@
{
"cell_type": "markdown",
"metadata": {
"id": "d8cd12648da4"
"id": "3de7470326a2"
},
"source": [
"## Overview\n",
"\n",
"This notebook demonstrates deploying the pre-trained [BLIP2](https://huggingface.co/Salesforce/blip2-opt-2.7b) model on Vertex AI for online prediction.\n",
"This notebook demonstrates how to deploy a **Blip 2-Opt- 2. 7-B** open model on Google Cloud Vertex AI.\n",
"\n",
"### Objective\n",
"### Objectives\n",
"\n",
"- Upload the model to [Model Registry](https://cloud.google.com/vertex-ai/docs/model-registry/introduction).\n",
"- Deploy the model on [Endpoint](https://cloud.google.com/vertex-ai/docs/predictions/using-private-endpoints).\n",
"- Run online predictions for image captioning.\n",
"- Deploy Blip 2-Opt- 2. 7-B using containerized backends like [vLLM](https://github.com/vllm-project/vllm) on GPU.\n",
"- Use the deployed model to serve chat completion requests for both text and multimodal inputs.\n",
"\n",
"### File a bug\n",
"### File a Bug\n",
"\n",
"File a bug on [GitHub](https://github.com/GoogleCloudPlatform/vertex-ai-samples/issues/new) if you encounter any issue with the notebook.\n",
"If you encounter issues with this notebook, report them on [GitHub](https://github.com/GoogleCloudPlatform/vertex-ai-samples/issues/new).\n",
"\n",
"### Costs\n",
"\n",
"This tutorial uses billable components of Google Cloud:\n",
"\n",
"* Vertex AI\n",
"* Cloud Storage\n",
"- Vertex AI\n",
"- Cloud Storage\n",
"\n",
"Learn about [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing), [Cloud Storage pricing](https://cloud.google.com/storage/pricing), and use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to generate a cost estimate based on your projected usage."
"Refer to the [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing) and [Cloud Storage pricing](https://cloud.google.com/storage/pricing) pages for more information. Use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to estimate your projected costs."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "264c07757582"
"id": "jeYw-Czg-DFy"
},
"source": [
"## Before you begin"
"## Get Started"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "d73ffa0c0b83"
"id": "KgyhGvEzBDkj"
},
"source": [
"### Colab only"
"### Install Vertex AI SDK and other required packages"
]
},
{
@@ -104,27 +103,325 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "2707b02ef5df"
"id": "iCacdLqG-IsH"
},
"outputs": [],
"source": [
"# @title Setup Google Cloud project\n",
"%pip install --upgrade --force-reinstall --quiet 'google-cloud-aiplatform>=1.106.0' 'openai' 'google-auth==2.27.0' 'requests==2.32.3'"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "HUKCrpBy-3yf"
},
"source": [
"### Authenticate the Notebook Environment (Colab only)\n",
"\n",
"# @markdown 1. [Make sure that billing is enabled for your project](https://cloud.google.com/billing/docs/how-to/modify-project).\n",
"If you're running this notebook in Google Colab, run the following cell to authenticate."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "JXwCT1kn-3Gu"
},
"outputs": [],
"source": [
"import sys\n",
"\n",
"# @markdown 2. **[Optional]** Set region. If not set, the region will be set automatically according to Colab Enterprise environment.\n",
"if \"google.colab\" in sys.modules:\n",
" from google.colab import auth\n",
"\n",
"REGION = \"\" # @param {type:\"string\"}\n",
" auth.authenticate_user()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "AcW2nwB8-7yC"
},
"source": [
"### Set Google Cloud Project Information\n",
"\n",
"# Upgrade Vertex AI SDK.\n",
"! pip3 install --upgrade --quiet 'google-cloud-aiplatform==1.103.0'\n",
"To get started with Vertex AI, ensure you have an existing Google Cloud project and that the [Vertex AI API is enabled](https://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com).\n",
"\n",
"See the guide on [setting up your project and development environment](https://cloud.google.com/vertex-ai/docs/start/cloud-environment). Also confirm that [billing is enabled](https://cloud.google.com/billing/docs/how-to/modify-project).\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "eIVLp0oE--k-"
},
"outputs": [],
"source": [
"# Use the environment variable if the user doesn't provide Project ID.\n",
"import os\n",
"\n",
"import vertexai\n",
"\n",
"PROJECT_ID = \"[your-project-id]\" # @param {type: \"string\", placeholder: \"[your-project-id]\", isTemplate: true}\n",
"\n",
"if not PROJECT_ID or PROJECT_ID == \"[your-project-id]\":\n",
" PROJECT_ID = str(os.environ.get(\"GOOGLE_CLOUD_PROJECT\"))\n",
"\n",
"REGION = os.environ.get(\"GOOGLE_CLOUD_REGION\", \"us-central1\")\n",
"\n",
"vertexai.init(project=PROJECT_ID, location=REGION)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "Q0CXrvcZH_aw"
},
"source": [
"### Import libraries"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "3G2UXB82ICs6"
},
"outputs": [],
"source": [
"from vertexai import model_garden"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "upYRiGtP_-iN"
},
"source": [
"## Deploy model"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "H2WC_0hXDVXc"
},
"source": [
"### Choose model variant"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "u41zbNa2EoFq"
},
"source": [
"You can proceed with the default model variant or select a different one."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "-fgC4NLSDkF7"
},
"outputs": [],
"source": [
"model_version = \"blip2-opt-2.7b\" # @param [\"blip2-opt-2.7b\", \"blip2-opt-2.7b-image-to-text\", \"blip2-opt-2.7b-visual-question-answering\"] {isTemplate:true}\n",
"MODEL_NAME = f\"salesforce/blip2-opt-2.7-b@{model_version}\""
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "VRnUgU8LF3_i"
},
"source": [
"To see all deployable model variants available in Model Garden, use:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "-QLd-wshF6sB"
},
"outputs": [],
"source": [
"all_model_versions = model_garden.list_deployable_models(\n",
" model_filter=\"blip2-opt-2.7-b\", list_hf_models=False\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "N0UeFHa2GO63"
},
"source": [
"Once you've selected a model variant, initialize it:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "GZiV3trBBcA3"
},
"outputs": [],
"source": [
"model = model_garden.OpenModel(MODEL_NAME)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "-0cL378wFlvf"
},
"source": [
"### Check the Deployment Configuration\n",
"\n",
"Use the `list_deploy_options()` method to view the verified deployment configurations for your selected model. This helps ensure you have sufficient resources (e.g., GPU quota) available to deploy it."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "zm73g7vFFm9N"
},
"outputs": [],
"source": [
"deploy_options = model.list_deploy_options(concise=True)\n",
"print(deploy_options)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "WjV499VsGwrD"
},
"source": [
"### Deploy the Model\n",
"\n",
"Now that you’ve reviewed the deployment options, use the `deploy()` method to serve the selected open model to a Vertex AI endpoint. Deployment time may vary depending on the model size and infrastructure requirements.\n",
"\n",
"> **Note**: If the model requires accepting a license agreement (EULA), set the `accept_eula=True` flag in the deploy call. Set `use_dedicated_endpoint` to False if you don't want to use [dedicated endpoint](https://cloud.google.com/vertex-ai/docs/general/deployment#create-dedicated-endpoint)."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "wX1itVTvXdEP"
},
"outputs": [],
"source": [
"use_dedicated_endpoint = True"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "MRmPFEPoGzsB"
},
"outputs": [],
"source": [
"endpoint = model.deploy(\n",
" accept_eula=True,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "PHBtn8DQp-ID"
},
"source": [
"Alternatively, you can select one of the verified deployment configurations listed above."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "ADsJG8JYqI6c"
},
"outputs": [],
"source": [
"endpoint = model.deploy(\n",
" accept_eula=True,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
" serving_container_image_uri=\"us-docker.pkg.dev/vertex-ai/vertex-vision-model-garden-dockers/pytorch-transformers-serve\",\n",
" machine_type=\"n1-standard-8\",\n",
" accelerator_type=\"NVIDIA_TESLA_T4\",\n",
" accelerator_count=1,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "kqSUK2CwsImi"
},
"source": [
"To further customize your deployment, you can configure:\n",
"\n",
"- **Compute Resources**: Machine type, replica count (min/max), accelerator type and quantity.\n",
"- **Infrastructure**: Use Spot VMs, reservation affinity, or dedicated endpoints.\n",
"- **Serving Container**: Customize container image, ports, health checks, and environment variables.\n",
"\n",
"See the [Model Garden SDK README](https://github.com/googleapis/python-aiplatform/blob/main/vertexai/model_garden/README.md) for advanced configuration options."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "b2-D20D_W3XG"
},
"outputs": [],
"source": [
"# @title Predict"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "NSpOSQuLCNog"
},
"source": [
"NOTE: The model weights will be downloaded after the deployment succeeds. Thus additional 10 minutes of waiting time is needed **after** the above model deployment step succeeds and before you run the next step below. Otherwise you might see a `ServiceUnavailable: 503 502:Bad Gateway` error when you send requests to the endpoint."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "eUoMSibECGiY"
},
"outputs": [],
"source": [
"import importlib\n",
"import os\n",
"from io import BytesIO\n",
"\n",
"import requests\n",
"from google.cloud import aiplatform\n",
"from PIL import Image\n",
"\n",
"if os.environ.get(\"VERTEX_PRODUCT\") != \"COLAB_ENTERPRISE\":\n",
@@ -132,157 +429,10 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
"\n",
"\n",
"# Get the default cloud project id.\n",
"PROJECT_ID = os.environ[\"GOOGLE_CLOUD_PROJECT\"]\n",
"\n",
"# Get the default region for launching jobs.\n",
"if not REGION:\n",
" REGION = os.environ[\"GOOGLE_CLOUD_REGION\"]\n",
"\n",
"# Initialize Vertex AI API.\n",
"print(\"Initializing Vertex AI API.\")\n",
"aiplatform.init(project=PROJECT_ID, location=REGION)\n",
"\n",
"! gcloud config set project $PROJECT_ID\n",
"\n",
"import vertexai\n",
"\n",
"vertexai.init(\n",
" project=PROJECT_ID,\n",
" location=REGION,\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "cac4478ae098"
},
"outputs": [],
"source": [
"# @title Select model parameters\n",
"\n",
"MODEL_ID = \"Salesforce/blip2-opt-2.7b\"\n",
"TASK = \"image-to-text\"\n",
"\n",
"# @markdown Set the accelerator type.\n",
"accelerator_type = \"NVIDIA_L4\" # @param[\"NVIDIA_TESLA_T4\", \"NVIDIA_L4\"]\n",
"\n",
"if accelerator_type == \"NVIDIA_TESLA_T4\":\n",
" machine_type = \"n1-standard-8\"\n",
" accelerator_count = 1\n",
"elif accelerator_type == \"NVIDIA_L4\":\n",
" machine_type = \"g2-standard-12\"\n",
" accelerator_count = 1\n",
"else:\n",
" print(f\"Unsupported accelerator type: {accelerator_type}\")\n",
"\n",
"\n",
"# The pre-built serving docker image.\n",
"# The model artifacts are embedded within the container, except for model weights which will be downloaded during deployment.\n",
"SERVE_DOCKER_URI = \"us-docker.pkg.dev/deeplearning-platform-release/vertex-model-garden/pytorch-inference.cu125.0-4.ubuntu2204.py310:model-garden.pytorch-inference-0-4-gpu-release_20250708.04_p0\"\n",
"\n",
"# @markdown Set use_dedicated_endpoint to False if you don't want to use [dedicated endpoint](https://cloud.google.com/vertex-ai/docs/general/deployment#create-dedicated-endpoint). Note that [dedicated endpoint does not support VPC Service Controls](https://cloud.google.com/vertex-ai/docs/predictions/choose-endpoint-type), uncheck the box if you are using VPC-SC.\n",
"use_dedicated_endpoint = True # @param {type:\"boolean\"}\n",
"\n",
"\n",
"def deploy_model(\n",
" model_name: str,\n",
" model_id: str,\n",
" task: str,\n",
" machine_type: str = \"n1-standard-8\",\n",
" accelerator_type: str = \"NVIDIA_TESLA_T4\",\n",
" accelerator_count: int = 1,\n",
" use_dedicated_endpoint: bool = False,\n",
"):\n",
" model_name = \"blip2\"\n",
" endpoint = aiplatform.Endpoint.create(\n",
" display_name=f\"{TASK}-endpoint\",\n",
" dedicated_endpoint_enabled=use_dedicated_endpoint,\n",
" )\n",
" serving_env = {\n",
" \"MODEL_ID\": model_id,\n",
" \"TASK\": task,\n",
" \"DEPLOY_SOURCE\": \"notebook\",\n",
" }\n",
" # If the model_id is a GCS path, use artifact_uri to pass it to serving docker.\n",
" artifact_uri = model_id if model_id.startswith(\"gs://\") else None\n",
" model = aiplatform.Model.upload(\n",
" display_name=model_name,\n",
" serving_container_image_uri=SERVE_DOCKER_URI,\n",
" serving_container_ports=[7080],\n",
" serving_container_predict_route=\"/predict\",\n",
" serving_container_health_route=\"/health\",\n",
" serving_container_environment_variables=serving_env,\n",
" artifact_uri=artifact_uri,\n",
" model_garden_source_model_name=\"publishers/salesforce/models/blip2-opt-2.7-b\",\n",
" )\n",
" model.deploy(\n",
" endpoint=endpoint,\n",
" machine_type=\"n1-standard-8\",\n",
" accelerator_type=\"NVIDIA_TESLA_T4\",\n",
" accelerator_count=1,\n",
" deploy_request_timeout=1800,\n",
" system_labels={\n",
" \"NOTEBOOK_NAME\": \"model_garden_pytorch_blip2.ipynb\",\n",
" \"NOTEBOOK_ENVIRONMENT\": common_util.get_deploy_source(),\n",
" },\n",
" )\n",
" return model, endpoint"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "1hyUb_79_MQm"
},
"source": [
"#### Image Captioning"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "_a8X09zz4Xgi"
},
"outputs": [],
"source": [
"# @title Deploy\n",
"\n",
"LABEL = \"blip2\"\n",
"models[LABEL], endpoints[LABEL] = deploy_model(\n",
" model_name=common_util.get_job_name_with_datetime(prefix=MODEL_ID),\n",
" model_id=MODEL_ID,\n",
" task=TASK,\n",
" machine_type=machine_type,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=1,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
")\n",
"\n",
"model = models[LABEL]\n",
"endpoint = endpoints[LABEL]"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "80b3fd2ace09"
},
"source": [
"NOTE: The model weights will be downloaded after the deployment succeeds. Thus additional 10 minutes of waiting time is needed **after** the above model deployment step succeeds and before you run the next step below. Otherwise you might see a `ServiceUnavailable: 503 502:Bad Gateway` error when you send requests to the endpoint."
]
},
{
"cell_type": "code",
"execution_count": null,
@@ -292,7 +442,10 @@
},
"outputs": [],
"source": [
"# @title Predict\n",
"# @title Image Captioning\n",
"\n",
"if \"visual-question-answering\" in MODEL_NAME:\n",
" raise ValueError(\"Use VQA (Visual-Question-Answering) section instead.\")\n",
"\n",
"INPUT_IMAGE = \"http://images.cocodataset.org/val2017/000000039769.jpg\" # @param\n",
"\n",
@@ -317,50 +470,6 @@
"print(preds)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "-CBm0Inc_WBo"
},
"source": [
"#### VQA (Visual-Question-Answering)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "26018d961cf9"
},
"outputs": [],
"source": [
"# @title Deploy\n",
"\n",
"TASK = \"visual-question-answering\"\n",
"model, endpoint = deploy_model(\n",
" model_name=common_util.get_job_name_with_datetime(prefix=MODEL_ID),\n",
" model_id=MODEL_ID,\n",
" task=TASK,\n",
" machine_type=machine_type,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
")\n",
"\n",
"model = models[LABEL]\n",
"endpoint = endpoints[LABEL]"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "WLEcOG_VO61Y"
},
"source": [
"NOTE: The model weights will be downloaded after the deployment succeeds. Thus additional 10 minutes of waiting time is needed **after** the above model deployment step succeeds and before you run the next step below. Otherwise you might see a `ServiceUnavailable: 503 502:Bad Gateway` error when you send requests to the endpoint."
]
},
{
"cell_type": "code",
"execution_count": null,
@@ -370,7 +479,10 @@
},
"outputs": [],
"source": [
"# @title Predict\n",
"# @title VQA (Visual-Question-Answering)\n",
"\n",
"if \"visual-question-answering\" not in MODEL_NAME:\n",
" raise ValueError(\"Use Image Captioning section instead.\")\n",
"\n",
"INPUT_IMAGE = \"https://media.newyorker.com/cartoons/63dc6847be24a6a76d90eb99/master/w_1160,c_limit/230213_a26611_838.jpg\" # @param\n",
"\n",
@@ -394,7 +506,7 @@
"id": "UBMw4dOe_nMQ"
},
"source": [
"#### Delete the models and endpoints"
"#### Delete the endpoint"
]
},
{
@@ -406,16 +518,10 @@
},
"outputs": [],
"source": [
"# @markdown Delete the experiment models and endpoints to recycle the resources\n",
"# @markdown and avoid unnecessary continuous charges that may incur.\n",
"# @markdown Delete the endpoint.\n",
"\n",
"# Undeploy model and delete endpoint.\n",
"for endpoint in endpoints.values():\n",
" endpoint.delete(force=True)\n",
"\n",
"# Delete models.\n",
"for model in models.values():\n",
" model.delete()"
"if endpoint:\n",
" endpoint.delete(force=True)"
]
}
],
@@ -54,40 +54,48 @@
{
"cell_type": "markdown",
"metadata": {
"id": "d8cd12648da4"
"id": "3de7470326a2"
},
"source": [
"## Overview\n",
"\n",
"This notebook demonstrates deploying the pre-trained [BLIP Image Captioning](https://huggingface.co/Salesforce/blip-image-captioning-base) model on Vertex AI for online prediction.\n",
"This notebook demonstrates how to deploy a **Blip-Image-Captioning-Base** open model on Google Cloud Vertex AI.\n",
"\n",
"### Objective\n",
"### Objectives\n",
"\n",
"- Upload the model to [Model Registry](https://cloud.google.com/vertex-ai/docs/model-registry/introduction).\n",
"- Deploy the model on [Endpoint](https://cloud.google.com/vertex-ai/docs/predictions/using-private-endpoints).\n",
"- Run online predictions for image captioning.\n",
"- Deploy Blip-Image-Captioning-Base using containerized backends like [vLLM](https://github.com/vllm-project/vllm) on GPU.\n",
"- Use the deployed model to serve chat completion requests for both text and multimodal inputs.\n",
"\n",
"### File a bug\n",
"### File a Bug\n",
"\n",
"File a bug on [GitHub](https://github.com/GoogleCloudPlatform/vertex-ai-samples/issues/new) if you encounter any issue with the notebook.\n",
"If you encounter issues with this notebook, report them on [GitHub](https://github.com/GoogleCloudPlatform/vertex-ai-samples/issues/new).\n",
"\n",
"### Costs\n",
"\n",
"This tutorial uses billable components of Google Cloud:\n",
"\n",
"* Vertex AI\n",
"* Cloud Storage\n",
"- Vertex AI\n",
"- Cloud Storage\n",
"\n",
"Learn about [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing), [Cloud Storage pricing](https://cloud.google.com/storage/pricing), and use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to generate a cost estimate based on your projected usage."
"Refer to the [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing) and [Cloud Storage pricing](https://cloud.google.com/storage/pricing) pages for more information. Use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to estimate your projected costs."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "G3GYgynjaKG8"
"id": "jeYw-Czg-DFy"
},
"source": [
"## Before you begin"
"## Get Started"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "KgyhGvEzBDkj"
},
"source": [
"### Install Vertex AI SDK and other required packages"
]
},
{
@@ -95,71 +103,85 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "cDtpu7ZVaTFb"
"id": "iCacdLqG-IsH"
},
"outputs": [],
"source": [
"# @title Setup Google Cloud project\n",
"%pip install --upgrade --force-reinstall --quiet 'google-cloud-aiplatform>=1.106.0' 'openai' 'google-auth' 'requests'"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "HUKCrpBy-3yf"
},
"source": [
"### Authenticate the Notebook Environment (Colab only)\n",
"\n",
"# @markdown 1. [Make sure that billing is enabled for your project](https://cloud.google.com/billing/docs/how-to/modify-project).\n",
"If you're running this notebook in Google Colab, run the following cell to authenticate."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "JXwCT1kn-3Gu"
},
"outputs": [],
"source": [
"import sys\n",
"\n",
"# @markdown 2. **[Optional]** Set region. If not set, the region will be set automatically according to Colab Enterprise environment.\n",
"if \"google.colab\" in sys.modules:\n",
" from google.colab import auth\n",
"\n",
"REGION = \"\" # @param {type:\"string\"}\n",
" auth.authenticate_user()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "AcW2nwB8-7yC"
},
"source": [
"### Set Google Cloud Project Information\n",
"\n",
"# @markdown 3. If you want to run predictions with A100 80GB or H100 GPUs, we recommend using the regions listed below. **NOTE:** Make sure you have associated quota in selected regions. Click the links to see your current quota for each GPU type: [Nvidia A100 80GB](https://console.cloud.google.com/iam-admin/quotas?metric=aiplatform.googleapis.com%2Fcustom_model_serving_nvidia_a100_80gb_gpus), [Nvidia H100 80GB](https://console.cloud.google.com/iam-admin/quotas?metric=aiplatform.googleapis.com%2Fcustom_model_serving_nvidia_h100_gpus). You can request for quota following the instructions at [\"Request a higher quota\"](https://cloud.google.com/docs/quota/view-manage#requesting_higher_quota).\n",
"To get started with Vertex AI, ensure you have an existing Google Cloud project and that the [Vertex AI API is enabled](https://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com).\n",
"\n",
"# @markdown > | Machine Type | Accelerator Type | Recommended Regions |\n",
"# @markdown | ----------- | ----------- | ----------- |\n",
"# @markdown | a2-ultragpu-1g | 1 NVIDIA_A100_80GB | us-central1, us-east4, europe-west4, asia-southeast1, us-east4 |\n",
"# @markdown | a3-highgpu-2g | 2 NVIDIA_H100_80GB | us-west1, asia-southeast1, europe-west4 |\n",
"# @markdown | a3-highgpu-4g | 4 NVIDIA_H100_80GB | us-west1, asia-southeast1, europe-west4 |\n",
"# @markdown | a3-highgpu-8g | 8 NVIDIA_H100_80GB | us-central1, europe-west4, us-west1, asia-southeast1 |\n",
"\n",
"# Import the necessary packages\n",
"\n",
"# Upgrade Vertex AI SDK.\n",
"! pip3 install --upgrade --quiet 'google-cloud-aiplatform==1.103.0'\n",
"\n",
"# Import the necessary packages.\n",
"import importlib\n",
"See the guide on [setting up your project and development environment](https://cloud.google.com/vertex-ai/docs/start/cloud-environment). Also confirm that [billing is enabled](https://cloud.google.com/billing/docs/how-to/modify-project).\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "eIVLp0oE--k-"
},
"outputs": [],
"source": [
"# Use the environment variable if the user doesn't provide Project ID.\n",
"import os\n",
"from typing import Tuple\n",
"\n",
"from google.cloud import aiplatform\n",
"\n",
"if os.environ.get(\"VERTEX_PRODUCT\") != \"COLAB_ENTERPRISE\":\n",
" ! pip install --upgrade tensorflow\n",
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
")\n",
"\n",
"LABEL = \"transformers_gpu\"\n",
"models, endpoints = {}, {}\n",
"\n",
"\n",
"# Get the default cloud project id.\n",
"PROJECT_ID = os.environ[\"GOOGLE_CLOUD_PROJECT\"]\n",
"\n",
"# Get the default region for launching jobs.\n",
"if not REGION:\n",
" REGION = os.environ[\"GOOGLE_CLOUD_REGION\"]\n",
"\n",
"# Initialize Vertex AI API.\n",
"print(\"Initializing Vertex AI API.\")\n",
"aiplatform.init(project=PROJECT_ID, location=REGION)\n",
"\n",
"! gcloud config set project $PROJECT_ID\n",
"import vertexai\n",
"\n",
"vertexai.init(\n",
" project=PROJECT_ID,\n",
" location=REGION,\n",
")\n",
"PROJECT_ID = \"[your-project-id]\" # @param {type: \"string\", placeholder: \"[your-project-id]\", isTemplate: true}\n",
"\n",
"# @markdown Click \"Show Code\" to see more details."
"if not PROJECT_ID or PROJECT_ID == \"[your-project-id]\":\n",
" PROJECT_ID = str(os.environ.get(\"GOOGLE_CLOUD_PROJECT\"))\n",
"\n",
"REGION = os.environ.get(\"GOOGLE_CLOUD_REGION\", \"us-central1\")\n",
"\n",
"vertexai.init(project=PROJECT_ID, location=REGION)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "Q0CXrvcZH_aw"
},
"source": [
"### Import libraries"
]
},
{
@@ -167,35 +189,38 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "RYfNg5r8ezx1"
"id": "3G2UXB82ICs6"
},
"outputs": [],
"source": [
"# @title Select the model parameters\n",
"\n",
"MODEL_ID = \"Salesforce/blip-image-captioning-base\"\n",
"TASK = \"image-to-text\"\n",
"\n",
"# Set the machine specifications.\n",
"accelerator_type = \"NVIDIA_L4\" # @param [\"NVIDIA_L4\", \"NVIDIA_TESLA_V100\", \"NVIDIA_TESLA_T4\"]\n",
"if accelerator_type == \"NVIDIA_L4\":\n",
" accelerator_count = 1\n",
" machine_type = \"g2-standard-8\"\n",
"elif accelerator_type == \"NVIDIA_TESLA_V100\":\n",
" accelerator_count = 1\n",
" machine_type = \"n1-standard-4\"\n",
"elif accelerator_type == \"NVIDIA_TESLA_T4\":\n",
" accelerator_count = 1\n",
" machine_type = \"n1-standard-8\"\n",
"else:\n",
" raise ValueError(\n",
" f\"Recommended machine settings not found for: {accelerator_type}. To use another accelerator type, edit this code block to pass in an appropriate `training_machine_type`, `training_accelerator_type`, and `per_node_accelerator_count` by clicking `Show Code` and then modifying the code.\"\n",
" )\n",
"\n",
"# @markdown Set use_dedicated_endpoint to False if you don't want to use [dedicated endpoint](https://cloud.google.com/vertex-ai/docs/general/deployment#create-dedicated-endpoint). Note that [dedicated endpoint does not support VPC Service Controls](https://cloud.google.com/vertex-ai/docs/predictions/choose-endpoint-type), uncheck the box if you are using VPC-SC.\n",
"use_dedicated_endpoint = True # @param {type:\"boolean\"}\n",
"\n",
"# @markdown Click \"Show Code\" to see more details."
"from vertexai import model_garden"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "upYRiGtP_-iN"
},
"source": [
"## Deploy model"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "H2WC_0hXDVXc"
},
"source": [
"### Choose model variant"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "u41zbNa2EoFq"
},
"source": [
"You can proceed with the default model variant or select a different one."
]
},
{
@@ -203,75 +228,118 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "cac4478ae098"
"id": "-fgC4NLSDkF7"
},
"outputs": [],
"source": [
"# @title Deploy with customized configs\n",
"model_version = \"blip-image-captioning-base\" # @param [\"blip-image-captioning-base\"] {isTemplate:true}\n",
"MODEL_NAME = f\"salesforce/blip-image-captioning-base@{model_version}\""
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "VRnUgU8LF3_i"
},
"source": [
"To see all deployable model variants available in Model Garden, use:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "-QLd-wshF6sB"
},
"outputs": [],
"source": [
"all_model_versions = model_garden.list_deployable_models(\n",
" model_filter=\"blip-image-captioning-base\", list_hf_models=False\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "N0UeFHa2GO63"
},
"source": [
"Once you've selected a model variant, initialize it:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "GZiV3trBBcA3"
},
"outputs": [],
"source": [
"model = model_garden.OpenModel(MODEL_NAME)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "-0cL378wFlvf"
},
"source": [
"### Check the Deployment Configuration\n",
"\n",
"# The pre-built serving docker image. It contains serving scripts and models.\n",
"SERVE_DOCKER_URI = \"us-docker.pkg.dev/deeplearning-platform-release/vertex-model-garden/pytorch-inference.cu125.0-4.ubuntu2204.py310:model-garden.pytorch-inference-0-4-gpu-release_20250708.04_p0\"\n",
"Use the `list_deploy_options()` method to view the verified deployment configurations for your selected model. This helps ensure you have sufficient resources (e.g., GPU quota) available to deploy it."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "zm73g7vFFm9N"
},
"outputs": [],
"source": [
"deploy_options = model.list_deploy_options(concise=True)\n",
"print(deploy_options)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "WjV499VsGwrD"
},
"source": [
"### Deploy the Model\n",
"\n",
"Now that you’ve reviewed the deployment options, use the `deploy()` method to serve the selected open model to a Vertex AI endpoint. Deployment time may vary depending on the model size and infrastructure requirements.\n",
"\n",
"def deploy_model(\n",
" model_name: str,\n",
" model_id: str,\n",
" task: str,\n",
" machine_type: str = \"g2-standard-8\",\n",
" accelerator_type: str = \"NVIDIA_L4\",\n",
" accelerator_count: int = 1,\n",
" use_dedicated_endpoint: bool = True,\n",
") -> Tuple[aiplatform.Model, aiplatform.Endpoint]:\n",
" model_name = \"blip-image-captioning\"\n",
" endpoint = aiplatform.Endpoint.create(\n",
" display_name=f\"{model_name}-endpoint\",\n",
" dedicated_endpoint_enabled=use_dedicated_endpoint,\n",
" )\n",
" serving_env = {\n",
" \"MODEL_ID\": model_id,\n",
" \"TASK\": task,\n",
" \"DEPLOY_SOURCE\": \"notebook\",\n",
" }\n",
" # If the model_id is a GCS path, use artifact_uri to pass it to serving docker.\n",
" artifact_uri = model_id if model_id.startswith(\"gs://\") else None\n",
" model = aiplatform.Model.upload(\n",
" display_name=model_name,\n",
" serving_container_image_uri=SERVE_DOCKER_URI,\n",
" serving_container_ports=[7080],\n",
" serving_container_predict_route=\"/predict\",\n",
" serving_container_health_route=\"/health\",\n",
" serving_container_environment_variables=serving_env,\n",
" artifact_uri=artifact_uri,\n",
" model_garden_source_model_name=\"publishers/salesforce/models/blip-image-captioning-base\",\n",
" )\n",
" model.deploy(\n",
" endpoint=endpoint,\n",
" machine_type=machine_type,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
" deploy_request_timeout=1800,\n",
" system_labels={\n",
" \"NOTEBOOK_NAME\": \"model_garden_pytorch_blip_image_captioning.ipynb\"\n",
" },\n",
" )\n",
" return model, endpoint\n",
"\n",
"\n",
"common_util.check_quota(\n",
" project_id=PROJECT_ID,\n",
" region=REGION,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
" is_for_training=False,\n",
")\n",
"\n",
"models[LABEL], endpoints[LABEL] = deploy_model(\n",
" model_name=common_util.get_job_name_with_datetime(prefix=\"blip-image-captioning\"),\n",
" model_id=MODEL_ID,\n",
" task=TASK,\n",
" machine_type=machine_type,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
"> **Note**: If the model requires accepting a license agreement (EULA), set the `accept_eula=True` flag in the deploy call. Set `use_dedicated_endpoint` to False if you don't want to use [dedicated endpoint](https://cloud.google.com/vertex-ai/docs/general/deployment#create-dedicated-endpoint)."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "wX1itVTvXdEP"
},
"outputs": [],
"source": [
"use_dedicated_endpoint = True"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "MRmPFEPoGzsB"
},
"outputs": [],
"source": [
"endpoint = model.deploy(\n",
" accept_eula=True,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
")"
]
@@ -279,10 +347,44 @@
{
"cell_type": "markdown",
"metadata": {
"id": "80b3fd2ace09"
"id": "PHBtn8DQp-ID"
},
"source": [
"NOTE: The model weights will be downloaded after the deployment succeeds. Thus additional 5 minutes of waiting time is needed **after** the above model deployment step succeeds and before you run the next step below. Otherwise you might see a `ServiceUnavailable: 503 502:Bad Gateway` error when you send requests to the endpoint."
"Alternatively, you can select one of the verified deployment configurations listed above."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "ADsJG8JYqI6c"
},
"outputs": [],
"source": [
"endpoint = model.deploy(\n",
" accept_eula=True,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
" serving_container_image_uri=\"us-docker.pkg.dev/deeplearning-platform-release/vertex-model-garden/pytorch-inference.cu125.0-4.ubuntu2204.py310:model-garden.pytorch-inference-0-4-gpu-release_20250708.04_p0\",\n",
" machine_type=\"n1-standard-8\",\n",
" accelerator_type=\"NVIDIA_TESLA_T4\",\n",
" accelerator_count=1,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "kqSUK2CwsImi"
},
"source": [
"To further customize your deployment, you can configure:\n",
"\n",
"- **Compute Resources**: Machine type, replica count (min/max), accelerator type and quantity.\n",
"- **Infrastructure**: Use Spot VMs, reservation affinity, or dedicated endpoints.\n",
"- **Serving Container**: Customize container image, ports, health checks, and environment variables.\n",
"\n",
"See the [Model Garden SDK README](https://github.com/googleapis/python-aiplatform/blob/main/vertexai/model_garden/README.md) for advanced configuration options."
]
},
{
@@ -296,19 +398,31 @@
"source": [
"# @title Predict\n",
"\n",
"import importlib\n",
"\n",
"from IPython.display import Image, display\n",
"\n",
"# Import the necessary packages.\n",
"! rm -rf vertex-ai-samples && git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"! cd vertex-ai-samples\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"INPUT_IMAGE = \"http://images.cocodataset.org/val2017/000000039769.jpg\" # @param {type: \"string\"}\n",
"image = common_util.download_image(INPUT_IMAGE)\n",
"display(image)\n",
"\n",
"\n",
"display(Image(INPUT_IMAGE))\n",
"\n",
"instances = [\n",
" {\"image\": common_util.image_to_base64(image), \"text\": \"unused\"},\n",
"]\n",
"\n",
"preds = (\n",
" endpoints[LABEL]\n",
" .predict(instances=instances, use_dedicated_endpoint=use_dedicated_endpoint)\n",
" .predictions\n",
")\n",
"preds = endpoint.predict(\n",
" instances=instances, use_dedicated_endpoint=use_dedicated_endpoint\n",
").predictions\n",
"print(preds)"
]
},
@@ -322,16 +436,10 @@
"outputs": [],
"source": [
"# @title Clean up resources\n",
"# @markdown Delete the experiment models and endpoints to recycle the resources\n",
"# @markdown and avoid unnecessary continuous charges that may incur.\n",
"# @markdown Delete the endpoint.\n",
"\n",
"# Undeploy model and delete endpoint.\n",
"for endpoint in endpoints.values():\n",
" endpoint.delete(force=True)\n",
"\n",
"# Delete models.\n",
"for model in models.values():\n",
" model.delete()"
"if endpoint:\n",
" endpoint.delete(force=True)"
]
}
],
@@ -120,7 +120,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"\n",
@@ -131,7 +131,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"LABEL = \"vllm_gpu\"\n",
@@ -258,7 +258,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -267,6 +266,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -4,11 +4,12 @@
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "7d9bbf86da5e"
"cellView": "form",
"id": "ebvioouiVV_9"
},
"outputs": [],
"source": [
"# Copyright 2024 Google LLC\n",
"# Copyright 2025 Google LLC\n",
"#\n",
"# Licensed under the Apache License, Version 2.0 (the \"License\");\n",
"# you may not use this file except in compliance with the License.\n",
@@ -31,19 +32,18 @@
"source": [
"# Vertex AI Model Garden - ControlNet\n",
"\n",
"<table align=\"left\">\n",
" <td>\n",
"<table><tbody><tr>\n",
" <td style=\"text-align: center\">\n",
" <a href=\"https://console.cloud.google.com/vertex-ai/colab/import/https:%2F%2Fraw.githubusercontent.com%2FGoogleCloudPlatform%2Fvertex-ai-samples%2Fmain%2Fnotebooks%2Fcommunity%2Fmodel_garden%2Fmodel_garden_pytorch_controlnet.ipynb\">\n",
" <img alt=\"Google Cloud Colab Enterprise logo\" src=\"https://lh3.googleusercontent.com/JmcxdQi-qOpctIvWKgPtrzZdJJK-J3sWE1RsfjZNwshCFgE_9fULcNpuXYTilIR2hjwN\" width=\"32px\"><br> Run in Colab Enterprise\n",
" </a>\n",
" </td>\n",
" <td>\n",
" <td style=\"text-align: center\">\n",
" <a href=\"https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/model_garden/model_garden_pytorch_controlnet.ipynb\">\n",
" <img src=\"https://github.githubassets.com/assets/GitHub-Mark-ea2971cee799.png\" alt=\"GitHub logo\"><br>\n",
" View on GitHub\n",
" <img alt=\"GitHub logo\" src=\"https://github.githubassets.com/assets/GitHub-Mark-ea2971cee799.png\" width=\"32px\"><br> View on GitHub\n",
" </a>\n",
" </td>\n",
"</table>"
"</tr></tbody></table>"
]
},
{
@@ -54,31 +54,43 @@
"source": [
"## Overview\n",
"\n",
"This notebook demonstrates deploying the [ControlNet](https://huggingface.co/lllyasviel/ControlNet) model on Vertex AI for online prediction.\n",
"This notebook demonstrates how to deploy a **Control-Net** open model on Google Cloud Vertex AI.\n",
"\n",
"### Objective\n",
"### Objectives\n",
"\n",
"- Upload the model to [Model Registry](https://cloud.google.com/vertex-ai/docs/model-registry/introduction).\n",
"- Deploy the model on [Endpoint](https://cloud.google.com/vertex-ai/docs/predictions/using-private-endpoints).\n",
"- Run online predictions for text-guided-image-to-image.\n",
"- Deploy Control-Net using containerized backends like [vLLM](https://github.com/vllm-project/vllm) on GPU.\n",
"- Use the deployed model to serve chat completion requests for both text and multimodal inputs.\n",
"\n",
"### File a Bug\n",
"\n",
"If you encounter issues with this notebook, report them on [GitHub](https://github.com/GoogleCloudPlatform/vertex-ai-samples/issues/new).\n",
"\n",
"### Costs\n",
"\n",
"This tutorial uses billable components of Google Cloud:\n",
"\n",
"* Vertex AI\n",
"* Cloud Storage\n",
"- Vertex AI\n",
"- Cloud Storage\n",
"\n",
"Learn about [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing) and [Cloud Storage pricing](https://cloud.google.com/storage/pricing), and use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to generate a cost estimate based on your projected usage."
"Refer to the [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing) and [Cloud Storage pricing](https://cloud.google.com/storage/pricing) pages for more information. Use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to estimate your projected costs."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "264c07757582"
"id": "jeYw-Czg-DFy"
},
"source": [
"## Run the notebook"
"## Get Started"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "KgyhGvEzBDkj"
},
"source": [
"### Install Vertex AI SDK and other required packages"
]
},
{
@@ -86,147 +98,52 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "2707b02ef5df"
"id": "iCacdLqG-IsH"
},
"outputs": [],
"source": [
"# @title Setup Google Cloud project\n",
"%pip install --upgrade --force-reinstall --quiet 'google-cloud-aiplatform>=1.106.0' 'openai' 'google-auth==2.27.0' 'requests==2.32.3'"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "HUKCrpBy-3yf"
},
"source": [
"### Authenticate the Notebook Environment (Colab only)\n",
"\n",
"# @markdown 1. [Make sure that billing is enabled for your project](https://cloud.google.com/billing/docs/how-to/modify-project).\n",
"\n",
"# @markdown 2. [Optional] [Create a Cloud Storage bucket](https://cloud.google.com/storage/docs/creating-buckets) for storing experiment outputs. Set the BUCKET_URI for the experiment environment. The specified Cloud Storage bucket (`BUCKET_URI`) should be located in the same region as where the notebook was launched. Note that a multi-region bucket (eg. \"us\") is not considered a match for a single region covered by the multi-region range (eg. \"us-central1\"). If not set, a unique GCS bucket will be created instead.\n",
"\n",
"import base64\n",
"import os\n",
"If you're running this notebook in Google Colab, run the following cell to authenticate."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "JXwCT1kn-3Gu"
},
"outputs": [],
"source": [
"import sys\n",
"import uuid\n",
"from datetime import datetime\n",
"from io import BytesIO\n",
"\n",
"import cv2\n",
"import numpy as np\n",
"import requests\n",
"from google.cloud import aiplatform\n",
"from PIL import Image\n",
"\n",
"# Get the default cloud project id.\n",
"PROJECT_ID = os.environ[\"GOOGLE_CLOUD_PROJECT\"]\n",
"\n",
"# Get the default region for launching jobs.\n",
"REGION = os.environ[\"GOOGLE_CLOUD_REGION\"]\n",
"\n",
"# Enable the Vertex AI API and Compute Engine API, if not already.\n",
"! gcloud services enable aiplatform.googleapis.com compute.googleapis.com\n",
"\n",
"# Cloud Storage bucket for storing the experiment artifacts.\n",
"# A unique GCS bucket will be created for the purpose of this notebook. If you\n",
"# prefer using your own GCS bucket, please change the value yourself below.\n",
"now = datetime.now().strftime(\"%Y%m%d%H%M%S\")\n",
"BUCKET_URI = \"gs://\" # @param {type: \"string\"}\n",
"BUCKET_NAME = \"/\".join(BUCKET_URI.split(\"/\")[:3])\n",
"assert BUCKET_URI.startswith(\"gs://\"), \"BUCKET_URI must start with `gs://`.\"\n",
"\n",
"# Create a unique GCS bucket for this notebook, if not specified by the user.\n",
"if BUCKET_URI is None or BUCKET_URI.strip() == \"\" or BUCKET_URI == \"gs://\":\n",
" BUCKET_URI = f\"gs://{PROJECT_ID}-tmp-{now}-{str(uuid.uuid4())[:4]}\"\n",
" BUCKET_NAME = \"/\".join(BUCKET_URI.split(\"/\")[:3])\n",
" ! gsutil mb -l {REGION} {BUCKET_URI}\n",
"else:\n",
" shell_output = ! gsutil ls -Lb {BUCKET_NAME} | grep \"Location constraint:\" | sed \"s/Location constraint://\"\n",
" bucket_region = shell_output[0].strip().lower()\n",
" if bucket_region != REGION:\n",
" raise ValueError(\n",
" \"Bucket region %s is different from notebook region %s.\"\n",
" % (bucket_region, REGION)\n",
" )\n",
"\n",
"print(f\"Using this GCS Bucket: {BUCKET_URI}\")\n",
"\n",
"# Set up the default SERVICE_ACCOUNT.\n",
"SERVICE_ACCOUNT = None\n",
"shell_output = ! gcloud projects describe $PROJECT_ID\n",
"project_number = shell_output[-1].split(\":\")[1].strip().replace(\"'\", \"\")\n",
"SERVICE_ACCOUNT = f\"{project_number}-compute@developer.gserviceaccount.com\"\n",
"\n",
"print(\"Using this default Service Account:\", SERVICE_ACCOUNT)\n",
"\n",
"# Provision permissions to the SERVICE_ACCOUNT with the GCS bucket\n",
"! gsutil iam ch serviceAccount:{SERVICE_ACCOUNT}:roles/storage.admin $BUCKET_NAME\n",
"\n",
"if \"google.colab\" in sys.modules:\n",
" from google.colab import auth\n",
"\n",
" auth.authenticate_user(project_id=PROJECT_ID)\n",
" auth.authenticate_user()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "AcW2nwB8-7yC"
},
"source": [
"### Set Google Cloud Project Information\n",
"\n",
"To get started with Vertex AI, ensure you have an existing Google Cloud project and that the [Vertex AI API is enabled](https://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com).\n",
"\n",
"# The pre-built serving docker image. It contains serving scripts and models.\n",
"SERVE_DOCKER_URI = \"us-docker.pkg.dev/vertex-ai/vertex-vision-model-garden-dockers/pytorch-diffusers-serve-opt:20240605_1400_RC00\"\n",
"\n",
"\n",
"# Define common functions.\n",
"def download_image(url):\n",
" response = requests.get(url)\n",
" return Image.open(BytesIO(response.content))\n",
"\n",
"\n",
"def image_to_base64(image, format=\"JPEG\"):\n",
" buffer = BytesIO()\n",
" image.save(buffer, format=format)\n",
" image_str = base64.b64encode(buffer.getvalue()).decode(\"utf-8\")\n",
" return image_str\n",
"\n",
"\n",
"def base64_to_image(image_str):\n",
" image = Image.open(BytesIO(base64.b64decode(image_str)))\n",
" return image\n",
"\n",
"\n",
"def image_grid(imgs, rows=2, cols=2):\n",
" w, h = imgs[0].size\n",
" grid = Image.new(\"RGB\", size=(cols * w, rows * h), color=(255, 255, 255))\n",
" for i, img in enumerate(imgs):\n",
" grid.paste(img, box=(i % cols * w + 10 * i, i // cols * h))\n",
" return grid\n",
"\n",
"\n",
"def canny(image):\n",
" image = np.array(image)\n",
" image = cv2.Canny(image, 100, 200)\n",
" image = image[:, :, None]\n",
" image = np.concatenate([image, image, image], axis=2)\n",
" image = Image.fromarray(image)\n",
" return image\n",
"\n",
"\n",
"def deploy_model(model_id, task):\n",
" model_name = \"controlnet\"\n",
" endpoint = aiplatform.Endpoint.create(display_name=f\"{model_name}-endpoint\")\n",
" serving_env = {\n",
" \"MODEL_ID\": model_id,\n",
" \"TASK\": task,\n",
" \"DEPLOY_SOURCE\": \"notebook\",\n",
" }\n",
" model = aiplatform.Model.upload(\n",
" display_name=model_name,\n",
" serving_container_image_uri=SERVE_DOCKER_URI,\n",
" serving_container_ports=[7080],\n",
" serving_container_predict_route=\"/predictions/diffusers_serving\",\n",
" serving_container_health_route=\"/ping\",\n",
" serving_container_environment_variables=serving_env,\n",
" model_garden_source_model_name=\"publishers/lllyasviel/models/control-net\"\n",
" )\n",
" model.deploy(\n",
" endpoint=endpoint,\n",
" machine_type=\"g2-standard-8\",\n",
" accelerator_type=\"NVIDIA_L4\",\n",
" accelerator_count=1,\n",
" deploy_request_timeout=1800,\n",
" service_account=SERVICE_ACCOUNT,\n",
" system_labels={\n",
" \"NOTEBOOK_NAME\": \"model_garden_pytorch_controlnet.ipynb\"\n",
" },\n",
" )\n",
" return model, endpoint"
"See the guide on [setting up your project and development environment](https://cloud.google.com/vertex-ai/docs/start/cloud-environment). Also confirm that [billing is enabled](https://cloud.google.com/billing/docs/how-to/modify-project).\n"
]
},
{
@@ -234,21 +151,249 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "a9bbe0f7c237"
"id": "eIVLp0oE--k-"
},
"outputs": [],
"source": [
"# @title Upload and deploy model\n",
"# Use the environment variable if the user doesn't provide Project ID.\n",
"import os\n",
"\n",
"# @markdown This step deploys the pre-trained [lllyasviel/sd-controlnet-canny](https://huggingface.co/lllyasviel/sd-controlnet-canny) model for the text-guided image-to-image task.\n",
"import vertexai\n",
"\n",
"# @markdown The model deployment step will take ~15 minutes to complete.\n",
"PROJECT_ID = \"[your-project-id]\" # @param {type: \"string\", placeholder: \"[your-project-id]\", isTemplate: true}\n",
"\n",
"model, endpoint = deploy_model(\n",
" model_id=\"lllyasviel/sd-controlnet-canny\", task=\"controlnet\"\n",
"if not PROJECT_ID or PROJECT_ID == \"[your-project-id]\":\n",
" PROJECT_ID = str(os.environ.get(\"GOOGLE_CLOUD_PROJECT\"))\n",
"\n",
"REGION = os.environ.get(\"GOOGLE_CLOUD_REGION\", \"us-central1\")\n",
"\n",
"vertexai.init(project=PROJECT_ID, location=REGION)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "Q0CXrvcZH_aw"
},
"source": [
"### Import libraries"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "3G2UXB82ICs6"
},
"outputs": [],
"source": [
"from vertexai import model_garden"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "upYRiGtP_-iN"
},
"source": [
"## Deploy model"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "H2WC_0hXDVXc"
},
"source": [
"### Choose model variant"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "u41zbNa2EoFq"
},
"source": [
"You can proceed with the default model variant or select a different one."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "-fgC4NLSDkF7"
},
"outputs": [],
"source": [
"model_version = \"sd-controlnet-canny\" # @param [\"sd-controlnet-canny\"] {isTemplate:true}\n",
"MODEL_NAME = f\"lllyasviel/control-net@{model_version}\""
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "VRnUgU8LF3_i"
},
"source": [
"To see all deployable model variants available in Model Garden, use:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "-QLd-wshF6sB"
},
"outputs": [],
"source": [
"all_model_versions = model_garden.list_deployable_models(\n",
" model_filter=\"control-net\", list_hf_models=False\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "N0UeFHa2GO63"
},
"source": [
"Once you've selected a model variant, initialize it:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "GZiV3trBBcA3"
},
"outputs": [],
"source": [
"model = model_garden.OpenModel(MODEL_NAME)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "-0cL378wFlvf"
},
"source": [
"### Check the Deployment Configuration\n",
"\n",
"Use the `list_deploy_options()` method to view the verified deployment configurations for your selected model. This helps ensure you have sufficient resources (e.g., GPU quota) available to deploy it."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "zm73g7vFFm9N"
},
"outputs": [],
"source": [
"deploy_options = model.list_deploy_options(concise=True)\n",
"print(deploy_options)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "WjV499VsGwrD"
},
"source": [
"### Deploy the Model\n",
"\n",
"Now that you’ve reviewed the deployment options, use the `deploy()` method to serve the selected open model to a Vertex AI endpoint. Deployment time may vary depending on the model size and infrastructure requirements.\n",
"\n",
"> **Note**: If the model requires accepting a license agreement (EULA), set the `accept_eula=True` flag in the deploy call. Set `use_dedicated_endpoint` to False if you don't want to use [dedicated endpoint](https://cloud.google.com/vertex-ai/docs/general/deployment#create-dedicated-endpoint)."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "wX1itVTvXdEP"
},
"outputs": [],
"source": [
"use_dedicated_endpoint = True"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "MRmPFEPoGzsB"
},
"outputs": [],
"source": [
"endpoint = model.deploy(\n",
" accept_eula=True,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "PHBtn8DQp-ID"
},
"source": [
"Alternatively, you can select one of the verified deployment configurations listed above."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "ADsJG8JYqI6c"
},
"outputs": [],
"source": [
"endpoint = model.deploy(\n",
" accept_eula=True,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
" serving_container_image_uri=\"us-docker.pkg.dev/vertex-ai/vertex-vision-model-garden-dockers/pytorch-diffusers-serve-opt:20240605_1400_RC00\",\n",
" machine_type=\"g2-standard-8\",\n",
" accelerator_type=\"NVIDIA_L4\",\n",
" accelerator_count=1,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "kqSUK2CwsImi"
},
"source": [
"To further customize your deployment, you can configure:\n",
"\n",
"- **Compute Resources**: Machine type, replica count (min/max), accelerator type and quantity.\n",
"- **Infrastructure**: Use Spot VMs, reservation affinity, or dedicated endpoints.\n",
"- **Serving Container**: Customize container image, ports, health checks, and environment variables.\n",
"\n",
"See the [Model Garden SDK README](https://github.com/googleapis/python-aiplatform/blob/main/vertexai/model_garden/README.md) for advanced configuration options."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "Pt1O-tITu5xL"
},
"outputs": [],
"source": [
"! pip install --upgrade 'opencv-python'"
]
},
{
"cell_type": "code",
"execution_count": null,
@@ -266,25 +411,49 @@
"\n",
"# @markdown You may adjust the parameters below to achieve best image quality.\n",
"\n",
"import importlib\n",
"\n",
"import cv2\n",
"import numpy as np\n",
"from PIL import Image\n",
"\n",
"# Import the necessary packages.\n",
"! rm -rf vertex-ai-samples && git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"! cd vertex-ai-samples\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"\n",
"def canny(image):\n",
" image = np.array(image)\n",
" image = cv2.Canny(image, 100, 200)\n",
" image = image[:, :, None]\n",
" image = np.concatenate([image, image, image], axis=2)\n",
" image = Image.fromarray(image)\n",
" return image\n",
"\n",
"\n",
"prompt = \"bird\" # @param {type: \"string\"}\n",
"image = \"https://huggingface.co/takuma104/controlnet_dev/resolve/main/gen_compare/output_images/diffusers/output_bird_canny_1.png\" # @param {type: \"string\"}\n",
"num_inference_steps = 25 # @param {type:\"number\"}\n",
"\n",
"init_image = download_image(image)\n",
"init_image = common_util.download_image(image)\n",
"canny_image = canny(init_image)\n",
"\n",
"instances = [\n",
" {\n",
" \"prompt\": prompt,\n",
" \"image\": image_to_base64(canny_image),\n",
" \"image\": common_util.image_to_base64(canny_image),\n",
" \"num_inference_steps\": num_inference_steps,\n",
" },\n",
"]\n",
"response = endpoint.predict(instances=instances)\n",
"images = [base64_to_image(image) for image in response.predictions]\n",
"images = [common_util.base64_to_image(image) for image in response.predictions]\n",
"new_image = images[0]\n",
"\n",
"image_grid([init_image, canny_image, new_image], rows=1, cols=3)"
"common_util.image_grid([init_image, canny_image, new_image], rows=1, cols=3)"
]
},
{
@@ -298,19 +467,10 @@
"source": [
"# @title Clean up resources\n",
"\n",
"# @markdown Delete the experiment models and endpoints to recycle the resources\n",
"# @markdown and avoid unnecessary continouous charges that may incur.\n",
"# @markdown Delete the endpoint.\n",
"\n",
"# Undeploy model and delete endpoint.\n",
"endpoint.delete(force=True)\n",
"\n",
"# Delete models.\n",
"model.delete()\n",
"\n",
"# Delete bucket.\n",
"delete_bucket = False # @param {type:\"boolean\"}\n",
"if delete_bucket:\n",
" ! gsutil -m rm -r $BUCKET_NAME"
"if endpoint:\n",
" endpoint.delete(force=True)"
]
}
],
@@ -124,7 +124,7 @@
"from IPython.display import Audio\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -89,24 +89,6 @@
"## Before you begin"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "ax7zWynUDcjk"
},
"outputs": [],
"source": [
"# @title Request for quota\n",
"\n",
"# @markdown To deploy the largest variants of the DeepSeek models, you need 1 host of 8 x H200 machine, or 2 hosts of 8 x H100 machines (which gives a total of 16 x H100s). Check that you have sufficient quota:\n",
"# @markdown - For Spot VM quota, check [`CustomModelServingPreemptibleH100GPUsPerProjectPerRegion`](https://console.cloud.google.com/iam-admin/quotas?metric=aiplatform.googleapis.com%2Fcustom_model_serving_preemptible_nvidia_h100_gpus). H200 GPUs are currently not available in Spot VM quota.\n",
"# @markdown - For regular VM quota, check [`CustomModelServingH200GPUsPerProjectPerRegion`](https://console.cloud.google.com/iam-admin/quotas?metric=aiplatform.googleapis.com%2Fcustom_model_serving_nvidia_h200_gpus) and [`CustomModelServingH100GPUsPerProjectPerRegion`](https://console.cloud.google.com/iam-admin/quotas?metric=aiplatform.googleapis.com%2Fcustom_model_serving_nvidia_h100_gpus).\n",
"#\n",
"# @markdown If you don't have sufficient quota, request for quota following the instructions at [\"Request a higher quota\"](https://cloud.google.com/docs/quota/view-manage#requesting_higher_quota)."
]
},
{
"cell_type": "code",
"execution_count": null,
@@ -120,16 +102,16 @@
"\n",
"# @markdown 1. [Make sure that billing is enabled for your project](https://cloud.google.com/billing/docs/how-to/modify-project).\n",
"\n",
"# @markdown 2. **[Optional]** Set region. If left unchanged, the region defaults to us-east4 for using H200 GPUs.\n",
"# @markdown 2. **[Optional]** Set region. If not set, the region will be set automatically according to Colab Enterprise environment.\n",
"\n",
"REGION = \"us-east4\" # @param {type:\"string\"}\n",
"REGION = \"\" # @param {type:\"string\"}\n",
"\n",
"# @markdown 3. If you want to run predictions with H100 GPUs, we recommend using the regions listed below. **NOTE:** Make sure you have associated quota in selected regions. Click the links to see your current quota for Spot VM H100 GPUs: [`CustomModelServingPreemptibleH100GPUsPerProjectPerRegion`](https://console.cloud.google.com/iam-admin/quotas?metric=aiplatform.googleapis.com%2Fcustom_model_serving_preemptible_nvidia_h100_gpus) and regular VM H100s: [`CustomModelServingH100GPUsPerProjectPerRegion`](https://console.cloud.google.com/iam-admin/quotas?metric=aiplatform.googleapis.com%2Fcustom_model_serving_nvidia_h100_gpus)..\n",
"# @markdown 3. If you want to run predictions with H100 GPUs or H200 GPUs, we recommend using the regions listed below. **NOTE:** Make sure you have associated quota in selected regions. Click the links to see your current quota for H100s: [`CustomModelServingH100GPUsPerProjectPerRegion`](https://console.cloud.google.com/iam-admin/quotas?metric=aiplatform.googleapis.com%2Fcustom_model_serving_nvidia_h100_gpus) and H200s: [`CustomModelServingH200GPUsPerProjectPerRegion`](https://console.cloud.google.com/iam-admin/quotas?metric=aiplatform.googleapis.com%2Fcustom_model_serving_nvidia_h200_gpus). You can request for quota following the instructions at [\"Request a higher quota\"](https://cloud.google.com/docs/quota/view-manage#requesting_higher_quota).\n",
"\n",
"# @markdown > | Machine Type | Accelerator Type | Recommended Regions |\n",
"# @markdown | Machine Type | Accelerator Type | Recommended Regions |\n",
"# @markdown | ----------- | ----------- | ----------- |\n",
"# @markdown | a3-highgpu-8g (Spot VM) | 8 NVIDIA_H100_80GB | us-central1, europe-west4, asia-southeast1 |\n",
"# @markdown | a3-highgpu-8g (regular VM) | 8 NVIDIA_H100_80GB | us-central1, europe-west4, us-west1, asia-southeast1 |\n",
"# @markdown | a3-highgpu-8g | 8 NVIDIA_H100_80GB | asia-southeast1, europe-west4, us-central1, us-east5, us-west1 |\n",
"# @markdown | a3-ultragpu-8g | 8 NVIDIA_H200_141GB | asia-south2, us-south1 |\n",
"\n",
"# Upgrade Vertex AI SDK.\n",
"! pip3 install --upgrade --quiet 'google-cloud-aiplatform==1.103.0'\n",
@@ -151,38 +133,9 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"\n",
"def check_quota(\n",
" project_id: str,\n",
" region: str,\n",
" resource_id: str,\n",
" accelerator_count: int,\n",
"):\n",
" \"\"\"Checks if the project and the region has the required quota.\"\"\"\n",
" quota = common_util.get_quota(project_id, region, resource_id)\n",
" quota_request_instruction = (\n",
" \"Either use \"\n",
" \"a different region or request additional quota. Follow \"\n",
" \"instructions here \"\n",
" \"https://cloud.google.com/docs/quotas/view-manage#requesting_higher_quota\"\n",
" \" to check quota in a region or request additional quota for \"\n",
" \"your project.\"\n",
" )\n",
" if quota == -1:\n",
" raise ValueError(\n",
" f\"Quota not found for: {resource_id} in {region}.\"\n",
" f\" {quota_request_instruction}\"\n",
" )\n",
" if quota < accelerator_count:\n",
" raise ValueError(\n",
" f\"Quota not enough for {resource_id} in {region}: {quota} <\"\n",
" f\" {accelerator_count}. {quota_request_instruction}\"\n",
" )\n",
"\n",
"\n",
"LABEL = \"vllm_gpu\"\n",
"models, endpoints = {}, {}\n",
"\n",
@@ -276,23 +229,17 @@
"if accelerator_type == \"NVIDIA_H200_141GB\":\n",
" machine_type = \"a3-ultragpu-8g\"\n",
" multihost_gpu_node_count = 1\n",
" if is_spot:\n",
" raise ValueError(\"H200 GPUs are currently not available in Spot VM quota.\")\n",
" else:\n",
" resource_id = \"custom_model_serving_nvidia_h200_gpus\"\n",
"else:\n",
" machine_type = \"a3-highgpu-8g\"\n",
" multihost_gpu_node_count = 2\n",
" if is_spot:\n",
" resource_id = \"custom_model_serving_preemptible_nvidia_h100_gpus\"\n",
" else:\n",
" resource_id = \"custom_model_serving_nvidia_h100_gpus\"\n",
"\n",
"check_quota(\n",
"common_util.check_quota(\n",
" project_id=PROJECT_ID,\n",
" region=REGION,\n",
" resource_id=resource_id,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=int(accelerator_count * multihost_gpu_node_count),\n",
" is_for_training=False,\n",
" is_spot=is_spot,\n",
")\n",
"\n",
"if accelerator_type == \"NVIDIA_H200_141GB\":\n",
@@ -425,7 +372,6 @@
" f\"--tensor-parallel-size={int(accelerator_count * multihost_gpu_node_count / pipeline_parallel_size)}\",\n",
" f\"--pipeline-parallel-size={pipeline_parallel_size}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--kv-cache-dtype={kv_cache_dtype}\",\n",
@@ -438,6 +384,9 @@
" if multihost_gpu_node_count > 1:\n",
" vllm_args = [\"/vllm-workspace/ray_launcher.sh\"] + vllm_args\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -819,23 +768,17 @@
"if accelerator_type == \"NVIDIA_H200_141GB\":\n",
" machine_type = \"a3-ultragpu-8g\"\n",
" multihost_gpu_node_count = 1\n",
" if is_spot:\n",
" raise ValueError(\"H200 GPUs are currently not available in Spot VM quota.\")\n",
" else:\n",
" resource_id = \"custom_model_serving_nvidia_h200_gpus\"\n",
"else:\n",
" machine_type = \"a3-highgpu-8g\"\n",
" multihost_gpu_node_count = 2\n",
" if is_spot:\n",
" resource_id = \"custom_model_serving_preemptible_nvidia_h100_gpus\"\n",
" else:\n",
" resource_id = \"custom_model_serving_nvidia_h100_gpus\"\n",
"\n",
"check_quota(\n",
"common_util.check_quota(\n",
" project_id=PROJECT_ID,\n",
" region=REGION,\n",
" resource_id=resource_id,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=int(accelerator_count * multihost_gpu_node_count),\n",
" is_for_training=False,\n",
" is_spot=is_spot,\n",
")\n",
"\n",
"# @markdown The maximum context length 163840 is supported in the following configurations.\n",
@@ -1332,15 +1275,16 @@
"if trtllm_accelerator_type == \"NVIDIA_H200_141GB\":\n",
" machine_type = \"a3-ultragpu-8g\"\n",
" multihost_gpu_node_count = 1\n",
" resource_id = \"custom_model_serving_nvidia_h200_gpus\"\n",
"else:\n",
" raise ValueError(\"Only NVIDIA_H200_141GB is supported for DeepSeek-R1.\")\n",
"\n",
"check_quota(\n",
"common_util.check_quota(\n",
" project_id=PROJECT_ID,\n",
" region=trtllm_region,\n",
" resource_id=resource_id,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=int(accelerator_count * multihost_gpu_node_count),\n",
" is_for_training=False,\n",
" is_spot=is_spot,\n",
")\n",
"\n",
"# 18K context length. This is the maximum supported by the current version of TensorRT-LLM on DeepSeek V3/R1 models.\n",
@@ -141,38 +141,10 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"\n",
"def check_quota(\n",
" project_id: str,\n",
" region: str,\n",
" resource_id: str,\n",
" accelerator_count: int,\n",
"):\n",
" \"\"\"Checks if the project and the region has the required quota.\"\"\"\n",
" quota = common_util.get_quota(project_id, region, resource_id)\n",
" quota_request_instruction = (\n",
" \"Either use \"\n",
" \"a different region or request additional quota. Follow \"\n",
" \"instructions here \"\n",
" \"https://cloud.google.com/docs/quotas/view-manage#requesting_higher_quota\"\n",
" \" to check quota in a region or request additional quota for \"\n",
" \"your project.\"\n",
" )\n",
" if quota == -1:\n",
" raise ValueError(\n",
" f\"Quota not found for: {resource_id} in {region}.\"\n",
" f\" {quota_request_instruction}\"\n",
" )\n",
" if quota < accelerator_count:\n",
" raise ValueError(\n",
" f\"Quota not enough for {resource_id} in {region}: {quota} <\"\n",
" f\" {accelerator_count}. {quota_request_instruction}\"\n",
" )\n",
"\n",
"\n",
"LABEL = \"vllm_gpu\"\n",
"models, endpoints = {}, {}\n",
"\n",
@@ -254,15 +226,15 @@
"if accelerator_type == \"NVIDIA_GB200\":\n",
" accelerator_count = 4\n",
" machine_type = \"a4x-highgpu-4g\"\n",
" resource_id = \"custom_model_serving_nvidia_gb200_gpus\"\n",
"else:\n",
" raise ValueError(\"Sample deployment options are not available.\")\n",
"\n",
"check_quota(\n",
"common_util.check_quota(\n",
" project_id=PROJECT_ID,\n",
" region=REGION,\n",
" resource_id=resource_id,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
" is_for_training=False,\n",
")\n",
"\n",
"max_model_len = 131072\n",
@@ -337,7 +309,6 @@
" f\"--tensor-parallel-size={int(accelerator_count * multihost_gpu_node_count / pipeline_parallel_size)}\",\n",
" f\"--pipeline-parallel-size={pipeline_parallel_size}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--kv-cache-dtype={kv_cache_dtype}\",\n",
@@ -350,6 +321,9 @@
" if multihost_gpu_node_count > 1:\n",
" vllm_args = [\"/vllm-workspace/ray_launcher.sh\"] + vllm_args\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -169,7 +169,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -129,7 +129,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -136,7 +136,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -129,7 +129,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -120,7 +120,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"\n",
@@ -155,7 +155,7 @@
" custom_job as gca_custom_job_compat\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -728,7 +728,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -737,6 +736,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -0,0 +1,569 @@
{
"cells": [
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "SgQ6t5bqZVlH"
},
"outputs": [],
"source": [
"# Copyright 2025 Google LLC\n",
"#\n",
"# Licensed under the Apache License, Version 2.0 (the \"License\");\n",
"# you may not use this file except in compliance with the License.\n",
"# You may obtain a copy of the License at\n",
"#\n",
"# https://www.apache.org/licenses/LICENSE-2.0\n",
"#\n",
"# Unless required by applicable law or agreed to in writing, software\n",
"# distributed under the License is distributed on an \"AS IS\" BASIS,\n",
"# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.\n",
"# See the License for the specific language governing permissions and\n",
"# limitations under the License."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "99c1c3fc2ca5"
},
"source": [
"# Vertex AI Model Garden - GPT OSS (Deployment)\n",
"\n",
"<table><tbody><tr>\n",
" <td style=\"text-align: center\">\n",
" <a href=\"https://console.cloud.google.com/vertex-ai/notebooks/deploy-notebook?download_url=https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/main/notebooks/community/model_garden/model_garden_pytorch_gpt_oss_deployment.ipynb\">\n",
" <img alt=\"Workbench logo\" src=\"https://lh3.googleusercontent.com/UiNooY4LUgW_oTvpsNhPpQzsstV5W8F7rYgxgGBD85cWJoLmrOzhVs_ksK_vgx40SHs7jCqkTkCk=e14-rj-sc0xffffff-h130-w32\" width=\"32px\"><br> Run in Workbench\n",
" </a>\n",
" </td>\n",
" <td style=\"text-align: center\">\n",
" <a href=\"https://console.cloud.google.com/vertex-ai/colab/import/https:%2F%2Fraw.githubusercontent.com%2FGoogleCloudPlatform%2Fvertex-ai-samples%2Fmain%2Fnotebooks%2Fcommunity%2Fmodel_garden%2Fmodel_garden_pytorch_gpt_oss_deployment.ipynb\">\n",
" <img alt=\"Google Cloud Colab Enterprise logo\" src=\"https://lh3.googleusercontent.com/JmcxdQi-qOpctIvWKgPtrzZdJJK-J3sWE1RsfjZNwshCFgE_9fULcNpuXYTilIR2hjwN\" width=\"32px\"><br> Run in Colab Enterprise\n",
" </a>\n",
" </td>\n",
" <td style=\"text-align: center\">\n",
" <a href=\"https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/model_garden/model_garden_pytorch_gpt_oss_deployment.ipynb\">\n",
" <img alt=\"GitHub logo\" src=\"https://github.githubassets.com/assets/GitHub-Mark-ea2971cee799.png\" width=\"32px\"><br> View on GitHub\n",
" </a>\n",
" </td>\n",
"</tr></tbody></table>"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "3de7470326a2"
},
"source": [
"## Overview\n",
"\n",
"This notebook demonstrates how to deploy a **gpt-oss** open model on Google Cloud Vertex AI.\n",
"\n",
"### Objectives\n",
"\n",
"- Deploy gpt-oss using containerized backends like [vLLM](https://github.com/vllm-project/vllm) on GPU.\n",
"- Use the deployed model to serve chat completion requests for both text and multimodal inputs.\n",
"\n",
"### File a Bug\n",
"\n",
"If you encounter issues with this notebook, report them on [GitHub](https://github.com/GoogleCloudPlatform/vertex-ai-samples/issues/new).\n",
"\n",
"### Costs\n",
"\n",
"This tutorial uses billable components of Google Cloud:\n",
"\n",
"- Vertex AI\n",
"- Cloud Storage\n",
"\n",
"Refer to the [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing) and [Cloud Storage pricing](https://cloud.google.com/storage/pricing) pages for more information. Use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to estimate your projected costs."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "jeYw-Czg-DFy"
},
"source": [
"## Get Started"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "KgyhGvEzBDkj"
},
"source": [
"### Install Vertex AI SDK and other required packages"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "iCacdLqG-IsH"
},
"outputs": [],
"source": [
"%pip install --upgrade --force-reinstall --quiet 'google-cloud-aiplatform>=1.106.0' 'openai' 'google-auth==2.27.0' 'requests==2.32.3'"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "HUKCrpBy-3yf"
},
"source": [
"### Authenticate the Notebook Environment (Colab only)\n",
"\n",
"If you're running this notebook in Google Colab, run the following cell to authenticate."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "JXwCT1kn-3Gu"
},
"outputs": [],
"source": [
"import sys\n",
"\n",
"if \"google.colab\" in sys.modules:\n",
" from google.colab import auth\n",
"\n",
" auth.authenticate_user()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "AcW2nwB8-7yC"
},
"source": [
"### Set Google Cloud Project Information\n",
"\n",
"To get started with Vertex AI, ensure you have an existing Google Cloud project and that the [Vertex AI API is enabled](https://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com).\n",
"\n",
"See the guide on [setting up your project and development environment](https://cloud.google.com/vertex-ai/docs/start/cloud-environment). Also confirm that [billing is enabled](https://cloud.google.com/billing/docs/how-to/modify-project).\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "eIVLp0oE--k-"
},
"outputs": [],
"source": [
"# Use the environment variable if the user doesn't provide Project ID.\n",
"import os\n",
"\n",
"import vertexai\n",
"\n",
"PROJECT_ID = \"[your-project-id]\" # @param {type: \"string\", placeholder: \"[your-project-id]\", isTemplate: true}\n",
"\n",
"if not PROJECT_ID or PROJECT_ID == \"[your-project-id]\":\n",
" PROJECT_ID = str(os.environ.get(\"GOOGLE_CLOUD_PROJECT\"))\n",
"\n",
"REGION = os.environ.get(\"GOOGLE_CLOUD_REGION\", \"us-central1\")\n",
"\n",
"vertexai.init(project=PROJECT_ID, location=REGION)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "Q0CXrvcZH_aw"
},
"source": [
"### Import libraries"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "3G2UXB82ICs6"
},
"outputs": [],
"source": [
"from vertexai import model_garden"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "upYRiGtP_-iN"
},
"source": [
"## Deploy model"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "H2WC_0hXDVXc"
},
"source": [
"### Choose model variant"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "u41zbNa2EoFq"
},
"source": [
"You can proceed with the default model variant or select a different one."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "-fgC4NLSDkF7"
},
"outputs": [],
"source": [
"model_version = \"gpt-oss-120b\" # @param [\"gpt-oss-120b\", \"gpt-oss-20b\"] {isTemplate:true}\n",
"MODEL_NAME = f\"openai/gpt-oss@{model_version}\""
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "VRnUgU8LF3_i"
},
"source": [
"To see all deployable model variants available in Model Garden, use:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "-QLd-wshF6sB"
},
"outputs": [],
"source": [
"all_model_versions = model_garden.list_deployable_models(\n",
" model_filter=\"gpt-oss\", list_hf_models=False\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "N0UeFHa2GO63"
},
"source": [
"Once you've selected a model variant, initialize it:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "GZiV3trBBcA3"
},
"outputs": [],
"source": [
"model = model_garden.OpenModel(MODEL_NAME)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "-0cL378wFlvf"
},
"source": [
"### Check the Deployment Configuration\n",
"\n",
"Use the `list_deploy_options()` method to view the verified deployment configurations for your selected model. This helps ensure you have sufficient resources (e.g., GPU quota) available to deploy it."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "zm73g7vFFm9N"
},
"outputs": [],
"source": [
"deploy_options = model.list_deploy_options(concise=True)\n",
"print(deploy_options)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "WjV499VsGwrD"
},
"source": [
"### Deploy the Model\n",
"\n",
"Now that you’ve reviewed the deployment options, use the `deploy()` method to serve the selected open model to a Vertex AI endpoint. Deployment time may vary depending on the model size and infrastructure requirements.\n",
"\n",
"> **Note**: If the model requires accepting a license agreement (EULA), set the `accept_eula=True` flag in the deploy call. Set `use_dedicated_endpoint` to False if you don't want to use [dedicated endpoint](https://cloud.google.com/vertex-ai/docs/general/deployment#create-dedicated-endpoint)."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "wX1itVTvXdEP"
},
"outputs": [],
"source": [
"use_dedicated_endpoint = True"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "MRmPFEPoGzsB"
},
"outputs": [],
"source": [
"endpoint = model.deploy(\n",
" accept_eula=True,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "PHBtn8DQp-ID"
},
"source": [
"Alternatively, you can select one of the verified deployment configurations listed above."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "ADsJG8JYqI6c"
},
"outputs": [],
"source": [
"endpoint = model.deploy(\n",
" accept_eula=True,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
" serving_container_image_uri=\"us-docker.pkg.dev/vertex-ai/vertex-vision-model-garden-dockers/pytorch-vllm-serve:20250807_0916_RC01_maas\",\n",
" machine_type=\"a3-highgpu-2g\",\n",
" accelerator_type=\"NVIDIA_H100_80GB\",\n",
" accelerator_count=2,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "kqSUK2CwsImi"
},
"source": [
"To further customize your deployment, you can configure:\n",
"\n",
"- **Compute Resources**: Machine type, replica count (min/max), accelerator type and quantity.\n",
"- **Infrastructure**: Use Spot VMs, reservation affinity, or dedicated endpoints.\n",
"- **Serving Container**: Customize container image, ports, health checks, and environment variables.\n",
"\n",
"See the [Model Garden SDK README](https://github.com/googleapis/python-aiplatform/blob/main/vertexai/model_garden/README.md) for advanced configuration options."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "rDHsCOqvFYBi"
},
"outputs": [],
"source": [
"# @title Raw predict\n",
"\n",
"# @markdown Once deployment succeeds, you can send requests to the endpoint with text prompts. Sampling parameters supported by vLLM can be found [here](https://docs.vllm.ai/en/latest/dev/sampling_params.html).\n",
"\n",
"# @markdown Example:\n",
"\n",
"# @markdown ```\n",
"# @markdown Human: What is a car?\n",
"# @markdown Assistant: A car, or a motor car, is a road-connected human-transportation system used to move people or goods from one place to another. The term also encompasses a wide range of vehicles, including motorboats, trains, and aircrafts. Cars typically have four wheels, a cabin for passengers, and an engine or motor. They have been around since the early 19th century and are now one of the most popular forms of transportation, used for daily commuting, shopping, and other purposes.\n",
"# @markdown ```\n",
"# @markdown Additionally, you can moderate the generated text with Vertex AI. See [Moderate text documentation](https://cloud.google.com/natural-language/docs/moderating-text) for more details.\n",
"\n",
"# Loads an existing endpoint instance using the endpoint name:\n",
"# - Using `endpoint_name = endpoint.name` allows us to get the\n",
"# endpoint name of the endpoint `endpoint` created in the cell\n",
"# above.\n",
"# - Alternatively, you can set `endpoint_name = \"1234567890123456789\"` to load\n",
"# an existing endpoint with the ID 1234567890123456789.\n",
"# You may uncomment the code below to load an existing endpoint.\n",
"\n",
"# endpoint_name = \"\" # @param {type:\"string\"}\n",
"# aip_endpoint_name = (\n",
"# f\"projects/{PROJECT_ID}/locations/{REGION}/endpoints/{endpoint_name}\"\n",
"# )\n",
"# endpoint = aiplatform.Endpoint(aip_endpoint_name)\n",
"\n",
"prompt = \"What is a car?\" # @param {type: \"string\"}\n",
"# @markdown If you encounter an issue like `ServiceUnavailable: 503 Took too long to respond when processing`, you can reduce the maximum number of output tokens, by lowering `max_tokens`.\n",
"max_tokens = 50 # @param {type:\"integer\"}\n",
"temperature = 1.0 # @param {type:\"number\"}\n",
"top_p = 1.0 # @param {type:\"number\"}\n",
"top_k = 1 # @param {type:\"integer\"}\n",
"# @markdown Set `raw_response` to `True` to obtain the raw model output. Set `raw_response` to `False` to apply additional formatting in the structure of `\"Prompt:\\n{prompt.strip()}\\nOutput:\\n{output}\"`.\n",
"raw_response = False # @param {type:\"boolean\"}\n",
"\n",
"# Overrides parameters for inferences.\n",
"instances = [\n",
" {\n",
" \"prompt\": prompt,\n",
" \"max_tokens\": max_tokens,\n",
" \"temperature\": temperature,\n",
" \"top_p\": top_p,\n",
" \"top_k\": top_k,\n",
" \"raw_response\": raw_response,\n",
" },\n",
"]\n",
"response = endpoint.predict(\n",
" instances=instances, use_dedicated_endpoint=use_dedicated_endpoint\n",
")\n",
"\n",
"for prediction in response.predictions:\n",
" print(prediction)\n",
"\n",
"# @markdown Click \"Show Code\" to see more details."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "LSG9ITWTbTb7"
},
"outputs": [],
"source": [
"# @title Chat completion\n",
"\n",
"if use_dedicated_endpoint:\n",
" DEDICATED_ENDPOINT_DNS = endpoint.gca_resource.dedicated_endpoint_dns\n",
"ENDPOINT_RESOURCE_NAME = endpoint.resource_name\n",
"\n",
"# @title Chat Completions Inference\n",
"\n",
"# @markdown Once deployment succeeds, you can send requests to the endpoint using the OpenAI SDK.\n",
"\n",
"# @markdown First you will need to install the SDK and some auth-related dependencies.\n",
"\n",
"! pip install -qU openai google-auth requests\n",
"\n",
"# @markdown Next fill out some request parameters:\n",
"\n",
"user_message = \"How is your day going?\" # @param {type: \"string\"}\n",
"# @markdown If you encounter the issue like `ServiceUnavailable: 503 Took too long to respond when processing`, you can reduce the maximum number of output tokens, such as set `max_tokens` as 20.\n",
"max_tokens = 50 # @param {type: \"integer\"}\n",
"temperature = 1.0 # @param {type: \"number\"}\n",
"stream = False # @param {type: \"boolean\"}\n",
"\n",
"# @markdown Now we can send a request.\n",
"\n",
"import google.auth\n",
"import openai\n",
"\n",
"creds, project = google.auth.default()\n",
"auth_req = google.auth.transport.requests.Request()\n",
"creds.refresh(auth_req)\n",
"\n",
"BASE_URL = (\n",
" f\"https://{REGION}-aiplatform.googleapis.com/v1beta1/{ENDPOINT_RESOURCE_NAME}\"\n",
")\n",
"try:\n",
" if use_dedicated_endpoint:\n",
" BASE_URL = f\"https://{DEDICATED_ENDPOINT_DNS}/v1beta1/{ENDPOINT_RESOURCE_NAME}\"\n",
"except NameError:\n",
" pass\n",
"\n",
"client = openai.OpenAI(base_url=BASE_URL, api_key=creds.token)\n",
"\n",
"model_response = client.chat.completions.create(\n",
" model=\"\",\n",
" messages=[{\"role\": \"user\", \"content\": user_message}],\n",
" temperature=temperature,\n",
" max_tokens=max_tokens,\n",
" stream=stream,\n",
")\n",
"\n",
"if stream:\n",
" usage = None\n",
" contents = []\n",
" for chunk in model_response:\n",
" if chunk.usage is not None:\n",
" usage = chunk.usage\n",
" continue\n",
" print(chunk.choices[0].delta.content, end=\"\")\n",
" contents.append(chunk.choices[0].delta.content)\n",
" print(f\"\\n\\n{usage}\")\n",
"else:\n",
" print(model_response)\n",
"\n",
"# @markdown Click \"Show Code\" to see more details."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "tqtxJakIapIg"
},
"source": [
"## Clean up resources"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "kzgEmmd0aiUM"
},
"outputs": [],
"source": [
"# @title Delete the endpoint\n",
"\n",
"# @markdown Delete the endpoint.\n",
"\n",
"if endpoint:\n",
" endpoint.delete(force=True)"
]
}
],
"metadata": {
"colab": {
"name": "model_garden_pytorch_gpt_oss_deployment.ipynb",
"toc_visible": true
},
"kernelspec": {
"display_name": "Python 3",
"name": "python3"
}
},
"nbformat": 4,
"nbformat_minor": 0
}
@@ -129,7 +129,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -143,7 +143,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"\n",
@@ -130,7 +130,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"LABEL = \"instant_id_gpu\"\n",
@@ -312,7 +312,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -321,6 +320,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -107,7 +107,7 @@
"\n",
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"# Get the default cloud project id.\n",
@@ -131,7 +131,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"LABEL = \"diffusers_gpu\"\n",
@@ -136,7 +136,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"LABEL = \"vllm_gpu\"\n",
@@ -398,7 +398,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -407,6 +406,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -115,7 +115,7 @@
"from google.cloud import aiplatform, storage\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"# Get the default cloud project id.\n",
@@ -138,7 +138,7 @@
"from google.cloud import aiplatform\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"# Get the default cloud project id.\n",
@@ -164,7 +164,7 @@
"from google.cloud import aiplatform\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -510,7 +510,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -519,6 +518,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -159,7 +159,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -687,7 +687,9 @@
" .get(\"show-faster-deployment-option\")\n",
" == \"true\"\n",
" ]\n",
" if fast_deploy_config:\n",
" if len(fast_deploy_config) > 1:\n",
" fast_deploy_config = fast_deploy_config[1]\n",
" elif fast_deploy_config:\n",
" fast_deploy_config = fast_deploy_config[0]\n",
" else:\n",
" raise ValueError(\n",
@@ -1096,7 +1098,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -1105,6 +1106,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -141,7 +141,7 @@
"\n",
"# Import the necessary packages\n",
"! rm -rf vertex-ai-samples && git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"! cd vertex-ai-samples && git reset --hard c45f6a4f4d32e31a050f0e4ba52824b0caf4eda3\n",
"! cd vertex-ai-samples && git reset --hard 7ae13b346a72ee2a2dc8152dd40c6ddd72d6c810\n",
"\n",
"import datetime\n",
"import importlib\n",
@@ -154,7 +154,7 @@
" custom_job as gca_custom_job_compat\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -893,7 +893,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -902,6 +901,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -130,7 +130,7 @@
"models, endpoints = {}, {}\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"# Get the default cloud project id.\n",
@@ -142,7 +142,7 @@
"models, endpoints = {}, {}\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"# Get the default cloud project id.\n",
@@ -164,7 +164,7 @@
"from google.cloud import aiplatform\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -510,7 +510,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -519,6 +518,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -154,7 +154,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -1016,7 +1016,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -1025,6 +1024,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -59,34 +59,43 @@
"source": [
"## Overview\n",
"\n",
"This notebook demonstrates downloading, deploying, and serving prebuilt Llama 3.3 model with [vLLM](https://github.com/vllm-project/vllm) or [TensorRT-LLM](https://github.com/NVIDIA/TensorRT-LLM).\n",
"This notebook demonstrates how to deploy a **llama3.3** open model on Google Cloud Vertex AI.\n",
"\n",
"### Objectives\n",
"\n",
"### Objective\n",
"- Deploy llama3.3 using containerized backends like [vLLM](https://github.com/vllm-project/vllm) on GPU.\n",
"- Use the deployed model to serve chat completion requests for both text and multimodal inputs.\n",
"\n",
"- Deploy Llama 3.3 70B Instruct with vLLM (optionally with dynamic LoRA adapters) or TensorRT-LLM on GPU.\n",
"### File a Bug\n",
"\n",
"### File a bug\n",
"\n",
"File a bug on [GitHub](https://github.com/GoogleCloudPlatform/vertex-ai-samples/issues/new) if you encounter any issue with the notebook.\n",
"If you encounter issues with this notebook, report them on [GitHub](https://github.com/GoogleCloudPlatform/vertex-ai-samples/issues/new).\n",
"\n",
"### Costs\n",
"\n",
"This tutorial uses billable components of Google Cloud:\n",
"\n",
"* Vertex AI\n",
"* Cloud Storage\n",
"- Vertex AI\n",
"- Cloud Storage\n",
"\n",
"Learn about [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing), [Cloud Storage pricing](https://cloud.google.com/storage/pricing), and use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to generate a cost estimate based on your projected usage."
"Refer to the [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing) and [Cloud Storage pricing](https://cloud.google.com/storage/pricing) pages for more information. Use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to estimate your projected costs."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "264c07757582"
"id": "jeYw-Czg-DFy"
},
"source": [
"## Before you begin"
"## Get Started"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "KgyhGvEzBDkj"
},
"source": [
"### Install Vertex AI SDK and other required packages"
]
},
{
@@ -94,13 +103,22 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "ax7zWynUDcjk"
"id": "iCacdLqG-IsH"
},
"outputs": [],
"source": [
"# @title Request for quota\n",
"%pip install --upgrade --force-reinstall --quiet 'google-cloud-aiplatform>=1.106.0' 'openai' 'google-auth==2.27.0' 'requests==2.32.3'"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "HUKCrpBy-3yf"
},
"source": [
"### Authenticate the Notebook Environment (Colab only)\n",
"\n",
"# @markdown By default, the quota for H100 deployment `Custom model serving per region` is 0. You need to request for H100 quota following the instructions at [\"Request a higher quota\"](https://cloud.google.com/docs/quota/view-manage#requesting_higher_quota)."
"If you're running this notebook in Google Colab, run the following cell to authenticate."
]
},
{
@@ -108,102 +126,62 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "YXFGIp1l-qtT"
"id": "JXwCT1kn-3Gu"
},
"outputs": [],
"source": [
"# @title Setup Google Cloud project\n",
"import sys\n",
"\n",
"# @markdown 1. [Make sure that billing is enabled for your project](https://cloud.google.com/billing/docs/how-to/modify-project).\n",
"if \"google.colab\" in sys.modules:\n",
" from google.colab import auth\n",
"\n",
"# @markdown 2. **[Optional]** Set region. If not set, the region will be set automatically according to Colab Enterprise environment.\n",
" auth.authenticate_user()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "AcW2nwB8-7yC"
},
"source": [
"### Set Google Cloud Project Information\n",
"\n",
"REGION = \"\" # @param {type:\"string\"}\n",
"To get started with Vertex AI, ensure you have an existing Google Cloud project and that the [Vertex AI API is enabled](https://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com).\n",
"\n",
"# @markdown 3. If you want to run predictions with A100 80GB or H100 GPUs, we recommend using the regions listed below. **NOTE:** Make sure you have associated quota in selected regions. Click the links to see your current quota for each GPU type: [Nvidia A100 80GB](https://console.cloud.google.com/iam-admin/quotas?metric=aiplatform.googleapis.com%2Fcustom_model_serving_nvidia_a100_80gb_gpus), [Nvidia H100 80GB](https://console.cloud.google.com/iam-admin/quotas?metric=aiplatform.googleapis.com%2Fcustom_model_serving_nvidia_h100_gpus). You can request for quota following the instructions at [\"Request a higher quota\"](https://cloud.google.com/docs/quota/view-manage#requesting_higher_quota).\n",
"\n",
"# @markdown > | Machine Type | Accelerator Type | Recommended Regions |\n",
"# @markdown | ----------- | ----------- | ----------- |\n",
"# @markdown | a2-ultragpu-1g | 1 NVIDIA_A100_80GB | us-central1, us-east4, europe-west4, asia-southeast1, us-east4 |\n",
"# @markdown | a3-highgpu-2g | 2 NVIDIA_H100_80GB | us-west1, asia-southeast1, europe-west4 |\n",
"# @markdown | a3-highgpu-4g | 4 NVIDIA_H100_80GB | us-west1, asia-southeast1, europe-west4 |\n",
"# @markdown | a3-highgpu-8g | 8 NVIDIA_H100_80GB | us-central1, europe-west4, us-west1, asia-southeast1 |\n",
"\n",
"# Import the necessary packages\n",
"\n",
"# Upgrade Vertex AI SDK.\n",
"! pip3 install --upgrade --quiet 'google-cloud-aiplatform==1.103.0'\n",
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"import importlib\n",
"import json\n",
"See the guide on [setting up your project and development environment](https://cloud.google.com/vertex-ai/docs/start/cloud-environment). Also confirm that [billing is enabled](https://cloud.google.com/billing/docs/how-to/modify-project).\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "eIVLp0oE--k-"
},
"outputs": [],
"source": [
"# Use the environment variable if the user doesn't provide Project ID.\n",
"import os\n",
"import re\n",
"import time\n",
"from typing import Tuple\n",
"\n",
"import requests\n",
"from google import auth\n",
"from google.cloud import aiplatform\n",
"\n",
"if os.environ.get(\"VERTEX_PRODUCT\") != \"COLAB_ENTERPRISE\":\n",
" ! pip install --upgrade tensorflow\n",
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
")\n",
"\n",
"LABEL = \"vllm_gpu\"\n",
"models, endpoints = {}, {}\n",
"\n",
"# Get the default cloud project id.\n",
"PROJECT_ID = os.environ[\"GOOGLE_CLOUD_PROJECT\"]\n",
"\n",
"# Get the default region for launching jobs.\n",
"if not REGION:\n",
" REGION = os.environ[\"GOOGLE_CLOUD_REGION\"]\n",
"\n",
"# Initialize Vertex AI API.\n",
"print(\"Initializing Vertex AI API.\")\n",
"aiplatform.init(project=PROJECT_ID, location=REGION)\n",
"\n",
"! gcloud config set project $PROJECT_ID\n",
"import vertexai\n",
"\n",
"vertexai.init(\n",
" project=PROJECT_ID,\n",
" location=REGION,\n",
")\n",
"PROJECT_ID = \"[your-project-id]\" # @param {type: \"string\", placeholder: \"[your-project-id]\", isTemplate: true}\n",
"\n",
"# @markdown # Access Llama 3.3 models on Vertex AI for serving\n",
"# @markdown The original models from Meta are converted into the Hugging Face format for serving in Vertex AI.\n",
"# @markdown Accept the model agreement to access the models:\n",
"# @markdown 1. Open the [Llama 3.3 model card](https://console.cloud.google.com/vertex-ai/publishers/meta/model-garden/llama3-3) from [Vertex AI Model Garden](https://cloud.google.com/model-garden).\n",
"# @markdown 2. Review and accept the agreement in the pop-up window on the model card page. If you have previously accepted the model agreement, there will not be a pop-up window on the model card page and this step is not needed.\n",
"# @markdown 3. After accepting the agreement of Llama 3.3, a `gs://` URI containing Llama 3.3 models will be shared.\n",
"# @markdown 4. Paste the URI in the `VERTEX_AI_MODEL_GARDEN_LLAMA_3_3` field below.\n",
"if not PROJECT_ID or PROJECT_ID == \"[your-project-id]\":\n",
" PROJECT_ID = str(os.environ.get(\"GOOGLE_CLOUD_PROJECT\"))\n",
"\n",
"REGION = os.environ.get(\"GOOGLE_CLOUD_REGION\", \"us-central1\")\n",
"\n",
"VERTEX_AI_MODEL_GARDEN_LLAMA_3_3 = \"\" # @param {type:\"string\", isTemplate:true}\n",
"assert (\n",
" VERTEX_AI_MODEL_GARDEN_LLAMA_3_3\n",
"), \"Click the agreement of Llama 3.3 in Vertex AI Model Garden, and get the GCS path of Llama 3.3 model artifacts.\"\n",
"parsed_gcs_url = re.search(\"gs://.*?(/)?(?=[ ]|$)\", VERTEX_AI_MODEL_GARDEN_LLAMA_3_3)\n",
"if parsed_gcs_url:\n",
" VERTEX_AI_MODEL_GARDEN_LLAMA3 = parsed_gcs_url.group()\n",
"assert VERTEX_AI_MODEL_GARDEN_LLAMA3.startswith(\n",
" \"gs://\"\n",
"), \"VERTEX_AI_MODEL_GARDEN_LLAMA3 is expected to be a GCS URI and must start with `gs://`.\""
"vertexai.init(project=PROJECT_ID, location=REGION)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "z-XybZjtgF9M"
"id": "Q0CXrvcZH_aw"
},
"source": [
"## Deploy Llama 3.3 70B Instruct with vLLM"
"### Import libraries"
]
},
{
@@ -211,77 +189,83 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "E8OiHHNNE_wj"
"id": "3G2UXB82ICs6"
},
"outputs": [],
"source": [
"# @title Select the model variants\n",
"\n",
"# @markdown Set the model to deploy.\n",
"\n",
"base_model_name = \"Llama-3.3-70B-Instruct\" # @param [\"Llama-3.3-70B-Instruct\"] {isTemplate:true}\n",
"\n",
"model_id = os.path.join(VERTEX_AI_MODEL_GARDEN_LLAMA_3_3, base_model_name)\n",
"ENABLE_DYNAMIC_LORA = True # @param {type:\"boolean\", isTemplate:true}\n",
"hf_model_id = \"meta-llama/\" + base_model_name\n",
"version_id = \"llama-3.3-70b-instruct\"\n",
"PUBLISHER_MODEL_NAME = f\"publishers/meta/models/llama3-3@{version_id}\"\n",
"\n",
"accelerator_type = \"NVIDIA_H100_80GB\" # @param [\"NVIDIA_H100_80GB\", \"NVIDIA_L4\"]\n",
"\n",
"# The pre-built serving docker images.\n",
"VLLM_DOCKER_URI = \"us-docker.pkg.dev/vertex-ai/vertex-vision-model-garden-dockers/pytorch-vllm-serve:20250114_0916_RC00_maas\"\n",
"\n",
"# @markdown Set use_dedicated_endpoint to False if you don't want to use [dedicated endpoint](https://cloud.google.com/vertex-ai/docs/general/deployment#create-dedicated-endpoint). Note that [dedicated endpoint does not support VPC Service Controls](https://cloud.google.com/vertex-ai/docs/predictions/choose-endpoint-type), uncheck the box if you are using VPC-SC.\n",
"use_dedicated_endpoint = True # @param {type:\"boolean\"}\n",
"\n",
"# @markdown Find Vertex AI prediction supported accelerators and regions at https://cloud.google.com/vertex-ai/docs/predictions/configure-compute.\n",
"if accelerator_type == \"NVIDIA_L4\":\n",
" machine_type = \"g2-standard-96\"\n",
" accelerator_count = 8\n",
" max_loras = 1\n",
"elif accelerator_type == \"NVIDIA_H100_80GB\":\n",
" machine_type = \"a3-highgpu-4g\"\n",
" accelerator_count = 4\n",
" max_loras = 1\n",
"else:\n",
" raise ValueError(\n",
" f\"Recommended GPU setting not found for: {accelerator_type} and {base_model_name}.\"\n",
" )\n",
"\n",
"common_util.check_quota(\n",
" project_id=PROJECT_ID,\n",
" region=REGION,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
" is_for_training=False,\n",
"from vertexai import model_garden"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "upYRiGtP_-iN"
},
"source": [
"## Deploy model"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "H2WC_0hXDVXc"
},
"source": [
"### Choose model variant"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "u41zbNa2EoFq"
},
"source": [
"You can proceed with the default model variant or select a different one."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "-fgC4NLSDkF7"
},
"outputs": [],
"source": [
"model_version = \"llama-3.3-70b-instruct\" # @param [\"llama-3.3-70b-instruct\", \"llama-3.3-70b-instruct-fp8\"] {isTemplate:true}\n",
"MODEL_NAME = f\"meta/llama3-3@{model_version}\""
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "VRnUgU8LF3_i"
},
"source": [
"To see all deployable model variants available in Model Garden, use:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "-QLd-wshF6sB"
},
"outputs": [],
"source": [
"all_model_versions = model_garden.list_deployable_models(\n",
" model_filter=\"llama3-3\", list_hf_models=False\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"cell_type": "markdown",
"metadata": {
"cellView": "form",
"id": "EYEJBfsbNNmR"
"id": "N0UeFHa2GO63"
},
"outputs": [],
"source": [
"# @title [Option 1] Deploy with Model Garden SDK\n",
"\n",
"# @markdown Deploy with Gen AI model-centric SDK. This section uploads the prebuilt model to Model Registry and deploys it to a Vertex AI Endpoint. It takes 15 minutes to 1 hour to finish depending on the size of the model. See [use open models with Vertex AI](https://cloud.google.com/vertex-ai/generative-ai/docs/open-models/use-open-models) for documentation on other use cases.\n",
"from vertexai import model_garden\n",
"\n",
"model = model_garden.OpenModel(PUBLISHER_MODEL_NAME)\n",
"endpoints[LABEL] = model.deploy(\n",
" machine_type=machine_type,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
" accept_eula=True, # Accept the End User License Agreement (EULA) on the model card before deploy. Otherwise, the deployment will be forbidden.\n",
")\n",
"\n",
"endpoint = endpoints[LABEL]"
"Once you've selected a model variant, initialize it:"
]
},
{
@@ -289,243 +273,118 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "Hl4l047_NSb3"
"id": "GZiV3trBBcA3"
},
"outputs": [],
"source": [
"# @title [Option 2] Deploy with customized configs\n",
"model = model_garden.OpenModel(MODEL_NAME)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "-0cL378wFlvf"
},
"source": [
"### Check the Deployment Configuration\n",
"\n",
"# @markdown This section uploads Llama 3.3 to Model Registry and deploys it to a Vertex AI Endpoint. It takes ~30 minutes.\n",
"Use the `list_deploy_options()` method to view the verified deployment configurations for your selected model. This helps ensure you have sufficient resources (e.g., GPU quota) available to deploy it."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "zm73g7vFFm9N"
},
"outputs": [],
"source": [
"deploy_options = model.list_deploy_options(concise=True)\n",
"print(deploy_options)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "WjV499VsGwrD"
},
"source": [
"### Deploy the Model\n",
"\n",
"# @markdown The serving efficiency of L4 GPUs is inferior to that of H100 GPUs, but L4 GPUs are nevertheless good serving solutions if you do not have H100 quota.\n",
"Now that you’ve reviewed the deployment options, use the `deploy()` method to serve the selected open model to a Vertex AI endpoint. Deployment time may vary depending on the model size and infrastructure requirements.\n",
"\n",
"# @markdown H100 is hard to get for now. It's recommended to use the deployment button in the model card. You can still try to deploy H100 endpoint through the notebook, but there is a chance that resource is not available.\n",
"\n",
"gpu_memory_utilization = 0.95\n",
"max_model_len = 8192 # Maximum context length.\n",
"\n",
"# @markdown Choose whether to use a [Spot VM](https://cloud.google.com/compute/docs/instances/spot) for the deployment.\n",
"is_spot = False # @param {type:\"boolean\"}\n",
"\n",
"# @markdown To enable the auto-scaling in deployment, you can set the following options:\n",
"\n",
"min_replica_count = 1 # @param {type:\"integer\"}\n",
"max_replica_count = 1 # @param {type:\"integer\"}\n",
"required_replica_count = 1 # @param {type:\"integer\"}\n",
"\n",
"# @markdown Set the target of GPU duty cycle or CPU usage between 1 and 100 for auto-scaling.\n",
"autoscale_by_gpu_duty_cycle_target = 0 # @param {type:\"integer\"}\n",
"autoscale_by_cpu_usage_target = 0 # @param {type:\"integer\"}\n",
"\n",
"# @markdown Note: GPU duty cycle is not the most accurate metric for scaling workloads. More advanced auto-scaling metrics are coming soon. See [the public doc](https://cloud.google.com/vertex-ai/docs/reference/rest/v1/DedicatedResources#AutoscalingMetricSpec) for more details.\n",
"\n",
"\n",
"def deploy_model_vllm(\n",
" model_name: str,\n",
" model_id: str,\n",
" publisher: str,\n",
" publisher_model_id: str,\n",
" base_model_id: str = None,\n",
" machine_type: str = \"g2-standard-8\",\n",
" accelerator_type: str = \"NVIDIA_L4\",\n",
" accelerator_count: int = 1,\n",
" gpu_memory_utilization: float = 0.9,\n",
" max_model_len: int = 4096,\n",
" dtype: str = \"auto\",\n",
" enable_trust_remote_code: bool = False,\n",
" enforce_eager: bool = False,\n",
" enable_lora: bool = False,\n",
" enable_chunked_prefill: bool = False,\n",
" enable_prefix_cache: bool = False,\n",
" host_prefix_kv_cache_utilization_target: float = 0.0,\n",
" max_loras: int = 1,\n",
" max_cpu_loras: int = 8,\n",
" use_dedicated_endpoint: bool = False,\n",
" max_num_seqs: int = 256,\n",
" model_type: str = None,\n",
" enable_llama_tool_parser: bool = False,\n",
" min_replica_count: int = 1,\n",
" max_replica_count: int = 1,\n",
" required_replica_count: int = 1,\n",
" autoscale_by_gpu_duty_cycle_target: int = 0,\n",
" autoscale_by_cpu_usage_target: int = 0,\n",
" is_spot: bool = False,\n",
") -> Tuple[aiplatform.Model, aiplatform.Endpoint]:\n",
" \"\"\"Deploys trained models with vLLM into Vertex AI.\"\"\"\n",
" endpoint = aiplatform.Endpoint.create(\n",
" display_name=f\"{model_name}-endpoint\",\n",
" dedicated_endpoint_enabled=use_dedicated_endpoint,\n",
" )\n",
"\n",
" if not base_model_id:\n",
" base_model_id = model_id\n",
"\n",
" # See https://docs.vllm.ai/en/latest/models/engine_args.html for a list of possible arguments with descriptions.\n",
" vllm_args = [\n",
" \"python\",\n",
" \"-m\",\n",
" \"vllm.entrypoints.api_server\",\n",
" \"--host=0.0.0.0\",\n",
" \"--port=8080\",\n",
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
" f\"--max-cpu-loras={max_cpu_loras}\",\n",
" f\"--max-num-seqs={max_num_seqs}\",\n",
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
" if enforce_eager:\n",
" vllm_args.append(\"--enforce-eager\")\n",
"\n",
" if enable_lora:\n",
" vllm_args.append(\"--enable-lora\")\n",
"\n",
" if enable_chunked_prefill:\n",
" vllm_args.append(\"--enable-chunked-prefill\")\n",
"\n",
" if enable_prefix_cache:\n",
" vllm_args.append(\"--enable-prefix-caching\")\n",
"\n",
" if 0 < host_prefix_kv_cache_utilization_target < 1:\n",
" vllm_args.append(\n",
" f\"--host-prefix-kv-cache-utilization-target={host_prefix_kv_cache_utilization_target}\"\n",
" )\n",
"\n",
" if model_type:\n",
" vllm_args.append(f\"--model-type={model_type}\")\n",
"\n",
" if enable_llama_tool_parser:\n",
" if \"Llama-4\" not in model_id:\n",
" vllm_args.append(\"--enable-auto-tool-choice\")\n",
" vllm_args.append(\"--tool-call-parser=vertex-llama-3\")\n",
" else:\n",
" vllm_args.append(\"--enable-auto-tool-choice\")\n",
" vllm_args.append(\"--tool-call-parser=llama3_json\")\n",
"\n",
" env_vars = {\n",
" \"MODEL_ID\": base_model_id,\n",
" \"DEPLOY_SOURCE\": \"notebook\",\n",
" }\n",
"\n",
" # HF_TOKEN is not a compulsory field and may not be defined.\n",
" try:\n",
" if HF_TOKEN:\n",
" env_vars[\"HF_TOKEN\"] = HF_TOKEN\n",
" except NameError:\n",
" pass\n",
"\n",
" model = aiplatform.Model.upload(\n",
" display_name=model_name,\n",
" serving_container_image_uri=VLLM_DOCKER_URI,\n",
" serving_container_args=vllm_args,\n",
" serving_container_ports=[8080],\n",
" serving_container_predict_route=\"/generate\",\n",
" serving_container_health_route=\"/ping\",\n",
" serving_container_environment_variables=env_vars,\n",
" serving_container_shared_memory_size_mb=(16 * 1024), # 16 GB\n",
" serving_container_deployment_timeout=7200,\n",
" model_garden_source_model_name=(\n",
" f\"publishers/{publisher}/models/{publisher_model_id}\"\n",
" ),\n",
" )\n",
" print(\n",
" f\"Deploying {model_name} on {machine_type} with {accelerator_count} {accelerator_type} GPU(s).\"\n",
" )\n",
"\n",
" creds, _ = auth.default()\n",
" auth_req = auth.transport.requests.Request()\n",
" creds.refresh(auth_req)\n",
"\n",
" url = f\"https://{REGION}-aiplatform.googleapis.com/ui/projects/{PROJECT_ID}/locations/{REGION}/endpoints/{endpoint.name}:deployModel\"\n",
" headers = {\n",
" \"Content-Type\": \"application/json\",\n",
" \"Authorization\": f\"Bearer {creds.token}\",\n",
" }\n",
" data = {\n",
" \"deployedModel\": {\n",
" \"model\": model.resource_name,\n",
" \"displayName\": model_name,\n",
" \"dedicatedResources\": {\n",
" \"machineSpec\": {\n",
" \"machineType\": machine_type,\n",
" \"acceleratorType\": accelerator_type,\n",
" \"acceleratorCount\": accelerator_count,\n",
" },\n",
" \"minReplicaCount\": min_replica_count,\n",
" \"requiredReplicaCount\": required_replica_count,\n",
" \"maxReplicaCount\": max_replica_count,\n",
" },\n",
" \"system_labels\": {\n",
" \"NOTEBOOK_NAME\": \"model_garden_pytorch_llama3_3_deployment.ipynb\",\n",
" \"NOTEBOOK_ENVIRONMENT\": common_util.get_deploy_source(),\n",
" },\n",
" },\n",
" }\n",
" if is_spot:\n",
" data[\"deployedModel\"][\"dedicatedResources\"][\"spot\"] = True\n",
" if autoscale_by_gpu_duty_cycle_target > 0 or autoscale_by_cpu_usage_target > 0:\n",
" data[\"deployedModel\"][\"dedicatedResources\"][\"autoscalingMetricSpecs\"] = []\n",
" if autoscale_by_gpu_duty_cycle_target > 0:\n",
" data[\"deployedModel\"][\"dedicatedResources\"][\n",
" \"autoscalingMetricSpecs\"\n",
" ].append(\n",
" {\n",
" \"metricName\": \"aiplatform.googleapis.com/prediction/online/accelerator/duty_cycle\",\n",
" \"target\": autoscale_by_gpu_duty_cycle_target,\n",
" }\n",
" )\n",
" if autoscale_by_cpu_usage_target > 0:\n",
" data[\"deployedModel\"][\"dedicatedResources\"][\n",
" \"autoscalingMetricSpecs\"\n",
" ].append(\n",
" {\n",
" \"metricName\": \"aiplatform.googleapis.com/prediction/online/cpu/utilization\",\n",
" \"target\": autoscale_by_cpu_usage_target,\n",
" }\n",
" )\n",
" response = requests.post(url, headers=headers, json=data)\n",
" print(f\"Deploy Model response: {response.json()}\")\n",
" if response.status_code != 200 or \"name\" not in response.json():\n",
" raise ValueError(f\"Failed to deploy model: {response.text}\")\n",
" common_util.poll_and_wait(response.json()[\"name\"], REGION, 7200)\n",
" print(\"endpoint_name:\", endpoint.name)\n",
"\n",
" return model, endpoint\n",
"\n",
"\n",
"models[LABEL], endpoints[LABEL] = deploy_model_vllm(\n",
" model_name=common_util.get_job_name_with_datetime(prefix=\"llama3-3-serve\"),\n",
" model_id=model_id,\n",
" publisher=\"meta\",\n",
" publisher_model_id=\"llama3-3\",\n",
" base_model_id=hf_model_id,\n",
" machine_type=machine_type,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
" gpu_memory_utilization=gpu_memory_utilization,\n",
" max_model_len=max_model_len,\n",
" max_loras=max_loras,\n",
" enforce_eager=True,\n",
" enable_lora=ENABLE_DYNAMIC_LORA,\n",
" enable_chunked_prefill=not ENABLE_DYNAMIC_LORA,\n",
"> **Note**: If the model requires accepting a license agreement (EULA), set the `accept_eula=True` flag in the deploy call. Set `use_dedicated_endpoint` to False if you don't want to use [dedicated endpoint](https://cloud.google.com/vertex-ai/docs/general/deployment#create-dedicated-endpoint)."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "wX1itVTvXdEP"
},
"outputs": [],
"source": [
"use_dedicated_endpoint = True"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "MRmPFEPoGzsB"
},
"outputs": [],
"source": [
"endpoint = model.deploy(\n",
" accept_eula=True,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
" enable_llama_tool_parser=True,\n",
" min_replica_count=min_replica_count,\n",
" max_replica_count=max_replica_count,\n",
" required_replica_count=required_replica_count,\n",
" autoscale_by_gpu_duty_cycle_target=autoscale_by_gpu_duty_cycle_target,\n",
" autoscale_by_cpu_usage_target=autoscale_by_cpu_usage_target,\n",
" is_spot=is_spot,\n",
")\n",
"# @markdown Click \"Show Code\" to see more details."
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "PHBtn8DQp-ID"
},
"source": [
"Alternatively, you can select one of the verified deployment configurations listed above."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "ADsJG8JYqI6c"
},
"outputs": [],
"source": [
"endpoint = model.deploy(\n",
" accept_eula=True,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
" serving_container_image_uri=\"us-docker.pkg.dev/vertex-ai-restricted/vertex-vision-model-garden-dockers/hex-llm-serve:stable\",\n",
" machine_type=\"ct6e-standard-8t\",\n",
" accelerator_type=\"ACCELERATOR_TYPE_UNSPECIFIED\",\n",
" accelerator_count=0,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "kqSUK2CwsImi"
},
"source": [
"To further customize your deployment, you can configure:\n",
"\n",
"- **Compute Resources**: Machine type, replica count (min/max), accelerator type and quantity.\n",
"- **Infrastructure**: Use Spot VMs, reservation affinity, or dedicated endpoints.\n",
"- **Serving Container**: Customize container image, ports, health checks, and environment variables.\n",
"\n",
"See the [Model Garden SDK README](https://github.com/googleapis/python-aiplatform/blob/main/vertexai/model_garden/README.md) for advanced configuration options."
]
},
{
@@ -545,13 +404,24 @@
"\n",
"# @markdown ```\n",
"# @markdown Human: What is a car?\n",
"# @markdown Assistant: A car, or a motor car, is a road-connected human-transportation system used to move people or goods from one place to another.\n",
"# @markdown Assistant: A car, or a motor car, is a road-connected human-transportation system used to move people or goods from one place to another. The term also encompasses a wide range of vehicles, including motorboats, trains, and aircrafts. Cars typically have four wheels, a cabin for passengers, and an engine or motor. They have been around since the early 19th century and are now one of the most popular forms of transportation, used for daily commuting, shopping, and other purposes.\n",
"# @markdown ```\n",
"\n",
"# @markdown Optionally, you can apply LoRA weights to prediction. Set `lora_id` to be either a GCS URI or a HuggingFace repo containing the LoRA weight.\n",
"\n",
"# @markdown Additionally, you can moderate the generated text with Vertex AI. See [Moderate text documentation](https://cloud.google.com/natural-language/docs/moderating-text) for more details.\n",
"\n",
"# Loads an existing endpoint instance using the endpoint name:\n",
"# - Using `endpoint_name = endpoint.name` allows us to get the\n",
"# endpoint name of the endpoint `endpoint` created in the cell\n",
"# above.\n",
"# - Alternatively, you can set `endpoint_name = \"1234567890123456789\"` to load\n",
"# an existing endpoint with the ID 1234567890123456789.\n",
"# You may uncomment the code below to load an existing endpoint.\n",
"\n",
"# endpoint_name = \"\" # @param {type:\"string\"}\n",
"# aip_endpoint_name = (\n",
"# f\"projects/{PROJECT_ID}/locations/{REGION}/endpoints/{endpoint_name}\"\n",
"# )\n",
"# endpoint = aiplatform.Endpoint(aip_endpoint_name)\n",
"\n",
"prompt = \"What is a car?\" # @param {type: \"string\"}\n",
"# @markdown If you encounter an issue like `ServiceUnavailable: 503 Took too long to respond when processing`, you can reduce the maximum number of output tokens, by lowering `max_tokens`.\n",
"max_tokens = 50 # @param {type:\"integer\"}\n",
@@ -560,21 +430,19 @@
"top_k = 1 # @param {type:\"integer\"}\n",
"# @markdown Set `raw_response` to `True` to obtain the raw model output. Set `raw_response` to `False` to apply additional formatting in the structure of `\"Prompt:\\n{prompt.strip()}\\nOutput:\\n{output}\"`.\n",
"raw_response = False # @param {type:\"boolean\"}\n",
"lora_id = \"\" # @param {type:\"string\", isTemplate: true}\n",
"\n",
"# Overrides parameters for inferences.\n",
"instance = {\n",
" \"prompt\": prompt,\n",
" \"max_tokens\": max_tokens,\n",
" \"temperature\": temperature,\n",
" \"top_p\": top_p,\n",
" \"top_k\": top_k,\n",
" \"raw_response\": raw_response,\n",
"}\n",
"if lora_id:\n",
" instance[\"dynamic-lora\"] = lora_id\n",
"instances = [instance]\n",
"response = endpoints[\"vllm_gpu\"].predict(\n",
"instances = [\n",
" {\n",
" \"prompt\": prompt,\n",
" \"max_tokens\": max_tokens,\n",
" \"temperature\": temperature,\n",
" \"top_p\": top_p,\n",
" \"top_k\": top_k,\n",
" \"raw_response\": raw_response,\n",
" },\n",
"]\n",
"response = endpoint.predict(\n",
" instances=instances, use_dedicated_endpoint=use_dedicated_endpoint\n",
")\n",
"\n",
@@ -596,8 +464,8 @@
"# @title Chat completion\n",
"\n",
"if use_dedicated_endpoint:\n",
" DEDICATED_ENDPOINT_DNS = endpoints[\"vllm_gpu\"].gca_resource.dedicated_endpoint_dns\n",
"ENDPOINT_RESOURCE_NAME = endpoints[\"vllm_gpu\"].resource_name\n",
" DEDICATED_ENDPOINT_DNS = endpoint.gca_resource.dedicated_endpoint_dns\n",
"ENDPOINT_RESOURCE_NAME = endpoint.resource_name\n",
"\n",
"# @title Chat Completions Inference\n",
"\n",
@@ -962,18 +830,12 @@
},
"outputs": [],
"source": [
"# @title Delete the models and endpoints\n",
"# @title Delete the endpoints\n",
"\n",
"# @markdown Delete the experiment models and endpoints to recycle the resources\n",
"# @markdown and avoid unnecessary continuous charges that may incur.\n",
"# @markdown Delete the endpoint.\n",
"\n",
"# Undeploy model and delete endpoint.\n",
"for endpoint in endpoints.values():\n",
" endpoint.delete(force=True)\n",
"\n",
"# Delete models.\n",
"for model in models.values():\n",
" model.delete()"
"if endpoint:\n",
" endpoint.delete(force=True)"
]
}
],
@@ -141,7 +141,7 @@
"\n",
"# Import the necessary packages.\n",
"! rm -rf vertex-ai-samples && git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"! cd vertex-ai-samples && git reset --hard c45f6a4f4d32e31a050f0e4ba52824b0caf4eda3\n",
"! cd vertex-ai-samples && git reset --hard 7ae13b346a72ee2a2dc8152dd40c6ddd72d6c810\n",
"\n",
"import datetime\n",
"import importlib\n",
@@ -154,7 +154,7 @@
" custom_job as gca_custom_job_compat\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -869,7 +869,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -878,6 +877,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -130,7 +130,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -372,7 +372,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -381,6 +380,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -131,7 +131,7 @@
"from google.cloud import aiplatform\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -637,7 +637,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -646,6 +645,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
File diff suppressed because it is too large Load Diff
@@ -54,40 +54,48 @@
{
"cell_type": "markdown",
"metadata": {
"id": "cbDI9ag4oR4C"
"id": "3de7470326a2"
},
"source": [
"## Overview\n",
"\n",
"This notebook demonstrates deploying prebuilt [LLaVA models](https://huggingface.co/LLaVA-hf) with [vLLM](https://github.com/vllm-project/vllm) to improve serving throughput.\n",
"This notebook demonstrates how to deploy a **Pytorch-Llava** open model on Google Cloud Vertex AI.\n",
"\n",
"### Objectives\n",
"\n",
"### Objective\n",
"- Deploy Pytorch-Llava using containerized backends like [vLLM](https://github.com/vllm-project/vllm) on GPU.\n",
"- Use the deployed model to serve chat completion requests for both text and multimodal inputs.\n",
"\n",
"- Download and deploy prebuilt LLaVA models\n",
"- Deploy LLaVA with [vLLM](https://github.com/vllm-project/vllm) to improve serving throughput\n",
"### File a Bug\n",
"\n",
"### File a bug\n",
"\n",
"File a bug on [GitHub](https://github.com/GoogleCloudPlatform/vertex-ai-samples/issues/new) if you encounter any issue with the notebook.\n",
"If you encounter issues with this notebook, report them on [GitHub](https://github.com/GoogleCloudPlatform/vertex-ai-samples/issues/new).\n",
"\n",
"### Costs\n",
"\n",
"This tutorial uses billable components of Google Cloud:\n",
"\n",
"* Vertex AI\n",
"* Cloud Storage\n",
"- Vertex AI\n",
"- Cloud Storage\n",
"\n",
"Learn about [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing), [Cloud Storage pricing](https://cloud.google.com/storage/pricing), and use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to generate a cost estimate based on your projected usage."
"Refer to the [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing) and [Cloud Storage pricing](https://cloud.google.com/storage/pricing) pages for more information. Use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to estimate your projected costs."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "hQJWRopioSKT"
"id": "jeYw-Czg-DFy"
},
"source": [
"## Before you begin"
"## Get Started"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "KgyhGvEzBDkj"
},
"source": [
"### Install Vertex AI SDK and other required packages"
]
},
{
@@ -95,84 +103,85 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "J_jmxcIZoSxU"
"id": "iCacdLqG-IsH"
},
"outputs": [],
"source": [
"# @title Setup Google Cloud project\n",
"%pip install --upgrade --force-reinstall --quiet 'google-cloud-aiplatform>=1.106.0' 'openai' 'google-auth==2.27.0' 'requests==2.32.3'"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "HUKCrpBy-3yf"
},
"source": [
"### Authenticate the Notebook Environment (Colab only)\n",
"\n",
"# @markdown 1. [Make sure that billing is enabled for your project](https://cloud.google.com/billing/docs/how-to/modify-project).\n",
"If you're running this notebook in Google Colab, run the following cell to authenticate."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "JXwCT1kn-3Gu"
},
"outputs": [],
"source": [
"import sys\n",
"\n",
"# @markdown 2. **[Optional]** Set region. If not set, the region will be set automatically according to Colab Enterprise environment.\n",
"if \"google.colab\" in sys.modules:\n",
" from google.colab import auth\n",
"\n",
"REGION = \"\" # @param {type:\"string\"}\n",
" auth.authenticate_user()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "AcW2nwB8-7yC"
},
"source": [
"### Set Google Cloud Project Information\n",
"\n",
"# @markdown 3. If you want to run predictions with A100 80GB or H100 GPUs, we recommend using the regions listed below. **NOTE:** Make sure you have associated quota in selected regions. Click the links to see your current quota for each GPU type: [Nvidia A100 80GB](https://console.cloud.google.com/iam-admin/quotas?metric=aiplatform.googleapis.com%2Fcustom_model_serving_nvidia_a100_80gb_gpus), [Nvidia H100 80GB](https://console.cloud.google.com/iam-admin/quotas?metric=aiplatform.googleapis.com%2Fcustom_model_serving_nvidia_h100_gpus). You can request for quota following the instructions at [\"Request a higher quota\"](https://cloud.google.com/docs/quota/view-manage#requesting_higher_quota).\n",
"To get started with Vertex AI, ensure you have an existing Google Cloud project and that the [Vertex AI API is enabled](https://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com).\n",
"\n",
"# @markdown > | Machine Type | Accelerator Type | Recommended Regions |\n",
"# @markdown | ----------- | ----------- | ----------- |\n",
"# @markdown | a2-ultragpu-1g | 1 NVIDIA_A100_80GB | us-central1, us-east4, europe-west4, asia-southeast1, us-east4 |\n",
"# @markdown | a3-highgpu-2g | 2 NVIDIA_H100_80GB | us-west1, asia-southeast1, europe-west4 |\n",
"# @markdown | a3-highgpu-4g | 4 NVIDIA_H100_80GB | us-west1, asia-southeast1, europe-west4 |\n",
"# @markdown | a3-highgpu-8g | 8 NVIDIA_H100_80GB | us-central1, europe-west4, us-west1, asia-southeast1 |\n",
"\n",
"# Upgrade Vertex AI SDK.\n",
"! pip3 install --upgrade --quiet 'google-cloud-aiplatform==1.103.0'\n",
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"import importlib\n",
"See the guide on [setting up your project and development environment](https://cloud.google.com/vertex-ai/docs/start/cloud-environment). Also confirm that [billing is enabled](https://cloud.google.com/billing/docs/how-to/modify-project).\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "eIVLp0oE--k-"
},
"outputs": [],
"source": [
"# Use the environment variable if the user doesn't provide Project ID.\n",
"import os\n",
"from typing import Tuple\n",
"\n",
"from google.cloud import aiplatform\n",
"\n",
"if os.environ.get(\"VERTEX_PRODUCT\") != \"COLAB_ENTERPRISE\":\n",
" ! pip install --upgrade tensorflow\n",
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
")\n",
"\n",
"\n",
"LABEL = \"vllm_gpu\"\n",
"models, endpoints = {}, {}\n",
"\n",
"# Get the default cloud project id.\n",
"PROJECT_ID = os.environ[\"GOOGLE_CLOUD_PROJECT\"]\n",
"\n",
"# Get the default region for launching jobs.\n",
"if not REGION:\n",
" REGION = os.environ[\"GOOGLE_CLOUD_REGION\"]\n",
"\n",
"# Initialize Vertex AI API.\n",
"print(\"Initializing Vertex AI API.\")\n",
"aiplatform.init(project=PROJECT_ID, location=REGION)\n",
"\n",
"! gcloud config set project $PROJECT_ID\n",
"import vertexai\n",
"\n",
"vertexai.init(\n",
" project=PROJECT_ID,\n",
" location=REGION,\n",
")\n",
"PROJECT_ID = \"[your-project-id]\" # @param {type: \"string\", placeholder: \"[your-project-id]\", isTemplate: true}\n",
"\n",
"# @markdown You must provide a Hugging Face User Access Token (with read access) to access models from Hugging Face. You can follow the [Hugging Face documentation](https://huggingface.co/docs/hub/en/security-tokens) to create a **read** access token and put it in the `HF_TOKEN` field below.\n",
"if not PROJECT_ID or PROJECT_ID == \"[your-project-id]\":\n",
" PROJECT_ID = str(os.environ.get(\"GOOGLE_CLOUD_PROJECT\"))\n",
"\n",
"HF_TOKEN = \"\" # @param {type:\"string\"}\n",
"assert HF_TOKEN, \"Provide a read HF_TOKEN to load models from Hugging Face.\"\n",
"REGION = os.environ.get(\"GOOGLE_CLOUD_REGION\", \"us-central1\")\n",
"\n",
"\n",
"# @markdown Click \"Show code\" to see more details."
"vertexai.init(project=PROJECT_ID, location=REGION)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "cEIT7Ol_zs9L"
"id": "Q0CXrvcZH_aw"
},
"source": [
"## Deploy prebuilt LLaVA models"
"### Import libraries"
]
},
{
@@ -180,206 +189,202 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "ixW6s1W5_oWN"
"id": "3G2UXB82ICs6"
},
"outputs": [],
"source": [
"# @title Select the model variants\n",
"\n",
"# @markdown Set the model to deploy.\n",
"base_model_name = \"llava-1.5-7b-hf\" # @param [\"llava-1.5-7b-hf\", \"llava-1.5-13b-hf\", \"llava-v1.6-mistral-7b-hf\", \"llava-v1.6-vicuna-7b-hf\"]\n",
"MODEL_ID = \"llava-hf/\" + base_model_name\n",
"PUBLISHER_MODEL_NAME = f\"publishers/liuhaotian/models/pytorch-llava@{base_model_name}\"\n",
"\n",
"accelerator_type = \"NVIDIA_L4\" # @param [\"NVIDIA_L4\", \"NVIDIA_TESLA_T4\", \"NVIDIA_TESLA_V100\", \"NVIDIA_TESLA_A100\", \"NVIDIA_A100_80GB\"]\n",
"accelerator_count = 1\n",
"if accelerator_type == \"NVIDIA_L4\":\n",
" machine_type = \"g2-standard-12\"\n",
"elif accelerator_type == \"NVIDIA_TESLA_T4\":\n",
" machine_type = \"n1-standard-8\"\n",
"elif accelerator_type == \"NVIDIA_TESLA_V100\":\n",
" machine_type = \"n1-standard-8\"\n",
"elif accelerator_type == \"NVIDIA_TESLA_A100\":\n",
" machine_type = \"a2-highgpu-1g\"\n",
"elif accelerator_type == \"NVIDIA_A100_80GB\":\n",
" machine_type = \"a2-ultragpu-1g\"\n",
"else:\n",
" raise ValueError(\n",
" f\"Recommended GPU setting not found for: {accelerator_type} and {base_model_name}.\"\n",
" )\n",
"\n",
"# @markdown Set use_dedicated_endpoint to False if you don't want to use [dedicated endpoint](https://cloud.google.com/vertex-ai/docs/general/deployment#create-dedicated-endpoint). Note that [dedicated endpoint does not support VPC Service Controls](https://cloud.google.com/vertex-ai/docs/predictions/choose-endpoint-type), uncheck the box if you are using VPC-SC.\n",
"use_dedicated_endpoint = True # @param {type:\"boolean\"}\n",
"\n",
"# The pre-built serving docker images.\n",
"VLLM_DOCKER_URI = \"us-docker.pkg.dev/vertex-ai/vertex-vision-model-garden-dockers/pytorch-vllm-serve:20241007_2233_RC00\"\n",
"\n",
"common_util.check_quota(\n",
" project_id=PROJECT_ID,\n",
" region=REGION,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
" is_for_training=False,\n",
"from vertexai import model_garden"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "upYRiGtP_-iN"
},
"source": [
"## Deploy model"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "H2WC_0hXDVXc"
},
"source": [
"### Choose model variant"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "u41zbNa2EoFq"
},
"source": [
"You can proceed with the default model variant or select a different one."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "-fgC4NLSDkF7"
},
"outputs": [],
"source": [
"model_version = \"llava-1.5-7b-hf\" # @param [\"llama3-llava-next-8b-hf\", \"llava-1.5-13b-hf\", \"llava-1.5-7b-hf\", \"llava-v1.5-7b\", \"llava-v1.6-34b-hf\", \"llava-v1.6-mistral-7b-hf\", \"llava-v1.6-vicuna-13b-hf\", \"llava-v1.6-vicuna-7b-hf\"] {isTemplate:true}\n",
"MODEL_NAME = f\"liuhaotian/pytorch-llava@{model_version}\""
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "VRnUgU8LF3_i"
},
"source": [
"To see all deployable model variants available in Model Garden, use:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "-QLd-wshF6sB"
},
"outputs": [],
"source": [
"all_model_versions = model_garden.list_deployable_models(\n",
" model_filter=\"pytorch-llava\", list_hf_models=False\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "N0UeFHa2GO63"
},
"source": [
"Once you've selected a model variant, initialize it:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "USB7dvYqvNdu"
"id": "GZiV3trBBcA3"
},
"outputs": [],
"source": [
"# @title Deploy with customized configs\n",
"model = model_garden.OpenModel(MODEL_NAME)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "-0cL378wFlvf"
},
"source": [
"### Check the Deployment Configuration\n",
"\n",
"# @markdown This section uploads prebuilt LLaVA models to Model Registry and deploys it to a Vertex AI Endpoint. It takes 15 to 30 minutes to finish depending on the size of the model.\n",
"Use the `list_deploy_options()` method to view the verified deployment configurations for your selected model. This helps ensure you have sufficient resources (e.g., GPU quota) available to deploy it."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "zm73g7vFFm9N"
},
"outputs": [],
"source": [
"deploy_options = model.list_deploy_options(concise=True)\n",
"print(deploy_options)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "WjV499VsGwrD"
},
"source": [
"### Deploy the Model\n",
"\n",
"vllm_dtype = \"bfloat16\"\n",
"max_model_len = 4096\n",
"gpu_memory_utilization = 0.9\n",
"Now that you’ve reviewed the deployment options, use the `deploy()` method to serve the selected open model to a Vertex AI endpoint. Deployment time may vary depending on the model size and infrastructure requirements.\n",
"\n",
"> **Note**: If the model requires accepting a license agreement (EULA), set the `accept_eula=True` flag in the deploy call. Set `use_dedicated_endpoint` to False if you don't want to use [dedicated endpoint](https://cloud.google.com/vertex-ai/docs/general/deployment#create-dedicated-endpoint)."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "wX1itVTvXdEP"
},
"outputs": [],
"source": [
"use_dedicated_endpoint = True"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "MRmPFEPoGzsB"
},
"outputs": [],
"source": [
"endpoint = model.deploy(\n",
" accept_eula=True,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "PHBtn8DQp-ID"
},
"source": [
"Alternatively, you can select one of the verified deployment configurations listed above."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "ADsJG8JYqI6c"
},
"outputs": [],
"source": [
"endpoint = model.deploy(\n",
" accept_eula=True,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
" serving_container_image_uri=\"us-docker.pkg.dev/vertex-ai/vertex-vision-model-garden-dockers/pytorch-vllm-serve:20250601_0916_RC01\",\n",
" machine_type=\"g2-standard-12\",\n",
" accelerator_type=\"NVIDIA_L4\",\n",
" accelerator_count=1,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "kqSUK2CwsImi"
},
"source": [
"To further customize your deployment, you can configure:\n",
"\n",
"def deploy_model_vllm(\n",
" model_name: str,\n",
" model_id: str,\n",
" publisher: str,\n",
" publisher_model_id: str,\n",
" base_model_id: str = None,\n",
" machine_type: str = \"g2-standard-8\",\n",
" accelerator_type: str = \"NVIDIA_L4\",\n",
" accelerator_count: int = 1,\n",
" gpu_memory_utilization: float = 0.9,\n",
" max_model_len: int = 4096,\n",
" dtype: str = \"auto\",\n",
" enable_trust_remote_code: bool = False,\n",
" enforce_eager: bool = False,\n",
" enable_lora: bool = False,\n",
" enable_chunked_prefill: bool = False,\n",
" enable_prefix_cache: bool = False,\n",
" host_prefix_kv_cache_utilization_target: float = 0.0,\n",
" max_loras: int = 1,\n",
" max_cpu_loras: int = 8,\n",
" use_dedicated_endpoint: bool = False,\n",
" max_num_seqs: int = 256,\n",
" model_type: str = None,\n",
" enable_llama_tool_parser: bool = False,\n",
") -> Tuple[aiplatform.Model, aiplatform.Endpoint]:\n",
" \"\"\"Deploys trained models with vLLM into Vertex AI.\"\"\"\n",
" endpoint = aiplatform.Endpoint.create(\n",
" display_name=f\"{model_name}-endpoint\",\n",
" dedicated_endpoint_enabled=use_dedicated_endpoint,\n",
" )\n",
"- **Compute Resources**: Machine type, replica count (min/max), accelerator type and quantity.\n",
"- **Infrastructure**: Use Spot VMs, reservation affinity, or dedicated endpoints.\n",
"- **Serving Container**: Customize container image, ports, health checks, and environment variables.\n",
"\n",
" if not base_model_id:\n",
" base_model_id = model_id\n",
"\n",
" # See https://docs.vllm.ai/en/latest/models/engine_args.html for a list of possible arguments with descriptions.\n",
" vllm_args = [\n",
" \"python\",\n",
" \"-m\",\n",
" \"vllm.entrypoints.api_server\",\n",
" \"--host=0.0.0.0\",\n",
" \"--port=8080\",\n",
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
" f\"--max-cpu-loras={max_cpu_loras}\",\n",
" f\"--max-num-seqs={max_num_seqs}\",\n",
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
" if enforce_eager:\n",
" vllm_args.append(\"--enforce-eager\")\n",
"\n",
" if enable_lora:\n",
" vllm_args.append(\"--enable-lora\")\n",
"\n",
" if enable_chunked_prefill:\n",
" vllm_args.append(\"--enable-chunked-prefill\")\n",
"\n",
" if enable_prefix_cache:\n",
" vllm_args.append(\"--enable-prefix-caching\")\n",
"\n",
" if 0 < host_prefix_kv_cache_utilization_target < 1:\n",
" vllm_args.append(\n",
" f\"--host-prefix-kv-cache-utilization-target={host_prefix_kv_cache_utilization_target}\"\n",
" )\n",
"\n",
" if model_type:\n",
" vllm_args.append(f\"--model-type={model_type}\")\n",
"\n",
" if enable_llama_tool_parser:\n",
" vllm_args.append(\"--enable-auto-tool-choice\")\n",
" vllm_args.append(\"--tool-call-parser=vertex-llama-3\")\n",
"\n",
" env_vars = {\n",
" \"MODEL_ID\": base_model_id,\n",
" \"DEPLOY_SOURCE\": \"notebook\",\n",
" }\n",
"\n",
" # HF_TOKEN is not a compulsory field and may not be defined.\n",
" try:\n",
" if HF_TOKEN:\n",
" env_vars[\"HF_TOKEN\"] = HF_TOKEN\n",
" except NameError:\n",
" pass\n",
"\n",
" model = aiplatform.Model.upload(\n",
" display_name=model_name,\n",
" serving_container_image_uri=VLLM_DOCKER_URI,\n",
" serving_container_args=vllm_args,\n",
" serving_container_ports=[8080],\n",
" serving_container_predict_route=\"/generate\",\n",
" serving_container_health_route=\"/ping\",\n",
" serving_container_environment_variables=env_vars,\n",
" serving_container_shared_memory_size_mb=(16 * 1024), # 16 GB\n",
" serving_container_deployment_timeout=7200,\n",
" model_garden_source_model_name=(\n",
" f\"publishers/{publisher}/models/{publisher_model_id}\"\n",
" ),\n",
" )\n",
" print(\n",
" f\"Deploying {model_name} on {machine_type} with {accelerator_count} {accelerator_type} GPU(s).\"\n",
" )\n",
" model.deploy(\n",
" endpoint=endpoint,\n",
" machine_type=machine_type,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
" deploy_request_timeout=1800,\n",
" system_labels={\n",
" \"NOTEBOOK_NAME\": \"model_garden_pytorch_llava.ipynb\",\n",
" \"NOTEBOOK_ENVIRONMENT\": common_util.get_deploy_source(),\n",
" },\n",
" )\n",
" print(\"endpoint_name:\", endpoint.name)\n",
"\n",
" return model, endpoint\n",
"\n",
"\n",
"models[LABEL], endpoints[LABEL] = deploy_model_vllm(\n",
" model_name=common_util.get_job_name_with_datetime(prefix=MODEL_ID),\n",
" model_id=MODEL_ID,\n",
" publisher=\"liuhaotian\",\n",
" publisher_model_id=\"pytorch-llava\",\n",
" machine_type=machine_type,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
" max_model_len=max_model_len,\n",
" gpu_memory_utilization=gpu_memory_utilization,\n",
" dtype=vllm_dtype,\n",
")\n",
"\n",
"# @markdown Click \"Show code\" to see more details."
"See the [Model Garden SDK README](https://github.com/googleapis/python-aiplatform/blob/main/vertexai/model_garden/README.md) for advanced configuration options."
]
},
{
@@ -429,7 +434,7 @@
" \"top_p\": top_p,\n",
" },\n",
"]\n",
"response = endpoints[LABEL].predict(instances=instances)\n",
"response = endpoint.predict(instances=instances)\n",
"\n",
"for raw_prediction in response.predictions:\n",
" prediction = raw_prediction.split(\"Output:\")\n",
@@ -457,16 +462,10 @@
"outputs": [],
"source": [
"# @title Delete the models and endpoints\n",
"# @markdown Delete the experiment models and endpoints to recycle the resources\n",
"# @markdown and avoid unnecessary continuous charges that may incur.\n",
"# @markdown Delete the endpoint.\n",
"\n",
"# Undeploy model and delete endpoint.\n",
"for endpoint in endpoints.values():\n",
" endpoint.delete(force=True)\n",
"\n",
"# Delete models.\n",
"for model in models.values():\n",
" model.delete()"
"if endpoint:\n",
" endpoint.delete(force=True)"
]
}
],
@@ -138,7 +138,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"LABEL = \"vllm_gpu\"\n",
@@ -364,7 +364,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -373,6 +372,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -156,7 +156,7 @@
" custom_job as gca_custom_job_compat\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -701,7 +701,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -710,6 +709,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -143,7 +143,7 @@
"from google.cloud import aiplatform\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -316,7 +316,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -325,6 +324,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -156,7 +156,7 @@
" custom_job as gca_custom_job_compat\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -706,7 +706,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -715,6 +714,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -130,7 +130,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"\n",
@@ -141,7 +141,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"LABEL = \"pytorch_inference_gpu\"\n",
@@ -140,7 +140,7 @@
"\n",
"# Import the necessary packages\n",
"! rm -rf vertex-ai-samples && git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"! cd vertex-ai-samples && git reset --hard f147b50332ec996f1c309653c35b91a5eed7824e\n",
"! cd vertex-ai-samples && git reset --hard 7ae13b346a72ee2a2dc8152dd40c6ddd72d6c810\n",
"\n",
"import datetime\n",
"import importlib\n",
@@ -153,7 +153,7 @@
" custom_job as gca_custom_job_compat\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
@@ -825,7 +825,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -834,6 +833,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -134,7 +134,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"\n",
@@ -359,7 +359,6 @@
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
@@ -368,6 +367,9 @@
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if gpu_memory_utilization:\n",
" vllm_args.append(f\"--gpu-memory-utilization={gpu_memory_utilization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
@@ -148,7 +148,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"LABEL = \"sglang_gpu\"\n",
@@ -59,34 +59,43 @@
"source": [
"## Overview\n",
"\n",
"This notebook demonstrates serving Qwen3 models with [SGLang](https://github.com/sgl-project/sglang). [Qwen3](https://huggingface.co/collections/Qwen/qwen3-67dd247413f0e2e4f653967f) is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support\n",
"This notebook demonstrates how to deploy a **Qwen 3** open model on Google Cloud Vertex AI.\n",
"\n",
"### Objectives\n",
"\n",
"### Objective\n",
"- Deploy Qwen 3 using containerized backends like [vLLM](https://github.com/vllm-project/vllm) on GPU.\n",
"- Use the deployed model to serve chat completion requests for both text and multimodal inputs.\n",
"\n",
"- Deploy Qwen3 with SGLang on GPU using single-host serving, and [Spot VMs](https://cloud.google.com/compute/docs/instances/spot) (Optional).\n",
"### File a Bug\n",
"\n",
"### File a bug\n",
"\n",
"File a bug on [GitHub](https://github.com/GoogleCloudPlatform/vertex-ai-samples/issues/new) if you encounter any issue with the notebook.\n",
"If you encounter issues with this notebook, report them on [GitHub](https://github.com/GoogleCloudPlatform/vertex-ai-samples/issues/new).\n",
"\n",
"### Costs\n",
"\n",
"This tutorial uses billable components of Google Cloud:\n",
"\n",
"* Vertex AI\n",
"* Cloud Storage\n",
"- Vertex AI\n",
"- Cloud Storage\n",
"\n",
"Learn about [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing), [Cloud Storage pricing](https://cloud.google.com/storage/pricing), and use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to generate a cost estimate based on your projected usage."
"Refer to the [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing) and [Cloud Storage pricing](https://cloud.google.com/storage/pricing) pages for more information. Use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to estimate your projected costs."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "264c07757582"
"id": "jeYw-Czg-DFy"
},
"source": [
"## Before you begin"
"## Get Started"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "KgyhGvEzBDkj"
},
"source": [
"### Install Vertex AI SDK and other required packages"
]
},
{
@@ -94,17 +103,22 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "ax7zWynUDcjk"
"id": "iCacdLqG-IsH"
},
"outputs": [],
"source": [
"# @title Request for quota\n",
"%pip install --upgrade --force-reinstall --quiet 'google-cloud-aiplatform>=1.106.0' 'openai' 'google-auth==2.27.0' 'requests==2.32.3'"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "HUKCrpBy-3yf"
},
"source": [
"### Authenticate the Notebook Environment (Colab only)\n",
"\n",
"# @markdown To deploy the largest variant of the Qwen3 models, you need 1 host of 8 x H100 machine. Check that you have sufficient quota:\n",
"# @markdown - For Spot VM quota, check [`CustomModelServingPreemptibleH100GPUsPerProjectPerRegion`](https://console.cloud.google.com/iam-admin/quotas?metric=aiplatform.googleapis.com%2Fcustom_model_serving_preemptible_nvidia_h100_gpus).\n",
"# @markdown - For regular VM quota, check [`CustomModelServingH100GPUsPerProjectPerRegion`](https://console.cloud.google.com/iam-admin/quotas?metric=aiplatform.googleapis.com%2Fcustom_model_serving_nvidia_h100_gpus).\n",
"#\n",
"# @markdown If you don't have sufficient quota, request for quota following the instructions at [\"Request a higher quota\"](https://cloud.google.com/docs/quota/view-manage#requesting_higher_quota)."
"If you're running this notebook in Google Colab, run the following cell to authenticate."
]
},
{
@@ -112,109 +126,146 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "YXFGIp1l-qtT"
"id": "JXwCT1kn-3Gu"
},
"outputs": [],
"source": [
"# @title Setup Google Cloud project\n",
"import sys\n",
"\n",
"# @markdown 1. [Make sure that billing is enabled for your project](https://cloud.google.com/billing/docs/how-to/modify-project).\n",
"if \"google.colab\" in sys.modules:\n",
" from google.colab import auth\n",
"\n",
"# @markdown 2. **[Optional]** Set region. If not set, the region will be set automatically according to Colab Enterprise environment.\n",
" auth.authenticate_user()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "AcW2nwB8-7yC"
},
"source": [
"### Set Google Cloud Project Information\n",
"\n",
"REGION = \"\" # @param {type:\"string\"}\n",
"To get started with Vertex AI, ensure you have an existing Google Cloud project and that the [Vertex AI API is enabled](https://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com).\n",
"\n",
"# @markdown 3. If you want to run predictions with A100 80GB or H100 GPUs, we recommend using the regions listed below. **NOTE:** Make sure you have associated quota in selected regions. Click the links to see your current quota for each GPU type: [Nvidia A100 80GB](https://console.cloud.google.com/iam-admin/quotas?metric=aiplatform.googleapis.com%2Fcustom_model_serving_nvidia_a100_80gb_gpus), [Nvidia H100 80GB](https://console.cloud.google.com/iam-admin/quotas?metric=aiplatform.googleapis.com%2Fcustom_model_serving_nvidia_h100_gpus). You can request for quota following the instructions at [\"Request a higher quota\"](https://cloud.google.com/docs/quota/view-manage#requesting_higher_quota).\n",
"\n",
"# @markdown > | Machine Type | Accelerator Type | Recommended Regions |\n",
"# @markdown | ----------- | ----------- | ----------- |\n",
"# @markdown | a2-ultragpu-1g | 1 NVIDIA_A100_80GB | us-central1, us-east4, europe-west4, asia-southeast1, us-east4 |\n",
"# @markdown | a3-highgpu-2g | 2 NVIDIA_H100_80GB | us-west1, asia-southeast1, europe-west4 |\n",
"# @markdown | a3-highgpu-4g | 4 NVIDIA_H100_80GB | us-west1, asia-southeast1, europe-west4 |\n",
"# @markdown | a3-highgpu-8g | 8 NVIDIA_H100_80GB | us-central1, europe-west4, us-west1, asia-southeast1 |\n",
"\n",
"# Upgrade Vertex AI SDK.\n",
"! pip3 install --upgrade --quiet 'google-cloud-aiplatform==1.103.0'\n",
"\n",
"# Import the necessary packages\n",
"import importlib\n",
"See the guide on [setting up your project and development environment](https://cloud.google.com/vertex-ai/docs/start/cloud-environment). Also confirm that [billing is enabled](https://cloud.google.com/billing/docs/how-to/modify-project).\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "eIVLp0oE--k-"
},
"outputs": [],
"source": [
"# Use the environment variable if the user doesn't provide Project ID.\n",
"import os\n",
"import time\n",
"from typing import Tuple\n",
"\n",
"import requests\n",
"from google import auth\n",
"from google.cloud import aiplatform\n",
"\n",
"# Upgrade Vertex AI SDK.\n",
"if os.environ.get(\"VERTEX_PRODUCT\") != \"COLAB_ENTERPRISE\":\n",
" ! pip install --upgrade tensorflow\n",
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
")\n",
"\n",
"\n",
"def check_quota(\n",
" project_id: str,\n",
" region: str,\n",
" resource_id: str,\n",
" accelerator_count: int,\n",
"):\n",
" \"\"\"Checks if the project and the region has the required quota.\"\"\"\n",
" quota = common_util.get_quota(project_id, region, resource_id)\n",
" quota_request_instruction = (\n",
" \"Either use \"\n",
" \"a different region or request additional quota. Follow \"\n",
" \"instructions here \"\n",
" \"https://cloud.google.com/docs/quotas/view-manage#requesting_higher_quota\"\n",
" \" to check quota in a region or request additional quota for \"\n",
" \"your project.\"\n",
" )\n",
" if quota == -1:\n",
" raise ValueError(\n",
" f\"Quota not found for: {resource_id} in {region}.\"\n",
" f\" {quota_request_instruction}\"\n",
" )\n",
" if quota < accelerator_count:\n",
" raise ValueError(\n",
" f\"Quota not enough for {resource_id} in {region}: {quota} <\"\n",
" f\" {accelerator_count}. {quota_request_instruction}\"\n",
" )\n",
"\n",
"\n",
"LABEL = \"sglang_gpu\"\n",
"models, endpoints = {}, {}\n",
"\n",
"# Get the default cloud project id.\n",
"PROJECT_ID = os.environ[\"GOOGLE_CLOUD_PROJECT\"]\n",
"\n",
"# Get the default region for launching jobs.\n",
"if not REGION:\n",
" REGION = os.environ[\"GOOGLE_CLOUD_REGION\"]\n",
"\n",
"# Initialize Vertex AI API.\n",
"print(\"Initializing Vertex AI API.\")\n",
"aiplatform.init(project=PROJECT_ID, location=REGION)\n",
"\n",
"! gcloud config set project $PROJECT_ID\n",
"\n",
"import vertexai\n",
"\n",
"vertexai.init(\n",
" project=PROJECT_ID,\n",
" location=REGION,\n",
"PROJECT_ID = \"[your-project-id]\" # @param {type: \"string\", placeholder: \"[your-project-id]\", isTemplate: true}\n",
"\n",
"if not PROJECT_ID or PROJECT_ID == \"[your-project-id]\":\n",
" PROJECT_ID = str(os.environ.get(\"GOOGLE_CLOUD_PROJECT\"))\n",
"\n",
"REGION = os.environ.get(\"GOOGLE_CLOUD_REGION\", \"us-central1\")\n",
"\n",
"vertexai.init(project=PROJECT_ID, location=REGION)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "Q0CXrvcZH_aw"
},
"source": [
"### Import libraries"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "3G2UXB82ICs6"
},
"outputs": [],
"source": [
"from vertexai import model_garden"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "upYRiGtP_-iN"
},
"source": [
"## Deploy model"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "H2WC_0hXDVXc"
},
"source": [
"### Choose model variant"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "u41zbNa2EoFq"
},
"source": [
"You can proceed with the default model variant or select a different one."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "-fgC4NLSDkF7"
},
"outputs": [],
"source": [
"model_version = \"qwen3-4b\" # @param [\"qwen3-0.6b\", \"qwen3-0.6b-base\", \"qwen3-0.6b-fp8\", \"qwen3-1.7b\", \"qwen3-1.7b-base\", \"qwen3-1.7b-fp8\", \"qwen3-14b\", \"qwen3-14b-base\", \"qwen3-14b-fp8\", \"qwen3-235b-a22b\", \"qwen3-235b-a22b-fp8\", \"qwen3-235b-a22b-instruct-2507\", \"qwen3-235b-a22b-instruct-2507-fp8\", \"qwen3-235b-a22b-thinking-2507\", \"qwen3-235b-a22b-thinking-2507-fp8\", \"qwen3-30b-a3b\", \"qwen3-30b-a3b-base\", \"qwen3-30b-a3b-fp8\", \"qwen3-30b-a3b-instruct-2507\", \"qwen3-30b-a3b-instruct-2507-fp8\", \"qwen3-30b-a3b-thinking-2507\", \"qwen3-30b-a3b-thinking-2507-fp8\", \"qwen3-32b\", \"qwen3-32b-fp8\", \"qwen3-4b\", \"qwen3-4b-base\", \"qwen3-4b-fp8\", \"qwen3-4b-instruct-2507\", \"qwen3-4b-instruct-2507-fp8\", \"qwen3-4b-thinking-2507\", \"qwen3-4b-thinking-2507-fp8\", \"qwen3-8b\", \"qwen3-8b-base\", \"qwen3-8b-fp8\"] {isTemplate:true}\n",
"MODEL_NAME = f\"qwen/qwen3@{model_version}\""
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "VRnUgU8LF3_i"
},
"source": [
"To see all deployable model variants available in Model Garden, use:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "-QLd-wshF6sB"
},
"outputs": [],
"source": [
"all_model_versions = model_garden.list_deployable_models(\n",
" model_filter=\"qwen3\", list_hf_models=False\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "3zJJDmldn7rw"
"id": "N0UeFHa2GO63"
},
"source": [
"## Deploy Qwen3 with SGLang"
"Once you've selected a model variant, initialize it:"
]
},
{
@@ -222,110 +273,22 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "_3Swj3pxn7rw"
"id": "GZiV3trBBcA3"
},
"outputs": [],
"source": [
"# @title Select the model variants\n",
"model = model_garden.OpenModel(MODEL_NAME)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "-0cL378wFlvf"
},
"source": [
"### Check the Deployment Configuration\n",
"\n",
"# @markdown Set the model to deploy.\n",
"\n",
"base_model_name = \"Qwen3-235B-A22B-Instruct-2507\" # @param [\"Qwen3-235B-A22B-Instruct-2507\", \"Qwen3-235B-A22B-Instruct-2507-FP8\", \"Qwen3-235B-A22B-Thinking-2507\", \"Qwen3-235B-A22B-Thinking-2507-FP8\", \"Qwen3-235B-A22B\", \"Qwen3-235B-A22B-FP8\", \"Qwen3-30B-A3B\", \"Qwen3-30B-A3B-Base\", \"Qwen/Qwen3-30B-A3B-FP8\", \"Qwen3-32B\", \"Qwen/Qwen3-32B-FP8\", \"Qwen3-14B\", \"Qwen3-14B-Base\", \"Qwen/Qwen3-14B-FP8\", \"Qwen3-8B\", \"Qwen3-8B-Base\", \"Qwen/Qwen3-8B-FP8\", \"Qwen3-4B\", \"Qwen3-4B-Base\", \"Qwen/Qwen3-4B-FP8\", \"Qwen3-1.7B\", \"Qwen3-1.7B-Base\", \"Qwen/Qwen3-1.7B-FP8\", \"Qwen3-0.6B\", \"Qwen3-0.6B-Base\", \"Qwen/Qwen3-0.6B-FP8\"] {isTemplate:true}\n",
"model_id = \"Qwen/\" + base_model_name\n",
"hf_model_id = model_id\n",
"\n",
"# The pre-built serving docker images.\n",
"SGLANG_DOCKER_URI = \"us-docker.pkg.dev/deeplearning-platform-release/vertex-model-garden/sglang-serve.cu124.0-4.ubuntu2204.py310:model-garden.sglang-0-4-release_20250718.00_p0\"\n",
"\n",
"# @markdown Choose whether to use a [Spot VM](https://cloud.google.com/compute/docs/instances/spot) for the deployment.\n",
"is_spot = False # @param {type:\"boolean\"}\n",
"\n",
"# @markdown Set use_dedicated_endpoint to False if you don't want to use [dedicated endpoint](https://cloud.google.com/vertex-ai/docs/general/deployment#create-dedicated-endpoint). Note that [dedicated endpoint does not support VPC Service Controls](https://cloud.google.com/vertex-ai/docs/predictions/choose-endpoint-type), uncheck the box if you are using VPC-SC.\n",
"use_dedicated_endpoint = True # @param {type:\"boolean\"}\n",
"\n",
"# @markdown Find Vertex AI prediction supported accelerators and regions at https://cloud.google.com/vertex-ai/docs/predictions/configure-compute.\n",
"accelerator_type = \"NVIDIA_H100_80GB\" # @param [\"NVIDIA_H100_80GB\", \"NVIDIA_L4\"] {isTemplate:true}\n",
"\n",
"PUBLISHER_MODEL_NAME = f\"publishers/qwen/models/qwen3@{base_model_name.lower()}\"\n",
"\n",
"if accelerator_type == \"NVIDIA_H100_80GB\":\n",
" if is_spot:\n",
" resource_id = \"custom_model_serving_preemptible_nvidia_h100_gpus\"\n",
" else:\n",
" resource_id = \"custom_model_serving_nvidia_h100_gpus\"\n",
" if base_model_name in [\n",
" \"Qwen3-235B-A22B\",\n",
" \"Qwen3-235B-A22B-Instruct-2507\",\n",
" \"Qwen3-235B-A22B-Thinking-2507\",\n",
" ]:\n",
" machine_type = \"a3-highgpu-8g\"\n",
" accelerator_count = 8\n",
" model_id = f\"gs://vertex-model-garden-restricted-us/qwen3/{base_model_name}\"\n",
" elif base_model_name in [\n",
" \"Qwen3-235B-A22B-FP8\",\n",
" \"Qwen3-235B-A22B-Instruct-2507-FP8\",\n",
" \"Qwen3-235B-A22B-Thinking-2507-FP8\",\n",
" ]:\n",
" machine_type = \"a3-highgpu-4g\"\n",
" accelerator_count = 4\n",
" else:\n",
" machine_type = \"a3-highgpu-1g\"\n",
" accelerator_count = 1\n",
"elif accelerator_type == \"NVIDIA_L4\":\n",
" if is_spot:\n",
" resource_id = \"custom_model_serving_preemptible_nvidia_l4_gpus\"\n",
" else:\n",
" resource_id = \"custom_model_serving_nvidia_l4_gpus\"\n",
" if base_model_name in [\n",
" \"Qwen3-235B-A22B\",\n",
" \"Qwen3-235B-A22B-FP8\",\n",
" \"Qwen3-235B-A22B-Instruct-2507\",\n",
" \"Qwen3-235B-A22B-Instruct-2507-FP8\",\n",
" \"Qwen3-235B-A22B-Thinking-2507\",\n",
" \"Qwen3-235B-A22B-Thinking-2507-FP8\",\n",
" ]:\n",
" raise ValueError(\"L4s are insufficient to serve Qwen3-235B-A22B models.\")\n",
" elif base_model_name in [\"Qwen3-30B-A3B\", \"Qwen3-30B-A3B-Base\", \"Qwen3-32B\"]:\n",
" machine_type = \"g2-standard-48\"\n",
" accelerator_count = 4\n",
" elif base_model_name in [\n",
" \"Qwen3-14B\",\n",
" \"Qwen3-14B-Base\",\n",
" \"Qwen3-30B-A3B-FP8\",\n",
" \"Qwen3-32B-FP8\",\n",
" ]:\n",
" machine_type = \"g2-standard-24\"\n",
" accelerator_count = 2\n",
" elif base_model_name in [\n",
" \"Qwen3-8B\",\n",
" \"Qwen3-8B-Base\",\n",
" \"Qwen3-4B\",\n",
" \"Qwen3-4B-Base\",\n",
" \"Qwen3-1.7B\",\n",
" \"Qwen3-1.7B-Base\",\n",
" \"Qwen3-0.6B\",\n",
" \"Qwen3-0.6B-Base\",\n",
" \"Qwen/Qwen3-14B-FP8\",\n",
" \"Qwen/Qwen3-8B-FP8\",\n",
" \"Qwen/Qwen3-4B-FP8\",\n",
" \"Qwen/Qwen3-1.7B-FP8\",\n",
" \"Qwen/Qwen3-0.6B-FP8\",\n",
" ]:\n",
" machine_type = \"g2-standard-12\"\n",
" accelerator_count = 1\n",
" else:\n",
" raise ValueError(f\"Recommended GPU setting not found for: {base_model_name}.\")\n",
"else:\n",
" raise ValueError(f\"Recommended GPU setting not found for: {base_model_name}.\")\n",
"\n",
"check_quota(\n",
" project_id=PROJECT_ID,\n",
" region=REGION,\n",
" resource_id=resource_id,\n",
" accelerator_count=accelerator_count,\n",
")\n",
"\n",
"# @markdown Click \"Show Code\" to see more details."
"Use the `list_deploy_options()` method to view the verified deployment configurations for your selected model. This helps ensure you have sufficient resources (e.g., GPU quota) available to deploy it."
]
},
{
@@ -333,29 +296,61 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "omW0LaC8wWz5"
"id": "zm73g7vFFm9N"
},
"outputs": [],
"source": [
"# @title [Option 1] Deploy with Model Garden SDK\n",
"# @markdown Deploy with Gen AI model-centric SDK. This section uploads the prebuilt model to Model Registry and deploys it to a Vertex AI Endpoint. It takes 15 minutes to 1 hour to finish depending on the size of the model. See [use open models with Vertex AI](https://cloud.google.com/vertex-ai/generative-ai/docs/open-models/use-open-models) for documentation on other use cases.\n",
"deploy_request_timeout = 1800 # 30 minutes\n",
"from vertexai import model_garden\n",
"deploy_options = model.list_deploy_options(concise=True)\n",
"print(deploy_options)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "WjV499VsGwrD"
},
"source": [
"### Deploy the Model\n",
"\n",
"model = model_garden.OpenModel(PUBLISHER_MODEL_NAME)\n",
"endpoints[LABEL] = model.deploy(\n",
" machine_type=machine_type,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
"Now that you’ve reviewed the deployment options, use the `deploy()` method to serve the selected open model to a Vertex AI endpoint. Deployment time may vary depending on the model size and infrastructure requirements.\n",
"\n",
"> **Note**: If the model requires accepting a license agreement (EULA), set the `accept_eula=True` flag in the deploy call. Set `use_dedicated_endpoint` to False if you don't want to use [dedicated endpoint](https://cloud.google.com/vertex-ai/docs/general/deployment#create-dedicated-endpoint)."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "wX1itVTvXdEP"
},
"outputs": [],
"source": [
"use_dedicated_endpoint = True"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "MRmPFEPoGzsB"
},
"outputs": [],
"source": [
"endpoint = model.deploy(\n",
" accept_eula=True,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
" spot=is_spot,\n",
" deploy_request_timeout=deploy_request_timeout,\n",
" accept_eula=False,\n",
")\n",
"\n",
"endpoint = endpoints[LABEL]\n",
"\n",
"# @markdown Click \"Show Code\" to see more details."
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "PHBtn8DQp-ID"
},
"source": [
"Alternatively, you can select one of the verified deployment configurations listed above."
]
},
{
@@ -363,249 +358,33 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "3m-tDxgawYhU"
"id": "ADsJG8JYqI6c"
},
"outputs": [],
"source": [
"# @title [Option 2] Deploy with customized configs\n",
"\n",
"# @markdown This section uploads Qwen3 models to Model Registry and deploys them to a Vertex Prediction Endpoint. It takes ~1 hour to finish.\n",
"\n",
"# @markdown It's recommended to use the region selected by the deployment button on the model card. If the deployment button is not available, it's recommended to stay with the default region of the notebook.\n",
"\n",
"\n",
"def poll_operation(op_name: str) -> bool: # noqa: F811\n",
" creds, _ = auth.default()\n",
" auth_req = auth.transport.requests.Request()\n",
" creds.refresh(auth_req)\n",
" headers = {\n",
" \"Authorization\": f\"Bearer {creds.token}\",\n",
" }\n",
" get_resp = requests.get(\n",
" f\"https://{REGION}-aiplatform.googleapis.com/ui/{op_name}\",\n",
" headers=headers,\n",
" )\n",
" opjs = get_resp.json()\n",
" if \"error\" in opjs:\n",
" raise ValueError(f\"Operation failed: {opjs['error']}\")\n",
" return opjs.get(\"done\", False)\n",
"\n",
"\n",
"def poll_and_wait(op_name: str, total_wait: int, interval: int = 60): # noqa: F811\n",
" waited = 0\n",
" while not poll_operation(op_name):\n",
" if waited > total_wait:\n",
" raise TimeoutError(\"Operation timed out\")\n",
" print(\n",
" f\"\\rStill waiting for operation... Waited time in second: {waited:<6}\",\n",
" end=\"\",\n",
" flush=True,\n",
" )\n",
" waited += interval\n",
" time.sleep(interval)\n",
"\n",
"\n",
"def deploy_model_sglang_multihost(\n",
" model_name: str,\n",
" model_id: str,\n",
" publisher: str,\n",
" publisher_model_id: str,\n",
" service_account: str = \"\",\n",
" base_model_id: str = \"\",\n",
" machine_type: str = \"g2-standard-8\",\n",
" accelerator_type: str = \"NVIDIA_L4\",\n",
" accelerator_count: int = 1,\n",
" multihost_gpu_node_count: int = 1,\n",
" gpu_memory_utilization: float | None = None,\n",
" context_length: int | None = None,\n",
" dtype: str | None = None,\n",
" quantization: str | None = None,\n",
" enable_trust_remote_code: bool = False,\n",
" enable_torch_compile: bool = False,\n",
" torch_compile_max_bs: int | None = None,\n",
" attention_backend: str = \"\",\n",
" enable_flashinfer_mla: bool = False,\n",
" disable_cuda_graph: bool = False,\n",
" speculative_algorithm: str | None = None,\n",
" speculative_draft_model_path: str = \"\",\n",
" speculative_num_steps: int = 3,\n",
" speculative_eagle_topk: int = 1,\n",
" speculative_num_draft_tokens: int = 4,\n",
" enable_jit_deepgemm: bool = False,\n",
" enable_dp_attention: bool = False,\n",
" dp_size: int = 1,\n",
" enable_multimodal: bool = False,\n",
" use_dedicated_endpoint: bool = False,\n",
" max_num_seqs: int | None = None,\n",
" is_spot: bool = True,\n",
" tool_call_parser: str | None = None,\n",
") -> Tuple[aiplatform.Model, aiplatform.Endpoint]:\n",
" \"\"\"Deploys trained models with SGLang into Vertex AI.\"\"\"\n",
" endpoint = aiplatform.Endpoint.create(\n",
" display_name=f\"{model_name}-endpoint\",\n",
" dedicated_endpoint_enabled=use_dedicated_endpoint,\n",
" )\n",
"\n",
" if not base_model_id:\n",
" base_model_id = model_id\n",
"\n",
" # See https://docs.sglang.ai/backend/server_arguments.html for a list of possible arguments with descriptions.\n",
" sglang_args = [\n",
" f\"--model={model_id}\",\n",
" f\"--tp={accelerator_count * multihost_gpu_node_count}\",\n",
" f\"--dp={dp_size}\",\n",
" ]\n",
"\n",
" if context_length:\n",
" sglang_args.append(f\"--context-length={context_length}\")\n",
"\n",
" if gpu_memory_utilization:\n",
" sglang_args.append(f\"--mem-fraction-static={gpu_memory_utilization}\")\n",
"\n",
" if max_num_seqs:\n",
" sglang_args.append(f\"--max-running-requests={max_num_seqs}\")\n",
"\n",
" if dtype:\n",
" sglang_args.append(f\"--dtype={dtype}\")\n",
"\n",
" if quantization:\n",
" sglang_args.append(f\"--quantization={quantization}\")\n",
"\n",
" if enable_trust_remote_code:\n",
" sglang_args.append(\"--trust-remote-code\")\n",
"\n",
" if enable_torch_compile:\n",
" sglang_args.append(\"--enable-torch-compile\")\n",
" if torch_compile_max_bs:\n",
" sglang_args.append(f\"--torch-compile-max-bs={torch_compile_max_bs}\")\n",
"\n",
" if attention_backend:\n",
" sglang_args.append(f\"--attention-backend={attention_backend}\")\n",
"\n",
" if enable_flashinfer_mla:\n",
" sglang_args.append(\"--enable-flashinfer-mla\")\n",
"\n",
" if disable_cuda_graph:\n",
" sglang_args.append(\"--disable-cuda-graph\")\n",
"\n",
" if speculative_algorithm:\n",
" sglang_args.append(f\"--speculative-algorithm={speculative_algorithm}\")\n",
" sglang_args.append(\n",
" f\"--speculative-draft-model-path={speculative_draft_model_path}\"\n",
" )\n",
" sglang_args.append(f\"--speculative-num-steps={speculative_num_steps}\")\n",
" sglang_args.append(f\"--speculative-eagle-topk={speculative_eagle_topk}\")\n",
" sglang_args.append(\n",
" f\"--speculative-num-draft-tokens={speculative_num_draft_tokens}\"\n",
" )\n",
"\n",
" if enable_dp_attention:\n",
" sglang_args.append(\"--enable-dp-attention\")\n",
"\n",
" if enable_multimodal:\n",
" sglang_args.append(\"--enable-multimodal\")\n",
"\n",
" if tool_call_parser:\n",
" sglang_args.append(f\"--tool-call-parser={tool_call_parser}\")\n",
"\n",
" env_vars = {\n",
" \"MODEL_ID\": base_model_id,\n",
" \"DEPLOY_SOURCE\": \"notebook\",\n",
" }\n",
"\n",
" if enable_jit_deepgemm:\n",
" env_vars[\"SGL_ENABLE_JIT_DEEPGEMM\"] = \"1\"\n",
"\n",
" # HF_TOKEN is not a compulsory field and may not be defined.\n",
" try:\n",
" if HF_TOKEN:\n",
" env_vars[\"HF_TOKEN\"] = HF_TOKEN\n",
" except NameError:\n",
" pass\n",
"\n",
" model = aiplatform.Model.upload(\n",
" display_name=model_name,\n",
" serving_container_image_uri=SGLANG_DOCKER_URI,\n",
" serving_container_args=sglang_args,\n",
" serving_container_ports=[30000],\n",
" serving_container_predict_route=\"/vertex_generate\",\n",
" serving_container_health_route=\"/health\",\n",
" serving_container_environment_variables=env_vars,\n",
" serving_container_shared_memory_size_mb=(16 * 1024), # 16 GB\n",
" serving_container_deployment_timeout=7200,\n",
" model_garden_source_model_name=(\n",
" f\"publishers/{publisher}/models/{publisher_model_id}\"\n",
" ),\n",
" )\n",
" print(\n",
" f\"Deploying {model_name} on {machine_type} with {int(accelerator_count * multihost_gpu_node_count)} {accelerator_type} GPU(s).\"\n",
" )\n",
"\n",
" creds, _ = auth.default()\n",
" auth_req = auth.transport.requests.Request()\n",
" creds.refresh(auth_req)\n",
"\n",
" url = f\"https://{REGION}-aiplatform.googleapis.com/ui/projects/{PROJECT_ID}/locations/{REGION}/endpoints/{endpoint.name}:deployModel\"\n",
" headers = {\n",
" \"Content-Type\": \"application/json\",\n",
" \"Authorization\": f\"Bearer {creds.token}\",\n",
" }\n",
" data = {\n",
" \"deployedModel\": {\n",
" \"model\": model.resource_name,\n",
" \"displayName\": model_name,\n",
" \"dedicatedResources\": {\n",
" \"machineSpec\": {\n",
" \"machineType\": machine_type,\n",
" \"multihostGpuNodeCount\": multihost_gpu_node_count,\n",
" \"acceleratorType\": accelerator_type,\n",
" \"acceleratorCount\": accelerator_count,\n",
" },\n",
" \"minReplicaCount\": 1,\n",
" \"maxReplicaCount\": 1,\n",
" },\n",
" \"system_labels\": {\n",
" \"NOTEBOOK_NAME\": \"model_garden_pytorch_qwen3_deployment.ipynb\",\n",
" \"NOTEBOOK_ENVIRONMENT\": common_util.get_deploy_source(),\n",
" },\n",
" },\n",
" }\n",
" if service_account:\n",
" data[\"deployedModel\"][\"serviceAccount\"] = service_account\n",
" if is_spot:\n",
" data[\"deployedModel\"][\"dedicatedResources\"][\"spot\"] = True\n",
" response = requests.post(url, headers=headers, json=data)\n",
" print(f\"Deploy Model response: {response.json()}\")\n",
" if response.status_code != 200 or \"name\" not in response.json():\n",
" raise ValueError(f\"Failed to deploy model: {response.text}\")\n",
" poll_and_wait(response.json()[\"name\"], 7200)\n",
" print(\"endpoint_name:\", endpoint.name)\n",
"\n",
" return model, endpoint\n",
"\n",
"\n",
"common_util.check_quota(\n",
" project_id=PROJECT_ID,\n",
" region=REGION,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
" is_for_training=False,\n",
")\n",
"\n",
"models[LABEL], endpoints[LABEL] = deploy_model_sglang_multihost(\n",
" model_name=common_util.get_job_name_with_datetime(prefix=\"qwen3-serve\"),\n",
" model_id=model_id,\n",
" publisher=\"qwen\",\n",
" publisher_model_id=\"qwen3\",\n",
" base_model_id=hf_model_id,\n",
" machine_type=machine_type,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
"endpoint = model.deploy(\n",
" accept_eula=True,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
" is_spot=is_spot,\n",
")\n",
" serving_container_image_uri=\"us-docker.pkg.dev/deeplearning-platform-release/vertex-model-garden/sglang-serve.cu124.0-4.ubuntu2204.py310:20250428-1803-rc0\",\n",
" machine_type=\"a3-highgpu-2g\",\n",
" accelerator_type=\"NVIDIA_H100_80GB\",\n",
" accelerator_count=2,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "kqSUK2CwsImi"
},
"source": [
"To further customize your deployment, you can configure:\n",
"\n",
"# @markdown Click \"Show Code\" to see more details."
"- **Compute Resources**: Machine type, replica count (min/max), accelerator type and quantity.\n",
"- **Infrastructure**: Use Spot VMs, reservation affinity, or dedicated endpoints.\n",
"- **Serving Container**: Customize container image, ports, health checks, and environment variables.\n",
"\n",
"See the [Model Garden SDK README](https://github.com/googleapis/python-aiplatform/blob/main/vertexai/model_garden/README.md) for advanced configuration options."
]
},
{
@@ -683,7 +462,7 @@
" \"min_p\": min_p,\n",
" }\n",
"}\n",
"response = endpoints[\"sglang_gpu\"].predict(\n",
"response = endpoint.predict(\n",
" instances=instances,\n",
" parameters=parameters,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
@@ -707,8 +486,8 @@
"# @title Chat completion\n",
"\n",
"if use_dedicated_endpoint:\n",
" DEDICATED_ENDPOINT_DNS = endpoints[\"sglang_gpu\"].gca_resource.dedicated_endpoint_dns\n",
"ENDPOINT_RESOURCE_NAME = endpoints[\"sglang_gpu\"].resource_name\n",
" DEDICATED_ENDPOINT_DNS = endpoint.gca_resource.dedicated_endpoint_dns\n",
"ENDPOINT_RESOURCE_NAME = endpoint.resource_name\n",
"\n",
"# @markdown Because the Qwen3 models generate detailed reasoning steps, the output is expected to be long. We recommend using streaming for a better generation experience.\n",
"# @title Chat Completions Inference\n",
@@ -789,18 +568,12 @@
},
"outputs": [],
"source": [
"# @title Delete the models and endpoints\n",
"# @title Delete the endpoints\n",
"\n",
"# @markdown Delete the experiment models and endpoints to recycle the resources\n",
"# @markdown and avoid unnecessary continuous charges that may incur.\n",
"# @markdown Delete the endpoint.\n",
"\n",
"# Undeploy model and delete endpoint.\n",
"for endpoint in endpoints.values():\n",
" endpoint.delete(force=True)\n",
"\n",
"# Delete models.\n",
"for model in models.values():\n",
" model.delete()"
"if endpoint:\n",
" endpoint.delete(force=True)"
]
}
],
@@ -49,35 +49,48 @@
{
"cell_type": "markdown",
"metadata": {
"id": "cbDI9ag4oR4C"
"id": "3de7470326a2"
},
"source": [
"## Overview\n",
"\n",
"This notebook demonstrates deploying prebuilt [QwQ](https://huggingface.co/collections/Qwen/qwq-674762b79b75eac01735070a) models with [vLLM](https://github.com/vllm-project/vllm) to improve serving throughput.\n",
"This notebook demonstrates how to deploy a **Qwq** open model on Google Cloud Vertex AI.\n",
"\n",
"### Objectives\n",
"\n",
"### Objective\n",
"- Deploy Qwq using containerized backends like [vLLM](https://github.com/vllm-project/vllm) on GPU.\n",
"- Use the deployed model to serve chat completion requests for both text and multimodal inputs.\n",
"\n",
"- Download and deploy prebuilt QwQ models with vLLM.\n",
"### File a Bug\n",
"\n",
"If you encounter issues with this notebook, report them on [GitHub](https://github.com/GoogleCloudPlatform/vertex-ai-samples/issues/new).\n",
"\n",
"### Costs\n",
"\n",
"This tutorial uses billable components of Google Cloud:\n",
"\n",
"* Vertex AI\n",
"* Cloud Storage\n",
"- Vertex AI\n",
"- Cloud Storage\n",
"\n",
"Learn about [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing), [Cloud Storage pricing](https://cloud.google.com/storage/pricing), and use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to generate a cost estimate based on your projected usage."
"Refer to the [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing) and [Cloud Storage pricing](https://cloud.google.com/storage/pricing) pages for more information. Use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to estimate your projected costs."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "hQJWRopioSKT"
"id": "jeYw-Czg-DFy"
},
"source": [
"## Before you begin"
"## Get Started"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "KgyhGvEzBDkj"
},
"source": [
"### Install Vertex AI SDK and other required packages"
]
},
{
@@ -85,119 +98,85 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "J_jmxcIZoSxU"
"id": "iCacdLqG-IsH"
},
"outputs": [],
"source": [
"# @title Setup Google Cloud project\n",
"%pip install --upgrade --force-reinstall --quiet 'google-cloud-aiplatform>=1.106.0' 'openai' 'google-auth==2.27.0' 'requests==2.32.3'"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "HUKCrpBy-3yf"
},
"source": [
"### Authenticate the Notebook Environment (Colab only)\n",
"\n",
"# @markdown 1. [Make sure that billing is enabled for your project](https://cloud.google.com/billing/docs/how-to/modify-project).\n",
"If you're running this notebook in Google Colab, run the following cell to authenticate."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "JXwCT1kn-3Gu"
},
"outputs": [],
"source": [
"import sys\n",
"\n",
"# @markdown 2. **[Optional]** [Create a Cloud Storage bucket](https://cloud.google.com/storage/docs/creating-buckets) for storing experiment outputs. Set the BUCKET_URI for the experiment environment. The specified Cloud Storage bucket (`BUCKET_URI`) should be located in the same region as where the notebook was launched. Note that a multi-region bucket (eg. \"us\") is not considered a match for a single region covered by the multi-region range (eg. \"us-central1\"). If not set, a unique GCS bucket will be created instead.\n",
"if \"google.colab\" in sys.modules:\n",
" from google.colab import auth\n",
"\n",
"BUCKET_URI = \"gs://\" # @param {type:\"string\"}\n",
" auth.authenticate_user()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "AcW2nwB8-7yC"
},
"source": [
"### Set Google Cloud Project Information\n",
"\n",
"# @markdown 3. **[Optional]** Set region. If not set, the region will be set automatically according to Colab Enterprise environment.\n",
"To get started with Vertex AI, ensure you have an existing Google Cloud project and that the [Vertex AI API is enabled](https://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com).\n",
"\n",
"REGION = \"\" # @param {type:\"string\"}\n",
"\n",
"# @markdown 4. If you want to run predictions with A100 80GB or H100 GPUs, we recommend using the regions listed below. **NOTE:** Make sure you have associated quota in selected regions. Click the links to see your current quota for each GPU type: [Nvidia A100 80GB](https://console.cloud.google.com/iam-admin/quotas?metric=aiplatform.googleapis.com%2Fcustom_model_serving_nvidia_a100_80gb_gpus), [Nvidia H100 80GB](https://console.cloud.google.com/iam-admin/quotas?metric=aiplatform.googleapis.com%2Fcustom_model_serving_nvidia_h100_gpus). You can request for quota following the instructions at [\"Request a higher quota\"](https://cloud.google.com/docs/quota/view-manage#requesting_higher_quota).\n",
"\n",
"# @markdown > | Machine Type | Accelerator Type | Recommended Regions |\n",
"# @markdown | ----------- | ----------- | ----------- |\n",
"# @markdown | a2-ultragpu-1g | 1 NVIDIA_A100_80GB | us-central1, us-east4, europe-west4, asia-southeast1, us-east4 |\n",
"# @markdown | a3-highgpu-2g | 2 NVIDIA_H100_80GB | us-west1, asia-southeast1, europe-west4 |\n",
"# @markdown | a3-highgpu-4g | 4 NVIDIA_H100_80GB | us-west1, asia-southeast1, europe-west4 |\n",
"# @markdown | a3-highgpu-8g | 8 NVIDIA_H100_80GB | us-central1, europe-west4, us-west1, asia-southeast1 |\n",
"\n",
"# Upgrade Vertex AI SDK.\n",
"! pip3 install --upgrade --quiet 'google-cloud-aiplatform>=1.64.0'\n",
"\n",
"# Import the necessary packages\n",
"import datetime\n",
"import importlib\n",
"See the guide on [setting up your project and development environment](https://cloud.google.com/vertex-ai/docs/start/cloud-environment). Also confirm that [billing is enabled](https://cloud.google.com/billing/docs/how-to/modify-project).\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "eIVLp0oE--k-"
},
"outputs": [],
"source": [
"# Use the environment variable if the user doesn't provide Project ID.\n",
"import os\n",
"import uuid\n",
"from typing import Tuple\n",
"\n",
"from google.cloud import aiplatform\n",
"import vertexai\n",
"\n",
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"PROJECT_ID = \"[your-project-id]\" # @param {type: \"string\", placeholder: \"[your-project-id]\", isTemplate: true}\n",
"\n",
"models, endpoints = {}, {}\n",
"if not PROJECT_ID or PROJECT_ID == \"[your-project-id]\":\n",
" PROJECT_ID = str(os.environ.get(\"GOOGLE_CLOUD_PROJECT\"))\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
")\n",
"REGION = os.environ.get(\"GOOGLE_CLOUD_REGION\", \"us-central1\")\n",
"\n",
"# Get the default cloud project id.\n",
"PROJECT_ID = os.environ[\"GOOGLE_CLOUD_PROJECT\"]\n",
"\n",
"# Get the default region for launching jobs.\n",
"if not REGION:\n",
" if not os.environ.get(\"GOOGLE_CLOUD_REGION\"):\n",
" raise ValueError(\n",
" \"REGION must be set. See\"\n",
" \" https://cloud.google.com/vertex-ai/docs/general/locations for\"\n",
" \" available cloud locations.\"\n",
" )\n",
" REGION = os.environ[\"GOOGLE_CLOUD_REGION\"]\n",
"\n",
"# Enable the Vertex AI API and Compute Engine API, if not already.\n",
"print(\"Enabling Vertex AI API and Compute Engine API.\")\n",
"! gcloud services enable aiplatform.googleapis.com compute.googleapis.com\n",
"\n",
"# Cloud Storage bucket for storing the experiment artifacts.\n",
"# A unique GCS bucket will be created for the purpose of this notebook. If you\n",
"# prefer using your own GCS bucket, change the value yourself below.\n",
"now = datetime.datetime.now().strftime(\"%Y%m%d%H%M%S\")\n",
"BUCKET_NAME = \"/\".join(BUCKET_URI.split(\"/\")[:3])\n",
"\n",
"if BUCKET_URI is None or BUCKET_URI.strip() == \"\" or BUCKET_URI == \"gs://\":\n",
" BUCKET_URI = f\"gs://{PROJECT_ID}-tmp-{now}-{str(uuid.uuid4())[:4]}\"\n",
" BUCKET_NAME = \"/\".join(BUCKET_URI.split(\"/\")[:3])\n",
" ! gsutil mb -l {REGION} {BUCKET_URI}\n",
"else:\n",
" assert BUCKET_URI.startswith(\"gs://\"), \"BUCKET_URI must start with `gs://`.\"\n",
" shell_output = ! gsutil ls -Lb {BUCKET_NAME} | grep \"Location constraint:\" | sed \"s/Location constraint://\"\n",
" bucket_region = shell_output[0].strip().lower()\n",
" if bucket_region != REGION:\n",
" raise ValueError(\n",
" \"Bucket region %s is different from notebook region %s\"\n",
" % (bucket_region, REGION)\n",
" )\n",
"print(f\"Using this GCS Bucket: {BUCKET_URI}\")\n",
"\n",
"STAGING_BUCKET = os.path.join(BUCKET_URI, \"temporal\")\n",
"MODEL_BUCKET = os.path.join(BUCKET_URI, \"qwq\")\n",
"\n",
"\n",
"# Initialize Vertex AI API.\n",
"print(\"Initializing Vertex AI API.\")\n",
"aiplatform.init(project=PROJECT_ID, location=REGION, staging_bucket=STAGING_BUCKET)\n",
"\n",
"# Gets the default SERVICE_ACCOUNT.\n",
"shell_output = ! gcloud projects describe $PROJECT_ID\n",
"project_number = shell_output[-1].split(\":\")[1].strip().replace(\"'\", \"\")\n",
"SERVICE_ACCOUNT = f\"{project_number}-compute@developer.gserviceaccount.com\"\n",
"print(\"Using this default Service Account:\", SERVICE_ACCOUNT)\n",
"\n",
"\n",
"# Provision permissions to the SERVICE_ACCOUNT with the GCS bucket\n",
"! gsutil iam ch serviceAccount:{SERVICE_ACCOUNT}:roles/storage.admin $BUCKET_NAME\n",
"\n",
"! gcloud config set project $PROJECT_ID\n",
"! gcloud projects add-iam-policy-binding --no-user-output-enabled {PROJECT_ID} --member=serviceAccount:{SERVICE_ACCOUNT} --role=\"roles/storage.admin\"\n",
"! gcloud projects add-iam-policy-binding --no-user-output-enabled {PROJECT_ID} --member=serviceAccount:{SERVICE_ACCOUNT} --role=\"roles/aiplatform.user\""
"vertexai.init(project=PROJECT_ID, location=REGION)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "z9UuiysLu_gB"
"id": "Q0CXrvcZH_aw"
},
"source": [
"## Deploy prebuilt QwQ models on vLLM"
"### Import libraries"
]
},
{
@@ -205,206 +184,202 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "USB7dvYqvNdu"
"id": "3G2UXB82ICs6"
},
"outputs": [],
"source": [
"# @title Deploy\n",
"from vertexai import model_garden"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "upYRiGtP_-iN"
},
"source": [
"## Deploy model"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "H2WC_0hXDVXc"
},
"source": [
"### Choose model variant"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "u41zbNa2EoFq"
},
"source": [
"You can proceed with the default model variant or select a different one."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "-fgC4NLSDkF7"
},
"outputs": [],
"source": [
"model_version = \"qwq-32b\" # @param [\"qwq-32b\"] {isTemplate:true}\n",
"MODEL_NAME = f\"qwen/qwq@{model_version}\""
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "VRnUgU8LF3_i"
},
"source": [
"To see all deployable model variants available in Model Garden, use:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "-QLd-wshF6sB"
},
"outputs": [],
"source": [
"all_model_versions = model_garden.list_deployable_models(\n",
" model_filter=\"qwq\", list_hf_models=False\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "N0UeFHa2GO63"
},
"source": [
"Once you've selected a model variant, initialize it:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "GZiV3trBBcA3"
},
"outputs": [],
"source": [
"model = model_garden.OpenModel(MODEL_NAME)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "-0cL378wFlvf"
},
"source": [
"### Check the Deployment Configuration\n",
"\n",
"# @markdown This section uploads prebuilt QwQ models to Model Registry and deploys it to a Vertex AI Endpoint. It takes 15 to 30 minutes to finish depending on the size of the model.\n",
"Use the `list_deploy_options()` method to view the verified deployment configurations for your selected model. This helps ensure you have sufficient resources (e.g., GPU quota) available to deploy it."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "zm73g7vFFm9N"
},
"outputs": [],
"source": [
"deploy_options = model.list_deploy_options(concise=True)\n",
"print(deploy_options)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "WjV499VsGwrD"
},
"source": [
"### Deploy the Model\n",
"\n",
"MODEL_ID = \"QwQ-32B\" # @param [\"QwQ-32B\"] {isTemplate: true}\n",
"model_path_prefix = \"Qwen\"\n",
"model_id = os.path.join(model_path_prefix, MODEL_ID)\n",
"Now that you’ve reviewed the deployment options, use the `deploy()` method to serve the selected open model to a Vertex AI endpoint. Deployment time may vary depending on the model size and infrastructure requirements.\n",
"\n",
"# The pre-built serving docker image for vLLM.\n",
"VLLM_DOCKER_URI = \"us-docker.pkg.dev/vertex-ai/vertex-vision-model-garden-dockers/pytorch-vllm-serve:20250506_0916_RC01\"\n",
"\n",
"# @markdown Set `use_dedicated_endpoint` to False if you don't want to use [dedicated endpoint](https://cloud.google.com/vertex-ai/docs/general/deployment#create-dedicated-endpoint).\n",
"use_dedicated_endpoint = True # @param {type:\"boolean\"}\n",
"\n",
"accelerator_type = \"NVIDIA_L4\" # @param [\"NVIDIA_L4\", \"NVIDIA_H100_80GB\"] {isTemplate: true}\n",
"\n",
"# @markdown For inputs exceeding 8,192 tokens, enable [YaRN](https://arxiv.org/abs/2309.00071) on the deployment to improve the model's ability to capture long-sequence information effectively.\n",
"# @markdown Enabling YaRN will allow for 128k context-length but may require more GPUs.\n",
"enable_yarn_scaling = False # @param {type:\"boolean\"}\n",
"\n",
"if enable_yarn_scaling:\n",
" max_model_len = 131072\n",
"else:\n",
" max_model_len = 32768\n",
"\n",
"if accelerator_type == \"NVIDIA_L4\":\n",
" if enable_yarn_scaling:\n",
" accelerator_count = 8\n",
" machine_type = \"g2-standard-96\"\n",
" else:\n",
" accelerator_count = 4\n",
" machine_type = \"g2-standard-48\"\n",
"elif accelerator_type == \"NVIDIA_H100_80GB\":\n",
" accelerator_count = 2\n",
" machine_type = \"a3-highgpu-2g\"\n",
"else:\n",
" raise ValueError(\n",
" \"Recommended machine settings not found for accelerator type: %s\"\n",
" % accelerator_type\n",
" )\n",
"\n",
"\n",
"common_util.check_quota(\n",
" project_id=PROJECT_ID,\n",
" region=REGION,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
" is_for_training=False,\n",
")\n",
"\n",
"\n",
"def deploy_model_vllm(\n",
" model_name: str,\n",
" model_id: str,\n",
" publisher: str,\n",
" publisher_model_id: str,\n",
" service_account: str,\n",
" base_model_id: str = None,\n",
" machine_type: str = \"g2-standard-8\",\n",
" accelerator_type: str = \"NVIDIA_L4\",\n",
" accelerator_count: int = 1,\n",
" gpu_memory_utilization: float = 0.9,\n",
" max_model_len: int = 4096,\n",
" dtype: str = \"auto\",\n",
" enable_trust_remote_code: bool = False,\n",
" enforce_eager: bool = False,\n",
" enable_lora: bool = False,\n",
" enable_chunked_prefill: bool = False,\n",
" enable_prefix_cache: bool = False,\n",
" host_prefix_kv_cache_utilization_target: float = 0.0,\n",
" max_loras: int = 1,\n",
" max_cpu_loras: int = 8,\n",
" use_dedicated_endpoint: bool = False,\n",
" max_num_seqs: int = 256,\n",
" model_type: str = None,\n",
" enable_yarn_scaling: bool = False,\n",
") -> Tuple[aiplatform.Model, aiplatform.Endpoint]:\n",
" \"\"\"Deploys trained models with vLLM into Vertex AI.\"\"\"\n",
" endpoint = aiplatform.Endpoint.create(\n",
" display_name=f\"{model_name}-endpoint\",\n",
" dedicated_endpoint_enabled=use_dedicated_endpoint,\n",
" )\n",
"\n",
" if not base_model_id:\n",
" base_model_id = model_id\n",
"\n",
" # See https://docs.vllm.ai/en/latest/models/engine_args.html for a list of possible arguments with descriptions.\n",
" vllm_args = [\n",
" \"python\",\n",
" \"-m\",\n",
" \"vllm.entrypoints.api_server\",\n",
" \"--host=0.0.0.0\",\n",
" \"--port=8080\",\n",
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={accelerator_count}\",\n",
" \"--swap-space=16\",\n",
" f\"--gpu-memory-utilization={gpu_memory_utilization}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" f\"--dtype={dtype}\",\n",
" f\"--max-loras={max_loras}\",\n",
" f\"--max-cpu-loras={max_cpu_loras}\",\n",
" f\"--max-num-seqs={max_num_seqs}\",\n",
" \"--disable-log-stats\",\n",
" ]\n",
"\n",
" if enable_trust_remote_code:\n",
" vllm_args.append(\"--trust-remote-code\")\n",
"\n",
" if enforce_eager:\n",
" vllm_args.append(\"--enforce-eager\")\n",
"\n",
" if enable_lora:\n",
" vllm_args.append(\"--enable-lora\")\n",
"\n",
" if enable_chunked_prefill:\n",
" vllm_args.append(\"--enable-chunked-prefill\")\n",
"\n",
" if enable_prefix_cache:\n",
" vllm_args.append(\"--enable-prefix-caching\")\n",
"\n",
" if 0 < host_prefix_kv_cache_utilization_target < 1:\n",
" vllm_args.append(\n",
" f\"--host-prefix-kv-cache-utilization-target={host_prefix_kv_cache_utilization_target}\"\n",
" )\n",
"\n",
" if enable_yarn_scaling:\n",
" vllm_args.append(\n",
" '--rope-scaling=\\'{\"factor\": 4.0, \"original_max_position_embeddings\": 32768, \"rope_type\": \"yarn\"}\\''\n",
" )\n",
"\n",
" if model_type:\n",
" vllm_args.append(f\"--model-type={model_type}\")\n",
"\n",
" env_vars = {\n",
" \"MODEL_ID\": base_model_id,\n",
" \"DEPLOY_SOURCE\": common_util.get_deploy_source(),\n",
" }\n",
"\n",
" # HF_TOKEN is not a compulsory field and may not be defined.\n",
" try:\n",
" if HF_TOKEN:\n",
" env_vars[\"HF_TOKEN\"] = HF_TOKEN\n",
" except NameError:\n",
" pass\n",
"\n",
" model = aiplatform.Model.upload(\n",
" display_name=model_name,\n",
" serving_container_image_uri=VLLM_DOCKER_URI,\n",
" serving_container_args=vllm_args,\n",
" serving_container_ports=[8080],\n",
" serving_container_predict_route=\"/generate\",\n",
" serving_container_health_route=\"/ping\",\n",
" serving_container_environment_variables=env_vars,\n",
" serving_container_shared_memory_size_mb=(16 * 1024), # 16 GB\n",
" serving_container_deployment_timeout=7200,\n",
" model_garden_source_model_name=(\n",
" f\"publishers/{publisher}/models/{publisher_model_id}\"\n",
" ),\n",
" )\n",
" print(\n",
" f\"Deploying {model_name} on {machine_type} with {accelerator_count} {accelerator_type} GPU(s).\"\n",
" )\n",
" model.deploy(\n",
" endpoint=endpoint,\n",
" machine_type=machine_type,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
" deploy_request_timeout=1800,\n",
" service_account=service_account,\n",
" system_labels={\n",
" \"NOTEBOOK_NAME\": \"model_garden_pytorch_qwq_deployment.ipynb\",\n",
" },\n",
" )\n",
" print(\"endpoint_name:\", endpoint.name)\n",
"\n",
" return model, endpoint\n",
"\n",
"\n",
"models[\"vllm_gpu\"], endpoints[\"vllm_gpu\"] = deploy_model_vllm(\n",
" model_name=common_util.get_job_name_with_datetime(prefix=MODEL_ID),\n",
" model_id=model_id,\n",
" publisher=\"qwen\",\n",
" publisher_model_id=\"qwq\",\n",
" service_account=SERVICE_ACCOUNT,\n",
" machine_type=machine_type,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
" max_model_len=max_model_len,\n",
" gpu_memory_utilization=0.9,\n",
" enable_chunked_prefill=True,\n",
"> **Note**: If the model requires accepting a license agreement (EULA), set the `accept_eula=True` flag in the deploy call. Set `use_dedicated_endpoint` to False if you don't want to use [dedicated endpoint](https://cloud.google.com/vertex-ai/docs/general/deployment#create-dedicated-endpoint)."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "wX1itVTvXdEP"
},
"outputs": [],
"source": [
"use_dedicated_endpoint = True"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "MRmPFEPoGzsB"
},
"outputs": [],
"source": [
"endpoint = model.deploy(\n",
" accept_eula=True,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
" max_num_seqs=128,\n",
" enable_yarn_scaling=enable_yarn_scaling,\n",
")\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "PHBtn8DQp-ID"
},
"source": [
"Alternatively, you can select one of the verified deployment configurations listed above."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "ADsJG8JYqI6c"
},
"outputs": [],
"source": [
"endpoint = model.deploy(\n",
" accept_eula=True,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
" serving_container_image_uri=\"us-docker.pkg.dev/vertex-ai/vertex-vision-model-garden-dockers/pytorch-vllm-serve:20250506_0916_RC01\",\n",
" machine_type=\"a3-highgpu-2g\",\n",
" accelerator_type=\"NVIDIA_H100_80GB\",\n",
" accelerator_count=2,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "kqSUK2CwsImi"
},
"source": [
"To further customize your deployment, you can configure:\n",
"\n",
"# @markdown Click \"Show Code\" to see more details."
"- **Compute Resources**: Machine type, replica count (min/max), accelerator type and quantity.\n",
"- **Infrastructure**: Use Spot VMs, reservation affinity, or dedicated endpoints.\n",
"- **Serving Container**: Customize container image, ports, health checks, and environment variables.\n",
"\n",
"See the [Model Garden SDK README](https://github.com/googleapis/python-aiplatform/blob/main/vertexai/model_garden/README.md) for advanced configuration options."
]
},
{
@@ -475,7 +450,7 @@
" \"top_k\": top_k,\n",
" },\n",
"]\n",
"response = endpoints[\"vllm_gpu\"].predict(\n",
"response = endpoint.predict(\n",
" instances=instances, use_dedicated_endpoint=use_dedicated_endpoint\n",
")\n",
"\n",
@@ -503,20 +478,11 @@
"outputs": [],
"source": [
"# @title Delete the models and endpoints\n",
"# @markdown Delete the experiment models and endpoints to recycle the resources\n",
"# @markdown and avoid unnecessary continuous charges that may incur.\n",
"\n",
"# Undeploy model and delete endpoint.\n",
"for endpoint in endpoints.values():\n",
" endpoint.delete(force=True)\n",
"# @markdown Delete the endpoint.\n",
"\n",
"# Delete models.\n",
"for model in models.values():\n",
" model.delete()\n",
"\n",
"delete_bucket = False # @param {type:\"boolean\"}\n",
"if delete_bucket:\n",
" ! gsutil -m rm -r $BUCKET_NAME"
"if endpoint:\n",
" endpoint.delete(force=True)"
]
}
],
@@ -54,45 +54,48 @@
{
"cell_type": "markdown",
"metadata": {
"id": "JbmPgTp2LRCY"
"id": "3de7470326a2"
},
"source": [
"## Overview\n",
"\n",
"This notebook demonstrates using the [huggingface/transformers](https://github.com/huggingface/transformers) framework to serve Segment Anything Model (SAM) models and deploy them for online prediction on Vertex AI.\n",
"This notebook demonstrates how to deploy a **Segment-Anything** open model on Google Cloud Vertex AI.\n",
"\n",
"Following the notebook you will conduct experiments using the pre-built docker image on Vertex AI.\n",
"### Objectives\n",
"\n",
"- With the pre-built docker images, you can **deploy** models for the following tasks:\n",
" - Mask Generation\n",
"- Deploy Segment-Anything using containerized backends like [vLLM](https://github.com/vllm-project/vllm) on GPU.\n",
"- Use the deployed model to serve chat completion requests for both text and multimodal inputs.\n",
"\n",
"### Objective\n",
"### File a Bug\n",
"\n",
"- Upload the model to [Model Registry](https://cloud.google.com/vertex-ai/docs/model-registry/introduction).\n",
"- Deploy the model on [Endpoint](https://cloud.google.com/vertex-ai/docs/predictions/using-private-endpoints).\n",
"- Run online predictions for image captioning.\n",
"\n",
"### File a bug\n",
"\n",
"File a bug on [GitHub](https://github.com/GoogleCloudPlatform/vertex-ai-samples/issues/new) if you encounter any issue with the notebook.\n",
"If you encounter issues with this notebook, report them on [GitHub](https://github.com/GoogleCloudPlatform/vertex-ai-samples/issues/new).\n",
"\n",
"### Costs\n",
"\n",
"This tutorial uses billable components of Google Cloud:\n",
"\n",
"* Vertex AI\n",
"* Cloud Storage\n",
"- Vertex AI\n",
"- Cloud Storage\n",
"\n",
"Learn about [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing), [Cloud Storage pricing](https://cloud.google.com/storage/pricing), and use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to generate a cost estimate based on your projected usage."
"Refer to the [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing) and [Cloud Storage pricing](https://cloud.google.com/storage/pricing) pages for more information. Use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to estimate your projected costs."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "78f72e0a-52e5-4de5-ac0f-2171b3493825"
"id": "jeYw-Czg-DFy"
},
"source": [
"## Before you begin"
"## Get Started"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "KgyhGvEzBDkj"
},
"source": [
"### Install Vertex AI SDK and other required packages"
]
},
{
@@ -100,63 +103,85 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "m-Ql0m9edvcA"
"id": "iCacdLqG-IsH"
},
"outputs": [],
"source": [
"# @title Setup Google Cloud project\n",
"%pip install --upgrade --force-reinstall --quiet 'google-cloud-aiplatform>=1.106.0' 'openai' 'google-auth==2.27.0' 'requests==2.32.3'"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "HUKCrpBy-3yf"
},
"source": [
"### Authenticate the Notebook Environment (Colab only)\n",
"\n",
"# @markdown 1. [Make sure that billing is enabled for your project](https://cloud.google.com/billing/docs/how-to/modify-project).\n",
"If you're running this notebook in Google Colab, run the following cell to authenticate."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "JXwCT1kn-3Gu"
},
"outputs": [],
"source": [
"import sys\n",
"\n",
"# @markdown 2. **[Optional]** Set region. If not set, the region will be set automatically according to Colab Enterprise environment.\n",
"if \"google.colab\" in sys.modules:\n",
" from google.colab import auth\n",
"\n",
"REGION = \"\" # @param {type:\"string\"}\n",
" auth.authenticate_user()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "AcW2nwB8-7yC"
},
"source": [
"### Set Google Cloud Project Information\n",
"\n",
"# Upgrade Vertex AI SDK.\n",
"! pip3 install --upgrade --quiet 'google-cloud-aiplatform==1.103.0'\n",
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"To get started with Vertex AI, ensure you have an existing Google Cloud project and that the [Vertex AI API is enabled](https://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com).\n",
"\n",
"import gc\n",
"import importlib\n",
"See the guide on [setting up your project and development environment](https://cloud.google.com/vertex-ai/docs/start/cloud-environment). Also confirm that [billing is enabled](https://cloud.google.com/billing/docs/how-to/modify-project).\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "eIVLp0oE--k-"
},
"outputs": [],
"source": [
"# Use the environment variable if the user doesn't provide Project ID.\n",
"import os\n",
"\n",
"import matplotlib.pyplot as plt\n",
"import numpy as np\n",
"from google.cloud import aiplatform\n",
"\n",
"if os.environ.get(\"VERTEX_PRODUCT\") != \"COLAB_ENTERPRISE\":\n",
" ! pip install --upgrade tensorflow\n",
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
")\n",
"\n",
"models, endpoints = {}, {}\n",
"\n",
"\n",
"# Get the default cloud project id.\n",
"PROJECT_ID = os.environ[\"GOOGLE_CLOUD_PROJECT\"]\n",
"\n",
"# Get the default region for launching jobs.\n",
"if not REGION:\n",
" REGION = os.environ[\"GOOGLE_CLOUD_REGION\"]\n",
"\n",
"# Initialize Vertex AI API.\n",
"print(\"Initializing Vertex AI API.\")\n",
"aiplatform.init(project=PROJECT_ID, location=REGION)\n",
"\n",
"! gcloud config set project $PROJECT_ID\n",
"\n",
"import vertexai\n",
"\n",
"vertexai.init(\n",
" project=PROJECT_ID,\n",
" location=REGION,\n",
")\n",
"PROJECT_ID = \"[your-project-id]\" # @param {type: \"string\", placeholder: \"[your-project-id]\", isTemplate: true}\n",
"\n",
"base_model_name = \"sam-vit-large\"\n",
"PUBLISHER_MODEL_NAME = f\"publishers/meta/models/segment-anything@{base_model_name}\""
"if not PROJECT_ID or PROJECT_ID == \"[your-project-id]\":\n",
" PROJECT_ID = str(os.environ.get(\"GOOGLE_CLOUD_PROJECT\"))\n",
"\n",
"REGION = os.environ.get(\"GOOGLE_CLOUD_REGION\", \"us-central1\")\n",
"\n",
"vertexai.init(project=PROJECT_ID, location=REGION)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "Q0CXrvcZH_aw"
},
"source": [
"### Import libraries"
]
},
{
@@ -164,68 +189,38 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "ot2HhTqxRYri"
"id": "3G2UXB82ICs6"
},
"outputs": [],
"source": [
"# @title Deploy\n",
"\n",
"# The pre-built serving docker image.\n",
"# The model artifacts are embedded within the container, except for model weights which will be downloaded during deployment.\n",
"SERVE_DOCKER_URI = \"us-docker.pkg.dev/deeplearning-platform-release/vertex-model-garden/pytorch-inference.cu125.0-4.ubuntu2204.py310\"\n",
"\n",
"# @markdown Set the accelerator type.\n",
"accelerator_type = \"NVIDIA_L4\" # @param[\"NVIDIA_TESLA_V100\", \"NVIDIA_L4\"]\n",
"\n",
"if accelerator_type == \"NVIDIA_TESLA_V100\":\n",
" machine_type = \"n1-standard-8\"\n",
" accelerator_count = 1\n",
"elif accelerator_type == \"NVIDIA_L4\":\n",
" machine_type = \"g2-standard-12\"\n",
" accelerator_count = 1\n",
"else:\n",
" print(f\"Unsupported accelerator type: {accelerator_type}\")\n",
"\n",
"MODEL_ID = \"facebook/sam-vit-large\"\n",
"task = \"mask-generation\"\n",
"\n",
"# @markdown Set use_dedicated_endpoint to False if you don't want to use [dedicated endpoint](https://cloud.google.com/vertex-ai/docs/general/deployment#create-dedicated-endpoint). Note that [dedicated endpoint does not support VPC Service Controls](https://cloud.google.com/vertex-ai/docs/predictions/choose-endpoint-type), uncheck the box if you are using VPC-SC.\n",
"use_dedicated_endpoint = True # @param {type:\"boolean\"}\n",
"\n",
"\n",
"def deploy_model(\n",
" task, display_name, model_id, machine_type, accelerator_type, accelerator_count\n",
"):\n",
" endpoint = aiplatform.Endpoint.create(\n",
" display_name=common_util.get_job_name_with_datetime(prefix=task),\n",
" dedicated_endpoint_enabled=use_dedicated_endpoint,\n",
" )\n",
" serving_env = {\n",
" \"MODEL_ID\": model_id,\n",
" \"TASK\": task,\n",
" \"DEPLOY_SOURCE\": \"notebook\",\n",
" }\n",
" model = aiplatform.Model.upload(\n",
" display_name=task,\n",
" serving_container_image_uri=SERVE_DOCKER_URI,\n",
" serving_container_ports=[7080],\n",
" serving_container_predict_route=\"/predict\",\n",
" serving_container_health_route=\"/ping\",\n",
" serving_container_environment_variables=serving_env,\n",
" model_garden_source_model_name=\"publishers/meta/models/segment-anything\",\n",
" )\n",
" model.deploy(\n",
" endpoint=endpoint,\n",
" machine_type=machine_type,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=1,\n",
" deploy_request_timeout=1800,\n",
" system_labels={\n",
" \"NOTEBOOK_NAME\": \"model_garden_pytorch_sam.ipynb\",\n",
" \"NOTEBOOK_ENVIRONMENT\": common_util.get_deploy_source(),\n",
" },\n",
" )\n",
" return model, endpoint"
"from vertexai import model_garden"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "upYRiGtP_-iN"
},
"source": [
"## Deploy model"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "H2WC_0hXDVXc"
},
"source": [
"### Choose model variant"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "u41zbNa2EoFq"
},
"source": [
"You can proceed with the default model variant or select a different one."
]
},
{
@@ -233,26 +228,129 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "HF8tcQUFG7jB"
"id": "-fgC4NLSDkF7"
},
"outputs": [],
"source": [
"# @title [Option 1] Deploy with Model Garden SDK\n",
"model_version = \"sam-vit-large\" # @param [\"sam-vit-large\"] {isTemplate:true}\n",
"MODEL_NAME = f\"meta/segment-anything@{model_version}\""
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "VRnUgU8LF3_i"
},
"source": [
"To see all deployable model variants available in Model Garden, use:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "-QLd-wshF6sB"
},
"outputs": [],
"source": [
"all_model_versions = model_garden.list_deployable_models(\n",
" model_filter=\"segment-anything\", list_hf_models=False\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "N0UeFHa2GO63"
},
"source": [
"Once you've selected a model variant, initialize it:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "GZiV3trBBcA3"
},
"outputs": [],
"source": [
"model = model_garden.OpenModel(MODEL_NAME)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "-0cL378wFlvf"
},
"source": [
"### Check the Deployment Configuration\n",
"\n",
"LABEL = \"sdk-deploy\"\n",
"# @markdown Deploy with Gen AI model-centric SDK. This section uploads the prebuilt model to Model Registry and deploys it to a Vertex AI Endpoint. It takes 15 minutes to 1 hour to finish depending on the size of the model. See [use open models with Vertex AI](https://cloud.google.com/vertex-ai/generative-ai/docs/open-models/use-open-models) for documentation on other use cases.\n",
"from vertexai import model_garden\n",
"Use the `list_deploy_options()` method to view the verified deployment configurations for your selected model. This helps ensure you have sufficient resources (e.g., GPU quota) available to deploy it."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "zm73g7vFFm9N"
},
"outputs": [],
"source": [
"deploy_options = model.list_deploy_options(concise=True)\n",
"print(deploy_options)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "WjV499VsGwrD"
},
"source": [
"### Deploy the Model\n",
"\n",
"model = model_garden.OpenModel(PUBLISHER_MODEL_NAME)\n",
"endpoints[LABEL] = model.deploy(\n",
" machine_type=machine_type,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
"Now that you’ve reviewed the deployment options, use the `deploy()` method to serve the selected open model to a Vertex AI endpoint. Deployment time may vary depending on the model size and infrastructure requirements.\n",
"\n",
"> **Note**: If the model requires accepting a license agreement (EULA), set the `accept_eula=True` flag in the deploy call. Set `use_dedicated_endpoint` to False if you don't want to use [dedicated endpoint](https://cloud.google.com/vertex-ai/docs/general/deployment#create-dedicated-endpoint)."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "wX1itVTvXdEP"
},
"outputs": [],
"source": [
"use_dedicated_endpoint = True"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "MRmPFEPoGzsB"
},
"outputs": [],
"source": [
"endpoint = model.deploy(\n",
" accept_eula=True,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
" accept_eula=True, # Accept the End User License Agreement (EULA) on the model card before deploy. Otherwise, the deployment will be forbidden.\n",
")\n",
"\n",
"endpoint = endpoints[LABEL]"
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "PHBtn8DQp-ID"
},
"source": [
"Alternatively, you can select one of the verified deployment configurations listed above."
]
},
{
@@ -260,36 +358,33 @@
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "lyZSkLIwG7jB"
"id": "ADsJG8JYqI6c"
},
"outputs": [],
"source": [
"# @title [Option 2] Deploy with customized configs\n",
"\n",
"# @markdown This section deploys a pre-trained `sam-vit-large` model on Model Registry by using 1 L4 Machine.\n",
"\n",
"# @markdown The model deploy step will take around 20 minutes to complete.\n",
"\n",
"common_util.check_quota(\n",
" project_id=PROJECT_ID,\n",
" region=REGION,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
" is_for_training=False,\n",
")\n",
"\n",
"LABEL = \"custom-deploy\"\n",
"models[LABEL], endpoints[LABEL] = deploy_model(\n",
" task=task,\n",
" display_name=common_util.get_job_name_with_datetime(prefix=task),\n",
" model_id=\"facebook/sam-vit-large\",\n",
" machine_type=machine_type,\n",
" accelerator_type=accelerator_type,\n",
"endpoint = model.deploy(\n",
" accept_eula=True,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
" serving_container_image_uri=\"us-docker.pkg.dev/deeplearning-platform-release/vertex-model-garden/pytorch-inference.cu125.0-4.ubuntu2204.py310\",\n",
" machine_type=\"n1-standard-8\",\n",
" accelerator_type=\"NVIDIA_TESLA_V100\",\n",
" accelerator_count=1,\n",
")\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "kqSUK2CwsImi"
},
"source": [
"To further customize your deployment, you can configure:\n",
"\n",
"model = models[LABEL]\n",
"endpoint = endpoints[LABEL]"
"- **Compute Resources**: Machine type, replica count (min/max), accelerator type and quantity.\n",
"- **Infrastructure**: Use Spot VMs, reservation affinity, or dedicated endpoints.\n",
"- **Serving Container**: Customize container image, ports, health checks, and environment variables.\n",
"\n",
"See the [Model Garden SDK README](https://github.com/googleapis/python-aiplatform/blob/main/vertexai/model_garden/README.md) for advanced configuration options."
]
},
{
@@ -303,6 +398,20 @@
"source": [
"# @title Predict\n",
"\n",
"import gc\n",
"import importlib\n",
"\n",
"import matplotlib.pyplot as plt\n",
"import numpy as np\n",
"\n",
"# Import the necessary packages.\n",
"! rm -rf vertex-ai-samples && git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"! cd vertex-ai-samples\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"input_image1 = \"http://images.cocodataset.org/val2017/000000039769.jpg\" # @param {type:\"string\"}\n",
"input_image2 = \"http://images.cocodataset.org/val2017/000000000285.jpg\" # @param {type:\"string\"}\n",
"\n",
@@ -368,16 +477,12 @@
},
"outputs": [],
"source": [
"# @markdown Delete the experiment models and endpoints to recycle the resources\n",
"# @markdown and avoid unnecessary continuous charges that may incur.\n",
"# @title Delete the models and endpoints\n",
"\n",
"# Undeploy model and delete endpoint.\n",
"for endpoint in endpoints.values():\n",
" endpoint.delete(force=True)\n",
"# @markdown Delete the endpoint.\n",
"\n",
"# Delete models.\n",
"for model in models.values():\n",
" model.delete()"
"if endpoint:\n",
" endpoint.delete(force=True)"
]
}
],
@@ -121,7 +121,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"\n",
@@ -124,7 +124,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"\n",
@@ -132,7 +132,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"\n",
@@ -124,7 +124,7 @@
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.community-content.vertex_model_garden.model_oss.notebook_util.common_util\"\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"# Get the default cloud project id.\n",

Some files were not shown because too many files have changed in this diff Show More