Compare commits

...
Author SHA1 Message Date
Aaron DietzandGitHub 9fc480b81d Update notebook_template_review.py
Minor change to help us update references to "custom training" to specify "serverless training"
2025-12-17 14:01:17 -08:00
Sam-DecigaandGitHub babeba9f02 feat:NVIDIA Nemotron Nano v2 12B VL - 2025-12-TBD (#4391)
* feat:NVIDIA Nemotron Nano v2 12B VL - 2025-12-TBD

* refactor:Reformat Notebook

* refactor:Reformat Notebook
2025-12-17 18:49:58 +00:00
Vertex MG TeamandCopybara-Service 85c649dd26 Add Llama 3.3 TPU7x deployment notebook.
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 845481841
2025-12-16 16:50:45 -08:00
Vertex MG TeamandCopybara-Service b7135ae1f0 Add Llama 3.3 TPU7x deployment notebook.
PiperOrigin-RevId: 845481841
2025-12-16 16:38:07 -08:00
Vertex MG TeamandCopybara-Service 2f5119a266 Allows the user to select spot VM for deployment
PiperOrigin-RevId: 845064375
2025-12-15 21:23:11 -08:00
Vertex MG TeamandCopybara-Service 23e64ca76f fix: Update TimesFM 2.0 deployment notebook to use GCS path as MODEL_ID
PiperOrigin-RevId: 844826585
2025-12-15 10:25:10 -08:00
gurusai-voletiandGitHub 0ba5a62cc9 Fix ci workflow to use python 3.13 to avoid linter issues (#4397)
* update

* use python 3.13
2025-12-15 13:54:49 +00:00
Damodar PanigrahiandGitHub 0be2c6fd0c feat: authenticate using sa (#4394) 2025-12-12 15:38:25 -05:00
Vertex MG TeamandCopybara-Service cef4928c49 Allows the user to select spot VM for deployment
PiperOrigin-RevId: 843098639
2025-12-11 01:03:35 -08:00
Vertex MG TeamandCopybara-Service 9bb8107110 Add notebook for using Deepseek 3.2 model on Vertex AI.
PiperOrigin-RevId: 842763792
2025-12-10 09:42:39 -08:00
Vertex MG TeamandCopybara-Service b075990d88 Weekly update the vllm/hf-tei/hf-inference-toolkit containers.
PiperOrigin-RevId: 842339077
2025-12-09 12:04:54 -08:00
Ravi DalalGitHubgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
6a83c4c695 Updated notebook comment for custom vllm container image (#4385)
* updated comment for custom vllm container image

* Update notebooks/official/prediction/vertexai_serving_vllm/vertexai_serving_vllm_cpu_llama3_2_3B.ipynb

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

---------

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-12-05 20:26:32 +00:00
Damodar PanigrahiGitHubgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
820c0f8db4 Update Use Case description and API change to accept GCP auth token (#4384)
* pass auth_token in create_http_client

* lint on the notebook

* feat:Removed the last update date

* Update notebooks/community/alphagenome/README.md

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update notebooks/community/alphagenome/cloudai_alphagenome_vai_quickstart.ipynb

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

---------

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-12-05 18:30:23 +00:00
Damodar PanigrahiGitHubgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
4ab197a4ba AlphaGenome GCP API with quickstart.ipynb and README.md (#4378)
* AlphaGenome GCP API  with quickstart.ipynb and README.md

* Update notebooks/community/alphagenome/README.md

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update notebooks/community/alphagenome/README.md

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update notebooks/community/alphagenome/README.md

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update notebooks/community/alphagenome/README.md

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update notebooks/community/alphagenome/README.md

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update notebooks/community/alphagenome/README.md

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update cloudai_alphagenome_vai_quickstart.ipynb

lint errors

* Update cloudai_alphagenome_vai_quickstart.ipynb

lint errors

* lint errors

* lint errors

* lint import order

* lint errors

* lint import order

* lint import

* Update cloudai_alphagenome_vai_quickstart.ipynb format

* Update cloudai_alphagenome_vai_quickstart.ipynb remove hardcoded url

---------

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-11-27 18:56:07 +00:00
Damodar PanigrahiandGitHub 2a5877fbd1 Update CODEOWNERS (#4380)
* Update CODEOWNERS

* Update CODEOWNERS
2025-11-27 18:30:33 +00:00
27 changed files with 8494 additions and 22 deletions
+1 -1
View File
@@ -9,7 +9,7 @@ jobs:
- name: Set up Python
uses: actions/setup-python@v6
with:
python-version: '3.x'
python-version: '3.13'
- name: Fetch pull request branch
uses: actions/checkout@v4
with:
+1
View File
@@ -26,6 +26,7 @@
/vertex_endpoints/find_ideal_machine_type/find_ideal_machine_type/find_ideal_machine_type.ipynb @entrpn
/vertex_endpoints/nvidia-triton/nvidia-triton-custom-container-prediction.ipynb @RajeshThallam
/vertex_endpoints/optimized_tensorflow_runtime @vlasenkoalexey
/notebooks/community/alphagenome/cloudai_alphagenome_vai_quickstart.ipynb @dpanigra
/notebooks/community/ml_ops/stage2/get_started_with_visionapi_and_automl.ipynb @mansari
/notebooks/community/neo4j/graph_paysim.ipynb @benofben @laeg
/notebooks/community/ml_ops/stage1/get_started_with_visionapi_and_vertex_datasets.ipynb @mansari
+93
View File
@@ -0,0 +1,93 @@
![AlphaGenome header image](https://raw.githubusercontent.com/google-deepmind/alphagenome/refs/heads/main/docs/source/_static/header.png)
# AlphaGenome
[**Overview**](#overview) | [**Use Cases**](#use-cases) | [**Documentation**](#documentation) | [**Pricing**](#pricing) | [**Quick start**](#quick-start)
## Overview
**Disclaimer:** *Experimental*.
*The AlphaGenome Private Preview is a "Pre-GA Offering" subject to the "Pre-GA
Offerings Terms" in the General Service Terms section of the Google Cloud
[Service Specific Terms](https://cloud.google.com/terms/service-terms). It is
also a “Generative AI Preview Product” as defined in and subject to the
[Additional Terms for Generative AI Preview Products](https://cloud.google.com/trustedtester/aitos?e=48754805&hl=en).
Pre-GA products are available "as is" and might have limited support. For more
information, see the [launch stage](https://cloud.google.com/products?e=48754805#product-launch-stages)
descriptions.*
Access to the AlphaGenome model capabilities requires application and approval.
Users must be added to an allowlist to use the service.
If you are interested in applying to the program, **Request Access** above.
&nbsp;
AlphaGenome is Google DeepMind’s unifying model for deciphering the regulatory
code within DNA sequences.
AlphaGenome offers multimodal predictions, encompassing diverse functional
outputs such as gene expression, splicing patterns, chromatin features, and
contact maps (see diagram below). The model analyzes DNA sequences of up to 1
million base pairs in length and can deliver predictions at single base-pair
resolution for most outputs.
Training data was sourced from large public consortia including
[ENCODE](http://encodeproject.org/), [GTEx](https://www.gtexportal.org/),
[4D Nucleome](https://4dnucleome.org/) and
[FANTOM5](https://fantom.gsc.riken.jp/5/), which experimentally measured these
properties covering important modalities of gene regulation across hundreds of
human and mouse cell types and tissues.
![Diagram showing an overview of the AlphaGenome model architecture and its inputs/outputs](https://www.alphagenomedocs.com/_images/model_overview.png)
## Use Cases
* **Sequence-to-function predictions:** Predict multiple functional tracks (such as gene expression, splicing) from DNA sequences across a wide variety of tissues and cell types.
* **Variant effect scoring:** Assess the impact of genetic variants by comparing predictions for the reference and alternative alleles and summarising the differences between them.
* **Identify functional regions:** Use in silico mutagenesis (ISM) to identify functionally important regions in the DNA sequence.
* **Human and mouse capability:** Generate predictions for both human and mouse genomes.
## Documentation
This API provides access to AlphaGenome, Google DeepMind's unifying model for
deciphering the regulatory code within DNA sequences. AlphaGenome offers
multimodal predictions, encompassing diverse functional outputs including gene
expression, splicing patterns, chromatin features, and contact maps (see diagram
below). The model analyzes up to 1 million base pairs of DNA sequence and can
deliver predictions at single base-pair resolution for most modalities.
AlphaGenome achieves state-of-the-art performance across a range of genomic
prediction benchmarks, including diverse variant effect prediction tasks.
The Google Cloud API for AlphaGenome provides a way for Google Cloud customers
to explore the AlphaGenome API for commercial use cases. This API is in private
preview (Request Access above). Once allowlisted, customers can access the API
directly or use the [colab](cloudai_alphagenome_vai_quickstart.ipynb).
### Acknowledgements
*Avsec, Ž., Latysheva, N., Cheng, J., Novati, G., Taylor, K. R., Ward, T., ... Kohli, P. (2025). AlphaGenome: advancing regulatory variant effect prediction with a unified DNA sequence model. bioRxiv.* [https://doi.org/10.1101/2025.06.25.661532](https://doi.org/10.1101/2025.06.25.661532)
### Contact
If you have any questions on using these models on Google Cloud please contact:
[alphagenome-cloud-external@google.com](mailto:alphagenome-cloud-external@google.com) or join the community [Discourse](https://www.alphagenomecommunity.com/) for more generic questions on AlphaGenome.
### Links
* Read our [paper](https://doi.org/10.1101/2025.06.25.661532)
* Read our [blog post](https://deepmind.google/discover/blog/alphagenome-ai-for-better-understanding-the-genome)
* Join the [community](https://www.alphagenomecommunity.com/)
* Check out the [AlphaGenome 101 Video](https://youtu.be/Xbvloe13nak)
## Pricing
Access to AlphaGenome on Vertex AI is currently restricted.
To utilize these models via this service:
* You must **Request Access** using your Google contact.
* Your application will be reviewed, and if approved, you will be **added to
an allowlist**.
* Only allowlisted users can access the API
* **Pricing information** will be shared directly with users upon approval
and placement on the allowlist.
## Quick start
The quickest way to get started with the AlphaGenome in Google Cloud Platform is to run [our example notebook](cloudai_alphagenome_vai_quickstart.ipynb) in [Google Colab](https://colab.research.google.com/).
File diff suppressed because one or more lines are too long
@@ -546,6 +546,7 @@ def get_quota_id(
"NVIDIA_GB200": "B200GPUs",
"NVIDIA_TESLA_T4": "T4GPUs",
"NVIDIA_RTX_PRO_6000": "RTXPRO6000GPUs",
"TPU_7x": "7XTPU",
"TPU_V6e": "V6ETPU",
"TPU_V5e": "V5ETPU",
"TPU_V3": "V3TPUs",
@@ -546,6 +546,7 @@ def get_quota_id(
"NVIDIA_GB200": "B200GPUs",
"NVIDIA_TESLA_T4": "T4GPUs",
"NVIDIA_RTX_PRO_6000": "RTXPRO6000GPUs",
"TPU_7x": "7XTPU",
"TPU_V6e": "V6ETPU",
"TPU_V5e": "V5ETPU",
"TPU_V3": "V3TPUs",
@@ -1105,6 +1105,7 @@
" max_num_seqs: int = 256,\n",
" model_type: str = None,\n",
" enable_llama_tool_parser: bool = False,\n",
" is_spot: bool = False,\n",
") -> Tuple[aiplatform.Model, aiplatform.Endpoint]:\n",
" \"\"\"Deploys trained models with vLLM into Vertex AI.\"\"\"\n",
" endpoint = aiplatform.Endpoint.create(\n",
@@ -1199,6 +1200,7 @@
" accelerator_count=accelerator_count,\n",
" deploy_request_timeout=1800,\n",
" service_account=service_account,\n",
" spot=is_spot,\n",
" system_labels={\n",
" \"NOTEBOOK_NAME\": \"model_garden_axolotl_finetuning.ipynb\",\n",
" \"NOTEBOOK_ENVIRONMENT\": get_deploy_source(),\n",
@@ -109,7 +109,8 @@
"! pip install --upgrade --quiet gcsfs==2024.3.1\n",
"! pip install --upgrade --quiet accelerate==0.34.2\n",
"! pip install --upgrade --quiet transformers==4.47.1\n",
"! pip install --upgrade --quiet datasets==2.20.0"
"! pip install --upgrade --quiet datasets==2.20.0\n",
"! pip install --upgrade --quiet google-cloud-aiplatform==1.130.0"
]
},
{
@@ -839,6 +840,7 @@
" max_num_seqs: int = 256,\n",
" model_type: str = None,\n",
" enable_llama_tool_parser: bool = False,\n",
" is_spot: bool = False,\n",
") -> Tuple[aiplatform.Model, aiplatform.Endpoint]:\n",
" \"\"\"Deploys trained models with vLLM into Vertex AI.\"\"\"\n",
" endpoint = aiplatform.Endpoint.create(\n",
@@ -933,6 +935,7 @@
" accelerator_count=accelerator_count,\n",
" deploy_request_timeout=1800,\n",
" service_account=service_account,\n",
" spot=is_spot,\n",
" system_labels={\n",
" \"NOTEBOOK_NAME\": \"model_garden_gemma2_finetuning_on_vertex.ipynb\",\n",
" \"NOTEBOOK_ENVIRONMENT\": get_deploy_source(),\n",
@@ -97,7 +97,7 @@
"# @title Install Python Packages for Finetuning\n",
"\n",
"# @markdown 1. Install google-cloud-aiplatform package and restart the session if instructed.\n",
"! pip install --upgrade --quiet 'google-cloud-aiplatform>=1.66.0'\n",
"! pip install --upgrade --quiet google-cloud-aiplatform==1.130.0\n",
"\n",
"# @markdown 2. Install packages to validate dataset with template.\n",
"! pip install --upgrade --quiet accelerate==0.31.0\n",
@@ -700,6 +700,7 @@
" max_num_seqs: int = 256,\n",
" model_type: str = None,\n",
" enable_llama_tool_parser: bool = False,\n",
" is_spot: bool = False,\n",
") -> Tuple[aiplatform.Model, aiplatform.Endpoint]:\n",
" \"\"\"Deploys trained models with vLLM into Vertex AI.\"\"\"\n",
" endpoint = aiplatform.Endpoint.create(\n",
@@ -794,6 +795,7 @@
" accelerator_count=accelerator_count,\n",
" deploy_request_timeout=1800,\n",
" service_account=service_account,\n",
" spot=is_spot,\n",
" system_labels={\n",
" \"NOTEBOOK_NAME\": \"model_garden_gemma_finetuning_on_vertex.ipynb\",\n",
" \"NOTEBOOK_ENVIRONMENT\": get_deploy_source(),\n",
@@ -163,7 +163,7 @@
"TASK = \"text-classification\" # @param {type: \"string\", isTemplate: true}\n",
"\n",
"# The pre-built serving docker images for Hugging Face Pytorch Inference.\n",
"SERVE_DOCKER_URI = \"us-docker.pkg.dev/deeplearning-platform-release/vertex-model-garden/hf-inference-toolkit.cu125.0-1.ubuntu2204.py311:model-garden.hf-inference-toolkit-0-1-release_20251115.00_p0\"\n",
"SERVE_DOCKER_URI = \"us-docker.pkg.dev/deeplearning-platform-release/vertex-model-garden/hf-inference-toolkit.cu125.0-1.ubuntu2204.py311:model-garden.hf-inference-toolkit-0-1-release_20251206.00_p0\"\n",
"\n",
"machine_type = \"g2-standard-8\" # @param {type: \"string\", isTemplate: true}\n",
"accelerator_type = \"NVIDIA_L4\" # @param [\"NVIDIA_L4\", \"None\"] {isTemplate: true}\n",
@@ -214,7 +214,7 @@
"HUGGING_FACE_MODEL_ID = \"Qwen/Qwen3-Embedding-8B\" # @param {type: \"string\", isTemplate: true}\n",
"\n",
"# The pre-built serving docker images for TEI.\n",
"TEI_DOCKER_URI = \"us-docker.pkg.dev/deeplearning-platform-release/vertex-model-garden/hf-tei.cu125.0-1.ubuntu2204.py310:model-garden.hf-tei-0-1-release_20251030.00_p0\"\n",
"TEI_DOCKER_URI = \"us-docker.pkg.dev/deeplearning-platform-release/vertex-model-garden/hf-tei.cu125.0-1.ubuntu2204.py310:model-garden.hf-tei-0-1-release_20251205.00_p0\"\n",
"\n",
"machine_type = \"g2-standard-8\" # @param {type: \"string\", isTemplate: true}\n",
"accelerator_type = \"NVIDIA_L4\" # @param [\"NVIDIA_L4\", \"None\"] {isTemplate: true}\n",
@@ -164,7 +164,7 @@
"HF_TOKEN = \"\" # @param {type:\"string\", isTemplate: true}\n",
"\n",
"# The pre-built vLLM serving docker image.\n",
"VLLM_DOCKER_URI = \"us-docker.pkg.dev/vertex-ai/vertex-vision-model-garden-dockers/pytorch-vllm-serve:20251114_0916_RC01\"\n",
"VLLM_DOCKER_URI = \"us-docker.pkg.dev/vertex-ai/vertex-vision-model-garden-dockers/pytorch-vllm-serve:20251205_0916_RC01\"\n",
"SERVING_CONTAINER_IMAGE_URI = VLLM_DOCKER_URI\n",
"LABEL = \"vllm\"\n",
"\n",
@@ -148,7 +148,8 @@
"! pip install --quiet accelerate==0.31.0\n",
"! pip install --quiet transformers==4.43.1\n",
"! pip install --quiet datasets==2.19.2\n",
"! pip install --quiet tensorflow==2.18.0"
"! pip install --quiet tensorflow==2.18.0\n",
"! pip install --upgrade --quiet google-cloud-aiplatform==1.130.0"
]
},
{
@@ -1366,6 +1367,7 @@
" max_num_seqs: int = 256,\n",
" model_type: str = None,\n",
" enable_llama_tool_parser: bool = False,\n",
" is_spot: bool = False,\n",
") -> Tuple[aiplatform.Model, aiplatform.Endpoint]:\n",
" \"\"\"Deploys trained models with vLLM into Vertex AI.\"\"\"\n",
" endpoint = aiplatform.Endpoint.create(\n",
@@ -1460,6 +1462,7 @@
" accelerator_count=accelerator_count,\n",
" deploy_request_timeout=1800,\n",
" service_account=service_account,\n",
" spot=is_spot,\n",
" system_labels={\n",
" \"NOTEBOOK_NAME\": \"model_garden_llama3_1_finetuning_with_workbench.ipynb\",\n",
" \"NOTEBOOK_ENVIRONMENT\": get_deploy_source(),\n",
@@ -0,0 +1,822 @@
{
"cells": [
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "ur8xi4C7S06n"
},
"outputs": [],
"source": [
"# Copyright 2025 Google LLC\n",
"#\n",
"# Licensed under the Apache License, Version 2.0 (the \"License\");\n",
"# you may not use this file except in compliance with the License.\n",
"# You may obtain a copy of the License at\n",
"#\n",
"# https://www.apache.org/licenses/LICENSE-2.0\n",
"#\n",
"# Unless required by applicable law or agreed to in writing, software\n",
"# distributed under the License is distributed on an \"AS IS\" BASIS,\n",
"# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.\n",
"# See the License for the specific language governing permissions and\n",
"# limitations under the License."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "JAPoU8Sm5E6e"
},
"source": [
"# Vertex AI Model Garden - Get started with DeepSeek-V3.2 models\n",
"\n",
"<table align=\"left\">\n",
" <td style=\"text-align: center\">\n",
" <a href=\"https://colab.research.google.com/github/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/model_garden/model_garden_openai_api_deepseek3_2.ipynb\">\n",
" <img src=\"https://cloud.google.com/ml-engine/images/colab-logo-32px.png\" alt=\"Google Colaboratory logo\"><br> Open in Colab\n",
" </a>\n",
" </td>\n",
" <td style=\"text-align: center\">\n",
" <a href=\"https://console.cloud.google.com/vertex-ai/colab/import/https:%2F%2Fraw.githubusercontent.com%2FGoogleCloudPlatform%2Fvertex-ai-samples%2Fmain%2Fnotebooks%2Fcommunity%2Fmodel_garden%2Fmodel_garden_openai_api_deepseek3_2.ipynb\"\">\n",
" <img width=\"32px\" src=\"https://cloud.google.com/ml-engine/images/colab-enterprise-logo-32px.png\" alt=\"Google Cloud Colab Enterprise logo\"><br> Open in Colab Enterprise\n",
" </a>\n",
" </td> \n",
" <td style=\"text-align: center\">\n",
" <a href=\"https://console.cloud.google.com/vertex-ai/workbench/deploy-notebook?download_url=https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/model_garden/model_garden_openai_api_deepseek3_2.ipynb\">\n",
" <img src=\"https://lh3.googleusercontent.com/UiNooY4LUgW_oTvpsNhPpQzsstV5W8F7rYgxgGBD85cWJoLmrOzhVs_ksK_vgx40SHs7jCqkTkCk=e14-rj-sc0xffffff-h130-w32\" alt=\"Vertex AI logo\"><br> Open in Workbench\n",
" </a>\n",
" </td>\n",
" <td style=\"text-align: center\">\n",
" <a href=\"https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/model_garden/model_garden_openai_api_deepseek3_2.ipynb\">\n",
" <img src=\"https://cloud.google.com/ml-engine/images/github-logo-32px.png\" alt=\"GitHub logo\"><br> View on GitHub\n",
" </a>\n",
" </td>\n",
"</table>"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "tvgnzT1CKxrO"
},
"source": [
"## Overview\n",
"\n",
"This notebook demonstrates how to get started with using the OpenAI library and demonstrates how to use DeepSeek-V3.2 models as Model-as-service (MaaS) for building translation chain and document question-answer.\n",
"\n",
"### Objective\n",
"\n",
"- Configure OpenAI SDK for the DeepSeek-V3.2 Completions API\n",
"- Chat with DeepSeek-V3.2 models with different prompts and model parameters, and apply Llama Guard for safeguarding\n",
"- Build with DeepSeek-V3.2 models\n",
" - Translation Chain.\n",
"\n",
"### Costs\n",
"\n",
"This tutorial uses billable components of Google Cloud:\n",
"\n",
"* Vertex AI\n",
"* Cloud Storage\n",
"\n",
"Learn about [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing), [Cloud Storage pricing](https://cloud.google.com/storage/pricing), and use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to generate a cost estimate based on your projected usage."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "61RBz8LLbxCR"
},
"source": [
"## Get started"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "No17Cw5hgx12"
},
"source": [
"### Install Vertex AI SDK for Python and other required packages\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "tFy3H3aPgx12"
},
"outputs": [],
"source": [
"! pip3 install --upgrade --quiet google-cloud-aiplatform[langchain] openai\n",
"! pip3 install --upgrade --quiet langchain-openai"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "R5Xep4W9lq-Z"
},
"source": [
"### Restart runtime (Colab only)\n",
"\n",
"To use the newly installed packages, you must restart the runtime on Google Colab."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "XRvKdaPDTznN"
},
"outputs": [],
"source": [
"import sys\n",
"\n",
"if \"google.colab\" in sys.modules:\n",
"\n",
" import IPython\n",
"\n",
" app = IPython.Application.instance()\n",
" app.kernel.do_shutdown(True)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "SbmM4z7FOBpM"
},
"source": [
"<div class=\"alert alert-block alert-warning\">\n",
"<b>⚠️ The kernel is going to restart. Wait until it's finished before continuing to the next step. ⚠️</b>\n",
"</div>\n"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "dmWOrTJ3gx13"
},
"source": [
"### Authenticate your notebook environment (Colab only)\n",
"\n",
"Authenticate your environment on Google Colab.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "NyKGtVQjgx13"
},
"outputs": [],
"source": [
"import sys\n",
"\n",
"if \"google.colab\" in sys.modules:\n",
"\n",
" from google.colab import auth\n",
"\n",
" auth.authenticate_user()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "DF4l8DTdWgPY"
},
"source": [
"### Set Google Cloud project information\n",
"\n",
"To get started using Vertex AI, you must have an existing Google Cloud project and [enable the Vertex AI API](https://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com). Learn more about [setting up a project and a development environment](https://cloud.google.com/vertex-ai/docs/start/cloud-environment)."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "Nqwi-5ufWp_B"
},
"outputs": [],
"source": [
"PROJECT_ID = \"<YOUR PROJECT ID>\" # @param {type:\"string\"}\n",
"\n",
"LOCATION = \"global\" # @param {type:\"string\"}"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "jVYoyDl165EE"
},
"source": [
"### Import libraries\n",
"\n",
"Import libraries to use in this tutorial."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "c1tEW-U968h8"
},
"outputs": [],
"source": [
"# Chat completions API\n",
"import openai\n",
"from google.auth import default, transport\n",
"from langchain import PromptTemplate\n",
"# Build\n",
"from langchain_openai import ChatOpenAI"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "uqYCG2Fw7D3L"
},
"source": [
"### Configure OpenAI SDK for the DeepSeek-V3.2 Chat Completions API\n",
"\n",
"To configure the OpenAI SDK for the DeepSeek-V3.2 Chat Completions API, you need to request the access token and initialize the client pointing to the DeepSeek-V3.2 endpoint.\n"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "W0K6VSJRHhH2"
},
"source": [
"#### Authentication\n",
"\n",
"You can request an access token from the default credentials for the current environment. Note that the access token lives for [1 hour by default](https://cloud.google.com/docs/authentication/token-types#at-lifetime); after expiration, it must be refreshed.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "i0qceuiQEPHv"
},
"outputs": [],
"source": [
"credentials, _ = default()\n",
"auth_request = transport.requests.Request()\n",
"credentials.refresh(auth_request)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "Q04wJmA0HT6X"
},
"source": [
"Then configure the OpenAI SDK to point to the DeepSeek-V3.2 Chat Completions API endpoint.\n",
"\n",
"Notice, only `global` is supported region for DeepSeek-V3.2 models using Model-as-a-Service (MaaS)."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "c-MRhsnlj6iw"
},
"outputs": [],
"source": [
"MODEL_LOCATION = \"global\"\n",
"\n",
"client = openai.OpenAI(\n",
" base_url=f\"https://aiplatform.googleapis.com/v1/projects/{PROJECT_ID}/locations/{MODEL_LOCATION}/endpoints/openapi/chat/completions?\",\n",
" api_key=credentials.token,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "UGokrtdiIHrX"
},
"source": [
"#### DeepSeek-V3.2 Model\n",
"\n",
"This tutorial uses DeepSeek-V3.2 using Model-as-a-Service (MaaS). Using Model-as-a-Service (MaaS), you can access DeepSeek-V3.2 model in just a few clicks without any setup or infrastructure hassles. Model-as-a-Service (MaaS) integrates [Llama Guard](https://huggingface.co/meta-llama/Llama-Guard-3-8B) as a safety filter. It is switched on by default and can be switched off. Llama Guard enables us to safeguard model inputs and outputs. If a response is filtered, it will be populated with a `finish_reason` field (with value `content_filtered`) and a `refusal` field (stating the filtering reason)."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "r7OhyH46H2H5"
},
"outputs": [],
"source": [
"MODEL_ID = \"deepseek-ai/deepseek-v3.2-maas\" # @param {type:\"string\"} [\"deepseek-ai/deepseek-v3.2-maas\"]"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "1xD62NTpqHXd"
},
"source": [
"### Chat with DeepSeek-V3.2\n",
"\n",
"Use the Chat Completions API to send a request to the DeepSeek-V3.2model."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "tkyp9kZSuJGx"
},
"source": [
"#### Hello, DeepSeek-V3.2!"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "CKVOZ1HEqRbY"
},
"outputs": [],
"source": [
"apply_llama_guard = True # @param {type:\"boolean\"}\n",
"\n",
"response = client.chat.completions.create(\n",
" model=MODEL_ID,\n",
" messages=[{\"role\": \"user\", \"content\": \"Hello, Deepseek!\"}],\n",
" extra_body={\n",
" \"extra_body\": {\n",
" \"google\": {\n",
" \"model_safety_settings\": {\n",
" \"enabled\": apply_llama_guard,\n",
" \"llama_guard_settings\": {},\n",
" }\n",
" }\n",
" }\n",
" },\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "LxpdxYCxH51u"
},
"outputs": [],
"source": [
"print(response.choices[0].message.content)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "B1rKbHUQt605"
},
"source": [
"#### Ask DeepSeek-V3.2 using different model configuration\n",
"\n",
"Use the following parameters to generate different answers:\n",
"\n",
"* `temperature` to control the randomness of the response\n",
"* `max_tokens` to limit the response length\n",
"* `top_p` to control the quality of the response\n",
"* `stream` to stream the response back or not\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "owv-5Sz5rIEU"
},
"outputs": [],
"source": [
"temperature = 1.0 # @param {type:\"number\"}\n",
"max_tokens = 256 # @param {type:\"integer\"}\n",
"top_p = 1.0 # @param {type:\"number\"}\n",
"stream = True # @param {type:\"boolean\"}"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "a-qBuhcK-G1V"
},
"source": [
"Get the answer."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "O1YU8bSivH0B"
},
"outputs": [],
"source": [
"apply_llama_guard = True # @param {type:\"boolean\"}\n",
"\n",
"response = client.chat.completions.create(\n",
" model=MODEL_ID,\n",
" messages=[\n",
" {\"role\": \"user\", \"content\": \"What is Vertex AI?\"},\n",
" {\"role\": \"assistant\", \"content\": \"Sure, Vertex AI is:\"},\n",
" ],\n",
" temperature=temperature,\n",
" max_tokens=max_tokens,\n",
" top_p=top_p,\n",
" stream=stream,\n",
" extra_body={\n",
" \"extra_body\": {\n",
" \"google\": {\n",
" \"model_safety_settings\": {\n",
" \"enabled\": apply_llama_guard,\n",
" \"llama_guard_settings\": {},\n",
" }\n",
" }\n",
" }\n",
" },\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "-o9-gF0U-Kba"
},
"source": [
"Depending if `stream` parameter is enabled or not, you can print the response entirely or chunk by chunk."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "CoDHLGhyyt8d"
},
"outputs": [],
"source": [
"if stream:\n",
" for chunk in response:\n",
" if chunk.choices:\n",
" print(chunk.choices[0].delta.content, end=\"\")\n",
"else:\n",
" print(response.choices[0].message.content)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "BkoaelaKxm1r"
},
"source": [
"#### Use DeepSeek-V3.2 with different tasks\n",
"\n",
"In this section, you will use DeepSeek-V3.2 to perform different tasks including text generation, text summarization, and code generation.\n",
"\n",
"For each task, you'll define a different prompt and submit a request to the model as you did before."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "-en7AYQDyONt"
},
"source": [
"##### Text Generation"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "QANInNvizWbi"
},
"outputs": [],
"source": [
"prompt = \"Write a poem about a cat who loves to code\""
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "2x8ML1Y_yfom"
},
"outputs": [],
"source": [
"apply_llama_guard = True # @param {type:\"boolean\"}\n",
"\n",
"response = client.chat.completions.create(\n",
" model=MODEL_ID,\n",
" messages=[\n",
" {\"role\": \"user\", \"content\": prompt},\n",
" ],\n",
" extra_body={\n",
" \"extra_body\": {\n",
" \"google\": {\n",
" \"model_safety_settings\": {\n",
" \"enabled\": apply_llama_guard,\n",
" \"llama_guard_settings\": {},\n",
" }\n",
" }\n",
" }\n",
" },\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "qQ6UUgpHztXZ"
},
"outputs": [],
"source": [
"print(response.choices[0].message.content)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "kBLESIw4zhto"
},
"source": [
"##### Text summarization"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "UrAybklfzhtz"
},
"outputs": [],
"source": [
"article = \"\"\"\n",
"Vertex AI: Google's Unified Platform for Machine Learning\n",
"\n",
"Google Cloud's Vertex AI is a comprehensive platform that simplifies the process of building, deploying, and managing machine learning (ML) models and AI applications. It provides a single environment for all your AI needs, from data preparation to model deployment and monitoring.\n",
"\n",
"Vertex AI offers a range of features to cater to various user levels, including:\n",
"\n",
"AutoML: This feature allows you to train models on tabular, image, text, or video data without writing code. It's ideal for users without extensive ML expertise.\n",
"Custom Training: For advanced users, Vertex AI provides custom training options, allowing you to use your preferred ML framework and write your own code.\n",
"Model Garden: This feature lets you discover, test, and deploy pre-trained models from Vertex AI and open-source sources.\n",
"Generative AI: Access Google's powerful large language models (LLMs) to generate text, code, images, and speech, which can be customized and deployed for your applications.\n",
"Vertex AI seamlessly integrates with other Google Cloud services like BigQuery for data warehousing, Cloud Storage for data management, and Cloud AI Platform for custom model training. It provides managed infrastructure that can be tailored to your performance and budget needs.\n",
"\n",
"Whether you're a seasoned data scientist or just starting out with AI, Vertex AI simplifies the entire ML lifecycle and empowers you to build and deploy AI solutions effectively.\n",
"\"\"\"\n",
"\n",
"\n",
"prompt = (\"Summarize the following article in one sentence: \" + article).replace(\n",
" \"\\n\", \"\"\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "VJYIZbGyzhtz"
},
"outputs": [],
"source": [
"apply_llama_guard = True # @param {type:\"boolean\"}\n",
"\n",
"response = client.chat.completions.create(\n",
" model=MODEL_ID,\n",
" messages=[\n",
" {\"role\": \"user\", \"content\": prompt},\n",
" ],\n",
" extra_body={\n",
" \"extra_body\": {\n",
" \"google\": {\n",
" \"model_safety_settings\": {\n",
" \"enabled\": apply_llama_guard,\n",
" \"llama_guard_settings\": {},\n",
" }\n",
" }\n",
" }\n",
" },\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "9vfWA2i9zwOZ"
},
"outputs": [],
"source": [
"print(response.choices[0].message.content)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "-_mNB0Fezh6G"
},
"source": [
"##### Code generation"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "ltbhrGiwzh6H"
},
"outputs": [],
"source": [
"prompt = \"Write a Python function that takes a list of numbers and returns the average. Include error handling for empty lists.\""
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "q66yE4Pszh6H"
},
"outputs": [],
"source": [
"apply_llama_guard = True # @param {type:\"boolean\"}\n",
"\n",
"response = client.chat.completions.create(\n",
" model=MODEL_ID,\n",
" messages=[\n",
" {\"role\": \"user\", \"content\": prompt},\n",
" ],\n",
" extra_body={\n",
" \"extra_body\": {\n",
" \"google\": {\n",
" \"model_safety_settings\": {\n",
" \"enabled\": apply_llama_guard,\n",
" \"llama_guard_settings\": {},\n",
" }\n",
" }\n",
" }\n",
" },\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "mcAf5tXrtPIu"
},
"outputs": [],
"source": [
"print(response.choices[0].message.content)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "gnrXpv5Y3yFK"
},
"source": [
"### Build with DeepSeek-V3.2\n",
"\n",
"In this section, you use DeepSeek-V3.2 to build a translation simple applications.\n",
"\n",
"**Translation Chain** to translate text across multiple languages using DeepSeek-V3.2 and LangChain Expression Language (LCEL).\n"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "IowGcZq95HqZ"
},
"source": [
"#### Translation chain\n",
"\n",
"In this scenario, you use LangChain Expression Language (LCEL) to build a simple chain which translates some `text_to_translate` to the specified `target_language`."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "yLAwaYPzFqDQ"
},
"source": [
"##### Initialize the chat interface and the translation prompt template using LangChain"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "CE2KxIrG5xKC"
},
"outputs": [],
"source": [
"llm = ChatOpenAI(\n",
" model=MODEL_ID,\n",
" base_url=f\"https://aiplatform.googleapis.com/v1/projects/{PROJECT_ID}/locations/{MODEL_LOCATION}/endpoints/openapi/chat/completions?\",\n",
" api_key=credentials.token,\n",
")\n",
"\n",
"template = \"\"\"Translate the following {text} to {target_language}:\"\"\"\n",
"\n",
"prompt = PromptTemplate(input_variables=[\"text\", \"target_language\"], template=template)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "odTLqzLiF8h_"
},
"source": [
"##### Initialize the chain"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "74Mywe4W9MmE"
},
"outputs": [],
"source": [
"chain = prompt | llm"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "QE3sizbzGFER"
},
"source": [
"##### Translate a text"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "SRJ-xuSI9ZQl"
},
"outputs": [],
"source": [
"text_to_translate = \"Hello Deepseek!\" # @param {type:\"string\"}\n",
"target_language = \"Italian\" # @param {type:\"string\"}\n",
"\n",
"response = chain.invoke({\"text\": text_to_translate, \"target_language\": target_language})"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "yYzc1kCjEHGP"
},
"outputs": [],
"source": [
"print(response.content)"
]
}
],
"metadata": {
"colab": {
"name": "model_garden_openai_api_deepseek3_2.ipynb",
"toc_visible": true
},
"kernelspec": {
"display_name": "Python 3",
"name": "python3"
}
},
"nbformat": 4,
"nbformat_minor": 0
}
@@ -97,7 +97,7 @@
"# @title Install Python Packages for Finetuning\n",
"\n",
"# @markdown 1. Install google-cloud-aiplatform package and restart the session if instructed.\n",
"! pip install --upgrade --quiet 'google-cloud-aiplatform>=1.66.0'\n",
"! pip install --upgrade --quiet google-cloud-aiplatform==1.130.0\n",
"\n",
"# @markdown 2. Install packages to validate dataset with template.\n",
"! pip install --upgrade --quiet accelerate==0.31.0\n",
@@ -708,6 +708,7 @@
" max_num_seqs: int = 256,\n",
" model_type: str = None,\n",
" enable_llama_tool_parser: bool = False,\n",
" is_spot: bool = False,\n",
") -> Tuple[aiplatform.Model, aiplatform.Endpoint]:\n",
" \"\"\"Deploys trained models with vLLM into Vertex AI.\"\"\"\n",
" endpoint = aiplatform.Endpoint.create(\n",
@@ -802,6 +803,7 @@
" accelerator_count=accelerator_count,\n",
" deploy_request_timeout=1800,\n",
" service_account=service_account,\n",
" spot=is_spot,\n",
" system_labels={\n",
" \"NOTEBOOK_NAME\": \"model_garden_pytorch_gemma_peft_finetuning_hf.ipynb\",\n",
" \"NOTEBOOK_ENVIRONMENT\": get_deploy_source(),\n",
@@ -104,7 +104,8 @@
"! pip install --upgrade --quiet gcsfs==2024.3.1\n",
"! pip install --upgrade --quiet accelerate==0.34.2\n",
"! pip install --upgrade --quiet transformers==4.47.1\n",
"! pip install --upgrade --quiet datasets==2.20.0"
"! pip install --upgrade --quiet datasets==2.20.0\n",
"! pip install --upgrade --quiet google-cloud-aiplatform==1.130.0"
]
},
{
@@ -816,6 +817,9 @@
"# The pre-built serving docker image for vLLM.\n",
"VLLM_DOCKER_URI = \"us-docker.pkg.dev/vertex-ai/vertex-vision-model-garden-dockers/pytorch-vllm-serve:20250116_0916_RC00\"\n",
"\n",
"# @markdown Choose whether to use a [Spot VM](https://cloud.google.com/compute/docs/instances/spot) for the deployment.\n",
"is_spot = False # @param {type:\"boolean\"}\n",
"\n",
"# @markdown Set the Deployment Region. If not set, it will be set to default region.\n",
"DEPLOY_REGION = \"\" # @param {type: \"string\"}\n",
"if not DEPLOY_REGION:\n",
@@ -895,6 +899,7 @@
" max_num_seqs: int = 256,\n",
" model_type: str = None,\n",
" enable_llama_tool_parser: bool = False,\n",
" is_spot: bool = False,\n",
") -> Tuple[aiplatform.Model, aiplatform.Endpoint]:\n",
" \"\"\"Deploys trained models with vLLM into Vertex AI.\"\"\"\n",
" endpoint = aiplatform.Endpoint.create(\n",
@@ -989,6 +994,7 @@
" accelerator_count=accelerator_count,\n",
" deploy_request_timeout=1800,\n",
" service_account=service_account,\n",
" spot=is_spot,\n",
" system_labels={\n",
" \"NOTEBOOK_NAME\": \"model_garden_pytorch_llama3_1_finetuning.ipynb\",\n",
" \"NOTEBOOK_ENVIRONMENT\": get_deploy_source(),\n",
@@ -1047,6 +1053,7 @@
" max_model_len=max_model_len,\n",
" enable_lora=True,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
" is_spot=is_spot,\n",
")\n",
"\n",
"# @markdown Click \"Show Code\" to see more details."
@@ -104,7 +104,8 @@
"! pip install --upgrade --quiet gcsfs==2024.3.1\n",
"! pip install --upgrade --quiet accelerate==0.34.2\n",
"! pip install --upgrade --quiet transformers==4.47.1\n",
"! pip install --upgrade --quiet datasets==2.20.0"
"! pip install --upgrade --quiet datasets==2.20.0\n",
"! pip install --upgrade --quiet google-cloud-aiplatform==1.130.0"
]
},
{
@@ -849,6 +850,7 @@
" max_num_seqs: int = 256,\n",
" model_type: str = None,\n",
" enable_llama_tool_parser: bool = False,\n",
" is_spot: bool = False,\n",
") -> Tuple[aiplatform.Model, aiplatform.Endpoint]:\n",
" \"\"\"Deploys trained models with vLLM into Vertex AI.\"\"\"\n",
" endpoint = aiplatform.Endpoint.create(\n",
@@ -943,6 +945,7 @@
" accelerator_count=accelerator_count,\n",
" deploy_request_timeout=1800,\n",
" service_account=service_account,\n",
" spot=is_spot,\n",
" system_labels={\n",
" \"NOTEBOOK_NAME\": \"model_garden_pytorch_llama3_3_finetuning.ipynb\",\n",
" \"NOTEBOOK_ENVIRONMENT\": get_deploy_source(),\n",
@@ -0,0 +1,553 @@
{
"cells": [
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "SgQ6t5bqZVlH"
},
"outputs": [],
"source": [
"# Copyright 2025 Google LLC\n",
"#\n",
"# Licensed under the Apache License, Version 2.0 (the \"License\");\n",
"# you may not use this file except in compliance with the License.\n",
"# You may obtain a copy of the License at\n",
"#\n",
"# https://www.apache.org/licenses/LICENSE-2.0\n",
"#\n",
"# Unless required by applicable law or agreed to in writing, software\n",
"# distributed under the License is distributed on an \"AS IS\" BASIS,\n",
"# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.\n",
"# See the License for the specific language governing permissions and\n",
"# limitations under the License."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "99c1c3fc2ca5"
},
"source": [
"# Vertex AI Model Garden - Llama 3.3 (Deployment on TPU7x)\n",
"\n",
"<table><tbody><tr>\n",
" <td style=\"text-align: center\">\n",
" <a href=\"https://console.cloud.google.com/vertex-ai/notebooks/deploy-notebook?download_url=https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/main/notebooks/community/model_garden/model_garden_pytorch_llama3_3_tpu7x_deployment.ipynb\">\n",
" <img alt=\"Workbench logo\" src=\"https://lh3.googleusercontent.com/UiNooY4LUgW_oTvpsNhPpQzsstV5W8F7rYgxgGBD85cWJoLmrOzhVs_ksK_vgx40SHs7jCqkTkCk=e14-rj-sc0xffffff-h130-w32\" width=\"32px\"><br> Run in Workbench\n",
" </a>\n",
" </td>\n",
" <td style=\"text-align: center\">\n",
" <a href=\"https://console.cloud.google.com/vertex-ai/colab/import/https:%2F%2Fraw.githubusercontent.com%2FGoogleCloudPlatform%2Fvertex-ai-samples%2Fmain%2Fnotebooks%2Fcommunity%2Fmodel_garden%2Fmodel_garden_pytorch_llama3_3_tpu7x_deployment.ipynb\">\n",
" <img alt=\"Google Cloud Colab Enterprise logo\" src=\"https://lh3.googleusercontent.com/JmcxdQi-qOpctIvWKgPtrzZdJJK-J3sWE1RsfjZNwshCFgE_9fULcNpuXYTilIR2hjwN\" width=\"32px\"><br> Run in Colab Enterprise\n",
" </a>\n",
" </td>\n",
" <td style=\"text-align: center\">\n",
" <a href=\"https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/model_garden/model_garden_pytorch_llama3_3_tpu7x_deployment.ipynb\">\n",
" <img alt=\"GitHub logo\" src=\"https://github.githubassets.com/assets/GitHub-Mark-ea2971cee799.png\" width=\"32px\"><br> View on GitHub\n",
" </a>\n",
" </td>\n",
"</tr></tbody></table>"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "3de7470326a2"
},
"source": [
"## Overview\n",
"\n",
"This notebook demonstrates serving [meta-llama/Llama-3.3-70B-Instruct](https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct) models with [vLLM TPU](https://github.com/vllm-project/vllm) on [TPU7x](https://docs.cloud.google.com/tpu/docs/tpu7x) machines.\n",
"\n",
"### Objective\n",
"\n",
"- Deploy meta-llama/Llama-3.3-70B-Instruct on TPU7x machines with vLLM TPU.\n",
"\n",
"### File a bug\n",
"\n",
"File a bug on [GitHub](https://github.com/GoogleCloudPlatform/vertex-ai-samples/issues/new) if you encounter any issue with the notebook.\n",
"\n",
"### Costs\n",
"\n",
"This tutorial uses billable components of Google Cloud:\n",
"\n",
"* Vertex AI\n",
"* Cloud Storage\n",
"\n",
"Learn about [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing), [Cloud Storage pricing](https://cloud.google.com/storage/pricing), and use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to generate a cost estimate based on your projected usage."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "264c07757582"
},
"source": [
"## Before you begin"
]
},
{
"cell_type": "code",
"execution_count": 1,
"metadata": {
"cellView": "form",
"id": "ax7zWynUDcjk"
},
"outputs": [],
"source": [
"# @title Request for quota\n",
"\n",
"# @markdown To deploy with TPU7x machines, check that you have sufficient quota: [CustomModelServing7XTPUPerProjectPerRegion](https://console.cloud.google.com/iam-admin/quotas?metric=aiplatform.googleapis.com%2Fcustom_model_serving_tpu7x). Find the available region(s) [here](https://cloud.google.com/vertex-ai/docs/general/locations#region_considerations).\n",
"\n",
"# @markdown If you don't have sufficient quota, request for quota following the instructions at [\"Request a higher quota\"](https://cloud.google.com/docs/quota/view-manage#requesting_higher_quota).\n",
"\n",
"# @markdown You can also use Compute Engine reservations with Vertex Prediction following the instructions [here](https://cloud.google.com/vertex-ai/docs/predictions/use-reservations). Note that the GCE quota for the shared reservation will be managed separately. Shared reservation is the only GCE consumption mode."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "YXFGIp1l-qtT"
},
"outputs": [],
"source": [
"# @title Setup Google Cloud project\n",
"\n",
"# @markdown 1. [Make sure that billing is enabled for your project](https://cloud.google.com/billing/docs/how-to/modify-project).\n",
"\n",
"# @markdown 2. **[Optional]** Set region. If not set, the region will be set automatically according to Colab Enterprise environment.\n",
"\n",
"REGION = \"\" # @param {type:\"string\"}\n",
"\n",
"# Upgrade Vertex AI SDK.\n",
"! pip3 install --upgrade --quiet 'google-cloud-aiplatform==1.103.0'\n",
"\n",
"# Import the necessary packages\n",
"import importlib # noqa: F401\n",
"import os # noqa: F401\n",
"from typing import Tuple # noqa: F401\n",
"\n",
"from google.cloud import aiplatform # noqa: F401\n",
"\n",
"# Upgrade Vertex AI SDK.\n",
"if os.environ.get(\"VERTEX_PRODUCT\") != \"COLAB_ENTERPRISE\":\n",
" ! pip install --upgrade tensorflow\n",
"! git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git\n",
"\n",
"common_util = importlib.import_module(\n",
" \"vertex-ai-samples.notebooks.community.model_garden.docker_source_codes.notebook_util.common_util\"\n",
")\n",
"\n",
"LABEL = \"vllm_tpu\"\n",
"models, endpoints = {}, {}\n",
"\n",
"# Get the default cloud project id.\n",
"PROJECT_ID = os.environ[\"GOOGLE_CLOUD_PROJECT\"]\n",
"\n",
"# Get the default region for launching jobs.\n",
"if not REGION:\n",
" REGION = os.environ[\"GOOGLE_CLOUD_REGION\"]\n",
"\n",
"# Initialize Vertex AI API.\n",
"print(\"Initializing Vertex AI API.\")\n",
"aiplatform.init(project=PROJECT_ID, location=REGION)\n",
"\n",
"! gcloud config set project $PROJECT_ID\n",
"\n",
"import vertexai\n",
"\n",
"vertexai.init(\n",
" project=PROJECT_ID,\n",
" location=REGION,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "z-XybZjtgF9M"
},
"source": [
"## Deploy Llama 3.3 with vLLM TPU"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "E8OiHHNNE_wj"
},
"outputs": [],
"source": [
"# @title Set the model variants\n",
"\n",
"# @markdown Set the model to deploy.\n",
"\n",
"base_model_name = \"Llama-3.3-70B-Instruct\" # @param [\"Llama-3.3-70B-Instruct\"] {isTemplate:true}\n",
"hf_model_id = \"meta-llama/\" + base_model_name\n",
"model_user_id = \"llama3-3\"\n",
"model_id = f\"gs://vertex-model-garden-restricted-us/llama3.3/{base_model_name}\"\n",
"\n",
"PUBLISHER_MODEL_NAME = (\n",
" f\"publishers/meta/models/{model_user_id}@{base_model_name.lower()}\"\n",
")\n",
"\n",
"# @markdown Set use_dedicated_endpoint to False if you don't want to use [dedicated endpoint](https://cloud.google.com/vertex-ai/docs/general/deployment#create-dedicated-endpoint). Note that [dedicated endpoint does not support VPC Service Controls](https://cloud.google.com/vertex-ai/docs/predictions/choose-endpoint-type), uncheck the box if you are using VPC-SC.\n",
"use_dedicated_endpoint = True # @param {type:\"boolean\"}"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "acd75fc92341"
},
"outputs": [],
"source": [
"# @title Deploy with customized configs\n",
"\n",
"# @markdown This section uploads the Llama-3.3-70B-Instruct model to Model Registry and deploys them to a Vertex Prediction Endpoint. It takes ~1 hour to finish.\n",
"\n",
"# @markdown The pre-built serving docker image.\n",
"vLLM_TPU_DOCKER_URI = (\n",
" \"us-docker.pkg.dev/vertex-ai/vertex-vision-model-garden-dockers/vllm-serve:tpu7x\"\n",
")\n",
"\n",
"# @markdown Find Vertex AI prediction supported accelerators and regions at https://cloud.google.com/vertex-ai/docs/predictions/configure-compute.\n",
"machine_type = \"tpu7x-standard-4t\" # @param [\"tpu7x-standard-1t\"] {isTemplate:true}\n",
"if machine_type == \"tpu7x-standard-4t\":\n",
" accelerator_type = \"TPU_7x\"\n",
" accelerator_count = 4\n",
" tensor_parallel_size = 8\n",
"else:\n",
" raise ValueError(\"Sample deployment options are not available.\")\n",
"\n",
"common_util.check_quota(\n",
" project_id=PROJECT_ID,\n",
" region=REGION,\n",
" accelerator_type=accelerator_type,\n",
" accelerator_count=accelerator_count,\n",
" is_for_training=False,\n",
")\n",
"\n",
"TPU_DEPLOYMENT_REGION = REGION\n",
"max_model_len = 65536\n",
"max_num_seqs = 128\n",
"max_num_batched_tokens = 1024\n",
"\n",
"\n",
"def deploy_model_vllm_tpu(\n",
" model_name: str,\n",
" model_id: str,\n",
" publisher: str,\n",
" publisher_model_id: str,\n",
" base_model_id: str = None,\n",
" tensor_parallel_size: int = 1,\n",
" machine_type: str = \"ct6e-standard-1t\",\n",
" tpu_topology: str = \"1x1\",\n",
" max_model_len: int = 4096,\n",
" max_num_seqs: int = None,\n",
" max_num_batched_tokens: int = None,\n",
" enable_chunked_prefill: bool = False,\n",
" enable_prefix_cache: bool = False,\n",
" endpoint_id: str = \"\",\n",
" min_replica_count: int = 1,\n",
" max_replica_count: int = 1,\n",
" use_dedicated_endpoint: bool = False,\n",
" model_type: str = None,\n",
") -> Tuple[aiplatform.Model, aiplatform.Endpoint]:\n",
" \"\"\"Deploys models with vLLM on TPU in Vertex AI.\"\"\"\n",
" if endpoint_id:\n",
" aip_endpoint_name = (\n",
" f\"projects/{PROJECT_ID}/locations/{REGION}/endpoints/{endpoint_id}\"\n",
" )\n",
" endpoint = aiplatform.Endpoint(aip_endpoint_name)\n",
" else:\n",
" endpoint = aiplatform.Endpoint.create(\n",
" display_name=f\"{model_name}-endpoint\",\n",
" location=TPU_DEPLOYMENT_REGION,\n",
" dedicated_endpoint_enabled=use_dedicated_endpoint,\n",
" )\n",
"\n",
" if not base_model_id:\n",
" base_model_id = model_id\n",
"\n",
" if not tensor_parallel_size:\n",
" tensor_parallel_size = int(machine_type[-2])\n",
"\n",
" num_hosts = int(tpu_topology.split(\"x\")[0])\n",
"\n",
" vllmtpu_args = [\n",
" \"python\",\n",
" \"-m\",\n",
" \"vllm.entrypoints.api_server\",\n",
" \"--host=0.0.0.0\",\n",
" \"--port=7080\",\n",
" f\"--model={model_id}\",\n",
" f\"--tensor-parallel-size={tensor_parallel_size}\",\n",
" f\"--max-model-len={max_model_len}\",\n",
" ]\n",
"\n",
" if enable_chunked_prefill:\n",
" vllmtpu_args.append(\"--enable-chunked-prefill\")\n",
"\n",
" if enable_prefix_cache:\n",
" vllmtpu_args.append(\"--enable-prefix-caching\")\n",
"\n",
" if max_num_seqs is not None:\n",
" vllmtpu_args.append(f\"--max-num-seqs={max_num_seqs}\")\n",
"\n",
" if max_num_batched_tokens is not None:\n",
" vllmtpu_args.append(f\"--max-num-batched-tokens={max_num_batched_tokens}\")\n",
"\n",
" env_vars = {\n",
" \"MODEL_ID\": base_model_id,\n",
" \"DEPLOY_SOURCE\": \"notebook\",\n",
" \"VLLM_USE_V1\": \"1\",\n",
" }\n",
"\n",
" # HF_TOKEN is not a compulsory field and may not be defined.\n",
" try:\n",
" if HF_TOKEN:\n",
" env_vars[\"HF_TOKEN\"] = HF_TOKEN\n",
" except NameError:\n",
" pass\n",
"\n",
" model = aiplatform.Model.upload(\n",
" display_name=model_name,\n",
" serving_container_image_uri=vLLM_TPU_DOCKER_URI,\n",
" serving_container_args=vllmtpu_args,\n",
" serving_container_ports=[7080],\n",
" serving_container_predict_route=\"/generate\",\n",
" serving_container_health_route=\"/ping\",\n",
" serving_container_environment_variables=env_vars,\n",
" serving_container_shared_memory_size_mb=(16 * 1024), # 16 GB\n",
" serving_container_deployment_timeout=7200,\n",
" model_garden_source_model_name=(\n",
" f\"publishers/{publisher}/models/{publisher_model_id}\"\n",
" ),\n",
" location=TPU_DEPLOYMENT_REGION,\n",
" )\n",
"\n",
" model.deploy(\n",
" endpoint=endpoint,\n",
" machine_type=machine_type,\n",
" tpu_topology=tpu_topology if num_hosts > 1 else None,\n",
" deploy_request_timeout=1800,\n",
" min_replica_count=min_replica_count,\n",
" max_replica_count=max_replica_count,\n",
" system_labels={\n",
" \"NOTEBOOK_NAME\": \"model_garden_pytorch_llama3_3_tpu7x_deployment.ipynb\",\n",
" },\n",
" )\n",
" return model, endpoint\n",
"\n",
"\n",
"models[\"vllm_tpu\"], endpoints[\"vllm_tpu\"] = deploy_model_vllm_tpu(\n",
" model_name=common_util.get_job_name_with_datetime(prefix=\"llama3-3-serve\"),\n",
" model_id=model_id,\n",
" publisher=\"meta\",\n",
" publisher_model_id=\"llama3-3\",\n",
" base_model_id=hf_model_id,\n",
" tensor_parallel_size=tensor_parallel_size,\n",
" machine_type=machine_type,\n",
" max_model_len=max_model_len,\n",
" max_num_seqs=max_num_seqs,\n",
" max_num_batched_tokens=max_num_batched_tokens,\n",
" use_dedicated_endpoint=use_dedicated_endpoint,\n",
")\n",
"# @markdown Click \"Show Code\" to see more details."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "rDHsCOqvFYBi"
},
"outputs": [],
"source": [
"# @title Raw predict\n",
"\n",
"# @markdown Once deployment succeeds, you can send requests to the endpoint with text prompts. Sampling parameters supported by vLLM can be found [here](https://docs.vllm.ai/en/latest/dev/sampling_params.html).\n",
"\n",
"# @markdown Example:\n",
"\n",
"# @markdown ```\n",
"# @markdown Human: What is a car?\n",
"# @markdown Assistant: A car, or a motor car, is a road-connected human-transportation system used to move people or goods from one place to another. The term also encompasses a wide range of vehicles, including motorboats, trains, and aircrafts. Cars typically have four wheels, a cabin for passengers, and an engine or motor. They have been around since the early 19th century and are now one of the most popular forms of transportation, used for daily commuting, shopping, and other purposes.\n",
"# @markdown ```\n",
"# @markdown Additionally, you can moderate the generated text with Vertex AI. See [Moderate text documentation](https://cloud.google.com/natural-language/docs/moderating-text) for more details.\n",
"\n",
"# Loads an existing endpoint instance using the endpoint name:\n",
"# - Using `endpoint_name = endpoint.name` allows us to get the\n",
"# endpoint name of the endpoint `endpoint` created in the cell\n",
"# above.\n",
"# - Alternatively, you can set `endpoint_name = \"1234567890123456789\"` to load\n",
"# an existing endpoint with the ID 1234567890123456789.\n",
"# You may uncomment the code below to load an existing endpoint.\n",
"\n",
"# endpoint_name = \"\" # @param {type:\"string\"}\n",
"# aip_endpoint_name = (\n",
"# f\"projects/{PROJECT_ID}/locations/{REGION}/endpoints/{endpoint_name}\"\n",
"# )\n",
"# endpoint = aiplatform.Endpoint(aip_endpoint_name)\n",
"\n",
"prompt = \"What is a car?\" # @param {type: \"string\"}\n",
"# @markdown If you encounter an issue like `ServiceUnavailable: 503 Took too long to respond when processing`, you can reduce the maximum number of output tokens, by lowering `max_tokens`.\n",
"max_tokens = 50 # @param {type:\"integer\"}\n",
"temperature = 1.0 # @param {type:\"number\"}\n",
"top_p = 1.0 # @param {type:\"number\"}\n",
"top_k = 1 # @param {type:\"integer\"}\n",
"# @markdown Set `raw_response` to `True` to obtain the raw model output. Set `raw_response` to `False` to apply additional formatting in the structure of `\"Prompt:\\n{prompt.strip()}\\nOutput:\\n{output}\"`.\n",
"raw_response = False # @param {type:\"boolean\"}\n",
"\n",
"# Overrides parameters for inferences.\n",
"instances = [\n",
" {\n",
" \"prompt\": prompt,\n",
" \"max_tokens\": max_tokens,\n",
" \"temperature\": temperature,\n",
" \"top_p\": top_p,\n",
" \"top_k\": top_k,\n",
" \"raw_response\": raw_response,\n",
" },\n",
"]\n",
"response = endpoints[\"vllm_tpu\"].predict(\n",
" instances=instances, use_dedicated_endpoint=use_dedicated_endpoint\n",
")\n",
"\n",
"for prediction in response.predictions:\n",
" print(prediction)\n",
"\n",
"# @markdown Click \"Show Code\" to see more details."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "LSG9ITWTbTb7"
},
"outputs": [],
"source": [
"# @title Chat completion\n",
"\n",
"if use_dedicated_endpoint:\n",
" DEDICATED_ENDPOINT_DNS = endpoints[\"vllm_tpu\"].gca_resource.dedicated_endpoint_dns\n",
"ENDPOINT_RESOURCE_NAME = endpoints[\"vllm_tpu\"].resource_name\n",
"\n",
"# @title Chat Completions Inference\n",
"\n",
"# @markdown Once deployment succeeds, you can send requests to the endpoint using the OpenAI SDK.\n",
"\n",
"# @markdown First you will need to install the SDK and some auth-related dependencies.\n",
"\n",
"! pip install -qU openai google-auth requests\n",
"\n",
"# @markdown Next fill out some request parameters:\n",
"\n",
"user_message = \"How is your day going?\" # @param {type: \"string\"}\n",
"# @markdown If you encounter the issue like `ServiceUnavailable: 503 Took too long to respond when processing`, you can reduce the maximum number of output tokens, such as set `max_tokens` as 20.\n",
"max_tokens = 50 # @param {type: \"integer\"}\n",
"temperature = 1.0 # @param {type: \"number\"}\n",
"stream = False # @param {type: \"boolean\"}\n",
"\n",
"# @markdown Now we can send a request.\n",
"\n",
"import google.auth\n",
"import openai\n",
"\n",
"creds, project = google.auth.default()\n",
"auth_req = google.auth.transport.requests.Request()\n",
"creds.refresh(auth_req)\n",
"\n",
"BASE_URL = (\n",
" f\"https://{REGION}-aiplatform.googleapis.com/v1beta1/{ENDPOINT_RESOURCE_NAME}\"\n",
")\n",
"try:\n",
" if use_dedicated_endpoint:\n",
" BASE_URL = f\"https://{DEDICATED_ENDPOINT_DNS}/v1beta1/{ENDPOINT_RESOURCE_NAME}\"\n",
"except NameError:\n",
" pass\n",
"\n",
"client = openai.OpenAI(base_url=BASE_URL, api_key=creds.token)\n",
"\n",
"model_response = client.chat.completions.create(\n",
" model=\"\",\n",
" messages=[{\"role\": \"user\", \"content\": user_message}],\n",
" temperature=temperature,\n",
" max_tokens=max_tokens,\n",
" stream=stream,\n",
")\n",
"\n",
"if stream:\n",
" usage = None\n",
" contents = []\n",
" for chunk in model_response:\n",
" if chunk.usage is not None:\n",
" usage = chunk.usage\n",
" continue\n",
" print(chunk.choices[0].delta.content, end=\"\")\n",
" contents.append(chunk.choices[0].delta.content)\n",
" print(f\"\\n\\n{usage}\")\n",
"else:\n",
" print(model_response)\n",
"\n",
"# @markdown Click \"Show Code\" to see more details."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "tqtxJakIapIg"
},
"source": [
"## Clean up resources"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"cellView": "form",
"id": "kzgEmmd0aiUM"
},
"outputs": [],
"source": [
"# @title Delete the models and endpoints\n",
"\n",
"# @markdown Delete the experiment models and endpoints to recycle the resources\n",
"# @markdown and avoid unnecessary continuous charges that may incur.\n",
"\n",
"# Undeploy model and delete endpoint.\n",
"for endpoint in endpoints.values():\n",
" endpoint.delete(force=True)\n",
"\n",
"# Delete models.\n",
"for model in models.values():\n",
" model.delete()"
]
}
],
"metadata": {
"colab": {
"name": "model_garden_pytorch_llama3_3_tpu7x_deployment.ipynb",
"toc_visible": true
},
"kernelspec": {
"display_name": "Python 3",
"name": "python3"
}
},
"nbformat": 4,
"nbformat_minor": 0
}
@@ -97,7 +97,7 @@
"# @title Install Python Packages for Finetuning\n",
"\n",
"# @markdown 1. Install google-cloud-aiplatform package and restart the session if instructed.\n",
"! pip install --upgrade --quiet 'google-cloud-aiplatform>=1.66.0'\n",
"! pip install --upgrade --quiet google-cloud-aiplatform==1.130.0\n",
"\n",
"# @markdown 2. Install packages to validate dataset with template.\n",
"! pip install --upgrade --quiet gcsfs==2024.3.1\n",
@@ -681,6 +681,7 @@
" max_num_seqs: int = 256,\n",
" model_type: str = None,\n",
" enable_llama_tool_parser: bool = False,\n",
" is_spot: bool = False,\n",
") -> Tuple[aiplatform.Model, aiplatform.Endpoint]:\n",
" \"\"\"Deploys trained models with vLLM into Vertex AI.\"\"\"\n",
" endpoint = aiplatform.Endpoint.create(\n",
@@ -775,6 +776,7 @@
" accelerator_count=accelerator_count,\n",
" deploy_request_timeout=1800,\n",
" service_account=service_account,\n",
" spot=is_spot,\n",
" system_labels={\n",
" \"NOTEBOOK_NAME\": \"model_garden_pytorch_mistral_peft_tuning.ipynb\",\n",
" \"NOTEBOOK_ENVIRONMENT\": get_deploy_source(),\n",
@@ -97,7 +97,7 @@
"# @title Install Python Packages for Finetuning\n",
"\n",
"# @markdown 1. Install google-cloud-aiplatform package and restart the session if instructed.\n",
"! pip install --upgrade --quiet 'google-cloud-aiplatform>=1.66.0'\n",
"! pip install --upgrade --quiet google-cloud-aiplatform==1.130.0\n",
"\n",
"# @markdown 2. Install packages to validate dataset with template.\n",
"! pip install --upgrade --quiet gcsfs==2024.3.1\n",
@@ -686,6 +686,7 @@
" max_num_seqs: int = 256,\n",
" model_type: str = None,\n",
" enable_llama_tool_parser: bool = False,\n",
" is_spot: bool = False,\n",
") -> Tuple[aiplatform.Model, aiplatform.Endpoint]:\n",
" \"\"\"Deploys trained models with vLLM into Vertex AI.\"\"\"\n",
" endpoint = aiplatform.Endpoint.create(\n",
@@ -780,6 +781,7 @@
" accelerator_count=accelerator_count,\n",
" deploy_request_timeout=1800,\n",
" service_account=service_account,\n",
" spot=is_spot,\n",
" system_labels={\n",
" \"NOTEBOOK_NAME\": \"model_garden_pytorch_mixtral_peft_tuning.ipynb\",\n",
" \"NOTEBOOK_ENVIRONMENT\": get_deploy_source(),\n",
@@ -204,7 +204,6 @@
"\n",
"VERTEX_AI_MODEL_GARDEN_TIMESFM = \"gs://vertex-model-garden-public-us/timesfm\" # @param {type:\"string\", isTemplate:true} [\"gs://vertex-model-garden-public-us/timesfm\", \"gs://vertex-model-garden-public-eu/timesfm\", \"gs://vertex-model-garden-public-asia/timesfm\"]\n",
"MODEL_VARIANT = \"timesfm-2.0-500m-jax\" # @param [\"timesfm-2.0-500m-jax\"]\n",
"hf_model_id = \"google/\" + MODEL_VARIANT\n",
"\n",
"\n",
"print(\n",
@@ -320,7 +319,6 @@
"\n",
"def deploy_model(\n",
" model_name: str,\n",
" base_model_id: str,\n",
" checkpoint_path: str,\n",
" horizon: str,\n",
" machine_type: str = \"g2-standard-8\",\n",
@@ -347,13 +345,12 @@
"\n",
" model = aiplatform.Model.upload(\n",
" display_name=model_name_with_time,\n",
" artifact_uri=checkpoint_path,\n",
" serving_container_image_uri=SERVE_DOCKER_URI,\n",
" serving_container_ports=[8080],\n",
" serving_container_predict_route=\"/predict\",\n",
" serving_container_health_route=\"/health\",\n",
" serving_container_environment_variables={\n",
" \"MODEL_ID\": base_model_id,\n",
" \"MODEL_ID\": checkpoint_path,\n",
" \"DEPLOY_SOURCE\": deploy_source,\n",
" \"TIMESFM_HORIZON\": str(horizon),\n",
" \"TIMESFM_BACKEND\": timesfm_backend,\n",
@@ -385,7 +382,6 @@
"\n",
"models[\"timesfm\"], endpoints[\"timesfm\"] = deploy_model(\n",
" model_name=f\"timesfm-{MODEL_VARIANT}\",\n",
" base_model_id=hf_model_id,\n",
" checkpoint_path=checkpoint_path,\n",
" horizon=horizon,\n",
" machine_type=machine_type,\n",
+1 -1
View File
@@ -227,7 +227,7 @@ def parse_dir(directory: str) -> int:
elif tag == 'bigquery_ml':
tag = 'BigQuery ML'
elif tag == 'custom':
tag = 'Vertex AI Training'
tag = 'Vertex AI serverless training'
elif tag == 'experiments':
tag = 'Vertex AI Experiments'
elif tag == 'explainable_ai':
File diff suppressed because it is too large Load Diff
@@ -232,7 +232,7 @@
"source": [
"## Create vLLM Customer Container Image for Vertex AI\n",
"\n",
"Vertex AI requires [requests](https://cloud.google.com/vertex-ai/docs/predictions/custom-container-requirements#inference) and [responses](https://cloud.google.com/vertex-ai/docs/predictions/custom-container-requirements#response_requirements) in specific formats. vLLM API server implements OpenAI API protocol and therefore, it does not support the Vertex AI request and response requirements. Therefore, the vLLM API server (vllm.entrypoints.openai.api_server.py) needs to be updated to support Vertex AI request and response formats."
"The vLLM custom container image requires gcloud SDK so that it can download the model from Google Cloud Storage when Hugging Face token is not provided. The `entrypoint` script has also been updated to enable this feature."
]
},
{
@@ -232,7 +232,7 @@
"source": [
"## Create vLLM Customer Container Image for Vertex AI\n",
"\n",
"Vertex AI requires [requests](https://cloud.google.com/vertex-ai/docs/predictions/custom-container-requirements#inference) and [responses](https://cloud.google.com/vertex-ai/docs/predictions/custom-container-requirements#response_requirements) in specific formats. vLLM API server implements OpenAI API protocol and therefore, it does not support the Vertex AI request and response requirements. Therefore, the vLLM API server (vllm.entrypoints.openai.api_server.py) needs to be updated to support Vertex AI request and response formats."
"The vLLM custom container image requires gcloud SDK so that it can download the model from Google Cloud Storage when Hugging Face token is not provided. The `entrypoint` script has also been updated to enable this feature."
]
},
{
@@ -372,7 +372,7 @@
"source": [
"## Create vLLM Customer Container Image for Vertex AI\n",
"\n",
"Vertex AI requires [requests](https://cloud.google.com/vertex-ai/docs/predictions/custom-container-requirements#inference) and [responses](https://cloud.google.com/vertex-ai/docs/predictions/custom-container-requirements#response_requirements) in specific formats. vLLM API server implements OpenAI API protocol and therefore, it does not support the Vertex AI request and response requirements. Therefore, the vLLM API server (vllm.entrypoints.openai.api_server.py) needs to be updated to support Vertex AI request and response formats."
"The vLLM custom container image requires gcloud SDK so that it can download the model from Google Cloud Storage when Hugging Face token is not provided. The `entrypoint` script has also been updated to enable this feature."
]
},
{
@@ -232,7 +232,7 @@
"source": [
"## Create vLLM Customer Container Image for Vertex AI\n",
"\n",
"Vertex AI requires [requests](https://cloud.google.com/vertex-ai/docs/predictions/custom-container-requirements#inference) and [responses](https://cloud.google.com/vertex-ai/docs/predictions/custom-container-requirements#response_requirements) in specific formats. vLLM API server implements OpenAI API protocol and therefore, it does not support the Vertex AI request and response requirements. Therefore, the vLLM API server (vllm.entrypoints.openai.api_server.py) needs to be updated to support Vertex AI request and response formats."
"The vLLM custom container image requires gcloud SDK so that it can download the model from Google Cloud Storage when Hugging Face token is not provided. The `entrypoint` script has also been updated to enable this feature."
]
},
{