Compare commits

...
Author SHA1 Message Date
Andrew FerlitschandGitHub f2371b4f7d Merge branch 'main' into update_model_eval 2022-04-21 11:36:37 -07:00
Andrew Ferlitsch 20cb46cc29 feat: improve notebook for metric compare 2022-04-21 18:34:55 +00:00
Andrew Ferlitsch db1827cb74 feat: improve notebook for metric compare 2022-04-21 18:33:55 +00:00
0137cd106e adds Colab part and minor changes to ml_ops/stage2/get_started_vertex_experiments notebook in community folder (#493)
* updates the get-started-vertex-experiments notebook in the community folder

* ran linter test

* adds the costs section

* ran linter test

* adds colab part and minor changes

* ran linter test

Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
2022-04-21 09:20:29 -07:00
8b3b63d714 Adds Colab part and minor changes to ml_ops/stage2/get_started_automl_training notebook (#492)
* updates the get-started-automl-training notebook

* ran linter test

* adds --user flag during installation step

* ran linter test

* updates the clean up step

* ran linter test

* adds colab part and minor changes

* ran linter test

Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-04-21 09:19:34 -07:00
bf79916f29 Adds Colab part to ml_ops/stage1/get_started_bq_datasets notebook (#490)
* adds the updated mlops-stage1-get_started_bq_datasets notebook to the official branch and removes it from the community branch

* removes second instance of create_bigquery_dataset() function

* ran linter test successfully

* adds costs section

* ran linter test successfully

* updates the dependency installation step and GCS bucket explanation

* ran linter test

* adds pyarrow to the installations

* ran linter test

* removes unnecessary installations + adds silent install + moves the notebook back from official to community folder + adds IS_TESTING condition during clean-up

* ran linter test

* resolves the move up?? comment and builtin comment

* ran linter test

* updates textual content about package installation

* ran linter test

* resolves the future-tense and  dependency installations comments

* ran linter test

* updates the header according to template

* ran linter test

* adds Colab part and minor changes

* ran linter test

* updates the enable apis step in setup project section

* ran linter test

* changes vertex to vertex ai

* ran linter test

* moves temporary BQ table deletion outside the delete_storage condition

* ran linter test

Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
2022-04-21 09:17:49 -07:00
6bf462f79d Notebook fix to handle GCS outputs and resolve (AutoML Tabular Forecasting notebook error: no row field 'name' #453) (#477)
* notebook fix to handle gcs output

* linter test

* minor bug fix and markup added

* linter test

Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
2022-04-21 09:13:52 -07:00
Andrew FerlitschandGitHub 106cdee495 feat: update model eval metrics for comparison (#498)
* feat: improve notebook for metric compare

* feat: improve notebook for metric compare

* feat: improve notebook for metric compare
2022-04-20 19:21:23 -07:00
Andrew Ferlitsch b702bf3a9e feat: improve notebook for metric compare 2022-04-20 22:17:18 +00:00
Andrew Ferlitsch 36ea560aad feat: improve notebook for metric compare 2022-04-20 22:16:57 +00:00
Andrew Ferlitsch 48e744d004 feat: improve notebook for metric compare 2022-04-20 22:15:55 +00:00
dfb7301733 Inardini - feature store demo blog review (#484)
* review for blog

* linter code passed

Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
2022-04-18 15:17:17 -07:00
da707b2cbc Made minor changes to custom-tabular-bq-managed-dataset file (#473)
* modified notebook

* modified file

* ran linter test

Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
2022-04-18 15:07:07 -07:00
873ba9dde9 Made minor changes to get_started_with_rapid_prototyping_bqml_automl file (#459)
* modified file

* made linter changes

* made changes

* linter test issues resolved

* ran linter test

Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
2022-04-18 15:06:27 -07:00
Andrew FerlitschandGitHub 15c452f39d fix: Issue 451, add db-dtypes to requirements (#458)
* fix: issue 451

* fix: issue 451
2022-04-18 12:14:56 -07:00
9 changed files with 474 additions and 165 deletions
Binary file not shown.

Before

Width:  |  Height:  |  Size: 92 KiB

After

Width:  |  Height:  |  Size: 122 KiB

@@ -32,21 +32,22 @@
"<table align=\"left\">\n",
"\n",
" <td>\n",
" <a href=\"https://colab.research.google.com/github/GoogleCloudPlatform/vertex-ai-samples/blob/main/vertex-ai-samples/notebooks/community/feature_store/mobile_gaming_feature_store.ipynb\">\n",
" <a href=\"https://colab.research.google.com/github/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/notebook_template.ipynb\">\n",
" <img src=\"https://cloud.google.com/ml-engine/images/colab-logo-32px.png\" alt=\"Colab logo\"> Run in Colab\n",
" </a>\n",
" </td>\n",
" <td>\n",
" <a href=\"https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/vertex-ai-samples/notebooks/community/feature_store/mobile_gaming_feature_store.ipynb\">\n",
" <a href=\"https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/notebook_template.ipynb\">\n",
" <img src=\"https://cloud.google.com/ml-engine/images/github-logo-32px.png\" alt=\"GitHub logo\">\n",
" View on GitHub\n",
" </a>\n",
" </td>\n",
" <td>\n",
" <a href=\"https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/feature_store/mobile_gaming/mobile_gaming_feature_store.ipynb\">\n",
" <a href=\"https://console.cloud.google.com/vertex-ai/workbench/deploy-notebook?download_url=https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/notebook_template.ipynb\">\n",
" <img src=\"https://lh3.googleusercontent.com/UiNooY4LUgW_oTvpsNhPpQzsstV5W8F7rYgxgGBD85cWJoLmrOzhVs_ksK_vgx40SHs7jCqkTkCk=e14-rj-sc0xffffff-h130-w32\" alt=\"Vertex AI logo\">\n",
" Open in Vertex AI Workbench\n",
" </a>\n",
" </td>\n",
" </td> \n",
"</table>\n"
]
},
@@ -306,7 +307,7 @@
"outputs": [],
"source": [
"if PROJECT_ID == \"\" or PROJECT_ID is None:\n",
" PROJECT_ID = \"inardini-playground\" # @param {type:\"string\"}"
" PROJECT_ID = \"\" # @param {type:\"string\"}"
]
},
{
@@ -2070,7 +2071,7 @@
},
"outputs": [],
"source": [
"TRAIN_JOB_RESOURCE_NAME = \"projects/309823771116/locations/us-central1/customJobs/1016630991030059008\" # @param {type:\"string\"}"
"TRAIN_JOB_RESOURCE_NAME = \"\" # @param {type:\"string\"}"
]
},
{
@@ -2174,7 +2175,7 @@
"\n",
"Below you can see how it works\n",
"\n",
"<img src=\"./assets/online_serving_5.png\">\n",
"<img src=\"./assets/online_serving_5.png\" width=\"600\">\n",
"\n",
"But think about those features for a second. \n",
"\n",
@@ -38,6 +38,11 @@
" View on GitHub\n",
" </a>\n",
" </td>\n",
" <td>\n",
" <a href=\"https://colab.research.google.com/github/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/ml_ops/stage1/get_started_bq_datasets.ipynb\">\n",
" <img src=\"https://cloud.google.com/ml-engine/images/colab-logo-32px.png\\\" alt=\"Colab logo\"> Run in Colab\n",
" </a>\n",
" </td>\n",
" <td>\n",
" <a href=\"https://console.cloud.google.com/vertex-ai/workbench/deploy-notebook?download_url=https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/ml_ops/stage1/get_started_bq_datasets.ipynb\">\n",
" <img src=\"https://lh3.googleusercontent.com/UiNooY4LUgW_oTvpsNhPpQzsstV5W8F7rYgxgGBD85cWJoLmrOzhVs_ksK_vgx40SHs7jCqkTkCk=e14-rj-sc0xffffff-h130-w32\" alt=\"Vertex AI logo\">\n",
@@ -237,7 +242,7 @@
"\n",
"1. [Make sure that billing is enabled for your project](https://cloud.google.com/billing/docs/how-to/modify-project).\n",
"\n",
"1. [Enable the Vertex AI API](https://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com). {TODO: Update the APIs needed for your tutorial. Edit the API names, and update the link to append the API IDs, separating each one with a comma. For example, container.googleapis.com,cloudbuild.googleapis.com}\n",
"1. [Enable the Vertex AI, BigQuery, Compute Engine and Cloud Storage APIs](https://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com,bigquery,compute_component,storage_component).\n",
"\n",
"1. If you are running this notebook locally, you need to install the [Cloud SDK](https://cloud.google.com/sdk).\n",
"\n",
@@ -353,6 +358,66 @@
"TIMESTAMP = datetime.now().strftime(\"%Y%m%d%H%M%S\")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "77c385f0db59"
},
"source": [
"### Authenticate your Google Cloud account\n",
"\n",
"**If you are using Google Cloud Notebooks**, your environment is already authenticated. Skip this step.\n",
"\n",
"**If you are using Colab**, run the cell below and follow the instructions when prompted to authenticate your account via oAuth.\n",
"\n",
"**Otherwise**, follow these steps:\n",
"\n",
"In the Cloud Console, go to the [Create service account key](https://console.cloud.google.com/apis/credentials/serviceaccountkey) page.\n",
"\n",
"1. **Click Create service account**.\n",
"\n",
"2. In the **Service account name** field, enter a name, and click **Create**.\n",
"\n",
"3. In the **Grant this service account access to project** section, click the Role drop-down list. Type \"Vertex AI\" into the filter box, and select **Vertex AI Administrator**. Type \"Storage Object Admin\" into the filter box, and select **Storage Object Admin**.\n",
"\n",
"4. Click Create. A JSON file that contains your key downloads to your local environment.\n",
"\n",
"5. Enter the path to your service account key as the GOOGLE_APPLICATION_CREDENTIALS variable in the cell below and run the cell."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "535223fa4b84"
},
"outputs": [],
"source": [
"import os\n",
"import sys\n",
"\n",
"# If you are running this notebook in Colab, run this cell and follow the\n",
"# instructions to authenticate your GCP account. This provides access to your\n",
"# Cloud Storage bucket and lets you submit training jobs and prediction\n",
"# requests.\n",
"\n",
"# The Google Cloud Notebook product has specific requirements\n",
"IS_GOOGLE_CLOUD_NOTEBOOK = os.path.exists(\"/opt/deeplearning/metadata/env_version\")\n",
"\n",
"# If on Google Cloud Notebooks, then don't execute this code\n",
"if not IS_GOOGLE_CLOUD_NOTEBOOK:\n",
" if \"google.colab\" in sys.modules:\n",
" from google.colab import auth as google_auth\n",
"\n",
" google_auth.authenticate_user()\n",
"\n",
" # If you are running this notebook locally, replace the string below with the\n",
" # path to your service account key and run this cell to authenticate your GCP\n",
" # account.\n",
" elif not os.getenv(\"IS_TESTING\"):\n",
" %env GOOGLE_APPLICATION_CREDENTIALS ''"
]
},
{
"cell_type": "markdown",
"metadata": {
@@ -1112,7 +1177,9 @@
"\n",
"- Vertex AI Dataset resource\n",
"- Cloud Storage Bucket\n",
"- BigQuery Dataset"
"- BigQuery Dataset\n",
"\n",
"Set `delete_storage` to _True_ to delete the storage resources used in this notebook."
]
},
{
@@ -1127,13 +1194,15 @@
"\n",
"# Delete the dataset using the Vertex dataset object\n",
"dataset.delete()\n",
"# Delete the temporary BigQuery dataset\n",
"! bq rm -r -f $PROJECT_ID:$DATASET_ID\n",
"\n",
"if os.getenv(\"IS_TESTING\"):\n",
"delete_storage = False\n",
"if delete_storage or os.getenv(\"IS_TESTING\"):\n",
" # Delete the created GCS bucket\n",
" ! gsutil rm -r $BUCKET_URI\n",
" # Delete the created BigQuery datasets\n",
" ! bq rm -r -f $PROJECT_ID:$BQ_MY_DATASET\n",
" ! bq rm -r -f $PROJECT_ID:$DATASET_ID"
" ! bq rm -r -f $PROJECT_ID:$BQ_MY_DATASET"
]
}
],
@@ -38,6 +38,11 @@
" View on GitHub\n",
" </a>\n",
" </td>\n",
" <td>\n",
" <a href=\"https://colab.research.google.com/github/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/ml_ops/stage2/get_started_automl_training.ipynb\">\n",
" <img src=\"https://cloud.google.com/ml-engine/images/colab-logo-32px.png\\\" alt=\"Colab logo\"> Run in Colab\n",
" </a>\n",
" </td>\n",
" <td>\n",
" <a href=\"https://console.cloud.google.com/vertex-ai/workbench/deploy-notebook?download_url=https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/ml_ops/stage2/get_started_automl_training.ipynb\">\n",
" <img src=\"https://lh3.googleusercontent.com/UiNooY4LUgW_oTvpsNhPpQzsstV5W8F7rYgxgGBD85cWJoLmrOzhVs_ksK_vgx40SHs7jCqkTkCk=e14-rj-sc0xffffff-h130-w32\" alt=\"Vertex AI logo\">\n",
@@ -228,6 +233,23 @@
"id": "project_id"
},
"source": [
"### Set up your Google Cloud project\n",
"\n",
"**The following steps are required, regardless of your notebook environment.**\n",
"\n",
"1. [Select or create a Google Cloud project](https://console.cloud.google.com/cloud-resource-manager). When you first create an account, you get a $300 free credit towards your compute/storage costs.\n",
"\n",
"1. [Make sure that billing is enabled for your project](https://cloud.google.com/billing/docs/how-to/modify-project).\n",
"\n",
"1. [Enable the Vertex AI, Compute Engine and Cloud Storage APIs](https://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com,compute_component,storage_component).\n",
"\n",
"1. If you are running this notebook locally, you need to install the [Cloud SDK](https://cloud.google.com/sdk).\n",
"\n",
"1. Enter your project ID in the cell below. Then run the cell to make sure the\n",
"Cloud SDK uses the right project for all the commands in this notebook.\n",
"\n",
"**Note**: Jupyter runs lines prefixed with `!` as shell commands, and it interpolates Python variables prefixed with `$` into these commands.\n",
"\n",
"#### Set your project ID\n",
"\n",
"**If you don't know your project ID**, you may be able to get your project ID using `gcloud`."
@@ -328,6 +350,63 @@
"TIMESTAMP = datetime.now().strftime(\"%Y%m%d%H%M%S\")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "3ffa6b6c7cdb"
},
"source": [
"### Authenticate your Google Cloud account\n",
"\n",
"**If you are using Google Cloud Notebooks**, your environment is already authenticated. Skip this step.\n",
"\n",
"**If you are using Colab**, run the cell below and follow the instructions when prompted to authenticate your account via oAuth.\n",
"\n",
"**Otherwise**, follow these steps:\n",
"\n",
"In the Cloud Console, go to the [Create service account key](https://console.cloud.google.com/apis/credentials/serviceaccountkey) page.\n",
"\n",
"1. **Click Create service account**.\n",
"\n",
"2. In the **Service account name** field, enter a name, and click **Create**.\n",
"\n",
"3. In the **Grant this service account access to project** section, click the Role drop-down list. Type \"Vertex AI\" into the filter box, and select **Vertex AI Administrator**. Type \"Storage Object Admin\" into the filter box, and select **Storage Object Admin**.\n",
"\n",
"4. Click Create. A JSON file that contains your key downloads to your local environment.\n",
"\n",
"5. Enter the path to your service account key as the GOOGLE_APPLICATION_CREDENTIALS variable in the cell below and run the cell."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "2b72272258fc"
},
"outputs": [],
"source": [
"# If you are running this notebook in Colab, run this cell and follow the\n",
"# instructions to authenticate your GCP account. This provides access to your\n",
"# Cloud Storage bucket and lets you submit training jobs and prediction\n",
"# requests.\n",
"\n",
"import os\n",
"import sys\n",
"\n",
"# If on Google Cloud Notebook, then don't execute this code\n",
"if not os.path.exists(\"/opt/deeplearning/metadata/env_version\"):\n",
" if \"google.colab\" in sys.modules:\n",
" from google.colab import auth as google_auth\n",
"\n",
" google_auth.authenticate_user()\n",
"\n",
" # If you are running this notebook locally, replace the string below with the\n",
" # path to your service account key and run this cell to authenticate your GCP\n",
" # account.\n",
" elif not os.getenv(\"IS_TESTING\"):\n",
" %env GOOGLE_APPLICATION_CREDENTIALS ''"
]
},
{
"cell_type": "markdown",
"metadata": {
@@ -38,6 +38,11 @@
" View on GitHub\n",
" </a>\n",
" </td>\n",
" <td>\n",
" <a href=\"https://colab.research.google.com/github/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/ml_ops/stage2/get_started_vertex_experiments.ipynb\">\n",
" <img src=\"https://cloud.google.com/ml-engine/images/colab-logo-32px.png\\\" alt=\"Colab logo\"> Run in Colab\n",
" </a>\n",
" </td>\n",
" <td>\n",
" <a href=\"https://console.cloud.google.com/vertex-ai/workbench/deploy-notebook?download_url=https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/ml_ops/stage2/get_started_vertex_experiments.ipynb\">\n",
" <img src=\"https://lh3.googleusercontent.com/UiNooY4LUgW_oTvpsNhPpQzsstV5W8F7rYgxgGBD85cWJoLmrOzhVs_ksK_vgx40SHs7jCqkTkCk=e14-rj-sc0xffffff-h130-w32\" alt=\"Vertex AI logo\">\n",
@@ -183,6 +188,24 @@
"id": "project_id"
},
"source": [
"### Set up your Google Cloud project\n",
"\n",
"**The following steps are required, regardless of your notebook environment.**\n",
"\n",
"1. [Select or create a Google Cloud project](https://console.cloud.google.com/cloud-resource-manager). When you first create an account, you get a $300 free credit towards your compute/storage costs.\n",
"\n",
"1. [Make sure that billing is enabled for your project](https://cloud.google.com/billing/docs/how-to/modify-project).\n",
"\n",
"1. [Enable the Vertex AI, Compute Engine, Cloud Storage and Cloud Logging APIs](https://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com,compute_component,storage_component,logging).\n",
"\n",
"1. If you are running this notebook locally, you need to install the [Cloud SDK](https://cloud.google.com/sdk).\n",
"\n",
"1. Enter your project ID in the cell below. Then run the cell to make sure the\n",
"Cloud SDK uses the right project for all the commands in this notebook.\n",
"\n",
"**Note**: Jupyter runs lines prefixed with `!` as shell commands, and it interpolates Python variables prefixed with `$` into these commands.\n",
"\n",
"\n",
"#### Set your project ID\n",
"\n",
"**If you don't know your project ID**, you may be able to get your project ID using `gcloud`."
@@ -283,6 +306,63 @@
"TIMESTAMP = datetime.now().strftime(\"%Y%m%d%H%M%S\")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "f3bd8c0d0469"
},
"source": [
"### Authenticate your Google Cloud account\n",
"\n",
"**If you are using Google Cloud Notebooks**, your environment is already authenticated. Skip this step.\n",
"\n",
"**If you are using Colab**, run the cell below and follow the instructions when prompted to authenticate your account via oAuth.\n",
"\n",
"**Otherwise**, follow these steps:\n",
"\n",
"In the Cloud Console, go to the [Create service account key](https://console.cloud.google.com/apis/credentials/serviceaccountkey) page.\n",
"\n",
"1. **Click Create service account**.\n",
"\n",
"2. In the **Service account name** field, enter a name, and click **Create**.\n",
"\n",
"3. In the **Grant this service account access to project** section, click the Role drop-down list. Type \"Vertex AI\" into the filter box, and select **Vertex AI Administrator**. Type \"Storage Object Admin\" into the filter box, and select **Storage Object Admin**.\n",
"\n",
"4. Click Create. A JSON file that contains your key downloads to your local environment.\n",
"\n",
"5. Enter the path to your service account key as the GOOGLE_APPLICATION_CREDENTIALS variable in the cell below and run the cell."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "e0953a00668e"
},
"outputs": [],
"source": [
"# If you are running this notebook in Colab, run this cell and follow the\n",
"# instructions to authenticate your GCP account. This provides access to your\n",
"# Cloud Storage bucket and lets you submit training jobs and prediction\n",
"# requests.\n",
"\n",
"import os\n",
"import sys\n",
"\n",
"# If on Google Cloud Notebook, then don't execute this code\n",
"if not os.path.exists(\"/opt/deeplearning/metadata/env_version\"):\n",
" if \"google.colab\" in sys.modules:\n",
" from google.colab import auth as google_auth\n",
"\n",
" google_auth.authenticate_user()\n",
"\n",
" # If you are running this notebook locally, replace the string below with the\n",
" # path to your service account key and run this cell to authenticate your GCP\n",
" # account.\n",
" elif not os.getenv(\"IS_TESTING\"):\n",
" %env GOOGLE_APPLICATION_CREDENTIALS ''"
]
},
{
"cell_type": "markdown",
"metadata": {
@@ -505,7 +585,7 @@
"from google.cloud.logging.handlers import CloudLoggingHandler\n",
"\n",
"# Connect to the Cloud Logging service\n",
"cl_client = google.cloud.logging.Client()\n",
"cl_client = google.cloud.logging.Client(project=PROJECT_ID)\n",
"handler = CloudLoggingHandler(cl_client, name=\"mylog\")\n",
"\n",
"# Create a logger instance and logging level\n",
@@ -612,7 +692,7 @@
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "bd7fb247cbae"
"id": "1ed46e349cf2"
},
"outputs": [],
"source": [
@@ -32,14 +32,20 @@
"# E2E ML on GCP: MLOps stage 3 : Get started with rapid prototyping with AutoML and BQML\n",
"<table align=\"left\">\n",
" <td>\n",
" <a href=\"https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/ml_ops/stage6/get_started_with_rapid_prototyping.ipynb\">\n",
" <a href=\"https://colab.research.google.com/github/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/ml_ops/stage3/get_started_with_rapid_prototyping_bqml_automl.ipynb\">\n",
" <img src=\"https://cloud.google.com/ml-engine/images/colab-logo-32px.png\" alt=\"Colab logo\"> Run in Colab\n",
" </a>\n",
" </td>\n",
" <td>\n",
" <a href=\"https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/ml_ops/stage3/get_started_with_rapid_prototyping_bqml_automl.ipynb\">\n",
" <img src=\"https://cloud.google.com/ml-engine/images/github-logo-32px.png\" alt=\"GitHub logo\">\n",
" View on GitHub\n",
" </a>\n",
" </td>\n",
" <td>\n",
" <a href=\"https://console.cloud.google.com/ai/platform/notebooks/deploy-notebook?download_url=https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/ml_ops/stage6/get_started_with_rapid_prototyping.ipynb\">\n",
" Open in Vertex Workbench\n",
" <a href=\"https://console.cloud.google.com/vertex-ai/workbench/deploy-notebook?download_url=https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/ml_ops/stage3/get_started_with_rapid_prototyping_bqml_automl.ipynb\">\n",
" <img src=\"https://lh3.googleusercontent.com/UiNooY4LUgW_oTvpsNhPpQzsstV5W8F7rYgxgGBD85cWJoLmrOzhVs_ksK_vgx40SHs7jCqkTkCk=e14-rj-sc0xffffff-h130-w32\" alt=\"Vertex AI logo\">\n",
" Open in Vertex AI Workbench\n",
" </a>\n",
" </td>\n",
"</table>\n",
@@ -252,9 +258,7 @@
"if os.path.exists(\"/opt/deeplearning/metadata/env_version\"):\n",
" USER_FLAG = \"--user\"\n",
"else:\n",
" USER_FLAG = \"\"\n",
"\n",
"! pip3 install --quiet --upgrade google-cloud-aiplatform {USER_FLAG}"
" USER_FLAG = \"\""
]
},
{
@@ -274,6 +278,7 @@
},
"outputs": [],
"source": [
"! pip3 install --quiet --upgrade google-cloud-aiplatform {USER_FLAG}\n",
"! pip3 install {USER_FLAG} --quiet -U google-cloud-pipeline-components==1.0 kfp\n",
"! pip3 install {USER_FLAG} --quiet --upgrade google-cloud-bigquery"
]
@@ -414,6 +419,8 @@
},
"outputs": [],
"source": [
"import os\n",
"\n",
"PROJECT_ID = \"[your-project-id]\" # @param {type:\"string\"}"
]
},
@@ -471,7 +478,9 @@
},
"outputs": [],
"source": [
"REGION = \"us-central1\" # @param {type: \"string\"}"
"REGION = \"[your-region]\" # @param {type:\"string\"}\n",
"if REGION == \"[your-region]\":\n",
" REGION = \"us-central1\""
]
},
{
@@ -667,6 +676,7 @@
"from typing import NamedTuple\n",
"\n",
"import google.cloud.aiplatform as aip\n",
"from google.cloud import bigquery\n",
"from kfp import dsl\n",
"from kfp.v2 import compiler\n",
"from kfp.v2.dsl import Artifact, Input, Metrics, Output, component"
@@ -1429,6 +1439,17 @@
"- `validate_infrastructure`: Validate the deployed model serving infrastructure."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "040e82bc1646"
},
"outputs": [],
"source": [
"DISPLAY_NAME = \"rapid-prototyping\""
]
},
{
"cell_type": "code",
"execution_count": null,
@@ -1453,8 +1474,7 @@
" from google_cloud_pipeline_components.types import artifact_types\n",
" from google_cloud_pipeline_components.v1.bigquery import (\n",
" BigqueryCreateModelJobOp, BigqueryEvaluateModelJobOp,\n",
" BigqueryExportModelJobOp, BigqueryPredictModelJobOp,\n",
" BigqueryQueryJobOp)\n",
" BigqueryExportModelJobOp)\n",
" from google_cloud_pipeline_components.v1.endpoint import (EndpointCreateOp,\n",
" ModelDeployOp)\n",
" from google_cloud_pipeline_components.v1.model import ModelUploadOp\n",
@@ -1670,7 +1690,7 @@
"PIPELINE_ROOT = f\"{BUCKET_URI}/pipeline_root\"\n",
"image_prefix = REGION.split(\"-\")[0]\n",
"BQML_SERVING_CONTAINER_IMAGE_URI = (\n",
" f\"{image_prefix}-docker.pkg.dev/vertex-ai/prediction/tf2-cpu.2-6:latest\"\n",
" f\"{image_prefix}-docker.pkg.dev/vertex-ai/prediction/tf2-cpu.2-8:latest\"\n",
")\n",
"\n",
"BQ_DATASET = \"rapid_prototype\" # j90wipxexhrgq3cquanc5\" # @param {type:\"string\"}\n",
@@ -1678,7 +1698,6 @@
"BQ_LOCATION = BQ_LOCATION.upper()\n",
"BQML_EXPORT_LOCATION = f\"{BUCKET_URI}/artifacts/bqml\"\n",
"\n",
"DISPLAY_NAME = \"rapid-prototyping\"\n",
"ENDPOINT_DISPLAY_NAME = f\"{DISPLAY_NAME}_endpoint\"\n",
"\n",
"compiler.Compiler().compile(\n",
@@ -1706,7 +1725,7 @@
" template_path=PIPELINE_JSON_PKG_PATH,\n",
" pipeline_root=PIPELINE_ROOT,\n",
" parameter_values=pipeline_params,\n",
" enable_caching=True,\n",
" enable_caching=False,\n",
")\n",
"\n",
"response = pipeline_job.submit()"
@@ -1758,96 +1777,72 @@
},
"outputs": [],
"source": [
"delete = True # set to True if you want to delete resources created in this tutorial.\n",
"delete_bucket = True\n",
"\n",
"print(\"Will delete endpoint\")\n",
"\n",
"delete_vertex_dataset = True and delete\n",
"delete_pipeline = True and delete\n",
"delete_model = True and delete\n",
"delete_endpoint = True and delete\n",
"delete_batchjob = True and delete\n",
"delete_bucket = True and delete\n",
"delete_bq_dataset = True and delete\n",
"endpoints = aip.Endpoint.list(\n",
" filter=f\"display_name={DISPLAY_NAME}_endpoint\", order_by=\"create_time\"\n",
")\n",
"endpoint = endpoints[0]\n",
"endpoint.undeploy_all()\n",
"aip.Endpoint.delete(endpoint.resource_name)\n",
"print(\"Deleted endpoint:\", endpoint)\n",
"\n",
"try:\n",
" if delete_endpoint and \"DISPLAY_NAME\" in globals():\n",
" print(\"Will delete endpoint\")\n",
" endpoints = aip.Endpoint.list(\n",
" filter=f\"display_name={DISPLAY_NAME}_endpoint\", order_by=\"create_time\"\n",
" )\n",
" endpoint = endpoints[0]\n",
" endpoint.undeploy_all()\n",
" aip.Endpoint.delete(endpoint.resource_name)\n",
" print(\"Deleted endpoint:\", endpoint)\n",
"except Exception as e:\n",
" print(e)\n",
"\n",
"if delete_model and \"DISPLAY_NAME\" in globals():\n",
" print(\"Will delete models\")\n",
" suffix_list = [\"bqml\", \"automl\", \"best\"]\n",
" for suffix in suffix_list:\n",
" try:\n",
" model_display_name = f\"{DISPLAY_NAME}_{suffix}\"\n",
" print(\"Will delete model with name \" + model_display_name)\n",
" models = aip.Model.list(\n",
" filter=f\"display_name={model_display_name}\", order_by=\"create_time\"\n",
" )\n",
"\n",
" model = models[0]\n",
" aip.Model.delete(model)\n",
" print(\"Deleted model:\", model)\n",
" except Exception as e:\n",
" print(e)\n",
"\n",
"if delete_vertex_dataset and \"DISPLAY_NAME\" in globals():\n",
" print(\"Will delete Vertex dataset\")\n",
"print(\"Will delete models\")\n",
"suffix_list = [\"bqml\", \"automl\", \"best\"]\n",
"for suffix in suffix_list:\n",
" try:\n",
" datasets = aip.TabularDataset.list(\n",
" filter=f\"display_name={DISPLAY_NAME}\", order_by=\"create_time\"\n",
" model_display_name = f\"{DISPLAY_NAME}_{suffix}\"\n",
" print(\"Will delete model with name \" + model_display_name)\n",
" models = aip.Model.list(\n",
" filter=f\"display_name={model_display_name}\", order_by=\"create_time\"\n",
" )\n",
"\n",
" dataset = datasets[0]\n",
" aip.TabularDataset.delete(dataset)\n",
" print(\"Deleted Vertex dataset:\", dataset)\n",
" model = models[0]\n",
" aip.Model.delete(model)\n",
" print(\"Deleted model:\", model)\n",
" except Exception as e:\n",
" print(e)\n",
"\n",
"\n",
"try:\n",
" if delete_pipeline and \"DISPLAY_NAME\" in globals():\n",
" pipelines = aip.PipelineJob.list(\n",
" filter=f\"pipeline_name={DISPLAY_NAME}\", order_by=\"create_time\"\n",
" )\n",
" pipeline = pipelines[0]\n",
" aip.PipelineJob.delete(pipeline)\n",
" print(\"Deleted pipeline:\", pipeline)\n",
"except Exception as e:\n",
" print(e)\n",
"print(\"Will delete Vertex dataset\")\n",
"\n",
"if delete_bq_dataset and \"DISPLAY_NAME\" in globals():\n",
" from google.cloud import bigquery\n",
"datasets = aip.TabularDataset.list(\n",
" filter=f\"display_name={DISPLAY_NAME}\", order_by=\"create_time\"\n",
")\n",
"\n",
" try:\n",
" # Construct a BigQuery client object.\n",
"\n",
" bq_client = bigquery.Client(project=PROJECT_ID, location=BQ_LOCATION)\n",
"\n",
" # TODO(developer): Set model_id to the ID of the model to fetch.\n",
" dataset_id = f\"{PROJECT_ID}.{BQ_DATASET}\"\n",
"\n",
" print(f\"Will delete BQ dataset '{dataset_id}' from location {BQ_LOCATION}.\")\n",
" # Use the delete_contents parameter to delete a dataset and its contents.\n",
" # Use the not_found_ok parameter to not receive an error if the dataset has already been deleted.\n",
" bq_client.delete_dataset(\n",
" dataset_id, delete_contents=True, not_found_ok=True\n",
" ) # Make an API request.\n",
"\n",
" print(f\"Deleted BQ dataset '{dataset_id}' from location {BQ_LOCATION}.\")\n",
" except Exception as e:\n",
" print(e)\n",
"dataset = datasets[0]\n",
"aip.TabularDataset.delete(dataset)\n",
"print(\"Deleted Vertex dataset:\", dataset)\n",
"\n",
"\n",
"if delete_bucket and \"BUCKET_URI\" in globals():\n",
"pipelines = aip.PipelineJob.list(\n",
" filter=f\"pipeline_name={DISPLAY_NAME}\", order_by=\"create_time\"\n",
")\n",
"pipeline = pipelines[0]\n",
"aip.PipelineJob.delete(pipeline)\n",
"print(\"Deleted pipeline:\", pipeline)\n",
"\n",
"\n",
"# Construct a BigQuery client object.\n",
"\n",
"bq_client = bigquery.Client(project=PROJECT_ID, location=BQ_LOCATION)\n",
"\n",
"# TODO(developer): Set dataset_id to the ID of the dataset to fetch.\n",
"dataset_id = f\"{PROJECT_ID}.{BQ_DATASET}\"\n",
"\n",
"print(f\"Will delete BQ dataset '{dataset_id}' from location {BQ_LOCATION}.\")\n",
"# Use the delete_contents parameter to delete a dataset and its contents.\n",
"# Use the not_found_ok parameter to not receive an error if the dataset has already been deleted.\n",
"bq_client.delete_dataset(\n",
" dataset_id, delete_contents=True, not_found_ok=True\n",
") # Make an API request.\n",
"\n",
"print(f\"Deleted BQ dataset '{dataset_id}' from location {BQ_LOCATION}.\")\n",
"\n",
"\n",
"if delete_bucket or os.getenv(\"IS_TESTING\"):\n",
" ! gsutil rm -r $BUCKET_URI"
]
}
@@ -8,7 +8,7 @@
},
"outputs": [],
"source": [
"# Copyright 2021 Google LLC\n",
"# Copyright 2022 Google LLC\n",
"#\n",
"# Licensed under the Apache License, Version 2.0 (the \"License\");\n",
"# you may not use this file except in compliance with the License.\n",
@@ -45,6 +45,7 @@
" </td>\n",
" <td>\n",
" <a href=\"https://console.cloud.google.com/vertex-ai/notebooks/deploy-notebook?download_url=https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/master/notebooks/official/automl/sdk_automl_tabular_forecasting_batch.ipynb\">\n",
" <img src=\"https://lh3.googleusercontent.com/UiNooY4LUgW_oTvpsNhPpQzsstV5W8F7rYgxgGBD85cWJoLmrOzhVs_ksK_vgx40SHs7jCqkTkCk=e14-rj-sc0xffffff-h130-w32\" alt=\"Vertex AI logo\">\n",
" Open in Vertex AI Workbench\n",
" </a>\n",
" </td>\n",
@@ -173,19 +174,8 @@
"else:\n",
" USER_FLAG = \"\"\n",
"\n",
"! pip3 install --upgrade google-cloud-aiplatform $USER_FLAG"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "install_tensorflow"
},
"outputs": [],
"source": [
"if os.getenv(\"IS_TESTING\"):\n",
" ! pip3 install --upgrade tensorflow $USER_FLAG"
"! pip3 install --upgrade google-cloud-aiplatform $USER_FLAG\n",
"! pip3 install --upgrade tensorflow $USER_FLAG"
]
},
{
@@ -419,7 +409,7 @@
},
"outputs": [],
"source": [
"BUCKET_NAME = \"gs://[your-bucket-name]\" # @param {type:\"string\"}"
"BUCKET_URI = \"gs://[your-bucket-name]\" # @param {type:\"string\"}"
]
},
{
@@ -430,8 +420,8 @@
},
"outputs": [],
"source": [
"if BUCKET_NAME == \"\" or BUCKET_NAME is None or BUCKET_NAME == \"gs://[your-bucket-name]\":\n",
" BUCKET_NAME = \"gs://\" + PROJECT_ID + \"aip-\" + TIMESTAMP"
"if BUCKET_URI == \"\" or BUCKET_URI is None or BUCKET_URI == \"gs://[your-bucket-name]\":\n",
" BUCKET_URI = \"gs://\" + PROJECT_ID + \"aip-\" + TIMESTAMP"
]
},
{
@@ -451,7 +441,7 @@
},
"outputs": [],
"source": [
"! gsutil mb -l $REGION $BUCKET_NAME"
"! gsutil mb -l $REGION $BUCKET_URI"
]
},
{
@@ -471,7 +461,7 @@
},
"outputs": [],
"source": [
"! gsutil ls -al $BUCKET_NAME"
"! gsutil ls -al $BUCKET_URI"
]
},
{
@@ -516,7 +506,7 @@
},
"outputs": [],
"source": [
"aiplatform.init(project=PROJECT_ID, staging_bucket=BUCKET_NAME)"
"aiplatform.init(project=PROJECT_ID, staging_bucket=BUCKET_URI)"
]
},
{
@@ -798,6 +788,15 @@
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "99b7a9287ba6"
},
"source": [
"`batch_predict` can export predictions either to BigQuery or GCS. The BQ option is commented out below and the predictions will be exported to the BUCKET_URI."
]
},
{
"cell_type": "code",
"execution_count": null,
@@ -814,7 +813,8 @@
" job_display_name=f\"iowa_liquor_sales_forecasting_predictions_{TIMESTAMP}\",\n",
" bigquery_source=PREDICTION_DATASET_BQ_PATH,\n",
" instances_format=\"bigquery\",\n",
" bigquery_destination_prefix=batch_predict_bq_output_uri_prefix,\n",
" # bigquery_destination_prefix=batch_predict_bq_output_uri_prefix,\n",
" gcs_destination_prefix=BUCKET_URI,\n",
" predictions_format=\"bigquery\",\n",
" sync=False,\n",
")\n",
@@ -1018,6 +1018,9 @@
},
"outputs": [],
"source": [
"# Set this to true only if you'd like to delete your bucket\n",
"delete_bucket = False\n",
"\n",
"# Delete dataset\n",
"dataset.delete()\n",
"\n",
@@ -1030,8 +1033,8 @@
"# Delete batch prediction job\n",
"batch_prediction_job.delete()\n",
"\n",
"if os.getenv(\"IS_TESTING\"):\n",
" ! gsutil rm -r $BUCKET_NAME"
"if delete_bucket or os.getenv(\"IS_TESTING\"):\n",
" ! gsutil rm -r $BUCKET_URI"
]
}
],
@@ -8,7 +8,7 @@
},
"outputs": [],
"source": [
"# Copyright 2020 Google LLC\n",
"# Copyright 2022 Google LLC\n",
"#\n",
"# Licensed under the Apache License, Version 2.0 (the \"License\");\n",
"# you may not use this file except in compliance with the License.\n",
@@ -43,6 +43,12 @@
" View on GitHub\n",
" </a>\n",
" </td>\n",
" <td>\n",
" <a href=\"https://console.cloud.google.com/vertex-ai/workbench/deploy-notebook?download_url=https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/official/custom/custom-tabular-bq-managed-dataset.ipynb\">\n",
" <img src=\"https://lh3.googleusercontent.com/UiNooY4LUgW_oTvpsNhPpQzsstV5W8F7rYgxgGBD85cWJoLmrOzhVs_ksK_vgx40SHs7jCqkTkCk=e14-rj-sc0xffffff-h130-w32\" alt=\"Vertex AI logo\">\n",
" Open in Vertex AI Workbench\n",
" </a>\n",
" </td> \n",
"</table>"
]
},
@@ -55,7 +61,7 @@
"## Overview\n",
"\n",
"\n",
"This tutorial demonstrates how to use the Vertex SDK for Python to train and deploy a custom tabular classification model for online prediction."
"This tutorial demonstrates how to use the Vertex AI SDK for Python to train and deploy a custom tabular classification model for online prediction."
]
},
{
@@ -116,7 +122,7 @@
"source": [
"## Installation\n",
"\n",
"Install the latest version of Vertex SDK for Python."
"Install the latest version of Vertex AI SDK for Python."
]
},
{
@@ -425,8 +431,8 @@
},
"outputs": [],
"source": [
"BUCKET_NAME = \"\" # @param {type:\"string\"}\n",
"REGION = \"us-central1\" # @param {type:\"string\"}"
"BUCKET_URI = \"gs://[your-bucket-name]\"\n",
"REGION = \"[your-region]\" # @param {type:\"string\"}"
]
},
{
@@ -437,8 +443,11 @@
},
"outputs": [],
"source": [
"if BUCKET_NAME == \"\" or BUCKET_NAME is None or BUCKET_NAME == \"gs://[your-bucket-name]\":\n",
" BUCKET_NAME = \"gs://\" + PROJECT_ID + \"aip-\" + TIMESTAMP"
"if BUCKET_URI == \"\" or BUCKET_URI is None or BUCKET_URI == \"gs://[your-bucket-name]\":\n",
" BUCKET_URI = \"gs://\" + PROJECT_ID + \"aip-\" + TIMESTAMP\n",
"\n",
"if REGION == \"[your-region]\":\n",
" REGION = \"us-central1\""
]
},
{
@@ -458,7 +467,7 @@
},
"outputs": [],
"source": [
"! gsutil mb -l $REGION $BUCKET_NAME"
"! gsutil mb -l $REGION $BUCKET_URI\n"
]
},
{
@@ -478,7 +487,7 @@
},
"outputs": [],
"source": [
"! gsutil ls -al $BUCKET_NAME"
"! gsutil ls -al $BUCKET_URI"
]
},
{
@@ -498,9 +507,9 @@
"id": "import_aip"
},
"source": [
"### Import Vertex SDK for Python\n",
"### Import Vertex AI SDK for Python\n",
"\n",
"Import the Vertex SDK for Python into your Python environment and initialize it."
"Import the Vertex AI SDK for Python into your Python environment and initialize it."
]
},
{
@@ -511,13 +520,15 @@
},
"outputs": [],
"source": [
"import json\n",
"import os\n",
"import sys\n",
"\n",
"from google.cloud import aiplatform\n",
"import numpy as np\n",
"from google.cloud import aiplatform, bigquery\n",
"from google.cloud.aiplatform import gapic as aip\n",
"\n",
"aiplatform.init(project=PROJECT_ID, location=REGION, staging_bucket=BUCKET_NAME)"
"aiplatform.init(project=PROJECT_ID, location=REGION, staging_bucket=BUCKET_URI)"
]
},
{
@@ -575,8 +586,8 @@
},
"outputs": [],
"source": [
"TRAIN_VERSION = \"tf-gpu.2-4\"\n",
"DEPLOY_VERSION = \"tf2-gpu.2-4\"\n",
"TRAIN_VERSION = \"tf-gpu.2-8\"\n",
"DEPLOY_VERSION = \"tf2-gpu.2-8\"\n",
"\n",
"TRAIN_IMAGE = \"us-docker.pkg.dev/vertex-ai/training/{}:latest\".format(TRAIN_VERSION)\n",
"DEPLOY_IMAGE = \"us-docker.pkg.dev/vertex-ai/prediction/{}:latest\".format(DEPLOY_VERSION)\n",
@@ -665,11 +676,7 @@
},
"outputs": [],
"source": [
"import json\n",
"\n",
"import numpy as np\n",
"# Calculate mean and std across all rows\n",
"from google.cloud import bigquery\n",
"\n",
"NA_VALUES = [\"NA\", \".\"]\n",
"\n",
@@ -724,7 +731,7 @@
" json.dump(mean_and_std, outfile)\n",
"\n",
"# Save to the staging bucket\n",
"! gsutil cp {MEAN_AND_STD_JSON_FILE} {BUCKET_NAME}"
"! gsutil cp {MEAN_AND_STD_JSON_FILE} {BUCKET_URI}"
]
},
{
@@ -807,7 +814,7 @@
" \"--epochs=\" + str(EPOCHS),\n",
" \"--batch_size=\" + str(BATCH_SIZE),\n",
" \"--distribute=\" + TRAIN_STRATEGY,\n",
" \"--mean_and_std_json_file=\" + f\"{BUCKET_NAME}/{MEAN_AND_STD_JSON_FILE}\",\n",
" \"--mean_and_std_json_file=\" + f\"{BUCKET_URI}/{MEAN_AND_STD_JSON_FILE}\",\n",
"]"
]
},
@@ -853,9 +860,9 @@
"from google.cloud import storage\n",
"\n",
"# Read environmental variables\n",
"training_data_uri = os.environ[\"AIP_TRAINING_DATA_URI\"]\n",
"validation_data_uri = os.environ[\"AIP_VALIDATION_DATA_URI\"]\n",
"test_data_uri = os.environ[\"AIP_TEST_DATA_URI\"]\n",
"training_data_uri = os.getenv(\"AIP_TRAINING_DATA_URI\")\n",
"validation_data_uri = os.getenv(\"AIP_VALIDATION_DATA_URI\")\n",
"test_data_uri = os.getenv(\"AIP_TEST_DATA_URI\")\n",
"\n",
"# Read args\n",
"parser = argparse.ArgumentParser()\n",
@@ -1128,7 +1135,7 @@
"# Train the model\n",
"model.fit(dataset_train, epochs=args.epochs, validation_data=dataset_validation)\n",
"\n",
"tf.saved_model.save(model, os.environ[\"AIP_MODEL_DIR\"])\n",
"tf.saved_model.save(model, os.getenv(\"AIP_MODEL_DIR\"))\n",
"\n",
"df_test.head()"
]
@@ -1175,7 +1182,7 @@
" display_name=JOB_NAME,\n",
" script_path=\"task.py\",\n",
" container_uri=TRAIN_IMAGE,\n",
" requirements=[\"google-cloud-bigquery>=2.20.0\"],\n",
" requirements=[\"google-cloud-bigquery>=2.20.0\", \"db-dtypes\"],\n",
" model_serving_container_image_uri=DEPLOY_IMAGE,\n",
")\n",
"\n",
@@ -1501,10 +1508,6 @@
},
"outputs": [],
"source": [
"delete_training_job = True\n",
"delete_model = True\n",
"delete_endpoint = True\n",
"\n",
"# Warning: Setting this to true will delete everything in your bucket\n",
"delete_bucket = False\n",
"\n",
@@ -1517,8 +1520,8 @@
"# Delete the endpoint\n",
"endpoint.delete()\n",
"\n",
"if delete_bucket and \"BUCKET_NAME\" in globals():\n",
" ! gsutil -m rm -r $BUCKET_NAME"
"if delete_bucket or os.getenv(\"IS_TESTING\"):\n",
" ! gsutil rm -r $BUCKET_URI"
]
}
],
@@ -90,7 +90,8 @@
"\n",
"- Upload a pre-trained model as a `Model` resource.\n",
"- Run a `BatchPredictionJob` on the `Model` resource with ground truth data.\n",
"- Generate Evaluation metrics about the `Model`.\n"
"- Generate evaluation `Metrics` artifact about the `Model` resource.\n",
"- Compare the evaluation metrics to a threshold.\n"
]
},
{
@@ -154,7 +155,7 @@
"source": [
"## Installation\n",
"\n",
"Install the latest version of Vertex SDK for Python."
"Install the latest version of Vertex AI SDK for Python."
]
},
{
@@ -193,7 +194,8 @@
},
"outputs": [],
"source": [
"! pip3 install --upgrade google-cloud-pipeline-components $USER_FLAG"
"! pip3 install --upgrade google-cloud-pipeline-components $USER_FLAG\n",
"! pip3 install --upgrade kfp $USER_FLAG"
]
},
{
@@ -231,7 +233,7 @@
"id": "check_versions"
},
"source": [
"Check the versions of the packages you installed. The KFP SDK version should be >=1.6."
"Check the versions of the packages you installed. "
]
},
{
@@ -540,10 +542,21 @@
"):\n",
" # Get your GCP project id from gcloud\n",
" shell_output = !gcloud auth list 2>/dev/null\n",
" SERVICE_ACCOUNT = shell_output[2].replace('*', '').strip()\n",
" SERVICE_ACCOUNT = shell_output[2].replace(\"*\", \"\").strip()\n",
" print(\"Service Account:\", SERVICE_ACCOUNT)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "9ca2fb92cb31"
},
"outputs": [],
"source": [
"shell_output[2].replace(\"*\", \"\").strip()"
]
},
{
"cell_type": "markdown",
"metadata": {
@@ -588,7 +601,8 @@
},
"outputs": [],
"source": [
"import google.cloud.aiplatform as aip"
"import google.cloud.aiplatform as aip\n",
"from kfp.v2.dsl import Input, Metrics, component"
]
},
{
@@ -635,6 +649,59 @@
"aip.init(project=PROJECT_ID, staging_bucket=BUCKET_NAME, location=REGION)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "d7a5cc8c7f3c"
},
"source": [
"## Create component for comparing evalution metrics to a threshold\n",
"\n",
"First, you create your own component that will take as input the evaluation metrics artifact and make a comparison to a threshold and return a yes/no decision that could be used in a subsequent dsl.Condition() to decide whether the model should proceed to the next step -- e.g., online deployment.\n",
"\n",
"The component takes the following parameters:\n",
"\n",
"- `eval_metrics`: The evaluation metrics artifact returned from `ModelEvaluation` component.\n",
"- `metric_name`: The key name for the metric entry to make the comparison to.\n",
"- `threshold`: The threshold for the metric value for a yes/no decision."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "5cf109a9bcca"
},
"outputs": [],
"source": [
"@component()\n",
"def compare(eval_metrics: Input[Metrics], metric_name: str, threshold: float) -> str:\n",
" path = eval_metrics.path\n",
" # print(\"PATH\", path)\n",
"\n",
" gs_prefix = \"gs://\"\n",
" gcsfuse_prefix = \"/gcs/\"\n",
" if path.startswith(gs_prefix):\n",
" path = path.replace(gs_prefix, gcsfuse_prefix)\n",
"\n",
" import json\n",
"\n",
" with open(path, \"r\") as f:\n",
" data = json.load(f)\n",
"\n",
" slices = data[\"slicedMetrics\"]\n",
" # print(\"# slices\", len(slices))\n",
"\n",
" metrics = slices[0][\"metrics\"][\"classification\"]\n",
" # print(\"METRIC KEYS\", metrics.keys())\n",
"\n",
" value = metrics[metric_name]\n",
" if value > threshold:\n",
" return \"true\"\n",
"\n",
" return \"false\""
]
},
{
"cell_type": "markdown",
"metadata": {
@@ -684,6 +751,8 @@
"\n",
"@kfp.dsl.pipeline(name=\"upload-evaluate-\" + TIMESTAMP)\n",
"def pipeline(\n",
" metric: str,\n",
" threshold: float,\n",
" project: str = PROJECT_ID,\n",
" model_display_name: str = MODEL_DISPLAY_NAME,\n",
" batch_prediction_display_name: str = BATCH_PREDICTION_DISPLAY_NAME,\n",
@@ -723,7 +792,7 @@
" machine_type=\"n1-standard-32\",\n",
" )\n",
"\n",
" evaluation_op(\n",
" eval_task = evaluation_op(\n",
" project=project,\n",
" root_dir=WORKING_DIR,\n",
" problem_type=\"classification\",\n",
@@ -732,6 +801,12 @@
" class_names=[\"0\", \"1\"],\n",
" predictions_format=\"jsonl\",\n",
" batch_prediction_job=batch_prediction_task.outputs[\"batchpredictionjob\"],\n",
" )\n",
"\n",
" _ = compare(\n",
" eval_metrics=eval_task.outputs[\"evaluation_metrics\"],\n",
" metric_name=metric,\n",
" threshold=threshold,\n",
" )"
]
},
@@ -787,6 +862,7 @@
" display_name=DISPLAY_NAME,\n",
" template_path=\"evaluation_demo_pipeline.json\",\n",
" pipeline_root=PIPELINE_ROOT,\n",
" parameter_values={\"metric\": \"auPrc\", \"threshold\": 0.95},\n",
" enable_caching=True,\n",
")\n",
"\n",
@@ -896,7 +972,10 @@
"artifacts = print_pipeline_output(job, \"model-batch-predict\")\n",
"print(\"\\n\\n\")\n",
"print(\"model-evaluation\")\n",
"metrics = print_pipeline_output(job, \"model-evaluation\")"
"metrics = print_pipeline_output(job, \"model-evaluation\")\n",
"print(\"\\n\\n\")\n",
"print(\"compare\")\n",
"artifacts = print_pipeline_output(job, \"compare\")"
]
},
{