Compare commits

...
Author SHA1 Message Date
Andrew Ferlitsch fb474b51ce Merge branch 'xgboost_security' of https://github.com/GoogleCloudPlatform/vertex-ai-samples into xgboost_security 2023-04-06 20:18:42 +00:00
Andrew Ferlitsch ad5cbfe6b2 fix: lint 2023-04-06 20:18:09 +00:00
Andrew FerlitschandGitHub 343e67b352 another attempt to fix the timeout 2023-04-06 09:39:37 -07:00
Andrew Ferlitsch 062bf19730 fix: CI/CD test 2023-03-31 16:52:43 +00:00
Andrew FerlitschandGitHub b3e4ce9a27 Update xgboost_data_parallel_training_on_cpu_using_dask.ipynb 2023-03-30 10:08:13 -07:00
Andrew FerlitschandGitHub c0bce32ecb reduce rounds < 24hrs 2023-03-27 12:40:56 -07:00
Andrew FerlitschandGitHub 2b576bd055 reduce rounds to get run time < 24hrs 2023-03-24 10:36:13 -07:00
Andrew FerlitschandGitHub bef12ed726 further reduce training time to get under 24hrs 2023-03-23 10:13:22 -07:00
Andrew Ferlitsch 662097b2b6 fix: reduce test time < 24hrs 2023-03-22 17:04:19 +00:00
Andrew Ferlitsch 4665910e94 fix: not deleting docker tmp image 2023-03-20 22:10:01 +00:00
gericdongandGitHub a56efdcec7 Merge pull request #1622 from GoogleCloudPlatform/pytorch_nccl
fix: missing code for nccl version
2023-03-20 19:01:55 +00:00
Andrew FerlitschandGitHub 6247fbb96f Merge pull request #1516 from sarahcdugan/patch-2
Update bqml_vertexai_model_registry.ipynb
2023-03-20 18:02:17 +00:00
Andrew Ferlitsch ad339286b0 fix: missing code for nccl version 2023-03-20 17:56:18 +00:00
sarahcdugan 94eef657ee Removed an incorrect comma 2023-03-20 16:17:05 +00:00
gericdongandGitHub b92337699a Merge pull request #1619 from GoogleCloudPlatform/sklearn_sa
fix: add missing set sa
2023-03-17 19:27:19 +00:00
gericdongandGitHub 4dcc5413cf Merge pull request #1618 from GoogleCloudPlatform/xgboost_sa_2
fix: add missing set sa
2023-03-17 18:52:15 +00:00
gericdongandGitHub 5f47ba8023 Merge pull request #1617 from GoogleCloudPlatform/xgboost_sa
fix: add missing set sa
2023-03-17 16:51:01 +00:00
Andrew Ferlitsch 698503e73d fix: add missing set sa 2023-03-17 16:14:55 +00:00
Andrew Ferlitsch e1a15c4bc9 fix: add missing set sa 2023-03-17 16:11:08 +00:00
Andrew Ferlitsch 0260d79703 fix: add missing set sa 2023-03-17 16:07:35 +00:00
Eric SchmidtandGitHub 16712e53ba Merge pull request #1614 from GoogleCloudPlatform/hier_pred
fix: correct the steps
2023-03-17 16:02:07 +00:00
gericdongandGitHub 401064a06c Merge pull request #1616 from GoogleCloudPlatform/project_id
fix: remove hw project id
2023-03-17 15:47:20 +00:00
Ivan CheungandGitHub 28d29b4691 Merge pull request #1613 from GoogleCloudPlatform/imkc--stackoverflow-redis
Added redis support to stackoverflow matching engine demo
2023-03-17 15:43:14 +00:00
ivanmkc@google.com 20902244de Ran linter 2023-03-16 23:49:17 -04:00
Andrew Ferlitsch ff5939aa8b fix: remove hw project id 2023-03-16 19:47:36 +00:00
sarahcduganandGitHub fb6527f66a Update bqml_vertexai_model_registry.ipynb 2023-03-16 12:59:55 -05:00
Andrew FerlitschandGitHub 55f8a6f78a Merge pull request #1608 from iversonic/patch-2
Fix a typo in the title of the tutorial
2023-03-16 17:50:59 +00:00
Andrew FerlitschandGitHub 397285f4bf Update custom_tabular_train_batch_pred_bq_pipeline.ipynb 2023-03-16 09:49:44 -07:00
Andrew Ferlitsch ea3167b8f7 fix: correct the steps 2023-03-15 20:56:43 +00:00
ivanmkc@google.com c3526504d8 Added redis info 2023-03-15 14:54:42 -04:00
Andrew FerlitschandGitHub 9a409b9011 Merge pull request #1610 from GoogleCloudPlatform/sklearn_2
fix: issue 1251
2023-03-15 17:33:11 +00:00
Andrew FerlitschandGitHub c9cca725c6 Merge pull request #1609 from GoogleCloudPlatform/sklearn_1
fix: issue 1251
2023-03-15 17:32:54 +00:00
Andrew FerlitschandGitHub 3ddc77293b Merge pull request #1612 from GoogleCloudPlatform/rate_limit
fix: lower rate limit
2023-03-15 17:32:21 +00:00
Andrew FerlitschandGitHub 7fa90ee179 Merge pull request #1607 from GoogleCloudPlatform/contributing
fix: one notebook rule
2023-03-15 16:11:02 +00:00
Andrew Ferlitsch b88a775d33 fix: lower rate limit 2023-03-15 15:52:39 +00:00
Andrew Ferlitsch 765d6ee296 fix: issue 1251 2023-03-14 22:10:48 +00:00
Andrew Ferlitsch aec5fbfd6f fix: issue 1251 2023-03-14 22:06:54 +00:00
Mark IversonandGitHub bcba9b5ea2 Fix a typo in the title of the tutorial 2023-03-14 14:52:25 -07:00
Andrew Ferlitsch fe42cb6ebd fix: one notebook rule 2023-03-14 21:46:43 +00:00
Andrew FerlitschandGitHub 1a3cbb4cf0 fix corrupted format 2023-02-15 07:24:17 -08:00
sarahcduganandGitHub e7e7a8e22c Update bqml_vertexai_model_registry.ipynb
XAI is now available for BQML models added to the Vertex AI Model Registry
2023-02-11 13:19:46 -06:00
14 changed files with 2108 additions and 1267 deletions
@@ -156,7 +156,7 @@ def _create_tag(filepath: str) -> str:
return tag
rate_limit = RateLimit(max_count=50, per=60, greedy=True)
rate_limit = RateLimit(max_count=25, per=60, greedy=True)
def process_and_execute_notebook(
+3 -1
View File
@@ -44,8 +44,10 @@ Finally, run this code block to check for errors. Each step will attempt to
automatically fix any issues. If the fixes can't be performed automatically,
then you will need to manually address them before submitting your PR.
Note: For official, only submit one notebook per PR.
```shell
docker run -v ${PWD}:/setup/app gcr.io/cloud-devrel-public-resources/notebook_linter:latest <your_notebooks>
docker run -v ${PWD}:/setup/app gcr.io/cloud-devrel-public-resources/notebook_linter:latest your_notebook
```
## Code Reviews
@@ -560,6 +560,30 @@
" print(\"Service Account:\", SERVICE_ACCOUNT)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "set_service_account:pipelines"
},
"source": [
"#### Set service account access for Vertex AI Pipelines\n",
"\n",
"Run the following commands to grant your service account access to read and write pipeline artifacts in the bucket that you created in the previous step -- you only need to run these once per service account."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "set_service_account:pipelines"
},
"outputs": [],
"source": [
"! gsutil iam ch serviceAccount:{SERVICE_ACCOUNT}:roles/storage.objectCreator $BUCKET_URI\n",
"\n",
"! gsutil iam ch serviceAccount:{SERVICE_ACCOUNT}:roles/storage.objectViewer $BUCKET_URI"
]
},
{
"cell_type": "markdown",
"metadata": {
@@ -557,6 +557,30 @@
" print(\"Service Account:\", SERVICE_ACCOUNT)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "set_service_account:pipelines"
},
"source": [
"#### Set service account access for Vertex AI Pipelines\n",
"\n",
"Run the following commands to grant your service account access to read and write pipeline artifacts in the bucket that you created in the previous step -- you only need to run these once per service account."
]
},
{
"cell_type": "code",
"execution_count": 13,
"metadata": {
"id": "set_service_account:pipelines"
},
"outputs": [],
"source": [
"! gsutil iam ch serviceAccount:{SERVICE_ACCOUNT}:roles/storage.objectCreator $BUCKET_URI\n",
"\n",
"! gsutil iam ch serviceAccount:{SERVICE_ACCOUNT}:roles/storage.objectViewer $BUCKET_URI"
]
},
{
"cell_type": "markdown",
"metadata": {
@@ -557,6 +557,30 @@
" print(\"Service Account:\", SERVICE_ACCOUNT)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "set_service_account:pipelines"
},
"source": [
"#### Set service account access for Vertex AI Pipelines\n",
"\n",
"Run the following commands to grant your service account access to read and write pipeline artifacts in the bucket that you created in the previous step -- you only need to run these once per service account."
]
},
{
"cell_type": "code",
"execution_count": 13,
"metadata": {
"id": "set_service_account:pipelines"
},
"outputs": [],
"source": [
"! gsutil iam ch serviceAccount:{SERVICE_ACCOUNT}:roles/storage.objectCreator $BUCKET_URI\n",
"\n",
"! gsutil iam ch serviceAccount:{SERVICE_ACCOUNT}:roles/storage.objectViewer $BUCKET_URI"
]
},
{
"cell_type": "markdown",
"metadata": {
@@ -87,9 +87,7 @@
"- Create a Vertex AI `TimeSeriesDataset` resource.\n",
"- Train the model.\n",
"- View the model evaluation.\n",
"- Deploy the `Model` resource to a serving `Endpoint` resource.\n",
"- Make a prediction.\n",
"- Undeploy the `Model`."
"- Make a batch prediction."
]
},
{
@@ -150,6 +150,28 @@
" tensorflow-hub"
]
},
{
"attachments": {},
"cell_type": "markdown",
"metadata": {
"id": "54ac7ebac10b"
},
"source": [
"Install the latest version of Redis for low-latency data retrieval"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "e3dd53b3c06c"
},
"outputs": [],
"source": [
"# Install the redis package\n",
"! pip install --upgrade redis"
]
},
{
"cell_type": "markdown",
"metadata": {
@@ -1015,6 +1037,7 @@
]
},
{
"attachments": {},
"cell_type": "markdown",
"metadata": {
"id": "6LCGvBNvBd8D"
@@ -1022,7 +1045,9 @@
"source": [
"## Create Online Queries\n",
"\n",
"After you built your indexes, you may query against the deployed index to find nearest neighbors."
"After you built your indexes, you may query against the deployed index to find nearest neighbors.\n",
"\n",
"Note: For the DOT_PRODUCT_DISTANCE distance type, the \"distance\" property returned with each MatchNeighbor actually refers to the similarity."
]
},
{
@@ -1087,6 +1112,100 @@
" )"
]
},
{
"attachments": {},
"cell_type": "markdown",
"metadata": {
"id": "05514825ba7d"
},
"source": [
"## Storing and retrieving titles from a Redis data store\n",
"When you productionize this code into a service, you will need to convert the nearest nearest id's returned from Vertex AI Matching Engine into data usable by downstream services.\n",
"\n",
"In this case, you'll need to convert the id's to titles.\n",
"\n",
"You can use Google Cloud's Memorystore to deploy a managed Redis instance to save the id-title key-value pairs.\n",
"\n",
"See more information on [Memorystore](https://cloud.google.com/memorystore/docs/redis/create-manage-instances?hl=en)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "5d2b240f0d52"
},
"outputs": [],
"source": [
"REDIS_INSTANCE_NAME = \"stackoverflow-questions-unique\"\n",
"\n",
"# Create a Redis instance\n",
"! gcloud redis instances create '{REDIS_INSTANCE_NAME}' --size=5 --region={REGION} --network={VPC_NETWORK_FULL} --connect-mode=private-service-access"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "371ccc0d2eb2"
},
"outputs": [],
"source": [
"# Get host and port info\n",
"REDIS_HOST = ! gcloud redis instances list --filter=\"INSTANCE_NAME:'{REDIS_INSTANCE_NAME}'\" --region {REGION} --format='value(HOST)'\n",
"REDIS_PORT = ! gcloud redis instances list --filter=\"INSTANCE_NAME:'{REDIS_INSTANCE_NAME}'\" --region {REGION} --format='value(PORT)'\n",
"\n",
"if isinstance(REDIS_HOST, list):\n",
" REDIS_HOST = REDIS_HOST[0]\n",
"\n",
"if isinstance(REDIS_PORT, list):\n",
" REDIS_PORT = REDIS_PORT[0]\n",
"\n",
"print(f\"REDIS_HOST = {REDIS_HOST}\")\n",
"print(f\"REDIS_PORT = {REDIS_PORT}\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "73796089386a"
},
"outputs": [],
"source": [
"# Connect to the instance\n",
"import redis\n",
"\n",
"redis_client = redis.StrictRedis(host=REDIS_HOST, port=REDIS_PORT)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "f000f5432d13"
},
"outputs": [],
"source": [
"# Convert the id -> title relationship into a dict and write to redis\n",
"redis_client.mset({str(id): str(title) for id, title in zip(df.id, df.title)})"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "b1f8b396aeb1"
},
"outputs": [],
"source": [
"# Verify that redis can retrieve the correct information\n",
"[\n",
" f\"Actual = {title}, Retrieved = {redis_client.get(str(id))}\"\n",
" for id, title in list(zip(df.id, df.title))[:10]\n",
"]"
]
},
{
"cell_type": "markdown",
"metadata": {
@@ -1123,6 +1242,18 @@
"# Delete indexes\n",
"tree_ah_index.delete()"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "d2fcf9468031"
},
"outputs": [],
"source": [
"# Delete redis instance\n",
"! gcloud redis instances delete '{REDIS_INSTANCE_NAME}' --region {REGION} --quiet"
]
}
],
"metadata": {
@@ -719,7 +719,7 @@
"outputs": [],
"source": [
"@component(\n",
" packages_to_install=[\"sklearn\", \"pandas\", \"joblib\"],\n",
" packages_to_install=[\"scikit-learn\", \"pandas\", \"joblib\"],\n",
" base_image=\"python:3.9\",\n",
" output_component_file=\"beans_model_component.yaml\",\n",
")\n",
File diff suppressed because it is too large Load Diff
@@ -29,7 +29,7 @@
"id": "978ab06a7e3f"
},
"source": [
"# Vertex AI Pipelines: Training and batch prediction with BigQuery source and destinantion for a custom tabular classification model \n",
"# Vertex AI Pipelines: Training and batch prediction with BigQuery source and destination for a custom tabular classification model \n",
"\n",
"<table align=\"left\">\n",
"\n",
@@ -225,7 +225,8 @@
" pandas \\\n",
" pyarrow \\\n",
" kfp \\\n",
" google-cloud-pipeline-components {USER_FLAG} -q"
" google-cloud-pipeline-components {USER_FLAG} -q \n",
"! pip3 install db-dtypes {USER_FLAG} -q"
]
},
{
@@ -713,7 +713,7 @@
"outputs": [],
"source": [
"@component(\n",
" packages_to_install=[\"sklearn\"],\n",
" packages_to_install=[\"scikit-learn\"],\n",
" base_image=\"python:3.9\",\n",
" output_component_file=\"wine_classification_component.yaml\",\n",
")\n",
@@ -758,7 +758,7 @@
},
"outputs": [],
"source": [
"@component(packages_to_install=[\"sklearn\"], base_image=\"python:3.9\")\n",
"@component(packages_to_install=[\"scikit-learn\"], base_image=\"python:3.9\")\n",
"def iris_sgdclassifier(\n",
" test_samples_fraction: float,\n",
" metricsc: Output[ClassificationMetrics],\n",
@@ -805,7 +805,7 @@
"outputs": [],
"source": [
"@component(\n",
" packages_to_install=[\"sklearn\"],\n",
" packages_to_install=[\"scikit-learn\"],\n",
" base_image=\"python:3.9\",\n",
")\n",
"def iris_logregression(\n",
@@ -207,7 +207,7 @@
},
"outputs": [],
"source": [
"PROJECT_ID = \"andy-1234-221921\" # @param {type:\"string\"}\n",
"PROJECT_ID = \"[your-project-id]\" # @param {type:\"string\"}\n",
"\n",
"# Set the project id\n",
"! gcloud config set project {PROJECT_ID}"
@@ -67,7 +67,6 @@
]
},
{
"attachments": {},
"cell_type": "markdown",
"metadata": {
"id": "AksIKBzZ-nre"
@@ -552,6 +551,8 @@
" 'nthread': 1\n",
"}\n",
"\n",
"ROUNDS = 2\n",
"\n",
"\n",
"def square(x):\n",
" return x ** 2\n",
@@ -615,7 +616,7 @@
" wait(y)\n",
" dtrain = DaskDMatrix(client, X, y)\n",
"\n",
" output = xgb.dask.train(client, XGB_PARAMS, dtrain, num_boost_round=100, evals=[(dtrain, 'train')])\n",
" output = xgb.dask.train(client, XGB_PARAMS, dtrain, num_boost_round=ROUNDS, evals=[(dtrain, 'train')])\n",
" print(\"Output: {}\".format(output), flush=True)\n",
" print(\"Saving file to: {}\".format(MODEL_FILE), flush=True)\n",
" output['booster'].save_model(MODEL_FILE)\n",
@@ -902,14 +903,15 @@
" container_uri=TRAIN_IMAGE,\n",
")\n",
"\n",
"custom_container_training_job.run(\n",
" base_output_dir=gcs_output_uri_prefix,\n",
" replica_count=replica_count,\n",
" machine_type=machine_type,\n",
" enable_dashboard_access=True,\n",
" enable_web_access=True,\n",
" sync=False,\n",
")"
"if not os.getenv(\"IS_TESTING\"):\n",
" custom_container_training_job.run(\n",
" base_output_dir=gcs_output_uri_prefix,\n",
" replica_count=replica_count,\n",
" machine_type=machine_type,\n",
" enable_dashboard_access=True,\n",
" enable_web_access=True,\n",
" sync=False,\n",
" )"
]
},
{
@@ -942,8 +944,9 @@
},
"outputs": [],
"source": [
"print(f\"Custom Training Job Name: {custom_container_training_job.resource_name}\")\n",
"print(f\"GCS Output URI Prefix: {gcs_output_uri_prefix}\")"
"if not os.getenv(\"IS_TESTING\"):\n",
" print(f\"Custom Training Job Name: {custom_container_training_job.resource_name}\")\n",
" print(f\"GCS Output URI Prefix: {gcs_output_uri_prefix}\")"
]
},
{
@@ -963,9 +966,10 @@
},
"outputs": [],
"source": [
"print(\n",
" f\"Custom Training Job URI: {custom_container_training_job._custom_job_console_uri()}\"\n",
")"
"if not os.getenv(\"IS_TESTING\"):\n",
" print(\n",
" f\"Custom Training Job URI: {custom_container_training_job._custom_job_console_uri()}\"\n",
" )"
]
},
{
@@ -985,7 +989,10 @@
},
"outputs": [],
"source": [
"print(f\"Web Access and Dashboard URIs: {custom_container_training_job.web_access_uris}\")"
"if not os.getenv(\"IS_TESTING\"):\n",
" print(\n",
" f\"Web Access and Dashboard URIs: {custom_container_training_job.web_access_uris}\"\n",
" )"
]
},
{
@@ -1066,7 +1073,8 @@
},
"outputs": [],
"source": [
"! gcloud ai custom-jobs create --region=us-central1 --config=config.yaml --display-name={display_name}"
"if not os.getenv(\"IS_TESTING\"):\n",
" ! gcloud ai custom-jobs create --region=us-central1 --config=config.yaml --display-name={display_name}"
]
},
{
@@ -1187,10 +1195,7 @@
"To clean up all Google Cloud resources used in this project, you can [delete the Google Cloud\n",
"project](https://cloud.google.com/resource-manager/docs/creating-managing-projects#shutting_down_projects) you used for the tutorial.\n",
"\n",
"Otherwise, you can delete the individual resources you created in this tutorial:\n",
"\n",
"- Cloud Storage Bucket\n",
"- Cloud Vertex Training Job"
"Otherwise, you can delete the individual resources you created in this tutorial."
]
},
{
@@ -1201,9 +1206,6 @@
},
"outputs": [],
"source": [
"import logging\n",
"import traceback\n",
"\n",
"# Set this to true only if you'd like to delete your bucket\n",
"delete_bucket = False\n",
"\n",
@@ -1215,8 +1217,10 @@
"try:\n",
" custom_container_training_job.delete()\n",
"except Exception as e:\n",
" logging.error(traceback.format_exc())\n",
" print(e)"
" print(e)\n",
"\n",
"# delete the docker image\n",
"! gcloud artifacts repositories delete --location {REGION} {TRAIN_IMAGE}"
]
}
],