Compare commits

..
Author SHA1 Message Date
ivanmkc@google.com 5679e46a12 Addressed TW comments 2023-03-07 12:01:41 -05:00
ivanmkc@google.com d54b3845fe Fixed notebooks/official/ml_metadata/sdk-metric-parameter-tracking-for-locally-trained-models.ipynb 2023-03-07 10:45:42 -05:00
ivanmkc@google.com be20a635e4 Fixed missing dependency 2023-03-07 08:42:35 -05:00
ivanmkc@google.com b42b6c6fb8 Added missing cells 2023-03-06 23:34:27 -05:00
ivanmkc@google.com 73fbc762fe Added tqdm to requirements.txt 2023-03-06 20:00:53 -05:00
ivanmkc@google.com 994a86d07a Fixed bugs 2023-03-06 16:53:46 -05:00
ivanmkc@google.com 71c1eca210 Fixed predictions 2023-03-06 16:16:39 -05:00
ivanmkc@google.com 1c431bdb85 Updated links 2023-03-06 14:41:50 -05:00
ivanmkc@google.com 096a5d069e Linted 2023-03-06 14:38:29 -05:00
ivanmkc@google.com 3ade1ab265 Added stackoverflow embeddings notebook 2023-03-06 10:45:06 -05:00
Andrew FerlitschandGitHub 8593308244 Merge pull request #1569 from ninataneja/final-doc-change
Update dashboard instructions
2023-03-03 22:29:28 +00:00
Nina Taneja 806801b4ea Fix print error 2023-03-03 21:15:33 +00:00
Nina Taneja 3acd72ec81 Fix lint error 2023-03-03 20:49:42 +00:00
Nina Taneja 12347f3ce2 Add error handling for delete job 2023-03-03 20:45:29 +00:00
Nina Taneja f9c4e32088 Add sleep for async job 2023-03-03 20:11:00 +00:00
Nina Taneja 78da2206b0 Update dashboard instructions 2023-03-03 18:50:21 +00:00
gericdongandGitHub 46fa993732 Merge pull request #1559 from GoogleCloudPlatform/ml_ops_registry
feat: add notebook for model versioning
2023-03-03 18:43:05 +00:00
Andrew FerlitschandGitHub df0c0d209b Merge pull request #1568 from kthytang/fs-integration
feat: notebook for prediction feature store integration
2023-03-03 18:32:05 +00:00
kthytang 2e8d6239df chore: run python3.9 -m tensorflow_docs.tools.nbfmt --remove_outputs "$notebook" 2023-03-03 10:02:46 -08:00
Andrew Ferlitsch 1841201fee fix: dep issue 2023-03-03 17:20:44 +00:00
kthytang f10857f299 chore: address comments 2023-03-03 07:17:50 -08:00
kthytang e3bcff62fc chore: fix lint 2023-03-02 14:18:47 -08:00
kthytang ce3a439d06 feat: notebook for prediction feature store integration 2023-03-02 14:04:38 -08:00
Andrew FerlitschandGitHub b47d4b46f3 Merge pull request #1567 from GoogleCloudPlatform/reznitskii-patch-12
Fixed typo
2023-03-02 21:40:56 +00:00
Andrew FerlitschandGitHub 1cc87860c4 Merge pull request #1566 from GoogleCloudPlatform/reznitskii-patch-11
Fixed typo and link
2023-03-02 21:40:06 +00:00
reznitskiiandGitHub 5d6aa4479d Update automl_video_classification_model_evaluation.ipynb 2023-03-02 15:11:07 -05:00
reznitskiiandGitHub 5a25b06f2d Update UJ14 Vertex SDK AutoML Video Classification.ipynb 2023-03-02 15:09:45 -05:00
Andrew Ferlitsch bb055ed061 fix: cleanup 2023-03-02 18:26:45 +00:00
Andrew FerlitschandGitHub 974610a555 Merge pull request #1562 from GoogleCloudPlatform/reznitskii-patch-8
Updated link
2023-03-02 08:25:51 +00:00
Andrew FerlitschandGitHub d78574e640 Merge pull request #1565 from ninataneja/dask-sdk
Add SDK support for Dask dashboard to Training
2023-03-02 08:25:22 +00:00
Nina Taneja ff367ae9f5 Fixed formatting problem 2023-03-02 01:34:46 +00:00
Nina Taneja 40b5e74645 Addressed formatting and wording changes 2023-03-02 01:26:52 +00:00
Nina Taneja dee509e8d7 Add SDK support for Dask dashboard to Training 2023-03-01 23:39:06 +00:00
Andrew Ferlitsch fd5921fa32 fix: invalid alias 2023-03-01 22:55:39 +00:00
reznitskiiandGitHub c74a44a4a8 Update sdk_automl_tabular_classification_online_explain.ipynb 2023-03-01 17:22:30 -05:00
Andrew FerlitschandGitHub 367c985642 Merge pull request #1560 from GoogleCloudPlatform/autoindex_tensorboard
fix: tensorboard branding
2023-03-01 22:03:24 +00:00
Andrew Ferlitsch f7a970e15b fix: TIMESTAMP 2023-03-01 21:53:51 +00:00
Andrew Ferlitsch 9a914a5af4 fix: tensorboard branding 2023-03-01 21:46:57 +00:00
Andrew Ferlitsch e5315be85a feat: add notebook for model versioning 2023-03-01 20:54:17 +00:00
Andrew FerlitschandGitHub 6daf663a69 Merge pull request #1557 from GoogleCloudPlatform/custom_eval
feat: notebook for custom evaluations
2023-03-01 20:15:56 +00:00
Andrew Ferlitsch 45a65f1db8 fix: install issue 2023-03-01 20:05:54 +00:00
Andrew Ferlitsch 240c291728 feat: add eval on versioned model 2023-03-01 19:21:21 +00:00
gericdongandGitHub 4aec14c576 Merge pull request #1558 from GoogleCloudPlatform/andrewferlitsch-patch-4
tmp file added by mistake
2023-03-01 18:34:45 +00:00
Andrew FerlitschandGitHub f7d4d3a9c3 tmp file added by mistake 2023-03-01 10:08:20 -08:00
Andrew Ferlitsch 3c64a8aa58 fix: review nits 2023-03-01 18:04:34 +00:00
Andrew Ferlitsch 5f692ea299 fix: missing installs 2023-03-01 16:06:54 +00:00
Andrew FerlitschandGitHub ab1e97ac18 Merge pull request #1553 from GoogleCloudPlatform/autoindex_max_3
fix: 5 branding bugs
2023-03-01 16:04:23 +00:00
Andrew FerlitschandGitHub 78630ae2c0 Merge pull request #1552 from GoogleCloudPlatform/reznitskii-patch-3
Fixed typo
2023-03-01 16:03:51 +00:00
Andrew Ferlitsch 0f48628782 feat: notebook for custom evaluations 2023-03-01 00:52:20 +00:00
Andrew Ferlitsch 93e9fbac92 fix: 5 branding bugs 2023-02-28 20:42:03 +00:00
Andrew FerlitschandGitHub 231b2ef02b Merge pull request #1556 from GoogleCloudPlatform/reznitskii-patch-6
Fixed typo
2023-02-28 20:40:04 +00:00
Andrew FerlitschandGitHub 7132c12831 Merge pull request #1555 from GoogleCloudPlatform/reznitskii-patch-5
Fixed typos
2023-02-28 20:31:30 +00:00
reznitskiiandGitHub 2b547e8279 Update forecasting-retail-demand.ipynb 2023-02-28 15:24:25 -05:00
reznitskiiandGitHub 8b84524244 Update ai-explanations-tabnet-algorithm.ipynb 2023-02-28 15:21:53 -05:00
Andrew Ferlitsch 60fe1bd2c6 fix: 5 branding bugs 2023-02-28 20:16:57 +00:00
reznitskiiandGitHub 0edea80ffa Update custom_tabular_regression_model_evaluation.ipynb 2023-02-28 15:02:47 -05:00
Yvonne LiandGitHub 493e50a999 Merge pull request #1550 from GoogleCloudPlatform/autoindex_max_2
fix: extra period in link
2023-02-28 19:50:08 +00:00
Andrew Ferlitsch 8440e7f164 fix: extra period in link 2023-02-28 19:44:33 +00:00
gericdongandGitHub 906aa91fe7 Merge pull request #1549 from GoogleCloudPlatform/mv_pytorch
fix: reorg
2023-02-28 18:41:15 +00:00
gericdongandGitHub 34e282172f Merge pull request #1548 from GoogleCloudPlatform/rm_pytorch_folder
fix: reorg
2023-02-28 18:29:13 +00:00
Andrew Ferlitsch bddf642b58 fix: reorg 2023-02-28 18:24:55 +00:00
Andrew Ferlitsch 25b89c497f fix: reorg 2023-02-28 18:19:55 +00:00
Andrew FerlitschandGitHub 23a51dbcaa Merge pull request #1526 from GoogleCloudPlatform/doc_tag_1
update tag/linkback #1 AutoML Video
2023-02-28 18:15:56 +00:00
Andrew FerlitschandGitHub 26761198d7 Merge pull request #1547 from GoogleCloudPlatform/autoindex_max_1
fix: web index tune
2023-02-28 17:39:34 +00:00
Andrew Ferlitsch 5b4e88e791 fix: web index tune 2023-02-28 17:29:45 +00:00
Andrew Ferlitsch 83cc796687 fix: workaround for running > 24hrs 2023-02-28 17:05:34 +00:00
Andrew FerlitschandGitHub 3e29cc62c7 Merge pull request #1536 from GoogleCloudPlatform/doc_tag_11
update tag/linkback #11 AutoML Video
2023-02-28 17:02:09 +00:00
Andrew FerlitschandGitHub 2b1bb77f4b Merge pull request #1535 from GoogleCloudPlatform/doc_tag_10
update tag/linkback #10 AutoMLVideo
2023-02-28 16:55:55 +00:00
Andrew Ferlitsch 7144b87f01 fix: workaround for running > 24hrs 2023-02-28 16:50:51 +00:00
Andrew Ferlitsch ebf4d6d8be fix: workaround for running > 24hrs 2023-02-28 16:47:14 +00:00
Andrew FerlitschandGitHub 2ccd913e2e Merge pull request #1528 from GoogleCloudPlatform/doc_tag_3
update tag/linkback #3 AutoML Video
2023-02-28 16:45:02 +00:00
Andrew Ferlitsch d777dd8625 fix: workaround for running > 24hrs 2023-02-28 16:32:16 +00:00
Andrew FerlitschandGitHub d912d3d4d8 Merge pull request #1545 from GoogleCloudPlatform/reznitskii-patch-1
Fixed typo
2023-02-28 02:29:29 +00:00
reznitskiiandGitHub bafe623590 Update automl_tabular_regression_model_evaluation.ipynb 2023-02-27 17:34:34 -05:00
gericdongandGitHub 152077a823 Merge pull request #1543 from GoogleCloudPlatform/guidelines
feat: add authoring guidelines
2023-02-27 20:43:43 +00:00
Andrew Ferlitsch 81c94f9711 feat: add authoring guidelines 2023-02-27 19:53:10 +00:00
Andrew FerlitschandGitHub 897e8e3e47 Merge pull request #1541 from GoogleCloudPlatform/template_linkback
fix: add tag/linkback
2023-02-27 18:06:24 +00:00
Andrew Ferlitsch ddc125da69 fix: smaller dataset 2023-02-27 16:16:13 +00:00
Eric SchmidtandGitHub abb66ade41 Merge pull request #1542 from gericdong/b270683209
fix: bad links in notebook
2023-02-24 17:06:07 +00:00
gericdong 2915824641 Lint 2023-02-24 09:03:19 -05:00
gericdong 44447637bd fixed bad links 2023-02-24 08:59:24 -05:00
Andrew FerlitschandGitHub 6a8345446d Merge pull request #1504 from reznitskii/b267661933-2
Replaced Vertex AI Training linkbacks with Custom training
2023-02-24 02:13:42 +00:00
Andrew Ferlitsch 27a7b8b1da fix: add tag/linkback 2023-02-24 00:22:18 +00:00
Andrew FerlitschandGitHub 04471e7104 Merge pull request #1540 from GoogleCloudPlatform/issue_1522
fix: link
2023-02-23 17:43:06 +00:00
Andrew FerlitschandGitHub e78ea4e805 Merge pull request #1538 from wintwoo/dataproc
Specify Dataproc Serverless Runtime version to use for batch workloads.
2023-02-23 04:14:06 +00:00
Andrew Ferlitsch 7487a14783 fix: link 2023-02-22 23:07:37 +00:00
Andrew FerlitschandGitHub 757d0c5087 change copyright back to 2022. Policy is year is the year first authored 2023-02-22 15:00:49 -08:00
Andrew FerlitschandGitHub 60cb93a9a9 Delete pytorch-text-sentiment-classification-custom-train-deploy.ipynb 2023-02-22 14:43:12 -08:00
Andrew FerlitschandGitHub 0fb72773f4 Merge pull request #1533 from GoogleCloudPlatform/doc_tag_8
update tag/linkback #8
2023-02-22 22:21:16 +00:00
Andrew FerlitschandGitHub 5d2fc74f78 Merge pull request #1532 from GoogleCloudPlatform/doc_tag_7
update tag/linkback #7
2023-02-22 22:20:58 +00:00
Andrew FerlitschandGitHub 1cc26481b2 Merge pull request #1531 from GoogleCloudPlatform/doc_tag_6
update tag/linkback #6
2023-02-22 22:20:50 +00:00
Andrew FerlitschandGitHub 85b0346eab Merge pull request #1530 from GoogleCloudPlatform/doc_tag_5
update tag/linkback #5
2023-02-22 22:20:24 +00:00
Andrew FerlitschandGitHub 910cbacb54 Merge pull request #1529 from GoogleCloudPlatform/doc_tag_4
update tag/linkback #4
2023-02-22 22:19:44 +00:00
Andrew FerlitschandGitHub c7206e3cc8 Merge pull request #1527 from GoogleCloudPlatform/doc_tag_2
update tag/linkback #2
2023-02-22 21:46:40 +00:00
Win Woo 12d721d33e Specify Dataproc runtime versions to use for batch jobs 2023-02-22 02:46:00 +00:00
Andrew FerlitschandGitHub 16df2216d1 Merge pull request #1525 from GoogleCloudPlatform/triton_ensenble_2
feat: triton ensemble
2023-02-21 21:32:16 +00:00
Andrew Ferlitsch fb26b7213c fix: links 2023-02-21 21:30:12 +00:00
Andrew Ferlitsch d90f52e122 fix: links 2023-02-21 19:47:23 +00:00
Andrew Ferlitsch 652f34b814 update tag/linkback 2023-02-21 18:37:09 +00:00
Andrew Ferlitsch c84f8738f8 update tag/linkback 2023-02-21 18:16:13 +00:00
Andrew Ferlitsch 9d72bbf237 update tag/linkback 2023-02-21 18:02:11 +00:00
Andrew Ferlitsch 85df50d06f update tag/linkback 2023-02-21 17:39:51 +00:00
Andrew Ferlitsch 9ab6cddade update tag/linkback 2023-02-21 17:33:56 +00:00
Andrew Ferlitsch 02ffffaeb7 update tag/linkback 2023-02-21 17:30:03 +00:00
Andrew Ferlitsch dfc5635d34 update tag/linkback 2023-02-21 17:25:40 +00:00
Andrew Ferlitsch 90b686cf69 update tag/linkback 2023-02-21 17:02:33 +00:00
Andrew Ferlitsch ad63153778 update tag/linkback 2023-02-21 16:58:56 +00:00
Andrew Ferlitsch 7858f8644e update tag/linkback 2023-02-21 16:55:10 +00:00
Andrew Ferlitsch 3df4098364 update tag/linkback 2023-02-21 16:46:03 +00:00
Andrew Ferlitsch ae3c1877e4 update tag/linkback 2023-02-21 16:41:17 +00:00
Andrew FerlitschandGitHub 2798ab1e02 set timeout 2023-02-21 08:33:40 -08:00
Andrew FerlitschandGitHub 7bd2c567a1 Merge pull request #1490 from GoogleCloudPlatform/imkc--matching-engine-embedding-tweak
Fixed typo in sdk_matching_engine_for_indexing.ipynb
2023-02-18 00:48:17 +00:00
Andrew Ferlitsch 56503cce6a feat: triton ensemble 2023-02-17 22:10:20 +00:00
Andrew FerlitschandGitHub a5a9f53a32 Merge pull request #1515 from junyanxu/add_experimental_info_to_automl_image_montioring_notebook
Add experimental information to automl image classifcation monitoring…
2023-02-17 15:52:07 +00:00
Junyan Xu 469e8436e9 Format the automl image online for model monitoring 2023-02-16 22:37:50 +00:00
Junyan Xu e9f9a5c29b Merge branch 'add_experimental_info_to_automl_image_montioring_notebook' of https://github.com/junyanxu/vertex-ai-samples into add_experimental_info_to_automl_image_montioring_notebook 2023-02-16 17:19:49 +00:00
Junyan Xu 102400a67f update online pip install package 2023-02-16 17:19:00 +00:00
Andrew FerlitschandGitHub 0ae53f47d4 Update get_started_with_model_monitoring_automl_image_batch.ipynb 2023-02-16 08:27:40 -08:00
Eric SchmidtandGitHub ff3e0b4784 Merge pull request #1519 from GoogleCloudPlatform/autoindex_official_6
tune: web index
2023-02-15 21:03:10 +00:00
Andrew Ferlitsch 0f4546c18c tune: web index 2023-02-15 20:56:47 +00:00
Andrew FerlitschandGitHub ee381e15f5 Update get_started_with_model_monitoring_automl_image_batch.ipynb 2023-02-15 12:13:36 -08:00
Andrew FerlitschandGitHub 4ddeeb290d Merge pull request #1295 from Ark-kun/Train_tabular_models
Train tabular models with many frameworks and import to Vertex AI using Pipelines
2023-02-15 19:54:11 +00:00
Andrew FerlitschandGitHub 7e99b07440 Merge branch 'main' into Train_tabular_models 2023-02-15 11:53:17 -08:00
Andrew FerlitschandGitHub bc1ecd6260 Merge pull request #1334 from renovate-bot/renovate/isort-5.x
chore(deps): update dependency isort to v5.12.0
2023-02-15 19:49:21 +00:00
Andrew FerlitschandGitHub f662ffb3e1 Merge pull request #1328 from renovate-bot/renovate/black-22.x
chore(deps): update dependency black to v22.12.0
2023-02-15 19:48:46 +00:00
Andrew FerlitschandGitHub c40c0e5247 Merge pull request #1227 from sudarshan-SpringML/auto_tab_on_vertex_pipeline
Update the file automl_tabular_on_vertex_pipelines
2023-02-15 19:32:19 +00:00
Eric SchmidtandGitHub bd3a0f5af0 Merge pull request #1518 from GoogleCloudPlatform/issue_265061259
fix: issue
2023-02-15 17:46:37 +00:00
Andrew Ferlitsch 99f318333b fix: issue 2023-02-15 17:10:28 +00:00
Andrew FerlitschandGitHub 11e7ba47a4 Merge pull request #1470 from GoogleCloudPlatform/dependabot/pip/community-content/pytorch_image_classification_single_gpu_with_vertex_sdk_and_torchserve/trainer/torch-1.13.1
Build(deps): Bump torch from 1.8.1 to 1.13.1 in /community-content/pytorch_image_classification_single_gpu_with_vertex_sdk_and_torchserve/trainer
2023-02-15 16:54:21 +00:00
Andrew FerlitschandGitHub 399427c2de Merge pull request #1495 from TheMichaelHu/mh-prophet
Reduce cost of running prophet notebook
2023-02-15 15:43:17 +00:00
gericdongandGitHub 0f116ab253 Merge pull request #1517 from GoogleCloudPlatform/stable-diffusion-fixes
chore: revisions to Stable Diffusion and TorchServe nb
2023-02-15 13:33:54 +00:00
Michael Hu 304000d719 use n1-standard-2s 2023-02-14 22:18:13 -05:00
Eric Schmidt 36df462615 chore: revisions to Stable Diffusion and TorchServe nb 2023-02-15 02:32:02 +00:00
Andrew FerlitschandGitHub 909f771bfd Merge pull request #1514 from abcdefgs0324/pytorch_ga
Update wording for pre-built pytorch images on Vertex Prediction.
2023-02-14 17:07:03 +00:00
Eric SchmidtandGitHub e9bef3542d Merge pull request #1510 from GoogleCloudPlatform/stable-diffusion-try2
feat: adds stable diffusion notebook with PyTorch serving
2023-02-13 20:33:24 +00:00
Eric Schmidt 037a3b041a linting 2023-02-13 20:31:22 +00:00
Eric Schmidt f55e6c60cd per reviewer 2023-02-13 18:11:09 +00:00
Eric Schmidt 2cd815c640 light edit 2023-02-10 23:00:11 +00:00
Eric Schmidt b1c0cc9a9d linter 2023-02-10 22:44:33 +00:00
Eric Schmidt af63d1c0e2 Revised notebook to use existing model 2023-02-10 22:40:38 +00:00
Junyan Xu c51a82febb format notebook 2023-02-10 19:15:15 +00:00
Junyan Xu 8e6e4729aa Add experimental information to automl image classifcation monitoring notebook 2023-02-10 18:48:09 +00:00
Eric Schmidt 0e97a83836 revisions 2023-02-10 17:06:33 +00:00
Eric Schmidt f6645e0125 moved notebook 2023-02-10 17:01:13 +00:00
Chun-Hsiang Wang 52385a6071 samples: Updated wording and removed preview email. 2023-02-10 00:43:43 +00:00
Chun-Hsiang WangandGitHub 2365d733c4 Merge branch 'GoogleCloudPlatform:main' into pytorch_ga 2023-02-09 12:55:15 -08:00
Eric Schmidt 71e6423066 iter 2023-02-09 18:06:24 +00:00
Eric Schmidt 139ed95ffc iter 2023-02-09 18:02:07 +00:00
Eric Schmidt cafd192417 deleted notebooks from old location 2023-02-09 17:32:32 +00:00
Eric Schmidt 0d7ec7cd60 iter 2023-02-09 17:21:45 +00:00
Eric Schmidt 7fd934045a moved location of notebook 2023-02-09 17:21:02 +00:00
Max Reznitskii 4e82877269 Reverting changes to failing notebooks 2023-02-09 16:42:44 +00:00
Andrew FerlitschandGitHub 6ef111144d fix: lost updates (#1513) 2023-02-08 18:35:05 -05:00
Eric Schmidt e9b8aa02e9 linter 2023-02-07 14:27:36 -08:00
Eric Schmidt ac1af33c8c feat: adds stable diffusion notebook with PyTorch serving 2023-02-07 21:50:22 +00:00
Andrew FerlitschandGitHub 5bc18b01e4 feat: MM for automl image (#1483)
* feat: MM for automl image

* feat: MM for automl image

* fix: missing import for testing

* fix: testing

* fix: test timing issues

* debug: timing

* test: fix timing issue

* tune: updates from TW for web index

* fix: code review
2023-02-07 14:21:16 -05:00
Andrew FerlitschandGitHub e98b9d6eb4 fix: bad link (#1507) 2023-02-07 10:52:10 -08:00
Andrew FerlitschandGitHub 5e6b8bf597 fix: bad link (#1508) 2023-02-07 10:51:28 -08:00
Andrew FerlitschandGitHub 5b39e7d995 fix: bad link (#1506) 2023-02-07 10:51:08 -08:00
Max Reznitskii c4f08589a9 Reverting changes to notebooks that fail tests 2023-02-03 19:21:48 +00:00
04c6ff4ec7 Fixed AutoML Tabular linkbacks. Linkbacks now refer to specific tabular data tasks. (#1503)
Co-authored-by: Max Reznitskii <reznitskii@google.com>
2023-02-03 10:35:37 -08:00
Max Reznitskii bcb0b19dc4 Replaced Vertex AI Training linkbacks with Custom training 2023-02-02 23:04:08 +00:00
Ivan CheungGitHubivanmkc@google.com <ivanmkc@google.com>
5586fd7c4d Fixed comment about GCS (#1500)
Co-authored-by: ivanmkc@google.com <ivanmkc@google.com>
2023-02-02 14:30:02 -08:00
gericdongandGitHub 30e747b966 correct/remove invalid github usernames (#1502) 2023-02-02 13:24:44 -08:00
Andrew FerlitschandGitHub 080e2b5bb5 fix: missed updates (#1499) 2023-02-01 15:52:22 -05:00
Michael Hu e908774b5b doc updates 2023-01-31 13:23:25 -05:00
Ivan CheungGitHubivanmkc@google.com <ivanmkc@google.com>
60e4416a7e cleanup: remove 3 deprecated notebooks (#1497)
Co-authored-by: ivanmkc@google.com <ivanmkc@google.com>
2023-01-31 00:07:38 -08:00
gericdongandGitHub b207b270b4 feat: enable TensorBoard profiler for custom training with prebuilt container (#1494)
* feat: enable TensorBoard profiler for custom training with prebuilt container

* Fixed package install error

* addressed review comments
2023-01-30 09:31:51 -08:00
Andrew FerlitschandGitHub 497e93aba1 fix: issue 263246858 (#1496) 2023-01-30 12:30:04 -05:00
Ivan CheungGitHubivanmkc@google.com <ivanmkc@google.com>
9a6c36d016 Fixed cleanup code for matching engine index endpoint (#1493)
Co-authored-by: ivanmkc@google.com <ivanmkc@google.com>
2023-01-30 09:52:32 -05:00
Renovate Bot c889a57c9e chore(deps): update dependency isort to v5.12.0 2023-01-28 18:21:42 +00:00
Michael Hu dbe5d61929 Reduce cost of running prophet notebook 2023-01-27 20:18:33 -05:00
Andrew FerlitschandGitHub c180408f41 feat: model monitoring (#1488)
* fix: review

* fix: testing
2023-01-26 00:26:53 -08:00
gericdongandGitHub 869b19d342 Add new notebook to support the XAI zero metadata config feature (#1484)
* feat: add new notebook to support the XAI zero metadata config feature

* add missing packages

* Attempt to fix issue of -- user install not performed in the env

* Fixed package issues

* Addressed review comments

* Addressed review comments

* Addressed review comments
2023-01-25 10:02:50 -05:00
Michael HuandGitHub 3d967b180d add prophet on vertex pipelines notebook (#1320)
* add prophet on vertex pipelines notebook

* update notebook

* add explicit bq dependency

* add more explanations for what the pipeline is doing

* oops

* oops

* Update overview and add parameter descriptions

* foo

* foo

* add more parameters and types

* remove future tense and fix links

* fix formatting

* fix docs
2023-01-24 22:41:19 -08:00
halio-gandGitHub 49710a9225 Improve the training code to support the non-distributed job and add … (#1489)
* Improve the training code to support the non-distributed job and add the dashboard access.

* format the notebook.

* Use the 8888 instead of getting the env since DASHBOARD_PORT is not populated in the pipeline.

* Resolved the pull request comments.
2023-01-24 15:59:02 -08:00
Andrew FerlitschandGitHub 9d31463585 fix: TW updates (#1492) 2023-01-24 18:58:40 -05:00
ivanmkc@google.com bcd86a3707 Fixed typo for index display name 2023-01-25 08:56:59 +09:00
Mend RenovateandGitHub 9335ea3591 chore(deps): update dependency flake8 to v6 (#1298) 2023-01-24 11:31:29 -08:00
Andrew FerlitschandGitHub 0f7343feee migration: experiments (#1487)
* migration: experiments

* fix: review
2023-01-24 12:27:19 -05:00
Ivan NardiniandGitHub 12cd965ce6 new demand forecasting pipeline notebook (#1439)
* new demand forecasting pipeline notebook

* linter passed

* review notebook

* linter passed

* review notebook

* linter passed

* andy review

* linter passed
2023-01-24 08:44:14 -08:00
ivanmkc@google.com 4e5c75fed6 Renamed json to jsonl 2023-01-24 21:08:08 +09:00
Ivan Cheung b28f941abe Fixed typo 2023-01-24 21:02:46 +09:00
32632711ff Add co-hosting model notebook (#713)
* Add notebook for co-hosting model

* Add notebook for co-hosting model

* Change co-hosting model notebook inline link to officical

Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
Co-authored-by: Eric Schmidt <em.schmidt78@gmail.com>
2023-01-23 10:29:13 -08:00
junkourataandGitHub d676d87cce feat: Add E2E notebook featuring Vertex Feature Store, Training and Prediction (#1398)
* Add E2E tutorial for Feature Store

* Add Codeowner and fix the formatting and lining.

* Fixed lint
2023-01-23 10:24:51 -08:00
Douglass ChenandGitHub 153a8044b8 Update Colab notebooks' default fields and URLs (#1485)
* Add Cloud natural language pipeline colab notebook

* Add ready-to-go text classification pipeline colab notebook

* Ran reformatting scripts on text classification pipeline colab notebooks

* Update CODEOWNERS files

* Fix order of cells in cloud_natural_language_pipeline.ipynb

* Remove unused variables via linter for text classification colabs; fix classification variable for preprocessing component

* Minor fix: remove GCPC version requirement

* Minor fix: remove outputs

* fix formatting with nbfmt

* move ready-to-go pipeline to notebooks/community

* fix link

* update CODEOWNERS

* move text classification colabs to notebooks/community/pipelines

* Address initial comments on NL notebook

* Remove commented lines in NL notebook

* minor cell formatting

* clear outputs

* minor changes to NL notebook

* address comments for ready-to-go pipeline

* run linter locally

* add pipeline description to NL pipeline

* run linter locally (PR check could not lint)

* Add cell to examine metrics, update kernel restart cell from official template

* lint

* Update default fields and URLs in NL notebook

* Fix URLs in ready to go notebook

* run linter
2023-01-23 10:00:25 -08:00
Andrew FerlitschandGitHub c0d9250416 migration: labeling (#1479) 2023-01-23 08:58:28 -08:00
fd16e39f91 Added notebook demonstrating hyperparameter tuning using tensorboard (#1451)
* Added notebook demonstrating hyperparameter tuning using tensorboard

* Added notebook demonstrating hyperparameter tuning using tensorboard - linter finished

* Added notebook demonstrating hyperparameter tuning using tensorboard - first round of comments resolved

* Added notebook demonstrating hyperparameter tuning using tensorboard - fixing uncomment error

* fixing comment and lint error

* Jack's comments resolved

* fixing the cell that caused CI/CDtest error

* attempt to fix CI/CD issue with loading tensorboard

* attempt to fix TF import error

* fix CI/CD issues

* fix CI/CD issue

Co-authored-by: Andrew Ferlitsch <aferlitsch@gmail.com>
2023-01-19 18:13:02 -08:00
dependabot[bot]andGitHub 6628866130 Build(deps): Bump torch
Bumps [torch](https://github.com/pytorch/pytorch) from 1.8.1 to 1.13.1.
- [Release notes](https://github.com/pytorch/pytorch/releases)
- [Changelog](https://github.com/pytorch/pytorch/blob/master/RELEASE.md)
- [Commits](https://github.com/pytorch/pytorch/compare/v1.8.1...v1.13.1)

---
updated-dependencies:
- dependency-name: torch
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
2023-01-13 17:52:59 +00:00
wintwooandGitHub f681879a8a Merge branch 'GoogleCloudPlatform:main' into dataproc 2022-12-28 11:23:34 +11:00
Renovate Bot e01174f169 chore(deps): update dependency black to v22.12.0 2022-12-09 16:42:13 +00:00
Win Woo 62c903f47d Addressed PR feedback 2022-12-08 03:01:56 +00:00
Alexey Volkov da77846b88 fix: Fixed the Scikit-learn components 2022-12-06 15:42:27 -08:00
Chun-Hsiang Wang 7c123e68d3 samples: Remove experimental text from Pytorch sample. 2022-11-30 07:46:08 +00:00
Win Woo 9f26e8385d Show how to use experiments 2022-11-18 10:02:02 +00:00
uday kumarandGitHub 0fffaee6dc Merge branch 'main' into auto_tab_on_vertex_pipeline 2022-11-14 15:42:17 +05:30
udaypunna c34b7ab651 linter test 2022-11-14 07:47:02 +00:00
udaypunna d57d9be9d4 downgraded gcpc version 2022-11-14 07:46:18 +00:00
udaypunna 88f3ecf567 downgraded gcpc version 2022-11-14 07:45:12 +00:00
uday kumarandGitHub ecfbeaa8ce Merge branch 'main' into auto_tab_on_vertex_pipeline 2022-11-10 15:46:16 +05:30
udaypunna 3e5a319455 ran linter test 2022-11-10 10:13:43 +00:00
udaypunna f35ce3432a gcpc version downgraded 2022-11-10 10:12:56 +00:00
udaypunna 2e9ff4a941 linter test 2022-11-09 14:40:27 +00:00
udaypunna 61fbd37fa4 service account and version chnages 2022-11-09 14:39:38 +00:00
udaypunna 474e4e602b ran linter test 2022-11-09 14:14:47 +00:00
udaypunna e409cf710b added service account and downgraded gcpc version 2022-11-09 14:14:06 +00:00
udaypunna e9bac09dce ran linter test 2022-11-08 05:40:59 +00:00
udaypunna b799aad08a gcpc version and textual corrections 2022-11-08 05:40:09 +00:00
udaypunna 4d6bd1afff ran linter test 2022-11-07 07:46:14 +00:00
udaypunna b4040ff01d textual corrections 2022-11-07 07:45:25 +00:00
Alexey Volkov c29e4fde26 Train tabular models with many frameworks and import to Vertex AI using Pipelines
These pipelines were previously in community content.
We'd like to move them to the official folder.

These pipelines are:

* Working out of the box (code runs with zero modifications)
* End-to-end (from nothing to a Vertex Model)
* Feature multiple ML frameworks (TensorFlow, PyTorch, XGBoost, Scikit-learn)
* Feature multiple training objectives: tabular classification and tabular regression

The main files are Python-based pipeline code (`pipeline.py`).
2022-11-03 01:53:55 -07:00
udaypunna c6fdad745f linter test 2022-10-18 12:01:03 +00:00
udaypunna 3c60c2cb97 DAG issues 2022-10-18 12:00:18 +00:00
udaypunna 278a014817 ran linter test 2022-10-17 06:36:00 +00:00
udaypunna 89ec43c706 Cloud Storage bucket permission issues resolved 2022-10-17 06:35:26 +00:00
udaypunna ba6be7ee99 Cloud Storage bucket permission issues resolved 2022-10-17 06:30:12 +00:00
91 changed files with 19733 additions and 2945 deletions
@@ -103,7 +103,6 @@ class EndpointResourceCleanupManager(VertexAIResourceCleanupManager):
models.id for models in resource._gca_resource.deployed_models
]:
resource._undeploy(deployed_model_id=deployed_model_id)
resource.delete(force=True)
@@ -117,3 +116,7 @@ class MatchingEngineIndexResourceCleanupManager(VertexAIResourceCleanupManager):
class MatchingEngineIndexEndpointResourceCleanupManager(VertexAIResourceCleanupManager):
vertex_ai_resource = aiplatform.MatchingEngineIndexEndpoint
def delete(self, resource):
resource.undeploy_all()
resource.delete(force=True)
+1
View File
@@ -11,3 +11,4 @@ google-cloud-storage
google-cloud-build
ratemate
GitPython
tqdm
+3 -3
View File
@@ -2,9 +2,9 @@ git+https://github.com/tensorflow/docs
ipython
jupyter
nbconvert
black==22.10.0
black==22.12.0
pyupgrade==2.38.4
isort==5.10.1
flake8==4.0.1
isort==5.12.0
flake8==6.0.0
nbqa==1.5.3
@@ -1,3 +1,3 @@
torch==1.8.1
torch==1.13.1
torchvision==0.9.1
tensorboard==2.5.0
@@ -31,17 +31,7 @@
"source": [
"# Deploying a PyTorch Text Classification Model on [Vertex AI](https://cloud.google.com/vertex-ai)\n",
"\n",
"**This is an Experimental release**, covered by the Pre-GA Offerings Terms of your Google Cloud Platform [Terms of Service](https://cloud.google.com/terms).\n",
"\n",
"Experiments are focused on validating a prototype and are not guaranteed to be released. They are not intended for production use or covered by any SLA, support obligation, or deprecation policy and might be subject to backward-incompatible changes.\n",
"\n",
"**Kindly drop us a note before you run any scale tests.**\n",
"\n",
"**Do not hesitate to contact vertexai-prediction-preview-feedback@google.com if you have any questions or run into any issues.**\n",
"\n",
"The usage of the product is free during the Experimental release period: you will still incur charges for other GCP products usage, such as storage.\n",
"\n",
"The projects need to be added to the allowlist in order to deploy PyTorch models using Vertex AI Prediction pre-built PyTorch images. If you are interested in the feature, please send an email to vertexai-prediction-preview-feedback@google.com to provide your project numbers OR project ids."
"**Kindly reach out to Vertex AI before you run any scale tests or you have any questions.**\n"
]
},
{
+1 -1
View File
@@ -40,4 +40,4 @@
/notebooks/community/pipelines/google_cloud_pipeline_components_bqml_pipeline_anomaly_detection.ipynb @inardini
/notebooks/community/pipelines/google_cloud_pipeline_components_cloud_natural_language_pipeline.ipynb @Narwhalprime
/notebooks/community/pipelines/google_cloud_pipeline_components_ready_to_go_text_classification_pipeline.ipynb @Narwhalprime
/notebooks/community/feature_store/get_started_vertex_feature_store.ipynb @junkourata
File diff suppressed because it is too large Load Diff
@@ -1333,7 +1333,6 @@
"Next, you compile the pipeline and then exeute it. The pipeline takes the following parameters, which are passed as the dictionary `parameter_values`:\n",
"\n",
"- `display_name`: A human readable name for the pipeline job.\n",
"- `import_file`: The Cloud Storage location to the dataset.\n",
"- `worker_pool_specs`: The the machine and container, and auto-scaling requirements, as well as command line arguments.\n",
"- `study_spec_metrics`: The metrics to optimize in the study trials.\n",
"- `study_spec_parameters`: The parameters to tune."
@@ -54,18 +54,20 @@
{
"cell_type": "markdown",
"metadata": {
"id": "tvgnzT1CKxrO"
"id": "239ba71252d3"
},
"source": [
"## Overview\n",
"\n",
"This notebook shows how to use `Vertex AI Pipelines` and `BigQuery ML pipeline components` to train and evaluate a demand forecasting model.\n",
"\n",
"### Dataset\n",
"\n",
"The dataset is a modified version of the dataset in [Build and visualize demand forecast predictions using Datastream, Dataflow, BigQuery ML, and Looker\n",
"](https://cloud.google.com/architecture/build-visualize-demand-forecast-prediction-datastream-dataflow-bigqueryml-looker) solution architecture\n",
"\n",
"This notebook shows how to use `Vertex AI Pipelines` and `BigQuery ML pipeline components` to train and evaluate a demand forecasting model."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "25c28706c23e"
},
"source": [
"### Objective\n",
"\n",
"In this tutorial, you learn how to train and evaluate a BigQuery ML model using Vertex AI Pipelines and BigQuery ML pipeline components. \n",
@@ -87,8 +89,27 @@
" - Generate the ARIMA Plus forecasts\n",
" - Generate the ARIMA PLUS forecast explainations\n",
"- Compile the pipeline.\n",
"- Execute the pipeline.\n",
"- Execute the pipeline."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "586acfa9b502"
},
"source": [
"### Dataset\n",
"\n",
"The dataset is a modified version of the dataset in [Build and visualize demand forecast predictions using Datastream, Dataflow, BigQuery ML, and Looker\n",
"](https://cloud.google.com/architecture/build-visualize-demand-forecast-prediction-datastream-dataflow-bigqueryml-looker) solution architecture\n"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "tvgnzT1CKxrO"
},
"source": [
"### Costs \n",
"\n",
"This tutorial uses billable components of Google Cloud:\n",
@@ -352,9 +373,8 @@
"id": "06571eb4063b"
},
"source": [
"#### Timestamp\n",
"\n",
"If you are in a live tutorial session, you might be using a shared test account or project. To avoid name collisions between users on resources created, you create a timestamp for each instance session, and append it onto the name of resources you create in this tutorial."
"#### UUID\n",
"If you are in a live tutorial session, you might be using a shared test account or project. To avoid name collisions between users on resources created, you create a uuid for each instance session, and append it onto the name of resources you create in this tutorial."
]
},
{
@@ -365,9 +385,16 @@
},
"outputs": [],
"source": [
"from datetime import datetime\n",
"import random\n",
"import string\n",
"\n",
"TIMESTAMP = datetime.now().strftime(\"%Y%m%d%H%M%S\")"
"\n",
"# Generate a uuid of a specifed length(default=8)\n",
"def generate_uuid(length: int = 8) -> str:\n",
" return \"\".join(random.choices(string.ascii_lowercase + string.digits, k=length))\n",
"\n",
"\n",
"UUID = generate_uuid()"
]
},
{
@@ -485,7 +512,7 @@
"outputs": [],
"source": [
"if BUCKET_NAME == \"\" or BUCKET_NAME is None or BUCKET_NAME == \"[your-bucket-name]\":\n",
" BUCKET_NAME = PROJECT_ID + \"-aip-\" + TIMESTAMP\n",
" BUCKET_NAME = PROJECT_ID + \"-aip-\" + UUID\n",
" BUCKET_URI = f\"gs://{BUCKET_NAME}\""
]
},
@@ -706,6 +733,7 @@
"KFP_COMPONENTS_PATH = \"components\"\n",
"PIPELINES_PATH = \"pipelines\"\n",
"\n",
"! mkdir -m 777 -p {DATA_PATH}\n",
"! mkdir -m 777 -p {KFP_COMPONENTS_PATH}\n",
"! mkdir -m 777 -p {PIPELINES_PATH}"
]
@@ -771,7 +799,7 @@
" --location={LOCATION} \\\n",
" --source_format=CSV \\\n",
" --skip_leading_rows=1\\\n",
" fast_fresh.orders_{TIMESTAMP} \\\n",
" fast_fresh.orders_{UUID} \\\n",
" {RAW_DATA_URI} \\\n",
" time_of_sale:DATETIME,order_id:INTEGER,product_name:STRING,price:NUMERIC,quantity:NUMERIC,payment_method:STRING,store_id:INTEGER,user_id:INTEGER"
]
@@ -782,7 +810,7 @@
"id": "ZrgOD30o7HcL"
},
"source": [
"## BQML Training Formalization\n",
"## BigQuery ML Training Formalization\n",
"\n",
"In the next cells, you build the components and pipeline to train and evaluate the BQML demand forecasting model."
]
@@ -820,13 +848,13 @@
"BQ_EVALUATE_MODEL_TABLE_PREFIX = \"orders_arima_model_evaluate\"\n",
"BQ_FORECAST_TABLE_PREFIX = \"orders_arima_forecast\"\n",
"BQ_EXPLAIN_FORECAST_TABLE_PREFIX = \"orders_arima_explain_forecast\"\n",
"BQ_ORDERS_TABLE = f\"{BQ_ORDERS_TABLE_PREFIX}_{TIMESTAMP}\"\n",
"BQ_TRAINING_TABLE = f\"{BQ_TRAINING_TABLE_PREFIX}_{TIMESTAMP}\"\n",
"BQ_MODEL_TABLE = f\"{BQ_MODEL_TABLE_PREFIX}_{TIMESTAMP}\"\n",
"BQ_EVALUATE_TS_TABLE = f\"{BQ_EVALUATE_TS_TABLE_PREFIX}_{TIMESTAMP}\"\n",
"BQ_EVALUATE_MODEL_TABLE = f\"{BQ_EVALUATE_MODEL_TABLE_PREFIX}_{TIMESTAMP}\"\n",
"BQ_FORECAST_TABLE = f\"{BQ_FORECAST_TABLE_PREFIX}_{TIMESTAMP}\"\n",
"BQ_EXPLAIN_FORECAST_TABLE = f\"{BQ_EXPLAIN_FORECAST_TABLE_PREFIX}_{TIMESTAMP}\"\n",
"BQ_ORDERS_TABLE = f\"{BQ_ORDERS_TABLE_PREFIX}_{UUID}\"\n",
"BQ_TRAINING_TABLE = f\"{BQ_TRAINING_TABLE_PREFIX}_{UUID}\"\n",
"BQ_MODEL_TABLE = f\"{BQ_MODEL_TABLE_PREFIX}_{UUID}\"\n",
"BQ_EVALUATE_TS_TABLE = f\"{BQ_EVALUATE_TS_TABLE_PREFIX}_{UUID}\"\n",
"BQ_EVALUATE_MODEL_TABLE = f\"{BQ_EVALUATE_MODEL_TABLE_PREFIX}_{UUID}\"\n",
"BQ_FORECAST_TABLE = f\"{BQ_FORECAST_TABLE_PREFIX}_{UUID}\"\n",
"BQ_EXPLAIN_FORECAST_TABLE = f\"{BQ_EXPLAIN_FORECAST_TABLE_PREFIX}_{UUID}\"\n",
"\n",
"BQ_TRAIN_CONFIGURATION = {\n",
" \"destinationTable\": {\n",
@@ -1022,7 +1050,7 @@
"id": "pcSL1FHk69KT"
},
"source": [
"### Build the BQML training pipeline\n",
"### Build the BigQuery ML training pipeline\n",
"\n",
"Define your workflow using Kubeflow Pipelines DSL package. \n",
"\n",
@@ -1094,8 +1122,8 @@
" location=location,\n",
" ).set_display_name(\"get train data\")\n",
"\n",
" # Train the ARIMA PLUS model\n",
" bq_arima_model_op = (\n",
" # Run an ARIMA PLUS experiment\n",
" bq_arima_model_exp_op = (\n",
" BigqueryCreateModelJobOp(\n",
" query=f\"\"\"\n",
" -- create model table\n",
@@ -1104,10 +1132,7 @@
" MODEL_TYPE = \\'ARIMA_PLUS\\',\n",
" TIME_SERIES_TIMESTAMP_COL = \\'hourly_timestamp\\',\n",
" TIME_SERIES_DATA_COL = \\'total_sold\\',\n",
" TIME_SERIES_ID_COL = [\\'product_name\\'],\n",
" MODEL_REGISTRY = \\'vertex_ai\\',\n",
" VERTEX_AI_MODEL_ID = \\'order_demand_forecasting\\',\n",
" VERTEX_AI_MODEL_VERSION_ALIASES = [\\'staging\\']\n",
" TIME_SERIES_ID_COL = [\\'product_name\\']\n",
" ) AS\n",
" SELECT\n",
" hourly_timestamp,\n",
@@ -1119,7 +1144,7 @@
" project=project,\n",
" location=location,\n",
" )\n",
" .set_display_name(\"train arima plus model\")\n",
" .set_display_name(\"run arima+ model experiment\")\n",
" .after(create_training_dataset_op)\n",
" )\n",
"\n",
@@ -1128,12 +1153,12 @@
" BigqueryMLArimaEvaluateJobOp(\n",
" project=project,\n",
" location=location,\n",
" model=bq_arima_model_op.outputs[\"model\"],\n",
" model=bq_arima_model_exp_op.outputs[\"model\"],\n",
" show_all_candidate_models=False,\n",
" job_configuration_query=bq_evaluate_time_series_configuration,\n",
" )\n",
" .set_display_name(\"evaluate arima plus time series\")\n",
" .after(bq_arima_model_op)\n",
" .after(bq_arima_model_exp_op)\n",
" )\n",
"\n",
" # Evaluate ARIMA Plus model\n",
@@ -1141,12 +1166,12 @@
" BigqueryEvaluateModelJobOp(\n",
" project=project,\n",
" location=location,\n",
" model=bq_arima_model_op.outputs[\"model\"],\n",
" model=bq_arima_model_exp_op.outputs[\"model\"],\n",
" query_statement=f\"\"\"SELECT * FROM `{project}.{bq_dataset}.{bq_training_table}` WHERE split='TEST'\"\"\",\n",
" job_configuration_query=bq_evaluate_model_configuration,\n",
" )\n",
" .set_display_name(\"evaluate arima plus model\")\n",
" .after(bq_arima_model_op)\n",
" .after(bq_arima_model_exp_op)\n",
" )\n",
"\n",
" # Plot model metrics\n",
@@ -1164,6 +1189,34 @@
" < PERF_THRESHOLD,\n",
" name=\"avg. mae good\",\n",
" ):\n",
" # Train the ARIMA PLUS model\n",
" bq_arima_model_op = (\n",
" BigqueryCreateModelJobOp(\n",
" query=f\"\"\"\n",
" -- create model table\n",
" CREATE OR REPLACE MODEL `{project}.{bq_dataset}.{bq_model_table}`\n",
" OPTIONS(\n",
" MODEL_TYPE = \\'ARIMA_PLUS\\',\n",
" TIME_SERIES_TIMESTAMP_COL = \\'hourly_timestamp\\',\n",
" TIME_SERIES_DATA_COL = \\'total_sold\\',\n",
" TIME_SERIES_ID_COL = [\\'product_name\\'],\n",
" MODEL_REGISTRY = \\'vertex_ai\\',\n",
" VERTEX_AI_MODEL_ID = \\'order_demand_forecasting\\',\n",
" VERTEX_AI_MODEL_VERSION_ALIASES = [\\'staging\\']\n",
" ) AS\n",
" SELECT\n",
" DATETIME_TRUNC(time_of_sale, HOUR) as hourly_timestamp,\n",
" product_name,\n",
" SUM(quantity) AS total_sold,\n",
" FROM `{project}.{bq_dataset}.{bq_orders_table}`\n",
" GROUP BY hourly_timestamp, product_name;\n",
" \"\"\",\n",
" project=project,\n",
" location=location,\n",
" )\n",
" .set_display_name(\"train arima+ model\")\n",
" .after(get_evaluation_model_metrics_op)\n",
" )\n",
"\n",
" # Generate the ARIMA PLUS forecasts\n",
" bq_arima_forecast_op = (\n",
@@ -1224,7 +1277,7 @@
"source": [
"### Execute your pipeline\n",
"\n",
"Next, you execute the pipeline. It takes the following parameters which we set as default:\n",
"Next, we execute the pipeline. It takes the following parameters which we set as default:\n",
"\n",
"- `bq_dataset`: The BigQuery dataset to train on.\n",
"- `bq_orders_table` : The BigQuery table of raw data.\n",
@@ -1266,7 +1319,7 @@
"source": [
"### View BigQuery ML training pipeline results\n",
"\n",
"Finally, you will view the artifact outputs of each task in the pipeline."
"Finally, you view the artifact outputs of each task in the pipeline."
]
},
{
@@ -1342,8 +1395,8 @@
"print(\"bigquery-ml-arima-evaluate-job\")\n",
"artifacts = print_pipeline_output(bqml_pipeline, \"bigquery-ml-arima-evaluate-job\")\n",
"print(\"\\n\\n\")\n",
"print(\"get-model-evaluation-metrics\")\n",
"artifacts = print_pipeline_output(bqml_pipeline, \"get-model-evaluation-metrics\")\n",
"print(\"bigquery-evaluate-model-job\")\n",
"artifacts = print_pipeline_output(bqml_pipeline, \"bigquery-evaluate-model-job\")\n",
"print(\"\\n\\n\")\n",
"print(\"bigquery-forecast-model-job\")\n",
"artifacts = print_pipeline_output(bqml_pipeline, \"bigquery-forecast-model-job\")\n",
@@ -1407,7 +1460,8 @@
"\n",
"# Remove local resorces\n",
"! rm -rf {KFP_COMPONENTS_PATH}\n",
"! rm -rf {PIPELINES_PATH}"
"! rm -rf {PIPELINES_PATH}\n",
"! rm -rf {DATA_PATH}"
]
}
],
@@ -42,12 +42,12 @@
"<table align=\"left\">\n",
"\n",
" <td>\n",
" <a href=\"https://colab.research.google.com/github/GoogleCloudPlatform/vertex-ai-samples/blob/master/notebooks/community/natural_language/cloud_natural_language_pipeline.ipynb\">\n",
" <a href=\"https://colab.research.google.com/github/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/pipelines/google_cloud_pipeline_components_cloud_natural_language_pipeline.ipynb\">\n",
" <img src=\"https://cloud.google.com/ml-engine/images/colab-logo-32px.png\" alt=\"Colab logo\"> Run in Colab\n",
" </a>\n",
" </td>\n",
" <td>\n",
" <a href=\"https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/natural_language/cloud_natural_language_pipeline.ipynb\">\n",
" <a href=\"https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/pipelines/google_cloud_pipeline_components_cloud_natural_language_pipeline.ipynb\">\n",
" <img src=\"https://cloud.google.com/ml-engine/images/github-logo-32px.png\" alt=\"GitHub logo\">\n",
" View on GitHub\n",
" </a>\n",
@@ -328,7 +328,7 @@
},
"outputs": [],
"source": [
"PROJECT_ID = \"cloud-ml-language-test\" # @param {type:\"string\"}\n",
"PROJECT_ID = \"your-project-id\" # @param {type:\"string\"}\n",
"if PROJECT_ID == \"\" or PROJECT_ID is None or PROJECT_ID == \"[your-project-id]\":\n",
" # Get your GCP project id from gcloud\n",
" shell_output = !gcloud config list --format 'value(core.project)' 2>/dev/null\n",
@@ -368,7 +368,7 @@
"source": [
"REGION = \"us\" # @param {type:\"string\"}\n",
"LOCATION = \"us-central1\" # @param {type:\"string\"}\n",
"TRAINING_DATA_LOCATION = \"gs://dougchen-20221130-pipeline-colab-test/data-00001-of-00001.jsonl\" # @param {type:\"string\"}\n",
"TRAINING_DATA_LOCATION = \"gs://your-training-data-location\" # @param {type:\"string\"}\n",
"TASK_TYPE = \"CLASSIFICATION\" # @param [\"CLASSIFICATION\", \"MULTILABEL_CLASSIFICATION\"]"
]
},
@@ -740,7 +740,7 @@
"\n",
"This sends a create pipeline job request to Vertex Pipelines. Note that this task run synchronously and may take a while to complete.\n",
"\n",
"You may view the progress of the job at any time by clicking on the generated links (after \"View Pipeline Job\" in the console output of the cell below). Once the pipeline finishes, you may examine the artifacts produced from this pipeline. See "
"You may view the progress of the job at any time by clicking on the generated links (after \"View Pipeline Job\" in the console output of the cell below). Once the pipeline finishes, you may examine the artifacts produced from this pipeline."
]
},
{
@@ -42,12 +42,12 @@
"<table align=\"left\">\n",
"\n",
" <td>\n",
" <a href=\"https://colab.research.google.com/github/GoogleCloudPlatform/vertex-ai-samples/blob/master/community-content/ready_to_go_text_classification_pipeline/ready_to_go_text_classification_pipeline.ipynb\">\n",
" <a href=\"https://colab.research.google.com/github/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/pipelines/google_cloud_pipeline_components_ready_to_go_text_classification_pipeline.ipynb\">\n",
" <img src=\"https://cloud.google.com/ml-engine/images/colab-logo-32px.png\" alt=\"Colab logo\"> Run in Colab\n",
" </a>\n",
" </td>\n",
" <td>\n",
" <a href=\"https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/community-content/ready_to_go_text_classification_pipeline/ready_to_go_text_classification_pipeline.ipynb\">\n",
" <a href=\"https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/pipelines/google_cloud_pipeline_components_ready_to_go_text_classification_pipeline.ipynb\">\n",
" <img src=\"https://cloud.google.com/ml-engine/images/github-logo-32px.png\" alt=\"GitHub logo\">\n",
" View on GitHub\n",
" </a>\n",
@@ -706,7 +706,7 @@
")\n",
"\n",
"# The GCS directory for keeping staging files for model evaluation.\n",
"ROOT_DIR = 'f\"{BASE_OUTPUT_DIR}/root\"' # @param {type:\"string\"}"
"ROOT_DIR = \"'f\\\"{BASE_OUTPUT_DIR}/root\\\"'\" # @param {type:\"string\"}"
]
},
{
File diff suppressed because one or more lines are too long
@@ -0,0 +1,987 @@
{
"cells": [
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "18ebbd838e32"
},
"outputs": [],
"source": [
"# Copyright 2023 Google LLC\n",
"#\n",
"# Licensed under the Apache License, Version 2.0 (the \"License\");\n",
"# you may not use this file except in compliance with the License.\n",
"# You may obtain a copy of the License at\n",
"#\n",
"# https://www.apache.org/licenses/LICENSE-2.0\n",
"#\n",
"# Unless required by applicable law or agreed to in writing, software\n",
"# distributed under the License is distributed on an \"AS IS\" BASIS,\n",
"# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.\n",
"# See the License for the specific language governing permissions and\n",
"# limitations under the License."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "219f1b1fe8fe"
},
"source": [
"# Deploy and host a Stable Diffusion model on Vertex AI\n"
]
},
{
"attachments": {},
"cell_type": "markdown",
"metadata": {
"id": "JAPoU8Sm5E6e"
},
"source": [
"<table align=\"left\">\n",
"\n",
" <td>\n",
" <a href=\"https://colab.research.google.com/github/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/vertex_endpoints/torchserve/dreambooth_stablediffusion.ipynb\">\n",
" <img src=\"https://cloud.google.com/ml-engine/images/colab-logo-32px.png\" alt=\"Colab logo\"> Run in Colab\n",
" </a>\n",
" </td>\n",
" <td>\n",
" <a href=\"https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/vertex_endpoints/torchserve/dreambooth_stablediffusion.ipynb\">\n",
" <img src=\"https://cloud.google.com/ml-engine/images/github-logo-32px.png\" alt=\"GitHub logo\">\n",
" View on GitHub\n",
" </a>\n",
" </td>\n",
" <td>\n",
" <a href=\"https://console.cloud.google.com/vertex-ai/workbench/deploy-notebook?download_url=https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/vertex_endpoints/torchserve/dreambooth_stablediffusion.ipynb\">\n",
" <img src=\"https://lh3.googleusercontent.com/UiNooY4LUgW_oTvpsNhPpQzsstV5W8F7rYgxgGBD85cWJoLmrOzhVs_ksK_vgx40SHs7jCqkTkCk=e14-rj-sc0xffffff-h130-w32\" alt=\"Vertex AI logo\">\n",
" Open in Vertex AI Workbench\n",
" </a>\n",
" </td>\n",
"</table>"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "fce05a8186d6"
},
"source": [
"## Overview\n",
"\n",
"This notebook demonstrates how to deploy and host a fine-tuned [Stable Diffusion 1.5](https://huggingface.co/runwayml/stable-diffusion-v1-5) model on Vertex AI. For hosting, you use the PyTorch 3 container built for Vertex AI with [TorchServe](https://pytorch.org/serve/index.html)."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "c76216b03fec"
},
"source": [
"### Objective\n",
"\n",
"In this tutorial, you learn how to host and deploy a Stable Diffusion 1.5 model on Vertex AI.\n",
"\n",
"This tutorial uses the following Google Cloud ML services:\n",
"\n",
"+ Vertex AI `Model` resource\n",
"+ Vertex AI `Endpoint` resource\n",
"\n",
"The steps performed include:\n",
"\n",
"+ Create a `torchserve` handler for responding to prediction requests.\n",
"+ Upload a Stable Diffusion 1.5 model on a prebuilt PyTorch container in Vertex AI.\n",
"+ Deploy a model to a Vertex AI Endpoint.\n",
"+ Send requests to the endpoint and parse the responses using Vertex AI Prediction service."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "c6deba5a8557"
},
"source": [
"### Model\n",
"\n",
"This notebook uses a collection of model artifacts fine-tuned to generate images of a small dog. These are the same images used in the original [DreamBooth paper](https://dreambooth.github.io/)."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "911dc651ea9c"
},
"source": [
"### Costs\n",
"\n",
"This tutorial uses billable components of Google Cloud:\n",
"\n",
"* Vertex AI models\n",
"* Vertex AI endpoints\n",
"* Vertex AI prediction\n",
"* Cloud Storage\n",
"* (Optionally) Vertex AI Workbench\n",
"\n",
"Learn about [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing) and [Cloud Storage pricing](https://cloud.google.com/storage/pricing), and use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to generate a cost estimate based on your projected usage."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "0d36bd3d53fa"
},
"source": [
"## Hardware requirements\n",
"\n",
"This notebook requires that you use a GPU with a sufficient amount of VRAM available. It was tested on a `NVIDIA Tesla A100 GPU` with 85 GB of VRAM. Run the following cell to ensure that you have the correct hardware configuration."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "3dd4022552e5"
},
"outputs": [],
"source": [
"!nvidia-smi --query-gpu=name,memory.total,memory.free --format=\"csv,noheader\""
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "a782627e5f73"
},
"source": [
"### Create a user-managed notebook on Vertex AI\n",
"\n",
"If you are using Vertex AI Workbench, you can create a notebook with the correct configuration by doing the following:\n",
"\n",
"+ Go to [Vertex AI Workbench](https://console.cloud.google.com/vertex-ai/workbench/user-managed) in the Google Cloud Console.\n",
"+ Click **New Notebook** and then click **PyTorch 1.13** > **With 1 NVIDIA T4**.\n",
"+ In the **New notebook** dialog box, click **Advanced Options**. The **Create a user-managed notebook** page opens up.\n",
"+ In the **Create a user-managed notebook** page, do the following:\n",
" * In the **Notebook name** box, type a name for your notebook, for example \"my-stablediffusion-nb\".\n",
" * In the **Machine type** drop-down, select **A2 highgpu** > **a2-highgpu-1g**.\n",
" * In the **GPU type** drop-down, select **NVIDIA Tesla A100**.\n",
" * Check the box next to **Install NVIDIA GPU driver automatically for me**\n",
" * Expand **Disk(s)** and do the following:\n",
" - Under **Boot disk type**, select **SSD Persistent Disk**.\n",
" - Under **Data disk type**, select **SSD Persistent Disk**.\n",
" * Click **Create**."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "FAPoU8Sm5E6e"
},
"source": [
"<div style=\"background:#feefe3; padding:5px; color:#aa0000\">\n",
"<strong>Caution:</strong> Using a Vertex AI Workbench notebook with the above configuration can increase your costs significantly. You can estimate your costs using the <a href=\"https://cloud.google.com/products/calculator\"><u>costs calculator</u></a>.</div>"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "0bb4201cc99a"
},
"source": [
"## Installation\n",
"\n",
"Install the following packages required to execute this notebook.\n",
"\n",
"**Note**: You might need to change the version of PyTorch (`torch`) installed by `pip`."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "9c769df171a6"
},
"outputs": [],
"source": [
"%%writefile requirements.txt\n",
"diffusers\n",
"ftfy\n",
"google-cloud-aiplatform\n",
"gradio\n",
"ninja\n",
"tensorboard==1.15.0\n",
"torch\n",
"torchaudio\n",
"torchvision\n",
"torchserve\n",
"torch-model-archiver\n",
"torch-workflow-archiver\n",
"transformers"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "e46804ac90d8"
},
"outputs": [],
"source": [
"%pip install -r requirements.txt"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "58707a750154"
},
"source": [
"### Colab only: Uncomment the following cell to restart the kernel."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "77c11549298a"
},
"outputs": [],
"source": [
"# Automatically restart kernel after installs so that your environment can access the new packages\n",
"# import IPython\n",
"\n",
"# app = IPython.Application.instance()\n",
"# app.kernel.do_shutdown(True)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "294df346a918"
},
"source": [
"## Before you begin"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "BF1j6f9HApxa"
},
"source": [
"### Set up your Google Cloud project\n",
"\n",
"**The following steps are required, regardless of your notebook environment.**\n",
"\n",
"1. [Select or create a Google Cloud project](https://console.cloud.google.com/cloud-resource-manager). When you first create an account, you get a $300 free credit towards your compute/storage costs.\n",
"\n",
"2. [Make sure that billing is enabled for your project](https://cloud.google.com/billing/docs/how-to/modify-project).\n",
"\n",
"3. [Enable the Vertex AI API](https://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com).\n",
"\n",
"4. If you are running this notebook locally, you need to install the [Cloud SDK](https://cloud.google.com/sdk)."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "WReHDGG5g0XY"
},
"source": [
"#### Set your project ID\n",
"\n",
"**If you don't know your project ID**, try the following:\n",
"* Run `gcloud config list`.\n",
"* Run `gcloud projects list`.\n",
"* See the support page: [Locate the project ID](https://support.google.com/googleapi/answer/7014113)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "oM1iC_MfAts1"
},
"outputs": [],
"source": [
"PROJECT_ID = \"[your-project-id]\" # @param {type:\"string\"}\n",
"\n",
"# Set the project id\n",
"! gcloud config set project {PROJECT_ID}"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "region"
},
"source": [
"#### Region\n",
"\n",
"You can also change the `REGION` variable used by Vertex AI. Learn more about [Vertex AI regions](https://cloud.google.com/vertex-ai/docs/general/locations)."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "region"
},
"outputs": [],
"source": [
"REGION = \"us-central1\""
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "sBCra4QMA2wR"
},
"source": [
"### Authenticate your Google Cloud account\n",
"\n",
"Depending on your Jupyter environment, you may have to manually authenticate. Follow the relevant instructions below."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "74ccc9e52986"
},
"source": [
"**1. Vertex AI Workbench**\n",
"* Do nothing as you are already authenticated."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "de775a3773ba"
},
"source": [
"**2. Local JupyterLab instance, uncomment and run:**"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "254614fa0c46"
},
"outputs": [],
"source": [
"# ! gcloud auth login"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "ef21552ccea8"
},
"source": [
"**3. Colab, uncomment and run:**"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "603adbbf0532"
},
"outputs": [],
"source": [
"# from google.colab import auth\n",
"# auth.authenticate_user()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "f6b2ccc891ed"
},
"source": [
"**4. Service account or other**\n",
"* See how to grant Cloud Storage permissions to your service account at https://cloud.google.com/storage/docs/gsutil/commands/iam#ch-examples."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "cb5c4ca3e851"
},
"source": [
"### Import libraries"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "7348591eda51"
},
"outputs": [],
"source": [
"import base64\n",
"import math\n",
"\n",
"import torch\n",
"from diffusers import StableDiffusionPipeline\n",
"from google.cloud import aiplatform\n",
"from IPython import display\n",
"from PIL import Image\n",
"from torch import autocast"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "697566b5f660"
},
"source": [
"## Optional: View model inferences\n",
"\n",
"Before uploading the model to Vertex AI, you can review the expected output from the model. The model used in this notebook is available for your use and can be downloaded from Cloud Storage. This download may take a few minutes to complete."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "d63df8d91215"
},
"outputs": [],
"source": [
"!gsutil -m cp gs://cloud-samples-data/vertex-ai/model-deployment/models/stable-diffusion/model_artifacts.zip \\\n",
" ."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "89d613ae9573"
},
"outputs": [],
"source": [
"!unzip model_artifacts.zip"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "9192ea4f3b57"
},
"source": [
"### Create new images\n",
"\n",
"With everything in place, you can now generate new images from the Stable Diffusion model. First you must load your model into a `StableDiffusionPipeline`."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "cd1f30223b79"
},
"outputs": [],
"source": [
"model_path = \"model_artifacts\"\n",
"\n",
"pipe = StableDiffusionPipeline.from_pretrained(\n",
" model_path, torch_dtype=torch.float16\n",
").to(\"cuda\")\n",
"\n",
"g_cuda = None"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "dd91520f58cf"
},
"outputs": [],
"source": [
"g_cuda = torch.Generator(device=\"cuda\")\n",
"seed = 52362\n",
"g_cuda.manual_seed(seed)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "1e9c5d734b9f"
},
"source": [
"With the model loaded into a `StableDiffusionPipeline`, you can now generate results (inferences) from the model. Each set of inference requires an input (called a [prompt](https://learnprompting.org/)) that specifies what the model should create.\n",
"\n",
"You can also vary other inputs into the model, as shown in the following cell."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "a28b73de55ce"
},
"outputs": [],
"source": [
"prompt = \"photo of examplePup dog in a Monet style\"\n",
"\n",
"num_samples = 4\n",
"num_batches = 1\n",
"num_columns = 2\n",
"guidance_scale = 10\n",
"num_inference_steps = 50\n",
"height = 512\n",
"width = 512"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "11a7bb79de59"
},
"outputs": [],
"source": [
"def image_grid(imgs, cols):\n",
" total = len(imgs)\n",
" rows = math.ceil(total / cols)\n",
"\n",
" w, h = imgs[0].size\n",
" grid = Image.new(\"RGB\", size=(cols * w, rows * h))\n",
" grid_w, grid_h = grid.size\n",
"\n",
" for i, img in enumerate(imgs):\n",
" grid.paste(img, box=(i % cols * w, i // cols * h))\n",
" return grid\n",
"\n",
"\n",
"all_images = []\n",
"for _ in range(num_batches):\n",
" with autocast(\"cuda\"):\n",
" images = pipe(\n",
" [prompt] * num_samples,\n",
" height=height,\n",
" width=width,\n",
" num_inference_steps=num_inference_steps,\n",
" guidance_scale=guidance_scale,\n",
" ).images\n",
" all_images.extend(images)\n",
"\n",
"\n",
"grid = image_grid(all_images, num_columns)\n",
"grid"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "bbc72963ff88"
},
"source": [
"## Deploy the model to Vertex AI\n",
"\n",
"You can host your Stable Diffusion 1.5 model on a Vertex AI endpoint where you can get inferences from it online. Uploading your model is a four step process: \n",
"\n",
"1. Create a custom TorchServe handler.\n",
"1. Upload the model artifacts onto Cloud Storage.\n",
"2. Create a Vertex AI model with the model artifacts and a prebuilt PyTorch container image.\n",
"3. Deploy the Vertex AI model onto an endpoint."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "eafbb0e0-40e6-43a0-a38e-edc54323da51"
},
"source": [
"### Create the custom TorchServe handler\n",
"\n",
"The model deployed to Vertex AI uses [TorchServe](https://pytorch.org/serve/) to handle requests and return responses from the model. You must create a custom TorchServe handler to include in with the model artifacts uploaded to Vertex AI.\n",
"\n",
"The handler file should be included in the directory with the other model artifacts."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "94567a87-9d74-4c87-a749-306ddaf01b61"
},
"outputs": [],
"source": [
"%%writefile model_artifacts/handler.py\n",
"\n",
"\"\"\"Customized handler for Stable Diffusion 1.5.\"\"\"\n",
"import base64\n",
"import logging\n",
"from io import BytesIO\n",
"\n",
"import torch\n",
"from diffusers import EulerDiscreteScheduler\n",
"from diffusers import StableDiffusionPipeline\n",
"from ts.torch_handler.base_handler import BaseHandler\n",
"\n",
"logger = logging.getLogger(__name__)\n",
"model_id = 'runwayml/stable-diffusion-v1-5'\n",
"\n",
"\n",
"class ModelHandler(BaseHandler):\n",
"\n",
" def __init__(self):\n",
" self.initialized = False\n",
" self.map_location = None\n",
" self.device = None\n",
" self.use_gpu = True\n",
" self.store_avg = True\n",
" self.pipe = None\n",
"\n",
" def initialize(self, context):\n",
" \"\"\"Initializes the pipe.\"\"\"\n",
" properties = context.system_properties\n",
" gpu_id = properties.get('gpu_id')\n",
"\n",
" self.map_location, self.device, self.use_gpu = \\\n",
" ('cuda', torch.device('cuda:' + str(gpu_id)),\n",
" True) if torch.cuda.is_available() else \\\n",
" ('cpu', torch.device('cpu'), False)\n",
"\n",
" # Use the Euler scheduler here instead\n",
" scheduler = EulerDiscreteScheduler.from_pretrained(model_id,\n",
" subfolder='scheduler')\n",
" pipe = StableDiffusionPipeline.from_pretrained(model_id,\n",
" scheduler=scheduler,\n",
" torch_dtype=torch.float16)\n",
" pipe = pipe.to('cuda')\n",
" # Uncomment the following line to reduce the GPU memory usage.\n",
" # pipe.enable_attention_slicing()\n",
" self.pipe = pipe\n",
"\n",
" self.initialized = True\n",
"\n",
" def preprocess(self, requests):\n",
" \"\"\"Noting to do here.\"\"\"\n",
" logger.info('requests: %s', requests)\n",
" return requests\n",
"\n",
" def inference(self, preprocessed_data, *args, **kwargs):\n",
" \"\"\"Run the inference.\"\"\"\n",
" images = []\n",
" for pd in preprocessed_data:\n",
" prompt = pd['prompt']\n",
" images.extend(self.pipe(prompt).images)\n",
" return images\n",
"\n",
" def postprocess(self, output_batch):\n",
" \"\"\"Converts the images to base64 string.\"\"\"\n",
" postprocessed_data = []\n",
" for op in output_batch:\n",
" fp = BytesIO()\n",
" op.save(fp, format='JPEG')\n",
" postprocessed_data.append(base64.b64encode(fp.getvalue()).decode('utf-8'))\n",
" fp.close()\n",
" return postprocessed_data\n"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "6ace1dac0af0"
},
"source": [
"After creating the handler file, you must package the handler as a model archiver (MAR) file. The output file must be named 'model.mar'."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "67707f95d440"
},
"outputs": [],
"source": [
"!torch-model-archiver \\\n",
" -f \\\n",
" --model-name model \\\n",
" --version 1.0 \\\n",
" --handler model_artifacts/handler.py \\\n",
" --export-path model_artifacts"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "ffab030f4bc8"
},
"source": [
"### Upload the model artifacts to Cloud Storage\n",
"\n",
"Create a new folder in your Cloud Storage bucket to hold the model artifacts"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "MzGDU7TWdts_"
},
"outputs": [],
"source": [
"BUCKET_NAME = \"your-bucket-name-unique\" # @param {type:\"string\"}\n",
"BUCKET_URI = f\"gs://{BUCKET_NAME}/\"\n",
"FULL_GCS_PATH = f\"{BUCKET_URI}model_artifacts\""
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "-EcIXiGsCePi"
},
"source": [
"**Only if your bucket doesn't already exist**: Run the following cell to create your Cloud Storage bucket."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "NIq7R4HZCfIc"
},
"outputs": [],
"source": [
"! gsutil mb -l $REGION -p $PROJECT_ID $BUCKET_URI"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "971232e28657"
},
"source": [
"Next, upload the model archive file and your trained Stable Diffusion 1.5 model to the folder on Cloud Storage."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "ef6baf44c808"
},
"outputs": [],
"source": [
"!gsutil cp -r model_artifacts $BUCKET_URI"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "402370ca9396"
},
"source": [
"### Create the Vertex AI model\n",
"\n",
"Once you've uploaded the model artifacts into a Cloud Storage bucket, you can create a new Vertex AI model. This notebook uses the [Vertex AI SDK](https://cloud.google.com/vertex-ai/docs/start/use-vertex-ai-python-sdk) to create the model."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "b6c58a74a0fd"
},
"outputs": [],
"source": [
"PYTORCH_PREDICTION_IMAGE_URI = (\n",
" \"us-docker.pkg.dev/vertex-ai/prediction/pytorch-gpu.1-12:latest\"\n",
")\n",
"APP_NAME = \"my-stable-diffusion\"\n",
"VERSION = 1\n",
"MODEL_DISPLAY_NAME = \"stable_diffusion_1_5-unique\"\n",
"MODEL_DESCRIPTION = \"stable_diffusion_1_5 container\"\n",
"ENDPOINT_DISPLAY_NAME = f\"{APP_NAME}-endpoint\""
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "07c3503a1a2e"
},
"outputs": [],
"source": [
"aiplatform.init(project=PROJECT_ID, location=REGION, staging_bucket=BUCKET_NAME)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "FAPoU8Sm5E6e"
},
"source": [
"<div style=\"background:#e3effe; padding:5px; color:#0000aa\">\n",
"<strong>Note:</strong> The next cell fails if you haven't <a href=\"https://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com\"><u>enabled the Vertex API</u></a>.</div>"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "a776324dd16f"
},
"outputs": [],
"source": [
"model = aiplatform.Model.upload(\n",
" display_name=MODEL_DISPLAY_NAME,\n",
" description=MODEL_DESCRIPTION,\n",
" serving_container_image_uri=PYTORCH_PREDICTION_IMAGE_URI,\n",
" artifact_uri=FULL_GCS_PATH,\n",
")\n",
"\n",
"model.wait()\n",
"\n",
"print(model.display_name)\n",
"print(model.resource_name)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "0fbc3d371574"
},
"source": [
"### Deploy the model to an endpoint\n",
"\n",
"To get online preductions from your Stable Diffusion 2.0 model, you must [deploy it to a Vertex AI endpoint](https://cloud.google.com/vertex-ai/docs/predictions/overview). You can again use the Vertex AI SDK to create the endpoint and deploy your model."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "ab29f0a770cb"
},
"outputs": [],
"source": [
"endpoint = aiplatform.Endpoint.create(display_name=ENDPOINT_DISPLAY_NAME)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "25f703df88c7"
},
"outputs": [],
"source": [
"model.deploy(\n",
" endpoint=endpoint,\n",
" deployed_model_display_name=MODEL_DISPLAY_NAME,\n",
" machine_type=\"n1-standard-8\",\n",
" accelerator_type=\"NVIDIA_TESLA_P100\",\n",
" accelerator_count=1,\n",
" traffic_percentage=100,\n",
" deploy_request_timeout=1200,\n",
" sync=True,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "c9fc560df0a9"
},
"source": [
"The previous cell, which deploys your model to the endpoint, can take a while to complete. If the previous cell times out before returning, your endpoint might still be successfully deployed to an endpoint. Check the [Cloud Console](https://console.cloud.google.com/vertex-ai/endpoints) to verify the results.\n",
"\n",
"You can also extend the time to wait for deployment by changing the `deploy_request_timeout` argument passed to `model.deploy()`."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "88a5304dfdc9"
},
"source": [
"## Get online predictions\n",
"\n",
"Finally, with your Stable Diffusion 1.5 model deployed to a Vertex AI endpoint, you can now get online predictions from it. Using the Vertex AI SDK, you only need a few lines of code to get an inference."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "0d6bc4aa34d6"
},
"outputs": [],
"source": [
"instances = [{\"prompt\": \"An examplePup dog with a baseball jersey.\"}]\n",
"response = endpoint.predict(instances=instances)\n",
"\n",
"with open(\"img5.jpg\", \"wb\") as g:\n",
" g.write(base64.b64decode(response.predictions[0]))"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "65bafefda60c"
},
"outputs": [],
"source": [
"display.Image(\"img5.jpg\")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "TpV-iwP9qw9c"
},
"source": [
"## Cleaning up\n",
"\n",
"To clean up all Google Cloud resources used in this project, you can [delete the Google Cloud\n",
"project](https://cloud.google.com/resource-manager/docs/creating-managing-projects#shutting_down_projects) you used for the tutorial.\n",
"\n",
"Otherwise, you can delete the individual resources you created in this tutorial:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "sx_vKniMq9ZX"
},
"outputs": [],
"source": [
"import os\n",
"\n",
"# Delete endpoint resource\n",
"endpoint.undeploy_all()\n",
"endpoint.delete()\n",
"\n",
"# Delete model resource\n",
"model.delete()\n",
"\n",
"# Delete Cloud Storage objects that were created\n",
"delete_bucket = False\n",
"if delete_bucket or os.getenv(\"IS_TESTING\"):\n",
" ! gsutil -m rm -r $BUCKET_URI"
]
}
],
"metadata": {
"colab": {
"name": "dreambooth_stablediffusion.ipynb",
"toc_visible": true
},
"kernelspec": {
"display_name": "Python 3",
"name": "python3"
}
},
"nbformat": 4,
"nbformat_minor": 0
}
+3 -1
View File
@@ -74,7 +74,9 @@
"source": [
"## Overview\n",
"\n",
"{TODO: Include a paragraph or two explaining what this example demonstrates, who should be interested in it, and what you need to know before you get started.}"
"{TODO: Include a paragraph or two explaining what this example demonstrates, who should be interested in it, and what you need to know before you get started.}\n",
"\n",
"Learn more about [web-doc-title](linkback-to-webdoc-page). {TODO: if more than one primary feature, add tag/linkback for each one}"
]
},
{
+57 -16
View File
@@ -695,7 +695,26 @@ class ObjectiveRule(NotebookRule):
ret = notebook.report_error(ErrorCode.ERROR_OBJECTIVE_MISSING_DESC, "Objective section missing desc")
else:
self.desc = self.desc.lstrip()
sentences = self.desc.split('.')
bracket = False
paren = False
sentences = ""
for _ in range(len(self.desc)):
if self.desc[_] == '[':
bracket = True
continue
elif self.desc[_] == ']':
bracket = False
continue
elif self.desc[_] == '(':
paren = True
elif self.desc[_] == ')':
paren = False
continue
if not paren:
sentences += self.desc[_]
sentences = sentences.split('.')
if len(sentences) > 1:
self.desc = sentences[0] + '.\n'
if self.desc.startswith('In this tutorial, you learn') or self.desc.startswith('In this notebook, you learn'):
@@ -1121,32 +1140,44 @@ def add_index(path: str,
print(f' {tag.strip()}<br/>\n')
print(' </td>')
print(' <td>')
print(f' <b>{title}</b><br/>\n')
print(f' <b>{title}</b>. ')
if args.desc:
desc = replace_cl(desc.replace('`', ''))
print('<br/>')
print(f' {desc}<br/>\n')
print(f' {desc}\n')
if args.steps:
steps = replace_cl(steps.replace('\n', '<br/>').replace('-', '&nbsp;&nbsp;-').replace('**', '').replace('*', '&nbsp;&nbsp;-').replace('`', ''))
print('<br/>' + steps + '<br/>')
if args.linkback and linkbacks:
num = len(tags)
for _ in range(num):
if linkbacks[_].startswith("vertex-ai"):
print(f'<br/> Learn more about <a href="https://cloud.google.com/{linkbacks[_]}" target="_blank">{replace_cl(tags[_])}</a>.\n')
print(f' Learn more about <a href="https://cloud.google.com/{linkbacks[_]}" target="_blank">{replace_cl(tags[_])}</a>.\n')
else:
print(f'<br/> Learn more about <a href="{linkbacks[_]}" target="_blank">{replace_cl(tags[_])}</a>.\n')
print(f' Learn more about <a href="{linkbacks[_]}" target="_blank">{replace_cl(tags[_])}</a>.\n')
if args.steps:
print("<devsite-expandable>\n")
print(' <p class="showalways">Tutorial steps</p>\n')
print(' <ul>\n')
if ":" in steps:
steps = steps.split(':')[1].replace('*', '').replace('-', '').strip().split('\n')
else:
steps = []
for step in steps:
print(f' <li>{replace_cl(step)}</li>\n')
print(' </ul>\n')
print("</devsite-expandable>\n")
print(' </td>')
print(' <td>')
if colab_link:
print(f' <a href="{colab_link}" target="_blank">Colab</a><br/>\n')
print(f' <a href="{colab_link}" target="_blank" class="external" track-type="notebookTutorial" track-name="colabLink">Colab</a><br/>\n')
if git_link:
print(f' <a href="{git_link}" target="_blank">GitHub</a><br/>\n')
print(f' <a href="{git_link}" target="_blank" class="external" track-type="notebookTutorial" track-name="gitHubLink">GitHub</a><br/>\n')
if workbench_link:
print(f' <a href="{workbench_link}" target="_blank">Vertex AI Workbench</a><br/>\n')
print(f' <a href="{workbench_link}" target="_blank" class="external" track-type="notebookTutorial" track-name="workbenchLink">Vertex AI Workbench</a><br/>\n')
print(' </td>')
print(' </tr>\n')
elif args.repo:
@@ -1209,14 +1240,16 @@ def replace_cl(text : str ) -> str:
'Vertex AI Prediction': '{{vertex_prediction_name}}',
'Vertex TensorBoard': '{{vertex_tensorboard_name}}',
'Vertex AI TensorBoard': '{{vertex_tensorboard_name}}',
'TensorBoard': '{{vertex_tensorboard_name}}',
'Tensorboard': '{{vertex_tensorboard_name}}',
'Vertex ML Metadata': '{{vertex_metadata_name}}',
'Vertex Pipelines': '{{vertex_pipelines_name}}',
'Vertex AI Pipelines': '{{vertex_pipelines_name}}',
'Vertex AI Data Labeling': '{{vertex_data_labeling_name}}',
'Vertex AI Experiments': '{{vertex_experiments_name}}',
'Vertex Experiments': '{{vertex_experiments_name}}',
'Vertex AI Matching Engine': '{vertex_matching_engine_name}}',
'Vertex Matching Engine': '{vertex_matching_engine_name}}',
'Vertex AI Matching Engine': '{{vertex_matching_engine_name}}',
'Vertex Matching Engine': '{{vertex_matching_engine_name}}',
'Vertex Model Monitoring': '{{vertex_model_monitoring_name}}',
'Vertex AI Model Monitoring': '{{vertex_model_monitoring_name}}',
'Vertex Feature Store': '{{vertex_featurestore_name}}',
@@ -1234,6 +1267,8 @@ def replace_cl(text : str ) -> str:
'Vertex AI': '{{vertex_ai_name}}',
'Cloud Storage': '{{storage_name}}',
'GCS': '{{storage_name}}',
'GCP': '{{gcp_name}}',
'TensorFlow Enterprise': '{{tf4gcp_name}}',
'TensorFlow': '{{tensorflow_name}}',
}
@@ -1282,9 +1317,14 @@ if args.web:
print('}')
print('</style>')
print('<table>')
print(' <th width="180px">Services</th>')
print(' <th>Description</th>')
print(' <th width="80px">Open in</th>')
print(' <thead>')
print(' <tr>')
print(' <th width="180px">Services</th>')
print(' <th>Description</th>')
print(' <th width="80px">Open in</th>')
print(' </tr>')
print(' </thead>')
print(' <tbody class="list">')
if args.notebook_dir:
if not os.path.isdir(args.notebook_dir):
@@ -1320,6 +1360,7 @@ else:
exit(1)
if args.web:
print(' </tbody>\n')
print('</table>\n')
exit(exit_code)
@@ -23,7 +23,62 @@
- Incorrect examples: "Let's update the field", "We'll update the field", "The user should update the field"
- **Googlers**: Please follow our [branding guidelines](http://goto/cloud-branding).
### Code
## Authoring guidelines
### Focus
Notebooks for official are expected to be narrow focused, which highlight a subset of features of a Vertex AI product/service.
The product/feature is to be highlighted in the Overview section. For example:
```
This tutorial demonstrates using Vertex AI Training to train an XGBoost model using a XGBoost pre-built training container.
```
In the above example, the Vertex AI product/service is `Vertex AI Training` and the feature is `XGBoost pre-built training container`.
### Scope
Notebooks for official are expected to be narrow in scope, without extra extraneous steps. For example, if the notebook is about training, we discourage ending the notebook with deploying the model and doing an online/batch prediction. On the later, we recommend a separate notebook about prediction that uses a pretrained model.
#### Training
Notebooks for training should be constructed as follows:
1. If the training script(s) are small, embed them in the notebook and use %writefile to store them locally.
2. If the training script(s) are large, store them in our public bucket: gs://cloud-samples-data/vertex-ai/dataset-management/script, and use !wgets to retrieve and store the script locally.
3. Train the model using the Vertex AI SDK methods for custom training.
4. Preferrable have the service upload the trained model to the Vertex AI Model Registry.
4. Have the script do an evaluation.
5. Retrieve the evaluation metrics and attach them as an artifact to the corresponding entry in the Model Registry.
6. Optionally, download the model artifacts and test locally -- i.e., make a local prediction request.
#### Evaluation
Notebooks for evaluation should be constructed as follows:
1. Use a pretrained model from a public repository.
2. Upload the pretrained model to the Vertex AI Model Registry.
3. Perform a model evaluation.
4. Review the model evaluation.
4. Attach the model evaluation to the corresponding entry in the Model Registry.
#### Prediction
Notebooks for prediction should be constructed as follows:
1. Use a pretrained model from a public repository.
2. If relevant, attach a serving function to the model artifacts.
3. Upload the pretrained model to the Vertex AI Model Registry.
4. For online:<br/>
A. Deploy the model.<br/>
B. Perform an online prediction.</br>
C. Review the result.
5. For batch:<br/>
A. Perform a batch prediction.</br/>
B. Review the result.
## Code
- Put all your installs and imports in a setup section.
- Save the notebook with the Table of Contents open.
@@ -31,7 +86,7 @@
- Follow the [Google Python Style guide](https://github.com/google/styleguide/blob/gh-pages/pyguide.md) and write readable code.
- Keep cells small (max ~20 lines).
## TensorFlow code style
### TensorFlow code style
Use the highest level API that gets the job done (unless the goal is to demonstrate the low level API). For example, when using Tensorflow:
+4 -7
View File
@@ -29,17 +29,14 @@
/pipelines/google_cloud_pipelines_dataproc_tabular @inardini
/automl/automl_forecasting_bqml_arima_plus_comparison.ipynb @TheMichaelHu
/automl/automl_tabular_on_vertex_pipelines.ipynb @helinwang
/custom/custom_training_tensorboard_profiler.ipynb @itseric
/custom/custom_training_tensorboard_profiler.ipynb @gericdong
/custom/get_started_with_vertex_endpoint_and_shared_vm.ipynb @andrewferlitsch
/workbench/spark/spark_sample_notebook.ipynb @bradmiro
/workbench/spark/spark_ml.ipynb @bradmiro
/model_registry/bqml_vertexai_model_registry.ipynb @soheilazangeneh
/workbench/exploratory_data_analysis/explore_data_in_bigquery_with_workbench.ipynb @alokpattani
/model_evaluation/automl_tabular_classification_model_evaluation.ipynb @soheilazangeneh
/model_evaluation/automl_tabular_regression_model_evaluation.ipynb @soheilazangeneh
/tabular_workflows/tabnet_on_vertex_pipelines.ipynb @sakagarwal
/tabular_workflows/wide_and_deep_on_vertex_pipelines.ipynb @sakagarwal
/tabular_workflows/prophet_on_vertex_pipelines.ipynb @TheMichaelHu
/model_evaluation/custom_tabular_classification_model_evaluation.ipynb @soheilazangeneh
/sdk/SDK_FBProphet_Forecasting_Online.ipynb @brianchunkang
/automl/sdk_automl_forecasting_hierarchical_batch.ipynb @ivanmkc
/prediction/custom_batch_prediction_feature_filter.ipynb @soheilazangeneh
/feature_store/feature_store_streaming_ingestion_sdk.ipynb @soheilazangeneh
/pipelines/Train_tabular_models_with_many_frameworks_and_import_to_Vertex_AI_using_Pipelines @Ark-kun
@@ -170,7 +170,7 @@
},
"outputs": [],
"source": [
"!pip install -U google-cloud-pipeline-components -q"
"!pip install -U google-cloud-pipeline-components==1.0.25 -q"
]
},
{
@@ -333,7 +333,10 @@
},
"outputs": [],
"source": [
"REGION = \"[your-region]\" # @param {type: \"string\"}"
"REGION = \"us-central1\" # @param {type: \"string\"}\n",
"\n",
"if REGION == \"[your-region]\":\n",
" REGION = \"us-central1\""
]
},
{
@@ -379,8 +382,11 @@
"import os\n",
"import sys\n",
"\n",
"# If on Google Cloud Notebook, then don't execute this code\n",
"if not os.path.exists(\"/opt/deeplearning/metadata/env_version\"):\n",
"# If on Vertex AI Workbench, then don't execute this code\n",
"IS_COLAB = \"google.colab\" in sys.modules\n",
"if not os.path.exists(\"/opt/deeplearning/metadata/env_version\") and not os.getenv(\n",
" \"DL_ANACONDA_HOME\"\n",
"):\n",
" if \"google.colab\" in sys.modules:\n",
" from google.colab import auth as google_auth\n",
"\n",
@@ -388,10 +394,9 @@
"\n",
" # If you are running this notebook locally, replace the string below with the\n",
" # path to your service account key and run this cell to authenticate your GCP\n",
" # account. Alternatively, you may edit this notebook to authenticate using\n",
" # gcloud.\n",
" # account.\n",
" elif not os.getenv(\"IS_TESTING\"):\n",
" %env GOOGLE_APPLICATION_CREDENTIALS ''"
" %env GOOGLE_APPLICATION_CREDENTIALS '[your-service-account-key-path]'"
]
},
{
@@ -478,6 +483,77 @@
"! gsutil ls -al $BUCKET_URI"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "44accda192d5"
},
"source": [
"#### Service Account\n",
"\n",
"You use a service account to create Vertex AI Pipeline jobs. If you do not want to use your project's Compute Engine service account, set `SERVICE_ACCOUNT` to another service account ID."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "e0c9c4f84849"
},
"outputs": [],
"source": [
"SERVICE_ACCOUNT = \"[your-service-account]\""
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "604ae09ab6d3"
},
"outputs": [],
"source": [
"if (\n",
" SERVICE_ACCOUNT == \"\"\n",
" or SERVICE_ACCOUNT is None\n",
" or SERVICE_ACCOUNT == \"[your-service-account]\"\n",
"):\n",
" # Get your service account from gcloud\n",
" if not IS_COLAB:\n",
" shell_output = !gcloud auth list 2>/dev/null\n",
" SERVICE_ACCOUNT = shell_output[2].replace(\"*\", \"\").strip()\n",
"\n",
" else: # IS_COLAB:\n",
" shell_output = ! gcloud projects describe $PROJECT_ID\n",
" project_number = shell_output[-1].split(\":\")[1].strip().replace(\"'\", \"\")\n",
" SERVICE_ACCOUNT = f\"{project_number}-compute@developer.gserviceaccount.com\"\n",
"\n",
" print(\"Service Account:\", SERVICE_ACCOUNT)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "d1ecb60964d5"
},
"source": [
"#### Set service account access for Vertex AI Pipelines\n",
"Run the following commands to grant your service account access to read and write pipeline artifacts in the bucket that you created in the previous step. You only need to run this step once per service account."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "a592f0a380c2"
},
"outputs": [],
"source": [
"! gsutil iam ch serviceAccount:{SERVICE_ACCOUNT}:roles/storage.objectCreator $BUCKET_URI\n",
"\n",
"! gsutil iam ch serviceAccount:{SERVICE_ACCOUNT}:roles/storage.objectViewer $BUCKET_URI"
]
},
{
"cell_type": "markdown",
"metadata": {
@@ -836,6 +912,7 @@
"\n",
"job.run()\n",
"\n",
"\n",
"pipeline_task_details = job.gca_resource.job_detail.task_details\n",
"\n",
"if export_additional_model_without_custom_ops:\n",
@@ -877,6 +954,7 @@
"stage_1_tuner_task = get_task_detail(\n",
" pipeline_task_details, \"automl-tabular-stage-1-tuner\"\n",
")\n",
"\n",
"stage_1_tuning_result_artifact_uri = (\n",
" stage_1_tuner_task.outputs[\"tuning_result_output\"].artifacts[0].uri\n",
")"
@@ -64,7 +64,7 @@
"\n",
"This tutorial demonstrates how to use AutoML in production. This tutorial covers get started with AutoML training.\n",
"\n",
"Learn more about [Vertex AI for AutoML](https://cloud.google.com/vertex-ai/docs/start/automl-users)."
"Learn more about [AutoML training](https://cloud.google.com/vertex-ai/docs/training-overview)."
]
},
{
@@ -872,6 +872,18 @@
"The execution of the training pipeline will take upto > 30 minutes."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "3eaba926cdfa"
},
"outputs": [],
"source": [
"if os.getenv(\"IS_TESTING\"):\n",
" sys.exit(0)"
]
},
{
"cell_type": "code",
"execution_count": null,
@@ -833,6 +833,7 @@
]
},
{
"attachments": {},
"cell_type": "markdown",
"metadata": {
"id": "99b7a9287ba6"
@@ -840,7 +841,7 @@
"source": [
"For AutoML models, manual scaling can be adjusted by setting both min and max nodes i.e., `starting_replica_count` and `max_replica_count` as the same value(in this example, set to 1). The node count can be increased or decreased as required by load.\n",
" \n",
"`batch_predict` can export predictions either to BigQuery or GCS. The BigQuery options are commented out below and the predictions will be exported to the BUCKET_URI."
"`batch_predict` can export predictions either to BigQuery or GCS. This example exports to BigQuery."
]
},
{
@@ -60,7 +60,7 @@
"\n",
"This tutorial walks you through building a custom container to serve a scikit-learn model on Vertex AI. You use the FastAPI Python web server framework to create a prediction and health endpoint. You also incorporate a pre-processor from training pipeline into your online serving application.\n",
"\n",
"Learn more about [Vertex AI Training](https://cloud.google.com/vertex-ai/docs/training/custom-training) and [Vertex AI Prediction](https://cloud.google.com/vertex-ai/docs/predictions/get-predictions)."
"Learn more about [Custom training](https://cloud.google.com/vertex-ai/docs/training/custom-training) and [Vertex AI Prediction](https://cloud.google.com/vertex-ai/docs/predictions/get-predictions)."
]
},
{
@@ -64,7 +64,7 @@
"\n",
"Learn more about serving an FBProphet model from this [article on testdriven.io: Deploying and Hosting a Machine Learning Model with FastAPI and Heroku](https://testdriven.io/blog/fastapi-machine-learning/).\n",
"\n",
"Learn more about [Vertex AI Training](https://cloud.google.com/vertex-ai/docs/training/custom-training) and [Vertex AI Prediction](https://cloud.google.com/vertex-ai/docs/predictions/get-predictions).\n"
"Learn more about [Custom training](https://cloud.google.com/vertex-ai/docs/training/custom-training) and [Vertex AI Prediction](https://cloud.google.com/vertex-ai/docs/predictions/get-predictions).\n"
]
},
{
@@ -147,8 +147,7 @@
"# Install the packages\n",
"! pip3 install --upgrade google-cloud-aiplatform \\\n",
" google-cloud-storage \\\n",
" google-cloud-bigquery \\\n",
" pyarrow"
" 'google-cloud-bigquery[pandas]'"
]
},
{
@@ -632,12 +631,12 @@
" df_train_x, df_train_y = df_train, df_train.pop(LABEL_COLUMN)\n",
" df_validation_x, df_validation_y = df_validation, df_validation.pop(LABEL_COLUMN)\n",
"\n",
" y_train = np.asarray(df_train_y).astype(\"float32\")\n",
" y_validation = np.asarray(df_validation_y).astype(\"float32\")\n",
" y_train = tf.convert_to_tensor(np.asarray(df_train_y).astype(\"float32\"))\n",
" y_validation = tf.convert_to_tensor(np.asarray(df_validation_y).astype(\"float32\"))\n",
"\n",
" # Convert to numpy representation\n",
" x_train = np.asarray(df_train_x) \n",
" x_test = np.asarray(df_validation_x)\n",
" x_train = tf.convert_to_tensor(np.asarray(df_train_x).astype(\"float32\"))\n",
" x_test = tf.convert_to_tensor(np.asarray(df_validation_x).astype(\"float32\"))\n",
"\n",
" # Convert to one-hot representation\n",
" num_species = len(df_train_y.unique())\n",
@@ -735,7 +734,7 @@
" display_name=JOB_NAME,\n",
" script_path=\"task.py\",\n",
" container_uri=\"us-docker.pkg.dev/vertex-ai/training/tf-cpu.2-8:latest\",\n",
" requirements=[\"google-cloud-bigquery>=2.20.0\", \"db-dtypes\"],\n",
" requirements=[\"google-cloud-bigquery[pandas]\", \"protobuf<3.20.0\"],\n",
" model_serving_container_image_uri=\"us-docker.pkg.dev/vertex-ai/prediction/tf2-cpu.2-8:latest\",\n",
")\n",
"\n",
File diff suppressed because it is too large Load Diff
@@ -64,7 +64,7 @@
"\n",
"This tutorial demonstrates how to use the Vertex AI SDK for Python to train and deploy a custom image classification model for batch prediction.\n",
"\n",
"Learn more about [Vertex AI Training](https://cloud.google.com/vertex-ai/docs/training/custom-training) and [Vertex AI Batch Prediction](https://cloud.google.com/vertex-ai/docs/tabular-data/classification-regression/get-batch-predictions)."
"Learn more about [Custom training](https://cloud.google.com/vertex-ai/docs/training/custom-training) and [Vertex AI Batch Prediction](https://cloud.google.com/vertex-ai/docs/tabular-data/classification-regression/get-batch-predictions)."
]
},
{
@@ -64,7 +64,7 @@
"\n",
"This tutorial demonstrates how to use the Vertex AI SDK for Python to train and deploy a custom image classification model for online prediction.\n",
"\n",
"Learn more about [Vertex AI Training](https://cloud.google.com/vertex-ai/docs/training/custom-training) and [Vertex AI Prediction](https://cloud.google.com/vertex-ai/docs/predictions/get-predictions)."
"Learn more about [Custom training](https://cloud.google.com/vertex-ai/docs/training/custom-training) and [Vertex AI Prediction](https://cloud.google.com/vertex-ai/docs/predictions/get-predictions)."
]
},
{
@@ -352,7 +352,10 @@
},
"outputs": [],
"source": [
"EMAIL = \"[your-email-address]\" # @param {type: \"string\"}"
"EMAIL = \"[your-email-address]\" # @param {type: \"string\"}\n",
"\n",
"if os.getenv(\"IS_TESTING\"):\n",
" EMAIL = \"noreply@google.com\""
]
},
{
File diff suppressed because it is too large Load Diff
@@ -64,7 +64,7 @@
"\n",
"This tutorial demonstrates how to use the Vertex AI SDK to create tabular binary classification models and do batch prediction with explanation using a Google Cloud [AutoML](https://cloud.google.com/vertex-ai/docs/start/automl-users) model.\n",
"\n",
"Learn more about [AutoML Tabular](https://cloud.google.com/vertex-ai/docs/tabular-data/overview) and [Vertex Explainable AI](https://cloud.google.com/vertex-ai/docs/explainable-ai/overview)."
"Learn more about [Classification for tabular data](https://cloud.google.com/vertex-ai/docs/tabular-data/classification-regression/overview). Learn more about [Vertex Explainable AI](https://cloud.google.com/vertex-ai/docs/explainable-ai/overview)."
]
},
{
@@ -64,7 +64,7 @@
"\n",
"This tutorial demonstrates how to use the Vertex AI SDK to create tabular classification models and do online prediction with explanation using a Google Cloud [AutoML](https://cloud.google.com/vertex-ai/docs/start/automl-users) model.\n",
"\n",
"Learn more about [AutoML Tabular](https://cloud.google.com/vertex-ai/docs/tabular-data/overview) and [Vertex Explainable AI](https://cloud.google.com/vertex-ai/docs/explainable-ai/overview)."
"Learn more about [Classification for tabular data](https://cloud.google.com/vertex-ai/docs/tabular-data/classification-regression/overview) and [Vertex Explainable AI](https://cloud.google.com/vertex-ai/docs/explainable-ai/overview)."
]
},
{
File diff suppressed because one or more lines are too long
File diff suppressed because it is too large Load Diff
@@ -292,8 +292,7 @@
},
"outputs": [],
"source": [
"# VPC_NETWORK = \"[your-vpc-network-name]\" # @param {type:\"string\"}\n",
"VPC_NETWORK = \"matching-engine-test\" # @param {type:\"string\"}\n",
"VPC_NETWORK = \"[your-vpc-network-name]\" # @param {type:\"string\"}\n",
"\n",
"PEERING_RANGE_NAME = \"ann-haystack-range\""
]
@@ -544,7 +543,7 @@
" json.dumps(\n",
" {\n",
" \"id\": str(index),\n",
" \"embedding\": [str(value) for value in train[index]],\n",
" \"embedding\": [str(value) for value in embedding],\n",
" \"restricts\": [\n",
" {\n",
" \"namespace\": \"class\",\n",
@@ -646,7 +645,7 @@
"outputs": [],
"source": [
"tree_ah_index = aiplatform.MatchingEngineIndex.create_tree_ah_index(\n",
" display_name=DISPLAY_NAME_BRUTE_FORCE,\n",
" display_name=DISPLAY_NAME,\n",
" contents_delta_uri=EMBEDDINGS_INITIAL_URI,\n",
" dimensions=DIMENSIONS,\n",
" approximate_neighbors_count=150,\n",
@@ -712,7 +711,7 @@
"outputs": [],
"source": [
"brute_force_index = aiplatform.MatchingEngineIndex.create_brute_force_index(\n",
" display_name=DISPLAY_NAME,\n",
" display_name=DISPLAY_NAME_BRUTE_FORCE,\n",
" contents_delta_uri=EMBEDDINGS_INITIAL_URI,\n",
" dimensions=DIMENSIONS,\n",
" distance_measure_type=\"DOT_PRODUCT_DISTANCE\",\n",
File diff suppressed because it is too large Load Diff
@@ -64,7 +64,7 @@
"\n",
"This tutorial demonstrates how to use the Vertex AI SDK for Python to train and deploy an AutoML image classification model.\n",
"\n",
"Learn more about [Migrate to Vertex AI](https://cloud.google.com/vertex-ai/docs/start/migrating-to-vertex-ai) and [AutoML Image](https://cloud.google.com/vertex-ai/docs/tutorials/image-recognition-automl/training)."
"Learn more about [Migrate to Vertex AI](https://cloud.google.com/vertex-ai/docs/start/migrating-to-vertex-ai) and [Classification for image data](https://cloud.google.com/vertex-ai/docs/training-overview#classification_for_images)."
]
},
{
@@ -64,7 +64,7 @@
"\n",
"This tutorial demonstrates how to use the Vertex AI SDK for Python to train and deploy a custom tabular classification scikit-learn model for batch prediction.\n",
"\n",
"Learn more about [Migrate to Vertex AI](https://cloud.google.com/vertex-ai/docs/start/migrating-to-vertex-ai) and [Vertex AI Training](https://cloud.google.com/vertex-ai/docs/training/custom-training)."
"Learn more about [Migrate to Vertex AI](https://cloud.google.com/vertex-ai/docs/start/migrating-to-vertex-ai) and [Custom training](https://cloud.google.com/vertex-ai/docs/training/custom-training)."
]
},
{
@@ -65,7 +65,7 @@
"\n",
"This tutorial demonstrates how to use the Vertex AI SDK for Python to hyperparamer tune a custom tabular classification TemsorFlow model.\n",
"\n",
"Learn more about [Migrate to Vertex AI](https://cloud.google.com/vertex-ai/docs/training/hyperparameter-tuning-overview) and [Vertex AI Training](https://cloud.google.com/vertex-ai/docs/training/custom-training)."
"Learn more about [Migrate to Vertex AI](https://cloud.google.com/vertex-ai/docs/training/hyperparameter-tuning-overview) and [Custom training](https://cloud.google.com/vertex-ai/docs/training/custom-training)."
]
},
{
@@ -29,7 +29,7 @@
"id": "title:migration,new"
},
"source": [
"# Vertex AI: Vertex AI Migration: AutoML Video Classificaton\n",
"# Vertex AI: Vertex AI Migration: AutoML Video Classification\n",
"\n",
"<table align=\"left\">\n",
"\n",
@@ -64,7 +64,7 @@
"\n",
"This tutorial demonstrates how to use the Vertex AI SDK for Python to train a AutoML video classification model and do a batch prediction.\n",
"\n",
"Learn more about [Migrate to Vertex AI](https://cloud.google.com/vertex-ai/docs/start/migrating-to-vertex-ai) and [AutoML Video](https://cloud.google.com/vertex-ai/docs/tutorials/video-classification-automl/training)."
"Learn more about [Migrate to Vertex AI](https://cloud.google.com/vertex-ai/docs/start/migrating-to-vertex-ai) and [Classification for video data](https://cloud.google.com/vertex-ai/docs/training-overview#classification_for_videos)."
]
},
{
@@ -568,7 +568,7 @@
},
"outputs": [],
"source": [
"IMPORT_FILE = \"gs://automl-video-demo-data/hmdb_split1_5classes_train_inf.csv\""
"IMPORT_FILE = \"gs://automl-video-demo-data/hmdb_split1_train_40_mp4_gs.csv\""
]
},
{
@@ -604,6 +604,19 @@
"! gsutil cat $FILE | head"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "f483d8d80f64"
},
"outputs": [],
"source": [
"import datetime\n",
"\n",
"start = datetime.datetime.now()"
]
},
{
"cell_type": "markdown",
"metadata": {
@@ -757,7 +770,19 @@
"\n",
"The `run` method when completed returns the `Model` resource.\n",
"\n",
"The execution of the training pipeline will take upto 20 minutes."
"The execution of the training pipeline may take over 24 hrs."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "8f135100f5d9"
},
"outputs": [],
"source": [
"if os.getenv(\"IS_TESTING\"):\n",
" sys.exit(0)"
]
},
{
@@ -806,6 +831,18 @@
" INFO:google.cloud.aiplatform.training_jobs:Model available at projects/759209241365/locations/us-central1/models/1899701006099283968"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "48ca7f0980e0"
},
"outputs": [],
"source": [
"end = datetime.datetime.now()\n",
"end - start"
]
},
{
"cell_type": "markdown",
"metadata": {
@@ -64,7 +64,7 @@
"\n",
"This tutorial demonstrates how to use the Vertex AI SDK for Python to train a AutoML video object tracking model and do a batch prediction.\n",
"\n",
"Learn more about [Migrate to Vertex AI](https://cloud.google.com/vertex-ai/docs/start/migrating-to-vertex-ai) and [AutoML Video](https://cloud.google.com/video-intelligence/automl/object-tracking/docs/index-object-tracking)."
"Learn more about [Migrate to Vertex AI](https://cloud.google.com/vertex-ai/docs/start/migrating-to-vertex-ai) and [Object tracking for video data](https://cloud.google.com/vertex-ai/docs/training-overview#object_tracking_for_videos)."
]
},
{
@@ -64,7 +64,7 @@
"\n",
"This tutorial demonstrates how to use the Vertex AI SDK for Python to train using a pre-built container and deploy a custom image classification model for online and batch prediction.\n",
"\n",
"Learn more about [Migrate to Vertex AI](https://cloud.google.com/vertex-ai/docs/start/migrating-to-vertex-ai) and [Vertex AI Training](https://cloud.google.com/vertex-ai/docs/training/custom-training)."
"Learn more about [Migrate to Vertex AI](https://cloud.google.com/vertex-ai/docs/start/migrating-to-vertex-ai) and [Custom training](https://cloud.google.com/vertex-ai/docs/training/custom-training)."
]
},
{
@@ -63,7 +63,7 @@
"\n",
"This notebook demonstrates training a custom image classification model using Tensorflow and Vertex AI SDK by creating a custom training container. Additionally, the notebooks also deploys the trained model to Vertex AI and predictions are generated from it.\n",
"\n",
"Learn more about [Migrate to Vertex AI](https://cloud.google.com/vertex-ai/docs/start/migrating-to-vertex-ai) and [Vertex AI Training](https://cloud.google.com/vertex-ai/docs/training/custom-training)."
"Learn more about [Migrate to Vertex AI](https://cloud.google.com/vertex-ai/docs/start/migrating-to-vertex-ai) and [Custom training](https://cloud.google.com/vertex-ai/docs/training/custom-training)."
]
},
{
@@ -64,7 +64,7 @@
"\n",
"This tutorial demonstrates how to use the Vertex AI SDK to create tabular binary classification models and do online prediction using a Google Cloud [AutoML](https://cloud.google.com/vertex-ai/docs/start/automl-users) model.\n",
"\n",
"Learn more about [Migrate to Vertex AI](https://cloud.google.com/vertex-ai/docs/start/migrating-to-vertex-ai) and [AutoML Tabular](https://cloud.google.com/vertex-ai/docs/start/automl-users#tables)."
"Learn more about [Migrate to Vertex AI](https://cloud.google.com/vertex-ai/docs/start/migrating-to-vertex-ai). Learn more about [Classification for tabular data](https://cloud.google.com/vertex-ai/docs/tabular-data/classification-regression/overview)."
]
},
{
@@ -65,7 +65,7 @@
"\n",
"This tutorial demonstrates how to use the Vertex AI SDK for Python to train and deploy an AutoML object detection model.\n",
"\n",
"Learn more about [Migrate to Vertex AI](https://cloud.google.com/vertex-ai/docs/start/migrating-to-vertex-ai) and [AutoML Image](https://cloud.google.com/vertex-ai/docs/tutorials/image-recognition-automl/training)."
"Learn more about [Migrate to Vertex AI](https://cloud.google.com/vertex-ai/docs/start/migrating-to-vertex-ai) and [Object detection for image data](https://cloud.google.com/vertex-ai/docs/training-overview#object_detection_for_images)."
]
},
{
@@ -61,11 +61,11 @@
"source": [
"## Overview\n",
"\n",
"This notebook demonstrates how to create an AutoML Video Classification Model, with a Vertex AI video dataset, and how to serve the model for batch prediction. It requires you provide a bucket where the dataset will be stored.\n",
"This notebook demonstrates how to create an AutoML Text Classification Model, with a Vertex AI text dataset, and how to serve the model for batch prediction. It requires you provide a bucket where the dataset will be stored.\n",
"\n",
"Note: you may incur charges for training, prediction, storage or usage of other GCP products in connection with testing this SDK.\n",
"\n",
"Learn more about [Migrate to Vertex AI](https://cloud.google.com/vertex-ai/docs/start/migrating-to-vertex-ai) and [AutoML Text](https://cloud.google.com/vertex-ai/docs/text-data/classification/prepare-data)."
"Learn more about [Migrate to Vertex AI](https://cloud.google.com/vertex-ai/docs/start/migrating-to-vertex-ai) and [Classification for text data](https://cloud.google.com/vertex-ai/docs/training-overview#classification_for_text)."
]
},
{
@@ -76,7 +76,7 @@
"source": [
"### Objective\n",
"\n",
"The objective of this notebook is to build a AutoML Video Classification Model. The following steps have been followed:\n",
"The objective of this notebook is to build a AutoML Text Classification Model. The following steps have been followed:\n",
"This tutorial uses the following Google Cloud ML services :\n",
"\n",
"* Vertex AI Dataset resource\n",
@@ -87,11 +87,11 @@
"The steps performed include the following:\n",
"\n",
"* Set your task name, and GCS prefix\n",
"* Copy AutoML video demo train data for creating managed dataset\n",
"* Copy AutoML text demo train data for creating managed dataset\n",
"* Create a dataset on Vertex AI.\n",
"* Configure a training job\n",
"* Launch a training job and create a model on Vertex AI\n",
"* Copy AutoML Video Demo Prediction Data for creating batch prediction job\n",
"* Copy AutoML Text Demo Prediction Data for creating batch prediction job\n",
"* Perform batch prediction job on the model"
]
},
@@ -65,7 +65,7 @@
"\n",
"Note: you may incur charges for training, prediction, storage or usage of other Google Cloud products in connection with testing this SDK.\n",
"\n",
"Learn more about [Migrate to Vertex AI](https://cloud.google.com/vertex-ai/docs/start/migrating-to-vertex-ai) and [AutoML Text](https://cloud.google.com/vertex-ai/docs/text-data/entity-extraction/prepare-data)."
"Learn more about [Migrate to Vertex AI](https://cloud.google.com/vertex-ai/docs/start/migrating-to-vertex-ai) and [Entity extraction for text data](https://cloud.google.com/vertex-ai/docs/training-overview#entity_extraction_for_text).\n"
]
},
{
@@ -76,7 +76,7 @@
"source": [
"### Objective\n",
"\n",
"The objective of this notebook is to build a AutoML Text Entity Extraction Model. The following steps have been followed:\n",
"The objective of this notebook is to build a AutoML Text Entity Extraction model. The following steps have been followed:\n",
"This tutorial uses the following Google Cloud ML services :\n",
"\n",
"* Vertex AI Dataset resource\n",
@@ -66,7 +66,7 @@
"\n",
"Note: you may incur charges for training, prediction, storage or usage of other Google Cloud products in connection with testing this SDK.\n",
"\n",
"Learn more about [Migrate to Vertex AI](https://cloud.google.com/vertex-ai/docs/start/migrating-to-vertex-ai) and [AutoML Text](https://cloud.google.com/vertex-ai/docs/text-data/sentiment-analysis/prepare-data)."
"Learn more about [Migrate to Vertex AI](https://cloud.google.com/vertex-ai/docs/start/migrating-to-vertex-ai) and [Sentiment analysis for text data](https://cloud.google.com/vertex-ai/docs/training-overview#sentiment_analysis_for_text)."
]
},
{
@@ -64,7 +64,7 @@
"\n",
"This tutorial demonstrates how to use the Vertex AI SDK for Python to train and deploy a custom tabular classification XGBoost model for batch prediction.\n",
"\n",
"Learn more about [Migrate to Vertex AI](https://cloud.google.com/vertex-ai/docs/start/migrating-to-vertex-ai) and [Vertex AI Training](https://cloud.google.com/vertex-ai/docs/training/custom-training)."
"Learn more about [Migrate to Vertex AI](https://cloud.google.com/vertex-ai/docs/start/migrating-to-vertex-ai) and [Custom training](https://cloud.google.com/vertex-ai/docs/training/custom-training)."
]
},
{
@@ -63,7 +63,7 @@
"\n",
"This notebook demonstrates how to track metrics and parameters for Vertex AI custom training jobs, and how to perform detailed analysis using this data.\n",
"\n",
"Learn more about [Vertex ML Metadata](https://cloud.google.com/vertex-ai/docs/ml-metadata) and [Vertex AI Training](https://cloud.google.com/vertex-ai/docs/training/custom-training)."
"Learn more about [Vertex ML Metadata](https://cloud.google.com/vertex-ai/docs/ml-metadata) and [Custom training](https://cloud.google.com/vertex-ai/docs/training/custom-training)."
]
},
{
@@ -140,7 +140,7 @@
"outputs": [],
"source": [
"! pip install --upgrade -q google-cloud-aiplatform \\\n",
" tensorflow==2.8 -q"
" tensorflow -q"
]
},
{
@@ -63,7 +63,7 @@
"\n",
"This notebook demonstrates how to use the Vertex AI classification model evaluation component to evaluate an AutoML Tabular classification model. Model evaluation helps determine your model's performance based on the evaluation metrics and improve the model whenever necessary. \n",
"\n",
"Learn more about [Vertex AI Model Evaluation](https://cloud.google.com/vertex-ai/docs/evaluation/introduction) and [AutoML Tabular](https://cloud.google.com/vertex-ai/docs/start/automl-users#tables)."
"Learn more about [Vertex AI Model Evaluation](https://cloud.google.com/vertex-ai/docs/evaluation/introduction). Learn more about [Classification for tabular data](https://cloud.google.com/vertex-ai/docs/tabular-data/classification-regression/overview)."
]
},
{
@@ -63,7 +63,7 @@
"\n",
"This notebook demonstrates how to use Vertex AI regression model evaluation component to evaluate an AutoML Tabular regression model. Model evaluation helps you determine your model performance based on the evaluation metrics and improve the model if necessary. \n",
"\n",
"Learn more about [Vertex AI Model Evaluation](https://cloud.google.com/vertex-ai/docs/evaluation/introduction) and [AutoML Tabular](https://cloud.google.com/vertex-ai/docs/start/automl-users#tables)."
"Learn more about [Vertex AI Model Evaluation](https://cloud.google.com/vertex-ai/docs/evaluation/introduction). Learn more about [Regression for tabular data](https://cloud.google.com/vertex-ai/docs/tabular-data/classification-regression/overview)."
]
},
{
@@ -92,7 +92,7 @@
"- Run the `AutoMLTabularTrainingJob` which returns a model\n",
"- Import a pre-trained `AutoML model resource` into the pipeline\n",
"- Run a `batch prediction` job in the pipeline\n",
"- Evaulate the AutoML model using the `regression evaluation component`\n",
"- Evaluate the AutoML model using the `regression evaluation component`\n",
"- Import the Regression Metrics to the AutoML model resource"
]
},
@@ -63,7 +63,7 @@
"\n",
"This notebook demonstrates how to use the Vertex AI classification model evaluation component to evaluate an AutoML video classification model. Model evaluation helps you determine your model performance based on the evaluation metrics and improve the model if necessary. \n",
"\n",
"Learn more about [Vertex AI Model Evaluation](https://cloud.google.com/vertex-ai/docs/evaluation/introduction) and [AutoML Video](https://cloud.google.com/vertex-ai/docs/video-data/classification/prepare-data)."
"Learn more about [Vertex AI Model Evaluation](https://cloud.google.com/vertex-ai/docs/evaluation/introduction) and [Classification for video data](https://cloud.google.com/vertex-ai/docs/training-overview#classification_for_videos)."
]
},
{
@@ -92,7 +92,7 @@
"- Train a Automl Video Classification model on the `Vertex AI Dataset` resource.\n",
"- Import the trained `AutoML Vertex AI Model resource` into the pipeline.\n",
"- Run a batch prediction job inside the pipeline.\n",
"- Evaulate the AutoML model using the classification evaluation component.\n",
"- Evaluate the AutoML model using the classification evaluation component.\n",
"- Import the classification metrics to the AutoML Vertex AI Model resource."
]
},
@@ -792,7 +792,19 @@
"\n",
"The `run` method when completed returns the `Model` resource.\n",
"\n",
"The execution of the training pipeline can take over 3 hours to complete."
"The execution of the training pipeline can take over 24 hours to complete."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "3eaba926cdfa"
},
"outputs": [],
"source": [
"if os.getenv(\"IS_TESTING\"):\n",
" sys.exit(0)"
]
},
{
@@ -63,7 +63,7 @@
"\n",
"This notebook demonstrates how to use the Vertex AI regression model evaluation component to evaluate a custom regression model. Model evaluation helps you determine your model performance based on the evaluation metrics and improve the model if necessary. \n",
"\n",
"Learn more about [Vertex AI Model Evaluation](https://cloud.google.com/vertex-ai/docs/evaluation/introduction) and [Vertex AI Training](https://cloud.google.com/vertex-ai/docs/training/custom-training)."
"Learn more about [Vertex AI Model Evaluation](https://cloud.google.com/vertex-ai/docs/evaluation/introduction) and [Custom training](https://cloud.google.com/vertex-ai/docs/training/custom-training)."
]
},
{
@@ -93,7 +93,7 @@
"- Upload the model as a Vertex AI Model resource.\n",
"- Import a pre-trained `Vertex AI model resource` into the pipeline.\n",
"- Run a `batch prediction` job in the pipeline.\n",
"- Evaulate the model using the `regression evaluation component`.\n",
"- Evaluate the model using the `regression evaluation component`.\n",
"- Import the Regression Metrics to the Vertex AI model resource."
]
},
@@ -0,0 +1,865 @@
{
"cells": [
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "ur8xi4C7S06n"
},
"outputs": [],
"source": [
"# Copyright 2023 Google LLC\n",
"#\n",
"# Licensed under the Apache License, Version 2.0 (the \"License\");\n",
"# you may not use this file except in compliance with the License.\n",
"# You may obtain a copy of the License at\n",
"#\n",
"# https://www.apache.org/licenses/LICENSE-2.0\n",
"#\n",
"# Unless required by applicable law or agreed to in writing, software\n",
"# distributed under the License is distributed on an \"AS IS\" BASIS,\n",
"# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.\n",
"# See the License for the specific language governing permissions and\n",
"# limitations under the License."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "JAPoU8Sm5E6e"
},
"source": [
"# Get started with importing a custom model evaluation to the Vertex AI Model Registry\n",
"\n",
"<table align=\"left\">\n",
"\n",
" <td>\n",
" <a href=\"https://colab.research.google.com/github/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/official/model_evaluation/get_started_with_custom_model_evaluation_import.ipynb\">\n",
" <img src=\"https://cloud.google.com/ml-engine/images/colab-logo-32px.png\" alt=\"Colab logo\"> Run in Colab\n",
" </a>\n",
" </td>\n",
" <td>\n",
" <a href=\"https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/official/model_evaluation/get_started_with_custom_model_evaluation_import.ipynb\">\n",
" <img src=\"https://cloud.google.com/ml-engine/images/github-logo-32px.png\" alt=\"GitHub logo\">\n",
" View on GitHub\n",
" </a>\n",
" </td>\n",
" <td>\n",
" <a href=\"https://console.cloud.google.com/vertex-ai/workbench/deploy-notebook?download_url=https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/main/notebooks/official/model_evaluation/get_started_with_custom_model_evaluation_import.ipynb\">\n",
" <img src=\"https://lh3.googleusercontent.com/UiNooY4LUgW_oTvpsNhPpQzsstV5W8F7rYgxgGBD85cWJoLmrOzhVs_ksK_vgx40SHs7jCqkTkCk=e14-rj-sc0xffffff-h130-w32\" alt=\"Vertex AI logo\">\n",
" Open in Vertex AI Workbench\n",
" </a>\n",
" </td> \n",
"</table>"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "24743cf4a1e1"
},
"source": [
"**_NOTE_**: This notebook has been tested in the following environment:\n",
"\n",
"* Python version = 3.9"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "tvgnzT1CKxrO"
},
"source": [
"## Overview\n",
"\n",
"This tutorial shows how to use Vertex AI Model Evaluation to import a custom model evaluation to an existing Vertex AI Model Registry entry."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "d975e698c9a4"
},
"source": [
"### Objective\n",
"\n",
"In this tutorial, you learn how to construct and upload a custom model evaluation, and upload the custom model evaluation to a Model resource entry in Vertex AI Model Registry.\n",
"\n",
"This tutorial uses the following Google Cloud ML services and resources:\n",
"\n",
"- Vertex AI Model Evaluation\n",
"- Vertex AI Model Registry\n",
"\n",
"The steps performed include:\n",
"\n",
"- Import a pretrained (blessed) model to the Vertex AI Model Registry.\n",
"- Construct a custom model evaluation.\n",
"- Import the model evaluation metrics to the corresponding model in the Vertex AI Model Registry.\n",
"- List the model evaluation for the corresponding model in the Vertex AI Model Registry.\n",
"- Construct a second custom model evaluation.\n",
"- Import the second model evaluation metrics to the corresponding model in the Vertex AI Model Registry.\n",
"- List the second model evaluation for the corresponding model in the Vertex AI Model Registry.\n",
"\n",
"Learn more about [Model Evaluation in Vertex AI](https://cloud.google.com/vertex-ai/docs/evaluation/introduction)."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "08d289fa873f"
},
"source": [
"### Model\n",
"\n",
"This tutorial uses a pre-trained image classification model from TensorFlow Hub, which is trained on ImageNet dataset.\n",
"\n",
"Learn more about [ResNet V2 pretained model](https://tfhub.dev/google/imagenet/resnet_v2_101/classification/5). "
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "aed92deeb4a0"
},
"source": [
"### Costs \n",
"\n",
"This tutorial uses billable components of Google Cloud:\n",
"\n",
"* Vertex AI\n",
"* Cloud Storage\n",
"\n",
"Learn about [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing),\n",
"and [Cloud Storage pricing](https://cloud.google.com/storage/pricing), \n",
"and use the [Pricing Calculator](https://cloud.google.com/products/calculator/)\n",
"to generate a cost estimate based on your projected usage."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "i7EUnXsZhAGF"
},
"source": [
"## Installation\n",
"\n",
"Install the following packages required to execute this notebook. "
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "2b4ef9b72d43"
},
"outputs": [],
"source": [
"# Install the packages\n",
"USER=''\n",
"! pip3 install {USER} --upgrade google-cloud-aiplatform \\\n",
" tensorflow==2.5 \\\n",
" tensorflow-hub"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "58707a750154"
},
"source": [
"### Colab only: Uncomment the following cell to restart the kernel."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "f200f10a1da3"
},
"outputs": [],
"source": [
"# Automatically restart kernel after installs so that your environment can access the new packages\n",
"# import IPython\n",
"\n",
"# app = IPython.Application.instance()\n",
"# app.kernel.do_shutdown(True)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "BF1j6f9HApxa"
},
"source": [
"## Before you begin\n",
"\n",
"### Set up your Google Cloud project\n",
"\n",
"**The following steps are required, regardless of your notebook environment.**\n",
"\n",
"1. [Select or create a Google Cloud project](https://console.cloud.google.com/cloud-resource-manager). When you first create an account, you get a $300 free credit towards your compute/storage costs.\n",
"\n",
"2. [Make sure that billing is enabled for your project](https://cloud.google.com/billing/docs/how-to/modify-project).\n",
"\n",
"3. [Enable the Vertex AI API]\n",
"\n",
"4. If you are running this notebook locally, you need to install the [Cloud SDK](https://cloud.google.com/sdk)."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "WReHDGG5g0XY"
},
"source": [
"#### Set your project ID\n",
"\n",
"**If you don't know your project ID**, try the following:\n",
"* Run `gcloud config list`.\n",
"* Run `gcloud projects list`.\n",
"* See the support page: [Locate the project ID](https://support.google.com/googleapi/answer/7014113)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "oM1iC_MfAts1"
},
"outputs": [],
"source": [
"PROJECT_ID = \"[your-project-id]\" # @param {type:\"string\"}\n",
"\n",
"# Set the project id\n",
"! gcloud config set project {PROJECT_ID}"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "region"
},
"source": [
"#### Region\n",
"\n",
"You can also change the `REGION` variable used by Vertex AI. Learn more about [Vertex AI regions](https://cloud.google.com/vertex-ai/docs/general/locations)."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "region"
},
"outputs": [],
"source": [
"REGION = \"us-central1\" # @param {type: \"string\"}"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "sBCra4QMA2wR"
},
"source": [
"### Authenticate your Google Cloud account\n",
"\n",
"Depending on your Jupyter environment, you may have to manually authenticate. Follow the relevant instructions below."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "74ccc9e52986"
},
"source": [
"**1. Vertex AI Workbench**\n",
"* Do nothing as you are already authenticated."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "de775a3773ba"
},
"source": [
"**2. Local JupyterLab instance, uncomment and run:**"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "254614fa0c46"
},
"outputs": [],
"source": [
"# ! gcloud auth login"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "ef21552ccea8"
},
"source": [
"**3. Colab, uncomment and run:**"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "603adbbf0532"
},
"outputs": [],
"source": [
"# from google.colab import auth\n",
"# auth.authenticate_user()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "f6b2ccc891ed"
},
"source": [
"**4. Service account or other**\n",
"* See how to grant Cloud Storage permissions to your service account at https://cloud.google.com/storage/docs/gsutil/commands/iam#ch-examples."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "zgPO1eR3CYjk"
},
"source": [
"### Create a Cloud Storage bucket\n",
"\n",
"Create a storage bucket to store intermediate artifacts such as datasets.\n",
"\n",
"- *{Note to notebook author: For any user-provided strings that need to be unique (like bucket names or model ID's), append \"-unique\" to the end so proper testing can occur}*"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "MzGDU7TWdts_"
},
"outputs": [],
"source": [
"BUCKET_URI = f\"gs://your-bucket-name-unique-{PROJECT_ID}\" # @param {type:\"string\"}"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "-EcIXiGsCePi"
},
"source": [
"**Only if your bucket doesn't already exist**: Run the following cell to create your Cloud Storage bucket."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "NIq7R4HZCfIc"
},
"outputs": [],
"source": [
"! gsutil mb -l $REGION -p $PROJECT_ID $BUCKET_URI"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "960505627ddf"
},
"source": [
"### Import libraries"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "PyQmSRbKA8r-"
},
"outputs": [],
"source": [
"import tensorflow as tf\n",
"import tensorflow_hub as hub\n",
"from google.cloud import aiplatform\n",
"from google.cloud.aiplatform import gapic"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "init_aip:mbsdk,all"
},
"source": [
"### Initialize Vertex AI SDK for Python\n",
"\n",
"Initialize the Vertex AI SDK for Python for your project."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "init_aip:mbsdk,all"
},
"outputs": [],
"source": [
"aiplatform.init(project=PROJECT_ID, location=REGION, staging_bucket=BUCKET_URI)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "accelerators:training,cpu,prediction,cpu,mbsdk"
},
"source": [
"#### Set hardware accelerators\n",
"\n",
"You can set hardware accelerators for prediction.\n",
"\n",
"Set the variables `DEPLOY_GPU/DEPLOY_NGPU` to use a container image supporting a GPU and the number of GPUs allocated to the virtual machine (VM) instance. For example, to use a GPU container image with 4 Nvidia Telsa K80 GPUs allocated to each VM, you would specify:\n",
"\n",
" (aip.gapic.AcceleratorType.NVIDIA_TESLA_K80, 4)\n",
"\n",
"\n",
"Otherwise specify `(None, None)` to use a container image to run on a CPU.\n",
"\n",
"Learn more about [hardware accelerator support for your region](https://cloud.google.com/vertex-ai/docs/general/locations#accelerators), and [GPU pricing](https://cloud.google.com/compute/gpus-pricing)."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "accelerators:training,cpu,prediction,cpu,mbsdk"
},
"outputs": [],
"source": [
"DEPLOY_GPU, DEPLOY_NGPU = (None, None)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "container:training,prediction"
},
"source": [
"#### Set pre-built containers\n",
"\n",
"Set the pre-built Docker container image for prediction.\n",
"\n",
"For the latest list, see [Pre-built containers for prediction](https://cloud.google.com/ai-platform-unified/docs/predictions/pre-built-containers)."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "container:training,prediction"
},
"outputs": [],
"source": [
"TF = \"2.5\".replace(\".\", \"-\")\n",
"\n",
"if DEPLOY_GPU:\n",
" DEPLOY_VERSION = \"tf2-gpu.{}\".format(TF)\n",
"else:\n",
" DEPLOY_VERSION = \"tf2-cpu.{}\".format(TF)\n",
"\n",
"DEPLOY_IMAGE = \"{}-docker.pkg.dev/vertex-ai/prediction/{}:latest\".format(\n",
" REGION.split(\"-\")[0], DEPLOY_VERSION\n",
")\n",
"\n",
"print(\"Deployment:\", DEPLOY_IMAGE, DEPLOY_GPU, DEPLOY_NGPU)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "d8128b8ff025"
},
"source": [
"## Get pretrained model from TensorFlow Hub\n",
"\n",
"For demonstration purposes, this tutorial uses a pretrained model from TensorFlow Hub (TFHub), which is then uploaded to a `Vertex AI Model` resource. Once you have a `Vertex AI Model` resource, the model can be deployed to a `Vertex AI Endpoint` resource.\n",
"\n",
"### Download the pretrained model\n",
"\n",
"First, you download the pretrained model from TensorFlow Hub. The model gets downloaded as a TF.Keras layer. To finalize the model, in this example, you create a `Sequential()` model with the downloaded TFHub model as a layer, and specify the input shape to the model."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "c55fa4c826f7"
},
"outputs": [],
"source": [
"tfhub_model = tf.keras.Sequential(\n",
" [hub.KerasLayer(\"https://tfhub.dev/google/imagenet/resnet_v2_101/classification/5\")]\n",
")\n",
"\n",
"tfhub_model.build([None, 32, 32, 3])\n",
"\n",
"tfhub_model.summary()"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "64618c713db9"
},
"outputs": [],
"source": [
"MODEL_DIR = BUCKET_URI + \"/model\"\n",
"tfhub_model.save(MODEL_DIR)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "e8ce91147c93"
},
"source": [
"### Upload the TensorFlow Hub model to a `Vertex AI Model` resource\n",
"\n",
"Finally, you upload the model artifacts from the TFHub model into a `Vertex AI Model` resource using the method `upload()`, with the following parameters:\n",
"\n",
"- `display_name`: A human readable name for the `Model` resource.\n",
"- `artifact_uri`: The Cloud Storage location of the model package.\n",
"- `serving_container_image_uri`: The serving container image.\n",
"\n",
"Uploading a model into a Vertex AI Model resource returns a long running operation, since it may take a few moments. \n",
"\n",
"*Note:* When you upload the model artifacts to a `Vertex AI Model` resource, you specify the corresponding deployment container image."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "ad61e1429512"
},
"outputs": [],
"source": [
"model = aiplatform.Model.upload(\n",
" display_name=\"resnet\",\n",
" artifact_uri=MODEL_DIR,\n",
" serving_container_image_uri=DEPLOY_IMAGE,\n",
" is_default_version=True,\n",
" version_aliases=[\"v1\"],\n",
")\n",
"\n",
"print(model)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "c11e98ef5391"
},
"source": [
"## Introduction to custom model evaluations\n",
"\n",
"When training a custom model, one generally performs some form of an evaluation of the trained model. Your custom model evaluation can then be imported to the corresponding model in the Vertex AI Model Registry using the `import_model_evaluation()` method. Once imported, the custom model evaluation can be subsequently retreived with the `list_model_evaluations()` method. \n",
"\n",
"The Vertex AI Model Registry supports importing multiple model evaluations for a model where each evaluation is distinquished by a unique `display_name`.\n",
"\n",
"\n",
"### Create a model evaluation\n",
"\n",
"First, you create a model evaluation in a format that corresponds to one of the predefined schemas for model evaluations. In this example, you use the schema for a classification metric, and specify the following subset of evaluation metrics as a dictionary:\n",
"\n",
"- `logLoss`: The log loss.\n",
"- `auPrc`: The accuracy.\n",
"\n",
"You then construct the `ModelEvaluation` object with the following parameters:\n",
"\n",
"- `display_name`: The human readable name for the evaluation metric.\n",
"- `metrics_schema_uri`: The schema for the specific type of evaluation metrics.\n",
"- `metrics`: The dictionary with the evaluation metrics.\n",
"\n",
"Learn more about [Schemas for evaluation metrics](https://cloud.google.com/vertex-ai/docs/evaluation/introduction#features)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "e9af222db292"
},
"outputs": [],
"source": [
"metrics = {\"logLoss\": 1.4, \"auPrc\": 0.85}\n",
"print(metrics)\n",
"\n",
"model_eval = gapic.ModelEvaluation(\n",
" display_name=\"eval\",\n",
" metrics_schema_uri=\"gs://google-cloud-aiplatform/schema/modelevaluation/classification_metrics_1.0.0.yaml\",\n",
" metrics=metrics,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "68870bd8194d"
},
"source": [
"### Upload the evaluation metrics to the Model Registry\n",
"\n",
"Next, upload the model's evaluation from the custom training job to the corresponding entry in the Vertex AI Model Registry.\n",
"\n",
"Currently, there is not yet support for this method in the SDK. Instead, you use the lower level GAPIC API interface."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "3044788848af"
},
"outputs": [],
"source": [
"API_ENDPOINT = f\"{REGION}-aiplatform.googleapis.com\"\n",
"client = gapic.ModelServiceClient(client_options={\"api_endpoint\": API_ENDPOINT})\n",
"\n",
"client.import_model_evaluation(parent=model.resource_name, model_evaluation=model_eval)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "337319a19efe"
},
"source": [
"### List the custom model evaluation\n",
"\n",
"Now that your custom metric has been uploaded to the corresponding model in the Vertex AI Model Registry, you can retriv"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "1b3374e988cf"
},
"outputs": [],
"source": [
"evaluation = model.list_model_evaluations()[0]\n",
"print(evaluation.gca_resource)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "68870bd8194d"
},
"source": [
"### Upload a second evaluation metrics to the Model Registry\n",
"\n",
"Next, upload a second model evaluation to the corresponding entry in the Vertex AI Model Registry. In this example, we refer to first evaluation metric as `eval` (from training) and the second as `prod` (from production data)."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "9b864c0ba07f"
},
"outputs": [],
"source": [
"metrics = {\"logLoss\": 1.2, \"auPrc\": 0.87}\n",
"print(metrics)\n",
"\n",
"model_prod = gapic.ModelEvaluation(\n",
" display_name=\"prod\",\n",
" metrics_schema_uri=\"gs://google-cloud-aiplatform/schema/modelevaluation/classification_metrics_1.0.0.yaml\",\n",
" metrics=metrics,\n",
")\n",
"\n",
"client.import_model_evaluation(parent=model.resource_name, model_evaluation=model_prod)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "337319a19efe"
},
"source": [
"### List a specific custom model evaluation\n",
"\n",
"Now that your custom second metric has been uploaded to the corresponding model in the Vertex AI Model Registry, you can retrieve this specific evaluation by filtering using the `display_name`:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "1b3374e988cf"
},
"outputs": [],
"source": [
"evaluations = model.list_model_evaluations()\n",
"for evaluation in evaluations:\n",
" if evaluation.display_name == \"prod\":\n",
" print(evaluation.gca_resource)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "b3a3993dc813"
},
"source": [
"### Upload version 2 of the TFHub model to the `Vertex AI Model Registry`\n",
"\n",
"Next, you upload the second version of the TFHub model as a `Model` resource in the `Vertex AI Model Registry`, with the additional following parameters:\n",
"\n",
"- `parent_model`: The existing `Model` resource for which to add this model as the next model version.\n",
"- `is_default_version`: Whether this will be the default version for the `Model` resource. In this example, you change from the default from the first version to the second version of the model.\n",
"- `version_ailiases`: User defined list of alternative alias names for the model version, such as `production`.\n",
"- `version_description`: User description of the model version.\n",
"\n",
"When a subsequent model version is created in the `Vertex AI Model Registry`, the property `version_id` will automatically be incremented. In this example, it will be set to 2 (2nd version)."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "852be5e5f964"
},
"outputs": [],
"source": [
"model_v2 = aiplatform.Model.upload(\n",
" display_name=\"resnet\",\n",
" artifact_uri=MODEL_DIR,\n",
" serving_container_image_uri=DEPLOY_IMAGE,\n",
" parent_model=model.resource_name,\n",
" is_default_version=True,\n",
" version_aliases=[\"v2\"],\n",
" version_description=\"This is the second version of the model\",\n",
")\n",
"\n",
"print(model_v2)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "68870bd8194d"
},
"source": [
"### Upload an evaluation metrics for version 2 of the model to the Model Registry\n",
"\n",
"Next, upload a model evaluation to the corresponding model version in the Vertex AI Model Registry. Note, you referenced `model_v2.resource_name` to refer to version 2 of this model."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "9b864c0ba07f"
},
"outputs": [],
"source": [
"metrics = {\"logLoss\": 1.0, \"auPrc\": 0.91}\n",
"print(metrics)\n",
"\n",
"model_eval = gapic.ModelEvaluation(\n",
" display_name=\"eval\",\n",
" metrics_schema_uri=\"gs://google-cloud-aiplatform/schema/modelevaluation/classification_metrics_1.0.0.yaml\",\n",
" metrics=metrics,\n",
")\n",
"\n",
"client.import_model_evaluation(\n",
" parent=model_v2.resource_name, model_evaluation=model_eval\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "d4e72b21b06b"
},
"source": [
"### List the evaluations for both versions of the model\n",
"\n",
"Finally, list the number of evaluations for both versions of the model."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "369a23d392a4"
},
"outputs": [],
"source": [
"evaluations = model.list_model_evaluations()\n",
"print(\"Model v1 no. of evaluations\", len(evaluations))\n",
"evaluations = model_v2.list_model_evaluations()\n",
"print(\"Model v2 no. of evaluations\", len(evaluations))"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "TpV-iwP9qw9c"
},
"source": [
"## Cleaning up\n",
"\n",
"To clean up all Google Cloud resources used in this project, you can [delete the Google Cloud\n",
"project](https://cloud.google.com/resource-manager/docs/creating-managing-projects#shutting_down_projects) you used for the tutorial.\n",
"\n",
"Otherwise, you can delete the individual resources you created in this tutorial:\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "sx_vKniMq9ZX"
},
"outputs": [],
"source": [
"import os\n",
"\n",
"# Delete model resource\n",
"model.delete()\n",
"\n",
"# Delete Cloud Storage objects that were created\n",
"delete_bucket = False\n",
"if delete_bucket or os.getenv(\"IS_TESTING\"):\n",
" ! gsutil -m rm -r $BUCKET_URI"
]
}
],
"metadata": {
"colab": {
"collapsed_sections": [],
"name": "get_started_with_custom_model_evaluation_import.ipynb",
"toc_visible": true
},
"kernelspec": {
"display_name": "Python 3",
"name": "python3"
}
},
"nbformat": 4,
"nbformat_minor": 0
}
@@ -31,7 +31,7 @@
"id": "fsv4jGuU89rX"
},
"source": [
"# E2E ML on GCP: MLOps stage 7 : monitoring: Vertex AI Model Monitoring for AutoML tabular models\n",
"# Vertex AI Model Monitoring for AutoML tabular models\n",
"\n",
"<table align=\"left\">\n",
" <td>\n",
@@ -91,7 +91,6 @@
"- Deploy the `Model` resource to the `Endpoint` resource.\n",
"- Configure the `Endpoint` resource for model monitoring.\n",
"- Generate synthetic prediction requests for skew.\n",
"- Wait for email alert notification.\n",
"- Generate synthetic prediction requests for drift.\n",
"- Wait for email alert notification.\n",
"\n",
@@ -106,7 +105,7 @@
"source": [
"### Dataset\n",
"\n",
"The dataset used for this tutorial is the GSOD dataset from [BigQuery public datasets](https://cloud.google.com/bigquery/public-data). The version of this dataset you use only the fields year, month and day to predict the value of mean daily temperature (mean_temp)."
"The dataset used for this tutorial is the GSOD dataset from [BigQuery public datasets](https://cloud.google.com/bigquery/public-data). In this notebook, you use only the fields year, month and day from the dataset to predict the value of mean daily temperature (mean_temp)."
]
},
{
@@ -563,7 +562,7 @@
"source": [
"### Create BigQuery client\n",
"\n",
"In this tutorial, you use data from the same public BigQuery table that was used to train the pre-trained model. You create a client interface, which you subsequently use to access the data."
"In this tutorial, you explore the monitoring data stored in BigQuery. You create a client interface, which you subsequently use to access the data."
]
},
{
@@ -668,23 +667,6 @@
"print(dataset.resource_name)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "set_transformations:gsod"
},
"outputs": [],
"source": [
"TRANSFORMATIONS = [\n",
" {\"auto\": {\"column_name\": \"year\"}},\n",
" {\"auto\": {\"column_name\": \"month\"}},\n",
" {\"auto\": {\"column_name\": \"day\"}},\n",
"]\n",
"\n",
"label_column = \"mean_temp\""
]
},
{
"cell_type": "markdown",
"metadata": {
@@ -703,7 +685,7 @@
"- `optimization_prediction_type`: The type task to train the model for.\n",
" - `classification`: A tabuar classification model.\n",
" - `regression`: A tabular regression model.\n",
"- `column_transformations`: (Optional): Transformations to apply to the input columns\n",
"- `column_transformations`: (Optional): Transformations to apply to the input columns. In this example, you set the column transformations to use the default transformation based on their data type.\n",
"- `optimization_objective`: The optimization objective to minimize or maximize.\n",
" - binary classification:\n",
" - `minimize-log-loss`\n",
@@ -719,6 +701,23 @@
" - `minimize-rmsle`"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "set_transformations:gsod"
},
"outputs": [],
"source": [
"TRANSFORMATIONS = [\n",
" {\"auto\": {\"column_name\": \"year\"}},\n",
" {\"auto\": {\"column_name\": \"month\"}},\n",
" {\"auto\": {\"column_name\": \"day\"}},\n",
"]\n",
"\n",
"label_column = \"mean_temp\""
]
},
{
"cell_type": "code",
"execution_count": null,
@@ -1065,7 +1064,9 @@
"You are receiving this mail because you are using the Vertex AI Model Monitoring service.\n",
"This mail is to inform you that we received your request to set up drift or skew detection for the Prediction Endpoint listed below. Starting from now, incoming prediction requests will be sampled and logged for analysis.\n",
"Raw requests and responses will be collected from prediction service and saved in bq://[your-project-id].model_deployment_monitoring_[endpoint-id].serving_predict .\n",
"</blockquote>"
"</blockquote>\n",
"\n",
"*Note:* You do not need to wait for the email notification to continue to the next step."
]
},
{
@@ -1130,7 +1131,7 @@
"\n",
"Next, you extract the first 1000 instances from the BigQuery training table to use for prediction requests. You modify the data (synthetic) to trigger the skew detection in the prediction requests from the training distribution versus serving distribution, as follows:\n",
"\n",
"- `year`: Set all values to 3 (was 2)."
"- `year`: Set all values to 3."
]
},
{
@@ -1275,7 +1276,7 @@
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "b91a0e19ff8b"
"id": "2e64ffaae2de"
},
"outputs": [],
"source": [
@@ -1283,41 +1284,6 @@
" time.sleep(60 * 45)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "2b5859ea4ae9"
},
"source": [
"### Logging sampled requests\n",
"\n",
"On the next monitoring interval, the sampled predictions are then copied over to the BigQuery logging table. Once the entries are in the BigQuery table, the monitoring service will analyze the sampled data.\n",
"\n",
"Next, you wait for the logged entres to appear in the BigQuery table used for logging prediction samples. Since you sent 1000 prediction requests, with 50% sampling, you should see around 1000 entries."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "bd177a8decbb"
},
"outputs": [],
"source": [
"while True:\n",
" time.sleep(180)\n",
"\n",
" ENDPOINT_ID = endpoint.resource_name.split(\"/\")[-1]\n",
"\n",
" table = bigquery.TableReference.from_string(\n",
" f\"{PROJECT_ID}.model_deployment_monitoring_{ENDPOINT_ID}.serving_predict\"\n",
" )\n",
" rows = bqclient.list_rows(table)\n",
" print(rows.total_rows)\n",
" if rows.total_rows > 505:\n",
" break"
]
},
{
"cell_type": "markdown",
"metadata": {
@@ -1419,7 +1385,7 @@
" )\n",
" rows = bqclient.list_rows(table)\n",
" print(rows.total_rows)\n",
" if rows.total_rows > 1050:\n",
" if rows.total_rows > 505:\n",
" break"
]
},
@@ -55,7 +55,11 @@
"</table>\n",
"<br/><br/><br/>\n",
"\n",
"**This feature is in private preview**. You can request access to model monitoring for AutoML image models with this [Form](https://docs.google.com/forms/d/1aTitFrlNRlUAF_gvbMFQWDY0Oq8T_fUrN1m79atF8xQ/edit?resourcekey=0-psKsTGtVUFFUBDL2timIQA)."
"**This feature is in experimental**. You can request access to model monitoring for AutoML image models with this [Form](https://docs.google.com/forms/d/1aTitFrlNRlUAF_gvbMFQWDY0Oq8T_fUrN1m79atF8xQ/edit?resourcekey=0-psKsTGtVUFFUBDL2timIQA).\n",
"\n",
"**LEGAL NOTICE**\n",
"\n",
"This is an Experimental release. Experiments are focused on validating a prototype and are not guaranteed to be released. Experiments are covered by the [Pre-GA Offerings Terms](https://cloud.google.com/terms/service-terms) of the Google Cloud Platform Terms of Service. They are not intended for production use or covered by any SLA, support obligation, or deprecation policy and might be subject to backward-incompatible changes.\n"
]
},
{
@@ -188,7 +192,8 @@
"\n",
"! pip3 install {USER_FLAG} -q -U google-cloud-aiplatform \"shapely<2\" \\\n",
" tensorflow==2.7 \\\n",
" google-api-core==2.10"
" google-api-core==2.10 \\\n",
" protobuf==3.20.3"
]
},
{
File diff suppressed because one or more lines are too long
File diff suppressed because it is too large Load Diff
+25 -1
View File
@@ -304,5 +304,29 @@ The steps performed include:
&nbsp;&nbsp;&nbsp;Learn more about [AutoML components](https://cloud.google.com/vertex-ai/docs/pipelines/vertex-automl-component).
&nbsp;&nbsp;&nbsp;Learn more about [BigQuery ML components](https://cloud.google.com/vertex-ai/docs/pipelines/bigqueryml-component).
[Train custom tabular ML models with many frameworks and import to Vertex AI using Vertex Pipelines](https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/main/notebooks/official/pipelines/Train_tabular_models_with_many_frameworks_and_import_to_Vertex_AI_using_Pipelines)
Learn how to build a pipeline that does the following:
* Ingest data
* Transform data
* Clean up data
* Split data into train/test subsets
* Configure model
* Train model using multiple ML frameworks
* Import model into Vertex Model Registry
* [Optional] Deploy model to Vertex Endpoints for serving
Included pipelines:
* Train ML model
* * Tabular classification
* * * TensorFlow
* * * PyTorch
* * * XGBoost
* * * Scikit-learn
* * Tabular regression
* * * TensorFlow
* * * PyTorch
* * * XGBoost
* * * Scikit-learn
@@ -0,0 +1,73 @@
# python3 -m pip install "kfp<2.0.0" "google-cloud-aiplatform>=1.16.0" --upgrade --quiet
from kfp import components
# %% Loading components
download_from_gcs_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/storage/download/component.yaml")
select_columns_using_Pandas_on_CSV_data_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/pandas/Select_columns/in_CSV_format/component.yaml")
fill_all_missing_values_using_Pandas_on_CSV_data_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/pandas/Fill_all_missing_values/in_CSV_format/component.yaml")
binarize_column_using_Pandas_on_CSV_data_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/pandas/Binarize_column/in_CSV_format/component.yaml")
train_logistic_regression_model_using_scikit_learn_from_CSV_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/ML_frameworks/Scikit_learn/Train_logistic_regression_model/from_CSV/component.yaml")
upload_Scikit_learn_pickle_model_to_Google_Cloud_Vertex_AI_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/1f5cf6e06409b704064b2086c0a705e4e6b4fcde/community-content/pipeline_components/google-cloud/Vertex_AI/Models/Upload_Scikit-learn_pickle_model/component.yaml")
deploy_model_to_endpoint_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/Vertex_AI/Models/Deploy_to_endpoint/component.yaml")
# %% Pipeline definition
def train_tabular_classification_logistic_regression_model_using_Scikit_learn_pipeline():
dataset_gcs_uri = "gs://ml-pipeline-dataset/Chicago_taxi_trips/chicago_taxi_trips_2019-01-01_-_2019-02-01_limit=10000.csv"
feature_columns = ["trip_seconds", "trip_miles", "pickup_community_area", "dropoff_community_area", "fare", "tolls", "extras"] # Excluded "trip_total"
label_column = "tips"
# Deploying the model might incur additional costs over time
deploy_model = False
classification_label_column = "class"
all_columns = [label_column] + feature_columns
training_data = download_from_gcs_op(
gcs_path=dataset_gcs_uri
).outputs["Data"]
training_data = select_columns_using_Pandas_on_CSV_data_op(
table=training_data,
column_names=all_columns,
).outputs["transformed_table"]
# Cleaning the NaN values.
training_data = fill_all_missing_values_using_Pandas_on_CSV_data_op(
table=training_data,
replacement_value="0",
#replacement_type_name="float",
).outputs["transformed_table"]
classification_training_data = binarize_column_using_Pandas_on_CSV_data_op(
table=training_data,
column_name=label_column,
predicate="> 0",
new_column_name=classification_label_column,
).outputs["transformed_table"]
model = train_logistic_regression_model_using_scikit_learn_from_CSV_op(
dataset=classification_training_data,
label_column_name=classification_label_column,
# Optional:
#penalty="l2",
#solver="lbfgs",
#max_iterations=100,
#multi_class_mode="auto",
#random_seed=0,
).outputs["model"]
vertex_model_name = upload_Scikit_learn_pickle_model_to_Google_Cloud_Vertex_AI_op(
model=model,
).outputs["model_name"]
# Deploying the model might incur additional costs over time
if deploy_model:
sklearn_vertex_endpoint_name = deploy_model_to_endpoint_op(
model_name=vertex_model_name,
).outputs["endpoint_name"]
pipeline_func = train_tabular_classification_logistic_regression_model_using_Scikit_learn_pipeline
# %% Pipeline submission
if __name__ == '__main__':
from google.cloud import aiplatform
aiplatform.PipelineJob.from_pipeline_func(pipeline_func=pipeline_func).submit()
@@ -0,0 +1,95 @@
# python3 -m pip install "kfp<2.0.0" "google-cloud-aiplatform>=1.16.0" --upgrade --quiet
from kfp import components
# %% Loading components
download_from_gcs_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/storage/download/component.yaml")
select_columns_using_Pandas_on_CSV_data_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/pandas/Select_columns/in_CSV_format/component.yaml")
fill_all_missing_values_using_Pandas_on_CSV_data_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/pandas/Fill_all_missing_values/in_CSV_format/component.yaml")
binarize_column_using_Pandas_on_CSV_data_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/pandas/Binarize_column/in_CSV_format/component.yaml")
create_fully_connected_pytorch_network_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/PyTorch/Create_fully_connected_network/component.yaml")
train_pytorch_model_from_csv_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/PyTorch/Train_PyTorch_model/from_CSV/component.yaml")
create_pytorch_model_archive_with_base_handler_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/PyTorch/Create_PyTorch_Model_Archive/with_base_handler/component.yaml")
upload_PyTorch_model_archive_to_Google_Cloud_Vertex_AI_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/Vertex_AI/Models/Upload_PyTorch_model_archive/component.yaml")
deploy_model_to_endpoint_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/Vertex_AI/Models/Deploy_to_endpoint/component.yaml")
# %% Pipeline definition
def train_tabular_classification_model_using_PyTorch_pipeline():
dataset_gcs_uri = "gs://ml-pipeline-dataset/Chicago_taxi_trips/chicago_taxi_trips_2019-01-01_-_2019-02-01_limit=10000.csv"
feature_columns = ["trip_seconds", "trip_miles", "pickup_community_area", "dropoff_community_area", "fare", "tolls", "extras"] # Excluded "trip_total"
label_column = "tips"
# Deploying the model might incur additional costs over time
deploy_model = False
classification_label_column = "class"
all_columns = [label_column] + feature_columns
training_data = download_from_gcs_op(
gcs_path=dataset_gcs_uri
).outputs["Data"]
training_data = select_columns_using_Pandas_on_CSV_data_op(
table=training_data,
column_names=all_columns,
).outputs["transformed_table"]
# Cleaning the NaN values.
training_data = fill_all_missing_values_using_Pandas_on_CSV_data_op(
table=training_data,
replacement_value="0",
#replacement_type_name="float",
).outputs["transformed_table"]
classification_training_data = binarize_column_using_Pandas_on_CSV_data_op(
table=training_data,
column_name=label_column,
predicate=" > 0",
new_column_name=classification_label_column,
).outputs["transformed_table"]
network = create_fully_connected_pytorch_network_op(
input_size=len(feature_columns),
# Optional:
hidden_layer_sizes=[10],
activation_name="elu",
output_activation_name="sigmoid",
# output_size=1,
).outputs["model"]
model = train_pytorch_model_from_csv_op(
model=network,
training_data=classification_training_data,
label_column_name=classification_label_column,
loss_function_name="binary_cross_entropy",
# Optional:
#number_of_epochs=1,
#learning_rate=0.1,
#optimizer_name="Adadelta",
#optimizer_parameters={},
#batch_size=32,
#batch_log_interval=100,
#random_seed=0,
).outputs["trained_model"]
model_archive = create_pytorch_model_archive_with_base_handler_op(
model=model,
# Optional:
# model_name="model",
# model_version="1.0",
).outputs["Model archive"]
vertex_model_name = upload_PyTorch_model_archive_to_Google_Cloud_Vertex_AI_op(
model_archive=model_archive,
).outputs["model_name"]
# Deploying the model might incur additional costs over time
if deploy_model:
vertex_endpoint_name = deploy_model_to_endpoint_op(
model_name=vertex_model_name,
).outputs["endpoint_name"]
pipeline_func=train_tabular_classification_model_using_PyTorch_pipeline
# %% Pipeline submission
if __name__ == '__main__':
from google.cloud import aiplatform
aiplatform.PipelineJob.from_pipeline_func(pipeline_func=pipeline_func).submit()
@@ -0,0 +1,106 @@
# python3 -m pip install "kfp<2.0.0" "google-cloud-aiplatform>=1.16.0" --upgrade --quiet
from kfp import components
# %% Loading components
download_from_gcs_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/storage/download/component.yaml")
select_columns_using_Pandas_on_CSV_data_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/pandas/Select_columns/in_CSV_format/component.yaml")
fill_all_missing_values_using_Pandas_on_CSV_data_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/pandas/Fill_all_missing_values/in_CSV_format/component.yaml")
binarize_column_using_Pandas_on_CSV_data_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/pandas/Binarize_column/in_CSV_format/component.yaml")
split_rows_into_subsets_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/dataset_manipulation/Split_rows_into_subsets/in_CSV/component.yaml")
create_fully_connected_tensorflow_network_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/tensorflow/Create_fully_connected_network/component.yaml")
train_model_using_Keras_on_CSV_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/tensorflow/Train_model_using_Keras/on_CSV/component.yaml")
predict_with_TensorFlow_model_on_CSV_data_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/tensorflow/Predict/on_CSV/component.yaml")
upload_Tensorflow_model_to_Google_Cloud_Vertex_AI_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/Vertex_AI/Models/Upload_Tensorflow_model/component.yaml")
deploy_model_to_endpoint_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/Vertex_AI/Models/Deploy_to_endpoint/component.yaml")
# %% Pipeline definition
def train_tabular_classification_model_using_TensorFlow_pipeline():
dataset_gcs_uri = "gs://ml-pipeline-dataset/Chicago_taxi_trips/chicago_taxi_trips_2019-01-01_-_2019-02-01_limit=10000.csv"
feature_columns = ["trip_seconds", "trip_miles", "pickup_community_area", "dropoff_community_area", "fare", "tolls", "extras"] # Excluded "trip_total"
label_column = "tips"
training_set_fraction = 0.8
# Deploying the model might incur additional costs over time
deploy_model = False
classification_label_column = "class"
all_columns = [label_column] + feature_columns
dataset = download_from_gcs_op(
gcs_path=dataset_gcs_uri
).outputs["Data"]
dataset = select_columns_using_Pandas_on_CSV_data_op(
table=dataset,
column_names=all_columns,
).outputs["transformed_table"]
dataset = fill_all_missing_values_using_Pandas_on_CSV_data_op(
table=dataset,
replacement_value="0",
# # Optional:
# column_names=None, # =[...]
).outputs["transformed_table"]
classification_dataset = binarize_column_using_Pandas_on_CSV_data_op(
table=dataset,
column_name=label_column,
predicate=" > 0",
new_column_name=classification_label_column,
).outputs["transformed_table"]
split_task = split_rows_into_subsets_op(
table=classification_dataset,
fraction_1=training_set_fraction,
)
classification_training_data = split_task.outputs["split_1"]
classification_testing_data = split_task.outputs["split_2"]
network = create_fully_connected_tensorflow_network_op(
input_size=len(feature_columns),
# Optional:
hidden_layer_sizes=[10],
activation_name="elu",
output_activation_name="sigmoid",
# output_size=1,
).outputs["model"]
model = train_model_using_Keras_on_CSV_op(
training_data=classification_training_data,
model=network,
label_column_name=classification_label_column,
# Optional:
loss_function_name="binary_crossentropy",
number_of_epochs=10,
#learning_rate=0.1,
#optimizer_name="Adadelta",
#optimizer_parameters={},
#batch_size=32,
#metric_names=["mean_absolute_error"],
#random_seed=0,
).outputs["trained_model"]
predictions = predict_with_TensorFlow_model_on_CSV_data_op(
dataset=classification_testing_data,
model=model,
# label_column_name needs to be set when doing prediction on a dataset that has labels
label_column_name=classification_label_column,
# Optional:
# batch_size=1000,
).outputs["predictions"]
vertex_model_name = upload_Tensorflow_model_to_Google_Cloud_Vertex_AI_op(
model=model,
).outputs["model_name"]
# Deploying the model might incur additional costs over time
if deploy_model:
vertex_endpoint_name = deploy_model_to_endpoint_op(
model_name=vertex_model_name,
).outputs["endpoint_name"]
pipeline_func = train_tabular_classification_model_using_TensorFlow_pipeline
# %% Pipeline submission
if __name__ == '__main__':
from google.cloud import aiplatform
aiplatform.PipelineJob.from_pipeline_func(pipeline_func=pipeline_func).submit()
@@ -0,0 +1,94 @@
# python3 -m pip install "kfp<2.0.0" "google-cloud-aiplatform>=1.16.0" --upgrade --quiet
from kfp import components
# %% Loading components
download_from_gcs_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/storage/download/component.yaml")
select_columns_using_Pandas_on_CSV_data_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/pandas/Select_columns/in_CSV_format/component.yaml")
fill_all_missing_values_using_Pandas_on_CSV_data_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/pandas/Fill_all_missing_values/in_CSV_format/component.yaml")
binarize_column_using_Pandas_on_CSV_data_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/pandas/Binarize_column/in_CSV_format/component.yaml")
split_rows_into_subsets_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/dataset_manipulation/Split_rows_into_subsets/in_CSV/component.yaml")
train_XGBoost_model_on_CSV_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/XGBoost/Train/component.yaml")
xgboost_predict_on_CSV_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/XGBoost/Predict/component.yaml")
upload_XGBoost_model_to_Google_Cloud_Vertex_AI_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/Vertex_AI/Models/Upload_XGBoost_model/component.yaml")
deploy_model_to_endpoint_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/Vertex_AI/Models/Deploy_to_endpoint/component.yaml")
# %% Pipeline definition
def train_tabular_classification_model_using_XGBoost_pipeline():
dataset_gcs_uri = "gs://ml-pipeline-dataset/Chicago_taxi_trips/chicago_taxi_trips_2019-01-01_-_2019-02-01_limit=10000.csv"
feature_columns = ["trip_seconds", "trip_miles", "pickup_community_area", "dropoff_community_area", "fare", "tolls", "extras"] # Excluded "trip_total"
label_column = "tips"
training_set_fraction = 0.8
# Deploying the model might incur additional costs over time
deploy_model = False
classification_label_column = "class"
all_columns = [label_column] + feature_columns
dataset = download_from_gcs_op(
gcs_path=dataset_gcs_uri
).outputs["Data"]
dataset = select_columns_using_Pandas_on_CSV_data_op(
table=dataset,
column_names=all_columns,
).outputs["transformed_table"]
dataset = fill_all_missing_values_using_Pandas_on_CSV_data_op(
table=dataset,
replacement_value="0",
# # Optional:
# column_names=None, # =[...]
).outputs["transformed_table"]
classification_dataset = binarize_column_using_Pandas_on_CSV_data_op(
table=dataset,
column_name=label_column,
predicate="> 0",
new_column_name=classification_label_column,
).outputs["transformed_table"]
split_task = split_rows_into_subsets_op(
table=classification_dataset,
fraction_1=training_set_fraction,
)
classification_training_data = split_task.outputs["split_1"]
classification_testing_data = split_task.outputs["split_2"]
model = train_XGBoost_model_on_CSV_op(
training_data=classification_training_data,
label_column_name=classification_label_column,
objective="binary:logistic",
# Optional:
#starting_model=None,
#num_iterations=10,
#booster_params={},
#booster="gbtree",
#learning_rate=0.3,
#min_split_loss=0,
#max_depth=6,
).outputs["model"]
# Predicting on the testing data
predictions = xgboost_predict_on_CSV_op(
data=classification_testing_data,
model=model,
# label_column needs to be set when doing prediction on a dataset that has labels
label_column_name=classification_label_column,
).outputs["predictions"]
vertex_model_name = upload_XGBoost_model_to_Google_Cloud_Vertex_AI_op(
model=model,
).outputs["model_name"]
# Deploying the model might incur additional costs over time
if deploy_model:
vertex_endpoint_name = deploy_model_to_endpoint_op(
model_name=vertex_model_name,
).outputs["endpoint_name"]
pipeline_func = train_tabular_classification_model_using_XGBoost_pipeline
# %% Pipeline submission
if __name__ == '__main__':
from google.cloud import aiplatform
aiplatform.PipelineJob.from_pipeline_func(pipeline_func=pipeline_func).submit()
@@ -0,0 +1,224 @@
# python3 -m pip install "kfp<2.0.0" "google-cloud-aiplatform>=1.16.0" --upgrade --quiet
from kfp import components
# %% Loading components
download_from_gcs_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/storage/download/component.yaml")
select_columns_using_Pandas_on_CSV_data_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/pandas/Select_columns/in_CSV_format/component.yaml")
fill_all_missing_values_using_Pandas_on_CSV_data_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/pandas/Fill_all_missing_values/in_CSV_format/component.yaml")
binarize_column_using_Pandas_on_CSV_data_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/pandas/Binarize_column/in_CSV_format/component.yaml")
split_rows_into_subsets_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/dataset_manipulation/Split_rows_into_subsets/in_CSV/component.yaml")
# TensorFlow
create_fully_connected_tensorflow_network_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/tensorflow/Create_fully_connected_network/component.yaml")
train_model_using_Keras_on_CSV_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/tensorflow/Train_model_using_Keras/on_CSV/component.yaml")
predict_with_TensorFlow_model_on_CSV_data_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/tensorflow/Predict/on_CSV/component.yaml")
upload_Tensorflow_model_to_Google_Cloud_Vertex_AI_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/Vertex_AI/Models/Upload_Tensorflow_model/component.yaml")
# PyTorch
create_fully_connected_pytorch_network_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/PyTorch/Create_fully_connected_network/component.yaml")
train_pytorch_model_from_csv_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/PyTorch/Train_PyTorch_model/from_CSV/component.yaml")
create_pytorch_model_archive_with_base_handler_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/PyTorch/Create_PyTorch_Model_Archive/with_base_handler/component.yaml")
upload_PyTorch_model_archive_to_Google_Cloud_Vertex_AI_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/Vertex_AI/Models/Upload_PyTorch_model_archive/component.yaml")
# XGBoost
train_XGBoost_model_on_CSV_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/XGBoost/Train/component.yaml")
xgboost_predict_on_CSV_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/XGBoost/Predict/component.yaml")
upload_XGBoost_model_to_Google_Cloud_Vertex_AI_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/Vertex_AI/Models/Upload_XGBoost_model/component.yaml")
# Scikit-learn
#train_linear_regression_model_using_scikit_learn_from_CSV_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/1f5cf6e06409b704064b2086c0a705e4e6b4fcde/community-content/pipeline_components/ML_frameworks/Scikit_learn/Train_linear_regression_model/from_CSV/component.yaml")
train_logistic_regression_model_using_scikit_learn_from_CSV_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/1f5cf6e06409b704064b2086c0a705e4e6b4fcde/community-content/pipeline_components/ML_frameworks/Scikit_learn/Train_logistic_regression_model/from_CSV/component.yaml")
upload_Scikit_learn_pickle_model_to_Google_Cloud_Vertex_AI_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/Vertex_AI/Models/Upload_Scikit-learn_pickle_model/component.yaml")
# Vertex AI
deploy_model_to_endpoint_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/Vertex_AI/Models/Deploy_to_endpoint/component.yaml")
# %% Pipeline definition
def train_tabular_classification_model_using_all_frameworks_pipeline():
dataset_gcs_uri = "gs://ml-pipeline-dataset/Chicago_taxi_trips/chicago_taxi_trips_2019-01-01_-_2019-02-01_limit=10000.csv"
feature_columns = ["trip_seconds", "trip_miles", "pickup_community_area", "dropoff_community_area", "fare", "tolls", "extras"] # Excluded "trip_total"
label_column = "tips"
training_set_fraction = 0.8
# Deploying the model might incur additional costs over time
deploy_model = False
classification_label_column = "class"
all_columns = [label_column] + feature_columns
dataset = download_from_gcs_op(
gcs_path=dataset_gcs_uri
).outputs["Data"]
dataset = select_columns_using_Pandas_on_CSV_data_op(
table=dataset,
column_names=all_columns,
).outputs["transformed_table"]
dataset = fill_all_missing_values_using_Pandas_on_CSV_data_op(
table=dataset,
replacement_value="0",
# # Optional:
# column_names=None, # =[...]
).outputs["transformed_table"]
classification_dataset = binarize_column_using_Pandas_on_CSV_data_op(
table=dataset,
column_name=label_column,
predicate=" > 0",
new_column_name=classification_label_column,
).outputs["transformed_table"]
split_task = split_rows_into_subsets_op(
table=classification_dataset,
fraction_1=training_set_fraction,
)
classification_training_data = split_task.outputs["split_1"]
classification_testing_data = split_task.outputs["split_2"]
# TensorFlow
tensorflow_network = create_fully_connected_tensorflow_network_op(
input_size=len(feature_columns),
# Optional:
hidden_layer_sizes=[10],
activation_name="elu",
output_activation_name="sigmoid",
# output_size=1,
).outputs["model"]
tensorflow_model = train_model_using_Keras_on_CSV_op(
training_data=classification_training_data,
model=tensorflow_network,
label_column_name=classification_label_column,
# Optional:
loss_function_name="binary_crossentropy",
number_of_epochs=10,
#learning_rate=0.1,
#optimizer_name="Adadelta",
#optimizer_parameters={},
#batch_size=32,
#metric_names=["mean_absolute_error"],
#random_seed=0,
).outputs["trained_model"]
tensorflow_predictions = predict_with_TensorFlow_model_on_CSV_data_op(
dataset=classification_testing_data,
model=tensorflow_model,
# label_column_name needs to be set when doing prediction on a dataset that has labels
label_column_name=classification_label_column,
# Optional:
# batch_size=1000,
).outputs["predictions"]
tensorflow_vertex_model_name = upload_Tensorflow_model_to_Google_Cloud_Vertex_AI_op(
model=tensorflow_model,
).outputs["model_name"]
# Deploying the model might incur additional costs over time
if deploy_model:
tensorflow_vertex_endpoint_name = deploy_model_to_endpoint_op(
model_name=tensorflow_vertex_model_name,
).outputs["endpoint_name"]
# PyTorch
pytorch_network = create_fully_connected_pytorch_network_op(
input_size=len(feature_columns),
# Optional:
hidden_layer_sizes=[10],
activation_name="elu",
output_activation_name="sigmoid",
# output_size=1,
).outputs["model"]
pytorch_model = train_pytorch_model_from_csv_op(
model=pytorch_network,
training_data=classification_training_data,
label_column_name=classification_label_column,
loss_function_name="binary_cross_entropy",
# Optional:
#number_of_epochs=1,
#learning_rate=0.1,
#optimizer_name="Adadelta",
#optimizer_parameters={},
#batch_size=32,
#batch_log_interval=100,
#random_seed=0,
).outputs["trained_model"]
pytorch_model_archive = create_pytorch_model_archive_with_base_handler_op(
model=pytorch_model,
# Optional:
# model_name="model",
# model_version="1.0",
).outputs["Model archive"]
pytorch_vertex_model_name = upload_PyTorch_model_archive_to_Google_Cloud_Vertex_AI_op(
model_archive=pytorch_model_archive,
).outputs["model_name"]
# Deploying the model might incur additional costs over time
if deploy_model:
pytorch_vertex_endpoint_name = deploy_model_to_endpoint_op(
model_name=pytorch_vertex_model_name,
).outputs["endpoint_name"]
# XGBoost
xgboost_model = train_XGBoost_model_on_CSV_op(
training_data=classification_training_data,
label_column_name=classification_label_column,
objective="binary:logistic",
# Optional:
#starting_model=None,
#num_iterations=10,
#booster_params={},
#booster="gbtree",
#learning_rate=0.3,
#min_split_loss=0,
#max_depth=6,
).outputs["model"]
# Predicting on the testing data
xgboost_predictions = xgboost_predict_on_CSV_op(
data=classification_testing_data,
model=xgboost_model,
# label_column needs to be set when doing prediction on a dataset that has labels
label_column_name=classification_label_column,
).outputs["predictions"]
xgboost_vertex_model_name = upload_XGBoost_model_to_Google_Cloud_Vertex_AI_op(
model=xgboost_model,
).outputs["model_name"]
# Deploying the model might incur additional costs over time
if deploy_model:
xgboost_vertex_endpoint_name = deploy_model_to_endpoint_op(
model_name=xgboost_vertex_model_name,
).outputs["endpoint_name"]
# Scikit-learn
sklearn_model = train_logistic_regression_model_using_scikit_learn_from_CSV_op(
dataset=classification_training_data,
label_column_name=classification_label_column,
# Optional:
#penalty="l2",
#solver="lbfgs",
#max_iterations=100,
#multi_class_mode="auto",
#random_seed=0,
).outputs["model"]
sklearn_vertex_model_name = upload_Scikit_learn_pickle_model_to_Google_Cloud_Vertex_AI_op(
model=sklearn_model,
).outputs["model_name"]
# Deploying the model might incur additional costs over time
if deploy_model:
sklearn_vertex_endpoint_name = deploy_model_to_endpoint_op(
model_name=sklearn_vertex_model_name,
).outputs["endpoint_name"]
pipeline_func=train_tabular_classification_model_using_all_frameworks_pipeline
# %% Pipeline submission
if __name__ == '__main__':
from google.cloud import aiplatform
aiplatform.PipelineJob.from_pipeline_func(pipeline_func=pipeline_func).submit()
@@ -0,0 +1,57 @@
# python3 -m pip install "kfp<2.0.0" "google-cloud-aiplatform>=1.16.0" --upgrade --quiet
from kfp import components
# %% Loading components
download_from_gcs_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/storage/download/component.yaml")
select_columns_using_Pandas_on_CSV_data_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/pandas/Select_columns/in_CSV_format/component.yaml")
fill_all_missing_values_using_Pandas_on_CSV_data_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/pandas/Fill_all_missing_values/in_CSV_format/component.yaml")
train_linear_regression_model_using_scikit_learn_from_CSV_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/1f5cf6e06409b704064b2086c0a705e4e6b4fcde/community-content/pipeline_components/ML_frameworks/Scikit_learn/Train_linear_regression_model/from_CSV/component.yaml")
upload_Scikit_learn_pickle_model_to_Google_Cloud_Vertex_AI_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/Vertex_AI/Models/Upload_Scikit-learn_pickle_model/component.yaml")
deploy_model_to_endpoint_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/Vertex_AI/Models/Deploy_to_endpoint/component.yaml")
# %% Pipeline definition
def train_tabular_regression_linear_model_using_Scikit_learn_pipeline():
dataset_gcs_uri = "gs://ml-pipeline-dataset/Chicago_taxi_trips/chicago_taxi_trips_2019-01-01_-_2019-02-01_limit=10000.csv"
feature_columns = ["trip_seconds", "trip_miles", "pickup_community_area", "dropoff_community_area", "fare", "tolls", "extras"] # Excluded "trip_total"
label_column = "tips"
all_columns = [label_column] + feature_columns
# Deploying the model might incur additional costs over time
deploy_model = False
training_data = download_from_gcs_op(
gcs_path=dataset_gcs_uri
).outputs["Data"]
training_data = select_columns_using_Pandas_on_CSV_data_op(
table=training_data,
column_names=all_columns,
).outputs["transformed_table"]
# Cleaning the NaN values.
training_data = fill_all_missing_values_using_Pandas_on_CSV_data_op(
table=training_data,
replacement_value="0",
#replacement_type_name="float",
).outputs["transformed_table"]
model = train_linear_regression_model_using_scikit_learn_from_CSV_op(
dataset=training_data,
label_column_name=label_column,
).outputs["model"]
vertex_model_name = upload_Scikit_learn_pickle_model_to_Google_Cloud_Vertex_AI_op(
model=model,
).outputs["model_name"]
# Deploying the model might incur additional costs over time
if deploy_model:
sklearn_vertex_endpoint_name = deploy_model_to_endpoint_op(
model_name=vertex_model_name,
).outputs["endpoint_name"]
pipeline_func = train_tabular_regression_linear_model_using_Scikit_learn_pipeline
# %% Pipeline submission
if __name__ == '__main__':
from google.cloud import aiplatform
aiplatform.PipelineJob.from_pipeline_func(pipeline_func=pipeline_func).submit()
@@ -0,0 +1,85 @@
# python3 -m pip install "kfp<2.0.0" "google-cloud-aiplatform>=1.16.0" --upgrade --quiet
from kfp import components
# %% Loading components
download_from_gcs_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/storage/download/component.yaml")
select_columns_using_Pandas_on_CSV_data_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/pandas/Select_columns/in_CSV_format/component.yaml")
fill_all_missing_values_using_Pandas_on_CSV_data_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/pandas/Fill_all_missing_values/in_CSV_format/component.yaml")
create_fully_connected_pytorch_network_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/PyTorch/Create_fully_connected_network/component.yaml")
train_pytorch_model_from_csv_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/PyTorch/Train_PyTorch_model/from_CSV/component.yaml")
create_pytorch_model_archive_with_base_handler_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/PyTorch/Create_PyTorch_Model_Archive/with_base_handler/component.yaml")
upload_PyTorch_model_archive_to_Google_Cloud_Vertex_AI_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/Vertex_AI/Models/Upload_PyTorch_model_archive/component.yaml")
deploy_model_to_endpoint_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/Vertex_AI/Models/Deploy_to_endpoint/component.yaml")
# %% Pipeline definition
def train_tabular_regression_model_using_PyTorch_pipeline():
dataset_gcs_uri = "gs://ml-pipeline-dataset/Chicago_taxi_trips/chicago_taxi_trips_2019-01-01_-_2019-02-01_limit=10000.csv"
feature_columns = ["trip_seconds", "trip_miles", "pickup_community_area", "dropoff_community_area", "fare", "tolls", "extras"] # Excluded "trip_total"
label_column = "tips"
all_columns = [label_column] + feature_columns
# Deploying the model might incur additional costs over time
deploy_model = False
training_data = download_from_gcs_op(
gcs_path=dataset_gcs_uri
).outputs["Data"]
training_data = select_columns_using_Pandas_on_CSV_data_op(
table=training_data,
column_names=all_columns,
).outputs["transformed_table"]
# Cleaning the NaN values.
training_data = fill_all_missing_values_using_Pandas_on_CSV_data_op(
table=training_data,
replacement_value="0",
#replacement_type_name="float",
).outputs["transformed_table"]
network = create_fully_connected_pytorch_network_op(
input_size=len(feature_columns),
# Optional:
hidden_layer_sizes=[10],
activation_name="elu",
# output_activation_name=None,
# output_size=1,
).outputs["model"]
model = train_pytorch_model_from_csv_op(
model=network,
training_data=training_data,
label_column_name=label_column,
# Optional:
#loss_function_name="mse_loss",
#number_of_epochs=1,
#learning_rate=0.1,
#optimizer_name="Adadelta",
#optimizer_parameters={},
#batch_size=32,
#batch_log_interval=100,
#random_seed=0,
).outputs["trained_model"]
model_archive = create_pytorch_model_archive_with_base_handler_op(
model=model,
# Optional:
# model_name="model",
# model_version="1.0",
).outputs["Model archive"]
vertex_model_name = upload_PyTorch_model_archive_to_Google_Cloud_Vertex_AI_op(
model_archive=model_archive,
).outputs["model_name"]
# Deploying the model might incur additional costs over time
if deploy_model:
vertex_endpoint_name = deploy_model_to_endpoint_op(
model_name=vertex_model_name,
).outputs["endpoint_name"]
pipeline_func=train_tabular_regression_model_using_PyTorch_pipeline
# %% Pipeline submission
if __name__ == '__main__':
from google.cloud import aiplatform
aiplatform.PipelineJob.from_pipeline_func(pipeline_func=pipeline_func).submit()
@@ -0,0 +1,97 @@
# python3 -m pip install "kfp<2.0.0" "google-cloud-aiplatform>=1.16.0" --upgrade --quiet
from kfp import components
# %% Loading components
download_from_gcs_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/storage/download/component.yaml")
select_columns_using_Pandas_on_CSV_data_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/pandas/Select_columns/in_CSV_format/component.yaml")
fill_all_missing_values_using_Pandas_on_CSV_data_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/pandas/Fill_all_missing_values/in_CSV_format/component.yaml")
split_rows_into_subsets_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/dataset_manipulation/Split_rows_into_subsets/in_CSV/component.yaml")
create_fully_connected_tensorflow_network_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/tensorflow/Create_fully_connected_network/component.yaml")
train_model_using_Keras_on_CSV_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/tensorflow/Train_model_using_Keras/on_CSV/component.yaml")
predict_with_TensorFlow_model_on_CSV_data_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/tensorflow/Predict/on_CSV/component.yaml")
upload_Tensorflow_model_to_Google_Cloud_Vertex_AI_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/Vertex_AI/Models/Upload_Tensorflow_model/component.yaml")
deploy_model_to_endpoint_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/Vertex_AI/Models/Deploy_to_endpoint/component.yaml")
# %% Pipeline definition
def train_tabular_regression_model_using_Tensorflow_pipeline():
dataset_gcs_uri = "gs://ml-pipeline-dataset/Chicago_taxi_trips/chicago_taxi_trips_2019-01-01_-_2019-02-01_limit=10000.csv"
feature_columns = ["trip_seconds", "trip_miles", "pickup_community_area", "dropoff_community_area", "fare", "tolls", "extras"] # Excluded "trip_total"
label_column = "tips"
training_set_fraction = 0.8
# Deploying the model might incur additional costs over time
deploy_model = False
all_columns = [label_column] + feature_columns
dataset = download_from_gcs_op(
gcs_path=dataset_gcs_uri
).outputs["Data"]
dataset = select_columns_using_Pandas_on_CSV_data_op(
table=dataset,
column_names=all_columns,
).outputs["transformed_table"]
dataset = fill_all_missing_values_using_Pandas_on_CSV_data_op(
table=dataset,
replacement_value="0",
# # Optional:
# column_names=None, # =[...]
).outputs["transformed_table"]
split_task = split_rows_into_subsets_op(
table=dataset,
fraction_1=training_set_fraction,
)
training_data = split_task.outputs["split_1"]
testing_data = split_task.outputs["split_2"]
network = create_fully_connected_tensorflow_network_op(
input_size=len(feature_columns),
# Optional:
hidden_layer_sizes=[10],
activation_name="elu",
# output_activation_name=None,
# output_size=1,
).outputs["model"]
model = train_model_using_Keras_on_CSV_op(
training_data=training_data,
model=network,
label_column_name=label_column,
# Optional:
#loss_function_name="mean_squared_error",
number_of_epochs=10,
#learning_rate=0.1,
#optimizer_name="Adadelta",
#optimizer_parameters={},
#batch_size=32,
metric_names=["mean_absolute_error"],
#random_seed=0,
).outputs["trained_model"]
predictions = predict_with_TensorFlow_model_on_CSV_data_op(
dataset=testing_data,
model=model,
# label_column_name needs to be set when doing prediction on a dataset that has labels
label_column_name=label_column,
# Optional:
# batch_size=1000,
).outputs["predictions"]
vertex_model_name = upload_Tensorflow_model_to_Google_Cloud_Vertex_AI_op(
model=model,
).outputs["model_name"]
# Deploying the model might incur additional costs over time
if deploy_model:
vertex_endpoint_name = deploy_model_to_endpoint_op(
model_name=vertex_model_name,
).outputs["endpoint_name"]
pipeline_func=train_tabular_regression_model_using_Tensorflow_pipeline
# %% Pipeline submission
if __name__ == '__main__':
from google.cloud import aiplatform
aiplatform.PipelineJob.from_pipeline_func(pipeline_func=pipeline_func).submit()
@@ -0,0 +1,85 @@
# python3 -m pip install "kfp<2.0.0" "google-cloud-aiplatform>=1.16.0" --upgrade --quiet
from kfp import components
# %% Loading components
download_from_gcs_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/storage/download/component.yaml")
select_columns_using_Pandas_on_CSV_data_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/pandas/Select_columns/in_CSV_format/component.yaml")
fill_all_missing_values_using_Pandas_on_CSV_data_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/pandas/Fill_all_missing_values/in_CSV_format/component.yaml")
split_rows_into_subsets_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/dataset_manipulation/Split_rows_into_subsets/in_CSV/component.yaml")
train_XGBoost_model_on_CSV_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/XGBoost/Train/component.yaml")
xgboost_predict_on_CSV_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/XGBoost/Predict/component.yaml")
upload_XGBoost_model_to_Google_Cloud_Vertex_AI_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/Vertex_AI/Models/Upload_XGBoost_model/component.yaml")
deploy_model_to_endpoint_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/Vertex_AI/Models/Deploy_to_endpoint/component.yaml")
# %% Pipeline definition
def train_tabular_regression_model_using_XGBoost_pipeline():
dataset_gcs_uri = "gs://ml-pipeline-dataset/Chicago_taxi_trips/chicago_taxi_trips_2019-01-01_-_2019-02-01_limit=10000.csv"
feature_columns = ["trip_seconds", "trip_miles", "pickup_community_area", "dropoff_community_area", "fare", "tolls", "extras"] # Excluded "trip_total"
label_column = "tips"
training_set_fraction = 0.8
# Deploying the model might incur additional costs over time
deploy_model = False
all_columns = [label_column] + feature_columns
dataset = download_from_gcs_op(
gcs_path=dataset_gcs_uri
).outputs["Data"]
dataset = select_columns_using_Pandas_on_CSV_data_op(
table=dataset,
column_names=all_columns,
).outputs["transformed_table"]
dataset = fill_all_missing_values_using_Pandas_on_CSV_data_op(
table=dataset,
replacement_value="0",
# # Optional:
# column_names=None, # =[...]
).outputs["transformed_table"]
split_task = split_rows_into_subsets_op(
table=dataset,
fraction_1=training_set_fraction,
)
training_data = split_task.outputs["split_1"]
testing_data = split_task.outputs["split_2"]
model = train_XGBoost_model_on_CSV_op(
training_data=training_data,
label_column_name=label_column,
# Optional:
#starting_model=None,
#num_iterations=10,
#booster_params={},
#objective="reg:squarederror",
#booster="gbtree",
#learning_rate=0.3,
#min_split_loss=0,
#max_depth=6,
).outputs["model"]
# Predicting on the testing data
predictions = xgboost_predict_on_CSV_op(
data=testing_data,
model=model,
# label_column needs to be set when doing prediction on a dataset that has labels
label_column_name=label_column,
).outputs["predictions"]
vertex_model_name = upload_XGBoost_model_to_Google_Cloud_Vertex_AI_op(
model=model,
).outputs["model_name"]
# Deploying the model might incur additional costs over time
if deploy_model:
vertex_endpoint_name = deploy_model_to_endpoint_op(
model_name=vertex_model_name,
).outputs["endpoint_name"]
pipeline_func = train_tabular_regression_model_using_XGBoost_pipeline
# %% Pipeline submission
if __name__ == '__main__':
from google.cloud import aiplatform
aiplatform.PipelineJob.from_pipeline_func(pipeline_func=pipeline_func).submit()
@@ -0,0 +1,208 @@
# python3 -m pip install "kfp<2.0.0" "google-cloud-aiplatform>=1.16.0" --upgrade --quiet
from kfp import components
# %% Loading components
download_from_gcs_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/storage/download/component.yaml")
select_columns_using_Pandas_on_CSV_data_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/pandas/Select_columns/in_CSV_format/component.yaml")
fill_all_missing_values_using_Pandas_on_CSV_data_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/pandas/Fill_all_missing_values/in_CSV_format/component.yaml")
split_rows_into_subsets_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/dataset_manipulation/Split_rows_into_subsets/in_CSV/component.yaml")
# TensorFlow
create_fully_connected_tensorflow_network_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/tensorflow/Create_fully_connected_network/component.yaml")
train_model_using_Keras_on_CSV_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/tensorflow/Train_model_using_Keras/on_CSV/component.yaml")
predict_with_TensorFlow_model_on_CSV_data_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/tensorflow/Predict/on_CSV/component.yaml")
upload_Tensorflow_model_to_Google_Cloud_Vertex_AI_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/Vertex_AI/Models/Upload_Tensorflow_model/component.yaml")
# PyTorch
create_fully_connected_pytorch_network_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/PyTorch/Create_fully_connected_network/component.yaml")
train_pytorch_model_from_csv_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/PyTorch/Train_PyTorch_model/from_CSV/component.yaml")
create_pytorch_model_archive_with_base_handler_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/PyTorch/Create_PyTorch_Model_Archive/with_base_handler/component.yaml")
upload_PyTorch_model_archive_to_Google_Cloud_Vertex_AI_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/Vertex_AI/Models/Upload_PyTorch_model_archive/component.yaml")
# XGBoost
train_XGBoost_model_on_CSV_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/XGBoost/Train/component.yaml")
xgboost_predict_on_CSV_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/XGBoost/Predict/component.yaml")
upload_XGBoost_model_to_Google_Cloud_Vertex_AI_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/Vertex_AI/Models/Upload_XGBoost_model/component.yaml")
# Scikit-learn
train_linear_regression_model_using_scikit_learn_from_CSV_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/1f5cf6e06409b704064b2086c0a705e4e6b4fcde/community-content/pipeline_components/ML_frameworks/Scikit_learn/Train_linear_regression_model/from_CSV/component.yaml")
upload_Scikit_learn_pickle_model_to_Google_Cloud_Vertex_AI_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/Vertex_AI/Models/Upload_Scikit-learn_pickle_model/component.yaml")
# Vertex AI
deploy_model_to_endpoint_op = components.load_component_from_url("https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/399405402d95f4a011e2d2e967c96f8508ba5688/community-content/pipeline_components/google-cloud/Vertex_AI/Models/Deploy_to_endpoint/component.yaml")
# %% Pipeline definition
def train_tabular_regression_model_using_all_frameworks_pipeline():
dataset_gcs_uri = "gs://ml-pipeline-dataset/Chicago_taxi_trips/chicago_taxi_trips_2019-01-01_-_2019-02-01_limit=10000.csv"
feature_columns = ["trip_seconds", "trip_miles", "pickup_community_area", "dropoff_community_area", "fare", "tolls", "extras"] # Excluded "trip_total"
label_column = "tips"
training_set_fraction = 0.8
# Deploying the model might incur additional costs over time
deploy_model = False
all_columns = [label_column] + feature_columns
dataset = download_from_gcs_op(
gcs_path=dataset_gcs_uri
).outputs["Data"]
dataset = select_columns_using_Pandas_on_CSV_data_op(
table=dataset,
column_names=all_columns,
).outputs["transformed_table"]
dataset = fill_all_missing_values_using_Pandas_on_CSV_data_op(
table=dataset,
replacement_value="0",
# # Optional:
# column_names=None, # =[...]
).outputs["transformed_table"]
split_task = split_rows_into_subsets_op(
table=dataset,
fraction_1=training_set_fraction,
)
training_data = split_task.outputs["split_1"]
testing_data = split_task.outputs["split_2"]
# TensorFlow
tensorflow_network = create_fully_connected_tensorflow_network_op(
input_size=len(feature_columns),
# Optional:
hidden_layer_sizes=[10],
activation_name="elu",
# output_activation_name=None,
# output_size=1,
).outputs["model"]
tensorflow_model = train_model_using_Keras_on_CSV_op(
training_data=training_data,
model=tensorflow_network,
label_column_name=label_column,
# Optional:
#loss_function_name="mean_squared_error",
number_of_epochs=10,
#learning_rate=0.1,
#optimizer_name="Adadelta",
#optimizer_parameters={},
#batch_size=32,
metric_names=["mean_absolute_error"],
#random_seed=0,
).outputs["trained_model"]
tensorflow_predictions = predict_with_TensorFlow_model_on_CSV_data_op(
dataset=testing_data,
model=tensorflow_model,
# label_column_name needs to be set when doing prediction on a dataset that has labels
label_column_name=label_column,
# Optional:
# batch_size=1000,
).outputs["predictions"]
tensorflow_vertex_model_name = upload_Tensorflow_model_to_Google_Cloud_Vertex_AI_op(
model=tensorflow_model,
).outputs["model_name"]
# Deploying the model might incur additional costs over time
if deploy_model:
tensorflow_vertex_endpoint_name = deploy_model_to_endpoint_op(
model_name=tensorflow_vertex_model_name,
).outputs["endpoint_name"]
# PyTorch
pytorch_network = create_fully_connected_pytorch_network_op(
input_size=len(feature_columns),
# Optional:
hidden_layer_sizes=[10],
activation_name="elu",
# output_activation_name=None,
# output_size=1,
).outputs["model"]
pytorch_model = train_pytorch_model_from_csv_op(
model=pytorch_network,
training_data=training_data,
label_column_name=label_column,
# Optional:
#loss_function_name="mse_loss",
#number_of_epochs=1,
#learning_rate=0.1,
#optimizer_name="Adadelta",
#optimizer_parameters={},
#batch_size=32,
#batch_log_interval=100,
#random_seed=0,
).outputs["trained_model"]
pytorch_model_archive = create_pytorch_model_archive_with_base_handler_op(
model=pytorch_model,
# Optional:
# model_name="model",
# model_version="1.0",
).outputs["Model archive"]
pytorch_vertex_model_name = upload_PyTorch_model_archive_to_Google_Cloud_Vertex_AI_op(
model_archive=pytorch_model_archive,
).outputs["model_name"]
# Deploying the model might incur additional costs over time
if deploy_model:
pytorch_vertex_endpoint_name = deploy_model_to_endpoint_op(
model_name=pytorch_vertex_model_name,
).outputs["endpoint_name"]
# XGBoost
xgboost_model = train_XGBoost_model_on_CSV_op(
training_data=training_data,
label_column_name=label_column,
# Optional:
#starting_model=None,
#num_iterations=10,
#booster_params={},
#objective="reg:squarederror",
#booster="gbtree",
#learning_rate=0.3,
#min_split_loss=0,
#max_depth=6,
).outputs["model"]
# Predicting on the testing data
xgboost_predictions = xgboost_predict_on_CSV_op(
data=testing_data,
model=xgboost_model,
# label_column needs to be set when doing prediction on a dataset that has labels
label_column_name=label_column,
).outputs["predictions"]
xgboost_vertex_model_name = upload_XGBoost_model_to_Google_Cloud_Vertex_AI_op(
model=xgboost_model,
).outputs["model_name"]
# Deploying the model might incur additional costs over time
if deploy_model:
xgboost_vertex_endpoint_name = deploy_model_to_endpoint_op(
model_name=xgboost_vertex_model_name,
).outputs["endpoint_name"]
# Scikit-learn
sklearn_model = train_linear_regression_model_using_scikit_learn_from_CSV_op(
dataset=training_data,
label_column_name=label_column,
).outputs["model"]
sklearn_vertex_model_name = upload_Scikit_learn_pickle_model_to_Google_Cloud_Vertex_AI_op(
model=sklearn_model,
).outputs["model_name"]
# Deploying the model might incur additional costs over time
if deploy_model:
sklearn_vertex_endpoint_name = deploy_model_to_endpoint_op(
model_name=sklearn_vertex_model_name,
).outputs["endpoint_name"]
pipeline_func=train_tabular_regression_model_using_all_frameworks_pipeline
# %% Pipeline submission
if __name__ == '__main__':
from google.cloud import aiplatform
aiplatform.PipelineJob.from_pipeline_func(pipeline_func=pipeline_func).submit()
@@ -68,7 +68,7 @@
"\n",
"<a href=\"https://storage.googleapis.com/amy-jo/images/mp/beans.png\" target=\"_blank\"><img src=\"https://storage.googleapis.com/amy-jo/images/mp/beans.png\" width=\"95%\"/></a>\n",
"\n",
"Learn more about [Vertex AI Pipelines](https://cloud.google.com/vertex-ai/docs/pipelines/introduction) and [AutoML components](https://cloud.google.com/vertex-ai/docs/pipelines/vertex-automl-component)."
"Learn more about [Vertex AI Pipelines](https://cloud.google.com/vertex-ai/docs/pipelines/introduction) and [AutoML components](https://cloud.google.com/vertex-ai/docs/pipelines/vertex-automl-component). Learn more about [Classification for tabular data](https://cloud.google.com/vertex-ai/docs/tabular-data/classification-regression/overview)."
]
},
{
@@ -65,7 +65,7 @@
"\n",
"This tutorial demonstrates how to use Vertex AI Pipelines with pre-built Google Cloud Pipeline Components for custom training.\n",
"\n",
"Learn more about [Vertex AI Pipelines](https://cloud.google.com/vertex-ai/docs/pipelines/introduction) and [Vertex AI Training components](https://cloud.google.com/vertex-ai/docs/training/create-training-pipeline)."
"Learn more about [Vertex AI Pipelines](https://cloud.google.com/vertex-ai/docs/pipelines/introduction) and [Custom training components](https://cloud.google.com/vertex-ai/docs/training/create-training-pipeline)."
]
},
{
@@ -63,7 +63,7 @@
"\n",
"This notebook shows how to use the components defined in [`google_cloud_pipeline_components`](https://github.com/kubeflow/pipelines/tree/master/components/google-cloud) to build an AutoML tabular regression workflow on Vertex AI Pipelines.\n",
"\n",
"Learn more about [Vertex AI Pipelines](https://cloud.google.com/vertex-ai/docs/pipelines/introduction) and [AutoML components](https://cloud.google.com/vertex-ai/docs/pipelines/vertex-automl-component)."
"Learn more about [Vertex AI Pipelines](https://cloud.google.com/vertex-ai/docs/pipelines/introduction) and [AutoML components](https://cloud.google.com/vertex-ai/docs/pipelines/vertex-automl-component). Learn more about [Regression for tabular data](https://cloud.google.com/vertex-ai/docs/tabular-data/classification-regression/overview)."
]
},
{
@@ -784,7 +784,7 @@
"outputs": [],
"source": [
"# Setup\n",
"SPARK_VERSION = \"3.1.2\"\n",
"DATAPROC_RUNTIME_VERSION = \"1.1.3\"\n",
"SRC = path(\"src\")\n",
"BUILD_PATH = path(\"build\")\n",
"DELIVERABLES = path(\"deliverables\")\n",
@@ -856,7 +856,7 @@
" HPT_BUNDLE_URI,\n",
"]\n",
"HPT_RUNTIME_PROPERTIES = {\n",
" \"spark.jars.packages\": \"ml.combust.mleap:mleap-spark-base_2.12:0.20.0,ml.combust.mleap:mleap-spark_2.12:0.20.0\"\n",
" \"spark.jars.packages\": \"ml.combust.mleap:mleap-spark-base_2.12:0.21.1,ml.combust.mleap:mleap-spark_2.12:0.21.1\"\n",
"}\n",
"\n",
"# Experiment\n",
@@ -2465,6 +2465,7 @@
" deploy_model: bool = DEPLOY_MODEL,\n",
" artifact_uri: str = ARTIFACT_URI,\n",
" serving_image_uri: str = SERVING_IMAGE_URI,\n",
" dataproc_runtime_version: str = DATAPROC_RUNTIME_VERSION,\n",
"):\n",
" from google_cloud_pipeline_components.v1.dataproc import \\\n",
" DataprocPySparkBatchOp\n",
@@ -2485,6 +2486,7 @@
" main_python_file_uri=preprocessing_main_python_file_uri,\n",
" args=build_preprocessing_args_op.output,\n",
" subnetwork_uri=subnetwork_uri,\n",
" runtime_config_version=dataproc_runtime_version,\n",
" ).after(build_preprocessing_args_op)\n",
"\n",
" # build training data args\n",
@@ -2502,6 +2504,7 @@
" main_python_file_uri=training_main_python_file_uri,\n",
" args=build_training_args_op.output,\n",
" subnetwork_uri=subnetwork_uri,\n",
" runtime_config_version=dataproc_runtime_version,\n",
" ).after(build_training_args_op)\n",
"\n",
" evaluate_model_op = evaluate_model(metrics_uri=metrics_path).after(\n",
@@ -2530,6 +2533,8 @@
" args=build_hpt_args_op.output,\n",
" runtime_config_properties=HPT_RUNTIME_PROPERTIES,\n",
" subnetwork_uri=subnetwork_uri,\n",
" # TODO: change to Dataproc Serverless Runtime 1.1.x image when MLeap supports Spark 3.3\n",
" runtime_config_version=\"1.0.29\",\n",
" ).after(model_training_op)\n",
"\n",
" # evaluate condition to upload and deploy model to Vertex AI\n",
@@ -63,7 +63,7 @@
"\n",
"This notebook shows how to use the components defined in [`google_cloud_pipeline_components`](https://github.com/kubeflow/pipelines/tree/master/components/google-cloud) to build a [Vertex AI Pipelines](https://cloud.google.com/vertex-ai/docs/pipelines) workflow that trains a [custom model](https://cloud.google.com/vertex-ai/docs/training/containers-overview), uploads the model as a `Model` resource, creates an `Endpoint` resource, and deploys the `Model` resource to the `Endpoint` resource.\n",
"\n",
"Learn more about [Vertex AI Pipelines](https://cloud.google.com/vertex-ai/docs/pipelines/introduction) and [Vertex AI Training components](https://cloud.google.com/vertex-ai/docs/pipelines/customjob-component)."
"Learn more about [Vertex AI Pipelines](https://cloud.google.com/vertex-ai/docs/pipelines/introduction) and [Custom training components](https://cloud.google.com/vertex-ai/docs/pipelines/customjob-component)."
]
},
{
@@ -94,7 +94,7 @@
"- Compile the KFP pipeline.\n",
"- Execute the KFP pipeline using `Vertex AI Pipelines`\n",
"\n",
"The components are [documented here](https://google-cloud-pipeline-components.readthedocs.io/en/latest/google_cloud_pipeline_components.aiplatform.html#module-google_cloud_pipeline_components.aiplatform).\n",
"The components are [documented here](https://google-cloud-pipeline-components.readthedocs.io/en/google-cloud-pipeline-components-1.0.39/google_cloud_pipeline_components.aiplatform.html).\n",
"(From that page, see also the `CustomPythonPackageTrainingJobRunOp` and `CustomContainerTrainingJobRunOp` components, which similarly run 'custom' training, but as with the related `google.cloud.aiplatform.CustomContainerTrainingJob` and `google.cloud.aiplatform.CustomPythonPackageTrainingJob` methods from the [Vertex AI SDK](https://googleapis.dev/python/aiplatform/latest/aiplatform.html), also upload the trained model)."
]
},
@@ -207,7 +207,7 @@
},
"outputs": [],
"source": [
"PROJECT_ID = \"[your-project-id]\" # @param {type:\"string\"}\n",
"PROJECT_ID = \"andy-1234-221921\" # @param {type:\"string\"}\n",
"\n",
"# Set the project id\n",
"! gcloud config set project {PROJECT_ID}"
@@ -590,8 +590,9 @@
"outputs": [],
"source": [
"@component\n",
"def consumer(text1: str, text2: str, text3: str):\n",
" print(f\"text1: {text1}; text2: {text2}; text3: {text3}\")"
"def consumer(text1: str, text2: str, text3: str) -> str:\n",
" print(f\"text1: {text1}; text2: {text2}; text3: {text3}\")\n",
" return f\"text1: {text1}; text2: {text2}; text3: {text3}\""
]
},
{
-20
View File
@@ -1,20 +0,0 @@
[Training, tuning and deploying a PyTorch text sentiment classification model on Vertex AI](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/official/pytorch/pytorch-text-sentiment-classification-custom-train-deploy.ipynb)
```
Learn to build, train, tune and deploy a PyTorch model on [Vertex AI](https://cloud.
The steps performed include:
- Create training package for the text classification model.
- Train the model with custom training on Vertex AI.
- Check the created model artifacts.
- Create a custom container for predictions.
- Deploy the trained model to a Vertex AI Endpoint using the custom container for predictions.
- Send online prediction requests to the deployed model and validate.
- Clean up the resources created in this notebook.
```
&nbsp;&nbsp;&nbsp;Learn more about [Vertex AI Training](https://cloud.google.com/vertex-ai/docs/training/custom-training).
@@ -64,7 +64,7 @@
"\n",
"Note: you may incur charges for training, prediction, storage or usage of other GCP products in connection with testing this SDK.\n",
"\n",
"Learn more about [AutoML Video](https://cloud.google.com/vertex-ai/docs/tutorials/video-classification-automl/training)."
"Learn more about [Classification for video data](https://cloud.google.com/vertex-ai/docs/training-overview#classification_for_videos)."
]
},
{
@@ -773,7 +773,19 @@
"\n",
"The `run` method when completed returns the `Model` resource.\n",
"\n",
"The execution of the training pipeline can take over 2 hours to complete."
"The execution of the training pipeline can take over 24 hours to complete."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "8f135100f5d9"
},
"outputs": [],
"source": [
"if os.getenv(\"IS_TESTING\"):\n",
" sys.exit(0)"
]
},
{
@@ -64,7 +64,7 @@
"\n",
"Note: You may incur charges for training, prediction, storage or usage of other GCP products in connection with testing this SDK.\n",
"\n",
"Learn more about [Vertex AI Training](https://cloud.google.com/vertex-ai/docs/training/custom-training)."
"Learn more about [Custom training](https://cloud.google.com/vertex-ai/docs/training/custom-training)."
]
},
{
@@ -29,7 +29,7 @@
"id": "JAPoU8Sm5E6e"
},
"source": [
"# Vertex AI Explainations with TabNet models\n",
"# Vertex AI Explanations with TabNet models\n",
"\n",
"<table align=\"left\">\n",
"\n",
@@ -86,7 +86,7 @@
"- TabNet builtin algorithm\n",
"\n",
"The steps performed are:\n",
"* Setup the the project.\n",
"* Setup the project.\n",
"* Download the prediction data of pretrain model onf Syn2 data.\n",
"* Visualize and understand the feature importance based on the masks output.\n",
"* Clean up the resource created by this tutorial."
File diff suppressed because it is too large Load Diff
@@ -91,7 +91,7 @@
"With Vertex AI TensorBoard, you can track, visualize, and compare\n",
"ML experiments and share them with your team.\n",
"\n",
"Learn more about [Vertex AI TensorBoard](https://cloud.google.com/vertex-ai/docs/experiments/tensorboard-overview) and [Vertex AI Training](https://cloud.google.com/vertex-ai/docs/training/custom-training)."
"Learn more about [Vertex AI TensorBoard](https://cloud.google.com/vertex-ai/docs/experiments/tensorboard-overview) and [Custom training](https://cloud.google.com/vertex-ai/docs/training/custom-training)."
]
},
{
@@ -91,7 +91,7 @@
"With Vertex AI TensorBoard, you can track, visualize, and compare\n",
"ML experiments and share them with your team.\n",
"\n",
"Learn more about [Vertex AI TensorBoard](https://cloud.google.com/vertex-ai/docs/experiments/tensorboard-overview) and [Vertex AI Training](https://cloud.google.com/vertex-ai/docs/training/custom-training)."
"Learn more about [Vertex AI TensorBoard](https://cloud.google.com/vertex-ai/docs/experiments/tensorboard-overview) and [Custom training](https://cloud.google.com/vertex-ai/docs/training/custom-training)."
]
},
{
@@ -0,0 +1,794 @@
{
"cells": [
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "ur8xi4C7S06n"
},
"outputs": [],
"source": [
"# Copyright 2022 Google LLC\n",
"#\n",
"# Licensed under the Apache License, Version 2.0 (the \"License\");\n",
"# you may not use this file except in compliance with the License.\n",
"# You may obtain a copy of the License at\n",
"#\n",
"# https://www.apache.org/licenses/LICENSE-2.0\n",
"#\n",
"# Unless required by applicable law or agreed to in writing, software\n",
"# distributed under the License is distributed on an \"AS IS\" BASIS,\n",
"# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.\n",
"# See the License for the specific language governing permissions and\n",
"# limitations under the License."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "JAPoU8Sm5E6e"
},
"source": [
"# Vertex AI TensorBoard Hyperparameter Tuning with the HParams Dashboard\n",
"\n",
"<table align=\"left\">\n",
"\n",
" <td>\n",
" <a href=\"https://colab.research.google.com/github/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/official/tensorboard/tensorboard_hyperparameter_tuning_with_hparams.ipynb\">\n",
" <img src=\"https://cloud.google.com/ml-engine/images/colab-logo-32px.png\" alt=\"Colab logo\"> Run in Colab\n",
" </a>\n",
" </td>\n",
" <td>\n",
" <a href=\"https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/official/tensorboard/tensorboard_hyperparameter_tuning_with_hparams.ipynb\">\n",
" <img src=\"https://cloud.google.com/ml-engine/images/github-logo-32px.png\" alt=\"GitHub logo\">\n",
" View on GitHub\n",
" </a>\n",
" </td>\n",
" <td>\n",
" <a href=\"https://console.cloud.google.com/vertex-ai/workbench/deploy-notebook?download_url=https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/main/notebooks/official/tensorboard/tensorboard_hyperparameter_tuning_with_hparams.ipynb\">\n",
" <img src=\"https://lh3.googleusercontent.com/UiNooY4LUgW_oTvpsNhPpQzsstV5W8F7rYgxgGBD85cWJoLmrOzhVs_ksK_vgx40SHs7jCqkTkCk=e14-rj-sc0xffffff-h130-w32\" alt=\"Vertex AI logo\">\n",
" Open in Vertex AI Workbench\n",
" </a>\n",
" </td>\n",
"</table>"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "24743cf4a1e1"
},
"source": [
"**_NOTE_**: This notebook has been tested in the following environments:\n",
"\n",
"* Python version = 3.8"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "tvgnzT1CKxrO"
},
"source": [
"## Overview\n",
"\n",
"### What is Vertex AI TensorBoard\n",
"\n",
"[Open source TensorBoard](https://www.tensorflow.org/tensorboard/get_started)\n",
"(TB) is a Google open source project for machine learning experiment\n",
"visualization. Vertex AI TensorBoard is an enterprise-ready managed\n",
"version of TensorBoard.\n",
"\n",
"Vertex AI TensorBoard provides various detailed visualizations, including the following:\n",
"\n",
"* Tracking and visualizing metrics, such as loss and accuracy over time.\n",
"* Visualizing model computational graphs (ops and layers).\n",
"* Viewing histograms of weights, biases, or other tensors as they change over time.\n",
"* Projecting embeddings to a lower dimensional space.\n",
"* Displaying image, text, and audio samples.\n",
"\n",
"In addition to the powerful visualizations from\n",
"TensorBoard, Vertex AI TensorBoard provides the following benefits:\n",
"\n",
"* A persistent, shareable link to your experiment's dashboard.\n",
"\n",
"* A searchable list of all experiments in a project.\n",
"\n",
"* Integrations with Vertex AI services for model training.\n",
"\n",
"* Enterprise-grade security, privacy, and compliance.\n",
"\n",
"With Vertex AI TensorBoard, you can track, visualize, and compare\n",
"ML experiments and share them with your team.\n",
"\n",
"Learn more about [Vertex AI TensorBoard](https://cloud.google.com/vertex-ai/docs/experiments/tensorboard-overview)."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "d975e698c9a4"
},
"source": [
"### Objective\n",
"\n",
"This tutorial shows you how to log hyperparameter experiment results in TensorFlow and visualize the results in TensorBoard's Hparams dashboard.\n",
"\n",
"This tutorial uses the following Vertex AI services and resources:\n",
"\n",
"- Vertex AI TensorBoard\n",
"\n",
"The steps performed include:\n",
"\n",
"* Adapt TensorFlow runs to log hyperparameters and metrics.\n",
"* Start runs and log them all under one parent directory.\n",
"* Visualize the results in TensorBoard's HParams dashboard."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "08d289fa873f"
},
"source": [
"### Dataset\n",
"\n",
"This tutorial uses the [FashionMNIST](https://github.com/zalandoresearch/fashion-mnist) dataset.\n"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "aed92deeb4a0"
},
"source": [
"### Costs\n",
"\n",
"This tutorial uses the following billable components of Google Cloud:\n",
"\n",
"* Vertex AI\n",
"\n",
"Learn about [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing),\n",
"and use the [Pricing Calculator](https://cloud.google.com/products/calculator/)\n",
"to generate a cost estimate based on your projected usage."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "1fD9UZaygyPG"
},
"source": [
"## Set up your local development environment\n",
"\n",
"**If you are using Colab or Vertex AI Workbench**, your environment already meets all the requirements to run this notebook. You can skip this step.\n",
"\n",
"Otherwise, make sure your environment meets this notebook's requirements. You need the following:\n",
"\n",
"- Git\n",
"- Python 3\n",
"- virtualenv\n",
"- Jupyter notebook running in a virtual environment with Python 3\n",
"\n",
"To quickly set up your environment to meet the requirements of this tutorial, perform the following:\n",
"\n",
"1. [Install and initialize the SDK](https://cloud.google.com/sdk/docs/).\n",
"\n",
"2. [Install Python 3](https://cloud.google.com/python/setup#installing_python).\n",
"\n",
"3. [Install virtualenv](https://cloud.google.com/python/setup#installing_and_using_virtualenv) and create a virtual environment that uses Python 3 and activate the virtual environment.\n",
"\n",
"4. Install Jupyter by running the following command in a terminal shell:\n",
"<br> `pip3 install jupyter`\n",
"\n",
"5. Launch Jupyter by running the following command in a terminal shell: <br> `jupyter notebook`\n",
"\n",
"6. Open this tutorial notebook in the Jupyter Notebook Dashboard."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "i7EUnXsZhAGF"
},
"source": [
"## Install dependencies\n",
"\n",
"Install the following packages required to run this tutorial notebook."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "th7tWguZiSN2"
},
"outputs": [],
"source": [
"import os\n",
"\n",
"# The Vertex AI Workbench Notebook product has specific requirements\n",
"IS_WORKBENCH_NOTEBOOK = os.getenv(\"DL_ANACONDA_HOME\")\n",
"IS_USER_MANAGED_WORKBENCH_NOTEBOOK = os.path.exists(\n",
" \"/opt/deeplearning/metadata/env_version\"\n",
")\n",
"\n",
"# Vertex AI Notebook requires dependencies to be installed with '--user'\n",
"USER_FLAG = \"\"\n",
"if IS_WORKBENCH_NOTEBOOK:\n",
" USER_FLAG = \"--user\"\n",
"\n",
"! pip3 install --upgrade google-cloud-aiplatform[tensorboard] tensorflow==2.7 {USER_FLAG} -q"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "58707a750154"
},
"source": [
"### Colab only: Uncomment the following cell to restart the kernel."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "f200f10a1da3"
},
"outputs": [],
"source": [
"# Automatically restart kernel after installs so that your environment can access the new packages\n",
"# import IPython\n",
"\n",
"# app = IPython.Application.instance()\n",
"# app.kernel.do_shutdown(True)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "BF1j6f9HApxa"
},
"source": [
"## Before you begin\n",
"\n",
"### Set up your Google Cloud project\n",
"\n",
"**The following steps are required, regardless of your notebook environment.**\n",
"\n",
"1. [Select or create a Google Cloud project](https://console.cloud.google.com/cloud-resource-manager). When you first create an account, you get a $300 free credit towards your compute/storage costs.\n",
"\n",
"2. [Make sure that billing is enabled for your project](https://cloud.google.com/billing/docs/how-to/modify-project).\n",
"\n",
"3. [Enable the Vertex AI API](https://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com).\n",
"\n",
"4. If you are running this notebook locally, install the [Cloud SDK](https://cloud.google.com/sdk)."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "WReHDGG5g0XY"
},
"source": [
"#### Set your project ID\n",
"\n",
"**If you don't know your project ID**, try the following:\n",
"* Run `gcloud config list`.\n",
"* Run `gcloud projects list`.\n",
"* See the support page: [Locate the project ID](https://support.google.com/googleapi/answer/7014113)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "oM1iC_MfAts1"
},
"outputs": [],
"source": [
"PROJECT_ID = \"[your-project-id]\" # @param {type:\"string\"}\n",
"\n",
"# Set the project id\n",
"! gcloud config set project {PROJECT_ID}"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "region"
},
"source": [
"#### Set the region\n",
"\n",
"**Optional**: Update the 'REGION' variable to specify the region that you want to use. Learn more about [Vertex AI regions](https://cloud.google.com/vertex-ai/docs/general/locations)."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "nsN5NJKSu-GU"
},
"outputs": [],
"source": [
"REGION = \"us-central1\" # @param {type: \"string\"}"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "sBCra4QMA2wR"
},
"source": [
"### Authenticate your Google Cloud account\n",
"\n",
"To authenticate your Google Cloud account, follow the instructions for your Jupyter environment:"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "74ccc9e52986"
},
"source": [
"* **Vertex AI Workbench**\n",
"<br>You are already authenticated."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "de775a3773ba"
},
"source": [
"* **Local JupyterLab instance**\n",
"<br>Uncomment and run the following code:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "254614fa0c46"
},
"outputs": [],
"source": [
"# ! gcloud auth login"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "ef21552ccea8"
},
"source": [
"* **Colab**\n",
"<br>Uncomment and run the following code:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "603adbbf0532"
},
"outputs": [],
"source": [
"# from google.colab import auth\n",
"\n",
"# auth.authenticate_user()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "960505627ddf"
},
"source": [
"### Import libraries"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "PyQmSRbKA8r-"
},
"outputs": [],
"source": [
"from google.cloud import aiplatform"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "init_aip:mbsdk,all"
},
"source": [
"### Initialize the Vertex AI SDK for Python\n",
"\n",
"Initialize the Vertex AI SDK for Python for your project."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "KllitKlIu-GW"
},
"outputs": [],
"source": [
"aiplatform.init(project=PROJECT_ID, location=REGION)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "WjWD61gONRkw"
},
"source": [
"### Load TensorBoard and TensorFlow components\n",
"\n",
"Load the TensorBoard notebook extension and import TensorFlow and the TensorBoard HParams plugin.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "KSayPNqxfJC_"
},
"outputs": [],
"source": [
"# Load the TensorBoard notebook extension\n",
"%load_ext tensorboard\n",
"\n",
"# Clear any logs from previous runs\n",
"!rm -rf ./logs/\n",
"\n",
"# Import TensorFlow and the TensorBoard HParams plugin\n",
"import tensorflow as tf\n",
"from tensorboard.plugins.hparams import api as hp"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "KJ4zE7rYfcvb"
},
"source": [
"### Download dataset\n",
"\n",
"Download the [FashionMNIST](https://github.com/zalandoresearch/fashion-mnist) dataset and scale it."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "vHME9wnnfiMr"
},
"outputs": [],
"source": [
"fashion_mnist = tf.keras.datasets.fashion_mnist\n",
"\n",
"(x_train, y_train), (x_test, y_test) = fashion_mnist.load_data()\n",
"x_train, x_test = x_train / 255.0, x_test / 255.0"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "ofGSMru5r4kP"
},
"source": [
"## Set up the experiment\n",
"\n",
"Run an experiment by specifying values for the following hyperparameters:\n",
"\n",
"* Number of units in the first dense layer\n",
"* Dropout rate in the dropout layer\n",
"* Optimizer\n",
"\n",
"Specify the hyperparameter values for the experiment in TensorBoard.\n",
"\n",
"*Optional*: For more fine grained filtering of hyperparameters in the UI, provide domain information and specify which metrics should be displayed."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "IG5sPLBAcDRy"
},
"outputs": [],
"source": [
"HP_NUM_UNITS = hp.HParam(\"num_units\", hp.Discrete([16, 32]))\n",
"HP_DROPOUT = hp.HParam(\"dropout\", hp.RealInterval(0.1, 0.2))\n",
"HP_OPTIMIZER = hp.HParam(\"optimizer\", hp.Discrete([\"adam\", \"sgd\"]))\n",
"\n",
"METRIC_ACCURACY = \"accuracy\"\n",
"\n",
"with tf.summary.create_file_writer(\"logs/hparam_tuning\").as_default():\n",
" hp.hparams_config(\n",
" hparams=[HP_NUM_UNITS, HP_DROPOUT, HP_OPTIMIZER],\n",
" metrics=[hp.Metric(METRIC_ACCURACY, display_name=\"Accuracy\")],\n",
" )"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "cLNgBNA6srlk"
},
"source": [
"## Adapt TensorFlow runs to log hyperparameters and metrics\n",
"\n",
"The model will be quite simple: two dense layers with a dropout layer between them. The training code will look familiar, although the hyperparameters are no longer hardcoded. Instead, the hyperparameters are provided in an `hparams` dictionary and used throughout the training function:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "C-RSsrF4u-Fq"
},
"outputs": [],
"source": [
"def train_test_model(hparams):\n",
" model = tf.keras.models.Sequential(\n",
" [\n",
" tf.keras.layers.Flatten(),\n",
" tf.keras.layers.Dense(hparams[HP_NUM_UNITS], activation=tf.nn.relu),\n",
" tf.keras.layers.Dropout(hparams[HP_DROPOUT]),\n",
" tf.keras.layers.Dense(10, activation=tf.nn.softmax),\n",
" ]\n",
" )\n",
" model.compile(\n",
" optimizer=hparams[HP_OPTIMIZER],\n",
" loss=\"sparse_categorical_crossentropy\",\n",
" metrics=[\"accuracy\"],\n",
" )\n",
"\n",
" model.fit(\n",
" x_train, y_train, epochs=1\n",
" ) # Run with 1 epoch to speed things up for demo purposes\n",
" _, accuracy = model.evaluate(x_test, y_test)\n",
" return accuracy"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "Esz3uqqCvLoK"
},
"source": [
"For each run, log an hparams summary with the hyperparameters and final accuracy:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "HwR1PAv1vPER"
},
"outputs": [],
"source": [
"def run(run_dir, hparams):\n",
" with tf.summary.create_file_writer(run_dir).as_default():\n",
" hp.hparams(hparams) # record the values used in this trial\n",
" accuracy = train_test_model(hparams)\n",
" tf.summary.scalar(METRIC_ACCURACY, accuracy, step=1)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "0V_8soFFvU7b"
},
"source": [
"## Start runs and log them all under one parent directory\n",
"\n",
"You can now try multiple experiments, training each one with a different set of hyperparameters.\n",
"\n",
"For simplicity, use a grid search: try all combinations of the discrete parameters and just the lower and upper bounds of the real-valued parameter. For more complex scenarios, it might be more effective to choose each hyperparameter value randomly (this is called a random search). There are more advanced methods that can be used.\n",
"\n",
"Run a few experiments, which will take a few minutes:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "6r2oO_PVvbdL"
},
"outputs": [],
"source": [
"session_num = 0\n",
"\n",
"for num_units in HP_NUM_UNITS.domain.values:\n",
" for dropout_rate in (HP_DROPOUT.domain.min_value, HP_DROPOUT.domain.max_value):\n",
" for optimizer in HP_OPTIMIZER.domain.values:\n",
" hparams = {\n",
" HP_NUM_UNITS: num_units,\n",
" HP_DROPOUT: dropout_rate,\n",
" HP_OPTIMIZER: optimizer,\n",
" }\n",
" run_name = \"run-%d\" % session_num\n",
" print(\"--- Starting trial: %s\" % run_name)\n",
" print({h.name: hparams[h] for h in hparams})\n",
" run(\"logs/hparam_tuning/\" + run_name, hparams)\n",
" session_num += 1"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "6FJJwCclvslF"
},
"source": [
"## Visualize the results in Vertex AI TensorBoard's HParams tab"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "BkbB5GEI3Ge3"
},
"source": [
"### Create Vertex AI Tensorboard\n",
"A Vertex AI TensorBoard instance, which is a regionalized resource storing your Vertex AI TensorBoard experiments, must be created before the experiments can be visualized. You can create multiple instances in a project. [documentation instructions](https://cloud.google.com/vertex-ai/docs/experiments/tensorboard-overview).\n",
"\n",
"Create a TensorBoard instance to be used by the training job."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "lQ-d3j-I3ZWV"
},
"outputs": [],
"source": [
"TENSORBOARD_NAME = \"[your-tensorboard-name]\" # @param {type:\"string\"}\n",
"\n",
"if (\n",
" TENSORBOARD_NAME == \"\"\n",
" or TENSORBOARD_NAME is None\n",
" or TENSORBOARD_NAME == \"[your-tensorboard-name]\"\n",
"):\n",
" TENSORBOARD_NAME = PROJECT_ID + \"-tb-\"\n",
"\n",
"tensorboard = aiplatform.Tensorboard.create(\n",
" display_name=TENSORBOARD_NAME, project=PROJECT_ID, location=REGION\n",
")\n",
"TENSORBOARD_RESOURCE_NAME = tensorboard.gca_resource.name\n",
"print(\"TensorBoard resource name:\", TENSORBOARD_RESOURCE_NAME)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "27rERDqeJ2nE"
},
"source": [
"Set your TensorBoard Experiment name."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "4OU4TMtFCn0_"
},
"outputs": [],
"source": [
"from datetime import datetime\n",
"\n",
"EXPERIMENT_NAME = \"[your-experiment-run-name]\" # @param {type:\"string\"}\n",
"\n",
"if (\n",
" EXPERIMENT_NAME == \"\"\n",
" or EXPERIMENT_NAME is None\n",
" or EXPERIMENT_NAME == \"[your-experiment-run-name]\"\n",
"):\n",
" EXPERIMENT_NAME = \"experiment\" + datetime.now().strftime(\"%H-%M-%S\")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "f1D2oU3K8Ys0"
},
"source": [
"Upload the log to your Vertex AI TensorBoard"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "TyXFVQuRv0-X"
},
"outputs": [],
"source": [
"!tb-gcp-uploader --one_shot=True --tensorboard_resource_name=$TENSORBOARD_RESOURCE_NAME --logdir=\"logs/hparam_tuning/\" --experiment_name=$EXPERIMENT_NAME"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "OFe3qRyh9Wjl"
},
"source": [
"Click the generated TensorBoard link and click on \"HParams\" at the top.\n",
"\n",
"The left pane of the dashboard provides filtering capabilities that are active across all the views in the HParams dashboard:\n",
"\n",
"- Filter which hyperparameters/metrics are shown in the dashboard\n",
"- Filter which hyperparameter/metrics values are shown in the dashboard\n",
"- Filter on run status (running, success, ...)\n",
"- Sort by hyperparameter/metric in the table view\n",
"- Number of session groups to show (useful for performance when there are many experiments)\n",
"\n",
"The HParams dashboard has three different views, with various useful information:\n",
"\n",
"* The **Table View** lists the runs, their hyperparameters, and their metrics.\n",
"* The **Parallel Coordinates View** shows each run as a line going through an axis for each hyperparemeter and metric. Click and drag the mouse on any axis to mark a region which will highlight only the runs that pass through it. This can be useful for identifying which groups of hyperparameters are most important. The axes themselves can be re-ordered by dragging them.\n",
"* The **Scatter Plot View** shows plots comparing each hyperparameter/metric with each metric. This can help identify correlations. Click and drag to select a region in a specific plot and highlight those sessions across the other plots.\n",
"\n",
"A table row, a parallel coordinates line, and a scatter plot market can be clicked to see a plot of the metrics as a function of training steps for that session (although in this tutorial only one step is used for each run)."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "TpV-iwP9qw9c"
},
"source": [
"## Cleaning up\n",
"\n",
"To clean up all Google Cloud resources used in this project, you can [delete the Google Cloud\n",
"project](https://cloud.google.com/resource-manager/docs/creating-managing-projects#shutting_down_projects) you used for the tutorial.\n",
"\n",
"Otherwise, you can delete the individual resources you created in this tutorial:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "sx_vKniMq9ZX"
},
"outputs": [],
"source": [
"import os\n",
"\n",
"# Delete endpoint resource\n",
"# e.g. `endpoint.delete()`\n",
"\n",
"# Delete model resource\n",
"# e.g. `model.delete()`\n",
"\n",
"# Delete Cloud Storage objects that were created\n",
"delete_bucket = False\n",
"if delete_bucket or os.getenv(\"IS_TESTING\"):\n",
" ! gsutil -m rm -r $BUCKET_URI"
]
}
],
"metadata": {
"colab": {
"name": "tensorboard_hyperparameter_tuning_with_hparams.ipynb",
"toc_visible": true
},
"kernelspec": {
"display_name": "Python 3",
"name": "python3"
}
},
"nbformat": 4,
"nbformat_minor": 0
}
@@ -0,0 +1,869 @@
{
"cells": [
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "ur8xi4C7S06n"
},
"outputs": [],
"source": [
"# Copyright 2022 Google LLC\n",
"#\n",
"# Licensed under the Apache License, Version 2.0 (the \"License\");\n",
"# you may not use this file except in compliance with the License.\n",
"# You may obtain a copy of the License at\n",
"#\n",
"# https://www.apache.org/licenses/LICENSE-2.0\n",
"#\n",
"# Unless required by applicable law or agreed to in writing, software\n",
"# distributed under the License is distributed on an \"AS IS\" BASIS,\n",
"# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.\n",
"# See the License for the specific language governing permissions and\n",
"# limitations under the License."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "l2mMvIUG9meX"
},
"source": [
"# Profile model training performance using Vertex AI TensorBoard Profiler in custom training with prebuilt container\n",
"\n",
"<table align=\"left\">\n",
"\n",
" <td>\n",
" <a href=\"https://colab.research.google.com/github/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/official/tensorboard/tensorboard_profiler_custom_training_with_prebuilt_container.ipynb\">\n",
" <img src=\"https://cloud.google.com/ml-engine/images/colab-logo-32px.png\" alt=\"Colab logo\"> Run in Colab\n",
" </a>\n",
" </td>\n",
" <td>\n",
" <a href=\"https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/official/tensorboard/tensorboard_profiler_custom_training_with_prebuilt_container.ipynb\">\n",
" <img src=\"https://cloud.google.com/ml-engine/images/github-logo-32px.png\" alt=\"GitHub logo\">\n",
" View on GitHub\n",
" </a>\n",
" </td>\n",
" <td>\n",
" <a href=\"https://console.cloud.google.com/vertex-ai/workbench/deploy-notebook?download_url=https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/main/notebooks/official/tensorboard/tensorboard_profiler_custom_training_with_prebuilt_container.ipynb\">\n",
" <img src=\"https://lh3.googleusercontent.com/UiNooY4LUgW_oTvpsNhPpQzsstV5W8F7rYgxgGBD85cWJoLmrOzhVs_ksK_vgx40SHs7jCqkTkCk=e14-rj-sc0xffffff-h130-w32\" alt=\"Vertex AI logo\">\n",
" Open in Vertex AI Workbench\n",
" </a>\n",
" </td>\n",
"</table>"
]
},
{
"attachments": {},
"cell_type": "markdown",
"metadata": {
"id": "tvgnzT1CKxrO"
},
"source": [
"## Overview\n",
"\n",
"The TensorFlow Profiler is a powerful tool that can help you to diagnose and debug performance bottlenecks, and make your model train faster. This tutorial demonstrates how to enable the TensorBoard Profiler in Vertex AI for custom training with a prebuilt container.\n",
"\n",
"Learn more about [Vertex AI TensorBoard Profiler](https://cloud.google.com/vertex-ai/docs/experiments/tensorboard-profiler)."
]
},
{
"attachments": {},
"cell_type": "markdown",
"metadata": {
"id": "dmfmQL6w84pS"
},
"source": [
"### Objective\n",
"\n",
"In this tutorial, you learn how to enable the TensorBoard Profiler in Vertex AI for custom training jobs with a prebuilt container.\n",
"\n",
"This tutorial uses the following Google Cloud AI services:\n",
"\n",
"- Vertex AI Training\n",
"- Vertex AI TensorBoard\n",
"\n",
"The steps performed include:\n",
"\n",
"- Prepare your custom training code and load your training code as a Python package to a prebuilt container\n",
"- Create and run a custom training job that enables the TensorBoard Profiler\n",
"- View the TensorBoard Profiler dashboard to debug your model training performance\n"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "zfXf0r-K81Y-"
},
"source": [
"### Dataset\n",
"\n",
"The dataset used for this tutorial is the [mnist dataset](https://www.tensorflow.org/datasets/catalog/mnist) from [TensorFlow Datasets](https://www.tensorflow.org/datasets/catalog/overview).\n"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "I3KFLvpq87rs"
},
"source": [
"### Costs\n",
"\n",
"This tutorial uses billable components of Google Cloud:\n",
"\n",
"* Vertex AI\n",
"* Cloud Storage\n",
"\n",
"Learn about [Vertex AI\n",
"pricing](https://cloud.google.com/vertex-ai/pricing) and [Cloud Storage\n",
"pricing](https://cloud.google.com/storage/pricing), and use the [Pricing\n",
"Calculator](https://cloud.google.com/products/calculator/)\n",
"to generate a cost estimate based on your projected usage."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "ze4-nDLfK4pw"
},
"source": [
"## Installation\n",
"\n",
"Install the following packages required to execute this notebook."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "2b4ef9b72d43"
},
"outputs": [],
"source": [
"! pip3 install --upgrade --quiet google-cloud-aiplatform"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "aUw6ibN-n5Za"
},
"source": [
"### Colab only: Uncomment the following cell to restart the kernel."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "FM12wbWhn7w0"
},
"outputs": [],
"source": [
"# Automatically restart kernel after installs so that your environment can access the new packages\n",
"# import IPython\n",
"\n",
"# app = IPython.Application.instance()\n",
"# app.kernel.do_shutdown(True)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "LgFWLeJfoGQu"
},
"source": [
"## Before you begin\n",
"\n",
"### Set up your Google Cloud project\n",
"\n",
"**The following steps are required, regardless of your notebook environment.**\n",
"\n",
"1. [Select or create a Google Cloud project](https://console.cloud.google.com/cloud-resource-manager). When you first create an account, you get a $300 free credit towards your compute/storage costs.\n",
"\n",
"2. [Make sure that billing is enabled for your project](https://cloud.google.com/billing/docs/how-to/modify-project).\n",
"\n",
"3. [Enable the Vertex AI API](https://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com).\n",
"\n",
"4. If you are running this notebook locally, you need to install the [Cloud SDK](https://cloud.google.com/sdk)."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "8ckyxpX_oSzD"
},
"source": [
"#### Set your project ID\n",
"\n",
"**If you don't know your project ID**, try the following:\n",
"* Run `gcloud config list`.\n",
"* Run `gcloud projects list`.\n",
"* See the support page: [Locate the project ID](https://support.google.com/googleapi/answer/7014113)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "zY8DKBoVoVy3"
},
"outputs": [],
"source": [
"PROJECT_ID = \"[your-project-id]\" # @param {type:\"string\"}\n",
"\n",
"# Set the project id\n",
"! gcloud config set project {PROJECT_ID}"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "mSQjVQmMosMl"
},
"source": [
"#### Region\n",
"\n",
"You can also change the `REGION` variable used by Vertex AI. Learn more about [Vertex AI regions](https://cloud.google.com/vertex-ai/docs/general/locations)."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "Se9FWWhLotvB"
},
"outputs": [],
"source": [
"REGION = \"us-central1\" # @param {type:\"string\"}"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "IfJRIMBpo5Pg"
},
"source": [
"### Authenticate your Google Cloud account\n",
"\n",
"Depending on your Jupyter environment, you may have to manually authenticate. Follow the relevant instructions below."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "acFN0s3So9-Y"
},
"source": [
"**1. Vertex AI Workbench**\n",
"* Do nothing as you are already authenticated."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "dQ_mNwuapE5T"
},
"source": [
"**2. Local JupyterLab instance, uncomment and run:**"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "cR_MzpknpGgM"
},
"outputs": [],
"source": [
"# ! gcloud auth login"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "h-MuVI_ypJfw"
},
"source": [
"**3. Colab, uncomment and run:**"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "BeaQlCwMpQUT"
},
"outputs": [],
"source": [
"# from google.colab import auth\n",
"# auth.authenticate_user()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "3ivZkPUjpaFz"
},
"source": [
"**4. Setup service account and permissions**\n",
"\n",
"A service account will be used to create custom training jobs. If you do not want to use your project's Compute Engine service account, set SERVICE_ACCOUNT to another service account ID. You can create a service account by following the [instructions](https://cloud.google.com/iam/docs/creating-managing-service-accounts#creating)."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "vYE3b942wza4"
},
"outputs": [],
"source": [
"SERVICE_ACCOUNT = \"[your-service-account]\" # @param {type:\"string\"}"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "WWIxsCJFCg5Z"
},
"outputs": [],
"source": [
"# Grant Cloud Storage permission.\n",
"! gcloud projects add-iam-policy-binding $PROJECT_ID \\\n",
" --member=\"serviceAccount:$SERVICE_ACCOUNT\" \\\n",
" --role=\"roles/storage.admin\" \\\n",
" --quiet\n",
"\n",
"# Grant AI Platform permission.\n",
"! gcloud projects add-iam-policy-binding $PROJECT_ID \\\n",
" --member=\"serviceAccount:$SERVICE_ACCOUNT\" \\\n",
" --role=\"roles/aiplatform.user\" \\\n",
" --quiet\n",
"\n",
"! gcloud projects get-iam-policy $PROJECT_ID \\\n",
" --filter=bindings.members:serviceAccount:$SERVICE_ACCOUNT"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "OKtKGmr9pfr6"
},
"source": [
"### Create a Cloud Storage bucket\n",
"\n",
"Create a storage bucket to store intermediate artifacts such as datasets."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "In3aQanwYjFB"
},
"outputs": [],
"source": [
"BUCKET_URI = \"gs://your-bucket-name-unique\" # @param {type:\"string\"}"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "GOaOsIjxp0oB"
},
"source": [
"**Only if your bucket doesn't already exist**: Run the following cell to create your Cloud Storage bucket."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "Wn5QiIl2p16e"
},
"outputs": [],
"source": [
"! gsutil mb -l $REGION -p $PROJECT_ID $BUCKET_URI"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "ankcS-vtp7Wv"
},
"source": [
"### Import libraries"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "WffSImMvp-Po"
},
"outputs": [],
"source": [
"from google.cloud import aiplatform"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "OMrAJ8RGqBQu"
},
"source": [
"### Initialize Vertex AI SDK for Python\n",
"\n",
"Initialize the Vertex AI SDK for Python for your project and corresponding bucket."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "AWRzBFExqERG"
},
"outputs": [],
"source": [
"aiplatform.init(project=PROJECT_ID, location=REGION, staging_bucket=BUCKET_URI)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "-ayTbNdi62_t"
},
"source": [
"### Create a TensorBoard instance\n",
"\n",
"A Vertex AI TensorBoard instance, which is a regionalized resource storing your Vertex AI TensorBoard experiments, must be created before the experiments can be visualized. You can create multiple instances in a project. You can use command `gcloud ai tensorboards list` to get a list of your existing TensorBoard instances."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "9c3QrDTZdaxk"
},
"source": [
"#### Set your TensorBoard instance display name\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "azlwb__AX8gs"
},
"outputs": [],
"source": [
"TENSORBOARD_NAME = \"your-tensorboard-unique\" # @param {type:\"string\"}"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "vJrWKK0mY7H7"
},
"source": [
"#### Create a TensorBoard instance\n",
"\n",
"If you don't have a TensorBoard instance, create one by running the following cell:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "JqVNsRFrc_78"
},
"outputs": [],
"source": [
"tensorboard = aiplatform.Tensorboard.create(\n",
" display_name=TENSORBOARD_NAME, project=PROJECT_ID, location=REGION\n",
")\n",
"\n",
"TENSORBOARD_INSTANCE_NAME = tensorboard.resource_name\n",
"print(\"TensorBoard instance name:\", TENSORBOARD_INSTANCE_NAME)"
]
},
{
"attachments": {},
"cell_type": "markdown",
"metadata": {
"id": "yoR29gW2S24w"
},
"source": [
"## Train a model\n",
"\n",
"To train a model using your custom training code, choose one of the following options:\n",
"\n",
"- **Prebuilt container**: Load your custom training code as a Python package to a prebuilt container image from Google Cloud.\n",
"\n",
"- **Custom container**: Create your own container image that contains your custom training code.\n",
"\n",
"In this tutorial, you will train a custom model using a prebuilt container."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "syw3GabNGgJz"
},
"source": [
"### Examine the training package\n",
"\n",
"#### Package layout\n",
"\n",
"Before you start the training, let's take a look at how a Python package is assembled for a custom training job. When extracted, the package contains the following:\n",
"\n",
"- PKG-INFO\n",
"- README.md\n",
"- setup.cfg\n",
"- setup.py\n",
"- trainer\n",
" - \\_\\_init\\_\\_.py\n",
" - task.py\n",
"\n",
"The files `setup.cfg` and `setup.py` are the instructions for installing the package into the operating environment of the docker image."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "b58ZAbysGkRo"
},
"outputs": [],
"source": [
"PYTHON_PACKAGE_APPLICATION_DIR = \"app\"\n",
"\n",
"source_package_file_name = f\"{PYTHON_PACKAGE_APPLICATION_DIR}/dist/trainer-0.1.tar.gz\"\n",
"python_package_gcs_uri = f\"{BUCKET_URI}/trainer-0.1.tar.gz\"\n",
"\n",
"# Make folder for Python training script\n",
"! rm -rf {PYTHON_PACKAGE_APPLICATION_DIR}\n",
"! mkdir {PYTHON_PACKAGE_APPLICATION_DIR}\n",
"\n",
"# Add package information\n",
"! touch {PYTHON_PACKAGE_APPLICATION_DIR}/README.md\n",
"\n",
"# Make the training subfolder\n",
"! mkdir {PYTHON_PACKAGE_APPLICATION_DIR}/trainer\n",
"! touch {PYTHON_PACKAGE_APPLICATION_DIR}/trainer/__init__.py"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "lj7hIeAXGrzg"
},
"outputs": [],
"source": [
"%%writefile ./{PYTHON_PACKAGE_APPLICATION_DIR}/setup.py\n",
"\n",
"from setuptools import find_packages\n",
"from setuptools import setup\n",
"import setuptools\n",
"\n",
"from distutils.command.build import build as _build\n",
"import subprocess\n",
"\n",
"REQUIRED_PACKAGES = [\n",
" 'google-cloud-aiplatform[cloud_profiler]>=1.20.0',\n",
"]\n",
"\n",
"setup(\n",
" install_requires=REQUIRED_PACKAGES,\n",
" packages=find_packages(),\n",
" include_package_data=True,\n",
" name='trainer',\n",
" version='0.1',\n",
" url=\"wwww.google.com\",\n",
" description='Vertex AI | Training | Python Package'\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "hyAwgsoQmaYI"
},
"source": [
"#### Prepare the training script\n",
"\n",
"The file `trainer/task.py` is the Python script for executing the custom training job.\n",
"\n",
"Your training code must be configured to write TensorBoard logs to a Cloud Storage bucket, the location of which Vertex AI Training automatically makes available through a predefined environment variable, `AIP_TENSORBOARD_LOG_DIR`. This can usually be done by providing `os.environ['AIP_TENSORBOARD_LOG_DIR']` as the log directory to the open source TensorBoard log writing APIs. For example, in TensorFlow 2.x, you can use following code to create a `tensorboard_callback`:\n",
"\n",
" tensorboard_callback = tf.keras.callbacks.TensorBoard(\n",
" log_dir=os.environ['AIP_TENSORBOARD_LOG_DIR'],\n",
" histogram_freq=1)\n",
"`AIP_TENSORBOARD_LOG_DIR` is in the `BASE_OUTPUT_DIR` that you provide when creating the custom training job.\n",
"\n",
"To enable Vertex AI TensorBoard Profiler for your training job, add the following to your training script:\n",
"\n",
"Add the cloud_profiler import at your top level imports:\n",
"\n",
" from google.cloud.aiplatform.training_utils import cloud_profiler\n",
"\n",
"\n",
"Initialize the cloud_profiler plugin by adding:\n",
"\n",
"\n",
" cloud_profiler.init()"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "8JCgWW7Au1w8"
},
"outputs": [],
"source": [
"%%writefile ./{PYTHON_PACKAGE_APPLICATION_DIR}/trainer/task.py\n",
"\n",
"import tensorflow as tf\n",
"import argparse\n",
"import os\n",
"import sys, traceback\n",
"from google.cloud.aiplatform.training_utils import cloud_profiler\n",
"\n",
"\"\"\"Train an mnist model and use cloud_profiler for profiling.\"\"\"\n",
"\n",
"def _create_model():\n",
" model = tf.keras.models.Sequential(\n",
" [\n",
" tf.keras.layers.Flatten(input_shape=(28, 28)),\n",
" tf.keras.layers.Dense(128, activation=\"relu\"),\n",
" tf.keras.layers.Dropout(0.2),\n",
" tf.keras.layers.Dense(10),\n",
" ]\n",
" )\n",
" return model\n",
"\n",
"\n",
"def main(args):\n",
" print('Initialize the profiler ...')\n",
" cloud_profiler.init()\n",
" print('The profiler initiated.')\n",
"\n",
" print('Loading and preprocessing data ...')\n",
" mnist = tf.keras.datasets.mnist\n",
"\n",
" (x_train, y_train), (x_test, y_test) = mnist.load_data()\n",
" x_train, x_test = x_train / 255.0, x_test / 255.0\n",
"\n",
" print('Creating and training model ...')\n",
"\n",
" model = _create_model()\n",
" model.compile(\n",
" optimizer=\"adam\",\n",
" loss=tf.keras.losses.sparse_categorical_crossentropy,\n",
" metrics=[\"accuracy\"],\n",
" )\n",
"\n",
" log_dir = \"logs\"\n",
" if 'AIP_TENSORBOARD_LOG_DIR' in os.environ:\n",
" log_dir = os.environ['AIP_TENSORBOARD_LOG_DIR']\n",
"\n",
" print('Setting up the TensorBoard callback ...')\n",
" tensorboard_callback = tf.keras.callbacks.TensorBoard(\n",
" log_dir=log_dir,\n",
" histogram_freq=1)\n",
"\n",
" print('Training model ...')\n",
" model.fit(\n",
" x_train,\n",
" y_train,\n",
" epochs=args.epochs,\n",
" verbose=0,\n",
" callbacks=[tensorboard_callback],\n",
" )\n",
" print('Training completed.')\n",
"\n",
" print('Saving model ...')\n",
"\n",
" model_dir = \"model\"\n",
" if 'AIP_MODEL_DIR' in os.environ:\n",
" model_dir = os.environ['AIP_MODEL_DIR']\n",
" tf.saved_model.save(model, model_dir)\n",
"\n",
" print('Model saved at ' + model_dir)\n",
"\n",
"\n",
"if __name__ == \"__main__\":\n",
" parser = argparse.ArgumentParser()\n",
" parser.add_argument(\n",
" \"--epochs\", type=int, default=100, help=\"Number of epochs to run model.\"\n",
" )\n",
"\n",
" args = parser.parse_args()\n",
" main(args)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "ihYFahRAr6sj"
},
"source": [
"#### Create a source distribution\n",
"\n",
"You create a source distribution with your training application and upload the source distribution to your Cloud Storage bucket."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "-XhccshCHQeb"
},
"outputs": [],
"source": [
"!cd {PYTHON_PACKAGE_APPLICATION_DIR} && python3 setup.py sdist --formats=gztar\n",
"\n",
"!gsutil cp {source_package_file_name} {python_package_gcs_uri}\n",
"\n",
"!gsutil ls -l {python_package_gcs_uri}"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "k4e6OYmimqTR"
},
"source": [
"### Create and run the custom training job\n",
"\n",
"Configure a [custom job](https://cloud.google.com/vertex-ai/docs/training/create-custom-job) with the [pre-built container](https://cloud.google.com/vertex-ai/docs/training/pre-built-containers) image for training code packaged as Python source distribution."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "t8GeVXjWHxuZ"
},
"outputs": [],
"source": [
"JOB_NAME = \"tensorboard-job-unique\"\n",
"MACHINE_TYPE = \"n1-standard-4\"\n",
"TRAIN_IMAGE = \"us-docker.pkg.dev/vertex-ai/training/tf-cpu.2-9:latest\"\n",
"base_output_dir = f\"{BUCKET_URI}/{JOB_NAME}\"\n",
"python_module_name = \"trainer.task\"\n",
"\n",
"EPOCHS = 20\n",
"training_args = [\n",
" \"--epochs=\" + str(EPOCHS),\n",
"]"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "B3JC7T3bH9Vy"
},
"outputs": [],
"source": [
"job = aiplatform.CustomPythonPackageTrainingJob(\n",
" display_name=JOB_NAME,\n",
" python_package_gcs_uri=python_package_gcs_uri,\n",
" python_module_name=python_module_name,\n",
" container_uri=TRAIN_IMAGE,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "51hKGTbU32Eg"
},
"source": [
"#### Run the custom training job\n",
"\n",
"Next, you run the custom job to start the training job by invoking the method `run`.\n",
"\n",
"**NOTE:** When using Vertex AI SDK for Python for submitting a training job, it creates a [training pipeline](https://cloud.google.com/vertex-ai/docs/training/create-training-pipeline) which launches the custom job on Vertex AI Training service."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "oIyfos1rIAx2"
},
"outputs": [],
"source": [
"job.run(\n",
" replica_count=1,\n",
" machine_type=MACHINE_TYPE,\n",
" base_output_dir=base_output_dir,\n",
" tensorboard=TENSORBOARD_INSTANCE_NAME,\n",
" service_account=SERVICE_ACCOUNT,\n",
" args=training_args,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "JkEe2Nb_85UD"
},
"source": [
"## View the TensorBoard Profiler dashboard\n",
"\n",
"When the custom job state switches to `Running`, you can access the Vertex AI TensorBoard Profiler dashboard through the Custom jobs page or the Experiments page on the Google Cloud console.\n",
"\n",
"The Google Cloud guide to [Profile model training performance using Profiler](https://cloud.google.com/vertex-ai/docs/experiments/tensorboard-profiler) provides detailed instructions for accessing the Vertex AI TensorBoard Profiler dashboard and capturing a profiling session.\n"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "TpV-iwP9qw9c"
},
"source": [
"## Cleaning up\n",
"\n",
"To clean up all Google Cloud resources used in this project, you can [delete the Google Cloud\n",
"project](https://cloud.google.com/resource-manager/docs/creating-managing-projects#shutting_down_projects) you used for the tutorial.\n",
"\n",
"Otherwise, you can delete the individual resources you created in this tutorial.\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "WR-ZhQ9XwpRI"
},
"outputs": [],
"source": [
"delete_bucket = False\n",
"\n",
"job.delete()\n",
"tensorboard.delete()\n",
"\n",
"if delete_bucket and \"BUCKET_URI\" in globals():\n",
" ! gsutil -m rm -r $BUCKET_URI"
]
}
],
"metadata": {
"colab": {
"name": "tensorboard_profiler_custom_training_with_prebuilt_container.ipynb",
"toc_visible": true
},
"kernelspec": {
"display_name": "Python 3",
"name": "python3"
}
},
"nbformat": 4,
"nbformat_minor": 0
}
@@ -72,7 +72,7 @@
"\n",
"This tutorial demonstrates how to create a multi-node, distributed image classification using PyTorch on Vertex AI SDK with GPU. This can help your training job scale to handle large amounts of data.\n",
"\n",
"Learn more about [Vertex AI Training](https://cloud.google.com/vertex-ai/docs/training/custom-training)."
"Learn more about [Custom training](https://cloud.google.com/vertex-ai/docs/training/custom-training)."
]
},
{
@@ -33,18 +33,18 @@
"\n",
"<table align=\"left\">\n",
" <td>\n",
" <a href=\"https://colab.research.google.com/github/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/official/pytorch/pytorch-text-sentiment-classification-custom-train-deploy.ipynb\">\n",
" <a href=\"https://colab.research.google.com/github/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/official/training/pytorch-text-sentiment-classification-custom-train-deploy.ipynb\">\n",
" <img src=\"https://cloud.google.com/ml-engine/images/colab-logo-32px.png\" alt=\"Colab logo\"> Run in Colab\n",
" </a>\n",
" </td>\n",
" <td>\n",
" <a href=\"https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/official/pytorch/pytorch-text-sentiment-classification-custom-train-deploy.ipynb\">\n",
" <a href=\"https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/official/training/pytorch-text-sentiment-classification-custom-train-deploy.ipynb\">\n",
" <img src=\"https://cloud.google.com/ml-engine/images/github-logo-32px.png\" alt=\"GitHub logo\">\n",
" View on GitHub\n",
" </a>\n",
" </td>\n",
" <td>\n",
"<a href=\"https://console.cloud.google.com/vertex-ai/workbench/deploy-notebook?download_url=https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/main/notebooks/official/pytorch/pytorch-text-sentiment-classification-custom-train-deploy.ipynb\" target='_blank'>\n",
"<a href=\"https://console.cloud.google.com/vertex-ai/workbench/deploy-notebook?download_url=https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/main/notebooks/official/training/pytorch-text-sentiment-classification-custom-train-deploy.ipynb\" target='_blank'>\n",
" <img src=\"https://lh3.googleusercontent.com/UiNooY4LUgW_oTvpsNhPpQzsstV5W8F7rYgxgGBD85cWJoLmrOzhVs_ksK_vgx40SHs7jCqkTkCk=e14-rj-sc0xffffff-h130-w32\" alt=\"Vertex AI logo\">\n",
" Open in Vertex AI Workbench\n",
" </a>\n",
@@ -65,7 +65,7 @@
"\n",
"You can find more details about the model at [Hugging Face Hub](https://huggingface.co/bert-base-cased). For more notebooks with the state of the art PyTorch/Tensorflow/JAX, you can explore [Hugging FaceNotebooks](https://huggingface.co/transformers/notebooks.html).\n",
"\n",
"Learn more about [Vertex AI Training](https://cloud.google.com/vertex-ai/docs/training/custom-training)."
"Learn more about [Custom training](https://cloud.google.com/vertex-ai/docs/training/custom-training)."
]
},
{
@@ -1301,12 +1301,26 @@
"- `machine_type`: Mahcine type on which the job needs to run.\n",
"- `accelerator_type`: Hardware accelerator type for running the job. One of _ACCELERATOR_TYPE_UNSPECIFIED_,\n",
" _NVIDIA_TESLA_K80_, _NVIDIA_TESLA_P100_, _NVIDIA_TESLA_V100_, _NVIDIA_TESLA_P4_,\n",
" _NVIDIA_TESLA_T4_.\n",
" _NVIDIA_TESLA_T4_, _NVIDIA_TELSA_A100_\n",
"- `accelerator_count`: The number of accelerators to attach to a worker replica.\n",
"- `replica_count`: The number of worker replicas.\n",
"- `args`: Command line arguments to be passed to the Python script.\n",
"\n",
"Learn more about Vertex AI's [Custom Python-Package Trainining](https://cloud.google.com/python/docs/reference/aiplatform/latest/google.cloud.aiplatform.CustomPythonPackageTrainingJob)."
"Learn more about Vertex AI's [Custom Python-Package Trainining](https://cloud.google.com/python/docs/reference/aiplatform/latest/google.cloud.aiplatform.CustomPythonPackageTrainingJob).\n",
"\n",
"*Note*: This training job may take over 24 hours."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "b3709beb73e2"
},
"outputs": [],
"source": [
"if os.getenv(\"IS_TESTING\"):\n",
" sys.exit(0)"
]
},
{
@@ -61,9 +61,9 @@
"source": [
"## Overview\n",
"\n",
"This tutorial shows you how to create a distributed custom training job on Vertex AI that can handle large amounts of training data. \n",
"This tutorial shows you how to create a distributed custom training job on Vertex AI that can handle large amounts of training data.\n",
"\n",
"Learn more about [Vertex AI Training](https://cloud.google.com/vertex-ai/docs/training/custom-training)."
"Learn more about [Custom training](https://cloud.google.com/vertex-ai/docs/training/custom-training)."
]
},
{
@@ -78,8 +78,7 @@
"\n",
"This tutorial uses the following Google Cloud ML services:\n",
"\n",
"- `Vertex AI SDK`\n",
"- `CustomContainerTrainingJob`\n",
"- `Vertex AI Training`\n",
"- `Artifact Registry`\n",
"\n",
"The steps performed include:\n",
@@ -98,7 +97,7 @@
"source": [
"### Dataset\n",
"\n",
"This tutorial uses the <a href=\"https://scikit-learn.org/stable/auto_examples/datasets/plot_iris_dataset.html\">IRIS dataset</a>, which consists of different types of irises.\n"
"This tutorial uses the <a href=\"https://scikit-learn.org/stable/auto_examples/datasets/plot_iris_dataset.html\">IRIS dataset</a>, which predicts the iris species.\n"
]
},
{
@@ -108,7 +107,7 @@
},
"source": [
"### Costs\n",
" \n",
"\n",
"This tutorial uses billable components of Google Cloud:\n",
"\n",
"* Vertex AI\n",
@@ -477,7 +476,7 @@
"id": "Xx_z9JQlrNwG"
},
"source": [
"# Create a custom training Python package \n",
"# Create a custom training Python package\n",
"\n",
"Before you can perform local training, you must a create a training script file and a docker file.\n",
"\n",
@@ -492,17 +491,7 @@
},
"outputs": [],
"source": [
"PYTHON_PACKAGE_APPLICATION_DIR = \"trainer\""
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "yjeHKqHwr4rV"
},
"outputs": [],
"source": [
"PYTHON_PACKAGE_APPLICATION_DIR = \"trainer\"\n",
"!mkdir -p $PYTHON_PACKAGE_APPLICATION_DIR"
]
},
@@ -574,14 +563,28 @@
" \"\"\"\n",
" return subprocess.check_call(cmd, stdout=sys.stdout, stderr=sys.stderr, shell=True)\n",
"\n",
"\n",
"def get_chief_ip(cluster_config_dict):\n",
" ip_address = cluster_config_dict['cluster']['workerpool0'][0].split(\":\")[0]\n",
" if 'workerpool0' in cluster_config_dict['cluster']:\n",
" ip_address = cluster_config_dict['cluster']['workerpool0'][0].split(\":\")[0]\n",
" else:\n",
" # if the job is not distributed, 'chief' will be populated instead of\n",
" # workerpool0.\n",
" ip_address = cluster_config_dict['cluster']['chief'][0].split(\":\")[0]\n",
"\n",
" print('The ip address of workerpool 0 is : {}'.format(ip_address))\n",
" return ip_address\n",
"\n",
"def get_chief_port(cluster_config_dict):\n",
" print(\"The open port is: {}\".format(cluster_config_dict['open_ports'][0]))\n",
" return cluster_config_dict['open_ports'][0]\n",
"\n",
" if \"open_ports\" in cluster_config_dict:\n",
" port = cluster_config_dict['open_ports'][0]\n",
" else:\n",
" # Use any port for the non-distributed job.\n",
" port = 7777\n",
" print(\"The open port is: {}\".format(port))\n",
"\n",
" return port\n",
"\n",
"if __name__ == '__main__':\n",
" cluster_config_str = os.environ.get('CLUSTER_SPEC')\n",
@@ -599,7 +602,7 @@
" proc_scheduler = launch('dask-scheduler --dashboard --dashboard-address 8888 --port {} &'.format(chief_port))\n",
" print('Done the dask scheduler.', flush=True)\n",
"\n",
" client = Client(chief_address)\n",
" client = Client(chief_address, timeout=1200)\n",
" print('Waiting the scheduler to be connected.', flush=True)\n",
" client.wait_for_workers(1)\n",
"\n",
@@ -610,7 +613,7 @@
" wait(X)\n",
" wait(y)\n",
" dtrain = DaskDMatrix(client, X, y)\n",
" \n",
"\n",
" output = xgb.dask.train(client, XGB_PARAMS, dtrain, num_boost_round=100, evals=[(dtrain, 'train')])\n",
" print(\"Output: {}\".format(output), flush=True)\n",
" print(\"Saving file to: {}\".format(MODEL_FILE), flush=True)\n",
@@ -623,6 +626,8 @@
" blob.upload_from_filename(MODEL_FILE)\n",
" print(\"Saved file to: {}/{}\".format(MODEL_DIR, MODEL_FILE), flush=True)\n",
"\n",
" # Waiting 10 mins to connect the Dask dashboard\n",
" time.sleep(60 * 10)\n",
" client.shutdown()\n",
"\n",
" else:\n",
@@ -630,7 +635,10 @@
" client = Client(chief_address, timeout=1200)\n",
" print('client: {}.'.format(client), flush=True)\n",
" launch('dask-worker {}'.format(chief_address))\n",
" print('Done with the dask worker.', flush=True)\n"
" print('Done with the dask worker.', flush=True)\n",
"\n",
" # Waiting 10 mins to connect the Dask dashboard\n",
" time.sleep(60 * 10)\n"
]
},
{
@@ -639,7 +647,8 @@
"id": "MxsT4Vaos2W5"
},
"source": [
"### Write the docker file"
"### Write the docker file\n",
"The docker file is used to build the custom training container and passed to the Vertex Training."
]
},
{
@@ -654,14 +663,20 @@
"FROM us-docker.pkg.dev/vertex-ai/training/tf-cpu.2-9:latest\n",
"WORKDIR /root\n",
"\n",
"# Update the keyring in order to run apt-get update.\n",
"RUN rm -rf /usr/share/keyrings/cloud.google.gpg\n",
"RUN rm -rf /etc/apt/sources.list.d/google-cloud-sdk.list\n",
"RUN curl https://packages.cloud.google.com/apt/doc/apt-key.gpg | sudo apt-key add -\n",
"RUN echo \"deb https://packages.cloud.google.com/apt cloud-sdk main\" | sudo tee -a /etc/apt/sources.list.d/google-cloud-sdk.list\n",
"\n",
"RUN apt-get update\n",
"RUN apt-get install -y telnet netcat iputils-ping net-tools\n",
"RUN python3.8 -m pip install dask==2022.7.1 distributed==2022.7.1 bokeh==2.1.1 dask-cuda --upgrade\n",
"RUN python3.8 -m pip install 'xgboost>=1.4.2' 'dask-ml[complete]==2022.5.27' #'dask[complete]==2022.7,1' --upgrade\n",
"RUN python3.8 -m pip install 'xgboost>=1.4.2' 'dask-ml[complete]==2022.5.27' 'dask[complete]==2022.7.1' --upgrade\n",
"RUN python3.8 -m pip install dask==2022.7.1 distributed==2022.7.1 bokeh==2.4.3 dask-cuda==22.8.0 --upgrade\n",
"RUN python3.8 -m pip install gcsfs --upgrade\n",
"\n",
"\n",
"## Make sure gsutil will use the default service account\n",
"# Make sure gsutil will use the default service account\n",
"RUN echo '[GoogleCompute]\\nservice_account = default' > /etc/boto.cfg\n",
"\n",
"# Copies the trainer code\n",
@@ -736,10 +751,10 @@
},
"outputs": [],
"source": [
"DEPLOY_IMAGE = (\n",
"TRAIN_IMAGE = (\n",
" f\"{REGION}-docker.pkg.dev/\" + PROJECT_ID + f\"/{PRIVATE_REPO}\" + \"/dask_support\"\n",
")\n",
"print(\"Deployment:\", DEPLOY_IMAGE)"
"print(\"Deployment:\", TRAIN_IMAGE)"
]
},
{
@@ -788,8 +803,8 @@
"outputs": [],
"source": [
"if not IS_COLAB:\n",
" ! docker build -t $DEPLOY_IMAGE -f Dockerfile .\n",
" ! docker push $DEPLOY_IMAGE"
" ! docker build -t $TRAIN_IMAGE -f Dockerfile .\n",
" ! docker push $TRAIN_IMAGE"
]
},
{
@@ -812,7 +827,16 @@
"outputs": [],
"source": [
"if IS_COLAB:\n",
" ! gcloud builds submit --timeout=1800s --region={REGION} --tag $DEPLOY_IMAGE"
" ! gcloud builds submit --timeout=1800s --region={REGION} --tag $TRAIN_IMAGE"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "41ZgDYPNfEvt"
},
"source": [
"## Run training job with SDK (Option 1) or with gcloud (Option 2)"
]
},
{
@@ -821,7 +845,7 @@
"id": "PtUycdZhCJvQ"
},
"source": [
"### Initialize Vertex AI SDK"
"### 1.1 Initialize Vertex AI SDK"
]
},
{
@@ -845,7 +869,16 @@
"id": "_GB2j39BCXiy"
},
"source": [
"### Run a Vertex AI SDK CustomContainerTrainingJob"
"### 1.2 Run a Vertex AI SDK CustomContainerTrainingJob"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "7udaU3jxfKs8"
},
"source": [
"You can specify the fields enable_web_access and enable_dashboard_access. The enable_web_access enables the interactive shell for the job and enable_dashboard_access allows the dask dashboard to be accessed."
]
},
{
@@ -860,20 +893,46 @@
"replica_count = 2\n",
"machine_type = \"n1-standard-4\"\n",
"display_name = \"test_display_name\"\n",
"DEPLOY_IMAGE = \"us-docker.pkg.dev/vertex-ai/prediction/tf2-cpu.2-8:latest\"\n",
"\n",
"custom_container_training_job = aiplatform.CustomContainerTrainingJob(\n",
" display_name=display_name,\n",
" model_serving_container_image_uri=\"us-docker.pkg.dev/vertex-ai/prediction/tf2-cpu.2-8:latest\",\n",
" container_uri=DEPLOY_IMAGE,\n",
" model_serving_container_image_uri=DEPLOY_IMAGE,\n",
" container_uri=TRAIN_IMAGE,\n",
")\n",
"\n",
"custom_container_training_job.run(\n",
" base_output_dir=gcs_output_uri_prefix,\n",
" replica_count=replica_count,\n",
" machine_type=machine_type,\n",
" enable_dashboard_access=True,\n",
" enable_web_access=True,\n",
" sync=False,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "ROqdvgBKFGKL"
},
"source": [
"Wait for a few minutes for the Custom Job to start"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "rBYbvim_FJz0"
},
"outputs": [],
"source": [
"import time\n",
"\n",
"time.sleep(60 * 3)"
]
},
{
"cell_type": "code",
"execution_count": null,
@@ -886,6 +945,216 @@
"print(f\"GCS Output URI Prefix: {gcs_output_uri_prefix}\")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "-0EJdcjPyjDK"
},
"source": [
"You can access the link to the Custom Job in the Cloud Console UI here:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "ZMZPrcqsyj0e"
},
"outputs": [],
"source": [
"print(\n",
" f\"Custom Training Job URI: {custom_container_training_job._custom_job_console_uri()}\"\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "7McPKJGxymLU"
},
"source": [
"Once the job is in the state \"RUNNING\", you can access the web access and dashboard access URIs here:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "tiHdwJs1yoE0"
},
"outputs": [],
"source": [
"print(f\"Web Access and Dashboard URIs: {custom_container_training_job.web_access_uris}\")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "9RDh5CzNywIA"
},
"source": [
"The interactive shell has the key with the format \"workerpool0-0\", while the dashboard uri has the key with the format \"workerpool0-0:\" + port number (workerpool0-0:8888 in this example). On the page for your Custom Job in the Cloud Console UI, you can also \"Launch web terminal\" for \"workerpool0-0\" for web access, or click \"Launch web terminal\" for \"workerpool0-0:\" + port number for dashboard access.\n",
"\n",
"Note that you can only access an interactive shell and dashboard while the job is running. If you don't see Launch web terminal in the UI or the URIs in the output of the Web Access and Dashboard URIs command, this might be because Vertex AI hasn't started running your job yet, or because the job has already finished or failed. If the job's Status is Queued or Pending, wait a minute; then try refreshing the page, or trying the command again."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "tVktIbToRpmR"
},
"source": [
"### 2. Run a CustomContainerTraining Job with gcloud"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "uVvxLj8GRsM6"
},
"source": [
"You can also create a training job with the gcloud command. With the gcloud command, you can specify the field enableWebAccess and enableDashboardAccess. The enableWebAccess enables the interactive shell for the job and enableDashboardAccess allows the dask dashboard to be accessed."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "pkOQtyDsRwsS"
},
"outputs": [],
"source": [
"%%bash -s \"$BUCKET_URI/output\" \"$TRAIN_IMAGE\"\n",
"\n",
"cat <<EOF >config.yaml\n",
"enableDashboardAccess: true\n",
"enableWebAccess: true\n",
"# Creates two worker pool. The first worker pool is a chief and the second is\n",
"# a worker.\n",
"workerPoolSpecs:\n",
" - machineSpec:\n",
" machineType: n1-standard-8\n",
" replicaCount: 1\n",
" containerSpec:\n",
" imageUri: $2\n",
" - machineSpec:\n",
" machineType: n1-standard-8\n",
" replicaCount: 1\n",
" containerSpec:\n",
" imageUri: $2\n",
"baseOutputDirectory:\n",
" outputUriPrefix: $1\n",
"EOF\n",
"cat config.yaml"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "d5FLoTWzSNw7"
},
"source": [
"The following command creates a training job."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "1MPj-NnpSQ1U"
},
"outputs": [],
"source": [
"! gcloud ai custom-jobs create --region=us-central1 --config=config.yaml --display-name={display_name}"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "HAYihHT9ff9q"
},
"source": [
"#### Access the dashboard and interactive shell for a gcloud custom job\n",
"\n",
"Once the job is created, you can access the web access URI and dashboard access URI by using the `gcloud ai custom-jobs describe` command to print the field webAccessUris. The interactive shell has the key with the format \"workerpool0-0\", while the dashboard uri has the key with the format \"workerpool0-0:\" + port number (workerpool0-0:8888 in this example).\n",
"\n",
"You also can find the links in the Cloud Console UI. In the Cloud Console UI, in the Vertex AI section, go to Training and then Custom Jobs. Click on the name of your custom training job. On the page for your job, click \"Launch web terminal\" for \"workerpool0-0\" for web access, or click \"Launch web terminal\" for \"workerpool0-0:\" + port number for dashboard access.\n",
"\n",
"Note that you can only access an interactive shell and dashboard while the job is running. If you don't see Launch web terminal in the UI or the URIs in the output of the gcloud command, this might be because Vertex AI hasn't started running your job yet, or because the job has already finished or failed. If the job's Status is Queued or Pending, wait a minute; then try refreshing the page, or trying the gcloud command again."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "FqFwDvWCSYFX"
},
"source": [
"#### Troubleshooting"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "MjbPElukSiZt"
},
"source": [
"The [interactive shell](https://cloud.google.com/vertex-ai/docs/training/monitor-debug-interactive-shell) can be used to debugging the access of the dask dashboard. You can get the dashboard point by the following command."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "z9eNOtYUTzJW"
},
"outputs": [],
"source": [
"# Note the following command should run inside the interactive shell.\n",
"# printenv | grep AIP_DASHBOARD_PORT"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "2SCRLCkpUDNM"
},
"source": [
"Then you can check if there are dashboard monitoring the port."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "fK74qU78ULPS"
},
"outputs": [],
"source": [
"# Note the following command should run inside the interactive shell.\n",
"# netstat -ntlp"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "8ucKHMGFUUF4"
},
"source": [
"You can manually turn up the dashboard instance."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "gJrjiePvUip0"
},
"outputs": [],
"source": [
"# Note the following command should run inside the interactive shell.\n",
"# dask-scheduler --dashboard-address :port_number"
]
},
{
"cell_type": "markdown",
"metadata": {
@@ -919,7 +1188,8 @@
"\n",
"Otherwise, you can delete the individual resources you created in this tutorial:\n",
"\n",
"- Cloud Storage Bucket"
"- Cloud Storage Bucket\n",
"- Cloud Vertex Training Job"
]
},
{
@@ -930,19 +1200,27 @@
},
"outputs": [],
"source": [
"import logging\n",
"import traceback\n",
"\n",
"# Set this to true only if you'd like to delete your bucket\n",
"delete_bucket = False\n",
"\n",
"! gsutil rm -rf $gcs_output_uri_prefix\n",
"\n",
"if delete_bucket or os.getenv(\"IS_TESTING\"):\n",
" ! gsutil rm -r $BUCKET_URI"
" ! gsutil rm -r $BUCKET_URI\n",
"\n",
"try:\n",
" custom_container_training_job.delete()\n",
"except Exception as e:\n",
" logging.error(traceback.format_exc())\n",
" print(e)"
]
}
],
"metadata": {
"colab": {
"collapsed_sections": [],
"name": "xgboost_data_parallel_training_on_cpu_using_dask.ipynb",
"toc_visible": true
},
@@ -89,7 +89,7 @@
"* Model with BigQuery and the ARIMA model\n",
"* Evaluate the model\n",
"* Evaluate the model results using BigQuery ML (on training data)\n",
"* Evalute the model results - MAE, MAPE, MSE, RMSE (on test data)\n",
"* Evaluate the model results - MAE, MAPE, MSE, RMSE (on test data)\n",
"* Use the executor feature"
]
},
@@ -63,7 +63,7 @@
"\n",
"This tutorial shows you how to build, deploy, and analyze predictions from a simple [random forest](https://en.wikipedia.org/wiki/Random_forest) model using tools like scikit-learn, Vertex AI, and the [What-IF Tool (WIT)](https://cloud.google.com/ai-platform/prediction/docs/using-what-if-tool) on a synthetic fraud transaction dataset to solve a financial fraud detection problem.\n",
"\n",
"Learn more about [Vertex AI Workbench](https://cloud.google.com/vertex-ai/docs/workbench/introduction) and [Vertex AI Training](https://cloud.google.com/vertex-ai/docs/training/custom-training)."
"Learn more about [Vertex AI Workbench](https://cloud.google.com/vertex-ai/docs/workbench/introduction) and [Custom training](https://cloud.google.com/vertex-ai/docs/training/custom-training)."
]
},
{