Compare commits

...
Author SHA1 Message Date
Andrew FerlitschandGitHub ccd89d1384 fix: replace os.environ with os.getenv 2022-03-18 15:45:23 -07:00
Andrew FerlitschandGitHub 0980f4e1b7 fix: replace os.environ with os.getenv 2022-03-18 15:42:23 -07:00
Andrew FerlitschandGitHub e5bbae8023 fix: replace os.environ with os.getenv 2022-03-18 15:41:43 -07:00
Andrew FerlitschandGitHub efaa340a98 fix: replace os.environ with os.getenv 2022-03-18 15:40:50 -07:00
Andrew FerlitschandGitHub a132f83e30 fix: replace os.environ with os.getenv 2022-03-18 15:40:07 -07:00
Andrew FerlitschandGitHub e338d7187b fix: replace os.environ with os.getenv 2022-03-18 15:39:11 -07:00
Andrew FerlitschandGitHub 78e3cb6170 fix: replace os.environ with os.getenv 2022-03-18 15:37:56 -07:00
Andrew FerlitschandGitHub 77f81441eb fix: replace os.environ with os.getenv 2022-03-18 15:36:50 -07:00
Andrew FerlitschandGitHub 3b7ca24a9c fix: replace os.environ with os.getenv 2022-03-18 15:35:31 -07:00
daa64efd40 fix: replace os.environ with os.getenv (#393)
* fix: use os.getenv()

* fix: use os.getenv()

* fix: use os.getenv()

* fix: use os.getenv()

* fix: use os.getenv()

Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-03-18 16:44:34 -05:00
Andrew FerlitschandGitHub a37deabe27 Update README.md 2022-03-18 13:15:34 -07:00
Andrew FerlitschandGitHub 6009ef0def Update README.md 2022-03-18 13:14:34 -07:00
Andrew FerlitschandGitHub edc644b0d0 Update README.md 2022-03-18 13:13:15 -07:00
Andrew FerlitschandGitHub 1297af8baf Create README.md 2022-03-18 13:11:38 -07:00
Andrew FerlitschandGitHub b8b1b6675b Update README.md 2022-03-18 13:09:51 -07:00
Andrew FerlitschandGitHub 7d3b7abc44 fix: review updates (#395)
* feat: add CPR notebook

* feat: add CPR notebook

* feat: add CPR notebook

* feat: add CPR notebook

* feat: finalize CPR notebook

* feat: finalize CPR notebook

* feat: notebook for bqml+automl

* feat: notebook for bqml+automl

* review: edits per Erwin review

* review: edits per Erwin review
2022-03-18 12:09:53 -07:00
Rajesh ThallamandGitHub 9a572f298e [community-content] Notebook to demo NVIDIA Triton Inference Server on Vertex AI Prediction (#391)
* Add NVIDIA Triton on Vertex AI Prediction official notebook

* Add NVIDIA Triton on Vertex AI Prediction official notebook

* Add NVIDIA Triton on Vertex AI Prediction official notebook

* Add NVIDIA Triton on Vertex AI Prediction community notebook

* Add NVIDIA Triton on Vertex AI Prediction community notebook

* Add NVIDIA Triton on Vertex AI Prediction community notebook

* Fixes based on feedback to NVIDIA Triton on Vertex AI Prediction community notebook
2022-03-18 11:31:01 -05:00
b53ca9e678 Tabnet (#375)
* Start a new branch for TabNet tutorial.

* Clean version Created using Colaboratory

* Created using Colaboratory

* Remove unused import

* format lint

* Remove unused import

* Created using Colaboratory

* Remove unused import

* Fix the first iteration of reviewing except the image location

* add import

* Update the image to vertex

* Force delete the BQ to avoid waiting

* Add codeowner for TabNet

* Remove - from folder name

Co-authored-by: Long Le <longtle@google.com>
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-03-17 17:20:01 -05:00
Andrew FerlitschandGitHub 9aceec161a Update README.md 2022-03-17 14:50:24 -07:00
Andrew FerlitschandGitHub 98b186ed55 feat: notebook for bqml+automl combined (#390)
* feat: add CPR notebook

* feat: add CPR notebook

* feat: add CPR notebook

* feat: add CPR notebook

* feat: finalize CPR notebook

* feat: finalize CPR notebook

* feat: notebook for bqml+automl

* feat: notebook for bqml+automl
2022-03-17 14:46:29 -07:00
sudarshan-SpringMLandGitHub f23ee1b5a8 Made small changes to Get started vertex vizier notebook (#389)
* modified notebook

* ran linter test

* modified notebook changed copyright licence year

* ran linter test
2022-03-17 10:03:19 -05:00
c6d779f1fc Featurestore colab update (#387)
* update featurestore colab comment

* delete some changes

* delete some changes 2

Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-03-16 21:31:22 -05:00
Andrew FerlitschandGitHub 8908b27b08 feat: finish CPR notebook (#388)
* feat: add CPR notebook

* feat: add CPR notebook

* feat: add CPR notebook

* feat: add CPR notebook

* feat: finalize CPR notebook

* feat: finalize CPR notebook
2022-03-16 14:00:03 -07:00
Andrew FerlitschandGitHub 8947c9b116 feat: more CPR (#385)
* feat: add CPR notebook

* feat: add CPR notebook

* feat: add CPR notebook

* feat: add CPR notebook
2022-03-15 11:56:29 -07:00
Andrew FerlitschandGitHub 9464caac6e Update README.md 2022-03-14 17:40:15 -07:00
Andrew FerlitschandGitHub 98ce91c575 feat: add CPR notebook (#384)
* feat: add CPR notebook

* feat: add CPR notebook
2022-03-14 17:38:47 -07:00
Andrew FerlitschandGitHub 07f8feda3d Update README.md 2022-03-14 11:19:31 -07:00
Andrew FerlitschandGitHub a4e0496ff5 feat: CMEK training (#383)
* feat: add FS from panda

* feat: add FS from panda

* feat: add CMEK example

* feat: add CMEK example
2022-03-14 11:15:35 -07:00
cf162c02c8 chore(deps): update dependency pyupgrade to v2.31.1 (#381)
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-03-14 09:34:58 -05:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>Karl Weinmeister
831aae94df build(deps): bump pillow (#378)
Bumps [pillow](https://github.com/python-pillow/Pillow) from 9.0.0 to 9.0.1.
- [Release notes](https://github.com/python-pillow/Pillow/releases)
- [Changelog](https://github.com/python-pillow/Pillow/blob/main/CHANGES.rst)
- [Commits](https://github.com/python-pillow/Pillow/compare/9.0.0...9.0.1)

---
updated-dependencies:
- dependency-name: pillow
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>

Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-03-14 09:33:36 -05:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>Karl Weinmeister
3fd9f28778 build(deps): bump pillow (#379)
Bumps [pillow](https://github.com/python-pillow/Pillow) from 9.0.0 to 9.0.1.
- [Release notes](https://github.com/python-pillow/Pillow/releases)
- [Changelog](https://github.com/python-pillow/Pillow/blob/main/CHANGES.rst)
- [Commits](https://github.com/python-pillow/Pillow/compare/9.0.0...9.0.1)

---
updated-dependencies:
- dependency-name: pillow
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>

Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-03-14 09:32:20 -05:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
a656e8e2a8 build(deps): bump pillow (#380)
Bumps [pillow](https://github.com/python-pillow/Pillow) from 9.0.0 to 9.0.1.
- [Release notes](https://github.com/python-pillow/Pillow/releases)
- [Changelog](https://github.com/python-pillow/Pillow/blob/main/CHANGES.rst)
- [Commits](https://github.com/python-pillow/Pillow/compare/9.0.0...9.0.1)

---
updated-dependencies:
- dependency-name: pillow
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>

Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2022-03-14 09:30:19 -05:00
Morgan DuandGitHub 2f02152703 fix: pip install google-cloud-aiplatform (#376) 2022-03-10 12:00:08 -06:00
9d08f8ce67 Using BQML 1st-party components and upgrading to 1.0.0 of google-cloud-pipeline-components (#372)
* Using BQML 1st-party components and 1.0.0 of google-cloud-pipeline-components

* Using BQML 1st-party components and upgrading to 1.0.0 of google-cloud-pipeline-components

* Using BQML 1st-party components and upgrading to 1.0.0 of google-cloud-pipeline-components

* Using BQML 1st-party components and upgrading to 1.0.0 of google-cloud-pipeline-components

* Using BQML components and upgrade to 1.0.0 of GCPC

Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-03-09 13:34:12 -06:00
ec6d508793 adds automl-text-sentiment-analysis-online notebook (#293)
* adds automl-text-sentiment-analysis-online notebook

* adds the cleaned up automl-text-sentiment-analysis notebook after running linter test

* adds textual content on what the dataset predicts in the dataset section

* ran the linter test after the update

* adds textual content on what the dataset predicts in the dataset section

* ran the linter test after the update

* corrects the IMPORT_FILE parameter in the notebook

* ran linter test after update

* deletes the source file from the community/sdk folder

* updates the colab, git & workbench links in the notebook

* ran linter test

* updates the license year to 2022 and simplifies the clean-up step for bucket-deletion

* ran linter test

* adds TESTING env condition while deleting the buckets

* ran linter test successfully

Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-03-09 13:09:10 -06:00
WhiteSource RenovateandGitHub 772e35ea75 chore(deps): update dependency nbqa to v1.3.1 (#374) 2022-03-09 08:24:16 -06:00
Andrew FerlitschandGitHub 1d28f886c8 feat: add example of FS values from dataframe (#373)
* feat: add FS from panda

* feat: add FS from panda
2022-03-08 17:44:59 -08:00
Andrew FerlitschandGitHub d3dc8aeb9a fix: missing create dataset schema (#371)
* fix: missing dataset create

* fix: missing dataset create
2022-03-07 11:53:39 -08:00
WhiteSource RenovateandGitHub a07d762934 chore(deps): update dependency nbqa to v1.3.0 (#368) 2022-03-07 09:53:42 -06:00
Andrew FerlitschandGitHub 85ac9e127d update: v1 (#366)
* fix: v1 upgrades

* fix: v1 upgrades

* fix: v1 upgrades

* fix: v1 upgrades

* update: v1

* update: v1
2022-03-04 17:47:56 -08:00
Gal ZahaviandGitHub 011c2823ff Update CODEOWNERS (#361) 2022-03-04 22:43:10 +02:00
Karl WeinmeisterandGitHub f28a94f03f docs: Add visualization of repo structure to README.md (#364) 2022-03-04 12:46:26 -06:00
5ecfc80cb9 Adds minor changes(license year and clean-up step) to Sdk automl video action recognition batch notebook (#355)
* adds the automl-video-action-recognition-notebook

* ran linter test

* fixes the dag variable by replacing with job

* ran linter test

* fixes the dag variable by replacing with job

* ran linter test

* corrects the IMPORT_FILE parameter in the notebook

* ran linter test after update

* updates the colab, git & vertex-ai links

* ran linter test

* updates the license year to 2022 and simplifies the lean-up step for bucket created

* ran linter test

* removes the file from the community folder

* adds the TESTING env condition while deleting the buckets

* ran linter test successfully

* adds TESTING env condition while deleting the bucket

* ran linter test successfully

Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-03-04 12:44:39 -06:00
Andrew FerlitschandGitHub e5e36ba050 fix: upgrades to v1 (#360)
* fix: v1 upgrades

* fix: v1 upgrades

* fix: v1 upgrades

* fix: v1 upgrades
2022-03-03 20:05:36 -08:00
Andrew FerlitschandGitHub 0ad9116d6a fix: v1 upgrades (#359)
* fix: v1 upgrades

* fix: v1 upgrades
2022-03-03 17:49:45 -08:00
0516032443 Update CODEOWNERS (#354)
* Update CODEOWNERS

* Update CODEOWNERS

Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-03-03 10:23:19 -06:00
95d211c90f Fixing minor issues in sdk_automl_text_entity_extraction_online.ipynb (#329)
* modified colab,github,vertexAI links and added vertex logo

* ran linter

* resolved comments

* ran linter

Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-03-03 10:17:48 -06:00
12d6a75ef7 Fixing minor issues in sdk_automl_video_object_tracking_batch.ipynb (#328)
* changed master to main for links and added vertex AI logo

* ran linter

* resolved comments

* ran linter

Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-03-03 10:07:50 -06:00
Andrew FerlitschandGitHub 06153dc373 Update README.md 2022-03-02 11:45:01 -08:00
Andrew FerlitschandGitHub 02afa91fc3 Ml ops 6v3 (#352)
* fix: eval comp improvement

* fix: eval comp improvement

* feat: add to KFP get started
2022-03-02 10:38:42 -08:00
Karl WeinmeisterandGitHub c8b212789f Revert "ci: Apply filter to format_and_lint_job (#348)" (#351)
This reverts commit 95256d3fcf.
2022-03-02 09:44:03 -06:00
Karl WeinmeisterandGitHub 95256d3fcf ci: Apply filter to format_and_lint_job (#348)
* ci: Apply filter to format_and_lint_job

Only run when PR contains a notebook file

* Minor fix
2022-03-02 09:29:38 -06:00
Karl WeinmeisterandGitHub 8a6d174c99 Create README.md for notebooks folder 2022-03-01 19:44:41 -06:00
WhiteSource RenovateandGitHub d36cf7f662 chore(deps): update actions/setup-python action to v3 (#337) 2022-03-01 19:04:32 -06:00
WhiteSource RenovateandGitHub 1b02a542c8 chore(deps): update actions/checkout action to v3 (#347) 2022-03-01 19:02:39 -06:00
Ivan CheungandGitHub 45430bb010 Fixed minor issues (#345) 2022-03-01 14:56:50 -06:00
Ivan CheungandGitHub be2a139ade Delete notebooks/community/feature_store/assets directory 2022-02-28 20:21:17 -05:00
70d77b24f6 feature store e2e with assets (#343)
* A notebook that shows Vertex AI feature store capabilities in a real-world scenario (#296)

* A notebook that shows Vertex AI feature store capabilities in a real-world scenario

* new notebook version

* fix CODEOWNERS

* comment to the feature store monitoring api

* format notebook

* fix CODEOWNERS

* fix CODEOWNERS as required

* new version

* new notebook version

* notebook cleaning

* new update

* add fix to pass lint test

* resolve conflict

* import libraries fix

* update image

* update notebook

* fix comment

* new notebook version

* new notebook and assets

Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>

* Added files in their old folder

* Deleted unneeded file

* Ran linter

* Fixed CODEOWNERS

Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
Co-authored-by: ivanmkc <ivans.mailbox@gmail.com>
2022-02-28 19:23:15 -05:00
Ivan CheungandGitHub 50c25d6d7b Revert "A notebook that shows Vertex AI feature store capabilities in a real-world scenario (#296)" (#338)
This reverts commit 5bcdc0bc64.
2022-02-28 11:01:39 -05:00
5bcdc0bc64 A notebook that shows Vertex AI feature store capabilities in a real-world scenario (#296)
* A notebook that shows Vertex AI feature store capabilities in a real-world scenario

* new notebook version

* fix CODEOWNERS

* comment to the feature store monitoring api

* format notebook

* fix CODEOWNERS

* fix CODEOWNERS as required

* new version

* new notebook version

* notebook cleaning

* new update

* add fix to pass lint test

* resolve conflict

* import libraries fix

* update image

* update notebook

* fix comment

* new notebook version

* new notebook and assets

Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
2022-02-28 07:47:26 -08:00
eff0f95b58 Automl links fix (#330)
* fixing links to open notebook - main and images

* linter test changes

Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-02-27 15:57:25 -06:00
50b53d31bd Jfacevedo endpoints tfhub obj detect (#319)
* Deploying TF Hub object detection model using Vertex endpoints

* Add user to codeowners

* fix path in CODEOWNERS

* clear all outputs

* run linter

* manual lint fix

* fix more linting errors

* order imports in alphabetical order

* run linter

* made changes requested on feedback

* automate fetching endpoint model id

* fix hardcoded value in bash command

* generalize region endpoint and project in bash cell

* retrieve endpoint and model ids programatically

* fix formatting

* run linter

* Remove pipfile

Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-02-27 15:55:54 -06:00
Andrew FerlitschandGitHub b5a56852f3 fix: automl eval component improvement (#333)
* fix: eval comp improvement

* fix: eval comp improvement
2022-02-25 12:07:06 -08:00
b7486e34ad Sdk automl video action recognition batch (#310)
* adds the automl-video-action-recognition-notebook

* ran linter test

* fixes the dag variable by replacing with job

* ran linter test

* fixes the dag variable by replacing with job

* ran linter test

* corrects the IMPORT_FILE parameter in the notebook

* ran linter test after update

* updates the colab, git & vertex-ai links

* ran linter test

Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-02-25 11:31:19 -06:00
3edc5f1425 Notebook template (#326)
* modified vertex AI link

* minor change

* changed master to main and added vertex logo

* ran lint

Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-02-25 10:16:34 -06:00
65b4b73cb5 Add google_cloud_pipeline_components_model_upload_predict_evaluate notebook (#288)
* Add google_cloud_pipeline_components_model_upload_predict_evaluate notebook.ipynb

* format with linter

* add import for tensorflow when in the testing environment

* linter

* fix dependency issues for testing env

* address comments

* eval component  does not output gcp_resources yet, still in experimental

* added location to aip.init

* add deletion for model and batch prediction jobs

* typo, missed a comma.

* linter

* Remove tensorflow import + use gsutil to check if artifacts exist.

Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-02-24 15:36:35 -08:00
Ivan CheungandGitHub 186c08e8c3 Update sdk_matching_engine_for_indexing.ipynb 2022-02-24 17:47:53 -05:00
Ivan CheungandGitHub 7721aa0def Added matching engine with SDK notebook (#327)
* Matching engine

* Added notebook

* Updated CODEOWNERS

* Fixed links

* Added logo

* Renamed Workbench
2022-02-24 14:57:16 -05:00
4987c60e03 updating gcloud usage because a flag was renamed (#320)
Co-authored-by: Yicheng Fang <yichengfang@google.com>
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-02-24 09:33:33 -06:00
Andrew FerlitschandGitHub cd845f7fdd feat: start on model eval notebook (#325)
* fix: bqml export format

* fix: bqml export format

* feat: start on custom model eval

* feat: start on custom model eval
2022-02-23 12:56:01 -08:00
Andrew FerlitschandGitHub 9e84d9e782 fix: bqml doc for exporting model (#324)
* fix: bqml export format

* fix: bqml export format
2022-02-23 12:44:43 -08:00
nayaknishantandGitHub 23c7fcc97f fix: changed bucket URL from pantheon to console (#323)
* moving REGION up

* moving REGION up and csv file name

* fix: changed bucket URL to console
2022-02-23 13:17:12 -06:00
e4024efbc7 feat: add community notebook for Vertex AI SDK Feature Store with Pandas (#311)
* feat: add sdk-feature-store-pandas notebook

* fix: add ldap to codeowners

* feat: add sdk-feature-store-pandas notebook

* fix: add ldap to codeowners

* fix: lint

* fix: addressed feedback

* fix: format

* fix: lint

* Linted

Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
Co-authored-by: ivanmkc <ivans.mailbox@gmail.com>
2022-02-23 10:22:28 -08:00
Ivan CheungandGitHub c4d53108af Matching Engine: Updated location for data (#322)
Switched to gs://cloud-samples-data/vertex-ai/matching_engine/glove-100-angular.hdf5
2022-02-23 12:03:48 -05:00
be8fe3564d Sdk automl video classification batch (#295)
* notebook refresh from vertex ai sdk project batch 1

* successfully ran linter test

* removed global variable import file

* update with linter test changes

* removing community version of dk_automl_video_classification_batch.ipynb

Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
2022-02-22 21:11:42 -08:00
1f95775057 Sdk automl text entity extraction online (#283)
* added notebook

* ran linter

* fix aip not defined error

* ran lint

* resolved git comments

* ran linter

* deleted file in community folder and removed globals

* ran linter

Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
2022-02-22 21:07:55 -08:00
f2a4dd875e Sdk automl video object tracking batch (#281)
* added notebook

* changed folder

* reinstalled linter

* ran linter

* pulled new changes and merged

* resolved comments

* resolved comments

* ran linter

* deleted file in community folder and removed globals in file

* ran linter

Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-02-22 21:06:52 -08:00
Aaron DietzandGitHub 67dd2300c8 updating text (#309) 2022-02-22 17:17:00 -05:00
Aaron DietzandGitHub b5391b06b4 Pricing optimization update (#308)
* updating text

* rename

* rename
2022-02-22 17:13:16 -05:00
Aaron DietzandGitHub 54f2c71c13 updating text (#307) 2022-02-22 16:51:54 -05:00
Aaron DietzandGitHub 6697900126 updating text (#306) 2022-02-22 16:39:46 -05:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
9ff3400b44 build(deps): bump tensorflow (#316)
Bumps [tensorflow](https://github.com/tensorflow/tensorflow) from 2.5.0 to 2.5.3.
- [Release notes](https://github.com/tensorflow/tensorflow/releases)
- [Changelog](https://github.com/tensorflow/tensorflow/blob/master/RELEASE.md)
- [Commits](https://github.com/tensorflow/tensorflow/compare/v2.5.0...v2.5.3)

---
updated-dependencies:
- dependency-name: tensorflow
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>

Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2022-02-18 13:01:06 -06:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
3ff0726ebf build(deps): bump tensorflow (#315)
Bumps [tensorflow](https://github.com/tensorflow/tensorflow) from 2.5.2 to 2.5.3.
- [Release notes](https://github.com/tensorflow/tensorflow/releases)
- [Changelog](https://github.com/tensorflow/tensorflow/blob/master/RELEASE.md)
- [Commits](https://github.com/tensorflow/tensorflow/compare/v2.5.2...v2.5.3)

---
updated-dependencies:
- dependency-name: tensorflow
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>

Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2022-02-18 08:54:19 -06:00
Amy WuandGitHub c96c939dfe Fix RL samples (#270)
* Update step_by_step sample

* Update mlops sample and fix worker_pool_specs issue for thr trainer component

* Fix lint

* Fix lint

* Fix lint

* Fix import order

* Fix nbfmt

* Update component.yaml

* Format notebook

* Update component.yaml url

* Lint
2022-02-17 16:37:52 -08:00
719cf280c9 Updating markdown text to meet higher standard (#305)
* Updating markdown text to meet higher standard

* formatted nb

Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-02-17 17:45:15 -06:00
4ec6e2df04 [community-content] PyTorch on Google Cloud Vertex AI - Fixes based on feedback (#290)
* PyTorch on Vertex - Updated to match GCPC v0.2.2 API

* PyTorch on Vertex - Fixes based on review comments

* PyTorch on Vertex - fixes based on review

* PyTorch on Vertex - linter fixes

* PyTorch on Vertex - fixes based on feedback

* PyTorch on Vertex - fixes based on feedback

Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
2022-02-17 17:41:48 -06:00
031a9190c3 automl-tabular-classification.ipynb: Added missing code and removed unneeded text (#255)
* Added missing code and removed unneeded text

* Ran linter

* Ran linter

Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
2022-02-17 08:03:16 -08:00
5204dcf327 update training and batch predict notebook with importer (#303)
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
2022-02-16 18:16:56 -08:00
cad623ef84 update hp tuning sample to use importer (#302)
* update hp tuning sample to use importer

* Update get_started_with_hpt_pipeline_components.ipynb

* Update get_started_with_hpt_pipeline_components.ipynb

* Update get_started_with_hpt_pipeline_components.ipynb

Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
2022-02-16 18:16:23 -08:00
a59f58f8b6 Use importer for the bqml (#299)
* Use importer for the bqml

* Update get_started_with_bqml_pipeline_components.ipynb

Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
2022-02-16 18:00:44 -08:00
nayaknishantandGitHub d89c613f5d moving REGION up (#300)
* moving REGION up

* moving REGION up and csv file name
2022-02-16 17:11:46 -08:00
Andrew FerlitschandGitHub fa265ddb2f feat: fix for sklearn XAI (#301)
* feat: get started XAI

* feat: get started XAI

* feat: XAI with sklearn

* feat: XAI with sklearn

* feat: add covert component example

* feat: add covert component example

* feat: upgrade FS to SDK

* feat: upgrade FS to SDK

* feat: update to v1

* feat: update to v1

* fix: XAI for sklearn

* fix: XAI for sklearn
2022-02-16 17:00:16 -08:00
ec3dd04935 SDK Featurestore notebook (#284)
* SDK Featurestore notebook

* fixed issues, tried to make notebook more readable, style

* removed previous notebook

* moved sdk-feature-store to community (for now)

* made fixes

* moved BQ output table cells down

Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
Co-authored-by: Morgan Du <morgandu@google.com>
2022-02-16 16:04:43 -08:00
Andrew FerlitschandGitHub 0c83e81410 fix: 0.3.0 breaking change 2022-02-16 14:13:57 -08:00
Andrew FerlitschandGitHub 7808a843cc fix: breaking 0.3.0 change 2022-02-16 14:02:36 -08:00
Andrew FerlitschandGitHub 615d7706af fix: breaking change in 0.3.0 2022-02-16 13:53:35 -08:00
Ivan CheungandGitHub f05ca4d06a Added project to BQ client instantiation (#285) 2022-02-16 16:27:08 -05:00
Andrew FerlitschandGitHub cb4145e2b6 test: fix for testing (#298)
* test: fixes for testing

* test: fixes for testing
2022-02-16 11:23:45 -08:00
6917c9aa7b adding initial sample notebook for Tensorboard (#286)
Co-authored-by: Yicheng Fang <yichengfang@google.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
2022-02-16 09:45:01 -08:00
Andrew FerlitschandGitHub c1d2451656 feat: update to v1 (#294)
* feat: update to v1

* feat: update to v1

* feat: update to v1

* feat: update to v1
2022-02-16 09:35:41 -08:00
Ivan CheungandGitHub c41ec1fabd Cleaned up cleanup section (#291) 2022-02-16 10:28:42 -05:00
Ivan CheungandGitHub 273c91882e Delete setup_env.md 2022-02-15 20:45:06 -05:00
Ivan CheungandGitHub ad6b5e5830 Update test_notebook_vm.txt 2022-02-15 18:17:11 -05:00
Ivan CheungandGitHub d441eb9d7a Update test_notebook_vm.txt 2022-02-15 18:14:37 -05:00
Andrew FerlitschandGitHub 139d805c9f feat: update to v1 (#292)
* feat: get started XAI

* feat: get started XAI

* feat: XAI with sklearn

* feat: XAI with sklearn

* feat: add covert component example

* feat: add covert component example

* feat: upgrade FS to SDK

* feat: upgrade FS to SDK

* feat: update to v1

* feat: update to v1
2022-02-15 11:20:22 -08:00
Andrew FerlitschandGitHub 783347fc8e feat: update FS notebook to use SDK (#287)
* feat: get started XAI

* feat: get started XAI

* feat: XAI with sklearn

* feat: XAI with sklearn

* feat: add covert component example

* feat: add covert component example

* feat: upgrade FS to SDK

* feat: upgrade FS to SDK
2022-02-14 13:21:28 -08:00
6d722d081d Fix links (#274)
* Fixes and renames link to launch automl-text-classification.pynb in Vertex AI Workbench

* Fixes and renames link to launch sdk_automl_tabular_forecasting_batch.pynb in Vertex AI Workbench

* Fixes links for launching notebook in Vertex AI Workbench for Explainable AI samples

* Fixes link for launching notebook in Vertex AI Workbench for model monitoring sample

* Fixes links to launch pipelines notebook samples

* Fixed lint problem in automl-text-classification.ipynb

* Autofixed lint errors

Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-02-14 11:48:41 -05:00
Andrew FerlitschandGitHub 33a8c6ca0e feat: add create component example (#279)
* feat: get started XAI

* feat: get started XAI

* feat: XAI with sklearn

* feat: XAI with sklearn

* feat: add covert component example

* feat: add covert component example
2022-02-10 14:23:22 -08:00
Karl WeinmeisterandGitHub ae043400f5 Add survey to README.md (#272) 2022-02-10 14:52:33 -06:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>Karl Weinmeister
e41f96b31f build(deps): bump tensorflow (#275)
Bumps [tensorflow](https://github.com/tensorflow/tensorflow) from 2.5.2 to 2.5.3.
- [Release notes](https://github.com/tensorflow/tensorflow/releases)
- [Changelog](https://github.com/tensorflow/tensorflow/blob/master/RELEASE.md)
- [Commits](https://github.com/tensorflow/tensorflow/compare/v2.5.2...v2.5.3)

---
updated-dependencies:
- dependency-name: tensorflow
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>

Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-02-10 09:37:36 -06:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>Karl Weinmeister
7220f15158 build(deps): bump tensorflow (#276)
Bumps [tensorflow](https://github.com/tensorflow/tensorflow) from 2.5.2 to 2.5.3.
- [Release notes](https://github.com/tensorflow/tensorflow/releases)
- [Changelog](https://github.com/tensorflow/tensorflow/blob/master/RELEASE.md)
- [Commits](https://github.com/tensorflow/tensorflow/compare/v2.5.2...v2.5.3)

---
updated-dependencies:
- dependency-name: tensorflow
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>

Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-02-10 09:33:59 -06:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>Karl Weinmeister
b8d7eaa767 build(deps): bump tensorflow (#278)
Bumps [tensorflow](https://github.com/tensorflow/tensorflow) from 2.5.2 to 2.5.3.
- [Release notes](https://github.com/tensorflow/tensorflow/releases)
- [Changelog](https://github.com/tensorflow/tensorflow/blob/master/RELEASE.md)
- [Commits](https://github.com/tensorflow/tensorflow/compare/v2.5.2...v2.5.3)

---
updated-dependencies:
- dependency-name: tensorflow
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>

Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-02-10 09:31:20 -06:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
a612b463a8 build(deps): bump tensorflow (#277)
Bumps [tensorflow](https://github.com/tensorflow/tensorflow) from 2.5.2 to 2.5.3.
- [Release notes](https://github.com/tensorflow/tensorflow/releases)
- [Changelog](https://github.com/tensorflow/tensorflow/blob/master/RELEASE.md)
- [Commits](https://github.com/tensorflow/tensorflow/compare/v2.5.2...v2.5.3)

---
updated-dependencies:
- dependency-name: tensorflow
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>

Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2022-02-10 09:16:10 -06:00
Andrew FerlitschandGitHub 493e7e81b5 Update README.md 2022-02-09 10:19:58 -08:00
Andrew FerlitschandGitHub b6cf0dbefd feat: XAI with sklearn (#273)
* feat: get started XAI

* feat: get started XAI

* feat: XAI with sklearn

* feat: XAI with sklearn
2022-02-09 10:17:16 -08:00
Brian KangandGitHub 5d5c08f9e7 Pytorch lightning sdk example - Missed Notebook added (#269)
* Added example for PyTorch lightning distributed training of a ResNet model

* Revert "Added example for PyTorch lightning distributed training of a ResNet model"

This reverts commit bcbe832c51.

* Added example for PyTorch lightning distributed training of a ResNet model

* Revert "Added example for PyTorch lightning distributed training of a ResNet model"

This reverts commit 2a792cb7ac.

* Added example for PyTorch lightning distributed training of a ResNet model

* Added Notebook for PyTorch lightning distributed training of a ResNet model

* Added Notebook for PyTorch lightning distributed training of a ResNet model

* Notebook updates after review

* Notebook updates after review

* Added example for PyTorch lightning distributed training of a ResNet model

* Revert "Added example for PyTorch lightning distributed training of a ResNet model"

This reverts commit bcbe832c51.

* Added example for PyTorch lightning distributed training of a ResNet model

* Revert "Added example for PyTorch lightning distributed training of a ResNet model"

This reverts commit 2a792cb7ac.

* Added example for PyTorch lightning distributed training of a ResNet model

* Added Notebook for PyTorch lightning distributed training of a ResNet model

* Added Notebook for PyTorch lightning distributed training of a ResNet model

* Notebook updates after review

* Notebook updates after review

* Adjust Tensorboard to TensorBoard

* Adjust Tensorboard to TensorBoard

* Revert "Adjust Tensorboard to TensorBoard"

This reverts commit 9aac52e358b4ccc27a9a5a9e3bc5ef3305553462.

* Adjust Tensorboard to TensorBoard

* Adjust Tensorboard to TensorBoard

* Adjust Tensorboard to TensorBoard and and run lint
2022-02-08 16:36:30 -08:00
Andrew FerlitschandGitHub bfdfac0f81 feat: XAI notebook (#271)
* feat: get started XAI

* feat: get started XAI
2022-02-08 14:13:05 -08:00
Andrew FerlitschandGitHub 9c72b7e3e7 Update README.md 2022-02-08 12:57:22 -08:00
Andrew FerlitschandGitHub 46519e5c64 Update README.md 2022-02-08 12:26:57 -08:00
Brian KangandGitHub 0ff961e203 Pytorch lightning sdk example (#268)
* Added example for PyTorch lightning distributed training of a ResNet model

* Revert "Added example for PyTorch lightning distributed training of a ResNet model"

This reverts commit bcbe832c51.

* Added example for PyTorch lightning distributed training of a ResNet model

* Revert "Added example for PyTorch lightning distributed training of a ResNet model"

This reverts commit 2a792cb7ac.

* Added example for PyTorch lightning distributed training of a ResNet model
2022-02-07 17:36:40 -08:00
Andrew FerlitschandGitHub cbc17c6832 feat: use gcsfuse (#265)
* feat: friday update

* feat: friday update

* feat: hpt notebook

* feat: hpt notebook

* feat: sklearn

* feat: sklearn

* feat: xgb training

* feat: xgb training

* fix: predict on exported BQML

* fix: predict on exported BQML

* feat: add Pytorch notebook

* feat: add Pytorch notebook

* feat: add R notebook

* feat: add R notebook

* fix: spelling

* fix: spelling

* feat: add batch and TPU

* feat: add batch and TPU

* feat: use gcsfuse

* feat: use gcsfuse
2022-02-04 11:17:51 -08:00
Andrew FerlitschandGitHub 93e5b15cba Update README.md 2022-02-04 11:09:31 -08:00
Andrew FerlitschandGitHub 081e65d076 Update README.md 2022-02-04 10:09:45 -08:00
Ivan CheungandGitHub 4ab2cfb713 feat: Added test_notebook_vm.txt to only test fast-running notebooks. Also fix private_pool not being used for tests. (#248)
* feat: Added test_notebook_vm.txt

* Propagate private pool to child builds

* Fixed workerpool issue

* Removed private pool requirement

* Added default pool

* Fixed private pool region issue

* Added comment

* Made private_pool_id arg optionally present

* Fixed when private pool is optional

* Fix regional issue bug

* Fixed pandas requirement

* Removed flaky notebook
2022-02-03 15:48:40 -05:00
Andrew FerlitschandGitHub a73335c0af feat: add batch and TPU (#262)
* feat: friday update

* feat: friday update

* feat: hpt notebook

* feat: hpt notebook

* feat: sklearn

* feat: sklearn

* feat: xgb training

* feat: xgb training

* fix: predict on exported BQML

* fix: predict on exported BQML

* feat: add Pytorch notebook

* feat: add Pytorch notebook

* feat: add R notebook

* feat: add R notebook

* fix: spelling

* fix: spelling

* feat: add batch and TPU

* feat: add batch and TPU
2022-02-02 17:45:31 -08:00
0f3e257773 Pricingoptimization (#249)
* added notebook

* ran lint

* updated main

* ran lint

Co-authored-by: Ivan Cheung <ivans.mailbox@gmail.com>
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-02-02 13:25:41 -05:00
Andrew FerlitschandGitHub 57d734d84f fix: spelling (#260)
* feat: friday update

* feat: friday update

* feat: hpt notebook

* feat: hpt notebook

* feat: sklearn

* feat: sklearn

* feat: xgb training

* feat: xgb training

* fix: predict on exported BQML

* fix: predict on exported BQML

* feat: add Pytorch notebook

* feat: add Pytorch notebook

* feat: add R notebook

* feat: add R notebook

* fix: spelling

* fix: spelling
2022-02-01 17:07:10 -08:00
Andrew FerlitschandGitHub 07f2e9c999 Update README.md 2022-02-01 12:15:31 -08:00
Andrew FerlitschandGitHub 401883cae3 feat: add R notebook (#259)
* feat: friday update

* feat: friday update

* feat: hpt notebook

* feat: hpt notebook

* feat: sklearn

* feat: sklearn

* feat: xgb training

* feat: xgb training

* fix: predict on exported BQML

* fix: predict on exported BQML

* feat: add Pytorch notebook

* feat: add Pytorch notebook

* feat: add R notebook

* feat: add R notebook
2022-02-01 12:11:26 -08:00
7bc92e1e3c Update google_cloud_pipeline_components_automl_images.ipynb (#252)
Co-authored-by: Ivan Cheung <ivans.mailbox@gmail.com>
2022-02-01 13:33:22 -06:00
de5f8b0653 Removed unneeded lines in linter (#257)
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-02-01 12:19:19 -05:00
cfa73ed53b chore(deps): update dependency black to v22 (#254)
Co-authored-by: Ivan Cheung <ivans.mailbox@gmail.com>
2022-01-31 13:50:57 -06:00
Andrew FerlitschandGitHub 67370bb1c7 Update README.md 2022-01-31 10:39:28 -08:00
Andrew FerlitschandGitHub f60593255a feat: add Pytorch notebook (#256)
* feat: friday update

* feat: friday update

* feat: hpt notebook

* feat: hpt notebook

* feat: sklearn

* feat: sklearn

* feat: xgb training

* feat: xgb training

* fix: predict on exported BQML

* fix: predict on exported BQML

* feat: add Pytorch notebook

* feat: add Pytorch notebook
2022-01-31 10:38:28 -08:00
Andrew FerlitschandGitHub c6f9b97615 test: rm caching of components (#247)
* fix: testing

* fix: testing

* fix: not installing pandas
2022-01-31 13:33:18 -05:00
Andrew FerlitschandGitHub 6161a394c2 fix: predict on exported BQML model (#250)
* feat: friday update

* feat: friday update

* feat: hpt notebook

* feat: hpt notebook

* feat: sklearn

* feat: sklearn

* feat: xgb training

* feat: xgb training

* fix: predict on exported BQML

* fix: predict on exported BQML
2022-01-27 10:35:20 -08:00
ivanmkc b35cd42015 Added exception check for deletion method 2022-01-27 11:54:29 -05:00
Ivan CheungandGitHub 5e323993db [WIP] Relax requirements for DLVM's (#238)
* Update requirements.txt

* Relax all CI requirements
2022-01-26 16:47:09 -05:00
Andrew FerlitschandGitHub f1623e419e Update README.md 2022-01-26 12:10:16 -08:00
Andrew FerlitschandGitHub a3047fb1bb feat: xgboost notebook (#246)
* feat: friday update

* feat: friday update

* feat: hpt notebook

* feat: hpt notebook

* feat: sklearn

* feat: sklearn

* feat: xgb training

* feat: xgb training
2022-01-26 12:09:07 -08:00
Karl WeinmeisterandGitHub b4d02f486e Update CODEOWNERS with new community blog post (#244) 2022-01-26 09:16:48 -08:00
9e24893b9b feat: Add sentiment analysis notebook (#195)
* added notebook

* ran lint

* resolved comments

* ran lint

* resolved comments

* ran lint

Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: Ivan Cheung <ivans.mailbox@gmail.com>
2022-01-26 08:56:40 -06:00
47dec6ecef chore(deps): update dependency pyupgrade to v2.31.0 (#241)
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-01-26 08:53:32 -06:00
059fea672c chore(deps): update dependency nbqa to v1.2.3 (#240)
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-01-26 08:49:47 -06:00
389e804426 [community-content] PyTorch on Google Cloud Vertex AI - Blog related notebook and scripts (#236)
* PyTorch on Vertex - Updated to match GCPC v0.2.2 API

* PyTorch on Vertex - Fixes based on review comments

Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-01-26 08:46:09 -06:00
Andrew FerlitschandGitHub 087a638c18 Update README.md 2022-01-25 22:16:55 -08:00
Andrew FerlitschandGitHub 97757c74ca feat: sklearn training (#243)
* feat: friday update

* feat: friday update

* feat: hpt notebook

* feat: hpt notebook

* feat: sklearn

* feat: sklearn
2022-01-25 22:15:51 -08:00
Andrew FerlitschandGitHub 08f3ad583b feat: finish hpt components notebook (#239)
* feat: friday update

* feat: friday update

* feat: hpt notebook

* feat: hpt notebook
2022-01-25 10:02:28 -08:00
Karl WeinmeisterandGitHub 16f01d31d3 chore: Update notebook template license date to 2022 (#237)
* chore: Update notebook template license date to 2022

* chore: Add extra space to address lint error
2022-01-25 11:20:00 -06:00
e0f6c66351 chore(deps): update dependency pandas to v1.4.0 (#233)
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-01-24 13:58:50 -06:00
Ivan CheungandGitHub 3733b28772 Updated requirements (#235) 2022-01-24 14:54:31 -05:00
24f8912134 feat: Add inventory prediction notebook (#216)
* made changes

* made changes

* ran lint

Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: Ivan Cheung <ivans.mailbox@gmail.com>
2022-01-24 14:07:42 -05:00
WhiteSource RenovateandGitHub cdc8847f9b chore(deps): update dependency papermill to v2.3.4 (#234) 2022-01-24 10:46:31 -06:00
Andrew FerlitschandGitHub 25bd4b9eb5 Update README.md 2022-01-21 19:05:35 -08:00
Andrew FerlitschandGitHub d476191252 feat: add notebooks (#232)
* feat: friday update

* feat: friday update
2022-01-21 17:14:31 -08:00
nicainandGitHub bf35d6a07c Add metadata to hide cells (#226) 2022-01-21 14:57:24 -08:00
5378d38a05 chore(deps): update dependency ipython to v8.0.1 (#222)
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-01-21 11:15:14 -06:00
Ivan CheungandGitHub 85984f1173 Fixed broken Colab link 2022-01-20 17:52:56 -05:00
Ivan CheungandGitHub 92bb40349f Cleaned up cloud build by renaming files and moving methods around (#219)
* Cleaned up cloud build

* Fixed bug

* More renaming and comment cleanup

* Added helper file

* Fixed missing return
2022-01-20 15:55:38 -05:00
Rajesh ThallamandGitHub fcee9bf738 [community-content] PyTorch on Google Cloud Vertex AI - Blog related notebook and scripts (#227)
* PyTorch on Vertex - Adding pipelines notebook

* PyTorch on Vertex - Updating pipelines notebook

* PyTorch on Vertex - Updating pipelines notebook, clearing outputs

* PyTorch on Vertex - Updating pipelines notebook

* PyTorch on Vertex - reverting serving Dockerfile changes

* PyTorch on Vertex - linting fixes

* PyTorch on Vertex - updates to README

* PyTorch on Vertex - linter related fixes
2022-01-20 13:20:00 -06:00
nicainandGitHub cc2f3408ad Update Run After --> Run Selected Cell and All Below (#228) 2022-01-20 08:07:31 -06:00
nicainandGitHub 89fb218041 Alphafold on gcp tutorial (#221)
* Original alphafold Dockerfile and notebook as a starting point

* Migrate dependencies from notebook into Dockerfile

* adding build_docker.sh script for building docker image

* remove cell, and replace with accelerator configuration cell. Also remove dependency on google.colab

* dockerfile remove redundant

* Changing text in launch button

* add vertexai.png

* update notebook, including permalink to vertexai.png

* updating notebook markdown and correcting the vertexai.png image display

* add div brackets and fix broken launch link

* table instead of div

* width=40

* resizing vertexai image

* updated FAQ

* updated Licence

* add back in the output_file zip

* launch button at top of notebook

* default workdir aligned with JuptyerLab home directory

* intro paragraph

* update CPU instructions

* exchange notebook title and launch header

* Dockerfile license

* collapsing cells

* splitting out sequences into un-collapsed cell

* Update licence

* update download instructions text

* Remove "double-click" text

* hide cells

* Increasing indent to pass linting for alphafold_on_gcp (#224)

* increasing indent to pass linting

* more linting

* wild

* AMBER relaxation f-string

* hidden cells

* wild commit

* move links to main
2022-01-19 18:47:21 -08:00
Andrew FerlitschandGitHub 4d0c3781e5 Update README.md 2022-01-19 12:30:01 -08:00
Andrew FerlitschandGitHub e2e319ac7f Update README.md 2022-01-19 11:08:55 -08:00
Andrew FerlitschandGitHub 63508cad3f feat: add ml metadata notebook (#223)
* weekly updates

* weekly updates

* feat: add metadata notebook

* feat: add metadata notebook
2022-01-19 10:58:03 -08:00
0b6eb6eed1 Updates to automl tabular beans pipeline (#220)
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-01-18 15:34:57 -06:00
Polong LinandGitHub a41cfdeba5 Fix typo: Wikepedia --> Wikipedia (#217) 2022-01-18 08:28:04 -06:00
7a6763ca33 chore(deps): update dependency numpy to v1.22.1 (#214)
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-01-14 17:10:52 -06:00
Sara RobinsonandGitHub 2a6e19e4f9 Update MLMD notebook to use pipeline submit() method (#215) 2022-01-14 17:09:45 -06:00
9f9526b722 Delete Dockerfile (#198)
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-01-14 14:28:30 -05:00
Ivan CheungandGitHub 267684b7c3 Added private pool to cloud build config (#211) 2022-01-14 14:27:34 -05:00
Andrew FerlitschandGitHub b74dd3d576 Update README.md 2022-01-13 16:22:14 -08:00
Andrew FerlitschandGitHub 278ae48842 Update README.md 2022-01-13 16:21:28 -08:00
Andrew FerlitschandGitHub 189ca12627 Update README.md 2022-01-13 16:20:39 -08:00
Andrew FerlitschandGitHub 5108f57ba8 feat: weekly updates (#212)
* weekly updates

* weekly updates
2022-01-13 16:19:21 -08:00
6bace666e4 chore(deps): update dependency ipython to v8 (#209)
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-01-13 15:24:51 -06:00
4c35616d30 fix: Increase timeout on explainability notebooks to prevent build failure (#205)
* test: update timeout of testing

* test: increase timeout for testing

Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-01-13 15:13:20 -06:00
cadbc733b8 chore(deps): update dependency numpy to v1.22.0 (#207)
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-01-13 14:19:43 -06:00
ee6088957d Update Swivel sample and Matching Engine sample (#193)
* Update swivel and matchine engine samples

* Fix lint

* format notebooks

* fix lint

* fix lint

* upgrade google-python-api-client for testing pipeline

* upgrade google-api-core for testing pipeline

* install tensorflow after other required packages

* fix lint

* upgrade google-auth for testing

* install tensorflow in testing env

* upgrade pip with user flag

* remove kfp as a dependency

* separate matching engine notebook into another commit

* fix service account extraction

* service account is optional so comment it

* submit pipeline without specifying service account

* clarify how TensorBoard relates to Vertex ML metadata

* add reference to Vertex ML metadata

* fix typo

Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
2022-01-12 17:51:23 -08:00
58118c89f7 Chicago taxi fare prediction (#196)
* added notebook

* made changes

* ran linter

* small changes

* made changes

* ran linter

* ran linter

Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-01-11 16:51:19 -06:00
108112c9a0 Predictive maintaince usecase (#194)
* added

* ran linter

* comments resolved

* ran linter

Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-01-11 16:49:09 -06:00
Andrew FerlitschandGitHub 4f586719e7 Update README.md 2022-01-11 12:57:16 -08:00
31e4985480 chore(deps): update dependency nbconvert to v6.4.0 (#197)
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-01-11 13:32:39 -06:00
672b8bd262 chore(deps): update dependency numpy to v1.21.5 (#191)
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-01-11 13:28:22 -06:00
Andrew FerlitschandGitHub c64c185407 Update README.md 2022-01-10 15:52:05 -08:00
Andrew FerlitschandGitHub 19fd6aa4e3 Add files via upload 2022-01-10 15:48:47 -08:00
Andrew FerlitschandGitHub 55746d05f3 Update README.md 2022-01-10 14:08:44 -08:00
Andrew FerlitschandGitHub ad8d6382d4 Update README.md 2022-01-10 14:07:04 -08:00
Andrew FerlitschandGitHub 5114a6c09c Update README.md 2022-01-10 14:05:09 -08:00
Andrew FerlitschandGitHub f405c109e2 fix: git links (#206)
* https://github.com/GoogleCloudPlatform/vertex-ai-samples.git
'

* feat: updates

* feat: add BQML components

* fix: git links

* fix: git links
2022-01-10 13:08:04 -08:00
Andrew FerlitschandGitHub ff646fee07 Delete mlops_data_management.ipynb 2022-01-10 12:55:36 -08:00
Andrew FerlitschandGitHub e517a8998c Update README.md 2022-01-06 11:30:03 -08:00
Andrew FerlitschandGitHub 776a699a76 feat: BQML (#201)
* https://github.com/GoogleCloudPlatform/vertex-ai-samples.git
'

* feat: updates

* feat: add BQML components
2022-01-06 11:25:46 -08:00
a942a63959 chore(deps): update dependency ipython to v7.31.0 (#199)
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2022-01-05 19:22:15 -06:00
Andrew FerlitschandGitHub 4b275e3370 feat: updates (#200)
* https://github.com/GoogleCloudPlatform/vertex-ai-samples.git
'

* feat: updates
2022-01-05 09:59:12 -08:00
Karl WeinmeisterandGitHub 7bb6787800 docs: Update incorrect quote mark in CONTRIBUTING.md (#190) 2021-12-21 10:20:42 -08:00
Andrew FerlitschandGitHub d314dda749 fix: friday update (#192)
* friday updates

* friday updates
2021-12-20 10:37:30 -08:00
3e955d0b83 chore(deps): update dependency pandas to v1.3.5 (#189)
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2021-12-16 17:20:10 -06:00
WhiteSource RenovateandGitHub ae058696eb chore(deps): update dependency matplotlib to v3.5.1 (#188) 2021-12-16 17:14:30 -06:00
fa62eef2d1 chore(deps): pin dependencies (#8)
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2021-12-16 17:06:22 -06:00
Karl WeinmeisterandGitHub 16b9dcc09f build: Add tensorflow pip install to migration notebook (#185)
* Add tensorflow pip install to migration notebook

The build is failing due to tensorflow missing. Adding this dependency.

* build: Testing change to notebook test to resolve build error

* build: Reverted change to base branch
2021-12-16 14:30:41 -08:00
Ben LackeyandGitHub ed559bd9f4 Change URL to root repo (#183) 2021-12-14 08:42:18 -06:00
Andrew FerlitschandGitHub ddb825addf fix: MLops update (#182)
* fix: friday update

* fix: friday update

* fix: friday update

* fix: friday update

* fix: friday update
2021-12-13 10:42:29 -08:00
Andrew FerlitschandGitHub b78dfa38ef Update README.md 2021-12-10 12:35:54 -08:00
Andrew FerlitschandGitHub 101abf0566 fix: MLops 3v4 update (#181)
* fix: friday update

* fix: friday update

* fix: friday update
2021-12-10 10:43:14 -08:00
Karl WeinmeisterandGitHub d121d4a51d Update CODEOWNERS to reflect current ownership (#179) 2021-12-09 11:22:48 -08:00
Ben LackeyandGitHub 5d933b9b13 Add Neo4j notebook (#173)
* Add Neo4j notebook

* delete extra line

* Across this notebook, the dataframe assignment and display occurs both within and outside with statements. Consider following the pattern of the 2nd query, where it is outside.

* Typo: unlabled

* AutoML Tables is no longer a standalone product in Vertex AI, so I suggest the naming "Vertex AI for AutoML tabular data."
2021-12-09 13:03:31 -06:00
Karl WeinmeisterandGitHub bdd0b0453a Update linter requirements.txt (#176)
Updating linter dependencies based on conversation in #173
2021-12-08 14:39:12 -08:00
639ecb962e Change to pip3 and add to PATH (#169)
Hi.  I tried to run this all through cloud shell.  pip makes to Python 2.0 there and the command fails.  So, this PR has pip3.

Also, the cloud shell path didn't know about the directory where all this stuff installed, so I've added a command to add it to PATH.

Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2021-12-07 13:54:23 -06:00
Karl WeinmeisterandGitHub b3adc0bbf3 Update .github files for main branch conversion (#170) 2021-12-07 13:25:06 -06:00
36455b8125 Pipelines rui (#158)
* feat: customm train gcpc

* feat: customm train gcpc

* feat: customm train gcpc

* feat: updates for Karl review

* fix: build issues

* fix: build issues

* fix: lint issue

* fix: build issue

* fix: build issue

* fixL]: build issue

* added back google-python-api-client

* Updated google-api-core dependency

* Fixed quote issue

* Update location of google-api-core install

* Add force reinstall flag

* Separate TF and google-api-core installs

* add missing comma

* Added reinstall of google-auth

* Added force-reinstall for Tensorflow dependencies

* fix missing comma

* trying upgrade of google core packages

* Added diagnostic step and added httplib2

* Added more debugging statements

* Try changing user_flag when IS_TESTING set

* Update google-api-python-client dependency

* Add google-auth and force reinstall

* Update google-api-python-client

* Removed diagnostics and simplified test environment installs

* fix: TW review updates

* fix: TW review

* debug: notebook testing

* TW review update

* TW  review update

* test: force retest

* fix: lint

Co-authored-by: Karl Weinmeister <kweinmeister@google.com>
2021-12-03 11:13:12 -08:00
Andrew FerlitschandGitHub 1a748c7d1c Update mlops_experimentation.ipynb 2021-12-03 10:29:50 -08:00
Andrew FerlitschandGitHub 53734bf984 Ml opml ops 3v4 (#167)
* friday updates

* friday updates

* friday updates

* friday updates

* fix: custom train

* fix: custom train

* remove debug info

* fix: remove debug

* fix: remove debug

* fix: remove debug

* fix: remove debug

* fix: remove debug
2021-12-03 10:28:47 -08:00
Andrew FerlitschandGitHub 858ed794e6 feat: MLops 3v4 (#166)
* friday updates

* friday updates

* friday updates

* friday updates

* fix: custom train

* fix: custom train

* remove debug info

* fix: remove debug
2021-12-03 10:18:18 -08:00
Andrew FerlitschandGitHub e90e0fa7bd Update README.md 2021-11-30 18:42:23 -08:00
Andrew FerlitschandGitHub 81bbcce326 fix: custom train (#165)
* friday updates

* friday updates

* friday updates

* friday updates

* fix: custom train

* fix: custom train
2021-11-30 18:41:28 -08:00
Brian KangandGitHub d271648779 TPU pipeline to community folder (#164)
* TPU pipeline to community folder

* Changes per review by Andrew F

* Lint test

* Moved to community/pipelines

* Updated codeowners to reference pipelines folder

* Update to codeowners
2021-11-30 10:35:19 -05:00
Andrew FerlitschandGitHub f832c5d120 feat: stage 3 updates (#163)
* friday updates

* friday updates

* friday updates

* friday updates
2021-11-29 12:28:11 -08:00
Andrew FerlitschandGitHub 18f12353a7 feat: stage 3 friday updates (#160)
* feat: friday updates

* feat: friday updates

* feat: Friday updates

* feat: Friday updates
2021-11-19 15:59:27 -08:00
Karl WeinmeisterandGitHub a79a283a77 build: Add step to upgrade pip (#159)
* build: Add step to upgrade pip

* Changed location for running pip upgrade

* Removed whitespace
2021-11-19 15:02:48 -06:00
f67e0f0531 Add ML Metadata + Pipelines notebook (#59)
* feat: MLMD + Pipelines notebook

* Updates from notebook execution test

* Update with changes from linter

* Update MLMD notebook from feedback, upgrade to latest sdk versions

* Add metadata notebook to codeowners file

* Resolve merge conflicts with codeowners

* Update KFP and Vertex SDK versions

* Run the linter

Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
2021-11-18 12:23:42 -06:00
Andrew FerlitschandGitHub d6b6eb9483 Update README.md 2021-11-16 16:29:19 -08:00
096e3e1090 Moved forecasting to official folder (#84)
* Added forecasting notebook

* Fixed forecasting notebook

* Ran linter

* Small fix

* Fixed dataset variable name conflict

* Ran linter

* Fixed cleanup bug

* fix: remove online prediction reference

* fix: rephrase title to just AutoML tabular forecasting

* fix: update training time to one hour

Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
2021-11-16 16:24:06 -08:00
Andrew FerlitschandGitHub 400ff8c4a4 feat: get started BQML (#157)
* feat: stage3 updates

* feat: stage3 updates

* feat: stage3 updates

* feat: stage3 updates

* feat: add BML get started

* feat: add BML get started
2021-11-16 16:22:28 -08:00
Karl WeinmeisterandGitHub 1b0a3e8f90 Linting updates for PR #154 (#155)
* Removed extra scaling call in image notebooks

* formatting/linting updates
2021-11-16 15:05:26 -08:00
Andrew FerlitschandGitHub 91e17e3d80 Add files via upload 2021-11-16 10:02:37 -08:00
Andrew FerlitschandGitHub 40b80f357d Update README.md 2021-11-16 08:57:51 -08:00
Andrew FerlitschandGitHub 868423fd09 feat: updates for stage3 (#156)
* feat: stage3 updates

* feat: stage3 updates

* feat: stage3 updates

* feat: stage3 updates
2021-11-16 08:56:00 -08:00
Karl WeinmeisterandGitHub 644612fde9 Removed extra scaling call in image notebooks (#154) 2021-11-15 17:45:02 -08:00
Ivan CheungandGitHub a2ae130378 Fixed edge case where no notebooks are detected (#152) 2021-11-15 13:11:17 -06:00
Karl WeinmeisterandGitHub dddd43c191 Revert "Temporary CI change to exempt failed notebooks (#148)" (#150)
This reverts commit 3632c5b849.
2021-11-15 13:03:46 -06:00
Andrew FerlitschandGitHub 6589d56e4c fix: install TF 2021-11-15 10:11:33 -08:00
Andrew FerlitschandGitHub 5d5c4b8e9d fix: add install for TF 2021-11-15 10:10:15 -08:00
Andrew FerlitschandGitHub d380f45485 fix: misspelling 2021-11-15 09:46:13 -08:00
Andrew FerlitschandGitHub 33d8fed6d7 fix: IS_TESTING 2021-11-15 09:32:24 -08:00
Andrew FerlitschandGitHub 6e59dec4f0 fix: IS_TESTING 2021-11-15 09:31:00 -08:00
Andrew FerlitschandGitHub fe802e118c fix: IS_TESTING 2021-11-15 09:30:06 -08:00
Andrew FerlitschandGitHub d018865a7c fix: IS_TESTING 2021-11-15 09:29:23 -08:00
Andrew FerlitschandGitHub 9c0f8fc5bf fix: IS_TESTING 2021-11-15 09:27:19 -08:00
Andrew FerlitschandGitHub bc58e05d2a fix: IS_TESTING 2021-11-15 09:26:26 -08:00
Andrew FerlitschandGitHub 977a4657fa fix: IS_TESTING 2021-11-15 09:10:46 -08:00
Andrew FerlitschandGitHub e8cbe3c5c5 fix: IS_TESTING 2021-11-15 09:09:40 -08:00
Andrew FerlitschandGitHub 8b114d3eec fix: IS_TESTING 2021-11-15 09:07:18 -08:00
Ivan CheungandGitHub 740958005a Parallel custom job notebook execution (#120)
* Initial commit

* Added parallel execution
2021-11-12 18:31:36 -05:00
Karl WeinmeisterandGitHub d65e074163 Updated doc locations for PolicyBot compliance (#149)
* Updated doc locations for policybot compliance

* Removed docs folder with duplicates
2021-11-12 16:28:02 -06:00
google-cloud-policy-bot[bot]GitHubgoogle-cloud-policy-bot[bot] <80869356+google-cloud-policy-bot[bot]@users.noreply.github.com>
816c8f2309 chore: add SECURITY.md (#145)
Co-authored-by: google-cloud-policy-bot[bot] <80869356+google-cloud-policy-bot[bot]@users.noreply.github.com>
2021-11-12 11:32:03 -06:00
google-cloud-policy-bot[bot]GitHubgoogle-cloud-policy-bot[bot] <80869356+google-cloud-policy-bot[bot]@users.noreply.github.com>
a9e0b99992 chore: add CONTRIBUTING.md (#146)
Co-authored-by: google-cloud-policy-bot[bot] <80869356+google-cloud-policy-bot[bot]@users.noreply.github.com>
2021-11-12 11:31:52 -06:00
google-cloud-policy-bot[bot]GitHubgoogle-cloud-policy-bot[bot] <80869356+google-cloud-policy-bot[bot]@users.noreply.github.com>
a61802821d chore: add a Code of Conduct (#144)
Co-authored-by: google-cloud-policy-bot[bot] <80869356+google-cloud-policy-bot[bot]@users.noreply.github.com>
2021-11-12 11:31:39 -06:00
Alec GlassfordandGitHub b078e50cc0 Relax language re: multi-region bucket for training (#142)
I believe you *can* use a multi-region; it just might not work as well (e.g. for latency, cost) as a regional bucket where you are doing other operations. This removes in inaccurate statement.
2021-11-12 11:22:30 -06:00
Karl WeinmeisterandGitHub 3632c5b849 Temporary CI change to exempt failed notebooks (#148)
#120 is failing due to existing notebooks failing. This change will temporarily reduce the number of folders checked, to enable this PR to be merged, which enables weekly regression tests. Then, we will address the issues with the notebooks and remove these exemptions.
2021-11-12 09:11:09 -08:00
Andrew FerlitschandGitHub 5201dc9828 Update README.md 2021-11-11 16:29:56 -08:00
Andrew FerlitschandGitHub e0861b00e4 Create README.md 2021-11-11 16:25:27 -08:00
Andrew FerlitschandGitHub 2d4cff3c2e Update README.md 2021-11-11 16:21:44 -08:00
Andrew FerlitschandGitHub 428f861b9b feat: stage 3 notebooks (#147)
* feat: stage 3 notebooks

* feat: stage 3 notebooks
2021-11-11 16:21:03 -08:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
31487745d6 build(deps): bump tensorflow (#138)
Bumps [tensorflow](https://github.com/tensorflow/tensorflow) from 2.4.1 to 2.5.2.
- [Release notes](https://github.com/tensorflow/tensorflow/releases)
- [Changelog](https://github.com/tensorflow/tensorflow/blob/master/RELEASE.md)
- [Commits](https://github.com/tensorflow/tensorflow/compare/v2.4.1...v2.5.2)

---
updated-dependencies:
- dependency-name: tensorflow
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>

Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2021-11-10 18:02:45 -06:00
Karl WeinmeisterandGitHub d1ccf37e68 Notebook fixes to work in Workbench (#141) 2021-11-10 18:01:12 -06:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
91440922b5 build(deps): bump tensorflow (#136)
Bumps [tensorflow](https://github.com/tensorflow/tensorflow) from 2.5.0 to 2.5.2.
- [Release notes](https://github.com/tensorflow/tensorflow/releases)
- [Changelog](https://github.com/tensorflow/tensorflow/blob/master/RELEASE.md)
- [Commits](https://github.com/tensorflow/tensorflow/compare/v2.5.0...v2.5.2)

---
updated-dependencies:
- dependency-name: tensorflow
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>

Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2021-11-10 15:23:36 -06:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
e15cc7e3cd build(deps): bump tensorflow (#137)
Bumps [tensorflow](https://github.com/tensorflow/tensorflow) from 2.5.1 to 2.5.2.
- [Release notes](https://github.com/tensorflow/tensorflow/releases)
- [Changelog](https://github.com/tensorflow/tensorflow/blob/master/RELEASE.md)
- [Commits](https://github.com/tensorflow/tensorflow/compare/v2.5.1...v2.5.2)

---
updated-dependencies:
- dependency-name: tensorflow
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>

Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2021-11-10 15:22:56 -06:00
c6e767edaf Train and deploy end-to-end sklearn text classifier (#122)
* added sklearn-example

* nb formatting

* added readme

* added endpoint to notebook

* nb formatting

* added endpoint prediction example

* nb formatting

* fixed request

* minor woring fixes in docstrings

* added codeowner and included PR feedback

Co-authored-by: Maximilian Engelhardt <maximilian.engelhardt@ing.com>
2021-11-10 15:20:57 -06:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
176c79033d build(deps): bump tensorflow (#135)
Bumps [tensorflow](https://github.com/tensorflow/tensorflow) from 2.5.0 to 2.5.2.
- [Release notes](https://github.com/tensorflow/tensorflow/releases)
- [Changelog](https://github.com/tensorflow/tensorflow/blob/master/RELEASE.md)
- [Commits](https://github.com/tensorflow/tensorflow/compare/v2.5.0...v2.5.2)

---
updated-dependencies:
- dependency-name: tensorflow
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>

Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2021-11-10 15:02:46 -06:00
199 changed files with 70603 additions and 4121 deletions
-194
View File
@@ -1,194 +0,0 @@
# Copyright 2020 Google LLC
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
# We want to use LTS ubuntu from our mirror because dockerhub has a
# rate limit.
# FROM mirror.gcr.io/library/ubuntu:18.04
# However, now the above image is not working, we're using our own cache
FROM gcr.io/cloud-devrel-kokoro-resources/ubuntu:20.04
ENV DEBIAN_FRONTEND noninteractive
# Ensure local Python is preferred over distribution Python.
ENV PATH /usr/local/bin:$PATH
# http://bugs.python.org/issue19846
# At the moment, setting "LANG=C" on a Linux system fundamentally breaks
# Python 3.
ENV LANG C.UTF-8
# Install dependencies.
RUN apt-get update \
&& apt-get install -y --no-install-recommends \
apt-transport-https \
build-essential \
ca-certificates \
curl \
dirmngr \
git \
gcc \
gpg-agent \
graphviz \
libbz2-dev \
libdb5.3-dev \
libexpat1-dev \
libffi-dev \
liblzma-dev \
libmagickwand-dev \
libmemcached-dev \
libpython3-dev \
libreadline-dev \
libsnappy-dev \
libssl-dev \
libsqlite3-dev \
portaudio19-dev \
pkg-config \
redis-server \
software-properties-common \
ssh \
sudo \
systemd \
tcl \
tcl-dev \
tk \
tk-dev \
uuid-dev \
wget \
zlib1g-dev \
&& apt-get clean autoclean \
&& apt-get autoremove -y \
&& rm -rf /var/lib/apt/lists/* \
&& rm -f /var/cache/apt/archives/*.deb
# Install docker
RUN curl -fsSL https://download.docker.com/linux/ubuntu/gpg | sudo apt-key add -
RUN add-apt-repository \
"deb [arch=amd64] https://download.docker.com/linux/ubuntu \
$(lsb_release -cs) \
stable"
RUN apt-get update \
&& apt-get install -y --no-install-recommends \
docker-ce \
&& apt-get clean autoclean \
&& apt-get autoremove -y \
&& rm -rf /var/lib/apt/lists/* \
&& rm -f /var/cache/apt/archives/*.deb
# Install Bazel for compiling Tink in Cloud SQL Client Side Encryption Samples
# TODO: Delete this section once google/tink#483 is resolved
RUN apt install -y curl gpgconf gpg \
&& curl -fsSL https://bazel.build/bazel-release.pub.gpg | gpg --dearmor > bazel.gpg \
&& mv bazel.gpg /etc/apt/trusted.gpg.d/ \
&& echo "deb [arch=amd64] https://storage.googleapis.com/bazel-apt stable jdk1.8" | sudo tee /etc/apt/sources.list.d/bazel.list \
&& apt update && apt install -y bazel \
&& apt-get clean autoclean \
&& apt-get autoremove -y \
&& rm -rf /var/lib/apt/lists/* \
&& rm -f /var/cache/apt/archives/*.deb
# Install Microsoft ODBC 17 Driver and unixodbc for testing SQL Server samples
RUN curl https://packages.microsoft.com/keys/microsoft.asc | apt-key add - \
&& curl https://packages.microsoft.com/config/ubuntu/20.04/prod.list > /etc/apt/sources.list.d/mssql-release.list \
&& apt-get update \
&& ACCEPT_EULA=Y apt-get install -y --no-install-recommends \
msodbcsql17 \
unixodbc-dev \
&& apt-get clean autoclean \
&& apt-get autoremove -y \
&& rm -rf /var/lib/apt/lists/* \
&& rm -f /var/cache/apt/archives/*.deb
COPY fetch_gpg_keys.sh /tmp
# Install the desired versions of Python.
RUN set -ex \
&& export GNUPGHOME="$(mktemp -d)" \
&& echo "disable-ipv6" >> "${GNUPGHOME}/dirmngr.conf" \
&& /tmp/fetch_gpg_keys.sh \
&& for PYTHON_VERSION in 2.7.18 3.6.13 3.7.10 3.8.8 3.9.2; do \
wget --no-check-certificate -O python-${PYTHON_VERSION}.tar.xz "https://www.python.org/ftp/python/${PYTHON_VERSION%%[a-z]*}/Python-$PYTHON_VERSION.tar.xz" \
&& wget --no-check-certificate -O python-${PYTHON_VERSION}.tar.xz.asc "https://www.python.org/ftp/python/${PYTHON_VERSION%%[a-z]*}/Python-$PYTHON_VERSION.tar.xz.asc" \
&& gpg --batch --verify python-${PYTHON_VERSION}.tar.xz.asc python-${PYTHON_VERSION}.tar.xz \
&& rm -r python-${PYTHON_VERSION}.tar.xz.asc \
&& mkdir -p /usr/src/python-${PYTHON_VERSION} \
&& tar -xJC /usr/src/python-${PYTHON_VERSION} --strip-components=1 -f python-${PYTHON_VERSION}.tar.xz \
&& rm python-${PYTHON_VERSION}.tar.xz \
&& cd /usr/src/python-${PYTHON_VERSION} \
&& ./configure \
--enable-shared \
# This works only on Python 2.7 and throws a warning on every other
# version, but seems otherwise harmless.
--enable-unicode=ucs4 \
--with-system-ffi \
--without-ensurepip \
&& make -j$(nproc) \
&& make install \
&& ldconfig \
; done \
&& rm -rf "${GNUPGHOME}" \
&& rm -rf /usr/src/python* \
&& rm -rf ~/.cache/
# Install pip on Python 3.6 only.
# If the environment variable is called "PIP_VERSION", pip explodes with
# "ValueError: invalid truth value '<VERSION>'"
ENV PYTHON_PIP_VERSION 20.2.4
RUN wget --no-check-certificate -O /tmp/get-pip.py 'https://bootstrap.pypa.io/get-pip.py' \
&& python3.6 /tmp/get-pip.py "pip==$PYTHON_PIP_VERSION" \
# we use "--force-reinstall" for the case where the version of pip we're trying to install is the same as the version bundled with Python
# ("Requirement already up-to-date: pip==8.1.2 in /usr/local/lib/python3.6/site-packages")
# https://github.com/docker-library/python/pull/143#issuecomment-241032683
&& pip3 install --no-cache-dir --upgrade --force-reinstall "pip==$PYTHON_PIP_VERSION" \
# then we use "pip list" to ensure we don't have more than one pip version installed
# https://github.com/docker-library/python/pull/100
&& [ "$(pip list |tac|tac| awk -F '[ ()]+' '$1 == "pip" { print $2; exit }')" = "$PYTHON_PIP_VERSION" ]
# Ensure Pip for python3
RUN python3 /tmp/get-pip.py
RUN rm /tmp/get-pip.py
# Install "virtualenv", since the vast majority of users of this image
# will want it.
RUN pip install --no-cache-dir virtualenv
# Setup Cloud SDK
ENV CLOUD_SDK_VERSION 339.0.0
# Use system python for cloud sdk.
ENV CLOUDSDK_PYTHON python3.6
RUN wget https://dl.google.com/dl/cloudsdk/channels/rapid/downloads/google-cloud-sdk-$CLOUD_SDK_VERSION-linux-x86_64.tar.gz
RUN tar xzf google-cloud-sdk-$CLOUD_SDK_VERSION-linux-x86_64.tar.gz
RUN /google-cloud-sdk/install.sh
ENV PATH /google-cloud-sdk/bin:$PATH
# Enable redis-server on boot.
RUN sudo systemctl enable redis-server.service
# Create a user and allow sudo
# kbuilder uid on the default Kokoro image
ARG UID=1000
ARG USERNAME=kbuilder
# Add a new user to the container image.
# This is needed for ssh and sudo access.
# Add a new user with the caller's uid and the username.
RUN useradd -d /h -u ${UID} ${USERNAME}
# Allow nopasswd sudo
RUN echo "${USERNAME} ALL=(ALL) NOPASSWD:ALL" >> /etc/sudoers
CMD ["python3.6"]
-175
View File
@@ -1,175 +0,0 @@
#!/usr/bin/env python
# Copyright 2021 Google LLC
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
import json
import sys
import nbformat
import os
import errno
from NotebookProcessors import RemoveNoExecuteCells, UpdateVariablesPreprocessor
from typing import Dict, Tuple
import papermill as pm
import shutil
import virtualenv
import uuid
from jupyter_client.kernelspecapp import KernelSpecManager
# This script is used to execute a notebook and write out the output notebook.
# The replaces calling the nbconvert via command-line, which doesn't write the output notebook correctly when there are errors during execution.
STAGING_FOLDER = "staging"
ENVIRONMENTS_PATH = "environments"
KERNELS_SPECS_PATH = "kernel_specs"
def create_and_install_kernel() -> Tuple[str, str]:
# Create environment
kernel_name = str(uuid.uuid4())
env_name = f"{ENVIRONMENTS_PATH}/{kernel_name}"
# venv.create(env_name, system_site_packages=True, with_pip=True)
virtualenv.cli_run([env_name, "--system-site-packages"])
# Create kernel spec
kernel_spec = {
"argv": [
f"{env_name}/bin/python",
"-m",
"ipykernel_launcher",
"-f",
"{connection_file}",
],
"display_name": "Python 3",
"language": "python",
}
kernel_spec_folder = os.path.join(KERNELS_SPECS_PATH, kernel_name)
kernel_spec_file = os.path.join(kernel_spec_folder, "kernel.json")
# Create kernel spec folder
if not os.path.exists(os.path.dirname(kernel_spec_file)):
try:
os.makedirs(os.path.dirname(kernel_spec_file))
except OSError as exc: # Guard against race condition
if exc.errno != errno.EEXIST:
raise
with open(kernel_spec_file, mode="w", encoding="utf-8") as f:
json.dump(kernel_spec, f)
# Install kernel
kernel_spec_manager = KernelSpecManager()
kernel_spec_manager.install_kernel_spec(
source_dir=kernel_spec_folder, kernel_name=kernel_name
)
return kernel_name, env_name
def execute_notebook(
notebook_file_path: str,
output_file_folder: str,
replacement_map: Dict[str, str],
should_log_output: bool,
should_use_new_kernel: bool,
):
# Create staging directory if it doesn't exist
staging_file_path = f"{STAGING_FOLDER}/{notebook_file_path}"
if not os.path.exists(os.path.dirname(staging_file_path)):
try:
os.makedirs(os.path.dirname(staging_file_path))
except OSError as exc: # Guard against race condition
if exc.errno != errno.EEXIST:
raise
file_name = os.path.basename(os.path.normpath(notebook_file_path))
# Create environments folder
if not os.path.exists(ENVIRONMENTS_PATH):
try:
os.makedirs(ENVIRONMENTS_PATH)
except OSError as exc: # Guard against race condition
if exc.errno != errno.EEXIST:
raise
# Create and install kernel
kernel_name = next(
iter(KernelSpecManager().find_kernel_specs().keys()), None
) # Find first existing kernel and use as default
env_name = None
if should_use_new_kernel:
kernel_name, env_name = create_and_install_kernel()
# Read notebook
with open(notebook_file_path) as f:
nb = nbformat.read(f, as_version=4)
has_error = False
# Execute notebook
try:
# Create preprocessors
remove_no_execute_cells_preprocessor = RemoveNoExecuteCells()
update_variables_preprocessor = UpdateVariablesPreprocessor(
replacement_map=replacement_map
)
# Use no-execute preprocessor
(
nb,
resources,
) = remove_no_execute_cells_preprocessor.preprocess(nb)
(nb, resources) = update_variables_preprocessor.preprocess(nb, resources)
# print(f"Staging modified notebook to: {staging_file_path}")
with open(staging_file_path, mode="w", encoding="utf-8") as f:
nbformat.write(nb, f)
# Execute notebook
pm.execute_notebook(
input_path=staging_file_path,
output_path=staging_file_path,
kernel_name=kernel_name,
progress_bar=should_log_output,
request_save_on_cell_execute=should_log_output,
log_output=should_log_output,
stdout_file=sys.stdout if should_log_output else None,
stderr_file=sys.stderr if should_log_output else None,
)
except Exception:
# print(f"Error executing the notebook: {notebook_file_path}.\n\n")
has_error = True
raise
finally:
# Clear env
if env_name is not None:
shutil.rmtree(path=env_name)
# Copy execute notebook
output_file_path = os.path.join(
output_file_folder, "failure" if has_error else "success", file_name
)
# Create directories if they don't exist
if not os.path.exists(os.path.dirname(output_file_path)):
try:
os.makedirs(os.path.dirname(output_file_path))
except OSError as exc: # Guard against race condition
if exc.errno != errno.EEXIST:
raise
# print(f"Writing output to: {output_file_path}")
shutil.move(staging_file_path, output_file_path)
+25 -22
View File
@@ -1,42 +1,45 @@
from typing import List
from resource_cleanup_manager import (
ResourceCleanupManager,
DatasetResourceCleanupManager,
EndpointResourceCleanupManager,
ModelResourceCleanupManager,
ResourceCleanupManager,
DatasetResourceCleanupManager,
EndpointResourceCleanupManager,
ModelResourceCleanupManager,
)
def run_cleanup_managers(managers: List[ResourceCleanupManager], is_dry_run: bool):
for manager in managers:
type_name = manager.type_name
for manager in managers:
type_name = manager.type_name
print(f"Fetching {type_name}'s...")
resources = manager.list()
print(f"Found {len(resources)} {type_name}'s")
for resource in resources:
if not manager.is_deletable(resource):
continue
print(f"Fetching {type_name}'s...")
resources = manager.list()
print(f"Found {len(resources)} {type_name}'s")
for resource in resources:
if not manager.is_deletable(resource):
continue
if is_dry_run:
resource_name = manager.resource_name(resource)
print(f"Will delete '{type_name}': {resource_name}")
else:
manager.delete(resource)
if is_dry_run:
resource_name = manager.resource_name(resource)
print(f"Will delete '{type_name}': {resource_name}")
else:
try:
manager.delete(resource)
except Exception as exception:
print(exception)
print("")
print("")
is_dry_run = False
if is_dry_run:
print("Starting cleanup in dry run mode...")
print("Starting cleanup in dry run mode...")
# List of all cleanup managers
managers = [
DatasetResourceCleanupManager(),
EndpointResourceCleanupManager(),
ModelResourceCleanupManager(),
DatasetResourceCleanupManager(),
EndpointResourceCleanupManager(),
ModelResourceCleanupManager(),
]
run_cleanup_managers(managers=managers, is_dry_run=is_dry_run)
+107
View File
@@ -0,0 +1,107 @@
#!/usr/bin/env python
# Copyright 2021 Google LLC
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
"""A CLI to process changed notebooks and execute them on Google Cloud Build"""
import argparse
import pathlib
import execute_changed_notebooks_helper
def str2bool(v):
if isinstance(v, bool):
return v
if v.lower() in ("yes", "true", "t", "y", "1"):
return True
elif v.lower() in ("no", "false", "f", "n", "0"):
return False
else:
raise argparse.ArgumentTypeError("Boolean value expected.")
parser = argparse.ArgumentParser(description="Run changed notebooks.")
parser.add_argument(
"--test_paths_file",
type=pathlib.Path,
help="The path to the file that has newline-limited folders of notebooks that should be tested.",
required=True,
)
parser.add_argument(
"--base_branch",
help="The base git branch to diff against to find changed files.",
required=False,
)
parser.add_argument(
"--container_uri",
type=str,
help="The container uri to run each notebook in.",
required=True,
)
parser.add_argument(
"--variable_project_id",
type=str,
help="The GCP project id. This is used to inject a variable value into the notebook before running.",
required=True,
)
parser.add_argument(
"--variable_region",
type=str,
help="The GCP region. This is used to inject a variable value into the notebook before running.",
required=True,
)
parser.add_argument(
"--staging_bucket",
type=str,
help="The GCP directory for staging temporary files.",
required=True,
)
parser.add_argument(
"--artifacts_bucket",
type=str,
help="The GCP directory for storing executed notebooks.",
required=True,
)
parser.add_argument(
"--private_pool_id",
type=str,
help="The private pool id.",
required=False,
)
parser.add_argument(
"--should_parallelize",
type=str2bool,
nargs="?",
const=True,
default=True,
help="Should run notebooks in parallel.",
)
args = parser.parse_args()
notebooks = execute_changed_notebooks_helper.get_changed_notebooks(
test_paths_file=args.test_paths_file,
base_branch=args.base_branch,
)
execute_changed_notebooks_helper.process_and_execute_notebooks(
notebooks=notebooks,
container_uri=args.container_uri,
staging_bucket=args.staging_bucket,
artifacts_bucket=args.artifacts_bucket,
variable_project_id=args.variable_project_id,
variable_region=args.variable_region,
private_pool_id=args.private_pool_id if not "default" else None,
should_parallelize=args.should_parallelize,
)
@@ -13,30 +13,22 @@
# See the License for the specific language governing permissions and
# limitations under the License.
import argparse
import concurrent
import dataclasses
import datetime
import functools
import pathlib
import os
import pathlib
import nbformat
import re
import subprocess
from pathlib import Path
from typing import List, Optional
import concurrent
from tabulate import tabulate
import operator
import ExecuteNotebook
def str2bool(v):
if isinstance(v, bool):
return v
if v.lower() in ("yes", "true", "t", "y", "1"):
return True
elif v.lower() in ("no", "false", "f", "n", "0"):
return False
else:
raise argparse.ArgumentTypeError("Boolean value expected.")
import execute_notebook_remote
from utils import util, NotebookProcessors
from google.cloud.devtools.cloudbuild_v1.types import BuildOperationMetadata
def format_timedelta(delta: datetime.timedelta) -> str:
@@ -62,49 +54,138 @@ def format_timedelta(delta: datetime.timedelta) -> str:
@dataclasses.dataclass
class NotebookExecutionResult:
notebook: str
name: str
duration: datetime.timedelta
is_pass: bool
log_url: str
output_uri: str
build_id: str
error_message: Optional[str]
def execute_notebook(
artifacts_path: str,
def _process_notebook(
notebook_path: str,
variable_project_id: str,
variable_region: str,
should_log_output: bool,
should_use_new_kernel: bool,
):
# Read notebook
with open(notebook_path) as f:
nb = nbformat.read(f, as_version=4)
# Create preprocessors
remove_no_execute_cells_preprocessor = NotebookProcessors.RemoveNoExecuteCells()
update_variables_preprocessor = NotebookProcessors.UpdateVariablesPreprocessor(
replacement_map={
"PROJECT_ID": variable_project_id,
"REGION": variable_region,
},
)
# Use no-execute preprocessor
(
nb,
resources,
) = remove_no_execute_cells_preprocessor.preprocess(nb)
(nb, resources) = update_variables_preprocessor.preprocess(nb, resources)
with open(notebook_path, mode="w", encoding="utf-8") as new_file:
nbformat.write(nb, new_file)
def _create_tag(filepath: str) -> str:
tag = os.path.basename(os.path.normpath(filepath))
tag = re.sub("[^0-9a-zA-Z_.-]+", "-", tag)
if tag.startswith(".") or tag.startswith("-"):
tag = tag[1:]
return tag
def process_and_execute_notebook(
container_uri: str,
staging_bucket: str,
artifacts_bucket: str,
variable_project_id: str,
variable_region: str,
private_pool_id: Optional[str],
notebook: str,
should_get_tail_logs: bool = False,
) -> NotebookExecutionResult:
print(f"Running notebook: {notebook}")
# Create paths
notebook_output_uri = "/".join([artifacts_bucket, pathlib.Path(notebook).name])
# Create tag from notebook
tag = _create_tag(filepath=notebook)
result = NotebookExecutionResult(
notebook=notebook,
name=tag,
duration=datetime.timedelta(seconds=0),
is_pass=False,
output_uri=notebook_output_uri,
log_url="",
build_id="",
error_message=None,
)
# TODO: Handle cases where multiple notebooks have the same name
time_start = datetime.datetime.now()
operation = None
try:
ExecuteNotebook.execute_notebook(
notebook_file_path=notebook,
output_file_folder=artifacts_path,
replacement_map={
"PROJECT_ID": variable_project_id,
"REGION": variable_region,
},
should_log_output=should_log_output,
should_use_new_kernel=should_use_new_kernel,
# Pre-process notebook by substituting variable names
_process_notebook(
notebook_path=notebook,
variable_project_id=variable_project_id,
variable_region=variable_region,
)
# Upload the pre-processed code to a GCS bucket
code_archive_uri = util.archive_code_and_upload(staging_bucket=staging_bucket)
operation = execute_notebook_remote.execute_notebook_remote(
code_archive_uri=code_archive_uri,
notebook_uri=notebook,
notebook_output_uri=notebook_output_uri,
container_uri=container_uri,
tag=tag,
region=variable_region,
private_pool_id=private_pool_id,
)
operation_metadata = BuildOperationMetadata(mapping=operation.metadata)
result.build_id = operation_metadata.build.id
result.log_url = operation_metadata.build.log_url
# Block and wait for the result
operation_result = operation.result()
result.duration = datetime.datetime.now() - time_start
result.is_pass = True
print(f"{notebook} PASSED in {format_timedelta(result.duration)}.")
except Exception as error:
result.error_message = str(error)
if operation and should_get_tail_logs:
# Extract the logs
logs_bucket = operation_metadata.build.logs_bucket
# Download tail end of logs file
log_file_uri = f"{logs_bucket}/log-{result.build_id}.txt"
# Use gcloud to get tail
try:
result.error_message = subprocess.check_output(
["gsutil", "cat", "-r", "-1000", log_file_uri], encoding="UTF-8"
)
except Exception as error:
result.error_message = str(error)
result.duration = datetime.datetime.now() - time_start
result.is_pass = False
result.error_message = str(error)
print(
f"{notebook} FAILED in {format_timedelta(result.duration)}: {result.error_message}"
)
@@ -112,43 +193,13 @@ def execute_notebook(
return result
def run_changed_notebooks(
def get_changed_notebooks(
test_paths_file: str,
base_branch: Optional[str],
output_folder: str,
variable_project_id: str,
variable_region: str,
should_parallelize: bool,
should_use_separate_kernels: bool,
):
base_branch: Optional[str] = None,
) -> List[str]:
"""
Run the notebooks that exist under the folders defined in the test_paths_file.
It only runs notebooks that have differences from the Git base_branch.
The executed notebooks are saved in the output_folder.
Variables are also injected into the notebooks such as the variable_project_id and variable_region.
Args:
test_paths_file (str):
Required. The new-line delimited file to folders and files that need checking.
Folders are checked recursively.
base_branch (str):
Optional. If provided, only the files that have changed from the base_branch will be checked.
If not provided, all files will be checked.
output_folder (str):
Required. The folder to write executed notebooks to.
variable_project_id (str):
Required. The value for PROJECT_ID to inject into notebooks.
variable_region (str):
Required. The value for REGION to inject into notebooks.
should_parallelize (bool):
Required. Should run notebooks in parallel using a thread pool as opposed to in sequence.
should_use_separate_kernels (bool):
Note: Dependencies don't install correctly when this is set to True
See https://github.com/nteract/papermill/issues/625
Required. Should run each notebook in a separate and independent virtual environment.
Get the notebooks that exist under the folders defined in the test_paths_file.
It only returns notebooks that have differences from the Git base_branch.
"""
test_paths = []
@@ -176,14 +227,47 @@ def run_changed_notebooks(
notebooks = notebooks.decode("utf-8").split("\n")
notebooks = [notebook for notebook in notebooks if notebook.endswith(".ipynb")]
notebooks = [notebook for notebook in notebooks if len(notebook) > 0]
notebooks = [notebook for notebook in notebooks if Path(notebook).exists()]
notebooks = [notebook for notebook in notebooks if pathlib.Path(notebook).exists()]
# Create paths
artifacts_path = Path(output_folder)
artifacts_path.mkdir(parents=True, exist_ok=True)
artifacts_path.joinpath("success").mkdir(parents=True, exist_ok=True)
artifacts_path.joinpath("failure").mkdir(parents=True, exist_ok=True)
return notebooks
def process_and_execute_notebooks(
notebooks: List[str],
container_uri: str,
staging_bucket: str,
artifacts_bucket: str,
variable_project_id: str,
variable_region: str,
private_pool_id: Optional[str],
should_parallelize: bool,
):
"""
Run the notebooks that exist under the folders defined in the test_paths_file.
It only runs notebooks that have differences from the Git base_branch.
The executed notebooks are saved in the artifacts_bucket.
Variables are also injected into the notebooks such as the variable_project_id and variable_region.
Args:
test_paths_file (str):
Required. The new-line delimited file to folders and files that need checking.
Folders are checked recursively.
base_branch (str):
Optional. If provided, only the files that have changed from the base_branch will be checked.
If not provided, all files will be checked.
staging_bucket (str):
Required. The GCS staging bucket to write source code to.
artifacts_bucket (str):
Required. The GCS staging bucket to write executed notebooks to.
variable_project_id (str):
Required. The value for PROJECT_ID to inject into notebooks.
variable_region (str):
Required. The value for REGION to inject into notebooks.
should_parallelize (bool):
Required. Should run notebooks in parallel using a thread pool as opposed to in sequence.
"""
notebook_execution_results: List[NotebookExecutionResult] = []
if len(notebooks) > 0:
@@ -197,25 +281,27 @@ def run_changed_notebooks(
notebook_execution_results = list(
executor.map(
functools.partial(
execute_notebook,
artifacts_path,
process_and_execute_notebook,
container_uri,
staging_bucket,
artifacts_bucket,
variable_project_id,
variable_region,
False,
should_use_separate_kernels,
private_pool_id,
),
notebooks,
)
)
else:
notebook_execution_results = [
execute_notebook(
artifacts_path=artifacts_path,
process_and_execute_notebook(
container_uri=container_uri,
staging_bucket=staging_bucket,
artifacts_bucket=artifacts_bucket,
variable_project_id=variable_project_id,
variable_region=variable_region,
private_pool_id=private_pool_id,
notebook=notebook,
should_log_output=True,
should_use_new_kernel=should_use_separate_kernels,
)
for notebook in notebooks
]
@@ -224,89 +310,38 @@ def run_changed_notebooks(
print("\n=== RESULTS ===\n")
notebooks_sorted = sorted(
results_sorted = sorted(
notebook_execution_results,
key=lambda result: result.is_pass,
reverse=True,
)
# Print results
print(
tabulate(
[
[
os.path.basename(os.path.normpath(result.notebook)),
result.name,
"PASSED" if result.is_pass else "FAILED",
format_timedelta(result.duration),
result.error_message or "--",
result.log_url,
]
for result in notebooks_sorted
for result in results_sorted
],
headers=["file", "status", "duration", "error"],
headers=["build_tag", "status", "duration", "log_url"],
)
)
print("\n=== END RESULTS===\n")
total_notebook_duration = functools.reduce(
operator.add,
[datetime.timedelta(seconds=0)]
+ [result.duration for result in results_sorted],
)
parser = argparse.ArgumentParser(description="Run changed notebooks.")
parser.add_argument(
"--test_paths_file",
type=pathlib.Path,
help="The path to the file that has newline-limited folders of notebooks that should be tested.",
required=True,
)
parser.add_argument(
"--base_branch",
help="The base git branch to diff against to find changed files.",
required=False,
)
parser.add_argument(
"--output_folder",
type=pathlib.Path,
help="The path to the folder to store executed notebooks.",
required=True,
)
parser.add_argument(
"--variable_project_id",
type=str,
help="The GCP project id. This is used to inject a variable value into the notebook before running.",
required=True,
)
parser.add_argument(
"--variable_region",
type=str,
help="The GCP region. This is used to inject a variable value into the notebook before running.",
required=True,
)
print(f"Cumulative notebook duration: {format_timedelta(total_notebook_duration)}")
# Note: Dependencies don't install correctly when this is set to True
parser.add_argument(
"--should_parallelize",
type=str2bool,
nargs="?",
const=True,
default=False,
help="Should run notebooks in parallel.",
)
# Note: This isn't guaranteed to work correctly due to existing Papermill issue
# See https://github.com/nteract/papermill/issues/625
parser.add_argument(
"--should_use_separate_kernels",
type=str2bool,
nargs="?",
const=True,
default=False,
help="(Experimental) Should run each notebook in a separate and independent virtual environment.",
)
args = parser.parse_args()
run_changed_notebooks(
test_paths_file=args.test_paths_file,
base_branch=args.base_branch,
output_folder=args.output_folder,
variable_project_id=args.variable_project_id,
variable_region=args.variable_region,
should_parallelize=args.should_parallelize,
should_use_separate_kernels=args.should_use_separate_kernels,
)
# Raise error if any notebooks failed
if not all([result.is_pass for result in results_sorted]):
raise RuntimeError("Notebook failures detected. See logs for details")
+40
View File
@@ -0,0 +1,40 @@
#!/usr/bin/env python
# Copyright 2021 Google LLC
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
"""A CLI to download (optional) and run a single notebook locally"""
import argparse
import execute_notebook_helper
parser = argparse.ArgumentParser(description="Run a single notebook locally.")
parser.add_argument(
"--notebook_source",
type=str,
help="Local filepath or GCS URI to notebook.",
required=True,
)
parser.add_argument(
"--output_file_or_uri",
type=str,
help="Local file or GCS URI to save executed notebook to.",
required=True,
)
args = parser.parse_args()
execute_notebook_helper.execute_notebook(
notebook_source=args.notebook_source,
output_file_or_uri=args.output_file_or_uri,
should_log_output=True,
)
+91
View File
@@ -0,0 +1,91 @@
#!/usr/bin/env python
# Copyright 2021 Google LLC
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
"""Methods to run a notebook locally"""
import sys
import os
import errno
import papermill as pm
import shutil
from utils import util
from google.cloud.aiplatform import utils
# This script is used to execute a notebook and write out the output notebook.
def execute_notebook(
notebook_source: str,
output_file_or_uri: str,
should_log_output: bool,
):
"""Execute a single notebook using Papermill"""
file_name = os.path.basename(os.path.normpath(notebook_source))
# Download notebook if it's a GCS URI
if notebook_source.startswith("gs://"):
# Extract uri components
bucket_name, prefix = utils.extract_bucket_and_prefix_from_gcs_path(
notebook_source
)
# Download remote notebook to local file system
notebook_source = file_name
util.download_file(
bucket_name=bucket_name, blob_name=prefix, destination_file=notebook_source
)
execution_exception = None
# Execute notebook
try:
# Execute notebook
pm.execute_notebook(
input_path=notebook_source,
output_path=notebook_source,
progress_bar=should_log_output,
request_save_on_cell_execute=should_log_output,
log_output=should_log_output,
stdout_file=sys.stdout if should_log_output else None,
stderr_file=sys.stderr if should_log_output else None,
)
except Exception as exception:
execution_exception = exception
finally:
# Copy executed notebook
if output_file_or_uri.startswith("gs://"):
# Upload to GCS path
util.upload_file(notebook_source, remote_file_path=output_file_or_uri)
print("\n=== EXECUTION FINISHED ===\n")
print(
f"Please debug the executed notebook by downloading: {output_file_or_uri}"
)
print("\n======\n")
else:
# Create directories if they don't exist
if not os.path.exists(os.path.dirname(output_file_or_uri)):
try:
os.makedirs(os.path.dirname(output_file_or_uri))
except OSError as exc: # Guard against race condition
if exc.errno != errno.EEXIST:
raise
print(f"Writing output to: {output_file_or_uri}")
shutil.move(notebook_source, output_file_or_uri)
if execution_exception:
raise execution_exception
+101
View File
@@ -0,0 +1,101 @@
#!/usr/bin/env python
# Copyright 2021 Google LLC
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
"""Methods to run a notebook on Google Cloud Build"""
from re import sub
from google.protobuf import duration_pb2
from yaml.loader import FullLoader
import google.auth
from google.cloud.devtools import cloudbuild_v1
from google.cloud.devtools.cloudbuild_v1.types import Source, StorageSource
from typing import Optional
import yaml
from google.cloud.aiplatform import utils
from google.api_core import operation, client_options
CLOUD_BUILD_FILEPATH = ".cloud-build/notebook-execution-test-cloudbuild-single.yaml"
TIMEOUT_IN_SECONDS = 86400
SERVICE_BASE_PATH = "cloudbuild.googleapis.com"
def execute_notebook_remote(
code_archive_uri: str,
notebook_uri: str,
notebook_output_uri: str,
container_uri: str,
region: str,
private_pool_id: Optional[str],
tag: Optional[str],
) -> operation.Operation:
"""Create and execute a single notebook on Google Cloud Build"""
# Load build steps from YAML
cloudbuild_config = yaml.load(open(CLOUD_BUILD_FILEPATH), Loader=FullLoader)
substitutions = {
"_PYTHON_IMAGE": container_uri,
"_NOTEBOOK_GCS_URI": notebook_uri,
"_NOTEBOOK_OUTPUT_GCS_URI": notebook_output_uri,
}
build = cloudbuild_v1.Build()
options: Optional[client_options.ClientOptions] = None
if private_pool_id:
substitutions["_PRIVATE_POOL_NAME"] = private_pool_id
build.options = cloudbuild_config["options"]
# Switch to the regional endpoint of the pool
options = client_options.ClientOptions(
api_endpoint=f"{region}-{SERVICE_BASE_PATH}"
)
# Authorize the client with Google defaults
credentials, project_id = google.auth.default()
client = cloudbuild_v1.services.cloud_build.CloudBuildClient(client_options=options)
(
source_archived_file_gcs_bucket,
source_archived_file_gcs_object,
) = utils.extract_bucket_and_prefix_from_gcs_path(code_archive_uri)
build.source = Source(
storage_source=StorageSource(
bucket=source_archived_file_gcs_bucket,
object_=source_archived_file_gcs_object,
)
)
build.steps = cloudbuild_config["steps"]
build.substitutions = substitutions
build.timeout = duration_pb2.Duration(seconds=TIMEOUT_IN_SECONDS)
build.queue_ttl = duration_pb2.Duration(seconds=TIMEOUT_IN_SECONDS)
if tag:
build.tags = [tag]
operation = client.create_build(project_id=project_id, build=build)
# Print the in-progress operation
# print("IN PROGRESS:")
# print(operation.metadata)
# Print the completed status
# print("RESULT:", result.status)
return operation
@@ -0,0 +1,31 @@
steps:
# Show the gcloud info and check if gcloud exists
- name: ${_PYTHON_IMAGE}
entrypoint: /bin/sh
args:
- -c
- 'gcloud config list'
# Check the Python version
- name: ${_PYTHON_IMAGE}
entrypoint: /bin/sh
args:
- -c
- 'python3 .cloud-build/CheckPythonVersion.py'
# Install Python dependencies
- name: ${_PYTHON_IMAGE}
entrypoint: /bin/sh
args:
- -c
- 'python3 -m pip install -U pip && python3 -m pip install -U --user -r .cloud-build/requirements.txt'
# Install Python dependencies and run testing script
- name: ${_PYTHON_IMAGE}
entrypoint: /bin/sh
args:
- -c
- 'python3 -m pip install -U pip && python3 -m pip freeze && python3 .cloud-build/execute_notebook_cli.py --notebook_source "${_NOTEBOOK_GCS_URI}" --output_file_or_uri "${_NOTEBOOK_OUTPUT_GCS_URI}"'
env:
- 'IS_TESTING=1'
timeout: 86400s
options:
pool:
name: ${_PRIVATE_POOL_NAME}
@@ -19,26 +19,20 @@ steps:
- 'if [ -n "${_BASE_BRANCH}" ]; then git fetch origin "${_BASE_BRANCH}":refs/remotes/origin/"${_BASE_BRANCH}"; else echo "Skipping fetch."; fi'
# Install Python dependencies
- name: ${_PYTHON_IMAGE}
entrypoint: pip
args: ['install', '--upgrade', '--user', '--requirement', '.cloud-build/requirements.txt']
entrypoint: /bin/sh
args:
- -c
- 'python3 -m pip install -U pip && python3 -m pip install -U --user -r .cloud-build/requirements.txt'
# Install Python dependencies and run testing script
# TODO: Only pass in private_pool_id if it is set
- name: ${_PYTHON_IMAGE}
entrypoint: /bin/sh
args:
- -c
- 'python3 -m pip freeze && python3 .cloud-build/ExecuteChangedNotebooks.py --test_paths_file "${_TEST_PATHS_FILE}" --base_branch "${_FORCED_BASE_BRANCH}" --output_folder ${BUILD_ID} --variable_project_id ${PROJECT_ID} --variable_region ${_GCP_REGION}'
- 'python3 -m pip install -U pip && python3 -m pip freeze && python3 .cloud-build/execute_changed_notebooks_cli.py --test_paths_file "${_TEST_PATHS_FILE}" --base_branch "${_FORCED_BASE_BRANCH}" --container_uri ${_PYTHON_IMAGE} --staging_bucket ${_GCS_STAGING_BUCKET} --artifacts_bucket ${_GCS_STAGING_BUCKET}/executed_notebooks/PR_${_PR_NUMBER}/BUILD_${BUILD_ID} --variable_project_id ${PROJECT_ID} --variable_region ${_GCP_REGION} `if [ ! -z "${_PRIVATE_POOL_NAME}" ]; then echo "--private_pool_id ${_PRIVATE_POOL_NAME}"; fi`'
env:
- 'IS_TESTING=1'
# Manually copy artifacts to GCS
- name: gcr.io/cloud-builders/gsutil
entrypoint: /bin/sh
args:
- -c
- 'if [ $(ls -pR "/workspace/${BUILD_ID}" | grep -v / | grep -v ^$ | wc -l) -ne 0 ]; then gsutil -m -q rsync -r "/workspace/${BUILD_ID}" "gs://${_GCS_ARTIFACTS_BUCKET}/test-artifacts/PR_${_PR_NUMBER}/BUILD_${BUILD_ID}/"; else echo "No artifacts to copy."; fi'
# Fail if there is anything in the failure folder
- name: ${_PYTHON_IMAGE}
entrypoint: /bin/sh
args:
- -c
- 'echo "Download executed notebooks with this command: \"mkdir -p artifacts && gsutil rsync -r gs://${_GCS_ARTIFACTS_BUCKET}/test-artifacts/PR_${_PR_NUMBER}/BUILD_${BUILD_ID} artifacts/\"" && if [ "$(ls -A /workspace/${BUILD_ID}/failure | wc -l)" -ne 0 ]; then exit 1; else exit 0; fi'
timeout: 86400s
options:
pool:
name: ${_PRIVATE_POOL_NAME}
+11 -7
View File
@@ -1,8 +1,12 @@
ipython>=7.0
jupyter>=1.0
nbconvert>=6.0
papermill>=2.3
numpy>=1.19
pandas>=1.2
matplotlib>=3.4
ipython
numpy
jupyter
nbconvert
papermill
pandas
matplotlib
tabulate
google-cloud-aiplatform
google-cloud-storage
google-cloud-build
gcloud
+5
View File
@@ -0,0 +1,5 @@
notebooks/official/vizier/gapic-vizier-multi-objective-optimization.ipynb
notebooks/official/pipelines/lightweight_functions_component_io_kfp.ipynb
notebooks/official/matching_engine/intro-swivel.ipynb
notebooks/official/ml_metadata/sdk-metric-parameter-tracking-for-locally-trained-models.ipynb
notebooks/official/pipelines/metrics_viz_run_compare_kfp.ipynb
@@ -15,7 +15,7 @@
from nbconvert.preprocessors import Preprocessor
from typing import Dict
import UpdateNotebookVariables
from . import UpdateNotebookVariables as update_notebook_variables
class RemoveNoExecuteCells(Preprocessor):
@@ -41,7 +41,7 @@ class UpdateVariablesPreprocessor(Preprocessor):
# VARIABLE_NAME = '[description]'
for variable_name, variable_value in replacement_map.items():
content = UpdateNotebookVariables.get_updated_value(
content = update_notebook_variables.get_updated_value(
content=content,
variable_name=variable_name,
variable_value=variable_value,
View File
+60
View File
@@ -0,0 +1,60 @@
from datetime import datetime
from typing import Optional
from google.cloud import storage
from google.cloud.aiplatform import utils
from google.auth import credentials as auth_credentials
import os
import subprocess
import tarfile
import uuid
def download_file(bucket_name: str, blob_name: str, destination_file: str) -> str:
"""Copies a remote GCS file to a local path"""
remote_file_path = "".join(["gs://", "/".join([bucket_name, blob_name])])
subprocess.check_output(
["gsutil", "cp", remote_file_path, destination_file], encoding="UTF-8"
)
return destination_file
def upload_file(
local_file_path: str,
remote_file_path: str,
) -> str:
"""Copies a local file to a GCS path"""
subprocess.check_output(
["gsutil", "cp", local_file_path, remote_file_path], encoding="UTF-8"
)
return remote_file_path
def archive_code_and_upload(staging_bucket: str):
# Archive all source in current directory
unique_id = uuid.uuid4()
source_archived_file = f"source_archived_{unique_id}.tar.gz"
git_files = subprocess.check_output(
["git", "ls-tree", "-r", "HEAD", "--name-only"], encoding="UTF-8"
).split("\n")
with tarfile.open(source_archived_file, "w:gz") as tar:
for file in git_files:
if len(file) > 0 and os.path.exists(file):
tar.add(file)
# Upload archive to GCS bucket
source_archived_file_gcs = upload_file(
local_file_path=f"{source_archived_file}",
remote_file_path="/".join(
[staging_bucket, "code_archives", source_archived_file]
),
)
print(f"Uploaded source code archive to {source_archived_file_gcs}")
return source_archived_file_gcs
+10 -10
View File
@@ -1,18 +1,18 @@
If you are opening a PR for `Official Notebooks` under the [notebooks/official](https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/master/notebooks/official) folder, follow this mandatory checklist:
- [ ] Use the [notebook template](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/master/notebooks/notebook_template.ipynb) as a starting point.
If you are opening a PR for `Official Notebooks` under the [notebooks/official](https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/main/notebooks/official) folder, follow this mandatory checklist:
- [ ] Use the [notebook template](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/notebook_template.ipynb) as a starting point.
- [ ] Follow the style and grammar rules outlined in the above notebook template.
- [ ] Verify the notebook runs successfully in Colab since the automated tests cannot guarantee this even when it passes.
- [ ] Passes all the required automated checks. You can locally test for formatting and linting with these [instructions](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/master/docs/contributing.md#code-quality-checks).
- [ ] Passes all the required automated checks. You can locally test for formatting and linting with these [instructions](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/docs/contributing.md#code-quality-checks).
- [ ] You have consulted with a tech writer to see if tech writer review is necessary. If so, the notebook has been reviewed by a tech writer, and they have approved it.
- [ ] This notebook has been added to the [CODEOWNERS](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/master/docs/CODEOWNERS) file under `# Official Notebooks` section, pointing to the author or the author's team.
- [ ] This notebook has been added to the [CODEOWNERS](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/docs/CODEOWNERS) file under `# Official Notebooks` section, pointing to the author or the author's team.
- [ ] The Jupyter notebook cleans up any artifacts it has created (datasets, ML models, endpoints, etc) so as not to eat up unnecessary resources.
If you are opening a PR for `Community Notebooks` under the [notebooks/community](https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/master/notebooks/community) folder:
- [ ] This notebook has been added to the [CODEOWNERS](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/master/docs/CODEOWNERS) file under the `# Community Notebooks` section, pointing to the author or the author's team.
- [ ] Passes all the required formatting and linting checks. You can locally test with these [instructions](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/master/docs/contributing.md#code-quality-checks).
If you are opening a PR for `Community Notebooks` under the [notebooks/community](https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/main/notebooks/community) folder:
- [ ] This notebook has been added to the [CODEOWNERS](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/docs/CODEOWNERS) file under the `# Community Notebooks` section, pointing to the author or the author's team.
- [ ] Passes all the required formatting and linting checks. You can locally test with these [instructions](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/docs/contributing.md#code-quality-checks).
If you are opening a PR for `Community Content` under the [community-content](https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/master/community-content) folder:
If you are opening a PR for `Community Content` under the [community-content](https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/main/community-content) folder:
- [ ] Make sure your main `Content Directory Name` is descriptive, informative, and includes some of the key products and attributes of your content, so that it is differentiable from other content
- [ ] The main content directory has been added to the [CODEOWNERS](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/master/docs/CODEOWNERS) file under the `# Community Content` section, pointing to the author or the author's team.
- [ ] Passes all the required formatting and linting checks. You can locally test with these [instructions](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/master/docs/contributing.md#code-quality-checks).
- [ ] The main content directory has been added to the [CODEOWNERS](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/docs/CODEOWNERS) file under the `# Community Content` section, pointing to the author or the author's team.
- [ ] Passes all the required formatting and linting checks. You can locally test with these [instructions](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/docs/contributing.md#code-quality-checks).
+4 -4
View File
@@ -7,13 +7,13 @@ jobs:
runs-on: ubuntu-latest
steps:
- name: Set up Python
uses: actions/setup-python@v2
uses: actions/setup-python@v3
- name: Fetch pull request branch
uses: actions/checkout@v2
uses: actions/checkout@v3
with:
fetch-depth: 0
- name: Fetch base master branch
run: git fetch -u "$GITHUB_SERVER_URL/$GITHUB_REPOSITORY" master:master
- name: Fetch base main branch
run: git fetch -u "$GITHUB_SERVER_URL/$GITHUB_REPOSITORY" main:main
- name: Install requirements
run: python3 -m pip install -U -r .github/workflows/linter/requirements.txt
- name: Format and lint notebooks
+5 -5
View File
@@ -2,8 +2,8 @@ git+https://github.com/tensorflow/docs
ipython
jupyter
nbconvert
black==20.8b1
pyupgrade==2.7.3
isort==5.6.4
flake8==3.9.0
nbqa==0.6.0
black==22.1.0
pyupgrade==2.31.1
isort==5.10.1
flake8==4.0.1
nbqa==1.3.1
+6 -6
View File
@@ -13,7 +13,7 @@
# See the License for the specific language governing permissions and
# limitations under the License.
# This script automatically formats and lints all notebooks that have changed from the head of the master branch.
# This script automatically formats and lints all notebooks that have changed from the head of the main branch.
#
# Options:
# -t: Test-mode. Only test if format and linting are required but make no changes to files.
@@ -52,7 +52,7 @@ echo "Test mode: $is_test"
notebooks=()
while read -r file || [ -n "$line" ]; do
notebooks+=("$file")
done < <(git diff --name-only master... | grep '\.ipynb$')
done < <(git diff --name-only main... | grep '\.ipynb$')
problematic_notebooks=()
if [ ${#notebooks[@]} -gt 0 ]; then
@@ -84,19 +84,19 @@ if [ ${#notebooks[@]} -gt 0 ]; then
FLAKE8_RTN=$?
else
echo "Running black..."
python3 -m nbqa black "$notebook" --nbqa-mutate
python3 -m nbqa black "$notebook"
BLACK_RTN=$?
echo "Running pyupgrade..."
python3 -m nbqa pyupgrade "$notebook" --nbqa-mutate
python3 -m nbqa pyupgrade "$notebook"
PYUPGRADE_RTN=$?
echo "Running isort..."
python3 -m nbqa isort "$notebook" --nbqa-mutate
python3 -m nbqa isort "$notebook"
ISORT_RTN=$?
echo "Running nbfmt..."
python3 -m tensorflow_docs.tools.nbfmt --remove_outputs "$notebook"
NBFMT_RTN=$?
echo "Running flake8..."
python3 -m nbqa flake8 "$notebook" --show-source --extend-ignore=W391,E501,F821,E402,F404,W503,E203,E722,W293,W291 --nbqa-mutate
python3 -m nbqa flake8 "$notebook" --show-source --extend-ignore=W391,E501,F821,E402,F404,W503,E203,E722,W293,W291
FLAKE8_RTN=$?
fi
+6
View File
@@ -0,0 +1,6 @@
# See https://help.github.com/en/articles/about-code-owners
# for more info about CODEOWNERS file.
# These owners will be the default owners for everything in
# the repo. Unless a later match takes precedence.
* @GoogleCloudPlatform/vertex-ai-samples-owners
+43
View File
@@ -0,0 +1,43 @@
# Contributor Code of Conduct
As contributors and maintainers of this project,
and in the interest of fostering an open and welcoming community,
we pledge to respect all people who contribute through reporting issues,
posting feature requests, updating documentation,
submitting pull requests or patches, and other activities.
We are committed to making participation in this project
a harassment-free experience for everyone,
regardless of level of experience, gender, gender identity and expression,
sexual orientation, disability, personal appearance,
body size, race, ethnicity, age, religion, or nationality.
Examples of unacceptable behavior by participants include:
* The use of sexualized language or imagery
* Personal attacks
* Trolling or insulting/derogatory comments
* Public or private harassment
* Publishing other's private information,
such as physical or electronic
addresses, without explicit permission
* Other unethical or unprofessional conduct.
Project maintainers have the right and responsibility to remove, edit, or reject
comments, commits, code, wiki edits, issues, and other contributions
that are not aligned to this Code of Conduct.
By adopting this Code of Conduct,
project maintainers commit themselves to fairly and consistently
applying these principles to every aspect of managing this project.
Project maintainers who do not follow or enforce the Code of Conduct
may be permanently removed from the project team.
This code of conduct applies both within project spaces and in public spaces
when an individual is representing the project or its community.
Instances of abusive, harassing, or otherwise unacceptable behavior
may be reported by opening an issue
or contacting one or more of the project maintainers.
This Code of Conduct is adapted from the [Contributor Covenant](http://contributor-covenant.org), version 1.2.0,
available at [http://contributor-covenant.org/version/1/2/0/](http://contributor-covenant.org/version/1/2/0/)
+10 -2
View File
@@ -23,11 +23,19 @@ you can follow these steps.
From a command-line terminal (e.g. from Vertex Workbench or locally), install
the code analysis tools:
```shell
pip install -U nbqa black flake8 isort pyupgrade git+https://github.com/tensorflow/docs
pip3 install --user -U nbqa black flake8 isort pyupgrade git+https://github.com/tensorflow/docs
```
You'll likely need to add the directory where these were installed to your PATH:
```shell
export PATH="$HOME/.local/bin:$PATH"
```
Then, set an environment variable for your notebook (or directory):
```shell
export notebook="your-notebook.ipynb"
```
@@ -35,6 +43,7 @@ export notebook="your-notebook.ipynb"
Finally, run this code block to check for errors. Each step will attempt to
automatically fix any issues. If the fixes can't be performed automatically,
then you will need to manually address them before submitting your PR.
```shell
nbqa black "$notebook"
nbqa pyupgrade "$notebook"
@@ -43,7 +52,6 @@ python3 -m tensorflow_docs.tools.nbfmt --remove_outputs "$notebook"
nbqa flake8 "$notebook" --extend-ignore=W391,E501,F821,E402,F404,W503,E203,E722,W293,W291
```
## Code Reviews
All submissions, including submissions by project members, require review. We
+19 -2
View File
@@ -6,15 +6,32 @@ Welcome to the Google Cloud [Vertex AI](https://cloud.google.com/vertex-ai/docs/
## Overview
The repository contains [Notebooks](https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/main/notebooks) and [Community Content](https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/master/community-content) that demonstrate how to develop and manage ML workflows using Google Cloud Vertex AI.
The repository contains [notebooks](https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/master/notebooks) and [community content](https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/master/community-content) that demonstrate how to develop and manage ML workflows using Google Cloud Vertex AI.
## Repository structure
```bash
├── community-content - Sample code and tutorials contributed by the community
├── notebooks
│ ├── community - Notebooks contributed by the community
│ ├── official - Notebooks demonstrating use of each Vertex AI service
│ │ ├── automl
│ │ ├── custom
│ │ ├── ...
```
## Contributing
Contributions welcome! See the [Contributing Guide](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/docs/contributing.md).
Contributions welcome! See the [Contributing Guide](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/master/CONTRIBUTING.md).
## Getting help
Please use the [issues page](https://github.com/GoogleCloudPlatform/vertex-ai-samples/issues) to provide feedback or submit a bug report.
## Disclaimer
This is not an officially supported Google product. The code in this repository is for demonstrative purposes only.
## Feedback
Please feel free to fill out our [survey](https://bit.ly/vertex-ai-samples-survey) to give us feedback on the repo and its content.
+7
View File
@@ -0,0 +1,7 @@
# Security Policy
To report a security issue, please use [g.co/vulnz](https://g.co/vulnz).
The Google Security Team will respond within 5 working days of your report on g.co/vulnz.
We use g.co/vulnz for our intake, and do coordination and disclosure here using GitHub Security Advisory to privately discuss and fix the issue.
+2
View File
@@ -1,3 +1,5 @@
* @vertex-ai-samples-contributors @GoogleCloudPlatform/cloudml-samples-owners
/tf_agents_bandits_movie_recommendation_with_kfp_and_vertex_sdk @yinghsienwu
/pytorch_text_classification_using_vertex_sdk_and_gcloud @RajeshThallam
/pytorch_text_classification_using_vertex_sdk_and_gcloud @RajeshThallam @ultrons
/sklearn_text_classification_from_script_using_vertex_sdk @maxhardt
@@ -0,0 +1,824 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {
"id": "pc5-mbsX9PZC"
},
"source": [
"# AlphaFold On Vertex AI Workbench\n",
"\n",
"[Vertex AI Workbench](https://cloud.google.com/vertex-ai/docs/workbench) offers an end-to-end notebook-based production environment that can be preconfigured with the runtime dependencies necessary to run AlphaFold on Vertex AI. With [User-Managed Notebooks](https://cloud.google.com/vertex-ai/docs/workbench/user-managed/introduction), you can configure a GPU accelerator to run AlphaFold using Tensorflow, without having to install and manage drivers or JupyterLab instances. This notebook allows you to easily predict the structure of a protein using a slightly simplified version of [AlphaFold v2.1.0](https://doi.org/10.1038/s41586-021-03819-2). \n",
"\n",
"## ![](https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/main/community-content/alphafold_on_workbench/vertexai_40.png) [Launch this Notebook in Vertex AI Workbench](https://console.cloud.google.com/vertex-ai/workbench/deploy-notebook?download_url=https://github.com/GoogleCloudPlatform/vertex-ai-samples/raw/main/community-content/alphafold_on_workbench/AlphaFold.ipynb)\n",
"\n",
"**Differences to AlphaFold v2.1.0**\n",
"\n",
"In comparison to AlphaFold v2.1.0, this notebook notebook uses **no templates (homologous structures)** and a selected portion of the [BFD database](https://bfd.mmseqs.com/). We have validated these changes on several thousand recent PDB structures. While accuracy will be near-identical to the full AlphaFold system on many targets, a small fraction have a large drop in accuracy due to the smaller MSA and lack of templates. For best reliability, we recommend instead using the [full open source AlphaFold](https://github.com/deepmind/alphafold/), or the [AlphaFold Protein Structure Database](https://alphafold.ebi.ac.uk/).\n",
"\n",
"**This notebook has an small drop in average accuracy for multimers compared to local AlphaFold installation, for full multimer accuracy it is highly recommended to run [AlphaFold locally](https://github.com/deepmind/alphafold#running-alphafold).** Moreover, the AlphaFold-Multimer requires searching for MSA for every unique sequence in the complex, hence it is substantially slower. If your notebook times-out due to slow multimer MSA search, we recommend running AlphaFold locally.\n",
"\n",
"Please note that this notebook is provided as an early-access prototype and is not a finished product. It is provided for theoretical modelling only and caution should be exercised in its use. \n",
"\n",
"**Citing this work**\n",
"\n",
"Any publication that discloses findings arising from using this notebook should [cite](https://github.com/deepmind/alphafold/#citing-this-work) the [AlphaFold paper](https://doi.org/10.1038/s41586-021-03819-2).\n",
"\n",
"**Licenses**\n",
"\n",
"This Colab uses the [AlphaFold model parameters](https://github.com/deepmind/alphafold/#model-parameters-license) which are subject to the Creative Commons Attribution 4.0 International ([CC BY 4.0](https://creativecommons.org/licenses/by/4.0/legalcode)) license. The Colab itself is provided under the [Apache 2.0 license](https://www.apache.org/licenses/LICENSE-2.0). See the full license statement below.\n",
"\n",
"\n",
"**More information**\n",
"\n",
"You can find more information about how AlphaFold works in the following papers:\n",
"\n",
"* [AlphaFold methods paper](https://www.nature.com/articles/s41586-021-03819-2)\n",
"* [AlphaFold predictions of the human proteome paper](https://www.nature.com/articles/s41586-021-03828-1)\n",
"* [AlphaFold-Multimer paper](https://www.biorxiv.org/content/10.1101/2021.10.04.463034v1)\n",
"\n",
"FAQ on how to interpret AlphaFold predictions are [here](https://alphafold.ebi.ac.uk/faq)."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "b7a02613eb1a"
},
"source": [
"## Download AlphaFold Data"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"collapsed": true,
"jupyter": {
"source_hidden": true
},
"cellView": "form",
"id": "woIxeCPygt7K"
},
"outputs": [],
"source": [
"import os\n",
"import subprocess\n",
"import sys\n",
"\n",
"import alphafold.common\n",
"import tqdm.notebook\n",
"from IPython.utils import io\n",
"\n",
"TQDM_BAR_FORMAT = (\n",
" \"{l_bar}{bar}| {n_fmt}/{total_fmt} [elapsed: {elapsed} remaining: {remaining}]\"\n",
")\n",
"\n",
"SOURCE_URL = (\n",
" \"https://storage.googleapis.com/alphafold/alphafold_params_colab_2022-01-19.tar\"\n",
")\n",
"PARAMS_DIR = \"alphafold/data/params\"\n",
"PARAMS_PATH = os.path.join(PARAMS_DIR, os.path.basename(SOURCE_URL))\n",
"ALPHAFOLD_COMMON_DIR = os.path.dirname(alphafold.common.__file__)\n",
"\n",
"try:\n",
" with tqdm.notebook.tqdm(total=100, bar_format=TQDM_BAR_FORMAT) as pbar:\n",
" with io.capture_output() as captured:\n",
"\n",
" # Download and store stereo_chemical_props.txt\n",
" !mkdir -p ~/content/alphafold/alphafold/common\n",
" !mkdir -p /opt/conda/lib/python3.7/site-packages/alphafold/common/\n",
" !wget -q -P ~/content/alphafold/alphafold/common https://git.scicore.unibas.ch/schwede/openstructure/-/raw/7102c63615b64735c4941278d92b554ec94415f8/modules/mol/alg/src/stereo_chemical_props.txt\n",
" pbar.update(18)\n",
" !cp -f ~/content/alphafold/alphafold/common/stereo_chemical_props.txt \"{ALPHAFOLD_COMMON_DIR}\"\n",
"\n",
" # Download alphafold_params_colab_2021-10-27.tar\n",
" !mkdir --parents \"{PARAMS_DIR}\"\n",
" !wget -O \"{PARAMS_PATH}\" \"{SOURCE_URL}\"\n",
" pbar.update(27)\n",
"\n",
" # Un-tar alphafold_params_colab_2021-10-27.tar\n",
" !tar --extract --verbose --file=\"{PARAMS_PATH}\" --directory=\"{PARAMS_DIR}\" --preserve-permissions\n",
" # !rm \"{PARAMS_PATH}\"\n",
" pbar.update(55)\n",
"\n",
"except subprocess.CalledProcessError:\n",
" print(captured)\n",
" raise"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "d8926b7d5529"
},
"source": [
"## Configure GPU Acceleration"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"collapsed": true,
"jupyter": {
"source_hidden": true
},
"cellView": "form",
"id": "VzJ5iMjTtoZw"
},
"outputs": [],
"source": [
"# Confirm accelerator configuration\n",
"import jax\n",
"\n",
"if jax.local_devices()[0].platform == \"tpu\":\n",
" raise RuntimeError(\n",
" \"TPU runtime not supported. Please configure GPU acceleration on the VM.\"\n",
" )\n",
"elif jax.local_devices()[0].platform == \"cpu\":\n",
" print(\n",
" \"CPU-only runtime is not recommended, because prediction execution will be slow. For better performance, consider GPU acceleration on the VM.\"\n",
" )\n",
"else:\n",
" print(f\"Running with {jax.local_devices()[0].device_kind} GPU\")\n",
"\n",
"# Make sure all necessary environment variables are set.\n",
"import os\n",
"\n",
"os.environ[\"TF_FORCE_UNIFIED_MEMORY\"] = \"1\"\n",
"os.environ[\"XLA_PYTHON_CLIENT_MEM_FRACTION\"] = \"2.0\""
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "W4JpOs6oA-QS"
},
"source": [
"## Making a prediction\n",
"\n",
"Please paste the sequence of your protein in the text box below, then run the remaining cells via _Run_ > _Run Selected Cell and All Below_. You can also run the cells individually by pressing the _Play_ button on the left.\n",
"\n",
"Note that the search against databases and the actual prediction can take some time, from minutes to hours, depending on the length of the protein and what type of GPU you allocate (see FAQ below).\n",
"\n",
"To start, enter the amino acid sequence(s) to fold ⬇️\n",
"\n",
"If you enter only a single sequence, the monomer model will be used. If you enter multiple sequences, the multimer model will be used."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "b310d44229d0"
},
"outputs": [],
"source": [
"# Input sequences (type: str)\n",
"sequence_1 = \"MAAHKGAEHHHKAAEHHEQAAKHHHAAAEHHEKGEHEQAAHHADTAYAHHKHAEEHAAQAAKHDAEHHAPKPH\"\n",
"sequence_2 = \"\"\n",
"sequence_3 = \"\"\n",
"sequence_4 = \"\"\n",
"sequence_5 = \"\"\n",
"sequence_6 = \"\"\n",
"sequence_7 = \"\"\n",
"sequence_8 = \"\""
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"collapsed": true,
"jupyter": {
"source_hidden": true
},
"cellView": "form",
"id": "rowN0bVYLe9n"
},
"outputs": [],
"source": [
"from alphafold.notebooks import notebook_utils\n",
"\n",
"input_sequences = (\n",
" sequence_1,\n",
" sequence_2,\n",
" sequence_3,\n",
" sequence_4,\n",
" sequence_5,\n",
" sequence_6,\n",
" sequence_7,\n",
" sequence_8,\n",
")\n",
"\n",
"# If folding a complex target and all the input sequences are\n",
"# prokaryotic then set `is_prokaryotic` to `True`. Set to `False`\n",
"# otherwise or if the origin is unknown.\n",
"\n",
"is_prokaryote = False # @param {type:\"boolean\"}\n",
"\n",
"MIN_SINGLE_SEQUENCE_LENGTH = 16\n",
"MAX_SINGLE_SEQUENCE_LENGTH = 2500\n",
"MAX_MULTIMER_LENGTH = 2500\n",
"\n",
"# Validate the input.\n",
"sequences, model_type_to_use = notebook_utils.validate_input(\n",
" input_sequences=input_sequences,\n",
" min_length=MIN_SINGLE_SEQUENCE_LENGTH,\n",
" max_length=MAX_SINGLE_SEQUENCE_LENGTH,\n",
" max_multimer_length=MAX_MULTIMER_LENGTH,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "db551d4877ea"
},
"source": [
"## Search against genetic databases\n",
"\n",
"Once this cell has been executed, you will see statistics about the multiple sequence alignment (MSA) that will be used by AlphaFold. In particular, you’ll see how well each residue is covered by similar sequences in the MSA."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"collapsed": true,
"jupyter": {
"source_hidden": true
},
"cellView": "form",
"id": "2tTeTTsLKPjB"
},
"outputs": [],
"source": [
"import collections\n",
"import copy\n",
"import random\n",
"from concurrent import futures\n",
"from urllib import request\n",
"\n",
"import matplotlib.pyplot as plt\n",
"import numpy as np\n",
"import py3Dmol\n",
"from alphafold.common import protein\n",
"from alphafold.data import (feature_processing, msa_pairing, pipeline,\n",
" pipeline_multimer)\n",
"from alphafold.data.tools import jackhmmer\n",
"from alphafold.model import config, data, model\n",
"from alphafold.relax import relax, utils\n",
"from IPython import display\n",
"from ipywidgets import GridspecLayout, Output\n",
"\n",
"# Color bands for visualizing plddt\n",
"PLDDT_BANDS = [\n",
" (0, 50, \"#FF7D45\"),\n",
" (50, 70, \"#FFDB13\"),\n",
" (70, 90, \"#65CBF3\"),\n",
" (90, 100, \"#0053D6\"),\n",
"]\n",
"\n",
"# --- Find the closest source ---\n",
"test_url_pattern = (\n",
" \"https://storage.googleapis.com/alphafold-colab{:s}/latest/uniref90_2021_03.fasta.1\"\n",
")\n",
"ex = futures.ThreadPoolExecutor(3)\n",
"\n",
"\n",
"def fetch(source):\n",
" request.urlretrieve(test_url_pattern.format(source))\n",
" return source\n",
"\n",
"\n",
"fs = [ex.submit(fetch, source) for source in [\"\", \"-europe\", \"-asia\"]]\n",
"source = None\n",
"for f in futures.as_completed(fs):\n",
" source = f.result()\n",
" ex.shutdown()\n",
" break\n",
"\n",
"JACKHMMER_BINARY_PATH = \"/usr/bin/jackhmmer\"\n",
"DB_ROOT_PATH = f\"https://storage.googleapis.com/alphafold-colab{source}/latest/\"\n",
"# The z_value is the number of sequences in a database.\n",
"MSA_DATABASES = [\n",
" {\n",
" \"db_name\": \"uniref90\",\n",
" \"db_path\": f\"{DB_ROOT_PATH}uniref90_2021_03.fasta\",\n",
" \"num_streamed_chunks\": 59,\n",
" \"z_value\": 135_301_051,\n",
" },\n",
" {\n",
" \"db_name\": \"smallbfd\",\n",
" \"db_path\": f\"{DB_ROOT_PATH}bfd-first_non_consensus_sequences.fasta\",\n",
" \"num_streamed_chunks\": 17,\n",
" \"z_value\": 65_984_053,\n",
" },\n",
" {\n",
" \"db_name\": \"mgnify\",\n",
" \"db_path\": f\"{DB_ROOT_PATH}mgy_clusters_2019_05.fasta\",\n",
" \"num_streamed_chunks\": 71,\n",
" \"z_value\": 304_820_129,\n",
" },\n",
"]\n",
"\n",
"# Search UniProt and construct the all_seq features only for heteromers, not homomers.\n",
"if model_type_to_use == notebook_utils.ModelType.MULTIMER and len(set(sequences)) > 1:\n",
" MSA_DATABASES.extend(\n",
" [\n",
" # Swiss-Prot and TrEMBL are concatenated together as UniProt.\n",
" {\n",
" \"db_name\": \"uniprot\",\n",
" \"db_path\": f\"{DB_ROOT_PATH}uniprot_2021_03.fasta\",\n",
" \"num_streamed_chunks\": 98,\n",
" \"z_value\": 219_174_961 + 565_254,\n",
" },\n",
" ]\n",
" )\n",
"\n",
"TOTAL_JACKHMMER_CHUNKS = sum(cfg[\"num_streamed_chunks\"] for cfg in MSA_DATABASES)\n",
"\n",
"MAX_HITS = {\n",
" \"uniref90\": 10_000,\n",
" \"smallbfd\": 5_000,\n",
" \"mgnify\": 501,\n",
" \"uniprot\": 50_000,\n",
"}\n",
"\n",
"\n",
"def get_msa(fasta_path):\n",
" \"\"\"Searches for MSA for the given sequence using chunked Jackhmmer search.\"\"\"\n",
"\n",
" # Run the search against chunks of genetic databases.\n",
" raw_msa_results = collections.defaultdict(list)\n",
" with tqdm.notebook.tqdm(\n",
" total=TOTAL_JACKHMMER_CHUNKS, bar_format=TQDM_BAR_FORMAT\n",
" ) as pbar:\n",
"\n",
" def jackhmmer_chunk_callback(i):\n",
" pbar.update(n=1)\n",
"\n",
" for db_config in MSA_DATABASES:\n",
" db_name = db_config[\"db_name\"]\n",
" pbar.set_description(f\"Searching {db_name}\")\n",
" jackhmmer_runner = jackhmmer.Jackhmmer(\n",
" binary_path=JACKHMMER_BINARY_PATH,\n",
" database_path=db_config[\"db_path\"],\n",
" get_tblout=True,\n",
" num_streamed_chunks=db_config[\"num_streamed_chunks\"],\n",
" streaming_callback=jackhmmer_chunk_callback,\n",
" z_value=db_config[\"z_value\"],\n",
" )\n",
" # Group the results by database name.\n",
" raw_msa_results[db_name].extend(jackhmmer_runner.query(fasta_path))\n",
"\n",
" return raw_msa_results\n",
"\n",
"\n",
"features_for_chain = {}\n",
"raw_msa_results_for_sequence = {}\n",
"for sequence_index, sequence in enumerate(sequences, start=1):\n",
" print(f\"\\nGetting MSA for sequence {sequence_index}\")\n",
"\n",
" fasta_path = f\"target_{sequence_index}.fasta\"\n",
" with open(fasta_path, \"wt\") as f:\n",
" f.write(f\">query\\n{sequence}\")\n",
"\n",
" # Don't do redundant work for multiple copies of the same chain in the multimer.\n",
" if sequence not in raw_msa_results_for_sequence:\n",
" raw_msa_results = get_msa(fasta_path=fasta_path)\n",
" raw_msa_results_for_sequence[sequence] = raw_msa_results\n",
" else:\n",
" raw_msa_results = copy.deepcopy(raw_msa_results_for_sequence[sequence])\n",
"\n",
" # Extract the MSAs from the Stockholm files.\n",
" # NB: deduplication happens later in pipeline.make_msa_features.\n",
" single_chain_msas = []\n",
" uniprot_msa = None\n",
" for db_name, db_results in raw_msa_results.items():\n",
" merged_msa = notebook_utils.merge_chunked_msa(\n",
" results=db_results, max_hits=MAX_HITS.get(db_name)\n",
" )\n",
" if merged_msa.sequences and db_name != \"uniprot\":\n",
" single_chain_msas.append(merged_msa)\n",
" msa_size = len(set(merged_msa.sequences))\n",
" print(\n",
" f\"{msa_size} unique sequences found in {db_name} for sequence {sequence_index}\"\n",
" )\n",
" elif merged_msa.sequences and db_name == \"uniprot\":\n",
" uniprot_msa = merged_msa\n",
"\n",
" notebook_utils.show_msa_info(\n",
" single_chain_msas=single_chain_msas, sequence_index=sequence_index\n",
" )\n",
"\n",
" # Turn the raw data into model features.\n",
" feature_dict = {}\n",
" feature_dict.update(\n",
" pipeline.make_sequence_features(\n",
" sequence=sequence, description=\"query\", num_res=len(sequence)\n",
" )\n",
" )\n",
" feature_dict.update(pipeline.make_msa_features(msas=single_chain_msas))\n",
" # We don't use templates in AlphaFold notebook, add only empty placeholder features.\n",
" feature_dict.update(\n",
" notebook_utils.empty_placeholder_template_features(\n",
" num_templates=0, num_res=len(sequence)\n",
" )\n",
" )\n",
"\n",
" # Construct the all_seq features only for heteromers, not homomers.\n",
" if (\n",
" model_type_to_use == notebook_utils.ModelType.MULTIMER\n",
" and len(set(sequences)) > 1\n",
" ):\n",
" valid_feats = msa_pairing.MSA_FEATURES + (\n",
" \"msa_uniprot_accession_identifiers\",\n",
" \"msa_species_identifiers\",\n",
" )\n",
" all_seq_features = {\n",
" f\"{k}_all_seq\": v\n",
" for k, v in pipeline.make_msa_features([uniprot_msa]).items()\n",
" if k in valid_feats\n",
" }\n",
" feature_dict.update(all_seq_features)\n",
"\n",
" features_for_chain[protein.PDB_CHAIN_IDS[sequence_index - 1]] = feature_dict\n",
"\n",
"\n",
"# Do further feature post-processing depending on the model type.\n",
"if model_type_to_use == notebook_utils.ModelType.MONOMER:\n",
" np_example = features_for_chain[protein.PDB_CHAIN_IDS[0]]\n",
"\n",
"elif model_type_to_use == notebook_utils.ModelType.MULTIMER:\n",
" all_chain_features = {}\n",
" for chain_id, chain_features in features_for_chain.items():\n",
" all_chain_features[chain_id] = pipeline_multimer.convert_monomer_features(\n",
" chain_features, chain_id\n",
" )\n",
"\n",
" all_chain_features = pipeline_multimer.add_assembly_features(all_chain_features)\n",
"\n",
" np_example = feature_processing.pair_and_merge(\n",
" all_chain_features=all_chain_features, is_prokaryote=is_prokaryote\n",
" )\n",
"\n",
" # Pad MSA to avoid zero-sized extra_msa.\n",
" np_example = pipeline_multimer.pad_msa(np_example, min_num_seq=512)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "9640643486bd"
},
"source": [
"## Run AlphaFold\n",
"\n",
"Once this cell has been executed, a zip-archive \"prediction.zip\" with the obtained prediction will be saved on the VM, and available for download to your computer in the sidebar. In case you are having issues with the relaxation stage, you can disable it below. Warning: This means that the prediction might have distracting small stereochemical violations."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"collapsed": true,
"jupyter": {
"source_hidden": true
},
"cellView": "form",
"id": "XUo6foMQxwS2"
},
"outputs": [],
"source": [
"run_relax = True\n",
"\n",
"# --- Run the model ---\n",
"if model_type_to_use == notebook_utils.ModelType.MONOMER:\n",
" model_names = config.MODEL_PRESETS[\"monomer\"] + (\"model_2_ptm\",)\n",
"elif model_type_to_use == notebook_utils.ModelType.MULTIMER:\n",
" model_names = config.MODEL_PRESETS[\"multimer\"]\n",
"\n",
"output_dir = \"prediction\"\n",
"os.makedirs(output_dir, exist_ok=True)\n",
"\n",
"plddts = {}\n",
"ranking_confidences = {}\n",
"pae_outputs = {}\n",
"unrelaxed_proteins = {}\n",
"\n",
"with tqdm.notebook.tqdm(total=len(model_names) + 1, bar_format=TQDM_BAR_FORMAT) as pbar:\n",
" for model_name in model_names:\n",
" pbar.set_description(f\"Running {model_name}\")\n",
"\n",
" cfg = config.model_config(model_name)\n",
" if model_type_to_use == notebook_utils.ModelType.MONOMER:\n",
" cfg.data.eval.num_ensemble = 1\n",
" elif model_type_to_use == notebook_utils.ModelType.MULTIMER:\n",
" cfg.model.num_ensemble_eval = 1\n",
" params = data.get_model_haiku_params(model_name, \"./alphafold/data\")\n",
" model_runner = model.RunModel(cfg, params)\n",
" processed_feature_dict = model_runner.process_features(\n",
" np_example, random_seed=0\n",
" )\n",
" prediction = model_runner.predict(\n",
" processed_feature_dict, random_seed=random.randrange(sys.maxsize)\n",
" )\n",
"\n",
" mean_plddt = prediction[\"plddt\"].mean()\n",
"\n",
" if model_type_to_use == notebook_utils.ModelType.MONOMER:\n",
" if \"predicted_aligned_error\" in prediction:\n",
" pae_outputs[model_name] = (\n",
" prediction[\"predicted_aligned_error\"],\n",
" prediction[\"max_predicted_aligned_error\"],\n",
" )\n",
" else:\n",
" # Monomer models are sorted by mean pLDDT. Do not put monomer pTM models here as they\n",
" # should never get selected.\n",
" ranking_confidences[model_name] = prediction[\"ranking_confidence\"]\n",
" plddts[model_name] = prediction[\"plddt\"]\n",
" elif model_type_to_use == notebook_utils.ModelType.MULTIMER:\n",
" # Multimer models are sorted by pTM+ipTM.\n",
" ranking_confidences[model_name] = prediction[\"ranking_confidence\"]\n",
" plddts[model_name] = prediction[\"plddt\"]\n",
" pae_outputs[model_name] = (\n",
" prediction[\"predicted_aligned_error\"],\n",
" prediction[\"max_predicted_aligned_error\"],\n",
" )\n",
"\n",
" # Set the b-factors to the per-residue plddt.\n",
" final_atom_mask = prediction[\"structure_module\"][\"final_atom_mask\"]\n",
" b_factors = prediction[\"plddt\"][:, None] * final_atom_mask\n",
" unrelaxed_protein = protein.from_prediction(\n",
" processed_feature_dict,\n",
" prediction,\n",
" b_factors=b_factors,\n",
" remove_leading_feature_dimension=(\n",
" model_type_to_use == notebook_utils.ModelType.MONOMER\n",
" ),\n",
" )\n",
" unrelaxed_proteins[model_name] = unrelaxed_protein\n",
"\n",
" # Delete unused outputs to save memory.\n",
" del model_runner\n",
" del params\n",
" del prediction\n",
" pbar.update(n=1)\n",
"\n",
" # --- AMBER relax the best model ---\n",
"\n",
" # Find the best model according to the mean pLDDT.\n",
" best_model_name = max(\n",
" ranking_confidences.keys(), key=lambda x: ranking_confidences[x]\n",
" )\n",
"\n",
" if run_relax:\n",
" pbar.set_description(\"AMBER relaxation\")\n",
" amber_relaxer = relax.AmberRelaxation(\n",
" max_iterations=0,\n",
" tolerance=2.39,\n",
" stiffness=10.0,\n",
" exclude_residues=[],\n",
" max_outer_iterations=3,\n",
" )\n",
" relaxed_pdb, _, _ = amber_relaxer.process(\n",
" prot=unrelaxed_proteins[best_model_name]\n",
" )\n",
" else:\n",
" print(\"Warning: Running without the relaxation stage.\")\n",
" relaxed_pdb = protein.to_pdb(unrelaxed_proteins[best_model_name])\n",
" pbar.update(n=1) # Finished AMBER relax.\n",
"\n",
"# Construct multiclass b-factors to indicate confidence bands\n",
"# 0=very low, 1=low, 2=confident, 3=very high\n",
"banded_b_factors = []\n",
"for plddt in plddts[best_model_name]:\n",
" for idx, (min_val, max_val, _) in enumerate(PLDDT_BANDS):\n",
" if plddt >= min_val and plddt <= max_val:\n",
" banded_b_factors.append(idx)\n",
" break\n",
"banded_b_factors = np.array(banded_b_factors)[:, None] * final_atom_mask\n",
"to_visualize_pdb = utils.overwrite_b_factors(relaxed_pdb, banded_b_factors)\n",
"\n",
"\n",
"# Write out the prediction\n",
"pred_output_path = os.path.join(output_dir, \"selected_prediction.pdb\")\n",
"with open(pred_output_path, \"w\") as f:\n",
" f.write(relaxed_pdb)\n",
"\n",
"\n",
"# --- Visualise the prediction & confidence ---\n",
"show_sidechains = True\n",
"\n",
"\n",
"def plot_plddt_legend():\n",
" \"\"\"Plots the legend for pLDDT.\"\"\"\n",
" thresh = [\n",
" \"Very low (pLDDT < 50)\",\n",
" \"Low (70 > pLDDT > 50)\",\n",
" \"Confident (90 > pLDDT > 70)\",\n",
" \"Very high (pLDDT > 90)\",\n",
" ]\n",
"\n",
" colors = [x[2] for x in PLDDT_BANDS]\n",
"\n",
" plt.figure(figsize=(2, 2))\n",
" for c in colors:\n",
" plt.bar(0, 0, color=c)\n",
" plt.legend(thresh, frameon=False, loc=\"center\", fontsize=20)\n",
" plt.xticks([])\n",
" plt.yticks([])\n",
" ax = plt.gca()\n",
" ax.spines[\"right\"].set_visible(False)\n",
" ax.spines[\"top\"].set_visible(False)\n",
" ax.spines[\"left\"].set_visible(False)\n",
" ax.spines[\"bottom\"].set_visible(False)\n",
" plt.title(\"Model Confidence\", fontsize=20, pad=20)\n",
" return plt\n",
"\n",
"\n",
"# Show the structure coloured by chain if the multimer model has been used.\n",
"if model_type_to_use == notebook_utils.ModelType.MULTIMER:\n",
" multichain_view = py3Dmol.view(width=800, height=600)\n",
" multichain_view.addModelsAsFrames(to_visualize_pdb)\n",
" multichain_style = {\"cartoon\": {\"colorscheme\": \"chain\"}}\n",
" multichain_view.setStyle({\"model\": -1}, multichain_style)\n",
" multichain_view.zoomTo()\n",
" multichain_view.show()\n",
"\n",
"# Color the structure by per-residue pLDDT\n",
"color_map = {i: bands[2] for i, bands in enumerate(PLDDT_BANDS)}\n",
"view = py3Dmol.view(width=800, height=600)\n",
"view.addModelsAsFrames(to_visualize_pdb)\n",
"style = {\"cartoon\": {\"colorscheme\": {\"prop\": \"b\", \"map\": color_map}}}\n",
"if show_sidechains:\n",
" style[\"stick\"] = {}\n",
"view.setStyle({\"model\": -1}, style)\n",
"view.zoomTo()\n",
"\n",
"grid = GridspecLayout(1, 2)\n",
"out = Output()\n",
"with out:\n",
" view.show()\n",
"grid[0, 0] = out\n",
"\n",
"out = Output()\n",
"with out:\n",
" plot_plddt_legend().show()\n",
"grid[0, 1] = out\n",
"\n",
"display.display(grid)\n",
"\n",
"# Display pLDDT and predicted aligned error (if output by the model).\n",
"if pae_outputs:\n",
" num_plots = 2\n",
"else:\n",
" num_plots = 1\n",
"\n",
"plt.figure(figsize=[8 * num_plots, 6])\n",
"plt.subplot(1, num_plots, 1)\n",
"plt.plot(plddts[best_model_name])\n",
"plt.title(\"Predicted LDDT\")\n",
"plt.xlabel(\"Residue\")\n",
"plt.ylabel(\"pLDDT\")\n",
"\n",
"if num_plots == 2:\n",
" plt.subplot(1, 2, 2)\n",
" pae, max_pae = list(pae_outputs.values())[0]\n",
" plt.imshow(pae, vmin=0.0, vmax=max_pae, cmap=\"Greens_r\")\n",
" plt.colorbar(fraction=0.046, pad=0.04)\n",
"\n",
" # Display lines at chain boundaries.\n",
" best_unrelaxed_prot = unrelaxed_proteins[best_model_name]\n",
" total_num_res = best_unrelaxed_prot.residue_index.shape[-1]\n",
" chain_ids = best_unrelaxed_prot.chain_index\n",
" for chain_boundary in np.nonzero(chain_ids[:-1] - chain_ids[1:]):\n",
" if chain_boundary.size:\n",
" plt.plot([0, total_num_res], [chain_boundary, chain_boundary], color=\"red\")\n",
" plt.plot([chain_boundary, chain_boundary], [0, total_num_res], color=\"red\")\n",
"\n",
" plt.title(\"Predicted Aligned Error\")\n",
" plt.xlabel(\"Scored residue\")\n",
" plt.ylabel(\"Aligned residue\")\n",
"\n",
"# Save the predicted aligned error (if it exists).\n",
"pae_output_path = os.path.join(output_dir, \"predicted_aligned_error.json\")\n",
"if pae_outputs:\n",
" # Save predicted aligned error in the same format as the AF EMBL DB.\n",
" pae_data = notebook_utils.get_pae_json(pae=pae, max_pae=max_pae.item())\n",
" with open(pae_output_path, \"w\") as f:\n",
" f.write(pae_data)\n",
"\n",
"!zip -q -r {output_dir}.zip {output_dir}"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "lUQAn5LYC5n4"
},
"source": [
"### Interpreting the prediction\n",
"\n",
"In general predicted LDDT (pLDDT) is best used for intra-domain confidence, whereas Predicted Aligned Error (PAE) is best used for determining between domain or between chain confidence.\n",
"\n",
"Please see the [AlphaFold methods paper](https://www.nature.com/articles/s41586-021-03819-2), the [AlphaFold predictions of the human proteome paper](https://www.nature.com/articles/s41586-021-03828-1), and the [AlphaFold-Multimer paper](https://www.biorxiv.org/content/10.1101/2021.10.04.463034v1) as well as [our FAQ](https://alphafold.ebi.ac.uk/faq) on how to interpret AlphaFold predictions."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "jeb2z8DIA4om"
},
"source": [
"## FAQ & Troubleshooting\n",
"\n",
"\n",
"* How do I get a predicted protein structure for my protein?\n",
" * Connect the notebook to the Jupyter kernel \"Python 3 (ipykernel)\".\n",
" * Paste the amino acid sequence of your protein (without any headers) into the variable sequence_1 in \"Making a Prediction\".\n",
" * Run all cells in the notebook, either by running them individually or via \"Kernel\"/\"Restart Kernel and Run All Cells...\"\n",
" * The predicted protein structure will be downloaded once all cells have been executed. Note: This can take minutes to hours - see below.\n",
"* How long will this take?\n",
" * The search against genetic databases can take minutes to hours.\n",
" * Running AlphaFold and generating the prediction can take minutes to hours, depending on the length of your protein and on which GPU-type your VM has access to.\n",
"* My notebook no longer seems to be doing anything, what should I do?\n",
" * Some steps may take minutes to hours to complete.\n",
" * If nothing happens or if you receive an error message, try restarting your notebook runtime via \"Kernel\"/\"Restart Kernel and Run All Cells...\".\n",
" * If this doesn’t help, try resetting restarting your VM inside the GCloud Console (\"Compute Engine\"/\"VM Instances\").\n",
"* How does this compare to the open-source version of AlphaFold?\n",
" * This notebook version of AlphaFold searches a selected portion of the BFD dataset and currently doesn’t use templates, so its accuracy is reduced in comparison to the full version of AlphaFold that is described in the [AlphaFold paper](https://doi.org/10.1038/s41586-021-03819-2) and [Github repo](https://github.com/deepmind/alphafold/) (the full version is available via the inference script).\n",
"* I received a warning “Notebook requires high RAM”, what do I do?\n",
" * In the \"Compute Engine\"/\"VM Instances\" Console menu, you can reconfigure the host VM settings. See [Changing the machine type of a VM instance](https://cloud.google.com/compute/docs/instances/changing-machine-type-of-stopped-instance) for instructions.\n",
"* Does this tool install anything on my computer?\n",
" * No, everything happens in the VM instance within your Google Cloud project.\n",
"* How should I share feedback and bug reports?\n",
" * Please share any feedback and bug reports as an [issue](https://github.com/GoogleCloudPlatform/vertex-ai-samples/issues) on Github.\n",
"\n",
"\n",
"## Related work\n",
"\n",
"Take a look at these Colab notebooks provided by the community (please note that these notebooks may vary from our validated AlphaFold system and we cannot guarantee their accuracy):\n",
"\n",
"* The [ColabFold AlphaFold2 notebook](https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/AlphaFold2.ipynb) by Sergey Ovchinnikov, Milot Mirdita and Martin Steinegger, which uses an API hosted at the Södinglab based on the MMseqs2 server ([Mirdita et al. 2019, Bioinformatics](https://academic.oup.com/bioinformatics/article/35/16/2856/5280135)) for the multiple sequence alignment creation.\n"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "YfPhvYgKC81B"
},
"source": [
"# License and Disclaimer\n",
"\n",
"This is not an officially-supported Google product.\n",
"\n",
"This notebook and other information provided is for theoretical modelling only, caution should be exercised in its use. It is provided ‘as-is’ without any warranty of any kind, whether expressed or implied. Information is not intended to be a substitute for professional medical advice, diagnosis, or treatment, and does not constitute medical or other professional advice.\n",
"\n",
"Copyright 2021 DeepMind Technologies Limited.\n",
"\n",
"\n",
"## AlphaFold Code License\n",
"\n",
"Licensed under the Apache License, Version 2.0 (the \"License\"); you may not use this file except in compliance with the License. You may obtain a copy of the License at https://www.apache.org/licenses/LICENSE-2.0.\n",
"\n",
"Unless required by applicable law or agreed to in writing, software distributed under the License is distributed on an \"AS IS\" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. See the License for the specific language governing permissions and limitations under the License.\n",
"\n",
"## Model Parameters License\n",
"\n",
"The AlphaFold parameters are made available under the terms of the Creative Commons Attribution 4.0 International (CC BY 4.0) license. You can find details at: https://creativecommons.org/licenses/by/4.0/legalcode\n",
"\n",
"\n",
"## Third-party software\n",
"\n",
"Use of the third-party software, libraries or code referred to in the [Acknowledgements section](https://github.com/deepmind/alphafold/#acknowledgements) in the AlphaFold README may be governed by separate terms and conditions or license provisions. Your use of the third-party software, libraries or code is subject to any such terms and you should check that you can comply with any applicable restrictions or terms and conditions before use.\n",
"\n",
"\n",
"## Mirrored Databases\n",
"\n",
"The following databases have been mirrored by DeepMind, and are available with reference to the following:\n",
"* UniProt: v2021\\_03 (unmodified), by The UniProt Consortium, available under a [Creative Commons Attribution-NoDerivatives 4.0 International License](http://creativecommons.org/licenses/by-nd/4.0/).\n",
"* UniRef90: v2021\\_03 (unmodified), by The UniProt Consortium, available under a [Creative Commons Attribution-NoDerivatives 4.0 International License](http://creativecommons.org/licenses/by-nd/4.0/).\n",
"* MGnify: v2019\\_05 (unmodified), by Mitchell AL et al., available free of all copyright restrictions and made fully and freely available for both non-commercial and commercial use under [CC0 1.0 Universal (CC0 1.0) Public Domain Dedication](https://creativecommons.org/publicdomain/zero/1.0/).\n",
"* BFD: (modified), by Steinegger M. and Söding J., modified by DeepMind, available under a [Creative Commons Attribution-ShareAlike 4.0 International License](https://creativecommons.org/licenses/by/4.0/). See the Methods section of the [AlphaFold proteome paper](https://www.nature.com/articles/s41586-021-03828-1) for details."
]
}
],
"metadata": {
"accelerator": "GPU",
"colab": {
"collapsed_sections": [],
"name": "AlphaFold.ipynb",
"toc_visible": true
},
"kernelspec": {
"display_name": "Python 3",
"name": "python3"
}
},
"nbformat": 4,
"nbformat_minor": 0
}
@@ -0,0 +1,82 @@
# Copyright 2022 Google LLC
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
ARG CUDA_MAJOR=11
ARG CUDA_MINOR=0
FROM gcr.io/deeplearning-platform-release/base-cu110
ARG CUDA_MAJOR
ARG CUDA_MINOR
SHELL ["/bin/bash", "-c"]
RUN apt-get update && DEBIAN_FRONTEND=noninteractive apt-get install -y \
build-essential \
cmake \
cuda-command-line-tools-${CUDA_MAJOR}-${CUDA_MINOR} \
git \
hmmer \
kalign \
tzdata \
wget \
&& rm -rf /var/lib/apt/lists/*
# Compile HHsuite from source.
RUN git clone --branch v3.3.0 https://github.com/soedinglab/hh-suite.git /tmp/hh-suite \
&& mkdir /tmp/hh-suite/build \
&& pushd /tmp/hh-suite/build \
&& cmake -DCMAKE_INSTALL_PREFIX=/opt/hhsuite .. \
&& make -j 4 && make install \
&& ln -s /opt/hhsuite/bin/* /usr/bin \
&& popd \
&& rm -rf /tmp/hh-suite
ENV PATH="/opt/conda/bin:$PATH"
RUN conda update -qy conda \
&& conda install -y -c conda-forge \
openmm=7.5.1 \
cudatoolkit==${CUDA_VERSION} \
pdbfixer \
pip \
python=3.7
COPY . /app/alphafold
# Install pip packages.
RUN pip3 install --upgrade pip \
&& pip3 install -r /app/alphafold/requirements.txt \
&& pip3 install py3Dmol tqdm \
&& pip3 install --upgrade jax==0.2.14 jaxlib==0.1.69+cuda${CUDA_MAJOR}${CUDA_MINOR} -f \
https://storage.googleapis.com/jax-releases/jax_releases.html
# Install alphafold.
WORKDIR /app/alphafold
RUN python setup.py install
# Apply OpenMM patch.
WORKDIR /opt/conda/lib/python3.7/site-packages
RUN patch -p0 < /app/alphafold/docker/openmm.patch
# Creating a tmp location for jackhmmr; not mounting through to host though.
RUN sudo mkdir -m 777 --parents /tmp/ramdisk
# We need to run `ldconfig` first to ensure GPUs are visible, due to some quirk
# with Debian. See https://github.com/NVIDIA/nvidia-docker/issues/1399 for
# details.
# ENTRYPOINT does not support easily running multiple commands, so instead we
# write a shell script to wrap them up.
WORKDIR /home/jupyter
RUN echo '#!/bin/bash\nldconfig\n\'
+20
View File
@@ -0,0 +1,20 @@
#!/usr/bin/env bash
set -e
# Prod (Publicly viewable)
PROJECT=cloud-devrel-public-resources
REPOSITORY=alphafold
LOCAL_IMAGE=alphafold-on-gcp
REMOTE_IMAGE=${LOCAL_IMAGE?}
TAG=latest
REGISTRY="us-west1-docker.pkg.dev/${PROJECT?}/${REPOSITORY?}/${REMOTE_IMAGE?}:${TAG?}"
git clone https://github.com/deepmind/alphafold.git
cp Dockerfile alphafold/docker/Dockerfile
cp AlphaFold.ipynb alphafold/notebooks/AlphaFold.ipynb
cd alphafold && sudo docker build --tag ${LOCAL_IMAGE?}:${TAG?} -f docker/Dockerfile .
sudo docker tag ${LOCAL_IMAGE?}:${TAG?} ${REGISTRY?}
sudo docker push ${REGISTRY?}
Binary file not shown.

After

Width:  |  Height:  |  Size: 3.1 KiB

@@ -1,6 +1,6 @@
# PyTorch on Google Cloud: Text Classification
In the PyTorch on Google Cloud series of blog posts, we aim to share how to build, train and deploy PyTorch models at scale and how to create reproducible machine learning pipelines on Google Cloud with [Vertex AI](https://cloud.google.com/vertex-ai).
In the PyTorch on Google Cloud series of blog posts, we aim to share how to build, train, deploy and orchestrate PyTorch models at scale and how to create reproducible machine learning pipelines on Google Cloud with [Vertex AI](https://cloud.google.com/vertex-ai).
This tutorial on text classification shows how to train a PyTorch based text classification model by fine tuning a pre-trained Huggingface Transformers model and deploy the model on [Vertex AI](https://cloud.google.com/vertex-ai/docs/start/client-libraries#python) using Vertex SDK and [`gcloud ai`](https://cloud.google.com/sdk/gcloud/reference/beta/ai).
@@ -9,6 +9,7 @@ This tutorial on text classification shows how to train a PyTorch based text cla
| <h4>Notebook</h4> | <h4>Description</h4> |
| :-------- | :------- |
| [pytorch-text-classification-vertex-ai-train-tune-deploy.ipynb](./pytorch-text-classification-vertex-ai-train-tune-deploy.ipynb) | Notebook to show training, hyper-parameter tuning and deploying a PyTorch model on Vertex AI |
| [pytorch-text-classification-vertex-ai-pipelines.ipynb](./pytorch-text-classification-vertex-ai-pipelines.ipynb) | Notebook to show orchestration of PyTorch ML workflows on Vertex AI Pipelines using Kubeflow Pipelines SDK |
## Folders
@@ -1,6 +1,7 @@
# Use pytorch GPU base image
FROM gcr.io/cloud-aiplatform/training/pytorch-gpu.1-7
# FROM gcr.io/cloud-aiplatform/training/pytorch-gpu.1-7
FROM us-docker.pkg.dev/vertex-ai/training/pytorch-gpu.1-10:latest
# set working directory
WORKDIR /app
@@ -22,15 +22,18 @@ PROJECT_ID=$(gcloud config list --format 'value(core.project)')
# BUCKET_NAME: Change to your bucket name.
BUCKET_NAME="[your-bucket-name]" # <-- CHANGE TO YOUR BUCKET NAME
BUCKET_NAME=cloud-ai-platform-2f444b6a-a742-444b-b91a-c7519f51bd77
# validate bucket name
if [ "${BUCKET_NAME}" = "[your-bucket-name]" ]
then
echo "[ERROR] INVALID VALUE: Please update the variable BUCKET_NAME with valid Cloud Storage bucket name. Exiting the script..."
exit 1
fi
# JOB_NAME: the name of your job running on AI Platform.
JOB_PREFIX="finetuned-bert-classifier-pytorch-cstm-cntr-"
JOB_PREFIX="finetuned-bert-classifier-pytorch-cstm-cntr"
JOB_NAME=${JOB_PREFIX}-$(date +%Y%m%d%H%M%S)-custom-job
# This can be a GCS location to a zipped and uploaded package
PACKAGE_PATH=./trainer
# REGION: select a region from https://cloud.google.com/vertex-ai/docs/general/locations#available_regions
# or use the default '`us-central1`'. The region is where the job will be run.
REGION="us-central1"
@@ -41,11 +44,8 @@ JOB_DIR=gs://${BUCKET_NAME}/${JOB_PREFIX}/models/${JOB_NAME}
# IMAGE_REPO_NAME: set a local repo name to distinquish our image
IMAGE_REPO_NAME=pytorch_gpu_train_finetuned-bert-classifier
# IMAGE_TAG: an easily identifiable tag for your docker image
IMAGE_TAG=latest
# IMAGE_URI: the complete URI location for Cloud Container Registry
CUSTOM_TRAIN_IMAGE_URI=gcr.io/${PROJECT_ID}/${IMAGE_REPO_NAME}:${IMAGE_TAG}
CUSTOM_TRAIN_IMAGE_URI=gcr.io/${PROJECT_ID}/${IMAGE_REPO_NAME}
# Build the docker image
docker build --no-cache -f Dockerfile -t $CUSTOM_TRAIN_IMAGE_URI ../python_package
@@ -53,11 +53,19 @@ docker build --no-cache -f Dockerfile -t $CUSTOM_TRAIN_IMAGE_URI ../python_packa
# Deploy the docker image to Cloud Container Registry
docker push ${CUSTOM_TRAIN_IMAGE_URI}
# worker pool spec
worker_pool_spec="\
replica-count=1,\
machine-type=n1-standard-8,\
accelerator-type=NVIDIA_TESLA_V100,\
accelerator-count=1,\
container-image-uri=${CUSTOM_TRAIN_IMAGE_URI}"
# Submit Custom Job to Vertex AI
gcloud beta ai custom-jobs create \
--display-name=${JOB_NAME} \
--region ${REGION} \
--worker-pool-spec=replica-count=1,machine-type='n1-standard-8',accelerator-type='NVIDIA_TESLA_V100',accelerator-count=1,container-image-uri=${CUSTOM_TRAIN_IMAGE_URI} \
--worker-pool-spec="${worker_pool_spec}" \
--args="--model-name","finetuned-bert-classifier","--job-dir",$JOB_DIR
echo "After the job is completed successfully, model files will be saved at $JOB_DIR/"
Binary file not shown.

After

Width:  |  Height:  |  Size: 45 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 37 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 248 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 38 KiB

@@ -2,10 +2,13 @@
FROM pytorch/torchserve:latest-cpu
# install dependencies
RUN python3 -m pip install --upgrade pip
RUN pip3 install transformers
USER model-server
# copy model artifacts, custom handler and other dependencies
COPY ./custom_text_handler.py /home/model-server/
COPY ./custom_handler.py /home/model-server/
COPY ./index_to_name.json /home/model-server/
COPY ./model/finetuned-bert-classifier/ /home/model-server/
@@ -21,7 +24,7 @@ EXPOSE 7080
EXPOSE 7081
# create model archive file packaging model artifacts and dependencies
RUN torch-model-archiver -f --model-name=finetuned-bert-classifier --version=1.0 --serialized-file=/home/model-server/pytorch_model.bin --handler=/home/model-server/custom_text_handler.py --extra-files "/home/model-server/config.json,/home/model-server/tokenizer.json,/home/model-server/training_args.bin,/home/model-server/tokenizer_config.json,/home/model-server/special_tokens_map.json,/home/model-server/vocab.txt,/home/model-server/index_to_name.json" --export-path=/home/model-server/model-store
RUN torch-model-archiver -f --model-name=finetuned-bert-classifier --version=1.0 --serialized-file=/home/model-server/pytorch_model.bin --handler=/home/model-server/custom_handler.py --extra-files "/home/model-server/config.json,/home/model-server/tokenizer.json,/home/model-server/training_args.bin,/home/model-server/tokenizer_config.json,/home/model-server/special_tokens_map.json,/home/model-server/vocab.txt,/home/model-server/index_to_name.json" --export-path=/home/model-server/model-store
# run Torchserve HTTP serve to respond to prediction requests
CMD ["torchserve", "--start", "--ts-config=/home/model-server/config.properties", "--models", "finetuned-bert-classifier=finetuned-bert-classifier.mar", "--model-store", "/home/model-server/model-store"]
@@ -0,0 +1,37 @@
FROM pytorch/torchserve:latest-cpu
USER root
# run and update some basic packages software packages, including security libs
RUN apt-get update && apt-get install -y software-properties-common && add-apt-repository -y ppa:ubuntu-toolchain-r/test && apt-get update && apt-get install -y gcc-9 g++-9 apt-transport-https ca-certificates gnupg curl
# Install gcloud tools for gsutil as well as debugging
RUN echo "deb [signed-by=/usr/share/keyrings/cloud.google.gpg] http://packages.cloud.google.com/apt cloud-sdk main" | tee -a /etc/apt/sources.list.d/google-cloud-sdk.list && curl https://packages.cloud.google.com/apt/doc/apt-key.gpg | apt-key --keyring /usr/share/keyrings/cloud.google.gpg add - && apt-get update -y && apt-get install google-cloud-sdk -y
USER model-server
# install dependencies
RUN python3 -m pip install --upgrade pip
RUN pip3 install transformers
ARG MODEL_NAME=finetuned-bert-classifier
ENV MODEL_NAME="${MODEL_NAME}"
# health and prediction listener ports
ARG AIP_HTTP_PORT=7080
ENV AIP_HTTP_PORT="${AIP_HTTP_PORT}"
ARG MODEL_MGMT_PORT=7081
# expose health and prediction listener ports from the image
EXPOSE "${AIP_HTTP_PORT}"
EXPOSE "${MODEL_MGMT_PORT}"
EXPOSE 8080 8081 8082 7070 7071
# create torchserve configuration file
USER root
RUN echo "service_envelope=json\n" "inference_address=http://0.0.0.0:${AIP_HTTP_PORT}\n" "management_address=http://0.0.0.0:${MODEL_MGMT_PORT}" >> /home/model-server/config.properties
USER model-server
# run Torchserve HTTP serve to respond to prediction requests
CMD ["echo", "AIP_STORAGE_URI=${AIP_STORAGE_URI}", ";", "gsutil", "cp", "-r", "${AIP_STORAGE_URI}/${MODEL_NAME}.mar", "/home/model-server/model-store/", ";", "ls", "-ltr", "/home/model-server/model-store/", ";", "torchserve", "--start", "--ts-config=/home/model-server/config.properties", "--models", "${MODEL_NAME}=${MODEL_NAME}.mar", "--model-store", "/home/model-server/model-store"]
@@ -52,7 +52,8 @@ class TransformersClassifierHandler(BaseHandler):
with open(mapping_file_path) as f:
self.mapping = json.load(f)
else:
logger.warning('Missing the index_to_name.json file. Inference output will not include class name.')
logger.warning('Missing the index_to_name.json file. Inference output will default.')
self.mapping = {"0": "Negative", "1": "Positive"}
self.initialized = True
@@ -88,4 +89,3 @@ class TransformersClassifierHandler(BaseHandler):
def postprocess(self, inference_output):
return inference_output
@@ -19,13 +19,19 @@ echo "Submitting Custom Job to Vertex AI to train PyTorch model"
# BUCKET_NAME: Change to your bucket name
BUCKET_NAME="[your-bucket-name]" # <-- CHANGE TO YOUR BUCKET NAME
BUCKET_NAME="cloud-ai-platform-2f444b6a-a742-444b-b91a-c7519f51bd77"
# validate bucket name
if [ "${BUCKET_NAME}" = "[your-bucket-name]" ]
then
echo "[ERROR] INVALID VALUE: Please update the variable BUCKET_NAME with valid Cloud Storage bucket name. Exiting the script..."
exit 1
fi
# The PyTorch image provided by Vertex AI Training.
IMAGE_URI="us-docker.pkg.dev/vertex-ai/training/pytorch-gpu.1-7:latest"
# JOB_NAME: the name of your job running on Vertex AI.
JOB_PREFIX="finetuned-bert-classifier-pytorch-pkg-ar-"
JOB_PREFIX="finetuned-bert-classifier-pytorch-pkg-ar"
JOB_NAME=${JOB_PREFIX}-$(date +%Y%m%d%H%M%S)-custom-job
# REGION: select a region from https://cloud.google.com/vertex-ai/docs/general/locations#available_regions
@@ -35,19 +41,21 @@ REGION="us-central1"
# JOB_DIR: Where to store prepared package and upload output model.
JOB_DIR=gs://${BUCKET_NAME}/${JOB_PREFIX}/model/${JOB_NAME}
# validate bucket name
if [ "${BUCKET_NAME}" = "[your-bucket-name]" ]
then
echo "[ERROR] INVALID VALUE: Please update the variable BUCKET_NAME with valid Cloud Storage bucket name. Exiting the script..."
exit 1
fi
# worker pool spec
worker_pool_spec="\
replica-count=1,\
machine-type=n1-standard-8,\
accelerator-type=NVIDIA_TESLA_V100,\
accelerator-count=1,\
executor-image-uri=${IMAGE_URI},\
python-module=trainer.task,\
local-package-path=../python_package/"
# Submit Custom Job to Vertex AI
gcloud beta ai custom-jobs create \
--display-name=${JOB_NAME} \
--region ${REGION} \
--python-package-uris=${PACKAGE_PATH} \
--worker-pool-spec=replica-count=1,machine-type='n1-standard-8',accelerator-type='NVIDIA_TESLA_V100',accelerator-count=1,executor-image-uri=${IMAGE_URI},python-module='trainer.task',local-package-path="../python_package/" \
--worker-pool-spec="${worker_pool_spec}" \
--args="--model-name","finetuned-bert-classifier","--job-dir",$JOB_DIR
echo "After the job is completed successfully, model files will be saved at $JOB_DIR/"
@@ -122,6 +122,9 @@ def run(args):
# Train / Test the model
trainer = train(args, text_classifier, train_dataset, test_dataset)
metrics = trainer.evaluate(eval_dataset=test_dataset)
trainer.save_metrics("all", metrics)
# Export the trained model
trainer.save_model(os.path.join("/tmp", args.model_name))
@@ -63,20 +63,20 @@
"- [Training](#Training)\n",
" - [Run Training Locally in the Notebook](#Training-locally-in-the-notebook)\n",
" - [Run Training Job on Vertex AI](#Training-on-Vertex-AI)\n",
" - [Training with pre-built container](#Run-Custom-Job-on-Vertex-Training-with-a-pre-built-container)\n",
" - [Training with custom container](#Run-Custom-Job-on-Vertex-Training-with-custom-container)\n",
" - [Training with pre-built container](#Run-Custom-Job-on-Vertex-AI-Training-with-a-pre-built-container)\n",
" - [Training with custom container](#Run-Custom-Job-on-Vertex-AI-Training-with-custom-container)\n",
"- [Tuning](#Hyperparameter-Tuning) \n",
" - [Run Hyperparameter Tuning job on Vertex AI](#Run-Hyperparameter-Tuning-Job-on-Vertex-AI)\n",
"- [Deploying](#Deploying)\n",
" - [Deploying model on Vertex Predictions with custom container](#Deploying-model-on-Vertex-Predictions-with-custom-container)\n",
" - [Deploying model on Vertex AI Predictions with custom container](#Deploying-model-on-Vertex AI-Predictions-with-custom-container)\n",
"\n",
"### Costs \n",
"\n",
"This tutorial uses billable components of Google Cloud Platform (GCP):\n",
"\n",
"* [Notebooks](https://cloud.google.com/notebooks)\n",
"* [Vertex Training](https://cloud.google.com/vertex-ai/docs/training/custom-training)\n",
"* [Vertex Predictions](https://cloud.google.com/vertex-ai/docs/predictions/getting-predictions)\n",
"* [Vertex AI Workbench](https://cloud.google.com/vertex-ai-workbench)\n",
"* [Vertex AI Training](https://cloud.google.com/vertex-ai/docs/training/custom-training)\n",
"* [Vertex AI Predictions](https://cloud.google.com/vertex-ai/docs/predictions/getting-predictions)\n",
"* [Cloud Storage](https://cloud.google.com/storage)\n",
"* [Container Registry](https://cloud.google.com/container-registry)\n",
"* [Cloud Build](https://cloud.google.com/build) *[Optional]*\n",
@@ -202,9 +202,9 @@
"id": "e0c1dcadc2c8"
},
"source": [
"We will be using [Vertex SDK for Python](https://cloud.google.com/vertex-ai/docs/start/client-libraries#python) to interact with Vertex AI services. The high-level `aiplatform` library is designed to simplify common data science workflows by using wrapper classes and opinionated defaults. \n",
"We will be using [Vertex AI SDK for Python](https://cloud.google.com/vertex-ai/docs/start/client-libraries#python) to interact with Vertex AI services. The high-level `aiplatform` library is designed to simplify common data science workflows by using wrapper classes and opinionated defaults. \n",
"\n",
"#### Install Vertex SDK for Python"
"#### Install Vertex AI SDK for Python"
]
},
{
@@ -1199,7 +1199,7 @@
"source": [
"### Run predictions locally with sample examples\n",
"\n",
"Using the trained model, we can predict the sentiment label for an input text after applying the preprocessing function that was used during the training. We will run the predictions locally in the notebook and later show how you can deploy the model to an endpoint using [TorchServe](https://pytorch.org/serve/) on Vertex Predictions."
"Using the trained model, we can predict the sentiment label for an input text after applying the preprocessing function that was used during the training. We will run the predictions locally in the notebook and later show how you can deploy the model to an endpoint using [TorchServe](https://pytorch.org/serve/) on Vertex AI Predictions."
]
},
{
@@ -1382,7 +1382,7 @@
"id": "f7466d414a0e"
},
"source": [
"### Run Custom Job on Vertex Training with a pre-built container"
"### Run Custom Job on Vertex AI Training with a pre-built container"
]
},
{
@@ -1395,7 +1395,7 @@
"\n",
"In this notebook, we are using Hugging Face Datasets and fine tuning a transformer model from Hugging Face Transformers Library for sentiment analysis task using PyTorch. We will use [pre-built container for PyTorch](https://cloud.google.com/vertex-ai/docs/training/pre-built-containers#pytorch) and package the training application code by adding standard Python dependencies - `transformers`, `datasets` and `tqdm` - in the `setup.py` file. \n",
"\n",
"![Training with Prebuilt Containers on Vertex Training](./images/training-with-prebuilt-containers-on-vertex-training.png)"
"![Training with Prebuilt Containers on Vertex AI Training](./images/training-with-prebuilt-containers-on-vertex-training.png)"
]
},
{
@@ -1569,7 +1569,7 @@
"source": [
"#### **Run custom training job on Vertex AI**\n",
"\n",
"We use [Vertex SDK for Python](https://cloud.google.com/vertex-ai/docs/start/client-libraries#client_libraries) to create and submit training job to the Vertex training service."
"We use [Vertex AI SDK for Python](https://cloud.google.com/vertex-ai/docs/start/client-libraries#client_libraries) to create and submit training job to the Vertex AI training service."
]
},
{
@@ -1578,7 +1578,7 @@
"id": "5d2957ef04fd"
},
"source": [
"##### **Initialize the Vertex SDK for Python**"
"##### **Initialize the Vertex AI SDK for Python**"
]
},
{
@@ -1598,7 +1598,7 @@
"id": "6b0fed34b728"
},
"source": [
"##### **Configure and submit Custom Job to Vertex Training service**"
"##### **Configure and submit Custom Job to Vertex AI Training service**"
]
},
{
@@ -1609,7 +1609,7 @@
"source": [
"Configure a [Custom Job](https://cloud.google.com/vertex-ai/docs/training/create-custom-job) with the [pre-built container](https://cloud.google.com/vertex-ai/docs/training/pre-built-containers) image for PyTorch and training code packaged as Python source distribution. \n",
"\n",
"**NOTE:** When using Vertex SDK for Python for submitting a training job, it creates a [Training Pipeline](https://cloud.google.com/vertex-ai/docs/training/create-training-pipeline) which launches the Custom Job on Vertex Training service."
"**NOTE:** When using Vertex AI SDK for Python for submitting a training job, it creates a [Training Pipeline](https://cloud.google.com/vertex-ai/docs/training/create-training-pipeline) which launches the Custom Job on Vertex AI Training service."
]
},
{
@@ -1686,7 +1686,7 @@
"\n",
"You can monitor the custom job launched from Cloud Console following the link [here](https://console.cloud.google.com/vertex-ai/training/training-pipelines/) or use gcloud CLI command [`gcloud beta ai custom-jobs stream-logs`](https://cloud.google.com/sdk/gcloud/reference/beta/ai/custom-jobs/stream-logs)\n",
"\n",
"![Monitor custom job progress in Vertex Training](./images/vertex-training-monitor-custom-job.png)"
"![Monitor custom job progress in Vertex AI Training](./images/vertex-training-monitor-custom-job.png)"
]
},
{
@@ -1798,7 +1798,7 @@
"id": "c170d386492b"
},
"source": [
"### Run Custom Job on Vertex Training with custom container"
"### Run Custom Job on Vertex AI Training with custom container"
]
},
{
@@ -1807,7 +1807,7 @@
"id": "035227b6e581"
},
"source": [
"To create a [training job with custom container](https://cloud.google.com/vertex-ai/docs/training/create-custom-container?hl=hr), you define a `Dockerfile` to install or add the dependencies required for the training job. Then, you build and test your Docker image locally to verify, push the image to Container Registry and submit a Custom Job to Vertex Training service.\n",
"To create a [training job with custom container](https://cloud.google.com/vertex-ai/docs/training/create-custom-container?hl=hr), you define a `Dockerfile` to install or add the dependencies required for the training job. Then, you build and test your Docker image locally to verify, push the image to Container Registry and submit a Custom Job to Vertex AI Training service.\n",
"\n",
"![Training with custom containers on Vertex AI](./images/training-with-custom-containers-on-vertex-training.png)"
]
@@ -1834,7 +1834,7 @@
"%%writefile ./custom_container/Dockerfile\n",
"\n",
"# Use pytorch GPU base image\n",
"FROM gcr.io/cloud-aiplatform/training/pytorch-gpu.1-7\n",
"FROM us-docker.pkg.dev/vertex-ai/training/pytorch-gpu.1-10:latest\n",
"\n",
"# set working directory\n",
"WORKDIR /app\n",
@@ -1968,7 +1968,7 @@
"id": "a23e5e34bea9"
},
"source": [
"##### **Initialize the Vertex SDK for Python**"
"##### **Initialize the Vertex AI SDK for Python**"
]
},
{
@@ -1988,11 +1988,11 @@
"id": "abf1fa4085cb"
},
"source": [
"##### **Configure and submit Custom Job to Vertex Training service**\n",
"##### **Configure and submit Custom Job to Vertex AI Training service**\n",
"\n",
"Configure a [Custom Job](https://cloud.google.com/vertex-ai/docs/training/create-custom-job) with the [custom container](https://cloud.google.com/vertex-ai/docs/training/create-custom-container) image with training code and other dependencies\n",
"\n",
"**NOTE:** When using Vertex SDK for Python for submitting a training job, it creates a [Training Pipeline](https://cloud.google.com/vertex-ai/docs/training/create-training-pipeline) which launches the Custom Job to train on Vertex Training."
"**NOTE:** When using Vertex AI SDK for Python for submitting a training job, it creates a [Training Pipeline](https://cloud.google.com/vertex-ai/docs/training/create-training-pipeline) which launches the Custom Job to train on Vertex AI Training."
]
},
{
@@ -2044,7 +2044,7 @@
},
"outputs": [],
"source": [
"# submit the custom job to Vertex training service\n",
"# submit the custom job to Vertex AI training service\n",
"model = job.run(\n",
" replica_count=1,\n",
" machine_type=\"n1-standard-8\",\n",
@@ -2065,7 +2065,7 @@
"\n",
"You can monitor the custom job launched from Cloud Console following the link [here](https://console.cloud.google.com/vertex-ai/training/training-pipelines/) or use gcloud CLI command [`gcloud beta ai custom-jobs stream-logs`](https://cloud.google.com/sdk/gcloud/reference/beta/ai/custom-jobs/stream-logs)\n",
"\n",
"![Monitor custom job progress in Vertex Training](./images/vertex-training-monitor-custom-job-container.png)"
"![Monitor custom job progress in Vertex AI Training](./images/vertex-training-monitor-custom-job-container.png)"
]
},
{
@@ -2148,11 +2148,11 @@
"id": "ba6122f929e3"
},
"source": [
"The training application code for fine-tuning a transformer model for sentiment analysis task uses hyperparameters such as learning rate and weight decay. These hyperparameters control the behavior of the training algorithm and can have a significant effect on the performance of the resulting model. This part of the notebook show how you can automate tuning these hyperparameters with Vertex Training service.\n",
"The training application code for fine-tuning a transformer model for sentiment analysis task uses hyperparameters such as learning rate and weight decay. These hyperparameters control the behavior of the training algorithm and can have a significant effect on the performance of the resulting model. This part of the notebook show how you can automate tuning these hyperparameters with Vertex AI Training service.\n",
"\n",
"We submit a [Hyperparameter Tuning job](https://cloud.google.com/vertex-ai/docs/training/hyperparameter-tuning-overview) to Vertex Training service by packaging the training application code and dependencies in a Docker container and push the container to Google Container Registry, similar to running a Custom Job on Vertex AI with Custom Container.\n",
"We submit a [Hyperparameter Tuning job](https://cloud.google.com/vertex-ai/docs/training/hyperparameter-tuning-overview) to Vertex AI Training service by packaging the training application code and dependencies in a Docker container and push the container to Google Container Registry, similar to running a Custom Job on Vertex AI with Custom Container.\n",
"\n",
"![Hyperparameter Tuning with Custom Containers on Vertex Training](./images/hp-tuning-with-custom-containers-on-vertex-training.png)"
"![Hyperparameter Tuning with Custom Containers on Vertex AI Training](./images/hp-tuning-with-custom-containers-on-vertex-training.png)"
]
},
{
@@ -2163,7 +2163,7 @@
"source": [
"### How hyperparameter tuning works in Vertex AI?\n",
"\n",
"Following are the high level steps involved in running a Hyperparameter Tuning job on Vertex Training service:\n",
"Following are the high level steps involved in running a Hyperparameter Tuning job on Vertex AI Training service:\n",
"\n",
"- You define the hyperparameters to tune the model along with the metric (or goal) to optimize\n",
"- Vertex AI runs multiple trials of your training application with the hyperparameters and limits you specified - maximum number of trials to run and number of parallel trials. \n",
@@ -2297,7 +2297,7 @@
"source": [
"### Run Hyperparameter Tuning Job on Vertex AI\n",
"\n",
"Before submitting the hyperparameter tuning job to Vertex AI, push the custom container image with training application to Google Cloud Container Registry and then submit the job to Vertex AI. We will be using the same image used for running Custom Job on Vertex Training service."
"Before submitting the hyperparameter tuning job to Vertex AI, push the custom container image with training application to Google Cloud Container Registry and then submit the job to Vertex AI. We will be using the same image used for running Custom Job on Vertex AI Training service."
]
},
{
@@ -2326,7 +2326,7 @@
"id": "f60fab07d67c"
},
"source": [
"##### **Initialize the Vertex SDK for Python**"
"##### **Initialize the Vertex AI SDK for Python**"
]
},
{
@@ -2346,7 +2346,7 @@
"id": "6652aa63ddff"
},
"source": [
"##### **Configure and submit Hyperparameter Tuning Job to Vertex Training service**\n",
"##### **Configure and submit Hyperparameter Tuning Job to Vertex AI Training service**\n",
"\n",
"Configure a [Hyperparameter Tuning Job](https://cloud.google.com/vertex-ai/docs/training/using-hyperparameter-tuning) with the [custom container](https://cloud.google.com/vertex-ai/docs/training/create-custom-container) image with training code and other dependencies.\n",
"\n",
@@ -2374,7 +2374,7 @@
"id": "9d46db3a8b23"
},
"source": [
"Define the training arguments with `hp-tune` argument set to `y` so that training application code can report metrics to Vertex"
"Define the training arguments with `hp-tune` argument set to `y` so that training application code can report metrics to Vertex AI"
]
},
{
@@ -2548,7 +2548,7 @@
"\n",
"You can monitor the hyperparameter tuning job launched from Cloud Console following the link [here](https://console.cloud.google.com/vertex-ai/training/hyperparameter-tuning-jobs/) or use gcloud CLI command [`gcloud beta ai custom-jobs stream-logs`](https://cloud.google.com/sdk/gcloud/reference/beta/ai/custom-jobs/stream-logs)\n",
"\n",
"![Monitor hyperparameter tuning job progress in Vertex Training](./images/vertex-training-monitor-hptuning-job-container.png)"
"![Monitor hyperparameter tuning job progress in Vertex AI Training](./images/vertex-training-monitor-hptuning-job-container.png)"
]
},
{
@@ -2557,7 +2557,7 @@
"id": "ba934b434f03"
},
"source": [
"After the job is finished, you can view and format the results of the hyperparameter tuning Trials (run by Vertex Training service) as a Pandas dataframe"
"After the job is finished, you can view and format the results of the hyperparameter tuning Trials (run by Vertex AI Training service) as a Pandas dataframe"
]
},
{
@@ -2612,7 +2612,7 @@
"id": "5dbccb2b7d32"
},
"source": [
"Now from the results of Trials, you can pick the best performing Trial to deploy to Vertex Predictions"
"Now from the results of Trials, you can pick the best performing Trial to deploy to Vertex AI Predictions"
]
},
{
@@ -2701,8 +2701,8 @@
"JOB_NAME=${JOB_PREFIX}-pytorch-hptune-$(date +%Y%m%d%H%M%S)\n",
"echo \"Launching hyperparameter tuning job with display name as \"$JOB_NAME\n",
"\n",
"# BUCKET_NAME: Change to your bucket name\n",
"BUCKET_NAME=$1 # <-- CHANGE TO YOUR BUCKET NAME\n",
"# BUCKET_NAME is a required parameter to run the cell.\n",
"BUCKET_NAME=$1\n",
"\n",
"# APP_NAME: get application name\n",
"APP_NAME=$2\n",
@@ -2711,7 +2711,7 @@
"JOB_DIR=${BUCKET_NAME}/${JOB_PREFIX}/model/${JOB_NAME}\n",
"\n",
"# custom container image URI\n",
"CUSTOM_TRAIN_IMAGE_URI=f'gcr.io/'${PROJECT_ID}'/pytorch_gpu_train_'${APP_NAME}\n",
"CUSTOM_TRAIN_IMAGE_URI='gcr.io/'${PROJECT_ID}'/pytorch_gpu_train_'${APP_NAME}\n",
"\n",
"# ========================================================\n",
"# create hyperparameter tuning configuration file\n",
@@ -2772,20 +2772,20 @@
"source": [
"## Deploying\n",
"\n",
"Deploying a PyTorch model on [Vertex Predictions](https://cloud.google.com/vertex-ai/docs/predictions/getting-predictions) requires to use a custom container that serves online predictions. You will deploy a container running [PyTorch's TorchServe](https://pytorch.org/serve/) tool in order to serve predictions from a fine-tuned transformer model from Hugging Face Transformers for sentiment analysis task. You can then use Vertex Predictions to classify sentiment of input texts. \n",
"Deploying a PyTorch model on [Vertex AI Predictions](https://cloud.google.com/vertex-ai/docs/predictions/getting-predictions) requires to use a custom container that serves online predictions. You will deploy a container running [PyTorch's TorchServe](https://pytorch.org/serve/) tool in order to serve predictions from a fine-tuned transformer model from Hugging Face Transformers for sentiment analysis task. You can then use Vertex AI Predictions to classify sentiment of input texts. \n",
"\n",
"### Deploying model on Vertex Predictions with custom container\n",
"### Deploying model on Vertex AI Predictions with custom container\n",
"\n",
"To use a custom container to serve predictions from a PyTorch model, you must provide Vertex AI with a Docker container image that runs an HTTP server, such as TorchServe in this case. Please refer to [documentation](https://cloud.google.com/vertex-ai/docs/predictions/custom-container-requirements) that describes the container image requirements to be compatible with Vertex Predictions.\n",
"To use a custom container to serve predictions from a PyTorch model, you must provide Vertex AI with a Docker container image that runs an HTTP server, such as TorchServe in this case. Please refer to [documentation](https://cloud.google.com/vertex-ai/docs/predictions/custom-container-requirements) that describes the container image requirements to be compatible with Vertex AI Predictions.\n",
"\n",
"![Serving with Custom Containers on Vertex Predictions](./images/serve-pytorch-model-on-vertex-predictions-with-custom-containers.png)\n",
"![Serving with Custom Containers on Vertex AI Predictions](./images/serve-pytorch-model-on-vertex-predictions-with-custom-containers.png)\n",
"\n",
"Essentially, to deploy a PyTorch model on Vertex Predictions following are the steps:\n",
"Essentially, to deploy a PyTorch model on Vertex AI Predictions following are the steps:\n",
"\n",
"1. Package the trained model artifacts including [default](https://pytorch.org/serve/#default-handlers) or [custom](https://pytorch.org/serve/custom_service.html) handlers by creating an archive file using [Torch model archiver](https://github.com/pytorch/serve/tree/master/model-archiver)\n",
"2. Build a [custom container](https://cloud.google.com/vertex-ai/docs/predictions/custom-container-requirements) compatible with Vertex Predictions to serve the model using Torchserve\n",
"3. Upload the model with custom container image to serve predictions as a Vertex Model resource\n",
"4. Create a Vertex Endpoint and [deploy the model](https://cloud.google.com/vertex-ai/docs/predictions/deploy-model-api) resource"
"2. Build a [custom container](https://cloud.google.com/vertex-ai/docs/predictions/custom-container-requirements) compatible with Vertex AI Predictions to serve the model using Torchserve\n",
"3. Upload the model with custom container image to serve predictions as a Vertex AI Model resource\n",
"4. Create a Vertex AI Endpoint and [deploy the model](https://cloud.google.com/vertex-ai/docs/predictions/deploy-model-api) resource"
]
},
{
@@ -2815,7 +2815,7 @@
},
"outputs": [],
"source": [
"%%writefile predictor/custom_text_handler.py\n",
"%%writefile predictor/custom_handler.py\n",
"\n",
"import os\n",
"import json\n",
@@ -2870,7 +2870,8 @@
" with open(mapping_file_path) as f:\n",
" self.mapping = json.load(f)\n",
" else:\n",
" logger.warning('Missing the index_to_name.json file. Inference output will not include class name.')\n",
" logger.warning('Missing the index_to_name.json file. Inference output will default.')\n",
" self.mapping = {\"0\": \"Negative\", \"1\": \"Positive\"}\n",
"\n",
" self.initialized = True\n",
"\n",
@@ -3047,10 +3048,13 @@
"FROM pytorch/torchserve:latest-cpu\n",
"\n",
"# install dependencies\n",
"RUN python3 -m pip install --upgrade pip\n",
"RUN pip3 install transformers\n",
"\n",
"USER model-server\n",
"\n",
"# copy model artifacts, custom handler and other dependencies\n",
"COPY ./custom_text_handler.py /home/model-server/\n",
"COPY ./custom_handler.py /home/model-server/\n",
"COPY ./index_to_name.json /home/model-server/\n",
"COPY ./model/$APP_NAME/ /home/model-server/\n",
"\n",
@@ -3070,7 +3074,7 @@
" --model-name=$APP_NAME \\\n",
" --version=1.0 \\\n",
" --serialized-file=/home/model-server/pytorch_model.bin \\\n",
" --handler=/home/model-server/custom_text_handler.py \\\n",
" --handler=/home/model-server/custom_handler.py \\\n",
" --extra-files \"/home/model-server/config.json,/home/model-server/tokenizer.json,/home/model-server/training_args.bin,/home/model-server/tokenizer_config.json,/home/model-server/special_tokens_map.json,/home/model-server/vocab.txt,/home/model-server/index_to_name.json\" \\\n",
" --export-path=/home/model-server/model-store\n",
"\n",
@@ -3129,7 +3133,7 @@
"source": [
"#### **Run the container locally** ***[Optional]***\n",
"\n",
"Before push the container image to Container Registry to use it with Vertex Predictions, you can run it as a container in your local environment to verify that the server works as expected"
"Before push the container image to Container Registry to use it with Vertex AI Predictions, you can run it as a container in your local environment to verify that the server works as expected"
]
},
{
@@ -3267,9 +3271,9 @@
"id": "69477b3a00c0"
},
"source": [
"#### **Deploying the serving container to Vertex Predictions**\n",
"#### **Deploying the serving container to Vertex AI Predictions**\n",
"\n",
"We create a model resource on Vertex AI and deploy the model to a Vertex Endpoints. You must deploy a model to an endpoint before using the model. The deployed model runs the custom container image to serve predictions. "
"We create a model resource on Vertex AI and deploy the model to a Vertex AI Endpoints. You must deploy a model to an endpoint before using the model. The deployed model runs the custom container image to serve predictions. "
]
},
{
@@ -3300,7 +3304,7 @@
"id": "a3da91e19af4"
},
"source": [
"##### **Initialize the Vertex SDK for Python**"
"##### **Initialize the Vertex AI SDK for Python**"
]
},
{
@@ -3437,7 +3441,7 @@
"id": "bc4673478269"
},
"source": [
"#### **Invoking the Endpoint with deployed Model using Vertex SDK to make predictions**"
"#### **Invoking the Endpoint with deployed Model using Vertex AI SDK to make predictions**"
]
},
{
@@ -3487,7 +3491,7 @@
"source": [
"##### **Formatting input for online prediction**\n",
"\n",
"For online prediction requests, the prediction input instances must be formatted as JSON with base64 encoding as shown here:\n",
"This notebook uses [Torchserve's KServe based inference API](https://pytorch.org/serve/inference_api.html#kserve-inference-api) which is also [Vertex AI Predictions compatible format](https://cloud.google.com/vertex-ai/docs/predictions/custom-container-requirements#prediction). For online prediction requests, format the prediction input instances as JSON with base64 encoding as shown here:\n",
"\n",
"```\n",
"[\n",
@@ -3560,9 +3564,9 @@
},
"source": [
"##### ***[Optional]*** **Make prediction requests using gcloud CLI**\n",
"You can also call the Vertex Endpoint to make predictions using [`gcloud beta ai endpoints predict`](https://cloud.google.com/sdk/gcloud/reference/beta/ai/endpoints/predict). \n",
"You can also call the Vertex AI Endpoint to make predictions using [`gcloud beta ai endpoints predict`](https://cloud.google.com/sdk/gcloud/reference/beta/ai/endpoints/predict). \n",
"\n",
"The following cell shows how to make a prediction request to Vertex Endpoints using `gcloud` CLI: "
"The following cell shows how to make a prediction request to Vertex AI Endpoints using `gcloud` CLI: "
]
},
{
@@ -3653,12 +3657,12 @@
},
"outputs": [],
"source": [
"delete_custom_job = True\n",
"delete_hp_tuning_job = True\n",
"delete_custom_job = False\n",
"delete_hp_tuning_job = False\n",
"delete_endpoint = True\n",
"delete_model = True\n",
"delete_bucket = True\n",
"delete_image = True"
"delete_model = False\n",
"delete_bucket = False\n",
"delete_image = False"
]
},
{
@@ -3686,7 +3690,7 @@
"\n",
"client_options = {\"api_endpoint\": API_ENDPOINT}\n",
"\n",
"# Initialize Vertex SDK\n",
"# Initialize Vertex AI SDK\n",
"aiplatform.init(project=PROJECT_ID, staging_bucket=BUCKET_NAME)"
]
},
@@ -3924,7 +3928,7 @@
"if delete_bucket and \"BUCKET_NAME\" in globals():\n",
" print(f\"Deleting all contents from the bucket {BUCKET_NAME}\")\n",
"\n",
" shell_output=! gsutil du -as $BUCKET_NAME\n",
" shell_output = ! gsutil du -as $BUCKET_NAME\n",
" print(\n",
" f\"Size of the bucket {BUCKET_NAME} before deleting = {shell_output[0].split()[0]} bytes\"\n",
" )\n",
@@ -3932,7 +3936,7 @@
" # uncomment below line to delete contents of the bucket\n",
" # ! gsutil rm -r $BUCKET_NAME\n",
"\n",
" shell_output=! gsutil du -as $BUCKET_NAME\n",
" shell_output = ! gsutil du -as $BUCKET_NAME\n",
" if float(shell_output[0].split()[0]) > 0:\n",
" print(\n",
" \"PLEASE UNCOMMENT LINE TO DELETE BUCKET. CONTENT FROM THE BUCKET NOT DELETED\"\n",
@@ -0,0 +1,28 @@
# Train and deploy a scikit-learn model with Vertex AI
This repository shows how to train and deploy a text classifier using scikit-learn and Vertex AI.
The main used Vertex AI features are:
- Vertex AI Custom Training
- Vertex AI Model
- Vertex AI Endpoint
Further used GCP services are:
- Google Cloud Logging
- Google Cloud Storage
## Repository
├── README.md
├── create_job.ipynb # <-- creates the training job and deploys the model
├── requirements.txt # <-- requirements for deploying the job
└── task.py # <-- contains the training application
## Training job overview
The training job performs the following steps:
1. Downloads the `NewsAggregator` dataset from the UCI Machine Learning Repository
2. Trains and evaluates a classifier using scikit-learn
3. Exports model and evaluation artifacts to GCS
4. Deploys the model as a `Vertex AI Endpoint`
@@ -0,0 +1,290 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {
"id": "72b875d67303"
},
"source": [
"# Create and run a custom Vertex AI Training Job from a local script"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "398b976f501e"
},
"source": [
"## Install Vertex AI Python Client"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "2162adc1d8fe"
},
"outputs": [],
"source": [
"!pip install -r requirements.txt --upgrade"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "d4bf4ab70e65"
},
"source": [
"## GCP authentication"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "41084de2e96a"
},
"outputs": [],
"source": [
"import os\n",
"\n",
"os.environ[\n",
" \"GOOGLE_APPLICATION_CREDENTIALS\"\n",
"] = \"\" # TODO: path to credentials .json file"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "1ce7efb99a95"
},
"source": [
"## Create the custom Vertex AI Training Job\n",
"\n",
"1. Define the custom job parameters\n",
"2. Submit the job to create a `Vertex AI Model`"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "c04b9efb5eb3"
},
"outputs": [],
"source": [
"# Import the Vertex AI SDK (Python Client)\n",
"from google.cloud import aiplatform"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "b56f5fdcefd6"
},
"outputs": [],
"source": [
"# Project meta data\n",
"PROJECT_ID = \"\" # TODO\n",
"REGION = \"\" # TODO e.g. europe\n",
"ZONE = \"\" # TODO e.g. west4\n",
"LOCATION = f\"{REGION}-{ZONE}\""
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "4af8cfb1e1a1"
},
"outputs": [],
"source": [
"aiplatform.init(project=PROJECT_ID, location=LOCATION)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "e23d0161b489"
},
"outputs": [],
"source": [
"# Variables for specifying the job\n",
"DISPLAY_NAME = (\n",
" \"news-classifier-training\" # TODO: How the job is displayed on Vertex AI GUI\n",
")\n",
"SCRIPT_PATH = \"./task.py\" # Path to local training script\n",
"STAGING_BUCKET = (\n",
" \"\" # TODO GCS URI where meta data and artifacts are stored for this job\n",
")\n",
"MODEL_TRAINING_IMAGE = f\"{REGION}-docker.pkg.dev/vertex-ai/training/scikit-learn-cpu.0-23:latest\" # Pre-built training image\n",
"REQUIREMENTS = [\"wget\"] # Additional requirements not already part of the base image\n",
"# !Required if the Training Pipeline produces a managed Vertex AI Model!\n",
"MODEL_SERVING_IMAGE = f\"{REGION}-docker.pkg.dev/vertex-ai/prediction/sklearn-cpu.0-23:latest\" # Pre-built serving image"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "61565ec3e6de"
},
"outputs": [],
"source": [
"# Job definition\n",
"custom_training_job = aiplatform.CustomTrainingJob(\n",
" project=PROJECT_ID,\n",
" location=LOCATION,\n",
" display_name=DISPLAY_NAME,\n",
" script_path=SCRIPT_PATH,\n",
" staging_bucket=STAGING_BUCKET,\n",
" container_uri=MODEL_TRAINING_IMAGE,\n",
" requirements=REQUIREMENTS,\n",
" model_serving_container_image_uri=MODEL_SERVING_IMAGE,\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "8c8d4bd78688"
},
"outputs": [],
"source": [
"# Variables for running the job\n",
"MACHINE_TYPE = \"n1-standard-4\" # Standard VM with 4 CPUs\n",
"# !Required if the Training Pipeline produces a managed Vertex AI Model!\n",
"MODEL_DISPLAY_NAME = (\n",
" \"news-classifier-model\" # TODO: Name for the resulting managed Vertex AI Model.\n",
")\n",
"# Note that a single job may produce multiple models (e.g. one per run).\n",
"# The url to download the training data from.\n",
"DATASET_URL = \"https://archive.ics.uci.edu/ml/machine-learning-databases/00359/NewsAggregatorDataset.zip\""
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "13d31a7d0bb5"
},
"outputs": [],
"source": [
"# Run the job\n",
"model = custom_training_job.run(\n",
" machine_type=MACHINE_TYPE,\n",
" model_display_name=MODEL_DISPLAY_NAME,\n",
" args=[f\"--dataset_url={DATASET_URL}\", f\"--project_id={PROJECT_ID}\"],\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "8026ac119722"
},
"outputs": [],
"source": [
"MODEL_RESOURCE_NAME = model.resource_name"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "2d985fca37f3"
},
"source": [
"## Deploy model to Vertex AI Endpoint\n",
"\n",
"1. Retrieve the registered `Vertex AI Model`\n",
"2. Deploy the model to a new `Vertex AI Endpoint`\n",
"3. Get some test predictions from the endpoint"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "58c08631f94c"
},
"outputs": [],
"source": [
"ENDPOINT_DISPLAY_NAME = \"news-classifier-endpoint\" # TODO\n",
"MACHINE_TYPE_SERVING = \"n1-standard-2\""
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "b54867240880"
},
"outputs": [],
"source": [
"endpoint = aiplatform.Endpoint.create(\n",
" display_name=ENDPOINT_DISPLAY_NAME,\n",
" location=LOCATION,\n",
" project=PROJECT_ID,\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "a76e7c721b88"
},
"outputs": [],
"source": [
"model = aiplatform.Model(model_name=MODEL_RESOURCE_NAME)"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "d06714302f81"
},
"outputs": [],
"source": [
"model.deploy(\n",
" endpoint=endpoint,\n",
" deployed_model_display_name=MODEL_DISPLAY_NAME,\n",
" machine_type=MACHINE_TYPE,\n",
" traffic_percentage=100,\n",
" min_replica_count=1,\n",
" max_replica_count=1,\n",
" accelerator_type=None,\n",
" accelerator_count=None,\n",
")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "5b7eca31d80e"
},
"outputs": [],
"source": [
"endpoint.predict(instances={\"instances\": [\"A news headline to be classified\"]})"
]
}
],
"metadata": {
"colab": {
"name": "create_job.ipynb",
"toc_visible": true
},
"kernelspec": {
"display_name": "Python 3",
"name": "python3"
}
},
"nbformat": 4,
"nbformat_minor": 0
}
@@ -0,0 +1,2 @@
google-cloud-aiplatform
ipykernel
@@ -0,0 +1,167 @@
import argparse
import logging
import os
import pickle
import zipfile
from typing import List, Tuple
import pandas as pd
import wget
from google.cloud import storage
from google.cloud.logging import Client as LogClient
from sklearn.feature_extraction.text import CountVectorizer, TfidfTransformer
from sklearn.model_selection import train_test_split
from sklearn.naive_bayes import MultinomialNB
from sklearn.pipeline import Pipeline
def download_dataset_from_url(url: str) -> pd.DataFrame:
"""Downloads and unzips the dataset from `url` and reads it with pandas.
Args:
url (str, optional): URL to the dataset.
"""
zip_filepath = wget.download(url, out=".")
with zipfile.ZipFile(zip_filepath, "r") as zf:
zf.extract(path=".", member="newsCorpora.csv")
COLUMN_NAMES = ["id", "title", "url", "publisher",
"category", "story", "hostname", "timestamp"]
return pd.read_csv(
"newsCorpora.csv", delimiter="\t", names=COLUMN_NAMES, index_col=0
)
def get_train_test_data(dataframe: pd.DataFrame, test_size: float = 0.2
) -> Tuple[List, List, List, List]:
"""Splits the news dataset into train and test features and labels.
Args:
news (pd.DataFrame): The dataset as pandas DataFrame.
test_size (float): The size in percent of the test data.
Returns:
Tuple[List, List, List, List]: Tuple with train and test data
"""
train, test = train_test_split(dataframe, test_size=test_size)
x_train, y_train = train["title"].values, train["category"].values
x_test, y_test = test["title"].values, test["category"].values
return x_train, y_train, x_test, y_test
def export_model_to_gcs(fitted_pipeline: Pipeline, gcs_uri: str) -> str:
"""Exports trained pipeline to GCS
Parameters:
fitted_pipeline (sklearn.pipelines.Pipeline): the Pipeline object
with data already fitted (trained pipeline object).
gcs_uri (str): GCS path to store the trained pipeline
i.e gs://example_bucket/training-job.
Returns:
export_path (str): Model GCS location
"""
artifact_filename = 'model.pkl'
# Save model artifact to local filesystem (doesn't persist)
local_path = artifact_filename
with open(local_path, 'wb') as model_file:
pickle.dump(fitted_pipeline, model_file)
# Upload model artifact to Cloud Storage
storage_path = os.path.join(gcs_uri, artifact_filename)
blob = storage.blob.Blob.from_string(storage_path, client=storage.Client())
blob.upload_from_filename(local_path)
def export_evaluation_report_to_gcs(report: str, gcs_uri: str) -> None:
"""
Exports training job report to GCS
Parameters:
report (str): Full report in text to sent to GCS
gcs_uri (str): GCS path to store the report
i.e gs://example_bucket/training-job
"""
artifact_filename = 'report.txt'
# Upload model artifact to Cloud Storage
storage_path = os.path.join(gcs_uri, artifact_filename)
blob = storage.blob.Blob.from_string(storage_path, client=storage.Client())
blob.upload_from_string(report)
def train_and_score(X_train: List, y_train: List, X_test: List, y_test: List
) -> Tuple[Pipeline, float]:
"""Trains and cross-validates a text classifier pipeline.
Args:
X_train (List): Train features as list of strings.
y_train (List): Train labels as list of strings.
X_test (List): Test labels as list of strings.
y_test (List): Test labels as list of strings.
Returns:
Tuple[Pipeline, float]: Fitted pipeline and mean accuracy.
"""
pipeline = Pipeline([
("vectorizer", CountVectorizer()),
("tfidf", TfidfTransformer()),
("naivebayes", MultinomialNB()),
])
pipeline.fit(X_train, y_train)
score = pipeline.score(X_test, y_test)
return pipeline, score
# Define all the command line arguments your model can accept for training
if __name__ == "__main__":
parser = argparse.ArgumentParser()
parser.add_argument(
"--dataset_url",
help="Download url for the training data.",
type=str
)
parser.add_argument(
"--project_id",
help="GCP project id for cloud logging.",
type=str
)
args = parser.parse_args()
arguments = args.__dict__
# set up the GCP logger
client = LogClient(project=arguments["project_id"])
client.setup_logging(log_level=logging.INFO)
logging.info("Starting custom training job.")
# download the data from url
logging.info("Downloading training data from: {}".format(arguments["dataset_url"]))
dataframe = download_dataset_from_url(arguments["dataset_url"])
train_test_data = get_train_test_data(dataframe)
# train and cross validate
logging.info("Training started ...")
model, score = train_and_score(*train_test_data)
logging.info(f"Training completed with model score: {score}")
# export model to gcs
_gcs_uri = os.environ["AIP_MODEL_DIR"]
logging.info("Exporting model artifacts ...")
export_model_to_gcs(model, _gcs_uri)
export_evaluation_report_to_gcs(str(score), _gcs_uri)
logging.info(f"Exported model artifacts to GCS bucket: {_gcs_uri}")
@@ -188,11 +188,14 @@
},
"outputs": [],
"source": [
"! pip3 install {USER_FLAG} google-cloud-aiplatform==1.0.1\n",
"! pip3 install {USER_FLAG} google-cloud-pipeline-components==0.1.3\n",
"! pip3 install {USER_FLAG} google-cloud-aiplatform\n",
"! pip3 install {USER_FLAG} google-cloud-pipeline-components\n",
"! pip3 install {USER_FLAG} --upgrade kfp\n",
"! pip3 install {USER_FLAG} numpy==1.20.3\n",
"! pip3 install {USER_FLAG} --upgrade tensorflow"
"! pip3 install {USER_FLAG} numpy\n",
"! pip3 install {USER_FLAG} --upgrade tensorflow\n",
"! pip3 install {USER_FLAG} --upgrade pillow\n",
"! pip3 install {USER_FLAG} --upgrade tf-agents\n",
"! pip3 install {USER_FLAG} --upgrade fastapi"
]
},
{
@@ -287,7 +290,7 @@
"\n",
"# Get your Google Cloud project ID from gcloud\n",
"if not os.getenv(\"IS_TESTING\"):\n",
" shell_output=!gcloud config list --format 'value(core.project)' 2>/dev/null\n",
" shell_output = !gcloud config list --format 'value(core.project)' 2>/dev/null\n",
" PROJECT_ID = shell_output[0]\n",
" print(\"Project ID: \", PROJECT_ID)"
]
@@ -518,6 +521,7 @@
"import os\n",
"import sys\n",
"\n",
"from google.cloud import aiplatform\n",
"from google_cloud_pipeline_components import aiplatform as gcc_aip\n",
"from kfp.v2 import compiler, dsl\n",
"from kfp.v2.google.client import AIPlatformClient"
@@ -561,13 +565,34 @@
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "H3530hdGGilo"
"id": "895ac243c125"
},
"outputs": [],
"source": [
"# Dataset parameters\n",
"RAW_DATA_PATH = \"gs://cloud-samples-data/vertex-ai/community-content/tf_agents_bandits_movie_recommendation_with_kfp_and_vertex_sdk/u.data\" # Location of the MovieLens 100K dataset's \"u.data\" file.\n",
"\n",
"RAW_DATA_PATH = \"gs://[your-bucket-name]/raw_data/u.data\" # @param {type:\"string\"}"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "62bfb9a820f6"
},
"outputs": [],
"source": [
"# Download the sample data into your RAW_DATA_PATH\n",
"! gsutil cp \"gs://cloud-samples-data/vertex-ai/community-content/tf_agents_bandits_movie_recommendation_with_kfp_and_vertex_sdk/u.data\" $RAW_DATA_PATH"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "H3530hdGGilo"
},
"outputs": [],
"source": [
"# Pipeline parameters\n",
"PIPELINE_NAME = \"movielens-pipeline\" # Pipeline display name.\n",
"ENABLE_CACHING = False # Whether to enable execution caching for the pipeline.\n",
@@ -635,7 +660,7 @@
"source": [
"#### Run unit tests on the Generator component\n",
"\n",
"Before running the command, fill in `RAW_DATA_PATH` in [`src/generator/test_generator_component.py`](src/generator/test_generator_component.py)."
"Before running the command, you should update the `RAW_DATA_PATH` in [`src/generator/test_generator_component.py`](src/generator/test_generator_component.py)."
]
},
{
@@ -713,12 +738,12 @@
"TRAINING_ARTIFACTS_DIR = (\n",
" f\"{BUCKET_NAME}/artifacts\" # Root directory for training artifacts.\n",
")\n",
"TRAINING_REPLICA_COUNT = \"1\" # Number of replica to run the custom training job.\n",
"TRAINING_REPLICA_COUNT = 1 # Number of replica to run the custom training job.\n",
"TRAINING_MACHINE_TYPE = (\n",
" \"n1-standard-4\" # Type of machine to run the custom training job.\n",
")\n",
"TRAINING_ACCELERATOR_TYPE = \"ACCELERATOR_TYPE_UNSPECIFIED\" # Type of accelerators to run the custom training job.\n",
"TRAINING_ACCELERATOR_COUNT = \"0\" # Number of accelerators for the custom training job."
"TRAINING_ACCELERATOR_COUNT = 0 # Number of accelerators for the custom training job."
]
},
{
@@ -769,8 +794,12 @@
"TRAINED_POLICY_DISPLAY_NAME = (\n",
" \"movielens-trained-policy\" # Display name of the uploaded and deployed policy.\n",
")\n",
"TRAFFIC_SPLIT = {\"0\": 100}\n",
"ENDPOINT_DISPLAY_NAME = \"movielens-endpoint\" # Display name of the prediction endpoint.\n",
"ENDPOINT_MACHINE_TYPE = \"n1-standard-4\" # Type of machine of the prediction endpoint."
"ENDPOINT_MACHINE_TYPE = \"n1-standard-4\" # Type of machine of the prediction endpoint.\n",
"ENDPOINT_REPLICA_COUNT = 1 # Number of replicas of the prediction endpoint.\n",
"ENDPOINT_ACCELERATOR_TYPE = \"ACCELERATOR_TYPE_UNSPECIFIED\" # Type of accelerators to run the custom training job.\n",
"ENDPOINT_ACCELERATOR_COUNT = 0 # Number of accelerators for the custom training job."
]
},
{
@@ -900,16 +929,17 @@
},
"outputs": [],
"source": [
"from google_cloud_pipeline_components.experimental.custom_job import utils\n",
"from kfp.components import load_component_from_url\n",
"\n",
"generate_op = load_component_from_url(\n",
" \"https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/68d6cf46ee22a9b9295d62ea71996150baf8db94/community-content/tf_agents_bandits_movie_recommendation_with_kfp_and_vertex_sdk/mlops_pipeline_tf_agents_bandits_movie_recommendation/src/generator/component.yaml\"\n",
" \"https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/62a2a7611499490b4b04d731d48a7ba87c2d636f/community-content/tf_agents_bandits_movie_recommendation_with_kfp_and_vertex_sdk/mlops_pipeline_tf_agents_bandits_movie_recommendation/src/generator/component.yaml\"\n",
")\n",
"ingest_op = load_component_from_url(\n",
" \"https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/68d6cf46ee22a9b9295d62ea71996150baf8db94/community-content/tf_agents_bandits_movie_recommendation_with_kfp_and_vertex_sdk/mlops_pipeline_tf_agents_bandits_movie_recommendation/src/ingester/component.yaml\"\n",
" \"https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/62a2a7611499490b4b04d731d48a7ba87c2d636f/community-content/tf_agents_bandits_movie_recommendation_with_kfp_and_vertex_sdk/mlops_pipeline_tf_agents_bandits_movie_recommendation/src/ingester/component.yaml\"\n",
")\n",
"train_op = load_component_from_url(\n",
" \"https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/68d6cf46ee22a9b9295d62ea71996150baf8db94/community-content/tf_agents_bandits_movie_recommendation_with_kfp_and_vertex_sdk/mlops_pipeline_tf_agents_bandits_movie_recommendation/src/trainer/component.yaml\"\n",
" \"https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/62a2a7611499490b4b04d731d48a7ba87c2d636f/community-content/tf_agents_bandits_movie_recommendation_with_kfp_and_vertex_sdk/mlops_pipeline_tf_agents_bandits_movie_recommendation/src/trainer/component.yaml\"\n",
")\n",
"\n",
"\n",
@@ -978,7 +1008,7 @@
" bigquery_location=bigquery_location,\n",
" bigquery_table_id=bigquery_table_id,\n",
" )\n",
"\n",
" \n",
" # Run the Ingester component.\n",
" ingest_task = ingest_op(\n",
" project_id=project_id,\n",
@@ -988,7 +1018,16 @@
" )\n",
"\n",
" # Run the Trainer component and submit custom job to Vertex AI.\n",
" train_task = train_op(\n",
" # Convert the train_op component into a Vertex AI Custom Job pre-built component\n",
" custom_job_training_op = utils.create_custom_training_job_op_from_component(\n",
" component_spec=train_op,\n",
" replica_count=TRAINING_REPLICA_COUNT,\n",
" machine_type=TRAINING_MACHINE_TYPE,\n",
" accelerator_type=TRAINING_ACCELERATOR_TYPE,\n",
" accelerator_count=TRAINING_ACCELERATOR_COUNT,\n",
" )\n",
"\n",
" train_task = custom_job_training_op(\n",
" training_artifacts_dir=training_artifacts_dir,\n",
" tfrecord_file=ingest_task.outputs[\"tfrecord_file\"],\n",
" num_epochs=num_epochs,\n",
@@ -996,28 +1035,10 @@
" num_actions=num_actions,\n",
" tikhonov_weight=tikhonov_weight,\n",
" agent_alpha=agent_alpha,\n",
" project=PROJECT_ID,\n",
" location=REGION,\n",
" )\n",
"\n",
" worker_pool_specs = [\n",
" {\n",
" \"containerSpec\": {\n",
" \"imageUri\": train_task.container.image,\n",
" },\n",
" \"replicaCount\": TRAINING_REPLICA_COUNT,\n",
" \"machineSpec\": {\n",
" \"machineType\": TRAINING_MACHINE_TYPE,\n",
" \"acceleratorType\": TRAINING_ACCELERATOR_TYPE,\n",
" \"acceleratorCount\": TRAINING_ACCELERATOR_COUNT,\n",
" },\n",
" },\n",
" ]\n",
" train_task.custom_job_spec = {\n",
" \"displayName\": train_task.name,\n",
" \"jobSpec\": {\n",
" \"workerPoolSpecs\": worker_pool_specs,\n",
" },\n",
" }\n",
"\n",
" # Run the Deployer components.\n",
" # Upload the trained policy as a model.\n",
" model_upload_op = gcc_aip.ModelUploadOp(\n",
@@ -1034,11 +1055,14 @@
" # Deploy the uploaded, trained policy to the created endpoint. (This operation\n",
" # has to occur after both model uploading and endpoint creation complete.)\n",
" gcc_aip.ModelDeployOp(\n",
" project=project_id,\n",
" endpoint=endpoint_create_op.outputs[\"endpoint\"],\n",
" model=model_upload_op.outputs[\"model\"],\n",
" deployed_model_display_name=TRAINED_POLICY_DISPLAY_NAME,\n",
" machine_type=ENDPOINT_MACHINE_TYPE,\n",
" traffic_split=TRAFFIC_SPLIT,\n",
" dedicated_resources_machine_type=ENDPOINT_MACHINE_TYPE,\n",
" dedicated_resources_accelerator_type=ENDPOINT_ACCELERATOR_TYPE,\n",
" dedicated_resources_accelerator_count=ENDPOINT_ACCELERATOR_COUNT,\n",
" dedicated_resources_min_replica_count=ENDPOINT_REPLICA_COUNT,\n",
" )"
]
},
@@ -1053,12 +1077,11 @@
"# Compile the authored pipeline.\n",
"compiler.Compiler().compile(pipeline_func=pipeline, package_path=PIPELINE_SPEC_PATH)\n",
"\n",
"# Createa Vertex AI client.\n",
"api_client = AIPlatformClient(project_id=PROJECT_ID, region=REGION)\n",
"\n",
"# Create a pipeline run job.\n",
"response = api_client.create_run_from_job_spec(\n",
" job_spec_path=PIPELINE_SPEC_PATH,\n",
"job = aiplatform.PipelineJob(\n",
" display_name=f\"{PIPELINE_NAME}-startup\",\n",
" template_path=PIPELINE_SPEC_PATH,\n",
" pipeline_root=PIPELINE_ROOT,\n",
" parameter_values={\n",
" # Pipeline configs\n",
" \"project_id\": PROJECT_ID,\n",
@@ -1070,7 +1093,9 @@
" \"bigquery_table_id\": BIGQUERY_TABLE_ID,\n",
" },\n",
" enable_caching=ENABLE_CACHING,\n",
")"
")\n",
"\n",
"job.run()"
]
},
{
@@ -1111,7 +1136,11 @@
"SIMULATOR_SCHEDULE = \"*/5 * * * *\" # Cloud Scheduler cron job schedule for the Simulator. Eg. \"*/5 * * * *\" means every 5 mins.\n",
"SIMULATOR_SCHEDULER_MESSAGE = (\n",
" \"simulator-message\" # Cloud Scheduler message for the Simulator.\n",
")"
")\n",
"# TF-Agents RL configs\n",
"BATCH_SIZE = 8\n",
"RANK_K = 20\n",
"NUM_ACTIONS = 20"
]
},
{
@@ -1221,7 +1250,7 @@
},
"outputs": [],
"source": [
"endpoints = ! gcloud beta ai endpoints list \\\n",
"endpoints = ! gcloud ai endpoints list \\\n",
" --region=$REGION \\\n",
" --filter=display_name=$ENDPOINT_DISPLAY_NAME\n",
"print(\"\\n\".join(endpoints), \"\\n\")\n",
@@ -1424,13 +1453,11 @@
},
"outputs": [],
"source": [
"from kfp.components import load_component_from_url\n",
"\n",
"ingest_op = load_component_from_url(\n",
" \"https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/68d6cf46ee22a9b9295d62ea71996150baf8db94/community-content/tf_agents_bandits_movie_recommendation_with_kfp_and_vertex_sdk/mlops_pipeline_tf_agents_bandits_movie_recommendation/src/ingester/component.yaml\"\n",
" \"https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/62a2a7611499490b4b04d731d48a7ba87c2d636f/community-content/tf_agents_bandits_movie_recommendation_with_kfp_and_vertex_sdk/mlops_pipeline_tf_agents_bandits_movie_recommendation/src/ingester/component.yaml\"\n",
")\n",
"train_op = load_component_from_url(\n",
" \"https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/68d6cf46ee22a9b9295d62ea71996150baf8db94/community-content/tf_agents_bandits_movie_recommendation_with_kfp_and_vertex_sdk/mlops_pipeline_tf_agents_bandits_movie_recommendation/src/trainer/component.yaml\"\n",
" \"https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/62a2a7611499490b4b04d731d48a7ba87c2d636f/community-content/tf_agents_bandits_movie_recommendation_with_kfp_and_vertex_sdk/mlops_pipeline_tf_agents_bandits_movie_recommendation/src/trainer/component.yaml\"\n",
")\n",
"\n",
"\n",
@@ -1481,7 +1508,16 @@
" )\n",
"\n",
" # Run the Trainer component and submit custom job to Vertex AI.\n",
" train_task = train_op(\n",
" # Convert the train_op component into a Vertex AI Custom Job pre-built component\n",
" custom_job_training_op = utils.create_custom_training_job_op_from_component(\n",
" component_spec=train_op,\n",
" replica_count=TRAINING_REPLICA_COUNT,\n",
" machine_type=TRAINING_MACHINE_TYPE,\n",
" accelerator_type=TRAINING_ACCELERATOR_TYPE,\n",
" accelerator_count=TRAINING_ACCELERATOR_COUNT,\n",
" )\n",
"\n",
" train_task = custom_job_training_op(\n",
" training_artifacts_dir=training_artifacts_dir,\n",
" tfrecord_file=ingest_task.outputs[\"tfrecord_file\"],\n",
" num_epochs=num_epochs,\n",
@@ -1489,28 +1525,10 @@
" num_actions=num_actions,\n",
" tikhonov_weight=tikhonov_weight,\n",
" agent_alpha=agent_alpha,\n",
" project=PROJECT_ID,\n",
" location=REGION,\n",
" )\n",
"\n",
" worker_pool_specs = [\n",
" {\n",
" \"containerSpec\": {\n",
" \"imageUri\": train_task.container.image,\n",
" },\n",
" \"replicaCount\": TRAINING_REPLICA_COUNT,\n",
" \"machineSpec\": {\n",
" \"machineType\": TRAINING_MACHINE_TYPE,\n",
" \"acceleratorType\": TRAINING_ACCELERATOR_TYPE,\n",
" \"acceleratorCount\": TRAINING_ACCELERATOR_COUNT,\n",
" },\n",
" },\n",
" ]\n",
" train_task.custom_job_spec = {\n",
" \"displayName\": train_task.name,\n",
" \"jobSpec\": {\n",
" \"workerPoolSpecs\": worker_pool_specs,\n",
" },\n",
" }\n",
"\n",
" # Run the Deployer components.\n",
" # Upload the trained policy as a model.\n",
" model_upload_op = gcc_aip.ModelUploadOp(\n",
@@ -1527,11 +1545,13 @@
" # Deploy the uploaded, trained policy to the created endpoint. (This operation\n",
" # has to occur after both model uploading and endpoint creation complete.)\n",
" gcc_aip.ModelDeployOp(\n",
" project=project_id,\n",
" endpoint=endpoint_create_op.outputs[\"endpoint\"],\n",
" model=model_upload_op.outputs[\"model\"],\n",
" deployed_model_display_name=TRAINED_POLICY_DISPLAY_NAME,\n",
" machine_type=ENDPOINT_MACHINE_TYPE,\n",
" dedicated_resources_machine_type=ENDPOINT_MACHINE_TYPE,\n",
" dedicated_resources_accelerator_type=ENDPOINT_ACCELERATOR_TYPE,\n",
" dedicated_resources_accelerator_count=ENDPOINT_ACCELERATOR_COUNT,\n",
" dedicated_resources_min_replica_count=ENDPOINT_REPLICA_COUNT,\n",
" )"
]
},
@@ -39,14 +39,15 @@ outputs:
- {name: bigquery_table_id, type: String}
implementation:
container:
image: tensorflow/tensorflow:2.5.0
image: python:3.7
command:
- sh
- -c
- (PIP_DISABLE_PIP_VERSION_CHECK=1 python3 -m pip install --quiet --no-warn-script-location
'google-cloud-bigquery==2.20.0' 'tensorflow==2.5.0' 'tf-agents==0.8.0' || PIP_DISABLE_PIP_VERSION_CHECK=1
python3 -m pip install --quiet --no-warn-script-location 'google-cloud-bigquery==2.20.0'
'tensorflow==2.5.0' 'tf-agents==0.8.0' --user) && "$0" "$@"
'google-cloud-bigquery==2.20.0' 'pillow' 'tensorflow==2.5.0' 'tf-agents==0.8.0'
|| PIP_DISABLE_PIP_VERSION_CHECK=1 python3 -m pip install --quiet --no-warn-script-location
'google-cloud-bigquery==2.20.0' 'pillow' 'tensorflow==2.5.0' 'tf-agents==0.8.0'
--user) && "$0" "$@"
- sh
- -ec
- |
@@ -296,7 +297,8 @@ implementation:
def _serialize_str(str_value: str) -> str:
if not isinstance(str_value, str):
raise TypeError('Value "{}" has type "{}" instead of str.'.format(str(str_value), str(type(str_value))))
raise TypeError('Value "{}" has type "{}" instead of str.'.format(
str(str_value), str(type(str_value))))
return str_value
import argparse
@@ -20,7 +20,7 @@ outputs:
- {name: tfrecord_file, type: String}
implementation:
container:
image: tensorflow/tensorflow:2.5.0
image: python:3.7
command:
- sh
- -c
@@ -187,7 +187,8 @@ implementation:
def _serialize_str(str_value: str) -> str:
if not isinstance(str_value, str):
raise TypeError('Value "{}" has type "{}" instead of str.'.format(str(str_value), str(type(str_value))))
raise TypeError('Value "{}" has type "{}" instead of str.'.format(
str(str_value), str(type(str_value))))
return str_value
import argparse
@@ -1,3 +1,4 @@
google-cloud-bigquery==2.20.0
tensorflow==2.5.0
tensorflow==2.5.3
pillow==9.0.1
tf-agents==0.8.0
@@ -1,2 +1,4 @@
google-cloud-pubsub==2.5.0
pillow==9.0.1
tf-agents==0.8.0
tensorflow==2.5.0
tensorflow==2.5.3
@@ -0,0 +1,5 @@
dataclasses==0.6
google-cloud-aiplatform==1.8.1
tensorflow==2.5.3
pillow==9.0.1
tf-agents==0.8.0
@@ -27,14 +27,14 @@ outputs:
- {name: training_artifacts_dir, type: String}
implementation:
container:
image: tensorflow/tensorflow:2.5.0
image: python:3.7
command:
- sh
- -c
- (PIP_DISABLE_PIP_VERSION_CHECK=1 python3 -m pip install --quiet --no-warn-script-location
'tensorflow==2.5.0' 'tf-agents==0.8.0' || PIP_DISABLE_PIP_VERSION_CHECK=1 python3
-m pip install --quiet --no-warn-script-location 'tensorflow==2.5.0' 'tf-agents==0.8.0'
--user) && "$0" "$@"
'tensorflow==2.5.0' 'tf-agents==0.8.0' 'Pillow' || PIP_DISABLE_PIP_VERSION_CHECK=1
python3 -m pip install --quiet --no-warn-script-location 'tensorflow==2.5.0'
'tf-agents==0.8.0' 'Pillow' --user) && "$0" "$@"
- sh
- -ec
- |
@@ -270,7 +270,8 @@ implementation:
def _serialize_str(str_value: str) -> str:
if not isinstance(str_value, str):
raise TypeError('Value "{}" has type "{}" instead of str.'.format(str(str_value), str(type(str_value))))
raise TypeError('Value "{}" has type "{}" instead of str.'.format(
str(str_value), str(type(str_value))))
return str_value
import argparse
@@ -22,13 +22,13 @@ from src.training import task
# Paths and configurations
DATA_PATH = "gs://[your-bucket-name]/[your-dataset-dir]/u.data" # FILL IN
DATA_PATH = "gs://[your-bucket-name]/artifacts/u.data" # FILL IN
ROOT_DIR = "gs://[your-bucket-name]/artifacts" # FILL IN
ARTIFACTS_DIR = "gs://[your-bucket-name]/artifacts" # FILL IN
PROFILER_DIR = "gs://[your-bucket-name]/profiler" # FILL IN
HPTUNING_RESULT_DIR = "[your-hptuning-result-dir]/" # FILL IN
HPTUNING_RESULT_PATH = os.path.join(HPTUNING_RESULT_DIR,
"[your-file-name].json") # FILL IN
"result.json") # FILL IN
RAW_BUCKET_NAME = "[your-hptuning-result-bucket-name]" # FILL IN
# Hyperparameters
@@ -1 +1 @@
tensorflow==2.4.1
tensorflow==2.5.3
@@ -8,7 +8,7 @@
},
"outputs": [],
"source": [
"# Copyright 2021 Google LLC\n",
"# Copyright 2022 Google LLC\n",
"#\n",
"# Licensed under the Apache License, Version 2.0 (the \"License\");\n",
"# you may not use this file except in compliance with the License.\n",
@@ -113,8 +113,8 @@
},
"outputs": [],
"source": [
"! gcloud beta ai custom-jobs local-run \\\n",
" --base-image=$BASE_IMAGE_URI \\\n",
"! gcloud ai custom-jobs local-run \\\n",
" --executor-image-uri=$BASE_IMAGE_URI \\\n",
" --script=$SCRIPT_PATH \\\n",
" --output-image-uri=$OUTPUT_IMAGE_NAME \\\n",
" -- \\\n",
-93
View File
@@ -1,93 +0,0 @@
# Code of Conduct
## Our Pledge
In the interest of fostering an open and welcoming environment, we as
contributors and maintainers pledge to making participation in our project and
our community a harassment-free experience for everyone, regardless of age, body
size, disability, ethnicity, gender identity and expression, level of
experience, education, socio-economic status, nationality, personal appearance,
race, religion, or sexual identity and orientation.
## Our Standards
Examples of behavior that contributes to creating a positive environment
include:
* Using welcoming and inclusive language
* Being respectful of differing viewpoints and experiences
* Gracefully accepting constructive criticism
* Focusing on what is best for the community
* Showing empathy towards other community members
Examples of unacceptable behavior by participants include:
* The use of sexualized language or imagery and unwelcome sexual attention or
advances
* Trolling, insulting/derogatory comments, and personal or political attacks
* Public or private harassment
* Publishing others' private information, such as a physical or electronic
address, without explicit permission
* Other conduct which could reasonably be considered inappropriate in a
professional setting
## Our Responsibilities
Project maintainers are responsible for clarifying the standards of acceptable
behavior and are expected to take appropriate and fair corrective action in
response to any instances of unacceptable behavior.
Project maintainers have the right and responsibility to remove, edit, or reject
comments, commits, code, wiki edits, issues, and other contributions that are
not aligned to this Code of Conduct, or to ban temporarily or permanently any
contributor for other behaviors that they deem inappropriate, threatening,
offensive, or harmful.
## Scope
This Code of Conduct applies both within project spaces and in public spaces
when an individual is representing the project or its community. Examples of
representing a project or community include using an official project e-mail
address, posting via an official social media account, or acting as an appointed
representative at an online or offline event. Representation of a project may be
further defined and clarified by project maintainers.
This Code of Conduct also applies outside the project spaces when the Project
Steward has a reasonable belief that an individual's behavior may have a
negative impact on the project or its community.
## Conflict Resolution
We do not believe that all conflict is bad; healthy debate and disagreement
often yield positive results. However, it is never okay to be disrespectful or
to engage in behavior that violates the project’s code of conduct.
If you see someone violating the code of conduct, you are encouraged to address
the behavior directly with those involved. Many issues can be resolved quickly
and easily, and this gives people more control over the outcome of their
dispute. If you are unable to resolve the matter for any reason, or if the
behavior is threatening or harassing, report it. We are dedicated to providing
an environment where participants feel welcome and safe.
Reports should be directed to *[PROJECT STEWARD NAME(s) AND EMAIL(s)]*, the
Project Steward(s) for *[PROJECT NAME]*. It is the Project Steward’s duty to
receive and address reported violations of the code of conduct. They will then
work with a committee consisting of representatives from the Open Source
Programs Office and the Google Open Source Strategy team. If for any reason you
are uncomfortable reaching out to the Project Steward, please email
opensource@google.com.
We will investigate every complaint, but you may not receive a direct response.
We will use our discretion in determining when and how to follow up on reported
incidents, which may range from not taking action to permanent expulsion from
the project and project-sponsored spaces. We will notify the accused of the
report and provide them an opportunity to discuss it before any action is taken.
The identity of the reporter will be omitted from the details of the report
supplied to the accused. In potentially harmful situations, such as ongoing
harassment or threats to anyone's safety, we may take action without notice.
## Attribution
This Code of Conduct is adapted from the Contributor Covenant, version 1.4,
available at
https://www.contributor-covenant.org/version/1/4/code-of-conduct.html
+5
View File
@@ -0,0 +1,5 @@
The [official](https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/main/notebooks/official) folder contains notebooks organized by Google Cloud product.
The [community](https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/main/notebooks/community) folder contains notebooks that aren't officially supported by Google.
Contributions to the repo should use the [notebook template](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/notebook_template.ipynb) as a starting point.
+14 -9
View File
@@ -3,14 +3,19 @@
# @global-owner1 and @global-owner2 will be requested for
# review when someone opens a pull request.
/sdk/sdk_* @aferlitsch
/gapic @aferlitsch
/ml_ops @aferlitsch
/model_monitoring/* @mco
/sdk/sdk_* @andrewferlitsch
/gapic @andrewferlitsch
/ml_ops @andrewferlitsch
/model_monitoring/* @mco-gh
/structured_data/rapid_prototyping_* @rafael-carvalho
/managed_notebooks/ @notebooks-team
/sdk/SDK_FBProphet_Forecasting_Online.ipynb @brianchunkang
/sdk/SDK_AutoML_Forecasting_Model_Training_Example.ipynb @thehardikv
/sdk/sdk_automl_forecasting_evaluating_a_model.ipynb @thehardikv
/matching_engine @yinghsienwu
/managed_notebooks/
/sdk/SDK_FBProphet_Forecasting_Online.ipynb @brianchunkang
/pipelines/google_cloud_pipeline_components_TPU_model_train_upload_deploy.ipynb @brianchunkang
/matching_engine/sdk_matching_engine_for_indexing.ipynb @ivanmkc
/matching_engine/matching_engine_for_indexing.ipynb @yinghsienwu
/sdk/pytorch_lightning_custom_container_training.ipynb @brianchunkang
/tensorboard @yfang1
/feature_store @nayaknishant @morgandu
/vertex_endpoints/tf_hub_obj_detection/deploy_tfhub_object_detection_on_vertex_endpoints.ipynb @entrpn
/vertex_endpoints/nvidia-triton/nvidia-triton-custom-container-prediction.ipynb @RajeshThallam
Binary file not shown.

After

Width:  |  Height:  |  Size: 153 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 138 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 88 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 182 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 142 KiB

File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
@@ -480,8 +480,7 @@
"\n",
"from google.cloud.aiplatform import gapic as aip\n",
"from google.protobuf import json_format\n",
"from google.protobuf.json_format import MessageToJson, ParseDict\n",
"from google.protobuf.struct_pb2 import Struct, Value"
"from google.protobuf.struct_pb2 import Value"
]
},
{
@@ -1496,9 +1495,8 @@
"\n",
"When you send a prediction or explanation request, the content of the request is base 64 decoded into a Tensorflow string (`tf.string`), which is passed to the serving function (`serving_fn`). The serving function preprocesses the `tf.string` into raw (uncompressed) numpy bytes (`preprocess_fn`) to match the input requirements of the model:\n",
"- `io.decode_jpeg`- Decompresses the JPG image which is returned as a Tensorflow tensor with three channels (RGB).\n",
"- `image.convert_image_dtype` - Changes integer pixel values to float 32.\n",
"- `image.convert_image_dtype` - Changes integer pixel values to float 32, and rescales pixel data between 0 and 1.\n",
"- `image.resize` - Resizes the image to match the input shape for the model.\n",
"- `resized / 255.0` - Rescales (normalization) the pixel data between 0 and 1.\n",
"\n",
"At this point, the data can be passed to the model (`m_call`)."
]
@@ -1518,8 +1516,7 @@
" decoded = tf.io.decode_jpeg(bytes_input, channels=3)\n",
" decoded = tf.image.convert_image_dtype(decoded, tf.float32)\n",
" resized = tf.image.resize(decoded, size=(32, 32))\n",
" rescale = tf.cast(resized / 255.0, tf.float32)\n",
" return rescale\n",
" return resized\n",
"\n",
"\n",
"@tf.function(input_signature=[tf.TensorSpec([None], tf.string)])\n",
@@ -505,8 +505,7 @@
"\n",
"import google.cloud.aiplatform_v1beta1 as aip\n",
"from google.protobuf import json_format\n",
"from google.protobuf.json_format import MessageToJson, ParseDict\n",
"from google.protobuf.struct_pb2 import Struct, Value"
"from google.protobuf.struct_pb2 import Value"
]
},
{
@@ -1521,9 +1520,8 @@
"\n",
"When you send a prediction or explanation request, the content of the request is base 64 decoded into a Tensorflow string (`tf.string`), which is passed to the serving function (`serving_fn`). The serving function preprocesses the `tf.string` into raw (uncompressed) numpy bytes (`preprocess_fn`) to match the input requirements of the model:\n",
"- `io.decode_jpeg`- Decompresses the JPG image which is returned as a Tensorflow tensor with three channels (RGB).\n",
"- `image.convert_image_dtype` - Changes integer pixel values to float 32.\n",
"- `image.convert_image_dtype` - Changes integer pixel values to float 32, and rescales pixel data between 0 and 1.\n",
"- `image.resize` - Resizes the image to match the input shape for the model.\n",
"- `resized / 255.0` - Rescales (normalization) the pixel data between 0 and 1.\n",
"\n",
"At this point, the data can be passed to the model (`m_call`).\n",
"\n",
@@ -1552,8 +1550,7 @@
" decoded = tf.io.decode_jpeg(bytes_input, channels=3)\n",
" decoded = tf.image.convert_image_dtype(decoded, tf.float32)\n",
" resized = tf.image.resize(decoded, size=(32, 32))\n",
" rescale = tf.cast(resized / 255.0, tf.float32)\n",
" return rescale\n",
" return resized\n",
"\n",
"\n",
"@tf.function(input_signature=[tf.TensorSpec([None], tf.string)])\n",
@@ -482,8 +482,7 @@
"\n",
"from google.cloud.aiplatform import gapic as aip\n",
"from google.protobuf import json_format\n",
"from google.protobuf.json_format import MessageToJson, ParseDict\n",
"from google.protobuf.struct_pb2 import Struct, Value"
"from google.protobuf.struct_pb2 import Value"
]
},
{
@@ -1498,9 +1497,8 @@
"\n",
"When you send a prediction or explanation request, the content of the request is base 64 decoded into a Tensorflow string (`tf.string`), which is passed to the serving function (`serving_fn`). The serving function preprocesses the `tf.string` into raw (uncompressed) numpy bytes (`preprocess_fn`) to match the input requirements of the model:\n",
"- `io.decode_jpeg`- Decompresses the JPG image which is returned as a Tensorflow tensor with three channels (RGB).\n",
"- `image.convert_image_dtype` - Changes integer pixel values to float 32.\n",
"- `image.convert_image_dtype` - Changes integer pixel values to float 32, and rescales pixel data between 0 and 1.\n",
"- `image.resize` - Resizes the image to match the input shape for the model.\n",
"- `resized / 255.0` - Rescales (normalization) the pixel data between 0 and 1.\n",
"\n",
"At this point, the data can be passed to the model (`m_call`)."
]
@@ -1520,8 +1518,7 @@
" decoded = tf.io.decode_jpeg(bytes_input, channels=3)\n",
" decoded = tf.image.convert_image_dtype(decoded, tf.float32)\n",
" resized = tf.image.resize(decoded, size=(32, 32))\n",
" rescale = tf.cast(resized / 255.0, tf.float32)\n",
" return rescale\n",
" return resized\n",
"\n",
"\n",
"@tf.function(input_signature=[tf.TensorSpec([None], tf.string)])\n",
@@ -483,8 +483,7 @@
"\n",
"from google.cloud.aiplatform import gapic as aip\n",
"from google.protobuf import json_format\n",
"from google.protobuf.json_format import MessageToJson, ParseDict\n",
"from google.protobuf.struct_pb2 import Struct, Value"
"from google.protobuf.struct_pb2 import Value"
]
},
{
@@ -1710,7 +1709,7 @@
"outputs": [],
"source": [
"import tensorflow as tf\n",
"from tensorflow.keras import Input, Model\n",
"from tensorflow.keras import Model\n",
"from tensorflow.keras.layers import Lambda\n",
"\n",
"softmax = model_A.outputs[0]\n",
@@ -1761,9 +1760,8 @@
"\n",
"When you send a prediction or explanation request, the content of the request is base 64 decoded into a Tensorflow string (`tf.string`), which is passed to the serving function (`serving_fn`). The serving function preprocesses the `tf.string` into raw (uncompressed) numpy bytes (`preprocess_fn`) to match the input requirements of the model:\n",
"- `io.decode_jpeg`- Decompresses the JPG image which is returned as a Tensorflow tensor with three channels (RGB).\n",
"- `image.convert_image_dtype` - Changes integer pixel values to float 32.\n",
"- `image.convert_image_dtype` - Changes integer pixel values to float 32, and rescales pixel data between 0 and 1.\n",
"- `image.resize` - Resizes the image to match the input shape for the model.\n",
"- `resized / 255.0` - Rescales (normalization) the pixel data between 0 and 1.\n",
"\n",
"At this point, the data can be passed to the model (`m_call`)."
]
@@ -1783,8 +1781,7 @@
" decoded = tf.io.decode_jpeg(bytes_input, channels=3)\n",
" decoded = tf.image.convert_image_dtype(decoded, tf.float32)\n",
" resized = tf.image.resize(decoded, size=(32, 32))\n",
" rescale = tf.cast(resized / 255.0, tf.float32)\n",
" return rescale\n",
" return resized\n",
"\n",
"\n",
"@tf.function(input_signature=[tf.TensorSpec([None], tf.string)])\n",
@@ -1824,16 +1821,15 @@
"CONCRETE_INPUT = \"numpy_inputs\"\n",
"\n",
"\n",
"def _preprocess(bytes_input):\n",
"def _preprocess(bytes_input): # noqa: 811\n",
" decoded = tf.io.decode_jpeg(bytes_input, channels=3)\n",
" decoded = tf.image.convert_image_dtype(decoded, tf.float32)\n",
" resized = tf.image.resize(decoded, size=(32, 32))\n",
" rescale = tf.cast(resized / 255.0, tf.float32)\n",
" return rescale\n",
" return resized\n",
"\n",
"\n",
"@tf.function(input_signature=[tf.TensorSpec([None], tf.string)])\n",
"def preprocess_fn(bytes_inputs):\n",
"def preprocess_fn(bytes_inputs): # noqa: 811\n",
" decoded_images = tf.map_fn(\n",
" _preprocess, bytes_inputs, dtype=tf.float32, back_prop=False\n",
" )\n",
@@ -482,8 +482,7 @@
"\n",
"from google.cloud.aiplatform import gapic as aip\n",
"from google.protobuf import json_format\n",
"from google.protobuf.json_format import MessageToJson, ParseDict\n",
"from google.protobuf.struct_pb2 import Struct, Value"
"from google.protobuf.struct_pb2 import Value"
]
},
{
@@ -1600,9 +1599,8 @@
"\n",
"When you send a prediction or explanation request, the content of the request is base 64 decoded into a Tensorflow string (`tf.string`), which is passed to the serving function (`serving_fn`). The serving function preprocesses the `tf.string` into raw (uncompressed) numpy bytes (`preprocess_fn`) to match the input requirements of the model:\n",
"- `io.decode_jpeg`- Decompresses the JPG image which is returned as a Tensorflow tensor with three channels (RGB).\n",
"- `image.convert_image_dtype` - Changes integer pixel values to float 32.\n",
"- `image.convert_image_dtype` - Changes integer pixel values to float 32, and rescales pixel data between 0 and 1.\n",
"- `image.resize` - Resizes the image to match the input shape for the model.\n",
"- `resized / 255.0` - Rescales (normalization) the pixel data between 0 and 1.\n",
"\n",
"At this point, the data can be passed to the model (`m_call`)."
]
@@ -1622,8 +1620,7 @@
" decoded = tf.io.decode_jpeg(bytes_input, channels=3)\n",
" decoded = tf.image.convert_image_dtype(decoded, tf.float32)\n",
" resized = tf.image.resize(decoded, size=(32, 32))\n",
" rescale = tf.cast(resized / 255.0, tf.float32)\n",
" return rescale\n",
" return resized\n",
"\n",
"\n",
"@tf.function(input_signature=[tf.TensorSpec([None], tf.string)])\n",
@@ -507,8 +507,7 @@
"\n",
"import google.cloud.aiplatform_v1beta1 as aip\n",
"from google.protobuf import json_format\n",
"from google.protobuf.json_format import MessageToJson, ParseDict\n",
"from google.protobuf.struct_pb2 import Struct, Value"
"from google.protobuf.struct_pb2 import Value"
]
},
{
@@ -1523,9 +1522,8 @@
"\n",
"When you send a prediction or explanation request, the content of the request is base 64 decoded into a Tensorflow string (`tf.string`), which is passed to the serving function (`serving_fn`). The serving function preprocesses the `tf.string` into raw (uncompressed) numpy bytes (`preprocess_fn`) to match the input requirements of the model:\n",
"- `io.decode_jpeg`- Decompresses the JPG image which is returned as a Tensorflow tensor with three channels (RGB).\n",
"- `image.convert_image_dtype` - Changes integer pixel values to float 32.\n",
"- `image.convert_image_dtype` - Changes integer pixel values to float 32, and rescales pixel data between 0 and 1.\n",
"- `image.resize` - Resizes the image to match the input shape for the model.\n",
"- `resized / 255.0` - Rescales (normalization) the pixel data between 0 and 1.\n",
"\n",
"At this point, the data can be passed to the model (`m_call`).\n",
"\n",
@@ -1554,8 +1552,7 @@
" decoded = tf.io.decode_jpeg(bytes_input, channels=3)\n",
" decoded = tf.image.convert_image_dtype(decoded, tf.float32)\n",
" resized = tf.image.resize(decoded, size=(32, 32))\n",
" rescale = tf.cast(resized / 255.0, tf.float32)\n",
" return rescale\n",
" return resized\n",
"\n",
"\n",
"@tf.function(input_signature=[tf.TensorSpec([None], tf.string)])\n",
@@ -484,8 +484,7 @@
"\n",
"from google.cloud.aiplatform import gapic as aip\n",
"from google.protobuf import json_format\n",
"from google.protobuf.json_format import MessageToJson, ParseDict\n",
"from google.protobuf.struct_pb2 import Struct, Value"
"from google.protobuf.struct_pb2 import Value"
]
},
{
@@ -2103,9 +2102,8 @@
"\n",
"When you send a prediction or explanation request, the content of the request is base 64 decoded into a Tensorflow string (`tf.string`), which is passed to the serving function (`serving_fn`). The serving function preprocesses the `tf.string` into raw (uncompressed) numpy bytes (`preprocess_fn`) to match the input requirements of the model:\n",
"- `io.decode_jpeg`- Decompresses the JPG image which is returned as a Tensorflow tensor with three channels (RGB).\n",
"- `image.convert_image_dtype` - Changes integer pixel values to float 32.\n",
"- `image.convert_image_dtype` - Changes integer pixel values to float 32, and rescales pixel data between 0 and 1.\n",
"- `image.resize` - Resizes the image to match the input shape for the model.\n",
"- `resized / 255.0` - Rescales (normalization) the pixel data between 0 and 1.\n",
"\n",
"At this point, the data can be passed to the model (`m_call`)."
]
@@ -2125,8 +2123,7 @@
" decoded = tf.io.decode_jpeg(bytes_input, channels=3)\n",
" decoded = tf.image.convert_image_dtype(decoded, tf.float32)\n",
" resized = tf.image.resize(decoded, size=(128, 128))\n",
" rescale = tf.cast(resized / 255.0, tf.float32)\n",
" return rescale\n",
" return resized\n",
"\n",
"\n",
"@tf.function(input_signature=[tf.TensorSpec([None], tf.string)])\n",
@@ -483,8 +483,7 @@
"\n",
"from google.cloud.aiplatform import gapic as aip\n",
"from google.protobuf import json_format\n",
"from google.protobuf.json_format import MessageToJson, ParseDict\n",
"from google.protobuf.struct_pb2 import Struct, Value"
"from google.protobuf.struct_pb2 import Value"
]
},
{
@@ -1247,9 +1246,6 @@
},
"outputs": [],
"source": [
"from google.protobuf import json_format\n",
"from google.protobuf.struct_pb2 import Value\n",
"\n",
"MODEL_NAME = \"custom_pipeline-\" + TIMESTAMP\n",
"PIPELINE_DISPLAY_NAME = \"custom-training-pipeline\" + TIMESTAMP\n",
"\n",
@@ -1567,9 +1563,8 @@
"\n",
"When you send a prediction or explanation request, the content of the request is base 64 decoded into a Tensorflow string (`tf.string`), which is passed to the serving function (`serving_fn`). The serving function preprocesses the `tf.string` into raw (uncompressed) numpy bytes (`preprocess_fn`) to match the input requirements of the model:\n",
"- `io.decode_jpeg`- Decompresses the JPG image which is returned as a Tensorflow tensor with three channels (RGB).\n",
"- `image.convert_image_dtype` - Changes integer pixel values to float 32.\n",
"- `image.convert_image_dtype` - Changes integer pixel values to float 32, and rescales pixel data between 0 and 1.\n",
"- `image.resize` - Resizes the image to match the input shape for the model.\n",
"- `resized / 255.0` - Rescales (normalization) the pixel data between 0 and 1.\n",
"\n",
"At this point, the data can be passed to the model (`m_call`)."
]
@@ -1589,8 +1584,7 @@
" decoded = tf.io.decode_jpeg(bytes_input, channels=3)\n",
" decoded = tf.image.convert_image_dtype(decoded, tf.float32)\n",
" resized = tf.image.resize(decoded, size=(32, 32))\n",
" rescale = tf.cast(resized / 255.0, tf.float32)\n",
" return rescale\n",
" return resized\n",
"\n",
"\n",
"@tf.function(input_signature=[tf.TensorSpec([None], tf.string)])\n",
@@ -482,8 +482,7 @@
"\n",
"from google.cloud.aiplatform import gapic as aip\n",
"from google.protobuf import json_format\n",
"from google.protobuf.json_format import MessageToJson, ParseDict\n",
"from google.protobuf.struct_pb2 import Struct, Value"
"from google.protobuf.struct_pb2 import Value"
]
},
{
@@ -1559,9 +1558,8 @@
"\n",
"When you send a prediction or explanation request, the content of the request is base 64 decoded into a Tensorflow string (`tf.string`), which is passed to the serving function (`serving_fn`). The serving function preprocesses the `tf.string` into raw (uncompressed) numpy bytes (`preprocess_fn`) to match the input requirements of the model:\n",
"- `io.decode_jpeg`- Decompresses the JPG image which is returned as a Tensorflow tensor with three channels (RGB).\n",
"- `image.convert_image_dtype` - Changes integer pixel values to float 32.\n",
"- `image.convert_image_dtype` - Changes integer pixel values to float 32, and rescales pixel data between 0 and 1.\n",
"- `image.resize` - Resizes the image to match the input shape for the model.\n",
"- `resized / 255.0` - Rescales (normalization) the pixel data between 0 and 1.\n",
"\n",
"At this point, the data can be passed to the model (`m_call`)."
]
@@ -1581,8 +1579,7 @@
" decoded = tf.io.decode_jpeg(bytes_input, channels=3)\n",
" decoded = tf.image.convert_image_dtype(decoded, tf.float32)\n",
" resized = tf.image.resize(decoded, size=(32, 32))\n",
" rescale = tf.cast(resized / 255.0, tf.float32)\n",
" return rescale\n",
" return resized\n",
"\n",
"\n",
"@tf.function(input_signature=[tf.TensorSpec([None], tf.string)])\n",
@@ -483,8 +483,7 @@
"\n",
"from google.cloud.aiplatform import gapic as aip\n",
"from google.protobuf import json_format\n",
"from google.protobuf.json_format import MessageToJson, ParseDict\n",
"from google.protobuf.struct_pb2 import Struct, Value"
"from google.protobuf.struct_pb2 import Value"
]
},
{
@@ -1501,9 +1500,8 @@
"\n",
"When you send a prediction or explanation request, the content of the request is base 64 decoded into a Tensorflow string (`tf.string`), which is passed to the serving function (`serving_fn`). The serving function preprocesses the `tf.string` into raw (uncompressed) numpy bytes (`preprocess_fn`) to match the input requirements of the model:\n",
"- `io.decode_jpeg`- Decompresses the JPG image which is returned as a Tensorflow tensor with three channels (RGB).\n",
"- `image.convert_image_dtype` - Changes integer pixel values to float 32.\n",
"- `image.convert_image_dtype` - Changes integer pixel values to float 32, and rescales pixel data between 0 and 1.\n",
"- `image.resize` - Resizes the image to match the input shape for the model.\n",
"- `resized / 255.0` - Rescales (normalization) the pixel data between 0 and 1.\n",
"\n",
"At this point, the data can be passed to the model (`m_call`).\n",
"\n",
@@ -1529,8 +1527,7 @@
" decoded = tf.io.decode_jpeg(bytes_input, channels=3)\n",
" decoded = tf.image.convert_image_dtype(decoded, tf.float32)\n",
" resized = tf.image.resize(decoded, size=(16, 16))\n",
" rescale = tf.cast(resized / 255.0, tf.float32)\n",
" return rescale\n",
" return resized\n",
"\n",
"\n",
"@tf.function(input_signature=[tf.TensorSpec([None], tf.string)])\n",
File diff suppressed because it is too large Load Diff
Binary file not shown.

After

Width:  |  Height:  |  Size: 382 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 445 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 63 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 59 KiB

@@ -0,0 +1,823 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {
"id": "d1cc1c1fa076"
},
"source": [
"# Pricing Optimization \n",
"## Table of contents\n",
"* [Overview](#section-1)\n",
"* [Dataset](#section-2)\n",
"* [Objective](#section-3)\n",
"* [Costs](#section-4)\n",
"* [Create a BigQuery dataset](#section-5)\n",
"* [Load the dataset from Cloud Storage](#section-6)\n",
"* [Data analysis](#section-7)\n",
"* [Preprocess the data for training](#section-8)\n",
"* [Train the model using BigQuery ML](#section-9)\n",
"* [Generate forecasts from the model](#section-10)\n",
"* [Interpret the results to choose the best price](#section-11)\n",
"* [Clean up](#section-12)\n",
"\n",
"## Overview\n",
"<a name=\"section-1\"></a>\n",
"\n",
"This notebook demonstrates analysis of pricing optimization on [CDM Pricing Data](https://github.com/trifacta/trifacta-google-cloud/tree/main/design-pattern-pricing-optimization) and automating the workflow using Vertex AI Workbench managed notebooks.\n",
"\n",
"*Note: This notebook file was developed to run in a [Vertex AI Workbench managed notebooks](https://console.cloud.google.com/vertex-ai/workbench/list/managed) instance using the Python (Local) kernel. Some components of this notebook may not work in other notebook environments.*\n",
"\n",
"## Dataset\n",
"<a name=\"section-2\"></a>\n",
"\n",
"The dataset used in this notebook is a part of the [CDM Pricing dataset](https://github.com/trifacta/trifacta-google-cloud/blob/main/design-pattern-pricing-optimization/CDM_Pricing_large_table.csv), which consists of product sales information on specified dates.\n",
"\n",
"## Objective\n",
"<a name=\"section-3\"></a>\n",
"\n",
"The objective of this notebook is to build a pricing optimization model using Vertex AI. The following steps have been followed: \n",
"\n",
"- Load the required dataset from a Cloud Storage bucket.\n",
"- Analyze the fields present in the dataset.\n",
"- Process the data to build a model.\n",
"- Build a BigQuery ML forecast model on the processed data.\n",
"- Get forecasted values from the BigQuery ML model.\n",
"- Interpret the forecasts to identify the best prices.\n",
"- Clean up.\n",
"\n",
"## Costs\n",
"<a name=\"section-4\"></a>\n",
"\n",
"This tutorial uses the following billable components of Google Cloud:\n",
"\n",
"- Vertex AI\n",
"- BigQuery\n",
"- Cloud Storage\n",
"\n",
"\n",
"Learn about [Vertex AI\n",
"pricing](https://cloud.google.com/vertex-ai/pricing), [BigQuery pricing](https://cloud.google.com/bigquery/pricing) and [Cloud Storage\n",
"pricing](https://cloud.google.com/storage/pricing), and use the [Pricing\n",
"Calculator](https://cloud.google.com/products/calculator/)\n",
"to generate a cost estimate based on your projected usage.\n"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "5ed1f5e85640"
},
"source": [
"## Before you begin\n",
"\n",
"### Set your project ID\n",
"\n",
"**If you don't know your project ID**, you may be able to get your project ID using `gcloud`."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "c3f30148b66d"
},
"outputs": [],
"source": [
"import os\n",
"\n",
"PROJECT_ID = \"\"\n",
"\n",
"# Get your Google Cloud project ID from gcloud\n",
"if not os.getenv(\"IS_TESTING\"):\n",
" shell_output = !gcloud config list --format 'value(core.project)' 2>/dev/null\n",
" PROJECT_ID = shell_output[0]\n",
" print(\"Project ID: \", PROJECT_ID)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "750bf2883c2d"
},
"source": [
"Otherwise, set your project ID here."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "3c6db1ca88b9"
},
"outputs": [],
"source": [
"if PROJECT_ID == \"\" or PROJECT_ID is None:\n",
" PROJECT_ID = \"[your-project-id]\" # @param {type:\"string\"}"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "2a1c270c7d34"
},
"source": [
"### Import the required libraries and define constants\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "acc6fac1fa55"
},
"outputs": [],
"source": [
"import matplotlib.pyplot as plt\n",
"import pandas as pd\n",
"import seaborn as sns\n",
"from google.cloud import bigquery\n",
"from google.cloud.bigquery import Client"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "a06006dff8f9"
},
"outputs": [],
"source": [
"DATASET = \"[your-bigquery-dataset-id]\" # set the BigQuery dataset-id\n",
"TRAINING_DATA_TABLE = \"[your-bigquery-table-id-to-store-the-training-data]\" # set the BigQuery table-id to store the training data"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "016c3d47cc69"
},
"source": [
"## Create a BigQuery dataset\n",
"<a name=\"section-5\"></a>\n"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "12ccd8d7956e"
},
"source": [
"#@bigquery\n",
"-- create a dataset in BigQuery\n",
"\n",
"CREATE SCHEMA pricing_optimization\n",
"OPTIONS(\n",
" location=\"us\"\n",
" )"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "c106b978a79b"
},
"source": [
"## Load the dataset from Cloud Storage\n",
"<a name=\"section-6\"></a>\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "8aeae9da9796"
},
"outputs": [],
"source": [
"DATA_LOCATION = \"gs://cloud-samples-data/ai-platform-unified/datasets/tabular/cdm_pricing_large_table.csv\"\n",
"df = pd.read_csv(DATA_LOCATION)\n",
"print(df.shape)\n",
"df.head()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "7b98d5f09842"
},
"source": [
"You will build a forecast model on this data and thus determine the best price for a product. For this type of model, you will not be using many fields: only the sales and price related ones. For the current execrcise, focus on the following fields:\n",
"\n",
"- `Product_ID`\n",
"- `Customer_Hierarchy`\n",
"- `Fiscal_Date`\n",
"- `List_Price_Converged`\n",
"- `Invoiced_quantity_in_Pieces`\n",
"- `Net_Sales`\n",
"\n",
"## Data Analysis\n",
"<a name=\"section-7\"></a>\n",
"\n",
"First, explore the data and distributions.\n",
"\n",
"Select the required columns from the dataframe."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "af4b41c5eb1f"
},
"outputs": [],
"source": [
"id_col = \"Product_ID\"\n",
"date_col = \"Fiscal_Date\"\n",
"categ_cols = [\"Customer_Hierarchy\"]\n",
"num_cols = [\"List_Price_Converged\", \"Invoiced_quantity_in_Pieces\", \"Net_Sales\"]\n",
"\n",
"df = df[[id_col, date_col] + categ_cols + num_cols].copy()\n",
"df.head()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "3d780043ee5b"
},
"source": [
"Check the column types and null values in the dataframe."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "f54c445a1288"
},
"outputs": [],
"source": [
"df.info()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "cd817b414c4d"
},
"source": [
"This data description reveals that there are no null values in the data. Also, the field `Fiscal_Date` which is a date field is loaded as an object type. \n",
"\n",
"Change the type of the date field to datetime."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "b160fac085c8"
},
"outputs": [],
"source": [
"df[\"Fiscal_Date\"] = pd.to_datetime(df[\"Fiscal_Date\"], infer_datetime_format=True)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "fb4778578064"
},
"source": [
"Plot the distributions for the categorical fields."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "dd0467cd57c3"
},
"outputs": [],
"source": [
"for i in categ_cols:\n",
" df[i].value_counts(normalize=True).plot(kind=\"bar\")\n",
" plt.title(i)\n",
" plt.show()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "145deed255e0"
},
"source": [
"Plot the distributions for the numerical fields."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "f934137c6d82"
},
"outputs": [],
"source": [
"for i in num_cols:\n",
" _, ax = plt.subplots(1, 2, figsize=(10, 4))\n",
" df[i].plot(kind=\"box\", ax=ax[0])\n",
" df[i].plot(kind=\"hist\", ax=ax[1])\n",
" ax[0].set_title(i + \"-Boxplot\")\n",
" ax[1].set_title(i + \"-Histogram\")\n",
" plt.show()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "f9b9c2e58380"
},
"source": [
"Check the maximum date and minimum date in Fiscal_Date column."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "2a10aa689f9d"
},
"outputs": [],
"source": [
"print(df[\"Fiscal_Date\"].max())\n",
"print(df[\"Fiscal_Date\"].min())"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "4834f63e2e59"
},
"source": [
"Check the product distribution across each category."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "4664877f5304"
},
"outputs": [],
"source": [
"grp_cols = [\"Customer_Hierarchy\", \"Product_ID\"]\n",
"grp_df = df[grp_cols].groupby(by=grp_cols).count().reset_index()\n",
"grp_df.groupby(\"Customer_Hierarchy\").nunique()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "01ed02b9c8fd"
},
"source": [
"Check the percentage changes in the orders based on the percentage changes in the price."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "0b2c428cb135"
},
"outputs": [],
"source": [
"# aggregate the data\n",
"df_aggr = (\n",
" df.groupby([\"Product_ID\", \"List_Price_Converged\"])\n",
" .agg({\"Fiscal_Date\": min, \"Invoiced_quantity_in_Pieces\": sum, \"Net_Sales\": sum})\n",
" .reset_index()\n",
")\n",
"# rename the aggregated columns\n",
"df_aggr.rename(\n",
" columns={\n",
" \"Fiscal_Date\": \"First_price_date\",\n",
" \"Invoiced_quantity_in_Pieces\": \"Total_ordered_pieces\",\n",
" \"Net_Sales\": \"Total_net_sales\",\n",
" },\n",
" inplace=True,\n",
")\n",
"\n",
"# sort values chronologically\n",
"df_aggr.sort_values(by=[\"Product_ID\", \"First_price_date\"], inplace=True)\n",
"df_aggr.reset_index(drop=True, inplace=True)\n",
"\n",
"# add columns for previous values\n",
"df_aggr[\"Previous_List\"] = df_aggr.groupby([\"Product_ID\"])[\n",
" \"List_Price_Converged\"\n",
"].shift()\n",
"df_aggr[\"Previous_Total_ordered_pieces\"] = df_aggr.groupby([\"Product_ID\"])[\n",
" \"Total_ordered_pieces\"\n",
"].shift()\n",
"\n",
"# average price change across sku's\n",
"df_aggr[\"price_change_perc\"] = (\n",
" (df_aggr[\"List_Price_Converged\"] - df_aggr[\"Previous_List\"])\n",
" / df_aggr[\"Previous_List\"].fillna(0)\n",
" * 100\n",
")\n",
"df_aggr[\"order_change_perc\"] = (\n",
" (df_aggr[\"Total_ordered_pieces\"] - df_aggr[\"Previous_Total_ordered_pieces\"])\n",
" / df_aggr[\"Previous_Total_ordered_pieces\"].fillna(0)\n",
" * 100\n",
")\n",
"\n",
"# plot a scatterplot to visualize the changes\n",
"sns.scatterplot(\n",
" x=\"price_change_perc\",\n",
" y=\"order_change_perc\",\n",
" data=df_aggr,\n",
" hue=\"Product_ID\",\n",
" legend=False,\n",
")\n",
"plt.title(\"Percentage of change in price vs order\")\n",
"plt.show()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "8259e916fe25"
},
"source": [
"For most of the products, the percentage change in orders are high where the percentage changes in the prices are low. This suggests that too much change in the prices can affect the number of orders. \n",
"\n",
"**Note**: There seem to be some outliers in the data as percentage changes greater than 800 are found. In the current exercise, do not take any manual measures to deal with outliers as you will create a BigQuery ML timeseries model that already deals with outliers.\n",
"\n",
"## Preprocess the data for training\n",
"<a name=\"section-8\"></a>\n",
"\n",
"Check which `Product_ID`'s have the maximum orders."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "f5cbc7709c6a"
},
"outputs": [],
"source": [
"df_orders = df.groupby([\"Product_ID\", \"Customer_Hierarchy\"], as_index=False)[\n",
" \"Invoiced_quantity_in_Pieces\"\n",
"].sum()\n",
"df_orders.loc[\n",
" df_orders.groupby(\"Customer_Hierarchy\")[\"Invoiced_quantity_in_Pieces\"].idxmax()\n",
"]"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "fd6d227e513e"
},
"source": [
"From the above result, you can infer the following:\n",
"\n",
"- Under the **Food** category, **SKU 62** has the maximum orders.\n",
"- Under the **Manufacturing** category, **SKU 17** has the maximum orders.\n",
"- Under the **Paper** category, **SKU 107** has the maximum orders.\n",
"- Under the **Publishing** category, **SKU 8** has the maximum orders.\n",
"- Under the **Utilities** category, **SKU 140** has the maximum orders.\n",
"\n",
"Given that there are too many ids and only a few records for most of them, consider only the above `Product_ID`s for which there are a maximum number of orders. \n",
"\n",
"**Note**: The `Invoiced_quantity_in_Pieces` field seems to be a *float* type rather than an *int* type as it should be. This could be because the data itself might be averaged in the first place."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "2dbc0d64d157"
},
"source": [
"Check the various prices available for these `Product_ID`s."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "acc1dbd2d838"
},
"outputs": [],
"source": [
"df_type_food = df[(df[\"Product_ID\"] == \"SKU 62\") & (df[\"Customer_Hierarchy\"] == \"Food\")]\n",
"print(\"Food :\")\n",
"print(df_type_food[\"List_Price_Converged\"].value_counts())\n",
"df_type_manuf = df[\n",
" (df[\"Product_ID\"] == \"SKU 17\") & (df[\"Customer_Hierarchy\"] == \"Manufacturing\")\n",
"]\n",
"print(\"Manufacturing :\")\n",
"print(df_type_manuf[\"List_Price_Converged\"].value_counts())\n",
"df_type_paper = df[\n",
" (df[\"Product_ID\"] == \"SKU 107\") & (df[\"Customer_Hierarchy\"] == \"Paper\")\n",
"]\n",
"print(\"Paper :\")\n",
"print(df_type_paper[\"List_Price_Converged\"].value_counts())\n",
"df_type_pub = df[\n",
" (df[\"Product_ID\"] == \"SKU 8\") & (df[\"Customer_Hierarchy\"] == \"Publishing\")\n",
"]\n",
"print(\"Publishing :\")\n",
"print(df_type_pub[\"List_Price_Converged\"].value_counts())\n",
"df_type_util = df[\n",
" (df[\"Product_ID\"] == \"SKU 140\") & (df[\"Customer_Hierarchy\"] == \"Utilities\")\n",
"]\n",
"print(\"Utilities :\")\n",
"print(df_type_util[\"List_Price_Converged\"].value_counts())"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "f023af578c0f"
},
"source": [
"In the publishing category, `Product_ID` `SKU 8` and `SKU 17` are less than or equal to two different prices in the entire data and so you will exclude them and consider the rest for building the forecast model. The idea here is to train a forecast model on the timeseries data for products with different prices.\n",
"\n",
"Join the data for all the `Product_ID`s into one dataframe and remove duplicate records."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "a44771cc4c20"
},
"outputs": [],
"source": [
"df_final = pd.concat([df_type_food, df_type_paper, df_type_util])\n",
"df_final = (\n",
" df_final[\n",
" [\n",
" \"Product_ID\",\n",
" \"Fiscal_Date\",\n",
" \"Customer_Hierarchy\",\n",
" \"List_Price_Converged\",\n",
" \"Invoiced_quantity_in_Pieces\",\n",
" ]\n",
" ]\n",
" .drop_duplicates()\n",
" .reset_index(drop=True)\n",
")\n",
"df_final.head()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "add5063df368"
},
"source": [
"Save the data to a BigQuery table."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "fd82ba56571f"
},
"outputs": [],
"source": [
"bq_client = bigquery.Client(project=PROJECT_ID)\n",
"\n",
"job_config = bigquery.LoadJobConfig(\n",
" # Specify a (partial) schema. All columns are always written to the\n",
" # table. The schema is used to assist in data type definitions.\n",
" schema=[\n",
" bigquery.SchemaField(\"Product_ID\", bigquery.enums.SqlTypeNames.STRING),\n",
" bigquery.SchemaField(\"Fiscal_Date\", bigquery.enums.SqlTypeNames.DATE),\n",
" bigquery.SchemaField(\"List_Price_Converged\", bigquery.enums.SqlTypeNames.FLOAT),\n",
" bigquery.SchemaField(\n",
" \"Invoiced_quantity_in_Pieces\", bigquery.enums.SqlTypeNames.FLOAT\n",
" ),\n",
" ],\n",
" # Optionally, set the write disposition. BigQuery appends loaded rows\n",
" # to an existing table by default, but with WRITE_TRUNCATE write\n",
" # disposition it replaces the table with the loaded data.\n",
" write_disposition=\"WRITE_TRUNCATE\",\n",
")\n",
"\n",
"# save the dataframe to a table in the created dataset\n",
"job = bq_client.load_table_from_dataframe(\n",
" df_final,\n",
" \"{}.{}.{}\".format(PROJECT_ID, DATASET, TRAINING_DATA_TABLE),\n",
" job_config=job_config,\n",
") # Make an API request.\n",
"job.result() # Wait for the job to complete."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "fca77641b03b"
},
"source": [
"# Train the model using BigQuery ML\n",
"<a name=\"section-9\"></a>\n",
"\n",
"Train an [Arima-Plus](https://cloud.google.com/bigquery-ml/docs/reference/standard-sql/bigqueryml-syntax-create-time-series) model on the data using BigQuery ML."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "cded27507891"
},
"source": [
"#@bigquery\n",
"create or replace model pricing_optimization.bqml_arima\n",
"options\n",
" (model_type = 'ARIMA_PLUS',\n",
" time_series_timestamp_col = 'Fiscal_Date',\n",
" time_series_data_col = 'Invoiced_quantity_in_Pieces',\n",
" time_series_id_col = 'ID'\n",
" ) as\n",
"select\n",
" Fiscal_Date,\n",
" Concat(Product_ID,\"_\" ,Cast(List_Price_Converged as string)) as ID,\n",
" Invoiced_quantity_in_Pieces\n",
"from\n",
" pricing_optimization.TRAINING_DATA\n"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "332fd11ff32b"
},
"source": [
"## Generate forecasts from the model\n",
"<a name=\"section-10\"></a>\n",
"\n",
"Predict the sales for the next 30 days for each id and save to a dataframe."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "ef926cdbf28e"
},
"outputs": [],
"source": [
"client = Client()\n",
"\n",
"query = '''\n",
"DECLARE HORIZON STRING DEFAULT \"30\"; #number of values to forecast\n",
"DECLARE CONFIDENCE_LEVEL STRING DEFAULT \"0.90\"; ## required confidence level\n",
"\n",
"EXECUTE IMMEDIATE format(\"\"\"\n",
" SELECT\n",
" *\n",
" FROM \n",
" ML.FORECAST(MODEL pricing_optimization.bqml_arima, \n",
" STRUCT(%s AS horizon, \n",
" %s AS confidence_level)\n",
" )\n",
" \"\"\",HORIZON,CONFIDENCE_LEVEL)'''\n",
"job = client.query(query)\n",
"dfforecast = job.to_dataframe()\n",
"dfforecast.head()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "608c7de72dae"
},
"source": [
"## Interpret the results to choose the best price\n",
"<a name=\"section-11\"></a>\n",
"\n",
"Calculate average forecast values for the forecast duration."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "e1e193680400"
},
"outputs": [],
"source": [
"dfforecast_avg = (\n",
" dfforecast[[\"ID\", \"forecast_value\"]].groupby(\"ID\", as_index=False).mean()\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "5ce395d652a3"
},
"source": [
"Extract the ID and Price fields from the ID field."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "452c56fa58ed"
},
"outputs": [],
"source": [
"dfforecast_avg[\"Product_ID\"] = dfforecast_avg[\"ID\"].apply(lambda x: x.split(\"_\")[0])\n",
"dfforecast_avg[\"Price\"] = dfforecast_avg[\"ID\"].apply(lambda x: x.split(\"_\")[1])"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "3cee67f4028f"
},
"source": [
"Plot the average forecasted sales vs. the price of the product."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "fb351c8f383d"
},
"outputs": [],
"source": [
"for i in dfforecast_avg[\"Product_ID\"].unique():\n",
" dfforecast_avg[dfforecast_avg[\"Product_ID\"] == i].set_index(\"Price\").sort_values(\n",
" \"forecast_value\"\n",
" ).plot(kind=\"bar\")\n",
" plt.title(\"Price vs. Average Sales for \" + i)\n",
" plt.show()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "67ff3acc74a5"
},
"source": [
"Based on the plots for price vs. the average forecasted orders, it can be said that to use the maximum orders, each of the considered `Product_ID`s can follow the below prices:\n",
"\n",
"- SKU 107's price range can be from 4.44 - 4.73 units\n",
"- SKU 140's price can be 1.95 units\n",
"- SKU 62's price can be 4.23 units\n",
"\n",
"\n",
"## Clean Up\n",
"<a name=\"section-12\"></a>\n",
"\n",
"To clean up all Google Cloud resources used in this project, you can [delete the Google Cloud project](https://cloud.google.com/resource-manager/docs/creating-managing-projects#shutting_down_projects) you used for the tutorial.\n",
"\n",
"Otherwise, you can delete the individual resources you created in this tutorial. The following code deletes the entire dataset."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "d78908b8134d"
},
"outputs": [],
"source": [
"# Construct a BigQuery client object.\n",
"client = bigquery.Client()\n",
"\n",
"# TODO(developer): Set model_id to the ID of the model to fetch.\n",
"dataset_id = \"{PROJECT}.{DATASET}\".format(PROJECT=PROJECT_ID, DATASET=DATASET)\n",
"\n",
"# Use the delete_contents parameter to delete a dataset and its contents.\n",
"# Use the not_found_ok parameter to not receive an error if the dataset has already been deleted.\n",
"client.delete_dataset(\n",
" dataset_id, delete_contents=True, not_found_ok=True\n",
") # Make an API request.\n",
"\n",
"print(\"Deleted dataset '{}'.\".format(dataset_id))"
]
}
],
"metadata": {
"colab": {
"name": "pricing-optimization.ipynb",
"toc_visible": true
},
"kernelspec": {
"display_name": "Python 3",
"name": "python3"
}
},
"nbformat": 4,
"nbformat_minor": 0
}
@@ -459,7 +459,7 @@
},
"outputs": [],
"source": [
"! gsutil cp gs://cloud-samples-data/ai-platform-unified/matching_engine/glove-100-angular.hdf5 ."
"! gsutil cp gs://cloud-samples-data/vertex-ai/matching_engine/glove-100-angular.hdf5 ."
]
},
{
File diff suppressed because it is too large Load Diff
+3 -3
View File
@@ -10,9 +10,9 @@ The purpose of this set of notebooks and markdown files is to demonstrate Google
1. [Data Management](stage1)
2. [Experimentation](stage2)
3. Formalization
4. Evaluation
3. [Formalization](stage3)
4. [Evaluation](stage4)
5. Deployment
6. Serving
6. [Serving](stage6)
7. Monitoring
8. Continuous Training
@@ -30,10 +30,66 @@ The first stage in MLOps is the collection and preparation for the purpose of de
[Get Started with BQ datasets](get_started_bq_datasets.ipynb)
```
The steps performed include:
- Create a Vertex AI `Dataset` resource from `BigQuery` table -- compatible for `AutoML` training.
- Extract a copy of the dataset from `BigQuery` to a CSV file in Cloud Storage -- compatible for `AutoML` or custom training.
- Select rows from a `BigQuery` dataset into a `pandas` dataframe -- compatible for custom training.
- Select rows from a `BigQuery` dataset into a `tf.data.Dataset` -- compatible for custom training `TensorFlow` models.
- Select rows from extracted CSV files into a `tf.data.Dataset` -- compatible for custom training `TensorFlow` models.
- Create a `BigQuery` dataset from CSV files.
- Extract data from `BigQuery` table into a `DMatrix` -- compatible for custom training `XGBoost` models.
```
[Get Started with Vertex datasets](get_started_vertex_datasets.ipynb)
```
The steps performed include:
- Create a Vertex AI `Dataset` resource for:
- image data
- text data
- video data
- tabular data
- forecasting data
- Search `Dataset` resources using a filter.
- Read a sample of a `BigQuery` dataset into a dataframe.
- Generate statistics and data schema using TensorFlow Data Validation from the samples in the dataframe.
- Detect anomalies in new data using TensorFlow Data Validation.
- Generate a TFRecord feature specification using TensorFlow Transform from the data schema.
- Export a dataset and convert to TFRecords.
```
[Get Started with Dataflow](get_started_dataflow.ipynb)
```
The steps performed include:
- Offline preprocessing of data:
- Serially - w/o dataflow
- Parallel - with dataflow
- Upstream preprocessing of data:
- tabular data
- image data
```
### E2E Stage Example
[Stage 1: Data Management](mlops_data_management.ipynb)
```
The steps performed include:
- Explore and visualize the data.
- Create a Vertex AI `Dataset` resource from `BigQuery` table -- for AutoML training.
- Extract a copy of the dataset to a CSV file in Cloud Storage.
- Create a Vertex AI `Dataset` resource from CSV files -- alternative for AutoML training.
- Read a sample of the `BigQuery` dataset into a dataframe.
- Generate statistics and data schema using TensorFlow Data Validation from the samples in the dataframe.
- Generate a TFRecord feature specification using TensorFlow Data Validation from the data schema.
- Preprocess a portion of the BigQuery data using `Dataflow` -- for custom training.
```
@@ -8,7 +8,7 @@
},
"outputs": [],
"source": [
"# Copyright 2020 Google LLC\n",
"# Copyright 2021 Google LLC\n",
"#\n",
"# Licensed under the Apache License, Version 2.0 (the \"License\");\n",
"# you may not use this file except in compliance with the License.\n",
@@ -33,13 +33,13 @@
"\n",
"<table align=\"left\">\n",
" <td>\n",
" <a href=\"https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/master/notebooks/official/automl/ml_ops_stage1/get_started_bq_datasets.ipynb\">\n",
" <a href=\"https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/ml_ops/stage1/get_started_bq_datasets.ipynb\">\n",
" <img src=\"https://cloud.google.com/ml-engine/images/github-logo-32px.png\" alt=\"GitHub logo\">\n",
" View on GitHub\n",
" </a>\n",
" </td>\n",
" <td>\n",
" <a href=\"https://console.cloud.google.com/ai/platform/notebooks/deploy-notebook?download_url=https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/master/notebooks/official/automl/ml_ops_stage1/get_started_bq_datasets.ipynb\">\n",
" <a href=\"https://console.cloud.google.com/ai/platform/notebooks/deploy-notebook?download_url=https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/ml_ops/stage1/get_started_bq_datasets.ipynb\">\n",
" Open in Google Cloud Notebooks\n",
" </a>\n",
" </td>\n",
@@ -82,12 +82,12 @@
"\n",
"This tutorial uses the following Google Cloud ML services:\n",
"\n",
"- `Vertex Datasets`\n",
"- `Vertex AI Datasets`\n",
"- `BigQuery Datasets`\n",
"\n",
"The steps performed include:\n",
"\n",
"- Create a Vertex `Dataset` resource from `BigQuery` table -- compatible for `AutoML` training.\n",
"- Create a Vertex AI `Dataset` resource from `BigQuery` table -- compatible for `AutoML` training.\n",
"- Extract a copy of the dataset from `BigQuery` to a CSV file in Cloud Storage -- compatible for `AutoML` or custom training.\n",
"- Select rows from a `BigQuery` dataset into a `pandas` dataframe -- compatible for custom training.\n",
"- Select rows from a `BigQuery` dataset into a `tf.data.Dataset` -- compatible for custom training `TensorFlow` models.\n",
@@ -107,7 +107,7 @@
"When doing E2E MLOps on Google Cloud, the following best practices with structured (tabular) data in BigQuery:\n",
"\n",
"- For AutoML training:\n",
" - Create a managed dataset with Vertex `TabularDataset`.\n",
" - Create a managed dataset with Vertex AI `TabularDataset`.\n",
" - Use the BigQuery table as the input to the dataset.\n",
" - Specify columns and columns transformations when running the AutoML training pipeline job.\n",
"\n",
@@ -164,11 +164,13 @@
" ! pip3 install -U tensorflow-transform==1.2 $USER_FLAG\n",
" ! pip3 install -U tensorflow-io==0.18 $USER_FLAG\n",
" ! pip3 install --upgrade google-cloud-aiplatform[tensorboard] $USER_FLAG\n",
" ! pip3 install --upgrade google-cloud-pipeline-components $USER_FLAG\n",
" ! pip3 install --upgrade google-cloud-bigquery $USER_FLAG\n",
" ! pip3 install --upgrade google-cloud-logging $USER_FLAG\n",
" ! pip3 install --upgrade apache-beam[gcp] $USER_FLAG\n",
" ! pip3 install --upgrade pyarrow $USER_FLAG\n",
" ! pip3 install --upgrade cloudml-hypertune $USER_FLAG"
" ! pip3 install --upgrade cloudml-hypertune $USER_FLAG\n",
" ! pip3 install --upgrade kfp $USER_FLAG"
]
},
{
@@ -285,7 +287,7 @@
"\n",
"You may not use a multi-regional bucket for training with Vertex AI. Not all regions provide support for all Vertex AI services.\n",
"\n",
"Learn more about [Vertex AI regions](https://cloud.google.com/vertex-ai/docs/general/locations)"
"Learn more about [Vertex AI regions](https://cloud.google.com/vertex-ai/docs/general/locations)."
]
},
{
@@ -501,9 +503,9 @@
"id": "init_aip:mbsdk,region"
},
"source": [
"### Initialize Vertex SDK for Python\n",
"### Initialize Vertex AI SDK for Python\n",
"\n",
"Initialize the Vertex SDK for Python for your project and corresponding bucket."
"Initialize the Vertex AI SDK for Python for your project and corresponding bucket."
]
},
{
@@ -1188,6 +1190,7 @@
" # Delete the endpoint using the Vertex endpoint object\n",
" try:\n",
" if \"endpoint\" in globals():\n",
" endpoint.undeploy_all()\n",
" endpoint.delete()\n",
" except Exception as e:\n",
" print(e)\n",
@@ -33,13 +33,13 @@
"\n",
"<table align=\"left\">\n",
" <td>\n",
" <a href=\"https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/master/notebooks/official/automl/ml_ops_stage1/get_started_dataflow.ipynb\">\n",
" <a href=\"https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/ml_ops/stage1/get_started_dataflow.ipynb\">\n",
" <img src=\"https://cloud.google.com/ml-engine/images/github-logo-32px.png\" alt=\"GitHub logo\">\n",
" View on GitHub\n",
" </a>\n",
" </td>\n",
" <td>\n",
" <a href=\"https://console.cloud.google.com/ai/platform/notebooks/deploy-notebook?download_url=https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/master/notebooks/official/automl/ml_ops_stage1/get_started_dataflow.ipynb\">\n",
" <a href=\"https://console.cloud.google.com/ai/platform/notebooks/deploy-notebook?download_url=https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/ml_ops/stage1/get_started_dataflow.ipynb\">\n",
" Open in Google Cloud Notebooks\n",
" </a>\n",
" </td>\n",
@@ -157,11 +157,13 @@
" ! pip3 install -U tensorflow-transform==1.2 $USER_FLAG\n",
" ! pip3 install -U tensorflow-io==0.18 $USER_FLAG\n",
" ! pip3 install --upgrade google-cloud-aiplatform[tensorboard] $USER_FLAG\n",
" ! pip3 install --upgrade google-cloud-pipeline-components $USER_FLAG\n",
" ! pip3 install --upgrade google-cloud-bigquery $USER_FLAG\n",
" ! pip3 install --upgrade google-cloud-logging $USER_FLAG\n",
" ! pip3 install --upgrade apache-beam[gcp] $USER_FLAG\n",
" ! pip3 install --upgrade pyarrow $USER_FLAG\n",
" ! pip3 install --upgrade cloudml-hypertune $USER_FLAG"
" ! pip3 install --upgrade cloudml-hypertune $USER_FLAG\n",
" ! pip3 install --upgrade kfp $USER_FLAG"
]
},
{
@@ -258,7 +260,7 @@
"\n",
"You may not use a multi-regional bucket for training with Vertex AI. Not all regions provide support for all Vertex AI services.\n",
"\n",
"Learn more about [Vertex AI regions](https://cloud.google.com/vertex-ai/docs/general/locations)"
"Learn more about [Vertex AI regions](https://cloud.google.com/vertex-ai/docs/general/locations)."
]
},
{
@@ -474,7 +476,7 @@
"id": "import_numpy"
},
"source": [
"#### Import pandas\n",
"#### Import numpy\n",
"\n",
"Import the numpy package into your Python environment."
]
@@ -540,9 +542,9 @@
"id": "init_aip:mbsdk,region"
},
"source": [
"### Initialize Vertex SDK for Python\n",
"### Initialize Vertex AI SDK for Python\n",
"\n",
"Initialize the Vertex SDK for Python for your project and corresponding bucket."
"Initialize the Vertex AI SDK for Python for your project and corresponding bucket."
]
},
{
@@ -674,6 +676,31 @@
"dataframe[\"station_number\"] = pd.to_numeric(dataframe[\"station_number\"])"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "bqml_create_dataset"
},
"source": [
"### Create BQ dataset resource\n",
"\n",
"First, you create an empty dataset resource in your project."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "bqml_create_dataset"
},
"outputs": [],
"source": [
"BQ_MY_DATASET = 'samples'\n",
"BQ_MY_TABLE = 'gsod'\n",
"! bq --location=US mk -d \\\n",
"$PROJECT_ID:$BQ_MY_DATASET"
]
},
{
"cell_type": "code",
"execution_count": null,
@@ -1201,6 +1228,7 @@
" # Delete the endpoint using the Vertex endpoint object\n",
" try:\n",
" if \"endpoint\" in globals():\n",
" endpoint.undeploy_all()\n",
" endpoint.delete()\n",
" except Exception as e:\n",
" print(e)\n",

Some files were not shown because too many files have changed in this diff Show More