* Add a colab to show how to integrate the training job with Dask.
* Reformat the notebook xgboost_data_parallel_training_on_cpu_using_dask
* Add the code owner of the sample training/xgboost_data_parallel_training_on_cpu_using_dask.ipynb
* Changed the project id to [your-project-id].
* Fixed the issue for Non colab.
* Adding the sample of converting the Vertex Vizier SDK with Open source Vizier.
* Add the owner for conversions_vertex_vizier_and_open_source_vizier.ipynb
* Addressed the comments in the xgboost_data_parallel_training_on_cpu_using_dask
* Fixed the format of xgboost_data_parallel_training_on_cpu_using_dask
* Addressed the comments in the training/xgboost_data_parallel_training_on_cpu_using_dask.ipynb
* Add explanation that Docker is not available on Colab.
* Add before docker command.
* add timeout in the worker to wait for the scheduler.
* Addressed the comments in the pr.
* Addressed the comments in the pr.
* added model evaluation component
* linter test cases
* linter test cases
* model_name param issues resolved
* linter test case
* import issues resloved
* linter test cases
* made review changes
* made review changes
* ran linter test
* made review changes
* made review changes
* made review changes
* linter test
* ran linter test
* made review changes
* ran linter test
* made review changes
* ran linter test
* linter test
* review changes
* ran linter test
* added The links for Colab, Github and Workbench
* added The links for Colab, Github and Workbench
* ran linter test
* added The links for Colab, Github and Workbench
* ran linter test
* added The links for Colab, Github and Workbench
* ran linter test
* notebook title changed
* ran linter test
* text changes and made review changes
* ran linter test
* made review changes
* linter test
* made review changes
* linter test
* content changes
* ran linter test
* textual corrections and links
* ran linter test
Co-authored-by: Ivan Cheung <ivans.mailbox@gmail.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
Co-authored-by: Soheila Zangeneh <49654056+soheilazangeneh@users.noreply.github.com>
* Add a colab to show how to integrate the training job with Dask.
* Reformat the notebook xgboost_data_parallel_training_on_cpu_using_dask
* Add the code owner of the sample training/xgboost_data_parallel_training_on_cpu_using_dask.ipynb
* Changed the project id to [your-project-id].
* Fixed the issue for Non colab.
* Addressed the comments in the xgboost_data_parallel_training_on_cpu_using_dask
* Fixed the format of xgboost_data_parallel_training_on_cpu_using_dask
* Addressed the comments in the training/xgboost_data_parallel_training_on_cpu_using_dask.ipynb
* Add explanation that Docker is not available on Colab.
* Add before docker command.
* add timeout in the worker to wait for the scheduler.
* cleans up the notebook,replaces docker with cloud build, textual edits still in progress
* cleans up the notebook
* ran linter test
* changes tf train/serve version to 2.9
* ran linter test
* adds '=' to fix a typo
* ran linter test
* updates the opencv installation package
* ran linter test
* updates the opencv installation dependencies
* ran linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* feat: add autogen index
* fix: missed the REAME
* fix: update autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* fix: incorrect linking for index
* feat: add autogen index
* fix: missed the REAME
* fix: update autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* fix: missed the REAME
* fix: update autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* fix: missed the REAME
* fix: update autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* fix: missed the REAME
* fix: update autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* fix: missed the REAME
* fix: update autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* fix: missed the REAME
* fix: update autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* fix: missed the REAME
* fix: update autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* fix: missed the REAME
* fix: update autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* fix: missed the REAME
* fix: update autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* fix: missed the REAME
* fix: update autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* fix: missed the REAME
* fix: update autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* fix: missed the REAME
* fix: update autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* fix: missed the REAME
* fix: update autogen index
* feat: add autogen index
* feat: add autogen index
* feat: add autogen index
* fix: objective conformance
* fix: objective conformance
* fix: objective conformance
* fix: objective conformance
* fix: objective conformance
* fix: bad links and objective
* fix: bad links and objective
* fix: bad links and objective
* fix: bad links and objective
* fix: bad links and objective
* fix: bad links and objective
* fix: bad links and objective
* fix: bad links and objective
* fix: objective conformance
* fix: objective conformance
* fix: objective conformance
* fix: objective conformance
* fix: objective conformance
* fix: bad links and objective
* fix: bad links and objective
* fix: bad links and objective
* fix: bad links and objective
* fix: bad links and objective
* fix: bad links and objective
* fix: bad links and objective
* fix: objective conformance
* fix: objective conformance
* fix: objective conformance
* fix: objective conformance
* fix: objective conformance
* fix: bad links and objective
* fix: bad links and objective
* fix: bad links and objective
* fix: bad links and objective
* fix: bad links and objective
* fix: bad links and objective
* fix: objective conformance
* fix: objective conformance
* fix: objective conformance
* fix: objective conformance
* fix: objective conformance
* fix: bad links and objective
* fix: bad links and objective
* fix: bad links and objective
* fix: bad links and objective
* fix: bad links and objective
* feat: added new notebook for PyTorch distributed training on reduction server
* Fixed linter errors
* Addressed review comments
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* resloved tf library issues
* ran linter test
* made review changes
* ran linter test
* made review changes
* ran linter
* made review changes
* linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* Model evaluation for text classification using Automl
* Removed unused library
* Ran Linter Test
* Commented text
* Commented Text
* Ran Linter test
* Made changes suggested in review
* Removed unused import
* Ran Linter Test
* Removed dataflow parameters as mentioned in the review
* Ran Linter Test
* Made review changes and changed notebook links in the beginning of the notebook
* Ran Linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* Add notebook for Wide & Deep on Vertex Pipelines
* Address comments
* Clear output
* Fix deletion logic
* Run linter
* Rename custom_job and hpt_job to pipeline_job
* Use Vertex SDK for deletion instead of gcloud
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* Added notebook for TabNet on Vertex Pipelines
* Added notebook for TabNet on Vertex Pipelines
* Added notebook for TabNet on Vertex Pipelines
* Run linter
* Ran linter
* Ran linter again
* Addressed comments
* Addressed more comments
* Remove parameter definitions, reference documentation, add details under CustomJob/HPT job sections
* Address more comments
* Fix tests
* Bug fix
* Rework TabNet HPT job section
* Addressed more comments
* Update documentation links
* Fix tests
* Fix tests
* Linter and some changes
* Address more comments
* Update CustomJob description to align with documentation
* Use bank-marketing dataset instead of safedriver
* Rename variables
* Rename variables
* Update gcs location to official one
* Changes to make notebook run with updated SDK
* Bug fix
* Rewording
* Fix workbench link
* Update notebook format to align with template
* Fix deletion logic
* Run linter
* Rename custom_job and hpt_job to pipeline_job
* Use Vertex SDK for deletion instead of gcloud
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: Ivan Cheung <ivans.mailbox@gmail.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* Made changes in accordance with Notebook template
* Ran Linter Test
* Attached UUID to BQ_DATSET and use bq to delete dataset
* Removed unused imports
* Ran Linter Test
* Fixed BQ_DATASET
* Fixed BQ_DATASET
* Ran linter test
* Made review changes
* Removed unused library
* Ran linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* add The links for Colab, Github and Workbench
* added The links for Colab, Github and Workbench
* ran linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* modified notebook telecom-subscriber-churn-prediction.ipynb
* ran linter
* changes suggested by andrew are done
* ran linter
* changes suggested by andrew done
* ran linter
* resolved error
* ran linter
* fixed a monor bug
* ran linter
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* added model evaluation component
* linter test cases
* linter test cases
* model_name param issues resolved
* linter test case
* import issues resloved
* linter test cases
* made review changes
* made review changes
* ran linter test
* made review changes
* made review changes
* made review changes
* linter test
* ran linter test
* made review changes
* ran linter test
* made review changes
* ran linter test
* linter test
* review changes
* ran linter test
* added The links for Colab, Github and Workbench
* added The links for Colab, Github and Workbench
* ran linter test
* added The links for Colab, Github and Workbench
* ran linter test
* added The links for Colab, Github and Workbench
* ran linter test
Co-authored-by: Ivan Cheung <ivans.mailbox@gmail.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* new notebook
* linter test passed
* linter test passed
* andy review
* linter test passed
* add codeowner
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* added model evaluation component
* linter test cases
* linter test cases
* model_name param issues resolved
* linter test case
* import issues resloved
* linter test cases
* made review changes
* made review changes
* ran linter test
* made review changes
* made review changes
* made review changes
* linter test
* ran linter test
* made review changes
* ran linter test
* made review changes
* ran linter test
* linter test
* review changes
* ran linter test
Co-authored-by: Ivan Cheung <ivans.mailbox@gmail.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* add new notebook
* linter test passed
* clean text
* linter test passed
* add code owner new model registry notebook
* add more description
* linter test passed
* align with new template
* linter test passed
* andy reviews
* linter test passed
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* Added changes in notebook
* Ran linter test
* Replaced Timestamp with UUID; Added 'delete-bucket' in cleanup; Added condition for repo creation and few other minor changes
* Ran Linter Test
* Made some minor changes to install packages
* Made minor changes to fix linter failed tests
* ran linter test
* Made some minor changes
* ran linter test
* addresses the review comments: fixes container build steps, license year, updates according to the template
* ran linter test
* removes beta from gcloud to avoid timeouts
* ran linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: SamyuktaDR <samyukta.dontireddy@springml.com>
Co-authored-by: SamyuktaDR <45586340+SamyuktaDR@users.noreply.github.com>
Co-authored-by: Krishna Chaitanya Movva <krishr2d2@gmail.com>
Co-authored-by: krishr2d2 <krishna.movva@springml.com>
* resolves the shell-output issue + updates the structure based on the template
* ran linter test
* fixes issues from review: textual updates, delete_bucket=False
* ran linter test
Co-authored-by: Ivan Cheung <ivans.mailbox@gmail.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* modified notebook according to notebook_template.
* ran linter
* Added create dataset step
* ran linter
* replaced hardcoded dataset name with a variable
* ran linter
* changes suggested by nadrew done
* ran linter
* followed prolong lin comments, data cleaning now done in bigquery
* ran linter
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
Co-authored-by: Ivan Cheung <ivans.mailbox@gmail.com>
* Add bqml-vertexai-model-registry notebook
* Run linter
* Add notebook to CODEOWNERS file
* Update the links
* Ran linter again
* Rename bigquey-ml folder to model-registry
* Add bigquery-ml folder
* Moved the notebook
* Deleted folder
* Resolve comments
* Use UUID
* Remove using existing endpoint
* Remove try statement
* Get model sample based on model's name
* Use job.result to check query job status
* Run linter
* Fix dataset not found error by adding region in bq client creation
* Revert changes
* Fix bq bugs
* Run linter
* Resolve comments
* Display dataframe
* Updated the codeowner file
* Updated the links and editted text
* Run linter
* Remove repeated resources
* Run linter
* Resolve comments
* Run linter
* Install pyarrow
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* Create a function to get notebook python version for execution
* Inject python version to yaml file
* Fix python version references
* Add python string to python version variable
* add python version extraction script (untested)
* Create a function to get notebook python version for execution
* Inject python version to yaml file
* Fix python version references
* Add python string to python version variable
* Remove one notebook condition
* Remove extra check and use python 3 as default version
* Use python3.9 as default version
* Update python version notebook parser
* Add python version to the notebook template
* Fix bug
* Update python version parser function
* Add python version to a notebook for testing
* Run linter
* Use regex in python version parser function
* Add new notebook for testing
* Use better variable name
* fix typo
* Use f string instead +
* Use python from env instead of using _PYTHON_VERSION
* Use simpler regex
* Add python version test notebook
* Fixed a mistake
* Updated notebook template with python version
* Fixed python version format
* Print log contents to stdout
* Remove failing notebook
* Run linter
* Remove extra steps in the printed log
* Edit comments
* Run linter
* Run linter
* Revert test changes
Co-authored-by: AG Sol <aarongabriel@google.com>
* new changes for sentiment analysis notebook
* new changes for sentiment analysis notebook
* changed back to year 2021 text
* changed back to year 2021 text
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* added file
* removed unnecessary imports
* ran linter
* removed extra batch prediction component
* ran linter
* moved file to official
* changed links to point to official
* ran linter
* Jason comments addressed
* ran linter
* comments addressed
* ran linter
* comments addresed
* ran linter
* added predictionschema instance schema files
* ran linter
* Followed Karen Lin's comments
* ran linter
* replaced old prebuilt containers with latest ones
* ran linter
Co-authored-by: Ivan Cheung <ivans.mailbox@gmail.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* Update the vizier codelab to replace the gapic library with new Vertex Vizier SDK.
* Added the [project_id] and [region] in the parameter field.
* Fixed the lint errors for vizier sample.
Co-authored-by: Ivan Cheung <ivans.mailbox@gmail.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* clean and update comparing_local_trained_models based on feedback
* linter test passed
* fix libraries
* linter test passed
* andy review fixes
* linter test passed
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* moves the sentiment_analysis notebook from community to official folder after making the updates
* removes unused modules
* ran linter test
* updates the dataset's GCS links and notebook links in the heading
* ran linter test
* fixes the typo(=)
* ran linter test
* removes wait() calls and IS_TESTING condition
* ran linter test
Co-authored-by: Ivan Cheung <ivans.mailbox@gmail.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* modified notebook
* ran linter
* tensorflow was used only for file reading.So replaced tensorflow with pandas
* ran linter
* made text changes
* ran linter
* latest andrew domments addressed
* ran linter
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* Replaced timestamp with UUID
* Ran Linter test
* Removed local kernel from metadata
* ran linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
Co-authored-by: Ivan Cheung <ivans.mailbox@gmail.com>
* add unfinished notebook on hpt using R
* clear output
* add working version of notebook
* finish R HPT notebook
* update CODEOWNERS
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* Add automl regression model eval first draft
* Remove extra file
* Pring evaluation results
* adds the automl-tabular-classification notebook in model_evaluation folder
* removes unnecessary imports
* adjusts the imports inside the pipeline
* adjusts the imports
* elaborates imports inside pipeline
* modified regression notebook
* renamed pipeline displayname to resolve error
* Add automl regression model eval first draft
* Remove extra file
* Pring evaluation results
* modified some text
* added suggested updates from review: remove dataflow params, add/change textual descriptions, add UUID
* removes the output from the notebooks
* removes the extra matplotlib import
* ran linter test
* addressed soheila's comments
* ran linter
* addresses the review comments
* ran linter test
* removes the artifacts comment
* ran linter test
* reviewed comments
* ran linter
* addresses review comments: textual updates, removes unnecessary parameters
* ran linter test
* addressed comments
* ran linter
* removed unwanted variables
* ran linter
* addresses the tech-writer's comments + updates the pipeline image with data-sampler task
* ran linter test
* Update text
* Move model eval folder to official
* Update CODEOWNERS
* Run linter
* Removed problem_type parameter
* Run linter
* addresses Andrew's review comments: textual updates and removes additional gcpc installation
* ran linter test
* comments addressed
* ran linter
* removed trailing comma on last parameter of trainingjob.run
* ran linter
Co-authored-by: krishr2d2 <krishna.movva@springml.com>
Co-authored-by: sudarshan-SpringML <sudarshan.c@springml.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* feat: notebook for custom text model batch prediction
* feat: notebook for custom text model batch prediction
Co-authored-by: gericdong <itseric@google.com>
* Add automl regression model eval first draft
* Remove extra file
* Pring evaluation results
* adds the automl-tabular-classification notebook in model_evaluation folder
* removes unnecessary imports
* adjusts the imports inside the pipeline
* adjusts the imports
* elaborates imports inside pipeline
* modified regression notebook
* renamed pipeline displayname to resolve error
* Add automl regression model eval first draft
* Remove extra file
* Pring evaluation results
* modified some text
* added suggested updates from review: remove dataflow params, add/change textual descriptions, add UUID
* removes the output from the notebooks
* removes the extra matplotlib import
* ran linter test
* addressed soheila's comments
* ran linter
* addresses the review comments
* ran linter test
* removes the artifacts comment
* ran linter test
* reviewed comments
* ran linter
* addresses review comments: textual updates, removes unnecessary parameters
* ran linter test
* addressed comments
* ran linter
* removed unwanted variables
* ran linter
* addresses the tech-writer's comments + updates the pipeline image with data-sampler task
* ran linter test
Co-authored-by: krishr2d2 <krishna.movva@springml.com>
Co-authored-by: sudarshan-SpringML <sudarshan.c@springml.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* move bqml-vertex notebook from community to official
* add to CODEOWNERS official
* fix errors for execution-test
* fix project_id line
* fix linting
* fixes re: comments from sarahcdugan
* fix links at top of notebook from community/ to official/
* added UUID to model name
* fix linting
* fix error in TIMESTAMP --> UUID
* fixing linting double space
* fixes re: ivanmkc comments
* fixed notebook after linting issues
* linting via cloud shell
* simplified run_bq_query function
* linting
Co-authored-by: Ivan Cheung <ivans.mailbox@gmail.com>
* Minor fixes for CPR Pytorch sample: Add missing test data, add auth info to readme, scrub private project and bucket names from config, tolerate missing config.json in unit tests.
* Minor fixes for CPR Pytorch sample: Add missing test data, add auth info to readme, scrub private project and bucket names from config, tolerate missing config.json in unit tests.
* Fix merge conflicts
* fix typo
* Point CPR links to main branch of SDK repo.
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* feat: add notebook for custom tabular batch predict
* feat: add notebook for custom tabular batch predict
* feat: add example for BQ input
* feat: add example for BQ input
* feat: add example for BQ input
* feat: add example for BQ input
* changed to andrew comments
* changes according to andrew comments
* changes according to andrew comments
* review changes
* review changes
* review changes
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* samples: Add a new sample for pre-built Pytorch deployments. It's
borrowed from the examples in community-content/pytorch_text_classification_using_vertex_sdk_and_gcloud.
* samples: Removed all training related stuff in the notebooks.
* samples: Fixed comments.
* samples: Updated readme.
* samples: Updated emails for Pytorch launch.
* Added condition to create Featurestore if it doesn't exist
* Ran Linter Test
* Made changes mentioned in review
* Ran Linter Test
* Attached uuid to featurestore_id to avoid error while creating featurestore with existing name
* Ran linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
less chance for an error and confusion in name clashing with the `datasets` pypi package also used in the notebook.
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* Create explore_data_in_bigquery_with_workbench.ipynb
Adding in notebook for exploratory data analysis as part of "Data to AI" effort. See this Colab for what this notebook looks like after it is run: https://colab.research.google.com/drive/1JeNeMtj2A_5P5vo9wxSkrwHQM5JQSoAu. Submitting it with outputs shown since a lot of this about interactive visualization, which can inspire folks to use/read the notebook beyond just the code.
* Update CODEOWNERS
Adding owner for forthcoming exploratory data analysis notebook
* Update CODEOWNERS
* Updating exploratory data analysis notebook with latest updates from linter/review
* Updated notebook formatting to try to pass format test
* Trying again to pass notebook formatting test
* Trying again to pass notebook formatting test
* Trying again to pass notebook formatting test
* Linted version of notebook & better project picker
* Uploading linted version from ivanmkc@
* Update CODEOWNERS with EDA notebook
* Updated notebook w/ Tech Writer edits, re-ran all
* 1-2 minor text updates, try to pass linter again
* Trying w/ updated linted file from ivanmkc@
Co-authored-by: Ivan Cheung <ivans.mailbox@gmail.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* new changes of build model notebook
* new changes of build model notebook
* linter test issues
* linter test issues
* review changes
* review changes
* review changes
* review changes
* review changes
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* added new cell for is_colab condition
* added new cell for is_colab condition
* changes andrew comments
* changes andrew comments
* review changes
* review changes
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* Made minor changes
* Ran linter test
* Made changes mentioned in the review
* Ran Linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* Changed prediction_output format from csv to jsonl to support generate_explanations
* ran linter test
* Made the changes as mentioned in the review
* Ran Linter Test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* Removed try except blocks from cleanup section
* Ran Linter test
* Made changes mentioned in review and removed globals
* Removed an unused variable
* Ran Linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* new notebook of classification beans
* new notebook of classification beans
* changes on andrew comments
* changes on andrew comments
* json file issues
* json file issue
* fixes issues from reviews: adds parameter descriptions, fixes clean up, textual updates and replaces gapic functionality
* ran linter test
* adds the missing machine-type parameter
* ran linter test
* retreives the metrics using dict method
* removes unused variables
* ran linter test
* replaces old code for resource-name with new one
* ran linter test
* updates fetching the resourceName from the training artifacts
* ran linter test
* adds wait method for endpoint deployment
* ran linter test
* removes the wait method
* ran linter test
* adds wait gcp resources component
* ran linter test
* updates colab link, removes wait component, sets force to true in delete endpoint step
* ran linter test
* adds endpoint.wait() method
* ran linter test
* moves model deletion down the endpoint deletion and removes endpoint.wait() method
* ran linter test
Co-authored-by: Krishna Chaitanya Movva <krishr2d2@gmail.com>
Co-authored-by: Krishna Chaitanya Movva <krishna.movva@springml.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* Add bqml-vertexai-model-registry notebook
* Run linter
* Add notebook to CODEOWNERS file
* Update the links
* Ran linter again
* Rename bigquey-ml folder to model-registry
* Add bigquery-ml folder
* Moved the notebook
* Deleted folder
* Resolve comments
* Use UUID
* Remove using existing endpoint
* Remove try statement
* Get model sample based on model's name
* Use job.result to check query job status
* Run linter
* Fix dataset not found error by adding region in bq client creation
* Revert changes
* Fix bq bugs
* Run linter
* Resolve comments
* Display dataframe
* Updated the codeowner file
* fix: notebook template tuning
* fix: notebook template tuning
* fix: possible confusion on when to wait for the email notification
* fix: possible confusion on when to wait for the email notification
* adds the updated predictive-maintenance (managed)notebook from community to official folder
* ran linter test
* resubmitting during phase2
* ran linter test
* addresses the review comments: updates based on the new template, sets delete_bucket to False
* ran linter test
* replaces timestamp with uuid
* ran linter test
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* fixes tensorflow version, replaces timestamp with uuid, adds service-account and minor textual changes
* removes second defnition of random library
* ran linter test
* uncomments the user flag and updates the installation command
* ran linter test
* removes the extra backslash
* ran linter test
* updates the installation step to fix long running compatibility checks
* ran linter test
* updates the METADATA path during installation steps
* ran linter test
* adds google-api-core version in the installation
* ran linter test
* removes METADATA step during installation
* ran linter test
* updates google api-core & auth versions
* ran linter test
* fixes tensorflow version, replaces timestamp with uuid, adds service-account and minor textual changes
* removes second defnition of random library
* ran linter test
* uncomments the user flag and updates the installation command
* ran linter test
* removes the extra backslash
* ran linter test
* updates the installation step to fix long running compatibility checks
* ran linter test
* updates the METADATA path during installation steps
* ran linter test
* adds google-api-core version in the installation
* ran linter test
* removes METADATA step during installation
* ran linter test
* updates google api-core & auth versions
* ran linter test
* fixes the issues from review: cell descriptions, parameter definitions, tense changes, 3rd person --> 2nd person, list model after pipeline run
* ran linter test
* removes the METADATA hack and updates the installations
* ran linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* changed based on andrew review comments
* changed based on andrew review comments
* import library issues
* import library issues
* import issues
* import issues
* modified notebook
* modified notebook
* added new notebook
* added new notebook
* new auto_ml_text_classifiation
* new auto_ml_text_classifiation
* new automl text classification
* linter test
* linter test
* changes on andrew comments
* changes on andrew comments
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* fixing links to open notebook - main and images
* linter test changes
* fixes the papermill execution error(hard-coded bucket link was the cause)
* ran linter test
* adds minor textual changes
* ran linter test
* fixes issues from review: future tense, copyright year, section placement, latest sdk methods, new updates from the template
* ran linter test
Co-authored-by: Manuel Amunategui <manuel.amunategui@springml.com>
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* made changes
* made changes
* ran linter test
* made changes
* ran linter
* made changes
* ran linter
* changes suggested by andrew done
* ran linter
* replaced timestamp with uuid
* ran linter
* changed bucket creation command according to template
* ran linter
* changed text in overview
* changed region cell from markdown to code
* made changes
* replaced dataset from constant to a variable
* replaced constant dataset_id with a variable
* ran linter
* changed suggested by andrew done
* ran linter
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* adds the missing delete_bucket variable, adds steps to configure SERVICE_ACCOUNT
* removes unnecessary random import
* ran linter test
* removes the src folder dependency to run on Colab, adds the pipeline.wait step, updates the cleanup steps
* ran linter test
* fixed issues from review: section posistions, tense changes, section descriptions, template updates
* ran linter test
* Added matching engine notebook official
* Ran linter
* Added matching engine to .cloud-build/test_notebook_vm.txt
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* new notebook of custom image classification
* new notebook of custome image classification
* andrew commented changes
* andrew commented changes
* andrew commented changes
* andrew commented changes
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* fixes exception(reg-test), replaces timestamp with uuid, minor changes
* ran linter test
* resolved review comments: license year, Vertex AI SDK, dataset after objective and future tense
* ran linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* deleted file in community and added file in official folder
* renamed file
* ran linter test
* renamed file
* ran linter
* made changes
* ran linter test
* made changes
* ran linter test
* made changes
* ran linter test
* made changes
* ran linter
* made change
* ran linter test
* changes suggested by andrew done
* ran linter
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* modified notebook according to notebook_template.
* ran linter
* Added create dataset step
* ran linter
* replaced hardcoded dataset name with a variable
* ran linter
* changes suggested by nadrew done
* ran linter
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* done UUID changes
* ran lintertest
* made changes in cleanup section
* ran lintertest
* done UUID changes
* ran lintertest
* made changes in cleanup section
* ran lintertest
* made changes in cleanup section
* Ran linter test
* made UUID changes
* RAN linter test
* Made Some minor Chanages notebook
* Ran Linter Test
* small changes made
* Ran Linter Test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* modified notebook according to template, tensorflow library is used only for file opening so instead of tf we used bucket.blob.download_as_string()
* ran linter
* all changes requested by andrew are done
* cleared all outputs
* making changes to run linter test
* making changes to run linter test
* ran linter
* removed region text in create bucket step
* ran linter
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* updates the configuring steps, replaces timestamp with uuid, expands the imports
* ran linter test
* separates the vertex-ai and bigquery initialization steps
* adds comment to cell_24
* adds blank line to cell_24:7:1
* adds blank line to cell_24:7:1
* ran linter test
* fixes aiplatform+bigquery installation compatibility issue
* ran linter test
* fixes installation dependencies
* ran linter test
* fixes the issues from the review: future tense, section positions, updates from the latest template
* ran linter test
* fixes the issues from the review: Code formatting, delete redundant cells, resource name changes
* ran linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* Created spark notebook with test code
* Implemented experimental code for poly_view
* implemented the table
* WIP-notebook
* moved experimental to tutorial
* delete experimental code and rename the notebook
* clear all outputs
* fix: delete outputs again
* fix: changed template to the newer one
* fix: modify link on workbench
* add creating a cluster
* Completed Before you begin part
* Change execution sequence
* completed write back process
* change order that switching kernel goes top
* modify pie chart to bar chart
* Completed write up part
* WIP: adding description and comments.
* WIP: delete outputs
* fix: nbqa done
* Completed the first draft
* Delete %%time from cells
* Apply changes as per the code review from Brad except SparkSql
* delete outputs
* Change SparkSQL to Spark API
* change label to xlabel
* fix: description in Dataset
* fix: change BUCKET_NAME to DATASET_NAME, link for the region, and add descriptions and examples for frequency table
* fix: move normalize_name to top of the cell
* fix: refactor udf functions and descriptions
* fix: add link for udf
* fix: description in Dataset
* fix: grammer
* fix: add declared in the sentence
* fix: small changes on grammar
* fix: delete string
* fix: change UserDefinedFunction to udf
* fix: as per TW's code review
* fix: reorder REGION and TIMESTAMP under Creating a GCS bucket
* fix: lint
* chore: add bmiro@ as a codeowner of this doc
* fix: change the variable to fix a bug
* fix: as per TW's second review
* fix: add installation part to pass the ci test
* fix: url for links to main
* fix: delete disabling API since it doesn't affect to the pricing
* fix: add conditions for CI test
* fix: changed jar for testing
* fix: add gcs connector
* fix: change writing method to direct
* fix: delete gcs connector
* fix: specify java folder
* fix: change java_home location
* fix: change unzip instruction
* fix: delete mono_ranking_avg_bytes from testing env
* fix: delete frequency_table from testing env
* fix: delete GCS bucket part
* fix: as per Brad's review
* fix: lint
* fix: add version
* fix: change comment
* fix: add package due to switching the kernel
* fix: delete dataproc cluster command
* fix: change link
* fix: change link
* fix: revert cluster deletion command
* fix: change parenthesis to encoded character
* fix: change the name of the notebook
* fix: change timestamp to UUID
* fix: change the link and add description
* fix: change metadata
* replaces timestamp with uuid #create *task #tag1 replace the TIMESTAMP with uuid in other official notebooks
* ran linter test
* updates the uuid code
* fixes the comment style highlighted through linter-test
* ran linter test
* adds length argument to uuid function defaulted to 8
* ran linter test
* Start a new branch for TabNet tutorial.
* format lint
* Clean version Created using Colaboratory
* Remove unused import
* Remove unused import
* Created using Colaboratory
* add import
* Add visualization for TabNet
* add gcs
* run format
* reformat
* reformat
* Rmove the - file
* run linter
* run linter
* Update the objective and data section
* Update the link.
* Update data description.
* Update data description.
Co-authored-by: Long Le <longtle@google.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* Added notebook demonstrating Tensorboard Custom Training with custom container.
* Added notebook demonstrating Tensorboard Custom Training with custom container.
* update codeowners file
* call Vertex API instead of gapic API
* resolve comments for custom container
* resolve comments and format
* resolve comments
* using --quiet for delete doctor repository
* address more comments
Co-authored-by: gericdong <itseric@google.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* Added notebook demonstrating Tensorboard Custom Training with prebuilt container
* Added notebook demonstrating Tensorboard Custom Training with prebuilt container
* small fix
* address comments
* format
* update project id to be [your-project-id], and populate tensorboard resource name automatically
* Added notebook demonstrating Tensorboard Custom Training with prebuilt container
* Added notebook demonstrating Tensorboard Custom Training with prebuilt container
* small fix
* address comments
* format
* fix typo for service account
* use vertex api instead of gapic api
* address comments
* minor fix
* minor fix for link
* minor fix
* resolve more comments
* a minor fix for comment
* format the notebook
* resolve comments
* address more comments
Co-authored-by: gericdong <itseric@google.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* Added notebook
* Made changes in installing packages code cell
* Ran linter test
* fixes the installation issues and updates some textual content
* fixes the # formatting for comments
* ran linter test
* adds pyarrow to the packages
* ran linter test
* replaces timestamp with uuid
* ran linter test
* updates the uuid code
* fixes the comment style highlighted through linter-test
* ran linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: krishr2d2 <krishna.movva@springml.com>
* notebook refresh from vertex ai sdk project
* linter test
* notebook refresh from vertex ai sdk project with trainer folder
* linter test
* add pyarrow
* modified notebook
* ran linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@gmail.com>
Co-authored-by: sudarshan-SpringML <sudarshan.c@springml.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* modified notebook
* small changes done
* modified notebook and moved notebook to official folder
* ran linter test
* resolved comments
* ran linter test
* sentence case heading added for some more text
* ran linter test
* made changes
* ran linter test
* made changes
* ran linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* multi_node_ddp_gloo_vertex_training_with_custom_container refresh and related trainer folder
* linter test
* various fixes and colab update
* linter test
* modified notebook
* modified notebook
* ran linter test
* Update multi_node_ddp_gloo_vertex_training_with_custom_container.ipynb
Co-authored-by: sudarshan-SpringML <sudarshan.c@springml.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@gmail.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* moving REGION up
* moving REGION up and csv file name
* fix: changed bucket URL to console
* removing TODOs from Tabnet notebook
* adding notebook and editing CODEOWNERS file
* fixing links
* adding to community because of test issue
* removing CODEOWNERS
* reverting CODEOWNERS
* linting?
* adding fixes
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
* feat: import automl tabular model
* feat: import automl tabular model
* feat: HPT for non-TF
* feat: HPT for non-TF
* fix: split guidelines from template
* fix: split guidelines from template
* fix: split guidelines from template
* upgrade: updates for new release
* upgrade: updates for new release
* update: tune title
* update: tune title
* feat: auto-discover
* fix: title
* fix: title
* fix: title
* fix: title
* fix: title
* fix: title
* feat: autodiscover
* feat: autodiscover
* fix: title
* fix: title
* feat: autodiscover
* fix: title
* fix: title
* feat: autodiscover
* fix: title
* fix: title
* feat: autodiscover
* fix: title
* fix: title
* fix: title
* fix: title
* feat: autodiscover
* fix: pinning
* fix: pinning
* adding new notebook on BQML online pred via Model Registry
* minor changes
* fixes to linting
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* feat: import automl tabular model
* feat: import automl tabular model
* feat: HPT for non-TF
* feat: HPT for non-TF
* fix: split guidelines from template
* fix: split guidelines from template
* fix: split guidelines from template
* upgrade: updates for new release
* upgrade: updates for new release
* update: tune title
* update: tune title
* feat: auto-discover
* fix: title
* fix: title
* fix: title
* fix: title
* fix: title
* fix: title
* feat: autodiscover
* feat: autodiscover
* fix: title
* fix: title
* feat: autodiscover
* fix: title
* fix: title
* feat: autodiscover
* fix: title
* fix: title
* feat: autodiscover
* fix: title
* fix: title
* fix: title
* fix: title
* feat: autodiscover
* feat: import automl tabular model
* feat: import automl tabular model
* feat: HPT for non-TF
* feat: HPT for non-TF
* fix: split guidelines from template
* fix: split guidelines from template
* fix: split guidelines from template
* upgrade: updates for new release
* upgrade: updates for new release
* update: tune title
* update: tune title
* feat: auto-discover
* fix: title
* fix: title
* fix: title
* fix: title
* fix: title
* fix: title
* feat: autodiscover
* feat: autodiscover
* fix: title
* fix: title
* feat: autodiscover
* fix: title
* fix: title
* feat: autodiscover
* fix: title
* fix: title
* feat: autodiscover
* fix: title
* fix: title
* fix: title
* fix: title
* feat: import automl tabular model
* feat: import automl tabular model
* feat: HPT for non-TF
* feat: HPT for non-TF
* fix: split guidelines from template
* fix: split guidelines from template
* fix: split guidelines from template
* upgrade: updates for new release
* upgrade: updates for new release
* update: tune title
* update: tune title
* feat: auto-discover
* fix: title
* fix: title
* fix: title
* fix: title
* fix: title
* fix: title
* feat: autodiscover
* feat: autodiscover
* fix: title
* fix: title
* feat: autodiscover
* fix: title
* fix: title
* feat: autodiscover
* fix: title
* fix: title
* feat: autodiscover
* fix: title
* fix: title
* feat: import automl tabular model
* feat: import automl tabular model
* feat: HPT for non-TF
* feat: HPT for non-TF
* fix: split guidelines from template
* fix: split guidelines from template
* fix: split guidelines from template
* upgrade: updates for new release
* upgrade: updates for new release
* update: tune title
* update: tune title
* feat: auto-discover
* fix: title
* fix: title
* fix: title
* fix: title
* fix: title
* fix: title
* feat: autodiscover
* feat: autodiscover
* fix: title
* fix: title
* feat: autodiscover
* fix: title
* fix: title
* feat: autodiscover
* fix: title
* fix: title
* feat: autodiscover
* feat: import automl tabular model
* feat: import automl tabular model
* feat: HPT for non-TF
* feat: HPT for non-TF
* fix: split guidelines from template
* fix: split guidelines from template
* fix: split guidelines from template
* upgrade: updates for new release
* upgrade: updates for new release
* update: tune title
* update: tune title
* feat: auto-discover
* fix: title
* fix: title
* fix: title
* fix: title
* fix: title
* fix: title
* feat: autodiscover
* feat: autodiscover
* fix: title
* fix: title
* feat: autodiscover
* fix: title
* fix: title
* feat: autodiscover
* fix: title
* fix: title
* feat: import automl tabular model
* feat: import automl tabular model
* feat: HPT for non-TF
* feat: HPT for non-TF
* fix: split guidelines from template
* fix: split guidelines from template
* fix: split guidelines from template
* upgrade: updates for new release
* upgrade: updates for new release
* update: tune title
* update: tune title
* feat: auto-discover
* fix: title
* fix: title
* fix: title
* fix: title
* fix: title
* fix: title
* feat: autodiscover
* feat: autodiscover
* fix: title
* fix: title
* feat: autodiscover
* fix: title
* fix: title
* feat: import automl tabular model
* feat: import automl tabular model
* feat: HPT for non-TF
* feat: HPT for non-TF
* fix: split guidelines from template
* fix: split guidelines from template
* fix: split guidelines from template
* upgrade: updates for new release
* upgrade: updates for new release
* update: tune title
* update: tune title
* feat: auto-discover
* fix: title
* fix: title
* fix: title
* fix: title
* fix: title
* fix: title
* feat: autodiscover
* feat: autodiscover
* fix: title
* fix: title
* feat: autodiscover
* feat: import automl tabular model
* feat: import automl tabular model
* feat: HPT for non-TF
* feat: HPT for non-TF
* fix: split guidelines from template
* fix: split guidelines from template
* fix: split guidelines from template
* upgrade: updates for new release
* upgrade: updates for new release
* update: tune title
* update: tune title
* feat: auto-discover
* fix: title
* fix: title
* fix: title
* fix: title
* fix: title
* fix: title
* feat: autodiscover
* feat: autodiscover
* fix: title
* fix: title
* feat: import automl tabular model
* feat: import automl tabular model
* feat: HPT for non-TF
* feat: HPT for non-TF
* fix: split guidelines from template
* fix: split guidelines from template
* fix: split guidelines from template
* upgrade: updates for new release
* upgrade: updates for new release
* update: tune title
* update: tune title
* feat: auto-discover
* fix: title
* fix: title
* fix: title
* fix: title
* fix: title
* fix: title
* feat: autodiscover
* feat: autodiscover
* feat: import automl tabular model
* feat: import automl tabular model
* feat: HPT for non-TF
* feat: HPT for non-TF
* fix: split guidelines from template
* fix: split guidelines from template
* fix: split guidelines from template
* upgrade: updates for new release
* upgrade: updates for new release
* update: tune title
* update: tune title
* feat: auto-discover
* fix: title
* fix: title
* fix: title
* fix: title
* fix: title
* fix: title
* feat: import automl tabular model
* feat: import automl tabular model
* feat: HPT for non-TF
* feat: HPT for non-TF
* fix: split guidelines from template
* fix: split guidelines from template
* fix: split guidelines from template
* upgrade: updates for new release
* upgrade: updates for new release
* update: tune title
* update: tune title
* feat: auto-discover
* fix: title
* fix: title
* fix: title
* fix: title
* fix: title
* fix: title
* feat: import automl tabular model
* feat: import automl tabular model
* feat: HPT for non-TF
* feat: HPT for non-TF
* fix: split guidelines from template
* fix: split guidelines from template
* fix: split guidelines from template
* upgrade: updates for new release
* upgrade: updates for new release
* update: tune title
* update: tune title
* feat: auto-discover
* fix: title
* fix: title
* feat: import automl tabular model
* feat: import automl tabular model
* feat: HPT for non-TF
* feat: HPT for non-TF
* fix: split guidelines from template
* fix: split guidelines from template
* fix: split guidelines from template
* upgrade: updates for new release
* upgrade: updates for new release
* update: tune title
* update: tune title
* feat: auto-discover
* fix: title
* fix: title
* feat: import automl tabular model
* feat: import automl tabular model
* feat: HPT for non-TF
* feat: HPT for non-TF
* fix: split guidelines from template
* fix: split guidelines from template
* fix: split guidelines from template
* upgrade: updates for new release
* upgrade: updates for new release
* update: tune title
* update: tune title
* feat: auto-discover
* feat: import automl tabular model
* feat: import automl tabular model
* feat: HPT for non-TF
* feat: HPT for non-TF
* fix: split guidelines from template
* fix: split guidelines from template
* fix: split guidelines from template
* upgrade: updates for new release
* upgrade: updates for new release
* update: tune title
* update: tune title
* feat: import automl tabular model
* feat: import automl tabular model
* feat: HPT for non-TF
* feat: HPT for non-TF
* fix: split guidelines from template
* fix: split guidelines from template
* fix: split guidelines from template
* upgrade: updates for new release
* upgrade: updates for new release
* New notebook to demonstrate how to enable TensorBoard Profiler
* Reformatted with Lint
* Changed service account handling and added a step to monitor job state
* Addressed review comments
* Addressed technical writerreview comments
* Addressed Ivan review comments
* Switched from GAPIC to Vertex SDK
* Removed an unused package
* Addressed review comments
* Add an example use case for custom prediction routines.
* Addressing some PR comments: reworded the readme in a few places, added a 'probe' command to build.py that sends a sample predict request, and pinned versions in requirements. Also fixed a bug where the artifacts_uri passed in during deployment on Vertex AI was not recognized as a directory.
* Autoformat code with black and fix a couple of typing errors.
* Addressing PR comments: Add deployment machine type to the config and add docstring to probe_prediction method.
* Update example to work with new LocalModel interface.
* Update example to work with new LocalModel interface.
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
* Added a service account injection
* Revert this
* Added service account injection
* Fixed cloud build file
* Added gcloud version debug info
* Fixed sa injection
* Removed test file
* Revert CODEOWNERS
* google_cloud_pipeline_components_bqml_pipeline_demand_forecasting notebook
* linter test to check with andy
* google_cloud_pipeline_components_bqml_pipeline_demand_forecasting notebook
* linter test to check with andy
* merge
* linter test minor fails. check with andy
* add code owner
* minor changes
* remove components
* linter test passed
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* moving REGION up
* moving REGION up and csv file name
* fix: changed bucket URL to console
* removing TODOs from Tabnet notebook
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* experiments cuj1 notebook release
* add andy reviews
* linter test passed
* align notebooks
* linter test passed
* minor changes
* linter test passed
* minor changes
* linter test passed
Co-authored-by: Ivan Cheung <ivans.mailbox@gmail.com>
Currently, there is no kernel_name. Hence, the execution test cannot run for notebooks that don't have kernels defined in their .ipynb file.
Side-note: We should use lint to remove the kernel_name from .ipynb as well, as it could include info specific to the author's environment.
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
* Added new Stage 1 notebook to create unlabelled
Vertex AI AutoML text entity extraction dataset
from collection of PDF files on Google Cloud Storage
* Linted notebook
* Removed TODOs
* Updates per PR comments
* Revered to multiple imports per line
* upgrade: current notebook standard
* upgrade: current notebook standard
* Update sdk_automl_tabular_binary_classification_batch_explain.ipynb
* fix: bucket nit
* upgrade: current notebook standard
* upgrade: current notebook standard
* Update google_cloud_pipeline_components_automl_tabular.ipynb
* fix: bucket
* fix: bucket
* upgrade: current notebook standard
* upgrade: current notebook standard
* Update google_cloud_pipeline_components_automl_images.ipynb
* fix: bucket
* fix: bucket
* feat: add example of import from dataframe
* feat: add example of import from dataframe
* update: change in required perms
* update: change in required perms
* review: updates from review
* review: updates from review
* updates: fine tuning
* feat: add example of import from dataframe
* feat: add example of import from dataframe
* update: change in required perms
* update: change in required perms
* review: updates from review
* review: updates from review
* feat: add example of import from dataframe
* feat: add example of import from dataframe
* update: change in required perms
* update: change in required perms
* Added official version of tabular regression batch bq
* Ran linter
* Fixed cleanup
* Additional cleanup
* Added working version
* Refactored and made work
* Ran linter and cleaned up
* Renamed aip to aiplatform
* Replaced online with batch
* Renamed notebook
* Ran linter and cleaned up
* Fixed bug
* Fixed SQL by adding backticks
* Install google-cloud-bigquery[all]
* Refactored datasets
* Ran linter
* Removed GCS cells
* Fixed import file
* Fixed SQL issues and added cleanup of training dataset
* Fixed hardcorded table
* Fixed brand names
* Fixed header
* Fixed results table
* Addressed tech writing review comments
* Ran linter
* added MLPerf benchmark reference and updated Criteo sample to use GRPC for stock containers
* addressed feedback for BERT sample and did similar changes to Criteo sample
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* updates and adds the telecom-subscriber-churn-prediction notebook to official and removes from the community
* ran linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* adds manual-scaling config and explanation to the notebook
* ran linter test after installing linter requirement updates
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* Delete revised version
* Copy notebook from /notebooks/official
* Renamed base notebook
* Added first version by mansari@
* Updated to revised version by andrewferlitsch@
* Added author / reviewer information
Added sample files
* Updated CODEOWNERS
* Fixed links for opening the notebook in Colab/Github/Vertex
Removed installation of and references to pandas
Fixed gcs_annotation_file_name string reference
* Fixed the links for opening notebook (again!)
* Added attribution and references
* Removed references as covered at top
* Updated installation commands to match
* Updated Vertex AI region name to be more clear
* Added db-types dependency for pandas operations
that are now failing
* Minor edits
* Combined package installation and
added a note to ignore the errors
* Minor edit to message
* Added special thanks to andrewferlitsch@
* Updated andrewferlitsch@ GithHub profile link
* Updated sample files URLs to absolute URLs
* Removed empty code block
* Added additional attribution (and the one that did not make it into previous commit!)
* Fixed multi-package import formatting
Switched to pandas instead of db-dtypes
* Fixed isort issue
* Removed unnecessary pandas import
* Formatted the notebook with nbfmt
* Additional notebook formatting
* Updated link to open in Vertex AI Workbench
to point to raw .ipynb file
* Fixed lint issues
* Formatting changes
Added additional APIs to be enabled
* Fixed sample dataset link to point to public version
* Fixed linting issues
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* Commit for lint
* Commit after name change
* Commit of notebook and CODEOWNERS
Added custom container with xai notebook, and explainable_ai folder in the community folder
* Removed extra copy of file
* Remove extra file
* Updated per review from DPE
* Lint test updates
* linter ran
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
When opening a PR, the CODEOWNERS and instructions hyperlinks throw a 404 error because they point to a URL that has been changed. Fixing these hyperlinks.
* fix: new template review updates
* fix: new template review updates
* mport -> import
* fix: dummy code sample required an import
dummy code samples (not otherwise part of template) -- should be self contained since they will be deleted by the template user.
* fix: added install for self-contained code passes ingestion test
* fix: example code (not otherwise part of template) not self-contained.
* fix: continue update so code example is self-contained
* update: numpy already installed in test env
* Delete revised version
* Copy notebook from /notebooks/official
* Renamed base notebook
* Added first version by mansari@
* Updated to revised version by andrewferlitsch@
* Added author / reviewer information
Added sample files
* Updated CODEOWNERS
* Fixed links for opening the notebook in Colab/Github/Vertex
Removed installation of and references to pandas
Fixed gcs_annotation_file_name string reference
* Fixed the links for opening notebook (again!)
* Added attribution and references
* Removed references as covered at top
* Updated installation commands to match
* Updated Vertex AI region name to be more clear
* Added db-types dependency for pandas operations
that are now failing
* Minor edits
* Combined package installation and
added a note to ignore the errors
* Minor edit to message
* Added special thanks to andrewferlitsch@
* Updated andrewferlitsch@ GithHub profile link
* Updated sample files URLs to absolute URLs
* Removed empty code block
* Added additional attribution (and the one that did not make it into previous commit!)
* Fixed multi-package import formatting
Switched to pandas instead of db-dtypes
* Fixed isort issue
* Removed unnecessary pandas import
* Formatted the notebook with nbfmt
* Additional notebook formatting
* Updated link to open in Vertex AI Workbench
to point to raw .ipynb file
* Fixed lint issues
* Formatting changes
Added additional APIs to be enabled
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* Delete revised version
* Copy notebook from /notebooks/official
* Renamed base notebook
* Added first version by mansari@
* Updated to revised version by andrewferlitsch@
* Added author / reviewer information
Added sample files
* Updated CODEOWNERS
* Fixed links for opening the notebook in Colab/Github/Vertex
Removed installation of and references to pandas
Fixed gcs_annotation_file_name string reference
* Fixed the links for opening notebook (again!)
* Added attribution and references
* Removed references as covered at top
* Updated installation commands to match
* Updated Vertex AI region name to be more clear
* Added db-types dependency for pandas operations
that are now failing
* Minor edits
* Combined package installation and
added a note to ignore the errors
* Minor edit to message
* Added special thanks to andrewferlitsch@
* Updated andrewferlitsch@ GithHub profile link
* Updated sample files URLs to absolute URLs
* Removed empty code block
* Added additional attribution (and the one that did not make it into previous commit!)
* Fixed multi-package import formatting
Switched to pandas instead of db-dtypes
* Fixed isort issue
* Removed unnecessary pandas import
* Formatted the notebook with nbfmt
* Additional notebook formatting
* Updated link to open in Vertex AI Workbench
to point to raw .ipynb file
* Fixed lint issues
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* feat: using Vision API for preprocessing data
* feat: using Vision API for preprocessing data
* feat: component vs job resource settings
* feat: component vs job resource settings
* feat: triton server
* feat: triton server
* feat: triton server
* feat: triton server
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* update: ModelEvaluation SDK
* update: ModelEvaluation SDK
* feat: matching engine
* feat: matching engine
* feat: wip: twotowers
* feat: wip: twotowers
* feat: using Vision API for preprocessing data
* feat: using Vision API for preprocessing data
* feat: component vs job resource settings
* feat: component vs job resource settings
* feat: triton server
* feat: triton server
* feat: triton server
* feat: triton server
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* update: ModelEvaluation SDK
* update: ModelEvaluation SDK
* feat: matching engine
* feat: matching engine
* Delete revised version
* Copy notebook from /notebooks/official
* Renamed base notebook
* Added first version by mansari@
* Updated to revised version by andrewferlitsch@
* Added author / reviewer information
Added sample files
* Updated CODEOWNERS
* Fixed links for opening the notebook in Colab/Github/Vertex
Removed installation of and references to pandas
Fixed gcs_annotation_file_name string reference
* Fixed the links for opening notebook (again!)
* Added attribution and references
* Removed references as covered at top
* Updated installation commands to match
* Updated Vertex AI region name to be more clear
* Added db-types dependency for pandas operations
that are now failing
* Minor edits
* Combined package installation and
added a note to ignore the errors
* Minor edit to message
* Added special thanks to andrewferlitsch@
* Updated andrewferlitsch@ GithHub profile link
* Updated sample files URLs to absolute URLs
* Removed empty code block
* Added additional attribution (and the one that did not make it into previous commit!)
* Fixed multi-package import formatting
Switched to pandas instead of db-dtypes
* Fixed isort issue
* Removed unnecessary pandas import
* Formatted the notebook with nbfmt
* Additional notebook formatting
* feat: using Vision API for preprocessing data
* feat: using Vision API for preprocessing data
* feat: component vs job resource settings
* feat: component vs job resource settings
* feat: triton server
* feat: triton server
* feat: triton server
* feat: triton server
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* update: ModelEvaluation SDK
* update: ModelEvaluation SDK
* feat: using Vision API for preprocessing data
* feat: using Vision API for preprocessing data
* feat: component vs job resource settings
* feat: component vs job resource settings
* feat: triton server
* feat: triton server
* feat: triton server
* feat: triton server
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* feat: using Vision API for preprocessing data
* feat: using Vision API for preprocessing data
* feat: component vs job resource settings
* feat: component vs job resource settings
* feat: triton server
* feat: triton server
* feat: triton server
* feat: triton server
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* feat: using Vision API for preprocessing data
* feat: using Vision API for preprocessing data
* feat: component vs job resource settings
* feat: component vs job resource settings
* feat: triton server
* feat: triton server
* feat: triton server
* feat: triton server
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* feat: using Vision API for preprocessing data
* feat: using Vision API for preprocessing data
* feat: component vs job resource settings
* feat: component vs job resource settings
* feat: triton server
* feat: triton server
* feat: triton server
* feat: triton server
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* feat: using Vision API for preprocessing data
* feat: using Vision API for preprocessing data
* feat: component vs job resource settings
* feat: component vs job resource settings
* feat: triton server
* feat: triton server
* feat: triton server
* feat: triton server
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* fix: check for workbench
* feat: using Vision API for preprocessing data
* feat: using Vision API for preprocessing data
* feat: component vs job resource settings
* feat: component vs job resource settings
* feat: triton server
* feat: triton server
* feat: triton server
* feat: triton server
* fix: check for workbench
* fix: check for workbench
* Delete revised version
* Copy notebook from /notebooks/official
* Renamed base notebook
* Added first version by mansari@
* Updated to revised version by andrewferlitsch@
* Added author / reviewer information
Added sample files
* Updated CODEOWNERS
* Fixed links for opening the notebook in Colab/Github/Vertex
Removed installation of and references to pandas
Fixed gcs_annotation_file_name string reference
* feat: using Vision API for preprocessing data
* feat: using Vision API for preprocessing data
* feat: component vs job resource settings
* feat: component vs job resource settings
* feat: triton server
* feat: triton server
* feat: triton server
* feat: triton server
* feat: using Vision API for preprocessing data
* feat: using Vision API for preprocessing data
* feat: component vs job resource settings
* feat: component vs job resource settings
* feat: triton server
* feat: triton server
* feat: using Vision API for preprocessing data
* feat: using Vision API for preprocessing data
* feat: component vs job resource settings
* feat: component vs job resource settings
* feat: add colab code for docker
* feat: add colab code for docker
* fix: add colab support for docker
* fix: add colab support for docker
* fix: delete tmp BQ model
* fix: delete tmp BQ model
* feat: GAPIC->SDK for private endpoints
* feat: GAPIC->SDK for private endpoints
* fix: add IS_COLAB flag
* fix: add IS_COLAB flag
* fix: add IS_COLAB flag
* fix: add IS_COLAB flag
* fix: add IS_COLAB flag
* fix: add IS_COLAB flag
* fix: IS_COLAB
* fix: IS_COLAB
* fix: IS_COLAB
* fix: IS_COLAB
* fix: IS_COLAB
* feat: more model eval work
* feat: more model eval work
* made changes
* ran linter test
* added minor changes
* ran linter test
* made changes
* ran linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* made changes
* ran linter test
* made minor changes
* ran linter test
* made changes
* ran linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* changes made
* ran linter test
* made minor changes
* ran linter test
* made changes
* ran linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* Add minor changes to get_started_with_dataflow_pipeline_components
* minor changes and tested
* remove variable dataflow_wait_op, since not used in other places.
* remove variable dataflow_wait_op, since not used in other places
* removed unused import
* Run linter test
* Add gcloud project set when using colab
* Run linter
* correct anem toColab logo Run in Colab
* run linter
* correct the list of items to remove
* Run linter
* Run liinter
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* Added vertexai notebook
* Ran the linter test
* Made the required changes based on the comments
* Ran linter test again
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* feat: add colab code for docker
* feat: add colab code for docker
* fix: add colab support for docker
* fix: add colab support for docker
* fix: delete tmp BQ model
* fix: delete tmp BQ model
* Add minor changes to get_started_vertex_datasets notebook
* run linter
* Run Linter test
* Add google authentication cell for colab execution
* run linter
* correct the project id definition
* Run linter
* Add project id cell
* run liner
* Added imports that are required
* run linter test
* Add gcloud project set
* Run linter
* add linter run
* Running linter test
* run linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* adds the ml_ops/stage2/get_Started_bqml_training notebook to official and removes the same from community folder
* ran linter test
* updates the textual content
* ran linter test
* moves the updated stage2/get-started-bqml notebook back to the communit folder
* ran linter test
* updates the header according to the template
* ran linter test
* adds colab part and minor changes
* ran linter test
* retains the newly added code lost in conflicts
* ran linter test
* converts vertex to vertex ai
* ran linter test
* moves deletion of temporary BQ table outside delete_storage condition
* ran linter test
* adds bigquery-storage dependency to the notebook tested on Colab
* ran linter test
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* modified file
* modified file
* ran linter test
* deleted file in community folder
* ran linter test
* changed folder name in links
* ran linter test
* resolved comments
* ran linter test
* modified file
* ran linter test
* deleted file in community folder
* modified notebook
* ran linter test
* renamed managed_notebooks folder to workbench
* ran linter
* resolved comments
* ran linter test
* pulled new version of branch
* ran linter again
* resolved comments
* ran linter test
* removed %%time and added --user flag to all pip installs
* ran linter test
* added debug statements
* ran linter test
* added verbose
* ran linter test
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* updates the get-started-vertex-experiments notebook in the community folder
* ran linter test
* adds the costs section
* ran linter test
* adds colab part and minor changes
* ran linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* updates the get-started-automl-training notebook
* ran linter test
* adds --user flag during installation step
* ran linter test
* updates the clean up step
* ran linter test
* adds colab part and minor changes
* ran linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
* adds the updated mlops-stage1-get_started_bq_datasets notebook to the official branch and removes it from the community branch
* removes second instance of create_bigquery_dataset() function
* ran linter test successfully
* adds costs section
* ran linter test successfully
* updates the dependency installation step and GCS bucket explanation
* ran linter test
* adds pyarrow to the installations
* ran linter test
* removes unnecessary installations + adds silent install + moves the notebook back from official to community folder + adds IS_TESTING condition during clean-up
* ran linter test
* resolves the move up?? comment and builtin comment
* ran linter test
* updates textual content about package installation
* ran linter test
* resolves the future-tense and dependency installations comments
* ran linter test
* updates the header according to template
* ran linter test
* adds Colab part and minor changes
* ran linter test
* updates the enable apis step in setup project section
* ran linter test
* changes vertex to vertex ai
* ran linter test
* moves temporary BQ table deletion outside the delete_storage condition
* ran linter test
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* modified file
* made linter changes
* made changes
* linter test issues resolved
* ran linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* Start a new branch for TabNet tutorial.
* Clean version Created using Colaboratory
* Created using Colaboratory
* Remove unused import
* format lint
* Remove unused import
* Created using Colaboratory
* Remove unused import
* Fix the first iteration of reviewing except the image location
* add import
* Update the image to vertex
* Force delete the BQ to avoid waiting
* Add codeowner for TabNet
* Remove - from folder name
* Add deployment in Vertex AI
* Add delete the resource
Co-authored-by: Long Le <longtle@google.com>
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* minor changes made to notebook
* ran lintertest
* added coment
* ran lintertest
* made changes sujjested in git review
* ran linter test
* changes done as per review
* ran lintertest
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* adding Vertex AI optimized TensorFlow runtime samples
* updated URLs, added code to import benchmark.py
* fixed 'Open in Vertex AI Workbench' links
* final cleanup
* added @vlasesnkoalexey as an owner of notebooks/community/vertex_endpoints/optimized_tensorflow_runtime
* rerun linter
* updates the get-started-automl-pipelines in the mlops/stage3 folder inside community folder
* replaces the unused variable deploy_op with _
* removes the unused Model import
* adds the costs section
* ran linter test
* adds Colab part to the notebook
* ran linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* updates the mlops/stage3/get_started_with_kubeflow_pipelines.ipynb notebook
* fixes unused variables
* fixes conflicting function names
* ran linter test
* adds colab changes
* ran linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* updates: adds delete-batch code + adds colab part + updates textual content
* sets delete_bucket to False as default
* ran linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* updates get-started-featurestore notebook in mlops/stage2
* ran linter test
* adds the colab changes and minor textual changes
* ran linter test
* adds the colab changes to the notebook and minor textual changes
* ran linter test
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* modified file
* run linter test
* run in colab
* added coment
* run lintertest
* changed as per review coments
* ran lintertest
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* add new notebook version
* linter test done. passed
* simple fix
* add images
* linter test done
* fix image name
* fix file name in the notebook
* linter code run. done
* linter code run. done
* name fixes. linter code done. passed.
* fix project id and region
* test done
* format
* linter test done.
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* updates the mlops/stage3/get_started_with_kubeflow_pipelines.ipynb notebook
* fixes unused variables
* fixes conflicting function names
* ran linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* updates the get-started-automl-pipelines in the mlops/stage3 folder inside community folder
* replaces the unused variable deploy_op with _
* removes the unused Model import
* adds the costs section
* ran linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* modified notebook
* linter test issues resolved
* ran linter test
* added colab option
* ran linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* updates the get-started-automl-training notebook
* ran linter test
* adds --user flag during installation step
* ran linter test
* updates the clean up step
* ran linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
* add dataproc components tabular notebook
* add src package
* add codeowner
* linter test done. almost ok except for the flake8 E231. need to follow up with andy
* fix typos based on andy review
* linter test done. review with andy
* hyperparameter_tuning_op fix
* project name
* add delete repo
* fix image
* linter test done
* fix image reference
* fix typo image reference
* minor fixes
* karl fixes
* karl fixes on links
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
* add new notebook version
* linter test done. passed
* simple fix
* add images
* linter test done
* fix image name
* fix file name in the notebook
* linter code run. done
* linter code run. done
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
* feat: add notebook for FastAPI server
* feat: add notebook for FastAPI server
* feat: add notebook for FastAPI server
* feat: notebook for private endpoints
* feat: notebook for private endpoints
* license tweak
* remove unused import json
* fixed a missing import
* add sleep(300) to test my theory
* add missing newline
* put sleep behind a conditional
* revert new notebook name to previous name for compatibility with extant links
* fix quoting syntax error
* reformatted due to relint
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
* Add example using auto scale
* Format with nbqa
* Complete sentence
* Give the sample for CreateFeaturestoreRequest only, instead of actual call to create FS to avoid duplicate resource or extra cleanup.
* Remove unused import
* Remove version pinning
* Add try block to avoid error when test was not cleanup properly.
* Lint
* Fix import
* Merge print lro result with the call in the same try block
* Fix typo
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: Morgan Du <morgandu@google.com>
In "Step by Step Guide to Building Reinforcement Learning Applications using Vertex AI", the replay_buffer was unbound if training_data_spec_transformation_fn was provided to the train() function
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
* update: more details to objective on endpoint notebook
* update: more details to objective on endpoint notebook
* update: more details to objective on endpoint notebook
* update: more details to objective on endpoint notebook
* update: workaround for blocking issue
* update: workaround for blocking issue
* fix: reconfigure endpoint
* fix: reconfigure endpoint
* feat: add get started with TF serving functions
* feat: add get started with TF serving functions
* feat: notebook for TF Serving
* feat: notebook for TF Serving
* bqml pipeline notebook for official blog
* add notebook to CODEOWNERS
* add author name
* requirements commenting fix
* linter test done
* unpin the maintenance version for kfp
* fix: install conflicts
* Update google_cloud_pipeline_components_bqml_text.ipynb
* add karl fix
* lint test done
* add andy fixes
* linter test done
* flip order of the special METADATA fix
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@gmail.com>
* update: more details to objective on endpoint notebook
* update: more details to objective on endpoint notebook
* update: more details to objective on endpoint notebook
* update: more details to objective on endpoint notebook
* update: workaround for blocking issue
* update: workaround for blocking issue
* fix: reconfigure endpoint
* fix: reconfigure endpoint
* feat: add get started with TF serving functions
* feat: add get started with TF serving functions
* Ml ops 7v2 (#429)
* update: more details to objective on endpoint notebook
* update: more details to objective on endpoint notebook
* update: more details to objective on endpoint notebook
* update: more details to objective on endpoint notebook
* Update README.md
* Add files via upload
* Update README.md
* Delete stage6b.png
* Delete stage6c.png
* Add files via upload
* Delete stage6b.png
* Delete stage6c.png
* Add files via upload
* Delete stage6b.png
* feat: new notebook on endpoints (#430)
* update: more details to objective on endpoint notebook
* update: more details to objective on endpoint notebook
* update: more details to objective on endpoint notebook
* update: more details to objective on endpoint notebook
* update: workaround for blocking issue
* update: workaround for blocking issue
* wrong location
* Create README.md
* Update README.md
* Update README.md
* Update README.md
* Update README.md
* fix: links
* fix: title
* fix: example for reconfiguring the traffic split (#431)
* update: more details to objective on endpoint notebook
* update: more details to objective on endpoint notebook
* update: more details to objective on endpoint notebook
* update: more details to objective on endpoint notebook
* update: workaround for blocking issue
* update: workaround for blocking issue
* fix: reconfigure endpoint
* fix: reconfigure endpoint
* Add section on granting Dataproc IAM roles.
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@gmail.com>
Co-authored-by: Win Woo <wwoo@google.com>
* update: more details to objective on endpoint notebook
* update: more details to objective on endpoint notebook
* update: more details to objective on endpoint notebook
* update: more details to objective on endpoint notebook
* update: workaround for blocking issue
* update: workaround for blocking issue
* fix: reconfigure endpoint
* fix: reconfigure endpoint
* update: more details to objective on endpoint notebook
* update: more details to objective on endpoint notebook
* update: more details to objective on endpoint notebook
* update: more details to objective on endpoint notebook
* update: workaround for blocking issue
* update: workaround for blocking issue
* update: more details to objective on endpoint notebook
* update: more details to objective on endpoint notebook
* update: more details to objective on endpoint notebook
* update: more details to objective on endpoint notebook
* feat: add data labeling notebook
* feat: add data labeling notebook
* update: details on dsl.Condition
* update: details on dsl.Condition
* feat: add TFHub model example
* feat: add TFHub model example
* adds the updated mlops-stage1-get_started_bq_datasets notebook to the official branch and removes it from the community branch
* removes second instance of create_bigquery_dataset() function
* ran linter test successfully
* adds costs section
* ran linter test successfully
* updates the dependency installation step and GCS bucket explanation
* ran linter test
* adds pyarrow to the installations
* ran linter test
* removes unnecessary installations + adds silent install + moves the notebook back from official to community folder + adds IS_TESTING condition during clean-up
* ran linter test
* resolves the move up?? comment and builtin comment
* ran linter test
* updates textual content about package installation
* ran linter test
* resolves the future-tense and dependency installations comments
* ran linter test
* updates the header according to template
* ran linter test
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
* add automl tabular regression online bq with minor changes
* Run Linter
* Fix errors from the CLA test
* run linter
* resolve issue.
* run Linter
* Merge
* test lint
* fix for linter test
* add automl tabular regression online bq with minor changes
* Run Linter
* Fix errors from the CLA test
* run linter
* resolve issue.
* run Linter
* Merge
* test lint
* fix for linter test
* Fix Bucket name variable
* run linter
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
* adds the ml_ops/stage2/get_Started_bqml_training notebook to official and removes the same from community folder
* ran linter test
* updates the textual content
* ran linter test
* moves the updated stage2/get-started-bqml notebook back to the communit folder
* ran linter test
* updates the header according to the template
* ran linter test
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
* update: for official
* update: for official
* cleanup: add deleting model/endpoint created from pipeline
* cleanup: add deleting model/endpoint created from pipeline
* feat: add dataproc notebook
* feat: add dataproc notebook
* update: for official
* update: for official
* cleanup: add deleting model/endpoint created from pipeline
* cleanup: add deleting model/endpoint created from pipeline
* Add minor changes to automl image object detection
* run linter
* Correct the milli nodes hours
* fix errors
* fix getenv
* Run linter
* remove tabular notebook, wrongly added
* Correct the bucket varible and minor changes to text
* Run linter
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* fix: use os.getenv()
* fix: use os.getenv()
* fix: use os.getenv()
* fix: use os.getenv()
* fix: use os.getenv()
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
* Add NVIDIA Triton on Vertex AI Prediction official notebook
* Add NVIDIA Triton on Vertex AI Prediction official notebook
* Add NVIDIA Triton on Vertex AI Prediction official notebook
* Add NVIDIA Triton on Vertex AI Prediction community notebook
* Add NVIDIA Triton on Vertex AI Prediction community notebook
* Add NVIDIA Triton on Vertex AI Prediction community notebook
* Fixes based on feedback to NVIDIA Triton on Vertex AI Prediction community notebook
* Start a new branch for TabNet tutorial.
* Clean version Created using Colaboratory
* Created using Colaboratory
* Remove unused import
* format lint
* Remove unused import
* Created using Colaboratory
* Remove unused import
* Fix the first iteration of reviewing except the image location
* add import
* Update the image to vertex
* Force delete the BQ to avoid waiting
* Add codeowner for TabNet
* Remove - from folder name
Co-authored-by: Long Le <longtle@google.com>
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
* Using BQML 1st-party components and 1.0.0 of google-cloud-pipeline-components
* Using BQML 1st-party components and upgrading to 1.0.0 of google-cloud-pipeline-components
* Using BQML 1st-party components and upgrading to 1.0.0 of google-cloud-pipeline-components
* Using BQML 1st-party components and upgrading to 1.0.0 of google-cloud-pipeline-components
* Using BQML components and upgrade to 1.0.0 of GCPC
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
* adds automl-text-sentiment-analysis-online notebook
* adds the cleaned up automl-text-sentiment-analysis notebook after running linter test
* adds textual content on what the dataset predicts in the dataset section
* ran the linter test after the update
* adds textual content on what the dataset predicts in the dataset section
* ran the linter test after the update
* corrects the IMPORT_FILE parameter in the notebook
* ran linter test after update
* deletes the source file from the community/sdk folder
* updates the colab, git & workbench links in the notebook
* ran linter test
* updates the license year to 2022 and simplifies the clean-up step for bucket-deletion
* ran linter test
* adds TESTING env condition while deleting the buckets
* ran linter test successfully
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
* adds the automl-video-action-recognition-notebook
* ran linter test
* fixes the dag variable by replacing with job
* ran linter test
* fixes the dag variable by replacing with job
* ran linter test
* corrects the IMPORT_FILE parameter in the notebook
* ran linter test after update
* updates the colab, git & vertex-ai links
* ran linter test
* updates the license year to 2022 and simplifies the lean-up step for bucket created
* ran linter test
* removes the file from the community folder
* adds the TESTING env condition while deleting the buckets
* ran linter test successfully
* adds TESTING env condition while deleting the bucket
* ran linter test successfully
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
* modified colab,github,vertexAI links and added vertex logo
* ran linter
* resolved comments
* ran linter
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
* changed master to main for links and added vertex AI logo
* ran linter
* resolved comments
* ran linter
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
* A notebook that shows Vertex AI feature store capabilities in a real-world scenario (#296)
* A notebook that shows Vertex AI feature store capabilities in a real-world scenario
* new notebook version
* fix CODEOWNERS
* comment to the feature store monitoring api
* format notebook
* fix CODEOWNERS
* fix CODEOWNERS as required
* new version
* new notebook version
* notebook cleaning
* new update
* add fix to pass lint test
* resolve conflict
* import libraries fix
* update image
* update notebook
* fix comment
* new notebook version
* new notebook and assets
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* Added files in their old folder
* Deleted unneeded file
* Ran linter
* Fixed CODEOWNERS
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
Co-authored-by: ivanmkc <ivans.mailbox@gmail.com>
* A notebook that shows Vertex AI feature store capabilities in a real-world scenario
* new notebook version
* fix CODEOWNERS
* comment to the feature store monitoring api
* format notebook
* fix CODEOWNERS
* fix CODEOWNERS as required
* new version
* new notebook version
* notebook cleaning
* new update
* add fix to pass lint test
* resolve conflict
* import libraries fix
* update image
* update notebook
* fix comment
* new notebook version
* new notebook and assets
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* Deploying TF Hub object detection model using Vertex endpoints
* Add user to codeowners
* fix path in CODEOWNERS
* clear all outputs
* run linter
* manual lint fix
* fix more linting errors
* order imports in alphabetical order
* run linter
* made changes requested on feedback
* automate fetching endpoint model id
* fix hardcoded value in bash command
* generalize region endpoint and project in bash cell
* retrieve endpoint and model ids programatically
* fix formatting
* run linter
* Remove pipfile
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
* adds the automl-video-action-recognition-notebook
* ran linter test
* fixes the dag variable by replacing with job
* ran linter test
* fixes the dag variable by replacing with job
* ran linter test
* corrects the IMPORT_FILE parameter in the notebook
* ran linter test after update
* updates the colab, git & vertex-ai links
* ran linter test
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
* Add google_cloud_pipeline_components_model_upload_predict_evaluate notebook.ipynb
* format with linter
* add import for tensorflow when in the testing environment
* linter
* fix dependency issues for testing env
* address comments
* eval component does not output gcp_resources yet, still in experimental
* added location to aip.init
* add deletion for model and batch prediction jobs
* typo, missed a comma.
* linter
* Remove tensorflow import + use gsutil to check if artifacts exist.
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
* notebook refresh from vertex ai sdk project batch 1
* successfully ran linter test
* removed global variable import file
* update with linter test changes
* removing community version of dk_automl_video_classification_batch.ipynb
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* added notebook
* ran linter
* fix aip not defined error
* ran lint
* resolved git comments
* ran linter
* deleted file in community folder and removed globals
* ran linter
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* added notebook
* changed folder
* reinstalled linter
* ran linter
* pulled new changes and merged
* resolved comments
* resolved comments
* ran linter
* deleted file in community folder and removed globals in file
* ran linter
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
* PyTorch on Vertex - Updated to match GCPC v0.2.2 API
* PyTorch on Vertex - Fixes based on review comments
* PyTorch on Vertex - fixes based on review
* PyTorch on Vertex - linter fixes
* PyTorch on Vertex - fixes based on feedback
* PyTorch on Vertex - fixes based on feedback
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
* feat: get started XAI
* feat: get started XAI
* feat: XAI with sklearn
* feat: XAI with sklearn
* feat: add covert component example
* feat: add covert component example
* feat: upgrade FS to SDK
* feat: upgrade FS to SDK
* feat: update to v1
* feat: update to v1
* fix: XAI for sklearn
* fix: XAI for sklearn
* SDK Featurestore notebook
* fixed issues, tried to make notebook more readable, style
* removed previous notebook
* moved sdk-feature-store to community (for now)
* made fixes
* moved BQ output table cells down
Co-authored-by: Andrew Ferlitsch <aferlitsch@google.com>
Co-authored-by: Morgan Du <morgandu@google.com>
* feat: get started XAI
* feat: get started XAI
* feat: XAI with sklearn
* feat: XAI with sklearn
* feat: add covert component example
* feat: add covert component example
* feat: upgrade FS to SDK
* feat: upgrade FS to SDK
* feat: update to v1
* feat: update to v1
* feat: get started XAI
* feat: get started XAI
* feat: XAI with sklearn
* feat: XAI with sklearn
* feat: add covert component example
* feat: add covert component example
* feat: upgrade FS to SDK
* feat: upgrade FS to SDK
* Fixes and renames link to launch automl-text-classification.pynb in Vertex AI Workbench
* Fixes and renames link to launch sdk_automl_tabular_forecasting_batch.pynb in Vertex AI Workbench
* Fixes links for launching notebook in Vertex AI Workbench for Explainable AI samples
* Fixes link for launching notebook in Vertex AI Workbench for model monitoring sample
* Fixes links to launch pipelines notebook samples
* Fixed lint problem in automl-text-classification.ipynb
* Autofixed lint errors
Co-authored-by: Karl Weinmeister <11586922+kweinmeister@users.noreply.github.com>
* feat: get started XAI
* feat: get started XAI
* feat: XAI with sklearn
* feat: XAI with sklearn
* feat: add covert component example
* feat: add covert component example
* Added example for PyTorch lightning distributed training of a ResNet model
* Revert "Added example for PyTorch lightning distributed training of a ResNet model"
This reverts commit bcbe832c51.
* Added example for PyTorch lightning distributed training of a ResNet model
* Revert "Added example for PyTorch lightning distributed training of a ResNet model"
This reverts commit 2a792cb7ac.
* Added example for PyTorch lightning distributed training of a ResNet model
* Added Notebook for PyTorch lightning distributed training of a ResNet model
* Added Notebook for PyTorch lightning distributed training of a ResNet model
* Notebook updates after review
* Notebook updates after review
* Added example for PyTorch lightning distributed training of a ResNet model
* Revert "Added example for PyTorch lightning distributed training of a ResNet model"
This reverts commit bcbe832c51.
* Added example for PyTorch lightning distributed training of a ResNet model
* Revert "Added example for PyTorch lightning distributed training of a ResNet model"
This reverts commit 2a792cb7ac.
* Added example for PyTorch lightning distributed training of a ResNet model
* Added Notebook for PyTorch lightning distributed training of a ResNet model
* Added Notebook for PyTorch lightning distributed training of a ResNet model
* Notebook updates after review
* Notebook updates after review
* Adjust Tensorboard to TensorBoard
* Adjust Tensorboard to TensorBoard
* Revert "Adjust Tensorboard to TensorBoard"
This reverts commit 9aac52e358b4ccc27a9a5a9e3bc5ef3305553462.
* Adjust Tensorboard to TensorBoard
* Adjust Tensorboard to TensorBoard
* Adjust Tensorboard to TensorBoard and and run lint
* Added example for PyTorch lightning distributed training of a ResNet model
* Revert "Added example for PyTorch lightning distributed training of a ResNet model"
This reverts commit bcbe832c51.
* Added example for PyTorch lightning distributed training of a ResNet model
* Revert "Added example for PyTorch lightning distributed training of a ResNet model"
This reverts commit 2a792cb7ac.
* Added example for PyTorch lightning distributed training of a ResNet model
help="The GCP region. This is used to inject a variable value into the notebook before running.",
required=True,
)
parser.add_argument(
"--variable_service_account",
type=str,
help="A service account. This is used to inject a variable value into the notebook before running. This is not the account that will run the notebook.",
required=True,
)
parser.add_argument(
"--variable_vpc_network",
type=str,
help="The full VPC network name. See https://cloud.google.com/compute/docs/networks-and-firewalls#networks. Format is projects/{project}/global/networks/{network}, where {project} is a project number, as in '12345', and {network} is network name. See <https://cloud.google.com/compute/docs/reference/rest/v1/networks/insert> for details. This is used to inject a variable value into the notebook before running.",
required=False,
)
parser.add_argument(
"--staging_bucket",
type=str,
@@ -73,6 +86,13 @@ parser.add_argument(
help="The GCP directory for storing executed notebooks.",
"**_NOTE_**: This notebook has been tested in the following environment:\n",
"\n",
"* Python version = 3.7\n",
"\n",
"**The following steps are required, regardless of your notebook environment.**\n",
"\n",
"1. [Select or create a Google Cloud project](https://console.cloud.google.com/cloud-resource-manager). When you first create an account, you get a $300 free credit towards your compute/storage costs.\n",
"\n",
"1. [Make sure that billing is enabled for your project](https://cloud.google.com/billing/docs/how-to/modify-project).\n",
"\n",
"1. [Enable the Vertex AI API and Compute Engine API](https://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com,compute_component).\n",
"\n",
"1. If you are running this notebook locally, you will need to install the [Cloud SDK](https://cloud.google.com/sdk).\n",
"\n",
"1. Enter your project ID in the cell below. Then run the cell to make sure the\n",
"Cloud SDK uses the right project for all the commands in this notebook.\n",
"\n",
"**Note**: Jupyter runs lines prefixed with `!` as shell commands, and it interpolates Python variables prefixed with `$` into these commands."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "c6516f90311b"
},
"outputs": [],
"source": [
"# test if the right python version is being used\n",
If you are opening a PR for `Official Notebooks` under the [notebooks/official](https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/main/notebooks/official) folder, follow this mandatory checklist:
**REQUIRED:** Add a summary of your PR here, typically including why the change is needed and what was changed. Include any design alternatives for discussion purposes.
<br>
--- YOUR PR SUMMARY GOES HERE ---
<br><br><br>
**REQUIRED:** Fill out the below checklists or remove if irrelevant
1. If you are opening a PR for `Official Notebooks` under the [notebooks/official](https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/main/notebooks/official) folder, follow this mandatory checklist:
- [ ] Use the [notebook template](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/notebook_template.ipynb) as a starting point.
- [ ] Follow the style and grammar rules outlined in the above notebook template.
- [ ] Verify the notebook runs successfully in Colab since the automated tests cannot guarantee this even when it passes.
- [ ] Passes all the required automated checks. You can locally test for formatting and linting with these [instructions](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/docs/contributing.md#code-quality-checks).
- [ ] Passes all the required automated checks. You can locally test for formatting and linting with these [instructions](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/CONTRIBUTING.md#code-quality-checks).
- [ ] You have consulted with a tech writer to see if tech writer review is necessary. If so, the notebook has been reviewed by a tech writer, and they have approved it.
- [ ] This notebook has been added to the [CODEOWNERS](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/docs/CODEOWNERS) file under `# Official Notebooks` section, pointing to the author or the author's team.
- [ ] This notebook has been added to the [CODEOWNERS](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/official/CODEOWNERS) file under the `Official Notebooks` section, pointing to the author or the author's team.
- [ ] The Jupyter notebook cleans up any artifacts it has created (datasets, ML models, endpoints, etc) so as not to eat up unnecessary resources.
<br>
If you are opening a PR for `Community Notebooks` under the [notebooks/community](https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/main/notebooks/community) folder:
- [ ] This notebook has been added to the [CODEOWNERS](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/docs/CODEOWNERS) file under the `# Community Notebooks` section, pointing to the author or the author's team.
- [ ] Passes all the required formatting and linting checks. You can locally test with these [instructions](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/docs/contributing.md#code-quality-checks).
2.If you are opening a PR for `Community Notebooks` under the [notebooks/community](https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/main/notebooks/community) folder:
- [ ] This notebook has been added to the [CODEOWNERS](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/community/CODEOWNERS) file under the `Community Notebooks` section, pointing to the author or the author's team.
- [ ] Passes all the required formatting and linting checks. You can locally test with these [instructions](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/CONTRIBUTING.md#code-quality-checks).
If you are opening a PR for `Community Content` under the [community-content](https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/main/community-content) folder:
<br>
3. If you are opening a PR for `Community Content` under the [community-content](https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/main/community-content) folder:
- [ ] Make sure your main `Content Directory Name` is descriptive, informative, and includes some of the key products and attributes of your content, so that it is differentiable from other content
- [ ] The main content directory has been added to the [CODEOWNERS](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/docs/CODEOWNERS) file under the `# Community Content` section, pointing to the author or the author's team.
- [ ] Passes all the required formatting and linting checks. You can locally test with these [instructions](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/docs/contributing.md#code-quality-checks).
- [ ] The main content directory has been added to the [CODEOWNERS](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/community-content/CODEOWNERS) file under the `Community Content` section, pointing to the author or the author's team.
- [ ] Passes all the required formatting and linting checks. You can locally test with these [instructions](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/CONTRIBUTING.md#code-quality-checks).
@@ -6,7 +6,19 @@ Welcome to the Google Cloud [Vertex AI](https://cloud.google.com/vertex-ai/docs/
## Overview
The repository contains [Notebooks](https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/master/notebooks) and [Community Content](https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/master/community-content) that demonstrate how to develop and manage ML workflows using Google Cloud Vertex AI.
The repository contains [notebooks](https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/master/notebooks) and [community content](https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/master/community-content) that demonstrate how to develop and manage ML workflows using Google Cloud Vertex AI.
## Repository structure
```bash
├── community-content - Sample code and tutorials contributed by the community
├── notebooks
│ ├── community - Notebooks contributed by the community
│ ├── official - Notebooks demonstrating use of each Vertex AI service
│ │ ├── automl
│ │ ├── custom
│ │ ├── ...
```
## Contributing
@@ -19,3 +31,7 @@ Please use the [issues page](https://github.com/GoogleCloudPlatform/vertex-ai-sa
## Disclaimer
This is not an officially supported Google product. The code in this repository is for demonstrative purposes only.
## Feedback
Please feel free to fill out our [survey](https://bit.ly/vertex-ai-samples-survey) to give us feedback on the repo and its content.
CPR ([custom prediction routines](https://github.com/googleapis/python-aiplatform/blob/main/google/cloud/aiplatform/prediction/README.md)) is a framework designed by Google Cloud developers to make it easier to combine machine learning models with custom preprocessing and postprocessing logic in a real-time serving application.
## Using this example
This code is a self-contained example of a custom model server project built using CPR.
As is, you can use it to serve the ViT-Small image classification model from Ross Wightman's [`timm`](https://github.com/rwightman/pytorch-image-models) library of image model implementations in PyTorch. Both CPU and GPU are supported.
You can also consider using the code here as a template for your own CPR project if you want to use a different model from `timm`, a different PyTorch model, or an entirely different framework.
### Requirements
In order to use this example, you'll need Docker and Python 3 installed on your system.
To get started, first create a virtual environment in an empty directory:
```sh
mkdir cpr-example
python3 -m venv cpr-example
cd cpr-example &&source bin/activate
```
Then, clone the [vertex-ai-samples repo](https://github.com/GoogleCloudPlatform/vertex-ai-samples) in that directory:
The `TimmPredictor` class in `timm_serving/predictor.py` implements most of the important logic for the server.
-`load(artifacts_dir)`: The predictor's `load` method is called when the server starts up in order to set up the predictor, usually by loading model weights and any artifacts needed for preprocessing and postprocessing. In this example, we initialize the saved model from the `state_dict.pth` file located inside the `artifacts_dir` folder and create the preprocessing transform from the model config.
-`preprocess`, `predict`, `postprocess`: These methods are applied in sequence to the deserialized JSON data from each request.
-`preprocess` decodes images from base64 and apply cropping, scaling and normalizing transforms.
-`predict` runs the ViT-Small model on the preprocessed images and returns class scores.
-`postprocess` finds the top five classes and packs the class names, probabilities, and indices in a serializable result.
### Building the container
To build the model server locally, run the build command:
```sh
python build.py build
```
You can edit configuration values such as the model server's base image, the name and tag assigned to the image, and the path where model weights are stored locally.
When you run the build command, model weights are downloaded and the model server container is built.
### Running local tests
`test.py` contains a suite of unit tests for the predictor as well as end-to-end tests for the model server.
- The infamous [mandrill](https://commons.wikimedia.org/wiki/File:Wikipedia-sipi-image-db-mandrill-4.2.03.png)
### Deploying to Vertex AI
Before uploading or deploying the container, you'll need to modify `config.py` to set appropriate values for:
-`project_id`: Your GCP project id.
-`region`: Region where the model will be uploaded and deployed.
-`repository`: [Artifact Registry repository](https://cloud.google.com/artifact-registry/docs/repositories/create-repos) in your project where the container image will be uploaded.
-`artifacts_gcs_dir`: Folder in a [Google Cloud Storage bucket](https://cloud.google.com/storage/docs/creating-buckets) where the model weights will be uploaded.
Once this is done, first upload the model:
```sh
python build.py upload
```
Then deploy it:
```sh
python build.py deploy
```
If you run the deploy command again, it will create a new endpoint. If you want to undeploy the model, you can do so using the Vertex AI dashboard on the Google Cloud console, or use `gcloud ai endpoints undeploy` from the command line.
After deploying successfully, you can run `python build.py probe` to send a sample request to the deployed model.
*Pluto* is a programming environment for Julia, designed to be interactive and helpful. It provides a familiar notebook interface but it is not a Jupyter notebook. The biggest difference is that Pluto notebooks are reactive, changing a variable or function in one cell causes the cells that depend on that variable or function to be reevaluated. Pluto also provides useful interaction mechanisms that allow users to dynamically interact with the notebooks computation state.
The JuliaCon 2020 presentation: [Interactive notebooks ~ Pluto.jl]() provides a good introduction to Pluto. The source is at [fonsp/Pluto.jl]()
# Install Pluto
## Create a Vertex AI JupyterLab Instance
1. From the [GCP console](https://console.cloud.google.com) "hamburger menu"
select Vertex AI > Workbench
2. Click NEW NOTEBOOK
* Choose Python 3 if you won't be using a GPU
* Choose Python 3 (CUDA Toolkit xx.y) if you do want use a GPU
3. Give the notebook an appropriate name
4. Edit Notebook properties if you have special requirements otherwise accept the defaults and click CREATE
5. When the notebook instance is ready click OPEN JUPYTERLAB
# PyTorch Deployment on Google Cloud: Text Classification
**This is an Experimental release**, covered by the Pre-GA Offerings Terms of your Google Cloud Platform [Terms of Service](https://cloud.google.com/terms).
Experiments are focused on validating a prototype and are not guaranteed to be released. They are not intended for production use or covered by any SLA, support obligation, or deprecation policy and might be subject to backward-incompatible changes.
**Kindly drop us a note before you run any scale tests.**
**Do not hesitate to contact vertexai-prediction-preview-feedback@google.com if you have any questions or run into any issues.**
The projects need to be added to the allowlist in order to deploy PyTorch models using Vertex AI Prediction pre-built PyTorch images. If you are interested in the feature, please send an email to vertexai-prediction-preview-feedback@google.com to provide your project numbers OR project ids.
## Overview
In the PyTorch on Google Cloud series of blog posts, we aim to share how to deploy PyTorch models at scale on [Vertex AI](https://cloud.google.com/vertex-ai).
This tutorial on text classification shows how to deploy a PyTorch based text classification model on [Vertex AI](https://cloud.google.com/vertex-ai/docs/start/client-libraries#python) using Vertex SDK and [`gcloud ai`](https://cloud.google.com/sdk/gcloud/reference/beta/ai).
## Notebooks
| <h4>Notebook</h4> | <h4>Description</h4> |
| :-------- | :------- |
| [pytorch-text-classification-vertex-ai-deploy.ipynb](./pytorch-text-classification-vertex-ai-deploy.ipynb) | Notebook to show deploying a PyTorch model on Vertex AI |
## Folders
| <h4>Folder Name</h4> | <h4>Description</h4> |
| :-------- | :------- |
| [`predictor`](./predictor) | Folder with custom prediction handler to deploy a PyTorch model to Vertex Prediction. In the [notebook](./pytorch-text-classification-vertex-ai-deploy.ipynb), this folder is used for deploying a PyTorch model on Vertex AI using Vertex Prediction pre-built PyTorch images |
" - [Run Training Locally in the Notebook](#Training-locally-in-the-notebook)\n",
" - [Run Training Job on Vertex AI](#Training-on-Vertex-AI)\n",
" - [Training with pre-built container](#Run-Custom-Job-on-Vertex-Training-with-a-pre-built-container)\n",
" - [Training with custom container](#Run-Custom-Job-on-Vertex-Training-with-custom-container)\n",
" - [Training with pre-built container](#Run-Custom-Job-on-Vertex-AI-Training-with-a-pre-built-container)\n",
" - [Training with custom container](#Run-Custom-Job-on-Vertex-AI-Training-with-custom-container)\n",
"- [Tuning](#Hyperparameter-Tuning) \n",
" - [Run Hyperparameter Tuning job on Vertex AI](#Run-Hyperparameter-Tuning-Job-on-Vertex-AI)\n",
"- [Deploying](#Deploying)\n",
" - [Deploying model on Vertex Predictions with custom container](#Deploying-model-on-Vertex-Predictions-with-custom-container)\n",
" - [Deploying model on Vertex AI Predictions with custom container](#Deploying-model-on-Vertex AI-Predictions-with-custom-container)\n",
"\n",
"### Costs \n",
"\n",
@@ -202,9 +202,9 @@
"id": "e0c1dcadc2c8"
},
"source": [
"We will be using [Vertex SDK for Python](https://cloud.google.com/vertex-ai/docs/start/client-libraries#python) to interact with Vertex AI services. The high-level `aiplatform` library is designed to simplify common data science workflows by using wrapper classes and opinionated defaults. \n",
"We will be using [Vertex AI SDK for Python](https://cloud.google.com/vertex-ai/docs/start/client-libraries#python) to interact with Vertex AI services. The high-level `aiplatform` library is designed to simplify common data science workflows by using wrapper classes and opinionated defaults. \n",
"\n",
"#### Install Vertex SDK for Python"
"#### Install Vertex AI SDK for Python"
]
},
{
@@ -658,8 +658,8 @@
},
"outputs": [],
"source": [
"datasets = load_dataset(\"imdb\")\n",
"datasets"
"dataset = load_dataset(\"imdb\")\n",
"dataset"
]
},
{
@@ -668,7 +668,7 @@
"id": "RzfPtOMoIrIu"
},
"source": [
"The `datasets` object itself is [`DatasetDict`](https://huggingface.co/docs/datasets/package_reference/main_classes.html#datasetdict), which contains one key for the training, validation and test set."
"The `dataset` object itself is [`DatasetDict`](https://huggingface.co/docs/datasets/package_reference/main_classes.html#datasetdict), which contains one key for the training, validation and test set."
]
},
{
@@ -681,12 +681,12 @@
"source": [
"print(\n",
" \"Total # of rows in training dataset {} and size {:5.2f} MB\".format(\n",
"### Run predictions locally with sample examples\n",
"\n",
"Using the trained model, we can predict the sentiment label for an input text after applying the preprocessing function that was used during the training. We will run the predictions locally in the notebook and later show how you can deploy the model to an endpoint using [TorchServe](https://pytorch.org/serve/) on Vertex Predictions."
"Using the trained model, we can predict the sentiment label for an input text after applying the preprocessing function that was used during the training. We will run the predictions locally in the notebook and later show how you can deploy the model to an endpoint using [TorchServe](https://pytorch.org/serve/) on Vertex AI Predictions."
]
},
{
@@ -1382,7 +1382,7 @@
"id": "f7466d414a0e"
},
"source": [
"### Run Custom Job on Vertex Training with a pre-built container"
"### Run Custom Job on Vertex AI Training with a pre-built container"
]
},
{
@@ -1395,7 +1395,7 @@
"\n",
"In this notebook, we are using Hugging Face Datasets and fine tuning a transformer model from Hugging Face Transformers Library for sentiment analysis task using PyTorch. We will use [pre-built container for PyTorch](https://cloud.google.com/vertex-ai/docs/training/pre-built-containers#pytorch) and package the training application code by adding standard Python dependencies - `transformers`, `datasets` and `tqdm` - in the `setup.py` file. \n",
"\n",
""
""
]
},
{
@@ -1569,7 +1569,7 @@
"source": [
"#### **Run custom training job on Vertex AI**\n",
"\n",
"We use [Vertex SDK for Python](https://cloud.google.com/vertex-ai/docs/start/client-libraries#client_libraries) to create and submit training job to the Vertex training service."
"We use [Vertex AI SDK for Python](https://cloud.google.com/vertex-ai/docs/start/client-libraries#client_libraries) to create and submit training job to the Vertex AI training service."
]
},
{
@@ -1578,7 +1578,7 @@
"id": "5d2957ef04fd"
},
"source": [
"##### **Initialize the Vertex SDK for Python**"
"##### **Initialize the Vertex AI SDK for Python**"
]
},
{
@@ -1598,7 +1598,7 @@
"id": "6b0fed34b728"
},
"source": [
"##### **Configure and submit Custom Job to Vertex Training service**"
"##### **Configure and submit Custom Job to Vertex AI Training service**"
]
},
{
@@ -1609,7 +1609,7 @@
"source": [
"Configure a [Custom Job](https://cloud.google.com/vertex-ai/docs/training/create-custom-job) with the [pre-built container](https://cloud.google.com/vertex-ai/docs/training/pre-built-containers) image for PyTorch and training code packaged as Python source distribution. \n",
"\n",
"**NOTE:** When using Vertex SDK for Python for submitting a training job, it creates a [Training Pipeline](https://cloud.google.com/vertex-ai/docs/training/create-training-pipeline) which launches the Custom Job on Vertex Training service."
"**NOTE:** When using Vertex AI SDK for Python for submitting a training job, it creates a [Training Pipeline](https://cloud.google.com/vertex-ai/docs/training/create-training-pipeline) which launches the Custom Job on Vertex AI Training service."
]
},
{
@@ -1686,7 +1686,7 @@
"\n",
"You can monitor the custom job launched from Cloud Console following the link [here](https://console.cloud.google.com/vertex-ai/training/training-pipelines/) or use gcloud CLI command [`gcloud beta ai custom-jobs stream-logs`](https://cloud.google.com/sdk/gcloud/reference/beta/ai/custom-jobs/stream-logs)\n",
"\n",
""
""
]
},
{
@@ -1798,7 +1798,7 @@
"id": "c170d386492b"
},
"source": [
"### Run Custom Job on Vertex Training with custom container"
"### Run Custom Job on Vertex AI Training with custom container"
]
},
{
@@ -1807,7 +1807,7 @@
"id": "035227b6e581"
},
"source": [
"To create a [training job with custom container](https://cloud.google.com/vertex-ai/docs/training/create-custom-container?hl=hr), you define a `Dockerfile` to install or add the dependencies required for the training job. Then, you build and test your Docker image locally to verify, push the image to Container Registry and submit a Custom Job to Vertex Training service.\n",
"To create a [training job with custom container](https://cloud.google.com/vertex-ai/docs/training/create-custom-container?hl=hr), you define a `Dockerfile` to install or add the dependencies required for the training job. Then, you build and test your Docker image locally to verify, push the image to Container Registry and submit a Custom Job to Vertex AI Training service.\n",
"\n",
""
"##### **Initialize the Vertex AI SDK for Python**"
]
},
{
@@ -1988,11 +1988,11 @@
"id": "abf1fa4085cb"
},
"source": [
"##### **Configure and submit Custom Job to Vertex Training service**\n",
"##### **Configure and submit Custom Job to Vertex AI Training service**\n",
"\n",
"Configure a [Custom Job](https://cloud.google.com/vertex-ai/docs/training/create-custom-job) with the [custom container](https://cloud.google.com/vertex-ai/docs/training/create-custom-container) image with training code and other dependencies\n",
"\n",
"**NOTE:** When using Vertex SDK for Python for submitting a training job, it creates a [Training Pipeline](https://cloud.google.com/vertex-ai/docs/training/create-training-pipeline) which launches the Custom Job to train on Vertex Training."
"**NOTE:** When using Vertex AI SDK for Python for submitting a training job, it creates a [Training Pipeline](https://cloud.google.com/vertex-ai/docs/training/create-training-pipeline) which launches the Custom Job to train on Vertex AI Training."
]
},
{
@@ -2044,7 +2044,7 @@
},
"outputs": [],
"source": [
"# submit the custom job to Vertex training service\n",
"# submit the custom job to Vertex AI training service\n",
"model = job.run(\n",
" replica_count=1,\n",
" machine_type=\"n1-standard-8\",\n",
@@ -2065,7 +2065,7 @@
"\n",
"You can monitor the custom job launched from Cloud Console following the link [here](https://console.cloud.google.com/vertex-ai/training/training-pipelines/) or use gcloud CLI command [`gcloud beta ai custom-jobs stream-logs`](https://cloud.google.com/sdk/gcloud/reference/beta/ai/custom-jobs/stream-logs)\n",
"\n",
""
""
]
},
{
@@ -2148,11 +2148,11 @@
"id": "ba6122f929e3"
},
"source": [
"The training application code for fine-tuning a transformer model for sentiment analysis task uses hyperparameters such as learning rate and weight decay. These hyperparameters control the behavior of the training algorithm and can have a significant effect on the performance of the resulting model. This part of the notebook show how you can automate tuning these hyperparameters with Vertex Training service.\n",
"The training application code for fine-tuning a transformer model for sentiment analysis task uses hyperparameters such as learning rate and weight decay. These hyperparameters control the behavior of the training algorithm and can have a significant effect on the performance of the resulting model. This part of the notebook show how you can automate tuning these hyperparameters with Vertex AI Training service.\n",
"\n",
"We submit a [Hyperparameter Tuning job](https://cloud.google.com/vertex-ai/docs/training/hyperparameter-tuning-overview) to Vertex Training service by packaging the training application code and dependencies in a Docker container and push the container to Google Container Registry, similar to running a Custom Job on Vertex AI with Custom Container.\n",
"We submit a [Hyperparameter Tuning job](https://cloud.google.com/vertex-ai/docs/training/hyperparameter-tuning-overview) to Vertex AI Training service by packaging the training application code and dependencies in a Docker container and push the container to Google Container Registry, similar to running a Custom Job on Vertex AI with Custom Container.\n",
"\n",
""
""
]
},
{
@@ -2163,7 +2163,7 @@
"source": [
"### How hyperparameter tuning works in Vertex AI?\n",
"\n",
"Following are the high level steps involved in running a Hyperparameter Tuning job on Vertex Training service:\n",
"Following are the high level steps involved in running a Hyperparameter Tuning job on Vertex AI Training service:\n",
"\n",
"- You define the hyperparameters to tune the model along with the metric (or goal) to optimize\n",
"- Vertex AI runs multiple trials of your training application with the hyperparameters and limits you specified - maximum number of trials to run and number of parallel trials. \n",
@@ -2297,7 +2297,7 @@
"source": [
"### Run Hyperparameter Tuning Job on Vertex AI\n",
"\n",
"Before submitting the hyperparameter tuning job to Vertex AI, push the custom container image with training application to Google Cloud Container Registry and then submit the job to Vertex AI. We will be using the same image used for running Custom Job on Vertex Training service."
"Before submitting the hyperparameter tuning job to Vertex AI, push the custom container image with training application to Google Cloud Container Registry and then submit the job to Vertex AI. We will be using the same image used for running Custom Job on Vertex AI Training service."
]
},
{
@@ -2326,7 +2326,7 @@
"id": "f60fab07d67c"
},
"source": [
"##### **Initialize the Vertex SDK for Python**"
"##### **Initialize the Vertex AI SDK for Python**"
]
},
{
@@ -2346,7 +2346,7 @@
"id": "6652aa63ddff"
},
"source": [
"##### **Configure and submit Hyperparameter Tuning Job to Vertex Training service**\n",
"##### **Configure and submit Hyperparameter Tuning Job to Vertex AI Training service**\n",
"\n",
"Configure a [Hyperparameter Tuning Job](https://cloud.google.com/vertex-ai/docs/training/using-hyperparameter-tuning) with the [custom container](https://cloud.google.com/vertex-ai/docs/training/create-custom-container) image with training code and other dependencies.\n",
"\n",
@@ -2374,7 +2374,7 @@
"id": "9d46db3a8b23"
},
"source": [
"Define the training arguments with `hp-tune` argument set to `y` so that training application code can report metrics to Vertex"
"Define the training arguments with `hp-tune` argument set to `y` so that training application code can report metrics to Vertex AI"
]
},
{
@@ -2548,7 +2548,7 @@
"\n",
"You can monitor the hyperparameter tuning job launched from Cloud Console following the link [here](https://console.cloud.google.com/vertex-ai/training/hyperparameter-tuning-jobs/) or use gcloud CLI command [`gcloud beta ai custom-jobs stream-logs`](https://cloud.google.com/sdk/gcloud/reference/beta/ai/custom-jobs/stream-logs)\n",
"\n",
""
""
]
},
{
@@ -2557,7 +2557,7 @@
"id": "ba934b434f03"
},
"source": [
"After the job is finished, you can view and format the results of the hyperparameter tuning Trials (run by Vertex Training service) as a Pandas dataframe"
"After the job is finished, you can view and format the results of the hyperparameter tuning Trials (run by Vertex AI Training service) as a Pandas dataframe"
]
},
{
@@ -2612,7 +2612,7 @@
"id": "5dbccb2b7d32"
},
"source": [
"Now from the results of Trials, you can pick the best performing Trial to deploy to Vertex Predictions"
"Now from the results of Trials, you can pick the best performing Trial to deploy to Vertex AI Predictions"
"Deploying a PyTorch model on [Vertex Predictions](https://cloud.google.com/vertex-ai/docs/predictions/getting-predictions) requires to use a custom container that serves online predictions. You will deploy a container running [PyTorch's TorchServe](https://pytorch.org/serve/) tool in order to serve predictions from a fine-tuned transformer model from Hugging Face Transformers for sentiment analysis task. You can then use Vertex Predictions to classify sentiment of input texts. \n",
"Deploying a PyTorch model on [Vertex AI Predictions](https://cloud.google.com/vertex-ai/docs/predictions/getting-predictions) requires to use a custom container that serves online predictions. You will deploy a container running [PyTorch's TorchServe](https://pytorch.org/serve/) tool in order to serve predictions from a fine-tuned transformer model from Hugging Face Transformers for sentiment analysis task. You can then use Vertex AI Predictions to classify sentiment of input texts. \n",
"\n",
"### Deploying model on Vertex Predictions with custom container\n",
"### Deploying model on Vertex AI Predictions with custom container\n",
"\n",
"To use a custom container to serve predictions from a PyTorch model, you must provide Vertex AI with a Docker container image that runs an HTTP server, such as TorchServe in this case. Please refer to [documentation](https://cloud.google.com/vertex-ai/docs/predictions/custom-container-requirements) that describes the container image requirements to be compatible with Vertex Predictions.\n",
"To use a custom container to serve predictions from a PyTorch model, you must provide Vertex AI with a Docker container image that runs an HTTP server, such as TorchServe in this case. Please refer to [documentation](https://cloud.google.com/vertex-ai/docs/predictions/custom-container-requirements) that describes the container image requirements to be compatible with Vertex AI Predictions.\n",
"\n",
"\n",
"\n",
"\n",
"Essentially, to deploy a PyTorch model on Vertex Predictions following are the steps:\n",
"Essentially, to deploy a PyTorch model on Vertex AI Predictions following are the steps:\n",
"\n",
"1. Package the trained model artifacts including [default](https://pytorch.org/serve/#default-handlers) or [custom](https://pytorch.org/serve/custom_service.html) handlers by creating an archive file using [Torch model archiver](https://github.com/pytorch/serve/tree/master/model-archiver)\n",
"2. Build a [custom container](https://cloud.google.com/vertex-ai/docs/predictions/custom-container-requirements) compatible with Vertex Predictions to serve the model using Torchserve\n",
"3. Upload the model with custom container image to serve predictions as a Vertex Model resource\n",
"4. Create a Vertex Endpoint and [deploy the model](https://cloud.google.com/vertex-ai/docs/predictions/deploy-model-api) resource"
"2. Build a [custom container](https://cloud.google.com/vertex-ai/docs/predictions/custom-container-requirements) compatible with Vertex AI Predictions to serve the model using Torchserve\n",
"3. Upload the model with custom container image to serve predictions as a Vertex AI Model resource\n",
"4. Create a Vertex AI Endpoint and [deploy the model](https://cloud.google.com/vertex-ai/docs/predictions/deploy-model-api) resource"
]
},
{
@@ -3048,10 +3048,13 @@
"FROM pytorch/torchserve:latest-cpu\n",
"\n",
"# install dependencies\n",
"RUN python3 -m pip install --upgrade pip\n",
"RUN pip3 install transformers\n",
"\n",
"USER model-server\n",
"\n",
"# copy model artifacts, custom handler and other dependencies\n",
"#### **Run the container locally** ***[Optional]***\n",
"\n",
"Before push the container image to Container Registry to use it with Vertex Predictions, you can run it as a container in your local environment to verify that the server works as expected"
"Before push the container image to Container Registry to use it with Vertex AI Predictions, you can run it as a container in your local environment to verify that the server works as expected"
]
},
{
@@ -3268,9 +3271,9 @@
"id": "69477b3a00c0"
},
"source": [
"#### **Deploying the serving container to Vertex Predictions**\n",
"#### **Deploying the serving container to Vertex AI Predictions**\n",
"\n",
"We create a model resource on Vertex AI and deploy the model to a Vertex Endpoints. You must deploy a model to an endpoint before using the model. The deployed model runs the custom container image to serve predictions. "
"We create a model resource on Vertex AI and deploy the model to a Vertex AI Endpoints. You must deploy a model to an endpoint before using the model. The deployed model runs the custom container image to serve predictions. "
]
},
{
@@ -3301,7 +3304,7 @@
"id": "a3da91e19af4"
},
"source": [
"##### **Initialize the Vertex SDK for Python**"
"##### **Initialize the Vertex AI SDK for Python**"
]
},
{
@@ -3438,7 +3441,7 @@
"id": "bc4673478269"
},
"source": [
"#### **Invoking the Endpoint with deployed Model using Vertex SDK to make predictions**"
"#### **Invoking the Endpoint with deployed Model using Vertex AI SDK to make predictions**"
]
},
{
@@ -3488,7 +3491,7 @@
"source": [
"##### **Formatting input for online prediction**\n",
"\n",
"For online prediction requests, the prediction input instances must be formatted as JSON with base64 encoding as shown here:\n",
"This notebook uses [Torchserve's KServe based inference API](https://pytorch.org/serve/inference_api.html#kserve-inference-api) which is also [Vertex AI Predictions compatible format](https://cloud.google.com/vertex-ai/docs/predictions/custom-container-requirements#prediction). For online prediction requests, format the prediction input instances as JSON with base64 encoding as shown here:\n",
"\n",
"```\n",
"[\n",
@@ -3561,9 +3564,9 @@
},
"source": [
"##### ***[Optional]*** **Make prediction requests using gcloud CLI**\n",
"You can also call the Vertex Endpoint to make predictions using [`gcloud beta ai endpoints predict`](https://cloud.google.com/sdk/gcloud/reference/beta/ai/endpoints/predict). \n",
"You can also call the Vertex AI Endpoint to make predictions using [`gcloud beta ai endpoints predict`](https://cloud.google.com/sdk/gcloud/reference/beta/ai/endpoints/predict). \n",
"\n",
"The following cell shows how to make a prediction request to Vertex Endpoints using `gcloud` CLI: "
"The following cell shows how to make a prediction request to Vertex AI Endpoints using `gcloud` CLI: "
"ENABLE_CACHING = False # Whether to enable execution caching for the pipeline.\n",
@@ -635,7 +660,7 @@
"source": [
"#### Run unit tests on the Generator component\n",
"\n",
"Before running the command, fill in `RAW_DATA_PATH` in [`src/generator/test_generator_component.py`](src/generator/test_generator_component.py)."
"Before running the command, you should update the `RAW_DATA_PATH` in [`src/generator/test_generator_component.py`](src/generator/test_generator_component.py)."
]
},
{
@@ -713,12 +738,12 @@
"TRAINING_ARTIFACTS_DIR = (\n",
" f\"{BUCKET_NAME}/artifacts\" # Root directory for training artifacts.\n",
")\n",
"TRAINING_REPLICA_COUNT = \"1\" # Number of replica to run the custom training job.\n",
"TRAINING_REPLICA_COUNT = 1 # Number of replica to run the custom training job.\n",
"TRAINING_MACHINE_TYPE = (\n",
" \"n1-standard-4\" # Type of machine to run the custom training job.\n",
")\n",
"TRAINING_ACCELERATOR_TYPE = \"ACCELERATOR_TYPE_UNSPECIFIED\" # Type of accelerators to run the custom training job.\n",
"TRAINING_ACCELERATOR_COUNT = \"0\" # Number of accelerators for the custom training job."
"TRAINING_ACCELERATOR_COUNT = 0 # Number of accelerators for the custom training job."
]
},
{
@@ -769,8 +794,12 @@
"TRAINED_POLICY_DISPLAY_NAME = (\n",
" \"movielens-trained-policy\" # Display name of the uploaded and deployed policy.\n",
")\n",
"TRAFFIC_SPLIT = {\"0\": 100}\n",
"ENDPOINT_DISPLAY_NAME = \"movielens-endpoint\" # Display name of the prediction endpoint.\n",
"ENDPOINT_MACHINE_TYPE = \"n1-standard-4\" # Type of machine of the prediction endpoint."
"ENDPOINT_MACHINE_TYPE = \"n1-standard-4\" # Type of machine of the prediction endpoint.\n",
"ENDPOINT_REPLICA_COUNT = 1 # Number of replicas of the prediction endpoint.\n",
"ENDPOINT_ACCELERATOR_TYPE = \"ACCELERATOR_TYPE_UNSPECIFIED\" # Type of accelerators to run the custom training job.\n",
"ENDPOINT_ACCELERATOR_COUNT = 0 # Number of accelerators for the custom training job."
The [official](https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/main/notebooks/official) folder contains notebooks organized by Google Cloud product. These are tested weekly and maintained by Google.
The [community](https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/main/notebooks/community) folder contains notebooks that may be created by Google or external contributors. They are not necessary maintained.
Contributions to the repo should use the [notebook template](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/main/notebooks/notebook_template.ipynb) as a starting point.
"If you are in a live tutorial session, you might be using a shared test account or project. To avoid name collisions between users on resources created, you create a uuid for each instance session, and append it onto the name of resources you create in this tutorial."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "3ee72715c0fd"
},
"outputs": [],
"source": [
"import random\n",
"import string\n",
"\n",
"\n",
"# Generate a uuid of a specifed length(default=8)\n",
"# Wait for LRO to finish and get the LRO result.\n",
"print(create_lro.result())"
" # Wait for LRO to finish and get the LRO result.\n",
" print(create_lro.result())\n",
"except Exception as e:\n",
" print(e)"
]
},
{
@@ -574,7 +597,7 @@
"id": "ag8pCQ7rNjVf"
},
"source": [
"You can use [GetFeaturestore](https://cloud.google.com/vertex-ai/docs/reference/rpc/google.cloud.aiplatform.v1beta1#google.cloud.aiplatform.v1beta1.FeaturestoreService.GetFeaturestore) or [ListFeaturestores](https://cloud.google.com/vertex-ai/docs/reference/rpc/google.cloud.aiplatform.v1beta1#google.cloud.aiplatform.v1beta1.FeaturestoreService.ListFeaturestores) to check if the Featurestore was successfully created. The following example gets the details of the Featurestore.\n"
"You can use [GetFeaturestore](https://cloud.google.com/vertex-ai/docs/reference/rpc/google.cloud.aiplatform.v1#google.cloud.aiplatform.v1.FeaturestoreService.GetFeaturestore) or [ListFeaturestores](https://cloud.google.com/vertex-ai/docs/reference/rpc/google.cloud.aiplatform.v1#google.cloud.aiplatform.v1.FeaturestoreService.ListFeaturestores) to check if the Featurestore was successfully created. The following example gets the details of the Featurestore.\n"
]
},
{
@@ -590,6 +613,41 @@
")"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "018ab19d934f"
},
"source": [
"Auto scaling is available in v1 since v1.11. Below is the example for the `CreateFeaturestoreRequest` with auto-scaling, use it with `aiplatform_v1.FeaturestoreServiceClient` to create Featurestore:"
" # Similarly, wait for EntityType creation operation.\n",
" print(movies_entity_type_lro.result())\n",
"except Exception as e:\n",
" print(e)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "dPkT7KDuEvWv"
},
"source": [
"Feature [monitoring](https://cloud.google.com/vertex-ai/docs/featurestore/monitoring) is in preview, so you need to use v1 Python. Import feature analysis is only available through SDK for now."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "9kiqrBN6E28r"
},
"outputs": [],
"source": [
"from google.cloud.aiplatform_v1 import \\\n",
" FeaturestoreServiceClient as v1_FeaturestoreServiceClient\n",
"from google.cloud.aiplatform_v1.types import entity_type as v1_entity_type_pb2\n",
"Feature [monitoring](https://cloud.google.com/vertex-ai/docs/featurestore/monitoring) is in preview, so you need to use v1beta1 Python. The easiest way to set this for now is using [console UI](https://console.cloud.google.com/vertex-ai/features). For completeness, below is example to do this using v1beta1 SDK\n"
"The easiest way to set up snapshot analysis for now is using [console UI](https://console.cloud.google.com/vertex-ai/features). For completeness, below is example to do this using v1 SDK.\n",
"\n",
"You can view monitoring statistics on [console UI](https://console.cloud.google.com/vertex-ai/features)."
" description=\"The average rating for the movie, range is [1.0-5.0]\",\n",
" ),\n",
" feature_id=\"average_rating\",\n",
" ),\n",
" ],\n",
" ).result()\n",
"except Exception as e:\n",
" print(e)"
]
},
{
@@ -782,8 +919,8 @@
"source": [
"## Search created features\n",
"\n",
"While the [ListFeatures](https://cloud.google.com/vertex-ai/docs/reference/rpc/google.cloud.aiplatform.v1beta1#google.cloud.aiplatform.v1beta1.FeaturestoreService.ListFeatures) method allows you to easily view all features of a single\n",
"entity type, the [SearchFeatures](https://cloud.google.com/vertex-ai/docs/reference/rpc/google.cloud.aiplatform.v1beta1#google.cloud.aiplatform.v1beta1.FeaturestoreService.SearchFeatures) method searches across all featurestores\n",
"While the [ListFeatures](https://cloud.google.com/vertex-ai/docs/reference/rpc/google.cloud.aiplatform.v1#google.cloud.aiplatform.v1.FeaturestoreService.ListFeatures) method allows you to easily view all features of a single\n",
"entity type, the [SearchFeatures](https://cloud.google.com/vertex-ai/docs/reference/rpc/google.cloud.aiplatform.v1#google.cloud.aiplatform.v1.FeaturestoreService.SearchFeatures) method searches across all featurestores\n",
"and entity types in a given location (such as `us-central1`). This can help you discover features that were created by someone else.\n",
"\n",
"You can query based on feature properties including feature ID, entity type ID,\n",
@@ -985,6 +1122,8 @@
" ],\n",
" feature_time_field=\"update_time\",\n",
" worker_count=1,\n",
" # Default is False. If True, the import feature analysis won't happen for this specific operation.\n",
"lets you serve feature values for small batches of entities. It's designed for latency-sensitive service, such as online model prediction. For example, for a movie service, you might want to quickly shows movies that the current user would most likely watch by using online predictions."
"* [Load the dataset from Cloud Storage](#section-6)\n",
"* [Data Analysis](#section-7)\n",
"* [Preprocess the data for training](#section-8)\n",
"* [Train the model using BigQuery ML](#section-9)\n",
"* [Generate forecasts from the model](#section-10)\n",
"* [Interpret the results to choose the best price](#section-11)\n",
"* [Clean Up](#section-12)\n",
"\n",
"\n",
"## Overview\n",
"<a name=\"section-1\"></a>\n",
"This notebook demonstrates analysis of pricing optimization on [CDM Pricing Data](https://github.com/trifacta/trifacta-google-cloud/tree/main/design-pattern-pricing-optimization) and automating the workflow using Vertex AI Workbench's managed notebooks.\n",
"\n",
"<b>Note</b>: This notebook is designed to run on managed notebooks instance of Vertex AI Workbench. Some components of this notebook may not work in other notebook environments.\n",
"\n",
"## Dataset\n",
"<a name=\"section-2\"></a>\n",
"The dataset used in this notebook is a part of the [CDM Pricing Data](https://github.com/trifacta/trifacta-google-cloud/blob/main/design-pattern-pricing-optimization/CDM_Pricing_large_table.csv) which consists of products sales information on specified dates.\n",
"\n",
"## Objective\n",
"<a name=\"section-3\"></a>\n",
"The objective of this notebook is to build a pricing optimization model using Vertex AI on GCP. The following steps have been followed in this usecase : \n",
"\n",
"- Load the required dataset from a Cloud Storage bucket.\n",
"- Analyze the fields present in the dataset.\n",
"- Process the data to build a model.\n",
"- Build a BigQuery ML forecast model on the processed data.\n",
"- Get forecasted values from the BigQuery ML model.\n",
"- Interpret the forecasts to identify best prices.\n",
"- Clean up.\n",
"\n",
"\n",
"## Costs\n",
"<a name=\"section-4\"></a>\n",
"This tutorial uses the following billable components of Google Cloud:\n",
"\n",
"- Vertex AI\n",
"- Bigquery\n",
"- Cloud Storage\n",
"\n",
"\n",
"Learn about [Vertex AI\n",
"pricing](https://cloud.google.com/vertex-ai/pricing), [Bigquery pricing](https://cloud.google.com/bigquery/pricing) and [Cloud Storage\n",
"pricing](https://cloud.google.com/storage/pricing), and use the [Pricing\n",
"We will build a forecast model on this data and thus determine the best price for a product. For this type of model, we may not be using many fields but just the sales and price related ones. For the current execrcise, we will just focus on the following fields :\n",
"- `Product_ID`\n",
"- `Customer_Hierarchy`\n",
"- `Fiscal_Date`\n",
"- `List_Price_Converged`\n",
"- `Invoiced_quantity_in_Pieces`\n",
"- `Net_Sales`\n",
"\n",
"\n",
"\n",
"## Data Analysis\n",
"<a name=\"section-7\"></a>\n",
"\n",
"First, we will explore the data and distributions.\n",
"Check the column types and null values in the dataframe."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "f54c445a1288"
},
"outputs": [],
"source": [
"df.info()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "cd817b414c4d"
},
"source": [
"This data description reveals that there are no null values in the data. Also, the field `Fiscal_Date` which is a date field is loaded as an object type. \n",
"# plot a scatterplot to visualize the changes\n",
"sns.scatterplot(\n",
" x=\"price_change_perc\",\n",
" y=\"order_change_perc\",\n",
" data=df_aggr,\n",
" hue=\"Product_ID\",\n",
" legend=False,\n",
")\n",
"plt.title(\"Percentage of change in price vs order\")\n",
"plt.show()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "8259e916fe25"
},
"source": [
"For most of the products, we see that the percentage change in orders are high where the percentage changes in the prices are low. This suggests that too much change in the prices can affect the number of orders. \n",
"\n",
"**Note**: There seem to be some outliers in the data as percentage changes greater than 800 are found and as evident from the box plots made earlier. In the current exercise, we will not take any manual measures to deal with outliers as we will create a BigQuery ML timeseries model that already deals with outliers.\n",
"\n",
"## Preprocess the data for training\n",
"<a name=\"section-8\"></a>\n",
"\n",
"Check which `Product_ID`s that have the maximum orders."
"From the above result, we can infer the following :\n",
"\n",
"- Under **Food** category, **SKU 62** has maximum orders.\n",
"- Under **Manufacturing** category, **SKU 17** has maximum orders.\n",
"- Under **Paper** category, **SKU 107** has maximum orders.\n",
"- Under **Publishing** category, **SKU 8** has maximum orders.\n",
"- Under **Utilities** category, **SKU 140** has maximum orders.\n",
"\n",
"Given there are too many ids and only a few records for most of them, we will consider only the above `Product_ID`s for which there are maximum number of orders. \n",
"\n",
"**Note**: The `Invoiced_quantity_in_Pieces` field seem to be a *float* type rather than an *int* type as it should be. This could be probably because of the data itself might be averaged in the first place."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "2dbc0d64d157"
},
"source": [
"Check the various prices available for these `Product_ID`s."
"In the publishing category, `Product_ID` `SKU 8` and `SKU 17` has less than or equal to two different prices in the entire data and so we will exclude them and consider the rest for building the forecast model. The idea here is to train a forecast model on the timeseries data for products with different prices.\n",
"\n",
"Join the data for all the `Product_ID`s into one dataframe and remove duplicate records."
"Train an [Arima-Plus](https://cloud.google.com/bigquery-ml/docs/reference/standard-sql/bigqueryml-syntax-create-time-series) model on the data using BigQuery ML."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "cded27507891"
},
"source": [
"#@bigquery\n",
"create or replace model pricing_optimization.bqml_arima\n",
" plt.title(\"Price vs. Average Sales for \" + i)\n",
" plt.show()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "67ff3acc74a5"
},
"source": [
"Based on the plots for price vs. the average forecasted orders, it can be said that to avail the maximum orders, each of the considered `Product_ID`s can follow the below prices :\n",
"- SKU 107's price range can be from 4.44 - 4.73 units\n",
"- SKU 140's price can be 1.95 units\n",
"- SKU 62's price can be 4.23 units\n",
"\n",
"\n",
"## Clean Up\n",
"<a name=\"section-12\"></a>\n",
"\n",
"To clean up all Google Cloud resources used in this project, you can [delete the Google Cloud project](https://cloud.google.com/resource-manager/docs/creating-managing-projects#shutting_down_projects) you used for the tutorial.\n",
"\n",
"Otherwise, you can delete the individual resources you created in this tutorial. The following code deletes the entire dataset."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "d78908b8134d"
},
"outputs": [],
"source": [
"# Construct a BigQuery client object.\n",
"client = bigquery.Client()\n",
"\n",
"# TODO(developer): Set model_id to the ID of the model to fetch.\n",
**The following steps are required, regardless of your notebook environment.**
1. [Select or create a Google Cloud project](https://console.cloud.google.com/cloud-resource-manager). When you first create an account, you get a $300 free credit towards your compute/storage costs.
1. [Make sure that billing is enabled for your project](https://cloud.google.com/billing/docs/how-to/modify-project).
1. [Enable the Vertex AI, BigQuery, Compute Engine and Cloud Storage APIs](https://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com,bigquery,compute_component,storage_component).
1. If you are running this notebook locally, you need to install the [Cloud SDK](https://cloud.google.com/sdk).
1. Enter your project ID in the cell below. Then run the cell to make sure the
Cloud SDK uses the right project for all the commands in this notebook.
**Note**: Jupyter runs lines prefixed with `!` as shell commands, and it interpolates Python variables prefixed with `$` into these commands.
### Set up your local development environment
**If you are using Colab or Vertex AI Workbench Notebooks**, your environment already meets all the requirements to run this notebook. You can skip this step.
**Otherwise**, make sure your environment meets this notebook's requirements. You need the following:
- The Cloud Storage SDK
- Python 3
- virtualenv
- Jupyter notebook running in a virtual environment with Python 3
The Cloud Storage guide to [Setting up a Python development environment](https://cloud.google.com/python/setup) and the [Jupyter installation guide](https://jupyter.org/install) provide detailed instructions for meeting these requirements. The following steps provide a condensed set of instructions:
1. [Install and initialize the SDK](https://cloud.google.com/sdk/docs/).
3. [Install virtualenv](https://cloud.google.com/python/setup#installing_and_using_virtualenv) and create a virtual environment that uses Python 3. Activate the virtual environment.
4. To install Jupyter, run `pip3 install jupyter` on the command-line in a terminal shell.
5. To launch Jupyter, run `jupyter notebook` on the command-line in a terminal shell.
6. Open this notebook in the Jupyter Notebook Dashboard.
@@ -22,31 +22,31 @@ The first stage in MLOps is the collection and preparation for the purpose of de
- Data is preprocessed for training and evaluation using Dataflow.
- Data augmentation is performed on-the-fly and is coupled with model feeding.
<img src='stage1.jpg'>
<img src='stage1v2.png'>
## Notebooks
### Get Started
[Get Started with BQ datasets](get_started_bq_datasets.ipynb)
[Get started with Dataflow](community/ml_ops/stage1/get_started_dataflow.ipynb)
In this tutorial, you learn how to use `Dataflow` for training with `Vertex AI`.
```
The steps performed include:
- Create a Vertex AI `Dataset` resource from `BigQuery` table -- compatible for `AutoML` training.
- Extract a copy of the dataset from `BigQuery` to a CSV file in Cloud Storage -- compatible for `AutoML` or custom training.
- Select rows from a `BigQuery` dataset into a `pandas` dataframe -- compatible for custom training.
- Select rows from a `BigQuery` dataset into a `tf.data.Dataset` -- compatible for custom training `TensorFlow` models.
- Select rows from extracted CSV files into a `tf.data.Dataset` -- compatible for custom training `TensorFlow` models.
- Create a `BigQuery` dataset from CSV files.
- Extract data from `BigQuery` table into a `DMatrix` -- compatible for custom training `XGBoost` models.
```
- Offline preprocessing of data:
- Serially - w/o dataflow
- Parallel - with dataflow
- Upstream preprocessing of data:
- tabular data
- image data
[Get Started with Vertex datasets](get_started_vertex_datasets.ipynb)
[Get started with Vertex AI datasets](community/ml_ops/stage1/get_started_vertex_datasets.ipynb)
In this tutorial, you learn how to use `Vertex AI Dataset` for training with `Vertex AI`.
```
The steps performed include:
- Create a Vertex AI `Dataset` resource for:
- image data
- text data
@@ -61,28 +61,54 @@ The steps performed include:
- Detect anomalies in new data using TensorFlow Data Validation.
- Generate a TFRecord feature specification using TensorFlow Transform from the data schema.
- Export a dataset and convert to TFRecords.
```
[Get Started with Dataflow](get_started_dataflow.ipynb)
[Get started with BigQuery datasets](community/ml_ops/stage1/get_started_bq_datasets.ipynb)
In this tutorial, you learn how to use `BigQuery` as a dataset for training with `Vertex AI`.
```
The steps performed include:
- Offline preprocessing of data:
- Serially - w/o dataflow
- Parallel - with dataflow
- Upstream preprocessing of data:
- tabular data
- image data
```
- Create a Vertex AI `Dataset` resource from `BigQuery` table -- compatible for `AutoML` training.
- Extract a copy of the dataset from `BigQuery` to a CSV file in Cloud Storage -- compatible for `AutoML` or custom training.
- Select rows from a `BigQuery` dataset into a `pandas` dataframe -- compatible for custom training.
- Select rows from a `BigQuery` dataset into a `tf.data.Dataset` -- compatible for custom training `TensorFlow` models.
- Select rows from extracted CSV files into a `tf.data.Dataset` -- compatible for custom training `TensorFlow` models.
- Create a `BigQuery` dataset from CSV files.
- Extract data from `BigQuery` table into a `DMatrix` -- compatible for custom training `XGBoost` models.
[Get started with Vertex AI Data Labeling](community/ml_ops/stage1/get_started_with_data_labeling.ipynb)
In this tutorial, you learn how to use the `Vertex AI Data Labeling` service.
The steps performed include:
- Create a Specialist Pool for data labelers.
- Create a data labeling job.
- Submit the data labeling job.
- List data labeling jobs.
- Cancel a data labeling job.
[Create an unlabelled Vertex AI AutoML text entity extraction dataset from PDFs using Vision API](community/ml_ops/stage1/get_started_with_visionapi_and_vertex_datasets.ipynb)
In this tutorial, you learn to use `Vision API` to extract text from PDF files stored on a Cloud Storage bucket. You then process the results and create an unlabelled `Vertex AI Dataset`, compatible with `AutoML`, for text entity extraction.
The steps performed include:
1. Using `Vision API` to perform Optical Character Recognition (OCR) to extract text from PDF files.
2. Processing the results and saving them to text files.
3. Generating a `Vertex AI Dataset` import file.
4. Creating a new unlabelled text entity extraction `Vertex AI Dataset` resource in `Vertex AI`.
### E2E Stage Example
[Stage 1: Data Management](mlops_data_management.ipynb)
```
The steps performed include:
- Explore and visualize the data.
- Create a Vertex AI `Dataset` resource from `BigQuery` table -- for AutoML training.
- Extract a copy of the dataset to a CSV file in Cloud Storage.
@@ -92,4 +118,3 @@ The steps performed include:
- Generate a TFRecord feature specification using TensorFlow Data Validation from the data schema.
- Preprocess a portion of the BigQuery data using `Dataflow` -- for custom training.
" <img src=\"https://lh3.googleusercontent.com/UiNooY4LUgW_oTvpsNhPpQzsstV5W8F7rYgxgGBD85cWJoLmrOzhVs_ksK_vgx40SHs7jCqkTkCk=e14-rj-sc0xffffff-h130-w32\" alt=\"Vertex AI logo\">\n",
" Open in Vertex AI Workbench\n",
" </a>\n",
" </td>\n",
"</table>\n",
@@ -59,17 +64,6 @@
"This tutorial demonstrates how to use Vertex AI for E2E MLOps on Google Cloud in production. This tutorial covers stage 1 : data management: get started with BigQuery datasets."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "dataset:gsod,lrg"
},
"source": [
"### Dataset\n",
"\n",
"The dataset used for this tutorial is the GSOD dataset from [BigQuery public datasets](https://cloud.google.com/bigquery/public-data). The version of the dataset you use only the fields year, month and day to predict the value of mean daily temperature (mean_temp)."
]
},
{
"cell_type": "markdown",
"metadata": {
@@ -104,7 +98,7 @@
"source": [
"### Recommendations\n",
"\n",
"When doing E2E MLOps on Google Cloud, the following best practices with structured (tabular) data in BigQuery:\n",
"When doing E2E MLOps on Google Cloud, following are the best practices when dealing with structured (tabular) data in BigQuery:\n",
"\n",
"- For AutoML training:\n",
" - Create a managed dataset with Vertex AI `TabularDataset`.\n",
@@ -124,20 +118,47 @@
" - Within the generator (upstream)\n",
" - Within the model (downstream)\n",
" - XGBoost model training:\n",
" - Use BigQuery ML builtin XGBoost training.\n",
" - Use BigQuery ML built-in XGBoost training.\n",
" - Alternatively, create a DMatrix generator from CSV files extracted from BigQuery table.\n",
" - Pytorch model training:\n",
" - PyTorch model training:\n",
" - Extract the BigQuery to a pandas dataframe.\n",
" - Preprocess the data in the dataframe.\n",
" - Create a DataLoader generator from the pandas dataframe.\n",
"\n",
"\n",
"- Alternately:\n",
"- Alternatively:\n",
" - Extract the BigQuery table to CSV files.\n",
" - Preprocess the CSV files.\n",
" - Create a tf.data.Dataset generator from the CSV files."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "dataset:gsod,lrg"
},
"source": [
"### Dataset\n",
"\n",
"The dataset used for this tutorial is the GSOD dataset from [BigQuery public datasets](https://cloud.google.com/bigquery/public-data). In this version of the dataset you consider the fields year, month and day to predict the value of mean daily temperature (mean_temp)."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "9e483012a752"
},
"source": [
"### Costs\n",
"This tutorial uses billable components of Google Cloud:\n",
"\n",
"- Vertex AI\n",
"- Cloud Storage\n",
"- BigQuery\n",
"\n",
"Learn about [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing), [Cloud Storage pricing](https://cloud.google.com/storage/pricing) and [BigQuery pricing](https://cloud.google.com/bigquery/pricing) and use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to generate a cost estimate based on your projected usage."
]
},
{
"cell_type": "markdown",
"metadata": {
@@ -146,7 +167,7 @@
"source": [
"## Installations\n",
"\n",
"Install *one time* the packages for executing the MLOps notebooks."
"Install the following packages to execute this notebook."
"If you are in a live tutorial session, you might be using a shared test account or project. To avoid name collisions between users on resources created, you create a timestamp for each instance session, and append the timestamp onto the name of resources you create in this tutorial."
"**The following steps are required, regardless of your notebook environment.**\n",
"\n",
"When you submit a custom training job using the Vertex SDK, you upload a Python package\n",
"containing your training code to a Cloud Storage bucket. Vertex AI runs\n",
"the code from this package. In this tutorial, Vertex AI also saves the\n",
"trained model that results from your job in the same bucket. You can then\n",
"create an `Endpoint` resource based on this output in order to serve\n",
"online predictions.\n",
"\n",
"Set the name of your Cloud Storage bucket below. Bucket names must be globally unique across all Google Cloud projects, including those outside of your organization."
"Currently, there is no direct data feeding connector between BigQuery and the open source XGBoost.\n",
"Currently, there is no direct data feeding connector between BigQuery and the open source XGBoost. The BigQuery ML service has a built-in XGBoost training module.\n",
"\n",
"The BigQuery ML service has XGBoost training builtin.\n",
"Alernatively, you extract the data either as a pandas dataframe or as CSV files. The extracted data is then given as an input to a `DMatrix` object when training the model.\n",
"\n",
"Alernatively, you extract the data either as a pandas dataframe or as CSV files. The extracted data is then inputted to a `DMatrix` object when training the model.\n",
"\n",
"Learn more about [Getting started with builtin XGBoost](https://cloud.google.com/ai-platform/training/docs/algorithms/xgboost-start)"
"Learn more about [Getting started with built-in XGBoost](https://cloud.google.com/ai-platform/training/docs/algorithms/xgboost-start)."
]
},
{
@@ -1096,7 +889,7 @@
"source": [
"### Read pandas table into XGboost DMatrix\n",
"\n",
"Next, you load the pandas dataframe into a `DMatrix` object. XGBoost does not support non-numeric inputs. Any column that is categorical will need to be one-hot encoded prior to loading the dataframe."
"Next, you load the pandas dataframe into a `DMatrix` object. XGBoost does not support non-numeric inputs. Any column that is categorical need to be one-hot encoded prior to loading the dataframe."
" <img src=\"https://lh3.googleusercontent.com/UiNooY4LUgW_oTvpsNhPpQzsstV5W8F7rYgxgGBD85cWJoLmrOzhVs_ksK_vgx40SHs7jCqkTkCk=e14-rj-sc0xffffff-h130-w32\" alt=\"Vertex AI logo\">\n",
" Open in Vertex AI Workbench\n",
" </a>\n",
" </td>\n",
"</table>\n",
@@ -59,17 +65,6 @@
"This tutorial demonstrates how to use Vertex AI for E2E MLOps on Google Cloud in production. This tutorial covers stage 1 : data management: get started with Dataflow."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "dataset:gsod,lrg"
},
"source": [
"### Dataset\n",
"\n",
"The dataset used for this tutorial is the GSOD dataset from [BigQuery public datasets](https://cloud.google.com/bigquery/public-data). The version of the dataset you use only the fields year, month and day to predict the value of mean daily temperature (mean_temp)."
]
},
{
"cell_type": "markdown",
"metadata": {
@@ -131,6 +126,34 @@
"Alternately for AutoML tabular model training, you can reconfigure the otherwise default preprocessing."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "dataset:gsod,lrg"
},
"source": [
"### Dataset\n",
"\n",
"The dataset used for this tutorial is the GSOD dataset from [BigQuery public datasets](https://cloud.google.com/bigquery/public-data). The version of the dataset you use only the fields year, month and day to predict the value of mean daily temperature (mean_temp)."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "9e483012a752"
},
"source": [
"### Costs\n",
"This tutorial uses billable components of Google Cloud:\n",
"\n",
"- Vertex AI\n",
"- Cloud Storage\n",
"- BigQuery\n",
"- Dataflow\n",
"\n",
"Learn about [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing), [Cloud Storage pricing](https://cloud.google.com/storage/pricing), [BigQuery pricing](https://cloud.google.com/bigquery/pricing), and [Dataflow pricing](https://cloud.google.com/dataflow/pricing) and use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to generate a cost estimate based on your projected usage."
]
},
{
"cell_type": "markdown",
"metadata": {
@@ -139,7 +162,7 @@
"source": [
"## Installations\n",
"\n",
"Install *one time* the packages for executing the MLOps notebooks."
"Install the following packages to execute this notebook."
"If you are in a live tutorial session, you might be using a shared test account or project. To avoid name collisions between users on resources created, you create a timestamp for each instance session, and append the timestamp onto the name of resources you create in this tutorial."
"**The following steps are required, regardless of your notebook environment.**\n",
"\n",
"When you submit a custom training job using the Vertex SDK, you upload a Python package\n",
"containing your training code to a Cloud Storage bucket. Vertex AI runs\n",
"the code from this package. In this tutorial, Vertex AI also saves the\n",
"trained model that results from your job in the same bucket. You can then\n",
"create an `Endpoint` resource based on this output in order to serve\n",
"online predictions.\n",
"\n",
"Set the name of your Cloud Storage bucket below. Bucket names must be globally unique across all Google Cloud projects, including those outside of your organization."
" <img src=\"https://lh3.googleusercontent.com/UiNooY4LUgW_oTvpsNhPpQzsstV5W8F7rYgxgGBD85cWJoLmrOzhVs_ksK_vgx40SHs7jCqkTkCk=e14-rj-sc0xffffff-h130-w32\" alt=\"Vertex AI logo\">\n",
" Open in Vertex AI Workbench\n",
" </a>\n",
" </td> \n",
"</table>\n",
"<br/><br/><br/>"
]
@@ -133,6 +139,33 @@
" - Create a tf.data.Dataset from the TFRecords."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "533dd6fe83c8"
},
"source": [
"### Datasets\n",
"\n",
"This tutorial uses a variety of public datasets to demonstrate using a `Vertex AI` managed dataset."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "9e483012a752"
},
"source": [
"### Costs\n",
"This tutorial uses billable components of Google Cloud:\n",
"\n",
"- Vertex AI\n",
"- Cloud Storage\n",
"- BigQuery\n",
"\n",
"Learn about [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing), [Cloud Storage pricing](https://cloud.google.com/storage/pricing) and [BigQuery pricing](https://cloud.google.com/bigquery/pricing) and use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to generate a cost estimate based on your projected usage."
]
},
{
"cell_type": "markdown",
"metadata": {
@@ -141,7 +174,7 @@
"source": [
"## Installations\n",
"\n",
"Install *one time* the packages for executing the MLOps notebooks."
"Install the packages required for executing this notebook."
"**The following steps are required, regardless of your notebook environment.**\n",
"\n",
"1. [Select or create a Google Cloud project](https://console.cloud.google.com/cloud-resource-manager). When you first create an account, you get a $300 free credit towards your compute/storage costs.\n",
"\n",
"1. [Make sure that billing is enabled for your project](https://cloud.google.com/billing/docs/how-to/modify-project).\n",
"\n",
"1. [Enable the Vertex AI API](https://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com). \n",
"\n",
"1. If you are running this notebook locally, you need to install the [Cloud SDK](https://cloud.google.com/sdk).\n",
"\n",
"1. Enter your project ID in the cell below. Then run the cell to make sure the\n",
"Cloud SDK uses the right project for all the commands in this notebook.\n",
"\n",
"**Note**: Jupyter runs lines prefixed with `!` as shell commands, and it interpolates Python variables prefixed with `$` into these commands."
"**If you are using Vertex AI Workbench Notebooks**, your environment is already authenticated. Skip this step.\n",
"\n",
"**If you are using Colab**, run the cell below and follow the instructions when prompted to authenticate your account via oAuth.\n",
"\n",
"**Otherwise**, follow these steps:\n",
"\n",
"In the Cloud Console, go to the [Create service account key](https://console.cloud.google.com/apis/credentials/serviceaccountkey) page.\n",
"\n",
"**Click Create service account**.\n",
"\n",
"In the **Service account name** field, enter a name, and click **Create**.\n",
"\n",
"In the **Grant this service account access to project** section, click the Role drop-down list. Type \"Vertex\" into the filter box, and select **Vertex Administrator**. Type \"Storage Object Admin\" into the filter box, and select **Storage Object Admin**.\n",
"\n",
"Click Create. A JSON file that contains your key downloads to your local environment.\n",
"\n",
"Enter the path to your service account key as the GOOGLE_APPLICATION_CREDENTIALS variable in the cell below and run the cell."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "89788a802687"
},
"outputs": [],
"source": [
"# If you are running this notebook in Colab, run this cell and follow the\n",
"# instructions to authenticate your GCP account. This provides access to your\n",
"# Cloud Storage bucket and lets you submit training jobs and prediction\n",
"# requests.\n",
"\n",
"import os\n",
"import sys\n",
"\n",
"# If on Vertex AI Workbench, then don't execute this code\n",
"IS_COLAB = \"google.colab\" in sys.modules\n",
"if not os.path.exists(\"/opt/deeplearning/metadata/env_version\") and not os.getenv(\n",
" \"DL_ANACONDA_HOME\"\n",
"):\n",
" if \"google.colab\" in sys.modules:\n",
" from google.colab import auth as google_auth\n",
"\n",
" google_auth.authenticate_user()\n",
"\n",
" # If you are running this notebook locally, replace the string below with the\n",
" # path to your service account key and run this cell to authenticate your GCP\n",
"Next, create the `Dataset` resource using the `create` method for the `TabularDataset` class for CSV input data, which takes the following parameters:\n",
"\n",
"- `display_name`: The human readable name for the `Dataset` resource.\n",
"- `gcs_source`: A list of one or more dataset index files to import the data items into the `Dataset` resource.\n",
"\n",
"Learn more about [TabularDataset from CSV files](https://cloud.google.com/vertex-ai/docs/datasets/create-dataset-api#aiplatform_create_dataset_tabular_gcs_sample-python)"
"Next, create the `Dataset` resource using the `create` method for the `TabularDataset` class, which takes the following parameters:\n",
"Next, create the `Dataset` resource using the `create` method for the `TabularDataset` class for BigQuery table input, which takes the following parameters:\n",
"\n",
"- `display_name`: The human readable name for the `Dataset` resource.\n",
"- `gcs_source`: A list of one or more dataset index files to import the data items into the `Dataset` resource.\n",
"- `labels`: User defined metadata. In this example, you store the location of the Cloud Storage bucket containing the user defined data.\n",
"- `bq_source`: A list of one or more BigQuery tables to import the data items into the `Dataset` resource.\n",
"\n",
"Learn more about [TabularDataset from CSV files](https://cloud.google.com/vertex-ai/docs/datasets/create-dataset-api#aiplatform_create_dataset_tabular_gcs_sample-python)"
"Learn more about [TabularDataset from BigQuery table](https://cloud.google.com/vertex-ai/docs/datasets/create-dataset-api#aiplatform_create_dataset_tabular_bigquery_sample-pythonn)"
"Next, create the `Dataset` resource using the `create_from_dataframe` method for the `TabularDataset` class for pandas dataframe input, which takes the following parameters:\n",
"\n",
"- `display_name`: The human readable name for the `Dataset` resource.\n",
"- `df_source`: The pandas dataframe to import the data items into the `Dataset` resource.\n",
"- `staging_path`: The BigQuery table to store the imported data."
" <img src=\"https://lh3.googleusercontent.com/UiNooY4LUgW_oTvpsNhPpQzsstV5W8F7rYgxgGBD85cWJoLmrOzhVs_ksK_vgx40SHs7jCqkTkCk=e14-rj-sc0xffffff-h130-w32\" alt=\"Vertex AI logo\">\n",
" Open in Vertex AI Workbench\n",
" </a>\n",
" </td>\n",
"</table>\n",
"<br/><br/><br/>"
"<br/><br/><br/>\n",
"\n",
"*Note: This notebook is not supported for execution in Colab*"
]
},
{
@@ -59,17 +62,6 @@
"This tutorial demonstrates how to use Vertex AI for E2E MLOps on Google Cloud in production. This tutorial covers stage 1 : data management."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "dataset:bq,chicago,lbn"
},
"source": [
"### Dataset\n",
"\n",
"The dataset used for this tutorial is the [Chicago Taxi](https://www.kaggle.com/chicago/chicago-taxi-trips-bq). The version of the dataset you will use in this tutorial is stored in a public BigQuery table. The trained model predicts whether someone would leave a tip for a taxi fare."
]
},
{
"cell_type": "markdown",
"metadata": {
@@ -114,6 +106,34 @@
" - Preprocess the data with `Dataflow`"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "dataset:bq,chicago,lbn"
},
"source": [
"### Dataset\n",
"\n",
"The dataset used for this tutorial is the [Chicago Taxi](https://www.kaggle.com/chicago/chicago-taxi-trips-bq). The version of the dataset used in this tutorial is stored in a public BigQuery table. The trained model predicts whether someone leaves a tip for a taxi fare."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "9e483012a752"
},
"source": [
"### Costs\n",
"This tutorial uses billable components of Google Cloud:\n",
"\n",
"- Vertex AI\n",
"- Cloud Storage\n",
"- BigQuery\n",
"- Dataflow\n",
"\n",
"Learn about [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing), [Cloud Storage pricing](https://cloud.google.com/storage/pricing), [BigQuery pricing](https://cloud.google.com/bigquery/pricing), and [Dataflow pricing](https://cloud.google.com/dataflow/pricing) and use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to generate a cost estimate based on your projected usage."
]
},
{
"cell_type": "markdown",
"metadata": {
@@ -133,20 +153,33 @@
},
"outputs": [],
"source": [
"import os\n",
"\n",
"# The Vertex AI Workbench Notebook product has specific requirements\n",
"IS_WORKBENCH_NOTEBOOK = os.getenv(\"DL_ANACONDA_HOME\") and not os.getenv(\"VIRTUAL_ENV\")\n",
"**The following steps are required, regardless of your notebook environment.**\n",
"\n",
"1. [Select or create a Google Cloud project](https://console.cloud.google.com/cloud-resource-manager). When you first create an account, you get a $300 free credit towards your compute/storage costs.\n",
"\n",
"1. [Make sure that billing is enabled for your project](https://cloud.google.com/billing/docs/how-to/modify-project).\n",
"\n",
"1. [Enable the Vertex AI, BigQuery, Compute Engine and Cloud Storage APIs](https://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com,bigquery,compute_component,storage_component).\n",
"\n",
"1. If you are running this notebook locally, you need to install the [Cloud SDK](https://cloud.google.com/sdk).\n",
"\n",
"1. Enter your project ID in the cell below. Then run the cell to make sure the\n",
"Cloud SDK uses the right project for all the commands in this notebook.\n",
"\n",
"**Note**: Jupyter runs lines prefixed with `!` as shell commands, and it interpolates Python variables prefixed with `$` into these commands."
"**If you are using Vertex AI Workbench Notebooks**, your environment is already authenticated. \n",
"\n",
"**If you are using Colab**, run the cell below and follow the instructions when prompted to authenticate your account via oAuth.\n",
"\n",
"**Otherwise**, follow these steps:\n",
"\n",
"In the Cloud Console, go to the [Create service account key](https://console.cloud.google.com/apis/credentials/serviceaccountkey) page.\n",
"\n",
"1. **Click Create service account**.\n",
"\n",
"2. In the **Service account name** field, enter a name, and click **Create**.\n",
"\n",
"3. In the **Grant this service account access to project** section, click the Role drop-down list. Type \"Vertex AI\" into the filter box, and select **Vertex AI Administrator**. Type \"Storage Object Admin\" into the filter box, and select **Storage Object Admin**.\n",
"\n",
"4. Click Create. A JSON file that contains your key downloads to your local environment.\n",
"\n",
"5. Enter the path to your service account key as the GOOGLE_APPLICATION_CREDENTIALS variable in the cell below and run the cell."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "535223fa4b84"
},
"outputs": [],
"source": [
"# If you are running this notebook in Colab, run this cell and follow the\n",
"# instructions to authenticate your GCP account. This provides access to your\n",
"# Cloud Storage bucket and lets you submit training jobs and prediction\n",
"# requests.\n",
"\n",
"import os\n",
"import sys\n",
"\n",
"# If on Vertex AI Workbench, then don't execute this code\n",
"IS_COLAB = \"google.colab\" in sys.modules\n",
"if not os.path.exists(\"/opt/deeplearning/metadata/env_version\") and not os.getenv(\n",
" \"DL_ANACONDA_HOME\"\n",
"):\n",
" if \"google.colab\" in sys.modules:\n",
" from google.colab import auth as google_auth\n",
"\n",
" google_auth.authenticate_user()\n",
"\n",
" # If you are running this notebook locally, replace the string below with the\n",
" # path to your service account key and run this cell to authenticate your GCP\n",
" # account.\n",
" elif not os.getenv(\"IS_TESTING\"):\n",
" %env GOOGLE_APPLICATION_CREDENTIALS ''"
]
},
{
"cell_type": "markdown",
"metadata": {
@@ -291,7 +413,7 @@
"\n",
"**The following steps are required, regardless of your notebook environment.**\n",
"\n",
"When you submit a custom training job using the Vertex SDK, you upload a Python package\n",
"When you submit a custom training job using the Vertex AI SDK, you upload a Python package\n",
"containing your training code to a Cloud Storage bucket. Vertex AI runs\n",
"the code from this package. In this tutorial, Vertex AI also saves the\n",
"trained model that results from your job in the same bucket. You can then\n",
@@ -25,65 +25,313 @@ The second stage in MLOps is experimenting in developing one or more baseline mo
- Use the What-if-Tool (WIT) to explore how the trained model would make predictions in different scenarios.
<img src='stage2.png'>
<img src='stage2v3.png'>
<br/>
<br/>
<br/>
<img src='stage2.2v1.png'>
## Notebooks
### Get Started
[Get Started with Vertex Experiments and Vertex ML Metadata](get_started_vertex_experiments.ipynb)
[Get started with Vertex AI Training for R](community/ml_ops/stage2/get_started_vertex_training_r.ipynb)
In this tutorial, you learn how to use `Vertex AI Training` for training a R custom model.
The steps performed include:
- Locally train an R model in a notebook using %%R magic commands
- Create a deployment image with trained R model and serving functions.
- Test the deployment image locally.
- Create a `Vertex AI Model` resource for the deployment image with embedded R model.
- Deploy the deployment image with embedded R model to a `Vertex AI Endpoint` resource.
- Test the deployment image with embedded R model.
- Create a R-to-Python training package.
- Create a training image for training the model.
- Train a R model using `Vertex AI Trainingh` service with the R-to-Python training package.
[Get started with Logging](community/ml_ops/stage2/get_started_with_logging.ipynb)
In this tutorial, you learn how to use Python and Cloud logging awhen training with `Vertex AI`.
```
The steps performed include:
- Use Python logging to log training configuration/results locally.
- Use Google Cloud Logging to log training configuration/results in cloud storage.
- Create a Vertex AI `Experiment` resource.
- Instantiate an experiment run.
- Log parameters for the run.
- Log metrics for the run.
- Display the logged experiment run.
```
[Get Started with Vertex TensorBoard](get_started_vertex_tensorboard.ipynb)
[Get started with Vertex AI Hyperparameter Tuning for XGBoost] (community/ml_ops/stage2/get_started_vertex_hpt_xgboost.ipynb)
In this tutorial, you learn how to use `Vertex AI Hyperparameter Tuning` for training a XGBoost custom model.
The steps performed include:
- Training using a Python package.
- Report accuracy when hyperparameter tuning.
- Save the model artifacts to Cloud Storage using GCSFuse.
- Create a `Vertex AI Model` resource.
[Get started with Vertex AI Training for XGBoost](community/ml_ops/stage2/get_started_vertex_training_xgboost.ipynb)
In this tutorial, you learn how to use `Vertex AI Training` for training a XGBoost custom model.
The steps performed include:
- Training using a Python package.
- Report accuracy when hyperparameter tuning.
- Save the model artifacts to Cloud Storage using GCSFuse.
- Create a `Vertex AI Model` resource.
[Get started with TabNet builtin algorithm for training tabular models](community/ml_ops/stage2/get_started_with_tabnet.ipynb)
In this notebook, you learn how to run `Vertex AI TabNet` built algorithm for training custom tabular models.
The steps performed include:
- Get the training data.
- Configure training parameters for the `Vertex AI TabNet` container.
- Train the model using `Vertex AI Training` using CSV data.
- Upload the model as a `Vertex AI Model` resource.
- Deploy the `Vertex AI Model` resource to a `Vertex AI Endpoint` resource.
- Make a prediction with the deployed model.
- Hyperparameter tuning the `Vertex AI TabNet` model.
- Train the model using `Vertex AI Training` using BigQuery table.
[Get started with prebuilt TFHub models](community/ml_ops/stage2/get_started_with_tfhub_models.ipynb)
In this tutorial, you learn how to use `Vertex AI Training` with prebuilt models from TensorFlow Hub.
The steps performed include:
- Download a TensorFlow Hub prebuilt model.
- Add the task component as a classifier for the CIFAR-10 dataset.
- Fine tune locally the model with transfer learning training.
- Construct a custom training script:
- Get training data from TensorFlow Datasets
- Get model architecture from TensorFlow Hub
- Train then model
- Save model artifacts and upload as Vertex AI Model resource.
[Get started with BigQuery ML Training](community/ml_ops/stage2/get_started_bqml_training.ipynb)
In this tutorial, you learn how to use `BigQueryML` (BQML) for training with `Vertex AI`.
The steps performed include:
- Create a local BigQuery table in your project
- Train a BQML model
- Evaluate the BQML model
- Export the BQML model as a cloud model
- Upload the exported model as a `Vertex AI Model` resource
- Hyperparameter tune a BQML model with `Vertex AI Vizier`
- Automatically register a BQML model to `Vertex AI Model Registry`
[Get started with Vertex AI Vizier](community/ml_ops/stage2/get_started_vertex_vizier.ipynb)
In this tutorial, you learn how to use `Vertex AI Vizier` for when training with `Vertex AI`.
The steps performed include:
- Hyperparameter tuning with Random algorithm.
- Hyperparameter tuning with Vizier (Bayesian) algorithm.
- Suggesting trials and updating results for Vizier study
[Get started with distributed training using DASK](community/ml_ops/stage2/get_started_with_distributed_training_xgboost.ipynb)
In this tutorial, you learn how to use `Vertex AI Training` for distributed training of XGBoost model using the OSS package DASK. Additionally, you learn to construct and deploy a custom serving container using a Flask web server.
The steps performed include:
- Construct an XGBoost training script using DASK for distributed training.
- Construct a custom training container.
- Configure a distributed custom training job.
- Execute the custom training job.
- Construct a custom serving container using Flask.
- Upload the trained XGBoost model as a `Vertex AI Model` resource.
- Create a `Vertex AI Endpoint` resource.
- Deploy the `Vertex AI Model` resource to `Vertex AI Endpoint` resource.
- Make a prediction.
[Get started with Vertex AI TensorBoard](community/ml_ops/stage2/get_started_vertex_tensorboard.ipynb)
In this tutorial, you learn how to use `Vertex AI TensorBoard` when training with `Vertex AI`.
```
The steps performed include:
- Create a TensorBoard callback when training a model.
- Using Tensorboard with locally trained model.
- Using TensorBoard with locally trained model.
- Using Vertex AI TensorBoard with Vertex AI Training.
```
[Get Started with Custom Training Packages (Tensorflow)](get_started_vertex_training.ipynb)
[Get started with Vertex AI Training for R using R Kernel](community/ml_ops/stage2/get_started_vertex_training_r_using_r_kernel.ipynb)
In this tutorial, you learn how to use `Vertex AI`, using an R kernel, for training and deploying an R custom model.
The steps performed include:
- Create a custom R training script
- Create a custom R serving script
- Create a custom R deployment (serving) container.
- Train the model using `Vertex AI` custom training.
- Create an `Endpoint` resouce.
- Deploy the `Model` resource (trained R model) to the `Endpoint` resource.
- Make an online prediction.
[Get started Vision API test preprocessing and AutoML text model generation](community/ml_ops/stage2/get_started_with_visionapi_and_automl.ipynb)
In this tutorial, you create an `AutoML` text entity extraction model pre-existing extracted data by generating a custom import file. You deploy this mode for online prediction from a Python script using the `BigQuery`, `Vision AI`, Cloud Storage and `Vertex AI SDK` for Python.
The steps performed include:
- Preprocess training files using `Vision AI` APIs to extract the text from PDF files.
- Create a custom import file that includes annotation data based on the sample `BigQuery` dataset.
- Create a `Vertex AI Dataset` resource.
- Train the model.
- View the model evaluation.
- Deploy the `Vertex AI Model` resource to a serving `Endpoint` resource.
- Make a prediction.
- Undeploy the `Model`.
[Get started with Vertex AI Experiments](community/ml_ops/stage2/get_started_vertex_experiments.ipynb)
In this tutorial, you learn how to use `Vertex AI Experiments` when training with `Vertex AI`.
The steps performed include:
- Local (notebook) Training
- Create an experiment
- Create a first run in the experiment
- Log parameters and metrics
- Create artifact lineage
- Visualize the experiment results
- Execute a second run
- Compare the two runs in the experiment
- Cloud (`Vertex AI`) Training
- Within the training script:
- Create an experiment
- Log parameters and metrics
- Create artifact lineage
- Create a `Vertex AI Training` custom job
- Execute the custom job
- Visualize the experiment results
[AutoML Image Classfication Training with Customer Managed Encryption Keys (CMEK)](community/ml_ops/stage2/get_started_with_cmek_training.ipynb)
In this tutorial, you learn how to use a customer managed encryption key (CMEK) for `Vertex AI AutoML` training.
The steps performed include:
- Creating a customer managed encryption key.
- Creating an image dataset with CMEK encryption.
- Train an AutoML model with CMEK encryption.
[Get started with Vertex AI Feature Store](community/ml_ops/stage2/get_started_vertex_feature_store.ipynb)
In this tutorial, you learn how to use `Vertex AI Feature Store` when training and predicting with `Vertex AI`.
The steps performed include:
- Creating a Vertex AI `Featurestore` resource.
- Creating `EntityType` resources for the `Featurestore` resource.
- Creating `Feature` resources for each `EntityType` resource.
- Import feature values (entity data items) into `Featurestore` resource.
- From a Cloud Storage location.
- From a pandas DataFrame.
- Perform online serving from a `Featurestore` resource.
- Perform batch serving from a `Featurestore` resource.
[Get started with AutoML Training](community/ml_ops/stage2/get_started_automl_training.ipynb)
In this tutorial, you learn how to use `AutoML` for training with `Vertex AI`.
The steps performed include:
- Train an image model
- Export the image model as an edge model
- Train a tabular model
- Export the tabular model as a cloud model
- Train a text model
- Train a video model
[Get started with Vertex AI Training for LightGBM](community/ml_ops/stage2/get_started_vertex_training_lightgbm.ipynb)
In this tutorial, you learn how to use `Vertex AI Training` for training a LightGBM custom model.
The steps performed include:
- Training using a Python package.
- Save the model artifacts to Cloud Storage using GCSFuse.
- Construct a FastAPI prediction server.
- Construct a Dockerfile deployment image.
- Test the deployment image locally.
- Create a `Vertex AI Model` resource.
[Get started with Vertex AI Training for Scikit-Learn](community/ml_ops/stage2/get_started_vertex_training_sklearn.ipynb)
In this tutorial, you learn how to use `Vertex AI Training` for training a Scikit-Learn custom model.
The steps performed include:
- Training using a Python package.
- Report accuracy when hyperparameter tuning.
- Save the model artifacts to Cloud Storage using GCSFuse.
- Create a `Vertex AI Model` resource.
[Get started with Vertex AI Training](community/ml_ops/stage2/get_started_vertex_training.ipynb)
In this tutorial, you learn how to use `Vertex AI Training` for custom models when training with `Vertex AI`.
```
The steps performed include:
- Training using a single Python script.
- Training using a Python package.
- Training using a custom training image.
- Laying out a training package.
```
[Get Started with Custom Training Packages (Scikit-Learn)](get_started_vertex_training_sklearn.ipynb)
[Get Started with Custom Training Packages (XGBoost)](get_started_vertex_training_xgboost.ipynb)
[Get started with Vertex AI Training for Pytorch](community/ml_ops/stage2/get_started_vertex_training_pytorch.ipynb)
[Get Started with Custom Training Packages (Pytorch)](get_started_vertex_training_pytorch.ipynb)
In this tutorial, you learn how to use `Vertex AI Training` for training a Pytorch custom model.
[Get Started with Custom Training Packages (R)](get_started_vertex_training_r.ipynb)
The steps performed include:
[Get Started with Distributed Training](get_started_vertex_distributed_training.ipynb)
- Single node training using a Python package.
- Report accuracy when hyperparameter tuning.
- Save the model artifacts to Cloud Storage using GCSFuse.
- Create a `Vertex AI Model` resource.
[Get Started with Vizier Hyperparameter Tuning](get_started_vertex_vizier.ipynb)
[Get started with Vertex AI Distributed Training](community/ml_ops/stage2/get_started_vertex_distributed_training.ipynb)
[Get Started with AutoML Training](get_started_automl_training.ipynb)
In this tutorial, you learn how to use `Vertex AI Distributed Training` for when training with `Vertex AI`.
[Get Started with BQML Training](get_started_bqml_training.ipynb)
The steps performed include:
[Get Started with Vertex Feature Store](get_started_vertex_feature_store.ipynb)
-`MirroredStrategy`: Train on a single VM with multiple GPUs.
-`MultiWorkerMirroredStrategy`: Train on multiple VMs with automatic setup of replicas.
-`MultiWorkerMirroredStrategy`: Train on multiple VMs with fine grain control of replicas.
-`ReductionServer`: Train on multiple VMS and sync updates across VMS with `Vertex AI Reduction Server`.
" <img src=\"https://lh3.googleusercontent.com/UiNooY4LUgW_oTvpsNhPpQzsstV5W8F7rYgxgGBD85cWJoLmrOzhVs_ksK_vgx40SHs7jCqkTkCk=e14-rj-sc0xffffff-h130-w32\" alt=\"Vertex AI logo\">\n",
" Open in Vertex AI Workbench\n",
" </a>\n",
" </td>\n",
"</table>\n",
@@ -59,17 +65,6 @@
"This tutorial demonstrates how to use Vertex AI for E2E MLOps on Google Cloud in production. This tutorial covers stage 2 : experimentation: get started with BigQuery ML Training."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "dataset:penguins,lcn,bq"
},
"source": [
"### Dataset\n",
"\n",
"The dataset used for this tutorial is the Penguins dataset from [BigQuery public datasets](https://cloud.google.com/bigquery/public-data). The version of the dataset predicts the species."
]
},
{
"cell_type": "markdown",
"metadata": {
@@ -78,22 +73,50 @@
"source": [
"### Objective\n",
"\n",
"In this tutorial, you learn how to use `BigQueryML` (BQML) for training with `Vertex AI`.\n",
"In this tutorial, you learn how to use `BigQueryML` for training with `Vertex AI`.\n",
"\n",
"This tutorial uses the following Google Cloud ML services:\n",
"\n",
"- `BigQueryML Training`\n",
"- `Vertex AI Model resource`\n",
"- `Vertex AI Vizier.\n",
"- `Vertex AI Vizier`\n",
"\n",
"The steps performed include:\n",
"\n",
"- Create a local BQ table in your project.\n",
"- Train a BQML model.\n",
"- Evaluate the BQML model.\n",
"- Export the BQML model as a cloud model.\n",
"- Upload the exported model as a Vertex AI Model resource.\n",
"- Hyperparameter tune a BQML model with Vertex AI Vizier."
"- Create a local BigQuery table in your project\n",
"- Train a BigQuery ML model\n",
"- Evaluate the BigQuery ML model\n",
"- Export the BigQuery ML model as a cloud model\n",
"- Upload the exported model as a `Vertex AI Model` resource\n",
"- Hyperparameter tune a BigQuery ML model with `Vertex AI Vizier`\n",
"- Automatically register a BigQuery ML model to `Vertex AI Model Registry`"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "dataset:penguins,lcn,bq"
},
"source": [
"### Dataset\n",
"\n",
"The dataset used for this tutorial is the Penguins dataset from [BigQuery public datasets](https://cloud.google.com/bigquery/public-data). This version of the dataset is used to predict the species of penguins from the available features like culmen-length, flipper-depth etc."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "81c777b8ad32"
},
"source": [
"### Costs\n",
"This tutorial uses billable components of Google Cloud:\n",
"\n",
"- Vertex AI\n",
"- Cloud Storage\n",
"- BigQuery\n",
"\n",
"Learn about [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing), [Cloud Storage pricing](https://cloud.google.com/storage/pricing) and [BigQuery pricing](https://cloud.google.com/bigquery/pricing) and use the [Pricing Calculator](https://cloud.google.com/products/calculator/) to generate a cost estimate based on your projected usage."
]
},
{
@@ -104,7 +127,7 @@
"source": [
"## Installations\n",
"\n",
"Install *one time* the packages for executing the MLOps notebooks."
"Install the following packages for executing this notebook."
"# Vertex AI Notebook requires dependencies to be installed with '--user'\n",
"USER_FLAG = \"\"\n",
"if IS_WORKBENCH_NOTEBOOK:\n",
" USER_FLAG = \"--user\"\n",
"\n",
"# Install the packages\n",
"! pip3 install --upgrade pyarrow \\\n",
" google-cloud-aiplatform \\\n",
" google-cloud-bigquery \\\n",
" google-cloud-bigquery-storage $USER_FLAG -q"
]
},
{
@@ -165,6 +190,32 @@
"metadata": {
"id": "project_id"
},
"source": [
"## Before you begin\n",
"\n",
"### Set up your Google Cloud project\n",
"\n",
"**The following steps are required, regardless of your notebook environment.**\n",
"\n",
"1. [Select or create a Google Cloud project](https://console.cloud.google.com/cloud-resource-manager). When you first create an account, you get a $300 free credit towards your compute/storage costs.\n",
"\n",
"1. [Make sure that billing is enabled for your project](https://cloud.google.com/billing/docs/how-to/modify-project).\n",
"\n",
"1. [Enable the Vertex AI, BigQuery, Compute Engine and Cloud Storage APIs](https://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com,bigquery,compute_component,storage_component).\n",
"\n",
"1. If you are running this notebook locally, you need to install the [Cloud SDK](https://cloud.google.com/sdk).\n",
"\n",
"1. Enter your project ID in the cell below. Then run the cell to make sure the\n",
"Cloud SDK uses the right project for all the commands in this notebook.\n",
"\n",
"**Note**: Jupyter runs lines prefixed with `!` as shell commands, and it interpolates Python variables prefixed with `$` into these commands."
"**If you are using Vertex AI Workbench Notebooks**, your environment is already authenticated. Skip this step.\n",
"\n",
"**If you are using Colab**, run the cell below and follow the instructions when prompted to authenticate your account via oAuth.\n",
"\n",
"**Otherwise**, follow these steps:\n",
"\n",
"In the Cloud Console, go to the [Create service account key](https://console.cloud.google.com/apis/credentials/serviceaccountkey) page.\n",
"\n",
"1. **Click Create service account**.\n",
"\n",
"2. In the **Service account name** field, enter a name, and click **Create**.\n",
"\n",
"3. In the **Grant this service account access to project** section, click the Role drop-down list. Type \"Vertex AI\" into the filter box, and select **Vertex AI Administrator**. Type \"Storage Object Admin\" into the filter box, and select **Storage Object Admin**.\n",
"\n",
"4. Click Create. A JSON file that contains your key downloads to your local environment.\n",
"\n",
"5. Enter the path to your service account key as the GOOGLE_APPLICATION_CREDENTIALS variable in the cell below and run the cell."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "e0953a00668e"
},
"outputs": [],
"source": [
"# If you are running this notebook in Colab, run this cell and follow the\n",
"# instructions to authenticate your GCP account. This provides access to your\n",
"# Cloud Storage bucket and lets you submit training jobs and prediction\n",
"# requests.\n",
"\n",
"import os\n",
"import sys\n",
"\n",
"# If on Vertex AI Workbench, then don't execute this code\n",
"IS_COLAB = False\n",
"if not os.path.exists(\"/opt/deeplearning/metadata/env_version\") and not os.getenv(\n",
" \"DL_ANACONDA_HOME\"\n",
"):\n",
" if \"google.colab\" in sys.modules:\n",
" IS_COLAB = True\n",
" from google.colab import auth as google_auth\n",
"\n",
" google_auth.authenticate_user()\n",
"\n",
" # If you are running this notebook locally, replace the string below with the\n",
" # path to your service account key and run this cell to authenticate your GCP\n",
"You use a service account to create Vertex AI Pipeline jobs. If you do not want to use your project's Compute Engine service account, set `SERVICE_ACCOUNT` to another service account ID."
"You can set hardware accelerators for prediction.\n",
"\n",
"Set the variable `DEPLOY_GPU/DEPLOY_NGPU` to use a container image supporting a GPU and the number of GPUs allocated to the virtual machine (VM) instance. For example, to use a GPU container image with 4 Nvidia Telsa K80 GPUs allocated to each VM, you would specify:\n",
"Set the pre-built Docker container image for prediction.\n",
"\n",
@@ -519,11 +660,11 @@
"id": "machine:prediction"
},
"source": [
"#### Set machine type\n",
"### Set machine type\n",
"\n",
"Next, set the machine type to use for prediction.\n",
"\n",
"- Set the variable `DEPLOY_COMPUTE` to configure the compute resources for the VM you will use for prediction.\n",
"- Set the variable `DEPLOY_COMPUTE` to configure the compute resources for the VM which is used for prediction.\n",
" - `machine type`\n",
" - `n1-standard`: 3.75GB of memory per vCPU.\n",
" - `n1-highmem`: 6.5GB of memory per vCPU\n",
@@ -557,7 +698,7 @@
"id": "bqml_intro"
},
"source": [
"## Bigquery ML introduction\n",
"## BigQuery ML introduction\n",
"\n",
"BigQuery ML (BQML) provides the capability to train ML tabular models, such as classification and regression, in BigQuery using SQL syntax.\n",
"\n",
@@ -582,9 +723,9 @@
"id": "bqml_create_dataset"
},
"source": [
"### Create BQ dataset/model resource\n",
"### Create BQ dataset resource\n",
"\n",
"First, you create a empty dataset/model resource in your project."
"First, you create an empty dataset resource in your project."
]
},
{
@@ -608,9 +749,9 @@
"id": "bqml_create_model"
},
"source": [
"### Train BQML model\n",
"### Train BigQuery ML model\n",
"\n",
"Next, you create and train a BQML tabular classification model from the public dataset penguins and store the model in your project using the `CREATE MODEL` statement. The model configuration is specified in the `OPTIONS` statement as follows:\n",
"Next, you create and train a BigQuery ML tabular classification model from the public dataset penguins and store the model in your project using the `CREATE MODEL` statement. The model configuration is specified in the `OPTIONS` statement as follows:\n",
"\n",
"- `model_type`: The type and archictecture of tabular model to train, e.g., DNN classification.\n",
"- `labels`: The column which are the labels.\n",
@@ -659,9 +800,9 @@
"id": "bqml_eval_model"
},
"source": [
"### Evaluate the BQML trained model\n",
"### Evaluate the trained BigQuery ML model\n",
"\n",
"Next, retrieve the model evaluation for the trained BQML model.\n",
"Next, retrieve the model evaluation for the trained BigQuery ML model.\n",
"\n",
"Learn more about [The ML.EVALUATE function](https://cloud.google.com/bigquery-ml/docs/reference/standard-sql/bigqueryml-syntax-evaluate)."
]
@@ -692,9 +833,9 @@
"id": "bqml_export_model"
},
"source": [
"### Export the model from BQML\n",
"### Export the model from BigQuery ML\n",
"\n",
"The model you trained in BQML is a TensorFlow model. Next, you will export the TensorFlow model artifacts in TF.SavedModel format."
"The model you trained in BigQuery ML is a TensorFlow model. Next, you export the TensorFlow model artifacts in TF.SavedModel format."
"## Upload the BigQuery ML model to a Vertex AI Model resource\n",
"\n",
"Finally, now that you have the BQML model exported as a TF.SavedModel format, you upload the model artifacts to Vertex AI Model resource, in the same way as if you were uploading a custom trained model."
"Finally, now that you have the BigQuery ML model exported, you upload the model artifacts to Vertex AI Model resource, in the same way as if you were uploading a custom trained model.\n",
"\n",
"Below is a partial list of mapping BigQuery ML model types to their corresponding exported model format:\n",
"- `deployed_model_display_name`: A human readable name for the deployed model.\n",
"- `traffic_split`: Percent of traffic at the endpoint that goes to this model, which is specified as a dictionary of one or more key/value pairs.\n",
"If only one model, then specify as { \"0\": 100 }, where \"0\" refers to this model being uploaded and 100 means 100% of the traffic.\n",
"If there are existing models on the endpoint, for which the traffic will be split, then use model_id to specify as { \"0\": percent, model_id: percent, ... }, where model_id is the model id of an existing model to the deployed endpoint. The percents must add up to 100.\n",
"If there are existing models on the endpoint, for which the traffic needs to be split, then use model_id to specify as { \"0\": percent, model_id: percent, ... }, where model_id is the model id of an existing model to the deployed endpoint. The percents must add up to 100.\n",
"- `machine_type`: The type of machine to use for training.\n",
"- `accelerator_type`: The hardware accelerator type.\n",
"- `accelerator_count`: The number of accelerators to attach to a worker replica.\n",
@@ -801,7 +958,7 @@
"id": "undeploy_model:mbsdk"
},
"source": [
"## Undeploy the model\n",
"#### Undeploy the model\n",
"\n",
"When you are done doing predictions, you undeploy the model from the `Endpoint` resouce. This deprovisions all compute resources and ends billing for the deployed model."
]
@@ -823,9 +980,9 @@
"id": "model_delete:mbsdk"
},
"source": [
"#### Delete the model\n",
"#### Delete the `Vertex AI Model` resource\n",
"\n",
"The method 'delete()' will delete the model."
"The method 'delete()' deletes the model."
]
},
{
@@ -839,15 +996,41 @@
"model.delete()"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "7890ae6f6410"
},
"source": [
"### Delete the `BigQuery ML` model\n",
"\n",
"Next, delete the `BigQuery ML` instance of the model."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "f0b6163e70c0"
},
"outputs": [],
"source": [
"MODEL_QUERY = f\"\"\"\n",
"DROP MODEL `{BQ_DATASET_NAME}.{MODEL_NAME}`\n",
"\"\"\"\n",
"\n",
"job = bqclient.query(MODEL_QUERY)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "bqml_create_model:vizier"
},
"source": [
"### Hyperparameter Tune and train a BQML model\n",
"### Hyperparameter Tune and train a BigQuery ML model\n",
"\n",
"Next, you train a BQML tabular classification model with hyperparameter tuning using the Vertex AI Vizier service. The hyperparameter settings are specified in the `OPTIONS` statement as follows:\n",
"Next, you train a BigQuery ML tabular classification model with hyperparameter tuning using the `Vertex AI Vizier` service. The hyperparameter settings are specified in the `OPTIONS` statement as follows:\n",
"\n",
"- `HPARAM_TUNING_ALGORITHM`: The algorithm for selecting the next trial parameters.\n",
"- `num_trials`: The number of trials.\n",
@@ -900,9 +1083,9 @@
"id": "bqml_eval_model"
},
"source": [
"### Evaluate the BQML trained model\n",
"### Evaluate the BigQuery ML trained model\n",
"\n",
"Next, retrieve the model evaluation for the trained BQML model.\n",
"Next, retrieve the model evaluation results for the trained BigQuery ML model.\n",
"\n",
"Learn more about [The ML.EVALUATE function](https://cloud.google.com/bigquery-ml/docs/reference/standard-sql/bigqueryml-syntax-evaluate)."
]
@@ -927,15 +1110,41 @@
"print(results)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "3f3cee1236b1"
},
"source": [
"### Delete the `BigQuery ML` model\n",
"\n",
"Next, delete the `BigQuery ML` instance of the model."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "957b7d841502"
},
"outputs": [],
"source": [
"MODEL_QUERY = f\"\"\"\n",
"DROP MODEL `{BQ_DATASET_NAME}.{MODEL_NAME}`\n",
"\"\"\"\n",
"\n",
"job = bqclient.query(MODEL_QUERY)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "bqml_create_model:xai"
},
"source": [
"### Train a BQML model with Explainability\n",
"### Train a BigQuery ML model with Explainability\n",
"\n",
"Next, you train the same BQML model, but this time you enable Vertex AI Explainability on the model predictions by adding the option:\n",
"Next, you train the same BigQuery ML model, but this time you enable Vertex AI Explainability on the model predictions by adding the option:\n",
"\n",
"- `ENABLE_GLOBAL_EXPLAIN`"
]
@@ -976,6 +1185,179 @@
"print(\"{} created in {}\".format(tblname, job.ended - job.started))"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "4def8aaf3398"
},
"source": [
"### Delete the `BigQuery ML` model\n",
"\n",
"Next, delete the `BigQuery ML` instance of the model."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "ff5b32618018"
},
"outputs": [],
"source": [
"MODEL_QUERY = f\"\"\"\n",
"DROP MODEL `{BQ_DATASET_NAME}.{MODEL_NAME}`\n",
"\"\"\"\n",
"\n",
"job = bqclient.query(MODEL_QUERY)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "2b4498ca6fea"
},
"source": [
"## Model Registry\n",
"\n",
"Alternatively, you can implicitly upload your BigQuery ML model as a `Vertex AI Model` resource with exporting and importing the model artifacts. In this method, you add additional options when training the model that tells BigQuery ML to automatically upload and register the trained model as a `Model` resource.\n",
"\n",
"### Setting permissions to automatically register the model\n",
"\n",
"You need to set some additional IAM permissions for BigQuery ML to automatically upload and register the model after training. Depending on your service account, the setting of the permissions below may fail. In this case, we recommend executing the permissions in a Cloud Shell.\n",
"\n",
"Learn more about [Setting permissions for Model Registry](https://cloud.google.com/bigquery-ml/docs/managing-models-vertex)\n"
" <img src=\"https://lh3.googleusercontent.com/UiNooY4LUgW_oTvpsNhPpQzsstV5W8F7rYgxgGBD85cWJoLmrOzhVs_ksK_vgx40SHs7jCqkTkCk=e14-rj-sc0xffffff-h130-w32\" alt=\"Vertex AI logo\">\n",
"Open in Vertex AI Workbench\n",
" </a>\n",
" </td>\n",
"</table>\n",
@@ -56,18 +62,7 @@
"## Overview\n",
"\n",
"\n",
"This tutorial demonstrates how to use Vertex AI for E2E MLOps on Google Cloud in production. This tutorial covers stage 2 : experimentation: get started with Vertex Distributed Training."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "dataset:custom,boston,lrg"
},
"source": [
"### Dataset\n",
"\n",
"The dataset used for this tutorial is the [Boston Housing Prices dataset](https://www.cs.toronto.edu/~delve/data/boston/bostonDetail.html). The version of the dataset you will use in this tutorial is built into TensorFlow. The trained model predicts the median price of a house in units of 1K USD."
"This tutorial demonstrates how to use Vertex AI for E2E MLOps on Google Cloud in production. This tutorial covers stage 2 : experimentation: get started with Vertex Distributed Training. Please note: There are incompatibilities between Colab and Docker and the Docker section may not work until resolved by the platform."
]
},
{
@@ -126,59 +121,86 @@
{
"cell_type": "markdown",
"metadata": {
"id": "install_mlops"
"id": "dataset:custom,boston,lrg"
},
"source": [
"## Installations\n",
"### Dataset\n",
"\n",
"Install *one time* the packages for executing the MLOps notebooks."
"The dataset used for this tutorial is the [Boston Housing Prices dataset](https://www.cs.toronto.edu/~delve/data/boston/bostonDetail.html). The version of the dataset you use in this tutorial is built into TensorFlow. The trained model predicts the median price of a house in units of 1K USD."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "restart"
"id": "d10166df7141"
},
"source": [
"### Restart the kernel\n",
"### Costs\n",
" \n",
"This tutorial uses billable components of Google Cloud:\n",
"\n",
"Once you've installed the additional packages, you need to restart the notebook kernel so it can find the packages."
"Vertex AI\n",
"Cloud Storage\n",
"\n",
"Learn about [Vertex AI pricing](https://cloud.google.com/vertex-ai/pricing) and [Cloud Storage pricing](https://cloud.google.com/storage/pricing), and use the [Pricing Calculator](https://cloud.google.com/products/calculator/),\n",
" to generate a cost estimate based on your projected usage."
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "XkYpRvOQyVYb"
},
"source": [
"## Installation\n",
"\n",
"Install the packages required for executing this notebook."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "restart"
"id": "xs_Kt8RcyXTC"
},
"outputs": [],
"source": [
"import os\n",
"\n",
"# The Vertex AI Workbench Notebook product has specific requirements\n",
"IS_WORKBENCH_NOTEBOOK = os.getenv(\"DL_ANACONDA_HOME\") and not os.getenv(\"VIRTUAL_ENV\")\n",
"After you install the additional packages, you need to restart the notebook kernel so it can find the packages."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"id": "zo3YFZXLzCRJ"
},
"outputs": [],
"source": [
"# Automatically restart kernel after installs\n",
"import os\n",
"\n",
"if not os.getenv(\"IS_TESTING\"):\n",
@@ -189,6 +211,32 @@
" app.kernel.do_shutdown(True)"
]
},
{
"cell_type": "markdown",
"metadata": {
"id": "84cd83853240"
},
"source": [
"## Before you begin\n",
"\n",
"### Set up your Google Cloud project\n",
"\n",
"**The following steps are required, regardless of your notebook environment.**\n",
"\n",
"1. [Select or create a Google Cloud project](https://console.cloud.google.com/cloud-resource-manager). When you first create an account, you get a $300 free credit towards your compute/storage costs.\n",
"\n",
"1. [Make sure that billing is enabled for your project](https://cloud.google.com/billing/docs/how-to/modify-project).\n",
"\n",
"1. [Enable the Vertex AI, BigQuery, Compute Engine and Cloud Storage APIs](https://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com,bigquery,compute_component,storage_component).\n",
"\n",
"1. If you are running this notebook locally, you need to install the [Cloud SDK](https://cloud.google.com/sdk).\n",
"\n",
"1. Enter your project ID in the cell below. Then run the cell to make sure the\n",
"Cloud SDK uses the right project for all the commands in this notebook.\n",
"\n",
"**Note**: Jupyter runs lines prefixed with `!` as shell commands, and it interpolates Python variables prefixed with `$` into these commands."
File diff suppressed because it is too large
Load Diff
Some files were not shown because too many files have changed in this diff
Show More
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.