Compare commits

...
Author SHA1 Message Date
Vertex MG TeamandCopybara-Service c1b292bf73 feat: Update the model id to a default value, this is not used when the deployed endpoint is passed in
PiperOrigin-RevId: 725755135
2025-02-11 13:41:52 -08:00
Vertex MG TeamandCopybara-Service ae44e54d57 feat: Add Reasoning Engine integration notebook
PiperOrigin-RevId: 725511227
2025-02-11 00:24:28 -08:00
Mend RenovateandGitHub 7835b84082 chore(deps): update dependency isort to v6 (#3810) 2025-02-10 20:39:00 +00:00
ethan-gordonandGitHub 30d6112517 set FeatureView IAM and Service Accounts as version v1 and remove preview note. (#3795) 2025-02-10 20:37:36 +00:00
Vertex MG TeamandCopybara-Service 32737fda1f Sync mistral and mixtral ft notebooks
PiperOrigin-RevId: 725241490
2025-02-10 09:20:36 -08:00
Vertex MG TeamandCopybara-Service e02adc516c feat: Update Reasoning Engine with Llama 3.1 models notebook.
PiperOrigin-RevId: 725038312
2025-02-09 19:53:21 -08:00
Vertex MG TeamandCopybara-Service 80320a9a1b Refactor the notebook
PiperOrigin-RevId: 724288214
2025-02-07 03:55:12 -08:00
Minwoo ParkandCopybara-Service 8c527a89ed Add Model Garden finetuning tutorial notebook.
PiperOrigin-RevId: 724140766
2025-02-06 17:56:55 -08:00
Vertex MG TeamandCopybara-Service 33a2b7eaf9 feat: Add Reasoning Engine with Llama 3.1 models notebook.
PiperOrigin-RevId: 724139178
2025-02-06 17:51:09 -08:00
Minwoo ParkandCopybara-Service dfa461dabe Add Model Garden finetuning tutorial notebook.
PiperOrigin-RevId: 724065750
2025-02-06 14:02:30 -08:00
Vertex MG TeamandCopybara-Service c80de903af Update PaliGemma 2 notebook and handler to use weights from GCS.
PiperOrigin-RevId: 724065613
2025-02-06 14:01:01 -08:00
Vertex MG TeamandCopybara-Service dd333b8fdd feat: Add Reasoning Engine with Llama 3.1 models notebook.
PiperOrigin-RevId: 723785091
2025-02-05 22:12:52 -08:00
Vertex MG TeamandCopybara-Service 911ef0cb69 A minor tweaking in the helper functions and a Bug fix in the Predict Section.
PiperOrigin-RevId: 723779782
2025-02-05 21:46:35 -08:00
Vertex MG TeamandCopybara-Service ac6f4d669d Remove Llama-Guard (llama3.1) models from notebook model list
PiperOrigin-RevId: 723282525
2025-02-04 17:21:55 -08:00
Dustin LuongandCopybara-Service abdd887fa1 Set model_garden_source_model_name for model_garden_phi4_deployment notebook.
PiperOrigin-RevId: 722840147
2025-02-03 16:25:21 -08:00
Dustin LuongandCopybara-Service 711a4f0b8c Set model_garden_source_model_name for deployment notebooks.
PiperOrigin-RevId: 722701496
2025-02-03 10:06:48 -08:00
Aiden010200andGitHub 07f30dfde5 Upload a SGD classifier predictor example (#3783)
This example uses aiplatform and scikit-learn library to provide a SGD classifier.
2025-02-03 16:09:56 +00:00
Mend RenovateandGitHub b33e3287b7 chore(deps): update dependency black to v25 (#3815) 2025-02-03 16:09:21 +00:00
Dustin LuongandCopybara-Service 3c06e4797a Set model_garden_source_model_name for some model garden deployment notebooks.
PiperOrigin-RevId: 721846318
2025-01-31 11:39:27 -08:00
Vertex MG TeamandCopybara-Service a5944510d6 Add a notebook about Model Garden advanced features, including prefix caching and speculative decoding.
PiperOrigin-RevId: 721844683
2025-01-31 11:35:01 -08:00
ethan-gordonandGitHub 6d74f87e22 Add notebook vertex_ai_feature_store_update_feature_monitor_feature_group_iam_and_service_agent.ipynb. (#3814)
This change also inserts a corresponding entry to CODEOWNERS.
2025-01-30 22:04:59 +00:00
Vertex MG TeamandCopybara-Service 14031b238c BiomedCLIP Deployment on Vertex Notebook
PiperOrigin-RevId: 720990850
2025-01-29 08:40:29 -08:00
Vertex MG TeamandCopybara-Service 99e578ca2e Add usage tracking labels for finetuning notebooks
PiperOrigin-RevId: 720826465
2025-01-28 22:03:39 -08:00
Vertex MG TeamandCopybara-Service 539683f573 Adding Phi-4 Colab deployment notebook
PiperOrigin-RevId: 720596224
2025-01-28 09:04:43 -08:00
Dustin LuongandCopybara-Service 6fabc23db6 Set model_garden_source_model_name for vllm deployments.
PiperOrigin-RevId: 720345103
2025-01-27 16:20:59 -08:00
Dustin LuongandCopybara-Service 18305e35ea Set model_garden_source_model_name for hexllm deployments.
PiperOrigin-RevId: 720341876
2025-01-27 16:09:59 -08:00
Dustin LuongandCopybara-Service b71f0ec2dc Set model_garden_source_model_name for optimized vllm deployments.
PiperOrigin-RevId: 720317678
2025-01-27 14:54:29 -08:00
Dustin LuongandCopybara-Service d7fe713f7f Set model_garden_source_model_name for TGI deployments.
PiperOrigin-RevId: 720312091
2025-01-27 14:40:10 -08:00
Dustin LuongandCopybara-Service 52925286ec Set model_garden_source_model_name for pytorch inference deployments.
PiperOrigin-RevId: 720311822
2025-01-27 14:38:52 -08:00
Dustin LuongandCopybara-Service b7b14ba7d8 Set model_garden_source_model_name for llama3
reference implementation deployment.

PiperOrigin-RevId: 720311748
2025-01-27 14:38:41 -08:00
Dustin LuongandCopybara-Service 5e3c07f72e Set model_garden_source_model_name for TEI deployments.
PiperOrigin-RevId: 720311569
2025-01-27 14:37:21 -08:00
Vertex MG TeamandCopybara-Service da630753ef Add Mistral and Llama3.1 8B serving notebooks
PiperOrigin-RevId: 719348288
2025-01-24 10:15:58 -08:00
Vertex MG TeamandCopybara-Service c27751e3aa Distinguish between the train and deploy machine specs for Gemma PEFT Finetuning on HF Notebook
PiperOrigin-RevId: 719309601
2025-01-24 08:16:52 -08:00
mumletandGitHub e9428949f1 fix: Update the model file path (#3778)
* fix: Update model path of model_monitoring_for_custom_model_online_prediction.ipynb

Update the unavailable model path for model_monitoring_for_custom_model_online_prediction.ipynb

* Update model_monitoring_for_custom_model_online_prediction.ipynb

Update the bq dataset uri
2025-01-23 23:45:17 +00:00
Vertex MG TeamandCopybara-Service 180c2834fc Fix minor lint issues
PiperOrigin-RevId: 718941819
2025-01-23 11:15:04 -08:00
Vertex MG TeamandCopybara-Service 77fd06d4c0 Fix MaaS requests in Llama Guard notebook.
PiperOrigin-RevId: 718856137
2025-01-23 07:23:05 -08:00
0943520f8d Update template location in dataset validation (#3799)
Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2025-01-23 13:26:44 +00:00
Vertex MG TeamandCopybara-Service c14ed381bf Stable Diffusion XL Finetuning Dreambooth Lora
PiperOrigin-RevId: 718724713
2025-01-22 23:26:09 -08:00
1f15188b6e Update common util and dataset validation util (#3796)
Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2025-01-23 02:27:23 +00:00
Bhaskar GoyalandGitHub 9ef41d55ba feat: Deprecate Claude 3 Sonnet (#3790) 2025-01-21 21:36:16 +00:00
Vertex MG TeamandCopybara-Service d82b5a4066 Enable H100 80GB DWS for 11B and 90B eval.
PiperOrigin-RevId: 717949505
2025-01-21 09:24:17 -08:00
Vertex MG TeamandCopybara-Service fdd4d37275 Paligemma 2 Deployment notebook
PiperOrigin-RevId: 717927000
2025-01-21 08:21:47 -08:00
Vertex MG TeamandCopybara-Service 81f700d508 Rename accelerator variables in the notebook
PiperOrigin-RevId: 717750767
2025-01-20 22:54:59 -08:00
Vertex MG TeamandCopybara-Service 32136a1894 BioGPT serving Notebook
PiperOrigin-RevId: 717710469
2025-01-20 20:13:51 -08:00
Vertex MG TeamandCopybara-Service c93a9c2099 Create Cloud translation and evaluation demo notebook
PiperOrigin-RevId: 715907272
2025-01-15 12:45:59 -08:00
Vertex MG TeamandCopybara-Service 8ece5ef3eb Segment Anything Model (SAM) Serving on Vertex AI Notebook
PiperOrigin-RevId: 715723816
2025-01-15 03:08:36 -08:00
Vertex MG TeamandCopybara-Service dee72afc70 Enable dedicate endpoint for Prompt Guard deployment
PiperOrigin-RevId: 715395374
2025-01-14 08:38:18 -08:00
Vertex MG TeamandCopybara-Service ea45e3dd5c Add usage tracking labels to all the finetuning notebook
PiperOrigin-RevId: 715229459
2025-01-13 21:41:30 -08:00
Vertex MG TeamandCopybara-Service 9ed8896350 vLLM supports GPU HBM + host memory prefix kv caching
PiperOrigin-RevId: 715213919
2025-01-13 20:30:33 -08:00
Vertex MG TeamandCopybara-Service 2709fcd8ed Fix the error when result is list, make it works for both dictionary and list
PiperOrigin-RevId: 715156151
2025-01-13 16:54:47 -08:00
Vertex MG TeamandCopybara-Service a7c1b9af3a Add publisherdb api call response check in case call fails
PiperOrigin-RevId: 715099165
2025-01-13 14:01:37 -08:00
Vertex MG TeamandCopybara-Service 952af223da Update license of the notebooks to 2025
PiperOrigin-RevId: 715041456
2025-01-13 11:17:00 -08:00
Bhaskar GoyalandGitHub d87b6b7463 feat: Add Codestral (25.01) model to mistral docs. (#3779) 2025-01-13 17:05:46 +00:00
Dustin LuongandCopybara-Service 25b3364236 No public description
PiperOrigin-RevId: 713694601
2025-01-09 09:13:27 -08:00
Vertex MG TeamandCopybara-Service 9f80c3cd9b Hex-LLM supports prefix caching as a GA feature
PiperOrigin-RevId: 713503181
2025-01-08 19:47:24 -08:00
Mend RenovateandGitHub 9f6ad8439c Update dependency pyupgrade to v3.19.1 (#3757) 2025-01-08 20:31:26 +00:00
Aiden010200andGitHub 64e9a4ae06 Upload a ResNet predictor example (#3765)
This example uses aiplatform and torch library to provide a ResNet predictor.
2025-01-08 20:30:52 +00:00
Dustin LuongandCopybara-Service 883e1e5fc1 Set system_labels in notebooks
PiperOrigin-RevId: 713040032
2025-01-07 14:20:25 -08:00
Vertex MG TeamandCopybara-Service 847a49f6d0 Update docker images to avoid 'tags' KeyError while loading HF dataset
PiperOrigin-RevId: 711729872
2025-01-03 06:09:28 -08:00
Vertex MG TeamandCopybara-Service 0b9a450b44 MediaPipe Text Classification notebook
PiperOrigin-RevId: 711656574
2025-01-03 00:25:31 -08:00
Vertex MG TeamandCopybara-Service ed72474525 pyTorch IMage Model notebook
PiperOrigin-RevId: 711343028
2025-01-02 00:53:05 -08:00
Vertex MG TeamandCopybara-Service 6b20a30669 TFVision Image segmentation notebook
PiperOrigin-RevId: 710942626
2024-12-31 04:20:47 -08:00
Vertex MG TeamandCopybara-Service f50f5f749d mediapipe face stylizer notebook
PiperOrigin-RevId: 710861510
2024-12-30 20:30:38 -08:00
alicechang0909andGitHub 9621877583 Created using Colab (#3763)
* Created using Colab

* Created using Colab

* Add Featurestore Monitoring functionalities - Fix lint error in import

* Fix lint with commands

* Address comments

* Update restart section to fix lint error.

* fix: fix lint errors

* Fix: Try to fix lint errors

* Fix: try fix lint with python commands

* fix: remove self link

* fix: try submit from workbench
2024-12-30 18:59:13 +00:00
Vertex MG TeamandCopybara-Service 425a5c94c2 Fix PIL issue FreeTypeFont object has no attribute getsize
PiperOrigin-RevId: 710698919
2024-12-30 06:09:47 -08:00
Vertex MG TeamandCopybara-Service 0b8b4b8ad6 movinet action recognition notebook
PiperOrigin-RevId: 709969002
2024-12-26 22:46:35 -08:00
Vertex MG TeamandCopybara-Service f641e5d221 Enable dedicate endpoint for pytorch llama3 deployment
PiperOrigin-RevId: 707697043
2024-12-18 16:07:12 -08:00
Vertex MG TeamandCopybara-Service d89b29f759 Add usage labels to finetuning notebook
PiperOrigin-RevId: 707402297
2024-12-17 22:37:52 -08:00
Vertex MG TeamandCopybara-Service df49e18ce9 Hex-LLM supports disaggregated serving as an experimental feature
PiperOrigin-RevId: 707332453
2024-12-17 18:19:49 -08:00
Changyu ZhuandCopybara-Service dae9e791df Fix missing import in Llama 3 finetuning notebook
PiperOrigin-RevId: 707326361
2024-12-17 18:02:31 -08:00
Changyu ZhuandCopybara-Service 16237de9bb Add fast deployment section to Llama 3.2 deployment notebook
PiperOrigin-RevId: 707274649
2024-12-17 15:32:57 -08:00
Vertex MG TeamandCopybara-Service f5e0394b10 Add H100 80 GB config for Llama 3
PiperOrigin-RevId: 706888765
2024-12-16 17:28:23 -08:00
Vertex MG TeamandCopybara-Service 775aa37b88 Update yolov8 model to use model and endpoint dictionary.
PiperOrigin-RevId: 706750375
2024-12-16 10:13:00 -08:00
Vertex MG TeamandCopybara-Service b43418b17b mediapipe Object detection notebook bug fix and re-formatting
PiperOrigin-RevId: 706673143
2024-12-16 05:26:30 -08:00
Changyu ZhuandCopybara-Service 20138d9333 Add fast deployment section to Llama 3.1 deployment notebook
PiperOrigin-RevId: 705931336
2024-12-13 10:44:11 -08:00
sageof6pathandGitHub fadcb1e618 Peft docker fix (#3751)
* Updated util files to fix peft docker

* fix imports

* fix imports fileutils.py
2024-12-13 13:08:34 +00:00
Vertex MG TeamandCopybara-Service f0ec2acbb8 Update Hex-LLM container URI.
PiperOrigin-RevId: 705656670
2024-12-12 15:53:15 -08:00
Aiden010200andGitHub cc38839e2c Upload xgbranker predictor example (#3737)
* Upload pipeline job example

* Upload pipeline example which can combine other pipeline examples.

* Upload xgbranker predictor example

This example uses aiplatform and xgboost to provide a xgbranker predictor.
2024-12-12 16:03:43 +00:00
Vertex MG TeamandCopybara-Service 94b6a9624b mediapipe Object detection notebook
PiperOrigin-RevId: 705082582
2024-12-11 06:20:13 -08:00
Vertex MG TeamandCopybara-Service 9f8cc0e625 Enable dedicate endpoint for phi3 deployment
PiperOrigin-RevId: 704865877
2024-12-10 15:25:47 -08:00
Vertex MG TeamandCopybara-Service 57cc004855 Update the HF TGI and pytorch-inference notebooks, with the latest container image version.
PiperOrigin-RevId: 704421485
2024-12-09 14:40:05 -08:00
Minwoo ParkandCopybara-Service b9549ccee7 Add Llama 3.3 finetuning notebook.
PiperOrigin-RevId: 703546040
2024-12-06 10:42:19 -08:00
Vertex MG TeamandCopybara-Service 6c7383beaf Adding Qwen2.5-Instruct-32B-AWQ TPU configs to Colab deployment notebook
PiperOrigin-RevId: 703541226
2024-12-06 10:26:35 -08:00
Vertex MG TeamandCopybara-Service 4bec5ef258 Add new Llama 3.3 deployment notebook.
PiperOrigin-RevId: 703537600
2024-12-06 10:15:16 -08:00
Vertex MG TeamandCopybara-Service c2bac62780 tfvision classification notebook
PiperOrigin-RevId: 702659350
2024-12-04 03:25:48 -08:00
Pedro MelendezandGitHub cb7c18e439 Added blog post URL (#3739) 2024-12-03 17:41:39 +00:00
Vertex MG TeamandCopybara-Service e1e4a2ba5e Add vLLM + TPU Llama 3.1 and Qwen 2.5 deployment notebook.
PiperOrigin-RevId: 702091534
2024-12-02 14:46:03 -08:00
Vertex MG TeamandCopybara-Service 2a981b9568 Adding chunked prefill vllm server arg.
PiperOrigin-RevId: 702022997
2024-12-02 11:02:06 -08:00
Vertex MG TeamandCopybara-Service fb10a66d12 A fix in the prediction section
PiperOrigin-RevId: 701847844
2024-12-01 23:09:39 -08:00
Vertex MG TeamandCopybara-Service 2b2019afa6 LLaVA Deployment notebook
PiperOrigin-RevId: 701314497
2024-11-29 10:23:01 -08:00
Vertex MG TeamandCopybara-Service 751b8e7ad5 Avoid copying model artifacts to local GCS and use VERTEX_AI_MODEL_GARDEN_LLAMA_3_1 directly
PiperOrigin-RevId: 700732639
2024-11-27 09:55:49 -08:00
Aiden010200andGitHub 516c395db0 Upload pipeline job example (#3719)
* Upload pipeline example which can combine other pipeline examples.
2024-11-26 15:32:54 +00:00
Vertex MG TeamandCopybara-Service 964c481ed1 Enable dedicate endpoint for timesfm deployment
PiperOrigin-RevId: 700190957
2024-11-25 20:33:57 -08:00
Vertex MG TeamandCopybara-Service 454a90bc33 Enable dedicate endpoint for huggingface tei deployment
PiperOrigin-RevId: 700166368
2024-11-25 18:25:35 -08:00
Vertex MG TeamandCopybara-Service ee43b8c0b7 Add 2H100/4H100 deploy options to llama notebooks
PiperOrigin-RevId: 700107023
2024-11-25 14:45:36 -08:00
Eric DongandGitHub 556f8510f3 fix: Remove example ouptut (#3732) 2024-11-25 21:21:49 +00:00
Eric DongandGitHub c1ba930618 fix: Update the model file path (#3731)
* fix: Update the model file path

* Update the model file path 2
2024-11-25 20:25:48 +00:00
Vertex MG TeamandCopybara-Service 2dc1e96207 Add Llama 3.2 serving notebook.
PiperOrigin-RevId: 700025379
2024-11-25 10:19:14 -08:00
Vertex MG TeamandCopybara-Service 68d5a6d2d2 Enable dedicate endpoint for Llama Guard deployment
PiperOrigin-RevId: 700022732
2024-11-25 10:11:56 -08:00
Vertex MG TeamandCopybara-Service ff2f20fc26 Adding Qwen2/Qwen2.5 TPU configs to Colab deployment notebook
PiperOrigin-RevId: 699209055
2024-11-22 10:07:41 -08:00
Vertex MG TeamandCopybara-Service dd157ca445 A minor fix in the prediction section
PiperOrigin-RevId: 699134138
2024-11-22 05:05:58 -08:00
Vertex MG TeamandCopybara-Service 65d86f57ca Update region suggestion for A100_80GB and H100_80GB gpus
PiperOrigin-RevId: 698841858
2024-11-21 10:55:34 -08:00
0dadbb8400 feat: adding support for Mistral Large 24.11 part2 (#3723)
* feat: adding support for Mistral Large 24.11 part2

* feat: adding support for Mistral Large 24.11 part2

---------

Co-authored-by: denisj3030 <denisj@google.com>
2024-11-21 16:13:13 +00:00
Aaron DietzandGitHub 115413601f Update spark_on_ray_on_vertex_ai.ipynb (#3705)
Fix links for opening the notebook
2024-11-21 15:43:47 +00:00
Vertex MG TeamandCopybara-Service 3e7d427a1b Download only the required files notebook
PiperOrigin-RevId: 698657845
2024-11-20 23:04:15 -08:00
Vertex MG TeamandCopybara-Service 194978fd1c A minor fix in the HexLLM deploy section
PiperOrigin-RevId: 698420315
2024-11-20 09:36:56 -08:00
Vertex MG TeamandCopybara-Service 877305d425 Fix lint issues
PiperOrigin-RevId: 698241868
2024-11-19 20:47:46 -08:00
Vertex MG TeamandCopybara-Service 37c88851bd Enable dedicate endpoint for model_garden_gemma_finetuning_on_vertex.ipynb
PiperOrigin-RevId: 697850787
2024-11-18 20:12:55 -08:00
Vertex MG TeamandCopybara-Service 3b17b30051 Deprecate the Pytorch PEFT notebook. Notebooks such as model_garden_pytorch_llama3_1_finetuning.ipynb demonstrate the usage of the peft in VMG.
PiperOrigin-RevId: 697850024
2024-11-18 20:10:01 -08:00
Vertex MG TeamandCopybara-Service 8883d8e211 Enable dedicate endpoint for model_garden_pytorch_llama3_1_deployment.ipynb
PiperOrigin-RevId: 697498389
2024-11-17 22:16:51 -08:00
Vertex MG TeamandCopybara-Service 0879ef0057 Fix timestamp parameter in prediction section.
PiperOrigin-RevId: 697114085
2024-11-16 00:29:07 -08:00
Vertex MG TeamandCopybara-Service 4adff04d06 Enable dedicate endpoint for model_garden_pytorch_mixtral_deployment.ipynb
PiperOrigin-RevId: 696942601
2024-11-15 11:07:19 -08:00
Vertex MG TeamandCopybara-Service e6a8641896 Enable dedicate endpoint for model_garden_pytorch_qwen2_deployment.ipynb
PiperOrigin-RevId: 696937628
2024-11-15 10:52:28 -08:00
Pedro MelendezandGitHub a740092ae2 Added notebook titled "backoff_and_retry_for_LLMs.ipynb" to the /notebooks/community/generative_ai/ directory (#3706)
* Adding backoff and retry notebook

* Formatted notebook

* Formatted notebook

* Formatted notebook

* Formatted notebook

* Changed URL

* Format changes

* Remoevd URL

* Added formatting

* Added codeowner entry

* Added summary

* Added note about costs

* Lint format
2024-11-15 18:23:02 +00:00
Eric DongandGitHub 6289d1f0a6 feat: exclude model garden dockerfilers from dependabot checks (#3712) 2024-11-15 15:05:21 +00:00
Eric DongandGitHub 1301bd2ad2 Revert "Bump deepspeed (#3658)" (#3711)
This reverts commit 6c9bdba210.
2024-11-14 16:51:13 +00:00
Vertex MG TeamandCopybara-Service 21d5f7edcb Download only the required files notebook
PiperOrigin-RevId: 695745528
2024-11-12 08:31:36 -08:00
Vertex MG TeamandCopybara-Service 7f13964632 Update the notebook to remove the duplicated model agreement step.
PiperOrigin-RevId: 695575396
2024-11-11 20:23:02 -08:00
Vertex MG TeamandCopybara-Service 231dbae9e9 Update the recursion mae local inference notebook, to always download the model weights from HuggingFace Hub, instead of directly load the model weights from HF. The latter, for some reason, cannot find the model.safetensors file from the repo.
PiperOrigin-RevId: 695573672
2024-11-11 20:13:58 -08:00
Vertex MG TeamandCopybara-Service bf0507bfda OWL-ViT2 Notebook
PiperOrigin-RevId: 695315779
2024-11-11 06:46:38 -08:00
Mend RenovateandGitHub 3361dc70d1 chore(deps): update dependency nbqa to v1.9.1 (#3699) 2024-11-11 13:29:15 +00:00
Mend RenovateandGitHub e05c83832d chore(config): migrate config renovate.json (#3677) 2024-11-11 13:28:31 +00:00
Aaron DietzandGitHub e2f1aaed9d Update feature_store_streaming_ingestion_sdk.ipynb (#3694)
Moved mention of Feature Store (Legacy) so that it shows up in our notebook description when we generate the list of notebooks.
2024-11-11 13:26:54 +00:00
5d71616e8b Add ability to copy specific model artifacts (#3700)
Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2024-11-11 13:25:24 +00:00
57f598052e Update usage tracking metrics in the finetuning notebooks (#3695)
Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2024-11-11 13:24:27 +00:00
Vertex MG TeamandCopybara-Service 439a686f16 Create a notebook for the image_feature_extraction_mae model.
PiperOrigin-RevId: 694586865
2024-11-08 11:54:42 -08:00
Vertex MG TeamandCopybara-Service b99a1d8f42 Create a notebook for local inference for the partner recursion mae model.
PiperOrigin-RevId: 694248672
2024-11-07 14:28:31 -08:00
Vertex MG TeamandCopybara-Service 2d1339731f Present --data-parallel-size for Hex-LLM deployment; add advanced config for CodeGemma.
PiperOrigin-RevId: 693598038
2024-11-05 22:48:19 -08:00
Vertex MG TeamandCopybara-Service c71edd71e7 OWL-ViT Notebook
PiperOrigin-RevId: 693140455
2024-11-04 17:09:00 -08:00
Vertex MG TeamandCopybara-Service e0d078b6bb The HF TGI notebook should use the TGI 2.3 serving container, which is the latest
PiperOrigin-RevId: 693031093
2024-11-04 11:15:58 -08:00
Sujit KhasnisandGitHub 267b45f6aa feat:claude notebook 3.5 update (#3690) 2024-11-04 18:43:21 +00:00
Vertex MG TeamandCopybara-Service f1bb2f7aea Stable Diffusion 2.1 Dreambooth Finetuning notebook
PiperOrigin-RevId: 691686520
2024-10-30 23:24:39 -07:00
Vertex MG TeamandCopybara-Service c2132249c9 Fix Llama 3.1 deployment notebook vLLM version.
PiperOrigin-RevId: 691669244
2024-10-30 21:58:30 -07:00
Vertex MG TeamandCopybara-Service d4fdc15892 Update vLLM version and embedded links in Llama 3.2 deployment notebook.
PiperOrigin-RevId: 691604159
2024-10-30 17:13:44 -07:00
Sujit KhasnisandGitHub 43bc13ee66 feat: Claude region update(euw1) (#3685) 2024-10-30 16:52:16 +00:00
f11502ec80 Fix pip dependency error (#3683)
Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2024-10-30 12:38:10 +00:00
Vertex MG TeamandCopybara-Service 7f7cf51dd7 Enable dedicate endpoint for Mistral deployment
PiperOrigin-RevId: 691250786
2024-10-29 19:44:48 -07:00
Vertex MG TeamandCopybara-Service f13e086012 Add optimized vLLM to Mixtral deployment notebook.
PiperOrigin-RevId: 691209007
2024-10-29 16:52:16 -07:00
Vertex MG TeamandCopybara-Service b995ea7adf Add optimized vLLM to Llama 3.1 deployment notebook.
PiperOrigin-RevId: 691207909
2024-10-29 16:48:35 -07:00
Vertex MG TeamandCopybara-Service b183a3d796 Update model_garden_pytorch_sd_xl_finetuning_dreambooth_lora.ipynb
PiperOrigin-RevId: 689953226
2024-10-25 16:47:05 -07:00
Vertex MG TeamandCopybara-Service e500467637 Update Mediapipe Image Generation notebook
PiperOrigin-RevId: 689923141
2024-10-25 14:59:38 -07:00
Vertex MG TeamandCopybara-Service ad6035ae80 No public description
PiperOrigin-RevId: 688364622
2024-10-25 14:59:28 -07:00
Ivan NardiniandGitHub 731bc7a780 fix: Adding autoscaling to RoV cluster management notebook (#3674)
* new notebook

* fix issue

* linter passed

* change pname

* linter passed

* remove script
2024-10-25 17:28:47 +00:00
lee1premiumandGitHub 6dd9c27781 feat: Getting tuned embeddings using text-embedding-005. (#3673)
* feat: Getting tuned embeddings using text-embedding-005.

* feat: Getting tuned embeddings using text-embedding-005.

* feat: Getting tuned embeddings using text-embedding-005.
2024-10-24 00:49:45 +00:00
Mend RenovateandGitHub 830dc9f1c3 chore(deps): update dependency pyupgrade to v3.19.0 (#3667) 2024-10-23 12:55:43 +00:00
lee1premiumandGitHub 23a1398504 feat: Getting embeddings using text-embedding-005. (#3672)
* feat: Getting embeddings using text-embedding-005.

* feat: Getting embeddings using text-embedding-005.

* feat: Getting embeddings using text-embedding-005.

* feat: Getting embeddings using text-embedding-005.
2024-10-23 12:55:15 +00:00
Sujit KhasnisandGitHub 7b86dec7f7 fix: update titles, headers (#3670) 2024-10-22 20:24:40 +00:00
Sujit KhasnisandGitHub 02f1edd942 fix: model ordering in dropdown (#3669) 2024-10-22 16:15:10 +00:00
Sujit KhasnisandGitHub b8a3fafb9a feat: Claude notebook update (#3668) 2024-10-22 15:51:51 +00:00
Tianrui YangandGitHub 26849e7f20 fea: add PSC example code in Feature Store embedding notebook (#3663) 2024-10-22 14:50:02 +00:00
Shawn YangandCopybara-Service a96c2926af fix: Fix Notebook format issue in Preview by removing output.
PiperOrigin-RevId: 688364596
2024-10-21 19:54:06 -07:00
Shawn YangandCopybara-Service 90eaf5c7fe fix: Fix Notebook format issue in Preview.
PiperOrigin-RevId: 688317050
2024-10-21 16:44:52 -07:00
Shawn YangandCopybara-Service d6ad52e2cd feat: Update Reasoning Engine + Llama 3.1 models notebook with Function Calling Agent.
PiperOrigin-RevId: 688245912
2024-10-21 13:09:51 -07:00
Vertex MG TeamandCopybara-Service e24c20b372 Adding Phi-3-medium TPU configs to Colab deployment notebook
PiperOrigin-RevId: 687430227
2024-10-18 14:43:49 -07:00
Vertex MG TeamandCopybara-Service fb1871d100 Update Llama 3.1 MaaS naming.
PiperOrigin-RevId: 687362552
2024-10-18 11:09:18 -07:00
Vertex MG TeamandCopybara-Service a7dd5aa5a9 Adding Phi-3-mini TPU configs to Colab deployment notebook
PiperOrigin-RevId: 687333561
2024-10-18 09:40:46 -07:00
Vertex MG TeamandCopybara-Service 307f1a8d41 Reformat Instant ID notebook
PiperOrigin-RevId: 687309283
2024-10-18 08:20:02 -07:00
Vertex MG TeamandCopybara-Service 53895497c0 Update chat completions URL for dedicated endpoint in Gemma and Llama deployment notebooks.
PiperOrigin-RevId: 687307214
2024-10-18 08:11:33 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
6c9bdba210 Bump deepspeed (#3658)
Bumps [deepspeed](https://github.com/microsoft/DeepSpeed) from 0.14.4 to 0.15.1.
- [Release notes](https://github.com/microsoft/DeepSpeed/releases)
- [Commits](https://github.com/microsoft/DeepSpeed/compare/v0.14.4...v0.15.1)

---
updated-dependencies:
- dependency-name: deepspeed
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2024-10-18 14:31:08 +00:00
Vertex MG TeamandCopybara-Service 133395a908 Use /tmp as the dataset directory since the notebook executor environment may not use /content as the base directory
PiperOrigin-RevId: 687194370
2024-10-18 00:25:47 -07:00
Vertex MG TeamandCopybara-Service bd29537128 Remove Weaviate and Pinecone notebooks as they are relocated to another directory.
PiperOrigin-RevId: 686788959
2024-10-16 23:39:33 -07:00
Vertex MG TeamandCopybara-Service 3ac1c34633 Reformat stable diffusion gradio notebook
PiperOrigin-RevId: 686463355
2024-10-16 04:53:50 -07:00
Vertex MG TeamandCopybara-Service 2e56046d25 Add Llama 3.2 evaluation notebook.
PiperOrigin-RevId: 686349456
2024-10-15 20:51:39 -07:00
Vertex MG TeamandCopybara-Service f6dd048f98 Update Gemma PEFT notebook to use the new training docker image and add more instructions.
PiperOrigin-RevId: 686347835
2024-10-15 20:44:39 -07:00
Vertex MG TeamandCopybara-Service 9da812bf2e Update model_garden_pytorch_bart_large_cnn.ipynb
PiperOrigin-RevId: 686230074
2024-10-15 14:00:45 -07:00
Sujit KhasnisandGitHub 11d361f727 refactor: verbiage updates, minor code updates (#3655) 2024-10-15 18:13:14 +00:00
Mend RenovateandGitHub de9dd0f850 Update python Docker tag to v3.13 (#3626) 2024-10-15 13:49:03 +00:00
482170765andGitHub f20e9590a5 Update a custom job example (#3639)
* This is a custom job example of kfp v2
2024-10-15 13:48:37 +00:00
Mend RenovateandGitHub 5443944739 chore(deps): update dependency pyupgrade to v3.18.0 (#3640) 2024-10-15 13:48:00 +00:00
Sujit KhasnisandGitHub f2ee8a582f feat: NVIDIA NIM on Vertex Ai walkthrough (#3579)
* feat: NVIDIA NIM on Vertex Ai walkthrough

* fix: PR comments resolved

* fix: PR comments resolved#2

* fix: excep handling

* fix: handle if model not uploaded
2024-10-15 13:47:25 +00:00
Vertex MG TeamandCopybara-Service 1d3d3980e4 Fix the issue that the model and endpoint are not stored in the dictionary.
PiperOrigin-RevId: 685966395
2024-10-14 22:50:05 -07:00
Vertex MG TeamandCopybara-Service f9b371e030 Update the notebook to use A100 80GB for vLLM deployment.
PiperOrigin-RevId: 685934295
2024-10-14 20:42:11 -07:00
Dustin LuongandCopybara-Service a1d4026020 Update the notebook to use the Vertex SDK to send requests to the deployed endpoint instead of openai SDK.
PiperOrigin-RevId: 685713862
2024-10-14 08:23:23 -07:00
Vertex MG TeamandCopybara-Service eb166ed3df Update the default model to whisper-large-v3-turbo and add placeholder for language.
PiperOrigin-RevId: 685618090
2024-10-14 01:37:22 -07:00
Vertex MG TeamandCopybara-Service 08278233ff Update Qwen2 deployment notebook to support A100, H100 and A100 80GB and fix lint issues.
PiperOrigin-RevId: 685555814
2024-10-13 20:41:34 -07:00
Vertex MG TeamandCopybara-Service 51bf2a38ff Minor fix
PiperOrigin-RevId: 685020524
2024-10-11 18:22:58 -07:00
Vertex MG TeamandCopybara-Service 2dede7552b Added more details about notebook parameters for predict section. Also gave storage.objectViewer access service account for buckets used for predict section.
PiperOrigin-RevId: 684700814
2024-10-10 22:12:39 -07:00
Vertex MG TeamandCopybara-Service 905e89ad64 Add new notebook for Knowledge Engine API with Pinecone.
PiperOrigin-RevId: 684666941
2024-10-10 19:56:49 -07:00
Vertex MG TeamandCopybara-Service ec297e663c Add labels to finetuning notebooks
PiperOrigin-RevId: 684655746
2024-10-10 19:11:32 -07:00
Vertex MG TeamandCopybara-Service d15caa9ed5 Enable dedicate endpoint for TGI Gemma2 predict and chat completion
PiperOrigin-RevId: 684623757
2024-10-10 16:55:28 -07:00
Vertex MG TeamandCopybara-Service df55634281 Added more details to deployment section of notebook.
PiperOrigin-RevId: 684343445
2024-10-10 01:19:34 -07:00
Vertex MG TeamandCopybara-Service af84cdc46b Fix the error on stable_diffusion_gradio notebook for text2image models.
PiperOrigin-RevId: 684108916
2024-10-09 11:26:08 -07:00
Shawn YangandCopybara-Service e946e2b304 feat: Add Reasoning Engine with Llama 3.1 models notebook.
PiperOrigin-RevId: 684083427
2024-10-09 10:16:47 -07:00
Vertex MG TeamandCopybara-Service d7550f4756 LLaVA Deployment notebook
PiperOrigin-RevId: 683832899
2024-10-08 18:20:17 -07:00
Vertex MG TeamandCopybara-Service 3088fc896d Update vLLM and chat completions prediction samples.
PiperOrigin-RevId: 683698377
2024-10-08 11:25:16 -07:00
Vertex MG TeamandCopybara-Service 75cf2fe9b9 Add Qwen2.5 related updates to the notebook
PiperOrigin-RevId: 683423186
2024-10-07 19:46:47 -07:00
Vertex MG TeamandCopybara-Service d68494bd3b Update Prompt Guard deployment notebook.
PiperOrigin-RevId: 683352100
2024-10-07 15:42:32 -07:00
02f66799c8 Add README.md for PEFT train docker template (#3625)
Co-authored-by: minwoopark <minwoopark@google.com>
2024-10-07 20:54:53 +00:00
482170765andGitHub 6b10fe9636 Upload classifier predictor sample of sklearn (#3623)
Upload a classifier predictor sample using scikit-learn lib.
2024-10-07 20:54:18 +00:00
Mend RenovateandGitHub 4c3693929d Update dependency black to v24.10.0 (#3624) 2024-10-07 20:49:06 +00:00
Vertex MG TeamandCopybara-Service 6bff6af34a Fix lint issues in the notebooks
PiperOrigin-RevId: 682565175
2024-10-04 22:19:13 -07:00
Vertex MG TeamandCopybara-Service 219474b39c Add new notebook for RAG API with Weaviate.
PiperOrigin-RevId: 682497053
2024-10-04 16:58:17 -07:00
Vertex MG TeamandCopybara-Service bc152ece40 Fix the RAG notebook format.
PiperOrigin-RevId: 682352935
2024-10-04 09:42:36 -07:00
Vertex MG TeamandCopybara-Service 761b918de1 Use 'import datetime' instead of 'from datetime import datetime'
PiperOrigin-RevId: 682309409
2024-10-04 07:14:55 -07:00
Vertex MG TeamandCopybara-Service 6a574ea15d Update the docker image for Gemma finetuning.
PiperOrigin-RevId: 682163952
2024-10-03 21:52:34 -07:00
Vertex MG TeamandCopybara-Service c5212982d0 Fix the RAG notebook link.
PiperOrigin-RevId: 682031626
2024-10-03 14:27:27 -07:00
Vertex MG TeamandCopybara-Service 3986128f78 Autogluon notebook
PiperOrigin-RevId: 681881078
2024-10-03 08:05:26 -07:00
Vertex MG TeamandCopybara-Service 575025a1cd Fix format issues
PiperOrigin-RevId: 681739757
2024-10-02 23:33:38 -07:00
Vertex MG TeamandCopybara-Service fe42990c9d Update Llama 2 evaluation notebook.
PiperOrigin-RevId: 681657733
2024-10-02 17:49:18 -07:00
Dustin LuongandCopybara-Service bd3283b2a6 Internal change.
PiperOrigin-RevId: 681516025
2024-10-02 17:49:03 -07:00
24c001eaf6 Update PEFT train docker code (#3613)
Co-authored-by: minwoopark <minwoopark@google.com>
2024-10-02 19:13:30 +00:00
Aaron DietzandGitHub e500e70580 Update spark_on_ray_on_vertex_ai.ipynb (#3605)
Revised the overview so our github notebook list script will pick up the "last line" in the overview and start including this notebook in the output.
2024-10-02 18:13:15 +00:00
Vertex MG TeamandCopybara-Service bfb3775813 Fix the region for llama3 hex-llm chat completion
PiperOrigin-RevId: 681511627
2024-10-02 10:47:45 -07:00
Vertex MG TeamandCopybara-Service 538432df5c Fix the region for llama3 hex-llm chat completion
PiperOrigin-RevId: 681167715
2024-10-01 14:33:19 -07:00
Minwoo ParkandCopybara-Service 7764173895 Minor copyright year update.
PiperOrigin-RevId: 681141148
2024-10-01 13:21:14 -07:00
Vertex MG TeamandCopybara-Service 142237c1f5 Add region suggestion for A100_80GB and H100_80GB to notebooks.
PiperOrigin-RevId: 681090786
2024-10-01 11:07:54 -07:00
Vertex MG TeamandCopybara-Service 791a66ec54 Fix formatting issue
PiperOrigin-RevId: 681076924
2024-10-01 10:35:31 -07:00
Vertex MG TeamandCopybara-Service 8379361eff Fix typo - Use "prompt" in image captioning sample request.
PiperOrigin-RevId: 681074023
2024-10-01 10:27:17 -07:00
Vertex MG TeamandCopybara-Service a9f6b2d7f0 Adding Phi-3.5-MoE-instruct variant to Phi-3 deployment notebook.
PiperOrigin-RevId: 681057221
2024-10-01 09:44:32 -07:00
Vertex MG TeamandCopybara-Service 8f077f9bc2 This notebook demonstrates deploying prebuilt Whisper Large models.
PiperOrigin-RevId: 681039051
2024-10-01 08:55:51 -07:00
Sujit KhasnisandGitHub c5ec7e5d83 feat: mistral ai sdk support for vertexai (#3601)
* feat: mistral ai sdk support fro vertexai

* feat: mistral ai sdk support fro vertexai, token fix

* feat: mistral ai sdk support fro vertexai, excep handling
2024-10-01 14:59:10 +00:00
Vertex MG TeamandCopybara-Service daf92aa672 Fix lint issue
PiperOrigin-RevId: 680835893
2024-09-30 20:57:00 -07:00
Bhaskar GoyalandGitHub 4070d8b0e5 Update region for Claude Haiku and Sonnet 3.5 (#3606) 2024-09-30 23:06:35 +00:00
Vertex MG TeamandCopybara-Service a31f1e037e Use dedicated endpoint as default for Gemma deployment on vertex
PiperOrigin-RevId: 680645756
2024-09-30 11:09:15 -07:00
sharkeshdandGitHub dc9d1032a2 Update Dockerfile (#3576)
This multi-stage approach keeps your final image clean and lightweight.
2024-09-30 17:45:37 +00:00
Aaron DietzandGitHub d5f93bf8a9 Update spark_on_ray_on_vertex_ai.ipynb (#3594)
Updated branding/name of Vertex AI Workbench, added "Overview" heading.

Why? Not having an "Overview" heading prevents this notebook from getting picked up in our notebook list output.
2024-09-30 17:45:06 +00:00
Aaron DietzandGitHub 6a3bc32e45 Update xai_text_classification_feature_attributions.ipynb (#3596)
Removed bolding that doesn't render properly when we port the content to our docs
2024-09-30 17:44:09 +00:00
Aaron DietzandGitHub 1d1d4a586d Update get_started_with_model_monitoring_setup.ipynb (#3597)
Removed bolding that doesn't render properly when we port the content to our docs
2024-09-30 17:43:26 +00:00
Aaron DietzandGitHub b3c7ecb8cd Update hyperparameter_tuning_xgboost.ipynb (#3598)
Removed bolding that doesn't render properly when we port the content to our docs
2024-09-30 17:42:47 +00:00
Aaron DietzandGitHub cec4e447ab Update chicago_taxi_fare_prediction.ipynb (#3599)
Removed bolding that doesn't render properly when we port the content to our docs
2024-09-30 17:41:55 +00:00
Vertex MG TeamandCopybara-Service e61be64040 Use dedicated endpoint as default for Gemma2 deployment on vertex
PiperOrigin-RevId: 680634868
2024-09-30 10:41:39 -07:00
Minwoo ParkandCopybara-Service 148a6fad99 Add instruction to run TensorBoard in Cloud Shell
PiperOrigin-RevId: 679746165
2024-09-27 15:20:04 -07:00
Aaron DietzandGitHub ca84581ed0 Fixed missing word in predictive_maintenance_usecase.ipynb (#3581)
Added a word to make a sentence parse correctly.
2024-09-26 20:49:03 +00:00
Vertex MG TeamandCopybara-Service 4d1c59cba4 Minor fix in llama3.2 notebook
PiperOrigin-RevId: 679190955
2024-09-26 09:59:49 -07:00
Vertex MG TeamandCopybara-Service 5cd0c0e782 Minor change to the VOT and ZipNeRF notebooks.
PiperOrigin-RevId: 679159158
2024-09-26 08:27:19 -07:00
Vertex MG TeamandCopybara-Service 402231e2b6 Update finetuning notebook with stable_20240909 training image
PiperOrigin-RevId: 679139242
2024-09-26 07:25:59 -07:00
Vertex MG TeamandCopybara-Service 38e1a46a7b Fix typo in vllm args
PiperOrigin-RevId: 679131835
2024-09-26 06:57:15 -07:00
0727e19520 Add vmg templates, dataset_validation_util and update common_util (#3586)
* Add vmg templates, dataset_validation_util and update common_util

* Add name to CODEOWNERS

* Update common_util.py

---------

Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2024-09-26 13:16:09 +00:00
Vertex MG TeamandCopybara-Service 30c3e627a7 Update the RAG notebook for Llama3 models.
PiperOrigin-RevId: 679007617
2024-09-25 23:31:06 -07:00
Vertex MG TeamandCopybara-Service bdea63ec41 Adding Phi-3.5-mini-instruct variant to Phi-3 deployment notebook.
PiperOrigin-RevId: 678888299
2024-09-25 16:24:22 -07:00
Vertex MG TeamandCopybara-Service 2e5410fe35 No public description
PiperOrigin-RevId: 678885403
2024-09-25 16:15:39 -07:00
Vertex MG TeamandCopybara-Service 6f9813cba5 Update sample requests in Llama 3.2 OpenAI MaaS notebook.
PiperOrigin-RevId: 678828203
2024-09-25 13:35:19 -07:00
Vertex MG TeamandCopybara-Service 115d8f991a Update Llama 3.2 OpenAI MaaS notebook.
PiperOrigin-RevId: 678781425
2024-09-25 11:31:37 -07:00
Vertex MG TeamandCopybara-Service 7fd31a65ae Support deploying llama 3.2 guard models on model garden.
PiperOrigin-RevId: 678760438
2024-09-25 10:44:40 -07:00
Vertex MG TeamandCopybara-Service 883cc93ab7 Add Llama 3.2 OpenAI MaaS notebook.
PiperOrigin-RevId: 678753976
2024-09-25 10:27:22 -07:00
Changyu ZhuandCopybara-Service 1363868542 Add streaming chat completions example to the HF TGI notebook
PiperOrigin-RevId: 678753372
2024-09-25 10:26:02 -07:00
Vertex MG TeamandCopybara-Service 09122d1479 Support deploying llama 3.2 models on model garden.
PiperOrigin-RevId: 678752790
2024-09-25 10:24:34 -07:00
Minwoo ParkandCopybara-Service 2adb19be7c Resolving conflict
PiperOrigin-RevId: 678701565
2024-09-25 08:00:32 -07:00
Aaron DietzandGitHub c588d81d02 Update notebook_template_review.py (#3578)
Removed "external" class for the "open-notebook-in..." links. Style guide indicates we should avoid using the "external" class.
2024-09-25 12:43:15 +00:00
ShunpeIIIandGitHub e51af44898 Remove the link of the notebook that has been moved to the community. (#3571) 2024-09-23 14:39:40 +00:00
Vertex MG TeamandCopybara-Service 0872ce0a87 Fix the pip install command in the notebooks.
PiperOrigin-RevId: 677767387
2024-09-23 06:31:07 -07:00
Vertex MG TeamandCopybara-Service 867d7b7410 Fix max_context_length in Qwen2 deployment notebook.
PiperOrigin-RevId: 677762785
2024-09-23 06:12:36 -07:00
Vertex MG TeamandCopybara-Service d8e628b3ab Add Flux gradio notebook
PiperOrigin-RevId: 676901084
2024-09-20 10:57:44 -07:00
Vertex MG TeamandCopybara-Service dd4767546e Update gradio notebooks for Instant ID and Stable Diffusion
PiperOrigin-RevId: 676476082
2024-09-19 10:45:16 -07:00
Vertex MG TeamandCopybara-Service d185eff6ca Add region option to model garden notebooks.
PiperOrigin-RevId: 676473760
2024-09-19 10:39:10 -07:00
Vertex MG TeamandCopybara-Service f00c9cdeef Update Instant ID notebook
PiperOrigin-RevId: 676466759
2024-09-19 10:21:26 -07:00
Vertex MG TeamandCopybara-Service 12bd07e8fc Fix use dedicated endpoint parameter type in Pytorch Gemma Serving
PiperOrigin-RevId: 676433256
2024-09-19 08:55:18 -07:00
Vertex MG TeamandCopybara-Service 61d866aef3 Update deploy function.
PiperOrigin-RevId: 676408091
2024-09-19 07:37:53 -07:00
Vertex MG TeamandCopybara-Service bb2c8c502d Fix use dedicated endpoint parameter type in Gemma Serving
PiperOrigin-RevId: 676121299
2024-09-18 14:00:08 -07:00
Vertex MG TeamandCopybara-Service ea7915b38a Use standard id as MODEL_ID.
PiperOrigin-RevId: 675778051
2024-09-17 18:04:39 -07:00
Dustin LuongandGitHub b5ee2ea16b Copy Vertex MG files to notebook folder (#3560) 2024-09-18 00:47:12 +00:00
Vertex MG TeamandCopybara-Service 6daebf78bc Support deploying Hex-LLM on multi-hosts TPU, like v5e-16.
PiperOrigin-RevId: 675736856
2024-09-17 15:42:15 -07:00
Changyu ZhuandCopybara-Service 2752658b6d Minor updates to the MoViNet notebooks
PiperOrigin-RevId: 675723676
2024-09-17 15:03:58 -07:00
Vertex MG TeamandCopybara-Service 55ed17d227 Add dedicated endpoint support to Gemma Serving
PiperOrigin-RevId: 675627850
2024-09-17 10:49:28 -07:00
yutatanamotoandGitHub 3808d495fd fix official sample notebook for vector search (#3500)
* fix restriction declaration (allow_list → allow) for vector search index

* fix folder name in CODEOWNERS (/matching_engine → /vector_search)
2024-09-17 12:45:20 +00:00
Vertex MG TeamandCopybara-Service 6c78212f92 Add enable_model_cpu_offload option for flux example in local inference notebook
PiperOrigin-RevId: 675227512
2024-09-16 11:31:55 -07:00
william-ChengChungChuandGitHub b6771091fa Upload classifier predictor sample of xgboost (#3552)
This example uses aiplatform and xgboost to provide a classifier predictor.
2024-09-16 14:57:00 +00:00
Vertex MG TeamandCopybara-Service 6a62443d01 Stable Diffusion v2.1 notebook
PiperOrigin-RevId: 674233624
2024-09-13 03:52:22 -07:00
Vertex MG TeamandCopybara-Service eb7d8456be Add instructions for applying Llama Guard on MaaS.
PiperOrigin-RevId: 674144050
2024-09-12 22:14:50 -07:00
Changyu ZhuandCopybara-Service df4adf877e internal change
PiperOrigin-RevId: 673999030
2024-09-12 22:14:36 -07:00
Dustin LuongandGitHub 75fe45e07c Copy Vertex MG files to notebook folder (#3542) 2024-09-12 21:20:54 +00:00
Ayush AgrawalandCopybara-Service 0f53fec3fd Add Weaviate Vector DB option for corpus creation to rag notebook
PiperOrigin-RevId: 673948076
2024-09-12 12:08:17 -07:00
Vertex MG TeamandCopybara-Service 3acee41659 Use standard id as MODEL_ID.
PiperOrigin-RevId: 673693818
2024-09-11 23:21:44 -07:00
Vertex MG TeamandCopybara-Service e5a016651e Use standard id as MODEL_ID.
PiperOrigin-RevId: 673693121
2024-09-11 23:19:28 -07:00
Vertex MG TeamandCopybara-Service 41add3b060 Stable Diffusion XL 1.0 notebook
PiperOrigin-RevId: 673679901
2024-09-11 22:25:20 -07:00
Vertex MG TeamandCopybara-Service d05c27e077 Stable Diffusion XL Lightning notebook
PiperOrigin-RevId: 673675311
2024-09-11 22:08:20 -07:00
Vertex MG TeamandCopybara-Service 370f6cd0bf Use standard id as MODEL_ID.
PiperOrigin-RevId: 673522391
2024-09-11 13:49:45 -07:00
Sujit KhasnisandGitHub c53d3f215b feat: add support for jamba-large in euw4 (#3540) 2024-09-11 20:01:41 +00:00
Vertex MG TeamandCopybara-Service 760ad8ff32 Update OpenAI chat completions MaaS notebook with new variants.
PiperOrigin-RevId: 673208459
2024-09-10 20:33:24 -07:00
Vertex MG TeamandCopybara-Service f128e42a57 Use standard id as MODEL_ID.
PiperOrigin-RevId: 673163937
2024-09-10 17:16:56 -07:00
Dustin LuongandCopybara-Service fc917137da Use vLLM docker to deploy dolly-v2 model.
PiperOrigin-RevId: 673051532
2024-09-10 12:00:12 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
d7225f4ec7 chore(deps): bump tensorflow (#3532)
Bumps [tensorflow](https://github.com/tensorflow/tensorflow) from 2.7.2 to 2.12.1.
- [Release notes](https://github.com/tensorflow/tensorflow/releases)
- [Changelog](https://github.com/tensorflow/tensorflow/blob/master/RELEASE.md)
- [Commits](https://github.com/tensorflow/tensorflow/compare/v2.7.2...v2.12.1)

---
updated-dependencies:
- dependency-name: tensorflow
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2024-09-10 13:57:47 +00:00
Changyu ZhuandCopybara-Service e32337ac46 Add a template HF Pytorch Inference model deployment notebook
PiperOrigin-RevId: 672727375
2024-09-09 17:18:55 -07:00
Vertex MG TeamandCopybara-Service df1e6ea815 Use standard id as MODEL_ID.
PiperOrigin-RevId: 672664027
2024-09-09 13:58:16 -07:00
Dustin LuongandCopybara-Service d6a8ed040a Update deploy function.
PiperOrigin-RevId: 672608260
2024-09-09 11:17:52 -07:00
Mend RenovateandGitHub edeac27417 chore(deps): update dependency flake8 to v7.1.1 (#3382) 2024-09-09 12:32:46 +00:00
Vertex MG TeamandCopybara-Service d4dfcad09e Use standard id as MODEL_ID.
PiperOrigin-RevId: 672348661
2024-09-08 16:41:53 -07:00
Vertex MG TeamandCopybara-Service fcb353af5e Use standard id as MODEL_ID.
PiperOrigin-RevId: 672347208
2024-09-08 16:30:50 -07:00
Vertex MG TeamandCopybara-Service 47d04c117d Fix some formatting issues in the timesfm templated notebook.
PiperOrigin-RevId: 671907130
2024-09-06 15:19:01 -07:00
Vertex MG TeamandCopybara-Service 94168e1e10 Adding Phi-3-mini-128k variant to Phi-3 deployment notebook.
PiperOrigin-RevId: 671736868
2024-09-06 06:38:17 -07:00
Kaushik KoiladaandGitHub 276ff3779e fix: colab enterprise link fix (#3509) 2024-09-06 12:33:04 +00:00
Changyu ZhuandCopybara-Service 6a858a7b1f Add a template HF TEI model deployment notebook
PiperOrigin-RevId: 671513324
2024-09-05 14:34:17 -07:00
Vertex MG TeamandCopybara-Service a18c7bea12 Keras Yolov8 notebook
PiperOrigin-RevId: 671432374
2024-09-05 10:54:05 -07:00
Vertex MG TeamandCopybara-Service e87127955b Fix typo in pytorch llava notebook.
PiperOrigin-RevId: 671419544
2024-09-05 10:21:12 -07:00
Vertex MG TeamandCopybara-Service 484e6536ec Remove runwayml/stable-diffusion-v1-5 and runwayml/stable-diffusion-inpainting artifacts in the notebooks:
PiperOrigin-RevId: 671385551
2024-09-05 08:39:48 -07:00
Coby BenvenisteandGitHub 1c25cf5de1 Add Target Modules flag to allow passing in the target modules to the lora fine tuning (#3441) 2024-09-05 13:31:11 +00:00
Liang LongandGitHub 66e7e4effc Upload a model evaluation example (#3507)
* Upload a model evaluation example by kfp v2.
2024-09-05 13:28:52 +00:00
Vertex MG TeamandCopybara-Service ffd8d88419 Update OpenAI chat completions MaaS notebook.
PiperOrigin-RevId: 671146111
2024-09-04 16:43:04 -07:00
Yichen ZhouandCopybara-Service 9cee408bf6 Update TimesFM notebook.
1. Updated the `predict` call with latest signatures.
2. Added new code examples for covariate support.

PiperOrigin-RevId: 671136154
2024-09-04 16:09:01 -07:00
yexing111andGitHub 9b23762537 Optimized colab changes for PSC ga (#3514) 2024-09-04 18:51:35 +00:00
Changyu ZhuandCopybara-Service 06f4470396 Update Hugging Face local inference notebook to include more examples
PiperOrigin-RevId: 671039421
2024-09-04 11:27:13 -07:00
Changyu ZhuandCopybara-Service 92e166dab2 Add a template HF TGI model deployment notebook
PiperOrigin-RevId: 671036522
2024-09-04 11:18:50 -07:00
Changyu ZhuandCopybara-Service 9eb589baed Add the Gemma-2-2b-it public endpoint to the steaming chat completions notebook
PiperOrigin-RevId: 671027759
2024-09-04 10:55:50 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2627643386 chore(deps): bump tensorflow (#3510)
Bumps [tensorflow](https://github.com/tensorflow/tensorflow) from 2.7.2 to 2.12.1.
- [Release notes](https://github.com/tensorflow/tensorflow/releases)
- [Changelog](https://github.com/tensorflow/tensorflow/blob/master/RELEASE.md)
- [Commits](https://github.com/tensorflow/tensorflow/compare/v2.7.2...v2.12.1)

---
updated-dependencies:
- dependency-name: tensorflow
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2024-09-04 12:28:29 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
962b0a606b chore(deps): bump torch (#3511)
Bumps [torch](https://github.com/pytorch/pytorch) from 1.13.1 to 2.2.0.
- [Release notes](https://github.com/pytorch/pytorch/releases)
- [Changelog](https://github.com/pytorch/pytorch/blob/main/RELEASE.md)
- [Commits](https://github.com/pytorch/pytorch/compare/v1.13.1...v2.2.0)

---
updated-dependencies:
- dependency-name: torch
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2024-09-04 12:28:09 +00:00
Mend RenovateandGitHub dbfabddd27 chore(deps): update dependency nbqa to v1.9.0 (#3490) 2024-09-04 12:11:13 +00:00
Vertex MG TeamandCopybara-Service f3d6c64b48 Add A100 40G as default deployment option for Flux serving notebook
PiperOrigin-RevId: 670756573
2024-09-03 17:09:13 -07:00
Vertex MG TeamandCopybara-Service 9020504954 Minor fix to the SD2.1-dreambooth notebook.
PiperOrigin-RevId: 670715655
2024-09-03 14:52:38 -07:00
Vertex MG TeamandCopybara-Service f96f75f8c1 Use standard id as MODEL_ID.
PiperOrigin-RevId: 670644580
2024-09-03 11:45:30 -07:00
Vertex MG TeamandCopybara-Service 9df704fb1c Update image task related notebooks
PiperOrigin-RevId: 670080374
2024-09-01 22:38:09 -07:00
Vertex MG TeamandCopybara-Service 390eea8216 Use standard id as MODEL_ID.
PiperOrigin-RevId: 669463640
2024-08-30 15:27:59 -07:00
Vertex MG TeamandCopybara-Service 1ca151a0ed Fix minor lint issues
PiperOrigin-RevId: 667824010
2024-08-26 20:52:52 -07:00
16a39c4d9e fix(egen): fixed the colab enterprise link (#3464)
* <Fix> Fixed the colab enterprise link.

* updated the prediction steps

---------

Co-authored-by: UBhavani <bhavani.ummadi@egen.ai>
2024-08-26 22:33:30 +00:00
Changyu ZhuandCopybara-Service b412207250 Add a Gradio notebook for chatting with instruction-tuned text generation models
PiperOrigin-RevId: 667678200
2024-08-26 12:41:01 -07:00
9c5edd3dda <Fix> Fixed the colab enterprise link. (#3465)
Co-authored-by: UBhavani <bhavani.ummadi@egen.ai>
2024-08-26 16:37:44 +00:00
Vertex MG TeamandCopybara-Service bcda7206d4 Add fill-mask notebook
PiperOrigin-RevId: 666923145
2024-08-23 14:25:34 -07:00
dd7d581af2 Update url to overview 2 (#3459)
* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

added colab enterprise logo and link

* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

hope I fixed the JSON issue

* fix: remove new changes

* Update vertex_ai_feature_store_based_llm_grounding_tutorial.ipynb

remove back ticks from product/feature names

* added command to install JDK

* Update fraud-detection-model.ipynb

update url to avoid redirect.

---------

Co-authored-by: Katie Nguyen <21978337+katiemn@users.noreply.github.com>
Co-authored-by: Ravi Dalal <ravidalal@google.com>
2024-08-23 21:22:13 +00:00
f10d009a66 Update url to overview 3 (#3460)
* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

added colab enterprise logo and link

* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

hope I fixed the JSON issue

* fix: remove new changes

* Update vertex_ai_feature_store_based_llm_grounding_tutorial.ipynb

remove back ticks from product/feature names

* added command to install JDK

* Update predictive_maintenance_usecase.ipynb

updated URL to avoid redirect

---------

Co-authored-by: Katie Nguyen <21978337+katiemn@users.noreply.github.com>
Co-authored-by: Ravi Dalal <ravidalal@google.com>
2024-08-23 21:21:30 +00:00
cb3b3ab26b Update url and other edits (#3461)
* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

added colab enterprise logo and link

* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

hope I fixed the JSON issue

* fix: remove new changes

* Update vertex_ai_feature_store_based_llm_grounding_tutorial.ipynb

remove back ticks from product/feature names

* added command to install JDK

* Update sdk-hyperparameter-tuning.ipynb

update url to avoid redirect. Made other edits, as well.

---------

Co-authored-by: Katie Nguyen <21978337+katiemn@users.noreply.github.com>
Co-authored-by: Ravi Dalal <ravidalal@google.com>
2024-08-23 21:20:29 +00:00
cf80c17db3 <Fix> Fixed the colab enterprise link. (#3468)
Co-authored-by: UBhavani <bhavani.ummadi@egen.ai>
2024-08-23 21:18:48 +00:00
622f39b59b <Fix> Fixed the colab enterprise link. (#3470)
Co-authored-by: UBhavani <bhavani.ummadi@egen.ai>
2024-08-23 21:18:05 +00:00
cefd548084 <Fix> Fixed the colab enterprise link. (#3471)
Co-authored-by: UBhavani <bhavani.ummadi@egen.ai>
2024-08-23 21:17:15 +00:00
40678a7bb4 Update url to overview (#3458)
* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

added colab enterprise logo and link

* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

hope I fixed the JSON issue

* fix: remove new changes

* Update vertex_ai_feature_store_based_llm_grounding_tutorial.ipynb

remove back ticks from product/feature names

* added command to install JDK

* Update training-multi-class-classification-model-for-ads-targeting-usecase.ipynb

update url to avoid redirect

---------

Co-authored-by: Katie Nguyen <21978337+katiemn@users.noreply.github.com>
Co-authored-by: Ravi Dalal <ravidalal@google.com>
2024-08-23 21:16:20 +00:00
6dbba5f51b <Fix> Fixed the colab enterprise link. (#3469)
Co-authored-by: UBhavani <bhavani.ummadi@egen.ai>
2024-08-23 21:14:18 +00:00
50ddff8ca9 <Fix> Fixed the colab enterprise link. (#3472)
Co-authored-by: UBhavani <bhavani.ummadi@egen.ai>
2024-08-23 21:13:29 +00:00
Vertex MG TeamandCopybara-Service 26ddd529ed Add Dynamic LoRA example to Stable Diffusion serving notebooks
PiperOrigin-RevId: 666913631
2024-08-23 13:58:08 -07:00
Vertex MG TeamandCopybara-Service 099b5c82bd Add Flux.1-schnell serving notebook
PiperOrigin-RevId: 666899568
2024-08-23 13:13:28 -07:00
Vertex MG TeamandCopybara-Service 0c1ff34e4c Update the finetuning notebook to use the pre-built training docker image.
PiperOrigin-RevId: 666853563
2024-08-23 10:59:04 -07:00
Vertex MG TeamandCopybara-Service 94ecca5bd8 Update the default machine type and accelerator
PiperOrigin-RevId: 666849592
2024-08-23 10:47:35 -07:00
Vertex MG TeamandCopybara-Service c97044ec62 Add chat completion to llama3.1 deployment notebook
PiperOrigin-RevId: 666846635
2024-08-23 10:39:31 -07:00
92f352dd83 feat:Add Starry Net tutorial (#3449)
Co-authored-by: Steve T <tsteve@google.com>
2024-08-22 21:58:43 +00:00
Vertex MG TeamandCopybara-Service 830c954b01 Update the default machine type and accelerator
PiperOrigin-RevId: 666415414
2024-08-22 10:59:29 -07:00
Vertex MG TeamandCopybara-Service c38fd04899 No public description
PiperOrigin-RevId: 666391609
2024-08-22 10:01:04 -07:00
Sujit KhasnisandGitHub 2d159448f1 fix: remove euw4 for jamba large (#3474) 2024-08-22 16:08:53 +00:00
Sujit KhasnisandGitHub 673689da92 feat: ai21 labs jamba mini and large models into ModelGarden (#3462)
* feat: ai21 labs jamba mini and large models

* feat: ai21 labs jamba updated codeowners

* fix: AI21 labs, added missing import

* fix: AI21 labs, suggestions verbiage

* fix: AI21 labs, suggestions verbiage

* fix: AI21 labs, suggestions verbiage #2
2024-08-22 15:38:31 +00:00
b98d0f7c29 Revert "Add chat completion to llama3.1 deployment notebook" (#3467)
This reverts commit 0a3ee2e642.

Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2024-08-22 14:11:47 +00:00
Vertex MG TeamandCopybara-Service 20a3c6a4b1 Minor fixes in paligemma deployment notebook
PiperOrigin-RevId: 666156172
2024-08-21 20:28:52 -07:00
Vertex MG TeamandCopybara-Service 0a3ee2e642 Add chat completion to llama3.1 deployment notebook
PiperOrigin-RevId: 666078263
2024-08-21 16:20:59 -07:00
Vertex MG TeamandCopybara-Service fd4bc08358 Make minor changes to pytorch prompt guard deployment notebook
PiperOrigin-RevId: 666055416
2024-08-21 15:19:30 -07:00
Vertex MG TeamandCopybara-Service 2a9f2c012b Add sample notebook for Hugging Face pytorch local inference
PiperOrigin-RevId: 665989183
2024-08-21 12:38:18 -07:00
Vertex MG TeamandCopybara-Service b50b61a2e8 Move deploy_model definition closer to the calling function
PiperOrigin-RevId: 665639076
2024-08-20 19:36:49 -07:00
Vertex MG TeamandCopybara-Service d60140fcf9 Update Mistral-7B PEFT notebook to use A100 80GB for finetuning and L4 for serving.
PiperOrigin-RevId: 665632534
2024-08-20 19:15:15 -07:00
Vertex MG TeamandCopybara-Service 5a4ddc2926 Add Prompt Guard deployment notebook.
PiperOrigin-RevId: 665489940
2024-08-20 12:58:36 -07:00
Kaushik KoiladaandGitHub 41c9354efa chore, refactor (egen): edits get_started_with_pytorch_rov notebook (#3376)
* chore, refactor: edits made according to the template

* chore: lint run

* fix: remove ray version

* fix: made dataset url to http to deal with job failure error

* chore: lint run

* chore: fixes markdown as per guide

* chore: lint

* chore, fix: adds testing code and also fix the error

* chore, fix: clear outputs adds retries and adds http dataset path in testing

* chore:  review comment addressed

* chore: lint run

* refactor: removes IS_TESTING flag

* chore, fix: Removes IS_TESTING, fixes packages installations, runs end to end

* chore: lint run
2024-08-20 17:24:37 +00:00
83c90b11dd Sdk automl image object detection batch online (#3325)
* chore,refactor(egen): hardcoded the version of tensorflow, added comments in the cleanup code, modified cleanup code, changed region variable name to location, replaced uuid with _unique and removed the code of uuid generation, refactored code according to template guidelines and performed linter test.

* fix(egen): added code for copying from one bucket to other for importing datset

* chore(egen): modified the markdowns of copying data between google cloud storage buckets step and performed linter test.

* chore(egen): Added gcsfs package in the installation step and performed linter test.

* fix(egen): defined the model display name variable and performed the linter test.

* chore(egen): Done changes according to @kittyabs review and performed linter test.

---------

Co-authored-by: sriramya2610 <sriramya.peddapally@egen.ai>
2024-08-20 17:22:11 +00:00
7af5359fa2 add codeowners for gemma2 finetuning (#3445)
Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2024-08-20 13:48:22 +00:00
df32205187 Publish gemma2 finetuning notebook (#3444)
PiperOrigin-RevId: 665333655

Co-authored-by: Vertex MG Team <vertex-mg-bot@google.com>
2024-08-20 13:46:40 +00:00
fcad5ddbc2 Add minor improvements to the notebooks (#3436)
PiperOrigin-RevId: 664637239

Co-authored-by: Vertex MG Team <vertex-mg-bot@google.com>
2024-08-20 13:45:51 +00:00
Vertex MG TeamandCopybara-Service 205fb8530f Updated max_model_len to 128000 for mistral nemo when gpu_memory_utilization is 0.9
PiperOrigin-RevId: 665330240
2024-08-20 06:40:53 -07:00
Vertex MG TeamandCopybara-Service e1705a7a82 No public description
PiperOrigin-RevId: 665330199
2024-08-20 06:36:18 -07:00
Vertex MG TeamandCopybara-Service a5df8e7b82 Support A100 and add quota check to stable diffusion notebooks
PiperOrigin-RevId: 665014758
2024-08-19 15:41:59 -07:00
Vertex MG TeamandCopybara-Service 73e48a95ae Migrate stable diffusion text-to-image notebooks to VMG pytorch inference docker
PiperOrigin-RevId: 664990787
2024-08-19 14:50:34 -07:00
Vertex MG TeamandCopybara-Service 7cd14000b9 No public description
PiperOrigin-RevId: 664949792
2024-08-19 13:33:15 -07:00
skarukasandGitHub 2eddf4f83e Add parameter descriptions for text embedding tuning Colab (#3416)
* Add learning_rate_multiplier and output_dimensionality to embedding tuning notebook.

* Reformat

* Add parameter descriptions to text embedding tuning sample Colab.

* Run formatter
2024-08-19 12:52:06 +00:00
Dustin LuongandCopybara-Service 8e7d030426 No public description
PiperOrigin-RevId: 663483238
2024-08-16 18:31:07 -07:00
kewentandGitHub e259870898 feat: add text-embedding-preview-0815 model to notebook (#3371)
* feat: add text-embedding-preview-0815 model to notebook

* fix lint error
2024-08-17 01:04:42 +00:00
Kaushik KoiladaandGitHub 29881a0b80 chore: deletes notebook (#3430) 2024-08-16 20:19:48 +00:00
Vertex MG TeamandCopybara-Service 0aa09a8524 No public description
PiperOrigin-RevId: 663385165
2024-08-16 11:21:34 -07:00
c362269e75 feat: add support for mistral nemo (#3431)
Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2024-08-16 13:57:16 +00:00
1b94ad8e59 Spark on rov (#3429)
* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

added colab enterprise logo and link

* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

hope I fixed the JSON issue

* fix: remove new changes

* Update vertex_ai_feature_store_based_llm_grounding_tutorial.ipynb

remove back ticks from product/feature names

* Update spark_on_ray_on_vertex_ai.ipynb

Edited link text and added another link to relevant documentation

* added command to install JDK

---------

Co-authored-by: Katie Nguyen <21978337+katiemn@users.noreply.github.com>
Co-authored-by: Ravi Dalal <ravidalal@google.com>
2024-08-16 00:34:10 +00:00
Kaushik KoiladaandGitHub 04cc933895 chore, refactor(egen): edits sdk_automl_forecasting_hierarchical_batch notebook (#3327)
* chore, refactor:  refactors and edits according to the template.

* chore: lint

* chore: update objective and other verbiage

* chore: edits title of the notebook
2024-08-14 23:59:28 +00:00
34ffeaab3a refactor(egen): template fixes, updates clean up steps (#3423)
* <Refactor> Refactored the notebook according to the template.

* Applied suggested edits.

---------

Co-authored-by: UBhavani <bhavani.ummadi@egen.ai>
2024-08-14 23:51:58 +00:00
85660584dd Update tensorboard objective 2 (#3428)
* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

added colab enterprise logo and link

* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

hope I fixed the JSON issue

* fix: remove new changes

* Update vertex_ai_feature_store_based_llm_grounding_tutorial.ipynb

remove back ticks from product/feature names

* Update tensorboard_profiler_custom_training.ipynb

Added reference to "Vertex AI TensorBoard" so the notebook will show up in the Notebook Tutorials page when filtering for Vertex AI TensorBoard.

---------

Co-authored-by: Katie Nguyen <21978337+katiemn@users.noreply.github.com>
2024-08-14 23:47:42 +00:00
dcc5ebabb3 Update tensorboard objective 1 (#3427)
* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

added colab enterprise logo and link

* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

hope I fixed the JSON issue

* fix: remove new changes

* Update vertex_ai_feature_store_based_llm_grounding_tutorial.ipynb

remove back ticks from product/feature names

* Update tensorboard_profiler_custom_training_with_prebuilt_container.ipynb

Added reference to "Vertex AI TensorBoard" so the notebook will show up in the Notebook Tutorials page when filtering for Vertex AI TensorBoard.

---------

Co-authored-by: Katie Nguyen <21978337+katiemn@users.noreply.github.com>
2024-08-14 23:46:52 +00:00
Kaushik KoiladaandGitHub 257003e709 chore: deletes notebook as it is no longer valid (#3425) 2024-08-14 17:53:18 +00:00
Kaushik KoiladaandGitHub b23538d19d chore: deletes notebook as it is no longer valid (#3426) 2024-08-14 17:52:21 +00:00
Katie NguyenandGitHub ba935bf993 fix: title modifications (#3424) 2024-08-14 07:21:46 +00:00
weiran-workandGitHub a9e59e9626 feat: Add quota check for restricted image (#3421) 2024-08-13 21:32:31 +00:00
cdca4a3c2e convert caption prompt to bool (#3422)
Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2024-08-13 21:13:56 +00:00
kittyabsandGitHub 78276da7df Update tensorboard_custom_training_with_prebuilt_container.ipynb (#3420)
Fixed typos and cleaned up product/service names
2024-08-13 21:11:50 +00:00
sen-samandGitHub dbd36f76c1 Update vertex_ai_feature_store_feature_view_service_agents.ipynb (#3419)
Made a minor fix to the agenda to make sure it's populated correctly in the list of notebooks in the public docs.
2024-08-13 21:10:39 +00:00
Kaushik KoiladaandGitHub a9332fdb51 chore, refactor(egen): refactors sdk_automl_image_object_detection_batch.ipynb notebook (#3281)
* chore, refactor: removes boiler plate, adds colab enterprise, changes region to location

* refactor: adds testing code and variables

* chore: adds test code and runs end to end

* chore: end to end test and remove testing code

* chore: lint test

* chore: rectify parameter explanation

* fix: adds code and relevant markdown to copy dataset to the project's bucket for dealing with access issue

* chore: removes code font for Dataset

* chore: lint

* chore: address review comments

* adds gcfs to the installation
2024-08-13 21:06:49 +00:00
85522dfe48 fix, chore, feat, refactor(Egen): Replace K80 with T4, add cleanup steps, refactors (#3272)
* fix, refactor, chore: follows new template, replaces K80 with T4, replace docker steps with cloud build, reorganize the sections, heading corrections

* chore: remove will and contracts 'do not'

* fix: replace K80 with T4

* feat: adds step to remove the training folder in the cleaning up section

* fix, chore: adds worker-pool-specs back in the pipeline as a global var, adds comments

* feat: adds a cleaning up step for artifact registry

* fix, chore: addresses the review comments, removes the mention of experimental feature in the markdown, replaces TPU_V3 with TPU_V2(not found error)

* chore: addresses review comments

* fix, chore, refactor: updates tensorflow version to 2.13, updates TPU driver libs, elaborates some steps, refactors the pipeline creation and run step to parameterize the arguments instead of using global vars

---------

Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-08-13 17:57:56 +00:00
4c7180bd10 refactor, chore(egen): Replaces K80 GPU with T4, package version updates, pre-built Docker container image for prediction update, other corrections from template (#3241)
* <refactor, chore> updates gcr to artifact registry, package version upgrades, updates prebuilt docker container image to 2.13, refactores notebook according to the template

* <refactor, chore> updates gcr to artifact registry, package version upgrades, updates prebuilt docker container image to 2.13, refactores notebook according to the template

* updated pip install statements

* Colab enterprise link fix

* Colab enterprise link fix

* Colab enterprise link fix

* markdown edits

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-08-13 17:44:07 +00:00
5659c96f0e chore, refactor (egen): refactors and edits model_monitoring.ipynb notebook (#3195)
* chore: remove boiler plate, add colab enterprise, and format according to the template

* chore, refactor: test end to end

* chore: lint test

* chore, fix: removes force protobuf for package compatibility issue and removes testing variable reference

* fix: changes import statement for execution

* adds protobuf in install to deal with build error

* fix: changes protobuf installation version

* chore: lint run

* Adds tensorflow in install, other fixes from template

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-08-13 17:40:19 +00:00
Aaron DietzandGitHub 26a25d0cf7 Update google_cloud_pipeline_components_automl_tabular.ipynb (#3418)
Restructured bullet points to remove nested bullets. The nested bullets don't render properly when we auto-generate our notebook list for our docs.
2024-08-12 21:26:15 +00:00
cef9c0659a chore: Adds a deprecation note (#3353)
Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-08-12 21:21:55 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
d14e71b0ee chore(deps): bump tensorflow (#3405)
Bumps [tensorflow](https://github.com/tensorflow/tensorflow) from 2.7.2 to 2.12.1.
- [Release notes](https://github.com/tensorflow/tensorflow/releases)
- [Changelog](https://github.com/tensorflow/tensorflow/blob/master/RELEASE.md)
- [Commits](https://github.com/tensorflow/tensorflow/compare/v2.7.2...v2.12.1)

---
updated-dependencies:
- dependency-name: tensorflow
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2024-08-12 14:42:00 +00:00
c3998c9898 chore, refactor(egen): New template edits, clean up code refactor, remove future tense (#3403)
* Adapted notebook with new notebook template

* Added variable value which is used in further steps

* Added required permissions for service account

* chore, refactor: Follows new template, removes future tense, restricted links, refactors the cleaning up section

* feat: adds '-m' while deleting the cloud storage bucket

* chore: addresses the review comments

* chore: removes version mention for bison models and adds references to model versions and supported rlhf models

---------

Co-authored-by: nileshspringml <nilesh.mahajan@egen.ai>
Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-08-12 14:37:52 +00:00
1ab9e4b6e3 Add common fn related to image task (#3417)
Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2024-08-12 14:33:12 +00:00
8a2b41c5f1 fix, refactor, chore: Fix docker steps, Colab steps, follows new template etc. (#3385)
* chore: adds Colab Enterprise link

* fix, refactor, chore: Fixes the docker container image creation steps and predictions step, refactors the code to skip unnecessary steps and markdown text correction and simplification

* fix, chore: fixes typos in serving script, markdown text corrections, adds custom folder removal step in the cleanup section

* fix: adds --project for Colab steps and replaces the docker build and docker push with gcloud builds submit command for Colab

---------

Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-08-09 21:39:47 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
c3966ae8f3 chore(deps): bump tensorflow (#3381)
Bumps [tensorflow](https://github.com/tensorflow/tensorflow) from 2.7.2 to 2.12.1.
- [Release notes](https://github.com/tensorflow/tensorflow/releases)
- [Changelog](https://github.com/tensorflow/tensorflow/blob/master/RELEASE.md)
- [Commits](https://github.com/tensorflow/tensorflow/compare/v2.7.2...v2.12.1)

---
updated-dependencies:
- dependency-name: tensorflow
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2024-08-09 21:38:30 +00:00
Ayush AgrawalandGitHub 7b4d745853 feat: Add Slack and Jira Source file uploads and 3P embedding model config example for corpus creation (#3404)
* feat: Add Slack and Jira source file imports and Endpoint resource embedding model configuration for corpus creation

* docs: Link reference to deploying 3P models to endpoint

* fix: update linting
2024-08-09 21:29:36 +00:00
Aaron DietzandGitHub 216931b90d Update sdk_automl_tabular_binary_classification_batch_explain.ipynb (#3409)
Converted two bullet points into paragraphs because the bullets messed with how the key tasks were imported into our docs.
2024-08-09 21:28:23 +00:00
Aaron DietzandGitHub de16c7623d Minor update to text in sdk_automl_video_classification_batch.ipynb (#3410)
Converted two bullets into paragraphs because the bullets messed with how the key tasks were rendered in our docs
2024-08-09 21:27:45 +00:00
Aaron DietzandGitHub 569db62fc7 Minor update to text in sdk_automl_video_action_recognition_batch.ipynb (#3411)
Converted two bullets into paragraphs because the bullets messed with how the key tasks were rendered in our docs
2024-08-09 20:55:09 +00:00
Aaron DietzandGitHub fd8a0a2d0c Minor update to text in automl_image_classification_batch_prediction.ipynb (#3412)
Converted two bullets into paragraphs because the bullets messed with how the key tasks were rendered in our docs
2024-08-09 20:54:27 +00:00
f5366faca7 Refactor paligemma and codegemma notebooks (#3408)
* refactor: refactor paligemma and codegemma notebooks

* fix lint issues

* fix lint issues

* fix lint issues

---------

Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2024-08-09 20:53:42 +00:00
Ayush AgrawalandGitHub 170205fe08 Update model_garden_rag.ipynb notebook metadata (#3406) 2024-08-09 20:52:52 +00:00
Aaron DietzandGitHub 3ece113c7a Minor update to text in automl_tabular_classification_beans.ipynb (#3413)
Updates objectives to eliminate nested bullets (those didn't get rendered in our docs) and removes some bolding that wasn't rendering properly in our docs
2024-08-09 20:52:07 +00:00
Aaron DietzandGitHub acee2878b0 Minor update to notebook_template_review.py (#3414)
Moved a couple name/variable swaps up before "Vertex AI" so that it'll catch them.
2024-08-09 20:51:14 +00:00
Kathy YuandGitHub bb0ded3c46 Update Llama 2 evaluation notebook. (#3415) 2024-08-09 20:50:47 +00:00
nathreya-googleandGitHub 4921cbce9b Add chat completions sample for llama3 deployment. (#3407) 2024-08-08 21:35:45 +00:00
Mend RenovateandGitHub dbcdef9588 chore(deps): update dependency nbqa to v1.8.7 (#3400) 2024-08-07 09:42:49 +00:00
cc53c8e9d9 refactor: refactor mistral and mixtral deployment notebooks (#3401)
Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2024-08-07 09:29:16 +00:00
7590168238 feat: add resize image function to common util (#3399)
Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2024-08-06 17:37:14 +00:00
5cac0e6d0b refactor,chore(egen): added colab enterprise link (#3398)
Co-authored-by: Jayakrishna2801 <jayakrishna.rajaboina@egen.ai>
2024-08-06 17:35:11 +00:00
c3d25502bf refactor,chore(egen): added colab enterprise link (#3394)
Co-authored-by: Jayakrishna2801 <jayakrishna.rajaboina@egen.ai>
2024-08-06 05:40:54 +00:00
Kaushik KoiladaandGitHub f34deeccc5 chore, refactor(egen): tpuv5e_gemma_peft_finetuning_and_serving notebook fix (#3159)
* chore: remove boilerplate and adds colab enterprise, and changes region to location

* chore: lint run

* chore: comments addressed and lint run
2024-08-06 05:34:49 +00:00
Kaushik KoiladaandGitHub d434b5ddbc chore,refactor(egen): formats and edits notebooks/official/training/tpuv5e_llama2_pytorch_finetuning_and_serving.ipynb notebook (#3144)
* chore: reformats the copyright and run buttons. adds colab enterprise

* chore: reorders and formats markdown sections and removes boilerplate

* chore: remove boilerplate and reformats based on template

* chore,refactor: adds colab enterprise and refactors the cells according to template

* fix: rectifies issue with testing
2024-08-06 05:33:58 +00:00
42a8600aa7 chore,refactor(egen): minor changes to the automl_tabular_on_vertex_pipelines notebook (#3298)
* chore,refactor(egen): Changed REGION variable name to LOCATION, defined two variables to get the pipelines names to be used in cleanup section, added cleanup code for deletion of pipelines and models, replaced uuid with unique,removed os.getenv(IS_TESTING) from the cleanup section, refactored code according to the template guidelines and performed linter test.

* fix(egen): changed the version of google-cloud-pipeline-components and performed linter test.

* chore(egen): changed the version of google cloud pipeline components package in installation step and performed linter test.

* fix(egen): changed the model-evaluation parameter to model-evaluation-2 in get_evaluation_metrics function and performed linter test.

* fix(egen): modified the code in get_feature_attributions helper function and performed linter test.

* fix(egen): modified code in cleanup section to delete pipeline jobs and performed linter test.

* fix(egen): renamed model-upload-2 to model-upload in cleanup code of automl tabular architecture pipeline and performed linter test.

* chore(egen): changed the colab enterprise link by renaming automml to automl in link and performed linter test.

---------

Co-authored-by: sriramya2610 <sriramya.peddapally@egen.ai>
2024-08-06 04:21:51 +00:00
Kaushik KoiladaandGitHub 5ec392769a chore, refactor(egen): refactors the SDK_Custom_Training_Python_Package_Managed_Text_Dataset_Tensorflow_Serving_Container notebook (#3224)
* chore: adds colab enterprise, removes boilerplate and edits according to template

* chore: run end to end notebook

* chore, refactor: formats, runs end to end

* chore: lint

* chore: addresses comments and changes headers according to guidelines

* chore:review comments addressed

* chore: addresses review comments

* chore: lint run pass
2024-08-06 04:18:11 +00:00
a008c7b4f9 refactor, chore(egen): Removes boilerplate, heading fixes, and other corrections from template (#3092)
* refactor, chore(egen): Removes boilerplate, heading fixes, and other corrections from template

* Updated minor template related issue

* Change variable name from REGION to LOCATION

* removed unated variable

* Made variable values more redable

* Removed unwated commentes from header

* fix, chore: replaces test sample file with eval sample file, adds comments to the cleaning up section

* fix: replaces text-bison@001 with tex-bison@002, reverts the post-tuning data sample to test sample

---------

Co-authored-by: Krishna Chaithanya Movva <krishna.movva@springml.com>
Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-08-06 03:52:26 +00:00
e06b958cd8 refactor,chore(egen) : refactored code according to template guidelines (#3285)
* refactor,chore(egen) : refactored code according to template guidelines

* refactor,chore(egen) : added code to delete locally generated files and refactored as per template guidelines

* refactor,chore(egen) : refactored according to template guidelines

* refcator(egen): refactored notebook according to template guidelines

* refactor,chore(egen): fix %%bigquery command usage

* refactor,chore(egen): removed hardcoded value

---------

Co-authored-by: Jayakrishna2801 <jayakrishna.rajaboina@egen.ai>
2024-08-06 03:47:17 +00:00
811d321451 refactor,chore(egen): added colab enterprise for custom_tabular_train_batch_pred_bq_pipeline notebook (#3387)
* refactor,chore(egen): added colab enterprise link for custom_tabular_train_batch_pred_bq_pipeline notebook

* refactor,chore(egen): added colab enterprise link for custom_tabular_train_batch_pred_bq_pipeline notebook

---------

Co-authored-by: Jayakrishna2801 <jayakrishna.rajaboina@egen.ai>
2024-08-06 03:38:12 +00:00
6aff88a23f refactor,chore(egen): added colab enterprise link (#3390)
Co-authored-by: Jayakrishna2801 <jayakrishna.rajaboina@egen.ai>
2024-08-06 03:36:50 +00:00
1b4f1ab9c5 refactor,chore(egen):added colab enterprise link (#3391)
Co-authored-by: Jayakrishna2801 <jayakrishna.rajaboina@egen.ai>
2024-08-06 03:36:07 +00:00
8afb0dd4b3 refactor,chore(egen): added colab enterprise link (#3393)
Co-authored-by: Jayakrishna2801 <jayakrishna.rajaboina@egen.ai>
2024-08-06 03:35:29 +00:00
94f3c08adb refactor,chore(egen): added colab enterprise link (#3392)
Co-authored-by: Jayakrishna2801 <jayakrishna.rajaboina@egen.ai>
2024-08-06 03:34:33 +00:00
sumanvitaandGitHub 598d9b0207 typo fix in markdown (#3383) 2024-08-06 00:13:07 +00:00
ac2d168c96 refactor, chore(egen): updates dataset import process, adds gcfs library, other corrections from template (#3354)
* <refactor, chore> data is copied to project's own bucket for importing into the dataset, added gcfs library, other corrections from template

* <refactor, chore> data is copied to project's own bucket for importing into the dataset, added gcfs library, other corrections from template

* markdown edits

* markdown edits

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-08-06 00:05:48 +00:00
6bba78f9e2 refactor,chore(egen): added colab enterprise link for automl_video_classification_model_evaluation notebook (#3384)
Co-authored-by: Jayakrishna2801 <jayakrishna.rajaboina@egen.ai>
2024-08-06 00:02:23 +00:00
eb8a0a93a2 refactor,chore(egen): added colab enterprise for lightweight_functions_component_io_kfp notebook (#3386)
Co-authored-by: Jayakrishna2801 <jayakrishna.rajaboina@egen.ai>
2024-08-06 00:01:16 +00:00
2921efcf64 refactor,chore(egen): added colab enterprise link (#3388)
Co-authored-by: Jayakrishna2801 <jayakrishna.rajaboina@egen.ai>
2024-08-05 23:59:33 +00:00
8b9afc3fb4 refactor,chore(egen): added colab enterprise link (#3389)
Co-authored-by: Jayakrishna2801 <jayakrishna.rajaboina@egen.ai>
2024-08-05 23:58:43 +00:00
7d2b5ac7a4 refactor,chore(egen): added colab enterprise link (#3395)
Co-authored-by: Jayakrishna2801 <jayakrishna.rajaboina@egen.ai>
2024-08-05 23:58:04 +00:00
db0758197f refactor,chore(egen): added colab enterprise link (#3397)
Co-authored-by: Jayakrishna2801 <jayakrishna.rajaboina@egen.ai>
2024-08-05 23:57:25 +00:00
69518e8834 chore(egen): modified colab link, colab enterprise link, workbench link and github link in markdown and perfomed linter test. (#3378)
Co-authored-by: sriramya2610 <sriramya.peddapally@egen.ai>
2024-08-05 23:44:25 +00:00
c36e5f204d Update new job name creation format (#3377)
Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2024-08-02 20:55:38 +00:00
Mend RenovateandGitHub 0be74ddfa5 chore(deps): update dependency black to v24.8.0 (#3379) 2024-08-02 20:55:04 +00:00
sefgsefgandGitHub efd93073d0 Linear regression predictor using sklearn (#3357)
* Upload torch transformers predictor sample

This sample uses the aiplatform SDK and torch library to implement transformers predictor.

* Linear regression predictor using sklearn

Upload a linear regression predictor sample using scikit-learn lib.
2024-08-01 13:50:05 +00:00
cc2011e354 Make mistral notebook tunable (#3374)
* Make mistral notebook tunable

* make minor improvements

* make minor improvements

---------

Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2024-08-01 13:47:50 +00:00
Huguens JeanandGitHub 772881903b [MG Model Team] Add environment variable for downscaling video frames prior to prediction. (#3370) 2024-08-01 13:45:27 +00:00
Xiang XuandGitHub 04f86e647d Fix machine_type in model_garden_pytorch_llama3_1_deployment (#3373) 2024-08-01 13:44:36 +00:00
de120ccd62 refactor(egen): Colab enterprise link fix (#3368)
* <refactor, chore> Updated prebuilt container image for prediction to 1.3, scikit-learn package updated to 2.5.1, other corrections from template

* <refactor> refactored notebok according to notebook template

* <refactor> refactored notebok according to notebook template

* <refactor> refactored notebok according to notebook template

* <refactor> refactored notebok according to notebook template

* <refactor> refactored notebok according to notebook template

* Colab enterprise link fix

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-07-31 23:39:32 +00:00
0e5c1cebdf Update url to /generative ai/docs (#3319)
* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

added colab enterprise logo and link

* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

hope I fixed the JSON issue

* fix: remove new changes

* Update vertex_ai_feature_store_based_llm_grounding_tutorial.ipynb

remove back ticks from product/feature names

* Update distillation.ipynb

update url to point to /vertex-ai/generative-ai/docs

---------

Co-authored-by: Katie Nguyen <21978337+katiemn@users.noreply.github.com>
2024-07-31 23:36:29 +00:00
cb0caee3ea Update gemma2 notebook to include 2b instructions. (#3369)
Co-authored-by: Pooya Moradi <pooyam@google.com>
2024-07-31 15:21:38 +00:00
Yashika GandhiandGitHub 8eedc45652 Feat: Adding Phi-3 deployment notebook (#3364)
* Feat: Adding Phi 3 deployment notebook

* Adding Phi 3 deployment notebook

* Adding Phi 3 deployment notebook

* lint
2024-07-31 15:05:12 +00:00
255b520c01 feat: add qwen2 deployment notebook (#3345)
* feat: add qwen2 deployment notebook

* fix lint issues

* Qwen2 additional minor improvements

* Fix lint issues

---------

Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2024-07-31 11:41:44 +00:00
sumanvitaandGitHub 5b079af33d fix, refactor, chore(egen): adds matplotlib library in installation step, adds code to delete locally generated files, refactors code as per new template guidelines (#3271)
* fix, refactor, chore(egen): adds matplotlib library in installation step, adds code to delete locally generated files, refactors code as per new template guidelines, contraction of words, performs linter test

* fix, refactor, chore(egen): adds matplotlib library in installation step, adds code to delete locally generated files, refactors code as per new template guidelines, contraction of words, performs linter test

* Adds deprecation note for user managed instances
2024-07-31 05:50:52 +00:00
bd89533c50 chore, feat, refactor(Egen): Clean up steps, template update etc. (#3341)
* chore, feat, refactor: Remove future tense, follows new notebook template, adds cleanup step for deleting the pipelines and models created, refactors the utility functions to fetch model resource

* fix: corrects the notebook name in the links

* chore: addresses review comments

* fix: loads the model from resource name before deletion

---------

Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-07-31 05:47:58 +00:00
Kaushik KoiladaandGitHub 6555569156 chore,refactor(egen): reformats get_started_with_custom_training_autologging_local_script.ipynb notebook (#3179)
* chore: lint edits

* fix: retrieval of custom job removed as there is error in the method being called and also existing job object exists
2024-07-31 05:41:54 +00:00
Kaushik KoiladaandGitHub a6aa61ae1d chore, refactor(egen):edits template of get_started_with_model_monitoring_automl_image_batch notebook (#3269)
* chore: adds the colab enterprise and formates all open in tabs

* chore: changes 'Run in colab' to 'Open in Colab'

* chore: splits content of first cell for formatting

* refactor: fixes pip  install and rectifies disable_early_stopping  parameter description

* fix: image dataset csv is changed because of permission issue when importing from the content of the csv

* chore: running end to end notebook

* chore: removes testing variables and lints

* chore: markdown edits and end to end test

* chore: lint changes

* refactor: adds try except block following all the other cell codes
2024-07-31 05:37:07 +00:00
Kaushik KoiladaandGitHub 1580c63617 chore, refactor(egen): refactors automl_image_object_detection_export_edge.ipynb notebook (#3297)
* chore, refactor: removes boilerplate, adds colab enterprise and makes changes according to template and authoring guide

* chore: lint

* fix: adds dataset copying code to fix data access issue

* chore: lint

* fix: adds gcsfs package to deal with the check error

* chore: clears outputs
2024-07-31 05:30:30 +00:00
e05e777d1d refactor, chore(egen): xgboost package version set to 1.7.1, updates serving container image to 1.7, deletes intermediate files, other fixes from template (#3306)
* <refactor, chore> xgboost package version set to 1.7.1, updates serving continer image to 1.7, deletes intermediate files, other fixes from template

* <refactor, chore> xgboost package version set to 1.7.1, updates serving continer image to 1.7, deletes intermediate files, other fixes from template

* <refactor, chore> xgboost package version set to 1.7.1, updates serving continer image to 1.7, deletes intermediate files, other fixes from template

* Colab logo fix

* Colab logo fix

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-07-31 05:27:40 +00:00
91efb27aea refactor,chore(egen): refactored code as per template guidelines (#3344)
* refactor,chore(egen): refactored code as per template guidelines

* refactor,chore(egen): fix region variable

* refactor(egen): refactored notebook as per template guielines

* refactor,chore(egen): performed linter test

---------

Co-authored-by: Jayakrishna2801 <jayakrishna.rajaboina@egen.ai>
2024-07-31 05:24:15 +00:00
b08b7a21f1 Fix url to gen ai (#3318)
* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

added colab enterprise logo and link

* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

hope I fixed the JSON issue

* fix: remove new changes

* Update vertex_ai_feature_store_based_llm_grounding_tutorial.ipynb

remove back ticks from product/feature names

* Update tune_peft.ipynb

 “fix: update links in notebook”

* Update tune_peft.ipynb

add colab enterprise link

* Update tune_peft.ipynb

removed link

* Update tune_peft.ipynb

updated the link to /generative-ai/docs/tune_peft.ipynb and other edits

* fix: linter errors

---------

Co-authored-by: Katie Nguyen <21978337+katiemn@users.noreply.github.com>
2024-07-31 05:18:50 +00:00
3b2adb2e1d refactor(egen): Colab enterprise link fix (#3356)
* <refactor> refactored notebook according to template

* Apply edits suggested by @kittyabs

* Apply suggested edits from @kittyabs review

* Colab enterprise link fix

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-07-31 05:16:41 +00:00
472fe95430 refactor(egen): Colab enterprise link fix (#3355)
* <refactor, chore> refactored notebook according to new template

* Applied suggested edits

* Colab enterprise link fix

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-07-31 05:15:59 +00:00
cb8ec73e0a refactor(egen): Colab enterprise link fix (#3358)
* <fix, refactor> fixed and refactored notebook according to the template

* Apply suggested edits from @kittyabs

* Apply suggested edits from @kittyabs

* Colab enterprise link fix

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-07-31 05:15:22 +00:00
510e3d08ce refactor(egen): Colab enterprise link fix (#3359)
* <refactor, chore, fix> package version updates, pipeline components documentation link update, importer_node import fix, machineSpec update

* <refactor, chore, fix> package version updates, pipeline components documentation link update, importer_node import fix, machineSpec update

* Colab enterprise link fix

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-07-31 05:14:41 +00:00
9e3edfc8e4 refactor(egen): Colab enterprise link fix (#3360)
* <refactor, chore> replaces K80 GPU with T4, updates pre-built Docker container image for training and prediction to 2.13, cleansup intermediate files, updates kfp and tensorflow versions, fixes minor spelling mistakes and contracts words, removes future tense

* <refactor, chore> replaces K80 GPU with T4, updates pre-built Docker container image for training and prediction to 2.13, cleansup intermediate files, updates kfp and tensorflow versions, fixes minor spelling mistakes and contracts words, removes future tense

* <refactor, chore> replaces K80 GPU with T4, updates pre-built Docker container image for training and prediction to 2.13, cleansup intermediate files, updates kfp and tensorflow versions, fixes minor spelling mistakes and contracts words, removes future tense

* Grammar fix

* lint fix

* Colab enterprise link fix

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-07-31 05:13:43 +00:00
72c3ad3e26 refactor(egen): Colab enterprise link fix (#3361)
* <refactore, chore>refactored notebook according to the template

* refactor: Apply markdown text edit

* source distribution fix

* source distribution fix

* Colab enterprise link fix

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-07-31 05:12:55 +00:00
7626cd7025 refactor(egen): Colab enterprise link fix (#3362)
* <refator, chore> Adds Tensorflow in installation section, corrections from template

* <refator, chore> Adds Tensorflow in installation section, corrections from template

* Colab enterprise link fix

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-07-31 05:12:15 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
c1f54a2de9 chore(deps): bump torch (#3333)
Bumps [torch](https://github.com/pytorch/pytorch) from 2.0.1+cu118 to 2.2.0.
- [Release notes](https://github.com/pytorch/pytorch/releases)
- [Changelog](https://github.com/pytorch/pytorch/blob/main/RELEASE.md)
- [Commits](https://github.com/pytorch/pytorch/commits/v2.2.0)

---
updated-dependencies:
- dependency-name: torch
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2024-07-31 02:00:58 +00:00
Gary WeiandGitHub 795f182de3 Remove the text-to-video notebook as we decide to hide the model card from the UI. (#3351) 2024-07-31 02:00:26 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
86670246d9 chore(deps): bump tensorflow (#3365)
Bumps [tensorflow](https://github.com/tensorflow/tensorflow) from 2.7.2 to 2.12.1.
- [Release notes](https://github.com/tensorflow/tensorflow/releases)
- [Changelog](https://github.com/tensorflow/tensorflow/blob/master/RELEASE.md)
- [Commits](https://github.com/tensorflow/tensorflow/compare/v2.7.2...v2.12.1)

---
updated-dependencies:
- dependency-name: tensorflow
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2024-07-31 01:59:34 +00:00
c5443d656c Support checking H100 quota (#3367)
Co-authored-by: minwoopark <minwoopark@google.com>
2024-07-31 01:59:09 +00:00
Sujit KhasnisandGitHub ba526e4ac4 feat: Mistral AI colab notebook; new examples for Code Gen (FIM), Function calling (#3363)
* feat: Mistral AI colab notebook; new examples for Code Gen (FIM), Function calling

* feat: Mistral AI colab notebook; new examples for Code Gen (FIM), Function calling, Value error fix

* feat: Mistral AI colab notebook; Additional examples minor updates / fixes

* feat: Mistral AI colab notebook;

* feat: Mistral AI colab notebook; Chat completion*s* fix
2024-07-30 20:39:39 +00:00
sumanvitaandGitHub e6360bb1e6 fix,chore,refactor(egen): modifies the value of IMPORT_FILE, hardcodes TF version to 2.15.1, removes import os and os.getenv(IS_TESTING), refactors code as per the template guidelines. (#3326) 2024-07-29 22:21:03 +00:00
e61b249410 refactor, chore(egen): template fixes, update and add new package (#3305)
* <Refactor, Chore> Refactored the notebook according to the template, updated and added new package.

* numpy version conflict

* Added numpy.

---------

Co-authored-by: UBhavani <bhavani.ummadi@egen.ai>
2024-07-29 22:16:52 +00:00
00f722def7 fix,chore,refactor(egen): minor changes to the automl training notebook. (#3304)
* fix,chore,refactor(egen): Changed REGION variable name to LOCATION, modified the import file of flowers dataset, added comments in cleanup section, removed os.getenv(IS_TESTING) from the cleanup section, refactored code according to template guidelines and performed linter test.

* chore(egen): Done changes according to @kittylabs and performed linter test.

---------

Co-authored-by: sriramya2610 <sriramya.peddapally@egen.ai>
2024-07-29 22:13:46 +00:00
Kaushik KoiladaandGitHub 68269b07b0 refactor, chore (egen): refactors sdk_automl_video_object_tracking_batch.ipynb notebook (#3296)
* chore, refactor: adds colab enterprise, removes boiler plate, refactors according to template

* chore: run end to end and lint

* chore: addresses reviw comments
2024-07-29 22:04:59 +00:00
e83ca6e8e9 refactor(egen): Colab enterprise link fix (#3337)
* <refactor> refactored notebook according to template

* <refactor> refactored notebook according to template

* <refactor> refactored notebook according to template

* <refactor>: refactored notebook according to new notebook template

* <refactor>: refactored notebook according to new notebook template

* <refactor>: refactored notebook according to new notebook template

* Colab enterprise link fix

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-07-29 21:57:29 +00:00
8a679b5e41 refactor(egen): Colab enterprise link fix (#3340)
* <refactor>: refactored code according to notebook template

* <refactor> refactored notebook according to template

* Colab enterprise link fix

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-07-29 21:56:47 +00:00
d30089754b refactor(egen): Colab enterprise link fix (#3347)
* <fix, chore, refactor> refactored notebook according to template

* lint fix

* Apply suggested edits from @kittyabs review

* Colab enterprise link fix

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-07-29 21:55:29 +00:00
df7df07acf refactor(egen): Colab enterprise link fix (#3346)
* <refactor, chore> refactored notebook according to template

* Apply suggested edits from @kittyabs review

* Colab enterprise link fix

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-07-29 21:54:12 +00:00
96fa6122cf refactor,chore(egen) : refactored code according to template guidelines (#3278)
* refactor,chore(egen) : refactored code according to template guidelines and removed unused imports

* refactor(egen) : removed hardcoded values

* refactor(egen) : performed linter test

* refcator(egen) : refcatored according to template guidelines

* refactor,chore(egen): refactored code as per guidelines

* refactor,chore(egen): refactored code as per guidelines

---------

Co-authored-by: Jayakrishna2801 <jayakrishna.rajaboina@egen.ai>
2024-07-29 21:53:33 +00:00
7995377baf refactor(egen): Colab enterprise link fix (#3348)
* <refactor>: refactored code according to notebook template

* <refactor>: refactored code according to notebook template

* <refactor>: refactored code according to notebook template

* <refactor,chore> refactored notebook according to template

* <refactor,chore> refactored notebook according to template

* fix for  docker repository creation in PR test environment

* <included IS_TESTING condition for docker repository

* lint fix

* Apply suggested edits from @kittyabs review

* Colab enterprise link fix

* Colab enterprise link fix

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-07-29 21:52:13 +00:00
Aaron DietzandGitHub 1b51866384 Update tensorboard_profiler_custom_training.ipynb (#3350)
Updated name of Cloud Profiler (used to be called various versions of Tensorboard Profiler etc.). It's Cloud Profiler on first use, Profiler (shortened) for further uses.
2024-07-29 21:51:18 +00:00
2edd90b3b5 refactor,chore(egen) : refactored code as per template guidelines (#3291)
* refactor,chore(egen) : refactored code as per template guidelines

* refactor(egen): refactored notebook as per template guidelines

---------

Co-authored-by: Jayakrishna2801 <jayakrishna.rajaboina@egen.ai>
2024-07-29 21:49:48 +00:00
44ffa09022 refactor(egen): template fixes, adds clean up steps (#3290)
* chore, refactor: updates bqml_arima_plus notebook with the template changes

* <Refactor> Refactored the notebook according to the template.

* Applied suggested edits.

* Added db-dtypes package installation.

* Added db-dtypes package installation.

* fixed the pred_pipeline job issue.

* fixed the pred_pipeline job issue.

---------

Co-authored-by: k-root <kaushik.koilada@egen.ai>
Co-authored-by: UBhavani <bhavani.ummadi@egen.ai>
2024-07-29 21:47:53 +00:00
Holt SkinnerandGitHub 7ea4669471 Update run_linter.sh (#3349)
Fix spelling error `notebooked` → `notebooks`
2024-07-29 15:48:53 +00:00
Mend RenovateandGitHub d24204a2f8 chore(deps): update dependency pyupgrade to v3.17.0 (#3336) 2024-07-29 15:34:40 +00:00
Kathy YuandGitHub 462f0b6276 Update vLLM version in Llama 3.1 and Guard deployment notebooks. (#3335) 2024-07-27 00:25:53 +00:00
Ravi DalalandGitHub 2111ade4a6 Update github id in CODEOWNERS file (#3334) 2024-07-26 17:45:40 +00:00
Kathy YuandGitHub 07579d4274 Update OpenAI API Llama 3.1 and RAG notebooks. (#3323) 2024-07-26 16:54:57 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
20b9538572 chore(deps): bump torch (#3328)
Bumps [torch](https://github.com/pytorch/pytorch) from 1.13.1 to 2.2.0.
- [Release notes](https://github.com/pytorch/pytorch/releases)
- [Changelog](https://github.com/pytorch/pytorch/blob/main/RELEASE.md)
- [Commits](https://github.com/pytorch/pytorch/compare/v1.13.1...v2.2.0)

---
updated-dependencies:
- dependency-name: torch
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2024-07-26 16:51:19 +00:00
Kathy YuandGitHub 9ec0d3336a Update Llama Guard and synthetic data generation notebooks. (#3332) 2024-07-26 16:50:39 +00:00
Kathy YuandGitHub af45c644e3 Add variant in Llama 3.1 deployment notebook. (#3329) 2024-07-26 16:49:47 +00:00
Kathy YuandGitHub 2d092701c4 Update vLLM version in Llama 3.1 and Guard deployment notebooks. Fix bugs. (#3321) 2024-07-24 22:48:18 +00:00
weiran-workandGitHub c9272f1f85 chore: remove outdated benchmark report (#3320) 2024-07-24 22:46:23 +00:00
a37ff0e1f4 refactor, chore(egen): Removes boilerplate, heading fixes, and other corrections from template (#3102)
* refactor, chore(egen): Removes boilerplate, heading fixes, and other corrections from template

* refactor, chore(egen): Removes boilerplate, heading fixes, and other corrections from template

* Fixed issue related to image building

* Did minor fixes regarding docker image path and bucket creation command

* Updated command of bucket creation

* Did major changes in docker related code

* Fixed issue raised on PR

* Removed output of executed cells

* fix, refactor, chore: updates the model saving location in training script, removes the unnecessary code for training, refactors and updates the training image creation section accordingly

* chore: updates explanation about python package in the overview section

* fix, chore: specifies the working dir while building and running the container, adds '.' in a sentence in Overview

---------

Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-07-24 20:05:34 +00:00
sumanvitaandGitHub 89e243899e refactor, chore(egen): Changes aiplatfrom import statement, refactors as per notebook template guidelines, linter performed (#3315)
* refactor, chore(egen):  Changes aiplatfrom import statement, refactors as per notebook template guidelines, linter performed

* case change
2024-07-24 19:57:26 +00:00
0181e7bc2a refactor(egen): Colab enterprise link fix (#3307)
* <refactor>: refactored code according to notebook template

* <refactor>: refactored code according to notebook template

* <refactor>: refactored code according to notebook template

* <refactor>: refactored code according to notebook template

* <refactor>: refactored code according to new notebook template

* <refactor>: refactored code according to new notebook template

* <refactor>: refactored code according to new notebook template

* <refactor>: refactored notebook according to new notebook template

* <refactor>: refactored notebook according to new notebook template

* Updated colab enterprise link

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-07-24 19:54:38 +00:00
75bd4bb561 fix: fixes the path in the colab enterprise link (#3303)
Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-07-24 19:52:20 +00:00
sumanvitaandGitHub 3959a5f1d7 refactor, chore(egen): replaces region with location, adds additional parameters in dataset_delete() function inside cleanup cell (#3300)
* refactor, chore(egen): replaces region with location, adds additional parameters in dataset_delete() function inside cleanup cell

* wording changes

* wording and case change
2024-07-24 19:50:57 +00:00
16e87c024a fix,chore,refactor(egen): minor changes to the automl_image_classification_online_prediction notebook. (#3299)
* fix,chore,refactor(egen):Changed REGION variable name to LOCATION, modified the import file, added version for the tensorflow package in installation step, added comments in cleanup section, removed os.getenv(IS_TESTING) from the cleanup section, refactored code according to the template guidelines and performed linter test.

* chore(egen): Done changes according to @kittyabs and performed linter test.

---------

Co-authored-by: sriramya2610 <sriramya.peddapally@egen.ai>
2024-07-24 19:48:17 +00:00
sumanvitaandGitHub 07045c6f3b fix, refactor, chore(egen): adds numpy==1.23.0 and replaces google-vizier==0.0.4 to resolve dependency errors, replaces REGION with LOCATION, refactors code (#3295)
* fix, refactor, chore(egen): adds numpy==1.23.0 and replaces google-vizier==0.0.4 to remove dependency errors, refactors code as per template guidelines, performs linter

* adds code highlight

* removes overview as per PR comments
2024-07-24 19:45:31 +00:00
sumanvitaandGitHub 76c089430c fix, chore, refactor(egen): Adds numpy==1.23.0, removes version ==0.0.4 from vizier installation, replaces K80 with T4, refactors as per template guidelines (#3292)
* fix, chore, refactor(egen): Adds numpy==1.23.0, removes version ==0.0.4 from vizier installation, replaces K80 with T4, refactors as per template guidelines

* removes use of future tense

* case change, wording changes as per PR comments
2024-07-24 19:43:26 +00:00
Xiang XuandGitHub 0de6d08a16 Fix input template in model_garden_pytorch_llama3_1_deployment (#3317) 2024-07-24 18:43:01 +00:00
d703f31f89 feat: Update vllm and peft docker URI. (#3314)
Co-authored-by: Weiran <weiranzhao@google.com>
2024-07-24 17:48:14 +00:00
Sujit KhasnisandGitHub 8963a9275f fix: typo in Mistral AI Colab ent link (#3316) 2024-07-24 15:55:27 +00:00
Sujit KhasnisandGitHub 623c15662a feat: Official notebook for Mistral AI Release 07/24 (#3308)
* feat: Official notebook for Mistral AI Release 07/24

* feat: Official notebook for Mistral AI Release 07/24,added links to Vertex, Public docs

* feat: Official notebook for Mistral AI Release 07/24; links reorged

* feat: Official notebook for Mistral AI Release 07/24;lint issue resolved

* feat: Official notebook for Mistral AI Release 07/24;large name change [2407]
2024-07-24 15:34:41 +00:00
Xiang XuandGitHub c62089b99d Fix request format in model_garden_pytorch_llama3_1_deployment (#3313) 2024-07-24 02:24:51 +00:00
Xiang XuandGitHub ed51eb4689 Fix endpoint in synthetic_data_generation_using_llama3_1.ipynb (#3312) 2024-07-23 18:35:21 +00:00
Ivan NardiniandGitHub 1915c8ca3e feat: add llama3_1 notebooks (#3311)
* add llama3_1 notebooks

* fix conflict
2024-07-23 16:00:01 +00:00
Kathy YuandGitHub 35db5c5889 Add Llama Guard, RAG, synthetic data generation notebooks. (#3310)
* Add Llama Guard, RAG, synthetica data generation notebooks.

* Fix linter
2024-07-23 15:27:37 +00:00
Xiang XuandGitHub 27f903deb0 Add llama3.1 finetune and deploy notebooks (#3309) 2024-07-23 15:15:47 +00:00
81d900051b refactor, chore(egen): Removes boilerplate, heading fixes, and other corrections from template (#3132)
* Adapted code with new notebook template

* Updated working libraries and replace all the REGION variable with LOCATION

* Testead code and gone through notebook template. Did required changes.

* Did required changes based on the feedbak given on PR

* Did required changes based on feedback given on PR

* fix, chore: upgrades the tensorflow version to the latest, reorganizes and rewords the heading structure to follow the tutorial flow, removes unnecessary code highlights, fixes typos

---------

Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-07-21 18:04:04 +00:00
c20d754717 chore(egen): Follows new template, removes IS_TESTING, REGION --> LOCATION (#3207)
* chore: Follows new template, removes IS_TESTING, relaces REGION with LOCATION

* chore: Cloud console --> Google Cloud console

* chore: addresses review comments, sets delete_bucket to True to remove the GCS bucket in the cleaning up step

---------

Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-07-21 17:57:56 +00:00
Kaushik KoiladaandGitHub b98aeb471f refactor, chore: (egen):sdk-automl-object-tracking-batch-prediction (#3249)
* chore: removes boiler plate, adds colab enterprise and edits acording to template

* chore: run end to end and reformat according to template

* chore: lint

* fix: removes testing induced error and runs lint

* chore: update REGION to LOCATION

* chore: lint test and update checks failure

* chore: address review comments
2024-07-21 17:52:56 +00:00
5cb93b266b fix,chore,refactor(egen): minor fix and changes to the model monitoring automl image online notebook. (#3268)
* fix,chore,refactor(egen): Added colab enterprise logo with link, heading changes according to template guidelines,changed REGION variable name to LOCATION, modified code in online prediction using the SDK interface, added the clean up code for endpoint and training job, added comments in cleanup code, refactored code according to template guidelines and performed linter test.

* chore(egen): removed version in the installation step and performed linter test

* chore,refactor(Egen):Done changes according to @kittyabs review and performed linter test.

---------

Co-authored-by: sriramya2610 <sriramya.peddapally@egen.ai>
2024-07-21 17:50:18 +00:00
67abf6ad8b chore,refactor(egen): Changed REGION variable name to LOCATION, removed IS_TESTING from the cleaning up section, refactored code according to template guidelines and performed linter test. (#3280)
* chore,refactor(egen): Changed REGION variable name to LOCATION, removed IS_TESTING from the cleaning up section, refactored code according to template guidelines and performed linter test.

* chore(egen): Done changes according to @kittyabs review and performed linter test.

---------

Co-authored-by: sriramya2610 <sriramya.peddapally@egen.ai>
2024-07-21 17:46:16 +00:00
3238d99f08 fix,chore,refactor(egen): Changed REGION variable name to LOCATION, changed CLUSTER_REGION variable name to CLUSTER_LOCATION, added gcloud command to enable dataproc cluster, refactored code according to the template guidelines and performed linter test. (#3284)
* fix,chore,refactor(egen): Changed REGION variable name to LOCATION, changed CLUSTER_REGION variable name to CLUSTER_LOCATION, added gcloud command to enable dataproc cluster, refactored code according to the template guidelines and performed linter test.

* chore(egen): Done changes according to @kittyabs review and performed linter test.

---------

Co-authored-by: sriramya2610 <sriramya.peddapally@egen.ai>
2024-07-21 17:44:03 +00:00
5388cd53f7 fix,chore,refactor(egen): Changed CLUSTER_REGION variable name to CLUSTER_LOCATION, added gcloud command to enable dataproc api, modified bigquery dataset name by replacing hyphens to underscores, removed uuid code generation and replaced uuid with unique, refactored code according to template guidelines and perfomed linter test. (#3286)
Co-authored-by: sriramya2610 <sriramya.peddapally@egen.ai>
2024-07-21 17:40:26 +00:00
Kaushik KoiladaandGitHub e5e68af5fa chore, refactor(egen): updates and refactors sdk_automl_video_classification_batch notebook (#3287)
* chore, refactor: adds colab enterprise and removes boilerplate

* chore: updates REGION to LOCATION and run end to end test

* chore: lint run
2024-07-21 17:38:29 +00:00
2b796ea454 <Refactor> Refactored the notebook according to the template. (#3288)
Co-authored-by: UBhavani <bhavani.ummadi@egen.ai>
2024-07-21 17:37:11 +00:00
66563cb700 refactor, chore(egen): refactored the notebook according to the template, updated and added new package (#3289)
* <Refactor, Chore> Refactored the notebook according to the template, updated and added new package.

* Applied suggested edits.

* Applied suggested edits.

---------

Co-authored-by: UBhavani <bhavani.ummadi@egen.ai>
2024-07-21 17:35:15 +00:00
Aaron DietzandGitHub 56516b496a Update tensorboard_profiler_custom_training_with_prebuilt_container.ipynb (#3294)
Updated name of Cloud Profiler (used to be called various versions of Tensorboard Profiler etc. It's Cloud Profiler on first use, Profiler (shortened) for further uses.
2024-07-21 17:31:02 +00:00
Ravi DalalandGitHub 5236aced75 updated instruction for retry block (#3293) 2024-07-19 17:32:30 +00:00
Ravi DalalandGitHub 61ea845e26 Spark on Ray on Vertex AI notebook (#3282)
* added example notebook for Spark on RoV

* added example notebook for Spark on RoV

* ran linter on spark_on_ray_on_vertex_ai.ipynb

* updated official CODEOWNERS file for spark on ray on vertex ai notebook

* fixed text

* fixed project and location variables for build

* lint run

* added docker authentication

* renamed docker repo

* added sdk version

* added quiet to docker authentication

* added explicit dependencies installation

* added google cloud aiplatform ray module installation

* added gcloud update and cleanup

* added a wait to avoid timeout error in the test build

* fixed cluster resource name in delete

* added timestamp suffix to cluster name

* added a 5 minutes wait after cluster creation

* address PR comments
2024-07-19 03:03:45 +00:00
21864fd3a5 refcator,chore(egen) : refactored code according to template guidelines (#3279)
* forecasting-retail-demand.ipynb

* refcator,chore(egen) : refactored code according to template guidelines

* refcator,chore(egen) : refactored code according to template guidelines

* refactor(egen) : refactored according to template guidelines

---------

Co-authored-by: Jayakrishna2801 <jayakrishna.rajaboina@egen.ai>
2024-07-19 02:11:11 +00:00
598c91ace0 refactor, chore(egen): Tensorflow version fix, grammar corrections, other corrections from template (#3277)
* <refator, chore> Adds Tensorflow in installation section, corrections from template

* <refator, chore> Adds Tensorflow in installation section, corrections from template

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-07-19 02:02:47 +00:00
sumanvitaandGitHub 91d027ee02 refactor, chore(egen): removes import os, refactors as per new template guidelines (#3275)
* refactor, chore(egen): removes import os, refactors as per new template guidelines

* changes did'nt to didn't

* changes made as per PR comments
2024-07-19 01:58:42 +00:00
0674f3cdec Refactor(egen): Corrections from template (#3270)
* <refactor, chore> Updated prebuilt container image for prediction to 1.3, scikit-learn package updated to 2.5.1, other corrections from template

* <refactor> refactored notebok according to notebook template

* <refactor> refactored notebok according to notebook template

* <refactor> refactored notebok according to notebook template

* <refactor> refactored notebok according to notebook template

* <refactor> refactored notebok according to notebook template

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-07-19 01:51:03 +00:00
sumanvitaandGitHub 59a1cd0a1a refactor, chore(egen): refactors code as per new template guidelines, removes future tense, replaces REGION with LOCATION, removes generate_uuid function (#3267)
* refactor, chore(egen): refactors code as per new template guidelines, removes future tense, replaces REGION with LOCATION, performs linter test

* you are changed to you're

* wording changes, performed linter test

* notebook changed to notebooks (plural)
2024-07-19 01:37:03 +00:00
f7937c7117 chore,refactor(egen): minor changes to prophet on vertex pipelines notebook. (#3263)
* chore,refactor(egen): Changed REGION variable name to LOCATION, changed DATA_REGION variable name to DATA_LOCATION, added cleanup code for pipeline jobs, batch prediction job, modified cleanup code for deletion of bigquery dataset, removed versions of packages in the install step, removed os.getenv(IS_TESTING) while cleanup bucket, refactored code according to template guidelines and performed linter test.

* chore,refactor(Egen):Done changes according to @kittyabs review and performed linter test.

* chore(egen): redefined the variables of training pipeline job name and prediction pipeline job name and performed linter test

* chore(egen): redefined the model variable in cleanup section and performed linter test

---------

Co-authored-by: sriramya2610 <sriramya.peddapally@egen.ai>
2024-07-19 01:27:19 +00:00
Kaushik KoiladaandGitHub abf3594267 refactor, chore (egen): refactors sdk_automl_video_action_recognition_batch notebook (#3276)
* chore, refactor: adds colab enterprise, remove boiler plate

* refactor: changes REGION to LOCATION

* chore: run end to end and remove testing variable

* chore: lint
2024-07-18 23:35:21 +00:00
Kaushik KoiladaandGitHub 3ab0d62efe chore, refactor (egen): refactors wide_and_deep_on_vertex_pipelines.ipynb notebook (#3273)
* chore, refactor: edits according to template, adds colab eneterprise and changes REGION to LOCATION

* chore: testing end to end

* chore: lint
2024-07-18 23:32:53 +00:00
Manu KumarandGitHub d893c7857c feat: add online serving w/multiple entities notebook (#3245) 2024-07-18 12:22:33 +00:00
Manu KumarandGitHub 463adaefb0 feat: add offline feature serving notebook (#3184) 2024-07-17 19:01:48 +00:00
Jose BracheandGitHub 8fc1b2fb42 feat: Adding a new persistent resources notebook that uses the Vertex AI SDK (#3274) 2024-07-17 12:55:28 +00:00
Aaron DietzandGitHub aa5ab3c64f Update get_started_with_custom_training_autologging_local_script.ipynb (#3236)
Fixed typo: paramenters --> parameters
2024-07-17 12:50:45 +00:00
siping-huandGitHub 10385c316b Update notebook to use Gemini model instead of text-bison because text-bison will be deprecated. (#3265)
* [AutoSxS] Replace 1p model `text-bison` to `Gemini` because text-bison will be deprecated.

* product name edit

* Replace gemini 1.0 pro to gemini 1.5 pro.

* Fix the error when downloading the public dataset.
2024-07-17 12:49:53 +00:00
6afa3968b6 fix,chore,refactor(egen): done minor changes to sdk automl image classification batch online notebook (#3261)
* fix,chore,refactor(egen): replaced import file of gcs with new one to create dataset, removed import statement of os module, removed os.getenv(IS_TESTING) while cleanup bucket, replaced UUID with unique and deleted code to generate UUID, changed aip to aiplatform, changed REGION to LOCATION, added endpoint.delete() to delete the endpoint, hardcoded TF version to 2.15.1, refactored code according to template guidelines and performed linter test.

* chore(egen): removed back ticks

---------

Co-authored-by: sriramya2610 <sriramya.peddapally@egen.ai>
2024-07-17 00:48:34 +00:00
6cc596accf refactor(egen) : refactored notebook according to template guidelines nd added missed modules and libraries (#3260)
* refactor(egen) : refactored notebook according to template guidelines and added missed imports and libraries

* refactor(egen) : added code to delete locally generated files

* refactor(egen) : refactored according to template guidelines

---------

Co-authored-by: Jayakrishna2801 <jayakrishna.rajaboina@egen.ai>
2024-07-17 00:43:56 +00:00
0dcce973bf chore,refactor(egen): Added opencv-python-headless and tensorflow==2.15.1 packages and cleanup code for local files and cloud storage bucket (#3258)
* chore,refactor(egen): Added opencv-python-headless and tensorflow==2.15.1 packages in installation step, Added import os statement in set machine type configuaration cell, Added cloud storage bucket and local files cleanup code, refactored code according to template guidelines and performed lintr test

* refactor(Egen):Done changes according to @kittyabs review and performed linter test.

* chore(egen): added IS_TESTING part while creating artifact repository and perfomred linter test

---------

Co-authored-by: sriramya2610 <sriramya.peddapally@egen.ai>
2024-07-17 00:37:55 +00:00
Kaushik KoiladaandGitHub 0791c52923 refactor, fix, chore(egen): edits get_started_bq_datasets (#3248)
* chore: refactor according to template, removes boilerplate, adds colab enterprise

* refactore: adds testing variables

* fix, chore: end to end testing with version change as fix

* chore: lint

* chore: addresses review comments and runs lint
2024-07-17 00:26:55 +00:00
575d2f9479 refactor, chore(egen): Removes boilerplate, heading fixes, and other corrections from template (#3071)
* Fixed code issue template issue in distillation file

* Did required changes in notebook template

* Did minor changes

* Added execption handling at cleanup step to handle error while performing cleanup

* Fixed issue based on feedback given on feedback

* fix, chore, refactor: removes hard-coded project-id, remove future tense and reorganizes the sections, refactors the cleaning up section

* fix, refactor, chore: cleans up the resources using display name rather than resource name, adds wait step to wait until the pipeline job is finished, updates the overview section to remove 'we'

* fix: runs _job.wait() instead of .wait() method for waiting, updates the var pipeline_job to pipeline

---------

Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-07-17 00:10:17 +00:00
sumanvitaandGitHub 50d70a3037 refactor, chore(egen): refactors code as per new template guidelines, adds code to delete custom job and locally generated files (#3205)
* refactor, chore(egen): refactors code as per new template guidelines, adds code to delete locally generated files

* added code to delete custom job in the cleanup section

* license year changed to 2022, removed you as per PR comments

* future to present tense

* adds tensorflow installation, protobuf version changes to resolve dependency issues
2024-07-17 00:01:11 +00:00
Aaron DietzandGitHub 25e0b91162 Fix typo in pytorch_gcs_data_training.ipynb (#3237)
Fixed typo: runing --> running
2024-07-15 14:12:10 +00:00
5d8489e3e0 chore, refactor, feat(egen): follows new template, simplifies code, adds cleanup step (#3259)
* chore, refactor, feat: follows new template, simplifies code for display-names, adds steps for deleting the resources in the cleaning up section

* chore: addresses the review comments

---------

Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-07-12 20:18:01 +00:00
sumanvitaandGitHub 8f2d26abb2 fix, refactor, chore(egen): adds endpoint.wait() to fix timeout error, hardcodes TF to 2.15.1, adds code in cleanup section to delete locally generated files, refactors code as per template guidelides, performs linter test (#3252)
* fix, refactor, chore(egen): adds endpoint.wait() to fix timeout error, hardcodes TF to 2.15.1, adds code in cleanup section to delete locally generated files, refactors code as per template guidelides, performs linter test

* replaces UUID with unique

* updates URL involving redirect
2024-07-12 20:13:27 +00:00
sumanvitaandGitHub fdba11d9b8 refacto, chore (egen): refactors code as per template guidelines, hardcodes TF version to 2.15.1,deletes locally generated files (#3251)
* refacto, chore (egen): refactors code as per template guidelines, hardcodes TF version to 2.15.1, adds code to delete locally generated files, markdown changes ,performs linter test

* adds space

* changed URL for notebook redirects
2024-07-12 19:58:56 +00:00
Katie NguyenandGitHub a537eab067 fix: branding corrections (#3256) 2024-07-12 17:56:16 +00:00
bcff5c3cb3 refactor, chore(egen): Removes boilerplate, heading fixes, and other corrections from template (#3018)
* Updated notebook templace and did minor changes in notebook comments

* Added link of colab enterprice

* Fixed the link related issue and removed unwated variable value

* Added below comment in notebook:
# @title Copyright & License (click to expand)

* Fixed the issue related to notebook template based on reviewers feedback

* Fixed issue based on feedback given on PR

* Fixed the issue based on feedback given on PR

* fix, chore: replace REGION with LOCATION, remove will, contracts 'is not'

---------

Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-07-12 17:03:25 +00:00
Alok PattaniandGitHub 6ef29df46f Updating dataset and other changes from review (#3257) 2024-07-12 16:50:08 +00:00
6c7fa4b3d8 refactor, chore(egen): Removes boilerplate, heading fixes, and other corrections from template (#3125)
* Adapted notebook with new notebook template

* Hardcode tensorflow version

* Removed unwated commentes from notebook

* Revert "Removed unwated commentes from notebook"

This reverts commit 1f120466b0.

* Perform lint code on notebook

* fix, chore: updates the deprecated matplotlib function, adds installation for matplotlib, removes future tense, removes 'we', removes try-except in the cleaning up section

---------

Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-07-12 05:28:10 +00:00
a7ed179e22 refactor,chore(egen) : refactored code according to template guidelines (#3254)
* refactor,chore(egen) : refactored code according to template guidelines and removed unused and deprecated code

* refactor,chore(egen) : refactored code according to template guidelines and removed unused and deprecated code

* refactor(egen) : refactored notebook according to template guidelines

---------

Co-authored-by: Jayakrishna2801 <jayakrishna.rajaboina@egen.ai>
2024-07-12 05:25:24 +00:00
16f117a7eb chore, refactor, feat(egen): Follows new template, markdown fixes, refactors (#3227)
* chore, refactor, feat: follows new template, fixes typos, updates dsl.Condition to ds.If, rewords the headings and organizes them as per the tutorial, adds a cleanup step for the pipeline file

* chore, fix: minor sentence corrections, removes undefined UUID parameter

* chore: addresses the review comments

* fix: fixes the var name pipeline --> pipeline_job

---------

Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-07-12 05:22:54 +00:00
Kaushik KoiladaandGitHub 236b2b751f chore, refactor, fix(egen): formats and fixes pytorch_distributed_training_reduction_server.ipynb (#3235)
* chore: removes boilerplate, adds colab enterprise, changes region to location, adds testing variables

* chore, fix: adds verification_mode to load_dataset to deal with error and runs end to end

* chore: removes testing code

* chore: lint

* chore: addresses review comments
2024-07-12 00:37:36 +00:00
f5d57558fe refactor, chore(egen): Replaces K80 GPU with T4, kfp and tensorflow versoin updates, pre-built Docker container image for training and prediction update, other corrections from template (#3219)
* <refactor, chore> replaces K80 GPU with T4, updates pre-built Docker container image for training and prediction to 2.13, cleansup intermediate files, updates kfp and tensorflow versions, fixes minor spelling mistakes and contracts words, removes future tense

* <refactor, chore> replaces K80 GPU with T4, updates pre-built Docker container image for training and prediction to 2.13, cleansup intermediate files, updates kfp and tensorflow versions, fixes minor spelling mistakes and contracts words, removes future tense

* <refactor, chore> replaces K80 GPU with T4, updates pre-built Docker container image for training and prediction to 2.13, cleansup intermediate files, updates kfp and tensorflow versions, fixes minor spelling mistakes and contracts words, removes future tense

* Grammar fix

* lint fix

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-07-12 00:34:20 +00:00
0f01fd7c18 chore(egen): follows new template, fixes typos, K80-->T4, REGION-->LOCATION etc. (#3214)
* chore: follows new template, fixes typos, removes unnecessary code-highlights, replaces K80 with T4, replaces REGION with LOCATION, removes IS_TESTING in cleaning up section, makes sentence/heading corrections and re-organizes some subsections as per the tutorial

* chore, refactor: minor markdown corrections, updates machine_type description and code to suit the explanation

* chore: addresses review comments and corrects 'uploading to a Vertex AI model resource' to 'uploading to Vertex AI Model Registry'

* chore: addresses the review comments

---------

Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-07-12 00:31:39 +00:00
4f7b63f3af refactor,chore(egen) : refcatored code according to template guidelines and added cleanup code (#3211)
* refactor,chore(egen) : refcatored code according to template guidelines and added cleanup code

* refactored code according to template guidelines

* refcatord code according to template guidelines

* refactor(egen) : refactored code according to template guidelines

* refactor(egen) : refactored code according to template guidelines

---------

Co-authored-by: Jayakrishna2801 <jayakrishna.rajaboina@egen.ai>
2024-07-12 00:23:53 +00:00
sumanvitaandGitHub 42bdad56e3 chore, refactor(Egen): adds code to delete batch prediction job and locally generated files, refactors code as per notebook template guidelines (#3202)
* refactor, chore(egen): refactors code as per new template guidelines ,hardcodes TF version to 2.15.1, changes K80 to T4

* adds code to delete batch prediction jobs in the clean up section

* Changed lower to upper case

* reverted license to 2022 as per comment

* case change, wording changes as per PR comments
2024-07-12 00:20:18 +00:00
dc702a614b refactor, chore(egen): Removes boilerplate, heading fixes, corrections from template. (#3181)
* <refactore, chore>refactored notebook according to the template

* refactor: Apply markdown text edit

* source distribution fix

* source distribution fix

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-07-12 00:12:27 +00:00
73517d4b40 chore(egen) : Adds deprecation note to the notebook (#3243)
* Added deprecation note

* format and lint fix

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-07-11 23:44:26 +00:00
e0a3e785ec refactor,chore(egen) : added delete experiment code in cleanup section and required packages in installation section (#3216)
* refactor,chore(egen) : added delete experiment code in cleanup section and added required packages in installation section, refactored code according to template guideline

* refactor,chore(egen): removed hardcoded values

* downgraded numpy version

* refactor(egen) : refactored code according to template guidelines

* refcator(egen) : refactored code accordi gto template guidelines

* refactor(egen) : refactored code according to template guidelines

---------

Co-authored-by: Jayakrishna2801 <jayakrishna.rajaboina@egen.ai>
2024-07-11 23:39:08 +00:00
b9ac5a0296 refcator, chore(egen) : refactored code according to template guidelines , performed linter test (#3215)
* refcator, chore(egen) : refactored code according to template guidelines

* chore(egen) : changed headings as per guidelines

* performs linter test

* refactored code accordig to template guidelines

* formatted according to template guidelines

* formatted according to template guidelines

* refcator(egen) : refactored code according to template guidelines

* refactor(egen) : added warning message for kernal restart

---------

Co-authored-by: Jayakrishna2801 <jayakrishna.rajaboina@egen.ai>
2024-07-11 23:36:01 +00:00
671e9f84fa refactor: refactor gemma notebooks (#3255)
Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2024-07-11 19:23:15 +00:00
9387236c01 feat: add a fn to resize an image (#3246)
Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2024-07-11 17:30:47 +00:00
8142ce631a Add instructions for securing more GPUs. (#3247)
Co-authored-by: minwoopark <minwoopark@google.com>
2024-07-11 17:30:11 +00:00
Huguens JeanandGitHub 8923b4a7bd [Vertex AI MG Team] Remove corp link in NeRF gradio application. (#3253) 2024-07-11 17:29:11 +00:00
803bc9b489 chore,refactor(egen): minor changes to sdk feature store notebook (#3240)
* chore,refactor(egen): Changed REGION variable name to LOCATION, Added cleanup code fro cloud storage bucket, refactored code according to the template and performed linter test

* chore(egen): replaced region variable with location and perfomred linter test

---------

Co-authored-by: sriramya2610 <sriramya.peddapally@egen.ai>
2024-07-10 22:26:31 +00:00
5dc036b311 refcator,chore(egen) : refcatored code according to template guidelines (#3234)
* refcator,chore(egen) : refcatored code according to template guidelines

* refcator(egen) : formatted accorded to template guidelines

* refactor(egen) : refactored code according to template guidelines,removed hardcoded values  and performed linter test

* refactor(egen) : added warning message for kernal restart

* refactor(egen) : refactored according to template guidelines

* refactor(egen) : refactored according to template guidelines

* refactor(egen) : refactored according to template guidelines

---------

Co-authored-by: Jayakrishna2801 <jayakrishna.rajaboina@egen.ai>
2024-07-10 22:24:11 +00:00
fd15c01586 chore,refaactor(egen): minor changes to sdk feature store pandas notebook (#3239)
* chore,refactor(egen): Changed REGION variable name to LOCATION, Removed os.getenv(IS_TESTING) while cleanup bucket, refactored code according to the template and performed linter test

* chore,refactor(egen): Changed REGION variable name to LOCATION, Removed os.getenv(IS_TESTING) while cleanup bucket, refactored code according to the template and performed linter test

* chore(egen): replaced region variable with location and perfomred linter test

---------

Co-authored-by: sriramya2610 <sriramya.peddapally@egen.ai>
2024-07-10 22:20:51 +00:00
sumanvitaandGitHub 84cadc2c9a fix,refactor,chore(egen): adds endpoint.wait() to resolve timeout error, hardcodes TF version to 2.15.1, refactors code as per template, performs linter test (#3231)
* fix,refactor,chore(egen): adds endpoint.wait() to resolve timeout error, hardcodes tf version to 2.15.1, refactors code as per tempalte, performs linter test

* contraction of words

* added code highlight
2024-07-10 22:19:19 +00:00
f694abc42e refactor,chore(egen) : refactored code according to template guidelines (#3230)
* refactor,chore(egen) : refactored code according to template guidelines

* added code to remove locally generated files

* refactor(egen) : refactored code according to template guidelines

---------

Co-authored-by: Jayakrishna2801 <jayakrishna.rajaboina@egen.ai>
2024-07-10 22:04:27 +00:00
Kaushik KoiladaandGitHub 99fa616a14 refactor, chore, fix(egen): edits google_cloud_pipeline_components_model_upload_predict_evaluate.ipynb notebook (#3229)
* chore: removes boiler plate and reformats according to the template

* chore: adds colab enterprise link to the notebook

* chore: testing end to end

* chore:end to end test with reformatting

* chore: lint

* chore: addresses comment on the markups
2024-07-10 22:02:12 +00:00
sumanvitaandGitHub 0583f152ac fix, chore, refactor(egen): hardcodes scikit-learn version to 1.2, changes python version from 3.9 to 3.10, adds numpy==1.26.4 installation, adds code to undeploy model from endpoints (#3228)
* fix, chore, refactor(egen): hardcodes scikit-learn version to 1.2, changes python version from 3.9 to 3.10, adds numpy==1.26.4 installation, adds code to undeploy model from endpoints, rusage of future tense

* wording changes

* markdown wording changes as per PR comments
2024-07-10 21:51:01 +00:00
sumanvitaandGitHub f22fee6f84 fix, refactor, chore(egen): removes keras3 dependency error while saving the model, refactors code as per template (#3222)
* fix, refactor, chore(egen): removes keras3 dependency error while saving the model, refactors code as per template, performs linter test

* set epochs to 14 as per original code

* changed CustomJob to Custom Job
2024-07-10 21:42:04 +00:00
650272e370 chore, fix(egen): follows new template, markdown updates and fixes (#3221)
* chore: follows new template, remove IS_TESTING, replace REGION with LOCATION, organizes headings and styles, removes unnecessary code highlights

* fix: removes USER var, adds IS_COLAB var

* chore: addresses the review comments

---------

Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-07-10 21:36:46 +00:00
8f9c783f97 refactor,chore(egen) : refactored code as per template guidelines , performed linter test (#3220)
* refactor,chore(egen) : refactored code as per template guidelines and performed linter test

* refcator(egen) : refcatored code according to template guidelines

* refcator(egen) : refcatored code according to template guidelines

* refcator(egen) : refcatored code according to template guidelines

* refactor(egen) : refactored code according to template guidelines

---------

Co-authored-by: Jayakrishna2801 <jayakrishna.rajaboina@egen.ai>
2024-07-10 21:27:54 +00:00
ed721006b7 refactor(egen): Automl video classification model evalution (#3217)
* refactor,chore(egen) : refactored according to template guidelines , performed linter test

* refactor,chore(egen) : removed hardcoded values , performed linter test

* refcatored according to template guidelines

* refactor(egen) : refactored code according to template guidelines

* refactor(egen) : refactored code according to template guidelines

---------

Co-authored-by: Jayakrishna2801 <jayakrishna.rajaboina@egen.ai>
2024-07-10 21:18:59 +00:00
f4b17078af chore,refactor(egen): Changed REGION variable name to LOCATION, replaced np.NaN with np.nan, refactored code according to the template and performed linter test (#3242)
Co-authored-by: sriramya2610 <sriramya.peddapally@egen.ai>
2024-07-10 21:11:17 +00:00
8429bba776 chore,refactor(egen):changed the versions of tensorflow, tensorflow-hub, apache_beam[gcp] and bs4 in requirements.txt and setup.py files, refactored code according to template guidelines (#3233)
Co-authored-by: sriramya2610 <sriramya.peddapally@egen.ai>
2024-07-10 21:05:47 +00:00
9e42108fd0 chore,refactor(egen): minor updates to Automl Tabular Classification Model Evaluation Notebook (#3218)
* chore,refactor(Egen): Removed the google-cloud-pipeline-components package version, IS_TESTING Variable and import statement of os module from the cleaning up section, Replaced REGION variable with LOCATION, refactored code according to template guidelines and performed linter test.

* chore(egen):renamed location variable to LOCATION

---------

Co-authored-by: sriramya2610 <sriramya.peddapally@egen.ai>
2024-07-10 18:09:16 +00:00
1593c06811 chore(egen):Added the note for deprecation of notebook and performed linter test (#3182)
* chore(egen):Added the note for deprecation of notebook and performed linter test

* chore(egen):changed the title of the link in the deprecated note

---------

Co-authored-by: sriramya2610 <sriramya.peddapally@egen.ai>
2024-07-10 14:11:14 +00:00
7ccb27b044 chore(egen) : added note to deprecated notebook and performed linter test (#3209)
* chore(egen) : added note to deprecated notebook

* chore(egen) : performed linter test

---------

Co-authored-by: Jayakrishna2801 <jayakrishna.rajaboina@egen.ai>
2024-07-10 14:10:38 +00:00
1a4db478de chore(egen) : added note to deprecated notebook and performed linter test (#3210)
* chore(egen) : added notes to deprecated notebook

* chore(egen) : performed linter test

* chore(egen) : added notes to deperecated notebook and performed linter test

* linter test

---------

Co-authored-by: Jayakrishna2801 <jayakrishna.rajaboina@egen.ai>
2024-07-09 21:03:53 +00:00
Kaushik KoiladaandGitHub 8e5fd0545d refactor, chore (Egen) : Fixes headings, removes boilerplate and other corrections from template (#3044)
* refactor: removes boilerplate code, and fixes heading

* refactor: change region to location

* Removes IS_TESTING variable

* cleanup of changes

* chore: lint test done

* chore: aligns the icons to the center

* chore:verbiage changes and end to end code execution

* chore: reformatted by lint test

* chore: edits future tenses and reformatted by lint test

* chorE: address review comments and change import statement based on lint test

* fix: error rectification, remove vague testing variables

* fix: rectifies testing induced error in notebook

* chore: lint
2024-07-09 20:47:48 +00:00
Liang WuandGitHub c74a714a51 Support Mistral-7B-v0.3 and Mistral-7B-Instruct-v0.3 in deployment notebook. (#3226) 2024-07-09 19:16:21 +00:00
Aaron DietzandGitHub 2104f4c478 Fix typo in sdk_pytorch_torchrun_custom_container_training_imagenet.ipynb (#3238)
Fixed typo: Github --> GitHub
2024-07-09 18:31:37 +00:00
2e330ab7ab fix,chore,refactor(Egen): replaced aip with aiplatform and k80 GPU with T4 GPU, changed the versions of images,.keras extension added when saving and uploading the model, removed import statement of os module, refactored code according to template guidelines and performed linter test. (#3213)
Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-07-09 03:34:01 +00:00
141e77caec refactor, chore, fix(egen): Pipeline documentation link update, removes GPU from machineSpec in the pipeline, package version updates, fixes import errors (#3206)
* <refactor, chore, fix> package version updates, pipeline components documentation link update, importer_node import fix, machineSpec update

* <refactor, chore, fix> package version updates, pipeline components documentation link update, importer_node import fix, machineSpec update

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-07-09 03:04:57 +00:00
Kaushik KoiladaandGitHub 70a65386d2 refactor, chore(egen): edits build_model_experimentation_lineage_with_prebuild_code notebook (#3203)
* chore,refactor: removes boiler plate, removes version of aiplatform package, adds colab enterprise and formats according to the template

* chore: run end to end

* chore: lint run
2024-07-09 02:42:42 +00:00
00729da920 refactor, chore(egen): Refactored code according to template guidelines (#3136)
* refactor, chore(egen): Refactored code according to template guidelines, performed linter test

* refactor, chore(egen): Removed vertexai SDK initiation in the beginning, performed linter test

* refactor, chore(egen): Made some grammatical changes in markdown script, performed linter test

* refactor(egen): modified code to delete locally created file, code to delete custom job

* comment change

---------

Co-authored-by: sumanvita-springml <sumanvita.kandregula@egen.ai>
2024-07-09 02:34:34 +00:00
Eric DongandGitHub 3e33b7e0c6 Update README.md (7) (#3223) 2024-07-09 02:09:34 +00:00
d57617c726 remove: remove model_garden_pytorch_mistral notebook (#3225)
Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2024-07-09 01:10:57 +00:00
Kaushik KoiladaandGitHub 24ff289855 chore, refactor(egen): format and refactor for get_started_with_model_monitoring_custom_tf_serving notebook (#3162)
* chore, refactor: adds colab enterprise, removes boilerplate, formats based on template, changes region to location

* chore,refactor: run end to end and format according to template

* chore: Lint test

* chore: comments on lower case addressed and lint run

* chore: comments on lower case addressed and lint run

* chore: address comment on lowercase of resource names
2024-07-05 17:41:44 +00:00
6183a71af4 refactor, chore, fix(egen): Documentation link update, package version updates, fixes import errors (#3208)
* <Refactor> Refactored the notebook according to the template.

* <refactor, chore, fix> Refactored the notebook according to the template, updated the documentation link and package version, fixed the import errors.

---------

Co-authored-by: UBhavani <bhavani.ummadi@egen.ai>
2024-07-04 18:53:37 +00:00
8eb8db2ef8 chore , refactor : removed ! rm lightweight_pipeline.json , Added cleanup code for deletion of pipeline and refactored code according to template guidelines , performed linter test (#3149)
* refactor : removed ! rm lightweight_pipeline.json

* chore,refactor : Added cleanup code for deletion of pipeline and refactored code according to template guidelines , performed linter test

* chore,refactor : Added cleanup code for deletion of pipeline and refactored code according to template guidelines , performed linter test

* refactor,chore : Added code for deletion of pipeline and refactored code according to template guidelines , performed linter test

* refactor,chore : refactored code according to template guidelines and downgraded numpy version

* refactor : removed project name and bucket name used for testing in local

* refactor : refactored code according to template guidelines

* performed linter test

---------

Co-authored-by: Jayakrishna2801 <jayakrishna.rajaboina@egen.ai>
2024-07-03 23:05:12 +00:00
a749fa3377 chore, refactor(egen): replace dsl.Condition with dsl.If, sentence corrections (#3199)
* refactor(egen): refacted markdown as per notebook template

* chore, refactor: removes future tense, minor markdown fixes, removes UUID and uses -unique

* chore, fix: replaces dsl.Condition with dsl.If, fixes some sentences and explanation

---------

Co-authored-by: sumanvita-springml <sumanvita.kandregula@egen.ai>
Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-07-03 13:45:04 +00:00
b2e9dbdb54 chore(egen):Added the note to the deprecation of the notebook and performed linter test (#3191)
Co-authored-by: sriramya2610 <sriramya.peddapally@egen.ai>
2024-07-03 00:36:47 +00:00
f537d8ef4b chore(egen):Added the note to the deprecation of notebook and performed linter test (#3190)
Co-authored-by: sriramya2610 <sriramya.peddapally@egen.ai>
2024-07-03 00:36:02 +00:00
4a3260d92b chore(egen) : Adds the note to the deprecation of notebook and performs linter test (#3189)
* chore(egen):Added the details of the deprecation of the notebook and performed linter test

* chore(egen):changed the title of the link in the deprecated note

---------

Co-authored-by: sriramya2610 <sriramya.peddapally@egen.ai>
2024-07-03 00:35:21 +00:00
sumanvitaandGitHub 952f2d0aa6 chore(egen): Adds note to the deprecated notebook (#3188)
* chore(egen): Adds note to the deprecated notebook

* removes space
2024-07-03 00:34:38 +00:00
sumanvitaandGitHub 7f920f173f chore(egen): Adds note to the deprecated notebook (#3187)
* chore(egen): Adds note to the deprecated notebook, performs linter test

* removes extra space from the note
2024-07-03 00:33:51 +00:00
sumanvitaandGitHub 656f8d27d1 chore(egen): Adds note to the deprecated notebook (#3186)
* chore(egen): Adds note to the deprecated notebook, performs linter test

* removes extra space from note
2024-07-03 00:32:34 +00:00
sumanvitaandGitHub 03153ca48b chore(egen): Adds note to the deprecated notebook (#3185)
* chore(egen): Adds note to the deprecateed notebook and performs linter test

* Removes extra space in note
2024-07-03 00:30:39 +00:00
cecef79ae5 refactor(egen): template fixes, adds clean up steps (#3160)
* <Refactor> Refactored the notebook according to the template.

* <Refactor> Refactored the notebook according to the template.

* applied suggested edits.

* Changed region to location.

---------

Co-authored-by: UBhavani <bhavani.ummadi@egen.ai>
2024-07-03 00:01:32 +00:00
Eric DongandGitHub 13acdda204 fix: remove --user from package install (#3196) 2024-07-02 23:36:53 +00:00
Gary WeiandGitHub 23ac65b212 Add a Colab notebook for local dreambooth finetune user experience. (#3183)
* Create a Gradio notebook for the new InstantId model.

* Add dreambooth finetune to the stable diffusion Gradio workshop notebook.

* Update the image generation Gradio notebook to support Dreambooth finetuning.

* linter update

* linter update

* minor fix to the instant-id notebook.

* Minor fix to the stable diffusion gradio notebook.

* Split the 'instant-id' deployment notebook prediction into two sections.

* add `deployment_source` to the notebook.

* Switch to `pytorch-diffusers-serve-opt` container to for diffusion lora serving.

* add the dreambooth_lora notebook.

* minor update.

* Parameterize the "show_debug_logs" to facilitate automatic test of the Gradio notebooks.

* Lint format.

* minor updates

* Delete the two deprecated SD1.5 and 2.1 notebooks, as they were no longer referenced on any model cards.

* Sync Colab notebooks between g3 and github.

* format changes

* format update.

* Improve the SDXL-dreambooth-lora finetune notebook CUJ.

* Update the diffusers serving docker image version to `20240605_1400_RC00` to resolve vulnerabilities.

* Add dreambooth-lora-sdxl task for SDXL base model in the dreambooth finetune Gradio notebook.

* Add dreambooth-lora-sdxl task for SDXL base model in the dreambooth finetune Gradio notebook.

* Add dreambooth-lora-sdxl task for SDXL base model in the dreambooth finetune Gradio notebook.

* Add a new notebook for local dreambooth finetune user experience.
2024-07-02 12:55:29 +00:00
df0e5a09cc chore,refactor(Egen): replaced aip with aiplatform and kfp.v2 with kfp, removed import statement of os module, refactored code according to template guidelines and performed linter test. (#3170)
* refractor,chore(egen): replaced kfp.v2 with kfp in compile step, refracted code according to template guidelines, performed linter test

* chore,refactor(Egen): replaced aip with aiplatform, removed import statement of os module, refactored code according to template guidelines and performed linter test.

* refactor(egen):refactored code by renaming REGION variable to LOCATION and performed linter test

---------

Co-authored-by: sriramya2610 <sriramya.peddapally@egen.ai>
2024-07-02 01:49:49 +00:00
bb3d3bde21 fix, chore, refactor(egen): REST api fixes, KFP v2 refactor, new template (#3180)
* fix, chore, refactor: fixes the issue with job creation request, refactors to use the latest SDKs, follows the new template, sentence corrections

* chore: removes future tense, minor sentence corrections

* chore: addresses the review comments

---------

Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-07-01 20:40:05 +00:00
Kaushik KoiladaandGitHub 30f2a2ba03 Egen fix/delete outdated tensorboard experiments (#3178)
* chore: removes boiler plate and changes region to location

* chore: end to end run and lint
2024-07-01 20:32:44 +00:00
Kaushik KoiladaandGitHub f628caeb6d chore, fix, refactor(egen): fixes get_started_with_vertex_experiments_autologging notebook (#3177)
* chore, refactor: adds colab enterprise and refactors according to template

* fix, refactor: fixes issue with loading input and output with the types expected

* chore: lint
2024-07-01 20:31:02 +00:00
7e35553dbf refactor(egen): Corrections from template. (#3172)
* <refactor> refactored notebook according to new template

* Apply suggested edits

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-07-01 20:25:43 +00:00
sumanvitaandGitHub baa167a8d5 chore, refactor(egen): hardcodes TF version to 2.15.1, refactors code as per template, changes future to present tense (#3171)
* chore, refactor(egen): hardcodes TF version to 2.15.1, refactors code as per template,changes future to present tense

* refactor(egen): removes hypen between words

* refactor(egen): wording change

* contracts words like it is to it's
2024-07-01 20:21:41 +00:00
sumanvitaandGitHub ee4e4bb2ce Fix, chore, refactor(Egen) : hardcodes tensorflow dependency to 2.15.1, K80 to T4, refactors code as per notebook template guidelines (#3161)
* refactor, chore(egen): refactored code as per new notebook template, changes REGION to LOCATION

* chore, refactor(egen): adds code for deletion of locally generated files and other resources, rephrases sentences

* refactor(egen): changes lower to uppercase according to PR comments
2024-07-01 20:13:28 +00:00
ed8a842525 refactor, chore(egen): Replaces K80 with T4 GPU, Documentation link update, corrections from template. (#3163)
* <refactor, chore> refactored notebook according to new template

* Applied suggested edits

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-07-01 20:07:33 +00:00
56faf31221 Egen reviewed explainable ai (#3173)
* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

added colab enterprise logo and link

* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

hope I fixed the JSON issue

* fix: remove new changes

* Update vertex_ai_feature_store_based_llm_grounding_tutorial.ipynb

remove back ticks from product/feature names

* Update xai_image_classification_feature_attributions.ipynb

removed back ticks from product names.

---------

Co-authored-by: Katie Nguyen <21978337+katiemn@users.noreply.github.com>
2024-07-01 20:01:51 +00:00
7a15b88071 feat: Add notebook/colab example for prediction PSC based private (#3080)
endpoint.

Co-authored-by: TJ(Tianjiao) Liu <tianjiaoliu@google.com>
2024-06-29 00:35:54 +00:00
Gary WeiandGitHub 7ceda5e4e6 Add the TGI serving section to the Gemma2 deloyment notebook. (#3175)
* Add the TGI serving section to the Gemma2 deloyment notebook.

* Update the TGI serving container URI.
2024-06-29 00:34:32 +00:00
2a649e0a2e Egen reviewed model monitoring (#3176)
* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

added colab enterprise logo and link

* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

hope I fixed the JSON issue

* fix: remove new changes

* Update vertex_ai_feature_store_based_llm_grounding_tutorial.ipynb

remove back ticks from product/feature names

* Update get_started_with_model_monitoring_custom.ipynb

small edits

---------

Co-authored-by: Katie Nguyen <21978337+katiemn@users.noreply.github.com>
2024-06-29 00:14:36 +00:00
kittyabsandGitHub b2b97dddb0 Update get_started_with_model_registry.ipynb (#3174)
made numerous edits
2024-06-28 19:15:42 +00:00
800e92597c Add a notebook for MaMMUT (#3169)
* feat: Add a notebook for MaMMUT

* feat: Update CODEOWNERS

* fix: Remove unused import

---------

Co-authored-by: Ivy Wang <jiananwang@google.com>
2024-06-28 18:17:16 +00:00
skarukasandGitHub b895a348cc Add learning_rate_multiplier and output_dimensionality parameters to the text embedding tuning notebook. (#3165)
* Add learning_rate_multiplier and output_dimensionality to embedding tuning notebook.

* Reformat
2024-06-28 13:58:03 +00:00
Ivan NardiniandGitHub 4e2c698029 fix: update the torch sample on Ray on Vertex AI (#3131)
* review the torch rov notebook

* linter passed

* fix typos

* linter passed

* fix issue

* linter passed

* fix typos

* fix typos
2024-06-28 13:56:01 +00:00
cddfb9cb09 Featurestore notebook egen edited (#3167)
* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

added colab enterprise logo and link

* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

hope I fixed the JSON issue

* fix: remove new changes

* Update vertex_ai_feature_store_based_llm_grounding_tutorial.ipynb

remove back ticks from product/feature names

* Update vertex_ai_feature_store_feature_view_service_agents.ipynb

remove back ticks

---------

Co-authored-by: Katie Nguyen <21978337+katiemn@users.noreply.github.com>
2024-06-27 23:50:58 +00:00
Kathy YuandGitHub 510eb3855c Add Hex-LLM TPU deployment to Gemma, Code Gemma and Gemma 2 notebooks. (#3168)
* Add Hex-LLM TPU deployment to Gemma, Code Gemma and Gemma 2 notebooks.

* Fix linter issues.
2024-06-27 23:31:56 +00:00
ethan-gordonandGitHub 7a96949b68 Add vertex_ai_feature_store_iam_policy notebook. (#2913)
* Add vertex_ai_feature_store_iam_policy notebook.

* update CODEOWNERS

* Fix formatting of vertex_ai_feature_store_iam_policy.ipynb
2024-06-27 20:36:48 +00:00
KCFindstrandGitHub 2b0dd757d5 Switch movinet serving notebooks to use port 8080 (#3157) 2024-06-27 20:34:51 +00:00
Huguens JeanandGitHub 90c82b7b7b [MG Model Team] Cleanup cloudnerf gradio notebook outputs. (#3151) 2024-06-27 20:27:02 +00:00
praccu-googleandGitHub 8607144ed4 Add hugging face token to mixtral example colab. (#3158) 2024-06-27 20:24:50 +00:00
a8f9cbbaa6 fix, refactor(egen): fixes data import issue, includes pipeline job deletion, corrections from template. (#3150)
* <fix, refactor> fixed and refactored notebook according to the template

* Apply suggested edits from @kittyabs

* Apply suggested edits from @kittyabs

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-06-27 20:15:29 +00:00
b96dd105d9 chore,refactor(egen) : Added clean up code for deletion of model endpoint and local files, refactored code according to template guidelines, performed linter test (#3147)
* refractor(egen): refracted code according to template guidelines, performed linter test

* refractor(egen):  refracted code by removing os.getenv(IS_TESTING) in clean up code of bucket

* chore,refactor(egen):grammar check according to template guidelines, refactored the code by adding clean up code for deletion of local files created and model  endpoint and performed linter test.

* refactor(egen):refactored code by modifying clean up code and performed linter test

---------

Co-authored-by: sriramya2610 <sriramya.peddapally@egen.ai>
2024-06-27 20:07:57 +00:00
6bd828833f refactor, chore(egen): Refactored code according to template guidelines (#3135)
* refactor, chore(egen): Refactored code according to template guidelines, performed linter test

* refactor, chore(egen): Removed vertexai SDK initiation at the beginning, performed linter test

* chore, fix: removes future tense(will), fixes workbench-specific dependency compatibility issue by fixing versions, removes Only at bucket creation, fixes categorical fields while encoding, adds comments in cleanup, adds project-id in gsutil command

---------

Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-06-27 20:03:44 +00:00
sumanvitaandGitHub 4d19f36d39 chore, refactor(egen) : changes file extension from .json to .yaml, refactored code as per template guidelines (#3146)
* fix, refactor, chore(egen): replaced .json to .yaml, refacted the code as per notebook template, performed linter test.

* refactor, chore(egen): Rephrases sentences, changes from google.cloud import aiplatform instead to import google.cloud.aiplatform as aip

* refactor(egen(egen): cleared cell output

* refactor, chore(egen): changes made as per PR comments, perfomed linter test
2024-06-27 20:02:21 +00:00
a877ca3501 fix,refactor,chore(egen): replaced np.float with np and K80 GPU with T4 GPU, refacted code as per template guidelines (#3138)
* fix,refractor,chore(egen): replaced np.float with np, refracted code according to template guidelines, performed linter test

* refractor(egen): refracted code according to template guidelines, performed linter test

* chore, fix: corrects/rewords some sentences, replaces K80 with T4

* chore(egen): spell check

* refactor, chore(egen): wording changes, contracts words like is not to is'nt, ran linter test

* refactor, chore(egen): rewording ,reintroduces import numpy as np in task.py

* fix(egen): removes delete_custom_folder variable in cleanup section

---------

Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-06-27 19:57:59 +00:00
fd1b0f8383 Featurestore notebook#3 (#3155)
* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

added colab enterprise logo and link

* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

hope I fixed the JSON issue

* fix: remove new changes

* Update online_feature_serving_and_fetching_bigquery_data_with_feature_store_optimized.ipynb

made edits and also corrected the URL involving "pantheon", changing it to: https://console.cloud.google.com

---------

Co-authored-by: Katie Nguyen <21978337+katiemn@users.noreply.github.com>
2024-06-27 18:27:55 +00:00
kittyabsandGitHub 4fe89b9849 Update online_feature_serving_and_fetching_bigquery_data_with_feature_store_bigtable.ipynb (#3164)
removed back ticks from product names
2024-06-27 18:20:31 +00:00
d973169c0c Featurestore notebook non egen (#3156)
* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

added colab enterprise logo and link

* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

hope I fixed the JSON issue

* fix: remove new changes

* Update online_feature_serving_and_fetching_bigquery_data_with_feature_store_bigtable.ipynb

Some small edits, but mainly, I changed the URL that used "pantheon" to https://console.cloud.google.com,
Note: egen has not reviewed this notebook yet, so I didn't do a more detailed edit. Will do that once they've updated the notebook.

---------

Co-authored-by: Katie Nguyen <21978337+katiemn@users.noreply.github.com>
2024-06-27 17:34:00 +00:00
47ef3e3f6c Featurestore notebook#2 (#3154)
* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

added colab enterprise logo and link

* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

hope I fixed the JSON issue

* Update online_feature_serving_and_vector_retrieval_bigquery_data_with_feature_store.ipynb

some small edits

* fix: remove new changes

---------

Co-authored-by: Katie Nguyen <21978337+katiemn@users.noreply.github.com>
2024-06-26 22:42:29 +00:00
d43c5fee3d chore, refactor(egen): follows new template, removes IS_TESTING, spell corrections (#3148)
* chore, refactor: follows new template, removes IS_TESTING in clean up, contracts words and some steps, spell correct, K80 to T4

* chore: addresses review comments

---------

Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-06-26 22:34:35 +00:00
01317fda2a fix, chore, refactor(egen): Fix TF version, replace K80 with T4 etc. (#3143)
* chore, fix, refactor: template fixes, replace K80 with T4, remove IS_TESTING, reword some sentences and headings

* fix, refactor, feat, chore: replace K80 with T4, fix the TF version, template based fixes, remove local files in the clean up step

* chore: addresses review comments

---------

Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-06-26 22:08:55 +00:00
sumanvitaandGitHub 7c91b6f354 refactor, chore(egen): Refactored code according to template guidelines, added code to delete job in cleanup section (#3139)
* refactor,chore(egen): refacted the code as per template guidelines, correct/reword some sentences, future tense removal

* refactor(egen): adds -p {PROJECT_ID} while creating bucket

* clears cell outputs

* refactor, chore(egen): removes extra spaces, changes does not to does'nt, performed linter test
2024-06-26 21:55:04 +00:00
3cff78caf0 refactor, chore, fix(egen): replace K80 with T4, template fixes (#3137)
* fix,refractor,chore(egen): refracted code according to template guidelines, performed linter testfixed and refactored notebook according to template

* refactor(egen): replaced the project_id with [you-project-id] according to template guidelines

* chore: template guideline fixes

* chore: corrects/rewords some sentences

* chore: minor markdown fixes

* chore(Egen):Done changes according to @kittyabs review and performed linter test.

* performed linter test

---------

Co-authored-by: sriramya2610 <sriramya.peddapally@egen.ai>
Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-06-26 21:39:13 +00:00
a1bbe56d90 Featurestore notebook (#3152)
* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

added colab enterprise logo and link

* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

hope I fixed the JSON issue

* Update online_feature_serving_and_vector_retrieval_bigquery_data_with_feature_store.ipynb

Made numerous edit and fixed "pantheon" link.

* fix: remove new changes

---------

Co-authored-by: Katie Nguyen <21978337+katiemn@users.noreply.github.com>
2024-06-26 21:24:17 +00:00
39b48e8ef5 refactor(egen): Removes boilerplate, heading fixes, corrections from template. (#3126)
* <refactor> refactored notebook according to template

* Apply edits suggested by @kittyabs

* Apply suggested edits from @kittyabs review

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-06-26 21:22:34 +00:00
Eric DongandGitHub a992a7b185 Update README.md 6 (#3153)
Add more examples
2024-06-26 20:49:11 +00:00
Huguens JeanandGitHub 65baa52c6c [MG Model Team] Add checks to train and rendering job buttons in ZipNeRF gradio app. (#3133)
* [MG Model Team] Add checks to train and rendering job buttons in ZipNeRF gradio notebook.

* Clear output of all cells.

* Validate scene name in colmap workshop.
2024-06-26 12:54:36 +00:00
4615216bc2 Update training docker tag (#3145)
Co-authored-by: minwoopark <minwoopark@google.com>
2024-06-26 12:42:24 +00:00
Kaushik KoiladaandGitHub 5b3510a634 refactor,fix,chore: rectifies the egen_fix/sdk_pytorch_torchrun_custom_container_training_imagenet notebook (#3142)
* chore: adds colab enterprise and updates styling for all the open in tabs

* chore: removes boilerplate and updates according to template

* refactor: changes Region to Location

* test: end to end notebook testing

* chore: removes wil

* end to end run successful. clearing outputs

* chore: lint test done

* chore: rearrage cell
2024-06-26 00:17:52 +00:00
nileshspringmlandGitHub 550c4208f3 refactor, chore(egen): Removes boilerplate, heading fixes, and other corrections from template (#3124)
* Adapted code with updated template

* Tested end to end code and added required code

* Worked on the feedback given on PR

* Did required changes based on feedback given on PR
2024-06-26 00:15:55 +00:00
Kaushik KoiladaandGitHub cb149b0afd refactor,fix,chore: fixes error and refactors sdk_vector_search_create_stack_overflow_embeddings_vertex notebook according to template (#3108)
* chore: lint test

* chore, fix: lint run and fixes REGION issue
2024-06-26 00:13:42 +00:00
Rohith AllaandGitHub 5bced71ba5 refactore, chore (egen): Refactored code according to the notebook template (#3081)
* refactore, chore (egen): Refactored code according to the notebook template, performed linter test

* refactor, chore(egen): Reverted changes regarding IS_TESTING, performed linter test

* refactor, chore (egen): Removed reference to Tensorboard billing since it is not true anymore, performed linter test

* refactor, chore (egen): Rectified project ID, performed linter test

* refactor: Rectified service account variable

* chore: Performed linter test

* refactor, chore(egen): Updated comments in cleaning up section, performed linter test
2024-06-26 00:12:01 +00:00
Rohith AllaandGitHub 8410630e94 refactor, chore(egen): Refactored code according to template guidelines, performed linter test (#3134) 2024-06-25 17:24:03 +00:00
Eric DongandGitHub 1980ec9007 Update README.md 5 (#3118)
Add examples section
2024-06-24 23:00:24 +00:00
Kelsi LakeyandGitHub d8c8049d7e Remove incorrect pricing information about Vertex AI Tensorboard (#3057)
* Update comparing_local_trained_models.ipynb

Remove note stating Vertex TensorBoard is $300/month.
Update Create Tensorboard section to use default Tensorboard and init() method.

* Update comparing_local_trained_models.ipynb

* Update comparing_local_trained_models.ipynb

Update delete tensorboard section

* Update comparing_local_trained_models.ipynb
2024-06-24 13:13:51 +00:00
Aiden010200andGitHub 9c2cc6d39f Upload asynchronous prediction sample. (#3122)
* Upload examples of kfp v2

* Upload run experiment example.

* Upload batch prediction job sample.

* Update recycling of computing resources

Recycling computing resources after predictions.

* Upload missing file

Add delete endpoint func to recycle resources.

* Upload asynchronous prediction sample.

Upload asynchronous prediction sample of kfp v2.
2024-06-24 13:12:54 +00:00
41d482d65d Support A100-80GB for checking quota (#3119)
Co-authored-by: minwoopark <minwoopark@google.com>
2024-06-24 13:11:19 +00:00
7132270fb2 chore, feat(egen): template fixes, adds clean up steps (#3117)
* chore, feat: template fixes, adds steps to remove training job, local files, and remove future tense and contract long words

* fix: replaces REGION with LOCATION while creating the bucket

* refactor: reduces the budget_milli_node_hours to 1000 from 8000

---------

Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-06-24 13:10:40 +00:00
64aa5a5c69 refactor, chore(egen): Replaces K80 with T4 GPUs, corrections from template. (#3115)
* <refactor, chore> refactored notebook according to template

* indentation fix

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-06-24 13:09:38 +00:00
nileshspringmlandGitHub 787e43df1b refactor, chore(egen): Removes boilerplate, heading fixes, and other corrections from template (#3114)
* Adapted notebook with new notebook template and removed unwated code.

* Did major changes in dockerfile

* Did minor changes in cleanup section
2024-06-24 13:08:42 +00:00
02bdc8e3cb fix, chore, refactor, feat(egen): Replaces K80 to T4, gcr with Artifact Registry etc. (#3111)
* fix, chore, refactor, feat: replaces K80 to T4, replaces gcr pushes to artifact registry pushes, template based fixes, clean up for local files, remove IS_TESTING

* fix: defines the IS_COLAB step before the condition

* fix: REGION is replaced by LOCATION while creating the artifact registry

---------

Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-06-24 13:07:48 +00:00
f4a703ea23 refactor, chore(egen): Fix broken links in Markdown cells, corrections from template. (#3110)
* <refactor, chore> refactored notebook according to template

* Apply suggested edits from @kittyabs review

* lint fix

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-06-24 13:05:29 +00:00
Rohith AllaandGitHub 30b02b94f5 refactor, chore(egen): Replaced TESLA_V100 with TESLA_T4 and refactored the notebook with template notebook (#3107)
* refactor, chore(egen): Replaced TESLA_V100 with TESLA_T4 and refactored the notebook with updated template, performed linter test

* Refactor, chore(egen): Made few markdown changes,REGION > LOCATION, performed linter test
2024-06-24 13:04:16 +00:00
nileshspringmlandGitHub e87e3c39be refactor, chore(egen): Removes boilerplate, heading fixes, and other corrections from template (#3094)
* Updated notebook template

* Code fix related to variable TIMESTAMP, added project set for colab

* Updated tensorflow library and exuection of prediction code in VPC network

* Perfomred lint test of code

* Did minro changes regarding os.getenv

* Fixed issue based on feedback on PR
2024-06-24 13:03:03 +00:00
220382d41f fix, refactor, chore(egen): Fix model deployment to endpoint using Deployment Resource Pool, Replaces K80 with T4 GPUs, Documentation link update, corrections from template. (#3093)
* <fix, chore, refactor> refactored notebook according to template

* lint fix

* Apply suggested edits from @kittyabs review

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-06-24 13:01:47 +00:00
nileshspringmlandGitHub 2347cc4b7d refactor, chore(egen): Removes boilerplate, heading fixes, and other corrections from template (#3101)
* Did required changes in prediction/llm_streaming_prediction.ipynb. Additionally, note that code is not executed as it may requried huge resources.

* Removed hardcoded project name

* Updated REGION to LOCATION

* Updated notebook template

* Removed unwated library and addded required variable to delete resources

* Removed unwated comment as it was creating issue with linter test
2024-06-24 12:59:38 +00:00
dc07492d8d Support A100 80GB (#3129)
Co-authored-by: minwoopark <minwoopark@google.com>
2024-06-24 12:58:45 +00:00
ebd738e863 fix, refactor, chore(egen): Fix TF version support, template format (#3120)
* refactor(egen): Refactored code according to the template guidelines

* fix, refactor: fixes the compatible tf version and removes unnecessary import in the clean up step

* fix: adds matplotlib in the installation step

---------

Co-authored-by: rohith-egen <rohith.alla@egen.ai>
Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-06-21 21:11:43 +00:00
Huguens JeanandGitHub 19973ced06 [MG Model Team] Add CamP ZipNeRF Gradio application notebook to model garden. (#3121)
* [MG Model Team] Add CamP ZipNeRF Gradio application notebook to model garden.

* [MG Model Team] Add CamP ZipNeRF Gradio application notebook to model garden.

* [MG Model Team] Add CamP ZipNeRF Gradio application notebook to model garden.

* [MG Model Team] Add CamP ZipNeRF Gradio application notebook to model garden.
2024-06-21 21:10:41 +00:00
KCFindstrandGitHub bc2960fcce Update the server port in model_garden_keras_stable_diffusion.ipynb (#3128) 2024-06-21 21:08:23 +00:00
lee1premiumandGitHub 421ec0dc26 feat: Use pipeline_job_name from a tuning result. (#3127)
* feat: Use pipeline_job_name.

* feat: Use pipeline_job_name.
2024-06-21 21:07:58 +00:00
8eb2aa93d5 Add new model Claude 3.5 Sonnet and update regions for other models (#3112)
* add new model and update region for others

* fix minor error

---------

Co-authored-by: Huy Ngo <huyngo@google.com>
2024-06-20 18:22:34 +00:00
sefgsefgandGitHub 3947a8bc24 Upload torch transformers predictor sample (#3087)
This sample uses the aiplatform SDK and torch library to implement transformers predictor.
2024-06-20 12:39:50 +00:00
Eric DongandGitHub 022e1c8ee7 Update README.md 4 (#3104)
Add Get started section
2024-06-18 21:47:37 +00:00
cec3f9dd55 refactor: refactor llama3 deployment nb (#3100)
Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2024-06-18 12:36:46 +00:00
70ebb04f06 Update llama3 finetuning notebook to use 4 A100s instead of 8 (#3105)
Co-authored-by: minwoopark <minwoopark@google.com>
2024-06-18 12:31:42 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
fd21267d70 chore(deps): bump scikit-learn (#3106)
Bumps [scikit-learn](https://github.com/scikit-learn/scikit-learn) from 1.3.2 to 1.5.0.
- [Release notes](https://github.com/scikit-learn/scikit-learn/releases)
- [Commits](https://github.com/scikit-learn/scikit-learn/compare/1.3.2...1.5.0)

---
updated-dependencies:
- dependency-name: scikit-learn
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2024-06-18 12:31:00 +00:00
Eric DongandGitHub 77e0230eb8 Update README.md 3 (#3103)
Add usage section
2024-06-17 14:51:12 +00:00
8895075d28 chore: renames 'matching_engine' to 'vector_search' and 'Matching Engine' to 'Vector Search' (#3096)
Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-06-17 12:40:24 +00:00
Katie NguyenandGitHub f381ae9817 feat: add colab enterprise link format to script (#3097) 2024-06-17 12:37:28 +00:00
Mend RenovateandGitHub 0b623038fa chore(deps): update dependency flake8 to v7.1.0 (#3099) 2024-06-17 12:36:41 +00:00
7cd354436e chore, refactor, fix, feat(egen): template structure+K80GPU+contract words+clean up (#3085)
* chore, refactor, fix, feat: template structure, removes redundant code, contracts content, replaces K80 with T4, cleans up local files

* chore,chore, fix: addresses review comments+ sets replica=1,accelerator_count=4

* fix, chore: reduces the GPU count to 1, rewords the title

---------

Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-06-14 21:12:17 +00:00
Eric DongandGitHub d7c334353c Update README.md 2 (#3091)
Continued to update README.
2024-06-14 17:10:37 +00:00
kittyabsandGitHub ba80e59b1e Update tensorboard_vertex_ai_pipelines_integration.ipynb (#3036)
* Update tensorboard_vertex_ai_pipelines_integration.ipynb

added colab enterprise link and icon. Deleted <br> line, too

* Update tensorboard_vertex_ai_pipelines_integration.ipynb

fixed - added missing .ipynb
2024-06-14 14:23:43 +00:00
62b3f0af1b Update llama3 finetuning notebook (#3090)
* Update llama3 finetuning notebook

* Update model_garden_pytorch_llama3_finetuning.ipynb

fix linter issue

* Update model_garden_pytorch_llama3_finetuning.ipynb

Remove unnecessary metadata

---------

Co-authored-by: minwoopark <minwoopark@google.com>
2024-06-13 23:54:58 +00:00
bd10dede25 refactor(egen): follows new template+sentence and other minor corrections (#3073)
* chore: restructures according to template, contracts text and cells, organizes headings

* refactor: removes the IS_TESTING conditions for steps involving redis instance

* chore: addresses review comments

* fix: adds back the IS_TESTING conditions for redis commands to skip in the test environment

* chore: removes duplicate comment

---------

Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-06-13 20:52:43 +00:00
7e8bf8a126 fix, refactor, chore(egen): follows new template, adds dataflow enable step, refactor (#3089)
* Updated template of notebook

* fix, refactor, chore: follows new template, adds dataflow enable step, cleanup steps for files, removes IS_TESTING in cleanup, REGION==>LOCATION

* chore: addresses the review comments

---------

Co-authored-by: nileshspringml <nilesh.mahajan@egen.ai>
Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-06-13 20:36:50 +00:00
nileshspringmlandGitHub 847d38642c refactor, chore(egen): Removes boilerplate, heading fixes, and other corrections from template (#3086)
* Did required changes in notebook template

* Service account permission changed as we don't need admin level access for this notebook

* Fixed issue based on PR feedback
2024-06-13 20:31:07 +00:00
Gary WeiandGitHub 697a4c6e88 Add dreambooth-lora-sdxl task for SDXL base model in the dreambooth finetune Gradio notebook. (#3082)
* Add controlnet-canny to the Gradio playground, and some additional UX enhancement.

* Minor fixes.

* Minor fixes

* Add additional document regarding the list of supported models, and some UI enhancement.

* Minor update to the hyperlink.

* Rewrite the SD2.1 dreambooth finetune notebook.

* Add code owners.

* Some minor changes to the stable diffusion 2.1 and sd-xl notebooks.

* some additional minor fixes.

* additional fixes.

* Create a notebook to demonstrate dreambooth LoRA finetune for SD-XL model.

* minor updates

* add to the codeowner list.

* merge conflict.

* minor fix to the Gradio UI workshop notebook.

* Some minor updates to the SD2.1 deployment notebook.

* Minor update the `sd-xl` deployment notebook, based on the QA feedback.

* Add a few community models to the Gradio workshop.

* Switch `mediapipe-train` docker container from `vertex-ai-restricted` to `vertex-ai`, in the `mediapipe-train` notebooks.

* Create a notebook for model `instantx/instantid`.

* Update Gradio notebook to use the latest Gradio version and fix some bugs.

1. Update Gradio version to 4.29.0, as it complains 3.50.0 is too old.
2. Uninstall nest-asyncio and uvloop as a workaround to b/339301920 and https://github.com/gradio-app/gradio/issues/8238#issuecomment-2101066984.

* Resolve merge conflict.

* minor updates.

* minor updates.

* Merge some SD notebook in g3 and github.

* Remove the unused variable in the controlnet notebook.

* minor updates.

* include the SD1.5 dreambooth notebook.

* Include the sd1.5 dreambooth notebook.

* Improve the stable diffusion dreambooth tuning CUJ in the Gradio notebook.

* minor update.

* Add dreambooth-lora-sdxl task for SDXL base model in the dreambooth finetune Gradio notebook.
2024-06-13 14:01:39 +00:00
f5c17c0700 refactor, chore, feat(egen): contracts content, steps & template based fixes (#3075)
* refactor, chore, feat: removes/contracts long explanations, template based fixes, adds step to remove locally generated files

* fix: adds the missing import, replaces code markdown with bold style at some places

* fix: replaces REGION with LOCATION

* chore: addresses review comments

---------

Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-06-12 22:37:30 +00:00
0e662eca2b refactor, chore(egen): Uses gcloud builds for building and pushing image to artifact registry, Removes boilerplate, heading fixes, and other corrections from template (#3074)
* <refactor>: refactored code according to notebook template

* <refactor>: refactored code according to notebook template

* <refactor>: refactored code according to notebook template

* <refactor,chore> refactored notebook according to template

* <refactor,chore> refactored notebook according to template

* fix for  docker repository creation in PR test environment

* <included IS_TESTING condition for docker repository

* lint fix

* Apply suggested edits from @kittyabs review

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-06-12 17:28:01 +00:00
Eric DongandGitHub d339690a4b Update README.md (#3088)
Update the overview with link to generative-ai repo
2024-06-12 17:08:54 +00:00
Mend RenovateandGitHub 134a6221e6 chore(deps): update dependency pyupgrade to v3.16.0 (#3066) 2024-06-11 19:14:25 +00:00
kittyabsandGitHub e79c4b5caa Update get_started_with_pytorch_rov.ipynb (#3034)
Update colab enterprise logo
2024-06-11 19:13:18 +00:00
Kathy YuandGitHub 535a6232d3 Remove legacy Llama 2 notebook. (#3083) 2024-06-11 19:11:11 +00:00
KCFindstrandGitHub f49e09698e Fix broken links in model garden notebooks (#3077)
* Fix broken links in notebooks

PiperOrigin-RevId: 641077957

* Revert llama3 deployment notebook change
2024-06-11 19:10:16 +00:00
Gary WeiandGitHub 792e3b16a5 Update the diffusers serving docker image to 20240605_1400_RC00 to resolve vulnerabilities. (#3079)
* Create a Gradio notebook for the new InstantId model.

* Add dreambooth finetune to the stable diffusion Gradio workshop notebook.

* Update the image generation Gradio notebook to support Dreambooth finetuning.

* linter update

* linter update

* minor fix to the instant-id notebook.

* Minor fix to the stable diffusion gradio notebook.

* Split the 'instant-id' deployment notebook prediction into two sections.

* add `deployment_source` to the notebook.

* Switch to `pytorch-diffusers-serve-opt` container to for diffusion lora serving.

* add the dreambooth_lora notebook.

* minor update.

* Parameterize the "show_debug_logs" to facilitate automatic test of the Gradio notebooks.

* Lint format.

* minor updates

* Delete the two deprecated SD1.5 and 2.1 notebooks, as they were no longer referenced on any model cards.

* Sync Colab notebooks between g3 and github.

* format changes

* format update.

* Improve the SDXL-dreambooth-lora finetune notebook CUJ.

* Update the diffusers serving docker image version to `20240605_1400_RC00` to resolve vulnerabilities.

* Add dreambooth-lora-sdxl task for SDXL base model in the dreambooth finetune Gradio notebook.

* Add dreambooth-lora-sdxl task for SDXL base model in the dreambooth finetune Gradio notebook.

* Add dreambooth-lora-sdxl task for SDXL base model in the dreambooth finetune Gradio notebook.
2024-06-11 19:09:30 +00:00
Kaushik KoiladaandGitHub fbcf3b4cc8 chore, refactor(Egen): restructure from template, remove boilerplate, and fix notebook (#3068)
* chore: updates copyright text, adds colab enterprise and for

* chore: removes boilerplate and changes colab authentication and get started section

* refactor: removes IS_TESTING and other test code

* refactor: removes IS_TESTING and other test code

* chore: clear all outputs and linter reformatting

* chore: removes code added for testing and other fixes

* chore: runs lint

* fix: fixes testing induced bug

* chore: review comments with header and copyright changes addressed
2024-06-11 18:27:36 +00:00
ac8f7ad196 chore, fix, feat(egen): follows template + fix docker run + adds clean up (#3072)
* chore, fix: restructures as per template, rewords some sentences, removes sudo in the option docker run

* feat: adds clean up step for local files

* chore: reverts lowercase to camelcase for artifact registry and other review comments

---------

Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-06-11 18:06:50 +00:00
Kaushik KoiladaandGitHub 4e21510eeb refactor, chore, fix (Egen) : removes boilerplate and fixes notebook pytorch_train_deploy_models_with_prebuilt_containers.ipynotebook (#3070)
* chore: updates the license

* chore: formats run in buttons and adds colab enterprise link

* chore: removes code font for names of products

* refactor: removes boilerplate and edits aiplatform initialization cell order

* fix: changes naming of app to deal with setup changing tar.gz file to canonical name and gsutil command not able to find it

* chore: linter run

* chore: edit all the future tense sentences

* chore: updates review comments
2024-06-11 18:01:11 +00:00
Rohith AllaandGitHub 04059ebadb refactor(egen): Refactored the Custom tabular bq managed dataset code as per template guidelines (#3053)
* refactor: Refactored custom-tabular-bq-managed-dataset code according to the template notebook

* refactor: Included markdown code in the beginning of the license cell to make it collapsible

* chore: linter test

* refactor, chore: Updated markdown according to the updated notebook template and performed linter test

* chore: Grouped all the imports present in the notebook

* chore: linter test

* refactor, chore(egen): REGION is replaced with LOCATION, linter test
2024-06-11 17:53:12 +00:00
Kaushik KoiladaandGitHub acb4fbb684 refactor, fix, chore(Egen): changes and refactors the get_started_vertex_training_xgboost.ipynb notebook (#3069)
* chore: updates license information, adds colab enterprise, and reformats run buttons

* chore: 'will' replaced appropriately

* chore: updates region and removes boilerplate

* fix: corrects the argument at pip install

* refactor: refactoring the cells for end to end functionality

* chore: changes verbiage in bucket creation

* chore: linter test done

* refactor: rearrage cell order for aiplatform initialization

* chore: review comments addressed

* chore: review comments addressed
2024-06-11 17:48:44 +00:00
Rohith AllaandGitHub 32934925f6 refactor(egen): Refactored custom batch prediction feature filter code according to template guidelines (#3052)
* refactor: Refactored custom_batch_prediction_feature_filter code according to the template notebook

* refactor: Updated markdown according to the updated notebook template and did linter test

* chore: Modified code according to updated template

* chore: linter test

* refactor, chore(egen): REGION is replaced with LOCATION, linter test
2024-06-11 17:40:16 +00:00
Rohith AllaandGitHub daab623cac refactor, chore: Updated markdown according to the updated notebook template and performed linter test (#3063) 2024-06-07 23:12:01 +00:00
Rohith AllaandGitHub dad35742de fix, refactor(egen): Replaced tf2-gpu.2-5 with tf2-gpu.2-6, Refactored the code as per template guidelines (#3064)
* refactor: Updated markdown according to the updated notebook template

* refactor, chore: Added colab enterprise logo url and performed  linter test

* refactor, chore: Reverted back the max trail count and parallel trail count values, performed linter test
2024-06-07 21:30:06 +00:00
Rohith AllaandGitHub f7b49ebb40 refactor(egen): Refactored the Autosxs check alignment against human preference data code as per template guidelines (#3062)
* refactor, chore: Updated markdown according to the updated notebook template and performed linter test

* Revert "refactor, chore: Updated markdown according to the updated notebook template and performed linter test"

This reverts commit c91fa09431.

* refactor, chore: Updated markdown according to the updated notebook template and performed linter test
2024-06-07 21:24:53 +00:00
nileshspringmlandGitHub 95ffea1915 refactor, chore(egen): Removes boilerplate, heading fixes, and other corrections from template (#3058)
* Did all the required changes in notebook

* Added below comments:
# @title Copyright & License (click to expand)

* Upadated the template of notebook

* Fixed issue based on the feedback given on PR
2024-06-07 21:20:07 +00:00
nileshspringmlandGitHub 7e6122a50d refactor, chore(egen): Removes boilerplate, heading fixes, and other corrections from template (#3046)
* refactor, chore(egen): Removes boilerplate, heading fixes, and other corrections from template

* Added Below line in comments:
# @title Copyright & License (click to expand)

* Fixed issue in notebook based on feedback given by reviewer in PR

* Fixed issue based on feedback given on PR
2024-06-07 21:15:20 +00:00
7afdeb05f5 chore, fix(egen): corrections from template, guidelines, & fixes docker tag command (#3031)
* chore: replaces K80 with T4, Cloud ML with Vertex AI, REGION with LOCATION, restructures from template, & removes future tense

* fix: fixes the docker tag command for colab

* chore, fix: addresses the review comments, tf is pinned to 2.15.1 as the latest tf causes issues

* chore: replaces of with or

* chore: removes the collapsed license comment

* fix: removes the test env specific package update step and contracts the installations into one step to keep tf as 2.15

---------

Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-06-07 21:10:09 +00:00
2df29ec118 chore, fix(egen): restructuring from template, markdown corrections & dataset object fix (#3028)
* chore: restructures according to the template, removes future tense, removes IS_TESTING, unused imports

* chore, fix: addresses the review comments, converts npy array to list for running in py-3.9

* chore: adds --it's-- in the sentence

* chore: Getting started -->  Get started

* chore: removes the collapsed license comment

* fix: extracts list from numpy objects instead of Dataset object

---------

Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-06-07 20:58:54 +00:00
nileshspringmlandGitHub c2cd389201 refactor, chore(egen): Removes boilerplate, heading fixes, and other corrections from template (#3020)
* Completed code fixing for file prediction/pytorch_image_classification_with_prebuilt_serving_containers.ipynb. Please note that we need to execute file end to end.

* Added Colab Enterprice link and testead once again with different environment as previously it was throwing an expecption of  libraries

* Added below comment:
# @title Copyright & License (click to expand)

* Fixed issue in notebook based on feedback given by reviewer in PR

* Fixed issue based on feedback give on PR
2024-06-07 20:54:10 +00:00
nileshspringmlandGitHub b30a230d1d refactor, chore(egen): Removes boilerplate, heading fixes, and other corrections from template (#3019)
* Fix code of file prediction/get_started_with_raw_predict.ipynb

* Added link of Colab Enterprise

* Updated create bucket command because it requried those changes to execute in google colab notebook

* Removed IS_TESTING environment variable

* Added below comment in notebook:
# @title Copyright & License (click to expand)

* Fixed issue in notebook based on feedback given by reviewer in PR

* Fixed issue based on the feedback given on PR
2024-06-07 20:51:24 +00:00
Aiden010200andGitHub c5be708d45 Update func of computing resources recycling (#3043)
* Upload examples of kfp v2

* Upload run experiment example.

* Upload batch prediction job sample.

* Update recycling of computing resources

Recycling computing resources after predictions.

* Upload missing file

Add delete endpoint func to recycle resources.
2024-06-07 12:39:33 +00:00
Gary WeiandGitHub bc577d9f79 Improve the SDXL-dreambooth-lora finetune notebook CUJ. (#3055)
* Create a Gradio notebook for the new InstantId model.

* Add dreambooth finetune to the stable diffusion Gradio workshop notebook.

* Update the image generation Gradio notebook to support Dreambooth finetuning.

* linter update

* linter update

* minor fix to the instant-id notebook.

* Minor fix to the stable diffusion gradio notebook.

* Split the 'instant-id' deployment notebook prediction into two sections.

* add `deployment_source` to the notebook.

* Switch to `pytorch-diffusers-serve-opt` container to for diffusion lora serving.

* add the dreambooth_lora notebook.

* minor update.

* Parameterize the "show_debug_logs" to facilitate automatic test of the Gradio notebooks.

* Lint format.

* minor updates

* Delete the two deprecated SD1.5 and 2.1 notebooks, as they were no longer referenced on any model cards.

* Sync Colab notebooks between g3 and github.

* format changes

* format update.

* Improve the SDXL-dreambooth-lora finetune notebook CUJ.
2024-06-07 12:36:14 +00:00
KCFindstrandGitHub 1fcafd3364 Fix accelerator_type variable name case in quota checks (#3061)
PiperOrigin-RevId: 640634353
2024-06-07 12:34:21 +00:00
be31756dac refactor, chore(egen): Claude documentation link update, Text explanations for required cells and other corrections from template (#3027)
* <refactor> refactored notebook according to template

* <refactor> refactored notebook according to template

* <refactor> refactored notebook according to template

* <refactor>: refactored notebook according to new notebook template

* <refactor>: refactored notebook according to new notebook template

* <refactor>: refactored notebook according to new notebook template

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-06-06 23:03:34 +00:00
868e49224b refactor(egen): Removes boilerplate, heading fixes, and other corrections from template (#3026)
* <refactor>: refactored code according to notebook template

* <refactor>: refactored code according to notebook template

* <refactor>: refactored code according to notebook template

* <refactor>: refactored code according to notebook template

* <refactor>: refactored code according to new notebook template

* <refactor>: refactored code according to new notebook template

* <refactor>: refactored code according to new notebook template

* <refactor>: refactored notebook according to new notebook template

* <refactor>: refactored notebook according to new notebook template

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-06-06 23:01:07 +00:00
nileshspringmlandGitHub 32caaaa7d9 Updated notebook template (#3059)
* Updated notebook template based on google team comment on PR

* Removed unwated comment line from notebook
2024-06-06 13:12:57 +00:00
Rohith AllaandGitHub 521f58b7e6 refactor(egen): Refactored Text embedding api semantic search with scann code according to template guidelines (#3024)
* <refactor>: Refactored text_embedding_api_semantic_search_with_scann code according to the notebook template

* chore: linter test

* refactor: Modified colab enterprise logo url and made necessary changes according to the updated notebook template

* chore: linter test
2024-06-05 20:59:54 +00:00
030024cc51 refactor, chore, fix(egen): GPU usage fixes, template & markdown fixes (#3050)
* refactor, chore, fix: replaces K80 with T4, fixes to follow the template, fixes to utilize the defined accelerators while training

* chore: addresses the review comments

---------

Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-06-05 20:51:27 +00:00
Rohith AllaandGitHub dc38b9e0f9 refactor(egen): Refactored text embedding new api code according to the template guidelines (#3051)
* refactor: Refactored text_embedding_new_api code according to the template notebook

* refactor: Modified colab enterprise logo url and made necessary changes according to the updated notebook template
2024-06-05 20:12:57 +00:00
Amy WuandGitHub cd95c35cc2 remove data science package samples since they are deprecated (#3056) 2024-06-04 23:49:20 +00:00
4e04f7d166 Template vertex ai (#3054)
* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

added colab enterprise logo and link

* Update build_model_experimentation_lineage_with_prebuild_code.ipynb

hope I fixed the JSON issue

* Update notebook_template.ipynb

- changed "Getting Started" to "Get started" to be in compliance with style guide
- added "for Python" to "Vertex AI SDK" to be in compliance with product guidelines

* fix: reset changes

---------

Co-authored-by: Katie Nguyen <21978337+katiemn@users.noreply.github.com>
2024-06-04 18:02:17 +00:00
bc14b3fa05 refactor, chore(egen): Removes boilerplate, heading fixes, and other corrections from template (#3045)
* <refactor>: refactored code according to notebook template

* <refactor> refactored notebook according to template

---------

Co-authored-by: SumanthKasula99 <sumanth.kasula@egen.ai>
2024-06-04 08:20:36 +00:00
db9c182716 refactor, chore(egen): Removes boilerplate, heading fixes, and other corrections from template (#3021)
* refactor: removes boilerplate code, future tenses, and fixes heading styles

* chore: replaces REGION with LOCATION to match the notebook template

* chore: markdown heading fixes according to guidelines

---------

Co-authored-by: rohith-egen <rohith.alla@egen.ai>
Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-06-03 23:48:13 +00:00
Rohith AllaandGitHub c240cf6a34 refactor(egen): Refactored the code as per template guidelines (#3015)
* refactored code according to template notebook

* refactored code according to template notebook
2024-06-03 23:46:29 +00:00
lee1premiumandGitHub 78d2d46f80 fix: Comment on Vertex AI workbench. (#3049)
* fix: Comment on Vertex AI workbench.

* fix: Comment on Vertex AI workbench.
2024-06-03 20:04:53 +00:00
54ae9a17b2 chore(egen): removes future tense, adds colab enterprise link, and colab only steps (#3025)
* chore: updates the notebook according to the latest template

* chore: linter test

* refactor: removes IS_TESTING, and os import

* ran linter test using linter.sh

---------

Co-authored-by: rohith-egen <rohith.alla@egen.ai>
Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-06-03 18:24:18 +00:00
kittyabsandGitHub 9251100d25 Update get_started_with_model_registry.ipynb (#3047)
updated icon url: https://cloud.google.com/ml-engine/images/colab-enterprise-logo-32px.png
2024-06-03 18:05:06 +00:00
eccecc8e2b Update ray_cluster_management.ipynb (#3033)
* Update ray_cluster_management.ipynb

Updated logo for Colab Enterprise

* fix: change accelerator type

* fix: alter accelerator type

---------

Co-authored-by: Katie Nguyen <21978337+katiemn@users.noreply.github.com>
2024-06-01 20:48:18 +00:00
263dba54bc fix, refactor(egen): Replaces K80 with T4 GPUs, corrections from template (#3013)
* fix: n1-standard-8 changed to n1-standard-16 and Tesla K80 changed to Tesla T4 + refactored code according to the template

* refactor: keeping the machine types same. Original issue with the K80 accelerators as they're no longer supported

---------

Co-authored-by: rohith-egen <rohith.alla@egen.ai>
Co-authored-by: krishr2d2 <krishna.movva@egen.ai>
2024-06-01 00:15:24 +00:00
86243a1cb7 chore: Refactor the instructions in model_monitoring_v2 notebooks (#3023)
Co-authored-by: SereniCode <binbinf@google.com>
2024-05-31 23:40:23 +00:00
kittyabsandGitHub dd7c60acff Update tensorboard_hyperparameter_tuning_with_hparams.ipynb (#3037)
Add colab enterprise link and logo
2024-05-31 23:24:01 +00:00
kittyabsandGitHub 6901cd24e5 Update tensorboard_profiler_custom_training_with_prebuilt_container.ipynb (#3038)
added colab enterprise link and logo
2024-05-31 23:21:50 +00:00
kittyabsandGitHub 3dae9238f5 Update notebook_template.ipynb (#3029)
update enterprise logo
2024-05-31 23:15:54 +00:00
kittyabsandGitHub 4e8bfb88b3 Update model_garden_gemma_fine_tuning_batch_deployment_on_rov.ipynb (#3032)
update colab enterprise icon
2024-05-31 23:14:54 +00:00
kittyabsandGitHub b6b921ef68 Experiments delete outdated experiments (#3041)
* Update tensorboard_custom_training_with_custom_container.ipynb

added colab enterprise link and logo

* Update comparing_local_trained_models.ipynb

Added colab enterprise logo and link. Also made some edits.

* Update delete_outdated_tensorboard_experiments.ipynb

Updated Colab Enterprise logo
2024-05-31 23:01:12 +00:00
kittyabsandGitHub faf68b59bb Update delete_outdated_tensorboard_experiments.ipynb (#3042)
updated colab enterprise logo
2024-05-31 22:57:11 +00:00
f9351c0e0b Update get_started_with_model_registry.ipynb (#3030)
* Update get_started_with_model_registry.ipynb

add colabe enterprise link

* fix: remove extra line break elements

---------

Co-authored-by: Katie Nguyen <21978337+katiemn@users.noreply.github.com>
2024-05-31 22:51:24 +00:00
kittyabsandGitHub 921b1be406 Update get_started_with_model_registry.ipynb (#3022)
small edits.
2024-05-30 20:01:30 +00:00
Yichen ZhouandGitHub a0ed217801 Create TimesFM notebook for Vertex Model Garden (#3017)
* feat: Create TimesFM notebook.

* fix: updated CODEOWNER of TimesFM notebook

* fix: remove empty line in CODEOWNERS

* Removed unused import inside TimesFM notebook.

* fix: remove unused import and re-format
2024-05-30 18:50:16 +00:00
Tianrui YangandGitHub 8bed394caf fea: upgrade to v1 API in feature store llm grounding tutorial. (#3009)
* Upgrade to v1 API in feature store llm grounding tutorial.

* Sleep for 5min before starting serving to wait for DNS to be ready.

* use data_key in fetch request
2024-05-30 12:44:30 +00:00
c306cfaae1 Add common util functions for notebooks. (#3014)
Co-authored-by: minwoopark <minwoopark@google.com>
2024-05-29 15:23:59 +00:00
fe97c65ea4 Update Gemma and PaliGemma notebooks (#3007)
Co-authored-by: minwoopark <minwoopark@google.com>
2024-05-28 19:35:32 +00:00
Katie NguyenandGitHub 9d0ec7cc62 fix: update colab enterprise link (#3012) 2024-05-28 19:33:43 +00:00
de5146905c fix declare -x AUTO_PROXY="https://proxyconfig.corp.google.com/proxy.pac" (#3011)
declare -x CHROME_REMOTE_DESKTOP_DEFAULT_DESKTOP_SIZES="1600x1200,3840x2160,3840x2560,5120x1440,2160x3840"
declare -x COLORTERM="truecolor"
declare -x CVS_RSH="ssh"
declare -x DBUS_SESSION_BUS_ADDRESS="unix:path=/run/user/809963/bus"
declare -x GOOGLE_CLOUD_DISABLE_DIRECT_PATH="truen"
declare -x HISTCONTROL="ignoredups"
declare -x HOME="/usr/local/google/home/minwoopark"
declare -x LANG="en_US.UTF-8"
declare -x LESSCLOSE="/usr/bin/lesspipe %s %s"
declare -x LESSOPEN="| /usr/bin/lesspipe %s"
declare -x LOGNAME="minwoopark"
declare -x LS_COLORS="rs=0:di=01;34:ln=01;36:mh=00:pi=40;33:so=01;35:do=01;35:bd=40;33;01:cd=40;33;01:or=40;31;01:mi=00:su=37;41:sg=30;43:ca=00:tw=30;42:ow=34;42:st=37;44:ex=01;32:*.tar=01;31:*.tgz=01;31:*.arc=01;31:*.arj=01;31:*.taz=01;31:*.lha=01;31:*.lz4=01;31:*.lzh=01;31:*.lzma=01;31:*.tlz=01;31:*.txz=01;31:*.tzo=01;31:*.t7z=01;31:*.zip=01;31:*.z=01;31:*.dz=01;31:*.gz=01;31:*.lrz=01;31:*.lz=01;31:*.lzo=01;31:*.xz=01;31:*.zst=01;31:*.tzst=01;31:*.bz2=01;31:*.bz=01;31:*.tbz=01;31:*.tbz2=01;31:*.tz=01;31:*.deb=01;31:*.rpm=01;31:*.jar=01;31:*.war=01;31:*.ear=01;31:*.sar=01;31:*.rar=01;31:*.alz=01;31:*.ace=01;31:*.zoo=01;31:*.cpio=01;31:*.7z=01;31:*.rz=01;31:*.cab=01;31:*.wim=01;31:*.swm=01;31:*.dwm=01;31:*.esd=01;31:*.avif=01;35:*.jpg=01;35:*.jpeg=01;35:*.mjpg=01;35:*.mjpeg=01;35:*.gif=01;35:*.bmp=01;35:*.pbm=01;35:*.pgm=01;35:*.ppm=01;35:*.tga=01;35:*.xbm=01;35:*.xpm=01;35:*.tif=01;35:*.tiff=01;35:*.png=01;35:*.svg=01;35:*.svgz=01;35:*.mng=01;35:*.pcx=01;35:*.mov=01;35:*.mpg=01;35:*.mpeg=01;35:*.m2v=01;35:*.mkv=01;35:*.webm=01;35:*.webp=01;35:*.ogm=01;35:*.mp4=01;35:*.m4v=01;35:*.mp4v=01;35:*.vob=01;35:*.qt=01;35:*.nuv=01;35:*.wmv=01;35:*.asf=01;35:*.rm=01;35:*.rmvb=01;35:*.flc=01;35:*.avi=01;35:*.fli=01;35:*.flv=01;35:*.gl=01;35:*.dl=01;35:*.xcf=01;35:*.xwd=01;35:*.yuv=01;35:*.cgm=01;35:*.emf=01;35:*.ogv=01;35:*.ogx=01;35:*.aac=00;36:*.au=00;36:*.flac=00;36:*.m4a=00;36:*.mid=00;36:*.midi=00;36:*.mka=00;36:*.mp3=00;36:*.mpc=00;36:*.ogg=00;36:*.ra=00;36:*.wav=00;36:*.oga=00;36:*.opus=00;36:*.spx=00;36:*.xspf=00;36:*~=00;90:*#=00;90:*.bak=00;90:*.crdownload=00;90:*.dpkg-dist=00;90:*.dpkg-new=00;90:*.dpkg-old=00;90:*.dpkg-tmp=00;90:*.old=00;90:*.orig=00;90:*.part=00;90:*.rej=00;90:*.rpmnew=00;90:*.rpmorig=00;90:*.rpmsave=00;90:*.swp=00;90:*.tmp=00;90:*.ucf-dist=00;90:*.ucf-new=00;90:*.ucf-old=00;90:"
declare -x MOTD_SHOWN="pam"
declare -x OLDPWD="/tmp/vertex-ai-samples"
declare -x P4CONFIG=".p4config"
declare -x P4MERGE="/google/src/files/head/depot/eng/perforce/mergep4.tcl"
declare -x PARINIT="rTbgqR B=.?_A_a Q=_s>|:"
declare -x PATH="/usr/local/google/home/minwoopark/.local/bin:/usr/lib/google-golang/bin:/usr/local/buildtools/java/jdk/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin"
declare -x PWD="/tmp/vertex-ai-samples/notebooks/community/model_garden"
declare -x PYTHONPATH="/usr/local/buildtools/current/sitecustomize"
declare -x RSYNC_RSH="ssh"
declare -x SHELL="/bin/bash"
declare -x SHLVL="1"
declare -x SK_SIGNING_PLUGIN="gnubbyagent"
declare -x SSH_AUTH_KEY=$'ecdsa-sha2-nistp256-cert-v01@openssh.com AAAAKGVjZHNhLXNoYTItbmlzdHAyNTYtY2VydC12MDFAb3BlbnNzaC5jb20AAAAgazEAFZ2eo797Q/VO6uG2HkkShArEPHLKfyjKs/qYq0cAAAAIbmlzdHAyNTYAAABBBJQ0tXAyF9ZYMb4PE8Ev0Qsmg94oZ0skXvnuEHewktkb5R0Iq/01UbI5EC+g4ivkVsT1BEWkfEvK5mkVBYQkKpcZAAAADCmtRAAAAAEAAAAabWlud29vcGFya0Bjb3JwLmdvb2dsZS5jb20AAAAjAAAACm1pbndvb3BhcmsAAAARZ29vZ2xlXG1pbndvb3BhcmsAAAAAZlX9cQAAAABmVxfdAAAAAAAAAMwAAAAYY2VydC1tZXRhZGF0YUBnb29nbGUuY29tAAAAKgAAACYIARIgB1wlRexvZHwhV2JNrjQ1fgtUHfbXqyGq0FjDNs8V9lUgBgAAABVwZXJtaXQtWDExLWZvcndhcmRpbmcAAAAAAAAAF3Blcm1pdC1hZ2VudC1mb3J3YXJkaW5nAAAAAAAAABZwZXJtaXQtcG9ydC1mb3J3YXJkaW5nAAAAAAAAAApwZXJtaXQtcHR5AAAAAAAAAA5wZXJtaXQtdXNlci1yYwAAAAAAAAAAAAABFwAAAAdzc2gtcnNhAAAAAwEAAQAAAQEAvN0ZS5b1OZYtoJ1PSKY4GIwjis1i4zZZ2MBdN/TEYqJIOVsfAtkDrhC9YGSVuyai/kOXwLLnFc5dVDRWHLDSBzoXEgl4QKCmNu9nneV/cMLEq4d03o1DPOSPQGJDq+wep4K9HuRwvzog6wTDA5Kp0loCnWY8MHTbt4S/O2Ro5mvF0x0ec9vccwW1KOtc/CydQiGmevBZOQOyXt8ZCZKEtSOTIPhAE55WK8agtMEsJlHRtcswSg2BJNJMSeUKgL1An/oCE9bKAME/zXVYVK5Fuv4epqccnd3sQW2T8qniOIcEDI4oybDejm6G8VPw/pxieSPbaFGftuLyR/rHS52OhwAAARQAAAAMcnNhLXNoYTItMjU2AAABAB6848edrFMkVv7srR+12gWCcYsAXadttaF9VBNSK7AojCKPQC5axHjdfmGQHkc6FDj5x76wFhW2QoBWF6EEullMA/BBYtUeOYYrRwIQxvG2EshSl6Qyx1JW9RmClgEQa13JDNS+sHaGih0TK/+h1Cll3WcGYL5bSlTjyI4QX9gcoTA12Ctd01xKiMuXTscx+IMw8nyl9hQQ4eucc4FW130gRglHBdqH9VNjPgFXnoffXfC8P7kBKhPfVk9CjNWl3VN7LwEvFXAT40EBqFeK/b+2pKTL9lDEeUkXxDCz39GHWBu8BYd186GVjjjF8hgK01I90/xGldt/cORgorzziko=\n'
declare -x SSH_AUTH_SOCK="/tmp/ssh-XXXXFPQFEs/agent.634059"
declare -x SSH_CLIENT="172.253.30.128 55315 22"
declare -x SSH_CONNECTION="172.253.30.128 55315 192.168.3.61 22"
declare -x SSH_TTY="/dev/pts/1"
declare -x TERM="xterm-256color"
declare -x USER="minwoopark"
declare -x X20_HOME="/google/data/rw/users/mi/minwoopark"
declare -x XDG_DATA_DIRS="/usr/share/gnome:/usr/local/share/:/usr/share/"
declare -x XDG_RUNTIME_DIR="/run/user/809963"
declare -x XDG_SESSION_CLASS="user"
declare -x XDG_SESSION_ID="c45"
declare -x XDG_SESSION_TYPE="tty"
declare -x service_endpoint="aiplatform.googleapis.com" error

Co-authored-by: minwoopark <minwoopark@google.com>
2024-05-28 16:41:55 +00:00
kittyabsandGitHub dd8d7946c4 Tensorboard-intro-4 (#2999)
* Update get_started_with_pytorch_rov.ipynb

This is a test

* update url to /tensorboard-introduction
2024-05-24 14:14:45 +00:00
kittyabsandGitHub 37c0c2c373 Tensorboard-intro-3 (#2997)
* Update get_started_with_pytorch_rov.ipynb

This is a test

* Update URL to point to /tensorboard-introduction
2024-05-24 14:14:13 +00:00
Tianrui YangandGitHub 974c3e866c Update embedding column type to float (#3006) 2024-05-24 14:13:20 +00:00
Kathy YuandGitHub d5b6b42615 Fix GCS bucket handling code in Mixtral deployment notebook. (#3008) 2024-05-24 14:12:49 +00:00
186e886edf feat: add model_monitoring_v2 notebooks (#2988)
Co-authored-by: SereniCode <binbinf@google.com>
2024-05-23 17:28:19 +00:00
Gary WeiandGitHub f597680d44 Improve the stable diffusion dreambooth tuning CUJ in the Gradio notebook: (#3004)
* Add controlnet-canny to the Gradio playground, and some additional UX enhancement.

* Minor fixes.

* Minor fixes

* Add additional document regarding the list of supported models, and some UI enhancement.

* Minor update to the hyperlink.

* Rewrite the SD2.1 dreambooth finetune notebook.

* Add code owners.

* Some minor changes to the stable diffusion 2.1 and sd-xl notebooks.

* some additional minor fixes.

* additional fixes.

* Create a notebook to demonstrate dreambooth LoRA finetune for SD-XL model.

* minor updates

* add to the codeowner list.

* merge conflict.

* minor fix to the Gradio UI workshop notebook.

* Some minor updates to the SD2.1 deployment notebook.

* Minor update the `sd-xl` deployment notebook, based on the QA feedback.

* Add a few community models to the Gradio workshop.

* Switch `mediapipe-train` docker container from `vertex-ai-restricted` to `vertex-ai`, in the `mediapipe-train` notebooks.

* Create a notebook for model `instantx/instantid`.

* Update Gradio notebook to use the latest Gradio version and fix some bugs.

1. Update Gradio version to 4.29.0, as it complains 3.50.0 is too old.
2. Uninstall nest-asyncio and uvloop as a workaround to b/339301920 and https://github.com/gradio-app/gradio/issues/8238#issuecomment-2101066984.

* Resolve merge conflict.

* minor updates.

* minor updates.

* Merge some SD notebook in g3 and github.

* Remove the unused variable in the controlnet notebook.

* minor updates.

* include the SD1.5 dreambooth notebook.

* Include the sd1.5 dreambooth notebook.

* Improve the stable diffusion dreambooth tuning CUJ in the Gradio notebook.

* minor update.
2024-05-23 17:10:40 +00:00
Gary WeiandGitHub 51315ecb2f sync the colab notebooks between g3 and github. (#3003)
* Create a Gradio notebook for the new InstantId model.

* Add dreambooth finetune to the stable diffusion Gradio workshop notebook.

* Update the image generation Gradio notebook to support Dreambooth finetuning.

* linter update

* linter update

* minor fix to the instant-id notebook.

* Minor fix to the stable diffusion gradio notebook.

* Split the 'instant-id' deployment notebook prediction into two sections.

* add `deployment_source` to the notebook.

* Switch to `pytorch-diffusers-serve-opt` container to for diffusion lora serving.

* add the dreambooth_lora notebook.

* minor update.

* Parameterize the "show_debug_logs" to facilitate automatic test of the Gradio notebooks.

* Lint format.

* minor updates

* Delete the two deprecated SD1.5 and 2.1 notebooks, as they were no longer referenced on any model cards.

* Sync Colab notebooks between g3 and github.

* format changes

* format update.
2024-05-23 17:08:23 +00:00
KCFindstrandGitHub 4a16d19cec Reformat llama3 deployment notebook and add quota check (#3001) 2024-05-23 17:07:50 +00:00
kittyabsandGitHub 452c77adb8 Ray-on-vertex-ai-public-access (#3002)
* Update get_started_with_pytorch_rov.ipynb

This is a test

* Addresses a feature request regarding the "warning" when using public access
2024-05-22 22:40:37 +00:00
lee1premiumandGitHub 558e6eed88 feat: Getting Tuned Text-Embeddings tutorial. (#3000)
* feat: Getting Tuned Text-Embeddings tutorial.

* feat: Getting Tuned Text-Embeddings tutorial.

* feat: Getting Tuned Text-Embeddings tutorial.
2024-05-22 21:03:44 +00:00
kittyabsandGitHub 82f32efc30 Tensorboard-intro-2 (#2996)
* Update get_started_with_pytorch_rov.ipynb

This is a test

* updated URL to point to tensorboard-introduction (and not "overview")

* change URL to /tensorboard-introduction
2024-05-21 23:38:53 +00:00
kittyabsandGitHub 4d2ea0d50f Tensorboard-introduction (#2994)
* Update get_started_with_pytorch_rov.ipynb

This is a test

* updated URL to point to tensorboard-introduction (and not "overview")
2024-05-21 20:45:18 +00:00
Manu KumarandGitHub c81c195d46 fix: use renamed create_feature_view() function in embedding notebook (#2993)
Increase sleep for DNS propagation to pass CI.
2024-05-21 20:43:19 +00:00
11d7c6dd7e Add checks to ensure the artifacts are copied to user bucket. (#2992)
Co-authored-by: minwoopark <minwoopark@google.com>
2024-05-21 12:30:21 +00:00
Gary WeiandGitHub 8bb6aa592a Merge a few SD notebooks in g3 and github (#2989)
* Add controlnet-canny to the Gradio playground, and some additional UX enhancement.

* Minor fixes.

* Minor fixes

* Add additional document regarding the list of supported models, and some UI enhancement.

* Minor update to the hyperlink.

* Rewrite the SD2.1 dreambooth finetune notebook.

* Add code owners.

* Some minor changes to the stable diffusion 2.1 and sd-xl notebooks.

* some additional minor fixes.

* additional fixes.

* Create a notebook to demonstrate dreambooth LoRA finetune for SD-XL model.

* minor updates

* add to the codeowner list.

* merge conflict.

* minor fix to the Gradio UI workshop notebook.

* Some minor updates to the SD2.1 deployment notebook.

* Minor update the `sd-xl` deployment notebook, based on the QA feedback.

* Add a few community models to the Gradio workshop.

* Switch `mediapipe-train` docker container from `vertex-ai-restricted` to `vertex-ai`, in the `mediapipe-train` notebooks.

* Create a notebook for model `instantx/instantid`.

* Update Gradio notebook to use the latest Gradio version and fix some bugs.

1. Update Gradio version to 4.29.0, as it complains 3.50.0 is too old.
2. Uninstall nest-asyncio and uvloop as a workaround to b/339301920 and https://github.com/gradio-app/gradio/issues/8238#issuecomment-2101066984.

* Resolve merge conflict.

* minor updates.

* minor updates.

* Merge some SD notebook in g3 and github.

* Remove the unused variable in the controlnet notebook.

* minor updates.

* include the SD1.5 dreambooth notebook.

* Include the sd1.5 dreambooth notebook.
2024-05-21 12:29:32 +00:00
Gary WeiandGitHub f5409efb30 Delete the two deprecated SD1.5 and 2.1 notebooks, as they were no longer referenced on any model cards. (#2986)
* Create a Gradio notebook for the new InstantId model.

* Add dreambooth finetune to the stable diffusion Gradio workshop notebook.

* Update the image generation Gradio notebook to support Dreambooth finetuning.

* linter update

* linter update

* minor fix to the instant-id notebook.

* Minor fix to the stable diffusion gradio notebook.

* Split the 'instant-id' deployment notebook prediction into two sections.

* add `deployment_source` to the notebook.

* Switch to `pytorch-diffusers-serve-opt` container to for diffusion lora serving.

* add the dreambooth_lora notebook.

* minor update.

* Parameterize the "show_debug_logs" to facilitate automatic test of the Gradio notebooks.

* Lint format.

* minor updates

* Delete the two deprecated SD1.5 and 2.1 notebooks, as they were no longer referenced on any model cards.
2024-05-21 12:27:59 +00:00
Manu KumarandGitHub 9fed51d88a fix: use renamed create_feature_view() function in optimized notebook (#2944)
Fx colab parameter usage - the linting/auto-format placed some variables
across multiple lines which doesn't work in colab. Make the FOS ID names
shorter to avoid this issue.

Also fix some usage for getting FOS - this can be done directly using
SDK constructor.

Set PSC allow list project to current project.

Increase sleep for DNS propagation to pass CI.
2024-05-16 19:41:07 +00:00
Huguens JeanandGitHub 5005b580d6 [Model Garden Team] Fix Gemma linter issue. (#2983)
* [Model Garden Team] Fix Gemma linter issue.

* [Model Garden Team] Fix Gemma linter issue.

* Update model_garden_gemma_evaluation.ipynb

* Update model_garden_gemma_evaluation.ipynb
2024-05-15 22:25:54 +00:00
dstnluong-googleandGitHub 80707d319d Fix Gemma finetuning notebook typo again. (#2984)
PiperOrigin-RevId: 634030805
2024-05-15 19:34:01 +00:00
Gary WeiandGitHub b630f8d5f0 Parameterize the "show_debug_logs" to facilitate automatic test of the Gradio notebooks. (#2975)
* Create a Gradio notebook for the new InstantId model.

* Add dreambooth finetune to the stable diffusion Gradio workshop notebook.

* Update the image generation Gradio notebook to support Dreambooth finetuning.

* linter update

* linter update

* minor fix to the instant-id notebook.

* Minor fix to the stable diffusion gradio notebook.

* Split the 'instant-id' deployment notebook prediction into two sections.

* add `deployment_source` to the notebook.

* Switch to `pytorch-diffusers-serve-opt` container to for diffusion lora serving.

* add the dreambooth_lora notebook.

* minor update.

* Parameterize the "show_debug_logs" to facilitate automatic test of the Gradio notebooks.

* Lint format.

* minor updates
2024-05-15 16:34:58 +00:00
Michael HuandGitHub 4fc807c570 Import AutoSxS pipeline from v1 directory (#2973) 2024-05-15 16:33:19 +00:00
chrisheechoandGitHub e1aef2f241 Update model_garden_gemma_fine_tuning_batch_deployment_on_rov.ipynb (#2982)
Need to update this headnode size per product requirement to not cause errors
2024-05-15 16:30:31 +00:00
dstnluong-googleandGitHub d5312bec5d Fix typo in Gemma finetuning notebook. (#2980)
PiperOrigin-RevId: 633712168
2024-05-15 00:32:24 +00:00
Louis LinandGitHub 5f86e65e6d fix: Broken 'processor' param due to linter (#2972)
* feat: Add notebook for E5 text embedding models

* feat: Add notebook for E5 text embedding models

* fix: Broken 'processor' param due to linter

* fix: Update the dev TEI docker images to the public ones
2024-05-14 22:08:53 +00:00
chrisheechoandGitHub f51d6c8e06 Update ray_cluster_management.ipynb (#2974)
Need to change the head node default to 16 to avoid error - i'm the PM for this product
2024-05-14 19:07:15 +00:00
84012f0e22 Add PaliGemma notebooks (#2971)
Co-authored-by: minwoopark <minwoopark@google.com>
2024-05-13 23:34:41 +00:00
dstnluong-googleandGitHub bec5a1f61a Sync GitHub repo (#2966)
PiperOrigin-RevId: 629577775
2024-05-13 21:33:53 +00:00
Louis LinandGitHub 67d14b1e91 feat: Add notebook for E5 text embedding models (#2961)
* feat: Add notebook for E5 text embedding models

* feat: Add notebook for E5 text embedding models
2024-05-11 13:01:26 +00:00
Gary WeiandGitHub 1d720eba33 Switch from pytorch-peft-serve to pytorch-diffusers-serve-opt container for diffusion model serving with lora. (#2950)
* Create a Gradio notebook for the new InstantId model.

* Add dreambooth finetune to the stable diffusion Gradio workshop notebook.

* Update the image generation Gradio notebook to support Dreambooth finetuning.

* linter update

* linter update

* minor fix to the instant-id notebook.

* Minor fix to the stable diffusion gradio notebook.

* Split the 'instant-id' deployment notebook prediction into two sections.

* add `deployment_source` to the notebook.

* Switch to `pytorch-diffusers-serve-opt` container to for diffusion lora serving.

* add the dreambooth_lora notebook.

* minor update.
2024-05-09 23:56:08 +00:00
Gary WeiandGitHub 246dc04784 Update the instant-id Gradio notebook to use the latest Gradio version and fix some bugs. (#2959)
* Add controlnet-canny to the Gradio playground, and some additional UX enhancement.

* Minor fixes.

* Minor fixes

* Add additional document regarding the list of supported models, and some UI enhancement.

* Minor update to the hyperlink.

* Rewrite the SD2.1 dreambooth finetune notebook.

* Add code owners.

* Some minor changes to the stable diffusion 2.1 and sd-xl notebooks.

* some additional minor fixes.

* additional fixes.

* Create a notebook to demonstrate dreambooth LoRA finetune for SD-XL model.

* minor updates

* add to the codeowner list.

* merge conflict.

* minor fix to the Gradio UI workshop notebook.

* Some minor updates to the SD2.1 deployment notebook.

* Minor update the `sd-xl` deployment notebook, based on the QA feedback.

* Add a few community models to the Gradio workshop.

* Switch `mediapipe-train` docker container from `vertex-ai-restricted` to `vertex-ai`, in the `mediapipe-train` notebooks.

* Create a notebook for model `instantx/instantid`.

* Create a Gradio notebook for the new InstantId model.

* Add dreambooth finetune to the stable diffusion Gradio workshop notebook.

* Update the image generation Gradio notebook to support Dreambooth finetuning.

* linter update

* linter update

* minor fix to the instant-id notebook.

* Minor fix to the stable diffusion gradio notebook.

* Update Gradio notebook to use the latest Gradio version and fix some bugs.

1. Update Gradio version to 4.29.0, as it complains 3.50.0 is too old.
2. Uninstall nest-asyncio and uvloop as a workaround to b/339301920 and https://github.com/gradio-app/gradio/issues/8238#issuecomment-2101066984.

* Resolve merge conflict.

* minor updates.

* minor updates.

* Update the instant-id Gradio notebook to use the latest Gradio version and fix some bugs.
2024-05-09 23:54:56 +00:00
Gary WeiandGitHub e08d445e17 Update Gradio notebook to use the latest Gradio version and fix some bugs. (#2958)
* Add controlnet-canny to the Gradio playground, and some additional UX enhancement.

* Minor fixes.

* Minor fixes

* Add additional document regarding the list of supported models, and some UI enhancement.

* Minor update to the hyperlink.

* Rewrite the SD2.1 dreambooth finetune notebook.

* Add code owners.

* Some minor changes to the stable diffusion 2.1 and sd-xl notebooks.

* some additional minor fixes.

* additional fixes.

* Create a notebook to demonstrate dreambooth LoRA finetune for SD-XL model.

* minor updates

* add to the codeowner list.

* merge conflict.

* minor fix to the Gradio UI workshop notebook.

* Some minor updates to the SD2.1 deployment notebook.

* Minor update the `sd-xl` deployment notebook, based on the QA feedback.

* Add a few community models to the Gradio workshop.

* Switch `mediapipe-train` docker container from `vertex-ai-restricted` to `vertex-ai`, in the `mediapipe-train` notebooks.

* Create a notebook for model `instantx/instantid`.

* Update Gradio notebook to use the latest Gradio version and fix some bugs.

1. Update Gradio version to 4.29.0, as it complains 3.50.0 is too old.
2. Uninstall nest-asyncio and uvloop as a workaround to b/339301920 and https://github.com/gradio-app/gradio/issues/8238#issuecomment-2101066984.

* Resolve merge conflict.

* minor updates.

* minor updates.
2024-05-09 23:54:11 +00:00
kewentandGitHub 859ed6a55d remove preview models in description (#2948) 2024-05-08 19:16:13 +00:00
dstnluong-googleandGitHub 44d11bee00 Set DEPLOY_SOURCE to MG notebooks. (#2947) 2024-05-08 18:53:13 +00:00
Gary WeiandGitHub 7c0213dfff Split the 'instant-id' deployment notebook prediction into two sections: (#2946)
* Create a Gradio notebook for the new InstantId model.

* Add dreambooth finetune to the stable diffusion Gradio workshop notebook.

* Update the image generation Gradio notebook to support Dreambooth finetuning.

* linter update

* linter update

* minor fix to the instant-id notebook.

* Minor fix to the stable diffusion gradio notebook.

* Split the 'instant-id' deployment notebook prediction into two sections.

* add `deployment_source` to the notebook.
2024-05-08 13:29:29 +00:00
467 changed files with 96508 additions and 58648 deletions
+10
View File
@@ -0,0 +1,10 @@
version: 2
updates:
# Ignore model garden dockerfiles:
- package-ecosystem: "npm"
directory: "/community-content/vertex_model_garden"
schedule:
interval: "monthly"
ignore:
- dependency-name: "*"
+1 -1
View File
@@ -4,7 +4,7 @@
# 2. To lint specific notebooks:
# docker run -v ${PWD}:/setup/app gcr.io/python-docs-samples-tests/notebook_linter:latest notebooks/1.ipynb notebooks/2.ipynb
FROM python:3.12
FROM python:3.13
WORKDIR setup
+5 -5
View File
@@ -2,9 +2,9 @@ git+https://github.com/tensorflow/docs
ipython
jupyter
nbconvert
black==24.4.2
pyupgrade==3.15.2
isort==5.13.2
flake8==7.0.0
nbqa==1.8.5
black==25.1.0
pyupgrade==3.19.1
isort==6.0.0
flake8==7.1.1
nbqa==1.9.1
+1 -1
View File
@@ -58,7 +58,7 @@ done
# Only check notebooks in test folders modified in this pull request.
# Note: Use process substitution to persist the data in the array
if [ ${#notebooks[@]} -eq 0 ]; then
echo "Checking for changed notebooked using git"
echo "Checking for changed notebooks using git"
while read -r file || [ -n "$line" ]; do
notebooks+=("$file")
done < <(git diff --name-only main... | grep '\.ipynb$')
+152 -13
View File
@@ -1,37 +1,176 @@
# Google Cloud Vertex AI Samples
# ![Google Cloud](https://avatars.githubusercontent.com/u/2810941?s=60&v=4) Google Cloud Vertex AI Samples
[![License](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](LICENSE)
Welcome to the Google Cloud [Vertex AI](https://cloud.google.com/vertex-ai/docs/) sample repository.
This repository contains notebooks, code samples, sample apps, and other resources that demonstrate how to use, develop and manage machine learning and generative AI workflows using Google Cloud Vertex AI.
## Overview
The repository contains [notebooks](https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/master/notebooks) and [community content](https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/master/community-content) that demonstrate how to develop and manage ML workflows using Google Cloud Vertex AI.
[Vertex AI](https://cloud.google.com/vertex-ai) is a fully-managed, unified AI development platform for building and using generative AI. This repository is designed to help you get started with Vertex AI. Whether you're new to Vertex AI or an experienced ML practitioner, you'll find valuable resources here.
For more Vertex AI Generative AI notebook samples, please visit the Vertex AI [Generative AI](https://github.com/GoogleCloudPlatform/generative-ai) GitHub repository.
## Explore, learn and contribute
You can explore, learn, and contribute to this repository to unleash the full potential of machine learning on Vertex AI!
### Explore and learn
Explore this repository, follow the links in the header section of each of the notebooks to -
![Colab](https://cloud.google.com/ml-engine/images/colab-logo-32px.png) Open and run the notebook in [Colab](https://colab.google/)\
![Colab Enterprise](https://cloud.google.com/ml-engine/images/colab-enterprise-logo-32px.png) Open and run the notebook in [Colab Enterprise](https://cloud.google.com/colab/docs/introduction)\
![Workbench](https://lh3.googleusercontent.com/UiNooY4LUgW_oTvpsNhPpQzsstV5W8F7rYgxgGBD85cWJoLmrOzhVs_ksK_vgx40SHs7jCqkTkCk=e14-rj-sc0xffffff-h130-w32) Open and run the notebook in [Vertex AI Workbench](https://cloud.google.com/vertex-ai/docs/workbench/introduction)\
![Github](https://cloud.google.com/ml-engine/images/github-logo-32px.png) View the notebook on Github
### Contribute
See the [Contributing Guide](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/master/CONTRIBUTING.md).
## Get started
To get started using Vertex AI, you must have a Google Cloud project.
- If you don't have a Google Cloud project, you can learn and build on GCP for free using [Free Trail](https://cloud.google.com/free).
- Once you have a Google Cloud project, you can learn more about [setting up a project and a development environment](https://cloud.google.com/vertex-ai/docs/start/cloud-environment).
## Repository structure
```bash
├── community-content - Sample code and tutorials contributed by the community
├── notebooks
│ ├── community - Notebooks contributed by the community
│ ├── official - Notebooks demonstrating use of each Vertex AI service
│ │ ├── automl
│ │ ├── custom
│ │ ├── ...
│ ├── community - Notebooks contributed by the community
│ │ ├── model_garden
│ │ ├── ...
├── community-content - Sample code and tutorials contributed by the community
```
## Examples
## Contributing
<!-- markdownlint-disable MD033 -->
<table>
Contributions welcome! See the [Contributing Guide](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/master/CONTRIBUTING.md).
<tr>
<th style="text-align: center;">Category</th>
<th style="text-align: center;">Product</th>
<th style="text-align: center;">Description</th>
</tr>
<tr>
<td>Model</td>
<td>
<a href="notebooks/community/model_garden"><code>Model Garden/</code></a>
</td>
<td>
Curated collection of first-party, open-source, and third-party models available on Vertex AI including Gemini, Gemma, Llama 3, Claude 3 and many more.
</td>
</tr>
<tr>
<td>Data</td>
<td>
<a href="notebooks/official/feature_store"><code>Feature Store/</code></a>
</td>
<td>
Set up and manage online serving using Vertex AI Feature Store.
</td>
</tr>
<tr>
<td></td>
<td>
<a href="notebooks/official/datasets"><code>datasets/</code></a>
</td>
<td>
Use BigQuery and Data Labeling service with Vertex AI.
</td>
</tr>
<tr>
<td>Model development</td>
<td>
<a href="notebooks/official/automl"><code>automl/</code></a>
</td>
<td>
Train and make predictions on AutoML models
</td>
</tr>
<tr>
<td></td>
<td>
<a href="notebooks/official/custom"><code>custom/</code></a>
</td>
<td>
Create, deploy and serve custom models on Vertex AI
</td>
</tr>
<tr>
<td></td>
<td>
<a href="notebooks/official/ray_on_vertex_ai"><code>ray_on_vertex_ai/</code></a>
</td>
<td>
Use Colab Enterprise and Vertex AI SDK for Python to connect to the Ray Cluster.
</td>
</tr>
<tr>
<td>Deploy and use</td>
<td>
<a href="notebooks/official/prediction"><code>prediction/</code></a>
</td>
<td>
Build, train and deploy models using prebuilt containers for custom training and prediction.
</td>
</tr>
<tr>
<td></td>
<td>
<a href="notebooks/official/model_registry"><code>model_registry/</code></a>
</td>
<td>
Use Model Registry to create and register a model.
</td>
</tr>
<tr>
<td></td>
<td>
<a href="notebooks/official/explainable_ai"><code>Explainable AI/</code></a>
</td>
<td>
Use Vertex Explainable AI's feature-based and example-based explanations to explain how or why a model produced a specific prediction.
</td>
</tr>
<tr>
<td></td>
<td>
<a href="notebooks/official/ml_metadata"><code>ml_metadata/</code></a>
</td>
<td>
Record the metadata and artifacts and query that metadata to help analyze, debug, and audit the performance of your ML system.
</td>
</tr>
<tr>
<td>Tools</td>
<td>
<a href="notebooks/official/pipelines"><code>Pipelines/</code></a>
</td>
<td>
Use `Vertex AI Pipelines` and `Google Cloud Pipeline Components` to build, tune, or deploy a custom model.
</td>
</tr>
</table>
<!-- markdownlint-enable MD033 -->
## Getting help
Please use the [issues page](https://github.com/GoogleCloudPlatform/vertex-ai-samples/issues) to provide feedback or submit a bug report.
## Get help
Please use the [Issues page](https://github.com/GoogleCloudPlatform/vertex-ai-samples/issues) to provide feedback or submit a bug report.
## Disclaimer
This is not an officially supported Google product. The code in this repository is for demonstrative purposes only.
## Feedback
Please feel free to fill out our [survey](https://bit.ly/vertex-ai-samples-survey) to give us feedback on the repo and its content.
## References
- [Vertex AI Jupyter Notebook tutorials](https://cloud.google.com/vertex-ai/docs/tutorials/jupyter-notebooks)
- Vertex AI [Generative AI](https://github.com/GoogleCloudPlatform/generative-ai) GitHub repository
- [Vertex AI documentaton](https://cloud.google.com/vertex-ai/docs)
+2
View File
@@ -10,6 +10,7 @@
/pipeline_components @Ark-kun
/pipeline_components/image_ml_model_training @lakeyk
/prediction_featurestore_integration @googleapis/vertex-prediction-team
/vertex_model_garden/model_oss/notebook_util @minwoo33park
/vertex_model_garden/model_oss/util @weigary
/vertex_model_garden/model_oss/diffusers @weigary
/vertex_model_garden/model_oss/keras @dstnluong-google
@@ -19,6 +20,7 @@
/vertex_model_garden/model_oss/movinet @KCFindstr
/vertex_model_garden/model_oss/data_converter @KCFindstr
/vertex_model_garden/model_oss/peft @weigary
/vertex_model_garden/model_oss/peft/templates @rayandasoriya
/vertex_model_garden/model_oss/lm-evaluation-harness @kathyyu-google
/vertex_model_garden/model_oss/tfvision @dstnluong-google
/vertex_model_garden/model_oss/fvlm @minwoo33park
@@ -1,16 +1,40 @@
FROM pytorch/pytorch:1.8.1-cuda11.1-cudnn8-runtime
# Stage 1: Build Environment
FROM pytorch/pytorch:1.8.1-cuda11.1-cudnn8-runtime AS builder
# Install necessary tools and dependencies
RUN apt-get update && \
apt-get install -y curl gnupg && \
echo "deb [signed-by=/usr/share/keyrings/cloud.google.gpg] http://packages.cloud.google.com/apt cloud-sdk main" | tee -a /etc/apt/sources.list.d/google-cloud-sdk.list && \
curl https://packages.cloud.google.com/apt/doc/apt-key.gpg | apt-key --keyring /usr/share/keyrings/cloud.google.gpg add - && \
curl https://packages.cloud.google.com/apt/doc/apt-key.gpg | apt-key --keyring /usr/share/keyrings/cloud.google.gpg add - && \
apt-get update -y && \
apt-get install google-cloud-sdk -y
apt-get install -y google-cloud-sdk
# Copy application code
COPY . /trainer
# Set working directory
WORKDIR /trainer
RUN pip install -r requirements.txt
# Install Python dependencies
RUN pip install --no-cache-dir -r requirements.txt
ENTRYPOINT ["python", "-m", "task"]
# Stage 2: Runtime Environment
FROM pytorch/pytorch:1.8.1-cuda11.1-cudnn8-runtime
# Install Google Cloud SDK
RUN apt-get update && \
apt-get install -y curl gnupg && \
echo "deb [signed-by=/usr/share/keyrings/cloud.google.gpg] http://packages.cloud.google.com/apt cloud-sdk main" | tee -a /etc/apt/sources.list.d/google-cloud-sdk.list && \
curl https://packages.cloud.google.com/apt/doc/apt-key.gpg | apt-key --keyring /usr/share/keyrings/cloud.google.gpg add - && \
apt-get update -y && \
apt-get install -y google-cloud-sdk && \
apt-get clean && rm -rf /var/lib/apt/lists/*
# Copy from the builder stage
COPY --from=builder /trainer /trainer
# Set working directory
WORKDIR /trainer
# Set the entry point
ENTRYPOINT ["python", "-m", "task"]
@@ -1,3 +1,3 @@
torch==1.13.1
torch==2.2.0
torchvision==0.9.1
tensorboard==2.5.0
@@ -1,3 +1,3 @@
torch==1.13.1
torch==2.2.0
torchvision==0.9.1
tensorboard==2.5.0
@@ -1,4 +1,4 @@
google-cloud-bigquery==2.20.0
tensorflow==2.7.2
tensorflow==2.12.1
pillow==10.3.0
tf-agents==0.8.0
@@ -1,4 +1,4 @@
google-cloud-pubsub==2.5.0
pillow==10.3.0
tf-agents==0.8.0
tensorflow==2.7.2
tensorflow==2.12.1
@@ -1,5 +1,5 @@
dataclasses==0.6
google-cloud-aiplatform==1.8.1
tensorflow==2.7.2
tensorflow==2.12.1
pillow==10.3.0
tf-agents==0.8.0
@@ -1 +1 @@
tensorflow==2.7.2
tensorflow==2.12.1
@@ -0,0 +1,15 @@
# Vertex AI custom prediction routines samples
## Overview
Vertex Custom Prediction Routines(CPR) simplify the process of building custom containers
and make local model testing easy. Here are the sameple codes for different libraries.
### Objectives
The objective is to provide various samples for Vertex Custom Prediction Routine(CPR).
### Supporting libraries
* torch
* sklearn
* xgboost
@@ -0,0 +1,33 @@
import numpy as np
import os
import pickle
from google.cloud.aiplatform.constants import prediction
from google.cloud.aiplatform.utils import prediction_utils
from google.cloud.aiplatform.prediction.predictor import Predictor
from sklearn.datasets import load_breast_cancer
from sklearn.linear_model import RidgeClassifier
class LinearRegressionPredictor(Predictor):
def __init__(self):
return
def load(self, artifacts_uri: str) -> None:
prediction_utils.download_model_artifacts(artifacts_uri)
if os.path.exists(prediction.MODEL_FILENAME_PKL):
self._model = pickle.load(open(prediction.MODEL_FILENAME_PKL, "rb"))
else:
self._model = RidgeClassifier()
X, y = load_breast_cancer(return_X_y=True)
self._model.fit(X, y)
def preprocess(self, prediction_input: dict) -> np.ndarray:
instances = prediction_input["instances"]
return np.asarray(instances)
def predict(self, instances: np.ndarray) -> np.ndarray:
return self._model.predict(instances)
def postprocess(self, prediction_results: np.ndarray) -> dict:
return {"predictions": prediction_results.tolist()}
@@ -0,0 +1,33 @@
import numpy as np
import os
import pickle
from google.cloud.aiplatform.constants import prediction
from google.cloud.aiplatform.utils import prediction_utils
from google.cloud.aiplatform.prediction.predictor import Predictor
from sklearn.datasets import make_blobs
from sklearn.linear_model import LinearRegression
class LinearRegressionPredictor(Predictor):
def __init__(self):
return
def load(self, artifacts_uri: str) -> None:
prediction_utils.download_model_artifacts(artifacts_uri)
if os.path.exists(prediction.MODEL_FILENAME_PKL):
self._model = pickle.load(open(prediction.MODEL_FILENAME_PKL, "rb"))
else:
self._model = LogisticRegression()
X, y = make_blobs(n_samples=100, centers=2, n_features=2, random_state=1)
self._model.fit(X, y)
def preprocess(self, prediction_input: dict) -> np.ndarray:
instances = prediction_input["instances"]
return np.asarray(instances)
def predict(self, instances: np.ndarray) -> np.ndarray:
return self._model.predict_proba(instances)
def postprocess(self, prediction_results: np.ndarray) -> dict:
return {"predictions": prediction_results.tolist()}
@@ -0,0 +1,33 @@
import numpy as np
import os
import pickle
from google.cloud.aiplatform.constants import prediction
from google.cloud.aiplatform.utils import prediction_utils
from google.cloud.aiplatform.prediction.predictor import Predictor
from sklearn.linear_model import SGDClassifier
class SGDClassifierPredictor(Predictor):
def __init__(self):
return
def load(self, artifacts_uri: str) -> None:
prediction_utils.download_model_artifacts(artifacts_uri)
if os.path.exists(prediction.MODEL_FILENAME_PKL):
self._model = pickle.load(open(prediction.MODEL_FILENAME_PKL, "rb"))
else:
self._model = SGDClassifier(max_iter=5)
X = [[0., 0.], [1., 1.]]
y = [0, 1]
self._model.fit(X, y)
def preprocess(self, prediction_input: dict) -> np.ndarray:
instances = prediction_input["instances"]
return np.asarray(instances)
def predict(self, instances: np.ndarray) -> np.ndarray:
return self._model.predict(instances)
def postprocess(self, prediction_results: np.ndarray) -> dict:
return {"predictions": prediction_results.tolist()}
@@ -0,0 +1,34 @@
import os
import torch
from google.cloud.aiplatform.utils import prediction_utils
from google.cloud.aiplatform.prediction.predictor import Predictor
from torchvision.models import detection, resnet50, ResNet50_Weights
from typing import Dict, List
class ResNetPredictor(Predictor):
def __init__(self):
return
def load(self, artifacts_uri: str) -> None:
prediction_utils.download_model_artifacts(artifacts_uri)
if os.path.exists("model.pth.tar"):
self.model = detection.fasterrcnn_resnet50_fpn(pretrained=True)
stat_dic = torch.load("model.pth.tar")
self.model.load_state_dict(stat_dic['state_dict'])
else:
weights = ResNet50_Weights.DEFAULT
self.model = resnet50(weights=weights)
self.model.eval()
def preprocess(self, prediction_input: dict) -> torch.Tensor:
instances = prediction_input["instances"]
return torch.Tensor(instances)
@torch.inference_mode()
def predict(self, instances: torch.Tensor) -> List[str]:
return self._model(instances)
def postprocess(self, prediction_results: List[str]) -> Dict:
return {"predictions": prediction_results}
@@ -0,0 +1,73 @@
import ast
import json
import os
import pickle
import torch
from google.cloud.aiplatform.utils import prediction_utils
from google.cloud.aiplatform.prediction.predictor import Predictor
from transformers import AutoModelForQuestionAnswering
from typing import Dict, List
class TorchTransformersPredictor(Predictor):
def __init__(self):
return
def load(self, artifacts_uri: str) -> None:
prediction_utils.download_model_artifacts(artifacts_uri)
if os.path.isfile("setup_config.json"):
with open("setup_config.json") as setup_config_file:
self.setup_config = json.load(setup_config_file)
if os.path.exists("model.pt"):
self.model = AutoModelForQuestionAnswering.from_pretrained("model.pt")
self.model.eval()
else:
raise ValueError("One of the following model files must be provided: model.pt.")
def preprocess(self, prediction_input: dict) -> torch.Tensor:
max_length = self.setup_config["max_length"]
instances = prediction_input["instances"]
question_context = ast.literal_eval(instances)
question = question_context["question"]
context = question_context["context"]
inputs = self.tokenizer.encode_plus(
question,
context,
max_length=int(max_length),
pad_to_max_length=True,
add_special_tokens=True,
return_tensors="pt",
)
input_ids = inputs["input_ids"]
attention_mask = inputs["attention_mask"]
return torch.Tensor(input_ids, attention_mask)
@torch.inference_mode()
def predict(self, instances: torch.Tensor) -> List[str]:
input_ids, attention_mask = instances
outputs = self._model(input_ids, attention_mask)
answer_start_scores = outputs.start_logits
answer_end_scores = outputs.end_logits
num_rows, num_cols = answer_start_scores.shape
inferences = []
for i in range(num_rows):
answer_start_scores_one_seq = answer_start_scores[i].unsqueeze(0)
answer_start = torch.argmax(answer_start_scores_one_seq)
answer_end_scores_one_seq = answer_end_scores[i].unsqueeze(0)
answer_end = torch.argmax(answer_end_scores_one_seq) + 1
prediction = self.tokenizer.convert_tokens_to_string(
self.tokenizer.convert_ids_to_tokens(
input_ids[i].tolist()[answer_start:answer_end]
)
)
inferences.append(prediction)
return inferences
def postprocess(self, prediction_results: List[str]) -> Dict:
return {"predictions": prediction_results}
@@ -0,0 +1,37 @@
import os
import numpy as np
import pickle
import xgboost as xgb
from google.cloud.aiplatform.constants import prediction
from google.cloud.aiplatform.utils import prediction_utils
from google.cloud.aiplatform.prediction.predictor import Predictor
from sklearn.datasets import make_blobs
from xgboost import XGBClassifier
class ClassifierPredictor(Predictor):
def __init__(self):
return
def load(self, artifacts_uri: str) -> None:
prediction_utils.download_model_artifacts(artifacts_uri)
if os.path.exists(prediction.MODEL_FILENAME_PKL):
booster = pickle.load(open(prediction.MODEL_FILENAME_PKL, "rb"))
else:
X, y = make_blobs(n_samples=100, centers=2, n_features=2, random_state=1)
model = XGBClassifier()
model.fit(X, y)
booster = model.get_booster()
self._booster = booster
def preprocess(self, prediction_input: dict) -> xgb.DMatrix:
instances = prediction_input["instances"]
return xgb.DMatrix(instances)
def predict(self, instances: xgb.DMatrix) -> np.ndarray:
return self._booster.predict(instances)
def postprocess(self, prediction_results: np.ndarray) -> dict:
return {"predictions": prediction_results.tolist()}
@@ -0,0 +1,41 @@
import os
import numpy as np
import pandas as pd
import pickle
import xgboost as xgb
from google.cloud.aiplatform.constants import prediction
from google.cloud.aiplatform.utils import prediction_utils
from google.cloud.aiplatform.prediction.predictor import Predictor
class XGBRankerPredictor(Predictor):
def __init__(self):
return
def load(self, artifacts_uri: str) -> None:
prediction_utils.download_model_artifacts(artifacts_uri)
if os.path.exists(prediction.MODEL_FILENAME_PKL):
booster = pickle.load(open(prediction.MODEL_FILENAME_PKL, "rb"))
self._booster = booster
else:
N = 500
dates = pd.date_range(start='2023-01-01', end='2023-01-12', periods=N)
X = pd.DataFrame(np.random.randn(N, 5), columns=list('ABCDE'), index=dates)
y = pd.Series(np.random.randint(0, 10, size=N), index=dates, name='label')
group = X.groupby(dates + pd.offsets.MonthEnd(0)).size()
sample_weight = pd.Series(np.arange(len(group)), index=group.index)
model = xgb.XGBRanker(objective='rank:pairwise', max_depth=3, learning_rate=0.1, booster='gbtree', tree_method='hist', n_jobs=4, n_estimators=50, enable_categorical=False, random_state=42)
model.fit(X=X, y=y, group=group, sample_weight=sample_weight, verbose=True)
booster = model.get_booster()
self._booster = booster
def preprocess(self, prediction_input: dict) -> xgb.DMatrix:
instances = prediction_input["instances"]
return xgb.DMatrix(instances)
def predict(self, instances: xgb.DMatrix) -> np.ndarray:
return self._booster.predict(instances, output_margin=False, ntree_limit=0)
def postprocess(self, prediction_results: np.ndarray) -> dict:
return {"predictions": prediction_results.tolist()}
@@ -1,227 +0,0 @@
# Benchmark report on fine tuning the OpenLLaMA 7B model on Google Cloud Vertex Model Garden
Gary Wei, Software Engineer, Google Cloud
Dustin Luong, Software Engineer, Google Cloud
Changyu Zhu, Software Engineer, Google Cloud
Genquan Duan, Software Engineer, Google Cloud
## Introduction
Fine-tuning of LLMs can be non-trivial to find an optimal configuration of
machine types, training parameters, and other hyperparameters that achieves a
good balance between cost efficiency and model performance. To facilitate users
in conducting tuning experiments, this report benchmarks OpenLLaMA 7B
fine-tuning on Google Cloud Vertex Model Garden, demonstrating both efficiency
and effectiveness. The observations are general and can be applied to other LLM
models.
We benchmarked fine tuning algorithms [LoRA](https://arxiv.org/abs/2106.09685)
and [QLoRA](https://arxiv.org/abs/2305.14314) supported by
[huggingface PEFT libraries](https://github.com/huggingface/peft). LoRA, short
for Low-Rank Adaptation of Large Language Models, is an improved fine tuning
method where instead of fine tuning all the weights that constitute the weight
matrix of the pre-trained large language model, two smaller matrices that
approximate this larger matrix are fine-tuned. QLoRA is an even more
memory-efficient version of LoRA, where the pretrained model is loaded to GPU
memory as quantized 4-bit weights, while preserving similar effectiveness to
LoRA. We also provide simple scripts and parameter settings to reproduce the
results reported in this report.
In general, there are many factors that affect the performance of fine-tuning
experiments, such as hardware settings, parameters, cost, and accuracy. It is
impractical to obtain benchmarks for all possible combinations of these factors.
Instead, we focus on tuning a subset of related parameters and evaluating their
impact on a set of chosen metrics. The evaluation metrics are GPU memory usage,
percentage of parameters tuned, tuning speed, cost, and accuracy. The tuning
parameters are batch size, lora rank, maximum sequence length, and maximum
training steps.
## Key takeaways
- **Use QLoRA to minimize the peak GPU requirements**: The QLoRA can
significantly reduce the peak GPU memory usage by ~75% compared to LoRA. For
OpenLLaMA7b, the peak memory is ~28G for LoRA and ~7G for QLoRA.
- **Use LoRA to maximize the tuning speed and minimize the tuning cost**: LoRA
is ~66% faster than QLoRA in fine tuning speed. LoRA/QLoRA tuning cost is
low generally, while LoRA is even ~40% cheaper than QLoRA with the same
parameters. Suggest to use QLoRA for limited GPU memories, and LoRA for
limited training budgets. For OpenLLaMA7b, the tuning speed for LoRA/QLoRA
~5 samples / 3 samples per second, and the tuning cost for LoRA/QLoRA in 500
steps is ~$1/$1.7 on `a2-highgpu-1g` with 1 A100 40G GPU. The tuning cost
for QLoRA in 500 steps is $6.75 on n1-standard-8 with 1 V100 GPU, while LoRA
could not run because of OOM.
- **Use QLoRA to tune models with large sequence lengths**. For OpenLLaMA7b,
the max sequence length for QLoRA can be 2048 when consuming 16.3G GPU,
while the max sequence length for LoRA is 512 when consuming 28.2G GPU, and
encounter OOM when max sequence length is 1024.
- **Both LoRA and QLoRA give similar accuracy improvement after fine tuning.**
For OpenLLaMA7b, both LoRA/QLoRA can improve the average accuracy by ~4%
evaluating on 3 typical tasks (ARC challenge, HellaSwag and TruthfulQA),
after training 1875 steps on dataset
[timdettmers/openassistant-guanaco](https://huggingface.co/datasets/timdettmers/openassistant-guanaco).
- **Use a big batch size if GPU memory is not a constraint**. For OpenLLaMA7b
with other default parameters, we suggest using a batch size as 24 for
QLoRA, but 2 for LoRA when tuning with 1 A100 40G. We also suggest using a
batch size as 8 for QLoRA when tuning with 1 V100. Tuning with LoRA and
batch size as 1 got OOM and we don't recommend tuning LoRA with 1 V100.
## Benchmark Details
### Experiment Setup
The benchmark dataset is
[timdettmers/openassistant-guanaco](https://huggingface.co/datasets/timdettmers/openassistant-guanaco).
The training dataset is directly downloaded from hugging face to the VM, before
every experiment.
The default tuning parameters during benchmark are:
- Host VM: a2-highgpu-1g
- Accelerator type: 1 A100 40G
- batch size: 2
- lora_rank: 16
- max_seq_length: 512
- precision_mode: float16
- max_train_steps: 500
For simplicity, we set the precision mode to `float16` when tuning LoRA models,
and set the precision to `4bit` for QLoRA.
Sample script to start fine tuning dockers in a VM on GCP.
```shell
IMAGE_TAG=us-docker.pkg.dev/vertex-ai/vertex-vision-model-garden-dockers/pytorch-peft-train:latest
docker run --runtime=nvidia -e NVIDIA_VISIBLE_DEVICES=0 \
--rm --name "test_gpu" -it --pull=always ${IMAGE_TAG} \
--task=instruct-lora \
--pretrained_model_id=openlm-research/open_llama_7b \
--dataset_name="timdettmers/openassistant-guanaco" \
--instruct_column_in_dataset="text" \
--precision_mode="float16" \
--output_dir=<OUTPUT DIR> \
--lora_rank=2 \
--max_sequence_length=512 \
--learning_rate=2e-4 \
--max_steps=50
```
### GPU Memory
In this benchmark, we investigated the impact of batch size, lora rank, and
maximum sequence length on GPU memory, and then made recommendations on the
maximum batch size for different GPUs.
#### Peak GPU memory by batch size (GB)
<img src="images/openllama_7b_fine_tune_benchmark_report/openllama-7b-peak-gpu-vs-batch-size.png" width="600">
- The QLoRA can significantly reduce the peak GPU memory usage by ~75%
compared to LoRA. The peak GPU memory is ~28G for LoRA and ~7G for QLoRA
when batch size is 2.
- QLoRA can support much larger batch sizes than LoRA
- We can use a batch size as 32 for QLoRA, but only 2 for LoRA on 1 A100
40G.
- We can use a batch size of 8 for QLoRA on 1 V100 GPU. LoRA will fail
with OOM even with a batch size of 1.
#### Peak GPU memory by LoRA rank (GB)
<img src="images/openllama_7b_fine_tune_benchmark_report/openllama-7b-peak-gpu-vs-lora-rank.png" width="600">
- Peak GPU memories are quite similar for different LoRA ranks for both
LoRA/QLoRA.
- The peak GPU memory increasing percentages are very small generally when
LoRA rank increases.
- The peak GPU memory increases from 28G with LoRA rank 4 to 29.09G with
LoRA rank 64, and the increasing percentage is only ~3.9%.
#### Peak GPU memory by max sequence length for LoRA/QLoRA (GB)
<img src="images/openllama_7b_fine_tune_benchmark_report/openllama-7b-peak-gpu-vs-max-seq-length.png" width="600">
- The peak GPU increases quickly when max sequence length increases for both
LoRA/QLoRA, and the increasing rate of LoRA is much faster than QLoRA.
- For LoRA tuning, the GPU memory increased from 20.5G (max sequence
length=256) to 28.2G (max sequence length=512), an increase of ~37%.
- For QLoRA tuning, the GPU memory increased from 6.94G (max sequence
length=256) to 7.57G (max sequence length=512), an increase of ~9%.
- The max sequence length for QLoRA can be 2048 when consuming 16.3G GPU,
while the max sequence length for LoRA is 512 when consuming 28.2G GPU, and
encounter OOM when max sequence length is 1024.
### Fine Tuning Parameters
This section shows the number/percentage of trainable parameters, and the sizes
of the fine tuned models. LoRA and QLoRA differ only in how they represent the
precision of their parameters. The total number of parameters and the number of
trainable parameters are the same for both methods.
| LoRA Rank | Finetuned parameters | Total parameters | Trainable Parameter Percentage | Fine tuned model size (MB) |
| --------- | -------------------- | ---------------- | ------------------------------ | -------------------------- |
| 8 | 2.00E+07 | 6.76E+09 | 0.3% | 76.4 |
| 16 | 4.00E+07 | 6.78E+09 | 0.6% | 152.65 |
| 32 | 8.00E+07 | 6.82E+09 | 1.2% | 305.15 |
| 64 | 1.60E+08 | 6.90E+09 | 2.3% | 610.15 |
LoRA/QLoRA tunes quite a small fraction (only 0.3% with LoRA rank=8) of all
parameters, and the tuned models are very small (only 76.4MB with LoRA rank=8).
### Fine Tuning Speed And Costs
The fine-tuning speed and cost are affected by various factors, such as the
GPUs, LoRA ranks, and max sequence lengths.
- LoRA is ~66% faster than QLoRA in fine tuning speed. The tuning speed for
LoRA/QLoRA ~5 samples / 3 samples per second on 1 A100 40G GPU
- Higher LoRA ranks, slower tuning speed for both LoRA/QLoRA.
- LoRA tuning speed reduces from ~5 samples per second with LoRA rank as 8
to ~4 samples per second with LoRA rank as 64, slowed down by 20%.
- QLoRA tuning speed reduces from ~3 samples per second with LoRA rank as
8 to ~2.5 samples per second with LoRA rank as 64, slowed down by 17%.
<img src="images/openllama_7b_fine_tune_benchmark_report/openllama-7b-tune-speed-vs-lora-rank.png" width="600">
- Longer sequence lengths, slower tuning speed.
- LoRA tuning speed reduces from ~5.56 samples per second with max
sequence length as 256 to ~4.84 samples per second with max sequence
length as 512 slowed down by 13%.
- LoRA tuning speed reduces from ~2.95 samples per second with max
sequence length as 256 to ~2.88 samples per second with max sequence
length as 512 slowed down by ~2.4%.
<img src="images/openllama_7b_fine_tune_benchmark_report/openllama-7b-tune-speed-lora-qlora.png" width="600">
- LoRA/QLoRA tuning cost is low generally, while LoRA is even ~40% cheaper
than QLoRA with the same parameters.
- The LoRA/QLoRA fine tuning cost for 500 steps is ~$1/$1.7 on 1 A100 40G.
- The tuning cost for QLoRA in 500 steps is $6.75 on n1-standard-8 with 1
V100 GPU, while LoRA could not run because of OOM.
<img src="images/openllama_7b_fine_tune_benchmark_report/openllama-7b-tune-cost-lora-qlora.png" width="600">
### Accuracy
We fine tuned Open Llama 7B model with
[timdettmers/openassistant-guanaco](https://huggingface.co/datasets/timdettmers/openassistant-guanaco),
and report accuracy similar to the
[HuggingFace leaderboard](https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard)
using
[Eleuther AI Language Model Evaluation Harness](https://github.com/EleutherAI/lm-evaluation-harness).
[HuggingFace leaderboard](https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard)
mainly compares models on ARC, HellaSwag, MMLU, and TruthfulQA. The authors did
not publish OpenLLaMA 7B on MMLU
([link](https://huggingface.co/openlm-research/open_llama_7b)). Therefore, we
only benchmark accuracies on ARC, HellaSwag, and TruthfulQA.
| | Mean | ARC | HellaSwag | TruthfulQA | Tuning Parameters |
| ------------------------------------------------------------ | ---- | ---- | --------- | ---------- | ------------------------------------------------------------ |
| OpenLLaMA7B ([Original Report](https://huggingface.co/openlm-research/open_llama_7b)) | 0.49 | 0.41 | 0.73 | 0.34 | n/a |
| OpenLLaMA7B ([Re-run with lm-evaluation-harness](https://github.com/EleutherAI/lm-evaluation-harness)) | 0.51 | 0.47 | 0.72 | 0.35 | n/a |
| OpenLLaMA7B+LoRA | 0.56 | 0.48 | 0.74 | 0.45 | LoRA Rank=16; Max Sequence Length=512;Learning Rate=1e-4; Train steps=1875 |
| OpenLLaMA7B+QLoRA | 0.53 | 0.45 | 0.73 | 0.42 | LoRA Rank=16; Max Sequence Length=512; Learning Rate=1e-4; Train steps=1875 |
- The base OpenLLaMA7B model gets better performance (2%) when using the
[Eleuther AI Language Model Evaluation Harness](https://github.com/EleutherAI/lm-evaluation-harness).
- LoRA/QLoRA can improve the performance by ~2-4% when trained for 1875 steps
with learning rate 1e-4.
@@ -1,6 +1,6 @@
--find-links https://download.pytorch.org/whl/torch_stable.html
torch==2.0.1+cu118
torch==2.2.0
numpy==1.26.1
absl_py==2.0.0
accelerate==0.24.0
@@ -16,7 +16,7 @@ Pillow==10.3.0
rawpy==0.18.1
scipy==1.11.3
scikit-image==0.22.0
scikit-learn==1.3.2
scikit-learn==1.5.0
tensorboard==2.15.0
tensorboardX==2.6.2.2
tqdm==4.66.3
@@ -0,0 +1,622 @@
"""Common util functions for notebook."""
import base64
import datetime
import io
import json
import os
import subprocess
from typing import Any, Dict, Sequence
from google.cloud import storage
import matplotlib.pyplot as plt
import numpy as np
from PIL import Image
import requests
import tensorflow as tf
import yaml
GCS_URI_PREFIX = "gs://"
CHECKPOINT_BUCKET = "gs://model_garden_checkpoints"
def convert_numpy_array_to_byte_string_via_tf_tensor(
np_array: np.ndarray,
) -> str:
"""Serializes a numpy array to tensor bytes.
Args:
np_array: A numpy array.
Returns:
A tensor bytes.
"""
tensor_array = tf.convert_to_tensor(np_array)
tensor_byte_string = tf.io.serialize_tensor(tensor_array)
return tensor_byte_string.numpy()
def get_jpeg_bytes(local_image_path: str, new_width: int = -1) -> bytes:
"""Returns jpeg bytes given an image path and resizes if required.
Args:
local_image_path: A string of local image path.
new_width: An integer of new image width.
Returns:
A jpeg bytes.
"""
image = Image.open(local_image_path)
if new_width <= 0:
new_image = image
else:
width, height = image.size
print("original input image size: ", width, " , ", height)
new_height = int(height * new_width / width)
print("new input image size: ", new_width, " , ", new_height)
new_image = image.resize((new_width, new_height))
buffered = io.BytesIO()
new_image.save(buffered, format="JPEG")
return buffered.getvalue()
def gcs_fuse_path(path: str) -> str:
"""Try to convert path to gcsfuse path if it starts with gs:// else do not modify it.
Args:
path: A string of path.
Returns:
A gcsfuse path.
"""
path = path.strip()
if path.startswith("gs://"):
return "/gcs/" + path[5:]
return path
def get_job_name_with_datetime(prefix: str) -> str:
"""Gets a job name by adding current time to prefix.
Args:
prefix: A string of job name prefix.
Returns:
A job name.
"""
now = datetime.datetime.now().strftime("%Y%m%d_%H%M%S")
job_name = f"{prefix}-{now}".replace("_", "-")
return job_name
def create_job_name(prefix: str) -> str:
"""Creates a job name.
Args:
prefix: A string of job name prefix.
Returns:
A job name.
"""
user = os.environ.get("USER")
now = datetime.datetime.now().strftime("%Y%m%d_%H%M%S")
job_name = f"{prefix}-{user}-{now}".replace("_", "-")
return job_name
def save_subset_annotation(
input_annotation_path: str, output_annotation_path: str
):
"""Saves a subset of COCO annotation json file with CCA 4.0 license.
Args:
input_annotation_path: A string of input annotation path.
output_annotation_path: A string of output annotation path.
"""
with open(input_annotation_path) as f:
coco_json = json.load(f)
img_ids = set()
images = []
annotations = []
for img in coco_json["images"]:
if img["license"] in [4, 5]: # CCA 4.0 license.
img_ids.add(img["id"])
images.append(img)
for ann in coco_json["annotations"]:
if ann["image_id"] in img_ids:
annotations.append(ann)
new_json = {
"info": coco_json["info"],
"licenses": coco_json["licenses"],
"images": images,
"annotations": annotations,
"categories": coco_json["categories"],
}
with open(output_annotation_path, "w") as f:
json.dump(new_json, f)
def image_to_base64(image: Any, image_format: str = "JPEG") -> str:
"""Converts an image to base64.
Args:
image: A PIL.Image instance.
image_format: A string of image format.
Returns:
A base64 string.
"""
buffer = io.BytesIO()
image.save(buffer, format=image_format)
image_str = base64.b64encode(buffer.getvalue()).decode("utf-8")
return image_str
def base64_to_image(image_str: str) -> Any:
"""Convert base64 encoded string to an image.
Args:
image_str: A string of base64 encoded image.
Returns:
A PIL.Image instance.
"""
image = Image.open(io.BytesIO(base64.b64decode(image_str)))
return image
def image_grid(imgs: Sequence[Any], rows: int = 2, cols: int = 2) -> Any:
"""Creates an image grid.
Args:
imgs: A list of PIL.Image instances.
rows: An integer of number of rows.
cols: An integer of number of columns.
Returns:
A PIL.Image instance.
"""
w, h = imgs[0].size
grid = Image.new(
mode="RGB", size=(cols * w + 10 * cols, rows * h), color=(255, 255, 255)
)
for i, img in enumerate(imgs):
grid.paste(img, box=(i % cols * w + 10 * i, i // cols * h))
return grid
def display_image(image: Any):
"""Displays an image.
Args:
image: A PIL.Image instance.
"""
_ = plt.figure(figsize=(20, 15))
plt.grid(False)
plt.imshow(image)
def download_gcs_file_to_local(gcs_uri: str, local_path: str):
"""Download a gcs file to a local path.
Args:
gcs_uri: A string of file path on GCS.
local_path: A string of local file path.
"""
if not gcs_uri.startswith(GCS_URI_PREFIX):
raise ValueError(
f"{gcs_uri} is not a GCS path starting with {GCS_URI_PREFIX}."
)
client = storage.Client()
os.makedirs(os.path.dirname(local_path), exist_ok=True)
with open(local_path, "wb") as f:
client.download_blob_to_file(gcs_uri, f)
def download_image(url: str) -> str:
"""Downloads an image from the given URL.
Args:
url: The URL of the image to download.
Returns:
base64 encoded image.
"""
response = requests.get(url)
return Image.open(io.BytesIO(response.content)) # pytype: disable=bad-return-type # pillow-102-upgrade
def resize_image(image: Any, new_width: int = 1000) -> Any:
"""Resizes an image to a certain width.
Args:
image: The image which has to be resized.
new_width: New width of the image.
Returns:
New resized image.
"""
width, height = image.size
new_height = int(height * new_width / width)
new_img = image.resize((new_width, new_height))
return new_img
def load_img(path: str) -> Any:
"""Reads image from path and return PIL.Image instance.
Args:
path: A string of image path.
Returns:
A PIL.Image instance.
"""
img = tf.io.read_file(path)
img = tf.image.decode_jpeg(img, channels=3)
return Image.fromarray(np.uint8(img)).convert("RGB")
def decode_image(
image_str_tensor: tf.string, new_height: int, new_width: int
) -> tf.float32:
"""Converts and resizes image bytes to image tensor.
Args:
image_str_tensor: A string of image bytes.
new_height: An integer of new image height.
new_width: An integer of new image width.
Returns:
An image tensor.
"""
image = tf.io.decode_image(image_str_tensor, 3, expand_animations=False)
image = tf.image.resize(image, (new_height, new_width))
return image
def get_label_map(label_map_yaml_filepath: str) -> Dict[int, str]:
"""Returns class id to label mapping given a filepath to the label map.
Args:
label_map_yaml_filepath: A string of label map yaml file path.
Returns:
A dictionary of class id to label mapping.
"""
with tf.io.gfile.GFile(label_map_yaml_filepath, "rb") as input_file:
label_map = yaml.safe_load(input_file.read())["label_map"]
return label_map
def get_prediction_instances(test_filepath: str, new_width: int = -1) -> Any:
"""Generate instance from image path to pass to Vertex AI Endpoint for prediction.
Args:
test_filepath: A string of test image path.
new_width: An integer of new image width.
Returns:
A list of instances.
"""
if new_width <= 0:
with tf.io.gfile.GFile(test_filepath, "rb") as input_file:
encoded_string = base64.b64encode(input_file.read()).decode("utf-8")
else:
img = load_img(test_filepath)
width, height = img.size
print("original input image size: ", width, " , ", height)
new_height = int(height * new_width / width)
new_img = img.resize((new_width, new_height))
print("resized input image size: ", new_width, " , ", new_height)
buffered = io.BytesIO()
new_img.save(buffered, format="JPEG")
encoded_string = base64.b64encode(buffered.getvalue()).decode("utf-8")
instances = [{
"encoded_image": {"b64": encoded_string},
}]
return instances
def vqa_predict(
endpoint: Any,
question_prompts: Sequence[str],
image: Any,
language_code: str = "en",
new_width: int = 1000,
) -> Sequence[str]:
"""Predicts the answer to a question about an image using an Endpoint."""
# Resize and convert image to base64 string.
resized_image = resize_image(image, new_width)
resized_image_base64 = image_to_base64(resized_image)
instances = []
if question_prompts:
# Format question prompt
question_prompt_format = "answer {} {}\n"
for question_prompt in question_prompts:
if question_prompt:
instances.append({
"prompt": question_prompt_format.format(
language_code, question_prompt
),
"image": resized_image_base64,
})
else:
instances.append({
"image": resized_image_base64,
})
response = endpoint.predict(instances=instances)
return [pred.get("response") for pred in response.predictions]
def caption_predict(
endpoint: Any,
language_code: str,
image: Any,
caption_prompt: bool = False,
new_width: int = 1000,
) -> str:
"""Predicts a caption for a given image using an Endpoint."""
# Resize and convert image to base64 string.
resized_image = resize_image(image, new_width)
resized_image_base64 = image_to_base64(resized_image)
instance = {"image": resized_image_base64}
if caption_prompt:
# Format caption prompt
caption_prompt_format = "caption {}\n"
instance["prompt"] = caption_prompt_format.format(language_code)
instances = [instance]
response = endpoint.predict(instances=instances)
return response.predictions[0].get("response")
def ocr_predict(
endpoint: Any,
ocr_prompt: str,
image: Any,
new_width: int = 1000,
) -> str:
"""Extracts text from a given image using an Endpoint."""
# Resize and convert image to base64 string.
resized_image = resize_image(image, new_width)
resized_image_base64 = image_to_base64(resized_image)
instance = {"image": resized_image_base64}
if ocr_prompt:
instance["prompt"] = ocr_prompt
instances = [instance]
response = endpoint.predict(instances=instances)
return response.predictions[0].get("response")
def detect_predict(
endpoint: Any,
detect_prompt: str,
image: Any,
new_width: int = 1000,
) -> str:
"""Predicts the answer to a question about an image using an Endpoint."""
# Resize and convert image to base64 string.
resized_image = resize_image(image, new_width)
resized_image_base64 = image_to_base64(resized_image)
instance = {"image": resized_image_base64}
if detect_prompt:
instance["prompt"] = detect_prompt
instances = [instance]
response = endpoint.predict(instances=instances)
return response.predictions[0].get("response")
def copy_model_artifacts(
model_id: str,
model_source: str,
model_destination: str,
) -> None:
"""Copies model artifacts from model_source to model_destination.
model_source and model_destination should be GCS path.
Args:
model_id: The model id.
model_source: The source of the model artifact.
model_destination: The destination of the model artifact.
"""
if not model_source.startswith(GCS_URI_PREFIX):
raise ValueError(
f"{model_source} is not a GCS path starting with {GCS_URI_PREFIX}."
)
if not model_destination.startswith(GCS_URI_PREFIX):
raise ValueError(
f"{model_destination} is not a GCS path starting with {GCS_URI_PREFIX}."
)
model_source = f"{model_source}/{model_id}"
model_destination = f"{model_destination}/{model_id}"
print("Copying model artifact from ", model_source, " to ", model_destination)
subprocess.check_output([
"gcloud",
"storage",
"cp",
"-r",
model_source,
model_destination,
])
def get_quota(project_id: str, region: str, resource_id: str) -> int:
"""Returns the quota for a resource in a region.
Args:
project_id: The project id.
region: The region.
resource_id: The resource id.
Returns:
The quota for the resource in the region. Returns -1 if can not figure out
the quota.
Raises:
RuntimeError: If the command to get quota fails.
"""
service_endpoint = "aiplatform.googleapis.com"
command = (
"gcloud alpha services quota list"
f" --service={service_endpoint} --consumer=projects/{project_id}"
f" --filter='{service_endpoint}/{resource_id}' --format=json"
)
process = subprocess.run(
command, shell=True, capture_output=True, text=True, check=True
)
if process.returncode == 0:
quota_data = json.loads(process.stdout)
else:
raise RuntimeError(f"Error fetching quota data: {process.stderr}")
if not quota_data or "consumerQuotaLimits" not in quota_data[0]:
return -1
if (
not quota_data[0]["consumerQuotaLimits"]
or "quotaBuckets" not in quota_data[0]["consumerQuotaLimits"][0]
):
return -1
all_regions_data = quota_data[0]["consumerQuotaLimits"][0]["quotaBuckets"]
for region_data in all_regions_data:
if (
region_data.get("dimensions")
and region_data["dimensions"]["region"] == region
):
if "effectiveLimit" in region_data:
return int(region_data["effectiveLimit"])
else:
return 0
return -1
def get_resource_id(
accelerator_type: str,
is_for_training: bool,
is_restricted_image: bool = False,
is_dynamic_workload_scheduler: bool = False,
) -> str:
"""Returns the resource id for a given accelerator type and the use case.
Args:
accelerator_type: The accelerator type.
is_for_training: Whether the resource is used for training. Set false for
serving use case.
is_restricted_image: Whether the image is hosted in `vertex-ai-restricted`.
is_dynamic_workload_scheduler: Whether the resource is used with Dynamic
Workload Scheduler.
Returns:
The resource id.
"""
accelerator_suffix_map = {
"NVIDIA_TESLA_V100": "nvidia_v100_gpus",
"NVIDIA_TESLA_P100": "nvidia_p100_gpus",
"NVIDIA_L4": "nvidia_l4_gpus",
"NVIDIA_TESLA_A100": "nvidia_a100_gpus",
"NVIDIA_A100_80GB": "nvidia_a100_80gb_gpus",
"NVIDIA_H100_80GB": "nvidia_h100_gpus",
"NVIDIA_TESLA_T4": "nvidia_t4_gpus",
"TPU_V5e": "tpu_v5e",
"TPU_V3": "tpu_v3",
}
default_training_accelerator_map = {
key: f"custom_model_training_{accelerator_suffix_map[key]}"
for key in accelerator_suffix_map
}
dws_training_accelerator_map = {
key: f"custom_model_training_preemptible_{accelerator_suffix_map[key]}"
for key in accelerator_suffix_map
}
restricted_image_training_accelerator_map = {
"NVIDIA_A100_80GB": "restricted_image_training_nvidia_a100_80gb_gpus",
}
serving_accelerator_map = {
key: f"custom_model_serving_{accelerator_suffix_map[key]}"
for key in accelerator_suffix_map
}
if is_for_training:
if is_restricted_image and is_dynamic_workload_scheduler:
raise ValueError(
"Dynamic Workload Scheduler does not work for restricted image"
" training."
)
training_accelerator_map = (
restricted_image_training_accelerator_map
if is_restricted_image
else default_training_accelerator_map
)
if accelerator_type in training_accelerator_map:
if is_dynamic_workload_scheduler:
return dws_training_accelerator_map[accelerator_type]
else:
return training_accelerator_map[accelerator_type]
else:
raise ValueError(
f"Could not find accelerator type: {accelerator_type} for training."
)
else:
if is_dynamic_workload_scheduler:
raise ValueError("Dynamic Workload Scheduler does not work for serving.")
if accelerator_type in serving_accelerator_map:
return serving_accelerator_map[accelerator_type]
else:
raise ValueError(
f"Could not find accelerator type: {accelerator_type} for serving."
)
def check_quota(
project_id: str,
region: str,
accelerator_type: str,
accelerator_count: int,
is_for_training: bool,
is_restricted_image: bool = False,
is_dynamic_workload_scheduler: bool = False,
):
"""Checks if the project and the region has the required quota."""
resource_id = get_resource_id(
accelerator_type,
is_for_training=is_for_training,
is_restricted_image=is_restricted_image,
is_dynamic_workload_scheduler=is_dynamic_workload_scheduler,
)
quota = get_quota(project_id, region, resource_id)
quota_request_instruction = (
"Either use "
"a different region or request additional quota. Follow "
"instructions here "
"https://cloud.google.com/docs/quotas/view-manage#requesting_higher_quota"
" to check quota in a region or request additional quota for "
"your project."
)
if quota == -1:
raise ValueError(
f"Quota not found for: {resource_id} in {region}."
f" {quota_request_instruction}"
)
if quota < accelerator_count:
raise ValueError(
f"Quota not enough for {resource_id} in {region}: {quota} <"
f" {accelerator_count}. {quota_request_instruction}"
)
@@ -0,0 +1,544 @@
"""Functions for dataset validation.
This tool is used to validate the dataset against the given template.
"""
import json
import multiprocessing
import os
import subprocess
from typing import Any, Callable, Dict, Union
from absl import logging
import accelerate
import datasets
import transformers
GCS_URI_PREFIX = "gs://"
GCSFUSE_URI_PREFIX = "/gcs/"
LOCAL_BASE_MODEL_DIR = "/tmp/base_model_dir"
LOCAL_TEMPLATE_DIR = "/tmp/template_dir"
_TEMPLATE_DIRNAME = "templates"
_VERTEX_AI_SAMPLES_GITHUB_REPO_NAME = "vertex-ai-samples"
_VERTEX_AI_SAMPLES_GITHUB_TEMPLATE_DIR = (
"community-content/vertex_model_garden/model_oss/peft/train/vmg/templates"
)
_MODELS_REQUIRING_PAD_TOKEN = ("llama", "falcon", "mistral", "mixtral")
_MODELS_REQUIRING_EOS_TOEKN = ("gemma-2b", "gemma-7b")
_DESCRIPTION_KEY = "description"
_SOURCE_KEY = "source"
_PROMPT_INPUT_KEY = "prompt_input"
_PROMPT_NO_INPUT_KEY = "prompt_no_input"
_RESPONSE_SEPARATOR = "response_separator"
_INSTRUCTION_SEPARATOR = "instruction_separator"
_CHAT_TEMPLATE_KEY = "chat_template"
_KNOWN_KEYS = (
_DESCRIPTION_KEY,
_SOURCE_KEY,
_PROMPT_INPUT_KEY,
_PROMPT_NO_INPUT_KEY,
_RESPONSE_SEPARATOR,
_INSTRUCTION_SEPARATOR,
_CHAT_TEMPLATE_KEY,
)
def is_gcs_path(input_path: str) -> bool:
"""Checks if the input path is a Google Cloud Storage (GCS) path.
Args:
input_path: The input path to be checked.
Returns:
True if the input path is a GCS path, False otherwise.
"""
return input_path is not None and input_path.startswith(GCS_URI_PREFIX)
def force_gcs_fuse_path(gcs_uri: str) -> str:
"""Converts gs:// uris to their /gcs/ equivalents. No-op for other uris.
Args:
gcs_uri: The GCS URI to convert.
Returns:
The converted GCS URI.
"""
if is_gcs_path(gcs_uri):
return GCSFUSE_URI_PREFIX + gcs_uri[len(GCS_URI_PREFIX) :]
else:
return gcs_uri
def download_gcs_uri_to_local(
gcs_uri: str, destination_dir: str = LOCAL_BASE_MODEL_DIR
) -> str:
"""Downloads GCS URI to local.
If GCS URI is a directory, gs://some/folder is downloaded to
/destination_dir/folder. If GCS URI is a file, gs://some/file is downloaded to
/destination_dir/file.
Args:
gcs_uri: GCS URI to download.
destination_dir: Local directory directory.
Returns:
Local path to target folder/file.
"""
target = os.path.join(
destination_dir,
os.path.basename(os.path.normpath(gcs_uri)),
)
if os.path.exists(target):
logging.info("File %s already exists.", target)
return target
if accelerate.PartialState().is_local_main_process:
logging.info(
"Downloading file(s) from %s to %s...", gcs_uri, destination_dir
)
if not os.path.exists(destination_dir):
os.mkdir(destination_dir)
subprocess.check_output([
"gsutil",
"-m",
"cp",
"-r",
gcs_uri,
destination_dir,
])
logging.info("Downloaded file(s) from %s to %s.", gcs_uri, destination_dir)
# Make sure ALL processes process to next step after data downloading is done.
# It matters for the main process to wait for other processes as well.
accelerate.PartialState().wait_for_everyone()
return target
def get_template(template_path: str) -> Dict[str, str]:
"""Gets the template dictionary given the file path.
Args:
template_path: Path to the template file.
Returns:
A dictionary of the template.
Raises:
ValueError: If the template file does not exist or contains unknown keys.
"""
if is_gcs_path(template_path):
template_path = force_gcs_fuse_path(template_path)
elif not os.path.isfile(template_path):
template_path = os.path.join(
os.path.dirname(__file__),
_TEMPLATE_DIRNAME,
template_path + ".json",
)
if not os.path.isfile(template_path):
raise ValueError(f"Template file {template_path} does not exist.")
with open(template_path, "r") as f:
template_json: dict[str, str] = json.load(f)
for key in template_json:
if key not in _KNOWN_KEYS:
raise ValueError(f"Unknown key {key} in template {template_path}.")
return template_json
def get_response_separator(template_json: Dict[str, str]) -> Union[str, None]:
return template_json.get(_RESPONSE_SEPARATOR, None)
def get_instruction_separator(
template_json: Dict[str, str],
) -> Union[str, None]:
return template_json.get(_INSTRUCTION_SEPARATOR, None)
def _format_template_fn(
template: str,
input_column: str,
tokenizer: transformers.PreTrainedTokenizer | None = None,
) -> Callable[[Dict[str, str]], Dict[str, str]]:
"""Formats a dataset example according to a template.
Args:
template: Name of the JSON template file under `templates/` or GCS path to
the template file.
input_column: The input column in the dataset to be used or updated by the
template. If it does not exist, the template's `prompt_no_input` will be
used, and the input_column will be created.
tokenizer: The tokenizer to use for chat_template templates.
Returns:
A function that formats data according to the template.
"""
template_json = get_template(template)
if _CHAT_TEMPLATE_KEY not in template_json:
def format_fn(example: Dict[str, str]) -> Dict[str, str]:
format_dict = {key: value for key, value in example.items()}
if format_dict.get(input_column):
format_str = template_json[_PROMPT_INPUT_KEY]
elif _PROMPT_NO_INPUT_KEY in template_json:
format_str = template_json[_PROMPT_NO_INPUT_KEY]
else:
raise KeyError(
f"The template {os.path.basename(template)} does not contain"
f" {_PROMPT_INPUT_KEY} or {_PROMPT_NO_INPUT_KEY} key."
)
try:
return {input_column: format_str.format(**format_dict)}
except KeyError as e:
raise KeyError(
f"The template {os.path.basename(template)} contains a key {e} in"
f" {_PROMPT_INPUT_KEY} or {_PROMPT_NO_INPUT_KEY} that does not"
" exist in the dataset example. The dataset example looks like"
f" {format_dict}."
) from e
return format_fn
elif (
_PROMPT_INPUT_KEY in template_json
or _PROMPT_NO_INPUT_KEY in template_json
):
raise ValueError(
f"chat_template templates do not support {_PROMPT_INPUT_KEY} or"
f" {_PROMPT_NO_INPUT_KEY} templates."
)
else:
if tokenizer is None:
raise ValueError("A tokenizer is required for chat_template templates.")
# Assign HuggingFace jinja template.
tokenizer.chat_template = template_json[_CHAT_TEMPLATE_KEY]
def format_fn(example: Dict[str, str]) -> Dict[str, str]:
try:
return {
input_column: tokenizer.apply_chat_template(
example[input_column],
tokenize=False,
add_generation_prompt=False,
)
}
except KeyError as e:
raise KeyError(
f"The template {os.path.basename(template)} contains a key {e} in"
f" {_CHAT_TEMPLATE_KEY} that does not exist in the dataset example."
) from e
return format_fn
def _get_split_string(
split: str,
dataset_percent: int | None = None,
dataset_k_rows: int | None = None,
) -> str:
"""Gets the formatted split string for the dataset.
This is used to format the split string as per
https://huggingface.co/docs/datasets/v2.21.0/loading#slice-splits. Also, this
function will only be used to load the partial dataset for validating the
dataset against the template.
Args:
split: Split of the dataset.
dataset_percent: The percentage of the dataset to load.
dataset_k_rows: The top k sequences to load from the dataset.
Returns:
A formatted split string.
"""
# Validate the dataset_percent and dataset_k_rows values.
if dataset_percent and dataset_k_rows:
raise ValueError(
"You can set either validate_percentage_of_dataset or"
" validate_k_rows_of_dataset, but not both."
)
if dataset_percent:
logging.info("Loading %d percent of the dataset...", dataset_percent)
return f"{split}[:{dataset_percent}%]"
if dataset_k_rows:
logging.info("Loading top %d rows of the dataset...", dataset_k_rows)
return f"{split}[:{dataset_k_rows}]"
return split
def _github_template_path(template: str) -> str:
"""Generates the path to the template in the Vertex AI Samples GitHub repo.
Args:
template: Name of the template.
Returns:
The path to the template in the Vertex AI Samples GitHub repo.
"""
# vertex-ai-samples directory may lie under separate directory depending on
# the scratch_dir parameter in the notebook execution environment.
vertex_ai_samples_abs_path = os.getcwd().split(
_VERTEX_AI_SAMPLES_GITHUB_REPO_NAME
)[0]
return os.path.join(
vertex_ai_samples_abs_path,
_VERTEX_AI_SAMPLES_GITHUB_REPO_NAME,
_VERTEX_AI_SAMPLES_GITHUB_TEMPLATE_DIR,
template + ".json",
)
def _get_dataset(
dataset_name: str,
split: str,
num_proc: int | None = None,
) -> datasets.DatasetDict:
"""Gets a dataset.
Args:
dataset_name: Name of the dataset or path to a custom dataset.
split: Split of the dataset.
num_proc: Number of processors to use.
Returns:
A dataset.
"""
dataset_name = force_gcs_fuse_path(dataset_name)
if os.path.isfile(dataset_name):
# Custom dataset.
return datasets.load_dataset(
"json",
data_files=[dataset_name],
split=split,
num_proc=num_proc,
)
# HF dataset.
return datasets.load_dataset(dataset_name, split=split, num_proc=num_proc)
def should_add_pad_token(model_id: str) -> bool:
"""Returns whether the model requires adding a special pad token.
Args:
model_id: The name of the model.
Returns:
True if the model requires adding a special pad token, False otherwise.
"""
return any(s.lower() in model_id.lower() for s in _MODELS_REQUIRING_PAD_TOKEN)
def should_add_eos_token(model_id: str) -> bool:
"""Returns whether the model requires adding a special eos token.
Args:
model_id: The name of the model.
Returns:
True if the model requires adding a special eos token, False otherwise.
"""
return any(m in model_id for m in _MODELS_REQUIRING_EOS_TOEKN)
def load_tokenizer(
pretrained_model_id: str,
padding_side: str | None = None,
access_token: str | None = None,
) -> transformers.AutoTokenizer:
"""Loads tokenizer based on `pretrained_model_id`.
Args:
pretrained_model_id: The name of the pretrained model.
padding_side: The side to pad the input on.
access_token: The access token to use for the tokenizer.
Returns:
The tokenizer.
"""
tokenizer_kwargs = {}
if should_add_eos_token(pretrained_model_id):
tokenizer_kwargs["add_eos_token"] = True
if padding_side:
tokenizer_kwargs["padding_side"] = padding_side
with accelerate.PartialState().local_main_process_first():
tokenizer = transformers.AutoTokenizer.from_pretrained(
pretrained_model_id,
trust_remote_code=False,
use_fast=True,
token=access_token,
**tokenizer_kwargs,
)
if should_add_pad_token(pretrained_model_id):
tokenizer.add_special_tokens({"pad_token": "[PAD]"})
return tokenizer
def get_filtered_dataset(
dataset: Any,
input_column: str,
max_seq_length: int,
tokenizer: transformers.PreTrainedTokenizer,
) -> Any:
"""Returns the dataset by removing examples that are longer than max_seq_length.
Args:
dataset: The dataset to filter.
input_column: The input column in the dataset to be used.
max_seq_length: The maximum sequence length.
tokenizer: The tokenizer.
"""
actual_dataset_length = len(dataset)
filtered_dataset = dataset.filter(
lambda x: len(tokenizer(x[input_column])["input_ids"]) <= max_seq_length
)
filtered_dataset_length = len(filtered_dataset)
if actual_dataset_length != filtered_dataset_length:
examples_removed_percent = (
(actual_dataset_length - filtered_dataset_length)
* 100
/ actual_dataset_length
)
logging.info(
"(%.2f%%) of examples token length is <= max-seq-length(%d); (%.2f%%) >"
" max-seq-length. Filtering out %d example(s) which are longer than"
" max-seq-length.",
100 - examples_removed_percent,
max_seq_length,
examples_removed_percent,
actual_dataset_length - filtered_dataset_length,
)
return filtered_dataset
def load_dataset_with_template(
dataset_name: str,
split: str,
input_column: str,
template: str = None,
tokenizer: transformers.PreTrainedTokenizer | None = None,
) -> Any:
"""Loads dataset with templates.
Args:
dataset_name: Name of the dataset or path to a custom dataset.
split: Split of the dataset.
input_column: The input column in the dataset to be used or updaded by the
template. If it does not exist, the template's `prompt_no_input` will be
used, and the input_column will be created.
template: Name of the JSON template file under `templates/` or GCS path to
the template file.
tokenizer: The tokenizer to use for chat_template templates.
Returns:
A dataset compatible with the template.
"""
dataset = _get_dataset(dataset_name, split=split)
if template:
dataset = dataset.map(
_format_template_fn(
template,
input_column=input_column,
tokenizer=tokenizer,
)
)
return dataset
def validate_dataset_with_template(
dataset_name: str,
split: str,
input_column: str,
template: str,
tokenizer: transformers.PreTrainedTokenizer | None = None,
max_seq_length: int | None = None,
use_multiprocessing: bool = False,
validate_percentage_of_dataset: int | None = None,
validate_k_rows_of_dataset: int | None = None,
) -> Any:
"""Validates dataset with templates.
This function will be used to load the dataset and validate it against the
template. In case of validation, we also allow the users to load the dataset
partially by allowing them to read x% or top k rows of the dataset. To
validate the dataset, the template file must be available in the GCS bucket
and the dataset must be available either in the GCS bucket or Hugging Face.
Args:
dataset_name: Name of the dataset or path to a custom dataset.
split: Split of the dataset.
input_column: The input column in the dataset to be used or updaded by the
template. If it does not exist, the template's `prompt_no_input` will be
used, and the input_column will be created.
template: Name of the JSON template file under `templates/` or GCS path to
the template file.
tokenizer: The tokenizer to use for chat_template templates.
max_seq_length: The maximum sequence length.
use_multiprocessing: If True, it will use multiprocessing to load the
dataset.
validate_percentage_of_dataset: The percentage of the dataset to load.
validate_k_rows_of_dataset: The top k sequences to load from the dataset.
Returns:
None if the validation is successful, otherwise returns the error message.
"""
if not template:
raise ValueError("template is required for validate_dataset.")
if not dataset_name:
raise ValueError("dataset_name is empty.")
if not split:
raise ValueError("split is empty.")
split = _get_split_string(
split,
validate_percentage_of_dataset,
validate_k_rows_of_dataset,
)
num_proc = multiprocessing.cpu_count() if use_multiprocessing else 1
# gcsfuse cannot be used from the notebook runtime env. Hence, we have
# to download dataset and template from gcs to local.
if is_gcs_path(dataset_name):
dataset_name = download_gcs_uri_to_local(dataset_name, LOCAL_BASE_MODEL_DIR)
if is_gcs_path(template):
template_path = download_gcs_uri_to_local(template, LOCAL_TEMPLATE_DIR)
elif os.path.isfile(_github_template_path(template)):
template_path = _github_template_path(template)
else:
raise ValueError(
f"Template file {template} does not exist. To validate the"
" dataset, please provide a valid GCS path for the template or a valid"
" template name from"
f" https://github.com/GoogleCloudPlatform/{_VERTEX_AI_SAMPLES_GITHUB_REPO_NAME}/tree/main/{_VERTEX_AI_SAMPLES_GITHUB_TEMPLATE_DIR}."
)
dataset = _get_dataset(dataset_name, split, num_proc).map(
_format_template_fn(
template_path,
input_column=input_column,
tokenizer=tokenizer,
)
)
if tokenizer is not None:
get_filtered_dataset(
dataset=dataset,
input_column=input_column,
max_seq_length=max_seq_length,
tokenizer=tokenizer,
)
print(
"Dataset {} is compatible with the {} template.".format(
os.path.basename(dataset_name), os.path.basename(template)
)
)
@@ -1,142 +0,0 @@
"""Causal language modeling with LoRA models."""
# pylint: disable=g-importing-member
from datasets import load_dataset
from peft import get_peft_model
from peft import LoraConfig
import torch
from torch import nn
import transformers
from transformers import AutoModelForCausalLM
from transformers import AutoTokenizer
from transformers import BitsAndBytesConfig
from transformers import TrainingArguments
from util import constants
def finetune_causal_language_modeling(
pretrained_model_id: str,
dataset_name: str,
output_dir: str,
precision_mode: str = None,
lora_rank: int = 16,
lora_alpha: int = 32,
lora_dropout: float = 0.05,
warmup_steps: int = 10,
max_steps: int = 10,
learning_rate: float = 2e-4,
local_pretrained_model_id: str = None,
) -> None:
"""Finetunes causal language modelings."""
if precision_mode == constants.PRECISION_MODE_32:
model = AutoModelForCausalLM.from_pretrained(
local_pretrained_model_id
if local_pretrained_model_id
else pretrained_model_id,
torch_dtype=torch.float32,
device_map="auto",
)
elif precision_mode == constants.PRECISION_MODE_16:
model = AutoModelForCausalLM.from_pretrained(
local_pretrained_model_id
if local_pretrained_model_id
else pretrained_model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)
elif precision_mode == constants.PRECISION_MODE_8:
quantization_config = BitsAndBytesConfig(
load_in_8bit=True, int8_threshold=0
)
model = AutoModelForCausalLM.from_pretrained(
local_pretrained_model_id
if local_pretrained_model_id
else pretrained_model_id,
torch_dtype=torch.float16,
device_map="auto",
quantization_config=quantization_config,
)
else:
quantization_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16,
)
model = AutoModelForCausalLM.from_pretrained(
local_pretrained_model_id
if local_pretrained_model_id
else pretrained_model_id,
device_map="auto",
torch_dtype=torch.bfloat16,
quantization_config=quantization_config,
)
tokenizer = AutoTokenizer.from_pretrained(
local_pretrained_model_id
if local_pretrained_model_id
else pretrained_model_id
)
if "llama" in pretrained_model_id:
tokenizer.pad_token = "[PAD]"
for param in model.parameters():
# Freezes the model - train adapters later.
param.requires_grad = False
if param.ndim == 1:
# Casts the small parameters (e.g. layernorm) to fp32 for stability.
param.data = param.data.to(torch.float32)
# Reduces the number of stored activations.
model.gradient_checkpointing_enable()
model.enable_input_require_grads()
class CastOutputToFloat(nn.Sequential):
def forward(self, x):
return super().forward(x).to(torch.float32)
model.lm_head = CastOutputToFloat(model.lm_head)
config = LoraConfig(
r=lora_rank,
lora_alpha=lora_alpha,
target_modules=["q_proj", "v_proj"],
lora_dropout=lora_dropout,
bias="none",
task_type="CAUSAL_LM",
)
model = get_peft_model(model, config)
model.print_trainable_parameters()
data = load_dataset(dataset_name)
data = data.map(
lambda samples: tokenizer(samples["quote"]),
batched=True,
)
trainer = transformers.Trainer(
model=model,
train_dataset=data["train"],
args=TrainingArguments(
per_device_train_batch_size=4,
gradient_accumulation_steps=4,
warmup_steps=warmup_steps,
max_steps=max_steps,
learning_rate=learning_rate,
fp16=True,
logging_steps=1,
output_dir=output_dir,
ddp_find_unused_parameters=False,
),
data_collator=transformers.DataCollatorForLanguageModeling(
tokenizer,
mlm=False,
),
)
# Silence the warnings. Please re-enable for inference!
model.config.use_cache = False
trainer.train()
model.save_pretrained(output_dir)
@@ -0,0 +1,28 @@
# Base on pytorch-cuda image.
FROM pytorch/pytorch:2.0.0-cuda11.7-cudnn8-devel
# Install tools.
ENV DEBIAN_FRONTEND=noninteractive
RUN apt-get update
RUN apt-get install -y --no-install-recommends apt-utils
RUN apt-get install -y --no-install-recommends curl
RUN apt-get install -y --no-install-recommends wget
RUN apt-get install -y --no-install-recommends git
# Install libraries.
ENV PIP_ROOT_USER_ACTION=ignore
RUN python3 -m pip install --upgrade pip
RUN pip install tokenizers==0.13.3
RUN pip install accelerate==0.21.0
RUN pip install sentencepiece==0.1.99
RUN pip install datasets==2.14.4
RUN pip install protobuf==4.24.1
# Install transformers
RUN git clone https://github.com/huggingface/transformers.git
WORKDIR transformers
# Pin the commit to add-code-llama 08/25/2023
RUN git reset --hard 015f8e110d270a0ad42de4ae5b98198d69eb1964
RUN pip install -e .
ENTRYPOINT ["python","src/transformers/models/llama/convert_llama_weights_to_hf.py"]
@@ -0,0 +1,22 @@
# Dockerfile for Language Model Conversion.
#
# To build:
# docker build -f model_oss/peft/dockerfile/conversion.Dockerfile . -t ${YOUR_IMAGE_TAG}
#
# To push to gcr:
# docker tag ${YOUR_IMAGE_TAG} gcr.io/${YOUR_PROJECT}/${YOUR_IMAGE_TAG}
# docker push gcr.io/${YOUR_PROJECT}/${YOUR_IMAGE_TAG}
FROM tensorflow/build:2.14-python3.8
RUN git clone https://github.com/facebookresearch/llama-recipes.git && \
cd llama-recipes && \
pip install -r requirements.txt && \
pip freeze | grep transformers && \
git clone https://github.com/huggingface/transformers.git && \
cd transformers && \
pip install protobuf
WORKDIR /llama-recipes/transformers
ENTRYPOINT ["python","src/transformers/models/llama/convert_llama_weights_to_hf.py"]
@@ -7,39 +7,40 @@
# docker tag ${YOUR_IMAGE_TAG} gcr.io/${YOUR_PROJECT}/${YOUR_IMAGE_TAG}
# docker push gcr.io/${YOUR_PROJECT}/${YOUR_IMAGE_TAG}
FROM pytorch/torchserve:0.7.0-gpu
FROM pytorch/torchserve:0.11.0-gpu
USER root
ENV infer_port=7080
ENV mng_port=7081
ENV model_name="peft_serving"
ENV INFER_PORT=7080
ENV MNG_PORT=7081
ENV MODEL="peft_serving"
ENV PATH="/home/model-server/:${PATH}"
RUN apt-get update && apt-get install -y --no-install-recommends \
RUN apt-get update && apt-get -y upgrade && apt-get install -y --no-install-recommends \
curl \
wget \
vim \
git \
git-lfs
RUN git lfs install
RUN apt-get autoremove -y
# Install libraries.
ENV PIP_ROOT_USER_ACTION=ignore
RUN python3 -m pip install --upgrade pip
RUN pip install --upgrade torch==2.0.1
RUN pip install --upgrade torch==2.0.1 --index-url https://download.pytorch.org/whl/cu118
RUN pip install torchvision==0.15.2
RUN pip install tokenizers==0.13.3
RUN pip install accelerate==0.21.0
RUN pip install sentencepiece==0.1.99
RUN pip install grpcio-status==1.33.2
RUN pip install protobuf==3.19.6
RUN python3 -m pip install --no-cache-dir git+https://github.com/huggingface/peft.git
RUN pip install peft==0.5.0
RUN pip install datasets==2.14.4
RUN pip install triton==2.0.0.dev20221120
RUN pip install triton==3.0.0
RUN pip install xformers==0.0.20
RUN pip install google-cloud-storage==2.7.0
RUN pip install absl-py==1.4.0
RUN pip install google-cloud-storage
RUN pip install absl-py
RUN pip install scipy==1.10.1
RUN pip install evaluate==0.4.0
RUN pip install scikit-learn==1.2.2
@@ -47,52 +48,43 @@ RUN pip install loralib==0.1.1
RUN pip install bitsandbytes==0.39.0
RUN pip install trl==0.4.4
RUN pip install einops==0.6.1
# Install diffusers from source.
RUN git clone --depth 1 --branch v0.16.1 https://github.com/huggingface/diffusers.git
WORKDIR diffusers
RUN pip install -e .
WORKDIR /home/model-server
# Install transformers from source.
RUN git clone --depth 1 --branch v4.31.0 https://github.com/huggingface/transformers.git
# The patch is used to change the transformers loading model behavior:
# 1) For models on Huggingface hub: if the model has multiple shards, each shard
# will be downloaded separately and get deleted after loading to GPU.
# 2) For models on local disk: if a model bin file is actually a text file
# recording a GCS path, the model file will be downloaded and get deleted
# after loading to GPU.
COPY model_oss/peft/hf_transformers_lazy_download.patch /home/model-server/hf_transformers_lazy_download.patch
WORKDIR transformers
RUN git apply /home/model-server/hf_transformers_lazy_download.patch
RUN pip install -e .
WORKDIR /home/model-server
RUN pip install optimum==1.13.2
RUN pip install auto-gptq==0.4.2
RUN pip install https://github.com/casper-hansen/AutoAWQ/releases/download/v0.1.7/autoawq-0.1.7+cu118-cp39-cp39-linux_x86_64.whl
RUN pip install diffusers==0.27.2
RUN pip install tiktoken==0.6.0
RUn pip install git+https://github.com/huggingface/transformers.git@76fa17c1663a0efeca7208c20579833365584889
RUN pip install pynvml==11.4.0
RUN pip install -i https://test.pypi.org/simple/ bitsandbytes
# Copy license.
WORKDIR /home/model-server
RUN wget https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/main/LICENSE
# Copy model artifacts.
COPY model_oss/peft/handler.py /home/model-server/handler.py
COPY model_oss/peft/config.properties /home/model-server/config.properties
COPY model_oss/util/ /home/model-server/util/
COPY model_oss/util/pytorch_startup_prober.sh /model_garden/scripts/pytorch_startup_prober.sh
ENV PYTHONPATH /home/model-server/
# Expose ports.
EXPOSE ${infer_port}
EXPOSE ${mng_port}
EXPOSE ${INFER_PORT}
EXPOSE ${MNG_PORT}
# Set environments.
ENV TASK "causal-language-modeling-lora"
ENV MODEL_ID "openlm-research/open_llama_7b"
ENV BASE_MODEL_ID ""
ENV MODEL_ID ""
ENV PRECISION_LOADING_MODE "float16"
ENV FINETUNED_LORA_MODEL_PATH ""
ENV TRUST_REMOTE_CODE ""
# Archive model artifacts and dependencies.
# Do not set --model-file and --serialized-file because model and checkpoint
# will be dynamically loaded in handler.py.
RUN torch-model-archiver \
--model-name=${model_name} \
--model-name=${MODEL} \
--version=1.0 \
--handler=/home/model-server/handler.py \
--runtime=python3 \
@@ -103,5 +95,5 @@ RUN torch-model-archiver \
# Run Torchserve HTTP serve to respond to prediction requests.
CMD ["torchserve", "--start", \
"--ts-config", "/home/model-server/config.properties", \
"--models", "${model_name}=${model_name}.mar", \
"--models", "${MODEL}=${MODEL}.mar", \
"--model-store", "/home/model-server/model-store"]
@@ -1,111 +0,0 @@
# Dockerfile for PEFT Training.
#
# To build:
# docker build -f model_oss/peft/dockerfile/train.Dockerfile . -t ${YOUR_IMAGE_TAG}
#
# To push to gcr:
# docker tag ${YOUR_IMAGE_TAG} gcr.io/${YOUR_PROJECT}/${YOUR_IMAGE_TAG}
# docker push gcr.io/${YOUR_PROJECT}/${YOUR_IMAGE_TAG}
# Builds GPU docker image of PyTorch
# Uses multi-staged approach to reduce size
# Stage 1
# Use base conda image to reduce time
FROM continuumio/miniconda3:latest AS compile-image
# Specify py version
ENV PYTHON_VERSION=3.8
# Install apt libs - copied from https://github.com/huggingface/accelerate/blob/main/docker/accelerate-gpu/Dockerfile
RUN apt-get update && \
apt-get install -y curl git wget software-properties-common git-lfs && \
apt-get clean && \
rm -rf /var/lib/apt/lists*
# Install audio-related libraries
RUN apt-get update && \
apt install -y ffmpeg
RUN apt install -y libsndfile1-dev
RUN git lfs install
# Create our conda env - copied from https://github.com/huggingface/accelerate/blob/main/docker/accelerate-gpu/Dockerfile
RUN conda create --name peft python=${PYTHON_VERSION} ipython jupyter pip
RUN python3 -m pip install --no-cache-dir --upgrade pip
# Below is copied from https://github.com/huggingface/accelerate/blob/main/docker/accelerate-gpu/Dockerfile
# We don't install pytorch here yet since CUDA isn't available
# instead we use the direct torch wheel
ENV PATH /opt/conda/envs/peft/bin:$PATH
# Activate our bash shell
RUN chsh -s /bin/bash
SHELL ["/bin/bash", "-c"]
# Activate the conda env and install transformers + accelerate from source
RUN source activate peft
RUN python3 -m pip install --no-cache-dir git+https://github.com/huggingface/transformers
RUN python3 -m pip install --no-cache-dir git+https://github.com/huggingface/accelerate
RUN python3 -m pip install --no-cache-dir git+https://github.com/huggingface/peft#egg=peft[test]
RUN python3 -m pip install --no-cache-dir bitsandbytes
# Stage 2
FROM nvidia/cuda:11.2.2-cudnn8-devel-ubuntu20.04 AS build-image
COPY --from=compile-image /opt/conda /opt/conda
ENV PATH /opt/conda/bin:$PATH
# Install apt libs
RUN apt-get update && \
apt-get install -y curl git wget vim && \
apt-get clean && \
rm -rf /var/lib/apt/lists*
# Copy license.
RUN wget https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/main/LICENSE
RUN echo "source activate peft" >> ~/.profile
# Install libraries.
RUN pip install --upgrade torch==2.0.1
RUN pip install torchvision==0.15.2
RUN pip install git+https://github.com/huggingface/transformers@de9255de27abfcae4a1f816b904915f0b1e23cd9
RUN pip install transformers -U
RUN pip install accelerate==0.21.0
RUN pip install sentencepiece==0.1.99
RUN pip install grpcio-status==1.33.2
RUN pip install protobuf==3.19.6
RUN python3 -m pip install --no-cache-dir git+https://github.com/huggingface/peft.git
RUN pip install datasets==2.9.0
RUN pip install triton==2.0.0.dev20221120
RUN pip install xformers==0.0.20
RUN pip install Jinja2==3.1.2
RUN pip install ftfy==6.1.1
RUN pip install cloudml-hypertune==0.1.0.dev6
RUN pip install tensorboard==2.12.0
RUN pip install scipy==1.10.1
RUN pip install evaluate==0.4.0
RUN pip install scikit-learn==1.2.2
RUN pip install loralib==0.1.1
RUN pip install bitsandbytes==0.39.0
RUN pip install trl==0.4.4
RUN pip install einops==0.6.1
RUN pip install google-cloud-storage==2.7.0
RUN git clone --depth 1 --branch v0.16.1 https://github.com/huggingface/diffusers.git
WORKDIR diffusers
RUN pip install -e .
# Switch to diffusers examples folder.
WORKDIR examples
# NOTE: use 'sed' to modify train_text_to_image_lora.py to
# fix the bug for accelerator.
RUN sed -i \
"s#logging_dir=logging_dir#project_dir=logging_dir#g" \
text_to_image/train_text_to_image_lora.py
# Config accelerate.
RUN mkdir -p ./vertex_vision_model_garden_peft/
COPY model_oss/peft/train.sh ./vertex_vision_model_garden_peft/train.sh
COPY model_oss/peft/*.py ./vertex_vision_model_garden_peft/
COPY model_oss/util /diffusers/examples/util
ENV PYTHONPATH /diffusers/examples/
# Generate accelerate config at the beginning of docker run.
ENTRYPOINT ["python3", "vertex_vision_model_garden_peft/main.py"]
@@ -72,15 +72,39 @@ class PeftHandler(BaseHandler):
"PRECISION_LOADING_MODE", constants.PRECISION_MODE_16
)
self.task = os.environ.get("TASK", CAUSAL_LANGUAGE_MODELING_LORA)
self.base_model_id = os.environ.get("BASE_MODEL_ID", None)
self.model_id = self.base_model_id
if not self.base_model_id:
self.model_id = os.environ.get("MODEL_ID", "")
trust_remote_code = os.environ.get("TRUST_REMOTE_CODE", None)
if trust_remote_code == "false":
self.trust_remote_code = False
else:
self.trust_remote_code = True
# If present, the path of the model in the container.
aip_storage_dir = os.environ.get("AIP_STORAGE_DIR", None)
# If present, the URI of the model in a google owned GCS bucket.
aip_storage_uri = os.environ.get("AIP_STORAGE_URI", None)
model_id = os.environ.get("MODEL_ID", None)
base_model_id = os.environ.get("BASE_MODEL_ID", None)
self.model_id = None
if aip_storage_dir:
self.model_id = aip_storage_dir
logging.info(f"Loaded base model from AIP_STORAGE_DIR: {self.model_id}.")
elif aip_storage_uri:
self.model_id = aip_storage_uri
logging.info(f"Loaded base model from AIP_STORAGE_URI: {self.model_id}.")
elif model_id:
self.model_id = model_id
logging.info(f"Loaded base model from MODEL_ID: {self.model_id}.")
elif base_model_id:
# Note: BASE_MODEL_ID has been unified with MODEL_ID.
# MODEL_ID should be used whenever possible.
self.model_id = base_model_id
logging.info(f"Loaded base model from BASE_MODEL_ID: {self.model_id}.")
self.quantization = os.environ.get("QUANTIZATION", None)
logging.info(f"Load base model id from MODEL_ID:{self.model_id}.")
if not self.model_id:
self.model_id = os.environ.get("AIP_STORAGE_URI", "")
logging.info(f"Load base model id from AIP_STORAGE_URI: {self.model_id}.")
if not self.model_id:
raise ValueError("Base model id is must be set.")
if fileutils.is_gcs_path(self.model_id):
@@ -101,8 +125,7 @@ class PeftHandler(BaseHandler):
logging.info(
f"Using task:{self.task}, base model:{self.model_id}, lora model:"
f" {self.finetuned_lora_model_path}, and precision"
f" {self.precision_mode}."
f" {self.finetuned_lora_model_path}, precision {self.precision_mode}."
)
self.pipeline = None
@@ -145,11 +168,18 @@ class PeftHandler(BaseHandler):
elif (
self.task == CAUSAL_LANGUAGE_MODELING_LORA or self.task == INSTRUCT_LORA
):
tokenizer = AutoTokenizer.from_pretrained(self.model_id)
tokenizer = AutoTokenizer.from_pretrained(
self.model_id,
trust_remote_code=self.trust_remote_code,
)
logging.debug("Initialized the tokenizer.")
if self.task == CAUSAL_LANGUAGE_MODELING_LORA:
if self.quantization == constants.AWQ:
model = AutoAWQForCausalLM.from_quantized(self.model_id)
model = AutoAWQForCausalLM.from_quantized(
self.model_id,
trust_remote_code=self.trust_remote_code,
)
elif self.quantization == constants.GPTQ or not self.quantization:
if self.precision_mode == constants.PRECISION_MODE_32:
model = AutoModelForCausalLM.from_pretrained(
@@ -157,6 +187,7 @@ class PeftHandler(BaseHandler):
return_dict=True,
torch_dtype=torch.float32,
device_map="auto",
trust_remote_code=self.trust_remote_code,
)
elif self.precision_mode == constants.PRECISION_MODE_16B:
model = AutoModelForCausalLM.from_pretrained(
@@ -164,6 +195,7 @@ class PeftHandler(BaseHandler):
return_dict=True,
torch_dtype=torch.bfloat16,
device_map="auto",
trust_remote_code=self.trust_remote_code,
)
elif self.precision_mode == constants.PRECISION_MODE_16:
model = AutoModelForCausalLM.from_pretrained(
@@ -171,6 +203,7 @@ class PeftHandler(BaseHandler):
return_dict=True,
torch_dtype=torch.float16,
device_map="auto",
trust_remote_code=self.trust_remote_code,
)
elif self.precision_mode == constants.PRECISION_MODE_8:
quantization_config = BitsAndBytesConfig(
@@ -182,6 +215,7 @@ class PeftHandler(BaseHandler):
torch_dtype=torch.float16,
device_map="auto",
quantization_config=quantization_config,
trust_remote_code=self.trust_remote_code,
)
else:
quantization_config = BitsAndBytesConfig(
@@ -195,6 +229,7 @@ class PeftHandler(BaseHandler):
device_map="auto",
torch_dtype=torch.bfloat16,
quantization_config=quantization_config,
trust_remote_code=self.trust_remote_code,
)
else:
raise ValueError(f"Invalid QUANTIZATION value: {self.quantization}")
@@ -203,14 +238,14 @@ class PeftHandler(BaseHandler):
model = AutoModelForCausalLM.from_pretrained(
self.model_id,
torch_dtype=torch.bfloat16,
trust_remote_code=True,
trust_remote_code=self.trust_remote_code,
device_map="auto",
)
except: # pylint: disable=bare-except
model = AutoModelForCausalLM.from_pretrained(
self.model_id,
torch_dtype=torch.bfloat16,
trust_remote_code=True,
trust_remote_code=self.trust_remote_code,
device_map="auto",
)
logging.debug("Initialized the base model.")
@@ -329,4 +364,4 @@ class PeftHandler(BaseHandler):
return f"Prompt:\n{prompt.strip()}\nOutput:\n{output}"
# pylint: enable=logging-fstring-interpolation
# pylint: enable=logging-fstring-interpolation
@@ -1,131 +0,0 @@
diff --git a/src/transformers/modeling_utils.py b/src/transformers/modeling_utils.py
index 45459ed..32527f4 100644
--- a/src/transformers/modeling_utils.py
+++ b/src/transformers/modeling_utils.py
@@ -32,6 +32,8 @@ import torch
from packaging import version
from torch import Tensor, nn
from torch.nn import CrossEntropyLoss
+from huggingface_hub import hf_hub_download
+from google.cloud import storage
from .activations import get_activation
from .configuration_utils import PretrainedConfig
@@ -442,6 +444,29 @@ def load_state_dict(checkpoint_file: Union[str, os.PathLike]):
"""
Reads a PyTorch checkpoint file, returning properly formatted errors if they arise.
"""
+ delete_download = False
+ tmp_dir = "/tmp/model"
+ os.makedirs(tmp_dir, exist_ok=True)
+ if isinstance(checkpoint_file, dict):
+ # Download model file from huggingface
+ print(f"==> Download model from HF: {checkpoint_file}")
+ checkpoint_file = hf_hub_download(
+ local_dir=tmp_dir, local_dir_use_symlinks=False, force_download=True, resume_download=True, **checkpoint_file)
+ delete_download = True
+ else:
+ with open(checkpoint_file, "rb") as f:
+ is_gcs_file = (f.read(2) == b"gs")
+ if is_gcs_file:
+ # Download model file from GCS
+ with open(checkpoint_file, "r") as f:
+ gcs_file = f.read()
+ checkpoint_file = os.path.join(tmp_dir, gcs_file.split("/")[-1])
+ print(f"==> Download model from GCS: {gcs_file} to: {checkpoint_file}")
+ client = storage.Client()
+ with open(checkpoint_file, 'wb') as f:
+ client.download_blob_to_file(gcs_file, f)
+ delete_download = True
+
if checkpoint_file.endswith(".safetensors") and is_safetensors_available():
# Check format of the archive
with safe_open(checkpoint_file, framework="pt") as f:
@@ -455,9 +480,9 @@ def load_state_dict(checkpoint_file: Union[str, os.PathLike]):
raise NotImplementedError(
f"Conversion from a {metadata['format']} safetensors archive to PyTorch is not implemented yet."
)
- return safe_load_file(checkpoint_file)
+ state_dict = safe_load_file(checkpoint_file)
try:
- return torch.load(checkpoint_file, map_location="cpu")
+ state_dict = torch.load(checkpoint_file, map_location="cpu")
except Exception as e:
try:
with open(checkpoint_file) as f:
@@ -478,6 +503,10 @@ def load_state_dict(checkpoint_file: Union[str, os.PathLike]):
f"at '{checkpoint_file}'. "
"If you tried to load a PyTorch model from a TF 2.0 checkpoint, please set from_tf=True."
)
+ if delete_download:
+ print(f"==> Delete downloaded model: {checkpoint_file}")
+ os.remove(checkpoint_file)
+ return state_dict
def set_initialized_submodules(model, state_dict_keys):
@@ -3179,7 +3208,10 @@ class PreTrainedModel(nn.Module, ModuleUtilsMixin, GenerationMixin, PushToHubMix
return mismatched_keys
if resolved_archive_file is not None:
- folder = os.path.sep.join(resolved_archive_file[0].split(os.path.sep)[:-1])
+ if isinstance(resolved_archive_file, str):
+ folder = os.path.sep.join(resolved_archive_file[0].split(os.path.sep)[:-1])
+ else:
+ folder = None
else:
folder = None
if device_map is not None and is_safetensors:
diff --git a/src/transformers/utils/hub.py b/src/transformers/utils/hub.py
index ffed743..4b15770 100644
--- a/src/transformers/utils/hub.py
+++ b/src/transformers/utils/hub.py
@@ -414,20 +414,34 @@ def cached_file(
user_agent = http_user_agent(user_agent)
try:
# Load from URL or cache if already cached
- resolved_file = hf_hub_download(
- path_or_repo_id,
- filename,
- subfolder=None if len(subfolder) == 0 else subfolder,
- repo_type=repo_type,
- revision=revision,
- cache_dir=cache_dir,
- user_agent=user_agent,
- force_download=force_download,
- proxies=proxies,
- resume_download=resume_download,
- use_auth_token=use_auth_token,
- local_files_only=local_files_only,
- )
+ if filename.endswith(".bin"):
+ # NOTE: To save disk we do not download bin file eagerly. Do not support safetensors.
+ resolved_file = dict(
+ repo_id=path_or_repo_id,
+ filename=filename,
+ subfolder=None if len(subfolder) == 0 else subfolder,
+ repo_type=repo_type,
+ revision=revision,
+ user_agent=user_agent,
+ proxies=proxies,
+ use_auth_token=use_auth_token,
+ )
+ print(f"--> Apply lazy download to bin file: {resolved_file}")
+ else:
+ resolved_file = hf_hub_download(
+ path_or_repo_id,
+ filename,
+ subfolder=None if len(subfolder) == 0 else subfolder,
+ repo_type=repo_type,
+ revision=revision,
+ cache_dir=cache_dir,
+ user_agent=user_agent,
+ force_download=force_download,
+ proxies=proxies,
+ resume_download=resume_download,
+ use_auth_token=use_auth_token,
+ local_files_only=local_files_only,
+ )
except RepositoryNotFoundError:
raise EnvironmentError(
@@ -1,97 +0,0 @@
"""Instruct/Chat with LoRA models."""
# pylint: disable=g-importing-member
from datasets import load_dataset
from peft import LoraConfig
import torch
from transformers import AutoModelForCausalLM
from transformers import AutoTokenizer
from transformers import BitsAndBytesConfig
from transformers import TrainingArguments
from trl import SFTTrainer
def finetune_instruct(
pretrained_model_id: str,
dataset_name: str,
output_dir: str,
lora_rank: int = 64,
lora_alpha: int = 16,
lora_dropout: float = 0.1,
warmup_ratio: int = 0.03,
max_steps: int = 10,
max_seq_length: int = 512,
learning_rate: float = 2e-4,
) -> None:
"""Finetunes instruct."""
dataset = load_dataset(dataset_name, split="train")
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.float16,
)
model = AutoModelForCausalLM.from_pretrained(
pretrained_model_id,
quantization_config=bnb_config,
trust_remote_code=True,
)
model.config.use_cache = False
tokenizer = AutoTokenizer.from_pretrained(
pretrained_model_id, trust_remote_code=True
)
tokenizer.pad_token = tokenizer.eos_token
peft_config = LoraConfig(
lora_alpha=lora_alpha,
lora_dropout=lora_dropout,
r=lora_rank,
bias="none",
task_type="CAUSAL_LM",
target_modules=[
"query_key_value",
"dense",
"dense_h_to_4h",
"dense_4h_to_h",
],
)
per_device_train_batch_size = 4
gradient_accumulation_steps = 4
optim = "paged_adamw_32bit"
save_steps = 10
logging_steps = 10
max_grad_norm = 0.3
lr_scheduler_type = "constant"
training_arguments = TrainingArguments(
output_dir=output_dir,
per_device_train_batch_size=per_device_train_batch_size,
gradient_accumulation_steps=gradient_accumulation_steps,
optim=optim,
save_steps=save_steps,
logging_steps=logging_steps,
learning_rate=learning_rate,
fp16=True,
max_grad_norm=max_grad_norm,
max_steps=max_steps,
warmup_ratio=warmup_ratio,
group_by_length=True,
lr_scheduler_type=lr_scheduler_type,
)
trainer = SFTTrainer(
model=model,
train_dataset=dataset,
peft_config=peft_config,
dataset_text_field="text",
max_seq_length=max_seq_length,
tokenizer=tokenizer,
args=training_arguments,
)
for name, module in trainer.model.named_modules():
if "norm" in name:
module = module.to(torch.float32)
trainer.train()
@@ -1,177 +0,0 @@
"""Main function to start PEFT finetuning."""
import subprocess
from absl import app
from absl import flags
from absl import logging
from peft import causal_language_modeling_lora
from peft import instruct_lora
from peft import sequence_classification_lora
from util import constants
from util import fileutils
_TASK = flags.DEFINE_string(
'task',
constants.CAUSAL_LANGUAGE_MODELING_LORA,
'The supported PEFT tasks.',
)
_PRETRAINED_MODEL_ID = flags.DEFINE_string(
'pretrained_model_id',
None,
'The pretrained model id. Supported models can be causal language modeling'
' models from https://github.com/huggingface/peft/tree/main.',
required=True,
)
_DATASET_NAME = flags.DEFINE_string(
'dataset_name',
None,
'The dataset name in huggingface.',
required=True,
)
_OUTPUT_DIR = flags.DEFINE_string(
'output_dir',
None,
'The output directory.',
required=True,
)
_PRECISION_MODE = flags.DEFINE_string(
'precision_mode',
constants.PRECISION_MODE_16,
'Supported finetuning precision_modes are `{}` and `{}`.'.format(
constants.PRECISION_MODE_8, constants.PRECISION_MODE_16
),
)
_LORA_RANK = flags.DEFINE_integer(
'lora_rank',
16,
'The rank of the update matrices, expressed in int. Lower rank results in'
' smaller update matrices with fewer trainable parameters, referring to'
' https://huggingface.co/docs/peft/conceptual_guides/lora.',
)
_LORA_ALPHA = flags.DEFINE_integer(
'lora_alpha',
32,
'LoRA scaling factor, referring to'
' https://huggingface.co/docs/peft/conceptual_guides/lora.',
)
_LORA_DROPOUT = flags.DEFINE_float(
'lora_dropout',
0.05,
'dropout probability of the LoRA layers, referring to'
' https://huggingface.co/docs/peft/task_guides/token-classification-lora.',
)
_WARMUP_STEPS = flags.DEFINE_integer(
'warmup_steps',
10,
'Number of steps for the warmup in the learning rate scheduler.',
)
_WARMUP_RATIO = flags.DEFINE_float(
'warmup_ratio',
0.03,
'The warmup ratio in the learning rate scheduler.',
)
_MAX_STEPS = flags.DEFINE_integer(
'max_steps',
10,
'Total number of training steps.',
)
_MAX_SEQ_LENGTH = flags.DEFINE_integer(
'max_seq_length',
512,
'The maximum sequence length.',
)
_NUM_EPOCHS = flags.DEFINE_integer(
'num_epochs',
20,
'The number of training epochs.',
)
_BATCH_SIZE = flags.DEFINE_integer(
'batch_size',
32,
'The batch size.',
)
_LEARNING_RATE = flags.DEFINE_float(
'learning_rate',
2e-4,
'The learning rate after the potential warmup period.',
)
def main(_) -> None:
task = _TASK.value
pretrained_model_id = _PRETRAINED_MODEL_ID.value
local_pretrained_model_id = None
if pretrained_model_id.startswith(constants.GCS_URI_PREFIX):
logging.info(
'Start to copy pretrained models locally: %s.', pretrained_model_id
)
fileutils.download_gcs_dir_to_local(
pretrained_model_id, constants.LOCAL_BASE_MODEL_DIR
)
local_pretrained_model_id = constants.LOCAL_BASE_MODEL_DIR
logging.info(
'Finished copying pretrained models locally to: %s.',
local_pretrained_model_id,
)
if task == constants.TEXT_TO_IMAGE_LORA:
subprocess.run(['/bin/bash', 'train.sh'], check=True)
elif task == constants.SEQUENCE_CLASSIFICATION_LORA:
sequence_classification_lora.finetune_sequence_classification(
pretrained_model_id=pretrained_model_id,
dataset_name=_DATASET_NAME.value,
output_dir=_OUTPUT_DIR.value,
lora_rank=_LORA_RANK.value,
lora_alpha=_LORA_ALPHA.value,
lora_dropout=_LORA_DROPOUT.value,
num_epochs=_NUM_EPOCHS.value,
batch_size=_BATCH_SIZE.value,
learning_rate=_LEARNING_RATE.value,
)
elif task == constants.CAUSAL_LANGUAGE_MODELING_LORA:
causal_language_modeling_lora.finetune_causal_language_modeling(
pretrained_model_id=pretrained_model_id,
dataset_name=_DATASET_NAME.value,
output_dir=_OUTPUT_DIR.value,
precision_mode=_PRECISION_MODE.value,
lora_rank=_LORA_RANK.value,
lora_alpha=_LORA_ALPHA.value,
lora_dropout=_LORA_DROPOUT.value,
warmup_steps=_WARMUP_STEPS.value,
max_steps=_MAX_STEPS.value,
learning_rate=_LEARNING_RATE.value,
local_pretrained_model_id=local_pretrained_model_id,
)
elif task == constants.INSTRUCT_LORA:
instruct_lora.finetune_instruct(
pretrained_model_id=pretrained_model_id,
dataset_name=_DATASET_NAME.value,
output_dir=_OUTPUT_DIR.value,
lora_rank=_LORA_RANK.value,
lora_alpha=_LORA_ALPHA.value,
lora_dropout=_LORA_DROPOUT.value,
warmup_ratio=_WARMUP_RATIO.value,
max_steps=_MAX_STEPS.value,
max_seq_length=_MAX_SEQ_LENGTH.value,
learning_rate=_LEARNING_RATE.value,
)
else:
raise ValueError('The task {} is not supported.'.format(task))
if __name__ == '__main__':
app.run(main)
@@ -1,6 +0,0 @@
#!/bin/bash
# Setup accelerate config before running trainer.
python -c "from accelerate.utils import write_basic_config; write_basic_config(mixed_precision='fp16')"
accelerate launch "$@"
@@ -0,0 +1,16 @@
# Dockerfile for axolotl training.
#
# To build:
# docker build -f model_oss/peft/train/axolotol/dockerfile/train.Dockerfile . -t ${YOUR_IMAGE_TAG}
#
# To push to gcr:
# docker tag ${YOUR_IMAGE_TAG} gcr.io/${YOUR_PROJECT}/${YOUR_IMAGE_TAG}
# docker push gcr.io/${YOUR_PROJECT}/${YOUR_IMAGE_TAG}
FROM winglian/axolotl:main-latest
RUN mkdir -p ./vertex_vision_model_garden/
COPY model_oss/peft/train/axolotl/*.py ./vertex_vision_model_garden/
ENTRYPOINT ["python3", "./vertex_vision_model_garden/train_entrypoint.py"]
@@ -0,0 +1,20 @@
#!/bin/bash
# Run copybara first:
# cloud/ml/applications/vision/model_garden/copybara/run_copybara_local.sh
# Run docker build:
# cloud/ml/applications/vision/model_garden/model_oss/peft/train/axolotl/scripts/build_train_docker.sh
set -x
COPYBARA_DIR="/tmp/train_docker/"
pushd "${COPYBARA_DIR}"
PROJECT="cloud-nas-260507"
IMAGE_TAG="gcr.io/${PROJECT}/axolotl-train:${USER}-test"
docker build -f model_oss/peft/train/axolotl/dockerfile/train.Dockerfile . -t "${IMAGE_TAG}"
docker push "${IMAGE_TAG}"
popd
@@ -0,0 +1,88 @@
"""Entrypoint for axolotl train docker."""
import argparse
import json
import os
import subprocess
def _get_multi_node_flags(cluster_spec: str) -> list[str]:
"""Returns the multi-node flags."""
print(f'CLUSTER_SPEC: {cluster_spec}')
cluster_data = json.loads(cluster_spec)
# Get primary node info
primary_node = cluster_data['cluster']['workerpool0'][0]
print(f'primary node: {primary_node}')
primary_node_addr, primary_node_port = primary_node.split(':')
print(f'primary node address: {primary_node_addr}')
print(f'primary node port: {primary_node_port}')
# Determine node rank of this machine
workerpool = cluster_data['task']['type']
if workerpool == 'workerpool0':
node_rank = 0
else:
node_rank = cluster_data['task']['index'] + 1
print(f'node rank: {node_rank}')
# Calculate total nodes
num_worker_nodes = len(cluster_data['cluster']['workerpool1'])
num_nodes = num_worker_nodes + 1 # Add 1 for the primary node
print(f'num nodes: {num_nodes}')
return [
f'--machine_rank={node_rank}',
f'--num_machines={num_nodes}',
f'--main_process_ip={primary_node_addr}',
f'--main_process_port={primary_node_port}',
'--max_restarts=0',
'--monitor_interval=120',
'--dynamo_backend=no',
]
def main() -> None:
parser = argparse.ArgumentParser()
parser.add_argument('--config_file')
parser.add_argument('--huggingface_access_token')
args, unknown = parser.parse_known_args()
accelerate_flags = []
if args.config_file:
accelerate_flags.append(f'--config_file={args.config_file}')
if cluster_spec := os.getenv('CLUSTER_SPEC', default=None):
print('========== Launch on cloud multi nodes ==========')
accelerate_flags.extend(_get_multi_node_flags(cluster_spec))
cmd = (
[
'accelerate',
'launch',
]
+ accelerate_flags
+ [
'-m',
'axolotl.cli.train',
]
+ unknown
)
print(f'{cmd=}', flush=True)
env = os.environ.copy()
if args.huggingface_access_token:
env['HF_TOKEN'] = args.huggingface_access_token
subprocess.run(
cmd,
check=True,
env=env,
)
if __name__ == '__main__':
main()
@@ -0,0 +1,45 @@
"""Class that bundles docker related flags."""
import getpass
import os
import pwd
class DockerCommandBuilder:
"""Bundle docker related flags."""
def __init__(self, docker_uri, shm_size='128gb'):
self._docker_uri = [docker_uri]
self._defaults = [
'docker',
'run',
'--gpus=all',
'--net=host',
'--rm',
f'--shm-size={shm_size}',
]
user = getpass.getuser()
# username ends with `_google_com` is managed by ldap and does not have a
# corresponding entry in /etc/passwd or /etc/group file. We cannot enable
# non-root docker user with below method.
if not user.endswith('_google_com'):
uid = os.getuid()
gid = pwd.getpwuid(uid).pw_gid
self._defaults += [
f'--user={uid}:{gid}',
'--volume=/etc/group:/etc/group:ro',
'--volume=/etc/passwd:/etc/passwd:ro',
]
self._env_vars = []
self._mount_maps = []
def add_env_var(self, var, val):
self._env_vars.append(f'--env={var}={val}')
def add_mount_map(self, host_path, docker_path):
self._mount_maps.append(f'--volume={host_path}:{docker_path}')
def build_cmd(self) -> str:
return self._defaults + self._env_vars + self._mount_maps + self._docker_uri
@@ -0,0 +1,429 @@
# pylint: disable=W,C,R
# DO NOT MODIFY: this file is auto-generated
# See go/vmg-oss-peft-tests#command-builder-genpy
class InstructLoraCommandBuilder:
def __init__(self):
self._config_file = None
self._task = None
self._pretrained_model_id = None
self._dataset_name = None
self._train_split_name = None
self._template = None
self._instruct_column_in_dataset = None
self._output_dir = None
self._merge_base_and_lora_output_dir = None
self._logging_output_dir = None
self._per_device_train_batch_size = None
self._gradient_accumulation_steps = None
self._lora_rank = None
self._lora_alpha = None
self._lora_dropout = None
self._max_steps = None
self._num_epochs = None
self._max_seq_length = None
self._learning_rate = None
self._lr_scheduler_type = None
self._precision_mode = None
self._train_precision = None
self._enable_gradient_checkpointing = None
self._use_example_packing = None
self._attn_implementation = None
self._optimizer = None
self._warmup_ratio = None
self._report_to = None
self._save_steps = None
self._logging_steps = None
self._huggingface_access_token = None
self._eval_dataset_path = None
self._eval_column = None
self._eval_template = None
self._eval_split = None
self._eval_steps = None
self._eval_tasks = None
self._eval_metric_name = None
self._completion_only = None
self._max_grad_norm = None
self._logger_level = None
self._benchmark_out_file = None
self._tuning_data_stats_file = None
self._enable_peft = None
self._merge_model_precision_mode = None
self._target_modules = None
@property
def config_file(self):
return self._config_file
@config_file.setter
def config_file(self, val: str):
self._config_file = val
@property
def task(self):
return self._task
@task.setter
def task(self, val: str):
self._task = val
@property
def pretrained_model_id(self):
return self._pretrained_model_id
@pretrained_model_id.setter
def pretrained_model_id(self, val: str):
self._pretrained_model_id = val
@property
def train_dataset(self):
return self._dataset_name
@train_dataset.setter
def train_dataset(self, val: str):
self._dataset_name = val
@property
def train_split_name(self):
return self._train_split_name
@train_split_name.setter
def train_split_name(self, val: str):
self._train_split_name = val
@property
def template(self):
return self._template
@template.setter
def template(self, val: str):
self._template = val
@property
def instruct_column(self):
return self._instruct_column_in_dataset
@instruct_column.setter
def instruct_column(self, val: str):
self._instruct_column_in_dataset = val
@property
def ckpt_dir(self):
return self._output_dir
@ckpt_dir.setter
def ckpt_dir(self, val: str):
self._output_dir = val
@property
def merged_model_dir(self):
return self._merge_base_and_lora_output_dir
@merged_model_dir.setter
def merged_model_dir(self, val: str):
self._merge_base_and_lora_output_dir = val
@property
def logging_dir(self):
return self._logging_output_dir
@logging_dir.setter
def logging_dir(self, val: str):
self._logging_output_dir = val
@property
def per_device_batch_size(self):
return self._per_device_train_batch_size
@per_device_batch_size.setter
def per_device_batch_size(self, val: int):
self._per_device_train_batch_size = val
@property
def gradient_accumulation_steps(self):
return self._gradient_accumulation_steps
@gradient_accumulation_steps.setter
def gradient_accumulation_steps(self, val: int):
self._gradient_accumulation_steps = val
@property
def lora_rank(self):
return self._lora_rank
@lora_rank.setter
def lora_rank(self, val: int):
self._lora_rank = val
@property
def lora_alpha(self):
return self._lora_alpha
@lora_alpha.setter
def lora_alpha(self, val: int):
self._lora_alpha = val
@property
def lora_dropout(self):
return self._lora_dropout
@lora_dropout.setter
def lora_dropout(self, val: float):
self._lora_dropout = val
@property
def max_steps(self):
return self._max_steps
@max_steps.setter
def max_steps(self, val: int):
self._max_steps = val
@property
def num_epochs(self):
return self._num_epochs
@num_epochs.setter
def num_epochs(self, val: float):
self._num_epochs = val
@property
def max_seq_length(self):
return self._max_seq_length
@max_seq_length.setter
def max_seq_length(self, val: int):
self._max_seq_length = val
@property
def learning_rate(self):
return self._learning_rate
@learning_rate.setter
def learning_rate(self, val: float):
self._learning_rate = val
@property
def lr_scheduler_type(self):
return self._lr_scheduler_type
@lr_scheduler_type.setter
def lr_scheduler_type(self, val: str):
self._lr_scheduler_type = val
@property
def load_precision(self):
return self._precision_mode
@load_precision.setter
def load_precision(self, val: str):
self._precision_mode = val
@property
def train_precision(self):
return self._train_precision
@train_precision.setter
def train_precision(self, val: str):
self._train_precision = val
@property
def gradient_checkpointing(self):
return self._enable_gradient_checkpointing
@gradient_checkpointing.setter
def gradient_checkpointing(self, val: bool):
self._enable_gradient_checkpointing = val
@property
def example_packing(self):
return self._use_example_packing
@example_packing.setter
def example_packing(self, val: bool):
self._use_example_packing = val
@property
def attn_implementation(self):
return self._attn_implementation
@attn_implementation.setter
def attn_implementation(self, val: str):
self._attn_implementation = val
@property
def optimizer(self):
return self._optimizer
@optimizer.setter
def optimizer(self, val: str):
self._optimizer = val
@property
def warmup_ratio(self):
return self._warmup_ratio
@warmup_ratio.setter
def warmup_ratio(self, val: float):
self._warmup_ratio = val
@property
def report_to(self):
return self._report_to
@report_to.setter
def report_to(self, val: str):
self._report_to = val
@property
def save_steps(self):
return self._save_steps
@save_steps.setter
def save_steps(self, val: int):
self._save_steps = val
@property
def logging_steps(self):
return self._logging_steps
@logging_steps.setter
def logging_steps(self, val: int):
self._logging_steps = val
@property
def huggingface_access_token(self):
return self._huggingface_access_token
@huggingface_access_token.setter
def huggingface_access_token(self, val: str):
self._huggingface_access_token = val
@property
def eval_dataset(self):
return self._eval_dataset_path
@eval_dataset.setter
def eval_dataset(self, val: str):
self._eval_dataset_path = val
@property
def eval_instruct_column(self):
return self._eval_column
@eval_instruct_column.setter
def eval_instruct_column(self, val: str):
self._eval_column = val
@property
def eval_template(self):
return self._eval_template
@eval_template.setter
def eval_template(self, val: str):
self._eval_template = val
@property
def eval_split_name(self):
return self._eval_split
@eval_split_name.setter
def eval_split_name(self, val: str):
self._eval_split = val
@property
def eval_steps(self):
return self._eval_steps
@eval_steps.setter
def eval_steps(self, val: int):
self._eval_steps = val
@property
def eval_tasks(self):
return self._eval_tasks
@eval_tasks.setter
def eval_tasks(self, val: str):
self._eval_tasks = val
@property
def eval_metric_name(self):
return self._eval_metric_name
@eval_metric_name.setter
def eval_metric_name(self, val: str):
self._eval_metric_name = val
@property
def completion_only(self):
return self._completion_only
@completion_only.setter
def completion_only(self, val: bool):
self._completion_only = val
@property
def max_grad_norm(self):
return self._max_grad_norm
@max_grad_norm.setter
def max_grad_norm(self, val: float):
self._max_grad_norm = val
@property
def logger_level(self):
return self._logger_level
@logger_level.setter
def logger_level(self, val: str):
self._logger_level = val
@property
def benchmark_out_file(self):
return self._benchmark_out_file
@benchmark_out_file.setter
def benchmark_out_file(self, val: str):
self._benchmark_out_file = val
@property
def tuning_data_stats_file(self):
return self._tuning_data_stats_file
@tuning_data_stats_file.setter
def tuning_data_stats_file(self, val: str):
self._tuning_data_stats_file = val
@property
def enable_peft(self):
return self._enable_peft
@enable_peft.setter
def enable_peft(self, val: bool):
self._enable_peft = val
@property
def merge_model_precision_mode(self):
return self._merge_model_precision_mode
@merge_model_precision_mode.setter
def merge_model_precision_mode(self, val: str):
self._merge_model_precision_mode = val
@property
def target_modules(self):
return self._target_modules
@target_modules.setter
def target_modules(self, val: str):
self._target_modules = val
def build_cmd(self) -> str:
cmd = []
for k, v in self.__dict__.items():
if v is not None:
cmd.append(f'--{k[1:]}={v}')
return cmd
@@ -0,0 +1,142 @@
# pylint: disable=W,C,R
# DO NOT MODIFY: this file is auto-generated
# See go/vmg-oss-peft-tests#command-builder-genpy
class QuantizeModelCommandBuilder:
def __init__(self):
self._task = None
self._pretrained_model_id = None
self._quantization_method = None
self._quantization_precision_mode = None
self._quantization_dataset_name = None
self._text_column_in_quantization_dataset = None
self._quantization_output_dir = None
self._device_map = None
self._max_memory = None
self._group_size = None
self._desc_act = None
self._damp_percent = None
self._cache_examples_on_gpu = None
self._awq_version = None
@property
def task(self):
return self._task
@task.setter
def task(self, val: str):
self._task = val
@property
def pretrained_model_id(self):
return self._pretrained_model_id
@pretrained_model_id.setter
def pretrained_model_id(self, val: str):
self._pretrained_model_id = val
@property
def quantization_method(self):
return self._quantization_method
@quantization_method.setter
def quantization_method(self, val: str):
self._quantization_method = val
@property
def quantization_precision_mode(self):
return self._quantization_precision_mode
@quantization_precision_mode.setter
def quantization_precision_mode(self, val: str):
self._quantization_precision_mode = val
@property
def quantization_dataset_name(self):
return self._quantization_dataset_name
@quantization_dataset_name.setter
def quantization_dataset_name(self, val: str):
self._quantization_dataset_name = val
@property
def text_column_in_quantization_dataset(self):
return self._text_column_in_quantization_dataset
@text_column_in_quantization_dataset.setter
def text_column_in_quantization_dataset(self, val: str):
self._text_column_in_quantization_dataset = val
@property
def quantization_output_dir(self):
return self._quantization_output_dir
@quantization_output_dir.setter
def quantization_output_dir(self, val: str):
self._quantization_output_dir = val
@property
def device_map(self):
return self._device_map
@device_map.setter
def device_map(self, val: str):
self._device_map = val
@property
def max_memory(self):
return self._max_memory
@max_memory.setter
def max_memory(self, val: str):
self._max_memory = val
@property
def group_size(self):
return self._group_size
@group_size.setter
def group_size(self, val: int):
self._group_size = val
@property
def desc_act(self):
return self._desc_act
@desc_act.setter
def desc_act(self, val: bool):
self._desc_act = val
@property
def damp_percent(self):
return self._damp_percent
@damp_percent.setter
def damp_percent(self, val: float):
self._damp_percent = val
@property
def cache_examples_on_gpu(self):
return self._cache_examples_on_gpu
@cache_examples_on_gpu.setter
def cache_examples_on_gpu(self, val: bool):
self._cache_examples_on_gpu = val
@property
def awq_version(self):
return self._awq_version
@awq_version.setter
def awq_version(self, val: str):
self._awq_version = val
def build_cmd(self) -> str:
cmd = []
for k, v in self.__dict__.items():
if v is not None:
cmd.append(f'--{k[1:]}={v}')
return cmd
@@ -0,0 +1,106 @@
# pylint: disable=missing-function-docstring
# pylint: disable=missing-class-docstring
"""Tests adapters of PEFT train docker."""
import inspect
import os
import time
from absl.testing import absltest
from absl.testing import parameterized
import instruct_lora_command_builder as task_cmd_builder
from safetensors import safe_open
import test_util
class AdapterTest(test_util.TestBase):
# Needs to be accessible outside docker to check artifacts.
_TEST_OUTPUT_DIR = os.path.expanduser('~/output')
_MODULES_NEED_TO_BE_EXCLUDED_IN_ADAPTER = ['lm_head', 'embed_tokens']
@classmethod
def setUpClass(cls):
super().setUpClass()
cls.test_suite_output_dir = os.path.join(
cls._TEST_OUTPUT_DIR,
os.path.splitext(os.path.basename(__file__))[0],
cls.__class__.__name__,
)
def setUp(self):
super().setUp()
self.docker_builder.add_env_var('CUDA_VISIBLE_DEVICES', '0,1,2,3,4,5,6,7')
self.task_cmd_builder = task_cmd_builder.InstructLoraCommandBuilder()
self.task_cmd_builder.task = 'instruct-lora'
self.task_cmd_builder.per_device_batch_size = 1
self.task_cmd_builder.train_dataset = test_util.get_test_data_path(
'peft_train_sample.jsonl'
)
self.task_cmd_builder.train_split_name = 'train'
self.task_cmd_builder.instruct_column = 'input_text'
self.task_cmd_builder.template = 'llama3-text-bison'
self.task_cmd_builder.gradient_accumulation_steps = 1
self.task_cmd_builder.lora_rank = 16
self.task_cmd_builder.lora_alpha = 32
self.task_cmd_builder.lora_dropout = 0.05
self.task_cmd_builder.max_steps = 1
self.task_cmd_builder.max_seq_length = 256
self.task_cmd_builder.load_precision = '4bit'
self.task_cmd_builder.gradient_checkpointing = True
self.task_cmd_builder.attn_implementation = 'flash_attention_2'
self.task_cmd_builder.save_steps = 10
self.task_cmd_builder.max_steps = 3
self.task_cmd_builder.config_file = (
'vertex_vision_model_garden_peft/llama_fsdp_8gpu.yaml'
)
def setup_output_dir(self, testcase_name: str):
testcase_output_dir = os.path.join(
self.test_suite_output_dir, testcase_name, test_util.get_timestamp()
)
self.task_cmd_builder.ckpt_dir = os.path.join(
testcase_output_dir, 'adapter'
)
self.task_cmd_builder.logging_dir = os.path.join(
testcase_output_dir, 'logs'
)
def check_adapter_for_bad_modules(self, adapter_path):
unwanted_modules = set()
with safe_open(adapter_path, framework='pt', device='cpu') as f:
for key in f.keys():
for module in self._MODULES_NEED_TO_BE_EXCLUDED_IN_ADAPTER:
if module in key:
unwanted_modules.add(key)
assert (
not unwanted_modules
), f'Adapter includes unwanted modules: {unwanted_modules}'
@parameterized.named_parameters(
('llama3.1-8b', 'llama3.1-8b-hf'),
('llama3.1-70b', 'llama3.1-70b-hf'),
('llama2-7b', 'llama2-7b-hf'),
)
def test_llama_adapters(self, model_name):
test_function_name = inspect.stack()[0][3]
self.setup_output_dir(f'{test_function_name}-{model_name}')
self.task_cmd_builder.pretrained_model_id = (
test_util.get_pretrained_model_id(model_name)
)
start_time = time.time()
self.assertEqual(self.run_cmd(), 0)
end_time = time.time()
self.assertLess(end_time - start_time, 1000)
adapter = os.path.join(
self.task_cmd_builder.ckpt_dir,
'checkpoint-final/adapter_model.safetensors',
)
self.check_adapter_for_bad_modules(adapter)
if __name__ == '__main__':
absltest.main()
@@ -0,0 +1,290 @@
# pylint: disable=missing-function-docstring
# pylint: disable=missing-class-docstring
"""Tests various features of PEFT train docker."""
import os
import time
from absl.testing import absltest
from absl.testing import parameterized
import instruct_lora_command_builder as task_cmd_builder
import test_util
class GcsUploadDownloadTest(test_util.TestBase):
def setUp(self):
super().setUp()
self.docker_builder.add_env_var('CUDA_VISIBLE_DEVICES', '0')
self.task_cmd_builder = task_cmd_builder.InstructLoraCommandBuilder()
self.task_cmd_builder.task = 'instruct-lora'
self.task_cmd_builder.per_device_batch_size = 1
self.task_cmd_builder.train_dataset = test_util.get_test_data_path(
'peft_train_sample.jsonl'
)
self.task_cmd_builder.train_split_name = 'train'
self.task_cmd_builder.instruct_column = 'input_text'
self.task_cmd_builder.template = 'llama3-text-bison'
self.task_cmd_builder.gradient_accumulation_steps = 1
self.task_cmd_builder.lora_rank = 16
self.task_cmd_builder.lora_alpha = 32
self.task_cmd_builder.lora_dropout = 0.05
self.task_cmd_builder.max_steps = 1
self.task_cmd_builder.max_seq_length = 256
self.task_cmd_builder.load_precision = '4bit'
self.task_cmd_builder.gradient_checkpointing = True
self.task_cmd_builder.attn_implementation = 'flash_attention_2'
self.task_cmd_builder.ckpt_dir = '/tmp'
@parameterized.named_parameters(
(
'llama3_8b_gcs',
'gs://vertex-model-garden-public-us/llama3/llama3-8b-hf',
),
('llama2_7b_hf', 'NousResearch/Llama-2-7b-hf'),
)
def test_model_download_single_process(self, pretrained_model_id):
self.task_cmd_builder.pretrained_model_id = (
test_util.get_pretrained_model_id(pretrained_model_id)
)
start_time = time.time()
self.assertEqual(self.run_cmd(), 0)
end_time = time.time()
self.assertLess(end_time - start_time, 5 * 60.0)
@parameterized.named_parameters(
(
'llama3_8b_gcs',
'gs://vertex-model-garden-public-us/llama3/llama3-8b-hf',
),
('llama2_7b_hf', 'NousResearch/Llama-2-7b-hf'),
)
def test_model_download_multi_process(self, pretrained_model_id):
self.task_cmd_builder.pretrained_model_id = (
test_util.get_pretrained_model_id(pretrained_model_id)
)
self.task_cmd_builder.config_file = (
'vertex_vision_model_garden_peft/llama_fsdp_8gpu.yaml'
)
self.docker_builder.add_env_var('CUDA_VISIBLE_DEVICES', '0,1,2,3,4,5,6,7')
start_time = time.time()
self.assertEqual(self.run_cmd(), 0)
end_time = time.time()
self.assertLess(end_time - start_time, 5 * 60.0)
def test_70b_model_download(self):
self.task_cmd_builder.pretrained_model_id = (
'gs://vertex-model-garden-public-us/llama3/llama3-70b-hf'
)
start_time = time.time()
self.assertEqual(self.run_cmd(), 0)
end_time = time.time()
self.assertLess(end_time - start_time, 10 * 60.0)
@parameterized.named_parameters(
('merged-without-upload', '/tmp/merged'),
('merged-and-upload-to-gcs', 'gs://vmg-test-ttl-1y/tests/merged'),
)
def test_model_merge(self, output_dir):
self.task_cmd_builder.pretrained_model_id = (
test_util.get_pretrained_model_id('llama3.1-8b-hf')
)
ckpt_dir = os.path.join(
output_dir,
f'output-{test_util.get_timestamp()}',
)
self.task_cmd_builder.ckpt_dir = ckpt_dir
self.task_cmd_builder.merged_model_dir = os.path.join(ckpt_dir, 'merged')
self.task_cmd_builder.logging_dir = '/tmp/logging'
start_time = time.time()
self.assertEqual(self.run_cmd(), 0)
end_time = time.time()
self.assertLess(end_time - start_time, 5 * 60.0)
def test_model_fp8_conversion(self):
self.task_cmd_builder.pretrained_model_id = (
test_util.get_pretrained_model_id('llama3.1-8b-hf')
)
ckpt_dir = f'/tmp/output/output-{test_util.get_timestamp()}'
self.task_cmd_builder.ckpt_dir = ckpt_dir
self.task_cmd_builder.merged_model_dir = os.path.join(ckpt_dir, 'merged')
self.task_cmd_builder.logging_dir = os.path.join(ckpt_dir, 'logging')
self.task_cmd_builder.merge_model_precision_mode = 'float8'
start_time = time.time()
self.assertEqual(self.run_cmd(), 0)
end_time = time.time()
self.assertLess(end_time - start_time, 5 * 60.0)
@parameterized.named_parameters(
('merged-without-upload', '/tmp/merged'),
('merged-and-upload-to-gcs', 'gs://vmg-test-ttl-1y/tests/merged'),
)
def test_model_merge_and_upload_deepspeed(self, merged_model_dir):
self.task_cmd_builder.pretrained_model_id = (
test_util.get_pretrained_model_id('llama3.1-8b-hf')
)
self.task_cmd_builder.config_file = (
'vertex_vision_model_garden_peft/deepspeed_zero3_8gpu.yaml'
)
self.task_cmd_builder.merged_model_dir = os.path.join(
merged_model_dir, f'merged-{test_util.get_timestamp()}'
)
self.docker_builder.add_env_var('CUDA_VISIBLE_DEVICES', '0,1,2,3,4,5,6,7')
start_time = time.time()
self.assertEqual(self.run_cmd(), 0)
end_time = time.time()
self.assertLess(end_time - start_time, 5 * 60.0)
@parameterized.named_parameters(
('save-only-last', 10),
('save-multiple-times', 1),
)
def test_llama3_8b_save_and_merge_8_gpus_fsdp(self, save_steps):
self.task_cmd_builder.pretrained_model_id = (
test_util.get_pretrained_model_id('llama3.1-8b-hf')
)
self.task_cmd_builder.save_steps = save_steps
self.task_cmd_builder.max_steps = 3
self.task_cmd_builder.config_file = (
'vertex_vision_model_garden_peft/llama_fsdp_8gpu.yaml'
)
self.task_cmd_builder.merged_model_dir = '/tmp/merged'
self.docker_builder.add_env_var('CUDA_VISIBLE_DEVICES', '0,1,2,3,4,5,6,7')
start_time = time.time()
self.assertEqual(self.run_cmd(), 0)
end_time = time.time()
self.assertLess(end_time - start_time, 9 * 60.0)
class TemplateAndDataStatsTest(test_util.TestBase):
def setUp(self):
super().setUp()
self.docker_builder.add_env_var('CUDA_VISIBLE_DEVICES', '0')
self.task_cmd_builder = task_cmd_builder.InstructLoraCommandBuilder()
self.task_cmd_builder.task = 'instruct-lora'
self.task_cmd_builder.per_device_batch_size = 1
self.task_cmd_builder.gradient_accumulation_steps = 1
self.task_cmd_builder.lora_rank = 16
self.task_cmd_builder.lora_alpha = 32
self.task_cmd_builder.lora_dropout = 0.05
self.task_cmd_builder.max_steps = 1
self.task_cmd_builder.max_seq_length = 256
self.task_cmd_builder.load_precision = '4bit'
self.task_cmd_builder.gradient_checkpointing = True
self.task_cmd_builder.attn_implementation = 'flash_attention_2'
self.task_cmd_builder.ckpt_dir = '/tmp'
def test_openai_chat_template(self):
self.task_cmd_builder.pretrained_model_id = (
test_util.get_pretrained_model_id('llama3.1-8b-hf')
)
self.task_cmd_builder.train_dataset = test_util.get_test_data_path(
'openai-multi-chat-example-data.jsonl'
)
self.task_cmd_builder.train_split_name = 'train'
self.task_cmd_builder.instruct_column = 'messages'
self.task_cmd_builder.template = 'llama3'
self.assertEqual(self.run_cmd(), 0)
def test_openai_completion_template(self):
self.task_cmd_builder.pretrained_model_id = (
test_util.get_pretrained_model_id('llama3.1-8b-hf')
)
self.task_cmd_builder.train_dataset = test_util.get_test_data_path(
'openai-completion-example-data.jsonl'
)
self.task_cmd_builder.train_split_name = 'train'
self.task_cmd_builder.instruct_column = 'prompt'
self.task_cmd_builder.template = 'openai-completion'
self.assertEqual(self.run_cmd(), 0)
def test_data_stats_chat_template(self):
self.task_cmd_builder.pretrained_model_id = (
test_util.get_pretrained_model_id('llama3.1-8b-hf')
)
self.task_cmd_builder.config_file = (
'vertex_vision_model_garden_peft/llama_fsdp_8gpu.yaml'
)
self.task_cmd_builder.train_dataset = test_util.get_test_data_path(
'openai-multi-chat-example-data.jsonl'
)
self.task_cmd_builder.train_split_name = 'train'
self.task_cmd_builder.instruct_column = 'messages'
self.task_cmd_builder.template = 'llama3'
self.task_cmd_builder.tuning_data_stats_file = '/tmp/data-stats.json'
self.docker_builder.add_env_var('CUDA_VISIBLE_DEVICES', '0,1,2,3,4,5,6,7')
self.assertEqual(self.run_cmd(), 0)
def test_data_stats_completion_template(self):
self.task_cmd_builder.pretrained_model_id = (
test_util.get_pretrained_model_id('llama3.1-8b-hf')
)
self.task_cmd_builder.config_file = (
'vertex_vision_model_garden_peft/llama_fsdp_8gpu.yaml'
)
self.task_cmd_builder.train_dataset = test_util.get_test_data_path(
'openai-completion-example-data.jsonl'
)
self.task_cmd_builder.train_split_name = 'train'
self.task_cmd_builder.instruct_column = 'prompt'
self.task_cmd_builder.template = 'openai-completion'
self.task_cmd_builder.tuning_data_stats_file = '/tmp/data-stats.json'
self.docker_builder.add_env_var('CUDA_VISIBLE_DEVICES', '0,1,2,3,4,5,6,7')
self.assertEqual(self.run_cmd(), 0)
class TargetModulesTest(test_util.TestBase):
def setUp(self):
super().setUp()
self.docker_builder.add_env_var('CUDA_VISIBLE_DEVICES', '0')
self.task_cmd_builder = task_cmd_builder.InstructLoraCommandBuilder()
self.task_cmd_builder.task = 'instruct-lora'
self.task_cmd_builder.per_device_batch_size = 1
self.task_cmd_builder.train_dataset = test_util.get_test_data_path(
'peft_train_sample.jsonl'
)
self.task_cmd_builder.train_split_name = 'train'
self.task_cmd_builder.instruct_column = 'input_text'
self.task_cmd_builder.template = 'llama3-text-bison'
self.task_cmd_builder.gradient_accumulation_steps = 1
self.task_cmd_builder.lora_rank = 16
self.task_cmd_builder.lora_alpha = 32
self.task_cmd_builder.lora_dropout = 0.05
self.task_cmd_builder.max_steps = 1
self.task_cmd_builder.max_seq_length = 256
self.task_cmd_builder.load_precision = '4bit'
self.task_cmd_builder.gradient_checkpointing = True
self.task_cmd_builder.attn_implementation = 'flash_attention_2'
self.task_cmd_builder.ckpt_dir = '/tmp'
def test_target_modules(self):
self.task_cmd_builder.pretrained_model_id = (
test_util.get_pretrained_model_id('llama3.1-8b-hf')
)
self.task_cmd_builder.target_modules = 'q_proj, v_proj, k_proj'
self.assertEqual(self.run_cmd(), 0)
if __name__ == '__main__':
absltest.main()
@@ -0,0 +1,180 @@
# pylint: disable=missing-function-docstring
# pylint: disable=missing-class-docstring
"""Tests to check training throughput and GPU memory consumption."""
import os
import pathlib
from absl.testing import absltest
from absl.testing import parameterized
import instruct_lora_command_builder as task_cmd_builder
import test_util
class TrainerThroughputTest(test_util.TestBase):
_TEST_OUTPUT_DIR = os.path.expanduser('~/throughput_tests')
@classmethod
def setUpClass(cls):
super().setUpClass()
cls.test_suite_output_dir = os.path.join(
cls._TEST_OUTPUT_DIR, os.path.splitext(os.path.basename(__file__))[0]
)
if not os.path.isdir(cls.test_suite_output_dir):
pathlib.Path(cls.test_suite_output_dir).mkdir(parents=True)
def setUp(self):
super().setUp()
self.task_cmd_builder = task_cmd_builder.InstructLoraCommandBuilder()
self.task_cmd_builder.task = 'instruct-lora'
self.task_cmd_builder.per_device_batch_size = 1
self.task_cmd_builder.gradient_accumulation_steps = 1
self.task_cmd_builder.lora_rank = 16
self.task_cmd_builder.lora_alpha = 32
self.task_cmd_builder.lora_dropout = 0.05
self.task_cmd_builder.learning_rate = 5e-5
self.task_cmd_builder.warmup_ratio = 0.01
self.task_cmd_builder.max_steps = 10
self.task_cmd_builder.save_steps = 1000
self.task_cmd_builder.logging_steps = 1
self.task_cmd_builder.gradient_checkpointing = True
self.task_cmd_builder.attn_implementation = 'flash_attention_2'
self.task_cmd_builder.example_packing = True
self.task_cmd_builder.train_dataset = 'mlabonne/guanaco-llama2-1k'
self.task_cmd_builder.train_split_name = 'train'
self.task_cmd_builder.instruct_column = 'text'
self.task_cmd_builder.template = 'openassistant-guanaco'
self.task_cmd_builder.ckpt_dir = '/tmp/adapter'
self.task_cmd_builder.logging_dir = '/tmp/logs'
def run_cmd_and_handle_failure(self):
ret = self.run_cmd()
if ret != 0:
with open(self.task_cmd_builder.benchmark_out_file, 'a') as f:
max_seq_length = self.task_cmd_builder.max_seq_length
f.write(f'{max_seq_length/1024.0:.1f}k | failed | n/a\n')
return ret
@parameterized.product(
model_name=[
'llama3-70b-hf',
'llama3.1-70b-hf',
'Mistral-7B-v0.1',
'Mixtral-8x7B-v0.1',
'Gemma2-9b-it',
],
precision=['4bit', '8bit', 'bfloat16'],
max_seq_length=list(range(4 * 1024, 24 * 1024 + 1, 4 * 1024)),
)
def test_model_single_gpu(self, model_name, precision, max_seq_length):
self.task_cmd_builder.pretrained_model_id = (
test_util.get_pretrained_model_id(model_name)
)
self.task_cmd_builder.max_seq_length = max_seq_length
self.task_cmd_builder.load_precision = precision
self.task_cmd_builder.benchmark_out_file = os.path.join(
self.test_suite_output_dir, f'bm_{model_name}_{precision}.txt'
)
self.docker_builder.add_env_var('CUDA_VISIBLE_DEVICES', '0')
self.assertEqual(self.run_cmd_and_handle_failure(), 0)
@parameterized.product(
model_name=[
'llama3-70b-hf',
'llama3.1-70b-hf',
'Mistral-7B-v0.1',
'Mixtral-8x7B-v0.1',
'Gemma2-9b-it',
],
precision=['4bit', '8bit', 'bfloat16'],
max_seq_length=list(range(4 * 1024, 24 * 1024 + 1, 4 * 1024)),
num_gpus=[8],
config=['deepspeed_zero2', 'deepspeed_zero3'],
)
def test_model_multi_gpu_deepspeed(
self, model_name, precision, max_seq_length, num_gpus, config
):
self.assertTrue(num_gpus == 4 or num_gpus == 8)
self.task_cmd_builder.pretrained_model_id = (
test_util.get_pretrained_model_id(model_name)
)
self.task_cmd_builder.max_seq_length = max_seq_length
self.task_cmd_builder.load_precision = precision
self.task_cmd_builder.benchmark_out_file = os.path.join(
self.test_suite_output_dir,
f'bm_{config}_{num_gpus}gpu_{model_name}_{precision}.txt',
)
self.task_cmd_builder.config_file = (
f'vertex_vision_model_garden_peft/{config}_{num_gpus}gpu.yaml'
)
self.docker_builder.add_env_var(
'CUDA_VISIBLE_DEVICES', ','.join([str(x) for x in range(0, num_gpus)])
)
self.assertEqual(self.run_cmd_and_handle_failure(), 0)
@parameterized.product(
model_name=['llama3.1-70b-hf'],
precision=['4bit', '8bit', 'bfloat16'],
max_seq_length=list(range(4 * 1024, 24 * 1024 + 1, 4 * 1024)),
num_gpus=[8],
)
def test_model_multi_gpu_fsdp_lora(
self, model_name, precision, max_seq_length, num_gpus
):
self.task_cmd_builder.pretrained_model_id = (
test_util.get_pretrained_model_id(model_name)
)
self.task_cmd_builder.max_seq_length = max_seq_length
self.task_cmd_builder.load_precision = precision
self.task_cmd_builder.benchmark_out_file = os.path.join(
self.test_suite_output_dir,
f'bm_fsdp_{num_gpus}gpu_{model_name}_{precision}.txt',
)
self.task_cmd_builder.config_file = (
'vertex_vision_model_garden_peft/llama_fsdp_8gpu.yaml'
)
self.docker_builder.add_env_var(
'CUDA_VISIBLE_DEVICES', ','.join([str(x) for x in range(0, num_gpus)])
)
self.assertEqual(self.run_cmd_and_handle_failure(), 0)
@parameterized.product(
model_name=['llama3.1-70b-hf'],
precision=['bfloat16'],
max_seq_length=list(range(4 * 1024, 24 * 1024 + 1, 4 * 1024)),
num_gpus=[8],
)
def test_model_multi_gpu_fsdp_full_finetuning(
self, model_name, precision, max_seq_length, num_gpus
):
self.task_cmd_builder.pretrained_model_id = (
test_util.get_pretrained_model_id(model_name)
)
self.task_cmd_builder.max_seq_length = max_seq_length
self.task_cmd_builder.load_precision = precision
self.task_cmd_builder.benchmark_out_file = os.path.join(
self.test_suite_output_dir,
f'bm_fsdp_full_finetuning_{num_gpus}gpu_{model_name}_{precision}.txt',
)
self.task_cmd_builder.config_file = (
'vertex_vision_model_garden_peft/llama_fsdp_8gpu.yaml'
)
self.task_cmd_builder.enable_peft = False
self.docker_builder.add_env_var(
'CUDA_VISIBLE_DEVICES', ','.join([str(x) for x in range(0, num_gpus)])
)
self.assertEqual(self.run_cmd_and_handle_failure(), 0)
if __name__ == '__main__':
absltest.main()
@@ -0,0 +1,152 @@
# pylint: disable=missing-function-docstring
# pylint: disable=missing-class-docstring
"""Tests to make sure trained model achieves decent quality.
Right now, the metric is loss decreasing and we'll eyeball the TB graphs.
"""
import os
from absl.testing import absltest
from absl.testing import parameterized
import instruct_lora_command_builder as task_cmd_builder
import test_util
class TrainedModelQualityTest(test_util.TestBase):
_TEST_OUTPUT_DIR = os.path.expanduser('~/output')
@classmethod
def setUpClass(cls):
super().setUpClass()
cls.test_suite_output_dir = os.path.join(
cls._TEST_OUTPUT_DIR,
os.path.splitext(os.path.basename(__file__))[0],
cls.__class__.__name__,
)
def setUp(self):
super().setUp()
self.task_cmd_builder = task_cmd_builder.InstructLoraCommandBuilder()
self.task_cmd_builder.task = 'instruct-lora'
self.task_cmd_builder.eval_tasks = 'builtin_eval'
self.task_cmd_builder.eval_metric_name = 'loss'
self.task_cmd_builder.per_device_batch_size = 1
self.task_cmd_builder.gradient_accumulation_steps = 8
self.task_cmd_builder.lora_rank = 16
self.task_cmd_builder.lora_alpha = 32
self.task_cmd_builder.lora_dropout = 0.05
self.task_cmd_builder.learning_rate = 5e-5
self.task_cmd_builder.num_epochs = 2.0
self.task_cmd_builder.warmup_ratio = 0.01
self.task_cmd_builder.max_steps = -1
self.task_cmd_builder.save_steps = 10
self.task_cmd_builder.eval_steps = 10
self.task_cmd_builder.max_seq_length = 4096
self.task_cmd_builder.load_precision = '4bit'
self.task_cmd_builder.gradient_checkpointing = True
self.task_cmd_builder.completion_only = True
self.task_cmd_builder.attn_implementation = 'flash_attention_2'
self.task_cmd_builder.report_to = 'tensorboard'
def setup_output_dir(self, testcase_name: str):
testcase_output_dir = os.path.join(
self.test_suite_output_dir, testcase_name
)
self.task_cmd_builder.ckpt_dir = os.path.join(
testcase_output_dir, 'adapter'
)
self.task_cmd_builder.logging_dir = os.path.join(
testcase_output_dir, 'logs'
)
self.task_cmd_builder.merged_model_dir = os.path.join(
testcase_output_dir, 'merged'
)
@parameterized.named_parameters(
('llama3-8b', 'llama3-8b-hf'),
('llama3.1-8b', 'llama3.1-8b-hf'),
)
def test_8b_model_deepspeed(self, model_name):
self.setup_output_dir(f'test_deepspeed_{model_name}')
self.task_cmd_builder.pretrained_model_id = (
test_util.get_pretrained_model_id(model_name)
)
self.task_cmd_builder.train_dataset = test_util.get_test_data_path(
'peft_train_sample.jsonl'
)
self.task_cmd_builder.train_split_name = 'train'
self.task_cmd_builder.instruct_column = 'input_text'
self.task_cmd_builder.template = 'llama3-text-bison'
self.task_cmd_builder.eval_dataset = test_util.get_test_data_path(
'peft_eval_sample.jsonl'
)
self.task_cmd_builder.eval_split_name = 'train'
self.task_cmd_builder.eval_instruct_column = (
self.task_cmd_builder.instruct_column
)
self.task_cmd_builder.eval_template = self.task_cmd_builder.template
self.docker_builder.add_env_var('CUDA_VISIBLE_DEVICES', '0')
self.assertEqual(self.run_cmd(), 0)
@parameterized.named_parameters(
('llama3-70b', 'llama3-70b-hf'),
('llama3.1-70b', 'llama3.1-70b-hf'),
)
def test_70b_model_deepspeed(self, model_name):
self.setup_output_dir(f'test_deepspeed_{model_name}')
self.task_cmd_builder.pretrained_model_id = (
test_util.get_pretrained_model_id(model_name)
)
self.task_cmd_builder.config_file = (
'vertex_vision_model_garden_peft/deepspeed_zero2_8gpu.yaml'
)
self.task_cmd_builder.train_dataset = 'timdettmers/openassistant-guanaco'
self.task_cmd_builder.train_split_name = 'train'
self.task_cmd_builder.instruct_column = 'text'
self.task_cmd_builder.template = 'openassistant-guanaco'
self.task_cmd_builder.eval_dataset = self.task_cmd_builder.train_dataset
self.task_cmd_builder.eval_split_name = 'test'
self.task_cmd_builder.eval_instruct_column = (
self.task_cmd_builder.instruct_column
)
self.task_cmd_builder.eval_template = self.task_cmd_builder.template
self.docker_builder.add_env_var('CUDA_VISIBLE_DEVICES', '0,1,2,3,4,5,6,7')
self.assertEqual(self.run_cmd(), 0)
@parameterized.named_parameters(
('llama3-70b', 'llama3-70b-hf'),
('llama3.1-70b', 'llama3.1-70b-hf'),
)
def test_70b_model_fsdp(self, model_name):
self.setup_output_dir(f'test_fsdp_{model_name}')
self.task_cmd_builder.pretrained_model_id = (
test_util.get_pretrained_model_id(model_name)
)
self.task_cmd_builder.config_file = (
'vertex_vision_model_garden_peft/llama_fsdp_8gpu.yaml'
)
self.task_cmd_builder.train_dataset = 'timdettmers/openassistant-guanaco'
self.task_cmd_builder.train_split_name = 'train'
self.task_cmd_builder.instruct_column = 'text'
self.task_cmd_builder.template = 'openassistant-guanaco'
self.task_cmd_builder.eval_dataset = self.task_cmd_builder.train_dataset
self.task_cmd_builder.eval_split_name = 'test'
self.task_cmd_builder.eval_instruct_column = (
self.task_cmd_builder.instruct_column
)
self.task_cmd_builder.eval_template = self.task_cmd_builder.template
self.docker_builder.add_env_var('CUDA_VISIBLE_DEVICES', '0,1,2,3,4,5,6,7')
self.assertEqual(self.run_cmd(), 0)
if __name__ == '__main__':
absltest.main()
@@ -0,0 +1,49 @@
# pylint: disable=missing-function-docstring
# pylint: disable=missing-class-docstring
"""Tests quantize model task in PEFT docker."""
import os
import time
from absl.testing import absltest
import quantize_model_command_builder as task_cmd_builder
import test_util
class QuantizeModelTest(test_util.TestBase):
def setUp(self):
super().setUp()
self.docker_builder.add_env_var('CUDA_VISIBLE_DEVICES', '')
self.docker_builder.add_mount_map(
os.path.expanduser('~'), os.path.expanduser('~')
)
self.task_cmd_builder = task_cmd_builder.QuantizeModelCommandBuilder()
self.task_cmd_builder.task = 'quantize-model'
self.task_cmd_builder.pretrained_model_id = (
'gs://vertex-model-garden-public-us/llama3/llama3-8b-hf'
)
self.task_cmd_builder.quantization_method = 'awq'
self.task_cmd_builder.quantization_precision_mode = '4bit'
self.task_cmd_builder.quantization_dataset_name = 'pileval'
self.task_cmd_builder.text_column_in_quantization_dataset = 'text'
self.task_cmd_builder.quantization_output_dir = '~/llama3-8b-hf-quantized'
self.task_cmd_builder.device_map = None
self.task_cmd_builder.max_memory = None
self.task_cmd_builder.group_size = 128
self.task_cmd_builder.desc_act = False
self.task_cmd_builder.damp_percent = 0.1
self.task_cmd_builder.cache_examples_on_gpu = False
self.task_cmd_builder.awq_version = 'GEMM'
def test_llama3_8b_model_awq_quantization(self):
start_time = time.time()
self.assertEqual(self.run_cmd(), 0)
end_time = time.time()
self.assertLess(end_time - start_time, 1.5 * 60 * 60)
if __name__ == '__main__':
absltest.main()
@@ -0,0 +1,134 @@
"""Test util class."""
import datetime
import os
import signal
import subprocess
import sys
from absl import flags
from absl import logging
from absl.testing import parameterized
import docker_command_builder as docker_cmd_builder
_DOCKER_URI = flags.DEFINE_string(
'docker_uri', None, 'docker image uri', required=True
)
_DRY_RUN = flags.DEFINE_bool('dry_run', False, 'dry-run the commands')
_LOCAL_INPUT_DIR = flags.DEFINE_string(
'local_input_dir',
os.path.expanduser('~/test_input'),
'local directory for storing input data.',
)
_LOCAL_OUTPUT_DIR = flags.DEFINE_string(
'local_output_dir',
'/tmp',
'local directory for storing test output.',
)
_GCS_INPUT_DIR = flags.DEFINE_string(
'gcs_input_dir',
'gs://peft-docker-test',
'GCS directory that stores model checkpoint, dataset and etc.',
)
_GCS_OUTPUT_DIR = flags.DEFINE_string(
'gcs_output_dir',
'gs://peft-docker-test/output',
'GCS directory that stores test output.',
)
class TestBase(parameterized.TestCase):
"""Test base class that defines how to run commands."""
def setUp(self):
super().setUp()
self.docker_builder = docker_cmd_builder.DockerCommandBuilder(
_DOCKER_URI.value
)
self.docker_builder.add_mount_map(
os.path.expanduser('~'), os.path.expanduser('~')
)
self.docker_builder.add_mount_map(
self.local_input_dir(), self.local_input_dir()
)
self.task_cmd_builder = None
def cmd(self):
return self.docker_builder.build_cmd() + self.task_cmd_builder.build_cmd()
def run_cmd(self) -> int:
logging.info('running command: \n%s', ' \\\n'.join(self.cmd()))
if _DRY_RUN.value:
return 0
p = subprocess.Popen(self.cmd(), stdout=sys.stdout, stderr=sys.stderr)
try:
unused_output, unused_error = p.communicate()
return p.returncode
except KeyboardInterrupt:
p.send_signal(signal.SIGINT)
return 0
def gcs_output_dir(self):
return _GCS_OUTPUT_DIR.value
def local_output_dir(self):
return _LOCAL_OUTPUT_DIR.value
def local_input_dir(self):
return _LOCAL_INPUT_DIR.value
def get_timestamp():
return datetime.datetime.now(datetime.timezone.utc).strftime(
'%Y%m%d_%H%M%S%Z'
)
def get_test_data_path(name: str, download: bool = True) -> str:
"""Gets test data path.
Args:
name: name of the test data
download: if True, then download data from GCS and returns its local path.
Returns:
test data path.
"""
def _download_from_gcs(name):
if not os.path.exists(_LOCAL_INPUT_DIR.value):
os.mkdir(_LOCAL_INPUT_DIR.value)
subprocess.check_output([
'gsutil',
'-m',
'cp',
'-r',
os.path.join(_GCS_INPUT_DIR.value, name),
_LOCAL_INPUT_DIR.value,
])
if not download:
return os.path.join(_GCS_INPUT_DIR.value, name)
local_data = os.path.join(_LOCAL_INPUT_DIR.value, name)
if not os.path.exists(local_data):
_download_from_gcs(name)
return local_data
def get_pretrained_model_id(model_id: str) -> str:
# If `model_id` contains `/`, it is assumed to be HF model or model from GCS.
if '/' in model_id:
return model_id
return get_test_data_path(model_id, download=True)
@@ -0,0 +1,327 @@
"""Tests validate the dataset with template task in PEFT docker."""
from absl.testing import absltest
from absl.testing import parameterized
import test_util
import validate_dataset_with_template_command_builder as task_cmd_builder
class ValidateDatasetWithTemplateTest(test_util.TestBase):
"""Test the validate dataset with template task in PEFT docker."""
def setUp(self):
super().setUp()
self.task_cmd_builder = (
task_cmd_builder.ValidateDatasetWithTemplateCommandBuilder()
)
self.task_cmd_builder.task = "validate-dataset-with-template"
@parameterized.named_parameters(
dict(
testcase_name="valid_rows",
validate_top_k_rows=100,
expected_result=0,
),
dict(
testcase_name="negative_rows",
validate_top_k_rows=-10,
expected_result=0,
),
dict(
testcase_name="out_of_range_rows",
validate_top_k_rows=100000,
expected_result=1,
),
)
def test_validate_dataset_with_template_top_k_rows(
self,
validate_top_k_rows,
expected_result,
):
self.task_cmd_builder.dataset_name = "timdettmers/openassistant-guanaco"
self.task_cmd_builder.train_split_name = "train"
self.task_cmd_builder.instruct_column_in_dataset = "text"
self.task_cmd_builder.template = (
"gs://cloud-nas-260507-tmp-20240724/openassistant-guanaco.json"
)
self.task_cmd_builder.validate_percentage_of_dataset = None
self.task_cmd_builder.validate_k_rows_of_dataset = validate_top_k_rows
self.task_cmd_builder.use_multiprocessing = True
result = self.run_cmd()
self.assertEqual(result, expected_result)
@parameterized.named_parameters(
dict(
testcase_name="valid_positive_x_percent",
validate_percentage_of_dataset=10,
expected_result=0,
),
dict(
testcase_name="valid_negative_x_percent",
validate_percentage_of_dataset=-10,
expected_result=0,
),
dict(
testcase_name="invalid_positive_x_percent",
validate_percentage_of_dataset=110,
expected_result=1,
),
dict(
testcase_name="invalid_negative_x_percent",
validate_percentage_of_dataset=-110,
expected_result=1,
),
)
def test_validate_dataset_with_template_x_percent(
self,
validate_percentage_of_dataset,
expected_result,
):
self.task_cmd_builder.dataset_name = "timdettmers/openassistant-guanaco"
self.task_cmd_builder.train_split_name = "train"
self.task_cmd_builder.instruct_column_in_dataset = "text"
self.task_cmd_builder.template = (
"gs://cloud-nas-260507-tmp-20240724/openassistant-guanaco.json"
)
self.task_cmd_builder.validate_percentage_of_dataset = (
validate_percentage_of_dataset
)
self.task_cmd_builder.validate_k_rows_of_dataset = None
self.task_cmd_builder.use_multiprocessing = True
result = self.run_cmd()
self.assertEqual(result, expected_result)
@parameterized.named_parameters(
dict(
testcase_name="invalid_default_input_column",
dataset_name="timdettmers/openassistant-guanaco",
split="train",
input_column="",
template=(
"gs://cloud-nas-260507-tmp-20240724/openassistant-guanaco.json"
),
validate_percentage_of_dataset=None,
validate_top_k_rows=None,
use_multiprocessing=True,
expected_result=1,
),
dict(
testcase_name="invalid_percentage",
dataset_name="timdettmers/openassistant-guanaco",
split="train",
input_column="text",
template=(
"gs://cloud-nas-260507-tmp-20240724/openassistant-guanaco.json"
),
validate_percentage_of_dataset=110,
validate_top_k_rows=None,
use_multiprocessing=True,
expected_result=1,
),
dict(
testcase_name="negative_percentage",
dataset_name="timdettmers/openassistant-guanaco",
split="train",
input_column="text",
template=(
"gs://cloud-nas-260507-tmp-20240724/openassistant-guanaco.json"
),
validate_percentage_of_dataset=-110,
validate_top_k_rows=None,
use_multiprocessing=True,
expected_result=1,
),
dict(
testcase_name="empty_dataset",
dataset_name="",
split="train",
input_column="text",
template=(
"gs://cloud-nas-260507-tmp-20240724/openassistant-guanaco.json"
),
validate_percentage_of_dataset=None,
validate_top_k_rows=None,
use_multiprocessing=True,
expected_result=1,
),
dict(
testcase_name="empty_split",
dataset_name="timdettmers/openassistant-guanaco",
split="",
input_column="text",
template=(
"gs://cloud-nas-260507-tmp-20240724/openassistant-guanaco.json"
),
validate_percentage_of_dataset=None,
validate_top_k_rows=None,
use_multiprocessing=True,
expected_result=1,
),
dict(
testcase_name="empty_template",
dataset_name="timdettmers/openassistant-guanaco",
split="train",
input_column="text",
template="",
validate_percentage_of_dataset=None,
validate_top_k_rows=None,
use_multiprocessing=True,
expected_result=1,
),
dict(
testcase_name="wrong_gcs_template",
dataset_name="gs://cloud-nas-260507-tmp-20240724/model-evaluation/peft_train_sample.jsonl",
split="train",
input_column="text",
template="gs://cloud-nas-260507-tmp-20240724/sample-template.json",
validate_percentage_of_dataset=None,
validate_top_k_rows=None,
use_multiprocessing=True,
expected_result=1,
),
dict(
testcase_name="wrong_gcs_dataset_name",
dataset_name="gs://cloud-nas-260507-tmp-20240724/model-evaluation/peft-train_sample.jsonl",
split="train",
input_column="text",
template="gs://cloud-nas-260507-tmp-20240724/sample_template.json",
validate_percentage_of_dataset=None,
validate_top_k_rows=None,
use_multiprocessing=True,
expected_result=1,
),
)
def test_validate_dataset_with_template_invalid_input(
self,
dataset_name,
split,
input_column,
template,
validate_percentage_of_dataset,
validate_top_k_rows,
use_multiprocessing,
expected_result,
):
self.task_cmd_builder.dataset_name = dataset_name
self.task_cmd_builder.train_split_name = split
self.task_cmd_builder.instruct_column_in_dataset = input_column
self.task_cmd_builder.template = template
self.task_cmd_builder.validate_percentage_of_dataset = (
validate_percentage_of_dataset
)
self.task_cmd_builder.validate_k_rows_of_dataset = validate_top_k_rows
self.task_cmd_builder.use_multiprocessing = use_multiprocessing
result = self.run_cmd()
self.assertEqual(result, expected_result)
@parameterized.named_parameters(
dict(
testcase_name="full_hf_dataset_with_multiprocessing",
dataset_name="timdettmers/openassistant-guanaco",
template=(
"gs://cloud-nas-260507-tmp-20240724/openassistant-guanaco.json"
),
validate_percentage_of_dataset=None,
validate_top_k_rows=None,
use_multiprocessing=True,
expected_result=0,
),
dict(
testcase_name="full_gcs_dataset_with_multiprocessing",
dataset_name="gs://cloud-nas-260507-tmp-20240724/model-evaluation/peft_train_sample.jsonl",
template="gs://cloud-nas-260507-tmp-20240724/sample_template.json",
validate_percentage_of_dataset=None,
validate_top_k_rows=None,
use_multiprocessing=True,
expected_result=0,
),
dict(
testcase_name="half_dataset_with_multiprocessing",
dataset_name="timdettmers/openassistant-guanaco",
template=(
"gs://cloud-nas-260507-tmp-20240724/openassistant-guanaco.json"
),
validate_percentage_of_dataset=50,
validate_top_k_rows=None,
use_multiprocessing=True,
expected_result=0,
),
dict(
testcase_name="top_100_rows_with_multiprocessing",
dataset_name="timdettmers/openassistant-guanaco",
template=(
"gs://cloud-nas-260507-tmp-20240724/openassistant-guanaco.json"
),
validate_percentage_of_dataset=None,
validate_top_k_rows=100,
use_multiprocessing=True,
expected_result=0,
),
dict(
testcase_name="full_gcs_dataset_without_multiprocessing",
dataset_name="gs://cloud-nas-260507-tmp-20240724/model-evaluation/peft_train_sample.jsonl",
template="gs://cloud-nas-260507-tmp-20240724/sample_template.json",
validate_percentage_of_dataset=None,
validate_top_k_rows=None,
use_multiprocessing=False,
expected_result=0,
),
dict(
testcase_name="full_hf_dataset_without_multiprocessing",
dataset_name="timdettmers/openassistant-guanaco",
template=(
"gs://cloud-nas-260507-tmp-20240724/openassistant-guanaco.json"
),
validate_percentage_of_dataset=None,
validate_top_k_rows=None,
use_multiprocessing=False,
expected_result=0,
),
dict(
testcase_name="half_dataset_without_multiprocessing",
dataset_name="timdettmers/openassistant-guanaco",
template=(
"gs://cloud-nas-260507-tmp-20240724/openassistant-guanaco.json"
),
validate_percentage_of_dataset=50,
validate_top_k_rows=None,
use_multiprocessing=False,
expected_result=0,
),
dict(
testcase_name="top_100_rows_without_multiprocessing",
dataset_name="timdettmers/openassistant-guanaco",
template=(
"gs://cloud-nas-260507-tmp-20240724/openassistant-guanaco.json"
),
validate_percentage_of_dataset=None,
validate_top_k_rows=100,
use_multiprocessing=True,
expected_result=0,
),
)
def test_validate_dataset_with_template_multiprocessing_option(
self,
dataset_name,
template,
validate_percentage_of_dataset,
validate_top_k_rows,
use_multiprocessing,
expected_result,
):
self.task_cmd_builder.dataset_name = dataset_name
self.task_cmd_builder.train_split_name = "train"
self.task_cmd_builder.instruct_column_in_dataset = "text"
self.task_cmd_builder.template = template
self.task_cmd_builder.validate_percentage_of_dataset = (
validate_percentage_of_dataset
)
self.task_cmd_builder.validate_k_rows_of_dataset = validate_top_k_rows
self.task_cmd_builder.use_multiprocessing = use_multiprocessing
result = self.run_cmd()
self.assertEqual(result, expected_result)
if __name__ == "__main__":
absltest.main()
@@ -0,0 +1,109 @@
"""Tools to generate CommandBuilder class.
See go/vmg-oss-peft-tests#commandbuilder-class-generation for details.
"""
import argparse
import dataclasses
from typing import List
_DO_NOT_MODIFY_WARNING = """
# DO NOT MODIFY: this file is auto-generated
# See go/vmg-oss-peft-tests#command-builder-genpy
"""
_GETTER_TMPL = """
@property
def {}(self):
return self._{}
"""
_SETTER_TMPL = """
@{}.setter
def {}(self, val: {}):
self._{} = val
"""
_INIT_NAME = """
def __init__(self):"""
_INIT_FIELDS = """
self._{} = None"""
_BUILD_CMD = r"""
def build_cmd(self) -> str:
cmd = []
for k, v in self.__dict__.items():
if v is not None:
cmd.append(f'--{k[1:]}={v}')
return cmd
"""
@dataclasses.dataclass
class FlagInfo:
api_name: str
impl_name: str
arg_type: str
def get_flag_info(line: str) -> FlagInfo:
api_name, impl_name, arg_type = [x.strip() for x in line.split(',')]
return FlagInfo(api_name, impl_name, arg_type)
def gen_getter(info: FlagInfo) -> str:
return _GETTER_TMPL.format(info.api_name, info.impl_name)
def gen_setter(info: FlagInfo) -> str:
return _SETTER_TMPL.format(
info.api_name, info.api_name, info.arg_type, info.impl_name
)
def gen_init(infos: List[FlagInfo]) -> str:
fields = [_INIT_FIELDS.format(i.impl_name) for i in infos]
return ''.join([_INIT_NAME] + fields)
def main():
parser = argparse.ArgumentParser()
parser.add_argument(
'--flags_def', required=True, help='file path contain flags definition.'
)
parser.add_argument(
'--generated_file',
required=True,
help='file path to the generated command builder.',
)
parser.add_argument(
'--class_name',
required=True,
help='class name for command build',
)
args = parser.parse_args()
flags_info = []
with open(args.flags_def, 'r') as flags_f:
for line in flags_f:
if not line.startswith('#'):
flags_info.append(get_flag_info(line))
with open(args.generated_file, 'w') as gen_f:
# Disables pylint messages.
# See https://stackoverflow.com/a/43510297
print('# pylint: disable=W,C,R', file=gen_f)
print(_DO_NOT_MODIFY_WARNING, file=gen_f)
print(f'class {args.class_name}:', file=gen_f)
print(gen_init(flags_info), file=gen_f)
for info in flags_info:
print(gen_getter(info), file=gen_f)
print(gen_setter(info), file=gen_f)
print(_BUILD_CMD, file=gen_f)
print(f'file generated at {args.generated_file}')
if __name__ == '__main__':
main()
@@ -0,0 +1,50 @@
# api_name, impl_name, value_type
# eval related and etc.
config_file, config_file, str
task, task, str
pretrained_model_id, pretrained_model_id, str
train_dataset, dataset_name, str
train_split_name, train_split_name, str
template, template, str
instruct_column, instruct_column_in_dataset, str
ckpt_dir, output_dir, str
merged_model_dir, merge_base_and_lora_output_dir, str
logging_dir, logging_output_dir, str
per_device_batch_size, per_device_train_batch_size, int
gradient_accumulation_steps, gradient_accumulation_steps, int
lora_rank, lora_rank, int
lora_alpha, lora_alpha, int
lora_dropout, lora_dropout, float
max_steps, max_steps, int
num_epochs, num_epochs, float
max_seq_length, max_seq_length, int
learning_rate, learning_rate, float
lr_scheduler_type, lr_scheduler_type, str
load_precision, precision_mode, str
train_precision, train_precision, str
gradient_checkpointing, enable_gradient_checkpointing, bool
example_packing, use_example_packing, bool
attn_implementation, attn_implementation, str
optimizer, optimizer, str
warmup_ratio, warmup_ratio, float
report_to, report_to, str
save_steps, save_steps, int
logging_steps, logging_steps, int
huggingface_access_token, huggingface_access_token, str
eval_dataset, eval_dataset_path, str
eval_instruct_column, eval_column, str
eval_template, eval_template, str
eval_split_name, eval_split, str
eval_steps, eval_steps, int
eval_tasks, eval_tasks, str
eval_metric_name, eval_metric_name, str
completion_only, completion_only, bool
max_grad_norm, max_grad_norm, float
logger_level, logger_level, str
benchmark_out_file, benchmark_out_file, str
tuning_data_stats_file, tuning_data_stats_file, str
enable_peft, enable_peft, bool
merge_model_precision_mode, merge_model_precision_mode, str
target_modules, target_modules, str
@@ -0,0 +1,15 @@
# api_name, impl_name, value_type
task, task, str
pretrained_model_id, pretrained_model_id, str
quantization_method, quantization_method, str
quantization_precision_mode, quantization_precision_mode, str
quantization_dataset_name, quantization_dataset_name, str
text_column_in_quantization_dataset, text_column_in_quantization_dataset, str
quantization_output_dir, quantization_output_dir, str
device_map, device_map, str
max_memory, max_memory, str
group_size, group_size, int
desc_act, desc_act, bool
damp_percent, damp_percent, float
cache_examples_on_gpu, cache_examples_on_gpu, bool
awq_version, awq_version, str
@@ -0,0 +1,9 @@
# api_name, impl_name, value_type
task, task, str
template, template, str
dataset_name, dataset_name, str
train_split_name, train_split_name, str
instruct_column_in_dataset, instruct_column_in_dataset, str
use_multiprocessing, use_multiprocessing, bool
validate_k_rows_of_dataset, validate_k_rows_of_dataset, int
validate_percentage_of_dataset, validate_percentage_of_dataset, int
@@ -0,0 +1,88 @@
# pylint: disable=W,C,R
# DO NOT MODIFY: this file is auto-generated
# See go/vmg-oss-peft-tests#command-builder-genpy
class ValidateDatasetWithTemplateCommandBuilder:
def __init__(self):
self._task = None
self._template = None
self._dataset_name = None
self._train_split_name = None
self._instruct_column_in_dataset = None
self._use_multiprocessing = None
self._validate_k_rows_of_dataset = None
self._validate_percentage_of_dataset = None
@property
def task(self):
return self._task
@task.setter
def task(self, val: str):
self._task = val
@property
def template(self):
return self._template
@template.setter
def template(self, val: str):
self._template = val
@property
def dataset_name(self):
return self._dataset_name
@dataset_name.setter
def dataset_name(self, val: str):
self._dataset_name = val
@property
def train_split_name(self):
return self._train_split_name
@train_split_name.setter
def train_split_name(self, val: str):
self._train_split_name = val
@property
def instruct_column_in_dataset(self):
return self._instruct_column_in_dataset
@instruct_column_in_dataset.setter
def instruct_column_in_dataset(self, val: str):
self._instruct_column_in_dataset = val
@property
def use_multiprocessing(self):
return self._use_multiprocessing
@use_multiprocessing.setter
def use_multiprocessing(self, val: bool):
self._use_multiprocessing = val
@property
def validate_k_rows_of_dataset(self):
return self._validate_k_rows_of_dataset
@validate_k_rows_of_dataset.setter
def validate_k_rows_of_dataset(self, val: int):
self._validate_k_rows_of_dataset = val
@property
def validate_percentage_of_dataset(self):
return self._validate_percentage_of_dataset
@validate_percentage_of_dataset.setter
def validate_percentage_of_dataset(self, val: int):
self._validate_percentage_of_dataset = val
def build_cmd(self) -> str:
cmd = []
for k, v in self.__dict__.items():
if v is not None:
cmd.append(f'--{k[1:]}={v}')
return cmd
@@ -0,0 +1,86 @@
"""Different trainer callbacks for PEFT Trainer."""
import time
from absl import logging
import accelerate
from transformers import TrainingArguments
from transformers.trainer_callback import TrainerCallback
from transformers.trainer_callback import TrainerControl
from transformers.trainer_callback import TrainerState
from vertex_vision_model_garden_peft.train.vmg import utils
class TrainerStatsCallback(TrainerCallback):
"""Trainer callback to report trainer stats."""
def __init__(self, max_seq_length, filename=None):
self._max_seq_length = max_seq_length
self._filename = filename
self._partial_state = accelerate.PartialState()
self._start_time = float('nan')
self._prev_time = float('nan')
self._peak_mem = 0.0
self._avg_throughput = 0.0
def on_step_end(
self,
args: TrainingArguments,
state: TrainerState,
control: TrainerControl,
**kwargs,
):
if self._partial_state.is_main_process:
if state.global_step == 1:
self._prev_time = time.time()
delta_t = float('nan')
else:
cur_time = time.time()
delta_t = cur_time - self._prev_time
self._prev_time = cur_time
self._avg_throughput += (delta_t - self._avg_throughput) / (
state.global_step - 1
)
gpu_stats = utils.gpu_stats()
self._peak_mem = max(gpu_stats.total_mem, self._peak_mem)
logging.info(
'on_step_end: %s, throughput: %.2f s/it',
utils.gpu_stats_str(gpu_stats),
delta_t,
)
def on_train_begin(
self,
args: TrainingArguments,
state: TrainerState,
control: TrainerControl,
**kwargs,
):
if self._partial_state.is_main_process:
self._start_time = time.time()
logging.info('on_train_begin: %s', utils.gpu_stats_str())
def on_train_end(
self,
args: TrainingArguments,
state: TrainerState,
control: TrainerControl,
**kwargs,
):
if self._partial_state.is_main_process:
train_time = time.time() - self._start_time
logging.info(
'training time %.2f s, throughput: %.2f s/it, peak_mem: %.2f GB',
train_time,
self._avg_throughput,
self._peak_mem,
)
if self._filename:
with open(self._filename, 'a') as out_f:
out_f.write(
f'{self._max_seq_length/1024.0:.1f}k | {self._peak_mem:.2f} |'
f' {self._avg_throughput:.2f}\n'
)
@@ -0,0 +1,18 @@
group:
- vertex
task: custom_loglikelihood
dataset_path: json
dataset_name: null
output_type: loglikelihood
training_split: null
validation_split: null
test_split: test
doc_to_text: "Request: {{prompt}}\nResponse:"
doc_to_target: " {{ground_truth}}"
metric_list:
- metric: perplexity
aggregation: perplexity
higher_is_better: false
- metric: acc
aggregation: mean
higher_is_better: true
@@ -0,0 +1,17 @@
compute_environment: LOCAL_MACHINE
debug: false
distributed_type: MULTI_GPU
downcast_bf16: 'no'
enable_cpu_affinity: false
gpu_ids: all
machine_rank: 0
main_training_function: main
mixed_precision: fp16
num_machines: 1
num_processes: 4
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
@@ -0,0 +1,17 @@
compute_environment: LOCAL_MACHINE
debug: false
distributed_type: MULTI_GPU
downcast_bf16: 'no'
enable_cpu_affinity: false
gpu_ids: all
machine_rank: 0
main_training_function: main
mixed_precision: fp16
num_machines: 1
num_processes: 8
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
@@ -0,0 +1,17 @@
compute_environment: LOCAL_MACHINE
debug: false
deepspeed_config:
deepspeed_config_file: /diffusers/examples/vertex_vision_model_garden_peft/zero2.json
zero3_init_flag: true
distributed_type: DEEPSPEED
downcast_bf16: 'no'
machine_rank: 0
main_training_function: main
num_machines: 1
num_processes: 4
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
@@ -0,0 +1,17 @@
compute_environment: LOCAL_MACHINE
debug: false
deepspeed_config:
deepspeed_config_file: /diffusers/examples/vertex_vision_model_garden_peft/zero2.json
zero3_init_flag: true
distributed_type: DEEPSPEED
downcast_bf16: 'no'
machine_rank: 0
main_training_function: main
num_machines: 1
num_processes: 8
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
@@ -0,0 +1,17 @@
compute_environment: LOCAL_MACHINE
debug: false
deepspeed_config:
deepspeed_config_file: /diffusers/examples/vertex_vision_model_garden_peft/zero3.json
zero3_init_flag: true
distributed_type: DEEPSPEED
downcast_bf16: 'no'
machine_rank: 0
main_training_function: main
num_machines: 1
num_processes: 4
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
@@ -0,0 +1,17 @@
compute_environment: LOCAL_MACHINE
debug: false
deepspeed_config:
deepspeed_config_file: /diffusers/examples/vertex_vision_model_garden_peft/zero3.json
zero3_init_flag: true
distributed_type: DEEPSPEED
downcast_bf16: 'no'
machine_rank: 0
main_training_function: main
num_machines: 1
num_processes: 8
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
@@ -0,0 +1,28 @@
compute_environment: LOCAL_MACHINE
debug: false
distributed_type: FSDP
downcast_bf16: 'no'
enable_cpu_affinity: false
fsdp_config:
fsdp_auto_wrap_policy: TRANSFORMER_BASED_WRAP
fsdp_transformer_layer_cls_to_wrap: LlamaDecoderLayer
fsdp_backward_prefetch: NO_PREFETCH
fsdp_cpu_ram_efficient_loading: true
fsdp_forward_prefetch: false
fsdp_offload_params: true
fsdp_sharding_strategy: FULL_SHARD
fsdp_state_dict_type: SHARDED_STATE_DICT
fsdp_sync_module_states: true
fsdp_use_orig_params: false
fsdp_activation_checkpointing: false
main_training_function: main
mixed_precision: bf16
machine_rank: 0
num_machines: 16
num_processes: 128
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
@@ -0,0 +1,28 @@
compute_environment: LOCAL_MACHINE
debug: false
distributed_type: FSDP
downcast_bf16: 'no'
enable_cpu_affinity: false
fsdp_config:
fsdp_auto_wrap_policy: TRANSFORMER_BASED_WRAP
fsdp_transformer_layer_cls_to_wrap: LlamaDecoderLayer
fsdp_backward_prefetch: NO_PREFETCH
fsdp_cpu_ram_efficient_loading: true
fsdp_forward_prefetch: false
fsdp_offload_params: true
fsdp_sharding_strategy: FULL_SHARD
fsdp_state_dict_type: SHARDED_STATE_DICT
fsdp_sync_module_states: true
fsdp_use_orig_params: false
fsdp_activation_checkpointing: false
main_training_function: main
mixed_precision: bf16
machine_rank: 0
num_machines: 2
num_processes: 16
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
@@ -0,0 +1,28 @@
compute_environment: LOCAL_MACHINE
debug: false
distributed_type: FSDP
downcast_bf16: 'no'
enable_cpu_affinity: false
fsdp_config:
fsdp_auto_wrap_policy: TRANSFORMER_BASED_WRAP
fsdp_transformer_layer_cls_to_wrap: LlamaDecoderLayer
fsdp_backward_prefetch: NO_PREFETCH
fsdp_cpu_ram_efficient_loading: true
fsdp_forward_prefetch: false
fsdp_offload_params: true
fsdp_sharding_strategy: FULL_SHARD
fsdp_state_dict_type: SHARDED_STATE_DICT
fsdp_sync_module_states: true
fsdp_use_orig_params: false
fsdp_activation_checkpointing: false
main_training_function: main
mixed_precision: bf16
machine_rank: 0
num_machines: 3
num_processes: 24
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
@@ -0,0 +1,28 @@
compute_environment: LOCAL_MACHINE
debug: false
distributed_type: FSDP
downcast_bf16: 'no'
enable_cpu_affinity: false
fsdp_config:
fsdp_auto_wrap_policy: TRANSFORMER_BASED_WRAP
fsdp_transformer_layer_cls_to_wrap: LlamaDecoderLayer
fsdp_backward_prefetch: NO_PREFETCH
fsdp_cpu_ram_efficient_loading: true
fsdp_forward_prefetch: false
fsdp_offload_params: true
fsdp_sharding_strategy: FULL_SHARD
fsdp_state_dict_type: SHARDED_STATE_DICT
fsdp_sync_module_states: true
fsdp_use_orig_params: false
fsdp_activation_checkpointing: false
main_training_function: main
mixed_precision: bf16
machine_rank: 0
num_machines: 4
num_processes: 32
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
@@ -0,0 +1,28 @@
compute_environment: LOCAL_MACHINE
debug: false
distributed_type: FSDP
downcast_bf16: 'no'
enable_cpu_affinity: false
fsdp_config:
fsdp_auto_wrap_policy: TRANSFORMER_BASED_WRAP
fsdp_transformer_layer_cls_to_wrap: LlamaDecoderLayer
fsdp_backward_prefetch: NO_PREFETCH
fsdp_cpu_ram_efficient_loading: true
fsdp_forward_prefetch: false
fsdp_offload_params: true
fsdp_sharding_strategy: FULL_SHARD
fsdp_state_dict_type: SHARDED_STATE_DICT
fsdp_sync_module_states: true
fsdp_use_orig_params: false
fsdp_activation_checkpointing: false
main_training_function: main
mixed_precision: bf16
machine_rank: 0
num_machines: 1
num_processes: 8
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
@@ -0,0 +1,28 @@
compute_environment: LOCAL_MACHINE
debug: false
distributed_type: FSDP
downcast_bf16: 'no'
enable_cpu_affinity: false
fsdp_config:
fsdp_auto_wrap_policy: TRANSFORMER_BASED_WRAP
fsdp_transformer_layer_cls_to_wrap: LlamaDecoderLayer
fsdp_backward_prefetch: NO_PREFETCH
fsdp_cpu_ram_efficient_loading: true
fsdp_forward_prefetch: false
fsdp_offload_params: true
fsdp_sharding_strategy: HYBRID_SHARD
fsdp_state_dict_type: SHARDED_STATE_DICT
fsdp_sync_module_states: true
fsdp_use_orig_params: false
fsdp_activation_checkpointing: false
main_training_function: main
mixed_precision: bf16
machine_rank: 0
num_machines: 2
num_processes: 16
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
@@ -0,0 +1,28 @@
compute_environment: LOCAL_MACHINE
debug: false
distributed_type: FSDP
downcast_bf16: 'no'
enable_cpu_affinity: false
fsdp_config:
fsdp_auto_wrap_policy: TRANSFORMER_BASED_WRAP
fsdp_transformer_layer_cls_to_wrap: LlamaDecoderLayer
fsdp_backward_prefetch: NO_PREFETCH
fsdp_cpu_ram_efficient_loading: true
fsdp_forward_prefetch: false
fsdp_offload_params: true
fsdp_sharding_strategy: HYBRID_SHARD
fsdp_state_dict_type: SHARDED_STATE_DICT
fsdp_sync_module_states: true
fsdp_use_orig_params: false
fsdp_activation_checkpointing: false
main_training_function: main
mixed_precision: bf16
machine_rank: 0
num_machines: 3
num_processes: 24
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
@@ -0,0 +1,28 @@
compute_environment: LOCAL_MACHINE
debug: false
distributed_type: FSDP
downcast_bf16: 'no'
enable_cpu_affinity: false
fsdp_config:
fsdp_auto_wrap_policy: TRANSFORMER_BASED_WRAP
fsdp_transformer_layer_cls_to_wrap: LlamaDecoderLayer
fsdp_backward_prefetch: NO_PREFETCH
fsdp_cpu_ram_efficient_loading: true
fsdp_forward_prefetch: false
fsdp_offload_params: true
fsdp_sharding_strategy: HYBRID_SHARD
fsdp_state_dict_type: SHARDED_STATE_DICT
fsdp_sync_module_states: true
fsdp_use_orig_params: false
fsdp_activation_checkpointing: false
main_training_function: main
mixed_precision: bf16
machine_rank: 0
num_machines: 4
num_processes: 32
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
@@ -0,0 +1,24 @@
{
"zero_optimization": {
"stage": 2,
"contiguous_gradients": false,
"overlap_comm": false
},
"bf16": {
"enabled": "auto"
},
"fp16": {
"enabled": "auto",
"auto_cast": false,
"loss_scale": 0,
"initial_scale_power": 32,
"loss_scale_window": 1000,
"hysteresis": 2,
"min_loss_scale": 1
},
"gradient_accumulation_steps": "auto",
"gradient_clipping": "auto",
"train_batch_size": "auto",
"train_micro_batch_size_per_gpu": "auto",
"wall_clock_breakdown": false
}
@@ -0,0 +1,31 @@
{
"zero_optimization": {
"stage": 3,
"overlap_comm": false,
"contiguous_gradients": false,
"sub_group_size": 0,
"reduce_bucket_size": "auto",
"stage3_prefetch_bucket_size": "auto",
"stage3_param_persistence_threshold": "auto",
"stage3_max_live_parameters": 0,
"stage3_max_reuse_distance": 0,
"stage3_gather_16bit_weights_on_model_save": true
},
"bf16": {
"enabled": "auto"
},
"fp16": {
"enabled": "auto",
"auto_cast": false,
"loss_scale": 0,
"initial_scale_power": 32,
"loss_scale_window": 1000,
"hysteresis": 2,
"min_loss_scale": 1
},
"gradient_accumulation_steps": "auto",
"gradient_clipping": "auto",
"train_batch_size": "auto",
"train_micro_batch_size_per_gpu": "auto",
"wall_clock_breakdown": false
}
@@ -0,0 +1,43 @@
# Doc about format of conda environment file
# https://conda.io/projects/conda/en/latest/user-guide/tasks/manage-environments.html#create-env-file-manually
name: merge
channels:
- nodefaults
- conda-forge
dependencies:
- _libgcc_mutex=0.1=conda_forge
- _openmp_mutex=4.5=2_gnu
- bzip2=1.0.8=h4bc722e_7
- ca-certificates=2024.7.4=hbcca054_0
- ld_impl_linux-64=2.40=hf3520f5_7
- libffi=3.4.2=h7f98852_5
- libgcc-ng=14.1.0=h77fa898_0
- libgomp=14.1.0=h77fa898_0
- libnsl=2.0.1=hd590300_0
- libsqlite=3.46.0=hde9e2c9_0
- libuuid=2.38.1=h0b41bf4_0
- libxcrypt=4.4.36=hd590300_1
- libzlib=1.3.1=h4ab18f5_1
- ncurses=6.5=h59595ed_0
- openssl=3.3.1=h4bc722e_2
- pip=24.2=pyhd8ed1ab_0
- python=3.10.14=hd12c33a_0_cpython
- readline=8.2=h8228510_1
- setuptools=72.1.0=pyhd8ed1ab_0
- tk=8.6.13=noxft_h4845f30_101
- tzdata=2024a=h0c530f3_0
- wheel=0.44.0=pyhd8ed1ab_0
- xz=5.2.6=h166bdaf_0
- pip:
- --extra-index-url https://download.pytorch.org/whl/cu121
- absl-py==2.1.0
- accelerate==0.33.0 # Needed for fp8
- datasets==2.19.2
- fbgemm-gpu==0.8.0+cu121 # Needed for fp8
- kfp==2.5.0
- peft==0.12.0
- protobuf==3.20.3
- pynvml==11.5.3
- torch==2.4.0+cu121 # Needed for fp8
- transformers==4.43.1
- trl==0.9.6
@@ -0,0 +1,27 @@
# Doc about format of requirement file
# https://pip.pypa.io/en/stable/reference/requirements-file-format
--extra-index-url https://download.pytorch.org/whl/cu118
--extra-index-url https://huggingface.github.io/autogptq-index/whl/cu118/
# keep sorted
accelerate==0.31.0
auto_gptq==0.7.1+cu118
autoawq==0.2.5
bitsandbytes==0.43.2
cloudml-hypertune==0.1.0.dev6
datasets==2.19.2
deepspeed==0.14.4
diffusers==0.25.1
fsspec==2024.3.1
gcsfs==2024.3.1
lm_eval==0.4.3
ninja==1.11.1 # Needed to avoid `ninja 1.11.1.1 is not supported on this platform` error
optimum==1.17.1
peft==0.12.0
pynvml==11.5.3
torch==2.2.2+cu118
torchvision==0.17.2+cu118
transformers==4.43.1
trl==0.9.6
wandb==0.17.1
@@ -0,0 +1,76 @@
# Dockerfile for PEFT Training.
#
# To build:
# docker build -f model_oss/peft/dockerfile/train.Dockerfile . -t ${YOUR_IMAGE_TAG}
#
# To push to gcr:
# docker tag ${YOUR_IMAGE_TAG} gcr.io/${YOUR_PROJECT}/${YOUR_IMAGE_TAG}
# docker push gcr.io/${YOUR_PROJECT}/${YOUR_IMAGE_TAG}
# Picked from https://cloud.google.com/deep-learning-containers/docs/choosing-container#pytorch
FROM us-docker.pkg.dev/deeplearning-platform-release/gcr.io/pytorch-cu121.2-2.py310:m123
RUN apt-get update && \
apt-get upgrade -y && \
apt-get install -y curl git wget software-properties-common vim libaio-dev && \
apt-get clean && \
rm -rf /var/lib/apt/lists*
# Copy license.
RUN wget https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/main/LICENSE
# Install libraries.
ENV PIP_ROOT_USER_ACTION=ignore
RUN pip install --upgrade pip
# Prefer to install with requirement file as much as possible for reasons
# described in b/355034754.
COPY model_oss/peft/train/vmg/dockerfile/requirements.txt /tmp/requirements.txt
RUN pip install -r /tmp/requirements.txt
# flash-attn cannot be installed with the requirement file approach above
# because of the `no-build-isolation` requirement.
#
# It is OK to install it after other packages FOR NOW because it only has
# limited dependencies. And there's no concern about it overwriting previously
# installed packages.
# https://github.com/Dao-AILab/flash-attention/blob/v2.6.3/setup.py#L523
RUN pip install flash-attn==2.6.3 --no-build-isolation
# Install `diffusers` library as editable and in root folder (/) on purpose.
RUN git clone --depth 1 --branch v0.25.1 https://github.com/huggingface/diffusers.git
# Remove `diffusers` (NOTE that the dependency libraries are kept).
RUN pip uninstall -y diffusers
# Using `--no-deps` option to make sure previously installed packages are not
# overwritten.
RUN pip install --no-deps -e /diffusers
# Make sure there's no inconsistent pip libraries.
RUN pip check
# Install merge related packages in a separate env.
COPY model_oss/peft/train/vmg/dockerfile/merge_env.yaml /tmp/merge_env.yaml
RUN conda env create -n merge --yes --file /tmp/merge_env.yaml
RUN conda init
# Switch to diffusers examples folder.
WORKDIR /diffusers/examples
RUN mkdir -p ./vertex_vision_model_garden_peft/
COPY model_oss/peft/train/vmg/configs/* ./vertex_vision_model_garden_peft/
# custom `lm_eval` task.
ARG LM_EVAL_DIR=$(python -c 'import site; print(site.getsitepackages()[0])')/lm_eval
RUN mkdir -p $LM_EVAL_DIR/tasks/vertex && \
mv ./vertex_vision_model_garden_peft/custom_loglikelihood.yaml $LM_EVAL_DIR/tasks/vertex/
COPY model_oss/peft/train/vmg/*.py ./vertex_vision_model_garden_peft/train/vmg/
COPY model_oss/peft/train/vmg/templates /diffusers/examples/util/templates
COPY model_oss/util /diffusers/examples/util
COPY model_oss/notebook_util/dataset_validation_util.py /diffusers/examples/util
COPY model_oss/peft/train/tests/*.py ./vertex_vision_model_garden_peft/tests/
RUN chmod a+rwX -R /diffusers/examples/
ENV PYTHONPATH /diffusers/examples/
# Must disable torch XLA, otherwise runtime uses CPU even if GPU exists.
ENV USE_TORCH_XLA 0
ENTRYPOINT ["python3", "./vertex_vision_model_garden_peft/train/vmg/train_entrypoint.py"]
@@ -0,0 +1,183 @@
"""Library for running evaluations during training."""
import dataclasses
from typing import Any, Optional, Type
from absl import logging
import datasets
from lm_eval import evaluator
from lm_eval import tasks
from lm_eval import utils
from lm_eval.api import model as lm_model
from lm_eval.api import registry
from lm_eval.models import huggingface
from peft import peft_model
import transformers
from transformers import trainer
from util import dataset_validation_util
from util import constants
_DESCRIPTION_EVALUATION = "evaluation"
_BUILTIN_EVAL_TASK = "builtin_eval"
@dataclasses.dataclass(frozen=True)
class EvalConfig:
steps: int
tasks: list[str]
per_device_batch_size: int
num_fewshot: Optional[int]
limit: Optional[float]
metric_name: str
tokenize_dataset: bool
dataset_path: str = ""
split: str = "test"
template: str = ""
column: str = constants.DEFAULT_INSTRUCT_COLUMN_IN_DATASET
class PeftCausalLMModel(huggingface.HFLM):
"""PeftCausalLMModel that supports loading an in-memory model."""
AUTO_MODEL_CLASS = transformers.AutoModelForCausalLM
def __init__(
self,
model: peft_model.PeftModelForCausalLM,
tokenizer: transformers.PreTrainedTokenizerBase,
batch_size_per_gpu: int,
):
lm_model.LM.__init__(self)
self._model = model
self.tokenizer = tokenizer
self.vocab_size = tokenizer.vocab_size
tokenizer.pad_token_id = tokenizer.eos_token_id
self._config = model.config
self.batch_size_per_gpu = batch_size_per_gpu
self._device = model.device
self._max_length = None # Will be automatically determined from config.
self._add_special_tokens = (
None # Will be automatically determined from AUTO_MODEL_CLASS.
)
def create_trainer(
cls: Type[transformers.Trainer],
eval_config: Optional[EvalConfig],
tokenizer: Optional[transformers.PreTrainedTokenizerBase],
args: trainer.TrainingArguments,
**kwargs,
) -> transformers.Trainer:
"""Creates a trainer. If eval config is provided, injects evaluation loop."""
if not eval_config:
return cls(args=args, **kwargs)
args.eval_strategy = "steps"
args.eval_steps = eval_config.steps
args.per_device_eval_batch_size = eval_config.per_device_batch_size
kwargs["tokenizer"] = tokenizer
if eval_config.tasks == [_BUILTIN_EVAL_TASK]:
try:
eval_dataset = dataset_validation_util.load_dataset_with_template(
dataset_name=eval_config.dataset_path,
split=eval_config.split,
input_column=eval_config.column,
template=eval_config.template,
tokenizer=tokenizer,
)
if eval_config.limit is not None:
if eval_config.limit >= 1:
limit = int(eval_config.limit)
else:
limit = int(eval_config.limit * len(eval_dataset))
eval_dataset = eval_dataset.select(range(limit))
if eval_config.tokenize_dataset:
eval_dataset = eval_dataset.map(
lambda samples: tokenizer(samples[eval_config.column])
)
kwargs["eval_dataset"] = eval_dataset
except (OSError, ValueError, IndexError) as e:
logging.warning(
"Failed to load eval dataset %s. Evaluation will be skipped.\n%s",
eval_config.dataset_path,
e,
)
del args.evaluation_strategy
del args.eval_steps
del args.per_device_eval_batch_size
return cls(args=args, **kwargs)
class LMEvalTrainer(cls):
"""Trainer with lm_eval injected as the eval library."""
def __init__(self, **kwargs):
super().__init__(**kwargs)
task_names = utils.pattern_match(eval_config.tasks, registry.ALL_TASKS)
logging.info("Selected Eval Tasks: %s", task_names)
task_args = {}
if eval_config.num_fewshot is not None:
task_args["num_fewshot"] = eval_config.num_fewshot
if eval_config.dataset_path:
task_args["dataset_path"] = "json"
task_args["dataset_kwargs"] = {
"data_files": {"test": eval_config.dataset_path},
}
self._eval_task_dict = tasks.get_task_dict(task_names, **task_args)
def evaluation_loop(
self,
dataloader: trainer.DataLoader,
description: str,
prediction_loss_only: Optional[bool] = None,
ignore_keys: Optional[list[str]] = None,
metric_key_prefix: str = "eval",
) -> trainer.EvalLoopOutput:
"""Custom evaluation loop that invokes lm_eval."""
if description.lower() != _DESCRIPTION_EVALUATION:
return super().evaluation_loop(
dataloader,
description,
prediction_loss_only,
ignore_keys,
metric_key_prefix,
)
model = self._wrap_model(self.model, training=False)
lm = PeftCausalLMModel(
model,
self.tokenizer or self.data_collator.tokenizer,
eval_config.per_device_batch_size,
)
results: dict[str, Any] = evaluator.evaluate(
lm=lm,
task_dict=self._eval_task_dict,
limit=eval_config.limit,
)["results"]
metric_name = eval_config.metric_name
# Compute average value if there are multiple tasks.
metric_values: list[float] = []
for result in results.values():
for key, value in result.items():
if key.split(",")[0] == metric_name:
metric_values.append(value)
if not metric_values:
raise ValueError(
f"Metric {metric_name} not found in eval response: {results}"
)
metric_average = sum(metric_values) / len(metric_values)
logging.info("%s value: %f\n%s", metric_name, metric_average, results)
return trainer.EvalLoopOutput(
# Only metrics field is set. Other fields are dummy values.
predictions=None,
label_ids=None,
metrics={f"{metric_key_prefix}_{metric_name}": metric_average},
num_samples=0,
)
# Use empty eval dataset as a placeholder.
return LMEvalTrainer(
args=args, eval_dataset=datasets.Dataset.from_dict({"test": []}), **kwargs
)
@@ -0,0 +1,794 @@
"""Instruct/Chat with LoRA models."""
import dataclasses
import datetime
import json
import os
from typing import Any, Dict, Optional, Sequence
import warnings
from absl import app
from absl import flags
from absl import logging
from accelerate import DistributedType
from accelerate import PartialState
import bitsandbytes as bnb
import hypertune
from peft import get_peft_model
from peft import LoraConfig
import torch
from transformers import AutoModelForCausalLM
from transformers import TrainingArguments
from trl import DataCollatorForCompletionOnlyLM
from trl import SFTTrainer
import wandb
from util import dataset_validation_util
from vertex_vision_model_garden_peft.train.vmg import callbacks
from vertex_vision_model_garden_peft.train.vmg import eval_lib
from vertex_vision_model_garden_peft.train.vmg import utils
from util import constants
from util import fileutils
_PRETRAINED_MODEL_ID = flags.DEFINE_string(
'pretrained_model_id',
None,
'The pretrained model id. Supported models can be causal language modeling'
' models from https://github.com/huggingface/peft/tree/main. Note, there'
' might be different paddings for different models. This tool assumes the'
' pretrained_model_id contains model name, and then choose proper padding'
' methods. e.g. it must contain `llama` for `Llama2 models`.',
required=True,
)
_HUGGINGFACE_ACCESS_TOKEN = flags.DEFINE_string(
'huggingface_access_token',
None,
'The access token for loading huggingface gated models.',
)
_DATASET_NAME = flags.DEFINE_string(
'dataset_name',
None,
'The dataset name in huggingface.',
)
_OUTPUT_DIR = flags.DEFINE_string(
'output_dir',
None,
'The output directory.',
)
_LOGGING_OUTPUT_DIR = flags.DEFINE_string(
'logging_output_dir',
'',
'The logging output directory, which defaults to same as output_dir.',
)
_PRECISION_MODE = flags.DEFINE_enum(
'precision_mode',
constants.PRECISION_MODE_16,
[
constants.PRECISION_MODE_4,
constants.PRECISION_MODE_8,
constants.PRECISION_MODE_16,
constants.PRECISION_MODE_16B,
constants.PRECISION_MODE_32,
],
'Precision to load model weights for finetuning.',
)
_LORA_RANK = flags.DEFINE_integer(
'lora_rank',
16,
'The rank of the update matrices, expressed in int. Lower rank results in'
' smaller update matrices with fewer trainable parameters, referring to'
' https://huggingface.co/docs/peft/conceptual_guides/lora.',
)
_LORA_ALPHA = flags.DEFINE_integer(
'lora_alpha',
32,
'LoRA scaling factor, referring to'
' https://huggingface.co/docs/peft/conceptual_guides/lora.',
)
_LORA_DROPOUT = flags.DEFINE_float(
'lora_dropout',
0.05,
'dropout probability of the LoRA layers, referring to'
' https://huggingface.co/docs/peft/task_guides/token-classification-lora.',
)
_WARMUP_STEPS = flags.DEFINE_integer(
'warmup_steps',
10,
'Number of steps for the warmup in the learning rate scheduler.',
)
_WARMUP_RATIO = flags.DEFINE_float(
'warmup_ratio',
0.03,
'The warmup ratio in the learning rate scheduler.',
)
_WEIGHT_DECAY = flags.DEFINE_float(
'weight_decay',
0.001,
'The weight decay in the learning rate scheduler.',
)
_NUM_EPOCHS = flags.DEFINE_float(
'num_epochs',
None,
'The number of training epochs. Only used for'
' "sequence-classification-lora" with an integer value and for'
' "instruct-lora" with a float value allowed.',
)
_MAX_STEPS = flags.DEFINE_integer(
'max_steps',
None,
'Total number of training steps. Overrides num_epochs if set. Only used for'
' "instruct-lora."',
)
_MAX_SEQ_LENGTH = flags.DEFINE_integer(
'max_seq_length',
512,
'The maximum sequence length.',
)
_LEARNING_RATE = flags.DEFINE_float(
'learning_rate',
2e-4,
'The learning rate after the potential warmup period.',
)
_INSTRUCT_COLUMN_IN_DATASET = flags.DEFINE_string(
'instruct_column_in_dataset',
constants.DEFAULT_INSTRUCT_COLUMN_IN_DATASET,
'The instruct column in dataset.',
)
_REPORT_TO = flags.DEFINE_string(
'report_to',
constants.REPORT_TO_NONE,
'Where logging is reported to, which can be tensorboard or none.',
)
_PER_DEVICE_TRAIN_BATCH_SIZE = flags.DEFINE_integer(
'per_device_train_batch_size',
4,
'The per device train batch size.',
)
_GRADIENT_ACCUMULATION_STEPS = flags.DEFINE_integer(
'gradient_accumulation_steps',
4,
'The gradient accumulation steps.',
)
_ENABLE_GRADIENT_CHECKPOINTING = flags.DEFINE_boolean(
'enable_gradient_checkpointing',
False,
'Whether to enable gradient checkpointing.',
)
_ENABLE_PEFT = flags.DEFINE_boolean(
'enable_peft',
True,
'Whether to enable peft.',
)
_TEMPLATE = flags.DEFINE_string(
'template',
None,
'Template for formatting language model training data. Must be a filename'
' under `templates` folder, without `.json` extension, e.g. `alpaca`, or a'
' Cloud Storage URI to a JSON file.',
)
_OPTIMIZER = flags.DEFINE_string(
'optimizer',
'adamw_torch',
'The optimizer.',
)
_LR_SCHEDULER_TYPE = flags.DEFINE_string(
'lr_scheduler_type',
'cosine',
'The learning rate scheduler type.',
)
_SAVE_STEPS = flags.DEFINE_integer(
'save_steps',
10,
'The save steps.',
)
_LOGGING_STEPS = flags.DEFINE_integer(
'logging_steps',
10,
'The logging steps.',
)
_EVAL_STEPS = flags.DEFINE_integer(
'eval_steps',
10,
'The number of training steps between evaluations.',
)
_TRAIN_SPLIT_NAME = flags.DEFINE_string(
'train_split_name',
'train',
'The train split name.',
)
_EVAL_TASKS = flags.DEFINE_list(
'eval_tasks',
None,
'List of eval task names (can have wildcards) as in'
' https://github.com/EleutherAI/lm-evaluation-harness. Will not run'
' evaluation if not set. Runs the built-in trainer evaluation loop if set'
' to `builtin_eval`.',
)
_EVAL_PER_DEVICE_BATCH_SIZE = flags.DEFINE_integer(
'eval_per_device_batch_size',
1,
'The per device batch size for model evaluation.',
)
_EVAL_NUM_FEWSHOT = flags.DEFINE_integer(
'eval_num_fewshot',
None,
'Run N-shot language model evaluation. Not implemented in `builtin_eval`.',
)
_EVAL_LIMIT = flags.DEFINE_float(
'eval_limit',
None,
'Limit the number of examples per task. If <1, limit is a percentage of the'
' total number of examples.',
)
_EVAL_METRIC_NAME = flags.DEFINE_string(
'eval_metric_name',
'acc',
'The metric name to aggregate during model evaluation.',
)
_EVAL_DATASET_PATH = flags.DEFINE_string(
'eval_dataset_path',
None,
'Overrides the default evaluation dataset path. In `builtin_eval` mode,'
' this can be any Hugging Face dataset name.',
)
# We set the default eval split as `test`, based on observation from
# https://huggingface.co/datasets/timdettmers/openassistant-guanaco/viewer/default/test.
_EVAL_SPLIT = flags.DEFINE_string(
'eval_split',
'test',
'Eval split name in the eval dataset for `builtin_eval`.',
)
_EVAL_TEMPLATE = flags.DEFINE_string(
'eval_template',
None,
'Template for formatting language model evaluation data for `builtin_eval`.'
' Must be a filename under `templates` folder, without `.json` extension,'
' e.g. `alpaca`, or a Cloud Storage URI to a JSON file.',
)
_EVAL_COLUMN = flags.DEFINE_string(
'eval_column',
None,
'Eval column name in the eval dataset for `builtin_eval`.',
)
_TRAIN_PRECISION = flags.DEFINE_enum(
'train_precision',
constants.PRECISION_MODE_16B,
[
constants.PRECISION_MODE_16,
constants.PRECISION_MODE_16B,
constants.PRECISION_MODE_32,
],
'Precision to train the model.',
)
_USE_EXAMPLE_PACKING = flags.DEFINE_boolean(
'use_example_packing',
False,
'Enables example packing during training, which uses '
'`ConstantLengthDataset` under the hood.',
)
_COMPLETION_ONLY = flags.DEFINE_boolean(
'completion_only',
False,
'If set, it uses DataCollatorForCompletionOnlyLM to train the model on the'
' generated prompts only, i.e., masking out the input',
)
_ATTN_IMPLEMENTATION = flags.DEFINE_string(
'attn_implementation',
None,
'Attention implementation, can be `eager`, `sdpa` or `flash_attention_2`',
)
_MAX_GRAD_NORM = flags.DEFINE_float(
'max_grad_norm',
0.3,
'Maximum gradient norm used for gradient clipping',
)
_WARNINGS_FILTER = flags.DEFINE_string(
'warnings_filter',
'ignore',
'Warning filter as defined in '
'https://docs.python.org/3/library/warnings.html#the-warnings-filter',
)
_LOGGER_LEVEL = flags.DEFINE_string(
'logger_level',
'passive',
'logging level passed to TrainingArguments. Note that this is for python'
' logging module, NOT the one from absl',
)
_BENCHMARK_OUT_FILE = flags.DEFINE_string(
'benchmark_out_file', None, 'file path for writing benchmark result'
)
_NCCL_TIMEOUT = flags.DEFINE_integer(
'nccl_timeout', 6000, 'nccl timeout in seconds'
)
_TUNING_DATA_STATS_FILE = flags.DEFINE_string(
'tuning_data_stats_file', None, 'file path for writing tuning data stats.'
)
_TARGET_MODULES = flags.DEFINE_list(
'target_modules', None, 'The names of the modules to apply LoRA adapter to.'
)
@flags.multi_flags_validator(
[
_COMPLETION_ONLY.name,
_USE_EXAMPLE_PACKING.name,
],
message=(
'`use_example_packing=True` does not work with `completion_only=True`'
),
)
def check_example_packing(flags_dict: Dict[str, Any]) -> bool:
"""Check to make sure example packing is enabled properly.
Args:
flags_dict: Dictionary containing flags to check.
Returns:
If `use_example_packing` is set properly.
"""
if (
flags_dict[_COMPLETION_ONLY.name]
and flags_dict[_USE_EXAMPLE_PACKING.name]
):
return False
return True
@flags.multi_flags_validator(
[
_COMPLETION_ONLY.name,
_TEMPLATE.name,
],
message='`template` should be provided if using `completion_only=True`',
)
def check_completion_only(flags_dict: Dict[str, Any]) -> bool:
"""Check to make sure completion_only is enabled properly.
Args:
flags_dict: Dictionary containing flags to check.
Returns:
If `completion_only` is set properly
"""
if flags_dict[_COMPLETION_ONLY.name] and flags_dict[_TEMPLATE.name] is None:
return False
return True
# References:
# Huggingface SFT trainer example:
# https://github.com/huggingface/trl/blob/main/examples/scripts/sft_trainer.py.
# Huggingface sagemaker example:
# https://github.com/huggingface/notebooks/blob/main/sagemaker/28_train_llms_with_qlora/scripts/run_clm.py.
# Copied from https://github.com/artidoro/qlora/blob/main/qlora.py.
def find_all_linear_names(
model: AutoModelForCausalLM, precision_mode: str
) -> list[str]:
"""Finds all linear module names."""
if precision_mode == constants.PRECISION_MODE_4:
cls = bnb.nn.Linear4bit
elif precision_mode == constants.PRECISION_MODE_8:
cls = bnb.nn.Linear8bitLt
else:
cls = torch.nn.Linear
lora_module_names = set()
for name, module in model.named_modules():
if isinstance(module, cls):
names = name.split('.')
lora_module_names.add(names[0] if len(names) == 1 else names[-1])
if 'lm_head' in lora_module_names: # needed for 16-bit
lora_module_names.remove('lm_head')
return list(lora_module_names)
def finetune_instruct(
pretrained_model_id: str,
dataset_name: str,
output_dir: str,
logging_output_dir: str,
lora_rank: int = 64,
lora_alpha: int = 16,
lora_dropout: float = 0.1,
warmup_ratio: int = 0.03,
num_epochs: Optional[float] = None,
max_steps: Optional[int] = None,
warmup_steps: int = 10,
max_seq_length: int = 512,
learning_rate: float = 2e-4,
precision_mode: str = None,
instruct_column_in_dataset: str = constants.DEFAULT_INSTRUCT_COLUMN_IN_DATASET,
per_device_train_batch_size: int = 4,
gradient_accumulation_steps: int = 4,
optim: str = 'paged_adamw_32bit',
weight_decay: float = 0.001,
enable_gradient_checkpointing: bool = False,
enable_peft: bool = True,
template: str = None,
lr_scheduler_type: str = 'constant',
save_steps: int = 10,
logging_steps: int = 10,
train_split_name: str = 'train',
eval_config: Optional[eval_lib.EvalConfig] = None,
report_to: str = constants.REPORT_TO_NONE,
access_token: Optional[str] = None,
train_precision: str = constants.PRECISION_MODE_16B,
use_example_packing: bool = False,
attn_implementation: Optional[str] = None,
max_grad_norm: float = 0.3,
completion_only: bool = False,
logger_level: str = 'passive',
benchmark_out_file: Optional[str] = None,
tuning_data_stats_file: Optional[str] = None,
target_modules: Optional[str] = None,
) -> None:
"""Finetunes instruct."""
logging.info('on entering instruct_lora, %s', utils.gpu_stats_str())
gradient_checkpointing_kwargs = {}
# DDP provides limited support with the reentrant variant of gradient
# checkpoint [1]. Below is an indirect way of checking whether DDP will be
# used. It is "indirect" because there are complex logic under the hood of
# `SFTTrainer` and since those are not public API, they might change as we
# update the library.
if PartialState().distributed_type == DistributedType.MULTI_GPU:
gradient_checkpointing_kwargs['use_reentrant'] = False
tokenizer = utils.load_tokenizer(
pretrained_model_id,
'right',
access_token=access_token,
)
train_dataset = dataset_validation_util.load_dataset_with_template(
dataset_name,
split=train_split_name,
input_column=instruct_column_in_dataset,
template=template,
tokenizer=tokenizer,
)
if tuning_data_stats_file:
with PartialState().main_process_first():
effective_batch_size = (
per_device_train_batch_size
* gradient_accumulation_steps
* PartialState().num_processes
)
logging.info(
'getting tuning data stats with effective batch size %s',
effective_batch_size,
)
train_dataset_stats = utils.get_dataset_stats(
train_dataset,
tokenizer,
instruct_column_in_dataset,
effective_batch_size,
)
logging.info('stats: %s', train_dataset_stats)
tuning_data_stats_file = dataset_validation_util.force_gcs_fuse_path(
tuning_data_stats_file
)
with open(tuning_data_stats_file, 'w') as out_f:
json.dump(dataclasses.asdict(train_dataset_stats), out_f)
model = utils.load_model(
pretrained_model_id=pretrained_model_id,
tokenizer=tokenizer,
precision_mode=precision_mode,
enable_gradient_checkpointing=enable_gradient_checkpointing,
access_token=access_token,
attn_implementation=attn_implementation,
train_precision=train_precision,
)
if enable_peft:
if target_modules is None:
target_modules = find_all_linear_names(
model, precision_mode=precision_mode
)
logging.info('applying lora adapters to modules: %s', target_modules)
peft_config = LoraConfig(
lora_alpha=lora_alpha,
lora_dropout=lora_dropout,
r=lora_rank,
bias='none',
task_type='CAUSAL_LM',
target_modules=target_modules,
)
# If we pass in `peft_config` to SFTTrainer, it does a lot of magic under
# the hood, e.g., calling `prepare_model_for_kbit_training` before calling
# `get_peft_model`, which may revert other changes we did before. That's why
# we are calling `get_peft_model` explicitly here.
model = get_peft_model(model, peft_config)
# This is to work-around mix-precision training. This issue is not fixed as
# of transformers==4.41.2.
# See b/332760883#comment30 for more details.
if precision_mode in (
constants.PRECISION_MODE_16,
constants.PRECISION_MODE_16B,
):
for param in filter(lambda p: p.requires_grad, model.parameters()):
param.data = param.data.to(torch.float32)
if not logging_output_dir:
logging_output_dir = output_dir
# To use singleton PartialState() without re-initializing it. See
# b/357970482#comment3
accelerator_config = {'use_configured_state': True}
training_arguments = TrainingArguments(
report_to=report_to,
output_dir=output_dir,
per_device_train_batch_size=per_device_train_batch_size,
gradient_accumulation_steps=gradient_accumulation_steps,
optim=optim,
save_steps=save_steps,
save_strategy='steps',
save_total_limit=3,
logging_dir=os.path.join(logging_output_dir, 'logs'),
logging_steps=logging_steps,
learning_rate=learning_rate,
fp16=(train_precision == constants.PRECISION_MODE_16),
bf16=(train_precision == constants.PRECISION_MODE_16B),
max_grad_norm=max_grad_norm,
num_train_epochs=num_epochs if num_epochs else -1,
max_steps=max_steps if max_steps else -1,
warmup_ratio=warmup_ratio,
warmup_steps=warmup_steps,
group_by_length=False,
lr_scheduler_type=lr_scheduler_type,
gradient_checkpointing=enable_gradient_checkpointing,
gradient_checkpointing_kwargs=gradient_checkpointing_kwargs,
weight_decay=weight_decay,
log_level=logger_level,
accelerator_config=accelerator_config,
)
trainer_kwargs = {}
if completion_only and template:
template_json = dataset_validation_util.get_template(template_path=template)
instruction_sep = dataset_validation_util.get_instruction_separator(
template_json
)
response_sep = dataset_validation_util.get_response_separator(template_json)
if not response_sep:
raise ValueError(
'`response_separator` must be provided to use'
' `DataCollatorForCompletionOnlyLM`'
)
trainer_kwargs['data_collator'] = DataCollatorForCompletionOnlyLM(
instruction_template=instruction_sep,
response_template=response_sep,
tokenizer=tokenizer,
)
logging.info('using DataCollatorForCompletionOnlyLM')
trainer_stats_callback = callbacks.TrainerStatsCallback(
max_seq_length, benchmark_out_file
)
trainer = eval_lib.create_trainer(
cls=SFTTrainer,
eval_config=eval_config,
model=model,
train_dataset=train_dataset,
dataset_text_field=instruct_column_in_dataset,
max_seq_length=max_seq_length,
tokenizer=tokenizer,
args=training_arguments,
packing=use_example_packing,
callbacks=[trainer_stats_callback],
**trainer_kwargs,
)
# `eval_lib.create_trainer` might modify the training args. Printing here
# should capture what will be used by the trainer.
if PartialState().is_main_process:
logging.info('training args: %s', trainer.args)
if enable_peft:
trainer.model.print_trainable_parameters()
if trainer.is_fsdp_enabled:
logging.info('Trainer running with FSDP.')
elif trainer.is_deepspeed_enabled:
logging.info('Trainer running with DeepSpeed.')
else:
logging.info('Trainer running without parallelism.')
trainer.train()
# Always save the final checkpoint.
final_checkpoint = utils.get_final_checkpoint_path(output_dir)
logging.info('The final checkpoint is: %s.', final_checkpoint)
if trainer.is_fsdp_enabled:
trainer.accelerator.state.fsdp_plugin.set_state_dict_type('FULL_STATE_DICT')
# This method saves the sharded weights like `accelerator.save_state`, see
# https://huggingface.co/docs/accelerate/en/usage_guides/fsdp#saving-and-loading
trainer.save_model(output_dir)
model = trainer.model.cpu() # Avoids GPU OOM
state_dict = trainer.accelerator.get_state_dict(model)
# To aggregate the weights from all the devices, we need to use
# `state_dict=state_dict`.
model.save_pretrained(
final_checkpoint,
state_dict=state_dict,
is_main_process=PartialState().is_main_process,
save_embedding_layers=False, # Only pad token is added. See go/lora-adapter-pad-token #pylint: disable=line-too-long
)
model.cuda() # Move back to GPU to do eval.
else:
trainer.model.save_pretrained(
final_checkpoint,
is_main_process=PartialState().is_main_process,
save_embedding_layers=False, # Only pad token is added. See go/lora-adapter-pad-token #pylint: disable=line-too-long
)
if eval_config is not None and trainer.eval_dataset is not None:
metrics = trainer.evaluate(metric_key_prefix='eval')
# Both `log_metrics` and `save_metrics` are multiple process safe.
# https://github.com/huggingface/transformers/blob/v4.38.2/src/transformers/trainer_pt_utils.py#L911 #pylint: disable=line-too-long
# https://github.com/huggingface/transformers/blob/v4.38.2/src/transformers/trainer_pt_utils.py#L1001 #pylint: disable=line-too-long
trainer.log_metrics('eval', metrics)
trainer.save_metrics('eval', metrics)
if PartialState().is_main_process:
hp_metric = metrics[f'eval_{eval_config.metric_name}']
hpt = hypertune.HyperTune()
hpt.report_hyperparameter_tuning_metric(
hyperparameter_metric_tag=constants.HP_METRIC_TAG,
metric_value=hp_metric,
)
logging.info('Send HP metric: %f to hyperparameter tuning.', hp_metric)
PartialState().wait_for_everyone()
if not enable_peft:
tokenizer.save_pretrained(
final_checkpoint, is_main_process=PartialState().is_main_process
)
def main(unused_argv: Sequence[str]) -> None:
# This needs to be called before any other PartialState() calls.
utils.init_partial_state(
timeout=datetime.timedelta(seconds=_NCCL_TIMEOUT.value)
)
utils.print_library_versions()
warnings.simplefilter(_WARNINGS_FILTER.value)
pretrained_model_id = fileutils.force_gcs_path(_PRETRAINED_MODEL_ID.value)
if dataset_validation_util.is_gcs_path(pretrained_model_id):
pretrained_model_id = dataset_validation_util.download_gcs_uri_to_local(
pretrained_model_id
)
output_dir = utils.GcsOrLocalDirectory(
_OUTPUT_DIR.value, check_empty=True, upload_from_all_nodes=True
)
# GCS Fuse does not sync flushed files if not closed. See b/361771727.
logging_output_dir = fileutils.force_gcs_path(_LOGGING_OUTPUT_DIR.value)
# Creates evaluation config.
if _EVAL_TASKS.value:
eval_config = eval_lib.EvalConfig(
tasks=_EVAL_TASKS.value,
per_device_batch_size=_EVAL_PER_DEVICE_BATCH_SIZE.value,
num_fewshot=_EVAL_NUM_FEWSHOT.value,
limit=_EVAL_LIMIT.value,
metric_name=_EVAL_METRIC_NAME.value,
steps=_EVAL_STEPS.value,
dataset_path=dataset_validation_util.force_gcs_fuse_path(
_EVAL_DATASET_PATH.value
),
split=_EVAL_SPLIT.value,
template=_EVAL_TEMPLATE.value,
column=_EVAL_COLUMN.value,
tokenize_dataset=False,
)
else:
eval_config = None
if _REPORT_TO.value == constants.REPORT_TO_WANDB:
wandb.login()
finetune_instruct(
pretrained_model_id=pretrained_model_id,
dataset_name=_DATASET_NAME.value,
output_dir=output_dir.local_dir,
logging_output_dir=logging_output_dir,
precision_mode=_PRECISION_MODE.value,
lora_rank=_LORA_RANK.value,
lora_alpha=_LORA_ALPHA.value,
lora_dropout=_LORA_DROPOUT.value,
warmup_ratio=_WARMUP_RATIO.value,
num_epochs=_NUM_EPOCHS.value,
warmup_steps=_WARMUP_STEPS.value,
max_steps=_MAX_STEPS.value,
max_seq_length=_MAX_SEQ_LENGTH.value,
learning_rate=_LEARNING_RATE.value,
instruct_column_in_dataset=_INSTRUCT_COLUMN_IN_DATASET.value,
per_device_train_batch_size=_PER_DEVICE_TRAIN_BATCH_SIZE.value,
optim=_OPTIMIZER.value,
weight_decay=_WEIGHT_DECAY.value,
gradient_accumulation_steps=_GRADIENT_ACCUMULATION_STEPS.value,
enable_gradient_checkpointing=_ENABLE_GRADIENT_CHECKPOINTING.value,
enable_peft=_ENABLE_PEFT.value,
template=_TEMPLATE.value,
lr_scheduler_type=_LR_SCHEDULER_TYPE.value,
save_steps=_SAVE_STEPS.value,
logging_steps=_LOGGING_STEPS.value,
train_split_name=_TRAIN_SPLIT_NAME.value,
eval_config=eval_config,
report_to=_REPORT_TO.value,
access_token=_HUGGINGFACE_ACCESS_TOKEN.value,
train_precision=_TRAIN_PRECISION.value,
use_example_packing=_USE_EXAMPLE_PACKING.value,
attn_implementation=_ATTN_IMPLEMENTATION.value,
max_grad_norm=_MAX_GRAD_NORM.value,
completion_only=_COMPLETION_ONLY.value,
logger_level=_LOGGER_LEVEL.value,
benchmark_out_file=_BENCHMARK_OUT_FILE.value,
tuning_data_stats_file=_TUNING_DATA_STATS_FILE.value,
target_modules=_TARGET_MODULES.value,
)
# Frees the model from GPU.
utils.force_gc()
output_dir.upload_to_gcs(skip_if_exists=True)
if __name__ == '__main__':
app.run(main)
@@ -0,0 +1,132 @@
"""Script to merge PEFT adapter with base model."""
from typing import Any, Dict, Sequence
from absl import app
from absl import flags
from util import dataset_validation_util
from vertex_vision_model_garden_peft.train.vmg import utils
from util import constants
from util import fileutils
_PRETRAINED_MODEL_ID = flags.DEFINE_string(
'pretrained_model_id',
None,
'The pretrained model id. Supported models can be causal language modeling'
' models from https://github.com/huggingface/peft/tree/main. Note, there'
' might be different paddings for different models. This tool assumes the'
' pretrained_model_id contains model name, and then choose proper padding'
' methods. e.g. it must contain `llama` for `Llama2 models`.',
required=True,
)
_MERGE_BASE_AND_LORA_OUTPUT_DIR = flags.DEFINE_string(
'merge_base_and_lora_output_dir',
None,
'The directory to store the merged model with the base and lora adapter.',
)
_MERGE_MODEL_PRECISION_MODE = flags.DEFINE_enum(
'merge_model_precision_mode',
constants.PRECISION_MODE_16,
[
constants.PRECISION_MODE_4,
constants.PRECISION_MODE_8,
constants.PRECISION_MODE_FP8,
constants.PRECISION_MODE_16,
constants.PRECISION_MODE_16B,
constants.PRECISION_MODE_32,
],
'Merging model precision mode.',
)
_FINETUNED_LORA_MODEL_DIR = flags.DEFINE_string(
'finetuned_lora_model_dir',
None,
'The directory storing finetuned LoRA model weights.',
)
_RESTRICT_MODEL_UPLOAD_DOCKER_URI = flags.DEFINE_string(
'restrict_model_upload_docker_uri',
'',
'If set, mark output model as only uploadable to Model Registry with the'
' specified Docker URI.',
)
_EXECUTOR_INPUT = flags.DEFINE_string(
'executor_input',
'',
'For internal use. Kubeflow pipeline context when running trainer as part'
' of an internal pipeline.',
)
_HUGGINGFACE_ACCESS_TOKEN = flags.DEFINE_string(
'huggingface_access_token',
None,
'The access token for loading huggingface gated models.',
)
@flags.multi_flags_validator(
[
_PRETRAINED_MODEL_ID.name,
_FINETUNED_LORA_MODEL_DIR.name,
_MERGE_BASE_AND_LORA_OUTPUT_DIR.name,
],
)
def check_merge_lora_model_flags(flags_dict: Dict[str, Any]) -> bool:
"""Check if required flags are set on merge model LoRA task.
Args:
flags_dict: Dictionary containing task and flags to check.
Returns:
If required flags are not None.
"""
return all(map(lambda x: x is not None, flags_dict.values()))
def main(unused_argv: Sequence[str]) -> None:
pretrained_model_id = fileutils.force_gcs_path(_PRETRAINED_MODEL_ID.value)
if dataset_validation_util.is_gcs_path(pretrained_model_id):
pretrained_model_id = dataset_validation_util.download_gcs_uri_to_local(
pretrained_model_id
)
finetuned_lora_model_dir = utils.GcsOrLocalDirectory(
_FINETUNED_LORA_MODEL_DIR.value
)
merge_base_and_lora_output_dir = utils.GcsOrLocalDirectory(
_MERGE_BASE_AND_LORA_OUTPUT_DIR.value
)
utils.merge_causal_language_model_with_lora(
pretrained_model_id=pretrained_model_id,
precision_mode=_MERGE_MODEL_PRECISION_MODE.value,
finetuned_lora_model_dir=finetuned_lora_model_dir.local_dir,
merged_model_output_dir=merge_base_and_lora_output_dir.local_dir,
access_token=_HUGGINGFACE_ACCESS_TOKEN.value,
)
if _RESTRICT_MODEL_UPLOAD_DOCKER_URI.value:
utils.write_first_party_model_metadata(
merge_base_and_lora_output_dir.local_dir,
_RESTRICT_MODEL_UPLOAD_DOCKER_URI.value,
)
if _EXECUTOR_INPUT.value:
utils.write_kfp_outputs(
_EXECUTOR_INPUT.value,
{
'saved_model': _MERGE_BASE_AND_LORA_OUTPUT_DIR.value,
},
)
merge_base_and_lora_output_dir.upload_to_gcs(skip_if_exists=True)
if __name__ == '__main__':
app.run(main)
@@ -0,0 +1,349 @@
"""Quantizes the model."""
import json
import os
from typing import Any, Dict, List, Sequence, Union
from absl import app
from absl import flags
from absl import logging
from auto_gptq import AutoGPTQForCausalLM
from auto_gptq import BaseQuantizeConfig
from awq import AutoAWQForCausalLM
from optimum.gptq.data import get_dataset
from transformers import AutoTokenizer
from util import dataset_validation_util
from vertex_vision_model_garden_peft.train.vmg import utils
from util import constants
_PRETRAINED_MODEL_ID = flags.DEFINE_string(
'pretrained_model_id',
None,
'The pretrained model id. Supported models can be causal language modeling'
' models from https://github.com/huggingface/peft/tree/main. Note, there'
' might be different paddings for different models. This tool assumes the'
' pretrained_model_id contains model name, and then choose proper padding'
' methods. e.g. it must contain `llama` for `Llama2 models`.',
)
_QUANTIZATION_METHOD = flags.DEFINE_enum(
'quantization_method',
None,
[constants.GPTQ, constants.AWQ],
'The quantization method. Choose from ["gtpq", "awq"].',
)
_QUANTIZATION_PRECISION_MODE = flags.DEFINE_enum(
'quantization_precision_mode',
constants.PRECISION_MODE_4,
[
constants.PRECISION_MODE_8,
constants.PRECISION_MODE_4,
constants.PRECISION_MODE_3,
constants.PRECISION_MODE_2,
],
'Quantization precision mode.',
)
_QUANTIZATION_DATASET_NAME = flags.DEFINE_string(
'quantization_dataset_name',
None,
'The dataset used for quantization. You can provide your own dataset in a'
' list of string or just use the original datasets used in GPTQ paper'
' ["wikitext2","c4","c4-new","ptb","ptb-new"] for GPTQ quantization. Using'
" a dataset more appropriate to the model's training can improve"
' quantisation accuracy. Note that the GPTQ dataset is not the same as the'
' dataset used to train the model.',
)
_TEXT_COLUMN_IN_QUANTIZATION_DATASET = flags.DEFINE_string(
'text_column_in_quantization_dataset',
constants.DEFAULT_TEXT_COLUMN_IN_QUANTIZATION_DATASET,
'The text column in quantization dataset.',
)
_QUANTIZATION_OUTPUT_DIR = flags.DEFINE_string(
'quantization_output_dir',
None,
'The directory to store the quantized model.',
)
_QUANTIZATION_DEVICE_MAP = flags.DEFINE_string(
'device_map', None, 'The device map.'
)
_QUANTIZATION_MAX_MEMORY = flags.DEFINE_string(
'max_memory', None, 'The maximum memory.'
)
_GROUP_SIZE = flags.DEFINE_integer(
'group_size',
None,
'The group size to use for quantization. Recommended value is 128 and -1'
' uses per-column quantization. Higher numbers use less VRAM, but have'
' lower quantisation accuracy. "None" is the lowest possible value.',
)
_DESC_ACT = flags.DEFINE_boolean(
'desc_act',
False,
'Whether to quantize columns in order of decreasing activation size.'
' Setting it to False can significantly speed up inference but the'
' perplexity may become slightly worse. Also known as act-order.',
)
_DAMP_PERCENT = flags.DEFINE_float(
'damp_percent',
0.1,
'The percent of the average Hessian diagonal to use for dampening.',
)
_CACHE_EXAMPLES_ON_GPU = flags.DEFINE_boolean(
'cache_examples_on_gpu',
True,
'Whether to cache the examples on GPU. Disabling will reduce VRAM usage,'
' but increase quantization time.',
)
_AWQ_VERSION = flags.DEFINE_enum(
'awq_version',
constants.GEMM,
[constants.GEMM, constants.GEMV],
'The version of the AWQ to use. It determines how matrix multiplication'
' runs under the hood. GEMV is 20% faster than GEMM, only at batch size 1'
' (not good for large contexts). GEMM is much faster than FP16 at batch'
' sizes below 8 (good with large contexts).',
)
@flags.multi_flags_validator(
[
_PRETRAINED_MODEL_ID.name,
_QUANTIZATION_METHOD.name,
_QUANTIZATION_PRECISION_MODE.name,
_QUANTIZATION_DATASET_NAME.name,
_QUANTIZATION_OUTPUT_DIR.name,
],
)
def check_quantization_flags(flags_dict: Dict[str, Any]) -> bool:
"""Check if required flags are set on quantization task.
Args:
flags_dict: Dictionary containing task and flags to check.
Returns:
If required flags are not None.
"""
required_flags = [
_QUANTIZATION_METHOD.name,
_PRETRAINED_MODEL_ID.name,
_QUANTIZATION_PRECISION_MODE.name,
_QUANTIZATION_DATASET_NAME.name,
_QUANTIZATION_OUTPUT_DIR.name,
]
return all(map(lambda x: flags_dict[x] is not None, required_flags))
def quantize_model(
quantization_method: str,
pretrained_model_id: str,
quantization_output_dir: str,
quantization_precision_mode: str = None,
quantization_dataset_name: Union[List[str]] = None,
text_column_in_quantization_dataset: str = constants.DEFAULT_TEXT_COLUMN_IN_QUANTIZATION_DATASET,
group_size: int = None,
desc_act: bool = True,
damp_percent: float = 0.1,
awq_version: str = 'GEMM',
device_map: str = None,
max_memory: Dict[Any, str] = None,
cache_examples_on_gpu: bool = True,
) -> None:
"""Quantizes the model using `quantization_method`."""
if quantization_method == constants.GPTQ:
gptq_quantize_model(
pretrained_model_id=pretrained_model_id,
gptq_output_dir=quantization_output_dir,
gptq_precision_mode=quantization_precision_mode,
gptq_dataset_name=quantization_dataset_name,
group_size=group_size,
desc_act=desc_act,
damp_percent=damp_percent,
cache_examples_on_gpu=cache_examples_on_gpu,
)
elif quantization_method == constants.AWQ:
awq_quantize_model(
pretrained_model_id=pretrained_model_id,
quantization_output_dir=quantization_output_dir,
quantization_precision_mode=quantization_precision_mode,
quantization_dataset_name=quantization_dataset_name,
text_column_in_quantization_dataset=text_column_in_quantization_dataset,
group_size=group_size,
awq_version=awq_version,
device_map=device_map,
max_memory=max_memory,
)
def awq_quantize_model(
pretrained_model_id: str,
quantization_output_dir: str,
quantization_precision_mode: str = None,
quantization_dataset_name: Union[List[str]] = None,
text_column_in_quantization_dataset: str = constants.DEFAULT_TEXT_COLUMN_IN_QUANTIZATION_DATASET,
group_size: int = None,
awq_version: str = 'GEMM',
device_map: str = None,
max_memory: Dict[Any, str] = None,
) -> None:
"""Quantizes the model using AWQ."""
if quantization_precision_mode != constants.PRECISION_MODE_4:
raise ValueError(
f'Invalid precision mode: {quantization_precision_mode} for AWQ. 4bit'
' quantization must be used.'
)
else:
bits = 4
if not group_size:
group_size = 128
if not device_map:
device_map = 'cpu'
if dataset_validation_util.is_gcs_path(quantization_dataset_name):
logging.info('Using custom dataset: %s', quantization_dataset_name)
with open(
dataset_validation_util.force_gcs_fuse_path(quantization_dataset_name),
'r',
) as f:
quantization_dataset = [line.rstrip('\n') for line in f]
else:
quantization_dataset = quantization_dataset_name
quant_config = {
'zero_point': True,
'q_group_size': group_size,
'w_bit': bits,
'version': awq_version,
}
logging.info('Quantization config: %s', quant_config)
model = AutoAWQForCausalLM.from_pretrained(
pretrained_model_id,
trust_remote_code=True,
device_map=device_map,
max_memory=max_memory,
low_cpu_mem_usage=True,
)
tokenizer = AutoTokenizer.from_pretrained(
pretrained_model_id, trust_remote_code=True
)
model.quantize(
tokenizer,
quant_config=quant_config,
calib_data=quantization_dataset,
text_column=text_column_in_quantization_dataset,
)
model.save_quantized(quantization_output_dir)
tokenizer.save_pretrained(quantization_output_dir)
def gptq_quantize_model(
pretrained_model_id: str,
gptq_output_dir: str,
gptq_precision_mode: str = None,
gptq_dataset_name: Union[List[str]] = None,
group_size: int = -1,
desc_act: bool = False,
damp_percent: float = 0.1,
cache_examples_on_gpu: bool = True,
) -> None:
"""Quantizes the model using GPTQ."""
logging.info(
'PYTORCH_CUDA_ALLOC_CONF: %s',
os.environ.get('PYTORCH_CUDA_ALLOC_CONF', ''),
)
if dataset_validation_util.is_gcs_path(gptq_dataset_name):
logging.info('Using custom dataset: %s', gptq_dataset_name)
with open(
dataset_validation_util.force_gcs_fuse_path(gptq_dataset_name), 'r'
) as f:
gptq_dataset = [line.rstrip('\n') for line in f]
else:
gptq_dataset = gptq_dataset_name
if gptq_precision_mode == constants.PRECISION_MODE_8:
bits = 8
elif gptq_precision_mode == constants.PRECISION_MODE_4:
bits = 4
elif gptq_precision_mode == constants.PRECISION_MODE_3:
bits = 3
elif gptq_precision_mode == constants.PRECISION_MODE_2:
bits = 2
else:
raise ValueError(f'Invalid precision mode: {gptq_precision_mode} for GPTQ.')
if not group_size:
group_size = -1
tokenizer = AutoTokenizer.from_pretrained(pretrained_model_id)
gptq_dataset = get_dataset(gptq_dataset, tokenizer)
quantization_config = BaseQuantizeConfig(
bits=bits,
group_size=group_size,
damp_percent=damp_percent,
desc_act=desc_act,
)
logging.info('Quantization config: %s', quantization_config.to_dict())
model = AutoGPTQForCausalLM.from_pretrained(
pretrained_model_id,
quantization_config,
low_cpu_mem_usage=True,
torch_dtype='auto',
trust_remote_code=True,
)
model.quantize(
examples=gptq_dataset,
cache_examples_on_gpu=cache_examples_on_gpu,
)
if utils.should_add_pad_token(pretrained_model_id):
tokenizer.add_special_tokens({'pad_token': '[PAD]'})
model.resize_token_embeddings(len(tokenizer))
model.save_pretrained(gptq_output_dir)
tokenizer.save_pretrained(gptq_output_dir)
def main(unused_argv: Sequence[str]) -> None:
pretrained_model_id = _PRETRAINED_MODEL_ID.value
if dataset_validation_util.is_gcs_path(pretrained_model_id):
pretrained_model_id = dataset_validation_util.download_gcs_uri_to_local(
pretrained_model_id
)
pretrained_model_id = dataset_validation_util.force_gcs_fuse_path(
pretrained_model_id
)
if _QUANTIZATION_MAX_MEMORY.value:
max_memory = json.loads(_QUANTIZATION_MAX_MEMORY.value)
else:
max_memory = None
quantize_model(
quantization_method=_QUANTIZATION_METHOD.value,
pretrained_model_id=pretrained_model_id,
quantization_output_dir=_QUANTIZATION_OUTPUT_DIR.value,
quantization_precision_mode=_QUANTIZATION_PRECISION_MODE.value,
quantization_dataset_name=_QUANTIZATION_DATASET_NAME.value,
text_column_in_quantization_dataset=_TEXT_COLUMN_IN_QUANTIZATION_DATASET.value,
group_size=_GROUP_SIZE.value,
desc_act=_DESC_ACT.value,
damp_percent=_DAMP_PERCENT.value,
awq_version=_AWQ_VERSION.value,
device_map=_QUANTIZATION_DEVICE_MAP.value,
max_memory=max_memory,
cache_examples_on_gpu=_CACHE_EXAMPLES_ON_GPU.value,
)
if __name__ == '__main__':
app.run(main)
@@ -0,0 +1,21 @@
#!/bin/bash
# Run copybara first:
# cloud/ml/applications/vision/model_garden/copybara/run_copybara_local.sh
# Run docker build:
# cloud/ml/applications/vision/model_garden/model_oss/peft/train/vmg/scripts/build_train_docker.sh
set -x
set -e
COPYBARA_DIR="/tmp/train_docker/"
pushd "${COPYBARA_DIR}"
PROJECT="cloud-nas-260507"
IMAGE_TAG="gcr.io/${PROJECT}/pytorch-peft-train:${USER}-test"
docker build -f model_oss/peft/train/vmg/dockerfile/train.Dockerfile . -t "${IMAGE_TAG}"
docker push "${IMAGE_TAG}"
popd
@@ -1,7 +1,9 @@
"""Sequence classification with LoRA models."""
# pylint: disable=g-importing-member
from typing import Sequence
from absl import app
from absl import flags
from datasets import load_dataset
import evaluate
from peft import get_peft_model
@@ -14,6 +16,71 @@ from transformers import AutoModelForSequenceClassification
from transformers import AutoTokenizer
from transformers import get_linear_schedule_with_warmup
from util import dataset_validation_util
_PRETRAINED_MODEL_ID = flags.DEFINE_string(
"pretrained_model_id",
None,
"The pretrained model id. Supported models can be causal language modeling"
" models from https://github.com/huggingface/peft/tree/main. Note, there"
" might be different paddings for different models. This tool assumes the"
" pretrained_model_id contains model name, and then choose proper padding"
" methods. e.g. it must contain `llama` for `Llama2 models`.",
)
_OUTPUT_DIR = flags.DEFINE_string(
"output_dir",
None,
"The output directory.",
)
_DATASET_NAME = flags.DEFINE_string(
"dataset_name",
None,
"The dataset name in huggingface.",
)
_LORA_RANK = flags.DEFINE_integer(
"lora_rank",
16,
"The rank of the update matrices, expressed in int. Lower rank results in"
" smaller update matrices with fewer trainable parameters, referring to"
" https://huggingface.co/docs/peft/conceptual_guides/lora.",
)
_LORA_ALPHA = flags.DEFINE_integer(
"lora_alpha",
32,
"LoRA scaling factor, referring to"
" https://huggingface.co/docs/peft/conceptual_guides/lora.",
)
_LORA_DROPOUT = flags.DEFINE_float(
"lora_dropout",
0.05,
"dropout probability of the LoRA layers, referring to"
" https://huggingface.co/docs/peft/task_guides/token-classification-lora.",
)
_NUM_EPOCHS = flags.DEFINE_integer(
"num_epochs",
None,
"The number of training epochs.",
)
_BATCH_SIZE = flags.DEFINE_integer(
"batch_size",
32,
"The batch size.",
)
_LEARNING_RATE = flags.DEFINE_float(
"learning_rate",
2e-4,
"The learning rate after the potential warmup period.",
)
def finetune_sequence_classification(
pretrained_model_id: str,
@@ -131,3 +198,32 @@ def finetune_sequence_classification(
print(f"epoch {epoch}:", eval_metric)
model.save_pretrained(output_dir)
def main(unused_argv: Sequence[str]) -> None:
if dataset_validation_util.is_gcs_path(_PRETRAINED_MODEL_ID.value):
pretrained_model_id = dataset_validation_util.download_gcs_uri_to_local(
_PRETRAINED_MODEL_ID.value
)
else:
pretrained_model_id = _PRETRAINED_MODEL_ID.value
pretrained_model_path = dataset_validation_util.force_gcs_fuse_path(
pretrained_model_id
)
output_dir = dataset_validation_util.force_gcs_fuse_path(_OUTPUT_DIR.value)
finetune_sequence_classification(
pretrained_model_id=pretrained_model_path,
dataset_name=_DATASET_NAME.value,
output_dir=output_dir,
lora_rank=_LORA_RANK.value,
lora_alpha=_LORA_ALPHA.value,
lora_dropout=_LORA_DROPOUT.value,
num_epochs=int(_NUM_EPOCHS.value),
batch_size=_BATCH_SIZE.value,
learning_rate=_LEARNING_RATE.value,
)
if __name__ == "__main__":
app.run(main)
@@ -0,0 +1,98 @@
# Vertex Model Garden Training Dataset Template
## Overview
Vertex Model Garden training provides templates for streamlined preprocessing of
datasets. Although datasets often have intricate structures, the supported LLM
models accept only flat strings. A template facilitates parsing a dataset and
preprocessing it to be compatible with the model.
When fine-tuning a pretrained model, it is advisable to maintain the same format
as the original training data. A template helps replicate the format, ensuring
consistency and potentially enhancing the fine-tuning process.
Both multi-turn messages and single instruction-response pairs are supported.
Multi-turn messages are accommodated using a more general `chat_template` field,
whereas simple instruction-response pair datasets are supported through the
`prompt_input` field.
A template is a JSON file consisting of string key-value pairs. Refer to the
following for the definitions of the supported fields.
## Template field documentation
**description**: An explanation of the template.
**source**: Information about the origin of the template.
**chat_template**: A
[jinja template](https://jinja.palletsprojects.com/en/3.1.x/templates/) that can
be used to parse a chat dataset. This is the same format as
[HF chat templates](https://huggingface.co/docs/transformers/main/en/chat_templating).
To create a chat_template, use the `messages` variable to be filled with the
sample. The flag `--instruct_column_in_dataset` identifies which column will be
passed to the `messages` variable in the chat_template. This field is mutually
exclusive with `prompt_input` and `prompt_no_input`.
**prompt_input**: A string template that is used when value for the input column
exists in the sample. It should be able to be formatted with the
[str.format](https://docs.python.org/3/library/stdtypes.html#str.format) method.
The input column is specified with the flag `--instruct_column_in_dataset`. Used
for instruction dataset. This field is mutually exclusive with `chat_template`.
**prompt_no_input**: A string template that is used when value for the input
column does not exist in the sample. It should be able to be formatted with the
[str.format](https://docs.python.org/3/library/stdtypes.html#str.format) method.
The input column is specified with the flag `--instruct_column_in_dataset`. Used
for instruction dataset. This field is mutually exclusive with `chat_template`.
**instruction_separator**: A unique string used to indicate the start of the
instructions. If not specified, every token after response_separator will be
treated as a response, and every token before the first response_separator will
be treated as instruction.
**response_separator**: A unique string used to indicate the start of the
response. This field is required if `--completion_only` flag is set to `True`.
## Example templates
- See the list of all supported templates [here](https://github.com/GoogleCloudPlatform/vertex-ai-samples/tree/main/community-content/vertex_model_garden/model_oss/peft/train/vmg/templates).
- For an example with `chat_template` see the JSON template below.
```
{
"description": "Chat template used by Llama 3.",
"source": "https://huggingface.co/meta-llama/Meta-Llama-3-70B-Instruct/blob/a5a71a7527eac1d651bb145436c72026887fb68e/tokenizer_config.json#L2053",
"chat_template": "{% set loop_messages = messages %}{% for message in loop_messages %}{% set content = '<|start_header_id|>' + message['role'] + '<|end_header_id|>\n\n'+ message['content'] | trim + '<|eot_id|>' %}{% if loop.index0 == 0 %}{% set content = bos_token + content %}{% endif %}{{ content }}{% endfor %}{% if add_generation_prompt %}{{ '<|start_header_id|>assistant<|end_header_id|>\n\n' }}{% endif %}",
"instruction_separator": "<|start_header_id|>user<|end_header_id|>\n\n",
"response_separator": "<|start_header_id|>assistant<|end_header_id|>\n\n"
}
```
- For an example with `prompt_input` see the JSON template below. In this case
the flag `--instruct_column_in_dataset=text` should be set, and there must
be a column named `text` in the dataset.
```
{
"description": "Template for openassistant-guanaco dataset.",
"source": "https://huggingface.co/datasets/timdettmers/openassistant-guanaco",
"prompt_input": "{text}",
"instruction_separator": "### Human:",
"response_separator": "### Assistant:"
}
```
- For an example with `prompt_no_input` see the JSON template below. In this
case the flag `--instruct_column_in_dataset=input` should be set, and there
must be columns named `input` and `instruction` in the dataset.
```
{
"description": "Template used by Alpaca-LoRA.",
"source": "https://github.com/tloen/alpaca-lora/blob/main/templates/alpaca.json",
"prompt_input": "Below is an instruction that describes a task, paired with an input that provides further context. Write a response that appropriately completes the request.\n\n### Instruction:\n{instruction}\n\n### Input:\n{input}\n\n### Response:\n",
"prompt_no_input": "Below is an instruction that describes a task. Write a response that appropriately completes the request.\n\n### Instruction:\n{instruction}\n\n### Response:\n",
"response_separator": "### Response:"
}
```
@@ -0,0 +1,7 @@
{
"description": "Template used by Alpaca-LoRA.",
"source": "https://github.com/tloen/alpaca-lora/blob/main/templates/alpaca.json",
"prompt_input": "Below is an instruction that describes a task, paired with an input that provides further context. Write a response that appropriately completes the request.\n\n### Instruction:\n{instruction}\n\n### Input:\n{input}\n\n### Response:\n",
"prompt_no_input": "Below is an instruction that describes a task. Write a response that appropriately completes the request.\n\n### Instruction:\n{instruction}\n\n### Response:\n",
"response_separator": "### Response:"
}
@@ -0,0 +1,7 @@
{
"description": "A shorter template to experiment with.",
"source": "https://github.com/tloen/alpaca-lora/blob/main/templates/alpaca_short.json",
"prompt_input": "### Instruction:\n{instruction}\n\n### Input:\n{input}\n\n### Response:\n",
"prompt_no_input": "### Instruction:\n{instruction}\n\n### Response:\n",
"response_separator": "### Response:"
}
@@ -0,0 +1,7 @@
{
"description": "Chat template used by Gemma. 'assistant' role is replaced by 'model'",
"source": "https://huggingface.co/google/gemma-1.1-2b-it/blob/bf4924f313df5166dee1467161e886e55f2eb4d4/tokenizer_config.json#L1507",
"chat_template": "{{ bos_token }}{% if messages[0]['role'] == 'system' %}{{ raise_exception('System role not supported') }}{% endif %}{% for message in messages %}{% if (message['role'] == 'user') != (loop.index0 % 2 == 0) %}{{ raise_exception('Conversation roles must alternate user/assistant/user/assistant/...') }}{% endif %}{% if (message['role'] == 'assistant') %}{% set role = 'model' %}{% else %}{% set role = message['role'] %}{% endif %}{{ '<start_of_turn>' + role + '\n' + message['content'] | trim + '<end_of_turn>\n' }}{% endfor %}{% if add_generation_prompt %}{{'<start_of_turn>model\n'}}{% endif %}",
"instruction_separator": "<start_of_turn>user\n",
"response_separator": "<start_of_turn>model\n"
}
@@ -0,0 +1,7 @@
{
"description": "Template used by Llama 3, accepting text-bison format.",
"source": "https://cloud.google.com/vertex-ai/generative-ai/docs/models/tune-text-models-supervised#dataset-format",
"prompt_input": "\n\n<|start_header_id|>user<|end_header_id|>\n\n{input_text}<|eot_id|>\n\n<|start_header_id|>assistant<|end_header_id|>\n\n{output_text}<|eot_id|>",
"instruction_separator": "<|start_header_id|>user<|end_header_id|>\n\n",
"response_separator": "<|start_header_id|>assistant<|end_header_id|>\n\n"
}
@@ -0,0 +1,7 @@
{
"description": "Chat template used by Llama 3.",
"source": "https://huggingface.co/meta-llama/Meta-Llama-3-70B-Instruct/blob/a5a71a7527eac1d651bb145436c72026887fb68e/tokenizer_config.json#L2053",
"chat_template": "{% set loop_messages = messages %}{% for message in loop_messages %}{% set content = '<|start_header_id|>' + message['role'] + '<|end_header_id|>\n\n'+ message['content'] | trim + '<|eot_id|>' %}{% if loop.index0 == 0 %}{% set content = bos_token + content %}{% endif %}{{ content }}{% endfor %}{% if add_generation_prompt %}{{ '<|start_header_id|>assistant<|end_header_id|>\n\n' }}{% endif %}",
"instruction_separator": "<|start_header_id|>user<|end_header_id|>\n\n",
"response_separator": "<|start_header_id|>assistant<|end_header_id|>\n\n"
}
@@ -0,0 +1,7 @@
{
"description": "Chat template used by Mistral.",
"source": "https://github.com/OpenAccess-AI-Collective/axolotl/blob/main/src/axolotl/utils/chat_templates.py",
"chat_template": "{{ bos_token }}{% for message in messages %}{% if (message['role'] == 'user') != (loop.index0 % 2 == 0) %}{{ raise_exception('Conversation roles must alternate user/assistant/user/assistant/...') }}{% endif %}{% if message['role'] == 'user' %}{{ '[INST] ' + message['content'] + ' [/INST]' }}{% elif message['role'] == 'assistant' %}{{ message['content'] + eos_token}}{% else %}{{ raise_exception('Only user and assistant roles are supported!') }}{% endif %}{% endfor %}",
"instruction_separator": "[INST]",
"response_separator": "[/INST]"
}
@@ -0,0 +1,7 @@
{
"description": "Template used by openai completion.",
"source": "https://platform.openai.com/docs/api-reference/fine-tuning/completions-input",
"prompt_input": "\n\n<|start_header_id|>user<|end_header_id|>\n\n{prompt}<|eot_id|>\n\n<|start_header_id|>assistant<|end_header_id|>\n\n{completion}<|eot_id|>",
"instruction_separator": "<|start_header_id|>user<|end_header_id|>\n\n",
"response_separator": "<|start_header_id|>assistant<|end_header_id|>\n\n"
}
@@ -0,0 +1,7 @@
{
"description": "Template for openassistant-guanaco dataset.",
"source": "https://huggingface.co/datasets/timdettmers/openassistant-guanaco",
"prompt_input": "{text}",
"instruction_separator": "### Human:",
"response_separator": "### Assistant:"
}
@@ -0,0 +1,7 @@
{
"description": "Template used for chat based models.",
"source": "https://huggingface.co/meta-llama/Meta-Llama-3-70B-Instruct/blob/a5a71a7527eac1d651bb145436c72026887fb68e/tokenizer_config.json#L2053",
"chat_template": "{% set loop_messages = messages %}{% for message in loop_messages %}{% set content = '<|start_header_id|>' + message['role'] + '<|end_header_id|>\n\n'+ message['content'] | trim + '<|eot_id|>' %}{% if loop.index0 == 0 %}{% set content = bos_token + content %}{% endif %}{{ content }}{% endfor %}{% if add_generation_prompt %}{{ '<|start_header_id|>model<|end_header_id|>\n\n' }}{% endif %}",
"instruction_separator": "<|start_header_id|>user<|end_header_id|>\n\n",
"response_separator": "<|start_header_id|>model<|end_header_id|>\n\n"
}
@@ -0,0 +1,242 @@
"""Entrypoint for peft train docker.
Dispatches to different scripts based on `task` type.
For task type in `_TASK_TO_SCRIPT`, if `--config_file` is specified, the script
will dispatch the call to `accelerate`, which is friendly for multi-GPU
environment. Otherwise, `python3` is used.
"""
import argparse
import json
import os
import subprocess
from typing import List, Optional, Sequence
from absl import app
from absl import flags
from absl import logging
from util import dataset_validation_util
from vertex_vision_model_garden_peft.train.vmg import utils
from util import constants
from util import hypertune_utils
_TEXT_TO_IMAGE_TASKS_SCRIPTS = {
constants.TEXT_TO_IMAGE: 'text_to_image/train_text_to_image.py',
constants.TEXT_TO_IMAGE_LORA: 'text_to_image/train_text_to_image_lora.py',
constants.TEXT_TO_IMAGE_DREAMBOOTH: 'dreambooth/train_dreambooth.py',
constants.TEXT_TO_IMAGE_DREAMBOOTH_LORA: (
'dreambooth/train_dreambooth_lora.py'
),
constants.TEXT_TO_IMAGE_DREAMBOOTH_LORA_SDXL: (
'dreambooth/train_dreambooth_lora_sdxl.py'
),
}
_TASK_TO_SCRIPT = {
constants.INSTRUCT_LORA: (
'vertex_vision_model_garden_peft/train/vmg/instruct_lora.py'
),
constants.MERGE_CAUSAL_LANGUAGE_MODEL_LORA: 'vertex_vision_model_garden_peft/train/vmg/merge_causal_language_model_lora.py',
constants.QUANTIZE_MODEL: (
'vertex_vision_model_garden_peft/train/vmg/quantize_model.py'
),
constants.SEQUENCE_CLASSIFICATION_LORA: 'vertex_vision_model_garden_peft/train/vmg/sequence_classification_lora.py',
constants.VALIDATE_DATASET_WITH_TEMPLATE: 'vertex_vision_model_garden_peft/train/vmg/validate_dataset_with_template.py',
}
def launch_script_cmd(
script: str,
config_file: Optional[str],
accelerate_args: argparse.Namespace = argparse.Namespace(),
) -> List[str]:
"""Returns the command to launch the script."""
if config_file:
cmd = [
'accelerate',
'launch',
'--config_file={}'.format(config_file),
]
else:
cmd = ['python3']
_append_args_to_command_in_place(accelerate_args, cmd)
cmd.append(script)
return cmd
def _get_accelerate_args() -> argparse.Namespace:
"""Returns the accelerate args."""
# For the format of the cluster spec, see
# https://cloud.google.com/vertex-ai/docs/training/distributed-training#cluster-spec-format # pylint: disable=line-too-long
cluster_spec = os.getenv('CLUSTER_SPEC', default=None)
if not cluster_spec:
return argparse.Namespace()
logging.info('CLUSTER_SPEC: %s', cluster_spec)
cluster_data = json.loads(cluster_spec)
if (
'workerpool1' not in cluster_data['cluster']
or not cluster_data['cluster']['workerpool1']
):
return argparse.Namespace()
# Get primary node info
primary_node = cluster_data['cluster']['workerpool0'][0]
logging.info('primary node: %s', primary_node)
primary_node_addr, primary_node_port = primary_node.split(':')
logging.info('primary node address: %s', primary_node_addr)
logging.info('primary node port: %s', primary_node_port)
# Determine node rank of this machine
workerpool = cluster_data['task']['type']
if workerpool == 'workerpool0':
node_rank = 0
elif workerpool == 'workerpool1':
# Add 1 for the primary node, since `index` is the index of workerpool1.
node_rank = cluster_data['task']['index'] + 1
else:
raise ValueError(
'Only workerpool0 and workerpool1 are supported. Unknown workerpool:'
f' {workerpool}'
)
logging.info('node rank: %s', node_rank)
# Calculate total nodes
num_worker_nodes = len(cluster_data['cluster']['workerpool1'])
num_nodes = num_worker_nodes + 1 # Add 1 for the primary node
logging.info('num nodes: %s', num_nodes)
accelerate_args = argparse.Namespace()
accelerate_args.machine_rank = node_rank
accelerate_args.num_machines = num_nodes
accelerate_args.main_process_ip = primary_node_addr
accelerate_args.main_process_port = primary_node_port
accelerate_args.max_restarts = 0
accelerate_args.monitor_interval = 120
return accelerate_args
def _append_args_to_command_in_place(
args: argparse.Namespace, command: List[str]
):
for key, value in vars(args).items():
# If not specified, skip.
if value is not None:
command.append(f'--{key}={value}')
def _get_train_cmd_and_maybe_merge_cmd(
task: str, config_file: str, unknown: Sequence[str]
) -> Sequence[Sequence[str]]:
"""Returns the training command and maybe the merge command if applicable."""
# Only populated when multi-node is used.
accelerate_args = _get_accelerate_args()
training_cmd = launch_script_cmd(
_TASK_TO_SCRIPT[task],
config_file,
accelerate_args=accelerate_args,
)
# Training only flag.
train_parser = argparse.ArgumentParser()
train_parser.add_argument('--output_dir', required=True)
training_args, unknown = train_parser.parse_known_args(unknown)
# Checks for `hypertune_utils._ENVIRONMENT_VARIABLE_FOR_TRIAL_ID` env var and
# appends the trial id if it exists.
training_args.output_dir = hypertune_utils.maybe_append_trial_id(
dataset_validation_util.force_gcs_fuse_path(training_args.output_dir)
)
# Merge only flags.
merge_parser = argparse.ArgumentParser()
merge_parser.add_argument('--merge_model_precision_mode')
merge_parser.add_argument('--executor_input')
merge_parser.add_argument('--restrict_model_upload_docker_uri')
merge_parser.add_argument('--merge_base_and_lora_output_dir')
merge_args, unknown = merge_parser.parse_known_args(unknown)
# Common flags shared by merging and training.
common_parser = argparse.ArgumentParser()
common_parser.add_argument('--pretrained_model_id', required=True)
common_parser.add_argument('--huggingface_access_token')
common_args, remaining = common_parser.parse_known_args(unknown)
# Add flags for training.
_append_args_to_command_in_place(training_args, training_cmd)
_append_args_to_command_in_place(common_args, training_cmd)
training_cmd.extend(remaining) # Remaining args are passed to training cmd.
commands = [training_cmd]
# Only the main node runs merging.
if (
merge_args.merge_base_and_lora_output_dir
and getattr(accelerate_args, 'machine_rank', 0) == 0
):
lora_dir = utils.get_final_checkpoint_path(training_args.output_dir)
merge_cmd = [
'WORLD_SIZE=1', # To ignore other nodes in multi-node setting.
'python3',
_TASK_TO_SCRIPT[constants.MERGE_CAUSAL_LANGUAGE_MODEL_LORA],
f'--finetuned_lora_model_dir={lora_dir}',
]
_append_args_to_command_in_place(merge_args, merge_cmd)
_append_args_to_command_in_place(common_args, merge_cmd)
# Run in a conda environment.
conda_run_cmd = [
'/bin/bash',
'-c',
f'conda run -n merge {" ".join(merge_cmd)}',
]
commands.append(conda_run_cmd)
return commands
def main(unused_argv: Sequence[str]) -> None:
parser = argparse.ArgumentParser()
parser.add_argument('--config_file')
parser.add_argument('--task')
args, unknown = parser.parse_known_args()
task = args.task
if task in _TEXT_TO_IMAGE_TASKS_SCRIPTS:
# Setup accelerate config before running trainer.
config_gen_cmd = [
'python',
'-c',
(
'from accelerate.utils import write_basic_config;'
' write_basic_config(mixed_precision="fp16")'
),
]
task_cmd = [
'accelerate',
'launch',
_TEXT_TO_IMAGE_TASKS_SCRIPTS[task],
] + list(map(dataset_validation_util.force_gcs_fuse_path, unknown))
commands = [config_gen_cmd, task_cmd]
elif task in [constants.INSTRUCT_LORA]:
commands = _get_train_cmd_and_maybe_merge_cmd(
task=task, config_file=args.config_file, unknown=unknown
)
else:
assert task in _TASK_TO_SCRIPT
cmd = launch_script_cmd(_TASK_TO_SCRIPT[task], args.config_file)
cmd.extend(unknown)
commands = [cmd]
for cmd in commands:
logging.info('launching task=%s with cmd: \n%s', task, ' \\\n'.join(cmd))
subprocess.run(cmd, check=True)
if __name__ == '__main__':
app.run(main, flags_parser=lambda _args: flags.FLAGS(_args, known_only=True))
@@ -0,0 +1,742 @@
"""Common libraries for PEFT."""
import dataclasses
import datetime
import gc
import multiprocessing as mp
import os
import subprocess
from typing import Any, Dict, Optional, Sequence
from absl import logging
import accelerate
from accelerate import DistributedType
from accelerate import PartialState
from google.protobuf import json_format
from kfp.pipeline_spec import pipeline_spec_pb2
import numpy as np
import peft
from peft import PeftModel
from peft import prepare_model_for_kbit_training
import pynvml
import torch
import transformers
from transformers import AutoModelForCausalLM
from transformers import AutoTokenizer
from transformers import BitsAndBytesConfig
from transformers import FbgemmFp8Config
from transformers.integrations import is_deepspeed_zero3_enabled
import trl
from util import dataset_validation_util
from util import constants
from util import fileutils
_MODELS_REQUIRING_PAD_TOKEN = ("llama", "falcon", "mistral", "mixtral")
_MODELS_REQUIRING_EOS_TOEKN = ("gemma-2b", "gemma-7b")
_LLAMA_3_1_405B_MODEL_ID = "Meta-Llama-3.1-405B"
_LOCAL_MERGED_MODEL_DIR = "/tmp/merged_model"
class GcsOrLocalDirectory(os.PathLike):
"""A class to represent a directory with upload support if GCS path is given.
This class is used to represent a directory. It can be used for a temporary
local directory and for uploading files to the GCS directory later if the
given path is a GCS directory. If the given path is a local directory, a call
to gcs_dir attribute will raise an error. This class has multi-node and
multi-process support with accelerate.
Attributes:
local_dir: The local directory to store the files.
gcs_dir: The path to the GCS directory.
"""
def __init__(
self,
path: str,
check_empty: bool = False,
upload_from_all_nodes: bool = False,
):
"""Initializes the GcsOrLocalDirectory.
Args:
path: The path to the directory.
check_empty: If True, check if the GCS directory is empty. No-op for local
directory.
upload_from_all_nodes: If True, upload the local directory to GCS from all
nodes.
"""
if len(path) > 1:
path = path.rstrip("/")
self._upload_from_all_nodes = upload_from_all_nodes
if path.startswith(constants.GCS_URI_PREFIX) or path.startswith(
constants.GCSFUSE_URI_PREFIX
):
self._is_gcs_path = True
self._local_dir = _get_local_dir_from_gcs_dir(path)
self._gcs_dir = fileutils.force_gcs_path(path)
os.makedirs(self.local_dir, exist_ok=True)
with PartialState().main_process_first():
if (
check_empty
and PartialState().is_main_process
and not _is_gcs_dir_empty(self._gcs_dir)
):
raise ValueError(f"{self._gcs_dir} needs to be empty.")
else:
self._is_gcs_path = False
self._local_dir = path
self._gcs_dir = path
def __fspath__(self) -> str:
return self.local_dir
@property
def local_dir(self) -> str:
return self._local_dir
@property
def gcs_dir(self) -> str:
"""Returns the GCS directory path.
Returns:
The GCS directory path.
Raises:
ValueError: If the path is not a GCS path.
"""
if not self._is_gcs_path:
raise ValueError(f"{self._gcs_dir} is not a GCS path.")
return self._gcs_dir
def upload_to_gcs(
self,
skip_if_exists: bool = True,
force_upload: bool = False,
):
"""Uploads the local directory to GCS."""
if not self._is_gcs_path:
logging.info(
"Not uploading to GCS since %s is not a GCS path.", self.local_dir
)
return
if not os.listdir(self.local_dir):
logging.info("Not uploading to GCS since %s is empty.", self.local_dir)
return
target = os.path.dirname(self.gcs_dir) + "/"
# Avoid race condition uploading the same file from multiple processes.
with PartialState().main_process_first():
if not PartialState().is_local_main_process:
# Non local main processes don't upload.
pass
elif self._upload_from_all_nodes or PartialState().is_main_process:
logging.info("Uploading %s to %s...", self.local_dir, target)
cmd = [
"gsutil",
"-m",
"cp",
"-r",
]
if skip_if_exists:
cmd.append("-n")
if force_upload:
cmd.append("-f")
cmd.extend([self.local_dir, target])
subprocess.check_output(cmd)
logging.info("%s uploaded.", self.local_dir)
def _get_local_dir_from_gcs_dir(path: str) -> str:
return os.path.join(
constants.LOCAL_OUTPUT_DIR,
dataset_validation_util.force_gcs_fuse_path(path)[1:],
)
def _is_gcs_dir_empty(path: str) -> bool:
"""Checks if a GCS directory is empty.
Args:
path: The GCS directory path.
Returns:
True if the directory is empty.
Raises:
subprocess.CalledProcessError: If the gsutil command failure reason is not
because the dir is empty.
"""
path = path.rstrip("/") + "/"
try:
subprocess.check_output(["gsutil", "ls", path], stderr=subprocess.STDOUT)
except subprocess.CalledProcessError as e:
if (
str(e.output, encoding="utf-8")
== "CommandException: One or more URLs matched no objects.\n"
):
return True
else:
logging.info(str(e.output, encoding="utf-8"))
raise
else:
return False
def load_tokenizer(
pretrained_model_id: str,
padding_side: Optional[str] = None,
access_token: Optional[str] = None,
) -> AutoTokenizer:
"""Loads tokenizer based on `pretrained_model_id`."""
tokenizer_kwargs = {}
if should_add_eos_token(pretrained_model_id):
tokenizer_kwargs["add_eos_token"] = True
if padding_side:
tokenizer_kwargs["padding_side"] = padding_side
with PartialState().local_main_process_first():
tokenizer = AutoTokenizer.from_pretrained(
pretrained_model_id,
trust_remote_code=False,
use_fast=True,
token=access_token,
**tokenizer_kwargs,
)
if should_add_pad_token(pretrained_model_id):
tokenizer.add_special_tokens({"pad_token": "[PAD]"})
return tokenizer
def load_model(
pretrained_model_id: str,
tokenizer: AutoTokenizer,
precision_mode: str = None,
enable_gradient_checkpointing: bool = False,
gradient_checkpointing_kwargs: Optional[Dict[str, Any]] = None,
access_token: Optional[str] = None,
attn_implementation: Optional[str] = None,
train_precision: Optional[str] = None,
device_map: Optional[str] = None,
is_training: bool = True,
) -> AutoModelForCausalLM:
"""Loads models from the local dir if specified or from huggingface."""
# The `distributed_type` we got through `PartialState` is incorrect for FSDP.
# And that's why `Accelerator` is used here.
# See b/357138252 for more details.
accelerator = accelerate.Accelerator()
logging.info("using distributed_type %s", accelerator.distributed_type)
if device_map is None:
if accelerator.distributed_type == DistributedType.MULTI_GPU:
# https://github.com/artidoro/qlora/issues/186#issuecomment-1943045599
# and b/342038175.
device_map = {"": accelerator.process_index}
elif accelerator.distributed_type == DistributedType.DEEPSPEED:
# Deepspeed Zero3 does not allow setting device_map.
# https://github.com/huggingface/transformers/blob/v4.38.2/src/transformers/modeling_utils.py#L2941-L2943
device_map = None
elif accelerator.distributed_type == DistributedType.FSDP:
if precision_mode in [
constants.PRECISION_MODE_4,
constants.PRECISION_MODE_8,
]:
device_map = trl.get_kbit_device_map()
else:
device_map = None
elif (
accelerator.distributed_type == DistributedType.NO
and torch.cuda.device_count() > 1
):
# Setting device map to None to avoid using model parallelism (MP) when
# there are multiple GPUs, which can have very inefficient GPU utilization
# (b/342252819). This setting should trigger torch's nn.DataParallel
# instead, which has better GPU utilization.
device_map = None
else:
device_map = "auto"
logging.info("using device_map %s", device_map)
if train_precision == constants.PRECISION_MODE_32:
train_dtype = torch.float32
elif train_precision == constants.PRECISION_MODE_16:
train_dtype = torch.float16
elif train_precision == constants.PRECISION_MODE_16B:
train_dtype = torch.bfloat16
else:
train_dtype = "auto"
quantization_config = None
# Note: use_cache is False when enable gradient checkpointing.
if precision_mode == constants.PRECISION_MODE_32:
torch_dtype = torch.float32
elif precision_mode == constants.PRECISION_MODE_16:
torch_dtype = torch.float16
elif precision_mode == constants.PRECISION_MODE_16B:
torch_dtype = torch.bfloat16
elif precision_mode == constants.PRECISION_MODE_8:
quantization_config = BitsAndBytesConfig(
load_in_8bit=True, int8_threshold=0
)
torch_dtype = train_dtype
elif precision_mode == constants.PRECISION_MODE_4:
quantization_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_use_double_quant=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=train_dtype,
)
# `bnb_4bit_quant_storage` must be set when using FSDP.
# https://huggingface.co/docs/bitsandbytes/main/en/fsdp_qlora
if accelerator.distributed_type == DistributedType.FSDP:
quantization_config.bnb_4bit_quant_storage = train_dtype
torch_dtype = train_dtype
else:
raise ValueError(f"Invalid precision mode: {precision_mode}")
logging.info("using torch_type=%s", torch_dtype)
model = AutoModelForCausalLM.from_pretrained(
pretrained_model_id,
use_cache=not enable_gradient_checkpointing,
device_map=device_map,
torch_dtype=torch_dtype,
quantization_config=quantization_config,
trust_remote_code=True,
token=access_token,
attn_implementation=attn_implementation,
)
if precision_mode in (constants.PRECISION_MODE_4, constants.PRECISION_MODE_8):
model = prepare_model_for_kbit_training(
model,
use_gradient_checkpointing=enable_gradient_checkpointing,
gradient_checkpointing_kwargs=gradient_checkpointing_kwargs,
)
if enable_gradient_checkpointing:
model.gradient_checkpointing_enable(
gradient_checkpointing_kwargs=gradient_checkpointing_kwargs
)
# Flash attention only supports fp16 or bf16 [1].
# prepare_model_for_kbit_training will force cast some layers to float32 [2]
#
# [1]: https://github.com/Dao-AILab/flash-attention/issues/882
# [2]: https://github.com/huggingface/peft/blob/v0.10.0/src/peft/utils/other.py#L79-L81 # pylint: disable=line-too-long
if attn_implementation == "flash_attention_2" and precision_mode in (
constants.PRECISION_MODE_4,
constants.PRECISION_MODE_8,
):
for _, param in model.named_parameters():
if param.dtype == torch.float32:
param.data = param.data.to(torch_dtype)
if is_training:
# KV cache is useless during training
# https://stackoverflow.com/a/77408076
model.config.use_cache = False
if should_add_pad_token(pretrained_model_id):
model.resize_token_embeddings(len(tokenizer))
if is_training:
# The following is needed since we added a new token that needs to be
# learned.
# https://github.com/QwenLM/Qwen/issues/405#issuecomment-1751680291
model.enable_input_require_grads()
return model
def _merge_causal_language_model_with_lora_internal(
pretrained_model_id: str,
merge_precision_mode: str,
finetuned_lora_model_dir: str,
merged_model_output_dir: str,
access_token: Optional[str] = None,
) -> None:
"""Internal function to merges the base model with the lora adapter."""
logging.info("loading tokenizer...")
tokenizer = load_tokenizer(pretrained_model_id)
# Note: merging peft adapter requires loading model in 16 bits, so merging
# is done on CPU on purpose in case one GPU cannot hold the base model.
logging.info("loading model %s...", pretrained_model_id)
device_map = "cpu"
model = load_model(
pretrained_model_id=pretrained_model_id,
tokenizer=tokenizer,
precision_mode=merge_precision_mode,
access_token=access_token,
device_map=device_map,
is_training=False,
)
logging.info("loading LoRA model...")
model = PeftModel.from_pretrained(
model, finetuned_lora_model_dir, device_map=device_map
)
logging.info("merging base model with finetuned LoRA model...")
model = model.merge_and_unload()
logging.info("saving model to %s...", merged_model_output_dir)
model.save_pretrained(
merged_model_output_dir,
safe_serialization=False,
is_main_process=PartialState().is_main_process,
)
logging.info("saving tokenizer to %s...", merged_model_output_dir)
tokenizer.save_pretrained(
merged_model_output_dir,
is_main_process=PartialState().is_main_process,
)
def merge_causal_language_model_with_lora_fsdp(
pretrained_model_id: str,
merge_precision_mode: str,
finetuned_lora_model_dir: str,
merged_model_output_dir: str,
access_token: Optional[str] = None,
) -> None:
"""Merges the base model with the lora adapter for FSDP.
Only the main process should call this function.
Args:
pretrained_model_id: Predefined base model name or path to directory
containing model checkpoints.
merge_precision_mode: Precision mode for saving model weights.
finetuned_lora_model_dir: Path to directory containing PEFT-finetuned model
weights.
merged_model_output_dir: Path to directory to save the merged model.
access_token: Access token for accessing the model.
"""
assert PartialState().is_main_process
_merge_causal_language_model_with_lora_internal(
pretrained_model_id=pretrained_model_id,
merge_precision_mode=merge_precision_mode,
finetuned_lora_model_dir=finetuned_lora_model_dir,
merged_model_output_dir=merged_model_output_dir,
access_token=access_token,
)
def merge_causal_language_model_with_lora(
pretrained_model_id: str,
precision_mode: str,
finetuned_lora_model_dir: str,
merged_model_output_dir: str,
access_token: Optional[str] = None,
) -> None:
"""Merges the base model with the lora adapter."""
# Set merge related variables.
if precision_mode == constants.PRECISION_MODE_FP8:
# Merge as FP16. FP8 requires conversion after merge.
merge_precision_mode = constants.PRECISION_MODE_16
local_merged_model_dir = _LOCAL_MERGED_MODEL_DIR
else:
merge_precision_mode = precision_mode
local_merged_model_dir = merged_model_output_dir
if PartialState().is_main_process:
logging.info("Starting merging job...")
# When deepspeed Zero3 is enabled, users are not allowed to specify
# `device_map` when loading the model (even on CPU).
#
# To work-around this, we kick off another process (from the
# is_main_process) and set up the environment to avoid using Deepspeed when
# doing the merging.
if is_deepspeed_zero3_enabled():
ctx = mp.get_context("spawn")
os.environ["ACCELERATE_USE_DEEPSPEED"] = "false"
merge_job = ctx.Process(
target=_merge_causal_language_model_with_lora_internal,
args=(
pretrained_model_id,
merge_precision_mode,
finetuned_lora_model_dir,
local_merged_model_dir,
),
kwargs={
"access_token": access_token,
},
)
merge_job.start()
merge_job.join()
os.environ["ACCELERATE_USE_DEEPSPEED"] = "true"
else:
_merge_causal_language_model_with_lora_internal(
pretrained_model_id=pretrained_model_id,
merge_precision_mode=merge_precision_mode,
finetuned_lora_model_dir=finetuned_lora_model_dir,
merged_model_output_dir=local_merged_model_dir,
access_token=access_token,
)
logging.info("merging job is done")
# Wait for all processes to sync here.
PartialState().wait_for_everyone()
if precision_mode == constants.PRECISION_MODE_FP8:
convert_model_to_fp8(
pretrained_model_name_or_path=pretrained_model_id,
merged_model_output_dir=local_merged_model_dir,
quantized_model_output_dir=merged_model_output_dir,
access_token=access_token,
)
def convert_model_to_fp8(
pretrained_model_name_or_path: str,
merged_model_output_dir: str,
quantized_model_output_dir: str,
access_token: Optional[str] = None,
) -> None:
"""Converts the model to fp8.
Args:
pretrained_model_name_or_path: Original base model name or path.
merged_model_output_dir: Path to directory containing the merged model.
quantized_model_output_dir: Path to directory to save the quantized model.
access_token: Access token for accessing the model.
"""
if PartialState().is_main_process:
quantization_config = FbgemmFp8Config(
modules_to_not_convert=_maybe_get_modules_to_not_convert_by_model_id(
pretrained_model_name_or_path
)
)
quantized_model = AutoModelForCausalLM.from_pretrained(
merged_model_output_dir,
device_map="cpu",
quantization_config=quantization_config,
trust_remote_code=False,
token=access_token,
)
tokenizer = AutoTokenizer.from_pretrained(merged_model_output_dir)
quantized_model.save_pretrained(quantized_model_output_dir)
tokenizer.save_pretrained(quantized_model_output_dir)
PartialState().wait_for_everyone()
@dataclasses.dataclass
class TuningDataStats:
tuning_dataset_example_count: int
total_billable_token_count: int
tuning_step_count: int
def get_dataset_stats(
dataset: Any,
tokenizer: transformers.PreTrainedTokenizer,
column: str,
effective_batch_size: int,
) -> TuningDataStats:
"""Calculates dataset statistics, e.g., total number of tokens."""
tokenized_dataset = dataset.map(lambda x: tokenizer(x[column]))
inputs = tokenized_dataset["input_ids"]
tuning_dataset_example_count = int(len(inputs))
total_billable_token_count = int(np.sum([len(ex) for ex in inputs]))
tuning_step_count = (
tuning_dataset_example_count + effective_batch_size - 1
) // effective_batch_size
return TuningDataStats(
tuning_dataset_example_count,
total_billable_token_count,
tuning_step_count,
)
def force_gc():
"""Collects garbage immediately to release unused CPU/GPU resources."""
gc.collect()
torch.cuda.empty_cache()
def should_add_pad_token(model_id: str) -> bool:
"""Returns whether the model requires adding a special pad token."""
return any(s.lower() in model_id.lower() for s in _MODELS_REQUIRING_PAD_TOKEN)
def should_add_eos_token(model_id: str) -> bool:
"""Returns whether the model requires adding a special eos token."""
return any(m in model_id for m in _MODELS_REQUIRING_EOS_TOEKN)
def write_kfp_outputs(
executor_input: str, output_artifacts: Dict[str, str]
) -> None:
"""Writes KFP outputs given a dict of output artifact names and URIs."""
# Only the main process writes to avoid race condition.
if PartialState().is_main_process:
executor_input = json_format.Parse(
executor_input, pipeline_spec_pb2.ExecutorInput()
)
outputs = executor_input.outputs
# set all artifacts
for name, uri in output_artifacts.items():
artifact_list = outputs.artifacts.get(name)
if not artifact_list or not artifact_list.artifacts:
raise ValueError(f"Artifact name={name} does not exist.")
artifact_list.artifacts[0].uri = uri
# write output file
executor_output = pipeline_spec_pb2.ExecutorOutput(
artifacts=outputs.artifacts
)
os.makedirs(os.path.dirname(outputs.output_file), exist_ok=True)
with open(outputs.output_file, "w") as f:
f.write(json_format.MessageToJson(executor_output, indent=None))
# Wait for the main process to finish before moving on to the next task.
PartialState().wait_for_everyone()
def upload_local_dir_to_gcs(local_dir: str, gcs_path: str):
"""Uploads local dir to GCS."""
if PartialState().is_main_process:
logging.info("uploading %s to %s...", local_dir, gcs_path)
subprocess.check_output([
"gsutil",
"-m",
"cp",
"-r",
local_dir,
gcs_path,
])
logging.info("%s uploaded.", local_dir)
PartialState().wait_for_everyone()
def write_first_party_model_metadata(output_dir: str, docker_uri: str) -> None:
"""Multi-process friendly version of fileutils.write_first_party_model_metadata."""
if PartialState().is_main_process:
fileutils.write_first_party_model_metadata(output_dir, docker_uri)
PartialState().wait_for_everyone()
@dataclasses.dataclass
class GpuStats:
"""Holds information about GPU usage stats.
For memory related, see
https://pytorch.org/docs/stable/notes/cuda.html#cuda-memory-management
"""
# total memory
total_mem: float
# memory occupied.
occupied: float
# memory reserved, but not used.
unused: float
# nvidia-smi usually reports more memory usages than pytorch (for driver,
# kernel and etc). `smi_diff` tracks this difference.
smi_diff: float
# Gpu utilization.
util: float
# Allows unpacking operation like
# total_mem, occupied, unused, smi_diff, util = GpuStats(...)
# See https://stackoverflow.com/a/70753113
def __iter__(self):
return iter(dataclasses.astuple(self))
def gpu_stats() -> GpuStats:
"""Reports GPU memory usage and utilization."""
# See https://pytorch.org/docs/stable/notes/cuda.html#memory-management
bytes_per_gb = 1024.0**3
device = torch.cuda.current_device()
occupied = torch.cuda.memory_allocated(device) / bytes_per_gb
reserved = torch.cuda.memory_reserved(device) / bytes_per_gb
unused = reserved - occupied
def smi_mem(device):
try:
pynvml.nvmlInit()
handle = pynvml.nvmlDeviceGetHandleByIndex(device)
info = pynvml.nvmlDeviceGetMemoryInfo(handle)
return info.used / bytes_per_gb
except pynvml.NVMLError:
return 0.0
mem_used_smi = smi_mem(device)
smi_diff = mem_used_smi - reserved
util = torch.cuda.utilization(device)
return GpuStats(mem_used_smi, occupied, unused, smi_diff, util)
def gpu_stats_str(stats: Optional[GpuStats] = None) -> str:
if stats is None:
stats = gpu_stats()
total, occupied, unused, smi_diff, util = stats
return (
f"GPU memory: {total:.2f}({occupied=:.2f}, {unused=:.2f},"
f" {smi_diff=:.2f}) GB. Utilization: {util:.2f}%"
)
def init_partial_state(
timeout: datetime.timedelta = datetime.timedelta(seconds=600),
) -> None:
"""Initializes the partial state with timeout."""
# This needs to be called before any other PartialState() calls, and
# TrainingArguments needs `use_configured_state`. See b/357970482#comment3
# for more details.
PartialState(timeout=timeout)
def print_library_versions():
if PartialState().is_main_process:
logging.info("======================")
logging.info("library versions")
logging.info("======================")
logging.info("accelerate: %s", accelerate.__version__)
logging.info("peft: %s", peft.__version__)
logging.info("transformers: %s", transformers.__version__)
logging.info("trl: %s", trl.__version__)
PartialState().wait_for_everyone()
def get_final_checkpoint_path(output_dir: str) -> str:
"""Returns the final checkpoint path."""
return os.path.join(output_dir, constants.FINAL_CHECKPOINT_DIRNAME)
def _maybe_get_modules_to_not_convert_by_model_id(
pretrained_model_name_or_path: str,
) -> Optional[Sequence[str]]:
"""Returns the modules to not convert for the model."""
if _LLAMA_3_1_405B_MODEL_ID in pretrained_model_name_or_path:
return _get_llama_3_1_405b_modules_to_not_convert()
else:
return None
def _get_llama_3_1_405b_modules_to_not_convert() -> Sequence[str]:
"""Returns the modules to not convert for Llama 3.1 405B model."""
modules_to_not_convert = ["lm_head"]
for idx in range(126):
for proj_name in ["k_proj", "o_proj", "q_proj", "v_proj"]:
modules_to_not_convert.append(f"model.layers.{idx}.self_attn.{proj_name}")
for proj_name in ["down_proj", "gate_proj", "up_proj"]:
modules_to_not_convert.append(f"model.layers.0.mlp.{proj_name}")
modules_to_not_convert.append(f"model.layers.125.mlp.{proj_name}")
return tuple(modules_to_not_convert)
@@ -0,0 +1,77 @@
"""Validate the dataset with the template."""
from typing import Sequence
from absl import app
from absl import flags
from util import dataset_validation_util
from vertex_vision_model_garden_peft.train.vmg import utils
from util import constants
_DATASET_NAME = flags.DEFINE_string(
'dataset_name',
None,
'The dataset name in huggingface.',
required=True,
)
_TRAIN_SPLIT_NAME = flags.DEFINE_string(
'train_split_name',
'train',
'The train split name.',
)
_INSTRUCT_COLUMN_IN_DATASET = flags.DEFINE_string(
'instruct_column_in_dataset',
constants.DEFAULT_INSTRUCT_COLUMN_IN_DATASET,
'The instruct column in dataset.',
)
_TEMPLATE = flags.DEFINE_string(
'template',
None,
'Template for formatting language model training data. Must be a filename'
' under `templates` folder, without `.json` extension, e.g. `alpaca`, or a'
' Cloud Storage URI to a JSON file.',
required=True,
)
_VALIDATE_PERCENTAGE_OF_DATASET = flags.DEFINE_integer(
'validate_percentage_of_dataset',
None,
'The percentage of the dataset to validate with the template. If set to'
' -1, it loads the full dataset.',
)
_VALIDATE_K_ROWS_OF_DATASET = flags.DEFINE_integer(
'validate_k_rows_of_dataset',
None,
'The top k rows of the dataset to validate with the template. If set to -1,'
' it loads the full dataset.',
)
_USE_MULTIPROCESSING = flags.DEFINE_boolean(
'use_multiprocessing',
False,
'Whether to use multiprocessing for loading the dataset.',
)
def main(unused_argv: Sequence[str]) -> None:
utils.print_library_versions()
dataset_validation_util.validate_dataset_with_template(
dataset_name=_DATASET_NAME.value,
split=_TRAIN_SPLIT_NAME.value,
input_column=_INSTRUCT_COLUMN_IN_DATASET.value,
template=_TEMPLATE.value,
use_multiprocessing=_USE_MULTIPROCESSING.value,
validate_percentage_of_dataset=_VALIDATE_PERCENTAGE_OF_DATASET.value,
validate_k_rows_of_dataset=_VALIDATE_K_ROWS_OF_DATASET.value,
)
if __name__ == '__main__':
app.run(main)
@@ -1,6 +1,6 @@
"""Common utility lib for prediction on images."""
from typing import Any, Dict, List
from typing import Any, Dict, List, Tuple
import numpy as np
from PIL import Image
@@ -10,6 +10,28 @@ import yaml
from util import image_format_converter
def convert_list_to_label_map(
input_list: List[str],
) -> Tuple[Dict[str, Dict[int, str]], List[int]]:
"""Converts a list of labels to a dictionary and numerical encoding.
Args:
input_list: A list of strings representing class labels.
Returns:
A tuple containing:
label_map: A dictionary mapping unique labels to integer indices.
encoded_list: A list of integers corresponding to the labels in the input
list.
"""
unique_labels = set(input_list)
label_map_reverse = {label: idx for idx, label in enumerate(unique_labels)}
label_map = {idx: label for idx, label in enumerate(unique_labels)}
encoded_list = [label_map_reverse[label] for label in input_list]
return {"label_map": label_map}, encoded_list
def get_prediction_instances(image: Image.Image) -> List[Dict[str, Any]]:
"""Gets prediction instances.
@@ -40,14 +62,14 @@ def get_label_map(label_map_yaml_filepath: str) -> Dict[str, Any]:
def get_object_detection_endpoint_predictions(
detection_endpoint: ...,
detector_endpoint: ...,
input_image: np.ndarray,
detection_thresh: float = 0.2,
) -> np.ndarray:
"""Gets endpoint predictions.
Args:
detection_endpoint: image object detection endpoint.
detector_endpoint: image object detection endpoint.
input_image: Input image.
detection_thresh: Detection threshold.
@@ -55,9 +77,10 @@ def get_object_detection_endpoint_predictions(
Object detection predictions from endpoints.
"""
height, width, _ = input_image.shape
predictions = detection_endpoint.predict(
predictions = detector_endpoint.predict(
get_prediction_instances(Image.fromarray(input_image))
).predictions
detection_scores = np.array(predictions[0]["detection_scores"])
detection_classes = np.array(predictions[0]["detection_classes"])
detection_boxes = np.array(
@@ -66,6 +89,29 @@ def get_object_detection_endpoint_predictions(
for b in predictions[0]["detection_boxes"]
]
)
return merge_boxes_and_classes(
detection_scores, detection_boxes, detection_classes, detection_thresh
)
def merge_boxes_and_classes(
detection_scores: np.ndarray,
detection_boxes: np.ndarray,
detection_classes: np.ndarray,
detection_thresh: float = 0.2,
) -> np.ndarray:
"""Merges prediction boxes and classes.
Args:
detection_scores: array of detection scores.
detection_boxes: array of detection boxes.
detection_classes: array of detection classes.
detection_thresh: float indicating the detection threshold.
Returns:
preds_merge_cls: a numpy array containing the detection boxes, scores and
classes.
"""
thresh_indices = [
x for x, val in enumerate(detection_scores) if val > detection_thresh
]
@@ -76,4 +122,5 @@ def get_object_detection_endpoint_predictions(
preds_merge_cls = np.column_stack(
(preds_merge_conf, detection_classes[thresh_indices])
)
return preds_merge_cls
@@ -36,6 +36,12 @@ BEST_CKPT_METRIC_COMP = 'higher'
# Reported hyperparameter tuning metric tag.
HP_METRIC_TAG = 'model_performance'
HP_LOSS_TAG = 'model_loss'
# Reported places.
REPORT_TO_NONE = 'none'
REPORT_TO_WANDB = 'wandb'
REPORT_TO_TENSORBOARD = 'tensorboard'
# HPT trial prefix.
TRIAL_PREFIX = 'trial_'
@@ -45,7 +51,7 @@ ML_USE_TRAINING = 'training'
ML_USE_VALIDATION = 'validation'
ML_USE_TEST = 'test'
# COCO json keys
# COCO json keys.
COCO_JSON_ANNOTATIONS = 'annotations'
COCO_JSON_ANNOTATION_IMAGE_ID = 'image_id'
COCO_JSON_ANNOTATION_CATEGORY_ID = 'category_id'
@@ -60,26 +66,88 @@ COCO_JSON_IMAGE_HEIGHT = 'height'
COCO_JSON_IMAGE_COCO_URL = 'coco_url'
COCO_ANNOTATION_BBOX = 'bbox'
# GCS prefixes
# GCS prefixes.
GCS_URI_PREFIX = 'gs://'
GCSFUSE_URI_PREFIX = '/gcs/'
LOCAL_EVALUATION_RESULT_DIR = '/tmp/evaluation_result_dir'
LOCAL_MODEL_DIR = '/tmp/model_dir'
LOCAL_LORA_DIR = '/tmp/lora_dir'
LOCAL_BASE_MODEL_DIR = '/tmp/base_model_dir'
LOCAL_DATA_DIR = '/tmp/data'
LOCAL_OUTPUT_DIR = '/tmp/output_dir'
LOCAL_PREDICTION_RESULT_DIR = '/tmp/prediction_result_dir'
SHARED_MEM_DIR = '/dev/shm'
# Huggingface files.
HF_MODEL_WEIGHTS_SUFFIX = '.bin'
# PEFT finetuning constants.
TEXT_TO_IMAGE = 'text-to-image'
TEXT_TO_IMAGE_LORA = 'text-to-image-lora'
TEXT_TO_IMAGE_DREAMBOOTH = 'text-to-image-dreambooth'
TEXT_TO_IMAGE_DREAMBOOTH_LORA = 'text-to-image-dreambooth-lora'
TEXT_TO_IMAGE_DREAMBOOTH_LORA_SDXL = 'text-to-image-dreambooth-lora-sdxl'
SEQUENCE_CLASSIFICATION_LORA = 'sequence-classification-lora'
CAUSAL_LANGUAGE_MODELING_LORA = 'causal-language-modeling-lora'
MERGE_CAUSAL_LANGUAGE_MODEL_LORA = 'merge-causal-language-model-lora'
QUANTIZE_MODEL = 'quantize-model'
INSTRUCT_LORA = 'instruct-lora'
VALIDATE_DATASET_WITH_TEMPLATE = 'validate-dataset-with-template'
DEFAULT_TEXT_COLUMN_IN_DATASET = 'quote'
DEFAULT_TEXT_COLUMN_IN_QUANTIZATION_DATASET = 'text'
DEFAULT_INSTRUCT_COLUMN_IN_DATASET = 'text'
FINAL_CHECKPOINT_DIRNAME = 'checkpoint-final'
# ImageBind inference constants.
FEATURE_EMBEDDING_GENERATION = 'feature-embedding-generation'
ZERO_SHOT_CLASSIFICATION = 'zero-shot-classification'
# Precision modes for loading model weights.
PRECISION_MODE_2 = '2bit'
PRECISION_MODE_3 = '3bit'
PRECISION_MODE_4 = '4bit'
PRECISION_MODE_8 = '8bit'
PRECISION_MODE_FP8 = 'float8' # to use fbgemm_fp8 quantization
PRECISION_MODE_16 = 'float16'
PRECISION_MODE_16B = 'bfloat16'
PRECISION_MODE_32 = 'float32'
# Quantization modes.
GPTQ = 'gptq'
AWQ = 'awq'
# AWQ versions.
GEMM = 'GEMM'
GEMV = 'GEMV'
# Environment variable keys.
PRIVATE_BUCKET_ENV_KEY = 'AIP_PRIVATE_BUCKET_NAME'
# Kfp pipeline constants.
TFVISION_TRAIN_OUTPUT_ARTIFACT_NAME = 'checkpoint_dir'
# Vertex IOD type.
AUTOML = 'AUTOML'
MODEL_GARDEN = 'MODEL_GARDEN'
# LRU Disk Cache constants.
MD5_HASHMAP_FILENAME = 'md5_hashmap.json'
# Prediction request keys.
PREDICT_INSTANCE_KEY = 'instances'
PREDICT_INSTANCE_IMAGE_KEY = 'image'
PREDICT_INSTANCE_POSE_IMAGE_KEY = 'pose_image'
PREDICT_INSTANCE_TEXT_KEY = 'text'
PREDICT_INSTANCE_PROMPT_KEY = 'prompt'
PREDICT_PARAMETERS_KEY = 'parameters'
PREDICT_PARAMETERS_NUM_INFERENCE_STEPS_KEY = 'num_inference_steps'
PREDICT_PARAMETERS_HEIGHT_KEY = 'height'
PREDICT_PARAMETERS_WIDTH_KEY = 'width'
PREDICT_PARAMETERS_GUIDANCE_SCALE_KEY = 'guidance_scale'
PREDICT_PARAMETERS_NEGATIVE_PROMPT_KEY = 'negative_prompt'
PREDICT_PARAMETERS_LORA_ID_KEY = 'lora_id'
PREDICT_PARAMETERS_IGNORE_LORA_CACHE_KEY = 'ignore_lora_cache'
PREDICT_OUTPUT_KEY = 'output'
@@ -1,10 +1,10 @@
"""Fileutil lib to copy files between gcs and local."""
import glob
import fnmatch
import os
import pathlib
import shutil
from typing import Tuple
from typing import List, Optional, Tuple
import uuid
from absl import logging
@@ -13,6 +13,17 @@ from google.cloud import storage
from util import constants
_GCS_CLIENT = None
def _get_gcs_client() -> storage.Client:
"""Gets the default GCS client."""
global _GCS_CLIENT
if _GCS_CLIENT is None:
_GCS_CLIENT = storage.Client()
return _GCS_CLIENT
def generate_tmp_path(extension: str = '') -> str:
"""Generates a temporary file path with UUID.
@@ -36,6 +47,16 @@ def force_gcs_fuse_path(gcs_uri: str) -> str:
return gcs_uri
def force_gcs_path(uri: str) -> str:
"""Converts /gcs/ uris to their gs:// equivalents. No-op for other uris."""
if uri.startswith(constants.GCSFUSE_URI_PREFIX):
return uri.replace(
constants.GCSFUSE_URI_PREFIX, constants.GCS_URI_PREFIX, 1
)
else:
return uri
def download_gcs_file_to_local_dir(gcs_uri: str, local_dir: str):
"""Download a gcs file to a local dir.
@@ -62,15 +83,47 @@ def download_gcs_file_to_local(gcs_uri: str, local_path: str):
raise ValueError(
f'{gcs_uri} is not a GCS path starting with {constants.GCS_URI_PREFIX}.'
)
client = storage.Client()
client = _get_gcs_client()
os.makedirs(os.path.dirname(local_path), exist_ok=True)
with open(local_path, 'wb') as f:
client.download_blob_to_file(gcs_uri, f)
def download_gcs_file_list_to_local(
gcs_uri_list: List[str], local_dir: str
) -> List[str]:
"""Downloads a list of GCS files to a local directory.
Args:
gcs_uri_list: A list of GCS file paths.
local_dir: Local directory in which the GCS files are saved.
Returns:
The local file paths corresponding to the input GCS file paths.
Raises:
ValueError: An input file path is not a GCS path.
"""
local_paths = []
for gcs_uri in gcs_uri_list:
if not is_gcs_path(gcs_uri):
raise ValueError(
f'{gcs_uri} is not a GCS path starting with'
f' {constants.GCS_URI_PREFIX}.'
)
local_path = os.path.join(local_dir, gcs_uri.replace('gs://', ''))
download_gcs_file_to_local(gcs_uri, local_path)
local_paths.append(local_path)
return local_paths
def download_gcs_dir_to_local(
gcs_dir: str, local_dir: str, skip_hf_model_bin: bool = False
):
gcs_dir: str,
local_dir: str,
skip_hf_model_bin: bool = False,
allow_patterns: Optional[List[str]] = None,
log: bool = True,
) -> None:
"""Downloads files in a GCS directory to a local directory.
For example:
@@ -78,16 +131,21 @@ def download_gcs_dir_to_local(
gs://bucket/foo/a -> /tmp/bar/a
gs://bucket/foo/b/c -> /tmp/bar/b/c
Arguments:
Args:
gcs_dir: A string of directory path on GCS.
local_dir: A string of local directory path.
skip_hf_model_bin: True to skip downloading HF model bin files.
allow_patterns: A list of allowed patterns. If provided, only files matching
one or more patterns are downloaded.
log: True to log each downloaded file.
"""
if not is_gcs_path(gcs_dir):
raise ValueError(f'{gcs_dir} is not a GCS path starting with gs://.')
bucket_name = gcs_dir.split('/')[2]
prefix = gcs_dir[len(constants.GCS_URI_PREFIX + bucket_name) :].strip('/')
client = storage.Client()
prefix = (
gcs_dir[len(constants.GCS_URI_PREFIX + bucket_name) :].strip('/') + '/'
)
client = _get_gcs_client()
blobs = client.list_blobs(bucket_name, prefix=prefix)
for blob in blobs:
if blob.name[-1] == '/':
@@ -95,43 +153,63 @@ def download_gcs_dir_to_local(
file_path = blob.name[len(prefix) :].strip('/')
local_file_path = os.path.join(local_dir, file_path)
os.makedirs(os.path.dirname(local_file_path), exist_ok=True)
if allow_patterns and all(
[not fnmatch.fnmatch(file_path, p) for p in allow_patterns]
):
continue
if (
file_path.endswith(constants.HF_MODEL_WEIGHTS_SUFFIX)
and skip_hf_model_bin
):
logging.info('Skip downloading model bin %s', file_path)
if log:
logging.info('Skip downloading model bin %s', file_path)
with open(local_file_path, 'w') as f:
f.write(f'{constants.GCS_URI_PREFIX}{bucket_name}/{prefix}/{file_path}')
f.write(f'{constants.GCS_URI_PREFIX}{bucket_name}/{prefix}{file_path}')
else:
logging.info('Downloading %s to %s', file_path, local_file_path)
if log:
logging.info('Downloading %s to %s', file_path, local_file_path)
blob.download_to_filename(local_file_path)
def _get_relative_paths(base_dir: str) -> List[str]:
"""Gets relative paths of all files in a local base directory."""
path = pathlib.Path(base_dir)
relative_paths = []
for local_file in path.rglob('*'):
if os.path.isfile(local_file):
relative_path = os.path.relpath(local_file, base_dir)
relative_paths.append(relative_path)
return relative_paths
def _upload_local_files_to_gcs(
relative_paths: List[str], local_dir: str, gcs_dir: str
):
"""Uploads local files to gcs."""
bucket_name = gcs_dir.split('/')[2]
blob_dir = '/'.join(gcs_dir.split('/')[3:])
client = _get_gcs_client()
bucket = client.bucket(bucket_name)
for relative_path in relative_paths:
blob = bucket.blob(os.path.join(blob_dir, relative_path))
blob.upload_from_filename(os.path.join(local_dir, relative_path))
def upload_local_dir_to_gcs(local_dir: str, gcs_dir: str):
"""Uploads local dir to gcs.
For example:
upload_local_dir_to_gcs(/tmp/bar, gs://bucket/foo)
gs://bucket/foo/a -> /tmp/bar/a
gs://bucket/foo/b/c -> /tmp/bar/b/c
/tmp/bar/a -> gs://bucket/foo/a
/tmp/bar/b/c -> gs://bucket/foo/b/c
Arguments:
local_dir: A string of local directory path.
gcs_dir: A string of directory path on GCS.
"""
bucket_name = gcs_dir.split('/')[2]
blob_dir = '/'.join(gcs_dir.split('/')[3:])
client = storage.Client()
bucket = client.bucket(bucket_name)
for local_file in glob.glob(local_dir + '/**'):
if os.path.isfile(local_file):
logging.info(
'Uploading %s to %s',
local_file,
os.path.join(constants.GCS_URI_PREFIX, bucket_name, blob_dir),
)
blob = bucket.blob(os.path.join(blob_dir, os.path.basename(local_file)))
blob.upload_from_filename(local_file)
# Relative paths of all files in local_dir.
relative_paths = _get_relative_paths(local_dir)
_upload_local_files_to_gcs(relative_paths, local_dir, gcs_dir)
def upload_file_to_gcs_path(
@@ -155,7 +233,7 @@ def upload_file_to_gcs_path(
if not source_path_obj.exists():
raise RuntimeError(f'Source path does not exist: {source_path}')
storage_client = storage.Client()
storage_client = _get_gcs_client()
source_file_path = source_path
destination_file_uri = destination_uri
logging.info('Uploading "%s" to "%s"', source_file_path, destination_file_uri)
@@ -174,7 +252,9 @@ def is_gcs_path(input_path: str) -> bool:
Returns:
True if the input path is a GCS path, False otherwise.
"""
return input_path.startswith(constants.GCS_URI_PREFIX)
return input_path is not None and input_path.startswith(
constants.GCS_URI_PREFIX
)
def release_text_assets(
@@ -232,13 +312,10 @@ def download_video_from_gcs_to_local(video_file_path: str) -> Tuple[str, str]:
"""
_, local_video_file_name = os.path.split(video_file_path)
file_extension = os.path.splitext(video_file_path)[1]
if file_extension:
remote_video_file_name = local_video_file_name.replace(
file_extension, '_overlay.mp4'
)
else:
remote_video_file_name = local_video_file_name + '_overlay.mp4'
local_file_path = generate_tmp_path(file_extension)
remote_video_file_name = local_video_file_name.replace(
file_extension, '_overlay.mp4'
)
local_file_path = generate_tmp_path(os.path.splitext(video_file_path)[1])
logging.info('Downloading %s to %s...', video_file_path, local_file_path)
download_gcs_file_to_local(video_file_path, local_file_path)
return local_file_path, remote_video_file_name
@@ -254,10 +331,28 @@ def get_output_video_file(video_output_file_path: str) -> str:
str: Local video output file path.
"""
file_extension = os.path.splitext(video_output_file_path)[1]
if file_extension:
out_local_video_file_name = video_output_file_path.replace(
file_extension, '_overlay' + file_extension
)
else:
out_local_video_file_name = video_output_file_path + '_overlay'
out_local_video_file_name = video_output_file_path.replace(
file_extension, '_overlay' + file_extension
)
return out_local_video_file_name
def write_first_party_model_metadata(
output_path: str, required_container_uri: str
) -> None:
"""Write Vertex internal model metadata for first party artifacts."""
model_metadata_fname = 'model_metadata.jsonl'
if len(required_container_uri) > 126:
raise ValueError(f'Docker URI exceeds 126 chars: {required_container_uri}')
payload = '\n{}{}'.format( # serialized proto
chr(len(required_container_uri)),
required_container_uri,
)
os.makedirs(output_path, exist_ok=True)
output_dirs = [output_path]
if output_path.startswith('/gcs'):
# include all parent dirs, except "/", "/gcs"
output_dirs.extend([str(p) for p in pathlib.Path(output_path).parents][:-2])
for output_dir in output_dirs:
with open(os.path.join(output_dir, model_metadata_fname), 'w') as f:
f.write(payload)
@@ -20,3 +20,11 @@ def get_trial_id_from_environment() -> str:
_ENVIRONMENT_VARIABLE_FOR_TRIAL_ID,
)
return os.environ.get(_ENVIRONMENT_VARIABLE_FOR_TRIAL_ID, '0')
def maybe_append_trial_id(path: str) -> str:
"""Appends trial_N to path if running in a Hyperparameter Tuning Job."""
trial_id = os.environ.get(_ENVIRONMENT_VARIABLE_FOR_TRIAL_ID)
if trial_id is None:
return path
return os.path.join(path, f'trial_{trial_id}')
@@ -0,0 +1,32 @@
import numpy as np
from kfp.v2 import dsl
@dsl.component(base_image='python:3.8',packages_to_install=['google-cloud-aiplatform==1.36.0'])
def async_predict(
endpoint_id: str,
instances: dict,
) -> np.ndarray:
import numpy as np
from google.cloud import aiplatform
endpoint = aiplatform.Endpoint(endpoint_id)
response = await endpoint.predict_async(instances)
predictions = np.asarray(response.predictions)
print(predictions.tolist())
return predictions
@dsl.pipeline(name='async-prediction')
def pipeline_prediction():
project = "projects/990000000009/locations/us-west1"
endpoint_id = project + "/endpoints/2200000000000000002"
instances = [{
"key1": "value1",
"key2": 2
}]
async_predict(endpoint_id, instances)
if __name__ == "__main__":
from kfp.v2 import compiler
compiler.Compiler().compile(
pipeline_func=pipeline_prediction,
package_path='async_prediction.json')
@@ -0,0 +1,41 @@
from kfp.v2 import dsl
@dsl.component(base_image='python:3.8',packages_to_install=['google-cloud-aiplatform==1.36.0'])
def customjob(
project_id: str,
location: str,
staging_bucket: str,
experiment: str,
job_name: str,
script_path: str,
container_uri: str,
machine_type: str,
):
import os
from google.cloud import aiplatform
aiplatform.init(
project=project_id,
location=location,
staging_bucket=staging_bucket,
experiment=experiment,
)
job = aiplatform.CustomJob.from_local_script(
display_name=job_name,
script_path=os.path.join(os.getcwd(), script_path),
container_uri=container_uri,
machine_type=machine_type,
)
job.run()
@dsl.pipeline(name='run-customjob')
def pipeline_customjob():
customjob("990000000009", "us-west1", "gs://staging-bucket/customjob",
"run-experiment", "custom-job", "customjob.py",
"gcr.io/path/to/model_name:latest", "n1-standard-4")
if __name__ == "__main__":
from kfp.v2 import compiler
compiler.Compiler().compile(
pipeline_func=pipeline_customjob,
package_path='customjob.json')

Some files were not shown because too many files have changed in this diff Show More