Compare commits

...
Author SHA1 Message Date
Rayan DasoriyaandCopybara-Service 9cf8ce16fa Delete deprecated LoRA fine-tuning notebooks and related tutorials.
PiperOrigin-RevId: 976392412
2026-09-04 10:44:57 -07:00
Chun-Hsiang WangandGitHub 4b983a2701 feat: Claude Fable 5.1 Launch (#4581)
* feat: Claude Fable 5.1 Launch

* refactor: replace model/region if-elif chains with a dict lookup

Addresses review feedback on both Select Claude model cells. The mapping is
unchanged for all 20 models; only the lookup mechanism differs.

* chore: apply nbfmt

Runs the repo's own tensorflow-docs nbfmt over the notebook so the
'notebook format and lint' check passes.
2026-09-01 20:45:17 -04:00
Eric DongandGitHub 3b11c876bd fix: correct a typo in error message (#4577) 2026-08-25 17:03:21 -04:00
Mend RenovateandGitHub cc0d791ef2 Update dependency black to v26.5.1 (#4517) 2026-08-19 21:44:22 +00:00
Mend RenovateandGitHub df83a345bb Update dependency isort to v8 (#4444) 2026-08-19 20:52:37 +00:00
Mend RenovateandGitHub e6ded7beaa Update dependency pandas to v3.0.5 (#4491) 2026-08-19 20:51:20 +00:00
Mend RenovateandGitHub e64a4e89d5 chore(deps): update dependency google-cloud-aiplatform to v1.165.0 (#4457) 2026-08-19 20:50:48 +00:00
Mend RenovateandGitHub 7ac54985e4 chore(deps): update dependency smart_open to v8 (#4534) 2026-08-18 22:49:25 +00:00
Mend RenovateandGitHub 6ce96a08d3 Update dependency smart_open to v7.7.1 (#4494) 2026-08-18 21:16:42 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
756711b3c9 chore(deps): bump idna (#4518)
Bumps [idna](https://github.com/kjd/idna) from 3.10 to 3.15.
- [Release notes](https://github.com/kjd/idna/releases)
- [Changelog](https://github.com/kjd/idna/blob/master/HISTORY.md)
- [Commits](https://github.com/kjd/idna/compare/v3.10...v3.15)

---
updated-dependencies:
- dependency-name: idna
  dependency-version: '3.15'
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-18 21:15:21 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
a25d209139 chore(deps): bump torch (#4545)
Bumps [torch](https://github.com/pytorch/pytorch) from 2.8.0 to 2.13.0.
- [Release notes](https://github.com/pytorch/pytorch/releases)
- [Changelog](https://github.com/pytorch/pytorch/blob/main/RELEASE.md)
- [Commits](https://github.com/pytorch/pytorch/compare/v2.8.0...v2.13.0)

---
updated-dependencies:
- dependency-name: torch
  dependency-version: 2.13.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-18 21:14:39 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
59da536b9a chore(deps): bump pillow (#4548)
Bumps [pillow](https://github.com/python-pillow/Pillow) from 12.2.0 to 12.3.0.
- [Release notes](https://github.com/python-pillow/Pillow/releases)
- [Changelog](https://github.com/python-pillow/Pillow/blob/main/CHANGES.rst)
- [Commits](https://github.com/python-pillow/Pillow/compare/12.2.0...12.3.0)

---
updated-dependencies:
- dependency-name: pillow
  dependency-version: 12.3.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-18 21:13:58 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
1bc2839a2b chore(deps): bump pillow (#4568)
Bumps [pillow](https://github.com/python-pillow/Pillow) from 12.2.0 to 12.3.0.
- [Release notes](https://github.com/python-pillow/Pillow/releases)
- [Changelog](https://github.com/python-pillow/Pillow/blob/main/CHANGES.rst)
- [Commits](https://github.com/python-pillow/Pillow/compare/12.2.0...12.3.0)

---
updated-dependencies:
- dependency-name: pillow
  dependency-version: 12.3.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-18 21:13:27 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
0d5e268a1f chore(deps): bump urllib3 (#4512)
Bumps [urllib3](https://github.com/urllib3/urllib3) from 2.6.3 to 2.7.0.
- [Release notes](https://github.com/urllib3/urllib3/releases)
- [Changelog](https://github.com/urllib3/urllib3/blob/main/CHANGES.rst)
- [Commits](https://github.com/urllib3/urllib3/compare/2.6.3...2.7.0)

---
updated-dependencies:
- dependency-name: urllib3
  dependency-version: 2.7.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-18 21:12:21 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
1b019a76e4 Bump google-cloud-aiplatform (#4446)
Bumps [google-cloud-aiplatform](https://github.com/googleapis/python-aiplatform) from 1.92.0 to 1.133.0.
- [Release notes](https://github.com/googleapis/python-aiplatform/releases)
- [Changelog](https://github.com/googleapis/python-aiplatform/blob/main/CHANGELOG.md)
- [Commits](https://github.com/googleapis/python-aiplatform/compare/v1.92.0...v1.133.0)

---
updated-dependencies:
- dependency-name: google-cloud-aiplatform
  dependency-version: 1.133.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-18 21:11:56 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
87c1ed686a chore(deps): bump diffusers (#4510)
Bumps [diffusers](https://github.com/huggingface/diffusers) from 0.25.1 to 0.38.0.
- [Release notes](https://github.com/huggingface/diffusers/releases)
- [Commits](https://github.com/huggingface/diffusers/compare/v0.25.1...v0.38.0)

---
updated-dependencies:
- dependency-name: diffusers
  dependency-version: 0.38.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-18 21:11:22 +00:00
Mend RenovateandGitHub 187fdc526c Update dependency datasets to v5 (#4521) 2026-08-18 21:10:37 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
215c8eee3e chore(deps): bump urllib3 (#4513)
Bumps [urllib3](https://github.com/urllib3/urllib3) from 2.6.3 to 2.7.0.
- [Release notes](https://github.com/urllib3/urllib3/releases)
- [Changelog](https://github.com/urllib3/urllib3/blob/main/CHANGES.rst)
- [Commits](https://github.com/urllib3/urllib3/compare/2.6.3...2.7.0)

---
updated-dependencies:
- dependency-name: urllib3
  dependency-version: 2.7.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-18 21:10:01 +00:00
f90cd0d6ed Add AlphaFold 3 quickstart notebook (#4572)
* Add AlphaFold 3 quickstart notebook

* Update CODEOWNERS

---------

Co-authored-by: Amit Rai <raiamit@google.com>
2026-08-17 13:54:13 -07:00
Mend RenovateandGitHub 1985f06e99 Update dependency numpy to v2.5.2 (#4516) 2026-08-14 18:43:39 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
ff428dc589 chore(deps): bump torch (#4544)
Bumps [torch](https://github.com/pytorch/pytorch) from 2.7.0 to 2.13.0.
- [Release notes](https://github.com/pytorch/pytorch/releases)
- [Changelog](https://github.com/pytorch/pytorch/blob/main/RELEASE.md)
- [Commits](https://github.com/pytorch/pytorch/compare/v2.7.0...v2.13.0)

---
updated-dependencies:
- dependency-name: torch
  dependency-version: 2.13.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-14 18:41:44 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
f6124370b0 chore(deps): bump pillow (#4547)
Bumps [pillow](https://github.com/python-pillow/Pillow) from 12.1.1 to 12.3.0.
- [Release notes](https://github.com/python-pillow/Pillow/releases)
- [Changelog](https://github.com/python-pillow/Pillow/blob/main/CHANGES.rst)
- [Commits](https://github.com/python-pillow/Pillow/compare/12.1.1...12.3.0)

---
updated-dependencies:
- dependency-name: pillow
  dependency-version: 12.3.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-14 18:40:42 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
b8822f5008 chore(deps): bump pyasn1 (#4549)
Bumps [pyasn1](https://github.com/pyasn1/pyasn1) from 0.6.3 to 0.6.4.
- [Release notes](https://github.com/pyasn1/pyasn1/releases)
- [Changelog](https://github.com/pyasn1/pyasn1/blob/main/CHANGES.rst)
- [Commits](https://github.com/pyasn1/pyasn1/compare/v0.6.3...v0.6.4)

---
updated-dependencies:
- dependency-name: pyasn1
  dependency-version: 0.6.4
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-14 18:39:56 +00:00
Dustin LuongandCopybara-Service 8976c57b9c Update the Kimi-K3 deployment notebook image URI.
PiperOrigin-RevId: 964692807
2026-08-14 07:39:13 -07:00
gmaninatarajanandGitHub 3985da440e fix: Updated new whl file with SDK update to add interval_variants parameter to score_ism_variants() (#4565)
* fix: Updated new whl file with SDK update to add interval_variants parameter to score_ism_variants()

* fix: updating the whl file download cell
2026-08-11 19:56:05 -04:00
Damodar PanigrahiGitHubgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
1c9092ced3 refactor - restructure the notebook (#4564)
* refactor - restructure the notebook

* Update notebooks/community/weathernext/CUSTOM_INPUTS_GUIDE.md

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update notebooks/community/weathernext/weathernext_2_ic_pc.ipynb

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update notebooks/community/weathernext/weathernext_2_dws.ipynb

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

---------

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-08-07 20:53:54 +00:00
Damodar PanigrahiandGitHub 77b2af09ce feat: WN2 with GPU GA (#4563) 2026-08-07 17:27:35 +00:00
genquan9andGitHub c6d33c2a0d Add tau2-bench RL blog post to docs README (#4561) 2026-08-05 23:18:11 +00:00
genquan9andGitHub 89772b320e Fix inline math rendering: use span+1798467 for GitHub Pages MathJax (#4560) 2026-08-05 17:47:23 +00:00
genquan9andGitHub a1f6d2c069 Fix LaTeX rendering for Pass Rate formula (#4559)
Replace underscores in \text{num\_pass} with spaces to avoid
LaTeX math mode errors on GitHub rendering.
2026-08-05 17:30:29 +00:00
genquan9andGitHub 936a6adf77 Add multi-turn RL for tau2-bench technical report (#4558)
* Add multi-turn RL for tau2-bench technical report

Add technical report documenting multi-turn reinforcement learning
training pipeline for tau2-bench customer service benchmark, including
GRPO training, data synthesis pipeline, and evaluation results.

* Fix deprecated MathJax CDN and broken anchor link

- Remove deprecated cdn.mathjax.org script tag (GitHub renders LaTeX natively)
- Fix broken ToC anchor from #2-bench to #tau2-bench
2026-08-05 16:25:05 +00:00
Dustin LuongandCopybara-Service 37a85d53f4 No public description
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 958290655
2026-08-03 04:07:45 -07:00
Dustin LuongandCopybara-Service 0b0e362ab9 Add Kimi K3 Model Garden deployment notebook
PiperOrigin-RevId: 958290655
2026-08-03 04:06:44 -07:00
Damodar PanigrahiandGitHub 98103d462f test (#4554) 2026-07-30 23:37:05 +00:00
Oleh PrypinandCopybara-Service 9ea1cf3b86 No public description
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 955252640
2026-07-28 07:46:15 -07:00
Sam-DecigaandGitHub 8f3e6668e1 feat: Claude Opus 5 Launch (#4552) 2026-07-26 10:04:43 -04:00
Tianzi CaiandGitHub 5d9853db5c Fix formatting in Anthropic Claude intro notebook 2026-07-22 20:59:25 -07:00
Tianzi CaiandGitHub 003fb5121b Remove unused httpx imports and related comments 2026-07-22 20:56:48 -07:00
Tianzi CaiandGitHub a62695fb38 Update image URL and request handling in notebook (#4551)
* Update image URL and request handling in notebook

* Remove Colab link markdown cell

Removed markdown cell with Colab link from the notebook.

* Remove unused import
2026-07-22 23:52:09 +00:00
Sam-DecigaandGitHub 3c630fdbb8 feat: Claude-Sonnet5-Launch (#4537) 2026-06-30 16:31:27 -04:00
Damodar PanigrahiandGitHub 6ca1d899d6 feat: wn2 doc polished (#4531) 2026-06-23 23:22:47 +00:00
Damodar PanigrahiandGitHub 1894602fff feat: wn2 notebook (#4530) 2026-06-23 21:59:17 +00:00
Damodar PanigrahiandGitHub 31a52d6e92 feat: WeatherNext IC (#4523)
* feat: WeatherNext IC

* Fix: Replace weathernext_2_ic_early_access_program.ipynb symlink with actual notebook file

* fix: Replace Vertex Jobs with Gemini Enterprise Agent Platform Jobs in WeatherNext notebook

* fix: Correct typos, broken links, and apply linter formatting
2026-06-11 19:49:13 +00:00
0f9d9734c3 feat: Claude Fable 5 Launch (#4522)
Co-authored-by: Holt Skinner <13262395+holtskinner@users.noreply.github.com>
2026-06-09 14:47:46 -04:00
Vertex MG TeamandCopybara-Service e85cf9a174 Update link to Cloud Quotas page to correct location
PiperOrigin-RevId: 926490177
2026-06-03 23:21:57 -07:00
Sam-DecigaandGitHub b4c0bbc1a0 feat: Ant-Opus4.8 Launch (#4520) 2026-05-28 14:43:47 -04:00
Rayan DasoriyaandCopybara-Service 24244351cd Add a new notebook for OSS distillation feasibility study.
PiperOrigin-RevId: 917879385
2026-05-19 09:37:47 -07:00
Vertex MG TeamandCopybara-Service bf0e1300a9 No public description
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 878451476
2026-05-13 12:58:56 -07:00
chnduandGitHub cf048b6fe4 Add live_api skills that help the user build their own liveapi service (#4511)
* Add live_api skills that help the user to build their own liveapi service.

Implementation are based on websocket. Support different coding languages.

* Update based on review

* Fix typos

* Update vertex to gemini enterprise.
2026-05-11 17:19:12 +00:00
Mend RenovateandGitHub 8c8820ecfa chore(deps): update dependency numpy to v2.4.4 (#4466) 2026-05-06 14:24:57 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
a1a52d8145 chore(deps): bump requests (#4488)
Bumps [requests](https://github.com/psf/requests) from 2.32.4 to 2.33.0.
- [Release notes](https://github.com/psf/requests/releases)
- [Changelog](https://github.com/psf/requests/blob/main/HISTORY.md)
- [Commits](https://github.com/psf/requests/compare/v2.32.4...v2.33.0)

---
updated-dependencies:
- dependency-name: requests
  dependency-version: 2.33.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-06 14:22:23 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
a06ce545e7 chore(deps): bump pillow (#4497)
Bumps [pillow](https://github.com/python-pillow/Pillow) from 12.1.1 to 12.2.0.
- [Release notes](https://github.com/python-pillow/Pillow/releases)
- [Changelog](https://github.com/python-pillow/Pillow/blob/main/CHANGES.rst)
- [Commits](https://github.com/python-pillow/Pillow/compare/12.1.1...12.2.0)

---
updated-dependencies:
- dependency-name: pillow
  dependency-version: 12.2.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-06 14:19:28 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
913780c4cb chore(deps): bump pillow (#4509)
Bumps [pillow](https://github.com/python-pillow/Pillow) from 10.3.0 to 12.2.0.
- [Release notes](https://github.com/python-pillow/Pillow/releases)
- [Changelog](https://github.com/python-pillow/Pillow/blob/main/CHANGES.rst)
- [Commits](https://github.com/python-pillow/Pillow/compare/10.3.0...12.2.0)

---
updated-dependencies:
- dependency-name: pillow
  dependency-version: 12.2.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-06 14:14:14 +00:00
gmaninatarajanandGitHub 71be46e7d8 feat: Updated whl file and package name as part of Vertex Model Garden setup (#4508) 2026-04-30 08:28:04 -04:00
Mayank SharanandGitHub 849e88a627 Vtc blog 2 (#4507)
* Adding reviewed version of VTC blog 2

* GCA suggested fixes

* Updating readme to have links
2026-04-29 19:10:45 +00:00
Jason DaiandGitHub daf56bcd0b Create Eval Quality Flywheel Skill for preview (#4505) 2026-04-23 16:37:08 +00:00
Mayank SharanandGitHub 563f423b93 Adding reviewed version of VTC blog 2 (#4502)
* Adding reviewed version of VTC blog 2

* GCA suggested fixes
2026-04-16 21:12:03 +00:00
ian1780andGitHub 292e540e96 Update anthropic_claude_intro.ipynb (#4501)
add opus 4.7 multi region endpoint support b/491171457
2026-04-16 17:59:12 +00:00
Sam-DecigaandGitHub 6c6a703c5a Ant nickel (#4500)
* Feat: Anthropic Opus-4-7 launch

* Feat: Anthropic Opus-4-7 launch
2026-04-16 12:18:47 -04:00
Sam-DecigaandGitHub 7ef83c6f73 MARS8 new Asian regions (#4498) 2026-04-14 19:59:45 +00:00
Vertex MG TeamandCopybara-Service aba6598109 use old docker hash for whisper model deployment.
PiperOrigin-RevId: 892120947
2026-03-30 23:08:58 -07:00
Sam-DecigaandGitHub 88a6b8037e Refactor: Anthropic NB (#4489) 2026-03-26 14:49:59 -04:00
Eric DongandGitHub 8845f7ab27 Update Gemini model references and availability details
Update Gemini versions.
2026-03-26 09:40:01 -04:00
Eric DongandGitHub 3b2e711a16 Update fine-tuning model reference in README
Updated the fine-tuning model reference from Gemini 1.5 Pro to Gemini 2.5 Pro in the README.
2026-03-25 16:44:52 -04:00
Eric DongandGitHub 5c0629cdc7 Revise README for Agent Skills in Vertex AI
Updated terminology and formatting for clarity.
2026-03-25 15:18:03 -04:00
Eric DongandGitHub e107d30807 chore: Add detailed installation instructions for skills (#4487) 2026-03-25 15:06:59 -04:00
Eric DongandGitHub b98ab36913 refactor: Add tool configuation in skills readme (#4486)
* chore: Update vertex-ai Skills readme

* refactor: Add tool configuation in  skills readme
2026-03-25 13:41:45 -04:00
Eric DongandGitHub f1d90b5a71 chore: Update vertex-ai Skills readme (#4485) 2026-03-25 11:36:41 -04:00
Eric DongandGitHub f848db6132 Revise README title and formatting for emphasis
Updated the title and emphasized 'Skills' in the README.
2026-03-25 10:16:42 -04:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
cc9fffd945 chore(deps): bump pillow (#4482)
Bumps [pillow](https://github.com/python-pillow/Pillow) from 10.3.0 to 12.1.1.
- [Release notes](https://github.com/python-pillow/Pillow/releases)
- [Changelog](https://github.com/python-pillow/Pillow/blob/main/CHANGES.rst)
- [Commits](https://github.com/python-pillow/Pillow/compare/10.3.0...12.1.1)

---
updated-dependencies:
- dependency-name: pillow
  dependency-version: 12.1.1
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-03-24 15:13:00 +00:00
Eric DongandGitHub 7606a1de03 chore: Update the skills readme with instructions (#4484) 2026-03-24 10:44:28 -04:00
Eric DongandGitHub cca59aa753 Update README.md
Remove icons
2026-03-24 10:12:04 -04:00
Eric DongandGitHub a1907da27a chore: Update readme (#4483) 2026-03-24 10:00:50 -04:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
e8cb7738d0 Bump pillow (#4443)
Bumps [pillow](https://github.com/python-pillow/Pillow) from 10.3.0 to 12.1.1.
- [Release notes](https://github.com/python-pillow/Pillow/releases)
- [Changelog](https://github.com/python-pillow/Pillow/blob/main/CHANGES.rst)
- [Commits](https://github.com/python-pillow/Pillow/compare/10.3.0...12.1.1)

---
updated-dependencies:
- dependency-name: pillow
  dependency-version: 12.1.1
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-03-24 13:52:57 +00:00
Mend RenovateandGitHub 18e8d603de chore(deps): update dependency black to v26.3.1 [security] (#4470) 2026-03-24 13:52:02 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
1bd5901fdb Bump black (#4468)
Bumps [black](https://github.com/psf/black) from 25.1.0 to 26.3.1.
- [Release notes](https://github.com/psf/black/releases)
- [Changelog](https://github.com/psf/black/blob/main/CHANGES.md)
- [Commits](https://github.com/psf/black/compare/25.1.0...26.3.1)

---
updated-dependencies:
- dependency-name: black
  dependency-version: 26.3.1
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-03-24 13:51:38 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
5b245024cd Bump pyasn1 (#4474)
Bumps [pyasn1](https://github.com/pyasn1/pyasn1) from 0.6.2 to 0.6.3.
- [Release notes](https://github.com/pyasn1/pyasn1/releases)
- [Changelog](https://github.com/pyasn1/pyasn1/blob/main/CHANGES.rst)
- [Commits](https://github.com/pyasn1/pyasn1/compare/v0.6.2...v0.6.3)

---
updated-dependencies:
- dependency-name: pyasn1
  dependency-version: 0.6.3
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-03-24 13:50:31 +00:00
Eric DongandGitHub ba043c196c chore: Update skills readme with architecture (#4481) 2026-03-24 09:49:35 -04:00
Eric DongandGitHub 28ce8f6d7a chore: Update readme and template (#4480) 2026-03-24 09:42:45 -04:00
Eric DongandGitHub 3b5a8cad41 feat: Add Gen AI SDK skill for Vertex (#4479) 2026-03-24 09:29:01 -04:00
gmaninatarajanandGitHub c6d7971bc9 fix:Simplified authentication section and addressed timeout issues (#4478) 2026-03-23 18:50:35 -04:00
Eric DongandGitHub dbe28965cb feat: Add primary routing and readme for vertex ai skills (#4477) 2026-03-23 17:18:16 -04:00
Vertex MG TeamandCopybara-Service 0d34d6bbea update minimax m2 notebook.
PiperOrigin-RevId: 886885747
2026-03-20 11:14:55 -07:00
Sam-DecigaandGitHub a1d898f35e feat: Jina EmbV3 launch (#4476)
* feat: Jina EmbV3 launch

* feat: Jina EmbV3 launch
2026-03-20 08:19:38 -04:00
Eric DongandGitHub 772ee71bc3 feat: use Vertex AI MCP server (#4475)
* feat: use Vertex AI MCP server

* Address review comments
2026-03-19 11:02:28 -04:00
Sam-DecigaandGitHub 425851cedc feat: Nemotron3-Super model launch (#4473)
* feat: Nemotron3-Super model launch

* feat: Nemotron3-Super model launch

* feat: Nemotron3-Super model launch
2026-03-16 20:41:36 -04:00
Lav RaiandGitHub bcccbee164 Update distillation report. (#4472) 2026-03-16 18:40:10 +00:00
Lav RaiandGitHub 5ae325528a Add distillation report. (#4471) 2026-03-13 15:33:34 +00:00
vincentkt-googleandGitHub 86674effee Add and update existing vertex skills (#4467)
* Add and update existing vertex skills

- Add support for fine tuning for 1p gemini tuning
- Add support for deploying fine tuned model support
- Add support for running inference on MaaS models
- Add open model support for regions and cost estimating for 3p tuning

* fixing some of the commit errors

* updated scripts to use existing gemini 1.5 pro model

* swap gemini 1.5 pro to gemini 2.5 pro
2026-03-11 19:45:42 +00:00
vincentkt-googleandGitHub 8b4708c606 feat: add vertex ai skills to repo (#4454) 2026-03-05 17:50:41 +00:00
Yichen ZhouandCopybara-Service f3dd6cbca3 Update TimesFM-2.5 notebook for Model Garden.
PiperOrigin-RevId: 878726929
2026-03-04 16:42:55 -08:00
Rayan DasoriyaandCopybara-Service 1f9e93993c No public description
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 878215121
2026-03-03 18:26:58 -08:00
Vertex MG TeamandCopybara-Service 062835174e Updated the image default TAG to release
PiperOrigin-RevId: 877893209
2026-03-03 05:26:21 -08:00
Rayan DasoriyaandCopybara-Service cb4916f590 No public description
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 875506020
2026-02-25 21:44:41 -08:00
Damodar PanigrahiandGitHub 2933fe606b bug: remove A100 as recommended specs (#4450) 2026-02-25 14:05:36 +00:00
Damodar PanigrahiandGitHub b468809df7 bug: ahref update (#4449)
* bug: ahref update

* fix: linter

* fix: typo fix
2026-02-25 13:46:19 +00:00
Sam-DecigaandGitHub 0417d8b9c4 feat: Deprecate Claude 3 Haiku (#4448)
Deprecation start date: Feb. 23, 2026
End of Support date: Aug. 23, 2026
b/485993204
2026-02-23 20:57:43 -05:00
Damodar PanigrahiandGitHub 7750e83fbb fix: inference key change, finetuning jaxlib update (#4447) 2026-02-23 19:45:59 +00:00
Damodar PanigrahiandGitHub 649800e646 feat: Alphagenome finetuning notebook (#4445)
* feat: Alphagenome finetuning notebook

* Update cloudai_alphagenome_finetune.ipynb

Fixed the lint errors

* Update cloudai_alphagenome_finetune.ipynb

Fix lint errors

* feat: Add Alphagenome finetune

* feat: Include the Alphagenome finetuning notebook url in the readme. Update the codeowners

* feat: Add alphagegenome finetune notebook to the readme, add user to codeowners

* feat: fix spelling
2026-02-20 14:13:36 +00:00
Mend RenovateandGitHub 42b35056fa chore(deps): update dependency black to v26 (#4422) 2026-02-18 15:07:12 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
24974eda95 Bump protobuf (#4437)
Bumps [protobuf](https://github.com/protocolbuffers/protobuf) from 4.25.8 to 5.29.6.
- [Release notes](https://github.com/protocolbuffers/protobuf/releases)
- [Commits](https://github.com/protocolbuffers/protobuf/commits)

---
updated-dependencies:
- dependency-name: protobuf
  dependency-version: 5.29.6
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-02-18 15:06:37 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
ad41377783 Bump protobuf (#4438)
Bumps [protobuf](https://github.com/protocolbuffers/protobuf) from 4.25.8 to 5.29.6.
- [Release notes](https://github.com/protocolbuffers/protobuf/releases)
- [Commits](https://github.com/protocolbuffers/protobuf/commits)

---
updated-dependencies:
- dependency-name: protobuf
  dependency-version: 5.29.6
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-02-18 15:06:10 +00:00
Sam-DecigaandGitHub 36dea3ca01 feat: Anthropic Sonnet-4-6 launch (#4442)
Signed-off-by: Sam-Deciga <decigagarcia@google.com>
2026-02-17 14:02:30 -05:00
b75b2ea4d7 fix: updated with latest whl file version : alphagenome-0.4.2.6-py3-none-any.whl (#4439)
Co-authored-by: hyper-param <peeyusht@google.com>
2026-02-10 18:27:03 +00:00
Vertex MG TeamandCopybara-Service 28f7fc4445 Add SAM 3 notebook to Vertex AI Model Garden.
PiperOrigin-RevId: 866598579
2026-02-06 13:45:47 -08:00
Rayan DasoriyaandCopybara-Service bf2c1226fd Update the license year
PiperOrigin-RevId: 866214899
2026-02-05 19:07:14 -08:00
Vertex MG TeamandCopybara-Service 2990c53292 Added notebook sample for batch inference using the remote sensing VMG models
PiperOrigin-RevId: 866061602
2026-02-05 12:27:10 -08:00
Sam-DecigaandGitHub 531d9cfee0 feat: New Anthropic model (#4436) 2026-02-05 14:07:05 -05:00
Vertex MG TeamandCopybara-Service ff18ec7af5 Add --total-gpus to multi-model model-cohost deployment config.
PiperOrigin-RevId: 863070590
2026-01-29 22:39:06 -08:00
Sam-DecigaandGitHub 008eb409ef feat: NVIDIA-Llama-Nemotron-Super-49B (#4430) 2026-01-29 08:51:52 -05:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
b2dba4b568 Bump pyasn1 (#4421)
Bumps [pyasn1](https://github.com/pyasn1/pyasn1) from 0.6.1 to 0.6.2.
- [Release notes](https://github.com/pyasn1/pyasn1/releases)
- [Changelog](https://github.com/pyasn1/pyasn1/blob/main/CHANGES.rst)
- [Commits](https://github.com/pyasn1/pyasn1/compare/v0.6.1...v0.6.2)

---
updated-dependencies:
- dependency-name: pyasn1
  dependency-version: 0.6.2
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-01-27 20:14:34 +00:00
Sam-DecigaandGitHub 665547f790 feat:New MARS8 model (#4429)
* feat:New MARS8 model

* feat:New MARS8 model
2026-01-27 10:43:43 -05:00
Vertex MG TeamandCopybara-Service 5078c44eb8 Updated the bucket path for the remote sensing models.
PiperOrigin-RevId: 861234161
2026-01-26 09:47:01 -08:00
Bhaskar GoyalandGitHub da6e46531e feature: Retire Mistral 24.11 and Codestral 25.01 from Mistral Intro files. (#4427) 2026-01-23 18:47:36 +00:00
0a4091a3b1 fix: Fix json response parsing error (#4426)
Co-authored-by: hyper-param <peeyusht@google.com>
2026-01-23 13:58:03 +00:00
Sam-DecigaandGitHub 5afa83dd25 feat: MongoDB voyage-4 launch (#4419)
* feat: MongoDB voyage-4 launch

* feat: MongoDB voyage-4 launch

* feat: MongoDB voyage-4 launch

* feat: MongoDB voyage-4 launch
2026-01-16 16:26:05 -05:00
Sam-DecigaandGitHub 8cab85d6ad feat: MongoDB voyage-multimodal-3.5 launch (#4420)
* feat: MongoDB voyage-multimodal-3.5 launch

* feat: MongoDB voyage-multimodal-3.5 launch
2026-01-16 14:57:11 -05:00
Eric DongandGitHub a7f3940635 refactor: Update Github icon (#4418) 2026-01-13 13:36:20 -05:00
Mend RenovateandGitHub 5bb1a75a48 chore(deps): update dependency pyupgrade to v3.21.2 (#4359) 2026-01-13 18:26:35 +00:00
Damodar PanigrahiandGitHub 8fe4985aa8 feat: WeatherNext2 Initial Updates (#4417) 2026-01-13 13:24:48 -05:00
intentsolutions.ioandGitHub 996b6534d9 Add ADK inline source deployment tutorial for Agent Engine (#4393)
* Add ADK inline source deployment tutorial for Agent Engine

* fix: address Gemini review feedback

- Change model from gemini-2.0-flash to gemini-1.5-flash-001
- Improve exception handling with ZoneInfoNotFoundError

* fix: address Gemini code review feedback

- Use specific ZoneInfoNotFoundError exception instead of generic Exception
- Define REQUIREMENTS variable once and reuse to avoid duplication
- Keep generic Exception as fallback for unexpected errors

🤖 Generated with [Claude Code](https://claude.com/claude-code)

* style: fix notebook formatting via official linter
2026-01-13 08:56:07 -05:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
1061ae5348 Bump urllib3 (#4416)
Bumps [urllib3](https://github.com/urllib3/urllib3) from 2.6.0 to 2.6.3.
- [Release notes](https://github.com/urllib3/urllib3/releases)
- [Changelog](https://github.com/urllib3/urllib3/blob/main/CHANGES.rst)
- [Commits](https://github.com/urllib3/urllib3/compare/2.6.0...2.6.3)

---
updated-dependencies:
- dependency-name: urllib3
  dependency-version: 2.6.3
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-01-12 20:01:22 +00:00
Eric DongandGitHub 4fac3a630f refactor: reformat notebook template (#4414)
* refactor: reformat notebook template

* Update Python version to 3.12

* Fix an import

* Update licience year
2026-01-12 13:15:30 -05:00
Eric DongandGitHub 1f2db2903c refactor: Updated Github icons (#4413) 2026-01-12 10:10:34 -05:00
Sam-DecigaandGitHub 912a52de70 feat:MongoDB Voyage 3.5-Lite (#4403)
* feat:MongoDB Voyage 3.5-Lite

* chore:apply linter

* Update voyage-3.5-lite.ipynb

Fixing MODEL_NAME
2026-01-06 21:19:17 -05:00
Sam-DecigaandGitHub 5e29090e86 Model Deprecation (#4409)
Haiku 3.5 Model Deprecation
2026-01-05 16:30:23 -05:00
Mend RenovateandGitHub cd8fcd1839 Update actions/checkout action to v6 (#4374) 2026-01-05 14:20:36 +00:00
Mend RenovateandGitHub e51075ec4b chore(deps): update dependency black to v25.12.0 (#4360) 2026-01-05 14:18:08 +00:00
Mend RenovateandGitHub 9709c0dddb chore(deps): update dependency isort to v7 (#4289) 2026-01-05 14:14:26 +00:00
Mend RenovateandGitHub a21ae41762 Update dependency python to 3.14 (#4398) 2026-01-05 14:13:11 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
ff5ff7b609 Bump urllib3 (#4386)
Bumps [urllib3](https://github.com/urllib3/urllib3) from 2.5.0 to 2.6.0.
- [Release notes](https://github.com/urllib3/urllib3/releases)
- [Changelog](https://github.com/urllib3/urllib3/blob/main/CHANGES.rst)
- [Commits](https://github.com/urllib3/urllib3/compare/2.5.0...2.6.0)

---
updated-dependencies:
- dependency-name: urllib3
  dependency-version: 2.6.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-01-05 14:12:43 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
23748f443e Bump urllib3 (#4387)
Bumps [urllib3](https://github.com/urllib3/urllib3) from 2.5.0 to 2.6.0.
- [Release notes](https://github.com/urllib3/urllib3/releases)
- [Changelog](https://github.com/urllib3/urllib3/blob/main/CHANGES.rst)
- [Commits](https://github.com/urllib3/urllib3/compare/2.5.0...2.6.0)

---
updated-dependencies:
- dependency-name: urllib3
  dependency-version: 2.6.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-01-05 14:12:18 +00:00
Vertex MG TeamandCopybara-Service fc6b2167de ComfyUI tutorial notebook
PiperOrigin-RevId: 851380881
2026-01-02 10:20:16 -08:00
e79a45358c Migrate gsutil usage to gcloud storage (#4299)
* Migrate gsutil usage to gcloud storage

* changes for 4299

* Apply automated linter fixes

* update

* remove model_garden changes

* revert to main

* revert model garden file

* Update model_garden_weather_prediction_on_vertex.ipynb

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-31 08:37:14 -05:00
20d19fb11c Migrate gsutil usage to gcloud storage (#4331)
* Migrate gsutil usage to gcloud storage

* Manual Changes

* Changes for 4331

* Changes for 4331

* Removed changes model garden

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-31 08:36:46 -05:00
9c9f7a6e2a Migrate gsutil usage to gcloud storage (#4334)
* Migrate gsutil usage to gcloud storage

* Manual Changes

* changes for 4334

* fix linter issue\ for 4334

* manual changes

* Restore gcloud migration code

* Remove changes for model garden

* update

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-31 08:36:21 -05:00
ede41c2115 Migrate gsutil usage to gcloud storage (#4335)
* Migrate gsutil usage to gcloud storage

* remoed changes for model garden

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-31 08:35:53 -05:00
ca53786c04 Migrate gsutil usage to gcloud storage (#4322)
* Migrate gsutil usage to gcloud storage

* Changes for 4322

* fix linter issue for 4322

* removed changes for model garden

* Update model_garden_pytorch_gemma_peft_finetuning_hf.ipynb

* Update model_garden_pytorch_gemma_peft_finetuning_hf.ipynb

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-25 10:49:46 -05:00
419f8310c9 Migrate gsutil usage to gcloud storage (#4316)
* Migrate gsutil usage to gcloud storage

* Changes for 4316

* Changes for 4316

* fix linter issue for 4316

* removed changes for model garden

* removed changes for model garden

* Update model_garden_tfvision_image_classification.ipynb

* Update model_garden_tfvision_image_classification.ipynb

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-25 10:49:33 -05:00
996b690e03 Migrate gsutil usage to gcloud storage (#4317)
* Migrate gsutil usage to gcloud storage

* Changes for 4317

* Changes for 4317

* fix linter issue for 4317

* removed changes for model garden

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-25 15:48:30 +00:00
5b9d04d63a Migrate gsutil usage to gcloud storage (#4318)
* Migrate gsutil usage to gcloud storage

* Changes for 4318

* fix linter issue for 4318

* removed changes for model garden:

* Update model_garden_gemma2_deployment_on_vertex.ipynb

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-25 10:47:49 -05:00
9ed3c2f83d Migrate gsutil usage to gcloud storage (#4320)
* Migrate gsutil usage to gcloud storage

* Changes for 4320

* fix linter issue for 4320

* removed chnages for model garden

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-25 10:47:17 -05:00
5fc93c8bdb Migrate gsutil usage to gcloud storage (#4328)
* Migrate gsutil usage to gcloud storage

* PR changes for 4328

* Fix Linter issue for 4328

* changes removed model garden

* Update model_garden_jax_fvlm.ipynb

* Update model_garden_jax_fvlm.ipynb

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-25 10:46:52 -05:00
3cf46226e9 Migrate gsutil usage to gcloud storage (#4321)
* Migrate gsutil usage to gcloud storage

* Changes for 4321

* fix the linter issue for br 4321

* removed changes for model garden

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-25 10:46:04 -05:00
439f6a0cae Migrate gsutil usage to gcloud storage (#4329)
* Migrate gsutil usage to gcloud storage

* Manual Changes

* Manual Changes

* Linter fiex the issues for 4329

* Revert "Linter fiex the issues for 4329"

This reverts commit a1ed9c6f59.

* Revert "Manual Changes"

This reverts commit 8ca9b56c5b.

* changes removed model garden

* Update model_garden_gemma2_finetuning_on_vertex.ipynb

* Update model_garden_pytorch_llama3_3_finetuning.ipynb

* Update model_garden_pytorch_llama3_3_finetuning.ipynb

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-25 10:45:30 -05:00
993898bb71 Migrate gsutil usage to gcloud storage (#4330)
* Migrate gsutil usage to gcloud storage

* Manual Changes

* Removed blank line and spaces Manual Changes

* fix the Linter issue for 4330

* Changes for model garden

* Update model_garden_pytorch_llama3_1_finetuning.ipynb

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-25 15:45:00 +00:00
9e9e639375 Migrate gsutil usage to gcloud storage (#4323)
* Migrate gsutil usage to gcloud storage

* Manual Changes-Migrate gsutil usage to gcloud storage

* Manual Changes

* Fix linter issue for 4323

* removed changes for model garden

* Update model_garden_axolotl_gpt_oss_finetuning.ipynb

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-25 10:44:34 -05:00
633cf6a799 Migrate gsutil usage to gcloud storage (#4327)
* Migrate gsutil usage to gcloud storage

* Manual Changes-Migrate gsutil usage to gcloud storage

* Changes for 4327

* fixing linting erro

* removed changes of model garden

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-25 10:44:05 -05:00
8b618bc455 Migrate gsutil usage to gcloud storage (#4326)
* Migrate gsutil usage to gcloud storage

* Manual Changes-Updated the cell by replacing 'gsutil copy' with the correct 'gcloud storage cp'

* Manual Changes-Updated the cell by replacing 'gsutil copy' with the correct 'gcloud storage cp'

* Revert "Manual Changes-Updated the cell by replacing 'gsutil copy' with the correct 'gcloud storage cp'"

This reverts commit 175eaa4fe8.

* Manual Changes-Updated the cell by replacing 'gsutil copy' with the correct 'gcloud storage cp'

* Changes for 4326

* Changes for 4326

* Linter fix issue for 4326

* removed model garden changes

* Update model_garden_movinet_action_recognition.ipynb

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-25 10:43:31 -05:00
6bbe3bcfe0 Migrate gsutil usage to gcloud storage (#4336)
* Migrate gsutil usage to gcloud storage

* Manual Changes

* removed model garden changes

* Update model_garden_llama3_1_finetuning_with_workbench.ipynb

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-25 10:42:44 -05:00
Margubur RahmanandGitHub 814827ac19 Migrate gsutil usage to gcloud storage (#4332) 2025-12-25 10:42:12 -05:00
Margubur RahmanandGitHub 7a613785b9 Migrate gsutil usage to gcloud storage (#4333) 2025-12-25 15:41:38 +00:00
100243e90a Migrate gsutil usage to gcloud storage (#4338)
* Migrate gsutil usage to gcloud storage

* Manual Changes

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
2025-12-25 15:40:46 +00:00
6424515b03 Migrate gsutil usage to gcloud storage (#4339)
* Migrate gsutil usage to gcloud storage

* Manual Changes

* Manual Changes

* Fix: Updated gcloud storage command without formatting

* Manual Changes

* Manual Changes

* Revert "Manual Changes"

This reverts commit a7a7bda0f9.

* Manual Changes

* Revert "Manual Changes"

This reverts commit a7a7bda0f9.

* Manual Changes

* Revert "Manual Changes"

This reverts commit 71c777d5f1.

* Manual Changes

* Manual changes

* Changes for 4339

* Changes for 4339

* Changes for 4339

* Fix: Applied linter formatting and resolved style issues

* gcloud to gsutilchanges for  4339

* removed gsutil to gcloud migration

* Manual changes

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-25 15:40:00 +00:00
8d22b221b4 Migrate gsutil usage to gcloud storage (#4337)
* Migrate gsutil usage to gcloud storage

* Manual Changes

* Manual Changes

* Manual Changes

* Manual Changes

* Revert "Manual Changes"

This reverts commit 3ef23f1456.

* Manaul Changes

* Changes for 4337

* Fix: Resolved linter errors and formatted notebooks

* Revert "Fix: Resolved linter errors and formatted notebooks"

This reverts commit 7534760fe1.

* Revert "Changes for 4337"

This reverts commit f497568138.

* Changes for 4337

* gcloud to gsutil migration

* removed changes for model garden

* Update model_garden_axolotl_qwen3_finetuning.ipynb

* Update model_garden_axolotl_qwen3_finetuning.ipynb

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-25 10:39:31 -05:00
2a8ad7cdbb Migrate gsutil usage to gcloud storage (#4319)
* Migrate gsutil usage to gcloud storage

* changes for 4319

* changes for 4319

* Apply automated linter fixes

* remove unused import

* removed model_garden changes

* Update model_garden_pytorch_mixtral_peft_tuning.ipynb

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-24 09:41:58 -05:00
3239b301f2 Migrate gsutil usage to gcloud storage (#4315)
* Migrate gsutil usage to gcloud storage

* changes for 4315

* Apply automated linter fixes

* removed model_garden folder changes

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-24 09:41:28 -05:00
acb10d14b8 Migrate gsutil usage to gcloud storage (#4314)
* Migrate gsutil usage to gcloud storage

* changes for 4314

* Apply automated linter fixes

* remove unused imports

* removed model_garden changes

* updates

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-24 09:41:00 -05:00
778d145970 Migrate gsutil usage to gcloud storage (#4313)
* Migrate gsutil usage to gcloud storage

* changes for 4313

* Apply automated linter fixes

* remove model_garden updates

* remove model_garden updates

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-24 09:40:34 -05:00
4d00356f4b Migrate gsutil usage to gcloud storage (#4312)
* Migrate gsutil usage to gcloud storage

* changes for 4312

* Apply automated linter fixes

* remove model_garden folder changes

* remove model_garden folder changes

* Update model_garden_pytorch_llama2_peft_hyperparameter_tuning.ipynb

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-24 09:40:01 -05:00
0654305994 Migrate gsutil usage to gcloud storage (#4311)
* Migrate gsutil usage to gcloud storage

* changes for 4311

* Apply automated linter fixes

* removed unused imports

* remove model_garden changes

* Update model_garden_pytorch_zipnerf.ipynb

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-24 09:39:27 -05:00
c53f392c5a Migrate gsutil usage to gcloud storage (#4310)
* Migrate gsutil usage to gcloud storage

* changes for 4310

* Apply automated linter fixes

* removed model_garden folder changes

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-24 09:38:53 -05:00
f4b56e92ae Migrate gsutil usage to gcloud storage (#4308)
* Migrate gsutil usage to gcloud storage

* changes for 4308

* added = in command

* Apply automated linter fixes

* removed model_garden folder changes

* removed model_garden folder changes

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-24 09:38:24 -05:00
Margubur RahmanGitHubgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gurusai-voletigemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
45ec1cf18a Migrate gsutil usage to gcloud storage (#4307)
* Migrate gsutil usage to gcloud storage

* changes for 4307

* Apply suggestion from @gemini-code-assist[bot]

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Apply suggestion from @gemini-code-assist[bot]

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Apply automated linter fixes

* Update sdk_vector_search_for_indexing.ipynb

* Update sdk_vector_search_for_indexing.ipynb

* update

* Update model_garden_pytorch_falcon_instruct_quantization.ipynb

* Update model_garden_pipeline_templates_t5x.ipynb

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-12-24 14:37:45 +00:00
f6c8bcf937 Migrate gsutil usage to gcloud storage (#4301)
* Migrate gsutil usage to gcloud storage

* changes for 4301

* linter changes

* Revert "linter changes"

This reverts commit 6665133b2c.

* Apply automated linter fixes

* Update training-multi-class-classification-model-for-ads-targeting-usecase.ipynb

* Update training-multi-class-classification-model-for-ads-targeting-usecase.ipynb

* removed model_garden folder changes

* Update model_garden_mediapipe_object_detection.ipynb

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-24 09:37:11 -05:00
Rayan DasoriyaandCopybara-Service 9e590d5a9f Update notebooks based on latest deployment options
PiperOrigin-RevId: 847555297
2025-12-21 19:32:54 -08:00
46e0ea4f1c Migrate gsutil usage to gcloud storage (#4309)
* Migrate gsutil usage to gcloud storage

* changes for 4309

* changes for 4309

* Apply automated linter fixes

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-20 11:34:42 -05:00
300fce6b9f Migrate gsutil usage to gcloud storage (#4305)
* Migrate gsutil usage to gcloud storage

* Apply automated linter fixes

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-20 11:34:10 -05:00
Vertex MG TeamandCopybara-Service 52e3066c38 No public description
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 845978812
2025-12-19 10:42:55 -08:00
27ebf52198 Migrate gsutil usage to gcloud storage (#4296)
* Migrate gsutil usage to gcloud storage

* updates

* update

* Update llm_streaming_prediction.ipynb

* revert change

* removed model_garden folder changes

* removed model_garden folder changes

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-19 11:57:20 -05:00
ca7d4e153e Migrate gsutil usage to gcloud storage (#4297)
* Migrate gsutil usage to gcloud storage

* changes for 4297

* Apply automated linter fixes

* removed changes from model_garden

* removed changes from model_garden

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-19 11:56:49 -05:00
5b6c766629 Migrate gsutil usage to gcloud storage (#4298)
* Migrate gsutil usage to gcloud storage

* update

* linter changes for 4298

* Revert "linter changes for 4298"

This reverts commit c7a00a9710.

* Linter fixes

* update

* update

* Update distributed_hyperparameter_tuning.ipynb

* removed changes in model_garden folder

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-19 11:56:15 -05:00
Margubur RahmanandGitHub 5efa51206f Migrate gsutil usage to gcloud storage (#4324) 2025-12-19 15:11:28 +00:00
Margubur RahmanandGitHub e604a4d43e Migrate gsutil usage to gcloud storage (#4341) 2025-12-19 15:10:52 +00:00
23af5373ec Migrate gsutil usage to gcloud storage (#4295)
* Migrate gsutil usage to gcloud storage

* removed notes

* Apply automated linter fixes

* Revert "Apply automated linter fixes"

This reverts commit 3f6c4c0d36.

* linter changes

* update

* update

* update

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-19 10:09:05 -05:00
7577c0b1fc Migrate gsutil usage to gcloud storage (#4292)
* Migrate gsutil usage to gcloud storage

* Manual change

* update

* Update NotebookProcessors.py

* Update NotebookProcessors.py

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-19 15:07:50 +00:00
b648f9e73b Migrate gsutil usage to gcloud storage (#4306)
* Migrate gsutil usage to gcloud storage

* changes for 4306

* Apply automated linter fixes

* revert gsutil to gcloud

* Update model_garden_pytorch_stable_diffusion_custom.ipynb

* Update model_garden_pytorch_stable_diffusion_xl_lcm.ipynb

* Update model_garden_pytorch_deployed_model_agent_engine.ipynb

* Update model_garden_pytorch_blip_vqa.ipynb

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-19 15:07:21 +00:00
Margubur RahmanandGitHub 966bbc49a7 Migrate gsutil usage to gcloud storage (#4340) 2025-12-18 13:50:35 -05:00
f5d341ae45 Migrate gsutil usage to gcloud storage (#4303)
* Migrate gsutil usage to gcloud storage

* changes for 4303

* Apply automated linter fixes

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-18 13:49:38 -05:00
70770a50c7 Migrate gsutil usage to gcloud storage (#4302)
* Migrate gsutil usage to gcloud storage

* changes for 4302

* changes for 4302

* removed note

* linter changes

* Revert "linter changes"

This reverts commit a9544e8251.

* Apply automated linter fixes

* Update lightweight_functions_component_io_kfp.ipynb

* Update lightweight_functions_component_io_kfp.ipynb

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-18 13:48:39 -05:00
Matej AleksandrovandCopybara-Service f181c39cbf No public description
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 845941350
2025-12-18 10:04:35 -08:00
Rayan DasoriyaandCopybara-Service a5637f87f2 Add notebook for T5Gemma 2 local inference
PiperOrigin-RevId: 846315479
2025-12-18 10:03:26 -08:00
0103299084 Migrate gsutil usage to gcloud storage (#4304)
* Migrate gsutil usage to gcloud storage

* changes for 4304

* Apply automated linter fixes

* remove unused import

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-18 14:00:23 +00:00
1867536d76 Migrate gsutil usage to gcloud storage (#4294)
* Migrate gsutil usage to gcloud storage

* updates

* removed notes added by agent

* removed note

* linter changes

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-18 13:59:10 +00:00
0edae683e7 Migrate gsutil usage to gcloud storage (#4293)
* Migrate gsutil usage to gcloud storage

* update

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-18 13:58:03 +00:00
Matej AleksandrovandCopybara-Service 8471b5cb6f No public description
PiperOrigin-RevId: 845941350
2025-12-17 15:25:49 -08:00
Aaron DietzandGitHub 87f540ac53 Update notebook_template_review.py (#4405)
Minor change to help us update references to "custom training" to specify "serverless training"
2025-12-17 22:08:00 +00:00
Sam-DecigaandGitHub babeba9f02 feat:NVIDIA Nemotron Nano v2 12B VL - 2025-12-TBD (#4391)
* feat:NVIDIA Nemotron Nano v2 12B VL - 2025-12-TBD

* refactor:Reformat Notebook

* refactor:Reformat Notebook
2025-12-17 18:49:58 +00:00
Vertex MG TeamandCopybara-Service 85c649dd26 Add Llama 3.3 TPU7x deployment notebook.
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 845481841
2025-12-16 16:50:45 -08:00
Vertex MG TeamandCopybara-Service b7135ae1f0 Add Llama 3.3 TPU7x deployment notebook.
PiperOrigin-RevId: 845481841
2025-12-16 16:38:07 -08:00
Vertex MG TeamandCopybara-Service 2f5119a266 Allows the user to select spot VM for deployment
PiperOrigin-RevId: 845064375
2025-12-15 21:23:11 -08:00
Vertex MG TeamandCopybara-Service 23e64ca76f fix: Update TimesFM 2.0 deployment notebook to use GCS path as MODEL_ID
PiperOrigin-RevId: 844826585
2025-12-15 10:25:10 -08:00
gurusai-voletiandGitHub 0ba5a62cc9 Fix ci workflow to use python 3.13 to avoid linter issues (#4397)
* update

* use python 3.13
2025-12-15 13:54:49 +00:00
Damodar PanigrahiandGitHub 0be2c6fd0c feat: authenticate using sa (#4394) 2025-12-12 15:38:25 -05:00
Vertex MG TeamandCopybara-Service cef4928c49 Allows the user to select spot VM for deployment
PiperOrigin-RevId: 843098639
2025-12-11 01:03:35 -08:00
Vertex MG TeamandCopybara-Service 9bb8107110 Add notebook for using Deepseek 3.2 model on Vertex AI.
PiperOrigin-RevId: 842763792
2025-12-10 09:42:39 -08:00
Vertex MG TeamandCopybara-Service b075990d88 Weekly update the vllm/hf-tei/hf-inference-toolkit containers.
PiperOrigin-RevId: 842339077
2025-12-09 12:04:54 -08:00
Ravi DalalGitHubgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
6a83c4c695 Updated notebook comment for custom vllm container image (#4385)
* updated comment for custom vllm container image

* Update notebooks/official/prediction/vertexai_serving_vllm/vertexai_serving_vllm_cpu_llama3_2_3B.ipynb

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

---------

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-12-05 20:26:32 +00:00
Damodar PanigrahiGitHubgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
820c0f8db4 Update Use Case description and API change to accept GCP auth token (#4384)
* pass auth_token in create_http_client

* lint on the notebook

* feat:Removed the last update date

* Update notebooks/community/alphagenome/README.md

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update notebooks/community/alphagenome/cloudai_alphagenome_vai_quickstart.ipynb

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

---------

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-12-05 18:30:23 +00:00
Damodar PanigrahiGitHubgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
4ab197a4ba AlphaGenome GCP API with quickstart.ipynb and README.md (#4378)
* AlphaGenome GCP API  with quickstart.ipynb and README.md

* Update notebooks/community/alphagenome/README.md

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update notebooks/community/alphagenome/README.md

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update notebooks/community/alphagenome/README.md

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update notebooks/community/alphagenome/README.md

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update notebooks/community/alphagenome/README.md

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update notebooks/community/alphagenome/README.md

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update cloudai_alphagenome_vai_quickstart.ipynb

lint errors

* Update cloudai_alphagenome_vai_quickstart.ipynb

lint errors

* lint errors

* lint errors

* lint import order

* lint errors

* lint import order

* lint import

* Update cloudai_alphagenome_vai_quickstart.ipynb format

* Update cloudai_alphagenome_vai_quickstart.ipynb remove hardcoded url

---------

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-11-27 18:56:07 +00:00
Damodar PanigrahiandGitHub 2a5877fbd1 Update CODEOWNERS (#4380)
* Update CODEOWNERS

* Update CODEOWNERS
2025-11-27 18:30:33 +00:00
Vertex MG TeamandCopybara-Service 090e1d9fee Add dynamic model loading/unloading and instructions to model co-hosting notebook.
PiperOrigin-RevId: 837273935
2025-11-26 15:17:30 -08:00
Eric DongandGitHub 2e049d4830 feat: Add new supported model for Claude (#4377)
* feat: Add new supported model for Claude

* Fix elif

* Remove unused endpoint
2025-11-24 16:56:19 -05:00
Vertex MG TeamandCopybara-Service e3320d2126 Update vLLM docker URI in model co-hosting notebook.
PiperOrigin-RevId: 836250727
2025-11-24 09:14:59 -08:00
Vertex MG TeamandCopybara-Service 5fc0e03ca3 Use separate regions for training, evaluation, and deployment in Llama 3.1 finetuning notebook
PiperOrigin-RevId: 835030738
2025-11-20 20:31:36 -08:00
Vertex MG TeamandCopybara-Service ff2a16237d Minor typo fixes and prints
PiperOrigin-RevId: 834789027
2025-11-20 09:15:31 -08:00
Vertex MG TeamandCopybara-Service d26f081642 Weekly update the vllm/hf-tei/hf-inference-toolkit containers.
PiperOrigin-RevId: 833620797
2025-11-17 20:35:49 -08:00
Vertex MG TeamandCopybara-Service 8a0a39176c Some minor updates and refactoring
PiperOrigin-RevId: 832558814
2025-11-14 20:22:08 -08:00
Vertex MG TeamandCopybara-Service 646532ea69 Fixed the deployment quota check and modified the documentation.
PiperOrigin-RevId: 832174331
2025-11-13 23:14:25 -08:00
Vertex MG TeamandCopybara-Service 9e96a3da67 MiniMax-M2 deployment notebook
PiperOrigin-RevId: 831880572
2025-11-13 08:58:34 -08:00
Vertex MG TeamandCopybara-Service db34e1fbd5 Add multi-model benchmark utility and benchmark results to model co-hosting tutorial notebook.
PiperOrigin-RevId: 831592927
2025-11-12 16:54:24 -08:00
Sam-DecigaandGitHub 82308acbac Update anthropic_claude_3_intro.ipynb - Sonnet 3.7 Deprecation (#4364)
Given information above. Approved.
2025-11-12 12:32:57 -05:00
Vertex MG TeamandCopybara-Service ee0ba75d1e Weekly update the vllm/hf-tei/hf-inference-toolkit containers.
PiperOrigin-RevId: 830644861
2025-11-10 16:31:45 -08:00
Vertex MG TeamandCopybara-Service d61aedc721 Weekly update the vllm/hf-tei/hf-inference-toolkit containers.
PiperOrigin-RevId: 829629618
2025-11-07 17:08:47 -08:00
Vertex MG TeamandCopybara-Service 19f7f94af5 DeepSeek-OCR deployment notebook
PiperOrigin-RevId: 827772392
2025-11-03 21:08:50 -08:00
Vertex MG TeamandCopybara-Service 4e5ce9b226 Add single-model multi-replica & multi-model model co-hosting tutorial notebook.
PiperOrigin-RevId: 826614269
2025-10-31 13:46:14 -07:00
Vertex MG TeamandCopybara-Service 99938244f4 Add DWS to the 8B model in the Eval section
PiperOrigin-RevId: 825943414
2025-10-30 02:42:42 -07:00
Vertex MG TeamandCopybara-Service 0cc7be4a6a Added remote sensing deployment notebook
PiperOrigin-RevId: 825002026
2025-10-28 06:13:52 -07:00
Vertex MG TeamandCopybara-Service b6bde41850 Deepseek deployment v3_2 notebook
PiperOrigin-RevId: 824868883
2025-10-27 23:39:38 -07:00
Vertex MG TeamandCopybara-Service 52444a0933 Weekly update the vllm/hf-tei/hf-inference-toolkit containers.
PiperOrigin-RevId: 824690270
2025-10-27 14:57:44 -07:00
Vertex MG TeamandCopybara-Service bab9c398fd Add new variants to qwen3-vl
PiperOrigin-RevId: 823383168
2025-10-24 00:02:05 -07:00
Vertex MG TeamandCopybara-Service 447affcc93 Weekly update the vllm/hf-tei/hf-inference-toolkit containers.
PiperOrigin-RevId: 823276006
2025-10-23 18:49:58 -07:00
Bhaskar GoyalandGitHub 6132c37be0 Initiate Deprecation for Mistral Large (24.11) and Codestral (25.01) (#4347) 2025-10-23 16:38:09 +00:00
Vertex MG TeamandCopybara-Service 8d7f59aeec Update vLLM TPU deployment container image URI.
PiperOrigin-RevId: 822769821
2025-10-22 15:49:05 -07:00
Vertex MG TeamandCopybara-Service bd327ad424 Qwen3-VL deployment notebook
PiperOrigin-RevId: 822441900
2025-10-21 23:36:55 -07:00
Vertex MG TeamandCopybara-Service 3b1fbdb382 Update image in text+image chat completions requests in MG notebooks.
PiperOrigin-RevId: 821912015
2025-10-20 19:56:47 -07:00
Vertex MG TeamandCopybara-Service 954043a729 No public description
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 821735608
2025-10-20 12:10:35 -07:00
Vertex MG TeamandCopybara-Service 16ef9ee80e Add notebook for deploying GPT OSS models on G4 (RTX Pro 6000).
PiperOrigin-RevId: 821735608
2025-10-20 11:42:04 -07:00
Bhaskar GoyalandGitHub 79301b4a4d <feature> - Add Codestral 2 Model (#4291) 2025-10-16 22:05:58 +00:00
Vertex MG TeamandCopybara-Service 0b38d02e6f Add vLLM TPU deployment notebook for qwen3
PiperOrigin-RevId: 820273514
2025-10-16 09:41:13 -07:00
kthytangandGitHub c52ff25ba4 Haiku 4.5 update to anthropic_claude_intro.ipynb (#4300)
* Haiku 4.5 update to anthropic_claude_intro.ipynb

* Update anthropic_claude_intro.ipynb
2025-10-15 20:53:05 +00:00
Vertex MG TeamandCopybara-Service 424400bace Weekly update the vllm/hf-tei/hf-inference-toolkit containers.
PiperOrigin-RevId: 819784960
2025-10-15 09:16:29 -07:00
Mend RenovateandGitHub b1dfac2043 chore(deps): update actions/setup-python action to v6 (#4245) 2025-10-10 14:32:20 +00:00
Mend RenovateandGitHub b81ffcddab chore(deps): update python docker tag to v3.14 (#4284) 2025-10-10 14:29:19 +00:00
Mend RenovateandGitHub 21d8f144aa chore(deps): update dependency pyupgrade to v3.21.0 (#4287) 2025-10-10 14:29:11 +00:00
Vertex MG TeamandCopybara-Service 065a674305 Minor fixes for ollama deployment notebook
PiperOrigin-RevId: 817513895
2025-10-10 00:21:23 -07:00
haomengchaoandGitHub 571d498d08 feat: add notebook for VirtueAI model in Model Garden (#4280)
* feat: add virtueai's notebook for Model Garden

* feat: add virtueai's notebook for Model Garden with fixes

* feat: fix endpoint place holder to pass the test

* feat: fix endpoint place holder to pass the test

* feat: fix endpoint place holder to pass the test

* feat: fix a typo
2025-10-09 22:15:24 +00:00
Bhaskar GoyalandGitHub e936882123 <feature>: Medium 3 launch (#4279) 2025-10-09 16:42:11 +00:00
Vertex MG TeamandCopybara-Service f754f99052 Increase the dws max_wait_duration to 90 minutes
PiperOrigin-RevId: 817042685
2025-10-09 00:11:15 -07:00
Vertex MG TeamandCopybara-Service aa5523a5e9 Delete Gemma 3 peft finetuning notebooks.
PiperOrigin-RevId: 816744576
2025-10-08 09:39:54 -07:00
Vertex MG TeamandCopybara-Service 81393ede1a Weekly update the hf-tei/hf-inference-toolkit containers.
PiperOrigin-RevId: 815926264
2025-10-06 16:17:53 -07:00
Vertex MG TeamandCopybara-Service 5a1c0222da Qwen Image Deployment Notebook
PiperOrigin-RevId: 815781924
2025-10-06 10:19:30 -07:00
Vertex MG TeamandCopybara-Service 1f9326bd56 Weekly update the vllm container version.
PiperOrigin-RevId: 814830089
2025-10-03 14:24:00 -07:00
Vertex MG TeamandCopybara-Service a94cae2e79 Weekly update the vllm/hf-tei/hf-inference-toolkit containers.
PiperOrigin-RevId: 813495515
2025-09-30 17:28:38 -07:00
kthytangandGitHub cb861713c8 Update anthropic_claude_intro.ipynb (#4277) 2025-09-29 17:17:00 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
3a55087789 Bump urllib3 (#4257)
Bumps [urllib3](https://github.com/urllib3/urllib3) from 2.4.0 to 2.5.0.
- [Release notes](https://github.com/urllib3/urllib3/releases)
- [Changelog](https://github.com/urllib3/urllib3/blob/main/CHANGES.rst)
- [Commits](https://github.com/urllib3/urllib3/compare/2.4.0...2.5.0)

---
updated-dependencies:
- dependency-name: urllib3
  dependency-version: 2.5.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2025-09-29 16:59:52 +00:00
Vertex MG TeamandCopybara-Service d359b21f3e Weekly update the vllm/hf-tei/hf-inference-toolkit containers.
PiperOrigin-RevId: 811502904
2025-09-25 14:30:34 -07:00
Vertex MG TeamandCopybara-Service c7d4123b25 Print project and region information
PiperOrigin-RevId: 810501296
2025-09-23 10:54:22 -07:00
Vertex MG TeamandCopybara-Service 07a8bb2d0c Migrate Phi-4 notebook to use Model Garden SDK
PiperOrigin-RevId: 810268843
2025-09-22 20:52:05 -07:00
Vertex MG TeamandCopybara-Service f6b6f365b6 Migrate Ollama deploy notebook to use Model Garden SDK
PiperOrigin-RevId: 810150526
2025-09-22 14:13:49 -07:00
Vertex MG TeamandCopybara-Service fec825f9e5 Use dictionary for the deletion of multiple endpoints
PiperOrigin-RevId: 810066401
2025-09-22 10:26:45 -07:00
Vertex MG TeamandCopybara-Service 1e9bf72097 Blip2 notebook refactoring
PiperOrigin-RevId: 810050386
2025-09-22 09:47:46 -07:00
Vertex MG TeamandCopybara-Service 3fa1cf99ec Make 4b as the default model_version in the Gemma3 notebook
PiperOrigin-RevId: 809821706
2025-09-21 19:09:19 -07:00
Vertex MG TeamandCopybara-Service 2bcaf8abde Allow the user to enter Region
PiperOrigin-RevId: 808435945
2025-09-18 00:17:02 -07:00
Vertex MG TeamandCopybara-Service 759495a1f8 Migrate LaMa notebook to use Model Garden SDK
PiperOrigin-RevId: 808001560
2025-09-16 23:19:01 -07:00
Vertex MG TeamandCopybara-Service 9925e62c4b feat: Refactor to use deploy SDK.
PiperOrigin-RevId: 807803299
2025-09-16 12:35:20 -07:00
Vertex MG TeamandCopybara-Service 40ade71b35 Weekly update the vllm serving container version to 20250911_0916_RC01.
PiperOrigin-RevId: 807413330
2025-09-15 15:45:51 -07:00
Vertex MG TeamandCopybara-Service 65173071a5 Weekly update the serving container version for hf-inference-toolkit and hf-tei containers.
PiperOrigin-RevId: 807413263
2025-09-15 15:44:23 -07:00
Vertex MG TeamandCopybara-Service a514bb51c2 Allow the user to enter Region
PiperOrigin-RevId: 807182772
2025-09-15 04:21:01 -07:00
Rayan DasoriyaandCopybara-Service a6dc1b0f6d Fix notebook issues
PiperOrigin-RevId: 806407972
2025-09-12 13:33:39 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
75236998c7 Bump torch (#4239)
Bumps [torch](https://github.com/pytorch/pytorch) from 2.2.0 to 2.8.0.
- [Release notes](https://github.com/pytorch/pytorch/releases)
- [Changelog](https://github.com/pytorch/pytorch/blob/main/RELEASE.md)
- [Commits](https://github.com/pytorch/pytorch/compare/v2.2.0...v2.8.0)

---
updated-dependencies:
- dependency-name: torch
  dependency-version: 2.8.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2025-09-10 23:33:41 +00:00
MarkandGitHub 9c8a7808bf chore: Replace 'prediction' with 'inference' per urgent rebranding request (#4231) 2025-09-10 23:32:32 +00:00
Holt SkinnerandGitHub 86ce1576d2 Delete notebooks/official/model_evaluation/automl_video_classification_model_evaluation.ipynb (#4256) 2025-09-10 23:32:08 +00:00
Holt Skinner c2b743bfeb Removed Deprecated notebooks 2025-09-10 18:31:40 -05:00
Vertex MG TeamandCopybara-Service b04575d746 Fix axolotl gcs output path.
PiperOrigin-RevId: 805417425
2025-09-10 10:26:55 -07:00
Vertex MG TeamandCopybara-Service b90d16885e Weekly update the serving container version for hf-inference-toolkit and hf-tei containers.
PiperOrigin-RevId: 805081873
2025-09-09 15:12:41 -07:00
Holt SkinnerandGitHub 828f1a26fa Delete notebooks/official/pipelines/google_cloud_pipeline_components_automl_text.ipynb (#4254)
b/442906902
2025-09-09 20:43:03 +00:00
Holt SkinnerandGitHub 915e5edf0a chore: Remove AutoML Notebooks for deprecated Text and Video features (#4253)
* chore: Remove AutoML Notebooks for deprecated Text and Video features

* Remove remaining Text/Video samples
2025-09-09 16:49:06 +00:00
Vertex MG TeamandCopybara-Service 5d4c1be285 Migrate llava notebook to use Model Garden SDK
PiperOrigin-RevId: 804765256
2025-09-09 00:16:12 -07:00
Vertex MG TeamandCopybara-Service 275bfb8f69 Weekly update the vllm serving container version to 20250905_0916_RC01.
PiperOrigin-RevId: 804514266
2025-09-08 11:21:48 -07:00
Vertex MG TeamandCopybara-Service 93cb0ca3d9 Migrate blip image captioning notebook to use Model Garden SDK
PiperOrigin-RevId: 803091009
2025-09-04 10:49:28 -07:00
Vertex MG TeamandCopybara-Service 41d60d052d Add EmbeddingGemma local inference notebook.
PiperOrigin-RevId: 803038284
2025-09-04 08:30:34 -07:00
Vertex MG TeamandCopybara-Service d872cdcf7e Migrate Paligemma2 notebook to use Model Garden SDK
PiperOrigin-RevId: 802861534
2025-09-03 22:35:04 -07:00
Vertex MG TeamandCopybara-Service ada5e4a854 chore: fix google-auth and requests package version.
PiperOrigin-RevId: 802270767
2025-09-02 13:33:33 -07:00
Vertex MG TeamandCopybara-Service ecb32b099d Migrate Stable Diffusion Upscaler notebook to use Model Garden SDK
PiperOrigin-RevId: 801824204
2025-09-01 08:55:58 -07:00
Vertex MG TeamandCopybara-Service f8d09e8e9b Migrate Stable Diffusion XL Lightning notebook to use Model Garden SDK
PiperOrigin-RevId: 801782973
2025-09-01 06:06:41 -07:00
Vertex MG TeamandCopybara-Service 007df88fba Migrate Llama4 notebook to use Model Garden SDK
PiperOrigin-RevId: 801662862
2025-08-31 22:27:06 -07:00
Vertex MG TeamandCopybara-Service c70f3ef9a3 Weekly update vllm/hf-tei/hf-inference-toolkit container image versions.
PiperOrigin-RevId: 801034279
2025-08-29 14:39:29 -07:00
Vertex MG TeamandCopybara-Service 505e101452 Bug fix for Wan2.2
PiperOrigin-RevId: 800419261
2025-08-28 05:11:01 -07:00
Vertex MG TeamandCopybara-Service b77d51b58b Migrate QWEN3 notebook to use Model Garden SDK
PiperOrigin-RevId: 800346334
2025-08-28 00:57:59 -07:00
Rayan DasoriyaandCopybara-Service deaa1ccf2b Update quota check to use gcloud beta quotas info describe command
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 800261134
2025-08-27 19:19:54 -07:00
Vertex MG TeamandCopybara-Service f6e38860aa Update H100/H200 region recommendations in DeepSeek deployment notebook.
PiperOrigin-RevId: 799800908
2025-08-26 18:29:57 -07:00
Vertex MG TeamandCopybara-Service 4937e382b1 Bug fix for Wan2.1
PiperOrigin-RevId: 799116301
2025-08-25 07:40:20 -07:00
Vertex MG TeamandCopybara-Service 0fe2770947 Weekly update the vLLM container image version.
PiperOrigin-RevId: 798391915
2025-08-22 16:59:25 -07:00
Vertex MG TeamandCopybara-Service 80d7ee67d6 Weekly update the container version for hf-inference-toolkit and hf-tei.
PiperOrigin-RevId: 797920954
2025-08-21 14:45:09 -07:00
Vertex MG TeamandCopybara-Service f17e2d6c8d Migrate gpt-oss deploy notebook to use Model Garden SDK
PiperOrigin-RevId: 797824747
2025-08-21 10:40:45 -07:00
Vertex MG TeamandCopybara-Service ae1cd0ed08 Migrate llama3.3 deployment notebook to use Model Garden SDK
PiperOrigin-RevId: 797622523
2025-08-20 23:34:08 -07:00
denisj3030andGitHub a68b491edc dep35 (#4227) 2025-08-20 20:02:14 +00:00
Vertex MG TeamandCopybara-Service d39bed012a Migrate Controlnet notebook to use Model Garden SDK
PiperOrigin-RevId: 797349160
2025-08-20 09:37:09 -07:00
Vertex MG TeamandCopybara-Service f8c93de1c4 Migrate BLIP2 deploy notebook to use Model Garden SDK
PiperOrigin-RevId: 797131322
2025-08-19 20:34:06 -07:00
Vertex MG TeamandCopybara-Service 0b5d97aa3a use A100 machine by default for Llama3.1 fast deployment
PiperOrigin-RevId: 796788128
2025-08-19 02:51:44 -07:00
Vertex MG TeamandCopybara-Service 87f0785ed1 Migrate Segment Anything Model(SAM) notebook to use Model Garden SDK
PiperOrigin-RevId: 796693389
2025-08-18 20:55:11 -07:00
Vertex MG TeamandCopybara-Service 09c8ea616b Weekly update the vLLM container image version.
PiperOrigin-RevId: 796581324
2025-08-18 14:40:14 -07:00
Vertex MG TeamandCopybara-Service 6c11241cf7 Migrate BLIP2 VQA(Visual Question Answering) notebook to use Model Garden SDK
PiperOrigin-RevId: 796316856
2025-08-18 01:35:17 -07:00
Vertex MG TeamandCopybara-Service 9b7dc2e4fd Migrate BLIP2 VQA(Visual Question Answering) notebook to use Model Garden SDK
PiperOrigin-RevId: 795738195
2025-08-15 22:43:11 -07:00
Vertex MG TeamandCopybara-Service 9b3d43b1a6 Use correct notebook_util.
PiperOrigin-RevId: 794874566
2025-08-13 22:18:11 -07:00
Dustin LuongandCopybara-Service 57ee4e2eab Update Qwen3 deployment notebook with new variants and corrected model names.
PiperOrigin-RevId: 794407429
2025-08-12 22:41:45 -07:00
Dustin LuongandCopybara-Service 4d736d7992 Use correct notebook_util.
PiperOrigin-RevId: 794396544
2025-08-12 22:00:22 -07:00
Dustin LuongandCopybara-Service 69e92a650d Add gpt-oss-20b finetuning with lora on Vertex notebook.
PiperOrigin-RevId: 794317611
2025-08-12 16:57:36 -07:00
Vertex MG TeamandCopybara-Service 4ed979eec4 Migrate Stable Diffusion XL 1.0 notebook to use Model Garden SDK
PiperOrigin-RevId: 793954002
2025-08-11 23:22:40 -07:00
Vertex MG TeamandCopybara-Service 0344de8090 Wan Deployment Notebook
PiperOrigin-RevId: 793839574
2025-08-11 16:30:15 -07:00
Vertex MG TeamandCopybara-Service 83eca0f0cc Update the vLLM container image version.
PiperOrigin-RevId: 793738684
2025-08-11 11:47:34 -07:00
Vertex MG TeamandCopybara-Service 4619b272fa Migrate QwQ notebook to use Model Garden SDK.
PiperOrigin-RevId: 793694751
2025-08-11 10:05:17 -07:00
Vertex MG TeamandCopybara-Service 091fa33c01 Migrate Gemma3 notebook to use Model Garden SDK.
PiperOrigin-RevId: 793692575
2025-08-11 10:01:02 -07:00
Rayan DasoriyaandCopybara-Service 42a05c2b35 No public description
PiperOrigin-RevId: 793644258
2025-08-11 07:50:06 -07:00
Vertex MG TeamandCopybara-Service 8ac32fa42f Migrate Gemma3n notebook to use Model Garden SDK
PiperOrigin-RevId: 793583308
2025-08-11 04:09:41 -07:00
Rayan DasoriyaandCopybara-Service 0303057f11 Update the common util location in the notebooks
PiperOrigin-RevId: 793425712
2025-08-10 18:26:31 -07:00
Ravi DalalandGitHub 7ae13b346a added notebooks and dockerfiles for serving open models on vertexai using vllm custom containers (#4148)
* added notebooks and dockerfiles for serving open models on vertexai using vllm customer containers

* updated official codeowners

* fixed linting errors

* fixed linting errors

* moved notebooks

* ran linter

* fixed param type

* added some formatting

* added autoscaling configuration to model deployment

* fixed a heading

* moved notebooks and docker folder under prediction

* updated notebook repo paths

* switched to raw_predict to avoid code changes and rebuild

* removed linting errors

* fixed readme lint error

* fixed links

* removed dedicated_endpoint_enabled

* updated workdir path

* fixed cell type

* added license

* fixed linting error

* using cloud build container image build

* fixed linting issues

* fixes

* cloudbuild yaml

* fixed gemini review comments

* fixed linting errors

* fixed tpu_count type

* handled invalid device type

* optimized dockerfile run command

* updated dockerfile

* addressed review comments

* fixed linting errors

* added license to cloudbuild and dockerfile

* optimized image build code

* fixed image_name variable
2025-08-08 18:01:03 +00:00
Rayan DasoriyaandCopybara-Service 44390cbd99 Add a no-op message to check_quota if it fails.
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 792654121
2025-08-08 09:32:08 -07:00
Vertex MG TeamandCopybara-Service 58dab0b1bb E5 notebook
PiperOrigin-RevId: 791999324
2025-08-06 22:48:50 -07:00
Vertex MG TeamandCopybara-Service 8d36834fcd Add GPT OSS models deployment notebook.
PiperOrigin-RevId: 791864273
2025-08-06 15:17:54 -07:00
Dustin LuongandCopybara-Service 31ca3e36f4 Add Qwen3-30B-A3B instruct and thinking 2507 variants to notebook.
PiperOrigin-RevId: 791746627
2025-08-06 10:27:39 -07:00
Vertex MG TeamandCopybara-Service a72d7dc49f Fix custom dataset input for axolotl notebooks.
PiperOrigin-RevId: 791627090
2025-08-06 04:16:45 -07:00
denisj3030andGitHub 9efbd48233 cl41 (#4193) 2025-08-05 17:02:14 +00:00
Vertex MG TeamandCopybara-Service 166f0f8ce7 Update the vLLM container image version.
PiperOrigin-RevId: 790903846
2025-08-04 14:56:59 -07:00
Vertex MG TeamandCopybara-Service bd9f9675cf Add us-south1 region
PiperOrigin-RevId: 790762290
2025-08-04 08:35:53 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
42ec0e6ad9 Bump urllib3 (#4173)
Bumps [urllib3](https://github.com/urllib3/urllib3) from 2.0.7 to 2.5.0.
- [Release notes](https://github.com/urllib3/urllib3/releases)
- [Changelog](https://github.com/urllib3/urllib3/blob/main/CHANGES.rst)
- [Commits](https://github.com/urllib3/urllib3/compare/2.0.7...2.5.0)

---
updated-dependencies:
- dependency-name: urllib3
  dependency-version: 2.5.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2025-08-01 14:29:14 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
fde3d98a2a Bump requests (#4174)
Bumps [requests](https://github.com/psf/requests) from 2.32.3 to 2.32.4.
- [Release notes](https://github.com/psf/requests/releases)
- [Changelog](https://github.com/psf/requests/blob/main/HISTORY.md)
- [Commits](https://github.com/psf/requests/compare/v2.32.3...v2.32.4)

---
updated-dependencies:
- dependency-name: requests
  dependency-version: 2.32.4
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2025-08-01 14:28:53 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
b83f869a44 Bump protobuf (#4175)
Bumps [protobuf](https://github.com/protocolbuffers/protobuf) from 3.20.3 to 4.25.8.
- [Release notes](https://github.com/protocolbuffers/protobuf/releases)
- [Changelog](https://github.com/protocolbuffers/protobuf/blob/main/protobuf_release.bzl)
- [Commits](https://github.com/protocolbuffers/protobuf/compare/v3.20.3...v4.25.8)

---
updated-dependencies:
- dependency-name: protobuf
  dependency-version: 4.25.8
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2025-08-01 14:28:28 +00:00
Vertex MG TeamandCopybara-Service b918768776 add batch prediction
PiperOrigin-RevId: 789643160
2025-08-01 00:21:30 -07:00
Eric DongandGitHub a22aceb10c refactor: Batch update icon links for notebooks/community (#4188) 2025-07-31 15:41:21 +00:00
Dustin LuongandCopybara-Service c8b53a614e Fix text in TEI notebook to change instances of nomic-ai/nomic-embed-text-v1 to Qwen/Qwen3-Embedding-8B.
PiperOrigin-RevId: 789204281
2025-07-30 23:26:39 -07:00
Eric DongandGitHub 7be7d5be44 refactor: Batch update icon links for notebooks/official (#4187) 2025-07-31 00:08:25 +00:00
Eric DongandGitHub 02a030d6f1 fix: Update the icon links (#4186) 2025-07-30 20:37:13 +00:00
Dustin LuongandCopybara-Service fc8e9e4483 Set Qwen/Qwen3-Embedding-8B as example model in TEI notebook.
PiperOrigin-RevId: 788960826
2025-07-30 10:48:47 -07:00
Eliot LaidlawandGitHub fa019e051a Add CSM intro notebook (#4178)
* notebook

* linting

* add codeowner

* Add cleanup

* fixes

* Update copyright year
2025-07-30 16:36:51 +00:00
Dustin LuongandCopybara-Service 01d513eea3 Update the deployment notebook for Qwen3 models to include Qwen3-235B-Thinking-2507 and Qwen3-235B-Thinking-2507-FP8.
PiperOrigin-RevId: 788636454
2025-07-29 15:45:43 -07:00
Vertex MG TeamandCopybara-Service 928822cc0c Remove 8-bit mode for Qwen 2.5 finetuning notebook.
PiperOrigin-RevId: 788499530
2025-07-29 10:02:42 -07:00
Vertex MG TeamandCopybara-Service 1dfb4091ea notebook reformatting
PiperOrigin-RevId: 788430948
2025-07-29 06:34:26 -07:00
Dustin LuongandCopybara-Service ffb5ea7e88 Update HF TEI serving image to use new, FEDRamp compliant container.
PiperOrigin-RevId: 788203719
2025-07-28 16:36:17 -07:00
Vertex MG TeamandCopybara-Service dd2028d76c add batch prediction
PiperOrigin-RevId: 787943651
2025-07-28 04:02:27 -07:00
Vertex MG TeamandCopybara-Service 70daf2e605 Add Qwen3-Coder notebook
PiperOrigin-RevId: 787286540
2025-07-25 16:57:44 -07:00
Vertex MG TeamandCopybara-Service 42db3643d0 Fix pad tokens for custom weights.
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 787196797
2025-07-25 12:14:37 -07:00
Erwin HuizengaandGitHub 66ce2fa5b7 feat: Add sample for Vertex distributed training (#4163)
* feat: Add sample for Vertex distributed training

* refactor: Move distributed training to community content and add job config

* fix: Address review comments and update files

* minor fixes in the script

* updated codeowners
2025-07-23 21:15:09 +02:00
Dustin LuongandCopybara-Service 44a63c8186 Add Qwen3-235B-A22B-Instruct-2507 models to Qwen3 deployment notebook.
PiperOrigin-RevId: 786375387
2025-07-23 12:12:28 -07:00
Vertex MG TeamandCopybara-Service a29f7a376b Wan Deployment Notebook
PiperOrigin-RevId: 786069875
2025-07-22 18:05:53 -07:00
kittyabsandGitHub 66667ea1db Update ray_cluster_management.ipynb (#4171)
Updated to the current version of Ray supported.
2025-07-22 21:40:52 +00:00
Vertex MG TeamandCopybara-Service 0588a7b62d dedicated endpoints enabled
PiperOrigin-RevId: 785721997
2025-07-21 23:25:34 -07:00
Vertex MG TeamandCopybara-Service ba74664aa4 Update the hf-pytorch-inference notebook using the newly built in-house container hf-inference-toolkit.
PiperOrigin-RevId: 785559005
2025-07-21 13:42:01 -07:00
Vertex MG TeamandCopybara-Service 26011283d2 Updates to Flux.1 Schnell and CogVideoX-2b notebook
PiperOrigin-RevId: 785477091
2025-07-21 10:02:34 -07:00
Vertex MG TeamandCopybara-Service 737c635a59 Update auto-scaling documentation link in Model Garden notebooks.
PiperOrigin-RevId: 785450104
2025-07-21 08:40:26 -07:00
Vertex MG TeamandCopybara-Service e998333f34 Update SGLang version in Qwen3 deployment notebook.
PiperOrigin-RevId: 784744196
2025-07-18 16:57:58 -07:00
Changyu ZhuandCopybara-Service 6ea2b5dd24 Update OpenCLIP and BiomedCLIP serving container URI
PiperOrigin-RevId: 784702984
2025-07-18 14:27:46 -07:00
Vertex MG TeamandCopybara-Service 48f9d563a4 No public description
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 784702418
2025-07-18 14:26:19 -07:00
Vertex MG TeamandCopybara-Service 9167c42cbc Fix Gemma3 vllm deployment for 1b model
PiperOrigin-RevId: 784520440
2025-07-18 03:31:48 -07:00
Vertex MG TeamandCopybara-Service 9bd261b2c0 Refactoring the notebook
PiperOrigin-RevId: 784450805
2025-07-17 22:54:34 -07:00
Vertex MG TeamandCopybara-Service 12657051bd No public description
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 784221303
2025-07-17 10:06:59 -07:00
Vertex MG TeamandCopybara-Service b9471efe45 Add DeepSeek-R1-Distill a4x-highgpu-4g GB200 sample notebook.
PiperOrigin-RevId: 784214583
2025-07-17 09:45:17 -07:00
Vertex MG TeamandCopybara-Service 01d8d165da Refactor axolotl notebook.
PiperOrigin-RevId: 784071376
2025-07-17 01:05:08 -07:00
Vertex MG TeamandCopybara-Service de28ab9241 Update nllb notebook to use new container that is FedRamp compliant.
PiperOrigin-RevId: 783613303
2025-07-15 22:52:50 -07:00
Vertex MG TeamandCopybara-Service 52b91234e9 Update owl-vit notebook to use new container that is FedRamp compliant.
PiperOrigin-RevId: 783524069
2025-07-15 17:03:32 -07:00
Aaron DietzGitHubgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
eaa827959f Update notebook_template_review.py (#4158)
* Update notebook_template_review.py

Added icons to the "Open in" links in our generated list of notebook tutorials

* Update notebooks/notebook_template_review.py

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update notebooks/notebook_template_review.py

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update notebooks/notebook_template_review.py

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update notebooks/notebook_template_review.py

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update notebooks/notebook_template_review.py

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

---------

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-07-15 20:15:50 +00:00
Vertex MG TeamandCopybara-Service 489bdfc092 clean up section is added at the end of the every deployment section
PiperOrigin-RevId: 783221166
2025-07-15 01:00:01 -07:00
Vertex MG TeamandCopybara-Service 2631ce8f3b Refactoring the notebook
PiperOrigin-RevId: 783193023
2025-07-14 23:11:08 -07:00
Vertex MG TeamandCopybara-Service cb3fcaa27e Add support for multiple regions in Hex Deployment
PiperOrigin-RevId: 783168207
2025-07-14 21:26:30 -07:00
Vertex MG TeamandCopybara-Service 6a971663b3 Migrate pytorch-inference docker from 0.1 to 0.4 in notebooks
PiperOrigin-RevId: 783082929
2025-07-14 16:24:47 -07:00
Changyu ZhuandCopybara-Service e29e127d7b Migrate Blip and Blip2 notebooks to use the pytorch-inference container
PiperOrigin-RevId: 782961714
2025-07-14 10:41:52 -07:00
Dustin LuongandCopybara-Service 0ff91f926d Update SAM notebook to use new container that is FedRamp compliant.
PiperOrigin-RevId: 782652376
2025-07-13 14:06:26 -07:00
Vertex MG TeamandCopybara-Service ca19b8f8e7 In Hugging Face TEI notebook, add example to download hf model artifacts and upload to gcs.
PiperOrigin-RevId: 782119696
2025-07-11 14:58:51 -07:00
Rayan DasoriyaandCopybara-Service ad06948c12 Upgrade google-cloud-aiplatform to 1.103.0
PiperOrigin-RevId: 781812082
2025-07-10 21:15:40 -07:00
Dustin LuongandCopybara-Service 7ed8313d37 Add link to timesfm serving docker source code.
PiperOrigin-RevId: 781718474
2025-07-10 15:59:44 -07:00
Vertex MG TeamandCopybara-Service c559851f78 delete mediapipe training container source codes and notebooks.
PiperOrigin-RevId: 781618082
2025-07-10 11:41:52 -07:00
Yichen ZhouandCopybara-Service 7374698440 TimesFM 2.0 notebook demonstrating how to
1. manually deploy a 2.0 docker to an endpoint
2. query the endpoint

PiperOrigin-RevId: 781615204
2025-07-10 11:33:52 -07:00
Vertex MG TeamandCopybara-Service b77fb73ad0 request timeout set to 30 minutes
PiperOrigin-RevId: 781362956
2025-07-09 22:10:38 -07:00
Vertex MG TeamandCopybara-Service f5e9d7a9ed Remove TIMM notebook
PiperOrigin-RevId: 781186738
2025-07-09 13:28:23 -07:00
Vertex MG TeamandCopybara-Service e9c0a56b72 Create T5Gemma local inference notebook
PiperOrigin-RevId: 781066112
2025-07-09 08:32:54 -07:00
Vertex MG TeamandCopybara-Service d4545dbc61 update the training docker image
PiperOrigin-RevId: 780886915
2025-07-08 23:45:47 -07:00
Aaron DietzandGitHub d22db6d795 Update notebook_template_review.py (#4138)
Updates references to Vertex AI Prediction --> Vertex AI Inference (branding name change)
2025-07-07 18:00:41 +00:00
Vertex MG TeamandCopybara-Service 45af74953a delete eval and train job if the job was triggered
PiperOrigin-RevId: 779098485
2025-07-04 02:00:49 -07:00
Vertex MG TeamandCopybara-Service fd5574fa12 No public description
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 776672423
2025-07-02 19:34:48 -07:00
Vertex MG TeamandCopybara-Service 85fa955e6d Update vLLM container version in QwQ deployment notebook.
PiperOrigin-RevId: 778559325
2025-07-02 10:34:24 -07:00
Vertex MG TeamandCopybara-Service bbed90a494 Remove JAX Keras TPU train related contents from Gemma finetuning notebook. This path has been deprecated.
PiperOrigin-RevId: 778151963
2025-07-01 12:47:44 -07:00
denisj3030andGitHub a037d2bd78 marking opus 3 as deprecated (#4132) 2025-07-01 15:52:35 +00:00
Vertex MG TeamandCopybara-Service 5de7f31c07 Update SGLang container URI in Gemma 3n deployment notebook.
PiperOrigin-RevId: 776711854
2025-06-27 14:08:28 -07:00
Vertex MG TeamandCopybara-Service c8b50b4195 No public description
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 776379870
2025-06-27 10:07:48 -07:00
denisj3030andGitHub 4589293efc global endpoint opus4 (#4128) 2025-06-27 15:33:11 +00:00
Mend RenovateandGitHub bdf5745870 chore(deps): update dependency pyupgrade to v3.20.0 (#4071) 2025-06-26 18:29:58 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
a4b5c22aa2 Bump torch (#4083)
Bumps [torch](https://github.com/pytorch/pytorch) from 2.2.0 to 2.7.0.
- [Release notes](https://github.com/pytorch/pytorch/releases)
- [Changelog](https://github.com/pytorch/pytorch/blob/main/RELEASE.md)
- [Commits](https://github.com/pytorch/pytorch/compare/v2.2.0...v2.7.0)

---
updated-dependencies:
- dependency-name: torch
  dependency-version: 2.7.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2025-06-26 18:29:07 +00:00
Mend RenovateandGitHub ecfddc0edc chore(deps): update dependency flake8 to v7.3.0 (#4117) 2025-06-26 18:28:08 +00:00
Vertex MG TeamandCopybara-Service 385a8ca2ea Add vLLM + TPU Llama 3.1 and Qwen3 deployment notebook.
PiperOrigin-RevId: 776200752
2025-06-26 10:51:55 -07:00
Vertex MG TeamandCopybara-Service 0e91687156 Add Gemma 3n deployment notebook.
PiperOrigin-RevId: 776158991
2025-06-26 08:59:35 -07:00
Vertex MG TeamandCopybara-Service d04c79b378 additional check for fast deploy option
PiperOrigin-RevId: 775677059
2025-06-25 07:20:11 -07:00
Harizo RajaonaandGitHub ff8a9b9ac5 (WIP) [Mistral] - Add dedicated OCR notebook (#4088)
* Add dedicated OCR notebook

* - Remove OCR mentions in initial notebook
- Clear outputs and variable names in OCR notebook

* Fix

* Fix
2025-06-24 19:42:48 +00:00
Vertex MG TeamandCopybara-Service 5676d07dbd Add Gemma3 axolotl notebook.
PiperOrigin-RevId: 775074649
2025-06-23 22:45:29 -07:00
denisj3030andGitHub be1f7aa631 adding 2nd anthropic notebook (#4118)
* adding 2nd anthropic notebook

* adding 2nd anthropic notebook, links

* adding 2nd anthropic notebook, links fixed
2025-06-23 15:15:35 +00:00
Rayan DasoriyaandCopybara-Service 630de0bea8 Fix title link for workbench.
PiperOrigin-RevId: 774135635
2025-06-21 09:15:34 -07:00
Vertex MG TeamandCopybara-Service e7da210369 Remove the colab notebook for model vit-gpt2-image-captioning. The model card of which was deleted.
PiperOrigin-RevId: 773761048
2025-06-20 10:44:44 -07:00
Vertex MG TeamandCopybara-Service c57dd78a86 Formatting and refactoring of Imagebind notebook
PiperOrigin-RevId: 773680664
2025-06-20 06:45:30 -07:00
Rayan DasoriyaandCopybara-Service 4b5fe2c3cc Upgrade google-cloud-aiplatform version
PiperOrigin-RevId: 773655271
2025-06-20 05:15:21 -07:00
Vertex MG TeamandCopybara-Service 92e0c1b1c7 Formatting and refactoring of Detectron2 Notebook notebook
PiperOrigin-RevId: 772946944
2025-06-18 09:18:11 -07:00
e563a66114 Add support for tpu v6e quota check (#4111)
Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2025-06-17 13:01:19 +00:00
Vertex MG TeamandCopybara-Service de13d8c65c Llama 4 notebook fixes
PiperOrigin-RevId: 772098038
2025-06-16 11:09:45 -07:00
Vertex MG TeamandCopybara-Service a72b7bc9b0 The main changes:
- `get_deployment_pod_name` now extract the app selector to query the pods
- remove dependency to service, using instead pod port instead
- adds a `POD_PORT` as template variable to allow to pass the port from the UI

PiperOrigin-RevId: 771205143
2025-06-13 13:28:12 -07:00
Rayan DasoriyaandCopybara-Service e1cd3ce080 Fix broken github logo
PiperOrigin-RevId: 771161276
2025-06-13 11:23:05 -07:00
Rayan DasoriyaandCopybara-Service c18f3754e1 Update github logo link
PiperOrigin-RevId: 770866301
2025-06-12 18:00:09 -07:00
Rayan DasoriyaandCopybara-Service 844bfbd2d6 Fix notebook issues for kerasnlp to vertex ai
PiperOrigin-RevId: 770865678
2025-06-12 17:57:46 -07:00
Vertex MG TeamandCopybara-Service a9a512cdc4 Dedicated Endpoint Support for SDXL Dreambooth LoRA Finetuning
PiperOrigin-RevId: 770624167
2025-06-12 06:21:39 -07:00
Vertex MG TeamandCopybara-Service 51b2be8ab4 Formatting and refactoring of Falcon Instruct notebook
PiperOrigin-RevId: 770611011
2025-06-12 05:38:21 -07:00
Vertex MG TeamandCopybara-Service ca52e70dcf E5 Notebook to support dedicated endpoint
PiperOrigin-RevId: 770609921
2025-06-12 05:34:23 -07:00
Vertex MG TeamandCopybara-Service 2fe032efcb Updating Llama 4 Notebook
PiperOrigin-RevId: 770307483
2025-06-11 13:57:45 -07:00
Vertex MG TeamandCopybara-Service 1f3a7418f4 chore: Remove preview in import path
PiperOrigin-RevId: 770211946
2025-06-11 10:27:15 -07:00
Vertex MG TeamandCopybara-Service afacbe8ba6 Remove the usage of Service Account and support VPC-SC and refactoring
PiperOrigin-RevId: 769558936
2025-06-10 04:03:34 -07:00
Vertex MG TeamandCopybara-Service 74376f2787 Refactor qwen3 axolotl notebook.
PiperOrigin-RevId: 769218301
2025-06-09 11:12:47 -07:00
Vertex MG TeamandCopybara-Service abc0e9dec4 Formatting and refactoring
PiperOrigin-RevId: 769180142
2025-06-09 09:40:16 -07:00
d0fb60fb7d Add optional support for global quota check (#4096)
Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2025-06-06 12:20:26 +00:00
Vertex MG TeamandCopybara-Service 7d4fb0ff8d Formatting and refactoring
PiperOrigin-RevId: 767152796
2025-06-04 08:18:32 -07:00
Vertex MG TeamandCopybara-Service a6e69a4561 Formatting and refactoring of llama3_1 deployment
PiperOrigin-RevId: 767043085
2025-06-04 02:05:07 -07:00
Vertex MG TeamandCopybara-Service fdaa5a6b90 Formatting and refactoring of Qwen2 deployment notebook
PiperOrigin-RevId: 767042986
2025-06-04 02:03:46 -07:00
Vertex MG TeamandCopybara-Service 5ba56dfc71 Formatting and refactoring
PiperOrigin-RevId: 766516226
2025-06-03 00:05:09 -07:00
Vertex MG TeamandCopybara-Service 4c04abe724 Formatting and refactoring
PiperOrigin-RevId: 766506385
2025-06-02 23:31:56 -07:00
Genquan DuanandCopybara-Service 67a2f84f6e Add deployment examples of gemma3 to the agent notebook.
PiperOrigin-RevId: 766375463
2025-06-02 15:54:32 -07:00
Vertex MG TeamandCopybara-Service 53d58f16a4 Update deepseek deployment notebook to include DeepSeek-R1-0528.
PiperOrigin-RevId: 766335072
2025-06-02 14:06:15 -07:00
Genquan DuanandCopybara-Service f27aec1295 Add deployment examples of llama3/llama4/deepseek-r1 to the agent notebook.
PiperOrigin-RevId: 766321906
2025-06-02 13:33:37 -07:00
Rayan DasoriyaandCopybara-Service 559e476170 Add location specific param to trtllm deployment
PiperOrigin-RevId: 766116147
2025-06-02 03:24:41 -07:00
Vertex MG TeamandCopybara-Service cea8a8dfd9 Add failure check to local inference and local merge command.
PiperOrigin-RevId: 765481008
2025-05-30 23:27:06 -07:00
Vertex MG TeamandCopybara-Service 931ddb5fc0 Modify llama3.3 finetuning notebook to support DeepSeek-R1-Distill-Llama-70B.
PiperOrigin-RevId: 765450688
2025-05-30 21:08:58 -07:00
Vertex MG TeamandCopybara-Service 53d5932a63 Add deepseek-ai/DeepSeek-R1-0528-Qwen3-8B as a supported variant to the qwen3 tuning notebook.
PiperOrigin-RevId: 765286156
2025-05-30 12:11:36 -07:00
Vertex MG TeamandCopybara-Service adc0882b53 Add VPC SC feature and minor reformatting
PiperOrigin-RevId: 764996423
2025-05-29 20:19:46 -07:00
Vertex MG TeamandCopybara-Service daeb35986e Update working dir for qwen3 axolotl notebook.
PiperOrigin-RevId: 764995079
2025-05-29 20:15:10 -07:00
Vertex MG TeamandCopybara-Service 75a70e31c6 Add cli inference option to Axolotl qwen3 notebook.
PiperOrigin-RevId: 764993668
2025-05-29 20:10:12 -07:00
Vertex MG TeamandCopybara-Service f869a657d3 Formatting and refactoring
PiperOrigin-RevId: 764616237
2025-05-29 01:14:25 -07:00
Vertex MG TeamandCopybara-Service fcbb01480b Update axolotl notebook with latest successful configs.
PiperOrigin-RevId: 764071402
2025-05-27 20:51:00 -07:00
Vertex MG TeamandCopybara-Service 655634c236 [FIX] Minor fixes for the Cosmos 1.0 notebook
PiperOrigin-RevId: 764037909
2025-05-27 18:59:55 -07:00
Rayan DasoriyaandCopybara-Service b21fbe8b7e Use trtllm_region for trtllm deployment
PiperOrigin-RevId: 764029854
2025-05-27 18:26:35 -07:00
Genquan DuanandCopybara-Service 21bb972b85 rename notebook name as model_garden_integration_with_agent, which provides examples for adk/agent engine etc.
PiperOrigin-RevId: 763954229
2025-05-27 14:42:51 -07:00
Genquan DuanandCopybara-Service 0fcf2fb285 This notebook shows how to deploy OSS models and integrate with Agent Engine.
PiperOrigin-RevId: 763875104
2025-05-27 11:22:38 -07:00
denisj3030andGitHub 9ce448e0a4 Ocr small fix (#4070)
* adding mistral ocr

* adding mistral ocr

* fix missing ]

* fix codestral ver

* small fixed
2025-05-23 17:45:43 +00:00
denisj3030andGitHub ca8cb4480a adding mistral ocr (#4069)
* adding mistral ocr

* adding mistral ocr

* fix missing ]

* fix codestral ver
2025-05-23 16:03:27 +00:00
Rayan DasoriyaandCopybara-Service 5e4bdf4a2c Add different variable name for trtllm accelerator type
PiperOrigin-RevId: 762039341
2025-05-22 11:04:55 -07:00
denisj3030andGitHub 011422636f adding Anthropic v4 (#4065)
* adding v4

* adding v4 lint fix

* adding v4 lint fix

* adding v4 lint fix
2025-05-22 17:43:28 +00:00
Vertex MG TeamandCopybara-Service dace15a300 Publish Qwen3 Axolotl Notebook.
PiperOrigin-RevId: 761753984
2025-05-21 18:51:56 -07:00
Changyu ZhuandCopybara-Service f3134943a4 Update SGLang docker version in Mode Garden DeepSeek deployment notebook.
PiperOrigin-RevId: 761742957
2025-05-21 18:06:23 -07:00
skarukasandGitHub 5eb0a7e114 Update embedding notebooks to reference gemini-embedding-001 (#4058)
* Update embedding notebooks to reference gemini-embedding-001

* Update embedding notebooks to reference gemini-embedding-001
2025-05-21 12:52:44 +00:00
Rayan DasoriyaandCopybara-Service 48ead2727f Fix llama notebooks
PiperOrigin-RevId: 761112151
2025-05-20 09:22:11 -07:00
Vertex MG TeamandCopybara-Service dbb226a7a4 Update the vllm demo notebook to use the vanilla container image.
PiperOrigin-RevId: 760757673
2025-05-19 13:53:14 -07:00
Changyu ZhuandCopybara-Service 04352e92cc Add TensorRT-LLM deployment sample to Deepseek and Llama3.3 notebooks.
PiperOrigin-RevId: 759818622
2025-05-16 17:44:44 -07:00
Genquan DuanandCopybara-Service 23975591bc Add notebook examples for smooth integration of vmg oss llms + adk.
PiperOrigin-RevId: 759741078
2025-05-16 13:46:02 -07:00
Vertex MG TeamandCopybara-Service 9027adebc2 [VMG Tutorial] Create serving notebook tutorial for hexLLM Llama 3 deep dive.
PiperOrigin-RevId: 759632399
2025-05-16 08:52:42 -07:00
Vertex MG TeamandCopybara-Service ce05c8af80 Fix indentation error in the Fast deploy Chat Completion
PiperOrigin-RevId: 759431771
2025-05-15 21:18:07 -07:00
Vertex MG TeamandCopybara-Service cfb870a323 Updated Llama 4 notebook
PiperOrigin-RevId: 759146983
2025-05-15 07:53:14 -07:00
22709379dd feat: Add prediction dedicated endpoint colab sample (#3942)
* feat: Add prediction dedicated endpoint colab sample

* Update get_started_with_dedicated_endpoint.ipynb

---------

Co-authored-by: TJ(Tianjiao) Liu <tianjiaoliu@google.com>
2025-05-15 01:14:40 +00:00
Vertex MG TeamandCopybara-Service 4aafcfb40f Add qwen3 specific instruction to runtime creation.
PiperOrigin-RevId: 758737197
2025-05-14 10:21:56 -07:00
Rayan DasoriyaandCopybara-Service d53aa0c816 Fix gemma2 notebook
PiperOrigin-RevId: 758283299
2025-05-13 10:43:57 -07:00
Rayan DasoriyaandCopybara-Service 21976bf94f Clarify the HF token requirement
PiperOrigin-RevId: 758277776
2025-05-13 10:33:15 -07:00
Vertex MG TeamandCopybara-Service 434e8ac8fc [Fix] Add NVIDIA_A100_80GB option to vertex finetuning.
PiperOrigin-RevId: 757854822
2025-05-12 11:55:28 -07:00
Vertex MG TeamandCopybara-Service 2376532e3d Fix A100_80GB machine type.
PiperOrigin-RevId: 757734855
2025-05-12 06:17:50 -07:00
Vertex MG TeamandCopybara-Service 8b2fbe3f34 Make runtime connection instructions clearer for axolotl notebook.
PiperOrigin-RevId: 757720267
2025-05-12 05:29:51 -07:00
Vertex MG TeamandCopybara-Service f991ae44ab Fix local finetuning for axolotl.
PiperOrigin-RevId: 757655772
2025-05-12 02:01:51 -07:00
Dustin LuongandCopybara-Service 7eb8b76d19 Remove reference server deployment and inference from llama3.2 notebook.
PiperOrigin-RevId: 757455670
2025-05-11 11:01:47 -07:00
Vertex MG TeamandCopybara-Service 7da9c0b644 Remove a3-ultragpu from Axolotl notebook.
PiperOrigin-RevId: 756910136
2025-05-09 14:35:55 -07:00
Genquan DuanandCopybara-Service 46a75498d2 Update notebook to add more machine type suggestions, and support allowlisted real-time forecasting for graph_operational model.
PiperOrigin-RevId: 756860759
2025-05-09 12:17:35 -07:00
Vertex MG TeamandCopybara-Service c2208eb454 Fix hf cache dir for local finetuning and use gcsfuse for local training and merging of model.
PiperOrigin-RevId: 756772796
2025-05-09 08:15:32 -07:00
Vertex MG TeamandCopybara-Service f9c43d4a8a Add a3-ultragpu with dws to Axolotl notebook.
PiperOrigin-RevId: 756625405
2025-05-08 23:39:48 -07:00
Vertex MG TeamandCopybara-Service 32ae5b4af0 Add gpu type check in axolotl local run.
PiperOrigin-RevId: 756598801
2025-05-08 22:04:50 -07:00
Vertex MG TeamandCopybara-Service 2345f89625 Refactoring and minor fixes
PiperOrigin-RevId: 756570419
2025-05-08 20:25:13 -07:00
Vertex MG TeamandCopybara-Service 46a80b731b Add Dia-1.6B serving notebook
PiperOrigin-RevId: 756499377
2025-05-08 16:13:17 -07:00
Vertex MG TeamandCopybara-Service 2781808a96 Add n1+T4 as one deployment option for e5 model.
PiperOrigin-RevId: 756479268
2025-05-08 15:16:27 -07:00
Vertex MG TeamandCopybara-Service 5af1bc523d Fix Axolotl installation.
PiperOrigin-RevId: 756401696
2025-05-08 11:54:00 -07:00
Rayan DasoriyaandCopybara-Service 9742d29e51 Add eval harness support for Qwen2.5 finetuning notebook
PiperOrigin-RevId: 756350741
2025-05-08 09:50:30 -07:00
Vertex MG TeamandCopybara-Service 30fda48397 change default gpu type as empty to avoid unintended gpu runtime creation
PiperOrigin-RevId: 756199290
2025-05-08 01:17:09 -07:00
Vertex MG TeamandCopybara-Service 6888728be5 Fix dataset vars definition issue and make runtime creation session optional in Axolotl notebook.
PiperOrigin-RevId: 756138907
2025-05-07 21:40:57 -07:00
Vertex MG TeamandCopybara-Service 8c7fbc6210 Phi-4 reasoning variants
PiperOrigin-RevId: 756047269
2025-05-07 16:17:35 -07:00
Dustin LuongandCopybara-Service 4c4e224d31 Update Qwen3 deployment notebook with FP8 support.
PiperOrigin-RevId: 756041267
2025-05-07 15:59:11 -07:00
Vertex MG TeamandCopybara-Service ac98f72005 Fix runtime creation codes in Axolotl notebook to disallow L4x16 combination.
PiperOrigin-RevId: 755968781
2025-05-07 12:51:45 -07:00
Vertex MG TeamandCopybara-Service ddebceb70c Support Workbench in Gemma 2 tuning notebook
PiperOrigin-RevId: 755884581
2025-05-07 09:24:14 -07:00
Vertex MG TeamandCopybara-Service d8e5c3b461 Update Axolotl finetuning notebook to support Qwen3.
PiperOrigin-RevId: 755577906
2025-05-06 16:52:10 -07:00
Eric DongandGitHub 0bfe70afcd refactor: remove runtime reboot (#4026)
* refactor: remove runtime reboot

* Remove spaces

* Use %pip instead

* Update model path

* Update BQ path

* Downgrade numpy for backfoward compability
2025-05-05 21:12:55 +00:00
Ravi DalalandGitHub 7488a5dc27 upgraded spark on ray on vertex ai notebook to 2.42.0 version (#4027)
* upgraded spark on ray on vertex ai notebook to 2.42.0 version

* upgraded spark on ray on vertex ai notebook to 2.42.0 version

* upgraded spark on ray on vertex ai notebook to 2.42.0 version

* upgraded spark on ray on vertex ai notebook to 2.42.0 version
2025-05-05 19:05:30 +00:00
Vertex MG TeamandCopybara-Service 4914987998 Update Axolotl docker version for the finetuning notebook
PiperOrigin-RevId: 754984828
2025-05-05 10:20:32 -07:00
Vertex MG TeamandCopybara-Service abd30065d0 Add support for dedicated endpoint in paligemma finetuning
PiperOrigin-RevId: 753860258
2025-05-01 22:26:25 -07:00
Dustin LuongandCopybara-Service 68960f1221 Add SDK deploy option to Qwen3 deployment notebook.
PiperOrigin-RevId: 753636431
2025-05-01 10:17:08 -07:00
Vertex MG TeamandCopybara-Service 836d1ebbd1 No public description
PiperOrigin-RevId: 753635866
2025-05-01 10:15:36 -07:00
Vertex MG TeamandCopybara-Service f96c830a66 chore: Add back faster_deploy_enabled in deploy SDK.
PiperOrigin-RevId: 753454732
2025-04-30 23:17:36 -07:00
Dustin LuongandCopybara-Service ecd1fc28d2 Set model id to GCS bucket for Qwen3 235B model for better stability.
PiperOrigin-RevId: 753430192
2025-04-30 21:33:20 -07:00
Vertex MG TeamandCopybara-Service da0874682f Set dedicate endpoint as default for model_garden_finetuning_tutorial
PiperOrigin-RevId: 753294164
2025-04-30 13:38:52 -07:00
Dustin LuongandCopybara-Service e851c1ad99 Upload Qwen3 deployment notebook.
PiperOrigin-RevId: 753265446
2025-04-30 12:19:21 -07:00
Vertex MG TeamandCopybara-Service 93fd7088ba Add deployment options for the smaller Qwen 2.5 model versions too in the finetuning notebook.
PiperOrigin-RevId: 753246638
2025-04-30 11:30:02 -07:00
Vertex MG TeamandCopybara-Service cab440a05b Fixes for Llama Prompt Guard variant to deployment notebook.
PiperOrigin-RevId: 752852958
2025-04-29 13:08:40 -07:00
Vertex MG TeamandCopybara-Service f03f88a0f7 Add Llama Prompt Guard variant to deployment notebook.
PiperOrigin-RevId: 752802325
2025-04-29 10:58:25 -07:00
Vertex MG TeamandCopybara-Service 841fbf9f53 Add Llama Guard variant to deployment notebook.
PiperOrigin-RevId: 752801901
2025-04-29 10:56:47 -07:00
Vertex MG TeamandCopybara-Service bed0c09ad9 Add Model Garden deploy SDK
PiperOrigin-RevId: 752589085
2025-04-28 22:43:24 -07:00
Vertex MG TeamandCopybara-Service 2497476009 Remove 0.5B and 1.5B Qwen 2.5 models from the notebook
PiperOrigin-RevId: 752453505
2025-04-28 14:42:09 -07:00
Eric DongandGitHub 5e2841384f refactor: Use raw string in regex (#4016)
* refactor: Use raw string in regex

* np.NaN was removed in the NumPy 2.0

* Move default python version to 3.10

* Move back default python version to 3.9

* Remove version mistmatched tes notebooks
2025-04-28 19:54:44 +00:00
49de587b2c Add optional support for dedicated endpoint (#4014)
Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2025-04-28 12:03:33 +00:00
Vertex MG TeamandCopybara-Service feea47a206 Enable the use of dedicated Endpoints in instructpix2pix notebook
PiperOrigin-RevId: 751641417
2025-04-25 19:45:08 -07:00
Vertex MG TeamandCopybara-Service f4b1b277cf Add HiDream-I1 serving notebook
PiperOrigin-RevId: 751521086
2025-04-25 12:38:34 -07:00
Vertex MG TeamandCopybara-Service e4608c983b chore: Clean up endpoint resource names.
PiperOrigin-RevId: 751415593
2025-04-25 07:58:12 -07:00
Vertex MG TeamandCopybara-Service f147b50332 Support VPC-SC.
PiperOrigin-RevId: 751028021
2025-04-24 09:45:20 -07:00
Changyu ZhuandCopybara-Service b1b16f718b Fix missing positional argument in DeepSeek deployment notebook
PiperOrigin-RevId: 750742323
2025-04-23 15:19:37 -07:00
Vertex MG TeamandCopybara-Service 871eb25dc3 Deprecate the colab notebooks for a few models.
PiperOrigin-RevId: 750692428
2025-04-23 12:54:53 -07:00
denisj3030andGitHub 793515bac2 Launchpad qodo (#3994)
* launchpad notebooks

* Update ai21labs_intro.ipynb

* Update ai21labs_intro.ipynb

* launchpad

* audio play

* lint fixes

* Update ai21labs_intro.ipynb

* Update cambai_intro.ipynb

* Update cambai_intro.ipynb

* Update cambai_intro.ipynb

* qodo lint
2025-04-22 15:25:46 +00:00
Vertex MG TeamandCopybara-Service 4496842a86 Support VPC-SC. Add Model Garden deploy SDK
PiperOrigin-RevId: 750197006
2025-04-22 08:12:30 -07:00
Vertex MG TeamandCopybara-Service 0687238a97 Refactoring
PiperOrigin-RevId: 750127462
2025-04-22 03:47:19 -07:00
Vertex MG TeamandCopybara-Service cd9120bc15 Support VPC-SC. Add Model Garden deploy SDK
PiperOrigin-RevId: 749803092
2025-04-21 08:20:30 -07:00
Vertex MG TeamandCopybara-Service 77388979d7 Add Qwen 2.5 PEFT finetuning notebook
PiperOrigin-RevId: 748978717
2025-04-18 02:28:04 -07:00
Vertex MG TeamandCopybara-Service d9198306a8 Support dedicated endpoints
PiperOrigin-RevId: 748922184
2025-04-17 21:48:54 -07:00
Vertex MG TeamandCopybara-Service 093e32658b Fix check_quota fn for eval in finetuning notebooks
PiperOrigin-RevId: 748701270
2025-04-17 09:09:24 -07:00
Vertex MG TeamandCopybara-Service e6821c94c9 Add a link to file bugs
PiperOrigin-RevId: 748650287
2025-04-17 05:41:07 -07:00
Vertex MG TeamandCopybara-Service 2fbf8c4379 Workbench support and refactoring
PiperOrigin-RevId: 748167158
2025-04-15 23:09:41 -07:00
Vertex MG TeamandCopybara-Service 257c478ef3 Support VPC-SC. Add Model Garden deploy SDK
PiperOrigin-RevId: 748076170
2025-04-15 17:06:01 -07:00
Vertex MG TeamandCopybara-Service 9ce7e61434 Support VPC-SC. Add Model Garden deploy SDK
PiperOrigin-RevId: 748075941
2025-04-15 17:04:26 -07:00
denisj3030andGitHub 42fa7ac1a3 lint fixed (#3993)
* launchpad notebooks

* Update ai21labs_intro.ipynb

* Update ai21labs_intro.ipynb

* launchpad

* audio play

* lint fixes

* Update ai21labs_intro.ipynb

* Update cambai_intro.ipynb

* Update cambai_intro.ipynb

* Update cambai_intro.ipynb
2025-04-15 17:15:08 +00:00
85d43c76ca Add metadata filtering to image warehouse SDK notebook (#3959)
* Add metadata filtering to image warehouse SDK notebook.

* Add metadata filtering to image warehouse SDK notebook and installing dependencies

* linter fix

* Linter fix

* Fix formatting in image_warehouse_sdk notebook

* Fix formatting in image_warehouse_sdk notebook

---------

Co-authored-by: Yehia Elshater <elshater@google.com>
2025-04-15 12:31:27 +00:00
9f4d837e54 Update peft code for stable_20250409 (#3991)
Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2025-04-15 12:29:40 +00:00
Minwoo ParkandCopybara-Service 27db486275 Improve paligemma notebook documentation.
PiperOrigin-RevId: 747492281
2025-04-14 11:09:52 -07:00
Vertex MG TeamandCopybara-Service 2df60d7862 Fix dedicated endpoint codellama
PiperOrigin-RevId: 747448256
2025-04-14 09:22:18 -07:00
Vertex MG TeamandCopybara-Service 7850587517 Support VPC-SC and workbench
PiperOrigin-RevId: 747376863
2025-04-14 05:41:22 -07:00
Vertex MG TeamandCopybara-Service b924278b03 Add lm eval harness to the finetuning notebook
PiperOrigin-RevId: 747343434
2025-04-14 03:44:40 -07:00
Vertex MG TeamandCopybara-Service adb16ca8b4 Update cell ordering in Llama 4 MaaS notebook.
PiperOrigin-RevId: 746284188
2025-04-10 20:47:06 -07:00
Vertex MG TeamandCopybara-Service 42383514ff Add linter change for Llama 4 MaaS notebook.
PiperOrigin-RevId: 746223674
2025-04-10 16:48:23 -07:00
Vertex MG TeamandCopybara-Service 2609db529e Add Llama 4 MaaS Notebook
PiperOrigin-RevId: 746217133
2025-04-10 16:27:22 -07:00
Vertex MG TeamandCopybara-Service 632385c6bc Support VPC-SC and workbench
PiperOrigin-RevId: 746058497
2025-04-10 09:21:44 -07:00
Vertex MG TeamandCopybara-Service 4c4519f679 Support VPC-SC. Add Model Garden deploy SDK
PiperOrigin-RevId: 746057566
2025-04-10 09:18:53 -07:00
Vertex MG TeamandCopybara-Service 98cded3a79 Fix chat completion in model_garden_pytorch_llama3_2_deployment.ipynb
PiperOrigin-RevId: 745642276
2025-04-09 10:26:41 -07:00
Vertex MG TeamandCopybara-Service 8c37e3995b Set dedicate endpoint as default for some model garden samples
PiperOrigin-RevId: 745421163
2025-04-08 21:47:01 -07:00
Vertex MG TeamandCopybara-Service 1281bce438 see b/380317852 for details
PiperOrigin-RevId: 745345688
2025-04-08 17:02:32 -07:00
talshefandGitHub aaf0fd2ea8 Qodo intro notebook (#3968)
* update qodo notebook

* update qodo notebook

* update qodo notebook

* update qodo notebook

* update qodo notebook

* update qodo notebook

* update qodo notebook
2025-04-08 21:18:10 +00:00
Vertex MG TeamandCopybara-Service 8e249b61d9 Update Hugging Face vLLM deploy notebook title and documents.
PiperOrigin-RevId: 745208445
2025-04-08 10:48:41 -07:00
Vertex MG TeamandCopybara-Service 250418bab9 Set dedicate endpoint as default for some model garden samples
PiperOrigin-RevId: 745192594
2025-04-08 10:10:59 -07:00
Vertex MG TeamandCopybara-Service 70ec8b4ae2 Support VPC-SC. Add Model Garden deploy SDK
PiperOrigin-RevId: 745132382
2025-04-08 07:18:19 -07:00
Vertex MG TeamandCopybara-Service e48f3f05be Fix gemma3 finetuning.
PiperOrigin-RevId: 744956022
2025-04-07 20:45:44 -07:00
denisj3030andGitHub a60c0376c4 launchpad notebooks (#3964)
* launchpad notebooks

* Update ai21labs_intro.ipynb

* Update ai21labs_intro.ipynb

* launchpad

* audio play
2025-04-07 19:23:47 +00:00
denisj3030andGitHub 4bb5ffa1ec Update CODEOWNERS (#3970) 2025-04-07 18:53:44 +00:00
Vertex MG TeamandCopybara-Service 215491b24c Update polling time
PiperOrigin-RevId: 744731980
2025-04-07 08:32:12 -07:00
Vertex MG TeamandCopybara-Service e90c6925cb Update Llama 4 deployment notebook.
PiperOrigin-RevId: 744318800
2025-04-05 15:36:47 -07:00
Vertex MG TeamandCopybara-Service 8967e7301b Add Llama 4 deployment notebook
PiperOrigin-RevId: 744291525
2025-04-05 12:10:48 -07:00
Genquan DuanandCopybara-Service 0977d875ea Add forecasting with multi steps with animated visualization.
PiperOrigin-RevId: 744016523
2025-04-04 11:40:59 -07:00
Vertex MG TeamandCopybara-Service 2fec95e8c9 Support auto-scaling while deploying endpoints
PiperOrigin-RevId: 744004928
2025-04-04 11:06:36 -07:00
Vertex MG TeamandCopybara-Service 7665ae6d96 Support VPC-SC. Add Model Garden deploy SDK
PiperOrigin-RevId: 744004419
2025-04-04 11:04:41 -07:00
b58f02d5bf Add polling fn to common util (#3960)
Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2025-04-04 17:00:55 +00:00
Vertex MG TeamandCopybara-Service e09d6f40e5 reformatting of biomedclip
PiperOrigin-RevId: 743846093
2025-04-04 00:56:25 -07:00
Vertex MG TeamandCopybara-Service c45f6a4f4d Support VPC-SC and workbench
PiperOrigin-RevId: 743818321
2025-04-03 22:42:41 -07:00
Mend RenovateandGitHub ffaa5c114e chore(deps): update dependency flake8 to v7.2.0 (#3840) 2025-04-04 00:04:38 +00:00
Vertex MG TeamandCopybara-Service 62d8bbddba Workbench support
PiperOrigin-RevId: 743576188
2025-04-03 08:37:45 -07:00
Vertex MG TeamandCopybara-Service fe8ab55890 Add merge_model_precision_mode to the gemma2 finetuning.
PiperOrigin-RevId: 743487949
2025-04-03 03:19:17 -07:00
Vertex MG TeamandCopybara-Service 68e31b5407 Workbench support
PiperOrigin-RevId: 743436475
2025-04-03 00:08:55 -07:00
Vertex MG TeamandCopybara-Service 1640f568ea Fix lint issues
PiperOrigin-RevId: 743350440
2025-04-02 18:15:41 -07:00
Vertex MG TeamandCopybara-Service 15f058d9a2 chore: Remove deploy SDK due to timeout. Will add back when fixed.
PiperOrigin-RevId: 743199063
2025-04-02 10:49:11 -07:00
Vertex MG TeamandCopybara-Service e9c222737d chore: Remove faster_deploy_enabled in deploy SDK.
PiperOrigin-RevId: 743178007
2025-04-02 09:57:17 -07:00
Vertex MG TeamandCopybara-Service 4b48806393 Add workbench support
PiperOrigin-RevId: 743147221
2025-04-02 08:25:56 -07:00
Vertex MG TeamandCopybara-Service 2a2ab113c3 Mention the use of GCS path for autogluon train
PiperOrigin-RevId: 743128957
2025-04-02 07:30:38 -07:00
Vertex MG TeamandCopybara-Service db0f633415 Refactor the SD Inpainting notebook
PiperOrigin-RevId: 743108958
2025-04-02 06:22:17 -07:00
Vertex MG TeamandCopybara-Service b7f04de30d see design doc for more details: go/gke-model-ui-notebook-design
PiperOrigin-RevId: 742985852
2025-04-01 22:57:47 -07:00
Vertex MG TeamandCopybara-Service a89eadf7e8 chore: Update EULA comments.
PiperOrigin-RevId: 742890889
2025-04-01 16:58:26 -07:00
Changyu ZhuandCopybara-Service d10db455ea Enable speculative decoding draft model for DeepSeek-V3-0324 with SGLang
PiperOrigin-RevId: 742832025
2025-04-01 14:03:03 -07:00
Vertex MG TeamandCopybara-Service 1605c6884c Remove system_labels from OpenModel.deploy
PiperOrigin-RevId: 742760690
2025-04-01 10:49:29 -07:00
Vertex MG TeamandCopybara-Service b2661f9f69 Create Hugging Face vLLM deploy notebook
PiperOrigin-RevId: 742747326
2025-04-01 10:14:36 -07:00
Vertex MG TeamandCopybara-Service 8ad841b0f2 chore: Update SDK deploy timeout to 3 hours.
PiperOrigin-RevId: 742731078
2025-04-01 09:33:22 -07:00
Vertex MG TeamandCopybara-Service 3202f7602b Support VPC-SC. Add Model Garden deploy SDK
PiperOrigin-RevId: 742708341
2025-04-01 08:23:18 -07:00
Vertex MG TeamandCopybara-Service 6f8d986dc3 Add Model Garden deploy SDK
PiperOrigin-RevId: 742698763
2025-04-01 07:57:15 -07:00
Vertex MG TeamandCopybara-Service b5ce88df10 Add Model Garden deploy SDK
PiperOrigin-RevId: 742471878
2025-03-31 18:08:07 -07:00
Vertex MG TeamandCopybara-Service abdf3d73d6 Add Model Garden deploy SDK.
PiperOrigin-RevId: 742471204
2025-03-31 18:05:20 -07:00
Vertex MG TeamandCopybara-Service 0255d172f5 One template notebook for GKE Model UI
see design doc for more details: go/gke-model-ui-notebook-design

PiperOrigin-RevId: 742301507
2025-03-31 10:00:22 -07:00
Vertex MG TeamandCopybara-Service 8cebcbb984 Support VPC-SC. Add Model Garden deploy SDK
PiperOrigin-RevId: 742261473
2025-03-31 07:42:43 -07:00
Changyu ZhuandCopybara-Service 8eb680d0e9 Remove model_garden_source_model_name from Movinet VCN / VAR notebooks
PiperOrigin-RevId: 741586177
2025-03-28 11:14:10 -07:00
Genquan DuanandCopybara-Service 73c312e81a update weather next notebook of GenCast/GraphCast models
PiperOrigin-RevId: 741582377
2025-03-28 11:04:37 -07:00
Vertex MG TeamandCopybara-Service 0ba19214e1 Add Model Garden deploy SDK
PiperOrigin-RevId: 741371959
2025-03-27 20:13:50 -07:00
Vertex MG TeamandCopybara-Service 0568dc979a Add option for enable_llama_tool_parser
PiperOrigin-RevId: 741352395
2025-03-27 18:44:53 -07:00
Vertex MG TeamandCopybara-Service 0fda02c859 Add deploy source to the finetuning notebooks
PiperOrigin-RevId: 741207333
2025-03-27 10:46:00 -07:00
Vertex MG TeamandCopybara-Service 8411f656d0 Add Model Garden deploy SDK.
PiperOrigin-RevId: 741192166
2025-03-27 10:04:19 -07:00
Genquan DuanandCopybara-Service cc4a6675f9 update weather next notebook of GenCast/GraphCast models
PiperOrigin-RevId: 741186602
2025-03-27 09:47:24 -07:00
Vertex MG TeamandCopybara-Service 4e4bc7abe0 Fix minor lint issues
PiperOrigin-RevId: 740751344
2025-03-26 07:00:59 -07:00
Vertex MG TeamandCopybara-Service 95d06a3995 Update notebook with axolotl config suggestions.
PiperOrigin-RevId: 740418242
2025-03-25 11:18:36 -07:00
Vertex MG TeamandCopybara-Service 04b4d0ccd1 Add DeepSeek-V3-0324 for vLLM deployment in DeepSeek notebook.
PiperOrigin-RevId: 740405791
2025-03-25 10:47:12 -07:00
Vertex MG TeamandCopybara-Service f3c440359b chore: Update EULA comments.
PiperOrigin-RevId: 740400897
2025-03-25 10:34:02 -07:00
Changyu ZhuandCopybara-Service bacf337b4e Update #ModelGarden DeepSeek deployment notebook for SGLang 0.4.4.post1 and Deepseek-V3-0324 support
PiperOrigin-RevId: 740392018
2025-03-25 10:11:41 -07:00
Vertex MG TeamandCopybara-Service f5f5bc6a3c chore: Default accept EULA to True in notebooks.
PiperOrigin-RevId: 740116491
2025-03-24 16:43:16 -07:00
Vertex MG TeamandCopybara-Service da1a5fd6be Support VPC-SC.
PiperOrigin-RevId: 740043502
2025-03-24 13:06:13 -07:00
Vertex MG TeamandCopybara-Service fb6f778c6e Support VPC-SC.
PiperOrigin-RevId: 740042813
2025-03-24 13:03:56 -07:00
Dustin LuongandCopybara-Service c4b3f39bfe Remove raw_response field from request. Also fix chat completions format.
PiperOrigin-RevId: 740023921
2025-03-24 12:04:56 -07:00
Minwoo ParkandCopybara-Service 2b76b2618f Update llama3.3 finetuning notebook to use newer docker images.
PiperOrigin-RevId: 739987083
2025-03-24 10:23:03 -07:00
Vertex MG TeamandCopybara-Service c77125131c Support VPC-SC. Add Model Garden deploy SDK
PiperOrigin-RevId: 739555057
2025-03-22 17:07:10 -07:00
Vertex MG TeamandCopybara-Service a88aea1c2c Add system label notebook environment
PiperOrigin-RevId: 739481405
2025-03-22 07:54:19 -07:00
Vertex MG TeamandCopybara-Service 5066381dcc Fix paligemma2 notebook system label
PiperOrigin-RevId: 739380947
2025-03-21 20:38:50 -07:00
Changyu ZhuandCopybara-Service 5c54f6c5e9 Upgrade transformers version in ShieldGemma 2 local inference notebook
PiperOrigin-RevId: 739335379
2025-03-21 16:37:40 -07:00
Dustin LuongandCopybara-Service dc59488638 Add YaRN scaling to QwQ deployment notebook for 128k context length. Fix chatCompletions format.
PiperOrigin-RevId: 738871259
2025-03-20 11:09:46 -07:00
Vertex MG TeamandCopybara-Service 8ea6932dc2 Add local config support for axolotl.
PiperOrigin-RevId: 738854213
2025-03-20 10:27:22 -07:00
Vertex MG TeamandCopybara-Service 617d68a894 Update workbench link and region support for a3-highgpu-8g
PiperOrigin-RevId: 738724173
2025-03-20 02:34:17 -07:00
Vertex MG TeamandCopybara-Service 5845b12fba Fix SD XL notebook dataset
PiperOrigin-RevId: 738722681
2025-03-20 02:28:06 -07:00
Vertex MG TeamandCopybara-Service 1c87625014 Reformat the Pytorch Flux notebook
PiperOrigin-RevId: 738661642
2025-03-19 21:48:36 -07:00
Vertex MG TeamandCopybara-Service 2ba399778b Remove us east5 region for H100 MG deployment from notebooks.
PiperOrigin-RevId: 738464989
2025-03-19 11:02:11 -07:00
Changyu ZhuandCopybara-Service ddff8605c0 Add CSM-1B deployment notebook
PiperOrigin-RevId: 738459214
2025-03-19 10:46:59 -07:00
Vertex MG TeamandCopybara-Service 15ab7be0d8 fix deploy_source fn
PiperOrigin-RevId: 738438140
2025-03-19 09:55:52 -07:00
Vertex MG TeamandCopybara-Service 55f8adc328 Add Model Garden deploy SDK
PiperOrigin-RevId: 738434532
2025-03-19 09:45:53 -07:00
Vertex MG TeamandCopybara-Service 2ef0652cca Update DeepSeek notebook with new vLLM version, configs and H200 support.
PiperOrigin-RevId: 738408599
2025-03-19 08:29:00 -07:00
Dustin LuongandCopybara-Service 57f3cc3094 Add QwQ deployment notebook
PiperOrigin-RevId: 738183275
2025-03-18 16:29:26 -07:00
Vertex MG TeamandCopybara-Service 333ad532cb Add Model Garden deploy SDK
PiperOrigin-RevId: 738171512
2025-03-18 15:51:27 -07:00
Vertex MG TeamandCopybara-Service 1d55f3bf03 Fix minor lint issues
PiperOrigin-RevId: 738056173
2025-03-18 10:25:06 -07:00
Vertex MG TeamandCopybara-Service 5d3a71bd66 Add Model Garden deploy SDK
PiperOrigin-RevId: 737996111
2025-03-18 07:25:00 -07:00
Vertex MG TeamandCopybara-Service d7918ad939 Add Model Garden deploy SDK
PiperOrigin-RevId: 737989958
2025-03-18 07:02:36 -07:00
Vertex MG TeamandCopybara-Service cd468bc7a6 Add Model Garden deploy SDK.
PiperOrigin-RevId: 737988228
2025-03-18 06:56:17 -07:00
Vertex MG TeamandCopybara-Service 35f6daa68a Add support for L4 gpus with default quota and A100_40GB gpu with DWS.
PiperOrigin-RevId: 737865712
2025-03-17 22:10:21 -07:00
Vertex MG TeamandCopybara-Service a4d9e8c38e Support VPC-SC for the BiomedCLIP notebook.
PiperOrigin-RevId: 737819849
2025-03-17 18:22:59 -07:00
Vertex MG TeamandCopybara-Service 2e5c8082ab Add a reliable method for handling long-running (>10) prediction tasks using CURL (CLI) in Colab for enhanced stability.
PiperOrigin-RevId: 737809564
2025-03-17 17:36:10 -07:00
Vertex MG TeamandCopybara-Service f36b154181 Reduce the num_inference_steps from 30 to 25, which does not lose the generated video quality. There is a bug that the notebook does not support well if the total inference go beyond 10min.
PiperOrigin-RevId: 737808074
2025-03-17 17:28:58 -07:00
Vertex MG TeamandCopybara-Service 5f00f0713f Update deployment template to offload text encoder for 14B models.
PiperOrigin-RevId: 737777189
2025-03-17 15:40:35 -07:00
Vertex MG TeamandCopybara-Service f56007d36f Support Workbench in TGI and llama 3.3 serving notebooks
PiperOrigin-RevId: 737749714
2025-03-17 14:20:16 -07:00
Vertex MG TeamandCopybara-Service a165c2eece Create a colab notebook for nvidia cosmos model deployment on vertex.
PiperOrigin-RevId: 737740514
2025-03-17 13:54:55 -07:00
Vertex MG TeamandCopybara-Service 5b72950133 Add SpotVM and Reservations in-depth notebook with vLLM Llama-3.1 deployment tutorial notebook for model garden.
PiperOrigin-RevId: 737705120
2025-03-17 12:11:21 -07:00
8b2ffbad1e mistral small name update (#3886)
Co-authored-by: denisj3030 <denisj@google.com>
2025-03-17 16:31:18 +00:00
Genquan DuanandCopybara-Service b3b3d9294b demo notebook of GenCast/GraphCast models
PiperOrigin-RevId: 737641814
2025-03-17 09:26:01 -07:00
ec08071608 Mistral Small addition (#3885)
Co-authored-by: denisj3030 <denisj@google.com>
2025-03-17 16:18:50 +00:00
Vertex MG TeamandCopybara-Service 01e8b7e242 Fix pretrained model IDs in Gemma 3 deployment notebook.
PiperOrigin-RevId: 736740585
2025-03-13 21:37:49 -07:00
Changyu ZhuandCopybara-Service 1fdb83b791 Add (streaming) chat completion with SGLang to the DeepSeek deployment notebook
PiperOrigin-RevId: 736555836
2025-03-13 10:30:59 -07:00
Vertex MG TeamandCopybara-Service e1794a48a6 Support VPC-SC.
PiperOrigin-RevId: 736496220
2025-03-13 07:23:50 -07:00
Vertex MG TeamandCopybara-Service b5e105ae01 Update owlvit notebook to support Vertex Workbench.
PiperOrigin-RevId: 736335449
2025-03-12 18:35:05 -07:00
Vertex MG TeamandCopybara-Service ed1cd7e6bd Fixed gemma3 model name.
PiperOrigin-RevId: 736145845
2025-03-12 08:42:53 -07:00
Changyu ZhuandCopybara-Service e128823a12 Add ShieldGemma 2 local inference notebook
PiperOrigin-RevId: 736034669
2025-03-12 00:56:08 -07:00
Vertex MG TeamandCopybara-Service c7cbbf369a Fix context length in Gemma 3 deployment notebook.
PiperOrigin-RevId: 736033252
2025-03-12 00:49:02 -07:00
Vertex MG TeamandCopybara-Service 787c34d11f Add Gemma 3 finetuning notebook.
PiperOrigin-RevId: 736020252
2025-03-11 23:44:56 -07:00
Vertex MG TeamandCopybara-Service 841cfc3d09 Add Gemma 3 deployment notebook.
PiperOrigin-RevId: 736017148
2025-03-11 23:29:15 -07:00
040fc424fe Update common util (#3868)
Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2025-03-11 12:33:53 +00:00
Vertex MG TeamandCopybara-Service 69fab881a9 Ollama notebook for deepseek-r1-671b and deepseek-r1-1.5b
PiperOrigin-RevId: 735613073
2025-03-10 21:09:10 -07:00
Mend RenovateandGitHub f2ab60464a chore(deps): update dependency isort to v6.0.1 (#3854) 2025-03-08 01:51:13 +00:00
Franklin WhaiteandGitHub b5f634dd19 Update pandas version on xgboost_training component in kfp2_pipeline.ipynb (#3866)
The current version was causing a the below error when executing on xgboost_training component on vertex.

```
numpy.dtype size changed, may indicate binary incompatibility. Expected 96 from C header, got 88 from PyObject
```

Updating the pandas version as shown [here](https://github.com/numpy/numpy/issues/26710) resolved the issue
2025-03-08 01:49:52 +00:00
Vertex MG TeamandCopybara-Service 582a4542ae Colab notebook for CogVideoX-2b
PiperOrigin-RevId: 734587293
2025-03-07 09:35:56 -08:00
Vertex MG TeamandCopybara-Service 5e069e28eb Flux.1 Schnell config change
PiperOrigin-RevId: 734587136
2025-03-07 09:34:28 -08:00
22d27fafe7 chore: prediction psc private endpoint, add service attachment best practises (#3863)
Co-authored-by: TJ(Tianjiao) Liu <tianjiaoliu@google.com>
2025-03-07 14:20:42 +00:00
Changyu ZhuandCopybara-Service 4078d9449a Enable non-spot VM deployment in DeepSeek deploy notebooks
PiperOrigin-RevId: 734322869
2025-03-06 16:04:45 -08:00
Shawn YangandCopybara-Service 6703ad5d89 chore: Update AgentEngine notebooks.
PiperOrigin-RevId: 734200652
2025-03-06 10:37:06 -08:00
b2d0a04114 Update source code for stable_20250213 peft docker image (#3864)
Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2025-03-06 13:56:56 +00:00
Vertex MG TeamandCopybara-Service 4f866ccc89 Fix minor lint issues
PiperOrigin-RevId: 734055195
2025-03-06 02:25:35 -08:00
Vertex MG TeamandCopybara-Service 2bcd07214a updated with latest docker image.
PiperOrigin-RevId: 734037362
2025-03-06 01:22:32 -08:00
Changyu ZhuandCopybara-Service 429513816d Add SGLang section to deepseek deployment templated notebook
PiperOrigin-RevId: 733553038
2025-03-04 19:52:21 -08:00
Vertex MG TeamandCopybara-Service 4c96a1462e Update train image tag for tutorial notebook
PiperOrigin-RevId: 733546380
2025-03-04 19:15:53 -08:00
Vertex MG TeamandCopybara-Service 10ea4f8a4c Update DeepSeek notebook to support vLLM v0.7.2, and improve instructions and code samples.
PiperOrigin-RevId: 733511350
2025-03-04 16:46:17 -08:00
Vertex MG TeamandCopybara-Service 6dcd8cad37 Add Model Garden finetuning notebook for Llama3 using workbench.
PiperOrigin-RevId: 731363986
2025-02-26 10:14:32 -08:00
Vertex MG TeamandCopybara-Service 56b06420df Update server args in DeepSeek deployment notebook.
PiperOrigin-RevId: 730964887
2025-02-25 11:31:28 -08:00
2ee16893c4 update for satin (#3850)
Co-authored-by: denisj3030 <denisj@google.com>
2025-02-24 18:48:56 +00:00
Vertex MG TeamandCopybara-Service 0cf6d0235e Add vllm llama3.2 multimodal deep dive deployment tutorial notebook for model garden.
PiperOrigin-RevId: 729574413
2025-02-21 10:23:23 -08:00
Vertex MG TeamandCopybara-Service 495dd6a9a7 Add vllm llama3.1 text-only deep dive deployment tutorial notebook for model garden
devsite.

PiperOrigin-RevId: 729520321
2025-02-21 07:18:44 -08:00
Vertex MG TeamandCopybara-Service 0c7b9b96f1 Update DeepSeek notebook to use Spot VMs.
PiperOrigin-RevId: 729332283
2025-02-20 18:51:20 -08:00
Vertex MG TeamandCopybara-Service c8048bf7b9 Support VPC-SC.
PiperOrigin-RevId: 729238423
2025-02-20 13:49:34 -08:00
Vertex MG TeamandCopybara-Service cfd51e6c2b Add Model Garden finetuning notebook for Axolotl.
PiperOrigin-RevId: 729181397
2025-02-20 11:15:37 -08:00
Vertex MG TeamandCopybara-Service 7df853e17e Support VPC-SC.
PiperOrigin-RevId: 729181240
2025-02-20 11:14:09 -08:00
Vertex MG TeamandCopybara-Service b653153e2e Fix typo in PaliGemma2 notebook.
PiperOrigin-RevId: 729121693
2025-02-20 08:27:11 -08:00
Minwoo ParkandCopybara-Service 06d023f120 Update Vertex Model Garden Llama 3.1 and Gemma 2 finetuning notebooks with new training docker image.
PiperOrigin-RevId: 728801185
2025-02-19 14:29:14 -08:00
Vertex MG TeamandCopybara-Service d1f54aa2c8 Update PaliGemma 2 notebook.
PiperOrigin-RevId: 728699666
2025-02-19 09:16:44 -08:00
Minwoo ParkandCopybara-Service b0046bae76 Add Model Garden finetuning tutorial notebook.
PiperOrigin-RevId: 728675586
2025-02-19 08:03:38 -08:00
069d709edd update mistral sdk (#3836)
Co-authored-by: denisj3030 <denisj@google.com>
2025-02-13 19:19:47 +00:00
Vertex MG TeamandCopybara-Service a33a103c05 Fix service account setting with DeepSeek deployment notebook.
PiperOrigin-RevId: 726118186
2025-02-12 10:57:01 -08:00
Vertex MG TeamandCopybara-Service 44c29899e7 Add DeepSeek-V3/R1 multi-host deployment notebook.
PiperOrigin-RevId: 725810376
2025-02-11 16:14:48 -08:00
Vertex MG TeamandCopybara-Service ae44e54d57 feat: Add Reasoning Engine integration notebook
PiperOrigin-RevId: 725511227
2025-02-11 00:24:28 -08:00
Mend RenovateandGitHub 7835b84082 chore(deps): update dependency isort to v6 (#3810) 2025-02-10 20:39:00 +00:00
ethan-gordonandGitHub 30d6112517 set FeatureView IAM and Service Accounts as version v1 and remove preview note. (#3795) 2025-02-10 20:37:36 +00:00
Vertex MG TeamandCopybara-Service 32737fda1f Sync mistral and mixtral ft notebooks
PiperOrigin-RevId: 725241490
2025-02-10 09:20:36 -08:00
Vertex MG TeamandCopybara-Service e02adc516c feat: Update Reasoning Engine with Llama 3.1 models notebook.
PiperOrigin-RevId: 725038312
2025-02-09 19:53:21 -08:00
Vertex MG TeamandCopybara-Service 80320a9a1b Refactor the notebook
PiperOrigin-RevId: 724288214
2025-02-07 03:55:12 -08:00
Minwoo ParkandCopybara-Service 8c527a89ed Add Model Garden finetuning tutorial notebook.
PiperOrigin-RevId: 724140766
2025-02-06 17:56:55 -08:00
Vertex MG TeamandCopybara-Service 33a2b7eaf9 feat: Add Reasoning Engine with Llama 3.1 models notebook.
PiperOrigin-RevId: 724139178
2025-02-06 17:51:09 -08:00
Minwoo ParkandCopybara-Service dfa461dabe Add Model Garden finetuning tutorial notebook.
PiperOrigin-RevId: 724065750
2025-02-06 14:02:30 -08:00
Vertex MG TeamandCopybara-Service c80de903af Update PaliGemma 2 notebook and handler to use weights from GCS.
PiperOrigin-RevId: 724065613
2025-02-06 14:01:01 -08:00
Vertex MG TeamandCopybara-Service dd333b8fdd feat: Add Reasoning Engine with Llama 3.1 models notebook.
PiperOrigin-RevId: 723785091
2025-02-05 22:12:52 -08:00
Vertex MG TeamandCopybara-Service 911ef0cb69 A minor tweaking in the helper functions and a Bug fix in the Predict Section.
PiperOrigin-RevId: 723779782
2025-02-05 21:46:35 -08:00
Vertex MG TeamandCopybara-Service ac6f4d669d Remove Llama-Guard (llama3.1) models from notebook model list
PiperOrigin-RevId: 723282525
2025-02-04 17:21:55 -08:00
Dustin LuongandCopybara-Service abdd887fa1 Set model_garden_source_model_name for model_garden_phi4_deployment notebook.
PiperOrigin-RevId: 722840147
2025-02-03 16:25:21 -08:00
Dustin LuongandCopybara-Service 711a4f0b8c Set model_garden_source_model_name for deployment notebooks.
PiperOrigin-RevId: 722701496
2025-02-03 10:06:48 -08:00
Aiden010200andGitHub 07f30dfde5 Upload a SGD classifier predictor example (#3783)
This example uses aiplatform and scikit-learn library to provide a SGD classifier.
2025-02-03 16:09:56 +00:00
Mend RenovateandGitHub b33e3287b7 chore(deps): update dependency black to v25 (#3815) 2025-02-03 16:09:21 +00:00
Dustin LuongandCopybara-Service 3c06e4797a Set model_garden_source_model_name for some model garden deployment notebooks.
PiperOrigin-RevId: 721846318
2025-01-31 11:39:27 -08:00
Vertex MG TeamandCopybara-Service a5944510d6 Add a notebook about Model Garden advanced features, including prefix caching and speculative decoding.
PiperOrigin-RevId: 721844683
2025-01-31 11:35:01 -08:00
ethan-gordonandGitHub 6d74f87e22 Add notebook vertex_ai_feature_store_update_feature_monitor_feature_group_iam_and_service_agent.ipynb. (#3814)
This change also inserts a corresponding entry to CODEOWNERS.
2025-01-30 22:04:59 +00:00
Vertex MG TeamandCopybara-Service 14031b238c BiomedCLIP Deployment on Vertex Notebook
PiperOrigin-RevId: 720990850
2025-01-29 08:40:29 -08:00
Vertex MG TeamandCopybara-Service 99e578ca2e Add usage tracking labels for finetuning notebooks
PiperOrigin-RevId: 720826465
2025-01-28 22:03:39 -08:00
Vertex MG TeamandCopybara-Service 539683f573 Adding Phi-4 Colab deployment notebook
PiperOrigin-RevId: 720596224
2025-01-28 09:04:43 -08:00
Dustin LuongandCopybara-Service 6fabc23db6 Set model_garden_source_model_name for vllm deployments.
PiperOrigin-RevId: 720345103
2025-01-27 16:20:59 -08:00
Dustin LuongandCopybara-Service 18305e35ea Set model_garden_source_model_name for hexllm deployments.
PiperOrigin-RevId: 720341876
2025-01-27 16:09:59 -08:00
Dustin LuongandCopybara-Service b71f0ec2dc Set model_garden_source_model_name for optimized vllm deployments.
PiperOrigin-RevId: 720317678
2025-01-27 14:54:29 -08:00
Dustin LuongandCopybara-Service d7fe713f7f Set model_garden_source_model_name for TGI deployments.
PiperOrigin-RevId: 720312091
2025-01-27 14:40:10 -08:00
Dustin LuongandCopybara-Service 52925286ec Set model_garden_source_model_name for pytorch inference deployments.
PiperOrigin-RevId: 720311822
2025-01-27 14:38:52 -08:00
Dustin LuongandCopybara-Service b7b14ba7d8 Set model_garden_source_model_name for llama3
reference implementation deployment.

PiperOrigin-RevId: 720311748
2025-01-27 14:38:41 -08:00
Dustin LuongandCopybara-Service 5e3c07f72e Set model_garden_source_model_name for TEI deployments.
PiperOrigin-RevId: 720311569
2025-01-27 14:37:21 -08:00
Vertex MG TeamandCopybara-Service da630753ef Add Mistral and Llama3.1 8B serving notebooks
PiperOrigin-RevId: 719348288
2025-01-24 10:15:58 -08:00
Vertex MG TeamandCopybara-Service c27751e3aa Distinguish between the train and deploy machine specs for Gemma PEFT Finetuning on HF Notebook
PiperOrigin-RevId: 719309601
2025-01-24 08:16:52 -08:00
mumletandGitHub e9428949f1 fix: Update the model file path (#3778)
* fix: Update model path of model_monitoring_for_custom_model_online_prediction.ipynb

Update the unavailable model path for model_monitoring_for_custom_model_online_prediction.ipynb

* Update model_monitoring_for_custom_model_online_prediction.ipynb

Update the bq dataset uri
2025-01-23 23:45:17 +00:00
Vertex MG TeamandCopybara-Service 180c2834fc Fix minor lint issues
PiperOrigin-RevId: 718941819
2025-01-23 11:15:04 -08:00
Vertex MG TeamandCopybara-Service 77fd06d4c0 Fix MaaS requests in Llama Guard notebook.
PiperOrigin-RevId: 718856137
2025-01-23 07:23:05 -08:00
0943520f8d Update template location in dataset validation (#3799)
Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2025-01-23 13:26:44 +00:00
Vertex MG TeamandCopybara-Service c14ed381bf Stable Diffusion XL Finetuning Dreambooth Lora
PiperOrigin-RevId: 718724713
2025-01-22 23:26:09 -08:00
1f15188b6e Update common util and dataset validation util (#3796)
Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2025-01-23 02:27:23 +00:00
Bhaskar GoyalandGitHub 9ef41d55ba feat: Deprecate Claude 3 Sonnet (#3790) 2025-01-21 21:36:16 +00:00
Vertex MG TeamandCopybara-Service d82b5a4066 Enable H100 80GB DWS for 11B and 90B eval.
PiperOrigin-RevId: 717949505
2025-01-21 09:24:17 -08:00
Vertex MG TeamandCopybara-Service fdd4d37275 Paligemma 2 Deployment notebook
PiperOrigin-RevId: 717927000
2025-01-21 08:21:47 -08:00
Vertex MG TeamandCopybara-Service 81f700d508 Rename accelerator variables in the notebook
PiperOrigin-RevId: 717750767
2025-01-20 22:54:59 -08:00
Vertex MG TeamandCopybara-Service 32136a1894 BioGPT serving Notebook
PiperOrigin-RevId: 717710469
2025-01-20 20:13:51 -08:00
Vertex MG TeamandCopybara-Service c93a9c2099 Create Cloud translation and evaluation demo notebook
PiperOrigin-RevId: 715907272
2025-01-15 12:45:59 -08:00
Vertex MG TeamandCopybara-Service 8ece5ef3eb Segment Anything Model (SAM) Serving on Vertex AI Notebook
PiperOrigin-RevId: 715723816
2025-01-15 03:08:36 -08:00
Vertex MG TeamandCopybara-Service dee72afc70 Enable dedicate endpoint for Prompt Guard deployment
PiperOrigin-RevId: 715395374
2025-01-14 08:38:18 -08:00
Vertex MG TeamandCopybara-Service ea45e3dd5c Add usage tracking labels to all the finetuning notebook
PiperOrigin-RevId: 715229459
2025-01-13 21:41:30 -08:00
Vertex MG TeamandCopybara-Service 9ed8896350 vLLM supports GPU HBM + host memory prefix kv caching
PiperOrigin-RevId: 715213919
2025-01-13 20:30:33 -08:00
Vertex MG TeamandCopybara-Service 2709fcd8ed Fix the error when result is list, make it works for both dictionary and list
PiperOrigin-RevId: 715156151
2025-01-13 16:54:47 -08:00
Vertex MG TeamandCopybara-Service a7c1b9af3a Add publisherdb api call response check in case call fails
PiperOrigin-RevId: 715099165
2025-01-13 14:01:37 -08:00
Vertex MG TeamandCopybara-Service 952af223da Update license of the notebooks to 2025
PiperOrigin-RevId: 715041456
2025-01-13 11:17:00 -08:00
Bhaskar GoyalandGitHub d87b6b7463 feat: Add Codestral (25.01) model to mistral docs. (#3779) 2025-01-13 17:05:46 +00:00
Dustin LuongandCopybara-Service 25b3364236 No public description
PiperOrigin-RevId: 713694601
2025-01-09 09:13:27 -08:00
Vertex MG TeamandCopybara-Service 9f80c3cd9b Hex-LLM supports prefix caching as a GA feature
PiperOrigin-RevId: 713503181
2025-01-08 19:47:24 -08:00
Mend RenovateandGitHub 9f6ad8439c Update dependency pyupgrade to v3.19.1 (#3757) 2025-01-08 20:31:26 +00:00
Aiden010200andGitHub 64e9a4ae06 Upload a ResNet predictor example (#3765)
This example uses aiplatform and torch library to provide a ResNet predictor.
2025-01-08 20:30:52 +00:00
Dustin LuongandCopybara-Service 883e1e5fc1 Set system_labels in notebooks
PiperOrigin-RevId: 713040032
2025-01-07 14:20:25 -08:00
Vertex MG TeamandCopybara-Service 847a49f6d0 Update docker images to avoid 'tags' KeyError while loading HF dataset
PiperOrigin-RevId: 711729872
2025-01-03 06:09:28 -08:00
Vertex MG TeamandCopybara-Service 0b9a450b44 MediaPipe Text Classification notebook
PiperOrigin-RevId: 711656574
2025-01-03 00:25:31 -08:00
Vertex MG TeamandCopybara-Service ed72474525 pyTorch IMage Model notebook
PiperOrigin-RevId: 711343028
2025-01-02 00:53:05 -08:00
Vertex MG TeamandCopybara-Service 6b20a30669 TFVision Image segmentation notebook
PiperOrigin-RevId: 710942626
2024-12-31 04:20:47 -08:00
Vertex MG TeamandCopybara-Service f50f5f749d mediapipe face stylizer notebook
PiperOrigin-RevId: 710861510
2024-12-30 20:30:38 -08:00
alicechang0909andGitHub 9621877583 Created using Colab (#3763)
* Created using Colab

* Created using Colab

* Add Featurestore Monitoring functionalities - Fix lint error in import

* Fix lint with commands

* Address comments

* Update restart section to fix lint error.

* fix: fix lint errors

* Fix: Try to fix lint errors

* Fix: try fix lint with python commands

* fix: remove self link

* fix: try submit from workbench
2024-12-30 18:59:13 +00:00
Vertex MG TeamandCopybara-Service 425a5c94c2 Fix PIL issue FreeTypeFont object has no attribute getsize
PiperOrigin-RevId: 710698919
2024-12-30 06:09:47 -08:00
Vertex MG TeamandCopybara-Service 0b8b4b8ad6 movinet action recognition notebook
PiperOrigin-RevId: 709969002
2024-12-26 22:46:35 -08:00
Vertex MG TeamandCopybara-Service f641e5d221 Enable dedicate endpoint for pytorch llama3 deployment
PiperOrigin-RevId: 707697043
2024-12-18 16:07:12 -08:00
Vertex MG TeamandCopybara-Service d89b29f759 Add usage labels to finetuning notebook
PiperOrigin-RevId: 707402297
2024-12-17 22:37:52 -08:00
Vertex MG TeamandCopybara-Service df49e18ce9 Hex-LLM supports disaggregated serving as an experimental feature
PiperOrigin-RevId: 707332453
2024-12-17 18:19:49 -08:00
Changyu ZhuandCopybara-Service dae9e791df Fix missing import in Llama 3 finetuning notebook
PiperOrigin-RevId: 707326361
2024-12-17 18:02:31 -08:00
Changyu ZhuandCopybara-Service 16237de9bb Add fast deployment section to Llama 3.2 deployment notebook
PiperOrigin-RevId: 707274649
2024-12-17 15:32:57 -08:00
Vertex MG TeamandCopybara-Service f5e0394b10 Add H100 80 GB config for Llama 3
PiperOrigin-RevId: 706888765
2024-12-16 17:28:23 -08:00
Vertex MG TeamandCopybara-Service 775aa37b88 Update yolov8 model to use model and endpoint dictionary.
PiperOrigin-RevId: 706750375
2024-12-16 10:13:00 -08:00
Vertex MG TeamandCopybara-Service b43418b17b mediapipe Object detection notebook bug fix and re-formatting
PiperOrigin-RevId: 706673143
2024-12-16 05:26:30 -08:00
Changyu ZhuandCopybara-Service 20138d9333 Add fast deployment section to Llama 3.1 deployment notebook
PiperOrigin-RevId: 705931336
2024-12-13 10:44:11 -08:00
sageof6pathandGitHub fadcb1e618 Peft docker fix (#3751)
* Updated util files to fix peft docker

* fix imports

* fix imports fileutils.py
2024-12-13 13:08:34 +00:00
Vertex MG TeamandCopybara-Service f0ec2acbb8 Update Hex-LLM container URI.
PiperOrigin-RevId: 705656670
2024-12-12 15:53:15 -08:00
Aiden010200andGitHub cc38839e2c Upload xgbranker predictor example (#3737)
* Upload pipeline job example

* Upload pipeline example which can combine other pipeline examples.

* Upload xgbranker predictor example

This example uses aiplatform and xgboost to provide a xgbranker predictor.
2024-12-12 16:03:43 +00:00
Vertex MG TeamandCopybara-Service 94b6a9624b mediapipe Object detection notebook
PiperOrigin-RevId: 705082582
2024-12-11 06:20:13 -08:00
Vertex MG TeamandCopybara-Service 9f8cc0e625 Enable dedicate endpoint for phi3 deployment
PiperOrigin-RevId: 704865877
2024-12-10 15:25:47 -08:00
Vertex MG TeamandCopybara-Service 57cc004855 Update the HF TGI and pytorch-inference notebooks, with the latest container image version.
PiperOrigin-RevId: 704421485
2024-12-09 14:40:05 -08:00
Minwoo ParkandCopybara-Service b9549ccee7 Add Llama 3.3 finetuning notebook.
PiperOrigin-RevId: 703546040
2024-12-06 10:42:19 -08:00
Vertex MG TeamandCopybara-Service 6c7383beaf Adding Qwen2.5-Instruct-32B-AWQ TPU configs to Colab deployment notebook
PiperOrigin-RevId: 703541226
2024-12-06 10:26:35 -08:00
Vertex MG TeamandCopybara-Service 4bec5ef258 Add new Llama 3.3 deployment notebook.
PiperOrigin-RevId: 703537600
2024-12-06 10:15:16 -08:00
Vertex MG TeamandCopybara-Service c2bac62780 tfvision classification notebook
PiperOrigin-RevId: 702659350
2024-12-04 03:25:48 -08:00
Pedro MelendezandGitHub cb7c18e439 Added blog post URL (#3739) 2024-12-03 17:41:39 +00:00
Vertex MG TeamandCopybara-Service e1e4a2ba5e Add vLLM + TPU Llama 3.1 and Qwen 2.5 deployment notebook.
PiperOrigin-RevId: 702091534
2024-12-02 14:46:03 -08:00
Vertex MG TeamandCopybara-Service 2a981b9568 Adding chunked prefill vllm server arg.
PiperOrigin-RevId: 702022997
2024-12-02 11:02:06 -08:00
Vertex MG TeamandCopybara-Service fb10a66d12 A fix in the prediction section
PiperOrigin-RevId: 701847844
2024-12-01 23:09:39 -08:00
Vertex MG TeamandCopybara-Service 2b2019afa6 LLaVA Deployment notebook
PiperOrigin-RevId: 701314497
2024-11-29 10:23:01 -08:00
Vertex MG TeamandCopybara-Service 751b8e7ad5 Avoid copying model artifacts to local GCS and use VERTEX_AI_MODEL_GARDEN_LLAMA_3_1 directly
PiperOrigin-RevId: 700732639
2024-11-27 09:55:49 -08:00
Aiden010200andGitHub 516c395db0 Upload pipeline job example (#3719)
* Upload pipeline example which can combine other pipeline examples.
2024-11-26 15:32:54 +00:00
Vertex MG TeamandCopybara-Service 964c481ed1 Enable dedicate endpoint for timesfm deployment
PiperOrigin-RevId: 700190957
2024-11-25 20:33:57 -08:00
Vertex MG TeamandCopybara-Service 454a90bc33 Enable dedicate endpoint for huggingface tei deployment
PiperOrigin-RevId: 700166368
2024-11-25 18:25:35 -08:00
Vertex MG TeamandCopybara-Service ee43b8c0b7 Add 2H100/4H100 deploy options to llama notebooks
PiperOrigin-RevId: 700107023
2024-11-25 14:45:36 -08:00
Eric DongandGitHub 556f8510f3 fix: Remove example ouptut (#3732) 2024-11-25 21:21:49 +00:00
Eric DongandGitHub c1ba930618 fix: Update the model file path (#3731)
* fix: Update the model file path

* Update the model file path 2
2024-11-25 20:25:48 +00:00
Vertex MG TeamandCopybara-Service 2dc1e96207 Add Llama 3.2 serving notebook.
PiperOrigin-RevId: 700025379
2024-11-25 10:19:14 -08:00
Vertex MG TeamandCopybara-Service 68d5a6d2d2 Enable dedicate endpoint for Llama Guard deployment
PiperOrigin-RevId: 700022732
2024-11-25 10:11:56 -08:00
Vertex MG TeamandCopybara-Service ff2f20fc26 Adding Qwen2/Qwen2.5 TPU configs to Colab deployment notebook
PiperOrigin-RevId: 699209055
2024-11-22 10:07:41 -08:00
Vertex MG TeamandCopybara-Service dd157ca445 A minor fix in the prediction section
PiperOrigin-RevId: 699134138
2024-11-22 05:05:58 -08:00
Vertex MG TeamandCopybara-Service 65d86f57ca Update region suggestion for A100_80GB and H100_80GB gpus
PiperOrigin-RevId: 698841858
2024-11-21 10:55:34 -08:00
0dadbb8400 feat: adding support for Mistral Large 24.11 part2 (#3723)
* feat: adding support for Mistral Large 24.11 part2

* feat: adding support for Mistral Large 24.11 part2

---------

Co-authored-by: denisj3030 <denisj@google.com>
2024-11-21 16:13:13 +00:00
Aaron DietzandGitHub 115413601f Update spark_on_ray_on_vertex_ai.ipynb (#3705)
Fix links for opening the notebook
2024-11-21 15:43:47 +00:00
Vertex MG TeamandCopybara-Service 3e7d427a1b Download only the required files notebook
PiperOrigin-RevId: 698657845
2024-11-20 23:04:15 -08:00
Vertex MG TeamandCopybara-Service 194978fd1c A minor fix in the HexLLM deploy section
PiperOrigin-RevId: 698420315
2024-11-20 09:36:56 -08:00
Vertex MG TeamandCopybara-Service 877305d425 Fix lint issues
PiperOrigin-RevId: 698241868
2024-11-19 20:47:46 -08:00
Vertex MG TeamandCopybara-Service 37c88851bd Enable dedicate endpoint for model_garden_gemma_finetuning_on_vertex.ipynb
PiperOrigin-RevId: 697850787
2024-11-18 20:12:55 -08:00
Vertex MG TeamandCopybara-Service 3b17b30051 Deprecate the Pytorch PEFT notebook. Notebooks such as model_garden_pytorch_llama3_1_finetuning.ipynb demonstrate the usage of the peft in VMG.
PiperOrigin-RevId: 697850024
2024-11-18 20:10:01 -08:00
Vertex MG TeamandCopybara-Service 8883d8e211 Enable dedicate endpoint for model_garden_pytorch_llama3_1_deployment.ipynb
PiperOrigin-RevId: 697498389
2024-11-17 22:16:51 -08:00
Vertex MG TeamandCopybara-Service 0879ef0057 Fix timestamp parameter in prediction section.
PiperOrigin-RevId: 697114085
2024-11-16 00:29:07 -08:00
Vertex MG TeamandCopybara-Service 4adff04d06 Enable dedicate endpoint for model_garden_pytorch_mixtral_deployment.ipynb
PiperOrigin-RevId: 696942601
2024-11-15 11:07:19 -08:00
Vertex MG TeamandCopybara-Service e6a8641896 Enable dedicate endpoint for model_garden_pytorch_qwen2_deployment.ipynb
PiperOrigin-RevId: 696937628
2024-11-15 10:52:28 -08:00
Pedro MelendezandGitHub a740092ae2 Added notebook titled "backoff_and_retry_for_LLMs.ipynb" to the /notebooks/community/generative_ai/ directory (#3706)
* Adding backoff and retry notebook

* Formatted notebook

* Formatted notebook

* Formatted notebook

* Formatted notebook

* Changed URL

* Format changes

* Remoevd URL

* Added formatting

* Added codeowner entry

* Added summary

* Added note about costs

* Lint format
2024-11-15 18:23:02 +00:00
Eric DongandGitHub 6289d1f0a6 feat: exclude model garden dockerfilers from dependabot checks (#3712) 2024-11-15 15:05:21 +00:00
Eric DongandGitHub 1301bd2ad2 Revert "Bump deepspeed (#3658)" (#3711)
This reverts commit 6c9bdba210.
2024-11-14 16:51:13 +00:00
Vertex MG TeamandCopybara-Service 21d5f7edcb Download only the required files notebook
PiperOrigin-RevId: 695745528
2024-11-12 08:31:36 -08:00
Vertex MG TeamandCopybara-Service 7f13964632 Update the notebook to remove the duplicated model agreement step.
PiperOrigin-RevId: 695575396
2024-11-11 20:23:02 -08:00
Vertex MG TeamandCopybara-Service 231dbae9e9 Update the recursion mae local inference notebook, to always download the model weights from HuggingFace Hub, instead of directly load the model weights from HF. The latter, for some reason, cannot find the model.safetensors file from the repo.
PiperOrigin-RevId: 695573672
2024-11-11 20:13:58 -08:00
Vertex MG TeamandCopybara-Service bf0507bfda OWL-ViT2 Notebook
PiperOrigin-RevId: 695315779
2024-11-11 06:46:38 -08:00
Mend RenovateandGitHub 3361dc70d1 chore(deps): update dependency nbqa to v1.9.1 (#3699) 2024-11-11 13:29:15 +00:00
Mend RenovateandGitHub e05c83832d chore(config): migrate config renovate.json (#3677) 2024-11-11 13:28:31 +00:00
Aaron DietzandGitHub e2f1aaed9d Update feature_store_streaming_ingestion_sdk.ipynb (#3694)
Moved mention of Feature Store (Legacy) so that it shows up in our notebook description when we generate the list of notebooks.
2024-11-11 13:26:54 +00:00
5d71616e8b Add ability to copy specific model artifacts (#3700)
Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2024-11-11 13:25:24 +00:00
57f598052e Update usage tracking metrics in the finetuning notebooks (#3695)
Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2024-11-11 13:24:27 +00:00
Vertex MG TeamandCopybara-Service 439a686f16 Create a notebook for the image_feature_extraction_mae model.
PiperOrigin-RevId: 694586865
2024-11-08 11:54:42 -08:00
Vertex MG TeamandCopybara-Service b99a1d8f42 Create a notebook for local inference for the partner recursion mae model.
PiperOrigin-RevId: 694248672
2024-11-07 14:28:31 -08:00
Vertex MG TeamandCopybara-Service 2d1339731f Present --data-parallel-size for Hex-LLM deployment; add advanced config for CodeGemma.
PiperOrigin-RevId: 693598038
2024-11-05 22:48:19 -08:00
Vertex MG TeamandCopybara-Service c71edd71e7 OWL-ViT Notebook
PiperOrigin-RevId: 693140455
2024-11-04 17:09:00 -08:00
Vertex MG TeamandCopybara-Service e0d078b6bb The HF TGI notebook should use the TGI 2.3 serving container, which is the latest
PiperOrigin-RevId: 693031093
2024-11-04 11:15:58 -08:00
Sujit KhasnisandGitHub 267b45f6aa feat:claude notebook 3.5 update (#3690) 2024-11-04 18:43:21 +00:00
Vertex MG TeamandCopybara-Service f1bb2f7aea Stable Diffusion 2.1 Dreambooth Finetuning notebook
PiperOrigin-RevId: 691686520
2024-10-30 23:24:39 -07:00
Vertex MG TeamandCopybara-Service c2132249c9 Fix Llama 3.1 deployment notebook vLLM version.
PiperOrigin-RevId: 691669244
2024-10-30 21:58:30 -07:00
Vertex MG TeamandCopybara-Service d4fdc15892 Update vLLM version and embedded links in Llama 3.2 deployment notebook.
PiperOrigin-RevId: 691604159
2024-10-30 17:13:44 -07:00
Sujit KhasnisandGitHub 43bc13ee66 feat: Claude region update(euw1) (#3685) 2024-10-30 16:52:16 +00:00
f11502ec80 Fix pip dependency error (#3683)
Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2024-10-30 12:38:10 +00:00
Vertex MG TeamandCopybara-Service 7f7cf51dd7 Enable dedicate endpoint for Mistral deployment
PiperOrigin-RevId: 691250786
2024-10-29 19:44:48 -07:00
Vertex MG TeamandCopybara-Service f13e086012 Add optimized vLLM to Mixtral deployment notebook.
PiperOrigin-RevId: 691209007
2024-10-29 16:52:16 -07:00
Vertex MG TeamandCopybara-Service b995ea7adf Add optimized vLLM to Llama 3.1 deployment notebook.
PiperOrigin-RevId: 691207909
2024-10-29 16:48:35 -07:00
Vertex MG TeamandCopybara-Service b183a3d796 Update model_garden_pytorch_sd_xl_finetuning_dreambooth_lora.ipynb
PiperOrigin-RevId: 689953226
2024-10-25 16:47:05 -07:00
Vertex MG TeamandCopybara-Service e500467637 Update Mediapipe Image Generation notebook
PiperOrigin-RevId: 689923141
2024-10-25 14:59:38 -07:00
Vertex MG TeamandCopybara-Service ad6035ae80 No public description
PiperOrigin-RevId: 688364622
2024-10-25 14:59:28 -07:00
Ivan NardiniandGitHub 731bc7a780 fix: Adding autoscaling to RoV cluster management notebook (#3674)
* new notebook

* fix issue

* linter passed

* change pname

* linter passed

* remove script
2024-10-25 17:28:47 +00:00
lee1premiumandGitHub 6dd9c27781 feat: Getting tuned embeddings using text-embedding-005. (#3673)
* feat: Getting tuned embeddings using text-embedding-005.

* feat: Getting tuned embeddings using text-embedding-005.

* feat: Getting tuned embeddings using text-embedding-005.
2024-10-24 00:49:45 +00:00
Mend RenovateandGitHub 830dc9f1c3 chore(deps): update dependency pyupgrade to v3.19.0 (#3667) 2024-10-23 12:55:43 +00:00
lee1premiumandGitHub 23a1398504 feat: Getting embeddings using text-embedding-005. (#3672)
* feat: Getting embeddings using text-embedding-005.

* feat: Getting embeddings using text-embedding-005.

* feat: Getting embeddings using text-embedding-005.

* feat: Getting embeddings using text-embedding-005.
2024-10-23 12:55:15 +00:00
Sujit KhasnisandGitHub 7b86dec7f7 fix: update titles, headers (#3670) 2024-10-22 20:24:40 +00:00
Sujit KhasnisandGitHub 02f1edd942 fix: model ordering in dropdown (#3669) 2024-10-22 16:15:10 +00:00
Sujit KhasnisandGitHub b8a3fafb9a feat: Claude notebook update (#3668) 2024-10-22 15:51:51 +00:00
Tianrui YangandGitHub 26849e7f20 fea: add PSC example code in Feature Store embedding notebook (#3663) 2024-10-22 14:50:02 +00:00
Shawn YangandCopybara-Service a96c2926af fix: Fix Notebook format issue in Preview by removing output.
PiperOrigin-RevId: 688364596
2024-10-21 19:54:06 -07:00
Shawn YangandCopybara-Service 90eaf5c7fe fix: Fix Notebook format issue in Preview.
PiperOrigin-RevId: 688317050
2024-10-21 16:44:52 -07:00
Shawn YangandCopybara-Service d6ad52e2cd feat: Update Reasoning Engine + Llama 3.1 models notebook with Function Calling Agent.
PiperOrigin-RevId: 688245912
2024-10-21 13:09:51 -07:00
Vertex MG TeamandCopybara-Service e24c20b372 Adding Phi-3-medium TPU configs to Colab deployment notebook
PiperOrigin-RevId: 687430227
2024-10-18 14:43:49 -07:00
Vertex MG TeamandCopybara-Service fb1871d100 Update Llama 3.1 MaaS naming.
PiperOrigin-RevId: 687362552
2024-10-18 11:09:18 -07:00
Vertex MG TeamandCopybara-Service a7dd5aa5a9 Adding Phi-3-mini TPU configs to Colab deployment notebook
PiperOrigin-RevId: 687333561
2024-10-18 09:40:46 -07:00
Vertex MG TeamandCopybara-Service 307f1a8d41 Reformat Instant ID notebook
PiperOrigin-RevId: 687309283
2024-10-18 08:20:02 -07:00
Vertex MG TeamandCopybara-Service 53895497c0 Update chat completions URL for dedicated endpoint in Gemma and Llama deployment notebooks.
PiperOrigin-RevId: 687307214
2024-10-18 08:11:33 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
6c9bdba210 Bump deepspeed (#3658)
Bumps [deepspeed](https://github.com/microsoft/DeepSpeed) from 0.14.4 to 0.15.1.
- [Release notes](https://github.com/microsoft/DeepSpeed/releases)
- [Commits](https://github.com/microsoft/DeepSpeed/compare/v0.14.4...v0.15.1)

---
updated-dependencies:
- dependency-name: deepspeed
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2024-10-18 14:31:08 +00:00
Vertex MG TeamandCopybara-Service 133395a908 Use /tmp as the dataset directory since the notebook executor environment may not use /content as the base directory
PiperOrigin-RevId: 687194370
2024-10-18 00:25:47 -07:00
Vertex MG TeamandCopybara-Service bd29537128 Remove Weaviate and Pinecone notebooks as they are relocated to another directory.
PiperOrigin-RevId: 686788959
2024-10-16 23:39:33 -07:00
Vertex MG TeamandCopybara-Service 3ac1c34633 Reformat stable diffusion gradio notebook
PiperOrigin-RevId: 686463355
2024-10-16 04:53:50 -07:00
Vertex MG TeamandCopybara-Service 2e56046d25 Add Llama 3.2 evaluation notebook.
PiperOrigin-RevId: 686349456
2024-10-15 20:51:39 -07:00
Vertex MG TeamandCopybara-Service f6dd048f98 Update Gemma PEFT notebook to use the new training docker image and add more instructions.
PiperOrigin-RevId: 686347835
2024-10-15 20:44:39 -07:00
Vertex MG TeamandCopybara-Service 9da812bf2e Update model_garden_pytorch_bart_large_cnn.ipynb
PiperOrigin-RevId: 686230074
2024-10-15 14:00:45 -07:00
Sujit KhasnisandGitHub 11d361f727 refactor: verbiage updates, minor code updates (#3655) 2024-10-15 18:13:14 +00:00
Mend RenovateandGitHub de9dd0f850 Update python Docker tag to v3.13 (#3626) 2024-10-15 13:49:03 +00:00
482170765andGitHub f20e9590a5 Update a custom job example (#3639)
* This is a custom job example of kfp v2
2024-10-15 13:48:37 +00:00
Mend RenovateandGitHub 5443944739 chore(deps): update dependency pyupgrade to v3.18.0 (#3640) 2024-10-15 13:48:00 +00:00
Sujit KhasnisandGitHub f2ee8a582f feat: NVIDIA NIM on Vertex Ai walkthrough (#3579)
* feat: NVIDIA NIM on Vertex Ai walkthrough

* fix: PR comments resolved

* fix: PR comments resolved#2

* fix: excep handling

* fix: handle if model not uploaded
2024-10-15 13:47:25 +00:00
Vertex MG TeamandCopybara-Service 1d3d3980e4 Fix the issue that the model and endpoint are not stored in the dictionary.
PiperOrigin-RevId: 685966395
2024-10-14 22:50:05 -07:00
Vertex MG TeamandCopybara-Service f9b371e030 Update the notebook to use A100 80GB for vLLM deployment.
PiperOrigin-RevId: 685934295
2024-10-14 20:42:11 -07:00
Dustin LuongandCopybara-Service a1d4026020 Update the notebook to use the Vertex SDK to send requests to the deployed endpoint instead of openai SDK.
PiperOrigin-RevId: 685713862
2024-10-14 08:23:23 -07:00
Vertex MG TeamandCopybara-Service eb166ed3df Update the default model to whisper-large-v3-turbo and add placeholder for language.
PiperOrigin-RevId: 685618090
2024-10-14 01:37:22 -07:00
Vertex MG TeamandCopybara-Service 08278233ff Update Qwen2 deployment notebook to support A100, H100 and A100 80GB and fix lint issues.
PiperOrigin-RevId: 685555814
2024-10-13 20:41:34 -07:00
Vertex MG TeamandCopybara-Service 51bf2a38ff Minor fix
PiperOrigin-RevId: 685020524
2024-10-11 18:22:58 -07:00
Vertex MG TeamandCopybara-Service 2dede7552b Added more details about notebook parameters for predict section. Also gave storage.objectViewer access service account for buckets used for predict section.
PiperOrigin-RevId: 684700814
2024-10-10 22:12:39 -07:00
Vertex MG TeamandCopybara-Service 905e89ad64 Add new notebook for Knowledge Engine API with Pinecone.
PiperOrigin-RevId: 684666941
2024-10-10 19:56:49 -07:00
Vertex MG TeamandCopybara-Service ec297e663c Add labels to finetuning notebooks
PiperOrigin-RevId: 684655746
2024-10-10 19:11:32 -07:00
Vertex MG TeamandCopybara-Service d15caa9ed5 Enable dedicate endpoint for TGI Gemma2 predict and chat completion
PiperOrigin-RevId: 684623757
2024-10-10 16:55:28 -07:00
Vertex MG TeamandCopybara-Service df55634281 Added more details to deployment section of notebook.
PiperOrigin-RevId: 684343445
2024-10-10 01:19:34 -07:00
Vertex MG TeamandCopybara-Service af84cdc46b Fix the error on stable_diffusion_gradio notebook for text2image models.
PiperOrigin-RevId: 684108916
2024-10-09 11:26:08 -07:00
Shawn YangandCopybara-Service e946e2b304 feat: Add Reasoning Engine with Llama 3.1 models notebook.
PiperOrigin-RevId: 684083427
2024-10-09 10:16:47 -07:00
Vertex MG TeamandCopybara-Service d7550f4756 LLaVA Deployment notebook
PiperOrigin-RevId: 683832899
2024-10-08 18:20:17 -07:00
Vertex MG TeamandCopybara-Service 3088fc896d Update vLLM and chat completions prediction samples.
PiperOrigin-RevId: 683698377
2024-10-08 11:25:16 -07:00
Vertex MG TeamandCopybara-Service 75cf2fe9b9 Add Qwen2.5 related updates to the notebook
PiperOrigin-RevId: 683423186
2024-10-07 19:46:47 -07:00
Vertex MG TeamandCopybara-Service d68494bd3b Update Prompt Guard deployment notebook.
PiperOrigin-RevId: 683352100
2024-10-07 15:42:32 -07:00
02f66799c8 Add README.md for PEFT train docker template (#3625)
Co-authored-by: minwoopark <minwoopark@google.com>
2024-10-07 20:54:53 +00:00
482170765andGitHub 6b10fe9636 Upload classifier predictor sample of sklearn (#3623)
Upload a classifier predictor sample using scikit-learn lib.
2024-10-07 20:54:18 +00:00
Mend RenovateandGitHub 4c3693929d Update dependency black to v24.10.0 (#3624) 2024-10-07 20:49:06 +00:00
Vertex MG TeamandCopybara-Service 6bff6af34a Fix lint issues in the notebooks
PiperOrigin-RevId: 682565175
2024-10-04 22:19:13 -07:00
Vertex MG TeamandCopybara-Service 219474b39c Add new notebook for RAG API with Weaviate.
PiperOrigin-RevId: 682497053
2024-10-04 16:58:17 -07:00
Vertex MG TeamandCopybara-Service bc152ece40 Fix the RAG notebook format.
PiperOrigin-RevId: 682352935
2024-10-04 09:42:36 -07:00
Vertex MG TeamandCopybara-Service 761b918de1 Use 'import datetime' instead of 'from datetime import datetime'
PiperOrigin-RevId: 682309409
2024-10-04 07:14:55 -07:00
Vertex MG TeamandCopybara-Service 6a574ea15d Update the docker image for Gemma finetuning.
PiperOrigin-RevId: 682163952
2024-10-03 21:52:34 -07:00
Vertex MG TeamandCopybara-Service c5212982d0 Fix the RAG notebook link.
PiperOrigin-RevId: 682031626
2024-10-03 14:27:27 -07:00
Vertex MG TeamandCopybara-Service 3986128f78 Autogluon notebook
PiperOrigin-RevId: 681881078
2024-10-03 08:05:26 -07:00
Vertex MG TeamandCopybara-Service 575025a1cd Fix format issues
PiperOrigin-RevId: 681739757
2024-10-02 23:33:38 -07:00
Vertex MG TeamandCopybara-Service fe42990c9d Update Llama 2 evaluation notebook.
PiperOrigin-RevId: 681657733
2024-10-02 17:49:18 -07:00
Dustin LuongandCopybara-Service bd3283b2a6 Internal change.
PiperOrigin-RevId: 681516025
2024-10-02 17:49:03 -07:00
24c001eaf6 Update PEFT train docker code (#3613)
Co-authored-by: minwoopark <minwoopark@google.com>
2024-10-02 19:13:30 +00:00
Aaron DietzandGitHub e500e70580 Update spark_on_ray_on_vertex_ai.ipynb (#3605)
Revised the overview so our github notebook list script will pick up the "last line" in the overview and start including this notebook in the output.
2024-10-02 18:13:15 +00:00
Vertex MG TeamandCopybara-Service bfb3775813 Fix the region for llama3 hex-llm chat completion
PiperOrigin-RevId: 681511627
2024-10-02 10:47:45 -07:00
Vertex MG TeamandCopybara-Service 538432df5c Fix the region for llama3 hex-llm chat completion
PiperOrigin-RevId: 681167715
2024-10-01 14:33:19 -07:00
Minwoo ParkandCopybara-Service 7764173895 Minor copyright year update.
PiperOrigin-RevId: 681141148
2024-10-01 13:21:14 -07:00
Vertex MG TeamandCopybara-Service 142237c1f5 Add region suggestion for A100_80GB and H100_80GB to notebooks.
PiperOrigin-RevId: 681090786
2024-10-01 11:07:54 -07:00
Vertex MG TeamandCopybara-Service 791a66ec54 Fix formatting issue
PiperOrigin-RevId: 681076924
2024-10-01 10:35:31 -07:00
Vertex MG TeamandCopybara-Service 8379361eff Fix typo - Use "prompt" in image captioning sample request.
PiperOrigin-RevId: 681074023
2024-10-01 10:27:17 -07:00
Vertex MG TeamandCopybara-Service a9f6b2d7f0 Adding Phi-3.5-MoE-instruct variant to Phi-3 deployment notebook.
PiperOrigin-RevId: 681057221
2024-10-01 09:44:32 -07:00
Vertex MG TeamandCopybara-Service 8f077f9bc2 This notebook demonstrates deploying prebuilt Whisper Large models.
PiperOrigin-RevId: 681039051
2024-10-01 08:55:51 -07:00
Sujit KhasnisandGitHub c5ec7e5d83 feat: mistral ai sdk support for vertexai (#3601)
* feat: mistral ai sdk support fro vertexai

* feat: mistral ai sdk support fro vertexai, token fix

* feat: mistral ai sdk support fro vertexai, excep handling
2024-10-01 14:59:10 +00:00
Vertex MG TeamandCopybara-Service daf92aa672 Fix lint issue
PiperOrigin-RevId: 680835893
2024-09-30 20:57:00 -07:00
Bhaskar GoyalandGitHub 4070d8b0e5 Update region for Claude Haiku and Sonnet 3.5 (#3606) 2024-09-30 23:06:35 +00:00
Vertex MG TeamandCopybara-Service a31f1e037e Use dedicated endpoint as default for Gemma deployment on vertex
PiperOrigin-RevId: 680645756
2024-09-30 11:09:15 -07:00
sharkeshdandGitHub dc9d1032a2 Update Dockerfile (#3576)
This multi-stage approach keeps your final image clean and lightweight.
2024-09-30 17:45:37 +00:00
Aaron DietzandGitHub d5f93bf8a9 Update spark_on_ray_on_vertex_ai.ipynb (#3594)
Updated branding/name of Vertex AI Workbench, added "Overview" heading.

Why? Not having an "Overview" heading prevents this notebook from getting picked up in our notebook list output.
2024-09-30 17:45:06 +00:00
Aaron DietzandGitHub 6a3bc32e45 Update xai_text_classification_feature_attributions.ipynb (#3596)
Removed bolding that doesn't render properly when we port the content to our docs
2024-09-30 17:44:09 +00:00
Aaron DietzandGitHub 1d1d4a586d Update get_started_with_model_monitoring_setup.ipynb (#3597)
Removed bolding that doesn't render properly when we port the content to our docs
2024-09-30 17:43:26 +00:00
Aaron DietzandGitHub b3c7ecb8cd Update hyperparameter_tuning_xgboost.ipynb (#3598)
Removed bolding that doesn't render properly when we port the content to our docs
2024-09-30 17:42:47 +00:00
Aaron DietzandGitHub cec4e447ab Update chicago_taxi_fare_prediction.ipynb (#3599)
Removed bolding that doesn't render properly when we port the content to our docs
2024-09-30 17:41:55 +00:00
Vertex MG TeamandCopybara-Service e61be64040 Use dedicated endpoint as default for Gemma2 deployment on vertex
PiperOrigin-RevId: 680634868
2024-09-30 10:41:39 -07:00
Minwoo ParkandCopybara-Service 148a6fad99 Add instruction to run TensorBoard in Cloud Shell
PiperOrigin-RevId: 679746165
2024-09-27 15:20:04 -07:00
Aaron DietzandGitHub ca84581ed0 Fixed missing word in predictive_maintenance_usecase.ipynb (#3581)
Added a word to make a sentence parse correctly.
2024-09-26 20:49:03 +00:00
Vertex MG TeamandCopybara-Service 4d1c59cba4 Minor fix in llama3.2 notebook
PiperOrigin-RevId: 679190955
2024-09-26 09:59:49 -07:00
Vertex MG TeamandCopybara-Service 5cd0c0e782 Minor change to the VOT and ZipNeRF notebooks.
PiperOrigin-RevId: 679159158
2024-09-26 08:27:19 -07:00
Vertex MG TeamandCopybara-Service 402231e2b6 Update finetuning notebook with stable_20240909 training image
PiperOrigin-RevId: 679139242
2024-09-26 07:25:59 -07:00
Vertex MG TeamandCopybara-Service 38e1a46a7b Fix typo in vllm args
PiperOrigin-RevId: 679131835
2024-09-26 06:57:15 -07:00
0727e19520 Add vmg templates, dataset_validation_util and update common_util (#3586)
* Add vmg templates, dataset_validation_util and update common_util

* Add name to CODEOWNERS

* Update common_util.py

---------

Co-authored-by: Rayan Dasoriya <dasoriya@google.com>
2024-09-26 13:16:09 +00:00
Vertex MG TeamandCopybara-Service 30c3e627a7 Update the RAG notebook for Llama3 models.
PiperOrigin-RevId: 679007617
2024-09-25 23:31:06 -07:00
Vertex MG TeamandCopybara-Service bdea63ec41 Adding Phi-3.5-mini-instruct variant to Phi-3 deployment notebook.
PiperOrigin-RevId: 678888299
2024-09-25 16:24:22 -07:00
Vertex MG TeamandCopybara-Service 2e5410fe35 No public description
PiperOrigin-RevId: 678885403
2024-09-25 16:15:39 -07:00
Vertex MG TeamandCopybara-Service 6f9813cba5 Update sample requests in Llama 3.2 OpenAI MaaS notebook.
PiperOrigin-RevId: 678828203
2024-09-25 13:35:19 -07:00
Vertex MG TeamandCopybara-Service 115d8f991a Update Llama 3.2 OpenAI MaaS notebook.
PiperOrigin-RevId: 678781425
2024-09-25 11:31:37 -07:00
Vertex MG TeamandCopybara-Service 7fd31a65ae Support deploying llama 3.2 guard models on model garden.
PiperOrigin-RevId: 678760438
2024-09-25 10:44:40 -07:00
Vertex MG TeamandCopybara-Service 883cc93ab7 Add Llama 3.2 OpenAI MaaS notebook.
PiperOrigin-RevId: 678753976
2024-09-25 10:27:22 -07:00
Changyu ZhuandCopybara-Service 1363868542 Add streaming chat completions example to the HF TGI notebook
PiperOrigin-RevId: 678753372
2024-09-25 10:26:02 -07:00
Vertex MG TeamandCopybara-Service 09122d1479 Support deploying llama 3.2 models on model garden.
PiperOrigin-RevId: 678752790
2024-09-25 10:24:34 -07:00
Minwoo ParkandCopybara-Service 2adb19be7c Resolving conflict
PiperOrigin-RevId: 678701565
2024-09-25 08:00:32 -07:00
Aaron DietzandGitHub c588d81d02 Update notebook_template_review.py (#3578)
Removed "external" class for the "open-notebook-in..." links. Style guide indicates we should avoid using the "external" class.
2024-09-25 12:43:15 +00:00
ShunpeIIIandGitHub e51af44898 Remove the link of the notebook that has been moved to the community. (#3571) 2024-09-23 14:39:40 +00:00
Vertex MG TeamandCopybara-Service 0872ce0a87 Fix the pip install command in the notebooks.
PiperOrigin-RevId: 677767387
2024-09-23 06:31:07 -07:00
Vertex MG TeamandCopybara-Service 867d7b7410 Fix max_context_length in Qwen2 deployment notebook.
PiperOrigin-RevId: 677762785
2024-09-23 06:12:36 -07:00
Vertex MG TeamandCopybara-Service d8e628b3ab Add Flux gradio notebook
PiperOrigin-RevId: 676901084
2024-09-20 10:57:44 -07:00
Vertex MG TeamandCopybara-Service dd4767546e Update gradio notebooks for Instant ID and Stable Diffusion
PiperOrigin-RevId: 676476082
2024-09-19 10:45:16 -07:00
Vertex MG TeamandCopybara-Service d185eff6ca Add region option to model garden notebooks.
PiperOrigin-RevId: 676473760
2024-09-19 10:39:10 -07:00
Vertex MG TeamandCopybara-Service f00c9cdeef Update Instant ID notebook
PiperOrigin-RevId: 676466759
2024-09-19 10:21:26 -07:00
Vertex MG TeamandCopybara-Service 12bd07e8fc Fix use dedicated endpoint parameter type in Pytorch Gemma Serving
PiperOrigin-RevId: 676433256
2024-09-19 08:55:18 -07:00
Vertex MG TeamandCopybara-Service 61d866aef3 Update deploy function.
PiperOrigin-RevId: 676408091
2024-09-19 07:37:53 -07:00
Vertex MG TeamandCopybara-Service bb2c8c502d Fix use dedicated endpoint parameter type in Gemma Serving
PiperOrigin-RevId: 676121299
2024-09-18 14:00:08 -07:00
Vertex MG TeamandCopybara-Service ea7915b38a Use standard id as MODEL_ID.
PiperOrigin-RevId: 675778051
2024-09-17 18:04:39 -07:00
Dustin LuongandGitHub b5ee2ea16b Copy Vertex MG files to notebook folder (#3560) 2024-09-18 00:47:12 +00:00
Vertex MG TeamandCopybara-Service 6daebf78bc Support deploying Hex-LLM on multi-hosts TPU, like v5e-16.
PiperOrigin-RevId: 675736856
2024-09-17 15:42:15 -07:00
Changyu ZhuandCopybara-Service 2752658b6d Minor updates to the MoViNet notebooks
PiperOrigin-RevId: 675723676
2024-09-17 15:03:58 -07:00
Vertex MG TeamandCopybara-Service 55ed17d227 Add dedicated endpoint support to Gemma Serving
PiperOrigin-RevId: 675627850
2024-09-17 10:49:28 -07:00
yutatanamotoandGitHub 3808d495fd fix official sample notebook for vector search (#3500)
* fix restriction declaration (allow_list → allow) for vector search index

* fix folder name in CODEOWNERS (/matching_engine → /vector_search)
2024-09-17 12:45:20 +00:00
Vertex MG TeamandCopybara-Service 6c78212f92 Add enable_model_cpu_offload option for flux example in local inference notebook
PiperOrigin-RevId: 675227512
2024-09-16 11:31:55 -07:00
william-ChengChungChuandGitHub b6771091fa Upload classifier predictor sample of xgboost (#3552)
This example uses aiplatform and xgboost to provide a classifier predictor.
2024-09-16 14:57:00 +00:00
Vertex MG TeamandCopybara-Service 6a62443d01 Stable Diffusion v2.1 notebook
PiperOrigin-RevId: 674233624
2024-09-13 03:52:22 -07:00
Vertex MG TeamandCopybara-Service eb7d8456be Add instructions for applying Llama Guard on MaaS.
PiperOrigin-RevId: 674144050
2024-09-12 22:14:50 -07:00
Changyu ZhuandCopybara-Service df4adf877e internal change
PiperOrigin-RevId: 673999030
2024-09-12 22:14:36 -07:00
Dustin LuongandGitHub 75fe45e07c Copy Vertex MG files to notebook folder (#3542) 2024-09-12 21:20:54 +00:00
986 changed files with 137617 additions and 52657 deletions
@@ -238,7 +238,7 @@ def _get_notebook_python_version(notebook_path: str) -> str:
# Look for the python version specification pattern
re_match = re.search(
"python version = (\d+\.\d+)", markdown, flags=re.IGNORECASE
r"python version = (\d+\.\d+)", markdown, flags=re.IGNORECASE
)
if re_match:
# get the version number
@@ -365,7 +365,7 @@ def process_and_execute_notebook(
# Use gcloud to get tail
try:
result.error_message = subprocess.check_output(
["gsutil", "cat", "-r", "-1000", log_file_uri], encoding="UTF-8"
["gcloud", "storage", "cat", "--range", "-1000", log_file_uri], encoding="UTF-8"
)
except Exception as error:
result.error_message = str(error)
+2 -2
View File
@@ -56,8 +56,8 @@ def execute_notebook(
print("\n=== DOWNLOAD EXECUTED NOTEBOOK ===\n")
print(f"Please debug the executed notebook by downloading the executed notebook:")
print("Option 1. Using gsutil. Run the following command in your terminal.")
print(f'\tgsutil cp "{output_file_or_uri}" .')
print("Option 1. Using gcloud storage. Run the following command in your terminal.")
print(f'\tgcloud storage cp "{output_file_or_uri}" .')
print("Option 2. Using this link.")
print(f"\thttps://storage.googleapis.com/{output_file_or_uri[5:]}")
-2
View File
@@ -1,5 +1,3 @@
notebooks/official/vizier/gapic-vizier-multi-objective-optimization.ipynb
notebooks/official/pipelines/lightweight_functions_component_io_kfp.ipynb
notebooks/official/ml_metadata/sdk-metric-parameter-tracking-for-locally-trained-models.ipynb
notebooks/official/custom/custom-tabular-bq-managed-dataset.ipynb
.cloud-build/tests/python_version_test.ipynb
+1 -1
View File
@@ -108,7 +108,7 @@ class VertexAIInstallProprocessor(Preprocessor):
if "google-cloud-aiplatform" not in content:
return content
return (
f"gsutil cp {self.vertex_ai_wheel} google-cloud-aiplatform.whl\n" +
f"gcloud storage cp {self.vertex_ai_wheel} google-cloud-aiplatform.whl\n" +
content.replace("google-cloud-aiplatform\n", "google-cloud-aiplatform.whl\n")
.replace("google-cloud-aiplatform ", "google-cloud-aiplatform.whl ")
)
+2 -2
View File
@@ -15,7 +15,7 @@ def download_file(bucket_name: str, blob_name: str, destination_file: str) -> st
remote_file_path = "".join(["gs://", "/".join([bucket_name, blob_name])])
subprocess.check_output(
["gsutil", "cp", remote_file_path, destination_file], encoding="UTF-8"
["gcloud", "storage", "cp", remote_file_path, destination_file], encoding="UTF-8"
)
return destination_file
@@ -27,7 +27,7 @@ def upload_file(
) -> str:
"""Copies a local file to a GCS path"""
subprocess.check_output(
["gsutil", "cp", local_file_path, remote_file_path], encoding="UTF-8"
["gcloud", "storage", "cp", local_file_path, remote_file_path], encoding="UTF-8"
)
return remote_file_path
+10
View File
@@ -0,0 +1,10 @@
version: 2
updates:
# Ignore model garden dockerfiles:
- package-ecosystem: "npm"
directory: "/community-content/vertex_model_garden"
schedule:
interval: "monthly"
ignore:
- dependency-name: "*"
+3 -3
View File
@@ -7,11 +7,11 @@ jobs:
runs-on: ubuntu-latest
steps:
- name: Set up Python
uses: actions/setup-python@v5
uses: actions/setup-python@v6
with:
python-version: '3.x'
python-version: '3.12'
- name: Fetch pull request branch
uses: actions/checkout@v4
uses: actions/checkout@v6
with:
fetch-depth: 0
- name: Fetch base main branch
+1 -1
View File
@@ -4,7 +4,7 @@
# 2. To lint specific notebooks:
# docker run -v ${PWD}:/setup/app gcr.io/python-docs-samples-tests/notebook_linter:latest notebooks/1.ipynb notebooks/2.ipynb
FROM python:3.12
FROM python:3.14
WORKDIR setup
+5 -5
View File
@@ -2,9 +2,9 @@ git+https://github.com/tensorflow/docs
ipython
jupyter
nbconvert
black==24.8.0
pyupgrade==3.17.0
isort==5.13.2
flake8==7.1.1
nbqa==1.9.0
black==26.5.1
pyupgrade==3.21.2
isort==8.0.1
flake8==7.3.0
nbqa==1.9.1
+25 -9
View File
@@ -1,12 +1,12 @@
# ![Google Cloud](https://avatars.githubusercontent.com/u/2810941?s=60&v=4) Google Cloud Vertex AI Samples
This repository contains notebooks, code samples, sample apps, and other resources that demonstrate how to use, develop and manage machine learning and generative AI workflows using Google Cloud Vertex AI.
This repository contains notebooks, code samples, sample apps, skills, and other resources that demonstrate how to use, develop and manage machine learning and generative AI workflows using Google Cloud Vertex AI.
## Overview
[Vertex AI](https://cloud.google.com/vertex-ai) is a fully-managed, unified AI development platform for building and using generative AI. This repository is designed to help you get started with Vertex AI. Whether you're new to Vertex AI or an experienced ML practitioner, you'll find valuable resources here.
For more Vertex AI Generative AI notebook samples, please visit the Vertex AI [Generative AI](https://github.com/GoogleCloudPlatform/generative-ai) GitHub repository.
⚠️ For more Vertex AI Generative AI notebook samples, please visit the Vertex AI [Generative AI](https://github.com/GoogleCloudPlatform/generative-ai) GitHub repository.
## Explore, learn and contribute
@@ -16,11 +16,11 @@ You can explore, learn, and contribute to this repository to unleash the full po
Explore this repository, follow the links in the header section of each of the notebooks to -
![Colab](https://cloud.google.com/ml-engine/images/colab-logo-32px.png) Open and run the notebook in [Colab](https://colab.google/)\
![Colab Enterprise](https://cloud.google.com/ml-engine/images/colab-enterprise-logo-32px.png) Open and run the notebook in [Colab Enterprise](https://cloud.google.com/colab/docs/introduction)\
![Workbench](https://lh3.googleusercontent.com/UiNooY4LUgW_oTvpsNhPpQzsstV5W8F7rYgxgGBD85cWJoLmrOzhVs_ksK_vgx40SHs7jCqkTkCk=e14-rj-sc0xffffff-h130-w32) Open and run the notebook in [Vertex AI Workbench](https://cloud.google.com/vertex-ai/docs/workbench/introduction)\
![Github](https://cloud.google.com/ml-engine/images/github-logo-32px.png) View the notebook on Github
- Open and run the notebook in [Colab](https://colab.google/)
- Open and run the notebook in [Colab Enterprise](https://cloud.google.com/colab/docs/introduction)
- Open and run the notebook in [Vertex AI Workbench](https://cloud.google.com/vertex-ai/docs/workbench/introduction)
- View the notebook on Github
### Contribute
See the [Contributing Guide](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/master/CONTRIBUTING.md).
@@ -35,7 +35,7 @@ To get started using Vertex AI, you must have a Google Cloud project.
## Repository structure
```bash
```text
├── notebooks
│ ├── official - Notebooks demonstrating use of each Vertex AI service
│ │ ├── automl
@@ -45,7 +45,23 @@ To get started using Vertex AI, you must have a Google Cloud project.
│ │ ├── model_garden
│ │ ├── ...
├── community-content - Sample code and tutorials contributed by the community
├── docs - Deep-dive documentation and advanced setup guides
└── skills - Suite of AI Agent "Skills" for Vertex AI
├── README.md # Developer guide for Vertex AI skills
├── vertex-ai/ # Primary router for Vertex AI tasks
│ └── SKILL.md # Entry point that routes across capabilities
├── genai-sdk/ # Gemini API usage with Gen AI SDK
│ └── SKILL.md # Guides for Python, JS/TS, Go, Java, C#
├── vertex-deploy/ # Deploying models to Endpoints
│ └── SKILL.md # Commands for open models & custom weights
├── vertex-inference/ # Inferencing with GenAI models
│ └── SKILL.md # Code samples for Gemini and OpenMaaS
└── vertex-tuning/ # Secondary router for model fine-tuning
├── SKILL.md # Router for tuning tasks
├── gemini/ # Fine-tuning first-party Gemini models
│ └── SKILL.md
└── open-model/ # Fine-tuning third-party open models
└── SKILL.md
```
## Examples
+2
View File
@@ -20,6 +20,7 @@
/vertex_model_garden/model_oss/movinet @KCFindstr
/vertex_model_garden/model_oss/data_converter @KCFindstr
/vertex_model_garden/model_oss/peft @weigary
/vertex_model_garden/model_oss/peft/templates @rayandasoriya
/vertex_model_garden/model_oss/lm-evaluation-harness @kathyyu-google
/vertex_model_garden/model_oss/tfvision @dstnluong-google
/vertex_model_garden/model_oss/fvlm @minwoo33park
@@ -28,4 +29,5 @@
/vertex_model_garden/model_oss/vllm @kathyyu-google
/vertex_model_garden/benchmarking_reports @lavraicse
/vertex_model_garden/model_oss/autogluon @lavraicse
/vertex_distributed_training/a3mega/llama-3-8b-nemo-pretraining @mstyer-google @erwinh85 @mchrestkha
@@ -148,7 +148,7 @@ implementation:
# Downloading the model archive from GCS
# TODO: Fix gsutil bugs (requires project ID, has auth issues) and use gsutil instead.
# gsutil cp "$model_archive_uri" "$model_archive_local_path"
# gcloud storage cp "$model_archive_uri" "$model_archive_local_path"
pip install google-cloud-storage
python -c '
import sys
@@ -24,12 +24,12 @@ implementation:
# Checking whether the URI points to a single blob, a directory or a URI pattern
# URI points to a blob when that URI does not end with slash and listing that URI only yields the same URI
if [[ "$uri" != */ ]] && (gsutil ls "$uri" | grep --fixed-strings --line-regexp "$uri"); then
if [[ "$uri" != */ ]] && (gcloud storage ls "$uri" | grep --fixed-strings --line-regexp "$uri"); then
mkdir -p "$(dirname "$output_path")"
gsutil -m cp -r "$uri" "$output_path"
gcloud storage cp --recursive "$uri" "$output_path"
else
mkdir -p "$output_path" # When source path is a directory, gsutil requires the destination to also be a directory
gsutil -m rsync -r "$uri" "$output_path" # gsutil cp has different path handling than Linux cp. It always puts the source directory (name) inside the destination directory. gsutil rsync does not have that problem.
gcloud storage rsync --recursive "$uri" "$output_path" # gsutil cp has different path handling than Linux cp. It always puts the source directory (name) inside the destination directory. gsutil rsync does not have that problem.
fi
- inputValue: GCS path
- outputPath: Data
@@ -1,16 +1,40 @@
FROM pytorch/pytorch:1.8.1-cuda11.1-cudnn8-runtime
# Stage 1: Build Environment
FROM pytorch/pytorch:1.8.1-cuda11.1-cudnn8-runtime AS builder
# Install necessary tools and dependencies
RUN apt-get update && \
apt-get install -y curl gnupg && \
echo "deb [signed-by=/usr/share/keyrings/cloud.google.gpg] http://packages.cloud.google.com/apt cloud-sdk main" | tee -a /etc/apt/sources.list.d/google-cloud-sdk.list && \
curl https://packages.cloud.google.com/apt/doc/apt-key.gpg | apt-key --keyring /usr/share/keyrings/cloud.google.gpg add - && \
curl https://packages.cloud.google.com/apt/doc/apt-key.gpg | apt-key --keyring /usr/share/keyrings/cloud.google.gpg add - && \
apt-get update -y && \
apt-get install google-cloud-sdk -y
apt-get install -y google-cloud-sdk
# Copy application code
COPY . /trainer
# Set working directory
WORKDIR /trainer
RUN pip install -r requirements.txt
# Install Python dependencies
RUN pip install --no-cache-dir -r requirements.txt
ENTRYPOINT ["python", "-m", "task"]
# Stage 2: Runtime Environment
FROM pytorch/pytorch:1.8.1-cuda11.1-cudnn8-runtime
# Install Google Cloud SDK
RUN apt-get update && \
apt-get install -y curl gnupg && \
echo "deb [signed-by=/usr/share/keyrings/cloud.google.gpg] http://packages.cloud.google.com/apt cloud-sdk main" | tee -a /etc/apt/sources.list.d/google-cloud-sdk.list && \
curl https://packages.cloud.google.com/apt/doc/apt-key.gpg | apt-key --keyring /usr/share/keyrings/cloud.google.gpg add - && \
apt-get update -y && \
apt-get install -y google-cloud-sdk && \
apt-get clean && rm -rf /var/lib/apt/lists/*
# Copy from the builder stage
COPY --from=builder /trainer /trainer
# Set working directory
WORKDIR /trainer
# Set the entry point
ENTRYPOINT ["python", "-m", "task"]
@@ -1,3 +1,3 @@
torch==2.2.0
torch==2.13.0
torchvision==0.9.1
tensorboard==2.5.0
@@ -1,3 +1,3 @@
torch==2.2.0
torch==2.13.0
torchvision==0.9.1
tensorboard==2.5.0
@@ -110,7 +110,7 @@
},
"outputs": [],
"source": [
"! gsutil ls $gcs_output_uri_prefix"
"! gcloud storage ls $gcs_output_uri_prefix"
]
},
{
@@ -192,7 +192,7 @@
},
"outputs": [],
"source": [
"! gsutil cp -r $gcs_output_uri_prefix/model ./model_server/"
"! gcloud storage cp --recursive $gcs_output_uri_prefix/model ./model_server/"
]
},
{
@@ -556,7 +556,7 @@
},
"outputs": [],
"source": [
"! gsutil rm -rf $gcs_output_uri_prefix"
"! gcloud storage rm --recursive --continue-on-error $gcs_output_uri_prefix"
]
},
{
@@ -412,7 +412,7 @@
},
"outputs": [],
"source": [
"! gsutil ls $gcs_output_uri_prefix"
"! gcloud storage ls $gcs_output_uri_prefix"
]
}
],
@@ -77,4 +77,4 @@ echo "After the job is completed successfully, model files will be saved at $JOB
# # Verify the model was exported
# echo "Verify the model was exported:"
# gsutil ls ${JOB_DIR}/
# gcloud storage ls ${JOB_DIR}/
@@ -34,4 +34,4 @@ RUN echo "service_envelope=json\n" "inference_address=http://0.0.0.0:${AIP_H
USER model-server
# run Torchserve HTTP serve to respond to prediction requests
CMD ["echo", "AIP_STORAGE_URI=${AIP_STORAGE_URI}", ";", "gsutil", "cp", "-r", "${AIP_STORAGE_URI}/${MODEL_NAME}.mar", "/home/model-server/model-store/", ";", "ls", "-ltr", "/home/model-server/model-store/", ";", "torchserve", "--start", "--ts-config=/home/model-server/config.properties", "--models", "${MODEL_NAME}=${MODEL_NAME}.mar", "--model-store", "/home/model-server/model-store"]
CMD ["echo", "AIP_STORAGE_URI=${AIP_STORAGE_URI}", ";", "gcloud", "storage", "cp", "--recursive", "${AIP_STORAGE_URI}/${MODEL_NAME}.mar", "/home/model-server/model-store/", ";", "ls", "-ltr", "/home/model-server/model-store/", ";", "torchserve", "--start", "--ts-config=/home/model-server/config.properties", "--models", "${MODEL_NAME}=${MODEL_NAME}.mar", "--model-store", "/home/model-server/model-store"]
@@ -67,4 +67,4 @@ echo "After the job is completed successfully, model files will be saved at $JOB
# # Verify the model was exported
# echo "Verify the model was exported:"
# gsutil ls ${JOB_DIR}/
# gcloud storage ls ${JOB_DIR}/
@@ -478,8 +478,7 @@
},
"outputs": [],
"source": [
"! gsutil mb -l $REGION $BUCKET_NAME"
]
"! gcloud storage buckets create --location $REGION $BUCKET_NAME" ]
},
{
"cell_type": "markdown",
@@ -498,8 +497,7 @@
},
"outputs": [],
"source": [
"! gsutil ls -al $BUCKET_NAME"
]
"! gcloud storage ls --all-versions --long $BUCKET_NAME" ]
},
{
"cell_type": "markdown",
@@ -582,8 +580,7 @@
"outputs": [],
"source": [
"# Download the sample data into your RAW_DATA_PATH\n",
"! gsutil cp \"gs://cloud-samples-data/vertex-ai/community-content/tf_agents_bandits_movie_recommendation_with_kfp_and_vertex_sdk/u.data\" $RAW_DATA_PATH"
]
"! gcloud storage cp \"gs://cloud-samples-data/vertex-ai/community-content/tf_agents_bandits_movie_recommendation_with_kfp_and_vertex_sdk/u.data\" $RAW_DATA_PATH" ]
},
{
"cell_type": "code",
@@ -1621,9 +1618,7 @@
"! gcloud scheduler jobs delete $SIMULATOR_SCHEDULER_JOB --quiet\n",
"\n",
"# Delete Cloud Storage objects that were created.\n",
"! gsutil -m rm -r $PIPELINE_ROOT\n",
"! gsutil -m rm -r $TRAINING_ARTIFACTS_DIR"
]
"! gcloud storage rm --recursive $PIPELINE_ROOT\n", "! gcloud storage rm --recursive $TRAINING_ARTIFACTS_DIR" ]
}
],
"metadata": {
@@ -1,4 +1,4 @@
google-cloud-bigquery==2.20.0
tensorflow==2.12.1
pillow==10.3.0
pillow==12.3.0
tf-agents==0.8.0
@@ -1,4 +1,4 @@
google-cloud-pubsub==2.5.0
pillow==10.3.0
pillow==12.3.0
tf-agents==0.8.0
tensorflow==2.12.1
@@ -1,5 +1,5 @@
dataclasses==0.6
google-cloud-aiplatform==1.8.1
tensorflow==2.12.1
pillow==10.3.0
pillow==12.3.0
tf-agents==0.8.0
@@ -398,6 +398,7 @@
"if not IS_GOOGLE_CLOUD_NOTEBOOK:\n",
" if \"google.colab\" in sys.modules:\n",
" from google.colab import auth as google_auth\n",
"\n",
" google_auth.authenticate_user()\n",
"\n",
" # If you are running this notebook locally, replace the string below with the\n",
@@ -472,7 +473,7 @@
},
"outputs": [],
"source": [
"! gsutil mb -l $REGION $BUCKET_NAME"
"! gcloud storage buckets create --location $REGION $BUCKET_NAME"
]
},
{
@@ -492,7 +493,7 @@
},
"outputs": [],
"source": [
"! gsutil ls -al $BUCKET_NAME"
"! gcloud storage ls --all-versions --long $BUCKET_NAME"
]
},
{
@@ -565,7 +566,7 @@
"outputs": [],
"source": [
"# Copy the sample data into your DATA_PATH\n",
"! gsutil cp \"gs://cloud-samples-data/vertex-ai/community-content/tf_agents_bandits_movie_recommendation_with_kfp_and_vertex_sdk/u.data\" $DATA_PATH"
"! gcloud storage cp \"gs://cloud-samples-data/vertex-ai/community-content/tf_agents_bandits_movie_recommendation_with_kfp_and_vertex_sdk/u.data\" $DATA_PATH"
]
},
{
@@ -579,11 +580,15 @@
"# Set hyperparameters.\n",
"BATCH_SIZE = 8 # @param {type:\"integer\"} Training and prediction batch size.\n",
"TRAINING_LOOPS = 5 # @param {type:\"integer\"} Number of training iterations.\n",
"STEPS_PER_LOOP = 2 # @param {type:\"integer\"} Number of driver steps per training iteration.\n",
"STEPS_PER_LOOP = (\n",
" 2 # @param {type:\"integer\"} Number of driver steps per training iteration.\n",
")\n",
"\n",
"# Set MovieLens simulation environment parameters.\n",
"RANK_K = 20 # @param {type:\"integer\"} Rank for matrix factorization in the MovieLens environment; also the observation dimension.\n",
"NUM_ACTIONS = 20 # @param {type:\"integer\"} Number of actions (movie items) to choose from.\n",
"NUM_ACTIONS = (\n",
" 20 # @param {type:\"integer\"} Number of actions (movie items) to choose from.\n",
")\n",
"PER_ARM = False # Use the non-per-arm version of the MovieLens environment.\n",
"\n",
"# Set agent parameters.\n",
@@ -621,7 +626,8 @@
"source": [
"# Define RL environment.\n",
"env = movielens_py_environment.MovieLensPyEnvironment(\n",
" DATA_PATH, RANK_K, BATCH_SIZE, num_movies=NUM_ACTIONS, csv_delimiter=\"\\t\")\n",
" DATA_PATH, RANK_K, BATCH_SIZE, num_movies=NUM_ACTIONS, csv_delimiter=\"\\t\"\n",
")\n",
"environment = tf_py_environment.TFPyEnvironment(env)\n",
"\n",
"# Define RL agent/algorithm.\n",
@@ -631,7 +637,8 @@
" tikhonov_weight=TIKHONOV_WEIGHT,\n",
" alpha=AGENT_ALPHA,\n",
" dtype=tf.float32,\n",
" accepts_per_arm_features=PER_ARM)\n",
" accepts_per_arm_features=PER_ARM,\n",
")\n",
"print(\"TimeStep Spec (for each batch):\\n\", agent.time_step_spec, \"\\n\")\n",
"print(\"Action Spec (for each batch):\\n\", agent.action_spec, \"\\n\")\n",
"print(\"Reward Spec (for each batch):\\n\", environment.reward_spec(), \"\\n\")\n",
@@ -639,7 +646,8 @@
"# Define RL metric.\n",
"optimal_reward_fn = functools.partial(\n",
" environment_utilities.compute_optimal_reward_with_movielens_environment,\n",
" environment=environment)\n",
" environment=environment,\n",
")\n",
"regret_metric = tf_bandit_metrics.RegretMetric(optimal_reward_fn)\n",
"metrics = [regret_metric]"
]
@@ -704,35 +712,38 @@
" if training_data_spec_transformation_fn is None:\n",
" data_spec = agent.policy.trajectory_spec\n",
" else:\n",
" data_spec = training_data_spec_transformation_fn(\n",
" agent.policy.trajectory_spec)\n",
" replay_buffer = trainer.get_replay_buffer(data_spec, environment.batch_size,\n",
" steps_per_loop)\n",
" data_spec = training_data_spec_transformation_fn(agent.policy.trajectory_spec)\n",
" replay_buffer = trainer.get_replay_buffer(\n",
" data_spec, environment.batch_size, steps_per_loop\n",
" )\n",
"\n",
" # `step_metric` records the number of individual rounds of bandit interaction;\n",
" # that is, (number of trajectories) * batch_size.\n",
" step_metric = tf_metrics.EnvironmentSteps()\n",
" metrics = [\n",
" tf_metrics.NumberOfEpisodes(),\n",
" tf_metrics.AverageEpisodeLengthMetric(batch_size=environment.batch_size)\n",
" tf_metrics.AverageEpisodeLengthMetric(batch_size=environment.batch_size),\n",
" ]\n",
" if additional_metrics:\n",
" metrics += additional_metrics\n",
"\n",
" if isinstance(environment.reward_spec(), dict):\n",
" metrics += [tf_metrics.AverageReturnMultiMetric(\n",
" reward_spec=environment.reward_spec(),\n",
" batch_size=environment.batch_size)]\n",
" else:\n",
" metrics += [\n",
" tf_metrics.AverageReturnMetric(batch_size=environment.batch_size)]\n",
" tf_metrics.AverageReturnMultiMetric(\n",
" reward_spec=environment.reward_spec(), batch_size=environment.batch_size\n",
" )\n",
" ]\n",
" else:\n",
" metrics += [tf_metrics.AverageReturnMetric(batch_size=environment.batch_size)]\n",
"\n",
" # Store intermediate metric results, indexed by metric names.\n",
" metric_results = defaultdict(list)\n",
"\n",
" if training_data_spec_transformation_fn is not None:\n",
" def add_batch_fn(data): return replay_buffer.add_batch(training_data_spec_transformation_fn(data)) \n",
" \n",
"\n",
" def add_batch_fn(data):\n",
" return replay_buffer.add_batch(training_data_spec_transformation_fn(data))\n",
"\n",
" else:\n",
" add_batch_fn = replay_buffer.add_batch\n",
"\n",
@@ -742,10 +753,12 @@
" env=environment,\n",
" policy=agent.collect_policy,\n",
" num_steps=steps_per_loop * environment.batch_size,\n",
" observers=observers)\n",
" observers=observers,\n",
" )\n",
"\n",
" training_loop = trainer.get_training_loop_fn(\n",
" driver, replay_buffer, agent, steps_per_loop)\n",
" driver, replay_buffer, agent, steps_per_loop\n",
" )\n",
" saver = policy_saver.PolicySaver(agent.policy)\n",
"\n",
" for _ in range(training_loops):\n",
@@ -783,7 +796,8 @@
" environment=environment,\n",
" training_loops=TRAINING_LOOPS,\n",
" steps_per_loop=STEPS_PER_LOOP,\n",
" additional_metrics=metrics)\n",
" additional_metrics=metrics,\n",
")\n",
"\n",
"tf.profiler.experimental.stop()"
]
@@ -1092,11 +1106,15 @@
},
"outputs": [],
"source": [
"RUN_HYPERPARAMETER_TUNING = True # Execute hyperparameter tuning instead of regular training.\n",
"RUN_HYPERPARAMETER_TUNING = (\n",
" True # Execute hyperparameter tuning instead of regular training.\n",
")\n",
"TRAIN_WITH_BEST_HYPERPARAMETERS = False # Do not train.\n",
"\n",
"HPTUNING_RESULT_DIR = \"hptuning/\" # @param {type: \"string\"} Directory to store the best hyperparameter(s) in `BUCKET_NAME` and locally (temporarily).\n",
"HPTUNING_RESULT_PATH = os.path.join(HPTUNING_RESULT_DIR, \"result.json\") # @param {type: \"string\"} Path to the file containing the best hyperparameter(s)."
"HPTUNING_RESULT_PATH = os.path.join(\n",
" HPTUNING_RESULT_DIR, \"result.json\"\n",
") # @param {type: \"string\"} Path to the file containing the best hyperparameter(s)."
]
},
{
@@ -1124,7 +1142,7 @@
" image_uri: str,\n",
" args: List[str],\n",
" location: str = \"us-central1\",\n",
" api_endpoint: str = \"us-central1-aiplatform.googleapis.com\"\n",
" api_endpoint: str = \"us-central1-aiplatform.googleapis.com\",\n",
") -> None:\n",
" \"\"\"Creates a hyperparameter tuning job using a custom container.\n",
"\n",
@@ -1197,8 +1215,8 @@
"\n",
" # Create job\n",
" response = client.create_hyperparameter_tuning_job(\n",
" parent=parent,\n",
" hyperparameter_tuning_job=hyperparameter_tuning_job)\n",
" parent=parent, hyperparameter_tuning_job=hyperparameter_tuning_job\n",
" )\n",
" job_id = response.name.split(\"/\")[-1]\n",
" print(\"Job ID:\", job_id)\n",
" print(\"Job config:\", response)\n",
@@ -1242,7 +1260,8 @@
" image_uri=f\"gcr.io/{PROJECT_ID}/{HPTUNING_TRAINING_CONTAINER}:latest\",\n",
" args=args,\n",
" location=REGION,\n",
" api_endpoint=f\"{REGION}-aiplatform.googleapis.com\")"
" api_endpoint=f\"{REGION}-aiplatform.googleapis.com\",\n",
")"
]
},
{
@@ -1292,7 +1311,8 @@
" name = client.hyperparameter_tuning_job_path(\n",
" project=project,\n",
" location=location,\n",
" hyperparameter_tuning_job=hyperparameter_tuning_job_id)\n",
" hyperparameter_tuning_job=hyperparameter_tuning_job_id,\n",
" )\n",
" response = client.get_hyperparameter_tuning_job(name=name)\n",
" return response"
]
@@ -1313,7 +1333,8 @@
" location=REGION,\n",
" api_endpoint=f\"{REGION}-aiplatform.googleapis.com\")\n",
" if response.state.name == 'JOB_STATE_SUCCEEDED':\n",
" print(\"Job succeeded.\\nJob Time:\", response.update_time - response.create_time)\n",
" print(\"Job succeeded.\n",
"Job Time:\", response.update_time - response.create_time)\n",
" trials = response.trials\n",
" print(\"Trials:\", trials)\n",
" break\n",
@@ -1348,8 +1369,8 @@
"if trials:\n",
" # Dict mapping from metric names to the best metric values seen so far\n",
" best_objective_values = dict.fromkeys(\n",
" [metric.metric_id for metric in trials[0].final_measurement.metrics],\n",
" -np.inf)\n",
" [metric.metric_id for metric in trials[0].final_measurement.metrics], -np.inf\n",
" )\n",
" # Dict mapping from metric names to a list of the best combination(s) of\n",
" # hyperparameter(s). Each combination is a dict mapping from hyperparameter\n",
" # names to their values.\n",
@@ -1358,12 +1379,13 @@
" # `final_measurement` and `parameters` are `RepeatedComposite` objects.\n",
" # Reference the structure above to extract the value of your interest.\n",
" for metric in trial.final_measurement.metrics:\n",
" params = {\n",
" param.parameter_id: param.value for param in trial.parameters}\n",
" params = {param.parameter_id: param.value for param in trial.parameters}\n",
" if metric.value > best_objective_values[metric.metric_id]:\n",
" best_params[metric.metric_id] = [params]\n",
" elif metric.value == best_objective_values[metric.metric_id]:\n",
" best_params[param.parameter_id].append(params) # Handle cases where multiple hyperparameter values lead to the same performance.\n",
" best_params[param.parameter_id].append(\n",
" params\n",
" ) # Handle cases where multiple hyperparameter values lead to the same performance.\n",
" print(\"Best hyperparameter value(s):\")\n",
" for metric, params in best_params.items():\n",
" print(f\"Metric={metric}: {sorted(params)}\")\n",
@@ -1443,7 +1465,9 @@
},
"outputs": [],
"source": [
"PREDICTION_CONTAINER = \"prediction-custom-container\" # @param {type:\"string\"} Name of the container image."
"PREDICTION_CONTAINER = (\n",
" \"prediction-custom-container\" # @param {type:\"string\"} Name of the container image.\n",
")"
]
},
{
@@ -1475,7 +1499,7 @@
" machineType: 'E2_HIGHCPU_8'\"\"\".format(\n",
" PROJECT_ID=PROJECT_ID,\n",
" PREDICTION_CONTAINER=PREDICTION_CONTAINER,\n",
" ARTIFACTS_DIR=ARTIFACTS_DIR\n",
" ARTIFACTS_DIR=ARTIFACTS_DIR,\n",
")\n",
"\n",
"with open(\"cloudbuild.yaml\", \"w\") as fp:\n",
@@ -1592,8 +1616,12 @@
},
"outputs": [],
"source": [
"RUN_HYPERPARAMETER_TUNING = False # Execute regular training instead of hyperparameter tuning.\n",
"TRAIN_WITH_BEST_HYPERPARAMETERS = True # @param {type:\"bool\"} Whether to use learned hyperparameters in training."
"RUN_HYPERPARAMETER_TUNING = (\n",
" False # Execute regular training instead of hyperparameter tuning.\n",
")\n",
"TRAIN_WITH_BEST_HYPERPARAMETERS = (\n",
" True # @param {type:\"bool\"} Whether to use learned hyperparameters in training.\n",
")"
]
},
{
@@ -1633,10 +1661,12 @@
"job = aiplatform.CustomContainerTrainingJob(\n",
" display_name=\"train-movielens\",\n",
" container_uri=f\"gcr.io/{PROJECT_ID}/{HPTUNING_TRAINING_CONTAINER}:latest\",\n",
" command=[\"python3\", \"-m\", \"src.training.task\"] + args, # Pass in training arguments, including hyperparameters.\n",
" command=[\"python3\", \"-m\", \"src.training.task\"]\n",
" + args, # Pass in training arguments, including hyperparameters.\n",
" model_serving_container_image_uri=f\"gcr.io/{PROJECT_ID}/{PREDICTION_CONTAINER}:latest\",\n",
" model_serving_container_predict_route=\"/predict\",\n",
" model_serving_container_health_route=\"/health\")\n",
" model_serving_container_health_route=\"/health\",\n",
")\n",
"\n",
"print(\"Training Spec:\", job._managed_model)\n",
"\n",
@@ -1645,7 +1675,8 @@
" replica_count=1,\n",
" machine_type=\"n1-standard-4\",\n",
" accelerator_type=\"ACCELERATOR_TYPE_UNSPECIFIED\",\n",
" accelerator_count=0)"
" accelerator_count=0,\n",
")"
]
},
{
@@ -1784,7 +1815,7 @@
"! gcloud ai models delete $model.name --quiet\n",
"\n",
"# Delete Cloud Storage objects that were created\n",
"! gsutil -m rm -r $ARTIFACTS_DIR"
"! gcloud storage rm --recursive $ARTIFACTS_DIR"
]
}
],
@@ -324,7 +324,7 @@
},
"outputs": [],
"source": [
"! gsutil ls $gcs_output_uri_prefix"
"! gcloud storage ls $gcs_output_uri_prefix"
]
},
{
@@ -344,7 +344,7 @@
},
"outputs": [],
"source": [
"! gsutil rm -rf $gcs_output_uri_prefix"
"! gcloud storage rm --recursive --continue-on-error $gcs_output_uri_prefix"
]
}
],
@@ -328,7 +328,7 @@
},
"outputs": [],
"source": [
"! gsutil ls $gcs_output_uri_prefix"
"! gcloud storage ls $gcs_output_uri_prefix"
]
},
{
@@ -348,7 +348,7 @@
},
"outputs": [],
"source": [
"! gsutil rm -rf $gcs_output_uri_prefix"
"! gcloud storage rm --recursive --continue-on-error $gcs_output_uri_prefix"
]
}
],
@@ -341,7 +341,7 @@
},
"outputs": [],
"source": [
"! gsutil ls $gcs_output_uri_prefix"
"! gcloud storage ls $gcs_output_uri_prefix"
]
},
{
@@ -361,7 +361,7 @@
},
"outputs": [],
"source": [
"! gsutil rm -rf $gcs_output_uri_prefix"
"! gcloud storage rm --recursive --continue-on-error $gcs_output_uri_prefix"
]
}
],
@@ -0,0 +1,126 @@
# Vertex AI Training: Llama 3.1 8B pre-training using Nvidia A3 Mega VMs (H100)
This document provides a step-by-step guide for pre-training a Llama 3.1 8B model on the `en-wiki` dataset using multiple [Vertex AI Custom Training](https://cloud.google.com/vertex-ai/docs/training/overview) `a3-megagpu-8g` nodes.
We will use a custom container based on NVIDIA's [NeMo Framework](https://docs.nvidia.com/nemo-framework/user-guide/24.07/overview.html) to demonstrate a scalable, multi-node training workflow. All required artifacts and commands are included.
## 1. Prerequisites
### 1.1. Google Cloud Project setup
- **Enable APIs:** Ensure the Vertex AI API is [enabled for your project](http://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com).
- **H100 Mega Quota:** A3 Mega VMs are powered by H100 GPUs. Request quota for `custom_model_training_nvidia_h100_mega_gpus` in one of the [supported regions](https://cloud.google.com/vertex-ai/docs/general/locations#accelerator_support). If using Spot VMs, request `custom_model_training_preemptible_nvidia_h100_mega_gpus` quota instead.
- **Reservations (Optional but recommended):** For guaranteed capacity, [create a reservation](https://cloud.google.com/compute/docs/instances/reservations-shared) and ensure the reservation is shared with the Vertex AI service account. This guide requires a minimum of **16 H100 GPUs** (2 full A3 Mega nodes).
### 1.2. GCS bucket
Create a [Cloud Storage bucket](https://cloud.google.com/storage/docs/creating-buckets) in the same region where you have quota. If you're using Hierarchical Namespace for your bucket, you may need to update permissions of the Vertex AI Custom Code Service Agent .
This bucket is used for:
- Staging the training application.
- Storing model checkpoints and logs.
- Storing data if you use your own data.
## 2. Setup & configuration
### 2.1. Clone the repo
First clone the repo into your development environment.
```bash
git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git
```
Navigate to the root folder for this sample.
### 2.2. Environment Setup
First, configure your local environment. These variables are used in subsequent commands.
```bash
# Required: Update with your values
export PROJECT_ID="<your-project-id>"
export REPOSITORY="<your-artifact-registry-repo-name>" # e.g., "my-containers"
export BUCKET="<your-gcs-bucket-name>"
# Optional: Change if needed
export REGION="us-central1"
# --- Do not change the lines below ---
export ARTIFACT_REGISTRY="${REGION}-docker.pkg.dev/${PROJECT_ID}/${REPOSITORY}"
export REPO_ROOT=$(git rev-parse --show-toplevel)
```
## 3. Build and push a docker container image to Artifact Registry
Normally, you can use any custom training container on Vertex AI Training. In this example you build a NeMo Docker image that is based on the [Nvidia’s NeMo 24.09](https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo/tags) image. Use Cloud Build to build and push the container image.
This document picked NeMo as the demonstrating container since it’s a widely adopted GPU LLM training framework providing high performance and versatile training functionalities.
In addition to the base image, some customizations are included to form the final prebuilt image:
- Some dependencies are installed to integrate with Vertex AI Training.
- An entrypoint script that sets up required environments and calls the training job.
- Some patches are applied to the NeMo code to let it load the dataset from a GCS bucket.
Run this command to build the container and push the container into the Google Artifact Registry.
```bash
cd "${REPO_ROOT}/community-content/vertex-distributed-training/a3mega/llama-3-8b-nemo-pretraining"
export IMAGE_NAME="vertex-nemo-llama"
gcloud builds submit . \
--project="${PROJECT_ID}" \
--region="${REGION}" \
--config=docker/cloudbuild.yml \
--substitutions="_ARTIFACT_REGISTRY=${ARTIFACT_REGISTRY},_IMAGE_NAME=${IMAGE_NAME}" \
--timeout="2h" \
--machine-type="e2-highcpu-32"
```
## 4. Launch the Training Job
### 4.1. Job Configuration File
Once the container is built, update the job_config.json to set up the training job.
File: job_config.json
```json
{
"project_id": "<project-id>",
"region": "<region>",
"zone": "<zone if using reservation>",
"bucket": "<bucket>",
"dataset_bucket": "github-repo/data/third-party/enwiki-latest-pages-articles",
"image_uri": "<docker image uri from artifact registry>",
"strategy": "spot",
"nodes": "2",
"machine_type": "a3-megagpu-8g",
"gpu_type": "NVIDIA_H100_MEGA_80GB",
"gpus_per_node": "8",
"recipe_name": "llama3_1_8b_pretrain_a3mega",
"job_prefix": "vertex-spot-",
"reservation_name": ""
}
```
### 4.2 Launch the Training Job
First, create a Python virtual environment using your tool of choice, then install
the requirements specified in `requirements.txt`. Using `pip`, the command would be:
```bash
pip install -r requirements.txt
```
Now launch the Vertex AI training job using the provided Python script.
```bash
python3 scripts/launch.py --config_file=job_config.json
```
This script reads job_config.json, defines the cluster specification (2 nodes, 8 GPUs each), and submits the custom training job to Vertex AI.
## 5. Monitor and Clean Up
### 5.1. Monitoring
Vertex AI Console: Track the job's status in the Google Cloud Console under Vertex AI > Training > Custom Jobs.
Logs: View detailed logs in Cloud Logging by filtering for your job name.
Checkpoints: Model checkpoints are saved to your GCS bucket at the path specified in your training script's configuration.
### 5.2. Cleaning Up
To avoid ongoing charges, delete the resources you created:
- The Artifact Registry image.
- The contents of the GCS bucket (checkpoints, logs).
- The Vertex AI Custom Job will eventually complete or fail, incurring no further cost.
@@ -0,0 +1,265 @@
# Reference:
# https://github.com/NVIDIA/NeMo-Framework-Launcher/blob/24.07/launcher_scripts/conf/training/llama/llama3_1_8b.yaml
name: llama3_1_8b_pretrain_a3mega
restore_from_path: null # used when starting from a .nemo file
trainer:
devices: 8
num_nodes: 1
accelerator: gpu
precision: bf16
logger: false # logger provided by exp_manager
enable_checkpointing: false
use_distributed_sampler: false
max_epochs: -1 # PTL default. In practice, max_steps will be reached first.
max_steps: 30 # consumed_samples = global_step * micro_batch_size * data_parallel_size * accumulate_grad_batches
log_every_n_steps: 1
val_check_interval: null
limit_val_batches: 1
limit_test_batches: 1
accumulate_grad_batches: 1 # do not modify, grad acc is automatic for training megatron models
gradient_clip_val: 1.0
benchmark: false
enable_model_summary: false # default PTL callback for this does not support model parallelism, instead we log manually
exp_manager:
explicit_log_dir: null
exp_dir: /data
name: ${name}
create_dllogger_logger: true
dllogger_logger_kwargs:
verbose: true
stdout: true
json_file: "/data/dllogger.json"
create_wandb_logger: false
wandb_logger_kwargs:
project: null
name: null
resume_if_exists: true
resume_ignore_no_checkpoint: true
create_checkpoint_callback: false
checkpoint_callback_params:
monitor: val_loss
save_top_k: 3
mode: min
always_save_nemo: false # saves nemo file during validation, not implemented for model parallel
save_nemo_on_train_end: false # not recommended when training large models on clusters with short time limits
filename: 'megatron_gpt--{val_loss:.2f}-{step}-{consumed_samples}'
model_parallel_size: ${multiply:${model.tensor_model_parallel_size}, ${model.pipeline_model_parallel_size}}
seconds_to_sleep: 5 # Allows node_rank!=0 to sleep and let node0 to init, like preparing data
model:
mcore_gpt: true
# specify micro_batch_size, global_batch_size, and model parallelism
# gradient accumulation will be done automatically based on data_parallel_size
micro_batch_size: 1 # limited by GPU memory
global_batch_size: 1024 # will use more micro batches to reach global batch size
tensor_model_parallel_size: 1 # intra-layer model parallelism
pipeline_model_parallel_size: 2 # inter-layer model parallelism
context_parallel_size: 1
virtual_pipeline_model_parallel_size: null # interleaved pipeline
## Sequence Parallelism
# Makes tensor parallelism more memory efficient for LLMs (20B+) by parallelizing layer norms and dropout sequentially
# See Reducing Activation Recomputation in Large Transformer Models: https://arxiv.org/abs/2205.05198 for more details.
sequence_parallel: false
fsdp: false
fsdp_cpu_offload: true
fsdp_sharding_strategy: "full" # Method to shard model states. Available options are 'full', 'hybrid', and 'grad'.
fsdp_grad_reduce_dtype: "16" # Gradient reduction data type.
fsdp_sharded_checkpoint: false # Store and load FSDP shared checkpoint.
fsdp_use_orig_params: false # Set to True to use FSDP for specific peft scheme.
# Distributed checkpoint setup
dist_ckpt_format: "torch_dist" # Set to 'torch_dist' to use PyTorch distributed checkpoint format.
dist_ckpt_load_on_device: true # whether to load checkpoint weights directly on GPU or to CPU
dist_ckpt_parallel_save: true # if true, each worker will write its own part of the dist checkpoint
dist_ckpt_parallel_save_within_dp: false # if true, save will be parallelized only within a DP group (whole world otherwise), which might slightly reduce the save overhead
dist_ckpt_parallel_load: false # if true, each worker will load part of the dist checkpoint and exchange with NCCL. Might use some extra GPU memory
dist_ckpt_torch_dist_multiproc: 2 # number of extra processes per rank used during ckpt save with PyTorch distributed format
dist_ckpt_assume_constant_structure: false # set to True only if the state dict structure doesn't change within a single job. Allows caching some computation across checkpoint saves.
dist_ckpt_parallel_dist_opt: true # parallel save/load of a DistributedOptimizer. 'True' allows performant save and reshardable checkpoints. Set to 'False' only in order to minimize the number of checkpoint files.
dist_ckpt_load_strictness: null # defines checkpoint keys mismatch behavior (only during dist-ckpt load). Choices: assume_ok_unexpected (default - try loading without any check), log_all (log mismatches), raise_all (raise mismatches)
# model architecture
encoder_seq_length: 8192
max_position_embeddings: ${.encoder_seq_length}
num_layers: 32 # 8b: 32 | 70b: 80 | 405b: 126
hidden_size: 4096 # 8b: 4096 | 70b: 8192 | 405b: 16384
ffn_hidden_size: 14336 # 8b: 14336 | 70b: 28672 | 405b: 53248
num_attention_heads: 32 # 8b: 32 | 70b: 64 | 405b: 128
num_query_groups: 8 # Number of query groups for group query attention. If None, normal attention is used. 8b: 8 | 70b: 8 | 405b: 16
init_method_std: 0.01 # Standard deviation of the zero mean normal distribution used for weight initialization. 8b: 0.01 | 70b: 0.008944 | 405b: 0.02
use_scaled_init_method: true # use scaled residuals initialization
hidden_dropout: 0.0 # Dropout probability for hidden state transformer.
attention_dropout: 0.0 # Dropout probability for attention
ffn_dropout: 0.0 # Dropout probability in the feed-forward layer.
kv_channels: null # Projection weights dimension in multi-head attention. Set to hidden_size // num_attention_heads if null
apply_query_key_layer_scaling: true # scale Q * K^T by 1 / layer-number.
normalization: 'rmsnorm' # Normalization layer to use. Options are 'layernorm', 'rmsnorm'
layernorm_epsilon: 1e-5
do_layer_norm_weight_decay: false # True means weight decay on all params
make_vocab_size_divisible_by: 128 # Pad the vocab size to be divisible by this value for computation efficiency.
pre_process: true # add embedding
post_process: true # add pooler
persist_layer_norm: true # Use of persistent fused layer norm kernel.
bias: false # Whether to use bias terms in all weight matrices.
activation: 'fast-swiglu' # Options ['gelu', 'geglu', 'swiglu', 'reglu', 'squared-relu', 'fast-geglu', 'fast-swiglu', 'fast-reglu']
headscale: false # Whether to learn extra parameters that scale the output of the each self-attention head.
transformer_block_type: 'pre_ln' # Options ['pre_ln', 'post_ln', 'normformer']
openai_gelu: false # Use OpenAI's GELU instead of the default GeLU
normalize_attention_scores: true # Whether to scale the output Q * K^T by 1 / sqrt(hidden_size_per_head). This arg is provided as a configuration option mostly for compatibility with models that have been weight-converted from HF. You almost always want to se this to True.
position_embedding_type: 'rope' # Position embedding type. Options ['learned_absolute', 'rope']
rotary_percentage: 1.0 # If using position_embedding_type=rope, then the per head dim is multiplied by this.
attention_type: 'multihead' # Attention type. Options ['multihead']
share_embeddings_and_output_weights: false # Share embedding and output layer weights.
scale_positional_embedding: true # This is false for llama3 models. Only used for >= llama3.1.
# Use GPT2BPETokenizer for test, because the testing dataset is tokenized by this tokenizer.
# https://docs.nvidia.com/nemo-framework/user-guide/24.07/playbooks/singlenodepretrain.html#data-download-and-pre-processing
tokenizer:
library: megatron
type: GPT2BPETokenizer
model: null # /path/to/tokenizer.model
vocab_file: null
merge_file: null
delimiter: null # only used for tabular tokenizer
sentencepiece_legacy: false # Legacy=True allows you to add special tokens to sentencepiece tokenizers.
# Mixed precision
native_amp_init_scale: 4294967296 # 2 ** 32
native_amp_growth_interval: 1000
hysteresis: 2 # Gradient scale hysteresis
fp32_residual_connection: false # Move residual connections to fp32
fp16_lm_cross_entropy: false # Move the cross entropy unreduced loss calculation for lm head to fp16
# Megatron O2-style half-precision
megatron_amp_O2: true # Enable O2-level automatic mixed precision using main parameters
grad_allreduce_chunk_size_mb: 125
# Fusion
grad_div_ar_fusion: true # Fuse grad division into torch.distributed.all_reduce. Only used with O2 and no pipeline parallelism..
gradient_accumulation_fusion: true # Fuse weight gradient accumulation to GEMMs. Only used with pipeline parallelism and O2.
bias_activation_fusion: true # Use a kernel that fuses the bias addition from weight matrices with the subsequent activation function.
bias_dropout_add_fusion: true # Use a kernel that fuses the bias addition, dropout and residual connection addition.
masked_softmax_fusion: true # Use a kernel that fuses the attention softmax with it's mask.
apply_rope_fusion: true # Use a kernel to add rotary positional embeddings. Only used if position_embedding_type=rope
cross_entropy_loss_fusion: true
# Miscellaneous
seed: 1234
resume_from_checkpoint: null # manually set the checkpoint file to load from
use_cpu_initialization: false # Init weights on the CPU (slow for large models)
onnx_safe: false # Use work-arounds for known problems with Torch ONNX exporter.
apex_transformer_log_level: 30 # Python logging level displays logs with severity greater than or equal to this
gradient_as_bucket_view: true # PyTorch DDP argument. Allocate gradients in a contiguous bucket to save memory (less fragmentation and buffer memory)
sync_batch_comm: false # Enable stream synchronization after each p2p communication between pipeline stages
## Activation Checkpointing
# NeMo Megatron supports 'selective' activation checkpointing where only the memory intensive part of attention is checkpointed.
# These memory intensive activations are also less compute intensive which makes activation checkpointing more efficient for LLMs (20B+).
# See Reducing Activation Recomputation in Large Transformer Models: https://arxiv.org/abs/2205.05198 for more details.
# 'full' will checkpoint the entire transformer layer.
activations_checkpoint_granularity: null # 'selective' or 'full'
activations_checkpoint_method: null # 'uniform', 'block'
# 'uniform' divides the total number of transformer layers and checkpoints the input activation
# of each chunk at the specified granularity. When used with 'selective', 'uniform' checkpoints all attention blocks in the model.
# 'block' checkpoints the specified number of layers per pipeline stage at the specified granularity
activations_checkpoint_num_layers: null
# when using 'uniform' this creates groups of transformer layers to checkpoint. Usually set to 1. Increase to save more memory.
# when using 'block' this this will checkpoint the first activations_checkpoint_num_layers per pipeline stage.
num_micro_batches_with_partial_activation_checkpoints: null
# This feature is valid only when used with pipeline-model-parallelism.
# When an integer value is provided, it sets the number of micro-batches where only a partial number of Transformer layers get checkpointed
# and recomputed within a window of micro-batches. The rest of micro-batches in the window checkpoint all Transformer layers. The size of window is
# set by the maximum outstanding micro-batch backpropagations, which varies at different pipeline stages. The number of partial layers to checkpoint
# per micro-batch is set by 'activations_checkpoint_num_layers' with 'activations_checkpoint_method' of 'block'.
# This feature enables using activation checkpoint at a fraction of micro-batches up to the point of full GPU memory usage.
activations_checkpoint_layers_per_pipeline: null
# This feature is valid only when used with pipeline-model-parallelism.
# When an integer value (rounded down when float is given) is provided, it sets the number of Transformer layers to skip checkpointing at later
# pipeline stages. For example, 'activations_checkpoint_layers_per_pipeline' of 3 makes pipeline stage 1 to checkpoint 3 layers less than
# stage 0 and stage 2 to checkpoint 6 layers less stage 0, and so on. This is possible because later pipeline stage
# uses less GPU memory with fewer outstanding micro-batch backpropagations. Used with 'num_micro_batches_with_partial_activation_checkpoints',
# this feature removes most of activation checkpoints at the last pipeline stage, which is the critical execution path.
## Transformer Engine
transformer_engine: true
fp8: false # enables fp8 in TransformerLayer forward
fp8_e4m3: false # sets fp8_format = recipe.Format.E4M3
fp8_hybrid: false # sets fp8_format = recipe.Format.HYBRID
fp8_margin: 0 # scaling margin
fp8_interval: 1 # scaling update interval
fp8_amax_history_len: 1024 # Number of steps for which amax history is recorded per tensor
fp8_amax_compute_algo: 'max' # 'most_recent' or 'max'. Algorithm for computing amax from history
ub_tp_comm_overlap: false # do not turn on because of b/397797926
use_flash_attention: true
gc_interval: 100
## Offloading Activations/Weights to CPU
cpu_offloading: false
cpu_offloading_num_layers: ${sum:${.num_layers},-1} # This value should be between [1,num_layers-1] as we don't want to offload the final layer's activations and expose any offloading duration for the final layer
cpu_offloading_activations: true
cpu_offloading_weights: true
data:
# Path to data must be specified by the user.
# Supports List, String and Dictionary
# List : can override from the CLI: "model.data.data_prefix=[.5,/raid/data/pile/my-gpt3_00_text_document,.5,/raid/data/pile/my-gpt3_01_text_document]",
# Or see example below:
# data_prefix:
# - .5
# - /raid/data/pile/my-gpt3_00_text_document
# - .5
# - /raid/data/pile/my-gpt3_01_text_document
# Dictionary: can override from CLI "model.data.data_prefix"={"train":[1.0, /path/to/data], "validation":/path/to/data, "test":/path/to/test}
# Or see example below:
# "model.data.data_prefix: {train:[1.0,/path/to/data], validation:[/path/to/data], test:[/path/to/test]}"
data_prefix: [1.0, /data/hfbpe_gpt_training_data_text_document]
index_mapping_dir: null # path to save index mapping .npy files, by default will save in the same location as data_prefix
data_impl: mmap
splits_string: 900,50,50
seq_length: ${model.encoder_seq_length}
skip_warmup: true
num_workers: 2
dataloader_type: single # cyclic
reset_position_ids: false # Reset position ids after end-of-document token
reset_attention_mask: false # Reset attention mask after end-of-document token
eod_mask_loss: false # Mask loss for the end of document tokens
validation_drop_last: true # Set to false if the last partial validation samples is to be consumed
no_seqlen_plus_one_input_tokens: false # Set to True to disable fetching (sequence length + 1) input tokens, instead get (sequence length) input tokens and mask the last token
pad_samples_to_global_batch_size: false # Set to True if you want to pad the last partial batch with -1's to equal global batch size
shuffle_documents: true # Set to False to disable documents shuffling. Sample index will still be shuffled
# Nsys profiling options
nsys_profile:
enabled: false
start_step: 0 # Global batch to start profiling
end_step: 1 # Global batch to end profiling
ranks: [0] # Global rank IDs to profile
gen_shape: false # Generate model and kernel details including input shapes
memory_profile:
enabled: false
start_step: 0
end_step: 1
ranks: [0]
output_path: /data # Must be a dir
optim:
name: distributed_fused_adam # E.g., fused_adam or set _target_: torch.optim.AdamW field
lr: 2e-5
weight_decay: 0.01
betas:
- 0.9
- 0.98
bucket_cap_mb: 125
overlap_grad_sync: true
overlap_param_sync: true
contiguous_grad_buffer: true
contiguous_param_buffer: true
sched:
name: CosineAnnealing
warmup_steps: 400
constant_steps: 0
min_lr: 2e-6
@@ -0,0 +1,26 @@
# Copyright 2024 Google LLC
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
steps:
- name: 'gcr.io/cloud-builders/docker'
args:
- 'build'
- '--tag=${_ARTIFACT_REGISTRY}/${_IMAGE_NAME}'
- '--file=docker/vertex-dist-recipes.Dockerfile'
- '.'
automapSubstitutions: true
env:
- 'DOCKER_BUILDKIT=1'
images:
- '${_ARTIFACT_REGISTRY}/${_IMAGE_NAME}'
@@ -0,0 +1,41 @@
diff --git a/nemo/collections/nlp/parts/megatron_trainer_builder.py b/nemo/collections/nlp/parts/megatron_trainer_builder.py
index b2c85cde4..a3a9670c3 100644
--- a/nemo/collections/nlp/parts/megatron_trainer_builder.py
+++ b/nemo/collections/nlp/parts/megatron_trainer_builder.py
@@ -19,6 +19,7 @@ from lightning_fabric.utilities.exceptions import MisconfigurationException
from omegaconf import DictConfig
from pytorch_lightning import Trainer
from pytorch_lightning.callbacks import ModelSummary
+from pytorch_lightning.callbacks import Callback
from pytorch_lightning.plugins.environments import TorchElasticEnvironment
from nemo.collections.common.metrics.perf_metrics import FLOPsMeasurementCallback
@@ -38,6 +39,23 @@ from nemo.utils.callbacks.dist_ckpt_io import (
AsyncFinalizerCallback,
DistributedCheckpointIO,
)
+from vmg.util.device_stats import gpu_stats_str
+
+class GpuStatsMon(Callback):
+ def on_train_start(self, trainer, pl_module) -> None:
+ rank=pl_module.global_rank
+ print(f'train_start: {rank=} {gpu_stats_str()}', flush=True)
+
+ def on_train_batch_start(self, trainer, pl_module, batch, batch_idx) -> None:
+ rank=pl_module.global_rank
+ print(f'batch_start: {rank=} {gpu_stats_str()}', flush=True)
+
+ def on_train_batch_end(self, trainer, pl_module, outputs, batch, batch_idx) -> None:
+ rank=pl_module.global_rank
+ print(f'batch_end: {rank=} {gpu_stats_str()}', flush=True)
class MegatronTrainerBuilder:
@@ -178,6 +196,7 @@ class MegatronTrainerBuilder:
if self.cfg.get('exp_manager', {}).get('log_tflops_per_sec_per_gpu', True):
callbacks.append(FLOPsMeasurementCallback(self.cfg))
+ callbacks.append(GpuStatsMon())
return callbacks
def create_trainer(self, callbacks=None) -> Trainer:
@@ -0,0 +1,41 @@
diff -ruN old-datasets/blended_megatron_dataset_builder.py datasets/blended_megatron_dataset_builder.py
--- old-datasets/blended_megatron_dataset_builder.py 2025-05-02 04:08:45.369199665 +0000
+++ datasets/blended_megatron_dataset_builder.py 2025-05-02 04:10:47.369119891 +0000
@@ -2,6 +2,7 @@
import logging
import math
+import os
from concurrent.futures import ThreadPoolExecutor
from typing import Any, Callable, Iterable, List, Optional, Type, Union
@@ -353,7 +354,7 @@
num_dataset_builder_threads = self.config.num_dataset_builder_threads
if torch.distributed.is_initialized():
- rank = torch.distributed.get_rank()
+ rank = int(os.getenv("LOCAL_RANK", "0"))
# First, build on rank 0
if rank == 0:
num_workers = num_dataset_builder_threads
@@ -475,7 +476,7 @@
Optional[Union[DistributedDataset, Iterable]]: The DistributedDataset instantion, the Iterable instantiation, or None
"""
if torch.distributed.is_initialized():
- rank = torch.distributed.get_rank()
+ rank = int(os.getenv("LOCAL_RANK", "0"))
dataset = None
diff -ruN old-datasets/gpt_dataset.py datasets/gpt_dataset.py
--- old-datasets/gpt_dataset.py 2025-05-02 04:08:45.369199665 +0000
+++ datasets/gpt_dataset.py 2025-05-02 04:09:30.309170278 +0000
@@ -351,7 +351,7 @@
if not path_to_cache or (
not cache_hit
- and (not torch.distributed.is_initialized() or torch.distributed.get_rank() == 0)
+ and (not torch.distributed.is_initialized() or int(os.getenv("LOCAL_RANK", "0")) == 0)
):
log_single_rank(
@@ -0,0 +1,13 @@
diff --git a/scripts/checkpoint_converters/convert_llama_nemo_to_hf.py b/scripts/checkpoint_converters/convert_llama_nemo_to_hf.py
index 8da15148d..005cae6c9 100644
--- a/scripts/checkpoint_converters/convert_llama_nemo_to_hf.py
+++ b/scripts/checkpoint_converters/convert_llama_nemo_to_hf.py
@@ -104,6 +104,8 @@ def convert(input_nemo_file, output_hf_file, precision=None, cpu_only=False) ->
dummy_trainer = Trainer(devices=1, accelerator='cpu', strategy=NLPDDPStrategy())
model_config = MegatronGPTModel.restore_from(input_nemo_file, trainer=dummy_trainer, return_config=True)
model_config.tensor_model_parallel_size = 1
+ model_config.virtual_pipeline_model_parallel_size = None
+ model_config.sequence_parallel = False
model_config.pipeline_model_parallel_size = 1
if cpu_only:
map_location = torch.device('cpu')
@@ -0,0 +1,24 @@
diff --git a/examples/nlp/language_modeling/tuning/megatron_gpt_finetuning.py b/examples/nlp/language_modeling/tuning/megatron_gpt_finetuning.py
index bfe8ea359..dfeaf93b5 100644
--- a/examples/nlp/language_modeling/tuning/megatron_gpt_finetuning.py
+++ b/examples/nlp/language_modeling/tuning/megatron_gpt_finetuning.py
@@ -13,6 +13,8 @@
# limitations under the License.
import torch.multiprocessing as mp
+import torch.distributed as dist
+
from omegaconf.omegaconf import OmegaConf
from nemo.collections.nlp.models.language_modeling.megatron_gpt_sft_model import MegatronGPTSFTModel
@@ -76,6 +78,10 @@ def main(cfg) -> None:
trainer.fit(model)
+ if dist.is_available() and dist.is_initialized():
+ dist.barrier()
+ dist.destroy_process_group()
+
if __name__ == '__main__':
main()
@@ -0,0 +1,13 @@
diff --git a/src/utils/training_metrics/process_training_results.py b/src/utils/training_metrics/process_training_results.py
index 3e82a66..e61e1d8 100644
--- a/src/utils/training_metrics/process_training_results.py
+++ b/src/utils/training_metrics/process_training_results.py
@@ -134,7 +134,7 @@ def get_average_step_time(file: str, start_step: int, end_step: int) -> float:
for line in datajson:
if line.get("step") != "PARAMETER":
step = line.get("step")
- if step >= start_step and step <= end_step:
+ if step >= start_step and step <= end_step and "train_step_timing in s" in line["data"]:
time_step_accumulator += line["data"].get("train_step_timing in s")
num_steps += 1
if num_steps == 0:
@@ -0,0 +1,10 @@
dllogger@git+https://github.com/NVIDIA/dllogger@v1.0.0
# Fixing these libraries versions to avoid conflicting or broken packages.
immutabledict==4.2.1
protobuf==5.29.6
opencv-python-headless==4.11.0.86
docutils==0.16
urllib3==2.7.0
google-cloud-storage==3.0.0
retrying
@@ -0,0 +1,18 @@
# cuml-cu12==24.8.0 was installed in nemo:24.09
# Removing cuml=24.4.0 to avoid conflicting packages.
cudf==24.4.0
cugraph==24.4.0
cugraph-service-server==24.4.0
cuml==24.4.0
dask-cudf==24.4.0
raft-dask==24.4.0
cugraph-dgl==24.4.0
cugraph-pyg==24.4.0
# The following packages are removed temporarily to avoid conflicting packages
# and can be brought back if needed.
tensorrt-llm==0.12.0
img2dataset==1.45.0
Sphinx==8.1.3
sphinxcontrib-bibtex==2.6.3
torchx==0.7.0
nemo-run
@@ -0,0 +1,66 @@
# Dockerfile wrapping NeMo.
#
# To workaround base nemo docker image using too many layers, we use Multi-stage
# build to first collect the additional files we'll need.
FROM alpine:latest AS prep_files
WORKDIR /workspace
RUN mkdir -p configs vdt vdt/util
COPY scripts/*.py vdt/
COPY scripts/util/*.py vdt/util/
COPY configs/* configs/
COPY docker/patches/24.09/* vdt/patches/
RUN chmod a+rwX -R vdt
# Copy license.
RUN wget https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/main/LICENSE
# Available tags
# https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo/tags
# It installs NeMo source code in /opt/NeMo folder, with tag=r2.0.0
FROM nvcr.io/nvidia/nemo:24.09
RUN apt-get update && apt-get install -y sudo zsh tmux && \
rm -rf /var/lib/apt/lists*
RUN echo "deb [signed-by=/usr/share/keyrings/cloud.google.gpg] http://packages.cloud.google.com/apt cloud-sdk main" | \
tee -a /etc/apt/sources.list.d/google-cloud-sdk.list && \
curl https://packages.cloud.google.com/apt/doc/apt-key.gpg | \
apt-key --keyring /usr/share/keyrings/cloud.google.gpg add - && \
apt-get update -y && apt-get install google-cloud-sdk -y && \
rm -rf /var/lib/apt/lists*
# Install libraries with pip
ENV PIP_ROOT_USER_ACTION=ignore
# We expect this will be run in the root directory of the vertex-dist-recipes repo
ARG HOST_SRC_DIR="."
# The pre-installed NeMo introduces a lot of deps conflicts.
# We uninstall the confilicting libs and reinstall some of them as needed.
COPY ${HOST_SRC_DIR}/docker/uninstall.txt /tmp/uninstall.txt
RUN cat /tmp/uninstall.txt | grep -v '#' | xargs pip uninstall -y
COPY ${HOST_SRC_DIR}/docker/requirements.txt /tmp/requirements.txt
RUN pip install -r /tmp/requirements.txt
# Make sure there's no inconsistent pip libraries.
RUN pip check
WORKDIR /workspace
# Copy configs
COPY ${HOST_SRC_DIR}/configs/* /opt/NeMo/examples/nlp/language_modeling/conf/
# Copy all additional files we need from `prep_files` image.
COPY --from=prep_files /workspace/ .
# Install for `src/utils/training_metrics/process_training_results.py` to report
# throughput and MFU numbers.
RUN git clone https://github.com/AI-Hypercomputer/gpu-recipes.git
# This hack is needed for multi-node training while not using a sharing file system.
RUN patch --verbose -l -d /opt/megatron-lm/megatron/core/datasets -p1 -i /workspace/vdt/patches/local_rank.patch; \
git -C /workspace/gpu-recipes apply /workspace/vdt/patches/throughput_calc.patch; \
git -C /opt/NeMo apply /workspace/vdt/patches/nemo2hf.patch; \
git -C /opt/NeMo apply /workspace/vdt/patches/sigabort.patch;
# git -C /opt/NeMo apply /workspace/vdt/patches/gpu_stats.patch;
# Do not put an entrypoint here. Specify the entrypoint in the docker run script.
@@ -0,0 +1,16 @@
{
"project_id": "<your_project_id>",
"region": "us-central1",
"zone": "us-central1-c",
"bucket": "<your_bucket",
"dataset_bucket": "github-repo/data/third-party/enwiki-latest-pages-articles",
"image_uri": "<your_image_uri>",
"strategy": "spot",
"nodes": "2",
"machine_type": "a3-megagpu-8g",
"gpu_type": "NVIDIA_H100_MEGA_80GB",
"gpus_per_node": "8",
"recipe_name": "llama3_1_8b_pretrain_a3mega",
"job_prefix": "vertex-ai",
"reservation_name": ""
}
@@ -0,0 +1,49 @@
absl-py==2.2.2
annotated-types==0.7.0
anyio==4.9.0
black==26.3.1
cachetools==5.5.2
certifi==2025.4.26
charset-normalizer==3.4.2
click==8.1.8
docstring_parser==0.16
google-api-core==2.24.2
google-auth==2.40.1
google-cloud-aiplatform==1.133.0
google-cloud-bigquery==3.31.0
google-cloud-core==2.4.3
google-cloud-resource-manager==1.14.2
google-cloud-storage==2.19.0
google-crc32c==1.7.1
google-genai==1.14.0
google-resumable-media==2.7.2
googleapis-common-protos==1.70.0
grpc-google-iam-v1==0.14.2
grpcio==1.71.0
grpcio-status==1.71.0
h11==0.16.0
httpcore==1.0.9
httpx==0.28.1
idna==3.15
mypy_extensions==1.1.0
numpy==2.2.5
packaging==25.0
pathspec==0.12.1
platformdirs==4.3.8
proto-plus==1.26.1
protobuf==5.29.6
pyasn1==0.6.4
pyasn1_modules==0.4.2
pydantic==2.11.4
pydantic_core==2.33.2
python-dateutil==2.9.0.post0
pytz==2025.2
requests==2.33.0
rsa==4.9.1
shapely==2.1.0
six==1.17.0
sniffio==1.3.1
typing-inspection==0.4.0
typing_extensions==4.13.2
urllib3==2.7.0
websockets==15.0.1
@@ -0,0 +1,173 @@
"""Launch script for Vertex distributed training"""
# Copy the sample_job_config.json file to job_config.json
# to define the job parameters.
#
# Run like this:
#
# python3 vertex_dist_train/launch.py --config_file=job_config.json
#
import datetime
import json
import os
import pprint
from collections.abc import Sequence
from typing import Any, List
from absl import app, flags
from google.cloud import aiplatform
from google.cloud.aiplatform_v1.types.custom_job import Scheduling
from pytz import timezone
FLAGS = flags.FLAGS
flags.DEFINE_string("config_file", None, "Path to JSON config file")
flags.DEFINE_boolean(
"debug", False, "Debug mode: just print the command, don't run it."
)
def launch_job(
job_name: str,
project: str,
region: str,
gcs_bucket: str,
image_uri: str,
entrypoint_cmd: List[str],
trainer_args: List[Any],
num_nodes: int,
machine_type: str,
num_gpus_per_node: int,
gpu_type: str,
strategy: str,
reservation_name: str = "",
):
assert strategy in ("dws", "spot", "reservation")
aiplatform.init(
project=project, location=region, staging_bucket=gcs_bucket
)
train_job = aiplatform.CustomContainerTrainingJob(
display_name=job_name,
container_uri=image_uri,
command=entrypoint_cmd,
)
job_args = dict(
args=trainer_args,
enable_web_access=True,
replica_count=num_nodes,
machine_type=machine_type,
accelerator_type=gpu_type,
accelerator_count=num_gpus_per_node,
boot_disk_size_gb=1000,
restart_job_on_worker_restart=True,
#restart_job_on_worker_restart=False,
)
if strategy == "spot":
job_args.update({"scheduling_strategy": Scheduling.Strategy.SPOT.name})
elif strategy == "dws":
job_args.update(
{"scheduling_strategy": Scheduling.Strategy.FLEX_START.name}
)
elif strategy == "reservation":
assert reservation_name != "", (
"If using a reservation, provide the reservation_name in the "
"format `projects/{project_id_or_number}/zones/{zone}/"
"reservations/{reservation_name}`"
)
job_args.update(
{
"reservation_affinity_type": "SPECIFIC_RESERVATION",
"reservation_affinity_key": "compute.googleapis.com/reservation-name",
"reservation_affinity_values": [reservation_name],
}
)
pprint.pprint(job_args)
if not FLAGS.debug:
train_job.submit(**job_args)
def main(argv: Sequence[str]) -> None:
config_file_path = FLAGS.config_file
print(f"Reading job config from {config_file_path}")
with open(config_file_path, encoding="utf-8") as config_file:
config = json.load(config_file)
project_id = config["project_id"]
region = config["region"]
zone = config["zone"]
bucket = config["bucket"]
dataset_bucket = config["dataset_bucket"]
n_nodes = int(config["nodes"])
machine_type = config["machine_type"]
num_gpus_per_node = int(config["gpus_per_node"])
gpu_type = config["gpu_type"]
reservation_name = config.get("reservation_name")
reservation_full_name = (
f"projects/{project_id}/zones/{zone}/reservations/{reservation_name}"
if "reservation_name" in config
else ""
)
strategy = config["strategy"]
recipe_name = config["recipe_name"]
job_prefix = config["job_prefix"]
image_uri = config["image_uri"]
# Job name
timestamp = (
datetime.datetime.now()
.astimezone(timezone("US/Pacific"))
.strftime("%Y%m%d_%H%M%S")
)
job_name = f"{recipe_name}-{timestamp}"
if job_prefix:
job_name = f"{job_prefix}-{job_name}"
base_output_dir = os.path.join("/gcs", bucket, job_name)
# Training command and args
entrypoint_cmd = ["python3", "vdt/run.py"]
dataset_bucket = f"gs://{config['dataset_bucket']}"
trainer_args = [
f"--train_data_gcs={dataset_bucket}",
"/opt/NeMo/examples/nlp/language_modeling/megatron_gpt_pretraining.py",
"--config-path=conf/",
f"--config-name={recipe_name}.yaml",
f"exp_manager.explicit_log_dir={base_output_dir}",
f"exp_manager.dllogger_logger_kwargs.json_file={base_output_dir}/dllogger.json",
"+exp_manager.create_tensorboard_logger=true",
"exp_manager.create_checkpoint_callback=false",
f"trainer.num_nodes={n_nodes}",
f"trainer.devices={num_gpus_per_node}",
"trainer.max_steps=10",
"trainer.log_every_n_steps=1",
"model.tokenizer.vocab_file=/data/gpt2-vocab.json",
"model.tokenizer.merge_file=/data/gpt2-merges.txt",
"model.data.data_prefix=[1.0,/data/hfbpe_gpt_training_data_text_document]",
]
launch_job(
job_name=job_name,
project=project_id,
region=region,
gcs_bucket=bucket,
image_uri=image_uri,
entrypoint_cmd=entrypoint_cmd,
trainer_args=trainer_args,
num_nodes=n_nodes,
machine_type=machine_type,
num_gpus_per_node=num_gpus_per_node,
gpu_type=gpu_type,
strategy=strategy,
reservation_name=reservation_full_name,
)
if __name__ == "__main__":
app.run(main)
@@ -0,0 +1,85 @@
"""Entrypoint for Vertex Distributed Training container."""
import argparse
import os
import sys
from collections.abc import Sequence
from subprocess import STDOUT, check_output, run
from absl import app, flags, logging
from util import cluster_spec
from retrying import retry
# PyTorch barrier call which synchronizes all of the nodes before launching the training process.
# This makes sure that processes will block until all processes are ready.
# Improves the reliability of spot VM usage for multi-node training jobs
@retry(stop_max_attempt_number=100, wait_exponential_multiplier=1000)
def barrier_with_retry() -> None:
import torch
logging.info("Starting barrier on RANK {}".format(os.environ["RANK"]))
torch.distributed.init_process_group()
torch.distributed.barrier()
torch.distributed.destroy_process_group()
logging.info("Finished barrier on RANK {}".format(os.environ["RANK"]))
def main(unused_argv: Sequence[str]) -> None:
parser = argparse.ArgumentParser()
parser.add_argument(
"--train_data_gcs",
type=str,
help="Download training data from gcs path",
)
args, unknown = parser.parse_known_args()
for key, val in os.environ.items():
logging.info("ENV %s=%s", key, val)
if args.train_data_gcs:
local_dir = "/data"
if not os.path.exists(local_dir):
os.mkdir(local_dir)
logging.info("downloading %s to %s...", args.train_data_gcs, local_dir)
check_output(
[
"gcloud",
"storage",
"cp",
"-r",
f"{args.train_data_gcs}/*",
local_dir,
],
stderr=STDOUT,
)
logging.info("%s downloaded.", args.train_data_gcs)
primary_node_addr, primary_node_port, node_rank, num_nodes = (
cluster_spec.get_cluster_spec()
)
cmd = [
"torchrun",
"--nproc-per-node=8",
f"--nnodes={num_nodes}",
f"--node_rank={node_rank}",
]
if num_nodes > 1:
cmd += [
"--max-restarts=3",
"--rdzv-backend=static",
f'--rdzv_id={os.getenv("CLOUD_ML_JOB_ID", primary_node_port)}',
f"--rdzv-endpoint={primary_node_addr}:{primary_node_port}",
]
cmd += unknown
logging.info("launching with cmd: \n%s", " \\\n".join(cmd))
barrier_with_retry()
run(cmd, stdout=sys.stdout, stderr=sys.stdout, check=True)
if __name__ == "__main__":
logging.get_absl_handler().python_handler.stream = sys.stdout
app.run(
main, flags_parser=lambda _args: flags.FLAGS(_args, known_only=True)
)
@@ -0,0 +1,81 @@
"""Get cluster info from environment variables."""
import dataclasses
import json
import os
from absl import logging
@dataclasses.dataclass
class ClusterInfo:
"""Contains information about the cluster.
Attributes:
primary_node_addr: The address of the primary node.
primary_node_port: The port of the primary node.
node_rank: The rank of the node.
num_nodes: The number of nodes in the cluster.
"""
primary_node_addr: str | None = None
primary_node_port: str | None = None
node_rank: int = 0
num_nodes: int = 1
# Allows unpacking operation like
# primary_node_addr, primary_node_port, _, _ = ClusterInfo()
# See https://stackoverflow.com/a/70753113
def __iter__(self):
return iter(dataclasses.astuple(self))
def get_cluster_spec() -> ClusterInfo:
"""Parses CLUSTER_SPEC environment variable and returns the cluster info.
Returns:
A ClusterInfo object.
"""
cluster_spec = os.getenv("CLUSTER_SPEC", None)
# If CLUSTER_SPEC is not set, use individual vars to construct cluster info.
if not cluster_spec:
cluster_info = ClusterInfo(
primary_node_addr=os.getenv("MASTER_ADDR", None),
primary_node_port=os.getenv("MASTER_PORT", None),
node_rank=int(os.getenv("RANK", "0")),
num_nodes=int(os.getenv("NNODES", "1")),
)
return cluster_info
cluster_data = json.loads(cluster_spec)
# Get primary node info
primary_node = cluster_data["cluster"]["workerpool0"][0]
logging.info("primary node: %s", primary_node)
primary_node_addr, primary_node_port = primary_node.split(":")
logging.info("primary node address: %s", primary_node_addr)
logging.info("primary node port: %s", primary_node_port)
# Determine node rank of this machine
workerpool = cluster_data["task"]["type"]
if workerpool == "workerpool0":
node_rank = 0
elif workerpool == "workerpool1":
# Add 1 for the primary node, since `index` is the index of workerpool1.
node_rank = cluster_data["task"]["index"] + 1
else:
raise ValueError(
"Only workerpool0 and workerpool1 are supported. Unknown workerpool:"
f" {workerpool}"
)
logging.info("node rank: %s", node_rank)
# Calculate total nodes.
num_nodes = 1 # For the primary node.
if "workerpool1" in cluster_data["cluster"]:
num_nodes += len(cluster_data["cluster"]["workerpool1"])
logging.info("num nodes: %s", num_nodes)
return ClusterInfo(
primary_node_addr, primary_node_port, node_rank, num_nodes
)
@@ -0,0 +1,59 @@
"""Add tests for cluster_spec.py."""
import os
from . import cluster_spec
# TODO(styer): Use pytest instead
class ClusterSpecTest(googletest.TestCase):
def setUp(self):
super().setUp()
self.curr_env_var = os.environ.copy()
def tearDown(self):
super().tearDown()
os.environ = self.curr_env_var
def test_get_cluster_spec_from_env_vars(self):
os.environ["CLUSTER_SPEC"] = ""
os.environ["MASTER_ADDR"] = "127.0.0.1"
os.environ["MASTER_PORT"] = "8080"
os.environ["RANK"] = "0"
os.environ["NNODES"] = "2"
cluster_info = cluster_spec.get_cluster_spec()
self.assertEqual(cluster_info.primary_node_addr, "127.0.0.1")
self.assertEqual(cluster_info.primary_node_port, "8080")
self.assertEqual(cluster_info.node_rank, 0)
self.assertEqual(cluster_info.num_nodes, 2)
def test_get_cluster_spec_from_cluster_spec(self):
os.environ[
"CLUSTER_SPEC"
] = """
{
"cluster": {
"workerpool0": [
"127.0.0.1:8080"
],
"workerpool1": [
"127.0.0.2:8080",
"127.0.0.3:8080"
]
},
"task": {
"type": "workerpool1",
"index": 0
}
}
"""
cluster_info = cluster_spec.get_cluster_spec()
self.assertEqual(cluster_info.primary_node_addr, "127.0.0.1")
self.assertEqual(cluster_info.primary_node_port, "8080")
self.assertEqual(cluster_info.node_rank, 1)
self.assertEqual(cluster_info.num_nodes, 3)
if __name__ == "__main__":
googletest.main()
@@ -0,0 +1,33 @@
import numpy as np
import os
import pickle
from google.cloud.aiplatform.constants import prediction
from google.cloud.aiplatform.utils import prediction_utils
from google.cloud.aiplatform.prediction.predictor import Predictor
from sklearn.datasets import load_breast_cancer
from sklearn.linear_model import RidgeClassifier
class LinearRegressionPredictor(Predictor):
def __init__(self):
return
def load(self, artifacts_uri: str) -> None:
prediction_utils.download_model_artifacts(artifacts_uri)
if os.path.exists(prediction.MODEL_FILENAME_PKL):
self._model = pickle.load(open(prediction.MODEL_FILENAME_PKL, "rb"))
else:
self._model = RidgeClassifier()
X, y = load_breast_cancer(return_X_y=True)
self._model.fit(X, y)
def preprocess(self, prediction_input: dict) -> np.ndarray:
instances = prediction_input["instances"]
return np.asarray(instances)
def predict(self, instances: np.ndarray) -> np.ndarray:
return self._model.predict(instances)
def postprocess(self, prediction_results: np.ndarray) -> dict:
return {"predictions": prediction_results.tolist()}
@@ -0,0 +1,33 @@
import numpy as np
import os
import pickle
from google.cloud.aiplatform.constants import prediction
from google.cloud.aiplatform.utils import prediction_utils
from google.cloud.aiplatform.prediction.predictor import Predictor
from sklearn.linear_model import SGDClassifier
class SGDClassifierPredictor(Predictor):
def __init__(self):
return
def load(self, artifacts_uri: str) -> None:
prediction_utils.download_model_artifacts(artifacts_uri)
if os.path.exists(prediction.MODEL_FILENAME_PKL):
self._model = pickle.load(open(prediction.MODEL_FILENAME_PKL, "rb"))
else:
self._model = SGDClassifier(max_iter=5)
X = [[0., 0.], [1., 1.]]
y = [0, 1]
self._model.fit(X, y)
def preprocess(self, prediction_input: dict) -> np.ndarray:
instances = prediction_input["instances"]
return np.asarray(instances)
def predict(self, instances: np.ndarray) -> np.ndarray:
return self._model.predict(instances)
def postprocess(self, prediction_results: np.ndarray) -> dict:
return {"predictions": prediction_results.tolist()}
@@ -0,0 +1,34 @@
import os
import torch
from google.cloud.aiplatform.utils import prediction_utils
from google.cloud.aiplatform.prediction.predictor import Predictor
from torchvision.models import detection, resnet50, ResNet50_Weights
from typing import Dict, List
class ResNetPredictor(Predictor):
def __init__(self):
return
def load(self, artifacts_uri: str) -> None:
prediction_utils.download_model_artifacts(artifacts_uri)
if os.path.exists("model.pth.tar"):
self.model = detection.fasterrcnn_resnet50_fpn(pretrained=True)
stat_dic = torch.load("model.pth.tar")
self.model.load_state_dict(stat_dic['state_dict'])
else:
weights = ResNet50_Weights.DEFAULT
self.model = resnet50(weights=weights)
self.model.eval()
def preprocess(self, prediction_input: dict) -> torch.Tensor:
instances = prediction_input["instances"]
return torch.Tensor(instances)
@torch.inference_mode()
def predict(self, instances: torch.Tensor) -> List[str]:
return self._model(instances)
def postprocess(self, prediction_results: List[str]) -> Dict:
return {"predictions": prediction_results}
@@ -0,0 +1,37 @@
import os
import numpy as np
import pickle
import xgboost as xgb
from google.cloud.aiplatform.constants import prediction
from google.cloud.aiplatform.utils import prediction_utils
from google.cloud.aiplatform.prediction.predictor import Predictor
from sklearn.datasets import make_blobs
from xgboost import XGBClassifier
class ClassifierPredictor(Predictor):
def __init__(self):
return
def load(self, artifacts_uri: str) -> None:
prediction_utils.download_model_artifacts(artifacts_uri)
if os.path.exists(prediction.MODEL_FILENAME_PKL):
booster = pickle.load(open(prediction.MODEL_FILENAME_PKL, "rb"))
else:
X, y = make_blobs(n_samples=100, centers=2, n_features=2, random_state=1)
model = XGBClassifier()
model.fit(X, y)
booster = model.get_booster()
self._booster = booster
def preprocess(self, prediction_input: dict) -> xgb.DMatrix:
instances = prediction_input["instances"]
return xgb.DMatrix(instances)
def predict(self, instances: xgb.DMatrix) -> np.ndarray:
return self._booster.predict(instances)
def postprocess(self, prediction_results: np.ndarray) -> dict:
return {"predictions": prediction_results.tolist()}
@@ -0,0 +1,41 @@
import os
import numpy as np
import pandas as pd
import pickle
import xgboost as xgb
from google.cloud.aiplatform.constants import prediction
from google.cloud.aiplatform.utils import prediction_utils
from google.cloud.aiplatform.prediction.predictor import Predictor
class XGBRankerPredictor(Predictor):
def __init__(self):
return
def load(self, artifacts_uri: str) -> None:
prediction_utils.download_model_artifacts(artifacts_uri)
if os.path.exists(prediction.MODEL_FILENAME_PKL):
booster = pickle.load(open(prediction.MODEL_FILENAME_PKL, "rb"))
self._booster = booster
else:
N = 500
dates = pd.date_range(start='2023-01-01', end='2023-01-12', periods=N)
X = pd.DataFrame(np.random.randn(N, 5), columns=list('ABCDE'), index=dates)
y = pd.Series(np.random.randint(0, 10, size=N), index=dates, name='label')
group = X.groupby(dates + pd.offsets.MonthEnd(0)).size()
sample_weight = pd.Series(np.arange(len(group)), index=group.index)
model = xgb.XGBRanker(objective='rank:pairwise', max_depth=3, learning_rate=0.1, booster='gbtree', tree_method='hist', n_jobs=4, n_estimators=50, enable_categorical=False, random_state=42)
model.fit(X=X, y=y, group=group, sample_weight=sample_weight, verbose=True)
booster = model.get_booster()
self._booster = booster
def preprocess(self, prediction_input: dict) -> xgb.DMatrix:
instances = prediction_input["instances"]
return xgb.DMatrix(instances)
def predict(self, instances: xgb.DMatrix) -> np.ndarray:
return self._booster.predict(instances, output_margin=False, ntree_limit=0)
def postprocess(self, prediction_results: np.ndarray) -> dict:
return {"predictions": prediction_results.tolist()}
@@ -66,7 +66,7 @@ mkdir -p "$local_folder"
mkdir -p "$output_folder"
# Download the content from the GCS URI
gsutil -m cp -r "$gcs_dataset_path"/* "$local_folder/"
gcloud storage cp --recursive "$gcs_dataset_path"/* "$local_folder/"
# Process files in the local folder
for file in "$local_folder"/*; do
@@ -122,23 +122,23 @@ cp -r "$output_folder" "$images_folder"/images_2
pushd "$images_folder"/images_2
ls | xargs -P 8 -I {} mogrify -resize 50% {}
popd
gsutil -m cp -r "$images_folder"/images_2/* "$gcs_experiment_path"/data/images_2
gcloud storage cp --recursive "$images_folder"/images_2/* "$gcs_experiment_path"/data/images_2
cp -r "$output_folder" "$images_folder"/images_4
pushd "$images_folder"/images_4
ls | xargs -P 8 -I {} mogrify -resize 25% {}
popd
gsutil -m cp -r "$images_folder"/images_4/* "$gcs_experiment_path"/data/images_4
gcloud storage cp --recursive "$images_folder"/images_4/* "$gcs_experiment_path"/data/images_4
cp -r "$output_folder" "$images_folder"/images_8
pushd "$images_folder"/images_8
ls | xargs -P 8 -I {} mogrify -resize 12.5% {}
popd
gsutil -m cp "$images_folder"/images_8/* "$gcs_experiment_path"/data/images_8
gcloud storage cp "$images_folder"/images_8/* "$gcs_experiment_path"/data/images_8
# Copy images and sparse reconstruction files to gcs experiment folder.
gsutil -m cp "$images_folder"/images/* "$gcs_experiment_path"/data/images
gsutil -m cp -r "$local_folder"/sparse "$gcs_experiment_path"/data
gsutil -m cp "$local_folder"/database.db "$gcs_experiment_path"/data
gcloud storage cp "$images_folder"/images/* "$gcs_experiment_path"/data/images
gcloud storage cp --recursive "$local_folder"/sparse "$gcs_experiment_path"/data
gcloud storage cp "$local_folder"/database.db "$gcs_experiment_path"/data
echo "Processing complete."
@@ -99,14 +99,14 @@ create_dir_if_not_exists "$CHECKPOINTS_PATH"
touch "$local_experiment_path/$exp_folder_name/log_render.txt"
# Copy experiment from GCS bucket to local
gsutil -m cp -r "${args[-gcs_experiment_path]}/data" "$local_experiment_path/$exp_folder_name" || exit 1
gsutil -m cp -r "${args[-gcs_experiment_path]}/checkpoints/${training_job_name}/*" "$CHECKPOINTS_PATH" || exit 1
gcloud storage cp --recursive "${args[-gcs_experiment_path]}/data" "$local_experiment_path/$exp_folder_name" || exit 1
gcloud storage cp --recursive "${args[-gcs_experiment_path]}/checkpoints/${training_job_name}/*" "$CHECKPOINTS_PATH" || exit 1
# Check and copy keyframes file.
if [[ -n ${args[-gcs_keyframes_file]} ]]; then
keyframes_file_basename=$(basename "${args[-gcs_keyframes_file]}")
local_keyframes_file="$local_dataset_path/$keyframes_file_basename"
gsutil cp "${args[-gcs_keyframes_file]}" "$local_keyframes_file" || exit 1
gcloud storage cp "${args[-gcs_keyframes_file]}" "$local_keyframes_file" || exit 1
echo "Local keyframe file: $local_keyframes_file"
launch_rendering "$local_keyframes_file"
else
@@ -114,4 +114,4 @@ else
fi
# Copy rendered data back to GCS.
gsutil -m cp -r "$OUTPUT_RENDER_PATH" "${args[-gcs_experiment_path]}/render/${rendering_job_name}"
gcloud storage cp --recursive "$OUTPUT_RENDER_PATH" "${args[-gcs_experiment_path]}/render/${rendering_job_name}"
@@ -74,7 +74,7 @@ create_dir_if_not_exists "$local_experiment_path"
create_dir_if_not_exists "$local_experiment_path/$scene_folder_name"
# Copy experiment from GCS bucket to local.
gsutil -m cp -r "${gcs_experiment_path}/data" "$local_experiment_path/$scene_folder_name" || exit 1
gcloud storage cp --recursive "${gcs_experiment_path}/data" "$local_experiment_path/$scene_folder_name" || exit 1
echo "GCS Experiment: $gcs_experiment_path"
echo "Gin Config File: $gin_config_file"
@@ -89,6 +89,6 @@ accelerate launch train.py --gin_configs="$gin_config_file" \
--gin_bindings="Config.factor = ${factor}" \
--gin_bindings="Config.max_steps = ${max_training_steps}"
gsutil -m rm -r "${gcs_experiment_path}/checkpoints/${training_job_name}"
gsutil -m cp -r "$local_experiment_path/$scene_folder_name/config.gin" "${gcs_experiment_path}/${training_job_name}_config.gin"
gsutil -m cp -r "$local_experiment_path/$scene_folder_name/checkpoints/*/*" "${gcs_experiment_path}/checkpoints/${training_job_name}"
gcloud storage rm --recursive "${gcs_experiment_path}/checkpoints/${training_job_name}"
gcloud storage cp --recursive "$local_experiment_path/$scene_folder_name/config.gin" "${gcs_experiment_path}/${training_job_name}_config.gin"
gcloud storage cp --recursive "$local_experiment_path/$scene_folder_name/checkpoints/*/*" "${gcs_experiment_path}/checkpoints/${training_job_name}"
@@ -1,13 +1,16 @@
"""Common util functions for notebook."""
import base64
from collections.abc import Sequence
import datetime
import io
import json
import os
import subprocess
from typing import Any, Dict, Sequence
import time
from typing import Any
from google import auth
from google.cloud import storage
import matplotlib.pyplot as plt
import numpy as np
@@ -230,7 +233,7 @@ def download_image(url: str) -> str:
base64 encoded image.
"""
response = requests.get(url)
return Image.open(io.BytesIO(response.content))
return Image.open(io.BytesIO(response.content)) # pytype: disable=bad-return-type # pillow-102-upgrade
def resize_image(image: Any, new_width: int = 1000) -> Any:
@@ -281,7 +284,7 @@ def decode_image(
return image
def get_label_map(label_map_yaml_filepath: str) -> Dict[int, str]:
def get_label_map(label_map_yaml_filepath: str) -> dict[int, str]:
"""Returns class id to label mapping given a filepath to the label map.
Args:
@@ -331,6 +334,7 @@ def vqa_predict(
image: Any,
language_code: str = "en",
new_width: int = 1000,
use_dedicated_endpoint: bool = False,
) -> Sequence[str]:
"""Predicts the answer to a question about an image using an Endpoint."""
# Resize and convert image to base64 string.
@@ -354,7 +358,9 @@ def vqa_predict(
"image": resized_image_base64,
})
response = endpoint.predict(instances=instances)
response = endpoint.predict(
instances=instances, use_dedicated_endpoint=use_dedicated_endpoint
)
return [pred.get("response") for pred in response.predictions]
@@ -364,6 +370,7 @@ def caption_predict(
image: Any,
caption_prompt: bool = False,
new_width: int = 1000,
use_dedicated_endpoint: bool = False,
) -> str:
"""Predicts a caption for a given image using an Endpoint."""
# Resize and convert image to base64 string.
@@ -378,7 +385,9 @@ def caption_predict(
instance["prompt"] = caption_prompt_format.format(language_code)
instances = [instance]
response = endpoint.predict(instances=instances)
response = endpoint.predict(
instances=instances, use_dedicated_endpoint=use_dedicated_endpoint
)
return response.predictions[0].get("response")
@@ -387,6 +396,7 @@ def ocr_predict(
ocr_prompt: str,
image: Any,
new_width: int = 1000,
use_dedicated_endpoint: bool = False,
) -> str:
"""Extracts text from a given image using an Endpoint."""
# Resize and convert image to base64 string.
@@ -398,7 +408,9 @@ def ocr_predict(
instance["prompt"] = ocr_prompt
instances = [instance]
response = endpoint.predict(instances=instances)
response = endpoint.predict(
instances=instances, use_dedicated_endpoint=use_dedicated_endpoint
)
return response.predictions[0].get("response")
@@ -407,6 +419,7 @@ def detect_predict(
detect_prompt: str,
image: Any,
new_width: int = 1000,
use_dedicated_endpoint: bool = False,
) -> str:
"""Predicts the answer to a question about an image using an Endpoint."""
# Resize and convert image to base64 string.
@@ -418,10 +431,47 @@ def detect_predict(
instance["prompt"] = detect_prompt
instances = [instance]
response = endpoint.predict(instances=instances)
response = endpoint.predict(
instances=instances, use_dedicated_endpoint=use_dedicated_endpoint
)
return response.predictions[0].get("response")
def copy_model_artifacts(
model_id: str,
model_source: str,
model_destination: str,
) -> None:
"""Copies model artifacts from model_source to model_destination.
model_source and model_destination should be GCS path.
Args:
model_id: The model id.
model_source: The source of the model artifact.
model_destination: The destination of the model artifact.
"""
if not model_source.startswith(GCS_URI_PREFIX):
raise ValueError(
f"{model_source} is not a GCS path starting with {GCS_URI_PREFIX}."
)
if not model_destination.startswith(GCS_URI_PREFIX):
raise ValueError(
f"{model_destination} is not a GCS path starting with {GCS_URI_PREFIX}."
)
model_source = f"{model_source}/{model_id}"
model_destination = f"{model_destination}/{model_id}"
print("Copying model artifact from ", model_source, " to ", model_destination)
subprocess.check_output([
"gcloud",
"storage",
"cp",
"-r",
model_source,
model_destination,
])
def get_quota(project_id: str, region: str, resource_id: str) -> int:
"""Returns the quota for a resource in a region.
@@ -460,6 +510,17 @@ def get_quota(project_id: str, region: str, resource_id: str) -> int:
):
return -1
all_regions_data = quota_data[0]["consumerQuotaLimits"][0]["quotaBuckets"]
# If the quota data does not have dimensions, it is global quota. However,
# global quota may be overridden by regional quota. So we need to check the
# global quota first.
global_quota = -1
if (
all_regions_data
and "dimensions" not in all_regions_data[0]
and "effectiveLimit" in all_regions_data[0]
):
global_quota = int(all_regions_data[0]["effectiveLimit"])
for region_data in all_regions_data:
if (
region_data.get("dimensions")
@@ -469,13 +530,15 @@ def get_quota(project_id: str, region: str, resource_id: str) -> int:
return int(region_data["effectiveLimit"])
else:
return 0
return -1
return global_quota
def get_resource_id(
accelerator_type: str,
is_for_training: bool,
is_spot: bool = False,
is_restricted_image: bool = False,
is_dynamic_workload_scheduler: bool = False,
) -> str:
"""Returns the resource id for a given accelerator type and the use case.
@@ -483,48 +546,76 @@ def get_resource_id(
accelerator_type: The accelerator type.
is_for_training: Whether the resource is used for training. Set false for
serving use case.
is_spot: Whether the resource is used with Spot.
is_restricted_image: Whether the image is hosted in `vertex-ai-restricted`.
is_dynamic_workload_scheduler: Whether the resource is used with Dynamic
Workload Scheduler.
Returns:
The resource id.
"""
accelerator_suffix_map = {
"NVIDIA_TESLA_V100": "nvidia_v100_gpus",
"NVIDIA_TESLA_P100": "nvidia_p100_gpus",
"NVIDIA_L4": "nvidia_l4_gpus",
"NVIDIA_TESLA_A100": "nvidia_a100_gpus",
"NVIDIA_A100_80GB": "nvidia_a100_80gb_gpus",
"NVIDIA_H100_80GB": "nvidia_h100_gpus",
"NVIDIA_H100_MEGA_80GB": "nvidia_h100_mega_gpus",
"NVIDIA_H200_141GB": "nvidia_h200_gpus",
"NVIDIA_TESLA_T4": "nvidia_t4_gpus",
"TPU_V6e": "tpu_v6e",
"TPU_V5e": "tpu_v5e",
"TPU_V3": "tpu_v3",
}
default_training_accelerator_map = {
"NVIDIA_TESLA_V100": "custom_model_training_nvidia_v100_gpus",
"NVIDIA_L4": "custom_model_training_nvidia_l4_gpus",
"NVIDIA_TESLA_A100": "custom_model_training_nvidia_a100_gpus",
"NVIDIA_A100_80GB": "custom_model_training_nvidia_a100_80gb_gpus",
"NVIDIA_H100_80GB": "custom_model_training_nvidia_h100_gpus",
"NVIDIA_TESLA_T4": "custom_model_training_nvidia_t4_gpus",
"TPU_V5e": "custom_model_training_tpu_v5e",
"TPU_V3": "custom_model_training_tpu_v3",
key: f"custom_model_training_{accelerator_suffix_map[key]}"
for key in accelerator_suffix_map
}
dws_training_accelerator_map = {
key: f"custom_model_training_preemptible_{accelerator_suffix_map[key]}"
for key in accelerator_suffix_map
}
restricted_image_training_accelerator_map = {
"NVIDIA_A100_80GB": "restricted_image_training_nvidia_a100_80gb_gpus",
}
serving_accelerator_map = {
"NVIDIA_TESLA_V100": "custom_model_serving_nvidia_v100_gpus",
"NVIDIA_L4": "custom_model_serving_nvidia_l4_gpus",
"NVIDIA_TESLA_A100": "custom_model_serving_nvidia_a100_gpus",
"NVIDIA_A100_80GB": "custom_model_serving_nvidia_a100_80gb_gpus",
"NVIDIA_H100_80GB": "custom_model_serving_nvidia_h100_gpus",
"NVIDIA_TESLA_T4": "custom_model_serving_nvidia_t4_gpus",
"TPU_V5e": "custom_model_serving_tpu_v5e",
spot_serving_accelerator_map = {
key: f"custom_model_serving_preemptible_{accelerator_suffix_map[key]}"
for key in accelerator_suffix_map
}
serving_accelerator_map = {
key: f"custom_model_serving_{accelerator_suffix_map[key]}"
for key in accelerator_suffix_map
}
if is_for_training:
if is_restricted_image and is_dynamic_workload_scheduler:
raise ValueError(
"Dynamic Workload Scheduler does not work for restricted image"
" training."
)
training_accelerator_map = (
restricted_image_training_accelerator_map
if is_restricted_image
else default_training_accelerator_map
)
if accelerator_type in training_accelerator_map:
return training_accelerator_map[accelerator_type]
if is_dynamic_workload_scheduler:
return dws_training_accelerator_map[accelerator_type]
else:
return training_accelerator_map[accelerator_type]
else:
raise ValueError(
f"Could not find accelerator type: {accelerator_type} for training."
)
else:
if accelerator_type in serving_accelerator_map:
return serving_accelerator_map[accelerator_type]
if is_dynamic_workload_scheduler:
raise ValueError("Dynamic Workload Scheduler does not work for serving.")
accelerator_map = (
spot_serving_accelerator_map if is_spot else serving_accelerator_map
)
if accelerator_type in accelerator_map:
return accelerator_map[accelerator_type]
else:
raise ValueError(
f"Could not find accelerator type: {accelerator_type} for serving."
@@ -537,11 +628,30 @@ def check_quota(
accelerator_type: str,
accelerator_count: int,
is_for_training: bool,
is_spot: bool = False,
is_restricted_image: bool = False,
):
"""Checks if the project and the region has the required quota."""
is_dynamic_workload_scheduler: bool = False,
) -> None:
"""Checks if the project and the region has the required quota.
Args:
project_id: The project id.
region: The region.
accelerator_type: The accelerator type.
accelerator_count: The number of accelerators to check quota for.
is_for_training: Whether the resource is used for training. Set false for
serving use case.
is_spot: Whether the resource is used with Spot.
is_restricted_image: Whether the image is hosted in `vertex-ai-restricted`.
is_dynamic_workload_scheduler: Whether the resource is used with Dynamic
Workload Scheduler.
"""
resource_id = get_resource_id(
accelerator_type, is_for_training, is_restricted_image
accelerator_type,
is_for_training=is_for_training,
is_spot=is_spot,
is_restricted_image=is_restricted_image,
is_dynamic_workload_scheduler=is_dynamic_workload_scheduler,
)
quota = get_quota(project_id, region, resource_id)
quota_request_instruction = (
@@ -562,3 +672,76 @@ def check_quota(
f"Quota not enough for {resource_id} in {region}: {quota} <"
f" {accelerator_count}. {quota_request_instruction}"
)
def get_deploy_source() -> str:
"""Gets deploy_source string based on running environment."""
vertex_product = os.environ.get("VERTEX_PRODUCT", "")
match vertex_product:
case "COLAB_ENTERPRISE":
return "notebook_colab_enterprise"
case "WORKBENCH_INSTANCE":
return "notebook_workbench"
case _:
# Legacy workbench, legacy colab, or other custom environments.
return "notebook_environment_unspecified"
def _is_operation_done(op_name: str, region: str) -> bool:
"""Checks if the operation is done.
Args:
op_name: The name of the operation to poll.
region: The region of the operation.
Returns:
True if the operation is done, False otherwise.
Raises:
ValueError: If the operation failed.
"""
creds, _ = auth.default()
auth_req = auth.transport.requests.Request()
creds.refresh(auth_req)
headers = {
"Authorization": f"Bearer {creds.token}",
}
url = f"https://{region}-aiplatform.googleapis.com/ui/{op_name}"
response = requests.get(url, headers=headers)
operation_data = response.json()
if "error" in operation_data:
raise ValueError(f"Operation failed: {operation_data['error']}")
return operation_data.get("done", False)
def poll_and_wait(
op_name: str, region: str, total_wait: int, interval: int = 60
) -> None:
"""Polls the operation and waits for it to complete.
Args:
op_name: The name of the operation to poll.
region: The region of the operation.
total_wait: The total wait time in seconds.
interval: The interval between each poll in seconds.
Raises:
TimeoutError: If the operation times out.
"""
start_time = time.time()
while True:
if _is_operation_done(op_name, region):
break
time_elapsed = time.time() - start_time
if time_elapsed > total_wait:
raise TimeoutError(
f"Operation timed out after {int(time_elapsed)} seconds."
)
print(
"\rStill waiting for operation... Elapsed time in seconds:"
f" {int(time_elapsed):<6}",
end="",
flush=True,
)
time.sleep(interval)
@@ -0,0 +1,570 @@
"""Functions for dataset validation.
This tool is used to validate the dataset against the given template.
"""
import json
import multiprocessing
import os
import subprocess
from typing import Any, Callable, Dict, Tuple, Union
from absl import logging
import accelerate
import datasets
import transformers
GCS_URI_PREFIX = "gs://"
GCSFUSE_URI_PREFIX = "/gcs/"
LOCAL_BASE_MODEL_DIR = "/tmp/base_model_dir"
LOCAL_TEMPLATE_DIR = "/tmp/template_dir"
_TEMPLATE_DIRNAME = "templates"
_VERTEX_AI_SAMPLES_GITHUB_REPO_NAME = "vertex-ai-samples"
_VERTEX_AI_SAMPLES_GITHUB_TEMPLATE_DIR = (
"community-content/vertex_model_garden/model_oss/peft/train/vmg/templates"
)
_MODELS_REQUIRING_PAD_TOKEN = ("llama", "falcon", "mistral", "mixtral")
_MODELS_REQUIRING_EOS_TOEKN = ("gemma-2b", "gemma-7b")
_DESCRIPTION_KEY = "description"
_SOURCE_KEY = "source"
_PROMPT_INPUT_KEY = "prompt_input"
_PROMPT_NO_INPUT_KEY = "prompt_no_input"
_RESPONSE_SEPARATOR = "response_separator"
_INSTRUCTION_SEPARATOR = "instruction_separator"
_CHAT_TEMPLATE_KEY = "chat_template"
_KNOWN_KEYS = (
_DESCRIPTION_KEY,
_SOURCE_KEY,
_PROMPT_INPUT_KEY,
_PROMPT_NO_INPUT_KEY,
_RESPONSE_SEPARATOR,
_INSTRUCTION_SEPARATOR,
_CHAT_TEMPLATE_KEY,
)
def is_gcs_path(input_path: str) -> bool:
"""Checks if the input path is a Google Cloud Storage (GCS) path.
Args:
input_path: The input path to be checked.
Returns:
True if the input path is a GCS path, False otherwise.
"""
return input_path is not None and input_path.startswith(GCS_URI_PREFIX)
def force_gcs_fuse_path(gcs_uri: str) -> str:
"""Converts gs:// uris to their /gcs/ equivalents. No-op for other uris.
Args:
gcs_uri: The GCS URI to convert.
Returns:
The converted GCS URI.
"""
if is_gcs_path(gcs_uri):
return GCSFUSE_URI_PREFIX + gcs_uri[len(GCS_URI_PREFIX) :]
else:
return gcs_uri
def download_gcs_uri_to_local(
gcs_uri: str,
destination_dir: str = LOCAL_BASE_MODEL_DIR,
check_path_exists: bool = True,
) -> str:
"""Downloads GCS URI to local.
If GCS URI is a directory, gs://some/folder is downloaded to
/destination_dir/folder. If GCS URI is a file, gs://some/file is downloaded to
/destination_dir/file.
Args:
gcs_uri: GCS URI to download.
destination_dir: Local directory directory.
check_path_exists: Whether to check if the path exists.
Returns:
Local path to target folder/file.
"""
target = os.path.join(
destination_dir,
os.path.basename(os.path.normpath(gcs_uri)),
)
if check_path_exists and os.path.exists(target):
logging.info("File %s already exists.", target)
return target
if accelerate.PartialState().is_local_main_process:
logging.info(
"Downloading file(s) from %s to %s...", gcs_uri, destination_dir
)
if not os.path.exists(destination_dir):
os.mkdir(destination_dir)
subprocess.check_output([
"gcloud",
"storage",
"cp",
"--recursive",
gcs_uri,
destination_dir,
])
logging.info("Downloaded file(s) from %s to %s.", gcs_uri, destination_dir)
# Make sure ALL processes process to next step after data downloading is done.
# It matters for the main process to wait for other processes as well.
accelerate.PartialState().wait_for_everyone()
return target
def get_template(template_path: str) -> Dict[str, str]:
"""Gets the template dictionary given the file path.
Args:
template_path: Path to the template file.
Returns:
A dictionary of the template.
Raises:
ValueError: If the template file does not exist or contains unknown keys.
"""
if is_gcs_path(template_path):
template_path = force_gcs_fuse_path(template_path)
elif not os.path.isfile(template_path):
template_path = os.path.join(
os.path.dirname(__file__),
_TEMPLATE_DIRNAME,
template_path + ".json",
)
if not os.path.isfile(template_path):
raise ValueError(f"Template file {template_path} does not exist.")
with open(template_path, "r") as f:
template_json: dict[str, str] = json.load(f)
for key in template_json:
if key not in _KNOWN_KEYS:
raise ValueError(f"Unknown key {key} in template {template_path}.")
return template_json
def get_response_separator(template_json: Dict[str, str]) -> Union[str, None]:
return template_json.get(_RESPONSE_SEPARATOR, None)
def get_instruction_separator(
template_json: Dict[str, str],
) -> Union[str, None]:
return template_json.get(_INSTRUCTION_SEPARATOR, None)
def _format_template_fn(
template: str,
input_column: str,
tokenizer: transformers.PreTrainedTokenizer | None = None,
) -> Callable[[Dict[str, str]], Dict[str, str]]:
"""Formats a dataset example according to a template.
Args:
template: Name of the JSON template file under `templates/` or GCS path to
the template file.
input_column: The input column in the dataset to be used or updated by the
template. If it does not exist, the template's `prompt_no_input` will be
used, and the input_column will be created.
tokenizer: The tokenizer to use for chat_template templates.
Returns:
A function that formats data according to the template.
"""
template_json = get_template(template)
if _CHAT_TEMPLATE_KEY not in template_json:
def format_fn(example: Dict[str, str]) -> Dict[str, str]:
format_dict = {key: value for key, value in example.items()}
if format_dict.get(input_column):
format_str = template_json[_PROMPT_INPUT_KEY]
elif _PROMPT_NO_INPUT_KEY in template_json:
format_str = template_json[_PROMPT_NO_INPUT_KEY]
else:
raise KeyError(
f"The template {os.path.basename(template)} does not contain"
f" {_PROMPT_INPUT_KEY} or {_PROMPT_NO_INPUT_KEY} key."
)
try:
return {input_column: format_str.format(**format_dict)}
except KeyError as e:
raise KeyError(
f"The template {os.path.basename(template)} contains a key {e} in"
f" {_PROMPT_INPUT_KEY} or {_PROMPT_NO_INPUT_KEY} that does not"
" exist in the dataset example. The dataset example looks like"
f" {format_dict}."
) from e
return format_fn
elif (
_PROMPT_INPUT_KEY in template_json
or _PROMPT_NO_INPUT_KEY in template_json
):
raise ValueError(
f"chat_template templates do not support {_PROMPT_INPUT_KEY} or"
f" {_PROMPT_NO_INPUT_KEY} templates."
)
else:
if tokenizer is None:
raise ValueError("A tokenizer is required for chat_template templates.")
# Assign HuggingFace jinja template.
tokenizer.chat_template = template_json[_CHAT_TEMPLATE_KEY]
def format_fn(example: Dict[str, str]) -> Dict[str, str]:
try:
return {
input_column: tokenizer.apply_chat_template(
example[input_column],
tokenize=False,
add_generation_prompt=False,
)
}
except KeyError as e:
raise KeyError(
f"The template {os.path.basename(template)} contains a key {e} in"
f" {_CHAT_TEMPLATE_KEY} that does not exist in the dataset example."
) from e
return format_fn
def _get_split_string(
split: str,
dataset_percent: int | None = None,
dataset_k_rows: int | None = None,
) -> str:
"""Gets the formatted split string for the dataset.
This is used to format the split string as per
https://huggingface.co/docs/datasets/v2.21.0/loading#slice-splits. Also, this
function will only be used to load the partial dataset for validating the
dataset against the template.
Args:
split: Split of the dataset.
dataset_percent: The percentage of the dataset to load.
dataset_k_rows: The top k sequences to load from the dataset.
Returns:
A formatted split string.
"""
# Validate the dataset_percent and dataset_k_rows values.
if dataset_percent and dataset_k_rows:
raise ValueError(
"You can set either validate_percentage_of_dataset or"
" validate_k_rows_of_dataset, but not both."
)
if dataset_percent:
logging.info("Loading %d percent of the dataset...", dataset_percent)
return f"{split}[:{dataset_percent}%]"
if dataset_k_rows:
logging.info("Loading top %d rows of the dataset...", dataset_k_rows)
return f"{split}[:{dataset_k_rows}]"
return split
def _github_template_path(template: str) -> str:
"""Generates the path to the template in the Vertex AI Samples GitHub repo.
Args:
template: Name of the template.
Returns:
The path to the template in the Vertex AI Samples GitHub repo.
"""
# vertex-ai-samples directory may lie under separate directory depending on
# the scratch_dir parameter in the notebook execution environment.
vertex_ai_samples_abs_path = os.getcwd().split(
_VERTEX_AI_SAMPLES_GITHUB_REPO_NAME
)[0]
return os.path.join(
vertex_ai_samples_abs_path,
_VERTEX_AI_SAMPLES_GITHUB_REPO_NAME,
_VERTEX_AI_SAMPLES_GITHUB_TEMPLATE_DIR,
template + ".json",
)
def _get_dataset(
dataset_name: str,
split: str,
num_proc: int | None = None,
) -> datasets.DatasetDict:
"""Gets a dataset.
Args:
dataset_name: Name of the dataset or path to a custom dataset.
split: Split of the dataset.
num_proc: Number of processors to use.
Returns:
A dataset.
"""
dataset_name = force_gcs_fuse_path(dataset_name)
if os.path.isfile(dataset_name):
# Custom dataset.
return datasets.load_dataset(
"json",
data_files=[dataset_name],
split=split,
num_proc=num_proc,
)
# HF dataset.
return datasets.load_dataset(dataset_name, split=split, num_proc=num_proc)
def should_add_pad_token(model_id: str) -> bool:
"""Returns whether the model requires adding a special pad token.
Args:
model_id: The name of the model.
Returns:
True if the model requires adding a special pad token, False otherwise.
"""
return any(s.lower() in model_id.lower() for s in _MODELS_REQUIRING_PAD_TOKEN)
def should_add_eos_token(model_id: str) -> bool:
"""Returns whether the model requires adding a special eos token.
Args:
model_id: The name of the model.
Returns:
True if the model requires adding a special eos token, False otherwise.
"""
return any(m in model_id for m in _MODELS_REQUIRING_EOS_TOEKN)
def load_tokenizer(
pretrained_model_id: str,
padding_side: str | None = None,
access_token: str | None = None,
) -> transformers.AutoTokenizer:
"""Loads tokenizer based on `pretrained_model_id`.
Args:
pretrained_model_id: The name of the pretrained model.
padding_side: The side to pad the input on.
access_token: The access token to use for the tokenizer.
Returns:
The tokenizer.
"""
tokenizer_kwargs = {}
if should_add_eos_token(pretrained_model_id):
tokenizer_kwargs["add_eos_token"] = True
if padding_side:
tokenizer_kwargs["padding_side"] = padding_side
with accelerate.PartialState().local_main_process_first():
tokenizer = transformers.AutoTokenizer.from_pretrained(
pretrained_model_id,
trust_remote_code=False,
use_fast=True,
token=access_token,
**tokenizer_kwargs,
)
if should_add_pad_token(pretrained_model_id):
tokenizer.add_special_tokens({"pad_token": "[PAD]"})
return tokenizer
def get_filtered_dataset(
dataset: Any,
input_column: str,
max_seq_length: int,
tokenizer: transformers.PreTrainedTokenizer,
) -> Any:
"""Returns the dataset by removing examples that are longer than max_seq_length.
Args:
dataset: The dataset to filter.
input_column: The input column in the dataset to be used.
max_seq_length: The maximum sequence length.
tokenizer: The tokenizer.
"""
actual_dataset_length = len(dataset)
filtered_dataset = dataset.filter(
lambda x: len(tokenizer(x[input_column])["input_ids"]) <= max_seq_length
)
filtered_dataset_length = len(filtered_dataset)
if actual_dataset_length != filtered_dataset_length:
examples_removed_percent = (
(actual_dataset_length - filtered_dataset_length)
* 100
/ actual_dataset_length
)
logging.info(
"(%.2f%%) of examples token length is <= max-seq-length(%d); (%.2f%%) >"
" max-seq-length. Filtering out %d example(s) which are longer than"
" max-seq-length.",
100 - examples_removed_percent,
max_seq_length,
examples_removed_percent,
actual_dataset_length - filtered_dataset_length,
)
return filtered_dataset
def format_dataset(
dataset: datasets.Dataset,
input_column: str,
template: str = None,
tokenizer: transformers.PreTrainedTokenizer | None = None,
) -> datasets.Dataset:
"""Takes a raw dataset and formats it using a template and tokenizer.
Args:
dataset: The raw (unprocessed) dataset to format.
input_column: The input column in the dataset to be used or updaded by the
template. If it does not exist, the template's `prompt_no_input` will be
used, and the input_column will be created.
template: Name of the JSON template file under `templates/` or GCS path to
the template file.
tokenizer: The tokenizer to use for chat_template templates.
Returns:
A dataset compatible with the template.
"""
return dataset.map(
_format_template_fn(
template,
input_column=input_column,
tokenizer=tokenizer,
)
)
def load_dataset_with_template(
dataset_name: str,
split: str,
input_column: str,
template: str = None,
tokenizer: transformers.PreTrainedTokenizer | None = None,
) -> Tuple[Any, Any]:
"""Loads dataset with templates.
Args:
dataset_name: Name of the dataset or path to a custom dataset.
split: Split of the dataset.
input_column: The input column in the dataset to be used or updaded by the
template. If it does not exist, the template's `prompt_no_input` will be
used, and the input_column will be created.
template: Name of the JSON template file under `templates/` or GCS path to
the template file.
tokenizer: The tokenizer to use for chat_template templates.
Returns:
The raw dataset and the dataset compatible with the template.
"""
raw = _get_dataset(dataset_name, split=split)
if template:
templated = format_dataset(raw, input_column, template, tokenizer)
else:
templated = None
return raw, templated
def validate_dataset_with_template(
dataset_name: str,
split: str,
input_column: str,
template: str,
tokenizer: transformers.PreTrainedTokenizer | None = None,
max_seq_length: int | None = None,
use_multiprocessing: bool = False,
validate_percentage_of_dataset: int | None = None,
validate_k_rows_of_dataset: int | None = None,
) -> Any:
"""Validates dataset with templates.
This function will be used to load the dataset and validate it against the
template. In case of validation, we also allow the users to load the dataset
partially by allowing them to read x% or top k rows of the dataset. To
validate the dataset, the template file must be available in the GCS bucket
and the dataset must be available either in the GCS bucket or Hugging Face.
Args:
dataset_name: Name of the dataset or path to a custom dataset.
split: Split of the dataset.
input_column: The input column in the dataset to be used or updaded by the
template. If it does not exist, the template's `prompt_no_input` will be
used, and the input_column will be created.
template: Name of the JSON template file under `templates/` or GCS path to
the template file.
tokenizer: The tokenizer to use for chat_template templates.
max_seq_length: The maximum sequence length.
use_multiprocessing: If True, it will use multiprocessing to load the
dataset.
validate_percentage_of_dataset: The percentage of the dataset to load.
validate_k_rows_of_dataset: The top k sequences to load from the dataset.
Returns:
None if the validation is successful, otherwise returns the error message.
"""
if not template:
raise ValueError("template is required for validate_dataset.")
if not dataset_name:
raise ValueError("dataset_name is empty.")
if not split:
raise ValueError("split is empty.")
split = _get_split_string(
split,
validate_percentage_of_dataset,
validate_k_rows_of_dataset,
)
num_proc = multiprocessing.cpu_count() if use_multiprocessing else 1
# gcsfuse cannot be used from the notebook runtime env. Hence, we have
# to download dataset and template from gcs to local.
if is_gcs_path(dataset_name):
dataset_name = download_gcs_uri_to_local(dataset_name, LOCAL_BASE_MODEL_DIR)
if is_gcs_path(template):
template_path = download_gcs_uri_to_local(template, LOCAL_TEMPLATE_DIR)
elif os.path.isfile(_github_template_path(template)):
template_path = _github_template_path(template)
else:
raise ValueError(
f"Template file {template} does not exist. To validate the"
" dataset, please provide a valid GCS path for the template or a valid"
" template name from"
f" https://github.com/GoogleCloudPlatform/{_VERTEX_AI_SAMPLES_GITHUB_REPO_NAME}/tree/main/{_VERTEX_AI_SAMPLES_GITHUB_TEMPLATE_DIR}."
)
dataset = format_dataset(
_get_dataset(dataset_name, split, num_proc),
input_column,
template_path,
tokenizer,
)
if tokenizer is not None:
get_filtered_dataset(
dataset=dataset,
input_column=input_column,
max_seq_length=max_seq_length,
tokenizer=tokenizer,
)
print(
"Dataset {} is compatible with the {} template.".format(
os.path.basename(dataset_name), os.path.basename(template)
)
)
@@ -1,144 +0,0 @@
"""Causal language modeling with LoRA models."""
# pylint: disable=g-importing-member
from datasets import load_dataset
from peft import get_peft_model
from peft import LoraConfig
import torch
from torch import nn
import transformers
from transformers import AutoModelForCausalLM
from transformers import AutoTokenizer
from transformers import BitsAndBytesConfig
from transformers import TrainingArguments
from typing import List
from util import constants
def finetune_causal_language_modeling(
pretrained_model_id: str,
dataset_name: str,
output_dir: str,
precision_mode: str = None,
lora_rank: int = 16,
lora_alpha: int = 32,
lora_dropout: float = 0.05,
target_modules: List[str] = constants.CAUSAL_LANGUAGE_MODELING_LORA_TARGET_MODULES,
warmup_steps: int = 10,
max_steps: int = 10,
learning_rate: float = 2e-4,
local_pretrained_model_id: str = None,
) -> None:
"""Finetunes causal language modelings."""
if precision_mode == constants.PRECISION_MODE_32:
model = AutoModelForCausalLM.from_pretrained(
local_pretrained_model_id
if local_pretrained_model_id
else pretrained_model_id,
torch_dtype=torch.float32,
device_map="auto",
)
elif precision_mode == constants.PRECISION_MODE_16:
model = AutoModelForCausalLM.from_pretrained(
local_pretrained_model_id
if local_pretrained_model_id
else pretrained_model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)
elif precision_mode == constants.PRECISION_MODE_8:
quantization_config = BitsAndBytesConfig(
load_in_8bit=True, int8_threshold=0
)
model = AutoModelForCausalLM.from_pretrained(
local_pretrained_model_id
if local_pretrained_model_id
else pretrained_model_id,
torch_dtype=torch.float16,
device_map="auto",
quantization_config=quantization_config,
)
else:
quantization_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16,
)
model = AutoModelForCausalLM.from_pretrained(
local_pretrained_model_id
if local_pretrained_model_id
else pretrained_model_id,
device_map="auto",
torch_dtype=torch.bfloat16,
quantization_config=quantization_config,
)
tokenizer = AutoTokenizer.from_pretrained(
local_pretrained_model_id
if local_pretrained_model_id
else pretrained_model_id
)
if "llama" in pretrained_model_id:
tokenizer.pad_token = "[PAD]"
for param in model.parameters():
# Freezes the model - train adapters later.
param.requires_grad = False
if param.ndim == 1:
# Casts the small parameters (e.g. layernorm) to fp32 for stability.
param.data = param.data.to(torch.float32)
# Reduces the number of stored activations.
model.gradient_checkpointing_enable()
model.enable_input_require_grads()
class CastOutputToFloat(nn.Sequential):
def forward(self, x):
return super().forward(x).to(torch.float32)
model.lm_head = CastOutputToFloat(model.lm_head)
config = LoraConfig(
r=lora_rank,
lora_alpha=lora_alpha,
target_modules=target_modules,
lora_dropout=lora_dropout,
bias="none",
task_type="CAUSAL_LM",
)
model = get_peft_model(model, config)
model.print_trainable_parameters()
data = load_dataset(dataset_name)
data = data.map(
lambda samples: tokenizer(samples["quote"]),
batched=True,
)
trainer = transformers.Trainer(
model=model,
train_dataset=data["train"],
args=TrainingArguments(
per_device_train_batch_size=4,
gradient_accumulation_steps=4,
warmup_steps=warmup_steps,
max_steps=max_steps,
learning_rate=learning_rate,
fp16=True,
logging_steps=1,
output_dir=output_dir,
ddp_find_unused_parameters=False,
),
data_collator=transformers.DataCollatorForLanguageModeling(
tokenizer,
mlm=False,
),
)
# Silence the warnings. Please re-enable for inference!
model.config.use_cache = False
trainer.train()
model.save_pretrained(output_dir)
@@ -0,0 +1,28 @@
# Base on pytorch-cuda image.
FROM pytorch/pytorch:2.0.0-cuda11.7-cudnn8-devel
# Install tools.
ENV DEBIAN_FRONTEND=noninteractive
RUN apt-get update
RUN apt-get install -y --no-install-recommends apt-utils
RUN apt-get install -y --no-install-recommends curl
RUN apt-get install -y --no-install-recommends wget
RUN apt-get install -y --no-install-recommends git
# Install libraries.
ENV PIP_ROOT_USER_ACTION=ignore
RUN python3 -m pip install --upgrade pip
RUN pip install tokenizers==0.13.3
RUN pip install accelerate==0.21.0
RUN pip install sentencepiece==0.1.99
RUN pip install datasets==2.14.4
RUN pip install protobuf==4.24.1
# Install transformers
RUN git clone https://github.com/huggingface/transformers.git
WORKDIR transformers
# Pin the commit to add-code-llama 08/25/2023
RUN git reset --hard 015f8e110d270a0ad42de4ae5b98198d69eb1964
RUN pip install -e .
ENTRYPOINT ["python","src/transformers/models/llama/convert_llama_weights_to_hf.py"]
@@ -0,0 +1,22 @@
# Dockerfile for Language Model Conversion.
#
# To build:
# docker build -f model_oss/peft/dockerfile/conversion.Dockerfile . -t ${YOUR_IMAGE_TAG}
#
# To push to gcr:
# docker tag ${YOUR_IMAGE_TAG} gcr.io/${YOUR_PROJECT}/${YOUR_IMAGE_TAG}
# docker push gcr.io/${YOUR_PROJECT}/${YOUR_IMAGE_TAG}
FROM tensorflow/build:2.14-python3.8
RUN git clone https://github.com/facebookresearch/llama-recipes.git && \
cd llama-recipes && \
pip install -r requirements.txt && \
pip freeze | grep transformers && \
git clone https://github.com/huggingface/transformers.git && \
cd transformers && \
pip install protobuf
WORKDIR /llama-recipes/transformers
ENTRYPOINT ["python","src/transformers/models/llama/convert_llama_weights_to_hf.py"]
@@ -7,39 +7,40 @@
# docker tag ${YOUR_IMAGE_TAG} gcr.io/${YOUR_PROJECT}/${YOUR_IMAGE_TAG}
# docker push gcr.io/${YOUR_PROJECT}/${YOUR_IMAGE_TAG}
FROM pytorch/torchserve:0.7.0-gpu
FROM pytorch/torchserve:0.11.0-gpu
USER root
ENV infer_port=7080
ENV mng_port=7081
ENV model_name="peft_serving"
ENV INFER_PORT=7080
ENV MNG_PORT=7081
ENV MODEL="peft_serving"
ENV PATH="/home/model-server/:${PATH}"
RUN apt-get update && apt-get install -y --no-install-recommends \
RUN apt-get update && apt-get -y upgrade && apt-get install -y --no-install-recommends \
curl \
wget \
vim \
git \
git-lfs
RUN git lfs install
RUN apt-get autoremove -y
# Install libraries.
ENV PIP_ROOT_USER_ACTION=ignore
RUN python3 -m pip install --upgrade pip
RUN pip install --upgrade torch==2.0.1
RUN pip install --upgrade torch==2.0.1 --index-url https://download.pytorch.org/whl/cu118
RUN pip install torchvision==0.15.2
RUN pip install tokenizers==0.13.3
RUN pip install accelerate==0.21.0
RUN pip install sentencepiece==0.1.99
RUN pip install grpcio-status==1.33.2
RUN pip install protobuf==3.19.6
RUN python3 -m pip install --no-cache-dir git+https://github.com/huggingface/peft.git
RUN pip install peft==0.5.0
RUN pip install datasets==2.14.4
RUN pip install triton==2.0.0.dev20221120
RUN pip install triton==3.0.0
RUN pip install xformers==0.0.20
RUN pip install google-cloud-storage==2.7.0
RUN pip install absl-py==1.4.0
RUN pip install google-cloud-storage
RUN pip install absl-py
RUN pip install scipy==1.10.1
RUN pip install evaluate==0.4.0
RUN pip install scikit-learn==1.2.2
@@ -47,52 +48,43 @@ RUN pip install loralib==0.1.1
RUN pip install bitsandbytes==0.39.0
RUN pip install trl==0.4.4
RUN pip install einops==0.6.1
# Install diffusers from source.
RUN git clone --depth 1 --branch v0.16.1 https://github.com/huggingface/diffusers.git
WORKDIR diffusers
RUN pip install -e .
WORKDIR /home/model-server
# Install transformers from source.
RUN git clone --depth 1 --branch v4.31.0 https://github.com/huggingface/transformers.git
# The patch is used to change the transformers loading model behavior:
# 1) For models on Huggingface hub: if the model has multiple shards, each shard
# will be downloaded separately and get deleted after loading to GPU.
# 2) For models on local disk: if a model bin file is actually a text file
# recording a GCS path, the model file will be downloaded and get deleted
# after loading to GPU.
COPY model_oss/peft/hf_transformers_lazy_download.patch /home/model-server/hf_transformers_lazy_download.patch
WORKDIR transformers
RUN git apply /home/model-server/hf_transformers_lazy_download.patch
RUN pip install -e .
WORKDIR /home/model-server
RUN pip install optimum==1.13.2
RUN pip install auto-gptq==0.4.2
RUN pip install https://github.com/casper-hansen/AutoAWQ/releases/download/v0.1.7/autoawq-0.1.7+cu118-cp39-cp39-linux_x86_64.whl
RUN pip install diffusers==0.27.2
RUN pip install tiktoken==0.6.0
RUn pip install git+https://github.com/huggingface/transformers.git@76fa17c1663a0efeca7208c20579833365584889
RUN pip install pynvml==11.4.0
RUN pip install -i https://test.pypi.org/simple/ bitsandbytes
# Copy license.
WORKDIR /home/model-server
RUN wget https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/main/LICENSE
# Copy model artifacts.
COPY model_oss/peft/handler.py /home/model-server/handler.py
COPY model_oss/peft/config.properties /home/model-server/config.properties
COPY model_oss/util/ /home/model-server/util/
COPY model_oss/util/pytorch_startup_prober.sh /model_garden/scripts/pytorch_startup_prober.sh
ENV PYTHONPATH /home/model-server/
# Expose ports.
EXPOSE ${infer_port}
EXPOSE ${mng_port}
EXPOSE ${INFER_PORT}
EXPOSE ${MNG_PORT}
# Set environments.
ENV TASK "causal-language-modeling-lora"
ENV MODEL_ID "openlm-research/open_llama_7b"
ENV BASE_MODEL_ID ""
ENV MODEL_ID ""
ENV PRECISION_LOADING_MODE "float16"
ENV FINETUNED_LORA_MODEL_PATH ""
ENV TRUST_REMOTE_CODE ""
# Archive model artifacts and dependencies.
# Do not set --model-file and --serialized-file because model and checkpoint
# will be dynamically loaded in handler.py.
RUN torch-model-archiver \
--model-name=${model_name} \
--model-name=${MODEL} \
--version=1.0 \
--handler=/home/model-server/handler.py \
--runtime=python3 \
@@ -103,5 +95,5 @@ RUN torch-model-archiver \
# Run Torchserve HTTP serve to respond to prediction requests.
CMD ["torchserve", "--start", \
"--ts-config", "/home/model-server/config.properties", \
"--models", "${model_name}=${model_name}.mar", \
"--models", "${MODEL}=${MODEL}.mar", \
"--model-store", "/home/model-server/model-store"]
@@ -1,111 +0,0 @@
# Dockerfile for PEFT Training.
#
# To build:
# docker build -f model_oss/peft/dockerfile/train.Dockerfile . -t ${YOUR_IMAGE_TAG}
#
# To push to gcr:
# docker tag ${YOUR_IMAGE_TAG} gcr.io/${YOUR_PROJECT}/${YOUR_IMAGE_TAG}
# docker push gcr.io/${YOUR_PROJECT}/${YOUR_IMAGE_TAG}
# Builds GPU docker image of PyTorch
# Uses multi-staged approach to reduce size
# Stage 1
# Use base conda image to reduce time
FROM continuumio/miniconda3:latest AS compile-image
# Specify py version
ENV PYTHON_VERSION=3.8
# Install apt libs - copied from https://github.com/huggingface/accelerate/blob/main/docker/accelerate-gpu/Dockerfile
RUN apt-get update && \
apt-get install -y curl git wget software-properties-common git-lfs && \
apt-get clean && \
rm -rf /var/lib/apt/lists*
# Install audio-related libraries
RUN apt-get update && \
apt install -y ffmpeg
RUN apt install -y libsndfile1-dev
RUN git lfs install
# Create our conda env - copied from https://github.com/huggingface/accelerate/blob/main/docker/accelerate-gpu/Dockerfile
RUN conda create --name peft python=${PYTHON_VERSION} ipython jupyter pip
RUN python3 -m pip install --no-cache-dir --upgrade pip
# Below is copied from https://github.com/huggingface/accelerate/blob/main/docker/accelerate-gpu/Dockerfile
# We don't install pytorch here yet since CUDA isn't available
# instead we use the direct torch wheel
ENV PATH /opt/conda/envs/peft/bin:$PATH
# Activate our bash shell
RUN chsh -s /bin/bash
SHELL ["/bin/bash", "-c"]
# Activate the conda env and install transformers + accelerate from source
RUN source activate peft
RUN python3 -m pip install --no-cache-dir git+https://github.com/huggingface/transformers
RUN python3 -m pip install --no-cache-dir git+https://github.com/huggingface/accelerate
RUN python3 -m pip install --no-cache-dir git+https://github.com/huggingface/peft#egg=peft[test]
RUN python3 -m pip install --no-cache-dir bitsandbytes
# Stage 2
FROM nvidia/cuda:11.2.2-cudnn8-devel-ubuntu20.04 AS build-image
COPY --from=compile-image /opt/conda /opt/conda
ENV PATH /opt/conda/bin:$PATH
# Install apt libs
RUN apt-get update && \
apt-get install -y curl git wget vim && \
apt-get clean && \
rm -rf /var/lib/apt/lists*
# Copy license.
RUN wget https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/main/LICENSE
RUN echo "source activate peft" >> ~/.profile
# Install libraries.
RUN pip install --upgrade torch==2.0.1
RUN pip install torchvision==0.15.2
RUN pip install git+https://github.com/huggingface/transformers@de9255de27abfcae4a1f816b904915f0b1e23cd9
RUN pip install transformers -U
RUN pip install accelerate==0.21.0
RUN pip install sentencepiece==0.1.99
RUN pip install grpcio-status==1.33.2
RUN pip install protobuf==3.19.6
RUN python3 -m pip install --no-cache-dir git+https://github.com/huggingface/peft.git
RUN pip install datasets==2.9.0
RUN pip install triton==2.0.0.dev20221120
RUN pip install xformers==0.0.20
RUN pip install Jinja2==3.1.2
RUN pip install ftfy==6.1.1
RUN pip install cloudml-hypertune==0.1.0.dev6
RUN pip install tensorboard==2.12.0
RUN pip install scipy==1.10.1
RUN pip install evaluate==0.4.0
RUN pip install scikit-learn==1.2.2
RUN pip install loralib==0.1.1
RUN pip install bitsandbytes==0.39.0
RUN pip install trl==0.4.4
RUN pip install einops==0.6.1
RUN pip install google-cloud-storage==2.7.0
RUN git clone --depth 1 --branch v0.16.1 https://github.com/huggingface/diffusers.git
WORKDIR diffusers
RUN pip install -e .
# Switch to diffusers examples folder.
WORKDIR examples
# NOTE: use 'sed' to modify train_text_to_image_lora.py to
# fix the bug for accelerator.
RUN sed -i \
"s#logging_dir=logging_dir#project_dir=logging_dir#g" \
text_to_image/train_text_to_image_lora.py
# Config accelerate.
RUN mkdir -p ./vertex_vision_model_garden_peft/
COPY model_oss/peft/train.sh ./vertex_vision_model_garden_peft/train.sh
COPY model_oss/peft/*.py ./vertex_vision_model_garden_peft/
COPY model_oss/util /diffusers/examples/util
ENV PYTHONPATH /diffusers/examples/
# Generate accelerate config at the beginning of docker run.
ENTRYPOINT ["python3", "vertex_vision_model_garden_peft/main.py"]
@@ -72,15 +72,39 @@ class PeftHandler(BaseHandler):
"PRECISION_LOADING_MODE", constants.PRECISION_MODE_16
)
self.task = os.environ.get("TASK", CAUSAL_LANGUAGE_MODELING_LORA)
self.base_model_id = os.environ.get("BASE_MODEL_ID", None)
self.model_id = self.base_model_id
if not self.base_model_id:
self.model_id = os.environ.get("MODEL_ID", "")
trust_remote_code = os.environ.get("TRUST_REMOTE_CODE", None)
if trust_remote_code == "false":
self.trust_remote_code = False
else:
self.trust_remote_code = True
# If present, the path of the model in the container.
aip_storage_dir = os.environ.get("AIP_STORAGE_DIR", None)
# If present, the URI of the model in a google owned GCS bucket.
aip_storage_uri = os.environ.get("AIP_STORAGE_URI", None)
model_id = os.environ.get("MODEL_ID", None)
base_model_id = os.environ.get("BASE_MODEL_ID", None)
self.model_id = None
if aip_storage_dir:
self.model_id = aip_storage_dir
logging.info(f"Loaded base model from AIP_STORAGE_DIR: {self.model_id}.")
elif aip_storage_uri:
self.model_id = aip_storage_uri
logging.info(f"Loaded base model from AIP_STORAGE_URI: {self.model_id}.")
elif model_id:
self.model_id = model_id
logging.info(f"Loaded base model from MODEL_ID: {self.model_id}.")
elif base_model_id:
# Note: BASE_MODEL_ID has been unified with MODEL_ID.
# MODEL_ID should be used whenever possible.
self.model_id = base_model_id
logging.info(f"Loaded base model from BASE_MODEL_ID: {self.model_id}.")
self.quantization = os.environ.get("QUANTIZATION", None)
logging.info(f"Load base model id from MODEL_ID:{self.model_id}.")
if not self.model_id:
self.model_id = os.environ.get("AIP_STORAGE_URI", "")
logging.info(f"Load base model id from AIP_STORAGE_URI: {self.model_id}.")
if not self.model_id:
raise ValueError("Base model id is must be set.")
if fileutils.is_gcs_path(self.model_id):
@@ -101,8 +125,7 @@ class PeftHandler(BaseHandler):
logging.info(
f"Using task:{self.task}, base model:{self.model_id}, lora model:"
f" {self.finetuned_lora_model_path}, and precision"
f" {self.precision_mode}."
f" {self.finetuned_lora_model_path}, precision {self.precision_mode}."
)
self.pipeline = None
@@ -145,11 +168,18 @@ class PeftHandler(BaseHandler):
elif (
self.task == CAUSAL_LANGUAGE_MODELING_LORA or self.task == INSTRUCT_LORA
):
tokenizer = AutoTokenizer.from_pretrained(self.model_id)
tokenizer = AutoTokenizer.from_pretrained(
self.model_id,
trust_remote_code=self.trust_remote_code,
)
logging.debug("Initialized the tokenizer.")
if self.task == CAUSAL_LANGUAGE_MODELING_LORA:
if self.quantization == constants.AWQ:
model = AutoAWQForCausalLM.from_quantized(self.model_id)
model = AutoAWQForCausalLM.from_quantized(
self.model_id,
trust_remote_code=self.trust_remote_code,
)
elif self.quantization == constants.GPTQ or not self.quantization:
if self.precision_mode == constants.PRECISION_MODE_32:
model = AutoModelForCausalLM.from_pretrained(
@@ -157,6 +187,7 @@ class PeftHandler(BaseHandler):
return_dict=True,
torch_dtype=torch.float32,
device_map="auto",
trust_remote_code=self.trust_remote_code,
)
elif self.precision_mode == constants.PRECISION_MODE_16B:
model = AutoModelForCausalLM.from_pretrained(
@@ -164,6 +195,7 @@ class PeftHandler(BaseHandler):
return_dict=True,
torch_dtype=torch.bfloat16,
device_map="auto",
trust_remote_code=self.trust_remote_code,
)
elif self.precision_mode == constants.PRECISION_MODE_16:
model = AutoModelForCausalLM.from_pretrained(
@@ -171,6 +203,7 @@ class PeftHandler(BaseHandler):
return_dict=True,
torch_dtype=torch.float16,
device_map="auto",
trust_remote_code=self.trust_remote_code,
)
elif self.precision_mode == constants.PRECISION_MODE_8:
quantization_config = BitsAndBytesConfig(
@@ -182,6 +215,7 @@ class PeftHandler(BaseHandler):
torch_dtype=torch.float16,
device_map="auto",
quantization_config=quantization_config,
trust_remote_code=self.trust_remote_code,
)
else:
quantization_config = BitsAndBytesConfig(
@@ -195,6 +229,7 @@ class PeftHandler(BaseHandler):
device_map="auto",
torch_dtype=torch.bfloat16,
quantization_config=quantization_config,
trust_remote_code=self.trust_remote_code,
)
else:
raise ValueError(f"Invalid QUANTIZATION value: {self.quantization}")
@@ -203,14 +238,14 @@ class PeftHandler(BaseHandler):
model = AutoModelForCausalLM.from_pretrained(
self.model_id,
torch_dtype=torch.bfloat16,
trust_remote_code=True,
trust_remote_code=self.trust_remote_code,
device_map="auto",
)
except: # pylint: disable=bare-except
model = AutoModelForCausalLM.from_pretrained(
self.model_id,
torch_dtype=torch.bfloat16,
trust_remote_code=True,
trust_remote_code=self.trust_remote_code,
device_map="auto",
)
logging.debug("Initialized the base model.")
@@ -329,4 +364,4 @@ class PeftHandler(BaseHandler):
return f"Prompt:\n{prompt.strip()}\nOutput:\n{output}"
# pylint: enable=logging-fstring-interpolation
# pylint: enable=logging-fstring-interpolation
@@ -1,131 +0,0 @@
diff --git a/src/transformers/modeling_utils.py b/src/transformers/modeling_utils.py
index 45459ed..32527f4 100644
--- a/src/transformers/modeling_utils.py
+++ b/src/transformers/modeling_utils.py
@@ -32,6 +32,8 @@ import torch
from packaging import version
from torch import Tensor, nn
from torch.nn import CrossEntropyLoss
+from huggingface_hub import hf_hub_download
+from google.cloud import storage
from .activations import get_activation
from .configuration_utils import PretrainedConfig
@@ -442,6 +444,29 @@ def load_state_dict(checkpoint_file: Union[str, os.PathLike]):
"""
Reads a PyTorch checkpoint file, returning properly formatted errors if they arise.
"""
+ delete_download = False
+ tmp_dir = "/tmp/model"
+ os.makedirs(tmp_dir, exist_ok=True)
+ if isinstance(checkpoint_file, dict):
+ # Download model file from huggingface
+ print(f"==> Download model from HF: {checkpoint_file}")
+ checkpoint_file = hf_hub_download(
+ local_dir=tmp_dir, local_dir_use_symlinks=False, force_download=True, resume_download=True, **checkpoint_file)
+ delete_download = True
+ else:
+ with open(checkpoint_file, "rb") as f:
+ is_gcs_file = (f.read(2) == b"gs")
+ if is_gcs_file:
+ # Download model file from GCS
+ with open(checkpoint_file, "r") as f:
+ gcs_file = f.read()
+ checkpoint_file = os.path.join(tmp_dir, gcs_file.split("/")[-1])
+ print(f"==> Download model from GCS: {gcs_file} to: {checkpoint_file}")
+ client = storage.Client()
+ with open(checkpoint_file, 'wb') as f:
+ client.download_blob_to_file(gcs_file, f)
+ delete_download = True
+
if checkpoint_file.endswith(".safetensors") and is_safetensors_available():
# Check format of the archive
with safe_open(checkpoint_file, framework="pt") as f:
@@ -455,9 +480,9 @@ def load_state_dict(checkpoint_file: Union[str, os.PathLike]):
raise NotImplementedError(
f"Conversion from a {metadata['format']} safetensors archive to PyTorch is not implemented yet."
)
- return safe_load_file(checkpoint_file)
+ state_dict = safe_load_file(checkpoint_file)
try:
- return torch.load(checkpoint_file, map_location="cpu")
+ state_dict = torch.load(checkpoint_file, map_location="cpu")
except Exception as e:
try:
with open(checkpoint_file) as f:
@@ -478,6 +503,10 @@ def load_state_dict(checkpoint_file: Union[str, os.PathLike]):
f"at '{checkpoint_file}'. "
"If you tried to load a PyTorch model from a TF 2.0 checkpoint, please set from_tf=True."
)
+ if delete_download:
+ print(f"==> Delete downloaded model: {checkpoint_file}")
+ os.remove(checkpoint_file)
+ return state_dict
def set_initialized_submodules(model, state_dict_keys):
@@ -3179,7 +3208,10 @@ class PreTrainedModel(nn.Module, ModuleUtilsMixin, GenerationMixin, PushToHubMix
return mismatched_keys
if resolved_archive_file is not None:
- folder = os.path.sep.join(resolved_archive_file[0].split(os.path.sep)[:-1])
+ if isinstance(resolved_archive_file, str):
+ folder = os.path.sep.join(resolved_archive_file[0].split(os.path.sep)[:-1])
+ else:
+ folder = None
else:
folder = None
if device_map is not None and is_safetensors:
diff --git a/src/transformers/utils/hub.py b/src/transformers/utils/hub.py
index ffed743..4b15770 100644
--- a/src/transformers/utils/hub.py
+++ b/src/transformers/utils/hub.py
@@ -414,20 +414,34 @@ def cached_file(
user_agent = http_user_agent(user_agent)
try:
# Load from URL or cache if already cached
- resolved_file = hf_hub_download(
- path_or_repo_id,
- filename,
- subfolder=None if len(subfolder) == 0 else subfolder,
- repo_type=repo_type,
- revision=revision,
- cache_dir=cache_dir,
- user_agent=user_agent,
- force_download=force_download,
- proxies=proxies,
- resume_download=resume_download,
- use_auth_token=use_auth_token,
- local_files_only=local_files_only,
- )
+ if filename.endswith(".bin"):
+ # NOTE: To save disk we do not download bin file eagerly. Do not support safetensors.
+ resolved_file = dict(
+ repo_id=path_or_repo_id,
+ filename=filename,
+ subfolder=None if len(subfolder) == 0 else subfolder,
+ repo_type=repo_type,
+ revision=revision,
+ user_agent=user_agent,
+ proxies=proxies,
+ use_auth_token=use_auth_token,
+ )
+ print(f"--> Apply lazy download to bin file: {resolved_file}")
+ else:
+ resolved_file = hf_hub_download(
+ path_or_repo_id,
+ filename,
+ subfolder=None if len(subfolder) == 0 else subfolder,
+ repo_type=repo_type,
+ revision=revision,
+ cache_dir=cache_dir,
+ user_agent=user_agent,
+ force_download=force_download,
+ proxies=proxies,
+ resume_download=resume_download,
+ use_auth_token=use_auth_token,
+ local_files_only=local_files_only,
+ )
except RepositoryNotFoundError:
raise EnvironmentError(
@@ -1,95 +0,0 @@
"""Instruct/Chat with LoRA models."""
# pylint: disable=g-importing-member
from datasets import load_dataset
from peft import LoraConfig
import torch
from transformers import AutoModelForCausalLM
from transformers import AutoTokenizer
from transformers import BitsAndBytesConfig
from transformers import TrainingArguments
from trl import SFTTrainer
from typing import List
from util import constants
def finetune_instruct(
pretrained_model_id: str,
dataset_name: str,
output_dir: str,
lora_rank: int = 64,
lora_alpha: int = 16,
lora_dropout: float = 0.1,
target_modules: List[str] = constants.INSTRUCT_LORA_TARGET_MODULES,
warmup_ratio: int = 0.03,
max_steps: int = 10,
max_seq_length: int = 512,
learning_rate: float = 2e-4,
) -> None:
"""Finetunes instruct."""
dataset = load_dataset(dataset_name, split="train")
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.float16,
)
model = AutoModelForCausalLM.from_pretrained(
pretrained_model_id,
quantization_config=bnb_config,
trust_remote_code=True,
)
model.config.use_cache = False
tokenizer = AutoTokenizer.from_pretrained(
pretrained_model_id, trust_remote_code=True
)
tokenizer.pad_token = tokenizer.eos_token
peft_config = LoraConfig(
lora_alpha=lora_alpha,
lora_dropout=lora_dropout,
r=lora_rank,
bias="none",
task_type="CAUSAL_LM",
target_modules=target_modules,
)
per_device_train_batch_size = 4
gradient_accumulation_steps = 4
optim = "paged_adamw_32bit"
save_steps = 10
logging_steps = 10
max_grad_norm = 0.3
lr_scheduler_type = "constant"
training_arguments = TrainingArguments(
output_dir=output_dir,
per_device_train_batch_size=per_device_train_batch_size,
gradient_accumulation_steps=gradient_accumulation_steps,
optim=optim,
save_steps=save_steps,
logging_steps=logging_steps,
learning_rate=learning_rate,
fp16=True,
max_grad_norm=max_grad_norm,
max_steps=max_steps,
warmup_ratio=warmup_ratio,
group_by_length=True,
lr_scheduler_type=lr_scheduler_type,
)
trainer = SFTTrainer(
model=model,
train_dataset=dataset,
peft_config=peft_config,
dataset_text_field="text",
max_seq_length=max_seq_length,
tokenizer=tokenizer,
args=training_arguments,
)
for name, module in trainer.model.named_modules():
if "norm" in name:
module = module.to(torch.float32)
trainer.train()
@@ -1,185 +0,0 @@
"""Main function to start PEFT finetuning."""
import subprocess
from absl import app
from absl import flags
from absl import logging
from peft import causal_language_modeling_lora
from peft import instruct_lora
from peft import sequence_classification_lora
from util import constants
from util import fileutils
_TASK = flags.DEFINE_string(
'task',
constants.CAUSAL_LANGUAGE_MODELING_LORA,
'The supported PEFT tasks.',
)
_PRETRAINED_MODEL_ID = flags.DEFINE_string(
'pretrained_model_id',
None,
'The pretrained model id. Supported models can be causal language modeling'
' models from https://github.com/huggingface/peft/tree/main.',
required=True,
)
_DATASET_NAME = flags.DEFINE_string(
'dataset_name',
None,
'The dataset name in huggingface.',
required=True,
)
_OUTPUT_DIR = flags.DEFINE_string(
'output_dir',
None,
'The output directory.',
required=True,
)
_PRECISION_MODE = flags.DEFINE_string(
'precision_mode',
constants.PRECISION_MODE_16,
'Supported finetuning precision_modes are `{}` and `{}`.'.format(
constants.PRECISION_MODE_8, constants.PRECISION_MODE_16
),
)
_LORA_RANK = flags.DEFINE_integer(
'lora_rank',
16,
'The rank of the update matrices, expressed in int. Lower rank results in'
' smaller update matrices with fewer trainable parameters, referring to'
' https://huggingface.co/docs/peft/conceptual_guides/lora.',
)
_LORA_ALPHA = flags.DEFINE_integer(
'lora_alpha',
32,
'LoRA scaling factor, referring to'
' https://huggingface.co/docs/peft/conceptual_guides/lora.',
)
_LORA_DROPOUT = flags.DEFINE_float(
'lora_dropout',
0.05,
'dropout probability of the LoRA layers, referring to'
' https://huggingface.co/docs/peft/task_guides/token-classification-lora.',
)
_TARGET_MODULES = flags.DEFINE_list(
'target_modules',
constants.CAUSAL_LANGUAGE_MODELING_LORA_TARGET_MODULES,
'The comma separated list of target modules for LoRa training.',
)
_WARMUP_STEPS = flags.DEFINE_integer(
'warmup_steps',
10,
'Number of steps for the warmup in the learning rate scheduler.',
)
_WARMUP_RATIO = flags.DEFINE_float(
'warmup_ratio',
0.03,
'The warmup ratio in the learning rate scheduler.',
)
_MAX_STEPS = flags.DEFINE_integer(
'max_steps',
10,
'Total number of training steps.',
)
_MAX_SEQ_LENGTH = flags.DEFINE_integer(
'max_seq_length',
512,
'The maximum sequence length.',
)
_NUM_EPOCHS = flags.DEFINE_integer(
'num_epochs',
20,
'The number of training epochs.',
)
_BATCH_SIZE = flags.DEFINE_integer(
'batch_size',
32,
'The batch size.',
)
_LEARNING_RATE = flags.DEFINE_float(
'learning_rate',
2e-4,
'The learning rate after the potential warmup period.',
)
def main(_) -> None:
task = _TASK.value
pretrained_model_id = _PRETRAINED_MODEL_ID.value
local_pretrained_model_id = None
if pretrained_model_id.startswith(constants.GCS_URI_PREFIX):
logging.info(
'Start to copy pretrained models locally: %s.', pretrained_model_id
)
fileutils.download_gcs_dir_to_local(
pretrained_model_id, constants.LOCAL_BASE_MODEL_DIR
)
local_pretrained_model_id = constants.LOCAL_BASE_MODEL_DIR
logging.info(
'Finished copying pretrained models locally to: %s.',
local_pretrained_model_id,
)
if task == constants.TEXT_TO_IMAGE_LORA:
subprocess.run(['/bin/bash', 'train.sh'], check=True)
elif task == constants.SEQUENCE_CLASSIFICATION_LORA:
sequence_classification_lora.finetune_sequence_classification(
pretrained_model_id=pretrained_model_id,
dataset_name=_DATASET_NAME.value,
output_dir=_OUTPUT_DIR.value,
lora_rank=_LORA_RANK.value,
lora_alpha=_LORA_ALPHA.value,
lora_dropout=_LORA_DROPOUT.value,
num_epochs=_NUM_EPOCHS.value,
batch_size=_BATCH_SIZE.value,
learning_rate=_LEARNING_RATE.value,
)
elif task == constants.CAUSAL_LANGUAGE_MODELING_LORA:
causal_language_modeling_lora.finetune_causal_language_modeling(
pretrained_model_id=pretrained_model_id,
dataset_name=_DATASET_NAME.value,
output_dir=_OUTPUT_DIR.value,
precision_mode=_PRECISION_MODE.value,
lora_rank=_LORA_RANK.value,
lora_alpha=_LORA_ALPHA.value,
lora_dropout=_LORA_DROPOUT.value,
target_modules=_TARGET_MODULES.value,
warmup_steps=_WARMUP_STEPS.value,
max_steps=_MAX_STEPS.value,
learning_rate=_LEARNING_RATE.value,
local_pretrained_model_id=local_pretrained_model_id,
)
elif task == constants.INSTRUCT_LORA:
instruct_lora.finetune_instruct(
pretrained_model_id=pretrained_model_id,
dataset_name=_DATASET_NAME.value,
output_dir=_OUTPUT_DIR.value,
lora_rank=_LORA_RANK.value,
lora_alpha=_LORA_ALPHA.value,
lora_dropout=_LORA_DROPOUT.value,
target_modules=_TARGET_MODULES.value,
warmup_ratio=_WARMUP_RATIO.value,
max_steps=_MAX_STEPS.value,
max_seq_length=_MAX_SEQ_LENGTH.value,
learning_rate=_LEARNING_RATE.value,
)
else:
raise ValueError('The task {} is not supported.'.format(task))
if __name__ == '__main__':
app.run(main)
@@ -1,133 +0,0 @@
"""Sequence classification with LoRA models."""
# pylint: disable=g-importing-member
from datasets import load_dataset
import evaluate
from peft import get_peft_model
from peft import LoraConfig
import torch
from torch.optim import AdamW
from torch.utils.data import DataLoader
from tqdm import tqdm
from transformers import AutoModelForSequenceClassification
from transformers import AutoTokenizer
from transformers import get_linear_schedule_with_warmup
def finetune_sequence_classification(
pretrained_model_id: str,
dataset_name: str,
output_dir: str,
lora_rank: int = 8,
lora_alpha: int = 16,
lora_dropout: float = 0.1,
num_epochs: int = 20,
batch_size: int = 32,
learning_rate: float = 3e-4,
) -> None:
"""Finetunes sequence classification."""
task = "mrpc"
device = "cuda"
peft_config = LoraConfig(
task_type="SEQ_CLS",
inference_mode=False,
r=lora_rank,
lora_alpha=lora_alpha,
lora_dropout=lora_dropout,
)
if any(k in pretrained_model_id for k in ("gpt", "opt", "bloom")):
padding_side = "left"
else:
padding_side = "right"
tokenizer = AutoTokenizer.from_pretrained(
pretrained_model_id, padding_side=padding_side
)
if getattr(tokenizer, "pad_token_id") is None:
tokenizer.pad_token_id = tokenizer.eos_token_id
datasets = load_dataset(dataset_name, task)
metric = evaluate.load(dataset_name, task)
def tokenize_function(examples):
# max_length=None => use the model max length (it's actually the default)
outputs = tokenizer(
examples["sentence1"],
examples["sentence2"],
truncation=True,
max_length=None,
)
return outputs
tokenized_datasets = datasets.map(
tokenize_function,
batched=True,
remove_columns=["idx", "sentence1", "sentence2"],
)
# We also rename the 'label' column to 'labels' which is the expected name for
# labels by the models of the transformers library.
tokenized_datasets = tokenized_datasets.rename_column("label", "labels")
def collate_fn(examples):
return tokenizer.pad(examples, padding="longest", return_tensors="pt")
# Instantiate dataloaders.
train_dataloader = DataLoader(
tokenized_datasets["train"],
shuffle=True,
collate_fn=collate_fn,
batch_size=batch_size,
)
eval_dataloader = DataLoader(
tokenized_datasets["validation"],
shuffle=False,
collate_fn=collate_fn,
batch_size=batch_size,
)
model = AutoModelForSequenceClassification.from_pretrained(
pretrained_model_id, return_dict=True
)
model = get_peft_model(model, peft_config)
model.print_trainable_parameters()
optimizer = AdamW(params=model.parameters(), lr=learning_rate)
# Instantiate scheduler
lr_scheduler = get_linear_schedule_with_warmup(
optimizer=optimizer,
num_warmup_steps=0.06 * (len(train_dataloader) * num_epochs),
num_training_steps=(len(train_dataloader) * num_epochs),
)
model.to(device)
for epoch in range(num_epochs):
model.train()
for _, batch in enumerate(tqdm(train_dataloader)):
batch.to(device)
outputs = model(**batch)
loss = outputs.loss
loss.backward()
optimizer.step()
lr_scheduler.step()
optimizer.zero_grad()
model.eval()
for _, batch in enumerate(tqdm(eval_dataloader)):
batch.to(device)
with torch.no_grad():
outputs = model(**batch)
predictions = outputs.logits.argmax(dim=-1)
references = batch["labels"]
metric.add_batch(
predictions=predictions,
references=references,
)
eval_metric = metric.compute()
print(f"epoch {epoch}:", eval_metric)
model.save_pretrained(output_dir)
@@ -1,6 +0,0 @@
#!/bin/bash
# Setup accelerate config before running trainer.
python -c "from accelerate.utils import write_basic_config; write_basic_config(mixed_precision='fp16')"
accelerate launch "$@"
@@ -0,0 +1,16 @@
# Dockerfile for axolotl training.
#
# To build:
# docker build -f model_oss/peft/train/axolotol/dockerfile/train.Dockerfile . -t ${YOUR_IMAGE_TAG}
#
# To push to gcr:
# docker tag ${YOUR_IMAGE_TAG} gcr.io/${YOUR_PROJECT}/${YOUR_IMAGE_TAG}
# docker push gcr.io/${YOUR_PROJECT}/${YOUR_IMAGE_TAG}
FROM winglian/axolotl:main-latest
RUN mkdir -p ./vertex_vision_model_garden/
COPY model_oss/peft/train/axolotl/*.py ./vertex_vision_model_garden/
ENTRYPOINT ["python3", "./vertex_vision_model_garden/train_entrypoint.py"]
@@ -0,0 +1,20 @@
#!/bin/bash
# Run copybara first:
# cloud/ml/applications/vision/model_garden/copybara/run_copybara_local.sh
# Run docker build:
# cloud/ml/applications/vision/model_garden/model_oss/peft/train/axolotl/scripts/build_train_docker.sh
set -x
COPYBARA_DIR="/tmp/train_docker/"
pushd "${COPYBARA_DIR}"
PROJECT="cloud-nas-260507"
IMAGE_TAG="gcr.io/${PROJECT}/axolotl-train:${USER}-test"
docker build -f model_oss/peft/train/axolotl/dockerfile/train.Dockerfile . -t "${IMAGE_TAG}"
docker push "${IMAGE_TAG}"
popd
@@ -0,0 +1,88 @@
"""Entrypoint for axolotl train docker."""
import argparse
import json
import os
import subprocess
def _get_multi_node_flags(cluster_spec: str) -> list[str]:
"""Returns the multi-node flags."""
print(f'CLUSTER_SPEC: {cluster_spec}')
cluster_data = json.loads(cluster_spec)
# Get primary node info
primary_node = cluster_data['cluster']['workerpool0'][0]
print(f'primary node: {primary_node}')
primary_node_addr, primary_node_port = primary_node.split(':')
print(f'primary node address: {primary_node_addr}')
print(f'primary node port: {primary_node_port}')
# Determine node rank of this machine
workerpool = cluster_data['task']['type']
if workerpool == 'workerpool0':
node_rank = 0
else:
node_rank = cluster_data['task']['index'] + 1
print(f'node rank: {node_rank}')
# Calculate total nodes
num_worker_nodes = len(cluster_data['cluster']['workerpool1'])
num_nodes = num_worker_nodes + 1 # Add 1 for the primary node
print(f'num nodes: {num_nodes}')
return [
f'--machine_rank={node_rank}',
f'--num_machines={num_nodes}',
f'--main_process_ip={primary_node_addr}',
f'--main_process_port={primary_node_port}',
'--max_restarts=0',
'--monitor_interval=120',
'--dynamo_backend=no',
]
def main() -> None:
parser = argparse.ArgumentParser()
parser.add_argument('--config_file')
parser.add_argument('--huggingface_access_token')
args, unknown = parser.parse_known_args()
accelerate_flags = []
if args.config_file:
accelerate_flags.append(f'--config_file={args.config_file}')
if cluster_spec := os.getenv('CLUSTER_SPEC', default=None):
print('========== Launch on cloud multi nodes ==========')
accelerate_flags.extend(_get_multi_node_flags(cluster_spec))
cmd = (
[
'accelerate',
'launch',
]
+ accelerate_flags
+ [
'-m',
'axolotl.cli.train',
]
+ unknown
)
print(f'{cmd=}', flush=True)
env = os.environ.copy()
if args.huggingface_access_token:
env['HF_TOKEN'] = args.huggingface_access_token
subprocess.run(
cmd,
check=True,
env=env,
)
if __name__ == '__main__':
main()
@@ -0,0 +1,95 @@
"""Class that bundles docker related flags."""
import getpass
import os
import pwd
class CommandBuilder:
"""Base class for building commands."""
def __init__(self):
self._defaults = []
self._env_vars = {}
def add_env_var(self, var: str, val: str) -> None:
"""Add environment variable to the command.
Args:
var: environment variable name.
val: environment variable value.
"""
self._env_vars[var] = val
def add_mount_map(self, host_path, docker_path):
pass
class DockerCommandBuilder(CommandBuilder):
"""Bundle docker related flags."""
def __init__(self, docker_uri: str, shm_size: str = '128gb'):
super().__init__()
self._docker_uri = [docker_uri]
self.privilege_mode = []
self.entrypoint = []
self._defaults = [
'docker',
'run',
'--gpus=all',
'--net=host',
'--rm',
f'--shm-size={shm_size}',
]
self._mount_maps = []
user = getpass.getuser()
# username ends with `_google_com` is managed by ldap and does not have a
# corresponding entry in /etc/passwd or /etc/group file. We cannot enable
# non-root docker user with below method.
if not user.endswith('_google_com'):
uid = os.getuid()
gid = pwd.getpwuid(uid).pw_gid
self._defaults += [
f'--user={uid}:{gid}',
'--volume=/etc/group:/etc/group:ro',
'--volume=/etc/passwd:/etc/passwd:ro',
]
def add_mount_map(self, host_path, docker_path):
self._mount_maps.append(f'--volume={host_path}:{docker_path}')
def add_privilege_mode(self):
self.privilege_mode = ['--privileged']
def add_entrypoint(self, entrypoint: list[str]):
self.entrypoint = entrypoint
def build_cmd(self) -> str:
return (
self._defaults
+ [f'--env={var}={val}' for var, val in self._env_vars.items()]
+ self._mount_maps
+ self.privilege_mode
+ self._docker_uri
+ self.entrypoint
)
class PythonCommandBuilder(CommandBuilder):
"""Bundle Python test command related flags."""
def __init__(self):
super().__init__()
self._defaults = [
'python3',
'./vertex_vision_model_garden_peft/train/vmg/train_entrypoint.py',
]
def build_cmd(self) -> str:
os.environ.update(self._env_vars)
return self._defaults
def add_entrypoint(self, entrypoint: list[str]):
self._defaults = entrypoint
@@ -0,0 +1,471 @@
"""Test util class."""
import copy
import dataclasses
import datetime
import inspect
import os
import signal
import subprocess
import sys
from absl import flags
from absl import logging
from absl.testing import parameterized
import command_builder
import immutabledict
import torch
_DOCKER_URI = flags.DEFINE_string('docker_uri', None, 'docker image uri')
_DRY_RUN = flags.DEFINE_bool('dry_run', False, 'dry-run the commands')
_LOCAL_INPUT_DIR = flags.DEFINE_string(
'local_input_dir',
os.path.expanduser('~/test_input'),
'local directory for storing input data.',
)
_LOCAL_OUTPUT_DIR = flags.DEFINE_string(
'local_output_dir',
'/tmp',
'local directory for storing test output.',
)
_GCS_INPUT_DIR = flags.DEFINE_string(
'gcs_input_dir',
'gs://vmg-tuning-docker-test',
'GCS directory that stores model checkpoint, dataset and etc.',
)
_GCS_OUTPUT_DIR = flags.DEFINE_string(
'gcs_output_dir',
'gs://vmg-tuning-docker-test/output',
'GCS directory that stores test output.',
)
_GCS_TESTDATA_DIR = 'peft-train-image-test'
_THROUGHPUT_TEST_EXCEPTIONS = immutabledict.immutabledict({
('bm_deepspeed_zero3_8gpu_gemma-2-9b-it_4bit.txt', '12.0'): float('inf'),
('bm_fsdp_8gpu_llama3.1-70b-hf_4bit.txt', '20.0'): float('inf'),
('bm_deepspeed_zero2_8gpu_gemma-2-2b-it_bfloat16.txt', '12.0'): 20.0,
('bm_deepspeed_zero3_8gpu_gemma-2-2b-it_4bit.txt', '4.0'): 20.0,
('bm_deepspeed_zero3_8gpu_gemma-2-27b-it_4bit.txt', '4.0'): 20.0,
})
@dataclasses.dataclass
class BenchmarkStats:
"""Class to store the benchmark result.
Attributes:
peak_mem: peak memory in GB.
throughput: throughput in tokens/sec.
"""
peak_mem: float
throughput: float
class TestBase(parameterized.TestCase):
"""Test base class that defines how to run commands."""
def setUp(self):
super().setUp()
# Create a copy of the environment variables
self.old_env_var = copy.deepcopy(os.environ)
if _DOCKER_URI.value:
self.command_builder = command_builder.DockerCommandBuilder(
_DOCKER_URI.value
)
else:
self.command_builder = command_builder.PythonCommandBuilder()
self.command_builder.add_mount_map(
os.path.expanduser('~'), os.path.expanduser('~')
)
self.command_builder.add_mount_map(
self.local_input_dir(), self.local_input_dir()
)
self.task_cmd_builder = None
def tearDown(self):
super().tearDown()
# Restore the original environment variables
os.environ.clear()
os.environ.update(self.old_env_var)
def cmd(self):
return self.command_builder.build_cmd() + self.task_cmd_builder.build_cmd()
def run_cmd(self) -> int:
return run_cmd(self.cmd(), output_file=None)
def gcs_output_dir(self):
return _GCS_OUTPUT_DIR.value
def local_input_dir(self):
"""Returns local input dir in host/docker."""
return _LOCAL_INPUT_DIR.value
def local_output_dir(self):
"""Returns local output dir in host/docker."""
return _LOCAL_OUTPUT_DIR.value
def get_testcase_name(self):
"""Returns the function name at the calling site."""
# https://docs.python.org/3/library/inspect.html#inspect.FrameInfo
cur_frame = inspect.currentframe()
# https://stackoverflow.com/a/17366561
return cur_frame.f_back.f_code.co_name
def get_timestamp():
return datetime.datetime.now(datetime.timezone.utc).strftime(
'%Y%m%d_%H%M%S%Z'
)
def download_from_gcs(gcs_uri: str, local_dir: str):
if not os.path.exists(local_dir):
os.mkdir(local_dir)
subprocess.check_output([
'gcloud',
'storage',
'cp',
'-r',
gcs_uri,
local_dir,
])
def get_test_data_path(name: str, download: bool = True) -> str:
"""Gets test data path.
Args:
name: name of the test data
download: if True, then download data from GCS and returns its local path.
Returns:
test data path.
"""
if not download:
return os.path.join(_GCS_INPUT_DIR.value, name)
local_data = os.path.join(_LOCAL_INPUT_DIR.value, name)
if not os.path.exists(local_data):
# If `name` is a file in sub-folders, then create the sub-folders under
# `_LOCAL_INPUT_DIR`.
local_data_dir = os.path.dirname(local_data)
if not os.path.exists(local_data_dir):
os.makedirs(local_data_dir)
download_from_gcs(os.path.join(_GCS_INPUT_DIR.value, name), local_data_dir)
return local_data
def run_cmd(cmd: list[str], output_file: str = None) -> int:
"""Runs the command and returns the return code.
Args:
cmd: The command to run.
output_file: The file to write the output to.
Returns:
The return code of the command.
"""
logging.info('running command: \n%s', ' \\\n'.join(cmd))
if _DRY_RUN.value:
return 0
stdout = sys.stdout if output_file is None else open(output_file, 'w')
p = subprocess.Popen(cmd, stdout=stdout, stderr=sys.stderr)
try:
unused_output, unused_error = p.communicate()
return_code = p.returncode
except KeyboardInterrupt:
p.send_signal(signal.SIGINT)
return_code = 0
finally:
if output_file is not None:
stdout.close()
return return_code
def get_pretrained_model_name_or_path(model_id: str) -> str:
# If `model_id` contains `/`, it is assumed to be HF model or model from GCS.
if '/' in model_id:
return model_id
return get_test_data_path(model_id, download=True)
def is_gpu_h100():
"""Checks if the GPU is H100."""
return 'H100' in torch.cuda.get_device_name()
def is_gpu_a100():
"""Checks if the GPU is A100."""
return 'A100' in torch.cuda.get_device_name()
def _get_formatted_string(max_seq_length: int) -> str:
"""Returns the formatted string for max_seq_length.
Args:
max_seq_length: max sequence length to get the formatted string.
Returns:
formatted string for max_seq_length.
"""
return f'{max_seq_length/1024.0:.1f}'
def get_benchmark_results(
benchmark_file_path: str, max_seq_length: int
) -> BenchmarkStats:
"""Gets benchmark results from the benchmark file.
Args:
benchmark_file_path: path to the benchmark file.
max_seq_length: max sequence length to get the benchmark results.
Returns:
peak_mem: peak memory in GB.
throughput: throughput in tokens/sec.
"""
formatted_max_seq_length = _get_formatted_string(max_seq_length)
peak_mem, throughput = None, None
with open(benchmark_file_path, 'r') as f:
for line in f:
if line.startswith(formatted_max_seq_length):
metrics = line.split('|')
try:
peak_mem = float(metrics[1].strip())
except ValueError:
pass
try:
throughput = float(metrics[2].strip())
except ValueError:
pass
break
else:
logging.error(
'No metrics found for max_seq_length %s in %s',
formatted_max_seq_length,
benchmark_file_path,
)
return BenchmarkStats(peak_mem, throughput)
def print_benchmark_file(file_path: str) -> None:
"""Prints the contents of the file.
Args:
file_path: path to the file.
"""
with open(file_path, 'r') as f:
for line in f:
logging.info(line.strip())
def print_benchmark_results(
benchmark_file_path: str, benchmark_type: str
) -> None:
"""Prints the benchmark results.
Args:
benchmark_file_path: path to the benchmark file.
benchmark_type: type of the benchmark.
"""
benchmark_filename = os.path.basename(benchmark_file_path)
logging.info('--------------------------------------------------------------')
logging.info('%s benchmark for %s', benchmark_type, benchmark_filename)
logging.info('--------------------------------------------------------------')
print_benchmark_file(benchmark_file_path)
def _calculate_percent_change(
actual_value: float, expected_value: float
) -> float:
"""Calculates the percent change between the actual and expected values.
Args:
actual_value: actual value to compare.
expected_value: expected value to compare.
Returns:
percent change between the actual and expected values.
"""
return ((actual_value - expected_value) / expected_value) * 100.0
def compare_benchmark_results(
expected_benchmark_file_path: str,
actual_benchmark_file_path: str,
allowed_threshold: float,
max_seq_length: int,
) -> bool:
"""Compares if the benchmark results are the similar.
Args:
expected_benchmark_file_path: path to the expected benchmark file.
actual_benchmark_file_path: path to the actual benchmark file.
allowed_threshold: allowed percent range of the benchmark results.
max_seq_length: max sequence length to get the benchmark results.
Returns:
True if the benchmark results are the similar, False otherwise.
"""
benchmark_filename = os.path.basename(expected_benchmark_file_path)
expected_results = get_benchmark_results(
expected_benchmark_file_path, max_seq_length
)
expected_peak_mem, expected_throughput = (
expected_results.peak_mem,
expected_results.throughput,
)
actual_results = get_benchmark_results(
actual_benchmark_file_path, max_seq_length
)
actual_peak_mem, actual_throughput = (
actual_results.peak_mem,
actual_results.throughput,
)
formatted_max_seq_length = _get_formatted_string(max_seq_length)
# Case 1: both peak mem and throughput are None(ideally due to OOM)
if expected_peak_mem is None and actual_peak_mem is None:
logging.info(
'Both peak mem and throughput are None for max_seq_length %d.',
max_seq_length,
)
return True
check_oom_exception = _THROUGHPUT_TEST_EXCEPTIONS.get(
(benchmark_filename, formatted_max_seq_length), 0.0
) == float('inf')
# Case 2: When something strated to fail recently, or something which failed
# before but is working now.
if expected_peak_mem is None and actual_peak_mem is not None:
if check_oom_exception:
return True
logging.error(
'One of the failing benchmarks in %s is passing now for max_seq_length'
' %d. The expected peak mem and throughput are None, but the actual'
' peak mem is %f and actual throughput is %f',
benchmark_filename,
max_seq_length,
actual_peak_mem,
actual_throughput,
)
return False
if actual_peak_mem is None and expected_peak_mem is not None:
if check_oom_exception:
return True
logging.error(
'One of the passing benchmarks in %s is failing now for max_seq_length'
' %d. The actual peak mem and throughput are None, but the expected'
' peak mem is %f and expected throughput is %f',
benchmark_filename,
max_seq_length,
expected_peak_mem,
expected_throughput,
)
return False
# Case 3: When both actual peak mem and throughput lies within the range
# of their respective expected values.
mem_percent_change = _calculate_percent_change(
actual_peak_mem, expected_peak_mem
)
throughput_percent_change = _calculate_percent_change(
actual_throughput, expected_throughput
)
allowed_threshold = _THROUGHPUT_TEST_EXCEPTIONS.get(
(benchmark_filename, formatted_max_seq_length), allowed_threshold
)
if abs(mem_percent_change) > allowed_threshold:
logging.error(
'The peak memory is changing by more than %f%% for max_seq_length %d.'
' Expected: %f, Actual: %f',
allowed_threshold,
max_seq_length,
expected_peak_mem,
actual_peak_mem,
)
return False
if abs(throughput_percent_change) > allowed_threshold:
logging.error(
'The throughput is changing by more than %f%% for max_seq_length %d.'
' Expected throughput: %f, Actual throughput: %f',
allowed_threshold,
max_seq_length,
expected_throughput,
actual_throughput,
)
return False
return True
def check_benchmark_results(
actual_benchmark_file_path: str,
model_family: str,
allowed_threshold: float,
max_seq_length: int,
) -> bool:
"""Checks the benchmark result between the actual and expected benchmark files.
Args:
actual_benchmark_file_path: path to the actual benchmark file.
model_family: family of the model.
allowed_threshold: allowed range of the benchmark results in percent.
max_seq_length: max sequence length to get the benchmark results.
Returns:
True if the benchmark results are the similar, False otherwise.
"""
benchmark_filename = os.path.basename(actual_benchmark_file_path)
get_test_data_path(_GCS_TESTDATA_DIR)
expected_benchmark_file_path = os.path.join(
_LOCAL_INPUT_DIR.value,
_GCS_TESTDATA_DIR,
model_family,
benchmark_filename,
)
print_benchmark_results(expected_benchmark_file_path, 'Expected')
print_benchmark_results(actual_benchmark_file_path, 'Actual')
return compare_benchmark_results(
expected_benchmark_file_path,
actual_benchmark_file_path,
allowed_threshold,
max_seq_length,
)
def list_gcs_directories(bucket: str, directory: str) -> list[str]:
"""Lists GCS files."""
output = subprocess.check_output([
'gcloud',
'storage',
'ls',
f'gs://{bucket}/{directory}',
])
return output.decode('utf-8').splitlines()
def delete_gcs_object(gcs_directory: str):
"""Deletes GCS object."""
subprocess.check_output([
'gcloud',
'storage',
'rm',
'-r',
f'{gcs_directory}',
])
@@ -0,0 +1,79 @@
"""Get cluster info from environment variables."""
import dataclasses
import json
import os
from absl import logging
@dataclasses.dataclass
class ClusterInfo:
"""Contains information about the cluster.
Attributes:
primary_node_addr: The address of the primary node.
primary_node_port: The port of the primary node.
node_rank: The rank of the node.
num_nodes: The number of nodes in the cluster.
"""
primary_node_addr: str | None = None
primary_node_port: str | None = None
node_rank: int = 0
num_nodes: int = 1
# Allows unpacking operation like
# primary_node_addr, primary_node_port, _, _ = ClusterInfo()
# See https://stackoverflow.com/a/70753113
def __iter__(self):
return iter(dataclasses.astuple(self))
def get_cluster_spec() -> ClusterInfo:
"""Parses CLUSTER_SPEC environment variable and returns the cluster info.
Returns:
A ClusterInfo object.
"""
cluster_spec = os.getenv('CLUSTER_SPEC', None)
# If CLUSTER_SPEC is not set, use individual vars to construct cluster info.
if not cluster_spec:
cluster_info = ClusterInfo(
primary_node_addr=os.getenv('MASTER_ADDR', None),
primary_node_port=os.getenv('MASTER_PORT', None),
node_rank=int(os.getenv('RANK', '0')),
num_nodes=int(os.getenv('NNODES', '1')),
)
return cluster_info
cluster_data = json.loads(cluster_spec)
# Get primary node info
primary_node = cluster_data['cluster']['workerpool0'][0]
logging.info('primary node: %s', primary_node)
primary_node_addr, primary_node_port = primary_node.split(':')
logging.info('primary node address: %s', primary_node_addr)
logging.info('primary node port: %s', primary_node_port)
# Determine node rank of this machine
workerpool = cluster_data['task']['type']
if workerpool == 'workerpool0':
node_rank = 0
elif workerpool == 'workerpool1':
# Add 1 for the primary node, since `index` is the index of workerpool1.
node_rank = cluster_data['task']['index'] + 1
else:
raise ValueError(
'Only workerpool0 and workerpool1 are supported. Unknown workerpool:'
f' {workerpool}'
)
logging.info('node rank: %s', node_rank)
# Calculate total nodes.
num_nodes = 1 # For the primary node.
if 'workerpool1' in cluster_data['cluster']:
num_nodes += len(cluster_data['cluster']['workerpool1'])
logging.info('num nodes: %s', num_nodes)
return ClusterInfo(primary_node_addr, primary_node_port, node_rank, num_nodes)
@@ -0,0 +1,24 @@
"""Utility functions."""
import logging
import subprocess
import sys
import time
def run_cmd(cmd: list[str]) -> float:
"""Runs the command and logs the output.
Args:
cmd: The command to run.
Returns:
The time it took to run the command.
"""
cmd_str = ' \\\n'.join(cmd)
logging.info('launching cmd: \n%s', cmd_str)
start_time = time.time()
subprocess.run(cmd, stdout=sys.stdout, stderr=sys.stdout, check=True)
elapsed_time = round(time.time() - start_time, 2)
logging.info('Command %s finished in %0.2f seconds.', cmd_str, elapsed_time)
return elapsed_time
@@ -0,0 +1,197 @@
"""Calculate dataset statistics like token, example and character counts."""
from collections.abc import Mapping, Sequence
import dataclasses
import json
from typing import Any
import datasets
import numpy as np
import transformers
from util import dataset_validation_util
_MAX_NUM_DATASET_SAMPLES = 6
@dataclasses.dataclass
class SupervisedTuningDatasetBucket:
"""Represents a histogram bucket for tuning dataset distribution stats."""
count: float = 0
left: float = 0
right: float = 0
@dataclasses.dataclass
class SupervisedTuningDatasetDistribution:
"""Represents a histogram with summary statistics for tuning dataset distribution stats."""
sum: int = 0
billable_sum: int = 0
min: float = 0
max: float = 0
mean: float = 0
median: float = 0
p5: float = 0
p95: float = 0
buckets: list[SupervisedTuningDatasetBucket] = dataclasses.field(
default_factory=list
)
# Represents detailed tuning dataset statistics.
@dataclasses.dataclass
class SupervisedTuningDataStats:
"""Represents detailed tuning dataset stats."""
tuning_dataset_example_count: int = 0
total_tuning_character_count: int = 0
total_billable_token_count: int = 0
tuning_step_count: int = 0
# Represents a histogram and some summary statistics of the number of input
# tokens across examples.
user_input_token_distribution: SupervisedTuningDatasetDistribution | None = (
None
)
# Represents a histogram and some summary statistics for the number of output
# tokens across examples.
user_output_token_distribution: SupervisedTuningDatasetDistribution | None = (
None
)
# Represents the number of "messages" (a single-turn conversation will have a
# single message) across examples.
user_message_per_example_distribution: (
SupervisedTuningDatasetDistribution | None
) = None
user_dataset_examples: list[str] = dataclasses.field(default_factory=list)
def get_dataset_stats(
*,
raw: Any,
templated: Any,
template: str,
tokenizer: transformers.PreTrainedTokenizer,
column: str,
effective_batch_size: int,
) -> Mapping[str, Any]:
"""Calculates dataset statistics for managed fine-tuning, e.g., total number of tokens."""
tokenized_dataset = templated.map(lambda x: tokenizer(x[column]))
inputs = tokenized_dataset["input_ids"]
tuning_dataset_example_count = int(len(inputs))
total_billable_token_count = int(np.sum([len(ex) for ex in inputs]))
total_tuning_character_count = int(
np.sum([len(ex[column]) for ex in templated])
)
tuning_step_count = (
tuning_dataset_example_count + effective_batch_size - 1
) // effective_batch_size
# Assume that data is represented as ChatCompletions or Vertex Text-Bison
# formats to extract per-example input/output tokens.
user_inputs = []
user_outputs = []
user_input_messages_counts = []
for ex in raw:
if "messages" in ex:
messages = ex["messages"]
if messages:
# For ChatCompletions assume the last turn (i.e. the instruction
# response) is the expected output.
user_inputs.append({**ex, "messages": messages[:-1]})
user_outputs.append({**ex, "messages": messages[-1:]})
# Exclude everything but the last message for the number of input
# messages.
user_input_messages_counts.append(len(messages[:-1]))
elif "input_text" in ex:
# For Vertex Text-Bison, the `output_text` field is the expected output.
user_inputs.append({**ex, "output_text": ""})
user_outputs.append(
{**ex, "input_text": ex["output_text"], "output_text": ""}
)
# Vertex Text-Bison goes from input -> output; i.e. there is only a single
# input "message".
user_input_messages_counts.append(1)
def calc_histogram(
counts: Sequence[int],
) -> SupervisedTuningDatasetDistribution:
mean = np.mean(counts)
median = np.median(counts).item()
max_count = np.max(counts).item()
min_count = np.min(counts).item()
count_sum = np.sum(counts).item()
p5 = np.percentile(counts, 0.05).item()
p95 = np.percentile(counts, 0.95).item()
hist, bin_edges = np.histogram(counts, bins=10)
return SupervisedTuningDatasetDistribution(
sum=count_sum,
billable_sum=count_sum,
min=min_count,
max=max_count,
mean=mean,
median=median,
p5=p5,
p95=p95,
buckets=[
SupervisedTuningDatasetBucket(
count=hist[i].item(),
left=bin_edges[i].item(),
right=bin_edges[i + 1].item(),
)
for i in range(len(hist))
],
)
# Tokenize input and output messages separately to generate separate summary
# statistics about them.
user_input_token_distribution = None
if user_inputs:
user_input_dataset = dataset_validation_util.format_dataset(
datasets.Dataset.from_list(user_inputs), column, template, tokenizer
)
user_input_tokenized_dataset = user_input_dataset.map(
lambda x: tokenizer(x[column])
)
user_input_tokens = user_input_tokenized_dataset["input_ids"]
user_input_token_counts = np.array([len(ex) for ex in user_input_tokens])
user_input_token_distribution = calc_histogram(user_input_token_counts)
user_output_token_distribution = None
if user_outputs:
user_output_dataset = dataset_validation_util.format_dataset(
datasets.Dataset.from_list(user_outputs), column, template, tokenizer
)
user_output_tokenized_dataset = user_output_dataset.map(
lambda x: tokenizer(x[column])
)
user_output_tokens = user_output_tokenized_dataset["input_ids"]
user_output_token_counts = np.array([len(ex) for ex in user_output_tokens])
user_output_token_distribution = calc_histogram(user_output_token_counts)
user_messages_per_example_distribution = None
if user_input_messages_counts:
user_input_messages_counts = np.array(user_input_messages_counts)
user_messages_per_example_distribution = calc_histogram(
user_input_messages_counts
)
user_dataset_examples = [
json.dumps(ex)
for ex in raw.shuffle().select(
range(min(len(raw), _MAX_NUM_DATASET_SAMPLES))
)
]
dataset_stats = SupervisedTuningDataStats(
tuning_dataset_example_count=tuning_dataset_example_count,
total_tuning_character_count=total_tuning_character_count,
total_billable_token_count=total_billable_token_count,
tuning_step_count=tuning_step_count,
user_input_token_distribution=user_input_token_distribution,
user_output_token_distribution=user_output_token_distribution,
user_message_per_example_distribution=user_messages_per_example_distribution,
user_dataset_examples=user_dataset_examples,
)
return dataclasses.asdict(dataset_stats)
@@ -0,0 +1,140 @@
"""Util functions for reporting device (GPU, CPU) stats."""
import dataclasses
import psutil
import pynvml
import torch
@dataclasses.dataclass
class GpuStats:
"""Holds information about GPU usage stats.
For memory related, see
https://pytorch.org/docs/stable/notes/cuda.html#cuda-memory-management
"""
# device id
device_id: int
# memory reserved.
reserved: float
# memory occupied.
occupied: float
# memory reserved, but not used.
unused: float
# nvidia-smi usually reports more memory usages than pytorch (for driver,
# kernel and etc). `smi_diff` tracks this difference.
smi_diff: float
# Gpu utilization.
util: float
# Allows unpacking operation like
# device_id, reserved, occupied, unused, smi_diff, util = GpuStats(...)
# See https://stackoverflow.com/a/70753113
def __iter__(self):
return iter(dataclasses.astuple(self))
def gpu_stats() -> GpuStats:
"""Reports GPU memory usage and utilization."""
# See https://pytorch.org/docs/stable/notes/cuda.html#memory-management
bytes_per_gb = 1024.0**3
device = torch.cuda.current_device()
occupied = torch.cuda.memory_allocated(device) / bytes_per_gb
reserved = torch.cuda.memory_reserved(device) / bytes_per_gb
unused = reserved - occupied
def smi_mem(device):
try:
pynvml.nvmlInit()
handle = pynvml.nvmlDeviceGetHandleByIndex(device)
info = pynvml.nvmlDeviceGetMemoryInfo(handle)
return info.used / bytes_per_gb
except pynvml.NVMLError:
return 0.0
mem_used_smi = smi_mem(device)
smi_diff = mem_used_smi - reserved
util = torch.cuda.utilization(device)
return GpuStats(device, reserved, occupied, unused, smi_diff, util)
def gpu_stats_str(stats: GpuStats | None = None) -> str:
if stats is None:
stats = gpu_stats()
device, reserved, occupied, unused, smi_diff, util = stats
return (
f"GPU ({device=}) memory: {reserved:.2f}({occupied=:.2f}, {unused=:.2f}),"
f" {smi_diff=:.2f} GB. Utilization: {util:.2f}%"
)
@dataclasses.dataclass
class CpuStats:
"""Holds information about CPU usage stats."""
# Total CPU virtual memory i.e. virtual memory allocated + unallocated.
total_virtual_mem: float
# CPU virtual memory available for use.
unallocated_virtual_mem: float
# CPU virtual memory already used.
allocated_virtual_mem: float
# Total CPU swap memory i.e. swap memory allocated + unallocated.
total_swap_mem: float
# CPU swap memory available for use.
unallocated_swap_mem: float
# CPU swap memory already used.
allocated_swap_mem: float
# CPU utilization percentage.
utilization: float
def cpu_stats() -> CpuStats:
"""Reports CPU memory usage and utilization."""
# https://psutil.readthedocs.io/en/latest/#memory
gb = 1024.0**3
vmem = psutil.virtual_memory()
vmem_total = vmem.total / gb
vmem_available = vmem.available / gb
vmem_used = vmem_total - vmem_available
smem = psutil.swap_memory()
swap_total = smem.total / gb
swap_free = smem.free / gb
swap_used = smem.used / gb
# https://psutil.readthedocs.io/en/latest/#psutil.cpu_percent
cpu_util = psutil.cpu_percent(interval=1e-6)
return CpuStats(
total_virtual_mem=vmem_total,
unallocated_virtual_mem=vmem_available,
allocated_virtual_mem=vmem_used,
total_swap_mem=swap_total,
unallocated_swap_mem=swap_free,
allocated_swap_mem=swap_used,
utilization=cpu_util,
)
def cpu_stats_str(stats: CpuStats | None = None) -> str:
"""Returns a string representation of the CPU stats."""
if stats is None:
stats = cpu_stats()
total, occupied, unused = (
stats.total_virtual_mem,
stats.allocated_virtual_mem,
stats.unallocated_virtual_mem,
)
virtual_mem = (
f"CPU virtual memory: {total:.2f}({occupied=:.2f}, {unused=:.2f}) GB"
)
total, occupied, unused = (
stats.total_swap_mem,
stats.allocated_swap_mem,
stats.unallocated_swap_mem,
)
swap_mem = f"CPU swap memory: {total:.2f}({occupied=:.2f}, {unused=:.2f}) GB"
percent = stats.utilization
return f"{virtual_mem} {swap_mem} CPU Utilization: {percent:.2f}%"
@@ -0,0 +1,126 @@
"""Different trainer callbacks for PEFT Trainer."""
from collections.abc import MutableMapping
import math
import time
from absl import logging
import accelerate
from transformers import TrainingArguments
from transformers.trainer_callback import TrainerCallback
from transformers.trainer_callback import TrainerControl
from transformers.trainer_callback import TrainerState
from util import device_stats
class TrainerStatsCallback(TrainerCallback):
"""Trainer callback to report trainer stats."""
def __init__(self, max_seq_length, filename=None):
self._max_seq_length = max_seq_length
self._filename = filename
self._partial_state = accelerate.PartialState()
self._start_time = float('nan')
self._prev_time = float('nan')
self._peak_mem = 0.0
self._avg_throughput = 0.0
def on_log(
self,
args: TrainingArguments,
state: TrainerState,
control: TrainerControl,
logs: MutableMapping[str, float] | None = None,
**kwargs,
) -> None:
"""Calculates perplexity from train loss.
Args:
args: Arguments passed to the trainer.
state: State of the trainer.
control: Control of the trainer.
logs: A dict of logs from the training loop.
**kwargs: Additional keyword arguments, not used in this callback.
"""
del kwargs # Unused.
if self._partial_state.is_main_process:
train_loss = logs.get('loss') if logs is not None else None
if train_loss is not None:
perplexity = round(float(math.exp(train_loss)), 4)
logs['perplexity'] = perplexity
def on_step_end(
self,
args: TrainingArguments,
state: TrainerState,
control: TrainerControl,
**kwargs,
):
if self._partial_state.is_main_process:
if state.global_step == 1:
self._prev_time = time.time()
self._prev_num_token = state.num_input_tokens_seen
throughput = 0.0
else:
cur_time = time.time()
cur_num_token = state.num_input_tokens_seen
throughput = (cur_num_token - self._prev_num_token) / (
cur_time - self._prev_time
)
self._prev_time = cur_time
self._prev_num_token = cur_num_token
self._avg_throughput += (throughput - self._avg_throughput) / (
state.global_step - 1
)
gpu_stats = device_stats.gpu_stats()
self._peak_mem = max(
gpu_stats.reserved + gpu_stats.smi_diff, self._peak_mem
)
logging.info(
'on_step_end: Throughput: %.2f token/s. %s, %s',
throughput,
device_stats.gpu_stats_str(gpu_stats),
device_stats.cpu_stats_str(),
)
def on_train_begin(
self,
args: TrainingArguments,
state: TrainerState,
control: TrainerControl,
**kwargs,
):
if self._partial_state.is_main_process:
self._start_time = time.time()
logging.info(
'on_train_begin: %s, %s',
device_stats.gpu_stats_str(),
device_stats.cpu_stats_str(),
)
def on_train_end(
self,
args: TrainingArguments,
state: TrainerState,
control: TrainerControl,
**kwargs,
):
if self._partial_state.is_main_process:
train_time = time.time() - self._start_time
throughput = state.num_input_tokens_seen / train_time
logging.info(
'training time %.2f s, throughput (including overhead, e.g., ckpt'
' saving): %.2f token/s, peak_mem: %.2f GB',
train_time,
throughput,
self._peak_mem,
)
if self._filename:
with open(self._filename, 'a') as out_f:
out_f.write(
f'{self._max_seq_length/1024.0:.1f} | {self._peak_mem:.2f} |'
f' {self._avg_throughput:.2f}\n'
)
@@ -0,0 +1,18 @@
group:
- vertex
task: custom_loglikelihood
dataset_path: json
dataset_name: null
output_type: loglikelihood
training_split: null
validation_split: null
test_split: test
doc_to_text: "Request: {{prompt}}\nResponse:"
doc_to_target: " {{ground_truth}}"
metric_list:
- metric: perplexity
aggregation: perplexity
higher_is_better: false
- metric: acc
aggregation: mean
higher_is_better: true
@@ -0,0 +1,17 @@
compute_environment: LOCAL_MACHINE
debug: false
distributed_type: MULTI_GPU
downcast_bf16: 'no'
enable_cpu_affinity: false
gpu_ids: all
machine_rank: 0
main_training_function: main
mixed_precision: fp16
num_machines: 1
num_processes: 4
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
@@ -0,0 +1,17 @@
compute_environment: LOCAL_MACHINE
debug: false
distributed_type: MULTI_GPU
downcast_bf16: 'no'
enable_cpu_affinity: false
gpu_ids: all
machine_rank: 0
main_training_function: main
mixed_precision: fp16
num_machines: 1
num_processes: 8
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
@@ -0,0 +1,17 @@
compute_environment: LOCAL_MACHINE
debug: false
deepspeed_config:
deepspeed_config_file: /diffusers/examples/vertex_vision_model_garden_peft/zero2.json
zero3_init_flag: true
distributed_type: DEEPSPEED
downcast_bf16: 'no'
machine_rank: 0
main_training_function: main
num_machines: 1
num_processes: 4
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
@@ -0,0 +1,17 @@
compute_environment: LOCAL_MACHINE
debug: false
deepspeed_config:
deepspeed_config_file: /diffusers/examples/vertex_vision_model_garden_peft/zero2.json
zero3_init_flag: true
distributed_type: DEEPSPEED
downcast_bf16: 'no'
machine_rank: 0
main_training_function: main
num_machines: 1
num_processes: 8
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
@@ -0,0 +1,17 @@
compute_environment: LOCAL_MACHINE
debug: false
deepspeed_config:
deepspeed_config_file: /diffusers/examples/vertex_vision_model_garden_peft/zero3.json
zero3_init_flag: true
distributed_type: DEEPSPEED
downcast_bf16: 'no'
machine_rank: 0
main_training_function: main
num_machines: 1
num_processes: 4
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
@@ -0,0 +1,17 @@
compute_environment: LOCAL_MACHINE
debug: false
deepspeed_config:
deepspeed_config_file: /diffusers/examples/vertex_vision_model_garden_peft/zero3.json
zero3_init_flag: true
distributed_type: DEEPSPEED
downcast_bf16: 'no'
machine_rank: 0
main_training_function: main
num_machines: 1
num_processes: 8
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
@@ -0,0 +1,28 @@
compute_environment: LOCAL_MACHINE
debug: false
distributed_type: FSDP
downcast_bf16: 'no'
enable_cpu_affinity: false
fsdp_config:
fsdp_auto_wrap_policy: TRANSFORMER_BASED_WRAP
fsdp_transformer_layer_cls_to_wrap: Gemma2DecoderLayer
fsdp_backward_prefetch: NO_PREFETCH
fsdp_cpu_ram_efficient_loading: true
fsdp_forward_prefetch: false
fsdp_offload_params: true
fsdp_sharding_strategy: FULL_SHARD
fsdp_state_dict_type: SHARDED_STATE_DICT
fsdp_sync_module_states: true
fsdp_use_orig_params: false
fsdp_activation_checkpointing: false
main_training_function: main
mixed_precision: bf16
machine_rank: 0
num_machines: 1
num_processes: 8
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
@@ -0,0 +1,28 @@
compute_environment: LOCAL_MACHINE
debug: false
distributed_type: FSDP
downcast_bf16: 'no'
enable_cpu_affinity: false
fsdp_config:
fsdp_auto_wrap_policy: TRANSFORMER_BASED_WRAP
fsdp_transformer_layer_cls_to_wrap: LlamaDecoderLayer
fsdp_backward_prefetch: NO_PREFETCH
fsdp_cpu_ram_efficient_loading: true
fsdp_forward_prefetch: false
fsdp_offload_params: true
fsdp_sharding_strategy: FULL_SHARD
fsdp_state_dict_type: SHARDED_STATE_DICT
fsdp_sync_module_states: true
fsdp_use_orig_params: false
fsdp_activation_checkpointing: false
main_training_function: main
mixed_precision: bf16
machine_rank: 0
num_machines: 16
num_processes: 128
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
@@ -0,0 +1,28 @@
compute_environment: LOCAL_MACHINE
debug: false
distributed_type: FSDP
downcast_bf16: 'no'
enable_cpu_affinity: false
fsdp_config:
fsdp_auto_wrap_policy: TRANSFORMER_BASED_WRAP
fsdp_transformer_layer_cls_to_wrap: LlamaDecoderLayer
fsdp_backward_prefetch: NO_PREFETCH
fsdp_cpu_ram_efficient_loading: true
fsdp_forward_prefetch: false
fsdp_offload_params: true
fsdp_sharding_strategy: FULL_SHARD
fsdp_state_dict_type: SHARDED_STATE_DICT
fsdp_sync_module_states: true
fsdp_use_orig_params: false
fsdp_activation_checkpointing: false
main_training_function: main
mixed_precision: bf16
machine_rank: 0
num_machines: 2
num_processes: 16
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
@@ -0,0 +1,28 @@
compute_environment: LOCAL_MACHINE
debug: false
distributed_type: FSDP
downcast_bf16: 'no'
enable_cpu_affinity: false
fsdp_config:
fsdp_auto_wrap_policy: TRANSFORMER_BASED_WRAP
fsdp_transformer_layer_cls_to_wrap: LlamaDecoderLayer
fsdp_backward_prefetch: NO_PREFETCH
fsdp_cpu_ram_efficient_loading: true
fsdp_forward_prefetch: false
fsdp_offload_params: true
fsdp_sharding_strategy: FULL_SHARD
fsdp_state_dict_type: SHARDED_STATE_DICT
fsdp_sync_module_states: true
fsdp_use_orig_params: false
fsdp_activation_checkpointing: false
main_training_function: main
mixed_precision: bf16
machine_rank: 0
num_machines: 3
num_processes: 24
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
@@ -0,0 +1,28 @@
compute_environment: LOCAL_MACHINE
debug: false
distributed_type: FSDP
downcast_bf16: 'no'
enable_cpu_affinity: false
fsdp_config:
fsdp_auto_wrap_policy: TRANSFORMER_BASED_WRAP
fsdp_transformer_layer_cls_to_wrap: LlamaDecoderLayer
fsdp_backward_prefetch: NO_PREFETCH
fsdp_cpu_ram_efficient_loading: true
fsdp_forward_prefetch: false
fsdp_offload_params: true
fsdp_sharding_strategy: FULL_SHARD
fsdp_state_dict_type: SHARDED_STATE_DICT
fsdp_sync_module_states: true
fsdp_use_orig_params: false
fsdp_activation_checkpointing: false
main_training_function: main
mixed_precision: bf16
machine_rank: 0
num_machines: 4
num_processes: 32
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
@@ -0,0 +1,28 @@
compute_environment: LOCAL_MACHINE
debug: false
distributed_type: FSDP
downcast_bf16: 'no'
enable_cpu_affinity: false
fsdp_config:
fsdp_auto_wrap_policy: TRANSFORMER_BASED_WRAP
fsdp_transformer_layer_cls_to_wrap: LlamaDecoderLayer
fsdp_backward_prefetch: NO_PREFETCH
fsdp_cpu_ram_efficient_loading: true
fsdp_forward_prefetch: false
fsdp_offload_params: true
fsdp_sharding_strategy: FULL_SHARD
fsdp_state_dict_type: SHARDED_STATE_DICT
fsdp_sync_module_states: true
fsdp_use_orig_params: false
fsdp_activation_checkpointing: false
main_training_function: main
mixed_precision: bf16
machine_rank: 0
num_machines: 1
num_processes: 8
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
@@ -0,0 +1,28 @@
compute_environment: LOCAL_MACHINE
debug: false
distributed_type: FSDP
downcast_bf16: 'no'
enable_cpu_affinity: false
fsdp_config:
fsdp_auto_wrap_policy: TRANSFORMER_BASED_WRAP
fsdp_transformer_layer_cls_to_wrap: LlamaDecoderLayer
fsdp_backward_prefetch: NO_PREFETCH
fsdp_cpu_ram_efficient_loading: true
fsdp_forward_prefetch: false
fsdp_offload_params: true
fsdp_sharding_strategy: HYBRID_SHARD
fsdp_state_dict_type: SHARDED_STATE_DICT
fsdp_sync_module_states: true
fsdp_use_orig_params: false
fsdp_activation_checkpointing: false
main_training_function: main
mixed_precision: bf16
machine_rank: 0
num_machines: 2
num_processes: 16
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
@@ -0,0 +1,28 @@
compute_environment: LOCAL_MACHINE
debug: false
distributed_type: FSDP
downcast_bf16: 'no'
enable_cpu_affinity: false
fsdp_config:
fsdp_auto_wrap_policy: TRANSFORMER_BASED_WRAP
fsdp_transformer_layer_cls_to_wrap: LlamaDecoderLayer
fsdp_backward_prefetch: NO_PREFETCH
fsdp_cpu_ram_efficient_loading: true
fsdp_forward_prefetch: false
fsdp_offload_params: true
fsdp_sharding_strategy: HYBRID_SHARD
fsdp_state_dict_type: SHARDED_STATE_DICT
fsdp_sync_module_states: true
fsdp_use_orig_params: false
fsdp_activation_checkpointing: false
main_training_function: main
mixed_precision: bf16
machine_rank: 0
num_machines: 3
num_processes: 24
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
@@ -0,0 +1,28 @@
compute_environment: LOCAL_MACHINE
debug: false
distributed_type: FSDP
downcast_bf16: 'no'
enable_cpu_affinity: false
fsdp_config:
fsdp_auto_wrap_policy: TRANSFORMER_BASED_WRAP
fsdp_transformer_layer_cls_to_wrap: LlamaDecoderLayer
fsdp_backward_prefetch: NO_PREFETCH
fsdp_cpu_ram_efficient_loading: true
fsdp_forward_prefetch: false
fsdp_offload_params: true
fsdp_sharding_strategy: HYBRID_SHARD
fsdp_state_dict_type: SHARDED_STATE_DICT
fsdp_sync_module_states: true
fsdp_use_orig_params: false
fsdp_activation_checkpointing: false
main_training_function: main
mixed_precision: bf16
machine_rank: 0
num_machines: 4
num_processes: 32
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
@@ -0,0 +1,28 @@
compute_environment: LOCAL_MACHINE
debug: false
distributed_type: FSDP
downcast_bf16: 'no'
enable_cpu_affinity: false
fsdp_config:
fsdp_auto_wrap_policy: TRANSFORMER_BASED_WRAP
fsdp_transformer_layer_cls_to_wrap: Qwen2DecoderLayer
fsdp_backward_prefetch: NO_PREFETCH
fsdp_cpu_ram_efficient_loading: true
fsdp_forward_prefetch: false
fsdp_offload_params: true
fsdp_sharding_strategy: FULL_SHARD
fsdp_state_dict_type: SHARDED_STATE_DICT
fsdp_sync_module_states: true
fsdp_use_orig_params: false
fsdp_activation_checkpointing: false
main_training_function: main
mixed_precision: bf16
machine_rank: 0
num_machines: 1
num_processes: 8
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
@@ -0,0 +1,24 @@
{
"zero_optimization": {
"stage": 2,
"contiguous_gradients": false,
"overlap_comm": false
},
"bf16": {
"enabled": "auto"
},
"fp16": {
"enabled": "auto",
"auto_cast": false,
"loss_scale": 0,
"initial_scale_power": 32,
"loss_scale_window": 1000,
"hysteresis": 2,
"min_loss_scale": 1
},
"gradient_accumulation_steps": "auto",
"gradient_clipping": "auto",
"train_batch_size": "auto",
"train_micro_batch_size_per_gpu": "auto",
"wall_clock_breakdown": false
}
@@ -0,0 +1,31 @@
{
"zero_optimization": {
"stage": 3,
"overlap_comm": false,
"contiguous_gradients": false,
"sub_group_size": 0,
"reduce_bucket_size": "auto",
"stage3_prefetch_bucket_size": "auto",
"stage3_param_persistence_threshold": "auto",
"stage3_max_live_parameters": 0,
"stage3_max_reuse_distance": 0,
"stage3_gather_16bit_weights_on_model_save": true
},
"bf16": {
"enabled": "auto"
},
"fp16": {
"enabled": "auto",
"auto_cast": false,
"loss_scale": 0,
"initial_scale_power": 32,
"loss_scale_window": 1000,
"hysteresis": 2,
"min_loss_scale": 1
},
"gradient_accumulation_steps": "auto",
"gradient_clipping": "auto",
"train_batch_size": "auto",
"train_micro_batch_size_per_gpu": "auto",
"wall_clock_breakdown": false
}
@@ -0,0 +1,43 @@
# Doc about format of conda environment file
# https://conda.io/projects/conda/en/latest/user-guide/tasks/manage-environments.html#create-env-file-manually
name: merge
channels:
- nodefaults
- conda-forge
dependencies:
- _libgcc_mutex=0.1=conda_forge
- _openmp_mutex=4.5=2_gnu
- bzip2=1.0.8=h4bc722e_7
- ca-certificates=2024.7.4=hbcca054_0
- ld_impl_linux-64=2.40=hf3520f5_7
- libffi=3.4.2=h7f98852_5
- libgcc-ng=14.1.0=h77fa898_0
- libgomp=14.1.0=h77fa898_0
- libnsl=2.0.1=hd590300_0
- libsqlite=3.46.0=hde9e2c9_0
- libuuid=2.38.1=h0b41bf4_0
- libxcrypt=4.4.36=hd590300_1
- libzlib=1.3.1=h4ab18f5_1
- ncurses=6.5=h59595ed_0
- openssl=3.3.1=h4bc722e_2
- pip=24.2=pyhd8ed1ab_0
- python=3.10.14=hd12c33a_0_cpython
- readline=8.2=h8228510_1
- setuptools=72.1.0=pyhd8ed1ab_0
- tk=8.6.13=noxft_h4845f30_101
- tzdata=2024a=h0c530f3_0
- wheel=0.44.0=pyhd8ed1ab_0
- xz=5.2.6=h166bdaf_0
- pip:
- --extra-index-url https://download.pytorch.org/whl/cu121
- absl-py==2.1.0
- accelerate==0.34.2 # Needed for fp8
- datasets==2.19.2
- fbgemm-gpu==0.8.0+cu121 # Needed for fp8
- kfp==2.5.0
- peft==0.12.0
- protobuf==3.20.3
- pynvml==11.5.3
- torch==2.4.0+cu121 # Needed for fp8
- transformers==4.47.1
- trl==0.11.2
@@ -0,0 +1,32 @@
# Doc about format of requirement file
# https://pip.pypa.io/en/stable/reference/requirements-file-format
--extra-index-url https://download.pytorch.org/whl/cu118
--extra-index-url https://huggingface.github.io/autogptq-index/whl/cu118/
# keep sorted
accelerate==0.34.2
auto_gptq==0.7.1+cu118
autoawq==0.2.8
bitsandbytes==0.43.2
cloudml-hypertune==0.1.0.dev6
datasets==2.20.0
deepspeed==0.15.2
diffusers==0.38.0
evaluate==0.4.3
fsspec==2024.3.1
gcsfs==2024.3.1
immutabledict==4.2.1
ninja==1.11.1 # Needed to avoid `ninja 1.11.1.1 is not supported on this platform` error
nltk==3.9.1
optimum==1.17.1
peft==0.12.0
pynvml==11.5.3
rouge_score==0.1.2
torch==2.2.2+cu118
torchvision==0.17.2+cu118
transformers==4.47.1
trl==0.11.2
wandb==0.17.1
ydata-profiling==4.7.0 # Upgrade the version from 4.6.0 to 4.7.0 to fix the old `pydantic` package error.
psutil==6.0.0

Some files were not shown because too many files have changed in this diff Show More