Compare commits

...
Author SHA1 Message Date
Rayan DasoriyaandCopybara-Service 9cf8ce16fa Delete deprecated LoRA fine-tuning notebooks and related tutorials.
PiperOrigin-RevId: 976392412
2026-09-04 10:44:57 -07:00
Chun-Hsiang WangandGitHub 4b983a2701 feat: Claude Fable 5.1 Launch (#4581)
* feat: Claude Fable 5.1 Launch

* refactor: replace model/region if-elif chains with a dict lookup

Addresses review feedback on both Select Claude model cells. The mapping is
unchanged for all 20 models; only the lookup mechanism differs.

* chore: apply nbfmt

Runs the repo's own tensorflow-docs nbfmt over the notebook so the
'notebook format and lint' check passes.
2026-09-01 20:45:17 -04:00
Eric DongandGitHub 3b11c876bd fix: correct a typo in error message (#4577) 2026-08-25 17:03:21 -04:00
Mend RenovateandGitHub cc0d791ef2 Update dependency black to v26.5.1 (#4517) 2026-08-19 21:44:22 +00:00
Mend RenovateandGitHub df83a345bb Update dependency isort to v8 (#4444) 2026-08-19 20:52:37 +00:00
Mend RenovateandGitHub e6ded7beaa Update dependency pandas to v3.0.5 (#4491) 2026-08-19 20:51:20 +00:00
Mend RenovateandGitHub e64a4e89d5 chore(deps): update dependency google-cloud-aiplatform to v1.165.0 (#4457) 2026-08-19 20:50:48 +00:00
Mend RenovateandGitHub 7ac54985e4 chore(deps): update dependency smart_open to v8 (#4534) 2026-08-18 22:49:25 +00:00
Mend RenovateandGitHub 6ce96a08d3 Update dependency smart_open to v7.7.1 (#4494) 2026-08-18 21:16:42 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
756711b3c9 chore(deps): bump idna (#4518)
Bumps [idna](https://github.com/kjd/idna) from 3.10 to 3.15.
- [Release notes](https://github.com/kjd/idna/releases)
- [Changelog](https://github.com/kjd/idna/blob/master/HISTORY.md)
- [Commits](https://github.com/kjd/idna/compare/v3.10...v3.15)

---
updated-dependencies:
- dependency-name: idna
  dependency-version: '3.15'
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-18 21:15:21 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
a25d209139 chore(deps): bump torch (#4545)
Bumps [torch](https://github.com/pytorch/pytorch) from 2.8.0 to 2.13.0.
- [Release notes](https://github.com/pytorch/pytorch/releases)
- [Changelog](https://github.com/pytorch/pytorch/blob/main/RELEASE.md)
- [Commits](https://github.com/pytorch/pytorch/compare/v2.8.0...v2.13.0)

---
updated-dependencies:
- dependency-name: torch
  dependency-version: 2.13.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-18 21:14:39 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
59da536b9a chore(deps): bump pillow (#4548)
Bumps [pillow](https://github.com/python-pillow/Pillow) from 12.2.0 to 12.3.0.
- [Release notes](https://github.com/python-pillow/Pillow/releases)
- [Changelog](https://github.com/python-pillow/Pillow/blob/main/CHANGES.rst)
- [Commits](https://github.com/python-pillow/Pillow/compare/12.2.0...12.3.0)

---
updated-dependencies:
- dependency-name: pillow
  dependency-version: 12.3.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-18 21:13:58 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
1bc2839a2b chore(deps): bump pillow (#4568)
Bumps [pillow](https://github.com/python-pillow/Pillow) from 12.2.0 to 12.3.0.
- [Release notes](https://github.com/python-pillow/Pillow/releases)
- [Changelog](https://github.com/python-pillow/Pillow/blob/main/CHANGES.rst)
- [Commits](https://github.com/python-pillow/Pillow/compare/12.2.0...12.3.0)

---
updated-dependencies:
- dependency-name: pillow
  dependency-version: 12.3.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-18 21:13:27 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
0d5e268a1f chore(deps): bump urllib3 (#4512)
Bumps [urllib3](https://github.com/urllib3/urllib3) from 2.6.3 to 2.7.0.
- [Release notes](https://github.com/urllib3/urllib3/releases)
- [Changelog](https://github.com/urllib3/urllib3/blob/main/CHANGES.rst)
- [Commits](https://github.com/urllib3/urllib3/compare/2.6.3...2.7.0)

---
updated-dependencies:
- dependency-name: urllib3
  dependency-version: 2.7.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-18 21:12:21 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
1b019a76e4 Bump google-cloud-aiplatform (#4446)
Bumps [google-cloud-aiplatform](https://github.com/googleapis/python-aiplatform) from 1.92.0 to 1.133.0.
- [Release notes](https://github.com/googleapis/python-aiplatform/releases)
- [Changelog](https://github.com/googleapis/python-aiplatform/blob/main/CHANGELOG.md)
- [Commits](https://github.com/googleapis/python-aiplatform/compare/v1.92.0...v1.133.0)

---
updated-dependencies:
- dependency-name: google-cloud-aiplatform
  dependency-version: 1.133.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-18 21:11:56 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
87c1ed686a chore(deps): bump diffusers (#4510)
Bumps [diffusers](https://github.com/huggingface/diffusers) from 0.25.1 to 0.38.0.
- [Release notes](https://github.com/huggingface/diffusers/releases)
- [Commits](https://github.com/huggingface/diffusers/compare/v0.25.1...v0.38.0)

---
updated-dependencies:
- dependency-name: diffusers
  dependency-version: 0.38.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-18 21:11:22 +00:00
Mend RenovateandGitHub 187fdc526c Update dependency datasets to v5 (#4521) 2026-08-18 21:10:37 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
215c8eee3e chore(deps): bump urllib3 (#4513)
Bumps [urllib3](https://github.com/urllib3/urllib3) from 2.6.3 to 2.7.0.
- [Release notes](https://github.com/urllib3/urllib3/releases)
- [Changelog](https://github.com/urllib3/urllib3/blob/main/CHANGES.rst)
- [Commits](https://github.com/urllib3/urllib3/compare/2.6.3...2.7.0)

---
updated-dependencies:
- dependency-name: urllib3
  dependency-version: 2.7.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-18 21:10:01 +00:00
f90cd0d6ed Add AlphaFold 3 quickstart notebook (#4572)
* Add AlphaFold 3 quickstart notebook

* Update CODEOWNERS

---------

Co-authored-by: Amit Rai <raiamit@google.com>
2026-08-17 13:54:13 -07:00
Mend RenovateandGitHub 1985f06e99 Update dependency numpy to v2.5.2 (#4516) 2026-08-14 18:43:39 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
ff428dc589 chore(deps): bump torch (#4544)
Bumps [torch](https://github.com/pytorch/pytorch) from 2.7.0 to 2.13.0.
- [Release notes](https://github.com/pytorch/pytorch/releases)
- [Changelog](https://github.com/pytorch/pytorch/blob/main/RELEASE.md)
- [Commits](https://github.com/pytorch/pytorch/compare/v2.7.0...v2.13.0)

---
updated-dependencies:
- dependency-name: torch
  dependency-version: 2.13.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-14 18:41:44 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
f6124370b0 chore(deps): bump pillow (#4547)
Bumps [pillow](https://github.com/python-pillow/Pillow) from 12.1.1 to 12.3.0.
- [Release notes](https://github.com/python-pillow/Pillow/releases)
- [Changelog](https://github.com/python-pillow/Pillow/blob/main/CHANGES.rst)
- [Commits](https://github.com/python-pillow/Pillow/compare/12.1.1...12.3.0)

---
updated-dependencies:
- dependency-name: pillow
  dependency-version: 12.3.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-14 18:40:42 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
b8822f5008 chore(deps): bump pyasn1 (#4549)
Bumps [pyasn1](https://github.com/pyasn1/pyasn1) from 0.6.3 to 0.6.4.
- [Release notes](https://github.com/pyasn1/pyasn1/releases)
- [Changelog](https://github.com/pyasn1/pyasn1/blob/main/CHANGES.rst)
- [Commits](https://github.com/pyasn1/pyasn1/compare/v0.6.3...v0.6.4)

---
updated-dependencies:
- dependency-name: pyasn1
  dependency-version: 0.6.4
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-14 18:39:56 +00:00
Dustin LuongandCopybara-Service 8976c57b9c Update the Kimi-K3 deployment notebook image URI.
PiperOrigin-RevId: 964692807
2026-08-14 07:39:13 -07:00
gmaninatarajanandGitHub 3985da440e fix: Updated new whl file with SDK update to add interval_variants parameter to score_ism_variants() (#4565)
* fix: Updated new whl file with SDK update to add interval_variants parameter to score_ism_variants()

* fix: updating the whl file download cell
2026-08-11 19:56:05 -04:00
Damodar PanigrahiGitHubgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
1c9092ced3 refactor - restructure the notebook (#4564)
* refactor - restructure the notebook

* Update notebooks/community/weathernext/CUSTOM_INPUTS_GUIDE.md

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update notebooks/community/weathernext/weathernext_2_ic_pc.ipynb

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update notebooks/community/weathernext/weathernext_2_dws.ipynb

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

---------

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-08-07 20:53:54 +00:00
Damodar PanigrahiandGitHub 77b2af09ce feat: WN2 with GPU GA (#4563) 2026-08-07 17:27:35 +00:00
genquan9andGitHub c6d33c2a0d Add tau2-bench RL blog post to docs README (#4561) 2026-08-05 23:18:11 +00:00
genquan9andGitHub 89772b320e Fix inline math rendering: use span+1798467 for GitHub Pages MathJax (#4560) 2026-08-05 17:47:23 +00:00
genquan9andGitHub a1f6d2c069 Fix LaTeX rendering for Pass Rate formula (#4559)
Replace underscores in \text{num\_pass} with spaces to avoid
LaTeX math mode errors on GitHub rendering.
2026-08-05 17:30:29 +00:00
genquan9andGitHub 936a6adf77 Add multi-turn RL for tau2-bench technical report (#4558)
* Add multi-turn RL for tau2-bench technical report

Add technical report documenting multi-turn reinforcement learning
training pipeline for tau2-bench customer service benchmark, including
GRPO training, data synthesis pipeline, and evaluation results.

* Fix deprecated MathJax CDN and broken anchor link

- Remove deprecated cdn.mathjax.org script tag (GitHub renders LaTeX natively)
- Fix broken ToC anchor from #2-bench to #tau2-bench
2026-08-05 16:25:05 +00:00
Dustin LuongandCopybara-Service 37a85d53f4 No public description
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 958290655
2026-08-03 04:07:45 -07:00
Dustin LuongandCopybara-Service 0b0e362ab9 Add Kimi K3 Model Garden deployment notebook
PiperOrigin-RevId: 958290655
2026-08-03 04:06:44 -07:00
Damodar PanigrahiandGitHub 98103d462f test (#4554) 2026-07-30 23:37:05 +00:00
Oleh PrypinandCopybara-Service 9ea1cf3b86 No public description
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 955252640
2026-07-28 07:46:15 -07:00
Sam-DecigaandGitHub 8f3e6668e1 feat: Claude Opus 5 Launch (#4552) 2026-07-26 10:04:43 -04:00
Tianzi CaiandGitHub 5d9853db5c Fix formatting in Anthropic Claude intro notebook 2026-07-22 20:59:25 -07:00
Tianzi CaiandGitHub 003fb5121b Remove unused httpx imports and related comments 2026-07-22 20:56:48 -07:00
Tianzi CaiandGitHub a62695fb38 Update image URL and request handling in notebook (#4551)
* Update image URL and request handling in notebook

* Remove Colab link markdown cell

Removed markdown cell with Colab link from the notebook.

* Remove unused import
2026-07-22 23:52:09 +00:00
Sam-DecigaandGitHub 3c630fdbb8 feat: Claude-Sonnet5-Launch (#4537) 2026-06-30 16:31:27 -04:00
Damodar PanigrahiandGitHub 6ca1d899d6 feat: wn2 doc polished (#4531) 2026-06-23 23:22:47 +00:00
Damodar PanigrahiandGitHub 1894602fff feat: wn2 notebook (#4530) 2026-06-23 21:59:17 +00:00
Damodar PanigrahiandGitHub 31a52d6e92 feat: WeatherNext IC (#4523)
* feat: WeatherNext IC

* Fix: Replace weathernext_2_ic_early_access_program.ipynb symlink with actual notebook file

* fix: Replace Vertex Jobs with Gemini Enterprise Agent Platform Jobs in WeatherNext notebook

* fix: Correct typos, broken links, and apply linter formatting
2026-06-11 19:49:13 +00:00
0f9d9734c3 feat: Claude Fable 5 Launch (#4522)
Co-authored-by: Holt Skinner <13262395+holtskinner@users.noreply.github.com>
2026-06-09 14:47:46 -04:00
Vertex MG TeamandCopybara-Service e85cf9a174 Update link to Cloud Quotas page to correct location
PiperOrigin-RevId: 926490177
2026-06-03 23:21:57 -07:00
Sam-DecigaandGitHub b4c0bbc1a0 feat: Ant-Opus4.8 Launch (#4520) 2026-05-28 14:43:47 -04:00
Rayan DasoriyaandCopybara-Service 24244351cd Add a new notebook for OSS distillation feasibility study.
PiperOrigin-RevId: 917879385
2026-05-19 09:37:47 -07:00
Vertex MG TeamandCopybara-Service bf0e1300a9 No public description
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 878451476
2026-05-13 12:58:56 -07:00
chnduandGitHub cf048b6fe4 Add live_api skills that help the user build their own liveapi service (#4511)
* Add live_api skills that help the user to build their own liveapi service.

Implementation are based on websocket. Support different coding languages.

* Update based on review

* Fix typos

* Update vertex to gemini enterprise.
2026-05-11 17:19:12 +00:00
Mend RenovateandGitHub 8c8820ecfa chore(deps): update dependency numpy to v2.4.4 (#4466) 2026-05-06 14:24:57 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
a1a52d8145 chore(deps): bump requests (#4488)
Bumps [requests](https://github.com/psf/requests) from 2.32.4 to 2.33.0.
- [Release notes](https://github.com/psf/requests/releases)
- [Changelog](https://github.com/psf/requests/blob/main/HISTORY.md)
- [Commits](https://github.com/psf/requests/compare/v2.32.4...v2.33.0)

---
updated-dependencies:
- dependency-name: requests
  dependency-version: 2.33.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-06 14:22:23 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
a06ce545e7 chore(deps): bump pillow (#4497)
Bumps [pillow](https://github.com/python-pillow/Pillow) from 12.1.1 to 12.2.0.
- [Release notes](https://github.com/python-pillow/Pillow/releases)
- [Changelog](https://github.com/python-pillow/Pillow/blob/main/CHANGES.rst)
- [Commits](https://github.com/python-pillow/Pillow/compare/12.1.1...12.2.0)

---
updated-dependencies:
- dependency-name: pillow
  dependency-version: 12.2.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-06 14:19:28 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
913780c4cb chore(deps): bump pillow (#4509)
Bumps [pillow](https://github.com/python-pillow/Pillow) from 10.3.0 to 12.2.0.
- [Release notes](https://github.com/python-pillow/Pillow/releases)
- [Changelog](https://github.com/python-pillow/Pillow/blob/main/CHANGES.rst)
- [Commits](https://github.com/python-pillow/Pillow/compare/10.3.0...12.2.0)

---
updated-dependencies:
- dependency-name: pillow
  dependency-version: 12.2.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-06 14:14:14 +00:00
gmaninatarajanandGitHub 71be46e7d8 feat: Updated whl file and package name as part of Vertex Model Garden setup (#4508) 2026-04-30 08:28:04 -04:00
Mayank SharanandGitHub 849e88a627 Vtc blog 2 (#4507)
* Adding reviewed version of VTC blog 2

* GCA suggested fixes

* Updating readme to have links
2026-04-29 19:10:45 +00:00
Jason DaiandGitHub daf56bcd0b Create Eval Quality Flywheel Skill for preview (#4505) 2026-04-23 16:37:08 +00:00
Mayank SharanandGitHub 563f423b93 Adding reviewed version of VTC blog 2 (#4502)
* Adding reviewed version of VTC blog 2

* GCA suggested fixes
2026-04-16 21:12:03 +00:00
ian1780andGitHub 292e540e96 Update anthropic_claude_intro.ipynb (#4501)
add opus 4.7 multi region endpoint support b/491171457
2026-04-16 17:59:12 +00:00
Sam-DecigaandGitHub 6c6a703c5a Ant nickel (#4500)
* Feat: Anthropic Opus-4-7 launch

* Feat: Anthropic Opus-4-7 launch
2026-04-16 12:18:47 -04:00
Sam-DecigaandGitHub 7ef83c6f73 MARS8 new Asian regions (#4498) 2026-04-14 19:59:45 +00:00
Vertex MG TeamandCopybara-Service aba6598109 use old docker hash for whisper model deployment.
PiperOrigin-RevId: 892120947
2026-03-30 23:08:58 -07:00
Sam-DecigaandGitHub 88a6b8037e Refactor: Anthropic NB (#4489) 2026-03-26 14:49:59 -04:00
Eric DongandGitHub 8845f7ab27 Update Gemini model references and availability details
Update Gemini versions.
2026-03-26 09:40:01 -04:00
Eric DongandGitHub 3b2e711a16 Update fine-tuning model reference in README
Updated the fine-tuning model reference from Gemini 1.5 Pro to Gemini 2.5 Pro in the README.
2026-03-25 16:44:52 -04:00
Eric DongandGitHub 5c0629cdc7 Revise README for Agent Skills in Vertex AI
Updated terminology and formatting for clarity.
2026-03-25 15:18:03 -04:00
Eric DongandGitHub e107d30807 chore: Add detailed installation instructions for skills (#4487) 2026-03-25 15:06:59 -04:00
Eric DongandGitHub b98ab36913 refactor: Add tool configuation in skills readme (#4486)
* chore: Update vertex-ai Skills readme

* refactor: Add tool configuation in  skills readme
2026-03-25 13:41:45 -04:00
Eric DongandGitHub f1d90b5a71 chore: Update vertex-ai Skills readme (#4485) 2026-03-25 11:36:41 -04:00
Eric DongandGitHub f848db6132 Revise README title and formatting for emphasis
Updated the title and emphasized 'Skills' in the README.
2026-03-25 10:16:42 -04:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
cc9fffd945 chore(deps): bump pillow (#4482)
Bumps [pillow](https://github.com/python-pillow/Pillow) from 10.3.0 to 12.1.1.
- [Release notes](https://github.com/python-pillow/Pillow/releases)
- [Changelog](https://github.com/python-pillow/Pillow/blob/main/CHANGES.rst)
- [Commits](https://github.com/python-pillow/Pillow/compare/10.3.0...12.1.1)

---
updated-dependencies:
- dependency-name: pillow
  dependency-version: 12.1.1
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-03-24 15:13:00 +00:00
Eric DongandGitHub 7606a1de03 chore: Update the skills readme with instructions (#4484) 2026-03-24 10:44:28 -04:00
Eric DongandGitHub cca59aa753 Update README.md
Remove icons
2026-03-24 10:12:04 -04:00
Eric DongandGitHub a1907da27a chore: Update readme (#4483) 2026-03-24 10:00:50 -04:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
e8cb7738d0 Bump pillow (#4443)
Bumps [pillow](https://github.com/python-pillow/Pillow) from 10.3.0 to 12.1.1.
- [Release notes](https://github.com/python-pillow/Pillow/releases)
- [Changelog](https://github.com/python-pillow/Pillow/blob/main/CHANGES.rst)
- [Commits](https://github.com/python-pillow/Pillow/compare/10.3.0...12.1.1)

---
updated-dependencies:
- dependency-name: pillow
  dependency-version: 12.1.1
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-03-24 13:52:57 +00:00
Mend RenovateandGitHub 18e8d603de chore(deps): update dependency black to v26.3.1 [security] (#4470) 2026-03-24 13:52:02 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
1bd5901fdb Bump black (#4468)
Bumps [black](https://github.com/psf/black) from 25.1.0 to 26.3.1.
- [Release notes](https://github.com/psf/black/releases)
- [Changelog](https://github.com/psf/black/blob/main/CHANGES.md)
- [Commits](https://github.com/psf/black/compare/25.1.0...26.3.1)

---
updated-dependencies:
- dependency-name: black
  dependency-version: 26.3.1
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-03-24 13:51:38 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
5b245024cd Bump pyasn1 (#4474)
Bumps [pyasn1](https://github.com/pyasn1/pyasn1) from 0.6.2 to 0.6.3.
- [Release notes](https://github.com/pyasn1/pyasn1/releases)
- [Changelog](https://github.com/pyasn1/pyasn1/blob/main/CHANGES.rst)
- [Commits](https://github.com/pyasn1/pyasn1/compare/v0.6.2...v0.6.3)

---
updated-dependencies:
- dependency-name: pyasn1
  dependency-version: 0.6.3
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-03-24 13:50:31 +00:00
Eric DongandGitHub ba043c196c chore: Update skills readme with architecture (#4481) 2026-03-24 09:49:35 -04:00
Eric DongandGitHub 28ce8f6d7a chore: Update readme and template (#4480) 2026-03-24 09:42:45 -04:00
Eric DongandGitHub 3b5a8cad41 feat: Add Gen AI SDK skill for Vertex (#4479) 2026-03-24 09:29:01 -04:00
gmaninatarajanandGitHub c6d7971bc9 fix:Simplified authentication section and addressed timeout issues (#4478) 2026-03-23 18:50:35 -04:00
Eric DongandGitHub dbe28965cb feat: Add primary routing and readme for vertex ai skills (#4477) 2026-03-23 17:18:16 -04:00
Vertex MG TeamandCopybara-Service 0d34d6bbea update minimax m2 notebook.
PiperOrigin-RevId: 886885747
2026-03-20 11:14:55 -07:00
Sam-DecigaandGitHub a1d898f35e feat: Jina EmbV3 launch (#4476)
* feat: Jina EmbV3 launch

* feat: Jina EmbV3 launch
2026-03-20 08:19:38 -04:00
Eric DongandGitHub 772ee71bc3 feat: use Vertex AI MCP server (#4475)
* feat: use Vertex AI MCP server

* Address review comments
2026-03-19 11:02:28 -04:00
Sam-DecigaandGitHub 425851cedc feat: Nemotron3-Super model launch (#4473)
* feat: Nemotron3-Super model launch

* feat: Nemotron3-Super model launch

* feat: Nemotron3-Super model launch
2026-03-16 20:41:36 -04:00
Lav RaiandGitHub bcccbee164 Update distillation report. (#4472) 2026-03-16 18:40:10 +00:00
Lav RaiandGitHub 5ae325528a Add distillation report. (#4471) 2026-03-13 15:33:34 +00:00
vincentkt-googleandGitHub 86674effee Add and update existing vertex skills (#4467)
* Add and update existing vertex skills

- Add support for fine tuning for 1p gemini tuning
- Add support for deploying fine tuned model support
- Add support for running inference on MaaS models
- Add open model support for regions and cost estimating for 3p tuning

* fixing some of the commit errors

* updated scripts to use existing gemini 1.5 pro model

* swap gemini 1.5 pro to gemini 2.5 pro
2026-03-11 19:45:42 +00:00
vincentkt-googleandGitHub 8b4708c606 feat: add vertex ai skills to repo (#4454) 2026-03-05 17:50:41 +00:00
Yichen ZhouandCopybara-Service f3dd6cbca3 Update TimesFM-2.5 notebook for Model Garden.
PiperOrigin-RevId: 878726929
2026-03-04 16:42:55 -08:00
Rayan DasoriyaandCopybara-Service 1f9e93993c No public description
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 878215121
2026-03-03 18:26:58 -08:00
Vertex MG TeamandCopybara-Service 062835174e Updated the image default TAG to release
PiperOrigin-RevId: 877893209
2026-03-03 05:26:21 -08:00
Rayan DasoriyaandCopybara-Service cb4916f590 No public description
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 875506020
2026-02-25 21:44:41 -08:00
Damodar PanigrahiandGitHub 2933fe606b bug: remove A100 as recommended specs (#4450) 2026-02-25 14:05:36 +00:00
Damodar PanigrahiandGitHub b468809df7 bug: ahref update (#4449)
* bug: ahref update

* fix: linter

* fix: typo fix
2026-02-25 13:46:19 +00:00
Sam-DecigaandGitHub 0417d8b9c4 feat: Deprecate Claude 3 Haiku (#4448)
Deprecation start date: Feb. 23, 2026
End of Support date: Aug. 23, 2026
b/485993204
2026-02-23 20:57:43 -05:00
Damodar PanigrahiandGitHub 7750e83fbb fix: inference key change, finetuning jaxlib update (#4447) 2026-02-23 19:45:59 +00:00
Damodar PanigrahiandGitHub 649800e646 feat: Alphagenome finetuning notebook (#4445)
* feat: Alphagenome finetuning notebook

* Update cloudai_alphagenome_finetune.ipynb

Fixed the lint errors

* Update cloudai_alphagenome_finetune.ipynb

Fix lint errors

* feat: Add Alphagenome finetune

* feat: Include the Alphagenome finetuning notebook url in the readme. Update the codeowners

* feat: Add alphagegenome finetune notebook to the readme, add user to codeowners

* feat: fix spelling
2026-02-20 14:13:36 +00:00
Mend RenovateandGitHub 42b35056fa chore(deps): update dependency black to v26 (#4422) 2026-02-18 15:07:12 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
24974eda95 Bump protobuf (#4437)
Bumps [protobuf](https://github.com/protocolbuffers/protobuf) from 4.25.8 to 5.29.6.
- [Release notes](https://github.com/protocolbuffers/protobuf/releases)
- [Commits](https://github.com/protocolbuffers/protobuf/commits)

---
updated-dependencies:
- dependency-name: protobuf
  dependency-version: 5.29.6
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-02-18 15:06:37 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
ad41377783 Bump protobuf (#4438)
Bumps [protobuf](https://github.com/protocolbuffers/protobuf) from 4.25.8 to 5.29.6.
- [Release notes](https://github.com/protocolbuffers/protobuf/releases)
- [Commits](https://github.com/protocolbuffers/protobuf/commits)

---
updated-dependencies:
- dependency-name: protobuf
  dependency-version: 5.29.6
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-02-18 15:06:10 +00:00
Sam-DecigaandGitHub 36dea3ca01 feat: Anthropic Sonnet-4-6 launch (#4442)
Signed-off-by: Sam-Deciga <decigagarcia@google.com>
2026-02-17 14:02:30 -05:00
b75b2ea4d7 fix: updated with latest whl file version : alphagenome-0.4.2.6-py3-none-any.whl (#4439)
Co-authored-by: hyper-param <peeyusht@google.com>
2026-02-10 18:27:03 +00:00
Vertex MG TeamandCopybara-Service 28f7fc4445 Add SAM 3 notebook to Vertex AI Model Garden.
PiperOrigin-RevId: 866598579
2026-02-06 13:45:47 -08:00
Rayan DasoriyaandCopybara-Service bf2c1226fd Update the license year
PiperOrigin-RevId: 866214899
2026-02-05 19:07:14 -08:00
Vertex MG TeamandCopybara-Service 2990c53292 Added notebook sample for batch inference using the remote sensing VMG models
PiperOrigin-RevId: 866061602
2026-02-05 12:27:10 -08:00
Sam-DecigaandGitHub 531d9cfee0 feat: New Anthropic model (#4436) 2026-02-05 14:07:05 -05:00
Vertex MG TeamandCopybara-Service ff18ec7af5 Add --total-gpus to multi-model model-cohost deployment config.
PiperOrigin-RevId: 863070590
2026-01-29 22:39:06 -08:00
Sam-DecigaandGitHub 008eb409ef feat: NVIDIA-Llama-Nemotron-Super-49B (#4430) 2026-01-29 08:51:52 -05:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
b2dba4b568 Bump pyasn1 (#4421)
Bumps [pyasn1](https://github.com/pyasn1/pyasn1) from 0.6.1 to 0.6.2.
- [Release notes](https://github.com/pyasn1/pyasn1/releases)
- [Changelog](https://github.com/pyasn1/pyasn1/blob/main/CHANGES.rst)
- [Commits](https://github.com/pyasn1/pyasn1/compare/v0.6.1...v0.6.2)

---
updated-dependencies:
- dependency-name: pyasn1
  dependency-version: 0.6.2
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-01-27 20:14:34 +00:00
Sam-DecigaandGitHub 665547f790 feat:New MARS8 model (#4429)
* feat:New MARS8 model

* feat:New MARS8 model
2026-01-27 10:43:43 -05:00
Vertex MG TeamandCopybara-Service 5078c44eb8 Updated the bucket path for the remote sensing models.
PiperOrigin-RevId: 861234161
2026-01-26 09:47:01 -08:00
Bhaskar GoyalandGitHub da6e46531e feature: Retire Mistral 24.11 and Codestral 25.01 from Mistral Intro files. (#4427) 2026-01-23 18:47:36 +00:00
0a4091a3b1 fix: Fix json response parsing error (#4426)
Co-authored-by: hyper-param <peeyusht@google.com>
2026-01-23 13:58:03 +00:00
Sam-DecigaandGitHub 5afa83dd25 feat: MongoDB voyage-4 launch (#4419)
* feat: MongoDB voyage-4 launch

* feat: MongoDB voyage-4 launch

* feat: MongoDB voyage-4 launch

* feat: MongoDB voyage-4 launch
2026-01-16 16:26:05 -05:00
Sam-DecigaandGitHub 8cab85d6ad feat: MongoDB voyage-multimodal-3.5 launch (#4420)
* feat: MongoDB voyage-multimodal-3.5 launch

* feat: MongoDB voyage-multimodal-3.5 launch
2026-01-16 14:57:11 -05:00
Eric DongandGitHub a7f3940635 refactor: Update Github icon (#4418) 2026-01-13 13:36:20 -05:00
Mend RenovateandGitHub 5bb1a75a48 chore(deps): update dependency pyupgrade to v3.21.2 (#4359) 2026-01-13 18:26:35 +00:00
Damodar PanigrahiandGitHub 8fe4985aa8 feat: WeatherNext2 Initial Updates (#4417) 2026-01-13 13:24:48 -05:00
intentsolutions.ioandGitHub 996b6534d9 Add ADK inline source deployment tutorial for Agent Engine (#4393)
* Add ADK inline source deployment tutorial for Agent Engine

* fix: address Gemini review feedback

- Change model from gemini-2.0-flash to gemini-1.5-flash-001
- Improve exception handling with ZoneInfoNotFoundError

* fix: address Gemini code review feedback

- Use specific ZoneInfoNotFoundError exception instead of generic Exception
- Define REQUIREMENTS variable once and reuse to avoid duplication
- Keep generic Exception as fallback for unexpected errors

🤖 Generated with [Claude Code](https://claude.com/claude-code)

* style: fix notebook formatting via official linter
2026-01-13 08:56:07 -05:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
1061ae5348 Bump urllib3 (#4416)
Bumps [urllib3](https://github.com/urllib3/urllib3) from 2.6.0 to 2.6.3.
- [Release notes](https://github.com/urllib3/urllib3/releases)
- [Changelog](https://github.com/urllib3/urllib3/blob/main/CHANGES.rst)
- [Commits](https://github.com/urllib3/urllib3/compare/2.6.0...2.6.3)

---
updated-dependencies:
- dependency-name: urllib3
  dependency-version: 2.6.3
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-01-12 20:01:22 +00:00
Eric DongandGitHub 4fac3a630f refactor: reformat notebook template (#4414)
* refactor: reformat notebook template

* Update Python version to 3.12

* Fix an import

* Update licience year
2026-01-12 13:15:30 -05:00
Eric DongandGitHub 1f2db2903c refactor: Updated Github icons (#4413) 2026-01-12 10:10:34 -05:00
Sam-DecigaandGitHub 912a52de70 feat:MongoDB Voyage 3.5-Lite (#4403)
* feat:MongoDB Voyage 3.5-Lite

* chore:apply linter

* Update voyage-3.5-lite.ipynb

Fixing MODEL_NAME
2026-01-06 21:19:17 -05:00
Sam-DecigaandGitHub 5e29090e86 Model Deprecation (#4409)
Haiku 3.5 Model Deprecation
2026-01-05 16:30:23 -05:00
Mend RenovateandGitHub cd8fcd1839 Update actions/checkout action to v6 (#4374) 2026-01-05 14:20:36 +00:00
Mend RenovateandGitHub e51075ec4b chore(deps): update dependency black to v25.12.0 (#4360) 2026-01-05 14:18:08 +00:00
Mend RenovateandGitHub 9709c0dddb chore(deps): update dependency isort to v7 (#4289) 2026-01-05 14:14:26 +00:00
Mend RenovateandGitHub a21ae41762 Update dependency python to 3.14 (#4398) 2026-01-05 14:13:11 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
ff5ff7b609 Bump urllib3 (#4386)
Bumps [urllib3](https://github.com/urllib3/urllib3) from 2.5.0 to 2.6.0.
- [Release notes](https://github.com/urllib3/urllib3/releases)
- [Changelog](https://github.com/urllib3/urllib3/blob/main/CHANGES.rst)
- [Commits](https://github.com/urllib3/urllib3/compare/2.5.0...2.6.0)

---
updated-dependencies:
- dependency-name: urllib3
  dependency-version: 2.6.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-01-05 14:12:43 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
23748f443e Bump urllib3 (#4387)
Bumps [urllib3](https://github.com/urllib3/urllib3) from 2.5.0 to 2.6.0.
- [Release notes](https://github.com/urllib3/urllib3/releases)
- [Changelog](https://github.com/urllib3/urllib3/blob/main/CHANGES.rst)
- [Commits](https://github.com/urllib3/urllib3/compare/2.5.0...2.6.0)

---
updated-dependencies:
- dependency-name: urllib3
  dependency-version: 2.6.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-01-05 14:12:18 +00:00
Vertex MG TeamandCopybara-Service fc6b2167de ComfyUI tutorial notebook
PiperOrigin-RevId: 851380881
2026-01-02 10:20:16 -08:00
e79a45358c Migrate gsutil usage to gcloud storage (#4299)
* Migrate gsutil usage to gcloud storage

* changes for 4299

* Apply automated linter fixes

* update

* remove model_garden changes

* revert to main

* revert model garden file

* Update model_garden_weather_prediction_on_vertex.ipynb

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-31 08:37:14 -05:00
20d19fb11c Migrate gsutil usage to gcloud storage (#4331)
* Migrate gsutil usage to gcloud storage

* Manual Changes

* Changes for 4331

* Changes for 4331

* Removed changes model garden

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-31 08:36:46 -05:00
9c9f7a6e2a Migrate gsutil usage to gcloud storage (#4334)
* Migrate gsutil usage to gcloud storage

* Manual Changes

* changes for 4334

* fix linter issue\ for 4334

* manual changes

* Restore gcloud migration code

* Remove changes for model garden

* update

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-31 08:36:21 -05:00
ede41c2115 Migrate gsutil usage to gcloud storage (#4335)
* Migrate gsutil usage to gcloud storage

* remoed changes for model garden

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-31 08:35:53 -05:00
ca53786c04 Migrate gsutil usage to gcloud storage (#4322)
* Migrate gsutil usage to gcloud storage

* Changes for 4322

* fix linter issue for 4322

* removed changes for model garden

* Update model_garden_pytorch_gemma_peft_finetuning_hf.ipynb

* Update model_garden_pytorch_gemma_peft_finetuning_hf.ipynb

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-25 10:49:46 -05:00
419f8310c9 Migrate gsutil usage to gcloud storage (#4316)
* Migrate gsutil usage to gcloud storage

* Changes for 4316

* Changes for 4316

* fix linter issue for 4316

* removed changes for model garden

* removed changes for model garden

* Update model_garden_tfvision_image_classification.ipynb

* Update model_garden_tfvision_image_classification.ipynb

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-25 10:49:33 -05:00
996b690e03 Migrate gsutil usage to gcloud storage (#4317)
* Migrate gsutil usage to gcloud storage

* Changes for 4317

* Changes for 4317

* fix linter issue for 4317

* removed changes for model garden

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-25 15:48:30 +00:00
5b9d04d63a Migrate gsutil usage to gcloud storage (#4318)
* Migrate gsutil usage to gcloud storage

* Changes for 4318

* fix linter issue for 4318

* removed changes for model garden:

* Update model_garden_gemma2_deployment_on_vertex.ipynb

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-25 10:47:49 -05:00
9ed3c2f83d Migrate gsutil usage to gcloud storage (#4320)
* Migrate gsutil usage to gcloud storage

* Changes for 4320

* fix linter issue for 4320

* removed chnages for model garden

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-25 10:47:17 -05:00
5fc93c8bdb Migrate gsutil usage to gcloud storage (#4328)
* Migrate gsutil usage to gcloud storage

* PR changes for 4328

* Fix Linter issue for 4328

* changes removed model garden

* Update model_garden_jax_fvlm.ipynb

* Update model_garden_jax_fvlm.ipynb

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-25 10:46:52 -05:00
3cf46226e9 Migrate gsutil usage to gcloud storage (#4321)
* Migrate gsutil usage to gcloud storage

* Changes for 4321

* fix the linter issue for br 4321

* removed changes for model garden

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-25 10:46:04 -05:00
439f6a0cae Migrate gsutil usage to gcloud storage (#4329)
* Migrate gsutil usage to gcloud storage

* Manual Changes

* Manual Changes

* Linter fiex the issues for 4329

* Revert "Linter fiex the issues for 4329"

This reverts commit a1ed9c6f59.

* Revert "Manual Changes"

This reverts commit 8ca9b56c5b.

* changes removed model garden

* Update model_garden_gemma2_finetuning_on_vertex.ipynb

* Update model_garden_pytorch_llama3_3_finetuning.ipynb

* Update model_garden_pytorch_llama3_3_finetuning.ipynb

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-25 10:45:30 -05:00
993898bb71 Migrate gsutil usage to gcloud storage (#4330)
* Migrate gsutil usage to gcloud storage

* Manual Changes

* Removed blank line and spaces Manual Changes

* fix the Linter issue for 4330

* Changes for model garden

* Update model_garden_pytorch_llama3_1_finetuning.ipynb

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-25 15:45:00 +00:00
9e9e639375 Migrate gsutil usage to gcloud storage (#4323)
* Migrate gsutil usage to gcloud storage

* Manual Changes-Migrate gsutil usage to gcloud storage

* Manual Changes

* Fix linter issue for 4323

* removed changes for model garden

* Update model_garden_axolotl_gpt_oss_finetuning.ipynb

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-25 10:44:34 -05:00
633cf6a799 Migrate gsutil usage to gcloud storage (#4327)
* Migrate gsutil usage to gcloud storage

* Manual Changes-Migrate gsutil usage to gcloud storage

* Changes for 4327

* fixing linting erro

* removed changes of model garden

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-25 10:44:05 -05:00
8b618bc455 Migrate gsutil usage to gcloud storage (#4326)
* Migrate gsutil usage to gcloud storage

* Manual Changes-Updated the cell by replacing 'gsutil copy' with the correct 'gcloud storage cp'

* Manual Changes-Updated the cell by replacing 'gsutil copy' with the correct 'gcloud storage cp'

* Revert "Manual Changes-Updated the cell by replacing 'gsutil copy' with the correct 'gcloud storage cp'"

This reverts commit 175eaa4fe8.

* Manual Changes-Updated the cell by replacing 'gsutil copy' with the correct 'gcloud storage cp'

* Changes for 4326

* Changes for 4326

* Linter fix issue for 4326

* removed model garden changes

* Update model_garden_movinet_action_recognition.ipynb

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-25 10:43:31 -05:00
6bbe3bcfe0 Migrate gsutil usage to gcloud storage (#4336)
* Migrate gsutil usage to gcloud storage

* Manual Changes

* removed model garden changes

* Update model_garden_llama3_1_finetuning_with_workbench.ipynb

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-25 10:42:44 -05:00
Margubur RahmanandGitHub 814827ac19 Migrate gsutil usage to gcloud storage (#4332) 2025-12-25 10:42:12 -05:00
Margubur RahmanandGitHub 7a613785b9 Migrate gsutil usage to gcloud storage (#4333) 2025-12-25 15:41:38 +00:00
100243e90a Migrate gsutil usage to gcloud storage (#4338)
* Migrate gsutil usage to gcloud storage

* Manual Changes

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
2025-12-25 15:40:46 +00:00
6424515b03 Migrate gsutil usage to gcloud storage (#4339)
* Migrate gsutil usage to gcloud storage

* Manual Changes

* Manual Changes

* Fix: Updated gcloud storage command without formatting

* Manual Changes

* Manual Changes

* Revert "Manual Changes"

This reverts commit a7a7bda0f9.

* Manual Changes

* Revert "Manual Changes"

This reverts commit a7a7bda0f9.

* Manual Changes

* Revert "Manual Changes"

This reverts commit 71c777d5f1.

* Manual Changes

* Manual changes

* Changes for 4339

* Changes for 4339

* Changes for 4339

* Fix: Applied linter formatting and resolved style issues

* gcloud to gsutilchanges for  4339

* removed gsutil to gcloud migration

* Manual changes

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-25 15:40:00 +00:00
8d22b221b4 Migrate gsutil usage to gcloud storage (#4337)
* Migrate gsutil usage to gcloud storage

* Manual Changes

* Manual Changes

* Manual Changes

* Manual Changes

* Revert "Manual Changes"

This reverts commit 3ef23f1456.

* Manaul Changes

* Changes for 4337

* Fix: Resolved linter errors and formatted notebooks

* Revert "Fix: Resolved linter errors and formatted notebooks"

This reverts commit 7534760fe1.

* Revert "Changes for 4337"

This reverts commit f497568138.

* Changes for 4337

* gcloud to gsutil migration

* removed changes for model garden

* Update model_garden_axolotl_qwen3_finetuning.ipynb

* Update model_garden_axolotl_qwen3_finetuning.ipynb

---------

Co-authored-by: bhandarivijay <bhandarivijay@google.com>
Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-25 10:39:31 -05:00
2a8ad7cdbb Migrate gsutil usage to gcloud storage (#4319)
* Migrate gsutil usage to gcloud storage

* changes for 4319

* changes for 4319

* Apply automated linter fixes

* remove unused import

* removed model_garden changes

* Update model_garden_pytorch_mixtral_peft_tuning.ipynb

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-24 09:41:58 -05:00
3239b301f2 Migrate gsutil usage to gcloud storage (#4315)
* Migrate gsutil usage to gcloud storage

* changes for 4315

* Apply automated linter fixes

* removed model_garden folder changes

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-24 09:41:28 -05:00
acb10d14b8 Migrate gsutil usage to gcloud storage (#4314)
* Migrate gsutil usage to gcloud storage

* changes for 4314

* Apply automated linter fixes

* remove unused imports

* removed model_garden changes

* updates

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-24 09:41:00 -05:00
778d145970 Migrate gsutil usage to gcloud storage (#4313)
* Migrate gsutil usage to gcloud storage

* changes for 4313

* Apply automated linter fixes

* remove model_garden updates

* remove model_garden updates

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-24 09:40:34 -05:00
4d00356f4b Migrate gsutil usage to gcloud storage (#4312)
* Migrate gsutil usage to gcloud storage

* changes for 4312

* Apply automated linter fixes

* remove model_garden folder changes

* remove model_garden folder changes

* Update model_garden_pytorch_llama2_peft_hyperparameter_tuning.ipynb

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-24 09:40:01 -05:00
0654305994 Migrate gsutil usage to gcloud storage (#4311)
* Migrate gsutil usage to gcloud storage

* changes for 4311

* Apply automated linter fixes

* removed unused imports

* remove model_garden changes

* Update model_garden_pytorch_zipnerf.ipynb

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-24 09:39:27 -05:00
c53f392c5a Migrate gsutil usage to gcloud storage (#4310)
* Migrate gsutil usage to gcloud storage

* changes for 4310

* Apply automated linter fixes

* removed model_garden folder changes

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-24 09:38:53 -05:00
f4b56e92ae Migrate gsutil usage to gcloud storage (#4308)
* Migrate gsutil usage to gcloud storage

* changes for 4308

* added = in command

* Apply automated linter fixes

* removed model_garden folder changes

* removed model_garden folder changes

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-24 09:38:24 -05:00
Margubur RahmanGitHubgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gurusai-voletigemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
45ec1cf18a Migrate gsutil usage to gcloud storage (#4307)
* Migrate gsutil usage to gcloud storage

* changes for 4307

* Apply suggestion from @gemini-code-assist[bot]

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Apply suggestion from @gemini-code-assist[bot]

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Apply automated linter fixes

* Update sdk_vector_search_for_indexing.ipynb

* Update sdk_vector_search_for_indexing.ipynb

* update

* Update model_garden_pytorch_falcon_instruct_quantization.ipynb

* Update model_garden_pipeline_templates_t5x.ipynb

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-12-24 14:37:45 +00:00
f6c8bcf937 Migrate gsutil usage to gcloud storage (#4301)
* Migrate gsutil usage to gcloud storage

* changes for 4301

* linter changes

* Revert "linter changes"

This reverts commit 6665133b2c.

* Apply automated linter fixes

* Update training-multi-class-classification-model-for-ads-targeting-usecase.ipynb

* Update training-multi-class-classification-model-for-ads-targeting-usecase.ipynb

* removed model_garden folder changes

* Update model_garden_mediapipe_object_detection.ipynb

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-24 09:37:11 -05:00
Rayan DasoriyaandCopybara-Service 9e590d5a9f Update notebooks based on latest deployment options
PiperOrigin-RevId: 847555297
2025-12-21 19:32:54 -08:00
46e0ea4f1c Migrate gsutil usage to gcloud storage (#4309)
* Migrate gsutil usage to gcloud storage

* changes for 4309

* changes for 4309

* Apply automated linter fixes

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-20 11:34:42 -05:00
300fce6b9f Migrate gsutil usage to gcloud storage (#4305)
* Migrate gsutil usage to gcloud storage

* Apply automated linter fixes

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-20 11:34:10 -05:00
Vertex MG TeamandCopybara-Service 52e3066c38 No public description
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 845978812
2025-12-19 10:42:55 -08:00
27ebf52198 Migrate gsutil usage to gcloud storage (#4296)
* Migrate gsutil usage to gcloud storage

* updates

* update

* Update llm_streaming_prediction.ipynb

* revert change

* removed model_garden folder changes

* removed model_garden folder changes

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-19 11:57:20 -05:00
ca7d4e153e Migrate gsutil usage to gcloud storage (#4297)
* Migrate gsutil usage to gcloud storage

* changes for 4297

* Apply automated linter fixes

* removed changes from model_garden

* removed changes from model_garden

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-19 11:56:49 -05:00
5b6c766629 Migrate gsutil usage to gcloud storage (#4298)
* Migrate gsutil usage to gcloud storage

* update

* linter changes for 4298

* Revert "linter changes for 4298"

This reverts commit c7a00a9710.

* Linter fixes

* update

* update

* Update distributed_hyperparameter_tuning.ipynb

* removed changes in model_garden folder

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-19 11:56:15 -05:00
Margubur RahmanandGitHub 5efa51206f Migrate gsutil usage to gcloud storage (#4324) 2025-12-19 15:11:28 +00:00
Margubur RahmanandGitHub e604a4d43e Migrate gsutil usage to gcloud storage (#4341) 2025-12-19 15:10:52 +00:00
23af5373ec Migrate gsutil usage to gcloud storage (#4295)
* Migrate gsutil usage to gcloud storage

* removed notes

* Apply automated linter fixes

* Revert "Apply automated linter fixes"

This reverts commit 3f6c4c0d36.

* linter changes

* update

* update

* update

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-19 10:09:05 -05:00
7577c0b1fc Migrate gsutil usage to gcloud storage (#4292)
* Migrate gsutil usage to gcloud storage

* Manual change

* update

* Update NotebookProcessors.py

* Update NotebookProcessors.py

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-19 15:07:50 +00:00
b648f9e73b Migrate gsutil usage to gcloud storage (#4306)
* Migrate gsutil usage to gcloud storage

* changes for 4306

* Apply automated linter fixes

* revert gsutil to gcloud

* Update model_garden_pytorch_stable_diffusion_custom.ipynb

* Update model_garden_pytorch_stable_diffusion_xl_lcm.ipynb

* Update model_garden_pytorch_deployed_model_agent_engine.ipynb

* Update model_garden_pytorch_blip_vqa.ipynb

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-19 15:07:21 +00:00
Margubur RahmanandGitHub 966bbc49a7 Migrate gsutil usage to gcloud storage (#4340) 2025-12-18 13:50:35 -05:00
f5d341ae45 Migrate gsutil usage to gcloud storage (#4303)
* Migrate gsutil usage to gcloud storage

* changes for 4303

* Apply automated linter fixes

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-18 13:49:38 -05:00
70770a50c7 Migrate gsutil usage to gcloud storage (#4302)
* Migrate gsutil usage to gcloud storage

* changes for 4302

* changes for 4302

* removed note

* linter changes

* Revert "linter changes"

This reverts commit a9544e8251.

* Apply automated linter fixes

* Update lightweight_functions_component_io_kfp.ipynb

* Update lightweight_functions_component_io_kfp.ipynb

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-18 13:48:39 -05:00
Matej AleksandrovandCopybara-Service f181c39cbf No public description
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 845941350
2025-12-18 10:04:35 -08:00
Rayan DasoriyaandCopybara-Service a5637f87f2 Add notebook for T5Gemma 2 local inference
PiperOrigin-RevId: 846315479
2025-12-18 10:03:26 -08:00
0103299084 Migrate gsutil usage to gcloud storage (#4304)
* Migrate gsutil usage to gcloud storage

* changes for 4304

* Apply automated linter fixes

* remove unused import

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-18 14:00:23 +00:00
1867536d76 Migrate gsutil usage to gcloud storage (#4294)
* Migrate gsutil usage to gcloud storage

* updates

* removed notes added by agent

* removed note

* linter changes

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-18 13:59:10 +00:00
0edae683e7 Migrate gsutil usage to gcloud storage (#4293)
* Migrate gsutil usage to gcloud storage

* update

---------

Co-authored-by: gurusai-voleti <gvoleti@google.com>
2025-12-18 13:58:03 +00:00
Matej AleksandrovandCopybara-Service 8471b5cb6f No public description
PiperOrigin-RevId: 845941350
2025-12-17 15:25:49 -08:00
Aaron DietzandGitHub 87f540ac53 Update notebook_template_review.py (#4405)
Minor change to help us update references to "custom training" to specify "serverless training"
2025-12-17 22:08:00 +00:00
Sam-DecigaandGitHub babeba9f02 feat:NVIDIA Nemotron Nano v2 12B VL - 2025-12-TBD (#4391)
* feat:NVIDIA Nemotron Nano v2 12B VL - 2025-12-TBD

* refactor:Reformat Notebook

* refactor:Reformat Notebook
2025-12-17 18:49:58 +00:00
Vertex MG TeamandCopybara-Service 85c649dd26 Add Llama 3.3 TPU7x deployment notebook.
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 845481841
2025-12-16 16:50:45 -08:00
Vertex MG TeamandCopybara-Service b7135ae1f0 Add Llama 3.3 TPU7x deployment notebook.
PiperOrigin-RevId: 845481841
2025-12-16 16:38:07 -08:00
Vertex MG TeamandCopybara-Service 2f5119a266 Allows the user to select spot VM for deployment
PiperOrigin-RevId: 845064375
2025-12-15 21:23:11 -08:00
Vertex MG TeamandCopybara-Service 23e64ca76f fix: Update TimesFM 2.0 deployment notebook to use GCS path as MODEL_ID
PiperOrigin-RevId: 844826585
2025-12-15 10:25:10 -08:00
gurusai-voletiandGitHub 0ba5a62cc9 Fix ci workflow to use python 3.13 to avoid linter issues (#4397)
* update

* use python 3.13
2025-12-15 13:54:49 +00:00
Damodar PanigrahiandGitHub 0be2c6fd0c feat: authenticate using sa (#4394) 2025-12-12 15:38:25 -05:00
Vertex MG TeamandCopybara-Service cef4928c49 Allows the user to select spot VM for deployment
PiperOrigin-RevId: 843098639
2025-12-11 01:03:35 -08:00
Vertex MG TeamandCopybara-Service 9bb8107110 Add notebook for using Deepseek 3.2 model on Vertex AI.
PiperOrigin-RevId: 842763792
2025-12-10 09:42:39 -08:00
Vertex MG TeamandCopybara-Service b075990d88 Weekly update the vllm/hf-tei/hf-inference-toolkit containers.
PiperOrigin-RevId: 842339077
2025-12-09 12:04:54 -08:00
Ravi DalalGitHubgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
6a83c4c695 Updated notebook comment for custom vllm container image (#4385)
* updated comment for custom vllm container image

* Update notebooks/official/prediction/vertexai_serving_vllm/vertexai_serving_vllm_cpu_llama3_2_3B.ipynb

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

---------

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-12-05 20:26:32 +00:00
Damodar PanigrahiGitHubgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
820c0f8db4 Update Use Case description and API change to accept GCP auth token (#4384)
* pass auth_token in create_http_client

* lint on the notebook

* feat:Removed the last update date

* Update notebooks/community/alphagenome/README.md

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update notebooks/community/alphagenome/cloudai_alphagenome_vai_quickstart.ipynb

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

---------

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-12-05 18:30:23 +00:00
Damodar PanigrahiGitHubgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
4ab197a4ba AlphaGenome GCP API with quickstart.ipynb and README.md (#4378)
* AlphaGenome GCP API  with quickstart.ipynb and README.md

* Update notebooks/community/alphagenome/README.md

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update notebooks/community/alphagenome/README.md

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update notebooks/community/alphagenome/README.md

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update notebooks/community/alphagenome/README.md

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update notebooks/community/alphagenome/README.md

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update notebooks/community/alphagenome/README.md

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Update cloudai_alphagenome_vai_quickstart.ipynb

lint errors

* Update cloudai_alphagenome_vai_quickstart.ipynb

lint errors

* lint errors

* lint errors

* lint import order

* lint errors

* lint import order

* lint import

* Update cloudai_alphagenome_vai_quickstart.ipynb format

* Update cloudai_alphagenome_vai_quickstart.ipynb remove hardcoded url

---------

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-11-27 18:56:07 +00:00
Damodar PanigrahiandGitHub 2a5877fbd1 Update CODEOWNERS (#4380)
* Update CODEOWNERS

* Update CODEOWNERS
2025-11-27 18:30:33 +00:00
Vertex MG TeamandCopybara-Service 090e1d9fee Add dynamic model loading/unloading and instructions to model co-hosting notebook.
PiperOrigin-RevId: 837273935
2025-11-26 15:17:30 -08:00
Eric DongandGitHub 2e049d4830 feat: Add new supported model for Claude (#4377)
* feat: Add new supported model for Claude

* Fix elif

* Remove unused endpoint
2025-11-24 16:56:19 -05:00
Vertex MG TeamandCopybara-Service e3320d2126 Update vLLM docker URI in model co-hosting notebook.
PiperOrigin-RevId: 836250727
2025-11-24 09:14:59 -08:00
Vertex MG TeamandCopybara-Service 5fc0e03ca3 Use separate regions for training, evaluation, and deployment in Llama 3.1 finetuning notebook
PiperOrigin-RevId: 835030738
2025-11-20 20:31:36 -08:00
Vertex MG TeamandCopybara-Service ff2a16237d Minor typo fixes and prints
PiperOrigin-RevId: 834789027
2025-11-20 09:15:31 -08:00
Vertex MG TeamandCopybara-Service d26f081642 Weekly update the vllm/hf-tei/hf-inference-toolkit containers.
PiperOrigin-RevId: 833620797
2025-11-17 20:35:49 -08:00
Vertex MG TeamandCopybara-Service 8a0a39176c Some minor updates and refactoring
PiperOrigin-RevId: 832558814
2025-11-14 20:22:08 -08:00
Vertex MG TeamandCopybara-Service 646532ea69 Fixed the deployment quota check and modified the documentation.
PiperOrigin-RevId: 832174331
2025-11-13 23:14:25 -08:00
Vertex MG TeamandCopybara-Service 9e96a3da67 MiniMax-M2 deployment notebook
PiperOrigin-RevId: 831880572
2025-11-13 08:58:34 -08:00
Vertex MG TeamandCopybara-Service db34e1fbd5 Add multi-model benchmark utility and benchmark results to model co-hosting tutorial notebook.
PiperOrigin-RevId: 831592927
2025-11-12 16:54:24 -08:00
Sam-DecigaandGitHub 82308acbac Update anthropic_claude_3_intro.ipynb - Sonnet 3.7 Deprecation (#4364)
Given information above. Approved.
2025-11-12 12:32:57 -05:00
Vertex MG TeamandCopybara-Service ee0ba75d1e Weekly update the vllm/hf-tei/hf-inference-toolkit containers.
PiperOrigin-RevId: 830644861
2025-11-10 16:31:45 -08:00
Vertex MG TeamandCopybara-Service d61aedc721 Weekly update the vllm/hf-tei/hf-inference-toolkit containers.
PiperOrigin-RevId: 829629618
2025-11-07 17:08:47 -08:00
Vertex MG TeamandCopybara-Service 19f7f94af5 DeepSeek-OCR deployment notebook
PiperOrigin-RevId: 827772392
2025-11-03 21:08:50 -08:00
Vertex MG TeamandCopybara-Service 4e5ce9b226 Add single-model multi-replica & multi-model model co-hosting tutorial notebook.
PiperOrigin-RevId: 826614269
2025-10-31 13:46:14 -07:00
Vertex MG TeamandCopybara-Service 99938244f4 Add DWS to the 8B model in the Eval section
PiperOrigin-RevId: 825943414
2025-10-30 02:42:42 -07:00
Vertex MG TeamandCopybara-Service 0cc7be4a6a Added remote sensing deployment notebook
PiperOrigin-RevId: 825002026
2025-10-28 06:13:52 -07:00
Vertex MG TeamandCopybara-Service b6bde41850 Deepseek deployment v3_2 notebook
PiperOrigin-RevId: 824868883
2025-10-27 23:39:38 -07:00
Vertex MG TeamandCopybara-Service 52444a0933 Weekly update the vllm/hf-tei/hf-inference-toolkit containers.
PiperOrigin-RevId: 824690270
2025-10-27 14:57:44 -07:00
Vertex MG TeamandCopybara-Service bab9c398fd Add new variants to qwen3-vl
PiperOrigin-RevId: 823383168
2025-10-24 00:02:05 -07:00
Vertex MG TeamandCopybara-Service 447affcc93 Weekly update the vllm/hf-tei/hf-inference-toolkit containers.
PiperOrigin-RevId: 823276006
2025-10-23 18:49:58 -07:00
Bhaskar GoyalandGitHub 6132c37be0 Initiate Deprecation for Mistral Large (24.11) and Codestral (25.01) (#4347) 2025-10-23 16:38:09 +00:00
Vertex MG TeamandCopybara-Service 8d7f59aeec Update vLLM TPU deployment container image URI.
PiperOrigin-RevId: 822769821
2025-10-22 15:49:05 -07:00
Vertex MG TeamandCopybara-Service bd327ad424 Qwen3-VL deployment notebook
PiperOrigin-RevId: 822441900
2025-10-21 23:36:55 -07:00
Vertex MG TeamandCopybara-Service 3b1fbdb382 Update image in text+image chat completions requests in MG notebooks.
PiperOrigin-RevId: 821912015
2025-10-20 19:56:47 -07:00
Vertex MG TeamandCopybara-Service 954043a729 No public description
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 821735608
2025-10-20 12:10:35 -07:00
Vertex MG TeamandCopybara-Service 16ef9ee80e Add notebook for deploying GPT OSS models on G4 (RTX Pro 6000).
PiperOrigin-RevId: 821735608
2025-10-20 11:42:04 -07:00
Bhaskar GoyalandGitHub 79301b4a4d <feature> - Add Codestral 2 Model (#4291) 2025-10-16 22:05:58 +00:00
Vertex MG TeamandCopybara-Service 0b38d02e6f Add vLLM TPU deployment notebook for qwen3
PiperOrigin-RevId: 820273514
2025-10-16 09:41:13 -07:00
kthytangandGitHub c52ff25ba4 Haiku 4.5 update to anthropic_claude_intro.ipynb (#4300)
* Haiku 4.5 update to anthropic_claude_intro.ipynb

* Update anthropic_claude_intro.ipynb
2025-10-15 20:53:05 +00:00
Vertex MG TeamandCopybara-Service 424400bace Weekly update the vllm/hf-tei/hf-inference-toolkit containers.
PiperOrigin-RevId: 819784960
2025-10-15 09:16:29 -07:00
Mend RenovateandGitHub b1dfac2043 chore(deps): update actions/setup-python action to v6 (#4245) 2025-10-10 14:32:20 +00:00
Mend RenovateandGitHub b81ffcddab chore(deps): update python docker tag to v3.14 (#4284) 2025-10-10 14:29:19 +00:00
Mend RenovateandGitHub 21d8f144aa chore(deps): update dependency pyupgrade to v3.21.0 (#4287) 2025-10-10 14:29:11 +00:00
Vertex MG TeamandCopybara-Service 065a674305 Minor fixes for ollama deployment notebook
PiperOrigin-RevId: 817513895
2025-10-10 00:21:23 -07:00
haomengchaoandGitHub 571d498d08 feat: add notebook for VirtueAI model in Model Garden (#4280)
* feat: add virtueai's notebook for Model Garden

* feat: add virtueai's notebook for Model Garden with fixes

* feat: fix endpoint place holder to pass the test

* feat: fix endpoint place holder to pass the test

* feat: fix endpoint place holder to pass the test

* feat: fix a typo
2025-10-09 22:15:24 +00:00
Bhaskar GoyalandGitHub e936882123 <feature>: Medium 3 launch (#4279) 2025-10-09 16:42:11 +00:00
Vertex MG TeamandCopybara-Service f754f99052 Increase the dws max_wait_duration to 90 minutes
PiperOrigin-RevId: 817042685
2025-10-09 00:11:15 -07:00
Vertex MG TeamandCopybara-Service aa5523a5e9 Delete Gemma 3 peft finetuning notebooks.
PiperOrigin-RevId: 816744576
2025-10-08 09:39:54 -07:00
Vertex MG TeamandCopybara-Service 81393ede1a Weekly update the hf-tei/hf-inference-toolkit containers.
PiperOrigin-RevId: 815926264
2025-10-06 16:17:53 -07:00
Vertex MG TeamandCopybara-Service 5a1c0222da Qwen Image Deployment Notebook
PiperOrigin-RevId: 815781924
2025-10-06 10:19:30 -07:00
Vertex MG TeamandCopybara-Service 1f9326bd56 Weekly update the vllm container version.
PiperOrigin-RevId: 814830089
2025-10-03 14:24:00 -07:00
Vertex MG TeamandCopybara-Service a94cae2e79 Weekly update the vllm/hf-tei/hf-inference-toolkit containers.
PiperOrigin-RevId: 813495515
2025-09-30 17:28:38 -07:00
kthytangandGitHub cb861713c8 Update anthropic_claude_intro.ipynb (#4277) 2025-09-29 17:17:00 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
3a55087789 Bump urllib3 (#4257)
Bumps [urllib3](https://github.com/urllib3/urllib3) from 2.4.0 to 2.5.0.
- [Release notes](https://github.com/urllib3/urllib3/releases)
- [Changelog](https://github.com/urllib3/urllib3/blob/main/CHANGES.rst)
- [Commits](https://github.com/urllib3/urllib3/compare/2.4.0...2.5.0)

---
updated-dependencies:
- dependency-name: urllib3
  dependency-version: 2.5.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2025-09-29 16:59:52 +00:00
Vertex MG TeamandCopybara-Service d359b21f3e Weekly update the vllm/hf-tei/hf-inference-toolkit containers.
PiperOrigin-RevId: 811502904
2025-09-25 14:30:34 -07:00
Vertex MG TeamandCopybara-Service c7d4123b25 Print project and region information
PiperOrigin-RevId: 810501296
2025-09-23 10:54:22 -07:00
Vertex MG TeamandCopybara-Service 07a8bb2d0c Migrate Phi-4 notebook to use Model Garden SDK
PiperOrigin-RevId: 810268843
2025-09-22 20:52:05 -07:00
Vertex MG TeamandCopybara-Service f6b6f365b6 Migrate Ollama deploy notebook to use Model Garden SDK
PiperOrigin-RevId: 810150526
2025-09-22 14:13:49 -07:00
Vertex MG TeamandCopybara-Service fec825f9e5 Use dictionary for the deletion of multiple endpoints
PiperOrigin-RevId: 810066401
2025-09-22 10:26:45 -07:00
Vertex MG TeamandCopybara-Service 1e9bf72097 Blip2 notebook refactoring
PiperOrigin-RevId: 810050386
2025-09-22 09:47:46 -07:00
Vertex MG TeamandCopybara-Service 3fa1cf99ec Make 4b as the default model_version in the Gemma3 notebook
PiperOrigin-RevId: 809821706
2025-09-21 19:09:19 -07:00
Vertex MG TeamandCopybara-Service 2bcaf8abde Allow the user to enter Region
PiperOrigin-RevId: 808435945
2025-09-18 00:17:02 -07:00
Vertex MG TeamandCopybara-Service 759495a1f8 Migrate LaMa notebook to use Model Garden SDK
PiperOrigin-RevId: 808001560
2025-09-16 23:19:01 -07:00
Vertex MG TeamandCopybara-Service 9925e62c4b feat: Refactor to use deploy SDK.
PiperOrigin-RevId: 807803299
2025-09-16 12:35:20 -07:00
Vertex MG TeamandCopybara-Service 40ade71b35 Weekly update the vllm serving container version to 20250911_0916_RC01.
PiperOrigin-RevId: 807413330
2025-09-15 15:45:51 -07:00
Vertex MG TeamandCopybara-Service 65173071a5 Weekly update the serving container version for hf-inference-toolkit and hf-tei containers.
PiperOrigin-RevId: 807413263
2025-09-15 15:44:23 -07:00
Vertex MG TeamandCopybara-Service a514bb51c2 Allow the user to enter Region
PiperOrigin-RevId: 807182772
2025-09-15 04:21:01 -07:00
Rayan DasoriyaandCopybara-Service a6dc1b0f6d Fix notebook issues
PiperOrigin-RevId: 806407972
2025-09-12 13:33:39 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
75236998c7 Bump torch (#4239)
Bumps [torch](https://github.com/pytorch/pytorch) from 2.2.0 to 2.8.0.
- [Release notes](https://github.com/pytorch/pytorch/releases)
- [Changelog](https://github.com/pytorch/pytorch/blob/main/RELEASE.md)
- [Commits](https://github.com/pytorch/pytorch/compare/v2.2.0...v2.8.0)

---
updated-dependencies:
- dependency-name: torch
  dependency-version: 2.8.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2025-09-10 23:33:41 +00:00
MarkandGitHub 9c8a7808bf chore: Replace 'prediction' with 'inference' per urgent rebranding request (#4231) 2025-09-10 23:32:32 +00:00
Holt SkinnerandGitHub 86ce1576d2 Delete notebooks/official/model_evaluation/automl_video_classification_model_evaluation.ipynb (#4256) 2025-09-10 23:32:08 +00:00
Holt Skinner c2b743bfeb Removed Deprecated notebooks 2025-09-10 18:31:40 -05:00
Vertex MG TeamandCopybara-Service b04575d746 Fix axolotl gcs output path.
PiperOrigin-RevId: 805417425
2025-09-10 10:26:55 -07:00
Vertex MG TeamandCopybara-Service b90d16885e Weekly update the serving container version for hf-inference-toolkit and hf-tei containers.
PiperOrigin-RevId: 805081873
2025-09-09 15:12:41 -07:00
Holt SkinnerandGitHub 828f1a26fa Delete notebooks/official/pipelines/google_cloud_pipeline_components_automl_text.ipynb (#4254)
b/442906902
2025-09-09 20:43:03 +00:00
Holt SkinnerandGitHub 915e5edf0a chore: Remove AutoML Notebooks for deprecated Text and Video features (#4253)
* chore: Remove AutoML Notebooks for deprecated Text and Video features

* Remove remaining Text/Video samples
2025-09-09 16:49:06 +00:00
Vertex MG TeamandCopybara-Service 5d4c1be285 Migrate llava notebook to use Model Garden SDK
PiperOrigin-RevId: 804765256
2025-09-09 00:16:12 -07:00
Vertex MG TeamandCopybara-Service 275bfb8f69 Weekly update the vllm serving container version to 20250905_0916_RC01.
PiperOrigin-RevId: 804514266
2025-09-08 11:21:48 -07:00
Vertex MG TeamandCopybara-Service 93cb0ca3d9 Migrate blip image captioning notebook to use Model Garden SDK
PiperOrigin-RevId: 803091009
2025-09-04 10:49:28 -07:00
Vertex MG TeamandCopybara-Service 41d60d052d Add EmbeddingGemma local inference notebook.
PiperOrigin-RevId: 803038284
2025-09-04 08:30:34 -07:00
Vertex MG TeamandCopybara-Service d872cdcf7e Migrate Paligemma2 notebook to use Model Garden SDK
PiperOrigin-RevId: 802861534
2025-09-03 22:35:04 -07:00
Vertex MG TeamandCopybara-Service ada5e4a854 chore: fix google-auth and requests package version.
PiperOrigin-RevId: 802270767
2025-09-02 13:33:33 -07:00
Vertex MG TeamandCopybara-Service ecb32b099d Migrate Stable Diffusion Upscaler notebook to use Model Garden SDK
PiperOrigin-RevId: 801824204
2025-09-01 08:55:58 -07:00
Vertex MG TeamandCopybara-Service f8d09e8e9b Migrate Stable Diffusion XL Lightning notebook to use Model Garden SDK
PiperOrigin-RevId: 801782973
2025-09-01 06:06:41 -07:00
Vertex MG TeamandCopybara-Service 007df88fba Migrate Llama4 notebook to use Model Garden SDK
PiperOrigin-RevId: 801662862
2025-08-31 22:27:06 -07:00
Vertex MG TeamandCopybara-Service c70f3ef9a3 Weekly update vllm/hf-tei/hf-inference-toolkit container image versions.
PiperOrigin-RevId: 801034279
2025-08-29 14:39:29 -07:00
Vertex MG TeamandCopybara-Service 505e101452 Bug fix for Wan2.2
PiperOrigin-RevId: 800419261
2025-08-28 05:11:01 -07:00
Vertex MG TeamandCopybara-Service b77d51b58b Migrate QWEN3 notebook to use Model Garden SDK
PiperOrigin-RevId: 800346334
2025-08-28 00:57:59 -07:00
Rayan DasoriyaandCopybara-Service deaa1ccf2b Update quota check to use gcloud beta quotas info describe command
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 800261134
2025-08-27 19:19:54 -07:00
Vertex MG TeamandCopybara-Service f6e38860aa Update H100/H200 region recommendations in DeepSeek deployment notebook.
PiperOrigin-RevId: 799800908
2025-08-26 18:29:57 -07:00
Vertex MG TeamandCopybara-Service 4937e382b1 Bug fix for Wan2.1
PiperOrigin-RevId: 799116301
2025-08-25 07:40:20 -07:00
Vertex MG TeamandCopybara-Service 0fe2770947 Weekly update the vLLM container image version.
PiperOrigin-RevId: 798391915
2025-08-22 16:59:25 -07:00
Vertex MG TeamandCopybara-Service 80d7ee67d6 Weekly update the container version for hf-inference-toolkit and hf-tei.
PiperOrigin-RevId: 797920954
2025-08-21 14:45:09 -07:00
Vertex MG TeamandCopybara-Service f17e2d6c8d Migrate gpt-oss deploy notebook to use Model Garden SDK
PiperOrigin-RevId: 797824747
2025-08-21 10:40:45 -07:00
Vertex MG TeamandCopybara-Service ae1cd0ed08 Migrate llama3.3 deployment notebook to use Model Garden SDK
PiperOrigin-RevId: 797622523
2025-08-20 23:34:08 -07:00
denisj3030andGitHub a68b491edc dep35 (#4227) 2025-08-20 20:02:14 +00:00
Vertex MG TeamandCopybara-Service d39bed012a Migrate Controlnet notebook to use Model Garden SDK
PiperOrigin-RevId: 797349160
2025-08-20 09:37:09 -07:00
Vertex MG TeamandCopybara-Service f8c93de1c4 Migrate BLIP2 deploy notebook to use Model Garden SDK
PiperOrigin-RevId: 797131322
2025-08-19 20:34:06 -07:00
Vertex MG TeamandCopybara-Service 0b5d97aa3a use A100 machine by default for Llama3.1 fast deployment
PiperOrigin-RevId: 796788128
2025-08-19 02:51:44 -07:00
Vertex MG TeamandCopybara-Service 87f0785ed1 Migrate Segment Anything Model(SAM) notebook to use Model Garden SDK
PiperOrigin-RevId: 796693389
2025-08-18 20:55:11 -07:00
Vertex MG TeamandCopybara-Service 09c8ea616b Weekly update the vLLM container image version.
PiperOrigin-RevId: 796581324
2025-08-18 14:40:14 -07:00
Vertex MG TeamandCopybara-Service 6c11241cf7 Migrate BLIP2 VQA(Visual Question Answering) notebook to use Model Garden SDK
PiperOrigin-RevId: 796316856
2025-08-18 01:35:17 -07:00
Vertex MG TeamandCopybara-Service 9b7dc2e4fd Migrate BLIP2 VQA(Visual Question Answering) notebook to use Model Garden SDK
PiperOrigin-RevId: 795738195
2025-08-15 22:43:11 -07:00
Vertex MG TeamandCopybara-Service 9b3d43b1a6 Use correct notebook_util.
PiperOrigin-RevId: 794874566
2025-08-13 22:18:11 -07:00
Dustin LuongandCopybara-Service 57ee4e2eab Update Qwen3 deployment notebook with new variants and corrected model names.
PiperOrigin-RevId: 794407429
2025-08-12 22:41:45 -07:00
Dustin LuongandCopybara-Service 4d736d7992 Use correct notebook_util.
PiperOrigin-RevId: 794396544
2025-08-12 22:00:22 -07:00
Dustin LuongandCopybara-Service 69e92a650d Add gpt-oss-20b finetuning with lora on Vertex notebook.
PiperOrigin-RevId: 794317611
2025-08-12 16:57:36 -07:00
Vertex MG TeamandCopybara-Service 4ed979eec4 Migrate Stable Diffusion XL 1.0 notebook to use Model Garden SDK
PiperOrigin-RevId: 793954002
2025-08-11 23:22:40 -07:00
Vertex MG TeamandCopybara-Service 0344de8090 Wan Deployment Notebook
PiperOrigin-RevId: 793839574
2025-08-11 16:30:15 -07:00
Vertex MG TeamandCopybara-Service 83eca0f0cc Update the vLLM container image version.
PiperOrigin-RevId: 793738684
2025-08-11 11:47:34 -07:00
Vertex MG TeamandCopybara-Service 4619b272fa Migrate QwQ notebook to use Model Garden SDK.
PiperOrigin-RevId: 793694751
2025-08-11 10:05:17 -07:00
Vertex MG TeamandCopybara-Service 091fa33c01 Migrate Gemma3 notebook to use Model Garden SDK.
PiperOrigin-RevId: 793692575
2025-08-11 10:01:02 -07:00
Rayan DasoriyaandCopybara-Service 42a05c2b35 No public description
PiperOrigin-RevId: 793644258
2025-08-11 07:50:06 -07:00
Vertex MG TeamandCopybara-Service 8ac32fa42f Migrate Gemma3n notebook to use Model Garden SDK
PiperOrigin-RevId: 793583308
2025-08-11 04:09:41 -07:00
Rayan DasoriyaandCopybara-Service 0303057f11 Update the common util location in the notebooks
PiperOrigin-RevId: 793425712
2025-08-10 18:26:31 -07:00
Ravi DalalandGitHub 7ae13b346a added notebooks and dockerfiles for serving open models on vertexai using vllm custom containers (#4148)
* added notebooks and dockerfiles for serving open models on vertexai using vllm customer containers

* updated official codeowners

* fixed linting errors

* fixed linting errors

* moved notebooks

* ran linter

* fixed param type

* added some formatting

* added autoscaling configuration to model deployment

* fixed a heading

* moved notebooks and docker folder under prediction

* updated notebook repo paths

* switched to raw_predict to avoid code changes and rebuild

* removed linting errors

* fixed readme lint error

* fixed links

* removed dedicated_endpoint_enabled

* updated workdir path

* fixed cell type

* added license

* fixed linting error

* using cloud build container image build

* fixed linting issues

* fixes

* cloudbuild yaml

* fixed gemini review comments

* fixed linting errors

* fixed tpu_count type

* handled invalid device type

* optimized dockerfile run command

* updated dockerfile

* addressed review comments

* fixed linting errors

* added license to cloudbuild and dockerfile

* optimized image build code

* fixed image_name variable
2025-08-08 18:01:03 +00:00
Rayan DasoriyaandCopybara-Service 44390cbd99 Add a no-op message to check_quota if it fails.
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 792654121
2025-08-08 09:32:08 -07:00
Vertex MG TeamandCopybara-Service 58dab0b1bb E5 notebook
PiperOrigin-RevId: 791999324
2025-08-06 22:48:50 -07:00
Vertex MG TeamandCopybara-Service 8d36834fcd Add GPT OSS models deployment notebook.
PiperOrigin-RevId: 791864273
2025-08-06 15:17:54 -07:00
Dustin LuongandCopybara-Service 31ca3e36f4 Add Qwen3-30B-A3B instruct and thinking 2507 variants to notebook.
PiperOrigin-RevId: 791746627
2025-08-06 10:27:39 -07:00
Vertex MG TeamandCopybara-Service a72d7dc49f Fix custom dataset input for axolotl notebooks.
PiperOrigin-RevId: 791627090
2025-08-06 04:16:45 -07:00
denisj3030andGitHub 9efbd48233 cl41 (#4193) 2025-08-05 17:02:14 +00:00
Vertex MG TeamandCopybara-Service 166f0f8ce7 Update the vLLM container image version.
PiperOrigin-RevId: 790903846
2025-08-04 14:56:59 -07:00
Vertex MG TeamandCopybara-Service bd9f9675cf Add us-south1 region
PiperOrigin-RevId: 790762290
2025-08-04 08:35:53 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
42ec0e6ad9 Bump urllib3 (#4173)
Bumps [urllib3](https://github.com/urllib3/urllib3) from 2.0.7 to 2.5.0.
- [Release notes](https://github.com/urllib3/urllib3/releases)
- [Changelog](https://github.com/urllib3/urllib3/blob/main/CHANGES.rst)
- [Commits](https://github.com/urllib3/urllib3/compare/2.0.7...2.5.0)

---
updated-dependencies:
- dependency-name: urllib3
  dependency-version: 2.5.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2025-08-01 14:29:14 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
fde3d98a2a Bump requests (#4174)
Bumps [requests](https://github.com/psf/requests) from 2.32.3 to 2.32.4.
- [Release notes](https://github.com/psf/requests/releases)
- [Changelog](https://github.com/psf/requests/blob/main/HISTORY.md)
- [Commits](https://github.com/psf/requests/compare/v2.32.3...v2.32.4)

---
updated-dependencies:
- dependency-name: requests
  dependency-version: 2.32.4
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2025-08-01 14:28:53 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
b83f869a44 Bump protobuf (#4175)
Bumps [protobuf](https://github.com/protocolbuffers/protobuf) from 3.20.3 to 4.25.8.
- [Release notes](https://github.com/protocolbuffers/protobuf/releases)
- [Changelog](https://github.com/protocolbuffers/protobuf/blob/main/protobuf_release.bzl)
- [Commits](https://github.com/protocolbuffers/protobuf/compare/v3.20.3...v4.25.8)

---
updated-dependencies:
- dependency-name: protobuf
  dependency-version: 4.25.8
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2025-08-01 14:28:28 +00:00
Vertex MG TeamandCopybara-Service b918768776 add batch prediction
PiperOrigin-RevId: 789643160
2025-08-01 00:21:30 -07:00
Eric DongandGitHub a22aceb10c refactor: Batch update icon links for notebooks/community (#4188) 2025-07-31 15:41:21 +00:00
Dustin LuongandCopybara-Service c8b53a614e Fix text in TEI notebook to change instances of nomic-ai/nomic-embed-text-v1 to Qwen/Qwen3-Embedding-8B.
PiperOrigin-RevId: 789204281
2025-07-30 23:26:39 -07:00
Eric DongandGitHub 7be7d5be44 refactor: Batch update icon links for notebooks/official (#4187) 2025-07-31 00:08:25 +00:00
Eric DongandGitHub 02a030d6f1 fix: Update the icon links (#4186) 2025-07-30 20:37:13 +00:00
Dustin LuongandCopybara-Service fc8e9e4483 Set Qwen/Qwen3-Embedding-8B as example model in TEI notebook.
PiperOrigin-RevId: 788960826
2025-07-30 10:48:47 -07:00
Eliot LaidlawandGitHub fa019e051a Add CSM intro notebook (#4178)
* notebook

* linting

* add codeowner

* Add cleanup

* fixes

* Update copyright year
2025-07-30 16:36:51 +00:00
Dustin LuongandCopybara-Service 01d513eea3 Update the deployment notebook for Qwen3 models to include Qwen3-235B-Thinking-2507 and Qwen3-235B-Thinking-2507-FP8.
PiperOrigin-RevId: 788636454
2025-07-29 15:45:43 -07:00
Vertex MG TeamandCopybara-Service 928822cc0c Remove 8-bit mode for Qwen 2.5 finetuning notebook.
PiperOrigin-RevId: 788499530
2025-07-29 10:02:42 -07:00
Vertex MG TeamandCopybara-Service 1dfb4091ea notebook reformatting
PiperOrigin-RevId: 788430948
2025-07-29 06:34:26 -07:00
Dustin LuongandCopybara-Service ffb5ea7e88 Update HF TEI serving image to use new, FEDRamp compliant container.
PiperOrigin-RevId: 788203719
2025-07-28 16:36:17 -07:00
Vertex MG TeamandCopybara-Service dd2028d76c add batch prediction
PiperOrigin-RevId: 787943651
2025-07-28 04:02:27 -07:00
Vertex MG TeamandCopybara-Service 70daf2e605 Add Qwen3-Coder notebook
PiperOrigin-RevId: 787286540
2025-07-25 16:57:44 -07:00
Vertex MG TeamandCopybara-Service 42db3643d0 Fix pad tokens for custom weights.
MG_DOCKER_CODES_PIPER_ORIGIN_REV_ID: 787196797
2025-07-25 12:14:37 -07:00
Erwin HuizengaandGitHub 66ce2fa5b7 feat: Add sample for Vertex distributed training (#4163)
* feat: Add sample for Vertex distributed training

* refactor: Move distributed training to community content and add job config

* fix: Address review comments and update files

* minor fixes in the script

* updated codeowners
2025-07-23 21:15:09 +02:00
Dustin LuongandCopybara-Service 44a63c8186 Add Qwen3-235B-A22B-Instruct-2507 models to Qwen3 deployment notebook.
PiperOrigin-RevId: 786375387
2025-07-23 12:12:28 -07:00
Vertex MG TeamandCopybara-Service a29f7a376b Wan Deployment Notebook
PiperOrigin-RevId: 786069875
2025-07-22 18:05:53 -07:00
kittyabsandGitHub 66667ea1db Update ray_cluster_management.ipynb (#4171)
Updated to the current version of Ray supported.
2025-07-22 21:40:52 +00:00
Vertex MG TeamandCopybara-Service 0588a7b62d dedicated endpoints enabled
PiperOrigin-RevId: 785721997
2025-07-21 23:25:34 -07:00
Vertex MG TeamandCopybara-Service ba74664aa4 Update the hf-pytorch-inference notebook using the newly built in-house container hf-inference-toolkit.
PiperOrigin-RevId: 785559005
2025-07-21 13:42:01 -07:00
Vertex MG TeamandCopybara-Service 26011283d2 Updates to Flux.1 Schnell and CogVideoX-2b notebook
PiperOrigin-RevId: 785477091
2025-07-21 10:02:34 -07:00
855 changed files with 66885 additions and 52429 deletions
@@ -365,7 +365,7 @@ def process_and_execute_notebook(
# Use gcloud to get tail
try:
result.error_message = subprocess.check_output(
["gsutil", "cat", "-r", "-1000", log_file_uri], encoding="UTF-8"
["gcloud", "storage", "cat", "--range", "-1000", log_file_uri], encoding="UTF-8"
)
except Exception as error:
result.error_message = str(error)
+2 -2
View File
@@ -56,8 +56,8 @@ def execute_notebook(
print("\n=== DOWNLOAD EXECUTED NOTEBOOK ===\n")
print(f"Please debug the executed notebook by downloading the executed notebook:")
print("Option 1. Using gsutil. Run the following command in your terminal.")
print(f'\tgsutil cp "{output_file_or_uri}" .')
print("Option 1. Using gcloud storage. Run the following command in your terminal.")
print(f'\tgcloud storage cp "{output_file_or_uri}" .')
print("Option 2. Using this link.")
print(f"\thttps://storage.googleapis.com/{output_file_or_uri[5:]}")
+1 -1
View File
@@ -108,7 +108,7 @@ class VertexAIInstallProprocessor(Preprocessor):
if "google-cloud-aiplatform" not in content:
return content
return (
f"gsutil cp {self.vertex_ai_wheel} google-cloud-aiplatform.whl\n" +
f"gcloud storage cp {self.vertex_ai_wheel} google-cloud-aiplatform.whl\n" +
content.replace("google-cloud-aiplatform\n", "google-cloud-aiplatform.whl\n")
.replace("google-cloud-aiplatform ", "google-cloud-aiplatform.whl ")
)
+2 -2
View File
@@ -15,7 +15,7 @@ def download_file(bucket_name: str, blob_name: str, destination_file: str) -> st
remote_file_path = "".join(["gs://", "/".join([bucket_name, blob_name])])
subprocess.check_output(
["gsutil", "cp", remote_file_path, destination_file], encoding="UTF-8"
["gcloud", "storage", "cp", remote_file_path, destination_file], encoding="UTF-8"
)
return destination_file
@@ -27,7 +27,7 @@ def upload_file(
) -> str:
"""Copies a local file to a GCS path"""
subprocess.check_output(
["gsutil", "cp", local_file_path, remote_file_path], encoding="UTF-8"
["gcloud", "storage", "cp", local_file_path, remote_file_path], encoding="UTF-8"
)
return remote_file_path
+3 -3
View File
@@ -7,11 +7,11 @@ jobs:
runs-on: ubuntu-latest
steps:
- name: Set up Python
uses: actions/setup-python@v5
uses: actions/setup-python@v6
with:
python-version: '3.x'
python-version: '3.12'
- name: Fetch pull request branch
uses: actions/checkout@v4
uses: actions/checkout@v6
with:
fetch-depth: 0
- name: Fetch base main branch
+1 -1
View File
@@ -4,7 +4,7 @@
# 2. To lint specific notebooks:
# docker run -v ${PWD}:/setup/app gcr.io/python-docs-samples-tests/notebook_linter:latest notebooks/1.ipynb notebooks/2.ipynb
FROM python:3.13
FROM python:3.14
WORKDIR setup
+3 -3
View File
@@ -2,9 +2,9 @@ git+https://github.com/tensorflow/docs
ipython
jupyter
nbconvert
black==25.1.0
pyupgrade==3.20.0
isort==6.0.1
black==26.5.1
pyupgrade==3.21.2
isort==8.0.1
flake8==7.3.0
nbqa==1.9.1
+25 -9
View File
@@ -1,12 +1,12 @@
# ![Google Cloud](https://avatars.githubusercontent.com/u/2810941?s=60&v=4) Google Cloud Vertex AI Samples
This repository contains notebooks, code samples, sample apps, and other resources that demonstrate how to use, develop and manage machine learning and generative AI workflows using Google Cloud Vertex AI.
This repository contains notebooks, code samples, sample apps, skills, and other resources that demonstrate how to use, develop and manage machine learning and generative AI workflows using Google Cloud Vertex AI.
## Overview
[Vertex AI](https://cloud.google.com/vertex-ai) is a fully-managed, unified AI development platform for building and using generative AI. This repository is designed to help you get started with Vertex AI. Whether you're new to Vertex AI or an experienced ML practitioner, you'll find valuable resources here.
For more Vertex AI Generative AI notebook samples, please visit the Vertex AI [Generative AI](https://github.com/GoogleCloudPlatform/generative-ai) GitHub repository.
⚠️ For more Vertex AI Generative AI notebook samples, please visit the Vertex AI [Generative AI](https://github.com/GoogleCloudPlatform/generative-ai) GitHub repository.
## Explore, learn and contribute
@@ -16,11 +16,11 @@ You can explore, learn, and contribute to this repository to unleash the full po
Explore this repository, follow the links in the header section of each of the notebooks to -
![Colab](https://cloud.google.com/ml-engine/images/colab-logo-32px.png) Open and run the notebook in [Colab](https://colab.google/)\
![Colab Enterprise](https://cloud.google.com/ml-engine/images/colab-enterprise-logo-32px.png) Open and run the notebook in [Colab Enterprise](https://cloud.google.com/colab/docs/introduction)\
![Workbench](https://lh3.googleusercontent.com/UiNooY4LUgW_oTvpsNhPpQzsstV5W8F7rYgxgGBD85cWJoLmrOzhVs_ksK_vgx40SHs7jCqkTkCk=e14-rj-sc0xffffff-h130-w32) Open and run the notebook in [Vertex AI Workbench](https://cloud.google.com/vertex-ai/docs/workbench/introduction)\
![Github](https://cloud.google.com/ml-engine/images/github-logo-32px.png) View the notebook on Github
- Open and run the notebook in [Colab](https://colab.google/)
- Open and run the notebook in [Colab Enterprise](https://cloud.google.com/colab/docs/introduction)
- Open and run the notebook in [Vertex AI Workbench](https://cloud.google.com/vertex-ai/docs/workbench/introduction)
- View the notebook on Github
### Contribute
See the [Contributing Guide](https://github.com/GoogleCloudPlatform/vertex-ai-samples/blob/master/CONTRIBUTING.md).
@@ -35,7 +35,7 @@ To get started using Vertex AI, you must have a Google Cloud project.
## Repository structure
```bash
```text
├── notebooks
│ ├── official - Notebooks demonstrating use of each Vertex AI service
│ │ ├── automl
@@ -45,7 +45,23 @@ To get started using Vertex AI, you must have a Google Cloud project.
│ │ ├── model_garden
│ │ ├── ...
├── community-content - Sample code and tutorials contributed by the community
├── docs - Deep-dive documentation and advanced setup guides
└── skills - Suite of AI Agent "Skills" for Vertex AI
├── README.md # Developer guide for Vertex AI skills
├── vertex-ai/ # Primary router for Vertex AI tasks
│ └── SKILL.md # Entry point that routes across capabilities
├── genai-sdk/ # Gemini API usage with Gen AI SDK
│ └── SKILL.md # Guides for Python, JS/TS, Go, Java, C#
├── vertex-deploy/ # Deploying models to Endpoints
│ └── SKILL.md # Commands for open models & custom weights
├── vertex-inference/ # Inferencing with GenAI models
│ └── SKILL.md # Code samples for Gemini and OpenMaaS
└── vertex-tuning/ # Secondary router for model fine-tuning
├── SKILL.md # Router for tuning tasks
├── gemini/ # Fine-tuning first-party Gemini models
│ └── SKILL.md
└── open-model/ # Fine-tuning third-party open models
└── SKILL.md
```
## Examples
+1
View File
@@ -29,4 +29,5 @@
/vertex_model_garden/model_oss/vllm @kathyyu-google
/vertex_model_garden/benchmarking_reports @lavraicse
/vertex_model_garden/model_oss/autogluon @lavraicse
/vertex_distributed_training/a3mega/llama-3-8b-nemo-pretraining @mstyer-google @erwinh85 @mchrestkha
@@ -148,7 +148,7 @@ implementation:
# Downloading the model archive from GCS
# TODO: Fix gsutil bugs (requires project ID, has auth issues) and use gsutil instead.
# gsutil cp "$model_archive_uri" "$model_archive_local_path"
# gcloud storage cp "$model_archive_uri" "$model_archive_local_path"
pip install google-cloud-storage
python -c '
import sys
@@ -24,12 +24,12 @@ implementation:
# Checking whether the URI points to a single blob, a directory or a URI pattern
# URI points to a blob when that URI does not end with slash and listing that URI only yields the same URI
if [[ "$uri" != */ ]] && (gsutil ls "$uri" | grep --fixed-strings --line-regexp "$uri"); then
if [[ "$uri" != */ ]] && (gcloud storage ls "$uri" | grep --fixed-strings --line-regexp "$uri"); then
mkdir -p "$(dirname "$output_path")"
gsutil -m cp -r "$uri" "$output_path"
gcloud storage cp --recursive "$uri" "$output_path"
else
mkdir -p "$output_path" # When source path is a directory, gsutil requires the destination to also be a directory
gsutil -m rsync -r "$uri" "$output_path" # gsutil cp has different path handling than Linux cp. It always puts the source directory (name) inside the destination directory. gsutil rsync does not have that problem.
gcloud storage rsync --recursive "$uri" "$output_path" # gsutil cp has different path handling than Linux cp. It always puts the source directory (name) inside the destination directory. gsutil rsync does not have that problem.
fi
- inputValue: GCS path
- outputPath: Data
@@ -1,3 +1,3 @@
torch==2.2.0
torch==2.13.0
torchvision==0.9.1
tensorboard==2.5.0
@@ -1,3 +1,3 @@
torch==2.7.0
torch==2.13.0
torchvision==0.9.1
tensorboard==2.5.0
@@ -110,7 +110,7 @@
},
"outputs": [],
"source": [
"! gsutil ls $gcs_output_uri_prefix"
"! gcloud storage ls $gcs_output_uri_prefix"
]
},
{
@@ -192,7 +192,7 @@
},
"outputs": [],
"source": [
"! gsutil cp -r $gcs_output_uri_prefix/model ./model_server/"
"! gcloud storage cp --recursive $gcs_output_uri_prefix/model ./model_server/"
]
},
{
@@ -556,7 +556,7 @@
},
"outputs": [],
"source": [
"! gsutil rm -rf $gcs_output_uri_prefix"
"! gcloud storage rm --recursive --continue-on-error $gcs_output_uri_prefix"
]
},
{
@@ -412,7 +412,7 @@
},
"outputs": [],
"source": [
"! gsutil ls $gcs_output_uri_prefix"
"! gcloud storage ls $gcs_output_uri_prefix"
]
}
],
@@ -77,4 +77,4 @@ echo "After the job is completed successfully, model files will be saved at $JOB
# # Verify the model was exported
# echo "Verify the model was exported:"
# gsutil ls ${JOB_DIR}/
# gcloud storage ls ${JOB_DIR}/
@@ -34,4 +34,4 @@ RUN echo "service_envelope=json\n" "inference_address=http://0.0.0.0:${AIP_H
USER model-server
# run Torchserve HTTP serve to respond to prediction requests
CMD ["echo", "AIP_STORAGE_URI=${AIP_STORAGE_URI}", ";", "gsutil", "cp", "-r", "${AIP_STORAGE_URI}/${MODEL_NAME}.mar", "/home/model-server/model-store/", ";", "ls", "-ltr", "/home/model-server/model-store/", ";", "torchserve", "--start", "--ts-config=/home/model-server/config.properties", "--models", "${MODEL_NAME}=${MODEL_NAME}.mar", "--model-store", "/home/model-server/model-store"]
CMD ["echo", "AIP_STORAGE_URI=${AIP_STORAGE_URI}", ";", "gcloud", "storage", "cp", "--recursive", "${AIP_STORAGE_URI}/${MODEL_NAME}.mar", "/home/model-server/model-store/", ";", "ls", "-ltr", "/home/model-server/model-store/", ";", "torchserve", "--start", "--ts-config=/home/model-server/config.properties", "--models", "${MODEL_NAME}=${MODEL_NAME}.mar", "--model-store", "/home/model-server/model-store"]
@@ -67,4 +67,4 @@ echo "After the job is completed successfully, model files will be saved at $JOB
# # Verify the model was exported
# echo "Verify the model was exported:"
# gsutil ls ${JOB_DIR}/
# gcloud storage ls ${JOB_DIR}/
@@ -478,8 +478,7 @@
},
"outputs": [],
"source": [
"! gsutil mb -l $REGION $BUCKET_NAME"
]
"! gcloud storage buckets create --location $REGION $BUCKET_NAME" ]
},
{
"cell_type": "markdown",
@@ -498,8 +497,7 @@
},
"outputs": [],
"source": [
"! gsutil ls -al $BUCKET_NAME"
]
"! gcloud storage ls --all-versions --long $BUCKET_NAME" ]
},
{
"cell_type": "markdown",
@@ -582,8 +580,7 @@
"outputs": [],
"source": [
"# Download the sample data into your RAW_DATA_PATH\n",
"! gsutil cp \"gs://cloud-samples-data/vertex-ai/community-content/tf_agents_bandits_movie_recommendation_with_kfp_and_vertex_sdk/u.data\" $RAW_DATA_PATH"
]
"! gcloud storage cp \"gs://cloud-samples-data/vertex-ai/community-content/tf_agents_bandits_movie_recommendation_with_kfp_and_vertex_sdk/u.data\" $RAW_DATA_PATH" ]
},
{
"cell_type": "code",
@@ -1621,9 +1618,7 @@
"! gcloud scheduler jobs delete $SIMULATOR_SCHEDULER_JOB --quiet\n",
"\n",
"# Delete Cloud Storage objects that were created.\n",
"! gsutil -m rm -r $PIPELINE_ROOT\n",
"! gsutil -m rm -r $TRAINING_ARTIFACTS_DIR"
]
"! gcloud storage rm --recursive $PIPELINE_ROOT\n", "! gcloud storage rm --recursive $TRAINING_ARTIFACTS_DIR" ]
}
],
"metadata": {
@@ -1,4 +1,4 @@
google-cloud-bigquery==2.20.0
tensorflow==2.12.1
pillow==10.3.0
pillow==12.3.0
tf-agents==0.8.0
@@ -1,4 +1,4 @@
google-cloud-pubsub==2.5.0
pillow==10.3.0
pillow==12.3.0
tf-agents==0.8.0
tensorflow==2.12.1
@@ -1,5 +1,5 @@
dataclasses==0.6
google-cloud-aiplatform==1.8.1
tensorflow==2.12.1
pillow==10.3.0
pillow==12.3.0
tf-agents==0.8.0
@@ -398,6 +398,7 @@
"if not IS_GOOGLE_CLOUD_NOTEBOOK:\n",
" if \"google.colab\" in sys.modules:\n",
" from google.colab import auth as google_auth\n",
"\n",
" google_auth.authenticate_user()\n",
"\n",
" # If you are running this notebook locally, replace the string below with the\n",
@@ -472,7 +473,7 @@
},
"outputs": [],
"source": [
"! gsutil mb -l $REGION $BUCKET_NAME"
"! gcloud storage buckets create --location $REGION $BUCKET_NAME"
]
},
{
@@ -492,7 +493,7 @@
},
"outputs": [],
"source": [
"! gsutil ls -al $BUCKET_NAME"
"! gcloud storage ls --all-versions --long $BUCKET_NAME"
]
},
{
@@ -565,7 +566,7 @@
"outputs": [],
"source": [
"# Copy the sample data into your DATA_PATH\n",
"! gsutil cp \"gs://cloud-samples-data/vertex-ai/community-content/tf_agents_bandits_movie_recommendation_with_kfp_and_vertex_sdk/u.data\" $DATA_PATH"
"! gcloud storage cp \"gs://cloud-samples-data/vertex-ai/community-content/tf_agents_bandits_movie_recommendation_with_kfp_and_vertex_sdk/u.data\" $DATA_PATH"
]
},
{
@@ -579,11 +580,15 @@
"# Set hyperparameters.\n",
"BATCH_SIZE = 8 # @param {type:\"integer\"} Training and prediction batch size.\n",
"TRAINING_LOOPS = 5 # @param {type:\"integer\"} Number of training iterations.\n",
"STEPS_PER_LOOP = 2 # @param {type:\"integer\"} Number of driver steps per training iteration.\n",
"STEPS_PER_LOOP = (\n",
" 2 # @param {type:\"integer\"} Number of driver steps per training iteration.\n",
")\n",
"\n",
"# Set MovieLens simulation environment parameters.\n",
"RANK_K = 20 # @param {type:\"integer\"} Rank for matrix factorization in the MovieLens environment; also the observation dimension.\n",
"NUM_ACTIONS = 20 # @param {type:\"integer\"} Number of actions (movie items) to choose from.\n",
"NUM_ACTIONS = (\n",
" 20 # @param {type:\"integer\"} Number of actions (movie items) to choose from.\n",
")\n",
"PER_ARM = False # Use the non-per-arm version of the MovieLens environment.\n",
"\n",
"# Set agent parameters.\n",
@@ -621,7 +626,8 @@
"source": [
"# Define RL environment.\n",
"env = movielens_py_environment.MovieLensPyEnvironment(\n",
" DATA_PATH, RANK_K, BATCH_SIZE, num_movies=NUM_ACTIONS, csv_delimiter=\"\\t\")\n",
" DATA_PATH, RANK_K, BATCH_SIZE, num_movies=NUM_ACTIONS, csv_delimiter=\"\\t\"\n",
")\n",
"environment = tf_py_environment.TFPyEnvironment(env)\n",
"\n",
"# Define RL agent/algorithm.\n",
@@ -631,7 +637,8 @@
" tikhonov_weight=TIKHONOV_WEIGHT,\n",
" alpha=AGENT_ALPHA,\n",
" dtype=tf.float32,\n",
" accepts_per_arm_features=PER_ARM)\n",
" accepts_per_arm_features=PER_ARM,\n",
")\n",
"print(\"TimeStep Spec (for each batch):\\n\", agent.time_step_spec, \"\\n\")\n",
"print(\"Action Spec (for each batch):\\n\", agent.action_spec, \"\\n\")\n",
"print(\"Reward Spec (for each batch):\\n\", environment.reward_spec(), \"\\n\")\n",
@@ -639,7 +646,8 @@
"# Define RL metric.\n",
"optimal_reward_fn = functools.partial(\n",
" environment_utilities.compute_optimal_reward_with_movielens_environment,\n",
" environment=environment)\n",
" environment=environment,\n",
")\n",
"regret_metric = tf_bandit_metrics.RegretMetric(optimal_reward_fn)\n",
"metrics = [regret_metric]"
]
@@ -704,35 +712,38 @@
" if training_data_spec_transformation_fn is None:\n",
" data_spec = agent.policy.trajectory_spec\n",
" else:\n",
" data_spec = training_data_spec_transformation_fn(\n",
" agent.policy.trajectory_spec)\n",
" replay_buffer = trainer.get_replay_buffer(data_spec, environment.batch_size,\n",
" steps_per_loop)\n",
" data_spec = training_data_spec_transformation_fn(agent.policy.trajectory_spec)\n",
" replay_buffer = trainer.get_replay_buffer(\n",
" data_spec, environment.batch_size, steps_per_loop\n",
" )\n",
"\n",
" # `step_metric` records the number of individual rounds of bandit interaction;\n",
" # that is, (number of trajectories) * batch_size.\n",
" step_metric = tf_metrics.EnvironmentSteps()\n",
" metrics = [\n",
" tf_metrics.NumberOfEpisodes(),\n",
" tf_metrics.AverageEpisodeLengthMetric(batch_size=environment.batch_size)\n",
" tf_metrics.AverageEpisodeLengthMetric(batch_size=environment.batch_size),\n",
" ]\n",
" if additional_metrics:\n",
" metrics += additional_metrics\n",
"\n",
" if isinstance(environment.reward_spec(), dict):\n",
" metrics += [tf_metrics.AverageReturnMultiMetric(\n",
" reward_spec=environment.reward_spec(),\n",
" batch_size=environment.batch_size)]\n",
" else:\n",
" metrics += [\n",
" tf_metrics.AverageReturnMetric(batch_size=environment.batch_size)]\n",
" tf_metrics.AverageReturnMultiMetric(\n",
" reward_spec=environment.reward_spec(), batch_size=environment.batch_size\n",
" )\n",
" ]\n",
" else:\n",
" metrics += [tf_metrics.AverageReturnMetric(batch_size=environment.batch_size)]\n",
"\n",
" # Store intermediate metric results, indexed by metric names.\n",
" metric_results = defaultdict(list)\n",
"\n",
" if training_data_spec_transformation_fn is not None:\n",
" def add_batch_fn(data): return replay_buffer.add_batch(training_data_spec_transformation_fn(data)) \n",
" \n",
"\n",
" def add_batch_fn(data):\n",
" return replay_buffer.add_batch(training_data_spec_transformation_fn(data))\n",
"\n",
" else:\n",
" add_batch_fn = replay_buffer.add_batch\n",
"\n",
@@ -742,10 +753,12 @@
" env=environment,\n",
" policy=agent.collect_policy,\n",
" num_steps=steps_per_loop * environment.batch_size,\n",
" observers=observers)\n",
" observers=observers,\n",
" )\n",
"\n",
" training_loop = trainer.get_training_loop_fn(\n",
" driver, replay_buffer, agent, steps_per_loop)\n",
" driver, replay_buffer, agent, steps_per_loop\n",
" )\n",
" saver = policy_saver.PolicySaver(agent.policy)\n",
"\n",
" for _ in range(training_loops):\n",
@@ -783,7 +796,8 @@
" environment=environment,\n",
" training_loops=TRAINING_LOOPS,\n",
" steps_per_loop=STEPS_PER_LOOP,\n",
" additional_metrics=metrics)\n",
" additional_metrics=metrics,\n",
")\n",
"\n",
"tf.profiler.experimental.stop()"
]
@@ -1092,11 +1106,15 @@
},
"outputs": [],
"source": [
"RUN_HYPERPARAMETER_TUNING = True # Execute hyperparameter tuning instead of regular training.\n",
"RUN_HYPERPARAMETER_TUNING = (\n",
" True # Execute hyperparameter tuning instead of regular training.\n",
")\n",
"TRAIN_WITH_BEST_HYPERPARAMETERS = False # Do not train.\n",
"\n",
"HPTUNING_RESULT_DIR = \"hptuning/\" # @param {type: \"string\"} Directory to store the best hyperparameter(s) in `BUCKET_NAME` and locally (temporarily).\n",
"HPTUNING_RESULT_PATH = os.path.join(HPTUNING_RESULT_DIR, \"result.json\") # @param {type: \"string\"} Path to the file containing the best hyperparameter(s)."
"HPTUNING_RESULT_PATH = os.path.join(\n",
" HPTUNING_RESULT_DIR, \"result.json\"\n",
") # @param {type: \"string\"} Path to the file containing the best hyperparameter(s)."
]
},
{
@@ -1124,7 +1142,7 @@
" image_uri: str,\n",
" args: List[str],\n",
" location: str = \"us-central1\",\n",
" api_endpoint: str = \"us-central1-aiplatform.googleapis.com\"\n",
" api_endpoint: str = \"us-central1-aiplatform.googleapis.com\",\n",
") -> None:\n",
" \"\"\"Creates a hyperparameter tuning job using a custom container.\n",
"\n",
@@ -1197,8 +1215,8 @@
"\n",
" # Create job\n",
" response = client.create_hyperparameter_tuning_job(\n",
" parent=parent,\n",
" hyperparameter_tuning_job=hyperparameter_tuning_job)\n",
" parent=parent, hyperparameter_tuning_job=hyperparameter_tuning_job\n",
" )\n",
" job_id = response.name.split(\"/\")[-1]\n",
" print(\"Job ID:\", job_id)\n",
" print(\"Job config:\", response)\n",
@@ -1242,7 +1260,8 @@
" image_uri=f\"gcr.io/{PROJECT_ID}/{HPTUNING_TRAINING_CONTAINER}:latest\",\n",
" args=args,\n",
" location=REGION,\n",
" api_endpoint=f\"{REGION}-aiplatform.googleapis.com\")"
" api_endpoint=f\"{REGION}-aiplatform.googleapis.com\",\n",
")"
]
},
{
@@ -1292,7 +1311,8 @@
" name = client.hyperparameter_tuning_job_path(\n",
" project=project,\n",
" location=location,\n",
" hyperparameter_tuning_job=hyperparameter_tuning_job_id)\n",
" hyperparameter_tuning_job=hyperparameter_tuning_job_id,\n",
" )\n",
" response = client.get_hyperparameter_tuning_job(name=name)\n",
" return response"
]
@@ -1313,7 +1333,8 @@
" location=REGION,\n",
" api_endpoint=f\"{REGION}-aiplatform.googleapis.com\")\n",
" if response.state.name == 'JOB_STATE_SUCCEEDED':\n",
" print(\"Job succeeded.\\nJob Time:\", response.update_time - response.create_time)\n",
" print(\"Job succeeded.\n",
"Job Time:\", response.update_time - response.create_time)\n",
" trials = response.trials\n",
" print(\"Trials:\", trials)\n",
" break\n",
@@ -1348,8 +1369,8 @@
"if trials:\n",
" # Dict mapping from metric names to the best metric values seen so far\n",
" best_objective_values = dict.fromkeys(\n",
" [metric.metric_id for metric in trials[0].final_measurement.metrics],\n",
" -np.inf)\n",
" [metric.metric_id for metric in trials[0].final_measurement.metrics], -np.inf\n",
" )\n",
" # Dict mapping from metric names to a list of the best combination(s) of\n",
" # hyperparameter(s). Each combination is a dict mapping from hyperparameter\n",
" # names to their values.\n",
@@ -1358,12 +1379,13 @@
" # `final_measurement` and `parameters` are `RepeatedComposite` objects.\n",
" # Reference the structure above to extract the value of your interest.\n",
" for metric in trial.final_measurement.metrics:\n",
" params = {\n",
" param.parameter_id: param.value for param in trial.parameters}\n",
" params = {param.parameter_id: param.value for param in trial.parameters}\n",
" if metric.value > best_objective_values[metric.metric_id]:\n",
" best_params[metric.metric_id] = [params]\n",
" elif metric.value == best_objective_values[metric.metric_id]:\n",
" best_params[param.parameter_id].append(params) # Handle cases where multiple hyperparameter values lead to the same performance.\n",
" best_params[param.parameter_id].append(\n",
" params\n",
" ) # Handle cases where multiple hyperparameter values lead to the same performance.\n",
" print(\"Best hyperparameter value(s):\")\n",
" for metric, params in best_params.items():\n",
" print(f\"Metric={metric}: {sorted(params)}\")\n",
@@ -1443,7 +1465,9 @@
},
"outputs": [],
"source": [
"PREDICTION_CONTAINER = \"prediction-custom-container\" # @param {type:\"string\"} Name of the container image."
"PREDICTION_CONTAINER = (\n",
" \"prediction-custom-container\" # @param {type:\"string\"} Name of the container image.\n",
")"
]
},
{
@@ -1475,7 +1499,7 @@
" machineType: 'E2_HIGHCPU_8'\"\"\".format(\n",
" PROJECT_ID=PROJECT_ID,\n",
" PREDICTION_CONTAINER=PREDICTION_CONTAINER,\n",
" ARTIFACTS_DIR=ARTIFACTS_DIR\n",
" ARTIFACTS_DIR=ARTIFACTS_DIR,\n",
")\n",
"\n",
"with open(\"cloudbuild.yaml\", \"w\") as fp:\n",
@@ -1592,8 +1616,12 @@
},
"outputs": [],
"source": [
"RUN_HYPERPARAMETER_TUNING = False # Execute regular training instead of hyperparameter tuning.\n",
"TRAIN_WITH_BEST_HYPERPARAMETERS = True # @param {type:\"bool\"} Whether to use learned hyperparameters in training."
"RUN_HYPERPARAMETER_TUNING = (\n",
" False # Execute regular training instead of hyperparameter tuning.\n",
")\n",
"TRAIN_WITH_BEST_HYPERPARAMETERS = (\n",
" True # @param {type:\"bool\"} Whether to use learned hyperparameters in training.\n",
")"
]
},
{
@@ -1633,10 +1661,12 @@
"job = aiplatform.CustomContainerTrainingJob(\n",
" display_name=\"train-movielens\",\n",
" container_uri=f\"gcr.io/{PROJECT_ID}/{HPTUNING_TRAINING_CONTAINER}:latest\",\n",
" command=[\"python3\", \"-m\", \"src.training.task\"] + args, # Pass in training arguments, including hyperparameters.\n",
" command=[\"python3\", \"-m\", \"src.training.task\"]\n",
" + args, # Pass in training arguments, including hyperparameters.\n",
" model_serving_container_image_uri=f\"gcr.io/{PROJECT_ID}/{PREDICTION_CONTAINER}:latest\",\n",
" model_serving_container_predict_route=\"/predict\",\n",
" model_serving_container_health_route=\"/health\")\n",
" model_serving_container_health_route=\"/health\",\n",
")\n",
"\n",
"print(\"Training Spec:\", job._managed_model)\n",
"\n",
@@ -1645,7 +1675,8 @@
" replica_count=1,\n",
" machine_type=\"n1-standard-4\",\n",
" accelerator_type=\"ACCELERATOR_TYPE_UNSPECIFIED\",\n",
" accelerator_count=0)"
" accelerator_count=0,\n",
")"
]
},
{
@@ -1784,7 +1815,7 @@
"! gcloud ai models delete $model.name --quiet\n",
"\n",
"# Delete Cloud Storage objects that were created\n",
"! gsutil -m rm -r $ARTIFACTS_DIR"
"! gcloud storage rm --recursive $ARTIFACTS_DIR"
]
}
],
@@ -324,7 +324,7 @@
},
"outputs": [],
"source": [
"! gsutil ls $gcs_output_uri_prefix"
"! gcloud storage ls $gcs_output_uri_prefix"
]
},
{
@@ -344,7 +344,7 @@
},
"outputs": [],
"source": [
"! gsutil rm -rf $gcs_output_uri_prefix"
"! gcloud storage rm --recursive --continue-on-error $gcs_output_uri_prefix"
]
}
],
@@ -328,7 +328,7 @@
},
"outputs": [],
"source": [
"! gsutil ls $gcs_output_uri_prefix"
"! gcloud storage ls $gcs_output_uri_prefix"
]
},
{
@@ -348,7 +348,7 @@
},
"outputs": [],
"source": [
"! gsutil rm -rf $gcs_output_uri_prefix"
"! gcloud storage rm --recursive --continue-on-error $gcs_output_uri_prefix"
]
}
],
@@ -341,7 +341,7 @@
},
"outputs": [],
"source": [
"! gsutil ls $gcs_output_uri_prefix"
"! gcloud storage ls $gcs_output_uri_prefix"
]
},
{
@@ -361,7 +361,7 @@
},
"outputs": [],
"source": [
"! gsutil rm -rf $gcs_output_uri_prefix"
"! gcloud storage rm --recursive --continue-on-error $gcs_output_uri_prefix"
]
}
],
@@ -0,0 +1,126 @@
# Vertex AI Training: Llama 3.1 8B pre-training using Nvidia A3 Mega VMs (H100)
This document provides a step-by-step guide for pre-training a Llama 3.1 8B model on the `en-wiki` dataset using multiple [Vertex AI Custom Training](https://cloud.google.com/vertex-ai/docs/training/overview) `a3-megagpu-8g` nodes.
We will use a custom container based on NVIDIA's [NeMo Framework](https://docs.nvidia.com/nemo-framework/user-guide/24.07/overview.html) to demonstrate a scalable, multi-node training workflow. All required artifacts and commands are included.
## 1. Prerequisites
### 1.1. Google Cloud Project setup
- **Enable APIs:** Ensure the Vertex AI API is [enabled for your project](http://console.cloud.google.com/flows/enableapi?apiid=aiplatform.googleapis.com).
- **H100 Mega Quota:** A3 Mega VMs are powered by H100 GPUs. Request quota for `custom_model_training_nvidia_h100_mega_gpus` in one of the [supported regions](https://cloud.google.com/vertex-ai/docs/general/locations#accelerator_support). If using Spot VMs, request `custom_model_training_preemptible_nvidia_h100_mega_gpus` quota instead.
- **Reservations (Optional but recommended):** For guaranteed capacity, [create a reservation](https://cloud.google.com/compute/docs/instances/reservations-shared) and ensure the reservation is shared with the Vertex AI service account. This guide requires a minimum of **16 H100 GPUs** (2 full A3 Mega nodes).
### 1.2. GCS bucket
Create a [Cloud Storage bucket](https://cloud.google.com/storage/docs/creating-buckets) in the same region where you have quota. If you're using Hierarchical Namespace for your bucket, you may need to update permissions of the Vertex AI Custom Code Service Agent .
This bucket is used for:
- Staging the training application.
- Storing model checkpoints and logs.
- Storing data if you use your own data.
## 2. Setup & configuration
### 2.1. Clone the repo
First clone the repo into your development environment.
```bash
git clone https://github.com/GoogleCloudPlatform/vertex-ai-samples.git
```
Navigate to the root folder for this sample.
### 2.2. Environment Setup
First, configure your local environment. These variables are used in subsequent commands.
```bash
# Required: Update with your values
export PROJECT_ID="<your-project-id>"
export REPOSITORY="<your-artifact-registry-repo-name>" # e.g., "my-containers"
export BUCKET="<your-gcs-bucket-name>"
# Optional: Change if needed
export REGION="us-central1"
# --- Do not change the lines below ---
export ARTIFACT_REGISTRY="${REGION}-docker.pkg.dev/${PROJECT_ID}/${REPOSITORY}"
export REPO_ROOT=$(git rev-parse --show-toplevel)
```
## 3. Build and push a docker container image to Artifact Registry
Normally, you can use any custom training container on Vertex AI Training. In this example you build a NeMo Docker image that is based on the [Nvidia’s NeMo 24.09](https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo/tags) image. Use Cloud Build to build and push the container image.
This document picked NeMo as the demonstrating container since it’s a widely adopted GPU LLM training framework providing high performance and versatile training functionalities.
In addition to the base image, some customizations are included to form the final prebuilt image:
- Some dependencies are installed to integrate with Vertex AI Training.
- An entrypoint script that sets up required environments and calls the training job.
- Some patches are applied to the NeMo code to let it load the dataset from a GCS bucket.
Run this command to build the container and push the container into the Google Artifact Registry.
```bash
cd "${REPO_ROOT}/community-content/vertex-distributed-training/a3mega/llama-3-8b-nemo-pretraining"
export IMAGE_NAME="vertex-nemo-llama"
gcloud builds submit . \
--project="${PROJECT_ID}" \
--region="${REGION}" \
--config=docker/cloudbuild.yml \
--substitutions="_ARTIFACT_REGISTRY=${ARTIFACT_REGISTRY},_IMAGE_NAME=${IMAGE_NAME}" \
--timeout="2h" \
--machine-type="e2-highcpu-32"
```
## 4. Launch the Training Job
### 4.1. Job Configuration File
Once the container is built, update the job_config.json to set up the training job.
File: job_config.json
```json
{
"project_id": "<project-id>",
"region": "<region>",
"zone": "<zone if using reservation>",
"bucket": "<bucket>",
"dataset_bucket": "github-repo/data/third-party/enwiki-latest-pages-articles",
"image_uri": "<docker image uri from artifact registry>",
"strategy": "spot",
"nodes": "2",
"machine_type": "a3-megagpu-8g",
"gpu_type": "NVIDIA_H100_MEGA_80GB",
"gpus_per_node": "8",
"recipe_name": "llama3_1_8b_pretrain_a3mega",
"job_prefix": "vertex-spot-",
"reservation_name": ""
}
```
### 4.2 Launch the Training Job
First, create a Python virtual environment using your tool of choice, then install
the requirements specified in `requirements.txt`. Using `pip`, the command would be:
```bash
pip install -r requirements.txt
```
Now launch the Vertex AI training job using the provided Python script.
```bash
python3 scripts/launch.py --config_file=job_config.json
```
This script reads job_config.json, defines the cluster specification (2 nodes, 8 GPUs each), and submits the custom training job to Vertex AI.
## 5. Monitor and Clean Up
### 5.1. Monitoring
Vertex AI Console: Track the job's status in the Google Cloud Console under Vertex AI > Training > Custom Jobs.
Logs: View detailed logs in Cloud Logging by filtering for your job name.
Checkpoints: Model checkpoints are saved to your GCS bucket at the path specified in your training script's configuration.
### 5.2. Cleaning Up
To avoid ongoing charges, delete the resources you created:
- The Artifact Registry image.
- The contents of the GCS bucket (checkpoints, logs).
- The Vertex AI Custom Job will eventually complete or fail, incurring no further cost.
@@ -0,0 +1,265 @@
# Reference:
# https://github.com/NVIDIA/NeMo-Framework-Launcher/blob/24.07/launcher_scripts/conf/training/llama/llama3_1_8b.yaml
name: llama3_1_8b_pretrain_a3mega
restore_from_path: null # used when starting from a .nemo file
trainer:
devices: 8
num_nodes: 1
accelerator: gpu
precision: bf16
logger: false # logger provided by exp_manager
enable_checkpointing: false
use_distributed_sampler: false
max_epochs: -1 # PTL default. In practice, max_steps will be reached first.
max_steps: 30 # consumed_samples = global_step * micro_batch_size * data_parallel_size * accumulate_grad_batches
log_every_n_steps: 1
val_check_interval: null
limit_val_batches: 1
limit_test_batches: 1
accumulate_grad_batches: 1 # do not modify, grad acc is automatic for training megatron models
gradient_clip_val: 1.0
benchmark: false
enable_model_summary: false # default PTL callback for this does not support model parallelism, instead we log manually
exp_manager:
explicit_log_dir: null
exp_dir: /data
name: ${name}
create_dllogger_logger: true
dllogger_logger_kwargs:
verbose: true
stdout: true
json_file: "/data/dllogger.json"
create_wandb_logger: false
wandb_logger_kwargs:
project: null
name: null
resume_if_exists: true
resume_ignore_no_checkpoint: true
create_checkpoint_callback: false
checkpoint_callback_params:
monitor: val_loss
save_top_k: 3
mode: min
always_save_nemo: false # saves nemo file during validation, not implemented for model parallel
save_nemo_on_train_end: false # not recommended when training large models on clusters with short time limits
filename: 'megatron_gpt--{val_loss:.2f}-{step}-{consumed_samples}'
model_parallel_size: ${multiply:${model.tensor_model_parallel_size}, ${model.pipeline_model_parallel_size}}
seconds_to_sleep: 5 # Allows node_rank!=0 to sleep and let node0 to init, like preparing data
model:
mcore_gpt: true
# specify micro_batch_size, global_batch_size, and model parallelism
# gradient accumulation will be done automatically based on data_parallel_size
micro_batch_size: 1 # limited by GPU memory
global_batch_size: 1024 # will use more micro batches to reach global batch size
tensor_model_parallel_size: 1 # intra-layer model parallelism
pipeline_model_parallel_size: 2 # inter-layer model parallelism
context_parallel_size: 1
virtual_pipeline_model_parallel_size: null # interleaved pipeline
## Sequence Parallelism
# Makes tensor parallelism more memory efficient for LLMs (20B+) by parallelizing layer norms and dropout sequentially
# See Reducing Activation Recomputation in Large Transformer Models: https://arxiv.org/abs/2205.05198 for more details.
sequence_parallel: false
fsdp: false
fsdp_cpu_offload: true
fsdp_sharding_strategy: "full" # Method to shard model states. Available options are 'full', 'hybrid', and 'grad'.
fsdp_grad_reduce_dtype: "16" # Gradient reduction data type.
fsdp_sharded_checkpoint: false # Store and load FSDP shared checkpoint.
fsdp_use_orig_params: false # Set to True to use FSDP for specific peft scheme.
# Distributed checkpoint setup
dist_ckpt_format: "torch_dist" # Set to 'torch_dist' to use PyTorch distributed checkpoint format.
dist_ckpt_load_on_device: true # whether to load checkpoint weights directly on GPU or to CPU
dist_ckpt_parallel_save: true # if true, each worker will write its own part of the dist checkpoint
dist_ckpt_parallel_save_within_dp: false # if true, save will be parallelized only within a DP group (whole world otherwise), which might slightly reduce the save overhead
dist_ckpt_parallel_load: false # if true, each worker will load part of the dist checkpoint and exchange with NCCL. Might use some extra GPU memory
dist_ckpt_torch_dist_multiproc: 2 # number of extra processes per rank used during ckpt save with PyTorch distributed format
dist_ckpt_assume_constant_structure: false # set to True only if the state dict structure doesn't change within a single job. Allows caching some computation across checkpoint saves.
dist_ckpt_parallel_dist_opt: true # parallel save/load of a DistributedOptimizer. 'True' allows performant save and reshardable checkpoints. Set to 'False' only in order to minimize the number of checkpoint files.
dist_ckpt_load_strictness: null # defines checkpoint keys mismatch behavior (only during dist-ckpt load). Choices: assume_ok_unexpected (default - try loading without any check), log_all (log mismatches), raise_all (raise mismatches)
# model architecture
encoder_seq_length: 8192
max_position_embeddings: ${.encoder_seq_length}
num_layers: 32 # 8b: 32 | 70b: 80 | 405b: 126
hidden_size: 4096 # 8b: 4096 | 70b: 8192 | 405b: 16384
ffn_hidden_size: 14336 # 8b: 14336 | 70b: 28672 | 405b: 53248
num_attention_heads: 32 # 8b: 32 | 70b: 64 | 405b: 128
num_query_groups: 8 # Number of query groups for group query attention. If None, normal attention is used. 8b: 8 | 70b: 8 | 405b: 16
init_method_std: 0.01 # Standard deviation of the zero mean normal distribution used for weight initialization. 8b: 0.01 | 70b: 0.008944 | 405b: 0.02
use_scaled_init_method: true # use scaled residuals initialization
hidden_dropout: 0.0 # Dropout probability for hidden state transformer.
attention_dropout: 0.0 # Dropout probability for attention
ffn_dropout: 0.0 # Dropout probability in the feed-forward layer.
kv_channels: null # Projection weights dimension in multi-head attention. Set to hidden_size // num_attention_heads if null
apply_query_key_layer_scaling: true # scale Q * K^T by 1 / layer-number.
normalization: 'rmsnorm' # Normalization layer to use. Options are 'layernorm', 'rmsnorm'
layernorm_epsilon: 1e-5
do_layer_norm_weight_decay: false # True means weight decay on all params
make_vocab_size_divisible_by: 128 # Pad the vocab size to be divisible by this value for computation efficiency.
pre_process: true # add embedding
post_process: true # add pooler
persist_layer_norm: true # Use of persistent fused layer norm kernel.
bias: false # Whether to use bias terms in all weight matrices.
activation: 'fast-swiglu' # Options ['gelu', 'geglu', 'swiglu', 'reglu', 'squared-relu', 'fast-geglu', 'fast-swiglu', 'fast-reglu']
headscale: false # Whether to learn extra parameters that scale the output of the each self-attention head.
transformer_block_type: 'pre_ln' # Options ['pre_ln', 'post_ln', 'normformer']
openai_gelu: false # Use OpenAI's GELU instead of the default GeLU
normalize_attention_scores: true # Whether to scale the output Q * K^T by 1 / sqrt(hidden_size_per_head). This arg is provided as a configuration option mostly for compatibility with models that have been weight-converted from HF. You almost always want to se this to True.
position_embedding_type: 'rope' # Position embedding type. Options ['learned_absolute', 'rope']
rotary_percentage: 1.0 # If using position_embedding_type=rope, then the per head dim is multiplied by this.
attention_type: 'multihead' # Attention type. Options ['multihead']
share_embeddings_and_output_weights: false # Share embedding and output layer weights.
scale_positional_embedding: true # This is false for llama3 models. Only used for >= llama3.1.
# Use GPT2BPETokenizer for test, because the testing dataset is tokenized by this tokenizer.
# https://docs.nvidia.com/nemo-framework/user-guide/24.07/playbooks/singlenodepretrain.html#data-download-and-pre-processing
tokenizer:
library: megatron
type: GPT2BPETokenizer
model: null # /path/to/tokenizer.model
vocab_file: null
merge_file: null
delimiter: null # only used for tabular tokenizer
sentencepiece_legacy: false # Legacy=True allows you to add special tokens to sentencepiece tokenizers.
# Mixed precision
native_amp_init_scale: 4294967296 # 2 ** 32
native_amp_growth_interval: 1000
hysteresis: 2 # Gradient scale hysteresis
fp32_residual_connection: false # Move residual connections to fp32
fp16_lm_cross_entropy: false # Move the cross entropy unreduced loss calculation for lm head to fp16
# Megatron O2-style half-precision
megatron_amp_O2: true # Enable O2-level automatic mixed precision using main parameters
grad_allreduce_chunk_size_mb: 125
# Fusion
grad_div_ar_fusion: true # Fuse grad division into torch.distributed.all_reduce. Only used with O2 and no pipeline parallelism..
gradient_accumulation_fusion: true # Fuse weight gradient accumulation to GEMMs. Only used with pipeline parallelism and O2.
bias_activation_fusion: true # Use a kernel that fuses the bias addition from weight matrices with the subsequent activation function.
bias_dropout_add_fusion: true # Use a kernel that fuses the bias addition, dropout and residual connection addition.
masked_softmax_fusion: true # Use a kernel that fuses the attention softmax with it's mask.
apply_rope_fusion: true # Use a kernel to add rotary positional embeddings. Only used if position_embedding_type=rope
cross_entropy_loss_fusion: true
# Miscellaneous
seed: 1234
resume_from_checkpoint: null # manually set the checkpoint file to load from
use_cpu_initialization: false # Init weights on the CPU (slow for large models)
onnx_safe: false # Use work-arounds for known problems with Torch ONNX exporter.
apex_transformer_log_level: 30 # Python logging level displays logs with severity greater than or equal to this
gradient_as_bucket_view: true # PyTorch DDP argument. Allocate gradients in a contiguous bucket to save memory (less fragmentation and buffer memory)
sync_batch_comm: false # Enable stream synchronization after each p2p communication between pipeline stages
## Activation Checkpointing
# NeMo Megatron supports 'selective' activation checkpointing where only the memory intensive part of attention is checkpointed.
# These memory intensive activations are also less compute intensive which makes activation checkpointing more efficient for LLMs (20B+).
# See Reducing Activation Recomputation in Large Transformer Models: https://arxiv.org/abs/2205.05198 for more details.
# 'full' will checkpoint the entire transformer layer.
activations_checkpoint_granularity: null # 'selective' or 'full'
activations_checkpoint_method: null # 'uniform', 'block'
# 'uniform' divides the total number of transformer layers and checkpoints the input activation
# of each chunk at the specified granularity. When used with 'selective', 'uniform' checkpoints all attention blocks in the model.
# 'block' checkpoints the specified number of layers per pipeline stage at the specified granularity
activations_checkpoint_num_layers: null
# when using 'uniform' this creates groups of transformer layers to checkpoint. Usually set to 1. Increase to save more memory.
# when using 'block' this this will checkpoint the first activations_checkpoint_num_layers per pipeline stage.
num_micro_batches_with_partial_activation_checkpoints: null
# This feature is valid only when used with pipeline-model-parallelism.
# When an integer value is provided, it sets the number of micro-batches where only a partial number of Transformer layers get checkpointed
# and recomputed within a window of micro-batches. The rest of micro-batches in the window checkpoint all Transformer layers. The size of window is
# set by the maximum outstanding micro-batch backpropagations, which varies at different pipeline stages. The number of partial layers to checkpoint
# per micro-batch is set by 'activations_checkpoint_num_layers' with 'activations_checkpoint_method' of 'block'.
# This feature enables using activation checkpoint at a fraction of micro-batches up to the point of full GPU memory usage.
activations_checkpoint_layers_per_pipeline: null
# This feature is valid only when used with pipeline-model-parallelism.
# When an integer value (rounded down when float is given) is provided, it sets the number of Transformer layers to skip checkpointing at later
# pipeline stages. For example, 'activations_checkpoint_layers_per_pipeline' of 3 makes pipeline stage 1 to checkpoint 3 layers less than
# stage 0 and stage 2 to checkpoint 6 layers less stage 0, and so on. This is possible because later pipeline stage
# uses less GPU memory with fewer outstanding micro-batch backpropagations. Used with 'num_micro_batches_with_partial_activation_checkpoints',
# this feature removes most of activation checkpoints at the last pipeline stage, which is the critical execution path.
## Transformer Engine
transformer_engine: true
fp8: false # enables fp8 in TransformerLayer forward
fp8_e4m3: false # sets fp8_format = recipe.Format.E4M3
fp8_hybrid: false # sets fp8_format = recipe.Format.HYBRID
fp8_margin: 0 # scaling margin
fp8_interval: 1 # scaling update interval
fp8_amax_history_len: 1024 # Number of steps for which amax history is recorded per tensor
fp8_amax_compute_algo: 'max' # 'most_recent' or 'max'. Algorithm for computing amax from history
ub_tp_comm_overlap: false # do not turn on because of b/397797926
use_flash_attention: true
gc_interval: 100
## Offloading Activations/Weights to CPU
cpu_offloading: false
cpu_offloading_num_layers: ${sum:${.num_layers},-1} # This value should be between [1,num_layers-1] as we don't want to offload the final layer's activations and expose any offloading duration for the final layer
cpu_offloading_activations: true
cpu_offloading_weights: true
data:
# Path to data must be specified by the user.
# Supports List, String and Dictionary
# List : can override from the CLI: "model.data.data_prefix=[.5,/raid/data/pile/my-gpt3_00_text_document,.5,/raid/data/pile/my-gpt3_01_text_document]",
# Or see example below:
# data_prefix:
# - .5
# - /raid/data/pile/my-gpt3_00_text_document
# - .5
# - /raid/data/pile/my-gpt3_01_text_document
# Dictionary: can override from CLI "model.data.data_prefix"={"train":[1.0, /path/to/data], "validation":/path/to/data, "test":/path/to/test}
# Or see example below:
# "model.data.data_prefix: {train:[1.0,/path/to/data], validation:[/path/to/data], test:[/path/to/test]}"
data_prefix: [1.0, /data/hfbpe_gpt_training_data_text_document]
index_mapping_dir: null # path to save index mapping .npy files, by default will save in the same location as data_prefix
data_impl: mmap
splits_string: 900,50,50
seq_length: ${model.encoder_seq_length}
skip_warmup: true
num_workers: 2
dataloader_type: single # cyclic
reset_position_ids: false # Reset position ids after end-of-document token
reset_attention_mask: false # Reset attention mask after end-of-document token
eod_mask_loss: false # Mask loss for the end of document tokens
validation_drop_last: true # Set to false if the last partial validation samples is to be consumed
no_seqlen_plus_one_input_tokens: false # Set to True to disable fetching (sequence length + 1) input tokens, instead get (sequence length) input tokens and mask the last token
pad_samples_to_global_batch_size: false # Set to True if you want to pad the last partial batch with -1's to equal global batch size
shuffle_documents: true # Set to False to disable documents shuffling. Sample index will still be shuffled
# Nsys profiling options
nsys_profile:
enabled: false
start_step: 0 # Global batch to start profiling
end_step: 1 # Global batch to end profiling
ranks: [0] # Global rank IDs to profile
gen_shape: false # Generate model and kernel details including input shapes
memory_profile:
enabled: false
start_step: 0
end_step: 1
ranks: [0]
output_path: /data # Must be a dir
optim:
name: distributed_fused_adam # E.g., fused_adam or set _target_: torch.optim.AdamW field
lr: 2e-5
weight_decay: 0.01
betas:
- 0.9
- 0.98
bucket_cap_mb: 125
overlap_grad_sync: true
overlap_param_sync: true
contiguous_grad_buffer: true
contiguous_param_buffer: true
sched:
name: CosineAnnealing
warmup_steps: 400
constant_steps: 0
min_lr: 2e-6
@@ -0,0 +1,26 @@
# Copyright 2024 Google LLC
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
steps:
- name: 'gcr.io/cloud-builders/docker'
args:
- 'build'
- '--tag=${_ARTIFACT_REGISTRY}/${_IMAGE_NAME}'
- '--file=docker/vertex-dist-recipes.Dockerfile'
- '.'
automapSubstitutions: true
env:
- 'DOCKER_BUILDKIT=1'
images:
- '${_ARTIFACT_REGISTRY}/${_IMAGE_NAME}'
@@ -0,0 +1,41 @@
diff --git a/nemo/collections/nlp/parts/megatron_trainer_builder.py b/nemo/collections/nlp/parts/megatron_trainer_builder.py
index b2c85cde4..a3a9670c3 100644
--- a/nemo/collections/nlp/parts/megatron_trainer_builder.py
+++ b/nemo/collections/nlp/parts/megatron_trainer_builder.py
@@ -19,6 +19,7 @@ from lightning_fabric.utilities.exceptions import MisconfigurationException
from omegaconf import DictConfig
from pytorch_lightning import Trainer
from pytorch_lightning.callbacks import ModelSummary
+from pytorch_lightning.callbacks import Callback
from pytorch_lightning.plugins.environments import TorchElasticEnvironment
from nemo.collections.common.metrics.perf_metrics import FLOPsMeasurementCallback
@@ -38,6 +39,23 @@ from nemo.utils.callbacks.dist_ckpt_io import (
AsyncFinalizerCallback,
DistributedCheckpointIO,
)
+from vmg.util.device_stats import gpu_stats_str
+
+class GpuStatsMon(Callback):
+ def on_train_start(self, trainer, pl_module) -> None:
+ rank=pl_module.global_rank
+ print(f'train_start: {rank=} {gpu_stats_str()}', flush=True)
+
+ def on_train_batch_start(self, trainer, pl_module, batch, batch_idx) -> None:
+ rank=pl_module.global_rank
+ print(f'batch_start: {rank=} {gpu_stats_str()}', flush=True)
+
+ def on_train_batch_end(self, trainer, pl_module, outputs, batch, batch_idx) -> None:
+ rank=pl_module.global_rank
+ print(f'batch_end: {rank=} {gpu_stats_str()}', flush=True)
class MegatronTrainerBuilder:
@@ -178,6 +196,7 @@ class MegatronTrainerBuilder:
if self.cfg.get('exp_manager', {}).get('log_tflops_per_sec_per_gpu', True):
callbacks.append(FLOPsMeasurementCallback(self.cfg))
+ callbacks.append(GpuStatsMon())
return callbacks
def create_trainer(self, callbacks=None) -> Trainer:
@@ -0,0 +1,41 @@
diff -ruN old-datasets/blended_megatron_dataset_builder.py datasets/blended_megatron_dataset_builder.py
--- old-datasets/blended_megatron_dataset_builder.py 2025-05-02 04:08:45.369199665 +0000
+++ datasets/blended_megatron_dataset_builder.py 2025-05-02 04:10:47.369119891 +0000
@@ -2,6 +2,7 @@
import logging
import math
+import os
from concurrent.futures import ThreadPoolExecutor
from typing import Any, Callable, Iterable, List, Optional, Type, Union
@@ -353,7 +354,7 @@
num_dataset_builder_threads = self.config.num_dataset_builder_threads
if torch.distributed.is_initialized():
- rank = torch.distributed.get_rank()
+ rank = int(os.getenv("LOCAL_RANK", "0"))
# First, build on rank 0
if rank == 0:
num_workers = num_dataset_builder_threads
@@ -475,7 +476,7 @@
Optional[Union[DistributedDataset, Iterable]]: The DistributedDataset instantion, the Iterable instantiation, or None
"""
if torch.distributed.is_initialized():
- rank = torch.distributed.get_rank()
+ rank = int(os.getenv("LOCAL_RANK", "0"))
dataset = None
diff -ruN old-datasets/gpt_dataset.py datasets/gpt_dataset.py
--- old-datasets/gpt_dataset.py 2025-05-02 04:08:45.369199665 +0000
+++ datasets/gpt_dataset.py 2025-05-02 04:09:30.309170278 +0000
@@ -351,7 +351,7 @@
if not path_to_cache or (
not cache_hit
- and (not torch.distributed.is_initialized() or torch.distributed.get_rank() == 0)
+ and (not torch.distributed.is_initialized() or int(os.getenv("LOCAL_RANK", "0")) == 0)
):
log_single_rank(
@@ -0,0 +1,13 @@
diff --git a/scripts/checkpoint_converters/convert_llama_nemo_to_hf.py b/scripts/checkpoint_converters/convert_llama_nemo_to_hf.py
index 8da15148d..005cae6c9 100644
--- a/scripts/checkpoint_converters/convert_llama_nemo_to_hf.py
+++ b/scripts/checkpoint_converters/convert_llama_nemo_to_hf.py
@@ -104,6 +104,8 @@ def convert(input_nemo_file, output_hf_file, precision=None, cpu_only=False) ->
dummy_trainer = Trainer(devices=1, accelerator='cpu', strategy=NLPDDPStrategy())
model_config = MegatronGPTModel.restore_from(input_nemo_file, trainer=dummy_trainer, return_config=True)
model_config.tensor_model_parallel_size = 1
+ model_config.virtual_pipeline_model_parallel_size = None
+ model_config.sequence_parallel = False
model_config.pipeline_model_parallel_size = 1
if cpu_only:
map_location = torch.device('cpu')
@@ -0,0 +1,24 @@
diff --git a/examples/nlp/language_modeling/tuning/megatron_gpt_finetuning.py b/examples/nlp/language_modeling/tuning/megatron_gpt_finetuning.py
index bfe8ea359..dfeaf93b5 100644
--- a/examples/nlp/language_modeling/tuning/megatron_gpt_finetuning.py
+++ b/examples/nlp/language_modeling/tuning/megatron_gpt_finetuning.py
@@ -13,6 +13,8 @@
# limitations under the License.
import torch.multiprocessing as mp
+import torch.distributed as dist
+
from omegaconf.omegaconf import OmegaConf
from nemo.collections.nlp.models.language_modeling.megatron_gpt_sft_model import MegatronGPTSFTModel
@@ -76,6 +78,10 @@ def main(cfg) -> None:
trainer.fit(model)
+ if dist.is_available() and dist.is_initialized():
+ dist.barrier()
+ dist.destroy_process_group()
+
if __name__ == '__main__':
main()
@@ -0,0 +1,13 @@
diff --git a/src/utils/training_metrics/process_training_results.py b/src/utils/training_metrics/process_training_results.py
index 3e82a66..e61e1d8 100644
--- a/src/utils/training_metrics/process_training_results.py
+++ b/src/utils/training_metrics/process_training_results.py
@@ -134,7 +134,7 @@ def get_average_step_time(file: str, start_step: int, end_step: int) -> float:
for line in datajson:
if line.get("step") != "PARAMETER":
step = line.get("step")
- if step >= start_step and step <= end_step:
+ if step >= start_step and step <= end_step and "train_step_timing in s" in line["data"]:
time_step_accumulator += line["data"].get("train_step_timing in s")
num_steps += 1
if num_steps == 0:
@@ -0,0 +1,10 @@
dllogger@git+https://github.com/NVIDIA/dllogger@v1.0.0
# Fixing these libraries versions to avoid conflicting or broken packages.
immutabledict==4.2.1
protobuf==5.29.6
opencv-python-headless==4.11.0.86
docutils==0.16
urllib3==2.7.0
google-cloud-storage==3.0.0
retrying
@@ -0,0 +1,18 @@
# cuml-cu12==24.8.0 was installed in nemo:24.09
# Removing cuml=24.4.0 to avoid conflicting packages.
cudf==24.4.0
cugraph==24.4.0
cugraph-service-server==24.4.0
cuml==24.4.0
dask-cudf==24.4.0
raft-dask==24.4.0
cugraph-dgl==24.4.0
cugraph-pyg==24.4.0
# The following packages are removed temporarily to avoid conflicting packages
# and can be brought back if needed.
tensorrt-llm==0.12.0
img2dataset==1.45.0
Sphinx==8.1.3
sphinxcontrib-bibtex==2.6.3
torchx==0.7.0
nemo-run
@@ -0,0 +1,66 @@
# Dockerfile wrapping NeMo.
#
# To workaround base nemo docker image using too many layers, we use Multi-stage
# build to first collect the additional files we'll need.
FROM alpine:latest AS prep_files
WORKDIR /workspace
RUN mkdir -p configs vdt vdt/util
COPY scripts/*.py vdt/
COPY scripts/util/*.py vdt/util/
COPY configs/* configs/
COPY docker/patches/24.09/* vdt/patches/
RUN chmod a+rwX -R vdt
# Copy license.
RUN wget https://raw.githubusercontent.com/GoogleCloudPlatform/vertex-ai-samples/main/LICENSE
# Available tags
# https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo/tags
# It installs NeMo source code in /opt/NeMo folder, with tag=r2.0.0
FROM nvcr.io/nvidia/nemo:24.09
RUN apt-get update && apt-get install -y sudo zsh tmux && \
rm -rf /var/lib/apt/lists*
RUN echo "deb [signed-by=/usr/share/keyrings/cloud.google.gpg] http://packages.cloud.google.com/apt cloud-sdk main" | \
tee -a /etc/apt/sources.list.d/google-cloud-sdk.list && \
curl https://packages.cloud.google.com/apt/doc/apt-key.gpg | \
apt-key --keyring /usr/share/keyrings/cloud.google.gpg add - && \
apt-get update -y && apt-get install google-cloud-sdk -y && \
rm -rf /var/lib/apt/lists*
# Install libraries with pip
ENV PIP_ROOT_USER_ACTION=ignore
# We expect this will be run in the root directory of the vertex-dist-recipes repo
ARG HOST_SRC_DIR="."
# The pre-installed NeMo introduces a lot of deps conflicts.
# We uninstall the confilicting libs and reinstall some of them as needed.
COPY ${HOST_SRC_DIR}/docker/uninstall.txt /tmp/uninstall.txt
RUN cat /tmp/uninstall.txt | grep -v '#' | xargs pip uninstall -y
COPY ${HOST_SRC_DIR}/docker/requirements.txt /tmp/requirements.txt
RUN pip install -r /tmp/requirements.txt
# Make sure there's no inconsistent pip libraries.
RUN pip check
WORKDIR /workspace
# Copy configs
COPY ${HOST_SRC_DIR}/configs/* /opt/NeMo/examples/nlp/language_modeling/conf/
# Copy all additional files we need from `prep_files` image.
COPY --from=prep_files /workspace/ .
# Install for `src/utils/training_metrics/process_training_results.py` to report
# throughput and MFU numbers.
RUN git clone https://github.com/AI-Hypercomputer/gpu-recipes.git
# This hack is needed for multi-node training while not using a sharing file system.
RUN patch --verbose -l -d /opt/megatron-lm/megatron/core/datasets -p1 -i /workspace/vdt/patches/local_rank.patch; \
git -C /workspace/gpu-recipes apply /workspace/vdt/patches/throughput_calc.patch; \
git -C /opt/NeMo apply /workspace/vdt/patches/nemo2hf.patch; \
git -C /opt/NeMo apply /workspace/vdt/patches/sigabort.patch;
# git -C /opt/NeMo apply /workspace/vdt/patches/gpu_stats.patch;
# Do not put an entrypoint here. Specify the entrypoint in the docker run script.
@@ -0,0 +1,16 @@
{
"project_id": "<your_project_id>",
"region": "us-central1",
"zone": "us-central1-c",
"bucket": "<your_bucket",
"dataset_bucket": "github-repo/data/third-party/enwiki-latest-pages-articles",
"image_uri": "<your_image_uri>",
"strategy": "spot",
"nodes": "2",
"machine_type": "a3-megagpu-8g",
"gpu_type": "NVIDIA_H100_MEGA_80GB",
"gpus_per_node": "8",
"recipe_name": "llama3_1_8b_pretrain_a3mega",
"job_prefix": "vertex-ai",
"reservation_name": ""
}
@@ -0,0 +1,49 @@
absl-py==2.2.2
annotated-types==0.7.0
anyio==4.9.0
black==26.3.1
cachetools==5.5.2
certifi==2025.4.26
charset-normalizer==3.4.2
click==8.1.8
docstring_parser==0.16
google-api-core==2.24.2
google-auth==2.40.1
google-cloud-aiplatform==1.133.0
google-cloud-bigquery==3.31.0
google-cloud-core==2.4.3
google-cloud-resource-manager==1.14.2
google-cloud-storage==2.19.0
google-crc32c==1.7.1
google-genai==1.14.0
google-resumable-media==2.7.2
googleapis-common-protos==1.70.0
grpc-google-iam-v1==0.14.2
grpcio==1.71.0
grpcio-status==1.71.0
h11==0.16.0
httpcore==1.0.9
httpx==0.28.1
idna==3.15
mypy_extensions==1.1.0
numpy==2.2.5
packaging==25.0
pathspec==0.12.1
platformdirs==4.3.8
proto-plus==1.26.1
protobuf==5.29.6
pyasn1==0.6.4
pyasn1_modules==0.4.2
pydantic==2.11.4
pydantic_core==2.33.2
python-dateutil==2.9.0.post0
pytz==2025.2
requests==2.33.0
rsa==4.9.1
shapely==2.1.0
six==1.17.0
sniffio==1.3.1
typing-inspection==0.4.0
typing_extensions==4.13.2
urllib3==2.7.0
websockets==15.0.1
@@ -0,0 +1,173 @@
"""Launch script for Vertex distributed training"""
# Copy the sample_job_config.json file to job_config.json
# to define the job parameters.
#
# Run like this:
#
# python3 vertex_dist_train/launch.py --config_file=job_config.json
#
import datetime
import json
import os
import pprint
from collections.abc import Sequence
from typing import Any, List
from absl import app, flags
from google.cloud import aiplatform
from google.cloud.aiplatform_v1.types.custom_job import Scheduling
from pytz import timezone
FLAGS = flags.FLAGS
flags.DEFINE_string("config_file", None, "Path to JSON config file")
flags.DEFINE_boolean(
"debug", False, "Debug mode: just print the command, don't run it."
)
def launch_job(
job_name: str,
project: str,
region: str,
gcs_bucket: str,
image_uri: str,
entrypoint_cmd: List[str],
trainer_args: List[Any],
num_nodes: int,
machine_type: str,
num_gpus_per_node: int,
gpu_type: str,
strategy: str,
reservation_name: str = "",
):
assert strategy in ("dws", "spot", "reservation")
aiplatform.init(
project=project, location=region, staging_bucket=gcs_bucket
)
train_job = aiplatform.CustomContainerTrainingJob(
display_name=job_name,
container_uri=image_uri,
command=entrypoint_cmd,
)
job_args = dict(
args=trainer_args,
enable_web_access=True,
replica_count=num_nodes,
machine_type=machine_type,
accelerator_type=gpu_type,
accelerator_count=num_gpus_per_node,
boot_disk_size_gb=1000,
restart_job_on_worker_restart=True,
#restart_job_on_worker_restart=False,
)
if strategy == "spot":
job_args.update({"scheduling_strategy": Scheduling.Strategy.SPOT.name})
elif strategy == "dws":
job_args.update(
{"scheduling_strategy": Scheduling.Strategy.FLEX_START.name}
)
elif strategy == "reservation":
assert reservation_name != "", (
"If using a reservation, provide the reservation_name in the "
"format `projects/{project_id_or_number}/zones/{zone}/"
"reservations/{reservation_name}`"
)
job_args.update(
{
"reservation_affinity_type": "SPECIFIC_RESERVATION",
"reservation_affinity_key": "compute.googleapis.com/reservation-name",
"reservation_affinity_values": [reservation_name],
}
)
pprint.pprint(job_args)
if not FLAGS.debug:
train_job.submit(**job_args)
def main(argv: Sequence[str]) -> None:
config_file_path = FLAGS.config_file
print(f"Reading job config from {config_file_path}")
with open(config_file_path, encoding="utf-8") as config_file:
config = json.load(config_file)
project_id = config["project_id"]
region = config["region"]
zone = config["zone"]
bucket = config["bucket"]
dataset_bucket = config["dataset_bucket"]
n_nodes = int(config["nodes"])
machine_type = config["machine_type"]
num_gpus_per_node = int(config["gpus_per_node"])
gpu_type = config["gpu_type"]
reservation_name = config.get("reservation_name")
reservation_full_name = (
f"projects/{project_id}/zones/{zone}/reservations/{reservation_name}"
if "reservation_name" in config
else ""
)
strategy = config["strategy"]
recipe_name = config["recipe_name"]
job_prefix = config["job_prefix"]
image_uri = config["image_uri"]
# Job name
timestamp = (
datetime.datetime.now()
.astimezone(timezone("US/Pacific"))
.strftime("%Y%m%d_%H%M%S")
)
job_name = f"{recipe_name}-{timestamp}"
if job_prefix:
job_name = f"{job_prefix}-{job_name}"
base_output_dir = os.path.join("/gcs", bucket, job_name)
# Training command and args
entrypoint_cmd = ["python3", "vdt/run.py"]
dataset_bucket = f"gs://{config['dataset_bucket']}"
trainer_args = [
f"--train_data_gcs={dataset_bucket}",
"/opt/NeMo/examples/nlp/language_modeling/megatron_gpt_pretraining.py",
"--config-path=conf/",
f"--config-name={recipe_name}.yaml",
f"exp_manager.explicit_log_dir={base_output_dir}",
f"exp_manager.dllogger_logger_kwargs.json_file={base_output_dir}/dllogger.json",
"+exp_manager.create_tensorboard_logger=true",
"exp_manager.create_checkpoint_callback=false",
f"trainer.num_nodes={n_nodes}",
f"trainer.devices={num_gpus_per_node}",
"trainer.max_steps=10",
"trainer.log_every_n_steps=1",
"model.tokenizer.vocab_file=/data/gpt2-vocab.json",
"model.tokenizer.merge_file=/data/gpt2-merges.txt",
"model.data.data_prefix=[1.0,/data/hfbpe_gpt_training_data_text_document]",
]
launch_job(
job_name=job_name,
project=project_id,
region=region,
gcs_bucket=bucket,
image_uri=image_uri,
entrypoint_cmd=entrypoint_cmd,
trainer_args=trainer_args,
num_nodes=n_nodes,
machine_type=machine_type,
num_gpus_per_node=num_gpus_per_node,
gpu_type=gpu_type,
strategy=strategy,
reservation_name=reservation_full_name,
)
if __name__ == "__main__":
app.run(main)
@@ -0,0 +1,85 @@
"""Entrypoint for Vertex Distributed Training container."""
import argparse
import os
import sys
from collections.abc import Sequence
from subprocess import STDOUT, check_output, run
from absl import app, flags, logging
from util import cluster_spec
from retrying import retry
# PyTorch barrier call which synchronizes all of the nodes before launching the training process.
# This makes sure that processes will block until all processes are ready.
# Improves the reliability of spot VM usage for multi-node training jobs
@retry(stop_max_attempt_number=100, wait_exponential_multiplier=1000)
def barrier_with_retry() -> None:
import torch
logging.info("Starting barrier on RANK {}".format(os.environ["RANK"]))
torch.distributed.init_process_group()
torch.distributed.barrier()
torch.distributed.destroy_process_group()
logging.info("Finished barrier on RANK {}".format(os.environ["RANK"]))
def main(unused_argv: Sequence[str]) -> None:
parser = argparse.ArgumentParser()
parser.add_argument(
"--train_data_gcs",
type=str,
help="Download training data from gcs path",
)
args, unknown = parser.parse_known_args()
for key, val in os.environ.items():
logging.info("ENV %s=%s", key, val)
if args.train_data_gcs:
local_dir = "/data"
if not os.path.exists(local_dir):
os.mkdir(local_dir)
logging.info("downloading %s to %s...", args.train_data_gcs, local_dir)
check_output(
[
"gcloud",
"storage",
"cp",
"-r",
f"{args.train_data_gcs}/*",
local_dir,
],
stderr=STDOUT,
)
logging.info("%s downloaded.", args.train_data_gcs)
primary_node_addr, primary_node_port, node_rank, num_nodes = (
cluster_spec.get_cluster_spec()
)
cmd = [
"torchrun",
"--nproc-per-node=8",
f"--nnodes={num_nodes}",
f"--node_rank={node_rank}",
]
if num_nodes > 1:
cmd += [
"--max-restarts=3",
"--rdzv-backend=static",
f'--rdzv_id={os.getenv("CLOUD_ML_JOB_ID", primary_node_port)}',
f"--rdzv-endpoint={primary_node_addr}:{primary_node_port}",
]
cmd += unknown
logging.info("launching with cmd: \n%s", " \\\n".join(cmd))
barrier_with_retry()
run(cmd, stdout=sys.stdout, stderr=sys.stdout, check=True)
if __name__ == "__main__":
logging.get_absl_handler().python_handler.stream = sys.stdout
app.run(
main, flags_parser=lambda _args: flags.FLAGS(_args, known_only=True)
)
@@ -0,0 +1,81 @@
"""Get cluster info from environment variables."""
import dataclasses
import json
import os
from absl import logging
@dataclasses.dataclass
class ClusterInfo:
"""Contains information about the cluster.
Attributes:
primary_node_addr: The address of the primary node.
primary_node_port: The port of the primary node.
node_rank: The rank of the node.
num_nodes: The number of nodes in the cluster.
"""
primary_node_addr: str | None = None
primary_node_port: str | None = None
node_rank: int = 0
num_nodes: int = 1
# Allows unpacking operation like
# primary_node_addr, primary_node_port, _, _ = ClusterInfo()
# See https://stackoverflow.com/a/70753113
def __iter__(self):
return iter(dataclasses.astuple(self))
def get_cluster_spec() -> ClusterInfo:
"""Parses CLUSTER_SPEC environment variable and returns the cluster info.
Returns:
A ClusterInfo object.
"""
cluster_spec = os.getenv("CLUSTER_SPEC", None)
# If CLUSTER_SPEC is not set, use individual vars to construct cluster info.
if not cluster_spec:
cluster_info = ClusterInfo(
primary_node_addr=os.getenv("MASTER_ADDR", None),
primary_node_port=os.getenv("MASTER_PORT", None),
node_rank=int(os.getenv("RANK", "0")),
num_nodes=int(os.getenv("NNODES", "1")),
)
return cluster_info
cluster_data = json.loads(cluster_spec)
# Get primary node info
primary_node = cluster_data["cluster"]["workerpool0"][0]
logging.info("primary node: %s", primary_node)
primary_node_addr, primary_node_port = primary_node.split(":")
logging.info("primary node address: %s", primary_node_addr)
logging.info("primary node port: %s", primary_node_port)
# Determine node rank of this machine
workerpool = cluster_data["task"]["type"]
if workerpool == "workerpool0":
node_rank = 0
elif workerpool == "workerpool1":
# Add 1 for the primary node, since `index` is the index of workerpool1.
node_rank = cluster_data["task"]["index"] + 1
else:
raise ValueError(
"Only workerpool0 and workerpool1 are supported. Unknown workerpool:"
f" {workerpool}"
)
logging.info("node rank: %s", node_rank)
# Calculate total nodes.
num_nodes = 1 # For the primary node.
if "workerpool1" in cluster_data["cluster"]:
num_nodes += len(cluster_data["cluster"]["workerpool1"])
logging.info("num nodes: %s", num_nodes)
return ClusterInfo(
primary_node_addr, primary_node_port, node_rank, num_nodes
)
@@ -0,0 +1,59 @@
"""Add tests for cluster_spec.py."""
import os
from . import cluster_spec
# TODO(styer): Use pytest instead
class ClusterSpecTest(googletest.TestCase):
def setUp(self):
super().setUp()
self.curr_env_var = os.environ.copy()
def tearDown(self):
super().tearDown()
os.environ = self.curr_env_var
def test_get_cluster_spec_from_env_vars(self):
os.environ["CLUSTER_SPEC"] = ""
os.environ["MASTER_ADDR"] = "127.0.0.1"
os.environ["MASTER_PORT"] = "8080"
os.environ["RANK"] = "0"
os.environ["NNODES"] = "2"
cluster_info = cluster_spec.get_cluster_spec()
self.assertEqual(cluster_info.primary_node_addr, "127.0.0.1")
self.assertEqual(cluster_info.primary_node_port, "8080")
self.assertEqual(cluster_info.node_rank, 0)
self.assertEqual(cluster_info.num_nodes, 2)
def test_get_cluster_spec_from_cluster_spec(self):
os.environ[
"CLUSTER_SPEC"
] = """
{
"cluster": {
"workerpool0": [
"127.0.0.1:8080"
],
"workerpool1": [
"127.0.0.2:8080",
"127.0.0.3:8080"
]
},
"task": {
"type": "workerpool1",
"index": 0
}
}
"""
cluster_info = cluster_spec.get_cluster_spec()
self.assertEqual(cluster_info.primary_node_addr, "127.0.0.1")
self.assertEqual(cluster_info.primary_node_port, "8080")
self.assertEqual(cluster_info.node_rank, 1)
self.assertEqual(cluster_info.num_nodes, 3)
if __name__ == "__main__":
googletest.main()
@@ -66,7 +66,7 @@ mkdir -p "$local_folder"
mkdir -p "$output_folder"
# Download the content from the GCS URI
gsutil -m cp -r "$gcs_dataset_path"/* "$local_folder/"
gcloud storage cp --recursive "$gcs_dataset_path"/* "$local_folder/"
# Process files in the local folder
for file in "$local_folder"/*; do
@@ -122,23 +122,23 @@ cp -r "$output_folder" "$images_folder"/images_2
pushd "$images_folder"/images_2
ls | xargs -P 8 -I {} mogrify -resize 50% {}
popd
gsutil -m cp -r "$images_folder"/images_2/* "$gcs_experiment_path"/data/images_2
gcloud storage cp --recursive "$images_folder"/images_2/* "$gcs_experiment_path"/data/images_2
cp -r "$output_folder" "$images_folder"/images_4
pushd "$images_folder"/images_4
ls | xargs -P 8 -I {} mogrify -resize 25% {}
popd
gsutil -m cp -r "$images_folder"/images_4/* "$gcs_experiment_path"/data/images_4
gcloud storage cp --recursive "$images_folder"/images_4/* "$gcs_experiment_path"/data/images_4
cp -r "$output_folder" "$images_folder"/images_8
pushd "$images_folder"/images_8
ls | xargs -P 8 -I {} mogrify -resize 12.5% {}
popd
gsutil -m cp "$images_folder"/images_8/* "$gcs_experiment_path"/data/images_8
gcloud storage cp "$images_folder"/images_8/* "$gcs_experiment_path"/data/images_8
# Copy images and sparse reconstruction files to gcs experiment folder.
gsutil -m cp "$images_folder"/images/* "$gcs_experiment_path"/data/images
gsutil -m cp -r "$local_folder"/sparse "$gcs_experiment_path"/data
gsutil -m cp "$local_folder"/database.db "$gcs_experiment_path"/data
gcloud storage cp "$images_folder"/images/* "$gcs_experiment_path"/data/images
gcloud storage cp --recursive "$local_folder"/sparse "$gcs_experiment_path"/data
gcloud storage cp "$local_folder"/database.db "$gcs_experiment_path"/data
echo "Processing complete."
@@ -99,14 +99,14 @@ create_dir_if_not_exists "$CHECKPOINTS_PATH"
touch "$local_experiment_path/$exp_folder_name/log_render.txt"
# Copy experiment from GCS bucket to local
gsutil -m cp -r "${args[-gcs_experiment_path]}/data" "$local_experiment_path/$exp_folder_name" || exit 1
gsutil -m cp -r "${args[-gcs_experiment_path]}/checkpoints/${training_job_name}/*" "$CHECKPOINTS_PATH" || exit 1
gcloud storage cp --recursive "${args[-gcs_experiment_path]}/data" "$local_experiment_path/$exp_folder_name" || exit 1
gcloud storage cp --recursive "${args[-gcs_experiment_path]}/checkpoints/${training_job_name}/*" "$CHECKPOINTS_PATH" || exit 1
# Check and copy keyframes file.
if [[ -n ${args[-gcs_keyframes_file]} ]]; then
keyframes_file_basename=$(basename "${args[-gcs_keyframes_file]}")
local_keyframes_file="$local_dataset_path/$keyframes_file_basename"
gsutil cp "${args[-gcs_keyframes_file]}" "$local_keyframes_file" || exit 1
gcloud storage cp "${args[-gcs_keyframes_file]}" "$local_keyframes_file" || exit 1
echo "Local keyframe file: $local_keyframes_file"
launch_rendering "$local_keyframes_file"
else
@@ -114,4 +114,4 @@ else
fi
# Copy rendered data back to GCS.
gsutil -m cp -r "$OUTPUT_RENDER_PATH" "${args[-gcs_experiment_path]}/render/${rendering_job_name}"
gcloud storage cp --recursive "$OUTPUT_RENDER_PATH" "${args[-gcs_experiment_path]}/render/${rendering_job_name}"
@@ -74,7 +74,7 @@ create_dir_if_not_exists "$local_experiment_path"
create_dir_if_not_exists "$local_experiment_path/$scene_folder_name"
# Copy experiment from GCS bucket to local.
gsutil -m cp -r "${gcs_experiment_path}/data" "$local_experiment_path/$scene_folder_name" || exit 1
gcloud storage cp --recursive "${gcs_experiment_path}/data" "$local_experiment_path/$scene_folder_name" || exit 1
echo "GCS Experiment: $gcs_experiment_path"
echo "Gin Config File: $gin_config_file"
@@ -89,6 +89,6 @@ accelerate launch train.py --gin_configs="$gin_config_file" \
--gin_bindings="Config.factor = ${factor}" \
--gin_bindings="Config.max_steps = ${max_training_steps}"
gsutil -m rm -r "${gcs_experiment_path}/checkpoints/${training_job_name}"
gsutil -m cp -r "$local_experiment_path/$scene_folder_name/config.gin" "${gcs_experiment_path}/${training_job_name}_config.gin"
gsutil -m cp -r "$local_experiment_path/$scene_folder_name/checkpoints/*/*" "${gcs_experiment_path}/checkpoints/${training_job_name}"
gcloud storage rm --recursive "${gcs_experiment_path}/checkpoints/${training_job_name}"
gcloud storage cp --recursive "$local_experiment_path/$scene_folder_name/config.gin" "${gcs_experiment_path}/${training_job_name}_config.gin"
gcloud storage cp --recursive "$local_experiment_path/$scene_folder_name/checkpoints/*/*" "${gcs_experiment_path}/checkpoints/${training_job_name}"
@@ -102,10 +102,10 @@ def download_gcs_uri_to_local(
if not os.path.exists(destination_dir):
os.mkdir(destination_dir)
subprocess.check_output([
"gsutil",
"-m",
"gcloud",
"storage",
"cp",
"-r",
"--recursive",
gcs_uri,
destination_dir,
])
@@ -12,7 +12,7 @@ bitsandbytes==0.43.2
cloudml-hypertune==0.1.0.dev6
datasets==2.20.0
deepspeed==0.15.2
diffusers==0.25.1
diffusers==0.38.0
evaluate==0.4.3
fsspec==2024.3.1
gcsfs==2024.3.1
+11
View File
@@ -0,0 +1,11 @@
# Agent Platform Training Clusters Blog Series
This directory contains deep-dive documentation, extended guides, and architectural references for Google Cloud Agent Platform Training Clusters.
## Contents
- **`vertex-training-cluster/`**: Documentation and setup guides for configuring and managing Agent Platform Training Clusters.
## Blog Posts
- [Model Distillation Best Practices](https://googlecloudplatform.github.io/vertex-ai-samples/vertex-training-cluster/model_distillation_best_practices): Explores off-policy model distillation, dataset curation, and hyperparameter scaling laws for training student models on Vertex AI.
- [Forgetting Mitigation via Data Mixing](https://googlecloudplatform.github.io/vertex-ai-samples/vertex-training-cluster/forgetting_mitigation_data_mixing): Discusses catastrophic forgetting in model fine-tuning and how to mitigate it using multi-domain data mixing on Vertex AI.
- [Multi-Turn Reinforcement Learning for τ²-bench](https://googlecloudplatform.github.io/vertex-ai-samples/vertex-training-cluster/multi_turn_reinforcement_learning_for_tau2_bench): Explores multi-turn RL training for tool-calling agents using GRPO on the τ²-bench customer service benchmark with NeMo RL.
File diff suppressed because one or more lines are too long
@@ -0,0 +1,448 @@
<script type="text/javascript" async
src="https://cdn.mathjax.org/mathjax/latest/MathJax.js?config=TeX-MML-AM_CHTML">
</script><br><br>
# VTC Multi-Domain Dataset: Mitigating Catastrophic Forgetting with Data Mixing
**Author:** [Mayank Sharan](mailto:mayanksharan@google.com)
## Table of Contents
* [Intro](#intro)
* [Background](#background)
* [Dataset Curation](#dataset-selection)
* [Forgetting Mitigation Best Practices](#forgetting-mitigation-best-practices)
* [Experimental Setup](#experimental-setup)
* [Mitigating Forgetting](#mitigating-forgetting)
* [Mixing Ratios](#mixing-ratios)
* [Different Starting Models](#different-starting-models)
* [Acknowledgements](#acknowledgements)
* [References](#references)
## Intro
In this entry of our blog series on model training best practices for Vertex AI Training Cluster (VTC) customers, we talk about catastrophic forgetting and how to mitigate it. We focus on tuning public models using supervised fine tuning (SFT) with a specialized domain dataset. With both open and closed source models performing well on general tasks the primary goal of training one's own models is to improve the performance on specialized tasks. This typically comes at the cost of the model forgetting general capabilities which can severely limit the utility of the trained model.
There are many possible interventions to limit forgetting, the most effective is mixing the target dataset with the actual dataset used in the model’s training. Since this is not available even for the most open source models, we have curated a multi-domain dataset that delivers the same benefits. This allows Vertex AI Training Cluster (VTC) customers to maintain and surpass frontier level model capabilities while training to further performance on specialized tasks.
<figure align="center" id="fig-teaser">
<table align="center" width="80%">
<tr>
<td align="center" width="100%">
<img src="images_data_mixing/teaser_forgetting.png" width="100%"><br>
</td>
</tr>
</table>
<figcaption align="left">
<sub><b>Figure 1: Impact of mixing VTC Post Training dataset on Forgetting (8B model). </b> <i>Comparing SFT runs using only a specialized target dataset (MedMCQA) vs a mix of the target dataset and the VTC Post training dataset. Forgetting across all non-target domains is significantly mitigated with no performance loss on the target metric. Qwen3 Public here is the instruction tuned public Qwen3 8B model and the other two models are trained starting from the base Qwen3 8B model using only the target dataset and a mix of target dataset with the VTC dataset.</i></sub>
</figcaption>
</figure>
We provide a thorough set of experiments to serve as a guide for reducing forgetting while post training the Qwen3 open-weight thinking model family, beginning from their base pre-trained checkpoints. Furthermore, we demonstrate the value our datasets provide across model sizes often surpassing the performance of the official Qwen3 models while preserving performance on the specialized task (See [Figure 1](#fig-teaser)). The Qwen3 family was specifically chosen for this study because its diverse range of parameter counts and the availability of both pre-trained and post-trained checkpoints provide an ideal environment for high-fidelity scaling analysis.
To ensure our findings can be applied to a broad set of applications we validate our findings across five model sizes: 0.6B, 1.7B, 4B, 8B and 14B parameters. To support our VTC community in accelerating their own development, all code, datasets, and experiment configurations used in this blog are being made available for use in your training workloads.
## Background
Loss landscapes for neural networks have always been a complex multidimensional manifold rather than the simple convex ones that gradient descent is built for. Forgetting is a well known phenomenon in model customization, the first academically recorded instance being (McCloskey and Cohen, 1989) [<a href="#ref1">1</a>]. These manifolds have become even more complex with the introduction of Large Language Models where the number of parameters being optimized are typically in the billions. This makes it hard to mathematically grasp issues like forgetting. [Figure 2](#fig-loss-landscape) demonstrates a geometric understanding of why forgetting happens and how data mixing can mitigate it.
<figure align="center" id="fig-loss-landscape">
<table align="center" width="80%">
<tr>
<td align="center" width="100%">
<img src="images_data_mixing/background_loss_landscape.png" width="100%"><br>
</td>
</tr>
</table>
<figcaption align="left">
<sub><b>Figure 2: Geometric Interpretation of Data Mixing to Mitigate Forgetting. </b> <i>Fine tuning objectives being meaningfully out of distribution from the pre-trained model often drives forgetting. Mixing in a dataset similar to the model distribution adjusts the objective enough to learn the new task without as much forgetting.</i></sub>
</figcaption>
</figure>
## Dataset Selection
Our primary requirements for a target dataset to run experiments to validate this were:
1. It should be out of distribution to cause forgetting
2. It should have an evaluation metric that it directly improves
3. It should be able to train the model to perform better than the counterpart generalist model
A good heuristic to determine where the data lies with respect to the model distribution is by calculating perplexity on samples from the dataset. Assuming
- <span>$$X={x_1, x_2, \dots, x_N}$$</span> is a dataset sample represented as sequence of tokens
- <span>$$P(x_i \mid x_{<i})$$</span> is the model likelihood of the i-th token given the sample till that token
Then the perplexity for this sample can be calculated as follows:
$$\begin{align*}
& ppl(X) = \exp \left( -\frac{1}{N} \sum_{i=1}^{N} \log P(x_i \mid x_{<i}) \right) \\
& = \exp \left( -\frac{1}{N} \log (\prod_{i=1}^{N} P(x_i \mid x_{<i})) \right)
\end{align*} $$
The product form of the equation shows that this is a direct measure of the joint probability of this sequence of tokens according to the model. Since this computation has a balancing negative sign to account for the negative log value a lower joint probability results in a higher perplexity value and vice versa. We evaluated the following datasets as out-of-distribution candidates:
- [MedMCQA](https://huggingface.co/datasets/syz-ml2025/medmcqa) : Multiple Choice Questions (MCQ) dataset focusing on the medical domain
- [BirdSQL](https://huggingface.co/datasets/birdsql/bird23-train-filtered) : Text to SQL generation dataset
- [HardGen](https://huggingface.co/datasets/Bingguang/HardGen) : Function calling dataset
We also calculate perplexity on [OpenR1-Math-220k](https://huggingface.co/datasets/open-r1/OpenR1-Math-220k) to provide a reference as we expect this to be in distribution for the model given the Qwen3 models are particularly strong in the math domain.
<table id="tab-perplexity" style="margin-left:auto; margin-right:auto;">
<thead>
<tr>
<th>Dataset \ Model</th>
<th>Qwen3-0.6B</th>
<th>Qwen3-8B</th>
<th>Ours-0.6B</th>
<th>Ours-8B</th>
</tr>
</thead>
<tbody>
<tr>
<td>MedMCQA</td>
<td>63.00</td>
<td>66.00</td>
<td>42.00</td>
<td>20.75</td>
</tr>
<tr>
<td>BirdSQL</td>
<td>38.50</td>
<td>55.50</td>
<td>45.50</td>
<td>17.00</td>
</tr>
<tr>
<td>HardGen</td>
<td>2.23</td>
<td>2.03</td>
<td>2.28</td>
<td>1.79</td>
</tr>
<tr>
<td>OpenR1-Math</td>
<td>8.63</td>
<td>9.75</td>
<td>6.44</td>
<td>5.34</td>
</tr>
</tbody>
<caption style="text-align: left;"><b>Table 1:</b> Perplexity score analysis with the public instruction tuned Qwen3 models and Qwen3 base models trained using the VTC dataset (Ours) to identify a suitable target dataset.</caption>
</table>
We see from [Table 1](#tab-perplexity) that OpenR1-Math-220k as we expected has low perplexity scores and HardGen shows an even lower perplexity score eliminating it from consideration. MedMCQA samples have high perplexity scores across all considered models. This dataset also has the advantage of a straightforward evaluation metric as we can use the validation split in the form of an MCQ verified evaluation.
Based on this analysis we choose MedMCQA as our target dataset for these experiments. Additionally, since we are training a thinking model and the dataset does not have thinking traces we use the Qwen3-235B model to inject thinking traces into the training samples.
## Forgetting Mitigation Best Practices
### Experimental Setup
#### Dataset Mixing
We tested the impact of how forgetting responds to mixing the base dataset in different ratios with the target dataset. The base dataset here refers to the multi domain SFT dataset we have developed (see our [distillation blog post](https://googlecloudplatform.github.io/vertex-ai-samples/vertex-training-cluster/model_distillation_best_practices) [<a href="#ref2">2</a>] for details of the generation process) that can replicate and on certain metrics beat the public Qwen3 models. The target dataset here refers to the MedMCQA dataset. It is important to understand that in all mixing scenarios where the target dataset is present we will use the complete target dataset as that is the reasonable course of action we expect any customer to take. This leads to the total number of samples used in training varying based on the mixing ratio.
We run 2 baseline experiments for each model size: using only the base dataset and only the target dataset. The mixing experiments are the base dataset being mixed in ratios of 0.9:0.1, 0.75:0.25 and 0.5:0.5. (0.9:0.1 means 90% of samples are from the base in-distribution dataset, while 10% are from the target out-of-distribution dataset.)
The base dataset is randomly subsampled for each of these experiments. For simpler reference and analysis let’s define a mixing ratio <span>$$0 \le \alpha < 1$$</span>, such that the final dataset mixture includes <span>$$ N'_{B} = \frac{\alpha}{1 - \alpha} N_T$$</span> samples from the base dataset where <span>$$N_T$$</span> is the number of samples in the target dataset. In each of these mixtures the complete target dataset is used, contributing <span>$$N_T$$</span> samples for a total training dataset size of <span>$$\frac{N_T}{1 - \alpha}$$</span>.
Since, our target dataset has 182,712 samples, this means that:
- 0.9:0.1 ratio (<span>$$\alpha = 0.9$$</span>) : Uses a total of 1,827,120 training samples
- 0.75:0.25 ratio (<span>$$\alpha = 0.75$$</span>) : Uses a total of 730,849 training samples
- 0.5:0.5 ratio (<span>$$\alpha = 0.5$$</span>) : Uses a total of 365,425 training samples
#### Evaluation
<table id="tab-eval-setup" style="margin-left:auto; margin-right:auto;">
<thead>
<tr>
<th>Capabilities</th>
<th>Benchmarks</th>
<th># Test Samples</th>
<th>Eval Metrics</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="7">Math</td>
<td>AIME 24</td>
<td>30</td>
<td>pass@1 (average of 10)</td>
</tr>
<tr>
<td>AIME 25</td>
<td>30</td>
<td>pass@1 (average of 10)</td>
</tr>
<tr>
<td>BeyondAIME</td>
<td>100</td>
<td>pass@1 (average of 5)</td>
</tr>
<tr>
<td>Math 500</td>
<td>500</td>
<td>pass@1</td>
</tr>
<tr>
<td>HMMT 25</td>
<td>30</td>
<td>pass@1 (average of 10)</td>
</tr>
<tr>
<td>BRUMO 25</td>
<td>30</td>
<td>pass@1 (average of 10)</td>
</tr>
<tr>
<td>CMIMC 25</td>
<td>40</td>
<td>pass@1 (average of 10)</td>
</tr>
<tr>
<td rowspan="3">Science</td>
<td>GPQA</td>
<td>448</td>
<td>pass@1 (average of 5)</td>
</tr>
<tr>
<td>MMLU</td>
<td>14042</td>
<td>pass@1</td>
</tr>
<tr>
<td>MMLU Pro</td>
<td>12032</td>
<td>pass@1</td>
</tr>
<tr>
<td rowspan="2">Coding</td>
<td>HumanEval</td>
<td>164</td>
<td>pass@1 (average of 5)</td>
</tr>
<tr>
<td>LiveCodeBench v6</td>
<td>175</td>
<td>pass@1 (average of 5)</td>
</tr>
<tr>
<td>Instruction Following</td>
<td>IFEval</td>
<td>541</td>
<td>pass@1 (Strict Accuracy)</td>
</tr>
<tr>
<td>Reasoning</td>
<td>ARC-AGI 1</td>
<td>400</td>
<td>pass@1 (average of 5)</td>
</tr>
<tr>
<td>Medical (Target domain)</td>
<td>MedMCQA</td>
<td>4183</td>
<td>pass@1</td>
</tr>
</tbody>
<caption style="text-align: left;"><b>Table 2:</b> Comprehensive overview of task domains, evaluation benchmarks, and associated performance metrics.</caption>
</table>
Our evaluation benchmarks and metrics are detailed in [Table 2](#tab-eval-setup). To ensure statistical reliability on smaller datasets, we report metrics averaged over multiple independent runs to mitigate variance. For each domain with multiple evaluations, we utilize the average score across the core benchmarks as our primary performance indicator. To maintain a consistent comparison, both our trained models and the official Qwen3 thinking models were evaluated using standardized sampling parameters — `Temperature=0.6`, `Top-P=0.95`, `Top-K=20` and `Max-tokens=32768` — aligning with the recommended [best practices](https://huggingface.co/Qwen/Qwen3-14B#best-practices) from the official Qwen3 model card.
Note that we have separated MedMCQA as a target metric instead of including it in the Science domain. This is to ensure clear outcomes from our experiments and to demonstrate impacts on model performance without any interference.
#### Training
##### Vertex AI Training Cluster
All experiments and results presented were orchestrated using the [Vertex AI Training Cluster (VTC)](https://docs.cloud.google.com/vertex-ai/docs/training/training-clusters/overview). VTC is a managed Google Cloud service designed to simplify and accelerate large-scale AI workloads. It provides a simple managed user experience that enables optimized GPU scheduling, automated fault tolerance, high hardware resiliency, quick start recipes and science tooling which drastically reduces the time from cluster setup to production training and speeds up experimentation.
##### Training Framework and Hyperparameters
We utilize NVIDIA [NeMo RL](https://github.com/NVIDIA-NeMo/RL), an open library from the [NVIDIA NeMo framework](https://github.com/NVIDIA-NeMo/) as the primary training library, leveraging the Megatron backend for distributed scaling. Models are initialized from a Qwen3 Base checkpoint and fine-tuned with a 32,768 context window on curated datasets. Optimization is handled via AdamW (<span>$$\beta_1=0.9$$</span>, <span>$$\beta_2=0.95$$</span>, weight decay=0.1) using a linear warmup and cosine decay schedule. All training is conducted using BF16 mixed precision. There are many model sizes and dataset mixes used in the experimentation so the maximum learning rate is guided by learning rate scaling laws (see [distillation blog post](https://googlecloudplatform.github.io/vertex-ai-samples/vertex-training-cluster/model_distillation_best_practices#hyperparameter-scaling) [<a href="#ref2">2</a>] for more) available as a part of VTC. The value is validated by testing slight adjustments from the recommended value for each dataset mixture.
### Mitigating Forgetting
All models in this experiment are trained starting from the Qwen3 base checkpoint. We explore the impact of dataset mixing by comparing the public Qwen3 instruction tuned model performance with our two baselines — model trained with only the target dataset and model trained only with the base dataset — and with a model trained using a 0.9 ratio mix.
<figure align="center" id="fig3_data_mixing">
<table align="center" width="100%">
<tr>
<td align="center" width="50%">
<img src="images_data_mixing/fig3_math.png" width="100%"><br>
<sub><b>(a)</b> Math</sub>
</td>
<td align="center" width="50%">
<img src="images_data_mixing/fig3_science.png" width="100%"><br>
<sub><b>(b)</b> Science</sub>
</td>
</tr>
<tr>
<td align="center" width="50%">
<img src="images_data_mixing/fig3_coding.png" width="100%"><br>
<sub><b>(c)</b> Coding</sub>
</td>
<td align="center" width="50%">
<img src="images_data_mixing/fig3_ifeval.png" width="100%"><br>
<sub><b>(d)</b> IFEval</sub>
</td>
</tr>
<tr>
<td align="center" width="50%">
<img src="images_data_mixing/fig3_arc_agi.png" width="100%"><br>
<sub><b>(e)</b> ARC-AGI</sub>
</td>
<td align="center" width="50%">
<img src="images_data_mixing/fig3_medmcqa.png" width="100%"><br>
<sub><b>(f)</b> MedMCQA</sub>
</td>
</tr>
</table>
<figcaption align="left">
<sub><b>Figure 3: Performance with and without Data Mixing.</b> <i>A comparison across (a) Math, (b) Science, (c) Coding, (d) IFEval, (e) ARC-AGI and (f) MedMCQA benchmarks showing how data mixing impacts forgetting and performance on the target metric.</i></sub>
</figcaption>
</figure>
[Figure 3](#fig3_data_mixing) shows that for all non-target metrics other than Science using just the target dataset shows significant forgetting. Math and ARC-AGI are almost completely forgotten for all model sizes up to 8B parameters. The mixed dataset recovers the performance to similar levels as the base dataset. The base dataset delivers performance comparable to the public model in all domains and significantly better on ARC-AGI.
The Science domain evaluations do not suffer severe forgetting likely because MedMCQA is very close to this domain. In fact, for the 8B and 14B sizes due to these transfer learning dynamics the <span>$$\alpha = 0.9$$</span> model outperforms both the public instruction-tuned and the base dataset (<span>$$\alpha = 1$$</span>) models.
Performance on the target metric of MedMCQA follows expected behavior with best results achieved by the model when trained only with the target dataset. It is important to note that the <span>$$\alpha = 0.9$$</span> model for all sizes is still significantly better than the public instruction-tuned and base dataset (<span>$$\alpha = 1$$</span>) model and for all sizes other than the 0.6B mostly maintains the performance gains of the target dataset (<span>$$\alpha = 0$$</span>) model.
#### Key Observations
Combining these conclusions we can see that mixing with our base dataset:
- Matches and outperforms the public instruction tuned model on general tasks.
- Preserves the gains beyond the public model on target tasks.
- Provides additional gains on tasks from a similar domain.
### Mixing Ratios
Now that we know that mixing the base dataset almost eliminates forgetting it is important to understand how performance changes for different mixing configurations. This is also important to examine as it determines training length and hence the cost. We will compare models trained only with the target dataset to models trained using dataset mixes with <span>$$\alpha = 0.5, 0.75, 0.9$$</span>. The ratio mentioned here refers to the proportion of the dataset from the base dataset.
<figure align="center" id="fig4_mixing_ratios">
<table align="center" width="100%">
<tr>
<td align="center" width="50%">
<img src="images_data_mixing/fig4_math.png" width="100%"><br>
<sub><b>(a)</b> Math</sub>
</td>
<td align="center" width="50%">
<img src="images_data_mixing/fig4_science.png" width="100%"><br>
<sub><b>(b)</b> Science</sub>
</td>
</tr>
<tr>
<td align="center" width="50%">
<img src="images_data_mixing/fig4_coding.png" width="100%"><br>
<sub><b>(c)</b> Coding</sub>
</td>
<td align="center" width="50%">
<img src="images_data_mixing/fig4_ifeval.png" width="100%"><br>
<sub><b>(d)</b> IFEval</sub>
</td>
</tr>
<tr>
<td align="center" width="50%">
<img src="images_data_mixing/fig4_arc_agi.png" width="100%"><br>
<sub><b>(e)</b> ARC-AGI</sub>
</td>
<td align="center" width="50%">
<img src="images_data_mixing/fig4_medmcqa.png" width="100%"><br>
<sub><b>(f)</b> MedMCQA</sub>
</td>
</tr>
</table>
<figcaption align="left">
<sub><b>Figure 4: Performance across Mixing Ratios.</b> <i>A comparison across (a) Math, (b) Science, (c) Coding, (d) IFEval, (e) ARC-AGI and (f) MedMCQA benchmarks showing how dataset mixing ratios impact forgetting and performance on the target metric.</i></sub>
</figcaption>
</figure>
[Figure 4](#fig4_mixing_ratios) shows that for all non-target metrics mixing helps achieve better performance than just using the target dataset even with a <span>$$\alpha = 0.5$$</span> mix. As expected the performance on non target metrics worsens as we lower the ratio of the base dataset. This effect is more pronounced in the smaller size models and for datasets like ARC-AGI where the mixed training provides a lot more gain. These patterns confirm that the gains on non target metrics are directly correlated to the base dataset.
The effect while present for Science domain metrics is much less pronounced due to the cross domain characteristics. Even with lower ratios the performance for models 4B and larger holds, confirming that our target dataset of MedMCQA here contributes to limiting forgetting for this domain.
The performance on the target metric, MedMCQA, stays mostly consistent with dips mostly when going from <span>$$\alpha = 0.75$$</span> mix to <span>$$\alpha = 0.5$$</span> mix. This aligns well as in all cases we are doing a complete epoch on the target dataset. The performance mostly holding at mixing ratios indicates that the tradeoff on the target metrics is relatively low even at an aggressive mixing ratio like 0.5.
#### Key Observations
The mixing ratio comparison shows us that:
- A mixing ratio of 0.9 is the best for achieving gains on target tasks and limiting forgetting.
- A mixing ratio of even 0.5 limits forgetting well while only doubling the token budget compared to training without any mixing.
### Different Starting Models
We have trained all our models starting from Qwen3 base checkpoints. A natural question here might be: What happens if we train starting from the instruction tuned public Qwen3 checkpoints for our target task? In this section we examine this question and compare the instruction-tuned model tuned with the target dataset and an <span>$$\alpha = 0.9$$</span> mix to the instruction-tuned model itself and the base model tuned with an <span>$$\alpha = 0.9$$</span> mix.
<figure align="center" id="fig5_starting_models">
<table align="center" width="100%">
<tr>
<td align="center" width="50%">
<img src="images_data_mixing/fig5_math.png" width="100%"><br>
<sub><b>(a)</b> Math</sub>
</td>
<td align="center" width="50%">
<img src="images_data_mixing/fig5_science.png" width="100%"><br>
<sub><b>(b)</b> Science</sub>
</td>
</tr>
<tr>
<td align="center" width="50%">
<img src="images_data_mixing/fig5_coding.png" width="100%"><br>
<sub><b>(c)</b> Coding</sub>
</td>
<td align="center" width="50%">
<img src="images_data_mixing/fig5_ifeval.png" width="100%"><br>
<sub><b>(d)</b> IFEval</sub>
</td>
</tr>
<tr>
<td align="center" width="50%">
<img src="images_data_mixing/fig5_arc_agi.png" width="100%"><br>
<sub><b>(e)</b> ARC-AGI</sub>
</td>
<td align="center" width="50%">
<img src="images_data_mixing/fig5_medmcqa.png" width="100%"><br>
<sub><b>(f)</b> MedMCQA</sub>
</td>
</tr>
</table>
<figcaption align="left">
<sub><b>Figure 5: Performance across Starting Models.</b> <i>A comparison across (a) Math, (b) Science, (c) Coding, (d) IFEval, (e) ARC-AGI and (f) MedMCQA benchmarks showing how different starting models impact forgetting and performance on the target metric. Qwen 3 Public is the public instruction tuned Qwen3 model, &alpha;=0 (IT) and &alpha;=0.9 (IT) are the public instruction-tuned Qwen3 model trained only with the target dataset and the &alpha;=0.9 mixed dataset. &alpha;=0.9 (Base) is the base Qwen3 model trained on a 90% VTC dataset and 10% target dataset mix.</i></sub>
</figcaption>
</figure>
In [Figure 5](#fig5_starting_models), among the non-target metrics other than science we see a common trend that starting with the IT model and using only the target dataset (<span>$$\alpha = 0$$</span>) shows severe forgetting. The base model and the instruction-tuned model trained using the <span>$$\alpha = 0.9$$</span> mix match or surpass the performance of the public model. This shows that starting with an instruction-tuned model while better than starting with the base model is still not a solution to forgetting. This also shows the high quality of our dataset that it can provide further gains on the public instruction-tuned model.
Science domain metrics show different trends based on the model size. The advantage of data mixing is much more apparent in 0.6B and 1.7B models. Overall though there are no disadvantages to mixing across all model sizes. The IT model demonstrating significant forgetting is a clear indication that cross domain characteristics of our target dataset are not enough to mitigate forgetting on its own.
The performance of the target metric, MedMCQA, shows no additional gain when we train using only the target dataset except for the 0.6B model, whether the starting model is a base model or the IT model. For all model sizes other than the 0.6B model we also see that the <span>$$\alpha = 0.9$$</span> mix trained model does not lose any meaningful performance compared to the target dataset only trained models. All the models trained using the target dataset clearly improve on the public model.
#### Key Observations
The comparison of different starting models shows us:
- Using the instruction-tuned model as the starting model is better than the Base model.
- The IT model also shows catastrophic forgetting and loses performance on non target metrics.
- The <span>$$\alpha = 0.9$$</span> mix avoids forgetting even with the instruction-tuned starting model showing its robustness.
## Acknowledgements
We would like to express our sincere gratitude to the NVIDIA NeMo RL team–specifically Terry Kong– for their invaluable support throughout this project.
We would also like to express our gratitude to our VTC teammates: Mohammadreza Mohseni, Weiran Zhao, Fei Xia, Youbao Tang, Xuehan Xiong, Joseph Pagadora, Jiuqiang Tang, Bo Wu, Lav Rai, and Minwoo Park for developing the underlying datasets, providing infrastructure support, feedback, and insightful discussions throughout the project. We also thank Ting Yu, Shengyang Dai, Peng Xu, and Saurabh Tiwary for their leadership and support.
## References
<a id="ref1"></a>[1] McCloskey, Michael, and Neal J. Cohen. "Catastrophic interference in connectionist networks: The sequential learning problem." Psychology of learning and motivation. Vol. 24. Academic Press, 1989. 109-165.
<a id="ref2"></a>[2] Google Cloud. "Model Distillation Best Practices." Vertex AI Training Cluster Samples. Google, 2026. https://googlecloudplatform.github.io/vertex-ai-samples/vertex-training-cluster/model_distillation_best_practices.
Binary file not shown.

After

Width:  |  Height:  |  Size: 9.0 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 5.6 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 48 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 259 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 295 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 40 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 8.1 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 190 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 189 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 154 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 195 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 195 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 191 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 184 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 152 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 167 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 158 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 177 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 180 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 193 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 229 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 230 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 232 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 232 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 234 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 241 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 231 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 219 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 202 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 234 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 270 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 230 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 262 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 211 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 209 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 189 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 215 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 210 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 212 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 206 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 190 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 184 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 194 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 213 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 214 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 219 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 107 KiB

Some files were not shown because too many files have changed in this diff Show More