Deploy Gemma with Google Cloud  |  Google AI for Developers Skip to main content Models Gemini About Docs API reference Pricing Imagen About Docs Pricing Veo About Docs Pricing Gemma About Docs Gemmaverse Solutions Build with Gemini Gemini API Google AI Studio Customize Gemma open models Gemma open models Multi-framework with Keras Fine-tune in Colab Run on-device Google AI Edge Gemini Nano on Android Chrome built-in web APIs Build responsibly Responsible GenAI Toolkit Secure AI Framework Code assistance Android Studio Chrome DevTools Colab Firebase Google Cloud JetBrains Jules VS Code Community Google AI Forum Gemini for Research / English Deutsch Español – América Latina Français Indonesia Italiano Polski Português – Brasil Shqip Tiếng Việt Türkçe Русский עברית العربيّة فارسی हिंदी বাংলা ภาษาไทย 中文 – 简体 中文 – 繁體 日本語 한국어 Sign in Gemma Gemma Docs Models More Gemma Docs Solutions More Code assistance More Community More Overview Get started Releases Models Core Gemma Overview Gemma 4 model card Gemma 3 model card Gemma 2 model card Gemma 1 model card Core Variants Gemma 3n Overview Model card DiffusionGemma Overview Model card Diffusion Explained Generate Output FunctionGemma Overview Model card Formatting and best practices Function calling with Hugging Face Transformers Full function calling sequence with FunctionGemma Fine-tune FunctionGemma EmbeddingGemma Overview Model card Generate embeddings with Sentence Transformers Fine-tune EmbeddingGemma PaliGemma Overview v2 model card v1 model card Generate output with Keras Fine-tune with JAX and Flax Prompt and system instructions ShieldGemma Overview ShieldGemma 2 Model card ShieldGemma 1 Model card Run Gemma Fundamentals Overview Prompt Formatting Legacy Gemma setup [Gemma 1, 2, and 3] Legacy Prompt and system instructions [Gemma 1, 2, and 3] Run locally with a Chat UI or integrate via API LM Studio Ollama Run efficiently on Edge LiteRT-LM Llama.cpp MLX Build/Train in Python Tunix (Tune-in-JAX) Hugging Face Transformers Keras Unsloth Deploy to Production / Enterprise Gemini API Google Cloud Cloud GKE Multi-Token Prediction (MTP) Overview Hugging Face Transformers Core Capabilities Text Basic and multi-turn chat Function calling Visual data Overview Image understanding Video understanding Audio data Thinking Tuning guides Overview Tune using Hugging Face Transformers and QLoRA Vision Tune using Hugging Face Transformers and QLoRA Full model fine-tune using Hugging Face Transformers Tune using Gemma library Research and tools RecurrentGemma Overview Inference using JAX and Flax Fine-tune using JAX and Flax Model card DataGemma Gemma Scope Gemma-APS Community Gemmaverse Discord Legal Terms of use Gemma 4 license Prohibited use Intended use statement Gemini About Docs API reference Pricing Imagen About Docs Pricing Veo About Docs Pricing Gemma About Docs Gemmaverse Build with Gemini Gemini API Google AI Studio Customize Gemma open models Gemma open models Multi-framework with Keras Fine-tune in Colab Run on-device Google AI Edge Gemini Nano on Android Chrome built-in web APIs Build responsibly Responsible GenAI Toolkit Secure AI Framework Android Studio Chrome DevTools Colab Firebase Google Cloud JetBrains Jules VS Code Google AI Forum Gemini for Research Gemma 4 released with text, audio and image input and long up to 256K context window! Learn more Home Gemma Models Docs Send feedback Deploy Gemma with Google Cloud The Google Cloud platform provides many options for deploying, serving, and fine-tuning Gemma 4 open models, including the following: Gemini Enterprise Agent Platform Cloud Run Google Kubernetes Engine (GKE) Agent Development Kit (ADK) Gemini Enterprise Agent Platform Training Clusters MaxText vLLM with TPUs Sovereign Cloud Gemini Enterprise Agent Platform Gemini Enterprise Agent Platform is a Google Cloud platform for rapidly building and scaling machine learning projects. Gemma 4 is available in Model Garden, a curated collection of models on Gemini Enterprise Agent Platform. You can test and deploy models directly from the console. To learn more, refer to the following pages: Agent Platform overview: Get started with Gemini Enterprise Agent Platform. Gemma with Gemini Enterprise Agent Platform: Use Gemma open models with Gemini Enterprise Agent Platform. Cloud Run Cloud Run is a fully managed platform to run your code or containers on top of Google's highly scalable infrastructure. Deploy Gemma 4 on Cloud Run using GPUs for scale-to-zero, pay-per-use inference. For larger mode sizes, leverage advanced configurations with RTX 6000 Pro GPUs and Model Streaming. Google Kubernetes Engine (GKE) Google Kubernetes Engine (GKE) is a managed Kubernetes service from Google Cloud. Run Gemma 4 on GKE for enterprise-grade container orchestration. Use TPUs and GPUs to serve models with high throughput and low latency. Agent Development Kit (ADK) Build and orchestrate AI agents with Gemma 4 and the Agent Development Kit (ADK). Gemma 4's strong reasoning and function-calling capabilities make it ideal for agentic workflows. Gemini Enterprise Agent Platform Training Clusters Fine-tune Gemma 4 using Gemini Enterprise Agent Platform Training Clusters. Training Clusters provides optimized infrastructure for large-scale training and fine-tuning of open models. vLLM with TPUs Serve Gemma 4 on Google Cloud TPUs for state-of-the-art serving performance. MaxText Gemma 4 is supported in MaxText, a high-performance, arbitrary-sized JAX LLM implementation for Google Cloud TPUs. Sovereign Cloud Gemma 4 is available on Sovereign Cloud solutions, providing enhanced control and compliance for sensitive workloads. Send feedback Except as otherwise noted, the content of this page is licensed under the Creative Commons Attribution 4.0 License, and code samples are licensed under the Apache 2.0 License. For details, see the Google Developers Site Policies. Java is a registered trademark of Oracle and/or its affiliates. Last updated 2026-07-02 UTC. Need to tell us more? [[["Easy to understand","easyToUnderstand","thumb-up"],["Solved my problem","solvedMyProblem","thumb-up"],["Other","otherUp","thumb-up"]],[["Missing the information I need","missingTheInformationINeed","thumb-down"],["Too complicated / too many steps","tooComplicatedTooManySteps","thumb-down"],["Out of date","outOfDate","thumb-down"],["Samples / code issue","samplesCodeIssue","thumb-down"],["Other","otherDown","thumb-down"]],["Last updated 2026-07-02 UTC."],[],[]] Terms Privacy Manage cookies English Deutsch Español – América Latina Français Indonesia Italiano Polski Português – Brasil Shqip Tiếng Việt Türkçe Русский עברית العربيّة فارسی हिंदी বাংলা ภาษาไทย 中文 – 简体 中文 – 繁體 日本語 한국어