Compatible Google Models
You can import large language models from Hugging Face and OCI Object Storage buckets into OCI Generative AI, create endpoints for those models, and use them in the Generative AI service.
MedGemma
Google MedGemma 27B Text IT is a 27-billion-parameter language model based on the Gemma 3 architecture. Trained on medical literature and clinical records, the model is instruction-tuned (IT), enabling it to better follow user requests and help perform healthcare-focused tasks such as medical question answering, clinical support, and summarizing.
| Hugging Face Model ID | Model Capability | Minimum Dedicated AI Cluster Unit Shape |
|---|---|---|
| google/medgemma-27b-text-it | TEXT_TO_TEXT |
|
Gemma
| Hugging Face Model ID | Model Capability | Minimum Dedicated AI Cluster Unit Shape |
|---|---|---|
| google/gemma-4-12B-it | IMAGE_TEXT_TO_TEXT |
|
| google/gemma-4-26B-A4B | IMAGE_TEXT_TO_TEXT |
|
| google/gemma-4-31B-it | IMAGE_TEXT_TO_TEXT |
|
| google/gemma-3-4b-it | IMAGE_TEXT_TO_TEXT | A100_80G_X1 |
| google/gemma-3-270m-it | TEXT_TO_TEXT | A100_80G_X1 |
| google/gemma-3-27b-it | IMAGE_TEXT_TO_TEXT | A100_80G_X2 |
| google/gemma-3-12b-it | IMAGE_TEXT_TO_TEXT |
|
| google/gemma-3-1b-it | TEXT_TO_TEXT | A100_80G_X1 |
| google/gemma-2-9b-it | TEXT_TO_TEXT | A100_80G_X1 |
| google/gemma-2-27b-it | TEXT_TO_TEXT | A100_80G_X2 |
| google/gemma-2-2b-it | TEXT_TO_TEXT | A100_80G_X1 |
- For imported models, you can use the native context length specified by the model provider. However, the effective maximum context length is limited by the underlying hardware setup that you select for the hosting dedicated AI clusters in OCI Generative AI. To take full advantage of a model's native context length, you might need to provision more hardware resources.
- Use the fine-tuned models only if they match the compatible base model's transformer version and have a parameter count within ±10% of the original.
- For available hardware shapes, see Hardware Unit Shapes for Imported Models.
- For steps on how to deploy the imported models, see Managing Imported Models.
Billing for Dedicated AI Clusters Hosting Imported Models
Dedicated AI clusters that host imported models don't require the 744-unit-hour minimum commitment that applies to dedicated AI clusters hosting pretrained models available in OCI Generative AI. Instead, an imported-model dedicated AI cluster has a minimum billable duration of one hour. For example, a cluster that uses one AI unit and runs for 30 minutes is billed for one AI unit-hour.
After the first hour, charges are based on the cluster's actual running time and are prorated to the millisecond. For example, a cluster that runs for one hour and one minute is billed for that duration, rather than for two full hours.
OCI sends usage records to Metering and Billing every five minutes while the cluster is running. If the cluster is deleted before the one-hour minimum is reached, OCI continues reporting billable usage until the minimum is met. If the cluster runs for more than one hour, OCI continues sending usage records every five minutes until the cluster is deleted.