Compatible MiniMax Models
You can import large language models from Hugging Face and OCI Object Storage buckets into OCI Generative AI, create endpoints for those models, and use them in the Generative AI service.
MiniMax M3
The MiniMax-M3-MXFP8 model is the MXFP8 quantized variant of MiniMax M3, a native multimodal model with a one million token context. The model has about 428 billion total parameters with about 23 billion activated parameters and uses MiniMax Sparse Attention (MSA) for efficient long-context processing. This model is optimized for coding, long-horizon agentic workflows, and collaborative productivity tasks.
| Hugging Face Model ID | Model Capability | Minimum Dedicated AI Cluster Unit Shape |
|---|---|---|
| MiniMaxAI/MiniMax-M3-MXFP8 | TEXT_TO_TEXT |
|
MiniMax M2
The MiniMax M2 text-to-text models are optimized for coding, complex reasoning, and agentic workflows such as tool use, search, and productivity tasks. MiniMax-M2 is a Mixture-of-Experts (MoE) model designed for efficient coding and agentic performance, and later MiniMax-M2 models extend this focus to more advanced software engineering and professional-work tasks. For more details, see MiniMax in the Hugging Face documentation.
| Hugging Face Model ID | Model Capability | Minimum Dedicated AI Cluster Unit Shape |
|---|---|---|
| MiniMaxAI/MiniMax-M2.7 | TEXT_TO_TEXT |
|
| MiniMaxAI/MiniMax-M2.5 | TEXT_TO_TEXT |
|
| MiniMaxAI/MiniMax-M2 | TEXT_TO_TEXT |
|
- For imported models, you can use the native context length specified by the model provider. However, the effective maximum context length is limited by the underlying hardware setup that you select for the hosting dedicated AI clusters in OCI Generative AI. To take full advantage of a model's native context length, you might need to provision more hardware resources.
- Use the fine-tuned models only if they match the compatible base model's transformer version and have a parameter count within ±10% of the original.
- For available hardware shapes, see Hardware Unit Shapes for Imported Models.
- For steps on how to deploy the imported models, see Managing Imported Models.
Billing for Dedicated AI Clusters Hosting Imported Models
Dedicated AI clusters that host imported models don't require the 744-unit-hour minimum commitment that applies to dedicated AI clusters hosting pretrained models available in OCI Generative AI. Instead, an imported-model dedicated AI cluster has a minimum billable duration of one hour. For example, a cluster that uses one AI unit and runs for 30 minutes is billed for one AI unit-hour.
After the first hour, charges are based on the cluster's actual running time and are prorated to the millisecond. For example, a cluster that runs for one hour and one minute is billed for that duration, rather than for two full hours.
OCI sends usage records to Metering and Billing every five minutes while the cluster is running. If the cluster is deleted before the one-hour minimum is reached, OCI continues reporting billable usage until the minimum is met. If the cluster runs for more than one hour, OCI continues sending usage records every five minutes until the cluster is deleted.