MiniMax H3 Model Enhances Multimodal Generative Capabilities
The MiniMax H3 is a general-purpose, omni-modal generative system developed by MiniMaxAI, designed to handle multimodal contexts including text, images, video, and audio. It can generate video with stereo audio at resolutions up to 2K and durations of up to 15 seconds. The system is built to understand and generate complex multimodal instructions, offering broad capabilities at the pre-training stage. The model supports various input and output specifications, including different aspect ratios and resolutions, and is capable of processing multiple languages. The system consists of three main modules: H3-Context-IR, H3-Base, and H3-Regenerate-2K, each contributing to the model's ability to produce high-quality outputs. The model is not yet open-sourced, but an API is available for validating results.