Model Overview

Model Details

To view model details, please visit the Model Marketplace.

Text Models

Text models are designed to process and generate natural language, with capabilities in language understanding and reasoning. They can automatically process large volumes of text and perform logical inference. Zhipu's text models combine powerful language generation and reasoning capabilities, enabling them not only to understand and generate text but also to perform advanced reasoning and judgment.

Vision Models

Vision models process visual information such as images and videos and are widely used for recognition, analysis, and decision-making tasks. Visual understanding models focus on interpreting image content, such as identifying objects, scenes, and relationships. Visual reasoning models go a step further by combining visual and language information to perform logical judgment, causal analysis, and multi-step reasoning. They are commonly used for visual question answering, image captioning, multimodal alignment, and other complex tasks.

Image Generation Models

Image generation models learn from large-scale image datasets to generate high-quality images from text prompts. They are widely used in visual content creation, game art, product design, medical image synthesis, and other fields.

Video Generation Models

Video generation models learn from temporal visual data to generate dynamic video content from text, images, or other video materials. They are widely used in film production, virtual humans, animation generation, digital marketing, and other fields.