Skip to main content
POST

Model selection

Both models share the same capabilities and parameters and generate only 1 image per request. Refer to model pricing for actual prices.
This endpoint is asynchronous. Retrieve the task ID from data[0].task_id after submission, then use task queries for results. Poll every 3–5 seconds with an overall wait timeout of 3 minutes. Stop when the status is completed or failed.

Request headers

string
required
Bearer authentication in the format Bearer <token>, where <token> is your APIMart API Key.

Request parameters

string
required
Model ID: mai-image-2.6 or mai-image-2.6-flash.
string
required
Image description or editing instructions. Supports Chinese and English, up to approximately 32,000 tokens (not characters).
string
default:"1:1"
Accepts an aspect ratio (such as 16:9), pixel dimensions (such as 1536x1024), or auto.
  • Aspect ratio: any integer ratio from 1:4 to 4:1, used together with resolution.
  • Pixel dimensions: accepts widthxheight, width*height, or width×height. In this mode, resolution does not determine dimensions.
  • auto: the model selects an aspect ratio based on the prompt.
Text-to-image only. When reference images are provided, the model determines output dimensions.
string
default:"1K"
Supports 1K and 2K, including lowercase. Other tiers such as 4K are unsupported and return HTTP 400.For text-to-image with an aspect ratio, this parameter selects the size tier. It does not determine dimensions when exact pixels are used. It cannot set image-to-image output dimensions.
integer
Exact pixel width. Must be provided together with height. This pair takes precedence over size and resolution for text-to-image dimensions.Both width and height must be at least 768, with no more than 2,359,296 total pixels. Use multiples of 32; otherwise each output dimension is rounded down to a multiple of 32.This parameter does not determine image-to-image output dimensions.
integer
Exact pixel height. Must be provided together with width and follows the constraints above. Does not determine image-to-image output dimensions.
string[]
Reference image list, up to 5 images. Omit for text-to-image; provide one for single-image editing or multiple for composition.Each entry supports a publicly accessible HTTP(S) image URL or a Base64 Data URL such as data:image/png;base64,....JPEG and PNG are supported; WEBP and GIF are automatically converted to PNG. Image URLs must be publicly accessible or the task will fail.Image-to-image output dimensions are determined by the model from the references, at about 1 million pixels with a similar aspect ratio. size, resolution, width, and height cannot set these dimensions.
boolean
default:"false"
Set to true to let the model choose an aspect ratio based on the prompt, equivalent to size: "auto".
boolean
default:"false"
Set to true to retrieve real-time information before generation, useful for images involving real people, places, or events.
integer
default:"1"
Only 1 is supported. Submit separate tasks for multiple images. Values above 1 return HTTP 400.

Text-to-image dimensions

Text-to-image dimension precedence: paired width / height → exact-pixel size → aspect-ratio size combined with resolution.

Resolution tiers and aspect ratios

Dimensions are converted to multiples of 32. Since the shorter side must be at least 768, extreme aspect ratios can exceed approximately 1 million pixels even at 1K. Billing uses the token count corresponding to the actual output pixels.

Exact-pixel constraints

  • Both width and height must be at least 768.
  • Width × height must not exceed 2,359,296 (1536 × 1536).
  • Each dimension is rounded down to a multiple of 32. For example, 1000x1000 produces 992x992. Use multiples of 32 for exact dimensions.
Dimensions such as 1536x1024, 2048x1152, and 3072x768 are supported. 512x512 is rejected because the sides are too small; 2048x2048 exceeds the total pixel limit.
The maximum applies to total pixels, not a 1536 limit on each side. Therefore, 2048x1152 and 3072x768 are valid, but the 4K tier is unsupported. These dimension settings apply only to text-to-image.

Request examples

Exact pixels and web grounding

Single-image editing

Multi-image composition

Replace the example image URLs with accessible URLs.

Unsupported parameters

quality, style, background, output_format, response_format, and mask_url are unsupported and ignored if provided. Output is always PNG. Mask-based editing is not supported.

Submission response

integer
Response status code. 200 indicates success.
array
Task submission result.

Query task results

Example success response (the image URL is a placeholder):

Billing

Billed by actual input and output token usage. Refer to model pricing for unit prices.
  • Image output tokens = actual output width × height ÷ 1024. For example, 1024×1024 corresponds to 1024 tokens; 1536×1536 to 2304 tokens.
  • Input tokens per reference image are approximately its width × height ÷ 1024. Text prompts also count toward input usage.
  • An amount based on the tier is charged upfront, then adjusted after success to actual token usage with a refund or additional charge.
  • Failed tasks receive an automatic full refund. Parameter errors rejected at submission create no task and incur no charge.

Common errors

If a task fails, check image download or content safety errors. Revise the prompt or references before retrying. Editing photorealistic images involving minors may be blocked by content safety policies.