Concepts

Providers & Models

The core registry contains eight provider lanes: OpenAI OAuth and API, Grok OAuth and API, Gemini/Antigravity CLI, direct Gemini API, AtlasCloud API, and MiniMax API. Runway and Higgsfield are separate MCP-backed integrations described below.

Provider paths

  • provider: "oauth" — uses the local Codex OAuth proxy. The default path; no API key needed.
  • provider: "api" — calls the OpenAI Responses API with the hosted image_generation tool. Requires OPENAI_API_KEY.
  • provider: "grok" — starts bundled progrok, runs mandatory xAI Web Search and a planner pass (default grok-4.5, configurable in settings or via --planner-model), then calls xAI Images API. Grok 4.3 remains an explicit compatibility override.
  • provider: "grok-api" — same xAI pipeline as grok but uses a direct XAI_API_KEY instead of the OAuth proxy.
  • provider: "gemini-api" — calls the Google Gemini image API directly. Supports two models (nano-banana-2 and nano-banana-pro), aspect ratio and resolution controls, and two auth modes: a GEMINI_API_KEY or a Vertex AI service-account JSON (VERTEX_SERVICE_ACCOUNT_JSON). When both are configured, Vertex AI takes priority unless overridden by the last-saved auth mode. Cost varies by model and resolution (see Gemini API section below).
  • provider: "agy" — spawns the Antigravity CLI (agy -p) to generate via Google Gemini (nano-banana-2). Fixed 1024×1024 JPEG output, max 3 refs. Free (no token cost).
  • provider: "atlascloud" — calls AtlasCloud with ATLASCLOUD_API_KEY for GPT Image 2 generation and edit, with up to 10 references.
  • provider: "minimax" — calls MiniMax with MINIMAX_API_KEY for image-01 and image-01-live, with one reference.

The core lanes cover Classic, Node, and Agent Mode. Agent Mode is web-UI only; Classic and Node also have CLI commands.

Per-request override

Generation commands use the eight explicit core lane IDs below. Where a command still accepts auto, it is a routing mode, not a ninth provider lane.

ValueBehavior
autoRouting mode where accepted; not a provider lane.
oauthForce the local OAuth proxy path.
apiForce the API-key Responses path; requires a configured key.
grokForce the bundled xAI path through 127.0.0.1:18645; run ima2 grok login once to authorize.
grok-apiSame xAI pipeline but authenticates with a direct XAI_API_KEY.
gemini-apiDirect Google Gemini API. Requires GEMINI_API_KEY or VERTEX_SERVICE_ACCOUNT_JSON. Supports aspect ratio and resolution controls. No quality/format/moderation/multimode controls.
agySpawn Antigravity CLI for Gemini image generation. Requires agy binary installed. 1024×1024 fixed, JPEG, max 3 refs, no quality/size/mask controls.
atlascloudDirect AtlasCloud API. Requires ATLASCLOUD_API_KEY.
minimaxDirect MiniMax API. Requires MINIMAX_API_KEY.

Models

The app defaults to gpt-5.6-luna for image generation and Prompt Builder planning. Older supported models remain explicit compatibility choices.

ModelUse
gpt-5.6-lunaCurrent image and Prompt Builder default.
gpt-5.6-terra / gpt-5.6-solCurrent GPT-5.6 alternatives when available.
gpt-5.4-miniCompatibility model.
gpt-5.4Recommended balanced choice.
gpt-5.5Strongest quality when your Codex CLI/OAuth backend supports it. May use more quota or need an updated Codex CLI.
grok-imagine-imageCompatible fast Grok image model.
grok-imagine-image-qualityDefault Grok image model ("Grok+" / Best in the UI).
grok-imagine-videoBase model for Ref2V, edit, and extension compatibility paths.
grok-imagine-video-1.5Default Grok video model for prompt-only T2V and single-image/frame I2V, including 1080p when supported.
nano-banana-2Gemini Flash image model (maps to gemini-3.1-flash-image). Used by both gemini-api and agy providers.
nano-banana-proGemini Pro image model (maps to gemini-3-pro-image). Available on the gemini-api provider only.

The app also exposes quality (low, medium, high) and moderation (auto, low) controls. Reasoning effort accepts none, low, medium, high, and xhigh.

Persisted defaults. ima2 defaults set model gpt-5.5 and ima2 defaults set reasoning high write both OAuth and API-provider default keys, so your "default model" stays one concept across provider paths.

Gemini API provider

The gemini-api provider calls Google's image generation API directly without the Antigravity CLI. It supports two models selectable in the UI or via --model:

  • nano-banana-2 (default) — Gemini Flash; faster, lower cost.
  • nano-banana-pro — Gemini Pro; higher quality, higher cost.

Auth modes. Configure either a Gemini API key (GEMINI_API_KEY env var or via the Settings UI) or a Google Cloud service-account JSON (VERTEX_SERVICE_ACCOUNT_JSON env var or via the Settings UI Vertex JSON input). When both are present the server respects the last-saved mode (geminiAuthMode: "apikey" or "vertex" in config), defaulting to Vertex when both are configured and no mode is saved.

Aspect ratio and resolution. The direct API path exposes 10 aspect ratios (1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9) and 4 resolution tiers (512px, 1K, 2K, 4K). Selections are mapped to exact pixel dimensions and passed as aspect_ratio and image_size protobuf enum values. The Vertex AI path ignores these controls — the Vertex endpoint does not accept the response_format field, so output defaults to 1K / 1:1 regardless.

Per-image cost estimates (based on output token counts × official rates):

Model512px1K2K4K
nano-banana-2 (Flash 3.1, $60/1M tok)$0.045$0.067$0.101$0.151
nano-banana-pro (Pro 3, $120/1M tok)$0.134$0.134$0.134$0.240

The agy provider (Antigravity CLI) uses nano-banana-2 only, at fixed 1024×1024 with no resolution or aspect controls, and is estimated as free (no API token charge in the cost estimator).

Grok pipeline

Grok Classic, Node, and Agent requests run a three-step pipeline: mandatory xAI Web Search, planner pass (default grok-4.5, overridable via IMA2_GROK_PLANNER_MODEL, settings UI, or --planner-model on video commands) with an English final image prompt, then xAI image creation. Grok 4.3 remains available as a compatibility override. Text-only requests use /v1/images/generations; requests with reference images, a Node parent image, or an Agent current image use /v1/images/edits so image-to-image context is preserved. Grok accepts up to three total input images in this path.

ima2 maps OpenAI-style sizes to xAI aspect_ratio and resolution controls. Grok mask edit is not wired in this release and returns GROK_MASK_UNSUPPORTED.

Model and size pickers. The UI exposes a two-button image model picker ("Grok" / Fast = grok-imagine-image; "Grok+" / Best = grok-imagine-image-quality) and a size picker with native xAI aspect_ratio and resolution (1k/2k) values.

Billing and quota. When Grok is authorized, GET /api/quota returns a grok object with a monthly usage bar and a billing field (usedUsd / limitUsd) displayed as "$used/$limit" in the QuotaCard header (e.g. "$134.80/$1500.00").

Switch Account. The QuotaCard exposes a "Switch Account" button for Grok that starts a server-side xAI device-code flow (POST /api/auth/switchGET /api/auth/switch/:sessionId). The button opens the verification URL in a new tab, displays the user code, and the server polls until complete. The same flow is available for Codex/GPT OAuth.

Grok video

Grok video generation defaults to canonical grok-imagine-video-1.5 ("Grok V1.5"). The base grok-imagine-video model remains available for Ref2V, edit, and extension. A two-button video model picker at the top of the video controls panel lets you switch between them. Three modes are auto-detected from reference count: text-to-video (0 refs), image-to-video (1 ref), and reference-to-video (2–7 refs, max 10s). Controls include duration (1–15s), resolution (480p, 720p, and 1080p for 1.5 single-image/frame I2V), and aspect ratio. The old grok-imagine-video-1.5-preview value is accepted as a compatibility alias. 1.5 does not add Ref2V, V2V edit, or extension support, so those routes remain base-model only. Choose the planner model in video settings or pass --planner-model on CLI video commands. The endpoint POST /api/video/generate streams SSE events: planning → submitted → progress → done. From the CLI: ima2 video "prompt" --model grok/grok-imagine-video-1.5 --duration 5 --resolution 720p (or set a persistent target once with ima2 defaults set video grok/grok-imagine-video-1.5).

MCP lanes (Runway, Higgsfield)

Since 3.0.0 the CLI routes two additional lanes through remote MCP providers: runway (Gen-4/4.5, Veo 3.1, Seedance 2, Kling — image and video) and higgsfield (catalog-only until a paid plan unlocks execution). Inspect live lane status and per-model capabilities with ima2 models, then target them like any other lane: ima2 video "prompt" --model runway/veo-3.1 --duration 8. MCP jobs are asynchronous — the CLI submits to POST /api/mcp/generate and waits on the shared SSE event stream until the file lands in the gallery.

API-provider defaults

When the API path is used without explicit options, these defaults apply:

VariableDefault
IMA2_API_IMAGE_MODEL_DEFAULTgpt-5.6-luna
IMA2_API_REASONING_EFFORTlow
IMA2_API_IMAGE_SIZE1024x1024
IMA2_API_ALLOW_WEB_SEARCHtrue

See Configuration for the full environment table.