Grok Imagine is xAI’s AI video generation model for creating high-quality videos with realistic motion, native audio, and synchronized speech. Grok Imagine 1.5 introduces faster generation speeds, improved motion physics, and stronger visual consistency for modern AI video applications.
No image yet. Submit the form to generate image.
Explore different use cases and parameter configurations
Transparent pricing with no hidden fees. Pay as you go.
| Rule & Modality | Channel | Credits | Price (USD) | Official / Reference Price | Daily Savings |
|---|---|---|---|---|---|
text-to-image imagegrokenable_pro: false | Starter | 0.45per run | $0.002 | xAI$0.022per run | Based on 10,000 images/dayxAI costs $199.550/day moreApiPass is 90.70% lower |
text-to-image-quality imagegrokenable_pro: true | Starter | 0.6per run | $0.003 | xAI$0.050per run | Based on 10,000 images/dayxAI costs $472.640/day moreApiPass is 94.55% lower |
Complete guide to using grok-imagine
Generate high-fidelity imagery directly from text prompts using the grok-imagine/text-to-image endpoint, seamlessly accessible through the ApiPass platform. The underlying Grok model interprets intricate visual descriptions with high precision, offering native support for multiple aspect ratios and an advanced quality mode designed for production-ready rendering.

ApiPass acts as a high-concurrency API proxy for xAI's Grok Imagine model, allowing you to generate advanced text-to-image assets via standard HTTP requests. ApiPass wraps the upstream infrastructure into a stable, asynchronous workflow, handling rate limits, billing, and global routing out of the box so you can plug Grok's visual capabilities straight into your platforms, developer solutions, or automated backend pipelines.
Generate portraits, environments, products, concept art, editorial scenes, and abstract visuals by writing clear English descriptions.
Use 2:3 for portrait orientation, 3:2 for landscape orientation, 1:1 for square format, 16:9 for widescreen images, and 9:16 for tall vertical images.
Create a task first, then receive the completed result by callback or query the task record endpoint with the returned taskId.
For production integrations, configure callBackUrl when creating a task so ApiPass can send task result updates automatically.
Create warm, atmospheric portraits with soft lighting, film grain, period styling, and shallow depth of field.
Generate clean commercial-style images with studio lighting, simple backgrounds, negative space, and polished surface details.
Turn descriptive ideas into high-impact visuals for mood boards, design exploration, creative pitches, and visual research.
The workflow is simple: write a prompt, choose optional settings, create a task, and retrieve the generated image URLs.
Include the subject, setting, composition, style, lighting, colors, mood, and important details. The prompt parameter is required and supports up to 5000 characters.
Set aspect_ratio to 2:3, 3:2, 1:1, 16:9, or 9:16. Leave it empty to use the default 1:1 square format.
Set enable_pro=false for speed mode, or enable_pro=true for quality mode.
Submit a POST request to the ApiPass endpoint /api/v1/jobs/createTask. Use the returned taskId to query /api/v1/jobs/recordInfo, or configure callBackUrl to receive results automatically.
Grok Text-to-Image generation is asynchronous. Create a generation task first, then receive the final image URLs through callback notifications or polling.
POST https://api.apipass.dev/api/v1/jobs/createTask with Authorization: Bearer YOUR_API_KEY and Content-Type: application/json. The request body top level is fixed to model and input; callBackUrl is optional.
Use the model value grok-imagine/text-to-image for this endpoint.
GET https://api.apipass.dev/api/v1/jobs/recordInfo?taskId=YOUR_TASK_ID to check status and retrieve output.urls when callback notifications are not configured.
When callBackUrl is configured, ApiPass sends task updates to that URL. A successful task returns output.urls containing generated image result URLs.
{"model":"grok-imagine/text-to-image","channel":"auto","callBackUrl":"https://your-domain.com/api/callback","input":{"prompt":"Cinematic portrait of a woman sitting by a vinyl record player, retro living room background, soft ambient lighting, warm earthy tones, nostalgic 1970s wardrobe, reflective mood, gentle film grain texture, shallow depth of field, vintage editorial photography style.","aspect_ratio":"3:2","enable_pro":false}}
Generate shareable images, thumbnails, post graphics, campaign concepts, and vertical visuals for social channels.
Embed prompt-based image generation into developer platforms for dynamic banners, profile assets, graphics tools, personalized visuals, and automated user-generated content experiences.
Create polished product scenes, advertising mockups, packaging concepts, and presentation assets with controlled lighting and composition.
Quickly explore art directions, character ideas, color palettes, scene layouts, editorial concepts, and visual design research.
Built on xAI’s advanced generative architecture, the model excels at deep semantic compliance, translating intricate, multi-layered English prompt descriptions into high-fidelity visual assets without drift.
ApiPass wraps the raw Grok infrastructure into a fully-managed, high-concurrency proxy layer, eliminating upstream rate limiting spikes and ensuring enterprise-grade uptime for scaling commercial networks.
The endpoint natively supports five core production aspect ratios at the API level: 1:1 (square), 16:9 (widescreen), 9:16 (vertical), 3:2 (landscape), and 2:3 (portrait).
No. The ApiPass integration utilizes an asynchronous task architecture. The API returns an immediate 202 Accepted response with a unique taskId, keeping your main backend threads completely unblocked.
Yes. Pass an optional seed integer in your request payload to lock the model's initial noise state. This ensures structural consistency across multiple generations when debugging prompts or tuning specific layout variables.
Selecting premium triggers extended inference steps on the Grok backend. It enhances fine-grained spatial accuracy, texel density, and typographic readability inside images, though it marginally adds to the rendering lifecycle duration.
All APIs require authentication via Bearer Token.
Authorization: Bearer
Create an asynchronous Grok Imagine text-to-image generation task from a text prompt.
1{
2 "model": "grok-imagine/text-to-image",
3 "channel": "auto",
4 "callBackUrl": "https://your-domain.com/api/callback",
5 "input": {
6 "prompt": "Cinematic portrait of a woman sitting by a vinyl record player, retro living room background, soft ambient lighting, warm earthy tones, nostalgic 1970s wardrobe, reflective mood, gentle film grain texture, shallow depth of field, vintage editorial photography style.",
7 "aspect_ratio": "3:2",
8 "enable_pro": false
9 }
10}modelRequiredstringModel endpoint name. Use grok-imagine/text-to-image for this endpoint.
grok-imagine/text-to-image
channelOptionalstringDefault channel is auto. APIPASS automatically allocates tasks across available providers based on real-time pricing and stability metrics. starter is ultra-low-cost with limited quotas and weaker stability; regular is standard and much cheaper than official APIs with moderate stability; official uses the model native API with high stability, fast task execution, and official pricing.
Default channel is auto. APIPASS automatically allocates tasks across available providers based on real-time pricing and stability metrics; starter is ultra-low-cost with limited quotas and weaker stability; regular is standard and much cheaper than official APIs with moderate stability; official uses the model native API with high stability, fast task execution, and official pricing.
Available options:
auto
callBackUrlOptionalstringCallback URL used to receive automatic task result notifications when generation completes. Recommended for production use instead of polling the status endpoint.
https://your-domain.com/api/callback
promptRequiredstringText prompt describing the desired image. Be detailed and specific about the visual elements, including composition, style, lighting, mood, and other details. Maximum length: 5000 characters. English prompts are supported.
Cinematic portrait of a woman sitting by a vinyl record player, retro living room background, soft ambient lighting, warm earthy tones, nostalgic 1970s wardrobe, reflective mood, gentle film grain texture, shallow depth of field, vintage editorial photography style.
aspect_ratioOptionalstringSpecifies the width-to-height ratio of the generated image. Default: 1:1.
2:3 is portrait orientation, 3:2 is landscape orientation, 1:1 is square format, 16:9 is widescreen format, and 9:16 is tall vertical format.
Available options:
3:2
enable_proOptionalbooleanControls the request processing strategy. false enables speed mode, where the system prioritizes response time and throughput for latency-sensitive scenarios. true enables quality mode, where the system prioritizes processing quality and precision for scenarios requiring higher accuracy.
false
1curl -X POST "https://api.apipass.dev/api/v1/jobs/createTask" \
2 -H "Authorization: Bearer YOUR_API_KEY" \
3 -H "Content-Type: application/json" \
4 -d '{
5 "model": "grok-imagine/text-to-image",
6 "channel": "auto",
7 "callBackUrl": "https://your-domain.com/api/callback",
8 "input": {
9 "aspect_ratio": "3:2",
10 "prompt": "Cinematic portrait of a woman sitting by a vinyl record player, retro living room background, soft ambient lighting, warm earthy tones, nostalgic 1970s wardrobe, reflective mood, gentle film grain texture, shallow depth of field, vintage editorial photography style.",
11 "enable_pro": false
12 }
13}'1{
2 "taskId": "task_xxx",
3 "status": "queued"
4}taskIdUnique task ID for the created generation job. Use this ID to query task status and results.
statusInitial task status after creation, such as queued.

flux
flux
Generate photorealistic images with robust text
Starting from
0 credits
Nano Banana 2 Lite, aka Gemini 3.1 Flash-Lite Image, is a newly released image model from Google DeepMind built for fast 1K image generation, efficient creative iteration, prompt-based image editing, and responsive visual output in product-facing creative scenarios.
Starting from
0 credits
OpenAI
OpenAI
Generate intentionally designed visuals with less prompting using the latest OpenAI GPT Image 2 model.
Starting from
0 credits
Qwen
Qwen
Starting from
0 credits
Nano Banana now allows you to freely draw on edited images, bringing you a completely new experience. Simply click on the image to draw whatever you want. You can also add text and details, and Gemini will do the rest.
Starting from
0 credits
Fast image generation model built on Gemini 3.1 Flash Image. Supports text-to-image, image editing, multi-image fusion with up to 14 reference images, Google Search grounding, and outputs up to 4K resolution.
Starting from
0 credits
seedream
seedream
Starting from
0 credits
qwen
qwen
Generate realistic 2K images with structured text
Starting from
0 credits
seedream
seedream
Starting from
0 credits
OpenAI
OpenAI
Bring your ideas to life with sharper, faster, more precise image generation using the GPT Image 2.5 API.
Starting from
0 credits
wan
wan
Generate images with cinematic precision through high-fidelity text-to-image synthesis and structural editing.
Starting from
0 credits
seedream
seedream
Gemini said Generate images with lightning-fast efficiency and high-fidelity detail through a streamlined, lightweight architecture optimized for instant creative expression.
Starting from
0 credits
qwen
qwen
Generate realistic 2K images with structured text
Starting from
0 credits
Luma
Luma
Starting from
0 credits
seedream
seedream
Starting from
0 credits