Spicy API enables AI agents to generate and customize AI-powered adult-oriented imagery using text prompts and image-based inputs, with options for controlling visual styles and creative outputs.
Encrypted at rest, isolated from the model
Resolved from an AES-256-GCM vault at the moment of the call and attached to the request — the model never sees the secrets.
Try asking
Every model with its kind (image, image-edit, video, video-edit, chat, speech, realtime, transcription, embedding), USD price (per image, per second by resolution or quality, per 1M tokens, per 10,000 characters of speech input, or per minute of audio), limits (image sizes, outputs per call, resolutions, clip lengths, aspect ratios, prompt character caps), whether it needs or accepts an input image or references, and example clips. Video rows carry `inputs`: which of first frame, last frame, reference images, clip and audio the model takes, `inputs.audio.mode` (driving: the clip is the soundtrack and the mouth follows it; reference: the model generates sound and uses the clip for voice, tone and beat; none) and `inputs.combinations`, the valid input combinations in plain sentences. Read those before generate_video. `silent: true` marks a video model that renders without sound; `bills_input_video_seconds: true` one that also bills the seconds of an input clip. Video-edit rows carry `video_edit` (instruction or motion mode, clip lengths, reference images, whether input seconds are billed) for edit_video. Speech rows carry `speech`: preset voices, whether `instructions` are accepted, custom voices, languages, inline `tags` and voice creation fees. Realtime rows (spicy-live-1) are live voice calls over a WebSocket: a server or app opens a session with POST /v1/realtime/sessions and connects the returned ticket URL; they cannot be driven from a tool call. Prices here are exactly what the API bills. Works with a $0 balance.
USD cost of a request before making it, from the same price table the gateway bills: images by count, video by seconds and resolution (plus the input clip or action clip on models that bill it), video edits by clip length (input plus output seconds for instruction edits, output seconds for motion transfer), chat by tokens, speech by characters of input (plus instructions), transcription by seconds of audio, embeddings by tokens, live calls by minutes. Use it to quote the user and to check against get_account before a spend. Works with a $0 balance.
Balance in USD, spend this calendar month, the monthly spend limit if set, whether approvals are skipped for agents, whether the account has ever topped up, the webhook URL, and whether this key is a sandbox key. Call it first and before any spend.
The account's generated images, newest first. These are the only URLs accepted as image inputs by edit_image, generate_video, edit_video and create_embeddings (no uploads, no third-party URLs). Use it to pick a first frame to animate or a performer for motion transfer.
The account's saved characters: reusable identities built from its own generated images. Pass a character id as `character` to generate_image, edit_image or generate_video to get the same person again across images, edits and video. No uploads are involved; a character can only be made from images this account generated. The exception is `imported: true` characters: a registered business can apply in the dashboard (Settings, Character import: https://www.spicyapi.com/dashboard/import) to bring its own platform's existing AI-generated characters over with POST /v1/characters/import. There is no tool for importing, and an agent cannot apply for the user. One with `status: "under_review"` cannot be used until a person approves it.
Save a reusable identity from 1 to 3 of the account's own generated images of the same person (list_images or a generate_image result). Renders one neutral reference sheet (billed as one image edit) and returns the character id to pass as `character` elsewhere. Up to 50 per account. Guarded: without approval_token it returns {error: "approval_required", summary, approval_token}; show the summary (it carries the cost), get the user's yes, call again with the token.
Remove a saved character. Its reference sheet stays in the image library. Guarded: without approval_token it returns {error: "approval_required", summary, approval_token}; show the summary (it carries the cost), get the user's yes, call again with the token.
Text to image, synchronous (10 to 30 seconds), returns durable CDN URLs reusable as image_url for edits and video. Or an image action: model spicy-image-action-1 with an action from list_image_actions and image_url of the woman (one of list_images) puts her face, hair and skin tone into that ready-made explicit scene, no prompt (a minute or two); style studio gives the finished picture a warm glamour colour grade, and anime, 3d or cartoon put her into the drawn version of the scene, for a drawn woman. Or a reference: model spicy-image-reference-1 with reference_image (a picture the user supplies, https or data URL) describes it (faces left out) and draws a brand new image from that description, new people with the same look, and returns the description to reuse as a prompt (30 to 60 seconds, one image). One to six images per call, billed per image at the model's rate; the exact amount is cost_usd in the result. Prompts are screened first and outputs after (withheld images are refunded and counted in `withheld`); blocked prompts cost nothing. Every request depicts adults only. Guarded: without approval_token it returns {error: "approval_required", summary, approval_token}; show the summary (it carries the cost), get the user's yes, call again with the token.
Prompt-driven edit of one to three images the account generated (list_images), for outfit, pose and scene changes. Synchronous; returns one to six durable CDN URLs, billed per output image (withheld outputs are refunded). Guarded: without approval_token it returns {error: "approval_required", summary, approval_token}; show the summary (it carries the cost), get the user's yes, call again with the token.
Image to video, text to video or references to video. Asynchronous: returns a task id at once with the price already debited; poll get_job every 10 to 15 seconds until succeeded (output.video_url, usage.output_seconds) or failed (refunded). Every input must be the account's own: images from list_images, clips from its own finished tasks (video_url), audio from upload_audio or generate_speech. Price = seconds x the per-second rate for the resolution; models with bills_input_video_seconds (spicy-motion-3, -fast) also bill the seconds of a video_url clip or of an action clip. duration -1 on spicy-motion-3 lets the model pick the length, debited at the maximum and refunded down to what is rendered. Text to video: spicy-motion-3, -fast, spicy-video-1, spicy-cinema-1. Image to video: spicy-motion-2, spicy-cinema-1-image, spicy-motion-3. Character or references to video: spicy-character-video-1, spicy-cinema-1-character (up to 9 references named Image 1, Image 2 in the prompt), spicy-motion-3. Cheap previews: spicy-motion-draft-1 is a silent draft tier from a first frame; render the keeper on spicy-motion-2 or 3. The spicy-cinema-1 family runs 3 to 15 seconds at 480P to 1080P with native audio always on, no negative_prompt, no enhance_prompt and no audio inputs. Which inputs may travel together differs per model: read list_models inputs.combinations first. Audio comes in two modes (inputs.audio.mode): driving (spicy-motion-2, spicy-video-1: the clip is the soundtrack, the mouth follows it, works with a first frame) and reference (spicy-motion-3, spicy-motion-3-fast: the model generates sound and uses the clip for voice, tone and beat; only with a text prompt or references, never with a first frame). An invalid combination is rejected here before any spend with the valid combinations spelled out. Actions (spicy-motion-3, -fast): pass an action id from list_actions and its reference clip of that sex act supplies the camera, pose and motion second for second while image_url supplies the woman, e.g. {model: "spicy-motion-3", action: "pov-missionary", image_url: <one of list_images>, prompt: "hotel room at night, red lingerie"}. To change a finished clip or transfer motion onto a person, use edit_video. Guarded: without approval_token it returns {error: "approval_required", summary, approval_token}; show the summary (it carries the cost), get the user's yes, call again with the token.
Change a finished clip, or make a person perform a motion. Asynchronous like generate_video: returns a task with the price debited; poll get_job every 10 to 15 seconds (refunded if it fails). Instruction edits (spicy-video-edit-1: 2 to 10 second clips, output the clip's length; spicy-cinema-1-edit: 3 to 30 second clips, the first 15 seconds are edited and returned): video_url (one of the account's finished video outputs) plus a prompt describing the change, optionally reference_image_urls for an outfit, prop or style (the account's own images), keep_audio to keep the soundtrack. Billed per second of input plus output at the resolution's rate, debited at the clip's length and settled to what is rendered. Motion transfer (spicy-animate-1): image_url (the performer, one of the account's images, full body and dressed works best) plus a motion id from list_motions or a video_url of the account's own clip (2 to 30 s); the output is as long as the motion clip, billed per output second by quality. Motion transfer refuses explicit or nude images and clips (refunded, error_code blocked). Guarded: without approval_token it returns {error: "approval_required", summary, approval_token}; show the summary (it carries the cost), get the user's yes, call again with the token.
The motion library for motion transfer on edit_video (spicy-animate-1): id, label, description, preview url, duration_s and aspect_ratio of each clip. Every clip is 9:16, full body, fixed camera and was generated with SpicyAPI video models (no real performer footage). The output is as long as the clip; match the performer image to its framing for best results. Works with a $0 balance.
The action library for generate_video with action (spicy-motion-3 and spicy-motion-3-fast): id, label, category (sex, anal, oral, handjob, cumshot, solo, trans), url (the reference clip, 4 to 6 seconds), preview_url (the full 8 second action), poster_url, duration_s (the reference clip's length) and aspect_ratio. The clip supplies the camera, pose and motion of one sex act; the output runs 8 seconds by default and carries the act on past the clip. The caller's image supplies the woman. Every clip was generated with SpicyAPI video models (no real performer footage). Its seconds are billed as input seconds. Works with a $0 balance.
The image action library for generate_image on spicy-image-action-1: id, label, category (sex, anal, oral, handjob, cumshot, solo, trans), image_url (the still), preview_url (a small copy), size (the output width*height) and styles (the still redrawn as anime, 3d and cartoon, for style on a drawn woman). 68 ready-made explicit stills; most share their id with a video action (list_actions). The still supplies the act, pose, framing, room and her body; the caller's image supplies her face, hair and skin tone. Every still was generated with AI (no real performers). Billed per image at the model's rate. Works with a $0 balance.
The art styles for `style` on generate_image and generate_video: id, label, description and preview_url (a sample image). photorealistic, studio (glossy glamour), anime, 3d (animated-film CGI) and cartoon (2D adult cartoon). A style only adds prompt text, so the model, seed and price are unchanged. Works with a $0 balance.
Bring in a WAV or MP3 (2 to 30 seconds, 15 MB max) from a public https URL for lip-sync and audio-driven video. The clip is transcribed and the transcript screened like a prompt; flat $0.01 for that screening. A clip that fails the screen is discarded with the same 422 shape as a blocked prompt; a 503 means screening was unavailable and nothing was charged. The returned url is what audio_url and audio_urls on generate_video accept; external audio URLs are rejected there. Guarded: without approval_token it returns {error: "approval_required", summary, approval_token}; show the summary (it carries the cost), get the user's yes, call again with the token.
The account's audio clips, newest first: uploads and generated speech (source upload or speech), with id, url, duration_s, mime and transcript. These urls are the only values generate_video accepts as audio_url or in audio_urls.
Forget an uploaded clip so it can no longer be used as a video input. Videos already made with it are unaffected. Guarded: without approval_token it returns {error: "approval_required", summary, approval_token}; show the summary (it carries the cost), get the user's yes, call again with the token.
Speech to text from a public https URL of an audio file (wav, mp3, m4a/mp4, ogg/opus, flac, webm and more; up to 10 MB and 5 minutes). Synchronous; returns {id, text, language, duration_s, cost_usd}. Billed by the second at the model's per-minute rate (minimum charge applies); the audio is not stored and nothing is screened, since nothing is generated. Use it for voice notes in a companion app or to read back a clip. Guarded: without approval_token it returns {error: "approval_required", summary, approval_token}; show the summary (it carries the cost), get the user's yes, call again with the token.
Text to speech, synchronous. Returns {id, url, duration_s, voice, format, cost_usd}: a durable clip (WAV 24 kHz mono, or MP3 on the spicy-voice-2 family) that is also one of the account's audio clips, so its url goes straight into generate_video as audio_url (driving audio on spicy-motion-2: the mouth follows the words, with a first frame; reference audio on spicy-motion-3: voice, tone and beat, no first frame). Video audio must be 2 to 30 seconds. Models: spicy-voice-1 (preset voices), spicy-voice-1-expressive (preset voices plus instructions for emotion, pace and delivery), spicy-voice-1-custom (the account's own vc_... voices from create_voice), spicy-voice-2 and spicy-voice-2-flash (their own preset voices, instructions, languages Auto, English and Chinese, up to 5,000 characters, and inline tags inside input: control tags switch delivery until the next tag, [sad] [amazed] [deep and loud shouting] [trembling] [angry] [excited] [sarcastic] [curious] [like dracula] [bored] [tired] [scornful] [shouting] [asmr] [panicked] [mischievously] [empathetic] [whispers] [reluctantly] [crying] [serious] [very slowly] [very fast]; sound tags insert a sound, [gasp] [sighing] [clears throat] [giggles] [laughing] [cough] [snorts]; e.g. "[excited]Hey you.[giggles] Come here.[whispers] Closer."). Tags count as billed characters. Input caps: 600 characters on the spicy-voice-1 family, 5,000 on spicy-voice-2. Billed per 10,000 characters of input plus instructions (CJK ideographs count 2), settled down to the billed count. Streaming (stream: true) is a REST feature for apps (POST /v1/audio/speech); this tool always returns a stored clip. The text is screened like a prompt; blocked text costs nothing. Guarded: without approval_token it returns {error: "approval_required", summary, approval_token}; show the summary (it carries the cost), get the user's yes, call again with the token.
Voices for generate_speech: every preset voice with the models that speak it (spicy-voice-1 has them all, spicy-voice-1-expressive a subset), then the account's designed and cloned voices (vc_... ids, spoken by spicy-voice-1-custom), newest first. Works with a $0 balance.
Create a custom voice for spicy-voice-1-custom, returned as a vc_... id. Design: name plus a description of vocal qualities (gender, age range, pitch, pace, tone, accent); a description that names, references or imitates a real person is refused. Clone: name plus audio_url (one of the account's clips from list_audio or upload_audio, 10 MB max) plus consent: true, the user's attestation that the voice is their own or that they hold the speaker's written consent; it is stored with the voice. Never set consent without the user confirming it. One-off fee per voice, refunded if creation fails. Guarded: without approval_token it returns {error: "approval_required", summary, approval_token}; show the summary (it carries the cost), get the user's yes, call again with the token.
Delete a designed or cloned voice (vc_...) so it can no longer be spoken. Clips already generated with it are unaffected. Guarded: without approval_token it returns {error: "approval_required", summary, approval_token}; show the summary (it carries the cost), get the user's yes, call again with the token.
Embedding vectors for search, memory and recommendations, OpenAI-compatible. spicy-embed-1: text, input a string or up to 10 strings (8,192 tokens each), dimensions 64, 128, 256, 512, 768, 1024 (default), 1536 or 2048. spicy-embed-vision-1: text and images in one 768-dimension space, input items are strings, {text} or {image} where every image is one of the account's own generated images (list_images); use it for "more like this" over the account's outputs, tagging and near-duplicate detection. Returns {data: [{index, embedding}], usage, cost_usd}; billed per token (minimum charge applies). Vectors are long: ask for small dimensions (256) when you only need to compare a few items here. Guarded: without approval_token it returns {error: "approval_required", summary, approval_token}; show the summary (it carries the cost), get the user's yes, call again with the token.
Status of a video task from generate_video or edit_video (queued, processing, finalizing, succeeded with output.video_url and usage, or failed and refunded) or of a logged image, chat, speech, transcription or embedding request by id.
Recent requests on the account (images, edits, video tasks including edits and motion transfer, chat, speech, voice creation, transcription, embeddings, live calls), newest first: id, type, model, status, cost_usd, output URL when finished, error when failed, and whether it came from the playground, the API or an agent.
Spend and request counts over the last N days, totalled and broken down by model, type and source (playground, api, agent), with the error count. Reads the account's request log.
Mint a new key for this account (plaintext returned once, stored hashed; at most 500 active keys). sandbox: true makes a key that answers everything from fixtures and bills nothing. Keys made with a sandbox key are always sandbox keys. Keep keys in environment variables, never in chat. Guarded: without approval_token it returns {error: "approval_required", summary, approval_token}; show the summary (it carries the cost), get the user's yes, call again with the token.
Permanently revoke a key by id (the id is returned by create_api_key and shown on the API keys page). Applications using it stop working at once. Guarded: without approval_token it returns {error: "approval_required", summary, approval_token}; show the summary (it carries the cost), get the user's yes, call again with the token.
Checkout URL to add funds by card (hosted checkout) or crypto for a preset amount ($50, $100, $250, $500, $1000; minimum $50). The USER opens it and pays in the browser; never enter payment details yourself. The balance is credited by the payment webhook; call get_account afterwards. If it returns commitment_required, the user must first sign the one-page Customer Commitment Letter on the Billing page; an agent cannot sign it. Works with a $0 balance.
Cap what this account can spend per calendar month (UTC), enforced at debit time on every surface: requests past the cap return 402 without charging. Lowering or setting a limit never needs approval; raising or removing one does. null removes the limit.
The acceptable use policy as markdown: prohibited content (minors in any form, images or video of real people with or without consent, non-consensual scenarios and the rest), age-verification duties for the customer's product, enforcement. Read it before the first generation and tell the user what their product must do. Works with a $0 balance.
One endpoint, the same key, whichever client you use.
~/Library/Application Support/Claude/claude_desktop_config.json (Mac) · %APPDATA%\Claude\claude_desktop_config.json (Windows)
Replace API_KEY with your own key.
Already have an "mcpServers" section in your config? Just add the server entry inside it.
Discovery, routing, credentials, tool scoping and execution logs all happen at the gateway→connections stay ACTIVE with no work from you
Spicy MCP runs through a gateway that holds the credentials, scopes the access and records every call.
Managed auth, hosted MCP servers, and every Gmail tool your agent needs.
Free to start.