This guide is for anyone opening OpenCake for the first time. OpenCake is an AI creative production platform for making product ads, UGC videos, product visuals, actor-led clips, motion-design reels, explainers, translated and lip-synced videos, captions, voiceovers, and reusable campaign assets. It combines a conversational creative agent, direct model access, reusable products and actors, specialized production workspaces, and a Library that keeps the useful results connected.
The shortest way to understand OpenCake is this: use Agent Cake when you want help planning and coordinating a complete job; use Studio when you want direct control over a specific image, video, audio, or editing model; create reusable Products and Actors when identity matters; and use the Library to carry strong assets from one step into the next.
Start with the dashboard
The dashboard is where you move between OpenCake workspaces. The correct starting point depends on how much of the production plan you already know. If you have an outcome but not the exact workflow, start with Agent Cake. If you already know which model you want and what inputs it needs, open Studio. Use a dedicated workspace when the job has a purpose-built interface, such as Video Dubbing, Kling 3.0 Motion Control, Captions, file conversion, compression, or metadata removal.
- Use Agent Cake when you want to describe a complete outcome in plain language and have the system plan the production path.
- Use Studio when you know the model or output type you want and prefer to control references and settings directly.
- Use Actors when you need a reusable person or character for videos.
- Use Products when you want product details to stay consistent across outputs.
- Use Kling 3.0 Motion Control when you want one identity across many video motions.
- Use Video Dubbing when you want to translate spoken footage and rebuild lip movement for another language.
- Use Captions when a finished video needs styled subtitles for social platforms.
- Use utility models when you need to clean, upscale, strip, convert, compress, or prepare media.
- Use Library when you need to find, reuse, download, or reference assets.
Agent Cake
Agent Cake is OpenCake’s conversational creative-production agent. Instead of beginning with a model card and a collection of parameters, you begin with the result: “Create a 15-second unboxing ad for this product,” “Turn this topic into a one-minute animated explainer,” or “Adapt the structure of this reference ad without copying its execution.” Agent Cake can identify the inputs, load a specialized production skill, plan the work, use OpenCake tools, and keep the conversation connected to the generated media.
Agent Cake is not one image or video model. It is the planning and tool-using layer above OpenCake’s models, skills, Products, Actors, research tools, media analysis, captions, dubbing, montage assembly, and Library. That distinction is useful for beginners: you do not need to memorize which model supports reference video or first-and-last frames before you can describe the job.
What Agent Cake can do
- Plan and produce UGC videos, product reviews, unboxing videos, product tutorials, SaaS presenter ads, virtual try-on concepts, commercials, cinematic sequences, podcasts, explainers, faceless YouTube videos, cartoons, and motion-design reels.
- Analyze supplied images and videos, transcribe media, detect scene cuts, extract public web pages, research the web, and use those findings as production context.
- Work with saved Products, OpenCake Actors, uploaded files, and existing Library assets instead of treating each request as an isolated prompt.
- Create or reuse Actors, generate images, video, and audio, assemble ordered clips, add captions, dub videos, and run supported media utilities.
- Show structured creative questions when a missing decision would materially change the result.
- Prepare an exact credit quote before paid media generation and keep completed outputs attached to the conversation and stored in the Library.
Ask mode and Yolo mode
Agent Cake has two generation-approval modes. Ask mode pauses after preparing the generation quote so you can approve the exact work and credit amount. Yolo mode can continue automatically after quoting. New users should begin with Ask mode: it makes the creative plan and cost visible before production. Yolo mode is better reserved for a workflow you understand and have already reviewed.
How to brief Agent Cake
A strong Agent Cake brief names the business goal, deliverable, audience, destination, source assets, factual boundaries, and review preference. For example: “Create three 15-second 9:16 UGC concepts for this saved coffee grinder. The audience is apartment coffee drinkers. Use one consistent actor, show only the product features in the supplied page, and show me the scripts and storyboards before quoting generation.”
Attach assets from the Library, Products, or Actors with the plus button in the composer. Explain what each attachment should control. “Use this actor for identity,” “use this product image for packaging,” and “take only the camera movement from this video” are clearer than attaching several references without roles.
Studio
Studio is the direct model workspace in OpenCake. Use it when you want to choose a specific image, video, audio, lip-sync, or media-processing model yourself. Studio builds its model picker from the live OpenCake catalog, renders the controls supported by the selected model, validates attached media, and updates the credit quote as the request changes.
The Studio composer keeps the prompt and attached references available while you move between Studio and the Library. You can upload a new file or reuse a Product, Actor, image, video, or audio asset. Models accept different combinations, so the visible controls and required media change with the selected card. The complete public Studio catalog below contains 32 entries as of August 24, 2026.
Creative video models in Studio
| Model | Maker | Best starting point |
|---|---|---|
| FLUX 3 | Black Forest Labs | Native-audio video, ordered keyframes, source-video continuation, and draft-to-full-quality production. |
| MiniMax H3 | MiniMax | Multimodal video using text, first and last frames, images, videos, and audio as separately assigned references. |
| Seedance 2.5 | ByteDance | Text, first/last frames, multimodal references, video editing, and extension with synchronized audio. |
| Seedance 2.0 | ByteDance | General cinematic generation, multi-reference control, actor/product work, dialogue, editing, and extension. |
| Veo 3.1 | Text, image, first/last-frame, reference, and extension workflows with optional synchronized audio and several quality tiers. | |
| LTX 2.3 | Lightricks | Text, image, and audio-driven video with portrait and high-resolution output options. |
| Kling 3.0 | Kuaishou | Fluid cinematic motion from text or images, native audio, and Standard, Pro, or 4K routes. |
| Happy Horse | Alibaba | Text, first-frame, multi-reference, and source-video editing with native audio and multilingual lip sync. |
| Grok Imagine Video 1.5 | xAI | Short image-to-video clips with included generated audio and straightforward duration and resolution controls. |
| Gemini Omni Flash | Short native-audio video from text, a seed image, reference images, or a source-video edit. |
Creative image models in Studio
| Model | Maker | Best starting point |
|---|---|---|
| Nano Banana 2 | Image generation, editing, and compositions using as many as 14 reference images. | |
| Nano Banana Pro | Typography, identity preservation, and precise multi-reference image editing. | |
| Midjourney V8.1 | Midjourney | Natural-language image generation, four-image Fast batches, Draft exploration, HD output, and strong art direction. |
| GPT Image 2 | OpenAI | Detailed image generation, readable typography, posters, storyboards, character sheets, and reference-based edits. |
| Ideogram V4 | Ideogram | Posters, logos, images, and restyling where crisp typography is important. |
| MAI Image 2.5 | Microsoft | Photorealistic text-to-image generation and single-image editing for design-ready visuals. |
| Seedream 5.0 Pro | ByteDance | Flagship image generation and multi-image editing. |
| Recraft V4.1 | Recraft | Logos, icons, brand graphics, typography-aware layouts, raster artwork, and true SVG vector output. |
| Kling O3 Omni-Image | Kuaishou | Reference-grounded generation and editing with up to 10 images, series output, 4K, typography, and optional face control. |
Audio models in Studio
| Model | Maker | Best starting point |
|---|---|---|
| Seed Audio 1.0 | ByteDance | Speech, dialogue, short audio scenes, preset voices, and reference-guided audio generation. |
| Fish Audio S2.1 Pro | Fish Audio | Expressive multilingual text-to-speech and authorized cloned-voice workflows. |
Editing and utility models in Studio
| Studio entry | Media | What it does |
|---|---|---|
| Remove Background | Image | Removes an image background and returns a transparent asset. |
| Image Upscaler | Image | Enhances images for larger or sharper exports, including 4K or 8K routes. |
| Image Watermark Remover | Image | Removes watermarks, logos, text overlays, or marks from images you have the right to edit. |
| Video Background Remover | Video | Separates the foreground and can return transparent WebM video. |
| FLUX Video Upscale | Video | Enhances an MP4 at 1.5–3× with source-faithful Precise or detail-generating Creative processing for 1080p, 2K, or 4K delivery. |
| Video Watermark Remover | Video | Removes watermarks, logos, captions, or on-screen text from video you have the right to edit. |
| Audio Remover | Video | Returns a silent copy of a supplied video. |
| Change Aspect Ratio | Video | Uses LTX 2.3 Reframe to generate scene-matched content around the original frame instead of simply cropping it. |
| VEED Lip Sync v2 | Video | Synchronizes a speaker’s mouth to supplied audio without generating a replacement voice. |
| Sync Lipsync v3 | Image or video | Drives advanced facial motion from audio using either a source video or a single still image. |
The catalog changes as models and provider routes evolve. Treat the Studio picker and its live controls as authoritative at generation time. A model name alone does not guarantee that every mode, duration, aspect ratio, resolution, reference type, or audio option is available in every configuration.
How to choose a Studio model
- Start with the required input. A source image, first and last frames, several identity references, a video to edit, or audio to drive motion immediately narrows the useful choices.
- Choose for the deliverable, not the newest name. Recraft is useful for an SVG logo; Fish Audio is useful for speech; a cinematic video model is not a substitute for either.
- Use the cheapest suitable quality tier to test composition, timing, and prompt direction before increasing resolution or duration.
- Keep reference roles explicit. Identity, product, location, first frame, last frame, movement, style, and voice are different controls even when they are all attached media.
- Compare models with the same brief and references when testing. If the prompt, duration, aspect ratio, and inputs all change, the comparison does not reveal which model caused the difference.
Using references in Studio
A product image, actor sheet, first frame, last frame, source video, or audio file usually provides more control than prompt text alone. OpenCake represents reusable references with mention tokens such as @actor1, @product1, @image1, @video1, and @audio1 when the selected model supports that media. The token tells the prompt which attachment it is discussing; the attached file supplies the actual pixels, motion, or sound.
Do not attach the maximum number of references by default. Begin with the smallest set that controls the result. For example, a UGC product clip may need one Actor and one Product. Add a location reference only if the setting must match, a video only if its movement matters, and audio only when it has a defined role such as voice, cadence, ambience, or music.
Which OpenCake workspace should you use?
| If you are starting with… | Open… | Why |
|---|---|---|
| A business goal or rough creative idea | Agent Cake | It can research, ask for missing decisions, plan several production stages, and use the appropriate tools. |
| A chosen model and prepared inputs | Studio | It gives you direct access to that model’s supported modes, references, settings, and live quote. |
| A person or character you will reuse | Actors | It turns identity into a saved reference that can travel across compatible workflows. |
| A product you will advertise repeatedly | Products | It keeps the product reference available without making you upload it for every generation. |
| A performance whose motion should drive another identity | Kling 3.0 Motion Control | Its interface is built around a face reference plus a driving video. |
| A finished spoken video for another market | Video Dubbing | It translates speech, rebuilds lip movement, and can add styled captions. |
| A finished social clip that needs subtitles | Captions | It provides transcription, bilingual options, styling, and a visual preview. |
| An existing file that needs technical preparation | Studio utilities or a dedicated file tool | Background removal, upscaling, reframing, conversion, compression, and metadata removal do not need a new creative concept. |
Start with the outcome and let Agent Cake coordinate the workflow. Open Agent Cake
Choose a model and control its supported inputs and settings directly. Open Studio
Actors
Actors are reusable identity references for AI-generated content. OpenCake includes a roster of 299+ built-in AI actors, and you can create a custom Actor when you have an authorized person or original character to preserve. Actors are useful for UGC ads, product demonstrations, spokesperson clips, tutorials, lifestyle scenes, virtual influencers, and recurring brand characters.
An Actor is an identity anchor, not a promise that every model will reproduce every detail perfectly. Use a clear, unobstructed face; avoid beauty filters, extreme expressions, motion blur, and dramatic colored light in the identity source. A three-view character sheet is especially useful because it gives compatible models front, profile, and rear information in one reference. Keep wardrobe, styling, age, and distinguishing features explicit when continuity matters.
Save an Actor when the same identity should appear across a campaign. Then test one short shot before requesting many variants. Reusing the same Actor, product, aspect ratio, and visual direction makes it easier to diagnose whether a change came from the prompt, the model, or the reference. Only create or animate a real person when you have the necessary consent and usage rights.
Choose a built-in Actor or create an authorized reusable identity. Browse OpenCake Actors
Products
Products are reusable visual references for physical goods, ecommerce items, food, packaging, apps, and branded objects. Saving a Product creates a stable source asset that Agent Cake and compatible Studio models can reuse in packshots, lifestyle images, UGC scenes, tutorials, unboxings, try-on concepts, and campaign variations.
Begin with a large, sharp product image in which the full silhouette, materials, proportions, label, and important controls are readable. A neutral background and honest perspective are usually more useful than an already stylized ad. If the back, interior, texture, or scale matters, keep additional approved views in the Library and assign each reference a specific role in the prompt.
The image controls appearance; your brief controls facts. State the exact feature, quantity, compatibility, ingredient, offer, or call to action that may appear. Do not ask a model to infer claims from unreadable packaging or invent a product function. Review labels, logos, hands, contact points, and product geometry before publishing any generated commercial asset.
Save a product reference once and reuse it across OpenCake workflows. Open Products
Kling 3.0 Motion Control
Kling 3.0 Motion Control lets you add a face, choose a video, and place that identity into the same movement. The face can come from an uploaded image, your Library, or an OpenCake actor. The video provides pose, motion, framing, timing, and scene.
Use Kling 3.0 Motion Control when you want a founder, creator, spokesperson, Actor, or brand character to inherit an existing performance without reshooting every variation. A clean face source helps identity; a stable driving clip with a visible subject helps motion transfer. The driving video should already contain the timing, body action, framing, and performance you want—the tool transfers that structure rather than directing a completely new scene from text.
Motion Control is a dedicated OpenCake workspace, not one of the 32 public Studio entries listed above. Test a short, simple movement before attempting fast turns, occluded faces, hand-to-face contact, or multiple people. Always use likenesses and performances you have the right to transform.
Combine an authorized identity with a driving performance. Open Kling 3.0 Motion Control
Video Dubbing
Video Dubbing turns a spoken source video into a translated version that preserves the character of the performance and rebuilds lip movement for the target speech. Choose or upload a video, select from the live catalog of 190+ languages and regional dialects, optionally add styled or bilingual captions, review the duration-based credit quote, and generate a new video that saves to your Library. Source videos can be up to three minutes long.
Use Video Dubbing for founder ads, UGC clips, product demonstrations, tutorials, spokesperson videos, and customer education content that already works visually but needs to reach another language market. Clear speech, one main speaker, a visible mouth, and a steady front or three-quarter camera angle usually make the strongest source material.
- Choose a finished spoken video from My Library or upload a new clip.
- Select the target language or regional dialect for the translated voice and lip movement.
- Enable captions when you want translated or bilingual on-screen text in the same job.
- Review the duration-based credit quote before generating.
- Generate the dubbed version and wait for it to appear in My Library.
- Review the translation, pacing, product claims, pronunciation, and visual synchronization before publishing.
Localize an existing spoken video and optionally add captions. Open Video Dubbing
Library
The Library is where your reusable assets live. It stores generated outputs, uploaded files, products, actors, images, videos, and audio. Use the Library when you want to reuse an asset in a new prompt, download an output, organize creative material, or choose a reference for another workflow.
Think of the Library as production memory rather than a download folder. A strong image can become a first frame. A useful clip can become a motion reference. A voice track can drive lip sync. A product render can become the base for a future campaign. Agent Cake thread outputs remain connected to the work that produced them, while reusable Products and Actors provide stable anchors for future jobs.
Name and organize assets around their future use: product, campaign, market, format, and version are more useful than “final-2.” Keep the clean source, the generated intermediate, and the approved export when each has a distinct role. This makes later resizing, dubbing, captioning, editing, and repurposing faster because you can return to the correct stage instead of degrading a compressed final file.
Find and reuse generated, uploaded, and saved campaign assets. Open My Library
Captions
Captions is the workspace for turning a finished clip into a captioned social video. Upload or select a video, choose a caption style, preview the typography on the video, and render a new copy with timed subtitles.
Use Captions for TikTok, Reels, Shorts, UGC ads, founder videos, product demos, tutorials, and talking-head clips where viewers may watch without sound. The workspace supports style presets, transcription-language selection, a second bilingual line, animation, position, color, size, and uppercase controls. Preview the treatment against the actual footage so text does not cover a face, product, interface, or platform control area.
- Choose a source video from your Library or upload a new clip.
- Pick a style such as Viral Pop, Documentary Box, or Minimal.
- Set the language and enable bilingual captions when you want two caption lines.
- Preview the caption look before rendering.
- Generate the captioned video and save the result back to your Library.
Transcribe and style subtitles for a finished video. Open Captions
Dedicated file tools outside Studio
Studio already contains the background removal, watermark removal, upscaling, audio removal, reframing, and lip-sync entries listed earlier. OpenCake also has dedicated tools for file operations that do not require a generative model. Use these when a good asset needs technical preparation or safer delivery rather than a creative change.
- File Converter changes a file format for an upload requirement, editing system, browser, or delivery channel.
- File Compressor reduces file size when a platform limit, delivery speed, or storage target matters.
- Metadata Strip removes embedded metadata before a file is shared outside your team.
A practical finishing chain is: preserve the clean master in the Library, make the necessary creative edit, render captions or dubbing, then convert or compress the delivery copy. Strip metadata from the file you intend to share, not from the only master. Repeatedly compressing or converting the same export can introduce visible quality loss.
Credits
OpenCake uses credits because different models and settings have different costs. Duration, quality, resolution, provider, input type, and workflow can all affect the credit estimate. Check the in-app estimate before generating.
In Studio, changing duration, resolution, quality tier, output count, or model route can update the quote. In Agent Cake Ask mode, review the proposed generation and quote before approval. A good beginner budget separates exploration from production: use short or lower-cost tests to confirm the idea, composition, identity, and movement; spend on higher quality or longer duration only after those decisions are stable.
Two good first workflows
Route one: let Agent Cake plan the job
- Save one clean Product and choose one Actor if the concept needs a person.
- Open Agent Cake in Ask mode and attach those references from the plus menu.
- Describe the outcome, audience, channel, format, duration, product facts, and what you want to approve before generation.
- Answer any focused creative questions and review the proposed scripts, shots, or plan.
- Approve the quote only when the planned generation matches the job.
- Review the result for factual accuracy, identity, product fidelity, timing, and publishability.
- Ask for a specific revision or reuse the strongest asset in the next step.
Route two: control a model directly in Studio
- Choose the output type and identify the minimum references it needs.
- Open Studio and select the model whose supported inputs fit that job.
- Attach a Product, Actor, image, video, or audio file and state the role of each reference.
- Set aspect ratio, duration, quality, resolution, and other visible controls.
- Review the live quote and generate the smallest useful test.
- Save the useful result to the Library, then change one meaningful variable for the next comparison.
Prompting basics
Agent Cake briefs and Studio prompts have different jobs. An Agent Cake brief can describe a complete business outcome because the agent can plan several steps. A Studio prompt should direct one selected model and one generation. In either case, replace empty praise such as “make it amazing” with observable details: who or what appears, what changes over time, what the camera does, what must remain accurate, and what must not appear.
A useful Studio video-prompt shape is: subject and references, action, setting, shot and camera movement, light, sound or dialogue, duration, and constraints. Example: “@actor1 places @product1 beside the sink, turns the front label toward camera, then demonstrates the top control. Chest-up handheld phone shot in a real apartment kitchen, soft morning window light, direct conversational delivery, preserve the supplied packaging and label, no captions or extra products.”
For image work, describe composition and hierarchy instead of motion. For editing, say what must change and what must remain untouched. For audio, write the exact spoken text when wording matters and separately describe voice, pace, emotion, pronunciation, ambience, and pauses. Never rely on generated visuals or speech to establish product facts that were not in your source material.
Beginner mistakes to avoid
- Starting in Studio with a model name instead of an input and deliverable. First decide whether you need text-to-video, image-to-video, editing, reference control, audio, or a utility.
- Attaching many references without saying what each controls. Identity, product appearance, movement, setting, style, and sound should have explicit roles.
- Using a small, obstructed, filtered, or already distorted source when product or identity accuracy matters.
- Changing the model, prompt, references, settings, and duration at once, then being unable to tell what improved the result.
- Paying for high resolution or long duration before the composition and motion work at a smaller test scale.
- Treating the first plausible output as publish-ready. Check claims, labels, logos, anatomy, continuity, speech, captions, safe areas, and export specifications.
- Using a face, voice, performance, logo, brand asset, copyrighted work, or source video without the necessary permission.
A practical campaign pattern
A reusable OpenCake campaign often follows this chain: save the approved Product and Actor; use Agent Cake to develop concepts, scripts, and shot plans; generate key visuals or individual clips in Agent Cake or Studio; keep approved intermediates in the Library; assemble or adapt the sequence; localize the proven edit with Video Dubbing; add captions; then convert, compress, or strip metadata from the delivery copy. You rarely need every stage, but thinking in stages prevents one overloaded prompt from trying to solve research, direction, generation, editing, localization, and export at once.
The fastest way to get value
Do not try to learn all 32 Studio entries and every workflow at once. Start with one real deliverable, one clean source asset, and one review criterion. If you know the outcome but not the production route, begin in Agent Cake. If you already know the model and inputs, begin in Studio. Generate a small test, keep the useful intermediate, and build the next variation from evidence rather than restarting from scratch. OpenCake becomes more valuable as your Library develops into a set of approved products, identities, frames, clips, voices, and campaign components that can be recombined.
The model picker will continue to evolve, but the durable workflow stays the same: give every reference a role, separate creative decisions from technical finishing, review facts and rights before publishing, and preserve the assets that make the next production faster.