Docs

Images & video

Generate images and video of your boo from Media. Every render keeps your boo's face and body consistent, using its reference photo as the identity.

Confirm your age

Media generation is for adults only. The first time you visit, you'll confirm you're 18 or older and acknowledge the Terms of Service — a one-time check per account.

Pick a boo

Choose which boo you're generating as from the selector at the top of the form. Its reference photo and voice are used automatically — upload your own image or voice track only when you want to override them for a single render.

Images

Describe the shot you want — pose, outfit, setting, style — and Boobot renders it using your boo's reference image automatically. You can also:

  • Swap in a different reference image for a single render
  • Add a negative prompt to steer away from unwanted details
  • Reuse a seed to reproduce a previous render exactly

Images are uncensored for boos marked NSFW.

Video

Three ways to bring a boo to life on video:

  • Lip Sync — a talking-head avatar with audio-conditioned lip-sync. Provide a voice track (and optional background music) and your boo speaks it. The clip runs as long as your voice track.
  • Bring to Life — one image in, and the model invents the movement and its own speech. Describe the action and camera movement, pick a length of 5–20 seconds, and add a voiceover to play over it.
  • Motion Capture — upload a video and its movement, timing and sound are copied onto your boo. The clip is exactly as long as the video you upload, up to 30 seconds, so that upload is also what it costs.

All three support a voiceover and burned-in captions, so a finished clip is ready to post without any external editor. Lip Sync and Bring to Life also take background music and ambience — generated, uploaded, or both (see below).

Bring to Life invents its own speech. The model generates audio as well as movement, and that audio can include the boo appearing to talk — words you didn't write, under any voiceover you add. If you need your boo to say something specific and say it in sync, use Lip Sync.

Picking between Bring to Life and Motion Capture. Bring to Life invents the movement from your description, so it's the one to use when you have an idea but no footage. Motion Capture copies movement from a real clip, so it's the one to use when you want a specific walk, dance or gesture — and because the motion is real, it stays at natural speed however long the clip runs.

Two things to know about Motion Capture: match the framing of your uploaded video to your boo's photo (a full-body dance clip needs a full-body photo, or only the head will move), and note that the output's shape comes from the Aspect ratio setting, not from the video you upload.

Uploaded videos can be up to 100MB. Renders are 720p.

Background music & ambience

Lip Sync and Bring to Life carry a Music & Ambience section with three tracks, and they layer together — you can run all three at once under the voice:

  • Background Music — describe the music you want and Boobot generates it.
  • Ambient Sound — the same, for room tone and atmosphere.
  • Your Own Track — a file you upload yourself.

Each track has its own Volume, Fade in and Fade out. A track longer than the clip is cut off at the end; a shorter one simply stops early.

Your Own Track takes MP3, WAV, M4A, AAC and OGG — or a video file, if the sound you want is already inside a clip you have. Only the sound is used: the picture is stripped out as soon as you upload, so what gets stored is a small audio file rather than the whole video. A video with no audio track is rejected outright rather than quietly added as a silent one.

Motion Capture has no Music & Ambience section, because the video you upload already brings its own sound.

Reference photo resolution

Your reference photo sets the detail ceiling for every render made from it. Before generating, Boobot crops it to the aspect ratio you picked — and if it's smaller than that crop, it gets scaled up, which adds size but no detail.

The uploader shows the photo's true pixel size, and tells you when it's being upscaled:

432 × 768 → upscaled to 1080 × 1920

To avoid it, pick a photo at least this big for the ratio you're rendering:

Aspect ratio Minimum photo
9:16 (vertical) 1080 × 1920
16:9 (widescreen) 1920 × 1080
1:1 (square) 1920 × 1920

A photo has to clear both numbers, not just one. If you want a single photo that works for any ratio, use 1920 × 1920.

Your renders

Every image and video you generate is saved to the Renders tab — your full history, newest first. From there you can:

  • View — open any render full-size, with its prompt and timestamp.
  • Remix — click Edit on a render to reload its exact settings into the form so you can tweak and regenerate.
  • Download — save the file locally.
  • Make public — flip a render public to get a shareable link anyone can open, even signed out.
  • Delete — remove a render for good.

Edits that don't cost tokens

Not every edit needs the video generated again. If the only things you changed sit on top of the finished clip, Boobot re-uses the video it already made and just redoes the final mix — free, no tokens — saving the result as a new entry in your Renders tab and leaving the original alone. That covers:

  • background music, ambient sound and your own uploaded track
  • B-Roll
  • burned-in captions
  • the voiceover, on Motion Capture only (there the voice plays over the clip; on the other flows it drives the render itself)

Change anything the model actually saw — the prompt, the reference photo, the aspect ratio, the length, the lip-sync voice — and it regenerates from scratch at the usual cost.

Captions on older renders. Captions are built from word timings captured when the voice was generated, and every render keeps its own copy — so you can add them later, however old the render is. Renders made before this was introduced are the exception: their timings were only held for seven days, and if that window has passed Boobot will tell you so and ask you to regenerate the voice, rather than quietly handing back a clip with no captions.

Tokens

Every render spends tokens from your balance:

Render Cost
Image 10 tokens
Video — Lip Sync or Bring to Life 10 tokens per second (50 / 100 / 150 for a 5s / 10s / 15s clip)
Video — Motion Capture 12 tokens per second (60 / 120 / 180 for a 5s / 10s / 15s clip)

Motion Capture costs a little more per second because the model behind it does. Its length isn't something you pick — it's the length of the video you upload, rounded up to the next whole second, and the uploader tells you the resulting cost before you generate.

Tokens come from a one-time pack — see Pricing & billing.