Post-production, without the queue
Voice, transcription and generated video running on your own hardware — so unreleased material never leaves your network.
Six places it replaces a booking
Narration at draft speed
Generate a scratch or final voice track in seconds instead of scheduling a booth, then regenerate when the script changes.
PiperConsistent character voices
Bind a voice to a character so it stays identical across episodes, without depending on one performer’s availability.
GPT-SoVITSSubtitles from the same clock
Synthesis returns per-word timings, so captions and audio come from one source rather than being aligned afterwards.
Word timingsTranscription for the edit
Turn raw footage audio into searchable text with a self-hosted model, so unreleased material never leaves your infrastructure.
WhisperPrevisualisation video
Generate short video from a prompt or a still for mood, pitch and blocking work before anything is shot.
WAN 2.2Lip-synced dialogue
Drive a talking-head clip from an audio track for ADR previews or localisation tests.
MuseTalk
One character, from table read to cut
A character here is one record: a locked face, a bound voice, and timings that come back with the audio. Change the script and the clip regenerates — same voice, same face, same clock.
The voice itself takes about ten seconds of reference to clone — how cloning works.
- The voice is bound to the character. It stays identical across episodes and does not depend on one performer’s calendar.
- The retake is a text edit. Regenerate the line in seconds when the script changes, at draft or final quality.
- Captions share the clock. Per-word timings arrive with the audio — approximate by design, and labelled that way.
- Pre-vis without a shoot. Text or image to video for mood, pitch and blocking before anything is filmed.
- Scriptthe retake is a text edit
- VoicedPiper · GPT-SoVITS
- Timedper word · approximate
- Lip-syncedMuseTalk
The stack, with its labels on
| Capability | Engine | Status |
|---|---|---|
| Text to speech | Piper | Shipped |
| Voice cloning | GPT-SoVITS · ~10 s reference | Shipped |
| Transcription | Whisper | Shipped |
| Lip-sync video | MuseTalk | Shipped |
| Text & image to video | WAN 2.2 | Shipped |
| Word timestamps | Piper | Approximate |
| Full-body animation | MusePose | Partial |
This is an adult-focused platform
The models here are uncensored by design and the platform is built for adult content. The capabilities above work perfectly well for mainstream production, but you should know what you are buying into and what your own compliance and brand-safety review will need to cover.
Performers’ voices may only be cloned with documented, revocable consent — see the acceptable use policy. For deployment inside your own infrastructure, see Enterprise.
What an evaluator asks first
Can this run inside our own network?
Yes. Every model already runs in a container; an enterprise deployment means running those containers on your hardware, in your network, under your change control. Unreleased material never crosses a boundary you do not operate.
What do we need to clone a performer’s voice?
About ten seconds of clean reference audio, and the performer’s documented, revocable consent. The consent is a terms obligation, not a technical control — keep your own records, because the platform does not keep them for you.
Are the word timings frame-accurate?
No — they are approximate. Alignment is interpolated rather than forced, so treat the timings as close enough for captions and lip-sync drafts, not as a frame-exact conform.
Is full-body animation production-ready?
Not yet — it is partial. The face pipeline is what is production-ready today; pose-driven full-body work is in progress and is labelled that way everywhere on this site.
Is the platform suitable for mainstream production?
The capabilities work perfectly well for mainstream work, but the platform is adult-focused and its models are uncensored by design. Your compliance and brand-safety review should know that before you commit, which is why it is stated on this page.