Making Xoopah’s AI UGC video good enough to publish
Creative view rate went from 45% to 54%, downloads from 12.5% to 22%, and each video got ~$0.50 cheaper to make. I re-engineered Xoopah’s UGC video pipeline end to end, until the output was something a small business would actually put its brand on.
- Role
- Senior Product Designer, end to end
- Company
- Xoopah, under Spursol
- Timeline
- ~3 weeks focused
- Team
- PM · Engineering · Creative
- Stack
- Nano Banana · ElevenLabs v3 · Kling · FFmpeg
Rented · 3rd-party
Owned · new pipeline
VS
Flat, not expressive
Expressive, controllable
The short version: Xoopah’s AI UGC feature was acquiring users but not keeping them, the avatar videos weren’t good enough to publish as real ads. Handed the feature end to end, I ran the R&D, made the call to stop renting a third-party engine, and designed + shipped a custom AI pipeline, validated with a working prototype I built myself before a single engineer touched it.
Xoopah is an AI creatives-and-ads tool that lets small businesses generate ad creatives and run campaigns from one interface. It sits under Spursol, funded in-house by the group’s revenue products, so we optimize for real users and unit economics, not investor demos.
The AI Video UGC feature already existed, built on a third-party avatar API. My PM handed me its improvement end to end: why to rebuild it, what to change, whether it was even feasible, the scope handed to engineering, the design, shadowing the build, and the distribution strategy.
Acquisition was healthy in Mixpanel, ads kept pulling users in, and the “tried it” line kept climbing. But almost nobody came back. So we went to the source: interviews on Maze and usertesting.com, Spursol’s own creative designers, Clarity session recordings. One message came back over and over.
“The avatars don’t feel natural enough to put our brand on. We tried it once and weren’t convinced.”
The insight that reframed the project: this wasn’t a UI problem. You can’t design your way out of an engine that doesn’t produce publishable output. The fix had to go deeper than the screens, the engine itself.
I tested the landscape first, general video models, the incumbent engine, and avatar-specific tools, against a quality bar I wrote before touching anything. Then the call: keep leaning on the rented engine, or build our own. The problem with renting was no differentiation and no control. Any user could subscribe to that third-party tool directly, so where was Xoopah’s edge? Expressiveness comes from emotion in the script and a detailed generation prompt, both impossible when a third party owns the engine.
Image → Voice
The base still from Nano Banana is passed to ElevenLabs v3, which generates the voiceover for the script.
Voice → Avatar
Image + voiceover drive Kling Avatar, which produces a lip-synced talking-head performance.
Avatar → Export
FFmpeg stitches the clips, trims dead frames, and compresses the final render for delivery.
The trade-off I owned: generation would take ~1–3 minutes longer. Our north star was output users would bet their brand on, and quality wins word of mouth. Latency is a solvable, temporary problem. A weak product is not.
Getting buy-in: I didn’t argue in the abstract. I built an internal testing tool with Claude Code that hit all the APIs, produced real videos, and put a direct output comparison in front of leadership. When they worried the extra minutes would spike drop-off, I answered with concrete mitigations (email-when-ready, progress timers) and showed the latency was temporary, solvable by hosting models in-house.
The unlock: in the old flow, users could never preview the voiceover or sync assets to it. The rented engine returned the finished video all at once, at the very end, after a blind “continue.” My pipeline generates the voice separately, so users finally hear the voiceover and arrange their assets to match it, the right visual appearing exactly when it’s mentioned.
Hardest UX problem: letting users align assets to voiceover duration without it feeling technical. We’d considered not giving them the control at all. User feedback said otherwise, so I designed the timeline to make it feel simple.
- Output finally cleared the quality bar. Good enough that our own marketing team picked one as a live creative, the first time this feature’s output made it into a real campaign.
- Engagement with the output rose. Creative view rate 45% → 54%, download rate 12.5% → 22% period-over-period.
- Repeat usage is healthy. Users average 2.25 generations each (Mixpanel, 60-day).
- Unit economics improved. ~$0.50 lower cost per video than the rented approach, while gaining control, customization and expression.
The strategic win: we moved from renting to owning the pipeline, free to experiment without a third-party dependency, with a roadmap that was impossible before: any photo → talking video, and self-avatars.
What I’d do differently: instrument retention before shipping, not after. We optimized hard for output quality, the right call, but launched without the return-rate tracking that would prove the quality brought users back. A quality bet is only as persuasive as the measurement wrapped around it.
Thanks for reading.
If you made it this far, we’d probably enjoy building something together.









