Making Xoopah’s AI UGC video good enough to publish
Creative view rate went from 45% to 54%, downloads from 12.5% to 22%, and each video got ~$0.50 cheaper to make. I re-engineered Xoopah’s UGC video pipeline end to end, until the output was something a small business would actually put its brand on.
- Role
- Senior Product Designer, end to end
- Company
- Xoopah, under Spursol
- Timeline
- ~3 weeks focused
- Team
- PM · Engineering · Creative
- Stack
- Nano Banana · ElevenLabs v3 · Kling · FFmpeg
Rented · 3rd-party
Owned · new pipeline
VS
Flat, not expressive
Expressive, controllable
The short version: Xoopah’s AI UGC feature was acquiring users but not keeping them, the avatar videos weren’t good enough to publish as real ads. Handed the feature end to end, I ran the R&D, made the call to stop renting a third-party engine, and designed + shipped a custom AI pipeline, validated with a working prototype I built myself before a single engineer touched it.
Xoopah is an AI creatives-and-ads tool that lets small businesses generate ad creatives and run campaigns from one simple interface. It sits under Spursol, funded in-house by the group’s revenue products, so we optimize for real users and unit economics, not investor demos.
I’ve been a Senior Product Designer there for two years. The AI Video UGC feature already existed, built on a third-party avatar API. My PM handed me its improvement end to end: why to rebuild it, what to change, whether it was even feasible, the scope handed to engineering, the design, shadowing the build, and the distribution strategy for performance marketing.
We watched the funnel closely in Mixpanel. Acquisition was healthy, ads kept pulling new users in, and the “tried it” line kept climbing. But almost nobody came back.
So we went to the source: interviews on Maze and usertesting.com, feedback from Spursol’s own creative designers, Clarity session recordings. One message came back over and over.
“The avatars don’t feel natural enough to put our brand on. We tried it once and weren’t convinced.”
The insight that reframed the project: this wasn’t a UI problem. You can’t design your way out of an engine that doesn’t produce publishable output. The fix had to go deeper than the screens, the engine itself.
- Retention / repeat usage climbs
- Video downloads & published creatives go up
- Quality judged by return usage and conversion, not first-try volume
I also wrote a concrete quality bar before touching a tool: higher resolution, real facial expression, smooth delivery, human (not robotic) voice, an overall professional-feeling creative.
Before proposing anything, I went wide. In Figma Weave, with Claude generating avatar-image and video prompts, I tested Seedance 2.0, Kling v1.6 / v3 / Lip-Sync / Avatars, Creatify Aurora, Higgsfield, Veo 3, Gemini Nano Banana 2, and ElevenLabs v2 & v3.
The general-purpose video models produced stunning individual clips, and failed the business test: $2–3 per 15-second clip, a ~15s length cap when a real creative runs 15–45s, voices that turned robotic across stitched clips, and latency that multiplied with every stitch. I killed that direction on evidence, not taste.
Then I hit Kling’s Avatar model: image + voiceover + prompt → an expressive avatar video. No hard length cap, controllable quality, $1 per 30s. And because its API accepted our own image, voice, and prompt, it opened three doors the rented engine never could: fully custom avatars, future self-avatars, our own voices.
Models under test
Output
We could keep leaning on the rented engine, or build our own. The problem with renting: no differentiation and no control. Any user could subscribe to that third-party tool directly, so where was Xoopah’s edge? Expressiveness comes from emotion in the script and a detailed generation prompt, both impossible when a third party owns the engine.
Image → Voice
The base still from Nano Banana is passed to ElevenLabs v3, which generates the voiceover for the script.
Voice → Avatar
Image + voiceover drive Kling Avatar, which produces a lip-synced talking-head performance.
Avatar → Export
FFmpeg stitches the clips, trims dead frames, and compresses the final render for delivery.
The trade-off I owned: generation would take ~1–3 minutes longer than the rented engine. I weighed it deliberately, our north star was output users would bet their brand on, and quality wins word of mouth. Latency is a solvable, temporary problem. A weak product is not.
Getting buy-in: I didn’t argue in the abstract. I built an internal testing tool with Claude Code that hit all the APIs, produced real videos, and put a direct output comparison in front of leadership , new pipeline vs. the alternatives, alongside the wins in control, quality, customization, and cost.
The pushback that could’ve killed it: leadership worried the extra minutes would spike drop-off. I won it three ways, concrete mitigations (email-when-ready, in-app progress timers), reframing the goal (a full funnel is worthless if quality doesn’t bring people back), and showing latency was temporary, solvable by hosting models in-house.
The new pipeline didn’t just improve quality, it unlocked a fundamentally better UX:
- 01Goal & offer, the user states their ad-creative goal and any specific offer.
- 02Script & avatar, a script is auto-generated, and the user can regenerate, edit, rewrite, or enhance it with emotion tags, then pick a presenter from our avatar library.
- 03Background slideshow, the user uploads media into a timeline, previews the voiceover, adjusts music, and previews the full background video.
The unlock: in the old flow, users could never preview the voiceover or sync assets to it, the rented engine returned the finished video all at once, at the very end, after a blind “continue.” My pipeline generates the voice separately, so users finally hear the voiceover and arrange their assets to match it, the right visual appears exactly when it’s mentioned.
Hardest UX problem: letting users align assets to voiceover duration without it feeling technical. We’d considered not giving them the control at all, user feedback said otherwise, so I designed the timeline to make it feel simple.
Calls I’m glad I made: transitions between slideshow assets (small call, big lift, the whole video suddenly felt alive), and rewriting the vague tone/messaging options in script generation into clear, confident choices.
In the product: the clever part wasn’t “using AI”, it was architecting a custom pipeline instead of renting one, which handed us control over expression, cost, and differentiation. Emotion tags made speech expressive, and generation prompts made the avatar actually perform the script’s emotion.
In my own process:
- Claude-in-Chrome, competitor teardowns and market research.
- Founder skills in Claude, so it reasoned across business and design, not just aesthetics.
- Claude as prompt engineer, every prompt in the chain, avatars, video models, script generator, emotion enhancer.
- Claude Code, built the internal API testing tool that de-risked the whole bet.
- Design in code, no Figma handoff, our design system packaged as an npm module + Figma MCP; I vibe-coded the flow in Angular and handed engineers a working coded reference.
- Claude for the PRD, and the distribution strategy behind the launch.
“AI is a force multiplier with human direction. I offloaded the no-thinking-required work and spent my judgment where it mattered: what to build, what to cut, where quality couldn’t slip.”
Verified numbers and directional signals, separately, no inflation.
- Output finally cleared the quality bar. The rebuilt pipeline produced avatars users would actually publish, good enough that our own marketing team picked one as a live creative, the first time this feature’s output made it into a real campaign.
- Repeat usage is healthy. Users average 2.25 generations each (Mixpanel, 60-day). Trailing 90 days: 184 signed up → 118 started a generation → 147 finished an ad.
- Engagement with the output rose. Creative view rate 45% → 54%, download rate 12.5% → 22% period-over-period.
- Unit economics improved. ~$0.50 lower cost per video than the rented approach, while gaining control, customization, and expression.
The strategic win: we moved from renting to owning the pipeline , free to experiment without a third-party dependency, with a roadmap that was impossible before: any photo → talking video, and self-avatars.
7-day return rate is still ~0% and retention isn’t fully instrumented, no lifecycle email at launch, so it’s too early to claim a retention win. The credible story is leading indicators: repeat generation, engagement, and output quality. The retention instrumentation + lifecycle email is scoped as the immediate next step.
Leading this end to end, not just the design, changed how I work. I ran it like an entrepreneur trying to make something succeed. When the priority is shipping fast and failing fast, process finesse matters less than execution, collaboration, and judgment.
What I’d do differently: instrument retention before shipping, not after. We optimized hard for output quality, the right call, but launched without the lifecycle email and return-rate tracking that would prove the quality brought users back. A quality bet is only as persuasive as the measurement wrapped around it, and the evidence loop belongs in the same sprint as the feature.
Where I’d take it next: the pipeline is a platform, not a feature, any-photo-to-talking-video and self-avatars give Xoopah a scalable line of video capabilities it didn’t have before.
Thanks for reading.
If you made it this far, we’d probably enjoy building something together.











