Generative AI Development for Image, Video, Speech and Music

How we put image, video, speech and music generation into products: model and pipeline choice, safety layers, rights and provenance, queueing and GPU cost.

A triangle, circle and square projected into an image frame, a film strip and a sound waveform.

You want your product to create media, not just display it: product photos from a single shot, short video clips from a script, a narrated version of every article, background music that fits a user’s edit. Generative AI development is the work of turning a model that does this well in a demo into a feature that runs reliably, at a predictable cost, with safety and rights questions considered before launch rather than after.

Good looks like this: results are consistent enough that users trust the button, the wait is short and honest about progress, harmful requests are stopped at the input and at the output, every result can be traced to the model and settings that made it, and GPU spending follows usage instead of idle capacity.

What we build with generative models

  • Image. Product and marketing visuals for e-commerce: background replacement, variations for different channels, upscaling, and inpainting to repair or extend a photo. Also personalized imagery and design tools inside apps.
  • Video. Short clips from images or text, animated product previews, and editing aids such as automatic captions and reframing for different aspect ratios.
  • Speech. Text-to-speech narration for articles and accessibility, voices for conversational interfaces, and dubbing into other languages, usually paired with transcription and translation.
  • Music and sound. Background tracks and sound effects for user-created content, where licensing questions deserve particular care.

These features rarely stand alone. They sit inside a media product, a store or a creative app, often next to text features built on language models, which our overview of AI app development covers.

Generation is not always the answer. If you need a handful of assets that must be exactly right, commission them from a designer or license them. Generation pays off when you need volume, variation or personalization that people could not produce by hand at a sensible cost.

Choosing models and building the pipeline

There are three ways to run generation, and many products combine them:

  • Hosted APIs are the fastest to adopt and need no infrastructure, but give you less control over the model, its updates and where data is processed.
  • Open-weight models on your own GPUs give full control, custom fine-tuning and stable behavior, at the cost of running inference infrastructure.
  • Managed inference platforms sit in between: you choose and customize the model, and they run the GPUs.

Read each model’s license before building on it. Commercial use terms vary widely, and some open-weight models restrict commercial use or attach conditions to it.

A production feature is usually a pipeline rather than a single model. A product photo feature might segment the product, generate a new background, relight and composite the result, upscale it, run moderation and attach provenance metadata. We prototype image pipelines quickly in node-based tools such as ComfyUI, then implement the production version as versioned code in which each stage can be tested on its own.

When a brand needs a consistent style or a specific product rendered faithfully, lightweight fine-tuning with LoRA adapters, trained on images you have the rights to use, usually beats ever-longer prompts. We choose between candidate models through blind side-by-side comparisons on your own content, rated by your team.

Prompt and parameter design

Unless your product is a tool for creative professionals, users should not have to write prompts. We translate intent into structured controls, such as style presets, reference images, aspect ratio, duration, voice and mood, and turn those into prompts through versioned templates. A language model can help by expanding a short request into a detailed prompt, using the practices described in LLM integration.

Parameters matter as much as words. Control inputs such as pose, depth or edge maps keep a composition stable, step counts and resolution trade quality against time, and fixed seeds, where the model exposes them, make results reproducible enough to debug.

Every output is stored with its full recipe: model and version, template version, seed, parameters and inputs. That recipe lets you reproduce a result, explain it, or find every asset made with a model you later decide to retire.

Safety, moderation and human review

Some users will try to make a generative feature produce harmful content, so safety works in layers:

Generated shapes falling through stacked filter layers, with rejected pieces diverted and one set aside for human review.
  1. Input checks. Classify prompts and uploaded images before generation, block known abuse patterns and rate-limit accounts that probe the filters.
  2. Generation constraints. Use the model’s own safety settings and limit what templates can request.
  3. Output checks. Run classifiers on every image, clip or audio file for sexual content, violence and the other categories your policy covers, and match hashes against known child sexual abuse material through established industry programs.
  4. Likeness and voice rules. No generation of identifiable real people without their consent, and voice cloning only with verified, recorded consent from the speaker.
  5. Reporting. Let users flag outputs, and handle every report through a defined takedown process.

Human review covers the rest: a queue for borderline cases, and review before publication for anything that goes out under your brand. Thresholds are tuned on your own data, because false positives frustrate legitimate users and false negatives harm people. Reviewers also need good tools and limits on their exposure to disturbing material.

Rights and provenance

We are engineers, not lawyers, and nothing here is legal advice. There are questions every generative product should put to counsel early: the license terms of each model, what is known about its training data, who owns or can protect generated output in your markets, how close outputs may come to existing works, trademarks or artists’ styles, and what consent voices and likenesses require. Music raises these questions most sharply.

Our part is to make the answers enforceable in software. We record which model and license produced every output, restrict features to the models cleared for your use, attach provenance metadata such as C2PA Content Credentials, apply visible labels where your policy calls for them, use invisible watermarking where the model or platform supports it, and keep consent records for every cloned voice or trained likeness.

Latency, queueing and GPU cost

Generation is slow compared with a normal web request: seconds for a fast image model, often minutes for video. That calls for an asynchronous design, the kind of queue-based system we build in backend development. The API creates a job and returns at once, a queue feeds GPU workers, results go to object storage behind a CDN, and the client hears about progress through WebSockets, server-sent events or a push notification.

Jobs flowing from a queue to scalable GPU workers, then to storage and a phone, with progress reported back.

Jobs support cancellation, safe retries and priority lanes, so paying customers are not stuck behind a bulk run. Where a model can produce a quick low-resolution preview, users see that first.

GPU cost is driven by time on the GPU, not by the number of requests. The main levers are:

  • Autoscaling on queue depth, with a deliberate choice between slow cold starts and paying for idle capacity.
  • Smaller or distilled models where quality allows, and batching compatible jobs.
  • Limits on resolution, duration and retries for each plan.
  • Caching and reusing outputs that many users request.
  • Hosted APIs for low or spiky volume, and your own GPUs once demand is steady enough to keep them busy.

We track cost per accepted output, including the generations users discarded, because that is what the feature really costs.

How a generative AI development project runs

Discovery defines the use case, the quality bar, the safety policy and the list of rights questions for your counsel. We then prototype two or three candidate pipelines and run a blind comparison with your team on your own content.

The build covers the queue and workers, storage and delivery, moderation, the review queue and an interface that handles waiting well. QA tests output quality on a fixed set of inputs, safety with adversarial prompts, and the whole system under load. Launch comes with quotas and a gradual rollout, and we iterate on real usage and cost data. You work with a business analyst, a designer, engineers and QA testers, see regular demos and written reports, and own the code, pipelines, templates and evaluation sets.

Frequently asked questions

Should we use a hosted API or run our own models?

Start with hosted APIs unless data residency, customization or licensing rules them out. Move steady, high-volume workloads to your own GPUs when the numbers show it pays, and keep the pipeline flexible enough to switch.

Can the output match our brand’s style?

Yes, through reference images, carefully built presets and fine-tuned adapters trained on assets you have the rights to use. We test style consistency on a fixed set of inputs every time something in the pipeline changes.

Who owns the generated content?

That depends on the model’s terms and the law in your markets, so it is a question for your lawyers. We give them the technical facts they need, such as which model, license and inputs produced each output.

Can users clone their own voice?

It can be done responsibly, with verified consent from the speaker, safeguards against cloning someone else’s voice and clear limits on how the voice is used. Without those safeguards, we recommend against offering it.

If you are planning a feature that creates images, video, voice or music, the useful first step is a short prototype on your own content, with safety and cost on the table from day one. Tell us about your product, and we will show you what a realistic pipeline for it looks like.