LLM Integration: Adding Language Models to Your Existing Product

How we add language model features to an existing product: swappable models, retrieval over your data, structured outputs, evals in CI and production safeguards.

A model module sliding into a socket on an existing backend, with a spare module ready to be swapped in.

You already have a product, a backend and users, and you want to add a feature that reads or writes language: summarize a case file, answer questions from your help center, turn inbound emails into structured tickets, draft a reply for a support agent to approve. LLM integration is the engineering that makes such a feature a dependable part of your system instead of a demo wired to an API key.

Good looks like this. The feature has a clear contract with the rest of your code, and you can change the model without a rewrite. Every prompt change is tested against real examples before it ships. Cost and latency are visible per feature, failures degrade gracefully, and personal data goes only where you decided it should go.

Where LLM integration fits in your system

First, check that the feature needs a model at all. Exact lookups, calculations and validations belong in ordinary code, and our overview of AI app development covers which problems suit a model and which do not.

When a model is the right tool, we treat it like any other external dependency, such as a payment provider: called from your backend, never directly from a browser or mobile app, with credentials kept on the server. Most integrations take one of three shapes:

  • Interactive and streamed, for drafting, rewriting and question answering, where the user watches the output arrive.
  • Background jobs, for extracting, classifying and summarizing documents through the queue your backend already uses, with results written back to your database.
  • Event-driven enrichment, where a new record triggers tagging, routing or indexing for search.

The integration usually lives in a small module with a narrow interface, such as a function that takes a support email and returns a typed ticket. The rest of your code never sees prompts or provider details. We build these modules in whatever your backend runs on, and our article on backend development covers the systems around them.

Choosing a model and keeping it replaceable

Model choice is an engineering decision made with your own examples, not leaderboards. We compare candidates on quality against your evaluation set, latency, cost per task, context size, data-handling terms, regional availability and reliability. Sometimes the answer is two models: a small, fast one for routine requests and a larger one for hard cases, with routing between them.

Hosted APIs are the usual starting point because they need no infrastructure. Open-weight models that you run yourself make sense when data cannot leave your environment, when volume is steady enough to keep GPUs busy, or when you need full control over the model. They also make you responsible for serving, scaling and upgrades.

We put the provider behind an internal interface defined around your tasks, not a generic chat wrapper. Providers differ in structured output, tool calling and caching, and a lowest-common-denominator abstraction throws those differences away. Where providers offer pinned model versions, we pin them, so a model upgrade becomes a tested change like any other dependency upgrade.

Retrieval over your own data

Most useful features need knowledge the model does not have: your documentation, your policies, the user’s own records. Retrieval-augmented generation supplies it at request time, and answer quality depends more on retrieval than on the model.

Documents split into chunks, searched two ways, ranked, filtered by permission and passed on to the model.
  • Ingestion. Parse documents properly, including tables and headings, split them along their structure rather than at fixed lengths, and store each chunk with metadata such as source, date, tenant and access rights.
  • Search. Combine keyword and vector search, since each finds what the other misses, then rerank the candidates. Vector support in the database you already run, such as pgvector for PostgreSQL, is often enough before a dedicated vector store is justified.
  • Permissions. Filter results by what the user may see before anything reaches the model. Asking the model to withhold restricted content is not access control.
  • Freshness. Re-index when source content changes, triggered by events rather than occasional full rebuilds.

We evaluate retrieval on its own by checking whether the right passages come back for each test question. When the data is structured, such as orders, balances or appointments, we skip embeddings and let the model query your existing APIs instead.

Structured outputs and tool calling

When the result feeds your code, free text is the wrong format. We request output that matches a JSON schema, using the provider’s structured output mode where available, and validate it anyway: first against the schema, then against business rules such as dates in range, totals that add up and categories that exist. A failed validation triggers a retry with the error attached, then a fallback or a human review queue.

Loose model output pressed through a schema template into validated shapes, with a misfit piece sent back for a retry.

Tool calling lets the model request an action, such as looking up an order or checking availability, which your code executes with the user’s permissions and the same validation as any other request. Treat everything the model reads as untrusted, because text inside a document or email can try to redirect it; model output alone must never authorize a sensitive operation. Once a feature chooses its own sequence of actions, it is an agent, and AI agent development covers the extra controls that requires.

Prompts, evals and regression tests

Prompts are code. They live in your repository as templates with typed variables, go through code review and ship together with the model and parameters they were tested with. Every production call logs its prompt version, so any output can be traced to the configuration that produced it.

The evaluation set comes from real, anonymized inputs with expected results agreed with your domain experts. Deterministic checks score schema validity and exact fields, while a grading model scores open-ended answers against a rubric once its judgments match those of human reviewers.

The suite runs in continuous integration on every change to a prompt, model or retrieval setting, and a drop below agreed thresholds blocks the release. Because outputs vary between runs, important cases run several times and we look at the spread, not a single result.

Running it in production

Caching

Exact-match caching serves repeated requests, like the same document summarized twice, without a new call. Several providers can also cache a long shared prompt prefix, which rewards putting stable instructions and reference material first. Semantic caching, which reuses answers to similar questions, is reserved for cases where a slightly wrong reuse is harmless.

Rate limits and fallbacks

Providers enforce quotas and occasionally slow down or fail. Every call gets a timeout, retries with exponential backoff and jitter, and a circuit breaker. Background work runs through queues with controlled concurrency, so a batch job cannot starve interactive users. A fallback model, validated on the same evaluation set, takes over when the primary is unavailable, and if both fail, the product offers a clear path without AI instead of an error.

Personal data

We send only the fields a task needs, redact or pseudonymize identifiers where the task allows, choose provider settings and regions that match your obligations, and apply the same redaction and retention limits to logs and traces.

Observability

Each request produces a trace with the prompt version, model, retrieved sources, tokens, latency, cost, validation result and user feedback. Dashboards show quality signals, spend and errors per feature, with alerts on spikes.

How an engagement works

We start small: one feature, one success metric and access to real examples. A business analyst works with your domain experts to define correct behavior and build the first evaluation set, an engineer prototypes against it, and we review the results together before committing to production work.

The build team is usually one or more backend engineers with QA, plus a designer when the feature has a user interface. You get regular demos and written reports, and at handover everything is in your repository: the module, prompts, evaluation suite wired into continuous integration, dashboards and a runbook.

Frequently asked questions

Can you work with the stack we already have?

Yes. The integration sits behind a narrow interface in your backend, whatever language and framework it uses, and follows your existing conventions for queues, configuration and deployment.

Do we need a vector database?

Not necessarily. Many products start with vector search inside their existing database, or with good keyword search. A dedicated vector store earns its place at larger scale or with specialized search needs.

What happens when a provider retires a model?

Because models are pinned and wrapped behind an interface, the replacement is a tested change: run the evaluation suite on candidates, adjust prompts where needed and roll out behind a flag.

Will the provider train on our data?

That depends on the provider, the plan and the settings. We review the terms with you, configure retention and training options accordingly, and self-host a model when your requirements rule out third-party processing.

How do we keep costs under control?

Set a budget per feature, route easy requests to smaller models, keep context lean, cache where it is safe and batch background work. Per-feature cost dashboards show overspending as it happens, not at the end of the month.

If you know the feature you want and the system it has to fit into, tell us about your product, and we will propose how to integrate a model safely, measurably and without locking you into one provider.