Back

RAG Is a Product Design Problem, Not an Engineering Problem

5 MINS

RAG Is a Product Design Problem, Not an Engineering Problem

Every team building an AI product hits the same wall. The model is fine. The infrastructure is fine. The vibes from the demo are fine. And then real users start asking real questions and the answers get oddly confident, oddly wrong, and oddly hard to debug. The reflex is to call in the engineers and "fix the RAG." It's almost always the wrong reflex. RAG is, primarily, a *product design* problem.

What teams get wrong

The standard RAG conversation in most rooms is about chunking strategy, embedding choice, and vector store latency. Those matter. They are not the headline.

The headline is: what does the user think they're asking, and what is the system actually retrieving? That gap is a product question, not an engineering one. You can have a perfect retrieval pipeline that pulls beautifully from the wrong corpus and fails users every time.

The mismatches I've seen most often:

The user asks a "how do I…" question. The retrieval is optimised for "what is…" content.
The user is in the middle of a workflow. The retrieval doesn't know that and pulls top-of-funnel docs.
The user wants the *current* policy. The retrieval doesn't distinguish between live and deprecated content. None of those are model problems. All of them are product problems pretending to be tech problems.

The corpus is your real product

I now think about the RAG corpus the way I used to think about a product catalogue. It needs:

Curators. Real owners who decide what's in and what's out.
Versioning. A way to deprecate stale content without nuking the index.
A taxonomy. Tags or metadata that map to actual user intents.
A feedback loop. A way for users (and reviewers) to say "this answer was wrong because the source was bad." Most teams I see have a Google Drive of PDFs, an embedding job, and a hope. The drift starts on day one and compounds quietly until someone in QA finally catches it. The teams that ship reliable RAG products have a content team, even if it's only one person with that mandate part-time. That single hire is usually worth more than upgrading the model.

What the UX has to do for you

Even a perfectly-curated corpus will produce wrong answers sometimes. The UX has to do the load-bearing work the retrieval can't:

Show the source. Always. If you can't surface the source, your retrieval isn't ready.
Make disagreement cheap. A one-click "this is wrong" with a reason taxonomy is worth more than any benchmark.
Refuse confidently. A clear "I don't know" beats a fluent guess. Train your team to celebrate refusals, not penalise them. If the UI hides the seams, the seams will tear under real load. If the UI exposes the seams thoughtfully, users build a calibrated trust — they know when to lean on the system and when to push back.

The honest test

Before launching any RAG feature, I run one test: ask 30 real questions from real users and read every answer with the source side-by-side.

Not aggregate metrics. Not benchmarks. Thirty answers, by hand, on a Friday afternoon. If the answers feel right *and* the sources are defensible, you're close. If only one of those is true, you're shipping a problem.

RAG looks like a tech stack. It's actually a content operation, a UX, and a feedback loop wearing a tech costume. PMs who lead with that framing build products that survive contact with users. The rest spend the next quarter explaining why the model "got worse."

The model didn't get worse. The corpus did.

Background

Palak skipped presentations and built real AI products.

Palak Jain was part of the March 2026 cohort at Curious PM, alongside 17 other talented participants.