Why it answers wrong about your company
You ask a model about your product warranty. It gives a coherent answer: "two years, as required by EU law". Except you offer three years, and it says so on your site. The model did not lie on purpose — it never had your text. It filled the gap with what is most likely in general.
That is the problem RAG (Retrieval-Augmented Generation) solves. The idea, without jargon: before answering, the system searches your documents and hands the model the relevant passages. The model answers from those.
How it works, step by step
- Chunking. Your documents — web pages, PDFs, procedures, FAQs — are cut into pieces of a few paragraphs each. Every piece keeps its source: which document, which section, what date.
- Embeddings. Each piece is turned into a list of numbers that describes its meaning. Two texts that say the same thing in different words get similar numbers. That is what allows search by meaning, not just by exact words.
- Search. When a question arrives it is converted the same way, and the system finds the closest pieces — usually 3 to 10. Good systems combine meaning-based search with classic keyword search, because product names and codes are found better exactly.
- Grounded answer. The model receives the question + the retrieved pieces + the instruction: "answer only from these texts; if the answer is not there, say you do not know". The answer can cite its source.
The model learned nothing new. It received the right material at the right moment. That is why RAG is cheaper and safer than "training" a model on your data: change a document and the answer changes today, not in a month.
What makes RAG good
- Clean sources. One true document per topic. If the price appears in three places with three values, the system will pick one at random.
- Freshness. Someone owns keeping documents current and there is a process: price changes → document changes → index rebuilds. Automatically, ideally.
- Citations. Every answer says where it came from. The customer can check, and you can see immediately which document produced a wrong answer.
- "I don't know". The most valuable behaviour. A system that says "I do not have that information, let me connect you with a colleague" beats one that answers everything.
- A narrow domain. Your company's assistant does not need to know the capital of Mongolia. A polite refusal outside its domain cuts risk enormously.
What breaks it
- Contradicting documents. The 2024 offer and the 2026 offer in the same index. The model gets both and blends them into an answer that is neither. Delete or archive what is no longer valid.
- Chunks too big or too small. Too big — you bring in noise; too small — you lose context ("the price above" without the "above").
- Questions that need calculation or aggregation. "How many orders did we have in June?" is not a RAG question. It is a database question. Good systems know to route them differently.
A real example: Florentin AI
The assistant on florentinpurcea.com and in the portal (Ask) is built exactly this way: a deterministic knowledge base — the Universal price, what a session includes, how approval works, what the systems do — plus a model that phrases the answer from it. When a question falls outside what it knows, it says so and offers a path to a human. It does not improvise a price or promise a timeline that is not in the base. Questions it could not answer are collected as a list of gaps we fill periodically — that is how the base grows, with verified answers, not with guesses.
How to evaluate before launch: 30 real questions
Do not test with questions you made up. Take 30 real questions from emails, WhatsApp and phone calls in the last two weeks. For each, write the correct answer and its source. Then run the system and record:
| Result | What it means |
|---|---|
| Correct + correct source | Good |
| Correct, wrong source | Luck — it will break |
| "I don't know" about something that is in the documents | A search or chunking problem |
| Invented answer | The worst — hold the launch until it is gone |
Repeat the test after every major document change. The 30 questions become your regression test.
What to remember
- The model gets your company wrong because it never had your text.
- RAG searches your documents and hands the model the relevant passages before it answers.
- Clean, current sources with citations and an honest "I don't know" make the difference.
- Broken PDFs and contradicting documents ruin any system.
- Test with 30 real questions before launch and after every major change.
Check yourself
Frequently asked
Is RAG the same as training the model on my data?
No. Training changes the model and is slow and expensive. RAG leaves the model unchanged and hands it the right documents at every question — change a document and the answer changes immediately.
What happens when I change a price?
If the process exists: the document is updated, the index is rebuilt, the answer is right from that moment. If it does not, the system will quote the old price forever. Freshness is a responsibility, not a feature.
Can RAG answer "how many orders did we have last month"?
Not directly. It is an aggregation question, for a database or a report. Good systems recognise it and route it to the right source instead of guessing from text.
Want Valhalla to check this for your business?
Valhalla Pulse scans the public signals of your website for free, in seconds.
Next lesson · in the "AI for business" path
AI agents vs. chatbots: when it is worth letting a model actA chatbot answers; an agent acts in your systems. What tools are, when the risk is worth it, guardrails in the right order, and the honest limits of agents in 2026.Lesson · 5 min · Intermediate · Updated 28 Sept 2026✓