Answers from your own documents and data, with the source shown every time.
A retrieval system that pulls the figure from the document that actually contains it, cites it, and says what it does not hold rather than guessing. Built for your material, not pointed at a folder.
What it covers
Retrieval built for your documents
Chunking, reranking and purpose built tools per document type. Factsheets, contracts, research notes and policies are not the same shape, so they are not retrieved the same way.
Figures locked in pictures
Much real data lives inside images of charts and tables. We make those answerable, which is usually the difference between a demo and a system.
Public named, confidential protected
Public sources are attributed by name. Confidential material is used and attributed to the process, never to a person.
Honest about gaps
The agent confirms it can answer before it tries, and declines in plain words when the material is not there.
How we do it
Corpus audit first
What you have, what shape it is in, and where the figures your users ask about actually live. This is where most RAG projects go wrong, so it is where we start.
A data layer, not a chatbot
One tool per source or view, bound to what that source can answer. The agent cannot ask what the data cannot answer.
Grounded, cited, calculated
Every figure comes from a query or an approved document, carries its source, and every sum goes through a calculator rather than the model’s head.
Marked by your team, then shipped
Real questions, marked right or wrong with the reason. The marked set runs before every change and lets you swap models later.
- 1
Corpus audit first
What you have, what shape it is in, and where the figures your users ask about actually live. This is where most RAG projects go wrong, so it is where we start.
- 2
A data layer, not a chatbot
One tool per source or view, bound to what that source can answer. The agent cannot ask what the data cannot answer.
- 3
Grounded, cited, calculated
Every figure comes from a query or an approved document, carries its source, and every sum goes through a calculator rather than the model’s head.
- 4
Marked by your team, then shipped
Real questions, marked right or wrong with the reason. The marked set runs before every change and lets you swap models later.
What you get
- A retrieval layer built for your document types, in your cloud
- An answering agent that cites every figure and declines what it cannot support
- A client marked evaluation set, run before every change
- Ingestion for your sources: drives, CMS, PDFs, databases
- Source code, documentation and a handover session
Where we have done it
OnlyFacts: 190 data tools, one per published view, a calculator for every sum, and not a single number from memory. The climate data agent that newsrooms quote.
OnlyFactsIndia Avenue: roughly 200 fund documents, figures locked in chart images, and an assistant under a financial services licence that never guesses a number.
India Avenue
Questions we get asked
Our documents are a mess. Is that a problem?
It is the normal state. The corpus audit tells us what is usable as is, what needs a purpose built tool, and what is not worth ingesting yet.
Will it hallucinate?
Which vector database?
Can it stay inside our network?
Also under ai engineering
Talk it through
Tell us what you need from rag and knowledge systems.
Thirty minutes with one of our architects. We will tell you whether it is a pilot, a build, or not worth doing yet.