Skip to Content
AiServa
Private RAG Knowledge Base

Company knowledge, answered with sources.

A private RAG knowledge base searches your own documents, then answers from what it found, citing file, page and section. On AiServa the whole pipeline, from OCR to the embedding model, vector store and search, runs on a server you own by default; a cloud embedding API or a Qdrant vector store are optional.

A monitor showing an AI answer with numbered citations linked to highlighted passages in three source documents, beside a compact AiServa server

Two phases: index once, answer many times.

Indexing, in the background

  1. cableCollect. Upload files or a ZIP into the right knowledge base, or save them straight from a task.
  2. document_scannerParse and OCR. Text, headings and tables out of PDF, Office, email and scans.
  3. segmentChunk. Split into self-contained passages that keep their headings.
  4. scatter_plotEmbed. Turn each passage into a vector with a local embedding model.
  5. storageStore. Save vectors, text and metadata (source, page, date, permissions) in the vector database, beside a BM25 keyword index.

Answering, per question

  1. help_outlineAsk. A signed-in user asks in English, Bahasa Malaysia or Chinese.
  2. manage_searchRetrieve. The question is embedded; vector and keyword search run together, filtered by what this user may see.
  3. mergeFuse. Both result lists are merged with reciprocal rank fusion.
  4. sortRerank. A cross-encoder rescores the top candidates for true relevance.
  5. format_quoteAnswer. The language model answers from the best passages only, with numbered citations, or says the documents do not cover it.

Answer generation time depends on the language model and the server it runs on; retrieval itself is typically the fast part.

Inside AiServa

Inside the AiServa Knowledge Engine.

The sections above describe RAG in general. This is what AiServa Knowledge Bases actually do, step by step. Document text, chunks and vectors are stored on your organisation's own AI server, not on AiServa's.

Printed documents on a pale oak table turning into small glowing passage cards that flow into a graphite AiServa server, with the top three ranked results highlighted in violet above it

Who Can Search What

Each knowledge base is Named People Only, Department or All Staff, with Allow or Deny per person or role. Restricted documents are filtered again by the server on every search.

50 MB

per original file

8,000,000

characters per document

5,000

PDF pages per file

25 / 100 MB

files / size per ZIP import

1,000

scanned pages per person per day

Reading the Files

PDFDOCXXLSXPPTXTXTMDCSVTSVJSONHTMLXMLLOG
  • PDFs are read in the browser, up to 5,000 pages. Scanned pages are sent as images to the knowledge base's image-reading model, which transcribes them.
  • Word, Excel and PowerPoint files are uploaded in pieces and parsed on the server.
  • An optional OCR Reader runs on your own server for scanned PDFs, up to 25 MB per file.
  • The original file is kept on your server with a SHA-256 fingerprint, so people can view, download or print the exact source.
  • Word and Excel files open in the browser with their layout, sheet tabs included; PowerPoint opens as its text; a cited PDF opens at the cited page.

Chunking That Respects Structure

  • Passages of about 1,200 characters, never cut in the middle of a sentence.
  • A table is kept whole, or split between rows with its header row repeated on each part.
  • The nearest heading is stored with each passage as its section.
  • At search time, up to 500 characters of the neighbouring text on each side are added back, so the model reads the passage in context.

How a Search Runs

  1. 01

    Two Searches

    Keyword (BM25) top 60 and meaning (vector) top 60, side by side.

  2. 02

    Fuse

    Reciprocal Rank Fusion with k = 60 merges both lists.

  3. 03

    Boost

    Extra credit for exact identifiers, word coverage and heading matches.

  4. 04

    Floor

    Passages below a similarity floor set for each embedding model are dropped.

  5. 05

    Spread

    At most 3 passages from any one document.

  6. 06

    Rerank

    Optional: a language model reorders the shortlist.

  7. 07

    Top 8

    The best 8 by default go to the answer model.

Each answer numbers its passages [1], [2] and so on, and only the sources the answer actually cites are kept under it: one chip per document with its match percentage, which opens the original. Name a knowledge base in the question and only that one is searched. Restricted documents a person may not see never take a place in the ranking.

ChoiceDefaultOptional
Embedding ModelA local model on your server: qwen3-embedding (0.6B, 4B or 8B), embeddinggemma, nomic-embed-text, mxbai-embed-large, bge-m3, snowflake-arctic-embed2 or granite-embeddingAny OpenAI-compatible embedding API with your organisation's own key, never for knowledge bases marked Local Only
Vector StoreBuilt-in store on your AI server, with the keyword index beside itQdrant: only vectors and ids are sent, never text; one collection per vector size; search falls back to the built-in store if Qdrant is unreachable. Local Only knowledge bases always stay built-in.
Managing DocumentsThe Knowledge Bases panel beside the AI Agent: search, upload, create, rename, move, copy, re-index, email or ZIP documents, with Read, Contribute or Manage rightsSave any file a task produced straight to a knowledge base from the Task Summary panel
Who SearchesThe AI Agent searches the knowledge bases a person may useSmart Routing adds a Knowledge Helper that searches for the Main Model and keeps only passages above a relevance floor

See every stage in the RAG Lab.

Ask a knowledge base one question and watch retrieval happen: the question's vector, the keyword and vector lists, how they were fused, what the floor dropped, how a reranker scored each passage, and the answer that cites them. It is the same search AI Agent runs, shown in full.

  1. 01

    Question

    Clause numbers and codes are also searched as exact phrases.

  2. 02

    Question Vector

    The embedding model, its dimensions and the first numbers of the vector.

  3. 03

    Keyword Search

    The BM25 list with each passage's score and the index query.

  4. 04

    Vector Search

    The nearest passages by cosine similarity, as a percentage.

  5. 05

    Fusion

    Keyword rank, vector rank, the fused score and what happened to each passage.

  6. 06

    Filter

    The similarity floor for this embedding model and the passages it dropped.

  7. 07

    Rerank

    An AI model scores each passage from 0 to 100 and reorders them.

  8. 08

    Retrieve

    The passages given to the model, with the text around each one.

  9. 09

    Answer

    A grounded answer with [n] citations linked to each passage.

Chunks and Vectors

Open any document to see how it was chunked: each passage with its page, section path, length and a preview of its vector, and write a cited summary of the whole document.

Tune Before You Roll Out

Try the floor AI Agent uses or a fixed one, rerank off, with the knowledge base's own model or with AI scores, and the top 3 to 10 passages, and see how the answer changes.

Same Rights, Same Limits

People see only the knowledge bases and documents they may read. A Local Only knowledge base never sends a passage to a cloud model. Model steps are limited to 200 per person and 2,000 per organisation a day.

The RAG Lab and cited answers, on real screens.

Captured from AiServa with a sample library of 31 policy, contract, finance and site documents.

AiServa RAG Lab running a question through the search pipeline: summary of the knowledge base, nine stages and the fusion table with keyword rank, vector rank and result
RAG Lab: every stage of one search, with the fusion table
AiServa RAG Lab answer with numbered citations and sources showing a match percentage for each passage
A grounded answer with [n] citations and match percentages
AiServa AI Agent comparing two versions of a leave policy in a table with a citation in each row and two Sources chips showing 69% and 68% match
AI Agent: sources under the answer with their match percentage
AiServa document viewer showing an Excel rate card in the browser with sheet tabs and Download, Print and Show as Text buttons
Word and Excel originals open in the browser with their layout

See All 40 Screens

Choosing the embedding model.

The embedding model decides what "similar" means. For Malaysian companies the deciding factor is usually language: a question typed in Bahasa Malaysia has to find a policy written in English, and a Chinese supplier email has to match an English specification. Any modern embedding model fits this architecture; the table shows popular open examples that run locally, not a fixed list.

Printed documents dissolving into glowing points of light that cluster above the table, illustrating text becoming embedding vectors
Example ModelDimensionsMax InputLanguagesWhy Pick It
BGE-M3 (BAAI)1,0248,192 tokens100+ multilingualDense, sparse and multi-vector retrieval from one model; long passages; strong cross-lingual matching.
multilingual-E5-large1,024512 tokensAbout 100Proven, well understood multilingual baseline for short passages.
Qwen3-Embedding (0.6B / 4B / 8B)up to 1,024 / 2,560 / 4,09632K tokens100+ multilingualScales from a small server to a large one; strong on Chinese and English; adjustable output dimensions.
nomic-embed-text v1.5768 (Matryoshka, down to 64)8,192 tokensMainly EnglishCompact and fast for English-only collections; dimensions can be truncated to save storage.

Figures are the models' published specifications. Other capable open families include GTE, Jina and Snowflake Arctic Embed. Cloud embedding APIs (for example from OpenAI, Cohere or Google) also exist, but they send your text outside. AiServa uses a local model by default; an organisation may choose an OpenAI-compatible embedding API for knowledge bases that are not marked Local Only. Final choice is made by measuring retrieval accuracy on your own documents and questions.

Match the Language Mix

Mixed English, Bahasa Malaysia and Chinese collections need a genuinely multilingual model, tested with questions in each language against documents in the others.

Match the Passage Length

Long policies and contracts benefit from models that accept long inputs; a model with a 512-token limit silently ignores anything beyond it.

Plan for Re-Embedding

Vectors from different models are not comparable. Changing model means re-embedding the collection, so choose with a test on your own documents first.

Knowledge bases, as your people see them.

AiServa Knowledge Bases page
Readers, document counts and server for each knowledge base.
AiServa Knowledge Health
Knowledge Health flags failed or stale documents.

See All 40 Screens

Chunking: the most underrated decision.

Retrieval can only return what chunking created. Cut a clause in half and neither half answers the question.

StrategyHow It WorksBest For
Fixed size with overlapEqual token windows, for example 500 tokens with 50 to 100 overlappingPlain text with little structure
Structure-awareSplit on headings, clauses and sections, keeping the heading path with each chunkPolicies, contracts, manuals, specifications
Table-preservingKeep each table, or each row with its header, as one unitInspection reports, price lists, BQs, schedules
Parent and childRetrieve on small chunks, hand the model the larger parent section for contextLong technical and legal documents
Per recordOne chunk per email, ticket, NCR or claim, with its fields as metadataRegisters, inboxes and case records

Choosing the vector store.

"Vector database" is a category, not a product. Vectors can live in a dedicated vector database, in a database you already run, in a search engine, or in an embedded library. The right kind depends on scale, filtering needs and what your IT team already operates; the brand comes second.

A compact graphite AI server with a violet backlit honeycomb front inside a tidy server cabinet, where the vector index lives
storage

Dedicated Vector Database

A server built only for vector search, with rich filtering and quantisation.

Examples: Qdrant, Milvus, Weaviate, Chroma

table_view

Vectors in Your Existing Database

A vector type added to a database you already back up, secure and monitor.

Examples: PostgreSQL + pgvector, MariaDB and MySQL vector types, SQL Server, Oracle

manage_search

Search Engine With Vectors

A full-text search engine that also stores vectors, so keyword and semantic search share one index.

Examples: Elasticsearch, OpenSearch, Apache Solr, Vespa

inventory_2

Embedded Library or File Store

Runs inside the application with no separate server; ideal for pilots and single servers.

Examples: LanceDB, FAISS, hnswlib, sqlite-vec

Popular Self-Hosted Examples, Compared

ExampleKindIndexStrengthsTypical Fit
PostgreSQL + pgvectorExtension to PostgreSQLHNSW, IVFFlatVectors beside relational data; SQL joins and filters; familiar backups and toolingTeams already on PostgreSQL; up to millions of passages
QdrantDedicated server (Rust)HNSW with payload filteringFast filtered search, scalar, product and binary quantisation, sparse vectors for hybrid searchMany departments, heavy permission filtering, large collections
LanceDBEmbedded, file basedIVF-PQ, disk basedNo separate server; low memory; data versioningA single compact server, pilots, departmental bases
MilvusDistributed clusterHNSW, IVF, DiskANN and moreHorizontal scale to very large collectionsCompany-wide, multi-server deployments with very large volumes
WeaviateDedicated server (Go)HNSW, flatBuilt-in hybrid search, multi-tenancy, pluggable modulesMulti-department knowledge bases wanting hybrid search out of the box
ChromaEmbedded or single serverHNSWVery simple developer API, quick to startPrototypes, pilots, small team knowledge bases
Elasticsearch / OpenSearchSearch engine with vector fieldsHNSW (k-NN)Mature BM25, aggregations and security, with vectors in the same engineCompanies already running Elastic or OpenSearch for search or logs
FAISSLibrary, not a databaseFlat, IVF, HNSW, PQVery fast raw similarity search, fine-grained controlCustom pipelines where storage and filtering are handled elsewhere

HNSW, the Most Common Index

HNSW builds a layered graph linking each vector to its nearest neighbours. A search enters at the sparse top layer and walks down toward the query, visiting a tiny fraction of vectors. Two settings trade accuracy for speed and memory: M, the links per node, and ef, how widely the search explores.

Quantisation in Plain Terms

Vectors are normally 32-bit floats. Scalar quantisation stores them as 8-bit integers, about a quarter of the memory; binary quantisation keeps one bit per dimension. Accuracy lost is usually recovered by rescoring the top results with the full vectors.

Search index and standing memory.

AiServa Vector Databases page
Built-in or Qdrant vector store.
AiServa Memories page
Standing instructions the AI reads before every task.

See All 40 Screens

Popular RAG stacks, as examples.

A RAG system is a set of layers, and each layer has several good options. These are common complete combinations seen in real deployments. None is required: the right stack is chosen per company, and components can be swapped later without starting over.

dns

Open Source, Fully Local

The most common private stack on a single AiServa server.

Orchestration
LlamaIndex or LangChain
Model serving
Ollama or llama.cpp
Embeddings
BGE-M3 or multilingual-E5
Vector store
Qdrant or Chroma
Reranker
bge-reranker
table_view

PostgreSQL-Centric

For teams that want one database to back up and secure.

Orchestration
LangChain, LlamaIndex or custom
Model serving
Ollama or vLLM
Embeddings
Any open model
Vector store
PostgreSQL + pgvector, with PostgreSQL full-text search for keywords
Reranker
Cross-encoder
manage_search

Search-Engine-Centric

For companies already running enterprise search.

Orchestration
Haystack or LangChain
Model serving
vLLM or Ollama
Embeddings
Any open model
Vector store
Elasticsearch or OpenSearch, BM25 and vectors in one index
Reranker
Cross-encoder
hub

High-Scale, Multi-Server

Company-wide deployments with very large document estates.

Orchestration
LlamaIndex, Haystack or custom services
Model serving
vLLM or Text Generation Inference
Embeddings
Qwen3-Embedding or BGE-M3
Vector store
Milvus or Weaviate clusters
Reranker
Dedicated reranker service
bolt

Lightweight Pilot

Smallest footprint, no extra servers, fast to stand up.

Orchestration
A lean Python pipeline
Model serving
llama.cpp or Ollama
Embeddings
A small open model such as nomic-embed-text
Vector store
LanceDB, FAISS or sqlite-vec
Reranker
Optional
cloud

Managed Cloud, for Comparison

Popular hosted options. Convenient, but documents and questions leave your premises.

Examples
Azure AI Search, Amazon Bedrock Knowledge Bases, Google Vertex AI Search
Vector services
Pinecone, managed Weaviate, managed Qdrant, Elastic Cloud
Embeddings
OpenAI, Cohere, Google embedding APIs
Trade-off
Less to run, but not private or air-gappable

The RAG Component Landscape

Every layer is interchangeable. AiServa deployments use the self-hosted column; the cloud column is shown so the names are familiar.

LayerSelf-Hosted Examples (Private)Managed Cloud Examples (Data Leaves Premises)
Parsing and OCRApache Tika, Unstructured, Docling, Tesseract, PaddleOCRAzure Document Intelligence, Amazon Textract, Google Document AI
Embedding modelBGE-M3, multilingual-E5, Qwen3-Embedding, nomic-embed-text, GTE, JinaOpenAI, Cohere, Google, Voyage embedding APIs
Vector storepgvector, Qdrant, Milvus, Weaviate, Chroma, LanceDB, Elasticsearch, OpenSearch, FAISSPinecone, Azure AI Search, Amazon OpenSearch Service, Vertex AI Vector Search
Rerankerbge-reranker, Jina reranker, ms-marco cross-encodersCohere Rerank, Voyage rerank
Language modelQwen, Llama, Gemma, Mistral families (open weight)GPT, Claude, Gemini (hosted)
Model servingOllama, llama.cpp, vLLM, Text Generation InferenceProvider APIs
OrchestrationLlamaIndex, LangChain, Haystack, custom codeManaged RAG services from cloud providers
EvaluationRAGAS, TruLens, DeepEval, custom harnessCloud evaluation tooling

Product names are trademarks of their owners and are listed as examples only. AiServa is not affiliated with, or endorsed by, any of them.

Size your vector index.

An estimate of passages and raw vector storage from your document estate. Real indexes add graph and metadata overhead, so plan headroom on top; sizing is confirmed during scoping.

vectors = documents × pages × chunks per page
storage = vectors × dimensions × bytes per value

Vectors

-

Raw Vector Storage

-

With 1.5× Headroom

-

What makes answers actually right.

join_inner

Hybrid Search

Vector search finds meaning; BM25 finds exact part numbers, clause numbers and names. Results are merged with reciprocal rank fusion, score = Σ 1 / (60 + rank).

sort

Reranking

A cross-encoder such as bge-reranker-v2-m3 reads question and passage together and reorders the top 20 to 50, so only the best few reach the answer model.

lock_person

Permission Filtering

Each chunk carries its source document's access list, and the filter is applied inside the search, never after, so restricted text never reaches the model.

update

Freshness

Changed files are re-indexed and superseded versions removed, so an answer never quotes last year's price list.

Three colleagues reviewing a printed sheet of test questions marked with ticks and crosses beside a laptop showing a retrieval accuracy chart
Accuracy is measured on real questions from your own documents, before rollout and after every change.

How We Measure It

MetricWhat It Tells You
Recall at kHow often the correct source passage is among the top k retrieved. If retrieval misses it, no model can answer correctly.
Mean reciprocal rank (MRR)How high the first correct passage ranks, on average. Higher means less noise reaches the model.
nDCGRanking quality when several passages are relevant to different degrees.
FaithfulnessWhether every statement in the answer is supported by the retrieved passages.
Citation accuracyWhether each citation points to the passage that actually supports the sentence.
Abstention rateHow often the assistant correctly says the documents do not cover a question, instead of guessing.

Common Failures, and the Fix

The right document exists but is never retrieved

Add hybrid search, fix chunk boundaries, or test a stronger multilingual embedding model.

Answers quote an out-of-date version

Re-index on change, store version and date metadata, and prefer the latest effective version.

Tables come back as jumbled text

Use layout-aware parsing that keeps table structure, and chunk per table or per row with headers.

The model adds facts not in the sources

Tighter grounded prompts, citation checks, and an explicit instruction to abstain when sources are silent.

Scanned documents are invisible

Run OCR during indexing and flag low-confidence pages for a human to check.

Users see answers from files they cannot open

Carry source permissions into every chunk and filter inside the vector search.

RAG or fine-tuning?

QuestionRAGFine-Tuning
Shows its sourcesYes, every answerNo
Picks up a changed documentOn re-index, minutesOnly after retraining
Respects document permissionsYes, per userNo, knowledge is in the weights
Good at house style and formatVia templates and promptsYes
Where to startHereLater, for style, if needed

RAG Questions, Answered

What is a RAG knowledge base?expand_more

A retrieval-augmented generation (RAG) knowledge base answers questions by first searching your own documents for the most relevant passages, then having a language model write an answer from only those passages, citing each source. It lets AI answer from company knowledge without retraining a model on it.

What is an embedding model?expand_more

An embedding model converts a passage of text into a vector, a list of several hundred to a few thousand numbers, positioned so that passages with similar meaning sit close together. Questions are embedded the same way, and the nearest passages are retrieved as candidate sources.

What is a vector database?expand_more

A vector database stores embeddings with their metadata and finds the vectors nearest to a query quickly, using an approximate nearest neighbour index such as HNSW rather than comparing against every vector. It can be a dedicated vector database, a vector extension of an existing database, a search engine with vector fields, or an embedded library. Popular examples include PostgreSQL with pgvector, Qdrant, Milvus, Weaviate, Chroma, Elasticsearch or OpenSearch, LanceDB and FAISS.

Is RAG better than fine-tuning a model on our documents?expand_more

For company knowledge, usually yes. RAG answers from the current version of each document, shows its sources, respects document permissions and updates as files change. Fine-tuning bakes knowledge into model weights, cannot cite sources, and must be repeated when documents change. The two can be combined, but RAG is the right starting point.

How much storage does a vector index need?expand_more

Raw vector storage is the number of passages multiplied by the embedding dimensions and the bytes per value. One million passages at 1,024 dimensions in 32-bit floats is about 4.1 GB, or about 1 GB with 8-bit scalar quantisation, before index overhead. The calculator on this page works it out for your numbers.

Does AiServa send documents to an embedding API?expand_more

Not by default. AiServa embeds with a local model on your own server. An organisation may choose an OpenAI-compatible embedding API with its own key for knowledge bases that are not marked Local Only; Local Only knowledge bases always embed locally and always use the built-in vector store.

Can I see how AiServa found a passage?expand_more

Yes. The RAG Lab follows one question through every stage: the question vector, the keyword and vector lists, reciprocal rank fusion, the similarity floor, an optional rerank with 0 to 100 scores, the passages retrieved and the cited answer. The Chunks and Vectors view shows how each document was split and stored.

Can staff open the document an answer cites?expand_more

Yes. Each cited document appears under the answer with its match percentage and opens the original: a PDF at the cited page, Word and Excel files with their layout in the browser, and PowerPoint as its text. Only people who may read the document can open it.

How do you measure whether RAG answers are right?expand_more

With a test set of real questions whose correct source passages are known. We measure retrieval recall at k, mean reciprocal rank, answer faithfulness to the retrieved passages and citation accuracy, before rollout and after every change to models, chunking or data.

Test RAG on your own documents.

Create a workspace, upload a folder of real documents, and ask twenty real questions. Check every answer against its citation. Free for 14 days.