Completed
AI Knowledge Base Chat
Upload PDFs, Word files and text documents, then ask questions about them in plain English. A FastAPI backend chunks and embeds each document with local Ollama models, stores the vectors in Qdrant and the metadata in SQL Server, and streams back answers grounded only in the retrieved passages.
PythonFastAPIOllamaQdrantSQL ServerLangChainReact 19TypeScriptTanStack Query
Source on GitHub βThe problem
Company knowledge lives in handbooks and policy PDFs that nobody reads end to end, and keyword search finds the file, not the answer. The goal: upload documents once, ask questions in plain English, and get answers drawn only from those documents β with local models, so nothing leaves the machine.
Engineering focus
- Retrieval-augmented generationPrimer
- Multi-format ingestion and chunking
- Local embeddings with Ollama
- Vector search in Qdrant
- Token streaming to React
- Upload validation
In use







How it works
- Background ingestion
- POST /api/documents saves the file and returns 202 with status pending. A background task then loads it, splits it into 500-character chunks with 100 overlap, embeds each one and upserts it β moving the row through processing to completed, or failed with the error recorded.
- Upload validation
- Extension allow-list, a 25 MB cap enforced while the file streams to disk, a %PDF- magic-byte check for PDFs, and a sanitised, UUID-prefixed filename. Anything rejected mid-stream is deleted, not left half-written.
- Grounded answers
- The question is embedded, the five nearest chunks come back from Qdrant, and llama3.2 is told to answer only from that context. An empty index returns a fixed message without calling the model; an unreachable Ollama returns a 503.
- Streaming
- /api/chat/stream returns plain-text tokens. The client reads the response stream, batches tokens into one re-render per animation frame, and aborts any earlier request when a new question is sent.
- Two stores, one delete
- SQL Server keeps each document's metadata, status and chunk text; Qdrant keeps the vectors. Deleting a document removes its SQL rows, its Qdrant points (filtered by doc_id) and the stored file.