A local Retrieval-Augmented Generation (RAG) API built with Python and FastAPI.
This project implements a complete RAG pipeline that retrieves relevant information from a custom knowledge base and uses a local LLM to generate grounded answers — running entirely on the local machine with zero cloud/API costs.
Documents
↓
Text Chunking
↓
Vector Embeddings
↓
ChromaDB
↓
Semantic Retrieval
↓
Relevant Context
↓
Prompt Augmentation
↓
Qwen LLM
↓
AI-Generated Answer