An agentic RAG pipeline for querying financial documents using natural language. Built as a portfolio project targeting financial data applications at enterprise scale.
- Ingestion: PDF loading (PyMuPDF) → overlap-aware chunking → sentence-transformer embeddings → FAISS + BM25 hybrid index
- Retrieval: LangGraph agent — query rewriting → hybrid search → cross-encoder reranking
- Generation: Grounded answer synthesis with source citations via Gemini
- Evaluation: RAGAS metrics — faithfulness, answer relevancy, context precision
Python · LangChain · LangGraph · FAISS · BM25 · Sentence-Transformers · RAGAS · Streamlit · Gemini API
git clone https://github.com/YOUR_USERNAME/financial-doc-qa.git cd financial-doc-qa
python -m venv venv venv\Scripts\activate # Windows source venv/bin/activate # Mac/Linux
pip install -r requirements.txt
set GOOGLE_API_KEY=your_gemini_key_here
Drop PDF files into data/documents/
python ingest.py
streamlit run app.py
This repo does not include documents or the prebuilt index (too large for GitHub).
Download any public SEC 10-K filing from google .
Or use Apple's 2023 annual report directly by google
Place downloaded PDFs inside data/documents/ then run:
python ingest.py
- "What was Apple's revenue growth in fiscal 2023?"
- "What risk factors did Microsoft highlight in their latest 10-K?"
- "How did operating margins change year over year?"
Run evaluation with: python -m evaluation.ragas_eval