IndicRAG — document QA for Indian languages
- Problem
- Retrieval tooling assumes English. Questions and documents in Telugu, Hindi, Tamil and others get poor matches and confident wrong answers.
- Approach
- BAAI/bge-m3 dense vectors fused with BM25 through Reciprocal Rank Fusion, reranked by bge-reranker-v2-m3. The v2 rewrite turned the pipeline into a six-node LangGraph agent with six tools and a reflexion step that scores each draft for faithfulness and completeness before answering.
- Result
- 12 languages, live search across arXiv, Semantic Scholar and OpenAlex alongside your own corpus, streamed answers, and failover across any OpenAI-compatible LLM provider. Now at v2.6, which adds index reconciliation and backups.
LangGraph · HuggingFace · ChromaDB · NLLB-200 · Gemini API · FastAPI