PRODUCTION
RAG SYSTEM

A production-grade Retrieval-Augmented Generation pipeline with hybrid BM25 + pgvector retrieval, RAGAS evaluation, LangSmith observability, and full GCP deployment via Terraform + GitHub Actions.

3×
Retrieval Modes
4
RAGAS Metrics
1024d
Embedding Dims
∞
Doc Capacity

ARCHITECTURE

Ingestion, retrieval, and evaluation pipelines running on GCP Cloud Run backed by PostgreSQL + pgvector and Redis caching.

CLIENT INTERNET HTTP CORE ENGINE FastAPI Document Loader PDF · TXT → Chunks Embedder BAAI/bge-large (1024d) RRF Fusion Reciprocal Rank Fusion RAGAS Evaluator Faithfulness · Relevancy dense sparse cache POSTGRES 16 pgvector cosine similarity Cloud SQL IN-PROCESS BM25 Index rank-bm25 · TF-IDF CACHING Redis 7 OBSERVABILITY LangSmith trace · eval · monitor CI/CD GitHub Actions test → build → push → deploy Terraform · Cloud Run GCP Cloud Run serverless · auto-scale deploy
📥
Ingestion Pipeline
Upload PDF/TXT → extract → chunk (fixed or semantic) → embed with BAAI/bge-large-en-v1.5 → store in pgvector + update BM25 index → invalidate Redis cache.
🔍
Hybrid Retrieval
Dense (pgvector cosine) + Sparse (BM25 TF-IDF) results fused via Reciprocal Rank Fusion. Redis caches query results with configurable TTL for sub-millisecond repeat queries.
📊
RAGAS Evaluation
Faithfulness, Answer Relevancy, Context Recall, Context Precision — computed per sample with LLM-as-judge (OpenAI) or extractive fallback. All runs traced to LangSmith.

TECH STACK

⚡
FastAPI
api/main.py
Backend
🐘
PostgreSQL 16
+ pgvector ext
Storage
🤗
BAAI/bge-large
1024-dim embeddings
ML Model
📈
BM25
rank-bm25 library
Retrieval
⚗️
RAGAS
evaluation framework
Eval
🔴
Redis 7
query + embedding cache
Cache
🔭
LangSmith
tracing + monitoring
Observability
☁️
GCP Cloud Run
serverless deployment
Infra
🏗️
Terraform
infrastructure as code
IaC
🔄
GitHub Actions
CI/CD pipeline
DevOps
🐳
Docker
containerization
Runtime
🧪
pytest
async test suite
Testing

LIVE DEMO

Connected to the production API. Upload a document, query it with three retrieval strategies, and evaluate quality with RAGAS.

Document Ingestion

Drop a file or browse

PDF or TXT · up to 50 MB

512
uploading... 0%
—
doc_id: —
—
Chunks
—
Strategy
—
ms
Recent Queries
No queries yet — run one below
Query Configuration
pgvector cosine + BM25 → RRF fusion
5
⌕

Results appear here after a query

Side-by-Side Mode Comparison

Run the same query across all three retrieval strategies simultaneously and compare ranking differences.

Dense —
Run comparison to see results
Sparse —
Run comparison to see results
Hybrid (RRF) —
Run comparison to see results
RAGAS Evaluation

Evaluate retrieval quality with four RAGAS metrics. Add multiple samples for aggregate scoring.

API ENDPOINTS

Full OpenAPI docs available at the live Swagger UI.

GET /health
Health check with dependency status for Postgres, Redis, and LangSmith.
statusstring"ok" when all services up
postgres_connectedbool
redis_connectedbool
POST /ingest/
Upload a PDF or TXT file. Returns document ID and chunk info.
fileFilePDF or plain text
chunking_strategystrfixed | semantic
chunk_sizeintcharacters per chunk
chunk_overlapintoverlap in characters
POST /query/
Retrieve relevant chunks using dense, sparse, or hybrid retrieval.
querystr1–1000 characters
retrieval_modestrdense | sparse | hybrid
top_kint1–50, default 5
score_thresholdfloat0.0–1.0
POST /evaluate/
Run RAGAS evaluation over live retrieval. Supports batch samples.
sampleslistquestion + optional ground_truth
retrieval_modestrdense | sparse | hybrid
top_kintchunks per sample