University RAG Platform is an enterprise-grade retrieval and question-answering system engineered for academic institutions. Beyond basic chatbot interfaces, it provides full admin observability, automated retrieval evaluation suites, and self-healing dual-engine fallbacks to query faculty rosters, research labs, and course syllabi in everyday English.
Most naive RAG setups fail in production: they hallucinate details, mix up similar departments (such as confusing the Computer Science department head with the IT head), or crash when cloud APIs hit rate limits. This platform eliminates those bottlenecks with a multi-stage retrieval architecture, automated quality benchmarking, and live telemetry controls.
Key Platform Capabilities
- Admin Telemetry & Observability Dashboard: Full administrative portal to inspect incoming queries, view confidence distributions, diagnose retrieval issues, and customize pipeline parameters in real time.
- Automated Evaluation Suite: Built-in benchmark runners that calculate retrieval accuracy (Mean Reciprocal Rank, Hit Rate, Noise Rate) and classification quality (Precision, Recall, F1-scores).
- Resilient Dual-LLM Architecture: Serves high-speed responses via Google Gemini in the cloud, with instant zero-downtime failover to an offline, quantized EXAONE model running locally on the CPU via Llama.cpp.
- Smart ML Pre-Filtering: Four independent classification models predict query type, intent, and department before vector search, locking in precise database filters to prevent cross-department bleed.
- Hybrid Lexical & Semantic Retrieval: In-memory BM25 keyword matching fused with dense vector search using Reciprocal Rank Fusion (RRF) for optimal candidate scoring.
- Full Knowledge Base Administration: REST endpoints and dashboard tools to ingest documents, update chunk metadata, purge stale records, and manage vector indices dynamically.
Performance Metrics
- ~85% Routing Accuracy: Intent and category classification precision across campus queries.
- ~0.90 Top Match Rate (MRR): Highly accurate first-rank retrieval across structured faculty and course datasets.
- Zero-Downtime Resilience: Seamless automatic switch to local offline CPU inference whenever external network limits occur.
Technical Architecture
1. Intent-Aware Document Chunking
Instead of arbitrary character splitting, documents are structured into a three-level hierarchy (Type ➔ Category ➔ Topic) and divided into three intent-specific formats:
list: Bullet points optimized for faculty rosters and course catalogs.count: Dedicated helper chunks (such as “Total faculty members: 15”) for counting queries.detail: Rich paragraphs describing specific labs, study facilities, and research initiatives.
2. Machine Learning Pre-Filter
Each query is vectorized through a stacked feature representation combining dense embeddings (384d) and sparse TF-IDF vectors. High-confidence predictions apply hard ChromaDB filters, while low-confidence queries gracefully fall back to broader semantic search.
3. Zero-Roundtrip Hybrid RRF Pipeline
The engine retrieves candidate chunks from ChromaDB, builds an in-memory BM25 index over the top candidates, and fuses lexical and semantic scores:
RRF_Score(doc) = (w_bm25 * 1 / (k + rank_bm25)) + (w_vector * 1 / (k + rank_vector))
4. Dual Cloud and Local Model Backends
- Cloud Mode: Google Gemini 2.5 Flash Lite for fast responses and structured summaries.
- Local Fallback: Quantized
EXAONE-3.5-2.4Bon CPU via Llama.cpp with custom prompt adaptation syntax for small model parameter spaces.
Tech Stack
- Backend & ML: FastAPI, Python, Scikit-learn, TF-IDF, Sentence-Transformers
- Vector Storage: ChromaDB
- Inference Backends: Google Gemini 2.5 Flash Lite, EXAONE-3.5-2.4B (Llama.cpp / GGUF)
- Observability & Auth: Express.js, JWT, Rate Limiting, Custom Evaluation Harness