The challenge
A regional bank's internal RAG assistant needed a rigorous way to measure and improve answer quality, retrieval relevance and agent performance. Quality was too subjective to govern, with responses stuck in the 35 to 40 second range.
What we did
Delivered an end-to-end evaluation framework for RAG and agentic AI: MLflow Evaluate and Traces, Mosaic AI Agent Evaluation, Review App workflows, Unity Catalog Vector Search benchmarking, hybrid retrieval and reranking assessments, offline and online evaluation, monitoring, and reusable implementation assets including LangGraph-based agent patterns.
The outcome
Cut response latency from 35 to 40 seconds to 15 to 18 seconds, roughly 55 percent faster at the midpoint, while making accuracy, retrieval relevance and agent performance continuously measurable.
Cut enterprise AI-assistant latency from 35–40 seconds to 15–18 seconds, roughly 55 percent faster, while enabling measurable improvement of quality and retrieval performance.
Evaluation pipelines scored retrieval relevance, parsing quality and end-to-end answer fidelity. Online and offline loops and tracing turned subjective assistant quality into governed, iterative improvement.