Q13
2025 Published

A Framework for Developing and Evaluating Trustworthy RAG-Based AI Agents in the Public Sector

Authors
Andrej Gono
Venue
PEFnet 2025: 29th European Scientific Conference of Doctoral Students

Abstract

RAG-based AI agents are increasingly deployed by public institutions, but the public sector has trustworthiness requirements that general-purpose AI evaluation frameworks do not address: explainability obligations, data residency rules, liability considerations, and the need to answer reliably from authoritative sources only.

This paper proposes a framework that structures the development and evaluation of RAG agents specifically for public sector use. The framework covers knowledge base construction from official sources, retrieval quality metrics relevant to citizen-facing tasks, guardrails against hallucination, and evaluation protocols that reflect real-world deployment conditions rather than benchmark datasets.

The work draws on practical experience from AI assistant deployments across Czech and Slovak municipalities and universities, where questions about social services, permits and university procedures require verifiable, source-grounded answers.