Relevance Eval: It takes a question and a context chunk, it then asks if the chunk can answer the question. It evaluates the quality of the retrieval process of RAG.
๐งช Test results of relevance prompt and Claude: 0.34 F1 not usable
Evaluating Relevance in RAG: Test Results and Findings | Arize AI Community