<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>RAG 评估 on Tony老师的博客</title><link>https://blog.tanteng.space/tags/rag-evaluation/</link><description>Recent content in RAG 评估 on Tony老师的博客</description><generator>Hugo</generator><language>zh</language><lastBuildDate>Thu, 20 Feb 2025 10:00:00 +0800</lastBuildDate><atom:link href="https://blog.tanteng.space/tags/rag-evaluation/index.xml" rel="self" type="application/rss+xml"/><item><title>RAG 效果评估体系：从人工标注到 LLM-as-Judge</title><link>https://blog.tanteng.space/2025/02/rag-evaluation-framework/</link><pubDate>Thu, 20 Feb 2025 10:00:00 +0800</pubDate><guid>https://blog.tanteng.space/2025/02/rag-evaluation-framework/</guid><description>&lt;blockquote&gt;
&lt;p&gt;上线一个 RAG 系统不写评估，就像把 SQL 代码 push 上生产但从不跑测试。LLM 输出有随机性，prompt 微调、embedding 换模型、rerank 加权调整都可能让质量断崖式下跌——没有自动化的回归测试，你永远不知道哪次迭代把检索质量搞砸了。&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;这篇文章不讲&amp;quot;为什么要做评估&amp;quot;，讲&lt;strong&gt;怎么搭一套真正能用的 RAG 评估体系&lt;/strong&gt;：&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;四种评估方法（人工 / 自动指标 / LLM-as-Judge / 端到端 A/B）的取舍&lt;/li&gt;
&lt;li&gt;Golden Dataset 怎么造、造多少、怎么维护&lt;/li&gt;
&lt;li&gt;RAGAS 四大核心指标的定义与局限&lt;/li&gt;
&lt;li&gt;LLM-as-Judge 与人类标注的相关系数证据&lt;/li&gt;
&lt;li&gt;实战代码：基于 RAGAS + DeepEval 的离线评估 pipeline&lt;/li&gt;
&lt;/ul&gt;</description></item></channel></rss>