<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Testing on Tony老师的博客</title><link>https://blog.tanteng.space/tags/testing/</link><description>Recent content in Testing on Tony老师的博客</description><generator>Hugo</generator><language>zh</language><lastBuildDate>Tue, 10 Jun 2025 10:00:00 +0800</lastBuildDate><atom:link href="https://blog.tanteng.space/tags/testing/index.xml" rel="self" type="application/rss+xml"/><item><title>AI Agent 评估体系：LLM-as-Judge 与回归测试集设计</title><link>https://blog.tanteng.space/2025/06/ai-agent-evaluation/</link><pubDate>Tue, 10 Jun 2025 10:00:00 +0800</pubDate><guid>https://blog.tanteng.space/2025/06/ai-agent-evaluation/</guid><description>&lt;blockquote&gt;
&lt;p&gt;上线一个 AI Agent 不写评估，就像把 SQL 代码 push 上生产但从不跑测试。LLM 的输出有随机性，prompt 微调一句可能让质量断崖式下跌——没有自动化的回归测试，&lt;strong&gt;你永远不知道哪次 commit 把产品搞砸了&lt;/strong&gt;。&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;这篇文章不讲&amp;quot;为什么要做评估&amp;quot;（这个地球人都知道），讲&lt;strong&gt;怎么搭一套真正能用的 Agent 评估体系&lt;/strong&gt;：&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;三种评估方法（确定性 / LLM-as-Judge / 人工）的取舍&lt;/li&gt;
&lt;li&gt;Golden Dataset 怎么造、造多少、怎么维护&lt;/li&gt;
&lt;li&gt;评估指标设计：从 RAG 召回到 Agent 工具调用&lt;/li&gt;
&lt;li&gt;CI/CD 集成：每次改 prompt 自动跑&lt;/li&gt;
&lt;li&gt;实战代码：基于 DeepEval 的回归 pipeline&lt;/li&gt;
&lt;/ul&gt;</description></item></channel></rss>