<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Evaluation on EnRedAndo Me - Carlos Prados</title><link>https://carlos.enredando.me/tags/evaluation/</link><description>Recent content in Evaluation on EnRedAndo Me - Carlos Prados</description><generator>Hugo -- gohugo.io</generator><language>en</language><managingEditor>mail@carlosprados.com (Carlos Prados)</managingEditor><webMaster>mail@carlosprados.com (Carlos Prados)</webMaster><copyright>© 2026 Carlos Prados</copyright><lastBuildDate>Tue, 15 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://carlos.enredando.me/tags/evaluation/index.xml" rel="self" type="application/rss+xml"/><item><title>Mastering Agentic AI: The Evaluation and Monitoring Pattern</title><link>https://carlos.enredando.me/posts/agentic-ai-evaluation/</link><pubDate>Tue, 15 Sep 2026 00:00:00 +0000</pubDate><author>mail@carlosprados.com (Carlos Prados)</author><guid>https://carlos.enredando.me/posts/agentic-ai-evaluation/</guid><description>&lt;p&gt;In my &lt;a href="https://carlos.enredando.me/posts/agentic-ai-guardrails/" &gt;previous post&lt;/a&gt;, we looked at Guardrails / Safety Patterns — the rails that keep an agent from doing something stupid or dangerous in the moment. Guardrails answer &amp;ldquo;is this single action safe right now?&amp;rdquo; But they don&amp;rsquo;t answer a different, harder question: &lt;em&gt;is this agent actually any good?&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;You can ship an agent that never violates a guardrail and still be slow, expensive, and wrong half the time. And the brutal part is that with a non-deterministic system, &amp;ldquo;wrong half the time&amp;rdquo; is invisible until you measure it. There&amp;rsquo;s no stack trace for a mediocre answer.&lt;/p&gt;</description><media:content xmlns:media="http://search.yahoo.com/mrss/" url="https://carlos.enredando.me/posts/agentic-ai-evaluation/featured.jpg"/></item></channel></rss>