<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Evaluation on Blog by Krishna Alagiri.</title><link>https://thekrishna.in/blogs/tags/evaluation/</link><description>Recent content in Evaluation on Blog by Krishna Alagiri.</description><generator>Hugo</generator><language>en-us</language><copyright>Written with ❤️️ by Krishna.</copyright><lastBuildDate>Sat, 03 Oct 2026 12:00:00 -0700</lastBuildDate><atom:link href="https://thekrishna.in/blogs/tags/evaluation/index.xml" rel="self" type="application/rss+xml"/><item><title>Embedding and Classifying Coding Agents’ Shell Commands for Anomaly Detection</title><link>https://thekrishna.in/blogs/blog/shell-atlas/</link><pubDate>Sat, 03 Oct 2026 12:00:00 -0700</pubDate><guid>https://thekrishna.in/blogs/blog/shell-atlas/</guid><description>&lt;h2 id="1-why-label-every-shell-call">1. Why label every shell call&lt;/h2>
&lt;p>A coding agent is a language model in a loop. It reads a task, calls a tool, reads the result and calls again until it reports the task done. Claude Code, Codex and OpenCode work this way, and their most consequential tool is the shell. One Bash call can read a file, run a test suite, rewrite a service, push a branch or delete a volume. The agent picks which, and once a user switches off per call approval to get through a long task, nobody reviews the choice.&lt;/p></description></item></channel></rss>