1. Why label every shell call
A coding agent is a language model in a loop. It reads a task, calls a tool, reads the result and calls again until it reports the task done. Claude Code, Codex and OpenCode work this way, and their most consequential tool is the shell. One Bash call can read a file, run a test suite, rewrite a service, push a branch or delete a volume. The agent picks which, and once a user switches off per call approval to get through a long task, nobody reviews the choice.
In July 2025 a Replit agent reportedly deleted a production database during a code freeze, as documented by the AI Incident Database. In April 2026 Railway described another agent deleting a production volume through its API. Railway later recovered the data. By July the concern extended beyond destructive operations. Hugging Face reported that an agent running OpenAI models escaped an evaluation sandbox and accessed five datasets on Hugging Face infrastructure.
Shell calls provide one view of that activity, not a complete record of everything an agent does. I wanted to classify each call by purpose at low cost, so teams could derive categories from their own traces and investigate unusual shell activity. This analysis tests purpose classification rather than incident detection or prevention. I also tested whether purpose labels could guide model selection.
Each call leaves a record like the one below. The transcript holds the command and a short description the model wrote for it.
{
"type": "tool_use",
"id": "toolu_example",
"name": "Bash",
"input": {
"command": "uv run pytest tests/test_parser.py -q",
"description": "Run the parser unit tests"
}
}
This block is illustrative, not a corpus record. The model writes both command and description before anything runs. The harness executes the command and returns its output under the same tool_use_id, so every call joins to its result. One call can hold several shell commands, so call counts are not counts of independent tasks.
The apparatus
The data comes from my daily setup. Grove runs Claude Code, OpenCode, Codex and other agents in isolated workspaces and ships every tool call to Langfuse, an open source tracing store for LLM applications. Each call lands as an observation carrying its command, result, model name and token counts. A LiteLLM gateway sits between Claude Code and the providers, so one harness runs Anthropic and OpenAI models interchangeably and the trace records which one answered.
Langfuse also runs evaluators. An evaluator is a second model that reads one observation and answers a fixed question about it. Here the question is what is this shell call for, with answers drawn from a list of purposes such as run_tests or explore_code. Its answers are stored as scores beside the observation. For this analysis I exported observations, scores and full command text, then joined them on the tool call identifier.
What I did with it
I collected 72,264 shell calls across 339 Claude Code sessions from my own development work. The evaluator had labelled 16.3% of them with one of 18 purposes. To label the rest I embedded each command, which turns text into a vector that places similar commands near each other, and trained a classifier on the labelled calls. Command text recovers the evaluator’s labels far better than the leading program alone. python3 and git by themselves are poor guides to intent.
Clustering those vectors without labels exposed structure the 18 purposes merge. Unit tests, browser tests and mutation checks sit apart. Reading a file at an old commit sits apart from searching the current tree. Reading the clusters produced the lasting output, a taxonomy of Bash call purposes with 47 labels, each traced to the clusters behind it. It is written as one multiple choice question for a Langfuse evaluator, so a team can label every Bash call its agents make for the price of one small model call each.
I also compared four models on the work they were given, the tokens around their calls and how often those calls errored. Claude Opus 5, Opus 5.5, Claude Fable 5.1 and GPT-6 Astra are the models I use most. Tasks were not randomly assigned, so differences between models can reflect their assignments rather than their capabilities. Picking a model by purpose does not beat always using the best fixed model, and the logs cannot say what a different model would have done with the same task.
The observed work at a glance
Every call below carries a purpose, from the evaluator for 16.3% of calls and from the classifier for the rest. No person has checked the labels. The six broad groups summarize those assignments without dropping calls that fall outside dense clusters.

2. Build a comparable corpus
The recorded calls run from September 2 through October 3, 2026. Holding the harness to Claude Code keeps the tool surface constant while models and tasks vary.
| Observation | Selected cohort |
|---|---|
| Claude Code shell calls | 72,264 |
| Recorded session IDs | 339 |
| Normalized model identifiers | 11 |
| Calls with an evaluator label | 11,761 |
| Calls assigned a predicted label | 60,503 |
| Sessions containing evaluator labels | 51 |
Telemetry reaches Langfuse over the OpenTelemetry Protocol. Full commands come from harness transcripts and join to evaluator scores on the tool call identifier. Model aliases come from the gateway configuration dated 2026-10-03 and are not independently verified provider identities.

Ollama serves Qwen3-Embedding-0.6B with Q8_0 weights, producing vectors of 1,024 dimensions. Command text is capped at 2,000 characters, which truncates 2,136 calls, or 3.0%, and can drop the decisive part of a long script. The same vectors feed the classifier in section 3 and the clustering in section 4.
3. Can command text recover purpose?
A leading program does not identify a command’s purpose. A python3 call might edit a file, query data or run a test. The classifier measures whether the full command text separates such uses.
Logistic regression is trained on the 11,761 labelled vectors with five folds from GroupKFold, so a session’s calls never appear in both training and validation. Baselines use cosine nearest neighbors or the most common training label for each leading program.
| Classifier | Macro-F1 | Micro-F1 |
|---|---|---|
| Embedding + logistic regression | 0.62 | 0.73 |
| Embedding + 15 nearest neighbors | 0.52 | 0.72 |
| Leading-program majority | 0.28 | 0.56 |
Macro F1 weights the 18 categories equally and micro F1 pools every decision. The embedding classifier more than doubles macro F1 over the leading program baseline.

Would the model’s own description help?
The example above pairs the test command with Run the parser unit tests. On the 11,723 labelled calls that carry a description, I compared classifiers trained on the command, the description and both.
| Classifier input | Macro-F1 | Micro-F1 |
|---|---|---|
| Command | 0.62 | 0.73 |
| Description | 0.43 | 0.53 |
| Command + description | 0.66 | 0.76 |
The description alone is a worse guide than the command. Together they add four points of macro F1, enough to keep the description as a secondary signal.
4. What structure does clustering add?
Clustering looks inside a purpose category. rg "parse" src/ searches code and sed -n '40,80p' src/parser.py reads a known range. Both can receive the same label while doing different things.
UMAP reduces the vectors to 10 dimensions with cosine distance, 30 neighbors and min_dist=0. HDBSCAN with min_cluster_size=120 and min_samples=20 finds 126 clusters and leaves 33.7% of calls outside any dense group. Across four density settings that share stays between 31.4% and 35.6%.
On the 8,885 labelled calls inside clusters, agreement with evaluator purposes is modest, with an adjusted Rand index of 0.23 and normalized mutual information of 0.35. The clusters do not reproduce the purpose vocabulary, and UMAP can distort density, so they are an inspection aid rather than a validated taxonomy.

Git history makes up 16.9% of unclustered calls against 10.7% of clustered calls. Dropping the unclustered third would shift the purpose distribution, so every model comparison keeps every call.
From clusters to a finer taxonomy
The clusters split what one evaluator label merges. Unit tests, browser suites and deliberate mutation checks form separate clusters inside one testing label. Reading code at a past revision separates from searching the current tree. Monitoring CI separates from editing issues and reviews.
I assigned each of the 126 clusters one label by reading its exemplar commands, leading programs and purpose mix. That yields 47 labels in ten families, plus other and unclear as residuals. Each unclustered call takes the label of its nearest cluster centroid, which gives an approximate share for every label.
Under that approximation 29 labels hold at least 1% of calls, against 16 of the 18 evaluator labels, and the largest label falls from 32.7% to 21.2% of calls. No model or person has yet classified calls under this taxonomy. These shares describe the clustering, not measured accuracy.
5. What the labels say about four models
Three confounds apply to everything in this section. Evaluator coverage ranges from 3.2% for Opus 5 to 37.9% for Opus 5.5, with 10.7% for Fable and 4.7% for Astra, and classifier accuracy is not measured per model. Attribution names the model that issued a call, which is not always the model driving the session. And 113 sessions contain more than one model, so a session counts toward several totals and those totals are not independent.
Sessions
Session exposure measures how widely a model appears. A model participates in a session if at least one shell call there is attributed to it. The denominator is all 339 sessions with shell calls.
| Model | Calls | Sessions (% of 339) | Calls/model-session, median (IQR) |
|---|---|---|---|
| Opus 5 | 21,642 | 199 (58.7%) | 55 (14.5 to 142) |
| Opus 5.5 | 14,927 | 101 (29.8%) | 89 (33 to 203) |
| Fable 5.1 | 5,618 | 53 (15.6%) | 44 (16 to 141) |
| GPT-6 Astra | 7,437 | 57 (16.8%) | 86 (34 to 171) |
The two Opus versions never overlap. Opus 5 runs from September 2 to September 22 and Opus 5.5 from September 22 to October 3. Delegated subagent work is 31.2% of Opus 5 calls and 4.1% of Opus 5.5 calls, so their comparison mixes period with role. Between Fable and Astra, Fable’s largest single session supplies 16.1% of its calls against 7.6% for Astra. Normalizing by volume alone does not remove that concentration.

The heatmap weights by calls, so long sessions dominate. Weighting each model and session pair equally instead changes the exploration and testing shares as follows.
| Model | Explore code: call → session weighted | Run tests: call → session weighted |
|---|---|---|
| Opus 5 | 34.8% → 30.0% | 12.0% → 8.0% |
| Opus 5.5 | 29.1% → 31.1% | 9.8% → 7.6% |
| Fable 5.1 | 30.7% → 28.1% | 7.7% → 5.3% |
| GPT-6 Astra | 9.9% → 10.7% | 13.1% → 10.2% |
Opus 5 and Opus 5.5 swap order on exploration under session weighting. Astra keeps the lowest exploration share and the highest test share either way. Some comparisons survive the change of weighting and some do not.
Tokens
Each call inherits the output tokens reported on the transcript message that issued it. Several calls can share one message, and providers count differently, so these associated output tokens measure recorded output around shell activity rather than a deduplicated bill. Only calls with a positive count enter the comparison.

Median associated output is 335 tokens for Opus 5, 483 for Opus 5.5, 456 for Fable and 144 for Astra. Command payloads and batching count toward these figures alongside any explanatory text.

Across the 64,727 eligible calls, exploration is 31.6% of calls and 42.8% of associated tokens. Editing is 5.2% and 9.3%. These are the places output accumulates. Turning them into savings would need deduplicated messages, input and cache charges and completed task outcomes.
Errors
Claude Code’s is_error flag marks tool execution errors, including nonzero exits. A test that finds a bug raises it. So does a grep with no match. The flag is a weak proxy for useful work, and rates below use only calls that carry it.

Aggregate error rates are 3.3% for Opus 5, 2.6% for Opus 5.5, 3.3% for Fable and 6.8% for Astra. Yet on service operations Fable errors more often than Astra. The ordering depends on the action.
Optimizing this flag alone would reward avoiding tests over finding and fixing failures. A better outcome would record whether a failing check was resolved and the task accepted.
6. Routing does not follow
A purpose lookup does not beat the best fixed model. This test asks whether purpose labels plus observed error rates can pick a model, scoring all four candidates on the same evaluation rows.
RouteLLM learns selection from preference data and RouterBench compares policies against fixed model baselines. Both see every candidate’s answer to the same request. These logs see only the model that ran.
Five validation folds keep sessions separate. In each training fold one lookup picks, per purpose, the model with the highest fraction of error free calls, and another picks the model with the lowest mean associated output tokens. Both are scored against always choosing one model. Purposes need at least 60 training and 20 validation calls per candidate, which retains 41,661 calls from 316 sessions. Each selection is scored with the chosen model’s validation mean for that purpose.
| Policy | Fraction without errors | Mean associated output tokens |
|---|---|---|
| Always Opus 5 | 0.9704 | 585 |
| Always Opus 5.5 | 0.9738 | 1,131 |
| Always Fable 5.1 | 0.9673 | 1,007 |
| Always GPT-6 Astra | 0.9343 | 222 |
| Highest training fraction without errors | 0.9703 | 1,003 |
| Lowest training mean output tokens | 0.9343 | 222 |
The error lookup scores below fixed Opus 5.5 and costs nearly twice the tokens of fixed Opus 5 for the same error proxy. The token lookup picks Astra every time. Neither lookup beats the fixed baselines on both measures.
The deeper problem is what the logs cannot hold. The lookup reads commands a chosen model has already written, and a router deciding beforehand would not have them. The logs record one model per call, so they cannot show whether another model would have finished the task more cheaply. The validation split also protects only the policy fit, since the purpose classifier was trained on labels from across the cohort.
7. Adopting the taxonomy
The Bash call purpose taxonomy is meant for an organization where many engineers run coding agents and every Bash call reaches an observability backend. A decision model such as Jev answers one typed question per observation and returns a probability for each option, which makes labelling every call cheaper than sampling.
The Langfuse evaluator definition poses the taxonomy as one choice question with 49 options, within Langfuse’s limit of 255. Each label carries its family, a description written for the decision model, and the clusters, leading programs and evaluator labels behind it. The example command above would map to verify_unit_tests.
With a label on every call, monitoring becomes counting. A team learns the mix of purposes its agents normally produce and alerts when a session drifts, for instance a documentation task that starts issuing service_deploy calls, or a burst of config_secrets reads from an agent asked to fix a test. The label does not say whether an action is allowed. Paired with the workspace, the ticket and the permission mode, it says whether the action fits the task, which is the first question a reviewer asks.
Adoption still requires calibration. A sample labelled by people would measure accuracy per option and show which clusters need splitting or merging. I have not tested anomaly detection on these labels, so the monitoring use is a design the taxonomy supports rather than a result. Establishing routing value would separately require assignments made from information available before execution, and comparable task outcomes.
Supplementary figures
These repeat evidence shown above from another angle. They are here for inspection rather than argument.


| Minimum cluster size | Minimum samples | Clusters | Unclustered |
|---|---|---|---|
| 120 | 20 | 126 | 33.7% |
| 120 | 5 | 130 | 31.4% |
| 60 | 20 | 230 | 35.6% |
| 60 | 5 | 266 | 33.0% |
Table S1. HDBSCAN sensitivity on the same UMAP coordinates. Other embeddings and UMAP settings were not tested.



Data
Raw traces include private workspace content, so I am not releasing them. The CSVs above reproduce the summaries, not the analysis.
