1. Why label every shell call

A coding agent is a language model in a loop. It reads a task, calls a tool, reads the result and calls again until it reports the task done. Claude Code, Codex and OpenCode work this way, and their most consequential tool is the shell. One Bash call can read a file, run a test suite, rewrite a service, push a branch or delete a volume. The agent picks which, and once a user switches off per call approval to get through a long task, nobody reviews the choice.

In July 2025 a Replit agent reportedly deleted a production database during a code freeze, as documented by the AI Incident Database. In April 2026 Railway described another agent deleting a production volume through its API. Railway later recovered the data. By July the concern extended beyond destructive operations. Hugging Face reported that an agent running OpenAI models escaped an evaluation sandbox and accessed five datasets on Hugging Face infrastructure.

Shell calls provide one view of that activity, not a complete record of everything an agent does. I wanted to classify each call by purpose at low cost, so teams could derive categories from their own traces and investigate unusual shell activity. This analysis tests purpose classification rather than incident detection or prevention. I also tested whether purpose labels could guide model selection.

Each call leaves a record like the one below. The transcript holds the command and a short description the model wrote for it.

{
  "type": "tool_use",
  "id": "toolu_example",
  "name": "Bash",
  "input": {
    "command": "uv run pytest tests/test_parser.py -q",
    "description": "Run the parser unit tests"
  }
}

This block is illustrative, not a corpus record. The model writes both command and description before anything runs. The harness executes the command and returns its output under the same tool_use_id, so every call joins to its result. One call can hold several shell commands, so call counts are not counts of independent tasks.

The apparatus

The data comes from my daily setup. Grove runs Claude Code, OpenCode, Codex and other agents in isolated workspaces and ships every tool call to Langfuse, an open source tracing store for LLM applications. Each call lands as an observation carrying its command, result, model name and token counts. A LiteLLM gateway sits between Claude Code and the providers, so one harness runs Anthropic and OpenAI models interchangeably and the trace records which one answered.

Langfuse also runs evaluators. An evaluator is a second model that reads one observation and answers a fixed question about it. Here the question is what is this shell call for, with answers drawn from a list of purposes such as run_tests or explore_code. Its answers are stored as scores beside the observation. For this analysis I exported observations, scores and full command text, then joined them on the tool call identifier.

What I did with it

I collected 72,264 shell calls across 339 Claude Code sessions from my own development work. The evaluator had labelled 16.3% of them with one of 18 purposes. To label the rest I embedded each command, which turns text into a vector that places similar commands near each other, and trained a classifier on the labelled calls. Command text recovers the evaluator’s labels far better than the leading program alone. python3 and git by themselves are poor guides to intent.

Clustering those vectors without labels exposed structure the 18 purposes merge. Unit tests, browser tests and mutation checks sit apart. Reading a file at an old commit sits apart from searching the current tree. Reading the clusters produced the lasting output, a taxonomy of Bash call purposes with 47 labels, each traced to the clusters behind it. It is written as one multiple choice question for a Langfuse evaluator, so a team can label every Bash call its agents make for the price of one small model call each.

I also compared four models on the work they were given, the tokens around their calls and how often those calls errored. Claude Opus 5, Opus 5.5, Claude Fable 5.1 and GPT-6 Astra are the models I use most. Tasks were not randomly assigned, so differences between models can reflect their assignments rather than their capabilities. Picking a model by purpose does not beat always using the best fixed model, and the logs cannot say what a different model would have done with the same task.

The observed work at a glance

Every call below carries a purpose, from the evaluator for 16.3% of calls and from the classifier for the rest. No person has checked the labels. The six broad groups summarize those assignments without dropping calls that fall outside dense clusters.

Purpose distributions for the full cohort, Claude Fable 5.1, and GPT-6 Astra
Fig. 1. Broad purpose groups for the cohort and two comparison models. Download the counts or the category key.

2. Build a comparable corpus

The recorded calls run from September 2 through October 3, 2026. Holding the harness to Claude Code keeps the tool surface constant while models and tasks vary.

ObservationSelected cohort
Claude Code shell calls72,264
Recorded session IDs339
Normalized model identifiers11
Calls with an evaluator label11,761
Calls assigned a predicted label60,503
Sessions containing evaluator labels51

Telemetry reaches Langfuse over the OpenTelemetry Protocol. Full commands come from harness transcripts and join to evaluator scores on the tool call identifier. Model aliases come from the gateway configuration dated 2026-10-03 and are not independently verified provider identities.

Weekly shell call exposure by recorded or resolved model version
Fig. 2. Model exposure changes across the observation window. Comparisons also reflect when each model was used.

Ollama serves Qwen3-Embedding-0.6B with Q8_0 weights, producing vectors of 1,024 dimensions. Command text is capped at 2,000 characters, which truncates 2,136 calls, or 3.0%, and can drop the decisive part of a long script. The same vectors feed the classifier in section 3 and the clustering in section 4.

3. Can command text recover purpose?

A leading program does not identify a command’s purpose. A python3 call might edit a file, query data or run a test. The classifier measures whether the full command text separates such uses.

Logistic regression is trained on the 11,761 labelled vectors with five folds from GroupKFold, so a session’s calls never appear in both training and validation. Baselines use cosine nearest neighbors or the most common training label for each leading program.

ClassifierMacro-F1Micro-F1
Embedding + logistic regression0.620.73
Embedding + 15 nearest neighbors0.520.72
Leading-program majority0.280.56

Macro F1 weights the 18 categories equally and micro F1 pools every decision. The embedding classifier more than doubles macro F1 over the leading program baseline.

Purpose predictions on validation sessions compared with evaluator labels
Fig. 3. Each row shows predictions as shares of its reference category. Exploration and git history are easier to recover than broad categories such as other.

Would the model’s own description help?

The example above pairs the test command with Run the parser unit tests. On the 11,723 labelled calls that carry a description, I compared classifiers trained on the command, the description and both.

Classifier inputMacro-F1Micro-F1
Command0.620.73
Description0.430.53
Command + description0.660.76

The description alone is a worse guide than the command. Together they add four points of macro F1, enough to keep the description as a secondary signal.

4. What structure does clustering add?

Clustering looks inside a purpose category. rg "parse" src/ searches code and sed -n '40,80p' src/parser.py reads a known range. Both can receive the same label while doing different things.

UMAP reduces the vectors to 10 dimensions with cosine distance, 30 neighbors and min_dist=0. HDBSCAN with min_cluster_size=120 and min_samples=20 finds 126 clusters and leaves 33.7% of calls outside any dense group. Across four density settings that share stays between 31.4% and 35.6%.

On the 8,885 labelled calls inside clusters, agreement with evaluator purposes is modest, with an adjusted Rand index of 0.23 and normalized mutual information of 0.35. The clusters do not reproduce the purpose vocabulary, and UMAP can distort density, so they are an inspection aid rather than a validated taxonomy.

Hierarchy of the larger command clusters
Fig. 4. Average linkage across the 34 clusters with at least 400 calls, which together cover 38% of the cohort. Leaves name the dominant purpose and leading program. A separate density cluster view shows membership and unclustered points.

Git history makes up 16.9% of unclustered calls against 10.7% of clustered calls. Dropping the unclustered third would shift the purpose distribution, so every model comparison keeps every call.

From clusters to a finer taxonomy

The clusters split what one evaluator label merges. Unit tests, browser suites and deliberate mutation checks form separate clusters inside one testing label. Reading code at a past revision separates from searching the current tree. Monitoring CI separates from editing issues and reviews.

I assigned each of the 126 clusters one label by reading its exemplar commands, leading programs and purpose mix. That yields 47 labels in ten families, plus other and unclear as residuals. Each unclustered call takes the label of its nearest cluster centroid, which gives an approximate share for every label.

Under that approximation 29 labels hold at least 1% of calls, against 16 of the 18 evaluator labels, and the largest label falls from 32.7% to 21.2% of calls. No model or person has yet classified calls under this taxonomy. These shares describe the clustering, not measured accuracy.

5. What the labels say about four models

Three confounds apply to everything in this section. Evaluator coverage ranges from 3.2% for Opus 5 to 37.9% for Opus 5.5, with 10.7% for Fable and 4.7% for Astra, and classifier accuracy is not measured per model. Attribution names the model that issued a call, which is not always the model driving the session. And 113 sessions contain more than one model, so a session counts toward several totals and those totals are not independent.

Sessions

Session exposure measures how widely a model appears. A model participates in a session if at least one shell call there is attributed to it. The denominator is all 339 sessions with shell calls.

ModelCallsSessions (% of 339)Calls/model-session, median (IQR)
Opus 521,642199 (58.7%)55 (14.5 to 142)
Opus 5.514,927101 (29.8%)89 (33 to 203)
Fable 5.15,61853 (15.6%)44 (16 to 141)
GPT-6 Astra7,43757 (16.8%)86 (34 to 171)

The two Opus versions never overlap. Opus 5 runs from September 2 to September 22 and Opus 5.5 from September 22 to October 3. Delegated subagent work is 31.2% of Opus 5 calls and 4.1% of Opus 5.5 calls, so their comparison mixes period with role. Between Fable and Astra, Fable’s largest single session supplies 16.1% of its calls against 7.6% for Astra. Normalizing by volume alone does not remove that concentration.

Purpose shares of calls for Opus 5, Opus 5.5, Fable 5.1, and GPT-6 Astra
Fig. 5. Each column sums to 100% of a model's calls across the 18 purpose categories. The pies repeat the broad groups from Figure 1.

The heatmap weights by calls, so long sessions dominate. Weighting each model and session pair equally instead changes the exploration and testing shares as follows.

ModelExplore code: call → session weightedRun tests: call → session weighted
Opus 534.8% → 30.0%12.0% → 8.0%
Opus 5.529.1% → 31.1%9.8% → 7.6%
Fable 5.130.7% → 28.1%7.7% → 5.3%
GPT-6 Astra9.9% → 10.7%13.1% → 10.2%

Opus 5 and Opus 5.5 swap order on exploration under session weighting. Astra keeps the lowest exploration share and the highest test share either way. Some comparisons survive the change of weighting and some do not.

Tokens

Each call inherits the output tokens reported on the transcript message that issued it. Several calls can share one message, and providers count differently, so these associated output tokens measure recorded output around shell activity rather than a deduplicated bill. Only calls with a positive count enter the comparison.

Associated output token distributions for the four comparison models
Fig. 6. Interquartile ranges of output tokens associated with each model's calls.

Median associated output is 335 tokens for Opus 5, 483 for Opus 5.5, 456 for Fable and 144 for Astra. Command payloads and batching count toward these figures alongside any explanatory text.

Purpose shares of calls and associated output tokens across the cohort
Fig. 7. Call share against associated output share by purpose. Tokens are counted once per call linked to a transcript message.

Across the 64,727 eligible calls, exploration is 31.6% of calls and 42.8% of associated tokens. Editing is 5.2% and 9.3%. These are the places output accumulates. Turning them into savings would need deduplicated messages, input and cache charges and completed task outcomes.

Errors

Claude Code’s is_error flag marks tool execution errors, including nonzero exits. A test that finds a bug raises it. So does a grep with no match. The flag is a weak proxy for useful work, and rates below use only calls that carry it.

Tool error rates by purpose for the four comparison models
Fig. 8. Error rates by purpose and model in cells with at least 80 calls. Task difficulty and the timing of checks remain uncontrolled.

Aggregate error rates are 3.3% for Opus 5, 2.6% for Opus 5.5, 3.3% for Fable and 6.8% for Astra. Yet on service operations Fable errors more often than Astra. The ordering depends on the action.

Optimizing this flag alone would reward avoiding tests over finding and fixing failures. A better outcome would record whether a failing check was resolved and the task accepted.

6. Routing does not follow

A purpose lookup does not beat the best fixed model. This test asks whether purpose labels plus observed error rates can pick a model, scoring all four candidates on the same evaluation rows.

RouteLLM learns selection from preference data and RouterBench compares policies against fixed model baselines. Both see every candidate’s answer to the same request. These logs see only the model that ran.

Five validation folds keep sessions separate. In each training fold one lookup picks, per purpose, the model with the highest fraction of error free calls, and another picks the model with the lowest mean associated output tokens. Both are scored against always choosing one model. Purposes need at least 60 training and 20 validation calls per candidate, which retains 41,661 calls from 316 sessions. Each selection is scored with the chosen model’s validation mean for that purpose.

PolicyFraction without errorsMean associated output tokens
Always Opus 50.9704585
Always Opus 5.50.97381,131
Always Fable 5.10.96731,007
Always GPT-6 Astra0.9343222
Highest training fraction without errors0.97031,003
Lowest training mean output tokens0.9343222

The error lookup scores below fixed Opus 5.5 and costs nearly twice the tokens of fixed Opus 5 for the same error proxy. The token lookup picks Astra every time. Neither lookup beats the fixed baselines on both measures.

The deeper problem is what the logs cannot hold. The lookup reads commands a chosen model has already written, and a router deciding beforehand would not have them. The logs record one model per call, so they cannot show whether another model would have finished the task more cheaply. The validation split also protects only the policy fit, since the purpose classifier was trained on labels from across the cohort.

7. Adopting the taxonomy

The Bash call purpose taxonomy is meant for an organization where many engineers run coding agents and every Bash call reaches an observability backend. A decision model such as Jev answers one typed question per observation and returns a probability for each option, which makes labelling every call cheaper than sampling.

The Langfuse evaluator definition poses the taxonomy as one choice question with 49 options, within Langfuse’s limit of 255. Each label carries its family, a description written for the decision model, and the clusters, leading programs and evaluator labels behind it. The example command above would map to verify_unit_tests.

With a label on every call, monitoring becomes counting. A team learns the mix of purposes its agents normally produce and alerts when a session drifts, for instance a documentation task that starts issuing service_deploy calls, or a burst of config_secrets reads from an agent asked to fix a test. The label does not say whether an action is allowed. Paired with the workspace, the ticket and the permission mode, it says whether the action fits the task, which is the first question a reviewer asks.

Adoption still requires calibration. A sample labelled by people would measure accuracy per option and show which clusters need splitting or merging. I have not tested anomaly detection on these labels, so the monitoring use is a design the taxonomy supports rather than a result. Establishing routing value would separately require assignments made from information available before execution, and comparable task outcomes.

Supplementary figures

These repeat evidence shown above from another angle. They are here for inspection rather than argument.

Shell call counts by model version in the Claude Code cohort
Fig. S1. Corpus composition by model version. The counts match the sessions table in section 5. Volume is not performance.
Label recovery from command and description inputs with their embedding similarity for Fable and Astra
Fig. S2. Panel A repeats the section 3 input comparison. Panel B measures similarity between command and description embeddings. Similarity does not establish that the description faithfully represents the command.
Minimum cluster sizeMinimum samplesClustersUnclustered
1202012633.7%
120513031.4%
602023035.6%
60526633.0%

Table S1. HDBSCAN sensitivity on the same UMAP coordinates. Other embeddings and UMAP settings were not tested.

Command embeddings colored by broad purpose including calls without density cluster membership
Fig. S3. The cover image at full size. Colors use the purpose assignments summarized in Figure 1. Coordinate extremes are cropped.
Median associated output tokens by purpose for the four comparison models
Fig. S4. Median associated output tokens by purpose and model. Editing can carry replacement file content in the command payload, so its output tokens include the material being written.
Observed tool error rates for the four comparison models
Fig. S5. Aggregate tool error rates with Wilson intervals under an independence assumption, which ignores dependence within sessions.

Data

Raw traces include private workspace content, so I am not releasing them. The CSVs above reproduce the summaries, not the analysis.

comments powered by Disqus