Overview
The assistant project is an open stack for putting long running agents to work across the systems you already run. It is self hosted and runs on any model . Every agent is scoped to the person it runs for.
Coding agents made engineers far more productive. Everyone else files a ticket and waits. A GTM engineer wants scheduled lead enrichment. A support lead wants one view over three tools. A new hire wants to ask a repository how it works. All three need an agent next to their own data, inside their own permissions. It has to live on the surface they already work in. I built one harness and pointed it at each job. Tasks, the wiki, search and apps are that harness over different data. The same session reaches a browser, a terminal, an Android phone, a team chat, an inbox and a Home Assistant microphone. An improvement to the loop reaches all of them.
Harness architecture and governance
Every agent at every depth runs one async agentic loop . The model returns tool calls through native function calling. The harness runs the permitted ones. The loop ends when a reply has no calls. Its tools read and edit files, run a shell and put a question back to the person. A language server reruns diagnostics after every edit. A long session does not overflow the context window. Compaction summarises older turns at 80 percent of the window and keeps the recent ones verbatim. The file being edited and the last error output stay in the prompt verbatim.
Work that can run at once goes to subagents in an orchestrator and worker pattern. A hypervisor owns their admission and caps delegation at depth five. A spawn past the ceiling is rejected, not queued, so a saturated fleet stays visible. Each child returns a structured result and the root folds them into one answer. A session can take its own worktree , so two never edit the same checkout. A Web IDE opens beside any session when you want the files yourself.

Giving people an agent is a permissions decision. The identity provider says who someone is and role based access control says what they may run, on a least privilege basis. Every write and shell call clears a rule or a hook . The shell runs in a kernel enforced Landlock sandbox , the same Linux security module behind Codex CLI’s Linux sandbox. Plan mode is the human in the loop gate. It holds an unfamiliar task at a drafted plan until a person approves it. Above those binary rules sit two semantic guardrails. A policy is a plain English rule checked by an isolated LLM judge call before a greylisted tool runs. It refuses a call on its content rather than its name. A monitor watches the whole agent tree across the session for budget, deadline or drift. It can inject a correction or stop the run. Every model call lands on one Langfuse trace with its token counts, so an audit answers who ran what and what it cost.
The harness loads what a team already has. CLAUDE.md, AGENTS.md, Agent Skills
, Model Context Protocol servers and Claude Code plugins
activate exactly as authored. Triggers
start a session with nobody watching, on a cron schedule, a CI run, a pull request or a webhook. A task that worked once under supervision becomes a job that runs every night. The harness is also an MCP server
, so Claude Code or Cursor can hand it a task.

An answer is not always text. When the result is tabular or a time series, the agent renders a widget inline in the chat. Structured outputs take a question and a JSON schema and return an object that fits it. Your own software gets an agent without building one.
Repository indexing, knowledge graph and question answering
Paste a repository URL and the harness generates its documentation . Tree sitter parses each file into symbols joined by call and import edges. A model generates entity labels for subsystems and concepts, each anchored to the symbols it was derived from. Notes attach to both and retire when their code changes. All three share one graph . Reindexing touches only the pages and nodes a change reached. A README’s concepts become entities even with no code behind them.
The same graph answers questions as graph based retrieval augmented generation. Probe agents enter at the symbols an embedding lookup ranks highest. They follow the edges out to callers and to the note explaining why. The answer cites a path and a line range. A durable fact from an answer is written back as a note. A new engineer reads the source instead of interrupting its author. A platform team gets current docs for every service.

Agentic applications and cross source search
Describe an app in one sentence and a builder agent writes it against real data. It derives the data model and writes a Streamlit front end that runs as WebAssembly in the browser. It also writes the pipelines that feed it. The steady state is plain code on a schedule and spends no tokens. A model returns only when a pipeline breaks. Every submit is a version and rollback is one step. A health panel shows freshness per pipeline and warns loudly. Whoever owns a metric owns its dashboard without a sprint.

Search uses the same graph idea on sources that are not code. Connect a CRM, a database or an MCP server. Its schema becomes a graph of what it can answer . No records enter it. A query is embedded and routed by nearest neighbour search before any data is fetched. Routing itself calls no language model. Probe agents run under least privilege, with only the sources and tools their route lists. A question spanning the CRM, the warehouse and the tickets returns one cited answer. Its confidence score counts how many independent routes returned corroborating records. Every run has a stable link that renders the same synthesis and trace.

