Skip to content

Architecture

helioai/
├── config.py               settings singleton, loaded once from .env
├── logging_config.py       structlog setup (console or JSON)
├── datastore.py            npz + manifest per session; the key to reproducible export
├── workspace.py            per-user, per-session directories
├── export.py               session → standalone .ipynb
├── indexer.py              speasy catalogue → ChromaDB
├── index_snapshot.py       export / fetch / import the prebuilt index (Hugging Face Hub)
├── mcp_server.py           MCP stdio + streamable HTTP
├── runtime/
│   ├── runner.py           Runner(policy).run(history) — the one agent loop
│   ├── policies.py         Policy — what makes a run the lead or a delegated role
│   └── context.py          RunContext — who runs, in which session, writing where
├── core/
│   ├── agent_loop.py       stream_chat — the lead: its policy, the task/skill tools, persistence
│   ├── sub_agents.py       stream_subagent — the roles: a policy each, a whitelist, a report
│   ├── tool_exec.py        tool-call mechanics the runner uses: summaries, artifacts, checks
│   ├── session.py          SQLite history + event journal per (user_id, session_id)
│   ├── skills_loader.py    markdown skills
│   ├── vision.py           stateless figure review side-call
│   ├── skills/             SKILL.md prompt assets
│   └── llm/
│       ├── base.py         Message, ToolCall, ToolDef, LLMClient, call_with_retry
│       ├── openai_compat.py  one client for every OpenAI-wire provider
│       ├── azure_openai.py   thin subclass (deployment routing, developer role)
│       ├── gemini.py         native google-genai client
│       └── factory.py        build_llm_client + the provider table
├── tools/
│   ├── registry.py         ToolRegistry — JSON dispatch to async functions
│   ├── results.py          ToolResult — a call's payload and the text the model reads
│   ├── setup.py            registers all 18 tools at import
│   ├── rag.py              hybrid BM25 + dense retrieval, fused by RRF
│   ├── speasy_tools.py     search, download, data-quality scan
│   ├── catalog_tools.py    AMDA catalogs, event timeseries
│   ├── plasmapy_tools.py   formulary wrappers
│   ├── sandbox.py          bubblewrap-isolated Python execution
│   ├── sandbox_helpers.py  coordinate transforms, boundary models
│   ├── recipes.py          recipe loading, and run_recipe: a recipe run as shipped
│   ├── literature.py       NASA ADS
│   └── mcp_client.py       mounts remote MCP servers into the registry
├── data/recipes/           shipped scientific recipes (inside the package)
└── interfaces/
    ├── errors.py           one actionable sentence for a failed turn, shared by all three
    ├── cli.py              readline CLI
    ├── jupyter_magic.py    %%helioai
    └── web/                FastAPI + SSE + vanilla JS

The agent loop

There is one loop, runtime.Runner. Each turn: send the compacted history plus tool definitions to the model, start the turn's tool calls together, dispatch each in the model's order, review the figures, emit the events, append the results, repeat until the model answers with text or the turn budget is hit. Every step yields an event, which is what lets all four interfaces render progress live from the same source.

What varies is a runtime.Policy: the lead's (stream_chat) shows every tool plus the task and skill tools, thinks aloud, persists the history and closes with the answer checks; a role's (stream_subagent) shows and allows only its whitelist, must call a tool on its first turn, and reports findings, summary and usage back to the lead instead of persisting anything. The two loops used to be copies of each other and drifted the way copies do; the wrappers are now a few dozen lines each.

A run carries a runtime.RunContext — user, session, session directory, agent, network flag. The runner binds it to the workspace contextvars for the duration of the run (the one place they are set), and the four tools that write — run_python, get_timeseries, get_events_timeseries, save_catalog — receive their directories from it as trusted arguments (tool_exec.trusted_args), so a tool called over MCP or from a test writes where its caller said. The contextvars remain the hot path for everything that only reads.

A lead turn closes with two judgements, both descriptive and neither blocking. runtime.validator.validate runs the answer checks in one place — catalogue ids, recipe bypass, the numbers in the prose, the figure reviews — and, when the model closed with final_answer(answer, claims), places each named number against the provenance ledger by name with a unit-aware tolerance; the result is one verdict event. runtime.plan.adherence compares the tools the lead called with the plan it presented and emits one plan_report. Both are journaled with the rest of the turn. A third, off by default, reads the question itself: with a judging backend and the judgment_intent experiment, core.judgment asks a System One judge what the question committed the answer to and core.joins places that contract against what the turn loaded and delivered — one intent event, observation only. See The judgment layer.

Registry

Tools are async functions registered with a JSON Schema:

@registry.register(name="plasma_beta", description="...", parameters={...})
async def plasma_beta(B_nT: float, n_cm3: float, T_eV: float) -> dict: ...

call_tool always awaits, so every tool must be async. Arguments starting with _ are rejected from model-supplied input — framework-injected parameters travel through a separate trusted channel so generated code cannot spoof them.

Every call returns a ToolResult (tools/results.py): the payload the tool produced, parsed once, and for_llm(), the exact text appended to the history. Readers — artifact extraction, the figure review, the MCP server's isError — work on the payload; nothing downstream re-parses the model's text to learn what a tool returned.

Storage

Everything is namespaced per user, then per session:

<data_dir>/users/<user_id>/workspace/<session>/   figures, scripts, npz, manifest.json
<data_dir>/users/<user_id>/catalogs/              saved catalogs
<data_dir>/chroma/                                the shared parameter index
<data_dir>/sessions.db                            SQLite history, usage and event journal

<data_dir> is <repo>/data from a clone and ~/.local/share/helioai when installed — see config._default_data_dir. The agent's hot path resolves paths from a contextvar; every other entry point passes user_id explicitly, because a contextvar set inside stream_chat is not visible to a CLI subcommand.

Sandbox

run_python spawns a subprocess. On Linux with bubblewrap available: read-only root, an environment allowlist rather than inherited os.environ (so API keys are unreadable), dropped privileges, and resource limits. Elsewhere it degrades to a plain subprocess with only the timeout and process-group kill — documented in SECURITY.md.

The subprocess always gets start_new_session=True. Without it, killing the process group on timeout kills the HelioAI server itself.

LLM providers

Groq, Ollama and Azure all speak the OpenAI chat-completions format, so they share OpenAICompatClient; a provider is a base_url entry in factory.OPENAI_COMPAT. Azure subclasses it for AsyncAzureOpenAI and two dialect quirks. Gemini keeps a native client because its wire format genuinely differs — notably it has no tool-call ids, so the client synthesises name::hex and parses the name back out.

Where to look first

To change Start at
how the agent decides core/agent_loop.py + core/skills/
what the agent can do tools/setup.py
how parameters are found tools/rag.py
what a notebook looks like export.py
how a provider is added core/llm/factory.py