Skip to content

Interfaces

See Interfaces for how to use each surface.

Error messages

helioai.interfaces.errors

One sentence for a failed turn, shared by the CLI, the web UI and Jupyter.

The first thing a new install does is fail: no key, a local server not started, a key pasted with a space. Each interface used to show that failure its own way — 264 lines of traceback in the terminal, the SDK's bare "Connection error." in the browser — and none of them said what to do next. The translation lives here, once, so the three surfaces say the same thing and the fix is written next to the problem.

setup_problem

setup_problem(provider: str | None = None) -> str | None

What stops provider from answering before a single request is sent, or None.

build_llm_client already refuses a missing key. A missing model it lets through, on purpose for OpenCode: the gateway has no sensible default, so its model is empty until the user picks one — and the request then fails at the gateway with an error that does not name the variable to set. Checked here, before the question is spent.

Source code in helioai/interfaces/errors.py
def setup_problem(provider: str | None = None) -> str | None:
    """What stops `provider` from answering before a single request is sent, or None.

    `build_llm_client` already refuses a missing key. A missing *model* it lets through,
    on purpose for OpenCode: the gateway has no sensible default, so its `model` is empty
    until the user picks one — and the request then fails at the gateway with an error
    that does not name the variable to set. Checked here, before the question is spent.
    """
    from helioai.config import settings

    p = _provider(provider)
    cfg = getattr(settings.llm, p, None)
    if p == "opencode" and cfg is not None and not cfg.model:
        return (
            "HELIOAI_OPENCODE_MODEL is not set — OpenCode has no default model. "
            "Set it to a model id from your OpenCode dashboard (e.g. deepseek-v4-pro)."
        )
    return None

describe_llm_error

describe_llm_error(exc: BaseException, provider: str | None = None) -> str

Turn an exception raised while answering a question into one actionable line.

Parameters:

Name Type Description Default
exc BaseException

What build_llm_client or the agent loop raised.

required
provider str | None

The provider the turn ran on; the configured one when omitted.

None

Returns:

Type Description
str

A sentence saying what failed and how to fix it. Errors this function does not

str

recognise keep their type and message, and point at DEBUG logging for the

str

traceback, rather than being dressed up as a diagnosis they are not.

Source code in helioai/interfaces/errors.py
def describe_llm_error(exc: BaseException, provider: str | None = None) -> str:
    """Turn an exception raised while answering a question into one actionable line.

    Args:
        exc: What `build_llm_client` or the agent loop raised.
        provider: The provider the turn ran on; the configured one when omitted.

    Returns:
        A sentence saying what failed and how to fix it. Errors this function does not
        recognise keep their type and message, and point at DEBUG logging for the
        traceback, rather than being dressed up as a diagnosis they are not.
    """
    p = _provider(provider)
    message = str(exc).strip()

    if isinstance(exc, RuntimeError) and not _status(exc):
        if " is not set" in message:
            message = message.replace(" is not set in .env", " is not set")
            return (
                f"{message}. Set it in your shell or in a .env file, or choose another "
                f"provider with HELIOAI_LLM_PROVIDER. {_DOCTOR}"
            )
        if message.startswith("Unknown LLM provider"):
            return f"{message}. {_DOCTOR}"
        return message

    name = type(exc).__name__
    if name in ("APIConnectionError", "APITimeoutError", "ConnectError", "ConnectTimeout"):
        where = _endpoint(p, exc)
        where = f" at {where}" if where else ""
        hint = (
            " Is Ollama running? Start it with `ollama serve`."
            if p == "ollama"
            else " Check your network connection and the endpoint URL."
        )
        verb = "timed out reaching" if "Timeout" in name else "could not reach"
        return f"HelioAI {verb} the {p} server{where}.{hint}"

    status = _status(exc)
    if status in (401, 403):
        key = _KEY_ENV.get(p)
        fix = f" — check {key}" if key else ""
        return f"The {p} server refused the credentials (HTTP {status}){fix}. {_DOCTOR}"
    if status == 404:
        model = _MODEL_ENV.get(p)
        fix = f" Check {model}." if model else ""
        return f"The {p} server does not know the requested model (HTTP 404).{fix}"
    if status == 429:
        return (
            f"The {p} server is rate-limiting this key or its quota is used up (HTTP 429). "
            "Wait a moment, or switch provider with HELIOAI_LLM_PROVIDER."
        )
    if status is not None and status >= 500:
        return f"The {p} server failed (HTTP {status}). This is on the provider's side; retry."
    if status is not None:
        return f"The {p} server rejected the request (HTTP {status}): {message}"

    detail = f"{name}: {message}" if message else name
    return f"Unexpected error — {detail}. Set HELIOAI_LOG_LEVEL=DEBUG to see the traceback."

Command line

helioai.interfaces.cli

Interactive CLI for HelioAI.

Usage

helioai # interactive session (/help lists its commands) helioai "your query" # one-shot query helioai --resume # pick a past session and continue it helioai history # list sessions helioai history delete # delete a session and its workspace helioai index [--rebuild --classify] # fetch the prebuilt index when empty, else (re)index locally; --classify fills measurement types and regions helioai index --download # replace the index with the one published for this release helioai index --export DIR # write the index as a snapshot (what CI publishes) helioai export [id] # export a session as a reproducible .ipynb helioai profile # edit the user profile helioai mcp-install [--write] # MCP client config pointing at this install helioai serve --web # web UI on :7890 (--host, --port) helioai serve # MCP server on stdio helioai serve --http # MCP server over HTTP on :8765 (--host, --port) helioai migrate-storage # move legacy data into the per-user layout helioai doctor [--online] # check this install: index, sandbox, keys, .env (--json)

Options

--session # continue a specific session --dev # supply the dev token, lifting the scope guard -h, --help # print this and exit -V, --version # print the version and exit

main

main() -> None

Entry point for the helioai command.

Routes subcommands (index, export, history, profile, serve, doctor, ...) and otherwise runs either a one-shot query or the interactive prompt.

--help is answered before anything else runs. The default branch of this router treats an unrecognised argument as a question, so until it was handled, helioai --help created a workspace and billed an LLM call to ask the model what --help meant — the first thing anyone types after pip install. Printing __doc__ keeps the help and the module's own documentation as one string. Only a standalone --help token counts: a quoted question that happens to contain the words is still a question.

Source code in helioai/interfaces/cli.py
def main() -> None:
    """Entry point for the `helioai` command.

    Routes subcommands (index, export, history, profile, serve, doctor, ...) and
    otherwise runs either a one-shot query or the interactive prompt.

    `--help` is answered before anything else runs. The default branch of this
    router treats an unrecognised argument as a question, so until it was
    handled, `helioai --help` created a workspace and billed an LLM call to ask
    the model what `--help` meant — the first thing anyone types after
    `pip install`. Printing `__doc__` keeps the help and the module's own
    documentation as one string. Only a standalone `--help` token counts: a quoted
    question that happens to contain the words is still a question.
    """
    global _SESSION_ID

    args = sys.argv[1:]
    if {"-h", "--help"} & set(args):
        _print(__doc__)
        return
    if {"-V", "--version"} & set(args):
        from helioai import __version__

        _print(f"helioai {__version__}")
        return

    options, rest = _global_options(args)
    mistake = _not_a_question(rest)
    if mistake:
        _print(mistake, file=sys.stderr)
        raise SystemExit(2)

    if rest and rest[0] in _COMMANDS and rest[0] != "doctor":
        # Storage commands need the user bound; doctor must not touch the workspace.
        from helioai.workspace import set_user

        set_user(_USER_ID)
        _run_command(rest[0], rest[1:])
        return
    if rest and rest[0] == "doctor":
        _run_command("doctor", rest[1:])
        return

    from helioai.config import dev_unlock, settings
    from helioai.workspace import cleanup_old_runs, set_user

    set_user(_USER_ID)
    cleanup_old_runs()
    restricted = not dev_unlock(settings.dev.token if options.dev else None)
    if options.session:
        _SESSION_ID = options.session

    if options.resume:
        session_id = _pick_session()
        if session_id:
            _SESSION_ID = session_id
        _interactive(restricted=restricted)
        return

    if not rest:
        _interactive(restricted=restricted)
        return

    if not asyncio.run(_run_query(" ".join(rest), restricted=restricted)):
        raise SystemExit(1)

Jupyter magic

helioai.interfaces.jupyter_magic

Jupyter IPython magics for HelioAI.

Load with

%load_ext helioai.interfaces.jupyter_magic

Cell magic

%%helioai solar wind density ACE 2005-01-17

Line magics

%helioai_session new|reset|delete %helioai_provider [opencode|groq|gemini|azure|ollama] %helioai_history %helioai_resume %helioai_export [session_id] %helioai_dev on|off — toggle dev mode (bypasses helio-only scope guardrail)

HelioAIMagics

Bases: Magics

IPython magics exposing the agent inside a notebook.

Source code in helioai/interfaces/jupyter_magic.py
@magics_class
class HelioAIMagics(Magics):
    """IPython magics exposing the agent inside a notebook."""

    def __init__(self, shell=None, **kwargs) -> None:
        super().__init__(shell, **kwargs)
        # The provider chosen with %helioai_provider, None until the user picks one.
        # Held on the instance and passed to the factory explicitly — the web UI does
        # the same with its request field. Writing HELIOAI_LLM_PROVIDER into the
        # environment, as this used to, changed nothing: settings reads it once, at import.
        self._provider: str | None = None

    @cell_magic
    def helioai(self, line: str, cell: str) -> None:
        """`%%helioai` — send a natural-language query to the agent.

        Figures render inline; parameter cards and catalog previews render as HTML.

        Example:
            %load_ext helioai.interfaces.jupyter_magic

            %%helioai
            Download ACE IMF for the 2015-03-17 storm, plot Bz and mark
            the shock arrival.
        """
        import helioai.tools.setup  # noqa: F401
        from helioai.core.agent_loop import stream_chat
        from helioai.interfaces.errors import describe_llm_error, setup_problem
        from helioai.logging_config import setup_logging

        setup_logging("WARNING", tracebacks=False)

        def _error(message: str) -> None:
            _render_jupyter_event({"event": "error", "data": {"message": message}})

        async def _run():
            problem = setup_problem(self._provider)
            if problem:
                _error(problem)
                return
            try:
                llm = _get_llm(self._provider)
            except RuntimeError as e:
                _error(describe_llm_error(e, self._provider))
                return
            try:
                async for ev in stream_chat(
                    llm, _USER_ID, _SESSION_ID, cell.strip(), restricted=_dev_restricted
                ):
                    _render_jupyter_event(ev)
            except Exception as e:
                _error(describe_llm_error(e, self._provider))
            finally:
                # Must happen inside this loop: the pool is bound to it, and
                # `_run_async` closes the loop the moment this returns.
                await llm.aclose()

        _run_async(_run())

    @line_magic
    def helioai_session(self, line: str) -> None:
        """`%helioai_session new|reset|delete <id>` — manage the active session.

        `new` starts a fresh conversation and keeps the previous one; `reset` deletes
        the current one first. The distinction matters at the end of a demo: the
        analysis just exported must survive the next question.
        """
        global _SESSION_ID
        parts = line.strip().split(maxsplit=1)
        cmd = parts[0] if parts else ""
        arg = parts[1] if len(parts) > 1 else ""

        if cmd == "new":
            _SESSION_ID = str(uuid.uuid4())
            print(f"New session. Id: {_SESSION_ID[:8]} (previous one kept).")
        elif cmd == "reset":
            from helioai.core.session import store

            store.reset(_USER_ID, _SESSION_ID)
            _SESSION_ID = str(uuid.uuid4())
            print(f"Session reset. New id: {_SESSION_ID[:8]}")
        elif cmd == "delete":
            from helioai.core.session import store
            from helioai.workspace import _root

            if not arg:
                print("Usage: %helioai_session delete <session_id_prefix>")
                return
            all_ids = store.all_sessions(_USER_ID)
            matches = [s for s in all_ids if s.startswith(arg)]
            if not matches:
                print(f"No session matching {arg!r}.")
                return
            sid = matches[0]
            wdir = store.get_workspace_dir(_USER_ID, sid)
            store.reset(_USER_ID, sid)
            if wdir:
                import shutil

                ws_path = _root() / wdir
                if ws_path.exists():
                    shutil.rmtree(ws_path, ignore_errors=True)
            if sid == _SESSION_ID:
                _SESSION_ID = str(uuid.uuid4())
                print(f"Current session deleted. New id: {_SESSION_ID[:8]}")
            else:
                print(f"Session {sid[:8]} deleted.")
        else:
            print(f"Unknown command: {cmd!r}. Use 'new', 'reset' or 'delete <id>'.")

    @line_magic
    def helioai_provider(self, line: str) -> None:
        """`%helioai_provider [name]` — show or switch the LLM provider for the next cells.

        The choice lives on this magics instance, so it lasts for the kernel and is
        handed to `build_llm_client` on every cell. The provider names come from the
        factory rather than a list kept here, which is how ollama went missing once.
        """
        from helioai.config import settings
        from helioai.core.llm.factory import OPENAI_COMPAT

        known = ["azure", "gemini", *OPENAI_COMPAT]
        provider = line.strip().lower()
        if not provider:
            print(
                f"Provider: {self._provider or settings.llm.provider!r}. Known: {' | '.join(known)}"
            )
            return
        if provider not in known:
            print(f"Unknown provider {provider!r}. Use: {' | '.join(known)}")
            return
        self._provider = provider
        print(f"Provider switched to {provider!r}.")

    @line_magic
    def helioai_history(self, line: str) -> None:
        """`%helioai_history` — list recent sessions."""
        from helioai.core.session import store

        summaries = store.list_summaries(_USER_ID)
        if not summaries:
            print("No history found.")
            return
        rows = "".join(
            f"<tr>"
            f"<td><code>{s['session_id'][:8]}</code></td>"
            f"<td>{s['updated_at'][:16].replace('T', ' ')}</td>"
            f"<td style='text-align:center'>{s['n_messages']}</td>"
            f"<td>{s['first_message']}</td>"
            f"</tr>"
            for s in summaries
        )
        display(
            HTML(
                "<table><thead><tr>"
                "<th>Session</th><th>Updated</th><th>Msgs</th><th>First message</th>"
                "</tr></thead><tbody>" + rows + "</tbody></table>"
            )
        )

    @line_magic
    def helioai_profile(self, line: str) -> None:
        """`%helioai_profile` — show or edit the user profile."""
        from helioai.workspace import user_home

        parts = line.strip().split(maxsplit=1)
        cmd = parts[0] if parts else ""
        arg = parts[1].strip().strip("\"'") if len(parts) > 1 else ""
        p = user_home(_USER_ID) / "profile.md"

        if cmd == "show":
            content = p.read_text(encoding="utf-8").strip() if p.exists() else ""
            display(Markdown(content if content else "_(profil vide)_"))
        elif cmd == "set":
            if not arg:
                print('Usage: %helioai_profile set "your preferences here"')
                return
            p.parent.mkdir(parents=True, exist_ok=True)
            with p.open("a", encoding="utf-8") as f:
                f.write(("\n" if p.stat().st_size > 0 else "") + arg + "\n")
            print(f"Profile updated ({p}).")
        else:
            print('Usage: %helioai_profile show | set "<text>"')

    @line_magic
    def helioai_export(self, line: str) -> None:
        """`%helioai_export` — export the session as a standalone notebook."""
        from helioai.core.session import store
        from helioai.export import export_session_notebook

        prefix = line.strip()
        session_id = _SESSION_ID
        if prefix:
            matches = [s for s in store.all_sessions(_USER_ID) if s.startswith(prefix)]
            if not matches:
                print(f"No session matching {prefix!r}.")
                return
            session_id = matches[0]
        path = export_session_notebook(_USER_ID, session_id)
        from IPython.display import FileLink

        display(FileLink(str(path), result_html_prefix="📓 Exported notebook: "))

    @line_magic
    def helioai_resume(self, line: str) -> None:
        """`%helioai_resume` — pick a previous session to continue."""
        global _SESSION_ID
        prefix = line.strip()
        if not prefix:
            print("Usage: %helioai_resume <session_id>")
            return
        from helioai.core.session import store

        all_ids = store.all_sessions(_USER_ID)
        matches = [s for s in all_ids if s.startswith(prefix)]
        if not matches:
            print(f"No session found matching {prefix!r}.")
            return
        _SESSION_ID = matches[0]
        msgs = store.get_or_create(_USER_ID, _SESSION_ID)
        print(f"Resumed session {_SESSION_ID[:8]} ({len(msgs)} messages).")

    @line_magic
    def helioai_dev(self, line: str) -> None:
        """`%helioai_dev <token>` — unlock unrestricted mode for this session."""
        global _dev_restricted
        from helioai.config import dev_unlock, settings

        cmd = line.strip().lower()
        if cmd == "on":
            if not dev_unlock(settings.dev.token):
                print("Dev token not configured or incorrect. Set HELIOAI_DEV_TOKEN in .env.")
                return
            _dev_restricted = False
            print("Dev mode ON — scope guardrail disabled.")
        elif cmd == "off":
            _dev_restricted = True
            print("Dev mode OFF — helio-only scope guardrail active.")
        else:
            status = "OFF (restricted)" if _dev_restricted else "ON (unrestricted)"
            print(f"Dev mode: {status}. Usage: %helioai_dev on | off")

helioai

helioai(line: str, cell: str) -> None

%%helioai — send a natural-language query to the agent.

Figures render inline; parameter cards and catalog previews render as HTML.

Example

%load_ext helioai.interfaces.jupyter_magic

%%helioai Download ACE IMF for the 2015-03-17 storm, plot Bz and mark the shock arrival.

Source code in helioai/interfaces/jupyter_magic.py
@cell_magic
def helioai(self, line: str, cell: str) -> None:
    """`%%helioai` — send a natural-language query to the agent.

    Figures render inline; parameter cards and catalog previews render as HTML.

    Example:
        %load_ext helioai.interfaces.jupyter_magic

        %%helioai
        Download ACE IMF for the 2015-03-17 storm, plot Bz and mark
        the shock arrival.
    """
    import helioai.tools.setup  # noqa: F401
    from helioai.core.agent_loop import stream_chat
    from helioai.interfaces.errors import describe_llm_error, setup_problem
    from helioai.logging_config import setup_logging

    setup_logging("WARNING", tracebacks=False)

    def _error(message: str) -> None:
        _render_jupyter_event({"event": "error", "data": {"message": message}})

    async def _run():
        problem = setup_problem(self._provider)
        if problem:
            _error(problem)
            return
        try:
            llm = _get_llm(self._provider)
        except RuntimeError as e:
            _error(describe_llm_error(e, self._provider))
            return
        try:
            async for ev in stream_chat(
                llm, _USER_ID, _SESSION_ID, cell.strip(), restricted=_dev_restricted
            ):
                _render_jupyter_event(ev)
        except Exception as e:
            _error(describe_llm_error(e, self._provider))
        finally:
            # Must happen inside this loop: the pool is bound to it, and
            # `_run_async` closes the loop the moment this returns.
            await llm.aclose()

    _run_async(_run())

helioai_session

helioai_session(line: str) -> None

%helioai_session new|reset|delete <id> — manage the active session.

new starts a fresh conversation and keeps the previous one; reset deletes the current one first. The distinction matters at the end of a demo: the analysis just exported must survive the next question.

Source code in helioai/interfaces/jupyter_magic.py
@line_magic
def helioai_session(self, line: str) -> None:
    """`%helioai_session new|reset|delete <id>` — manage the active session.

    `new` starts a fresh conversation and keeps the previous one; `reset` deletes
    the current one first. The distinction matters at the end of a demo: the
    analysis just exported must survive the next question.
    """
    global _SESSION_ID
    parts = line.strip().split(maxsplit=1)
    cmd = parts[0] if parts else ""
    arg = parts[1] if len(parts) > 1 else ""

    if cmd == "new":
        _SESSION_ID = str(uuid.uuid4())
        print(f"New session. Id: {_SESSION_ID[:8]} (previous one kept).")
    elif cmd == "reset":
        from helioai.core.session import store

        store.reset(_USER_ID, _SESSION_ID)
        _SESSION_ID = str(uuid.uuid4())
        print(f"Session reset. New id: {_SESSION_ID[:8]}")
    elif cmd == "delete":
        from helioai.core.session import store
        from helioai.workspace import _root

        if not arg:
            print("Usage: %helioai_session delete <session_id_prefix>")
            return
        all_ids = store.all_sessions(_USER_ID)
        matches = [s for s in all_ids if s.startswith(arg)]
        if not matches:
            print(f"No session matching {arg!r}.")
            return
        sid = matches[0]
        wdir = store.get_workspace_dir(_USER_ID, sid)
        store.reset(_USER_ID, sid)
        if wdir:
            import shutil

            ws_path = _root() / wdir
            if ws_path.exists():
                shutil.rmtree(ws_path, ignore_errors=True)
        if sid == _SESSION_ID:
            _SESSION_ID = str(uuid.uuid4())
            print(f"Current session deleted. New id: {_SESSION_ID[:8]}")
        else:
            print(f"Session {sid[:8]} deleted.")
    else:
        print(f"Unknown command: {cmd!r}. Use 'new', 'reset' or 'delete <id>'.")

helioai_provider

helioai_provider(line: str) -> None

%helioai_provider [name] — show or switch the LLM provider for the next cells.

The choice lives on this magics instance, so it lasts for the kernel and is handed to build_llm_client on every cell. The provider names come from the factory rather than a list kept here, which is how ollama went missing once.

Source code in helioai/interfaces/jupyter_magic.py
@line_magic
def helioai_provider(self, line: str) -> None:
    """`%helioai_provider [name]` — show or switch the LLM provider for the next cells.

    The choice lives on this magics instance, so it lasts for the kernel and is
    handed to `build_llm_client` on every cell. The provider names come from the
    factory rather than a list kept here, which is how ollama went missing once.
    """
    from helioai.config import settings
    from helioai.core.llm.factory import OPENAI_COMPAT

    known = ["azure", "gemini", *OPENAI_COMPAT]
    provider = line.strip().lower()
    if not provider:
        print(
            f"Provider: {self._provider or settings.llm.provider!r}. Known: {' | '.join(known)}"
        )
        return
    if provider not in known:
        print(f"Unknown provider {provider!r}. Use: {' | '.join(known)}")
        return
    self._provider = provider
    print(f"Provider switched to {provider!r}.")

helioai_history

helioai_history(line: str) -> None

%helioai_history — list recent sessions.

Source code in helioai/interfaces/jupyter_magic.py
@line_magic
def helioai_history(self, line: str) -> None:
    """`%helioai_history` — list recent sessions."""
    from helioai.core.session import store

    summaries = store.list_summaries(_USER_ID)
    if not summaries:
        print("No history found.")
        return
    rows = "".join(
        f"<tr>"
        f"<td><code>{s['session_id'][:8]}</code></td>"
        f"<td>{s['updated_at'][:16].replace('T', ' ')}</td>"
        f"<td style='text-align:center'>{s['n_messages']}</td>"
        f"<td>{s['first_message']}</td>"
        f"</tr>"
        for s in summaries
    )
    display(
        HTML(
            "<table><thead><tr>"
            "<th>Session</th><th>Updated</th><th>Msgs</th><th>First message</th>"
            "</tr></thead><tbody>" + rows + "</tbody></table>"
        )
    )

helioai_profile

helioai_profile(line: str) -> None

%helioai_profile — show or edit the user profile.

Source code in helioai/interfaces/jupyter_magic.py
@line_magic
def helioai_profile(self, line: str) -> None:
    """`%helioai_profile` — show or edit the user profile."""
    from helioai.workspace import user_home

    parts = line.strip().split(maxsplit=1)
    cmd = parts[0] if parts else ""
    arg = parts[1].strip().strip("\"'") if len(parts) > 1 else ""
    p = user_home(_USER_ID) / "profile.md"

    if cmd == "show":
        content = p.read_text(encoding="utf-8").strip() if p.exists() else ""
        display(Markdown(content if content else "_(profil vide)_"))
    elif cmd == "set":
        if not arg:
            print('Usage: %helioai_profile set "your preferences here"')
            return
        p.parent.mkdir(parents=True, exist_ok=True)
        with p.open("a", encoding="utf-8") as f:
            f.write(("\n" if p.stat().st_size > 0 else "") + arg + "\n")
        print(f"Profile updated ({p}).")
    else:
        print('Usage: %helioai_profile show | set "<text>"')

helioai_export

helioai_export(line: str) -> None

%helioai_export — export the session as a standalone notebook.

Source code in helioai/interfaces/jupyter_magic.py
@line_magic
def helioai_export(self, line: str) -> None:
    """`%helioai_export` — export the session as a standalone notebook."""
    from helioai.core.session import store
    from helioai.export import export_session_notebook

    prefix = line.strip()
    session_id = _SESSION_ID
    if prefix:
        matches = [s for s in store.all_sessions(_USER_ID) if s.startswith(prefix)]
        if not matches:
            print(f"No session matching {prefix!r}.")
            return
        session_id = matches[0]
    path = export_session_notebook(_USER_ID, session_id)
    from IPython.display import FileLink

    display(FileLink(str(path), result_html_prefix="📓 Exported notebook: "))

helioai_resume

helioai_resume(line: str) -> None

%helioai_resume — pick a previous session to continue.

Source code in helioai/interfaces/jupyter_magic.py
@line_magic
def helioai_resume(self, line: str) -> None:
    """`%helioai_resume` — pick a previous session to continue."""
    global _SESSION_ID
    prefix = line.strip()
    if not prefix:
        print("Usage: %helioai_resume <session_id>")
        return
    from helioai.core.session import store

    all_ids = store.all_sessions(_USER_ID)
    matches = [s for s in all_ids if s.startswith(prefix)]
    if not matches:
        print(f"No session found matching {prefix!r}.")
        return
    _SESSION_ID = matches[0]
    msgs = store.get_or_create(_USER_ID, _SESSION_ID)
    print(f"Resumed session {_SESSION_ID[:8]} ({len(msgs)} messages).")

helioai_dev

helioai_dev(line: str) -> None

%helioai_dev <token> — unlock unrestricted mode for this session.

Source code in helioai/interfaces/jupyter_magic.py
@line_magic
def helioai_dev(self, line: str) -> None:
    """`%helioai_dev <token>` — unlock unrestricted mode for this session."""
    global _dev_restricted
    from helioai.config import dev_unlock, settings

    cmd = line.strip().lower()
    if cmd == "on":
        if not dev_unlock(settings.dev.token):
            print("Dev token not configured or incorrect. Set HELIOAI_DEV_TOKEN in .env.")
            return
        _dev_restricted = False
        print("Dev mode ON — scope guardrail disabled.")
    elif cmd == "off":
        _dev_restricted = True
        print("Dev mode OFF — helio-only scope guardrail active.")
    else:
        status = "OFF (restricted)" if _dev_restricted else "ON (unrestricted)"
        print(f"Dev mode: {status}. Usage: %helioai_dev on | off")

load_ipython_extension

load_ipython_extension(ipython: Any) -> None

Register the HelioAI magics. Called by %load_ext.

Also sweeps expired session workspaces: a notebook kernel is the one HelioAI process the CLI's startup sweep never runs in.

Parameters:

Name Type Description Default
ipython Any

The active InteractiveShell, supplied by IPython itself. This hook name and signature are IPython's contract, not ours.

required
Source code in helioai/interfaces/jupyter_magic.py
def load_ipython_extension(ipython: Any) -> None:
    """Register the HelioAI magics. Called by `%load_ext`.

    Also sweeps expired session workspaces: a notebook kernel is the one HelioAI
    process the CLI's startup sweep never runs in.

    Args:
        ipython: The active InteractiveShell, supplied by IPython itself. This
            hook name and signature are IPython's contract, not ours.
    """
    from helioai.workspace import cleanup_old_runs

    ipython.register_magics(HelioAIMagics)
    cleanup_old_runs()

Web application

helioai.interfaces.web.app

FastAPI web interface for HelioAI.

Single-user, no auth. Streams agent events as SSE. Figures from the sandbox are served via /figure?path=.

require_user async

require_user(x_helio_token: str | None = Header(default=None)) -> str

Resolve the caller's user_id from the X-Helio-Token header.

No users configured (local dev) → single shared user, no auth. Once HELIOAI_USERS is set (deployment), a valid nominative token is required.

Parameters:

Name Type Description Default
x_helio_token str | None

The X-Helio-Token header, absent in local dev.

Header(default=None)

Returns:

Type Description
str

The user id owning storage for this request.

Raises:

Type Description
HTTPException

401 when users are configured and the token is unknown.

Source code in helioai/interfaces/web/app.py
async def require_user(x_helio_token: str | None = Header(default=None)) -> str:
    """Resolve the caller's user_id from the X-Helio-Token header.

    No users configured (local dev) → single shared user, no auth. Once
    HELIOAI_USERS is set (deployment), a valid nominative token is required.

    Args:
        x_helio_token: The `X-Helio-Token` header, absent in local dev.

    Returns:
        The user id owning storage for this request.

    Raises:
        HTTPException: 401 when users are configured and the token is unknown.
    """
    users = settings.web_auth.users
    if not users:
        return _DEFAULT_USER
    # Compared token by token in constant time, like the dev token and the MCP bearer:
    # a dict lookup leaks how much of a guess matched through its timing.
    if x_helio_token:
        for token, user_id in users.items():
            if hmac.compare_digest(token, x_helio_token):
                return user_id
    raise HTTPException(status_code=401, detail="Invalid or missing token")

harden_for_host

harden_for_host(app: FastAPI, host: str) -> FastAPI

Add the middleware a given bind address calls for, and return the app.

Kept apart from serve_web so a test can build exactly what uvicorn will serve: added inside serve_web, the host guard was never on the app the TestClient imported, and the DNS-rebinding defence went untested for a year.

Parameters:

Name Type Description Default
app FastAPI

The FastAPI application.

required
host str

The address about to be bound.

required

Returns:

Type Description
FastAPI

The same app, so the call reads as an expression.

Source code in helioai/interfaces/web/app.py
def harden_for_host(app: FastAPI, host: str) -> FastAPI:
    """Add the middleware a given bind address calls for, and return the app.

    Kept apart from `serve_web` so a test can build exactly what uvicorn will serve:
    added inside `serve_web`, the host guard was never on the `app` the TestClient
    imported, and the DNS-rebinding defence went untested for a year.

    Args:
        app: The FastAPI application.
        host: The address about to be bound.

    Returns:
        The same app, so the call reads as an expression.
    """
    if host in _LOOPBACK_HOSTS:
        # A loopback bind is not a boundary: any web page can resolve its own domain
        # to 127.0.0.1 and reach this server (DNS rebinding). Pinning Host costs
        # nothing here and CORS does not cover it.
        from starlette.middleware.trustedhost import TrustedHostMiddleware

        app.add_middleware(TrustedHostMiddleware, allowed_hosts=sorted(_LOOPBACK_HOSTS))
    else:
        log.warning("web_exposed_beyond_loopback", host=host)
    return app

index async

index()

Serve the single-page web UI.

Source code in helioai/interfaces/web/app.py
@app.get("/")
async def index():
    """Serve the single-page web UI."""
    return FileResponse(_STATIC / "index.html")

favicon async

favicon()

The icon browsers ask for at the root whatever the page links, and 404'd on.

Source code in helioai/interfaces/web/app.py
@app.get("/favicon.ico", include_in_schema=False)
async def favicon():
    """The icon browsers ask for at the root whatever the page links, and 404'd on."""
    return FileResponse(_STATIC / "favicon.ico", media_type="image/x-icon")

health async

health()

Liveness probe. Returns {"status": "ok"}.

Source code in helioai/interfaces/web/app.py
@app.get("/health")
async def health():
    """Liveness probe. Returns `{"status": "ok"}`."""
    return {"status": "ok"}

api_config async

api_config()

Server-side settings the UI cannot know on its own — never a secret, only whether one is set.

The provider selector used to default to whichever option came first in the markup — azure — and sent it on every message, so a server configured for another provider was quietly overridden by the browser. providers says which of them this server can actually reach, so the selector stops offering a key nobody set.

auth and dev_token decide what the sidebar's token field is for: with HELIOAI_USERS it is the access token every request needs, with only HELIOAI_DEV_TOKEN it is the optional dev token, and with neither it is hidden — a field asking a local user for a token that does not exist was the first thing a newcomer asked about.

Source code in helioai/interfaces/web/app.py
@app.get("/api/config")
async def api_config():
    """Server-side settings the UI cannot know on its own — never a secret, only whether
    one is set.

    The provider selector used to default to whichever option came first in the markup
    — `azure` — and sent it on every message, so a server configured for another
    provider was quietly overridden by the browser. `providers` says which of them this
    server can actually reach, so the selector stops offering a key nobody set.

    `auth` and `dev_token` decide what the sidebar's token field is for: with
    `HELIOAI_USERS` it is the access token every request needs, with only
    `HELIOAI_DEV_TOKEN` it is the optional dev token, and with neither it is hidden —
    a field asking a local user for a token that does not exist was the first thing a
    newcomer asked about.
    """
    return {
        "provider": settings.llm.provider,
        "providers": _provider_status(),
        "auth": bool(settings.web_auth.users),
        "dev_token": bool(settings.dev.token),
    }

chat_stream async

chat_stream(req: _ChatRequest, x_helio_dev_token: str | None = Header(default=None), user_id: str = Depends(require_user)) -> StreamingResponse

Stream one agent turn as Server-Sent Events.

Each agent event — tool calls, results, artifacts, sub-agent activity — is forwarded as it happens, which is what drives the live activity dock.

Parameters:

Name Type Description Default
req _ChatRequest

Body carrying the question and the session to continue.

required
x_helio_dev_token str | None

Optional dev token lifting the scope guardrail.

Header(default=None)
user_id str

Resolved by require_user.

Depends(require_user)

Returns:

Type Description
StreamingResponse

A StreamingResponse of SSE frames, one per agent event.

Source code in helioai/interfaces/web/app.py
@app.post("/chat/stream")
async def chat_stream(
    req: _ChatRequest,
    x_helio_dev_token: str | None = Header(default=None),
    user_id: str = Depends(require_user),
) -> StreamingResponse:
    """Stream one agent turn as Server-Sent Events.

    Each agent event — tool calls, results, artifacts, sub-agent activity — is
    forwarded as it happens, which is what drives the live activity dock.

    Args:
        req: Body carrying the question and the session to continue.
        x_helio_dev_token: Optional dev token lifting the scope guardrail.
        user_id: Resolved by `require_user`.

    Returns:
        A `StreamingResponse` of SSE frames, one per agent event.
    """
    # Authenticated nominative users are trusted → unrestricted; the legacy dev
    # token still unlocks scope when no users are configured (local dev).
    restricted = not (bool(settings.web_auth.users) or dev_unlock(x_helio_dev_token))

    # stream_chat serialises turns per session with a lock; answering 409 here is
    # only so a second tab fails fast instead of looking hung while it queues.
    if store.is_busy(user_id, req.session_id):
        raise HTTPException(status_code=409, detail="a reply is already streaming for this session")

    async def gen():
        llm = None
        try:
            problem = setup_problem(req.provider)
            if problem:
                raise RuntimeError(problem)
            llm = build_llm_client(req.provider)
            async for ev in stream_chat(
                llm, user_id, req.session_id, req.message, restricted=restricted
            ):
                yield f"data: {json.dumps(ev)}\n\n"
        except Exception as e:
            message = describe_llm_error(e, req.provider)
            yield f"data: {json.dumps({'event': 'error', 'data': {'message': message}})}\n\n"
        finally:
            # One client per request, so the pool has to be released per request —
            # including when the browser disconnects mid-stream and this generator
            # is closed early.
            if llm is not None:
                await llm.aclose()

    return StreamingResponse(
        gen(),
        media_type="text/event-stream",
        headers={"X-Accel-Buffering": "no", "Cache-Control": "no-cache"},
    )

me async

me(user_id: str = Depends(require_user)) -> dict

Who the caller is and what they have spent.

The first thing a per-user quota needs is a number to compare against; until now nothing summed the token counts the providers report. Totals for today (UTC-ish: the last 24 h), the last 30 days and all time.

Returns:

Type Description
dict

{"user_id", "usage": {"day", "month", "total"}} — each a dict of

dict

prompt/completion/cached tokens and call count.

Source code in helioai/interfaces/web/app.py
@app.get("/api/me")
async def me(user_id: str = Depends(require_user)) -> dict:
    """Who the caller is and what they have spent.

    The first thing a per-user quota needs is a number to compare against; until now
    nothing summed the token counts the providers report. Totals for today (UTC-ish:
    the last 24 h), the last 30 days and all time.

    Returns:
        `{"user_id", "usage": {"day", "month", "total"}}` — each a dict of
        prompt/completion/cached tokens and call count.
    """
    return {
        "user_id": user_id,
        "usage": {
            "day": store.usage_totals(user_id, since_days=1),
            "month": store.usage_totals(user_id, since_days=30),
            "total": store.usage_totals(user_id),
        },
    }

list_sessions async

list_sessions(user_id: str = Depends(require_user)) -> list

List the calling user's sessions, most recent first.

Parameters:

Name Type Description Default
user_id str

Resolved by require_user; scopes the listing, so no caller can enumerate another's sessions.

Depends(require_user)

Returns:

Type Description
list

Session summaries, newest first.

Source code in helioai/interfaces/web/app.py
@app.get("/api/sessions")
async def list_sessions(user_id: str = Depends(require_user)) -> list:
    """List the calling user's sessions, most recent first.

    Args:
        user_id: Resolved by `require_user`; scopes the listing, so no caller
            can enumerate another's sessions.

    Returns:
        Session summaries, newest first.
    """
    return store.list_summaries(user_id)

get_session_events async

get_session_events(session_id: str, user_id: str = Depends(require_user)) -> dict

Replay a session from its journal: every event the live stream showed, in order.

The browser renders these with the same function as the live stream, so a reloaded session shows the plan, the provenance verdict, the figure reviews and the sub-agent trace exactly as they appeared.

Parameters:

Name Type Description Default
session_id str

Session to replay.

required
user_id str

Resolved by require_user; a session belonging to anyone else reads as empty rather than as a 403, which says nothing about whether it exists.

Depends(require_user)

Returns:

Type Description
dict

{"events": [...]} — empty for a session recorded before the journal existed,

dict

which the browser then fetches through /messages.

Source code in helioai/interfaces/web/app.py
@app.get("/api/sessions/{session_id}/events")
async def get_session_events(session_id: str, user_id: str = Depends(require_user)) -> dict:
    """Replay a session from its journal: every event the live stream showed, in order.

    The browser renders these with the same function as the live stream, so a reloaded
    session shows the plan, the provenance verdict, the figure reviews and the
    sub-agent trace exactly as they appeared.

    Args:
        session_id: Session to replay.
        user_id: Resolved by `require_user`; a session belonging to anyone else reads
            as empty rather than as a 403, which says nothing about whether it exists.

    Returns:
        `{"events": [...]}` — empty for a session recorded before the journal existed,
        which the browser then fetches through `/messages`.
    """
    return {"events": store.events(user_id, session_id)}

get_session_messages async

get_session_messages(session_id: str, user_id: str = Depends(require_user)) -> dict

Replay a session recorded before the event journal, from its messages.

Kept for those sessions only — see legacy_replay. A session with a journal is served by /events.

Parameters:

Name Type Description Default
session_id str

Session to replay.

required
user_id str

Resolved by require_user; another user's session reads as empty.

Depends(require_user)

Returns:

Type Description
dict

{"messages": [...]} — the stored messages in order, each assistant entry

dict

carrying the artifacts the tool calls before it produced.

Source code in helioai/interfaces/web/app.py
@app.get("/api/sessions/{session_id}/messages")
async def get_session_messages(session_id: str, user_id: str = Depends(require_user)) -> dict:
    """Replay a session recorded before the event journal, from its messages.

    Kept for those sessions only — see `legacy_replay`. A session with a journal is
    served by `/events`.

    Args:
        session_id: Session to replay.
        user_id: Resolved by `require_user`; another user's session reads as empty.

    Returns:
        `{"messages": [...]}` — the stored messages in order, each assistant entry
        carrying the artifacts the tool calls before it produced.
    """
    return {"messages": messages_view(store.get_or_create(user_id, session_id))}

get_profile async

get_profile(user_id: str = Depends(require_user)) -> dict

Return the caller's profile markdown.

Parameters:

Name Type Description Default
user_id str

Resolved by require_user.

Depends(require_user)

Returns:

Type Description
dict

{"content": markdown}, empty for a user who has written none.

Source code in helioai/interfaces/web/app.py
@app.get("/api/profile")
async def get_profile(user_id: str = Depends(require_user)) -> dict:
    """Return the caller's profile markdown.

    Args:
        user_id: Resolved by `require_user`.

    Returns:
        `{"content": markdown}`, empty for a user who has written none.
    """
    p = _profile_path(user_id)
    content = p.read_text(encoding="utf-8").strip() if p.exists() else ""
    return {"content": content}

put_profile async

put_profile(body: _ProfileBody, user_id: str = Depends(require_user)) -> dict

Replace the caller's profile markdown.

Parameters:

Name Type Description Default
body _ProfileBody

New profile content, replacing the previous one wholesale.

required
user_id str

Resolved by require_user.

Depends(require_user)

Returns:

Type Description
dict

{"ok": True} once written.

Source code in helioai/interfaces/web/app.py
@app.put("/api/profile")
async def put_profile(body: _ProfileBody, user_id: str = Depends(require_user)) -> dict:
    """Replace the caller's profile markdown.

    Args:
        body: New profile content, replacing the previous one wholesale.
        user_id: Resolved by `require_user`.

    Returns:
        `{"ok": True}` once written.
    """
    p = _profile_path(user_id)
    p.parent.mkdir(parents=True, exist_ok=True)
    p.write_text(body.content, encoding="utf-8")
    return {"ok": True}

delete_session async

delete_session(session_id: str, user_id: str = Depends(require_user)) -> dict

Delete one of the caller's sessions and its workspace.

Parameters:

Name Type Description Default
session_id str

Session to delete. Sanitised through safe_id before it reaches the rmtree behind this route.

required
user_id str

Resolved by require_user; only the owner's tree is touched.

Depends(require_user)

Returns:

Type Description
dict

{"deleted": session_id}, whether or not anything existed — a caller

dict

learns nothing about other users' session ids from the answer.

Source code in helioai/interfaces/web/app.py
@app.delete("/api/sessions/{session_id}")
async def delete_session(session_id: str, user_id: str = Depends(require_user)) -> dict:
    """Delete one of the caller's sessions and its workspace.

    Args:
        session_id: Session to delete. Sanitised through `safe_id` before it
            reaches the `rmtree` behind this route.
        user_id: Resolved by `require_user`; only the owner's tree is touched.

    Returns:
        `{"deleted": session_id}`, whether or not anything existed — a caller
        learns nothing about other users' session ids from the answer.
    """
    wdir = store.get_workspace_dir(user_id, session_id)
    store.reset(user_id, session_id)
    if wdir:
        # Containment, not trust: the label is persisted data, and a row written by
        # an older build (before session ids were sanitised) would walk this rmtree
        # straight out of the user's home.
        ws_root = (user_home(user_id) / "workspace").resolve()
        ws_path = (ws_root / wdir).resolve()
        if ws_path.is_relative_to(ws_root) and ws_path.exists():
            shutil.rmtree(ws_path, ignore_errors=True)
    return {"deleted": session_id}

export_notebook async

export_notebook(session_id: str, user_id: str = Depends(require_user)) -> FileResponse

Export a session as a standalone .ipynb and return it.

Parameters:

Name Type Description Default
session_id str

Session to export.

required
user_id str

Resolved by require_user.

Depends(require_user)

Returns:

Type Description
FileResponse

The notebook as a file download, built in memory rather than written to

FileResponse

the workspace.

Source code in helioai/interfaces/web/app.py
@app.get("/api/export")
async def export_notebook(session_id: str, user_id: str = Depends(require_user)) -> FileResponse:
    """Export a session as a standalone `.ipynb` and return it.

    Args:
        session_id: Session to export.
        user_id: Resolved by `require_user`.

    Returns:
        The notebook as a file download, built in memory rather than written to
        the workspace.
    """
    from helioai.export import export_session_notebook

    if session_id not in store.all_sessions(user_id):
        raise HTTPException(status_code=404, detail="Unknown session")
    path = export_session_notebook(user_id, session_id)
    return FileResponse(
        path,
        media_type="application/x-ipynb+json",
        filename=path.name,
    )

serve_code async

serve_code(path: str, full: bool = False, user_id: str = Depends(require_user)) -> PlainTextResponse

Return a generated script, rewritten to standalone form.

Ownership is checked against the caller before anything is read, so a path outside the caller's workspace is a 404 rather than a leak. A run_recipe run of an unmodified shipped recipe is shown as its own lines, the recipe read from the installed package (export.recipe_run_view) — six hundred lines of recipe buried the five the model wrote — and the X-HelioAI-Full-Lines header tells the panel that the whole script is one full=true away.

Parameters:

Name Type Description Default
path str

Absolute path of the generated script, as the artifact reported it.

required
full bool

Return the whole script even when a short recipe view exists.

False
user_id str

Resolved by require_user.

Depends(require_user)

Returns:

Type Description
PlainTextResponse

The script rewritten to standalone speasy calls.

Raises:

Type Description
HTTPException

404 for a path outside the caller's workspace or absent.

Source code in helioai/interfaces/web/app.py
@app.get("/code")
async def serve_code(
    path: str, full: bool = False, user_id: str = Depends(require_user)
) -> PlainTextResponse:
    """Return a generated script, rewritten to standalone form.

    Ownership is checked against the caller before anything is read, so a path
    outside the caller's workspace is a 404 rather than a leak. A `run_recipe` run of an
    unmodified shipped recipe is shown as its own lines, the recipe read from the
    installed package (`export.recipe_run_view`) — six hundred lines of recipe buried the
    five the model wrote — and the `X-HelioAI-Full-Lines` header tells the panel that the
    whole script is one `full=true` away.

    Args:
        path: Absolute path of the generated script, as the artifact reported it.
        full: Return the whole script even when a short recipe view exists.
        user_id: Resolved by `require_user`.

    Returns:
        The script rewritten to standalone speasy calls.

    Raises:
        HTTPException: 404 for a path outside the caller's workspace or absent.
    """
    path = _relocated(user_id, path.strip())
    if not is_under_workspace(path) or not _owns_path(user_id, path):
        log.warning("code_rejected", path=path, reason="outside workspace or not owner")
        raise HTTPException(status_code=404, detail="Not found")
    p = Path(path).resolve()
    if p.suffix != ".py" or not p.is_file():
        log.warning("code_rejected", path=path, reason="file not found or not .py")
        raise HTTPException(status_code=404, detail="Not found")
    from helioai.datastore import read_manifest
    from helioai.export import recipe_run_view, to_standalone

    manifest = read_manifest(p.parent)
    code = p.read_text(encoding="utf-8")
    short = None if full else recipe_run_view(code, manifest)
    if short is not None:
        n_lines = str(code.count("\n") + 1)
        return PlainTextResponse(short, headers={"X-HelioAI-Full-Lines": n_lines})
    return PlainTextResponse(to_standalone(code, manifest, with_header=True))

serve_figure async

serve_figure(path: str, user_id: str = Depends(require_user)) -> FileResponse

Serve a figure (PNG or PDF) from the caller's workspace.

Parameters:

Name Type Description Default
path str

Absolute path of the figure, as the artifact reported it.

required
user_id str

Resolved by require_user.

Depends(require_user)

Returns:

Type Description
FileResponse

The file, with a content type derived from its extension. Only PNG and

FileResponse

PDF are served, so a traversal that reached another file type still

FileResponse

returns nothing.

Raises:

Type Description
HTTPException

404 outside the caller's workspace, or absent.

Source code in helioai/interfaces/web/app.py
@app.get("/figure")
async def serve_figure(path: str, user_id: str = Depends(require_user)) -> FileResponse:
    """Serve a figure (PNG or PDF) from the caller's workspace.

    Args:
        path: Absolute path of the figure, as the artifact reported it.
        user_id: Resolved by `require_user`.

    Returns:
        The file, with a content type derived from its extension. Only PNG and
        PDF are served, so a traversal that reached another file type still
        returns nothing.

    Raises:
        HTTPException: 404 outside the caller's workspace, or absent.
    """
    path = _relocated(user_id, path.strip())
    if not is_under_workspace(path) or not _owns_path(user_id, path):
        log.warning("figure_rejected", path=path, reason="outside workspace or not owner")
        raise HTTPException(status_code=404, detail="Not found")
    p = Path(path).resolve()
    media_type = _FIGURE_TYPES.get(p.suffix.lower())
    if media_type is None:
        log.warning("figure_rejected", path=path, reason="unsupported type")
        raise HTTPException(status_code=404, detail="Not found")
    if not p.is_file():
        log.warning("figure_rejected", path=path, reason="file not found")
        raise HTTPException(status_code=404, detail="Not found")
    return FileResponse(p, media_type=media_type)

serve_web

serve_web(host: str = '127.0.0.1', port: int = 7890) -> None

Run the web UI with uvicorn.

Binds to localhost by default. The open-source build ships no authentication and run_python executes model-written code, so do not expose this on a network without putting auth in front of it.

Parameters:

Name Type Description Default
host str

Bind address. Anything but loopback exposes an arbitrary code executor; read SECURITY.md before changing it.

'127.0.0.1'
port int

TCP port.

7890
Source code in helioai/interfaces/web/app.py
def serve_web(host: str = "127.0.0.1", port: int = 7890) -> None:
    """Run the web UI with uvicorn.

    Binds to localhost by default. The open-source build ships no authentication
    and `run_python` executes model-written code, so do not expose this on a
    network without putting auth in front of it.

    Args:
        host: Bind address. Anything but loopback exposes an arbitrary code
            executor; read SECURITY.md before changing it.
        port: TCP port.
    """
    import uvicorn

    from helioai.workspace import cleanup_old_runs

    refuse_unauthenticated_public_bind(host)
    harden_for_host(app, host)
    cleanup_old_runs()
    uvicorn.run(app, host=host, port=port)

refuse_unauthenticated_public_bind

refuse_unauthenticated_public_bind(host: str) -> None

Exit rather than serve run_python to a network with no one authenticated.

The same rule helioai-mcp --http applies to a bind without a token: a public address with no HELIOAI_USERS is a deployment error, and a warning someone might read after the fact is not a boundary. HELIOAI_ALLOW_UNAUTHENTICATED_PUBLIC=1 is the explicit opt-out for a container that binds 0.0.0.0 behind a loopback publish.

Parameters:

Name Type Description Default
host str

The address about to be bound.

required

Raises:

Type Description
SystemExit

On a non-loopback host with neither users nor the opt-out.

Source code in helioai/interfaces/web/app.py
def refuse_unauthenticated_public_bind(host: str) -> None:
    """Exit rather than serve `run_python` to a network with no one authenticated.

    The same rule `helioai-mcp --http` applies to a bind without a token: a public
    address with no `HELIOAI_USERS` is a deployment error, and a warning someone might
    read after the fact is not a boundary. `HELIOAI_ALLOW_UNAUTHENTICATED_PUBLIC=1` is
    the explicit opt-out for a container that binds 0.0.0.0 behind a loopback publish.

    Args:
        host: The address about to be bound.

    Raises:
        SystemExit: On a non-loopback host with neither users nor the opt-out.
    """
    if host in _LOOPBACK_HOSTS or settings.web_auth.users:
        return
    if settings.web_auth.allow_unauthenticated_public:
        log.warning("web_public_unauthenticated_by_choice", host=host)
        return
    log.error(
        "web_refused_without_auth",
        host=host,
        detail=(
            "set HELIOAI_USERS, bind to loopback, or set "
            "HELIOAI_ALLOW_UNAUTHENTICATED_PUBLIC=1 behind a loopback port publish"
        ),
    )
    raise SystemExit(1)

MCP server

helioai.mcp_server

MCP server for HelioAI — exposes registered tools and read-only resources (recipes, skills) via stdio or HTTP streamable transport.

Usage

helioai serve # stdio (Claude Desktop / claude CLI) helioai serve --http # HTTP streamable on 127.0.0.1:8765 helioai serve --http --host 0.0.0.0 --port 9000 # requires HELIOAI_MCP_TOKEN helioai-mcp [--http ...] # direct entry point, same flags

Skills are listed from a process-lifetime-cached index (skills_loader._discover is lru_cache'd): a skill added or edited after this process started is invisible until restart. Recipes re-glob the filesystem on every call and need no restart.

serve_stdio async

serve_stdio() -> None

Run the MCP server over stdio, for clients like Claude Desktop.

Blocks until the client closes the pipe. All registry tools and the recipe/skill resources are exposed — over stdio the client owns the process, so no auth applies.

Source code in helioai/mcp_server.py
async def serve_stdio() -> None:
    """Run the MCP server over stdio, for clients like Claude Desktop.

    Blocks until the client closes the pipe. All registry tools and the recipe/skill
    resources are exposed — over stdio the client owns the process, so no auth applies.
    """
    from mcp.server.stdio import stdio_server

    async with stdio_server() as (read, write):
        await server.run(read, write, _init_options())

build_http_app

build_http_app()

Build the streamable-HTTP ASGI app exposing the MCP server.

Returns a Starlette app mounting the MCP session manager at /mcp, wrapped in a Bearer-token check when HELIOAI_MCP_TOKEN is set, suitable for any ASGI server (serve_http wraps it in uvicorn).

Source code in helioai/mcp_server.py
def build_http_app():
    """Build the streamable-HTTP ASGI app exposing the MCP server.

    Returns a Starlette app mounting the MCP session manager at `/mcp`, wrapped in a
    Bearer-token check when `HELIOAI_MCP_TOKEN` is set, suitable for any ASGI server
    (`serve_http` wraps it in uvicorn).
    """
    from mcp.server.streamable_http_manager import StreamableHTTPSessionManager
    from starlette.applications import Starlette
    from starlette.routing import Mount

    manager = StreamableHTTPSessionManager(app=server, json_response=False, stateless=False)

    @contextlib.asynccontextmanager
    async def lifespan(app):
        async with manager.run():
            yield

    app = Starlette(routes=[Mount("/mcp", app=manager.handle_request)], lifespan=lifespan)
    return _require_bearer_token(app, settings.mcp.token)

serve_http

serve_http(host: str, port: int) -> None

Run the MCP server over streamable HTTP.

Parameters:

Name Type Description Default
host str

Bind address.

required
port int

TCP port.

required
Source code in helioai/mcp_server.py
def serve_http(host: str, port: int) -> None:
    """Run the MCP server over streamable HTTP.

    Args:
        host: Bind address.
        port: TCP port.
    """
    import uvicorn

    uvicorn.run(build_http_app(), host=host, port=port)

main

main() -> None

Entry point for the helioai-mcp command.

--help and --version are answered before anything else: an MCP server on stdio reads the terminal as its protocol stream, so the reflex helioai-mcp --help used to start a server that sat waiting for JSON-RPC, with nothing on screen to say so.

Example

helioai-mcp # stdio (Claude Desktop, claude CLI) helioai-mcp --http --port 8765 # streamable HTTP on 127.0.0.1:8765

Source code in helioai/mcp_server.py
def main() -> None:
    """Entry point for the `helioai-mcp` command.

    `--help` and `--version` are answered before anything else: an MCP server on stdio
    reads the terminal as its protocol stream, so the reflex `helioai-mcp --help` used
    to start a server that sat waiting for JSON-RPC, with nothing on screen to say so.

    Example:
        helioai-mcp                        # stdio (Claude Desktop, claude CLI)
        helioai-mcp --http --port 8765     # streamable HTTP on 127.0.0.1:8765
    """
    args = sys.argv[1:]
    if {"-h", "--help"} & set(args):
        print(_USAGE)
        return
    if {"-V", "--version"} & set(args):
        from helioai import __version__

        print(f"helioai-mcp {__version__}")
        return
    setup_logging("WARNING")
    from helioai.workspace import cleanup_old_runs

    cleanup_old_runs()
    if "--http" in args:
        host = _arg(args, "--host", "127.0.0.1")
        port = int(_arg(args, "--port", "8765"))
        if host not in _LOOPBACK_HOSTS and not settings.mcp.token:
            # run_python is arbitrary code execution. Binding off loopback with no
            # token is a deployment error, not a warning someone might read after
            # the fact — refuse to start instead of trusting that.
            get_logger(__name__).error(
                "mcp_http_refused_without_auth",
                host=host,
                port=port,
                detail="set HELIOAI_MCP_TOKEN or bind to loopback",
            )
            raise SystemExit(1)
        serve_http(host, port)
    else:
        asyncio.run(serve_stdio())

Indexer

helioai.indexer

Build the speasy catalog ChromaDB index.

Usage

helioai index # incremental (skip existing) helioai index --rebuild # wipe and rebuild

MEASUREMENT_TYPES module-attribute

MEASUREMENT_TYPES: tuple[str, ...] = ('MagneticField', 'ElectricField', 'ThermalPlasma', 'EnergeticParticles', 'IonComposition', 'Ephemeris', 'Waves', 'Spectrum', 'NeutralGas', 'InstrumentStatus', 'Irradiance', 'Radiance')

The SPASE MeasurementType vocabulary the index already carries (AMDA, CSA), as the closed set a classifier chooses from — so a filled field is usable as an exact filter.

REGIONS module-attribute

REGIONS: tuple[str, ...] = ('Sun', 'Sun.Corona', 'Heliosphere', 'Heliosphere.Inner', 'Heliosphere.NearEarth', 'Heliosphere.Remote1AU', 'Heliosphere.Outer', 'Earth', 'Earth.Magnetosphere', 'Earth.Magnetosheath', 'Earth.Magnetosphere.Polar', 'Earth.Magnetosphere.Magnetotail', 'Earth.Magnetosphere.RadiationBelt', 'Earth.NearSurface', 'Earth.NearSurface.Ionosphere', 'Earth.NearSurface.AuroralRegion', 'Earth.NearSurface.EquatorialRegion', 'Earth.NearSurface.PolarCap', 'Mercury', 'Venus', 'Mars', 'Jupiter', 'Jupiter.Io', 'Jupiter.Europa', 'Jupiter.Ganymede', 'Jupiter.Callisto', 'Saturn', 'Saturn.Enceladus', 'Uranus', 'Neptune', 'Pluto', 'Comet')

The SPASE Region vocabulary AMDA publishes as dataset targets (30 values on 8 435 products) plus the two the indexer's table uses and AMDA does not — the closed set a classifier chooses from.

SHIPPED_JUDGED module-attribute

SHIPPED_JUDGED = Path(__file__).parent / 'data' / JUDGED_FILE

Every question the judge has been asked about a product, with its answer, shipped with the package: one gzipped JSON line per product (id, name, then one {choice, confidence} per question asked — mtype, region), after a first line of provenance (meta: date, models, floors, count). It is the part of the index that cannot be rebuilt from code — 82 266 requests, US$ 2.4 on 2026-09-22 — kept as the judge's raw answers, not as decided fields, so the policy (floors, never overwriting a published label) lives in code and can change without asking again. Abstentions are in it too: a question already asked is not paid for twice. helioai index applies it to every product the archive left untyped; --classify asks only what no record answers and appends to the local copy.

build_index

build_index(rebuild: bool = False, batch_size: int = 128, verbose: bool = True, classify: bool = False) -> int

Walk the speasy inventory and index all parameters into ChromaDB.

Backs helioai index and must run once before search_parameters works; the index persists under settings.rag.chroma_dir.

Parameters:

Name Type Description Default
rebuild bool

Drop and re-create the collection instead of appending.

False
batch_size int

Documents per ChromaDB insert.

128
verbose bool

Print per-provider progress to stdout.

True
classify bool

Ask the judgment backend, before embedding, for the SPASE measurement type of every product the archive leaves untyped and for the SPASE region of every product whose region is the table's guess (--classify; needs HELIOAI_JUDGMENT_BACKEND=jev). See classify_products.

False

Returns:

Type Description
int

Number of parameters indexed (0 when speasy or chromadb is missing).

Example

build_index(rebuild=True) # equivalent to: helioai index --rebuild 82433

Source code in helioai/indexer.py
def build_index(
    rebuild: bool = False, batch_size: int = 128, verbose: bool = True, classify: bool = False
) -> int:
    """Walk the speasy inventory and index all parameters into ChromaDB.

    Backs `helioai index` and must run once before `search_parameters` works;
    the index persists under `settings.rag.chroma_dir`.

    Args:
        rebuild: Drop and re-create the collection instead of appending.
        batch_size: Documents per ChromaDB insert.
        verbose: Print per-provider progress to stdout.
        classify: Ask the judgment backend, before embedding, for the SPASE measurement
            type of every product the archive leaves untyped and for the SPASE region of
            every product whose region is the table's guess (`--classify`; needs
            `HELIOAI_JUDGMENT_BACKEND=jev`). See `classify_products`.

    Returns:
        Number of parameters indexed (0 when speasy or chromadb is missing).

    Example:
        >>> build_index(rebuild=True)   # equivalent to: helioai index --rebuild
        82433
    """
    try:
        import chromadb  # noqa: F401 — the index cannot be built without it; opened in open_collections
        import speasy as spz
        from sentence_transformers import SentenceTransformer
        from speasy.core.inventory.indexes import SpeasyIndex
    except ImportError as e:
        print(f"[indexer] Missing dependency: {e}")
        print("[indexer] Run: pip install speasy chromadb sentence-transformers")
        return 0

    from helioai.config import settings

    chroma_dir = settings.rag.chroma_dir
    collection_name = settings.rag.collection_name
    embed_model = settings.rag.embed_model

    if rebuild and chroma_dir.exists():
        records = chroma_dir / JUDGMENT_RECORDS
        kept = records.read_bytes() if records.exists() else None
        if verbose:
            print(f"[indexer] wiping {chroma_dir}")
        shutil.rmtree(chroma_dir)
        if kept is not None:
            chroma_dir.mkdir(parents=True, exist_ok=True)
            records.write_bytes(kept)

    chroma_dir.mkdir(parents=True, exist_ok=True)
    judged_meta, judged = load_judged(SHIPPED_JUDGED, local_judged_path())
    if verbose and judged:
        print(
            f"[indexer] {len(judged)} products already put to the judge "
            f"(snapshot {judged_meta.get('date', '?')}, {', '.join(judged_meta.get('models') or [])})"
        )

    if verbose:
        print(f"[indexer] loading embedding model {embed_model}…")
    model = SentenceTransformer(embed_model)

    _, (collection, catalog_collection) = open_collections(
        chroma_dir, [collection_name, settings.rag.catalogs_collection_name], verbose=verbose
    )

    existing_ids: set[str] = set()
    if not rebuild:
        try:
            existing_ids = set(collection.get(include=[])["ids"])
            if verbose and existing_ids:
                print(f"[indexer] {len(existing_ids)} existing entries — skipping")
        except Exception:
            pass

    if verbose:
        print("[indexer] walking speasy inventory…")

    docs: list[dict] = []
    tree = spz.inventories.tree

    for provider_attr, prefix in _PROVIDER_PREFIXES.items():
        provider_node = getattr(tree, provider_attr, None)
        if provider_node is None:
            continue
        before = len(docs)
        _walk(provider_node, prefix, docs, existing_ids, SpeasyIndex)
        if verbose:
            print(f"[indexer]   {prefix}: {len(docs) - before} new params")

    if verbose:
        print(f"[indexer] total new params to index: {len(docs)}")

    if not docs:
        if verbose:
            print("[indexer] up to date — nothing to index")
        return 0

    if judged:
        applied = apply_judged(docs, judged)
        if verbose:
            print(
                f"[indexer] the judge's recorded answers typed {applied} products, no request made"
            )
    if classify:
        import asyncio

        docs = asyncio.run(
            classify_products(
                docs, chroma_dir, judged=judged, judged_meta=judged_meta, verbose=verbose
            )
        )

    t0 = time.perf_counter()
    total = 0

    for i in range(0, len(docs), batch_size):
        batch = docs[i : i + batch_size]
        ids = [d["id"] for d in batch]
        texts = [d["text"] for d in batch]
        metas = [d["meta"] for d in batch]

        embeddings = model.encode(
            texts,
            batch_size=batch_size,
            show_progress_bar=False,
            convert_to_numpy=True,
            normalize_embeddings=True,
        ).tolist()

        collection.upsert(ids=ids, embeddings=embeddings, documents=texts, metadatas=metas)
        total += len(ids)
        if verbose:
            print(f"[indexer]   {total}/{len(docs)} indexed…", end="\r", flush=True)

    elapsed = time.perf_counter() - t0
    if verbose:
        print()
        print(f"[indexer] done: {total} params in {elapsed:.1f}s")
        print(f"[indexer] collection total: {collection.count()}")

    # Index catalogs + timetables into a separate collection
    cat_total = _build_catalog_index(
        model, catalog_collection, settings, rebuild=rebuild, verbose=verbose
    )

    return total + cat_total

local_judged_path

local_judged_path() -> Path

Where --classify writes the answers it obtains: beside the data, not the index.

The Chroma directory is wiped by --rebuild; the data root is not. A user with a key who classifies a new provider keeps those answers across every rebuild, and they take precedence over the shipped file for the same id.

Source code in helioai/indexer.py
def local_judged_path() -> Path:
    """Where `--classify` writes the answers it obtains: beside the data, not the index.

    The Chroma directory is wiped by `--rebuild`; the data root is not. A user with a key
    who classifies a new provider keeps those answers across every rebuild, and they take
    precedence over the shipped file for the same id.
    """
    from helioai.config import settings

    return Path(settings.data_dir) / JUDGED_FILE

load_judged

load_judged(*paths: Path) -> tuple[dict, dict[str, dict]]

Read the judge's recorded answers from each file in turn, later files overriding earlier ones id by id; a missing or unreadable file contributes nothing.

Returns:

Type Description
dict

(meta, records) — the provenance of the last file read that had any, and

dict[str, dict]

{id: {"name": …, "mtype": {"choice", "confidence"}, "region": {…}}}.

Source code in helioai/indexer.py
def load_judged(*paths: Path) -> tuple[dict, dict[str, dict]]:
    """Read the judge's recorded answers from each file in turn, later files overriding
    earlier ones id by id; a missing or unreadable file contributes nothing.

    Returns:
        `(meta, records)` — the provenance of the last file read that had any, and
        `{id: {"name": …, "mtype": {"choice", "confidence"}, "region": {…}}}`.
    """
    import gzip
    import json

    meta: dict = {}
    records: dict[str, dict] = {}
    for path in paths:
        try:
            with gzip.open(path, "rt", encoding="utf-8") as f:
                for line in f:
                    if not line.strip():
                        continue
                    row = json.loads(line)
                    if "meta" in row:
                        meta = dict(row["meta"])
                        continue
                    pid = row.get("id")
                    if pid:
                        records[pid] = {k: v for k, v in row.items() if k != "id"}
        except (OSError, ValueError):
            continue
    return meta, records

save_judged

save_judged(path: Path, meta: dict, records: dict[str, dict]) -> None

Write the answers as load_judged reads them, sorted by id, provenance first.

Source code in helioai/indexer.py
def save_judged(path: Path, meta: dict, records: dict[str, dict]) -> None:
    """Write the answers as `load_judged` reads them, sorted by id, provenance first."""
    import gzip
    import json

    path.parent.mkdir(parents=True, exist_ok=True)
    meta = {**meta, "asked": len(records)}
    with gzip.open(path, "wt", encoding="utf-8", compresslevel=9) as f:
        f.write(json.dumps({"meta": meta}, ensure_ascii=False) + "\n")
        for pid in sorted(records):
            f.write(json.dumps({"id": pid, **records[pid]}, ensure_ascii=False) + "\n")

apply_judged

apply_judged(docs: list[dict], judged: dict[str, dict]) -> int

Apply the recorded answers to every walked doc they name.

A record whose name no longer matches the product's is skipped: the id was reused for something else, and a type decided about the old content is not evidence about the new. A record without a name (none of the first pass had one) is applied.

Returns:

Type Description
int

How many docs received at least one field.

Source code in helioai/indexer.py
def apply_judged(docs: list[dict], judged: dict[str, dict]) -> int:
    """Apply the recorded answers to every walked doc they name.

    A record whose `name` no longer matches the product's is skipped: the id was reused
    for something else, and a type decided about the old content is not evidence about
    the new. A record without a name (none of the first pass had one) is applied.

    Returns:
        How many docs received at least one field.
    """
    applied = 0
    for doc in docs:
        record = judged.get(doc["id"])
        if not record:
            continue
        name = record.get("name")
        if name and doc["meta"].get("name") and name != doc["meta"]["name"]:
            continue
        applied += _apply_record(doc, record)
    return applied

classify_products async

classify_products(docs: list[dict], record_dir: Path | str, *, judged: dict[str, dict] | None = None, judged_meta: dict | None = None, verbose: bool = False) -> list[dict]

Ask the judge what no record has answered yet, apply it, and keep the answers.

Two closed questions per product: the SPASE measurement type where the archive left it empty, and the SPASE region where the indexer had only guessed. Neither answer ever overwrites anything the archive published.

Measurement type. The field is indexed on 15.6 % of the products — AMDA and CSA — and on none of CDA's 68 000, so every ranking signal built on it (_rerank_penalty) and every filter reaches a sixth of the catalogue. Measured on 2026-09-22 against 200 products the archive had labelled, label stripped from the text before asking: 76.5 % agreement, 89 % where the judge's confidence is at least 0.9 (72 % of the items) — and the remaining confident disagreements were the archive's errors (MMS FPI plasma moments labelled MagneticField, a JADE density labelled EnergeticParticles, a Langmuir-probe density labelled ElectricField). So the judge is better than its ground truth, and the floor is 0.9: below it the field stays empty — abstention is a type — and a published label the judge contradicts at or above it is kept and flagged as measurement_type_jev, for a person to adjudicate, never replaced.

Region. _get_region guesses from a 40-entry table matched as a substring; against AMDA's 8 435 published targets it agrees on 26.9 %, is silent on 41 % and wrong on 32 % ("ac" inside "cce_mepa_ion_act" made AMPTE/CCE a near-Earth heliospheric product). Measured the same day on 200 of those products, target stripped: the judge agrees exactly on 70 %, on the body (Earth, Jupiter, Heliosphere…) on 97.1 % at confidence ≥ 0.9, and where judge and table differ the judge is right 75 times to the table's one. Its confident disagreements with the archive are granularity, in both directions (Helios filed as Heliosphere, a Galileo Io flyby read as Jupiter), so a published target is never flagged — it stands. The table's guess is not a publication: at or above the floor the judge's region replaces it, or fills the silence, with region_source: "jev" and the confidence; below it the guess stays, marked as the guess it is.

What is asked. A product is asked only the questions no record in judged answers for it — the shipped file plus the local one — so a second pass over the same catalogue costs nothing, a new provider costs its own products, and adding a question costs one request per product for that question alone. Both sentences the text carried are stripped before asking, so the judge reads the product, not the labels. Every call is recorded to judgment_index.jsonl in the index directory with the product id as its key, and every answer — abstentions included — is appended to the local judged_products.jsonl.gz for the next rebuild. ~2.9 ¢ per 1 000 requests (metered 2026-09-22: 838 tokens a request, the instruction being most of it).

Parameters:

Name Type Description Default
docs list[dict]

{id, text, meta} as _walk collects them.

required
record_dir Path | str

Where the calls are recorded (the Chroma directory).

required
judged dict[str, dict] | None

The answers already on record, updated in place.

None
judged_meta dict | None

Their provenance, carried into the saved file.

None
verbose bool

Print the counts.

False

Returns:

Type Description
list[dict]

The same docs, metadata and text amended in place.

Source code in helioai/indexer.py
async def classify_products(
    docs: list[dict],
    record_dir: Path | str,
    *,
    judged: dict[str, dict] | None = None,
    judged_meta: dict | None = None,
    verbose: bool = False,
) -> list[dict]:
    """Ask the judge what no record has answered yet, apply it, and keep the answers.

    Two closed questions per product: the SPASE measurement type where the archive left it
    empty, and the SPASE region where the indexer had only guessed. Neither answer ever
    overwrites anything the archive published.

    **Measurement type.** The field is indexed on 15.6 % of the products — AMDA and CSA —
    and on none of CDA's 68 000, so every ranking signal built on it (`_rerank_penalty`)
    and every filter reaches a sixth of the catalogue. Measured on 2026-09-22 against 200
    products the archive had labelled, label stripped from the text before asking: 76.5 %
    agreement, **89 % where the judge's confidence is at least 0.9** (72 % of the items) —
    and the remaining confident disagreements were the archive's errors (MMS FPI plasma
    moments labelled MagneticField, a JADE density labelled EnergeticParticles, a
    Langmuir-probe density labelled ElectricField). So the judge is better than its ground
    truth, and the floor is 0.9: below it the field stays empty — abstention is a type —
    and a published label the judge contradicts at or above it is kept and flagged as
    `measurement_type_jev`, for a person to adjudicate, never replaced.

    **Region.** `_get_region` guesses from a 40-entry table matched as a substring; against
    AMDA's 8 435 published targets it agrees on 26.9 %, is silent on 41 % and wrong on 32 %
    ("ac" inside "cce_mepa_ion_act" made AMPTE/CCE a near-Earth heliospheric product).
    Measured the same day on 200 of those products, target stripped: the judge agrees
    exactly on 70 %, **on the body (Earth, Jupiter, Heliosphere…) on 97.1 % at confidence
    ≥ 0.9**, and where judge and table differ the judge is right 75 times to the table's
    one. Its confident disagreements with the archive are granularity, in both directions
    (Helios filed as Heliosphere, a Galileo Io flyby read as Jupiter), so a published target
    is never flagged — it stands. The table's guess is not a publication: at or above the
    floor the judge's region replaces it, or fills the silence, with `region_source: "jev"`
    and the confidence; below it the guess stays, marked as the guess it is.

    **What is asked.** A product is asked only the questions no record in `judged` answers
    for it — the shipped file plus the local one — so a second pass over the same
    catalogue costs nothing, a new provider costs its own products, and adding a question
    costs one request per product for that question alone. Both sentences the text carried
    are stripped before asking, so the judge reads the product, not the labels. Every call
    is recorded to `judgment_index.jsonl` in the index directory with the product id as its
    key, and every answer — abstentions included — is appended to the local
    `judged_products.jsonl.gz` for the next rebuild. ~2.9 ¢ per 1 000 requests (metered
    2026-09-22: 838 tokens a request, the instruction being most of it).

    Args:
        docs: `{id, text, meta}` as `_walk` collects them.
        record_dir: Where the calls are recorded (the Chroma directory).
        judged: The answers already on record, updated in place.
        judged_meta: Their provenance, carried into the saved file.
        verbose: Print the counts.

    Returns:
        The same docs, metadata and text amended in place.
    """
    from helioai.config import settings
    from helioai.core import judgment

    judged = {} if judged is None else judged
    if settings.judgment.backend == "null":
        if verbose:
            print(
                "[indexer] --classify needs HELIOAI_JUDGMENT_BACKEND=jev and TYPESAFE_API_KEY; "
                "skipping classification"
            )
        return docs
    questions = {
        "mtype": judgment.Choice(
            MEASUREMENT_TYPE_INSTRUCTIONS, MEASUREMENT_TYPES, floor=MEASUREMENT_TYPE_FLOOR
        ),
        "region": judgment.Choice(REGION_INSTRUCTIONS, REGIONS, floor=REGION_FLOOR),
    }
    groups: dict[tuple[str, ...], list[dict]] = {}
    for doc in docs:
        record = judged.get(doc["id"]) or {}
        missing = tuple(q for q in questions if q not in record)
        if missing:
            groups.setdefault(missing, []).append(doc)
    n_requests = sum(len(g) for g in groups.values())
    if verbose:
        print(
            f"[indexer] {len(docs) - n_requests} products fully on record; "
            f"asking {n_requests} requests: "
            + ", ".join(f"{len(g)} × {'+'.join(k)}" for k, g in groups.items())
        )
    models: set[str] = set(judged_meta.get("models") or []) if judged_meta else set()
    filled = flagged = 0
    for missing, group in groups.items():
        asked = {q: questions[q] for q in missing}
        states = [
            {"product": _REGION_SENTENCE.sub("", _LABEL_SENTENCE.sub(" ", d["text"])).strip()}
            for d in group
        ]
        answers = await judgment.batch(
            "index_classify",
            states,
            asked,
            record_to=Path(record_dir) / JUDGMENT_RECORDS,
            keys=[d["id"] for d in group],
        )
        for doc, answer in zip(group, answers, strict=True):
            if answer is None:
                continue
            record = judged.setdefault(doc["id"], {})
            record["name"] = doc["meta"].get("name") or record.get("name")
            for q in missing:
                raw = answer.raw.get(q) or {}
                if raw.get("choice") is not None:
                    record[q] = {
                        "choice": raw["choice"],
                        "confidence": round(float(raw.get("confidence") or 0.0), 3),
                    }
            models.add(answer.model)
            before = doc["meta"].get("measurement_type")
            if _apply_record(doc, record):
                if doc["meta"].get("measurement_type_source") == "jev" and not before:
                    filled += 1
                if "measurement_type_jev" in doc["meta"]:
                    flagged += 1
    if n_requests:
        meta = {
            **(judged_meta or {}),
            "date": time.strftime("%Y-%m-%d"),
            "models": sorted(models),
            "floors": {"mtype": MEASUREMENT_TYPE_FLOOR, "region": REGION_FLOOR},
        }
        save_judged(local_judged_path(), meta, judged)
    if verbose:
        print(
            f"[indexer] this pass: {filled} types filled, {flagged} published labels flagged, "
            f"{n_requests} requests, answers saved to {local_judged_path()}"
        )
    return docs

open_collections

open_collections(chroma_dir: Path | str, names: list[str], *, verbose: bool = False) -> tuple[Any, list[Any]]

Open or create the index's collections so that every write is persisted at once, and so that the dense search looks wide enough to find a near-twin.

Chroma's local HNSW segment persists to disk only every sync_threshold writes — 1000 by default. Whatever follows the last persist stays in the write-ahead log and is replayed into the in-memory graph at every process start, in an order that varies, so the graph varies and the ranking with it. Measured on 2026-09-22: the 325 SSCWeb trajectories added after the last persist gave five different dense top-50 lists in five processes for one query embedding, ssc/mms1 at rank 1 in four of them and absent from the fifth. The catalogue collection, 221 entries, had never been persisted at all. A threshold of one costs 0.11 s per batch on the full 82k index and leaves nothing to replay, so a search ranks the same in every process and a read-only process never writes to the index.

ef_search is raised from Chroma's 100 to 400. The catalogue is full of near-twins — 314 SSCWeb trajectories that differ by a spacecraft name, hundreds of housekeeping variables that differ by a suffix — and an approximate search with a narrow beam loses the exact twin: ssc/mms1, the true nearest neighbour of "MMS1 spacecraft position GSE 2019", was absent from the dense top-50 at 100 and is rank 1 at 400, for 0.7 → 1.2 ms per query (measured 2026-09-22). On the 30 HelioBench n1 queries the change moved recall@1 from 53.3 % to 56.7 %.

A collection created before these settings keeps the ones it was loaded with, so it is modified and the client reopened: the replay on reopen persists its tail. Reopening clears Chroma's process-wide client cache — fine in helioai index, and the reason this is not done lazily by a process that also serves searches.

Parameters:

Name Type Description Default
chroma_dir Path | str

The Chroma directory, created when absent.

required
names list[str]

Collection names, opened in order.

required
verbose bool

Print when a legacy collection is settled.

False

Returns:

Type Description
tuple[Any, list[Any]]

(client, collections) — the client the collections belong to.

Source code in helioai/indexer.py
def open_collections(
    chroma_dir: Path | str, names: list[str], *, verbose: bool = False
) -> tuple[Any, list[Any]]:
    """Open or create the index's collections so that every write is persisted at once,
    and so that the dense search looks wide enough to find a near-twin.

    Chroma's local HNSW segment persists to disk only every `sync_threshold` writes — 1000
    by default. Whatever follows the last persist stays in the write-ahead log and is
    replayed into the in-memory graph at every process start, in an order that varies, so
    the graph varies and the ranking with it. Measured on 2026-09-22: the 325 SSCWeb
    trajectories added after the last persist gave five different dense top-50 lists in five
    processes for one query embedding, `ssc/mms1` at rank 1 in four of them and absent from
    the fifth. The catalogue collection, 221 entries, had never been persisted at all. A
    threshold of one costs 0.11 s per batch on the full 82k index and leaves nothing to
    replay, so a search ranks the same in every process and a read-only process never
    writes to the index.

    `ef_search` is raised from Chroma's 100 to 400. The catalogue is full of near-twins —
    314 SSCWeb trajectories that differ by a spacecraft name, hundreds of housekeeping
    variables that differ by a suffix — and an approximate search with a narrow beam loses
    the exact twin: `ssc/mms1`, the true nearest neighbour of "MMS1 spacecraft position GSE
    2019", was absent from the dense top-50 at 100 and is rank 1 at 400, for 0.7 → 1.2 ms per
    query (measured 2026-09-22). On the 30 HelioBench n1 queries the change moved recall@1
    from 53.3 % to 56.7 %.

    A collection created before these settings keeps the ones it was loaded with, so it is
    modified and the client reopened: the replay on reopen persists its tail. Reopening
    clears Chroma's process-wide client cache — fine in `helioai index`, and the reason this
    is not done lazily by a process that also serves searches.

    Args:
        chroma_dir: The Chroma directory, created when absent.
        names: Collection names, opened in order.
        verbose: Print when a legacy collection is settled.

    Returns:
        `(client, collections)` — the client the collections belong to.
    """
    import chromadb
    from chromadb.api.client import SharedSystemClient

    wanted = {"sync_threshold": HNSW_SYNC_THRESHOLD, "ef_search": HNSW_EF_SEARCH}
    client = chromadb.PersistentClient(path=str(chroma_dir))
    collections = [
        client.get_or_create_collection(
            name=n, configuration={"hnsw": {"space": "cosine", **wanted}}
        )
        for n in names
    ]
    legacy = [
        c
        for c in collections
        if any((c.configuration_json.get("hnsw") or {}).get(k) != v for k, v in wanted.items())
    ]
    if not legacy:
        return client, collections
    for c in legacy:
        c.modify(configuration={"hnsw": wanted})
    SharedSystemClient.clear_system_cache()
    client = chromadb.PersistentClient(path=str(chroma_dir))
    collections = [client.get_collection(n) for n in names]
    for c in collections:
        c.count()
    if verbose:
        print(f"[indexer] settled {len(legacy)} collection(s): pending writes persisted")
    return client, collections

Index snapshots

helioai.index_snapshot

Share a built index: export it to files, fetch a published one, import it.

Building the index walks the speasy inventory and embeds ~83 000 products — 7 to 10 minutes on a recent machine, longer on a modest one, before a researcher can ask a first question. For a given catalogue and embedding model the result is the same on every machine, so CI builds it once per release (.github/workflows/index.yml) and publishes a snapshot on the Hugging Face Hub; helioai index on an empty index fetches that snapshot instead of building.

A snapshot is the content of the collections — ids, documents, metadata, embeddings — and not Chroma's directory: a store written by one Chroma release is not guaranteed to open in an older one, and chromadb carries no upper bound. Importing re-creates the collections through open_collections, with the HNSW settings a local build uses.

index_is_empty

index_is_empty(chroma_dir: Path | None = None) -> bool

Whether the product collection is absent or holds nothing.

This is the condition under which helioai index fetches the published snapshot rather than walking the inventory, and a fetch replaces the index. So a store that exists but cannot be read — corrupt, or written by a Chroma this one cannot open — is not empty: helioai index then builds on it and fails loudly, rather than discarding products a user built or classified.

Source code in helioai/index_snapshot.py
def index_is_empty(chroma_dir: Path | None = None) -> bool:
    """Whether the product collection is absent or holds nothing.

    This is the condition under which `helioai index` fetches the published snapshot
    rather than walking the inventory, and a fetch replaces the index. So a store that
    exists but cannot be read — corrupt, or written by a Chroma this one cannot open — is
    not empty: `helioai index` then builds on it and fails loudly, rather than discarding
    products a user built or classified.
    """
    from helioai.config import settings

    chroma_dir = Path(chroma_dir or settings.rag.chroma_dir)
    if not (chroma_dir / "chroma.sqlite3").exists():
        return True
    import chromadb

    name = settings.rag.collection_name
    try:
        client = chromadb.PersistentClient(path=str(chroma_dir))
        listed = {c if isinstance(c, str) else c.name for c in client.list_collections()}
        return name not in listed or client.get_collection(name).count() == 0
    except Exception:
        return False
    finally:
        _release_clients()

export_index

export_index(out_dir: Path, chroma_dir: Path | None = None) -> dict

Write the index as a snapshot: per collection a gzipped JSONL and a float32 .npy.

The records keep Chroma's order, so an import inserts in the order the build did. The manifest carries a SHA-256 per file — the import refuses a truncated download rather than serving a partial index — and the embedding model, since vectors from another model would be silently meaningless against this install's query embeddings.

Parameters:

Name Type Description Default
out_dir Path

Destination directory, created when absent.

required
chroma_dir Path | None

The index to export; settings.rag.chroma_dir by default.

None

Returns:

Type Description
dict

The manifest written to out_dir/manifest.json.

Raises:

Type Description
FileNotFoundError

No index at chroma_dir.

ValueError

The product collection is empty.

Source code in helioai/index_snapshot.py
def export_index(out_dir: Path, chroma_dir: Path | None = None) -> dict:
    """Write the index as a snapshot: per collection a gzipped JSONL and a float32 `.npy`.

    The records keep Chroma's order, so an import inserts in the order the build did. The
    manifest carries a SHA-256 per file — the import refuses a truncated download rather
    than serving a partial index — and the embedding model, since vectors from another
    model would be silently meaningless against this install's query embeddings.

    Args:
        out_dir: Destination directory, created when absent.
        chroma_dir: The index to export; `settings.rag.chroma_dir` by default.

    Returns:
        The manifest written to `out_dir/manifest.json`.

    Raises:
        FileNotFoundError: No index at `chroma_dir`.
        ValueError: The product collection is empty.
    """
    import chromadb

    from helioai import __version__
    from helioai.config import settings

    chroma_dir = Path(chroma_dir or settings.rag.chroma_dir)
    if not (chroma_dir / "chroma.sqlite3").exists():
        raise FileNotFoundError(f"no index at {chroma_dir}")
    out_dir = Path(out_dir)
    out_dir.mkdir(parents=True, exist_ok=True)

    client = chromadb.PersistentClient(path=str(chroma_dir))
    collections: dict[str, dict] = {}
    try:
        for role, name in _collection_names().items():
            try:
                collection = client.get_collection(name)
            except Exception:
                continue
            records = out_dir / f"{role}.jsonl.gz"
            vectors: list[np.ndarray] = []
            count = 0
            with (
                records.open("wb") as raw,
                gzip.GzipFile(fileobj=raw, mode="wb", mtime=0) as gz,
            ):
                for offset in range(0, collection.count(), _PAGE):
                    page = collection.get(
                        include=["documents", "metadatas", "embeddings"],
                        limit=_PAGE,
                        offset=offset,
                    )
                    for pid, doc, meta in zip(
                        page["ids"], page["documents"], page["metadatas"], strict=True
                    ):
                        line = {"id": pid, "document": doc, "metadata": meta}
                        gz.write((json.dumps(line, ensure_ascii=False) + "\n").encode())
                    vectors.append(np.asarray(page["embeddings"], dtype=np.float32))
                    count += len(page["ids"])
            matrix = np.concatenate(vectors) if vectors else np.empty((0, 0), np.float32)
            embeddings = out_dir / f"{role}.npy"
            np.save(embeddings, matrix)
            collections[role] = {
                "count": count,
                "dim": int(matrix.shape[1]),
                "files": {p.name: _sha256(p) for p in (records, embeddings)},
            }
    finally:
        _release_clients()

    if not collections.get("products", {}).get("count"):
        raise ValueError(f"the product collection at {chroma_dir} is empty: nothing to export")

    try:
        from importlib.metadata import version

        speasy_version = version("speasy")
    except Exception:
        speasy_version = None
    manifest = {
        "format": FORMAT,
        "helioai": __version__,
        "speasy": speasy_version,
        "embed_model": settings.rag.embed_model,
        "created": datetime.now(UTC).isoformat(timespec="seconds"),
        "collections": collections,
    }
    (out_dir / MANIFEST).write_text(json.dumps(manifest, indent=2) + "\n", encoding="utf-8")
    return manifest

import_index

import_index(snapshot_dir: Path, chroma_dir: Path | None = None, verbose: bool = True) -> int

Replace the index with a snapshot's content, and return the number of entries.

The collections are filled in a staging directory beside the index and swapped in only once complete: an interrupted import leaves the previous index — or none — and never a partial one that helioai index would then take for up to date and merely top up. The previous index is set aside as <index>.previous until the new one is in place, and put back if the swap fails — on Windows a store another process holds open cannot be renamed. A .previous already there, left by a crash, is never deleted: the import refuses instead. The judge's local answers (judgment_index.jsonl) are carried across the swap, as --rebuild does, since nothing but a paid request can recreate them.

Raises:

Type Description
ValueError

Unknown format, another embedding model, or a checksum mismatch.

Source code in helioai/index_snapshot.py
def import_index(snapshot_dir: Path, chroma_dir: Path | None = None, verbose: bool = True) -> int:
    """Replace the index with a snapshot's content, and return the number of entries.

    The collections are filled in a staging directory beside the index and swapped in
    only once complete: an interrupted import leaves the previous index — or none — and
    never a partial one that `helioai index` would then take for up to date and merely
    top up. The previous index is set aside as `<index>.previous` until the new one is in
    place, and put back if the swap fails — on Windows a store another process holds open
    cannot be renamed. A `.previous` already there, left by a crash, is never deleted:
    the import refuses instead. The judge's local answers (`judgment_index.jsonl`) are
    carried across the swap, as `--rebuild` does, since nothing but a paid request can
    recreate them.

    Raises:
        ValueError: Unknown format, another embedding model, or a checksum mismatch.
    """
    from helioai.config import settings
    from helioai.indexer import JUDGMENT_RECORDS, open_collections

    snapshot_dir = Path(snapshot_dir)
    chroma_dir = Path(chroma_dir or settings.rag.chroma_dir)
    manifest = _read_manifest(snapshot_dir)
    names = _collection_names()
    roles = [r for r in names if r in manifest["collections"]]

    staging = chroma_dir.with_name(chroma_dir.name + ".partial")
    if staging.exists():
        shutil.rmtree(staging)
    total = 0
    try:
        client, collections = open_collections(staging, [names[r] for r in roles])
        batch = min(client.get_max_batch_size(), _PAGE)
        for role, collection in zip(roles, collections, strict=True):
            with gzip.open(snapshot_dir / f"{role}.jsonl.gz", "rt", encoding="utf-8") as f:
                records = [json.loads(line) for line in f]
            vectors = np.load(snapshot_dir / f"{role}.npy")
            if (
                len(records) != len(vectors)
                or len(records) != manifest["collections"][role]["count"]
            ):
                raise ValueError(f"{role}: records, embeddings and manifest disagree on the count")
            for i in range(0, len(records), batch):
                chunk = records[i : i + batch]
                collection.upsert(
                    ids=[r["id"] for r in chunk],
                    documents=[r["document"] for r in chunk],
                    metadatas=[r["metadata"] or None for r in chunk],
                    embeddings=vectors[i : i + batch],
                )
                if verbose:
                    print(
                        f"[indexer]   {role}: {i + len(chunk)}/{len(records)}", end="\r", flush=True
                    )
            if verbose:
                print()
            total += len(records)
    except BaseException:
        _release_clients()
        _discard(staging)
        raise
    _release_clients()

    previous = chroma_dir.with_name(chroma_dir.name + ".previous")
    moved = False
    try:
        if chroma_dir.exists():
            kept = chroma_dir / JUDGMENT_RECORDS
            if kept.exists():
                shutil.copy2(kept, staging / JUDGMENT_RECORDS)
            _retry(chroma_dir.rename, previous)
            moved = True
        _retry(staging.rename, chroma_dir)
    except BaseException:
        if moved:
            _retry(previous.rename, chroma_dir)
        _discard(staging)
        raise
    if moved:
        _discard(previous)
    return total

download_index

download_index(repo: str, version: str | None = None) -> tuple[Path, str]

Fetch the snapshot published for this release, or the latest one when there is none.

CI tags each snapshot with the release it was built by (v0.4.0), so an installed release gets the index its own code describes. A development install has no tag of its own and gets main, the most recent snapshot: an index that at worst predates some describing change, which helioai index --rebuild catches up with. The Hub caches the files, so a second fetch of the same revision downloads nothing.

Returns:

Type Description
tuple[Path, str]

(local_dir, revision).

Source code in helioai/index_snapshot.py
def download_index(repo: str, version: str | None = None) -> tuple[Path, str]:
    """Fetch the snapshot published for this release, or the latest one when there is none.

    CI tags each snapshot with the release it was built by (`v0.4.0`), so an installed
    release gets the index its own code describes. A development install has no tag of
    its own and gets `main`, the most recent snapshot: an index that at worst predates
    some describing change, which `helioai index --rebuild` catches up with.
    The Hub caches the files, so a second fetch of the same revision downloads nothing.

    Returns:
        `(local_dir, revision)`.
    """
    from huggingface_hub import snapshot_download
    from huggingface_hub.errors import RevisionNotFoundError

    from helioai import __version__

    tag = f"v{version or __version__}"
    try:
        return Path(snapshot_download(repo, repo_type="dataset", revision=tag)), tag
    except RevisionNotFoundError:
        return Path(snapshot_download(repo, repo_type="dataset", revision="main")), "main"

fetch_index

fetch_index(verbose: bool = True) -> int

Download the published snapshot from settings.rag.index_repo and import it.

Returns:

Type Description
int

Number of entries imported.

Raises:

Type Description
ValueError

HELIOAI_INDEX_REPO is empty, or the snapshot is unusable.

Exception

Whatever the Hub raises — network, unknown repository.

Source code in helioai/index_snapshot.py
def fetch_index(verbose: bool = True) -> int:
    """Download the published snapshot from `settings.rag.index_repo` and import it.

    Returns:
        Number of entries imported.

    Raises:
        ValueError: `HELIOAI_INDEX_REPO` is empty, or the snapshot is unusable.
        Exception: Whatever the Hub raises — network, unknown repository.
    """
    from helioai.config import settings

    repo = settings.rag.index_repo
    if not repo:
        raise ValueError("HELIOAI_INDEX_REPO is empty: fetching a prebuilt index is disabled")
    if verbose:
        print(f"[indexer] fetching the prebuilt index from huggingface.co/datasets/{repo}…")
    local, revision = download_index(repo)
    if not (local / MANIFEST).exists():
        raise ValueError(f"no snapshot published on {repo}@{revision}")
    total = import_index(local, verbose=verbose)
    manifest = json.loads((local / MANIFEST).read_text(encoding="utf-8"))
    if verbose:
        print(
            f"[indexer] {total} entries from snapshot {revision} "
            f"(HelioAI {manifest.get('helioai')}, speasy {manifest.get('speasy')}, "
            f"built {manifest.get('created')})"
        )
        print("[indexer] run `helioai index` again to add what speasy published since")
    return total

publish_index

publish_index(snapshot_dir: Path, repo: str, tag: str | None = None) -> str

Upload a snapshot to main of the Hub dataset repo, and tag it with a release.

The upload is refused when the new snapshot holds fewer than MIN_KEPT of the products expected. The build walks five archives over the network, and one of them answering 502 for the length of the walk yields an index without that archive and no error; published, it would replace a complete index for every new install. Expected is what the dataset already holds, or — for the first publication, which happens unattended at a release tag — the products the shipped classification answered for (judged_products.jsonl.gz, one line per product the paid pass saw). A tag already on the dataset is looked up and moved to the new commit, so re-running a release publishes that release's index; nothing depends on which error the Hub raises for an absent tag. The dataset card (README.md, from .github/hf-index-card.md) goes with the snapshot when its directory holds one, so the card is reviewed in the repository rather than edited on the Hub.

Parameters:

Name Type Description Default
snapshot_dir Path

What export_index wrote.

required
repo str

owner/name of the dataset; the token comes from HF_TOKEN.

required
tag str | None

The release, v0.4.0 — what download_index asks for first.

None

Returns:

Type Description
str

The commit id on the Hub.

Source code in helioai/index_snapshot.py
def publish_index(snapshot_dir: Path, repo: str, tag: str | None = None) -> str:
    """Upload a snapshot to `main` of the Hub dataset `repo`, and tag it with a release.

    The upload is refused when the new snapshot holds fewer than `MIN_KEPT` of the
    products expected. The build walks five archives over the network, and one of them
    answering 502 for the length of the walk yields an index without that archive and no
    error; published, it would replace a complete index for every new install. Expected
    is what the dataset already holds, or — for the first publication, which happens
    unattended at a release tag — the products the shipped classification answered for
    (`judged_products.jsonl.gz`, one line per product the paid pass saw). A tag already
    on the dataset is looked up and moved to the new commit, so re-running a release
    publishes that release's index; nothing depends on which error the Hub raises for
    an absent tag. The dataset card (`README.md`, from `.github/hf-index-card.md`) goes
    with the snapshot when its directory holds one, so the card is reviewed in the
    repository rather than edited on the Hub.

    Args:
        snapshot_dir: What `export_index` wrote.
        repo: `owner/name` of the dataset; the token comes from `HF_TOKEN`.
        tag: The release, `v0.4.0` — what `download_index` asks for first.

    Returns:
        The commit id on the Hub.
    """
    from huggingface_hub import HfApi

    snapshot_dir = Path(snapshot_dir)
    manifest = json.loads((snapshot_dir / MANIFEST).read_text(encoding="utf-8"))
    count = manifest["collections"]["products"]["count"]
    api = HfApi()
    if api.file_exists(repo, MANIFEST, repo_type="dataset"):
        published = json.loads(
            Path(api.hf_hub_download(repo, MANIFEST, repo_type="dataset")).read_text("utf-8")
        )
        expected = published["collections"]["products"]["count"]
    else:
        expected = _shipped_product_count()
    if count < MIN_KEPT * expected:
        raise ValueError(
            f"{count} products against {expected} expected: an archive is missing, not publishing"
        )
    commit = api.upload_folder(
        repo_id=repo,
        repo_type="dataset",
        folder_path=snapshot_dir,
        allow_patterns=[MANIFEST, "README.md", "*.jsonl.gz", "*.npy"],
        commit_message=f"HelioAI {manifest['helioai']}: {count} products",
    )
    if tag:
        if tag in {t.name for t in api.list_repo_refs(repo, repo_type="dataset").tags}:
            api.delete_tag(repo, tag=tag, repo_type="dataset")
        api.create_tag(repo, tag=tag, revision=commit.oid, repo_type="dataset")
    return commit.oid